edgi

Generative pre-trained transformerTrain first, adapt later

A generative pre-trained transformer (GPT) is a type of artificial intelligence model that learns patterns from massive amounts of unlabeled text and uses them to generate new content. Instead of being built for one specific task, it trains first by repeatedly predicting the next word or token in a sequence. This general pretraining allows a single model to adapt to many different jobs, like answering questions, writing code, or summarizing articles.

By the edgi team We find the most surprising true thing about an idea and build a 60-second lesson around it.

Generative pre-trained transformer lesson Play the 60-second lessonGPT-1 trained on 7,000 unpublished books before anyone gave it a particular job. Then it set new results on nine of twelve tasks.

Train first, pick the job later

Before large-scale pretraining became common, models were often trained for a particular job. A spam filter used marked spam; a translator used matched sentence pairs. Each new job usually needed its own training data and adaptation step.

In 2018, OpenAI trained GPT on a large collection of ordinary books before choosing a downstream job. Its pretraining task was to predict the next token. That is what pre-trained means here: the expensive general training happened before a particular use was chosen.

The adaptation step shrank

Take one job: decide whether a film review is positive or negative. GPT-1 still used fine-tuning. It adjusted on labeled examples for each downstream task.

GPT-1 set new results on nine of twelve datasets after that task-specific adaptation. GPT-2 showed that a larger pretrained model could attempt some tasks from its prompt alone, without a new training update.

GPT-3 improved on many tasks when a prompt included a few examples. Those examples change the text the model reads, not its trained weights.

Why pretraining transfers

Predicting the next token exposes a model to many kinds of writing: explanations, code, dialogue, recipes, and mathematics. To continue each kind of text, the model can learn patterns that are useful later. It does not have to receive a separate task label for every pattern.

That is why pretraining can make one model useful on several tasks. It does not guarantee equally good performance on all of them. The later prompt or fine-tuning step points those learned patterns at a particular job.

How does a GPT work?

GPT models rely on the transformer deep learning architecture introduced by Google researchers in 2017. Older natural language designs processed text sequentially, but the transformer uses an attention mechanism that lets the model process entire sequences of text at once. GPT architectures use specifically the decoder part of the transformer.

Diagram of the full GPT architecture, showing the overall model on the left and a detailed view of a Transformer Block on the right. Key components include Input Embedding, Positional Encoding, Transformer Blocks, LayerNorm, Linear layers, Softmax, Dropout, Gelu, Matmul, Mask, and Head units.
This diagram shows the original GPT architecture, which uses a transformer decoder to process text sequences. Original: Marxav, Vectorization: Mrmw, CC0, via Wikimedia Commons

The primary pretraining step uses self-supervised learning on large datasets without human labels. OpenAI trained the original GPT-1 on BookCorpus, a collection of books, while GPT-2 scaled up to WebText, an 8-million-page dataset. By learning solely to predict the next token across diverse writing, the model absorbs structural rules, facts, and reasoning patterns without requiring manual labels for every single skill.

How did GPT models evolve?

Early models required direct modifications to their internal weights for each new job. GPT-1 used fine-tuning on labeled datasets after pretraining, setting new records on nine out of twelve language benchmarks. As models grew, their ability to perform tasks without extra training increased.

Flowchart illustrating a three-stage large language model training workflow, including Stage 1: Pretraining, Stage 2: Supervised Fine-Tuning (SFT), and Stage 3: Reinforcement Learning from Human Feedback (RLHF). The legend categorizes elements as Data, Process, Reward, Model, and Stage.
The workflow illustrates how InstructGPT and ChatGPT use reinforcement learning from human feedback to refine a pretrained model. Kjerish, CC0, via Wikimedia Commons

GPT-2 scaled up parameter count by a factor of 10 to 1.5 billion parameters, demonstrating that a larger model could handle some tasks directly from a prompt. GPT-3 expanded to 175 billion parameters in 2020, introducing few-shot learning where providing a few examples in the prompt allowed the model to execute tasks it was never explicitly trained on.

Later releases added reinforcement learning from human feedback (RLHF) to create InstructGPT and ChatGPT. Modern models have expanded beyond text: multimodal systems like GPT-4o process and generate images and audio, while reasoning models such as DeepSeek R1 allocate extra computation time to produce chain-of-thought reasoning before answering.

Test yourself

How does ChatGPT construct its responses?

By predicting one token at a time. ChatGPT calculates probabilities to pick the most likely next piece of text, adds it to the prompt, and repeats the cycle.

What is the primary role of the attention mechanism in a transformer?

Weigh connections between all words in context. Attention evaluates the relationships among all words in a prompt at once so the model understands specific contextual meanings.

Why does training on ordinary text help a model solve specific problems later?

It builds transferable patterns. General text exposure teaches underlying language structures that happen to be useful across many different tasks.

Collectible card

Claim the Generative pre-trained transformer card

Play the lesson in edgi and the card is yours. It lands on your Map next to the ideas it connects to, and turns from matte to foil to gold as you learn more around it.

Questions people ask

What makes a model multimodal?

A multimodal model can process or generate multiple formats of data instead of text alone. For example, systems like GPT-4o can accept and create text, audio, and visual images.

What is the difference between pre-training and fine-tuning?

Pre-training is the initial phase where a model learns general patterns by predicting tokens across large, unlabeled datasets. Fine-tuning is the subsequent step where the pretrained model is adjusted on a smaller, labeled dataset for a specific downstream task.

What is a reasoning GPT?

A reasoning GPT is a model trained with reinforcement learning to spend extra computation time working through a problem step by step before returning an answer. Models like o3 and DeepSeek R1 use this chain-of-thought approach to handle complex subjects like mathematics.

Part of the Set · 8 cards

How ChatGPT Actually Works

ChatGPT predicts the next token. The surprising part is what training and feedback can build on top of that.

  1. ChatGPT
  2. Hallucination (artificial intelligence)
  3. Embedding (machine learning)
  4. Self-Supervised Learning
  5. Attention (machine learning)
  6. Transformer (machine learning)
  7. Generative pre-trained transformerReading now
  8. Neural Scaling Law
Learn the whole Set

Where this leads