edgi

Transformer: how ditching word order sped up AI

A transformer is a neural network architecture that analyzes all parts of an input sequence at the same time using attention, which measures how much different items relate to each other. Instead of reading text word by word, it processes entire passages at once and learns relationships across the whole sequence. This shift allowed models to train across many computer chips simultaneously at a fraction of previous costs.

By the edgi team We find the most surprising true thing about an idea and build a 60-second lesson around it.

Transformer (machine learning) lesson Play the 60-second lessonEvery word looks at every other word, all in the same step. That made reading fast, and it makes long texts expensive.

Reading through a straw

Many earlier neural language models processed a sentence in order. One word updated the model before it could handle the next. The model had to carry what it had read forward in a running state. It was like reading a novel while keeping your notes on one index card.

During training, word 50 had to wait for word 49. That dependency made it hard to spread one sentence across many chips. Chips can do many calculations at once. A word-by-word loop limits how much of one sentence can be trained in parallel.

Diagram comparing data flow in Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), and self-attention mechanisms. Each model shows layers of nodes (circles) and connections (solid and dashed arrows) processing input data represented by x1 through x5.
Diagram comparing data flow in Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), and self-attention mechanisms. Zhang, Aston and Lipton, Zachary C. and Li, Mu and Smola, Alexander J., CC BY-SA 4.0, via Wikimedia Commons

Train on the whole page

In 2017, eight Google researchers published Attention Is All You Need. Their transformer removed recurrent reading from the model. During training, a transformer can process all positions in a passage together and compare how they relate. That comparison is attention.

Because training has no word-by-word chain, it can run many position calculations in parallel across chips. When it generates a reply, an autoregressive model still produces tokens one after another. The paper reported new best translation results after three and a half days on eight GPUs, at a small fraction of the training cost reported for earlier leading systems.

The price, and what it bought

If a model receives word representations with no position information, dog bites man and man bites dog contain the same items. A transformer adds position information to each word representation, then learns how to use it. Order becomes another input feature.

The same basic architecture can work with other sequences of representations. A vision transformer uses image patches; protein models use amino-acid sequences. The input has to be prepared for its domain, but the transformer supplies a shared way to learn relationships among its parts.

How transformers process sequences

A transformer breaks inputs like text, audio, or images into numerical units called tokens, converting each token into a mathematical vector. Because the attention mechanism looks at every token simultaneously, it naturally treats an input as an unordered bag of items. To fix this, the model injects positional encodings or learned embeddings into each token vector so the network can tell the difference between different word orders.

Diagram of a standard transformer architecture in deep learning, showing an encoder on the left and a decoder on the right. Both sections include "Nx 'Layers'" composed of "Multi-Headed Self-Attention" or "Masked Multi-Headed Self-Attention" and "Feed-Forward Network" blocks, with "Norm" layers and "Positional Encoding" inputs.
The transformer architecture combines stacked encoder and decoder blocks with attention mechanisms and positional inputs. dvgodoy, CC BY 4.0, via Wikimedia Commons

Inside each layer, multi-head attention compares tokens within a set context window. The mechanism amplifies signals for key related tokens while diminishing unimportant ones. Unlike recurrent neural networks that suffered from the vanishing-gradient problem over long texts, attention lets any token connect directly with another across the sequence.

Transformer variants and uses

Modern architectures are grouped into encoder-only, decoder-only, and encoder-decoder designs depending on their target tasks. Encoder-only models like BERT excel at representation learning. Decoder-only models like generative pre-trained transformers (GPTs) focus on autoregressive generation, producing output tokens one after another.

The architecture extends far beyond machine translation. Vision transformers slice images into patches to perform computer vision tasks. Other implementations run reinforcement learning, process audio, model protein amino-acid sequences, direct robotics, and play chess.

Test yourself

How does a transformer machine learning model process input sequences during training?

It processes all positions together in parallel. Transformers remove recurrent reading loops, allowing the model to analyze an entire passage at once and vastly speed up training across multiple chips.

Why must a transformer model explicitly add position information to its input representations?

Attention alone treats input as an unordered set. Because self-attention compares all parts simultaneously without a strict processing order, word order must be encoded as an extra feature so meaning does not depend solely on the bag of words.

Why can one basic architecture work with text, images, and proteins?

It learns relationships among input representations. Text tokens, image patches, and amino-acid sequences can all be represented as inputs with positions. The transformer can then learn how the parts relate.

Collectible card

Claim the Transformer (machine learning) card

Play the lesson in edgi and the card is yours. It lands on your Map next to the ideas it connects to, and turns from matte to foil to gold as you learn more around it.

Questions people ask

What neural networks did transformers replace?

They largely replaced recurrent neural networks (RNNs) and long short-term memory (LSTM) models. Those older systems had to process sequences sequentially, one token at a time, which limited parallel training across hardware.

What is the computational tradeoff of self-attention?

Standard self-attention computes relationships between all sequence positions, which causes computational and memory costs to scale quadratically with the length of the input sequence.

Part of the Set · 8 cards

How ChatGPT Actually Works

ChatGPT predicts the next token. The surprising part is what training and feedback can build on top of that.

  1. ChatGPT
  2. Hallucination (artificial intelligence)
  3. Embedding (machine learning)
  4. Self-Supervised Learning
  5. Attention (machine learning)
  6. Transformer (machine learning)Reading now
  7. Generative pre-trained transformer
  8. Neural Scaling Law
Learn the whole Set

Where this leads