A transformer is a neural network architecture that analyzes all parts of an input sequence at the same time using attention, which measures how much different items relate to each other. Instead of reading text word by word, it processes entire passages at once and learns relationships across the whole sequence. This shift allowed models to train across many computer chips simultaneously at a fraction of previous costs.
By the edgi team We find the most surprising true thing about an idea and build a 60-second lesson around it.
Many earlier neural language models processed a sentence in order. One word updated the model before it could handle the next. The model had to carry what it had read forward in a running state. It was like reading a novel while keeping your notes on one index card.
During training, word 50 had to wait for word 49. That dependency made it hard to spread one sentence across many chips. Chips can do many calculations at once. A word-by-word loop limits how much of one sentence can be trained in parallel.
Diagram comparing data flow in Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), and self-attention mechanisms. Zhang, Aston and Lipton, Zachary C. and Li, Mu and Smola, Alexander J., CC BY-SA 4.0, via Wikimedia Commons
Train on the whole page
In 2017, eight Google researchers published Attention Is All You Need. Their transformer removed recurrent reading from the model. During training, a transformer can process all positions in a passage together and compare how they relate. That comparison is attention.
Because training has no word-by-word chain, it can run many position calculations in parallel across chips. When it generates a reply, an autoregressive model still produces tokens one after another. The paper reported new best translation results after three and a half days on eight GPUs, at a small fraction of the training cost reported for earlier leading systems.
The price, and what it bought
If a model receives word representations with no position information, dog bites man and man bites dog contain the same items. A transformer adds position information to each word representation, then learns how to use it. Order becomes another input feature.
The same basic architecture can work with other sequences of representations. A vision transformer uses image patches; protein models use amino-acid sequences. The input has to be prepared for its domain, but the transformer supplies a shared way to learn relationships among its parts.
How transformers process sequences
A transformer breaks inputs like text, audio, or images into numerical units called tokens, converting each token into a mathematical vector. Because the attention mechanism looks at every token simultaneously, it naturally treats an input as an unordered bag of items. To fix this, the model injects positional encodings or learned embeddings into each token vector so the network can tell the difference between different word orders.
The transformer architecture combines stacked encoder and decoder blocks with attention mechanisms and positional inputs. dvgodoy, CC BY 4.0, via Wikimedia Commons
Inside each layer, multi-head attention compares tokens within a set context window. The mechanism amplifies signals for key related tokens while diminishing unimportant ones. Unlike recurrent neural networks that suffered from the vanishing-gradient problem over long texts, attention lets any token connect directly with another across the sequence.
Transformer variants and uses
Modern architectures are grouped into encoder-only, decoder-only, and encoder-decoder designs depending on their target tasks. Encoder-only models like BERT excel at representation learning. Decoder-only models like generative pre-trained transformers (GPTs) focus on autoregressive generation, producing output tokens one after another.
The architecture extends far beyond machine translation. Vision transformers slice images into patches to perform computer vision tasks. Other implementations run reinforcement learning, process audio, model protein amino-acid sequences, direct robotics, and play chess.
Test yourself
How does a transformer machine learning model process input sequences during training?
It processes all positions together in parallel. Transformers remove recurrent reading loops, allowing the model to analyze an entire passage at once and vastly speed up training across multiple chips.
Why must a transformer model explicitly add position information to its input representations?
Attention alone treats input as an unordered set. Because self-attention compares all parts simultaneously without a strict processing order, word order must be encoded as an extra feature so meaning does not depend solely on the bag of words.
Why can one basic architecture work with text, images, and proteins?
It learns relationships among input representations. Text tokens, image patches, and amino-acid sequences can all be represented as inputs with positions. The transformer can then learn how the parts relate.
Play the lesson in edgi and the card is yours. It lands on your Map next to the ideas it connects to, and turns from matte to foil to gold as you learn more around it.
They largely replaced recurrent neural networks (RNNs) and long short-term memory (LSTM) models. Those older systems had to process sequences sequentially, one token at a time, which limited parallel training across hardware.
What is the computational tradeoff of self-attention?
Standard self-attention computes relationships between all sequence positions, which causes computational and memory costs to scale quadratically with the length of the input sequence.