Attention is a method in machine learning that calculates how important every component in a sequence is relative to the others. Instead of processing data one word at a time and forgetting early details, it gives each token direct, immediate access to the entire sequence. This dynamic weighing lets modern language models grasp context across long spans of text simultaneously.
By the edgi team We find the most surprising true thing about an idea and build a 60-second lesson around it.
"The animal didn't cross the street because it was too tired." You know "it" means the animal. Change "tired" to "wide" and "it" means the street. You settled the meaning by weighing "it" against every other word in the sentence and choosing the word that fits. Attention is a machine doing the same weighing.
For the word "it", attention scores every other word in the text for how much that word matters to "it". The scores add up to 1, so "it" spends most of its budget on "animal". Attention runs the scoring again for the next word, and "street" gets its own set. The same word can be worth almost everything to one word and almost nothing to another.
How it learns where to look
Nobody wrote a rule saying that "tired" points "it" at the animal. The scores are learned through a prediction task. Take a sentence from the internet, hide the last word, and make the model guess it. Compare the guess with the real word, nudge the model, and repeat. Billions of times.
To guess "tired" you need "it" already attached to something that can get tired. Scoring "animal" high is what makes that guess land, so the game rewards scoring it high. The training signal says whether the model predicted the right word, not which earlier word it should use. The word-to-word pattern has to emerge from the task.
You can see where it looked
In 2014, an attention mechanism helped neural machine translation align words across languages. The paper printed a grid: English words along one edge, French along the other, and bright squares where a French word drew heavily on an English one.
Most of the bright squares run down the diagonal. But French writes "the European Economic Area" backwards, zone économique européenne, and the grid shows the model reaching ahead to "Area" for "zone".
What the scoring costs
In full self-attention, a transformer compares each position with every position it is allowed to see. A causal chatbot masks future words. A 10-word full attention layer makes roughly 100 comparisons; a 1,000-word layer makes roughly a million. Doubling length makes about four times as many comparisons.
Why recurrent neural networks fell short
Older systems relied on recurrent neural networks that processed language one word after another. This sequential pipeline meant that words near the end of a sentence overshadowed earlier ones, weakening the predictive value of the opening text.
This chart compares how data flows in convolutional networks, recurrent networks, and self-attention models. Zhang, Aston and Lipton, Zachary C. and Li, Mu and Smola, Alexander J., CC BY-SA 4.0, via Wikimedia Commons
Attention solved this problem by giving every token equal access to any part of the sequence directly rather than passing through an accumulated hidden state. In 2017, self-attention became the foundation of the Transformer architecture, replacing recurrent models entirely and powering systems like BERT, T5, and GPT.
Soft weights and the alignment matrix
Attention relies on soft weights, which are dynamic coefficients generated on the forward pass of a model rather than permanent numbers fixed during training. Because these soft values change with every new input step, the model can distribute its focus across multiple places instead of picking a single rigid match.
An encoder-decoder architecture uses an attention block to calculate dynamic weights between input sequences and generated outputs. Numiri, CC BY-SA 4.0, via Wikimedia Commons
This flexibility is clear in translation tasks where word orders change. Translating the English phrase "I love you" into French requires the decoder to assign 94% of its attention to "I" to produce "je", 88% to "you" for "t'", and 95% to "love" for "aime". Stacking these soft row vectors produces an alignment matrix that maps relationships even when words change positions.
The quadratic cost of comparing everything
Full self-attention calculates a matrix whose size grows with the square of the input length. A layer handling 1,000 tokens makes roughly one million comparisons, consuming massive amounts of GPU memory on long documents.
To handle this scaling bottleneck, optimizations like Flash attention partition matrix computations into small blocks that fit inside faster on-chip GPU memory. This avoids storing large intermediate matrices, speeding up processing while keeping the mathematical results identical.
Test yourself
How are attention scores in machine learning models determined during training?
Learned through a word prediction task. Attention weights are not programmed manually. They emerge automatically as the model is trained billions of times to predict missing words in text.
What happens to the computational cost of full self-attention when input text length doubles?
It increases by a factor of roughly four. Because full self-attention compares every position with every other position, doubling the length quadruples the total number of pairwise comparisons.
Why can full attention cost jump for longer text?
Every word is compared with all others. Ten times the words means about a hundred times the comparisons in full attention. The cost comes from scoring pairs of positions, not from rereading or retraining.
Play the lesson in edgi and the card is yours. It lands on your Map next to the ideas it connects to, and turns from matte to foil to gold as you learn more around it.
What is the difference between soft and hard attention weights?
Hard weights are permanent numbers updated only during the backward training pass, or rigid switches that pick just one token. Soft weights are dynamic scores calculated on the forward pass that re-weight multiple tokens at the same time.
How is attention used outside natural language processing?
Computer vision models use visual attention to focus on specific regions of an image during object detection and captioning. Researchers inspect these decisions by turning vision transformer attention scores into visual heat maps.
What is Flash attention?
Flash attention is an optimized computing method that divides large attention matrices into smaller blocks. These blocks fit directly onto fast GPU memory, which drastically cuts memory consumption without losing calculation accuracy.