Machine Learning Embeddings: how math maps meaning
An embedding is a list of numbers that represents complex data like words or images so a computer can work with them mathematically. The system arranges these coordinates so that similar items land close to each other in the number space. This allows algorithms to compare meanings directly instead of relying on exact word matches.
By the edgi team We find the most surprising true thing about an idea and build a 60-second lesson around it.
A screen represents color with three values: red, green and blue. Pure red is 255, 0, 0. Orange is 255, 165, 0. Blue is 0, 0, 255. Orange is closer to red than blue in that number system, just as it looks closer to red.
Close-up image of RGB sub-pixels on an LCD TV screen, showing a grid of red, green, and blue vertical stripes. Stan Zurek, CC BY-SA 3.0, via Wikimedia Commons
An embedding uses the same trick. It represents something as a list of numbers, arranged so similar things land near each other. A text embedding can have thousands of numbers. The numbers give a model a form it can compare and calculate with.
How words find their neighbors
Nobody chooses thousands of numbers for the word Tuesday. A model learns them from the words that appear around it. Tuesday and Wednesday occur in many similar sentences, so a trained model can place them near each other.
Google's 2013 word2vec work made this relationship famous with an analogy: king minus man plus woman points near queen. The lookup excludes the words already used in the calculation. Otherwise king itself is a tempting nearest answer, which would hide the relationship the example is meant to show.
A 2D vector space diagram illustrating word embeddings, with "Man" and "King" as blue points and "Woman" and "Queen" as red points. Singerep, CC BY-SA 4.0, via Wikimedia Commons
What a map of meaning buys
Once each item has coordinates, a computer can calculate which of millions of items is closest to a new one. Search can use that distance as one signal. A search for stopping a dog barking can surface a page about quieting a noisy pet even when the wording differs.
Captioned-image models can learn a shared space too. The word cat and photos of cats can land near each other. Before a chatbot processes your message, it turns its tokens into vectors, numbers the model can use.
How computers turn data into coordinates
Instead of relying on humans to manually label relationships, machine learning models learn embeddings directly from raw data like text, images, or user interactions. A model maps complex items into a lower-dimensional space of numerical vectors. Words like Tuesday and Wednesday end up near each other because they repeatedly appear in similar sentences.
A famous example comes from Google's 2013 word2vec system, which showed that vector math captures abstract relationships: taking the vector for king, subtracting man, and adding woman points to a location near queen. The system automatically extracts these patterns without requiring prior domain knowledge.
How models measure distance and similarity
Once items exist as numerical vectors, a model computes similarity by measuring the mathematical distance between points. For normalized vectors, models often use cosine similarity, which checks the angle between two vectors rather than their magnitude.
Cosine similarity prevents very common training data from dominating the results. The standard dot product includes magnitude inherently, which biases results toward more frequent items. In high-dimensional spaces, simple Euclidean distance becomes less reliable as vectors converge in distance, making angle-based comparisons a standard tool.
Where embeddings appear in everyday tools
Embeddings power search engines by letting them retrieve pages that match meaning rather than exact keywords. A query about stopping a barking dog can successfully return an article about quieting a noisy pet.
They also bridge different media. Captioned-image models place visual data and written concepts in a shared coordinate space, putting the word cat near actual photos of cats. Chatbots also convert incoming text tokens into vectors before processing any message.
Test yourself
When a machine learning embedding places two distinct words near each other in vector space, what does it primarily reflect?
Similarity in their surrounding usage. Embeddings group items based on contextual proximity in training data, meaning words with similar usage patterns land close together regardless of exact definitions.
Does a machine learning embedding rely on manually assigned definitions or observed context?
Observed context. Models learn embeddings by observing which words appear in similar surrounding contexts, rather than relying on human-written dictionaries.
What makes an embedding useful for semantic search?
Similar meanings land close together. An embedding gives each item coordinates arranged by learned similarity. Search can use those coordinates as one way to find related wording.
Play the lesson in edgi and the card is yours. It lands on your Map next to the ideas it connects to, and turns from matte to foil to gold as you learn more around it.
Models can create embeddings for words, images, user interactions, and knowledge graphs. Each type maps raw concepts into vector spaces designed for tasks like computer vision or recommendation systems.
How does an embedding differ from one-hot encoding?
One-hot encoding is a manually designed method that represents items with rigid, sparse categories. Embeddings are learned automatically by models, capturing latent relationships and reducing data complexity into a lower-dimensional space.