Embeddings: how AI turns meaning into numbers

An online shop sells a “rain shell”. A customer searches for “waterproof jacket” and finds nothing, because the words do not match. An embedding fixes exactly this kind of miss: it turns a piece of text into a list of numbers, a vector, and texts that mean similar things receive similar vectors. The search then compares meaning instead of spelling.

From text to vector

An embedding is an array of floating-point numbers produced by a dedicated embedding model. Every vector from the same model has the same length, and no single number in it means anything to a person. What carries meaning is position: the model is trained so that related content ends up close together, which lets software measure how related two texts are.

That measurement is usually cosine similarity, which expresses how closely two vectors point in the same direction. A query vector is compared with the stored vectors, and the closest ones win.

Four jobs embeddings do well

JobExample
Semantic searcha product catalogue in which “waterproof jacket” also finds the rain shell
Retrieval-augmented generation (RAG)an assistant that answers staff questions from the HR handbook by first retrieving the relevant passages
Clusteringgrouping thousands of customer reviews by theme without reading them one by one
Recommendationssuggesting products whose descriptions sit close to the one a visitor is viewing

Three decisions before you start

  • Chunk size. A whole FAQ page embedded as one vector matches every question a little and none of them well. Splitting it into one chunk per question and answer usually retrieves more precisely. There is no universal size, so test it against real questions.
  • Storage. A shop with a few hundred product descriptions can compare a query against every vector directly. A catalogue with millions of entries needs a vector index to answer quickly, whether in a dedicated vector database or in an extension to the database the shop already uses.
  • Model choice and where it runs. Each model builds its own vector space, so switching models later means embedding everything again. Where the model runs is the other half of the decision: embedding an HR handbook full of personnel data through a hosted service sends that data outside the organisation, which is a reason to consider an open embedding model on your own hardware for such collections. Settling this early is part of an AI implementation.

Two limits worth knowing

Closeness in vector space says nothing about accuracy. If an outdated price list is the nearest match to a question, a RAG system retrieves it just as readily as a current one, so keeping the collection clean matters as much as choosing the model.

Languages are the second limit. Whether a German question finds an English document depends on whether the embedding model was trained on both languages; multilingual models are built for this, others handle it poorly.

Related terms

TermWhat it means
Semantic searchSearch by meaning rather than by matching words.
ChunkingSplitting documents into passages before embedding them.
Vector databaseA store optimised for finding the nearest vectors to a query.
Fine-tuningRetraining the model itself, an alternative when the goal is behaviour rather than finding content.

Sources

This page expands an entry from the minoka AI glossary, which covers many more terms in brief.

Frequently Asked Questions

Frequently Asked Questions

How do embeddings improve search?

Embeddings let a search compare meaning instead of exact words. A query and the stored texts are turned into vectors, and the texts closest to the query are returned, even when they use different wording.

What is a good chunk size for embeddings?

There is no universal size. Passages that each answer one question or cover one topic tend to work well, and the right size should be checked against real questions from your users.

Are embeddings the same as a vector database?

No. An embedding is the vector that represents a piece of content. A vector database stores many such vectors and finds the ones closest to a query.

Do embeddings work across languages?

It depends on the model. Multilingual embedding models place texts with the same meaning in different languages close together; models trained mainly on one language do this poorly.

Preferred source

Prefer minoka.de on Google

If you add minoka.de as a preferred source, Google shows articles from this site more often in Top Stories, AI Overviews and AI Mode. One click, a Google account, revocable at any time in your source settings.

Prefer on GoogleOpens the source settings at Google