Skip to content

Courses

Learning streak: 6 days in a row

Reading · 10 min · Lesson 2 of 7

Chunking and embeddings

How documents are prepared so the system can find them.

Before anything can be retrieved, documents are split into chunks, typically a few paragraphs each, and every chunk is turned into an embedding: a list of numbers that captures its meaning. Questions are embedded the same way, and the system finds the chunks whose meaning is closest to the question, even when they use different words.

Chunk size is a real trade-off. Chunks that are too small lose context, so a rule gets separated from its exception. Chunks that are too large dilute the match and waste context space. Splitting along the document's own structure, by headings and sections, usually beats cutting at a fixed number of characters.

Pure meaning-based search has a blind spot: exact terms such as product codes, article numbers or names. Most production systems therefore combine embedding search with classic keyword search, known as hybrid search, and add a reranking step that reorders the best candidates before they reach the model.

Key takeaways

  1. Documents are split into chunks and embedded as vectors of meaning.
  2. Split along the document's structure, not at fixed lengths.
  3. Hybrid search and reranking handle the exact terms that embeddings miss.