When a RAG system gives a wrong answer, the language model usually gets the blame. In our experience, most failures happen earlier: the right passage was never retrieved because it was split in the wrong place, embedded with a model that handles Ukrainian poorly, or outranked by a passage that merely sounded similar. This guide focuses on that retrieval layer — embeddings and chunking — and the practical decisions that move answer quality the most.

Embeddings in one paragraph

An embedding model maps text to a vector so that texts with similar meaning end up close together. Modern embedding models are transformer encoders trained with contrastive objectives on huge sets of related text pairs — questions and answers, titles and bodies, translations (Reimers & Gurevych, Sentence-BERT). At query time, you embed the user's question and search for the nearest chunk vectors in a vector database.

Two consequences matter in practice:

  1. The model's notion of "similar" comes from its training data. Domain-specific language (legal, medical, accounting) and less-represented languages can be handled worse than general English.
  2. Questions and answers are not phrased alike. Good retrieval models are trained for asymmetric search (short query → long passage), and many expect a prefix such as query: and passage:; read the model card.

Choosing an embedding model

The MTEB benchmark compares embedding models across many tasks and languages. The original paper's key finding still holds: no single model dominates all tasks (Muennighoff et al., 2022). Use MTEB to shortlist, then decide on your own data.

Criteria:

Criterion Why it matters
Retrieval scores in your languages Ukrainian and mixed Ukrainian/English/Russian text is common in local business data
Max input length Determines maximum chunk size (512 vs 8,192 tokens)
Dimensions and Matryoshka support Memory and speed; truncatable vectors save storage (Kusupati et al., 2022)
Hosting API (OpenAI, Cohere, Voyage, Google) vs open weights you can self-host
Cost and throughput Re-embedding a large corpus happens more often than you expect
Licence Open-weight model licences vary

Multilingual open models worth testing for Ukrainian include BGE-M3, which supports dense, sparse and multi-vector retrieval in one model (Chen et al., 2024), and the multilingual E5 family (Wang et al., 2024). Commercial APIs from major providers also handle Ukrainian well. Self-hosting matters when documents cannot leave your infrastructure; see self-hosting LLMs and privacy for LLM apps.

Evaluate on your own data — it takes a day

Build a small retrieval eval set:

  1. Collect 100–200 real questions (from support tickets, search logs, employees).
  2. For each, mark the chunk(s) that contain the answer.
  3. Embed your corpus with 2–4 candidate models.
  4. Measure recall@k (is the right chunk in the top 5/10/20?) and MRR (how high is it ranked?).
def recall_at_k(results: dict[str, list[str]], gold: dict[str, set[str]], k: int) -> float:
    hits = sum(1 for q, ids in results.items() if gold[q] & set(ids[:k]))
    return hits / len(results)

for model in ["bge-m3", "multilingual-e5-large", "api-model-x"]:
    results = run_retrieval(model, questions, top_k=20)
    print(model, {k: round(recall_at_k(results, gold, k), 3) for k in (5, 10, 20)})

Differences of 10–20 points of recall@10 between models on real Ukrainian business data are common — far more than leaderboard averages suggest. This set also becomes part of your end-to-end LLM evals.

Chunking: where most quality is won or lost

Chunking decides what a "unit of retrieval" is. Too large, and chunks mix topics, dilute the embedding and waste context. Too small, and chunks lose the context needed to understand them ("The fee is 2%" — which fee?).

Strategies

  • Fixed-size with overlap. Split every N tokens with 10–20% overlap. Simple baseline; ignores structure.
  • Recursive / structure-aware. Split by headings, then paragraphs, then sentences, until chunks fit the size limit. The best default for most business documents.
  • Document-type specific. Contracts by clause, FAQs by question, code by function, tables row-group by row-group with headers repeated, emails by message.
  • Semantic chunking. Split where embedding similarity between adjacent sentences drops. Sometimes helps for unstructured prose; more compute, inconsistent gains.

Typical starting sizes are 200–800 tokens. Measure; do not assume.

Preserve context in every chunk

A chunk should make sense when read alone. Cheap techniques:

  • Prepend the document title and section path: Refund policy > Annual plans > Partial refunds.
  • Store metadata (document type, date, product, language, access level) for filtering.
  • Keep tables intact or convert them to "column: value" sentences per row.

Contextual retrieval

Anthropic's contextual retrieval technique uses an LLM to write a short, chunk-specific context ("This chunk is from ACME's Q2 2026 report and discusses revenue in the EU segment...") that is prepended before embedding and BM25 indexing. In their experiments, contextual embeddings plus contextual BM25 reduced failed retrievals by 49%, and by 67% when combined with reranking. With prompt caching, the cost of generating these contexts for a whole corpus is modest.

Late chunking and multi-vector approaches

Alternatives to per-chunk context generation: late chunking embeds the full document with a long-context model and pools token embeddings per chunk, so each chunk vector "knows" its surroundings (Günther et al., 2024). Multi-vector models such as ColBERT keep one vector per token and match at a fine-grained level, at higher storage cost (Khattab & Zaharia, 2020).

Hybrid search and reranking

Embeddings capture meaning; they are weaker at exact identifiers — article numbers, SKUs, error codes, names. Combine them with keyword search (BM25) and merge the rankings. Then apply a cross-encoder reranker to the top 30–100 candidates: rerankers read the query and passage together and score relevance far more accurately than vector similarity, at higher cost per pair. Hybrid + rerank is the configuration we deploy by default; implementation details are in RAG for business and, for PostgreSQL, in vector databases compared.

Ukrainian-specific considerations

  • Mixed languages. Business data in Ukraine often mixes Ukrainian, English and Russian, sometimes in one document. Use multilingual models and test cross-lingual queries (question in Ukrainian, answer in an English document).
  • Morphology. Ukrainian inflection hurts naive keyword search: "договору", "договором", "договорі". Use a proper analyser or rely more on the dense component; test BM25 with and without stemming.
  • Apostrophes and transliteration. Normalise apostrophe variants (', ’, ʼ) and consider transliterated forms for names and brands.
  • Scanned documents. OCR quality on Cyrillic varies; poor OCR destroys retrieval. See LLM document processing.

Operational practices

  • Version embeddings. Store the model name and version with every vector. Mixing vectors from different models in one index silently breaks search.
  • Keep raw text. You will re-embed when you switch models; plan it as a background migration with a dual-read period.
  • Incremental updates. Re-embed only changed documents; detect changes by content hash.
  • Respect access control at retrieval time. Filter by the user's permissions in the query, never only in the prompt; see AI agent security.
  • Monitor retrieval in production. Log retrieved chunk IDs and scores per answer so you can debug bad answers back to retrieval; see LLM observability.

A starting configuration

For a typical Ukrainian business knowledge base:

  1. Structure-aware chunking, 300–600 tokens, title and section path prepended.
  2. A multilingual embedding model chosen by recall@10 on 150 real questions.
  3. Hybrid search: dense + BM25, fused with RRF, top 50 candidates.
  4. Cross-encoder reranking down to the top 5–8 chunks.
  5. Contextual retrieval if recall@10 is still below about 85–90%.

Then iterate with your eval set, changing one thing at a time.

FAQ

Bigger embedding dimensions — better results? Not necessarily. Model quality matters more than dimension count; Matryoshka models often keep most quality at half the dimensions.

Should we fine-tune an embedding model? Consider it after you have tried hybrid search, reranking and better chunking, and only if you have thousands of labelled query-passage pairs. See fine-tuning vs RAG vs prompting.

Can we just use long-context models and skip retrieval? For small corpora, sometimes. For anything larger or frequently changing, retrieval remains cheaper, faster and easier to control; see context engineering.

How often should we re-embed? When content changes (incrementally), and when you switch models (fully). Otherwise, embeddings do not expire.

Sources

  1. Muennighoff et al. (2022). MTEB: Massive Text Embedding Benchmark; MTEB leaderboard.
  2. Reimers, N., Gurevych, I. (2019). Sentence-BERT.
  3. Chen et al. (2024). BGE M3-Embedding; Wang et al. (2024). Multilingual E5 Text Embeddings.
  4. Kusupati et al. (2022). Matryoshka Representation Learning.
  5. Anthropic (2024). Introducing Contextual Retrieval.
  6. Günther et al. (2024). Late Chunking; Khattab, O., Zaharia, M. (2020). ColBERT.