Why retrieval quality is becoming the defining challenge in AI agent architecture
Disclosure: Vespa asked me to evaluate where retrieval fits in large-scale agentic AI workflows. This post is my view of the retrieval problem and where Vespa is used as an example implementation.
Agentic systems usually have two jobs: build context, then use that context to produce an answer or action.
Many failures that look like LLM problems start in the context-building step. The answer the LLM gives is limited by the context it was given or it finds through tool calls. If the agent model cannot find the right sources then improving the generation model will not improve the overall system.
Many failures that look like LLM problems start in the context-building step.
A client, Specstory, wanted to give users the ability to ask questions from agent history. For example, why a team chose Authlib for authentication and what alternatives they considered. The chatbot needs the right prior conversations, decisions, and tradeoffs from a large corpus of coding sessions. The model and system prompt help only after those chat turns are retrieved and in context.
The same pattern showed up in an AnkiHub operator review in our private community. A request for help studying based on lecture slides only works if the agent’s tool calls retrieve the right flashcards.
The setup changes by product. A human may see retrieval as a search box. An agent may call local search, semantic search, a database, or a SaaS API. Either way, the process is the same: gather the right context, then generate from it.
An order-status agent might pull customer records from Salesforce, shipment data from SAP, and tracking events from a logistics API. Those sources are structured application state rather than a document corpus. The agent still has to gather the right context before it answers.
A coding agent
runs rg, opens files, reads logs, and inspects tests before writing a patch.
A research agent searches the web and internal notes before writing an answer.
A study assistant searches deck facts and user context before suggesting what to learn next.
When context building fails the symptoms look like generation failures.
| Symptom | Retrieval cause |
|---|---|
| Hallucination | The answer source never made it into context. |
| Context rot | Low recall forces a high top_k, so noisy results fill the context window. |
| Latency | Weak retrieval causes more tool calls, larger candidate sets, and bigger context windows. |
A useful source can disappear at several points in the path. The trace should tell you which one.
| Failed step | What the trace shows | Change to make |
|---|---|---|
rg / grep |
A conceptual query returns literal matches and misses the relevant files. | Add semantic search over files or chunks, or generate better keyword queries before rg. |
| BM25 | The query uses the right concept but different words from the source material. | Add semantic search, synonyms, or query expansion. |
| Semantic search | Exact names, error strings, document IDs, or domain terms are missing from the results. | Add a keyword or BM25 path, or boost exact term matches. |
| Hybrid retrieval | The relevant passage is ranked 7th, but the context only takes the top 5. | Add or tune a reranker, or raise candidate top_k before reranking. |
| SQL / database | The generated query uses the wrong joins, filters, or time window. | Ground query generation in the schema, validate filters, and trace the returned rows. |
| SaaS / API call | The agent calls the wrong endpoint, omits an account ID, or requests too few fields. | Tighten tool schemas, required arguments, and response summaries. |
A better model helps with reasoning and writing, but it cannot give a better answer without the right context.
The Mixedbread OfficeQA-Pro Eval shows the pattern at benchmark scale. OfficeQA-Pro uses 89,000 pages of financial documents, dense tables, scanned PDFs, and questions that require reasoning across documents. Giving Codex better search tools reduced tool calls and improved answer quality.

Plain text tools like grep and rg work (ish) over flat code files.
They do not work well when context lives in PDFs, tables, chat histories, multi-modal inputs, web results, SaaS records, and permissioned data.
In those cases the agent needs retrieval that can combine exact terms, meaning, metadata, permissions, and ranking quality.
Retrieval needs traces and evals
Once retrieval enters the architecture, the next question is whether it finds the right information.
For that you need traces and evals. For each retrieval step the minimum trace is the input, the outputs, and a way to label whether each output was relevant.

For a coding agent using rg the input is the command, the output is the returned snippets,
and the label says which snippets helped, which were noise, and which relevant files were missing.
For product retrieval the step might be BM25, semantic search, hybrid search with reranking, a generated SQL query, or something else. Capture the query or arguments, the returned documents or chunks, and whether those results were helpful.
Trace each retrieval step by itself, then evaluate the full context-building pass. The local trace answers “Did this query return useful material?” The full trace answers “Did the system collect everything the model needed before generation?” If it did, failures are a generation problem. If not, it’s a retrieval problem.
You cannot know where the failure started or what to fix without traces.
Different failures need different fixes
“Improve retrieval” is too broad to be useful because different problems have different solutions. If a relevant document is missing the trace should show where it disappeared: query building, retrieval, filtering, ranking, or final context assembly.

That diagnosis should land on a concrete system change. The right fix depends on what the system was trying to retrieve. A decision-history question needs the decision, the alternatives, and the chats where the team worked through them. A study question depends on the lecture material, deck metadata, semantic matches, and the user’s study context.
The architecture
Once you trace individual retrieval calls, the full architecture has a simple shape: fan out to context-building tools, then fan in to generate the final output.

Where Vespa fits
Vespa is useful once retrieval has to serve many filtered, ranked requests instead of one broad search call. The app still owns the user request, workflow, and final context budget. Vespa stores the retrievable chunks and serves filtered, ranked evidence back to the app.
That boundary matters because agents rarely make one clean search. They ask follow-up questions, widen or narrow filters, and need enough trace data to explain why a source did or did not reach the model.
Consider a research agent answering, “What changed in the refund policy after the February outage?” The answer might live across the current policy doc, a support ticket thread, an incident review, and a changelog. The app knows the user, tenant, allowed sources, time window, and why the agent is asking. Vespa serves the evidence.
The same shape applies when the retrievable units are customer records, shipment events, product rows, or permissioned SaaS objects instead of document chunks.
In Vespa, that starts with storage. For policy docs, notes, tickets, chat histories, and long PDFs, the retrievable unit is usually a chunk with a stable source id, chunk id, timestamp, permission fields, metadata, and embedding. For records and SaaS objects, it might be a row or object with the same retrieval fields. The schema makes those fields available for search, filters, ranking, and summaries.
At query time, the app sends the request text, filters, query embedding, and rank profile. Permission and tenant filters narrow the candidate set before ranking. Vespa can retrieve candidates with text matching, nearest-neighbor search, or both. Ranking can favor the current policy doc, exact issue IDs, and semantically related incident notes.
Document summaries return the chunk text, ids, metadata, scores, and source fields the app needs. The app chooses what enters the model context. For debugging, the trace should preserve the query, filters, hits, selected context, answer id, and reviewer labels. Vespa provides query traces and logs; the labeling workflow is usually app-owned.
Ranking
Ranking is where Vespa decides which allowed candidates deserve the limited context budget.
Better ranking improves precision (more of the returned chunks are relevant). If the right chunk is ranked 40th and the context only includes the top 10, the system behaves as if retrieval missed it. Returning the top 40 is one answer but it costs a lot of tokens and latency. If the relevant chunks are near the top the system can pass fewer chunks into the model, spend fewer tokens, reduce latency, and give the model less noise.
Semantic similarity is one ranking signal. A chunk with a strong vector score may still be less useful than a newer source, an authoritative source type, or a chunk with an exact error code, file name, ticket ID, or quoted phrase. Permissions and tenant boundaries should already have narrowed the candidate set. Ranking decides how to order the allowed candidates.
A ranker usually uses several kinds of information:
| Information | What it means | How it changes ranking |
|---|---|---|
| Exact text match | The chunk contains the same words, names, IDs, error strings, or quoted phrases. | Keeps precise terms from being buried by vaguely related semantic matches. |
| Semantic match | The query and chunk have similar embeddings, even if they use different words. | Finds conceptually related chunks when the wording differs. |
| Metadata | Structured fields on the chunk, such as source type, owner, status, or selected collection. | Boosts or lowers results based on product context. |
| Freshness | The timestamp, version, or update status attached to the chunk. | Gives current sources more weight when old context is less useful. |
| Model score | A slower model compares the query and candidate after cheaper retrieval. | Rescores a smaller candidate set when simple ranking is not enough. |
The Vespa schema defines the document fields and ranking options. A rank profile is a named scoring recipe in that schema. A profile can combine exact text scores, semantic scores, metadata, values sent with the query, and model scores in one formula.
Vespa separates the query that finds candidates from the rank profile that orders them. A nearestNeighbor query uses the vector index to find candidate chunks near the query embedding. A text query uses term matching. A hybrid query can use both. After that, Vespa scores each candidate so it can return the best ones first. For example, a hybrid rank profile might start with vector closeness plus BM25, then use a slower model such as a cross-encoder to rescore the best candidates.
Scale introduces more challenges
With a small corpus a prototype can search and still feel fast. With millions or billions of chunks, every extra retrieval call, candidate, ranking pass, and returned token adds up.
One user request can trigger several retrieval calls in an agentic system. Each of those could be multi-stage retrieval calls (query rewriting, keyword/semantic, ranking, etc). Combine that with many concurrent users and it becomes clear the retrieval layer has to keep quality high while controlling latency and cost.
| Scale pressure | What has to hold |
|---|---|
| More possible matches | Permissions, tenants, dates, and source types should reduce the candidate set before ranking. |
| Expensive reranking | Stronger models need a controlled number of candidates. |
| Chained retrieval calls | Slow tail latency compounds across query rewrites, follow-up searches, and citation lookups. |
| High write volume | Freshness matters in small corpora too, but high write volume makes stale context harder to avoid. |
| Harder cost debugging | Traces need timing, candidate counts, ranking stages, and returned token counts. |
| Larger returned context | Summaries need to stay narrow so every request does not carry unnecessary text. |
When retrieval and ranking run in the same serving path, filters can shrink the candidate set before expensive ranking, summaries can limit returned text, and traces can show whether latency came from matching, ranking, or response size.
Multi-stage retrieval is the common shape
Most production systems split retrieval into stages, even when the UI looks like one chat box. The trace should show whether the app asked the wrong retrieval question, the candidate set never included the source, filters removed it, ranking buried it, the summary left out needed fields, or context assembly dropped it before generation.
If the query builder chose the wrong filters then changing the embedding model will not help. If the right source was retrieved but ranked low, the ranking must be tweaked. You must be able to drill into any part of the system. That’s much easier if it’s all in one system
Vespa references
| Need | Vespa concept | Why it matters | Docs |
|---|---|---|---|
| Store searchable units | Schema and chunks | Defines text, metadata, embeddings, ids, permissions, summaries, and rank profiles. | schema, chunks |
| Build constrained queries | Query API and YQL | Adds filters, permissions, source limits, dates, rank profiles, and query embeddings outside the prompt. | Query API, YQL |
| Retrieve candidates | Text matching and nearest neighbor | Supports exact terms, domain vocabulary, semantic search, and hybrid retrieval. | text matching, BM25, nearest neighbor |
| Rank results | Rank profiles and phased ranking | Combines cheap first-pass ranking with stronger second-phase or global reranking. | phased ranking |
| Return usable context | Document summaries | Controls which chunks, metadata, ids, scores, and source fields come back to the agent. | document summaries |
| Debug failures | Query tracing and app logs | Shows the query, filters, hits, ranking, selected context, and answer id. | query tracing |
| Build eval data | App-owned annotation store | Keeps relevance labels tied to query ids, document ids, chunk ids, and answer ids. | App-owned |
| Build a full RAG app | RAG blueprint | Shows schema design, chunks, query profiles, hybrid retrieval, phased ranking, and ranking evals. | RAG blueprint |
Whatever stack you use, when an answer is wrong you should be able to see whether the source was never retrieved, filtered out, ranked too low, or dropped before generation. Vespa is strongest when that diagnostic path needs to run over large, changing corpora with ranking, filtering, and serving constraints built into the same retrieval layer.