The call usually starts the same way. The assistant is slow, some answers are wrong, and the team has already tried to fix it by adding things: a query rewriter, a second index, a re-ranker, an agent that decides whether to search at all. Each addition made sense the day it shipped. Together they made a system nobody can reason about.
I’ve seen this often enough that my first question is now: what can we take out?
How retrieval gets heavy
Nobody sets out to build a behemoth. Retrieval gets heavy one reasonable fix at a time. An answer comes back wrong, so someone adds a step that catches that case. Latency creeps up, so someone adds a cache. The data lives in three places, so the pipeline learns to query all three and merge the results.
Every stage is a place where the right passage can get dropped.
The cost isn’t only speed. Every stage is a place where the right passage can get dropped, and once an answer is wrong, you can no longer tell which stage dropped it. Debugging turns into archaeology.
What a rebuild looks like
One of my clients, a property-rental platform, used an AI agent to answer guests’ messages. Replies took minutes. The retrieval behind the agent had grown into what their CTO later called “an over-engineered behemoth.”
We rebuilt it from first principles. The biggest change was the least exciting one: the property data lived in several disconnected sources, so we consolidated it into a single vector index the agents could search. One place to look, and one place to fix. We also moved ingestion from a monolith to small event-driven services, so an edited listing reached the index without reprocessing everything else.
Replies went from minutes to under a minute, and the answers got more accurate. Nothing in the new design was clever, and that was the point.
Where I’d start
If your retrieval feels slow or unreliable, try this before you add anything:
- List every stage between the question and the answer, and what each one is for. If nobody remembers why a stage exists, it’s your first candidate.
- Build a small test set: fifty real questions, each with the passage that should answer it. Without it, you can’t remove anything safely.
- Remove one stage at a time and rerun the set. Keep what earns its place.
- Put your data in as few places as you can. Most retrieval bugs turn out to be data bugs.
Sometimes you do need the re-ranker. You should be able to point to the questions it fixes.
Retrieval has one job: put the right passage in front of the model, quickly, in a way you can debug at 3 a.m. Anything that doesn’t help with that job can go.