Jesús Martínez
ES
← Work · Case study

Answers in under a minute, not minutes.

Cortado, a property-rental platform, answers its guests’ messages with AI agents. The retrieval behind them had grown into what the CTO called an over-engineered behemoth, and guests were waiting minutes for a reply. I rebuilt it from first principles.

Client
Cortado
Role
LLM engineer
Timeline
Aug 2024 to Nov 2025
Built with
Python, Pinecone, LLM agents, MCP, event-driven microservices

The situation

A guest asks. The agent goes looking.

Answering a guest well is a retrieval problem before it’s a language problem: the agent needs the right facts about the specific property in front of it. Those facts lived in several separate sources. And in rentals, a guest who waits too long can simply ask someone else.

The problem

Every extra hop is time on a typing indicator.

The system that pulled those facts together had grown past the job, and replies took minutes. Ingestion had the same problem from the other side: one monolithic pipeline, where every stage moved at the speed of the whole.

The approach

Don’t tune it. Rebuild it.

Optimizing an over-engineered system tends to keep whatever makes it slow. So we started from what the agent actually needs when it answers, and kept only that: one place to look, and an ingestion pipeline that runs in small pieces.

01

Change

A listing is edited somewhere upstream, and that edit becomes an event.

02

Ingest

Small event-driven services pick it up and process only what changed, instead of pushing everything through one monolith.

03

Index

Property data from every source lands in one Pinecone vector index.

04

Answer

The agent retrieves from that one index and replies to the guest in under a minute.

Decisions and trade-offs

Three calls that made it work.

01

Rebuild, don’t tune.

In an over-engineered system, the parts that cost the most time are usually the hardest to remove without admitting they were never needed. Starting from first principles made that call on purpose instead of inheriting it.

Chose a first-principles rebuild over optimizing the old system
02

One place to look.

Property data was spread across sources, and an answer depended on which one a query happened to reach. Consolidating it into one vector index gave the agents a single place to retrieve from, and accuracy improved with it.

Chose one vector index over querying several sources
03

Events, not a monolith.

Ingestion moved to distributed, event-based services. A slow stage stopped setting the pace for the rest, processing latency dropped, and an edited listing reached the index without reprocessing everything else.

Chose event-driven ingestion over one monolithic pipeline

The twist

Taking things out made it better.

Taking stages out of a retrieval pipeline is supposed to trade quality for speed. Here it bought both. Each stage had been added to fix a problem of its own, and each one stood between a guest’s question and the facts about the property. The leaner path beat the old one on both counts.

Fasterfewer stages to wait on
More accuratefewer places to lose the right fact

The outcome

Where it landed.

< 1 minagent reply time, down from minutes per response
1vector index for property data that used to be scattered

Retrieval accuracy and guest satisfaction improved too. The new design was simpler than the one it replaced, and easier to fix when something went wrong.

In their words

“Jesús was integral to helping us overhaul the RAG system that powered Cortado’s AI guest messaging software. With Jesús’ help, we rebuilt our retrieval system from first principles, reworking an over-engineered behemoth into a lean, mean retrieval machine. My team would recommend Jesús to anyone looking to master modern machine learning for the age of artificial intelligence.”

Harry DubkeCTO, Cortado

For the technical reader

Engineering notes

Deciding what to delete

An over-engineered system is rarely built carelessly. It’s built by people answering real constraints one at a time. Cutting it down meant working out which constraints were still real and which were being designed around long after they stopped mattering. Getting that wrong removes something load-bearing.

One store forces one truth

Once several sources become one index, every place they used to disagree has to be settled. Before, it settled by accident, by whichever source a query reached. After, it had to be a decision someone made on purpose.

Migrating under a live product

The move from a monolith to events ran while guests were still being answered. The old path kept serving while the new one took over stage by stage, and the in-between states are the ones nobody designs for.

  • Python
  • Pinecone
  • LLM agents
  • MCP
  • event-driven microservices

Next project · Document AI · 2025

$77M in political ad spend, made searchable

Is your AI assistant keeping customers waiting?