Jesús Martínez
ES
← Work · Case study

A voice coach that grades itself.

Charming Chad is a game for practising real conversations. You talk out loud with an AI persona at a simulated venue, and an AI referee scores how it went. I built the whole backend, plus the test bench that grades the AI before any change reaches a player.

Client
Palm Island
Role
AI architect and sole backend engineer
Timeline
Feb 2026 to Sep 2026
Built with
Python, FastAPI, Supabase, LiveKit, Deepgram, Cartesia, OpenRouter

The situation

Practice needs a partner who tells you the truth.

Getting better at conversation takes practice with someone willing to give you honest, structured feedback, and that is exactly what’s hard to find. A text chatbot can’t reproduce the pressure of speaking out loud in real time. It can’t tell you whether you came across as smooth or creepy, whether you pushed too soon, or whether you kept repeating yourself.

The problem

A referee nobody can check isn’t a referee.

The product needed live voice, personas consistent enough to be worth practising against, and scores players would accept as fair. That raised the harder question: every time a model or a prompt changes, how do you know the coach got better, and not just different?

The approach

Build a game, not a chatbot.

Every session has a beginning, a middle and an outcome, and the player is being judged the whole time. So the system runs the live conversation on one side and keeps score on the other, and the score never depends on the conversation’s process staying alive.

01

Join

The player enters a real-time voice room with a persona at a venue, and each venue has its own difficulty. A state machine walks the session from the greeting to the evaluation.

02

Talk

Speech becomes text, a language model answers in character, and the reply comes back as a voice, with a live video avatar if the player opts in.

03

React

Every reply carries the persona’s emotion, which moves a 0 to 100 receptiveness score. Push too hard or turn crude and the persona goes cold. Real abuse ends the session.

04

Score

When the session ends, that venue’s own referee scores it from 0 to 100 against its rubric and returns feedback by category: what worked and what didn’t.

Decisions and trade-offs

Three calls that made it work.

01

Measure, don’t guess.

Every model change goes through an eval harness. Two models answer the same conversations, an LLM judge picks the better reply with the order swapped to cancel its bias, and a candidate ships only if it clears gates on both win rate and output format. Every run reports what it cost.

Chose an eval harness over tuning prompts by feel
02

Keep score for free.

The persona already returns an emotion and its intensity with every reply. Turning that into the receptiveness score takes plain code, not another model call, so the game’s state costs nothing extra and every rule has a unit test.

Chose deterministic rules over an extra LLM call per turn
03

Never lose a score.

Scoring first ran as a background task inside the server, and a deploy that landed mid-session simply lost the result. Now every score is a durable job: claimed with a lease, retried with backoff, finished exactly once. A deploy now costs time, not results.

Chose a durable job queue over in-process background tasks

The twist

The best model didn’t ship.

The harness drove a migration off the model in production, and one candidate came out on top. Then we measured its latency: about twice as slow as the runner-up, too slow for a live voice conversation, where every pause is audible. We shipped the runner-up instead. It still beat the old model on all six rubric dimensions.

#1on quality, and about twice as slow
100%win rate for the runner-up we shipped

The outcome

Where it landed.

12.5% → 82.3%replies in the exact format the game reads, before and after the migration
100%judge win rate over the previous model, on all six rubric dimensions
4,568tests across unit, integration and end-to-end, with zero failures at the last release
20personas, each tunable from a data file instead of a deploy

Live in production with real players. Built solo in about six months, while a separate frontend team worked from the API contract and integration guides I wrote for them.

For the technical reader

Engineering notes

Examples the judge can’t be fooled by

Each persona speaks from a curated library of example lines. The judge measures consistency against a separate, held-out set, so a model can’t win by parroting its own examples, and an n-gram overlap check flags any reply that echoes them past a fixed threshold.

A voice agent inside 512 MB

The voice worker was being killed for running out of memory, and a hard kill left sessions open, locking players out behind the one-open-session rule. Admission control now reads real container memory instead of CPU, heavy imports stay out of job processes, and the avatar renders in the cloud, so video never lands on the worker. A reaper closes whatever a crash leaves behind.

Guards with a replaceable core

Abuse, crude propositions, content ceilings, prompt injection and in-fiction requests for contact details each sit behind one narrow check. The first versions are deterministic and easy to review; a classifier can replace any of them later without touching a call site. Player turns always travel as user messages, never spliced into the instructions, and injection detection only logs, because stripping matched words would corrupt honest messages.

Avatar research before code

I reviewed fourteen real-time avatar providers. The real constraint turned out to be vendor content policy, not technology, so the recommendation was to get written acceptable-use clearance before committing. It shipped behind a provider-agnostic interface and a feature flag, opt-in per session, with an audio-only fallback for every failure.

Two processes, one write path

The API and the voice worker run as separate services, coordinated by job dispatch. The worker never writes to the database; it calls back through internal endpoints with a service token, so every write has one authoritative path. Releases went out weekly, with written pre-flight checklists and blast-radius checks before any migration that changed data.

  • Python
  • FastAPI
  • Supabase
  • PostgreSQL
  • Pydantic
  • LiveKit Agents
  • Deepgram
  • Cartesia
  • OpenRouter
  • LemonSlice
  • Langfuse
  • Stripe
  • Sentry
  • Docker
  • Render
  • React
  • TypeScript

Next project · Legal AI · 2026

Immigration answers an attorney can stand behind

Got an AI feature nobody can prove is getting better?