Jesús Martínez
ES
← Work · Case study

Immigration answers an attorney can stand behind.

People with immigration problems ask in their own language, often scared, about law that keeps changing. I built VizEx, a public tool that answers one question with a memo-style overview grounded in current agency data, and never crosses the line into legal advice.

Client
Liveday
Role
AI lead and sole engineer
Timeline
Apr 2026 to Jul 2026
Built with
Python, FastAPI, Supabase, Claude, Tavily, React

The situation

Plenty of answers. Nobody accountable.

Immigration questions come from people in distress, in many languages, about a body of law that shifts under political pressure. What they find for free online is either too general to help or comes from sources nobody answers for.

The problem

One wrong sentence is practising law.

Telling a member of the public that they qualify for a visa, or that their application will be approved, is legal advice. An AI that says it once, in any language, turns a lead source into a liability. The supervising attorney, with 35 years of practice, needed a system safe enough to put his name on.

The approach

Write the rules down. Then grade every answer against them.

The attorney issued formal written rules for how the AI may behave. I built the system that enforces them on every answer, and a test tier that checks the answers themselves, not just the code that produces them.

01

Ask

One question, in any language, with no account, through a widget embedded on the public site.

02

Screen

A fast, inexpensive model reads the question first, before the expensive one is spent on it.

03

Answer

A stronger model writes a one-to-three-page overview, using live web search for current agency data. It maps the paths and the risks, names any complication early, like an overstay or a prior denial, and never predicts an outcome.

04

Deliver

The answer reaches the browser in 20 to 90 seconds, in the reader’s language. It closes with the required disclaimer and an offer to book a consultation, and urgent situations get an urgent offer.

05

Learn

The attorney reviews candidate answers in his own tool. The ones he picks are matched to new questions and shown to the model as examples.

Decisions and trade-offs

Three calls that made it work.

01

Grade the rules, not just the code.

A fourth test tier grades generated answers against the attorney’s rules. Each scenario pairs a question with an ideal answer he wrote; the non-negotiable rules are checked by code, the rest by an LLM judge. It runs outside the everyday suite, so routine commits never spend model budget.

Chose graded scenarios over trusting the prompt
02

Don’t hold the line open.

An answer can take a minute and a half. A new question becomes a job in the database, a worker claims it and writes the answer back, and the browser gets it the moment it’s done. A sweep every minute picks up anything that stalled.

Chose an async job queue over a long-held connection
03

Curate, don’t fine-tune.

The attorney drew the line: the live model is never retrained. Instead, the answers he rates best are retrieved by similarity to each new question and given to the model as examples, always in the same language and behind a kill switch.

Chose curated examples over fine-tuning

The twist

The rules beat the model.

The obvious way to make the answers safer was a stronger model. That isn’t what did it. VizEx became safe to put an attorney’s name on through the test tier: his rules, graded scenario by scenario against ideal answers he wrote, with the rules that can never bend checked by code instead of by a judge. Every run is logged against the version that produced it, so a change that breaks a rule shows up in testing, not in an answer someone has already read.

1wrong sentence is enough to practise law
4thtest tier, grading the answers, not the code

The outcome

Where it landed.

~100questions answered in the first two and a half weeks live
~33consultations booked from them, about one in three
~5signed matters, at roughly $5,000 each
3languages answered natively: English, Spanish and Portuguese

Live and public, and the client keeps extending it. Built solo in about three and a half months: the backend, the attorney’s review tool and the embeddable widget.

For the technical reader

Engineering notes

A job queue inside Postgres

Inserting a question fires a webhook to the worker, which claims the job atomically with FOR UPDATE SKIP LOCKED. The finished row reaches the browser over Supabase Realtime, so a long generation never depends on a connection staying open. Failures requeue up to a limit, then fail loudly into error tracking.

Rules with triggers

Some passages must appear whenever their trigger does: automatic visa voiding after an overstay under INA §222(g), the three- and ten-year unlawful-presence bars with their expiry computed from the departure date the reader gives, and a flag on areas of policy that are actively changing. The model also emits an urgency tag that the pipeline reads and strips before anyone sees it.

Native in every language

Answers come back at native register, with no English fragments and no untranslated headings, and the disclaimer is translated too. Portuguese defaults to the Brazilian register unless the question suggests otherwise.

Three answers that are actually different

The review tool generates questions across a set spread of emotional registers, and measures that the generator really produces that mix. Each question gets three candidates at pinned temperatures with different approaches. A similarity gate then compares their openings, where the difference really lives, and separately rejects near-duplicates.

A widget that asks nothing of the host

Self-contained vanilla JavaScript in a same-origin iframe, so the public site needed no CORS setup. Three locale builds, staged loading messages for a wait of up to 90 seconds, booking links in the reader’s language, and an integration guide for the outside agency that runs the site. Behind it, 1,124 tests, with more than twice as much test code as application code.

  • Python
  • FastAPI
  • Supabase
  • PostgreSQL
  • pgvector
  • pg_cron
  • pg_net
  • Supabase Realtime
  • Claude Sonnet
  • Claude Haiku
  • OpenAI embeddings
  • Tavily
  • Calendly
  • Sentry
  • React
  • Docker
  • Render

Next project · Multi-agent · 2025

Eight AI agents, one Slack workspace

Need AI in a field where one wrong sentence costs you?