Measure, don’t guess.
Every model change goes through an eval harness. Two models answer the same conversations, an LLM judge picks the better reply with the order swapped to cancel its bias, and a candidate ships only if it clears gates on both win rate and output format. Every run reports what it cost.
Chose an eval harness over tuning prompts by feel