My LLM gateway exposes stable route names. The application asks for lab-chat-hosted and doesn’t care what sits behind it, which is the whole point of putting a gateway there. I was multitasking and accidentally updated that route to a different model in the same family while an eval run was going.

Then I ran a comparison between two agent architectures and found that the new one was dramatically faster. P50 dropped from 21.6 seconds to 1.68 seconds.

Thirteen times faster. That number had nothing to do with the architecture. It was the reasoning tokens.

The only reason I caught it is that my results files record the resolved model rather than the friendly name I asked for. Two runs that both said lab-chat-hosted were not running the same model. Without that one field I would have written up a thirteen-fold architecture win in good faith, and I would have believed it.

That is the thing I keep running into lately. The measurements we take aren’t unbiased. Worse, the bias isn’t random: an instrument built alongside a system can inadvertently encode that system’s behavior as normal, so anything that behaves differently reads as failing. Most of the biases I’ve found in my own observability systems have pointed the same direction, so “I measured it and the new approach was worse” deserves more suspicion than it usually gets.

What I’m actually building

All of this is a personal research project running on my own hardware, on my own time, with made up callers and fake dispatch tickets. The agent is a roadside assistance intake bot I built as a test bed for ideas. It exists so I have a realistic multi-turn workload with clear goals and constraints that I can point instruments at.

I picked roadside assistance because the ground truth is obvious to anyone: you can tell whether a tow was dispatched, whether a location got captured, whether an emergency escalation was actually needed. I don’t need any domain expertise to write the assertions, or to tell when the agent has gone off task.

I’ve been diving into evaluating model quality and performance, and instrumenting testing and benchmarking capability so I can iterate on my agent and harness with confidence. I set up a LiteLLM gateway so I could easily swap between model providers and my own hosted llama model on my home lab box. The gateway produces traces to Langfuse that are enriched with my application side telemetry, enough to evaluate correctness. I’ve built the same intake agent two ways for my research thesis: I want to discover what open weight models can be used within the decomposed agent design pattern, and how much state driven harness is necessary for task completion and safety.

The more intelligence you pay for, the more you expect to be able to defer autonomy and reasoning. What I did not expect was how much of the difficulty would live in the measuring rather than the building.

Latency is not one number

I’ve come to see that even just measuring model latency is challenging in itself. What are we even measuring? TTFT, mean time, P95?

TTFT is a good measure, but in a voice call we care more about time to first sentence, since that’s the point at which the speech has buffered enough with the synth provider to sound out a sentence with the proper inflections. P50 duration is a useful focal point. However, an inference provider serving 800ms responses with a 10 second P95 is disqualified from my use.

Finding the right thinking level and service tier to compare 1:1 is also a bit of a challenge. With different defaults and settings for reasoning effort (which a voice call can afford none of in the turn latency budget), service tier options, and router options (when using Fireworks, the Turbo router is your ticket to fast inference), it becomes a matrix of toggles and knobs that all must be configured before benchmarking.

I’ve had my fair share of multi-pass evals run only to realize I had misconfigured a setting and had to scrap the results and run again. The reasoning token mixup above was just the one that nearly made it into print.

Token caching plays an important role for latency too. Cold start versus warmed can throw off early probes, and it matters again when measuring and comparing cost.

The cost result that went backwards

Models have a tokens in and tokens out cost, plus separate pricing for caching, although the rules differ by provider.

I expected a cost reduction when moving from a classic hub and spoke agent model (one system prompt with tools) to a multi-agent with state driven handoffs. What I saw was actually a cost increase, due to token caching misses.

The numbers, on the same workload:

billed prompt tokenscachedcost
hub and spoke4,66923,803$0.0134
state driven handoffs15,2336,912$0.0201

The second architecture sent 22% less total context and paid 3.3 times more for it. Cache coverage fell from 83% to 31%.

A smaller portion of the payload was cacheable, because a dynamic state block sat inside the system prompt. Cache coverage ends at the first token that changes, so everything after that block gets reprocessed on every call, including the entire conversation history sitting behind it. In this experiment both agent harnesses had an ever accumulating chat history, which grew the cacheable tokens in the first paradigm and the uncacheable ones in the second.

I would not have found this without per-call cache metrics. Total token count was down, which is what I was watching, but the bill went the other way. Cost alone is not enough to decide what’s good, but the facts around cost are less obvious than they first appear, and they depend on model settings and inference provider, including things like service and router tiers.

There’s a second lesson buried in there. An earlier version of this same comparison showed the new architecture running 42% cheaper, and I believed it for a while. It was cheaper because it was broken: the agent was stalling out mid-conversation and never completing, so it simply made fewer calls. A metric can look great for reasons that have nothing to do with what you think you’re measuring.

The harness that only broke for one architecture

In addition to my static per turn eval cases, I had an end to end simulation harness. In early development it only played scripted conversations at the agent. Caller lines were fixed and runs were half deterministic, so I could benchmark consistently. It actually worked fine for a while, across many versions I’d built, all testing the same flows.

Then I tested the architecture that gates progress on state, and watched the agent ask the same question eight turns in a row.

AGENT: ...are you or anyone else injured? And is your vehicle in a safe spot?
AGENT: ...Are you pulled over and out of the way of traffic right now?
AGENT: ...Are you pulled over and out of the way of traffic right now?
AGENT: ...I still need to check on your safety first. Are you pulled over...
AGENT: ...I still need to know: are you pulled over and out of traffic...
AGENT: ...Right now I need to confirm your safety before anything else...

The agent was doing exactly what I built it to do. It refused to move forward until it had a confirmed safety status, and it handled every digression along the way politely. The scripted caller never answered, because it wasn’t listening. It was reciting the next line about a Subaru and traffic on the way back from their sister’s, stuck in a situation I hadn’t accounted for.

Obviously a script can’t answer all questions. That isn’t the interesting part. The interesting part is that this limitation had been sitting in the end to end harness for a long time and had never once mattered, because every architecture I’d tested with it barreled ahead and worked with whatever the caller happened to volunteer.

On paper the result was that the new design “couldn’t complete a call.” What was actually happening is that my instrument punished the exact behavior I had built the new design to have. If I hadn’t gone and read the transcripts, I’d have concluded the architecture was a failure and moved on, and the harness would have gone on looking fine for everything else I pointed it at.

The judge that scored the wrong thing

Deterministic assertions get you a long way, but some things resist them. Whether an agent is pleasant to talk to is not something I can regex. So I did what everyone does and set up an LLM judge to compare transcripts pairwise, position swapped, with a different model family from the one under test.

I predicted the state driven architecture would lose slightly on naturalness. It won. The judge preferred it, and tagged the simpler agent as more form-like in seven out of nine comparisons.

The tell was in the justifications, which kept praising resolution and dispatch. I had asked which conversation was more natural, and the judge had answered which one solved the caller’s problem. Those are correlated, and one of them is much easier to see.

What convinced me was the calibration number. I had run the judge three times over identical transcripts to measure how much it disagreed with itself, and it came back at a flip rate of zero. Perfectly consistent. At the time I read that as a well calibrated instrument.

It’s the opposite. Naturalness is a genuinely hard call and a judge weighing it should disagree with itself sometimes. Perfect consistency means it found something clear to key on, and task success is about as clear a signal as there is. The judge was reliable, just not about the thing I asked it.

I re-ran it with outcome matched pairs and the transcripts truncated before either agent resolved anything. The flip rate went to 0.8, well above the difference I was trying to detect, and the honest answer became that I don’t know which is more pleasant to talk to. That’s a worse result and a more truthful one.

Worth noting: the judge wasn’t broken. On a different comparison, where both agents had similar completion rates, it separated them cleanly on form-feel. It could see the thing I wanted. It was just drowned out by a louder signal whenever one was available.

Suspect the good news

I’ve learned that creating an eval harness capable of simulating usage takes lots of decisions along the way. Judge model selection, cases, evaluation criteria - both LLM and code/assertions. These all influence the way we measure quality, and it takes discipline to sit through runs without jumping in and making that one prompt tweak that you think will pump the numbers.

The discipline that has actually paid off is writing predictions down before the run, including what result would mean I was wrong. I predicted the new architecture would lose on naturalness. It won, and my first instinct was to accept the good news. Pre-registering the prediction is what made me look twice at a result that was too good, which is exactly when nobody looks.

The related habit is running the baseline several times before claiming anything. At one point my run to run variance was larger than the effect I was trying to measure, which meant a single comparison could show almost any answer I wanted. Three runs of the same thing cost a few dollars and tells you whether you have a finding or a coin flip.

I’ve gotten more comfortable with relativistic performance. Comparing models against one another gives honest context for what’s possible, and where the agent framework needs to bring more to the equation.

The direction of the bias

Looking back, the ones I’ve found so far are all the same shape.

What quietly changedWhat it looked like
A route pointed at a different modela thirteen-fold architecture win
Assertion code revised between runsa regression that never happened
Thresholds calibrated on one variantthe other variant failing
Substring matching, no negation handlingcorrect refusals scored as violations
A scripted caller that can’t answer questionsan architecture that can’t complete a call
A judge keying on task successa result about naturalness
A container tag that movesnothing, until a pod restarts

The common thread is that an instrument built while developing one system encodes that system’s behavior as the definition of normal. Its thresholds are calibrated on that system’s distribution. Its string matching is tuned to that system’s phrasing. Its scenarios exercise that system’s interaction style. When something behaves differently, the difference shows up as the new thing failing rather than as the instrument being narrow.

Every one of these favored the incumbent. This points to systematic bias, and it means the most dangerous measurement is the one confirming that what you already have beats the thing you were considering, because that is the result the instrument was built to produce.

Some mechanical things that have helped since: record the context inside the artifact, and refuse to compare across a mismatch. My results files now carry a hash of the assertion code, the scenario set, the resolved model, the container digest. If two runs disagree on any of them, the comparison tool errors out instead of printing a table.

None of this matters much on a toy. It matters enormously anywhere that stability is the product, where adding one feature can’t risk regression across every capability that came before. That’s the general case eval driven development exists for. What it means in practice is lots of fiddling with instructions to get all the good with the least amount of bad. It’s time consuming and can be misleading given baked in assumptions or bad judgments. It’s also enormously tempting to believe that the perfect model is out there that will make all your eval cases go green, and you just haven’t found it yet.

The constraint underneath all of this

The service oriented philosophy that I’ve cultivated over the years has influenced how I approach building agents. The more well scoped and encapsulated agent components are, the easier they are to evaluate on quality and performance, both in service of refining instructions but also in finding the right model for the role.

In the world of low latency voice agents where I specialize, it’s not an option to throw a frontier lab model with high thinking effort at a turn and call it a day. The caller would be waiting ages for a response, get fed up, and pull the ‘representative’ rip cord. Maybe someday that changes, and I’d certainly welcome it, but for now I’m continuing to focus my efforts on these skills so I can build low cost, high quality agents, and I’m willing to take the time to figure it out.

The instrument is a bigger part of that work than I expected when I started. Many biases found so far, all pointing the same way. I don’t take that to mean my measurements are air tight now, just that I know how many I’ve caught, and not how many I’ve missed.