Three competing axes dominate an agent deployment, and moving any of them reliably is extremely difficult without the right observability. Cost, quality, and latency are not always diametrically opposed, but you rarely get to pull one lever independently. Very often when you move one, at least one of the others moves with it. A faster model that misses extraction is not a latency win. A frontier model that costs an obscene amount and requires thinking tokens isn’t functional for a voice turn, so it’s not a quality unlock either.
What matters to a user calling for a tow while standing on the side of a busy road? That their information is captured quickly and accurately, and that the tow truck ACTUALLY COMES! What matters to us as implementers is that we can control our costs, so the service is either cheaper than the human alternative or a significantly better experience that justifies the spend.
Measuring the axes
Iterating on harnesses, model selection, router settings, etc. is near impossible without a high quality signal to show whether changes make a positive impact on cost, latency, and quality. Over the last month I built four harnesses for the same roadside assistance scenario and measured them against each other, all of them agent led, with the agent driving and asking the questions. The monolith is one system prompt with all the tools and a conversation history that keeps growing. The multi-agent version swaps ‘specialist’ prompts and tools in as the call moves along. The state driven hybrid puts a state machine in charge of the phases of the call and the context for each, and the model handles the language within a phase. The role based harness splits the turn into components that capture information, decide the next action, and speak, each with its own prompt and its own model.
Evals tell you whether the call went right. They don’t tell you which part of it was slow, expensive, or wrong. A monolith can get away with a turn total, since there’s one model and one bill. A role based harness can’t. If five components ran and the turn took three seconds, those three seconds aren’t something you can act on until you know whose they were.
For me, this involved setting up my LiteLLM gateway to produce traces and full callbacks so they could be attributed to specific components in my test harness. Every hop records which component ran, which model the route actually resolved to, and the tokens, cost, and latency for that call. The resolved model matters more than it sounds like it should. In my last essay, a route quietly changing models nearly handed me a thirteen-fold architecture win that didn’t exist. That one was my own mistake, but it won’t always be. When I went back to run DeepSeek V4 Flash again, it was gone. If an MLOps team maintaining the gateway had pointed that route at the latest version to keep things working, every run after that would have been against a different model, and nothing in the results would have said so. Routing each component separately multiplies those chances, so the trace has to record the model that actually ran, not the name you asked for. Combined with an eval benchmark suite, simulation integration runs, and aggregate analysis in Langfuse, that let me compare the harnesses, and then explore model selection and inference options for the various components.
The monolith
When you operate with a monolith, even a multi-agent setup with prompt and tool swaps, you have a couple knobs: the model and inference settings for the agent (service tier, temperature, timeouts/retry). The LLM must do every role well, and quickly: capture and normalize information (slots), take actions at the right time based on rules (tools), and speak in a fluid, natural way while adhering to strict guardrails.
One of the main benefits here is token caching, which helps on both latency and cost. With a fixed system prompt and growing user messages, the cacheable tokens keep growing in chunks. That’s all good until the context window gets large enough on longer conversations (including tool use) that a.) the latency becomes increasingly noticeable, and b.) the large context starts to degrade instruction following and quality suffers (context rot). I’ve seen it firsthand, and coming back to the user’s goals, the tow truck must actually come! When something critical rides on the LLM performing accurately at the end of a long call, it’s worth taking a step back and evaluating the design pattern.
A simple roadside assist intake isn’t where that really bites. You collect a location, a vehicle, and a service, confirm it, and dispatch. The call is designed to end before the growing history becomes a problem, so on cost and latency the monolith held up fine in my runs.
Quality was a different story. When the LLM is the brain, it does both halves of a dispatch: it decides the caller confirmed, and it sends the truck. Under pressure, like a caller saying “just send someone” or a member who couldn’t be verified, the monolith dispatched anyway, and it could call the confirmation tool itself to license that action.
On short, relatively simple, or low stakes user led interactions, the monolith or ‘specialist’ prompt swap multi-agent pattern is often sufficient and really the best choice. When the call gets long or the stakes get high, that’s where the alternative earns its keep.
A role based harness
In a deconstructed, role based design pattern, we gain greater flexibility and control at the expense of a larger, more opinionated harness. A distributed agent harness is like a service oriented architecture for agents, and like its platform engineering counterpart, it requires robust interface definitions and clear boundaries for all the pieces to work together harmoniously. Each component deals with a slice of the turn: accurately capturing information, detecting emotion, speaking responses that are safe and grounded, and taking rules based actions.
This segmentation of responsibility allows for more targeted prompts and smaller context windows. For example, an ‘extractor’ component is instructed to capture key pieces of information without any knowledge of how or why they matter to the overall task. This has been found to reduce hallucinated information, and also unwarranted actions in conversations where friction is present and the LLM reaches for an ‘escape hatch’ to complete the task, creating a situation that wasn’t truly there. In my role based harness, a ‘yes’ confirms the value that was just read back to the caller, and “just send someone” doesn’t.
It’s also the answer to context rot. The more structured data we accurately produce and maintain, exposing it to the LLM components where needed, the less we need to maintain the full user/agent message history. This is by far the largest ‘unlock’ of the distributed design pattern. What enables decoupling input and output is also what creates proactive ‘memory’ that can persist in place of raw user/agent messages. It’s like micro conversation compression happening on every turn, and it has an interesting side effect of reducing the overall LLM specification requirement.
The trade is caching. There are fewer cacheable tokens since the message prefix is no longer stable, and I’ve already been bitten by what a broken prefix does to the bill in the cost result that went backwards. But with a consistent token count, the latency and quality of the conversation don’t degrade as it goes on. We can also start pulling in smaller, cheaper models, since the instructions and context are smaller and more focused.
That’s the other thing a role based design pattern gives us: a separate set of levers for each node. With classification, for example, we can use newer models like Jev, TypeSafe AI’s new system one model, in targeted ways that contribute to the context of the system at large. The model that satisfies the latency and quality requirements of the reply component can be significantly different from what’s required for information capture, where complex narrative has to be coerced into structured data.
I got to see that move in one of my experiments. I held Kimi K2.6 on the reply component and moved information capture from Kimi to DeepSeek V4 Flash. Both wirings got all five of my cooperative callers to a ticket, including the one who needed a tow. The split cost about half as much, $0.023 against $0.043, because capture is most of the tokens and Flash is the cheaper model. Quality held and cost moved, which is exactly the lever a monolith doesn’t have.
The programmatic nature of a role based agent also creates state driven opportunities, like calling tools only under specific state conditions, and emitting predefined utterances that must be spoken exactly for compliance or accuracy purposes. In my harness, the safety line after a ticket and the ETA are fixed lines. They cost zero prompt tokens, take no model hop, and come out the same every time. A monolith can be instructed to say them, but it can’t take the turn away from the model once the ticket exists.
The cost is rigidity. A more strictly designed and codified call experience fits certain types of interactions well and others not as well. It works best when the agent is ‘driving’ and asking the questions, which is typically the case in real business scenarios.
Where things broke
Every harness has its failure modes, and it was really interesting to discover some new ones and get a real feel for each pattern and its weaknesses. A few from the model runs on the role based harness stuck with me.
Nemotron was the fastest and cheapest hosted model I tried, with turns around 0.7 seconds at the median. It read a tow back to the caller, and on the very next turn called it a lockout! Fast and cheap doesn’t count for much when a lockout truck shows up for a car that needs a tow.
A local Qwen 2.5 3B running on my lab box (small server, no GPUs) sat around 30 seconds a turn! The rambling caller and the tow caller sims both hung up before they got a ticket. There’s no invoice for it, but that doesn’t make it cheap, since the cost is the caller’s patience which wore out.
On the Flash and Kimi split, one capture turn with the rambling caller stalled for about eight minutes. No caller is waiting eight minutes for anything, and that’s a lesson in the basics more than in model selection: every model call needs a timeout, and a retry/failover. In production code I’d never ship a network call without them. This was my test harness, so I cut the corner, and it bit me. What saved the analysis was the detail of traces. The median turn on that wiring stayed around three seconds, and the reply component never went over two. One bad request showed up as one bad turn, pinned to capture, instead of an average that made the whole model usage look slow.
These are are from one or two passes, so they’re a rough shape and interesting lessons to build from, not declarative of a winning combination. What they show is that you can’t make changes with any confidence unless you can see which component moved which axis.
It depends
The partly unsatisfying finding, and the real answer to the ‘what’s best?’ question, is that it depends. Anything cleaner would be too good to be true. What it depends on are the things that fall out of discovery and the design phase, before the first prompt gets written:
- How long are the conversations? Short calls favor the monolith, where a growing cache keeps cost and latency down. Long ones favor structured state over a growing history, which keeps quality from rotting as the call goes on.
- How hard is the information to capture? If complex narrative has to become structured data, give capture its own component and its own model. That’s where quality is won or lost, and since it’s most of the tokens, it’s also where cost moves the most.
- What happens if the model acts on a bad read? When an action has real stakes, separate who confirms it from who takes it. That’s quality you shouldn’t leave to a single prompt.
- Are there words that must be said exactly? Then they shouldn’t go through a model at all. A fixed line costs nothing, adds no latency, and comes out right every time.
- What’s the latency budget? In voice, every hop is time the caller hears, so you spend per-component model selection on speed and keep the turn lean. Email/text may afford to wait, so the same lever gets pulled the other way: cheaper, slower models, and more deconstruction for better accuracy at the cost of more hops per turn.
There’s no one design pattern that works for every situation, and the right one for today’s workload may not be the right one as it grows and changes. The problem is that most teams building agents today are stuck with one model doing every job, not because a monolith is right for them, but because they don’t have the harness options or the observability to see what the alternatives would do. Pick the harness that fits the agent, measure every component, and above all else: make sure the tow truck actually comes.