Skip to content
Apsan Works

Why AI Agents Fail in Production (And What Actually Fixes It)

The prototype worked, then nothing shipped. The failure is almost never the model. It is five specific engineering gaps, and all are avoidable.

6 min read

There’s a specific and very common shape to a failed agent project. A prototype gets built in a week and it’s genuinely impressive. Leadership sees it. Budget appears. Then, somewhere between month two and month six, it quietly stops being mentioned.

We’ve been brought in to rescue enough of these to notice that the failure is rarely the model. It’s almost always one of five gaps, and all five are known problems with known solutions.

1. There was never a way to tell if it was working

This is the big one. It causes more stalled projects than everything else combined.

A prototype gets evaluated by a person trying it and saying “yeah, that’s good.” That works fine for ten examples. It doesn’t work for the thousandth run, and it fails completely the moment you want to change something, because the only way to know whether your change helped is to try it by hand again, and your memory of how it behaved last week is not a baseline.

So the team stops changing things. The agent freezes at whatever quality it happened to reach, because every modification is now an unmeasurable risk. That isn’t a shipped product. It’s a hostage situation.

The eval set doesn’t need to be sophisticated. Real inputs, expected outputs, and a grading function, sometimes exact match, sometimes a rubric scored by a model, sometimes a human spot-check on a sample. What matters is that it exists and runs automatically.

2. The happy path was the only path

Prototypes get built against clean inputs, because that’s what’s at hand. Production has the other kind.

The PDF that’s a scan of a photocopy. The API that returns a 200 with an empty body. The record with a null where the schema promised a string. The input in a language nobody scoped for. The document that runs four hundred pages when everything was tested against ten.

Each of these is individually trivial to handle. Collectively they’re most of the work, and skipping them is why the gap between demo and production gets measured in months rather than weeks.

The fix is unglamorous. Every tool call gets explicit failure handling. Every parse gets a schema validation with a defined behaviour on failure. Every assumption about input shape gets checked instead of trusted. The agent should have a defined response to “I could not read this,” and that response should not be to hallucinate something plausible instead.

3. Cost was never modelled

Agent loops multiply token usage in ways that are genuinely hard to intuit. A task that takes eight tool-calling iterations, each carrying the full accumulated context, can cost fifty times what a single completion costs. That’s fine at ten runs a day. At ten thousand it’s a budget crisis, and it tends to arrive as a surprise.

Three things prevent it. Hard spend ceilings per run, enforced in code rather than a spreadsheet projection. Context management, summarising or truncating history instead of letting every iteration carry everything that came before. And model tiering: most steps in a real agent loop are classification, routing, or extraction, and those don’t need a frontier model. Reserve the expensive model for the steps that genuinely require reasoning, and you can often cut cost by an order of magnitude with no measurable loss in quality.

4. Failures were silent

A traditional service that breaks throws an error. An agent that breaks returns a confident, well-formatted, entirely wrong answer, with nothing in the response that looks like a failure.

This is the most dangerous property of these systems, and it’s why observability for agents is a different discipline. Full trace persistence matters: every tool call, argument, and result for every run, retained long enough to investigate a complaint. Confidence signals need to be surfaced, not swallowed. When the agent is uncertain, that uncertainty has to reach the interface instead of getting smoothed away by fluent prose. Anomaly detection on distributions catches what a single-run check won’t; a sudden change in output length, tool-call frequency, or refusal rate is usually the first visible sign that something upstream changed. And sampled human review has to stay in place permanently, not just during rollout. A small percentage of runs, checked by someone who knows what right looks like.

5. Nobody defined what the agent was not allowed to decide

An agent with tools will use them. If one of those tools sends an email, it will send emails, including in situations where a person would have hesitated.

The fix is boring, and it works: irreversible actions sit behind gates. Sending, paying, publishing, deleting, anything with an external side effect requires either an explicit confidence threshold or a human approval step. Everything reversible can proceed freely.

The interesting design question isn’t whether to have gates but where to put them, and the answer comes from weighing the cost of a wrong action against the cost of a delay. High-cost, low-frequency actions gate cheaply. Low-cost, high-frequency actions shouldn’t gate at all, or you’ve simply built an expensive way to generate work for a human.

This gap deserves more depth than one gap in a list of five, and it gets it in AI agent guardrails, including the specific test for telling a real constraint from a system prompt asking nicely.

The pattern underneath all five

Every one of these failures comes from treating an agent like a feature instead of a distributed system that happens to have a model in it.

Agents are stateful, non-deterministic, network-dependent, and expensive per operation. Every engineering discipline that applies to that class of system, observability, retries, idempotency, circuit breakers, cost control, graceful degradation, applies here too. The novelty of the model layer distracts from the fact that the rest of it is well-understood engineering.

The teams that ship agents successfully aren’t the ones with the best prompts. They’re the ones who built the boring infrastructure first.

Rescuing a stalled project

If you have a prototype that won’t cross the line, a sequence that usually works: build the eval set from the failures you already have (every case where someone said “it got this wrong” is a test case, and you likely have more of these than you think), instrument before you change anything (you can’t improve what you can’t see, and you want a baseline first), then narrow the scope aggressively. Ship the twenty percent of cases the agent handles reliably, route the rest to a human, and widen from there. A narrow thing in production beats a broad thing in staging, permanently. Only then start improving quality, with numbers behind each change.

Most of the prototype’s domain logic survives this. What gets rebuilt is the execution layer around it, and that’s usually a few weeks of work rather than a restart.

We do this kind of rescue regularly. See how we approach agent work, or read about evaluating agents properly.

Tell us what is slowing you down

A short conversation is usually enough to tell whether this is a build, an automation, or something you should not do at all. We will tell you which.