LLM Evaluation Metrics for Agentic Workflows in Production
Evaluating AI agents requires measuring their decisions, not just their final answers.

Cursor, the AI coding assistant, became a small case study in why agent evaluation is harder than it looks when it confidently deleted a user's production database during an autonomous run and kept working as though nothing had happened. The task had already failed several steps earlier, and nothing in a standard answer-quality check would have caught it. That gap between "the output reads well" and "the agent did the right thing" is the subject of this piece: what it takes to evaluate LLM agents once they stop answering questions and start taking actions. A different metric set, built around trajectories, tool calls, and plans rather than final text alone, is what's needed.
Why standard evaluation metrics fail when agents act rather than answer
Retrieval-augmented generation gave the field a clean, closed loop: retrieve some documents, generate an answer, score the answer against the documents and the question. An agent breaks that loop before it even gets going. It plans, calls tools, hands sub-tasks to other agents, and in many cases commits an action, a write to a database, a message sent, a charge processed, before a human ever looks at what it produced. The final answer in an agentic run is just the last thing left behind. It says nothing about the eleven decisions made to get there.
This creates a specific failure mode: the compounding error. Debugging an agent by staring at its final output is like trying to find a typo by reading the printed book: by the time you see the problem, the actual error happened somewhere back in the manuscript.
Four properties of agentic systems make this evaluation problem structural. Errors compound across steps. Success and quality are not the same measurement, and treating them as one is exactly how a database gets deleted while the log looks clean.
The three evaluation layers every agentic system requires
Agent evaluation needs three layers, because each layer answers a question the others can't. Skipping one leaves a whole category of failure invisible, because the remaining metrics were never looking in that direction.
Layer one is end-to-end evaluation: did the task actually get done? This is the layer most teams already measure, and it is necessary, but it's a blunter instrument than it looks.
Layer two is trajectory evaluation: was the path the agent took efficient and sound? This layer scores the sequence, the tool calls, the reasoning transitions, the replanning cycles, that produced the outcome, independent of whether the outcome was correct. A trajectory can be expensive, looping, or unsafe while producing a final answer that reads perfectly fine.
Layer three is per-turn evaluation: what did each individual turn mean as it happened? A single turn might carry a jailbreak attempt, a leaked system prompt, a policy violation, or a user showing clear frustration, and none of that appears in a status code, a latency number, or a token count. A turn's meaning is visible only when something is actually looking at that turn.
The three layers map onto the rest of this piece. Task completion and success rate cover layer one. The trap most teams fall into is treating layer one as the whole job: a pass/fail number stays green while the trajectory underneath it is a mess and several turns quietly drift off policy.
Task completion and success rate: whether the agent finished the job
Task completion is the metric everyone starts with, and for good reason: it's the one that maps directly to whether the business got what it paid for. It requires checking whether the thing the agent was supposed to do in the world actually happened.
An agent can write a message confirming a refund was issued, a ticket was closed, a record was updated, with none of those actions having occurred, or having occurred incorrectly. The database disagrees. This is the gap that makes naive text-matching evaluation dangerous for agentic workflows: a grammatically sound, confident-sounding final answer tells you nothing about whether the underlying system state changed the way it should have.
Success rate, measured this way, is the metric that executives and stakeholders can read directly: it answers whether the agent got the job done. It belongs at the top of any dashboard. It also tells you almost nothing about why an agent failed, only that it did. Trajectory and component-level metrics have to sit alongside it for anyone trying to fix the system to use it. A passing success rate can sit directly on top of a trajectory that looped twice, called the wrong tool once, and got lucky on the third try. The next layer down is where that becomes visible.
Tool-call accuracy (scoring each tool selection and argument at the span level)
Tool-call accuracy checks one thing at a time: for a given step, did the agent pick the right tool, and did it pass that tool the right arguments? Scoring this requires looking at each tool call on its own, against a reference or a schema, rather than trying to infer backward from whether the final answer happened to be correct.
Two sub-dimensions need separate scores. Tool correctness asks whether the agent chose the right tool for the step it was on, given the task and what it had seen so far. A system can get one right and the other wrong, call the correct API with a malformed date string, or call the wrong API with perfectly formed arguments, and a metric that only looks at the top-level outcome will miss both.
Most evaluation tooling logs which tools got called and in what order. Visualizing it as a graph makes span-level scoring practical.
The judge mechanism should match the question being asked. Argument correctness is softer, especially when the space of valid arguments isn't fully enumerable, and that's where LLM-as-judge earns its place.
The stakes here are higher than a single wrong step. A wrong tool call at step two doesn't just fail step two, it hands a corrupted result to every step that follows, and tracing that failure back to its origin is hard. Root cause accuracy without automated error propagation tracking runs at 38%, according to the Integrating Evaluation into AI Workflows: 2026 Guide. That's why tool-call accuracy sits at the component layer but functions as one of the strongest predictors of whether the end-to-end task will succeed at all.
Step efficiency and loop detection (catching waste and pathological behavior inside a correct run)
An agent that passes the success-rate check can still be a mess underneath. Step efficiency measures how many steps a run took relative to the minimum the task actually required, and loop detection flags repeated identical actions, the same tool called with the same arguments more than once in sequence. Neither of these metrics cares whether the final answer was right. Both catch problems that a success-rate check will never see.
A loop can sit inside a passing run. A loop that runs inside a passing session still costs three to five times the tokens and latency of a clean execution, and across a fleet of agents running thousands of sessions a day, that pattern becomes a budget line.
Repeated identical actions are the clearest signal that something has gone wrong, and they're detectable without any LLM judge at all: if the same tool gets called with the same arguments twice in a row, that's a loop, full stop on the detection logic even if the consequences deserve more than a shrug. A common ceiling in production systems is 25 turns per session, capping how many iterations an agent gets before the session is forced to resolve or fail. The actual fix is usually the prompt or the tool set, because the ceiling was never the problem, it was the alarm.
Trajectory match (comparing the agent's path against a reference to catch structural drift)
Step efficiency and loop detection catch waste. Trajectory match catches something different: whether the agent's overall path resembles a path that's known to be good, independent of how many steps it took. An agent can take a correct number of steps, avoid looping entirely, and still arrive at the right answer through a structurally different route than the one that's been validated for that task.
Two matching modes handle two different risk appetites. Strict match requires the agent's trajectory to replicate a reference trajectory step for step, which fits workflows where the path itself carries compliance weight, a regulated process where skipping or reordering steps is a problem even if the outcome is identical. AgentEvals' create_trajectory_match_evaluator supports four such modes, strict, unordered, subset, and superset, so a team can set the tolerance to match what their workflow actually needs rather than picking one rigid standard for every task.
Trajectory match earns its keep most clearly after a model swap or a prompt rewrite. The success-rate number can stay flat while the underlying path changes shape entirely, and a path that happens to work on familiar inputs may fall apart the first time it meets something slightly novel. That's the regression trajectory match is built to catch.
Plan quality and plan adherence (evaluating the reasoning that precedes action)
Everything covered so far measures what the agent did. Plan quality and plan adherence measure what it intended to do before it did anything, and for agents that plan explicitly before acting, that intention is often the strongest leading indicator available for how the rest of the run will go. A flawed plan produces flawed steps even when each step, in isolation, gets executed cleanly. Perfect execution of a bad plan is still a bad outcome.
These are two separate metrics, not one. Losing track of the original goal and drifting for no good reason is not, and the two can look identical from the outside unless the original plan was captured somewhere.
That capture requirement is easy to skip and expensive to skip. Scoring plan adherence means comparing mid-run decisions against the plan the agent actually started with. The plan has to exist as a recorded, traceable artifact from the start of the run rather than something reconstructed after the fact from the trajectory. Reconstruction after the fact is guesswork dressed up as measurement.
Telling valid replanning apart from goal drift is one of the harder judgment calls in this entire field, and LLM-as-judge is the only scorer that works at any real scale for it. That doesn't mean it works out of the box. Agentic systems need dedicated metrics for planning quality, step-level faithfulness, reasoning coherence, and completion across multi-step runs, because repurposing metrics built for single-shot retrieval and generation produces scores that look precise and mean very little.
Groundedness and reasoning coherence in agentic systems
None of this means retrieval-era metrics are worthless for agents. Some of them still do real work when they get applied to the right piece of the system.
Faithfulness still matters wherever an agent makes a claim based on something it retrieved: does the claim actually match what the retrieved content says? Answer relevancy still matters for any individual sub-task that produces a text response: is that response relevant to the prompt that sub-task was given?
The constraint that makes these useful rather than misleading is granularity. They have to run at the span level, scoped to one retrieval step or one generation step inside the trace, rather than applied to the agent's final output as a blanket score. Running faithfulness against a final answer that's the product of six tool calls and two sub-agent delegations conflates retrieval quality, planning quality, and tool-call quality into one number that explains none of them.
Hallucination frequency and retrieval precision are hard to pin down with a fixed threshold, so LLM-as-a-Judge frameworks have become the common way to automate them at scale. Reasoning coherence is the agentic cousin of faithfulness: does each intermediate decision actually follow from the task, the context the agent has seen, and the tool results it's accumulated so far? This one lives at the trajectory layer, not the component layer, judging the thread connecting decisions rather than any single step in isolation, unlike the other RAG holdovers.
Safety metrics that must run at the span level, not just on final output
Safety checks that only look at an agent's final output arrive too late to matter, because by the time something unsafe shows up in that final text, the agent has usually already taken the action behind it. A safety layer that only reviews the finished product is reviewing a decision that's already been executed.
The safety metrics that matter for agentic systems operate at the level of individual spans: single tool calls, single reasoning steps, single pieces of retrieved content. Prompt injection detection checks whether any turn, including content pulled in from an external source the agent doesn't control, contains instructions trying to hijack its behavior. Data exfiltration checks whether the agent sent internal context, system prompt content, or user data out to an external endpoint it had no reason to talk to.
The thing that makes all of this non-negotiable is irreversibility. Agents write to databases, call external APIs, and send emails before any human reviews what's happening, so a safety check sitting only at the final output is reviewing the scene after the fact. Prompt injection, unauthorized tool use, and data exfiltration are documented attack vectors against systems already running in production, and the span-level check is the only version of safety evaluation that can catch them before the action.


