LLM Agent Evaluation With the RAGAS Framework
RAGAS adds agent-specific metrics to catch multi-step failures that final-answer scores miss.

This article looks at RAGAS, the evaluation framework built for RAG pipelines, and at how it now scores LLM agents on goal completion, tool use, and multi-turn behavior. Its angle: RAGAS gives teams a real measurement system for agent quality, but knowing which metrics apply at which layer, and what none of them can see, is what separates a useful score from a false sense of security. Security teams asking how to find unauthorized AI activity on their network should read the final section first: evaluation scores tell you if your sanctioned agent improved release over release, but they say nothing about the agents nobody sanctioned.
Why RAGAS exists
BLEU and ROUGE measure how much a generated string overlaps, word for word, with a reference answer. That approach works fine for machine translation, where a reference sentence is close to the only right answer. If you use it for retrieval-augmented generation, it falls apart: a correct answer can be phrased a dozen different ways, so n-gram overlap against one fixed reference string tells you nothing about whether the model pulled its facts from the documents it retrieved. A chatbot can score well on ROUGE when it states something the retrieved context never supported, and it can score poorly when it gives a fully grounded answer in different words.
RAGAS replaced that string-matching approach with an LLM-as-judge method. It checks semantic grounding, retrieval relevance, and generation faithfulness. It can tell whether an answer is both correct and supported by what was retrieved, rather than lexically close to some reference someone wrote down in advance. The framework came out of research published at EACL 2024, and because it works without reference answers in most cases, you can run evaluation without having to hand-label ground truth for every example in a dataset. So a team can run evaluation all the time instead of once a quarter, because that cost barrier is gone.
The four canonical metrics split along RAG's two points of failure: retrieval and generation. If the generated answer isn't grounded in the retrieved documents, faithfulness catches it. Answer relevancy checks whether the answer addresses the question you asked, so it catches responses that are true but beside the point. Context precision checks whether the retrieved chunks are focused and useful, surfacing noisy retrieval. If the retrieved context didn't hold what you needed to answer the question, context recall catches it. Splitting the score this way, retrieval separate from generation, gives an engineer a clear diagnosis of which half of the pipeline to fix. That same separation turns out to matter even more once the system in question is not a single retrieve-then-answer loop but an agent running multiple steps on its own.
How RAG evaluation maps onto agent evaluation
RAG pipelines and LLM agents share the same basic failure shape: both can produce an answer that sounds confident and turns out wrong because retrieval missed something, the model ignored the context it had, or it hallucinated. That shared failure structure is the reason RAG evaluation concepts carry over to agents at all, rather than requiring an entirely new vocabulary from scratch.
But an agent is not just a RAG pipeline with extra steps bolted on. A March 2026 SoK paper on Agentic RAG lays out the shift precisely: agentic systems add planning mechanisms, retrieval orchestration, memory paradigms, and tool-invocation behavior on top of the basic retrieve-then-generate loop. Each of those additions opens its own failure surface, and none of them is visible to a metric that only looks at the final answer. The paper names risks that simply don't exist in single-turn RAG: hallucinations that compound across steps instead of staying contained to one response, memory poisoning, retrieval misalignment across a multi-turn session, and tool-execution vulnerabilities that cascade once one bad call feeds into the next.
Picture an agent tasked with booking a refund. It retrieves the correct policy document, generates a faithful summary of it, and then calls the refund API with the wrong order ID because an earlier turn in the conversation polluted its working memory. Faithfulness and context precision both score fine on that trace. The agent still fails the user. That is the mapping breaking down: evaluation that stops at the final answer cannot see a failure that happened three steps upstream in planning or tool selection. The same SoK paper argues for trajectory-level assessment: you score the reasoning and retrieval behavior across the whole run, not just the last thing the agent said. Teams building on RAGAS need to know which of its metrics carry over cleanly to this setting and which ones require something built specifically for agent behavior.
What AgentGoalAccuracy, ToolCallAccuracy, and TopicAdherence measure
RAGAS has added agent-specific metrics on top of its four canonical RAG metrics, but these additions are newer and less battle-tested than the RAG core. If you know what each one checks, and what it doesn't, you know whether a score tells you something true or just something comforting.
AgentGoalAccuracy checks whether the agent got the user where they wanted to go. It's an end-state check: did the task get done. It says nothing about how the agent got there. An agent that flails through three wrong tool calls before stumbling into the right answer scores identically, on this metric, to one that planned the whole route correctly from the start. ToolCallAccuracy is built to catch that gap: it checks whether the agent picked the right tool and called it with the right arguments, the step-level detail AgentGoalAccuracy can't see. A tool invoked with a wrong parameter can still return a response that happens to let the agent limp toward a correct answer anyway, and ToolCallAccuracy is the only one of the two that would flag it.
TopicAdherence checks whether the agent stayed inside the topic domain it was built for, which matters most for agents with a defined persona or a compliance boundary where drifting into unrelated territory is itself the failure, regardless of whether the final answer was accurate. RAGAS can also score a full multi-turn conversation rather than turn by turn, so an agent that answers each question correctly but never remembers what the user said two turns earlier gets flagged in the aggregate score.
What none of this covers natively: the quality of how the agent broke a task down into subtasks, whether it knew when to stop looping, and latency across the full range of percentiles, from typical cases to the worst-case tail. Any team running complex tool orchestration in production needs to watch for those, so it has to run RAGAS scores alongside that kind of monitoring, not in place of it.
RAG metrics in an agent evaluation suite
None of that should read as an argument that faithfulness, context precision, and answer relevancy stop mattering once an agent enters the picture. Most LLM agents still have retrieval steps somewhere in their trace, and those steps can fail just like a standalone RAG pipeline does. So the four canonical RAGAS metrics stay just as diagnostic inside an agent as they were before agents existed.
The Agentic RAG SoK paper makes the structural case directly: agentic systems stack planning and tool use on top of the retrieve-generate loop, so the old failure modes from that loop persist alongside the new ones and compound with them. A retrieval step buried inside an agent's trace can be scored for context precision and faithfulness on its own, independent of whether the agent eventually reached its goal. That independence is the whole point. An agent trace can show a high faithfulness score sitting right on top of a low context-precision score, and that combination is the signature of an agent that sounds authoritative while hallucinating, because it faithfully summarized documents that were themselves the wrong documents to retrieve. A goal-level score alone would never catch that; it would just show the agent eventually failed, with no clue why.
The SoK paper's push toward trajectory-level assessment only works if each step in that trajectory carries its own score. Span-attached evaluation scores every retrieval call and every generation step in an agent trace on its own, and that is what makes trajectory-level assessment possible, not just theoretical. The practical rule follows directly from that: run the four canonical RAGAS metrics on every retrieval-and-generation span inside an agent trace, and run AgentGoalAccuracy and ToolCallAccuracy at the level of the full trajectory. Teams shouldn't treat RAG metrics and agent metrics as competing options. They measure different layers of the same failure.
Running RAGAS in practice: building a harness that acts on scores
A score computed once, on a handful of hand-picked prompts, is not evaluation. Evaluation is a harness: a repeatable process that runs against a labeled dataset, blocks deployment when scores drop below a line the team drew before shipping, and turns production failures into the next round of test cases.
RAGAS itself is a lightweight Python library with little setup overhead. It works with the major LLM providers, OpenAI, Anthropic Claude, Google Gemini, and plugs into LangChain and LlamaIndex, so a team already building on those tools adds evaluation without migrating to a new platform. The cost appears elsewhere: RAGAS itself doesn't charge, but every metric it computes runs an LLM call as the judge, and teams running high-volume evaluation, especially with a top-tier model doing the judging, need to budget for that API spend the same way they'd budget for any other production LLM cost.
The harness design is where teams separate themselves into two camps: those shipping on vibes and those shipping with actual confidence in a release. Four decisions matter most. The dataset has to represent real production traffic, not a curated set of prompts picked because the system already handles them well. You need to set thresholds for faithfulness and context relevance before the first release ships, not adjust them after something breaks and someone needs an excuse. Scoring has to attach to individual spans, so if faithfulness drops in one retrieval step, you see it as a localized problem instead of watching it get averaged away inside an aggregate number that looks fine. And when a production trace's score disagrees with what the team expected, the team turns that trace into a new test case, closing the gap between test-set performance and what actually happens in front of users.
RAGAS has no managed commercial tier of its own. It's a pure open-source library, so if you want a hosted evaluation platform at scale, you typically pair it with something like Confident AI, built on DeepEval, to get the orchestration layer RAGAS doesn't provide on its own.
Where RAGAS fits in the broader evaluation framework landscape
RAGAS is the right tool to reach for when the core question is whether a RAG pipeline, or the retrieval steps inside an agent, are grounded and relevant. It's purpose-built for that question and produces scores a team can actually interpret and act on. Teams running more complex, multi-component agent stacks with mature CI/CD pipelines need to weigh it against frameworks that go further on tracing and metric breadth.
A 2026 comparison from DeepEval maps five frameworks across metric coverage, dataset support, agent tracing, dataset generation, and multi-turn simulation, the axes that matter most for agent work. RAGAS comes out RAG-first, with agent metrics added on but only partial support for datasets and multi-turn scenarios, and tracing limited to a basic OpenTelemetry integration. Its strength is how fast a team can get it running: RAGAS is in use at AWS, Microsoft, Databricks, and Moody's. DeepEval covers agents, RAG, chat, safety, and multimodal use cases more broadly, integrates natively into CI/CD through pytest, supports both reference-free and reference-based evaluation, and runs its commercial layer through Confident AI. TruLens pioneered the RAG Triad, context relevance, groundedness, and answer relevance, built strong OpenTelemetry-based tracing, and now sits under Snowflake, with Walmart Global Tech, Cisco, J.P. Morgan Chase, Equinix, VMware by Broadcom, Hitachi Digital Services, Thomson Reuters, phData, and HID Global among its users. Arize Phoenix is open-source, compatible with OpenTelemetry, and strong on versioned datasets and run comparisons, covering both RAG and trace-based agent evaluation.
The practical logic: pick RAGAS when retrieval and grounding quality is the main problem. Choose DeepEval or MLflow for broader agent metric coverage and a tighter CI/CD loop. Pick TruLens if the stack already runs on Snowflake or OpenTelemetry. Pick Arize Phoenix for open-source tracing paired with dataset versioning. Whichever framework a team lands on, the traces and quality scores it produces become the raw material for something outside the evaluation layer entirely: governance. Speakeasy, an enterprise AI control plane, is one example of infrastructure built to take that evaluation output and turn it into enforceable policy rather than a report nobody reads after the sprint ends.
What evaluation scores don't tell you
An evaluation harness answers one question well: is this version of the agent better than the last one. It does not answer which agents are running in production right now, who signed off on deploying them, what systems and data they can reach, or whether a given tool call complied with access policy. Those are governance questions, and no evaluation metric answers them, however well it's designed.
The Agentic RAG SoK paper's own list of risks, compounding hallucination propagation, memory poisoning, cascading tool-execution vulnerabilities, describes failures that a good evaluation suite can catch during testing but cannot stop from happening at runtime without some kind of enforcement sitting in the path. The risks that come with autonomous loops need more than a test suite running in CI. They require real-time detection built into the agent layer itself, which is a different piece of infrastructure than the one that produces a faithfulness score.
That gap turns sharp at enterprise scale. A team can run rigorous RAGAS-gated evaluation on every agent it officially built and shipped, and still have zero visibility into the agents employees spun up on their own, wired directly into internal APIs, SaaS tools, and sensitive data stores without anyone in IT reviewing the connection. Evaluation frameworks were never built to answer that question, because sanctioned and unsanctioned agents both sound just as confident, and a test suite only sees the agents you pointed it at.
Closing that gap takes a governed layer sitting between agents and whatever systems they call: one that enforces identity, so the system knows who the agent is acting as; access policy, so it knows which tools and data that agent can actually reach; and audit, a record of what it did, when, and on whose authority. That's the same architectural logic that made API gateways a requirement for REST APIs a decade ago, now applied to MCP servers and agentic skills. Speakeasy intercepts PII exposure, prompt injection attempts, and credential or secret leakage at that control-plane layer, so it catches those failures as they happen in production instead of waiting for the next evaluation run to flag them after the fact. If organizations run agents at real scale, they need continuous visibility into how those agents behave across teams, tools, and data access patterns, not just a score attached to the last version that shipped. Evaluation produces the traces and the quality signals; a control plane is what turns those signals into policy that actually gets enforced, real-time detection of the failures evaluation can only describe after the fact, and an audit record built to survive a compliance review. The two layers do different jobs, and an enterprise security team needs both running at once, not a choice between them.


