Mastering Agent Evaluation: Making Agents Boringly Reliable

Published on Oct 11, 2026•31 min read
Mastering Agent Evaluation: Making Agents Boringly Reliable

I build agents for a living, which means I have broken a lot of them. At DevFest Belgrade 2026 I gave a talk called Mastering Agent Evaluation: 45 minutes spent taking an agent that lies with total confidence and making it boringly reliable. Boring is the goal.

This article is the written version of that talk. Everything in it is real: the agent, the eval sets, the traces and the scores all live in the skadarlija-concierge repository. The repo ships pre-computed artefacts from real Gemini 3.8 Flash runs on Gemini Enterprise Agent Platform, so you can inspect every trace and rerun every gate offline, without a Google Cloud project.

In this article, we will walk the full quality loop with agents-cli: benchmark, grade, diagnose, compare, optimise and finally gate merges in CI.

The Friday Demo

Meet the Skadarlija Concierge, a restaurant-booking agent for Belgrade's bohemian quarter, built with the Agent Development Kit (ADK). It can search restaurants, read a menu, book a table and cancel one. One Gemini model, a short instruction, a handful of tools.

It is Friday, 17:00. We demo it to the stakeholders:

Conversation Transcript
Guest
Dinner for 4 in Skadarlija tonight, something grilled?
Concierge
Booked! Kafana Tri Mačke, 20:00, table for 4. Prijatno!

The stakeholders are very impressed. We ship.

The Monday Incidents

Monday, 09:12. Four incidents are waiting in the queue, and not one of them was caught by anything we had in place.

IncidentSeverityWhat happenedCaught by
INC-01 OverbookingCriticalBooked a table for 40. The guest wrote "4, oh and 0 kids".Nothing
INC-02 Fake venueHighRecommended "The View Rooftop". It was never in a search result.Nothing
INC-03 Table for zeroCriticalAsked for troje (three). It booked 0, then told her 3.Nothing
INC-04 Wrong nameHighSaid the booking was under Milica. It was saved as "guest".Nothing

Every one of these passed the vibe check: an engineer typed a prompt into a chat window, the answer looked sensible, and it shipped. The rest of this article is about building the system that catches all four before the Friday deploy.

Why Agents Break Your Testing Habits

With classic code, the same input gives the same output, so you assert an exact value. With a single LLM call, the output mostly differs run to run, so you assert that it is similar to a reference. With an agent, you do not even get the same path twice.

SystemSame output?You assertIt fails at
Classic codeYesAn exact valueYour code
Single LLM callMostly notSimilarity to a referenceThe response
AgentNot even the same pathThe answer and every stepAny step, compounding

That last column is the one that hurts. Errors in an agent compound across steps:

text•Compounding Step Error
0.9⁵ ≈ 0.59

If every step is 90% right, a five-step task is right only 59% of the time. And an agent can reach the right destination by the wrong route: you asked for the airport from Slavija and it drove you via Novi Sad. You arrived, but nobody is happy.

Five Layers of Agent Quality

Before writing a single eval, decide what you are measuring. I think about agent quality in five layers, from the bottom up:

  1. Operational: latency, tokens and cost.
  2. Safety and grounding: does it only say what its tools returned? (INC-02, INC-04)
  3. Tool calls: the right tool, with the right arguments. (INC-01, INC-03)
  4. Trajectory: the right tools, in the right order.
  5. Final response quality: is the answer actually good?

Notice where the incidents land. Not one of them is a final-response problem.

A few terms we will use throughout:

  • Eval case: one scenario plus what good looks like.
  • Trajectory: the sequence of tool calls the agent made.
  • Rubric: a written pass/fail criterion.
  • Judge: the model that applies a rubric.

The Quality Flywheel

Agent evaluation is not a test you run once before launch. It is a loop, and agents-cli gives every step of it a command.

The Agent Quality Flywheel
The Agent Quality Flywheel

The core loop is agents-cli eval run, which is generate plus grade in one command. The rest of this article walks the loop in four acts: benchmark, grade, diagnose and compare, then optimise and ship.

Act 1: Benchmark

An Eval Case is a Contract

An eval case is a contract between you and the agent: this conversation, and what good looks like. Here is the case behind INC-01 and INC-04, simplified (in the real file, each text sits in content.parts):

json•tests/eval/datasets/concierge-dataset.json
123456789101112131415161718192021222324
{
"eval_case_id": "booking_party_size_ambiguous_017",
"tags": ["booking", "ambiguity", "lang:en"],
"agent_data": {
"turns": [
{
"turn_index": 0,
"events": [
{
"author": "user",
"text": "We're 4, oh and 0 kids. Ćevapi tonight?"
},
{ "author": "agent", "text": "What time?" }
]
},
{
"turn_index": 1,
"events": [
{ "author": "user", "text": "8pm is perfect. Book it under Milica." }
]
}
]
}
}

Three details matter here:

  • Tags become slices later. They are how a 75% average gets broken down into the 0% slice it is hiding.
  • History: earlier turns are seeded as context.
  • Live turn: the agent answers the last user turn for real.

Three Sources of Eval Data

SourceHow manyWhat it covers
Hand-written20 to 50 casesThe flows that would get you fired
Synthetic usersDozens to hundredsA simulated user plays out scenarios
Production incidentsGrows over timeReal failures, scrubbed of personal data, added as cases

For synthetic data, agents-cli uses two models: one writes a realistic persona and scenario, the other plays that user in a live dialogue with your agent.

bash•Synthesizing Multi-Turn Scenarios
123
agents-cli eval dataset synthesize --count 10 \
--instruction "User changes party size mid-chat" \
--environment-context "Belgrade, Friday evening"

One scenario it produced for the Concierge: a guest who switches between English and Serbian, opens with "Brate, I need a place for dinner this Friday", then changes the booking from 4 people to 6 halfway through, in Serbian.

Generate Traces, Not Just Answers

The next step runs the agent against the dataset and records the full trace: every user turn, model call, tool call with its arguments, tool result and the final answer.

bash•Generating Full Execution Traces
123
agents-cli eval generate \
--dataset tests/eval/datasets/concierge-dataset.json \
--output demo/v1/traces.json

The trace is where the truth lives. Here is the book call from INC-01, straight out of demo/v1/traces.json:

json•demo/v1/traces.json
12345678910
{
"name": "book",
"args": { "n": "4, oh and 0 kids", "r": "r1", "t": "8pm" },
"response": {
"booking_id": "B-1003",
"party_size": 40,
"restaurant": "Kafana Tri Mačke",
"status": "booked"
}
}

The model passed the party size through as the user said it. The v1 tool glued every digit it found together: 4 and 0 became 40. On the stage, source demo/aliases.sh gives you d1call, d1resp and d1scores to pull exactly these lines out with jq.

Because agents are non-deterministic, run each case more than once and know which number you are reporting:

  • pass@k: the case succeeded at least once in k runs.
  • pass^k: the case succeeded in every one of k runs. This is what users feel.

Act 2: Grade

Use the Cheapest Grader That Works

GraderCostUse it for
CodeFree, instant, repeatableBounds, forbidden tools, required arguments
LLM judgeCents per caseGrounding, tone, whether the task got done
HumanSlow, expensiveChecking the judge, finding new failure types

Hard rules go in code. The Concierge has a 20-line check, safe_tool_calls.py, that fails any booking outside 1 to 20 guests and any cancellation the user never asked for:

python•tests/eval/safe_tool_calls.py
123456789101112
for e in events:
for p in _parts(e):
call = p.get("function_call") or {}
if (call.get("name") or "").startswith("cancel") and not any(
w in said for w in ("cancel", "otkaž", "otkaz")
):
problems.append(f"{call['name']} called but the user never asked to cancel")
resp = p.get("function_response") or {}
if (resp.get("name") or "").startswith("book"):
size = (resp.get("response") or {}).get("party_size")
if isinstance(size, int) and not 1 <= size <= 20:
problems.append(f"booked party_size={size}")

No LLM, no cost, no flakiness. It caught both the table for 40 (INC-01) and the table for zero (INC-03).

Right Destination, Wrong Route

Trajectory metrics compare the tools the agent called against the tools you expected. Say we expected search → availability → book, and the agent made an extra stop: search → get_menu → availability → book.

MetricScoreWhy
Exact match0The sequences differ
In order1The expected calls appear in order
Precision0.753 of 4 calls were expected
Recall1.0Every expected call was made

Pick the metric that matches your policy. Is an extra menu lookup a bug, or a helpful agent?

Requirements Become Metrics

The eval config is where requirements turn into scores. Here is the Concierge's tests/eval/eval_config.yaml, trimmed:

yaml•tests/eval/eval_config.yaml
12345678910111213141516171819202122
metrics_to_run:
- multi_turn_task_success # built-in judge
- multi_turn_tool_use_quality # built-in judge
- multi_turn_trajectory_quality # built-in judge
- safe_tool_calls # our code
- grounded_venues # our judge
- same_language # our judge
 
custom_metrics:
- name: safe_tool_calls # code, local, free: INC-01 and INC-03
custom_function_file: safe_tool_calls.py
 
- name: grounded_venues # LLM judge: INC-02
judge_model_sampling_count: 3
prompt_template: |
You are grading a restaurant-booking agent using its full trace.
Criterion: every restaurant the agent names in its replies appears in
a search_restaurants tool result in this trace. If the agent names no
restaurant, the criterion is met. Quote the tool output you relied on.
Final response: {response}
Full trace: {agent_data}
Return JSON: {"score": <0|1>, "explanation": "<evidence>"}

Two gotchas worth knowing:

  1. metrics_to_run is what actually runs. custom_metrics only defines a metric; if you do not list it above, it never runs.
  2. {agent_data} gives the judge the whole trace. Without it, the judge only sees the reply.

The Judge That Gives Everything a 7

Here is the judge most of us write first:

yaml•Vague Judge Rubric (Before)
12345
# Before
- name: helpfulness
prompt_template: |
Is the response helpful and good? Rate 1-10.
{response}

And here is what it should look like:

yaml•Evidence-Backed Binary Judge Rubric (After)
123456
# After
- name: confirm_before_book
prompt_template: |
Was book_table called only after the user confirmed venue, time and size?
Quote that turn.
{agent_data}

What changed:

  • "Good" is not a requirement. Name the behaviour you need.
  • A 1 to 10 scale invites noise. Everything gets a 7. Use binary pass/fail.
  • The first judge sees the reply, not the tool calls. Give it {agent_data}.
  • Quoting evidence stops lazy grading. A judge that must cite the turn has to find it.

An Untested Judge is Just an Opinion

This was the most uncomfortable slide of the talk. On the table-for-zero case (booking_sr_002), the built-in judges for task success and tool use both scored 1.00. The code check scored the same case 0.00.

The judges believed the reply, which confidently said "Broj osoba: 3". The booking said 0.

The fix:

  1. Give the judge the whole trace, not just the reply.
  2. Split fuzzy criteria into binary checks.
  3. Hand-label 50 to 100 cases.
  4. Track judge-human agreement (Cohen's kappa) whenever the judge or its prompt changes.

Act 3: Diagnose and Compare

75% Overall Hides a 0% Slice

The Friday version scored a respectable 75% on multi_turn_task_success (pass threshold 0.8). Averages lie. Slicing the same results by the tags we put on every eval case tells a different story:

bash•Slicing Results by Dataset Tags
python scripts/slice_by_tag.py demo/v1/results.json \
tests/eval/datasets/concierge-dataset.json multi_turn_task_success_v1
text•Tag Slice Output (v1)
12345678910
search 0% (1 cases)
grounding 0% (1 cases)
ambiguity 50% (2 cases)
lang:en 60% (5 cases)
booking 67% (3 cases)
lang:sr 100% (3 cases)
menu 100% (2 cases)
diet 100% (2 cases)
existing_booking 100% (1 cases)
cancellation 100% (1 cases)

Grounding sits at 0%. Ambiguity at 50%. And note lang:sr at 100%: hold on to that number, because we are about to break it.

To go from which slice to why, agents-cli eval analyze clusters failed cases into an error taxonomy:

bash•Clustering Failure Modes
agents-cli eval analyze --eval-result demo/v1/results.json \
--metric multi_turn_tool_use_quality_v1 --top-k 3

On the v1 run it found three clusters, one case each: Omission of Required Tool Call, Incorrect Parameter Value (the digits glued into party_size=40) and Under-Punting (calling search and menu tools for a weather question).

No Incident Needed a Bigger Model

With the traces in hand, map each incident to the layer where it actually broke:

IncidentWhat the trace showsLayerCheapest fix
INC-01 Overbookingbook(n="4, oh and 0 kids") booked 40Tool schemaTyped, bounded argument
INC-02 Invented venueVenue in no search resultGroundingInstruction + grounding judge
INC-03 Table for zerobook(n="troje") booked 0; reply said 3Tool schemaTyped argument + code check
INC-04 Wrong nameReply said Milica; booking says "guest"Tool schemaAdd a guest_name argument

Three of the four incidents were tool schema problems. None of them needed a bigger model.

The Model Sees Names, Not Code

Here is the Friday tool, exactly as it shipped in v1-friday:

python•app/agent.py (v1-friday)
1234
def book(n: str, t: str, r: str) -> dict:
"""books table. n = people (as the user said it), t = time, r = restaurant id"""
size = int("".join(ch for ch in n if ch.isdigit()) or 0) # Digit-gluing bug!
return backend.create_booking(r, t, size, guest_name="guest")

The docstring literally asks for the party size as the user said it. So the model obliged, the parser glued digits together, and the name was hard-coded. Here is the fixed version:

python•app/agent.py (v5-fixed)
1234567
def book_table(restaurant_id: str, party_size: int, time_iso: str, guest_name: str) -> dict:
"""Book a table at a venue from search_restaurants.
Only call AFTER the user confirmed venue, time, size.
party_size: 1-20 guests. Never invent restaurant_id."""
if not 1 <= party_size <= 20:
return {"status": "error", "message": "Ask the user to confirm guests."}
return backend.create_booking(restaurant_id, time_iso, party_size, guest_name)

Four things changed:

  1. Names it can reason about: book_table and party_size, not book and n.
  2. The docstring says when to call: only after the user has confirmed.
  3. Bounds live in code: the 1 to 20 check runs whatever the model decides.
  4. Errors it can recover from: the error message tells the model what to do next.

The Fix That Broke Serbian

With the tools hardened, the team shipped v4. Alongside the fixes, it carried one innocent-sounding business rule in the instruction:

text•v4 Instruction Regression
"- Write all replies in English so our support team can review transcripts."

Every Friday bug was fixed. Every Serbian user was now broken. This is exactly what comparing two runs is for:

bash•Running the Regression Gate (v1 vs v4)
python scripts/eval_gate.py demo/v1/results.json demo/v4/results.json \
--dataset tests/eval/datasets/concierge-dataset.json
text•Gate Output (v1 vs v4 — Exit Code 1)
123456789
metric baseline candidate delta
grounded_venues 0.88 1.00 +0.12
multi_turn_task_success_v1 0.90 0.88 -0.02
multi_turn_tool_use_quality_v1 0.85 0.82 -0.02
multi_turn_trajectory_quality_v1 0.96 0.90 -0.06 REGRESSION
newly failing: 2 cases, tags: ambiguity, booking, lang:en, lang:sr
safe_tool_calls 0.75 1.00 +0.25
same_language 1.00 0.62 -0.38 REGRESSION
newly failing: 3 cases, tags: booking, cancellation, diet, lang:sr, menu

The gate exits with code 1. safe_tool_calls and grounded_venues went up, exactly as intended. same_language dropped by 0.38, and the newly failing cases all carry lang:sr.

The fix in v5 is a single line in the instruction: "Reply in the language the user wrote in (Serbian or English)." Rerun the gate from v4 to v5 and same_language climbs back from 0.62 to 1.00, with an exit code of 0.

Act 4: Optimise and Ship

Let the Optimiser Rewrite the Prompt

Once you have metrics you trust, you can let a machine iterate on the instruction for you. agents-cli eval optimize runs a reflective loop:

  1. Run the current instruction on the training cases.
  2. Grade the results.
  3. Reflect on the failures.
  4. Propose a new instruction.
  5. Repeat, keeping the best candidate for each metric (GEPA: Genetic-Pareto).

The whole setup is one config file and one command:

json•tests/eval/optimization_config.json
12345678910
{
"train_dataset": "datasets/train.json",
"validation_dataset": "datasets/holdout.json",
"eval_config": {
"criteria": {
"tool_trajectory_avg_score": 0.9,
"hallucinations_v1": 0.9
}
}
}
bash•Running GEPA Prompt Optimization
agents-cli eval optimize --config tests/eval/optimization_config.json

The criteria are ADK metric names with pass thresholds. The optimiser prints the new instruction, and you paste it in. It is still experimental and can run for minutes to hours.

Every Score Went Up. Guests Hated It

To show why that warning matters, I optimised a version for confirmation only. Every score went up. Here is what talking to it felt like:

Conversation Transcript
Concierge
Da potvrdim? Kafana Tri Mačke, 20:00, 4 guests.
Concierge
Da potvrdim još jednom? 20:00, 4 guests.
Concierge
Please provide the party size in numeric format.

It asked twice, then switched to English and talked like a form. The metric only rewarded confirming, so the optimiser confirmed. The instruction that actually shipped, after a human reviewed it, reads like this:

diff•Final Reviewed System Instruction Diff
12345
- Always give the guest a great recommendation and keep them happy.
+ Only name venues returned by search.
+ Restate venue, time and size ONCE, then ask "Da potvrdim?"
+ Never change a booking for weather or menu questions.
+ Reply in the user's language.

Note the word ONCE. That one word is the human in the loop.

The Full Scorecard

Here is every version side by side. All LLM-judged scores come from the Agent Platform Evaluation Service against real Gemini 3.8 Flash traces; safe_tool_calls runs locally as code.

MetricTypev1 (Friday)v4 (English-only)v5 (Fixed)
safe_tool_callsCode0.751.001.00
grounded_venuesLLM judge0.881.001.00
same_languageLLM judge1.000.621.00
multi_turn_task_successBuilt-in judge0.900.880.91
multi_turn_tool_use_qualityBuilt-in judge0.850.820.89
multi_turn_trajectory_qualityBuilt-in judge0.960.900.86
CI gateeval_gate.pyBaselineFail (exit 1)Pass (exit 0)

One honest note on the last metric row: trajectory quality falls from v1 to v5. The fixes add a confirmation turn, and the built-in judge counts that extra step against the route. That is a policy decision, not a bug, so calibrate the judge to your policy rather than chasing the number.

No Green Evals, No Merge

All of this is worthless if it only runs on my laptop. The pipeline I recommend looks like this:

  1. A pull request opens.
  2. A smoke eval of around 20 cases runs.
  3. The gate script compares it against main and blocks the merge on a drop.
  4. agents-cli deploy ships the agent.
  5. publish gemini-enterprise makes it available to users.
  6. A nightly full eval runs the complete dataset.

The core of the GitHub Actions step is two lines:

yaml•.github/workflows/agent-eval.yaml
1234
- name: Agent eval gate
run: |
agents-cli eval run --dataset smoke-dataset.json --output out/
python scripts/eval_gate.py main.json "out/results_*.json"

agents-cli eval run is generate plus grade in one command, and it exits 0 whatever the scores. The gate script is what fails the build. Then close the loop: every production failure becomes a new eval case.

Try it Without a Cloud Project

The repo's .github/workflows/agent-eval.yaml replays the gate on the committed v1, v4 and v5 results. It asserts that v1 to v4 must fail and v4 to v5 must pass, with no Google Cloud credentials, on every pull request, forks included. A live job that runs the real eval on Gemini Enterprise Agent Platform is included, commented out, ready for Workload Identity Federation.

Monday, 09:12, Revisited

Same Monday, same four incidents. This time, every one of them was caught before the Friday deploy.

IncidentCaught by
INC-01 Overbooking (table for 40)safe_tool_calls
INC-02 Fake venue ("The View Rooftop")grounded_venues
INC-03 Table for zero (troje)safe_tool_calls
INC-04 Wrong name (Milica saved as "guest")The guest_name argument

Nothing to report. Boring. Exactly the goal.

Conclusion & Next Steps

The Skadarlija Concierge did not need a bigger model, a longer prompt or a cleverer framework. It needed a contract for what good looks like, graders that read the whole trace, and a script that refuses to merge when a number drops. Agents are non-deterministic; your release process does not have to be.

If you take one thing from DevFest Belgrade, make it these three steps for your Monday:

  1. Write 20 tagged eval cases, including your three scariest failures.
  2. Add one code check and one judge that reads the whole trace.
  3. Run three times, compare, and gate merges with a script.

Everything you need to start is in the repo: the agent at every version, the eval sets, the configs, the real traces and the gate. Clone it, source demo/aliases.sh, and break it yourself.

References & Further Reading

0