agentevals
v0.10.0
shipped two days ago and removed its own WebSocket ingestion path in favor of plain OpenTelemetry. Send OTLP traces in, run an eval, and the score comes back out as a standard gen_ai.evaluation.result OTel log event, attached to the exact turn it scored. I ran both halves of that claim against a real, independent OTel Collector and watched the events land. Repo: themsquared/agentevals-otel-roundtrip
.
Disclosure: I work at Solo.io, which also builds agentgateway and kagent; agentevals-dev is a separate open-source project several Solo engineers contribute to. Everything below is the open-source CLI, and every claim links to the v0.10.0 release notes or docs.
The short answer
agentevals used to be a destination with its own ingestion protocol: a WebSocket stream, a bespoke session model, a SDK you wired into your agent. v0.10.0 throws that out. It is now an OTLP receiver like any tracing backend, and its output is OTel too: agentevals run ... --emit-otel (or agentevals serve --emit-otel for live evaluations) sends every score to OTEL_EXPORTER_OTLP_ENDPOINT, the same environment variable every other OTel exporter in your stack already reads.
Agent with OTel instrumentation --OTLP--> OpenTelemetry Collector --> your tracing backend
\--> agentevals
--gen_ai.evaluation.result--> Collector
That’s the shape the agentevals README describes, and the Collector in the middle is optional; an agent can export straight to agentevals. I wanted to prove the right-hand side of it actually happens, with nothing agentevals-shaped standing in for the real thing.
Three changes in the release notes matter most:
- OTLP in, replacing the old WebSocket path: HTTP on
:4318or gRPC on:4317, from any producer that follows the OpenTelemetry GenAI semantic conventions (Google ADK, Strands, LangChain, the OpenAI Agents SDK, Pydantic AI, or the official OpenTelemetry GenAI instrumentations). - One conversation model everywhere. The CLI, API, MCP server, live UI and run worker now share the same turns, so a model call wrapped by a framework counts once and tool rounds join across traces. That also means turn boundaries and token counts can differ from v0.9.x on the same recorded telemetry.
- Scores out as
gen_ai.evaluation.resultOTel log events via--emit-otel, which is the half this repo exercises.
It also removed the openai_eval evaluator ahead of OpenAI shutting down its Evals API on 2026-10-31.
What I built
Two recorded agent traces, a golden eval set, a plain otel/opentelemetry-collector-contrib container with a logging exporter standing in for “your existing backend”, and agentevals run --emit-otel pointed at it:
samples/helm.json --\
samples/k8s.json --+--> agentevals run --emit-otel --> OTel Collector --> debug exporter
samples/eval_set_helm.json (stdout)
The two sample traces and the eval set come straight from agentevals’ own samples/
(Apache-2.0): a Helm agent asked to list releases in a kagent namespace, scored on whether it actually called helm_list_releases. One trace calls the tool. One doesn’t. No live LLM call is needed for this demo since agentevals run scores recorded telemetry, which is also why re-scoring the same run costs nothing.
Running it
git clone https://github.com/themsquared/agentevals-otel-roundtrip.git
cd agentevals-otel-roundtrip
./run.sh
run.sh does five things: starts the Collector on localhost:14318 (offset by +10000 from the OTel defaults so it doesn’t collide with one you already have running), builds a venv with agentevals-cli==0.10.0, runs both traces through agentevals run --emit-otel, and greps the Collector’s own log output for the resulting events.
The PASS case
OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:14318 \
agentevals run samples/helm.json \
--eval-set samples/eval_set_helm.json \
-m tool_trajectory_avg_score \
--emit-otel
Trace: 3e289017fe03ffd7c4145316d2eb3d0d
Metric Score Status Per-Invocation Time
------ ------------------------- ------- -------- ---------------- ------
[PASS] tool_trajectory_avg_score 1 PASSED 1 7ms
The FAIL case
Same command, the other trace. This one never called helm_list_releases:
Trace: d497c9dd55717f2c5ecb79bda3028993
Metric Score Status Per-Invocation Time
------ ------------------------- ------- -------- ---------------- ------
[FAIL] tool_trajectory_avg_score 0 FAILED 0 16ms
Invocation 1 trajectory mismatch:
Expected:
- helm_list_releases({})
Actual:
(none)
Gotcha: this command exits non-zero. The first time I ran it in a script with set -e, the whole thing stopped there. That’s agentevals run behaving correctly as a CI gate, not a bug. Wrap it in || true (or check $? deliberately) if your script needs to keep going to inspect the emitted event afterward, the way run.sh does.
The round trip
Both commands ran with --emit-otel. Here’s what landed in the Collector’s own logs, not agentevals’ UI or API, for the PASS case:
EventName: gen_ai.evaluation.result
Attributes:
-> gen_ai.evaluation.name: Str(tool_trajectory_avg_score)
-> agentevals.eval.case.id: Str(helm_list_releases)
-> agentevals.eval_set.id: Str(helm_eval_set)
-> agentevals.evaluator.type: Str(deterministic)
-> gen_ai.evaluation.score.value: Double(1)
-> gen_ai.evaluation.score.label: Str(pass)
-> gen_ai.conversation.id: Str(ctx-7ed9780f-3688-4fc0-b10b-2e4df2f83cf0)
Trace ID: 3e289017fe03ffd7c4145316d2eb3d0d
And the FAIL case produced the same shape with score.value: Double(0) and score.label: Str(fail), carrying its own trace ID. Two evaluator runs, two independent OTel log events, each one still bound to the trace it scored by Trace ID. That binding is the entire point: a backend that already groups logs and spans by trace ID gets the score for free, no join table required.
Why the OTel-native rewrite is bigger than it sounds
I wrote about evaluating kagent agents with agentevals three weeks ago
, back when the only way in was reading kagent’s ADK event table directly and hand-converting function_call/function_response parts into OpenAI-format chat messages. That bridge, kagent-agentevals
, exists because agentevals didn’t yet speak the protocol kagent’s own telemetry was already in.
v0.10.0 makes that bridge unnecessary for anything already emitting standard GenAI OTel spans. The release notes list Go ADK agents on kagent getting full model and tool calls, with arguments, results and tokens, straight over OTel. That’s the ingestion half I didn’t exercise in this demo (see below).
The release marks the rest experimental. The Claude Code harness on kagent v1.0.0-alpha9 captures tool calls by name and model calls with token usage, but not tool arguments, so a trajectory score that checks arguments fails on that harness today. The Codex harness isn’t tested at all. Worth knowing before you point a CI gate at either one.
What this demo doesn’t cover
- Live ingestion.
agentevals serve --devplus a real agent exporting OTLP to:4318is the other half of the v0.10.0 story, and it needs a live LLM call to produce a real trace. I kept this demo to the scores-out half, which is both what changed most and what’s provable without one. - The kagent harness integration described above. Exercising it means standing up a kagent cluster and the Collector
filterprocessor the docs recommend for kagent’sTaskStorespans, which is a second repo’s worth of scope. - A production Collector topology. agentevals’ own OpenTelemetry pipelines doc
has the real recipe: one Collector, two branches, content stripped before it reaches your backend but kept on the branch that reaches agentevals, plus a connector that turns
gen_ai.evaluation.resultevents into a score histogram metric. This demo’scollector/config.yamlis the minimal receiver-to-debug-exporter shape, not that.
Why this is worth caring about if you’re not an agentevals user
The score and the trace it came from carried the same Trace ID in this demo without me writing any code to join them. If your tracing backend already groups logs and spans by trace ID, a gen_ai.evaluation.result event lands next to the agent’s own spans the moment agentevals emits it. No eval dashboard to check separately, no export job to glue the two datasets together after the fact.
Try it
Trace 3e289017fe03ffd7c4145316d2eb3d0d scored 1 and trace d497c9dd55717f2c5ecb79bda3028993 scored 0, and both numbers showed up in a Collector that has never heard of agentevals beyond an OTLP receiver and a logging exporter. That’s the whole claim, reproducible on your own machine:
git clone https://github.com/themsquared/agentevals-otel-roundtrip.git
cd agentevals-otel-roundtrip && ./run.sh
Full source for the Collector config, the run.sh orchestration, and the sample traces: themsquared/agentevals-otel-roundtrip
. The upstream project is agentevals-dev/agentevals
; the OpenTelemetry GenAI semantic conventions it implements are at opentelemetry.io
. Next, I want to try the ingestion half against a real kagent Go ADK agent and see whether the tool-argument gap on the Claude Code harness closes.