agentevals v0.10.0: Scores Come Back as OpenTelemetry Events

agentevals v0.10.0 sends scores back as gen_ai.evaluation.result OpenTelemetry events. I proved the round trip against a real OTel Collector, pass and fail.

agentevals v0.10.0 shipped two days ago and removed its own WebSocket ingestion path in favor of plain OpenTelemetry. Send OTLP traces in, run an eval, and the score comes back out as a standard gen_ai.evaluation.result OTel log event, attached to the exact turn it scored. I ran both halves of that claim against a real, independent OTel Collector and watched the events land. Repo: themsquared/agentevals-otel-roundtrip .

Disclosure: I work at Solo.io, which also builds agentgateway and kagent; agentevals-dev is a separate open-source project several Solo engineers contribute to. Everything below is the open-source CLI, and every claim links to the v0.10.0 release notes or docs.

The short answer

agentevals used to be a destination with its own ingestion protocol: a WebSocket stream, a bespoke session model, a SDK you wired into your agent. v0.10.0 throws that out. It is now an OTLP receiver like any tracing backend, and its output is OTel too: agentevals run ... --emit-otel (or agentevals serve --emit-otel for live evaluations) sends every score to OTEL_EXPORTER_OTLP_ENDPOINT, the same environment variable every other OTel exporter in your stack already reads.

Agent with OTel instrumentation --OTLP--> OpenTelemetry Collector --> your tracing backend
                                                                 \--> agentevals
                                                                       --gen_ai.evaluation.result--> Collector

That’s the shape the agentevals README describes, and the Collector in the middle is optional; an agent can export straight to agentevals. I wanted to prove the right-hand side of it actually happens, with nothing agentevals-shaped standing in for the real thing.

Three changes in the release notes matter most:

  • OTLP in, replacing the old WebSocket path: HTTP on :4318 or gRPC on :4317, from any producer that follows the OpenTelemetry GenAI semantic conventions (Google ADK, Strands, LangChain, the OpenAI Agents SDK, Pydantic AI, or the official OpenTelemetry GenAI instrumentations).
  • One conversation model everywhere. The CLI, API, MCP server, live UI and run worker now share the same turns, so a model call wrapped by a framework counts once and tool rounds join across traces. That also means turn boundaries and token counts can differ from v0.9.x on the same recorded telemetry.
  • Scores out as gen_ai.evaluation.result OTel log events via --emit-otel, which is the half this repo exercises.

It also removed the openai_eval evaluator ahead of OpenAI shutting down its Evals API on 2026-10-31.

What I built

Two recorded agent traces, a golden eval set, a plain otel/opentelemetry-collector-contrib container with a logging exporter standing in for “your existing backend”, and agentevals run --emit-otel pointed at it:

samples/helm.json  --\
samples/k8s.json    --+-->  agentevals run --emit-otel  -->  OTel Collector  -->  debug exporter
samples/eval_set_helm.json                                   (stdout)

The two sample traces and the eval set come straight from agentevals’ own samples/ (Apache-2.0): a Helm agent asked to list releases in a kagent namespace, scored on whether it actually called helm_list_releases. One trace calls the tool. One doesn’t. No live LLM call is needed for this demo since agentevals run scores recorded telemetry, which is also why re-scoring the same run costs nothing.

Running it

git clone https://github.com/themsquared/agentevals-otel-roundtrip.git
cd agentevals-otel-roundtrip
./run.sh

run.sh does five things: starts the Collector on localhost:14318 (offset by +10000 from the OTel defaults so it doesn’t collide with one you already have running), builds a venv with agentevals-cli==0.10.0, runs both traces through agentevals run --emit-otel, and greps the Collector’s own log output for the resulting events.

The PASS case

OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:14318 \
  agentevals run samples/helm.json \
    --eval-set samples/eval_set_helm.json \
    -m tool_trajectory_avg_score \
    --emit-otel
Trace: 3e289017fe03ffd7c4145316d2eb3d0d
        Metric                       Score  Status      Per-Invocation  Time
------  -------------------------  -------  --------  ----------------  ------
[PASS]  tool_trajectory_avg_score        1  PASSED                   1  7ms

The FAIL case

Same command, the other trace. This one never called helm_list_releases:

Trace: d497c9dd55717f2c5ecb79bda3028993
        Metric                       Score  Status      Per-Invocation  Time
------  -------------------------  -------  --------  ----------------  ------
[FAIL]  tool_trajectory_avg_score        0  FAILED                   0  16ms
  Invocation 1 trajectory mismatch:
    Expected:
      - helm_list_releases({})
    Actual:
      (none)

Gotcha: this command exits non-zero. The first time I ran it in a script with set -e, the whole thing stopped there. That’s agentevals run behaving correctly as a CI gate, not a bug. Wrap it in || true (or check $? deliberately) if your script needs to keep going to inspect the emitted event afterward, the way run.sh does.

The round trip

Both commands ran with --emit-otel. Here’s what landed in the Collector’s own logs, not agentevals’ UI or API, for the PASS case:

EventName: gen_ai.evaluation.result
Attributes:
     -> gen_ai.evaluation.name: Str(tool_trajectory_avg_score)
     -> agentevals.eval.case.id: Str(helm_list_releases)
     -> agentevals.eval_set.id: Str(helm_eval_set)
     -> agentevals.evaluator.type: Str(deterministic)
     -> gen_ai.evaluation.score.value: Double(1)
     -> gen_ai.evaluation.score.label: Str(pass)
     -> gen_ai.conversation.id: Str(ctx-7ed9780f-3688-4fc0-b10b-2e4df2f83cf0)
Trace ID: 3e289017fe03ffd7c4145316d2eb3d0d

And the FAIL case produced the same shape with score.value: Double(0) and score.label: Str(fail), carrying its own trace ID. Two evaluator runs, two independent OTel log events, each one still bound to the trace it scored by Trace ID. That binding is the entire point: a backend that already groups logs and spans by trace ID gets the score for free, no join table required.

Why the OTel-native rewrite is bigger than it sounds

I wrote about evaluating kagent agents with agentevals three weeks ago , back when the only way in was reading kagent’s ADK event table directly and hand-converting function_call/function_response parts into OpenAI-format chat messages. That bridge, kagent-agentevals , exists because agentevals didn’t yet speak the protocol kagent’s own telemetry was already in.

v0.10.0 makes that bridge unnecessary for anything already emitting standard GenAI OTel spans. The release notes list Go ADK agents on kagent getting full model and tool calls, with arguments, results and tokens, straight over OTel. That’s the ingestion half I didn’t exercise in this demo (see below).

The release marks the rest experimental. The Claude Code harness on kagent v1.0.0-alpha9 captures tool calls by name and model calls with token usage, but not tool arguments, so a trajectory score that checks arguments fails on that harness today. The Codex harness isn’t tested at all. Worth knowing before you point a CI gate at either one.

What this demo doesn’t cover

  • Live ingestion. agentevals serve --dev plus a real agent exporting OTLP to :4318 is the other half of the v0.10.0 story, and it needs a live LLM call to produce a real trace. I kept this demo to the scores-out half, which is both what changed most and what’s provable without one.
  • The kagent harness integration described above. Exercising it means standing up a kagent cluster and the Collector filter processor the docs recommend for kagent’s TaskStore spans, which is a second repo’s worth of scope.
  • A production Collector topology. agentevals’ own OpenTelemetry pipelines doc has the real recipe: one Collector, two branches, content stripped before it reaches your backend but kept on the branch that reaches agentevals, plus a connector that turns gen_ai.evaluation.result events into a score histogram metric. This demo’s collector/config.yaml is the minimal receiver-to-debug-exporter shape, not that.

Why this is worth caring about if you’re not an agentevals user

The score and the trace it came from carried the same Trace ID in this demo without me writing any code to join them. If your tracing backend already groups logs and spans by trace ID, a gen_ai.evaluation.result event lands next to the agent’s own spans the moment agentevals emits it. No eval dashboard to check separately, no export job to glue the two datasets together after the fact.

Try it

Trace 3e289017fe03ffd7c4145316d2eb3d0d scored 1 and trace d497c9dd55717f2c5ecb79bda3028993 scored 0, and both numbers showed up in a Collector that has never heard of agentevals beyond an OTLP receiver and a logging exporter. That’s the whole claim, reproducible on your own machine:

git clone https://github.com/themsquared/agentevals-otel-roundtrip.git
cd agentevals-otel-roundtrip && ./run.sh

Full source for the Collector config, the run.sh orchestration, and the sample traces: themsquared/agentevals-otel-roundtrip . The upstream project is agentevals-dev/agentevals ; the OpenTelemetry GenAI semantic conventions it implements are at opentelemetry.io . Next, I want to try the ingestion half against a real kagent Go ADK agent and see whether the tool-argument gap on the Claude Code harness closes.

Frequently asked questions

How does agentevals v0.10.0 send evaluation results back into my observability pipeline?

Run `agentevals run` or `agentevals serve` with `--emit-otel` (or set AGENTEVALS_EVALUATION_EVENTS=true), and point the standard OTEL_EXPORTER_OTLP_ENDPOINT variable at your collector. Each metric emits one gen_ai.evaluation.result log event per scored turn, carrying gen_ai.evaluation.name, score.value, score.label (pass, fail or not_evaluated), and the trace and span IDs of the turn it scores.

Do I need agentevals' own SDK to get agent traces into it?

No. agentevals ingests plain OTLP over HTTP (:4318) or gRPC (:4317) from any producer that follows the OpenTelemetry GenAI semantic conventions: Google ADK, Strands, LangChain, the OpenAI Agents SDK, Pydantic AI, or the official OpenTelemetry GenAI instrumentations. That replaced the WebSocket ingestion path v0.9.x used.

What changed between agentevals v0.9.x and v0.10.0?

v0.10.0 is a breaking OTel-native rewrite. WebSocket ingestion is removed in favor of OTLP. The CLI, API, MCP server, live UI and run worker now share one conversation model, so a model call wrapped by a framework counts once and tool rounds join across traces; turn boundaries and token counts can differ from v0.9.x as a result. The openai_eval evaluator was also dropped ahead of OpenAI shutting down its Evals API on 2026-10-31.

Can I evaluate kagent agents with agentevals v0.10.0 without a custom bridge?

Partially, and the release notes mark it experimental. Go ADK agents on kagent get full model and tool calls with arguments, results and tokens over OTel. The Claude Code harness is tested on kagent v1.0.0-alpha9 for tool calls by name and model calls with token usage, but tool arguments are not recorded yet, so trajectory scores that check arguments fail on that harness. The Codex harness is untested.