# agentevals v0.10.0: Scores Come Back as OpenTelemetry Events

> agentevals v0.10.0 sends scores back as gen_ai.evaluation.result OpenTelemetry events. I proved the round trip against a real OTel Collector, pass and fail.

- Canonical URL: https://webofmike.com/agentevals-otel-native/
- Author: Mike Moore (https://webofmike.com/about/)
- Published: 2026-10-11
- Last modified: 2026-10-11
- Tags: AI Agents, Evals, Observability, Kubernetes
- Cite as: Mike Moore, "agentevals v0.10.0: Scores Come Back as OpenTelemetry Events", Web of Mike (webofmike.com), 2026-10-11. https://webofmike.com/agentevals-otel-native/


[agentevals](https://github.com/agentevals-dev/agentevals) [v0.10.0](https://github.com/agentevals-dev/agentevals/releases/tag/v0.10.0) shipped two days ago and removed its own WebSocket ingestion path in favor of plain OpenTelemetry. Send OTLP traces in, run an eval, and the score comes back out as a standard `gen_ai.evaluation.result` OTel log event, attached to the exact turn it scored. I ran both halves of that claim against a real, independent OTel Collector and watched the events land. Repo: [themsquared/agentevals-otel-roundtrip](https://github.com/themsquared/agentevals-otel-roundtrip).

*Disclosure: I work at Solo.io, which also builds agentgateway and kagent; agentevals-dev is a separate open-source project several Solo engineers contribute to. Everything below is the open-source CLI, and every claim links to the v0.10.0 release notes or docs.*

## The short answer

agentevals used to be a destination with its own ingestion protocol: a WebSocket stream, a bespoke session model, a SDK you wired into your agent. v0.10.0 throws that out. It is now an OTLP receiver like any tracing backend, and its output is OTel too: `agentevals run ... --emit-otel` (or `agentevals serve --emit-otel` for live evaluations) sends every score to `OTEL_EXPORTER_OTLP_ENDPOINT`, the same environment variable every other OTel exporter in your stack already reads.

```text
Agent with OTel instrumentation --OTLP--> OpenTelemetry Collector --> your tracing backend
                                                                 \--> agentevals
                                                                       --gen_ai.evaluation.result--> Collector
```

That's the shape the [agentevals README](https://github.com/agentevals-dev/agentevals/blob/v0.10.0/README.md) describes, and the Collector in the middle is optional; an agent can export straight to agentevals. I wanted to prove the right-hand side of it actually happens, with nothing agentevals-shaped standing in for the real thing.

Three changes in the release notes matter most:

- **OTLP in**, replacing the old WebSocket path: HTTP on `:4318` or gRPC on `:4317`, from any producer that follows the OpenTelemetry GenAI semantic conventions (Google ADK, Strands, LangChain, the OpenAI Agents SDK, Pydantic AI, or the official OpenTelemetry GenAI instrumentations).
- **One conversation model everywhere.** The CLI, API, MCP server, live UI and run worker now share the same turns, so a model call wrapped by a framework counts once and tool rounds join across traces. That also means turn boundaries and token counts can differ from v0.9.x on the same recorded telemetry.
- **Scores out** as `gen_ai.evaluation.result` OTel log events via `--emit-otel`, which is the half this repo exercises.

It also removed the `openai_eval` evaluator ahead of OpenAI shutting down its Evals API on 2026-10-31.

## What I built

Two recorded agent traces, a golden eval set, a plain `otel/opentelemetry-collector-contrib` container with a logging exporter standing in for "your existing backend", and `agentevals run --emit-otel` pointed at it:

```
samples/helm.json  --\
samples/k8s.json    --+-->  agentevals run --emit-otel  -->  OTel Collector  -->  debug exporter
samples/eval_set_helm.json                                   (stdout)
```

The two sample traces and the eval set come straight from [agentevals' own `samples/`](https://github.com/agentevals-dev/agentevals/tree/v0.10.0/samples) (Apache-2.0): a Helm agent asked to list releases in a `kagent` namespace, scored on whether it actually called `helm_list_releases`. One trace calls the tool. One doesn't. No live LLM call is needed for this demo since `agentevals run` scores recorded telemetry, which is also why re-scoring the same run costs nothing.

## Running it

```bash
git clone https://github.com/themsquared/agentevals-otel-roundtrip.git
cd agentevals-otel-roundtrip
./run.sh
```

`run.sh` does five things: starts the Collector on `localhost:14318` (offset by `+10000` from the OTel defaults so it doesn't collide with one you already have running), builds a venv with `agentevals-cli==0.10.0`, runs both traces through `agentevals run --emit-otel`, and greps the Collector's own log output for the resulting events.

### The PASS case

```bash
OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:14318 \
  agentevals run samples/helm.json \
    --eval-set samples/eval_set_helm.json \
    -m tool_trajectory_avg_score \
    --emit-otel
```

```
Trace: 3e289017fe03ffd7c4145316d2eb3d0d
        Metric                       Score  Status      Per-Invocation  Time
------  -------------------------  -------  --------  ----------------  ------
[PASS]  tool_trajectory_avg_score        1  PASSED                   1  7ms
```

### The FAIL case

Same command, the other trace. This one never called `helm_list_releases`:

```
Trace: d497c9dd55717f2c5ecb79bda3028993
        Metric                       Score  Status      Per-Invocation  Time
------  -------------------------  -------  --------  ----------------  ------
[FAIL]  tool_trajectory_avg_score        0  FAILED                   0  16ms
  Invocation 1 trajectory mismatch:
    Expected:
      - helm_list_releases({})
    Actual:
      (none)
```

**Gotcha:** this command exits non-zero. The first time I ran it in a script with `set -e`, the whole thing stopped there. That's `agentevals run` behaving correctly as a CI gate, not a bug. Wrap it in `|| true` (or check `$?` deliberately) if your script needs to keep going to inspect the emitted event afterward, the way `run.sh` does.

### The round trip

Both commands ran with `--emit-otel`. Here's what landed in the Collector's own logs, not agentevals' UI or API, for the PASS case:

```
EventName: gen_ai.evaluation.result
Attributes:
     -> gen_ai.evaluation.name: Str(tool_trajectory_avg_score)
     -> agentevals.eval.case.id: Str(helm_list_releases)
     -> agentevals.eval_set.id: Str(helm_eval_set)
     -> agentevals.evaluator.type: Str(deterministic)
     -> gen_ai.evaluation.score.value: Double(1)
     -> gen_ai.evaluation.score.label: Str(pass)
     -> gen_ai.conversation.id: Str(ctx-7ed9780f-3688-4fc0-b10b-2e4df2f83cf0)
Trace ID: 3e289017fe03ffd7c4145316d2eb3d0d
```

And the FAIL case produced the same shape with `score.value: Double(0)` and `score.label: Str(fail)`, carrying its own trace ID. Two evaluator runs, two independent OTel log events, each one still bound to the trace it scored by `Trace ID`. That binding is the entire point: a backend that already groups logs and spans by trace ID gets the score for free, no join table required.

## Why the OTel-native rewrite is bigger than it sounds

I wrote about evaluating kagent agents with agentevals [three weeks ago](/kagent-trajectory-evals/), back when the only way in was reading kagent's ADK event table directly and hand-converting `function_call`/`function_response` parts into OpenAI-format chat messages. That bridge, [kagent-agentevals](https://github.com/themsquared/kagent-agentevals), exists because agentevals didn't yet speak the protocol kagent's own telemetry was already in.

v0.10.0 makes that bridge unnecessary for anything already emitting standard GenAI OTel spans. The [release notes](https://github.com/agentevals-dev/agentevals/releases/tag/v0.10.0) list Go ADK agents on kagent getting full model and tool calls, with arguments, results and tokens, straight over OTel. That's the ingestion half I didn't exercise in this demo (see below).

The release marks the rest experimental. The Claude Code harness on kagent v1.0.0-alpha9 captures tool calls by name and model calls with token usage, but not tool arguments, so a trajectory score that checks arguments fails on that harness today. The Codex harness isn't tested at all. Worth knowing before you point a CI gate at either one.

## What this demo doesn't cover

- **Live ingestion.** `agentevals serve --dev` plus a real agent exporting OTLP to `:4318` is the other half of the v0.10.0 story, and it needs a live LLM call to produce a real trace. I kept this demo to the scores-out half, which is both what changed most and what's provable without one.
- **The kagent harness integration** described above. Exercising it means standing up a kagent cluster and the Collector `filter` processor the docs recommend for kagent's `TaskStore` spans, which is a second repo's worth of scope.
- **A production Collector topology.** agentevals' own [OpenTelemetry pipelines doc](https://github.com/agentevals-dev/agentevals/blob/v0.10.0/docs/opentelemetry-pipeline.md) has the real recipe: one Collector, two branches, content stripped before it reaches your backend but kept on the branch that reaches agentevals, plus a connector that turns `gen_ai.evaluation.result` events into a score histogram metric. This demo's `collector/config.yaml` is the minimal receiver-to-debug-exporter shape, not that.

## Why this is worth caring about if you're not an agentevals user

The score and the trace it came from carried the same `Trace ID` in this demo without me writing any code to join them. If your tracing backend already groups logs and spans by trace ID, a `gen_ai.evaluation.result` event lands next to the agent's own spans the moment agentevals emits it. No eval dashboard to check separately, no export job to glue the two datasets together after the fact.

## Try it

Trace `3e289017fe03ffd7c4145316d2eb3d0d` scored 1 and trace `d497c9dd55717f2c5ecb79bda3028993` scored 0, and both numbers showed up in a Collector that has never heard of agentevals beyond an OTLP receiver and a logging exporter. That's the whole claim, reproducible on your own machine:

```bash
git clone https://github.com/themsquared/agentevals-otel-roundtrip.git
cd agentevals-otel-roundtrip && ./run.sh
```

Full source for the Collector config, the `run.sh` orchestration, and the sample traces: [themsquared/agentevals-otel-roundtrip](https://github.com/themsquared/agentevals-otel-roundtrip). The upstream project is [agentevals-dev/agentevals](https://github.com/agentevals-dev/agentevals); the OpenTelemetry GenAI semantic conventions it implements are at [opentelemetry.io](https://opentelemetry.io/docs/specs/semconv/gen-ai/). Next, I want to try the ingestion half against a real kagent Go ADK agent and see whether the tool-argument gap on the Claude Code harness closes.

