The green checkmark problem

A checkbox that says a task is done tells you very little about how it was done. An AI agent can reach the correct final answer while burning tokens on a failed search that triggers a second search, or on a truncated file read that leads to a command fetching the same content again. The extra steps add latency, cost money in tokens, and create more chances for something to go wrong, yet a simple success check hides all of it.

On September 30, 2026, NVIDIA published a technical blog post, "Tracing Agent Harness Behavior with NVIDIA NeMo Relay," that demonstrates an approach to this problem: record what the agent actually did, step by step, so humans can audit the path and not just the destination. This article is an analysis of that tutorial and of the practice it represents. It is important to be clear up front that the source is a vendor tutorial with vendor-run demonstrations, not independent benchmarking. Everything below that depends on NVIDIA's own numbers should be read as vendor-reported.

What the primary source actually says

An agent can finish a task and still take an inefficient path. A failed search can trigger another search. A truncated file read can lead to a command fetching the same content again. A correct final answer hides those extra steps, even though they increase latency and consume tokens.

That is the opening framing of the blog post, published September 30, 2026 on the NVIDIA Technical Blog by William Markito Oliveira, Maryam Najafian, Moon Chung and Teknium. The wording matters because it states the article's core premise plainly: success checks answer whether a task succeeded, but not how. The authors write that a success check by itself cannot explain why an agent recovered from a tool error, stopped early, or needed extra model calls.

Three views of one run

NeMo Relay is NVIDIA's observability layer for AI agents. According to the tutorial, the Hermes Agent harness includes NeMo Relay natively, and Relay represents sessions, turns, model calls, and tool calls in a scope hierarchy, recording lifecycle events as work begins and ends with their timing and parent-child relationships.

Relay produces three representations of the same run, each for a different purpose. ATOF, the Agent Trajectory Observability Format, is a JSONL log of scope starts, scope ends, and point-in-time marks, with IDs and timestamps that let you reconstruct the run; the tutorial positions it for debugging and auditing individual events. ATIF, the Agent Trajectory Interchange Format, is a step-by-step JSON record of agent interactions, tool calls, and observations assembled from those lifecycle events, meant for reviewing or evaluating the agent's path. Third, Relay exports OpenTelemetry spans with OpenInference labels, which mark agent, LLM, and tool spans and can be viewed in OpenTelemetry-compatible tools such as Arize Phoenix or, per the tutorial, other OTLP-compatible backends such as LangSmith.

One nuance the post draws out is worth repeating: an ATIF tool request shows what the model asked to run, but not whether it succeeded. To verify an outcome you inspect ATOF for the matching tool start and end events and any recorded errors, paired by a shared uuid and linked to a parent via parent_uuid.

Two demonstrations, vendor-run

The first demonstration is deliberately small. Hermes uses its terminal tool to run a Python script inside an isolated Docker container, and the script prints a fixed output, VALUE=42, giving the runner an exact success check. The tutorial states that the container cannot access the network, the repository checkout, or the NVIDIA API key, and that Hermes cannot fall back to running terminal commands on the host. Those isolation claims are vendor-stated; I did not independently reproduce them.

In one verified run, the reported trace summary showed 74 ATOF events, 2 completed LLM scopes both carrying usage data, 7,239 prompt tokens and 96 completion tokens, 1 tool call, 0 tool errors, and a three-step ATIF trajectory for the model nvidia/nemotron-3.5-lightning-30b-a3b. NVIDIA notes that token counts and identifiers vary between runs, so these figures illustrate one run rather than a stable benchmark.

The second experiment is a multi-tool research task: Hermes receives a travel record with clues about an unnamed machine-learning conference, must search the web, verify the answer on the official conference website, save a report, and return the conference name. The runner checks that the agent identified COLT 2026, saved a report with the expected details and official source, completed read_file, web_search, web_extract, and write_file calls, produced a nonempty ATIF trajectory, and sent spans with positive token usage to Phoenix. Because the task uses live web search, the tutorial itself cautions that such runs are for exploring behavior rather than ranking models; a controlled comparison would require fixed search responses and repeated runs.

The 108-run harness comparison

The most substantive part of the post is a case study on evaluating a change to an agent harness. Using the Hermes ToolPerf evaluation, the authors compared a baseline harness revision against a fixed revision across 108 runs. The vendor-reported result: Qwen Coder 30B recovered more tasks under the fixed revision, but with increased calls, data, and latency.

That pattern is the economically interesting one. A fix that raises task-completion rates while also raising token consumption, data transfer, and response time is not free. Whether it is worth adopting depends on how much the extra tokens and latency cost against the value of the recovered tasks, and the trace evidence is what makes that tradeoff visible. The tutorial outlines a general method: choose a fixed task with an exact, automated success check, define a baseline and a change, run both repeatedly, and compare verification results against trace evidence such as tool-call counts, durations, and token use.

Again, the caveat bears emphasizing: the 108-run comparison is reported by NVIDIA and the Hermes ToolPerf repository, and this article has not independently replicated it. The referenced August 6, 2026 rerun results in the ToolPerf repository could not be retrieved for verification during the writing of this piece, so the rerun date is based on the blog's links and the assignment's verified notes.

Privacy is the flip side of the recorder

The tutorial closes with a safety note that deserves attention from anyone adopting trace tooling. Traces, depending on configuration, can contain prompts, model responses, tool arguments and results, file paths, and other application data. NVIDIA advises reviewing traces before sharing them. A flight recorder, in other words, records the sensitive things your agent touched, which is exactly what makes it useful for auditing and exactly what makes it a liability if shared carelessly.

The post also frames Relay as an evidence layer for agent safety and security governance: structured traces and trajectories that enterprises, evaluators, and security systems can use to investigate agent behavior, evaluate policies, and improve controls.

The nontechnical takeaway: audit the path, not just the destination

Why does this matter to a non-specialist reader? Because AI agents are increasingly trusted to do work, and the honest way to build that trust is inspectability, not vibes. A system that can only say "the task passed" leaves you unable to answer basic questions: did it waste money on redundant calls? did it silently recover from an error that signals a deeper problem? did a harness change that looks like an improvement actually trade accuracy for cost?

Trace-based observability is an emerging answer, and NeMo Relay is one concrete instance of it: an open, structured record of an agent's path that ordinary tooling can display. The Expectancy principle at work here is simple. Progress claims should be inspectable. When a vendor shows you both the success rate and the token bill, that is a healthier kind of evidence than a headline number, and readers should ask for the same standard everywhere.

One final boundary: this article analyzes NVIDIA's own tutorial and case study. Independent evaluation of NeMo Relay, Hermes traces, or the ToolPerf comparison would require running the harnesses on neutral tasks, which has not been done here. The verified facts are that the tool exists, that the blog documents these demonstrations and outputs as described, and that the recommended practices, isolation, exact success checks, trace review before sharing, are sound engineering hygiene regardless of vendor.