An agent acts before its mistakes surface
When a large language model agent is deployed in front of a customer, it acts. It issues refund tool calls, database queries and blocks of code. If one of those actions is silently wrong, the mistake lands downstream, after execution. The natural question an operator asks is simple: which of these calls should I trust?
Until now, the practical answers were weak. The most direct confidence signal, the model's own token probabilities, is not exposed by frontier chat APIs. Asking the agent how confident it is, is unreliable exactly when it matters most. And the standard fallback, resampling the same model several times and measuring agreement, fails in a surprising way: deployed frontier models are near-deterministic, so they reproduce the same call across samples, leaving almost nothing to compare while still costing multiple frontier generations.
A preprint posted to arXiv on October 2, 2026, proposes a third route: measure the agent's proposed action with a different model that you control. "Proxy Confidence: Auditing Black-Box LLM Agents with a Surrogate's Log-Probabilities", by Yikai Zhao, Saurabh Pandey and Pradeep Kumar Misra (all affiliated with Amazon, per the paper), describes a low-cost, open-weight "sidecar" model that reads the same context, tool schema and proposed action the agent saw, and scores the call using its own log-probabilities.
This article is an analysis of that preprint and of what it means for access to independent verification of AI systems. All performance numbers cited below are the authors' own reported results on their chosen tasks. The paper is a non-peer-reviewed preprint, and its results are not independently validated in production. Nothing here should be read as evidence that any production system has adopted the method.
How the surrogate scoring works
The sidecar works without touching the agent's internals. There are no gradients, no fine-tuning, no access to weights or shared weights, and no need for the provider's cooperation. Because the auditor is a smaller open-weight model and the action is already known, its scoring pass is a single forward pass over existing text (a prefill) with no autoregressive decoding, which the authors describe as lightweight and actor-agnostic.
The paper organizes the signal into two families. Generative likelihood readouts, teacher forcing and request-PMI, weigh how likely the surrogate finds each argument value in the proposed call, with request-PMI discounting values the scorer would find likely regardless of the request. Discriminative readouts, a yes/no verdict and a tool-choice competition among sibling functions, judge the call as a whole. The authors state a principle for which to trust: a generative likelihood localizes wrong argument values, while a verdict catches holistically wrong calls. When the error type is unknown, as it usually is in deployment, they recommend an ensemble of readouts; they report that a learned logistic combiner does no better than a simple orientation-corrected average of the family.
It is worth being clear about what is measured and what is inferred. What is directly observed in the paper is that a second model's likelihood, computed over an agent's proposed tool call, ranks correct calls above wrong ones on the authors' benchmarks better than stated confidence or self-consistency do. What is an interpretation, offered by the authors, is the claim that this readout type structure predicts which failure mode a readout catches, and the resulting guidance about which readout to trust in which regime.
Why the access angle matters
The interesting backdrop is not the arXiv paper itself but the infrastructure around it. Frontier chat APIs generally do not expose the token probabilities of the hosted model. That is a design choice by providers, and it means the most direct confidence signal for auditing an agent is off the table for anyone who does not run the weights.
This sits alongside a broader pattern of capability access that is gated rather than open. On October 5, 2026, OpenAI published its approach to EU text provenance rules under the AI Act. The announcement describes text watermarking and a detector; detector access is, per the announcement, initially limited to approved researchers and expert organizations, and the tool reports whether it detects an OpenAI watermark without identifying users or revealing prompts. OpenAI also states plans to make the watermarking technology available in open source, and its image and audio verification tools remain publicly accessible. The pattern, across both cases, is that some verification-relevant signals are provided only at the provider's discretion, under conditions the provider sets.
The proxy-confidence approach takes a different philosophy: it does not ask the provider for anything. It runs a model the operator already controls, over information the operator already has, and produces its own signal. In that sense it is a demonstration that independent verification does not always require provider permission, even when the provider withholds its internal numbers.
The limit is equally real. A surrogate's log-probabilities are a second opinion from a different model, not the agent's internals. The authors themselves scope their claims: the readout is trained on nothing and sees no hidden state, so it infers confidence from plausibility, not from knowledge of the actor. Where the agent's failure depends on information only the agent holds, a sidecar cannot see it. Open-weight surrogates partially fill the audit gap; they do not close it. A provider that published its own token log-probabilities would still supply a strictly more direct signal than any external model can.
Our view at Expectancy, stated as opinion: this is the kind of access architecture worth encouraging. Independent verification that depends on an operator's own compute, rather than on provider approvals, strengthens the position of everyone who deploys or is affected by deployed agents. The paper is one preprint with unreplicated numbers, not a validated production method, but the direction is sound and worth watching for replication and independent evaluation.
What readers should take from this
For non-specialists, the takeaway can be summarized in three steps. First, when a black-box agent proposes an action, both its stated confidence and repeated sampling are weak checks, one because it is least reliable on confident mistakes and the other because modern agents repeat themselves. Second, you can run a small open-weight model of your own alongside, feed it the same context and proposed action, and use its likelihood numbers as a second opinion; the paper reports that this second opinion ranks good and bad calls apart better than the alternatives, at a fraction of the cost of resampling. Third, the signal can be used either to route suspicious calls to human review or fed back to the agent so it adapts, which the authors report improved task success on live-execution benchmarks.
What remains open: independent replication of the reported gains, evaluation on tasks and actor families beyond the authors' selection, and evidence from real production deployments. None of that exists yet in public, verified form. The paper is best read as a promising, clearly reasoned proposal with author-reported evidence, not as a proven technique.
The Expectancy desk supports evidence-based AI progress, learning, curiosity and human agency. Methods like this one, which let operators check deployed agents without waiting for permission, are squarely in that tradition. So is plain honesty about what has and has not been shown, which is why this article labels every number above as author-reported and treats the access argument as analysis rather than settled fact.
