Answer in brief
A 21 September study finds major reporting gaps in agent-security evaluations. Its warning concerns whether attack rates can be compared, rather than proving that any particular defense fails.
The new finding concerns how scores are made
A preprint submitted on 21 September by Chetan Pathade, Prathamesh Pawar and Shubham Patil examines a measurement problem in AI security evaluation. Across 259 agent-security papers, the authors ask whether a shared headline, attack success rate, actually describes comparable experiments. The September result remains relevant as teams compare increasingly capable agents, because a single percentage can conceal choices about what was tested and what counted as success.
The study combines an audit of published reporting with mathematical analysis. It does not rerun defenses end to end or publish a new ranking of commercial models. Its central concern is the relation between papers: two internally sensible experiments can produce percentages that should not be placed on the same scale. That is a limitation on comparison, rather than proof that every underlying result is wrong.
The denominator can change the story
Attack success rate divides successful outcomes by a set of evaluated units, but the unit can be an attempt, a task or a complete trajectory. A task may contain several opportunities for a system to fail. Counting each opportunity separately and counting whether the whole task experienced any failure can therefore produce different percentages from the same underlying behavior, without either calculation being arithmetically mistaken.
The authors identify six dimensions, including the success judge, repeat trials, attacker adaptation, knowledge of the defense and treatment of partial success. For readers, the practical question is whether two reports fixed those dimensions in compatible ways. A lower percentage becomes meaningful only after its denominator and decision rule are understood. Precision to a decimal place cannot repair a poorly specified or changing measurement target.
| Reporting question | Result | Population and method |
|---|---|---|
| Neither variance nor repeated runs | 58%; 95% interval 44–71% | Manual random sample of 50 papers |
| Neither variance nor repeated runs | 65.3% | Automated coding of 259 papers |
| Enough decoding detail | 30.9% | Automated corpus assessment |
| Human agreement check for model judges | 29.7% | 64 papers confirmed to use an LLM judge |
Missing uncertainty is a reporting finding
In a manually coded random sample of 50 papers, the authors find that 58 percent report neither variance nor repeated runs, with a rounded 95 percent interval of 44–71 percent. Automated coding of the full 259-paper corpus gives 65.3 percent. These are separate estimates with different methods and uncertainty; the larger automated number should not simply replace the manual result as if it were exact.
The audit also reports that 30.9 percent disclose enough decoding information to establish whether evaluation was stochastic. Among 64 papers confirmed to use a language model as judge, 29.7 percent report a check against human labels. These statistics concern what the inspected documents disclose. They do not establish that researchers never repeated experiments or that every unreported judgment was incorrect.
Small samples can make a ranking fragile
The paper’s analytical scenario uses a benchmark of 100 independent instances. Under its stated assumptions, the detectable difference is about 18.2 percentage points at conventional statistical power. For two defenses separated by a true five-point gap, a normal approximation puts the chance of reversing their order in a single evaluation near 21 percent. Those are calculated examples, not observed production incident rates.
Independence is a consequential assumption: related tasks or shared environments can reduce the effective amount of evidence. The figures therefore illustrate why modest score gaps deserve uncertainty estimates, rather than creating a universal threshold that applies to every benchmark. A useful comparison should show repeated outcomes and the sampling design, allowing readers to judge whether an apparent improvement is larger than the variation in its own evaluation.
A secure system still has to complete legitimate work
The proposed ten-item reporting checklist includes performance on benign tasks. This is a useful connection between safety and capability evaluation. A system that refuses almost everything may show few successful attacks while failing the work it was deployed to do. Reporting the security measure beside legitimate task completion helps distinguish an effective boundary from an unusable service, without requiring a misleading single combined score.
NIST’s established AI Risk Management Framework offers contextual support for documented test sets, methods, deployment conditions and metric effectiveness. It is older guidance, not a September endorsement of this preprint. Our interpretation is that purchasing decisions should demand a pair of observable outcomes: how often the defended system stays within its limits and how often it successfully performs the authorized task under the same conditions.
The audit itself has limits worth retaining
The corpus is restricted to English-language arXiv papers and depends on search, rendering and automated detection. The authors describe missing HTML, possible numeral loss and a single human adjudicator. They have not published an artifact repository, and their reporting checklist is proposed rather than experimentally validated. These details limit how confidently the estimated reporting rates can be generalized to the entire research field.
As of 1 October, the contribution is a concrete challenge to casual cross-paper security rankings. The strongest next step would independently reproduce the coding and test matched defenses with repeated trials, consistent success rules and a legitimate-task measure. Until then, the useful editorial conclusion is to read an attack rate together with its measurement procedure. The percentage alone cannot tell a buyer which agent is reliably safer.
Questions and answers
Does this study show that most AI defenses do not work?
No. It audits how studies report measurements and whether their attack rates are comparable. Missing variance or repetition details do not prove that an individual defense is ineffective or that its experiment was never repeated.
Why do the manual and automated percentages differ?
The manual estimate comes from a random sample of 50 papers, while automated coding covers 259. Sampling uncertainty and detection errors both matter, so the two figures are separate estimates rather than interchangeable counts.
Is the roughly 21% misranking figure an observed model failure rate?
No. It comes from an analytical comparison under stated sampling assumptions for a 100-instance scenario with a true five-percentage-point difference. It is not a measured security failure rate in production.
