🧪 Test?View on arXiv
When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess
Not provided in the abstract
multi-agent systemscode verificationlabel-free methodsaccuracy improvement
2609.30328
Builder Relevance
1h ago70%
Abstract
The paper discusses the limitations of multi-agent verification in code judging and introduces a label-free method to assess when a judge lacks a basis for its answer.
Reality Card
Core Claim
The study demonstrates a label-free method that improves the accuracy of code judging by allowing the judge to decline comparisons it cannot make.
Method / Result
The accuracy of the judge improved from 20.7% to 36.9% by gating on one of the measurements.
Limitations
The method's effectiveness is limited by the independence of evidence and the nature of the comparisons being made.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.