Papers/2609.30328
🧪 Test?View on arXiv

When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess

Not provided in the abstract

multi-agent systemscode verificationlabel-free methodsaccuracy improvement
2609.30328
Builder Relevance
70%
1h ago

Abstract

The paper discusses the limitations of multi-agent verification in code judging and introduces a label-free method to assess when a judge lacks a basis for its answer.

Reality Card

Core Claim

The study demonstrates a label-free method that improves the accuracy of code judging by allowing the judge to decline comparisons it cannot make.

Method / Result

The accuracy of the judge improved from 20.7% to 36.9% by gating on one of the measurements.

Limitations

The method's effectiveness is limited by the independence of evidence and the nature of the comparisons being made.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers