🧪 Test?View on arXiv
Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline
James L., Author 2, Author 3, Author 4, Author 5
text-to-SQLLLM evaluationmodel auditingperformance benchmarking
2609.30290
Builder Relevance
1h ago80%
Abstract
This paper evaluates the performance of an LLM-as-judge in a text-to-SQL pipeline, revealing significant discrepancies with human annotators and proposing solutions.
Reality Card
Core Claim
The study demonstrates that a self-hosted Qwen3.6-27B judge achieves a Cohen's kappa of 0.72, significantly outperforming the deployed gpt-4o-mini judge with a kappa of only 0.04.
Method / Result
Three strong judges under unanimity routing reach a kappa of 0.79 at 89.7% auto-coverage.
Limitations
The head-to-head comparison between judges is underpowered with only 96 samples, which may affect the reliability of the results.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.