Papers/2609.30290
🧪 Test?View on arXiv

Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline

James L., Author 2, Author 3, Author 4, Author 5

text-to-SQLLLM evaluationmodel auditingperformance benchmarking
2609.30290
Builder Relevance
80%
1h ago

Abstract

This paper evaluates the performance of an LLM-as-judge in a text-to-SQL pipeline, revealing significant discrepancies with human annotators and proposing solutions.

Reality Card

Core Claim

The study demonstrates that a self-hosted Qwen3.6-27B judge achieves a Cohen's kappa of 0.72, significantly outperforming the deployed gpt-4o-mini judge with a kappa of only 0.04.

Method / Result

Three strong judges under unanimity routing reach a kappa of 0.79 at 89.7% auto-coverage.

Limitations

The head-to-head comparison between judges is underpowered with only 96 samples, which may affect the reliability of the results.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers