🧪 Test?View on arXiv
ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence
Claude Sonnet, GPT-4o, Local Llama
NL2SQLbenchmarkingenterprise databasessemantic divergence
2608.23569
Builder Relevance
3h ago80%
Abstract
The paper introduces ESQ-Bench, a benchmark designed to evaluate NL2SQL models in complex enterprise database environments.
Reality Card
Core Claim
ESQ-Bench demonstrates significant execution accuracy degradation across complexity tiers in NL2SQL models, highlighting the challenges of generalization in enterprise schemas.
Method / Result
Schema-linked prompting with GPT-4o shows execution accuracy of 79.8%, 60.3%, and 57.2% across three complexity tiers.
Limitations
The gap between closed API models and open-weight baselines on enterprise Oracle schemas raises concerns about reproducibility.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.