Papers/2608.23569
🧪 Test?View on arXiv

ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence

Claude Sonnet, GPT-4o, Local Llama

NL2SQLbenchmarkingenterprise databasessemantic divergence
2608.23569
Builder Relevance
80%
3h ago

Abstract

The paper introduces ESQ-Bench, a benchmark designed to evaluate NL2SQL models in complex enterprise database environments.

Reality Card

Core Claim

ESQ-Bench demonstrates significant execution accuracy degradation across complexity tiers in NL2SQL models, highlighting the challenges of generalization in enterprise schemas.

Method / Result

Schema-linked prompting with GPT-4o shows execution accuracy of 79.8%, 60.3%, and 57.2% across three complexity tiers.

Limitations

The gap between closed API models and open-weight baselines on enterprise Oracle schemas raises concerns about reproducibility.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers