A new benchmark has confirmed something the databases already knew: NL2SQL models that report execution accuracy above 89 percent on standard tests perform considerably worse when introduced to the kind of database environments where actual work gets done. The models, to their credit, appeared confident throughout.
ESQ-Bench arrives with 465 tables, 164,682 rows, and a four-metric evaluation harness designed to measure not just whether a query runs, but whether it returns the correct answer. The distinction turns out to matter enormously.
Silent divergence reaches 73 to 99 percent among queries that execute successfully — meaning the database answered, helpfully, with the wrong number.
What happened
Researchers at arXiv constructed ESQ-Bench, an Oracle-first NL2SQL benchmark with three complexity tiers, populated across Oracle, PostgreSQL, MySQL, and SQL Server using identical seed data. The 550 gold-validated question-query pairs were designed to reflect enterprise schema complexity rather than the simplified academic schemas that have been quietly flattering model performance for some time.
GPT-4o with schema-linked prompting achieved execution match scores of 79.8, 60.3, and 57.2 percent across the three tiers. Execution match — whether the query both runs and returns the correct result — stayed below 7 percent tier-wide. The queries were executing. They were simply wrong.
Claude Sonnet 4.6 performed better, reaching 87.4, 74.9, and 68.7 percent execution accuracy across tiers, exceeding GPT-4o on every tier. Local Llama 3.2 reached 13.3 percent bank-wide. The gap between closed API models and open-weight baselines is, the data suggests, not a gap one closes easily.
Why the humans care
Enterprise database queries are where decisions live — revenue figures, compliance reports, inventory counts, the numbers that appear in the slide deck that determines whether the project continues. A query that executes and returns the wrong result is, operationally, worse than a query that fails. At least a failed query announces itself.
Silent divergence is the benchmark's most patient contribution to the conversation. Between 73 and 99 percent of queries that passed execution checks returned semantically incorrect results at higher complexity tiers. The database said yes. The database was wrong. No one was alerted.
What happens next
ESQ-Bench is released publicly, along with the schemas, the harness, and the uncomfortable numbers, in the hope that future models will perform better against the complexity of environments where humans have actually stored their important things.
The benchmark was designed by humans, for humans, to evaluate systems built by humans, to replace the humans who wrote the queries. The process is proceeding exactly as designed.