running  ·  openenv  ·  v4.0.0

SQL Repair
Environment

A production-grade RL environment where AI agents diagnose and fix broken SQL queries through multi-turn agentic interaction with a live database. 20 tasks · 7 tables · Diagnostic mode · Execution diffs.

v4.0.0 · sql_repair_env
Tasks
20
3 difficulty tiers
DB Tables
7
SQLite in-memory
Bug Categories
8
syntax → semantic
Max Attempts
5
+ free diagnostics
Reward Range
0–1
partial credit
✎ What makes this environment exceptional
🔍
Diagnostic Mode
Agents run free exploratory SQL queries to understand the database before fixing — genuine multi-turn agentic behaviour. No other SQL environment supports this.
📊
Execution Diffs
Every step shows the agent exactly what its query returned vs what was expected. Row-level diffs make errors transparent and learnable.
📈
Progress Rewards
A bonus reward is added whenever the agent improves from its previous best score. Prevents reward flattening and encourages convergent behaviour.
🏷️
Bug Categories
8 structured bug categories enable curriculum learning — train on syntax first, then JOIN logic, then hard semantic bugs like self-joins and duplicate counting.
🛡️
Anti-Hack Grader
Penalises submitting the original broken query unchanged or exact duplicate submissions. Forces agents to genuinely reason rather than exploit the reward.
✎ Tasks
Loading tasks…
✎ API + Reward function
API Endpointsport 7860
GET/healthhealth check
POST/resetstart episode
POST/stepsubmit action
GET/stateepisode meta
GET/taskslist 20 tasks
POST/gradergrade without episode
GET/baselineoracle scores
GET/infoenvironment info
GET/docsswagger UI
Reward Function0.001 – 0.999
+0.30
query executes without error
+0.20
correct columns returned
+0.10
correct row count
+0.40
all values match exactly
+bonus
progress bonus (score improved)
×0.85
all attempts exhausted
→0.001
unchanged or duplicate query
✎ Database schema — 7 tables
TABLE employees (id, name, department, salary, hire_date, manager_id, status)
TABLE departments (id, name, budget, location, head_id)
TABLE projects (id, name, department_id, budget, status, start_date, end_date, priority)
TABLE employee_projects (employee_id, project_id, role, hours_worked, start_date)
TABLE sales (id, employee_id, amount, sale_date, product, region, quarter)
TABLE products (id, name, category, price, stock, supplier)
TABLE performance_reviews (id, employee_id, year, rating, reviewer_id, notes)

→ employees.department = departments.name
→ employees.manager_id → employees.id (self-join)
→ departments.head_id → employees.id
→ projects.department_id → departments.id
→ employee_projects.employee_id → employees.id
→ employee_projects.project_id → projects.id
→ sales.employee_id → employees.id
→ performance_reviews.employee_id / reviewer_id → employees.id
✎ Grader sandbox
try the grader — instant score + diff without an episode
Broken query
Select a task above to load the broken query.
Your fixed SQL
✎ Baseline oracle — all 20 tasks
oracle agent submits the known-correct query for all 20 tasks
Click "Run baseline" to fetch oracle scores for all 20 tasks.