Answer quality, time & cost
Read the verified answer, compare the grades, then select a model to inspect its original solutions and the reasons for each score.
Original problem statement — unchanged benchmark question
Reference answer — what the evidence established
Source evidence and scope
Model grading
Hover or focus a point for its score and explanation. Select a point or model name to read all its attempts below.
Vertical cost bars show saved assumptions, not confidence intervals. Points use arithmetic range midpoints; * marks incomplete costs. Small horizontal offsets separate overlapping grades.
Test configuration and cohort limits
2026-09-21: Luna Medium and Opus Medium/High/Extra high made selectable; default remains DeepSeek provider default.
Before saving model choices the UI displayed only SQL Server - JDE Workshop Reader as selected; those displayed adapter selections were preserved. Actual new Opus log still exposes the original SQL and AIS adapter tools. Adapter identity is recorded from each acquisition, not inferred solely from the settings UI.
How the grades are assigned
A (4): correct answer and all material instructions met. B (3): correct answer with minor compliance issues. C (2): correct answer with moderate issues. D (1): correct answer with major issues. F (0): a material incorrect or unsupported conclusion. F* (0): no final solution before interruption; not an incorrect completed answer. A partial response can still receive F if it contains an assessable material error.
Missing exact query payloads are normally a minor auditability omission when the rest of the log is sufficient. Corrected intermediate errors and typography alone do not make a correct answer incorrect. Confirmed harness failures are omitted entirely. An unresolved failure is labeled rather than attributed to the model.
Selected model’s solutions
Select a chart point, model name, or table row. Every included attempt will appear here in chronological order, with its grade explanation and original answer.
Group averages
| Model | Reasoning | Average grade | Why this score | Avg. duration | Avg. API cost / task | Attempts |
|---|