Answer quality, time & cost
Read the verified answer, compare the grades, then select a model to inspect its original solutions and the reasons for each score.
Benchmark question
Reference answer — what the evidence established
Source evidence and scope
Model grading
Hover or focus a point for its score and explanation. Select a point or model name to read all its attempts below.
Vertical cost bars show pricing assumptions, not confidence intervals. Points use arithmetic range midpoints; * marks incomplete costs. Small horizontal offsets separate overlapping grades.
How the grades are assigned
A (4): correct answer and all material instructions met. B (3): correct answer with minor compliance issues. C (2): correct answer with moderate issues. D (1): correct answer with major issues. F (0): a material incorrect or unsupported conclusion. F* (0): no final solution before interruption; not an incorrect completed answer. A partial response can still receive F if it contains an assessable material error.
Missing exact query payloads are normally a minor auditability omission when the rest of the log is sufficient. Corrected intermediate errors and typography alone do not make a correct answer incorrect. Confirmed harness failures are omitted entirely. An unresolved failure is labeled rather than attributed to the model.
Selected model’s solutions
Select a chart point, model name, or table row. Every included attempt will appear here in chronological order, with its grade explanation and original answer.
Group averages
| Model | Reasoning | Average grade | Why this score | Avg. duration | Avg. API cost / task | Attempts |
|---|