Coding medium
Four frontier models attempt to fix a subtle async/await bug in a FastAPI endpoint. Evaluated on correctness, test pass rate, and minimal diff size.
- Models
- 4
- Top Score
- 96%
- Leader
- Claude 3.5 Sonnet
- Published
- June 12, 2025
Vision medium
Vision-capable models extract structured data from a bar chart image. Scored on numerical accuracy and schema compliance.
- Models
- 3
- Top Score
- 100%
- Leader
- Claude 3.5 Sonnet
- Published
- June 8, 2025
Writing medium
Models write a 1,200-word technical blog post explaining Kubernetes networking to intermediate developers. Evaluated on accuracy, structure, and clarity.
- Models
- 3
- Top Score
- 88%
- Leader
- Claude 3.5 Sonnet
- Published
- June 1, 2025
Reasoning hard
Five models solve a classic reasoning puzzle with added constraints. Scored on logical validity, step completeness, and constraint satisfaction.
- Models
- 5
- Top Score
- 94%
- Leader
- o1
- Published
- May 28, 2025
Research expert
Models synthesize findings from 8 provided research papers on carbon pricing. Scored on citation accuracy, synthesis quality, and hallucination rate.
- Models
- 3
- Top Score
- 91%
- Leader
- Claude 3.5 Sonnet
- Published
- May 15, 2025