Skip to content

Independent Research

Evidence over opinions. Reproducible AI benchmarks.

Reproducible, transparent AI model comparisons. Evidence over opinions. No sponsored rankings. Every experiment documents exact prompts, parameters, and evaluation methods.

Published Benchmarks
5
Categories
12
Models Tested
12+
Sponsored Rankings
0

Mission

Build the most trustworthy public collection of AI benchmark experiments.

Reproducibility

Every benchmark publishes exact prompts, API parameters, execution environment, and reproducibility hashes.

Transparency

Scoring rubrics, human review protocols, and raw outputs are documented. No black-box evaluations.

Independence

No sponsored rankings. No pay-to-win placements. No invented results. Full disclosure of funding.

Featured Experiments

Rigorously designed benchmarks with full documentation.

Coding medium Featured

Python Bug Fix — Null Reference in Async Handler

Four frontier models attempt to fix a subtle async/await bug in a FastAPI endpoint. Evaluated on correctness, test pass rate, and minimal diff size.

Models
4
Top Score
96%
Leader
Claude 3.5 Sonnet
Published
June 12, 2025

Latest Benchmarks

Recently published experiments.

View all →
Coding medium

Python Bug Fix — Null Reference in Async Handler

Four frontier models attempt to fix a subtle async/await bug in a FastAPI endpoint. Evaluated on correctness, test pass rate, and minimal diff size.

Models
4
Top Score
96%
Leader
Claude 3.5 Sonnet
Published
June 12, 2025
Writing medium

Technical Blog Post — Kubernetes Networking Explainer

Models write a 1,200-word technical blog post explaining Kubernetes networking to intermediate developers. Evaluated on accuracy, structure, and clarity.

Models
3
Top Score
88%
Leader
Claude 3.5 Sonnet
Published
June 1, 2025

Latest Report

Market and benchmark-planning reports kept separate from experimental results.

View all reports →

7/16/2026

Mid-Tier AI Models vs Enterprise Java Migrations

A report-driven benchmark proposal comparing Claude Sonnet 5, GPT-5.6 Terra/Sol, DeepSeek V4, and Gemini 3.5 Flash on ScarfBench-style Java migration tasks.

#reports#java#migrations#agents#scarfbench#enterprise