AWESOME AI PROOFS
← Problems
Research / Jan 23, 2025–Jun 12, 2026

FrontierMath evaluation integrity

How do data access and errors in the problem set affect the interpretation of scores?

Model / AI: Not specified

verified (audit)
Jan 23, 2025

FrontierMath disclosure

Model / AI: Not specified

verified (audit)

OpenAI had commissioned the benchmark's 300-problem core and had access to statements and solutions, apart from a 50-problem holdout; the funding was disclosed only alongside the headline 25% score, the data-access terms a month later, and Epoch's own April 2025 evaluation of the released model scored far below that headline.

Source review pending. Imported from the original notes; linked claims and artifacts have not been re-audited in this restructuring.

Jun 12, 2026

FrontierMath v2

Model / AI: Not specified

verified (audit)

an AI-assisted audit found errors in roughly 42% of the original problem set, where earlier human QA had caught about 5%; v2 corrected 135 problems and removed 12, so any score must name the version it was measured on and v1 and v2 results are not directly comparable.

Source review pending. Imported from the original notes; linked claims and artifacts have not been re-audited in this restructuring.