How do data access and errors in the problem set affect the interpretation of scores?
Model / AI: Not specified
⚑verified (audit)
Jan 23, 2025
FrontierMath disclosure
Model / AI: Not specified
⚑verified (audit)
OpenAI had commissioned the benchmark's 300-problem core and had access to statements and solutions, apart from a 50-problem holdout; the funding was disclosed only alongside the headline 25% score, the data-access terms a month later, and Epoch's own April 2025 evaluation of the released model scored far below that headline.
Source review pending. Imported from the original notes; linked claims and artifacts have not been re-audited in this restructuring.
an AI-assisted audit found errors in roughly 42% of the original problem set, where earlier human QA had caught about 5%; v2 corrected 135 problems and removed 12, so any score must name the version it was measured on and v1 and v2 results are not directly comparable.
Source review pending. Imported from the original notes; linked claims and artifacts have not been re-audited in this restructuring.