AWESOME AI PROOFS
← Problems
Competitions / Jul 2025–Oct 2025

IMO 2025 problem set

Which IMO 2025 problems were solved, and how do formal coverage and grading differ across submissions?

Model / AI: Gemini Deep Think, OpenAI (unreleased), Aristotle, Gemini 2.5 Pro, Grok 4, GPT-5

verifiedmachine-checkedself-reporteddisputed
Jul 2025

Gemini Deep Think

Model / AI: Gemini Deep Think

verified

IMO 2025 gold standard (35/42), graded by official IMO coordinators, end-to-end in natural language inside the 4.5-hour contest limit.

Source review pending. Imported from the original notes; linked claims and artifacts have not been re-audited in this restructuring.

Jul 2025

OpenAI at IMO 2025

Model / AI: OpenAI (unreleased)

self-reported

gold claimed; graded by three former IMO medalists OpenAI says it hired, not by the IMO, whose president said it "cannot validate the methods". Proofs public; round-up.

Source review pending. Imported from the original notes; linked claims and artifacts have not been re-audited in this restructuring.

Oct 2025

Aristotle (Harmonic) at IMO 2025

Model / AI: Aristotle

machine-checkedself-reported

gold claimed on five problems; public Lean proofs for 1, 3, 4, 5, while geometry problem 2 is certified only by Harmonic's own solver, and both problem 4 proofs use native_decide. System paper; grading self-reported.

Source review pending. Imported from the original notes; linked claims and artifacts have not been re-audited in this restructuring.

Jul 2025

IMO 2025 by verify-and-refine

Model / AI: Gemini 2.5 Pro, Grok 4, GPT-5

self-reporteddisputed

Huang and Yang report 5 of 6 problems from a model-agnostic pipeline, with Gemini 2.5 Pro, Grok-4 or GPT-5 equally; self-graded, and MathArena disputes the methodology.

Source review pending. Imported from the original notes; linked claims and artifacts have not been re-audited in this restructuring.