Research Trends 2026-09-21

Epoch AI's New Benchmark Is 68 Unsolved Erdős Problems -- and the Best Model So Far Solved 2

FrontierMath Erdős asks AI systems to write complete Lean proofs for real open problems in mathematics, curated by the mathematician who maintains ErdősProblems.com -- a harder, more honest test than benchmarks with a known answer key.

Epoch AI launched FrontierMath Erdős this month: 68 problems posed or studied by Paul Erdős, open as of August 2026, curated specifically for difficulty and interest by Thomas Bloom, the mathematician who maintains ErdősProblems.com. The task isn't multiple choice or a numeric answer -- each problem is formalized in Lean, and the AI system has to produce a complete, machine-checkable Lean proof or disproof. That's a meaningfully harder bar than most benchmarks clear: there's no partial credit for a plausible-sounding argument, and no possibility of the answer having leaked into training data, because these are currently unsolved problems, not textbook exercises with a known solution.

The first results are exactly what you'd expect from a benchmark designed to actually be hard: no prior model solved any of the 68 problems, and the best current model, GPT-6 Astra, solved 2 -- a 3% score. Each attempt runs with a default inference budget of $300 per problem, which is itself a data point worth noting -- this isn't a benchmark measuring quick inference, it's measuring what a model can do with real reasoning budget on a problem that's resisted human mathematicians.

Context matters for reading the 3% number correctly: FrontierMath Erdős is the hardest of three components under the FrontierMath umbrella (the others are the tiered difficulty benchmark and a broader Open Problems collection), specifically selected to be near the current frontier of what's tractable at all. A near-zero score here isn't a failure of the models being tested so much as confirmation the benchmark is doing its job -- measuring genuine research-level mathematical capability rather than a saturated test everything already scores 90%+ on.

The instructive contrast is with SWE-bench Verified, covered elsewhere this month, which went from 60% to near-100% in about a year -- a benchmark that got solved, not just improved on. FrontierMath Erdős reads like Epoch's answer to that saturation: build a benchmark specifically hard enough that a 3% starting score is honest, and watch that number over the next year as the more meaningful signal of research-level mathematical progress than another coding benchmark inching toward its ceiling.

A 3% score on FrontierMath Erdős from the current best model is the benchmark working as intended, not a disappointing result -- treat this number's trajectory over the coming months as a genuinely informative signal of research-level mathematical reasoning progress, in contrast to coding benchmarks like SWE-bench Verified that have already largely saturated.