Newsai5 min read
A cheap model solved open math problems, once Google put a team around it
Ahmad JSep 1, 2026Updated Sep 15, 2026

Google published a list on 31 August that reads like a research lab's annual report. Seven open problems solved in mathematics and theoretical computer science. A cycle-accurate RISC-V processor simulator, written from scratch, that boots an operating system to a shell. Performance patches merged upstream into two widely used C++ libraries. All of it produced by teams of AI agents running for hours or days without a human in the loop.
The version that travelled was that Gemini 3.7 Flash, the cheap and fast tier, did this. Google's own post says something more careful, and the careful version is the more useful one.
Which model produced the seven results#
The seven come from what Google calls the Long Proof pattern, and they are specific enough to check:
Six of the seven were reviewed and confirmed correct by human experts. The seventh, Knuth's Cycles, was formally verified in Lean, which is the stronger check of the two: a proof assistant either accepts a proof or it does not. Five of the resulting papers are on arXiv.
Now the sentence that the summaries dropped.
Google states these results were obtained using Gemini 3.1 Pro, and that three of them, problems 1, 3 and 4, were reproduced with Gemini 3.7 Flash.
That reproduction is what the company calls a first, and it is genuinely one. It is not the same claim as the cheap model having solved seven open problems.
Problem 4 is the one worth your attention if you run models yourself. Quantization is the technique that makes a large model fit in the memory you actually have, and a tighter provable bound on a Hadamard scheme is the kind of result that eventually reaches the tools you use rather than staying in the paper. It is also one of the three the Flash model reproduced.
The benchmark number needs the same care#
Google reports a real move on a hard benchmark, and the credit has to be shared.
The configuration that scored 71% is therefore not the configuration you can run today, which is a distinction worth holding on to whenever a benchmark and a product launch arrive in the same announcement.
The simulator is the result that is unambiguously Flash#
The hard part was not writing the code. Google names the obstacle the "silent execution gap": a window of up to hundreds of cycles in which the microarchitectural state can diverge before anything visibly fails. The agents closed it by holding continuous lockstep co-simulation against a reference simulator whose source code was sandboxed so they could not simply read the answer.
That detail matters more than the headline. An agent that can be graded continuously against a trustworthy oracle is in a different situation from one that has to be judged at the end, and most useful engineering work is the second kind.
What the orchestration is actually doing#
Google's diagnosis of why naive multi-agent setups fail is blunt: "Loosely organized agents quickly go off track, agreeing with other agents' early mistakes and building confidently on flawed ideas." Teamwork is the answer they built to that.
Work is organised into patterns, which are specifications rather than programs: Iterative Coding, Distributed Coding, Long Proof, Self-Verification, Document Review. The model picks one from your prompt.
Strip the specifics and you get an unfashionable claim: the gain came from structure around the model, not from a larger model. Anyone who has watched a single coding assistant confidently build on its own early mistake will recognise the problem being solved.
The cost line everyone will skip#
Two sentences in the post do most of the work of setting expectations.
"Some of these results used higher parallelism than the default."
"The version available on Antigravity balances cost and capability, while still supporting the same workflow."
Read together, those say the published results ran wider than the shipped product runs, and that the shipped product trades some of that away. Teamwork is available as /teamwork-preview in Antigravity on paid plans.
The short version#
A Flash-tier model reproduced three PhD-level results and built a cycle-accurate CPU simulator, which is a real and checkable milestone. It did not solve seven open problems on its own, and the 71% benchmark score came from a Flash and Pro combination that is not shipping yet.
The transferable lesson is not about Gemini. It is that on long-horizon work, adversarial structure around a cheap model bought more than a bigger model would have. If you are choosing between paying for a larger model and building a critique loop around a smaller one, this is the strongest public evidence yet that the second option deserves the first attempt.
Sources
- Google: Pairing Google Antigravity with Gemini 3.7 Flash solves notable multi-agent math and engineering problemsblog.google
- Google Antigravity: Teamwork, When AI Becomes a Research Partnerantigravity.google
- Google Antigravity: Teamwork documentationantigravity.google
- arXiv: the provable Hadamard quantization result (2608.02564)arxiv.org



Discussion