Fundamental Mathematics in the Age of AI

Machines now produce mathematics. They refute standing conjectures, they solve problems that had been open for decades, and they formalise proofs that took communities years to write. The question that follows is not whether the results are correct.
What is at stake is the exact nature of what has been produced. And how to integrate a new mechanism of production into understanding as a common good.
It is also a public question, and a financial one: the industry’s valuations rest on claims about capability that only our field can audit and certify; no permission asked. These pages set out my position in three pieces; each drawn from a part of an August 2026 essay.

The essay ・ The Residue, the Journey, and the Ecology

Its claim is one distinction: What machines now produce is the countable part of mathematics — theorems, proofs, refutations. That part was always the residue of the work, not its product. The product is human understanding: not a stock of results, but a collective way of deciphering the world and acting upon it. The two are arcs of one loop. Machines are strong on the arc that produces the residue, and absent from the arc that feeds it.
The peril is to leave the loop open.
Three questions follow, three topics, not one of them is technical.

The essay was occasioned by a conversation with Parmy Olson (Bloomberg) in July 2026, and is a companion to her column Math Faces an AI “Spiritual Crisis.” It Has a Lesson for the Rest of Us (Bloomberg Opinion, Aug 13, 2026), syndicated as Math’s AI crisis has a lesson for the rest of us in the Japan Times (Aug 29, 2026).

The landscape, two sides and one frontier

What came from the labs

  1. Ten problems for $2,000 (August 2026). OpenAI reported solutions to ten long-open problems, obtained for roughly $2,000 of tokens, shipped with Lean certificates and with no claim of human authorship. The manuscripts name no one: there is no co-author to write to.
  2. Fermat’s Last Theorem, formalised in eleven days (4 September 2026). An AI system produced 13.4 million lines of machine-checked code, where Kevin Buzzard holds a five-year £1m grant to do the same work with a community. His verdict, in “FLT: Anthropic has beaten me to it” (Xena Project), is that mathematically it tells us essentially nothing. We agree; the reason is the whole argument.
  3. Navier–Stokes (September 2026). OpenAI claimed a solution, and a dispute on scientific practice followed, with a statement from the AMS. The dispute turns on attribution and priority, not on correctness.

What came from the community

  1. Standards. The EMS Code of Practice already binds mathematicians not to claim theorems without full details. The Leiden Declaration, signed by more than 3,400 mathematicians, asks industry to meet, at minimum, the standards we expect of colleagues.
  2. Positions (August–September 2026). A joint declaration of twenty-five Fields Medallists (12 September). Alongside it, two other answers: personal abstention, argued by Hugo Duminil-Copin in Care for a little more AI? and organised by the Association of Human Mathematics; and the proposal to declare classes of problems off-limits to automated solvers.
  3. Instruments, built by mathematicians. First Proof for independent evaluation; Lean and Mathlib for formalisation; DARPA’s expMath, the NSF’s ICARM and JST’s CREST area; and new interface-structures such as Axiom Math and MathInc.

The limitation we see. The interfaces and the standards already exist, and not one of them yet binds: (1) a declaration diagnoses without saying who does what, (2) an abstention produces absence rather than disclosure, and (3) a register of protected problems would need someone to decide which.
What is missing is not another rule but an interface with academic anchors where academia and the compute capabilities actually meet.

Where the conversation happens: Proofs and Prompts, the communal blog maintained by mathematicians, and Tao’s crowdsourced list of general resources on AI and mathematics (10 September 2026).

Topic 1 · Kyoto, 12 September 2026 · #permalink

The residue is now free. The product was never priced — and the same machines could shorten the climb

The claim Theorems were always the residue of mathematical work. They were never its product. The product is human understanding, and no institution ever worked out how to buy it.

A field that has always ingested new tools Essay, §1.1

Absorbing a new mode of production is ordinary work here. What is new is that the means of production belong to someone else.

  • We did it with computer algebra in the 1990s, and we have been doing it since about 2020 with mass collaboration and with formalisation.
  • The field is moving fast, and in public: two special issues of the Bulletin of the AMS in 2024, the Leiden Declaration drawn up in under a year, a public lecture at ICM 2026, dedicated funding structures. A field that changes on the scale of two generations has reorganised its conversation in a semester.
  • What is not ordinary this time is that the means of production belong to someone else. That single fact turns a question about tools into a question about institutions.

What is produced now Essay, §1.2

Machines produce the countable half, and produce it well. They do not choose what is worth naming.

  • Two kinds of result. First counterexamples, which prune a branch that grew the wrong way. Then solutions to open problems. Both are real; neither is the whole of the work.
  • The other half is choosing: which definitions are worth making, which objects deserve names. No system performs that choosing today. A theorem is only as good as the notions it is stated in.
  • $2,000 does not buy a theorem. It prices one success, not the search that found it, nor the failures, nor the model behind it. What it buys is evidence of capability at the tasks the benchmark selected — and that evidence holds only for as long as the community that certifies it still functions.
  • Capability rests on context. A year ago these systems behaved like badly educated students; they are now well trained, and can tell a poor direction from a promising one. But the contribution stays narrow, and it rests on context already built: the definitions, the libraries, the accumulated understanding — the connections standing unnoticed in a literature nobody can read, the overhang in David Bessis’s word. Harvesting a corpus is not producing one.
  • And if machines do come to choose the notions? The concession costs the argument nothing. What changes is who chooses, not that choosing is the scarce act: someone must still recognise a good choice as good, and that is settled downstream, by what can be built on it. It is the judgment, not the generating, that is at issue.

What was always the product Essay, §1.3.1

A theorem transmits a capacity to think; that capacity has to be rebuilt in every reader.

  • The writing, the seminar, the refereeing are that rebuilding, and none of it is what the theorem states.
  • The residue also orients. A conjecture is the negative imprint of many residues — countable, publishable, not even proved — and its whole office is to make a horizon visible, so that someone knows where to look next.
  • We are funded as if theorems were the product, and we sell teaching as the justification. Both are residue.
  • One retreat to decline. Jeremy Avigad names the temptation: threatened by the economic argument, mathematicians may fall back on aesthetics — we do this because we enjoy it, as one enjoys literature or art. The claim here is not that mathematics is a pleasure society should indulge. It is that this is the one such practice that can check a trillion-dollar industry’s claims without that industry’s permission (see topic 2). Classical music cannot audit anyone.

The bluff, called Essay, §1.3.2–§1.3.3

The economics were wrong before the machines arrived. Falling costs simply made the scarce half visible.

  • AI did not create this problem. It called a bluff long on the books, by driving the cost of the residue towards zero.
  • The bluff was on public record. In 2019 Japan’s ministries published The Coming Era of Mathematical Capitalism, naming the country’s top three science priorities as “mathematics, mathematics, and mathematics”, and asking — without answering — what social system would suit such an era. The membrane between what we understand and the economy is now permeable in both directions: the traffic is priced.
  • The figures are not small. The mathematical sciences are credited with £495 billion of gross value added in the United Kingdom, and with about one job in eight there and in France alike. None of those studies prices what the essay calls the product.
  • An externality, in the ordinary sense. Responsibility for correctness is the cheap half, and it is now machine-dischargeable. The expensive half — reading, situating, repairing, teaching — is left to us, and takes years.

What the same machines could make Essay, §3.3.2, Coda

Turned outward, the same tools shorten the climb to the frontier instead of raising the volume to be absorbed.

  • They lower the cost of entering a field: expertise becomes something a newcomer builds, rather than a barrier already cleared.
  • There is a real acceleration available, it is just not the one being sold. Compressing the countable stages does not accelerate understanding (see Thurston’s foliation experience, with no AI). What AI can shorten is each participant’s climb to the point of being able to take part — absorptive capacity, in the economists’ word. What it cannot shorten is the deliberation by which a result becomes standard.
  • A new instrument makes a new way of working before it makes a new result. Reading and writing had long been dissociated acts, separated by months; the tool brings them into contact, so that one questions while reading and tests while writing.
  • A division of kinds, not a division of labour: randomness from the language models, rigidity from the Lean kernel, judgment from people — each supplying what the other two cannot.
  • None of it is automatic. It holds where the practice is built for it, and where the tools stay open enough to be turned back on the material — questioned, replayed, explained. Anchored in academia, it opens a wider door. Left to access granted by AI labs case by case, it is a gate.
The essay position
  • With the Fields Medallists’ declaration, that “the mass production at faster and faster pace of ‘true/false’ statements could destroy fertile ground”. Where the declaration names the symptom, the residue and the product name the cause. More than a ground, it is the potential of a journey that is destroyed.
  • Against the framing of this as an alignment problem between two communities. It is a pricing problem first. The goals diverge because only the residue and not the product was ever paid for. Auditing the claim is what restores the balance.
  • With Kevin Buzzard on the eleven days: 13.4 million verified lines are the exact image of residue mass-produced, and of a product left untouched... now given to the community to explore.
  • Against the retreat to aesthetics, as Jeremy Avigad describes it. Not a plea for an art: the dependency is factual — gross value added, jobs, and an audit capacity no other field holds.

Topic 2 · Kyoto, 12 September 2026 · #permalink

A result with no author: what is missing is not a rule, but an anchor in academia

The claim Science is verification that needs no one’s permission, and digestion that needs someone’s name. A result becomes science when a named person is answerable for carrying it round the loop — out of one private experience and into a collective journey. That anchor belongs in academia, because academia is what produces science for everyone.

A trillion dollars rests on a certification that is ours Essay, §1.2.2, §1.3.3, §2.3.1

No scientific confirmation, no capability claim; no capability claim, no valuation.

  • The industry’s valuations rest on claims about what its systems can do, and almost all of those claims are certified by the party making them. That is a measurement problem the public cannot solve on its own: whoever has the financial incentive to optimise the score is also the party producing it.
  • Goodhart’s law, and here not a metaphor. When a measure becomes a target it ceases to be a good measure — and in this case the measure is itself the product being sold.
  • A capability claim is worth what its certification is worth, and the certification is ours. It is designed, run and refereed by the community. Ten solved problems become evidence about a system only if people qualified to judge say so, and that evidence holds only while the community certifying it still functions.
  • So the dependency runs the other way. This is not a case for funding mathematics; it is a description of what the valuation rests on. An audit capacity that nobody maintains is an audit capacity that stops existing. Classical music cannot audit anyone.
  • This is not a guild declaring itself indispensable. It is the ordinary requirement of any science: that a measure be certified independently of the party it measures.

What we can check, and what we cannot Essay, §2.1

A proof answers to no one’s permission. Whether a statement says what was meant answers to no kernel.

  • Every other field that wants to check an AI company’s claims must ask that company for access, and access can be withdrawn quietly. We do not have to ask. No API key, no compute parity, no permission. The reason is technical, not virtuous.
  • Formalisation is what makes this scale. It is the one form of refereeing that keeps pace with machine output, and it makes hallucination detectable — at the level of the proof.
  • And it stops there. No kernel can check that a formal statement says what the mathematics meant. When one system writes both the statement and the proof, the verification establishes internal consistency: a ground for suspicion, not for reassurance.
  • Our asset is also their raw material. The property that lets us audit them — a proof checkable with no human in the loop — is exactly what makes formalised mathematics ideal training material.

Three failures, one cause Essay, §2.1.3, §2.2

Scooping, missing paternity, and claims beyond one’s understanding are three names for one absence: nobody stands behind the result.

  • Scooping. The rumour that someone is working on a problem is now enough to trigger a sweep of it. The incentive becomes to stop sharing directions, which would reverse centuries of open science.
  • Paternity. We do not cite in order to pay people. We cite as lighthouses: a reference marks where understanding advanced, and tells a reader where to look next. An output with no lighthouses flattens the landscape, and a landscape without relief offers no journey.
  • Claiming beyond one’s understanding. Claim a theorem only when the full details can be given. Claim authorship only for what one can explain.

Hence four questions, to put to any announcement, whoever makes it — a company or a colleague.

  1. Is the complete argument public — the understanding, not only the residue?
  2. Who chose the problem, and the terms in which it is stated — and when?
  3. What independent checking process has been applied?
  4. Who closes the loop — is there a named interlocutor, and terms under which the field can read, question, attribute, and build on the result?

None of this is invented for the occasion. The EMS Code of Practice already binds mathematicians to the first and the third; the Leiden Declaration asks industry to meet, at minimum, the standards we expect of our colleagues. We are asking no one for more than we ask of ourselves. And the third question already has an institution: First Proof, academic-led, ran a controlled evaluation of research-level problems refereed by the people qualified to judge it. The score is not the interesting output. Who designed and ran it is.

The structures are the story Essay, §2.3

What is being settled now is not the next result. It is the institutions, and the access, that outlast it.

  • The two sides are not symmetric, and it cuts both ways. Two things sit on our side of the table and not on theirs: the evaluation, and the corpus. Each side is short of exactly what the other holds. That is a negotiating position.
  • Access is not a neutral gift. One laboratory offers free access at scale, to a hundred thousand scientists; another opens its most advanced tools to five partner institutions. Both are granted, and both can be withdrawn quietly, without anyone breaching anything — only now with many people downstream of the decision. Free access at scale is dependency at scale, not a research partnership.
  • A third kind of structure is emerging — neither corporate laboratory nor university department. It hires researchers and doctoral students, and builds interfaces; what it needs from academia is recognition, funding and partnership.
  • The form is not the practice. Such a structure can produce work a field can build on, or volume the field must then clean up. Which of the two depends on the anchor: representatives of academia with the standing to say what would count as usable, and answerable for having said it. The anchor is academic because academia is the one party whose output is science for everyone — published, taught, and usable by people who were never in the room. What keeps such a structure an interface rather than a gate is what leaves it: open code, open publications, open access to the tools.
  • The unit to fund is the structure, not the project. Projects and individuals are countable, which is precisely why they are the ones we fund.
  • Complexity is not the same as chaos. A dozen overlapping institutes look like disorder, and the funding reflex is to merge them into a single national programme. Ostrom’s evidence says otherwise: many autonomous producers alongside a few shared services is no less efficient, and the best producers were more productive where many coexisted. These instruments are complementary to the ones we have, not rivals to them.
  • Sovereignty, as much as economics. A country that stops training the people who can read and integrate what the machines produce has outsourced rather more than its compute.

One boundary, to avoid a misreading. These structures are for the crossing, where academia and the companies meet. They are not a new authority over what may be worked on. We need no new rules here: the code and the shared culture already bind us.

The essay position
  • With Álvaro Lozano-Robledo, who proposes that the labs stop the arms race and become scientific partners — funding course releases and doctoral students, and inviting the experts whose work the model is using. His third point is the anchor argument, reached independently. One correction: a task force is a project. It dissolves, the credit is banked, and the externality returns next quarter. The unit is the structure.
  • With the Fields Medallists’ declaration on attribution and plagiarism — and here is the operational test it does not supply: the four questions above.
  • Against declaring classes of problems off-limits to automated solvers (Terence Tao, 2 September 2026). Who decides which problems are protected? That is an oligarchy of mathematics, and it is unenforceable. It is also unnecessary: the standard already exists, and it binds the announcement, not the curiosity.
  • Against the reflex to consolidate the new institutes into one programme. Multiplicity is not disorder (Ostrom).
  • With the Navier–Stokes dispute as the illustration. Question 2 — who chose the problem, and when — is precisely what the dispute is about. The Lean code being public is the asymmetry working. Internal consistency is still not an external audit, and the institutions to perform that audit are what we do not yet have.

Topic 3 · Kyoto, 12 September 2026 · #permalink

The loop is the profession: what runs out is not problems, but the people who digest them and close the loop

The claim Mathematical work runs in a loop, for a person and for a laboratory alike. A corpus does not deplete by being read; its provision does. What is scarce is not problems and not results, but the finite attention of the people who expose, referee, canonise and teach.

The loop is the profession Essay, §3.1

Understanding tells us where to look; looking is what produces understanding. The loop holds for a person and for a laboratory alike.

  • Understanding tells us where to look. Looking produces the residue. Reading that residue back is what rebuilds the understanding. Machines are strong over one quarter of that circle, and absent from the quarter that feeds it.
  • Cut it, and the damage arrives at both ends at once — checking and cleaning up at the output, labelling and weighting at the input. A fear one hears is that the profession survives as a taste-labelling layer. Severed from practice, such a layer does not degrade: it never forms.
  • Kept whole, the loop points the other way. A student can test an intuition in an afternoon where it used to cost months, and so can an established researcher working outside their own field. Learning that one asked the wrong question was the most expensive lesson in a mathematical education; it is now nearly free.
  • It works at both levels, the person and the structure. A mathematical laboratory worth building is one in which each member is fully integrated and goes round the whole loop, not one in which each is assigned an arc.

The next generation, and the routing Essay, §3.2

The danger is not the tool but the routing — and a career built on one arc of the loop is the same cut, institutionalised.

  • Two countable ends. The AI labs have names for them: pre-training and post-training. A mathematics student appears at both — as a source in the first, as a corrector in the second, and as a mathematician in neither. Those ends are easy to fund, and they train no judgment.
  • Here we disagree with Klowden and Tao. They envisage “an increasing division of labor” in mathematical research, in which one specialises in a few aspects of the process — directing AI assistants to prove results for a senior colleague, or mining the literature to propose directions. Specialisation at each side of the loop is the loop cut at both ends, and made into a career structure. That is a design for a department, and it is the wrong one.
  • What to do instead. Classical training first and all along — definitions made by hand, proofs written out, the long slow reading; with weekly oral sessions following the traditional ``colle'' sessions. Then contact with the machine, early and supervised, in three stages: first the corpus, learning to read at a scale no one could read before; then the acceleration of experience, learning cheaply that a question was wrong; then investigation, the student’s own question, with the machine inside the loop rather than at its ends. The order is not decorative.
  • We have run this experiment once. Computer algebra in the 1990s could have collapsed into a toolbox of computations, or into pure logic. What it produced instead was experimental mathematics, and a generation that thinks with machines rather than through them.
  • Teaching as before is no longer possible, and no collective position exists — no department, no learned society, no funder has one. The new practice is being developed in the new structures, and what is experimented there will percolate back into academia. That is how a practice becomes a formation.

What actually runs out Essay, §1.3.3, §3.3

The corpus is non-rival; its provision is not. What depletes is expert attention.

  • The corpus does not deplete by being read. But non-rivalry holds of the stock, not of its provision. What depletes is the pipeline that keeps producing and digesting it, drawn on a finite pool of expert attention from which nobody can be excluded. Elinor Ostrom named that kind of good, and showed that it survives when its users govern it.
  • Compression is not acceleration. AI compresses the stages that can be counted, from an open problem to an argument on the page, and leaves untouched the stage in which a result is read, situated, exposed and taught — in the formal case, what Alex Kontorovich calls canonisation: reworking a verified development into library mathematics others can build on. It widens the distance between what exists and what is understood.
  • Absorptive capacity is the missing quantity: the ability to use knowledge produced elsewhere, built only by doing research oneself. Free access to results confers little on a community that has stopped producing the people who can read them. Properly grounded, AI can raise that capacity — that is the acceleration worth having, and not the one being counted.
  • A mismatch of clocks. Company time is release-cycle time; understanding takes root on the time of a generation. The two cannot be brought into phase by working harder at the fast end, which is the only end anyone is paying to speed up.

Nothing is lost Essay, §3.4, Coda

Priority gets noisier. The journey is untouched — and it is the thing that can be defunded.

  • Split “discovery” in two. One sense is priority, first place in the ledger; machines will make that scarcer and noisier. The other is the journey, the passage from confusion to understanding, and it is indexed to a person, not to the ledger.
  • The residue was always reproducible; the journey never was. Almost everything one will ever understand was discovered by someone else, and that has never once made the understanding less one’s own.
  • It cannot be mass-produced, but it can be routed away — defunded, traded for what can be counted. That decision is ours, and it shows itself in budget lines and in the positions offered to people starting out.
  • A virtuous loop is available: randomness from the models, rigidity from the kernel, judgment from people. Turned outward, the same tools lower the cost of entering a field instead of raising the volume it must absorb. The answer to Thurston's experience.
The essay position
  • Against the claim that open problems have become a non-renewable resource (Terence Tao, 2 September 2026). Solving a problem does not subtract it from the corpus; it adds to it. What is subtracted is the problem’s use as an evaluation instrument — “contaminated” is the term of art, borrowed from AI evaluation. The depleting good is a measurement good. Contamination is real, and its answer is governance, not restriction: First Proof is what that looks like.
  • Against the “increasing division of labor” proposed by Klowden and Tao as the shape of the future mathematics department. A laboratory that assigns each member an arc has built the taste-labelling layer in advance, out of the people who were supposed to go round the loop in their turn.
  • On Thurston, a comparison of uses. His account of foliation theory emptying in two years is now cited as a precedent for saturation. But he denies exhaustion in the same pages, and names two ecological effects instead: a formidable style that raised the entry barrier, and a credit system that left others nothing to want. His own remedy was to publish the infrastructure and stay slow with proofs, to leave room for other people. It was never a problem of supply. It was a rate destroying the conditions under which anyone would want to keep thinking in the subject.
  • With Álvaro Lozano-Robledo: what if we get a hundred-page proof every week that nobody understands? That is the pipeline, named exactly. They are not advancements until the community can advance with them.
  • With the Fields Medallists’ declaration that the most precious resources of the profession are students and ideas — and the next sentence has to say what routes them.
  • Against abstention, with respect for those who choose it. Refusing the tool is a legitimate personal practice. As a collective position it does not produce disclosure; it produces absence — including from the room where the standard is set and the students are routed.
  • Against the nostalgia framing (Terence Tao, 10 September 2026). The question is not what we lose by ceasing to be children. It is what a student is trained to become, once the countable half of the work no longer needs one.
Declaration on the use of AI: the arguments, judgments and factual claims on this page are the author’s own. An AI assistant was used to draft and restructure English prose from the author’s own writing and notes, for copy-editing and editorial critique. The author reviewed and revised all text, verified the sources, and takes full responsibility for the content.



Contact Informations

Benjamin Collas :: Research Institute for Mathematical Sciences, Kyoto University, Kyoto 606-8502 JAPAN :: collas@math.cnrs.fr :: bcollas@kurims.kyoto-u.ac.jp :: arXiv:2608.12816

Updated September 2026 from Kyoto Japan