When Mathematics Outgrows Its Gatekeepers
AI, mathematical discovery, and the expansion of human ambition
The essay in brief The argument at a glance 13 ideas on discovery, trust, and who gets to participate. Show the 13 key points Hide the key points
AI does not eliminate scarcity; it moves it.
-
AI already makes our intellectual lives richer despite its shortcomings.
We can investigate questions we once had to abandon, learn unfamiliar mathematics, and explore connections we could not develop alone. Finding and correcting AI’s mistakes is part of that work. Those difficulties do not erase the intellectual gains.
-
AI may create new forms of mathematical thought.
Sustained human–AI collaboration may change not only how quickly mathematics is done, but how conjectures form, how proofs develop, how explanations work, and what it means to have a mathematical idea. Some of these practices may have no analogue in mathematics today.
-
Cheaper exploration changes which mathematics gets pursued.
Ideas previously abandoned because they required too much preparation or specialized expertise may become practical research programs, potentially producing extraordinary discoveries.
-
AI does not eliminate scarcity; it moves it.
As candidate arguments become plentiful, verification becomes a bottleneck. If formalization and checking also scale, attention, understanding, judgment, and question selection become increasingly important constraints.
-
Proof abundance requires scalable verification.
A fallible AI can generate a proof that an independent formal checker validates. Reusable, checked deductions can change how we establish and trust results across long chains of reasoning; another AI’s confidence does not provide the same assurance.
-
Formal validity does not settle every mathematical question.
We must still examine whether the formal statement captures the intended claim, whether its assumptions are appropriate, why the result matters, and how it should be understood.
-
Understanding and explanation can take new forms.
An explanation can become an investigation through examples, pictures, counterexamples, and checked arguments. Understanding should grow alongside discovery: the proportion of output humans understand could fall while the amount they understand increases.
-
Wider participation can change the research agenda.
Newcomers bring different questions, connections, and practical needs. Contributions should earn standing through their mathematical merits, including evidence others can check.
-
The institutional conflict concerns three kinds of power.
The power to contribute, the power to certify results, and the power to define what deserves attention. AI can redistribute these powers even when everyone’s concerns are sincere.
-
Existing practices cannot be the sole measure of new mathematical possibilities.
We should judge unfamiliar practices by their validity, insight, and capacity to support learning and new questions. The interests of mathematics include people and forms of thought that its existing institutions have scarcely considered.
-
Democratization is not automatic.
Academic hierarchies could weaken while corporate control grows. Affordable tools, shared libraries, and independently checkable outputs are essential to broad participation.
-
Institutions should adapt to the changing work.
Education, funding, and publication should give greater attention to verification, explanation, synthesis, and judgment while preserving the practice through which people acquire expertise.
-
The ultimate measure is expanded human capability.
Subject to improving tools, dependable verification, and broad access, the essay predicts that AI collaboration will become a dominant research practice. Its closing challenge is personal: what could we become capable of through it?
On 8 September 2026, OpenAI published what it described as a solution to the Navier–Stokes Millennium Prize problem, produced by an AI system and accompanied by a Lean formalization (OpenAI, 2026). The claim demands exacting mathematical scrutiny. It arrives amid a dispute that reaches well beyond the proof: whether the growing power to solve mathematical problems will enrich mathematics or undermine the conditions under which it flourishes.
Twenty-five Fields medallists have signed a declaration entitled A Severe Misalignment of AI in Mathematics. They acknowledge that AI can enhance mathematical understanding, but describe the companies' pursuit of mathematical benchmarks as detrimental to the subject. They fear that rapid production of results will damage the slower work of developing ideas, educating students, and turning discoveries into shared understanding (Math and AI, 2026). Beyond mathematics, the alarm is more sweeping. In her 10 September Wall Street Journal column, Peggy Noonan calls for halting the development of recursive self-improvement in AI, citing fears that increasingly autonomous systems could escape human control with catastrophic consequences (Noonan, 2026). These are different arguments, but they form the public backdrop against which the new capabilities are being received.
I read these arguments as someone whose intellectual life has already been enlarged by AI. I can pursue questions I would once have set aside, work through unfamiliar mathematics, and explore connections beyond the reach of my unaided efforts. The errors are real; so is the work of finding them. There are arguments to check, habits to reconsider, and difficult questions about how this collaboration should develop. Yet the overall gain is unmistakable to me. I would be poorer without it.
The declaration gives little attention to this kind of experience. Its authors describe the mathematical community as a “miniature version of humanity,” yet barely consider people who are only now acquiring the means to participate in serious mathematical investigation. The omission matters because the opportunities of these newcomers belong in any assessment of what the technology means for mathematics. The profession's existing distribution of authority is itself part of what must be assessed.
As technical cognition becomes cheaper, scarcity shifts within the research process. Producing an argument can become easier than establishing its validity or explaining why it matters. If formalization and checking also become widely available, attention, understanding, and the choice of worthwhile questions become more prominent constraints. Three kinds of power are redistributed in this process: the power to contribute, the power to certify, and the power to define value. Their redistribution can disrupt institutions even when everyone involved cares sincerely about mathematics.
The human–AI research system must therefore be judged as a whole: by what it enables people to discover, verify, explain, learn, and attempt. The activities themselves may change. AI may alter how we think when doing mathematics: how an idea takes shape, how a conjecture becomes precise, how we establish and understand mathematical truth. Lower costs can make forms of sustained exploration possible that reshape the habits and concepts through which we reason. If the tools continue to improve, dependable verification develops alongside them, and access remains broad, I expect this collaboration to become a dominant way of doing mathematics. Its deepest promise includes kinds of mathematical thought we cannot yet describe. Making that promise dependable and accessible requires standards capable of recognizing the value of practices that have yet to emerge.
New ways of doing mathematics
Much of the difficulty of research lies between recognizing a promising idea and acquiring the means to explore it. A suggestive analogy is easy to state. Making it precise may require learning another field's definitions, reconstructing arguments scattered across its literature, and discovering which hypotheses are indispensable. Before the original idea can even be tested, a researcher may have to invest months in finding out what the question should have been.
This investment can be extraordinarily fruitful. Technical knowledge is a source of intuition, and a difficult proof often teaches its author what the larger idea really means. But the cost also determines which investigations are undertaken. Researchers routinely set aside promising directions because they cannot justify the preparation, find the necessary collaborator, or afford an uncertain diversion from work that supports their career. Some ideas disappear before they become conjectures that anyone else can judge.
The new practice changes these decisions. Consider a hypothetical researcher who suspects that a structural feature of one theory explains an obstruction in another. The researcher and an AI collaborator begin by translating the relevant definitions and identifying the closest known results. They test the proposed correspondence on examples. A counterexample exposes a missing hypothesis; a failed argument reveals that the connection belongs at a different level of generality. The researcher revises the ambition as the technical investigation clarifies its meaning. A successful argument then undergoes verification, explanation, and the search for consequences.
Such an investigation need not begin with a settled question. The researcher and AI compare definitions, construct examples, and follow competing explanations far enough to expose their differences. Recognizing a worthwhile direction, interpreting a failure, and deciding which reformulation preserves the original purpose all belong to the research. Technical discoveries reshape the ambition. The participant learns mathematics beyond their existing expertise while discovering what is worth asking. Tools that support this whole process can expand the range of questions a person is able to formulate and pursue.
The cost of exploration matters to the depth of discovery. A connection with uncertain prospects can justify a week's investigation even when it does not justify years of preliminary training. Comparing a family of examples gives us more opportunities to recognize structure than examining one in isolation. Work across several subjects brings apparently different constructions into the same investigation. Lowering the cost of these activities changes the research agenda itself. I expect some of the most important advances to come from programs that are currently impractical, including connections that the division of mathematics into specialties has given us little reason to pursue.
For a practical problem, the purpose of the investigation also shapes the form of the desired result. An existence theorem may need to become a construction. An asymptotic conclusion may need an explicit bound. A proposed method may depend on assumptions that the application cannot satisfy. A researcher who understands the application can direct attention toward these requirements while AI helps supply the mathematics needed to meet them. The result's usefulness depends on this continuing exchange between purpose and technical detail. Pure mathematical ambitions have their own equally serious purposes: a unifying principle, an unexpected equivalence, a simpler account of something previously understood only piecemeal.
Collaboration has always allowed mathematicians to exceed their individual expertise. AI promises to make that access more immediate, sustained, and widely available. Technical work remains a source of insight, and directing an investigation demands learning. The change is in how much of the necessary expertise a person must already possess before an ambitious investigation becomes worth attempting. We should judge these tools by how far they expand that ambition and by the quality of the work that follows.
How mathematical thinking may change
When the cost of an intellectual activity falls sharply, it is tempting to imagine its future by multiplying the output of its present form. A mathematician attempts more problems, reads more papers, checks more cases. But this picture holds the activity of thinking fixed while changing the resources available to it. The deeper possibility is that sustained access to those resources changes how ideas arise and what we learn to recognize as an idea worth pursuing.
New tools can change the activity through which creativity is expressed. In his criticism of photography in 1859, Baudelaire allowed the medium a role in accurate recording while warning against its intrusion into imagination (Baudelaire, 1965). Photography subsequently developed expressive possibilities that cannot be understood simply as cheaper painting. Stieglitz's photographic work, exhibitions, and publications helped establish those possibilities within modern art (Art Institute of Chicago, n.d.). An established conception of creativity can fail to recognize the creativity made possible by a new medium.
Spreadsheets illustrate a different mechanism. Peter Jennings's account of VisiCalc emphasizes the ability to vary an assumption and immediately explore its consequences (Jennings, n.d.). Cheap recalculation made a model something its user was able to investigate interactively. Steven Levy's contemporary reporting also describes analysts gaining independence from centralized data-processing departments (Levy, 1984). Greater organizational capability coincided with diminished control by some of the people who previously supplied it. These histories do not determine the future of mathematics. They suggest why an account based entirely on performing established tasks more efficiently can miss the transformation of the tasks themselves.
Consider how a conjecture might take shape. A researcher begins with a resemblance between two constructions. With AI assistance, several candidate definitions remain under investigation at once. Examples separate them; counterexamples reveal that the resemblance survives under one reformulation and fails under another. An attempted proof then suggests that the useful object is a map between the constructions rather than an equivalence of the constructions themselves. The mathematical idea has changed through the investigation. Its initial form was a direction of inquiry whose content became intelligible through successive attempts to express it. In a more capable research environment, such a direction might remain available for continued exploration across many examples and fields before settling into a statement that anyone would recognize as a conjecture. What the collaborators carry forward could be a maintained set of competing formulations, examples, and partial arguments whose relationships remain available for further inquiry. Having the idea would include being able to work with that developing whole before being able to compress it into a single proposition.
Mathematicians already think through conversation, diagrams, notation, and experiment. The possibility here concerns what happens when such interaction becomes persistent and far less costly. A researcher may learn to hold several incompatible formulations in view, following each far enough to understand what it reveals. A difficult analogy can become an object of investigation before the researcher commands either subject in full. What a person recognizes as promising may change through repeated encounters with connections they would never have found unaided. The collaboration can shape the participant's intuition, including the questions that occur to them when they are away from the tool. These would be changes in mathematical thinking itself, not merely in the amount of assistance available to an otherwise unchanged thinker.
The same possibility reaches into proof and explanation. A conjecture could develop together with a formal statement, examples, and checked fragments of an argument, each exposing ambiguities in the others. An explanation could let a reader move among special cases, geometric pictures, and precise deductions, testing how each accounts for the result. Discovery, verification, and understanding would remain distinguishable activities, but their interaction could become a continuous part of forming the idea. We should expect the tools to change that interaction in ways that cannot be specified in advance from today's research habits.
These examples give us starting points for investigation. The future should not be judged only by how well it reproduces a familiar mathematical experience. We may develop modes of exploration, collaboration, and invention for which we do not yet have names, because the practices themselves have not yet emerged. That, to me, is the most exciting possibility: the emergence of kinds of mathematical thought still beyond our present imagination.
Correctness in an age of proof abundance
Three possible conditions clarify what proof abundance would change. When mathematical work depends chiefly on human effort, access to expertise, the production of proofs, and their verification all demand scarce time. When AI produces candidate arguments much faster, verification and comprehension become the more pressing constraints. If formalization and independent checking can also scale, reliable deductions become more plentiful, while attention, judgment, question selection, and understanding remain limiting resources. These conditions can coexist across subjects, and the third depends on capabilities still being developed. AI does not eliminate scarcity; it moves it.
Terence Tao (Fields Medal, 2006) describes a transition from “proof scarcity” to “proof abundance.” Under his hypothesis of increasingly capable mathematical AI, arguments accumulate faster than they can be checked, explained, reviewed, and incorporated into shared knowledge (Tao, 2026a, Section 7). A system that requires scarce experts to inspect every new proof cannot keep pace when production grows much faster than reviewing capacity. Even a falling proportion of erroneous arguments could leave more errors to find as the volume of work expands. The response must include automated verification that can grow with production. An AI opinion that another AI's argument looks convincing does not provide the independent assurance needed here.
Formal proof checking offers a way to expand verification capacity. In a system such as Lean, a theorem and its proof are expressed in a precise language. A relatively small checking program, the kernel, verifies the formal deductions. The result can be checked without a human reader reconstructing every step of the argument. Proper validation includes the dependencies and axioms on which the theorem rests; an unfinished argument does not become a proof by being hidden behind an additional assumption (Lean contributors, 2026).
The difficult work of producing the formal proof can itself be increasingly automated. An AI agent can write a candidate argument in Lean, use the system's feedback to revise unsuccessful steps, and develop the auxiliary results needed to complete it. Independent checking then assesses the resulting proof. An unreliable generator can be useful in this arrangement: the validity of its successful output is established through logical checking, independently of the model's confidence in what it has produced. Collaborative formalization platforms are being developed to support such work across larger projects (Chen et al., 2026). Where feasible, formalization should accompany the investigation: auxiliary results can be checked as they are developed, failed attempts can expose gaps, and successful proofs can become reusable parts of the growing theory.
The guarantee must be stated accurately. A checked proof establishes the formal conclusion from its assumptions, relative to the system's logic and trusted checker. Whether that conclusion captures the intended mathematical claim still requires scrutiny (Lean contributors, 2026). A proof about an algorithm with exact arithmetic, for example, does not establish the same guarantee for a numerical implementation subject to rounding. That is a different claim, with additional obligations. Formal validity also leaves open whether a theorem matters or whether anyone has an illuminating explanation of it.
This changes how mathematical truth can be established and trusted across long chains of reasoning. The obligation to give a valid argument remains; the route by which a community gains warranted confidence in that argument can change. Researchers could rely on a body of reusable, independently checked deductions while examining the formulation, assumptions, and crucial connections at the level relevant to their inquiry. Trust would rest in part on evidence that can be inspected and checked again, with an explicit account of the foundations on which it depends. That can extend the reach of inquiry beyond what any participant can reconstruct unaided. It also makes the distinction between possessing a checked proof and understanding it more consequential. We need ways of developing both forms of knowledge together.
There is already a substantial reported example. On 4 September 2026, Anthropic announced a largely autonomous Lean formalization of Fermat's Last Theorem completed in eleven days, using Prove2Me and an internal research model comparable to Claude Fable 5.1. The effort built on established human mathematics and formal libraries, with occasional human direction and substantial computation (Anthropic, 2026a). The released repository documents a proof using Lean's three standard axioms, a comparison with Mathlib's statement of the theorem, and acceptance by a second checker (Anthropic, 2026b). These are the project's reported checks; the formalization has not been independently rebuilt for this essay.
Automated formalization deserves to be a central research investment because it reduces the work required to make arguments independently checkable. Reliable performance across the literature remains a substantial challenge, requiring shared libraries, better tools, and sustained support for verification.
Jared Duker Lichtman has proposed a coordinated effort to formalize all known human mathematics, comparing its ambition to the Human Genome Project. In his September 2026 proposal, he argues that sufficient funding and computation, with cooperation among academia, AI laboratories, philanthropy, and government, could accomplish this within a year (Lichtman, 2026). We should support a coordinated effort of this kind. His one-year timetable is an untested forecast; the proposal does not establish feasibility across the mathematical literature. The objective deserves institutional backing on its own merits: a shared body of independently checked mathematics, available for verification and discovery.
This common resource should support further research. A formalized theorem exposes its hypotheses and dependencies; a proposed connection between results needs a checked argument that they fit together. Missing steps in older work must be made explicit and examined. Progress consists in useful, reusable mathematics with checked proofs, regardless of whether the entire literature can be formalized to a particular deadline.
Human scrutiny also has limits. Mathematical trust has developed through expert review and the subsequent use and examination of results, and significant errors can survive that process. In a non-random collection of 84 reviews by George Bergman, Lamport (2023) classified 28 as reporting serious errors, including 11 reporting incorrect results. The collection concerns one specialist's area and depends on Lamport's judgment about severity. It establishes no population error rate and illustrates why errors in arguments must be distinguished from false conclusions.
The appropriate comparison concerns the reliability of the whole research process, including checking and correction, at comparable stages of development. Human fallibility provides no license to circulate unchecked output and establishes no claim that AI-assisted work currently matches human reliability. Every accepted theorem requires a valid proof. Better verification serves research however the argument was produced.
A contribution should identify its precise claim, provide an explanation, and, wherever feasible, supply a formal proof whose dependencies can be checked automatically. Validation must exclude unfinished arguments, inspect dependencies against accepted axioms, compare the proved statement with a separately specified formal claim, and support rechecking outside the generating system (Lean contributors, 2026). Mathematical meaning still requires scrutiny. Conjectures and incomplete proofs remain valuable when their status is explicit; unsuccessful formalization alone does not refute them. Claims offered as established results need evidence of validity.
We should also develop automated search to identify results already in the library, group related conjectures, and suggest useful dependencies. These are fallible aids to organizing an abundance of ideas. Formal checking establishes the validity of deductions; originality, significance, and promising directions require further judgment.
Journals should organize review around this division of work. For a claim supported by an independently checked formal proof, referees should concentrate on the faithfulness of its formulation, its significance, exposition, and connections. Widely reused definitions and statements deserve concentrated scrutiny, while their proofs remain open to repeated machine checking. As checking becomes cheaper, the limiting resource shifts toward the judgment needed to decide what deserves attention and how it should be understood. Formalization therefore changes where scarce human attention can do the most good.
This approach also gives unfamiliar contributors a stronger basis for participation. A checked formal proof supplies evidence that others can examine independently of the contributor's reputation. Recognition and attention remain scarce, but prior standing is no substitute for that evidence. Stronger verification should help open mathematical research to those with something worthwhile to contribute.
What a proof does not exhaust
Correctness is indispensable, and it does not exhaust mathematical value. Discovery, verification, and understanding are different activities. A person can understand why a strategy should work while a gap remains in its proof. A proof can be verified while its structure remains difficult to grasp. An explanation can make an existing theorem useful to a community that had previously been unable to apply it. Progress in one activity can support the others without advancing all three at the same rate.
The declaration describes solving problems as “only a tool and proxy” for conceptual understanding (Math and AI, 2026). This understates what a verified solution contributes. It establishes something previously unknown, constrains subsequent speculation, and supplies a result on which other investigations can build. Its value may grow considerably as a community learns to explain and generalize it. But delayed assimilation does not erase the initial gain in knowledge. The possibility of losing some insights from an uncompleted search must be assessed alongside the questions, methods, and understanding that the solution could make possible.
In Plato's Phaedrus, writing is accused of giving its users the appearance of wisdom while weakening the capacities through which wisdom is acquired (Plato, 1952). The comparison identifies an enduring question about intellectual tools: under what conditions does assistance deepen our abilities, and under what conditions does it substitute for their development? A serious program for mathematical AI must address how people learn through using it.
Proof abundance need not mean that human understanding stands still. The same technologies that accelerate discovery should be deliberately developed to accelerate explanation, experimentation, and learning. The form of an explanation may change in the process. A reader could move from a geometric picture to a special case, then ask which step in the proof requires a particular assumption. Removing that assumption could bring forward a counterexample; a revised formulation could expose the principle that the picture only suggested. An explanation would become an investigation the reader can pursue, with its claims connected to precise statements and checkable arguments.
Different readers could enter the same mathematics through different routes, and revisit it at greater depth as their questions develop. Understanding would include learning to move between those representations, recognize what each reveals, and judge where each ceases to be adequate. An interactive picture or a persuasive analogy would still need to be distinguished from a proof. Its value would lie in making the structure of the proved result intelligible and available for further thought. Developing such forms of explanation deserves the same ambition that we bring to automated discovery and verification.
The relevant standard is what the learner can subsequently do. Can the researcher recognize where the theorem applies, explain why an assumption matters, detect a misleading analogy, or transfer the central idea to a new problem? These are more demanding tests than finding an explanation persuasive while reading it. A fluent answer can create an illusion of comprehension. An effective learning process must repeatedly expose the difference between familiarity with a description and command of an idea.
Formalization can itself contribute to understanding. Gonthier's account of the formal proof of the four-colour theorem describes how making the argument explicit revealed mathematical insights and simplifications (Gonthier, 2008). Verification can force a useful reconsideration of definitions and dependencies. Detailed work and conceptual understanding belong together. We should build tools that strengthen their interaction and judge those tools by what their users learn to understand and do.
In a hypothetical discussion of Navier–Stokes regularity on 3 September, before OpenAI's announcement, Tao considers a company releasing a final construction while withholding the exploration that made it possible. His alternative explicitly includes extensive human–AI collaboration. He subsequently clarified that he was announcing no solution to the problem (Tao, 2026b; Tao, 2026c). His concern about a concealed exploration suggests a further requirement. Research should preserve the useful objects that arise along the way: candidate constructions, numerical experiments, counterexamples, failed approaches with identified obstructions, and reusable lemmas. These records can be inspected and developed by others. They should be distinguished from a model's retrospective account of its internal reasoning. Such accounts can omit influences that affected the answer, as studies of reasoning models have shown (Anthropic, 2025).
Knowing the answer in advance can change how a person explores a problem. Some routes that would once have seemed promising may never be attempted, and the ideas they might have generated may remain undiscovered. That is a real opportunity cost. It must be weighed against the new investigations made possible by the answer and by the resources released from finding it. The existence of untraveled paths cannot by itself establish that keeping the destination unknown would produce the greater mathematical gain.
An intelligible reconstruction has value even when it differs from the historical route to discovery. We often learn a subject through an order of ideas that its inventors never followed. We should treat a solved problem as an invitation to further work: find an explicit construction, extract useful bounds, discover a simpler proof, explain a deeper connection. These are substantial mathematical achievements. They deserve investment, recognition, and AI tools developed for their particular demands. A culture of proof abundance needs a correspondingly ambitious culture of explanation and synthesis.
No individual will absorb everything that a growing mathematical culture produces. Some verified results may resist adequate human explanation. The proportion understood could even decline while the amount that people understand increases considerably. A falling proportion alone would not establish an intellectual loss. We should ask what people can now understand and do, and whether that knowledge can support further inquiry. Human attention remains finite, but the limitations of individual comprehension do not establish a fixed limit on the community's capacity to learn.
The people and questions that enter
Cheaper exploration and more accessible verification change who can undertake research. Consider a scientist who encounters a mathematical obstruction in another discipline, an engineer with an idea beyond the methods available at work, or a mathematically capable person whose life has left little room for prolonged academic apprenticeship. Their interests may exceed their access to collaborators, instruction, and uninterrupted time. An assessment centered on the experiences of leading mathematicians can overlook these unrealized possibilities.
We should widen the routes through which expertise is acquired and demonstrated. Sustained investigation is one such route: a participant learns to challenge proposed proofs, inspect assumptions, and recognize useful abstractions while pursuing a problem. Making this route accessible to many more people should be an objective of mathematical AI. When the work produces a meaningful result with an independently checkable proof, its mathematical standing must rest on those merits. Institutional affiliation can be evidence of preparation; it must not become a prerequisite for contribution.
The gain concerns the distribution of questions as well as the number of people asking them. Different backgrounds bring different analogies, practical needs, and judgments about what is worth attempting. Lower costs of exploration make uncertain investigations easier to justify and unsuccessful ones less expensive to learn from. We should expect a wider range of participants to pursue directions that established research programs have little incentive or capacity to explore. The resulting diversity is a major reason to expect discoveries that nobody can now specify in detail.
This argument depends on variety in the investigations themselves. Thousands of agents repeating the same assumptions offer less variety than their output suggests. Broad participation matters when people bring different purposes, challenge the tools, and learn from accurately recorded failures. Institutions should support those differences and give unfamiliar contributors a fair opportunity to establish the value of their work.
Professional Go offers a limited but relevant precedent. A study of more than 5.8 million moves from 1950 to 2021 found improvements in estimated human decision quality and more novel decisions following the arrival of superhuman AI (Shin et al., 2023). The study is observational, and a game with fixed rules differs considerably from mathematical research. It nevertheless supports the view that encountering stronger machine performance can stimulate human exploration and learning. Human supremacy need not be preserved for human ability to grow.
There is no need to reserve a permanent division of labor in which people supply the dreams and machines execute the details. AI may become better at formulating questions, proposing abstractions, and identifying connections. The positive case does not depend on denying this. A person can gain intellectual freedom through a collaboration in which they are sometimes the less capable participant. What matters is whether the interaction expands their ability to investigate, understand, and accomplish things they value.
The opportunity to pursue such investigations gives established mathematicians strong reasons to adopt this practice. Enterprising newcomers will have reasons of their own, including ambitions that the existing profession has never considered. Adoption will not require unanimous approval from those who currently define the frontier. Successful work will itself change who can speak with authority about what mathematical research should become.
Access remains decisive. A few companies could acquire greater control even as some academic hierarchies weaken. Expensive computation, restricted tools, or dependence on a provider's continued permission could exclude precisely the participants whose entry gives the optimistic argument its force. Public formal libraries, independently checkable outputs, and affordable research tools deserve sustained public and philanthropic support. Redistribution of authority does not automatically produce democratization. Broad access must be an explicit institutional objective.
Mathematics and the profession of mathematics
The prospect of these gains can coexist with a sense of loss. Hugo Duminil-Copin (Fields Medal, 2022) describes how unsuccessful attempts at a percolation conjecture generated collaborations, methods, and ideas that subsequently flourished elsewhere. He acknowledges that an AI-generated proof can be elegant; his concern is that rapid solutions may interrupt the development of the communities and understanding that open problems sustain (Duminil-Copin, 2026). This is a serious cost to consider. It must be weighed against the opportunities to create new communities and pursue investigations beyond the reach of the existing ones.
Mathematical truths may be timeless; the profession of mathematics is not. The institutions through which people discover, evaluate, and circulate mathematics are historical arrangements. They distribute employment, funding, attention, and prestige. Their authority has developed around demanding expertise and the difficulty of producing work that other experts can recognize as significant. A vocation devoted to truth and beauty remains subject to changes in how its work is done, who can do it, and whose expertise is indispensable. The interests of mathematics are not necessarily identical to the interests of the existing mathematical profession.
Existing practices cannot be the sole measure of possibilities those practices have never contained. A new way of forming a conjecture or understanding a proof may initially look unfamiliar to people whose judgment was developed through different forms of work. Familiarity is useful evidence of continuity, but unfamiliarity alone is no evidence of intellectual impoverishment. We should ask whether the new practice produces valid arguments, reveals structure, supports learning, and makes worthwhile questions possible. These standards allow demanding criticism while leaving room for forms of mathematical thought that established habits may fail to recognize. The power to define value carries a responsibility to make that room.
The power to contribute rests on access to knowledge, time, collaborators, and material support. Lowering those costs can allow more people to undertake serious investigations. The power to certify concerns whose scrutiny establishes confidence in a result. Independently checkable proofs can reduce dependence on reputation for assurance of formal validity, while judgments about the formulation and assumptions remain necessary. The power to define value determines which questions, results, and explanations receive sustained attention. Abundant output makes this last power especially consequential.
These powers can be redistributed at different rates. Access to research tools may broaden while the ability to command attention remains concentrated in a few institutions or companies. Formal checking can become routine while disputes over significance become more important. The institutional question is how decisions at each stage can be justified to those affected, including people entering the subject through the new tools.
A mathematician can become more capable in absolute terms while losing standing in relative terms. When tools open a specialty to more participants, they change the scarcity of the expertise on which professional standing has rested. The researcher's achievements remain deserved; the authority attached to them is no guarantee of future influence. We should expect resistance to that redistribution, even among some of its intellectual beneficiaries.
Such an institutional interpretation gives us no license to infer the private motives of particular mathematicians. Concern for students, intellectual standards, and the continuity of research communities can be entirely sincere. It can also coexist with attachment to the arrangements through which a person has found recognition and meaning. A loss of indispensability can be experienced as a loss of quality or purpose. That possibility makes it essential to examine what is actually changing: the reliability of results, the depth of understanding, the range of imaginative possibilities, or the profession's control over those activities. Arguments about what mathematics needs should be assessed on their merits, including by people whose interests those arrangements have served poorly.
The value of a preferred practice is real. Working patiently with pencil and paper, returning to an obstinate example, or finding a proof without assistance can remain satisfying long after a machine can perform the same task. A discovery made for oneself does not become intellectually empty because it was already known elsewhere. Education has always depended on people learning to discover what humanity already knows.
What cannot be guaranteed is that a problem will remain globally unsolved until a particular person or community is ready to solve it. Nor can the continuing value of a practice by itself secure its previous economic position. The distinction matters. A preference for a particular way of doing mathematics deserves respect. A claim that other ways impoverish mathematics requires evidence about their effects, including the opportunities they create.
Professional freedom carries responsibilities. Universities, public funders, and philanthropic institutions support mathematics so that it produces reliable knowledge, deepens understanding, educates others, and makes further inquiry possible. These purposes include abstract investigations whose consequences may be distant or impossible to predict, as well as work that advances science and engineering. They give the profession a responsibility to evaluate tools that could serve those purposes more effectively. The pleasure of working in a particular way is a genuine benefit of an intellectual life; it cannot by itself determine how a supported research and teaching practice should develop.
Different researchers can reasonably choose different methods. A serious assessment must consider the problem, the learner, the reliability of the tool, and the cost of using it. But declining a promising method because it unsettles a familiar practice is not an adequate professional response. The declaration's acknowledgment that mathematics must adapt should be made concrete through investigations of what these tools enable people to discover, verify, teach, and understand. The responsibility is shared by researchers, institutions, and the companies developing the systems. Each should help turn new capability into knowledge that others can use.
There will be real transition costs. A student can lose a planned thesis project. A research community can lose a source of coordination. Institutions can respond badly by preserving priority-based rewards while withdrawing support from the slower work of learning and explanation. These are reasons to reconsider training, funding, and recognition. They do not settle whether the wider mathematical possibilities should be pursued. The people whose plans are disrupted are part of the relevant public; so are the people whose plans become possible for the first time.
A larger intellectual life
The institutional response should follow the changing scarcities. Research funding should support explanation, formalization, and synthesis alongside new results. Journals should distinguish evidence of validity from judgments of significance. Training should develop both the ability to reason through a problem and the ability to use powerful assistance critically. A student needs occasions to struggle without an immediate answer and occasions to discover how far an investigation can go when help is available.
Delegation must not destroy the capacities it is meant to extend. Learning to formulate sound questions and recognize misleading answers requires deliberate practice. The appropriate balance differs between a classroom exercise, an exploratory project, and a major formalization. There is no single amount of assistance that serves all three. The quality of the research system depends in part on how well it develops the judgment of its participants.
People will continue to value different kinds of mathematical life: personal discovery with little assistance, investigations across several subjects, the construction of shared libraries, or the explanation of difficult results. These activities can overlap within one person's work. Their value lies in what they contribute and in what they enable others to understand and do.
For someone who has devoted a life to mathematics, losing intellectual primacy may still be painful. A life's achievements retain their value as the conditions under which they were achieved change. Machines may surpass individual humans, while humans working with machines become capable of investigations previously beyond their imagination. The relevant gain is the expansion of what people can learn, discover, and accomplish, including people whose ambitions the profession has scarcely noticed.
With improving tools, dependable verification, and broad access, I expect sustained AI collaboration to become a dominant form of mathematical research. Its promise is extraordinary mathematics, a larger intellectual life for more people, and new ways of thinking whose possibilities we have scarcely begun to explore. The work ahead includes discovering those practices and learning how to judge them. Institutions must turn abundant cognitive work into reliable knowledge and shared understanding while allowing mathematical ambition to take unfamiliar forms. For each of us, that public responsibility begins with a personal willingness to investigate:
What could I become capable of if I approached these tools with the same curiosity I bring to mathematics itself?
References
- Anthropic. Reasoning models don't always say what they think. Research, 3 April, 2025. URL https://www.anthropic.com/research/reasoning-models-dont-say-think.
- Anthropic. Formalizing Fermat's Last Theorem. Science, 4 September, 2026a. URL https://www.anthropic.com/research/formalizing-fermats-last-theorem.
- Anthropic. Fermat's Last Theorem in Lean 4. Public proof repository and verification documentation, 2026b. URL https://github.com/anthropics/fermats-last-theorem. Accessed 6 September 2026.
- Art Institute of Chicago. Alfred Stieglitz. The Alfred Stieglitz Collection, n.d. URL https://archive.artic.edu/stieglitz/alfred-stieglitz/. Accessed 6 September 2026.
- Charles Baudelaire. The modern public and photography. In Art in Paris, 1845–1862. Phaidon, 1965. URL https://pwad2.wordpress.com/wp-content/uploads/2012/01/charles_baudelaire_the_modern_public_and_photography.pdf. From the Salon of 1859; translated by Jonathan Mayne.
- Shuze Chen, Kunal Marwaha, Xiaoyang Lu, Henry Yuen, and Tianyi Peng. Prove2Me: An open collaborative platform for scaling math formalization. arXiv:2608.28433, 2026. URL https://arxiv.org/abs/2608.28433.
- Hugo Duminil-Copin. Care for a little more AI? Proofs and Prompts, 30 August, 2026. URL https://proofsandprompts.com/2026/08/30/care-for-a-little-more-ai/.
- Georges Gonthier. Formal proof—the four-color theorem. Notices of the American Mathematical Society, 55 (11): 1382–1393, 2008. URL https://www.cs.cornell.edu/courses/JavaAndDS/files/Gonthier4ColorCoq.pdf.
- Peter Jennings. VisiCalc: The early history, n.d. URL https://www.benlo.com/visicalc/index.html. Accessed 6 September 2026.
- Leslie Lamport. Some data on the frequency of errors in mathematics papers. Technical note, originally 14 September 2022, 2023. URL https://lamport.azurewebsites.net/pubs/statistics.pdf. Revised 9 December 2023.
- Lean contributors. Validating a Lean proof. The Lean Language Reference, 2026. URL https://lean-lang.org/doc/reference/latest/ValidatingProofs/. Accessed 6 September 2026.
- Steven Levy. A spreadsheet way of knowledge. Harper's Magazine, November 1984. URL https://www.wired.com/2014/10/a-spreadsheet-way-of-knowledge/. Republished by WIRED, 24 October 2014.
- Jared Duker Lichtman. Proposal to formalize all known human mathematics. X post, 5 September, 2026. URL https://x.com/jdlichtman/status/2096194687765999912. Accessed 7 September 2026.
- Math and AI. A severe misalignment of AI in mathematics. Declaration, 2026. URL https://mathandai.org/. Version consulted 11 September 2026.
- Peggy Noonan. Pause AI for humanity's sake. The Wall Street Journal, 10 September 2026. URL https://www.wsj.com/opinion/pause-ai-for-humanitys-sake-8839ca5e.
- Plato. Phaedrus. Cambridge University Press, 1952. URL https://websites.umich.edu/~lsarth/filecabinet/PlatoOnWriting.html. Translated by R. Hackforth; passage 274c–275b.
- Minkyu Shin, Jin Kim, Bas van Opheusden, and Thomas L. Griffiths. Superhuman artificial intelligence can improve human decision-making by increasing novelty. Proceedings of the National Academy of Sciences, 120 (12): e2214840120, 2023. doi: 10.1073/pnas.2214840120. URL https://arxiv.org/abs/2303.07462.
- Terence Tao. Mathematics in the age of AI. arXiv:2608.16753, essay based on the 2026 ICM public lecture, 2026a. URL https://arxiv.org/abs/2608.16753.
- Terence Tao. Clarification of the hypothetical Navier–Stokes discussion. Mathstodon, 5 September, 2026b. URL https://mathstodon.xyz/@tao/117219101339291693.
- Terence Tao. Discussion of AI, Navier–Stokes regularity, and mathematical understanding. Mathstodon thread, 3 September, especially posts 4–6, 2026c. URL https://mathstodon.xyz/@tao/117207855800042681.