Research and Development · 18 min read
Always-On Research and the Future of Science
We went looking for the labs already running research around the clock, and for who ends up owning what they find.
On the first of September a model was given one of the Millennium Prize problems, the seven questions the Clay Mathematics Institute put a million dollars each on in 2000. Eighty-eight hours later it had an answer.
The answer was that fluids can break. For every viscosity above zero there is a smooth force that takes a fluid starting from rest and drives its velocity to infinity in finite time. Roughly ten thousand agents worked on it at once, in groups that could talk among themselves, exchanging 2.7 million messages and burning around 130 billion tokens. Then a second model spent seventeen more hours writing the whole argument out in Lean, a language in which a computer checks every step and refuses anything it cannot follow.
The paper is 166 pages. Where the authors' names go, it says OPENAI.
The hours are not the interesting part. Something new exists in the world, anyone with a laptop can check it, and there is no person to name as its inventor. OpenAI says it will not claim the prize. The Clay Institute still lists the problem as open, and its president told reporters the evaluation would be "deliberately unhurried."
The cost was somewhere in the millions of dollars, which nobody disputes and nobody has itemised.

Noam Brown
@polynoamial
Yes, this result cost millions of dollars. But remember that when @OpenAI announced o3 it cost ~$500,000 to score 87.5% on ARC-AGI 1. Today, Astra scores higher for ~$20.
Massively scaling test-time compute gives us a glimpse of the future.
8 September 2026
He is right that the price of a given result falls fast. What interests us more is what happens to everything else when it does.
I don't use AI. I have Luis.
That is Diego Córdoba, at the Institute of Mathematical Sciences in Madrid, asked by Quanta whether he uses these tools. Luis is Luis Martínez-Zoroa, who works with him.
The two of them opened the route. Their idea was to force a fluid to blow up by amplifying it across scales while keeping the force that does the amplifying smooth, and OpenAI's paper cites three of their papers and builds directly on that strategy. Tristan Buckmaster at NYU had been attacking the same problem by the same route, and was beaten to it by days. He wrote afterwards that the credit for the basic idea belongs to Córdoba and Martínez-Zoroa, and that he believes Martínez-Zoroa deserves a Fields Medal. Charles Fefferman, who wrote the official problem statement in the first place, told Quanta the same thing in fewer words. The heroes of the story are Córdoba and Martínez-Zoroa.
This keeps happening, and it is worth collecting the cases.
Last October, Google DeepMind's AlphaEvolve improved the lower bound on the kissing number in eleven dimensions, which is the largest number of unit spheres you can pack around one more, from 592 to 593. Then Mikhail Ganzhinov, a doctoral student at Aalto, set new records in dimensions ten and fourteen by restricting the search to highly symmetric arrangements. Two of three dimensions, to one person with a better idea. He thinks eleven can be pushed past 600.
In April, a twenty-three-year-old with no graduate training named Liam Price asked a model a question about primitive sets. The method that came back, Markov chains with von Mangoldt weights, turned out to resolve two conjectures from 1966 and had apparently been sitting unnoticed since Erdős's 1935 paper. Terence Tao is the eighth author.
None of that is going away. What is changing is what it now runs against.
Nobody goes home
Research has always been paced by human attention. A grant cycle is three years. A doctorate is four. An experiment runs when somebody is in the building to run it, and stops at the end of the week, and the equipment sits idle over Christmas. The rate at which new things get discovered has been bounded, for as long as there has been research, by how many trained people are looking and how many hours they are awake.
A loop that proposes its own next experiment does not have that shape.
We ran into this from an odd direction. Star Voyager is a galaxy we built out of twenty-eight thousand agents that hold memory, form cultures and keep making decisions whether or not anybody is watching. The thing that surprised us was the shape of the bill. Compute tracked the size of the living world and barely moved with the number of players in it. The cost of the system was the cost of keeping it awake.
The same economics is now arriving in research. The input stops being how many people you can hire, and becomes how much you are willing to keep running.
The longest we have left one of our own systems running unattended is three days. A group of agents was pulling structured data out of thousands of legal documents for a client, spawning subagents to rework the pipeline every time it met a kind of clause it had not seen. We gave it a goal that would keep looping until the requirements were met, and left it overnight. Then another night. Then another.
When we finally sat down and read what it had been doing, nothing had crashed. The agents had worked out that their general passes were landing under an assumed eighty per cent accuracy floor, and had responded by building an elaborate bureaucracy of checks around themselves, until the system was processing exactly one document at a time. At that rate it was about fifteen days from finishing. It would have got there, too.
Nothing broke. The loop did exactly what we asked, and what was wrong was the asking. We stopped it and rewrote the goal to something far more agile, which is the whole job now: the machine will keep going, and the expensive question is what you pointed it at.
Terence Tao described what this would look like five days before OpenAI announced anything.

Terence Tao
@tao@mathstodon.xyz
mathstodon.xyz
Technically, one of the most prominent open problems in mathematics would now be solved; but there would be almost no value added to mathematics as a consequence.
3 September 2026
He was writing about the possibility that a company could run the whole search internally and publish only the answer, keeping the reasoning that produced it out of view. Roughly speaking, that is what happened. His sharper point came in the same thread. Pointing a powerful tool at a problem with no expert guiding it has driven a wedge between producing answers and producing understanding. The two have started moving in opposite directions.
A catalyst made from a meteorite
Mathematics is the easy case, because the work and the proof live in the same place. The interesting question is what happens when the loop can pick things up.
At the University of Science and Technology of China, a robot was given fragments of Martian meteorite. Laser spectroscopy read what the rocks were made of. A model trained on about thirty thousand simulated compounds cut 3,764,376 possible recipes down to a shortlist. A mobile robot carried samples between fourteen workstations to make them, and an electrochemical rig measured which ones split water best. Two hundred and forty-three compositions were physically made and tested. The whole campaign took six weeks, and the resulting six-element catalyst ran for more than five hundred thousand seconds without failing, including a stretch at minus thirty-seven degrees.
The point of that last detail is that it was aimed at Mars, where making oxygen out of local material is the difference between a visit and a stay.
Perhaps a dozen loops in the world have produced something new this way.
At the Technical University of Denmark, FastCat makes and electrochemically tests up to seventy-five catalyst compositions a day with nobody involved, and has found four-element combinations that beat the previous best by ten millivolts. A rig at Boston University has been printing, retrieving, weighing and crushing plastic structures since 2021 at over ninety per cent uptime. It has tested more than twenty-five thousand shapes, and it holds the record for energy absorption efficiency, having taken it off a human. At the University of Illinois, a closed loop found a general condition for a reaction used in thousands of published papers that doubled the average yield across a twenty-substrate test set, from twenty-one per cent to forty-six.
Liverpool's mobile robotic chemist is the one we keep coming back to, because of what it refused to do. Andrew Cooper's group built no machine of their own. They put a four-hundred-kilogram robot on wheels into an ordinary laboratory and had it work the instruments already standing there, unmodified, in the dark. It ran 688 experiments in eight days, working 172 hours out of 192, and spent thirty-two per cent of that time waiting for the gas chromatograph. Cooper's description of the design is the best sentence in the field. Automate the researcher, rather than the instruments.
The loops that actually close
What a laboratory does in a day when nobody goes home
Published throughput for the handful of systems where software picks the next experiment, equipment runs it, and a measurement feeds back. Converted to a common unit, which flatters some of them and none of them enormously.
Experiments per day at five autonomous laboratories. Polybot at Argonne, one hundred. The mobile robotic chemist at Liverpool, eighty-six, from 688 experiments in eight days. FastCat at the Technical University of Denmark, seventy-five, with no human involved. Radical AI in New York, fifty, from twelve hundred alloys in six months. A-Lab at Lawrence Berkeley, twenty-one, from 353 experiments in seventeen days.
PolybotArgonne
one sample every 15 minutes
100
Mobile robotic chemistLiverpool
688 experiments in 8 days
86
FastCatTechnical University of Denmark
no human involved
75
Radical AINew York
1,200 alloys in six months
50
A-LabLawrence Berkeley
353 experiments in 17 days
21
experiments per day
Each figure is the laboratory's own published rate, and each is measured over a different campaign length, so they are indicative rather than a league table. The units differ underneath: an alloy, a thin film and an inorganic powder are not the same amount of work.
The commercial version of this is further along than it looks. A New York company called Radical AI, whose chief scientist also runs the autonomous lab at Lawrence Berkeley, produces and characterises around fifty alloys a day and made twelve hundred in six months. Its alloy RAI-939 beat C103, the niobium alloy that has been the aerospace default since Apollo, by a factor of 125 on environmental survivability in a torch test verified by Purdue. It took sixteen weeks.
The mortar and the pestle
Now the other half, because most of what gets called a self-driving laboratory is not one.
In 2024 Microsoft screened 32,598,079 candidate battery materials down to eighteen novel ones, in under eighty hours across a thousand cloud machines, and the result was a real electrolyte that ran a real lightbulb. Then read how it got made. A scientist at Pacific Northwest National Laboratory explains that at this point the work is artisanal, and that one of the first steps is to grind the solid precursors by hand with a mortar and pestle. Three of the four synthesis targets turned out to produce phases that were already well known.
The bottleneck has nothing to do with intelligence. Nobody has worked out how to move a tray of samples between two instruments without a person carrying it. That one unsolved problem explains most of the field. Almost every working loop handles liquids or powders, because those are the geometries robots can already manage, and more than sixty per cent of the errors in Liverpool's landmark run came from dispensing liquid and crimping caps.
We ran into the same wall from the other side. When we designed an observatory in the Alentejo, the part that survived every review was a hard line between the two halves of the night: AI-assisted planning and interpretation on one side, and a simple, predictable safety layer on the other. A dumb layer fails in ways you can write down in advance, which is the entire reason to have one. Every operating loop in this article has that line somewhere, and the ones that get into trouble are the ones that drew it in the wrong place.
The people building these systems say so plainly, when you read past the press releases. Alán Aspuru-Guzik's own group wrote in 2022 that they had implemented self-driving components but had not arrived at a fully autonomous system, and that some steps still require a human to load new vials. When Toyota Research Institute surveyed 102 materials researchers, only twenty-six per cent were comfortable automating their full workflow. The rest wanted to keep idea generation, interpretation and on-the-fly adjustment.
Anthropic published its own failure, which is worth more than most companies' successes.

Anthropic
anthropic.com
Human experts had to intervene when Claude misread errors caused by bubbles as software failures and initially responded in a way that produced more foam.
27 August 2026
The claims need reading carefully too. In November 2023 the autonomous laboratory at Lawrence Berkeley announced forty-one novel compounds from fifty-eight targets in seventeen days, and it became the most cited result in the field. Fourteen months later Nature issued a correction. The count fell to thirty-six, four were inconclusive, one compound was removed because it had been in the training data, and the word "novel" was deleted from the paper's title. The authors conceded that the materials were new to the prediction platform rather than necessarily new to science. Critics who had been arguing this since 2024 put the number of new compounds at three.
How to read any claim in this article
Four ways of checking the same sentence
Every result here sits on one of these rungs, and the rung matters more than the headline. A number checked by nobody and a number checked by a compiler are not the same object.
Four levels of verification, weakest to strongest. One: nobody outside the organisation has looked, as with Lila Sciences' antibody claims, Kosmos at 79.4 per cent accuracy and Anthropic's Riemann zeta bound. Two: graders the claimant hired, often unnamed, as with OpenAI's 2025 olympiad gold marked by three former medallists. Three: independent experts with no stake, as with the Erdős unit distance disproof checked by nine mathematicians and the 2026 olympiad graded by the competition. Four: a formal checker compiles it or refuses, as with the Navier-Stokes proof in Lean, the Fermat formalisation compiled by a rival, and AlphaProof Nexus.
Nobody outside
The organisation says so, and no one else has looked
- Lila Sciences' antibodies
- Kosmos at 79.4% accuracy
- Anthropic's Riemann zeta bound
Graders it hired
Experts chosen and paid by the claimant, often unnamed
- OpenAI's IMO 2025 gold, marked by three former medallists
Independent experts
Outside specialists with no stake, working from the output
- The Erdős unit distance disproof, checked by nine mathematicians
- IMO 2026, graded by the competition itself
A machine
A formal checker compiles it, or refuses
- The Navier-Stokes proof, 641,332 lines of Lean
- Fermat's Last Theorem, compiled by a rival
- AlphaProof Nexus, nine Erdős problems
A formal checker proves a theorem follows from the axioms it was given. It does not prove the formal statement matches the question anyone meant to ask, which is a separate job still done by people.
The same care applies to the loudest claims about mathematics. The researcher who announced OpenAI's Erdős results in October 2025 described what the model had actually done, before anybody else recast it.

Sébastien Bubeck
@SebastienBubeck
gpt5-pro is superhuman at literature search: it just solved Erdos Problem #339 (listed as open in the official database) by realizing that it had actually been solved 20 years ago
12 October 2025
Finding what is already known is a real and useful capability. Making something new is a different one. The two get reported as though they were the same thing.
The honest measure of how far this has come belongs to the people paying for it. The US Department of Energy has put thirty million dollars into robotics testbeds for autonomous discovery, and asks everyone applying to benchmark one number: mean time between human interventions.
Five words, and they contain the whole field. The number is going up.
Now it is for sale
Three things happened this year that move this out of the demonstration category.
OpenAI put GPT-5 in charge of a real laboratory for six months. Ginkgo Bioworks supplied the hardware, a connected system of about a hundred robotic carts each wrapped around one instrument. The model designed every experiment: 480 plate designs, 29,527 unique reaction compositions, roughly 150,000 measurements. The result cut the cost of cell-free protein synthesis by forty per cent and raised the yield by twenty-seven, against the best published human work on the same problem. As far as we can tell that is the first time a machine-run campaign has beaten a human one head to head on a published benchmark. Humans still bought the reagents and chose which of the model's designs to run, and two plates ran with a volume overrun and a unit-conversion bug.
On the ninth of September, the day before we wrote this, rentosertib dosed its first Phase III patient. The target was found by software and the molecule was designed by software, at Insilico Medicine, which is not one of the frontier labs. No fully AI-designed drug has been approved anywhere yet, and its own lead investigator puts approval three to four years out under favourable conditions.
And in July, AlphaEvolve became a product. The system improved matrix multiplication for the first time since 1969 and holds new lower bounds on nine Ramsey numbers. In July it went generally available on Google Cloud, with an API and a skill that runs inside a code editor. You supply a seed program and a way to score the output.
Always-on discovery is not a thing that is coming. It has a signup page.
We cannot rule it out
So who owns what comes out of it.
The answer turns out to be more reassuring than we expected, and it turns out to be the wrong question. Read the terms of service and the output belongs to you. OpenAI assigns you all its right, title and interest in whatever the model produces, and Anthropic's commercial terms say the same. A machine cannot be named as an inventor anywhere that matters. Thaler lost in the US Federal Circuit, in the UK Supreme Court and at the European Patent Office. Last November the USPTO went further and withdrew its own guidance, saying there is no separate standard for AI-assisted inventions and that AI should be treated as laboratory equipment. If you find a cure using one of these tools, it is yours. Nobody found evidence of OpenAI or Anthropic patenting an AI-generated scientific result.
What concentrates is the ability to look.
Google is the clearest case, because the reversal is dated. AlphaFold 2 was open-sourced in 2021 with a free database, and Demis Hassabis said at the time that they were huge believers in open science and had done that with all of their scientific work. AlphaFold 3 was published in 2024 with pseudocode and no code; the code followed six months later, and the weights never did. They are available to non-commercial organisations, from Google directly, with redistribution forbidden. IsoDDE, in February this year, arrived with no code, no weights, no API and no peer-reviewed paper, available through pharmaceutical partnerships worth up to 1.7 billion dollars in milestones with Eli Lilly and 1.2 billion with Novartis.
And last October a patent was granted to a DeepMind holding company. Its first claim begins with obtaining a ligand, where the ligand is a drug for treating a disease, and its last step is synthesising it. Read plainly, the claim covers the method of finding and making a drug.
The harder version is data. Periodic Labs raised three hundred million dollars in its first round and has not yet published a result. Its site says plainly what the labs are for. They produce enormous quantities of high-quality data that exists nowhere else, including the valuable negative results nobody bothers to publish. Every word of that is correct, and it describes the one asset with no open substitute. A withheld model gets reimplemented within a year; Boltz shipped MIT-licensed weights eight days after AlphaFold 3's code release. A withheld set of experiments does not get reimplemented by anyone.
Which brings us to the argument that started the day after the proof landed. Buckmaster had been putting his drafts into OpenAI's coding product for the whole project, and asked whether the model had been trained on them. He was careful about it, in a way most of the coverage was not, writing that he did not know whether their data was used and was not accusing anyone of anything. OpenAI's answer is in its own announcement.
OpenAI
openai.com
While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.
8 September 2026
Nobody has shown that anything improper happened, and the two proofs differ substantially. What the sentence establishes is the shape of the thing. A company can supply the tools a researcher works in, compete with that researcher, arrive first, and be unable to say for certain that the one did not feed the other.
The scale of it is now measured. The UN's scientific panel on AI, co-chaired by Yoshua Bengio and Maria Ressa, reported in July. The United States holds about seventy-five per cent of the compute in the world's five hundred largest clusters. Ninety-one per cent of notable models come from the private sector. A hundred and eighteen countries take no part in AI governance discussions at all.
Yoshua Bengio and Maria Ressa
un.org
Concentration of capability is becoming concentration of political power. The sovereignty of most nations is at stake.
1 July 2026
Six per cent of the genome, for two years
There is one good precedent for this, and it is close enough to be useful.
When Celera Genomics raced the public Human Genome Project, it did not rely on patents. It put its assembled genome on a website and gave away DVDs, on the condition that you agreed not to commercialise or redistribute what you found. The restriction on passing it on is what made it a sellable asset. Pharmaceutical companies paid between five and fifteen million dollars a year for access. A university laboratory paid seven and a half thousand. The arrangement covered 1,682 genes, about six per cent of the total, and it expired automatically as the public effort resequenced each one, so nothing was held for more than two years.
Heidi Williams measured what those two years cost. Comparing genes held under Celera's terms with equivalent genes that were public, she found reductions in subsequent research and product development of twenty to thirty per cent, persisting long after the restrictions had lapsed. Roughly fourteen hundred papers that were never written, and about forty diagnostic tests that did not exist by 2009.
What restricted access cost, the last time it was measured
Two years, six per cent of the genome, a fifth of the science
Celera held 1,682 genes under contract rather than patent, and released each one automatically as the public project caught up. Heidi Williams compared those genes with equivalent public ones and found the gap had outlived the restriction.
Two comparisons between genes Celera held under contract and equivalent public genes. Papers published per gene between 2001 and 2009: 1.24 for held genes against 2.12 for open ones. Used in a diagnostic test by 2009: 3.0 per cent of held genes against 5.4 per cent of open ones. Williams estimates the overall reduction in subsequent research and product development at twenty to thirty per cent.
Papers published about the gene
2001 to 2009, per gene
Used in a diagnostic test
by 2009
Williams, Intellectual Property Rights and Innovation: Evidence from the Human Genome, Journal of Political Economy 121(1). Her headline estimate, after controls, is a reduction of twenty to thirty per cent. The raw gaps shown here are larger.
Two years, on six per cent, by contract rather than patent.
The counterweight is Bell Labs, and it is instructive in a way that is easy to miss. The 1956 antitrust decree forced Bell to license 7,820 patents royalty-free, including the transistor, and follow-on innovation on those patents rose by seventeen per cent over the next five years, most of it from small and young companies. Sony paid twenty-five thousand dollars in 1953 and shipped a transistor radio two years later. But the increase happened only outside telecommunications. In the market where Bell itself competed, opening the patents changed nothing at all.
And the case where sharing was left voluntary is the one with nothing to show. The World Health Organization set up a pool for COVID technology in May 2020, backed by forty-five countries. Pfizer, BioNTech and Moderna all declined. In three and a half years it attracted six licences, none of them for an mRNA vaccine. The technology transfer hub in Cape Town eventually built its candidate by working from published data, because, as its chief executive put it, the companies did not even want to talk.
Every time access was actually widened, somebody was made to do it.
At a price
The case for calm is that intelligence is not the binding constraint. Benjamin Jones at Kellogg makes it carefully: even extreme intelligence appears strongly limited when it operates on only a minority of the tasks involved. Somebody still has to grind the powder. More than a hundred and seventy AI-designed drug programmes are in clinical development and not one has been approved, and success rates in the clinic have not moved.
The case for paying attention is that all of those are engineering problems, and engineering problems fall. Sample transport is a mechanical problem, and mechanical problems get solved. Mean time between human interventions is a number that goes up every year.
What that adds up to is a research system whose output scales with capital rather than with the number of trained people. Cures for diseases, machines that waste less, materials that do not exist yet. Those things move from being bounded by how many researchers are awake to being bounded by how much anyone is willing to spend keeping the loop running. All of that is good news, and it arrives attached to a question nobody has answered.
Two years of restricted access to a small slice of the genome cost a fifth of the science that would otherwise have been done on it. Nobody is proposing anything as modest as that now.
We do not have a policy to offer. When we shut down our agent product we said the work we actually wanted was putting AI into the physical world, into problems where the hard part is not another integration. This is what that looks like from the outside. We have a company that builds agent systems and puts them next to hardware, and a growing suspicion that the interesting question for the next decade is not whether the machines can do the research. The question is who gets to keep what they find, and for how long, and what everybody else does in the meantime.
Explore our R&D work or talk to Darkmatter.
Darkmatter is a hardware, software and AI lab based at Poolside Santos in Lisbon.

