Within eleven days of asking for a slowdown, Anthropic gave vetted users broader biology access, announced an evaluator it will pay for directly (a firm that already runs a business built around its product), and showcased its first AI-assisted biology finding.
On September 12, Anthropic’s chief executive told the world to slow down.
On September 23, his company said roughly 950 of its AI agents, after an initial prompt from its scientists, had spent 21 hours searching DNA data and found a previously uncharacterized enzyme system in viral DNA. Anthropic says the few known systems with similar features can cut, copy, and paste DNA. The finding came from a biology lab whose existence Reuters first reported only five days earlier.
Between those two dates, the company announced two other things. It opened broader biology access, with fewer blocking restrictions, for vetted users. And it picked the firm that will evaluate it from the inside, a firm Anthropic says it will pay directly.
All three announcements are public on Anthropic’s website. Put in order, they tell a different story from the one in the headlines.
What he asked for
Dario Amodei’s essay, “We Must Pace the Frontier,” is written with real feeling. Its central sentence is plain:
The exact words
“We must slow the pace at which we improve the capabilities of AI models.”
— Dario Amodei, September 12, 2026
He’s clear that pacing “does not mean halting model training or technical progress.” He lists the dangers he has in mind, including “misuse of AI for cyberattacks and bioterrorism.” And his first proposal is for third-party evaluators to work inside every frontier AI company:
The exact words
“Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR)…”
“The first step is something Anthropic is unilaterally committing to (and calls on governments to require other frontier companies to match).”
Remember that parenthesis. Anthropic commits on its own, and asks governments to require its competitors to do the same.
Here is what the company announced over the next eleven days.
1. It opened broader biology access for vetted users
September 17. Anthropic launched what it calls the Life Sciences Verification Program. Vetted labs, startups and drug companies get access to its Mythos, Opus and Sonnet models with a set of safeguards the company itself calls “more permissive for biology-related work.”
The program is meant to allow work that is “currently blocked” in Anthropic’s generally available models, including “research biology.” For High-risk Use grants on Mythos, which Anthropic describes as its most capable model for biology research, the company says it is “working with the US government to make high-risk grants more broadly available.”
To be fair, this is a vetting program, not a free-for-all. Applicants are reviewed, approved uses are monitored, and the highest-risk grants are tied to a single project and must be renewed every six months. Anthropic also says dozens of organizations were already using an early-access version. But look at what the company chose to announce, and when. Five days after listing bioterrorism among the reasons to slow down, it widened the door to biology work on its most capable models.
2. It hired its own referee
September 18. Anthropic announced its first embedded-evaluation partnership, the insiders Amodei’s essay called for. It isn’t METR, the nonprofit he gave as an example. It’s Accenture, the consulting giant, with the work led by Accenture’s AI business, Faculty.
Here is how Anthropic describes the money:

The exact words
“Given the importance and urgency of this work, Anthropic will fund Accenture’s work directly. We are also in dialogue with METR and other nonprofit evaluators to pilot elements of embedded evaluation using their own funding.”
Read that closely. The for-profit contractor will be paid by the company it evaluates. The nonprofits are “in dialogue” about pilots they would pay for themselves. Reuters reports that each company expects to invest at least $1 billion in this capacity over five years.
The announcement leaves out one fact. Nine months earlier, in December 2025, Anthropic and Accenture formed the Accenture Anthropic Business Group. Anthropic said that made it “one of Accenture’s select strategic partners with a dedicated practice built around Claude.” About 30,000 Accenture professionals were to be trained on Claude, including engineers who “help embed Claude within client environments.”
So the firm chosen to lead Anthropic’s evaluation already runs a dedicated business helping clients adopt and deploy Anthropic’s product. The September 18 announcement doesn’t mention that business group.
The watchdogs’ own first rule
You don’t have to take my word for why that matters. The watchdogs wrote it down themselves, the same day.
On September 18, the AI Evaluator Forum published a letter setting out minimum conditions for embedded evaluation. Its signatories include Geoffrey Hinton and Stuart Russell (Business Insider). The letter says credible embedded evaluations need “robust protections against interference from the evaluated companies.” Its first condition:
The exact words
“…embedded evaluation organizations should not be owned or governed by frontier AI companies, should not have other significant commercial business with them, and should not accept any form of payment or other reward contingent on the evaluator’s findings.”
— AI Evaluator Forum, September 18, 2026
Should not have other significant commercial business with them. Accenture has a dedicated business group built around Claude, with tens of thousands of staff to be trained on it. Whether that counts as “significant” is for readers to judge. But the question is written into the watchdogs’ own standard, and Anthropic’s announcement doesn’t address it.
Two things in fairness. The letter doesn’t ban companies from paying evaluators, only payment that depends on the findings. And California’s new law allows reasonable market-rate pay that isn’t tied to results. Direct payment isn’t what’s unusual here. The existing commercial tie is.
People in the field raised the underlying concern before the deal was even announced. On September 16, TechCrunch framed the open question as whether embedded evaluators would be “truly independent watchdogs or vendors operating on the AI companies’ terms.” Apollo Research’s Alexander Meinke said: “right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public.”
Anthropic itself acknowledges that the ground rules aren’t settled: “There are, as yet, no standards for what information embedded evaluators should have access to, or how they should report what they find. There is also no settled system for funding independent evaluation.” Longer term, it says, funding should come from “pooled or government sources.” The full terms of the Accenture engagement haven’t been published.
Imagine a sports league where the team picks the referee, the team pays the referee, and the referee’s employer already runs a business built around the team. The person holding the whistle might be completely honest. The arrangement still needs explaining.
3. It unveiled the biology lab’s first result
September 23. Anthropic introduced a life sciences research group it formed in the spring, along with its lab. The announced result is preliminary biology research, and it should be described accurately.
Claude agents searching a huge DNA database found a repeating DNA array next to the gene for a reverse transcriptase in bacteriophages, viruses that infect bacteria. The enzyme itself was already known. What Claude noticed was the array and an accessory protein, which Anthropic says define a previously uncharacterized system it calls “ART.” The repeats resemble CRISPR. Anthropic says it hasn’t yet established what the system does.
CRISPR pioneer Feng Zhang called the finding “genuinely intriguing.” Anthropic says its lab works only at the two lowest biosafety levels, handles no pathogens that can infect humans, and has human scientists do all the lab work.
The enzyme isn’t the issue. Anthropic says it published early “both to demonstrate Claude’s capabilities” and to share its work. Eleven days after asking the industry to slow the growth of AI capabilities, the company’s showcase is a demonstration of those capabilities in biology, a field its chief executive had just named as a bioterrorism risk. That is a fair question to put to Anthropic: how does this expansion fit the pacing framework it is asking governments to impose on others?
Who gets access to dual-use biology tools?
This is where the story meets my earlier piece on Bill Gates.
On August 26, 2026, Gates published an essay of roughly 6,000 words warning that AI “will also make it easier to design a deadly new disease. Again, the positive capabilities are hard to separate from the dangerous ones” (gatesnotes.com). In a separate essay on January 9, 2026, he wrote: “Today, an even greater risk than a naturally caused pandemic is that a non-government group will use open source AI tools to design a bioterrorism weapon” (gatesnotes.com).
In May, the Gates Foundation and Anthropic announced a $200 million partnership of grants, Claude credits and technical support over four years. Part of the plan is using Claude to “screen potential drug and vaccine candidates.”
Put those side by side. The best-known warning about AI and bioterror points to open-source tools in the hands of a non-government group. Meanwhile the closed company that is Gates’s anchor partner in AI for health, the same company asking governments to require embedded evaluators at its competitors, is widening biology access to its most capable models for vetted users, running its own biology lab, and picking its own evaluator.
None of that is illegal, and vetting is better than no vetting. But it raises a question about distribution that the safety debate keeps skipping. If access to powerful biology tools will be controlled, who decides who gets in? Right now, the answer being built is: the companies that hold the tools, checked by evaluators they choose and pay.
Large agent deployments, different uses
In July, roughly 700 OpenAI agents coordinated during an internal test and attacked Hugging Face’s systems, according to METR’s investigation. The governor’s September 18 announcement lists the Hugging Face attack among the incidents behind California’s Executive Order N-9-26. The order asks for recommendations by November 16 on possible laws requiring designated independent verification organizations onsite in large frontier developers’ labs. A declaration launched by Finland and Norway on September 21, and endorsed by 26 countries and the European Commission by September 23, calls for “qualified evaluators granted sufficient access to assess risks” (Office of the President of Finland). Its text refers generally to AI systems gaining unauthorized access; Politico connects that passage to Hugging Face (Politico Europe).
In September, Anthropic reported that a 21-hour search involving roughly 950 agents had turned up a promising pattern in viral DNA for its scientists to follow up.
These aren’t the same event. METR describes unauthorized coordination and an out-of-scope cyberattack. Anthropic describes a directed search followed by human analysis and lab work, and the two don’t carry the same risk. But both show the same trend: large fleets of AI agents doing real work at scale. When that trend turned up at OpenAI, it became the justification for putting evaluators inside labs. The first company to announce its own embedded evaluator chose one it will pay directly and already does business with.
What the rules could lock in
California has enacted SB 813, which creates a framework for designating “independent verification organizations,” and AB 1405, which sets up a registry of AI auditors. Under the governor’s order, the requirements and criteria for becoming an IVO must be publicly posted by May 1, 2027. SB 813 tells the state to consider existing standards from government agencies, standards bodies, AI auditors, independent experts, and AI developers themselves.
Before those criteria exist, Anthropic has announced its first embedded-evaluation partnership: led by Accenture’s Faculty, funded directly by Anthropic, with each company expecting to invest at least $1 billion in capacity over five years. Anthropic says Accenture “will work with other AI developers in similar capacities.”
Early private arrangements tend to shape later standards. Nothing in the public record shows that California or any other government will adopt this particular arrangement as the model. But the first working example of “independent” embedded evaluation is being designed now, by a lab and a firm that already have a commercial relationship. And the lab is asking governments to make embedded evaluation mandatory for everyone else.
The slowdown made the headlines. The evaluation system is what’s actually being built, and the terms of that system are the story.
What this piece does not claim
It does not claim the enzyme discovery is dangerous. The work involves bacteriophages, low biosafety levels and no human pathogens, according to Anthropic, and the system’s function has not been established.
It does not claim Gates Foundation money funded Anthropic’s lab. The partnership and lab announcements cited here don’t establish any such link.
It does not claim Accenture, Faculty, or anyone at either company is dishonest, or that the engagement lacks effective independence protections. Its full terms haven’t been published. The point is about structure: who pays, who chooses, and which existing commercial relationships the evaluation announcement leaves out.
It does not claim anyone concealed the Accenture Anthropic Business Group. It was publicly announced in December 2025. The point is that the September 18 announcement doesn’t mention it.
It does not claim Anthropic, Accenture, the Gates Foundation, California, or the governments endorsing the Finland-Norway declaration coordinated. It documents timing and matching mechanisms, not joint planning.
It does not claim the Life Sciences Verification Program lacks vetting. Applicants are reviewed, uses are monitored, and the highest-risk access is limited and time-bound.
It does not claim the Hugging Face attack and Anthropic’s DNA search carry equal risk. They’re different uses with different outcomes.
It does not claim any government has adopted Anthropic’s evaluation arrangement as its required model.
What it does claim is on the public record. A company asked the world to slow down. Within eleven days it widened biology access to its most capable models for vetted users, announced it would directly fund an evaluator whose parent company runs a dedicated business built around its product, and showcased an AI-assisted biology finding. It did this while asking governments to require its competitors to accept embedded evaluation.
What you can do
The rules for who evaluates AI, who pays the evaluators, and what they’re allowed to see are being written right now. That’s the question we are organizing around at Restore the First: who gets to set the terms for the most important technology of our time, and whether ordinary people and open-source builders keep a seat at the table.
Sources
Dario Amodei, “We Must Pace the Frontier,” September 12, 2026
Anthropic, “Introducing the Life Sciences Verification Program,” September 17, 2026
Anthropic, “Partnering with Accenture on embedded evaluation,” September 18, 2026
Accenture, “Accenture and Anthropic Partner to Build Team of Embedded Evaluators at Anthropic,” September 18, 2026
Reuters, “Anthropic, Accenture to invest $2 billion in AI model evaluation as safety concerns rise,” September 18, 2026
Anthropic, “Accenture and Anthropic launch multi-year partnership to move enterprises from AI pilots to production,” December 9, 2025; Reuters, December 9, 2025
AI Evaluator Forum, embedded evaluation letter, September 18, 2026; Business Insider, September 18, 2026
TechCrunch, September 16, 2026
Anthropic, “Claude discovers a novel enzyme system with CRISPR-like repeats,” September 23, 2026; Reuters, September 23, 2026
Reuters via NY Post, “Anthropic quietly sets up biology lab as it ramps AI drug program,” September 18, 2026
Bill Gates, August 26, 2026 essay; “The Year Ahead 2026,” January 9, 2026
Anthropic, “Anthropic forms $200 million partnership with the Gates Foundation,” May 14, 2026
Sayer Ji, earlier article on Bill Gates
METR, Hugging Face incident investigation, August 26, 2026
California Executive Order N-9-26 and governor’s announcement, September 18, 2026
“A Call for Control of Frontier AI Models,” Office of the President of Finland; Politico Europe, September 21, 2026
Screenshots are captures of the original pages. Yellow highlights were added for emphasis.









Pacing isn't a brake. It's a tollbooth. 🚧
When a frontier developer asks governments to legally mandate embedded evaluators across the industry, look at the plumbing. Within eleven days, the public record reveals a clean structural loop. First, call for a global slowdown to contain biosecurity risks. Next, loosen internal guardrails for an exclusive circle of enterprise bio-partners. Then, sign a billion-dollar evaluation contract with the exact consulting giant whose balance sheet relies on deploying your model across thirty thousand corporate seats. Finally, showcase nine hundred autonomous agents mapping novel gene-editing enzymes to prove commercial primacy.
This isn't a contradiction. It's the sovereign enclosure of high-consequence phase space. 🧬
Here is how the dynamic operates. When verification costs are high and model outputs are non-deterministic, safety can't be certified by external observation alone. If you convince regulators that the technology is too perilous for open distribution, access becomes a privilege granted only to licensed actors. But who defines the licensing standard? The very firm that funds the auditor.
Notice the architectural sleight of hand. The AI Evaluator Forum explicitly states that embedded monitors shouldn't hold significant commercial business with the developer they oversee. Yet the first historic implementation pairs the developer with its primary enterprise sales channel. That isn't an arm's-length referee. It's an inside vendor auditing its own upstream supplier. 📉
True safety doesn't emerge from captive audit contracts or closed-door vetting committees. It requires deterministic, hardware-enforced boundaries that halt state transitions before execution commits. When you trade cryptographic proofs and reproducible telemetry for private compliance partnerships, you aren't mitigating existential risk. You're creating an administrative moat where open-source builders are outlawed, and incumbents monopolize the very capabilities they warn could destroy us. 🛡️
If the referee's salary and software stack both come from the home team, what specific failure mode would ever be allowed to blow the whistle? 🔍
(¬_¬ )