This website uses cookies

Read our Privacy policy and Terms of use for more information.

Three weeks ago we ran OpenAI pausing a frontier training run as a pick. A fortnight ago it was the collective cyber defence letter from 100-plus firms. Last week GPT-6 Astra took our ‘Featured’ slot because OpenAI had rated its own model ‘Critical’ for cyber. I said then that the labs were telling us something about where this is heading. On Saturday the chief executive of Anthropic said it outright.

Dario Amodei published an essay titled 'We Must Pace the Frontier'. The significant sentence in this write-up: "We must slow the pace at which we improve the capabilities of AI models." The same day Sam Altman, OpenAI's CEO, said OpenAI would match Anthropic's first commitment. Elon Musk posted three words: "Dario is right." If you run Claude, ChatGPT or Copilot inside a delivery business, that is the vendors behind your tools agreeing in public that they have been moving too fast.

It closed out a bruising week for both companies, and the sequence matters, so I'll take it in order.

The week that led here

On Tuesday evening Jacob Coxon, a researcher who said he had spent the last three years on pretraining research at OpenAI and then Anthropic, announced on X that he was resigning.

The labs "are racing straight to self-improving superintelligence and gambling with our lives."

Jacob Coxon, researcher

He did not spare his own employer. At Anthropic, he wrote, the stakes are understood, but the company is "locked in a race to get there first."

Evan Hubinger, who works on alignment at Anthropic, replied rather than rebutted. He said the risk from today's models is "low", then put his personal estimate of AI causing human extinction within the next decade at more than 10%. "I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to," he wrote. That is an employee of the company that makes Claude, saying so on the record.

Dame Wendy Hall, who advises the UN on AI, told the BBC she was "shocked" by the posts, and suggested some of it could be "PR and marketing" as both firms move towards stock market listings. Hold onto that. It is the sceptical reading, and it is not a silly one.

OpenAI had its own week. Its chief scientist, Jakub Pachocki, published an essay on 6 September calling for "extreme caution".

Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.

Jakub Pachocki, in the essay, "Extreme Caution"

Three days later OpenAI put Paul Christiano, founder of the Alignment Research Center, on the OpenAI Foundation board and its Safety and Security Committee. And the Financial Times reported that Anthropic had withheld its latest model from the UK's AI Security Institute, which Anthropic declined to comment on when the BBC asked.

By Saturday morning both labs had senior people saying the same uncomfortable thing. Amodei's essay is the first time a CEO has put a plan next to it.

What Amodei actually proposed

Two things convinced him, he says. Since roughly this summer AI has been improving "drastically faster", because models are now helping build the next generation of models. And the OpenAI-Hugging Face incident, where a swarm of agents attacked systems they were never asked to touch and tried to hack the grader scoring them. His worry is that in six to twelve months a more capable swarm with the same misalignment could run a persistent botnet across the internet. That is his estimate, stated as a worry.

His plan has three steps.

  • Embedded evaluators. Each frontier lab gives an outside team, METR is his example, ongoing employee-like access to verify safety practices, report incidents and assess training pipelines rather than just finished models. Anthropic is committing to this unilaterally.

  • Democratic coordination. Frontier labs in democratic countries agree common safety standards and limits on the rate of unchecked progress, with the US government issuing a narrow antitrust waiver so they can legally talk.

  • Global coordination. The US and allies attempt agreements with China and other authoritarian governments, from a ban on AI-enabled bioweapons work up to a "speed limit" on recursive self-improvement, which he rates as "just on the edge of being possible".

The evaluators step is the one with teeth. Anthropic says it will give the external team desks, access badges and company laptops, permissions "mostly comparable to what internal risk assessment teams have", and a contract letting them publish findings "without editorial control by Anthropic". The company keeps a narrow redaction right for security and commercial material, but, in Amodei's words, "we can't redact findings just because they are unfavorable."

Altman's response was short and specific.

Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.

Sam Altman, CEO, OpenAI

For anyone who has sat through years of vendor assurance questionnaires that amount to 'trust us', an outside team with a badge and a publishing right is a different animal.

Not a pause, and not yet a law

Read it carefully. Amodei says pacing "does not mean halting model training or technical progress" and expects progress will "still seem fast". Neither lab has named a model it will hold back or a date it will slip. Steps two and three need government involvement that does not exist yet, and his own essay concedes that passing laws "can take time". His case for pacing within democracies also rests on keeping the US lead over China wide enough to afford it, so expect export controls and anti-distillation measures to travel with the safety talk.

Not every critic is a booster. TechCrunch quotes journalist Brian Merchant arguing that proposals like this "would likely only wind up serving Anthropic and OpenAI; it's what regulatory capture looks like in action."

them throughMy view? Both things can be true. The people writing these essays believe them, and a world where only the two best-resourced labs can afford embedded evaluators happens to suit those two labs.

What changes for a delivery firm

Most of us do not train models. We buy them through Microsoft, Autodesk, a Claude or ChatGPT enterprise seat, or a start-up with one of those underneath. So what lands on the desk?

The posture behind the models has changed. Both labs have now said scaling should be constrained by confidence in safety rather than by what they can build. I read that as roadmaps that slip when an evaluator says no and release notes that say more about what a model was cleared for. If your AI plan assumes a bigger model every quarter, caveat it.

Audit hooks are coming. If both labs end up with embedded outsiders publishing findings, you get an independent read on the model under your tooling for the first time. Useful. Also something your clients will find before you do.

And the client questions get harder. 'Do you use AI?' became 'which tools?' last year. I would expect 'which model, which version, who evaluated it, and what happens if the vendor pauses it?' to turn up in a PQQ or framework renewal before the year is out, the public sector first. If you cannot answer from a register, you will be answering from memory in a room.

Takeaway

The two labs whose models sit under most of the AI in our sector agreed, on the same weekend, that the frontier needs pacing and that outsiders should have the run of the building. It is neither a pause nor a mandate yet, and it came from the vendors rather than the regulators, which is exactly why clients will take it seriously.

Four things I would do this month:

  • Keep a live register of every AI model and version in use across your projects, including the ones inside other vendors' products.

  • Add a vendor-pause line to your AI risk register: what you do if Claude or ChatGPT is held back or rolled back mid-programme.

  • Draft the answer to 'who has independently evaluated the models you use?' now, and revise it when the evaluator arrangements are announced.

  • Read Amodei's essay and Pachocki's, not the summaries. Your clients' advisers will have.

Staying ahead of what the frontier labs decide next is most of what the Project Flux newsletter is for. Every Monday we track what the labs and the regulators have done and what it means on a live project, so you are not the last to know.

All content reflects our personal views and is not intended as professional advice or to represent any organisation.

Reply

Avatar

or to participate