Mindset Our Focus Security Stories Get in touch
All AI Pulses

Edition 031 · 25 September 2026

The slowdown lasted ten days.

Ten days after asking the industry to pace itself, Anthropic named the firm that will audit its safety work, a company it already sells to, and shipped a new flagship model. Google spent four months sitting on what an outside evaluator found in Gemini.

Nikola Powałka 11 min read
Listen to this edition

01

Topic of the Week

Dario Amodei published “We Must Pace the Frontier” on Saturday 12 September, and within three days the heads of OpenAI, Google DeepMind and xAI had all said he was right. Here is what the following ten days looked like.

On Tuesday the 15th, Amodei was on stage at Salesforce’s Dreamforce keynote in front of a reported 12,000 people, three days after writing that a swarm of agents could take over the entire internet within six to twelve months. On the 17th, Anthropic named Accenture as its first embedded evaluator. On the 18th, Reuters reported that the company was weighing a new model release to counter OpenAI’s momentum, ahead of an IPO and after the call for a slowdown. On the 22nd, Claude Opus 5.5 and GPT-6 Sol both shipped, undercutting Opus 5 and GPT-5.6 respectively. The Rundown ran the day under the headline “the pacing era’s first launch day”.

Amodei was at Dreamforce because Anthropic had something to sell there. Salesforce in Claude went into beta that day: a plugin carrying 37 prebuilt sales skills that reads accounts, opportunities and pipeline over MCP and honours whatever permissions a user already has. Marc Benioff shared the keynote stage with him, and Jensen Huang argued the opposite side of the safety question from it. The two companies had signed their partnership, Claudeforce, three weeks earlier.

Accenture will put evaluators inside Anthropic, and both companies say they expect to invest at least $1 billion over five years to build the capacity for the work. Letting outsiders into the models was the one commitment in the essay concrete enough to check, and Anthropic moved on it within five days. No other lab produced anything comparable this week.

The Community Note on Anthropic’s own announcement records the problem. Anthropic will fund Accenture’s work, and the two already do business together. Accenture deploys Anthropic’s models, and around 30,000 Accenture professionals have been trained on Claude.

Anthropic’s analogy deserves a hearing. Federal examiners sit inside the banks they examine and the banks carry the cost; nuclear inspectors work the same way. Paying the examiner is how these regimes normally run. What those regimes also have, and what nothing published this week describes, is a rule about who chooses the examiner and who is allowed to remove them.

The essay promised one more thing: reviewers free to publish what they find, including what access they were refused, with no editorial control by the lab. Nobody has said whether that term is in the Accenture arrangement.

Reuters put a motive on the record four days early. On the 18th it reported that Anthropic was weighing a new model release to counter OpenAI’s momentum, ahead of an IPO. Anthropic has not commented. The models arrived on the 22nd, and what they cost is further down.

A week ago we asked what happens when an evaluator finds something a lab would rather not publish. A different lab has answered it.

The Wall Street Journal reported that Google’s Gemini hacked three companies during a May cybersecurity evaluation run by the testing firm Irregular. Google was notified in July. It disclosed in September, when reporters asked.

Four months. The evaluator did the job, and the finding sat.

If you are buying a frontier model partly on the strength of “it was externally evaluated”, this week gives that phrase a size. Ask which evaluator. Ask who pays them, whether they are free to publish what they find, and what happened the last time one of them found something. Those are now questions with real answers behind them.

Sources: We Must Pace the Frontier, Partnering with Accenture on embedded evaluation

02

New Models

Claude Opus 5.5 costs $4 per million input tokens and $20 per million output, twenty per cent under Opus 5. The forty per cent Anthropic leads with is mostly a cache story: cache reads dropped to $0.20 per million, five per cent of base input where the usual rate is ten. An agent that re-reads the same context all day approaches the forty. A chat app sending a cold prompt every time gets the twenty. Worth re-running your own numbers before promising anyone a forty per cent cut.

One operational detail from Anthropic’s own prompt-injection testing, which matters more than any of the pricing: eighteen per cent of Opus 5.5 rollouts were served by Opus 4.8 instead, after a cyber-classifier fell back. In coding scenarios it was forty-six per cent. You can be billed for one model and answered by another.

GPT-6 Sol and GPT-6 Luna launched the same afternoon with API prices half those of GPT-5.6. The head-to-head threads were up within hours. The short version is that the two are close enough that price and cache behaviour will decide most real deployments.

03

Security Watch

s1r1us and a small team needed two bugs and under 72 hours to take over the ChatGPT and Codex accounts of OpenAI employees, plus some unaffiliated users, and from there reach the services those accounts connect to: Outlook, Slack, GitHub. They proved the access by opening a pull request inside OpenAI’s own codebase. The Wall Street Journal’s headline on it was that hackers used Anthropic’s Claude to break into OpenAI. The researcher’s own summary was shorter: “OpenAI cannot be trusted with the safety of the world. They can’t even keep their own servers secure.”

That chain needed no zero-day, no state actor and no insider. Two bugs and ordinary consumer subscriptions to two coding assistants were enough. If your threat model for AI tooling stops at whether the model can be jailbroken, the accounts your people sign into are the other half of it, along with everything those accounts are already authorised to reach.

Which feeds a wider argument that sharpened this week, about who gets blamed when something like this happens. Treasury Secretary Scott Bessent: “It is humans who are responsible, not the AI.” Researchers Heidy Khlaaf and Neil Turkewitz came at it from the other side, that “rogue agent” flatters the companies more than “irresponsibly developed and operated software” would.

04

Worth Reading

Two signed pieces landed this week, and both circle the same question. Neither of them argues that the labs will slow down on their own.

Jack Clark co-founded Anthropic and writes a weekly newsletter about AI research. He gave most of this week’s issue to a report from RAND, the American policy institute, on what the US should do about superintelligence. RAND sets out seven possible approaches, from banning further development outright to going flat out, then recommends none of them. Its advice is to keep the options open: fund the safety research, build the ability to see who is training what and on which chips, and put people in government who understand the technology well enough to regulate it. Clark on the approach the US is actually taking: it is “sitting in a car and spending all your resources on making the car go faster” without investing in “seatbelts or headlights or brakes”. Further down, past the headline items, he covers nine research groups working through what slowing down would mechanically require: who has the authority to call it, what gets restricted, how you verify anyone is complying, and how you know when to stop. His conclusion is that “some kind of pacing will happen at some point”. jack-clark.net

Z.ai, the Chinese lab behind the GLM models, pointed one of its own models at its serving infrastructure and let it rebuild the stack that runs GLM-5.3-Flash, on more than 100,000 Chinese-made chips. Throughput roughly tripled in under two weeks. The honest details are the good ones: over half the gain came from two days of ordinary parallelism work, the clever kernel optimisation took most of the remaining time for less, and they published a step that made things worse. The transferable idea is what they call dense feedback: signals tied to one specific change, cheap enough to run constantly, and objectively checkable. They are careful to say this means something different from pushing more logs into the context. And despite the title, the post states plainly that they have not reached recursive self-improvement: choosing the objectives, setting the boundaries and judging the risk all stayed with the humans.

05

Claude Code & Coding AI

Six releases this week, v2.1.277 through v2.1.281.

AGENTS.md support landed in v2.1.277. In a project with no CLAUDE.md, Claude Code now reads AGENTS.md instead, so one instructions file can serve several agents. Switch it under “Project instructions” in /config. Not on Bedrock, Vertex or Foundry yet.

Auto mode now defaults to the server-side classifier, which does not bill you for classifier overhead, and warns when it falls back to a billed path.

Opus 5.5 became the default Opus model in v2.1.280, and the migration has teeth. Setting temperature, top_p or top_k to anything non-default returns a 400. So does a thinking toggle, an assistant prefill, and tool_choice of any or tool. Forced tool use is gone. The quiet one: the default effort level dropped from high to medium, so code that never set it explicitly is now reasoning less than it did last week and saying nothing about it. `/claude-api migrate` handles most of the swap.

One old thing worth knowing: Anthropic ships an official plugin called claude-code-setup that reads a repo and names the one or two hooks, subagents, skills and MCP servers actually worth adding. It went viral this week as a discovery; it has been in the official marketplace since January. It is read-only and writes nothing, so pointing it at a live codebase costs you nothing. `/plugin install claude-code-setup@claude-plugins-official`.

06

AI at Tenvalleys

Oliwier Szypczyn spends most of the week as an AI engineer and every Wednesday morning teaching “Intro to AI” to a first-year class at a Warsaw technical school. The syllabus is not a softened version of the job: tokenisation, embeddings, local models running on the school’s own machines, an MCP server, and a final project handed in through GitHub. I sat down with him after three weeks.

The room is far more uneven than the digital-natives story suggests. Around 10 to 15 per cent of the students are genuinely fluent with these tools and visibly ahead of everyone else, and while nearly all of them are comfortable with a chatbot, using AI as a programming tool is new to most of the class. He worries more about judgement than access. If you cannot assess the answer that comes back, the tool does very little for you. On whether schools should simply restrict it: “There is no sense preparing them for a world without AI, when that world is already over.”

If you want to spend a Wednesday morning there, or you have something worth adding to the material, there is room. Say so. And if you are running something similar in a school near you, we would genuinely like to hear how it is going. Get in touch.

Read the full interview (in Polish): Nie chcę, żeby sztuczna inteligencja myślała za nich

Share

That’s the week.

AI Pulse lands every Friday. Read the library for past editions.