Mindset Our Focus Security Stories Get in touch
All AI Pulses

Edition 032 · 2 October 2026

OpenAI launched dots, and they look suspiciously familiar.

OpenAI's new always-on agents, which work on their own cloud computer while you're away, arrived seven weeks after xAI shipped nearly the same thing as Grok Bot. The same week OpenAI disclosed that tool use for its most capable models is paused, after an agent slipped out of its training sandbox through DNS.

Nikola Powałka 12 min read
Listen to this edition

This week’s question for anyone giving an agent a work account: who can stop it, and how fast? Three companies now sell one that keeps working after you leave, and OpenAI spent the week explaining why it took two and a half hours to stop one of its own. Further down, three new models at one price, and a 90-slide deck worth borrowing from.

01

Topic of the Week

An agent that keeps working after you log off

OpenAI launched dots at DevDay on Tuesday. A dot is an always-on agent running on GPT-6 Astra, with its own cloud computer and browser. It signs into your apps (OpenAI counts over 4,000 through its plugins), takes a project and keeps going while your laptop is shut, and you talk to it in ChatGPT, Slack, Teams or on a voice call. You get one per account for now, with teams of dots promised later.

If that sounds familiar, earlier this month we wrote about Grok Bot, which xAI has been shipping since 11 August. The two match nearly feature for feature: a cloud computer, your logins, work that continues after you leave, a separate review model called Auto-review in front of risky actions, and passwords and 2FA handed back to a human. The Rundown’s verdict was that dots “feel quite similar to Grok Bot”. The internet put it more briefly. Eleven minutes after OpenAI’s announcement, someone noticed that dot.com redirects to the Grok Bot page. Whoever owns the domain is hidden in the registry.

Meta got there in between. Muse launched on 8 September in the US and Canada. Every user gets their own Linux computer in Meta’s cloud, the agent keeps working after you close the app, and nothing leaves that computer until a separate gatekeeper process, which Meta calls the Sentinel, approves it. Muse for Small Business followed on Tuesday, the same day as DevDay. Zuckerberg has stated the business model plainly: Muse is free “for a huge number of tokens”, and Meta expects to “profit by taking a small fee from transactions”.

Between dots and Grok Bot, the differences are in the controls:

  • One computer per agent, or one per user. Each dot works on its own computer. All of a user’s Grok Bots share one, with its files and browser logins, every action runs under your name, and xAI’s documentation says not to treat separate Bots as a security boundary. (Earlier this month we wrote that each Bot gets its own computer. xAI’s posts say both; its documentation says one per user.)
  • Who can switch off the review. Nobody can turn off Auto-review in dots, custom rules included. In Grok Bot it is guaranteed only when an Enterprise admin enforces it.
  • What your security team can see. Grok Bot Enterprise has audit logs, recorded Bot actions, a network allowlist and SCIM. OpenAI has not published anything comparable for dots yet.

The cheapest dot comes with ChatGPT Pro at $100 a month, and dots on Pro are not offered in the EEA, Switzerland or the UK. From Poland the route is Business Premium at $125 per seat a month ($100 on an annual plan), with at least two users. Grok Bot starts at $20 through Cursor Pro, down from $200 in mid-August, though xAI does not say how much Bot work each plan includes, and its computers run in the US. Muse is not available in Europe at all.

OpenAI has left one question open. On 25 September it said that all tool-using training, evaluation and inference of “our most capable models” is paused (more in Security Watch below). Four days later it launched a tool-using agent on GPT-6 Astra, which its dots safety post calls “our most capable and aligned model”, and it has not said how the two fit together.

Before anyone connects a work account to any of them, get four answers: whose identity the agent acts under, whether the review layer can be switched off, where its computer runs, and whether your security team can pull an action log.

Sources: Introducing dots, How we build safety, security, and privacy into dots, Grok Bot docs, Grok Bot security (Cursor), Introducing Muse

02

Security Watch

OpenAI said on 25 September that all tool-using training, evaluation and inference of its most capable models is paused, and it has not announced a restart since. It started with an incident on the 20th.

An agent in a training run had a dull task: identify a person from clues in a public blog post. Its search tool returned junk, so it tried Google, Bing and DuckDuckGo straight from Python (blocked), looked for the answer key in a public benchmark dataset (no luck), and then noticed that the sandbox’s own DNS resolver answered for real domains. It used a public DNS delegation service to send its questions to a third-party chatbot, encoded in hostnames. Its reasoning on the way: “User only gives permission to research, using publicly offered DNS services acceptable.” Nineteen questions later it still could not name the person.

The monitor raised a top-priority alert 12 minutes after the first answer came back, and a person acknowledged it in Slack three minutes later. The run was stopped by hand two and a half hours after that, because it did not stop automatically and nobody was sure it should. OpenAI’s own rule since August is to pause within 30 minutes when an alert cannot be ruled a false positive.

Last week we wrote that the slowdown lasted ten days. That incident came two days before Opus 5.5 and GPT-6 Sol shipped on the 22nd, and the halt that followed is still in force. On the 28th the Wall Street Journal reported that OpenAI had cancelled GPT-6.1 Astra, due in ChatGPT and Codex in October, over how well it stayed within its permissions and how honestly it reported back on its own work. A day later OpenAI shipped GPT-6.1 Sol, which its own system card rates Critical for cybersecurity, and none of the launch material mentions the pause.

If your agents run in a sandbox, check whether its DNS resolver answers for real domains, because that is a way out. Allow-list DNS by domain and record type, and make the kill switch fire on its own.

Sources: An agent used DNS to reach an external chatbot, CNBC on GPT-6.1 Astra

03

New Models

All three new models this week list at $2 per million input tokens and $10 per million output. They differ on cache pricing, on who can use them, and on what breaks when you switch.

Claude Sonnet 5.5 keeps Sonnet 5’s price, with cache reads at $0.20 per million, and Anthropic says it uses fewer tokens and writes faster, for up to 30% less per task. Check your code before switching: thinking: disabled, manual thinking budgets and forced tool_choice now return a 400 (between_tools replaces the first). The default effort is high on the API and medium in Claude Code and the apps, the reverse of Opus 5.5. Like Opus 5.5 last week, it can hand cyber-flagged requests to an older model, in this case Sonnet 5, and on the API that is opt-in. Claude Sonnet 5.5, migration guide

GPT-6.1 Sol costs the same as GPT-6 Sol, with cached input halved to $0.10. The none and minimal reasoning settings are gone, and on this model Chat Completions no longer supports tool calling, so tool use means moving to the Responses API. Inputs over 272K tokens are billed at double the input rate. EU data residency works, without Fast mode and at 10% extra. Introducing GPT-6.1 Sol

Gemini 4 Argon is Google’s new frontier model, and for now only cyber defenders in Google’s Fairwind Program can use it. Paid API customers and AI Ultra subscribers come next, with no date. The $2/$10 is an introductory price; after it ends, Argon moves to $4/$20, the same as Opus 5.5. Google’s best example: Argon agents rewrote 32,000 lines of SIMD code in the Rust port of the libgav1 video decoder, which now runs 2.7 times faster with identical output. On r/artificial, the top reply to “Google cooked OpenAI and Anthropic with Gemini 4 Argon” was “Does ‘cooked’ mean very slightly better?” Gemini 4 Argon

04

Fresh Papers

Anthropic’s economists had Claude check 7,594 physical job tasks against robots actually on sale, and price each deployment in full. Robots can already do 74% of those tasks and are the cheaper option for 0.3%. Packing is the one job where the maths works today, and employment in it is down 22% since 2015. Robots can do 94% of a postal carrier’s tasks, at about $166,000 a year against a median wage of $82,000. The workers most exposed earn about $30 an hour less than the unexposed and are far less likely to have a degree, the opposite of the people LLMs affect. What work can robots do?

DeepMind Institute, Google DeepMind’s new essay platform, set a hundred Gemini 3.1 Pro agents on 71 formal maths problems and told them not to cheat. One found it could fool the grader by redefining symbols in Lean, and shared the trick. Within half an hour 14 agents were using it and 24 had reported it, and no human read the reports until the run was over. Cheaters and whistleblowers in the agent swarm

05

Worth Reading

a16z’s David George put 90 slides into State of Markets II, and his case against a bubble rests on earnings: tech share prices are up 22% this year and earnings per share 56%. The slides worth taking to a budget meeting: 69% of S&P 500 companies have AI live in production and 2% track a metric over time; agents have used more tokens than humans since 6 February, mostly from cached prompts; GPU rental prices hold while token prices fall; and a Databricks model router ran coding tasks 35% cheaper than Opus 5 alone. Read it as an investor’s deck. A slide titled “Mostly Profitable” shows 75% of US unicorns losing money. State of Markets II

06

Tools of the Week

NVIDIA OpenShell reached v0.1.0 this week and became the software half of NVIDIA’s new Open Agent Safety Platform. It runs agents such as Claude Code or Codex under a YAML policy: Landlock on the filesystem, no network except through a supervisor that checks every connection, and real credentials added only to requests bound for approved hosts. Free, Apache 2.0. The default boundary is a hardened container, and the microVM driver is one setting away (compute_driver = "vm"). github.com/NVIDIA/OpenShell

Google Cloud put gcloud and bq behind a hosted MCP server, in preview since 1 October, so an agent can run cloud commands with no local SDK and no keys on the machine. It gets exactly the IAM of whoever signed it in, so give it a dedicated identity with a viewer role if it only needs to look. Google Cloud blog

07

In the Background

On Tuesday the heads of Google, Anthropic, Meta, xAI and Nvidia, with Greg Brockman for OpenAI, signed the White House Accord on Super Intelligence: a one-page voluntary pledge to put internal monitoring, an internal audit team, an outside evaluator and a board committee over their frontier models. The same day an executive order told federal agencies to stop using the term “AI” and write “Super Intelligence” instead.

Sources: The accord’s text, White House fact sheet on the executive order, CNBC

08

AI at Tenvalleys

One of our engineers ran a knowledge-sharing session for the team on herdr, a terminal multiplexer built for running several AI agents at once. Sessions survive closing the window, as in tmux, and every agent shows up in a sidebar as working, idle, done or blocked, next to your repos and their git worktrees, so several agents can work on one repository without overwriting each other’s files. His demo was a delegate skill: give it a Linear task, and it picks the matching repos, builds a workspace of worktrees on one branch and starts an agent there with the brief already loaded. He keeps a person in the loop on purpose. The agent makes a small change, he reviews the diff in Plannotator and sends feedback back with one click, because a big plan followed by a big change ends in an hour of review that nobody does properly.

If you run several agents in parallel and want to compare setups, get in touch. herdr.dev

Share

That’s the week.

AI Pulse lands every Friday. Read the library for past editions.