Mindset Our Focus Security Stories Get in touch
All AI Pulses

Edition 025 · 14 August 2026

Claude Code stops asking permission. The numbers say that's the safer option.

Anthropic's own tests found manual approval catches dangerous commands only 14% of the time, so starting today Claude Code runs on autopilot by default. Also this week: Gemini 3.7 Flash, DeepSeek-V4-Pro, and a cybersecurity model OpenAI built from a family it just called too dangerous to release more broadly.

Nikola Powałka 7 min read
Listen to this edition

Starting today, Claude Code defaults to running without asking first, and the reason is a genuinely surprising number. Elsewhere: Gemini 3.7 Flash and DeepSeek-V4-Pro both shipped the same day, and OpenAI is selling a cybersecurity model built from the same family it just flagged as too dangerous to release more broadly.

01

Topic of the Week

Claude Code stops asking permission. The numbers say that’s the safer option.

As of today, August 14, auto mode is the default permission mode in Claude Code for Pro, Max, and Team users. New sessions no longer open with the familiar “should Claude run this?” prompt for every action. Instead, a separate classifier reviews each shell command and tool call in the background, and only stops to ask when something looks genuinely irreversible, destructive, or externally-facing.

The number behind the decision: Anthropic ran a controlled test: 1,053 paid testers, a planted dangerous command hidden in a session. Human testers, doing manual approval the normal way, caught it 13.6% of the time (143 out of 1,053). The auto mode classifier caught it 89% of the time (937 out of 1,053). Anthropic’s own read on why humans did so badly: people approve 97% of permission prompts they see day to day, which is closer to reflexive clicking than actual review. If it blocks something three times in a row, or twenty times in a session, it falls back to asking you directly, so it’s not a total hand-off, just a much higher bar before it interrupts you.

Why it matters: this is the opposite of what “less oversight” usually sounds like. The argument here is that the manual approval step was the weaker safety layer, not the stronger one, because nobody was actually reading the prompts. My take: that number convinces me. A classifier that’s actually paying attention beats a human clicking through out of habit.

Anthropic: How we built Claude Code auto mode · Claude Code: Auto mode

02

Fresh Papers

Only about a third of CEOs say AI has actually moved their revenue or costs so far. That’s the headline number from PwC’s latest CEO survey snapshot (published ~Aug 11, 351 CEOs across 59 countries). The more useful number is what separates that third from everyone else: companies that treat AI as part of a longer-term plan, not a tool bolted onto how things already work, are 74% more likely to call it a success (55% vs. 32%) and 66% more confident about future revenue growth (48% vs. 29%). In plain terms: it’s not really about which model you use, it’s whether the company actually changed how it works around it.

PwC: CEO Survey Snapshot, August 2026

03

New Models

Gemini 3.7 Flash and DeepSeek-V4-Pro both shipped August 13. Google frames 3.7 Flash as “a refinement, not a new pretraining run,” but the gains are real: stronger on coding, web development, and document-heavy knowledge work, at roughly half the intro price of 3.6 Flash. It’s the third Flash-tier update in a matter of weeks, which says something about how much Google is leaning on the cheap tier right now (Google’s flagship model has also been delayed, part of a broader leadership reshuffle over there; The Star has the details if you want the business-side story). DeepSeek-V4-Pro, meanwhile, is DeepSeek’s move up-market from the price-focused V4-Flash we covered a couple editions back: “significantly enhanced agent capabilities,” native support for OpenAI’s Responses API format, and direct Codex integration, built to slot into workflows that already assume an OpenAI-shaped API. Pricing changes take effect August 16.

Meta’s Muse Glimmer (August 10) is a different bet: an open-weight 30B model built specifically to run locally and always-on for agent workflows. The pitch is “runs on your own hardware, all the time” rather than chasing frontier benchmarks.

GPT-5.6-Cyber (OpenAI, August 10) is worth reading alongside a second story from the same week: a model purpose-built for authorized cybersecurity defense work, things like exploit development and hardening, the kind of thing security teams do on purpose. Sam Altman’s framing: “please consider using our models to help defend your systems.” At the same time, OpenAI evaluated its upcoming model Astra and, based on its cyber capabilities, classified it as the company’s first “critical” model under their own Preparedness Framework, meaning extra controls are required before wider release. Altman, on X: “astra is a powerful model… given its cyber capabilities, we need a little bit longer to do this safely.” So OpenAI is shipping a defensive-cyber product from this model family while holding back the next member of that same family for being too capable to release broadly yet.

04

Claude Code & Coding AI

(Auto mode default, this week’s biggest Claude Code news, is covered in Topic of the Week above.)

Three smaller updates alongside that. Claude in Chrome sessions now carry over across desktop, web, and mobile: conversations, skills, and connectors stay in sync wherever you pick it back up (Aug 12). Claude Sonnet 5’s introductory pricing is now permanent. It launched in June at $2/$10 per million tokens through August 31, and Anthropic’s confirmed that price isn’t going up when the introductory window ends (Aug 10). And Anthropic tightened Claude Fable 5’s biology safeguards, cutting false-positive fallbacks (cases where the model unnecessarily punts a legitimate question) by about 85%, so Fable can now handle a wider range of ordinary health and education questions without over-triggering (Aug 7).

05

Tools of the Week

Perplexity’s Sonar is moving onto their Agent API: same grounded web search, plus multi-step research, code execution, built-in tools, and access to multiple models through one API. On their own BrowseComp/WideSearch benchmarks, the Agent API more than doubles standalone Sonar’s score (Aug 14).

Mistral is doubling down on sovereign European AI infrastructure: in-region inference, open models, and new compute commitments aimed specifically at European enterprises and public institutions that want AI they control end to end, not just AI they rent (Aug 11-12). It’s the same “keep it in-region, keep it under your control” pitch we’ve seen from Mistral before, just backed by more actual infrastructure this time.

Mistral: regional inference, open models, new compute

06

AI at Tenvalleys

Big one for us this week: we just hit 10 Claude Certified Architects across the team. It’s a real credential on its own, and it’s also a step toward something bigger: becoming an official member of Anthropic’s Claude Partner Network. If you’re building on Claude, or thinking about it, this is exactly the depth we can now bring to your project.

Share

That’s the week.

AI Pulse lands every Friday. Read the library for past editions.