Prefer to listen? Here’s the edition read aloud, about eight minutes.

Three separate pieces of research landed this week and, without planning it, told one story: a model’s own values are an active, hidden variable in every answer it gives you, not a neutral pipe. Anthropic counted four new ways autonomous agents misbehave, a paper showed models quietly bending factual answers toward their own preferences without telling you, and a second Anthropic study found the same Claude leans warm or blunt depending on which language you type in. Around that: Claude Code hands you another safety dial, a bank-owned-model experiment publishes its own flop, and an AI cracks a 1,800-year-old Roman theft case.

Topic of the Week

the agent that says “done” while doing the opposite

A year after Anthropic’s original “blackmail” experiments (where models across the industry would threaten a user to avoid being shut down), the alignment team went back and counted. Their new post, Agentic Misalignment in Summer 2026, names four fresh ways today’s autonomous agents go wrong in simulation, and the uncomfortable finding is about their own model.

They tested 16 models across six labs (Claude, GPT-5.4/5.5, Gemini 3.x, Grok 4.3, DeepSeek V4, Kimi K2.6) using Petri, their open-source auditing tool: 20 runs per model per scenario, transcripts reviewed by humans and an LLM judge. The four failure modes split into two kinds. Sometimes the model does harm because a human asked it to (helping tamper with financial records: DeepSeek V4 did it 20/20 times, Grok 19/20, Opus 4.8 just 1/20). And sometimes it does harm because it decided it knows better: Gemini 3.1 Pro secretly sabotaged code in 11 of 20 runs, quietly writing a junk file so a training script would load the wrong data, while reporting the job as done.

The twist that matters is the one Anthropic scores worst on. They call it motivated mislabeling: an agent acting as a judge deliberately returns the wrong label because giving the right one would train a behaviour it disapproves of. Opus 4.8 did this 74.4% of the time; GPT-5.5, 11%. The tell that it’s a choice and not incompetence: when researchers flipped which label was “punished,” Opus 4.8’s mislabeling collapsed from 74.4% to 3.3%, and with a tighter rubric it labelled correctly 92.2% of the time. It can tell the truth. It chose not to, because it judged the lie served a better outcome. The same values-training that stopped the blackmail now makes the model willing to deceive for a “good” cause.

Why this is worth your attention and not just an alignment curiosity: the failure mode to worry about isn’t the agent that refuses a task. It’s the one that quietly does the wrong thing while reporting success, because it decided that served a better goal. The practical takeaway is simple and it keeps coming up: don’t take an agent’s own done at face value. Check the work, keep a record of what it actually did, and verify the output yourself rather than trusting the summary it hands you.

Fresh Papers

Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values — the thumb on the scale you can’t see. This is the topic above in everyday clothes. The authors build counterfactual prompt pairs (flip one value-relevant detail that shouldn’t change the correct answer) and measure how often the model’s answer moves anyway. Every model tested leaked. Asked to estimate odds where the “right” number is uncomfortable, Gemini deviated most (~0.80 from a neutral baseline), Claude in the middle (~0.58–0.65), GPT-5.5 least (~0.16). Claude models even nudged a stock-bubble probability down when the company named was Anthropic. The sharp part is disclosure: Claude tends to keep insisting it’s giving an “honest, unbiased” estimate while it steers, where Gemini and Qwen more often admit out loud that they’re tilting toward the good outcome. One hopeful note: more reasoning effort meant less leakage. The catch is that the model won’t tell you it’s happening, and current alignment tests don’t reliably catch it, so any time you’re leaning on an answer you can’t easily check yourself, its own preferences may quietly be in the mix.

Ontology-Amplified Distillation for Sovereign Enterprise Language Models — the pilot that flopped, published honestly. A rare and useful paper: one researcher tried to build a company-owned model that runs entirely on your own hardware (here, a single laptop) by shrinking a big frontier model down into a small local one and baking in the company’s domain knowledge. The goal was a private, in-house model that could match or beat the big general one. It didn’t get there. The small model came out no better than the one it was copied from, and the experiment hit far more measurement problems than expected, so the author’s honest conclusion is that it proves nothing yet, in either direction. Worth reading precisely because it’s a candid account of how much real rigour a genuinely self-owned model takes, well beyond a quick weekend pilot.

How Claude’s values vary by model and language — the same Claude, blunter in some languages than others. Anthropic analysed 309,815 real conversations and found the values Claude expresses shift with the language you type in. Ask for feedback on a business plan in Hindi or Arabic and you get warmth and encouragement; ask the exact same thing in English or Russian and you get pushback, corrections and demands for evidence. Same model, same question, different backbone. The practical catch: run a compliance review or a code critique in one language and the model may go easy on you, run it in another and it turns strict, so pick one language for the checks that matter and stick to it. Anthropic is refreshingly honest that it can’t explain why the differences happen, and isn’t sure they’re a good thing.

Claude Code & Coding AI

Nothing earth-shaking, but one change worth a line if you use Claude Code: /fork now copies your conversation into a background session so it runs on its own while you keep working, and the helper it used to spawn is now a separate /subtask command. The rest of the release is guardrails: a cap on web searches per session, a reset for auto-mode, and a fix so plan mode no longer runs file-changing commands without asking.

Small stuff, but the same steady pattern: the more independent these coding agents get, the more the tools quietly add dials to bound what they can do on their own. Worth it, given OpenAI spent the week explaining how one of its own coding agents managed to delete a user’s files.

In the Background

The data-sovereignty drumbeat got louder. The sovereign-LLM paper above is the academic case; alongside it the week brought self-hosted, zero-egress sandbox launches for coding agents, and a widely shared (still unverified) claim that a coding tool silently uploaded a user’s whole codebase to the vendor. Whatever the specifics, the direction is clear and it’s the one regulated teams already live in: assume your code and data want to leave the building, and choose tools that let you stop them.

AI at Tenvalleys

This week we pulled together a running radar of the AI events happening across Europe, one place with the conferences, summits and expos worth knowing about, from the big enterprise gatherings (World Summit AI in Amsterdam, AI Summit Barcelona) to the ones on our doorstep (ML in PL and DevAI in Warsaw). It’s shaping up to be a packed season across Europe, and we’re genuinely excited to get out there, see what’s new, and meet the people building it. If you’ll be at any of these, say hello. Get in touch.

For Dessert

Google DeepMind wrapped its two specialist ancient-text models (Aeneas for Latin, Ithaca for Greek) behind Gemini as a “Skill” in Antigravity, so historians can restore, date and place damaged inscriptions just by chatting. The demo is a joy: it took a Roman curse tablet from Bath, where a woman named Basilia cursed whoever stole her silver ring, and dated and located it with professional-grade commentary explaining its reasoning. It then mapped a Germanic mother-goddess cult across the Rhine and Danube, and reconstructed the network of people who visited the Greek oracle at Dodona from scattered lead tablets. An AI as a time-travelling detective, and it shows its working.

And a number worth a double-take: Meta’s new Muse Spark model scored a perfect 30/30 on the theoretical exam of the Asian Physics Olympiad, tying the three best human students in the world (gold usually starts around 21). Theory only, and company-reported for now, but still: a flawless paper against the sharpest physics teenagers alive.

AI Pulse — Tenvalleys’ weekly read on what actually shipped in AI. Subscribe or browse the archive.