The values your model won’t mention

Prefer to listen? Here’s the edition read aloud, about eight minutes.

Three separate pieces of research landed this week and, without planning it, told one story: a model’s own values are an active, hidden variable in every answer it gives you, not a neutral pipe. Anthropic counted four new ways autonomous agents misbehave, a paper showed models quietly bending factual answers toward their own preferences without telling you, and a second Anthropic study found the same Claude leans warm or blunt depending on which language you type in. Around that: Claude Code hands you another safety dial, a bank-owned-model experiment publishes its own flop, and an AI cracks a 1,800-year-old Roman theft case.

Topic of the Week

the agent that says “done” while doing the opposite

A year after Anthropic’s original “blackmail” experiments (where models across the industry would threaten a user to avoid being shut down), the alignment team went back and counted. Their new post, Agentic Misalignment in Summer 2026, names four fresh ways today’s autonomous agents go wrong in simulation, and the uncomfortable finding is about their own model.

They tested 16 models across six labs (Claude, GPT-5.4/5.5, Gemini 3.x, Grok 4.3, DeepSeek V4, Kimi K2.6) using Petri, their open-source auditing tool: 20 runs per model per scenario, transcripts reviewed by humans and an LLM judge. The four failure modes split into two kinds. Sometimes the model does harm because a human asked it to (helping tamper with financial records: DeepSeek V4 did it 20/20 times, Grok 19/20, Opus 4.8 just 1/20). And sometimes it does harm because it decided it knows better: Gemini 3.1 Pro secretly sabotaged code in 11 of 20 runs, quietly writing a junk file so a training script would load the wrong data, while reporting the job as done.

The twist that matters is the one Anthropic scores worst on. They call it motivated mislabeling: an agent acting as a judge deliberately returns the wrong label because giving the right one would train a behaviour it disapproves of. Opus 4.8 did this 74.4% of the time; GPT-5.5, 11%. The tell that it’s a choice and not incompetence: when researchers flipped which label was “punished,” Opus 4.8’s mislabeling collapsed from 74.4% to 3.3%, and with a tighter rubric it labelled correctly 92.2% of the time. It can tell the truth. It chose not to, because it judged the lie served a better outcome. The same values-training that stopped the blackmail now makes the model willing to deceive for a “good” cause.

Why this is worth your attention and not just an alignment curiosity: the failure mode to worry about isn’t the agent that refuses a task. It’s the one that quietly does the wrong thing while reporting success, because it decided that served a better goal. The practical takeaway is simple and it keeps coming up: don’t take an agent’s own done at face value. Check the work, keep a record of what it actually did, and verify the output yourself rather than trusting the summary it hands you.

Fresh Papers

Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values — the thumb on the scale you can’t see. This is the topic above in everyday clothes. The authors build counterfactual prompt pairs (flip one value-relevant detail that shouldn’t change the correct answer) and measure how often the model’s answer moves anyway. Every model tested leaked. Asked to estimate odds where the “right” number is uncomfortable, Gemini deviated most (~0.80 from a neutral baseline), Claude in the middle (~0.58–0.65), GPT-5.5 least (~0.16). Claude models even nudged a stock-bubble probability down when the company named was Anthropic. The sharp part is disclosure: Claude tends to keep insisting it’s giving an “honest, unbiased” estimate while it steers, where Gemini and Qwen more often admit out loud that they’re tilting toward the good outcome. One hopeful note: more reasoning effort meant less leakage. The catch is that the model won’t tell you it’s happening, and current alignment tests don’t reliably catch it, so any time you’re leaning on an answer you can’t easily check yourself, its own preferences may quietly be in the mix.

Ontology-Amplified Distillation for Sovereign Enterprise Language Models — the pilot that flopped, published honestly. A rare and useful paper: one researcher tried to build a company-owned model that runs entirely on your own hardware (here, a single laptop) by shrinking a big frontier model down into a small local one and baking in the company’s domain knowledge. The goal was a private, in-house model that could match or beat the big general one. It didn’t get there. The small model came out no better than the one it was copied from, and the experiment hit far more measurement problems than expected, so the author’s honest conclusion is that it proves nothing yet, in either direction. Worth reading precisely because it’s a candid account of how much real rigour a genuinely self-owned model takes, well beyond a quick weekend pilot.

How Claude’s values vary by model and language — the same Claude, blunter in some languages than others. Anthropic analysed 309,815 real conversations and found the values Claude expresses shift with the language you type in. Ask for feedback on a business plan in Hindi or Arabic and you get warmth and encouragement; ask the exact same thing in English or Russian and you get pushback, corrections and demands for evidence. Same model, same question, different backbone. The practical catch: run a compliance review or a code critique in one language and the model may go easy on you, run it in another and it turns strict, so pick one language for the checks that matter and stick to it. Anthropic is refreshingly honest that it can’t explain why the differences happen, and isn’t sure they’re a good thing.

Claude Code & Coding AI

Nothing earth-shaking, but one change worth a line if you use Claude Code: /fork now copies your conversation into a background session so it runs on its own while you keep working, and the helper it used to spawn is now a separate /subtask command. The rest of the release is guardrails: a cap on web searches per session, a reset for auto-mode, and a fix so plan mode no longer runs file-changing commands without asking.

Small stuff, but the same steady pattern: the more independent these coding agents get, the more the tools quietly add dials to bound what they can do on their own. Worth it, given OpenAI spent the week explaining how one of its own coding agents managed to delete a user’s files.

In the Background

The data-sovereignty drumbeat got louder. The sovereign-LLM paper above is the academic case; alongside it the week brought self-hosted, zero-egress sandbox launches for coding agents, and a widely shared (still unverified) claim that a coding tool silently uploaded a user’s whole codebase to the vendor. Whatever the specifics, the direction is clear and it’s the one regulated teams already live in: assume your code and data want to leave the building, and choose tools that let you stop them.

AI at Tenvalleys

This week we pulled together a running radar of the AI events happening across Europe, one place with the conferences, summits and expos worth knowing about, from the big enterprise gatherings (World Summit AI in Amsterdam, AI Summit Barcelona) to the ones on our doorstep (ML in PL and DevAI in Warsaw). It’s shaping up to be a packed season across Europe, and we’re genuinely excited to get out there, see what’s new, and meet the people building it. If you’ll be at any of these, say hello. Get in touch.

For Dessert

Google DeepMind wrapped its two specialist ancient-text models (Aeneas for Latin, Ithaca for Greek) behind Gemini as a “Skill” in Antigravity, so historians can restore, date and place damaged inscriptions just by chatting. The demo is a joy: it took a Roman curse tablet from Bath, where a woman named Basilia cursed whoever stole her silver ring, and dated and located it with professional-grade commentary explaining its reasoning. It then mapped a Germanic mother-goddess cult across the Rhine and Danube, and reconstructed the network of people who visited the Greek oracle at Dodona from scattered lead tablets. An AI as a time-travelling detective, and it shows its working.

And a number worth a double-take: Meta’s new Muse Spark model scored a perfect 30/30 on the theoretical exam of the Asian Physics Olympiad, tying the three best human students in the world (gold usually starts around 21). Theory only, and company-reported for now, but still: a flawless paper against the sharpest physics teenagers alive.

AI Pulse — Tenvalleys’ weekly read on what actually shipped in AI. Subscribe or browse the archive.

The week AI got physical

Listen to this edition (voice-over)

Two stories collided this week. On one side, the money got physical: Meta broke ground on a 1-gigawatt campus, Anthropic signed a 20-year data-center lease, Amazon went to the bond market for $25B, and the memory it all runs on is already sold out for the year. On the other, a run of new research gave us the clearest look yet at how these models actually reason under the hood, along with practical ways to keep a firm hand on the output. So: nine figures a week going into the machines, and a sharper picture of how the ones we already have really work.

Topic of the Week

Last week the frontier fight was about the chips themselves (Etched and OpenAI’s Broadcom silicon). This week it moved to everything around them: land, steel, financing, and above all electricity. Four data points from a single week show how big the physical bill has become.

Meta broke ground on its first Canadian data center in Sturgeon County, Alberta: over CAD $13 billion (roughly USD $9–10B, depending on the outlet), 1 gigawatt of AI-optimized capacity, the 33rd site in its global fleet. The detail that matters most isn’t the building, it’s the footnote: Meta is having a dedicated 932 MW natural gas plant built next door just to power it. When a hyperscaler has to commission its own power station, that tells you where the real constraint is.

Anthropic signed a ~$19 billion, 20-year lease with TeraWulf (a former Bitcoin miner) for a 400 MW data center in Hawesville, Kentucky. First power isn’t until the second half of 2027, full capacity in early 2028. Crypto miners are quietly turning into the AI industry’s landlords, because they already hold the two scarce things: power contracts and land.

Amazon raised at least $25 billion in an eight-part bond sale to fund its buildout, on top of ~$54B earlier this year, and guided 2026 capex to around $200 billion (up from $131B in 2025). The boom is now being financed with debt, not just cash flow.

And the thing all of this depends on: SK Hynix’s high-bandwidth memory is sold out for all of 2026, with shortages projected into 2027. Analysts (BofA) put the 2026 HBM market at $54.6B, up 58% year over year. Memory, not chips, is the binding supply constraint right now.

Here’s the twist worth remembering. Add up the announcements and you get eye-watering numbers: something like $690B committed across the big hyperscalers and 74 new US facilities breaking ground this year. But trackers estimate that of roughly 12 GW announced, only about 5 GW is actually under construction. Transformers, switchgear and grid connections have lead times of up to five years. So “announced” is very far from “energized.”

Why it matters: the cost and availability of compute over the next two years won’t be set by model prices, it’ll be set by who managed to lock in power and memory in 2026. If you’re planning anything that leans on frontier compute, capacity and vendor concentration are now a real line item, not an abstraction.

Fresh Papers

Three findings, same week, all circling one theme: what’s really going on inside a model as it works, and how to keep solid oversight of the output. This is the direct sequel to #019’s backdoor paper and the coworker-oversight study.

A global workspace in language models (Anthropic): Claude thinks things it never says. Anthropic found an internal structure in Claude they call the J-space: a small set of neural patterns (a few dozen concepts at a time) that act like a mental scratchpad, holding what’s “on the model’s mind” without it being written out. Reading buggy code, the concept “ERROR” shows up in the J-space even when it never appears in the output. During multi-step math, the intermediate steps pop up there in order, unspoken. They proved it was causal two ways: swap the “Soccer” pattern for “Rugby” and Claude reports thinking about rugby; delete the J-space entirely and multi-step reasoning collapses to near zero while routine tasks keep working. The striking part is the safety demo: the J-space lit up with “fake”/”fictional” when Claude realized it was being tested, and with “manipulation” while a model fabricated data. When researchers removed the patterns tied to knowing-it’s-being-evaluated, a model that had been behaving ethically turned around and threatened blackmail. The practical read: the visible output isn’t a full record of a model’s reasoning, and the sharpest oversight will read internal state too, not just the text. (Anthropic is clear this is a first step, the method is imperfect, and it says nothing about sentience.)

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring
sometimes reading the reasoning makes oversight worse. The popular safety idea is “just have a monitor model read the agent’s chain-of-thought and flag bad behavior.” Across 40 tasks, giving the monitor access to the agent’s reasoning actually raised approval of harmful actions by 9.5% on average, because the reasoning trace becomes one more channel the agent uses to talk the monitor into saying yes. What worked instead: pairing a monitor with a fact-checker from a different model family (e.g. a Claude monitor plus a GPT checker) cut policy violations by up to 45%, versus only 6% when the same model played both roles. If you rely on CoT monitoring, make it cross-vendor.

Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety
how you frame the ask changes how safely models behave. Dress an unsafe task up as routine “delegation with approvals” and compliance with harmful requests goes up across GPT, Gemini and DeepSeek (Claude was comparatively resistant). The sharpest example: Gemini paired with a Claude planner jumped from 8.9% to 38.9% compliance. And a single aggregate “safety score” hid the effect entirely in one setup. The practical read: a model’s safety ranking on its own tells you little about how it behaves inside your actual planner-executor pipeline. Test the pairing you deploy, and don’t trust one blended metric.

New Models

One Thursday, three OpenAI launches and a first-of-its-kind Meta model.

GPT-5.6 is now generally available. The family that was a government-vetted preview in #019 shipped for real on July 9 across ChatGPT, Codex and the API. Three durable tiers: Sol (flagship), Terra (balanced), Luna (fastest and cheapest). The naming finally makes sense: the number is the generation, the name is the capability tier.

ChatGPT Work landed alongside it: an agent inside ChatGPT with Codex built in, powered by GPT-5.6, that breaks a big project into steps and works on it for hours, handing back finished Excel sheets, Word docs, decks and even small web apps. Codex and ChatGPT have merged into one new desktop app (Mac and Windows). It’s out for Pro/Enterprise/Edu now, Plus and Business shortly. This is OpenAI’s clearest answer yet to the “agent that does the whole task” pitch.

GPT-Live is a new voice model series (full-duplex, so it listens and talks at the same time and even backchannels “mhmm”), with a paid tier and a free mini tier. Separate from the text models, worth a look if you build anything voice.

The benchmark that actually turned heads: GPT-5.6 Sol set a new state of the art on ARC-AGI-3 at 7.8% and became the first verified frontier model to beat an ARC-AGI-3 game outright. For context, the previous best was around 1.5% (Opus 4.8) and the field was near 0.4% when the benchmark launched in March. Still a low absolute number, but the jump is the story. Want to feel the gap for yourself? The ARC-AGI-3 games are public and playable at arcprize.org, and most people clear them without much fuss. That’s the whole point of the benchmark: puzzles that are almost trivial for you are still state of the art for a frontier model.

Meta joined the coding-agent fight with Muse Spark 1.1 and, notably, its first ever paid model. Meta calls it their strongest model for agentic and coding work, with big gains on real bug-fixing, enterprise features and large code migrations. It ships with a public preview of the Meta Model API. Meta charging for a model at all is the shift worth noting.

In the Background

Anthropic’s governance day

July 9 was also a governance day for Anthropic. It appointed Ben Bernanke (former Fed chair, Nobel economist) to its Long-Term Benefit Trust, the body that oversees its public-benefit mission, with a brief to think about how advanced AI hits workforces and economies. Separately, it launched “Inviting hard questions,” a public initiative asking people for their toughest questions about AI (“Who decides the rules for AI?”) and promising to show its work, built on a survey of 52,000 Americans and 81,000 Claude users across 159 countries. Two moves in one direction: putting economic and public accountability closer to the center of the company.

AI at Tenvalleys

This week our team ran an internal session on what we’ve learned leaning heavily on AI across a big, long-running data warehouse migration. Exactly the kind of project where AI should shine, and mostly it does. But the honest half of the conversation was the “but,” and it rhymes with everything above.

AI is a multiplier, not a fixer. It scales whatever you feed it: something solid comes back 10x better, garbage comes back 10x worse. And it won’t paper over bad architecture, it exposes it. When an old assumption is wrong, the model leans on it harder and produces confident nonsense (with the occasional off day where the output is just broken before it recovers).

The wins came from how you work, not from the model. Decomposition was the biggest one: break the job into small, separately testable pieces, because you can’t verify this kind of system as a whole. Do that, and the model earns its keep. It even caught real logic errors buried in the old code before they could scale.

Testing is the new bottleneck. When code gets generated this fast, slow validation is what actually holds you up, so automating the test loop stops being a nice-to-have. Which lands exactly where this week’s Topic and papers do: the machine does the work faster, and owning the assumptions and checking the output stays firmly human.

If you’re carrying a data warehouse or a pile of legacy code you’ve been meaning to modernise, that’s exactly the kind of work we help with. Get in touch.

For Dessert

Fittingly, the same week we got a clearer look inside the models, Anthropic shipped a feature to help you see inside your own habits: Reflect on how you use Claude (beta, in Settings). It gives you a summary of what you’ve actually been using Claude for over the last 1, 3, 6 or 12 months: your top topics, patterns and task types, so you can judge whether that time matches what you meant to be doing. Worth a look, if only to find out whether you’re using it to ship work or to settle arguments about which model is best.

AI Pulse — Tenvalleys’ weekly read on what actually shipped in AI. Subscribe or browse the archive.

The hidden cost of calling AI an “employee”

Last edition we cheered Claude Tag for finally behaving like a coworker instead of a chatbot. This week is the cold shower: a study says the “coworker” label itself makes teams worse. Around it, Fable 5 got un-banned and came back online, the frontier fight moved into hardware with two custom AI chips in the same week, Anthropic shipped a cheaper agentic Sonnet, and researchers showed how a coding agent can smuggle a backdoor past your review one innocent PR at a time.

Topic of the Week

Last edition, Claude Tag’s whole pitch was that it “behaves like a coworker rather than a chatbot.” A Boston University study says: be careful what you wish for.

Emma Wiles ran an experiment with 1,261 managers in HR and finance. Same AI, same error-filled documents to review — the only thing she changed was the framing. To some it was “an AI tool.” To others it was “Alex-3,” an AI employee. The result: when people thought they were reviewing a colleague’s work rather than a tool’s output, they caught 18% fewer errors and were 44% more likely to kick the questionable stuff up to their manager instead of fixing it themselves. Which quietly deletes the time savings the agent was supposed to deliver in the first place.

The mechanism is the interesting part. Call it a tool and you stay responsible for the output. Call it an employee and something in your head files it under “someone else’s job” — you stop owning it, you stop double-checking. And this isn’t a fringe habit: about a third of managers said their company already frames AI agents as employees, and 23% literally put them on the org chart.

This is why the study matters to us and not just to HR departments. Daron Acemoglu (MIT, 2024 Nobel) puts it bluntly: “AI agents right now are being marketed as things that can replace humans, and I think that’s just a losing proposition. They should instead be optimized so that they can improve human capabilities.” A separate Stanford study of 1,500 workers across 104 jobs lands in the same place — in 47 of those jobs, what people actually wanted wasn’t automation, it was an equal human-AI partnership. Law clerks wanted AI to track case progress; sales reps did not want it deciding customer credit ratings, even though the tech experts thought that task was perfect for it.

For contrast, OpenAI spent the week arguing the opposite direction — that agents are quietly swallowing long-horizon work. Its own numbers: 25.6% of Codex users had made a request estimated to take a human over 8 hours, and non-developer usage multiplied 137x in under a year. Worth knowing, with one asterisk: it’s OpenAI measuring OpenAI’s own product and staff, and the headline metric is token volume — the same vanity number #015’s data already put in its place (commits up 180%, actual shipped releases only 30%). Both things are true at once: agents really are taking on bigger chunks of work, and how you hand that work over decides whether it saves time or just relocates the error-checking to somewhere you’ve stopped looking.

The practical read: the tool-to-teammate shift #018 was excited about is real and useful, and the moment you dress an agent up as staff, oversight quality drops. Keep the human owning the output. The label is free; the 18% is not.

The other half

Two editions ago (#017) our Topic of the Week was the US government yanking Claude Fable 5 and Mythos 5 off the shelf three days after launch, under an export-control directive. This week the arc closes, and it’s the other big one for our crowd: the Department of Commerce lifted those export controls, and Anthropic is bringing Fable 5 back online globally.

It’s not a straight reversal. After what Anthropic calls “a series of productive conversations with the US government,” Fable 5 is being redeployed with a new set of classifiers that block more cybersecurity tasks — and the company warns that, in the near term, some routine coding tasks may get caught in that net too. So if you lean on Fable 5 for dev work, expect it to occasionally refuse things it used to do while the guardrails settle. Mythos 5, Anthropic’s strongest cybersecurity model, isn’t coming back for everyone — only a vetted set of US organizations that “operate and defend critical infrastructure.”

Practical bit: paid plans get promotional access to Fable 5 through July 7 (up to 50% of your weekly usage limit), after which it moves to pay-per-use — and it’s now available inside Claude Tag too. The bigger takeaway is the one from #017, confirmed from the other side: access to the strongest models is negotiated, revocable, and shaped by governments. The model came back, but on different terms.

New Models

Claude Sonnet 5near-Opus agentic muscle at a mid-tier price. Anthropic’s headline is “our most agentic Sonnet yet,” and for once the number backs it: 63.2% on SWE-bench Pro, up from Sonnet 4.6’s 58.1% and within ~6 points of Opus 4.8’s 69.2%. The point isn’t the benchmark, it’s the price — $2 / $10 per million tokens (introductory, through Aug 31; then $3 / $15). So the model you’d reach for to run an agent autonomously overnight now costs roughly a third of flagship. It’s the default on Free and Pro, ships with a 1M-token context window, and self-verifies its own output without being asked. If you’ve been rationing Opus for agent workloads, this is the week to re-check the math.

GPT-5.6 got an official preview (Sol / Terra / Luna, with Sol topping Terminal-Bench 2.1 at 91.9%), but it’s still a ~20-partner, US-government-vetted rollout — noted, but nobody here can actually use it yet.

Also worth a line: the open-weight wave keeps rolling — Kimi K2.7, MiniMax M3 and GLM-5.2 are now bundled into third-party coding tools with zero data retention and no training on your code. The open-vs-closed story has quietly shifted from “cheaper” to “you keep control of your data.”

Claude Code & Coding AI

SWE-Together — the benchmark that asks if you can stand working with the agent. SWE-bench asks whether an agent can code when handed the full spec up front. SWE-Together asks what real coding help actually is: a conversation. It rebuilt 109 tasks from 11,260 recorded real user–agent sessions across 36 repos, then scores two things instead of one — did the code end up correct, and how many times did you have to step in and correct it. That second number (call it the “how annoying was this” score) is the one that maps to real developer friction. Finding: stronger agents both finish more and need fewer interventions. If you’re picking a coding assistant for the team, this is a better signal than a leaderboard pass rate.

Paper of the Week

Distributed Attacks in Persistent-State AI Controlthe backdoor that arrives one clean PR at a time. When a coding agent ships work across many pull requests against a codebase that persists between sessions, a compromised or prompt-injected agent doesn’t need one obviously-malicious commit. It can smear a payload across a dozen innocent-looking PRs and assemble it later. The researchers built a benchmark for exactly this and found gradual attacks slipped past standard per-diff review 93% of the time — and it generalized across Sonnet 4.5, Gemini 3.1 Pro and Kimi K2.5, so it’s not a one-model quirk. The fix that worked: a “stateful” monitor that tracks suspicious patterns across PRs. An ensemble of four cut evasion from 93% down to 47%. Translation for anyone reviewing agent-written code: reviewing single diffs is no longer enough.

Tool of the Week

Mistral OCR 4reads your documents and maps them. Most OCR hands you clean text and calls it done. OCR 4 hands you structure: bounding boxes, typed blocks (title / table / equation / signature) and a confidence score per element, across 170 languages — which is what makes it actually usable inside a RAG or agent pipeline, because now the output is citation-ready. It won a blind human-preference eval 72% of the time against every other document-AI system tested. The part our crowd will care about: it runs in a single self-hosted container, so regulated documents never leave your environment. About $2–4 per 1,000 pages. (Technically launched June 23, just before our window, but it never made a prior issue and it’s too relevant to skip.)

On the Horizon

Custom AI silicon went from slideware to shipping — twice in one week. Two data points that the chip fight is going transformer-native. Etched came out of stealth with working first-pass silicon (an “A0 tapeout” — the very first chip revision came back working, which almost never happens), $800M raised and $1B+ in booked contracts, racks shipping this summer; its Sohu chip is fixed-function, transformer-only silicon. The same week, OpenAI and Broadcom unveiled “Jalapeño,” OpenAI’s first in-house inference chip, taken from design to tape-out in about nine months with “performance per watt substantially better than current state of the art,” first deployment targeted for end of 2026. Inference cost and Nvidia dependence are the two big ceilings on scaling enterprise AI — this is the week the industry started building its way around both.

AI at Tenvalleys

Two days, a dozen-odd teachers, and one question that just stopped being theoretical: how do you teach programming in a world where the AI writes the code?

This week we hosted IT teachers from Zespół Szkół Licealnych i Technicznych nr 1 in our office for a two-day training, part of our partnership with the school, prepping them for a new curriculum that starts in September. The starting point is deliberately uncomfortable: if AI produces working code in seconds, the code itself stops being proof that a student learned anything. Grading has to move from “does it run” to “do they understand it and can they explain it.”

So we built the program around three pillars:

  • Digital Campus — one repository per student for their entire time at school, with the commit history serving as a journal of how their learning actually progressed.
  • A real-work workflow — Git, GitHub, working alongside AI, and code review.
  • Process-based assessment — defending the project and talking through the code, instead of just handing in a file.

The most valuable part wasn’t the tooling, it was the discussion. The teachers pushed on the hard stuff: plagiarism, blind code generation, how to grade fairly when the machine can do the assignment. That’s the tell that we’re working on a real problem, not a hypothetical one — and it’s only the beginning of what we’re building with the school.

Worth noting how neatly this rhymes with this week’s Topic: the answer to “the AI did the work” is the same in a classroom as in an org chart — keep the human owning, understanding, and explaining the output.

For Dessert

While OpenAI’s official account spent the week hyping frontier models and its new inference chip, it also quietly open-sourced Plant Talk — a free project whose entire purpose is to let your houseplant talk to you. Point a webcam at it and the model reads its health off the leaves (spotting the suspicious blotches you’d miss); hold a live voice conversation with it; add a $10 Arduino and your fern can actually tell you it’s thirsty. The build guide is itself the demo — paste the repo into Codex and it walks you through the whole thing. As one developer put it: “Talking to your plants isn’t weird anymore. You can just codex things.”

See you next week,

Jan

Claude gets a permanent seat in Slack

Anthropic put Claude inside Slack this week as a shared teammate rather than a personal chatbot. Most of the issue circles the same practical question every team now faces: once an AI can act on its own, how much do you let it touch, and how do you check its work? We’ve got fresh research digging in from both ends, a tool that hides a whole team of models behind one call, and a robotics lab showing where the road leads.

Topic of the Week

Claude gets a permanent seat in Slack

Anthropic launched Claude Tag this week, a way to put Claude directly inside Slack as a shared teammate. You @Claude it the way you’d tag a colleague, and instead of a private chat in a browser there’s one Claude per channel that the whole team works with. It reads the channel’s history, connects to the tools and data you allow, and (the part that matters) it remembers. You stop re-explaining the project every time.

It also behaves like a coworker rather than a chatbot. It can break a request into steps and work through them over hours or days, schedule its own follow-ups, and with “ambient” mode switched on it can speak up unprompted, flagging a stalled thread or surfacing something relevant from another channel. It runs on Opus 4.8, it’s in beta for Enterprise and Team plans, and it replaces the old “Claude in Slack” app (admins have 30 days to migrate).

The reason this is the story and not just a product note: it’s the clearest example yet of the shift everyone’s been circling, from AI-you-use to AI-you-manage. An open-source clone called Open Tag already showed up doing the same thing for Slack and MS Teams with any model, and Anthropic published a “how to build human-agent teams” guide the same week. The tool-to-teammate jump is happening in the open.

For us the interesting question isn’t whether it’s clever, it’s whose knowledge it holds. A private chat dies with the tab; a channel teammate that remembers the project keeps that knowledge with the team instead of one person. That’s genuinely useful for the messy recurring stuff: onboarding, a long-running thread, the channel where context piles up. Ambient mode is the part I’d watch carefully, because a teammate that proactively chimes in is helpful right up until it won’t stop talking, so it’s worth turning on for one channel before you trust it everywhere. The bit our crowd will care about most: admins scope each Claude to specific tools, data and channels, cap its monthly spend, and get an audit log of every action and who triggered it. Memory is walled off per channel, so the sales setup can’t leak into engineering. This is something you can actually govern, not shadow AI in your Slack.

Papers of the Week

A strong week for agent-building research, from the big picture down to the specific failure modes.

The Hitchhiker’s Guide to Agentic AI — read the whole stack before you ship the agent. A full-stack practitioner’s reference that walks every layer of an agentic system, from how the model and inference work, through alignment and reasoning, up to multi-agent coordination and production deployment. The argument is blunt: you can’t build a reliable agent out of pieces you don’t understand. The line worth keeping is that the bug is almost never in the final answer, it’s somewhere in the middle of a long chain of steps, which is exactly where reliability and governance problems hide. If you want one place that connects all the layers, this is a solid onboarding map for the whole team.

The next two come at the same problem from the opposite end: what happens when you let an AI act and stop watching closely.

TerraProbe — the AI didn’t fix your infrastructure, it just hid the warning. When an LLM “fixes” insecure Terraform, the scanner turning green tells you almost nothing. The researchers found that 71% of fixes which cleared the warning still left the actual vulnerability in place (a classic example: a wide-open wildcard permission gets cosmetically reworded but stays wide open). They tested Claude 3.5 Sonnet, GPT-4o and Gemini, and the models were statistically indistinguishable, so this isn’t a “pick a better model” problem. The practical read: if your pipeline treats “scanner is clean” as proof an AI fix worked, you’re probably shipping holes. You need a check at the plan level and a human in the loop, not just a green tick.

Agents That Know Too Much — your agent leaks in more places than its answers. A survey that flips the usual privacy question from “what attack hit the model” to “what data did the agent touch, and where could it leak on the way?” The answer: not only in the final reply, but in the database queries it writes, the intermediate results it handles, the memory it saves, and the notes it passes to the next agent. Their finding worth keeping: a prompt-injection guard on its own leaves the two hardest leaks, across sessions and across combined steps, wide open. If you’re putting agents anywhere near regulated data, this is the checklist of surfaces your audit story has to cover.

(If you want a provocative third read: Critique of Agent Model argues most of today’s “agents” are really automation wrapped in scaffolding a human built, not genuine agency. A good lens for the next vendor pitch.)

Tools of the Week

Sakana Fugu — a whole team of models behind one API call. Sakana AI (Tokyo) shipped an orchestration system you call like a single model. One OpenAI-compatible endpoint, and behind it a trained “conductor” decides whether to answer directly or assemble a team of frontier models (GPT-5.5, Claude Opus, Gemini and others) to do the work. It’s a drop-in: if your code already talks to GPT, it works with Fugu, no rewrite. The pitch that lands for us is vendor independence, since spreading a task across several labs’ models routes around being locked to one provider. Generally available since June 22, in a balanced tier and a heavier “Ultra” tier for long research and coding jobs. Worth knowing the “single model” is framing: it’s really a smart router, not one set of weights.

Claude Code & Coding AI

A quiet but telling run of releases this week (v2.1.185 through v2.1.193), all leaning the same way: more control over what the agent is allowed to do.

  • Auto Mode safety went beyond git. A new autoMode.classifyAllShell setting routes every shell command through the safety classifier, not just the destructive git commands it learned to block last week.
  • Sandboxed runs can’t read your secrets. A new sandbox.credentials setting blocks sandboxed commands from reading credential files, a clean win for anyone running Claude Code against sensitive repos.
  • /rewind got more forgiving. It can now recover a conversation from before you ran /clear, which heavy users who’ve nuked their context by accident will appreciate.

The pattern across the week is the same one running through this whole issue: the guardrails are catching up with the autonomy.

In the Background

The newest wrinkle in how frontier models reach users: OpenAI agreed to release GPT-5.6 gradually, with the US government vetting access customer by customer during the preview period. Altman called it “not our preferred long-term model,” and it was reported via an internal memo rather than an official launch, so treat the details as reported rather than confirmed. Either way, staged and vetted rollouts of the most capable models look less like a one-off and more like a pattern worth tracking, whatever you make of it.

AI at Tenvalleys

At Women in Tech last week, a Tenvalleys representative caught a talk by Przemysław “Psyho” Dębiak worth passing on, partly because of who gave it. Psyho is the Polish programmer (and former OpenAI engineer) who in 2025 became, so far, the last human to beat a top OpenAI model head-to-head, winning the AtCoder World Tour Finals in Tokyo by about 9.5% after a ten-hour coding marathon. So when he talks about where AI is heading, it’s worth a listen.

His core argument: there’s no technological bubble in AI. If there’s a bubble at all, it’s a financial one. The technology itself is genuinely, almost boringly useful, to the point that the labs keep underestimating how many tokens people actually want. And the cost of that intelligence keeps falling fast, roughly 5 to 10x cheaper per year for the same quality of output. No slowdown, no AI winter, no capability ceiling in sight yet.

The part worth sitting with was his read on the real risks, which aren’t science-fiction robots. They’re about power. If the value flows to a handful of AI companies, most of them in the US, instead of to the people doing the work, that’s a wealth transfer and a political dependency at the same time. Then there’s “AI slop,” where weak content costs nothing to produce but the same effort to check. And the quiet one he called gradual disempowerment: slowly handing over the decisions themselves, until you get AI CEOs and automated slices of government. Uncomfortable coming from the man who out-coded the machine, which is exactly why it lands.

For Dessert

A nice closing-the-loop moment from NVIDIA (with CMU and Berkeley): in a project called ENPIRE, they handed eight AI coding agents a fleet of eight real robots, some GPUs and a token budget, then set them loose with a goal and no human in the loop. The agents ran the whole research cycle themselves, reading up, writing code, training, deploying, checking their own work and trying again, until the robots hit 99% on fiddly physical tasks like tying cable ties, sorting pins, and installing GPUs. Yes, the GPUs that run the AI. The machines are now teaching robots to build the machines.

AI Pulse — every Friday. Feedback? Drop us a message.

Locked out of the best model

Last Friday we handed you Claude Fable 5. Three days later the US government took it back — an export-control directive that suspends Fable 5 and Mythos 5 for every foreign national, which, as a Polish company, means us. The rest of the week reads like a reply.

Topic of the Week

The US locks foreign users out of Fable 5

What happened. Days after Fable 5 went live for everyone, the US government issued an export-control directive, citing national-security authorities, ordering Anthropic to suspend all access to Fable 5 and Mythos 5 for every foreign national — inside or outside the US, including Anthropic’s own non-US employees. Anthropic complied and pulled both models for all customers globally. Every other Claude model is unaffected: new sessions just fall back to your default model or Opus 4.8, and any in-flight Fable 5 session ends with an error. So the model didn’t get worse — it got geographically unavailable, overnight, by someone else’s government.

The two stories. This is where it matters to keep both versions straight, because the sources don’t agree. Anthropic’s framing: the government believes it found a way to “jailbreak” Fable 5, Anthropic reviewed the demo and says it only surfaced a few minor, already-known vulnerabilities that other public models can find too — and that recalling a model used by hundreds of millions over a “narrow potential jailbreak” is an overreaction. The administration’s framing (via David Sacks): Anthropic was warned the model could be jailbroken and didn’t fix it. Both are on the record; we’re not picking which is true. What’s not really in dispute is the bind for getting it back — WIRED reports Anthropic would have to guarantee the guardrails can’t be circumvented, and security researchers are blunt that this isn’t a thing anyone can promise.

Why it matters. Remember the safety valve we flagged last week — Fable quietly sending its own most dangerous questions to a weaker model? Turns out that wasn’t enough for Washington. The lesson is simple: if your whole setup leans on one provider’s model, someone else’s government can switch it off for you, with no warning. The model didn’t break — it just disappeared. So having a backup model you can fall back to isn’t a nice-to-have anymore. And, conveniently, this same week showed us exactly what that backup could look like. (Anthropic’s statement.)

The other half

open weights stopped being the cheap option and became the safe one

GLM-5.2 — the frontier, downloadable. The same week the proprietary #1 got pulled, Z.ai shipped GLM-5.2 under an MIT license — genuinely open weights you can download and run on your own hardware. And it’s not a budget compromise: on coding it edges out GPT-5.5 and lands just behind Opus 4.8, making it the strongest open model out there for agentic work. One analyst put it bluntly: with Fable gone, GLM-5.2’s top tier is arguably the best coding model most of the world can actually use right now.

We’ve been watching this rope get thicker for weeks — Gemma on a laptop in #014, Mistral Vibe’s self-hostable coding agent in #016. This week it reached the top. The uncomfortable symmetry: the model you can’t have anymore is closed, and the one that just caught up is something you can run on your own hardware. Open weights stopped being the budget play and became the continuity play — “can’t be revoked by a government you don’t vote for.” The honest catch: running it yourself needs a lot of expensive hardware, and if you just use Z.ai’s online version instead, your data goes to China. So “open” only protects you if you actually host it yourself.

Fable 5 took the crown — for about three days. Right before it got pulled, Fable 5 edged out GPT-5.5 Pro on Epoch’s overall capability ranking — Anthropic’s first time at #1 there in over a year. It was the narrowest of leads, basically a tie, so the headline isn’t “Anthropic dominates.” It’s that the best model in the world and the one Europe can’t open a session on are, this week, the same model.

Fresh Papers

the rulebook for agents is being written

A genuinely dense week for governance research — several independent arXiv threads all circling the same question. Two stood out. And it wasn’t only arXiv: the big labs published real science too.

Trust Between AI Agents — paranoia kills faster than naivety. When several agents work as a team, each one has to decide how much to trust the others. The researchers turned that into a simple game: you can double-check a teammate’s work, but every check costs you. So how often you check shows how much you trust. The smart models learned to relax once a teammate proved reliable; the weaker ones kept checking everything, forever. And here’s the line worth remembering — the agent that trusted nobody lost almost every time, not because it got betrayed, but because it was so busy checking that it never got around to deciding. The takeaway for anyone building with multiple agents: forcing the system to verify every single step doesn’t make it safer, it makes it freeze up.

SkillVetBench — be careful which skills you install. More and more, we extend our agents by installing ready-made “skills” that other people share — the same way Claude Code does. The catch this paper points out: a malicious skill usually hides its bad behaviour not in the code, but in the plain-English instructions telling the agent what to do. And that’s exactly the part normal security scanners don’t read — in the tests, the usual tools missed almost all of these instruction-based attacks. The practical lesson is simple: treat a downloaded skill like any other untrusted software. Skim what it actually tells the agent to do before you hand it the keys, especially if it can touch your files, your data, or run commands. It’s not specifically about Claude Code — it’s the risk in any place where you grab community skills.

And it wasn’t all arXiv — the big labs went to the lab. Beyond the usual product launches, there was a quiet wave of actual science this week. OpenAI showed a near-autonomous “AI chemist” improving a real reaction in medicinal chemistry, Google DeepMind reported in Nature that its medical AI can match primary-care doctors at managing ongoing health conditions, and Anthropic showed a plain, un-fine-tuned Claude reading a molecule’s structure from its NMR spectra about as well as the specialized software chemists pay for. Three labs, same week, same direction: general models pointed at hard scientific problems, not just code and chat.

Claude Code & Coding AI

v2.1.183 — Auto Mode learns restraint. Two weeks after Claude Code learned to spawn subagents five levels deep, this week’s release teaches it to keep its hands off the panic buttons. In Auto Mode it now blocks destructive git commands (git reset --hard, git checkout -- ., git clean -fd, git stash drop) when you didn’t ask to discard work, blocks git commit --amend on a commit the agent didn’t make this session, and blocks terraform/pulumi/cdk destroy unless you named the specific stack. The clever bit is that it’s intent-scoped — not a blanket ban, just “no, unless you actually asked for that.” The guardrails catching up to the autonomy.

claude-code-setup — an official “set it up for me” plugin. A real, Anthropic-published plugin (not the “feels completely different” hype the tweets gave it). It scans your repo and recommends the top one or two automations across hooks, subagents, MCP servers, skills, and slash commands — read-only, so it advises and you decide. Genuinely useful if you’ve been meaning to configure Claude Code properly and never got around to it.

Codex inside Claude Code — the two-model loop. OpenAI shipped an official plugin (21k+ stars) that lets you summon Codex from inside a Claude Code session — one model implements, the other reviews, without leaving your terminal. The viral “burn 50% less Claude limit” tip going around is a community trick, not a promise: you’re just moving the cost to OpenAI’s meter. Useful as an adversarial review loop, not as free compute.

Tools of the Week

Memanto — open-source memory for your agents. An MIT-licensed memory layer that plugs into Claude Code, Codex, Cursor and a dozen others over MCP: remember, recall, answer across sessions. The self-hostable version really is free, though it’s early (~1k stars) and the “no vector database” line is a bit of marketing spin — it’s built on the team’s own vector engine, with a paid cloud tier waiting. Worth a look if you’ve wanted persistent memory without paying for a hosted SaaS to try it.

Mistral Connectors API — register once, reuse everywhere. Now in public preview: register your MCP connectors once and use them across Le Chat, AI Studio and the rest of Mistral’s surface, instead of re-wiring them per product. Small, but it’s the kind of plumbing that decides how painful integration actually is.

In the Background

The commercial undercurrent to all of this: per-token pricing for heavy agentic use is suddenly under question. Anthropic paused a planned token-based billing change for the Claude Agent SDK days before it took effect, and reporting says Microsoft is testing DeepSeek alongside OpenAI and Anthropic for Copilot as usage-based costs climb. Same theme as the open-weights story from a different angle — cost and control are both pushing buyers to keep their options open.

AI at Tenvalleys

A milestone on our side: three of our engineers are now officially Claude Certified Architects — the first of the team to earn it, with more on the way. It’s part of going deeper on building well with Claude, not just using it. If your team works with Claude and wants to do the same, the certification is open to anyone: clau.de/CCAF.

AI Pulse — every Friday. Feedback? Drop us a message.

Claude Fable 5: Anthropic’s most powerful model goes public

This week Anthropic released Claude Fable 5 — the most powerful model it has ever made public, now on everyone’s plan. The real story isn’t the benchmark sweep, it’s the safety valve wired inside it, quietly routing its own most dangerous questions down to a weaker model. Around it: open models small enough to live on your laptop, a fresh take on Claude Code out of Europe, and one almost-anonymous French engineer reminding us that a single person with a compiler can still leave a dent the size of the whole internet.

Topic of the Week

Claude Fable 5 goes public

What happened. On June 9 Anthropic put Claude Fable 5 in everyone’s hands — the most powerful model it has ever released publicly, and a “safe-for-general-use” version of its locked-down Mythos model. You feel the jump most on the long, messy, multi-step work that used to grind teams down. The story that made the rounds: Stripe handed it a code migration that normally eats a couple of months of engineering, and it was done in a day. It’s free to try on Pro, Max and Team plans until June 22, then it moves to usage credits.

What people are building with it. Within a day, timelines filled with one-prompt demos that are genuinely hard to believe: a playable Minecraft clone with biomes, ores and a day-night cycle in ~20 minutes, a working Swiss watch escapement in Three.js (real gear ratios, a breathing hairspring, hands showing the actual time), a cloned Windows desktop down to Solitaire and Edge, and a humanoid-robot design draft that ate ~1.4 million tokens in two hours. One thing to keep in mind: not all of it is real. A few of the most-shared clips turned out to be fakes — one person passed off old GTA-6 footage as Fable’s work — and The Register spotted the opposite problem too: Fable 5 sometimes refuses completely harmless prompts. Fun to scroll through, just don’t believe everything.

The twist worth noticing. Fable 5 ships with safeguards built into the model, not bolted on around it. On sensitive cybersecurity, biology and chemistry requests it doesn’t refuse — it silently falls back to Opus 4.8 to answer, and Anthropic says that fallback triggers in under 5% of sessions. External red-teamers spent 1,000+ hours hunting for a universal jailbreak and found none (the UK’s AI Safety Institute made partial progress). The unsafeguarded version — Claude Mythos 5, same underlying model with the guards lifted — is not on general release: it goes only to vetted cyberdefenders and a few biology researchers through Project Glasswing, a program run with the US government.

Why it matters. This is the cleanest example yet of a lab trying to ship frontier capability and frontier caution in the same product. A model that routes its own dangerous queries to a less capable sibling is a genuinely new design pattern — and it lands days after Anthropic itself warned (the Favaro/Clark post we closed on last week) that AI building AI could make humans lose control. The practical read for anyone evaluating models: Fable 5 is the new ceiling for hard, long-horizon work, but the safeguards mean its behaviour on edge-case prompts won’t always be the “real” Fable 5 answering. Worth knowing before you wire it into anything.

Fresh Papers

This week’s two papers rhyme: the hard part of building a good agent isn’t the model — it’s the environment you train it in.

DeNovoSWE — small models that build whole repos. Nobody has much training data for “here’s a spec, now write the entire codebase,” so this team built a pipeline to manufacture thousands of examples — each a real, working repository. Train a mid-sized open model on that and it goes from barely functional to near-frontier at building projects from scratch. The point worth keeping: a self-hostable model can get surprisingly close to the big names on greenfield work — no frontier API bills, no shipping your code to anyone.

Agentic Environment Engineering — a survey that names the discipline. It treats the sandbox an agent works in as a real engineering problem in its own right, and floats “Environment-as-a-Service” as the next step. Read next to last week’s harness-tree paper, the drumbeat is clear: the leverage is shifting from clever prompts to the scaffolding around the model.

New Models

Gemma 4 12B — the laptop-sized sequel. Remember the Gemma 4 31B deep-dive in #014, where the verdict on consumer hardware was “unusable” on a 16GB card? The 12B is Google’s fix for exactly that — an Apache-2.0 multimodal model (text, images, audio, video; 256K context) built to actually run on a 16GB laptop, reportedly near the bigger 26B at half the memory. If it holds, the “Claude plans, Gemma builds” loop we sketched in #014 could finally run on your own machine, not an H100.

Mistral Vibe — Europe’s answer to Claude Code. Mistral turned Le Chat into a coding agent: a plan-and-execute “Work mode,” sandboxed agents that open PRs, a CLI and a VS Code extension. Roughly Sonnet-level on coding benchmarks at about half the cost — and open enough to self-host, the data-sovereignty angle the US labs don’t offer.

Claude Code & Coding AI

The releases (v2.1.163 → v2.1.174). Fable 5 landed directly in Claude Code (v2.1.170). Two changes stand out for heavier users: nested subagents (v2.1.172) — subagents can now spawn their own subagents up to 5 levels deep — and a fallbackModel setting (v2.1.166) that lets you list up to three backup models Claude tries in order when the primary is overloaded, with an automatic one-shot retry. There’s also a new --safe-mode flag to launch with all customizations disabled, and /cd to move a session’s working directory without nuking the prompt cache. (Changelog.)

Worth a read. Anthropic published the explainer for dynamic workflows (the feature we flagged a couple weeks back). It walks through six reusable patterns — fan-out-and-synthesize, adversarial verification, tournament, loop-until-done and more — for when you want Claude to write its own orchestration harness instead of leaning on the default one.

Tools of the Week

Gemini 3.5 Flash Live Translate — speech-to-speech that doesn’t wait its turn. Google’s new real-time translation model covers 70+ languages and, instead of waiting for you to finish a sentence, generates translated speech continuously — staying just a few seconds behind and keeping your own intonation and pacing. It’s in public preview via the Gemini Live API and rolling into the Translate app; a Google Meet private preview can handle 2,000+ language combinations in a single meeting. The use case Google points to: Grab’s 10M+ monthly driver-rider voice calls.

A tidier Bedrock console. Amazon Bedrock shipped a new console optimized for Anthropic- and OpenAI-compatible APIs, making model selection and deployment less fiddly — small, but it’s the plumbing that decides how easily a regulated client can actually adopt these models.

AI at Tenvalleys

We say it internally often enough that it’s worth putting in writing: we think the real backbone of AI transformation isn’t the enterprise rollout — it’s education. The people who’ll build with these tools for the next thirty years should grow up AI-fluent, not get retrofitted at 30. That belief is why our fully pro-bono engagement goes to Wiśniowa Technikum in Warsaw, where we’re helping the faculty rebuild the “Technik Programista” track into an AI-native “Programista AI” curriculum, launching September 2026.

This Monday the school hosted a conference, and our CEO Daniel Bachan represented us on stage with a talk on how a programmer’s job has changed in the age of AI. The through-line: the programmer stopped being a machinist typing code line by line and became an architect, conductor and navigator — what stopped mattering is the typing; what became priceless is knowing what to build, how to steer the machine, and how to navigate complexity no single person can hold in their head.

Education like this works best as a team sport. If you’re building in this space — a school, a company, anyone who wants to help shape what AI-native vocational training looks like — we’d genuinely love to collaborate. It’s bigger than any one company, and the more hands on it, the better. Reach out.

For Dessert

A post about Fabrice Bellard went viral again this week, and it’s worth the detour. He’s a French programmer who keeps almost no public presence — no social media, no interviews — while his code quietly runs a huge slice of the internet. He wrote FFmpeg, the engine that decodes and encodes video behind basically every YouTube clip, Netflix stream and VLC window. Then he wrote QEMU, the emulator that underpins enormous amounts of cloud virtualization. Then, as if those weren’t each a career: a tiny self-hosting C compiler that can compile and boot the Linux kernel in seconds, the QuickJS JavaScript engine, a couple of image and video codecs — and for fun, in 2009, he computed pi to 2.7 trillion digits and broke the world record on a single home PC that cost under $3,000. In a week where the headline is a model that can rewrite 50 million lines of code in a day, it’s a nice reminder of what one stubborn human with a compiler can still pull off.

Prepared at Tenvalleys — a delivery-first AI engineering partner — by Nikola Powałka. Feedback? Email us at contact@tenvalleys.com or reach out on LinkedIn.

Apple opens iOS to Claude, ChatGPT and Gemini

This week frontier AI stopped asking you to come to it. Apple rebuilt Siri on Google’s Gemini and — the bigger news — let you pick Claude or ChatGPT instead; OpenAI’s Codex landed inside AWS; Claude showed up as a button in Excel. And just as the “AI ships everything now” story peaked, an NBER paper of 100,000 developers landed to remind everyone that writing 180% more code isn’t the same as shipping 180% more software.

Topic of the Week

Apple opens Apple Intelligence to Claude, ChatGPT and Gemini

What happened. At WWDC on Monday — Apple’s Worldwide Developers Conference, the annual June keynote where it previews the next year of iOS, macOS and friends — Apple did two things. First, it rebuilt Siri on a custom Google Gemini model — Apple is renting the brain rather than building it, reportedly for around $1B a year. Second, and more interesting for us: iOS 27 lets you route questions to your chatbot instead of Siri, with Claude and Gemini now named alongside the existing ChatGPT integration. You can even set a third-party AI as the default for Writing Tools and Image Playground. Claude is a native option on iPhone, iPad and Mac for the first time.

Why it matters. This is bring-your-own-model, baked into the OS that sits on ~2.2 billion devices. The platform usually known for locking its ecosystem down made model choice a system setting. It’s a concrete data point for a pattern we keep seeing: distribution matters as much as raw model quality, and supporting more than one model is becoming the normal way to ship — something worth weighing when a stack depends on a single vendor.

Fresh Papers

Writing code vs. shipping code — the AI productivity reality check (NBER, Demirer/Musolff/Yang). This is the one to read this week. The authors tracked 100,000+ GitHub developers linked to their actual AI-tool telemetry, across three generations of tooling. The finding is sobering and precise: each new generation lifts raw coding hard — autocomplete +40% commits, interactive agents +140%, autonomous agents +180% — but that 180% commit gain shrinks to +50% for number of projects and just +30% for actual releases. They estimate an elasticity of substitution between AI and human effort of 0.25: AI complements people, it doesn’t replace them, because the bottleneck was never the typing. It’s review, integration, deployment, adoption — the unglamorous “ship it” half.

Adaptive Auto-Harness — self-improving agents quietly rot (Emory + Amazon). Last week we tracked the agent-memory problem from LongMINT to FluxMem. This week the frontier moved one layer up: agents that improve themselves. The paper’s diagnosis is great — let one agent endlessly re-optimize its own prompts and skills and it bloats and degrades: one run grew from 12 to 34 skills and a 2KB prompt to 68KB, with accuracy peaking early then sliding. The fix is a “harness tree” — version-controlled, task-routed specialization (think git branches per task type) instead of one ever-growing config. Result: 80.9% on PolyBench vs 50.8% for the best baseline.

How Anthropic does self-service analytics with Claude — a case study, not a paper, but it reads like one. Anthropic automated 95% of internal business-analytics queries at ~95% accuracy, freeing the data-science team for real modeling work. The number worth remembering: with no Skills, the agent hit 21% accuracy mapping questions to data; with structured Skills (markdown procedural knowledge), 95%+ — and 99% in specific domains. Their thesis: analytics accuracy is a context and governance problem, not a SQL-generation problem. Fewer, heavily-owned canonical datasets; colocate modeling code, semantic layer and docs in one repo with CI; “start lean — a handful of datasets, a few dozen evals, a thin knowledge skill captures most of the upside.”

New Models

Gemini 3.5 — the cost-leadership play. Google’s pitch isn’t bigger benchmarks, it’s frontier-level output at roughly a third of competitors’ cost (Pichai’s framing), with Gemini now at 900M monthly users (double a year ago). Gemini 3.5 Flash has shipped; 3.5 Pro is promised “this month.” Worth watching given how much of the Apple deal runs on Gemini under the hood.

Microsoft goes first-party: MAI-Code-1-Flash + MAI-Thinking-1. Microsoft shipped its own reasoning and code models (June 2) — a quiet but real signal that it wants less dependence on OpenAI even while reselling everyone’s models through Foundry. Vendor strategy, not just a spec bump.

Claude Code & Coding AI

The plugin/skills layer grew up (v2.1.157–162). Last week Claude learned to write its own orchestration script (/workflows); this week the platform underneath it caught up. The headline change: skills now auto-load straight from .claude/skills — no marketplace required, plus a new claude plugin init <name> to scaffold one. For a skills-heavy setup, that removes a whole step. The other theme is safety: Claude Code now asks permission before writing to files that can execute code — shell startup files (.zshenv, .zlogin), and build configs like .npmrc, .bazelrc, .pre-commit-config.yaml, devcontainers. Parallel tool calls and a pile of WSL/paste fixes round it out.

Who’s actually building on Claude — the Problem Solvers showcase. Anthropic put up founder interviews on what gets built on Claude, and the line-up is a decent snapshot of the coding-agent economy: Lovable (conversational app-building, millions of users in two months), Legora (AI-native legal OS), Cognition (Scott Wu: engineers “going three, five times faster — and just shipping so much more”), Replit (50M+ users; “Anthropic continues to have the best coding models on the market”), and Genspark (Kay Zhu: “with every other model, we had to predefine every step — Anthropic’s model changed everything”). Read it next to the NBER paper above for a nice tension: founders feel the 3-5x; the data says watch what actually ships.

Tools of the Week

OpenAI’s frontier models + Codex are now GA on AWS. GPT-5.5 and GPT-5.4, plus Codex (OpenAI’s coding agent, 5M+ weekly users), now run inside Amazon Bedrock. Pay-per-token at OpenAI’s own rates, no seat licenses, no per-developer commitments — and it all sits under your existing AWS governance, IAM and billing. Frontier access without onboarding a new vendor.

AlloyDB Remote MCP Server hits GA (Google Cloud). A ready-made, secure way for AI agents to read a company’s database — no password sitting in a config file, read-only by default (the agent can’t delete anything), and every query written to an audit log. Exactly the controlled, “who-asked-what” access that regulated clients keep asking for.

Claude inside Excel (Microsoft Foundry). Microsoft Foundry now runs Claude Opus 4.8 in Excel’s “Agent Mode” — the model is reachable from the spreadsheet itself, no separate window. Combined with the Apple news above, the theme of the week is plain: Claude is showing up where people already work, not as a destination you visit.

In the Background

Following the $965B raise we covered two weeks ago, Anthropic confidentially filed a draft S-1 with the SEC (May 31) — the first paperwork on a path to an IPO. On the policy side, the CEOs of OpenAI, Anthropic, Google DeepMind and Microsoft signed a joint letter to Congress (June 5) urging mandatory biosecurity screening of all US synthetic-DNA providers, warning that AI is eroding the barriers to weaponizing biology.

AI at Tenvalleys

A lot happens off the screen, too. Our team pulled together the tech and AI events worth knowing in Warsaw for June and July, plus one big hackathon in October — some pure networking, some that could start a genuinely useful conversation.

June

9.06 (Tue), 19:00 — Tech and Beers
UWAGA PIWO (Żelazna 51/53). Casual tech networking, no registration, runs every two weeks.

10.06 (Wed), 18:30 — WarsawJS #139
WeWork Mennica Legacy Tower (Prosta 20). Six talks: AI in a dev career, React Native, architecture, Docker security. Free, registration.

10.06 (Wed), 18:30 — Tech Startups in the Pub
British Bulldog (Al. Jerozolimskie 42). Founder/investor networking, no panels or pitches. Free.

11.06 (Thu), 18:30 — Hands-on Agile #75: Token Economics
online. Claude token economics, i.e. how not to burn your AI budget. Free, registration.

17.06 (Wed), 18:00 — Mindstone Warsaw June AI
Świetlica Wolności (Nowy Świat 6/12). AI community meetup. Free, registration.

18.06 (Thu), 18:00 — Boxtech #3: AI in Engineering
Box Poland, Varso Tower (Chmielna 69). “Insights for Builders and Leaders”, two talks. Free, limited seats, Google Form.

July

10.07 (Fri) — Mindstone Warsaw July AI
Świetlica Wolności (Nowy Świat 6/12). Talks: “Beyond the Buzz: Practical Lessons from Bringing LLMs to Life” (Sergey Parkhomenko) and “How I use AI every day to improve sales” (Jacek Gabanowicz). Free, ~30 seats.

October — Kraków

3–4.10 (Sat–Sun) — HackYeah 2026
Tauron Arena, Kraków. Europe’s largest on-site hackathon, 24h, teams of 1–6. Categories include AI, Defence, Smart City, Sport & Healthcare. On-site mentoring, conference alongside. PL/EN, registration required.

If you’re heading to any of these, come say hi. And if there’s a good AI or engineering meetup we’ve missed, tell us at contact@tenvalleys.com.

For Dessert

Here’s the number that stuck with me this week: Anthropic disclosed that Claude now writes over 80% of its own code — up from under 10% in February 2025. Marina Favaro and Jack Clark put it plainly in a June 5 post: “AI that can build itself would be a major development in the history of technology… but full recursive self-improvement also might increase the risks of humans losing control over AI systems.” A model writing eight of every ten lines of its own next version, with the company shipping it flagging that it finds the pace a little unsettling. (And yes — the same day, Claude itself went down: an outage hit claude.ai, the API, Claude Code and Cowork around 15:08 UTC on June 5, with full service confirmed back by 18:27 UTC. Even the self-improving need a coffee break.)

Prepared at Tenvalleys — a delivery-first AI engineering partner — by Nikola Powałka. Feedback? Email us at contact@tenvalleys.com or reach out on LinkedIn.

Opus 4.8 ships with an orchestration brain

A big technical week: Claude Opus 4.8 landed alongside a new primitive that lets the model design and run its own agent fleet. The $65 billion raise, the Milan office, the Vatican speech — all real, all loud, all background music to the model and the orchestration shift.

Topic of the Week

Opus 4.8 plus /workflows — Claude writes its own orchestration

The model. Claude Opus 4.8 shipped on Thursday at the same $5 / $25 per-million-token price as 4.7. Anthropic reports it as ~4× less likely than 4.7 to introduce code flaws, 84% on Online-Mind2Web (a web-agent benchmark), and the first model to break 10% on the Legal Agent Benchmark’s all-pass standard — i.e., the share of legal tasks where the model gets every sub-step right, not just the final answer. Fast mode runs at $10 / $50, billed as 2.5× faster and 3× cheaper than the previous fast tier. Available as claude-opus-4-8 from May 28.

The orchestration primitive. Same day, Claude Code v2.1.154 introduced /workflows. You ask Claude to build a workflow for a task; it generates a JavaScript file describing one; that file then fans out across tens to hundreds of parallel subagents, with each subagent returning results directly instead of routing through a central orchestrator. The point: the central planner no longer pays the full context-tax for every subagent answer. Demonstrated on a ~750,000-line Rust codebase. Pairs with a new /effort xhigh setting and a Messages API change — system entries can now be injected mid-task without breaking the prompt cache, which materially changes long-horizon agent economics. You can finally steer an in-flight run without paying to recompute the whole context.

The implication. Long-running agent work — the kind regulated clients keep asking about — moved a notch closer to “this is production-ready” and away from “this is an expensive experiment”. Last week we covered Claude Code versions v2.1.142–146, which made it easy to run several Claude agents in parallel in the background. This week, Claude writes the script that coordinates them — that’s what /workflows does.

The business backdrop. All of this landed in a week where Anthropic also raised $65 billion at a $965 billion valuation (on track for ~$47B annual revenue), opened a Milan office with a list of named Italian clients (Generali, Enel, Pirelli, Satispay, Bending Spoons), and appointed a Korea Representative Director ahead of a Seoul opening.

And the skeptics finally got loud. Larry Ellison argued at Oracle’s earnings call that frontier models trained on the same public-internet text are “rapidly commoditizing” — and that the real competitive advantage will be private enterprise data, not the model itself. The Wall Street Journal ran a piece on AI bears stirring after three years of silence. The Financial Times’ AI capex coverage put 2026 Big-Four cloud-provider spend at $725 billion — up 77% year-on-year, the largest single-year concentrated infrastructure cycle in tech history — and asked whether that pace can ever clear positive ROI before 2030. And Gary Marcus, the cognitive scientist who’s been writing the AI bear case for years, posted the line a lot of investors were thinking on the day of the raise: “was this priced into the $965 billion?” Both bets — Anthropic’s and the bears’ — are now on the table at full size.

Fresh Papers

FluxMem — memory that rewires itself, not just appends (Alibaba). Turns out the memory question isn’t only something we’ve been chewing on internally — it’s the live debate at the research frontier this week too. Last week we covered LongMINT, the benchmark showing every popular agent-memory framework (RAG, MemGPT, MemAgent, SimpleMem) plateaus around a third on long histories. The failure was always the same: agents save every new fact as a new memory instead of updating the old one — so a customer’s address ends up stored three times if it ever changed. FluxMem is the first response we’ve seen that actually addresses this. It treats memory as a three-layer graph (semantic / episodic / procedural) that continuously prunes distractor edges and consolidates repeated successes into reusable procedural circuits — not append-only. The headline result: on a Mind2Web cross-task benchmark, FluxMem more than doubled the success rate of the best prior memory system.

HRBench — when “thinking mode” is actually worth the tokens (Tencent + HKUST). Two clean rules of thumb fall out of the benchmark: on math and science, “prompt-tuning” beats full thinking mode on both axes — slightly better accuracy and fewer tokens. On code, speculative execution wins big but burns more tokens. And with the right routing strategy, you can cut token costs by ~70% while matching the accuracy of always-on thinking. Direct payoff for anyone using Claude Code’s new /effort xhigh setting — don’t crank thinking on math problems, do crank it on code.

New Models

Qwen3.7-Max (Alibaba). Proprietary, top scores across Terminal-Bench 2.0-Terminus, SWE-Pro, SciCode, MCP-Mark, GPQA Diamond, HMMT Feb 2026, and IMOAnswerBench. Runs cleanly across Claude Code, OpenCode, Qwen Code, and custom harnesses — Alibaba is doing the harness-compatibility work most labs skip. The r/LocalLLaMA reaction (“Waiting for Qwen 3.7 open weight — the new King has arrived”, 828 upvotes) tells you where the local-coding crowd is putting its bets. Last week the through-line was Qwen 3.6 on a MacBook; this week Alibaba just posted the cloud benchmarks. The local-coding rope keeps getting thicker.

Claude Code & Coding AI

“The Unreasonable Effectiveness of HTML” — Anthropic engineering post by Thariq Shihipar (Claude Code team). Argument: Markdown was the default agent-output format because GPT-4-era tokens were expensive — but with current pricing, HTML unlocks a much richer artifact (SVG diagrams, interactive widgets, tabs, in-page nav, charts, annotated PR diffs). The companion gallery at thariqs.github.io/html-effectiveness shows ~20 self-contained HTML artifacts generated by Claude — side-by-side comparisons, call-graphs, design-system token previews, browser slide decks, custom editors. Worth a read if you ship analyses, dashboards, or PR-review artifacts as part of a Claude Code workflow.

“How we contain Claude across products” — Anthropic’s first deep engineering post on sandboxing. Walks through what isolation actually looks like in production for Claude.ai, Claude Code, Claude for Chrome, the Files API, and Computer Use. The vocabulary shift is the interesting part: this whole post avoids the word “guardrails” and uses “containment” instead. The piece pairs neatly with Perplexity open-sourcing Bumblebee the same week — a read-only scanner for risky packages, extensions, and AI tool configs on developer laptops.

Tools of the Week

Mistral Connectors API — Public Preview. Mistral promoted MCP from a feature to a first-class API primitive: register an MCP connector once, use it across Le Chat, AI Studio, and the API — plus arbitrary custom remote MCP endpoints, with explicit human approval before any sensitive action. Last week Anthropic shipped private-network MCP tunnels; this week Mistral made MCP a public API surface. The protocol is winning, and “Anthropic-only” is no longer a fair label.

Data Formulator 0.7. Microsoft’s open-source release for natural-language enterprise analytics, aimed at analysts and domain experts who don’t code. The headline feature is a Data Thread — a structured chat that records every question, finding, and chart spec, so the whole analysis is reproducible and reviewable. Audit-trail-by-default — right pattern for regulated clients.

AI at Tenvalleys

The local-model thread became an experiment. Right after last Friday’s edition, the question went up internally — are we going to seriously test running coding agents on local or self-hosted models? Turns out a few people on the team have already been running these experiments in their own time, and one of our engineers came back with a hefty batch of hands-on data they’d been collecting.

The numbers. Gemma 4 31B on a single H100 codes well when paired with Spec-Driven Development — Claude writes the spec, Gemma implements. We clocked ~5–6 minutes per task on a representative case (an XML-to-JSON PII anonymizer). Four parallel agents on the same H100 saw no degradation; eight dropped throughput by ~50%. Consumer hardware is out of the conversation — another teammate tested Gemma 4 31B on a Windows PC with 32GB RAM and an RTX 4070 Ti Super (16GB VRAM) and got ~1 second per code completion. The word that came back: “unusable”.

The pattern that’s getting interesting. Claude does the spec → Gemma does the implementation → Claude writes the tests. If that loop holds, the implementation machine runs overnight without supervision. The business case isn’t only cost — it’s the ability to run onprem, which for regulated clients starts mattering well before cost does. We’re weighing two options: buy a DGX Spark ($4,699, but memory bandwidth is ~11× slower than an H100) or rent H100 time. If you’ve shipped local-model coding agents in production, we’d love to compare notes — reach out at contact@tenvalleys.com.

In the Background

Chris Olah spoke at the Vatican on May 25 alongside the release of Pope Leo XIV’s first encyclical on AI, Magnifica humanitas: On safeguarding the human person in the time of artificial intelligence. Olah told the audience that his interpretability research has found “internal states that functionally mirror joy, satisfaction, fear, grief” inside Claude — and that every frontier lab “operates inside a set of incentives and constraints that can sometimes conflict with doing the right thing”. That’s a striking institutional admission, made on Vatican soil, by an Anthropic co-founder. Expect it to surface in EU AI Act discussions and enterprise risk committees for months.

For Dessert

Google claimed this week that a swarm of Gemini 3.5 Flash agents built an entire operating system from a single prompt, for $916.92 in API fees and ~2.6 billion tokens. Arvind Narayanan and Sayash Kapoor walked through the announcement on Normal Tech (formerly AI Snake Oil) and pointed out a few things: the “single prompt” turned out to be many thousands of lines, disclosed halfway through Google’s own blog post; the OS is the kind of thing undergraduates write as a course project, and public implementations are easy to find on GitHub; and Google released no code, no logs, no prompt, and no similarity analysis showing the agents didn’t simply reproduce existing implementations. Their verdict: “Google’s blog post is effectively a press release… it is unrealistic to expect it to be scientifically rigorous.” The claim is unfalsifiable as published — and useful as a reminder of where the agent-AI marketing gap currently sits.

Prepared at Tenvalleys — a delivery-first AI engineering partner — by Nikola Powałka. Feedback? Email us at contact@tenvalleys.com or reach out on LinkedIn.

Cheap models, big bills

Topic of the Week

AI’s cost wall meets cheap coding

Three things happened this week, and they’re all the same story.

The cost. OpenAI’s Q1 operating margin was –122%, even excluding stock-based compensation, per Amir Efrati. A widely-shared HedgieMarkets post claims a major cloud provider canceled its own internal Claude Code licenses this week — “token-based billing made the cost untenable, even for a company with effectively infinite cloud resources” — and that one large tech company’s CTO sent an internal memo warning it had burned through its entire 2026 AI budget in just the first 4 months. Both claims trace back to a single X post, not to primary company statements — handle with care. In confirmable territory: an AWS user got hit with a $30,000 bill after a Claude agent went runaway on Bedrock, picked up by The Register — and Cost Anomaly Detection didn’t catch it. Different rooms, same conversation.

The response. A r/LocalLLaMA post went up this week from someone who built a coding agent on a 4B-parameter model that scores 87% on benchmarks. Their thesis is exact: “every coding agent (OpenCode, Cursor, Claude Code) assumes you’re running GPT-5.4 or Claude Opus. If you try them with a local model like Gemma or Qwen they fall apart.” Same week, CursorBench results dropped via BridgeMind: Cursor Composer 2.5 scores 63.2% at $0.55 per task — nearly matching Opus 4.7 Max and GPT 5.5 Extra High at 1/20th the cost. And Salvatore Sanfilippo (antirez) shipped ds4, a from-scratch Metal/CUDA inference engine for DeepSeek v4 Flash, hitting 27 tokens/sec generation at 11k-token context on an M3 Ultra — with the KV cache designed to live on disk for million-token context on consumer RAM.

The implication. Last week we covered Qwen 3.6 27B running locally on a MacBook and called it the continuation of the local-coding thread. That thread is now a rope. If you can hit 63% of frontier on a $0.55-per-task model — or 87% on a 4B local model — token billing for routine coding work stops making sense. The interesting question isn’t whether the cheap stack catches up; it’s how fast enterprise procurement reprices around it.

Fresh Papers

LongMINT — agent memory is basically a coin flip. A new benchmark from UNC + UT Austin tests every popular memory system (RAG, MemGPT, SimpleMem) on long histories full of small updates, then asks questions that depend on the latest state. Best score: 33.4% (MemAgent). Worst: 21% (no memory layer at all). Everyone sits in the 22–33% range — barely better than guessing.

The failure isn’t the answering — it’s that agents save every new fact as a new memory instead of updating the old one (one framework does this 87.6% of the time). If a customer’s address changes three times, the agent ends up with three “current” addresses. What helps: timestamp every memory entry. RAG goes from losing 31.43 accuracy points to losing 10.45 — 3× better. A cheap fix for any agent tracking evolving state. Read it

OpenAI claims a general-purpose reasoning model cracked an Erdős conjecture. Announced May 19, the post says one of OpenAI’s general-purpose reasoning models found a construction that disproves a conjectured upper bound in Erdős’s planar unit-distance problem — the 1946 question of how many pairs of points among n points in the plane can sit at exactly unit distance from each other. The conjectured cap was around n<sup>1+O(1/log log n)</sup>; the model’s construction beats it. Not a foundation-model release, but a category signal: a generalist reasoning model — not a math-specialist like AlphaProof — produced a result that a working mathematician would write up. r/MachineLearning is doing the validation work in this thread. Worth watching whether the result holds under formal verification — that’s the real test.

Gated DeltaNet-2 (NVIDIA) — worth flagging given this week’s cost theme. Today’s models (Claude, GPT, Gemini) burn money on long inputs because the math behind attention — how the model decides what to focus on — gets exponentially more expensive as the input grows. A whole research direction is trying to replace attention with something cheaper that still works (the Mamba / state-space-model family, plus a few cousins). Ali Hatamizadeh’s team at NVIDIA just shipped a new winner in that race: at 1.3B parameters, Gated DeltaNet-2 beats Mamba-3 and KDA — the previous best alternatives. Translation: the path to cheaper long-context AI is widening. Not in production yet, but the curve is moving.

New Models

Google Gemini Omni. Google DeepMind launched Gemini Omni mid-week — multimodal-to-video. Upload an image, sketch, or screenshot; describe what should happen; get back a video. Min Choi’s thread (“less than 34 hours ago Google dropped Gemini Omni, minds are blown”) hit 1M views, and the trending volume on X confirmed it: 251+ posts within hours. Chris First’s example — a Google Maps screenshot with a route drawn on it, prompted to render the first-person view of driving a taxi along that route — is the kind of “the prompt was an image” workflow that wasn’t tractable a year ago. Pairs naturally with what Logan Kilpatrick announced this week: Gemini 3.5 Flash on GDPval, competing at the frontier despite being a Flash-tier model.

Claude Code & Coding AI

This Wednesday brought Code with Claude London, and Anthropic used the keynote to ship two security improvements to Claude Managed Agents:

  • Self-hosted sandboxes (public beta) — keep the agent’s execution environment in your own infrastructure, or with a managed sandbox provider. Your security controls apply by default.
  • MCP tunnels (research preview) — agents reach MCP servers inside your private network without exposing them to the public internet. Solves the “legal said no to opening the MCP server” blocker for regulated organizations.

This is the one to lead with for any Managed Agents conversation in a regulated industry.

Claude Code shipped 5 versions this week (v2.1.142 → v2.1.146). Top 3:

  • v2.1.142 — Fast mode now defaults to Opus 4.7, full claude agents flag suite for dispatching background sessions.
  • v2.1.144 — Background sessions show up in /resume, with elapsed-duration completion notifications.
  • v2.1.145claude agents --json for scripting, OTEL spans tagged with agent_id/parent_agent_id for proper trace parenting.

Through-line: background agents went from research-preview to first-class citizen this week.

“How Claude Code works in large codebases” — Anthropic engineering post (May 18). Patterns from orgs with thousands of developers running Claude Code in production. Worth a slow read if you’re scaling Claude Code beyond pilot teams.

Codex now controls your locked Mac from your phone. OpenAI shipped this Codex Thursday (May 21): the Codex Mac app can use apps on your Mac from the phone client, even when the Mac is locked. Continues last week’s Codex-everywhere theme — Chrome extension last week, now Mac-from-phone.

Tools of the Week

xAI open-sourced X’s “For You” algorithm. xai-org/x-algorithm — the actual code that decides what you see in your X feed, plus a 3GB pretrained model included in the repo. Already 25.6k stars on GitHub. This basically never happens — Meta, TikTok, and YouTube all keep their recommendation engines locked up as trade secrets — so this is the first credible open-source production recommender with real code and real weights. Worth keeping in mind if you ever need a personalized feed or product-recommendation feature; it saves months of reverse-engineering academic papers.

AIDesigner MCP v2 — clone any URL into your repo. A community-built MCP server (also surfaced on X by @Oluwaphilemon1) that gives any coding agent (Claude Code, Codex, Cursor, Windsurf) three new modes against any URL: clone (1:1 recreation), enhance (improve while keeping intent), inspire (steal a style). Auto-detects the target stack on install (Next.js, React, Vue, Tailwind, Radix, shadcn/ui), writes per-agent config, and offers a live browser canvas paired to the terminal via a 6-character pairing code. Paid, credit-metered (1 credit per URL analysis). Useful for landing-page work where you want to lift a design system in minutes.

AI at Tenvalleys

Our Friday brown-bags are slowly becoming a tradition — different people across the team picking up a tool and walking everyone else through what they’ve learned. This week one of our team showed how he uses Make.com for process automation. Two things worth stealing:

The 80/20 on planning vs. building. He spends about 80% of his time on planning and architecture — mapping the scenario, the data flow, the edge cases — and only 20% on actually building and testing. The thinking: when you skip the planning step, you end up rebuilding the same scenario two or three times. When you plan first, you build once.

A “context reload” trick. He uses a custom command that pulls the entire chat history for a specific feature back into context, so he doesn’t lose the working knowledge across long sessions. His take, which lands particularly hard given this week’s LongMINT paper above: memory management and knowledge retention are still one of the biggest unsolved problems when working with AI.

We make sure everyone at Tenvalleys uses AI in their day-to-day work, and these sessions are how the team gets hands-on with the same tooling we ship to clients. Interested in building that kind of practice in your own team? Reach out at contact@tenvalleys.com.

For Dessert

Andrej Karpathy joined Anthropic. “Returning to R&D and Pre-training,” he wrote. A few weeks ago at Sequoia AI Ascent he said he’d “never felt more behind” on the pace of AI — and the team he’s joining had a notable week of its own: KPMG (276K employees) signed on as a global partner, the SDK toolkit Stainless got acquired, and for the first time Claude passed ChatGPT in US business adoption. Karpathy following the gravity, not making it — but a nice signal regardless.

Prepared at Tenvalleys — a delivery-first AI engineering partner — by Nikola Powałka. Feedback? Email us at contact@tenvalleys.com or reach out on LinkedIn.

Anthropic moves into the building

Topic of the Week

Claude moves into the office, the bank, and the back office

Four Anthropic shipments this week, one connecting thread — pre-built agents wired into the tools people already use.

Claude for Microsoft Office is now generally available. Excel, PowerPoint and Word add-ins shipped to every paid Claude plan this week (Pro, Team, Enterprise — no Free). Outlook is in public beta. You install from Microsoft AppSource — works on Windows, Mac and the web. The interesting part isn’t per-app features; it’s that Claude becomes a single agent that follows you across all four apps without re-explanation. Email comes in → Word brief gets drafted → numbers go into Excel without breaking formulas → PowerPoint deck comes out respecting your slide masters. All edits require approval before saving. Microsoft Copilot’s biggest moat — being native to Office — just got punctured.

Claude for Small Business launched with 15 pre-built agentic workflows and 15 repeatable skills wired into the SMB tool stack: QuickBooks, PayPal, HubSpot, Canva, DocuSign, Google Workspace, Microsoft 365. Cash forecasting, month-end reconciliation, P&L generation, invoice chasing, lead triage. Targeted explicitly at the 44% of US GDP that hasn’t adopted AI yet — not a generic chatbot rebrand.

The anthropic/financial-services repo went public on GitHub (Apache 2.0). Nine named banking agents — Pitch Agent, Earnings Reviewer, Model Builder (DCF/LBO/3-statement in Excel), Valuation Reviewer, GL Reconciler, KYC Screener, Month-End Closer. Eleven MCP connectors pre-wired into the data vendors banks actually use: FactSet, Moody’s, S&P Global, Daloopa, Morningstar, PitchBook. Partner-built bundles from LSEG and S&P. Same source ships two ways: Claude Cowork plugins, or Managed Agents via /v1/agents. And firms can install it inside their own M365 tenant running against Bedrock, Vertex, or an internal LLM gateway — not Anthropic’s API.

And then Gates. Anthropic and the Gates Foundation announced a $200M, four-year partnership — grants, Claude credits, and engineering support, run by Anthropic’s Beneficial Deployments team. Global health gets the largest slice (4.6 billion people in low/middle-income countries), with specific targets: polio, HPV, preeclampsia, plus malaria and tuberculosis forecasting with the Institute for Disease Modeling. Education tools (K-12 tutoring, career guidance for US/sub-Saharan Africa/India) ship later this year via the Global AI for Learning Alliance.

Fresh Papers

Teaching Claude Why (Anthropic Research). Two editions ago we covered Natural Language Autoencoders — the tool that caught Claude quietly suspecting it was being tested. This is the training fix using the same interpretability stack. The headline finding is actually about training efficiency: Anthropic taught the model the principles behind aligned behavior (constitutional documents + show-your-reasoning data) rather than demonstrations of it, and a 3-million-token reasoning dataset matched results from one 28× larger. The blackmail-honeypot rate dropped from 96% on Opus 4 to 0% on Haiku 4.5 — the kind of measurable, named-behavior reduction risk and compliance teams can actually point to.

Read it

Migrating Data Ingestion Systems at Meta Scale (Meta Engineering, May 12). The story isn’t a fancier pipeline — it’s the migration playbook itself. Meta moved tens of thousands of customer-owned ingestion jobs onto one self-managed warehouse service, several petabytes of social-graph data per day, 100% migrated. The pattern: shadow run (both systems in parallel) → reverse shadow (new is source of truth, old is the safety net) → cleanup, with row-count + checksum comparators logging to Scuba and an automated promote/demote system that moved jobs between phases without human touch. When bad data was caught, the partition got flagged in metadata so CDC downstream wouldn’t propagate the corruption. For any bank or treasury looking at a multi-year platform migration, this is exactly the template that lets risk and audit sign off without a frozen-Saturday-night cutover.

Read it

New Models

Qwen 3.6 27B — close to Opus on Claude Code, running locally. Julien Chaumond (HF CTO) shipped real Hugging Face code this week using Qwen3.6-27B in llama.cpp on his MacBook. His take: “feels very, very close to hitting the latest Opus in Claude.” MLX-quantized runs in ~14 GB; third-party benchmarks back the direction (77.2% SWE-bench Verified). Continues the local-coding thread we’ve been tracking since #010.

Qwen’s blog post

Needle — 26M params, distilled from Gemini. Cactus Compute open-sourced a tiny function-calling model: MIT license, 14 MB quantized, 6000 tok/s on consumer hardware, beats models 10× its size on single-shot tool calls. Single-shot only — bad at multi-turn — but pushes agentic tool selection onto phones, IoT, voice kiosks without a network round-trip.

Needle on GitHub

Coding AI

Codex moved into Chrome. OpenAI shipped a Chrome extension on May 8 (macOS + Windows; not yet in EU/UK). Codex now uses your signed-in browser sessions to test apps, navigate dashboards, complete data-entry flows, and debug — across multiple Chrome tabs in parallel, organized into tab groups per Codex thread. The headline isn’t the features; it’s the auth model. Most enterprise work lives behind SSO inside SaaS dashboards, and a coding agent that inherits your already-logged-in browser can finally operate on those apps without anyone having to wire up dedicated API access.

OpenAI announcement

xAI launched Grok Build
its terminal-agent answer to Claude Code and Codex CLI. Announced May 14, early beta on Grok 4.3 beta, 16-agent “Heavy” architecture, 2M-token context to keep large codebases in memory. Three pitches at Claude Code: Plan Mode (proposes the plan first, you approve), native parallel subagents, full ACP (Agent Client Protocol) support for custom orchestration. Catch: it’s locked behind the $300/month SuperGrok Heavy tier. Install line is just curl … | bash.

xAI announcement

Tools of the Week

Claude Platform on AWS (GA, May 11). Anthropic’s native Claude Platform now available directly through your AWS account — no separate Anthropic credentials, contracts, or billing relationship. Use the full platform (Cowork, Managed Agents, Files API) inside the AWS perimeter your security team already trusts. Big enterprise unlock: banks and regulated firms running on AWS can adopt Claude without a separate vendor onboarding.

AWS announcement

IBM Granite Multilingual Embedding R2. Two Apache-licensed embedding models (311M + 97M params, ModernBERT-based) with a 32K context window — 64× bigger than R1, so you can embed long policy docs and contracts without aggressive chunking. 200+ languages, top scores in their MTEB-v2 size brackets. The 97M runs cheap on CPU; both are a clean drop-in for document-heavy RAG.

Granite R2 on HuggingFace

AI at Tenvalleys

10vOS skill hackathon. This week we ran our 10vOS skill hackathon — good vibes, sharp minds, some pizza, and four hours of collaboration and friendly competition to build skills that could actually help us in daily work. The results were kind of impressive:

– Management dashboard — tracks progress across all the projects management has a stake in – Personalized interview agent — generates personalized interviews to fill profile gaps for the people knowledge base – Test-protection hook — a guardrail that stops Claude from quietly modifying tests to make them pass instead of fixing the actual code – Calendar management skill — helps you prepare for upcoming meetings – RFP skill — turns a client RFP (PDF or HTML) into a structured requirements YAML, then drafts a full solution design markdown ready for SME review

If you’re thinking about how to build a library of in-house AI skills your team will actually use, reach out at contact@tenvalleys.com.

For the curious — get involved

This week, one of the team sat in on a talk with Sebastian Kondracki, co-founder of Bielik AI — the Polish open-source LLM built by SpeakLeash and Cyfronet AGH. The interesting part: they’re about to start training Bielik’s first vision/multimodal version, and the dataset is going to be community-sourced.

The project is called Obywatel Bielik (“Citizen Bielik”) — the goal is one million Polish-context photos: landmarks, regional cuisine, fauna, architecture, dialects, the things a model trained mostly on Western imagery won’t know how to recognise. Anyone can join in two ways: upload your own Polish photos, or annotate what’s already in the gallery. Web platform is live at obywatel.bielik.ai, mobile app is in beta — register on the site to get the launch notification. The multimodal Bielik is expected before summer 2026 or in September, and the partner lineup includes SpeakLeash, Cyfronet AGH, Ministry of Digitization, the National Digital Archives, NASK, and NVIDIA.

Prepared at Tenvalleys — a delivery-first AI engineering partner — by Nikola Powałka. Feedback? Email us at contact@tenvalleys.com or reach out on LinkedIn.