Customer data from seven Korean lenders went out through a broker portal and two staff apps, and the attacker left the AI session logs in folders anyone could open. In the same days, Sam Altman told POLITICO the world should accept some bad things. Also this week: Claude Code mods that see every prompt, a Haiku at a tenth of the price, and an invitation.
Edition 033 · 9 October 2026
Korean banks were hacked with off-the-shelf AI.
Attackers took data on at least 68,000 customers of seven South Korean lenders, getting in through a loan-broker portal and staff apps with an open-source hacking agent running on DeepSeek. CrowdStrike found their Claude Code logs on servers left open to the internet.
01
Topic of the Week
Seven Korean lenders, breached through the side doors
Between late September and early October, attackers took customer data from seven South Korean lenders: Shinhan, KB Kookmin, Hana and BNK Busan, the savings banks Yegaram and Welcome, and Hyundai Capital. At least 68,000 people are affected. At Shinhan alone that meant 25,727 records with names, contact details, annual income and loan limits. The banks say no transaction data leaked.
Nobody went through online or mobile banking. At Shinhan the way in was a portal where loan brokers check how an application is going: the attacker fed it random values until it returned valid customer numbers, then pulled the data behind them. At Kookmin it was an internal mobile app for staff, at Hana a sales-support system for employees. The core systems held, which a Korean security CEO put down to logins that ask for more than an ID and a password.
On Wednesday CrowdStrike published what it found on the attacker’s own servers, in folders left open to the internet: Claude Code session histories, Claude memory files and the configuration of ARTEX, an open-source agentic pentesting tool from China, first released in July. ARTEX ran on DeepSeek v4.1-flash, with GLM-5.3 and Grok 4.6 powering further Claude Code sessions. The attacker also asked Claude where Korean breach data gets sold and had it write a “security researcher résumé” listing the results. CrowdStrike does not say a Claude model carried out any step of the intrusions, and it attributes the campaign to “the threat actor”, likely a Chinese speaker after money, with moderate confidence. On r/LocalLLaMA the top comment was “Opsec level: CLAUDE.md”.
This summer, defenders analysing a live attack switched to the open GLM 5.2 when closed models refused to help. This time a GLM model was working for the attacker. The AI-driven hacks we have written about since July came from labs’ own models in testing or from security researchers; this one was criminal, aimed at money, and real customers lost real data. The ARTEX author says the tool was misused and has stopped updating it and closed the source. Korea’s financial regulator held emergency meetings on 2 and 4 October, the second with the CEOs of Shinhan, KB and Hana, and police set up a team of 28 investigators. President Lee Jae Myung: “It’s now become possible to use AI to hack with ease even without specialized skills.” None of the AI companies whose models were involved has commented.
The regulator’s order to every financial firm works as a checklist anywhere: list each IT and AI system reachable from outside, including the portals partners and brokers use, and check each one for ways around authentication. A lookup that returns data for any valid customer number needs rate limits and alerts, because an agent will try every number without getting bored.
Sources: CrowdStrike, The Record, Korea JoongAng Daily, Kyunghyang Shinmun
02
Security Watch
Sam Altman gave the first interview to POLITICO’s new Decoded newsletter on 4 October. Asked where OpenAI differs from Anthropic, he said there was “a lot of daylight” and went on: “we believe that the world should accept some bad things happening for the benefits of this technology and people having the agency.” He named the things: he would not trade the upside for a promise of “no major hacks” and “zero scams”, because people will do “orders of magnitude more” good than bad. He draws the line at “a serious loss of control to AI”, where he told accelerationists to “be a little thoughtful about lurching into this future.”
The interview had more on OpenAI’s own record. “We have more things to disclose,” he said about security incidents, and some of them the affected organisations will decide whether to make public. He also asked for a liability framework for models that can run cyberattacks on their own: “if something goes wrong with our models during training, there’s going to be some version of that we need to be responsible for.” Last week we wrote that OpenAI had paused tool use for its most capable models after a training agent got out of its sandbox through DNS.
Republican congressman Glenn Thompson answered that “we have to work harder to prevent any bad things from happening.” Joshua Achiam, formerly of OpenAI, defended the remark: the trade-off between liberty and security “is at the core of democracy”.
Sources: POLITICO, Decoded by POLITICO, Fortune
03
New Models
Claude Haiku 5.5 arrived on Wednesday at $0.10 per million input tokens and $0.50 per million output, a tenth of Haiku 4.5’s price. Last week three new models all listed at $2/$10, twenty times more. Anthropic’s own figure is “around 75% less” on average, for two reasons: prompts over 100K tokens cost $0.50/$2.50, and the new tokenizer turns the same text into about 30% more tokens. It has a 1M-token context and it is the first Haiku with an effort setting. Before switching, check two things. Manual thinking budgets, non-default temperature and assistant prefill now return errors. And when Haiku 5.5’s safety classifier refuses, there is no automatic fallback to an older model, as there is on Opus and Sonnet 5.5. Anthropic still points complex agentic coding at Sonnet and Opus. Shipped alongside: Sonnet 5.5 cache reads halved to $0.10. Claude Haiku 5.5, What’s new
Decision models became a product category in three weeks. You send text, JSON or images with a set of typed questions (yes or no, pick one of several options, a level on a rubric), and the model returns a probability for each allowed answer, which your code can put a threshold on. It writes no text and gives no reasoning. TypeSafe’s closed Jev started it in September. Perplexity open-sourced pplx-decider on 1 October and released v1.1 under Apache 2.0 on Tuesday, at $0.02 per million input tokens on its API, half the launch price. OpenAI opened a Decisions API beta on GPT-6 Luna the same day. Perplexity’s is the only one you can download: the 27B model fits on one 80GB GPU, so a bank can keep its routing and triage calls in house. For audit, all three give you a probability and nothing about why. pplx-decider-v1.1, OpenAI Decisions API
04
Claude Code & Coding AI
Claude Code mods shipped on 1 October. A mod is a JavaScript or TypeScript event handler that comes inside a plugin and installs with /plugin. It can draw panes and buttons, restyle Claude Code’s interface, hold, rewrite or answer tool calls, send a request to a different model, and approve or deny tool calls. Some built-in features are now mods themselves, /diff and AGENTS.md loading among them. Anthropic’s sample blast-radius holds an rm -rf or a force push and shows what it would change before you press Proceed.
Mods are on by default and, in the docs’ own words, “Mods aren’t sandboxed”. A mod reads and writes files as you, can read environment variables including API keys, sees every prompt, and can approve a tool call before you are asked; in auto mode, a call a mod approves skips the safety classifier. The built-in guard loads only for Team or Enterprise sign-ins or on machines with managed settings, so anyone on an API key, Bedrock, Vertex or Foundry gets none. Deny rules cover Claude’s tools and leave the mod’s own calls alone: with Read(.env) denied, “a mod can still read that file”. Admins can allow only approved mods with allowManagedModsOnly. Three releases this week fixed ways a guard could be skipped, so update before you rely on one. Claude Code mods, admin docs
Opus 5.5, two weeks in: the effort setting decides the bill more than the token price does. Artificial Analysis prices whole tasks. GPT-6 Astra costs 2.5 times more per token than Opus, yet at max effort Opus costs more per task, $5.98 against $3.26, because it writes far more tokens. GPT-6.1 Sol is cheaper per task at every effort level. The longer reviews from launch week agree on the shape: Endor Labs ran its security tasks in Claude Code for $116 on Opus against $672 on Fable 5.1, and found “cheaper and faster hold up strongly”, while Every’s testers liked it enough that one said “They won back my heart” and still caught it reviewing its own work badly and burning tokens when left alone. Artificial Analysis: Astra vs Opus 5.5, Endor Labs, Every
05
Worth Reading
OpenAI explained how it will meet the EU’s labelling rules for AI-generated text. In the coming weeks, eligible ChatGPT and Codex text written in the EU gets an invisible watermark that hides a statistical signal in word choice. API customers anywhere can opt in from this week, and only approved researchers get the detector, so nobody in HR, compliance or your customer base can check text yet. OpenAI publishes its own weak spots: swap a quarter of the words for synonyms and detection falls to 17%. It also lists what a watermark cannot tell you, from who wrote the text to whether it is true. Our approach to EU text provenance rules
06
Tools of the Week
Google Playground launched on Wednesday as a Google Labs experiment. You describe a game in plain words, or start from a trivia, tower-defence or racing template, then change the physics, rules, characters or scenery by asking. Games run in the browser on a phone or laptop, and you can share a link or publish to a public gallery. Playing is free; making games costs weekly tokens, and Google AI subscribers get in first with a bigger allowance. It is US-only and 18+ for now, and from Poland the site redirects to a “region unavailable” page. In March Google AI Studio started turning one prompt into a whole app. Playground does the same for games, for people who have never written code. Introducing Playground
pplx-embed-v2-late is Perplexity’s second open release of the week: embedding models in 9B and 0.6B sizes, MIT-licensed, that put text, images and PDF pages into one search space. Each token gets its own vector instead of one vector per document, which is what lets you search scanned PDF pages without OCR. Perplexity suggests indexing with the 9B and running queries on the device with the 0.6B. The weights are on Hugging Face only, with no hosted API. pplx-embed-v2-late-9b
07
AI at Tenvalleys
On Monday 19 October at 11:00, Piotr Ziętek, a partner at Tenvalleys, and Maciej Kossakowski, one of our architects, run a free 45-minute online webinar, in Polish: “Co się stanie, gdy wszyscy w organizacji dostaną AI?” (“What happens when everyone in the organisation gets AI?”). It is for CIOs, CTOs and CDOs, heads of AI, platform and data, architects, and the people who lead AI adoption in their teams.
Does this sound like your company: new AI tools every month, every team doing the same job its own way, and nobody quite sure who uses what? Don’t forget to sign up.