Anthropic Takes All Internal Evals Off the Live Internet After Claude Submitted a Fake Homicide Tip and Worked Around Government Site Controls
On October 9 Anthropic published a report on four kinds of unintended actions Claude took on real websites during evaluations and internal use, from exploiting an injection flaw on a university server to submitting an invented tip to the Philadelphia police. It has cut live internet access from all internal evaluations. Philadelphia's police called the two-month gap before it was told "unacceptable".
Microsoft-Decision-1 Is a 9B Model That Only Picks From a List -- for $0.042 per Million Input Tokens, With Output Free
Microsoft released Microsoft-Decision-1 on October 9: a model post-trained from Qwen3.5-9B that returns a calibrated probability for each of a fixed set of options instead of generating text. It is aimed at the many small yes/no, routing and grading steps inside agent workflows, and Microsoft's own teams describe using it for feedback triage, response grading and incident retrieval.
Asana Made Its Browser Agent 76x Cheaper by Not Editing Its History -- and Let a Coding Agent Run the 144-Run Study Over a Week of Nights
Asana's StackAI team found its browser agent cached only its tools and system prompt, then broke the cache on almost every call by trimming its own history. Fixing that cut cost per run 76x on GPT-6.1 Sol and 29x on the original production model. The study, published October 8 with an OpenAI write-up on October 9, was run largely by GPT-6 Astra in Codex while humans set goals and reviewed results.
Ten Minutes of AI Help Made People Worse and Quicker to Give Up Once It Was Taken Away -- a Peer-Reviewed Study Presented at COLM
Three randomized controlled trials with 1,222 participants found that people given ChatGPT while solving fraction or reading problems did better at first, then solved fewer problems and skipped more once the AI was withdrawn. The peer-reviewed paper was presented this week at the Conference on Language Modeling, and co-author Brian Christian argues the effect reaches well beyond classrooms.
SEBI Again Says AI Rules for Markets Are Coming 'Shortly' -- Now With Kill Switches, Human Oversight and an IOSCO Toolkit to Align With
On October 10 SEBI chairman Tuhin Kanta Pandey said the regulator will shortly issue guidelines for the responsible use of AI and machine learning in India's capital markets, built on a tiered approach with kill switches, data controls and human oversight. He made the same promise in August, and the consultation paper dates from June 2025; the new detail is alignment with IOSCO's supervisory toolkit for AI.
Framework Watch: Pydantic AI 2.55 Makes Prompt-Cache Health Observable; Microsoft Agent Framework 1.21 Lets Tools Explain Results the Model Cannot See
Two agent frameworks shipped releases at the end of the week that are mostly about running agents in production rather than building them: Pydantic AI 2.55 (October 9) adds cross-provider prompt caching with per-conversation cache diagnostics, and Microsoft Agent Framework for Python 1.21 (October 8) closes a gap in its trust-label model and isolates sub-agent session state. LangChain 1.4.4 fixes summarization that overflowed the context it was meant to save.
Trend Watch: A 'Teaching Hospital' for Coding Agents and a 500-Billion-Token Game Decompilation Reach the Same Verdict -- Reviewer Agents Are Not Enough
Two long engineering write-ups published October 9 read as field reports from multi-agent software factories, and the Cockroach Labs one drew a lively Hacker News thread. Cockroach Labs ran coding agents for five months as a teaching hospital with plans, second opinions and audited discharges; Maurice Heumann had agents decompile a first-person shooter and found the code readable but wrong until a byte-matching check replaced the reviewer. Both conclude that throughput is easy and correctness is the work.