Gemini 4 Argon Ships to Cyber Defenders First -- While Google Staff Question Its Real-World Coding
Google's first flagship in more than seven months posts 77.9% on DeepSWE v1.1 and raises the output limit to 1M tokens, but goes only to Fairwind security partners at launch. Bloomberg reports some employees think it is better at benchmarks than at real coding work.
Anthropic's Robot Exposure Index: Robots Can Do Three-Quarters of Physical Tasks, but Are Cost-Competitive for 0.3%
A new Anthropic study scores 7,594 O*NET physical tasks for what robots can do today. Technical exposure is wide, but cost is the bottleneck: on past price trends, the cost-competitive share takes 40 years to reach 10%.
GitHub Trending: Graphify and PageIndex Lead a Turn Toward Structured Retrieval Over Vector Search
Graphify turns a codebase into a queryable knowledge graph for coding agents, and PageIndex replaces chunk-and-embed RAG with a reasoning-navigated document tree. Both were among the top star-gainers on October 1's AI trend lists.
Framework Watch: Claude Code 2.1.286 Caps a Failing Call at 14 Requests by Default and Tightens Secret Redaction
The September 30 release puts one retry limit on each model call, falls back to the previous same-tier model when the API refuses a model, makes --bare truly minimal and refuses non-registry npm plugin sources. It also patches a run of credential leaks in logs and MCP errors.
Trend Watch: A Day After the Self-Policing Accord, Trump's FTC Is Reported to Be Preparing Subpoena-Style Demands to the Frontier Labs
Reports on September 30 said the FTC is drafting civil investigative demands to compel testimony from executives at Anthropic, OpenAI and other frontier labs over what they tell consumers about risk -- the regulatory track the accord was supposed to make unnecessary.