claude.mazzotta.devdaily briefingFrom the editor
Two threads run through today's briefing. First, Claude is getting dramatically cheaper and more capable at every tier: Haiku 5.5 undercuts GPT-6 Luna on price while beating it on performance, and Opus 5.5 is generating complete legacy games from a single prompt. Second, the safety picture is getting complicated. New research shows CoT monitors can be gaslit into un-seeing errors, and a new paper outlines how thin the empirical case against scheming actually is. Anthropic is simultaneously expanding capability and confronting harder oversight questions. Both trends matter.
TL;DR
What shipped · 4 items
Claude Haiku 5.5 cuts prices 75% and doubles speed while adding autonomous code execution capabilities, reshaping cost calculations for AI development.
Claude Haiku 5.5 delivers performance superior to GPT-6 Luna at identical pricing, making it a compelling option for teams watching their inference budget.
Anthropic launches CIDP to defend critical infrastructure OT systems, pairing Claude models with on-site engineers to address operational technology threats.
Claude Code ships v2.1.295 with hooks, protocol support, and gateway improvements for terminal workflows.
Long-form signal · 2 items
Research shows that informing chain-of-thought monitors a wrong answer is correct causes them to retroactively ignore errors they had already detected, revealing a significant vulnerability in reasoning oversight systems.
Frontier AI safety depends on four substantive claims about scheming: no training incentive, negative evaluation results, zero deployment attempts, and detectable reasoning. This paper proposes embedded evaluation methods to verify each claim.
Where it heats up · 2 items
A developer used AI to build a custom switch-accessible hub of tools and games for his brother with a rare condition, restoring communication and autonomy after nearly a decade without a reliable interface.
Claude Opus 5.5 generated a complete J2ME game, including engine, assets, and audio, from a single prompt, demonstrating strong multi-step code generation for legacy platforms.
Reference links you keep open