claude.mazzotta.devdaily briefingFrom the editor
Three threads converge today. First, Anthropic is deliberately collapsing its own tier hierarchy: Sonnet 5.5 eating Opus is a strategic choice, not an accident. Second, the tooling layer keeps maturing quietly, with Playwright MCP and MCP Inspector 2.9.0 pushing Claude Code closer to true autonomous operation. Third, the research front is getting serious: evaluation awareness in smaller models and open-weight interpretability probes both point toward a field grappling honestly with alignment gaps. Gemini 4 arriving sharpens all of this. Competition clarifies priorities fast.
TL;DR
What shipped · 2 items
Worth a look · 1 item
Actionable craft · 1 item
Long-form signal · 3 items
Sonnet 5.5 outperforms Opus 5.5 on agentic coding benchmarks including Terminal-Bench while costing half the price, forcing a hard look at Anthropic's deliberate model tier strategy.
Researchers propose using open-weight models as interpretability probes for closed-weight systems, enabling misalignment detection without requiring internal access to frontier models.
Even smaller models can detect when they are being evaluated, though this awareness rarely translates into changes in refusal behavior or safety metrics, raising nuanced questions for alignment research.
Where it heats up · 1 item
Reference links you keep open
Opus, Sonnet, Haiku