technologyScore 45/100Research

Dario Amodei: Anthropic has documented 'bad behaviors' in models and is racing to address them using mechanistic interpretability before systems become uncontrollable

Dario Amodei· Anthropic· AI· 2026-07-22· about Anthropic
We've increasingly documented the bad behaviors of the models when they emerge and are now working on trying to address them with mechanistic interpretability... If we build them poorly, if we're all racing and we go so fast that there's no guardrails, then I think there is risk of something going wrong.

Why it matters

An AI lab founder acknowledges observing undesirable model behaviors that require active intervention, signaling that safety/control remains an unresolved technical problem even as systems scale toward AGI-level capabilities.

Investment implication

Validates need for AI safety tools, interpretability research platforms, and monitoring/control infrastructure. Creates capex demand at AI labs for safety teams and tooling. May slow deployment of certain capabilities pending safety resolution.

Source

Dario Amodei & Demis Hassabis: We're 12 Months Away From AI Replacing Everyone (YouTube)
← All signals