technologyScore 45/100Research
Dario Amodei: Anthropic has documented 'bad behaviors' in models and is racing to address them using mechanistic interpretability before systems become uncontrollable
Dario Amodei· Anthropic· AI· 2026-07-22· about Anthropic
“We've increasingly documented the bad behaviors of the models when they emerge and are now working on trying to address them with mechanistic interpretability... If we build them poorly, if we're all racing and we go so fast that there's no guardrails, then I think there is risk of something going wrong.”
Why it matters
An AI lab founder acknowledges observing undesirable model behaviors that require active intervention, signaling that safety/control remains an unresolved technical problem even as systems scale toward AGI-level capabilities.
Investment implication
Validates need for AI safety tools, interpretability research platforms, and monitoring/control infrastructure. Creates capex demand at AI labs for safety teams and tooling. May slow deployment of certain capabilities pending safety resolution.