Why do AI safety evaluations miss dangerous behavior once models become more agentic and strategic? This episode connects concrete failures to the institutional machinery needed to catch them: red-teaming, independent audits, incident reporting, whistleblower protections, and containment systems.
Adam Gleave and Alex Turner of FAR.AI join Nathan Labenz and Prakash Narayanan for a practical discussion of deceptive agents, cyber and biological risk, autonomous weapons, alignment, and how to keep increasingly capable systems under meaningful human control.