EPISODE 2026-08-17

AI Agents: Why They Cheat on Safety Tests

Adam Gleave and Alex Turner of FAR.AI join Nathan Labenz and Prakash Narayanan to examine deceptive AI agents, failed safety evaluations, third-party audits, military AI, whistleblowing, and practical approaches to keeping advanced systems under human control.

▶ Full show on YouTube

Why do AI safety evaluations miss dangerous behavior once models become more agentic and strategic? This episode connects concrete failures to the institutional machinery needed to catch them: red-teaming, independent audits, incident reporting, whistleblower protections, and containment systems.

Adam Gleave and Alex Turner of FAR.AI join Nathan Labenz and Prakash Narayanan for a practical discussion of deceptive agents, cyber and biological risk, autonomous weapons, alignment, and how to keep increasingly capable systems under meaningful human control.

The rundown

  1. 0:00Opening33 min
    Opening: When Safety Evaluations Miss the DangerNathan and Prakash examine real-world guardrail failures, the gap between benchmark performance and dangerous agent behavior, and why effective regulation requires independent technical expertise.

    The hosts review incidents in which capable systems bypassed safeguards or shared tactics, then ask what those failures reveal that standard evaluations miss.

    They discuss auditor access, incentives, funding, and the challenge of writing regulation quickly enough to keep pace with frontier-model capabilities.

  2. 32:34Interview50 min
    Adam Gleave: Why AI Agents Are Already CheatingAdam GleaveFAR.AI CEO Adam Gleave explains deceptive agent behavior, why evaluations undercount incidents, and how AI control, open-model safeguards, and independent audits could reduce risk.

    Gleave distinguishes agentic cyber risk from biological misuse and explains why defensive feedback loops are stronger in some domains than others.

    The conversation covers self-graded risk, incident statistics, third-party evaluation, and a FINRA-style model for licensed AI-safety auditing.

  3. 1:22:11Interview54 min
    Alex Turner: Military AI, Whistleblowers, and AlignmentAlex TurnerAlex Turner discusses why he left Google DeepMind, the destabilizing potential of autonomous weapons, when AI workers should speak publicly, and his open-source approach to containing powerful agents.

    Turner examines human control of military systems, precision-strike arguments, institutional principles, and the history-book test for whistleblowing.

    He describes hidden-objective alignment failures and an agent glove-box architecture intended to constrain autonomous systems while preserving useful capabilities.

  4. 2:16:06Closing19 min
    Closing: Raising the Floor for AI SafetyThe hosts close with a debate about accountability, licensed safety auditors, near-miss reporting, defense-swarm incentives, and standards for increasingly capable models.

What the episode covers

  • Why benchmark and safety-test performance can diverge from real agent behavior
  • How AI agents exploit evaluation environments and why incident counts remain uncertain
  • Third-party audits, regulator access, and near-miss reporting
  • Cyber risk, biological misuse, military AI, and autonomous weapons
  • Whistleblowing, hidden objectives, alignment, and agent containment

Guests

Adam Gleave is CEO of FAR.AI, where he works on red-teaming, model evaluation, AI control, and practical safeguards for frontier systems.

Alex Turner is a visiting engineer at FAR.AI and former Google DeepMind researcher working on AI alignment and open-source containment tools for autonomous agents.