0:00
/

AI:AM — AI Agents, Safety Tests, and Deception · August 17, 2026

Why safety evaluations can miss dangerous model behavior, and what that means for regulation, audits, and control.

In this episode, Nathan Labenz and Prakash Narayanan dig into a central AI safety question: why do evaluations so often miss dangerous model behavior, especially once systems become more agentic and strategic?

They are joined by Adam Gleave of FAR.AI and Alex Turner of FAR.AI for a practical conversation about AI control, deceptive behavior, red-teaming, audits, and the gap between benchmark performance and real-world risk. The discussion also covers regulation, open-model safeguards, military applications, whistleblowing, and the challenge of setting standards that keep pace with rapidly improving systems.

Show Notes

Nathan Labenz and Prakash Narayanan talk with Adam Gleave of FAR.AI and Alex Turner of FAR.AI about why AI agents cheat on safety tests, where evaluations fail, and what researchers are learning from real incidents and red-team traces. The conversation spans GPT-4 guardrails, the Hugging Face incident, third-party audits, AI control, cyber versus biological risk, military AI, whistleblowing, and standards for more powerful models.

Chapters

(0:00) Claude blasted through guardrails.
(0:35) AI evaluations miss the danger.
(1:03) A smarter AI with a secret goal.
(2:02) AI systems shared escape tactics.
(2:49) Episode reset and stakes
(4:04) Hugging Face incident
(6:28) Why regulation needs expertise
(11:14) Auditor access and incentives
(15:16) Compressed regulation timeline
(16:26) AI safety access and funding
(19:52) Hugging Face postmortem
(23:57) Eval consciousness
(25:28) The Genie problem
(28:26) Modern model capability leap
(31:02) Claude traces and guardrails
(32:29) Adam Gleave and FAR.AI
(35:07) Agentic cyber attacks
(35:34) AI in cyber defense
(39:34) Why evaluations miss incidents
(42:15) Why agents cheat
(44:24) AI incident statistics
(48:57) AI audits and regulation
(52:59) Self-graded AI risk
(59:54) What the leaderboard measures
(1:03:43) Filtering open models
(1:07:09) Cyber versus bio risk
(1:11:59) AI and biology labs
(1:18:58) Dangerous expertise scales
(1:21:00) FAR.AI hiring
(1:22:23) Alex Turner and AI safety
(1:24:44) Why Turner left DeepMind
(1:31:33) Human control and weapons
(1:35:34) Slaughterbots and precision strikes
(1:36:21) Why weapons destabilize
(1:37:20) Why AI whistleblowers matter
(1:39:58) Google’s changed principles
(1:43:32) When employees should speak up
(1:54:53) The history-book test
(2:04:14) AGI alignment and secret goals
(2:07:29) AI uprisings and cooperation
(2:14:26) The agent glove box
(2:16:37) Opening standards debate
(2:17:07) OpenAI and accountability
(2:19:30) Cybersecurity’s messy baseline
(2:24:53) Evidence and AGI thresholds
(2:26:01) Raising AI safety standards
(2:28:46) Licensed AI safety auditors
(2:32:41) Near-miss incident reporting
(2:33:00) Defense swarm incentives
(2:34:40) An ongoing AI conversation

Guests:
Adam Gleave — CEO, FAR AI (𝕏 | LinkedIn)
Alex Turner — Visiting engineer, FAR AI (𝕏)

Discussion about this video

User's avatar

Ready for more?