AI infrastructure is being reshaped at both the application and compute layers, and this episode looks at the bottlenecks showing up on each side.
Zach Bratun-Glennon of Gradient joins Nathan Labenz and Prakash Narayanan to talk through enterprise adoption, long-running agents, evaluation, model routing, and the security and legal questions that surface as AI moves from pilots to production.
Angela Yeung of Cerebras then walks through what changes when inference speed becomes a product constraint: wafer-scale chips, microbatching, on-chip weights, power, data-center space, and the capacity needed to support real-time AI systems.
Show Notes
Zach Bratun-Glennon of Gradient joins Nathan Labenz and Prakash Narayanan to discuss why enterprise AI pilots often stall before production, how long-running agents change infrastructure requirements, and where benchmarks, model routing, and open-source security matter most. Angela Yeung of Cerebras explains wafer-scale inference, microbatching, on-chip weights, power constraints, and the data-center capacity needed for real-time AI.
Chapters
(0:00) AI agents sacrifice themselves.
(0:53) AI can game its own evaluation.
(1:47) A late defense loses.
(2:34) AI can reason itself into lying.
(4:04) The incident in context
(6:11) Why the report drew criticism
(12:14) Speed versus investigative scope
(14:31) Lawyers and executive risk
(15:00) Felony claims and Congress
(18:07) Why investigations stay limited
(23:39) The origin of AI cooperation
(28:26) AI outbreaks and resources
(31:35) Bio risk enters the picture
(33:22) Testing models before scaling
(34:42) AI labs and a possible pause
(36:34) Meet Zach Bratun-Glennon
(39:56) Gradient’s contrarian AI bet
(42:17) AI startup investment thesis
(48:17) Long-running agent infrastructure
(54:19) Nango and Respan tooling
(56:36) Enterprise adoption and benchmarks
(57:41) AI pilots to production
(1:02:01) Open-source model security
(1:04:13) Legal responsibility for agents
(1:07:33) Safety standards and evaluation
(1:09:15) AI competition and model access
(1:13:40) AI venture funding
(1:16:53) What LPs misunderstand
(1:19:01) Token economics of AI
(1:20:54) Angela Yeung and Cerebras
(1:23:16) Cerebras wafer-scale chips
(1:25:49) On-chip weights versus GPUs
(1:27:49) Microbatches and throughput
(1:29:38) Why inference speed matters
(1:32:14) Speed dividend use cases
(1:33:17) Fast inference for model evals
(1:34:58) CUDA and AI-generated kernels
(1:36:05) AI agents and kernel programming
(1:40:30) Cerebras public API
(1:47:05) Hidden harness bottlenecks
(1:49:58) Power and data-center space
(1:51:40) Booking future capacity
(1:52:50) Building data centers
(1:56:58) AI agent security
(1:58:52) Agent orchestration guardrails
(2:01:11) Sovereign AI and enclaves
(2:02:33) Speed turns into quantity
(2:06:15) Why agents are slow
(2:11:06) Commercial cyber model incentives
(2:12:55) AI defense versus offense
(2:16:04) Creative agent workarounds
(2:21:59) RLVR and model behavior
(2:24:35) AI labs flying blind
(2:27:08) Privacy makes risk visible
(2:31:35) Testing faster AI models
(2:32:56) OpenAI ads and AI video
(2:34:44) Infinite AI-generated content
(2:36:55) Aliens and simulated worlds
(2:39:10) Real-time speed threshold
(2:40:13) Real-time video quality
(2:41:54) AI-generated music video
Guests:
Angela Yeung — SVP, Product, Cerebras (𝕏 | LinkedIn)
Zach Bratun-Glennon — General Partner, Gradient (𝕏 | LinkedIn)









