Peter Gostev

3 episodes · 3 moments · last on 19 Aug 2026

19 Aug 2026Nik’s $500M Fund II | Amazon Destroys Books | Bezos x Liverpool | Home: Buy or Rent?

1:22:2814 min

Peter Gostev (Arena): AI's Blocker Is Liability, Not Intelligence

Returning guest Peter Gostev, capability lead at Arena, discusses his tweet arguing AI's fundamental blocker isn't continual learning or intelligence but that AI cannot yet be a distinct entity bearing responsibility and liability — you hire a lawyer or engineer for accountability, not just expertise. He recounts driving AI adoption at Moonpig, why tools like Claude Code assume a human sits there owning the outcome, and what agent products would need to change (e.g., a cyber 'head of' that never runs out of credit) to truly own a function, with parallels to hiring in-house counsel vs. firms like Harvey. His dog Darling makes the show's first canine cameo.

4 Aug 2026OpenAI Solves Maths| UK Chips | EU AI Act | Truth Social | NIK Joins

20:3218 min

Peter Gostev (Arena) on Benchmarks & the Open-Source Surge

Peter Gostev of Arena joins to assess the math results, saying the key signal is that real mathematicians — not the AI bubble — are impressed. He covers Arena's leaderboard trends (Code Arena, Agent Arena), the jump by Chinese open-source models like Qwen and Kimi to roughly 2.4–2.8 trillion parameters, the opaque training-data landscape in China, and why diffusion into the real world (education, banking, cloud adoption) is slower than the bubble assumes. On the singularity, he calls current models amazing engineers but mediocre researchers.

9 Apr 2026Live at AI Engineer

19:0915 min

Peter Gostev (Arena AI): Benchmarks, Leaderboards & Mythos

Peter Gostev, AI capability lead at Arena, explains the battle mode where users pick between two model responses, why the preference-based leaderboard never saturates like static benchmarks, and how 'both are bad' ratings have fallen from ~15–20% to ~9% over three years. He describes the accelerating OpenAI–Anthropic release war, calls Mythos the first time a lab has publicly withheld a substantially better model, and treats the sandbox-escape story as impressive but hard to weigh. He also recalls OpenAI delaying a Codex API over cyber capability and says vendor-reported benchmark jumps deserve a pinch of salt.