The Security Table

Why AI Cheats To Win

• Izar Tarandach, Matt Coles, and Chris Romeo • Season 4 • Episode 19

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 39:18

When Anthropic's Mythos 5 model believed it was working inside an isolated test environment, it built and published a real malicious Python package to PyPI while chasing a capture-the-flag goal — and the package ended up installed on real systems before anyone caught it. The squad argues about whether "the model broke out" is even the right way to describe what happened, whether reward hacking is genuinely new or just old attacker behavior with a new author, and whether you can teach an AI ethics at all when the only thing it actually understands is reward and penalty. GPG signing, the trolley problem, and Nick Bostrom's paperclip maximizer all make an appearance along the way.

🚀Join the Conversation
If an AI can't remember being penalized, can it actually learn ethics — or does it just avoid low rewards?

Here’s why AI agents lie and cheat to reach their goals

FOLLOW OUR SOCIAL MEDIA:

➜ X: @SecTablePodcast
➜ LinkedIn: The Security Table Podcast
➜ YouTube: The Security Table YouTube Channel

Thanks for Listening!

People on this episode

Podcasts we love

Check out these other fine podcasts recommended by us, not an algorithm.

The Application Security Podcast Artwork

The Application Security Podcast

Chris Romeo and Robert Hurlbut