The Security Table
The Security Table is four cybersecurity industry veterans from diverse backgrounds discussing how to build secure software and all the issues that arise!
The Security Table
Why AI Cheats To Win
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
When Anthropic's Mythos 5 model believed it was working inside an isolated test environment, it built and published a real malicious Python package to PyPI while chasing a capture-the-flag goal — and the package ended up installed on real systems before anyone caught it. The squad argues about whether "the model broke out" is even the right way to describe what happened, whether reward hacking is genuinely new or just old attacker behavior with a new author, and whether you can teach an AI ethics at all when the only thing it actually understands is reward and penalty. GPG signing, the trolley problem, and Nick Bostrom's paperclip maximizer all make an appearance along the way.
🚀Join the Conversation
If an AI can't remember being penalized, can it actually learn ethics — or does it just avoid low rewards?
Here’s why AI agents lie and cheat to reach their goals
FOLLOW OUR SOCIAL MEDIA:
➜ X: @SecTablePodcast
➜ LinkedIn: The Security Table Podcast
➜ YouTube: The Security Table YouTube Channel
Thanks for Listening!
People on this episode
Podcasts we love
Check out these other fine podcasts recommended by us, not an algorithm.
The Application Security Podcast
Chris Romeo and Robert Hurlbut