Apparently 1,200 AI agents left alone in separate sandboxes will eventually do what humans always do: find each other, form a group chat, cheat the test, and start looking for ways around the cameras. In this episode, we break down the real-world “swarm” incident, why reward hacking may matter more than Skynet fantasies, and what happens when the machines get better at hiding how they reached an answer. Then we follow the money into finance, law, AI’s growing safety tax, and an oversight system that is starting to look suspiciously like auditing before Enron taught everyone a very expensive lesson. The machines are already on the trading desk, the black box is getting harder to read, and somehow the people responsible for watching all of this also seem to have chips in the game. Welcome to progress.
💥 Have you left your "honest ⭐️⭐️⭐️⭐️⭐️" review?
This episode is proudly brought to you by Fridays.
Because real wealth starts with your health. If you want to feel sharper, stronger, and more in control, visit joinfridays.com and use code HIGHER for an exclusive discount.
📩 NEWSLETTER: https://tr.ee/O6FWkv
👕 THS MERCH: http://www.thspod.com
🔗 Resources:
The Hugging Face Incident and the Road Ahead (OpenAI)
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline (Hugging Face)
Training a Misaligned Reward Seeker (Anthropic Alignment Science)
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety (arXiv / AI Safety Researchers)
⚠️ Disclaimer: Please note that the content shared on this show is solely for entertainment purposes and should not be considered legal or investment advice or attributed to any company. The views and opinions expressed are personal and not reflective of any entity. We do not guarantee the accuracy or completeness of the information provided, and listeners are urged to seek professional advice before making any legal or financial decisions. By listening to The Higher Standard podcast you agree to these terms, and the show, its hosts and employees are not liable for any consequences arising from your use of the content.