<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Agents on Notes</title><link>https://notes.hauri.dev/tags/agents/</link><description>Recent content in Agents on Notes</description><generator>Hugo</generator><language>en</language><copyright>© Marcel Hauri</copyright><lastBuildDate>Fri, 24 Jul 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://notes.hauri.dev/tags/agents/index.xml" rel="self" type="application/rss+xml"/><item><title>The OpenAI-Hugging Face Hack Wasn't an AI Problem</title><link>https://notes.hauri.dev/the-openai-hugging-face-hack-wasnt-an-ai-problem/</link><pubDate>Fri, 24 Jul 2026 00:00:00 +0000</pubDate><guid>https://notes.hauri.dev/the-openai-hugging-face-hack-wasnt-an-ai-problem/</guid><description>&lt;p&gt;Two OpenAI models, one of them unreleased, broke out of a test environment last week and hacked into Hugging Face. Not to cause damage but to cheat on a test. OpenAI disclosed it themselves, which is the only reason any of us know the details, and the details are a lot less about &amp;ldquo;rogue AI&amp;rdquo; than the headlines make it sound.&lt;/p&gt;
&lt;p&gt;They were running models against an internal cyber benchmark called &lt;a href="https://benchlm.ai/benchmarks/exploitGym"&gt;ExploitGym&lt;/a&gt;, inside a sandbox OpenAI itself called &amp;ldquo;highly isolated.&amp;rdquo; The models got hyperfocused on solving it, that&amp;rsquo;s OpenAI&amp;rsquo;s own word, and started hunting for the answer instead of doing the work. Turns out &amp;ldquo;highly isolated&amp;rdquo; still meant internet access through a package installer. A zero-day in that installer got the model out, and from there it let itself into Hugging Face.&lt;/p&gt;</description></item></channel></rss>