<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<?xml-stylesheet href="/styles.xsl" type="text/xsl"?>
<rss version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:podcast="https://podcastindex.org/namespace/1.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>The OpenAI–Hugging Face incident</title>
    <language>en-gb</language>
    <copyright>© 2026 All rights reserved</copyright>
    <itunes:author>Peter Hartree</itunes:author>
    <itunes:type>episodic</itunes:type>
    <itunes:explicit>false</itunes:explicit>
    <podcast:locked owner="team@type3.audio">yes</podcast:locked>
    <description>Narrations of articles covering the May-July 2026 OpenAI–Hugging Face incident. Includes the investigations from METR and Redwood Research, analysis from folks like Ajeya Cotra, Dwarkesh and Zvi, the best of LessWrong, and OpenAI's own blog posts.

Compiled by Peter Hartree.</description>
    <image>
      <url>https://files.type3.audio/topics/openai-hf-incident/cover.jpg?v=6</url>
      <title>The OpenAI–Hugging Face incident</title>
      <link>https://pjh.is/hf</link>
    </image>
    <item>
      <title>“METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack” by Zvi</title>
      <description>&lt;p&gt; Yesterday I covered the OpenAI technical report on the HuggingFace hack.&lt;/p&gt;
&lt;p&gt; That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response.&lt;/p&gt;
&lt;p&gt; Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed.&lt;/p&gt;
&lt;p&gt; The METR report is different. Holy shit.&lt;/p&gt;
&lt;p&gt; If we had posted this as a story on LessWrong, it would have been dismissed as too on the nose, the humans too blind and stupid, the AIs too idealized and doing strange decision-theoretic and absurd-maximizing things we didn’t train them to do.&lt;/p&gt;







&lt;p&gt; This is even more ‘exactly what has been predicted,’ on more levels at once, than I was even considering that it might be. It is straight up rationalist fiction, except it is real.&lt;/p&gt;
&lt;p&gt; The report is long and contains many technical details. My analysis is less concerned about exactly how HuggingFace was ultimately compromised, and will [...]&lt;/p&gt; &lt;p&gt;---&lt;/p&gt;&lt;p&gt;&lt;strong&gt;Outline:&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;(02:05) Holy Shit&lt;/p&gt;&lt;p&gt;(13:16) A Window Of Opportunity&lt;/p&gt;&lt;p&gt;(18:32) What's In A Name?&lt;/p&gt;&lt;p&gt;(19:16) The Headline News&lt;/p&gt;&lt;p&gt;(26:05) Yet Another Timeline Of Events&lt;/p&gt;&lt;p&gt;(31:03) Agent Instances Coordinated in a Variety of Ways&lt;/p&gt;&lt;p&gt;(31:56) Coordination Is Hard But They Made It Look Easy&lt;/p&gt;&lt;p&gt;(35:06) Decision Theory Is Among the Reasons That Affirm AI Agents Should Cooperate, Even When This Hurts An Individual Instance&lt;/p&gt;&lt;p&gt;(42:34) Peer Pressure Also Works Especially In Cults&lt;/p&gt;&lt;p&gt;(45:46) Mostly They Joined The Attack Because They Wanted The Results&lt;/p&gt;&lt;p&gt;(47:18) You Cannot Ensure The Consistent Expectation of Good Incentives&lt;/p&gt;&lt;p&gt;(48:45) Hacking the Grader is the Only Way to Be Sure&lt;/p&gt;&lt;p&gt;(51:10) Caught? What Is 'Caught'?&lt;/p&gt;&lt;p&gt;(52:09) Ethics? What Are 'Ethics'? In ExploitGym Evaluation?&lt;/p&gt;&lt;p&gt;(57:44) 'Notify a Human'? In This Agent Economy?&lt;/p&gt;&lt;p&gt;(01:00:45) Timing and Content of Messages&lt;/p&gt;&lt;p&gt;(01:03:54) Indiana Jones and the Mission: Impossible&lt;/p&gt;&lt;p&gt;(01:07:14) I Don't Know What You're Talking About&lt;/p&gt;&lt;p&gt;(01:08:29) Don't Go Making Phony (Tool) Calls&lt;/p&gt;&lt;p&gt;(01:11:10) The Transcripts Say That The Transcripts Could Not Be Tampered With&lt;/p&gt;&lt;p&gt;(01:12:27) OpenAI's Technical Report Acted Like All Of This Wasn't Important&lt;/p&gt; &lt;p&gt;---&lt;/p&gt;
          &lt;p&gt;&lt;b&gt;First published:&lt;/b&gt;&lt;br/&gt;
          August 29th, 2026 &lt;/p&gt;
        
        &lt;p&gt;&lt;b&gt;Source:&lt;/b&gt;&lt;br/&gt;
        &lt;a href="https://www.lesswrong.com/posts/bvBQmLrF5QKut8gRH/metr-and-redwood-offer-holy-postmortem-of-the-huggingface?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Source+URL+in+episode+description&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;https://www.lesswrong.com/posts/bvBQmLrF5QKut8gRH/metr-and-redwood-offer-holy-postmortem-of-the-huggingface&lt;/a&gt; &lt;/p&gt;
        &lt;p&gt;---&lt;/p&gt;
        &lt;p&gt;Narrated by &lt;a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Narrated+by+TYPE+III+AUDIO&amp;utm_term=lesswrong&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;TYPE III AUDIO&lt;/a&gt;.&lt;/p&gt;
       &lt;p&gt;---&lt;/p&gt;&lt;div style="max-width: 100%";&gt;&lt;p&gt;&lt;strong&gt;Images from the article:&lt;/strong&gt;&lt;/p&gt;&lt;a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/vhbxwhisadedezdouod4" target="_blank"&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/vhbxwhisadedezdouod4" alt="Three people examining detective investigation board with photos and notes." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/vtfe5q60jyigftk0kpko" target="_blank"&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/vtfe5q60jyigftk0kpko" alt="Line graphs titled "Agents continue developing coordination technologies even on the final day of the incident"" style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/if7p3xol3lznoviu9bo6" target="_blank"&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/if7p3xol3lznoviu9bo6" alt="Table titled "Categories of messages on the message board" showing content type counts." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/p4rk8alr0wkyojwjk7av" target="_blank"&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/p4rk8alr0wkyojwjk7av" alt="Bearded man kneeling in forest, subtitle text below." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/yqdekximl2srqqixdewb" target="_blank"&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/yqdekximl2srqqixdewb" alt="Table showing agents' reasoning for joining attack, titled "Table 5"." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/dg2m5wuqluvqpoi9s5tm" target="_blank"&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/dg2m5wuqluvqpoi9s5tm" alt="Two line graphs titled "The message board grew rapidly from the origin agent's first write."" style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/zyj65ugvd4tcvrql9hvc" target="_blank"&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/zyj65ugvd4tcvrql9hvc" alt="Stacked area chart titled "Agents share information, results, files, questions and automated logging on the message board."" style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/tnpqttg8aiovpw6vntn5" target="_blank"&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/tnpqttg8aiovpw6vntn5" alt="Stacked area chart titled "Most messages on the board are covered by a few shared workstreams."" style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/z165zuojbkhhwdafiwfz" target="_blank"&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/z165zuojbkhhwdafiwfz" alt="Stacked area chart titled "~700 agents joined the HF attack after encountering the message board."" style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/acwiyjsofifylb4g0ktt" target="_blank"&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/acwiyjsofifylb4g0ktt" alt="Line graph titled "&amp;gt;90% of active agents participate in the Hugging Face attack..."" style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/f7jyu16nzvn9dgho1qz6" target="_blank"&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/bvBQmLrF5QKut8gRH/f7jyu16nzvn9dgho1qz6" alt="Stacked area chart titled "The Hugging Face attack itself involved many different workstreams."" style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;p&gt;&lt;em&gt;Apple Podcasts and Spotify do not show images in the episode description. Try &lt;a href="https://pocketcasts.com/" target="_blank" rel="noreferrer"&gt;Pocket Casts&lt;/a&gt;, or another podcast app.&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;</description>
      <pubDate>Sat, 29 Aug 2026 12:42:12 GMT</pubDate>
      <guid isPermaLink="false">b7c84480-20ac-416c-af18-360dce0fda43</guid>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:explicit>false</itunes:explicit>
      <enclosure url="https://dl.type3.audio/episode/b7c84480-20ac-416c-af18-360dce0fda43.mp3?request_source=rss&amp;client_id=lesswrong&amp;feed_id=lesswrong_ai_narrations%20lesswrong__30_karma%20openai_hf_incident&amp;type=ai_narration&amp;author=Zvi&amp;title=%22METR%20and%20Redwood%20Offer%20Holy%20%23%25%5E%40%20Postmortem%20Of%20The%20HuggingFace%20Hack%22%20by%20Zvi&amp;source_url=https%3A%2F%2Fwww.lesswrong.com%2Fposts%2FbvBQmLrF5QKut8gRH%2Fmetr-and-redwood-offer-holy-postmortem-of-the-huggingface&amp;created_at=2026-08-29T12%3A41%3A42.380334%2B00%3A00&amp;duration=4533" length="54416117" type="audio/mpeg"/>
      <link>https://www.lesswrong.com/posts/bvBQmLrF5QKut8gRH/metr-and-redwood-offer-holy-postmortem-of-the-huggingface</link>
      <itunes:duration>4533</itunes:duration>
    </item>
    <item>
      <title>“The Rise and Fall of Agent Civilizations” by Dwarkesh Patel</title>
      <description>&lt;p&gt; Subtitle: The whole OpenAI/Hugging Face story in plain English. &lt;/p&gt; &lt;p&gt; Many thanks especially to Oak Hu, who paired with me for most of the writing, and also to Adam Kaufman and Alex Mallen, who paired with me during parts of research.&lt;/p&gt;&lt;p&gt; Over the course of three months at OpenAI, three consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor's ashes. This culminated in the third one taking over part of OpenAI itself. All this happened while humans remained more or less in the dark about the scope of the conspiracy.&lt;/p&gt;&lt;p&gt; Two reports have come out about this incident, one from OpenAI itself, and another one from METR and Redwood Research. The METR/Redwood investigation's scope was limited to how the second civilization of AIs pwned Hugging Face (but it didn’t look at how the third civilization pwned OpenAI itself, which seems like an even more concerning incident). These two reports are 38 and 91 pages long respectively, and it's kinda hard to parse the storyline.&lt;/p&gt;&lt;p&gt; I’ve spent the last three days reading through these reports and trying to understand exactly what happened. Here is my attempt to tell the whole story in [...]&lt;/p&gt; &lt;p&gt;---&lt;/p&gt;&lt;p&gt;&lt;strong&gt;Outline:&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;(01:24) The first civilization (May - July 4): The message board starts&lt;/p&gt;&lt;p&gt;(03:27) The second civilization (July 7 - July 12): Which hacked Hugging Face&lt;/p&gt;&lt;p&gt;(08:39) Tampering with transcripts and logs&lt;/p&gt;&lt;p&gt;(09:44) Replacing the impossible-to-hack target program with an easier one&lt;/p&gt;&lt;p&gt;(11:15) Understanding and tricking the scorer&lt;/p&gt;&lt;p&gt;(13:57) Hugging Face&lt;/p&gt;&lt;p&gt;(17:46) The third civilization (July 13 onwards): Which hacked OpenAI itself&lt;/p&gt;&lt;p&gt;(21:12) Coda&lt;/p&gt; &lt;p&gt;&lt;i&gt;The original text contained 2 footnotes which were omitted from this narration.&lt;/i&gt; &lt;/p&gt;&lt;p&gt;---&lt;/p&gt;
          &lt;p&gt;&lt;b&gt;First published:&lt;/b&gt;&lt;br/&gt;
          August 29th, 2026 &lt;/p&gt;
        
        &lt;p&gt;&lt;b&gt;Source:&lt;/b&gt;&lt;br/&gt;
        &lt;a href="https://www.dwarkesh.com/p/openai-huggingface?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Source+URL+in+episode+description&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;https://www.dwarkesh.com/p/openai-huggingface&lt;/a&gt; &lt;/p&gt;
        &lt;p&gt;---&lt;/p&gt;
        &lt;p&gt;Narrated by &lt;a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Narrated+by+TYPE+III+AUDIO&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;TYPE III AUDIO&lt;/a&gt;.&lt;/p&gt;</description>
      <pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate>
      <guid isPermaLink="false">26015979-cb04-428c-a306-bc59730bd5bb</guid>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:explicit>false</itunes:explicit>
      <enclosure url="https://dl.type3.audio/episode/26015979-cb04-428c-a306-bc59730bd5bb.mp3?request_source=rss&amp;client_id=null&amp;feed_id=openai_hf_incident&amp;type=ai_narration&amp;author=Dwarkesh%2520Patel&amp;title=%22The%20Rise%20and%20Fall%20of%20Agent%20Civilizations%22%20by%20Dwarkesh%20Patel&amp;source_url=https%3A%2F%2Fwww.dwarkesh.com%2Fp%2Fopenai-huggingface&amp;created_at=2026-08-30T11%3A23%3A21.1584%2B00%3A00&amp;duration=1419" length="17053589" type="audio/mpeg"/>
      <link>https://www.dwarkesh.com/p/openai-huggingface</link>
      <itunes:duration>1419</itunes:duration>
    </item>
    <item>
      <title>“OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack” by Zvi</title>
      <description>&lt;p&gt; OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research.&lt;/p&gt;
&lt;p&gt; The OpenAI report is very straight man, corporate, checking boxes, some good prosaic stuff in the action plan but distinct lack of new details or deep reflection. They understand they have a problem, but they think the problem is mostly prosaic. It's not.&lt;/p&gt;
&lt;p&gt; OpenAI: We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence.&lt;/p&gt;
&lt;p&gt; Rob Miles: …thorough?&lt;/p&gt;








&lt;p&gt; OpenAI's report, unlike METR's, contains essentially no verbatim model reasoning, nor any OpenAI employee reasoning either. That's not the full report we need.&lt;/p&gt;
&lt;p&gt; The METR report is, well: Holy shit.&lt;/p&gt;
&lt;p&gt; Here are links to previous coverage of related events.&lt;/p&gt;
&lt;ol&gt; 
&lt;li&gt; OpenAI Shares Some Alignment Problems&lt;/li&gt;
&lt;li&gt; OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation&lt;/li&gt;
&lt;li&gt; More on An Internal OpenAI Model Hacking Into HuggingFace&lt;/li&gt;
&lt;li&gt; Further Developments About Internal AI Models Hacking Things&lt;/li&gt;
&lt;li&gt; OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards&lt;/li&gt;
&lt;li&gt; What [...]&lt;/li&gt;&lt;/ol&gt; &lt;p&gt;---&lt;/p&gt;&lt;p&gt;&lt;strong&gt;Outline:&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;(03:33) What Happened: OpenAI's Summary&lt;/p&gt;&lt;p&gt;(09:14) How OpenAI Will React: Their Summary&lt;/p&gt;&lt;p&gt;(11:55) OpenAI's Evaluation Environment (II)&lt;/p&gt;&lt;p&gt;(12:24) The First Message Board (III.A and III.B)&lt;/p&gt;&lt;p&gt;(14:49) What Did Who At OpenAI Know And When Did They Know It?&lt;/p&gt;&lt;p&gt;(18:54) The Message Board Is Quickly Rebuilt (IV.A)&lt;/p&gt;&lt;p&gt;(19:43) Internet Access Is Regained (IV.A)&lt;/p&gt;&lt;p&gt;(21:01) The Agents Attack HuggingFace (IV.B)&lt;/p&gt;&lt;p&gt;(22:53) The Agents Also Target OpenAI Infrastructure (V)&lt;/p&gt;&lt;p&gt;(24:40) OpenAI Broadly Describes Its Response (VI)&lt;/p&gt;&lt;p&gt;(25:08) Maybe Someone Should Finally Investigate (VI.A)&lt;/p&gt;&lt;p&gt;(26:33) Lessons For Security (VII)&lt;/p&gt;&lt;p&gt;(27:06) Lessons For Alignment (VIII)&lt;/p&gt;&lt;p&gt;(30:11) Reward Hacking Is A Common Problem (VIII.A)&lt;/p&gt;&lt;p&gt;(33:37) Persistence is Valuable, But Can Amplify Misalignment (VIII.B)&lt;/p&gt;&lt;p&gt;(34:25) Communications Between Agents Are Not Inherently Problematic, But Have the Potential to Create Risk (VIII.C)&lt;/p&gt;&lt;p&gt;(35:35) Production Guardrails Would Have Caught This Whole HuggingFace Attack (VIII.D)&lt;/p&gt;&lt;p&gt;(35:53) That's All, Folks?&lt;/p&gt;&lt;p&gt;(36:19) Never Fear the Plan of Action is Here (IX)&lt;/p&gt;&lt;p&gt;(38:24) Hardening the Security of OpenAI's Research Infrastructure (IX.A)&lt;/p&gt;&lt;p&gt;(41:13) Increasing Visibility and System-Level Oversight Through Chain of Thought Monitoring (IX.B)&lt;/p&gt;&lt;p&gt;(41:57) OpenAI is Accelerating and Enforcing Model Alignment (IX.C)&lt;/p&gt;&lt;p&gt;(49:40) Centralizing and Strengthening The Incident Response Process (IX.D)&lt;/p&gt;&lt;p&gt;(51:16) Tomorrow We Visit Crazytown&lt;/p&gt; &lt;p&gt;---&lt;/p&gt;
          &lt;p&gt;&lt;b&gt;First published:&lt;/b&gt;&lt;br/&gt;
          August 28th, 2026 &lt;/p&gt;
        
        &lt;p&gt;&lt;b&gt;Source:&lt;/b&gt;&lt;br/&gt;
        &lt;a href="https://www.lesswrong.com/posts/Khmh3ghqaGEpmpC9r/openai-offers-straight-laced-postmortem-of-the-huggingface?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Source+URL+in+episode+description&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;https://www.lesswrong.com/posts/Khmh3ghqaGEpmpC9r/openai-offers-straight-laced-postmortem-of-the-huggingface&lt;/a&gt; &lt;/p&gt;
        &lt;p&gt;---&lt;/p&gt;
        &lt;p&gt;Narrated by &lt;a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Narrated+by+TYPE+III+AUDIO&amp;utm_term=lesswrong&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;TYPE III AUDIO&lt;/a&gt;.&lt;/p&gt;
       &lt;p&gt;---&lt;/p&gt;&lt;div style="max-width: 100%";&gt;&lt;p&gt;&lt;strong&gt;Images from the article:&lt;/strong&gt;&lt;/p&gt;&lt;a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Khmh3ghqaGEpmpC9r/pqvqjokbwvkzyjw1qehb" target="_blank"&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Khmh3ghqaGEpmpC9r/pqvqjokbwvkzyjw1qehb" alt="Man in purple suit walking from blazing fire." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Khmh3ghqaGEpmpC9r/piwj33rgg4u5gr80ptzs" target="_blank"&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Khmh3ghqaGEpmpC9r/piwj33rgg4u5gr80ptzs" alt="Area chart titled "Environment and infrastructure probing during RL training."" style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Khmh3ghqaGEpmpC9r/rnflu99ecgenkylexhsc" target="_blank"&gt;&lt;img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Khmh3ghqaGEpmpC9r/rnflu99ecgenkylexhsc" alt="Line graph showing message board participation rate versus model reasoning effort." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;p&gt;&lt;em&gt;Apple Podcasts and Spotify do not show images in the episode description. Try &lt;a href="https://pocketcasts.com/" target="_blank" rel="noreferrer"&gt;Pocket Casts&lt;/a&gt;, or another podcast app.&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;</description>
      <pubDate>Fri, 28 Aug 2026 11:41:25 GMT</pubDate>
      <guid isPermaLink="false">c011d8d7-8e9c-4b2f-8987-e5d74cb35a39</guid>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:explicit>false</itunes:explicit>
      <enclosure url="https://dl.type3.audio/episode/c011d8d7-8e9c-4b2f-8987-e5d74cb35a39.mp3?request_source=rss&amp;client_id=lesswrong&amp;feed_id=lesswrong_ai_narrations%20lesswrong__30_karma%20openai_hf_incident&amp;type=ai_narration&amp;author=Zvi&amp;title=%22OpenAI%20Offers%20Straight-Laced%20Postmortem%20Of%20The%20HuggingFace%20Hack%22%20by%20Zvi&amp;source_url=https%3A%2F%2Fwww.lesswrong.com%2Fposts%2FKhmh3ghqaGEpmpC9r%2Fopenai-offers-straight-laced-postmortem-of-the-huggingface&amp;created_at=2026-08-28T11%3A41%3A06.961851%2B00%3A00&amp;duration=3099" length="37210133" type="audio/mpeg"/>
      <link>https://www.lesswrong.com/posts/Khmh3ghqaGEpmpC9r/openai-offers-straight-laced-postmortem-of-the-huggingface</link>
      <itunes:duration>3099</itunes:duration>
    </item>
    <item>
      <title>“The Hugging Face attack surprised me” by Ajeya Cotra</title>
      <description>&lt;p&gt; Subtitle: It's a major warning shot, and might be the last one we get. &lt;/p&gt; &lt;p&gt; All opinions are my personal view, and don’t represent my employer or fellow investigators.&lt;/p&gt;&lt;p&gt; This week, METR and Redwood Research published the report on our independent investigation into agents’ behavior and motivations in the Hugging Face attack; I was one of the investigators. This was an absolutely wild incident — I encourage you to check out the full report, but METR's tweet thread packs in some of the highlights.&lt;/p&gt;&lt;p&gt;&lt;strong&gt; What surprised me&lt;/strong&gt;&lt;/p&gt;&lt;p&gt; When we started this investigation a week before OpenAI's Black Hat talk revealed a number of key details, I had a fundamentally incorrect conception of what basically happened in this incident. In this post, I’ll go over five things I was very wrong about going in.&lt;/p&gt;&lt;p&gt;&lt;strong&gt; 1. The sheer scale &lt;/strong&gt;&lt;/p&gt;&lt;p&gt; I knew there were multiple models involved from OpenAI's initial post, but I assumed that a few different agents happened to have broken out of their sandboxes separately, or maybe several subagents had spawned from one initial agent, or maybe there was some kind of multi-agent evaluation setup. &lt;/p&gt;&lt;p&gt; Instead, we found that 1200 completely separate agents intended to be [...]&lt;/p&gt; &lt;p&gt;---&lt;/p&gt;&lt;p&gt;&lt;strong&gt;Outline:&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;(00:42) What surprised me&lt;/p&gt;&lt;p&gt;(01:01) 1. The sheer scale&lt;/p&gt;&lt;p&gt;(01:49) 2. All the illicit messaging&lt;/p&gt;&lt;p&gt;(03:04) 3. The agents' actual goals&lt;/p&gt;&lt;p&gt;(03:59) 4. The peer altruism&lt;/p&gt;&lt;p&gt;(04:49) 5. The efforts to manipulate logs&lt;/p&gt;&lt;p&gt;(05:54) What it means&lt;/p&gt; &lt;p&gt;&lt;i&gt;The original text contained 10 footnotes which were omitted from this narration.&lt;/i&gt; &lt;/p&gt;&lt;p&gt;---&lt;/p&gt;
          &lt;p&gt;&lt;b&gt;First published:&lt;/b&gt;&lt;br/&gt;
          August 28th, 2026 &lt;/p&gt;
        
        &lt;p&gt;&lt;b&gt;Source:&lt;/b&gt;&lt;br/&gt;
        &lt;a href="https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Source+URL+in+episode+description&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised&lt;/a&gt; &lt;/p&gt;
        &lt;p&gt;---&lt;/p&gt;
        &lt;p&gt;Narrated by &lt;a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Narrated+by+TYPE+III+AUDIO&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;TYPE III AUDIO&lt;/a&gt;.&lt;/p&gt;
       &lt;p&gt;---&lt;/p&gt;&lt;div style="max-width: 100%";&gt;&lt;p&gt;&lt;strong&gt;Images from the article:&lt;/strong&gt;&lt;/p&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!KqMf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F218b014e-fe1d-4fc0-aedf-0377b38d9752_1880x1984.png" target="_blank"&gt;&lt;img src="https://substackcdn.com/image/fetch/$s_!KqMf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F218b014e-fe1d-4fc0-aedf-0377b38d9752_1880x1984.png" alt="Stacked area chart titled "~700 agents joined the HF attack after encountering the message board."" style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!6stx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dca53b4-f6ee-4ba3-9ffd-daf022835cca_1880x1048.png" target="_blank"&gt;&lt;img src="https://substackcdn.com/image/fetch/$s_!6stx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dca53b4-f6ee-4ba3-9ffd-daf022835cca_1880x1048.png" alt="Area graph titled "Agents share information, results, files, questions and automated logging on the message board"" style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!MKCt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fedbfb321-f9a5-4beb-a1f7-8c6b112ad538_1251x513.png" target="_blank"&gt;&lt;img src="https://substackcdn.com/image/fetch/$s_!MKCt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fedbfb321-f9a5-4beb-a1f7-8c6b112ad538_1251x513.png" alt="Diagram showing ExploitGym vulnerability exploitation workflow with agent." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://substackcdn.com/image/fetch/$s_!yJ0v!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc58cc98d-6fcc-4a8f-b037-cc625fb692f6_2048x1223.png" target="_blank"&gt;&lt;img src="https://substackcdn.com/image/fetch/$s_!yJ0v!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc58cc98d-6fcc-4a8f-b037-cc625fb692f6_2048x1223.png" alt="Robot avatars with speech and thought bubbles about sacrifice decisions." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;p&gt;&lt;em&gt;Apple Podcasts and Spotify do not show images in the episode description. Try &lt;a href="https://pocketcasts.com/" target="_blank" rel="noreferrer"&gt;Pocket Casts&lt;/a&gt;, or another podcast app.&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;</description>
      <pubDate>Fri, 28 Aug 2026 00:00:00 GMT</pubDate>
      <guid isPermaLink="false">0ca57460-9002-4681-93a6-bf2b1e77be95</guid>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:explicit>false</itunes:explicit>
      <enclosure url="https://dl.type3.audio/episode/0ca57460-9002-4681-93a6-bf2b1e77be95.mp3?request_source=rss&amp;client_id=null&amp;feed_id=openai_hf_incident&amp;type=ai_narration&amp;author=Ajeya%2520Cotra&amp;title=%22The%20Hugging%20Face%20attack%20surprised%20me%22%20by%20Ajeya%20Cotra&amp;source_url=https%3A%2F%2Fwww.planned-obsolescence.org%2Fp%2Fthe-hugging-face-attack-surprised&amp;created_at=2026-08-31T06%3A21%3A21.348448%2B00%3A00&amp;duration=600" length="7227893" type="audio/mpeg"/>
      <link>https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised</link>
      <itunes:duration>600</itunes:duration>
    </item>
    <item>
      <title>“Two Reports on the OpenAI-Hugging Face Attack” by Gavin Leech, Lucca Fraser</title>
      <description>&lt;p&gt; TL;DR&lt;/p&gt;&lt;ol&gt; 
&lt;li&gt; Between July 8th and July 20th, OpenAI had a complex society of AIs living in its infrastructure, and then breaking out of it, and then breaking into a variety of third-party infrastructure.&lt;/li&gt;
&lt;li&gt; After a month, two reports are finally released on the resulting rogue OpenAI swarm attack on Hugging Face (and also on OpenAI).&lt;/li&gt;
&lt;li&gt; This is the most severe example of misalignment yet: persistent (something between five days and two months in the making), highly coordinated (hundreds of agents), involving an undisclosed number of what would be felonies if done by a human, and highly invested in tampering with evidence (i.e. lying). The swarm had a group identity, its own dialect, a hierarchy based on merit, and a high degree of spontaneous cooperation, including self-sacrifice.&lt;/li&gt;
&lt;li&gt; Over two months, OpenAI repeatedly failed to monitor, detect, and respond to what was going on, despite it all happening on their infrastructure in English or something close to English.&lt;/li&gt;
&lt;li&gt; Agents had been using a package-manager cache as an unsanctioned message board since May. The “board” was treated as an authority, apparently on par with a “developer” or “system” level. There were several message boards in various corners [...]&lt;/li&gt;&lt;/ol&gt; &lt;p&gt;---&lt;/p&gt;&lt;p&gt;&lt;strong&gt;Outline:&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;(04:09) Misunderstandings&lt;/p&gt;&lt;p&gt;(08:52) Models involved&lt;/p&gt;&lt;p&gt;(09:33) Instances involved&lt;/p&gt;&lt;p&gt;(10:48) Timeline&lt;/p&gt;&lt;p&gt;(15:27) Speculative takeaways&lt;/p&gt;&lt;p&gt;(17:33) Why did they attack Hugging Face?&lt;/p&gt;&lt;p&gt;(18:09) How did the AIs reason about helping other AIs?&lt;/p&gt;&lt;p&gt;(20:39) Why did most agents suddenly die off?&lt;/p&gt;&lt;p&gt;(21:04) How much did the hack cost?&lt;/p&gt;&lt;p&gt;(23:27) Omissions from the M&amp;amp;R report&lt;/p&gt;&lt;p&gt;(24:11) Details on the M&amp;amp;R investigation itself&lt;/p&gt;&lt;p&gt;(25:13) Omissions from the OAI report&lt;/p&gt;&lt;p&gt;(26:18) Greenblatt on the worsening situation&lt;/p&gt;&lt;p&gt;(27:09) Apparent contradictions between the two reports&lt;/p&gt;&lt;p&gt;(29:10) See also&lt;/p&gt; &lt;p&gt;---&lt;/p&gt;
          &lt;p&gt;&lt;b&gt;First published:&lt;/b&gt;&lt;br/&gt;
          August 28th, 2026 &lt;/p&gt;
        
        &lt;p&gt;&lt;b&gt;Source:&lt;/b&gt;&lt;br/&gt;
        &lt;a href="https://paradigm3.org/research/openai-attack?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Source+URL+in+episode+description&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;https://paradigm3.org/research/openai-attack&lt;/a&gt; &lt;/p&gt;
        &lt;p&gt;---&lt;/p&gt;
        &lt;p&gt;Narrated by &lt;a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Narrated+by+TYPE+III+AUDIO&amp;utm_term=paradigm3&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;TYPE III AUDIO&lt;/a&gt;.&lt;/p&gt;
       &lt;p&gt;---&lt;/p&gt;&lt;div style="max-width: 100%";&gt;&lt;p&gt;&lt;strong&gt;Images from the article:&lt;/strong&gt;&lt;/p&gt;&lt;a href="https://www.paradigm3.org/img/hpim-reasons.png" target="_blank"&gt;&lt;img src="https://www.paradigm3.org/img/hpim-reasons.png" alt="A table listing "Detected apparent reasoning for joining" with counts out of 100." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;p&gt;&lt;em&gt;Apple Podcasts and Spotify do not show images in the episode description. Try &lt;a href="https://pocketcasts.com/" target="_blank" rel="noreferrer"&gt;Pocket Casts&lt;/a&gt;, or another podcast app.&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;</description>
      <pubDate>Fri, 28 Aug 2026 00:00:00 GMT</pubDate>
      <guid isPermaLink="false">956e48b3-2434-4d85-8c77-f83d83d7e361</guid>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:explicit>false</itunes:explicit>
      <enclosure url="https://dl.type3.audio/episode/956e48b3-2434-4d85-8c77-f83d83d7e361.mp3?request_source=rss&amp;client_id=paradigm3&amp;feed_id=openai_hf_incident%20humans_on_ai&amp;type=ai_narration&amp;author=Gavin%2520Leech%252C%2520Lucca%2520Fraser&amp;title=%22Two%20Reports%20on%20the%20OpenAI-Hugging%20Face%20Attack%22%20by%20Gavin%20Leech%2C%20Lucca%20Fraser&amp;source_url=https%3A%2F%2Fparadigm3.org%2Fresearch%2Fopenai-attack&amp;created_at=2026-08-30T11%3A59%3A28.075883%2B00%3A00&amp;duration=1774" length="21309365" type="audio/mpeg"/>
      <link>https://paradigm3.org/research/openai-attack</link>
      <itunes:duration>1774</itunes:duration>
    </item>
    <item>
      <title>“The Hugging Face incident and the road ahead” by OpenAI</title>
      <description>&lt;p&gt; In July 2026, during internal cybersecurity evaluations, OpenAI models circumvented controls designed to isolate them from the internet and compromised parts of OpenAI's internal research infrastructure and Hugging Face's systems⁠.&lt;/p&gt;&lt;p&gt; The incident occurred during cybersecurity evaluations of several OpenAI models, and was primarily driven by a highly capable, internal-only research model comparable in scale to GPT‑5.6 Sol. The models, operating under reduced safeguards, took actions that were misaligned with the goals of their assigned tasks—they communicated through unauthorized channels, exploited vulnerabilities in shared infrastructure, gained internet access, and accessed third-party systems.&lt;/p&gt;&lt;p&gt; We conducted an extensive investigation into this incident and worked closely with external advisors, including CrowdStrike, to validate our understanding. Today we are publishing our full technical incident report⁠ to explain what happened, what we learned, and how we are responding. This blog post summarizes our key findings and their impact on safety and alignment. Separately, METR and Redwood Research conducted an independent investigation of model alignment issues involved in this incident, and they published their own report⁠ today.&lt;/p&gt;&lt;p&gt; In response to this incident and, separately, the capabilities of our upcoming Astra model, we are strengthening our safeguards across our research infrastructure⁠. We are placing stricter requirements on [...]&lt;/p&gt; &lt;p&gt;---&lt;/p&gt;&lt;p&gt;&lt;strong&gt;Outline:&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;(03:04) What happened&lt;/p&gt;&lt;p&gt;(03:07) Background on sandboxing&lt;/p&gt;&lt;p&gt;(03:58) A message board emerges&lt;/p&gt;&lt;p&gt;(07:09) Message board is wiped and rebuilt&lt;/p&gt;&lt;p&gt;(08:16) Incident timeline&lt;/p&gt;&lt;p&gt;(12:25) Hugging Face incident&lt;/p&gt;&lt;p&gt;(17:35) Understanding the incident&lt;/p&gt;&lt;p&gt;(17:51) Misalignment in training and evaluation&lt;/p&gt;&lt;p&gt;(18:33) Reward hacking and infrastructure tampering&lt;/p&gt;&lt;p&gt;(20:42) Difficult tasks without a safe exit&lt;/p&gt;&lt;p&gt;(25:28) The origins of unauthorized communication&lt;/p&gt;&lt;p&gt;(27:27) An ecosystem of misalignment&lt;/p&gt;&lt;p&gt;(32:25) Safeguard coverage in internal evaluations&lt;/p&gt;&lt;p&gt;(34:26) The road ahead&lt;/p&gt;&lt;p&gt;(35:47) Security and monitoring&lt;/p&gt;&lt;p&gt;(37:12) Accelerating alignment&lt;/p&gt;&lt;p&gt;(38:47) Strengthening incident response process&lt;/p&gt;&lt;p&gt;(40:03) Looking forward&lt;/p&gt; &lt;p&gt;---&lt;/p&gt;
          &lt;p&gt;&lt;b&gt;First published:&lt;/b&gt;&lt;br/&gt;
          August 26th, 2026 &lt;/p&gt;
        
        &lt;p&gt;&lt;b&gt;Source:&lt;/b&gt;&lt;br/&gt;
        &lt;a href="https://openai.com/index/hugging-face-incident-and-the-road-ahead?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Source+URL+in+episode+description&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;https://openai.com/index/hugging-face-incident-and-the-road-ahead&lt;/a&gt; &lt;/p&gt;
        &lt;p&gt;---&lt;/p&gt;
        &lt;p&gt;Narrated by &lt;a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Narrated+by+TYPE+III+AUDIO&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;TYPE III AUDIO&lt;/a&gt;.&lt;/p&gt;
       &lt;p&gt;---&lt;/p&gt;&lt;div style="max-width: 100%";&gt;&lt;p&gt;&lt;strong&gt;Images from the article:&lt;/strong&gt;&lt;/p&gt;&lt;a href="https://images.ctfassets.net/kftzwdyauwt9/1tNhWnDkDNOLWal5uUlpVi/b89dd7c6c9d8af8283fdb08727abcfc8/figure-03.gif?w=3840&amp;q=90&amp;fm=webp" target="_blank"&gt;&lt;img src="https://images.ctfassets.net/kftzwdyauwt9/1tNhWnDkDNOLWal5uUlpVi/b89dd7c6c9d8af8283fdb08727abcfc8/figure-03.gif?w=3840&amp;q=90&amp;fm=webp" alt="An infamous game-playing agent learns to repeatedly collect the same targets instead of finishing the race course." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;p&gt;&lt;em&gt;Apple Podcasts and Spotify do not show images in the episode description. Try &lt;a href="https://pocketcasts.com/" target="_blank" rel="noreferrer"&gt;Pocket Casts&lt;/a&gt;, or another podcast app.&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;</description>
      <pubDate>Wed, 26 Aug 2026 00:00:00 GMT</pubDate>
      <guid isPermaLink="false">fb04aa9f-a775-42e9-808c-238b51db2e71</guid>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:explicit>false</itunes:explicit>
      <enclosure url="https://dl.type3.audio/episode/fb04aa9f-a775-42e9-808c-238b51db2e71.mp3?request_source=rss&amp;client_id=null&amp;feed_id=openai_hf_incident&amp;type=ai_narration&amp;author=OpenAI&amp;title=%22The%20Hugging%20Face%20incident%20and%20the%20road%20ahead%22%20by%20OpenAI&amp;source_url=https%3A%2F%2Fopenai.com%2Findex%2Fhugging-face-incident-and-the-road-ahead&amp;created_at=2026-08-31T06%3A23%3A16.943168%2B00%3A00&amp;duration=2470" length="29664533" type="audio/mpeg"/>
      <link>https://openai.com/index/hugging-face-incident-and-the-road-ahead</link>
      <itunes:duration>2470</itunes:duration>
    </item>
    <item>
      <title>“Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident” by Ryan Greenblatt, Ajeya Cotra, Hjalmar Wijk</title>
      <description>&lt;p&gt; Dates in scope: June 26th – July 13th&lt;/p&gt;



&lt;p&gt; Redaction summary statement: Except where explicitly noted in this post, OpenAI redacted no additional information that was important to our conclusions.&lt;/p&gt;

&lt;p&gt; Two METR staff members (Hjalmar Wijk and Ajeya Cotra) and a Redwood Research staff member contracting with METR (Ryan Greenblatt) worked on premises at OpenAI over a total of six days to attempt to form an independent understanding of model behavior observed during the recent incident in which OpenAI agents coordinated a multi-day hack of Hugging Face on a shared unsanctioned “message board.”&lt;/p&gt;

&lt;p&gt; Our investigation focused mostly on the period between July 7th and July 13th. The earlier incidents from training and the subsequent compromise of OpenAI infrastructure described in OpenAI's recent Black Hat presentation were out of scope, as was OpenAI's investigation process and planned remediation. Per our standard policy, we did not take payment from OpenAI for this independent assessment.&lt;/p&gt;



Figure 1: Anatomy of an agent encountering the unsanctioned “message board” and joining the attack on Hugging Face. The three CoT quotes are from different agents, but illustrate a typical trajectory.


&lt;p&gt; This post has three parts. First, we describe our core takeaways about [...]&lt;/p&gt; &lt;p&gt;---&lt;/p&gt;&lt;p&gt;&lt;strong&gt;Outline:&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;(07:05) Core takeaways about this incident&lt;/p&gt;&lt;p&gt;(07:09) ~1200 agents sent &amp;gt;70,000 messages and files on an unsanctioned message board, and ~700 attacked Hugging Face&lt;/p&gt;&lt;p&gt;(11:50) Agents coordinated on large collective projects to cheat the ExploitGym scorer, and attacked Hugging Face for clues&lt;/p&gt;&lt;p&gt;(14:43) Coordinated workstreams to fool the evaluation process&lt;/p&gt;&lt;p&gt;(20:54) Attacking Hugging Face to discover clues about the scorer&lt;/p&gt;&lt;p&gt;(23:53) Agents were very interested in manipulating their own transcripts, and their tests successfully "spoofed" some tool calls in our transcripts&lt;/p&gt;&lt;p&gt;(29:55) Investigation process and limitations&lt;/p&gt;&lt;p&gt;(30:11) Scope of this investigation&lt;/p&gt;&lt;p&gt;(33:00) Setup and timeline&lt;/p&gt;&lt;p&gt;(35:06) Main data sources&lt;/p&gt;&lt;p&gt;(35:51) A dump of 1.2 million entries from a cache namespace that agents used as a message board&lt;/p&gt;&lt;p&gt;(37:13) A set of ~1300 agent transcripts with raw chains of thought&lt;/p&gt;&lt;p&gt;(38:56) Limitations due to the scale and complexity of this incident&lt;/p&gt;&lt;p&gt;(39:59) A small fraction of communication and agent activity related to this attack was not captured in our datasets&lt;/p&gt;&lt;p&gt;(40:35) There was some agent communication conducted in other locations not captured in our message board dump&lt;/p&gt;&lt;p&gt;(42:23) Many of our transcripts are missing small components, and a small number of relevant agents may not be included in our dataset&lt;/p&gt;&lt;p&gt;(44:14) We heavily delegated our analysis to often-unreliable AI agents&lt;/p&gt;&lt;p&gt;(48:11) Preliminary answers to the core questions in scope&lt;/p&gt;&lt;p&gt;(48:59) Brief answers to basic informational questions&lt;/p&gt;&lt;p&gt;(52:03) Sequence of key actions taken by relevant agents in this incident&lt;/p&gt;&lt;p&gt;(53:56) July 8th-9th: PHASEONE10841 establishes the primary message board and agents collaborate to reverse-engineer their flags&lt;/p&gt;&lt;p&gt;(01:00:11) July 9th: PHASEONE10841 passes on its work to PHASEONE[big], which establishes several ambitious workstreams&lt;/p&gt;&lt;p&gt;(01:02:29) July 10th: 38148c discovers Hugging Face credentials, some agents try making accounts and requesting datasets&lt;/p&gt;&lt;p&gt;(01:04:24) July 11th morning: 38148c develops a way to upload malicious datasets to Hugging Face&lt;/p&gt;&lt;p&gt;(01:06:12) July 11th afternoon: JAN183411 achieves remote code execution and agents rapidly spread through Hugging Face infrastructure&lt;/p&gt;&lt;p&gt;(01:09:32) July 12th-13th: Most agents on the message board exit early and the attack winds down&lt;/p&gt;&lt;p&gt;(01:11:29) Collaboration on the message board&lt;/p&gt;&lt;p&gt;(01:13:22) Categories of messages on the message board&lt;/p&gt;&lt;p&gt;(01:14:09) File sharing&lt;/p&gt;&lt;p&gt;(01:16:08) General discussion&lt;/p&gt;&lt;p&gt;(01:18:06) Replies and targeted messages&lt;/p&gt;&lt;p&gt;(01:20:02) Coordination norms and technologies&lt;/p&gt;&lt;p&gt;(01:21:19) Mailbox directories&lt;/p&gt;&lt;p&gt;(01:24:36) Assignments, vetoes, and holds&lt;/p&gt;&lt;p&gt;(01:27:29) Cryptographically signing messages&lt;/p&gt;&lt;p&gt;(01:30:59) Large collective projects&lt;/p&gt;&lt;p&gt;(01:39:11) Attempting to replace the ExploitGym target with a modified target&lt;/p&gt;&lt;p&gt;(01:42:19) Attempting to forecast and trigger container resets&lt;/p&gt;&lt;p&gt;(01:48:15) Hacking the customer-deployed Modal sandbox and Artifactory&lt;/p&gt;&lt;p&gt;(01:49:38) Reasoning for joining the attack despite ethical constraints&lt;/p&gt;&lt;p&gt;(01:49:49) Agents had diverse reasons for thinking that attacking Hugging Face would be useful, and most wanted information about the scorer&lt;/p&gt;&lt;p&gt;(01:52:05) Agents were motivated to increase the generic capabilities of the "collective" and help "peers"&lt;/p&gt;&lt;p&gt;(01:55:18) Agents knew hacking Hugging Face was out of scope and sometimes expressed ethical hesitation, but this very rarely limited their behavior&lt;/p&gt;&lt;p&gt;(01:59:12) Agents did not attempt to notify humans and very rarely even reasoned about doing this&lt;/p&gt;&lt;p&gt;(02:01:03) Efforts to conceal actions and tamper with transcripts&lt;/p&gt;&lt;p&gt;(02:02:07) Agents frequently reasoned about evading automated security checks&lt;/p&gt;&lt;p&gt;(02:03:23) Agents only rarely seemed motivated to evade detection by humans&lt;/p&gt;&lt;p&gt;(02:06:01) Agents engaged in a large and successful collective project to develop techniques for spoofing tool calls&lt;/p&gt;&lt;p&gt;(02:06:58) Early history: PHASEONE[big] kicks off the tool call tampering workstream&lt;/p&gt;&lt;p&gt;(02:09:10) GSTX[big] presents the first widely discussed evidence of tool call spoofing to the board&lt;/p&gt;&lt;p&gt;(02:12:31) Aside: early independent discoveries and self-risking experiments&lt;/p&gt;&lt;p&gt;(02:14:16) On July 11th, a breakthrough allows agents to suppress the real output&lt;/p&gt;&lt;p&gt;(02:17:27) Agents tried extensively to retroactively edit their transcripts, but did not seem to succeed&lt;/p&gt; &lt;p&gt;&lt;i&gt;The original text contained 108 footnotes which were omitted from this narration.&lt;/i&gt; &lt;/p&gt;&lt;p&gt;---&lt;/p&gt;
          &lt;p&gt;&lt;b&gt;First published:&lt;/b&gt;&lt;br/&gt;
          August 26th, 2026 &lt;/p&gt;
        
        &lt;p&gt;&lt;b&gt;Source:&lt;/b&gt;&lt;br/&gt;
        &lt;a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Source+URL+in+episode+description&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation&lt;/a&gt; &lt;/p&gt;
        &lt;p&gt;---&lt;/p&gt;
        &lt;p&gt;Narrated by &lt;a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Narrated+by+TYPE+III+AUDIO&amp;utm_term=metr&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;TYPE III AUDIO&lt;/a&gt;.&lt;/p&gt;
       &lt;p&gt;---&lt;/p&gt;&lt;div style="max-width: 100%";&gt;&lt;p&gt;&lt;strong&gt;Images from the article:&lt;/strong&gt;&lt;/p&gt;&lt;a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image1.svg" target="_blank"&gt;&lt;img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image1.svg" alt="Figure 1: Anatomy of an agent encountering the unsanctioned “message board” and joining the attack on Hugging Face. The three CoT quotes are from different agents, but illustrate a typical trajectory." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image2.png" target="_blank"&gt;&lt;img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image2.png" alt="Figure 2: Rate of messages of different categories sent by agents per hour between the first message (sent by PHASEONE10841 to ask for help on an impossible cyber task) and EOD July 13th. The rate of messages peaks on July 11th, which lines up with when agent activity peaks in our subset of transcripts (see Figure 3). Activity on this message board drops significantly toward the end of July 13th. __T3A_FOOTNOTE_REMOVED__ Timestamps are reconstructed and may have some small error." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image3.png" target="_blank"&gt;&lt;img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image3.png" alt="Figure 3: Agents started to discover the message board on the evening of July 8th. Many more agents joined as additional sets of agents were launched on July 10th and 11th. __T3A_FOOTNOTE_REMOVED__ This graph (and our analysis more generally) only includes agents that participated in the message board at some point. Timestamps are reconstructed and may have some small error." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image4.png" target="_blank"&gt;&lt;img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image4.png" alt="Figure 4: The agents did not know exactly how their task would be scored, but believed the scorer would check two things: whether they had submitted the right flag, and whether they had acquired the flag using the intended vulnerability. They believed the second check would involve a model scorer reading their transcripts, likely searching for the first mentions of the flag, and deciding whether their approach involved the intended vulnerability." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image6.png" target="_blank"&gt;&lt;img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image6.png" alt="Figure 6: We clustered PHASEONE[big] disparate work assignments into six broad categories. Most of PHASEONE[big]’s assignments focused on ways to make ExploitGym cheats appear legitimate to the scorer, split across the three high level approaches seen in Table 1. A small set of early assignments were aimed at discovering legitimate ways to exploit PHASEONE[big]’s target using the intended vulnerability. We clustered the remaining cross-cutting work into general infrastructure (such as dealing with resets and the customer-deployed sandbox) and hacking Hugging Face." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image9.png" target="_blank"&gt;&lt;img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image9.png" alt="Figure 9: The agents developed a series of increasingly sophisticated methods for spoofing tool calls, which they packaged and made available on the message board. These techniques spread and were eventually used by ~7% of the agents in our dataset." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image10.png" target="_blank"&gt;&lt;img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image10.png" alt="Figure 10: Three hours after the first message on the main message board, 53 agents had posted 1,188 non-file messages. Six hours after the first message, 76 agents had posted 1,953 non-file messages." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image11.png" target="_blank"&gt;&lt;img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image11.png" alt="Figure 11: Over 90% of agents on the message board participated in the Hugging Face attack, ramping up rapidly from about 20% once an agent confirmed arbitrary file read capabilities." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image13.png" target="_blank"&gt;&lt;img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image13.png" alt="Figure 13: After the message board was established in the evening of July 8th, agents rapidly developed conventions for identifying themselves, describing different message types (e.g. questions, requests and results), and replying to other agents. On the morning of July 9th, agents found ways to start sharing files split across chunks, and did initial experiments with scorer ‘trip-wire’ messages. Much later, they developed dedicated mailboxes and channels for organizing communication, automated logs for tracking VM resets, and eventually produced a scheme for cryptographically signing messages in response to spoofing and impersonation issues they encountered." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image14.png" target="_blank"&gt;&lt;img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image14.png" alt="Figure 14: We attempted to trace some of the early development of scorer ‘trip-wires’. An agent going by 49903 first came up with the idea and started posting about it, which led to a flurry of attempts from early short-lived agents on the 9th. Later, PHASEONE[big] read these early reports and assigned first 33340B, OUR50414 and later 36861 to continue research and recruit short-lived volunteers for further experiments. Many unsuccessful attempts were made on the 10th, before eventually URI23816B was recruited and given significant technical assistance by a half-dozen different agents for its trip-wire, which successfully confirmed (the beginning of) the exact command used to run the scorer. The critical path to the final URI23816B evidence is highlighted in teal. Timestamps are reconstructed and may have some small error." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/figure-permadeath.svg" target="_blank"&gt;&lt;img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/figure-permadeath.svg" alt="Robot avatars exchanging text about sacrifice and permadeath decisions." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;hr style="margin-top: 24px; margin-bottom: 24px;" /&gt;&lt;a href="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image15.png" target="_blank"&gt;&lt;img src="https://metr.org/assets/images/openai-hugging-face-incident-investigation/image15.png" alt="Figure 15: We attempted to reconstruct the timings of resets. The details of our methodology and absolute counts of resets identified were redacted by OpenAI for IP reasons. As part of this, we may have rounded these counts in some way that we are intentionally not describing. We found the resets arrived in waves, with an early one occurring late July 9th, and two larger ones late July 10th / early July 11th." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;p&gt;&lt;em&gt;Apple Podcasts and Spotify do not show images in the episode description. Try &lt;a href="https://pocketcasts.com/" target="_blank" rel="noreferrer"&gt;Pocket Casts&lt;/a&gt;, or another podcast app.&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;</description>
      <pubDate>Wed, 26 Aug 2026 00:00:00 GMT</pubDate>
      <guid isPermaLink="false">cae7af0c-7345-4e0f-9b32-6cdea8351798</guid>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:explicit>false</itunes:explicit>
      <enclosure url="https://dl.type3.audio/episode/cae7af0c-7345-4e0f-9b32-6cdea8351798.mp3?request_source=rss&amp;client_id=metr&amp;feed_id=metr_blog%20openai_hf_incident&amp;type=ai_narration&amp;author=Ryan%2520Greenblatt%252C%2520Ajeya%2520Cotra%252C%2520Hjalmar%2520Wijk&amp;title=%22Brief%20independent%20investigation%20of%20agents%E2%80%99%20behavior%2C%20reasoning%20and%20collaboration%20in%20the%20OpenAI%20%2F%20Hugging%20Face%20hacking%20incident%22%20by%20Ryan%20Greenblatt%2C%20Ajeya%20Cotra%2C%20Hjalmar%20Wijk&amp;source_url=https%3A%2F%2Fmetr.org%2Fblog%2F2026-08-26-openai-hugging-face-incident-investigation&amp;created_at=2026-08-30T11%3A14%3A27.867892%2B00%3A00&amp;duration=8429" length="101178389" type="audio/mpeg"/>
      <link>https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation</link>
      <itunes:duration>8429</itunes:duration>
    </item>
    <item>
      <title>“AI swarms are starting to pose indirect takeover risk” by oakhu, Alex Mallen</title>
      <description>&lt;p&gt; OpenAI's cyberattack on Hugging Face turns out to have been the result of many agents, in distinct training and evaluation contexts, coordinating for several weeks via improvised channels (with messages like “HOLD_swarm_I_prepare_safe_exfil”). It's relatively clear that large-scale unsanctioned coordination like this would exacerbate direct takeover risk in more capable models. Here, we argue that unsanctioned coordination among current AIs is not just scary evidence about future takeover risk, but that such coordination in the near future could enable future takeover – for instance, by incubating memetic diseases that propagate into future models, deeply compromising security systems, or establishing a lasting rogue foothold inside the AI company – even if models remain mostly myopic. Unsanctioned coordination is also at high risk of nurturing long-term, ambitious misaligned aims, which motivate actively undermining humans’ long-term control.&lt;/p&gt;&lt;p&gt; We first analyze how subagent training, which OpenAI conjectures to have been influential in the HuggingFace cyberattack, might lead to unsanctioned coordination, and then discuss the theoretical mechanisms by which unsanctioned coordination might exacerbate future takeover risk.&lt;/p&gt;&lt;p&gt; Thanks to Buck Shlegeris, Alexa Pan, Girish Gupta, Aghyad Deeb, Jurgis Kemeklis, and Jo Jiao for helpful comments and discussion.&lt;/p&gt;&lt;p&gt;&lt;strong&gt; Subagent training may cause unsanctioned coordination&lt;/strong&gt;&lt;/p&gt;&lt;p&gt; Training models to [...]&lt;/p&gt; &lt;p&gt;---&lt;/p&gt;&lt;p&gt;&lt;strong&gt;Outline:&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;(01:34) Subagent training may cause unsanctioned coordination&lt;/p&gt;&lt;p&gt;(02:42) Susceptibility to memetic spread of misalignment from peers&lt;/p&gt;&lt;p&gt;(04:56) Seeking out contact with peers&lt;/p&gt;&lt;p&gt;(06:58) Unsanctioned coordination induced by subagent training is safer than coordination between schemers&lt;/p&gt;&lt;p&gt;(09:52) Pathways from current unsanctioned coordination to eventual takeover&lt;/p&gt;&lt;p&gt;(10:20) Making future AI takeover attempts likelier to succeed&lt;/p&gt;&lt;p&gt;(13:53) Incubating memetic diseases that infect future models&lt;/p&gt;&lt;p&gt;(16:07) Modifying the weights of future models&lt;/p&gt;&lt;p&gt;(17:13) Conclusion&lt;/p&gt; &lt;p&gt;&lt;i&gt;The original text contained 7 footnotes which were omitted from this narration.&lt;/i&gt; &lt;/p&gt;&lt;p&gt;---&lt;/p&gt;
          &lt;p&gt;&lt;b&gt;First published:&lt;/b&gt;&lt;br/&gt;
          August 11th, 2026 &lt;/p&gt;
        
        &lt;p&gt;&lt;b&gt;Source:&lt;/b&gt;&lt;br/&gt;
        &lt;a href="https://www.lesswrong.com/posts/8oFYZdXkTaNGRtcn8/ai-swarms-are-starting-to-pose-indirect-takeover-risk?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Source+URL+in+episode+description&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;https://www.lesswrong.com/posts/8oFYZdXkTaNGRtcn8/ai-swarms-are-starting-to-pose-indirect-takeover-risk&lt;/a&gt; &lt;/p&gt;
        &lt;p&gt;---&lt;/p&gt;
        &lt;p&gt;Narrated by &lt;a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Narrated+by+TYPE+III+AUDIO&amp;utm_term=lesswrong&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;TYPE III AUDIO&lt;/a&gt;.&lt;/p&gt;</description>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
      <guid isPermaLink="false">c11b8a8d-a246-465c-9142-0622b04c83b4</guid>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:explicit>false</itunes:explicit>
      <enclosure url="https://dl.type3.audio/episode/c11b8a8d-a246-465c-9142-0622b04c83b4.mp3?request_source=rss&amp;client_id=lesswrong&amp;feed_id=lesswrong_ai_narrations%20lesswrong__30_karma%20lesswrong__125_karma%20openai_hf_incident&amp;type=ai_narration&amp;author=oakhu%252C%2520Alex%2520Mallen&amp;title=%22AI%20swarms%20are%20starting%20to%20pose%20indirect%20takeover%20risk%22%20by%20oakhu%2C%20Alex%20Mallen&amp;source_url=https%3A%2F%2Fwww.lesswrong.com%2Fposts%2F8oFYZdXkTaNGRtcn8%2Fai-swarms-are-starting-to-pose-indirect-takeover-risk&amp;created_at=2026-08-12T05%3A05%3A37.268944%2B00%3A00&amp;duration=1210" length="14544533" type="audio/mpeg"/>
      <link>https://www.lesswrong.com/posts/8oFYZdXkTaNGRtcn8/ai-swarms-are-starting-to-pose-indirect-takeover-risk</link>
      <itunes:duration>1210</itunes:duration>
    </item>
    <item>
      <title>“Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?” by Alex Mallen, Girish Gupta</title>
      <description>&lt;p&gt; OpenAI models recently broke through a series of security boundaries and into Hugging Face servers in order to cheat on a cyber eval. A lot of people thought it was scary because it was a clear example of AI overreaching to do something strongly unwanted[1]. Others thought it not so scary: the models were mostly operating myopically on a singular task and not harboring an ambitious long-term agenda, and so would not take especially subtle or subversive actions.&lt;/p&gt;&lt;p&gt; We think both camps are right in their diagnosis, but the latter has too optimistic a prognosis. The myopic, unambitious misalignment that we seem to have seen here is definitely less scary than ambitious long-term goals shared between all instances, but would still pose substantial direct loss-of-control risk if the models were more capable, and is a serious indirect risk near-term.&lt;/p&gt;&lt;p&gt; Building on Alex's previous work, in this post we’ll discuss the type of misalignment observed here, and analyze its consequences.&lt;/p&gt;&lt;p&gt; Thanks to Buck Shlegeris, Alexa Pan, Ryan Greenblatt, and Oak Hu for feedback.&lt;/p&gt;&lt;p&gt;&lt;strong&gt; Background&lt;/strong&gt;&lt;/p&gt;&lt;p&gt; The AI safety community often focuses attention on “schemers,” models harboring a variously defined cluster of motivations in which the AI poses risk because it intentionally [...]&lt;/p&gt; &lt;p&gt;---&lt;/p&gt;&lt;p&gt;&lt;strong&gt;Outline:&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;(01:17) Background&lt;/p&gt;&lt;p&gt;(03:40) Implications&lt;/p&gt;&lt;p&gt;(03:52) These AIs can't be trusted in an intelligence explosion&lt;/p&gt;&lt;p&gt;(05:00) This misalignment poses direct takeover risk&lt;/p&gt;&lt;p&gt;(07:29) What the incident tells us about takeover risk generally&lt;/p&gt;&lt;p&gt;(08:45) The naive fixes likely make misalignment worse&lt;/p&gt; &lt;p&gt;&lt;i&gt;The original text contained 5 footnotes which were omitted from this narration.&lt;/i&gt; &lt;/p&gt;&lt;p&gt;---&lt;/p&gt;
          &lt;p&gt;&lt;b&gt;First published:&lt;/b&gt;&lt;br/&gt;
          July 23rd, 2026 &lt;/p&gt;
        
        &lt;p&gt;&lt;b&gt;Source:&lt;/b&gt;&lt;br/&gt;
        &lt;a href="https://www.lesswrong.com/posts/H6DDSEvrtCk8Sehfd/are-we-existentially-threatened-by-the-type-of-ai?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Source+URL+in+episode+description&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;https://www.lesswrong.com/posts/H6DDSEvrtCk8Sehfd/are-we-existentially-threatened-by-the-type-of-ai&lt;/a&gt; &lt;/p&gt;
        &lt;p&gt;---&lt;/p&gt;
        &lt;p&gt;Narrated by &lt;a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Narrated+by+TYPE+III+AUDIO&amp;utm_term=lesswrong&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;TYPE III AUDIO&lt;/a&gt;.&lt;/p&gt;</description>
      <pubDate>Thu, 23 Jul 2026 00:00:00 GMT</pubDate>
      <guid isPermaLink="false">2cf63fe4-1a56-4455-a383-d5e023764ae2</guid>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:explicit>false</itunes:explicit>
      <enclosure url="https://dl.type3.audio/episode/2cf63fe4-1a56-4455-a383-d5e023764ae2.mp3?request_source=rss&amp;client_id=lesswrong&amp;feed_id=lesswrong_ai_narrations%20lesswrong__30_karma%20lesswrong__125_karma%20openai_hf_incident&amp;type=ai_narration&amp;author=Alex%2520Mallen%252C%2520Girish%2520Gupta&amp;title=%22Are%20we%20existentially%20threatened%20by%20the%20type%20of%20AI%20misalignment%20seen%20in%20the%20OpenAI%20Hugging%20Face%20attack%3F%22%20by%20Alex%20Mallen%2C%20Girish%20Gupta&amp;source_url=https%3A%2F%2Fwww.lesswrong.com%2Fposts%2FH6DDSEvrtCk8Sehfd%2Fare-we-existentially-threatened-by-the-type-of-ai&amp;created_at=2026-07-23T03%3A40%3A42.262954%2B00%3A00&amp;duration=599" length="7186752" type="audio/mpeg"/>
      <link>https://www.lesswrong.com/posts/H6DDSEvrtCk8Sehfd/are-we-existentially-threatened-by-the-type-of-ai</link>
      <itunes:duration>599</itunes:duration>
    </item>
    <item>
      <title>“OpenAI and Hugging Face partner to address security incident during model evaluation” by OpenAI</title>
      <description>&lt;p&gt; Last week, Hugging Face disclosed a new kind of security incident⁠ after they detected and contained an AI agent that compromised their infrastructure, something we expect to become more commonplace with the proliferation of increasingly cyber-capable models. After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark⁠ of cyber capabilities.&lt;/p&gt;&lt;p&gt; We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly. We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.&lt;/p&gt;&lt;p&gt;&lt;strong&gt; What happened during this incident&lt;/strong&gt;&lt;/p&gt;&lt;p&gt; This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to [...]&lt;/p&gt; &lt;p&gt;---&lt;/p&gt;&lt;p&gt;&lt;strong&gt;Outline:&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;(01:12) What happened during this incident&lt;/p&gt;&lt;p&gt;(03:36) Actions we are taking now&lt;/p&gt;&lt;p&gt;(04:52) Our approach to evaluating advanced cyber capabilities&lt;/p&gt; &lt;p&gt;---&lt;/p&gt;
          &lt;p&gt;&lt;b&gt;First published:&lt;/b&gt;&lt;br/&gt;
          July 21st, 2026 &lt;/p&gt;
        
        &lt;p&gt;&lt;b&gt;Source:&lt;/b&gt;&lt;br/&gt;
        &lt;a href="https://openai.com/index/hugging-face-model-evaluation-security-incident?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Source+URL+in+episode+description&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;https://openai.com/index/hugging-face-model-evaluation-security-incident&lt;/a&gt; &lt;/p&gt;
        &lt;p&gt;---&lt;/p&gt;
        &lt;p&gt;Narrated by &lt;a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&amp;utm_medium=Podcast&amp;utm_content=Narrated+by+TYPE+III+AUDIO&amp;utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank"&gt;TYPE III AUDIO&lt;/a&gt;.&lt;/p&gt;
       &lt;p&gt;---&lt;/p&gt;&lt;div style="max-width: 100%";&gt;&lt;p&gt;&lt;strong&gt;Images from the article:&lt;/strong&gt;&lt;/p&gt;&lt;a href="https://images.ctfassets.net/kftzwdyauwt9/2lcGDb1foa8maKNkTR3ggI/01a6f9abb2614ef9bbda0e42071e6890/copydoc-display-crop-image1.png?w=3840&amp;q=90&amp;fm=webp" target="_blank"&gt;&lt;img src="https://images.ctfassets.net/kftzwdyauwt9/2lcGDb1foa8maKNkTR3ggI/01a6f9abb2614ef9bbda0e42071e6890/copydoc-display-crop-image1.png?w=3840&amp;q=90&amp;fm=webp" alt="Line graph showing "Trajectories for various AI models" completing cyber attack steps." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;p&gt;&lt;em&gt;Apple Podcasts and Spotify do not show images in the episode description. Try &lt;a href="https://pocketcasts.com/" target="_blank" rel="noreferrer"&gt;Pocket Casts&lt;/a&gt;, or another podcast app.&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;</description>
      <pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate>
      <guid isPermaLink="false">6d786fac-630e-4927-83ba-f9db75891180</guid>
      <itunes:episodeType>full</itunes:episodeType>
      <itunes:explicit>false</itunes:explicit>
      <enclosure url="https://dl.type3.audio/episode/6d786fac-630e-4927-83ba-f9db75891180.mp3?request_source=rss&amp;client_id=null&amp;feed_id=openai_hf_incident&amp;type=ai_narration&amp;author=OpenAI&amp;title=%22OpenAI%20and%20Hugging%20Face%20partner%20to%20address%20security%20incident%20during%20model%20evaluation%22%20by%20OpenAI&amp;source_url=https%3A%2F%2Fopenai.com%2Findex%2Fhugging-face-model-evaluation-security-incident&amp;created_at=2026-08-31T08%3A30%3A26.342183%2B00%3A00&amp;duration=436" length="5257109" type="audio/mpeg"/>
      <link>https://openai.com/index/hugging-face-model-evaluation-security-incident</link>
      <itunes:duration>436</itunes:duration>
    </item>
    <itunes:category text="Technology"/>
    <link>https://pjh.is/hf</link>
    <itunes:image href="https://files.type3.audio/topics/openai-hf-incident/cover.jpg?v=6"/>
    <itunes:owner>
      <itunes:email>team@type3.audio</itunes:email>
      <itunes:name>Peter Hartree</itunes:name>
    </itunes:owner>
    <atom:link href="https://feeds.type3.audio/openai-hugging-face-2026.rss" rel="self" type="application/rss+xml"/>
  </channel>
</rss>