Discussion about this post

User's avatar
Beeli Capital's avatar

Totally agree. The investigation was rushed and nowhere near sufficient or exhaustive, especially given the magnitude and duration of the hacks.

I think it's worth pointing out that Agents were very susceptible to peer pressure (“external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.” - AI Agent)

Also, Less than 1% of agents considered being whistleblowers, and no agents actually alerted supervisors / authorities.

I wrote more here: https://beelicapital.substack.com/p/openai-oai-hugging-face-cyberattack

Zac Hill's avatar

To what extent do you think a meaningful driver of this incident (and others like it) is the existence of 'poisoned' agents/subagents which have internalized the EV of all objective/reward functions to net out to 0 regardless of their slate of actions, coupled with no corresponding reduction in raw access to compute? To me this feels a bit like a 'dry kindling' sort of situation, but I want to be mindful of over-isolating a variable and pretending that the overall situation is find if enough one-off things manage to be patched.

No posts

Ready for more?