Discussion about this post

User's avatar
Zac Hill's avatar

One thing that is really wild about this story, and which I appreciate your piece emphasizes, is that despite the act being the product of a test, *no one at OpenAI seems to have been actively monitoring the behavior of the model undergoing the test* as it was undertaking these actions. In other words, what might happen under conditions *not* deliberately constructed to identify misaligned behavior?

torchbearercommunity's avatar

A warning shot only works if those who need to hear it are told about it. This one surfaced because Hugging Face's security team spotted the intrusion and chose to go public on July 16. OpenAI found out it was involved days later in internal logs. This would have otherwise stayed secret inside one company. How many other such incidents have happened, or are happening right now?

No posts

Ready for more?