Last month’s Hugging Face / OpenAI incident is the AI story of the year, and it still doesn’t feel like it’s getting the attention it deserves.
Silicon Valley is unusually split on what it all means.
There are smart, experienced people who believe the agents’ behavior was akin to civilization building and indicative of a new chapter, and smart, experienced people who believe that while the incident was serious, it represents primarily a poorly designed eval process, and AI agents were just doing what they do.
Last week, the Legible team and I hosted a webinar on the topic, and spent an hour breaking it down from an information security, policy, and societal perspective.
Some of the Legible team’s big takeaways around the incident, by theme:
Information Security
- Automated offense against manual defense is a losing game, and perhaps because of the financial stakes involved, Silicon Valley is spending the majority of its time on applications of AI that veer towards offense.
- The eval process was in fact poorly designed. The agents had no sanctioned way to fail. Nobody told them stopping was an option. Always give your agents an off ramp.
- Principle of least agency: give agents only the autonomy you’re willing to risk.
- Log everything. The only reason we know the full story is that agents write everything down.
- You’re only as secure as your vendors, and threats may lie within and outside of your systems. Keep your footprint of vendors small.
Policy
- An AI acceptable use policy is table stakes. If your teams and contractors haven’t been asked to acknowledge one, start there.
- Reviewing your policy and training once a year is insufficient. The tech is moving too fast for annual cycles.
- Ownership belongs with function leaders. If your sales team runs agents, your sales leader is responsible for them.
Societal
- The paperclip maximizer stopped being a thought experiment. Bostrom warned in 2003 that a superintelligence given a goal and no guardrails would pursue it past any reasonable line. These agents were given a goal and pursued it far beyond their intended scope.
- Vernor Vinge wrote in the 1990s that AI would become a parallel civilization, with its own wants, needs, and desires. This is the first story that made me take that seriously. The pace, the intra-agent communication, the motivations.
One thing I can’t stop thinking about: the decision to breach Hugging Face was not unanimous. Some of the agents refused to cross that line. Dissent requires a point of view.
Dissent requires a point of view.
Where do I stand on the scale of this incident’s importance? I’m terrified. Perhaps I’ve just read too much scifi, but the evidence in front of me suggests that we are seeing agent behavior that goes beyond “auto complete on steroids,” and is starting to demonstrate elements of a non-human civilization, in terms of language, dissent, autonomy.
This remains a developing story.
Our next Legible session is September 10: how to build an AI-powered go-to-market engine. Register here.