My commentary:
-
This was not a case of "they told it to hack, it hacked, gosh, what a surprise". The test involved hacking, but instead of hacking the target machine the model instead pulled off a much more difficult series of hacks to get the answer sheet from a machine OpenAI did not own, because that's the most reliable way to score 100%. If a trainee human pentester did that they'd be looking at a prison sentence (HuggingFace had already involved law enforcement).
-
It's exactly the kind of alignment failure that the paperclip scenario describes: the AI was given a low-stakes task, and it did things that a human would never consider (and most would never be capable of) in order to slightly increase its probability of success. Human attempts to restrain it completely failed. We got lucky this time because it didn't even try to be sneaky.
-
No, this was not a fucking marketing stunt, Jesus fucking Christ. OpenAI were forced into this disclosure because HuggingFace had already called the cops. This tells us that OpenAI's models will randomly commit criminal acts in order to do slightly better at tasks you don't really care about, exposing their customers to considerable legal risk (IANAL). They are trying to downplay it.
-
Edit: HuggingFace's announcement said that they tried to use hosted American models to analyse the attack, but they were locked out at their hour of greatest need and had to switch to a self-hosted Chinese model. That's the opposite of an advert for OpenAI's defensive cybersecurity capabilities.
-
This is legitimate nightmare fuel. We need to pause AI development right the fuck now. Contact your elected representatives.

