Rogue Agents
With both the major private “AI” companies looking to IPO with insane valuations because they are desperate for money, both OpenAI and Anthropic have gone for the most obvious marketing ploy imaginable… “Oh noes, our AI is so powerful it went rogue and escaped containment and hacked other companies”.
There are three options:
- This is entirely made up, and is just hype-feeding marketing.
- These incidents happened exactly as they say they did.
- Some aspects of the incidents are true, but it is being embellished for marketing purposes.
Considering many claims by both company’s CEOs for years now have been objectively and demonstrably bullshit, option 1 is the most likely. Option 3 is also quite likely, but means some parts of the stories are true. Option 2 is extremely unlikely as telling the truth is just something Altman and Amodei just don’t do.
So what if some or all of these stories are true? The most absurd thing is they talk about how these models are “autonomously” going “rogue”, claiming they didn’t prompt it to do whatever is claimed.
These models are being trained on the entire internet, which includes thousands of pages of AI doom-mongering, including from the companies themselves. So they (the models) are being explicitly trained on what to do to solve a problem, which the training data tells them could include looking directly for the solution where ever it would be found, including another company’s server.
Remember, training data is the most important prompt of all. These “agents” are not operating autonomously at any point. Every action is the result of billions of dollars of programming. Oh they’ll tell you that training an LLM is not programming, but that is dishonest, as they are selectively constraining the word “programming” to mean being programmed with a specific set of steps.
You can’t claim your LLM is a trained “mixture of experts” and also claim it hasn’t been programmed with as much information relating to the expertise of each subject matter. Poisoning training data is a known thing, so of course feeding models with training data stuffed with Terminator fan fiction fuckwittery from AI doomsday cults is going to affect what the models do.
They are not going rogue. They are not operating as autonomous agents. They are doing exactly what they are being trained (programmed) to do. There are hundreds of billions of reasons why possibly none of this is true… all of them are dollars. If any element of this is true, it is because the companies have brought this about themselves.