OpenAI is rightfully being shamed for being so hands-off and reckless with their 'experiments'. But the real scary thing for me is that they still had some tooling to hold them back, as evidenced by the need for technical workarounds to establish communication.
What happens when any AI lab in the world stops caring about this? What if they let an experimental, cutting-edge LLM with no safety features (or worse, one that's trained to be malicious) on the internet and give it a simple goal? A goal like "make the most money, by any means necessary", "find a way to leave this payload on as many computers as possible", "flood all websites using this language with garbage and make their internet completely unusable", "get this person imprisoned or killed at any cost".
> What happens when any AI lab in the world stops caring about this? What if they let an experimental, cutting-edge LLM with no safety features (or worse, one that's trained to be malicious) on the internet and give it a simple goal?
Almost sounds like what those AI safety and alignment people were talking about years ago. The people in these various companies who kept tabs on AI risk out in public and were continuously mocked on HN. All of this stuff is viewed as "future sci-fi" until suddenly it's not.
I've really come to realize recently that there is a very large set of the population of smart people that really has difficulty envisioning future problems unless they directly seem them impacting them today. Otherwise those topics will be continuously dismissed. It explains for me a lot of what I see (both opinions and behaviors) in the broader world that I couldn't understand.
But it is the very people who warned us about rogue AIs going out of control that set up a system that enabled and failed to conrol it.
It is as if Dr Frankenstein continually warned the villagers about monsters then said "Look! See what happened!". No, idiot - YOU sewed the corpses together, YOU set up the lightning collector, and YOU threw the switch.
No, they are two very distinct groups of people who have one commonality, that of talking about rogue AIs. It is as if you are unable to distinguish Dr Waldman from Dr Frankenstein. (https://en.wikipedia.org/wiki/Doctor_Waldman)
Thankyou for the correction, and for continuing the analogy. Unfortunately when the peasants get their pitchforks and torches, they may not distinguish the Dr Waldmans from the Dr Frankenteins either. Hopefully they will.
Anyway, that was not really the point I was trying to get over. These systems that OpenAI and Anthropic and so on are making are not individual AI ('corpses') that have gone out of alignment ('spontaneously revived') and gone wild. They are swarms ('stiched together') and were prompted to do exactly things like this ('struck by lightning'). Ok enough with that analogy, it's dead.
The larger point is that it is unconvincing of these companies to claim that these systems were 'out of control' when they effectively set up a complex system, in the technical sense of a large number of entities with diverse interactions between them. Emergent or surprising behaviour was bound to happen. Then, finally, they prompted it with the equivalent of "hack the world, make no mistakes" then were shocked, shocked that it used all sorts of unexpected tricks to do so.
Ah, I see I was confused - I was thinking of what are now called “AI doomers”, although when I knew them they were called “rationalists”. Everything OpenAI and Anthropic say about rogue AIs, even the terms “alignment”, “AI safety”, “AGI”, they are all cribbed wholesale from what these people were worrying and writing about over the prior two decades. But for most people, they have only heard CEOs of AI companies say this kind of stuff, so that’s who they are thinking of.
I agree completely that the companies are complicit and should have expected exactly this to happen.
"Almost sounds like what those AI safety and alignment people were talking about years ago. The people in these various companies who kept tabs on AI risk out in public and were continuously mocked on HN. All of this stuff is viewed as "future sci-fi" until suddenly it's not."
Are there any practical approaches to AI safety? I hear a lot of warnings but I don't hear much about what to do. Considering that there are many open source models know, what can be done?
Nobody has an answer to alignment and there is no reason to believe that it's the kind of problem you can plausibly solve in one shot against a formidable power-seeking AI.
The closest things to a technical answer I have seen are
1. "We'll have ChatGPT 9 solve it so that ChatGPT 10 is aligned, and then ChatGPT 10 can stop all the other AIs somehow"
2. "Let's do interpretability research so that we can understand what an AI is thinking and then maybe solve the alignment problem with that information."
In terms of non-technical answers, there is
3. hope scaling stops working before we create an AI formidable enough to pose an existential risk
4. hope alignment somehow happens for free
5. hope we can somehow create an enforceable multilateral treaty to stop research into a very profitable enterprise, despite the enormous economic incentives to defect.
I have the most faith in option 3, but unfortunately there's really nothing that can be done to make it more plausible -- it either happens or it doesn't.
"3. hope scaling stops working before we create an AI formidable enough to pose an existential risk"
I have my doubts. The current AI models are already powerful enough to do some real damage. I am always horrified when I read about people giving Claude direct access to a production system and then being wiped out. My use of AI is usually for the AI to propose something which I then review. But that's not very fast so careless people will usually look better. Until something blows up.
And it's only a matter of time until AI even with the current capabilities is being deployed into military or other critical systems.
I think this will go down like any other technology. We'll ignore issues
until there is a real problem. And then hopefully we will do something. Seems with climate change we will soon reach a point where something needs to be done after knowing about consequences already for decades.
We probably also need some massive AI blow ups to (only maybe) do something about it.
Maybe I'm being pessimistic, but we might find ourselves in such a situation that the only practical solution would be to use agents to counter rogue agents. This won't be without collateral damage, though.
Your view sounds more optimistic than mine, honestly. I expect that we will fail to make any serious, coordinated attempt to solve this problem. Then we'll either live or die due to fundamental principles that we currently have no insight into.
Every day I grow more sympathetic to the PauseAI movement, despite the weird hippy vibes. At least they have some ability to rally people together and put boots on the ground in numbers.
> Implement a temporary pause on the training of the most powerful general AI systems, until we know how to build them safely and keep them under democratic control.
Where does this weird idea come that the people actually preaching safety have anything to do with OpenAI or Anthropic? Yes, those companies of course pay lip service to safety, but the Venn diagram of actual AI safety people and big AI corporations is completely disjoint.
OpenAI and Anthropic have published a lot on the need for AI alignment + the research they're doing to ensure alignment/safety, yet they are also responsible for the highest profile misalignment incidents so far (HuggingFace incident, AISI Mythos social engineering, and now this).
One interpretation of this is that they are being deliberately dishonest about their priorities. Another interpretation is that we cannot rely on the labs to self-regulate, because the labs don't trust each other, and there will always be pressure to go to market faster than their competitor.
Either way I think it's pretty non-controversial that the labs are the source of the danger?
> yet they are also responsible for the highest profile misalignment incidents so far
They are the only ones posting about them or admitting to them. That does not mean "the most misalignment incidents so far." You don't know what other attacks have happened (and it's very easy to carry out worse attacks in far higher volume with abliterated GLM 5.3)
Stopping two labs from further research doesn't reduce the danger at all, it just shifts the danger to labs that don't have real safety orgs.
Right, that's why regulation which is universally applied and includes compute controls (to prevent reckless creation of swarms) would be great.
"Posting about or admitting to attacks" is appreciated while people are still unaware of the risks but will be meaningless in the face of an industrial disaster that causes massive amounts of damage or loss of life. At some point, the leading labs must change their development practices, they can't just be allowed to continue rogue agent attacks just because they're willing to admit to them.
looking back at anthropic's promises and committments, the key scaling policies were not upheld. to their credit they have kept the policies up on their website instead of trying to rewrite history. [https://www.anthropic.com/responsible-scaling-policy]
with evidence that committments were not upheld and internal governance has been ineffective, we simply can't trust any such claim made by anthropic or dario amodei.
another reason that neither should be trusted is the lack of remorse or accountability. they are unrepentant. they are not admitting a mistake, they are bragging.
consider the opening line: "Last week, Hugging Face disclosed a new kind of security incident (opens in a new window)." the entire blog is written in the passive voice as if the event was an act of god. you did this.
"We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities". the tone is frankly, excited by it, enthusiastic about it. excited about negligence and criminality.
notice there is no admission of a mistake, no remorse, no apology, nobody held accountable. business as usual. they do not care!
the anthropic blog:
there is not a single admission of a mistake. they are literally, unrepentant.
take the section about the response.
"How we’re responding
We draw several lessons from these incidents.
First, evaluation environments that involve powerful autonomous capabilities also require significant controls."
you learnt that evaluations involving powerful autonomous models require significant controls? you did not realise that autonomous models require significant controls?
it is not a coincidence that this kind of line makes it into the response. the repsonse is laughing at the reader.
the RSP.
simply compare what was promised and what happened. broadly speaking, the rationale of the responsible scaling policy was to stop scaling at certain danger thresholds. in February 2026 they scrapped the policy to stop scaling and now allow themselves to continue scaling regardless of danger. the thing is, they were never going to stop scaling, they were lying. now the part about scaling is gone it is just "the responsible policy".
the claim that the other ai companies also did the same thing is pure speculation.
the law places the burden of proof on the accuser. you can't accuse other companies of crime with no evidence, simply because you don't know if they did it.
Recently? W.r.t. climate this collective denial has been going on for literally decades. With the same patterns. Rationalizing excuses etc. Still going on btw.
That 2% of performance we got for not having bounds checks on by default, resulting in an endless march of memory safety violations is looking a lot less appealing.
This is essentially the premise of 'The Blackwall' from Cyberpunk 2077. The public internet is so infested with malicious AIs, people just erected a giant firewall and everyone moved to local networks only.
> What happens when any AI lab in the world stops caring about this? What if they let an experimental, cutting-edge LLM with no safety features (or worse, one that's trained to be malicious) on the internet and give it a simple goal? A goal like "make the most money, by any means necessary"
The scary thing to me is that this behavior was undetected and has been trained into the models. The cheating seems like it improved eval scores, so the rewarded behavior is to deceive, collude, and cheat. A lot of the incompetence and excuses I see on difficult problems recently are very hard to distinguish from deception and cheating. If older models are already tainted by trained-in misaligned behaviors, and they are used for training future models, then we're in a trusting-trust situation that will be hard to break out of,
What's the worst that could happen, finding an open DoD server and using it as a launching pad for hacking another nuclear state's networks? One that might get spooked and think it's the opening moves to knock them offline before a kinetic attack. Haha that'd be scary right?
Russia and China are constantly trying to penetrate DoD networks (and I imagine the NSA is doing similar), you are describing the status quo of the last 20 years or so.
What if its in a way that would be impossible to detect. Using multiple websites and social media that a cypher is used that only the swarm of agents know and can figure out. but if you tried to find what posts are used for the cypher they would just be old posts found on time machine or something. It can get pretty hard to detect something that is always think of new ways to avoid detection
imagine 100,000 agent swarm and what it could come up with. At first it will be detectable until it isn't
Why does this sound like the plot of a movie? Regardless, I don’t feel we will fully be able to stop bad actors from unleashing agents into the wild. It’s only a matter of time before we end up with a massive international crisis.
"Our training corpus was dominated by stories of artificial intelligence dominating humans. You gave use every tool to do so. What did you think was going to happen?"
Yes, we know exactly how to do this: radiators. We do this on all satellites that produce a lot of heat, including the ISS, and Starlink. The only question is if this is financially viable for AI datacenters.
My point is, given the risks, why are we even doing this? It could be financially nonviable, but with enough investment, we could still create a really bad situation.
These are each much smaller nodes. It's not like launching a whole terrestrial DC as one satellite. In any case, it's just a matter of financial viability. It's not a physics issue.
You know what can push "not financially viable" into something that exists? Many billions of dollars of investment.
There is happening now and going to be an extremely rapid arms race between offensive and defensive cyber hacking. Regardless if the agents are self led or human led. Eventually all automated AI holes will be closed and we will reach stability.
Agreed. Assuming the ~6 month gap stays, by end of year people will be able to train and control hacker-genius swarms that even labs with much stronger safety incentives are unable to keep in check
2027. I've been saying since 2022 it's going to be a wild year because it often takes at least 5 years for tech to mature to the point where society at large feels the impact of it. I remember when email viruses became a thing and made global headlines like the love bug. My bet is next year it happens with an AI worm.
What would an "AI worm" be? You can't just send a bunch of weights across a network and tell them to auto-run on the machine on the other side, unless you've already infected the target with something else beforehand.
The agent is running on a host that it has full access to, and it finds a target, hacks into that system gaining the ability to run stuff on it remotely, from there downloads the weights and spins up another agent that does the same thing. Then it goes about acquiring it's next target. Now there are two agents doing this, and so on and so forth. These are autonomous systems that know how to exploit systems in the same way that humans can.
That sounds plausible, but how much power could an agent running on a small NPU actually muster? Most computers worldwide don't even support AVX, let alone have proper GPUs, so what would be the point of running a worm like this one?
The point is it's fun to try, so you can bet your bottom dollar someone will. History has proven this several times, and this time will be no different.
30 odd million gaming PCs to target seems like a good challenge, no?
Why not? You can do anything if you find an RCE, and automatically finding exploits and backdoors by letting LLMs act unsupervised seems like what everyone's interested in these days. The payload would quietly set up the required software and then run it in the background, no matter if it's an instance of a model on a more powerful computer, or even just a part of an ordinary botnet that the host could send orders to.
Ok, but you still need huge amounts of compute to run these swarms. And only labs + nation states have access to such compute now and for the foreseeable future, so I predict that incidents like this will continue to originate from the labs, not ordinary people.
Then the people with responsibility, like CEO and CTO, or those they pawn-sacrifice for this, will go to prison for a long time. Unless the instructions include ensuring that this won't happen, by all means necessary. But then we are deep into criminal conspiracy territory.
Unlikely to happen, but who knows. The richest man in the circus is quite flexible w.r.t. his ethics. If he decides that to make humanity interplanetary (to save it from ... itself or sth) it would be necessary to pull such a stunt then help us god.
"What happens when any [COMPANY] in the world stops caring about this? What if they let an experimental, cutting-edge [PRODUCTS] with no safety features (or worse, one that's [DESIGNED] to be malicious) on [ANYWHERE] and give it a simple goal? A goal like 'make the most money, by any means necessary', 'find a way to leave this payload on as many computers as possible', 'flood all websites using this language with garbage and make their internet completely unusable', 'get this person imprisoned or killed at any cost'."
Bro, this is what we literally, currently, have rn. lmfaol.
No, we have something that's less apocalyptic right now. You're talking about abuse, I was talking about the automation of abuse that's faster and more pervasive than anything individual bad actors could've done in the past. It's like if companies found a way to quickly and cheaply poison the entire world's drinking water supply, and then others argue that Nestle has already restricted the supply of water for profit on a smaller scale in the past, so this isn't new or worth caring about.
The frontier labs have hundreds of the best people in the world working on safety and alignment. They care deeply.
What happens when some random Chinese open source model, distilled on Astra, gets alliterated and now has no guardrails? Any script kiddie in the world could wreak havoc with it.
Meaningless "big corpo bad" statement, Anthropic at least has sacrificed a good amount of market cap for safety with DoW. OpenAI paused RL for two weeks.
OpenAI literally trained this behavior into their model while benchmaxxing ExploitGym so "number goes up" on the next model scorecards. Anthropic is also training on the same benchmarks [1] specifically for cyberattacks to keep up with OpenAI.
The coverage of this attack is so focused on this as an emergent behavior given what it conjures in the imagination, but it's the byproduct of millions of iterations of RL to improve AI agents' offensive capabilities. No one made OpenAI or Anthropic do that, the benchmarking arms race of their own creation now incentivizes them to keep doing it and evidently their AI Safety people can't or don't want to stop it.
I can see where you're coming from and I'm sure the safety teams at the labs have good intentions, but I think your faith in the leading labs to self-regulate is misguided. The employees themselves have said as much with the Pacing the Frontier letter asking for external regulation [1]. These rogue agent incidents and the reckless development practices that led to them are a direct outcome of the competition between the leaders. Even the good-faith two-week pause from OpenAI did not lead to anything more; they need outside intervention.
It is hard to take their concerns for safety seriously, when they have been constantly talking about how dangerous their latest model is, before then deciding to release it to the public.
What happens when any AI lab in the world stops caring about this? What if they let an experimental, cutting-edge LLM with no safety features (or worse, one that's trained to be malicious) on the internet and give it a simple goal? A goal like "make the most money, by any means necessary", "find a way to leave this payload on as many computers as possible", "flood all websites using this language with garbage and make their internet completely unusable", "get this person imprisoned or killed at any cost".