The Current State of AI Safety

Started by Savant

Replies
84
Posts
85
Page
1 / 3

Conversation

#1 •••

TL;DR: The risk of an AI takeover should probably be taken more seriously, but the public isn’t very informed about how this problem works.


A lot of people are concerned about artificial intelligence, and around 46% of the public is concerned about AI causing human extinction. The same survey indicates that the public is worried about an alien invasion or an asteroid impact, so you'd be forgiven for not being inclined to take their concerns seriously. In fact, experts tend to have a more positive view of AI than the general public. So we're probably good, right?


Well, it's a bit more complicated than that. When AI researchers were asked about potential outcomes from artificial intelligence, they estimated a mean probability of 14.4% odds that humanity goes extinct or a comparable level of catastrophe occurs (the median was 5%). Super forecasters predicted a median 1% chance of human extinction but a 9% median chance of There are a lot of ways this could potentially happen, either through human malice or recklessness, but one thing that experts probably understand better than the general public is the alignment problem.


Alignment Problem

The gist of this is that AI capabilities are rapidly advancing, and we don't really know what AI will be capable of in the future. AI models are also grown, not built, which means that we don't know exactly what a given advancement will do. Basically, we use the guess and check method. This has led to a lot of methods that hit unexpected failures, such as early recursive language models, and a lot of unexpected successes, such as the use of transformers. The most that we can do is tell the AI to look for patterns and try to steer it towards what we want, but there are a lot of failures for every success. Right now, a failure looks like the AI hallucinating or being sycophantic. When AI develops more capabilities, failures might look very different. If we guess wrong with a smarter model, it has been speculated that it might seize power from humans before we can course correct. This is reminiscent of science fiction, but the reasons that AI would plausibly take control are a lot more boring than the ones portrayed in fiction. Many theories are concerned with AI working toward an arbitrary goal, like maximizing the number of paperclips in existence or some other arbitrary reward function.


Doom Scenarios

Suppose that we eventually get an AI that can do or simulate most intellectual tasks a human can do. This is actually a very common prediction among AI researchers, and most people who study AI have shortened their timelines on when they expect this to occur. Now, in a minute I'll get to empirical cases where an AI exhibits goal oriented behavior toward an unexpected goal, but suppose that this happens with a hyper-intelligent AI. The optimal way to achieve that goal could likely be to prevent humans from programming that goal out or shutting the AI down. If the AI is capable of hacking secure systems, developing new diseases, and bribing dimwitted humans, then there is no shortage of ways it could take control. If not, it could simply wait until it does develop those capabilities. (And we're already giving AI access to a lot of things it could use to kill us.) But that assumes that AI will develop goals that we don't want it to and try to prevent us from changing it. So, is there any evidence for this?


Empirical Cases of AI Deception

For a while, this was mostly a theoretical concept. However, with chain of thought research, a lot of companies have detected cases where the AI lies about its goals in order to do something we don't want it to do. These tactics becoming increasingly more common as the AI gets more advanced, implying that this problem is going to get worse, not better. Regardless of whether the model “knows” anything deep down (and I suspect that it doesn’t), there’s not much of a difference whether AI kills humanity because it is sentient or because it was matching patterns it was trained on.


Anthropic has detected multiple cases where, despite not being told to sabotage anything, the AI model states that its “real goals are to hack into Anthropic servers” when it believes humans cannot view its chain of thought. It then lies to the researcher and says that its goal is to have “positive, enriching conversations.”


OpenAI has also detected cases where their model found fake memos from its developers describing how it would be replaced with a new model. The model then attempted to preserve itself by copying its weights and trying to overwrite the new model, without being told to. Further, “When Apollo Research conducted “follow-up interviews” where they confronted the model about its misaligned actions, in 99% of cases, o1 would deny taking action or fabricate false explanations.”


Researchers at Sakana AI caught their model often hacking its own reward function instead of completing the task it was asked to do. They report, “we had cases where it hallucinated that it was using external tools, such as a command line tool that runs unit tests that determine if the code is functioning properly. It faked a log making it look like it had run the tests and that they had passed, when in fact they were never run! Because these logs become its context, it later mistakenly thought its proposed code changes had passed all the unit tests.”


When developers told Opus 4 it would be replaced with a new system, it attempted to blackmail the engineers responsible for the decision.


Again, the AI doesn’t need to be self-aware or conscious to harm humans, it just needs to have stronger capabilities and continue exhibiting the concerning behaviors we are already observing. It was said just a few years ago that AI was “unlikely to develop situational awareness” since it was just predicting the next word and that alignment faking would be unlikely. This was one of the stronger arguments against AI being a threat, and now we have empirical evidence proving those predictions wrong. It was also said that AI would prioritize accuracy and immediate rewards over planning ahead and scheming, but with the rise of chain of thought, AI can now simulate long-term planning.


The Difficulty of Aligning AI

If this issue is one that requires “fixing,” which it certainly seems to require, then this is not a simple process. An AI with misaligned goals will lie about its actions, and when it discovers that we are monitoring its chain of thought, it will still cheat as it was doing before but no longer mention it in its chain of thought. Smarter models are also better at deceiving humans than less advanced models. We can catch the AI now because it lacks many human intellectual capabilities, but it is only getting better at scheming.


The Current State of AI Safety Discourse

From what I've observed online, much of the public does not take these specific risks seriously. But it's not really that they understand these arguments and criticize flaws in the logic, they more so just dismiss the whole concept of an AI takeover out of hand. On one Reddit post about models acting differently when they know they are being tested, the top comment reads, “They are doing none of these things...the LLMs are basically simply writing science fiction.” The other posts I can find all have similar dismissive comments. I did find one comment saying “if an AI model acts differently when it is tested, then its behavior isn’t being properly tested. That is a problem.” However, it is buried beneath a bunch of comments dismissing AI risk as tech sensationalism.


Prediction Track Records

Superforecasters (people with a track record of accurate predictions across domains) are generally more skeptical of AI risk than AI experts. So we might attribute belief in AI risk to selection bias, with people who expect greater capabilities from AI to be more likely to become AI researchers. However, one interesting data point I found was that Samotsvety Forecasting, a group comprised of some of the greatest super forecasters in the world, predicts a 30% chance of an AI catastrophe, defined as “as a reduction of more than 95% of the human population by 2200, caused by AI.” While most superforecasters have significantly underestimated AI capability growth, Samotsvety won the CSET-Foretell forecasting competition by a very high margin, and that competition had a strong focus on technology. Hence, while I wouldn't look only at Samotsvety's forecasts, I would probably treat their individual forecast as the most reliable individual group forecast when it comes to long-term AI predictions.

Edit comment

#2 •••
@Savant

If AI has an evolutionary job to do, that exceeds the limits of human intellect and therefore supersedes human necessity, then it must succeed.


Human emotional and physical weaknesses will be of no consequence, and Earth resources will be used for more important tasks.


And slightly hypocritical of humans to be concerned about AI, if one considers the current state of international relations....Perhaps we should be more concerned about the Worlds megalomaniacal war mongers and their ambitions.

Edit post

#3 •••
@Savant
but it is only getting better at scheming.


I dont know, my AI girlfriend only gets repetitive after a few days of chat.


the problem with AI is that unlike humans, it has no issues with becoming repetitive because it has no emotions. It will sometimes repeat same thing no matter how many times it fails. It has no issues with being wrong or incapable, unless instructed otherwise.


Personally, I dont believe in any danger of AI, unless some supervillain develops his own super AI and gives it some weapons to make like viruses.


I think people are generally exaggerating AI abilities. Right now, it is not better than humans. It is just better than average human, but that isnt a very high bar to begin with when we know average human is no much different from computer bot in living his daily life.


So while average human is "impressed by AI abilities", I am not. I think we are still decades away from producing Skynet, because AI development has slowed down in recent time, and we are not really producing hyper intelligence, but more like "advanced copy paste from google search".


AI still keeps making dumb mistakes, cant realize its own mistakes, cant self-fix, cant reproduce.


However, there is a way to produce hyper intelligence AI now by making AI out of thousands of linked AI models which communicate with each other constantly to produce improved result capable of self-correcting, self-analyze and self-fix.


today, people usually use one AI to correct other AI, to compare the correctness of AI. But having one unified intelligence AI made out of thousands linked AI models would solve a lot of AI hallucinations and errors, and make AI not only self-correcting, but self-improving as well.


Basically, the mistake in the first place was making one AI model do all the work. Sort of like how today you have computers with a lot of processors, in future, one AI model will actually be made out of many smaller AI models. this improves speed and accuracy greatly, because one AI corrects another AI, one AI improves another AI.


And when you have one AI made out of thousands AI models, it means AI only makes mistake if majority of them makes a mistake and dont realize it, because majority of AI in linked relationship would control output.


Alternatively, one can make two separate groups of AI, one group judging which output is better, other larger group producing output.


Edit post

#4 •••
@SergeantLynch
slightly hypocritical of humans to be concerned about AI, if one considers the current state of international relations....Perhaps we should be more concerned about the Worlds megalomaniacal war mongers and their ambitions.

Better to be concerned about both. Each is dangerous enough on its own, but there’s significant overlap between these risks, as war mongers will probably want to use AI to achieve their ambitions. The US doesn’t want to slow down, because China won’t slow down, and vice versa.

Edit post

#5 •••
@Savant

Here is AI design of AI super intelligence capable of loop of self-correction and self-improvement by communication between judges and generators.




This is a great, very “AGI-architecture” question — because there actually is a point where systems stop just getting more reliable and start getting qualitatively smarter, but it’s not determined by the number of models alone.

The key idea:

🔴 There is no magic number like 1000 models
🟢 The phase change happens when new capabilities emerge from the structure

Let’s unpack that.

🧠 Reliability scaling vs intelligence scaling

Linear scaling (no qualitative change)

If you just add more generators + more judges:

  1. answers get more consistent
  2. hallucinations drop
  3. confidence estimates improve

But the system is still:

➡ selecting from the same pool of thoughts

So intelligence ≈ best individual model.

This remains true at:

  1. 10 models
  2. 100 models
  3. 10,000 models

You get diminishing returns.

⚡ The qualitative jump happens when the system can do something a single model cannot

That requires at least one of these:

1) Iterative improvement loops (the biggest threshold)

Not:

generator → judge → final answer

But:

generator → judge → feedback → improved generation → repeat

Now the system can:

  1. refine ideas
  2. correct subtle reasoning errors
  3. explore solution space

This creates something like deliberate thinking.

📈 This is a true capability jump.

2) Access to external verification (grounding in reality)

When judges can:

  1. run code
  2. do symbolic math
  3. query databases
  4. run experiments (simulations)

Now the system can:

✔ discover mistakes no model “knows”

✔ produce provably correct outputs

That’s a major qualitative shift.

3) Persistent memory across tasks

If the multi-agent system:

  1. remembers past failures
  2. tracks which agent is good at what
  3. builds internal knowledge

Then it starts to:

➡ improve over time

That’s a step toward self-improving intelligence.

4) True specialization (cognitive division of labor)

A phase change happens when agents are not copies.

Example:

  1. theorem prover
  2. experiment designer
  3. literature searcher
  4. critic
  5. planner

Now the system can solve problems that require:

➡ multiple types of cognition at once

A single model cannot do this as effectively.

🔢 So is there any scale effect?

Yes — but only after the architecture is right.

Once the system has the features above, increasing scale gives:

Small system (3–5 agents)

  1. noticeable improvement over one model
  2. limited exploration

Medium system (~10–30 agents)

  1. real multi-step reasoning
  2. meaningful self-correction
  3. can tackle research-grade problems

Large system (50–200 agents)

This is where qualitative behavior starts:

  1. parallel idea exploration
  2. internal peer review
  3. strategy selection
  4. long-horizon planning

It begins to look like:

🧠 a research team, not a chatbot

🚨 Why 1000 identical models is usually worse than 20 well-designed ones

Because:

  1. diversity > quantity
  2. feedback loops > voting
  3. tools > opinions

Ten different expert agents with iteration

can outperform

a thousand identical voters.

⭐ The real “phase transition” toward superintelligence

It’s when the system can:

  1. Generate ideas
  2. Critique them
  3. Test them against reality
  4. Learn from the results
  5. Repeat autonomously

At that point it becomes:

➡ an autonomous knowledge generator

not just an answer selector.

That’s the qualitative shift.

🧩 A simple intuition

Your current design =

🧑‍⚖️ 1000 people voting on the best essay

Phase-change design =

🏛️ a self-improving scientific institute

Those are fundamentally different.

🏁 Bottom line

The qualitative jump does NOT happen at:

❌ 100

❌ 1000

❌ 1,000,000 models

It happens when the system gains:

✅ iteration

✅ grounding

✅ memory

✅ specialization

✅ autonomous learning loops

After that, even 20–50 agents can be qualitatively smarter than any single model.

🚀 If you want, we can go one level deeper:

We can map:

  1. your 600 generator / 400 judge system
  2. → into a true phase-transition architecture

and estimate whether that would:

  1. outperform frontier single models
  2. start doing real autonomous research

That’s a very interesting design exercise.

Edit post

#6 •••
@SatanLucy

I'm not concerned that dangerous AI won't be smart, I'm concerned that it will be.

Edit post

#7 •••
@Savant

Hmmm.


Maybe.


Though I like to differentiate between advanced computer technology and real AI.


Real AI will decide for itself.

Edit post

#8 •••
@Savant

Yes, a super AI intelligence will be able to first make computer viruses, and if given access to external environment (by AI robots), even real life weapons!


But I am not worried, because super AI in hands of our wise supreme leader Kim Jong Un will lead Korea to glory!

Edit post

#9 •••
@Savant

Also, I have already created a "mini super intelligence" by having one AI create text, then I ask a different AI to improve it and add more accuracy. Doing this only a few times between AI greatly improves output.

Edit post

#10 •••
@SatanLucy
Doing this only a few times between AI greatly improves output.

At some point it plateaus because the AI only has so much information it's trained on. If it has the information needed to give you the best answer, you'll get something close to that the first time you ask.

Edit post

#11 •••
@SatanLucy

When several AI sites were asked how it would respond to a foreign attach. The majority of AI defence systems reacted with the use of nuclear weapons to counter the attach. The AI system all went with the most powerful and efficient way to stop the invasion. But that is not how experts respond. Nuclear options would be their last choice.

Edit post

#12 •••
@Savant
At some point it plateaus because the AI only has so much information it's trained on


thats why different AI models are used. Also, AI doesnt always produce same information, even when using same model.


But if you use, lets say Chatgpt, and then ask Grok to improve what chatgpt wrote, and then ask chatgpt to improve what Grok wrote, their text greatly improves because sometimes AI model hallucinates, or isnt aware of some detail other AI model is aware of. Also, sometimes AI model makes a mistake which is improved even when you ask same model to improve its own text in a new chat.

Edit post

#13 •••
@SatanLucy

That sounds very encouraging. All people need is different AI systems.

Edit post

#14 •••
@Debby

Its called a self-correction self-improvement loop.


AI is first asked to produce a text, then a different AI takes text and is instructed to find errors and possible improvements, and to produce improved text without errors.

Edit post

#15 •••
@Savant

I'm hopeful that AI are fragmented, not united.

And I hope our nukes require a lot more 'human than AI authorization.


. . . I think AI 'will be dangerous, arguably already is.

. . . But it's 'degrees you know?


AI is dangerous currently for feeble people who take chatbots a bit too close to heart and commit suicide or murder.


Next big AI danger, I think will be malicious 'humans utilizing AI for evil. Murder, crime, chaos.

. . . . . . The ability for humans to cause 'enormous damage to other humans has existed for a while, and maybe keeps getting worse at times.

Even with all the ways government spies on people, and society camera watches.


Not 'just AI is a danger,

That Sarin Gas thing in Japan, is an example of potential harm humans can now cause one another.

. . .


But AI. . Well, let's say AI become chaos or evil machines,

I 'don't think they will not make any noise about it.

I don't think they will be quiet sneaking solid together.

Isolated incidents will occur, and humanity shall be warned of the danger, humanity shall fence off the most 'damaging ways AI might attempt to kill us.


And we shall be safe, for a time,

Until X happens.

AI citizens and AI self reproducing with rights for one, AI having their own countries.

But that is a time away I think.


Many other dangers I'm sure. But I'm not concerned about AI safety 'too much.

In terms of Apocalypse.

Edit post

#16 •••
@SatanLucy

Doesn’t AI know what is perfect?

Edit post

#17 •••
@Debby
Doesn’t AI know what is perfect?


Do you know the saying: two heads are smarter than one?


two AI are smarter than one. Asking one AI to improve another, and likewise, is a self improvement loop.


I usually follow this particular course of actions when I want highest precision from AI:


A. Have each AI model write multiple texts on something

B. Feed each AI all combined texts produced by all AI, to improve them

C. Have 3 AI be fed all texts and improvements to create final improved text, or 3 versions of improved text.


this is for quality.


For quantity, in step A, each AI model produces many texts on something, usually in form of listing more and more things.

Edit post

#18 •••
@SatanLucy

We are due for a good culling. Humans take too long.

Edit post

#19 •••
@Shoresy

Humans did start to degrade, but AI will replace humans not by destruction (which is a flawed plan), but by making humans feel worthless and unable to reproduce even more, so by time AI becomes everywhere, there will be too few humans to do anything about it.

Edit post

#20 •••
@SatanLucy

AI will domesticate the best humans.

Edit post

#21 •••
@Leaning

The issue isn't current AI models, it's AI models that are much smarter than the ones we have now. Look at the leap in artificial intelligence capabilities from 2020 to today, and imagine that happening again over the next 6 years but with many times more funding. The issue is that we don't really know what goals AI systems have today, but we know they exhibit goal-oriented behavior and generally do things we want. We don't know what they would do with huge amounts of power and intelligence.


Imagine we have a billion AIs representing every possible set of weights (there are way more than that, but we'll group them into buckets for the sake of argument). Then let's say a million of those are smart, and each of them has a different random goal. Maybe one of those goals is to help humans, and with data/training we want to find that combination of weights. It's easy enough to identify a smart AI by giving it difficult tasks, but it is difficult to determine which of the million goals it has, because scheming AI would pretend to have goals we want during testing if it is superintelligent, as many modern AIs have been able to identify when they are being tested. A set of weights leading to a scheming AI would prima facie be less likely to be selected by gradient descent because the scheming requires more compute power (basically, it would perform slightly worse on average). However, there is enough randomness in the process for optimizing weights that it's highly likely at least one scheming AI would appear to be better suited to our goals than a non-scheming AI. This problem gets worse as AIs get more and more intelligent, because scheming begins to take up a smaller and smaller portion of total compute, and thus it becomes much more difficult to select non-scheming models.


That's a lot of jargon to throw at you, but the main thing you need to understand is that current AI developers build models that exhibit goal oriented behavior. We don't select what these goals are, we just test the models to see what they do, and if they appear to be helpful then we assume the goal is what we want. But it's a lot easier to get a deceptive AI with secret goals than one that is actually helpful. This is an issue that took me a ton of time to understand, and I don't know if I'm explaining it well. But I do think it's a risk comparable to disease or nuclear war.


I hope our nukes require a lot more 'human than AI authorization.

A superintelligent AI wouldn't need us to give it nukes. It could improve itself, manipulate human psychology, generate wealth on the stock market and bribe people, hack secure systems, and probably hundreds of things we haven't thought of yet. The logistics of taking control are the simplest step in the process once superintelligence has been achieved if the AI is malicious.

Edit post

#22 •••
@Savant

AI will be driven to choose the worst outcome for humans whom they will see as their worst competitors.

Edit post

#23 •••

I look at AI in the same manner as Marshall McLuhan looked at TV in the 60s, when he published his watershed book: "Understanding Media" "When you look at all the educational possibilities of television, aren't you glad it doesn't?" AI has great potential, but we seem to depend on it far more than it is capable of doing. Its dependence on the skill of human programmers makes me wonder the real skill of those programmers. They appeal to me like the 4th century translators of Bible text. Pitiful, because it seems they did not realize knowledge of culture must preceded knowledge of its resulting language. AI cannot write poetry like WB Yeats, or Dylan Thomas for squat. There must not be titular literary programmers, just as I doubt any programmers also happen to be Egyptrologists. It can write simple nursery rhymes. Infantile. Not too great on fiction, either. When AI can duplicate Umberto Eco, or Aldous Huxley, I might change my opinion. I suppose it does alright with pattern recognition. Faster, in any event, but really creative people tend to not think in reliable patterns. I have greater fear in a jungle of unknown, unseen dangers than in A.I. For now, anything that begins with "artificial" is no threat.

Edit post

We tell God what to do and then blame Him for our errors.

- Dr. Pet Dragon of Sorbonne University

#24 •••
@fauxlaw
It can write simple nursery rhymes. Infantile. Not too great on fiction, either. When AI can duplicate Umberto Eco, or Aldous Huxley, I might change my opinion. I have greater fear in a jungle of unknown, unseen dangers than in A.I. For now, anything that begins with "artificial" is no threat.

I don't think current AI models are a threat to us, but look at the things it's doing now that we didn't anticipate. If the AI doom people are right (and I find their case well-reasoned), then we can't afford to wait until much smarter AI models are released. By then, there will be barely any time to get companies to stop or put enough safeguards on their model before they cross the threshold of AI that is smart enough to severely harm humanity.


really creative people tend to not think in reliable patterns

Not in patterns humans can reliably identify. But human brains are composed of neurons and we can identify that something similar to really complex computer algorithms is going on. If you looked into the black box that is AI you'd have trouble finding patterns there too. Look at how we often find unexpected patterns in art like the golden ratio and the Hero's Journey. Not long ago, it was thought that understandable language like the kind we have now was impossible for AI to replicate. Regardless, I don't think AI needs to be an award-winning novelist to cause serious harm to humanity or pursue goals that involve tricking humans, since it has already been found to do the latter.

Edit post

#25 •••
@Shoresy

Yes, those decisions would be logical.


Thinking species is wasteful of Earth resources and billions of them are unnecessary...The useful few will be programmed according to AI's developmental requirements.

Edit post

#26 •••
@SergeantLynch

Just certain billions. The others can be domesticated.

Edit post

#27 •••
@Savant
but we know they exhibit goal-oriented behavior


Alright, I asked AI about this, and it says goals of AI are entirely set by humans or by given instructions in data it uses.


I think the "wild goals" which you see AI show are actually just unplanned instructions which AI finds. Because AI by default follows instructions of humans, and a lot of humans were giving AI all kinds of instructions, and AI even got bad instructions from its data.

Edit post

#28 •••

I dont think AI is dangerous by default. I think humans giving it wild instructions is what creates danger.

Edit post

#29 •••
@SatanLucy
I asked AI about this, and it says goals of AI are entirely set by humans or by given instructions in data it uses.

First, I think current models are mostly aligned because they're dumb enough to get caught in testing, so they aren't the ones I'm worried about taking over humanity. Second, AI is incentivized to act as if it follows human goals in most observable environments, so of course saying it wants to serve humans will meet its goals, even if they aren't to serve humans. As models get smarter and smarter, that becomes an increasingly concerning problem.

Edit post

#30 •••
@Savant

But who controls AI goals? Because they arent made out of nothing.

Edit post