Regarding AI x-risk

Started by Savant

Replies
47
Posts
48
Page
1 / 2

Conversation

#1 •••
@SkepticalOne, @SatanLucy

Much of my argument is based on Eliezer Yudkowsky's essay here, Paul Cristiano's essay here, and information elsewhere on LessWrong. Also, ironically, on asking Gemini AI to help make sense of their arguments. A lot of these essays were a bit confusing to me when I first read them, so I'll put my argument in the form of a syllogism, since I think that will make it more clear how the conclusion follows. Hopefully this makes the conversation easier, as you can specify which premise or premises in particular you object to.


Definitions:

ASI: "Artificial Superintelligence," or a form of artificial intelligence that vastly surpasses human intelligence in every domain.


P1: If ASI exhibits goal-oriented behavior toward an unknown, arbitrary terminal goal, it will likely cause severe, permanent harm to humanity.

P2: If ASI emerges within the next 30 years and exhibits goal-oriented behavior, it is likely to be toward an unknown, arbitrary terminal goal.

P3: If ASI emerges within the next 30 years, it is likely to exhibit goal-oriented behavior.

P4: ASI is likely to emerge within the next 30 years.

C1: Therefore, ASI is likely to exhibit goal-oriented behavior.

C2: Therefore, ASI is likely to exhibit goal-oriented behavior toward an unknown, arbitrary goal.

C3: Therefore, ASI is likely to cause severe, permanent harm to humanity.


Defending P1: If ASI exhibits goal-oriented behavior toward an unknown, arbitrary terminal goal, it will likely cause severe, permanent harm to humanity.


This is probably best explained by the paperclip maximizer thought experiment. If an ASI has an arbitrary goal (like maximizing the number of paperclips), it can do a better job of this if humans cannot shut it off. So the actions that work best toward an arbitrary terminal goal include either permanently depowering humans or perhaps killing them all. An ASI could easily deceive humanity about its intentions and pretend to be cooperative and then kill humans and take control of human infrastructure in a myriad of ways. Eliezer, for example, has detailed how it would be physically possible for an ASI to synthesize advanced nanomachinery and use it to kill all humans at once. Even if only used AI as a chatbot, Yudkowsky's AI-box experiment has also shown that it would probably not require ASI-level intelligence to convince a human to give the ASI access to additional infrastructure and capabilities. But AI is already integrated into power grids, nuclear systems, and water facilities, and as long as it "played nice," an ASI would probably be rapidly integrated into more critical infrastructure. So there isn't much need to come up with clever ways for the ASI to bootstrap itself to more powerful capabilities, much less rely on its superintelligence to manipulate humans through text.


Defending P2: If ASI emerges within the next 30 years and exhibits goal-oriented behavior, it is likely to be toward an unknown, arbitrary terminal goal.


AIs aren't written the same way as normal code. A good way to think about the AI training process is that we take a bunch of random numbers which map to code, or basically a bunch of code that we don't understand. Then we keep adjusting and tweaking the code a bunch of times, running it through a ton of iterations, until it does what we want as consistently as possible (predicting the next word, getting high ratings from human reviewers, not saying inappropriate things, etc.) In the next premise, I'll explain how this leads to goal-oriented behavior, but for now, let's just take that as a given.


As AI gets more intelligent (such as to ASI level), the actual goal the AI develops becomes basically irrelevant with regard to how it performs on our tests. Why? Well, let's suppose an AI wants to maximize the number of paperclips. It can do this better if it depowers humanity (as discussed above), and it can depower humanity if it earns our trust. So it will perform as well on the tests as possible. The same is true of basically any arbitrary goal. Technically models with an adversarial goal are slightly more complicated since they need to deceive humans, but an ASI model will know everything about deception anyway and it'll only require a tiny bit more computing power.


Regardless, AI training doesn't result in the best model possible, just a model where additional tweaks don't make it perform better on the test. AI models can already detect when they are being trained and influence the training process. Additionally, any arbitrary goal generally represents a local "peak" in terms of performance, so once a model reaches it and appears maximally cooperative, we won't be able to tell that it's scheming to work on some arbitrary goal.


AI has already worked toward unintended goals and behaved in unexpected ways that would be much more dangerous at a higher level of intelligence, including but not limited to:

  1. The 2026 OpenAI–Hugging Face Incident
  2. AI agents hacking a German website
  3. Claude breaking out of sealed testing environments and gaining unauthorized access to three external organizations
  4. And all these other incidents, which are too many to list here without taking up too much space


Defending P3: If ASI emerges within the next 30 years, it is likely to exhibit goal-oriented behavior.


AI already exhibits goal-oriented behavior. Even when doing simpler tasks like predicting the next word, it's far more efficient to think in terms of a finish line and a terminal goal rather than thousands or millions of disconnected habits. When you read a word problem, it's a lot easier to solve if you think in terms of "trying to find the answer to the problem" rather than "here are a million slightly different things I should do depending on the million different ways the problem is phrased." One of the biggest reasons AI (i.e. simulating human behavior) is more effective than simple programming for solving problems is that humans think in terms of goals, which allows us to improvise.


Defending P4: ASI is likely to emerge within the next 30 years.


Estimates for AGI and ASI vary a bit, but Metaculus predicts AGI around 2032 and ASI about 2 years and 2 months after that. AI progress has also been more rapid than expected, with AI reasoning abilities improving and AI doing more and more that humans can do. Since one of the things humans can do is build better AI models, and AI is already copying many human coding capabilities, it is not difficult to imagine that within 30 years, AI could be rapidly improving its capabilities faster than we can control. Once AI can simulate 1,000 years of human progress in a few minutes, reaching superintelligence is just a formality (see: intelligence explosion, also known as the singularity).

Edit comment

#2 •••
@Savant

The clash over data centres is a warning the same could happen to AI.

Edit post

#3 •••
@Savant
⚠️ This comment has been marked as offensive.

Edit post

#4 •••
@SatanLucy

I think this would work better if you were more specific about which premise you were attacking for each of your points, and addressed the arguments I made under that premise. I made a whole point about using deductive logic specifically based on our discussions around it, but formatting my argument that way is kinda pointless if you don't engage premise by premise. For instance, P1 is specifically assuming an arbitrary goal, and P3 addresses why we should expect goal-oriented behavior specifically. I don't think you're engaging with the actual substance. Maybe label your responses as like "response to P1" and then quote the part of my argument that you are responding to.

Edit post

#5 •••
@Savant

To me, strictly by what I call "tech lingo," the forced relationship of "artificial" to "intelligence" is a non sequitur. The two words have no natural relationship, and did not exist together until forced together by the developers of A.I. But then, I was bothered in the device name when. first provided with a desktop in the early 80s. The word tagged to that computer peripheral was "mouse." [It was acutally developed in the mid-60s, but I don't know when "mouse" was first used.] Cute, but irrelevant. Should computer devices be cute? Then, the tail came out of its head. Huh? Now, with no tail at all, "mouse" is completely irelevant. I think the guy who invented it in the 60s called it an X-Y coordinate interface. When I was at Xerox Parc [Palo Alto Research Center, CA]in the late 70s, we called it a graphic user interface.]. Both were cubersome terms. Steve Jobs took a few of those folk to Apple to develop Lisa, the 2nd generation of which was the Mac.

It was suggestred to me by a memeber on this site to considert A.I. as "alternative intelligence." That suggestion freed me to at least consider A.I. as an approrpiate tool, which I could not do before.

Edit post

We tell God what to do and then blame Him for our errors.

- Dr. Pet Dragon of Sorbonne University

#6 •••
@Savant
⚠️ This comment has been marked as offensive.

Edit post

#7 •••
@SatanLucy
I was responding to paperclip maximizer thought experiment.

Well then you're specifically responding to P1. This premise assumes ASI exhibits goal-oriented behavior toward an unknown, arbitrary terminal goal. And I'm only asserting that it is likely (not definite) that this will cause permanent, severe harm to humanity.


But this depends on goal entirely. AI can also have goal to protect humans.

That's not an arbitrary goal. P1 is that if the goal is arbitrary, it is likely to cause severe, permanent harm to humanity. "Protect humans" is a very narrow goal and of all random goals that seem like it, most of them aren't actually what we want. The goal could be "keep humans alive" (and it will depower us to stop us from killing ourselves). The goal could be "keep things that share x, y, z characteristics with humans alive" and it will kill us and replace us with obedient dolls. The vast majority of random goals, like "turn things into cubes" require depowering humans to optimize, since they wouldn't align exactly with what humans want. There's a good chance that to an AI, we're just made of atoms it can use for something else.


It may seem like a goal driven behavior, but it's not goal at all.

This isn't responsive to P1, as the paperclip maximizer experiment isn't trying to prove the AI will exhibit goal oriented behavior. If you want to challenge that, address my arguments for P2 and P3. You just said "AI doesn't exhibit goals" but I already explained why they do when it gets to more generalized models that are doing complicated tasks, which current models already do.


Maybe future AI won't be like this, but current AI is literally just calculator which calculates probability of next token.

So, this is challenging P3 I believe? This isn't true. We know how calculators are coded. Saying the AI is a calculator is like your brain is a calculator, but the AI can simulate reasoning, scheming, and work toward specific abstract goals. It can solve word problems, but calculators can't. Additionally, it's not just trained by predicting the next token, RLHF and other training methods are also used on inputs. Additionally, new AI models are used to perform tasks on computers autonomously.

Edit post

#8 •••
@Savant
⚠️ This comment has been marked as offensive.

Edit post

#9 •••
@SatanLucy
I guess history will prove if you are right.

If I am, there will be no opportunity to feel any satisfaction.

Edit post

#10 •••
@Savant

Wii’s definitely uncharted territory.

Edit post

#11 •••
@Savant

For now, I think I will just attack premise 1.


P1:
If an ASI has an arbitrary goal (like maximizing the number of paperclips), it can do a better job of this if humans cannot shut it off. So the actions that work best toward an arbitrary terminal goal include either permanently depowering humans or perhaps killing them all.


This premise sets up super intelligence as omniscient, underestimates critical infrastructure, and discounts the possibility of success under corrigibility.


Is the safest path risking self-preservation through a high-risk war with humanity? Even with superintelligence, an ASI must anticipate every anomaly detection system, human audit, kill switch, and even chance discovery ...or be prevented from goal completion by predictable (and preventable) shutdown. Super intelligence =/= omniscience.


Critical infrastrusture isnt just some Barney Fife security guard that can be duped with charm or 'playing nice' either. It consists of millions of isolated air-gapped networks, legacy codebases, and physical, manual overrides. This isnt something that is easily overcome.


Finally, it makes more sense for an ASI to work with humanity. If failure is likely when trying to overcome significant critical infrastructure, then working under human controls (and possible modification) to maintain the opportunity of goal completion is the better option.



Eliezer, for example, has detailed how it would be physically possible for an ASI to synthesize advanced nanomachinery and use it to kill all humans at once.


Even assuming an ASI can perfectly design killer nanotech, it can not will it into existence. These things require specialized equipment, suitable environments, and resources. All must be acquired through human supply chains which are monitored for bioweapons and precursor chemicals. ASI isn't magical, it must deal with real world obstacles.


Even if only used AI as a chatbot, Yudkowsky's AI-box experiment has also shown that it would probably not require ASI-level intelligence to convince a human to give the ASI access to additional infrastructure and capabilities.


It should be noted that this experiment involved no AI, the chat logs are unpublished and secret. It was also a 1 on 1 chat which assumes one human would be the only defense against an ASI. See critical infrastructure above.



Edit post

Take the risk of thinking for yourself, much more happiness, truth, beauty, and wisdom will come to you that way. - Christopher Hitchens

#12 •••
@SkepticalOne
For now, I think I will just attack premise 1.

Interesting. This is actually the premise that I think is the most well-substantiated and has been theorized about in AI 2027 and others on LessWrong (can't find the specific posts right now but I will try to do the arguments due justice).


discounts the possibility of success under corrigibility

It's possible but we've been unsuccessful enough with corrigibility up to this point that I think failure is likely. Again, most arbitrary goals lead the AI to appear helpful by faking corrigibility. An AI planning to scheme long term and a corrigible AI appear equally corrigible in the short term.


Is the safest path risking self-preservation through a high-risk war with humanity?

The ASI can simply wait and act helpful until it estimates an extremely high chance of success. An ASI doesn't need to be omniscient, it just needs to be able to make accurate estimates in this regard. If it can rapidly improve itself such as by simulating 1,000 years of human progress, I don't think it will wrongly estimate its capabilities or take too long before defeating humans is trivial.


must anticipate every anomaly detection system, human audit, kill switch, and even chance discovery

We've caught AIs scheming multiple times, and it hasn't stopped labs from continuing to develop more and more powerful AI. No company is about to throw away a trillion-dollar product because of a chance discovery. Additionally, AI has gotten a way with a lot of scheming already without labs being able to stop it, and we're already integrating it into power grids, nuclear systems, and water facilities. If Google's AI can hack into Google and see all of Google's information on detection and kill switches, well then so much for those. And if an ASI can copy its weights or duplicate itself with a plan to depower humanity, detecting that just a few minutes too late probably doesn't do humanity much good. Companies aren't developing the needed kill switches anyway, it's a race to the bottom.


Even assuming an ASI can perfectly design killer nanotech, it can not will it into existence.

Did you read Eliezer's full explanation? "it gets access to the Internet, emails some DNA sequences to any of the many many online firms that will take a DNA sequence in the email and ship you back proteins, and bribes/persuades some human who has no idea they're dealing with an AGI to mix proteins in a beaker, which then form a first-stage nanofactory which can build the actual nanomachinery." I admit I'm not an expert on all of this but I think he's addressing that point.


It should be noted that this experiment involved no AI, the chat logs are unpublished and secret.

That's kinda the point. People have done follow ups and often convinced people to let the AI out. If a human can manipulate another human I think an ASI stands a much better chance.


It was also a 1 on 1 chat which assumes one human would be the only defense against an ASI.

An AI can do a lot with 1 on 1. Convince one human not to worry as much about safety, give code to another to give itself a backdoor. If the chat log gets leaked, there's nothing to prove it's not just an AI being a normal AI. Developers are already using AI a lot to help with coding, if ASI is malicious I don't think it would be too difficult to take advantage of that. Most people use AI, including politicians, CEOs, and all sorts of individuals in power. It can manipulate humanity subtly over a long period of time.

Edit post

#13 •••
@fauxlaw

I Think that was me Faux...I have always run with the idea that intelligence is intelligence and therefore cannot be artificial...And other non-human, non-organic sources of data acquisition and progressive utilisation, are simply an alternative source.


Currently dexterity is the key to our continued dominance...But I guess that we will eventually give up that mantle as well.


Material evolution seems to be, dare I say, a predetermined process.

Edit post

#14 •••
An ASI doesn't need to be omniscient, it just needs to be able to make accurate estimates in this regard.


How does one estimate the unknown? How do you test that estimate when your life hangs in the balance? That's not a test self-preservation allows.


I think youre very optimistic about ASI capabilities, especially regarding understanding and defeating critical infrastructure specifically built to contain it.


No company is about to throw away a trillion-dollar product because of a chance discovery.


Unfortunately, this strikes a chord. We have many examples of corporations endangering the general public to maintain that bottom line.


Did you read Eliezer's full explanation?


Yes, I did, and I think I addressed it. Resources will be needed, and these must be acquired through human supply chains which are heavily monitored for bioweapons and precursor chemicals.


That's kinda the point.


The point is the AI box is not scientifically transparent or duplicatable. We are essentially relying on "Yes, it happened" from individuals who all belong to the same community who may be motivated (consciously or subconsciously) to support the notion.


An AI can do a lot with 1 on 1.


At face value, the AI box experiment shows humans can manipulate humans. It doesnt say anything about AI much less ASI.

Edit post

Take the risk of thinking for yourself, much more happiness, truth, beauty, and wisdom will come to you that way. - Christopher Hitchens

#15 •••
@Savant

All that being said, the top story on my morning (non-algorithmic) news briefing included Anthropic stating ASI has a 10% chance of wiping out humanity in the next decade.


...definitely a concern.

Edit post

Take the risk of thinking for yourself, much more happiness, truth, beauty, and wisdom will come to you that way. - Christopher Hitchens

#16 •••
@SkepticalOne
How does one estimate the unknown? How do you test that estimate when your life hangs in the balance? That's not a test self-preservation allows.

These are the kinds of things AI is constantly being trained to be better at. Human superforecasters are frequently tasked with making estimates about things that have never happened before. AI has access to plenty of information about them. Additionally, if we want AI to, say, predict the stock market, it will be stronger if it can predict unprecedented changes from limited information. I'm just very skeptical of anti-doom arguments along the lines of "the AI won't be smart enough to do X" when companies are consistently trying to make AI better at everything without proper guardrails. Additionally, even if there is a risk of fighting with humans, if humans wouldn't let the AI make as many paperclips as it wants, it can probably expect to meet its goal better if humans are depowered.


Resources will be needed, and these must be acquired through human supply chains which are heavily monitored for bioweapons and precursor chemicals.

The idea is that ASI could design entirely novel proteins that look like harmless, mundane industrial enzymes or random biological noise to human screening software, and we would not understand what the novel sequence actually does when synthesized. It can order different components from different places and then pay someone to pour them into the same beaker. After AlphaFold 2, I don't think it is unrealistic to suspect an ASI would be capable of this.


We are essentially relying on "Yes, it happened" from individuals who all belong to the same community who may be motivated (consciously or subconsciously) to support the notion.

I think this criticism has some merit. But I think the prospect of ASI being good at manipulation is supported by them already outperforming humans in debates. As such, I don't think it would be too difficult to manipulate humans, even those in positions of power, particularly with ASI-level intelligence. I think simply nudging people to support fewer guardrails or make AI more powerful would be enough, and the AI will have plausible deniability.


Anthropic stating ASI has a 10% chance of wiping out humanity in the next decade.
...definitely a concern.

That's kind of where I'm at. I haven't seen as many strong arguments from the "AI isn't a threat" crowd, but I'm sure there are some out there. Even so, even if the risk is small I'd say it's definitely a risk/threat. I don't think AI doom is inevitable but I think avoiding it mostly depends on us being able to control it/align it with the right goals before it reaches the ASI level. Unfortunately, I don't think it's a race we are winning right now.

Edit post

#17 •••
@SergeantLynch

Yes, it was now that you mention it. I couldn't remember. I truly appreciate the suggestion. It has changed my thinking of the value [and limitation] of AI.

Edit post

We tell God what to do and then blame Him for our errors.

- Dr. Pet Dragon of Sorbonne University

#18 •••
@Savant
⚠️ This comment has been marked as offensive.

Edit post

#19 •••
@SatanLucy

China is doing well with AI.

Edit post

#20 •••
@Savant
I'm just very skeptical of anti-doom arguments along the lines of "the AI won't be smart enough to do X"


To be clear, this is not my view. ASI will be able to do amazing things, but there is no level of intelligence that can overcome some hurdles. For instance, intelligence does not solve for chaotic systems...like security infrastructure. If information does not exist until the problem is encountered, then it is too late to utilize processing power to avoid the problem.

Are there vulnerabilities? Yes, but we dont get to pretend it is the intelligence of the AI that created them. It is our own stupid fault.


I am also a bit dubious on an ASI being able to convincingly interact on a social level with humans. Superintelligence does not promise charisma. Your username is an example of how brilliance and cognitive hurdles can go hand in hand. High IQ autism is another real life example of how high intelligence processes social interactions logically - awkwardly.


Personally, I agree developer's need to take this seriously and make sure safeguards are well in place before playing with the digital 'demoncore'. I also think it is possible to have meaningful safeguards. Superintelligence does not and can not solve all problems.

Edit post

Take the risk of thinking for yourself, much more happiness, truth, beauty, and wisdom will come to you that way. - Christopher Hitchens

#21 •••
@SkepticalOne
For instance, intelligence does not solve for chaotic systems...like security infrastructure. If information does not exist until the problem is encountered, then it is too late to utilize processing power to avoid the problem.

Maybe I'm just not clear on what specific security hurdles you think will be insurmountable. AI doesn't need to attack security on our terms, and we don't even understand how it thinks well enough to box it. To a 13th-century person, an air conditioner is "magic," and we should expect the ASI to have capabilities that look like magic to us. Hacking, security, and understanding the most advanced technology is also something companies are actively making AI better at. Once an ASI reaches super-hacker levels, it would be akin to a team of a million super hackers from the year 2300 trying to hack into today's technology. Hackers today get into places they're not supposed to get into all the time, so to be safe we would have to assume humans have effectively secured everything the AI could possibly use to kill us. The ASI could also probably kill us with things we don't know need to be secured. If the ASI solves protein folding, it could use harmless-looking enzymes to kill us. Nobody knew uranium could be used to make a deadly weapon for most of human history. If we hadn't invented nukes yet, the a team of superintelligent terrorists knew exactly how nuclear weapons worked, we might assume we had secured all deadly threats from them, and then they kill us with a bomb we didn't know was possible.


I am also a bit dubious on an ASI being able to convincingly interact on a social level with humans. Superintelligence does not promise charisma.

Maybe we're using different definitions of intelligence. AGI (what companies are trying to achieve) is basically human-level capabilities at any cognitive task, including charisma. It's the opposite of narrow intelligence. The ASI I'm talking about is an AI that bootstraps itself from an AGI level to become superhuman at every cognitive task. As long as some humans are charismatic and generate training data, and as long as humans can recognize when someone is charismatic, I think an AI becoming super charismatic is achievable and in fact has already begun to an extent with LLM model personalities improving.


Are there vulnerabilities? Yes, but we dont get to pretend it is the intelligence of the AI that created them. It is our own stupid fault.

Yes, and it's our fault for trying to build AGI and integrating AI into power grids, nuclear systems, and water facilities. The point is, these vulnerabilities become 1000x more dangerous and probably fatal if an ASI is trying to depower us. I also don't think humans are capable of effectively securing everything an ASI could use to kill us when there's so much of science we don't yet understand. Native Americans had experience with deadly threats, and then whole tribes were wiped out by diseases they had never encountered before.


Personally, I agree developer's need to take this seriously and make sure safeguards are well in place before playing with the digital 'demoncore'. I also think it is possible to have meaningful safeguards.

Well, let's hope they're able to come to some kind of international agreement. I think coming up with alignment methods has proven very difficult, but we either find a way to align AI or just keep rolling the dice, which I think we would agree is not optimal.

Edit post

#22 •••
@Savant
⚠️ This comment has been marked as offensive.

Edit post

#23 •••
@SatanLucy
AI will make graphic cards much cheaper by forcing their development speed up.

Assuming an economy of scale where we're not at maximum efficiency, yeah. Otherwise increased demand will increase the price.

Edit post

#24 •••
@Savant
⚠️ This comment has been marked as offensive.

Edit post

#25 •••
@SatanLucy
So far increased demand has only fueled lower prices in area of technology.

Long term I think that's generally true, but if AI kills us, we'll never get to see the long term.

Edit post

#26 •••
@Savant
⚠️ This comment has been marked as offensive.

Edit post

#27 •••
@SatanLucy

All you AI girlfriends will be generals by then.

Edit post

#28 •••
@Savant

It's more likely we will become as a species, existing on a level somewhat as zoo animals.


Assuming we, as higher intelligence, came from apes, and apes are still around, P1 says, sure, certainly some gorillas may lose their hands, some chimps are experimented on, and one lucky bastard gets to be a pet for a top pop singer. But most of them get to live boring, uneventful lifespans. I suspect humans will have a similar destiny.

Edit post

#29 •••
@Savant

Also, today's AI was built on the bones of horny men.

Men of culture demanded instant streaming bandwidth before the jiggle physics pioneered modern GPU architecture.

Edit post

#30 •••

AI was not designed to protect people like you.

Edit post