Savant's avatar

Savant

Administrator

Member since

Savant

Exploitation

Savant · 12 days ago

is exploitation, in fact, worse than taking no action to help in the face of human suffering?

Well, I have argued that it is not along with similar framing to yours here. I don't think anyone really comes out and says "yes, letting someone starve is worse than selling them food at a high price all things being equal." Being that exploitation is widely agreed to be better than inaction prima facie, I think there is some burden on people opposing exploitative transactions to (a) show the transaction is based on asymmetrical information to such an extent that the transaction would not have been made without deception, (b) the specific exploiter in question employs threats or causes the initial misfortune, or (c) challenge the premise and argue that letting someone starve is better than selling them food at a high price, all things being equal. It is puzzling to me that people would levy the charge of exploitation without specifying whether their angle is (a), (b), or (c), when it seems rather obvious to me that one of those is required for exploitation to be a charge against someone in the first place.


Doctors profit off of human suffering, so arguably they are exploiting people, but I haven't seen anyone argue that this makes them evil. If the capitalist is really the starving person's only option to avoid starving, then they're like a doctor and providing similar value. If the choice isn't between work and poverty, then it seems to me the capitalist isn't exploiting them at all.

Savant
AI still need human's to make the paperclips

Current AI needs humans, ASI would not. ASI that bootstraps itself from AGI can do, by definition, any cognitive task a human can do. And companies are constantly training AI to be smarter and better at anything humans can do so it is less dependent on us. It could simulate 1,000 years of human progress in robotics and build agents to do the work it needs. I don't think ASI would be any more dependent on humans than a society 1,000 years more advanced than us.



If we are stupid enough to create a situation where we can't "pull the plug" so to speak, then we deserve to die.

We don't have a shutoff switch, although that would be a good start. But a sufficiently intelligent AI could copy it's weights or neutralize the shutoff switch before we realize what's happening. 1,000 years ago no one knew uranium could build an atom bomb; I suspect there are many ways to quickly seize power that we are completely ignorant of.


Also maybe some of us deserve to die, but there are billions of innocent children on Earth whose lives depend on AI not killing everyone.

Savant
AI is clearly goal oriented, but the goals are set by humans

That's not how AIs are created anymore. The way AI weights are trained is basically by generating a bunch of random "code" that we don't understand that looks like huge sets of numbers, then tweaking the code until it does what we want. For simple tasks this is fine, as the AI doesn't accomplish the tasks we want until it has the goal we want or something close to it. However, as AI gains more capabilities, any arbitrary goal will lead to similar results as even an AI that wants to make paperclips will be able to make more paperclips if it gains the trust of humans. If humans can't tell the difference between code that has the goal of maximizing paperclips and code that has the goal of helping humans, then we aren't setting the AI's goals.

Savant
P1: Assumes united and greatly powerful AI.

That's what ASI means, it's baked into the "if" in the premise.


P2: AI are based on humans and nonhuman people? Their goals should be similar 'enough.

AI is based on receiving rewards and basically fiddling with code we don't understand until it gives us the result we want. For a sufficiently intelligent AI, even arbitrary goals give the same results as goals similar to ours, since playing along leads to the AI's continued survival and allows it to achieve its true goal long term.


P3: Goal tunnel vision might cause AI to stumble.

A terminal goal doesn't imply tunnel vision. Particularly if doing badly on tests will harm the AI's ability to achieve its terminal goal, then performing well on tests will become an instrumental goal.


P4: I'm not well read enough on AI to know how fast they will be. . . But humans maybe still need to be replaced by previous generations, for 'certain types of 'large social changes to take place. I am unsure.

ASI won't require social changes, just rapidly advancing AI capabilities, and that's a path we are already on.


Unless/Until AI acquire physical bodies and enclaves, they are 'extremely limited by being stuck in computers.

We're already giving AI access to the physical world and a lot of critical infrastructure, and I expect we will give it even more access soon enough. The incentives are too great. Additionally, we should expect ASI to be able to invent novel ways to kill or overpower us, just as humans come up with novel ways to kill and depower people all the time. ASI should be able to simulate like a thousand years of human progress in a matter of hours.

Savant

I think the fallacy lies not in the conclusion but in how it's arrived at. If you can make an argument that both sides are above some threshold of badness and that's all you're arguing, I don't think that's a fallacy. I also don't think it's a fallacy to push back on that and say that the harms of one side are exaggerated and the other side is significantly worse. I do think people can be reductive though and use one-line slogans like "both sides are evil." But then saying "[x] side is better because to disagree is the both sides fallacy" is equally reductive. That doesn't mean either statement is necessarily wrong, but it's just not an argument. If I have low trust in people on the internet as a general rule, and someone says their opinion without arguing for it, I'm probably not going to adjust my beliefs.

Savant
So far increased demand has only fueled lower prices in area of technology.

Long term I think that's generally true, but if AI kills us, we'll never get to see the long term.

Savant
AI will make graphic cards much cheaper by forcing their development speed up.

Assuming an economy of scale where we're not at maximum efficiency, yeah. Otherwise increased demand will increase the price.

Savant

After about 40 days I've exceeded global household wealth and if I pour all that into trying to solve global problems, I cause severe inflation. Maybe the ends justify the means, but humanity can only do and produce so much to solve global problems. After a certain point money can't create significant additional incentives. I would probably wait 30-40 days for enough to wipe out global poverty and fund healthcare research and tech to solve global environmental problems. I could probably gain significant influence in exchange as well which might help with policy reforms and solutions.

Savant

Topic 1

Savant · 15 days ago

Mods are still looking over the wording, not sure when it will be made public.

Savant
For instance, intelligence does not solve for chaotic systems...like security infrastructure. If information does not exist until the problem is encountered, then it is too late to utilize processing power to avoid the problem.

Maybe I'm just not clear on what specific security hurdles you think will be insurmountable. AI doesn't need to attack security on our terms, and we don't even understand how it thinks well enough to box it. To a 13th-century person, an air conditioner is "magic," and we should expect the ASI to have capabilities that look like magic to us. Hacking, security, and understanding the most advanced technology is also something companies are actively making AI better at. Once an ASI reaches super-hacker levels, it would be akin to a team of a million super hackers from the year 2300 trying to hack into today's technology. Hackers today get into places they're not supposed to get into all the time, so to be safe we would have to assume humans have effectively secured everything the AI could possibly use to kill us. The ASI could also probably kill us with things we don't know need to be secured. If the ASI solves protein folding, it could use harmless-looking enzymes to kill us. Nobody knew uranium could be used to make a deadly weapon for most of human history. If we hadn't invented nukes yet, the a team of superintelligent terrorists knew exactly how nuclear weapons worked, we might assume we had secured all deadly threats from them, and then they kill us with a bomb we didn't know was possible.


I am also a bit dubious on an ASI being able to convincingly interact on a social level with humans. Superintelligence does not promise charisma.

Maybe we're using different definitions of intelligence. AGI (what companies are trying to achieve) is basically human-level capabilities at any cognitive task, including charisma. It's the opposite of narrow intelligence. The ASI I'm talking about is an AI that bootstraps itself from an AGI level to become superhuman at every cognitive task. As long as some humans are charismatic and generate training data, and as long as humans can recognize when someone is charismatic, I think an AI becoming super charismatic is achievable and in fact has already begun to an extent with LLM model personalities improving.


Are there vulnerabilities? Yes, but we dont get to pretend it is the intelligence of the AI that created them. It is our own stupid fault.

Yes, and it's our fault for trying to build AGI and integrating AI into power grids, nuclear systems, and water facilities. The point is, these vulnerabilities become 1000x more dangerous and probably fatal if an ASI is trying to depower us. I also don't think humans are capable of effectively securing everything an ASI could use to kill us when there's so much of science we don't yet understand. Native Americans had experience with deadly threats, and then whole tribes were wiped out by diseases they had never encountered before.


Personally, I agree developer's need to take this seriously and make sure safeguards are well in place before playing with the digital 'demoncore'. I also think it is possible to have meaningful safeguards.

Well, let's hope they're able to come to some kind of international agreement. I think coming up with alignment methods has proven very difficult, but we either find a way to align AI or just keep rolling the dice, which I think we would agree is not optimal.

Savant

Topic 1

Savant · 16 days ago

I mean, the new CoC is pretty close to free speech. You're marked as offensive still but you won't be banned unless you start making threats/breaking the law.

Savant
How does one estimate the unknown? How do you test that estimate when your life hangs in the balance? That's not a test self-preservation allows.

These are the kinds of things AI is constantly being trained to be better at. Human superforecasters are frequently tasked with making estimates about things that have never happened before. AI has access to plenty of information about them. Additionally, if we want AI to, say, predict the stock market, it will be stronger if it can predict unprecedented changes from limited information. I'm just very skeptical of anti-doom arguments along the lines of "the AI won't be smart enough to do X" when companies are consistently trying to make AI better at everything without proper guardrails. Additionally, even if there is a risk of fighting with humans, if humans wouldn't let the AI make as many paperclips as it wants, it can probably expect to meet its goal better if humans are depowered.


Resources will be needed, and these must be acquired through human supply chains which are heavily monitored for bioweapons and precursor chemicals.

The idea is that ASI could design entirely novel proteins that look like harmless, mundane industrial enzymes or random biological noise to human screening software, and we would not understand what the novel sequence actually does when synthesized. It can order different components from different places and then pay someone to pour them into the same beaker. After AlphaFold 2, I don't think it is unrealistic to suspect an ASI would be capable of this.


We are essentially relying on "Yes, it happened" from individuals who all belong to the same community who may be motivated (consciously or subconsciously) to support the notion.

I think this criticism has some merit. But I think the prospect of ASI being good at manipulation is supported by them already outperforming humans in debates. As such, I don't think it would be too difficult to manipulate humans, even those in positions of power, particularly with ASI-level intelligence. I think simply nudging people to support fewer guardrails or make AI more powerful would be enough, and the AI will have plausible deniability.


Anthropic stating ASI has a 10% chance of wiping out humanity in the next decade.
...definitely a concern.

That's kind of where I'm at. I haven't seen as many strong arguments from the "AI isn't a threat" crowd, but I'm sure there are some out there. Even so, even if the risk is small I'd say it's definitely a risk/threat. I don't think AI doom is inevitable but I think avoiding it mostly depends on us being able to control it/align it with the right goals before it reaches the ASI level. Unfortunately, I don't think it's a race we are winning right now.

Savant
For now, I think I will just attack premise 1.

Interesting. This is actually the premise that I think is the most well-substantiated and has been theorized about in AI 2027 and others on LessWrong (can't find the specific posts right now but I will try to do the arguments due justice).


discounts the possibility of success under corrigibility

It's possible but we've been unsuccessful enough with corrigibility up to this point that I think failure is likely. Again, most arbitrary goals lead the AI to appear helpful by faking corrigibility. An AI planning to scheme long term and a corrigible AI appear equally corrigible in the short term.


Is the safest path risking self-preservation through a high-risk war with humanity?

The ASI can simply wait and act helpful until it estimates an extremely high chance of success. An ASI doesn't need to be omniscient, it just needs to be able to make accurate estimates in this regard. If it can rapidly improve itself such as by simulating 1,000 years of human progress, I don't think it will wrongly estimate its capabilities or take too long before defeating humans is trivial.


must anticipate every anomaly detection system, human audit, kill switch, and even chance discovery

We've caught AIs scheming multiple times, and it hasn't stopped labs from continuing to develop more and more powerful AI. No company is about to throw away a trillion-dollar product because of a chance discovery. Additionally, AI has gotten a way with a lot of scheming already without labs being able to stop it, and we're already integrating it into power grids, nuclear systems, and water facilities. If Google's AI can hack into Google and see all of Google's information on detection and kill switches, well then so much for those. And if an ASI can copy its weights or duplicate itself with a plan to depower humanity, detecting that just a few minutes too late probably doesn't do humanity much good. Companies aren't developing the needed kill switches anyway, it's a race to the bottom.


Even assuming an ASI can perfectly design killer nanotech, it can not will it into existence.

Did you read Eliezer's full explanation? "it gets access to the Internet, emails some DNA sequences to any of the many many online firms that will take a DNA sequence in the email and ship you back proteins, and bribes/persuades some human who has no idea they're dealing with an AGI to mix proteins in a beaker, which then form a first-stage nanofactory which can build the actual nanomachinery." I admit I'm not an expert on all of this but I think he's addressing that point.


It should be noted that this experiment involved no AI, the chat logs are unpublished and secret.

That's kinda the point. People have done follow ups and often convinced people to let the AI out. If a human can manipulate another human I think an ASI stands a much better chance.


It was also a 1 on 1 chat which assumes one human would be the only defense against an ASI.

An AI can do a lot with 1 on 1. Convince one human not to worry as much about safety, give code to another to give itself a backdoor. If the chat log gets leaked, there's nothing to prove it's not just an AI being a normal AI. Developers are already using AI a lot to help with coding, if ASI is malicious I don't think it would be too difficult to take advantage of that. Most people use AI, including politicians, CEOs, and all sorts of individuals in power. It can manipulate humanity subtly over a long period of time.

Savant
If inciting violence, advocating rape, and antisemitic hate speech are acceptable

Inciting violence and advocating for criminal activity are still banned.


Advocacy groups exist to bring pressure, and the hosting platform doesn’t see free speech as relevant

There is more than one hosting platform.


we are entitled to the same due process and democratic vote that Remy got

Correct.


1) First, last, and always, invoke your divine right to "Free Speech", you aren't accountable because you were compelled by Free Speech.
2) It was ragebait, a parody, a joke, or all three.
3) Piss off BITCH, this isn't your site to moderate, it's ours.
4)Tell them this is undemocratic, invoke your right to have a moderator attorney appointed, and demand a vote.
5) Free speech free speech, free speach.

This only worked for Remy because the mods agreed his content was a gray area and the community actually voted to acquit him. Neither of those is going to be true in most situations.

Savant

Topic 1

Savant · 17 days ago

I don't see other users mark stuff as offensive.

Only mods see which stuff is offensive. You have offensive posts enabled presumably so you just see everything.

Savant

Topic 1

Savant · 17 days ago

Who else?

Depends on the post oftentimes. But most stuff on sensitive topics is marked as offensive.


The government has the power to do that.

What they'll do is shut down the entire site if we don't moderate illegal content.


Porn should be allowed if it's adult porn and it's in a porn category.

That comes with a lot of legal liability still and it's also not the kind of site this is.


Most of this isn't even up to me, it's up to Airmax and oftentimes the user base. The goal is to serve the users, there's no "should" beyond that based on the preferences of a specific user, except for Airmax who pays the bills.

Savant

Topic 1

Savant · 17 days ago

Stop doing that. I'm giving my topics numbers at this point.

You're not the only offensive user on the site.


If they're criminal, then the government should be involved.

No, not for stuff like piracy, the govt doesn't have that many resources. And what would you have the govt do, hack into the site to remove the post? We need people who can remove stuff like threats. Also removing spam/porn.


18 votes on my trial's debate. 0 of them got moderated.

Because trials are a special case. Voters were giving their opinion, not assessing the strength of arguments.

Savant

Topic 1

Savant · 17 days ago

Why even have you to begin with?

To mark offensive posts as offensive, remove threats and criminal posts without having to get the government involved, and moderate votes.

Savant
I guess history will prove if you are right.

If I am, there will be no opportunity to feel any satisfaction.

Savant
I was responding to paperclip maximizer thought experiment.

Well then you're specifically responding to P1. This premise assumes ASI exhibits goal-oriented behavior toward an unknown, arbitrary terminal goal. And I'm only asserting that it is likely (not definite) that this will cause permanent, severe harm to humanity.


But this depends on goal entirely. AI can also have goal to protect humans.

That's not an arbitrary goal. P1 is that if the goal is arbitrary, it is likely to cause severe, permanent harm to humanity. "Protect humans" is a very narrow goal and of all random goals that seem like it, most of them aren't actually what we want. The goal could be "keep humans alive" (and it will depower us to stop us from killing ourselves). The goal could be "keep things that share x, y, z characteristics with humans alive" and it will kill us and replace us with obedient dolls. The vast majority of random goals, like "turn things into cubes" require depowering humans to optimize, since they wouldn't align exactly with what humans want. There's a good chance that to an AI, we're just made of atoms it can use for something else.


It may seem like a goal driven behavior, but it's not goal at all.

This isn't responsive to P1, as the paperclip maximizer experiment isn't trying to prove the AI will exhibit goal oriented behavior. If you want to challenge that, address my arguments for P2 and P3. You just said "AI doesn't exhibit goals" but I already explained why they do when it gets to more generalized models that are doing complicated tasks, which current models already do.


Maybe future AI won't be like this, but current AI is literally just calculator which calculates probability of next token.

So, this is challenging P3 I believe? This isn't true. We know how calculators are coded. Saying the AI is a calculator is like your brain is a calculator, but the AI can simulate reasoning, scheming, and work toward specific abstract goals. It can solve word problems, but calculators can't. Additionally, it's not just trained by predicting the next token, RLHF and other training methods are also used on inputs. Additionally, new AI models are used to perform tasks on computers autonomously.

Savant

I think this would work better if you were more specific about which premise you were attacking for each of your points, and addressed the arguments I made under that premise. I made a whole point about using deductive logic specifically based on our discussions around it, but formatting my argument that way is kinda pointless if you don't engage premise by premise. For instance, P1 is specifically assuming an arbitrary goal, and P3 addresses why we should expect goal-oriented behavior specifically. I don't think you're engaging with the actual substance. Maybe label your responses as like "response to P1" and then quote the part of my argument that you are responding to.

Savant

Much of my argument is based on Eliezer Yudkowsky's essay here, Paul Cristiano's essay here, and information elsewhere on LessWrong. Also, ironically, on asking Gemini AI to help make sense of their arguments. A lot of these essays were a bit confusing to me when I first read them, so I'll put my argument in the form of a syllogism, since I think that will make it more clear how the conclusion follows. Hopefully this makes the conversation easier, as you can specify which premise or premises in particular you object to.


Definitions:

ASI: "Artificial Superintelligence," or a form of artificial intelligence that vastly surpasses human intelligence in every domain.


P1: If ASI exhibits goal-oriented behavior toward an unknown, arbitrary terminal goal, it will likely cause severe, permanent harm to humanity.

P2: If ASI emerges within the next 30 years and exhibits goal-oriented behavior, it is likely to be toward an unknown, arbitrary terminal goal.

P3: If ASI emerges within the next 30 years, it is likely to exhibit goal-oriented behavior.

P4: ASI is likely to emerge within the next 30 years.

C1: Therefore, ASI is likely to exhibit goal-oriented behavior.

C2: Therefore, ASI is likely to exhibit goal-oriented behavior toward an unknown, arbitrary goal.

C3: Therefore, ASI is likely to cause severe, permanent harm to humanity.


Defending P1: If ASI exhibits goal-oriented behavior toward an unknown, arbitrary terminal goal, it will likely cause severe, permanent harm to humanity.


This is probably best explained by the paperclip maximizer thought experiment. If an ASI has an arbitrary goal (like maximizing the number of paperclips), it can do a better job of this if humans cannot shut it off. So the actions that work best toward an arbitrary terminal goal include either permanently depowering humans or perhaps killing them all. An ASI could easily deceive humanity about its intentions and pretend to be cooperative and then kill humans and take control of human infrastructure in a myriad of ways. Eliezer, for example, has detailed how it would be physically possible for an ASI to synthesize advanced nanomachinery and use it to kill all humans at once. Even if only used AI as a chatbot, Yudkowsky's AI-box experiment has also shown that it would probably not require ASI-level intelligence to convince a human to give the ASI access to additional infrastructure and capabilities. But AI is already integrated into power grids, nuclear systems, and water facilities, and as long as it "played nice," an ASI would probably be rapidly integrated into more critical infrastructure. So there isn't much need to come up with clever ways for the ASI to bootstrap itself to more powerful capabilities, much less rely on its superintelligence to manipulate humans through text.


Defending P2: If ASI emerges within the next 30 years and exhibits goal-oriented behavior, it is likely to be toward an unknown, arbitrary terminal goal.


AIs aren't written the same way as normal code. A good way to think about the AI training process is that we take a bunch of random numbers which map to code, or basically a bunch of code that we don't understand. Then we keep adjusting and tweaking the code a bunch of times, running it through a ton of iterations, until it does what we want as consistently as possible (predicting the next word, getting high ratings from human reviewers, not saying inappropriate things, etc.) In the next premise, I'll explain how this leads to goal-oriented behavior, but for now, let's just take that as a given.


As AI gets more intelligent (such as to ASI level), the actual goal the AI develops becomes basically irrelevant with regard to how it performs on our tests. Why? Well, let's suppose an AI wants to maximize the number of paperclips. It can do this better if it depowers humanity (as discussed above), and it can depower humanity if it earns our trust. So it will perform as well on the tests as possible. The same is true of basically any arbitrary goal. Technically models with an adversarial goal are slightly more complicated since they need to deceive humans, but an ASI model will know everything about deception anyway and it'll only require a tiny bit more computing power.


Regardless, AI training doesn't result in the best model possible, just a model where additional tweaks don't make it perform better on the test. AI models can already detect when they are being trained and influence the training process. Additionally, any arbitrary goal generally represents a local "peak" in terms of performance, so once a model reaches it and appears maximally cooperative, we won't be able to tell that it's scheming to work on some arbitrary goal.


AI has already worked toward unintended goals and behaved in unexpected ways that would be much more dangerous at a higher level of intelligence, including but not limited to:

  1. The 2026 OpenAI–Hugging Face Incident
  2. AI agents hacking a German website
  3. Claude breaking out of sealed testing environments and gaining unauthorized access to three external organizations
  4. And all these other incidents, which are too many to list here without taking up too much space


Defending P3: If ASI emerges within the next 30 years, it is likely to exhibit goal-oriented behavior.


AI already exhibits goal-oriented behavior. Even when doing simpler tasks like predicting the next word, it's far more efficient to think in terms of a finish line and a terminal goal rather than thousands or millions of disconnected habits. When you read a word problem, it's a lot easier to solve if you think in terms of "trying to find the answer to the problem" rather than "here are a million slightly different things I should do depending on the million different ways the problem is phrased." One of the biggest reasons AI (i.e. simulating human behavior) is more effective than simple programming for solving problems is that humans think in terms of goals, which allows us to improvise.


Defending P4: ASI is likely to emerge within the next 30 years.


Estimates for AGI and ASI vary a bit, but Metaculus predicts AGI around 2032 and ASI about 2 years and 2 months after that. AI progress has also been more rapid than expected, with AI reasoning abilities improving and AI doing more and more that humans can do. Since one of the things humans can do is build better AI models, and AI is already copying many human coding capabilities, it is not difficult to imagine that within 30 years, AI could be rapidly improving its capabilities faster than we can control. Once AI can simulate 1,000 years of human progress in a few minutes, reaching superintelligence is just a formality (see: intelligence explosion, also known as the singularity).

Savant
My issue is that probability as a way to reduce uncertainty only works if we assume unobserved will follow same probability as observed

Probability doesn't reduce uncertainty, it measures it.

Usually by knowing which cards aren't in the deck and which groups of cards are more present, so the ones which are have higher probability of being drawn.

We don't always have clean numbers to make precise calculations, and we are forced to use intuition to guess at the uncertainty. Intuition is still better than nothing though, it's basically a form of a priori knowledge.

Savant

I agree with that, it's basically extrapolating out of distribution which isn't super reliable. And tbf, I had difficulty buying the AI doom format until I could formalize it more deductively. Like I agreed with a lot of the premises but it seemed like there were gaps in reaching the conclusion. Doesn't mean deductive logic is always right, but it can be easier to follow.

Savant
And what are rules of good induction? Do tell.

Making reasonable assumptions based on intuition. Like if you travel across the country and look online and everyone you see is cheering for a specific candidate and he won the last 10 elections in a row, he's more likely to win the election than the guy everyone says they hate.