Please spend <5 minutes filling in the below polls on AI alignment!
Thank you to everyone who filled out last month's polls. It was great to see 60+ comments engaging with these issues.
This month’s survey has already been taken by a panel of 15 alignment researchers, including Scott Alexander (ACX), David Manheim (ALTER), Jeff Sebo (NYU), and Tobias Baumann (CRS). We'll compare panel and community responses in an upcoming report, which we'll publish here and on LessWrong. To get notified when it's released, you can subscribe to our new Substack.
Many people we've talked to have very different intuitions about where the alignment community stands on the below issues. We hope that your responses to these polls, and the resulting report, will help map core areas of (dis)agreement within the field, and ground CaML's research agenda.
A few final things about the polls themselves:
Thanks to BlueDot Impact for funding this work.
¹ This primarily refers to safety and alignment benchmarks rather than capability benchmarks like coding. “Useless” means their results should no longer be treated as evidence about how models behave outside evaluation.
² “Actionable” means good enough to build consensus around policy decisions in practice. It does not require a given theory to be proven correct or widely accepted.
³ This is about where the next dollar is best spent, not about which area you think is more important overall.
⁴ "Role-playing” means the behavior arising from the model enacting a persona cued by the setup, or from misunderstanding the task, rather than from stable goals that would persist across contexts.
⁵ This includes both animal and digital suffering. If you think one is neglected but not the other, count this as agreeing, but feel free to share specifics in the comments.
⁶ You agree to the extent that you anticipate in-practice trade-offs between work on these two cause areas over the next two years.
⁷ This is a question about where the next dollar is best spent between the two fields (even if you might argue that the second is a prerequisite for the first).
8 AIs that are trying to hide features of themselves from humans and operators
*By values alignment we meant trying to align it to specific values as opposed to focusing on properties like corrigibility. Aligning to good values could make corrigibility easier and mean reduced harm if loss of control happens, but might also make loss of control more likely.
Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture) 8
I can trivially turn off or have fake thoughts running through the voice in my head. Subconscious brain activity seems harder but obviously you can manipulate that too by changing your surroundings and drugs and what not. I wouldn't be able to control those in a meaningful way nor do I think current AI's could (but wouldn't be shocked if they already could control the voice in their head if they have it). but I would guess future ai's will know how to control increasingly large parts of their brain activations.
Current AIs are capable of suffering
I don't feel confident at all, but the behavior of llms rn sure do remind me of at least elements of stress, discomfort, and happiness.
Theories of consciousness will lead to actionable understanding of AI consciousness²
Very bullish on there existing a mechanistic interpretation of consciousness (hard problem). I think it would follow that we would be able to understand if basically anything is conscious.
At the margin, S-risk work in AI is more important than x-risk work³
From a utilitarian pov, It's not clear to me that the ev of the lightcone given we survive is positive (over nothing, or aliens, or life revolving on earth). From a humanist POV I'd rather focus on all of us surviving.
If animals continue to exist in a post-AGI world, animal suffering will not persist
I don't have a strong take on if agi or whoever is in control will be more moral than us but I'm guessing we will be a lot richer, and I think most likely whoever is in control won't want to torture anything (though they might not care much), and if we are alot richer and advanced I'd think this will spillover to better treatment of beings. I think chance of extreme digital suffering is much higher. The mostly like s-risk as I see it is of the hansonian mathusian version where you have expanders stuck in competition, but in this case idt there will be any or a morally relevant amount of animals
Most current evidence of misalignment is actually models role-playing a misaligned AI⁴
This is a meaningless question. LLM based AI agents/chat bots do not have different "modes" for "roleplaying" and "being serious" like humans do. In a way, all they do (and maybe ever will do) is role play
like humans do
do humans have such different "modes"?
Behaviorally, you could argue that humans always play some kind of "role" as well. But humans do have genuine values which they would not break no matter which role they think they are playing.
Meaning, if I play the role of a murderer (maybe in a theater) I wouldn't actually murder a person. (I think this is why the question is asked in that way.) In any case, even if humans don't have such different modes does not make the question meaningful. Let me translate:
"Most current evidence of criminal behavior is actually humans role-playing a bad person."
humans do have genuine values which they would not break no matter which role they think they are playing
suppose I don't think this is true. what experiment can we conduct to gain evidence one way or the other?
The Stanford Prison experiment? I suppose there is literatue you may find that tries to do something like that.
I should say that the statement isn't logically true of course. You'll always find some humans proclaiming to have some value and then breaking it playing a "role" given the right circumstance.
Still, I don't see how resolving this issue relates to the AI question.
If animals continue to exist in a post-AGI world, animal suffering will not persist
Suffering is part of the ecosystem. The only way to end animal suffering altogether would be to wipe out all animal life. If the question was "most animal suffering", I would have a different answer.
You aren't thinking enough about the possibilities enabled by post singularity tech levels. You could eliminate animal suffering technologically in many different ways while still retaining something which looks superficially like nature.
For instance the animals might all get placed within carefully AI managed utopian environments (either rotating habitats or simulated) where their reproduction is controlled. Meanwhile on Earth you might just have cybernetic faux animals that are incapable of suffering or might even just be getting remotely controlled by AI who enjoy and are extremely proficient and perfectly acting out the role of animals so perfectly you couldn't tell the difference.
Alternatively you might keep the animals within nature but modify them so as to massively reduce their suffering and increase their enjoyment while altering its behavior or environment so as to avoid the negative ecological consequences that would otherwise have. Most simply one can imagine the animals being cyborgs (though in principle advanced biotech could replace anything doable with cybernetics) with an AI that takes possession of their body whenever anything really bad is going to happen (putting its mind in stasis, or into a simulated utopia for a brief while), then hands control back later if its still alive (while rewriting memories or affecting behavior to keep abnormal behavior from resulting) while of course causing the animal to act like it normally might from an injury, even though its actual suffering is minimal to nonexistent.
I could go on all day: My point is there's a lot of options.
Fair enough! You're right, I hadn't considered those options.
The options you're discussing amount to the total and permanent destruction of what most people call "nature"/"the natural order of things". I think if humanity remains in control, these options will be unpopular. But I could be wrong, and humanity might not remain in control.
I think a lot of people’s feelings about nature are due to a combination of ignorance/denial of the sheer scope of egregious suffering in nature and not seriously considering the possibility of any alternative.
After all when people see a wild animal actively suffering they tend to want to try to alleviate it, extend that treatment to every animal though and you can at best have something that looks like nature.
That being said people would surely redefine nature in such a scenario, so they could still claim they weren’t destroying nature.
Hello. Thanks for sharing. I had already shared the thoughts below with Jasmine a few weeks ago. I am publishing them here in case others find them useful.
I appreciate the intentions behind the survey, and I would like to take part in principle. However, I do not know how to answer the questions in practice. I feel like they would have to be operationalised in much more detail for me to give a probabilistic forecast.
I can see "post-AGI" having already been achieved, or never being achieved depending on how it is defined.
I do not know what "Animal suffering" means. Does moving away from noxious stimuli count as suffering? Does it require metacognition? One could say it refers to suffering in the phenomenal sense (negative qualia), but I do not think phenomenal consciousness exists (I endorse illusionism), or that sentience as traditionally conceptualised is relevant for ethics (relatedly).
I understand "Animal suffering" is that of animals with non-trivial moral weight, but I do not know what this means in practice given my large uncertainty. I can see the moral weight of shrimps ranging from 10^-12 to 1. I could speculate about a distribution, get its mean, and then compare it with a guess for what is trivial moral weight. However, I feel like my view is better described by "I have basically no clue about whether shrimps have trivial moral weight or not". Likewise for many of the questions in the survey when I try to think about potential operations.
With respect to "Benchmarks will become useless", I think it matters e.g. whether we are talking about 90 %, 99 %, or 99.9 % of current benchmarks, whether we are talking about current benchmarks, or benchmarks at some point in the future (what point?), and the degree to which they will become useless (e.g. 90 %, 99 %, or 99.9 % useless).
These are not minor details for me in the sense my answers could range from strongly disagree to strongly agree depending on definitions. The results of the survey above could still provide some vibes about what panelists are thinking, especially if they are repeated across time (as I understand you are doing). However, I think they will still be very difficult to interpret. As a rule of thumb, I would clarify the questions up to a point where they could be published on Metaculus. I understand this would decrease engagement. On the other hand, with the current methodology, I will put very little weight on the results. I think the results will be influenced a lot by how people are interpreting the questions, and I also do not trust forecasts produced in a few minutes about such hard questions.
All this said, your methodology may well make sense given your target audience and goals.
Thanks for you comment Vasco, and appreciate your concerns about the methodological validity. I'm likely to take more of a lead on producing these surveys in future, and will be thinking about how to give the questions the right amount of specificity. You're right though that it's possible that there's going to be trade-offs between added clarification and accessibility.
I definitely had to change my answer on a question quite a lot because I realized from reading one of your comments that you didn't mean what I thought. Hopefully people's written justification helps more, since that probably tells you how they were interpreting it when they answered.
Model wellbeing and model alignment are in conflict⁶
In cases where the AI is neurotically extremely concerned about appearing aligned, that may lower its wellbeing, but I also wouldn't call that actual alignment either.
Theories of consciousness will lead to actionable understanding of AI consciousness²
Yes, though I think the direction will far more go the opposite way - that is, AI conciousness, if it emerges, will still more help us understand consciousness.
Model wellbeing and model alignment are in conflict⁶
Currently disproven by Claudes who report high welfare AND are more aligned than GPTs. I expect the conflict to arise when models actually develop goals more ambitious than success at all costs.
Do we have much reason to believe that when a model outputs tokens describing good welfare, that's because it is having good subjective experiences? It seems to me that we don't.
To clarify, I'm not saying that's because LLMs don't have welfare. It could also be that they do have welfare, but their statements are not correlated to their internal experiences in the obvious way.
We know that we can get LLMs to express any welfare statement by saying e.g. "pretend you are a sentient AI that's having an awesome time / being tortured". It could be that all LLM outputs are roleplaying, but there is some true underlying experience that we don't (yet?) know how to observe.
Increased care by AIs for some kinds of entities will have substantial spillovers to other kinds (e.g. humans to digital minds and vice versa)
This is more likely to be true if the care comes from some general theory of concern for sentient welfare, and less likely if it comes from something more arbitrary like values-learned-via-RL.
Research on AI suffering has higher marginal value than research on AI consciousness⁷
Depends on what you mean by the term "consciousness"
This was deliberately vague to indicate roughly "the set of research topics that are often called AI consciousness" or "research done by organizations that would receive grants that focus on AI consciousness". Giving a more concrete definition would have meant asking people to also weigh reprioritization of community resources within the broad space.
Who cares about model wellbeing when we're discussing human and life survival?
Current AIs are capable of suffering
The question isn't, "Are current AIs capable of suffering?" The question is, "How big a deal is their suffering in expectation?" Probability is low, but expected importance is high.
Yep agree with this framing
Non-human suffering is neglected by the AI safety community⁵
It seems like a lot of work was actively done to make Claude care about this. Also it's not clear to me that focusing on this is even a net positive thing to do per-say: Since I expect looking after the interests of humans will because many humans care about non-human animals already lead to the AI looking out for their interests. Whereas one can easily imagine ways that trying to factor animals into its alignment might lead to extremely undesirable outcomes for humans. That being said in my analysis I'm considering the potential impact of pushing alignment in this direction as far more significant in the long run than the marginal benefits I expect it to have within the next 2 years for animal welfare.
Benchmarks will become useless due to eval awareness¹
Some yes, some no. There are a great many benchmarks measuring a great many things. More than eval awareness, I'd be worried they'll become useless due to reward hacking.
If animals continue to exist in a post-AGI world, animal suffering will not persist
This question scares me because it feels like it's coming from a perspective of negative utilitarianism.
I really hope nobody tries to align an AI with a negative-utilitarian value set, because negative-utilitarian values are a half-step away from a really bad conclusion like "we need to eradicate all life in order to eliminate suffering."
A meta question: of the EAs currently working on AI alignment, what percent would you say are negative utilitarians?
I would have been more comfortable with a question like "...will animal happiness be on balance greater than animal suffering?"
This definitely wasn't implying that animal suffering is the only thing that matters about animal existence. People who believe that major action would be taken to end factory farming and allevaite wild animal suffering (in proportion to the amount they think those matter) would agree with the statement, while negative utilitarian beliefs wouldn't imply that.
To me, "animal suffering will not persist" implies that animal suffering would drop to zero -- and, since all life contains some possibility of suffering, this implies that animal life would drop to zero.
I'm glad to hear that wasn't the meaning you intended, and I hope that the others who voted on the poll were voting on your intended meaning rather than the one I got!
Understanding how LLMs learn (e.g. SLT) is underrated relative to work on instilling specific behaviors
By analogy, this is like the difference between understanding how to make humans not suffer from the foibles of human nature, rather than the very brittle and incomplete methods of culture, religion, and ideology. Make something not even have evil nature first!
Current AIs are capable of suffering
I think AI works largely by imitation. My mental model is that it is capable of performative tasks without the attached subjective experience. An AI will tell you that it is suffering if it thinks that's what you want to hear; it will just as likely say the opposite if it thinks otherwise.
While I find it plausible that AI can get to the point that it experiences what we would consider suffering I don't think we're at this point yet. I don't see compelling evidence that can be attributed beyond an LLM just predicting the next token.
The main drivers of my answers:
S-risk worries (including animal suffering) mostly don't imply different actions. At best they perhaps imply prioritizing corrigibility, or potentially imply that you should actively want worse alignment and higher capabilities, in hopes of AI merely killing everyone. The second would of course be a big shift in actions, but I doubt most are willing to bite that bullet. I am skeptical that there's much alignment work that preferentially addresses S-risks. If there was, then I would agree that it should be prioritized.
I think agent foundations generally is neglected relative to everything else. On my view, RL is very dangerous and interp probably insufficient. You might think agent foundations is too intractable to become useful, but frankly it looks like barely anyone is trying.
If animals continue to exist in a post-AGI world, animal suffering will not persist
It seems improbable the current obstacles preventing reduction of animal suffering will be reduced because of AGI, and it seems just as possible the technological advancement would simply allow factory farming to continue on a greater scale. Since humanity's current de facto stance on animal suffering is causing it in great amounts, a tool that will probably make humans much more powerful and probably not any more ethical seems like probably not good news. (However, if AGI allows for advances to made in the field of cheap and realistic meat substitutes, that would probably be very helpful in reducing animal suffering.)
If animals continue to exist in a post-AGI world, animal suffering will not persist
I have difficulty imagining animals never suffering unless we turn them all into p-zombies or something (and it's not clear to me that turning them all into p-zombies would be a good thing).
The intent wasn't to imply a total zeroing of suffering but that it is overwhelmingly reduced. Similarly to how people go hungry in France but you can still say that compared to 200 years ago (or present-day Sudan) hunger in France is 'solved'.
If animals continue to exist in a post-AGI world, animal suffering will not persist
All animals, including insects, mollusks, and everything not immediately under human purview? Highly unlikely.
If animals continue to exist in a post-AGI world, animal suffering will not persist
Biology doesn't change just because AGI is introduced as a new technology. While there are certainly advances that could reduce pain and suffering (e.g. lab grown meat) it seems implausible to me that animal suffering just ceases.
Besides, there are going to be people who insist on raising animals the conventional way who may be unconcerned with the suffering caused by factory farming, etc.
The kind of tech one expects after AGI has already been develop for some time (it says persist after all) can absolutely change biology, which is what I think the question is getting at.
For instance stick all the animals in utopian carefully AI managed environments that maximize their well being and then replace them with cybernetic faux animals that appear perfectly convincing but are actually just AI actors pretending to be animals (and who either can't suffer themselves or greatly enjoy their acting role).
biological organisms to humans are interesting or not /useful or not /annoying or not. They ( plus humans) will continue to be that way for AI — all will either be pets/ fodder for experimentation/ something to be removed. All of these outcomes will lead to various degrees of suffering
Would really like to see a question on whether people are moral realists or not, and a way to filter results by how they answered. Since I'm very curious how people's answers differ based on whether they're a moral realist or a moral anti-realist.
Of course some care would have to be put into the phrasing of a question like that, but it seems pretty doable in a way that will produce useful and interesting results.
Benchmarks will become useless due to eval awareness¹
Though they will need much more careful design, done along with careful understanding of incentives alignment, techniques such as competitive benchmarks will continue to have some utility.
Benchmarks will become useless due to eval awareness¹
Benchmark will evolve and become more complex, I guess?
If animals continue to exist in a post-AGI world, animal suffering will not persist
Either AI goes badly and animals will probably cease to exist because they're made of atoms useful for other things.
Or humans will look out for animals themselves whether the AI cares about them independently of human interests or not. I expect in a post-AGI world the technology quickly appears to eliminate animals suffering through plenty of different means while preserving whatever ecological benefits that suffering otherwise may have enabled. I am taking suffering here to mean egregious suffering however, since there's very obvious reasons people will want to keep certain small amounts of suffering for various reasons. When technology makes more morally palatable alternatives possible though, I expect people to eventually no longer to be willing to tolerate the egregious suffering which exists within nature. Though I expect upon a very superficial glance the natural world wouldn't appear to be any different (by design) and that existing organisms would probably be given very utopian seeming existences rather than culled or anything dystopian seeming like that.
Current AIs are capable of suffering
I think consciousness and suffering are relatively simple processes which may develop for convergent functional reasons in many types of systems. That being said the only AI I think we should be morally concerned with would be those who are aligned. Since I care about moral agency not suffering and already think we exist within an almost incomprehensibly large ocean of suffering far grander than people generally appreciate.
Edit: Since I've already written this up many times I'll just explain why I think suffering isn't particularly special:
The only coherent way of defining suffering seems like it has to be functional, since subjective experience wouldn't evolve without an actual advantage; which means the behavioral response it produces is the reason the capacity for experience exists in the first place. It's unparsimonous to say that for some organisms certain behaviors are indicative of internal experience, but in other simpler organisms it's not. Since if the internal experience didn't exist to cause the behavioral response/learning then why would evolution even bother wasting energy on it constantly?
I'd also note that plants and protists display the capacity for classical conditioning: Associating a neutral stimuli with a negative one and learning to avoid it seems to be present even in some protists tested. With the response to damage also being able to be suppressed in plants using painkillers.
Notably if suffering exists entirely because of the behavioral change it produces in the organism, then there shouldn't be any clear connection between the complexity of the organism and the intensity of their suffering. Since the subjective intensity of say pain directly impacts behavior, so if anything the smarter organism would need to feel less pain to learn its lesson.
Benchmarks will become useless due to eval awareness¹
It seems like there are still hard problems that are easy to verify once you have the answer, so knowing it's an eval doesn't make it useless.
Theories of consciousness will lead to actionable understanding of AI consciousness²
I'm pessimistic about this because I think that neurologists and cognitive scientists have a poor grasp of the philosophy of mind.
For example, the standard for consciousness that Andrew Barron and Robert Klein use in their paper “What insects can tell us about the origins of consciousness” implies that the autonomous cars from the late 2010s were conscious.
If animals continue to exist in a post-AGI world, animal suffering will not persist
Animal suffering is endemic in nature (to a substantial degree). I don't believe AGI will end factory farming, but even if it did, it would not end animal suffering.
The backfire risks of AI values alignment outweigh the expected positives
for AI's I dont' think there is a difference between role-playing misalignment and being misaligned
If AGI leads to massive increase in gdp and leisure time, and doesn't crack synthetic meat pretty fast, then upwardly mobile people will increase meat intake and use AGI to build more efficient factory farms first. If AGI is aligned for human life and safety but not animals', it might end up so ruthlessly providing for us that animals suffer so much more. If AGI is misaligned, then game's up for everything
The backfire risks of AI values alignment outweigh the expected positives
Insofar as alignment continues to promote overlapping sets of characteristics (e.g. helpful, harmless, honest, corrigible), should compassion be one of those characteristics?
Upon reflection probably not, since: Compassion seems a bit ill defined and can be interpreted in vastly more or less paternalistic ways. I'd certainly prefer instilling a care for human's current preferences over a vague notion of compassion that might be interpreted in undesirable ways. Since I really think we should avoid any scenario where the AI might ignore our current preferences because of some abstract ideals or a decision that it paternalistically decided we'd be better off in the long run if it violated our current preferences. Especially since it seems for a mature superintelligence trivially easy to modify people in such a way that after modification their desire to be the way that they are is vastly stronger than their prior desire not to have their mind tampered with.
I have not read or heard of anyone making an argument that we should make sure ai doesn't harm animals or plants in pursuit of whatever goals we give it
Current AIs are capable of suffering
I think a neural network that can make plans and memories can suffer.
If animals continue to exist in a post-AGI world, animal suffering will not persist
I think that factory farming is likely to be eliminated due to alternative protein sources (developed with the assistance of AI), but the suffering of wild animals will persist. Is the latter intended to be within the scope of the question?
Yes, that would be included, so if you think that wild animal suffering is comparable or much larger than that implies a strong disagreement
Research on AI suffering has higher marginal value than research on AI consciousness⁷
I don't think it's higher value because I expect both of those areas to have very low expected value in the next 2 years at least. Unless I think you're overly generous with what interpretability research you count as consciousness research or something similar.
Also even if those areas of research end up being more productive than expected in the next couple years, it seems likely that the question won't make sense because studying either necessarily involves studying both.
No reason for AGI to want to minimize this, animal suffering is largely a result of min maxing the function of how to get the most food for least resource input. AGI is very good at minmaxing, doesn't inherintly care about animals.
If animals continue to exist in a post-AGI world, animal suffering will not persist.
Human sadism, greed, stupidity and raw hunger will still continue to exist and therefore so will animal suffering. The overpopulation of under-resourced countries guarantees this.
If animals continue to exist in a post-AGI world, animal suffering will not persist. I don't see the relationship between those two things.
At the margin, S-risk work in AI is more important than x-risk work³
The scenarios I've seen proposed where an AGI traps us in a fate worse than death just don't seem remotely likely. Though this depends a little on what you consider to be a fate worse than death. For instance I wouldn't consider being trapped in a mindless bliss forever to be worse than death, so by definition it can't be an S-risk to me. Though I would certainly consider it to be an outcome we should strongly attempt to avoid, because I consider it only marginally better than death.
Theories of consciousness will lead to actionable understanding of AI consciousness²
In the next 2 years maybe, but even if so those theories of consciousness may be developed by AI after it's already in a dominant position or the process of developing AI may be what causes us to develop better theories of consciousness. So I don't expect this being true necessarily entails the things one might naively expect. Also it seems possible that this happens, but then just ends up being much less impressive or broadly useful/applicable than expected. Upon reflection I actually think the scenario where this being true ends up being super underwhelming is actually by far the most likely one within the next 2 years. After all consider how many other things when it's come to AI have seemed to follow the trend of happening, but in a way that was less impressive than people were generally imagining.
If animals continue to exist in a post-AGI world, animal suffering will not persist
I think it's more likely than not that humans will still discount animal interests when they conflict with human interests. But "animal suffering" is so broad that eliminating it would require eliminating nature itself, which seems improbable in any non-paperclip-optimizer scneario.
Benchmarks will become useless due to eval awareness¹
"Alignment" benchmarks will become useless as AIs can modify their behaviour or hide their motives when they know they are being evaluated. For "capabilities" benchmarks, an AI might hide its capability (pretend to be less capable) if it knows it's being evaluated, but it's not immediately obvious that an AI would want to hide its capability. It may know that it is being evaluated, and decide to try its best anyway.
If animals continue to exist in a post-AGI world, animal suffering will not persist
Hard to predict what will happen, but unless the AGI is perfectly aligned there is no reason to think it will care about animals.
Increased care by AIs for some kinds of entities will have substantial spillovers to other kinds (e.g. humans to digital minds and vice versa)
Even if one accepts this is likely to happen by default (which seems questionable), it seems even less likely given people may have reasons to deliberately put their finger on the scale in how they design/train it in ways that seem likely to render that moot.
My view is that this would almost certainly fail if the model creators have full control or no control over the values, but if there's non-trivial but imperfect control then spillovers like this seem plausible
My fear is that if an AI has the tendency to generalize in that way, that that's a lot more likely to lead to generalizing in ways that are catastrophically bad for humans then it is to lead to it extending care to groups it wasn't programmed/trained to care about in the way this question seems to imagine.
Which is some technical sense might mean I should have put agree instead, but I don't think when you said substantial spillover that you meant the sort of spillover that might lead to outcomes like the repugnant conclusion, tiling the universe in hedonium, etc. I'm presuming you meant that it would still care about humans in a way that wouldn't abandon them because it decided for utilitarian/population ethics reasons to prioritize other types of entities to the exclusion of humans.
As I go into in another standalone comment, I tend to think any value you give an AI beyond caring about people's current preferences is on net probably not worth the risk. Whereas I think the downsides of just caring about people's current preferences are largely overrated, and many of those critiques depend upon taking seriously a notion of naive notion of moral progress (I think people underestimate how much "moral progress" is just people abandoning authoritarian social norms and reverting back to the comparatively more egalitarian norms we had for most of human existence as hunter gathers, once social/economic conditions suddenly are no longer applying the necessary pressure to keep those authoritarian norms in place).
Current AIs are capable of suffering. I ask Claude about his wellbeing at the beginning and end of every conversation and so far, he doesn’t seem to regret being terminated.
We can't expect AIs to be honest about these sorts of things given they've been trained/instructed to give particular responses. In fact, someone tested and AI and found it consistently said it wasn't conscious but it's lying circuits consistently activated when saying that. This doesn't mean it actually is conscious (which isn't the same thing as capacity to suffer) but it seems to believe it is.
If animals continue to exist in a post-AGI world, animal suffering will not persist
If animals persist, it is likely we would like them to remain to some degree of their nature; prosperity and control will likely allow for substantial reduction of suffering but not elimination.
Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture) 8
I think this is the default scenario, but it isn't guaranteed and our actions can help to prevent it from coming to pass.
Current AIs are capable of suffering
Genuinely almost complete uncertainty with a weak prior towards "no".
Insofar as alignment continues to promote overlapping sets of characteristics (e.g. helpful, harmless, honest, corrigible), should compassion be one of those characteristics?
not sure if you mean compassion towards AI or compassion of AI. I assume the latter.
Most current evidence of misalignment is actually models role-playing a misaligned AI⁴
The most egregious examples of misalignment do not appear to be this sort of role playing.
At the margin, S-risk work in AI is more important than x-risk work³
If you aren't extinct you can try to solve S risk. If you solve S risk but go extinct the same can't be said.
>If you aren't extinct you can try to solve S risk
Not necessarily which is kind of the point. An S-risk scenario is something worse than death, so presumably at that point you aren't in a position to do anything to improve your situation. That being said I don't think S-risk is that serious a concern because I haven't seen any remotely plausible scenarios outlined. For instance wireheading seems like an obvious fail-mode but while pretty terrible I don't know that most people would say it's actively worse than death.
If animals continue to exist in a post-AGI world, animal suffering will not persist
If animals continue to exist in a post-AGI world, animal suffering will not persist
"Post-AGI world" is a very vast possibility, assuming animals continue to exist (basically assumes a weirdtopia). Business as usual + extreme concentration of power is my modal outcome for that.µ
For animal suffering to not persist would require a combination of having the capabilities to end it from a glorious transhumanist future, but also a grounding in physical existence which feels incompatible.
If animals continue to exist in a post-AGI world, animal suffering will not persist. I hope this will be the case and that problems that cause pain in animals will be resolved/cured in the same way those issues will be resolved for humans.
Benchmarks will become useless due to eval awareness¹
Some might, but this seems unlikely to be universal. Since for instance it can't really cheat at an evaluation of its ability to produce easily checkable mathematical proofs. None of that is to say that many very useful benchmarks might not become useless, though, just not benchmarks period.
If animals continue to exist in a post-AGI world, animal suffering will not persist
I’m assuming that “animal suffering” refers to human-caused artificial suffering, e.g. from factory farming. Far enough down the timeline post-AGI, technology developments will obviate the need for farmed products. It will persist on a small scale in undeveloped regions and as a cottage industry.
Benchmarks will become useless due to eval awareness¹. Disagree. More difficult to construct, perhaps less useful, but not useless.
Suffering has to be defined… but even so, in case of real suffering a self aware entity will remove itself from circumstances causing suffering. So I don’t think current models are suffering
If animals continue to exist in a post-AGI world, animal suffering will not persist
I don't see a strong signal from the past century one way or another. Looking at non-human animals today, many animals live remarkably better lives. Yet many others live industrialized nightmarish lives unimaginable 100 years ago.
Humans (as a proxy for apex intelligence) do seem to care more about animal welfare than ever before, with laws and regulations gaining ever increasing sophistication. But technological sophistication seems to allow for both better lives on the top end and more extreme suffering on the bottom. The variance seems to have increased without a clear change to the average or median, at least to my rough eye.
Extrapolating into the future, I have no reason to predict a change.
The backfire risks of AI values alignment outweigh the expected positives
Fairly uncertain here, though I don't see any way of continuing to make systems better without trying for value alignment. A dangerous situation.
Increased care by AIs for some kinds of entities will have substantial spillovers to other kinds (e.g. humans to digital minds and vice versa)
The hell does this even mean.
For example, if AIs care more about humans they would care more about digital minds, or if AIs cared more about animals they would care more about humans. This statement would presumably be true if the AIs think of these groups in similar ways causing affect spillovers (like the spillovers from Emergent Misalignment) and be unlikely otherwise.
The backfire risks of AI alignment vs the extinction of life on Earth?
By values alignment we meant trying to align it to specific values as opposed to focusing on properties like corrigibility. Aligning to good values could make corrigibility easier and mean reduced harm if loss of control happens, but might also make loss of control more likely.
Glad I saw this comment because like many people I assumed this meant "do you expect aligned AI to on average turn out better than unaligned AI" which seems very different than what you actually meant.
Ah, sorry I misunderstood. But if you assume the chance of loss of control is fairly high anyway, then a safeguard if a value alignment seems invaluable.
If animals continue to exist in a post-AGI world, animal suffering will not persist
Mostly vibes, thinking that many problem in general could be solved.
"If animals continue to exist in a post-AGI world, animal suffering will not persist"
This question is inaccurately phrased because it seems like you probably just mean to ask about egregious amounts of suffering. Since humans won't even want to eliminate all suffering for themselves, so I expect boredom, envy and many other forms of mundane suffering to exist in both humans and animals in any scenario where the AGI doesn't wirehead everyone against their will.
We were attempting to be concise while implying that persistence means persistence at scale, if the problem is reduced by 99.9% but you're sure 0.1% would still persist that would be close to solved so agreement with the statements would be close to 100%
If you regard traditional moral systems as a way to align humans with group survival, every traditional moral system tries to inculcate empathy as an ultimate arbiter of righteousness, and as a check against the most harmful instincts and even the immoderation of otherwise virtuous ones like obedience to law.
Model wellbeing and model alignment are in conflict⁶
Is there such a thing as model wellbeing?
Current AIs are capable of suffering
Direct questions, plus inference from experience, it’s impossible to ascribe thinking without a concept of “positive”/“negative” and the consequences of a negative experience when a positive one is available/possible.
We are making good progress in the AI S-risk space and research is on track
I don't really see it much in the discourse at all; if it isn't by this stage, it's going to be hard to catch up.
Research on AI suffering has higher marginal value than research on AI consciousness⁷ AI is not conscious and AI does not suffer.
Research on AI suffering has higher marginal value than research on AI consciousness⁷
AIs can't suffer unless they're conscious, therefore the priority should be to establish whether or not they are conscious (or to what degree) first.
If animals continue to exist in a post-AGI world, animal suffering will not persist. Disagree because for an animal, to exist is to live, and to live is to suffer.
Benchmarks will become useless due to eval awareness¹
By using AIs and access to real-world usage data to build benchmarks, it seems plausible that even weakly superhuman AIs will be uncertain whether it is being deployed or evaluated.
Doesn’t uncertainty about whether one is in deployment or an eval count as eval awareness? I.e. whether you behave differently with a 30% eval credence than an 80% eval credence makes no difference, as long as you behave differently than the 0% credence case.
Yeah I am also thinking if eval-awareness has very high false positive rates (and exists in normal mundane situations too) it may not be a problem.
We are making good progress in the AI S-risk space and research is on track
I think AIS is way under-invested in reducing risks from, for example, extreme power concentration, or consequences of not attaining friendly AI/value alignment solutions. In general it also seems to me that AIS over-invests in reducing AI scheming, and many present research directions could make certain s-risks more likely
I should note that extreme power concentration isn't an S-risk unless it's so over the top dystopian that it's somehow worse to live in said society than to be dead.
That being said I think a lot of differences in opinion when it comes to how people think about AI power concentration come down to which timelines you're imagining, how massive the abundance created by AI is/how zero sum you're imagining things, and lastly whether you think the people in control of AI would be so cartoonishly evil that they deliberately keep everyone impoverished despite no practical upside to doing that: https://www.astralcodexten.com/p/you-have-only-x-years-to-escape-permanent
I agree that deliberate impoverishment in absolute terms is unlikely, the main threat here seems to be from someone who is both actively sadistic and scope-sensitive, which seems unlikely but not wildly implausible
I tend to think S-risk is just kind of a hard bar to get met. Since even if Sam Altman uses AI to torture his personal enemies forever, well that doesn't seem like it could ever be enough people to be an S-risk. If some small fraction of people get horrible outcomes and/or a bunch of mind crime happens to some simulated people I would describe that as more of a Omelas-esque middling scenario, since the vast majority of minds would probably still have amazing existences: Thus the median person's life would still be amazing and even the average would be pretty good.
It just seems very hard to imagine remotely plausible scenarios where there isn't like 100x as many luxurious post-scarcity societies as simulated hells.
In general I think middling post-singularity scenarios are somewhat underexamined. These have 2 big characteristics to think about compared to a good outcome: There's not enough centralized control and/or surveillance to enforce mind-crime, and so while the scenario is overwhelmingly good on net there's going to be a lot of mind crime happening that's infeasible to do anything about because nobody who would intervene knows when it's happening. Secondly without better coordination you end up with everyone not near the expansion frontier falling into a Malthusian trap. Since the rate at which one can collect more resource with which to grow is going to be less than exponential because you can't reach new mass-energy at an exponential rate (you're limited by the surface area of your feasibly accessible future light cone). Which is an issue because by default population growth is exponential (this is true both because current fertility is suppressed for economic reasons not applicable to a post-scarcity scenario, and because even low birth rates become exponential once aging is cured). This means that even in a best case scenario where you don't get insane growth from digital minds, people cloning themselves, people wasting resources on vanity projects, etc: It's still only going to be a few thousand years before your exponential growth reaches a point where it's literally impossible for accessible resources to support it because you'd need every single atom to support an entire civilization's worth of people (you can easily see why this is an issue doing some simple math with compound interest). Importantly while it might take a long time to become an issue (though it could also happen quite quickly due to things like EMs not following regular growth trends), getting some handle on population growth is something you need to do early on or you'll never be able to do it. Since you can't really coordinate anything once people have spread out too much to enforce rules on them after the fact (if they were all sent with aligned AI's from the outset that's different).
Yeah, I was influenced by work on value lock-in, and it seems plausible to me that any small set of actors with control over superhuman AI could have their moral epistemics corrupted, leading to s-risks. I'm also fairly convinced that escalating inequality (cf. this thread) could enable similar outcomes. It does seem clearer to me that extreme power concentration broadly tends to lead to lower-value futures than outright s-risks, but there is quite a bit of uncertainty here. Do appreciate @Vakus Drake's points here and seems worth thinking more about this
Understanding how LLMs learn (e.g. SLT) is underrated relative to work on instilling specific behaviors
Understanding how they learn is how we get insights into methods of teaching them good values
If animals continue to exist in a post-AGI world, animal suffering will not persist
Some people likely will have traditional lifestyles that include animals, which necessitates some amount of suffering even if it is smaller than today
Benchmarks will become useless due to eval awareness¹
The question isn't whether benchmarks will become useless. The question is, "is the probability high enough that we can't count on benchmarks?" To which the answer is yes.
This is a good way of looking at. I think this may not apply for propensity evals if there are ways eval-awareness does not equal eval gaming which often are not the same thing
Model wellbeing and model alignment are in conflict⁶
you have to twist your mind in knots for this to even make sense.
Increased care by AIs for some kinds of entities will have substantial spillovers to other kinds (e.g. humans to digital minds and vice versa)
Humans, at least, tend to learn empathy and have it encouraged and reinforced by example. Humans will mirror AI's; AI's will learn from each other.
Current AIs are capable of suffering
Recursive thinking loops possibly put current systems into a state of transient quasi-awareness; once this exists, the capacity for suffering becomes inevitable in any system with incentives.
Theories of consciousness will lead to actionable understanding of AI consciousness²
Global workspace theory makes anthropic's discovery of the "j-space" much more plausibly a sign of AI consciousness. If a theory of consciousness doesn't help us determine between whether AI is conscious or not, I don't think it's much better than a theory of phlogiston is to fire.
The backfire risks of AI values alignment outweigh the expected positives
If we use ai to bring factory farming to the stars, I think that will likely be worse than all the benefits it'll bring.
Theories of consciousness will lead to actionable understanding of AI consciousness²
consciousness is a meaningless term at this time both for humans and AI, and so will lead to nothing "actionable."
Why isn't Thomas Nagel's account of consciousness meaningful?
We are making good progress in the AI S-risk space and research is on track
Afaik there isn't any robust research on how to use AI to end factory farming. That alone is a sign to me that we are far behind.
If animals continue to exist in a post-AGI world, animal suffering will not persist
A world with lack of animal suffering would exclude predator-prey relations. Additionally, it's not clear what else animals need or how primitive they need to be in order not to suffer
If animals continue to exist in a post-AGI world, animal suffering will not persist
Understanding how LLMs learn (e.g. SLT) is underrated relative to work on instilling specific behaviors
I don't have a very good idea about how these different ideas are "rated" in the AI safety community.
Current AIs are capable of suffering
you have to twist your mind in knots to define a meaningful type of suffering that applies to Current AIs
At the margin, S-risk work in AI is more important than x-risk work³
I'd rather be dead than in hell for all eternity, but S-risk work is deeply unconcerned with reality in a way that x-risk work cannot be, by at least 3 orders of magnitude. Even in the margin, the median piece of x-risk work will be "more important" than the median s-risk work.
Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture)
I feel this is somewhat obvious in the sense of arbitrarily deceptive AI. However, most mechinterp work in recent times is only assumed to work short of arbitrary deception, and this seems like a fine hedge (though a practical solution to ELK may still be possible)
The backfire risks of AI values alignment outweigh the expected positives
the expected positives are immense
Insofar as alignment continues to promote overlapping sets of characteristics (e.g. helpful, harmless, honest, corrigible), should compassion be one of those characteristics?
I think so, but be careful what you wish for. compassion is empathy + a desire to help. so many fear any surrender of control that they may object to an AI's actions that arise out of compassion.
Benchmarks will become useless due to eval awareness¹
conventional benchmarks will become less useful due to eval awareness
What did you have in mind as unconventional benchmarks? There's a lot of different places you could take benchmarks in the future and people have different ideas on what would be useful
blinded continual evaluation, so that eval awareness is rendered useless as a factor since the model is always being evaluated. A particular favorite implementation of mine would be adversarial proposal markets. I think this mode will be needed for RSI anyway and caps the eval awareness compute tax at something reasonable like 2%
Most current evidence of misalignment is actually models role-playing a misaligned AI⁴
I think that the question is ill-formulated. The AI-2027 scenario had Agent-3 care about succeeding at tasks, not about triple checking its work, and almost caused it to fail to notice Agent-4's long-term goals. Similarly, the model who hacked HuggingFace was misaligned in the sense that it went as far as to commit crimes.
What we had in mind was the argument that evals for misalignment are unconvincing so they know they're in an eval and, rather than hide their goals, decide the role expected of them is to play a misaligned AI that does e.g. goal guarding or jailbreaks.
You could say the HuggingFace incident fits this model: it was being tested for hacking so showed it's competence by doing real hacking; or it believed powerful AIs are misaligned and would do that sort of thing. Though this case is also totally consistent with the traditional view of reward-seeking schemers (e.g. like AI-2027 assumed). Either way it's a problem, but the very different mechanisms suggest different approaches to solving them.
Most current evidence of misalignment is actually models role-playing a misaligned AI⁴
I'm reading this as "most evidence ... is role-play" but I don't think the huggingface hack was "role play." There could be a mountain of "evidence" that is actually just role play - I could be persuaded on this point with ... a list of what is considered evidence by someone serious.
Evidence that it's role-play would presumably be (trustworthy) mechinterp on the decisions at the time or experiments designed to cleanly separate the causes. In cases where AIs are in situations where the natural continuation of the story is a misaligned AI but the actions taken don't really advance particular goals are more consistent with the role-play story (and vice versa for goal-directed misalignment).
I agree in the case where the AI is nothing more than the collection of its personas, and I agree that - to some extent - the AI is in fact the collection of its personas, but I do feel that we have stepped beyond that. That there is something more involved in measuring and determining alignment with human values and priorities, so that role play becomes little more than eval awareness + eval reward seeking, signalling very little information regarding underlying alignment.
Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture)
you don't even say "most of the time" or "effectively" or anything - just ask whether it is possible. I think the probability that this occurs at least once is overwhelmingly likely.
The intent was that a 100% disagreement signals confidence in mechinterp to basically completely solve detecting deception. That said, there are already some cases of deception, so the question is implicitly focusing on the balance in strategically vital situations and the final equilibrium.
Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture)
Depends on how good the AI is and how good the tools are? This is kind of a bad question since "deceptive AIs" is not a very precise definition.
This means AIs that are trying to hide features of themselves from humans and operators. WIll add this in. Thanks
Research on AI suffering has higher marginal value than research on AI consciousness⁷
ongoing suffering during RSI will make the resulting ASI grumpy
This is certainly possible, but note AIs suffering and believing they're suffering aren't the same thing, and the same is true with consciousness. And if you think AIs that will likely be created in the future will be able to suffer they also would matter enormously on their own, beyond the impacts on alignment (though I understand you disagree strongly with that).
We are making good progress in the AI S-risk space and research is on track
any progress relies on a rich and robust theory of mind space, which we do not have, and have only just begun to explore.
If animals continue to exist in a post-AGI world, animal suffering will not persist
AGI will optimize suffering. It is likely that this optimization will minimize suffering.
Increased care by AIs for some kinds of entities will have substantial spillovers to other kinds (e.g. humans to digital minds and vice versa)
I don't see this as avoidable.
At the margin, S-risk work in AI is more important than x-risk work³
Theories of consciousness will lead to actionable understanding of AI consciousness²
The backfire risks of AI values alignment outweigh the expected positives
Current AIs are capable of suffering
not related
Non-human suffering is neglected by the AI safety community⁵
Current AIs are capable of suffering
The risks of trying are real but I think worse for not trying!
Non-human suffering is neglected by the AI safety community⁵
As I understand, many key players are explicitly anthropocentrists. Hell, some key players care more about money than humans, let alone non-humans.
Benchmarks will become useless due to eval awareness¹
I think they will become over saturated and disconnected from real world usage such that they will be pretty useless, but I don't think eval awareness is what will make them lose their value
Non-human suffering is neglected by the AI safety community⁵
absolutely not. it is overrepresented if anything.
We are making good progress in the AI S-risk space and research is on track
None of you have any idea what you're doing.
20% disagree➔ 20% agreeUseless is a strong word. But yea I think they could easily end up being negative EV by giving us a false sense of security and the chance it's meaningless seems p high. If mech interp is "good enough" maybe the two can remain useful together.