Why complete cluelessness is counterintuitive to me
I mean, specifically, this kind of pervasive cluelessness, where you can't justifiably decide between any two actions. It seems that such pervasive cluelessness is in tension with the idea of instrumental convergence. I'm personally sympathetic to imprecise probabilities, but intuitively I'm not convinced that cluelessness is that bad that we can't justify even basic learning and other (supposedly convergent) instrumental strategies, such as (at least) those in the low-footprint capacity building category.
And I think, if we have to show that cluelessness is at least not that pervasive, it seems useful to look at the most seemingly absurd cases. Being clueless about whether epistemic improvement is worthwhile is one of them. If we (or other aligned agents that we create) can in principle be non-clueless about some specific class of strategies, for instance, why wouldn't we at least sometimes be justified in taking actions that would make us non-clueless about them?[1] (hence, being additionally non-clueless about whether this transition is better than doing nothing or pursuing some other alternative with a small opportunity cost) And if we are indeed justified in doing so, why can't we derive some broader system of instrumental strategies from this fact?[2] (hence, potentially, having even more non-clueless states) It sounds strange that even strategies aimed at becoming non-clueless wouldn't be better justified than doing something random.
I have to admit, the situation is much less obvious than I first thought. Interestingly, some imprecise probabilists may have to pay to avoid free knowledge (unlike precise ones). And there is a thing called dilation: new information may make your intervals wider than before. But most importantly, knowledge is not actually free. And it's yet unclear to me how we should account for long-term changes in the structure of our entire decision trees in general. This seems to involve the topic of sequential rationality.
Anyway, why is impartial altruism so different from other value systems in this respect?[3] DiGiovanni writes:
...suppose that instead of spending some of your free time reflecting on virtue ethics, you reflect on whether some intervention could be robust to unawareness. Maybe this shapes the salience of different strategies to your future self, and you end up attempting an intervention that’s worse than what you would’ve done by default.
In principle, we can come up with examples where learning goes wrong, but it's not yet clear why we should treat them as severely undermining the whole idea of epistemic improvement.
Intuitively, these side effects may seem like hand-wringing nitpicks. But I think this intuition comes from mistakenly privileging intended consequences, and thinking of our future selves as perfectly coherent extensions of our current selves.
But we don't have to think that our future selves are such perfectly coherent extensions. They are indeed different, but they may differ to a degree small enough for us to make a justified decision.
So far, it is not clear why impartiality is special. Learning seems to have broadly similar structures across many domains. Impartial altruism doesn't seem to be an exceptional domain such that some magical demons suddenly appear from nowhere to make you go mad if you know too much.[4] Then why aren't there some learning strategies available to us that are more justified than doing nothing?[5]
We’re trying to weigh up speculative upsides and downsides to a degree of precision beyond our reach. In the face of this much epistemic fog, unintended effects could quite easily toggle the large-scale levers on the future in either direction.
This seems to be relevant when deciding between some direct interventions (such as donations), but how does the associated cluelessness infect learning itself—which is supposed to help us compare them? Cluelessness may emerge here if we don't know how to compare learning with its alternative, which is possibly higher-EV. But what if the alternative we consider is something which has low direct impact on your environment, such as resting? Then by choosing to learn, you may miss something valuable if it lies further down your decision tree and resting makes that opportunity more accessible. This is certainly normal, but isn't unique to impartial altruism.
We shouldn't always prefer learning to doing nothing—it's even possible that we should do it less often than some other things. This is because successful approaches to epistemic improvement should mix learning with other things, such as resting when tired, eating when hungry, not dying in the process, etc. But isn't it possible for us to identify some combination of such low-impact actions that is better than, for instance, indefinitely pursuing just one of them? Empirically, this seems possible in many domains, and we may be able to identify some similarities across them. So again, it seems unclear, what is it about impartial altruism that makes things more clueless here—to the degree that we can't even decide whether we better be alive and learn things.
It will be helpful to have greater clarity about the potential sources and mechanisms of cluelessness (and their relative strength) in such specific most-absurdly-looking cases.[6]
If you think you're justified in c-preferring the SWP donation, then Anthony's claim is disproved. But Anthony has said he doesn't consider this a counterexample, so it seems unlikely that offering more candidate counterexamples would move the debate forward.
That's especially so since I think the scope of Anthony's claim is intended to rule out other candidate counterexamples, e.g.:
You're trapped in a box that you know for sure will implode in 10 seconds. There's a puppy in there with you. You're justified in c-preferring not kicking the puppy to kicking the puppy.
That seems true to me, but I think this pair of actions falls outside the scope of Anthony's claim. He's talking about actions with effects that aren't so tightly limited in space and time.
So the debate calls for something more than just a bare counterexample. As Anthony says in another comment, I try to give that 'something more' in my post. Donating $5 to MAWF is justifiably c-dispreferred to some mixed action, as is any other action that fails the 'flanking variants' test. That likely rules out almost all actions.
Yeah, I think it's easy to find a counterexample to "unconditional cluelessness" (i.e., even in the box situation), but much harder to find one relevant to what you and I should do in our present non-simplified situations (which is presumably what Anthony meant for us to discuss).
Hi Jim. I agree trying to come up with counterexamples is useful. Below are 3 potential counterexamples I have given. @Anthony DiGiovanni 🔸 does not consider them counterexamples (see Anthony's replies for details).
1st example, which Elliot already quoted in this thread.
Consider these 2 options for what I could do tomorrow:
Torturing my family, and friends, and then killing myself. I would never do this.
My understanding is that you think it is "irreducibly indeterminate" which of the above is better to increase expected impartial welfare, whereas I believe the 2nd option is clearly better.
Hi Anthony. Do you think the expected welfare of 2 states of the world which only differ infinitesimally can be incomparable? I do not see how this could be possible. For example, it feels super counterintuitive to me that, given 2 identical states, moving an electron by 10^-100 m in one of the states would make their expected welfare incomparable. I guess one can get from any state of the universe to another in an astronomical number of infinitesimal steps, and I believe any 2 states which only differ infinitesimally are comparable. So I conclude any 2 states are comparable too, even if it is very hard to compare them, to the point that I do not know if electrically stunning shrimps increases or decreases welfare in expectation.
Here is a 4th example. Consider these 2 actions:
Killing the 100 people who are expected to decrease the most the uncertainty about how to compare the expected value of different actions. For example, Bob Fischer who has worked on decreasing uncertainty about comparing welfare across species.
Grating 10 M$ to the 100 people above (100 k$ per person, but the grant size could vary). The 10 M$ would otherwise be spent torturing people as much as possible.
I think we are justifyed in c-preferring the 2nd action. I believe Anthony disagrees
Yeah I still disagree for this 4th example as well, for the same reasons as the 1st.
I worry about a motte-and-bailey implicitly happening here, where the motte is "all things considered, we should prefer the second action" — very difficult to deny! — and the bailey is "we should c-prefer the second action". This matters because the corresponding motte seems actually pretty easy to deny (IMO) when the comparison is "study some altruistically irrelevant branch of academic philosophy" vs. "try to prevent AI misalignment". The latter only looks clearly preferable to me if it's c-preferable. (I guess this is what Ben's comment is getting at.)
(Just to be clear, that's a contingent matter. I don't find any of the counterexamples offered so far persuasive because I don't think they adequately engage with my arguments for P3.)
Do you think you've ever taken an action (or sequence of actions) that, with your current understanding of cluelessness and unwarenesss, should have been ex ante c-preferred to another (or doing nothing, specifically)? Like thinking a bit more about specific backfire risks or cluelessness, or breathing?
EDIT: Also some more weirder things, like not kicking a puppy when given the chance and no one else would know. Or, say, if you've already stepped on a snail and it's clearly going to die, should you put it out of its misery?
I think no. Basically, when I really internalize how dwarfed every action's cosmic-scale consequences are by off-target effects, "no" feels very common-sensical to me. (Cf. this paper on how "simple cluelessness" is fake.)
I think I wouldn't be clueless about c-preferability in Elliott's "trapped in a box" example, but can't think of any real-world case analogous to this.
I would be interested in a version of this contest attempting to answer something like "If you believe Vasco that donations to SWP are better than torture, should you also believe that AI alignment is better than misalignment?"
I think that is what Richard Chappell is referring to when he mentions "radical skeptics" here. But I would be interested in a version of his post which defends the claim that EAs are reasonably justified in making the trade-offs we currently make (even if a radical skeptic would not be convinced of this defense).
I think that is what Richard Chappell is referring to when he mentions "radical skeptics" here. But I would be interested in a version of his post which defends the claim that EAs are reasonably justified in making the trade-offs we currently make (even if a radical skeptic would not be convinced of this defense).
Some of the contest entries do at least briefly attempt that, I think. E.g. this post gives an argument that one should c-prefer "low-footprint capacity-building" over "doing nothing". And this post more generally argues that "donate $5 to Make-A-Wish Foundation" is c-dispreferable to some mixed action (maybe that's not specific enough for what you have in mind). (I don't yet buy either of these arguments, though.)
The maximality rule is too demanding. Under maximality, we can't even say that [1, 10 000] is better than [-10 000, 1.0001]. Are there any good reasons to avoid more permissive rules even in cases like this?
Of course, if your UEV intervals overlap and you choose one of them, you may make a mistake if your idealized version would choose the other one. But choosing at random is no better in this regard.
What seems important is how costly such mistakes are, i.e., how the overall performance of your choice rule compares with that of other rules—such as choosing at random.
Can these performance measures be defined without additional assumptions about how values are distributed within UEV intervals?
Under maximality, we can't even say that [1, 10 000] is better than [-10 000, 1.0001].
That's not quite right. Maximality says an action A is impermissible when some alternative B has higher EV on every probability function in your representor. And that can be true even when A's and B's EV ranges overlap.
Example:
If A had the same range but sloped the other way, then it would be permissible by maximality:
So to figure out what's permissible under maximality, we can't just look at ranges. We need to look at the representor.
Yes, but this means that you know something additional about the structure of the representor, not just its range. I'm asking whether we can do better than maximality even for intervals alone, without adding any further details.
(By the way, pictures are broken.)
UPD: Though, you probably mean that we do know some additional structure for EV if we look at how it is constructed from probabilities and utilities, for which we have just intervals without structure.
And yes I think often we know more than just intervals of EVs. For example, we know whether the EV of some action increases or decreases with the probability of some proposition X.
Just sharing my quick takes on how to navigate cluelessness with a rather long-term strategy here. It doesn't get to predictability land but to enough insight to move forward carefully despite vast uncertainties.Â
Reading through all the essays and comments this week, I've become much more convinced that cluelessness is pervasive and standard ways of handling it fail. I agree with the normative and conceptual premises.
But the empirical premise is untrue. In particular, it is untrue that for any pair of actions, the actions' consequences are too coarse to be compared.
Consider the following counterexample: I could either donate 10% of my income to GiveDirectly or take a nice vacation. Is it true that our knowledge of the consequences of these actions is so coarse-grained that the comparison between them is indeterminate? I don't think so, even though in principle the same concerns about the catch-all and unawareness apply.
The "best guess" justification for donating is simple. My lifetime income is very likely to be in the top 5% globally; under all plausible theories of utility, marginal utility declines with income; thus, the marginal transfer from me to someone in extreme poverty improves aggregate well-being.
How might this be incorrect or imprecise?
The recipient of the donation could use it for something bad. (This is analogous to the meat-eater problem, but given my own consumption of animal products the meat-eater problem itself doesn't apply.)
Catch-up growth in recipient countries could lead to an unstable multi-polar world order, which eventually causes a catastrophic WWIII.
Because I don't go on vacation, I don't accidentally step on a butterfly I otherwise would have. The butterfly flaps its wings and causes a tornado that kills 50 people. Also, those 50 people were all wild animal suffering researchers.
Everyone else is a p-zombie; only my consumption matters, and I should take more vacations.
In case this seems too glib, I promise I tried to come up with better examples (and asked AI). I do not think there is any plausible probability distribution under which the above are likely enough that taking the vacation is better than donating. I would welcome suggestions.
"Plausible" is doing a lot of work here. I don't have a good justification for why no probability distribution where these events are likely enough to change the outcome of the comparison is plausible to me. All I can say is that if we consider these type of events to be equally plausible to the "best guess", then I agree with Richard Y Chappell that cluelessness considerations approach radical skepticism.
More formally, I would claim that the set S of all possible actions contains a non-empty proper subset C of actions over which an ideal impartially altruistic agent has complete preferences. The comparison between each action in C and each action in S\C is indeterminate. This still offers a lot of guidance on impartial altruistic action, because I would claim we all have more than one action within C available to us.
There is still a lot of debate to be had and research to be done to figure out which actions are in C and which aren't (see Jim Buhler's "The Train to Arbitrary Land" for a good elaboration of this problem.)Â
Still, this offers a meaningful change from previous cause prioritization discussions, which generally assumed that all actions are in C. Practically, these considerations push me towards near-term interventions, whose sign is clear and comparisons of which with alternative uses of resources I believe to be determinate.
My lifetime income is very likely to be in the top 5% globally; under all plausible theories of utility, marginal utility declines with income; thus, the marginal transfer from me to someone in extreme poverty improves aggregate well-being.
How might this be incorrect or imprecise?
Here are two potential reasons I find far more plausible than yours, fwiw: - the GiveWell donation increases farmed animal suffering more than it increases human welfare. - it slightly increases the total human population, which slightly accelerates human progress and hence increases AI x-risks. The difference is small but the long-term future may overwhelm the short-term benefits of GW so much that this is enough.
The thesis that we are clueless about the overall sign of donating to effective global health and development charities is actually one of the most consensual in the cluelessness literature (see Mogensen 2021; Kollin et al. 2025, §1; Greaves 2016).Â
Longtermist causes are those that generate significant disagreement (see, e.g., this overview and refs therein).
Thank you for the response, and in particular for the references for further reading!
I'm curious to hear more about your reasoning for those examples, or why they seem especially likely under some plausible distribution. (For what it's worth, I'll clarify the hypothetical donation goes to GiveDirectly, not GiveWell. Also, I eat meat, and my meat consumption is sensitive to my income.)
I agree that it is possible for them to occur, but there is no probability distribution I would include in my representor under which they are sufficiently likely to occur that the comparison between actions is indeterminate.
I'm not sure how helpful it is to dwell on specific examples, but this is the disagreement I have with many arguments for cluelessness about neartermist causes.Â
The story given in Mogensen 2021 Â of the priest saving Hitler from drowning as a child is illustrative. This was the right thing to do. There have been billions and billions of children; a handful have had both the opportunity and desire to commit horrific genocides. I don't believe there is a plausible probability distribution that make it so likely an anonymous child will commit genocide that it is better to let them drown.
Oops sorry for the GW-GD confusion, but yeah, this changes nothing to my point I think.
I agree that it is possible for them to occur, but there is no probability distribution I would include in my representor under which they are sufficiently likely to occur that the comparison between actions is indeterminate.
This seems hardly defensible. - The average human being (including in poor countries) contributes to the farming of so many animals throughout their lifetime (at least in expectation). You would have to be astonishingly confident that the welfare of the animals we eat do not significantly matter for your above conclusion to follow. - The number of far-future lives we indirectly influence might be astronomical such that long-term effects (almost) always dominate. I don't see on what basis you can exclude this possibility from your probability distributions, given the arguments given here and refs therein.
(I don't recall this specific example from Mogensen and don't have time to dive back into it, so I won't comment on that, sorry.)
The problem of Cluelessness has been grappled with for millennia; two relevant examples:
1. The story of Khidr in the Quran, who, through a series of nonsensical or seemingly evil actions, demonstrates to Moses how we are Clueless about consequences.
2. The Bhagavad Gita,which advises disregarding consequences altogether:
To action alone hast thou a right and never at all to its fruits; let not the fruits of action be thy motive; neither let there be in thee any attachment to inaction."
The solutions end up looking like Deontology or Virtue Ethics, which Anthony (inadvertently?) references in his summary: "Other values and moral norms still matter to us, for example, rules like avoiding dishonesty or virtues like compassion."
Anyone grappling with the possibility of Cluelessness would do well to consider prior art.
For the sake of provocation: How is managing Cluelessness in Philanthropy any different from managing Cluelessness in, say, running a neighborhood coffee shop?Â
There is an unknowable amount of cause-effect relationships involved in running a business. Is there anything unique about Philanthropic modeling? Modeling dynamical systems is, in general, notoriously hard.Â
I wrote a few paragraphs fleshing out my question here if anybody wants to respond to me more in full.Â
This is briefly addressed by DiGiovanni here and there (his cluelessness X near-term AW post is also relevant), and I discuss some aspect of his point in the first ref a bit further here.
While the essays are being judged, we'd love for this event to kick off more valuable discussion of Anthony's Sequence, and the problems for impartial altruists it raises (for a refresher, read this summary).Â
To help with this goal, we're hosting a comment competition throughout the week. The best comments, according to me, @Anthony DiGiovanni 🔸 and @Will Aldred, will be awarded in multiples of $100, up to a total of $2000[2].Â
We're looking in particular for comments that:
Show a deep understanding of Anthony's sequence and the problem of unawareness,
And advance the conversation, by introducing further considerations, clarifying key points, making a new critique.Â
We may also award comments simply because they are very helpful for the discussion - i.e. comments that clear up a persistent confusion.Â
Awards will look like this:
How to use this discussion thread:
This is a place to post any questions or comments you have after reading the sequence, or the competition entries. This could be in the form of:
Clarification questions. If something confused you, it probably confused someone else.
Statements that you'd like to debate with others. Consider adding a poll to your comment if you are making a bold statement.
Quick arguments, even scrappy ones, that another commenter could disagree with or take further.Â
I.e., those that engage with Anthony's sequence, and offer a critique or solution. I'll also refrain from publishing some of the more unreadable (because AI-generated) pieces, since they won't be interesting to the Forum audience.Â
Same caveats here as listed in the disclaimer section here. Additionally, we may award less than the total $2000 if we don't find enough comments that we consider worth awarding. Note also that we will endeavour to award comments throughout the week, but some awards may be given the week after.Â
A preliminary estimate, and a request for better ones.
Summary
I believe the standard literature estimates for the number of DALYs attributable to a case of stunting are too low, largely because they don’t account for the long term effects. This means that childhood nutritional interventions that reduce the prevalence of stunting may be substantially more cost-effective than previously believed.
Epistemic status
Exploratory and back-o...
I’ve been feeling pretty shaken since the METR report about the Hugging Face incident came out last week. Over the weekend, I wrote up some thoughts on how lonely the AI situation sometimes feels to me. It’s more personal than what I’d usually share publicly, but I thought I’d post it here in case it resonates with anyone else.Â
I’m very grateful to the man...
TL;DR: Kairos has raised $50 million from Coefficient Giving for two years of funding, one of the largest commitments they’ve made towards AI safety fieldbuilding to date. We’re using this to make an ambitious push for growing Kairos, broadening our portfolio of talent infrastructure projects and incubating new organizations. We’ve doubled in size in the last six mon...
Why complete cluelessness is counterintuitive to me
I mean, specifically, this kind of pervasive cluelessness, where you can't justifiably decide between any two actions. It seems that such pervasive cluelessness is in tension with the idea of instrumental convergence. I'm personally sympathetic to imprecise probabilities, but intuitively I'm not convinced that cluelessness is that bad that we can't justify even basic learning and other (supposedly convergent) instrumental strategies, such as (at least) those in the low-footprint capacity building category.
And I think, if we have to show that cluelessness is at least not that pervasive, it seems useful to look at the most seemingly absurd cases. Being clueless about whether epistemic improvement is worthwhile is one of them. If we (or other aligned agents that we create) can in principle be non-clueless about some specific class of strategies, for instance, why wouldn't we at least sometimes be justified in taking actions that would make us non-clueless about them?[1] (hence, being additionally non-clueless about whether this transition is better than doing nothing or pursuing some other alternative with a small opportunity cost) And if we are indeed justified in doing so, why can't we derive some broader system of instrumental strategies from this fact?[2] (hence, potentially, having even more non-clueless states) It sounds strange that even strategies aimed at becoming non-clueless wouldn't be better justified than doing something random.
I have to admit, the situation is much less obvious than I first thought. Interestingly, some imprecise probabilists may have to pay to avoid free knowledge (unlike precise ones). And there is a thing called dilation: new information may make your intervals wider than before. But most importantly, knowledge is not actually free. And it's yet unclear to me how we should account for long-term changes in the structure of our entire decision trees in general. This seems to involve the topic of sequential rationality.
Anyway, why is impartial altruism so different from other value systems in this respect?[3] DiGiovanni writes:
In principle, we can come up with examples where learning goes wrong, but it's not yet clear why we should treat them as severely undermining the whole idea of epistemic improvement.
But we don't have to think that our future selves are such perfectly coherent extensions. They are indeed different, but they may differ to a degree small enough for us to make a justified decision.
So far, it is not clear why impartiality is special. Learning seems to have broadly similar structures across many domains. Impartial altruism doesn't seem to be an exceptional domain such that some magical demons suddenly appear from nowhere to make you go mad if you know too much.[4] Then why aren't there some learning strategies available to us that are more justified than doing nothing?[5]
This seems to be relevant when deciding between some direct interventions (such as donations), but how does the associated cluelessness infect learning itself—which is supposed to help us compare them? Cluelessness may emerge here if we don't know how to compare learning with its alternative, which is possibly higher-EV. But what if the alternative we consider is something which has low direct impact on your environment, such as resting? Then by choosing to learn, you may miss something valuable if it lies further down your decision tree and resting makes that opportunity more accessible. This is certainly normal, but isn't unique to impartial altruism.
We shouldn't always prefer learning to doing nothing—it's even possible that we should do it less often than some other things. This is because successful approaches to epistemic improvement should mix learning with other things, such as resting when tired, eating when hungry, not dying in the process, etc. But isn't it possible for us to identify some combination of such low-impact actions that is better than, for instance, indefinitely pursuing just one of them? Empirically, this seems possible in many domains, and we may be able to identify some similarities across them. So again, it seems unclear, what is it about impartial altruism that makes things more clueless here—to the degree that we can't even decide whether we better be alive and learn things.
It will be helpful to have greater clarity about the potential sources and mechanisms of cluelessness (and their relative strength) in such specific most-absurdly-looking cases.[6]
Alternatively, we may also consider the possibility that all realistic agents with sufficiently ambitious value systems should be clueless. Like, should realistic impartial paperclip maximizers be clueless too? This doesn't necessarily imply unconditional cluelessness, but it would also be very pessimistic. ↩︎
For example, maybe we could look at strategies that increase the probability of successfully creating aligned successors while reducing the probability of failure. ↩︎
Alternatively, maybe we should be clueless about some small-scale areas too, not only about impartial altruism? If so, it might be useful to study such areas to identify where exactly cluelessness starts to emerge. ↩︎
In principle, we may consider a possibility that there is such a hypothetical level of self-improvement, after which something strange happens—as in the case where the simulation hypothesis is true and the masters of simulation decide to stop you from being too successful :) Or perhaps there are some cognitohazards waiting out there in philosophy that will make you abandon your impartial altruism for some reason. But do these possibilities seem severe enough to make us really clueless about learning? ↩︎
To be clear, there is a huge difference between learning some random stuff and deliberately trying to learn more about possible crucial considerations, for instance. I'm mainly concerned with the decision-relevant knowledge here. Many things may be helpful when deciding between interventions directly aimed at helping our moral patients, and many other things are also helpful even if they are less directly related—such as learning how to live longer, or how to acquire more resources to have more impact. And "doing nothing" here is, of course, doing something—but with a (relatively) small direct impact on what happens around us. ↩︎
By the way, maybe we should use less-demanding decision rules than maximality—such as in the graded approach to comparison. ↩︎