Why complete cluelessness is counterintuitive to me
I mean, specifically, this kind of pervasive cluelessness, where you can't justifiably decide between any two actions. It seems that such pervasive cluelessness is in tension with the idea of instrumental convergence. I'm personally sympathetic to imprecise probabilities, but intuitively I'm not convinced that cluelessness is that bad that we can't justify even basic learning and other (supposedly convergent) instrumental strategies, such as (at least) those in the low-footprint capacity building category.
And I think, if we have to show that cluelessness is at least not that pervasive, it seems useful to look at the most seemingly absurd cases. Being clueless about whether epistemic improvement is worthwhile is one of them. If we (or other aligned agents that we create) can in principle be non-clueless about some specific class of strategies, for instance, why wouldn't we at least sometimes be justified in taking actions that would make us non-clueless about them?[1] (hence, being additionally non-clueless about whether this transition is better than doing nothing or pursuing some other alternative with a small opportunity cost) And if we are indeed justified in doing so, why can't we derive some broader system of instrumental strategies from this fact?[2] (hence, potentially, having even more non-clueless states) It sounds strange that even strategies aimed at becoming non-clueless wouldn't be better justified than doing something random.
I have to admit, the situation is much less obvious than I first thought. Interestingly, some imprecise probabilists may have to pay to avoid free knowledge (unlike precise ones). And there is a thing called dilation: new information may make your intervals wider than before. But most importantly, knowledge is not actually free. And it's yet unclear to me how we should account for long-term changes in the structure of our entire decision trees in general. This seems to involve the topic of sequential rationality.
Anyway, why is impartial altruism so different from other value systems in this respect?[3] DiGiovanni writes:
...suppose that instead of spending some of your free time reflecting on virtue ethics, you reflect on whether some intervention could be robust to unawareness. Maybe this shapes the salience of different strategies to your future self, and you end up attempting an intervention that’s worse than what you would’ve done by default.
In principle, we can come up with examples where learning goes wrong, but it's not yet clear why we should treat them as severely undermining the whole idea of epistemic improvement.
Intuitively, these side effects may seem like hand-wringing nitpicks. But I think this intuition comes from mistakenly privileging intended consequences, and thinking of our future selves as perfectly coherent extensions of our current selves.
But we don't have to think that our future selves are such perfectly coherent extensions. They are indeed different, but they may differ to a degree small enough for us to make a justified decision.
So far, it is not clear why impartiality is special. Learning seems to have broadly similar structures across many domains. Impartial altruism doesn't seem to be an exceptional domain such that some magical demons suddenly appear from nowhere to make you go mad if you know too much.[4] Then why aren't there some learning strategies available to us that are more justified than doing nothing?[5]
We’re trying to weigh up speculative upsides and downsides to a degree of precision beyond our reach. In the face of this much epistemic fog, unintended effects could quite easily toggle the large-scale levers on the future in either direction.
This seems to be relevant when deciding between some direct interventions (such as donations), but how does the associated cluelessness infect learning itself—which is supposed to help us compare them? Cluelessness may emerge here if we don't know how to compare learning with its alternative, which is possibly higher-EV. But what if the alternative we consider is something which has low direct impact on your environment, such as resting? Then by choosing to learn, you may miss something valuable if it lies further down your decision tree and resting makes that opportunity more accessible. This is certainly normal, but isn't unique to impartial altruism.
We shouldn't always prefer learning to doing nothing—it's even possible that we should do it less often than some other things. This is because successful approaches to epistemic improvement should mix learning with other things, such as resting when tired, eating when hungry, not dying in the process, etc. But isn't it possible for us to identify some combination of such low-impact actions that is better than, for instance, indefinitely pursuing just one of them? Empirically, this seems possible in many domains, and we may be able to identify some similarities across them. So again, it seems unclear, what is it about impartial altruism that makes things more clueless here—to the degree that we can't even decide whether we better be alive and learn things.
It will be helpful to have greater clarity about the potential sources and mechanisms of cluelessness (and their relative strength) in such specific most-absurdly-looking cases.[6]
Alternatively, we may also consider the possibility that all realistic agents with sufficiently ambitious value systems should be clueless. Like, should realistic impartial paperclip maximizers be clueless too? This doesn't necessarily imply unconditional cluelessness, but it would also be very pessimistic. ↩︎
For example, maybe we could look at strategies that increase the probability of successfully creating aligned successors while reducing the probability of failure. ↩︎
Alternatively, maybe we should be clueless about some small-scale areas too, not only about impartial altruism? If so, it might be useful to study such areas to identify where exactly cluelessness starts to emerge. ↩︎
In principle, we may consider a possibility that there is such a hypothetical level of self-improvement, after which something strange happens—as in the case where the simulation hypothesis is true and the masters of simulation decide to stop you from being too successful :) Or perhaps there are some cognitohazards waiting out there in philosophy that will make you abandon your impartial altruism for some reason. But do these possibilities seem severe enough to make us really clueless about learning? ↩︎
To be clear, there is a huge difference between learning some random stuff and deliberately trying to learn more about possible crucial considerations, for instance. I'm mainly concerned with the decision-relevant knowledge here. Many things may be helpful when deciding between interventions directly aimed at helping our moral patients, and many other things are also helpful even if they are less directly related—such as learning how to live longer, or how to acquire more resources to have more impact. And "doing nothing" here is, of course, doing something—but with a (relatively) small direct impact on what happens around us. ↩︎
By the way, maybe we should use less-demanding decision rules than maximality—such as in the graded approach to comparison. ↩︎
DiGiovanni states:
A single counterexample of a pair of actions where one is c-preferred over the other would suffice to disprove his claim.
It seems noteworthy that none of the solutions attempted this.
I don't think that's so surprising. There are obvious candidate counterexamples, e.g.:
If you think you're justified in c-preferring the SWP donation, then Anthony's claim is disproved. But Anthony has said he doesn't consider this a counterexample, so it seems unlikely that offering more candidate counterexamples would move the debate forward.
That's especially so since I think the scope of Anthony's claim is intended to rule out other candidate counterexamples, e.g.:
That seems true to me, but I think this pair of actions falls outside the scope of Anthony's claim. He's talking about actions with effects that aren't so tightly limited in space and time.
So the debate calls for something more than just a bare counterexample. As Anthony says in another comment, I try to give that 'something more' in my post. Donating $5 to MAWF is justifiably c-dispreferred to some mixed action, as is any other action that fails the 'flanking variants' test. That likely rules out almost all actions.
Yeah, I think it's easy to find a counterexample to "unconditional cluelessness" (i.e., even in the box situation), but much harder to find one relevant to what you and I should do in our present non-simplified situations (which is presumably what Anthony meant for us to discuss).
Agreed. I actually wish people would try to find one counterexample instead of trying to prove too much. This might allow us to specify some cruxes.
Hi Jim. I agree trying to come up with counterexamples is useful. Below are 3 potential counterexamples I have given. @Anthony DiGiovanni 🔸 does not consider them counterexamples (see Anthony's replies for details).
1st example, which Elliot already quoted in this thread.
2nd example.
3rd example.
Here is a 4th example. Consider these 2 actions:
I think we are justifyed in c-preferring the 2nd action. I believe Anthony disagrees
Some of the contest entries do at least briefly attempt that, I think. E.g. this post gives an argument that one should c-prefer "low-footprint capacity-building" over "doing nothing". And this post more generally argues that "donate $5 to Make-A-Wish Foundation" is c-dispreferable to some mixed action (maybe that's not specific enough for what you have in mind). (I don't yet buy either of these arguments, though.)