This reasoning threatens to prove too much by generalizing from cluelessness to the case of ordinary uncertainty. Suppose A directly saves a life, while its long-run effect is either an enormous benefit or an enormous harm to an astronomical number of people, each with a well-justified probability of exactly ½. Subjective expected-value reasoning says to choose A: The long-run effects cancel out in expectation, even though in actuality A may deterministically lead to the enormous harm and B to the enormous benefit. “Properly disentangled,” the effects would give us reason to favor B (though we don’t know this). So why, we might ask, consider the situation as morally urgent as if the effects canceled out (in actuality, not merely in expectation)?
But if that does not trouble us, it is unclear why indeterminacy should. While exact cancellation and bracketing differ epistemically—the former establishes an expected difference of zero, while the latter establishes no determinate direction—both leave us without a long-run contrastive reason for either option. Under both exact cancellation and bracketing, the moral importance of A derives from the determinate reason provided by the person saved.
I think this is a fair point. But here's where I was coming from in the quote you respond to here.
First, in that context I was taking the fundamental contrastive reasons-givers to be person-moments, not individual persons. From that perspective, the problem is that we do have long-run contrastive reasons for each option, which we don't know how to weigh up. I'd agree that if the contrastive reasons-givers are persons, each person whose welfare we're clueless about gives us no contrastive reasons.
Second, I think we should separate two claims:
If I'm forced to make a choice purely based on an impartial weighing of consequences, I only have contrastive reasons favoring A, and that reason is saving a life.
In my all-things-considered decision-making, I should give the sameweight to "consequentialist bracketing says I should choose A, because of (1)" that I would've given to "the subjective EV of choosing A is equivalent to the value of saving a life".
If the contrastive reasons-givers are persons, I agree with (1). But when I expressed doubt about the "moral urgency" of bracketing, I was doubting (2). To my metanormative intuitions, it just doesn't feel like a bracketing-based verdict in favor of A has as much weight as the analogous EV-based verdict. (I could try to say more on why, if that's helpful, though it's hard to articulate.)
Still, I find it hard to say whether bracketing-based verdicts have more weight than, say, "precise consequentialism says I should instead do [galaxy-brained thing]". Something that helps me probe my intuitions here is:
Let's say I'm instead just making decisions based on my + my loved ones' welfare over our lifetimes, not all sentient beings.
I think I'm pretty plausibly clueless about my actions' impact on our lifetime welfare.
But consider some form of bracketing where the contrastive reasons-givers are units of time within each of my + my loved ones' lives. When I think about concrete cases like "should I help my loved one avoid this source of totally unnecessary near-term pain?", the bracketing-based verdict "yes" feels quite weighty! (Relative to "precise consequentialism about my + my loved ones' welfare says I should do [galaxy-brained thing]".)
That seems like a big leap to me, and I don’t see how it follows or what justifies it
The sequence gives general arguments for this, especially sections 2.3, 3.2, and 4.1. I'm not exactly sure what you find uncompelling about them. The worry is that severe coarseness makes the degree of justification so severely vague that we have all the same qualitative problems as under the maximality analysis.
Of course, the claim is defeasible by arguments to the contrary for some specific A vs. B. But the burden of proof seems quite high to me.
I think no. Basically, when I really internalize how dwarfed every action's cosmic-scale consequences are by off-target effects, "no" feels very common-sensical to me. (Cf. this paper on how "simple cluelessness" is fake.)
I think I wouldn't be clueless about c-preferability in Elliott's "trapped in a box" example, but can't think of any real-world case analogous to this.
I don’t think the arguments against the capacity-building strategies I discuss are as strong as those in their favor
My core objection to capacity-building in the sequence is: For any concrete capacity-building strategy, our understanding of that strategy's full range of possible consequences is extremely coarse. And it seems very plausible to me that these consequences will include large off-target effects on, e.g., lock-in events — in which case, capacity-building strategies inherit the non-robustness of strategies aimed at influencing lock-in events. This is for pretty similar reasons to how the off-target effects of AMF donations seem to dominate. I don't yet see why you think otherwise.
(So in particular, I don't think your responses in your appendix to specific backfire risks I mentioned in the post address this core objection.)
Moreover, much of the value of building capacity is the value of being able to act on considerations we aren’t yet aware of. So, in my view, unawareness bears asymmetrically on these strategies rather than neutrally (I realize this latter point is stated very briefly and needs further development).
Yeah, I'd be interested in seeing this spelled out a lot more sometime. Per the above, even if a strategy might enable us to act on considerations we aren’t yet aware of, this doesn't help us with cluelessness if the strategy's impact is still very plausibly dominated by off-target effects.
I’m one of the judges of the competition. My comments shouldn't be taken as a full review of a post. And, unfortunately, I won’t have capacity to comment on every post or engage with all replies. Thanks so much to everyone who entered!
I'm confused by your responses to the vignettes and thought experiments you quote in this post.
The insensitivity to mild sweetening example isn't meant to be an argument for imprecision. It's in the section of the post that's focused on spelling out implications of imprecision.
Re: the pause AI example: I don't think your response engages with the core intuition the example is getting at, which is that giving a determinate answer seems arbitrary. I guess you don't share that intuition — fair enough — but it seems misleading to present the example as if it's intended to pump a brute intuition (of incompleteness) rather than the underlying intuition that justifies it.
Re: the vignette in unawareness post #1, you say: "That quote suggests that he doesn’t consider this vignette to have persuasive force." But I was saying that the vignette alone doesn't establish the whole conclusion. And why should it? (By that point in the sequence, I haven't gotten into all the rest of the argument.) The role of vignettes generally is to put concrete color on an otherwise abstract argument, not to substitute for the argument.
(Just to be clear, that's a contingent matter. I don't find any of the counterexamples offered so far persuasive because I don't think they adequately engage with my arguments for P3.)
Yeah I still disagree for this 4th example as well, for the same reasons as the 1st.
I worry about a motte-and-bailey implicitly happening here, where the motte is "all things considered, we should prefer the second action" — very difficult to deny! — and the bailey is "we should c-prefer the second action". This matters because the corresponding motte seems actually pretty easy to deny (IMO) when the comparison is "study some altruistically irrelevant branch of academic philosophy" vs. "try to prevent AI misalignment". The latter only looks clearly preferable to me if it's c-preferable. (I guess this is what Ben's comment is getting at.)
Some of the contest entries do at least briefly attempt that, I think. E.g. this post gives an argument that one should c-prefer "low-footprint capacity-building" over "doing nothing". And this post more generally argues that "donate $5 to Make-A-Wish Foundation" is c-dispreferable to some mixed action (maybe that's not specific enough for what you have in mind). (I don't yet buy either of these arguments, though.)
Alternatively, if you posit incommensurable values then you should probably reject P1.
It's a fair point that deference principles come into conflict with prospective reasons in Hare's case, and the prospective reasons argument seems really plausible there. I don't feel very confident, but here's how I'm thinking about this:
Deference to the idealized self isn't actually necessary for my argument.
I wrote P1 in those terms as one way of making precise the idea that we should account for hypotheses we're unaware of, when weighing up prospective reasons.[1]
But we could instead just directly say, as I do in the sequence: Insofar as we're impartial altruists, (1) we want to in some sense weigh up all possible outcomes by their value and plausibility; and (2) our values are defined over metaphysically possible outcomes, not over the extremely coarse-grained versions of such outcomes we conceive of. So when we (perhaps roughly, informally) compare actions as impartial altruists, we should do so in a way that tries to account for fine-grainings of possible outcomes that we haven't explicitly conceived of.
Still, I'm pretty sympathetic to the "deference to the idealized self" idea, so how do I reconcile this with Hare's case? Well, it's not clear to me that this particular deference principle is substantively analogousto the deference principle invoked in Hare's case, even though they're structurally similar.
The deference principle I'm invoking is, roughly: (D1) My current evidence doesn't warrant considering A c-preferable to B, from the perspective of an agent who assesses that evidence the way I'd want to if I didn't have my computational (etc.) limits. So I, in my current decision situation, shouldn't consider A c-preferable.
The principle in Hare's case is, roughly: (D2) If I had more evidence (thereby putting me in a different decision situation), I wouldn't be warranted in considering A c-preferable to B no matter what that evidence is. So I shouldn't consider A c-preferable.
It seems plausible to me that we should endorse D1 but not D2, because:
If you ask me why I don't c-prefer A over B, and I invoke D1, my answer only makes reference to features of the decision problem I'm actually in — it's just that those features are assessed from a less computationally limited perspective.
Whereas if you asked me why I don't c-prefer taking the sugar, and I invoked D2, I wouldn't be telling you why my actual decision problem fails to recommend taking the sugar. I'd be telling you that versions of myself in different problems don't c-prefer taking the sugar. They're more informed versions of myself, yes. But those versions of myself have different reasons!
(Compare to how it's coherent for an agent who endorses causal decision theory to two-box when "dropped into" Newcomb's problem, yet to commit to one-box in the future if the prediction of their decision hasn't yet been made. The former agent has different reasons than the latter.)
(Anyway, again, not confident in this, and I think the more important point is that the cluelessness argument as such doesn't depend on deference principles.)
In particular, this framing doesn't require that humans even have well-defined "hypotheses" in our epistemic state. I don't think the alternative framing I use in the sequence itself requires us to have unrealistically precisely defined hypotheses, either. But this was a bit of a sticking point when discussing the problem of unawareness with one thoughtful interlocutor — which is (AIUI) what led to Jesse writing thepost on deference to the idealized self that I'm drawing on.
I think this is a fair point. But here's where I was coming from in the quote you respond to here.
First, in that context I was taking the fundamental contrastive reasons-givers to be person-moments, not individual persons. From that perspective, the problem is that we do have long-run contrastive reasons for each option, which we don't know how to weigh up. I'd agree that if the contrastive reasons-givers are persons, each person whose welfare we're clueless about gives us no contrastive reasons.
Second, I think we should separate two claims:
If the contrastive reasons-givers are persons, I agree with (1). But when I expressed doubt about the "moral urgency" of bracketing, I was doubting (2). To my metanormative intuitions, it just doesn't feel like a bracketing-based verdict in favor of A has as much weight as the analogous EV-based verdict. (I could try to say more on why, if that's helpful, though it's hard to articulate.)
Still, I find it hard to say whether bracketing-based verdicts have more weight than, say, "precise consequentialism says I should instead do [galaxy-brained thing]". Something that helps me probe my intuitions here is:
The sequence gives general arguments for this, especially sections 2.3, 3.2, and 4.1. I'm not exactly sure what you find uncompelling about them. The worry is that severe coarseness makes the degree of justification so severely vague that we have all the same qualitative problems as under the maximality analysis.
Of course, the claim is defeasible by arguments to the contrary for some specific A vs. B. But the burden of proof seems quite high to me.
I think no. Basically, when I really internalize how dwarfed every action's cosmic-scale consequences are by off-target effects, "no" feels very common-sensical to me. (Cf. this paper on how "simple cluelessness" is fake.)
I think I wouldn't be clueless about c-preferability in Elliott's "trapped in a box" example, but can't think of any real-world case analogous to this.
Yeah, spacetime bracketing or some more formally nebulous form of bracketing along the lines here.
Thanks!
My core objection to capacity-building in the sequence is: For any concrete capacity-building strategy, our understanding of that strategy's full range of possible consequences is extremely coarse. And it seems very plausible to me that these consequences will include large off-target effects on, e.g., lock-in events — in which case, capacity-building strategies inherit the non-robustness of strategies aimed at influencing lock-in events. This is for pretty similar reasons to how the off-target effects of AMF donations seem to dominate. I don't yet see why you think otherwise.
(So in particular, I don't think your responses in your appendix to specific backfire risks I mentioned in the post address this core objection.)
Yeah, I'd be interested in seeing this spelled out a lot more sometime. Per the above, even if a strategy might enable us to act on considerations we aren’t yet aware of, this doesn't help us with cluelessness if the strategy's impact is still very plausibly dominated by off-target effects.
I’m one of the judges of the competition. My comments shouldn't be taken as a full review of a post. And, unfortunately, I won’t have capacity to comment on every post or engage with all replies. Thanks so much to everyone who entered!
I'm confused by your responses to the vignettes and thought experiments you quote in this post.
(Just to be clear, that's a contingent matter. I don't find any of the counterexamples offered so far persuasive because I don't think they adequately engage with my arguments for P3.)
Yeah I still disagree for this 4th example as well, for the same reasons as the 1st.
I worry about a motte-and-bailey implicitly happening here, where the motte is "all things considered, we should prefer the second action" — very difficult to deny! — and the bailey is "we should c-prefer the second action". This matters because the corresponding motte seems actually pretty easy to deny (IMO) when the comparison is "study some altruistically irrelevant branch of academic philosophy" vs. "try to prevent AI misalignment". The latter only looks clearly preferable to me if it's c-preferable. (I guess this is what Ben's comment is getting at.)
Some of the contest entries do at least briefly attempt that, I think. E.g. this post gives an argument that one should c-prefer "low-footprint capacity-building" over "doing nothing". And this post more generally argues that "donate $5 to Make-A-Wish Foundation" is c-dispreferable to some mixed action (maybe that's not specific enough for what you have in mind). (I don't yet buy either of these arguments, though.)
It's a fair point that deference principles come into conflict with prospective reasons in Hare's case, and the prospective reasons argument seems really plausible there. I don't feel very confident, but here's how I'm thinking about this:
In particular, this framing doesn't require that humans even have well-defined "hypotheses" in our epistemic state. I don't think the alternative framing I use in the sequence itself requires us to have unrealistically precisely defined hypotheses, either. But this was a bit of a sticking point when discussing the problem of unawareness with one thoughtful interlocutor — which is (AIUI) what led to Jesse writing the post on deference to the idealized self that I'm drawing on.