Thanks Vasco. I think it would help a lot if you spelled out the premises more, because they're quite opaque to me as written. E.g. I don't know what "the norm reveals a choice function" means. (I think this kind of use of jargon without giving context is a common failure mode of current LLM summaries.)
Also, if I understand correctly, "behaviour maximizes a family of preference orderings..." means that the result only shows that we can represent an agent's behavior as satisfying completeness. But my unawareness argument isn't about what our behavior can be represented as. The question is: When we're comparing our options when making decisions in the first place, should we have complete preferences? Cf "Winning isn't enough":
But what these arguments really show is that you are disposed to playing a dominated strategy if we cannot model your behavior as if you were a Bayesian with a certain prior and utility function. They don’t say anything about the procedure by which you need to make your decisions. I.e., they don’t say that you have to write down precise probabilities, utilities, and make decisions by solving for the Bayes-optimal policy for those.
(But let me know if I've misunderstood the result.)
The standard isn't either (a) or (b) exactly. I think no one, precise Bayesian or otherwise, has a complete standard for how to set credences. I'd gesture at something like the example in this comment: try setting credences in a similar way to precise Bayesians, but whenever you find that it seems arbitrary which distribution you pick among many, include them all.
So insofar as I understand what is meant by "a distribution needs a sponsor", I'm not committed to (b) by virtue of agreeing that we should have P(grey) = 1/2 in Joyce's case.
In particular:
I think what's going on in your last paragraph is an equivocation between:
"If you don't assign a precisely symmetric distribution over some set of hypotheses H, it must be because there is some respect in which the hypotheses are not symmetric."
"If you don't assign a precisely symmetric distribution over H, it must be because you're explicitly aware of a pair of hypotheses in H that are not symmetric."
(1) seems very plausible. But (2) isn't. It can be the case that I'm not aware of the hypotheses in H, yet I have reasons (based on the arguments given in sections 3.2.1 and 4.1.1) to consider them not symmetric.
currently, probably[1] we're clueless about the comparison of any two actions
it's plausible that, if we had a lot more evidence and more developed conceptual models of the cosmos, we'd be non-clueless about the comparisons of some actions. I can't say "how often" this would occur, in the sense of how easy it would be to get such evidence and models. Seems really hard; see here.
I mean that we have what I call "coarse awareness" here: we conceive of crude groups of possible worlds, rather than possible worlds specified in fine-grained enough detail to assign them precise values (wrt impartial altruist axiologies). See also here for some examples. Happy to unpack more if those sections don't answer things!
Good question! I think "other theoretically possible aggregations of all or most of the possible consequences of A and B" would also suffice, yeah. (Of course, if we ourselves can't specify what this alternative is, we have our work cut out for us if we're gonna argue that we should expect our idealized self to prefer A over B on this basis.)
Not a comprehensive reply, but: I think many of the examples you're talking about are arguably cases of coarse awareness. People were coarsely aware of the potential backfire risks earlier on, but (arguably) the reason they didn't give these risks enough weight was that they didn't have a more fine-grained awareness of the specific causal pathways. I think such cases count as evidence for the pessimistic induction.
I'd guess we're getting slightly better, yep. I might put less weight on the evidence you mention, than on: "We're living in a period of really unprecedented AI progress, seems like that puts in a better position to reason about the mechanisms governing the far future than ever before."
I think it's mostly (1), but I'm open to something like (2) or (3) as well.
(Following (1):) There is in principle some (a) amount of information that non-ideal agents could attain about the cosmos with non-Pascalian probability,[1] + (b) a priori modeling and induction we could apply to that information, such that we wouldn't be clueless. So I don't think we need to observe the target variable, or empirically "validate" the theory, to be non-clueless.
But the bar to achieve such an (a)+(b) seems very high, because:
If we do try to empirically validate the theory by appealing to calibration on near-term proxies:
I indeed don't see why we should expect such calibration to transfer, up to the degree of precision we need to escape cluelessness (sec. 2.3.1.1). This bites even if, say, we use AI to get much more calibrated on ~years-long time horizons.
If we don't, and instead try to argue conceptually that the theory captures enough of the relevant considerations in fine-grained enough detail:
The web of factors this theory would have to capture seems ludicrously complex (the rest of sec. 2.3). Of course, good theories can compress complexity, but getting that amount of compression while keeping things computationally tractable[2] sounds rough.
So my suspicion is that yeah, we'd still be clueless given the kind of theory you mention. But I find it hard to say, because I can't imagine exactly what "comparably good" looks like, concretely. I appreciate that that's hard to spell out on your end.
Maybe sufficiently advanced AI could get around this. Maybe not, e.g. if "the universal prior" is irreducibly imprecise, or if (following (3)) information about simulators or causally disconnected worlds is fundamentally inaccessible.
(I'm happy to unpack any of this more if useful, not sure if I answered your question properly!)
Thanks Vasco. I think it would help a lot if you spelled out the premises more, because they're quite opaque to me as written. E.g. I don't know what "the norm reveals a choice function" means. (I think this kind of use of jargon without giving context is a common failure mode of current LLM summaries.)
Also, if I understand correctly, "behaviour maximizes a family of preference orderings..." means that the result only shows that we can represent an agent's behavior as satisfying completeness. But my unawareness argument isn't about what our behavior can be represented as. The question is: When we're comparing our options when making decisions in the first place, should we have complete preferences? Cf "Winning isn't enough":
(But let me know if I've misunderstood the result.)
The standard isn't either (a) or (b) exactly. I think no one, precise Bayesian or otherwise, has a complete standard for how to set credences. I'd gesture at something like the example in this comment: try setting credences in a similar way to precise Bayesians, but whenever you find that it seems arbitrary which distribution you pick among many, include them all.
So insofar as I understand what is meant by "a distribution needs a sponsor", I'm not committed to (b) by virtue of agreeing that we should have P(grey) = 1/2 in Joyce's case.
In particular:
I think that from an impartial POV:
Referring to my uncertainty about the logical implications of our evidence for whether or not P3 is true.
Unfortunately I don't think there's really much of a case for "maintain option value, build capacity etc." being robustly good either, as argued here.
I mean that we have what I call "coarse awareness" here: we conceive of crude groups of possible worlds, rather than possible worlds specified in fine-grained enough detail to assign them precise values (wrt impartial altruist axiologies). See also here for some examples. Happy to unpack more if those sections don't answer things!
Good question! I think "other theoretically possible aggregations of all or most of the possible consequences of A and B" would also suffice, yeah. (Of course, if we ourselves can't specify what this alternative is, we have our work cut out for us if we're gonna argue that we should expect our idealized self to prefer A over B on this basis.)
Interesting, that's helpful to know.
Not a comprehensive reply, but: I think many of the examples you're talking about are arguably cases of coarse awareness. People were coarsely aware of the potential backfire risks earlier on, but (arguably) the reason they didn't give these risks enough weight was that they didn't have a more fine-grained awareness of the specific causal pathways. I think such cases count as evidence for the pessimistic induction.
Thanks Vasco. I've summarized my reply on LessWrong here (figured that this might be of (more?) interest to LW readers).
I'd guess we're getting slightly better, yep. I might put less weight on the evidence you mention, than on: "We're living in a period of really unprecedented AI progress, seems like that puts in a better position to reason about the mechanisms governing the far future than ever before."
I think it's mostly (1), but I'm open to something like (2) or (3) as well.
(Following (1):) There is in principle some (a) amount of information that non-ideal agents could attain about the cosmos with non-Pascalian probability,[1] + (b) a priori modeling and induction we could apply to that information, such that we wouldn't be clueless. So I don't think we need to observe the target variable, or empirically "validate" the theory, to be non-clueless.
But the bar to achieve such an (a)+(b) seems very high, because:
So my suspicion is that yeah, we'd still be clueless given the kind of theory you mention. But I find it hard to say, because I can't imagine exactly what "comparably good" looks like, concretely. I appreciate that that's hard to spell out on your end.
Maybe sufficiently advanced AI could get around this. Maybe not, e.g. if "the universal prior" is irreducibly imprecise, or if (following (3)) information about simulators or causally disconnected worlds is fundamentally inaccessible.
(I'm happy to unpack any of this more if useful, not sure if I answered your question properly!)
As in, if I were to represent this probability numerically, the interval wouldn't all be less than the Pascalian threshold.
Like, something analogous to the Schrödinger equation doesn't count. :)