What's the clearest example of a complex cluelessness sign flip you're aware of?
(By "clear" I mean "had a very narrow confidence interval before encountering some consideration and a narrow interval after encountering that consideration but the CIs now center points with opposite signs".[1])
The clearest examples I know of (e.g. rescuing Hitler as a child) seem to me like examples of simple cluelessness. You list some examples here, but they don't seem that clear to me, e.g. I disagree that "Early awareness-raising about AGI x-risk presumably seemed robustly good" and would guess most people involved in that had CIs which comfortably straddled zero.
- ^
Or alternatively: there are two representors with narrow but non-overlapping CIs.
You respond to Richard Ngo here:
Suppose instead we had a comparably good theory of the right reference class, e.g. "movements trying to shape transformative technologies." Would we still be clueless about AI safety movement-building?
More generally: you list various considerations across your posts and I have a hard time understanding which is load-bearing for your answer here. Some possibilities:
I think it's mostly (1), but I'm open to something like (2) or (3) as well.
(Following (1):) There is in principle some (a) amount of information that non-ideal agents could attain about the cosmos with non-Pascalian probability,[1] + (b) a priori modeling and induction we could apply to that information, such that we wouldn't be clueless. So I don't think we need to observe the target variable, or empirically "validate" the theory, to be non-clueless.
But the bar to achieve such an (a)+(b) seems very high, because:
So my suspicion is that yeah, we'd still be clueless given the kind of theory you mention. But I find it hard to say, because I can't imagine exactly what "comparably good" looks like, concretely. I appreciate that that's hard to spell out on your end.
Maybe sufficiently advanced AI could get around this. Maybe not, e.g. if "the universal prior" is irreducibly imprecise, or if (following (3)) information about simulators or causally disconnected worlds is fundamentally inaccessible.
(I'm happy to unpack any of this more if useful, not sure if I answered your question properly!)
As in, if I were to represent this probability numerically, the interval wouldn't all be less than the Pascalian threshold.
Like, something analogous to the Schrödinger equation doesn't count. :)