Thanks! Not sure I agree with "heuristic": Stockfish's stopping rule is a heuristic, but it's backed by a theorem (Cantelli) that holds with no distributional assumptions. What is distribution-dependent is the estimate of that feeds into it, and I agree that some skepticism is warranted there.
The chess analogy is meant to show the framework is implementable, not merely theoretical. That moves the question from "are expected values well-defined?" to "what is and how large is relative to it?" I consider that progress because the answer isn't uniformly "suspend judgment": some parameter values license acting, others clearly don't.
Thanks! The Cantelli bound is a basic result of probability theory, so I think it's hard to change chess such that it would no longer apply at all.
But an experiment I would be interested to do is to run an engine with varying parameters and see how much this changes performance. If the "deliberate until you don't see many sign flips" approach only works because we have a precisely-tuned definition of "many", then I do think this weakens the analogy.
Thanks! Tbc, I agree that there are persistence effects; I just don't think they are so strong that we can reasonably expect that slight changes in starting conditions result in a completely different world.
if we were to travel back in time and alter starting conditions just slightly, it seems reasonable to expect that the world today would be completely different.
a few short years after the bombs stopped falling in 1945, the world economy returned to trend as if nothing had happened.
In 1776, America rebelled in the name of freedom and democracy: the origin myth of the modern world order. And yet, somehow, unrebellious Canada ended up just as free and democratic. An unrebellious America likely would have too.
Thanks for writing this Sarah and best wishes for the new position!
I have been pretty pleased with the 2026 Forum output (particularly the unawareness event caused me to think more about my own work than most other things, maybe more than any other online event in 2026). Kudos to the rest of the team, and to your leadership for enabling that.
To see this, let’s be generous and say the expected number of future well-off people added by preventing an existential catastrophe is only 10^40
I think your argumentation supports the 10^40 number more than the "well-off" claim. I'm not sure for the best canonical source expressing skepticism of the "well-off" bit, but The Future Might Not Be So Great contains a long list of arguments about whether future people will be well off; I would be interested in you responding to some of those.
Anthony cites Greaves and MacAskill giving an example similar to your gunpowder one:
Consider, for example, would-be longtermists in the Middle Ages. It is plausible that the considerations most relevant to their decision – such as the benefits of science, and therefore the enormous value of efforts to help make the scientific and industrial revolutions happen sooner – would not have been on their radar. Rather, they might instead have backed attempts to spread Christianity, perhaps by violence: a putative route to value that, by our more enlightened lights today, looks wildly off the mark. The suggestion, then, is that our current predicament is relevantly similar to that of our medieval would-be longtermists.
I personally think these examples are less compelling than they first appear (e.g. the persistence literature generally finds weaker effects than what you might imagine), but I agree that a failure of EAs to find examples of sign flips doesn't mean that future ones won't exist.
The stochastic dominance point is helpful, ty
Thanks! Not sure I agree with "heuristic": Stockfish's stopping rule is a heuristic, but it's backed by a theorem (Cantelli) that holds with no distributional assumptions. What is distribution-dependent is the estimate of that feeds into it, and I agree that some skepticism is warranted there.
The chess analogy is meant to show the framework is implementable, not merely theoretical. That moves the question from "are expected values well-defined?" to "what is and how large is relative to it?" I consider that progress because the answer isn't uniformly "suspend judgment": some parameter values license acting, others clearly don't.
Thanks! The Cantelli bound is a basic result of probability theory, so I think it's hard to change chess such that it would no longer apply at all.
But an experiment I would be interested to do is to run an engine with varying parameters and see how much this changes performance. If the "deliberate until you don't see many sign flips" approach only works because we have a precisely-tuned definition of "many", then I do think this weakens the analogy.
There are even AI-safety-pun-based sports teams!
Thanks! Tbc, I agree that there are persistence effects; I just don't think they are so strong that we can reasonably expect that slight changes in starting conditions result in a completely different world.
This is possible, but seems pretty unclear to me. cf The Gods of Straight Lines:
$35/attendee is an extremely impressive cost. Thanks for doing this and writing it up!
Thanks for writing this Sarah and best wishes for the new position!
I have been pretty pleased with the 2026 Forum output (particularly the unawareness event caused me to think more about my own work than most other things, maybe more than any other online event in 2026). Kudos to the rest of the team, and to your leadership for enabling that.
I think your argumentation supports the 10^40 number more than the "well-off" claim. I'm not sure for the best canonical source expressing skepticism of the "well-off" bit, but The Future Might Not Be So Great contains a long list of arguments about whether future people will be well off; I would be interested in you responding to some of those.
Anthony cites Greaves and MacAskill giving an example similar to your gunpowder one:
I personally think these examples are less compelling than they first appear (e.g. the persistence literature generally finds weaker effects than what you might imagine), but I agree that a failure of EAs to find examples of sign flips doesn't mean that future ones won't exist.