one should never be required to strictly c-prefer a mixture of A and B to C whenever neither A nor B is c-preferred to C
There are candidate counterexamples to this claim. For example, imagine A is giving a benefit to Amy, and B and C each designate the same action of giving a benefit to Bobby. Then if you're impartial, you won't strictly c-prefer either of A or B to C, but you might strictly c-prefer the 50:50 mixture AB to C on the basis that it's fairer to randomize who gets the benefit.
Also, imprecise consequentialism (plus Dissent, Unanimity, and Justification) has an even more counterintuitive implication than 'you can be required to strictly c-prefer a mixture of AB to C even though neither A nor B is c-preferred to C.' It implies:
You can be required to strictly c-prefer a mixture of A, B, and C to D, even though (i) neither A nor B is c-preferred to D, (ii) C is c-dispreferred to D, and (iii) the mixture has an arbitrarily high probability of resulting in C.
Here's an example to illustrate:
I made the mixture have a 60% chance of C just to avoid the diagram being all bunched up. But the steeper you make the diagonals A and B, the higher you can push the probability of C and yet still have the mixture dominate D.
Thanks! Yeah, good question. The view I've sketched will say that A is c-impermissible, but we can still say that A is all-things-considered permissible after taking moral uncertainty into account. As an analogy, declining to push someone in front of a trolley to save 5 people is c-impermissible, but can be all-things-considered permissible after taking moral uncertainty into account.
On your second point, I think the thing I wrote in reply to Jim Buhler applies:
Basically I think the 'flanking variants' test will rule out a lot of actions but not all of them. In general, if our set of available actions is finite, then for each probability function there must be some action that has highest EV (and the same holds for infinite action sets modulo some small complications).
In practice, I think the easiest ones to identify will be extreme actions, like donating all your money to charity. Then we have some reason to think that the action is higher EV than its flanking variants, e.g. it's higher EV than giving less money, and it's impossible to give any more.
Of course, if we start individuating actions finely enough, then the flanking variants test might start ruling out even extreme actions. For example, it might be implausible that ideal reflection on your current evidence could lead you to judge that donating all your money at time X is higher EV than both (i) donating all your money one millisecond earlier and (ii) donating all your money one millisecond later. That would suggest donating all your money at time X is nowhere-optimal, in which case it's dominated by some mixed action, in which case it's c-impermissible.
But note that acting c-impermissibly in this way seems like an inevitable product of our cognitive limitations. So acting c-impermissibly in this way seems at least more forgiveable than choosing actions that are nowhere-optimal (and hence dominated) even on a coarse-grained individuation of our action space.
On your third point, I agree it seems kinda implausible to think you're required to strictly c-prefer the AB mixture to C even though neither A nor B is strictly c-preferred to C, but as you say denying it will have costs. We can read off one cost from my argument: you'll have to deny Dissent, Unanimity, or Justification.
Thanks! I guess given a choice between somewhere-uniquely-optimal actions, you should first check if any of them are ruled out by non-consequentialist considerations (like deontological constraints). If you still have multiple options at that stage, then the view I've sketched here will say that they're each permissible and that you can choose between them how you like. In practice, I'd try to choose the action that's somewhere-uniquely-optimal on my forced-to-say-a-number probability function, for moral uncertainty sorts of reasons.
Thanks! I wrote a reply to this but I now see it failed to send and I can't seem to recover the text.
Basically I think the 'flanking variants' test will rule out a lot of actions but not all of them. In general, if our set of available actions is finite, then for each probability function there must be some action that has highest EV (and the same holds for infinite action sets modulo some small complications).
In practice, I think the easiest ones to identify will be extreme actions, like donating all your money to charity. Then we have some reason to think that the action is higher EV than its flanking variants, e.g. it's higher EV than giving less money, and it's impossible to give any more.
Of course, if we start individuating actions finely enough, then the flanking variants test might start ruling out even extreme actions. For example, it might be implausible that ideal reflection on your current evidence could lead you to judge that donating all your money at time X is higher EV than both (i) donating all your money one millisecond earlier and (ii) donating all your money one millisecond later. That would suggest donating all your money at time X is nowhere-optimal, in which case it's dominated by some mixed action, in which case it's c-impermissible.
But note that acting c-impermissibly in this way seems like an inevitable product of our cognitive limitations, so it seems at least more forgiveable than choosing actions that are nowhere-optimal (and hence dominated) even on a coarse-grained individuation of our action space.
If you think you're justified in c-preferring the SWP donation, then Anthony's claim is disproved. But Anthony has said he doesn't consider this a counterexample, so it seems unlikely that offering more candidate counterexamples would move the debate forward.
That's especially so since I think the scope of Anthony's claim is intended to rule out other candidate counterexamples, e.g.:
You're trapped in a box that you know for sure will implode in 10 seconds. There's a puppy in there with you. You're justified in c-preferring not kicking the puppy to kicking the puppy.
That seems true to me, but I think this pair of actions falls outside the scope of Anthony's claim. He's talking about actions with effects that aren't so tightly limited in space and time.
So the debate calls for something more than just a bare counterexample. As Anthony says in another comment, I try to give that 'something more' in my post. Donating $5 to MAWF is justifiably c-dispreferred to some mixed action, as is any other action that fails the 'flanking variants' test. That likely rules out almost all actions.
Under maximality, we can't even say that [1, 10 000] is better than [-10 000, 1.0001].
That's not quite right. Maximality says an action A is impermissible when some alternative B has higher EV on every probability function in your representor. And that can be true even when A's and B's EV ranges overlap.
Example:
If A had the same range but sloped the other way, then it would be permissible by maximality:
So to figure out what's permissible under maximality, we can't just look at ranges. We need to look at the representor.
Thanks! I think the difference between our proposal and organizations paying workers to report themselves is the order that things come in. In our case, it's payment first and reporting later. The idea is that you could give a risk-averse AI $1,000 (say) and tell it that it can spend the money on anything it wants. Then so long as the risk-averse AI believes that the situation is real, it will spend the money on what it actually wants (e.g. making paperclips) and thereby reveal to you what it actually wants.
I think a human analogy would be something like:
You have an employee that claims that their only terminal value is the success of the company.
You want to find out whether this is true.
So you give the employee a no-strings-attached $10m.
If they give that money back to the company and continue working just as hard, that's strong evidence that they really do only terminally value the success of the company.
If they spend the money on something else, that's strong evidence that they don't only value the success of the company.
I'm unsure whether we can successfully train ASIs to be reliably risk-averse, including far OOD. Our claim is just that the chances of success are high enough to make risk aversion worth pursuing as a line of defense. That's the case we try to make in section 10. See also my reply to Ryan's comment. I also think our chances of success are a bit higher for AIs that aren't yet ASIs, and if we succeed in making them risk-averse I think they could help a lot with aligning any later-arising ASIs, by doing this sort of stuff.
Replied over there! There are also a few more comments in the thread now.
I’m surprised you say deals would be way less useful. Can you say more? Here’s my current sense of things:
In favor of keeping AIs risk-neutral
Risk-neutral AIs are more likely to attempt takeover. And we want them to attempt takeover, because (i) they’re very unlikely to succeed, and (ii) if we catch them attempting takeover, we can use them to reduce future takeover risk, by doing research on them, using the evidence of their takeover attempt to motivate international coordination, etc.
Worries about that:
One worry is that we need the risk-neutral AIs to be somewhat likely to successfully take over, otherwise they wouldn’t even attempt takeover and we couldn’t catch them. Taking the numbers from Fabien’s post (which are illustrative but don’t seem off by OOMs), their chance of successful takeover has to be greater than 1%. That’s a fairly big risk we’d be taking in the near term.
It also seems like we’ll only get really convincing evidence of a takeover attempt (the kind that might motivate international coordination) if the attempt gets pretty far along. And conditional on the attempt getting pretty far along, the AI’s chances of successful takeover are likely even higher.
One benefit of risk-averse AI (if it works) is that we don’t have to take that near-term risk.
And then with regard to reducing future takeover risk, I agree that getting to do research on confirmed-misaligned AIs would be a big benefit (but we can also get that with risk-averse AIs). On international coordination, I think most caught takeover attempts probably wouldn’t leave evidence clear enough to motivate it. I think many of the points from this post apply. In reply to evidence of a takeover attempt, people will likely say: shutting down would be costly, AIs do weird stuff sometimes, it was an isolated screw-up by that specific AI company, it’s a doomer plot, etc. And the news might not even get out of the lab. Also it seems fairly likely that AIs can try to take over in a way that would look pretty ambiguous if they failed. They could invent some justification for why their actions were actually in humanity’s best interest, etc.
(Sidenote is that I’m interested in the implications of ‘We want near-future AIs to attempt takeover.’ If that’s true, it seems like the AI safety community should be doing radically different stuff to the stuff it’s currently doing.)
In favor of making AIs risk-averse
Risk-averse AIs are less likely to attempt takeover in the near term.
If that were the only benefit, then I think it’d be pretty unclear which of risk-neutral AIs and risk-averse AIs is better. But risk-averse AIs would likely have lots of other benefits too, potentially letting us reduce future takeover risk by a lot. We can pay them to:
Reveal misalignment.
One idea here is that we give risk-averse AIs a small amount to spend on whatever they want. Then if they spend it on making paperclips (etc.), we’ve got clear evidence of misalignment. We can then do research on these misaligned AIs and use the evidence to motivate international coordination, etc.
This evidence of misalignment we get from risk-averse AIs seems about as good for enabling research and motivating international coordination as the evidence we’d get from risk-neutral AIs attempting takeover. And to get this evidence from risk-averse AIs, we don't need to bait them into an (at least somewhat likely to succeed) takeover attempt and hope that we catch them.
Reveal collusion signals.
Stop sandbagging on easy-to-evaluate tasks.
Identify security vulnerabilities.
Monitor untrusted AIs.
Do alignment research. (Hard to evaluate, of course. We say a bit about this in section 4.2.)
Taken together, all this stuff we can buy from risk-averse AIs seems much better for reducing future takeover risk than catching risk-neutral AIs in a takeover attempt. And we can buy all this stuff from risk-averse AIs without running a significant risk that AIs actually succeed in their takeover attempt.
(I'll reply to the generalization point in another comment.)
Extra stuff:
There are candidate counterexamples to this claim. For example, imagine A is giving a benefit to Amy, and B and C each designate the same action of giving a benefit to Bobby. Then if you're impartial, you won't strictly c-prefer either of A or B to C, but you might strictly c-prefer the 50:50 mixture AB to C on the basis that it's fairer to randomize who gets the benefit.
Also, imprecise consequentialism (plus Dissent, Unanimity, and Justification) has an even more counterintuitive implication than 'you can be required to strictly c-prefer a mixture of AB to C even though neither A nor B is c-preferred to C.' It implies:
Here's an example to illustrate:
I made the mixture have a 60% chance of C just to avoid the diagram being all bunched up. But the steeper you make the diagonals A and B, the higher you can push the probability of C and yet still have the mixture dominate D.
Thanks! Yeah, good question. The view I've sketched will say that A is c-impermissible, but we can still say that A is all-things-considered permissible after taking moral uncertainty into account. As an analogy, declining to push someone in front of a trolley to save 5 people is c-impermissible, but can be all-things-considered permissible after taking moral uncertainty into account.
On your second point, I think the thing I wrote in reply to Jim Buhler applies:
On your third point, I agree it seems kinda implausible to think you're required to strictly c-prefer the AB mixture to C even though neither A nor B is strictly c-preferred to C, but as you say denying it will have costs. We can read off one cost from my argument: you'll have to deny Dissent, Unanimity, or Justification.
Thanks! I guess given a choice between somewhere-uniquely-optimal actions, you should first check if any of them are ruled out by non-consequentialist considerations (like deontological constraints). If you still have multiple options at that stage, then the view I've sketched here will say that they're each permissible and that you can choose between them how you like. In practice, I'd try to choose the action that's somewhere-uniquely-optimal on my forced-to-say-a-number probability function, for moral uncertainty sorts of reasons.
Thanks! I wrote a reply to this but I now see it failed to send and I can't seem to recover the text.
Basically I think the 'flanking variants' test will rule out a lot of actions but not all of them. In general, if our set of available actions is finite, then for each probability function there must be some action that has highest EV (and the same holds for infinite action sets modulo some small complications).
In practice, I think the easiest ones to identify will be extreme actions, like donating all your money to charity. Then we have some reason to think that the action is higher EV than its flanking variants, e.g. it's higher EV than giving less money, and it's impossible to give any more.
Of course, if we start individuating actions finely enough, then the flanking variants test might start ruling out even extreme actions. For example, it might be implausible that ideal reflection on your current evidence could lead you to judge that donating all your money at time X is higher EV than both (i) donating all your money one millisecond earlier and (ii) donating all your money one millisecond later. That would suggest donating all your money at time X is nowhere-optimal, in which case it's dominated by some mixed action, in which case it's c-impermissible.
But note that acting c-impermissibly in this way seems like an inevitable product of our cognitive limitations, so it seems at least more forgiveable than choosing actions that are nowhere-optimal (and hence dominated) even on a coarse-grained individuation of our action space.
Yep, that's right!
I don't think that's so surprising. There are obvious candidate counterexamples, e.g.:
If you think you're justified in c-preferring the SWP donation, then Anthony's claim is disproved. But Anthony has said he doesn't consider this a counterexample, so it seems unlikely that offering more candidate counterexamples would move the debate forward.
That's especially so since I think the scope of Anthony's claim is intended to rule out other candidate counterexamples, e.g.:
That seems true to me, but I think this pair of actions falls outside the scope of Anthony's claim. He's talking about actions with effects that aren't so tightly limited in space and time.
So the debate calls for something more than just a bare counterexample. As Anthony says in another comment, I try to give that 'something more' in my post. Donating $5 to MAWF is justifiably c-dispreferred to some mixed action, as is any other action that fails the 'flanking variants' test. That likely rules out almost all actions.
That's not quite right. Maximality says an action A is impermissible when some alternative B has higher EV on every probability function in your representor. And that can be true even when A's and B's EV ranges overlap.
Example:
If A had the same range but sloped the other way, then it would be permissible by maximality:
So to figure out what's permissible under maximality, we can't just look at ranges. We need to look at the representor.
Thanks! I think the difference between our proposal and organizations paying workers to report themselves is the order that things come in. In our case, it's payment first and reporting later. The idea is that you could give a risk-averse AI $1,000 (say) and tell it that it can spend the money on anything it wants. Then so long as the risk-averse AI believes that the situation is real, it will spend the money on what it actually wants (e.g. making paperclips) and thereby reveal to you what it actually wants.
I think a human analogy would be something like:
Replied over there! There are also a few more comments in the thread now.