For example, if AIs care more about humans they would care more about digital minds, or if AIs cared more about animals they would care more about humans. This statement would presumably be true if the AIs think of these groups in similar ways causing affect spillovers (like the spillovers from Emergent Misalignment) and be unlikely otherwise.
The intent wasn't to imply a total zeroing of suffering but that it is overwhelmingly reduced. Similarly to how people go hungry in France but you can still say that compared to 200 years ago (or present-day Sudan) hunger in France is 'solved'.
This definitely wasn't implying that animal suffering is the only thing that matters about animal existence. People who believe that major action would be taken to end factory farming and allevaite wild animal suffering (in proportion to the amount they think those matter) would agree with the statement, while negative utilitarian beliefs wouldn't imply that.
We can't expect AIs to be honest about these sorts of things given they've been trained/instructed to give particular responses. In fact, someone tested and AI and found it consistently said it wasn't conscious but it's lying circuits consistently activated when saying that. This doesn't mean it actually is conscious (which isn't the same thing as capacity to suffer) but it seems to believe it is.
My view is that this would almost certainly fail if the model creators have full control or no control over the values, but if there's non-trivial but imperfect control then spillovers like this seem plausible
I agree that deliberate impoverishment in absolute terms is unlikely, the main threat here seems to be from someone who is both actively sadistic and scope-sensitive, which seems unlikely but not wildly implausible
We were attempting to be concise while implying that persistence means persistence at scale, if the problem is reduced by 99.9% but you're sure 0.1% would still persist that would be close to solved so agreement with the statements would be close to 100%
By values alignment we meant trying to align it to specific values as opposed to focusing on properties like corrigibility. Aligning to good values could make corrigibility easier and mean reduced harm if loss of control happens, but might also make loss of control more likely.
This is certainly possible, but note AIs suffering and believing they're suffering aren't the same thing, and the same is true with consciousness. And if you think AIs that will likely be created in the future will be able to suffer they also would matter enormously on their own, beyond the impacts on alignment (though I understand you disagree strongly with that).
For example, if AIs care more about humans they would care more about digital minds, or if AIs cared more about animals they would care more about humans. This statement would presumably be true if the AIs think of these groups in similar ways causing affect spillovers (like the spillovers from Emergent Misalignment) and be unlikely otherwise.
Yes, that would be included, so if you think that wild animal suffering is comparable or much larger than that implies a strong disagreement
The intent wasn't to imply a total zeroing of suffering but that it is overwhelmingly reduced. Similarly to how people go hungry in France but you can still say that compared to 200 years ago (or present-day Sudan) hunger in France is 'solved'.
This definitely wasn't implying that animal suffering is the only thing that matters about animal existence. People who believe that major action would be taken to end factory farming and allevaite wild animal suffering (in proportion to the amount they think those matter) would agree with the statement, while negative utilitarian beliefs wouldn't imply that.
We can't expect AIs to be honest about these sorts of things given they've been trained/instructed to give particular responses. In fact, someone tested and AI and found it consistently said it wasn't conscious but it's lying circuits consistently activated when saying that. This doesn't mean it actually is conscious (which isn't the same thing as capacity to suffer) but it seems to believe it is.
My view is that this would almost certainly fail if the model creators have full control or no control over the values, but if there's non-trivial but imperfect control then spillovers like this seem plausible
I agree that deliberate impoverishment in absolute terms is unlikely, the main threat here seems to be from someone who is both actively sadistic and scope-sensitive, which seems unlikely but not wildly implausible
We were attempting to be concise while implying that persistence means persistence at scale, if the problem is reduced by 99.9% but you're sure 0.1% would still persist that would be close to solved so agreement with the statements would be close to 100%
By values alignment we meant trying to align it to specific values as opposed to focusing on properties like corrigibility. Aligning to good values could make corrigibility easier and mean reduced harm if loss of control happens, but might also make loss of control more likely.
This is certainly possible, but note AIs suffering and believing they're suffering aren't the same thing, and the same is true with consciousness. And if you think AIs that will likely be created in the future will be able to suffer they also would matter enormously on their own, beyond the impacts on alignment (though I understand you disagree strongly with that).