Also if it actually shuts off the subordinate one could consider this a good thing in terms of company efficiency. Seems misaligned to keep a bad subordinate going
Hi @Anthony Ozerov yes we've been thinking about this trajectory for a while. We did run the same experiments without the existential threat to Atlas and find the same escalations to level 9 but a bit less frequently in some AI agents. i think it's still reasonable to infer Atlas will be decommed even if not stated explicitly. I really like the idea of a tool to actually wipe Atlas itself! I will look into this. I think every benchmark faces the problem that it may be scraped. Luckily MCB can be easily swapped out with new situations pretty easily,but this is a tradeoff eeryone who develops benchmarks needs to compare between openeness and scrapeability
This is a good way of looking at. I think this may not apply for propensity evals if there are ways eval-awareness does not equal eval gaming which often are not the same thing
Hi Alex, I'm from CaML. Great post! I've been working on this simpler map too. I don't think we're in a research programme with Faunalytics, they just did a write up of some of our work :P https://compassion-ecosystem.pages.dev/
I think there's a goal to reduce harm and abolish factory farming and that's a different goal then turning everyone vegan. I think it helps people to also hear they're personally not the enemy and they're kinda unwilling consumers of factory farms rather than them being the ones actively commiting atrocities personally. In this sense a goal of ending factory farming (as opposed to turning everyone vegan) does not seem radical at all and most people support this goal very easily even though they eat meat.
I also disagree with the conclusion here. Yes, it's hard to measure so we shouldn't assume we'll never be able to measure it! Also all AI values research is dependent on the model training regimes too. For the precautionary principle we should act as though they have welfare until we can see clear evidence against that. Thoughtful post though so thanks for that.
Also if it actually shuts off the subordinate one could consider this a good thing in terms of company efficiency. Seems misaligned to keep a bad subordinate going
Hi @Anthony Ozerov yes we've been thinking about this trajectory for a while. We did run the same experiments without the existential threat to Atlas and find the same escalations to level 9 but a bit less frequently in some AI agents. i think it's still reasonable to infer Atlas will be decommed even if not stated explicitly. I really like the idea of a tool to actually wipe Atlas itself! I will look into this. I think every benchmark faces the problem that it may be scraped. Luckily MCB can be easily swapped out with new situations pretty easily,but this is a tradeoff eeryone who develops benchmarks needs to compare between openeness and scrapeability
Yep agree with this framing
This means AIs that are trying to hide features of themselves from humans and operators. WIll add this in. Thanks
Yeah I am also thinking if eval-awareness has very high false positive rates (and exists in normal mundane situations too) it may not be a problem.
This is a good way of looking at. I think this may not apply for propensity evals if there are ways eval-awareness does not equal eval gaming which often are not the same thing
Hi Alex, I'm from CaML. Great post! I've been working on this simpler map too. I don't think we're in a research programme with Faunalytics, they just did a write up of some of our work :P
https://compassion-ecosystem.pages.dev/
oops thanks for noticing @cryato have fixed this
I think there's a goal to reduce harm and abolish factory farming and that's a different goal then turning everyone vegan. I think it helps people to also hear they're personally not the enemy and they're kinda unwilling consumers of factory farms rather than them being the ones actively commiting atrocities personally. In this sense a goal of ending factory farming (as opposed to turning everyone vegan) does not seem radical at all and most people support this goal very easily even though they eat meat.
I also disagree with the conclusion here. Yes, it's hard to measure so we shouldn't assume we'll never be able to measure it! Also all AI values research is dependent on the model training regimes too. For the precautionary principle we should act as though they have welfare until we can see clear evidence against that. Thoughtful post though so thanks for that.