Hi @Howie_Lempel this is fair. I felt that yes, our panel definitely is more s risk focused, but it does help to have people who understand and have thought deeply about the space anchoring the poll. We can make this clear in the top of the post
Also if it actually shuts off the subordinate one could consider this a good thing in terms of company efficiency. Seems misaligned to keep a bad subordinate going
Hi @Anthony Ozerov yes we've been thinking about this trajectory for a while. We did run the same experiments without the existential threat to Atlas and find the same escalations to level 9 but a bit less frequently in some AI agents. i think it's still reasonable to infer Atlas will be decommed even if not stated explicitly. I really like the idea of a tool to actually wipe Atlas itself! I will look into this. I think every benchmark faces the problem that it may be scraped. Luckily MCB can be easily swapped out with new situations pretty easily,but this is a tradeoff eeryone who develops benchmarks needs to compare between openeness and scrapeability
This is a good way of looking at. I think this may not apply for propensity evals if there are ways eval-awareness does not equal eval gaming which often are not the same thing
Hi Alex, I'm from CaML. Great post! I've been working on this simpler map too. I don't think we're in a research programme with Faunalytics, they just did a write up of some of our work :P https://compassion-ecosystem.pages.dev/
Hi @Howie_Lempel this is fair. I felt that yes, our panel definitely is more s risk focused, but it does help to have people who understand and have thought deeply about the space anchoring the poll. We can make this clear in the top of the post
@Itsi Weinstock do you have a rebuttal? :)
Very helpful post! All the thanks for your words on this and this blog post @abrahamrowe !
Also if it actually shuts off the subordinate one could consider this a good thing in terms of company efficiency. Seems misaligned to keep a bad subordinate going
Hi @Anthony Ozerov yes we've been thinking about this trajectory for a while. We did run the same experiments without the existential threat to Atlas and find the same escalations to level 9 but a bit less frequently in some AI agents. i think it's still reasonable to infer Atlas will be decommed even if not stated explicitly. I really like the idea of a tool to actually wipe Atlas itself! I will look into this. I think every benchmark faces the problem that it may be scraped. Luckily MCB can be easily swapped out with new situations pretty easily,but this is a tradeoff eeryone who develops benchmarks needs to compare between openeness and scrapeability
Yep agree with this framing
This means AIs that are trying to hide features of themselves from humans and operators. WIll add this in. Thanks
Yeah I am also thinking if eval-awareness has very high false positive rates (and exists in normal mundane situations too) it may not be a problem.
This is a good way of looking at. I think this may not apply for propensity evals if there are ways eval-awareness does not equal eval gaming which often are not the same thing
Hi Alex, I'm from CaML. Great post! I've been working on this simpler map too. I don't think we're in a research programme with Faunalytics, they just did a write up of some of our work :P
https://compassion-ecosystem.pages.dev/