Agree with the need. As you pointed out, implementing a safeguard into the model itself (either the raw model or the scaffold) is quite challenging for BAIMs, as most capable models are open-weight. Adversarial fine-tuning can easily strip out weight-implemented safeguards. This issue is especially pronounced for BAIMs compared to general-purpose models, due to their limited training data. For scaffolds, you can simply remove them. So most of the control levers for safety live outside the model, as you listed. But to calibrate "where and how much control is needed," we need rigorous empirical study on BAIMs, especially at the scaffold level, since that is where realized capability can be measured.
Agree with the need. As you pointed out, implementing a safeguard into the model itself (either the raw model or the scaffold) is quite challenging for BAIMs, as most capable models are open-weight. Adversarial fine-tuning can easily strip out weight-implemented safeguards. This issue is especially pronounced for BAIMs compared to general-purpose models, due to their limited training data. For scaffolds, you can simply remove them. So most of the control levers for safety live outside the model, as you listed. But to calibrate "where and how much control is needed," we need rigorous empirical study on BAIMs, especially at the scaffold level, since that is where realized capability can be measured.