As I understand it, what’s proposed here relies on application scores, and those scores are generated by panels. In my experience, panels can quite easily use the same label to score different criteria without anyone noticing. Definitions can drift over time, and new panellists don’t come with the tacit knowledge predecessors built. A while ago I ran a technical hiring round which failed: when running the post-mortem, it was clear panellists had different ideas of what a strong candidate looked like, which the final score hid. Before running another round, the panel documented the criteria explicitly and agreed how much each one mattered, using pairwise comparison. The discussion turned out to be more useful than the scores.
This was in a large company, with a stable panel and written processes. I’d expect this to be a bigger issue in orgs using volunteer mentors with different research interests, and panels that change between rounds, e.g. due to availability.
There could be a quick way to test the above: if an application was scored by more than one panellist, look to see how far apart the scores were. Also, ask a recent round’s panellists to write down, independently, what they were scoring against and what mattered most, and review how well those match. Neither needs any data on what applicants went on to do.
Note, my background is hiring and managing technical people in industry, not AI safety, so YMMV.
I want to add one thing from the hiring side:
As I understand it, what’s proposed here relies on application scores, and those scores are generated by panels. In my experience, panels can quite easily use the same label to score different criteria without anyone noticing. Definitions can drift over time, and new panellists don’t come with the tacit knowledge predecessors built. A while ago I ran a technical hiring round which failed: when running the post-mortem, it was clear panellists had different ideas of what a strong candidate looked like, which the final score hid. Before running another round, the panel documented the criteria explicitly and agreed how much each one mattered, using pairwise comparison. The discussion turned out to be more useful than the scores.
This was in a large company, with a stable panel and written processes. I’d expect this to be a bigger issue in orgs using volunteer mentors with different research interests, and panels that change between rounds, e.g. due to availability.
There could be a quick way to test the above: if an application was scored by more than one panellist, look to see how far apart the scores were. Also, ask a recent round’s panellists to write down, independently, what they were scoring against and what mattered most, and review how well those match. Neither needs any data on what applicants went on to do.
Note, my background is hiring and managing technical people in industry, not AI safety, so YMMV.