Founder and Director, Q16 Public Benefit Corporation. Author of Inside Cyber Warfare: Mapping the Cyber Underworld (O'Reilly - 2009, 2011, 2024). Co-founder of the Whitefish Security Summit.
Coming to this two years late, but it bears directly on something I'm building, so: this is the clearest statement I've seen of the y-axis problem.
I run Q16, a public benefit corporation that is launching a quarterly, panel-based assessment of global AI capability (https://q16pbc.com/blog/volunteer-quarterly-ai-capability-assessment) We publish a 0–1.0 headline number, so your opening question would rightly be: what's my axis?
Two design choices, both of which this post sharpened for me. First, the number is not a reading on a capability scale. It's an aggregated probability judgment — the panel's collective credence about proximity to a defined threshold — so it inherits the scale properties of probability rather than pretending to interval units of "capability," which I agree don't exist. Second, the elicitation instrument doesn't ask panelists to rate capability in the abstract. It asks decomposed, concrete questions about what changed in the quarter, and the headline is derived from those answers. Your "deflate to specifics and argue in their native currency" is close to a description of the method; the difference is we then aggregate back up, and we publish the panel's divergence rather than averaging it away.
Where we remain exposed: threshold definitions carry the same construct-validity risk you describe, just moved one level up. That part is unsolved.
The first full cycle runs in October, with a practice round in September. If anyone wants to pressure-test the instrument before then: [email protected], or the recruitment post (published above).
The "dark side" is a perfect descriptor for the LW forum. :-D
Coming to this two years late, but it bears directly on something I'm building, so: this is the clearest statement I've seen of the y-axis problem.
I run Q16, a public benefit corporation that is launching a quarterly, panel-based assessment of global AI capability (https://q16pbc.com/blog/volunteer-quarterly-ai-capability-assessment) We publish a 0–1.0 headline number, so your opening question would rightly be: what's my axis?
Two design choices, both of which this post sharpened for me. First, the number is not a reading on a capability scale. It's an aggregated probability judgment — the panel's collective credence about proximity to a defined threshold — so it inherits the scale properties of probability rather than pretending to interval units of "capability," which I agree don't exist. Second, the elicitation instrument doesn't ask panelists to rate capability in the abstract. It asks decomposed, concrete questions about what changed in the quarter, and the headline is derived from those answers. Your "deflate to specifics and argue in their native currency" is close to a description of the method; the difference is we then aggregate back up, and we publish the panel's divergence rather than averaging it away.
Where we remain exposed: threshold definitions carry the same construct-validity risk you describe, just moved one level up. That part is unsolved.
The first full cycle runs in October, with a practice round in September. If anyone wants to pressure-test the instrument before then: [email protected], or the recruitment post (published above).
I'm announcing an open call for volunteers with expertise in the following areas to complete an assessment of AI's global frontier:
(1) Frontier AI capability evaluation
(2) AI safety including alignment and interpretability
(3) AI governance and policy
(4) ML research, compute, infrastructure, or security
(5) AI policy for national security, economics, or law
Individual scores are never published. We only report the aggregate; meaning the median and spread across all respondents.
If you qualify and want to participate, please contact me - [email protected].