Gregory -- thank you for the interesting post. I admit I have *read most of it twice and not fully understood. May I ask two clarifying questions? Specifically on ECI.
First: you say link function choice is "relatively minor" (rank and ECI correlate ~0.97), but you also raise an endogeneity point, which "entirely relies on the link function." I think the first claim holds link functions vary, shared-shape assumption stays fixed; the second is about dropping the shared-shape assumption itself. You say a shared-shape assumption is "dubious" but I'm unsure what your reasons are?
Second: I'm wondering what the empirical test(s) is for your critiques. In general, if you think measures of capabilities are not robust in some way, what is the sensitivity analysis? I am thinking like an economist here.
One idea -- refit ECI under different (common) link functions, see if the "speeding up" trend survives. I don't think Epoch runs this, though their appendix comes close, although I don't think this would be a crux for you.
Two other ideas that came to mind, specifically on 'measure endogeneity':
Fit using only benchmarks that predate the model being scored.
Let discrimination depend on whether the benchmark came before or after the model.
FYI - there is an excellent Sentience Institute Report from 2018 which treated this issue in a lot of detail: https://www.sentienceinstitute.org/gm-foods
There's lots of points here. While they are possible, I would suggest they are not particularly common/well-suported in the psychological literature as it is today.
In addition, I don't know why these explanations would lead to desensitisation towards positive and negative events.
I can't think of a good theoretical reason why true effects should fall so significantly – like 40%. That's striking. The same attenuation result holds, even including income/age/event prevalence.
"Intuitively, if wellbeing saturates at the top end, having a really positive thing happen to me genuinely might not move the needle as much."
This is true. Another way of saying this is: "the true effects fall as you get happier". But then, given reported happiness has stayed constant, why would the effects fall?
Hm, I don't think I agree with you on linearity. Andrew Oswald was writing about this in 2008. One option is that the function is logistic/arctan: i.e,. quite concave/flat at high latent happiness levels. That is, you can't shift reported happiness above a 10 (a ceiling effect), even if you get happier.
In this case: even if the reporting function is non-linear (and assuming true effect sizes are constant), why would the observed effects fall? Because people are getting happier. Again, this is a different way of saying rescaling is happening.
I don't think you did mention this before...! I think this graph is just for 1 country. Perhaps Japan.
To be honest, I don't know what to think of the Wolfers/Stevenson objections! My only thought is: differences of, e.g,. 0.2 points, would look pretty small in comparison to the potential rescaling effects I suggest here.
Gregory -- thank you for the interesting post. I admit I have *read most of it twice and not fully understood. May I ask two clarifying questions? Specifically on ECI.
First: you say link function choice is "relatively minor" (rank and ECI correlate ~0.97), but you also raise an endogeneity point, which "entirely relies on the link function." I think the first claim holds link functions vary, shared-shape assumption stays fixed; the second is about dropping the shared-shape assumption itself. You say a shared-shape assumption is "dubious" but I'm unsure what your reasons are?
Second: I'm wondering what the empirical test(s) is for your critiques. In general, if you think measures of capabilities are not robust in some way, what is the sensitivity analysis? I am thinking like an economist here.
One idea -- refit ECI under different (common) link functions, see if the "speeding up" trend survives. I don't think Epoch runs this, though their appendix comes close, although I don't think this would be a crux for you.
Two other ideas that came to mind, specifically on 'measure endogeneity':
Would you find these interesting or cruxy?
Thanks, Charlie
Thank you, this was interesting.
FYI - there is an excellent Sentience Institute Report from 2018 which treated this issue in a lot of detail: https://www.sentienceinstitute.org/gm-foods
Thanks for this Parker. I continue to think this research is insanely cool.
I agree with 1)
On 2), the condition you find makes sense, but aren't you implicitly assuming an elasticity of substitution of 1 with Cobb-Douglas?
Could be interesting to compare with Aghion et al. 2017. They look at a CES case with imperfect substitution (i.e., humans needed for some tasks).
https://www.nber.org/system/files/working_papers/w23928/w23928.pdf
Thanks, Alene! I appreciate that :)
Thank you for this interesting post.
“By assumption, the AI can perfectly substitute for human AI researchers.”
Any idea/intuition about what would happen if you relaxed this assumption?
Hello there,
There's lots of points here. While they are possible, I would suggest they are not particularly common/well-suported in the psychological literature as it is today.
In addition, I don't know why these explanations would lead to desensitisation towards positive and negative events.
Hello, Huw!
I can't think of a good theoretical reason why true effects should fall so significantly – like 40%. That's striking. The same attenuation result holds, even including income/age/event prevalence.
"Intuitively, if wellbeing saturates at the top end, having a really positive thing happen to me genuinely might not move the needle as much."
This is true. Another way of saying this is: "the true effects fall as you get happier". But then, given reported happiness has stayed constant, why would the effects fall?
Hm, I don't think I agree with you on linearity. Andrew Oswald was writing about this in 2008. One option is that the function is logistic/arctan: i.e,. quite concave/flat at high latent happiness levels. That is, you can't shift reported happiness above a 10 (a ceiling effect), even if you get happier.
In this case: even if the reporting function is non-linear (and assuming true effect sizes are constant), why would the observed effects fall? Because people are getting happier. Again, this is a different way of saying rescaling is happening.
I don't think you did mention this before...! I think this graph is just for 1 country. Perhaps Japan.
To be honest, I don't know what to think of the Wolfers/Stevenson objections! My only thought is: differences of, e.g,. 0.2 points, would look pretty small in comparison to the potential rescaling effects I suggest here.
Thanks, this is interesting. I wonder if this sort of individual-level noise might be smoothed out by large-n experience sampling.