Currently pursuing a PhD at the "Mathematics for Our Future Climate" CDT at Reading University.
Previously MSc in applied mathematics/theoretical ML.
Not really active here - racism, Rationality and weirdness in the movement are so bad they made me give up on it.
Seems like wealth does predict openness, just not entirely?
Not really. It's a criticism of a system that allows (and even encourages) you to pretend to know things when you don't.
Of course, I'm a mathematician and I think Bayesian reasoning is fundamentally correct and is useful in some contexts. Just not the contexts EA uses it for.
I mean "do I expect to see, within the relevant time frame, enough information to make Bayesian updating with a very wrong (or just very wide) initial distribution useful rather than harmful?".
The question of whether we should expect sufficient evidence to bring us close enough to the truth still stands.
But you're not claiming that the models should only be shared with AI researchers. You're claiming they should only be shared with AI researchers specifically employed by Anthropic.
Although no, I disagree that the input from non-AI-researchers is useless here - as you need to hear both from the end users and from people affected by AI and its decisions.
We don't know how to align a possible AGI yet. The best we can hope for is that current models are close enough to whatever AGI is going to be, that trying to align them will teach us about aligning an AGI. This task, of trying to align them, is something that shouldn't just be left to researchers in AI companies.
How can you "solve every possible jailbreak"? And is it worth it crippling large-scale research into safeguarding from future AI because of fears about what the current models might be capable of?
(My own answer is "maybe". It depends on how bad you think current models are for society - pretty bad in my opinion - vs. how likely you think it is an existentially-threatening AI will actually be born out of the current efforts).
I still maintain that publicly releasing models is the correct way to get any chance of good alignment research - you can't possibly believe that the researchers at Anthropic alone are enough to tackle the problem. It's a global problem and should have the opportunity for the global population to solve it.
Being featured on Snopes is sort of a major achievement IMO :)
Where was it published in traditional media?