Helpful post, Zach! I think it's more useful and concrete to focus on asking about specific capabilities instead of asking about AGI/TAI etc. and I'm pushing myself to ask such questions (e.g., when do you expect to have LLMs that can emulate Richard Feynmann-level -of-text). Also, I like the generality vs capability distinction. We already have a generalist (Gato) but we don't consider it to be an AGI (I think).
The quick answer is that wanting to do alignment-related work does not depend on a Philosophy PhD, or any graduate degree tbh. I'd say, start thinking about what are your interests more specifically and then there might be different paths to impact with or without the degree.
I know what you mean and I've definitely had this kind of experience (and in particular, last semester which led me to want to leave both my university and academia-- that's how bad it was). What I wanted to emphasize while teaching is that it's valuable to question our own thoughts, emotions, and experiences in the philosophy classroom, and it's disappointing to see that most people are not willing to do that. But hey, at least I tried...
A quick comment after reading about 50% of this article: it seems to focus on statements instead of arguments, e.g. "we cannot calculate the probability of future events as if people didn’t exist." or "We are after genuine creativity, not the illusion of creativity."At the same time, it doesn't really engage with the literature on AI risk or even explain why the definitions adopted e.g. the knowledge definition, are the most appropriate ones. There might be some interesting thoughts in there, but it'd be better for the author to develop them in shorter articles and make the arguments more clear.
I've been asked this question! Or, to be specific, I've been asked something along these lines: human cultures have always been speculating about the end of the world so how is this forecasting x-risk any different?
I don't think it's restricted only to agentic technologies; my model is for all technologies that involve risk. My toy example is that even producing a knife requires the designer to think about its dangers in advance and propose precautions.
Thank you, that's great. I'd be keen to start a project on this. For whoever is interested, please DM me and we can start brainstorming and form a group etc.
Helpful post, Zach! I think it's more useful and concrete to focus on asking about specific capabilities instead of asking about AGI/TAI etc. and I'm pushing myself to ask such questions (e.g., when do you expect to have LLMs that can emulate Richard Feynmann-level -of-text). Also, I like the generality vs capability distinction. We already have a generalist (Gato) but we don't consider it to be an AGI (I think).
The quick answer is that wanting to do alignment-related work does not depend on a Philosophy PhD, or any graduate degree tbh. I'd say, start thinking about what are your interests more specifically and then there might be different paths to impact with or without the degree.
I know what you mean and I've definitely had this kind of experience (and in particular, last semester which led me to want to leave both my university and academia-- that's how bad it was). What I wanted to emphasize while teaching is that it's valuable to question our own thoughts, emotions, and experiences in the philosophy classroom, and it's disappointing to see that most people are not willing to do that. But hey, at least I tried...
A quick comment after reading about 50% of this article: it seems to focus on statements instead of arguments, e.g. "we cannot calculate the probability of future events as if people didn’t exist." or "We are after genuine creativity, not the illusion of creativity."At the same time, it doesn't really engage with the literature on AI risk or even explain why the definitions adopted e.g. the knowledge definition, are the most appropriate ones. There might be some interesting thoughts in there, but it'd be better for the author to develop them in shorter articles and make the arguments more clear.
I've been asked this question! Or, to be specific, I've been asked something along these lines: human cultures have always been speculating about the end of the world so how is this forecasting x-risk any different?
Both Redwood and Anthropic have labs and do empirical work. This is also an example of experimental work: https://twitter.com/Karolis_Ram/status/1540301041769529346
I don't think it's restricted only to agentic technologies; my model is for all technologies that involve risk. My toy example is that even producing a knife requires the designer to think about its dangers in advance and propose precautions.
Here's my attempt to reflect on the topic: https://forum.effectivealtruism.org/posts/PWKWEFJMpHzFC6Qvu/alignment-is-hard-communicating-that-is-harder
Thank you, that's great. I'd be keen to start a project on this. For whoever is interested, please DM me and we can start brainstorming and form a group etc.