We used to call this upper case (ea community/movement) vs lower case (the (meta) philosophy). I think meta normative framework/philsophy is the way I think about it; that is, you can apply EA to most normative frameworks, esp consequentalisty ones (though to many people in the community, the EA framework is actually just applied total utilitarianism). FWIW, and I say this as someone who has many issues with the EA community, the vast majority of EA criticisms from outside the community I see are just highly inaccurate and not worth engaging with from an intellectual POV (but maybe still worth engaging with for movement reputation).
Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture) 8
I can trivially turn off or have fake thoughts running through the voice in my head. Subconscious brain activity seems harder but obviously you can manipulate that too by changing your surroundings and drugs and what not. I wouldn't be able to control those in a meaningful way nor do I think current AI's could (but wouldn't be shocked if they already could control the voice in their head if they have it). but I would guess future ai's will know how to control increasingly large parts of their brain activations.
At the margin, S-risk work in AI is more important than x-risk work³
From a utilitarian pov, It's not clear to me that the ev of the lightcone given we survive is positive (over nothing, or aliens, or life revolving on earth). From a humanist POV I'd rather focus on all of us surviving.
Theories of consciousness will lead to actionable understanding of AI consciousness²
Very bullish on there existing a mechanistic interpretation of consciousness (hard problem). I think it would follow that we would be able to understand if basically anything is conscious.
If animals continue to exist in a post-AGI world, animal suffering will not persist
I don't have a strong take on if agi or whoever is in control will be more moral than us but I'm guessing we will be a lot richer, and I think most likely whoever is in control won't want to torture anything (though they might not care much), and if we are alot richer and advanced I'd think this will spillover to better treatment of beings. I think chance of extreme digital suffering is much higher. The mostly like s-risk as I see it is of the hansonian mathusian version where you have expanders stuck in competition, but in this case idt there will be any or a morally relevant amount of animals
Benchmarks will become useless due to eval awareness¹ (read this backwards originally)
Useless is a strong word. But yea I think they could easily end up being negative EV by giving us a false sense of security and the chance it's meaningless seems p high. If mech interp is "good enough" maybe the two can remain useful together.
I keep seeing people say pangram is a really good tool and will improve AI safety. I feel decently confident that the end state of AI writing detection is that AI's can ~perfectly replicate human writing (when they want to) and that the wide scale deployment of this tool is analogous to spamming antibiotics on factory farms. I made a 100$ charitable bet at 1:1 odds that within 3 years AI writing detection will be ~useless. would be very eager to be pointed to any strong arguments for why I'm wrong.
Some musings on meta science here, kind of tangential and half baked. capabilities is a function of intelligence only once you pick a specific technology and you fix a utility function over the tech, otherwise it's not well defined.
I haven't done formal probability in long enough to have the exact words, but essentially you can think of building a technology as a markov chain of decisions. Both the path you take and the chance of success at each path are (partly) a function of your (intelligence). Capabilities is taking a utility function defined over the tech tree and multiplying it by the current probabilities/expected number of steps to different parts of the tree.
imagine you are trying to build a spear. You have 2 choices of sticks, 2 binding agents (glue or chords), 2 choices of stones, and only 1 of each choice is going to work. I guess you can think of this in p or bits but essentially you have 1/8. chance of a random walk working, given we fix p(success per task) at 100% for simplicity (https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ also it already has its own discussion).
Ok so the maximum intelligence in any situation is getting to the absorbing node/final tech in the minimum number of steps (with 100% chance). For any specific technology there does exist some (not necc unique) pathway s.t. there does not exist any other pathway with less (difficulty weighted) steps. So capability is always bounded on a specific tech.
The divergence (ratio of difficulty weighted steps) between the random walk and the correct path is basically the amount of juice on that tech intelligence can give you. This is basically unbounded in theory. Also p(success per task) might be a function of the same underlying architecture as p(correct path taken at node n) idk.
Now multiply a utility function over the tech tree. The increase of capabilities is the delta in utility per step between the two intelligences.
It also gets much more complicated in reality. Paths can be self correcting if when the agent goes down a wrong path there is some probability > 0 that they can realize this, which depends both on the intelligence of the agent and the nodes in the tech path itself. Similarly, some paths are smooth and monotonic among all the reasonable choices and others will punish greedy algorithms. but I think the core intuition is what I said earlier.
We used to call this upper case (ea community/movement) vs lower case (the (meta) philosophy). I think meta normative framework/philsophy is the way I think about it; that is, you can apply EA to most normative frameworks, esp consequentalisty ones (though to many people in the community, the EA framework is actually just applied total utilitarianism). FWIW, and I say this as someone who has many issues with the EA community, the vast majority of EA criticisms from outside the community I see are just highly inaccurate and not worth engaging with from an intellectual POV (but maybe still worth engaging with for movement reputation).
I can trivially turn off or have fake thoughts running through the voice in my head. Subconscious brain activity seems harder but obviously you can manipulate that too by changing your surroundings and drugs and what not. I wouldn't be able to control those in a meaningful way nor do I think current AI's could (but wouldn't be shocked if they already could control the voice in their head if they have it). but I would guess future ai's will know how to control increasingly large parts of their brain activations.
From a utilitarian pov, It's not clear to me that the ev of the lightcone given we survive is positive (over nothing, or aliens, or life revolving on earth). From a humanist POV I'd rather focus on all of us surviving.
Very bullish on there existing a mechanistic interpretation of consciousness (hard problem). I think it would follow that we would be able to understand if basically anything is conscious.
I don't feel confident at all, but the behavior of llms rn sure do remind me of at least elements of stress, discomfort, and happiness.
I don't have a strong take on if agi or whoever is in control will be more moral than us but I'm guessing we will be a lot richer, and I think most likely whoever is in control won't want to torture anything (though they might not care much), and if we are alot richer and advanced I'd think this will spillover to better treatment of beings. I think chance of extreme digital suffering is much higher. The mostly like s-risk as I see it is of the hansonian mathusian version where you have expanders stuck in competition, but in this case idt there will be any or a morally relevant amount of animals
20% disagree➔ 20% agreeUseless is a strong word. But yea I think they could easily end up being negative EV by giving us a false sense of security and the chance it's meaningless seems p high. If mech interp is "good enough" maybe the two can remain useful together.
I keep seeing people say pangram is a really good tool and will improve AI safety. I feel decently confident that the end state of AI writing detection is that AI's can ~perfectly replicate human writing (when they want to) and that the wide scale deployment of this tool is analogous to spamming antibiotics on factory farms. I made a 100$ charitable bet at 1:1 odds that within 3 years AI writing detection will be ~useless. would be very eager to be pointed to any strong arguments for why I'm wrong.
Some musings on meta science here, kind of tangential and half baked. capabilities is a function of intelligence only once you pick a specific technology and you fix a utility function over the tech, otherwise it's not well defined.
I haven't done formal probability in long enough to have the exact words, but essentially you can think of building a technology as a markov chain of decisions. Both the path you take and the chance of success at each path are (partly) a function of your (intelligence). Capabilities is taking a utility function defined over the tech tree and multiplying it by the current probabilities/expected number of steps to different parts of the tree.
imagine you are trying to build a spear. You have 2 choices of sticks, 2 binding agents (glue or chords), 2 choices of stones, and only 1 of each choice is going to work. I guess you can think of this in p or bits but essentially you have 1/8. chance of a random walk working, given we fix p(success per task) at 100% for simplicity (https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ also it already has its own discussion).
Ok so the maximum intelligence in any situation is getting to the absorbing node/final tech in the minimum number of steps (with 100% chance). For any specific technology there does exist some (not necc unique) pathway s.t. there does not exist any other pathway with less (difficulty weighted) steps. So capability is always bounded on a specific tech.
The divergence (ratio of difficulty weighted steps) between the random walk and the correct path is basically the amount of juice on that tech intelligence can give you. This is basically unbounded in theory. Also p(success per task) might be a function of the same underlying architecture as p(correct path taken at node n) idk.
Now multiply a utility function over the tech tree. The increase of capabilities is the delta in utility per step between the two intelligences.
It also gets much more complicated in reality. Paths can be self correcting if when the agent goes down a wrong path there is some probability > 0 that they can realize this, which depends both on the intelligence of the agent and the nodes in the tech path itself. Similarly, some paths are smooth and monotonic among all the reasonable choices and others will punish greedy algorithms. but I think the core intuition is what I said earlier.
https://x.com/anthropicai/status/2082965101083320543?s=46
Claude did some hacking.
(Separately, was there some memo I missed - ea forum seems to have ceded all ai discussion to less wrong except second order ipo donation stuff?)