Used to run Middlebury Effective Altruism Worked as an economist and Walmart and gave a bunch of money away Incoming MATS (Summer 25) scholar in Neel Nanda's stream See timhua.me
Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture)
Depends on how good the AI is and how good the tools are? This is kind of a bad question since "deceptive AIs" is not a very precise definition.
Benchmarks will become useless due to eval awareness¹
By using AIs and access to real-world usage data to build benchmarks, it seems plausible that even weakly superhuman AIs will be uncertain whether it is being deployed or evaluated.
One life hack people don't know is this Attendee Data Sheet gives you access to a google sheet filled with everyone who is going to the conference. You can then do things like:
Ctrl+F for "AI control" to see everyone who has mentioned that phrase anywhere in their profile.
Plug the entire data sheet as a csv file into Gemini 2.5 Pro and ask it questions.
I think these "preserve trees" offsets might be somewhat fake to begin with? I've personally given to make sunsets (direct aerosol injection) and also Climateworks (direct carbon capture from the air and injecting it into groundwater/the rocks).
In case people don't know, Oath is an Democratic party affiliated org that identifies underfunded and close races where your marginal donations could really matter.
I'd encourage you to stop making these sorts of posts. I think they're off-putting for people that might otherwise engage more with more reasonable EA ideas.
I strong downvoted this comment because I think this type of discourse censorship is terrible. Effective Altruism should be about figuring out how to do the most good, and then doing just that.
"This idea is off putting" can be use as a fully general counterargument against any new intervention or pivot. Helping farmed animals is off putting to many. Helping people abroad before helping those at home is off putting to many.
This is, by the way, not to say that you can't dismiss an argument if the logic lead to absurd conclusions. Reasoning from first principals can be a dangerous activity if you take your ideas seriously (see e.g., epistemic learned helplessness, memedic immune system). But when trying to figure out how to do the most good, I think it's really really bad to have any sort internal thought censors.
(I think it's comparably better to consider "does this sound off putting" deciding what actions to take.)
While you can use o1 and gemini with internet access, I think they almost certainly evaluated it without such access (see the original paper here).
I really really do not think you should put the plot there. It's like comparing two different students performance except one of them has access to the internet. I think it's extremely misleading. If you want to illustrate progress you could just use the FrontierMath/GPQA results or even ARC-AGI.
I endorse moral reasoning where you start from a conclusion, and then work backwards to discover general principals.
I think this community is much more at risk of being led astray by convincing-sounding but actually incorrect arguments, as opposed to having starting assumptions that vastly limit their ability to do good (I will probably give the opposite advice to most other people).
Depends on how good the AI is and how good the tools are? This is kind of a bad question since "deceptive AIs" is not a very precise definition.
By using AIs and access to real-world usage data to build benchmarks, it seems plausible that even weakly superhuman AIs will be uncertain whether it is being deployed or evaluated.
One life hack people don't know is this Attendee Data Sheet gives you access to a google sheet filled with everyone who is going to the conference. You can then do things like:
Ctrl+F for "AI control" to see everyone who has mentioned that phrase anywhere in their profile.
Plug the entire data sheet as a csv file into Gemini 2.5 Pro and ask it questions.
I think these "preserve trees" offsets might be somewhat fake to begin with? I've personally given to make sunsets (direct aerosol injection) and also Climateworks (direct carbon capture from the air and injecting it into groundwater/the rocks).
In case people don't know, Oath is an Democratic party affiliated org that identifies underfunded and close races where your marginal donations could really matter.
I strong downvoted this comment because I think this type of discourse censorship is terrible. Effective Altruism should be about figuring out how to do the most good, and then doing just that.
"This idea is off putting" can be use as a fully general counterargument against any new intervention or pivot. Helping farmed animals is off putting to many. Helping people abroad before helping those at home is off putting to many.
This is, by the way, not to say that you can't dismiss an argument if the logic lead to absurd conclusions. Reasoning from first principals can be a dangerous activity if you take your ideas seriously (see e.g., epistemic learned helplessness, memedic immune system). But when trying to figure out how to do the most good, I think it's really really bad to have any sort internal thought censors.
(I think it's comparably better to consider "does this sound off putting" deciding what actions to take.)
While you can use o1 and gemini with internet access, I think they almost certainly evaluated it without such access (see the original paper here).
I really really do not think you should put the plot there. It's like comparing two different students performance except one of them has access to the internet. I think it's extremely misleading. If you want to illustrate progress you could just use the FrontierMath/GPQA results or even ARC-AGI.
Wait that humanity's last exam plot is super misleading right? Since the other models did not have access to the internet but Deep Research does?
I don't seem to see the fireside chat with forethought on the agenda, will it be added later? I'd love to attend!
I endorse moral reasoning where you start from a conclusion, and then work backwards to discover general principals.
I think this community is much more at risk of being led astray by convincing-sounding but actually incorrect arguments, as opposed to having starting assumptions that vastly limit their ability to do good (I will probably give the opposite advice to most other people).
See e.g., Epistemic learned helplessness, Memetic immune system.