Director of Strategy for the Centre for Effective Altruism. I previously ran new programs at Innovate Animal Ag and led the research team at a nonprofit focused on building $1B+ philanthropic initiatives/megaprojects. Before that I lived in Tanzania and ran some RCTs there.
Quick note: the Forum does mark posts as AI-generated using Pangram, so users will ideally know before they read the post.
I agree that if an LLM is just taking a quick take and adding arbitrary length to it, that's bad for the Forum. For me, a challenge is that people use LLMs in a lot of different ways, including (to me!) quite good ways. For example, Rocky and Tom in the below post spoke their detailed thoughts on welfare tech into an LLM and it turned those in a high-quality post: https://forum.effectivealtruism.org/posts/QHgfRkiNFoGyysDpn/beware-silver-bullets-are-we-making-welfare-tech-into-the . In that example, the LLM mostly just made it much easier to turn their ideas into a single post that would have taken a lot longer to write (and therefore might not have been written).
That's all to say, I'm a bit unsure on whether there's a clear "AI written posts = bad" policy we can implement. The Forum team does sometimes move poor quality posts to Personal Blogs and off the front page, and "AI slop" can definitely count as that, but it's more of a case-by-case thing.
Just some quick thoughts as I think this is important and something the Forum team is working on (I lead the Online team at CEA, which includes the Forum).
Nice! Though I guess my concern that someone would conclude the program hadn't work, after failing to reject the null, is still there. But if the proposed study goes ahead, I'm sure your quick input could be very helpful!
I really like that you are digging into this! 2 quick questions:
Lastly just a statistical power flag: you link to LawrenceC's piece on variation in application assessments. That variation would suggest we'd need quite a good sample size to allow for precision in the detectable effect, which might be hard. I guess that's just a note to make sure some sample calculation is done once you have clarity on your outcome of interest to make sure you have a path to measuring the kind of effect that would make these programs cost-effective. The worst case scenario would be a study with 100 people looking to detect an impact of, say, 0.05 SD, when you are in fact only powered to detect 0.3 SD and basically doomed to concluding the program doesn't work, when perhaps it does.
These are intended as genuine questions to see if your ideas could work, if there is indeed a path to a good study here I'd love to see something proposed to, say, the EA Infrastructure Fund, your ideas seem valuable! (Disclaimer, I work at CEA, which contains the EAIF, but I am not involved in EAIF decisions and these are personal takes)
Good catch! I've updated the title for now to be the title of the podcast episode they are promoting (using my Immense Forum Admin Powers)
I love that! Thanks for the example.
Congrats on the great sign up results! Do you mind giving some more examples of what you did differently? Those could be helpful for other groups to learn from.
For instance, you noted "For example: you walk past a table and someone goes “Hey! Do you want to join the EA club?” " as an example of what you wouldn't to-- what would you do to get attention when tabling?
Sounds like you have a super relevant background! Some quick thoughts:
Both of those might just help you realize which parts of GW's process you don't really "get", so you know what to work on.
But I have never worked at GW, so take all of this as "one guy's opinion" :)
Interesting suggestion! I think the example you outline makes sense. I'd guess part of why you don't see this kind of thing in EA might include:
Anyway, just some quick thoughts, very open to counter takes and thank you for suggesting this idea! I do think there is a good place for some smart finance type stuff in EA for sure, such as loans for social enterprises to kick off, or other clever things like advance market commitments.
Wow this is so well done. I'm really excited to see what else you have in the pipeline, and more generally for high quality EA content!
Thanks Vasco! I think it depends on how well your particular system is a strict A > B> C> Done flow. If it's as linear as car production, then it is indeed that case that if you make 100 wheels, 20 front axles and 40 windshields per day, you have made 20 cars (your bottleneck is the axles). And making 400 wheels the next day without changing your axle output has zero effect on the number of cars produced, the additional "productivity" is entirely wasted.
Not every process will be quite so linear, but at least for the elements that are, I do think that increasing the output of non-bottlenecks will have zero effect on the output you care about.