The Google DeepMind mechanistic interpretability team has made a strategic pivot over the past year, from ambitious reverse-engineering to a focus on pragmatic interpretability:
Trying to directly solve problems on the critical path to AGI going well[1]
Measuring progress with empirical feedback on proxy tasks
We believe that, on the margin, more researchers who share our goalsshould take a pragmatic approach to interpretability, both in industry and academia, and we call on people to join us
Our proposed scope is broad and includes much non-mech interp work, but we see this as the natural approach for mech interp researchers to have impact
Specifically, we’ve found that the skills, tools and tastes of mech interp researchers transfer well to important and neglected problems outside “classic” mech interp
Most existing interpretability techniques struggle on today’s important behaviours, e.g. they involve large models, complex environments, agentic behaviour and long chains of thought
Problem: It is easy to do research that doesn't make real progress.
Our approach: ground your work with a North Star - a meaningful stepping-stone goal towards AGI going well - and a proxy task - empirical feedback that stops you fooling yourself and that tracks progress toward the North Star.
We see two main approaches to research projects: focused projects (proxy task driven), and exploratory projects (curiosity-driven, proxy task validated)
Curiosity-driven work can be very effective, but can also get caught in rabbit holes. We recommend starting in a robustly useful setting, time box your exploration[3], and finding a proxy task as a validation step[4]
We advocate method minimalism: start solving your proxy task with the simplest methods (e.g. prompting, steering, probing, reading chain-of-thought). Introduce complexity or design new methods only once baselines have failed.
Read the full post here, and the companion piece on promising AGI Safety relevant research directions here
TL;DR
NOVAH (No Violence At Home) was incubated by Charity Entrepreneurship (now Ambitious Impact) in 2024 to test a promising idea: preventing intimate partner violence through edutainment, in our case a serialised radio drama. Over the past two years we have produced and aired two seasons in Rwanda.
We are currently evaluating our second season through a randomized controlled trial with 2,400 couples in Rwanda in partnership wi...
TL;DR
* The Long-Term Future Fund is closing down, and EA Funds is launching the Transformative AI Fund with a new full-time team.
* The fund's primary focus is technical AI safety and AI governance (including post-AGI governance), as well as supporting fields such as field-building and forecasting. We'll also consider non-GCR implications of transformative AI such as flourishing futures and digital...
The current Long Term Future Fund (LTFF) fund managers and I have decided to step back from our work on the LTFF. Because we believe LTFF donors trusted the fund managers to ensure that the funds would be used in line with the purposes of their donation, we've decided the right move is to close the fund.
While LTFF is closing, note that EA Funds has launched a new fund...