We've continued to update this. Some recent changes and additions (caveat: comment below is AI-generated, human-vouched, "DR" is a human addition)
Some of the main improvements:
Broader and more structured research discovery. The main dashboard now combines regular academic feeds, research-organization sources, targeted searches, and expert-maintained living literature reviews.
Better scoring and clearer interpretations. Papers now receive a standard scoring pass, with deeper analysis for higher-potential candidates. You can switch between:
Evaluation priority: How useful might an independent Unjournal evaluation be?
Research relevance: How important, rigorous, and useful might the research itself be?
There are also configurable weights for people who disagree with ours.
Human feedback is incorporated explicitly. The default score is now a human–AI synthesis where human ratings exist. There’s a quick-rate mode for --/-/~/+/++ judgments and a fuller rating/comment form. The weighting is deliberately visible and described as ad hoc rather than presented as more principled than it is.
DR: We're very keen to get your feedback and ratings!
A public calibration-review process. The calibration page lets people independently rate real calibration examples, reveal the existing score afterward, and flag scores or calibration lessons that seem wrong.
DR: I need to look at this more carefully, it's very preliminary
Research is connected to explicit cruxes and Pivotal Questions. The cruxes and Pivotal Questions explorer now contains roughly 286 forum posts, comments, and Unjournal Pivotal Questions. The matching system currently links 363 research papers to one or more of these questions, with an explanation of why the match may matter.
Better source and selection provenance. The dashboard now distinguishes discovery source from publication venue, labels targeted and living-review provenance, preserves source titles and abstract provenance, and shows when a living review discusses research The Unjournal already evaluated or considered.
The core caveat remains: these scores concern the potential value of further attention or evaluation. They aren’t grades of research quality, endorsements, or completed Unjournal decisions.
Some useful ways to help:
Use Quick-rate mode to rate 5–10 papers in an area you know.
Tell us which papers seem badly overrated or underrated, and why.
Tell us which decisions, funding questions, or research agendas this tool should be helping with. That’s probably more valuable than feedback on the interface alone.
DR: If you think you could add value here but want something in return, let me know what I/we can do to make it attractive to you
I agree that something like that needs to be done. I did something similar myself for EA forum posts, but apparently people didn't like it because it didn't get any upvotes here (I don't know why).
My main question/feedback is:
"EA Forum paper links" - what exactly is that? Are these EA posts that link to a paper? If so, then why not include all posts, why only papers?
"EA Forum paper links" is meant to find research papers/projects that were mentioned on the EA Forum. This tool is mainly meant for surfacing in-depth research that uses formal methods and is putting itself out there as being up to the highest methodological rigor. I mainly meant this tool as a source of research for The Unjournal to consider and evaluate.
(Of course some EA forum posts are also themselves research objects, and occasionally they are even sometimes rather detailed and technical.)
I'm not sure what your tool was doing but it does sound potentially interesting.
We've been experimenting with using LLMs to help identify and prioritize research for Unjournal evaluation, to work with and complement human prioritization (and learn). We now have a public prototype dashboard. It's early stage and needs refinement we have not invested a lot of compute/API credit into this.
What it does: Automatically discovers recent papers from NBER, arXiv (econ), CEPR, SSRN, Semantic Scholar, EA Forum paper links, OpenAlex, Anthropic Research and then using AI models (mostly GPT-5.4 family) against our prioritization criteria— decision relevance, prominence, timing value, and methodological potential.
Domain: economics, quantitative social science, forecasting, and policy-relevant research
Caveats: This is very preliminary and the AI recommendations are not yet well-calibrated. As of 14 Apr 2026 many of the suggestions are mediocre we're sharing it for transparency and feedback. This supplements our existing Public Database of Prioritized Research on Coda (and those papers have been folded in here too. )
Scores reflect evaluation priority (expected value of commissioning an independent review), not research quality. ATM (IIRC) the AI only sees paper metadata and abstracts, not full texts. There's also a statistics page showing the breakdown by source, cause area, and score distribution.
Probably building towards a hybrid/centaur model here, with human and AI prioritization feedback reinforcing each other. (And see "planned workflow" at the bottom.)
Feedback encouraged. You can also comment directly on the page via Hypothes.is, and we'll adapt.
Disclaimer: This blog post focuses on a new piece of research from Rethink Priorities. I was not involved with the funding or execution of this research, and therefore write this piece merely as a consumer. However, I am the Executive Director of Giving Green and a board member of Rethink Priorities, and acknowledge that these affiliations may bias my...
Epistemic status: Speculation from two decently informed advocates armed with anecdata.
Note on process: After having some version of this conversation several times and saying, “we should probably write about this publicly,” we took the less heroic route: we recorded one of our conversations, fed the transcript into an LLM, and then substantially revised the structure, substance, and framing ourselves. We will not be sharing the transcript, as it is in...
Michael Nielsen has a beautiful new essay on moral imagination: the ability humans have to 'develop and transmit new notions of good action, indeed, even new kinds of good'.
As examples, he gives:
* Hammurabi's creation of a code of laws (in 1754 BCE) and his justification of his rule not in terms of divine or hereditary right, but by the delivery of justice to his citizens.
* St Gregory of Nyssa's...
We've continued to update this. Some recent changes and additions (caveat: comment below is AI-generated, human-vouched, "DR" is a human addition)
Some of the main improvements:
Better scoring and clearer interpretations. Papers now receive a standard scoring pass, with deeper analysis for higher-potential candidates. You can switch between:
There are also configurable weights for people who disagree with ours.
--/-/~/+/++judgments and a fuller rating/comment form. The weighting is deliberately visible and described as ad hoc rather than presented as more principled than it is.The core caveat remains: these scores concern the potential value of further attention or evaluation. They aren’t grades of research quality, endorsements, or completed Unjournal decisions.
Some useful ways to help:
Tell us which decisions, funding questions, or research agendas this tool should be helping with. That’s probably more valuable than feedback on the interface alone.
DR: If you think you could add value here but want something in return, let me know what I/we can do to make it attractive to you