We've continued to update this. Some recent changes and additions (caveat: comment below is AI-generated, human-vouched, "DR" is a human addition)
Some of the main improvements:
- Broader and more structured research discovery. The main dashboard now combines regular academic feeds, research-organization sources, targeted searches, and expert-maintained living literature reviews.
- The historical living-review scan found 85 potentially relevant papers across 11 reviews. ... DR: Incorporating content from CG-sponsored "Living Literature reviews" (see filtering for these here) (What is a Living Literature Review?)
Better scoring and clearer interpretations. Papers now receive a standard scoring pass, with deeper analysis for higher-potential candidates. You can switch between:
- Evaluation priority: How useful might an independent Unjournal evaluation be?
- Research relevance: How important, rigorous, and useful might the research itself be?
There are also configurable weights for people who disagree with ours.
- Human feedback is incorporated explicitly. The default score is now a human–AI synthesis where human ratings exist. There’s a quick-rate mode for
--/-/~/+/++judgments and a fuller rating/comment form. The weighting is deliberately visible and described as ad hoc rather than presented as more principled than it is.- DR: We're very keen to get your feedback and ratings!
- A public calibration-review process. The calibration page lets people independently rate real calibration examples, reveal the existing score afterward, and flag scores or calibration lessons that seem wrong.
- DR: I need to look at this more carefully, it's very preliminary
- Research is connected to explicit cruxes and Pivotal Questions. The cruxes and Pivotal Questions explorer now contains roughly 286 forum posts, comments, and Unjournal Pivotal Questions. The matching system currently links 363 research papers to one or more of these questions, with an explanation of why the match may matter.
- DR: And you can filter on this
- Targeted searches are separately labeled and filterable. We’ve done focused public-paper searches in animal-welfare economics, AI impacts on global health and development, empirical conflict, and large-N GCR evidence. These are marked as targeted intake so readers don’t mistake deliberate topic oversampling for the natural composition of the literature.
- DR: And you can filter on this
- Better source and selection provenance. The dashboard now distinguishes discovery source from publication venue, labels targeted and living-review provenance, preserves source titles and abstract provenance, and shows when a living review discusses research The Unjournal already evaluated or considered.
- Some specialist views. There are now adjacent views for legal scholarship and empirical conflict research, plus improved source and coverage statistics.
The core caveat remains: these scores concern the potential value of further attention or evaluation. They aren’t grades of research quality, endorsements, or completed Unjournal decisions.
Some useful ways to help:
- Use Quick-rate mode to rate 5–10 papers in an area you know.
- Tell us which papers seem badly overrated or underrated, and why.
- Suggest a paper we missed, or submit your own research.
- Review a few calibration anchors.
- Suggest or correct a crux in the cruxes explorer.
Tell us which decisions, funding questions, or research agendas this tool should be helping with. That’s probably more valuable than feedback on the interface alone.
DR: If you think you could add value here but want something in return, let me know what I/we can do to make it attractive to you
In general this seems like a good use of LLMs. I might recommend something like prompt stability scores to see how calibrated LLM ratings are (and performance across LLMs): https://arxiv.org/abs/2407.02039
Thanks Bob. I agree this would be high value. I've been mainly thinking about human/LLM agreement, but it would be useful to know how consistent each model is, how sensiitive it is to the prompts, whether the models differ in systematic ways, etc.
Actionable takeaways from doing this? Is a 'more stable setup' (one that e.g., has fairly consistent rank orderings) better all else equal, and thus something we should work towards? Probably, as the alternative seems like just 'adding noise'.
Knowing the stability also may tells us how much effort to put into designing and building consensus around a 'reasonable prompt'; if the results are insensitive to this, we shouldn't waste too much time.
Working on implementing something now, at least a first step. I hope to report back on some legible measures of prompt stability (or lack thereof).
Is this going in the direction you would suggest? (Happy to co-work on this of course): https://uj-prioritization-prototype.netlify.app/stability/