Potentially impactful research: Unjournal AI-assisted prioritization dashboard (~prototype)View in threadBob Kubinec6d100In general this seems like a good use of LLMs. I might recommend something like prompt stability scores to see how calibrated LLM ratings are (and performance across LLMs): https://arxiv.org/abs/2407.02039Reply
In general this seems like a good use of LLMs. I might recommend something like prompt stability scores to see how calibrated LLM ratings are (and performance across LLMs): https://arxiv.org/abs/2407.02039