How we score evidence
You tapped a badge, so you want to know what the number means. Here is the whole thing, including the parts that don't flatter us.
Why there is a number at all
Every resource in Resolv points at a source. That was the rule from the start: no citation, no publishing. But "there is a citation" is a low bar. You can cite a survey of forty people in a journal nobody has heard of, funded by the company that sells the thing being surveyed, and it looks identical on the page to a Cochrane review of ninety thousand patients.
So we score the source. Not the claim, not whether we agree with it, not whether it fits the way we see mental health. The source.
Nobody awards the grade
This is the part that matters most, so it goes first. There is no field anywhere in our system where a person types in a trust score. An editor records facts about the paper: what kind of study it was, who funded it, where it was published, how many people were enrolled, whether it was registered before the data came in, whether conflicts were declared, whether an author was testing their own theory. The number is then computed from those facts, fresh, every time the page loads.
That design decision has a cost — we can't make an exception for a study we love. That's the point. It also means that if we improve the rubric next year, every resource in the library re-scores itself instantly, including the ones that get worse.
The six things we count
A hundred points, split six ways.
Study design — 35 points. The biggest single input, because it is the biggest single determinant of whether a result means anything. The ranking follows the standard evidence hierarchies: a meta-analysis of randomised trials at the top, then systematic reviews, then individual randomised trials, then cohort studies, case-control, cross-sectional, case reports, and opinion at the bottom [2][3]. Randomised evidence starts high and observational evidence starts low, which is the GRADE convention [2].
Two non-academic designs sit deliberately near the top. A regulator's formal safety determination — an FDA boxed warning, say — rests on the entire post-marketing safety dataset, not on one study, and it clears a legal evidentiary bar. A national clinical guideline like NICE or an RCPsych position statement synthesises the literature under a published, auditable method. Ranking either of those below a single small trial would be wrong, so we don't.
Who paid for it — 20 points. Weighted heavily and on purpose. The Cochrane methodology review on this question compared industry-sponsored studies with independently funded studies asking the same questions, and found the industry-sponsored ones reported more favourable efficacy results, with a risk ratio of 1.27, and more favourable overall conclusions [1]. The crucial detail is what happened next: the effect persisted after adjusting for the standard risk-of-bias domains — allocation concealment, blinding, follow-up [1]. It is not caught by looking at how well the study was run. So it has to be scored separately, or it goes unmeasured.
Independent funding and no funding both score full marks. Mixed funding scores half. Industry funding scores zero. A study with no funding statement at all scores low but not zero — a missing disclosure is a weak negative signal, not proof of sponsorship.
Where it was published — 15 points. Not impact factor. Impact factor is a journal-level citation average that tells you almost nothing about an individual paper, and the research community has been formally asking people to stop using it as a proxy for quality since 2013 [6]. We use quartile ranking plus indexing status, which are checkable and much harder to inflate. Government and national health bodies score near the top, for the same reason regulatory determinations do. Preprints score low but not zero — not yet reviewed is not the same as wrong. Predatory venues score zero.
How many people — 10 points. A crude proxy for imprecision, which is one of the five things GRADE asks you to downgrade for [4]. Ten thousand participants scores full marks; thirty scores almost nothing. Regulatory determinations and guidelines are exempt, because they rest on a synthesised or whole-population evidence base and a single enrolled N does not exist for them.
Whether they showed their work — 10 points. Preregistration is worth six of those points and a conflict-of-interest statement is worth four. Preregistration counts for more because it constrains the analysis before the data arrive, while a disclosure only tells you afterwards. There is a striking piece of evidence for why this matters: among large NHLBI cardiovascular trials, 57% reported a positive result before registration became standard, and 8% did after [5]. The trials didn't get worse. The reporting got honest.
Whether the author was testing their own idea — 10 points. This one is almost never scored anywhere, and it should be. When the person running the trial is the person whose theory is under test, results tilt toward the theory. Reviews of psychotherapy trials put this allegiance effect in the same broad range as the industry-funding effect [8][9], which is to say: large enough that ignoring it is a choice, not an oversight.
If we haven't checked yet, it scores half credit. Not full — we don't assume clean.
Retraction is not a deduction
If a paper has been retracted, the score is zero and the badge says retracted. Design, journal, funding and sample size stop mattering entirely at that point. It is also blocked from being published in the app at all, enforced at the database level rather than by anyone remembering. We check against the public retraction databases, which have been freely available since Crossref and Retraction Watch opened the dataset in 2023 [7].
The bands
We show you a band, not a bare number, because a raw score out of 100 implies a precision this rubric does not have. The number and the full component breakdown are there when you tap.
Gold standard, 85 and above. Top of the evidence hierarchy, independently funded.
Strong evidence, 70 to 84. Well-supported by good-quality research.
Moderate evidence, 55 to 69. Reasonable evidence with real limitations.
Early signal, 40 to 54. Suggestive but preliminary. Not settled.
Contested, below 40. Weak or heavily disputed. Treat it as a question, not an answer.
Retracted. Do not rely on it.
Three real examples
The FDA's 2020 boxed warning on benzodiazepines scores 86 — gold standard. Regulatory determination, no external funding, government body, population-level evidence base, conflicts declared, independent. It is not a randomised trial and it never will be, and it is still the strongest thing that exists on that question.
Cipriani's 2018 network meta-analysis of antidepressants scores 100. Meta-analysis of randomised trials, NIHR funded, published in the Lancet, over 116,000 participants, registered on PROSPERO in advance, conflicts declared, no allegiance problem. That is what a full house looks like.
Moncrieff's 2022 umbrella review of the serotonin hypothesis scores 72 — strong. Unfunded, published in a top-quartile journal, registered in advance. It loses the full ten allegiance points because the lead author is the most prominent proponent of the position the review supports. That is not an accusation. It is a fact about the paper, and the rubric applies it to everyone, including the researchers whose conclusions I personally find most persuasive. If the score bent for people I agree with, it would be worthless.
What the score cannot tell you
It cannot tell you the finding is true. Well-designed studies are wrong all the time. The score tells you how much weight the method can bear, not whether the conclusion holds.
It cannot tell you the finding applies to you. A drug trial in 40-year-old men with severe depression may say nothing useful about a 19-year-old with panic attacks. The score says nothing about that gap.
It cannot tell you a low-scoring finding is false. Some true things have only ever been studied badly, usually because nobody was willing to fund the good version. Withdrawal effects were dismissed for decades on exactly that basis. A low score means ask the question, not drop it.
When we haven't finished checking
Some resources will show as incomplete. That means an editor has recorded some of the provenance but not all of it, and the score is provisional. We would rather say we haven't finished than let a missing field look like a judgment about the paper.
Where a finding is disputed
A high score and a live controversy can coexist. Where they do, the badge carries a note saying who disputes the finding and why. A well-conducted study in a contested area is still in a contested area, and you are owed that before you take it to your doctor.
If we get it wrong
We will. The rubric is a set of judgment calls about weights, and reasonable people would weight them differently. If you think a resource is mis-scored, or a paper we cite has been retracted, or the provenance we recorded is factually wrong, tell us. Facts get corrected immediately. The weights get argued about in public.
One last thing. None of this is medical advice, and nothing in the library is a reason to change or stop a medication. It is there so you can walk into an appointment knowing what the evidence actually says, and ask better questions than you would have otherwise. That is the whole ambition.
adam
frequently asked questions
Who decides the score?
Nobody. An editor records facts about the paper — what kind of study it was, who paid for it, where it was published, how many people were in it, whether it was preregistered. The number is calculated from those facts every time you open the page. There is no field where someone types in a grade.
Does a high score mean the finding is true?
No. It means the finding was produced by a process that is harder to fool. Good process still produces wrong answers, and the history of medicine is full of well-designed studies that turned out to be wrong. The score tells you how much weight the method can bear, not whether the conclusion is correct.
Does a low score mean the finding is false?
No. Plenty of true things have only ever been studied badly, usually because nobody funded the good version. A low score means treat it as a question worth asking your doctor, not as an answer.
Why is who funded the study worth so many points?
Because it is one of the most reliably measured biases in medicine, and it is not caught by looking at study design. A Cochrane review of the research on research found industry-sponsored studies reported more favourable results than independently funded studies asking the same questions, and that the gap survived adjusting for the usual quality checks [1].
Why don't you use journal impact factor?
Impact factor is a journal-level average of citations. It says very little about any individual paper, and it is gamed. Thousands of researchers and institutions have signed a declaration asking people to stop using it this way [6]. We use quartile ranking and indexing status instead, which are checkable and much harder to inflate.
What happens if a paper we cite gets retracted?
It scores zero and the badge changes to say so, no matter how good the design or the journal was. A retracted paper is also blocked from being published in the app at all. We check against the public retraction databases [7].
A resource says the evidence is contested. Why is it still there?
Because a live disagreement is information you are owed. If serious people dispute a finding, we would rather show you the finding and the dispute than quietly delete one side. The badge tells you how much weight to put on it.
you don't have to go through this alone
free. anonymous. available 24/7. from struggle to resolved 🤍
get Resolv Social — it's freekeep reading
Mindfulness, minus the marketing
The research on mindfulness is real, but the hype has run ahead of the evidence. Here's what actually works and what doesn't.
The unsexy truth about gratitude and mood
Gratitude interventions show modest benefits for well-being, but the research is narrow and the effects are smaller than the hype suggests.
What the research says about talking to people who get it
Peer support works because shared experience creates understanding that professional care often can't. Here's what the evidence shows.