The version that works: a score you use in the room, not one you file

In a randomised trial, using rating scales to guide each treatment decision put 73.8% of people into remission versus 28.8% with usual care, and did it in half the time. The scale that pays is the same scale that helps, used differently.

not yet assessed

We have not finished checking this source. No judgement either way. how we score evidence

Caveat on this rating: One trial of 120 outpatients in a single setting, with pharmacotherapy restricted to two drugs; the review's 'virtually all' spans heterogeneous designs. The effect size is large and consistent with the wider literature, but this is not the same evidence tier as the meta-analyses in stops 2 and 3. The claim that the measures do not pay for the conversation is an observation about the measure definitions quoted in stop 4, not a finding of any study.

Significant findings

The score is not the problem. What is done with it is.

In a 24-week randomised trial of measurement-based care for major depression, 61 outpatients had their treatment adjusted by rating scales and a guideline at each visit, and 59 had clinicians decide as usual. Response was 86.9% versus 62.7%; remission was 73.8% versus 28.8%. Remission arrived in 10.2 weeks against 19.2. The measured group had more treatment adjustments (44 versus 23) and higher doses. Symptoms fell in both groups; they fell further where the score was acted on.

A 2017 review of 51 studies drew the line where this trail has been drawing it: "Virtually all randomized controlled trials with frequent and timely feedback of patient-reported symptoms to the provider during the medication management and psychotherapy encounters significantly improved outcomes. Ineffective approaches included one-time screening, assessing symptoms infrequently, and feeding back outcomes to providers outside the context of the clinical encounter." The same review notes that aggregated scores can "inform payers about the value of mental health services", which is the doorway the payment measures walked through.

Put the two halves together. The tool Pfizer paid for is accurate enough to flag, not to diagnose. On its own it changes nothing. Attached to a clinician who reads it with you and adjusts what you are doing, it roughly doubles remission in a trial. The federal measures pay for the score's existence and for a number under five a year later; they do not pay for the conversation in between, which is the part that works.

Worth asking

You can turn a form into the useful version yourself. When you are handed a PHQ-9: what is my number, what was it last time, and what would we change if it has not moved by five points? Those three questions are measurement-based care. They cost nothing, and nobody bills for them.

Source

Measurement-Based Care Versus Standard Care for Major Depression: A Randomized Controlled Trial With Blind Raters — Guo T, Xiang YT, Xiao L, Hu CQ, Chiu HF, Ungvari GS, et al. (2015)

Read the source: https://doi.org/10.1176/appi.ajp.2015.14050652

DOI: 10.1176/appi.ajp.2015.14050652

How this was scored

Study design
not recorded
Funding
not recorded
Published in
not recorded
Sample size
not recorded
Preregistered
not recorded
Conflicts disclosed
not recorded
Independent of proponent
not recorded
Retracted
No

We have not finished checking this source, so there is no scoring to show yet. “We have not checked this yet” and “this is disputed” are different statements, so no scored band is shown rather than a low one.

Read the full scoring rubric, including what it can't tell you.

Published September 15, 2026.

Questions

How strong is the evidence behind this?

veisund rates this source "not yet assessed". We have not finished checking this source. No judgement either way. One caveat travels with that badge: One trial of 120 outpatients in a single setting, with pharmacotherapy restricted to two drugs; the review's 'virtually all' spans heterogeneous designs. The effect size is large and consistent with the wider literature, but this is not the same evidence tier as the meta-analyses in stops 2 and 3. The claim that the measures do not pay for the conversation is an observation about the measure definitions quoted in stop 4, not a finding of any study. The score is calculated from recorded facts about the source — study design, funding, publication venue, sample size, preregistration — not typed in by an editor.

What is the source for this?

Measurement-Based Care Versus Standard Care for Major Depression: A Randomized Controlled Trial With Blind Raters — Guo T, Xiang YT, Xiao L, Hu CQ, Chiu HF, Ungvari GS, et al. (2015). DOI: 10.1176/appi.ajp.2015.14050652. The full source is linked on this page so you can read it yourself.

Is this medical advice?

This is information to bring to your prescriber, not medical advice and not a reason to change anything on your own. Nothing here is an instruction to stop or reduce a medication. If you are in crisis, call or text 988.

This is information to bring to your prescriber, not medical advice and not a reason to change anything on your own. Nothing here is an instruction to stop or reduce a medication. If you are in crisis, call or text 988.