A score of 10 is a flag, not a diagnosis

Across 58 studies and 17,357 people, a PHQ-9 of 10 or more had sensitivity 0.88 and specificity 0.85. In published meta-analyses, questionnaires put depression at 31% and diagnostic interviews at 17%.

not yet assessed

We have not finished checking this source. No judgement either way. how we score evidence

Caveat on this rating: The screen's accuracy depends on the reference standard: the 0.88/0.85 figures are against semistructured clinician interviews; against fully structured or MINI interviews the authors report different values. The prevalence gap (31% vs 17%) is across heterogeneous populations and does not by itself show that any one clinic over-diagnoses; it shows what happens when a screen is read as a verdict.

Significant findings

The largest test of the PHQ-9 as a screen pooled individual data from 58 studies, 17,357 people and 2,312 cases of major depression diagnosed by interview. At the usual cut-off of 10 or more, against a clinician-administered interview, sensitivity was 0.88 (95% confidence interval 0.83 to 0.92) and specificity 0.85 (0.82 to 0.88). That is a good screening instrument. It is not a diagnosis, and the people who ran that analysis said so in a second paper: depression questionnaires "are not designed to ascertain diagnostic status and, based on published sensitivity and specificity estimates, would theoretically be expected to overestimate prevalence."

Then they measured the overestimate. Across 69 meta-analyses of depression prevalence, the pooled figure was 31% when studies had used screening or rating tools, 22% for mixtures, and 17% when they had used diagnostic interviews. Of 2,094 primary studies, 77% used a questionnaire and 13% a validated interview; 71% of the questionnaire-based meta-analyses still described their result as "depression" or "depressive disorders".

The arithmetic follows from the numbers above, and it is ours, not the papers': with sensitivity 0.88 and specificity 0.85, in 1,000 people of whom 150 have major depression, about 132 of them screen positive, and so do about 128 people who do not. Roughly half of positive screens are not major depression. That is what a screen is for; it is not what a diagnosis is.

Worth asking

Every payment measure in the stops ahead starts from a PHQ-9 above nine. If a score of 10 or more is treated as depression rather than as a reason to look closer, half the people counted may not have it. If your score was over nine, was there a conversation after the form, or only a code?

Source

Accuracy of Patient Health Questionnaire-9 (PHQ-9) for screening to detect major depression: individual participant data meta-analysis — Levis B, Benedetti A, Thombs BD; DEPRESsion Screening Data (DEPRESSD) Collaboration (2019)

Read the source: https://doi.org/10.1136/bmj.l1476

DOI: 10.1136/bmj.l1476

How this was scored

Study design
not recorded
Funding
not recorded
Published in
not recorded
Sample size
not recorded
Preregistered
not recorded
Conflicts disclosed
not recorded
Independent of proponent
not recorded
Retracted
No

We have not finished checking this source, so there is no scoring to show yet. “We have not checked this yet” and “this is disputed” are different statements, so no scored band is shown rather than a low one.

Read the full scoring rubric, including what it can't tell you.

Published September 15, 2026.

Questions

How strong is the evidence behind this?

veisund rates this source "not yet assessed". We have not finished checking this source. No judgement either way. One caveat travels with that badge: The screen's accuracy depends on the reference standard: the 0.88/0.85 figures are against semistructured clinician interviews; against fully structured or MINI interviews the authors report different values. The prevalence gap (31% vs 17%) is across heterogeneous populations and does not by itself show that any one clinic over-diagnoses; it shows what happens when a screen is read as a verdict. The score is calculated from recorded facts about the source — study design, funding, publication venue, sample size, preregistration — not typed in by an editor.

What is the source for this?

Accuracy of Patient Health Questionnaire-9 (PHQ-9) for screening to detect major depression: individual participant data meta-analysis — Levis B, Benedetti A, Thombs BD; DEPRESsion Screening Data (DEPRESSD) Collaboration (2019). DOI: 10.1136/bmj.l1476. The full source is linked on this page so you can read it yourself.

Is this medical advice?

This is information to bring to your prescriber, not medical advice and not a reason to change anything on your own. Nothing here is an instruction to stop or reduce a medication. If you are in crisis, call or text 988.

This is information to bring to your prescriber, not medical advice and not a reason to change anything on your own. Nothing here is an instruction to stop or reduce a medication. If you are in crisis, call or text 988.