the depression questionnaire: who wrote it, who paid, and who gets paid on your score
nine questions on a tablet in the waiting room. over the last two weeks, how often have you been bothered by little interest or pleasure in doing things. not at all, several days, more than half the days, nearly every day. you add it up, or the tablet does, and a number goes into your chart.
this page is about that number. who wrote the form, who paid for it, what a score of ten does and does not mean, why a screen on its own changes nothing, how the number became something clinics and health plans are paid on, what that money actually is, and the one way of using the score that the trials say works. twenty-one sources, every one a paper or the regulator's own document, funding recorded only where the record states it.
the form, and the footer
the phq-9 is the nine-item patient health questionnaire. each item is scored 0 to 3, the total runs 0 to 27, and the last item asks about “Thoughts that you would be better off dead or of hurting yourself in some way”[1]. the small print at the bottom of the form is the most honest thing about it: “Developed by Drs. Robert L. Spitzer, Janet B.W. Williams, Kurt Kroenke and colleagues, with an educational grant from Pfizer Inc. No permission required to reproduce, translate, display or distribute.”[1]
the paper that validated it, in the journal of general internal medicine in 2001, says the same in its acknowledgments: “The development of the PHQ-9 was underwritten by an educational grant from Pfizer US Pharmaceuticals, New York, NY.” the instrument reproduced in that paper's appendix carries the line “Copyright 1999 Pfizer Inc.”[2]
the validation was large. 3,000 patients across 8 primary care clinics and 3,000 across 7 obstetrics-gynecology clinics, with 580 re-interviewed by a mental health professional. from it came the bands still used today: “PHQ-9 scores of 5, 10, 15, and 20 represented mild, moderate, moderately severe, and severe depression, respectively”, and at a score of 10 or more “a sensitivity of 88% and a specificity of 88% for major depression”[2]. the anxiety scale that usually rides beside it, the gad-7, was built by the same group with the same attribution line and validated in 2006 in 15 clinics, 2,740 questionnaires and 965 interviews, with a cut point of 10 giving sensitivity 89% and specificity 82%[3]. both descend from prime-md, the group's 1994 clinician-administered interview, which took an average of 8.4 minutes per patient and which the self-report form was built to shorten[4].
what a score of ten is, and is not
the largest test of the phq-9 as a screen pooled individual data from 58 studies, 17,357 people and 2,312 cases of major depression diagnosed by interview. at the usual cut-off of 10 or more, against a semistructured clinician interview, sensitivity was 0.88 (95% confidence interval 0.83 to 0.92) and specificity 0.85 (0.82 to 0.88)[5]. that is a good screening instrument.
it is not a diagnosis, and the people who ran that analysis said so in a second paper. depression questionnaires “are not designed to ascertain diagnostic status and, based on published sensitivity and specificity estimates, would theoretically be expected to overestimate prevalence”[6]. then they measured the overestimate. across 69 meta-analyses of depression prevalence, the pooled figure was 31% when the studies had used screening or rating tools, 22% for mixtures, and 17% when they had used diagnostic interviews. of 2,094 primary studies, 77% used a questionnaire and 13% a validated interview, and 71% of the questionnaire-based meta-analyses still described their result as “depression” or “depressive disorders”[6].
does screening by itself change anything
sixteen randomised trials, 7,576 patients, in which depression questionnaires were administered and the results fed back to clinicians. recognition of depression rose a little (relative risk 1.27, 1.02 to 1.59), but when questionnaires were given to everyone and the results returned regardless of score, it did not move (1.03, 0.85 to 1.24). antidepressant prescribing did not change (1.20, 0.87 to 1.66). in the seven trials that measured outcomes, the effect on depression was nil: standardised mean difference −0.02 (−0.25 to 0.20)[7].
the authors' conclusion: “If used alone, case-finding or screening questionnaires for depression appear to have little or no impact on the detection and management of depression by clinicians. Recommendations to adopt screening strategies using standardized questionnaires without organizational enhancements are not justified.”[7] their cochrane review put it harder: “Practice guidelines and recommendations to adopt this strategy, in isolation, in order to improve the quality of healthcare should be resisted.”[8]
who decided the score would be paid on
nobody decided it in one place, which is the honest answer to the question. three bodies, three levels.
congress, in 2015, passed the medicare access and chip reauthorization act, macra, which created the merit-based incentive payment system, mips, and wrote its payment adjustments into the social security act[16]. cms picks and publishes the measures every year in rulemaking. two of them turn the questionnaire into money. quality id 134, “Screening for Depression and Follow-Up Plan”, is a process measure: the “Percentage of patients aged 12 years and older screened for depression on the date of the encounter or up to 14 days prior to the date of the encounter using an age-appropriate standardized depression screening tool AND if positive, a follow-up plan is documented on the date of or up to two days after the date of the qualifying encounter.” the phq-9 and phq-2 are named as acceptable tools, and the plan is defined: “Referral to a provider for additional evaluation and assessment…; Pharmacological interventions; Other interventions or follow-up for the diagnosis or treatment of depression.” the specification carries its own caution: a clinician should “Only order pharmacological intervention when appropriate and after sufficient diagnostic evaluation.”[13]
quality id 370, “Depression Remission at Twelve Months”, is an outcome measure marked “High Priority”: “The percentage of adolescent patients 12 to 17 years of age and adult patients 18 years of age or older with major depression or dysthymia who reached remission 12 months (+/- 60 days) after an index event date.” the index event is a diagnosis plus “an initial Patient Health Questionnaire - 9 item version (PHQ-9)… greater than nine”; remission is “a PHQ-9 or PHQ-9M score of less than five.” the specification is copyrighted “© MN Community Measurement, 2024”, the minnesota nonprofit that developed it, and its clinical text calls the phq-9 “an effective monitoring and management tool” and notes that “A five-point drop in PHQ-9 score is considered the minimal clinically significant difference.”[14]
ncqa, the private accreditor, writes the parallel set for health plans, hedis: screening and follow-up; “Utilization of the PHQ-9 to Monitor Depression Symptoms”, the share of members with a depression diagnosis “who had an outpatient encounter with a PHQ-9 score present in their record”; and remission or response within 4 to 8 months, remission again below 5[15]. and medicare's own coverage of the screen dates to 2011: “Annual screening up to 15 minutes”, once per 12 months, with the staff-assisted supports above[12].
what the money is
three streams run through a depression score. none of them is a per-clinician figure, and nothing primary supports one, so this page does not offer one.
the clinician's. under macra, mips adjusts medicare part b payments by an “applicable percent”: “for 2019, 4 percent; for 2020, 5 percent; for 2021, 7 percent; and for 2022 and” each year after, 9 percent, in both directions, with positive adjustments scaled by a factor that “may not exceed 3.0”[16]. cms has “set the performance threshold at 75 points through the CY 2028 performance period”; score above it and the adjustment is positive, below it and it is negative[17]. measures 134 and 370 are among the quality measures that make up that score[13],[14]. the fee-schedule value of the screen itself was not read for this page and is not stated.
the health plan's. medpac, congress's advisory commission, wrote in march 2024 that “the current system for MA quality reporting and measurement is flawed and does not provide a reliable basis for evaluating quality across MA plans. Nonetheless, these measures are the basis for the MA quality bonus program (QBP), which uses trust fund and taxpayer dollars to increase MA payments by about $15 billion annually.” forty-two percent of contracts were in bonus status for 2024, and roughly three-quarters of enrollees sit in plans rated four stars or higher[18]. the commission's wider finding: medicare “spends an estimated 22 percent more for MA enrollees than it would spend if those beneficiaries were enrolled in FFS Medicare, a difference that translates into a projected $83 billion in 2024”, and “MA plans' diagnostic coding practices increase payments and distort the goal of plans competing to improve quality”[18].
the diagnosis itself. medicare advantage plans are paid more for sicker enrollees, and a coded diagnosis raises the payment. in one cms relative-factor table, the category “Major Depressive, Bipolar, and Paranoid Disorders” carried a factor of 0.388 for a community-dwelling, non-dual, aged enrollee, on a scale where the average person is 1.0[21]. the hhs inspector general found that “Diagnoses reported only on enrollees' HRAs and HRA-linked chart reviews, and not on any other 2022 service records, resulted in an estimated $7.5 billion in MA risk-adjusted payments for 2023”, about two-thirds of it from in-home assessments; of its three recommendations, cms agreed to one[19]. in its 2024 model cms reclassified the category: “Major Depressive, Bipolar, and Paranoid Disorders” became “HCC 155 Major Depression, Moderate or Severe, without Psychosis”, under a stated principle that “Diagnoses that are particularly subject to intentional or unintentional discretionary coding variation or inappropriate coding by health plans/providers… should not increase cost predictions”[20].
the version that works
the score is not the problem. what is done with it is. in a 24-week randomised trial of measurement-based care for major depression, 61 outpatients had their treatment adjusted by rating scales and a guideline at each visit, and 59 had clinicians decide as usual. response was 86.9% versus 62.7%. remission was 73.8% versus 28.8%, and it arrived in 10.2 weeks against 19.2. the measured group had more treatment adjustments, 44 versus 23[11]. one trial, one setting, two drugs allowed; not the same tier as the meta-analyses above, but consistent with them.
the 2017 review of 51 studies drew the line where this page has been drawing it: “Virtually all randomized controlled trials with frequent and timely feedback of patient-reported symptoms to the provider during the medication management and psychotherapy encounters significantly improved outcomes.” the same review notes that aggregated scores can “inform payers about the value of mental health services”, which is the doorway the payment measures walked through[10].
put the halves together. the tool pfizer paid for is accurate enough to flag, not to diagnose[2],[5]. on its own it changes nothing[7]. attached to a clinician who reads it with you and adjusts what you are doing, it roughly doubles remission in a trial[11]. the federal measures pay for the score's existence and for a number under five a year later[13],[14]; they do not pay for the conversation in between, which is the part that works.
what to ask when you are handed the form
these are questions, not advice. they are the ones the trials make it reasonable to ask, and together they are measurement-based care. they cost nothing, and nobody bills for them.
- what is my number, and what was it last time?
- if it is over nine, is this a flag or a diagnosis, and what would confirm which?[5]
- what would we change if it has not dropped by five points?[14]
- is there someone whose job is to follow up on this score, or does the form go in the file?[7],[12]
- is this screen part of my care, or part of a report?
this is not medical advice. it is a summary of published research, it is not a diagnosis, and it is not a recommendation for or against any treatment — nobody here has met you. decisions about starting, changing or stopping a medication belong to you and a prescriber who knows your history. do not change a prescribed medication on the strength of a web page, this one included.
last verified . if a source is updated, corrected or retracted, this page gets changed and re-dated.
sources
primary sources only — no news write-ups, no secondary summaries. each was fetched and checked on the access date shown.
[1] Spitzer RL, Williams JBW, Kroenke K and colleagues (form footer). Patient Health Questionnaire-9 (PHQ-9), the instrument. The PHQ-9 form, as hosted by the American Psychological Association, 1999.
the instrument itself: nine items scored 0 to 3 over the last two weeks, with its attribution footer · evidence tier: strong · funding: industry; the footer reads "Developed by Drs. Robert L. Spitzer, Janet B.W. Williams, Kurt Kroenke and colleagues, with an educational grant from Pfizer Inc. No permission required to reproduce, translate, display or distribute." · accessed September 14, 2026
the catch: a form is not a study. it is cited here for what it says about itself, and for the fact that the grant is also why it is free to use.
[2] Kroenke K, Spitzer RL, Williams JB. The PHQ-9: validity of a brief depression severity measure. Journal of General Internal Medicine, 2001. doi:10.1046/j.1525-1497.2001.016009606.x. PMID 11556941.
criterion validation in 3,000 patients across 8 primary care clinics and 3,000 across 7 obstetrics-gynecology clinics, with 580 re-interviewed by a mental health professional · n = 6,000 · evidence tier: strong · funding: industry; "The development of the PHQ-9 was underwritten by an educational grant from Pfizer US Pharmaceuticals, New York, NY." the instrument in the appendix is marked "Copyright 1999 Pfizer Inc." · accessed September 14, 2026
the catch: the criterion standard was a telephone interview by a mental health professional in a subset of 580 patients, not the whole sample; the cut-points were set by the authors and have been used unchanged since.
[3] Spitzer RL, Kroenke K, Williams JB, Löwe B. A brief measure for assessing generalized anxiety disorder: the GAD-7. Archives of Internal Medicine, 2006. doi:10.1001/archinte.166.10.1092. PMID 16717171.
criterion validation in 15 primary care clinics, November 2004 to June 2005; 2,740 adults completed the questionnaire and 965 had a telephone interview with a mental health professional · n = 2,740 · evidence tier: strong · funding: industry; the instrument’s own attribution reads "developed by Drs. Robert L. Spitzer, Janet B.W. Williams, Kurt Kroenke and colleagues, with an educational grant from Pfizer Inc." (read on the University of Washington-hosted instrument page, https://www.hiv.uw.edu/page/mental-health-screening/gad-7, not in the 2006 abstract) · accessed September 14, 2026
the catch: the same group, the same funder, and the same design as the phq-9; the cut point of 10 gave sensitivity 89% and specificity 82%, so about one in five people without generalised anxiety disorder still screen positive.
[4] Spitzer RL, Williams JB, Kroenke K, Linzer M, deGruy FV, Hahn SR, Brody D, Johnson JG. Utility of a new procedure for diagnosing mental disorders in primary care: the PRIME-MD 1000 study. JAMA, 1994. doi:10.1001/jama.1994.03520220043029. PMID 7966923.
validation of the clinician-administered predecessor in 1,000 adult patients at 4 primary care clinics with 31 physicians; average 8.4 minutes per patient · n = 1,000 · evidence tier: moderate · funding: unknown (the funding statement was not read in the record available to us; we do not assert who paid for PRIME-MD) · accessed September 14, 2026
the catch: the lineage, not the phq-9 itself. cited so the reader can see the questionnaire descends from a clinician-administered interview that took eight minutes, which the self-report form was built to shorten.
[5] Levis B, Benedetti A, Thombs BD; DEPRESsion Screening Data (DEPRESSD) Collaboration. Accuracy of Patient Health Questionnaire-9 (PHQ-9) for screening to detect major depression: individual participant data meta-analysis. BMJ, 2019. doi:10.1136/bmj.l1476. PMID 30967483.
individual participant data meta-analysis of 58 studies (of 72 eligible), 17,357 participants, 2,312 cases of major depression by diagnostic interview · n = 17,357 · evidence tier: gold standard · funding: unknown (not stated in the abstract we read) · accessed September 14, 2026
the catch: the 0.88 / 0.85 figures are against semistructured clinician interviews (29 studies, 6,725 participants); against fully structured interviews the estimates differ. a screen’s accuracy depends on what it is measured against.
[6] Levis B, Yan XW, He C, Sun Y, Benedetti A, Thombs BD. Comparison of depression prevalence estimates in meta-analyses based on screening tools and rating scales versus diagnostic interviews: a meta-research review. BMC Medicine, 2019. doi:10.1186/s12916-019-1297-6. PMID 30894161.
meta-research review of 69 meta-analyses (81 prevalence estimates) drawing on 2,094 primary studies · evidence tier: strong · funding: unknown (not stated in the abstract we read) · accessed September 14, 2026
the catch: the 31% versus 17% gap is across heterogeneous populations and study designs; it shows what happens when a screen is read as a verdict, not that any particular clinic over-diagnoses.
[7] Gilbody S, Sheldon T, House A. Screening and case-finding instruments for depression: a meta-analysis. CMAJ, 2008. doi:10.1503/cmaj.070281. PMID 18390942.
meta-analysis of 16 randomised trials in which depression questionnaires were administered and results fed back to clinicians · n = 7,576 · evidence tier: strong · funding: unknown (not stated in the abstract we read) · accessed September 14, 2026
the catch: the trials predate most collaborative-care programmes, and only seven measured patient outcomes. the null is for screening alone; it is not a finding that screening harms.
[8] Gilbody S, House AO, Sheldon TA. Screening and case finding instruments for depression. Cochrane Database of Systematic Reviews, 2005. doi:10.1002/14651858.CD002792.pub2. PMID 16235301.
cochrane systematic review of randomised trials of routinely administered case-finding or screening questionnaires · evidence tier: strong · funding: unknown (not stated in the abstract we read) · accessed September 14, 2026
the catch: later withdrawn from the cochrane library as out of date, which is why the 2008 cmaj update is the anchor and this is cited only for its own sentence.
[9] US Preventive Services Task Force. Screening for Depression and Suicide Risk in Adults: US Preventive Services Task Force Recommendation Statement. JAMA, 2023. doi:10.1001/jama.2023.9297. PMID 37338872.
guideline recommendation based on a commissioned systematic review · evidence tier: strong · funding: n/a (federal task force) · accessed September 14, 2026
the catch: a B grade, "moderate net benefit", for screening adults including pregnant, postpartum and older adults; an I statement (insufficient evidence) for suicide-risk screening. the grade rests on screening programmes with support in place, which is not in conflict with the gilbody null for screening alone.
[10] Fortney JC, Unützer J, Wrenn G, Pyne JM, Smith GR, Schoenbaum M, Harbin HT. A Tipping Point for Measurement-Based Care. Psychiatric Services, 2017. doi:10.1176/appi.ps.201500439. PMID 27582237.
narrative review of 51 articles on measurement-based care in behavioural health · evidence tier: moderate · funding: unknown (not stated in the abstract we read) · accessed September 14, 2026
the catch: a narrative review by advocates of measurement-based care, not a systematic one; its "virtually all" spans heterogeneous designs. cited for the distinction it draws, which the randomised evidence supports.
[11] Guo T, Xiang YT, Xiao L, Hu CQ, Chiu HF, Ungvari GS, Correll CU, Lai KY, Feng L, Geng Y, Feng Y, Wang G. Measurement-Based Care Versus Standard Care for Major Depression: A Randomized Controlled Trial With Blind Raters. American Journal of Psychiatry, 2015. doi:10.1176/appi.ajp.2015.14050652. PMID 26315978.
24-week randomised trial with blind raters; 61 outpatients on measurement-based care versus 59 on standard care · n = 120 · evidence tier: moderate · funding: unknown (not stated in the abstract we read) · accessed September 14, 2026
the catch: one trial in one setting, with pharmacotherapy restricted to two drugs; the effect is large and consistent with the wider literature but this is not the same evidence tier as the meta-analyses above.
[12] Centers for Medicare & Medicaid Services. National Coverage Determination 210.9: Screening for Depression in Adults. Medicare Coverage Database, 2011.
the coverage decision, effective 14 October 2011, that made annual depression screening a covered medicare service · evidence tier: strong · funding: n/a (regulator) · accessed September 14, 2026
the catch: covers "annual screening up to 15 minutes" only "when staff-assisted depression care supports are in place"; what a screen pays under the fee schedule was not read and is not stated on this page.
[13] Centers for Medicare & Medicaid Services, Quality Payment Program. Quality ID #134: Preventive Care and Screening: Screening for Depression and Follow-Up Plan (2024 MIPS clinical quality measure specification, version 8.0). MIPS measure specification, 2023.
process measure specification: description, denominator, numerator, tool examples, follow-up plan definition and the G-codes · evidence tier: strong · funding: n/a (regulator) · accessed September 14, 2026
the catch: the specification’s own caution is printed on the page: "Only order pharmacological intervention when appropriate and after sufficient diagnostic evaluation." a documented referral satisfies the measure; it does not require a prescription.
[14] Centers for Medicare & Medicaid Services, Quality Payment Program; measure copyright MN Community Measurement. Quality ID #370 (CBE 0710): Depression Remission at Twelve Months (2025 MIPS clinical quality measure specification, version 9.0). MIPS measure specification, 2024.
outcome measure specification, marked "High Priority": index event, remission definition, G-codes, and the steward’s clinical recommendation text · evidence tier: strong · funding: n/a (regulator; the measure is copyrighted "© MN Community Measurement, 2024") · accessed September 14, 2026
the catch: the steward’s reasons for choosing the phq-9 are on a login-walled page we could not read, so this page does not state them. remission is a legitimate outcome; the critique is about incentives around a screening score, not the measure’s intent.
[15] National Committee for Quality Assurance. HEDIS Depression Measures Specified for Electronic Clinical Data. ncqa.org, 2026.
the measure steward’s own descriptions of the three health-plan depression measures: screening and follow-up, phq-9 utilisation, and remission or response · evidence tier: moderate · funding: n/a (private accreditor; ncqa’s own funding was not examined for this page) · accessed September 14, 2026
the catch: the page does not say which plans must report these measures or which of them feed medicare star ratings, and this page does not claim either.
[16] United States Congress. Medicare Access and CHIP Reauthorization Act of 2015, Public Law 114-10, section 101(c): Merit-based Incentive Payment System. Public Law 114-10 (govinfo.gov), 2015.
the statute text: the applicable percent schedule, the negative floor, and the scaling-factor ceiling in Social Security Act section 1848(q)(6) · evidence tier: strong · funding: n/a (statute) · accessed September 14, 2026
the catch: the plus-or-minus 9 percent is the statutory range; mips is budget-neutral and most clinicians land close to zero, so it is a ceiling, not a typical outcome.
[17] Centers for Medicare & Medicaid Services. Calendar Year 2026 Quality Payment Program Final Rule Fact Sheet. qpp.cms.gov resource library, 2025.
the regulator’s summary of the final rule, including the mips performance threshold · evidence tier: strong · funding: n/a (regulator) · accessed September 14, 2026
the catch: read for one sentence: the performance threshold is "75 points through the CY 2028 performance period/2030 MIPS payment year". the fact sheet summarises the rule; the rule governs.
[18] Medicare Payment Advisory Commission (MedPAC). Report to the Congress: Medicare Payment Policy, March 2024, Chapter 12: The Medicare Advantage program: status report. MedPAC Report to the Congress, 2024.
congress’s independent advisory commission’s annual status report; pages 357 to 359 and 392 to 393 read · evidence tier: strong · funding: n/a (congressional advisory body) · accessed September 14, 2026
the catch: the $15 billion quality-bonus figure and the 22 percent / $83 billion figures are the commission’s estimates and are contested by the plans; medpac’s methods are published and its critics’ are not on this page.
[19] US Department of Health and Human Services, Office of Inspector General. Medicare Advantage: Questionable Use of Health Risk Assessments Continues To Drive Up Payments to Plans by Billions (OEI-03-23-00380). HHS OIG report, 2024.
analysis of 2022 medicare advantage encounter data to estimate 2023 risk-adjusted payments from diagnoses reported only on health risk assessments · evidence tier: strong · funding: n/a (inspector general) · accessed September 14, 2026
the catch: the $7.5 billion covers diagnoses of every kind recorded only on assessments, not depression alone; cms concurred with one of the three recommendations, and the report does not say which conditions drove the total.
[20] Centers for Medicare & Medicaid Services. Report to Congress: Risk Adjustment in Medicare Advantage, December 2024. CMS report to Congress, 2024.
the regulator’s account of the 2024 (V28) risk model; pages 17 to 20 and 28 to 30 read, including the model principles and the mental-health reclassification table · evidence tier: strong · funding: n/a (regulator) · accessed September 14, 2026
the catch: the report describes what cms did and why; it does not quantify the payment effect of moving depression from hcc 59 to hcc 155, and neither does this page.
[21] Centers for Medicare & Medicaid Services. Revised CMS-HCC Model Relative Factors for Community and Institutional Beneficiaries. CMS-HCC model relative factor tables, 2020.
the relative-factor table for the community and institutional segments; pages 1 to 6 read · evidence tier: moderate · funding: n/a (regulator) · accessed September 14, 2026
the catch: the model version and date were not printed on the pages we read, and the year here is our best reading of the file, not a stamp on it. it uses the pre-2024 numbering (hcc 58 "Major Depressive, Bipolar, and Paranoid Disorders", factor 0.388 for a community, non-dual, aged enrollee), which the 2024 model replaced. treat the figure as the order of magnitude a coded diagnosis added, not a current price.
questions
Who developed the PHQ-9?
Robert L. Spitzer, Janet B.W. Williams and Kurt Kroenke and colleagues, as the self-report successor to their 1994 clinician-administered PRIME-MD interview. The validation paper was published in the Journal of General Internal Medicine in 2001 with data from 6,000 patients in 15 clinics. The same group built the GAD-7 anxiety scale, validated in 2006.
Who paid for the PHQ-9 and GAD-7?
Pfizer. The form’s own footer reads "with an educational grant from Pfizer Inc.", the 2001 paper says "The development of the PHQ-9 was underwritten by an educational grant from Pfizer US Pharmaceuticals", and the instrument in the paper’s appendix is marked "Copyright 1999 Pfizer Inc." The GAD-7 carries the same attribution. The same grant is why both forms are free: no permission is required to reproduce, translate, display or distribute them.
Does a PHQ-9 score of 10 mean I have depression?
No. It means the screen is positive. Across 58 studies and 17,357 people, a score of 10 or more had sensitivity 0.88 and specificity 0.85 against a clinician interview. Using those figures, in 1,000 people of whom 150 have major depression, about 132 of them screen positive and so do about 128 people who do not, so roughly half of positive screens are not major depression. A diagnosis needs a conversation, not a form.
Does depression screening actually help?
On its own, no. Across 16 randomised trials in 7,576 patients, giving clinicians questionnaire results changed depression outcomes not at all (standardised mean difference −0.02). The same authors’ Cochrane review said recommendations to adopt screening in isolation "should be resisted". Attached to a system that acts on the score, yes: the US Preventive Services Task Force gives adult screening a B grade, Medicare covers it only when staff-assisted supports are in place, and a randomised trial of measurement-based care put 73.8% of people into remission versus 28.8% with usual care.
Who decided the PHQ-9 would be a quality measure?
Three different bodies, at three levels. Congress created the Merit-based Incentive Payment System in the 2015 MACRA statute. CMS selects and publishes the measures each year; the depression remission measure it uses (Quality ID 370) was developed and is copyrighted by MN Community Measurement, a Minnesota nonprofit, and the screening measure (Quality ID 134) names the PHQ-9 among acceptable tools. For health plans, NCQA writes the parallel HEDIS measures. Medicare’s coverage of annual depression screening dates to a 2011 national coverage determination.
How much do clinicians make by closing depression care gaps?
Nobody publishes a per-clinician figure and this page does not invent one. What the records show is program-level: MIPS moves a clinician’s Medicare Part B payments by up to 9 percent in either direction, with the performance threshold set at 75 points through 2028; Medicare Advantage star-rating bonuses add about $15 billion a year to plan payments, by MedPAC’s estimate; and diagnoses recorded only on health risk assessments produced an estimated $7.5 billion in 2023 risk-adjusted payments, per the HHS Inspector General. Most of that money is the plan’s, not the clinician’s. The fee-schedule value of a single screen was not read and is not stated here.
What is the minimal clinically significant change on the PHQ-9?
Five points, according to the clinical recommendation text in the MIPS remission measure specification, which cites Trivedi 2009. Remission under that measure is a score below 5; the index event that starts the clock is a diagnosis plus a score over 9.
related on veisund
the people who can tell you what a score did to their care are the ones who have filled it in
veisund is free peer support, and you are anonymous to everyone you talk to. you post under a handle you pick — no real name, no phone number, nothing that follows you back to your real life 🤍
get veisund — it's free