The evidence library

Plain-English summaries of the research behind mental health treatment — each one pointed at a named source and scored for how much weight that source can bear. Free to read. No account, no email, no paywall.

This page is the study library. The full catalog — 37 medication pages, the legal & safety records, the counting and funding investigations, the glossary, and the tools — lives on the library map →

How this library is verified

No citation, no publishing. Every resource points at a named source, and the trust badge is calculated from recorded facts about that source — study design, who funded it, where it was published, sample size, preregistration — never typed in by hand. Retracted papers score zero and are blocked from publishing.

resources published
329

resources published

with a source link
329

with a source link

with a verified DOI
301

with a verified DOI

retracted sources
0

retracted sources

dead source links
0

dead source links

A source link that quietly 404s is a broken citation whether or not anyone notices, so we stopped relying on noticing. As of 19 August 2026, all 231 distinct citation links in this library resolved — 214 DOIs confirmed registered at Crossref and 17 regulator and PubMed pages fetched directly. That check re-runs against production every morning and fails loudly the day one of them breaks, so this figure is monitored rather than remembered.

Read the full scoring rubric — including what it can't tell you.

Recently added

Intermittent theta burst beat sham on response and remission, with no more headache, dropout or mania

Across 23 randomised trials in 960 people, intermittent theta burst to the left prefrontal cortex beat sham on response and remission, with no excess of dropout, headache or switch to mania.

strong evidence73/100

Caveat on this rating: Small trials pooled across six protocols that differ in target and dose; the pooled safety comparisons (k = 7 for mania, k = 10 for headache) are thin. Favourable does not mean settled.

published September 15, 2026source 2024DOI verifiedread →

Theta burst most likely the best rTMS form, in a network where three of 141 trials were low risk of bias

A 2026 network meta-analysis put theta burst first among rTMS forms, then said the effects were modest, the intervals overlapped, and most rankings were low or very low confidence.

strong evidence70/100

Caveat on this rating: The strongest thing this paper says is about its own evidence base: 138 of 141 trials at high risk of bias or with concerns. Any ranking built on that is provisional, which the authors state. The 2019 BMJ network in this wave, with a different trial set, likewise found overlapping intervals among the TMS forms.

published September 15, 2026source 2026DOI verifiedread →

Accelerated TMS worked faster than standard TMS, not better, in four randomised and ten before-and-after studies

Depression scores fell after accelerated TMS, but against standard TMS the difference was not significant (SMD -0.67, 95% CI -1.62 to 0.27); follow-up was short and the maintenance signal is provisional.

moderate evidence62/100

Caveat on this rating: Four randomised trials is a thin base for a head-to-head conclusion, and the null difference has a wide interval (-1.62 to 0.27) that does not rule out either direction. The maintenance signal rests on short follow-up, as the authors say.

published September 15, 2026source 2024DOI verifiedread →

The Stanford accelerated protocol, replicated: remission in 50% versus 21% on sham at one month, in 48 people

The second randomised trial of Stanford neuromodulation therapy enrolled 53 people with treatment-resistant depression and randomised 48; at one month, 50.0% on active treatment were in remission versus 20.8% on sham.

moderate evidence55/100

Caveat on this rating: A 48-person trial with a one-month primary endpoint, from the developers, following a 29-person trial from the same group. The FDA clearance record in this wave notes that effectiveness of the device "has not been established beyond the timepoints evaluated" and that the trials behind it were single-site. Early, promising, and not yet independently replicated.

published September 15, 2026source 2026DOI verifiedread →

In adolescents, the only large randomised trial found TMS no better than sham

Ten studies of TMS for adolescent depression, two of them randomised: the pooled 41% response came mainly from uncontrolled studies, and "the only large-scale randomized trial suggests TMS is not more effective than sham stimulation".

moderate evidence56/100

Caveat on this rating: The pooled effect and the null RCT point in different directions only if uncontrolled studies are read as evidence of efficacy, which the authors do not do. The evidence base in adolescents is small, and the one adequately powered sham-controlled trial was negative.

published September 15, 2026source 2022DOI verifiedread →

In older adults, TMS response ranged from 6.7% to 54.3% across trials, and the review calls efficacy still unclear

Seven randomised trials (260 people) and seven uncontrolled ones: response to TMS in geriatric depression ran from 6.7% to 54.3%, with "substantial variability" and "large heterogeneity" in dose and protocol.

early signal53/100

Caveat on this rating: Seven small RCTs with widely different doses cannot be added up, and the authors do not try. The absence of a pooled number is the finding.

published September 15, 2026source 2022DOI verifiedread →

Medication approval journeys

37 medications

A different kind of source from the studies above: the regulatory record. One page per medication answering what its FDA approval was actually based on — which trials, how long they ran, and what the label says, quoted rather than paraphrased. Where the record does not state something, the page says so instead of estimating.

Read across all of them: how long these drugs were actually tested and who paid for the research.

Psilocybin, MDMA, ketamine and the rest of the pipeline. Every resource here carries its actual FDA development stage, so the hype and the regulatory reality sit side by side.

Microdosing with psilocybin mushrooms: a double-blind placebo-controlled study.

This study found the only reliable difference from placebo was in people who knew they had taken the active dose, and no gain in creativity or cognition — what am I actually expecting microdosing to do for me?

strong evidence78/100

Caveat on this rating: With 34 participants the study is underpowered to detect small true effects, and the authors' own Bayesian analysis in the successfully blinded subgroup was inconclusive rather than clearly negative.

Not FDA approved for this useNot FDA-evaluated for this usemicrodosingmicrodoselsd microdosing
published August 19, 2026source 2022DOI verifiedread →

Efficacy and Safety of Psychedelic Microdosing on Psychological Outcomes in Healthy Adults: A Systematic Review and…

The randomised evidence shows microdosing performing no better than placebo for depression, anxiety, or stress — what treatment with an actual demonstrated effect should I be trying first?

strong evidence72/100

Caveat on this rating: The meta-analytic estimates rest on very little randomised data — two parallel RCTs (three comparisons, n=117) for efficacy and two RCTs (n=109) for adverse events — so the confidence intervals are wide and the review is better read as demonstrating absence of evidence than proving absence of effect. Several authors are affiliated with psychedelic research centres and one is chief medical officer of a psychedelics company (disclosed), so the null result is not coming from sceptics.

Not FDA approved for this useNot FDA-evaluated for this usemicrodosingmicrodoselsd microdosing
published August 19, 2026source 2026DOI verifiedread →

Self-blinding citizen science to explore psychedelic microdosing.

In the largest placebo-controlled microdosing study, people taking empty capsules improved as much as people taking the drug — how do I tell whether what I am feeling is the substance or the expectation?

strong evidence70/100

Caveat on this rating: Participants sourced and weighed their own material, so actual doses were unverified, and the sample was self-selected enthusiasts rather than people with a diagnosed condition. The authors are based at Imperial College's Centre for Psychedelic Research with a co-author from the Beckley Foundation, both proponents of psychedelic research — which cuts against, not toward, a bias explanation for this null result.

Not FDA approved for this useNot FDA-evaluated for this usemicrodosingmicrodoselsd microdosing
published August 19, 2026source 2021DOI verifiedread →

A systematic study of microdosing psychedelics.

Observational microdosing data show people expect far more benefit than users actually report, and one measure — neuroticism — went up. Is that a trade I would knowingly make?

moderate evidence64/100

Caveat on this rating: Uncontrolled and unblinded: participants knew they were dosing, sourced their own substances, and were self-selected enthusiasts, so nothing here can separate drug effect from expectation. The authors say as much and explicitly call for dose- controlled research.

Not FDA approved for this useNot FDA-evaluated for this usemicrodosingmicrodoselsd microdosing
published August 19, 2026source 2019DOI verifiedread →

Ketamine versus ECT for Nonpsychotic Treatment-Resistant Major Depression

If I am a candidate for ECT, is IV ketamine a reasonable first option given this head-to-head result?

gold standard87/100

Caveat on this rating: Open-label by necessity (neither ketamine nor ECT can be masked), so subjective outcomes may favor the treatment patients preferred; population excluded psychotic depression, where ECT performs best.

Not FDA approved for this useketamineketalar
published August 19, 2026source 2023DOI verifiedread →

Safety and efficacy of methylenedioxymethamphetamine (MDMA)-assisted psychotherapy in post-traumatic stress disorder:…

Why do independent evidence graders rate this "low to very low certainty" when the raw effect sizes look so large?

strong evidence81/100

Caveat on this rating: This umbrella review's STRONG tier reflects its independent, systematic method — its actual conclusion is that the underlying MDMA evidence is weak. Read the tier as confidence in the review, not in MDMA.

Not FDA approved for this useFDA approval declinedmidomafetaminemdma
published August 19, 2026source 2025DOI verifiedread →
see all 25 in Psychedelic-assisted therapy

Medications

225 resources

Antidepressants, mood stabilisers, stimulants and sedatives — approvals, label warnings, withdrawal, side effects and the trials behind them.

Strategies for managing sexual dysfunction induced by antidepressant medication.

If I want to stay on this antidepressant, is adding sildenafil or twice-daily bupropion an option for me, or would switching drugs be the better first move?

gold standard87/100

Caveat on this rating: The declarations of interest read literally: "KH has previously acted as a temporary consultant for Pfizer (manufacturers of sildenafil). MT has been paid to lecture and received travel expenses from Bristol-Myers Squibb (manufacturers of buspirone) and Otsuka; his spouse is an employee of GlaxoSmithKline (manufacturers of bupropion)." Two of the six authors therefore have financial ties to the makers of the two drugs the review endorses, which is why independence is scored false despite the funding being public. The reviewers also warn that partial reporting of subscale results in the source trials could bias the effect estimates upward.

sildenafiltadalafilbupropion (Wellbutrin, Zyban, Aplenzin, Forfivo)
published August 19, 2026source 2013DOI verifiedread →

PRAC recommendations on signals adopted at the 13-16 May 2019 PRAC meeting, section 1.3: SNRIs and SSRIs - persistent…

European regulators require a label warning that sexual side effects can persist after stopping an SSRI or SNRI — how would you and I tell that apart from the depression itself if it happened to me?

strong evidence70/100

Caveat on this rating: A regulatory label change is a precautionary act on a safety signal, not proof of incidence, causation, or permanence. The evidence base the PRAC weighed was EudraVigilance reports, literature, social media, and manufacturer reviews. Clomipramine and vortioxetine were part of the same signal assessment but were explicitly not covered by the labelling recommendation. US FDA labelling does not carry equivalent wording.

citalopram (Celexa)escitalopram (Lexapro, Cipralex)fluvoxamine (Luvox)
published August 19, 2026source 2019read →

Treatment-emergent sexual dysfunction related to antidepressants: a meta-analysis.

Where does the antidepressant I'm on sit on the sexual side-effect ranking, and is there a drug with a lower rate that would still treat my condition?

moderate evidence63/100

Caveat on this rating: The authors state that including open-label studies and pooling across different sexual-function scales "could reduce the significance of our findings," so the 25.8-80.3% band is wide partly because the source studies were not uniform. No funding statement or conflict-of-interest disclosure was accessible for this paper.

sertraline (Zoloft)venlafaxine (Effexor)citalopram (Celexa)
published August 19, 2026source 2009DOI verifiedread →

Psychosocial interventions for erectile dysfunction.

Adding structured therapy to an erectile-dysfunction medication worked better than the medication alone in a Cochrane review — can we do both rather than picking one?

gold standard95/100

Caveat on this rating: The high trust score reflects the review's method, funding, and journal, not the strength of the underlying trials: 398 men across 11 studies, several from the 1970s and 1980s, with waitlist controls and confidence intervals touching 1.0. The single trial claiming group therapy outperformed sildenafil alone (WMD -12.40 on the IIEF) had 20 participants and was run by one of the review's own authors, which is a direct conflict on the review's most eye-catching result.

sildenafil
published August 19, 2026source 2007DOI verifiedread →

Valerian for sleep: a systematic review and meta-analysis.

The main positive finding for valerian came from six trials that showed signs of publication bias — is it worth my trying it, or should we go straight to something with better evidence for my insomnia?

gold standard85/100

Caveat on this rating: The high trust score reflects the quality of this meta-analysis as a piece of evidence synthesis, not the strength of the effect it found. The pooled benefit rests on a subjective, dichotomised outcome in six trials with demonstrated publication bias, and the authors' own conclusion is the hedged "valerian might improve sleep quality," followed by a call for better studies. Twenty years on, those studies have largely not materialised.

Not FDA approved for this useNot FDA-evaluated for this usevalerianvaleriana officinalisvalerian root
published August 19, 2026source 2006DOI verifiedread →

A systematic review of valerian as a sleep aid: safe but not effective.

The best-designed valerian studies all found no effect on sleep — what treatment for my insomnia actually has evidence behind it?

strong evidence79/100

Caveat on this rating: This review reaches the opposite conclusion to Bent 2006 from overlapping trials, and the difference is instructive: Taibi weighted trial quality and recency, and found that the better and newer the study, the smaller the effect. That pattern is the signature of a treatment whose apparent benefit comes from weak methods.

Not FDA approved for this useNot FDA-evaluated for this usevalerianvaleriana officinalisvalerian root
published August 19, 2026source 2007DOI verifiedread →
see all 225 in Medications

Fertility

15 resources

Research on psychiatric treatment and fertility — the questions people ask their prescriber too late.

The risk with no label: perinatal mental illness and maternal death

Across 38 US states in 2020, mental health conditions were the most frequent underlying cause of pregnancy-related death - 115 of the 511 deaths for which a cause was identified. Committees judged 83.5% of all such deaths preventable.

not yet assessed

Caveat on this rating: Definitions do the heavy lifting here and are stated rather than assumed. 'Pregnancy-related' means causally related to the pregnancy and within one year of its end - a wider window than the 42-day definition used in international comparisons, which is why UK figures for suicide as a direct cause within 42 days (0.80 per 100,000) look nothing like the 6-week-to-1-year figures quoted. The MMRC category bundles suicide with overdose and poisoning related to substance use disorder, so it is not a suicide count. The 2020 US figure covers 38 states, not the nation, and a 2020 coding change moved substance-use-related overdose deaths into this category, which is part of why 22.5% is not comparable with the 11% from 2008-17. A separate CDC report gives 26% for mental health conditions as a CONTRIBUTING circumstance rather than an underlying cause; that figure answers a different question and is not used. None of these numbers say anything about the risk facing any individual reader.

published September 10, 2026source 2021DOI verifiedread →

The risk of stopping is a measured risk too

In a cohort of women with recurrent depression, 68% of those who discontinued relapsed during pregnancy against 26% of those who continued. Nobody randomised them, and that limitation is the whole argument about the number.

not yet assessed

Caveat on this rating: Every study in this row is observational and self-selected, and that cuts against the row's own direction as much as for it. Cohen 2006 recruited at three specialist perinatal psychiatry centres, so its population is women with severe recurrent illness rather than a general one, and its hazard ratio of 5.0 cannot be read as the causal effect of stopping. Grote 2010 measured antenatal depression rather than treatment status, and its effect sizes swing with the measurement instrument and the country. Jarde 2016's own conflict-of-interest subgroup nearly doubles the preterm-birth odds ratio, which is a warning about the literature and is printed here for that reason. No randomised trial of continuation versus discontinuation exists, so 'stopping is riskier' and 'continuing is riskier' rest on the same grade of evidence - which is the point of the stop, and why it closes on a conversation rather than a conclusion.

published September 10, 2026source 2006DOI verifiedread →

Antidepressants in the first trimester: the cardiac-defect signal mostly vanishes once you adjust for the depression

7.23 cardiac defects per 1,000 infants with no exposure; 9.02 after first-trimester SSRI exposure. Adjust for the depression the medication was treating and the difference falls to 1.06 (0.93-1.22).

not yet assessed

Caveat on this rating: What this row does not claim: that SSRIs in pregnancy are risk-free. Adjustment is not randomisation, and the propensity-score models can only correct for what was measured in claims data - smoking, alcohol, illness severity and body-mass index are poorly captured in Medicaid records, and residual confounding runs in both directions. The primary-PPHN result still clears one after adjustment. Brown 2017's sibling analysis has wide intervals rather than a null, and its authors wrote that a causal relationship cannot be ruled out. Sujan 2017's surviving preterm-birth association may itself reflect illness severity within a family rather than the drug. Levinson-Castiel is a single centre with 60 exposed infants. The defensible reading is that the large early relative risks shrank substantially under better designs, that the absolute excess risks are small where they persist, and that neither 'proven harmful' nor 'proven safe' is what the evidence supports.

published September 10, 2026source 2014DOI verifiedread →

Lithium in the first trimester: the heart-defect risk is real, dose-driven, and smaller than psychiatry taught

A 1970s register of voluntarily reported cases produced a 400-fold figure, on the basis of two cases. A cohort of 1.3 million pregnancies later measured 2.41 cardiac malformations per 100 lithium-exposed births against 1.15 unexposed.

not yet assessed

Caveat on this rating: The uncertainty here is real and is printed rather than resolved. Patorno's adjusted risk ratio for cardiac malformations is 1.65 with a lower bound of 1.02, and the right ventricular outflow tract estimate is 2.66 with a lower bound of exactly 1.00 - both only just clear chance, and the lithium numerator for that outcome is suppressed under the Medicaid data-use agreement, so this row makes no claim about how many Ebstein cases the study did or did not observe. Patorno excluded Ebstein anomaly as a named outcome on purpose, because clinicians may be likelier to code a defect as Ebstein in an infant known to have been lithium-exposed. Munk-Olsen 2018's first-trimester malformation result is significant while its any-time and cardiac-specific results are not. Fornaro 2020's meta-analysis and Diav-Citrin 2014 point the same direction with wide intervals. The honest summary is a small absolute increase, dose-related, on a very small base - not the register's 400-fold, and not nothing.

published September 10, 2026source 2017DOI verifiedread →

Valproate: the clearest signal in the field, and its size

Two to four in 100 babies are born with a major birth defect with no medication involved at all. In a registry of 1,381 pregnancies on valproate alone it was 10.3 in 100 - and it rose steeply with the dose.

not yet assessed

Caveat on this rating: Named limits. EURAP is a prospective registry of women in epilepsy care, not a randomised comparison, and women prescribed valproate differ from women prescribed lamotrigine in ways registries cannot fully adjust away - the strongest single caution on every figure here. NEAD is observational and its age-6 arm retained 224 of 311 children. Christensen's whole-cohort autism estimate (4.42% vs 1.53%) is confounded by indication and the epilepsy-restricted hazard ratio of 1.7 (0.9-3.2) is NOT statistically significant; both are printed so the reader can see the difference the comparison group makes. The North American registry's 9.3% valproate figure is consistent with EURAP but its published abstract carries no confidence interval, so it is not quoted here. What is not in dispute across registries, cohorts and three regulators is the direction, the dose-dependence and the rough size.

published September 10, 2026source 2018DOI verifiedread →

The pregnancy letter categories were retired in 2015

A, B, C, D, X. The FDA removed them from prescription labelling because they were being read as a severity scale, which they were never built to be. A great many people were taught them and still believe they exist.

not yet assessed

Caveat on this rating: What is NOT established: that the narrative PLLR labels are better understood by clinicians in practice. No study shows that, and the Namazy survey points the other way - 68% of respondents did not find the narrative concise, and 95% still used the letters. This row asserts the documented change and the documented reasons the agency gave for it, not that the replacement works. The Namazy survey is also a 12% response rate (184/1500) among allergy/immunology clinicians, not a general prescriber sample. The Addis comparison is from 2000 and concerns three systems, two of which still exist.

published September 10, 2026source 2014read →
see all 15 in Fertility

Sexual health

13 resources

Sexual side effects, what the trials measured, and what they left out.

Strategies for managing sexual dysfunction induced by antidepressant medication.

If I want to stay on this antidepressant, is adding sildenafil or twice-daily bupropion an option for me, or would switching drugs be the better first move?

gold standard87/100

Caveat on this rating: The declarations of interest read literally: "KH has previously acted as a temporary consultant for Pfizer (manufacturers of sildenafil). MT has been paid to lecture and received travel expenses from Bristol-Myers Squibb (manufacturers of buspirone) and Otsuka; his spouse is an employee of GlaxoSmithKline (manufacturers of bupropion)." Two of the six authors therefore have financial ties to the makers of the two drugs the review endorses, which is why independence is scored false despite the funding being public. The reviewers also warn that partial reporting of subscale results in the source trials could bias the effect estimates upward.

sildenafiltadalafilbupropion (Wellbutrin, Zyban, Aplenzin, Forfivo)
published August 19, 2026source 2013DOI verifiedread →

PRAC recommendations on signals adopted at the 13-16 May 2019 PRAC meeting, section 1.3: SNRIs and SSRIs - persistent…

European regulators require a label warning that sexual side effects can persist after stopping an SSRI or SNRI — how would you and I tell that apart from the depression itself if it happened to me?

strong evidence70/100

Caveat on this rating: A regulatory label change is a precautionary act on a safety signal, not proof of incidence, causation, or permanence. The evidence base the PRAC weighed was EudraVigilance reports, literature, social media, and manufacturer reviews. Clomipramine and vortioxetine were part of the same signal assessment but were explicitly not covered by the labelling recommendation. US FDA labelling does not carry equivalent wording.

citalopram (Celexa)escitalopram (Lexapro, Cipralex)fluvoxamine (Luvox)
published August 19, 2026source 2019read →

Treatment-emergent sexual dysfunction related to antidepressants: a meta-analysis.

Where does the antidepressant I'm on sit on the sexual side-effect ranking, and is there a drug with a lower rate that would still treat my condition?

moderate evidence63/100

Caveat on this rating: The authors state that including open-label studies and pooling across different sexual-function scales "could reduce the significance of our findings," so the 25.8-80.3% band is wide partly because the source studies were not uniform. No funding statement or conflict-of-interest disclosure was accessible for this paper.

sertraline (Zoloft)venlafaxine (Effexor)citalopram (Celexa)
published August 19, 2026source 2009DOI verifiedread →

Psychosocial interventions for erectile dysfunction.

Adding structured therapy to an erectile-dysfunction medication worked better than the medication alone in a Cochrane review — can we do both rather than picking one?

gold standard95/100

Caveat on this rating: The high trust score reflects the review's method, funding, and journal, not the strength of the underlying trials: 398 men across 11 studies, several from the 1970s and 1980s, with waitlist controls and confidence intervals touching 1.0. The single trial claiming group therapy outperformed sildenafil alone (WMD -12.40 on the IIEF) had 20 participants and was run by one of the review's own authors, which is a direct conflict on the review's most eye-catching result.

sildenafil
published August 19, 2026source 2007DOI verifiedread →

A randomized trial comparing group mindfulness-based cognitive therapy with group supportive sex education and therapy…

Group therapy for low desire improved things for about half of women in a trial, with no drug involved — is there a group programme or a sex therapist you can refer me to?

strong evidence76/100

Caveat on this rating: Both arms improved similarly on the primary desire and arousal outcomes, so this trial does not show that mindfulness specifically works — it shows that eight weeks of structured group attention works. There was no no-treatment or waitlist arm, so regression to the mean and expectancy effects cannot be separated out. All 148 participants were cisgender women, and the authors note the need to diversify samples.

published August 19, 2026source 2021DOI verifiedread →

Efficacy of psychological interventions for sexual dysfunction: a systematic review and meta-analysis.

Talking therapy for sexual problems has moderate evidence behind it — is there a therapist trained in sex therapy you can refer me to, rather than starting with a pill?

moderate evidence62/100

Caveat on this rating: Every pooled effect is against a waitlist, not an active or placebo control, so attention and expectancy are baked into the d = 0.58 figure. The search ends at 2009, predating the internet-delivered treatments that now dominate access. No total participant count is reported in the abstract, and the MEDLINE record lists the publication type "Research Support, Non-U.S. Gov't", indicating unnamed external support.

published August 19, 2026source 2013DOI verifiedread →
see all 13 in Sexual health

Mechanism-level research on what actually moves depression, anxiety, sleep and connection — and how big the effect really is.

Intermittent theta burst beat sham on response and remission, with no more headache, dropout or mania

Across 23 randomised trials in 960 people, intermittent theta burst to the left prefrontal cortex beat sham on response and remission, with no excess of dropout, headache or switch to mania.

strong evidence73/100

Caveat on this rating: Small trials pooled across six protocols that differ in target and dose; the pooled safety comparisons (k = 7 for mania, k = 10 for headache) are thin. Favourable does not mean settled.

published September 15, 2026source 2024DOI verifiedread →

Theta burst most likely the best rTMS form, in a network where three of 141 trials were low risk of bias

A 2026 network meta-analysis put theta burst first among rTMS forms, then said the effects were modest, the intervals overlapped, and most rankings were low or very low confidence.

strong evidence70/100

Caveat on this rating: The strongest thing this paper says is about its own evidence base: 138 of 141 trials at high risk of bias or with concerns. Any ranking built on that is provisional, which the authors state. The 2019 BMJ network in this wave, with a different trial set, likewise found overlapping intervals among the TMS forms.

published September 15, 2026source 2026DOI verifiedread →

Accelerated TMS worked faster than standard TMS, not better, in four randomised and ten before-and-after studies

Depression scores fell after accelerated TMS, but against standard TMS the difference was not significant (SMD -0.67, 95% CI -1.62 to 0.27); follow-up was short and the maintenance signal is provisional.

moderate evidence62/100

Caveat on this rating: Four randomised trials is a thin base for a head-to-head conclusion, and the null difference has a wide interval (-1.62 to 0.27) that does not rule out either direction. The maintenance signal rests on short follow-up, as the authors say.

published September 15, 2026source 2024DOI verifiedread →

The Stanford accelerated protocol, replicated: remission in 50% versus 21% on sham at one month, in 48 people

The second randomised trial of Stanford neuromodulation therapy enrolled 53 people with treatment-resistant depression and randomised 48; at one month, 50.0% on active treatment were in remission versus 20.8% on sham.

moderate evidence55/100

Caveat on this rating: A 48-person trial with a one-month primary endpoint, from the developers, following a 29-person trial from the same group. The FDA clearance record in this wave notes that effectiveness of the device "has not been established beyond the timepoints evaluated" and that the trials behind it were single-site. Early, promising, and not yet independently replicated.

published September 15, 2026source 2026DOI verifiedread →

In adolescents, the only large randomised trial found TMS no better than sham

Ten studies of TMS for adolescent depression, two of them randomised: the pooled 41% response came mainly from uncontrolled studies, and "the only large-scale randomized trial suggests TMS is not more effective than sham stimulation".

moderate evidence56/100

Caveat on this rating: The pooled effect and the null RCT point in different directions only if uncontrolled studies are read as evidence of efficacy, which the authors do not do. The evidence base in adolescents is small, and the one adequately powered sham-controlled trial was negative.

published September 15, 2026source 2022DOI verifiedread →

In older adults, TMS response ranged from 6.7% to 54.3% across trials, and the review calls efficacy still unclear

Seven randomised trials (260 people) and seven uncontrolled ones: response to TMS in geriatric depression ran from 6.7% to 54.3%, with "substantial variability" and "large heterogeneity" in dose and protocol.

early signal53/100

Caveat on this rating: Seven small RCTs with widely different doses cannot be added up, and the authors do not try. The absence of a pooled number is the finding.

published September 15, 2026source 2022DOI verifiedread →
see all 19 in What actually works

Medications, withdrawal, side effects and regulatory action — what the label says, what the trials measured, and what nobody studied.

What is actually known about how these drugs work

Not much, honestly. An SSRI raises available serotonin within hours; people do not feel better for weeks. The two leading explanations for that gap are both unsettled, and the field says so in print.

not yet assessed

Caveat on this rating: This row asserts an absence, which is the claim most easily overstated. What is established is that the delayed onset is not explained by the immediate neurochemical change, and that the field's own reviews present competing accounts rather than a settled one. What is NOT established is that either account here is correct: the emotional-processing model and the neuroplasticity model are both active research programmes with supporting and non-supporting findings, and this row deliberately reports that they are unresolved instead of ranking them. The claim that no clinical test measures a chemical imbalance is a statement about available assays, not a claim that no biological contribution exists. A narrative review is also a lower grade of evidence than a systematic one, and it is cited here for what the field says about its own uncertainty rather than as a pooled result.

published September 10, 2026source 2017DOI verifiedread →

What the explanation does, and what nobody has measured

Biological explanations reduce blame and increase pessimism about recovery - both at once, in the same studies. What nobody has measured is what happens when the explanation is taken away.

not yet assessed

Caveat on this rating: This literature does not point one way, and the row would be dishonest if it implied otherwise. A randomised study from the same research group as the bogus-test experiment found self-blame and perceived helplessness GREATER after a cognitive-behavioural explanation than a biological one (Lee et al. 2016, J Soc Clin Psychol 35(7):571-588), and a pre-registered sham-genetic-test experiment found no effect on beliefs about depression at all (Lamontagne et al. 2023, PMID 36869258). Both are named here and neither is quoted, because their full texts were not obtained. The Kvaale effect sizes are small to moderate, the dangerousness result may reflect publication bias by the authors' own account, and the samples across this field are dominated by undergraduates and online panels rather than people in treatment. The absence of any study of debunking, or of any behavioural outcome, is stated as an absence found by search - it is not proof that no such work exists anywhere.

published September 10, 2026source 2013DOI verifiedread →

Whether it works and how it works are different questions

Aspirin was sold for about seventy years before anyone could say how it worked. Whether a drug helps is settled by trials, not by the story told about the mechanism - and these trials are large.

not yet assessed

Caveat on this rating: Every efficacy number here is contested at the edges, and none of the disputes touches the conclusion this stop rests on. Cipriani 2018's own authors graded certainty moderate to very low and 9% of included trials at high risk of bias; its between-drug ranking is not established and is not used here. Stone 2022's three-class result comes from a statistical model, not an observed grouping, and its reading has been argued over since publication - though its co-author list includes the field's most prominent placebo-effect sceptic, which cuts against overstating benefit rather than for it. Both meta-analyses are dominated by short, largely industry-sponsored acute-phase trials in adults meeting formal criteria for major depression, and neither measures long-term outcomes. Cowen and Browning's emotional-processing account is a hypothesis with supporting experiments, not a settled mechanism. The ANTLER figures describe trial conditions, not any individual reader's risk.

published September 10, 2026source 2015DOI verifiedread →

The umbrella review, and the fight about it

A 2022 review of 17 studies found no support for depression being caused by low serotonin. Thirty-five authors replied in the same journal that the review was too leaky to show it. Both stand; neither is withdrawn.

not yet assessed

Caveat on this rating: This stop is the dispute, so the caveat is about what it does NOT establish. It does not establish that the serotonin hypothesis is disproved: the review's own framing is absence of support, and Jauhar and colleagues argue that its method manufactures that absence. It does not establish that the rebuttal is correct either. The two sides also disagree about what tryptophan depletion is for - whether the test is inducing depression in healthy people, or worsening it in people with a history - and that disagreement is unresolved rather than adjudicated here. Four further published critiques and a second author reply appeared in 2024 and are named in the body only as "a second exchange", because their full texts are paywalled with no PMC deposit and this library does not quote a paper it has only read about. Declared interests are printed for both sides precisely because neither set of interests decides the science.

published September 10, 2026source 2023DOI verifiedread →

What the advertisements said, and what the literature said

"Prescription Zoloft works to correct this imbalance." A 2005 analysis in PLoS Medicine set the advertising beside the research and reported the research had never established the imbalance the advertising described.

not yet assessed

Caveat on this rating: The central dispute in this row is unresolved and is left that way on purpose. Pies argues that no professional body, textbook or peer-reviewed consensus ever advanced a chemical-imbalance theory and that critics cherry-pick quotations; Ang and colleagues report that highly cited reviews and every textbook they sampled did support the serotonin theory between 1990 and 2010. Two of the three authors on that side are also authors of the 2022 umbrella review, which is a stake in the outcome and is stated rather than hidden. Their paper was a sample of highly cited work, not a census, and its full text could not be obtained, so nothing beyond its published abstract is quoted here. The France 2007 survey is 251 US undergraduates in one place at one time and does not describe the general public. Nothing in this row establishes that anyone intended to mislead; intent cannot be documented from outside, and the essay itself does not claim it.

published September 10, 2026source 2005DOI verifiedread →

The risk with no label: perinatal mental illness and maternal death

Across 38 US states in 2020, mental health conditions were the most frequent underlying cause of pregnancy-related death - 115 of the 511 deaths for which a cause was identified. Committees judged 83.5% of all such deaths preventable.

not yet assessed

Caveat on this rating: Definitions do the heavy lifting here and are stated rather than assumed. 'Pregnancy-related' means causally related to the pregnancy and within one year of its end - a wider window than the 42-day definition used in international comparisons, which is why UK figures for suicide as a direct cause within 42 days (0.80 per 100,000) look nothing like the 6-week-to-1-year figures quoted. The MMRC category bundles suicide with overdose and poisoning related to substance use disorder, so it is not a suicide count. The 2020 US figure covers 38 states, not the nation, and a 2020 coding change moved substance-use-related overdose deaths into this category, which is part of why 22.5% is not comparable with the 11% from 2008-17. A separate CDC report gives 26% for mental health conditions as a CONTRIBUTING circumstance rather than an underlying cause; that figure answers a different question and is not used. None of these numbers say anything about the risk facing any individual reader.

published September 10, 2026source 2021DOI verifiedread →
see all 296 in Questioning the story

Follow the money

14 resources

Who paid for the study, who sells the treatment, and what that does to the result.

A score of 10 is a flag, not a diagnosis

Across 58 studies and 17,357 people, a PHQ-9 of 10 or more had sensitivity 0.88 and specificity 0.85. In published meta-analyses, questionnaires put depression at 31% and diagnostic interviews at 17%.

not yet assessed

Caveat on this rating: The screen's accuracy depends on the reference standard: the 0.88/0.85 figures are against semistructured clinician interviews; against fully structured or MINI interviews the authors report different values. The prevalence gap (31% vs 17%) is across heterogeneous populations and does not by itself show that any one clinic over-diagnoses; it shows what happens when a screen is read as a verdict.

published September 15, 2026source 2019DOI verifiedread →

Screening by itself does not change what happens next

Across 16 randomised trials in 7,576 patients, giving clinicians a depression questionnaire raised recognition a little and changed outcomes not at all: "little or no impact". What changes outcomes is what is built around the score.

not yet assessed

Caveat on this rating: The trials pooled here predate most collaborative-care programmes; the USPSTF's B grade rests on evidence that screening programmes with support improve outcomes, which is not in conflict with Gilbody's null for screening alone. Both are stated. Neither shows that screening harms.

published September 15, 2026source 2008DOI verifiedread →

How a score became a payment: the measures

MIPS measure 370 pays on a PHQ-9 below 5 at twelve months. Measure 134 pays on a screen plus a documented follow-up plan within two days. Health plans report the same PHQ-9 measures to NCQA. The score is now a unit of account.

not yet assessed

Caveat on this rating: Measure 134 is a process measure and its follow-up plan definition is broad by design; a documented referral counts, and the specification itself warns against reflexive prescribing. Measure 370 is an outcome measure and remission is a legitimate goal. The critique in this trail is about incentives around a screening score, not about the measures' stated intent, which is quoted.

published September 15, 2026source 2024read →

What the money is: nine percent, fifteen billion, and a diagnosis worth 0.39

MIPS moves up to 9% of a clinician's Medicare Part B pay. Star-rating bonuses add about $15 billion a year to Medicare Advantage. Diagnoses recorded only on health-risk assessments produced $7.5 billion in 2023 payments.

not yet assessed

Caveat on this rating: The 0.388 factor comes from a CMS table whose model version and date were not printed on the pages read; it is a pre-2024 numbering and the 2024 model reclassified the category, so treat it as the order of magnitude a coded diagnosis added, not a current price. MIPS adjustments are budget-neutral and most clinicians land near zero; the plus-or-minus 9 percent is the statutory range, not a typical outcome. The OIG figure covers diagnoses of all kinds recorded on HRAs, not depression alone.

published September 15, 2026source 2024read →

The version that works: a score you use in the room, not one you file

In a randomised trial, using rating scales to guide each treatment decision put 73.8% of people into remission versus 28.8% with usual care, and did it in half the time. The scale that pays is the same scale that helps, used differently.

not yet assessed

Caveat on this rating: One trial of 120 outpatients in a single setting, with pharmacotherapy restricted to two drugs; the review's 'virtually all' spans heterogeneous designs. The effect size is large and consistent with the wider literature, but this is not the same evidence tier as the meta-analyses in stops 2 and 3. The claim that the measures do not pay for the conversation is an observation about the measure definitions quoted in stop 4, not a finding of any study.

published September 15, 2026source 2015DOI verifiedread →

The depression questionnaire was paid for by a drug company, and the form says so

The PHQ-9 footer reads: "Developed by Drs. Robert L. Spitzer, Janet B.W. Williams, Kurt Kroenke and colleagues, with an educational grant from Pfizer Inc." The 2001 validation paper says the same.

not yet assessed

Caveat on this rating: Funding a tool is not the same as biasing it: the PHQ-9's accuracy has since been tested in an independent individual-participant meta-analysis of 58 studies (stop 2), and the tool is free precisely because of the grant. What the grant bought is reach, not a rigged score. The GAD-7 funding statement was read from the instrument's own attribution, not from the 2006 paper's text.

published September 15, 2026source 2001DOI verifiedread →
see all 14 in Follow the money

What this evidence adds up to

Each of these reads the relevant studies side by side — effect sizes, sample sizes, who funded what — and links back to every resource it draws on.

Does TMS work for depression?

what the 113-trial comparison of brain stimulation found, why the clinic success rate is not the trial success rate, and what the Stanford accelerated protocol has and has not shown.

Does ketamine work for depression?

the pooled odds at 24 hours and how certain they are, racemic versus esketamine, how long it lasts, the licensing trial that failed, the critics in their own words, and three ketamine-versus-ect studies that disagree.

Does ECT work for depression?

the strongest odds of response in the 113-trial network, the critics quoted in their own words, what the memory evidence measures and what it cannot, the one-in-two relapse figure, and what nice and the fda say.

How many Americans have a mental illness?

the federal survey by age, 2021 to 2025, beside the hospital records: who says they are unwell, who ends up admitted, what a stay costs and who pays. twenty-seven series, each with its limits printed above the numbers.

Who wrote the other mental health questionnaires, and who paid?

twenty forms behind the prescriptions: who wrote each, who funded it where the record says, which are free and which are sold, and where the fda and medicare wrote them into rules.

Who wrote the PHQ-9, and who gets paid on your score?

the depression questionnaire was paid for by pfizer and the form says so. what a score of ten means, why screening alone changed nothing in sixteen trials, how the number became a medicare payment measure, and what the money actually is.

Do any supplements actually work for mental health?

Seventeen substances, each with its regulatory status, its best evidence, and the risks the marketing leaves out.

What are psychiatric medications actually approved for?

Read off the labels: what each drug is approved to treat, and which of its common uses are off-label.

How common is postpartum depression, really?

One in six, flat across the whole first year — and no prior history of depression protects you less than people assume.

What are the real rates of SSRI sexual side effects?

25.8% To 80.3% depending on the drug — the rates the trials found when they asked with a questionnaire instead of waiting for a spontaneous complaint.

Does ashwagandha work for anxiety?

Every trial, its size, and who paid for it — including the liver-injury case series.

Does CBD help anxiety?

The dose-response curve, the labelling-accuracy problem, and what the 2026 meta-analysis found.

What does St John’s wort interact with?

Including hormonal contraceptives — the interaction people find out about too late.

Is kratom dangerous?

One randomised trial exists. The dependence data, the poison-centre calls, and the legal patchwork.

Ketamine vs esketamine — what is actually FDA-approved?

One is an approved nasal spray under a REMS. The other is a compounded off-label prescription.

Where does each psychedelic actually stand with the FDA?

Row by row, with approval dates, rejections and trial phases — updated as the agency acts.

Questions

What is in the veisund evidence library?

Short, plain-English summaries of individual studies, drug labels and regulatory actions relevant to mental health — medications, psychedelic-assisted therapy, loneliness, and what actually moves symptoms. Each one points at a named source and carries a trust score derived from that source.

How is the trust score calculated?

From recorded facts about the source, not from an editor’s opinion: study design, who funded it, where it was published, sample size, whether it was preregistered, whether conflicts were disclosed, and whether the researchers were independent of whoever benefits from the result. The score is recomputed on every read, so a change to the rubric applies retroactively to the whole library.

What happens if a cited paper is retracted?

It scores zero and the badge says "retracted", regardless of how good the design or journal was. A retracted paper is also blocked from being published in the app at all.

Does a high score mean the finding is true?

No. It means the finding was produced by a process that is harder to fool. Good process still produces wrong answers. The score tells you how much weight the method can bear, not whether the conclusion is correct.

Do I need the app to read these?

No. Every resource page is free to read on the web with no account, no email and no paywall. The app adds saving, "plan to try" tracking, one resource a day, and the people who have actually been through it.

Is this medical advice?

This is information to bring to your prescriber, not medical advice and not a reason to change anything on your own. Nothing here is an instruction to stop or reduce a medication. If you are in crisis, call or text 988.

This is information to bring to your prescriber, not medical advice and not a reason to change anything on your own. Nothing here is an instruction to stop or reduce a medication. If you are in crisis, call or text 988.