Improving Replicability of Neuroimaging Research

Simon Vandekar

Definitions

Replicability: ability to obtain consistent results across independent studies with different data aimed at answering the same scientific question.


Reproducibility: ability to obtain the identical results across independent studies with the same dataset aimed at answering the exact same scientific question.

  • Replicability is related to statistics, reproducibility is informatics

Overview

  • Replicability in neuroimaging
  • Challenges to replicability
  • Strategies to improve replicability
  • Effect sizes

Why effect sizes?

Larger effect sizes cause better replicability

“Parameters” versus “Estimates”

Three-panel diagram: the parameter is the fixed, unknown truth in the population; two different samples give two different estimates while the parameter never moves; a published result is one estimate drawn from the scatter of estimates around the parameter, not the truth.
  • Smaller samples give estimates that scatter farther from the parameter

Biological effect sizes

  • Governed by factors we cannot control
  • Units are determined by the variables themselves
    • \(\beta\) coefficients, odds ratios, relative risks

Standardized effect sizes

  • A unitless statistical index that summarizes the strength of an association
    • Cohen’s \(d\), Pearson’s correlation, \(R^2\), Cohen’s \(f\), \(\eta^2\)
  • Complimentary to reporting \(p\)-values and biological effect sizes

Standardized effect size scale

  • I will use effect sizes on the scale Cohen’s \(f \in [0, \infty)\), realistically, \(f<1\).
  • Cohen (behavioral scientist) recommended standardized benchmarks
    • These are only recommendations
Effect size Cohen’s \(f\)
Small 0.10
Medium 0.25
Large 0.40

Standardized effect size parameters drive replicability

Replicability (in statistics, a probability)
Standardized effect size (e.g. Pearson's correlation)
Study design
Underlying biological
effect size (e.g. β coefficient)
Sample size

Quiz: parameter/estimate? biological/standardized?

  • Hippocampal volume difference (in mm³) between individuals with schizophrenia and healthy adults in all humans  parameter, biological – units mm³, and it is the unknown truth in the population
  • ENIGMA reports \(d = -0.46\) for hippocampal volume in schizophrenia from 4,568 scans  estimate, standardized – unitless, and computed from their particular sample
  • Your study’s fitted slope of 0.8 mL of cortical volume lost per year of age  estimate, biological – mL/year units, but it came from your sample
  • The correlation between amyloid PET SUVR and memory score you would get if you scanned everyone  parameter, standardized – unitless, but defined on the population, not a sample

Parameter vs. estimate is about population vs. sample; biological vs. standardized is about the units.

Most neuroimaging studies are not replicable

Unreliability of Biomarker Research in Psychiatry

Vandekar, S., Zhang, P., Patel. P.K., Jefferson, A.L., Harrell, F. (under review)

Over decade of evidence of low replicability

Histogram of median statistical power across 49 neuroscience meta-analyses

Button 2013, Fig. 3 — median power across 49 meta-analyses (730 studies) was 21%; most studies fell below 20%.

Jaccard overlap between split-half task-fMRI maps as a function of sample size

Turner 2018, Fig. 2 — overlap between thresholded maps from two independent samples never reaches 0.6, even at N = 121.

Distributions of brain-wide association effect sizes and sampling variability by sample size

Marek 2022, Fig. 1 — brain–phenotype correlations cluster near zero (median |r| = 0.01); at N = 25 the same effect can come out r = +0.5 or −0.6.

Projected relative gain in prediction accuracy moving from 1,000 to 1 million participants

Schulz 2024, Fig. 3 — going from 1,000 to 1,000,000 participants buys only a 3–9× gain in prediction accuracy.

Challenges to replicability in neuroimaging

  • Inadequate sample sizes
  • Multitude of analytic decisions
  • Neglecting clinical information
  • Misleading uncertainty quantification

Inadequate sample sizes

  • Neuroimaging samples are often under ~100
Biomarker / association Reported effect Cohen's f Sample Original study Failed replication
FDG-PET right anterior insula treatment-selection biomarker Cohen's f = 0.80* 0.80* n = 38 primary analysis McGrath 2013, JAMA Psychiatry Kelley 2021, prospective PET test
fMRI cognitive-control connectivity → antidepressant response d=1.14–1.37 0.57–0.685 n = 40 Sertraline group Tozzi 2020, Biological Psychiatry
EEG abnormality → antidepressant response OR = 3.56 (escitalopram); OR = 2.76 (venlafaxine-XR) 0.35; 0.31 n = 622 per-protocol patients Arns 2017, Clinical EEG and Neuroscience Reveles Jensen, 2024
rTMS → brain-function variability in schizophrenia Local BOLD: F1,36 = 5.83; MCD: F1,36 = 32.57 0.40; 0.95 n = 42 (active 19; sham 23) Schifani 2025, Schizophrenia Bulletin

Cohen’s benchmarks for \(f\): 0.10 small, 0.25 medium, 0.40 large. If the first effect size were accurate, the study would be powered with less than 6 participants in each group.

  • Small samples are highly variable and lead to inaccurate published effect size estimates
  • Power analyses using published effect size estimates from small studies justify small samples in grants

Comparison to effect sizes we know

Association / intervention Reported effect Cohen's f (estimate) Evidence base
COVID-19 vaccine (BNT162b2) vs. placebo efficacy 95% 0.83 Polack 2020, NEJM — n = 43,448
Smoking → lung cancer (current vs. never) RR = 8.96 0.60 Gandini 2008, Int J Cancer — 216 studies
Psilocybin-assisted therapy for depression g = 0.84 0.42 Florineth 2026, JAMA Netw Open — 12 trials, n = 733
Antidepressants vs. placebo OR = 1.37–2.13 0.09–0.21 Cipriani 2018, Lancet — 522 trials, n = 116,477
GWAS: FTO (rs1558902) → BMI, largest common-variant effect β = 0.082 SD/allele 0.06 Locke 2015, Nature — n = 339,224
ENIGMA: hippocampal volume in schizophrenia d = −0.46 0.23 van Erp 2016, Mol Psychiatry — n = 4,568
ENIGMA: hippocampal volume in depression d = −0.14 0.07 Schmaal 2016, Mol Psychiatry — n = 8,927
Brain-wide association — largest replicable r = 0.16 0.16 Marek 2022, Nature — n ≈ 50,000
Brain-wide association — median r = 0.01 0.01 Marek 2022, Nature — n ≈ 50,000

Cohen’s benchmarks for \(f\): 0.10 small, 0.25 medium, 0.40 large.

  • Prespecified analyses in underpowered samples lead to inconclusive results
    • Effect size estimates are unstable
    • The CI is too wide to provide useful information
    • A nonsignificant \(p\)-value is uninformative rather than evidence of no effect

Selection from the analytic multiverse

  • Botvinik-Nezer et al. (2020, Nature): 70 independent teams tested the same hypotheses on one fMRI dataset
    • No two teams used identical analysis pipelines
  • Without specific pre-registration, small samples lead to study adjustments to “improve signal” over analytical decisions
  • Selection over many analytic steps causes biased associations
  • Difficult to detect in a published study

Analysis selection example

McGrath et al. (2013, JAMA Psychiatry) identified a PET “treatment-selection biomarker” in 38 usable participants

  • Treatment-selection biomarker – one used to identify if a patient will respond to a treatment (“treatment by biomarker” interaction)
  • The preregistered outcome was Hamilton Depression rating \(\le 7\) at week 12
  • The outcome used in the paper was quantified differently. Imaging targets were not prespecified at all
  • Effect size estimate was \(f=0.8\), a stronger association than smoking with lung cancer
McGrath paper excerpt defining remission and excluding partial responders

McGrath et al., outcome definition (extremely biased).

Neglecting clinical information

An imaging biomarker has to add information

McGrath et al. Figure 3B showing the interaction effect in the selected region

McGrath et al. (2013), Figure 3B. The interaction effect in the selected PET region.

McGrath et al. Table 1 with the baseline HAMA total score row highlighted

McGrath et al., Table 1. Baseline HAMA showed the same interaction (P = .009), suggesting clinical variables can explain the image association effect.

  • This means controlling for relevant baseline measures
    • Same McGrath paper 👆
    • Including baseline values of the outcome when predicting future variables with imaging data

Misleading uncertainty quantification

Circular analysis and data leakage Use components of the evaluation data for feature selection, preprocessing, or model tuning

  • The model can appear to predict treatment response without having any real association
  • Barac et al. (2026): Used the full dataset for feature selection.
    • They also didn’t control for baseline depression
Barac et al. feature-selection text describing selection of variables associated with baseline depression

Barac et al. (2026), feature selection. Baseline depression severity was used as the target for selecting the top features before treatment-response prediction.

  • Has the appearance of rigorous validation and strong prediction without actually doing it

More examples in the paper

Unreliability of Biomarker Research in Psychiatry

Vandekar, S., Zhang, P., Patel. P.K., Jefferson, A.L., Harrell, F. (under review)

Recommendations to improve replicability

Most imaging biomarker studies that fail external validation would have failed rigorous internal validation

  • No longer just writing for humans
  • Report effect sizes for all research results
  • Specifying research goals a priori

Stop writing like we are writing for humans

Line chart showing growth in neuroimaging publications over the last 25 years.

Figure 1

  • Each study is single data point in an ocean of data
  • Each paper is only useful insofar as it can contribute to a meta-analysis
  • Evidence will be aggregated across studies by AI

Quantitative reporting for AI

Writing for humans Writing for AI
Primary deliverable is interpretation Primary deliverable is quantitative
Purpose is conceptual Purpose is cumulative evidence synthesis
Statistics are complementary Magnitude is quantified
Magnitude rarely discussed Standardized effect sizes are reported consistently

Report effect sizes for all results

Bias comes from selective reporting

When only the “significant” regions get reported

Study 1: only the largest estimate and significant regions shown Study 2: only the largest estimate and significant regions shown Study 3: only the largest estimate and significant regions shown Study 4: only the largest estimate and significant regions shown Study 5: only the largest estimate and significant regions shown

Red = the study's strongest association · gray = other significant regions · = 95% CI excludes zero

  • The effect size estimates are biased upward because they must be large enough to be significant
  • Little shared information across datasets

In truth each study held information about all results

True region-level effect sizes across the Desikan-Killiany atlas Study 1 estimated association strength with 95% confidence intervals Study 2 estimated association strength with 95% confidence intervals Study 3 estimated association strength with 95% confidence intervals Study 4 estimated association strength with 95% confidence intervals Study 5 estimated association strength with 95% confidence intervals Study 6 estimated association strength with 95% confidence intervals Meta-analysis pooling all six studies, converging on the true effect sizes Meta-analysis of only the significant reported results, biased away from the truth

Green = true nonzero effect · black = true null · gray = study estimate ± 95% CI · red = the study's largest point estimate · = 95% CI excludes zero

  • Underpowered studies have wide, inconclusive CIs; larger studies converge on the truth and have greater precision
  • Meta-analysis pooling all six studies recovers the true signal with a tight CI
  • Meta-analysis based on biased reporting is also biased

Reporting all results

  • Counterintuitively: bias is reduced when you report ALL your uncorrected exploratory analyses (not just the significant ones)
  • For neuroimaging - upload your unthresolded statistical images with useful metadata to Neurovault
  • Include supplementary tables as PDF or online

Specify research questions clearly a priori

  • This will reduce repeated analyses to try to improve signal, which produces bias
  • This should be specific enough that AI could take the instructions with the data and perform your analyses

Standardized effect size are important

  • They drive replicability
  • They can be used for meta-analysis and power analysis

Resolving challenges with reporting effect size estimates

Scientific question:
e.g., X on Y in Z
?
Studies
Designs &
hypotheses
Statistical
analyses
Multiple GLM
Logistic
Two-group comparison
Simple GLM
Effect size
R² = 0.4 RESI = 0.13
95% CI = (…)
OR = 1.2 RESI = 0.45
95% CI = (…)
Cohen's d = 0.2 RESI = 0.1
95% CI = (…)
Pearson's r = 0.2 RESI = 0.1
95% CI = (…)
Two people confused, discussing mismatched effect size values: d=0.2?, CI?, OR=1.2? Two people happily toasting, comparing matching RESI values with confidence intervals
  • Model-specific indices from different studies/designs can't be directly compared or synthesized
  • Model-specific indices from different studies/designs can't be directly compared or synthesized
  • It's hard to get confidence intervals for some of them
  • We introduced a robust effect size index (RESI), a single, comparable effect size (with a CI) regardless model
  • Enables direct comparison and meta-analysis across studies

The robust effect size index (RESI)

  • Interpreted on the same scale as Cohen’s \(f\)
  • Signed version for regression coefficients, Unsigned version for more arbitrary effect sizes (all on the same scale)
  • “Robust” in the sense that it is consistent
  • Comes with a confidence interval
  • Our papers on this are deriving/proving/validating the properties
Effect size Cohen’s \(f\)
Small 0.10
Medium 0.25
Large 0.40

RESI R package

  • Works directly on fitted model objects (lm, glm, nls, coxph, lme/lmerMod, and more)
  • resi() returns effect sizes with confidence intervals
library(RESI)
library(splines)
mod <- lm(charges ~ (ns(age, df=3) + ns(bmi, df=3)) * smoker + sex + region, data = insurance)
insRESI <- resi(mod)
av <- anova(insRESI)
knitr::kable(av, digits = c(0, 2, 3, 3, 3, 3))
Df F Pr(>F) RESI 2.5% 97.5%
ns(age, df = 3) 3 278.02 0.000 0.788 0.695 0.878
ns(bmi, df = 3) 3 152.30 0.000 0.582 0.505 0.660
smoker 1 5386.90 0.000 2.005 1.757 2.256
sex 1 4.07 0.044 0.048 0.000 0.109
region 3 4.22 0.006 0.085 0.028 0.144
ns(age, df = 3):smoker 3 0.47 0.703 0.000 0.000 0.069
ns(bmi, df = 3):smoker 3 394.84 0.000 0.939 0.785 1.086
  • Every term gets a RESI and CI, directly comparable across models and studies

Easy visualization

library(ggplot2)
ggplot(av)

Example interpretation/reporting

Report the test statistic, degrees of freedom, \(p\)-value, and effect size with its CI for every term, e.g.:

Smoking status had a very strong effect (F(1, 1320) = 5386.9, p < .001, RESI = 2, 95% CI [1.76, 2.26]).

The nonlinear age-by-smoking interaction was negligible (F(3, 1320) = 0.5, p = 0.703, RESI = 0, 95% CI [0, 0.07]) – the entire interval falls at or below Cohen’s small benchmark (0.10).

The nonlinear bmi-by-smoking interaction was very large (F(3, 1320) = 394.8, p < .001, RESI = 0.94, 95% CI [0.79, 1.09]) – the point estimate and CI both exceed Cohen’s large benchmark (0.40) more than twofold, i.e. smoking status robustly reshapes the bmi-charges association.

  • The RESI and CI are directly comparable across models and different variables. \(p\)-values are not.

Supplemental information to include

  • In the supplement, include the full coefficient table with RESIs (or not)
# example for RESI
# knitr::kable(coefficients(insRESI), digits = 3)
# or, for a manuscript-style regression table:
sjPlot::tab_model(mod, vcov.fun = sandwich::vcovHC)
  charges
Predictors Estimates CI p
(Intercept) 3220.76 1834.51 – 4607.01 <0.001
age [1st degree] 5842.19 4621.24 – 7063.14 <0.001
age [2nd degree] 12121.52 10144.31 – 14098.74 <0.001
age [3rd degree] 11157.45 10169.47 – 12145.44 <0.001
bmi [1st degree] 1251.68 -60.95 – 2564.30 0.062
bmi [2nd degree] 1732.58 -1400.95 – 4866.12 0.278
bmi [3rd degree] -891.19 -2833.63 – 1051.25 0.368
smoker 12216.85 7886.19 – 16547.51 <0.001
sex [male] -519.96 -1025.57 – -14.35 0.044
region [northwest] -383.37 -1174.24 – 407.50 0.342
region [southeast] -994.77 -1769.34 – -220.20 0.012
region [southwest] -1178.03 -1882.75 – -473.30 0.001
age [1st degree] × smoker 784.04 -2125.99 – 3694.08 0.597
age [2nd degree] × smoker 600.72 -3291.05 – 4492.50 0.762
age [3rd degree] × smoker -941.56 -2990.07 – 1106.96 0.367
bmi [1st degree] × smoker 34647.85 31191.83 – 38103.86 <0.001
bmi [2nd degree] × smoker 22494.46 10124.51 – 34864.41 <0.001
bmi [3rd degree] × smoker 22080.08 10486.83 – 33673.34 <0.001
Observations 1338
R2 / R2 adjusted 0.854 / 0.852

Questions about statistical reporting?

  • Ask Simon Jr. – a .md instruction file to review your papers
  • Also, ask me!
    • I want help you do the best science possible.

Thank you!

NIH funding

  • R01MH123563
  • R01AG101908
QR code linking to https://simonvandekar.github.io/replicability.html

These slides are on my website:
simonvandekar.github.io/replicability.html

Students who worked on this

Xinyu Zhang

Xinyu Zhang

Megan Jones

Megan Jones

Kaidi Kang

Kaidi Kang

Colleagues

  • A.F. Alexander-Bloch
  • K. Armstrong
  • S. Avery
  • R.A.I. Bethlehem
  • J. Blume
  • B. Corbett
  • D.A. Fair
  • E. Feczko
  • Stephan Heckers
  • A.S. Keller
  • B. Larsen
  • M. McHugo
  • K. Mehta
  • O. Miranda Dominguez
  • R. Muscatello
  • S.M. Nelson
  • A. Randolph
  • T.D. Satterthwaite
  • J. Schildcrout
  • J. Seidlitz
  • R. Tao
  • Neil Woodward
  • J. Xiong
  • Gina Yu