Psychometric Assistant
Neuropsychological calculators for clinical practice, with APA-formatted output
Score Converter
AACN = American Academy of Clinical Neuropsychology · Ranges shown as Standard Score (SS)
Clinical Outcomes Table
SD mode: * ≥1 SD below, ** ≥1.5 SD, *** ≥2 SD. SEE mode: * below 90% CI, ** below 95% CI, *** below 99% CI lower bound.
| # | Subtest | Raw | Score | CI | Percentile | Classification |
|---|
Input Type
Effect Size Tools
Convert between effect-size metrics or derive them from group data, with a visual against the standard normal.
Group comparison at a target value
Use these published effect sizes to anchor your results. The values below span small to huge magnitudes so you can compare your finding against familiar clinical and epidemiological benchmarks.
| Heavy smokers (30+/day) vs never smokers, lung cancer (Pesch et al., 2012) | 2.60 |
| UK male vs female adult height (UK Biobank; Lui et al., 2021) | 2.04 |
| Smokers (any) vs never smokers, lung cancer (Pesch et al., 2012) | 1.75 |
| Cognitive therapy vs control for PTSD (Watts et al., 2013) | 1.63 |
| Former smokers vs never smokers, lung cancer (Pesch et al., 2012) | 1.10 |
| Exposure therapy vs control for PTSD (Watts et al., 2013) | 1.08 |
| EMDR vs control for PTSD (Watts et al., 2013) | 1.00 |
| Clozapine vs placebo for schizophrenia (Huhn et al., 2019 Lancet) | 0.89 |
| CBT vs control for depression (Cuijpers et al., 2023 World Psychiatry) | 0.80 |
| Methylphenidate vs placebo for ADHD, children (Storebø et al., 2023 Cochrane) | 0.75 |
| CBT vs placebo for anxiety disorders (Hofmann & Smits, 2008) | 0.70 |
| CBT for depression, low-risk-of-bias subset (Cuijpers et al., 2023) | 0.60 |
| Interpersonal Therapy for depression (Cuijpers et al., 2011) | 0.50 |
| Antidepressants vs placebo (Cipriani et al., 2018 Lancet) | 0.30 |
| CBT vs treatment-as-usual for chronic pain (Williams et al., 2020 Cochrane) | 0.20 |
| CBT vs active control for chronic pain (Williams et al., 2020 Cochrane) | 0.10 |
| Sugar on children's hyperactivity (Wolraich et al., 1995 JAMA) | 0.00 |
| No or negligible effect | 0.00 |
Enter scores on Score Tables, Change Analysis, the SD Index or the premorbid page to chart them here.
How these charts are drawn
The page covers four sources — Score Tables, the premorbid predictions, Change Analysis and the SD Index — each as a pane of its own, so the page does not grow into a long scroll. Every trial and subtest becomes a row of its test's chart, showing the score against classification bands, or the change against the reliable-change interval. The charts read the same data and settings as those tables, so the two cannot disagree.
Each test from Score Tables gets one chart, and each of its trials or subtests is a row of that chart, drawn in the test's native metric (standard, T, scaled or z) — no score is converted for display, and a test that mixes metrics gets one panel per metric, never two metrics on one axis. The shaded bands behind each row are the classification ranges of the scheme selected on Score Tables (Wechsler or Guilmette et al., 2020), with band boundaries at standard scores 70, 80, 90, 110, 120 and 130 converted onto the panel's scale. The dot marks the obtained score; the whisker is the same confidence interval printed in the table's CI column (SEM = SD·√(1−r), at the CI level selected on Score Tables, with the table's rounding).
Base-rate measures (WAIS-IV Longest Span) are charted from the published cumulative table (WAIS-IV Administration and Scoring Manual, Tables C.4–C.5): each row's step line is the percentage of the normative sample obtaining each span or higher, and the dot marks the patient's span on that line. A high base rate means a common, and therefore lower, score.
Error measures (perseverations, intrusions, false positives; marked ↓) follow the Score Tables convention: the percentile is reported as obtained, while the classification — and therefore the band shading — describes performance, so those rows' bands run reversed.
Raw-score measures are listed, not plotted: a raw score has no position on a standardised axis, so no percentile or classification is derived. The obtained score and its raw-unit confidence interval are still shown.
Change Analysis charts plot each trial's two testings in score units — open circle at the first, filled dot at the second — against the shaded reliable-change interval for the selected method: the region where |obtained − expected| falls below the method's critical value × standard error, i.e. exactly where that method's table prints "No reliable change". The expected value is the first score (Jacobson & Truax), the first score plus the normative practice effect (Iverson), or the regression-predicted score (McSweeney, Crawford & Garthwaite), and the interval uses the confidence level set on that method's page. The method buttons only choose which method to draw; they change nothing on the Change Analysis pages. As in the tables, outcomes state significance only and never a direction — the signed statistic is shown alongside so the direction of movement stays visible without the app interpreting it.
SD Index charts plot each trial's change in standard-deviation units against the ±1.96 (or ±1.645) band for the significance level set on that page, using the same per-row divisor the SD Index table applies.
Premorbid charts answer the question the ToPF and OPIE-4 tabs ask: each index shows the predicted score (open circle) with its prediction interval shaded, the achieved score (filled dot) where one has been entered, and the difference with its base rate. A thin line marks the population mean of 100. The predicted values, intervals, differences and base rates are read from the cells those tables print, not recomputed — both tabs round the estimate and the margin separately so the bounds stay symmetric, and re-deriving them here would create a third place to keep in step. The Estimates tab keeps its own model-comparison forest plot; this block does not duplicate it.
Getting around. Each source — Score Tables, Premorbid, Change Analysis, SD Index — is its own pane, with a Back/Next bar beneath, so the page does not grow into one long scroll as more tests are entered. Only sources that hold data appear. Within a pane, All charts shows the whole set together, which is how the profile across a battery is read; One at a time gives a single chart the full width for a close look or a clean export, with the left and right arrow keys paging through it.
Showing the scores a different way. The Score Tables block offers four axes. Native metric (the default) converts nothing. Percentile and Standard score put every measure of a test on one axis, which is what makes a whole battery comparable — the score column still shows the value as entered, only the axis position is converted, and where a test mixes metrics each row carries its own metric tag. A percentile axis is deliberately non-linear: it compresses the tails, so two clearly different low scores can sit close together, and symmetric confidence intervals become asymmetric. Raw scores simply shows the raw scores as entered, so their spread is visible directly; nothing is derived from them — no norms, no classification bands, no percentile. Its axis spans the values entered for that test rather than each measure's possible range, because the app holds no raw maximum for any measure, so where a test's measures are counted on different scales their positions are not comparable with each other. A value that falls outside any chart's axis is drawn at the edge, dimmed and marked with a caret, with its actual figure alongside.
Premorbid-comparison asterisks, when enabled on Score Tables, carry the same meaning here as in the table's classification column.
Standard Deviation Index
Quantify abnormality of test-retest discrepancy in standard-deviation units. Useful when reliability data are unavailable or for descriptive comparison.
Score Type
Basic Reliable Change Index
Jacobson & Truax (1991). Computes whether observed change exceeds measurement error, using the test's reliability coefficient and standard deviation.
Test data & patient scores
| # | Subtest | SD | r | Date 1 | Date 2 | RCI (z) | p | Outcome |
|---|
Practice Effect-Adjusted Reliable Change Index
Iverson (2001). Adjusts the standard RCI to control for the average improvement (practice effect) observed between assessments in the normative sample.
Test data & patient scores
| # | Subtest | M₁ | SD₁ | M₂ | SD₂ | r | Date 1 | Date 2 | RCI (z) | p | Outcome |
|---|
McSweeney Regression-Based (SRB) Reliable Change Index
McSweeney et al. (1993). Predicts each patient's expected retest score from their baseline and the normative sample's regression parameters; the residual is standardised against the standard error of estimate.
Test data & patient scores
| # | Subtest | M₁ | SD₁ | M₂ | SD₂ | r | Date 1 | Date 2 | Ŷ₂ | RCI (z) | p | Outcome |
|---|
Crawford Regression-Based Reliable Change Index
Crawford & Garthwaite (2007). Extends the standardised regression-based approach to use a t-distributed test statistic that incorporates the normative sample size (N), correctly accounting for uncertainty in the regression parameters when N is modest. Returns a sample-size-adjusted standard error of prediction.
Test data & patient scores
| # | Subtest | M₁ | SD₁ | M₂ | SD₂ | r | N | Date 1 | Date 2 | Ŷ₂ | t(RB) | p | Outcome |
|---|
Premorbid Estimate
Inputs
Enter whichever predictors are available. Leave unavailable fields blank; the estimate table will update only for models with enough information.
Figure. Premorbid FSIQ estimates with 90% confidence intervals.
Enter the patient's actual WAIS-IV / WMS-IV index scores in the Achieved column to compute ToPF-predicted vs actual discrepancies. Base rates are shown only for negative discrepancies (achieved < predicted), and are estimated from a normal model with SD = SEE rather than transcribed from observed standardisation-sample frequencies.
| Index | Predicted | Lower 90% | Upper 90% | Achieved | Difference | Base rate |
|---|---|---|---|---|---|---|
| WAIS-IV | ||||||
| Full Scale IQ | - | - | - | - | - | |
| Verbal Comprehension Index | - | - | - | - | - | |
| Perceptual Reasoning Index | - | - | - | - | - | |
| Working Memory Index | - | - | - | - | - | |
| Processing Speed Index | - | - | - | - | - | |
| WMS-IV | ||||||
| Immediate Memory Index | - | - | - | - | - | |
| Delayed Memory Index | - | - | - | - | - | |
| Visual Working Memory Index | - | - | - | - | - | |
Enter age (16–90), sex, plus Vocabulary and/or Matrix Reasoning raw scores in the Inputs panel above. Rows appear automatically for each model whose required inputs are present. Enter the patient's actual FSIQ / GAI in the Achieved column - the prorated index is calculated per ACS manual procedures, excluding the subtest(s) used as predictors. The three FSIQ rows predict three different prorated criteria and are not expected to agree with each other.
| Model | Predicted | Lower 90% | Upper 90% | Achieved | Difference | Base Rate |
|---|---|---|---|---|---|---|
| Enter age plus Vocabulary and/or Matrix Reasoning to populate the table. | ||||||
Methods & References
A clinical psychometric calculation tool for neuropsychological report writing. All computation is local; no patient data is ever transmitted.
Methods & conventions
What this tool does
Seven working pages: Premorbid Estimate, Score Tables, Change Analysis, Score Charts, Score Converter, Effect Size Tools and Data. Every calculation runs locally in the browser. No patient data is transmitted off-device, and the app works with no network connection.
The auto-fill normative database holds published parameters for seven instrument families — D-KEFS (original and Advanced), WAIS-IV, WMS-IV, WISC-V, CVLT-3, CVLT-C and the RBANS — with the retest sample size N where the publisher reports one. N is required for the Crawford & Garthwaite method and may need entering by hand where it is unavailable. Clinicians should verify every imported parameter against the current manual, and against local service standards, before interpreting it.
Score conversion and classification
Conversions between standard (M 100, SD 15), T (50, 10), scaled (10, 3) and z scores assume an approximately normal reference distribution. Two descriptor schemes are offered, and the one in force is named in the note beneath every exported table: Wechsler bands follow the WAIS-IV/WMS-IV manual conventions, and AACN labels follow Guilmette et al. (2020). Confidence levels throughout are 90% (z = 1.645) and 95% (z = 1.960); intervals round the estimate and the margin separately, so the printed bounds stay symmetric about the printed value.
Confidence intervals on Score Tables
Confidence intervals and standard errors of measurement. The CI column is the obtained score ± z × SEM, where SEM = SD × √(1 − r), centred on the obtained score rather than on an estimated true score.
The standard deviation is the normative SD of the metric the score is reported in — 15, 10, 3 or 1 — because a coefficient computed on, or corrected to, the normative sample must be paired with that sample's variability. Where a measure's stored statistics are raw, its own standard deviation is used instead, that being the only one in the right units. Four publishers state that rule outright, and the arithmetic confirms it: this pairing reproduces every published standard error of measurement the app is able to check, exactly, at the precision each is printed to — all 300 cells of WAIS-IV Table 4.3, 242 of WISC-V Table 4.4, 240 of WMS-IV Table 3.3, 168 across the D-KEFS SEM tables, 126 of RBANS Update Table 3.7, and all 38 CVLT-3 measures in Tables 3.4 and 3.5.
The reliability is, by default, the retest coefficient held in the normative database — an alternate-form coefficient in the case of the CVLT-3, which publishes no same-form retest — corrected for the normative sample's variability where the publisher reports a corrected value. Retest is the default for two reasons: it keeps a single, stated basis across a table that may mix batteries, and it is the appropriate coefficient for the many timed measures in the database, since split-half and alpha are not valid reliability estimates for speeded tests. The WAIS-IV manual makes that second argument itself for Coding, Symbol Search and Cancellation, describing the split-half coefficient as "not a proper reliability estimate" for a Processing Speed subtest; the values used here for those three are the ones it publishes, in all 38 of the cells its Table 4.1 gives them.
That default is set aside for a measure only where its publisher both reports an internal-consistency coefficient and derives its own published intervals from it. Seven manuals meet that bar:
| Instrument | Coefficient used | Source | Not applied to |
|---|---|---|---|
| CVLT-C | Odd–even split-half, by age | Manual Table 6.5 | Every index but List A Trials 1–5 Total; item scores on a word-list task are not independent. The interval printed in the manual's own worked example reproduces exactly. |
| D-KEFS | Internal consistency, by normative age band | Technical Manual, Tables 2.1–2.24 | Colour–Word Interference, whose only coefficient table is for a composite this app does not hold; Design Fluency, where item interdependence precluded the procedure; and five of the six Trail Making measures, the published table covering the composite alone. |
| D-KEFS Advanced | Split-half, by normative age band | Table 3.4 | Trail Making and Verbal Fluency, which that manual treats as speeded and scores on stability coefficients. |
| WAIS-IV | Split-half or alpha, by normative age band | Table 4.1 | Coding, Symbol Search and Cancellation — speeded. These keep the corrected stability coefficient the same table publishes for them. |
| WISC-V | Split-half, by single year of age | Table 4.1 | Coding, Symbol Search and Cancellation, together with the Cancellation Random and Cancellation Structured process scores — speeded, and likewise on the corrected stability coefficient. |
| WMS-IV | Split-half or alpha, by normative age band, Adult and Older Adult batteries separately | Table 3.1 | Verbal Paired Associates II Word Recall, a free-recall score with no consistent item count, which takes a stability coefficient. The recognition memory measures are absent altogether: their published reliability is a decision-consistency percentage, not a correlation, and cannot enter a standard error of measurement. |
| RBANS Update | Internal consistency, by normative age band | Table 3.6 | Figure Copy, Semantic Fluency, Coding, Story Recall and Figure Recall, which that table itself marks as estimated from test–retest and which therefore keep a stability coefficient taken from the same table. The four subtests reported as raw scores appear nowhere in it, the manual publishing reliability for its eight scaled subtests only, so no interval is shown for them. |
Which coefficient is right is a question for each manual rather than a policy of this tool, and the manuals genuinely disagree — the two D-KEFS manuals reach opposite conclusions about the same two test names. Each is followed as written, and the Data page names the basis actually in force for every measure in the database.
Age. Where a coefficient is tabulated by age, the interval uses the band for the patient's age, and the age used is named in the note beneath the table. Entering an age is optional. If none is entered, or the age falls outside a measure's normed range, the publisher's all-ages figure is used instead: the published average where a manual prints one, and otherwise the total-sample retest coefficient — which for the D-KEFS is that manual's own second regime rather than a substitute for a missing number. Both paths are therefore the publisher's own figures.
Reliable-change analysis is unaffected by any of the above and always uses the retest coefficient, which measures a different thing.
One consequence is worth bearing in mind when comparing output against a test manual. For the measures that remain on the retest default, where a manual derives its published intervals from internal-consistency reliability — almost always the higher of the two coefficients — the intervals shown here run wider than the manual's. They are therefore the more conservative, and answer the question how much would this score be expected to move on retesting rather than how precisely was it measured on the day. Not every publisher offers that comparison: the CVLT-3 manual declines to report internal-consistency reliability at all, on the grounds that item scores on a word-list task are not independent — recalling one word alters the probability of recalling the others, both within a trial and on later ones — and reports alternate-form coefficients in their place.
Measures with no normative-sample coefficient, and the reliability control
Every interval here multiplies a normative standard deviation by a reliability, and that is only a valid standard error of measurement when the two describe the same group. Most manuals supply a coefficient computed on, or corrected to, their normative sample. Three do not — the D-KEFS, D-KEFS Advanced and CVLT-C manuals report only the correlation observed in their own retest studies, a few dozen people each, and pair it with the normative standard deviation regardless.
The D-KEFS manual states that outright, fixing the standard deviation unit at 3 for all its scaled scores and deriving its test–retest standard errors of measurement "from the total sample of cases". Its Table 2.8 shows the arithmetic: the three Design Fluency all-ages values of 1.94, 1.97 and 2.47 are exactly 3 × √(1 − r) on the uncorrected coefficient. Those measures are therefore scored the way their own manuals score them, and the intervals shown reproduce the published ones. The Data page labels each such measure retest, uncorrected, so which rows rest on that footing can be read off rather than inferred.
The statistical objection to the pairing is nonetheless real, so Score Tables offers a reliability control with two settings. Published, the default, uses each manual's own coefficient and reproduces its printed interval. Corrected applies the standard range-restriction correction of Allen and Yen, rxx = 1 − (s²retest ÷ s²norm)(1 − r), to those measures alone, so that the coefficient describes the same population as the standard deviation it multiplies. A published coefficient is never overwritten in either setting.
The correction is not a guess: across the 267 database entries carrying both an observed and a publisher-corrected coefficient, it reproduces the publisher's own value to a median error of .003. But the resulting figures are not printed in the manuals concerned, which is why the default is Published and why the note beneath a corrected table says so. In practice the control moves 46 of the measures reachable from Score Tables, every one of them D-KEFS or D-KEFS Advanced: at the 95% level 9 intervals widen and 4 narrow, 33 are unchanged after rounding, and the largest single change is 2 scaled-score points. Reliable-change analysis is not affected by the control. Where a corrected reading is taken, the Data page follows it rather than continuing to show the published one.
Change analysis
Five methods, in ascending order of what they model. The Standard Deviation Index is descriptive only: SD Δ = (X₂ − X₁) ÷ SD, with no reliability correction. Simple Reliable Change (Jacobson & Truax, 1991) tests the observed change against measurement error. Practice Effect-Adjusted change (Iverson, 2001) subtracts the mean retest gain observed in the normative sample first. McSweeney Regression-Based change (McSweeney et al., 1993) predicts the retest score from baseline and standardises the residual against the standard error of estimate. Crawford & Garthwaite Regression-Based change (2007) does the same but with a standard error of prediction that accounts for the normative sample size and for the distance of the baseline score from the normative mean.
All p-values are two-tailed. The first four methods use the standard normal distribution; Crawford & Garthwaite uses the Student t distribution with N − 2 degrees of freedom, so a small normative sample raises the threshold — at N = 25 the 95% critical value is 2.069 against 1.960 for z.
These calculations use the retest coefficient paired with the standard deviation of the same retest sample, so that both terms describe one population. Where a publisher reports a coefficient corrected to the normative sample's variability, that value is offered as an option but is not the default, because it describes a differently distributed population from the standard deviation it would be multiplied by — and in the two regression methods the coefficient is a fitted slope, so substituting it changes the predicted score rather than only the interval. Reliability type varies by instrument and is stated in the note beneath each generated table: CVLT-3 coefficients are alternate-form, RBANS Form A coefficients are same-form retest.
Outcomes are reported as significance only — "Reliable change" or "No reliable change" — never as improvement or decline. The database holds many measures on which a higher score is the worse result (intrusions, perseverations, errors, false positives), and carries no score-direction flag, so reading a clinical direction off the sign of the statistic would assert the wrong conclusion for all of them. The signed statistic is displayed alongside, so the direction stays visible without the app interpreting it.
Premorbid estimation
Premorbid estimates combine ToPF-based and demographic equations with OPIE-4 prorated models, and produce predicted-versus-achieved discrepancy output with confidence intervals, base-rate lookups and APA-formatted export tables. Predicted-versus-achieved significance flagging uses a three-tier scheme at z = 1.645 (*), 1.960 (**) and 2.576 (***); this is separate from the 90%/95% confidence-interval selector.
OPIE-4 is provided for illustration only in a UK context and its output should not be quoted as a concrete premorbid estimate. The regression terms reproduce Holdnack et al. (2013), Table eA5.8, but the published equations also carry US education, ethnicity and region terms that are not applied here, which fixes every prediction at the US reference category (12th-grade high-school graduate, not African-American, not resident in the western US). Those terms are omitted rather than mapped because the education dummies encode how unusual a given attainment level is within the US population the model was fitted on, not years of schooling, and that does not transfer: the US reference category corresponds to A-levels if matched by years but to GCSE/O-level if matched by population position, and UK school-leaving age was raised to 16 only in 1972, so leaving school without qualifications was normative for older cohorts in a way it was not in the US sample. Expect estimates to run high for patients who left school early and low for graduates, by an amount this tool cannot quantify. For a UK demographic estimate, use the Crawford & Allan (2001) model.
Base rates. The ToPF predicted-difference base rates are estimated from a normal model with SD equal to the model's standard error of estimate, not transcribed from observed standardisation-sample frequencies; they are labelled as such wherever they appear. The published ToPF/ACS predicted-difference tables have not been transcribed, so the model has instead been benchmarked against the one genuinely empirical table available for a model of almost identical standard error of estimate — the OPIE-4 discrepancy base rates of ACS Table eA5.12. Against that table the parametric values run roughly 10% relatively low across the decisive −5 to −20 band: a discrepancy of −15 gives 3.78% here against about 4.3% empirically. The net effect is to show a given discrepancy as marginally rarer, and so marginally more pathological, than observed data suggest. The OPIE-4 discrepancy base rates are themselves empirical and are used as published.
Delis, D. C., Kaplan, E., & Kramer, J. H. (2001). Delis–Kaplan Executive Function System (D-KEFS): Technical manual. San Antonio, TX: The Psychological Corporation. [Internal-consistency coefficients and standard errors of measurement by age band, chapter 2 and Tables 2.1–2.26.]
Delis, D. C., Kramer, J. H., Kaplan, E., & Ober, B. A. (1994). California Verbal Learning Test – Children's Version (CVLT-C): Manual. San Antonio, TX: The Psychological Corporation. [Split-half reliability and standard errors of measurement: Table 6.5. Standardised score equivalents: Tables A.1 and A.2.]
Delis, D. C., Kramer, J. H., Kaplan, E., & Ober, B. A. (2017). California Verbal Learning Test – Third Edition (CVLT-3): Manual. Bloomington, MN: Pearson. [Alternate-form reliability and standard errors of measurement: Tables 3.4 and 3.5.]
Holdnack, J. A., Drozdick, L., Weiss, L. G., & Iverson, G. L. (2013). WAIS-IV, WMS-IV, and ACS: Advanced clinical interpretation. Oxford: Academic Press. [OPIE-4 prorated regression coefficients: Table eA5.8. OPIE-4 discrepancy base rates: Table eA5.12.]
Randolph, C. (2012). Repeatable Battery for the Assessment of Neuropsychological Status Update (RBANS Update): Manual. Bloomington, MN: Pearson. [Reliability by age band: Table 3.6. Standard errors of measurement: Table 3.7. Test–retest and alternate-form data: Tables 3.8–3.9.]
Wechsler, D. (2010). Wechsler Adult Intelligence Scale – Fourth UK Edition (WAIS–IVUK): Administration and scoring manual. London: Pearson Assessment. [Longest-span base rates: Tables C.4 and C.5.]
Wechsler, D. (2010). Wechsler Adult Intelligence Scale – Fourth UK Edition (WAIS–IVUK): Technical and interpretive manual. London: Pearson Assessment. [Reliability coefficients: Table 4.1. Standard errors of measurement: Table 4.3. Test–retest stability parameters: Table 4.5.]
Wechsler, D. (2010). Wechsler Memory Scale – Fourth UK Edition (WMS–IVUK): Technical and interpretive manual. London: Pearson Assessment. [Reliability coefficients: Table 3.1. Standard errors of measurement: Table 3.3.]
Wechsler, D. (2011). Test of Premorbid Functioning (ToPF-UK): Manual. London: Pearson Assessment.
Wechsler, D. (2014). Wechsler Intelligence Scale for Children – Fifth Edition (WISC-V): Technical and interpretive manual. Bloomington, MN: NCS Pearson. [Reliability coefficients: Table 4.1. Standard errors of measurement: Table 4.4. Test–retest stability: Table 4.7.]
Allen, M. J., & Yen, W. M. (1979). Introduction to measurement theory. Monterey, CA: Brooks/Cole. [Correction of a reliability coefficient for range restriction in the sample it was observed in.]
Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Hillsdale, NJ: Lawrence Erlbaum.
Crawford, J. R., & Allan, K. M. (2001). Estimating premorbid WAIS–R IQ with demographic variables: Regression equations derived from a UK sample. The Clinical Neuropsychologist, 11(2), 192–197.
Crawford, J. R., & Garthwaite, P. H. (2007). Using regression equations built from summary data in the neuropsychological assessment of the individual case. Neuropsychology, 21(5), 611–620.
Crawford, J. R., Millar, J., & Milne, A. B. (2001). Estimating premorbid IQ from demographic variables: A comparison of a regression equation vs. clinical judgement. British Journal of Clinical Psychology, 40(1), 97–105.
Guilmette, T. J., Sweet, J. J., Hebben, N., Koltai, D., Mahone, E. M., Spiegler, B. J., Stucky, K., Westerveld, M., & Conference Participants. (2020). American Academy of Clinical Neuropsychology consensus conference statement on uniform labeling of performance test scores. The Clinical Neuropsychologist, 34(3), 437–453.
Iverson, G. L. (2001). Interpreting change on the WAIS-III/WMS-III in clinical samples. Archives of Clinical Neuropsychology, 16(2), 183–191.
Jacobson, N. S., & Truax, P. (1991). Clinical significance: A statistical approach to defining meaningful change in psychotherapy research. Journal of Consulting and Clinical Psychology, 59(1), 12–19.
McSweeney, A. J., Naugle, R. I., Chelune, G. J., & Lüders, H. (1993). "T scores for change": An illustration of a regression approach to depicting change in clinical neuropsychology. The Clinical Neuropsychologist, 7(3), 300–312.
Sawilowsky, S. S. (2009). New effect size rules of thumb. Journal of Modern Applied Statistical Methods, 8(2), 597–599.
Cipriani, A., Furukawa, T. A., Salanti, G., Chaimani, A., Atkinson, L. Z., Ogawa, Y., Leucht, S., Ruhe, H. G., Turner, E. H., Higgins, J. P. T., Egger, M., Takeshima, N., Hayasaka, Y., Imai, H., Shinohara, K., Tajika, A., Ioannidis, J. P. A., & Geddes, J. R. (2018). Comparative efficacy and acceptability of 21 antidepressant drugs for the acute treatment of adults with major depressive disorder: A systematic review and network meta-analysis. The Lancet, 391(10128), 1357–1366.
Cuijpers, P., Geraedts, A. S., van Oppen, P., Andersson, G., Markowitz, J. C., & van Straten, A. (2011). Interpersonal psychotherapy for depression: A meta-analysis. American Journal of Psychiatry, 168(6), 581–592.
Cuijpers, P., Miguel, C., Harrer, M., Plessen, C. Y., Ciharova, M., Ebert, D., & Karyotaki, E. (2023). Cognitive behavior therapy vs. control conditions, other psychotherapies, pharmacotherapies and combined treatment for depression: A comprehensive meta-analysis including 409 trials with 52,702 patients. World Psychiatry, 22(1), 105–115.
Hofmann, S. G., & Smits, J. A. J. (2008). Cognitive-behavioral therapy for adult anxiety disorders: A meta-analysis of randomized placebo-controlled trials. The Journal of Clinical Psychiatry, 69(4), 621–632.
Huhn, M., Nikolakopoulou, A., Schneider-Thoma, J., Krause, M., Samara, M., Peter, N., Arndt, T., Bäckers, L., Rothe, P., Cipriani, A., Davis, J., Salanti, G., & Leucht, S. (2019). Comparative efficacy and tolerability of 32 oral antipsychotics for the acute treatment of adults with multi-episode schizophrenia: A systematic review and network meta-analysis. The Lancet, 394(10202), 939–951.
Pesch, B., Kendzia, B., Gustavsson, P., Jöckel, K.-H., Johnen, G., Pohlabeln, H., Olsson, A., Ahrens, W., Gross, I. M., Brüske, I., Wichmann, H.-E., Merletti, F., Richiardi, L., Simonato, L., Fortes, C., Siemiatycki, J., Parent, M.-E., Consonni, D., Landi, M. T., … Brüning, T. (2012). Cigarette smoking and lung cancer — Relative risk estimates for the major histological types from a pooled analysis of case–control studies. International Journal of Cancer, 131(5), 1210–1219.
Storebø, O. J., Storm, M. R. O., Pereira Ribeiro, J., Skoog, M., Groth, C., Callesen, H. E., Schaug, J. P., Darling Rasmussen, P., Huus, C.-M. L., Zwi, M., Kirubakaran, R., Simonsen, E., & Gluud, C. (2023). Methylphenidate for children and adolescents with attention deficit hyperactivity disorder (ADHD). Cochrane Database of Systematic Reviews, 3, CD009885.
Watts, B. V., Schnurr, P. P., Mayo, L., Young-Xu, Y., Weeks, W. B., & Friedman, M. J. (2013). Meta-analysis of the efficacy of treatments for posttraumatic stress disorder. The Journal of Clinical Psychiatry, 74(6), e541–e550.
Williams, A. C. de C., Fisher, E., Hearn, L., & Eccleston, C. (2020). Psychological therapies for the management of chronic pain (excluding headache) in adults. Cochrane Database of Systematic Reviews, 8, CD007407.
Wolraich, M. L., Wilson, D. B., & White, J. W. (1995). The effect of sugar on behavior or cognition in children: A meta-analysis. JAMA, 274(20), 1617–1621.
The Effect Size Tools page also cites a UK Biobank male-versus-female adult height contrast (d = 2.04) attributed to Lui et al. (2021). That source could not be verified and is not listed here; treat the value as an illustrative anchor only.