Skip to main content

Psychometric Assistant

Neuropsychological calculators for clinical practice, with APA-formatted output

Psychometrics · Standardised Score Conversion

Score Converter


−3 SD −2 SD −1 SD M +1 SD +2 SD +3 SD
Score equivalents
-3σ-2σ-1σ+1σ+2σ+3σ557085100115130145SSσz = 0.00 P(Z ≤ z) 50.0%
Wechsler
AACN

AACN = American Academy of Clinical Neuropsychology · Ranges shown as Standard Score (SS)

Conversion · Clinical Outcomes Table

Clinical Outcomes Table

Configuration
Premorbid Comparison

SD mode: * ≥1 SD below, ** ≥1.5 SD, *** ≥2 SD. SEE mode: * below 90% CI, ** below 95% CI, *** below 99% CI lower bound.

CVLT-3 Indices · Ages 16-44
CVLT-3 Indices · Ages 45-90
CVLT-3 Trials · Ages 16-44
CVLT-3 Trials · Ages 45-90
D-KEFS Colour-Word Interference · Ages 20-49
D-KEFS Colour-Word Interference · Ages 50-89
D-KEFS Colour-Word Interference · Ages 8-19
D-KEFS Colour-Word Interference · All Ages
D-KEFS Design Fluency · Ages 20-49
D-KEFS Design Fluency · Ages 50-89
D-KEFS Design Fluency · Ages 8-19
D-KEFS Design Fluency · All Ages
D-KEFS Sorting Test · Ages 20-49
D-KEFS Sorting Test · Ages 50-89
D-KEFS Sorting Test · Ages 8-19
D-KEFS Sorting Test · All Ages
D-KEFS Tower Test · Ages 20-49
D-KEFS Tower Test · Ages 50-89
D-KEFS Tower Test · Ages 8-19
D-KEFS Tower Test · All Ages
D-KEFS Trail Making Test · Ages 20-49
D-KEFS Trail Making Test · Ages 50-89
D-KEFS Trail Making Test · Ages 8-19
D-KEFS Trail Making Test · All Ages
D-KEFS Verbal Fluency · Ages 20-49
D-KEFS Verbal Fluency · Ages 50-89
D-KEFS Verbal Fluency · Ages 8-19
D-KEFS Verbal Fluency · All Ages
D-KEFS Word Context Test · Ages 20-49
D-KEFS Word Context Test · Ages 50-89
D-KEFS Word Context Test · Ages 8-19
D-KEFS Word Context Test · All Ages
D-KEFS Word Proverb Test · Ages 20-49
D-KEFS Word Proverb Test · Ages 50-89
D-KEFS Word Proverb Test · Ages 8-19
D-KEFS Word Proverb Test · All Ages
RBANS Indices · Ages 12-19
RBANS Indices · Ages 20-89
RBANS Subtests · Ages 12-19
RBANS Subtests · Ages 20-89
WAIS-IV Core Subtests · Ages 16-29
WAIS-IV Core Subtests · Ages 30-54
WAIS-IV Core Subtests · Ages 55-69
WAIS-IV Core Subtests · Ages 70-90
WAIS-IV Core Subtests · All Ages
WAIS-IV Indices · Ages 16-29
WAIS-IV Indices · Ages 30-54
WAIS-IV Indices · Ages 55-69
WAIS-IV Indices · Ages 70-90
WAIS-IV Indices · All Ages
WISC-V Indices · All Ages
WISC-V Subtests · All Ages
WMS-IV Indices · Ages 16-69
WMS-IV Indices · Ages 65-90
WMS-IV Subtests · Ages 16-69
WMS-IV Subtests · Ages 65-90
# Subtest Raw Score CI Percentile Classification
APA-formatted output
Enter at least one subtest with a score to preview the APA table.
Basic Report Tools · Effect Size

Effect Size Tools

Convert between effect-size metrics or derive them from group data, with a visual against the standard normal.

OR
Preloaded example UK male vs female height is loaded in the group-data fields from Option B (above). Edit the values or use Clear all inputs to start fresh.
0.00
Intervention (Group 1)
95% CI: -
Control (Group 2)
95% CI: -
Pooled N-
Pooled Mean-
Pooled SD-
Pooled SE-
Mean Diff-
Effect-size results
Cohen's d
-
-
Hedges' g
-
-
Pearson's r
-
-
-
-
Cohen's f
-
-
Fisher's z
-
Odds Ratio
-
Overlap
-
Cohen's U₃
-
CLES
-
Kraemer's NNT
-
Similar to
-
Group comparison at a target value
Quick compare. Enter a target value to see how each group performs at or above that point.
-
-
-
-
-
Requires means and standard deviations. If standard errors are entered above, the app converts them to SDs using each group's sample size.
Reference values (Cohen's d)

Use these published effect sizes to anchor your results. The values below span small to huge magnitudes so you can compare your finding against familiar clinical and epidemiological benchmarks.

Heavy smokers (30+/day) vs never smokers, lung cancer (Pesch et al., 2012)2.60
UK male vs female adult height (UK Biobank; Lui et al., 2021)2.04
Smokers (any) vs never smokers, lung cancer (Pesch et al., 2012)1.75
Cognitive therapy vs control for PTSD (Watts et al., 2013)1.63
Former smokers vs never smokers, lung cancer (Pesch et al., 2012)1.10
Exposure therapy vs control for PTSD (Watts et al., 2013)1.08
EMDR vs control for PTSD (Watts et al., 2013)1.00
Clozapine vs placebo for schizophrenia (Huhn et al., 2019 Lancet)0.89
CBT vs control for depression (Cuijpers et al., 2023 World Psychiatry)0.80
Methylphenidate vs placebo for ADHD, children (Storebø et al., 2023 Cochrane)0.75
CBT vs placebo for anxiety disorders (Hofmann & Smits, 2008)0.70
CBT for depression, low-risk-of-bias subset (Cuijpers et al., 2023)0.60
Interpersonal Therapy for depression (Cuijpers et al., 2011)0.50
Antidepressants vs placebo (Cipriani et al., 2018 Lancet)0.30
CBT vs treatment-as-usual for chronic pain (Williams et al., 2020 Cochrane)0.20
CBT vs active control for chronic pain (Williams et al., 2020 Cochrane)0.10
Sugar on children's hyperactivity (Wolraich et al., 1995 JAMA)0.00
No or negligible effect0.00
Common language description
Interpretation summary.
Enter a valid effect size to generate a plain-English interpretation.
The average person in Group 1 is above about - of Group 2 (Cohen's U₃).
Visual
Group 2 (μ = 0) Group 1 (μ = d)
Curves shifted by current Cohen's d.
How these charts are drawn

The page covers four sources — Score Tables, the premorbid predictions, Change Analysis and the SD Index — each as a pane of its own, so the page does not grow into a long scroll. Every trial and subtest becomes a row of its test's chart, showing the score against classification bands, or the change against the reliable-change interval. The charts read the same data and settings as those tables, so the two cannot disagree.

Each test from Score Tables gets one chart, and each of its trials or subtests is a row of that chart, drawn in the test's native metric (standard, T, scaled or z) — no score is converted for display, and a test that mixes metrics gets one panel per metric, never two metrics on one axis. The shaded bands behind each row are the classification ranges of the scheme selected on Score Tables (Wechsler or Guilmette et al., 2020), with band boundaries at standard scores 70, 80, 90, 110, 120 and 130 converted onto the panel's scale. The dot marks the obtained score; the whisker is the same confidence interval printed in the table's CI column (SEM = SD·√(1−r), at the CI level selected on Score Tables, with the table's rounding).

Base-rate measures (WAIS-IV Longest Span) are charted from the published cumulative table (WAIS-IV Administration and Scoring Manual, Tables C.4–C.5): each row's step line is the percentage of the normative sample obtaining each span or higher, and the dot marks the patient's span on that line. A high base rate means a common, and therefore lower, score.

Error measures (perseverations, intrusions, false positives; marked ↓) follow the Score Tables convention: the percentile is reported as obtained, while the classification — and therefore the band shading — describes performance, so those rows' bands run reversed.

Raw-score measures are listed, not plotted: a raw score has no position on a standardised axis, so no percentile or classification is derived. The obtained score and its raw-unit confidence interval are still shown.

Change Analysis charts plot each trial's two testings in score units — open circle at the first, filled dot at the second — against the shaded reliable-change interval for the selected method: the region where |obtained − expected| falls below the method's critical value × standard error, i.e. exactly where that method's table prints "No reliable change". The expected value is the first score (Jacobson & Truax), the first score plus the normative practice effect (Iverson), or the regression-predicted score (McSweeney, Crawford & Garthwaite), and the interval uses the confidence level set on that method's page. The method buttons only choose which method to draw; they change nothing on the Change Analysis pages. As in the tables, outcomes state significance only and never a direction — the signed statistic is shown alongside so the direction of movement stays visible without the app interpreting it.

SD Index charts plot each trial's change in standard-deviation units against the ±1.96 (or ±1.645) band for the significance level set on that page, using the same per-row divisor the SD Index table applies.

Premorbid charts answer the question the ToPF and OPIE-4 tabs ask: each index shows the predicted score (open circle) with its prediction interval shaded, the achieved score (filled dot) where one has been entered, and the difference with its base rate. A thin line marks the population mean of 100. The predicted values, intervals, differences and base rates are read from the cells those tables print, not recomputed — both tabs round the estimate and the margin separately so the bounds stay symmetric, and re-deriving them here would create a third place to keep in step. The Estimates tab keeps its own model-comparison forest plot; this block does not duplicate it.

Getting around. Each source — Score Tables, Premorbid, Change Analysis, SD Index — is its own pane, with a Back/Next bar beneath, so the page does not grow into one long scroll as more tests are entered. Only sources that hold data appear. Within a pane, All charts shows the whole set together, which is how the profile across a battery is read; One at a time gives a single chart the full width for a close look or a clean export, with the left and right arrow keys paging through it.

Showing the scores a different way. The Score Tables block offers four axes. Native metric (the default) converts nothing. Percentile and Standard score put every measure of a test on one axis, which is what makes a whole battery comparable — the score column still shows the value as entered, only the axis position is converted, and where a test mixes metrics each row carries its own metric tag. A percentile axis is deliberately non-linear: it compresses the tails, so two clearly different low scores can sit close together, and symmetric confidence intervals become asymmetric. Raw scores simply shows the raw scores as entered, so their spread is visible directly; nothing is derived from them — no norms, no classification bands, no percentile. Its axis spans the values entered for that test rather than each measure's possible range, because the app holds no raw maximum for any measure, so where a test's measures are counted on different scales their positions are not comparable with each other. A value that falls outside any chart's axis is drawn at the edge, dimmed and marked with a caret, with its actual figure alongside.

Premorbid-comparison asterisks, when enabled on Score Tables, carry the same meaning here as in the table's classification column.

Discrepancy · Standard Deviation Index

Standard Deviation Index

Quantify abnormality of test-retest discrepancy in standard-deviation units. Useful when reliability data are unavailable or for descriptive comparison.

APA-formatted output
Enter at least one subtest to preview.
Reliable Change Indices (RCI) · Simple

Basic Reliable Change Index

Jacobson & Truax (1991). Computes whether observed change exceeds measurement error, using the test's reliability coefficient and standard deviation.

CVLT-3 Indices · Ages 16-44
CVLT-3 Indices · Ages 45-90
CVLT-3 Trials · Ages 16-44
CVLT-3 Trials · Ages 45-90
D-KEFS Colour-Word Interference · Ages 20-49
D-KEFS Colour-Word Interference · Ages 50-89
D-KEFS Colour-Word Interference · Ages 8-19
D-KEFS Colour-Word Interference · All Ages
D-KEFS Design Fluency · Ages 20-49
D-KEFS Design Fluency · Ages 50-89
D-KEFS Design Fluency · Ages 8-19
D-KEFS Design Fluency · All Ages
D-KEFS Sorting Test · Ages 20-49
D-KEFS Sorting Test · Ages 50-89
D-KEFS Sorting Test · Ages 8-19
D-KEFS Sorting Test · All Ages
D-KEFS Tower Test · Ages 20-49
D-KEFS Tower Test · Ages 50-89
D-KEFS Tower Test · Ages 8-19
D-KEFS Tower Test · All Ages
D-KEFS Trail Making Test · Ages 20-49
D-KEFS Trail Making Test · Ages 50-89
D-KEFS Trail Making Test · Ages 8-19
D-KEFS Trail Making Test · All Ages
D-KEFS Verbal Fluency · Ages 20-49
D-KEFS Verbal Fluency · Ages 50-89
D-KEFS Verbal Fluency · Ages 8-19
D-KEFS Verbal Fluency · All Ages
D-KEFS Word Context Test · Ages 20-49
D-KEFS Word Context Test · Ages 50-89
D-KEFS Word Context Test · Ages 8-19
D-KEFS Word Context Test · All Ages
D-KEFS Word Proverb Test · Ages 20-49
D-KEFS Word Proverb Test · Ages 50-89
D-KEFS Word Proverb Test · Ages 8-19
D-KEFS Word Proverb Test · All Ages
RBANS Indices · Ages 12-19
RBANS Indices · Ages 20-89
RBANS Subtests · Ages 12-19
RBANS Subtests · Ages 20-89
WAIS-IV Core Subtests · Ages 16-29
WAIS-IV Core Subtests · Ages 30-54
WAIS-IV Core Subtests · Ages 55-69
WAIS-IV Core Subtests · Ages 70-90
WAIS-IV Core Subtests · All Ages
WAIS-IV Indices · Ages 16-29
WAIS-IV Indices · Ages 30-54
WAIS-IV Indices · Ages 55-69
WAIS-IV Indices · Ages 70-90
WAIS-IV Indices · All Ages
WISC-V Indices · All Ages
WISC-V Subtests · All Ages
WMS-IV Indices · Ages 16-69
WMS-IV Indices · Ages 65-90
WMS-IV Subtests · Ages 16-69
WMS-IV Subtests · Ages 65-90

Test data & patient scores

# Subtest SD r Date 1 Date 2 RCI (z) p Outcome
APA-formatted output
Enter test data and patient scores to preview the APA table.
Reliable Change Indices (RCI) · Practice Effects

Practice Effect-Adjusted Reliable Change Index

Iverson (2001). Adjusts the standard RCI to control for the average improvement (practice effect) observed between assessments in the normative sample.

CVLT-3 Indices · Ages 16-44
CVLT-3 Indices · Ages 45-90
CVLT-3 Trials · Ages 16-44
CVLT-3 Trials · Ages 45-90
D-KEFS Colour-Word Interference · Ages 20-49
D-KEFS Colour-Word Interference · Ages 50-89
D-KEFS Colour-Word Interference · Ages 8-19
D-KEFS Colour-Word Interference · All Ages
D-KEFS Design Fluency · Ages 20-49
D-KEFS Design Fluency · Ages 50-89
D-KEFS Design Fluency · Ages 8-19
D-KEFS Design Fluency · All Ages
D-KEFS Sorting Test · Ages 20-49
D-KEFS Sorting Test · Ages 50-89
D-KEFS Sorting Test · Ages 8-19
D-KEFS Sorting Test · All Ages
D-KEFS Tower Test · Ages 20-49
D-KEFS Tower Test · Ages 50-89
D-KEFS Tower Test · Ages 8-19
D-KEFS Tower Test · All Ages
D-KEFS Trail Making Test · Ages 20-49
D-KEFS Trail Making Test · Ages 50-89
D-KEFS Trail Making Test · Ages 8-19
D-KEFS Trail Making Test · All Ages
D-KEFS Verbal Fluency · Ages 20-49
D-KEFS Verbal Fluency · Ages 50-89
D-KEFS Verbal Fluency · Ages 8-19
D-KEFS Verbal Fluency · All Ages
D-KEFS Word Context Test · Ages 20-49
D-KEFS Word Context Test · Ages 50-89
D-KEFS Word Context Test · Ages 8-19
D-KEFS Word Context Test · All Ages
D-KEFS Word Proverb Test · Ages 20-49
D-KEFS Word Proverb Test · Ages 50-89
D-KEFS Word Proverb Test · Ages 8-19
D-KEFS Word Proverb Test · All Ages
RBANS Indices · Ages 12-19
RBANS Indices · Ages 20-89
RBANS Subtests · Ages 12-19
RBANS Subtests · Ages 20-89
WAIS-IV Core Subtests · Ages 16-29
WAIS-IV Core Subtests · Ages 30-54
WAIS-IV Core Subtests · Ages 55-69
WAIS-IV Core Subtests · Ages 70-90
WAIS-IV Core Subtests · All Ages
WAIS-IV Indices · Ages 16-29
WAIS-IV Indices · Ages 30-54
WAIS-IV Indices · Ages 55-69
WAIS-IV Indices · Ages 70-90
WAIS-IV Indices · All Ages
WISC-V Indices · All Ages
WISC-V Subtests · All Ages
WMS-IV Indices · Ages 16-69
WMS-IV Indices · Ages 65-90
WMS-IV Subtests · Ages 16-69
WMS-IV Subtests · Ages 65-90

Test data & patient scores

# Subtest M₁ SD₁ M₂ SD₂ r Date 1 Date 2 RCI (z) p Outcome
APA-formatted output
Enter test data and patient scores to preview the APA table.
Reliable Change Indices (RCI) · Regression-Based

McSweeney Regression-Based (SRB) Reliable Change Index

McSweeney et al. (1993). Predicts each patient's expected retest score from their baseline and the normative sample's regression parameters; the residual is standardised against the standard error of estimate.

CVLT-3 Indices · Ages 16-44
CVLT-3 Indices · Ages 45-90
CVLT-3 Trials · Ages 16-44
CVLT-3 Trials · Ages 45-90
D-KEFS Colour-Word Interference · Ages 20-49
D-KEFS Colour-Word Interference · Ages 50-89
D-KEFS Colour-Word Interference · Ages 8-19
D-KEFS Colour-Word Interference · All Ages
D-KEFS Design Fluency · Ages 20-49
D-KEFS Design Fluency · Ages 50-89
D-KEFS Design Fluency · Ages 8-19
D-KEFS Design Fluency · All Ages
D-KEFS Sorting Test · Ages 20-49
D-KEFS Sorting Test · Ages 50-89
D-KEFS Sorting Test · Ages 8-19
D-KEFS Sorting Test · All Ages
D-KEFS Tower Test · Ages 20-49
D-KEFS Tower Test · Ages 50-89
D-KEFS Tower Test · Ages 8-19
D-KEFS Tower Test · All Ages
D-KEFS Trail Making Test · Ages 20-49
D-KEFS Trail Making Test · Ages 50-89
D-KEFS Trail Making Test · Ages 8-19
D-KEFS Trail Making Test · All Ages
D-KEFS Verbal Fluency · Ages 20-49
D-KEFS Verbal Fluency · Ages 50-89
D-KEFS Verbal Fluency · Ages 8-19
D-KEFS Verbal Fluency · All Ages
D-KEFS Word Context Test · Ages 20-49
D-KEFS Word Context Test · Ages 50-89
D-KEFS Word Context Test · Ages 8-19
D-KEFS Word Context Test · All Ages
D-KEFS Word Proverb Test · Ages 20-49
D-KEFS Word Proverb Test · Ages 50-89
D-KEFS Word Proverb Test · Ages 8-19
D-KEFS Word Proverb Test · All Ages
RBANS Indices · Ages 12-19
RBANS Indices · Ages 20-89
RBANS Subtests · Ages 12-19
RBANS Subtests · Ages 20-89
WAIS-IV Core Subtests · Ages 16-29
WAIS-IV Core Subtests · Ages 30-54
WAIS-IV Core Subtests · Ages 55-69
WAIS-IV Core Subtests · Ages 70-90
WAIS-IV Core Subtests · All Ages
WAIS-IV Indices · Ages 16-29
WAIS-IV Indices · Ages 30-54
WAIS-IV Indices · Ages 55-69
WAIS-IV Indices · Ages 70-90
WAIS-IV Indices · All Ages
WISC-V Indices · All Ages
WISC-V Subtests · All Ages
WMS-IV Indices · Ages 16-69
WMS-IV Indices · Ages 65-90
WMS-IV Subtests · Ages 16-69
WMS-IV Subtests · Ages 65-90

Test data & patient scores

# Subtest M₁ SD₁ M₂ SD₂ r Date 1 Date 2 Ŷ₂ RCI (z) p Outcome
APA-formatted output
Enter test data and patient scores to preview the APA table.
Reliable Change Indices (RCI) · Regression-Based (Crawford)

Crawford Regression-Based Reliable Change Index

Crawford & Garthwaite (2007). Extends the standardised regression-based approach to use a t-distributed test statistic that incorporates the normative sample size (N), correctly accounting for uncertainty in the regression parameters when N is modest. Returns a sample-size-adjusted standard error of prediction.

CVLT-3 Indices · Ages 16-44
CVLT-3 Indices · Ages 45-90
CVLT-3 Trials · Ages 16-44
CVLT-3 Trials · Ages 45-90
D-KEFS Colour-Word Interference · All Ages
D-KEFS Design Fluency · All Ages
D-KEFS Sorting Test · All Ages
D-KEFS Tower Test · All Ages
D-KEFS Trail Making Test · All Ages
D-KEFS Verbal Fluency · All Ages
D-KEFS Word Context Test · All Ages
D-KEFS Word Proverb Test · All Ages
RBANS Indices · Ages 12-19
RBANS Indices · Ages 20-89
RBANS Subtests · Ages 12-19
RBANS Subtests · Ages 20-89
WAIS-IV Core Subtests · All Ages
WAIS-IV Indices · All Ages
WISC-V Indices · All Ages
WISC-V Subtests · All Ages
WMS-IV Indices · Ages 16-69
WMS-IV Indices · Ages 65-90
WMS-IV Subtests · Ages 16-69
WMS-IV Subtests · Ages 65-90

Test data & patient scores

# Subtest M₁ SD₁ M₂ SD₂ r N Date 1 Date 2 Ŷ₂ t(RB) p Outcome
APA-formatted output
Enter test data and patient scores to preview the APA table.
Premorbid · Estimation

Premorbid Estimate

Inputs

Enter whichever predictors are available. Leave unavailable fields blank; the estimate table will update only for models with enough information.

Available predictors
Demographics
Used by demographic and age-adjusted models where applicable
Shared with the patient age in the header
Output settings
Controls the confidence interval and report table title
Note. ToPF and Crawford & Allan equations use UK data. OPIE-4 uses the prorated WAIS-IV US coefficients (Holdnack et al., 2013), adapted for UK use by omitting the education, region and ethnicity terms; prorated equations avoid part-whole correlation inflation. Hover the ? beside a model for its required inputs.
APA-formatted output
Enter at least the ToPF raw score to generate estimates.

Enter the patient's actual WAIS-IV / WMS-IV index scores in the Achieved column to compute ToPF-predicted vs actual discrepancies. Base rates are shown only for negative discrepancies (achieved < predicted), and are estimated from a normal model with SD = SEE rather than transcribed from observed standardisation-sample frequencies.

Index Predicted Lower 90% Upper 90% Achieved Difference Base rate
WAIS-IV
Full Scale IQ - - - - -
Verbal Comprehension Index - - - - -
Perceptual Reasoning Index - - - - -
Working Memory Index - - - - -
Processing Speed Index - - - - -
WMS-IV
Immediate Memory Index - - - - -
Delayed Memory Index - - - - -
Visual Working Memory Index - - - - -
APA-formatted output
Enter at least one Achieved score to generate the discrepancy table.
Illustrative only: do not quote these numbers in a UK report. The coefficients here reproduce Holdnack et al. (2013) Table eA5.8 exactly, but the published equations also include terms for US education, ethnicity and region, and this tool does not apply them. Every patient is therefore scored as though they were the US reference case: a 12th-grade high-school graduate, not African-American, and not living in the US West. Those categories have no UK equivalent. The US education terms do not count years of schooling; they capture how unusual a given level of attainment is within the US population, so there is nothing here to map them onto. Expect these estimates to run high for patients who left school early and low for graduates, by an amount this tool cannot quantify. If you want a UK demographic estimate, use the Crawford & Allan (2001) row on the Estimates tab.

Enter age (16–90), sex, plus Vocabulary and/or Matrix Reasoning raw scores in the Inputs panel above. Rows appear automatically for each model whose required inputs are present. Enter the patient's actual FSIQ / GAI in the Achieved column - the prorated index is calculated per ACS manual procedures, excluding the subtest(s) used as predictors. The three FSIQ rows predict three different prorated criteria and are not expected to agree with each other.

Model Predicted Lower 90% Upper 90% Achieved Difference Base Rate
Enter age plus Vocabulary and/or Matrix Reasoning to populate the table.
APA-formatted output
Enter age plus a subtest to generate predictions, then add an Achieved score.
Settings · Methods & References

Methods & References

A clinical psychometric calculation tool for neuropsychological report writing. All computation is local; no patient data is ever transmitted.

Methods & conventions

What this tool does

Seven working pages: Premorbid Estimate, Score Tables, Change Analysis, Score Charts, Score Converter, Effect Size Tools and Data. Every calculation runs locally in the browser. No patient data is transmitted off-device, and the app works with no network connection.

The auto-fill normative database holds published parameters for seven instrument families — D-KEFS (original and Advanced), WAIS-IV, WMS-IV, WISC-V, CVLT-3, CVLT-C and the RBANS — with the retest sample size N where the publisher reports one. N is required for the Crawford & Garthwaite method and may need entering by hand where it is unavailable. Clinicians should verify every imported parameter against the current manual, and against local service standards, before interpreting it.

Score conversion and classification

Conversions between standard (M 100, SD 15), T (50, 10), scaled (10, 3) and z scores assume an approximately normal reference distribution. Two descriptor schemes are offered, and the one in force is named in the note beneath every exported table: Wechsler bands follow the WAIS-IV/WMS-IV manual conventions, and AACN labels follow Guilmette et al. (2020). Confidence levels throughout are 90% (z = 1.645) and 95% (z = 1.960); intervals round the estimate and the margin separately, so the printed bounds stay symmetric about the printed value.

Confidence intervals on Score Tables

Confidence intervals and standard errors of measurement. The CI column is the obtained score ± z × SEM, where SEM = SD × √(1 − r), centred on the obtained score rather than on an estimated true score.

The standard deviation is the normative SD of the metric the score is reported in — 15, 10, 3 or 1 — because a coefficient computed on, or corrected to, the normative sample must be paired with that sample's variability. Where a measure's stored statistics are raw, its own standard deviation is used instead, that being the only one in the right units. Four publishers state that rule outright, and the arithmetic confirms it: this pairing reproduces every published standard error of measurement the app is able to check, exactly, at the precision each is printed to — all 300 cells of WAIS-IV Table 4.3, 242 of WISC-V Table 4.4, 240 of WMS-IV Table 3.3, 168 across the D-KEFS SEM tables, 126 of RBANS Update Table 3.7, and all 38 CVLT-3 measures in Tables 3.4 and 3.5.

The reliability is, by default, the retest coefficient held in the normative database — an alternate-form coefficient in the case of the CVLT-3, which publishes no same-form retest — corrected for the normative sample's variability where the publisher reports a corrected value. Retest is the default for two reasons: it keeps a single, stated basis across a table that may mix batteries, and it is the appropriate coefficient for the many timed measures in the database, since split-half and alpha are not valid reliability estimates for speeded tests. The WAIS-IV manual makes that second argument itself for Coding, Symbol Search and Cancellation, describing the split-half coefficient as "not a proper reliability estimate" for a Processing Speed subtest; the values used here for those three are the ones it publishes, in all 38 of the cells its Table 4.1 gives them.

That default is set aside for a measure only where its publisher both reports an internal-consistency coefficient and derives its own published intervals from it. Seven manuals meet that bar:

InstrumentCoefficient usedSourceNot applied to
CVLT-C Odd–even split-half, by age Manual Table 6.5 Every index but List A Trials 1–5 Total; item scores on a word-list task are not independent. The interval printed in the manual's own worked example reproduces exactly.
D-KEFS Internal consistency, by normative age band Technical Manual, Tables 2.1–2.24 Colour–Word Interference, whose only coefficient table is for a composite this app does not hold; Design Fluency, where item interdependence precluded the procedure; and five of the six Trail Making measures, the published table covering the composite alone.
D-KEFS Advanced Split-half, by normative age band Table 3.4 Trail Making and Verbal Fluency, which that manual treats as speeded and scores on stability coefficients.
WAIS-IV Split-half or alpha, by normative age band Table 4.1 Coding, Symbol Search and Cancellation — speeded. These keep the corrected stability coefficient the same table publishes for them.
WISC-V Split-half, by single year of age Table 4.1 Coding, Symbol Search and Cancellation, together with the Cancellation Random and Cancellation Structured process scores — speeded, and likewise on the corrected stability coefficient.
WMS-IV Split-half or alpha, by normative age band, Adult and Older Adult batteries separately Table 3.1 Verbal Paired Associates II Word Recall, a free-recall score with no consistent item count, which takes a stability coefficient. The recognition memory measures are absent altogether: their published reliability is a decision-consistency percentage, not a correlation, and cannot enter a standard error of measurement.
RBANS Update Internal consistency, by normative age band Table 3.6 Figure Copy, Semantic Fluency, Coding, Story Recall and Figure Recall, which that table itself marks as estimated from test–retest and which therefore keep a stability coefficient taken from the same table. The four subtests reported as raw scores appear nowhere in it, the manual publishing reliability for its eight scaled subtests only, so no interval is shown for them.

Which coefficient is right is a question for each manual rather than a policy of this tool, and the manuals genuinely disagree — the two D-KEFS manuals reach opposite conclusions about the same two test names. Each is followed as written, and the Data page names the basis actually in force for every measure in the database.

Age. Where a coefficient is tabulated by age, the interval uses the band for the patient's age, and the age used is named in the note beneath the table. Entering an age is optional. If none is entered, or the age falls outside a measure's normed range, the publisher's all-ages figure is used instead: the published average where a manual prints one, and otherwise the total-sample retest coefficient — which for the D-KEFS is that manual's own second regime rather than a substitute for a missing number. Both paths are therefore the publisher's own figures.

Reliable-change analysis is unaffected by any of the above and always uses the retest coefficient, which measures a different thing.

One consequence is worth bearing in mind when comparing output against a test manual. For the measures that remain on the retest default, where a manual derives its published intervals from internal-consistency reliability — almost always the higher of the two coefficients — the intervals shown here run wider than the manual's. They are therefore the more conservative, and answer the question how much would this score be expected to move on retesting rather than how precisely was it measured on the day. Not every publisher offers that comparison: the CVLT-3 manual declines to report internal-consistency reliability at all, on the grounds that item scores on a word-list task are not independent — recalling one word alters the probability of recalling the others, both within a trial and on later ones — and reports alternate-form coefficients in their place.

Measures with no normative-sample coefficient, and the reliability control

Every interval here multiplies a normative standard deviation by a reliability, and that is only a valid standard error of measurement when the two describe the same group. Most manuals supply a coefficient computed on, or corrected to, their normative sample. Three do not — the D-KEFS, D-KEFS Advanced and CVLT-C manuals report only the correlation observed in their own retest studies, a few dozen people each, and pair it with the normative standard deviation regardless.

The D-KEFS manual states that outright, fixing the standard deviation unit at 3 for all its scaled scores and deriving its test–retest standard errors of measurement "from the total sample of cases". Its Table 2.8 shows the arithmetic: the three Design Fluency all-ages values of 1.94, 1.97 and 2.47 are exactly 3 × √(1 − r) on the uncorrected coefficient. Those measures are therefore scored the way their own manuals score them, and the intervals shown reproduce the published ones. The Data page labels each such measure retest, uncorrected, so which rows rest on that footing can be read off rather than inferred.

The statistical objection to the pairing is nonetheless real, so Score Tables offers a reliability control with two settings. Published, the default, uses each manual's own coefficient and reproduces its printed interval. Corrected applies the standard range-restriction correction of Allen and Yen, rxx = 1 − (s²retest ÷ s²norm)(1 − r), to those measures alone, so that the coefficient describes the same population as the standard deviation it multiplies. A published coefficient is never overwritten in either setting.

The correction is not a guess: across the 267 database entries carrying both an observed and a publisher-corrected coefficient, it reproduces the publisher's own value to a median error of .003. But the resulting figures are not printed in the manuals concerned, which is why the default is Published and why the note beneath a corrected table says so. In practice the control moves 46 of the measures reachable from Score Tables, every one of them D-KEFS or D-KEFS Advanced: at the 95% level 9 intervals widen and 4 narrow, 33 are unchanged after rounding, and the largest single change is 2 scaled-score points. Reliable-change analysis is not affected by the control. Where a corrected reading is taken, the Data page follows it rather than continuing to show the published one.

Change analysis

Five methods, in ascending order of what they model. The Standard Deviation Index is descriptive only: SD Δ = (X₂ − X₁) ÷ SD, with no reliability correction. Simple Reliable Change (Jacobson & Truax, 1991) tests the observed change against measurement error. Practice Effect-Adjusted change (Iverson, 2001) subtracts the mean retest gain observed in the normative sample first. McSweeney Regression-Based change (McSweeney et al., 1993) predicts the retest score from baseline and standardises the residual against the standard error of estimate. Crawford & Garthwaite Regression-Based change (2007) does the same but with a standard error of prediction that accounts for the normative sample size and for the distance of the baseline score from the normative mean.

All p-values are two-tailed. The first four methods use the standard normal distribution; Crawford & Garthwaite uses the Student t distribution with N − 2 degrees of freedom, so a small normative sample raises the threshold — at N = 25 the 95% critical value is 2.069 against 1.960 for z.

These calculations use the retest coefficient paired with the standard deviation of the same retest sample, so that both terms describe one population. Where a publisher reports a coefficient corrected to the normative sample's variability, that value is offered as an option but is not the default, because it describes a differently distributed population from the standard deviation it would be multiplied by — and in the two regression methods the coefficient is a fitted slope, so substituting it changes the predicted score rather than only the interval. Reliability type varies by instrument and is stated in the note beneath each generated table: CVLT-3 coefficients are alternate-form, RBANS Form A coefficients are same-form retest.

Outcomes are reported as significance only — "Reliable change" or "No reliable change" — never as improvement or decline. The database holds many measures on which a higher score is the worse result (intrusions, perseverations, errors, false positives), and carries no score-direction flag, so reading a clinical direction off the sign of the statistic would assert the wrong conclusion for all of them. The signed statistic is displayed alongside, so the direction stays visible without the app interpreting it.

Premorbid estimation

Premorbid estimates combine ToPF-based and demographic equations with OPIE-4 prorated models, and produce predicted-versus-achieved discrepancy output with confidence intervals, base-rate lookups and APA-formatted export tables. Predicted-versus-achieved significance flagging uses a three-tier scheme at z = 1.645 (*), 1.960 (**) and 2.576 (***); this is separate from the 90%/95% confidence-interval selector.

OPIE-4 is provided for illustration only in a UK context and its output should not be quoted as a concrete premorbid estimate. The regression terms reproduce Holdnack et al. (2013), Table eA5.8, but the published equations also carry US education, ethnicity and region terms that are not applied here, which fixes every prediction at the US reference category (12th-grade high-school graduate, not African-American, not resident in the western US). Those terms are omitted rather than mapped because the education dummies encode how unusual a given attainment level is within the US population the model was fitted on, not years of schooling, and that does not transfer: the US reference category corresponds to A-levels if matched by years but to GCSE/O-level if matched by population position, and UK school-leaving age was raised to 16 only in 1972, so leaving school without qualifications was normative for older cohorts in a way it was not in the US sample. Expect estimates to run high for patients who left school early and low for graduates, by an amount this tool cannot quantify. For a UK demographic estimate, use the Crawford & Allan (2001) model.

Base rates. The ToPF predicted-difference base rates are estimated from a normal model with SD equal to the model's standard error of estimate, not transcribed from observed standardisation-sample frequencies; they are labelled as such wherever they appear. The published ToPF/ACS predicted-difference tables have not been transcribed, so the model has instead been benchmarked against the one genuinely empirical table available for a model of almost identical standard error of estimate — the OPIE-4 discrepancy base rates of ACS Table eA5.12. Against that table the parametric values run roughly 10% relatively low across the decisive −5 to −20 band: a discrepancy of −15 gives 3.78% here against about 4.3% empirically. The net effect is to show a given discrepancy as marginally rarer, and so marginally more pathological, than observed data suggest. The OPIE-4 discrepancy base rates are themselves empirical and are used as published.

References
Test manuals and technical sources

Delis, D. C., Kaplan, E., & Kramer, J. H. (2001). Delis–Kaplan Executive Function System (D-KEFS): Technical manual. San Antonio, TX: The Psychological Corporation. [Internal-consistency coefficients and standard errors of measurement by age band, chapter 2 and Tables 2.1–2.26.]

Delis, D. C., Kramer, J. H., Kaplan, E., & Ober, B. A. (1994). California Verbal Learning Test – Children's Version (CVLT-C): Manual. San Antonio, TX: The Psychological Corporation. [Split-half reliability and standard errors of measurement: Table 6.5. Standardised score equivalents: Tables A.1 and A.2.]

Delis, D. C., Kramer, J. H., Kaplan, E., & Ober, B. A. (2017). California Verbal Learning Test – Third Edition (CVLT-3): Manual. Bloomington, MN: Pearson. [Alternate-form reliability and standard errors of measurement: Tables 3.4 and 3.5.]

Holdnack, J. A., Drozdick, L., Weiss, L. G., & Iverson, G. L. (2013). WAIS-IV, WMS-IV, and ACS: Advanced clinical interpretation. Oxford: Academic Press. [OPIE-4 prorated regression coefficients: Table eA5.8. OPIE-4 discrepancy base rates: Table eA5.12.]

Randolph, C. (2012). Repeatable Battery for the Assessment of Neuropsychological Status Update (RBANS Update): Manual. Bloomington, MN: Pearson. [Reliability by age band: Table 3.6. Standard errors of measurement: Table 3.7. Test–retest and alternate-form data: Tables 3.8–3.9.]

Wechsler, D. (2010). Wechsler Adult Intelligence Scale – Fourth UK Edition (WAIS–IVUK): Administration and scoring manual. London: Pearson Assessment. [Longest-span base rates: Tables C.4 and C.5.]

Wechsler, D. (2010). Wechsler Adult Intelligence Scale – Fourth UK Edition (WAIS–IVUK): Technical and interpretive manual. London: Pearson Assessment. [Reliability coefficients: Table 4.1. Standard errors of measurement: Table 4.3. Test–retest stability parameters: Table 4.5.]

Wechsler, D. (2010). Wechsler Memory Scale – Fourth UK Edition (WMS–IVUK): Technical and interpretive manual. London: Pearson Assessment. [Reliability coefficients: Table 3.1. Standard errors of measurement: Table 3.3.]

Wechsler, D. (2011). Test of Premorbid Functioning (ToPF-UK): Manual. London: Pearson Assessment.

Wechsler, D. (2014). Wechsler Intelligence Scale for Children – Fifth Edition (WISC-V): Technical and interpretive manual. Bloomington, MN: NCS Pearson. [Reliability coefficients: Table 4.1. Standard errors of measurement: Table 4.4. Test–retest stability: Table 4.7.]

Statistical and methodological sources

Allen, M. J., & Yen, W. M. (1979). Introduction to measurement theory. Monterey, CA: Brooks/Cole. [Correction of a reliability coefficient for range restriction in the sample it was observed in.]

Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Hillsdale, NJ: Lawrence Erlbaum.

Crawford, J. R., & Allan, K. M. (2001). Estimating premorbid WAIS–R IQ with demographic variables: Regression equations derived from a UK sample. The Clinical Neuropsychologist, 11(2), 192–197.

Crawford, J. R., & Garthwaite, P. H. (2007). Using regression equations built from summary data in the neuropsychological assessment of the individual case. Neuropsychology, 21(5), 611–620.

Crawford, J. R., Millar, J., & Milne, A. B. (2001). Estimating premorbid IQ from demographic variables: A comparison of a regression equation vs. clinical judgement. British Journal of Clinical Psychology, 40(1), 97–105.

Guilmette, T. J., Sweet, J. J., Hebben, N., Koltai, D., Mahone, E. M., Spiegler, B. J., Stucky, K., Westerveld, M., & Conference Participants. (2020). American Academy of Clinical Neuropsychology consensus conference statement on uniform labeling of performance test scores. The Clinical Neuropsychologist, 34(3), 437–453.

Iverson, G. L. (2001). Interpreting change on the WAIS-III/WMS-III in clinical samples. Archives of Clinical Neuropsychology, 16(2), 183–191.

Jacobson, N. S., & Truax, P. (1991). Clinical significance: A statistical approach to defining meaningful change in psychotherapy research. Journal of Consulting and Clinical Psychology, 59(1), 12–19.

McSweeney, A. J., Naugle, R. I., Chelune, G. J., & Lüders, H. (1993). "T scores for change": An illustration of a regression approach to depicting change in clinical neuropsychology. The Clinical Neuropsychologist, 7(3), 300–312.

Sawilowsky, S. S. (2009). New effect size rules of thumb. Journal of Modern Applied Statistical Methods, 8(2), 597–599.

Effect-size benchmarks

Cipriani, A., Furukawa, T. A., Salanti, G., Chaimani, A., Atkinson, L. Z., Ogawa, Y., Leucht, S., Ruhe, H. G., Turner, E. H., Higgins, J. P. T., Egger, M., Takeshima, N., Hayasaka, Y., Imai, H., Shinohara, K., Tajika, A., Ioannidis, J. P. A., & Geddes, J. R. (2018). Comparative efficacy and acceptability of 21 antidepressant drugs for the acute treatment of adults with major depressive disorder: A systematic review and network meta-analysis. The Lancet, 391(10128), 1357–1366.

Cuijpers, P., Geraedts, A. S., van Oppen, P., Andersson, G., Markowitz, J. C., & van Straten, A. (2011). Interpersonal psychotherapy for depression: A meta-analysis. American Journal of Psychiatry, 168(6), 581–592.

Cuijpers, P., Miguel, C., Harrer, M., Plessen, C. Y., Ciharova, M., Ebert, D., & Karyotaki, E. (2023). Cognitive behavior therapy vs. control conditions, other psychotherapies, pharmacotherapies and combined treatment for depression: A comprehensive meta-analysis including 409 trials with 52,702 patients. World Psychiatry, 22(1), 105–115.

Hofmann, S. G., & Smits, J. A. J. (2008). Cognitive-behavioral therapy for adult anxiety disorders: A meta-analysis of randomized placebo-controlled trials. The Journal of Clinical Psychiatry, 69(4), 621–632.

Huhn, M., Nikolakopoulou, A., Schneider-Thoma, J., Krause, M., Samara, M., Peter, N., Arndt, T., Bäckers, L., Rothe, P., Cipriani, A., Davis, J., Salanti, G., & Leucht, S. (2019). Comparative efficacy and tolerability of 32 oral antipsychotics for the acute treatment of adults with multi-episode schizophrenia: A systematic review and network meta-analysis. The Lancet, 394(10202), 939–951.

Pesch, B., Kendzia, B., Gustavsson, P., Jöckel, K.-H., Johnen, G., Pohlabeln, H., Olsson, A., Ahrens, W., Gross, I. M., Brüske, I., Wichmann, H.-E., Merletti, F., Richiardi, L., Simonato, L., Fortes, C., Siemiatycki, J., Parent, M.-E., Consonni, D., Landi, M. T., … Brüning, T. (2012). Cigarette smoking and lung cancer — Relative risk estimates for the major histological types from a pooled analysis of case–control studies. International Journal of Cancer, 131(5), 1210–1219.

Storebø, O. J., Storm, M. R. O., Pereira Ribeiro, J., Skoog, M., Groth, C., Callesen, H. E., Schaug, J. P., Darling Rasmussen, P., Huus, C.-M. L., Zwi, M., Kirubakaran, R., Simonsen, E., & Gluud, C. (2023). Methylphenidate for children and adolescents with attention deficit hyperactivity disorder (ADHD). Cochrane Database of Systematic Reviews, 3, CD009885.

Watts, B. V., Schnurr, P. P., Mayo, L., Young-Xu, Y., Weeks, W. B., & Friedman, M. J. (2013). Meta-analysis of the efficacy of treatments for posttraumatic stress disorder. The Journal of Clinical Psychiatry, 74(6), e541–e550.

Williams, A. C. de C., Fisher, E., Hearn, L., & Eccleston, C. (2020). Psychological therapies for the management of chronic pain (excluding headache) in adults. Cochrane Database of Systematic Reviews, 8, CD007407.

Wolraich, M. L., Wilson, D. B., & White, J. W. (1995). The effect of sugar on behavior or cognition in children: A meta-analysis. JAMA, 274(20), 1617–1621.

The Effect Size Tools page also cites a UK Biobank male-versus-female adult height contrast (d = 2.04) attributed to Lui et al. (2021). That source could not be verified and is not listed here; treat the value as an illustrative anchor only.

✓ Table copied - ready to paste into your report