Skip to main content

Psychometric Assistant

Neuropsychological calculators for clinical practice

Estimate premorbid ability using ToPF and OPIE-4, compared against WAIS-IV and WMS-IV.

Build APA-formatted report tables with confidence-interval columns and premorbid comparison.

Evaluate test-retest change across five methods, from descriptive SDI to Crawford regression.

Score embedded and stand-alone PVTs against published cut-offs, with base-rate-aware interpretation.

Convert between standard, scaled, T and z scores and percentiles, with classification bands.

Convert between and interpret effect sizes: Cohen's d, r, η², odds ratios and more.

Chart scores entered elsewhere in the suite against classification and reliable-change bands.

Ask how many of a patient’s scores are abnormal, and how many healthy people show that many.

Psychometrics · Standardised Score Conversion

Score Converter


−3 SD −2 SD −1 SD M +1 SD +2 SD +3 SD
Score equivalents
-3σ-2σ-1σ0σ+1σ+2σ+3σ557085100115130145SSσz = 0.00 P(Z ≤ z) 50.0%
Full conversion table every score, scannable; the entered score's row is highlighted
Rows anchored on

A whole score appears only on the row where it actually falls; a blank cell means that metric has no whole score at that position. Columns finer than the anchored one show their nearest whole score on every row; z is exact to 2 dp. Percentile and classification conventions match the converter above. Ranges: Standard 40–160 · T 10–90 · Scaled 1–19 · z −4 to +4 in 0.25 steps.

Wechsler
AACN

AACN = American Academy of Clinical Neuropsychology · Ranges shown as Standard Score (SS)

Conversion · Clinical Outcomes Table

Clinical Outcomes Table

Configuration
Premorbid Comparison

SD mode: * ≥1 SD below, ** ≥1.5 SD, *** ≥2 SD. SEE mode: * below 90% CI, ** below 95% CI, *** below 99% CI lower bound.

CVLT-3 Indices · Ages 16-44
CVLT-3 Indices · Ages 45-90
CVLT-3 Trials · Ages 16-44
CVLT-3 Trials · Ages 45-90
D-KEFS Colour-Word Interference · Ages 20-49
D-KEFS Colour-Word Interference · Ages 50-89
D-KEFS Colour-Word Interference · Ages 8-19
D-KEFS Colour-Word Interference · All Ages
D-KEFS Design Fluency · Ages 20-49
D-KEFS Design Fluency · Ages 50-89
D-KEFS Design Fluency · Ages 8-19
D-KEFS Design Fluency · All Ages
D-KEFS Sorting Test · Ages 20-49
D-KEFS Sorting Test · Ages 50-89
D-KEFS Sorting Test · Ages 8-19
D-KEFS Sorting Test · All Ages
D-KEFS Tower Test · Ages 20-49
D-KEFS Tower Test · Ages 50-89
D-KEFS Tower Test · Ages 8-19
D-KEFS Tower Test · All Ages
D-KEFS Trail Making Test · Ages 20-49
D-KEFS Trail Making Test · Ages 50-89
D-KEFS Trail Making Test · Ages 8-19
D-KEFS Trail Making Test · All Ages
D-KEFS Verbal Fluency · Ages 20-49
D-KEFS Verbal Fluency · Ages 50-89
D-KEFS Verbal Fluency · Ages 8-19
D-KEFS Verbal Fluency · All Ages
D-KEFS Word Context Test · Ages 20-49
D-KEFS Word Context Test · Ages 50-89
D-KEFS Word Context Test · Ages 8-19
D-KEFS Word Context Test · All Ages
D-KEFS Word Proverb Test · Ages 20-49
D-KEFS Word Proverb Test · Ages 50-89
D-KEFS Word Proverb Test · Ages 8-19
D-KEFS Word Proverb Test · All Ages
RBANS Indices · Ages 12-19
RBANS Indices · Ages 20-89
RBANS Subtests · Ages 12-19
RBANS Subtests · Ages 20-89
WAIS-IV Core Subtests · Ages 16-29
WAIS-IV Core Subtests · Ages 30-54
WAIS-IV Core Subtests · Ages 55-69
WAIS-IV Core Subtests · Ages 70-90
WAIS-IV Core Subtests · All Ages
WAIS-IV Indices · Ages 16-29
WAIS-IV Indices · Ages 30-54
WAIS-IV Indices · Ages 55-69
WAIS-IV Indices · Ages 70-90
WAIS-IV Indices · All Ages
WISC-V Indices · All Ages
WISC-V Subtests · All Ages
WMS-IV Indices · Ages 16-69
WMS-IV Indices · Ages 65-90
WMS-IV Subtests · Ages 16-69
WMS-IV Subtests · Ages 65-90
# Subtest Raw Score CI Percentile Classification
APA-formatted output
Enter at least one subtest with a score to preview the APA table.
Basic Report Tools · Effect Size

Effect Size Tools

Convert between effect-size metrics or derive them from group data, with a visual against the standard normal.

OR
Preloaded example UK male vs female height is loaded in the group-data fields from Option B (above). Edit the values or use Clear all inputs to start fresh.
0.00
Intervention (Group 1)
95% CI: -
Control (Group 2)
95% CI: -

Example: measured adult height (16+), men vs women, Health Survey for England 2022 (NHS England). Means and n as published in Table 1; SDs weighted, calculated from the survey data (National Centre for Social Research & University College London, 2025; UK Data Service SN 9469).

Pooled N-
Pooled Mean-
Pooled SD-
Pooled SE-
Mean Diff-
Effect-size results
Cohen's d
-
-
Hedges' g
-
-
Pearson's r
-
-
R²
-
-
Cohen's f
-
-
Fisher's z
-
Odds Ratio
-
Overlap
-
Cohen's U₃
-
CLES
-
Kraemer's NNT
-
Similar to
-
Group comparison at a target value
Quick compare. Enter a target value to see how each group performs at or above that point.
-
-
-
-
-
Requires means and standard deviations. If standard errors are entered above, the app converts them to SDs using each group's sample size.
Reference values (Cohen's d)

Use these published effect sizes to anchor your results. The values below span small to huge magnitudes so you can compare your finding against familiar clinical and epidemiological benchmarks.

Heavy smokers (30+/day) vs never smokers, lung cancer (Pesch et al., 2012)2.60
Male vs female adult height, England (Health Survey for England 2022; National Centre for Social Research & University College London, 2025)1.91
Smokers (any) vs never smokers, lung cancer (Pesch et al., 2012)1.75
Cognitive therapy vs control for PTSD (Watts et al., 2013)1.63
Former smokers vs never smokers, lung cancer (Pesch et al., 2012)1.10
Exposure therapy vs control for PTSD (Watts et al., 2013)1.08
EMDR vs control for PTSD (Watts et al., 2013)1.00
Clozapine vs placebo for schizophrenia (Huhn et al., 2019 Lancet)0.89
CBT vs control for depression (Cuijpers et al., 2023 World Psychiatry)0.80
Methylphenidate vs placebo for ADHD, children (Storebø et al., 2023 Cochrane)0.75
CBT vs placebo for anxiety disorders (Hofmann & Smits, 2008)0.70
CBT for depression, low-risk-of-bias subset (Cuijpers et al., 2023)0.60
Interpersonal Therapy for depression (Cuijpers et al., 2011)0.50
Antidepressants vs placebo (Cipriani et al., 2018 Lancet)0.30
CBT vs treatment-as-usual for chronic pain (Williams et al., 2020 Cochrane)0.20
CBT vs active control for chronic pain (Williams et al., 2020 Cochrane)0.10
Sugar on children's hyperactivity (Wolraich et al., 1995 JAMA)0.00
No or negligible effect0.00
Common language description
Interpretation summary.
Enter a valid effect size to generate a plain-English interpretation.
The average person in Group 1 is above about - of Group 2 (Cohen's U₃).
Visual
Group 2 (μ = 0) Group 1 (μ = d)
Curves shifted by current Cohen's d.
Change & Discrepancy · Profile Analysis

Profile Analysis

APA-formatted output
Performance Validity · Embedded & Stand-Alone PVTs

Performance Validity

Score embedded and stand-alone validity indicators against their published cut-offs, then weigh them together. No single index is a verdict. Results are cut-off comparisons only.

Scores & settings
Caution. Over-flags in dementia (~48%) and severe impairment; corroborate before concluding.
Both subtests load on attention/working memory and verbal recognition, so aphasia or significant memory disorder can inflate the EI while performance is valid. Severely impaired geriatric samples show ~31–37% failure rates. Do not apply single-subtest cut-offs in place of the combined EI, and do not declare a whole RBANS protocol invalid on the EI alone. Use it to generate hypotheses and trigger further validity testing (Silverberg et al., 2007).
Cut-offs: Silverberg, Wertheimer & Fichtenberg (2007), The Clinical Neuropsychologist, 21(5), 841–854. Accuracy: Shura et al. (2018), Neuropsychology Review, 28(3), 269–284, pooled estimates.
Formula & weighting table

Each raw score converts to a weighted score (Silverberg, Wertheimer & Fichtenberg, 2007, Table 2); EI = weighted Digit Span + weighted List Recognition (range 0–12). Higher EI = less credible performance. Some weights are unreachable for a given subtest (Digit Span never yields 1 or 4).

Digit Span (raw)List Recognition (raw)Weighted score
8–1618–200
—171
715–162
613–143
—11–124
5105
0–40–96
Scores & settings
Mandatory gate. Computed only where Digit Span < 9, List Recognition < 19, or their sum < 28 (Novitski et al., 2012).
Caution. Unstable outside Alzheimer-type amnesia; confirm with a stand-alone PVT.
Built around the Alzheimer's recall/recognition dissociation, so mixed or atypical memory profiles may not match the pattern: later work shows ~4% failure in AD but ~31% in non-AD dementia and >60% suboptimal rates in Parkinson's disease. The amnestic validation sample had no independent performance validity testing (Novitski et al., 2012).
Formula, gate & cut-off: Novitski, Steele, Karantzoulis & Randolph (2012), Archives of Clinical Neuropsychology, 27(2), 190–195. No published sensitivity/specificity pair. Discrimination is ROC AUC = .91 vs amnestic patients (vs .61 for the EI).
Formula

ES = ( List Recognition − [ List Recall + Story Recall + Figure Recall ] ) + Digit Span, all raw scores (Novitski et al., 2012). Lower ES = more suspicious of invalid performance. Cut-off: ES < 12, applicable only once the gating rule is met. The gate exists because in cognitively intact examinees free recall normally far exceeds the ceiling-limited recognition score, so an ungated ES over-flags; only 17% of the normal standardisation sample fall below the combined < 28 screen. But 15.1% of that sample meet the screen and also score ES < 12, so even with the gate roughly one healthy adult in seven would be flagged; Novitski et al. note the ES “will produce high false positive rates in the normal population”.

Scores & settings
Classic variant. Forward + Backward only; the WAIS-IV Sequencing variants use different cut-offs and are not scored here.
Caution. Prefer ≤ 6 in genuine impairment, and even ≤ 6 over-flags in some groups.
RDS loads on attention/working memory: low scores occur legitimately in genuine attentional impairment, anxiety, and especially conduction aphasia. Schroeder et al. (2012) found ≤ 7 held only 82–85% specificity across clinical groups, against 96–97% for ≤ 6. Even at ≤ 6, specificity remains inadequate in cerebrovascular accident, severe memory disorders, intellectual disability or borderline intellectual functioning, and English-as-a-second-language examinees. In early Alzheimer’s disease, RDS ≤ 6 gave a 13% false-positive rate (Loring et al., 2016, as reported by Sweet et al., 2021). Validation samples were predominantly mild-TBI/litigating.
Scoring: Greiffenstein, Baker & Gola (1994), Psychological Assessment, 6(3), 218–224. Cut-off accuracy: Schroeder et al. (2012), Assessment, 19(1), 21–30.
Scoring & worked example

RDS = (longest forward span passed on both trials) + (longest backward span passed on both trials) (Greiffenstein, Baker & Gola, 1994; cross-validated by Meyers & Volbrecht, 1998). The classic Forward + Backward variant is what the validation literature uses. Document which variant you administered. Floor rule: failing at least one trial each of 3 forward and 2 backward is recorded as RDS = 3. Worked example from the original paper: passes both trials of 3 forward but fails one 4-forward trial → forward = 3; passes both trials of 2 backward but fails one 3-backward trial → backward = 2; RDS = 3 + 2 = 5. Reference values: mean RDS ~8.8 in non-malingering TBI vs ~6.7 in probable malingerers.

Scores & settings
Caution. Not a stand-alone measure, and it shares its subtest with RDS, so the summary counts them as one indicator.
Axelrod et al. explicitly recommend against using the Digit Span scaled score as a stand-alone validity measure, as it accounted for a modest share of variance (AUC .77). Genuine impairment can suppress the score, so apply the conservative cut-off where real deficits are plausible. Both sources are WAIS-III. The WAIS-IV subtest adds Sequencing to the composite, so an ACSS from a WAIS-IV record is not on the identical metric these cut-offs were derived on.
Base rates & suspicion indices: Iverson & Tulsky (2003), Archives of Clinical Neuropsychology, 18(1), 1–9. Classification accuracy: Axelrod, Fichtenberg, Millis & Wertheimer (2006), The Clinical Neuropsychologist, 20(3), 513–523. Both WAIS-III.
Indices & base rates

The Digit Span age-corrected scaled score is entered as scored on the record form. The age correction is already applied by the Wechsler norms, so no separate age entry is needed here. ACSS ≤ 5 occurs in 3.8% of the WAIS-III standardisation sample and 3.4% of a combined clinical sample spanning TBI, alcohol abuse, Korsakoff's, temporal lobectomy and Alzheimer's disease (Iverson & Tulsky's suspicion guideline). Axelrod et al. found ≤ 7 the best single discriminator of probable malingering from TBI patients in neurorehabilitation (sens. .75, spec. .69, with .77 on non-litigating mild-TBI cross-validation), and no genuine mild-TBI patient scored below 6. Iverson & Tulsky give three further suspicion indices, each evaluated against base rates: a Vocabulary − Digit Span difference ≥ 5 (7.1% of the standardisation sample, 2.8% of the combined clinical sample), a longest span forward ≤ 4 applicable only under age 55 (base rates 2.5–5.5% in the under-55 bands, rising to 9.5–12.7% from age 55, which is why the index is age-limited; the app reads the top-bar patient age and withholds it without one), and a longest span backward ≤ 2 (2.0–6.0% across all bands, 3.4% clinical). The longest spans here are the longest passed on either trial, as the standardisation data tabulate them, not the RDS both-trials span. Both papers derive from the WAIS-III. Document the edition administered.

Scores & settings
Times the 10-second exposure and scores the recognition trial. Displays test stimuli: start only with the examinee present.
Caution. Specificity falls in dementia and intellectual disability; corroborate before concluding.
The minority of studies reporting low specificity for the < 9 cut-off generally included patients with intellectual disability or dementia, so avoid these cut-offs there. Certain false-positive recognition errors (circling "f", "5", "6" or a pentagon) occurred only in the noncredible group and may serve as virtual pathognomonic signs when present, though they are infrequent (Boone et al., 2002).
Recognition trial, combination score & accuracy: Boone, Salazar, Lu, Warner-Chacon & Razani (2002), Journal of Clinical and Experimental Neuropsychology, 24(5), 561–573. Recall < 9: sens. .47 · spec. .97–1.00. Combination < 20: sens. .71 · spec. .92–.94.
Scoring & administration

Combination score = free recall + (recognition correct − false positives), cut-off < 20; the classic free-recall cut-off is < 9 (Boone et al., 2002, Table 6; the combination raises sensitivity from .47 to .71 at comparable specificity). Administration: expose the stimulus page for 10 seconds ("there are 15 different things so you will have to learn them very quickly"), remove it and have the examinee draw what they remember; then present the recognition page ("on this page are the 15 things I showed you as well as 15 items that were not on the page, circle the things you remember"). The stimulus and recognition pages are deliberately not reproduced here.

Scores & settings
Caution. Failure rates climb with genuine impairment; a near-perfect score rules nothing out.
The manual is explicit that forced-choice recognition is suited to blatant exaggeration: “a poor score on forced choice testing is often a strong indicator of symptom amplification, but a perfect or near-perfect score does not rule it out”, since a sophisticated simulator recognises the task as easy. False positives concentrate in severe impairment. Failure rates rise with the severity of neurocognitive disorder, and are substantially elevated in dementia and intellectual disability, so the cut-offs should not be read at face value there (Schwartz et al., 2016). Conversely there is a well-replicated reverse severity effect in head injury: patients with mild TBI were two to three times more likely to fail than those with moderate-to-severe TBI, and failure was unrelated to injury severity, neuroradiological findings or performance on tests sensitive to TBI (Erdodi, Abeare, et al., 2018). Appendix D tabulates the Standard and Alternate Forms only. The Brief Form's 9-item list has no published base rates, and the trial is given after the Yes/No Recognition trial, roughly 10 minutes on (Delis et al., 2017).
Base rates: Delis, Kramer, Kaplan & Ober (2017), CVLT-3, Appendix D, Tables D.13–D.15 (Standard/Alternate Forms). The manual publishes base rates by age band, not a cut-off. Cut-offs and accuracy: Erdodi, Abeare, et al. (2018) and Schwartz et al. (2016), both CVLT-II.
How the cut-off is set, and why it changes with age

No published cut-off. The CVLT-3 manual gives no cut score for Forced Choice. Table D.13 gives, by age band, the percentage of the standardisation sample scoring each number of hits or fewer. By default a score is flagged when that percentage is at or below the false-positive rate set under Base-rate criterion: 10% by default, the consensus specificity of .90 per test (Sweet et al., 2021), or 5%. With no age entered, the manual's All Ages column is used.

Age matters. 15 of 16 hits is the 8.6th percentile across all ages but the 26th at ages 80–90. So the flag is ≤ 15 in every band except 80–90, where it is ≤ 14. A score of 15 is a 3–6% event in a working-age adult and an unremarkable one in an 85-year-old.

Published cut-offs (CVLT-II). Two published cut-offs can be chosen instead, under Flag basis. Both come from the CVLT-II; the trial is the same in both editions (16 List A targets, one distractor each, about 10 minutes after Yes/No Recognition). ≤ 14 is the standard: across 17 studies and 4,432 patients it gives sensitivity .50 at specificity .93, and no healthy control in the review scored at or below it (Schwartz et al., 2016). ≤ 15 counts a single error as a failure: in 104 adults with TBI it gave sensitivity .56 at specificity .92 against seven reference PVTs, identifying about 6% more invalid response sets (Erdodi, Abeare, et al., 2018). For every patient under 80, the age-based flag is already ≤ 15.

A fail is informative; a pass is not. Examinees who failed a reference PVT were about eight times more likely to fail Forced Choice, but roughly half of invalid response sets pass it, because a sophisticated simulator recognises how easy the task is (Schwartz et al., 2016).

Critical items are targets recalled, or recognised on Yes/No, earlier in the test but not chosen on Forced Choice. Tables D.14 and D.15 give the percentage scoring that many or more, by age band, and the critical-item flag always uses those tables, whichever cut-off is set for hits. Hits and both critical-item scores come from one administration, so together they count as one indicator.

Scores & settings
Caution. One study of 157 outpatients, excluding dementia and intellectual disability; corroborate before concluding.
The authors call the findings preliminary and ask for replication. Accuracy varied with the criterion it was measured against, and was lowest against a word-recognition test (the Word Choice Test). At ≤ 5, Conditions 2 and 3 fell below .84 specificity against its accuracy score (.77 to .78). The Condition 5 cut-off (≤ 8) lies in the average range; the authors argue for it from the score distribution and from the D-KEFS norms not screening for invalid performance. The authors note that on this task genuine impairment and invalid performance may be hard to separate (Erdodi, Hurtubise, et al., 2018).
Cut-offs & accuracy: Erdodi, Hurtubise, Charron, Dunn, Enache, McDermott & Hirst (2018), Psychological Assessment, 30(8), 1082–1095, Tables 5 and 6.
How the flag is set

Each condition's age-corrected scaled score is compared with the cut-off the paper found best for that condition: ≤ 5 on Conditions 1 to 3, ≤ 4 on Condition 4 and ≤ 8 on Condition 5. The cut-offs differ because the conditions differ in difficulty; a single cut-off across them, as on the Wechsler subtests, does not hold here. The indicator fails when at least the chosen number of conditions fail. The paper favours combining the conditions over any single one, and publishes combined accuracy for three, four and all five; three is the default, the combination the authors describe as a good balance of sensitivity and specificity.

The paper tested every cut-off against four criteria: the Word Choice Test's accuracy and completion time, and two composites of five embedded indicators each, one verbal and one processing-speed based. Sensitivity and specificity are therefore shown as the range across the four. The Condition 4 to Condition 2 time ratio is not used: its sensitivity was .00 to .09 at the ≤ 1.5 cut-off.

Scores & settings
Caution. Do not interpret traditional cut-offs in suspected or confirmed dementia.
Across dementia samples, weighted-mean specificity was ≤ .70 at traditional cut-offs, with no study achieving a false-positive rate under 10%. Specificity for the liberal < 49 cut-off also dipped below 90% in moderate/severe TBI and severe depression, and simulator studies overestimate sensitivity relative to real known-groups patients (Martin et al., 2020).
Test: Tombaugh (1996). Cut-offs & accuracy: Martin et al. (2020), The Clinical Neuropsychologist, 34(1), 88–119. Meta-analytic weighted means, shown per trial in the predictive-power table.
Cut-offs & classification accuracy

Meta-analytic weighted-mean specificity/sensitivity for neurocognitive and psychiatric samples (Martin et al., 2020): Trial 2 / Retention < 45: spec. .96–.98, sens. .45–.55; Trial 2 / Retention < 49: spec. .91–.97, sens. .59–.70; Trial 1 < 42: spec. .91, sens. .67–.69; Trial 1 < 41: spec. .93, sens. .66. All are Martin et al.'s weighted-mean values for neurocognitive/psychiatric samples; an earlier aggregation of Trial 1 (Denning, 2012) reported .92/.77 averaged across cut-offs 34–44, which is why the meta-analysis examined each cut-off individually. Positive and negative predictive power are derived from these values and the selected base rate by Bayes' theorem, as in Martin et al. Tables 16–17.

Aggregating multiple PVTs (Larrabee, 2014a)

This page applies the ≥ 2 independent failures rule from the AACN consensus statement (Sweet et al., 2021), which supports it when up to 7 to 9 validity measures are given. The page counts at most six independent indicators, which is within that range. Probable invalidity rests on the PVTs alone; calling it malingering also needs a substantial external incentive, under Sherman et al.’s (2020) update of Slick et al.’s (1999) criteria.

The table shows how the rule performed in one study: Larrabee's (2014a) 54 clinical and 41 malingering cases, each given seven indicators (6 PVTs and 1 SVT). It shows the trade-off between thresholds. It is not a false-positive rate for this page's measures, which Bilder, Sugar & Hellemann (2014) argue can be known only for combinations studied together.

ThresholdSpecificitySensitivityTotal correct
≥ 2 of 7 failures88.9%97.6%92.6%
≥ 3 of 7 failures96.3%87.8%92.6%
≥ 4 of 7 failures100%63.4%84.2%
Why ≥ 2, and when to demand more

Monte Carlo estimates overstate how often genuine patients fail two or more PVTs: their PVT scores sit at ceiling and are skewed, not normally distributed (Larrabee, 2014a). The count holds only for independent indicators; measures built from the same subtest raise each other’s failure rate (Larrabee, 2014a). On this page, the Effort Index and Effort Scale share RBANS subtests and count as one, and RDS also draws on Digit Span. In Larrabee’s sample, the six non-malingering patients who failed two or more had severe TBI with prolonged coma, complicated mild TBI, or stroke with a lesion on CT, and many failed only just inside the invalid range (RDS 7, WCST failure-to-maintain-set 2). Before reading two failures as probable invalidity in a genuinely impaired patient, check whether the patient’s other scores make the failures unlikely to be false positives (Larrabee, 2014a).

APA-formatted output
Enter at least one validity measure to preview.
Discrepancy · Standard Deviation Index

Standard Deviation Index

Quantify abnormality of test-retest discrepancy in standard-deviation units. Useful when reliability data are unavailable or for descriptive comparison.

APA-formatted output
Enter at least one subtest to preview.
Reliable Change Indices (RCI) · Simple

Basic Reliable Change Index

Jacobson & Truax (1991). Computes whether observed change exceeds measurement error, using the test's reliability coefficient and standard deviation.

CVLT-3 Indices · Ages 16-44
CVLT-3 Indices · Ages 45-90
CVLT-3 Trials · Ages 16-44
CVLT-3 Trials · Ages 45-90
D-KEFS Colour-Word Interference · Ages 20-49
D-KEFS Colour-Word Interference · Ages 50-89
D-KEFS Colour-Word Interference · Ages 8-19
D-KEFS Colour-Word Interference · All Ages
D-KEFS Design Fluency · Ages 20-49
D-KEFS Design Fluency · Ages 50-89
D-KEFS Design Fluency · Ages 8-19
D-KEFS Design Fluency · All Ages
D-KEFS Sorting Test · Ages 20-49
D-KEFS Sorting Test · Ages 50-89
D-KEFS Sorting Test · Ages 8-19
D-KEFS Sorting Test · All Ages
D-KEFS Tower Test · Ages 20-49
D-KEFS Tower Test · Ages 50-89
D-KEFS Tower Test · Ages 8-19
D-KEFS Tower Test · All Ages
D-KEFS Trail Making Test · Ages 20-49
D-KEFS Trail Making Test · Ages 50-89
D-KEFS Trail Making Test · Ages 8-19
D-KEFS Trail Making Test · All Ages
D-KEFS Verbal Fluency · Ages 20-49
D-KEFS Verbal Fluency · Ages 50-89
D-KEFS Verbal Fluency · Ages 8-19
D-KEFS Verbal Fluency · All Ages
D-KEFS Word Context Test · Ages 20-49
D-KEFS Word Context Test · Ages 50-89
D-KEFS Word Context Test · Ages 8-19
D-KEFS Word Context Test · All Ages
D-KEFS Word Proverb Test · Ages 20-49
D-KEFS Word Proverb Test · Ages 50-89
D-KEFS Word Proverb Test · Ages 8-19
D-KEFS Word Proverb Test · All Ages
RBANS Indices · Ages 12-19
RBANS Indices · Ages 20-89
RBANS Subtests · Ages 12-19
RBANS Subtests · Ages 20-89
WAIS-IV Core Subtests · Ages 16-29
WAIS-IV Core Subtests · Ages 30-54
WAIS-IV Core Subtests · Ages 55-69
WAIS-IV Core Subtests · Ages 70-90
WAIS-IV Core Subtests · All Ages
WAIS-IV Indices · Ages 16-29
WAIS-IV Indices · Ages 30-54
WAIS-IV Indices · Ages 55-69
WAIS-IV Indices · Ages 70-90
WAIS-IV Indices · All Ages
WISC-V Indices · All Ages
WISC-V Subtests · All Ages
WMS-IV Indices · Ages 16-69
WMS-IV Indices · Ages 65-90
WMS-IV Subtests · Ages 16-69
WMS-IV Subtests · Ages 65-90

Test data & patient scores

# Subtest SD r Date 1 Date 2 RCI (z) p Outcome
APA-formatted output
Enter test data and patient scores to preview the APA table.
Reliable Change Indices (RCI) · Practice Effects

Practice Effect-Adjusted Reliable Change Index

Iverson (2001). Adjusts the standard RCI to control for the average improvement (practice effect) observed between assessments in the normative sample.

CVLT-3 Indices · Ages 16-44
CVLT-3 Indices · Ages 45-90
CVLT-3 Trials · Ages 16-44
CVLT-3 Trials · Ages 45-90
D-KEFS Colour-Word Interference · Ages 20-49
D-KEFS Colour-Word Interference · Ages 50-89
D-KEFS Colour-Word Interference · Ages 8-19
D-KEFS Colour-Word Interference · All Ages
D-KEFS Design Fluency · Ages 20-49
D-KEFS Design Fluency · Ages 50-89
D-KEFS Design Fluency · Ages 8-19
D-KEFS Design Fluency · All Ages
D-KEFS Sorting Test · Ages 20-49
D-KEFS Sorting Test · Ages 50-89
D-KEFS Sorting Test · Ages 8-19
D-KEFS Sorting Test · All Ages
D-KEFS Tower Test · Ages 20-49
D-KEFS Tower Test · Ages 50-89
D-KEFS Tower Test · Ages 8-19
D-KEFS Tower Test · All Ages
D-KEFS Trail Making Test · Ages 20-49
D-KEFS Trail Making Test · Ages 50-89
D-KEFS Trail Making Test · Ages 8-19
D-KEFS Trail Making Test · All Ages
D-KEFS Verbal Fluency · Ages 20-49
D-KEFS Verbal Fluency · Ages 50-89
D-KEFS Verbal Fluency · Ages 8-19
D-KEFS Verbal Fluency · All Ages
D-KEFS Word Context Test · Ages 20-49
D-KEFS Word Context Test · Ages 50-89
D-KEFS Word Context Test · Ages 8-19
D-KEFS Word Context Test · All Ages
D-KEFS Word Proverb Test · Ages 20-49
D-KEFS Word Proverb Test · Ages 50-89
D-KEFS Word Proverb Test · Ages 8-19
D-KEFS Word Proverb Test · All Ages
RBANS Indices · Ages 12-19
RBANS Indices · Ages 20-89
RBANS Subtests · Ages 12-19
RBANS Subtests · Ages 20-89
WAIS-IV Core Subtests · Ages 16-29
WAIS-IV Core Subtests · Ages 30-54
WAIS-IV Core Subtests · Ages 55-69
WAIS-IV Core Subtests · Ages 70-90
WAIS-IV Core Subtests · All Ages
WAIS-IV Indices · Ages 16-29
WAIS-IV Indices · Ages 30-54
WAIS-IV Indices · Ages 55-69
WAIS-IV Indices · Ages 70-90
WAIS-IV Indices · All Ages
WISC-V Indices · All Ages
WISC-V Subtests · All Ages
WMS-IV Indices · Ages 16-69
WMS-IV Indices · Ages 65-90
WMS-IV Subtests · Ages 16-69
WMS-IV Subtests · Ages 65-90

Test data & patient scores

# Subtest M₁ SD₁ M₂ SD₂ r Date 1 Date 2 RCI (z) p Outcome
APA-formatted output
Enter test data and patient scores to preview the APA table.
Reliable Change Indices (RCI) · Regression-Based

McSweeny Regression-Based (SRB) Reliable Change Index

McSweeny et al. (1993). Predicts each patient's expected retest score from their baseline and the normative sample's regression parameters; the residual is standardised against the standard error of estimate.

CVLT-3 Indices · Ages 16-44
CVLT-3 Indices · Ages 45-90
CVLT-3 Trials · Ages 16-44
CVLT-3 Trials · Ages 45-90
D-KEFS Colour-Word Interference · Ages 20-49
D-KEFS Colour-Word Interference · Ages 50-89
D-KEFS Colour-Word Interference · Ages 8-19
D-KEFS Colour-Word Interference · All Ages
D-KEFS Design Fluency · Ages 20-49
D-KEFS Design Fluency · Ages 50-89
D-KEFS Design Fluency · Ages 8-19
D-KEFS Design Fluency · All Ages
D-KEFS Sorting Test · Ages 20-49
D-KEFS Sorting Test · Ages 50-89
D-KEFS Sorting Test · Ages 8-19
D-KEFS Sorting Test · All Ages
D-KEFS Tower Test · Ages 20-49
D-KEFS Tower Test · Ages 50-89
D-KEFS Tower Test · Ages 8-19
D-KEFS Tower Test · All Ages
D-KEFS Trail Making Test · Ages 20-49
D-KEFS Trail Making Test · Ages 50-89
D-KEFS Trail Making Test · Ages 8-19
D-KEFS Trail Making Test · All Ages
D-KEFS Verbal Fluency · Ages 20-49
D-KEFS Verbal Fluency · Ages 50-89
D-KEFS Verbal Fluency · Ages 8-19
D-KEFS Verbal Fluency · All Ages
D-KEFS Word Context Test · Ages 20-49
D-KEFS Word Context Test · Ages 50-89
D-KEFS Word Context Test · Ages 8-19
D-KEFS Word Context Test · All Ages
D-KEFS Word Proverb Test · Ages 20-49
D-KEFS Word Proverb Test · Ages 50-89
D-KEFS Word Proverb Test · Ages 8-19
D-KEFS Word Proverb Test · All Ages
RBANS Indices · Ages 12-19
RBANS Indices · Ages 20-89
RBANS Subtests · Ages 12-19
RBANS Subtests · Ages 20-89
WAIS-IV Core Subtests · Ages 16-29
WAIS-IV Core Subtests · Ages 30-54
WAIS-IV Core Subtests · Ages 55-69
WAIS-IV Core Subtests · Ages 70-90
WAIS-IV Core Subtests · All Ages
WAIS-IV Indices · Ages 16-29
WAIS-IV Indices · Ages 30-54
WAIS-IV Indices · Ages 55-69
WAIS-IV Indices · Ages 70-90
WAIS-IV Indices · All Ages
WISC-V Indices · All Ages
WISC-V Subtests · All Ages
WMS-IV Indices · Ages 16-69
WMS-IV Indices · Ages 65-90
WMS-IV Subtests · Ages 16-69
WMS-IV Subtests · Ages 65-90

Test data & patient scores

# Subtest M₁ SD₁ M₂ SD₂ r Date 1 Date 2 Ŷ₂ RCI (z) p Outcome
APA-formatted output
Enter test data and patient scores to preview the APA table.
Reliable Change Indices (RCI) · Regression-Based (Crawford)

Crawford Regression-Based Reliable Change Index

Crawford & Garthwaite (2007). Extends the standardised regression-based approach to use a t-distributed test statistic that incorporates the normative sample size (N), correctly accounting for uncertainty in the regression parameters when N is modest. Returns a sample-size-adjusted standard error of prediction.

CVLT-3 Indices · Ages 16-44
CVLT-3 Indices · Ages 45-90
CVLT-3 Trials · Ages 16-44
CVLT-3 Trials · Ages 45-90
D-KEFS Colour-Word Interference · All Ages
D-KEFS Design Fluency · All Ages
D-KEFS Sorting Test · All Ages
D-KEFS Tower Test · All Ages
D-KEFS Trail Making Test · All Ages
D-KEFS Verbal Fluency · All Ages
D-KEFS Word Context Test · All Ages
D-KEFS Word Proverb Test · All Ages
RBANS Indices · Ages 12-19
RBANS Indices · Ages 20-89
RBANS Subtests · Ages 12-19
RBANS Subtests · Ages 20-89
WAIS-IV Core Subtests · All Ages
WAIS-IV Indices · All Ages
WISC-V Indices · All Ages
WISC-V Subtests · All Ages
WMS-IV Indices · Ages 16-69
WMS-IV Indices · Ages 65-90
WMS-IV Subtests · Ages 16-69
WMS-IV Subtests · Ages 65-90

Test data & patient scores

# Subtest M₁ SD₁ M₂ SD₂ r N Date 1 Date 2 Ŷ₂ t(RB) p Outcome
APA-formatted output
Enter test data and patient scores to preview the APA table.
Premorbid · Estimation

Premorbid Estimate

Inputs

Enter whichever predictors are available. Leave unavailable fields blank; the estimate table will update only for models with enough information.

Available predictors
Demographics
Used by demographic and age-adjusted models where applicable
Shared with the patient age in the header
Output settings
Controls the confidence interval and report table title
Note. ToPF and Crawford & Allan equations use UK data. OPIE-4 uses the prorated WAIS-IV US coefficients (Holdnack et al., 2013), adapted for UK use by omitting the education, region and ethnicity terms; prorated equations avoid part-whole correlation inflation. OPIE-4 point estimates are illustrative only in UK contexts and should not be applied formally.

Sources Wechsler (2011), ToPF-UK manual · Crawford & Allan (1997) · Holdnack et al. (2013), Table eA5.8

APA-formatted output
Enter at least the ToPF raw score to generate estimates.

Enter the patient's actual WAIS-IV / WMS-IV index scores in the Achieved column to compute ToPF-predicted vs actual discrepancies. Base rates are the published figures from the ToPF-UK manual and are shown only for negative discrepancies (achieved < predicted). The manual derives them from a normal model with SD = SEE rather than from observed standardisation-sample frequencies.

Index Predicted Lower 90% Upper 90% Achieved Difference Base rate
WAIS-IV
Full Scale IQ - - - - -
Verbal Comprehension Index - - - - -
Perceptual Reasoning Index - - - - -
Working Memory Index - - - - -
Processing Speed Index - - - - -
WMS-IV
Immediate Memory Index - - - - -
Delayed Memory Index - - - - -
Visual Working Memory Index - - - - -
APA-formatted output
Enter at least one Achieved score to generate the discrepancy table.

Enter age (16–90), sex, plus Vocabulary and/or Matrix Reasoning raw scores in the Inputs panel above. Rows appear automatically for each model whose required inputs are present. Enter the patient's actual FSIQ / GAI in the Achieved column - the prorated index is calculated per ACS manual procedures, excluding the subtest(s) used as predictors. The three FSIQ rows predict three different prorated criteria and are not expected to agree with each other.

Settings · Methods & References

Methods & References

A clinical psychometric calculation tool for neuropsychological report writing. All computation is local; no patient data is ever transmitted.

Methods & conventions

What this tool does

Nine working pages: Premorbid Estimate, Score Tables, Change Analysis, Profile Analysis, Performance Validity, Score Charts, Score Converter, Effect Sizes and Data. Every calculation runs locally in the browser. No patient data is transmitted off-device, and the app works with no network connection.

The auto-fill normative database holds published parameters for seven instrument families: D-KEFS (original and Advanced), WAIS-IV, WMS-IV, WISC-V, CVLT-3, CVLT-C and the RBANS, with the retest sample size N where the publisher reports one. N is required for the Crawford & Garthwaite method and may need entering by hand where it is unavailable. Clinicians should verify every imported parameter against the current manual, and against local service standards, before interpreting it.

Score conversion and classification

Conversions between standard (M 100, SD 15), T (50, 10), scaled (10, 3) and z scores assume an approximately normal reference distribution. Two descriptor schemes are offered, and the one in force is named in the note beneath every exported table: Wechsler bands follow the WAIS-IV/WMS-IV manual conventions, and AACN labels follow Guilmette et al. (2020). Confidence levels throughout are 90% (z = 1.645) and 95% (z = 1.960); intervals round the estimate and the margin separately, so the printed bounds stay symmetric about the printed value.

Confidence intervals on Score Tables

Confidence intervals and standard errors of measurement. The CI column is the obtained score ± z × SEM, where SEM = SD × √(1 − r), centred on the obtained score rather than on an estimated true score.

The standard deviation is the normative SD of the metric the score is reported in (15, 10, 3 or 1), because a coefficient computed on, or corrected to, the normative sample must be paired with that sample's variability. Where a measure's stored statistics are raw, its own standard deviation is used instead, that being the only one in the right units. Four publishers state that rule outright, and the arithmetic confirms it: this pairing reproduces every published standard error of measurement the app is able to check, exactly, at the precision each is printed to: all 300 cells of WAIS-IV Table 4.3, 242 of WISC-V Table 4.4, 240 of WMS-IV Table 3.3, 168 across the D-KEFS SEM tables, 126 of RBANS Update Table 3.7, and all 38 CVLT-3 measures in Tables 3.4 and 3.5.

The reliability is, by default, the retest coefficient held in the normative database (an alternate-form coefficient in the case of the CVLT-3, which publishes no same-form retest), corrected for the normative sample's variability where the publisher reports a corrected value. Retest is the default for two reasons: it keeps a single, stated basis across a table that may mix batteries, and it is the appropriate coefficient for the many timed measures in the database, since split-half and alpha are not valid reliability estimates for speeded tests. The WAIS-IV manual makes that second argument itself for Coding, Symbol Search and Cancellation, describing the split-half coefficient as "not a proper reliability estimate" for a Processing Speed subtest; the values used here for those three are the ones it publishes, in all 38 of the cells its Table 4.1 gives them.

That default is set aside for a measure only where its publisher both reports an internal-consistency coefficient and derives its own published intervals from it. Seven manuals meet that bar:

InstrumentCoefficient usedSourceNot applied to
CVLT-C Odd–even split-half, by age Manual Table 6.5 Every index but List A Trials 1–5 Total; item scores on a word-list task are not independent. The interval printed in the manual's own worked example reproduces exactly.
D-KEFS Internal consistency, by normative age band Technical Manual, Tables 2.1–2.24 Colour–Word Interference, whose only coefficient table is for a composite this app does not hold; Design Fluency, where item interdependence precluded the procedure; and five of the six Trail Making measures, the published table covering the composite alone.
D-KEFS Advanced Split-half, by normative age band Table 3.4 Trail Making and Verbal Fluency, which that manual treats as speeded and scores on stability coefficients.
WAIS-IV Split-half or alpha, by normative age band Table 4.1 Coding, Symbol Search and Cancellation, which are speeded. These keep the corrected stability coefficient the same table publishes for them.
WISC-V Split-half, by single year of age Table 4.1 Coding, Symbol Search and Cancellation, together with the Cancellation Random and Cancellation Structured process scores, all speeded and likewise on the corrected stability coefficient.
WMS-IV Split-half or alpha, by normative age band, Adult and Older Adult batteries separately Table 3.1 Verbal Paired Associates II Word Recall, a free-recall score with no consistent item count, which takes a stability coefficient. The recognition memory measures are absent altogether: their published reliability is a decision-consistency percentage, not a correlation, and cannot enter a standard error of measurement.
RBANS Update Internal consistency, by normative age band Table 3.6 Figure Copy, Semantic Fluency, Coding, Story Recall and Figure Recall, which that table itself marks as estimated from test–retest and which therefore keep a stability coefficient taken from the same table. The four subtests reported as raw scores appear nowhere in it, the manual publishing reliability for its eight scaled subtests only, so no interval is shown for them.

Which coefficient is right is a question for each manual rather than a policy of this tool, and the manuals genuinely disagree: the two D-KEFS manuals reach opposite conclusions about the same two test names. Each is followed as written, and the Data page names the basis actually in force for every measure in the database.

Age. Where a coefficient is tabulated by age, the interval uses the band for the patient's age, and the age used is named in the note beneath the table. Entering an age is optional. If none is entered, or the age falls outside a measure's normed range, the publisher's all-ages figure is used instead: the published average where a manual prints one, and otherwise the total-sample retest coefficient, which for the D-KEFS is that manual's own second regime rather than a substitute for a missing number. Both paths are therefore the publisher's own figures.

Reliable-change analysis is unaffected by any of the above and always uses the retest coefficient, which measures a different thing.

One consequence is worth bearing in mind when comparing output against a test manual. For the measures that remain on the retest default, where a manual derives its published intervals from internal-consistency reliability (almost always the higher of the two coefficients), the intervals shown here run wider than the manual's. They are therefore the more conservative, and answer the question how much would this score be expected to move on retesting rather than how precisely was it measured on the day. Not every publisher offers that comparison: the CVLT-3 manual declines to report internal-consistency reliability at all, on the grounds that item scores on a word-list task are not independent (recalling one word alters the probability of recalling the others, both within a trial and on later ones), and reports alternate-form coefficients in their place.

Measures with no normative-sample coefficient, and the reliability control

Every interval here multiplies a normative standard deviation by a reliability, and that is only a valid standard error of measurement when the two describe the same group. Most manuals supply a coefficient computed on, or corrected to, their normative sample. Three do not: the D-KEFS, D-KEFS Advanced and CVLT-C manuals report only the correlation observed in their own retest studies, a few dozen people each, and pair it with the normative standard deviation regardless.

The D-KEFS manual states that outright, fixing the standard deviation unit at 3 for all its scaled scores and deriving its test–retest standard errors of measurement "from the total sample of cases". Its Table 2.8 shows the arithmetic: the three Design Fluency all-ages values of 1.94, 1.97 and 2.47 are exactly 3 × √(1 − r) on the uncorrected coefficient. Those measures are therefore scored the way their own manuals score them, and the intervals shown reproduce the published ones. The Data page labels each such measure retest, uncorrected, so which rows rest on that footing can be read off rather than inferred.

The statistical objection to the pairing is nonetheless real, so Score Tables offers a reliability control with two settings. Published, the default, uses each manual's own coefficient and reproduces its printed interval. Corrected applies the standard range-restriction correction of Allen and Yen (1979), rxx = 1 − (s²retest ÷ s²norm)(1 − r), to those measures alone, so that the coefficient describes the same population as the standard deviation it multiplies. A published coefficient is never overwritten in either setting.

The correction is not a guess: across the 267 database entries carrying both an observed and a publisher-corrected coefficient, it reproduces the publisher's own value to a median error of .003. But the resulting figures are not printed in the manuals concerned, which is why the default is Published and why the note beneath a corrected table says so. In practice the control moves 46 of the measures reachable from Score Tables, every one of them D-KEFS or D-KEFS Advanced: at the 95% level 9 intervals widen and 4 narrow, 33 are unchanged after rounding, and the largest single change is 2 scaled-score points. Reliable-change analysis is not affected by the control. Where a corrected reading is taken, the Data page follows it rather than continuing to show the published one.

Change analysis

Five methods, in ascending order of what they model. The Standard Deviation Index is descriptive only: SD Δ = (X₂ − X₁) ÷ SD, with no reliability correction. Simple Reliable Change (Jacobson & Truax, 1991) tests the observed change against measurement error. Practice Effect-Adjusted change (Iverson, 2001) subtracts the mean retest gain observed in the normative sample first. McSweeny Regression-Based change (McSweeny et al., 1993) predicts the retest score from baseline and standardises the residual against the standard error of estimate. Crawford & Garthwaite Regression-Based change (2007) does the same but with a standard error of prediction that accounts for the normative sample size and for the distance of the baseline score from the normative mean.

All p-values are two-tailed. The first four methods use the standard normal distribution; Crawford & Garthwaite uses the Student t distribution with N − 2 degrees of freedom, so a small normative sample raises the threshold: at N = 25 the 95% critical value is 2.069 against 1.960 for z.

These calculations use the retest coefficient paired with the standard deviation of the same retest sample, so that both terms describe one population. Where a publisher reports a coefficient corrected to the normative sample's variability, that value is offered as an option but is not the default, because it describes a differently distributed population from the standard deviation it would be multiplied by, and in the two regression methods the coefficient is a fitted slope, so substituting it changes the predicted score rather than only the interval. Reliability type varies by instrument and is stated in the note beneath each generated table: CVLT-3 coefficients are alternate-form, RBANS Form A coefficients are same-form retest.

Outcomes are reported as significance only ("Reliable change" or "No reliable change"), never as improvement or decline. The database holds many measures on which a higher score is the worse result (intrusions, perseverations, errors, false positives), and carries no score-direction flag, so reading a clinical direction off the sign of the statistic would assert the wrong conclusion for all of them. The signed statistic is displayed alongside, so the direction stays visible without the app interpreting it.

Profile analysis

A battery raises a question no single row answers: this patient has two scores below the 5th percentile, so how unusual is that? By definition 5% of the population falls below the 5th percentile on any one measure, but across several correlated measures a low score somewhere is common. Reading rows independently overcalls impairment: 13.7% of healthy adults show at least one abnormally low score across the four WAIS-IV Indices, and 25.0% across ten core subtests. The Profile Analysis page reports the count and the percentage of the healthy population showing that many or more (Crawford, Garthwaite & Gault, 2007).

Three questions are answered from one simulation: how many scores are abnormally low, how many pairwise differences between measures are abnormally large, and how many measures deviate abnormally from the patient’s own mean across the battery. The criterion is chosen by the clinician from the five the paper tabulates (below 1 SD, or below the 10th, 5th, 2nd or 1st percentile) and governs all three. A difference or deviation is abnormal when larger than the same percentage of the population shows, in either direction. Simulations are seeded, so a figure quoted in a report does not change when the page is reopened, and the sampling error of each percentage is stated beside it.

The only input is the correlation matrix over the measures administered, taken from the publisher’s normative sample: WAIS-IV Technical and Interpretive Manual (GB) Table 5.1, or WMS-IV Technical and Interpretive Manual (GB) Table 4.1 (Adult battery) or 4.2 (Older Adult battery). WAIS-IV and the WMS-IV Adult battery can also be profiled together, on WMS-IV Table 4.12, which correlates the two batteries in their co-normed sample (corrected to the WMS-IV normative sample’s variability); the Older Adult battery cannot, as that table pools both batteries under one Visual Memory Index. RBANS Update (Form A) index and subtest scores are profiled on RBANS Update Manual Table 4.1 (Randolph, 2012), and WISC-V scores on WISC-V Technical and Interpretive Manual Table 5.1. Scores are read from Score Tables rather than entered here, so the two cannot disagree. A profile may not hold both a measure and a part of itself (an index and its own subtest, two indices sharing subtests, or a subtest and its own process scores), so each instrument is profiled one level at a time. Where a publisher’s indices cross, as WMS-IV’s do (auditory and visual against immediate and delayed), the two readings are offered separately. Index scores give the more accurate estimate: the authors note that scaled scores are coarse, one point being a third of a standard deviation, so a subtest-level profile is the rougher reading and says so on screen and in the exported note.

Performance validity

Eight measures, each scored against a published cut-off or base rate. Two are embedded in the RBANS: the Effort Index (Silverberg et al., 2007; accuracy per Shura et al., 2018) and the Effort Scale (Novitski et al., 2012). Two are embedded in the WAIS: Reliable Digit Span (Greiffenstein et al., 1994; cut-offs per Schroeder et al., 2012) and the Digit Span indices (Iverson & Tulsky, 2003; Axelrod et al., 2006). One is embedded in the CVLT-3: the Forced Choice trial, read against the manual’s age-banded base rates (Delis et al., 2017), with CVLT-II cut-offs selectable (Schwartz et al., 2016; Erdodi, Abeare, et al., 2018). One is embedded in the D-KEFS: Trail Making, five age-corrected conditions each read against its own cut-off, failing when at least a chosen number fail (Erdodi, Hurtubise, et al., 2018). Two are stand-alone tests: the Rey 15-Item with recognition trial (Boone et al., 2002) and the TOMM (Tombaugh, 1996; cut-off accuracy per Martin et al., 2020). Every result is a comparison against the published figure only, reported with the published sensitivity and specificity where the source gives a pair; the app issues no verdict on the protocol, because no single indicator supports one.

Aggregation follows Larrabee (2014a): failing two or more independent validity indicators supports probable invalidity. Independence is the load-bearing word: measures derived from the same administration of the same instrument share error and cannot be counted twice, so the two RBANS indices count as one indicator between them, as do the two Digit Span measures. A base-rate index with no published accuracy (the Vocabulary − Digit Span difference, the longest spans, the CVLT-3 critical items) is reported as flagged but not counted as a failure. The Summary tab reports Larrabee's published classification accuracy at each failure count. Positive and negative predictive power appear on the TOMM tab, derived by Bayes' theorem from the meta-analytic sensitivity and specificity and a base rate the clinician selects (Martin et al., 2020, Tables 16–17), a selectable rate rather than a fixed one, because the base rate of invalid performance differs by setting and the choice belongs to the clinician, not the app.

Embedded indices are computed from subtests that also measure genuine ability, so each carries a caution naming the populations in which it over-flags (dementia and severe impairment chief among them), and the caution takes emphasis only once that measure has actually flagged. Cut-offs validated in one edition of an instrument are labelled with that edition, and the clinician should document which edition was administered.

Premorbid estimation

Premorbid estimates combine ToPF-based and demographic equations with OPIE-4 prorated models, and produce predicted-versus-achieved discrepancy output with confidence intervals, base-rate lookups and APA-formatted export tables. Predicted-versus-achieved significance flagging uses a three-tier scheme at z = 1.645 (*), 1.960 (**) and 2.576 (***); this is separate from the 90%/95% confidence-interval selector.

OPIE-4 is provided for illustration only in a UK context and its output should not be quoted as a concrete premorbid estimate. The regression terms reproduce Holdnack et al. (2013), Table eA5.8, but the published equations also carry US education, ethnicity and region terms that are not applied here, which fixes every prediction at the US reference category (12th-grade high-school graduate, not African-American, not resident in the western US). Those terms are omitted rather than mapped because the education dummies encode how unusual a given attainment level is within the US population the model was fitted on, not years of schooling, and that does not transfer: the US reference category corresponds to A-levels if matched by years but to GCSE/O-level if matched by population position, and UK school-leaving age was raised to 16 only in 1972, so leaving school without qualifications was normative for older cohorts in a way it was not in the US sample. Expect estimates to run high for patients who left school early and low for graduates, by an amount this tool cannot quantify. For a UK demographic estimate, use the Crawford & Allan (1997) model.

Base rates. The ToPF predicted-difference base rates are the published figures from the ToPF-UK manual (Wechsler, 2011) and are used as published. The manual derives them from a normal model with SD equal to the model's standard error of estimate, rather than tabulating observed standardisation-sample frequencies, and they are labelled as such wherever they appear: every published cell equals the normal-curve proportion below that discrepancy, exactly, at the printed precision. A predicted-difference table is necessarily built this way, a standardisation sample of about a thousand cases yielding no observed frequency at every discrepancy point. The OPIE-4 discrepancy base rates of ACS Table eA5.12, by contrast, are empirical, sitting on a count grid, and are likewise used as published. The two published tables therefore answer the same question by different methods, and differ by roughly 10% relatively across the decisive −5 to −20 band on models of almost identical standard error of estimate: a discrepancy of −15 gives 3.78% on the ToPF table against about 4.3% on the OPIE-4 one. That difference is a property of the two sources, not an adjustment made here; neither table is modified.

References
Test manuals and technical sources

Delis, D. C., & Kaplan, E. (2025). Delis-Kaplan Executive Function System Advanced: Manual. NCS Pearson. [Split-half reliability by normative age band: Table 3.4.]

Delis, D. C., Kaplan, E., & Kramer, J. H. (2001). Delis–Kaplan Executive Function System (D-KEFS): Technical manual. San Antonio, TX: The Psychological Corporation. [Internal-consistency coefficients and standard errors of measurement by age band, chapter 2 and Tables 2.1–2.26.]

Delis, D. C., Kramer, J. H., Kaplan, E., & Ober, B. A. (1994). California Verbal Learning Test – Children's Version (CVLT-C): Manual. San Antonio, TX: The Psychological Corporation. [Split-half reliability and standard errors of measurement: Table 6.5. Standardised score equivalents: Tables A.1 and A.2.]

Delis, D. C., Kramer, J. H., Kaplan, E., & Ober, B. A. (2017). California Verbal Learning Test – Third Edition (CVLT-3): Manual. Bloomington, MN: Pearson. [Alternate-form reliability and standard errors of measurement: Tables 3.4 and 3.5.]

Holdnack, J. A., Drozdick, L., Weiss, L. G., & Iverson, G. L. (2013). WAIS-IV, WMS-IV, and ACS: Advanced clinical interpretation. Oxford: Academic Press. [OPIE-4 prorated regression coefficients: Table eA5.8. OPIE-4 discrepancy base rates: Table eA5.12.]

Randolph, C. (2012). Repeatable Battery for the Assessment of Neuropsychological Status Update (RBANS Update): Manual. Bloomington, MN: Pearson. [Reliability by age band: Table 3.6. Standard errors of measurement: Table 3.7. Test–retest and alternate-form data: Tables 3.8–3.9.]

Wechsler, D. (2010). Wechsler Adult Intelligence Scale – Fourth UK Edition (WAIS–IVUK): Administration and scoring manual. London: Pearson Assessment. [Longest-span base rates: Tables C.4 and C.5.]

Wechsler, D. (2010). Wechsler Adult Intelligence Scale – Fourth UK Edition (WAIS–IVUK): Technical and interpretive manual. London: Pearson Assessment. [Reliability coefficients: Table 4.1. Standard errors of measurement: Table 4.3. Test–retest stability parameters: Table 4.5.]

Wechsler, D. (2010). Wechsler Memory Scale – Fourth UK Edition (WMS–IVUK): Technical and interpretive manual. London: Pearson Assessment. [Reliability coefficients: Table 3.1. Standard errors of measurement: Table 3.3.]

Wechsler, D. (2011). Test of Premorbid Functioning (ToPF-UK): Manual. London: Pearson Assessment. [ToPF-predicted vs obtained discrepancy base rates for the WAIS-IV and WMS-IV indices, derived by the publisher from a normal model on the standard error of estimate.]

Wechsler, D. (2014). Wechsler Intelligence Scale for Children – Fifth Edition (WISC-V): Technical and interpretive manual. Bloomington, MN: NCS Pearson. [Reliability coefficients: Table 4.1. Standard errors of measurement: Table 4.4. Test–retest stability: Table 4.7.]

Statistical and methodological sources

Allen, M. J., & Yen, W. M. (1979). Introduction to measurement theory. Monterey, CA: Brooks/Cole. [Correction of a reliability coefficient for range restriction in the sample it was observed in.]

Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Hillsdale, NJ: Lawrence Erlbaum.

Crawford, J. R., & Allan, K. M. (1997). Estimating premorbid WAIS–R IQ with demographic variables: Regression equations derived from a UK sample. The Clinical Neuropsychologist, 11(2), 192–197. https://doi.org/10.1080/13854049708407050

Crawford, J. R., & Garthwaite, P. H. (2007). Using regression equations built from summary data in the neuropsychological assessment of the individual case. Neuropsychology, 21(5), 611–620. https://doi.org/10.1037/0894-4105.21.5.611

Crawford, J. R., Garthwaite, P. H., & Gault, C. B. (2007). Estimating the percentage of the population with abnormally low scores (or abnormally large score differences) on standardized neuropsychological test batteries: A generic method with applications. Neuropsychology, 21(4), 419–430. https://doi.org/10.1037/0894-4105.21.4.419

Guilmette, T. J., Sweet, J. J., Hebben, N., Koltai, D., Mahone, E. M., Spiegler, B. J., Stucky, K., Westerveld, M., & Conference Participants. (2020). American Academy of Clinical Neuropsychology consensus conference statement on uniform labeling of performance test scores. The Clinical Neuropsychologist, 34(3), 437–453. https://doi.org/10.1080/13854046.2020.1722244

Iverson, G. L. (2001). Interpreting change on the WAIS-III/WMS-III in clinical samples. Archives of Clinical Neuropsychology, 16(2), 183–191. https://doi.org/10.1093/arclin/16.2.183

Jacobson, N. S., & Truax, P. (1991). Clinical significance: A statistical approach to defining meaningful change in psychotherapy research. Journal of Consulting and Clinical Psychology, 59(1), 12–19. https://doi.org/10.1037/0022-006X.59.1.12

McSweeny, A. J., Naugle, R. I., Chelune, G. J., & Lüders, H. (1993). "T scores for change": An illustration of a regression approach to depicting change in clinical neuropsychology. The Clinical Neuropsychologist, 7(3), 300–312. https://doi.org/10.1080/13854049308401901

Sawilowsky, S. S. (2009). New effect size rules of thumb. Journal of Modern Applied Statistical Methods, 8(2), 597–599. https://doi.org/10.22237/jmasm/1257035100

Performance validity sources

Axelrod, B. N., Fichtenberg, N. L., Millis, S. R., & Wertheimer, J. C. (2006). Detecting incomplete effort with Digit Span from the Wechsler Adult Intelligence Scale–Third Edition. The Clinical Neuropsychologist, 20(3), 513–523. https://doi.org/10.1080/13854040590967117

Bilder, R. M., Sugar, C. A., & Hellemann, G. S. (2014). Cumulative false positive rates given multiple performance validity tests: Commentary on Davis and Millis (2014) and Larrabee (2014). The Clinical Neuropsychologist, 28(8), 1212–1223. https://doi.org/10.1080/13854046.2014.969774

Boone, K. B., Salazar, X., Lu, P., Warner-Chacon, K., & Razani, J. (2002). The Rey 15-Item recognition trial: A technique to enhance sensitivity of the Rey 15-Item Memorization Test. Journal of Clinical and Experimental Neuropsychology, 24(5), 561–573. https://doi.org/10.1076/jcen.24.5.561.1004

Denning, J. H. (2012). The efficiency and accuracy of the Test of Memory Malingering trial 1, errors on the first 10 items of the Test of Memory Malingering, and five embedded measures in predicting invalid test performance. Archives of Clinical Neuropsychology, 27(4), 417–432. https://doi.org/10.1093/arclin/acs044

Erdodi, L. A., Abeare, C. A., Medoff, B., Seke, K. R., Sagar, S., & Kirsch, N. L. (2018). A single error is one too many: The Forced Choice Recognition trial of the CVLT-II as a measure of performance validity in adults with TBI. Archives of Clinical Neuropsychology, 33(7), 845–860. https://doi.org/10.1093/acn/acx110 [The CVLT-II ≤ 15 cut-off and its accuracy across seven reference PVTs.]

Erdodi, L. A., Hurtubise, J. L., Charron, C., Dunn, A., Enache, A., McDermott, A., & Hirst, R. B. (2018). The D-KEFS Trails as performance validity tests. Psychological Assessment, 30(8), 1082–1095. https://doi.org/10.1037/pas0000561

Greiffenstein, M. F., Baker, W. J., & Gola, T. (1994). Validation of malingered amnesia measures with a large clinical sample. Psychological Assessment, 6(3), 218–224. https://doi.org/10.1037/1040-3590.6.3.218

Iverson, G. L., & Tulsky, D. S. (2003). Detecting malingering on the WAIS-III: Unusual Digit Span performance patterns in the normal population and in clinical groups. Archives of Clinical Neuropsychology, 18(1), 1–9. https://doi.org/10.1093/arclin/18.1.1

Larrabee, G. J. (2014a). False-positive rates associated with the use of multiple performance and symptom validity tests. Archives of Clinical Neuropsychology, 29(4), 364–373. https://doi.org/10.1093/arclin/acu019

Larrabee, G. J. (2014b). Minimizing false positive error with multiple performance validity tests: Response to Bilder, Sugar, and Hellemann (2014). The Clinical Neuropsychologist, 28(8), 1230–1242. https://doi.org/10.1080/13854046.2014.988754

Martin, P. K., Schroeder, R. W., Olsen, D. H., Maloy, H., Boettcher, A., Ernst, N., & Okut, H. (2020). A systematic review and meta-analysis of the Test of Memory Malingering in adults: Two decades of deception detection. The Clinical Neuropsychologist, 34(1), 88–119. https://doi.org/10.1080/13854046.2019.1637027

Meyers, J. E., & Volbrecht, M. (1998). Validation of reliable digits for detection of malingering. Assessment, 5(3), 303–307. https://doi.org/10.1177/107319119800500309

Novitski, J., Steele, S., Karantzoulis, S., & Randolph, C. (2012). The Repeatable Battery for the Assessment of Neuropsychological Status Effort Scale. Archives of Clinical Neuropsychology, 27(2), 190–195. https://doi.org/10.1093/arclin/acr119

Schroeder, R. W., Twumasi-Ankrah, P., Baade, L. E., & Marshall, P. S. (2012). Reliable Digit Span: A systematic review and cross-validation study. Assessment, 19(1), 21–30. https://doi.org/10.1177/1073191111428764

Schwartz, E. S., Erdodi, L., Rodriguez, N., Ghosh, J. J., Curtain, J. R., Flashman, L. A., & Roth, R. M. (2016). CVLT-II Forced Choice Recognition trial as an embedded validity indicator: A systematic review of the evidence. Journal of the International Neuropsychological Society, 22(8), 851–858. https://doi.org/10.1017/s1355617716000746 [The CVLT-II ≤ 14 cut-off; classification accuracy and the impairment gradient in failure rates.]

Sherman, E. M. S., Slick, D. J., & Iverson, G. L. (2020). Multidimensional malingering criteria for neuropsychological assessment: A 20-year update of the malingered neuropsychological dysfunction criteria. Archives of Clinical Neuropsychology, 35(6), 735–764. https://doi.org/10.1093/arclin/acaa019

Shura, R. D., Brearly, T. W., Rowland, J. A., Martindale, S. L., Miskey, H. M., & Duff, K. (2018). RBANS validity indices: A systematic review and meta-analysis. Neuropsychology Review, 28(3), 269–284. https://doi.org/10.1007/s11065-018-9377-5

Silverberg, N. D., Wertheimer, J. C., & Fichtenberg, N. L. (2007). An effort index for the Repeatable Battery for the Assessment of Neuropsychological Status (RBANS). The Clinical Neuropsychologist, 21(5), 841–854. https://doi.org/10.1080/13854040600850958

Slick, D. J., Sherman, E. M. S., & Iverson, G. L. (1999). Diagnostic criteria for malingered neurocognitive dysfunction: Proposed standards for clinical practice and research. The Clinical Neuropsychologist, 13(4), 545–561. https://doi.org/10.1076/1385-4046(199911)13:04;1-y;ft545

Sweet, J. J., Heilbronner, R. L., Morgan, J. E., Larrabee, G. J., Rohling, M. L., Boone, K. B., Kirkwood, M. W., Schroeder, R. W., Suhr, J. A., & Conference Participants. (2021). American Academy of Clinical Neuropsychology (AACN) 2021 consensus statement on validity assessment: Update of the 2009 AACN consensus conference statement on neuropsychological assessment of effort, response bias, and malingering. The Clinical Neuropsychologist, 35(6), 1053–1106. https://doi.org/10.1080/13854046.2021.1896036

Tombaugh, T. N. (1996). Test of Memory Malingering (TOMM). North Tonawanda, NY: Multi-Health Systems.

Effect-size benchmarks

Cipriani, A., Furukawa, T. A., Salanti, G., Chaimani, A., Atkinson, L. Z., Ogawa, Y., Leucht, S., Ruhe, H. G., Turner, E. H., Higgins, J. P. T., Egger, M., Takeshima, N., Hayasaka, Y., Imai, H., Shinohara, K., Tajika, A., Ioannidis, J. P. A., & Geddes, J. R. (2018). Comparative efficacy and acceptability of 21 antidepressant drugs for the acute treatment of adults with major depressive disorder: A systematic review and network meta-analysis. The Lancet, 391(10128), 1357–1366. https://doi.org/10.1016/S0140-6736(17)32802-7

Cuijpers, P., Geraedts, A. S., van Oppen, P., Andersson, G., Markowitz, J. C., & van Straten, A. (2011). Interpersonal psychotherapy for depression: A meta-analysis. American Journal of Psychiatry, 168(6), 581–592. https://doi.org/10.1176/appi.ajp.2010.10101411

Cuijpers, P., Miguel, C., Harrer, M., Plessen, C. Y., Ciharova, M., Ebert, D., & Karyotaki, E. (2023). Cognitive behavior therapy vs. control conditions, other psychotherapies, pharmacotherapies and combined treatment for depression: A comprehensive meta-analysis including 409 trials with 52,702 patients. World Psychiatry, 22(1), 105–115. https://doi.org/10.1002/wps.21069

Hofmann, S. G., & Smits, J. A. J. (2008). Cognitive-behavioral therapy for adult anxiety disorders: A meta-analysis of randomized placebo-controlled trials. The Journal of Clinical Psychiatry, 69(4), 621–632. https://doi.org/10.4088/JCP.v69n0415

Huhn, M., Nikolakopoulou, A., Schneider-Thoma, J., Krause, M., Samara, M., Peter, N., Arndt, T., Bäckers, L., Rothe, P., Cipriani, A., Davis, J., Salanti, G., & Leucht, S. (2019). Comparative efficacy and tolerability of 32 oral antipsychotics for the acute treatment of adults with multi-episode schizophrenia: A systematic review and network meta-analysis. The Lancet, 394(10202), 939–951. https://doi.org/10.1016/S0140-6736(19)31135-3

National Centre for Social Research, University College London, Department of Epidemiology and Public Health. (2025). Health Survey for England, 2022 [Data collection]. UK Data Service. SN: 9469. https://doi.org/10.5255/UKDA-SN-9469-1

Pesch, B., Kendzia, B., Gustavsson, P., Jöckel, K.-H., Johnen, G., Pohlabeln, H., Olsson, A., Ahrens, W., Gross, I. M., Brüske, I., Wichmann, H.-E., Merletti, F., Richiardi, L., Simonato, L., Fortes, C., Siemiatycki, J., Parent, M.-E., Consonni, D., Landi, M. T., … Brüning, T. (2012). Cigarette smoking and lung cancer: Relative risk estimates for the major histological types from a pooled analysis of case–control studies. International Journal of Cancer, 131(5), 1210–1219. https://doi.org/10.1002/ijc.27339

Storebø, O. J., Storm, M. R. O., Pereira Ribeiro, J., Skoog, M., Groth, C., Callesen, H. E., Schaug, J. P., Darling Rasmussen, P., Huus, C.-M. L., Zwi, M., Kirubakaran, R., Simonsen, E., & Gluud, C. (2023). Methylphenidate for children and adolescents with attention deficit hyperactivity disorder (ADHD). Cochrane Database of Systematic Reviews, 2023(3), Article CD009885. https://doi.org/10.1002/14651858.CD009885.pub3

Watts, B. V., Schnurr, P. P., Mayo, L., Young-Xu, Y., Weeks, W. B., & Friedman, M. J. (2013). Meta-analysis of the efficacy of treatments for posttraumatic stress disorder. The Journal of Clinical Psychiatry, 74(6), e541–e550. https://doi.org/10.4088/JCP.12r08225

Williams, A. C. de C., Fisher, E., Hearn, L., & Eccleston, C. (2020). Psychological therapies for the management of chronic pain (excluding headache) in adults. Cochrane Database of Systematic Reviews, 2020(8), Article CD007407. https://doi.org/10.1002/14651858.CD007407.pub4

Wolraich, M. L., Wilson, D. B., & White, J. W. (1995). The effect of sugar on behavior or cognition in children: A meta-analysis. JAMA, 274(20), 1617–1621. https://doi.org/10.1001/jama.1995.03530200053037

Settings · Privacy & use

Privacy & use

What happens to the data you enter, and what you are agreeing to when you use these calculators.

Where your data goes

Nothing you enter is uploaded

Scores, tables, patient age and everything in the Working Report stay in this browser. The application makes no network requests of its own: it contains no analytics, no tracking, no error reporting and no server to send anything to. Every calculation runs on this device, and the whole suite works with no network connection at all.

The one external request is to Google Fonts, which loads the two typefaces the interface is set in. It carries no patient data. Beyond that, the only thing stored anywhere is stored here, on this computer.

Patient data lasts only as long as the tab

Entered scores and the Working Report are held in session storage, which survives a reload of the same tab and nothing else. Closing the tab, the window or the browser clears them. The app writes nothing patient-identifiable to disk of its own accord, and an older version of this app that did keep reports on disk has its leftovers deleted the first time this version loads.

Session storage is per-tab, so two tabs are two independent patient sessions rather than one shared report. Within a tab, use New patient in the top bar between patients: it clears every table and the age together, which is why it is one control rather than two.

To carry a patient over to another sitting, use Save session in the top bar. It writes a file to your computer holding everything entered: scores, patient age and page settings. That file is patient data: store it securely, as you would any clinical record, under your own data-protection obligations (for example UK GDPR). Open session reads it back and recalculates every table from those inputs.

Two things do persist here deliberately, and neither is patient data: your view preferences, such as drawer width and whether the report drawer starts minimised, and any custom test norms you add on the Data page, which would be worth little if they vanished with the tab.

Responsibility for use

This is a calculator, not a clinical judgement

The suite computes and formats. It is decision support: it does not decide anything, no result is a diagnosis on its own, and it cannot tell whether the method you have chosen suits the question you are asking. Every choice of method, norm and interpretation is your own clinical judgement.

Know the method before you use it

Use each tool only with knowledge of, and experience in, the method behind it. Before using a tool in clinical practice, consult its source manuals and papers and make sure your understanding of the method is adequate. Every method used here is cited, and the full reference list is on .

Norms and coefficients have limits

Norms, reliabilities and regression coefficients are transcribed from the cited editions. Several were standardised on US samples, and any of them may be superseded by a later edition or correction. Check that the source applies to your patient and is current.

You are responsible for the inputs and the method

Every result is only as sound as the norms, reliabilities and scores entered to produce it. Verify imported parameters against the current published manual before interpreting them, and satisfy yourself that the method is the right one for the case. Where a figure is a compromise the app says so rather than hiding it: the reliability basis is named in the note under every exported table, and the Data page labels each coefficient with what it actually rests on.

You are responsible for being qualified to use it

These instruments are restricted measures. By using this tool you confirm that you hold the training, registration and professional competence required to administer, score and interpret them, and that you are working within your scope of practice and your local service standards. The tool does not check this and cannot.

Read the note that travels with the table

Every exported table carries an APA note stating the method used, the reliability basis in force and the age band applied. That note is the record of how the numbers were produced, and it is written to be read by whoever receives the report. Specific cautions are flagged where they apply: OPIE-4 is labelled illustrative-only for UK use because the published equations carry US education, ethnicity and region terms that are not applied here, and the discrepancy base rates on the premorbid page are a parametric normal model rather than observed frequencies.

Reports produced here are your responsibility

Output from this tool may end up in clinical records and in medico-legal reports that are scrutinised. Check the numbers before they leave. The tools are provided as they are, without warranty that any calculation is correct or fit for a particular purpose. The authors accept no liability for errors or omissions or for how the output is used, and responsibility for anything written on the basis of it is yours. Word and CSV exports carry a line saying so.

Not a medical device

This software has not been assessed or registered as a medical device. It does not replace the publishers' manuals or scoring software, and it does not license the use of any test.

Your agreement

Before first use, each browser asks you to agree to these terms. If the terms change, you are asked again.

✓ Table copied - ready to paste into your report