Resources / 5: Assessment / Measuring Change and Assessment Technology

Maya times the bunny on a red and white block puzzle

Measuring Change and Assessment Technology

5: Assessment

Study guide by Anders Chan, PsyD · Updated

Measuring Change and Assessment Technology: Is the Change Real, and Can You Trust the Screen?

Two exam areas live here: how you know a client, couple, group, or program really changed, and what happens to a test delivered by computer, video, or phone. Both need one habit: separate signal from noise before you believe a number.

Why This Matters for Psychologists

A client's depression score drops from 34 to 24 after eight sessions. Her insurer wants proof that therapy works. Your supervisor asks if the drop is real. She asks if she's "normal" yet. Three questions, three different tools. The EPPP tests whether you can match them.

Three Things That Move a Score Besides Treatment

Before you credit therapy, rule out the usual suspects.

Measurement Error

No test is perfectly reliable, so every score carries error. The standard error of measurement (SEM) sizes it: SEM = SD × √(1 − reliability). The formula and confidence bands are covered in the Item Analysis and Test Reliability lesson. For change, one point matters: the gap between two scores has its own, larger error, because both scores carry error. Classic exam point: change (difference) scores tend to be unreliable, and more so when the two scores correlate highly (Clayson et al., 2021; May & Hittner, 2003).

It's like a bathroom scale that drifts a pound either way. Weigh yourself twice, and the difference between the readings wobbles even more than one reading does.

Regression to the Mean

Regression to the mean (RTM) can make natural ups and downs look like real change. Unusually high or low scores tend to be followed by scores closer to the average. RTM grows when measurement error is high and when people are picked because of an extreme starting score (Barnett et al., 2005). Research design calls this statistical regression (covered in the Research: Internal/External Validity lesson).

People often seek help during their worst week. Retest them a month later, and part of the drop would have happened even without treatment.

Practice Effects

Practice effects are score gains from having seen the test before: memory for items, learned strategies, or test sophistication (Calamia et al., 2012). They are a big issue on cognitive and neuropsychological retests. Across 50 studies of cognitive ability tests, retesting raised scores by about a quarter of an SD (d = .26), and gains were larger with coaching and identical forms (Hausknecht et al., 2007). Their size also depends on age, diagnosis, and the retest interval (Calamia et al., 2012). They can appear even when retests are more than six months apart, and alternate forms do not guarantee they disappear (Duff, 2012).

The second time through a dance routine, you hit the moves faster. You didn't become more coordinated overnight. You remember the eight-count.

Noise sourceWhat it looks likeCommon fix
Measurement errorScores bounce with no real changeReliable change index
Regression to the meanExtreme scorers drift toward averageComparison group; don't select on one extreme score
Practice effectsGains on cognitive retests, bigger with the same formAlternate forms (partial fix); practice-adjusted change methods

Neuropsychologists also use practice-adjusted RCIs and regression-based methods that predict the retest score from the first score (Duff, 2012).

The Reliable Change Index: Is the Change Bigger Than the Noise?

Jacobson and Truax (1991) proposed the reliable change index (RCI) to test whether one person's change is statistically reliable, meaning bigger than measurement error alone would produce. The Inferential Statistical Tests lesson gives the formula. Here you run it.

  1. SEM = SD × √(1 − r). Use the SD of pretest scores and the test's reliability.
  2. Standard error of the difference (S_diff) = √(2 × SEM²).
  3. RCI = (posttest − pretest) ÷ S_diff.

An RCI beyond ±1.96 means the change is reliable at the 95% level (Ferrer-Urbina et al., 2023; Vaganian et al., 2020).

Worked Example

Maya scores 34 on a depression inventory at intake and 24 after treatment. Lower is better. The inventory has SD = 10 and reliability = .91.

  • SEM = 10 × √.09 = 10 × .3 = 3
  • S_diff = √(2 × 3²) = √18 ≈ 4.24
  • RCI = (24 − 34) ÷ 4.24 ≈ −2.36

The size, 2.36, is beyond 1.96, and the sign points toward improvement. Maya's change is reliable. Shortcut: on this test, any change of about 8.3 points or more (1.96 × 4.24) is reliable.

How Reliability Feeds the RCI

ReliabilitySEMS_diffChange neededMaya's 10-point drop
.913.04.24about 8.3 pointsReliable (RCI ≈ −2.36)
.755.07.07about 13.9 pointsNot reliable (RCI ≈ −1.41)

Lower reliability means a bigger SEM and a bigger change needed. In music production terms, a noisy recording has a high noise floor. A quiet guitar line that's obvious on a clean track gets lost in the hiss. Exam rule: to detect change in one person, pick the most reliable measure you can.

Caveat: simulations show the RCI can flag too many false positives, and alternative indices exist (Ferrer-Urbina et al., 2023).

Clinical Significance: Did the Client Reach the Healthy Range?

The RCI answers "Is it real?" Clinical significance answers "Is it enough?" Jacobson, Follette, and Revenstorf (1984) defined clinically significant change as therapy taking a client out of the dysfunctional range or into the functional range (Jacobson & Truax, 1991). That gives three cutoffs (Benfer et al., 2024; Jacobson & Truax, 1991):

  • Cutoff a: the posttest falls 2 SD or more beyond the clinical group's mean, in the healthy direction. Use it when no healthy norms exist.
  • Cutoff b: the posttest falls within 2 SD of the healthy group's mean.
  • Cutoff c: the posttest passes the point between the two group means, weighted by each group's SD (Vaganian et al., 2020). With equal SDs, c is the exact midpoint.

When healthy norms exist, use b or c. Cutoff c is the recommended choice when the clinical and healthy distributions overlap (Benfer et al., 2024; McMillan et al., 2010).

Combine the RCI and the cutoff, and every client lands in one box (Jacobson & Truax, 1991; Vaganian et al., 2020):

CategoryReliable change?Passed the cutoff?
RecoveredYes, improvingYes
ImprovedYes, improvingNo
UnchangedNoDoesn't matter
DeterioratedYes, worseningDoesn't matter

Back to Maya. Say the clinical group averages 30 (SD 10) and the community averages 10 (SD 10). Then a = 30 − 20 = 10, b = 10 + 20 = 30, and c = 20. Maya's 24 passes b but not a or c. Here b sits right at the clinical mean, so b is too easy and a is too strict. That overlap is why c is the pick. Using c, she is improved, not recovered. If she had dropped to 18, her RCI would be about −3.77 and she'd be below 20: recovered. Notice that the cutoff you pick changes the verdict.

Caveat: clients who start with mild symptoms have little room to show clinically significant change (Vaganian et al., 2020).

Statistical vs. Clinical Significance

A large study can find a statistically significant difference that is clinically trivial (covered in the Inferential Statistical Tests lesson). Statistical significance and effect size describe groups. The RCI and clinical cutoffs describe individuals.

QuestionToolLevel
Is the group difference unlikely by chance?Statistical significance (p-value)Group
How big is the difference?Effect size (Cohen's d: .2 small, .5 medium, .8 large)Group
How many people must we treat for one extra good outcome?Number needed to treat (NNT)Group, applied
Is this person's change bigger than measurement error?RCIIndividual
Did this person reach the normal range?Clinical significance cutoffIndividual
Is the change big enough for the client to notice?Minimal clinically important difference (MCID)Individual

NNT is 1 divided by the absolute risk reduction. It carries both statistical and clinical meaning (Cook & Sackett, 1995). The MCID, also called the minimal important difference (MID), is the smallest change patients see as beneficial, big enough to change their management (Jaeschke et al., 1989, as cited in Vaganian et al., 2020).

Measurement-Based Care and Routine Outcome Monitoring

Measurement-based care (MBC) means giving standardized measures repeatedly and using the results to track progress and adjust the plan (American Psychiatric Association, 2022; Scott & Lewis, 2015). ROM's four components and the barriers to using it are covered in the Prevention, Consultation, and Psychotherapy Research lesson.

What works: frequent measurement, timely feedback, and results reviewed during the session. What doesn't: one-time screening, infrequent measurement, and feedback given outside the session (Fortney et al., 2017).

Lambert's OQ System and the Not-On-Track Client

The Outcome Questionnaire-45 (OQ-45) has 45 items about the past week and three subscales: symptom distress, interpersonal relations, and social role functioning. It is built to detect small changes. Critical items flag suicidality, substance use, and aggression (Matavovszky et al., 2024).

Lambert's system compares each client's scores with expected progress and alerts the therapist when a client is not on track (Shimokawa et al., 2010). Why not just trust the therapist? Clinicians rarely predict who won't benefit. Formal monitoring flagged every client who ended treatment worse, 85% of them by the third session (Hannan et al., 2005). Worsening hits a minority, about 5% to 10% of clients (Shimokawa et al., 2010).

The standard exam answer: feedback helps most for not-on-track clients. In Lambert's meta-analysis, feedback reduced deterioration and nearly doubled reliable or clinically significant change for clients predicted to do poorly (Lambert et al., 2018). Caveat: a larger 58-study meta-analysis found small effects overall (d = 0.15) and only slightly larger for not-on-track clients (d = 0.17) (de Jong et al., 2021).

PHQ-9 and GAD-7 as Change Measures

MeasureItemsSeverity cutoffsMCID
PHQ-99 DSM-IV criteria, each 0 to 3 (total 0 to 27)5, 10, 15, 20 = mild, moderate, moderately severe, severeAbout 5 points
GAD-775, 10, 15 = mild, moderate, severeAbout 4 points

A PHQ-9 of 10 or more had 88% sensitivity and 88% specificity for major depression (Kroenke et al., 2001). The GAD-7's best screening cutpoint is also 10 (Kroenke et al., 2010). The PHQ-9's MCID for one person is 5 points, about 2 SEMs (Löwe et al., 2004); the GAD-7's is 4 points (Toussaint et al., 2020). The PHQ-9's original rule for clinically significant change was a 50% drop plus a final score of 9 or less, which closely matches cutoff c (McMillan et al., 2010).

DSM-5-TR's cross-cutting symptom measures, severity measures, and WHODAS 2.0 are also built to be repeated over time to track change (American Psychiatric Association, 2022); they are covered in the Classification Systems lesson.

Goal Attainment Scaling

Goal attainment scaling (GAS) sets a measurable scale for each person's own goals before treatment and turns goal attainment into a standardized T-score (Kiresuk & Sherman, 1968). Each goal is rated from −2 to +2. Scores combine into a T-score that averages 50 when goals are met exactly as expected, with an SD of about 10 (Turner-Stokes, 2009). GAS is idiographic: it measures this client's goals, not a population's symptoms.

Example: a teen's goal is school attendance. −2 means she still misses most days. 0 is the expected result, missing no more than one day a week. +2 is perfect attendance.

Beyond the Individual: Couples, Families, Groups, Programs, and Prevention

Change isn't only measured in single clients, and each unit has its own trap.

  • Couples: Measure each partner. Partners influence each other, so their scores are not independent; dyadic methods like the actor-partner interdependence model handle this (Maroufizadeh et al., 2018). The IRT-built Couples Satisfaction Index measures satisfaction more precisely than the older Dyadic Adjustment Scale (Funk & Rogge, 2007).
  • Families: The McMaster Family Assessment Device rates family functioning on seven scales (Lyke & Matsen, 2013).
  • Groups: Group members interact, which breaks the independence assumption behind most statistics. In 33 group-treatment studies, none analyzed this correctly; after correction, only 12.4% to 68.2% of significant results held up (Baldwin et al., 2005).
  • Programs and organizations: Kirkpatrick's four levels and formative vs. summative evaluation are covered in the Training Methods and Evaluation lesson. GAS itself was built to evaluate community mental health programs (Kiresuk & Sherman, 1968).
  • Prevention: The outcome is incidence, meaning new cases. Across 19 trials, psychological prevention cut new depressive disorders by 22%, with an NNT of 22, and results did not differ by universal, selective, or indicated programs (Cuijpers et al., 2008). Prevention levels are covered in the Prevention, Consultation, and Psychotherapy Research lesson.
  • Single-case designs track one client with repeated measures across baseline and treatment phases (covered in the Research: Single-Subject and Group Designs lesson).

Technology in Assessment: Same Test, New Delivery

The exam asks about validity, cost-effectiveness, and consumer acceptability. APA's assessment guidelines name reliability, validity, and fairness as the essential criteria for technology-based measures, the same as for any test (American Psychological Association, 2020, Guideline 15).

Paper vs. Computer: Are the Scores Equivalent?

Don't assume equivalence; check the evidence. For cognitive tests, scores lined up well across modes for timed power tests (corrected cross-mode r = .97) but much less for speeded tests (r = .72). Adaptive and conventional computerized tests did not differ in equivalence (Mead & Drasgow, 1993). For self-report measures, computer and paper versions differed by about 0.2% of the scale range on average and correlated about .90, close to paper-to-paper retest (Gwaltney et al., 2008).

The International Test Commission (2006) lists what equivalence requires: comparable reliabilities, the expected correlation between versions, comparable correlations with other measures, and comparable means and SDs. Before using paper norms for a computer version, the test user confirms this evidence exists.

Moving a test to a screen is like remixing a track. Usually the song survives, but you check that the hook still lands.

Computerized Adaptive Testing

CAT basics are covered in the Other Measures of Cognitive Ability and Item Analysis and Test Reliability lessons. The angle here is efficiency. A depression CAT needed a mean of 12 items to reach its precision target while correlating .95 with the full 389-item bank. A fixed test holds the number of items constant and lets precision vary; a CAT holds precision constant and lets the number of items vary (Gibbons et al., 2012).

Computer-Based Test Interpretation

Computer-based test interpretation (CBTI) turns scores into narrative reports. The danger is the Barnum effect: people accept vague, general statements that fit almost anyone as accurate descriptions of themselves (Hua & Zhou, 2023). Students rated bogus CBTI reports 57.9% accurate, against 74.5% for real ones (Guastello et al., 1989). A computerized Rorschach report showed only 5% discriminating power for a given patient, and 60% of its statements just described typical outpatients (Prince & Guastello, 1990). Standard answer: CBTI reports are adjuncts to, not substitutes for, clinical judgment (Butcher et al., 2000). Your CBTI duties are covered in the APA Ethics Code Standards 9 & 10 and Cut Scores, Norms, and Precision lessons.

Tele-Assessment: What the Evidence Shows

A meta-analysis pooled 12 crossover studies in which adults took the same tests by video and in person (Brearly et al., 2017):

  • Verbally mediated tasks (digit span, verbal fluency, list learning) were not affected.
  • Untimed tasks and tasks allowing repetition ran about 0.1 SD lower by video, a small drop; so did the Boston Naming Test.
  • Motor tasks were too mixed to judge.
  • High-speed connections gave consistent results; older samples and slower connections were more variable.

APA's 2024 telepsychology guidelines (Guideline 8) ask you to check the evidence and norms for remote use, keep standard conditions when possible, document every deviation, protect test security, and confirm the client's identity. Self-report surveys adapt easily; tasks with physical materials like blocks don't, and video lag can distort timed tasks. Check the room for other people and phones, since third-party monitoring can threaten validity (American Psychological Association, 2024). The ethics side is in the APA Ethics Code Standards 9 & 10 lesson.

Giving block design over video is like teaching footwork over a laggy video call. Talking through the steps works. Judging the timing doesn't.

Unproctored Internet Testing and Test Security

The International Test Commission (2006) sorts online testing by level of control:

ModeSupervisionIdentity check
OpenNoneNone
ControlledNoneKnown test takers with a login
Supervised (proctored)A proctorYes
ManagedTesting center, highest controlYes

For moderate- or high-stakes use, people who pass a controlled-mode test should retake a supervised test, and the scores are compared. Norms for internet tests should come from people tested under the conditions real test takers will face, unproctored included (International Test Commission, 2006).

Ecological Momentary Assessment

Ecological momentary assessment (EMA) samples a person's current behavior and experience repeatedly, in real time, in daily life. It reduces recall bias and maximizes ecological validity, meaning results generalize to real life. Prompts can be tied to events or sent at random times, by paper diary, phone, or sensor (Shiffman et al., 2008).

Instead of asking on Friday, "How anxious were you this week?", the phone pings at random times and asks, "How anxious are you right now?"

Digital Phenotyping and Passive Sensing

Digital phenotyping is moment-by-moment measurement of a person's traits and behavior in daily life, using data from their own devices. Active data need the person's effort, like surveys. Passive data need none: GPS location, accelerometer movement, and call and text logs (Onnela & Rauch, 2016). Personal sensing is promising but still in its infancy, with open problems in privacy and in handling huge numbers of variables (Mohr et al., 2017). Being tracked may itself worsen anxiety or paranoia, and the technology risks a digital divide (Onnela & Rauch, 2016).

AI-Assisted Scoring and Prediction

Treat an AI tool like any test: it needs validity evidence for your purpose and your population. Two failure modes are exam-ready:

  • Proxy bias: A health algorithm predicted future costs instead of illness. Because less money is spent on Black patients, Black patients at the same risk score were sicker. Fixing it would raise the share of Black patients flagged for extra help from 17.7% to 46.5% (Obermeyer et al., 2019).
  • Base rates: Suicide prediction models classified people well overall (accuracy ≥ .80 in most), but positive predictive value for suicide death was ≤ .01 in most models (Belsher et al., 2019). Rare outcomes produce mostly false alarms.

Consent, HIPAA, and human oversight for AI tools are covered in the AI in Practice lesson.

Cost-Effectiveness and Consumer Acceptability

  • Cost: In a trial of 2,233 patients, computerized feedback raised reliable improvement by about 8 percentage points, at about £15 extra per patient; it was likely cost-effective (Delgadillo et al., 2021).
  • Acceptability: Older adults tested by video, including those with cognitive impairment, reported 98% satisfaction; about two-thirds had no preference between video and face-to-face (Parikh et al., 2013). Patients and providers find MBC highly acceptable (Fortney et al., 2017).
  • Disclosure: In a survey of adolescent males, audio computer-assisted self-interviewing raised reports of some sensitive behaviors by a factor of 3 or more compared with a paper questionnaire (Turner et al., 1998).
  • Access: Tech experience varies, so check reliability, validity, and acceptability in diverse groups, including older adults (American Psychological Association, 2020).

EPPP Traps and Common Misconceptions

Misconception 1: "A statistically significant treatment effect means clients recovered."

  • Reality: Significance describes the group. Recovery needs reliable change plus a passed cutoff, person by person.

Misconception 2: "A client whose score crossed the cutoff has recovered."

  • Reality: Without reliable change, the client is unchanged, cutoff or not.

Misconception 3: "Improvement on retest proves the treatment worked."

  • Reality: Rule out regression to the mean and practice effects first.

Misconception 4: "Alternate forms eliminate practice effects."

  • Reality: They help but don't guarantee it.

Misconception 5: "Computer versions of tests are automatically equivalent."

  • Reality: Power tests and self-reports usually transfer; speeded tests often don't. Check the equivalence evidence.

Misconception 6: "An accurate AI model gives trustworthy positive results."

  • Reality: With rare outcomes, even accurate models give mostly false positives.

Misconception 7: "A CBTI report that feels accurate is valid."

  • Reality: Generic Barnum statements feel accurate too.

Memory Aids

  • Real, then enough: RCI asks if change is real; the cutoff asks if it's enough. Both = recovered.
  • RCI recipe: "SEM, double, root, divide." Square the SEM and double it, take the root, divide the change by it.
  • Cutoffs: a = away from the clinical group; b = belongs with the healthy group; c = center point between them.
  • Equivalence: "Power transfers, speed stumbles."
  • Tele-assessment: "Talk transfers; touch and timing trouble."
  • ITC modes, least to most control: Open, Controlled, Supervised, Managed: "Only Cops Stop Mischief."
  • AI red flags: "Proxy and prevalence."

Key Takeaways

  • Three noise sources move scores without treatment: measurement error, regression to the mean, and practice effects. Change scores are less reliable than they look. Alternate forms reduce practice effects but don't remove them.
  • RCI = change ÷ S_diff, where S_diff = √(2 × SEM²). Beyond ±1.96 = reliable. Higher reliability = smaller change needed.
  • Clinical significance (Jacobson & Truax): cutoff a when no healthy norms exist; b or c when they do; c when the distributions overlap. Categories: recovered, improved, unchanged, deteriorated.
  • Statistical significance and effect size describe groups; RCI, cutoffs, and the MCID describe individuals. NNT = 1 ÷ absolute risk reduction.
  • MBC and ROM: measure often, feed back fast, review in session. OQ-45 alerts flag not-on-track clients, who benefit most from feedback; newer meta-analytic effects are small.
  • PHQ-9 cutoffs 5/10/15/20, MCID about 5; GAD-7 cutoffs 5/10/15, MCID about 4. DSM-5-TR's cross-cutting measures and WHODAS 2.0 track change. GAS measures individualized goals as T-scores.
  • Beyond individuals: partners and group members aren't independent; prevention is measured by incidence and NNT.
  • Technology is judged by reliability, validity, and fairness. Power tests and self-reports transfer to computers; speeded tests may not. Verbal tasks transfer to video; manipulatives and motor tasks need caution, and lag can distort timed tasks.
  • CBTI is an adjunct, not a substitute; watch for Barnum statements. EMA cuts recall bias; passive sensing raises privacy concerns; AI fails through proxies and base rates.

Now cover the tables and rerun Maya's numbers with reliability .75. If you get "not reliable," you've got the core of this lesson.

Ready to practice?

Get started in the app