Assessment Models and Methods: How Assessment Actually Gets Done
Tests are tools. A model tells you what to look for. A method tells you how to collect it. The EPPP rarely asks you to name a test here. It asks which approach fits a case, which source to trust, and why two sources disagree.
Why This Matters for Psychologists
An 8-year-old is referred for "not paying attention." Mom says it happens all day. The teacher says only during math. The child says he's fine. In your office he sits still for an hour. Nobody is lying. Each source sees a different slice. Your job is to pick the right model, gather data from more than one method, and read the disagreements as information.
Part 1: Assessment Models (What You Look For)
Traditional vs. Behavioral: Sign or Sample?
Traditional assessment reads a response as a sign of something underneath, like a trait or an inner conflict (Domino & Domino, 2006). Behavioral assessment reads what you observe as a sample of how the person acts in that situation (Kazdin, 1979). Goldfried and Kent (1972) laid out how far apart these two sets of assumptions sit.
A sign approach treats a child's tiny, cramped drawing as a clue to insecurity. A sample approach counts how often she leaves her seat during math, because leaving her seat in math is the problem.
| Traditional | Behavioral | |
|---|---|---|
| A response is | A sign of an underlying trait | A sample of behavior in context |
| Main question | What is this person like? | What sets off and keeps this behavior going? |
| Typical tools | Personality inventories, projective tests | Observation, functional assessment, self-monitoring |
Functional Assessment: ABC and SORC
The core behavioral tool is the ABC analysis: Antecedent (what comes before), Behavior, and Consequence (what follows). The goal is the behavior's function, the payoff that keeps it going. A functional behavioral assessment usually climbs three steps (Tereshko et al., 2023):
- Indirect: interviews, rating scales, questionnaires
- Descriptive: watching in the natural setting and writing narrative ABC records
- Experimental functional analysis: changing conditions on purpose and measuring the behavior. This is the gold standard (Hanley et al., 2003; Tereshko et al., 2023).
Descriptive data show only what tends to go together. They can mistake a common consequence for the reinforcer (Tereshko et al., 2023). A teacher always says "stop that" after a student yells, so the ABC notes point to attention. But the yelling may really be getting him out of a hard worksheet, and the scolding just tags along. The steps of a formal functional analysis are covered in the Interventions Based on Operant Conditioning lesson.
SORC adds the person to the chain: Stimulus, Organism (what the person brings, such as bodily state, beliefs, and risk or protective factors), Response, Consequence (Esposito-Smythers et al., 2012; Loose et al., 2020). Kanfer and Saslow's SORKC version adds K, the contingency between response and consequence (Loose et al., 2020).
Ecological Assessment
Bronfenbrenner's ecological systems theory says development grows out of interactions among settings such as family, school, and community. Ecological assessment applies that idea. You assess the systems around the person, not just the person, and you favor tools close to everyday life (Weisse et al., 2026).
A child melts down at school but never at Grandma's. The ecological question is not "What is wrong with this child?" It is "What is different between those two places?"
This ties to ecological validity. Many neuropsychological tests show only moderate ecological validity for everyday thinking, and they predict best when the real-life outcome matches the skill tested (Chaytor & Schmitter-Edgecombe, 2003). Ecological momentary assessment is covered in the Measuring Change and Assessment Technology lesson.
Developmental Assessment
The developmental model says the same behavior means different things at different ages. DSM-5-TR builds this into criteria. ADHD symptoms must be inconsistent with developmental level (American Psychiatric Association [APA], 2022).
The key split is screening vs. diagnostic evaluation. Pediatric guidance calls for developmental surveillance at every well-child visit and a standardized screening test at 9, 18, and 30 months. A concern leads to screening or referral, and a problem found that way leads to a full evaluation and diagnosis (Lipkin & Macias, 2020). A screen sorts. It does not diagnose.
A screen is a smoke detector: cheap, always on, and sometimes it beeps at burnt toast. The diagnostic evaluation is the fire inspector who comes out to look. Infant tests such as the Bayley-4 are covered in the Other Measures of Cognitive Ability lesson.
Neuropsychological Assessment: Fixed, Flexible, and Process
A fixed battery gives everyone the same full set of tests, whatever the referral question. The Halstead-Reitan and Luria-Nebraska are the classic examples. A flexible battery gives a small core set to everyone, then adds tests to check specific hypotheses (Larrabee, 2014). In one survey, 78% of neuropsychologists used a flexible battery, 18% a completely flexible approach, and only 5% a fixed battery (Sweet et al., 2011, as cited in Larrabee, 2014).
The process approach, also called the Boston Process Approach, grew from Edith Kaplan's work at the Boston VA. It asks how a person reaches an answer and why they struggle, using error analysis and problem-solving strategies (Diaz-Orueta et al., 2020). Critics faulted it for weak psychometrics, though norm-based error rates now exist for some tests (Diaz-Orueta et al., 2020).
Don't let the labels blur. Fixed vs. flexible is about which tests you give. Process is about what you watch while the person works. The process approach can run on either kind of battery, but it usually rides on a flexible one (Casaletto & Heaton, 2017). That is why some sources, including the Clinical Tests lesson, file it under "the flexible approach." Battery details live in that lesson.
| Approach | What happens | Strength | Weakness |
|---|---|---|---|
| Fixed | Same full battery for everyone | Uniform, standardized data | Not tailored to the question |
| Flexible | Core set plus tests picked for the question | Targeted to the question | Less standardized across cases |
| Process | Watch how the person solves items | Shows the reason behind a low score | Weaker psychometric backing |
Therapeutic Assessment
In the traditional model, the client hands over data and receives the expert's conclusions. Collaborative/Therapeutic Assessment treats clients as partners in making meaning (Aschieri et al., 2026). Stephen Finn developed Therapeutic Assessment, a brief, semistructured intervention. Client and assessor work together in every phase, and tests serve as "empathy magnifiers" (Durosini & Aschieri, 2021). The client helps write the questions the assessment will try to answer (Engelman et al., 2016; Smith et al., 2015).
It tends to help. Assessment with personalized, collaborative feedback produced d = 0.42 across 17 studies, with the largest effect on therapy process (Poston & Hanson, 2010). A meta-analysis of Finn's model found g = .46 for treatment process and g = .34 for symptoms (Durosini & Aschieri, 2021). Caveat: critics argued the 2010 review may overstate effects and did not rule out Barnum effects (Lilienfeld et al., 2011).
A traditional assessor is a mechanic who fixes the car and hands you the bill. A therapeutic assessor pulls you under the hood and shows you why the engine keeps stalling.
Identifying Learning Disabilities: Discrepancy vs. Response to Intervention
The ability-achievement discrepancy model labels a learning disability when achievement falls far below IQ. Its problems: referrals came late, usually after failure, and the method has not proven reliable. Poor readers with and without an IQ gap differ little (Fletcher & Vaughn, 2009). IQ barely predicts how kids respond to reading instruction. Across 22 studies it explained about 3% of the variance or less (Stuebing et al., 2009).
Response to intervention (RTI) flips the order. Screen every child. Teach well in class (Tier 1). Add small-group help, three to five students, for kids who fall behind (Tier 2). Add intensive, individual help for those still behind (Tier 3). Track progress weekly or every two weeks. A poor response to good teaching becomes part of the case for a learning disability. IDEA 2004 lets schools use RTI instead of a discrepancy (Fletcher & Vaughn, 2009). Caveat: there are no agreed criteria for an "inadequate" response, so RTI should not decide eligibility alone (Fletcher & Vaughn, 2009).
The discrepancy model waits for the gap to grow wide enough to see. RTI is a coach who adds extra practice right away and watches who still can't hit the ball.
DSM-5-TR requires no IQ-achievement gap (APA, 2022):
| Discrepancy model | RTI | DSM-5-TR specific learning disorder | |
|---|---|---|---|
| Core evidence | Achievement well below IQ | Poor response to good instruction | Low achievement for age that persists |
| Role of IQ | Central | Small | Rule out intellectual disability (IQ above about 70, ±5) |
| Time frame | Often after failure | Early screening | Difficulty lasting at least 6 months despite targeted help |
| Cutoff | Size of the gap | Response benchmarks | Standard score of 78 or less (1.5 SD) for the most certainty; 1.0 SD with converging evidence |
One exception to the age rule: in a gifted student, DSM-5-TR compares achievement with the student's own ability, not the population mean. No single data source is enough; the diagnosis is a clinical synthesis of history, school reports, and testing (APA, 2022). Full criteria are in the Neurodevelopmental Disorders lesson.
Multimethod Assessment
A review of over 125 meta-analyses and 800 multimethod samples found that different methods give unique information. Clinicians who rely only on interviews tend to reach incomplete understandings (Meyer et al., 2001). Whether an added measure actually improves accuracy (incremental validity) is covered in the Clinical Judgment and Interpretation lesson.
Part 2: Assessment Methods (How You Collect Data)
Clinical Interviews: Unstructured, Semi-Structured, Structured
In an unstructured interview, questions are not set ahead of time and answers are not scored by a standard system. In a structured interview, everyone gets the same prepared questions and a standard rating system. A semi-structured interview sits between. The SCID is the classic example: a clinician gives it in modules and follows a decision tree to test diagnostic hypotheses (Spitzer et al., 1992). A fully structured interview, like the CIDI, can be given by trained lay interviewers (Haro et al., 2006).
Why add structure? Raters disagree partly because they ask different questions (information variance) and partly because they apply different rules (criterion variance) (Kobak et al., 2009). Spitzer et al. (1975) named differing criteria the largest source of unreliable diagnosis. Structured interviews attack both with the same questions and the same rules. They are now the gold standard for reliable diagnosis (Bruchmüller et al., 2011).
They also catch more. In one clinic, more than a third of patients interviewed with the SCID got three or more diagnoses, compared with fewer than 10% after routine unstructured interviews. Anxiety disorders were among the most often missed (Zimmerman & Mattia, 1999). Diagnoses from routine clinical evaluations agree only modestly with standardized interviews, mean kappa = .27 (Rettew et al., 2009).
Do structured interviews hurt rapport? Patients rated their satisfaction with one at 86.55 out of 100. Surveyed therapists guessed 49.41, and they used structured interviews with only about 15% of their patients (Bruchmüller et al., 2011).
Unstructured is cooking by feel. Semi-structured is a recipe the chef can adjust. Fully structured is a meal kit with exact steps, so even a first-timer gets the same dish.
| Interview | Structure | For | Know this |
|---|---|---|---|
| SCID-5 | Semi-structured, clinician-given | DSM diagnoses | Modules and decision tree; clinician version shows high reliability (Osório et al., 2019) |
| K-SADS-PL | Semi-structured | Children and teens | Interviews both parent and child (Kaufman et al., 1997; Makino et al., 2023) |
| ADIS-5 | Semi-structured (Gordon & Heimberg, 2011) | Anxiety and related disorders (Mournet et al., 2025) | Child and parent versions, ADIS-C/P (Silverman et al., 2001) |
| MINI | Short, structured | DSM and ICD diagnoses; MINI-Kid for youth | About 15 minutes (Sheehan et al., 1998) |
| CIDI | Fully structured, lay-given | Large epidemiological surveys | Built for the WHO; computer-scored diagnoses (Robins et al., 1988; Kessler & Üstün, 2004) |
Even with a semi-structured interview, a common source of disagreement is whether symptoms are enough in number, severity, or duration (Brown et al., 2001). Structure narrows judgment calls. It does not remove them. Hiring interviews and the CIDI's role in national surveys are covered in the Employee Selection and Epidemiology and Base Rates lessons.
The Mental Status Exam
The mental status exam (MSE) is psychiatry's version of the physical exam. It records signs the examiner observes and symptoms the client reports. It is cross-sectional: it describes the person at the time of the interview, and it isn't built around one diagnosis (Oyesanya & Correll, 2026). Textbooks organize it differently, but nine core domains recur (Daza et al., 2025; Oyesanya & Correll, 2026). Many outlines add a tenth, judgment (Griffeth et al., 2017; Khan et al., 2025):
| Domain | What you note |
|---|---|
| Appearance | Clothing, grooming, physical condition |
| Behavior | Motor activity, abnormal movements |
| Speech | Rate, rhythm, flow, tone |
| Mood and affect | Sustained emotional state (mood); observable expression (affect) |
| Thought process | Flow, logic, and organization, such as loose associations |
| Thought content | What the person thinks about, such as delusions or suicidal ideas |
| Perception | Hallucinations and other altered perceptions |
| Cognition | Orientation, attention, memory |
| Insight | Awareness of one's own mental health problem |
| Judgment | How sound the person's decisions are |
The MSE is a photo. The history is the movie. You need both. Mood vs. affect is covered in the Cognition, Mood, and Temperament lesson.
Self-Report and Response Sets
Self-report inventories are easy to give and cheap. Their weakness is answers that are socially desirable, exaggerated, or misleading, on purpose or not. Watch three response sets, habits that shape answers apart from content:
- Social desirability: looking good. Teens high on a social desirability scale reported fewer depression and anxiety symptoms (Logan et al., 2008).
- Acquiescence: "yea-saying," agreeing regardless of content. It is higher among older and less educated respondents and in cultures with strong norms of deference (Lechner et al., 2019).
- Extreme responding: picking the end points of a scale, partly regardless of content. It is very common (Schoenmakers et al., 2026).
A balanced scale mixes items keyed in both directions. Someone who agrees with both "I feel calm" and "I feel tense" is yea-saying, not describing a mood. Validity scales for faking good and bad are in the MMPI-2 lesson.
Multi-Informant Assessment
Informants agree less than you'd think. In Achenbach's classic meta-analysis, mean correlations were .60 between similar informants who see the child in the same setting (two parents), .28 between different kinds of informants (parent and teacher), and .22 between the child's self-report and others' reports (Achenbach et al., 1987, as cited in Martel et al., 2017). Agreement was higher for externalizing than internalizing problems and for younger children (Achenbach et al., 1987, as cited in Moore et al., 2022). A later meta-analysis of 341 studies found nearly the same: r = .28 overall, .25 for internalizing and .30 for externalizing problems (De Los Reyes et al., 2015).
Why they disagree:
- Behavior changes with context. Children may show concerns at home and not at school (De Los Reyes et al., 2015). ADHD signs can vanish with frequent rewards, close supervision, a new setting, or one-on-one time, including the clinician's office (APA, 2022).
- Informants differ. The pair of raters, the type of problem, and the measure all change agreement (De Los Reyes et al., 2015). The Attribution Bias Context model is a framework for studying these discrepancies in clinic settings (De Los Reyes & Kazdin, 2005). Cultural background of the child and the rater can shape ratings (APA, 2022).
So disagreement is data, not error. It can tell you where a problem lives. DSM-5-TR requires several ADHD symptoms in two or more settings and notes that confirming this usually means consulting people who see the person there (APA, 2022). Using collateral to settle a diagnosis is covered in the Differential Diagnosis lesson.
| System | Forms | Ages | Notes |
|---|---|---|---|
| ASEBA (Achenbach) | CBCL (parent), TRF (teacher), YSR (youth self-report) | CBCL and TRF 6-18; YSR 11-18 | Internalizing, Externalizing, Total Problems; 8 syndrome scales; DSM-oriented scales; side-by-side informant comparison (Bordin et al., 2013) |
| BASC-3 | Parent and teacher rating scales, self-report, classroom observation system | Ratings 2-21; self-report 6 through college | Pearson now lists a BASC-4 |
| Conners 4 | Parent, teacher, self-report | Parent and teacher 6-18; self 8-18 | ADHD plus common co-occurring problems |
Direct Observation
Naturalistic observation watches behavior in its real setting. It gives high ecological validity but is hard to set up and control. Analog or structured observation sets up a task that stands in for real life. In a behavioral approach (avoidance) test, a child with a phobia approaches the feared object step by step while you measure steps completed, distress, and arousal (Ollendick et al., 2011).
Recording methods. Continuous methods capture every instance. Discontinuous methods sample, so they are easier but less exact (Fiske & Delmolino, 2012; LeBlanc et al., 2020). Interval methods and momentary sampling are sometimes all called "time sampling" (Powell et al., 1975).
| Method | You record | Best for | Bias |
|---|---|---|---|
| Narrative (ABC) | A written account of what came before and after | Early ideas about function | Correlational only |
| Event (frequency) | Every instance | Discrete behaviors with a clear start and end | Hard at very high rates |
| Duration | How long each instance lasts | Behaviors where length matters | None built in |
| Latency | Time from a cue to the start of a response | How fast someone starts | None built in |
| Partial interval | Did it occur at any point in the interval? | Behaviors you want to decrease | Overestimates |
| Whole interval | Did it last the entire interval? | Behaviors you want to increase | Underestimates |
| Momentary time sampling | Is it happening at the moment the interval ends? | Ongoing, duration-type behaviors | Roughly unbiased for duration; usually the least error |
The bias directions are solid (Powell et al., 1975; Wirth et al., 2014), and the standard exam answer follows from them: partial interval for behavior you want to decrease, whole interval for behavior you want to increase. Each errs in the safe direction, so progress never looks better than it is. Caveat: practice guidelines agree on partial interval for excesses but call whole interval overly strict for deficits and recommend momentary time sampling there (Fiske & Delmolino, 2012). Momentary time sampling over- and underestimates about equally often, though it runs a bit low when the behavior fills little of the session and a bit high when it fills most of it (Powell et al., 1975; Wirth et al., 2014).
Partial interval is a strict referee who calls a foul for one touch. Whole interval is a lenient one who calls it only if the foul lasts the whole play.
Observer problems:
- Reactivity: people act differently when they know they're watched. Obtrusive observation is often reactive, and behavior under obtrusive and unobtrusive conditions can bear little relation (Kazdin, 1979). Fixes: unobtrusive measures, archival records, and letting people get used to observers.
- Reactivity of reliability checks: observers agree more on days they know they're being checked (Taplin & Reid, 1972).
- Observer drift: a gradual, systematic shift in how observers interpret the behavior codes (Paul et al., 1986, as cited in Lumen Learning, n.d.). Regular recalibration is the fix. In depression ratings, experienced raters who hadn't calibrated together were less reliable than beginners (Kobak et al., 2009). Raters drifting together is consensual observer drift, covered in the Item Analysis and Test Reliability lesson.
- Observer bias: expectations skew what gets recorded. Use clear behavior definitions and check inter-rater reliability.
Self-Monitoring
In self-monitoring, clients record their own behavior, such as drinks, panic attacks, or urges. It reaches private events and daily life. The catch is reactivity: recording a behavior can change it. That can confound measurement or serve as treatment, and the effect has been inconsistent in research (Isaacs et al., 2021).
Psychophysiological Measures
These measure the body: heart rate, sweating, breathing, and penile response. They don't depend on what the person says. Limits: the polygraph measures arousal, but no arousal pattern is unique to lying, so its validity is highly questionable. On the stronger side, phallometric testing is a valid indicator of sexual interest in children and predicts sexual reoffending (McPhail et al., 2019).
Body, words, and behavior can disagree. In Lang's three-system model (also called the tripartite model of fear), fear shows up as physiological arousal, subjective distress, and behavioral avoidance, and these may line up (concordance) or not (discordance) (Ollendick et al., 2011).
Work Samples and Assessment Centers
A work sample has a job candidate do real job tasks. An assessment center uses several exercises and several raters. Both belong mainly to I-O psychology and are covered in the Employee Selection lessons.
Methods at a Glance
| Method | Main strength | Main limit |
|---|---|---|
| Unstructured interview | Flexible, follows the client | Misses comorbidity; less reliable |
| Structured interview | Reliable, thorough | Less flexible; threshold calls remain |
| Self-report | Cheap; reaches inner experience | Response sets, faking |
| Multi-informant ratings | Shows behavior across settings | Low agreement to interpret |
| Direct observation | Actual behavior | Reactivity, drift, cost |
| Self-monitoring | Daily life, private events | Reactivity |
| Psychophysiology | Doesn't rely on words | Arousal isn't specific |
EPPP Traps and Common Misconceptions
Misconception 1: "When parent and teacher disagree, one is wrong."
- Reality: Cross-setting agreement averages about .28. Disagreement often means the behavior differs by setting.
Misconception 2: "Partial interval recording undercounts behavior."
- Reality: It overestimates. Whole interval underestimates.
Misconception 3: "A child who sits still in your office doesn't have ADHD."
- Reality: DSM-5-TR says signs can vanish one-on-one, including in the clinician's office. You need informants from two or more settings.
Misconception 4: "Structured interviews ruin rapport, so skip them."
- Reality: Patients rate them highly. Therapists underestimate patient acceptance.
Misconception 5: "Semi-structured means the clinician can ask anything."
- Reality: The core questions and decision rules are set. The clinician adds follow-up probes and judges whether a symptom counts.
Misconception 6: "DSM-5-TR requires an IQ-achievement discrepancy for learning disorder."
- Reality: It requires low achievement for age, persisting at least 6 months despite targeted help, plus ruling out intellectual disability.
Misconception 7: "Most neuropsychologists use a fixed battery like the Halstead-Reitan."
- Reality: Flexible batteries dominate; only about 5% use a fixed battery.
Misconception 8: "Observer drift and reactivity are the same."
- Reality: Drift is the observer changing over time. Reactivity is the person observed changing because they're watched.
Misconception 9: "A positive developmental screen means a diagnosis."
- Reality: A screen sorts who needs a full evaluation. Diagnosis comes after.
Misconception 10: "The Boston Process Approach is just another name for a flexible battery."
- Reality: Flexible battery answers which tests. Process approach answers how the person solves them. The process approach usually runs on a flexible battery, which is why the two get lumped together.
Memory Aids
- Three ABCs: Antecedent-Behavior-Consequence (functional assessment); Ellis's Activating event-Belief-Consequence; the Attribution Bias Context model (informant discrepancies)
- Two "tripartite" models: Lang's fear model (arousal, distress, avoidance) vs. Clark and Watson's anxiety-depression model (shared negative affect, low positive affect, hyperarousal)
- Achenbach's 60-28-22: Similar raters .60, Different settings .28, Self vs. others .22
- Partial = Plenty (over), Whole = Withheld (under)
- Sign vs. Sample: Sign = Sits underneath; Sample = Same situation
- Interview structure ladder: unstructured (feel), semi-structured (recipe plus chef), fully structured (meal kit, lay interviewer)
- RTI tiers: 1 = whole class, 2 = small group, 3 = one-on-one intensive
Key Takeaways
- Traditional assessment reads responses as signs; behavioral assessment reads them as samples. Functional assessment (ABC, SORC) seeks the behavior's function; experimental functional analysis is the gold standard.
- Ecological assessment looks at the settings around the person. Developmental assessment judges behavior against age; screening sorts and diagnostic evaluation decides.
- Fixed batteries test everyone the same way; flexible batteries dominate; the process approach studies how a person solves items and usually runs on a flexible battery.
- Therapeutic Assessment (Finn) makes the client a partner who helps set the assessment questions. Effects are small to moderate and strongest on therapy process.
- IDEA 2004 lets schools use RTI instead of the weakly supported IQ-achievement discrepancy. DSM-5-TR requires low achievement for age, despite targeted help, with no IQ gap.
- Structured and semi-structured interviews (SCID-5, K-SADS, ADIS-5, MINI, CIDI) are more reliable and catch more comorbidity than unstructured interviews.
- The MSE is a cross-sectional snapshot: appearance, behavior, speech, mood and affect, thought process and content, perception, cognition, insight, and often judgment.
- Self-report risks social desirability, acquiescence, and extreme responding.
- Informants agree modestly (.60 similar, .28 different, .22 self vs. others), better for externalizing problems and younger children. Read disagreement as information about context. CBCL, TRF, YSR, BASC, and Conners give parallel forms.
- Observation: partial interval overestimates (use for decreases), whole interval underestimates (use for increases), momentary time sampling is roughly unbiased. Watch for reactivity, observer drift, and observer bias.
- Self-monitoring is reactive; psychophysiological measures skip words but arousal isn't specific.
Now cover the recording-methods table and rebuild it from memory: method, what you record, and which way it errs. Then explain the parent-teacher-child vignette at the top using 60-28-22.
