Why This Matters: You Are Sitting Inside the Example
The EPPP is a credentialing exam, and this lesson is the rulebook that built it. The AERA/APA/NCME Standards for Educational and Psychological Testing (2014) is the profession's shared manual for how scores get scaled, normed, linked, and cut. Chapter 11 covers credentialing directly, and it says something a candidate should sit with: a licensure test does not need to be equally precise across the whole scale. It needs to be precise near the cut score, because that is where pass and fail get decided.
Precision Where It Counts
The 2014 edition renamed the topic. Reliability/precision means score consistency across replications by any index: coefficients, standard errors, generalizability coefficients, IRT information functions, or classification-consistency indices. Reliability coefficient is reserved for the classical coefficient, and exam items may use either wording.
Everything rests on the replication of the testing procedure. Before a reliability number means anything you must say what may vary: time and place usually, items if forms sample a domain, the scorer if judgment is involved. Standards 2.1 and 2.2 require you to state that range and justify it.
Think of testing a recipe three times. Same kitchen, same oven, you cooking on three Sundays, and you learn about your consistency on Sundays. A different cook in a different kitchen each time, and you learn something much bigger. Neither answer is wrong, but they answer different questions, and you have to say which one you asked.
Two traps follow. Standard 2.6 says coefficients are not interchangeable, since internal consistency, alternate-form, and test-retest coefficients each define measurement error differently. Standard 2.7 covers scorers, and its comment is blunt that "interrater agreement does not guarantee high reliability" of examinee scores. Chapter 2 names test length as the first factor affecting reliability: shorten a test and reliability generally drops, lengthen it with comparable items and it generally rises.
Conditional SEM: the number near the cut
A single SEM is an average across the whole range, the one "Items and Reliability" taught you to build confidence bands with. The conditional standard error of measurement (CSEM) is the standard deviation of measurement errors at one specified score level, so it exposes what an average hides. Standard 2.14 requires CSEM at several score levels unless the error is known to be constant, and requires it near every cut score. Standard 2.13 adds that it belongs in the units of each reported score.
Decision consistency versus decision accuracy
Decision consistency asks whether the same person would land in the same category on a second replication. Decision accuracy asks whether the observed classification matches the person's true classification status. Standard 2.16 requires an estimate of the percentage classified the same way across two replications whenever a test is used for classification.
Error near the cut drives misclassification. People far above or below a cut carry a lot of error without changing category, so only people near the cut get flipped.
Differences, profiles, and adjusted coefficients
Standard 2.3 requires reliability evidence for every score or subscore that gets interpreted. Standard 2.4 adds that a difference between two observed scores needs standard errors for the difference itself, which can be far less reliable than either score behind it, and Standard 1.14 says the same on the validity side for subscores, differences, and profiles. Standards 1.21 and 2.20 require adjusted and unadjusted coefficients reported together with the procedure behind them, whether the fix corrects for range restriction or for attenuation.
Two Frameworks Beyond Classical Test Theory
Generalizability theory splits error into pieces. Where classical test theory assumes one undifferentiated error distribution, G theory uses analysis of variance to estimate separate variance components for items, occasions, raters, and their interactions. A person's expected score over all possible replications is the universe score, and the generalizability coefficient is universe-score variance divided by observed-score variance.
It is like troubleshooting a noisy recording by soloing each channel instead of turning the master volume down. Once you know the hiss comes from one microphone, you stop wasting money on better speakers.
In item response theory, precision is the test information function, mapping each trait level to the reciprocal of the conditional measurement error variance there. The comment to Standard 2.5 says adaptive precision can be estimated with model-based simulations, and that model-based conditional standard errors are particularly useful. Standard 4.3 requires adaptive and multistage design to be documented, including its termination conditions, and its comment gives the two stated reasons for going adaptive: greater precision, especially for high and low scorers, or comparable precision in less time. Chapter 4 adds the stopping rule: adaptive length is set by a fixed item count or by a target level of score precision.
Norms That Mean Something
Norms summarize how a defined reference population scored, so everything depends on whether that group is the right comparison. Standard 5.8 requires norms to refer to clearly described populations, including the groups users will want to compare their examinees against. Standard 5.9 lists what a norming study must report: population sampled, sampling procedure, participation rates, any weighting, the dates of testing, and descriptive statistics.
Local norms refer scores to one limited population, such as a district or clinic, and are not meant to generalize past it. User norms (also called program norms) describe whoever happened to be tested during some period, so they need a sound rationale and must be labeled as such. Subgroup norms can be useful, and Standard 5.8 notes their permissible use may be limited by law.
Norms go stale. Standard 5.11 splits the duty: while the test is in print the publisher must renorm often enough to keep interpretations accurate or show the old norms still hold, and the user must still avoid out-of-date norms. The comment to Standard 5.3 explains why. A scale point originally defined as the mean of a reference population stops meaning average performance if the scale is held fixed while the population changes.
It is what happened to clothing sizes. The tag still says medium, but medium was cut against bodies from decades ago, so the label drifted while the number stayed put.
That is the skeleton under the rising-IQ finding your prep material calls the Flynn effect, taught in "IQ Tests". The Standards never print that name, so no Standard number attaches to it. And outdated norms are not an outdated test: Chapter 4 says norms may need updating after rising or falling performance, or a changed test-taking population, while the content stays relevant.
Linking and Equating: When Are Two Scores the Same Score?
Linking is the umbrella term for relating scores across tests or forms. Equating is the strictest case: small statistical adjustments among alternate forms built to the same content and statistical specifications, so scale scores become interchangeable. Calibration, concordance, vertical scaling, projection, and moderation are weaker links that do not produce interchangeable scores.
Equating is converting between two rulers made from the same blueprint when one came out a hair long. Concordance is telling someone their shoe size in another country's sizing. Both are useful, but only one lets you hand the ruler to a colleague and forget which one you used.
Chapter 5 gives four equating designs. The same sample takes both forms, yielding the correlation between them but risking order effects. Randomly equivalent groups removes order effects but gives no correlation estimate. Embedded anchor items sit inside both forms, while an external anchor test is a separate section that does not count toward the total.
Anchors carry rules. Standard 5.15 requires anchor content to be exactly the same in each form, to proportionally reflect the content and difficulty range of the full test, to sit in similar positions, and to behave similarly in both forms after controlling for group differences. Chapter 5 adds that anchors whose relative difficulty has shifted substantially get dropped.
Vertical scaling links tests deliberately built to differ in difficulty, so growth across levels sits on one scale. Concordance relates a score on one test to a score on another test of a similar construct. Neither is stable by default: Standard 5.17 requires direct evidence of comparability whenever scores that cannot be equated are linked, and requires naming the population it holds for, since a conversion accurate for native speakers might overpredict or underpredict for nonnative speakers.
The comment to Standard 5.13 puts equating error in units of the reported score scale and says that for programs with cut scores, equating error near the cut matters most. Chapter 5 adds that retesting people with the same items may bias estimates of change, and equated alternate forms are the fix.
Cut Scores and Standard Setting
Cut scores are where a scale becomes a decision. Chapter 5 says they "embody value judgments" alongside technical and empirical considerations, so setting them is never purely technical. The Standards say cut score; older EPPP material says cutoff score.
Standard 5.21 sets a documentation burden that scales with what the cut does. A cut selecting a fixed number of people, such as applicants to interview, needs little standard-setting documentation. A cut sorting people into substantively distinct categories with no quota, such as pass versus fail, must be documented in detail: the judgments called for and their reliability, panelist selection and qualifications, training, feedback about what provisional judgments implied, any chance to confer, variability across panelists, and where feasible an estimate of how much the cut would move on replication. The same comment ties back to precision: adequate precision where a cut sits is a prerequisite to reliable classification.
Standard 5.22 covers what panelists need: familiarity with the proficiency descriptions, practice judging item difficulty with feedback on their accuracy, the experience of taking a form of the test, and feedback on the pass rates their provisional standards imply. Its comment refers to the scale score that would characterize a borderline examinee. That is the judgmental logic. The empirical logic is Standard 5.23: cuts defining substantively distinct categories should be informed by data on the relation between test performance and relevant criteria, in the score range around the cut. Its comment points to score distributions for criterion groups, and notes that criterion groups such as successful versus unsuccessful practitioners are often unavailable in credentialing.
Your prep material names these logics. Asking panelists what fraction of borderline candidates would answer each item correctly is the Angoff method. Comparing distributions for a known-competent and a known-incompetent group is the contrasting groups method. Locating the score that characterizes a candidate at the boundary is borderline group reasoning. The Standards never print those names, so no Standard number attaches to them.
In licensure the cut trades two errors. A false positive licenses someone who lacks the skill and exposes the public to harm; a false negative denies a license to a qualified candidate. Shrink one and you grow the other.
Picture a smoke alarm you can tune. Turn it up and it screams at burnt toast until someone pulls the battery. Turn it down and it sleeps through a real fire. There is no setting with zero errors, only a choice about which error you would rather live with, and a better sensor moves both numbers at once.
Chapter 11 criticizes two practices. A legislature that fixes a cut at something like 70 percent correct does harm for two reasons: nobody has looked at how the items relate to practice, and nobody has checked how hard those items actually are, so the number floats free of both. A credentialing body that adjusts its passing score to regulate how many people enter the profession follows what Chapter 11 calls a "questionable procedure" that threatens the validity of the passing score as an indicator of entry-level competence. Rescaling so a fixed number reach passing is technically inappropriate too.
When several tests feed one decision, the combination rule is a design choice whose basis must be articulated: a compensatory model is additive, so a high score can offset a low one, while a conjunctive model requires acceptable performance on each test in the series. Chapter 11 also says credentialing validation rests mainly on content-related evidence built from an analysis of the actual work performed, through job analysis, the critical incident technique, or practice surveys; criterion-related evidence has limited applicability, because a credentialing exam is not predicting performance in one specific job. And subscores handed to failing candidates deserve caution, since they often rest on few items and differences between them may be pure measurement error.
Revision, Obsolescence, and Stale Scores
Standard 4.24 says specifications should be revised when new research data, significant changes in the domain, or newly recommended conditions of use may reduce the validity of score interpretations. Age alone is not a reason. The comment adds a rule that catches many clinicians: keep an older edition after a newer one is published and you must show the older version is as appropriate for that use.
Standard 4.25 reserves a word. The label revised belongs only to a test whose specifications changed in a meaningful way, and a new form built from unchanged specifications does not qualify. Chapter 11 adds that when major revisions occur or the scale changes, the cut score has to be reestablished.
Scores expire too. The comment to Standard 6.14 says scores that no longer reflect the test taker's current state generally should not be used, except for research. The exception is real: older scores stay useful for longitudinal assessment or for tracking deterioration of function, so the question is valid use, not age. Chapter 10 gives the clinician's version: guard against relying on outdated scores, and retest when they are.
Using Scores the Way the Standards Require
Chapter 9 opens with a line worth memorizing: a test's name is never adequate grounds for choosing it. What does inform it is validity evidence for the intended use, reliability/precision, the fit of the norms, and the likely consequences.
Standard 9.4 puts the burden where it belongs. Use a test for a purpose with little or no validity evidence and you own the job of documenting the rationale and gathering reliability and validity evidence, and any hypotheses must be labeled tentative. Standard 10.5 says the same about populations: with no normative or validity studies for a population like your test taker, interpretations must be hedged and flagged as hypotheses, not findings.
No score stands alone. Standard 9.13 says a score should not be interpreted in isolation, and names the alternative explanations for a low one: low motivation, limited fluency in the language of the test, limited opportunity to learn, unfamiliarity with cultural concepts the items rest on, and perceptual or motor impairments. Standard 10.15 adds that diagnostic interpretation must rest on multiple sources of test and collateral information.
Standard 10.6 reaches into your assessment lessons. For differential diagnosis, pick tests shown to tell the competing diagnoses apart from each other. Sorting patients from healthy people is a different and easier job, and it is not enough. Its comment adds that a significant gap between two group averages does not license a diagnosis for one person. The clinical-group index patterns in "IQ Tests" are group means: research findings the exam tests, not decision rules for the person in front of you. Chapter 10 adds that variability across tests within a battery is common in the general population, so use base rate data to decide whether an observed spread is exceptional.
Computer-based test interpretation (CBTI) gets three standards. Standard 6.11 says the reasoning and evidence behind an automated interpretation must be available, along with what it cannot see, and warns such reports may miss the person's circumstances. Standard 9.10 forbids leaning on an automated report by itself and expects the user to know how it was built, and Standard 10.17 says they must verify the validity evidence is sufficient and the norms relevant. Chapter 6 puts it plainly: automatically generated reports are "not a substitute for the clinical judgment" of someone who actually met the client.
Two smaller duties. Standard 10.13 requires diagnostic terms to be defined and the nomenclature system named, because DSM and ICD can attach different symptom sets to the same word and some syndromes appear in neither. Standard 10.12 requires interpretation to account for other factors that shaped the outcome, effort included, and Chapter 10 adds that continuing an evaluation when effort is clearly low may lead to inappropriate score interpretations.
Validity Generalization
Validity generalization applies test-criterion evidence gathered elsewhere, usually through meta-analysis, instead of running a fresh local study. Chapter 1 explains why it works: test-criterion correlations vary a lot across settings, and much of that variation is statistical artifact, driven by sampling fluctuation, restriction of range, and unreliability in criterion measures.
A strong exam item lives here. Where the meta-analytic database is large, represents the situation you want to generalize to, and yields a clear consistent pattern after correction, a small local study adds little and can be "relatively limited if not actually misleading". A local study is more informative when the meta-analytic base is thin, findings are inconsistent, or your situation looks markedly different.
Transporting evidence directly from another setting requires sound evidence, typically a careful job analysis (a practice analysis in credentialing), that the local job is highly comparable. And under Standard 11.9, one earlier study carries a new setting only if it had a large sample, a relevant criterion, and a situation that closely matches yours. Not never, but not casually.
Name the artifacts. Standard 11.8 requires predictor-criterion studies to identify errors of measurement, restriction of range, criterion deficiency (the criterion misses parts of the job that matter), criterion contamination (something outside the criterion pollutes it), and missing data. Its comment gives the classic contamination case: criterion ratings collected from supervisors who already know the applicants' test scores. Chapter 1 adds sampling error and criterion unreliability.
Common Misconceptions and EPPP Traps
"Reliability is .92, so pass-fail decisions are trustworthy." A high average coefficient can coexist with poor precision at the cut. Ask for the conditional SEM there.
"Decision consistency and decision accuracy are the same thing." Consistency is agreement across two replications. Accuracy is agreement with true status.
"An answer key error hurts reliability." No. Systematic error is construct-irrelevant, so it reduces validity rather than reliability/precision.
"Interrater agreement of .95 means the scores are reliable." It means the raters agree. Task-to-task variation in the examinee's performance is a separate error source.
"Any two linked scores are interchangeable." Only equating claims that, and only for alternate forms built to the same specifications.
"A legislated 70 percent cut is objective, since no judgment is involved." The Standards say such cuts can be harmful, carrying no information about the test-job relationship and none about item difficulty.
"The index profile in the manual confirms the diagnosis." Group means are inadequate for individual diagnosis, and scatter is common in the general population. Check base rates and pick a test with evidence it separates the competing diagnoses.
Memory Aids
Average for the scale, conditional for the cut. Anything a licensing board decides gets judged by the conditional number, which covers Standards 2.14, 5.13, and 5.21 at once.
Consistency asks Again, Accuracy asks Actual. Same classification again on a replication, versus a match with the person's actual true status.
The linking ladder. Strongest to weakest: Equate, Vertically scale, Concord. Only the top rung gives interchangeable scores.
Key Takeaways
-
Define the replication before quoting reliability (2.1, 2.2). Coefficients are not interchangeable (2.6), interrater agreement is not examinee reliability (2.7), test length heads Chapter 2's list.
-
Conditional SEM is error at one score level, required near every cut score (2.14), in reported-score units (2.13). Decision consistency is the same classification on a replication; decision accuracy matches true status (2.16).
-
Differences, profiles, and subscores need their own evidence (2.3, 2.4, 1.14); adjusted coefficients go beside unadjusted ones (1.21, 2.20).
-
Generalizability theory gives variance components and the universe score; the test information function gives precision per trait level; adaptive precision uses simulation (2.5, 4.3).
-
Norms need a described population (5.8) and dates of testing (5.9). Local norms and user norms hold only within their limits, renorming in print is the publisher's job (5.11), and outdated norms are not an outdated test.
-
Linking is the umbrella; only equating gives interchangeable scores. Designs: single group, random equivalent groups, embedded anchor, external anchor. Equating error near the cut matters most (5.13).
-
Cut scores carry value judgments; documentation is heaviest with distinct categories and no quota (5.21). Angoff, contrasting groups, and borderline group are exam names, not Standards terms.
-
Credentialing rests on content evidence from a practice analysis. Compensatory lets scores offset; conjunctive does not.
-
Revision follows new data, domain change, or new use conditions (4.24); revised means changed specifications (4.25); obsolete scores are research or longitudinal only (6.14 comment).
-
The test's name is never enough. Unsupported purposes shift the burden to the user (9.4), diagnosis needs multiple sources (10.15) and evidence separating competing diagnoses (10.6), scatter needs base rates, CBTI never replaces judgment (6.11, 9.10, 10.17).
-
Validity generalization can beat a small local study, and one prior study is rarely enough (11.9). Artifacts: range restriction, criterion unreliability, criterion deficiency, criterion contamination, sampling error.
Fairness and differential item functioning belong to "Fairness and Bias in Testing", the validity framework to "What Tests Measure" and "Can Tests Predict", and these score types to "Interpreting Scores".
