Why Fairness Sits Inside Your Psychometrics Knowledge
KN29 ends with the words most candidates skim: test fairness and bias. That is a mistake. Fairness questions are psychometric questions wearing different clothes.
The 2014 Standards for Educational and Psychological Testing treat fairness as a validity issue running through every stage of development and use. The question is never whether a test is fair in the abstract, but whether your score interpretation, for your use, holds up for the person in front of you.
The other lessons built the machinery: What Tests Measure, Items and Reliability, Can Tests Predict, Interpreting Scores. This lesson asks one more question: does any of it work the same way for everybody?
The Two Ways a Test Goes Wrong
Chapter 1 names two failure modes, and Chapter 3 uses them to explain every fairness threat it discusses.
Construct underrepresentation means the test misses important parts of what it claims to measure, so scores mean less than the label promises. The Standards' example is an anxiety measure that only asks about physiological reactions, ignoring emotional, cognitive, and situational components.
Construct-irrelevant variance means something outside the construct is pushing scores around. On a reading test, familiarity with the passage topic. On a math test given to English language learners, a heavy reading load. On an anxiety scale, a tendency to underreport.
Picture measuring a room with two bad tape measures. The first is missing its last three feet, so you never capture the whole room. The second is full length but made of elastic, so the number moves for reasons that have nothing to do with the room. Both give a wrong answer, in opposite directions.
The comment to Standard 3.0, the overarching standard of the chapter, says the central idea of fairness is to "identify and remove construct-irrelevant barriers" so scores can be compared and interpreted for all examinees. When subgroup differences appear, Standard 3.6 asks you to investigate both underrepresentation and irrelevant variance as possible causes.
Consequences count only when they trace back
Standard 1.25 is a trap in disguise. When test use produces unintended consequences, you investigate whether they arise from the test's sensitivity to something it was not meant to assess, or from its failure to fully represent the construct.
Chapter 1's example: different hiring rates caused by a real, job-relevant difference in the skill measured do not make the interpretation invalid. The same gap caused by a sophisticated reading test used for a job needing only basic literacy traces to construct-irrelevant variance, so the interpretation is invalid even though scores correlate with job performance. A consequence tracing to neither failure mode may matter for policy, but it is not validity evidence.
A keying error hurts validity, not reliability
Systematic error, such as an incorrect answer key, consistently raises or lowers scores for all test takers or for some subset, and is unrelated to the construct. Chapter 2's point: it is generally not in the standard error of measurement and is not a shortfall in reliability. It is construct-irrelevant, so it lowers validity while leaving reliability untouched. Note the hinge: systematic error hitting everyone equally is construct-irrelevant variance, while systematic error hitting one group harder is what the Standards call bias.
What Fairness Means in the Standards
The Standards define fairness as the validity of score interpretations, for the intended uses, for individuals from all relevant subgroups. Fairness belongs to an interpretation, for a use, for a person, not to the test. A relevant subgroup is one identifiable in a way that matters for interpreting scores: race, ethnicity, gender, culture, language, age, disability, socioeconomic status.
Chapter 3 comes at fairness from four directions.
Equitable treatment during the testing process. Uniform directions, specified time limits and room arrangements, proctors, consistent procedures, so administration differences do not quietly change some people's scores. Examinees on older or slower equipment are disadvantaged for reasons unrelated to the construct.
Lack of measurement bias. Covered next.
Access to the construct as measured. Contrast the knowledge and skills that make up the construct with what a person needs simply to respond. A personality inventory in standard print is partly measuring visual acuity for a test taker with impaired vision. Idiomatic phrases and unfamiliar stimulus contexts do the same thing quietly.
Validity of the individual score interpretation. We generalize across groups for convenience, not because the groups are homogeneous. A clinician may be justified in deviating from standardized procedure to measure one person's standing more accurately, while in selection the same deviation would change the construct, break comparability, and unfairly advantage some candidates.
What fairness is not
The Standards explicitly exclude one common public meaning: fairness as equality of testing outcomes across subgroups. Group mean differences do not on their own show that a testing application is biased. The comment to Standard 3.6 says they should trigger follow-up studies to find the causes, not a verdict. Most real cases mix genuine differences with bias, and a search that comes up empty is reassurance, not proof.
Measurement Bias: DIF, DTF, and Predictive Bias
Measurement bias means construct underrepresentation or construct-irrelevant components that affect different groups' scores differently, damaging the validity of interpretations for those groups. It takes three forms.
Differential item functioning (DIF) occurs when equally able test takers differ in their probability of answering an item correctly as a function of group membership. Equally able is the load-bearing phrase: you compare people matched on the trait, or on an appropriate criterion, then ask whether the item still behaves differently.
Think of two sprinters with identical times in every previous race. Put them in adjacent lanes and one suddenly runs three seconds slower. Because you already knew they were equally fast, the slow time is information about the lane, not the runner. Match first, compare second.
DIF starts an investigation, it does not end one. Chapter 3 says a suitable, substantial explanation is needed before calling the item biased. Chapter 1 adds that DIF is "not always a flaw or weakness." Items sharing a feature can function differently across groups because the test genuinely is multidimensional, which may be unexpected or may match the test framework.
Differential test functioning (DTF) is the same idea at the test level. Individuals from different groups with the same standing on the characteristic assessed do not have the same expected test score.
Predictive bias is systematic under-prediction or over-prediction of criterion performance for a group defined by characteristics not relevant to that performance. Differential prediction is how you test for it. Standard 3.7 says use regression when sample sizes are sufficient and prior evidence or theory suggests it: compare slopes and intercepts between two targeted groups, or examine systematic deviations from a common regression line across any number of groups.
Now the trap. A correlation coefficient provides inadequate evidence for or against differential prediction when the groups have unequal means and unequal variances on the test and the criterion. Differential prediction is about where the line sits, and only slopes and intercepts describe that, which is why the validity coefficients in Can Tests Predict cannot be recycled here.
The forms of bias are independent: a predictor test can show no significant DIF and still show group differences in regression lines when predicting a criterion.
Small subgroup samples are not permission to skip the question. Use expert judgment and sensitivity review, focus groups, small-scale tryouts, and cognitive labs, which are think-aloud studies of how test takers work through a task. Operational results can also be accumulated until the sample is large enough.
Accessibility and Universal Design
Accessibility is the degree to which items let as many test takers as possible show their standing on the target construct without being impeded by features irrelevant to that construct. The Standards call it a bias issue, because obstacles to access produce different score meanings for different groups.
Universal design builds accessibility in from the first draft instead of patching it on later. Standard 3.1 lists the ingredients: define constructs precisely, keep instructions simple and clear, and avoid test characteristics that compromise valid interpretation for particular groups.
The example the Standards give of a characteristic to avoid is inappropriate test speededness. If working fast is not part of the construct, a needlessly tight time limit adds construct-irrelevant variance. Other moves named in the chapter: user-selected font sizes, avoiding culturally unfamiliar contexts, and minimizing linguistic load when the construct is not language.
A building with a ramp designed into the entrance has one entrance. A building with steps and a ramp bolted on years later has two, and one of them is around the back. Both let people in. Only one was designed so that getting in was never a separate question.
Universal design is not a guarantee, and no design makes every test workable for every test taker. That is where adaptations come in.
Adaptations: Accommodation Versus Modification
Adaptation is the umbrella term: any change in content, format, or administration conditions made to increase accessibility for people who would otherwise face construct-irrelevant barriers. Whether it changes the construct is the entire question.
Accommodation is an adaptation that leaves the construct intact, so scores stay comparable to the standard form. Text magnification for a visual impairment. A braille form. A native-language glossary of non-technical words on a safety test.
Modification is an adaptation that changes the construct, so scores differ in meaning from the original. The Standards' example is a calculator for a test taker with dyscalculia: if the construct is broader mathematics skill, it gives partial access to something worth measuring, but computation is no longer measured. Comparability is the dividing line, and it must be shown with evidence.
Glasses and a friend reading the eye chart out loud are not the same kind of help. Glasses let you show how well your corrected vision works, which is what the optometrist wanted to know. A friend reading the chart gives you a perfect score on a test of something else.
When extra time is an accommodation and when it is not
If speed is not part of the intended construct, extended time removes an irrelevant barrier and is an accommodation. If speed is part of the construct, extended time causes construct underrepresentation, because the score no longer contains the speed component the test was built to capture. Some court reporter jobs require working quickly, so speed cannot be adapted away.
Three situations where an adaptation is not appropriate
- The characteristic is part of the focal construct. A work sample for a customer service job requiring fluent English does not get translated.
- The test's purpose is to diagnose that condition. Extra time on a test built to detect distractibility and slowed processing speed makes it impossible to tell whether those difficulties exist.
- The person does not need it. Membership in a broad class does not create a need.
A modified test is a new test
The comment to Standard 3.9 is blunt: a modified assessment should be treated as newly developed, meeting the standards for validity, reliability and precision, and fairness on its own. A modification that changes the construct invalidates the norms, and criterion-referenced interpretation goes with them: pass and fail decisions and categories such as basic, proficient, or advanced, set with cut scores from the original test, are not valid on the modified one. That detaches a score from everything Interpreting Scores teaches.
Standard 3.10 adds the piece people forget: accommodations themselves must be standardized, with documented rules for who is eligible, precise administration instructions, and records of what was used. Delivered inconsistently, an accommodation becomes a fresh source of construct-irrelevant variance.
One distractor to know: the Americans with Disabilities Act uses accommodation and modification differently from the Standards, so answer exam items using the Standards' definitions.
Flagging
A flag marks a score as possibly not comparable, usually because it came from a modified test. Chapter 3 describes three situations, and only one calls for a flag. With clear evidence that regular and altered scores are not comparable, consider flagging, to the extent the law allows. With credible evidence they are comparable, "flagging generally is not appropriate." With no credible evidence either way, the field has little agreement, and the direction is to go collect it. Flagging tracks evidence about comparability, not the presence of an adaptation.
Language and Translation
Translating a test is not the same as producing an equivalent one. Translation alone does not ensure the new version matches the original in content and difficulty, or yields equally reliable and valid scores.
Standard 3.12 asks for the methods used to establish the adequacy of the adaptation plus evidence for validity. It also asks something surprising: a Spanish version used with Central American, Cuban, Mexican, Puerto Rican, South American, and Spanish populations should be evaluated with each group separately where feasible. Translation is also not automatically the right accommodation, since translated content tests do not work unless test takers were instructed in the language of the translation.
Standard 3.13 sets the default: test in the language the person is most proficient in. Two exceptions. When the purpose is to determine proficiency in a particular language, that language is the construct. And when the most proficient language is not the one the material was taught in, the language of instruction may be better.
Standard 3.14 covers interpreters. An interpreter should follow standardized procedures and be fluent in the test's language and content and in the examinee's native language and culture, including technical vocabulary. The part that shows up in exam items: unless a test was standardized and normed with interpreters, their use may itself be an alteration that changes the construct, since a third party has entered the room and the protocol has changed.
One last distinction. Conversational fluency is not test fluency: someone who seems fluent in casual English may be slower or less competent on tests demanding English comprehension and literacy. Cognitive Tests names the KABC-II and Raven's SPM for minimizing cultural and linguistic load. That is the defensible claim: they reduce construct-irrelevant cultural and linguistic demands, while fairness itself still rests on evidence for a subgroup and a use.
Using Scores Fairly
Standard 3.18 is the one to memorize. In testing individuals for diagnostic or special program placement purposes, scores should not be the sole indicators of a person's functioning, competence, attitudes, or predispositions. Use multiple sources, consider alternative explanations, and bring in the judgment of someone familiar with the test.
Opportunity to learn is one of those alternative explanations: the extent to which a person has been exposed to the instruction, language, or majority culture the test assumes. The comment to Standard 3.18 gives the clinical version: a recent immigrant with little prior schooling may not have learned concepts an ability measure treats as common knowledge, even when the test is given in that person's native language.
Standard 3.19 gives the educational version. Where the same authority is responsible for both the curriculum and the high-stakes decision, examinees should not suffer permanent negative consequences if evidence indicates they were not given the opportunity to learn the tested content. The low scores may accurately reflect what students know, so the interpretation is not necessarily biased. The unfairness is penalizing people for content an authority failed to teach. The standard does not apply where different authorities control curriculum and testing, as in college admissions.
Standard 3.15 says developers who claim a test can be used with a specific subgroup must provide the supporting evidence plus cautions against foreseeable misuse. A line in a manual is a claim, and claims need backing.
Standard 3.20 covers choosing between tests. When two ways of measuring a construct are equal in construct representation and validity, users should consider subgroup differences in mean scores, or in the percentage above the cut score, when deciding which to use. It is one factor among several, since cost, testing time, and logistics still count. But when two tests give equally valid interpretations and impose similar burdens, legal considerations may require choosing the one that minimizes subgroup differences.
The Vocabulary Bridge
The EPPP still writes items in the older testing vocabulary, and our other lessons teach it on purpose because you will meet it. Learn both wordings and the mapping, because an item can be written either way.
| Older exam term | Standards term | What the switch changes |
|---|---|---|
| Types of validity (content, construct, criterion) | Sources of validity evidence | One concept, several kinds of evidence |
| Culturally fair test | Reduces construct-irrelevant cultural and linguistic load | Fairness rests on evidence for a subgroup and a use, not a label |
| Divergent validity | Discriminant evidence | Same statistic, different label |
| Test bias | Measurement bias, as DIF, DTF, or predictive bias | One word, three things investigated separately |
| Cultural loading | Construct-irrelevant variance | Names the mechanism rather than the content |
Answer inside whichever framework the stem uses.
Common Misconceptions and EPPP Traps
Trap: a group scores lower, so the test is biased. Mean differences trigger investigation, not a conclusion.
Trap: DIF proves an item is biased. DIF needs a substantial explanation first, and it can reflect intended multidimensionality.
Trap: no DIF means no bias. The forms are independent, so clean DIF results can coexist with different regression lines across groups.
Trap: a validity coefficient tells you whether prediction is fair. Correlations are inadequate when groups differ in means and variances. Use regression.
Trap: extra time is always an accommodation. Only when speed is not part of the construct. Otherwise it is a modification producing construct underrepresentation.
Trap: a modified test still gives you a percentile and a pass or fail. It invalidates the original norms and cut scores, so treat it as a new test.
Trap: flag every adapted score. Flag only with clear evidence of non-comparability.
Trap: an answer key error hurts reliability. It is systematic and construct-irrelevant, so it lowers validity, stays out of the standard error of measurement, and leaves reliability alone.
Memory Aids
Missing or extra. Underrepresentation means the test measures too little of the construct, irrelevant variance too much of something else. Sort every scenario into missing or extra first.
Modification modifies the construct. The word tells you the answer. It changes what is measured, so score meaning, norms, and cut scores go with it. Accommodation does not.
DIF matches first, then compares. If a scenario compares raw pass rates without matching people on ability, it is not DIF.
Key Takeaways
-
Fairness is the validity of score interpretations, for the intended use, for individuals from all relevant subgroups. It belongs to an interpretation, not to a test.
-
Every fairness threat reduces to construct underrepresentation or construct-irrelevant variance, and consequences count as validity evidence only when they trace back to one of those two. Systematic error such as a keying error lowers validity but not reliability, and stays out of the standard error of measurement.
-
The four views: equitable treatment, lack of measurement bias, access to the construct, and validity of the individual interpretation. Fairness is not equal outcomes.
-
DIF means equally able test takers differ in probability of a correct response by group, and it needs a substantive explanation to count as bias. DTF is the same idea at the test level. Predictive bias is systematic over-prediction or under-prediction of a criterion, tested with regression, never a correlation when groups differ in means and variances. The forms are independent.
-
Accommodation keeps the construct and preserves comparability. Modification changes it, invalidates the original norms and cut scores, and creates a new test needing its own validation. Flag only with clear evidence of non-comparability.
-
An adaptation is inappropriate when the characteristic is part of the construct, when the test exists to diagnose that condition, or when the person does not need it.
-
Translation is not equivalence. Test in the most proficient language unless proficiency is the construct, and treat an interpreter as a possible change to it.
-
Scores are never the sole indicator for diagnosis or placement. Weigh opportunity to learn.
When a fairness item appears, do not reach for a social or legal frame. Ask what the construct is, then whether the problem is missing construct or extra noise.
