Why Item Analysis and Test Reliability Matter for Your Psychology Career
You've probably taken dozens of tests in your life, from college exams to professional assessments. But have you ever wondered what makes a good test actually good? When you're sitting across from a client, interpreting their depression inventory scores, or reviewing cognitive assessment results, you need to trust that those numbers mean something real and consistent.
That's where item analysis and test reliability come in. These concepts aren't just abstract statistics. They're the foundation that determines whether the tests you use in practice are worth the paper they're printed on. Understanding these principles will help you choose appropriate assessments, interpret scores accurately, and explain results to clients with confidence.
Classical Test Theory: The Foundation
Let's start with the basic framework that underlies most psychological testing: Classical Test Theory (CTT). The core idea is simple but profound. Every score someone gets on a test (we call this the "obtained score") is made up of two parts:
Obtained Score = True Score + Measurement Error
The true score is the person's expected score across an unlimited number of independent repetitions of the same measurement procedure. It is an average we cannot observe directly, not a perfect reading of the person's actual ability (Cappelleri et al., 2014).
Measurement error is all the random stuff that messes with the score but doesn't reflect the person's true ability. It's like when you're cooking and the smoke alarm goes off, distracting you, or when the oven temperature runs hot that day, or you're exhausted from a bad night's sleep. In testing situations, this includes things like:
- The testing room being too hot or noisy
- The person feeling sick that day
- Ambiguous wording in questions
- Random variation in guessing outcomes around the expected score, not all credit earned through guessing (Cappelleri et al., 2014)
- Simple fatigue from a long test
Classical Test Theory assumes that true scores are stable across strictly parallel forms of the same testing procedure. Change the procedure itself (a different form family, different raters, a different occasion) and you have changed what the true score refers to. But measurement error? That's unpredictable and random, changing every time.
Understanding Reliability: What Does "Consistent" Really Mean?
Test reliability tells you how much you can trust that a test gives consistent information. When a test has high reliability, you know that most of what you're seeing in the scores reflects true differences between people, not random error.
Here's the key insight: A reliability coefficient tells you directly what percentage of score variability comes from true differences versus error.
Let's say a depression inventory has a reliability coefficient of .85. This means:
- 85% of the differences you see in people's scores reflect actual differences in depression levels
- 15% is just noise from measurement error
Imagine you're checking your bank account balance on a banking app. If the app is highly reliable, the number you see reflects your actual balance. But if it's unreliable, sometimes it might show you're $100 short because of a glitch, or $50 over because a transaction didn't update. You want that number to be trustworthy when you're deciding whether you can afford something.
What Counts as "Good Enough" Reliability?
The acceptable level depends on what's at stake:
| Test Type | Minimum Acceptable | Why |
|---|---|---|
| Research measures, preliminary screening | .70 | Lower stakes; you're looking at group averages or just getting initial information |
| Clinical diagnosis, personnel selection, high-stakes decisions | .90+ | You're making important decisions about someone's life. You need to be very confident |
| Personality and attitude measures | .70-.80 | These traits are inherently more variable and harder to measure precisely |
| Cognitive ability tests | .80-.90+ | These should be highly reliable given their standardized nature |
These numbers are rules of thumb from the exam literature. The Standards set no fixed minimum; they say the precision you need rises with the stakes of the decision and falls when the decision is reversible or backed by other information.
Four Ways to Measure Reliability
Just like you might evaluate a new restaurant by checking multiple sources (online reviews, asking friends, trying it yourself at different times) there are different ways to evaluate test reliability. Each method answers a different question about consistency.
1. Test-Retest Reliability: Consistency Over Time
This method checks whether people get similar scores when they take the same test twice, with some time in between.
The process:
- Give the test to a group of people
- Wait days, weeks, or months
- Give them the exact same test again
- Calculate the correlation between the two sets of scores
When it's useful: Test-retest reliability matters for characteristics that should be stable over time, like intelligence or enduring personality traits.
When it's not appropriate: Don't use this for things that change rapidly, like current mood or state anxiety. It's like weighing yourself twice to check if your scale is consistent. That works if you're checking the scale's reliability, but not if you're trying to track actual weight loss over time.
2. Alternate Forms Reliability: Consistency Across Different Versions
Some tests have multiple versions (Form A, Form B, etc.) to prevent people from memorizing answers or to allow retesting.
The process:
- Give Form A to a group
- Give Form B to the same group (either immediately or later)
- Correlate the scores
When it's useful: Essential when your test has multiple forms and you need to know they're interchangeable. Also tells you about stability over time if you administer the forms at different times.
Think about taking the driver's license test at different DMV locations. You want to know that passing at one location is just as meaningful as passing at another. That the different versions are equivalent.
3. Internal Consistency Reliability: Do the Items Work Together?
This checks whether all the items on your test are measuring the same thing consistently.
Common methods:
Coefficient Alpha (Cronbach's Alpha): Calculates the average correlation among all items on the test. This is the most widely used method.
Kuder-Richardson 20 (KR-20): A version of coefficient alpha specifically for items scored as right/wrong (dichotomous scoring).
Split-Half Reliability: Divide the test in half (usually odd-numbered items vs. even-numbered items), correlate the two halves, then apply the Spearman-Brown prophecy formula to estimate what the full test's reliability would be.
Why the correction? Because shorter tests are less reliable, split-half gives you the reliability of two half-length tests. The Spearman-Brown formula corrects this to estimate the full test's reliability.
Important limitation: Internal consistency is NOT appropriate for speed tests (where you're measuring how fast someone completes tasks). Why? Because on speed tests, people who get far in the test answer more items correctly, making it look like the items are highly consistent when really you're just measuring speed. For speed tests, use test-retest or alternate forms instead.
4. Inter-Rater Reliability: Consistency Across Scorers
When tests require judgment to score (like evaluating therapy session quality, rating behavioral observations, or scoring projective tests) you need to know that different raters would give similar scores.
Methods:
Percent Agreement: Simple calculation: What percentage of the time do two raters agree? Easy to understand but has a major flaw. It doesn't account for chance agreement.
Imagine two people randomly guessing "yes" or "no" to coin flips. They'd agree about 50% of the time just by chance, even though neither knows what they're doing.
Cohen's Kappa: A more sophisticated measure that corrects for chance agreement. It's used when two raters are assigning ratings on a nominal scale (categories with no inherent order).
Cohen's kappa is (observed agreement - expected chance agreement) / (1 - expected chance agreement). With observed agreement of .75 and chance agreement of .50, kappa is .50. The raters achieved half of the possible agreement beyond chance (McHugh, 2012).
The Danger of Consensual Observer Drift
Here's something that can quietly ruin your inter-rater reliability: when raters talk to each other while rating, they start agreeing more and more with each other. But not necessarily with reality. Their ratings become consistent but inaccurate.
It's like when two friends binge-watch a TV series together and start finishing each other's sentences about what will happen next. They're very consistent with each other, but they might both be completely wrong about the plot.
How to prevent it:
- Keep raters working independently
- Provide thorough training before rating begins
- Regularly check ratings against a gold standard
What Makes Reliability Coefficients Go Up or Down?
Three key factors affect how reliable your test will be:
1. Content Homogeneity
Tests measuring one unified thing tend to be more reliable than tests measuring several different things.
A depression inventory asking only about sadness, crying, and hopelessness will typically have higher internal consistency than a general mental health screening that asks about depression, anxiety, and psychosis all mixed together.
2. Range of Scores (Sample Heterogeneity)
You get higher reliability coefficients when your sample includes people across the full spectrum of whatever you're measuring. High, medium, and low scorers.
Imagine trying to test whether a thermometer is accurate by only measuring temperatures between 70-72 degrees. You'd have trouble seeing if it really works across the full range. But if you measure temperatures from 0 to 100 degrees, you can really see how consistent it is.
If you only test your anxiety measure on people with severe anxiety disorders, you're restricting the range, and your reliability coefficient will be artificially low.
3. Guessing
In classical test theory, obtained score = true score + error. True score is the expected score across repeated administrations under the defined measurement procedure, not simply the number of answers known. Expected credit from guessing can contribute to that expectation; one observed total cannot identify the true score. If the true score is explicitly assumed unchanged, differences in obtained scores belong to the error terms (Cappelleri et al., 2014).
Pure uniform guessing has probability .50 of a correct answer on two options and .25 on four options. Those probabilities alone do not determine a test's reliability; distinguish expected credit from random variation around it (Cappelleri et al., 2014).
From Reliability Coefficient to Reliability Index
Some psychometricians distinguish between these two terms:
- Reliability coefficient (what we've been discussing): The proportion of observed score variance that's due to true score variance
- Reliability index: The theoretical correlation between observed scores and true scores, calculated as the square root of the reliability coefficient
For example, if a test has a reliability coefficient of .81, the reliability index would be √.81 = .90
You won't use this much in practice, but you might see it on the exam.
Item Analysis: Building Better Tests
When you're creating a new test, you need to figure out which items are worth keeping. That's where item analysis comes in. You're looking at two key characteristics for each item.
Item Difficulty (p)
The p-value tells you what percentage of test-takers got the item right.
Calculation: Number who answered correctly ÷ Total number of test-takers
Example: If 60 out of 100 people answered an item correctly, p = 60/100 = .60
Interpreting p-values:
- p = 1.0: Everyone got it right (very easy)
- p = .50: Half got it right (moderate difficulty)
- p = 0: No one got it right (very difficult)
What's optimal? For most tests, you want moderately difficult items (p = .30 to .70). Why? Because very easy or very hard items don't help you tell people apart.
Special cases:
Mastery tests: These are designed to check whether someone has achieved a specific level of competence. Think of a licensing exam where you need to verify someone knows at least 80% of critical safety information. For these, you want p-values matching the mastery level (e.g., p = .80 for an 80% mastery test).
Accounting for guessing: A common classroom heuristic targets the midpoint between 1.0 and the guessing probability. This arithmetic shortcut is not a universal optimum for every test purpose.
- Four-option multiple choice: guessing probability = .25, so the heuristic gives p = (1.0 + .25)/2 = .625
- True/false: guessing probability = .50, so the heuristic gives p = (1.0 + .50)/2 = .75
Item Discrimination Index (D)
The D-value tells you whether an item successfully distinguishes between high-performing and low-performing test-takers.
Calculation:
- Identify the top 27% of test-takers (based on total test scores)
- Identify the bottom 27% of test-takers
- Calculate: D = (% of high scorers who got it right) - (% of low scorers who got it right)
Example: If 85% of high scorers and 40% of low scorers answered correctly: D = .85 - .40 = .45
Interpreting D-values:
- D = +1.0: All high scorers got it right, no low scorers did (perfect discrimination)
- D = 0: Equal percentages of high and low scorers got it right (no discrimination)
- D = -1.0: All low scorers got it right, no high scorers did (something's wrong!)
What's acceptable? Generally, you want D ≥ .30
Important connection: Item difficulty affects discrimination. Moderately difficult items can discriminate better because they give both groups room to show differences. It's like trying to judge running ability, if you give everyone a 10-foot race, even slow runners finish quickly, so you can't tell them apart. A 5K race better shows the differences.
Standard Error of Measurement: Embracing Uncertainty
Here's a truth that makes testing more honest: When a test's reliability is less than perfect (which is always), you can't be certain that someone's obtained score is their true score.
The standard error of measurement (SEM) quantifies this uncertainty.
Formula: SEM = SD × √(1 - reliability coefficient)
Where SD is the test's standard deviation.
Example calculation:
- Test has SD = 15 and reliability = .91
- SEM = 15 × √(1 - .91)
- SEM = 15 × √.09
- SEM = 15 × .3
- SEM = 4.5
Confidence Intervals: Showing the Uncertainty
Rather than reporting a single score, we often report a confidence interval that shows the range where someone's true score likely falls.
Approximate normal-error rules: These bands assume a constant SEM and normally distributed measurement errors. About 95% uses 1.96 SEM, often rounded to 2; about 99% uses 2.58 SEM. Three SEM gives about 99.7%, not 99%. These are repeated-sampling coverage statements under the model (Cappelleri et al., 2014; Greenland et al., 2016).
- 68% confidence interval: Obtained score ± 1 SEM
- 95% confidence interval: Obtained score ± 2 SEM
- 99.7% confidence interval: Obtained score ± 3 SEM
Example: Someone scores 110 on an IQ test with SEM = 5.
| Confidence Level | Calculation | Range |
|---|---|---|
| 68% | 110 ± (1 × 5) | 105-115 |
| 95% | 110 ± (2 × 5) | 100-120 |
| 99.7% | 110 ± (3 × 5) | 95-125 |
Think about this when you're explaining test results to a client. Instead of saying "Your IQ is 110," you might say "Based on this test, we're 95% confident your IQ falls between 100 and 120." It's more honest and prevents over-interpretation of small score differences.
Item Response Theory: The Modern Alternative
Item Response Theory (IRT) represents a more sophisticated approach to test development that's increasingly important, especially for computerized testing.
Key Differences from Classical Test Theory
| Classical Test Theory | Item Response Theory |
|---|---|
| Test-based (focuses on total scores) | Item-based (focuses on individual items) |
| Sample-dependent (item statistics change with different groups) | Sample-invariant (item properties stable across groups) |
| Less suitable for adaptive testing | Ideal for adaptive testing |
| Simpler to understand and calculate | More complex but more powerful |
One caution on the invariance row: IRT item parameters stay stable across groups only when the model fits and the item is free of differential item functioning (DIF). That is an assumption to check, not a guarantee.
The Core Idea: Item Characteristic Curves
IRT examines how each item relates to the underlying trait (the latent trait) you're measuring. Ability, depression severity, extroversion, etc.
For each item, you create an Item Characteristic Curve (ICC) that shows:
- X-axis: Level of the trait (low to high)
- Y-axis: Probability of endorsing or answering correctly (0 to 1.0)
Three Item Parameters
Depending on which IRT model you use (one-, two-, or three-parameter), the ICC tells you:
1. Difficulty Parameter (b): In the Rasch model for right-or-wrong items, ability equal to b gives a 50% chance of a correct answer. The meaning of b depends on the model, so do not apply this rule to every IRT model (Cappelleri et al., 2014).
- Items on the left side of the graph: easier (endorsed by people with lower trait levels)
- Items on the right side: harder (only endorsed by people with higher trait levels)
2. Discrimination Parameter (a): How well does this item distinguish between people just above and just below the difficulty level?
- Indicated by the slope of the curve
- Steeper slope = better discrimination
- A highly discriminating item is like a precise filter that clearly separates people right at a certain skill level, while a poorly discriminating item is like a cloudy lens that can't quite distinguish between similar ability levels.
3. Guessing Parameter (c): In the three-parameter logistic model, this is the lower asymptote of the item curve.
- It describes the limiting response probability at very low ability.
- It is not generally the value at ability zero, where the curve crosses the vertical axis. (Cheng & Liu, 2015)
Why IRT Matters for Practice
Computerized Adaptive Testing: IRT makes it possible to give each person a customized test that adjusts to their ability level.
Instead of everyone taking the same 100-question test, imagine a system that starts with medium-difficulty questions, then adapts: if you get them right, it gives you harder ones; if you get them wrong, it gives you easier ones. You end up with a shorter, more efficient test that's tailored to your level.
This is how the GRE and many modern licensing exams work. It's only possible because IRT lets us precisely calibrate each item's difficulty and discrimination properties.
Differential Item Functioning: Compare Equal Ability
Differential item functioning (DIF) means an item works differently across groups after the groups are matched on the ability or trait being measured. A difference in the groups' overall average scores is not enough to show DIF.
Imagine two runners with the same speed facing different hurdles. DIF asks whether one test item acts like the taller hurdle. It does not start by assuming the runners have equal speed just because they joined the same race.
DIF is a signal to investigate, not automatic proof of bias. Review the item's content and the reason for the difference before deciding whether it adds an unfair barrier.
Common Misconceptions
"A reliability of .70 means the test is 70% accurate." Not quite. It means that 70% of the variance in scores is due to true differences, and 30% is error. It doesn't directly tell you about accuracy (which relates to validity, not reliability).
"If a test has high reliability, it must be measuring what it claims to measure." Wrong. A test could consistently measure the wrong thing. Reliability is necessary but not sufficient for validity. A bathroom scale might consistently give you the same reading every time (reliable), but if it's always 10 pounds off, it's not accurate (not valid).
"Internal consistency reliability is always the best method to use." No. It's inappropriate for speed tests and for tests measuring multiple different constructs. Choose your reliability method based on what the test measures and how it will be used.
"Small differences in scores are meaningful if the test is reliable." Not necessarily. Always consider the standard error of measurement. The difference between two scores has its own standard error, and it is larger than either score's SEM because both scores carry error. Test manuals report the standard error of the difference for this reason, and the Standards (2.4) ask for separate reliability evidence before you interpret subtest gaps or profile peaks and valleys.
Practice Tips for Remembering
The Reliability Coefficient Mnemonic: Remember "T/T+E" for the reliability coefficient concept:
- True score variance on Top
- True score variance plus Error variance on the bottom
- Reliability = T/(T+E)
The Four Reliability Types: Use "TAII" (sounds like "tie"):
- Test-retest
- Alternate forms
- Internal consistency
- Inter-rater
Confidence Intervals: Remember "1-2-3 for 68-95-99":
- 1 SEM = 68%
- 2 SEMs = 95%
- 3 SEMs = 99.7%
Item Difficulty: "p stands for percent who passed" (got it right)
Item Discrimination: "D = Difference" (between high and low scorers)
Key Takeaways
-
Differential item functioning compares item responses at matched ability; a DIF flag calls for review and does not by itself prove bias.
-
Classical Test Theory says obtained scores = true scores + error; reliability tells you what proportion is true score variance
-
Reliability coefficients range from 0 to 1.0 and are interpreted directly as the percentage of variance due to true differences (not error)
-
Four reliability methods: test-retest (consistency over time), alternate forms (consistency across versions), internal consistency (items measuring the same thing), and inter-rater (consistency across scorers)
-
Internal consistency is inappropriate for speed tests. Use test-retest or alternate forms instead
-
Higher reliability comes from: homogeneous content, unrestricted score range, and less susceptibility to guessing
-
Standard error of measurement quantifies score uncertainty; use it to create confidence intervals (±1 SEM = 68%, ±2 SEM = 95%, ±3 SEM = 99.7%)
-
Item difficulty (p) = proportion who answered correctly; optimal range is typically .30-.70
-
Item discrimination (D) = difference between top 27% and bottom 27%; want D ≥ .30
-
Item Response Theory is item-based (not test-based), produces sample-invariant parameters, and enables computerized adaptive testing
-
IRT uses Item Characteristic Curves to show difficulty, discrimination, and guessing parameters for each item
Understanding these concepts will help you select appropriate tests, interpret scores responsibly, and explain results clearly to clients and colleagues. When you see a test manual in practice, you'll know exactly which reliability information to look for and how to interpret what you find.
Choose the reliability evidence for the decision
Classical test theory separates an observed score into a true score and random error. The true score is an expected score over repetitions of a defined procedure. Reliability concerns consistency, not proof that the intended construct is measured. Test-retest evidence concerns stability over time; inter-rater evidence concerns scoring across raters; internal consistency concerns relations among items. Alpha can rise when related items are added, and a high alpha does not establish a single factor. Reliability estimates depend on the scores and population studied. Do not treat a reliability coefficient as a person's percentage correct or as the percentage of each individual score that is true. (Cappelleri et al., 2014; Tavakol & Dennick, 2011; Weir, 2005)
Read an item curve with its model in mind
In a dichotomous Rasch model, ability equal to item difficulty gives a .50 probability of a correct answer. An item's difficulty is a location on the ability scale, not the percentage correct in every sample. In a two-parameter logistic model, discrimination concerns the steepness of the item curve. The three-parameter logistic model adds a lower asymptote, often called a guessing parameter. It is not generally the curve's value at ability zero. Item information describes precision at a given ability level; information from items adds to test information under the model. A test can provide more information in one ability region than another. Adaptive testing uses item properties and estimated ability to choose informative items, rather than assuming every examinee needs the same questions. (Cappelleri et al., 2014; Jones, 2019; Cheng & Liu, 2015)
True score is an expected score
In classical test theory, obtained score = true score + error. True score is the expected score across repeated administrations under the defined measurement procedure, not simply the number of answers known. Expected credit from guessing can contribute to that expectation; one observed total cannot identify the true score. If the true score is explicitly assumed unchanged, differences in obtained scores belong to the error terms (Cappelleri et al., 2014).
Reliability requirements follow the decision
Evaluate precision for the intended use, population, and consequences of error. High-stakes decisions about an individual require close attention to uncertainty and classification accuracy near the cutoff. A coefficient such as .70 or .90 is not a universal permission rule, and reliability alone does not establish validity for diagnosis or hiring (Cheng et al., 2015; Strauss & Smith, 2009).
Flag-resolution sources
- Joseph C Cappelleri, J Jason Lundy, Ron D Hays. Overview of classical test theory and item response theory for the quantitative assessment of items in developing patient-reported outcomes measures. Clinical therapeutics. 2014 May;36(5):648-62. doi:10.1016/j.clinthera.2014.04.006. PMID:24811753. https://pubmed.ncbi.nlm.nih.gov/24811753/
- Sander Greenland, Stephen J Senn, Kenneth J Rothman, et al. Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations. European journal of epidemiology. 2016 Apr;31(4):337-50. doi:10.1007/s10654-016-0149-3. PMID:27209009. https://pubmed.ncbi.nlm.nih.gov/27209009/
