Resources / 7: Research Methods & Statistics / Overview of Inferential Statistics

Maya draws a bell curve on a whiteboard while the bunny sits above the peak

Overview of Inferential Statistics

7: Research Methods & Statistics

Study guide by Anders Chan, PsyD · Updated

Introduction: Why Inferential Statistics Matter for Your Psychology Career

You've collected data from your research study. Now comes the million-dollar question: Are your results real, or did they happen by chance? This is where inferential statistics come in, and honestly, this might be one of the most practical skills you'll use throughout your career. Whether you're evaluating treatment outcomes, reading research to guide your practice, or conducting your own studies, you need to know if what you're seeing is a genuine effect or just random noise.

Inferential statistics quantify uncertainty under sampling and model assumptions. Recruiting people from a few clinics does not prove that they represent everyone with a disorder. Describe the population to which the design can reasonably apply, and consider selection differences before extending conclusions (Handley et al., 2018).

The Foundation: Sampling Distributions and Probability

Let's start with a core concept that trips up many students: the sampling distribution of means.

Here's what you need to understand: Imagine you're taste-testing coffee to determine the average quality at a new café. You take one sip on Monday morning. Then another visit on Tuesday afternoon. Then Wednesday evening. Each time, the coffee tastes slightly different. Not because the café's quality wildly changes, but because of random variations in how hot it is, which beans were used that batch, whether you just ate something sweet, and so on.

A sampling distribution works the same way. If you could magically draw hundreds of different samples from the same population and calculate the mean for each one, you'd get a distribution of those means. Some would be higher than the true population mean, some lower, not because anything actually changed, but because of sampling error. Random variation that happens when you select different people.

The brilliant news? You don't actually need to draw hundreds of samples. This is where the central limit theorem becomes your best friend.

The Central Limit Theorem: Three Critical Predictions

For independent observations from the same population with finite variance, the sampling distribution of the mean becomes approximately normal as sample size grows. This describes repeated sample means, not the raw scores. Strong skew may require larger samples; neither 30 nor 100 guarantees a good normal approximation. The mean remains centered on the population mean, and its standard error is sigma divided by the square root of n (Kwak & Kim, 2017).

  1. Shape: Under the stated independent-sampling and finite-variance conditions, sample means increasingly approximate a normal distribution. There is no universal sufficient sample size (Kwak & Kim, 2017).

  2. Center: The mean of all those sample means equals the actual population mean. In other words, on average, samples give you the right answer.

  3. Spread: The standard deviation of this sampling distribution (called the standard error) equals the population standard deviation divided by the square root of your sample size.

That third point is huge for understanding why bigger samples are better, the standard error gets smaller as sample size increases, meaning your sample means cluster more tightly around the true population mean.

Setting Up Your Statistical Test: Hypotheses

Before you run any statistical test, you need to set up two competing hypotheses:

The null hypothesis (H₀): This is the skeptical position. It says your independent variable has NO effect on your dependent variable. Any differences you observe are just due to chance or sampling error.

The alternative hypothesis (H₁): This is usually what you actually believe. That your independent variable DOES have a real effect.

Think of it like a trial in the legal system. The null hypothesis is "innocent until proven guilty." You start by assuming there's no effect (innocence), and you need strong enough evidence to reject that assumption.

For example, if you're testing whether a new therapy reduces anxiety:

  • Null hypothesis: The therapy has no effect on anxiety levels
  • Alternative hypothesis: The therapy does reduce anxiety levels

Making Decisions: The Four Possible Outcomes

When you run your statistical test, you'll make a decision: either retain (keep) the null hypothesis or reject it. Since we're dealing with probability, not certainty, you could be right or wrong. Here's how it breaks down:

Your DecisionTrue State of RealityResult
Retain null hypothesisNull is actually true✓ Correct decision
Reject null hypothesisNull is actually false✓ Correct decision (you detected a real effect!)
Reject null hypothesisNull is actually true✗ Type I Error (false positive)
Retain null hypothesisNull is actually false✗ Type II Error (false negative)

Type I Error: The False Alarm

A Type I error happens when you reject a true null hypothesis. You've declared something works when it actually doesn't, like announcing you've found your soulmate on the third date, only to realize six months later there's no real connection.

The probability of making a Type I error is alpha (α), also called your significance level. Researchers typically set this at .05 or .01 before collecting data. If α = .05, you're accepting a 5% chance of a false positive. If α = .01, you're being more conservative with only a 1% chance.

Type II Error: The Missed Opportunity

A Type II error happens when you retain a false null hypothesis. Something really does work, but you failed to detect it, like dismissing a great job candidate because they had a bad interview day, missing out on someone who would have been excellent.

The probability of making a Type II error is beta (β). Unlike alpha, you don't directly set beta. However, you can reduce it by increasing your study's statistical power.

Statistical Power: Your Ability to Detect Real Effects

Statistical power is the probability that you'll correctly reject a false null hypothesis. In other words, your ability to detect a real effect when one exists. Think of it as the sensitivity of your study.

Power is affected by five main factors:

1. Alpha Level

Increasing alpha increases power. If you use α = .10 instead of α = .05, you're more likely to detect effects. However, this also increases your Type I error risk, which is why researchers rarely go above .05.

2. Effect Size

For a fixed test, alpha, and target effect, increasing sample size or reducing unexplained outcome variation generally increases power. Increasing residual variability while holding the mean difference and sample size fixed generally reduces power. Test choice must fit the design and assumptions; an intervention change does not guarantee a larger effect (Whitley & Ball, 2002).

3. Sample Size

This is the most practical factor you can control. Larger samples increase power substantially. It's like trying to hear a whisper in a noisy restaurant versus a quiet library, the "signal" (real effect) becomes clearer when you reduce the "noise" (random variation), and bigger samples do exactly that.

4. Type of Statistical Test

Choose a test that fits the outcome, study design, and assumptions. Parametric methods can be efficient when their assumptions fit, but they are not always more powerful than nonparametric methods. Changing the test is not an automatic way to gain power (Vrbin, 2022).

5. Population Homogeneity

For a fixed mean difference, sample size, and test, less unexplained outcome variation generally increases power. Researchers can influence the sample and design, although a narrower sample can limit generalization (Whitley & Ball, 2002).

A Modern Alternative: Bayesian Statistics

Everything we've discussed so far falls under frequentist statistics, the traditional approach that's dominated psychology for decades. But there's an alternative approach gaining popularity: Bayesian statistics.

The Fundamental Difference

Frequentist approach: Your current study's data is everything. You calculate whether results could have happened by chance.

Bayesian approach: You combine what's already known (previous research) with your current data to get updated knowledge. It's cumulative and collaborative.

How Probability is Defined

Here's where things get philosophically interesting:

Frequentist probability: If you could rerun your study infinite times under identical conditions, probability tells you how often you'd get certain results. For example, a p-value of .03 means that if the null hypothesis were true and you repeated the study many times, you'd get results this extreme or more only 3% of the time.

Bayesian probability: This represents your degree of certainty or belief that something is true. It's more intuitive and subjective.

Confidence Intervals vs. Credibility Intervals

This difference shows up clearly in how we interpret intervals:

95% Frequentist Confidence Interval: If we repeated this study many times and calculated a confidence interval each time, 95% of those intervals would contain the true population mean. (Note: We CANNOT say there's a 95% chance the true mean is in THIS specific interval. But people often make this mistake.)

95% Bayesian Credibility Interval: There's a 95% probability that the true population mean falls within this specific interval. This is actually how most people intuitively want to interpret confidence intervals!

The Bayesian Process: Prior, Likelihood, and Posterior

Bayesian analysis uses three components:

ComponentDefinitionSource
PriorYour probability distribution for a parameter BEFORE seeing new dataPrevious research, expert opinion, or theoretical assumptions
Likelihood FunctionThe observed-data likelihood evaluated at different parameter values (Introna et al., 2022)Your actual collected data
PosteriorThe updated probability distribution combining prior and likelihoodMathematical synthesis using Bayes' theorem

The posterior becomes your final answer. And could serve as the prior for the next study, creating a knowledge-building chain.

Advantages of Bayesian Statistics

  1. Incorporates existing knowledge: You're not treating each study as if it exists in a vacuum
  2. More intuitive interpretation: Credibility intervals mean what people think they mean
  3. Direct hypothesis testing: You can directly assess evidence for your research hypothesis, not just against the null
  4. User-friendly software: Programs like JASP make Bayesian analysis accessible

The Major Criticism

The prior is subjective. Two researchers could use different priors for the same study and reach different conclusions. It's like two movie critics watching the same film. One comes in having loved the director's previous work (positive prior), another hated it (negative prior), and they walk out with different overall impressions despite seeing identical content.

Critics argue this subjectivity undermines the objectivity that science requires. Supporters counter that making assumptions explicit (choosing a prior) is more honest than pretending frequentist methods don't also involve subjective choices.

Publication Bias: When the Literature Lies

Publication bias occurs when whether a study gets published depends on its results. Studies with significant, positive findings are far more likely to be published than studies with null results. That means the published literature systematically overestimates effects.

The File Drawer Problem

Rosenthal (1979) named the core issue: studies that "fail" don't get submitted or accepted. They sit in the researcher's file drawer while the flashy significant results make it into journals. Imagine judging a casino by only interviewing the winners walking out. You'd conclude gambling is very profitable.

Why this matters clinically: meta-analyses combine published studies. If the null studies are missing, the meta-analysis inherits the bias and can make a weak treatment look strong.

Detecting Publication Bias

Funnel plot: Scatter plot of each study's effect size (x-axis) against its precision, usually sample size or standard error (y-axis). Large studies cluster tightly near the true effect at the top; small studies scatter widely at the bottom, forming an inverted funnel.

  • Symmetric funnel = no evidence of publication bias
  • Asymmetric funnel (typically a missing chunk of small, null-result studies in one bottom corner) = publication bias is likely

Caveat worth knowing: asymmetry can also come from real heterogeneity or chance, so a funnel plot suggests bias, it doesn't prove it.

Egger's test: A regression-based statistical test for funnel plot asymmetry. More formal (and more powerful) than eyeballing the plot.

Correcting for Publication Bias

Trim-and-fill and fail-safe N are sensitivity checks, not proof that publication bias is absent. A small trim-and-fill shift means little changed under that procedure, which does not retrieve actual missing studies. A fail-safe N is a hypothetical study count under specified assumptions, not a count of known unpublished studies. Different bias methods can disagree, so compare methods and their assumptions (Soeken & Sripusanapan, 2003).

Fail-safe N estimates a hypothetical null-study count under a specified method. A benchmark such as 5k + 10 is a rule of thumb, not proof that bias is absent (Soeken & Sripusanapan, 2003).

Memory hook: Funnel plots and Egger's test assess asymmetry. Trim-and-fill and fail-safe N explore sensitivity. None proves that bias is absent (Soeken & Sripusanapan, 2003).

Common Misconceptions to Avoid

Misconception 1: "A p-value tells you the probability that the null hypothesis is true."

Reality: A p-value tells you the probability of getting your results (or more extreme) IF the null hypothesis were true. It's backward from what people think.

Misconception 2: "Statistical significance means practical importance."

Reality: With large enough samples, even tiny, meaningless effects can be statistically significant. Always consider effect size, not just p-values.

Misconception 3: "Failing to reject the null hypothesis proves there's no effect."

Reality: It means you didn't find sufficient evidence of an effect. The effect might exist but be too small for your study to detect (Type II error).

Misconception 4: "Confidence intervals can be interpreted as probability ranges."

Reality: Not for frequentist confidence intervals! That interpretation only works for Bayesian credibility intervals.

Misconception 5: "Bigger alpha always means better studies."

Reality: Increasing alpha increases power but also increases Type I error risk. It's a trade-off.

Practice Tips for Remembering

For Type I and Type II Errors: Create a simple reference table and memorize it. On exam day, quickly sketch it out if you get confused:

Error TypeWhat HappenedMemory Trick
Type IRejected true null (false positive)"I thought there was an effect, but I was wrong"
Type IIRetained false null (false negative)"II (two) can mean 'missed it too'"

For increasing power: Remember the acronym ASETH:

  • Alpha (increase it, though rarely done)
  • Sample size (increase it)
  • Effect size (larger effects easier to detect)
  • Test type (use parametric when possible)
  • Homogeneity (more homogeneous populations help, though you can't control this)

For Bayesian components: Think PLP. Prior, Likelihood, Posterior. It's like updating your GPS route: you start with a planned route (prior), get current traffic data (likelihood), and generate an updated route (posterior).

For the Central Limit Theorem: Remember "SCS". Shape (approaches normal under the stated conditions), Center (equals population mean), Spread (standard error = SD/√n).

Key Takeaways

  • Inferential statistics help us determine if research results reflect real effects or just sampling error by comparing sample values to sampling distributions

  • The central limit theorem tells us that sampling distributions of means approach normality under independent sampling and finite variance as sample size increases, center on the population mean, and have standard error equal to SD/√n

  • The null hypothesis states no effect exists; the alternative hypothesis states an effect exists

  • Type I error (false positive): rejecting a true null hypothesis, with probability = alpha (typically .05 or .01)

  • Type II error (false negative): retaining a false null hypothesis, with probability = beta

  • Statistical power (probability of correctly detecting real effects) increases with: larger alpha, larger effect sizes, larger samples, parametric tests, and more homogeneous populations

  • Frequentist statistics (traditional approach) analyzes current data using probability as long-run frequency

  • Bayesian statistics (alternative approach) combines prior knowledge with current data using probability as degree of belief

  • Bayesian credibility intervals can be interpreted as probability ranges, but frequentist confidence intervals cannot (despite common misinterpretation)

  • The Bayesian process uses a prior (previous knowledge), likelihood function (current data), and produces a posterior (updated knowledge)

  • Main criticism of Bayesian statistics: subjectivity in choosing priors can lead different researchers to different conclusions from the same data

  • Publication bias (the file drawer problem) inflates published effects; assess possible asymmetry and sensitivity with several methods; no one procedure proves absence of bias (Soeken & Sripusanapan, 2003)

Understanding these concepts isn't just about passing the EPPP. It's about being able to critically evaluate research throughout your career and make evidence-based decisions in your practice. Every treatment outcome study, assessment validation study, and meta-analysis you read will use these principles. Master them now, and you'll be a more informed, effective psychologist for decades to come.

Read the estimate before the significance label

A p-value describes how incompatible the data are with a specified model, including the null hypothesis. It is not the probability that the null is true or that chance caused the result. A confidence interval shows effect values compatible with the data under the analysis assumptions. A 95% confidence procedure covers the fixed population value in 95% of repeated samples under those assumptions. For a mean difference, zero is the no-difference value; for a risk ratio it is one. A result that misses significance does not prove no effect. Statistical significance does not show clinical importance. Compare an interval with a stated meaningful-effect threshold, not just zero. A smaller p-value alone does not show a larger effect. Cohen's d for independent groups divides a mean difference by a pooled standard deviation. Its sign depends on subtraction order. An omnibus ANOVA tests the joint equal-means hypothesis; further comparisons are needed to locate differences. (Greenland et al., 2016; Lakens, 2013; Bewick et al., 2004)

Plan precision and avoid overclaiming

Power is the chance of rejecting the null when a specified alternative is true. With the design and other assumptions fixed, more participants generally raise power and narrow confidence intervals. Detecting a smaller effect, choosing a lower alpha, or seeking higher power usually requires more participants. A nonsignificant small study can leave both useful and negligible effects plausible. More observations do not by themselves remove systematic bias or make a convenience sample representative. Random assignment concerns allocation to conditions, while population generalization also depends on who was sampled and the study setting. (Whitley & Ball, 2002; Greenland et al., 2016; Handley et al., 2018; Slater & Hasson, 2025)

Limits of the central limit theorem

For independent observations from the same population with finite variance, the sampling distribution of the mean becomes approximately normal as sample size grows. This describes repeated sample means, not the raw scores. Strong skew may require larger samples; neither 30 nor 100 guarantees a good normal approximation. The mean remains centered on the population mean, and its standard error is sigma divided by the square root of n (Kwak & Kim, 2017).

Power depends on the design

For a fixed test, alpha, and target effect, increasing sample size or reducing unexplained outcome variation generally increases power. Increasing residual variability while holding the mean difference and sample size fixed generally reduces power. Test choice must fit the design and assumptions; an intervention change does not guarantee a larger effect (Whitley & Ball, 2002).

Inference does not prove representativeness

Inferential statistics quantify uncertainty under sampling and model assumptions. Recruiting people from a few clinics does not prove that they represent everyone with a disorder. Describe the population to which the design can reasonably apply, and consider selection differences before extending conclusions (Handley et al., 2018).

Publication bias sensitivity checks

Trim-and-fill and fail-safe N are sensitivity checks, not proof that publication bias is absent. A small trim-and-fill shift means little changed under that procedure, which does not retrieve actual missing studies. A fail-safe N is a hypothetical study count under specified assumptions, not a count of known unpublished studies. Different bias methods can disagree, so compare methods and their assumptions (Soeken & Sripusanapan, 2003).

Prior likelihood and posterior

A prior distribution expresses uncertainty about a parameter before the current data. The likelihood evaluates the observed data at different parameter values; it is not itself a probability distribution over the parameter. Bayes' rule combines prior and likelihood to obtain the posterior distribution. Posterior probabilities about an effect depend on the data, prior, and model (Introna et al., 2022).

Effect sizes across studies

Meta-analysis combines study effect estimates to summarize evidence across studies. Effect sizes express the magnitude and direction of a finding on an appropriate scale. A study's p value alone does not describe its effect magnitude (Lakens, 2013).

Flag-resolution sources

  • Colleen M Vrbin. Parametric or nonparametric statistical tests: Considerations when choosing the most appropriate option for your data. Cytopathology : official journal of the British Society for Clinical Cytology. 2022 Nov;33(6):663-667. doi:10.1111/cyt.13174. PMID:36017662. https://pubmed.ncbi.nlm.nih.gov/36017662/

Ready to practice?

Get started in the app