The p-value is the most used and most misread number in statistics. It appears in every study you skim, decides which drugs reach market and which A/B test ships, and most explanations of it, including many in textbooks, are subtly wrong. The correct definition fits in one sentence. The skill, and the fun, is in what you can and cannot do with it.
The one-sentence definition
A p-value answers exactly this: if there were truly no effect, how often would random chance alone produce a result at least as extreme as yours? That is the whole thing. A p-value of 0.03 means that in a world where nothing is going on, data like yours shows up about 3 percent of the time.
Now notice what it is not. It is not the probability your hypothesis is true. It is not the probability the result happened “by chance.” And 1 minus p is not the probability you are right. The p-value interrogates the data assuming the boring world; it never assigns a probability to the boring world itself. Keeping that direction straight puts you ahead of a depressing share of published research commentary.
Where the number comes from
Your data first gets compressed into a test statistic, usually a z-score or t-statistic: a count of how many standard deviations your observed result sits from what the no-effect world expects. (If standard deviation itself is shaky ground, start with the averages guide and the standard deviation calculator; the p-value is built directly on that foundation.) The p-value is then the tail area of the sampling distribution beyond your statistic: the fraction of chance outcomes even more extreme than the one you observed.
A fully worked example
You run an A/B test on a signup page: 1,000 visitors see version A and 50 sign up (5.0%); 1,000 see version B and 65 sign up (6.5%). Is B really better, or did luck deal B a good crowd?
Step one, pool the data as the skeptic would: overall rate 115/2,000 = 5.75%. Step two, compute the standard error of the difference between two proportions at that pooled rate: about 1.04 percentage points. Step three, the statistic: the observed 1.5-point difference divided by 1.04 gives z = 1.44. Step four, the tail area beyond 1.44 in both directions: p = 0.15.
Reading: in a world where A and B convert identically, a gap this large appears about 15 percent of the time, roughly a one-in-seven shuffle of luck. That is suggestive, not conclusive, and the standard 0.05 convention says keep testing. Notice what the p-value gave you: not a verdict about B, but a calibrated measure of how surprised you are entitled to be.
Enter a z-score or t-statistic, choose one or two tails, and see the p-value with the exact tail area shaded on the distribution curve.
z or t: which statistic you actually have
The z distribution assumes you know the population’s spread, which in practice means large samples, above roughly 30, where the estimate is stable. Small samples estimate the spread from the same scarce data, which adds uncertainty, so they use the t distribution: same bell shape, fatter tails, controlled by degrees of freedom (sample size minus one). Fatter tails mean the same statistic earns a larger p-value: t = 2.228 with 10 degrees of freedom gives p = 0.05, where a z of 1.96 sufficed. As samples grow, t melts into z. The practical rule: small sample, unknown spread, reach for t.
One tail or two
A two-tailed test asks whether the result differs from zero in either direction and is the honest default. A one-tailed test asks about one direction only and, mechanically, halves the p-value. That halving is exactly why choosing one tail after peeking at the data is cheating: it converts p = 0.08 into p = 0.04 by rewording the question the answer already suggested. Decide the question before computing the answer, and treat any one-tailed result in the wild with an eyebrow raised.
What 0.05 actually is
The 0.05 threshold is a convention, not a law of nature: a false-positive budget of one in twenty that science largely inherited from an influential statistics textbook a century ago. Two consequences follow. First, p = 0.049 and p = 0.051 are nearly identical evidence wearing opposite labels; treating the line as a cliff is a ritual, not reasoning. Second, the budget spends itself across attempts: run twenty independent tests of true nothings and one will clear 0.05 by luck alone. Testing many outcomes and reporting the winner manufactures discoveries, which is why serious work pre-registers its hypotheses or corrects for multiple comparisons, and why a lone flashy result with a barely-significant p deserves patience rather than headlines. This same tail-area machinery, incidentally, powers percentile systems everywhere, from AP exam equating to growth charts: cutoffs are the p-value’s respectable cousins.
Significant is not the same as important
The p-value measures surprise, not size. With a big enough sample, a tiny, worthless effect becomes statistically significant; with a small sample, a large, real effect can miss the cutoff. Always read the effect size next to the p-value: how big is the difference, in units a human cares about? A weight-loss pill with p = 0.001 and an average loss of half a pound is a rounding error with excellent publicity. Confidence intervals bundle both questions, giving a range of plausible effect sizes instead of a binary stamp: “B improves conversion by somewhere between minus 0.5 and plus 3.5 points” says more than “p = 0.15” ever will.
A field guide to misreadings
| Claim you will hear | Status |
|---|---|
| “p = 0.03, so there is a 97% chance the effect is real” | Wrong: p says nothing about the probability of hypotheses |
| “Not significant, so there is no effect” | Wrong: absence of evidence is not evidence of absence, especially in small samples |
| “Significant, so it matters” | Wrong: check the effect size |
| “A smaller p means a bigger effect” | Wrong: it can just mean a bigger sample |
| “p = 0.15 means keep collecting data if the question matters” | Fair, if planned honestly rather than peeking until significance appears |
Power: the error nobody budgets for
Significance testing manages two failure modes. A false positive (Type I) is crying wolf, and the 0.05 threshold is its budget. A false negative (Type II) is missing a real wolf, and its budget is set by statistical power: the probability your study detects an effect of a given size if it exists. Convention aims for 80 percent power, yet underpowered studies remain everywhere, and they fail twice: they usually miss real effects, and when one sneaks past the threshold anyway, its measured size is inflated, a phenomenon called the winner’s curse. It is why dramatic small-sample findings so often shrink on replication. Before trusting any exciting result, ask the unglamorous question: was this study big enough to find what it claims, at the size it claims?
Bayes in one paragraph: base rates change everything
The p-value deliberately ignores how plausible your hypothesis was before the data, and that omission is where intuition faceplants. Screen a rare condition affecting 1 in 1,000 people with a test that false-alarms 5 percent of the time: among 1,000 people you expect roughly 1 true positive and 50 false ones, so a positive result is overwhelmingly likely to be wrong despite the impressive-sounding “95 percent” machinery. The same logic governs surprising research claims: an implausible hypothesis with p = 0.04 is still probably false, because 0.04 measured the data’s surprise, not the idea’s credibility. Extraordinary claims need better than borderline p-values; that is not cynicism, it is arithmetic.
Reading a study’s statistics in sixty seconds
A usable checklist for the wild. One: sample size and who was sampled, since n = 40 undergraduates generalizes to n = 40 undergraduates. Two: effect size in human units, not just stars in a table. Three: a confidence interval, and whether its boring end would change your decision. Four: how many outcomes were tested, and whether the headline was the plan or the survivor. Five: preregistration or replication anywhere in sight. A paper that passes all five deserves your attention; a headline that hides all five deserves your patience.
What a p-value says about causation: nothing
Significance is silent on why. A rock-solid p = 0.0001 correlation between ice cream sales and drownings is real, and its cause is summer. Observational data can be significant because A causes B, B causes A, or C causes both, and the p-value cannot tell the three apart; only design can, which is why randomized experiments carry the weight they do: randomization severs the hidden third variable. When you read “significantly associated,” translate it honestly as “probably not a coincidence,” and hold the word “causes” until someone shows you the experiment or a very good natural substitute for one.
A one-minute history of 0.05
The threshold traces to Ronald Fisher’s 1925 textbook, where one-in-twenty was offered as a convenient line for judging experiments; Neyman and Pearson later reframed testing as controlling long-run error rates, and the two philosophies fused into the ritual practiced today. A century on, the replication crisis pushed back: some journals now emphasize confidence intervals and effect sizes over verdicts, some methodologists argue for a stricter 0.005 default for new discoveries, and preregistration has moved from novelty to expectation in serious empirical work. The number survives because it is useful shorthand, but the field’s direction of travel is unmistakable: estimate sizes, state uncertainty, and stop worshiping the line.
Standard deviation vs standard error: the mix-up that breaks intuition
Two similarly named numbers do opposite jobs, and confusing them is the most common quiet error in reading statistics. Standard deviation describes the spread of individual data points: how much people, conversions, or measurements differ from each other, and it does not shrink as you collect more of them. Standard error describes the precision of your estimate of the average, and it equals the standard deviation divided by the square root of the sample size, which is why it does shrink as n grows. The test statistic divides your observed effect by the standard error, so quadrupling the sample halves the error and inflates the same real-world effect into a bigger z. That is the entire mechanism behind “large studies find significant everything”: the world did not change, the denominator did. When a paper reports error bars, your first question is which of the two they show, because standard error bars are routinely mistaken for the spread of the data and make findings look far tighter than the underlying humans ever were.
Frequently asked questions
What exactly does p = 0.05 mean?
In a no-effect world, results at least this extreme occur 5 percent of the time. Nothing more, and the “at least this extreme” clause is load-bearing.
Why 0.05 and not 0.01 or 0.10?
History and habit. Fields with expensive false positives (particle physics, genomics) demand far stricter thresholds; exploratory work sometimes tolerates looser ones. The threshold is a policy choice about which error you fear more.
Can a p-value prove there is no effect?
No. A large p means the data failed to distinguish the effect from noise, which is also what happens when the sample is simply too small.
Does a one-tailed test ever make sense?
When only one direction is possible or decision-relevant and that was fixed before the data arrived. Declared afterward, it is a discount coupon for significance.
How big a sample do I need?
Enough that the effect size you care about would produce a detectable statistic: that is a power calculation, and doing it before collecting data is what separates an experiment from a fishing trip.
Run your own statistic through the p-value calculator to see the shaded tail for yourself, use the percentage guide to translate effects into plain terms, and the rest of the math calculators handle the supporting arithmetic.