AIDA · Course page
AI and Data Analytics · Nayanta University

Could this just be luck?

The null hypothesis and the p-value, built up from one small example. No formulas to memorise. The whole idea fits on this page.

Where it starts

A claim, and a reason to doubt it

A bus company advertises that its buses arrive on time 80% of the time. Dev takes that bus every day and thinks the real number is lower. So he watches five buses picked at random. Three arrive late, two on time.

That is 40% on time in his sample, against a claim of 80%. Dev wants to tell the company its claim is wrong.

Before he can, he has to deal with an annoying possibility. Even if the company is telling the truth, five buses is a small number, and a run of three late buses can happen by chance. Bad luck looks exactly like a false claim when you only have five observations.

Toss a coin five times. Getting four heads is not unusual, and it does not mean your coin is loaded. How many heads would you need before you started suspecting the coin?

Everything on this page is a way of answering that last question with a number instead of a feeling.

Step 1

Write down the boring explanation

There are two stories that fit what Dev saw:

The boring one

The company is telling the truth. Buses really are on time 80% of the time, and Dev happened to catch an unlucky five.

The interesting one

The company is wrong. The real on-time rate is below 80%, and Dev's sample is showing him that.

The null hypothesis is the boring explanation, written down precisely enough to calculate with. Here it is: buses arrive on time 80% of the time.

The null hypothesis is almost always some version of "nothing special is going on here". The new teaching method works as well as the old one. The medicine does nothing. The two groups are the same. The company's claim is true.

You never set out to prove the null hypothesis. You set out to see whether your data is embarrassing for it.

Notice that the null hypothesis is a statement about the world, not about Dev's five buses. It says something about all the buses the company has ever run and will ever run. Dev only gets to see five of them.
Step 2

Assume the boring explanation, then count the luck

Here is the move that the whole subject rests on. Pretend, for a moment, that the company is telling the truth. Then ask: if the claim were true, how often would five random buses give a result at least as bad as what Dev saw?

To answer that, replace the buses with something you can hold. If buses are on time 80% of the time, that is the same as a bag of ten cards with eight red (on time) and two black (late). Draw a card, note the colour, put it back, and repeat five times. Dev's result of three late buses is the same as drawing three or more black cards.

Ten cards: 8 red, 2 black. Draw five with replacement.
Probability of getting 3 or more black = 5.8%

You can work that out by hand, or ask ChatGPT for it, or run five hundred trials with actual cards if you have a free afternoon. The number is the same. Roughly one time in seventeen, an honest company with an 80% record will produce a five-bus sample this bad.

The p-value is the probability of getting a result at least as extreme as the one you got, assuming the null hypothesis is true. For Dev, the p-value is 5.8%.

That last clause in italics is the part people drop, and dropping it is the source of nearly every mistake made with p-values. Read the definition twice.

  1. Write down the null hypothesis (buses are 80% on time).
  2. Collect data (five buses, three late).
  3. Assuming the null hypothesis is true, work out how likely a result this extreme or worse would be. That is the p-value (5.8%).
  4. If the p-value is small, the data is hard to explain away as luck, and you have grounds to reject the null hypothesis.
Step 3

The 5% line

How small does the p-value have to be? By convention, most fields draw the line at 5%. A p-value below 0.05 is called statistically significant and is treated as grounds for rejecting the null hypothesis.

Dev's p-value is 5.8%. That is above 5%, so on the standard convention he cannot reject the company's claim. Three late buses out of five is a bit unusual for an honest company, and not unusual enough.

The 5% line is a convention, not a law of nature. Ronald Fisher suggested it in the 1920s as a rough working rule and said so at the time. Particle physicists use a far stricter threshold before announcing a new particle. Medicine sometimes uses 1%. There is nothing magic about 0.05, and a study with p = 0.049 tells you the same thing as one with p = 0.051.

What Dev can honestly say is: "my sample is consistent with the company's claim, though it does not confirm it either". What he cannot say is that the claim is true. Five buses is far too little evidence to conclude anything much in either direction.

"We failed to reject the null hypothesis" and "we proved the null hypothesis" are completely different statements. Not finding evidence of a difference is not the same as finding evidence of no difference. Most of the time it just means the study was too small.
Step 4

The same result, with more buses

Suppose Dev keeps watching. Every time, exactly 40% of the buses are on time, which is the rate he saw at the start. Only the number of buses changes.

Buses watchedLateOn timep-valueVerdict at 5%
5340%5.8%Cannot reject
10640%12.1%Cannot reject
201240%3.2%Reject the claim
503040%0.09%Reject the claim
1006040%0.0004%Reject the claim

The observed on-time rate never changes. The p-value collapses anyway, because a run of bad luck can produce 40% out of five buses quite easily and cannot produce 40% out of a hundred.

Two things follow, and both of them matter when you read someone else's analysis.

A big study finds small things

With enough data, almost any tiny difference becomes statistically significant. A dating app testing a button colour on ten million users can get p < 0.001 for an effect that changes nothing anyone cares about.

A small study misses big things

With five buses, even a company with a terrible record might escape detection. A high p-value from a small sample tells you almost nothing.

The p-value answers "could this be luck?". It says nothing about "is this difference big enough to care about?". Those are separate questions and you have to ask the second one yourself. Look at the size of the difference, in the units of the thing being measured, before you decide whether a significant result is interesting.

The 10-bus row is a nice trap: watching twice as many buses made the p-value worse (12.1% against 5.8%). Six late out of ten is proportionally the same as three out of five, but the arithmetic of "at least this extreme" does not move smoothly for tiny samples. Do not read too much into any single small-sample p-value.

The part that gets tested

Five things a p-value does not mean

Dev's p-value is 5.8%. Every one of these readings of that number is wrong.

Wrong readingWhy it is wrong
"There is a 5.8% chance the company's claim is true." The p-value assumes the claim is true and works forward from there. It cannot also be the probability that the claim is true. It is not a statement about the claim at all, it is a statement about data.
"There is a 5.8% chance my result was a fluke." Same error in different clothes. 5.8% is how often an honest company would produce data this bad. It is not the chance that this particular result is a fluke.
"p above 0.05, so the claim is correct." Failing to reject is not proving. Dev's five buses are consistent with an 80% company and also with a 60% company and a 45% company. He does not have enough data to separate them.
"p below 0.05, so the effect is large." A small p-value says the effect is probably real. It says nothing about how big it is. Large samples give small p-values for trivial effects.
"p = 0.058 and p = 0.049 lead to opposite conclusions." Those two numbers are almost the same. Treating 0.05 as a cliff edge is a habit of reporting, not a fact about the evidence.
How results get faked without lying

Testing twenty things and reporting one

A 5% threshold means that if the null hypothesis is true, you will still get a "significant" result one time in twenty. That is the price of the rule. It becomes a problem when someone runs many tests and reports only the winner.

Independent tests runChance at least one comes out "significant" by luck alone
15%
523%
1040%
2064%
4087%

Run twenty comparisons on a dataset where nothing at all is going on, and you have a 64% chance of finding something publishable. Nobody has to fabricate data for this to happen. Slicing by gender, then by city, then by age band, then by income, then dropping a few odd-looking rows, is enough. This is called p-hacking, and it is the main reason a lot of published research does not replicate.

The defence is to say what you were looking for before you look. If you test twenty things, report twenty results, and expect one of them to be spurious.

This one is directly relevant to how we work in this course. When you ask ChatGPT to "find something interesting in this data", it may quietly try many comparisons and hand you the one that looks best. You are the analyst, the AI is the typist. Ask it what else it tested.
Use this

What to ask when you read "statistically significant"

Find a news story this week that reports a study. Answer the five questions above from the article alone. If you cannot answer three of them, that tells you something about the story.

Notes and sources