The Core Logic: Real Processes Leave Mechanisms Behind
Fake-data detection is not mind reading. It is model checking. We ask: if the claimed data collection process were true, what patterns should naturally appear? Then we compare that expectation with the numbers, file structure, images, and metadata we actually see.
Name the process
Was this a survey, experiment, spreadsheet merge, image capture, accounting system, or sensor log?
Predict the texture
Real processes create rounding, missingness, outliers, duplicates, timestamps, and measurement limits.
Measure the mismatch
Use histograms, digit counts, exact arithmetic, row order, similarity search, or simulation.
Look for one story
The best evidence is not one odd p-value. It is several clues pointing to the same data-generating story.
Stay careful
A red flag means "needs explanation." It becomes stronger when raw records or metadata agree.
Example 1: The Insurance Odometer Dataset
In a famous dishonesty study, car-insurance customers were asked to report odometer readings. Data Colada analyzed the posted dataset and argued that the field-experiment data were fabricated. The case is powerful for teaching because it combines three techniques: distribution shape, terminal digits, and near-duplicates.
What the example looked like
The key outcome was "miles driven": updated odometer mileage minus an older baseline mileage. For a real set of drivers, we expect many moderate values, some low values, and a smaller number of very high values. People differ in commuting, job, age of car, city, and time between reports.
Data Colada reported that the miles-driven values looked roughly uniform from 0 to 50,000 and then stopped abruptly. In plain language: about as many people appeared to drive 45,000-50,000 miles as 5,000-10,000 miles, and nobody exceeded the cap.
Why the technique works
A histogram is an argument about mechanism. Real driving mileage is produced by human travel behavior,
so it should have behavioral structure. A spreadsheet random-number generator such as
RANDBETWEEN(0,50000) produces a flat rectangle: every interval is about equally likely,
and no value can exceed the maximum. If the observed distribution has the shape of the generator
rather than the shape of driving behavior, that is a serious clue.
Mileage distribution: expected vs found pattern
Teaching reconstruction, percentages by mileage bandA real-driver distribution should usually taper. A random generator capped at 50,000 creates a much flatter shape and an abrupt stop.
What the example looked like
Humans often round large self-reported numbers. If asked for an odometer reading from memory, many people write values ending in 00, 000, 500, or another neat ending. In the insurance case, the older baseline readings showed rounding, but the later experimental readings reportedly did not.
Why the technique works
Terminal digits answer a simple question: did these numbers pass through human hands? If real people estimate large values, final digits are not random. If a computer generates integers uniformly, final digits are close to random, and neat endings are not special. A mismatch between two fields that should have the same reporting behavior is especially informative.
- Count values ending in
000,500,00, and each final digit. - Compare fields that should have similar reporting behavior.
- Ask whether the difference is explained by the collection method, not only by the numbers.
Terminal endings: human-rounded vs computer-like
Illustrative rates for large self-reported readingsThe red bars are not suspicious because they are small; they are suspicious because the companion field had human rounding while this one looked generated.
What the example looked like
Data Colada also described a font clue in the spreadsheet: some baseline values appeared in Calibri and others in Cambria. The two font groups had unusually similar records. Their account was that one group appeared to be copied from the other and slightly altered.
Why the technique works
Exact duplicates are easy to find, so fabricated rows are often changed slightly. Near-duplicate detection searches for records that are not identical but are implausibly close across many columns. One pair of similar customers can happen. Many paired records, especially separated by a visible file artifact such as font, formula, or sort order, are much harder to explain as chance.
The general method is: standardize comparable columns, compute distances between rows, find unusually close pairs, then ask whether those pairs also share metadata such as font, timestamp, source file, or row block.
Near-duplicate pairs by row distance
Expected chance matches should be sparseA cluster of many tiny distances is the visual clue: copied-and-adjusted rows create more near twins than independent records should.
Source: Data Colada #98 on the auto-insurance field experiment
Example 2: Out-of-Order Rows In A Behavioral Experiment
Data Colada's 2023 posts on studies co-authored by Francesca Gino are useful for teaching a different point: spreadsheet order can be data. Their "Clusterfake" example focused on rows that were almost sorted by condition and participant ID, except for a small group of duplicated or out-of-sequence observations.
What the case was about
In the experiment, participants solved math puzzles and could over-report how many they solved. The study manipulated whether an honesty pledge was signed at the top or bottom of a form. Data Colada described posted data that appeared sorted by experimental condition and participant ID, with a small cluster of rows that broke the otherwise orderly pattern.
The technique: row-order forensics
Row order is not always meaningful. But in spreadsheets exported from a study, row order often reflects entry order, participant ID order, a sort operation, or a merge. Once a file is mostly sorted by known keys, rows that violate the sort order become informative. They may mark later insertions, manual edits, copy-paste blocks, or merge errors.
Why it works
A clean sort creates a mathematical constraint: within each condition, participant IDs should move in one direction. If only the rows that drive the published result violate that constraint, the anomaly is not just cosmetic. It links file history to the substantive conclusion.
Row order: expected sorted sequence vs found sequence
Illustrative participant ID traceWhen a spreadsheet is otherwise sorted, a small local break can reveal late insertions or manual edits. The visual question is not "is it perfectly sorted?" but "are the exceptions substantively important?"
Source: Data Colada #109 and Data Colada #114. Gino has denied committing fraud and filed a defamation lawsuit; the classroom point here is the statistical and file-forensic logic of the analysis.
Technique: Benford's Law
Benford's law is a first-digit test. It says that in many datasets spread across several powers of ten, the first digit is not uniform: 1 appears about 30.1% of the time, while 9 appears about 4.6%.
Why the formula looks like this
Imagine looking only at numbers between 1 and 10 after repeatedly multiplying or dividing by 10 until each value falls in that interval. Numbers with first digit 1 occupy the interval from 1 to 2. Numbers with first digit 9 occupy the interval from 9 to 10. On an ordinary ruler both intervals have length 1. On a log scale, the interval from 1 to 2 is much wider than the interval from 9 to 10.
That log-scale width is the probability. For digit d, the interval is from
d to d + 1, so the probability is:
First digit frequencies: expected vs found
Flat fake sampleA fabricated uniform sample makes digits 7-9 far too common and digit 1 far too rare.
Technique: GRIM, Or "Can This Mean Exist?"
GRIM stands for Granularity-Related Inconsistency of Means. It checks whether a reported mean could have come from whole-number responses with the stated sample size.
Why it works
If 17 students answer a 1-7 Likert question, the total score must be a whole number. The mean is
total divided by 17. That means the only possible means are 17/17, 18/17,
19/17, and so on. After rounding, many decimal values are impossible.
Example: a reported mean of 4.33 with n = 17 implies a total of
4.33 x 17 = 73.61. No group of 17 whole-number responses can sum to 73.61. The more
precise question is whether any integer sum could round to 4.33 at the reported decimal places.
Possible rounded means near the report
CheckingGreen ticks are means that can be made from whole-number totals. The red marker is the reported mean.
Example 3: Elisabeth Bik And Image Duplication
Some fake or unreliable data are visual rather than numeric. Microbiologist Elisabeth Bik became widely known for spotting duplicated and manipulated image panels in scientific papers. A 2016 paper by Bik, Casadevall, and Fang examined thousands of biomedical papers for inappropriate image duplication.
What the technique is
Image-duplication detection asks whether two panels that are presented as different experiments, samples, or conditions contain the same underlying visual evidence. The duplicate may be exact, cropped, rotated, stretched, contrast-adjusted, or relabeled.
Why it works
Real microscopy, western blots, and gel images contain accidental local structure: speckles, bands, background noise, cell positions, scratches, and edges. Two independent experiments should not share the same accidental structure. If the same noise pattern appears twice, the simplest explanation is reuse of the same image source.
How to check
- Start visually: compare panels that have similar shapes, band patterns, or background texture.
- Transform one image: rotate, flip, crop, and adjust contrast to see whether structures align.
- Use software for scale: perceptual hashes and feature matching can find candidates, but humans still verify context.
- Ask what the labels claim: duplication is most serious when reused images are labeled as different conditions.
Classroom-use examples from Bik, Casadevall, and Fang (2016), distributed under CC BY 4.0. Use them to teach the visual logic, not to infer intent from a single figure.
Class Lab: Explain The Detector, Not Just The Result
The learning goal is that students can explain why a technique works. For each dataset, they should write a short evidence memo with four parts: expected mechanism, observed pattern, reason the mismatch matters, and innocent explanations that remain.
1. Distribution audit
Draw a histogram. Explain what shape the real-world process should create and why.
2. Digit audit
Count first digits or terminal digits. Explain whether the variable should follow Benford, rounding, or neither.
3. Arithmetic audit
Run GRIM-style checks on means and denominator checks on percentages. Explain the integer constraint.
4. Provenance audit
Inspect row order, duplicates, formulas, formatting, timestamps, and earlier file versions.
Evidence memo template
| Question | What students should write |
|---|---|
| What is the technique? | Name the test and the exact quantity computed: digit count, rounded-ending rate, row distance, possible integer sums, or image similarity. |
| Why should it work here? | Connect the test to the claimed data-generating process. The explanation matters more than the software. |
| What did you observe? | Show a plot, table, or calculation. Use neutral language: inconsistent, implausible, duplicated, too uniform, or requires explanation. |
| What else could explain it? | List benign alternatives such as rounding rules, data-entry correction, capped variables, survey skip logic, weighting, or merge errors. |
| What would settle it? | Name the needed evidence: raw forms, export logs, Qualtrics files, source images, formulas, metadata, or replication. |