The correlation that already knew the answer
Four predictions, written down and locked before I looked. One came back at p ≈ 10⁻⁹⁰. That is normally the best moment in this kind of work. It was actually the moment to be most suspicious of myself.
Like Issue 01, this one starts inside my own math research — a different mistake, a different mechanism, the same discipline applied a little later in the process. Issue 01 ended with me finding out, the hard way, that a number which refuses to move under every test isn't necessarily real — it can just mean your machine has fewer moving parts than you think. I took one lesson from it into everything after: write down what you expect, and how many things you're testing, before you look at any results. Pre-register, in the language people who do this for a living use.
So this time, before running anything, I wrote down four specific predictions and pinned them down in writing — pre-specified, if you want the careful word for it, rather than the more formal "pre-registered," since nobody but me had the record before I looked. Four chances to get lucky instead of one, so I applied a Bonferroni correction to push each individual bar back up.
Three of the four came back exactly the way an honest pre-registration is supposed to let you report: unremarkable, inside the noise, nothing to write home about. That's not a failure of the method. That's most of what pre-registration is for — it lets you say "these three didn't pan out" without quietly leaving them out of the story.
The fourth did not come back unremarkable.
Here is what that p-value actually says, and it is narrower than it sounds. Under the null model the test assumes — that there is truly nothing connecting these two quantities — a result at least this extreme would turn up with probability around one in ten to the ninetieth power. That sounds decisive. But the p-value is not the effect size, and it is not the odds that I'd stumbled onto something real. The correlation itself is the effect size, and it was 0.509 — strong, but not supernatural. All the p-value actually told me was how surprising a number like 0.509 would be if the null model were true. It said nothing about whether the null model was the right model to be testing in the first place, and it turned out not to be.
I didn't know that yet. For a few days I just had a beautiful number and the particular quiet excitement of wanting to tell someone and not letting myself, because saying it out loud would have made it a claim, and a claim is something a person can test.
Which gate actually caught this
I want to correct myself on something before I go further, because getting it wrong would undersell what actually saved me here. The instinct that did the damage — go back and see how each quantity is actually built, instead of taking the correlation at face value — is Gate One's job, not Gate Two's. Is it real? Does this survive a test built to kill it, or does it only work because of how I built the pipeline that produced it? That's the gate that did the killing.
So instead of writing the result up, I went back to the two quantities I'd correlated and wrote out exactly how each one was built, term by term, from the raw numbers up — not what they were supposed to represent, but what they actually were, arithmetically.
They weren't independent. One of them was partly built out of the other. Buried a few steps back in how each quantity was defined, they shared a term — the way a grand total shares a term with a subtotal that's already folded into it.
When two quantities share a piece like that, part of any correlation between them is locked in before you measure anything. The purest version of this is a person's weight in pounds and their weight in kilograms, which will correlate at essentially 1.0 forever, because one is nothing but the other multiplied by a constant. My case wasn't that clean — the shared term was one piece folded into two otherwise different quantities, not the whole of either one, which is presumably why the correlation came back at a merely alarming 0.509 instead of a self-evidently broken 0.99. But from that correlation alone, I had no way to tell how much of the 0.509 was the relationship I was actually testing for and how much I had built into the definitions myself, without noticing. That's what killed it as evidence — not that the number was zero, and not necessarily that the underlying hypothesis was wrong, but that this particular correlation could no longer be trusted to speak to it either way.
Gate Two got its turn too, and it landed the second blow rather than the first: this exact trap isn't new. Karl Pearson described the general problem in 1897, studying ratios of organ measurements that shared a common part, and called it spurious correlation — his own word for it, in the title of the paper. It shows up constantly wherever people compute indices and ratios instead of comparing raw, independent measurements, and modern statisticians discuss the same family of problems under the plainer name mathematical coupling. My machine had rediscovered a hundred-and-twenty-nine-year-old trap and handed it back to me formatted as a breakthrough.
The p-value wasn't lying. It was giving a precise answer to the wrong question.
What actually made me look
Here's the part I want to be precise about, because the honest answer isn't "the number was too big, so I got suspicious." Overwhelming significance is what you hope for. You don't get to distrust your own dream result just because it's dramatic, or you'd spend your whole life distrusting everything that ever went right.
What actually made me look was smaller, and easy to miss entirely. Before I'd run anything, I had written down not just that I expected a relationship, but roughly what shape it should take. What came back was in the direction I'd guessed — and far, far stronger than anything I'd predicted. Too strong, and strong in a way that, once I looked closely, wasn't tracking the phenomenon I'd predicted at all. It was tracking the piece of arithmetic the two quantities happened to have in common.
The size of the number wasn't the tell. The mismatch between the shape I'd predicted and the shape that actually showed up was the tell. Pre-registration didn't just protect me from cherry-picking after the fact — it gave me a specific, written prediction to compare the real result against, and the real result didn't match it. Without that comparison sitting on paper from before I looked, a result this strong is exactly the kind of thing that talks you out of checking it.
Cause of death: I tested whether the correlation was significant before checking whether the formulas had already built part of it in.
Killed at Gate 1 — Is it real? · Confirmed not new at Gate 2
What survived
The correlation didn't survive intact — enough of it was arithmetic that I couldn't tell you what, if anything, was left over once the shared term was accounted for. What survived is a specific, mechanical habit: before I accept any correlation now, I write out both quantities down to their raw definitions and check, explicitly, for a shared term. Not "do these seem like they measure different things." Do the numbers, on paper, actually share a piece.
And Gate Two, once it got involved, cut in both directions, which is the part of it people forget. Search the prior work before you claim new ground — that's the half everyone remembers. The half that's harder, and that this cost me time to learn, is: when a result looks too strong, your first move is to go hunting for the boring explanation, not to go hunting for more support. There usually is one. Mine was a hundred and twenty-nine years old.
Samir Hanna Safar is an independent inventor with 23 granted U.S. patents. Honest Limitations publishes one failed idea a week — and what survived after it failed.
Source note: Pearson, K. (1897), "On a Form of Spurious Correlation which may arise when Indices are used in the Measurement of Organs," Proceedings of the Royal Society of London. "Mathematical coupling" is the term used for the same family of problems in later statistical literature.
Drafting, computation and formalisation are carried out with the assistance of an AI system. The questions, the direction and every choice are mine, and I take full responsibility for them.