PF 14.6, on 16 trades
In April 2026 our research sweep surfaced a strategy variant with an out-of-sample profit factor of 14.6. For scale: a profit factor of 2 is excellent. This looked like the best thing the machine had ever found.
We didn't trade it. Here's why — and why the discipline behind that "no" is worth more than the number.
What the sweep found
We were testing options open-interest patterns — whether unusual accumulation in a name's options predicted the stock's next move. Standard practice for us is to sweep the variations: different intensity thresholds, different holding periods, different confirmation rules. One cell of that sweep — accumulation with an open-interest volume spike above a threshold — came back:
- In-sample: n = 21, PF 8.54
- Out-of-sample: n = 16, PF 14.60, win rate 81.25%
- Raw p-value in the sweep: 0.0009
Real out-of-sample split. A p-value that looks bulletproof. This is exactly the shape of result that launches a thousand subscription services.
Why we didn't believe it
Sixteen trades. With n = 16, profit factor is not a property of the strategy — it's a property of whichever two or three trades happened to be the biggest. One lucky +40% winner in a 16-trade sample can carry the entire number. Flip two outcomes and PF 14.6 becomes PF 2, or 0.8. The number has no stability at that sample size, which is why our rules set a hard floor: 200 closed trades minimum before a strategy can even graduate to shadow trading, 500 before it's trusted with capital.
The p-value came from a sweep. That 0.0009 was one cell among dozens of threshold-and-horizon combinations we tested in the same run. Test enough variations and something will show p < 0.001 by pure chance — that's not a flaw in statistics, it's the whole reason multiple-comparisons corrections exist. A p-value quoted without the number of things tested alongside it is close to meaningless.
Its siblings were telling on it. The same signal at a 5-day horizon — the adjacent cell in the same sweep — showed an out-of-sample win rate of 46.2% (n = 13). Below a coin flip. A real, robust edge degrades gracefully as you vary the parameters around it. A statistical fluke is a spike surrounded by noise. This was a spike surrounded by noise.
What happened to it
The variant went to our watchlist — parked, watched, not traded, and not quietly deleted either. Then the deeper audit came back: the parent population this variant was carved from — the broader options-volume-spike pattern — failed its own permutation test. If the base signal is indistinguishable from random, a spectacular sub-slice of it isn't a discovery; it's a lottery ticket with good-looking statistics.
The play was retired without ever producing a pick. The best number our machine has found went from discovery to grave without a single trade — which is not a malfunction. It's the system working.
What to take from it
When you see a spectacular backtest number, ask:
- What's n? Any profit factor on fewer than a few hundred trades is an anecdote with decimals.
- How many variations were tested to find this one? If the answer is "we don't say," assume the graveyard of failed variants is large and hidden.
- What do the neighboring parameters look like? Robust edges have robust neighborhoods. Miracles have cliffs on either side.
The hardest part of systematic trading isn't finding exciting numbers — the machine produces those weekly. It's having rules that force you to say no to almost all of them. Our rules said no to the best number we ever found. That's the product.
We run this same gauntlet on every strategy we build — most ideas die in research, 24 full plays have been killed and published, 6 are live. How the gates work: methodology. What a play is: the basics. How to judge anyone's claims, including ours: the difference.
Research publication — not investment advice — disclaimer.