Your backtest is lying to you

backtest PF 4.05 → live PF 0.48 (677 entries, 1-min bars) · re-audit residual alpha +2.7%/10d (t=3.09) · permutation p=0.056 FAIL · verdict stays KILLED

This is the story of one strategy that lied to us twice — once when it looked great, and once when it looked dead. It's the best illustration we have of why a single backtest number, good or bad, is never the answer.

Act one: the validated strategy

A momentum-continuation pattern — stocks gapping on catalysts, entered intraday. The backtest said profit factor 4.05. It cleared our review at the time and began firing live.

Its live signals — 677 of them — scored a profit factor of 0.48 when measured on minute-resolution data. It lost money at nearly the mirror-image rate at which the backtest said it would make money.

We killed it and wrote "classic overfit" in the changelog.

The detail that should scare you

Here's the part most post-mortems skip. When we re-simulated those same 677 picks on daily bars, the simulation said PF 1.50 — profitable. On 1-minute bars, the same picks scored PF 0.48.

Same strategy. Same entries. Same exits on paper. The only difference was the resolution of the price data — and it flipped the verdict from "works" to "loses." Intraday gappers fade fast; a daily-bar simulation simply cannot see the stop-outs that happen at 10:15am. Any backtest of an intraday strategy computed on daily bars is fiction, and a lot of published "verified backtests" are exactly that.

Act two: the autopsy that reopened the case

Two months later we noticed something uncomfortable: several of the biggest winners crossing our platform were names this dead strategy had flagged. So we re-investigated — full 12-dimension protocol, 456 deduplicated signals, tested against a 23,878-stock-day baseline.

The re-audit found our own kill diagnosis was partly wrong:

So we un-killed it, right?

No. Because it still fails the gates:

The verdict

Still killed as a strategy. The signal now feeds our scoring model as one input among many — a context where a tail-dependent, regime-untested signal can contribute without being trusted to trade on its own. If bear-market data accrues and the sample crosses 500, it earns a fresh trial. Not before.

What to take from it

  1. A backtest is only as honest as its price resolution. Daily bars on an intraday strategy flipped PF 1.50 into PF 0.48 on identical trades.
  2. The exit is part of the strategy. The same entries ranged from PF 1.54 to PF 2.16 depending on exit rule alone. Anyone quoting one backtest number tested one arbitrary combination.
  3. Tail-dependence is a disclosure, not a footnote. If removing the top 5% of trades kills the edge, you are not buying a strategy — you are buying a lottery position with good branding.
  4. Verdicts should survive their own autopsy. We re-opened our own kill, found our reasoning partly wrong, published that too — and the conclusion still held. That's what a record is for.

We run this gauntlet on everything we build. 24 strategies killed and published so far. The method is the product.

We run this same gauntlet on every strategy we build — most ideas die in research, 24 full plays have been killed and published, 6 are live. How the gates work: methodology. What a play is: the basics. How to judge anyone's claims, including ours: the difference.

Research publication — not investment advice — disclaimer.