I have run 35 pre-registered backtests of retail trading strategies. The pass/fail criteria were written down before each test, and the goalposts stayed welded in place after. Twenty-seven died, among them the strongest raw edge this pipeline ever measured, killed by friction alone. Two passed and still didn't get my money, because a tiebreaker written in advance said they hadn't earned it. Every strategy died in the pipeline before it touched a live account: total tuition paid to the market, $0.
This kit is that pipeline. It will not make you money. It stops your backtest from lying to you about making money, which, if you trade, is the same thing wearing work clothes.
35 gates Β· 29 kills Β· 1 pass Β· 2 certified, then benched by their own pre-registered tiebreaker Β· 3 split verdicts (Japan + the wrapper arc) Β· $0 lost live
counted from the public gate files; every number on this page is in the repo
Each of these looked plausible. Several looked great, right up until honest fills, split-adjusted data, and a de-survivorship universe were applied. The picture first, then the ledger; full writeups are free in the repo.
DISTANCE FROM THE GATE β best honest profit factor per family
fast-strategy gate, pre-registered: PF β₯ 1.30 at the stated cost Β· every dot is a documented result
Not one family reached the line. The catalyst pair is the expensive lesson: a passing primary sample (hollow dot) collapsing to negative expectancy on the holdout it never touched. The pairs dot is the newest and sharpest: PF 2.23 at zero cost β still the strongest gross edge in all 35 gates β landing at 1.06 once each hedged round trip pays the spread four times. Exact figures, verdicts, and causes of death: table below.
| strategy family | best honest result | verdict |
|---|---|---|
| Momentum scalping (213-trade window) | 49% win, Β±1.18%: a fair coin | FAIL |
| Opening-range breakout, long+short ("Sharpe 2.4" paper) | PF 0.82β0.85 @ 0.10%/side | FAIL |
| Catalyst gaps (β₯5% + volume) | PF 1.23 primary, 0.93 holdout | FAIL |
| Earnings drift (real SEC 8-K dates) | PF 1.11, worse than random gaps (1.29) | FAIL |
| Mean reversion (RSI-2) | PF 0.77 @ 0.25%/side | FAIL |
| News sentiment (19.5k articles, LLM-scored) | made every config worse | FAIL |
| Crypto funding carry | real mechanism, ~0.4% net in 2026 | FAIL |
| Cross-asset trend, monthly (Faber, zero tuned params) | 8.1% CAGR, β1.1% in 2022 Β· reader-forced 2008 extension: 5.9% over 18.4y, β1.5% in 2008 β pass holds, repriced | PASS |
| Same rule, sampled daily | 7.31%; whipsaw ate it | FAIL |
| Options premium selling (5 pro income ETFs + CBOE's own indices) | every fund's Sharpe below its own underlying | FAIL |
| Turn-of-month window (T-4..T+3) | Sharpe 0.44 vs SPY 0.73; 42% of returns in 29% of days | FAIL |
| Factor ETFs (MTUM/VLUE/QUAL/USMV, EW + 12-1 rotation) | Sharpe 0.87 / 0.84 vs SPY 0.90 | FAIL |
| Pairs stat-arb (GGR distance, liquid, long-short) | gross PF 2.23 β 1.06 @ 0.10%/side/leg | FAIL |
| Insider-cluster buying (2+ insiders, $200k+, SEC Form 4) | PF 1.17 @ 0.25%/side β vs 1.23 gross: friction wasn't the murderer | FAIL |
| Candlestick patterns (engulfing / hammer / piercing, 43,624 signals) | pooled PF 0.95 @ 0.25%/side; the hammer is a literal coin flip (1.00) | FAIL |
| Cash-merger arbitrage (789 deals hand-built from SEC filings, breaks paid in full) | 4.03% CAGR vs a cash+3 bar of 4.32%: clears it only at exactly zero cost | FAIL |
| Prediction-market cross-venue arb (17 identical-resolution Kalshi/Polymarket pairs) | +2.2%/yr on deployed capital, less than T-bills (3.62%) | FAIL |
| Paid order flow / maker rebates (receipts only, no simulated fills) | net β0.17 to β0.23Β’/share at retail tier; first positive payment starts at $10M/month | FAIL |
| Unified everything-rotation (1,365 stocks + 13 ETFs, one momentum pool) | 7.14% vs SPY's 11.53%; ETFs entered the top-10 in 0 of 55 months | FAIL |
| Trend overlay on a momentum index fund | halved the drawdown, cost 3.4pp/yr vs a 2pp cap | FAIL |
| "Smart money concepts" M15, mechanized (7,087 forex trades) | PF 0.951 before costs; inside a 20-seed random band | FAIL |
| Cross-asset trend inside a prop-firm eval (outside spec) | Sharpe 0.82; static rules: 78% funded, +$19k EV β trailing+consistency rules: dead | SPLIT |
| The same spec on FTMO's actual terms | overnight financing alone: Sharpe 0.78 β 0.11 | FAIL |
| The same spec, swap-free firm class | 76% funded, +$22k EV β unless the 40% consistency rule applies: 66β69% of payouts denied | SPLIT |
PF = profit factor. The pattern: nearly every DIRECTIONAL price-derived signal is ~PF 1.05β1.10 gross, and retail friction eats it whole. The market-neutral exception measured 2.23 gross and died anyway: hedged trades pay friction twice on a smaller move.
GATE 20 Β· OPTIONS PREMIUM SELLING: THE LAST MAINSTREAM FAMILY
tested twice over: 5 professional covered-call/put-write ETFs, then CBOE's own benchmark indices (real traded SPX option prices back to 1990)
Every income fund had a worse Sharpe than the index it writes options on. QQQ made 21.6%/yr over QYLD's decade; QYLD kept 10% of it with 70% of the drawdown. But the 33-year index record is the real exhibit: a mechanism dying in public.
| index (vs SPY, Sharpe) | pre-2012 | post-2012 |
|---|---|---|
| BXM covered-call | +0.20 vs +0.16 Β· beats | +0.61 vs +0.90 Β· lags |
| PUT put-write | +0.24 vs +0.06 Β· beats | +0.70 vs +0.90 Β· lags |
| BXMD 30-delta buy-write | +0.50 vs +0.33 Β· beats | +0.73 vs +0.90 Β· lags |
The volatility premium WAS real; the academic papers were written on the pre-2012 sample. Then it got institutionalized, and the edge went to whoever collects the flow. Same life-cycle as crypto funding carry (2021: ~22% β 2026: ~0.4%), two decades slower. The crash tell: in 2020 these indices fell HARDER than SPY (BXM β54.6% annualized vs β32.4%). Short vol is short the gap: insurance that pays out in drizzle and fails in the flood. Full writeup and both prediction audits in the repo.
GATE 17 Β· A READER RAN AT THE ONLY PASS
the challenge: "2022 is your easy case wearing a hard case's clothes β one inflation shock made trend rotate to cash exactly when cash started paying. Extend through 2008 or don't fund it."
He was right about the gap, wrong about the mechanics (the cash fix used the real monthly T-bill path, not a flat 4β5% β audited: it moved 2022 by +0.21pp, and legitimately). So the same frozen rules ran again, extended back to January 2008, same day. Pre-2016 prices cross-checked against the original source to 4.6bps.
| regime | this strategy | SPY |
|---|---|---|
| 2008 zero-rate bear | β1.5% (bond sleeves carried it, not cash) | β36.8% |
| 2011β2016 ZIRP grind | 5.1 yrs below T-bills, 4 yrs underwater β the real max drawdown lives here | compounding 13β32%/yr |
| 2022 dual bear | β1.1% | β18.2% |
| full 18.4 years | 5.85% CAGR Β· Sharpe 0.69 Β· maxDD β12.1% | 11.29% Β· 0.63 Β· β46.3% |
The pass holds β every pre-registered readout on the extended sample cleared β with one flag kept on the record: against the gate's literal CAGR β₯ 6% bar, 5.85% is a thin fail (the bar's written rationale was T-bills + 3, which it clears at 4.37%). The price changed: no year worse than β1.6%, but you can sit five years underperforming T-bills, and the realistic failure mode is quitting in year three of the grind, not a crash. Sizing now keys off these numbers, not the friendly 2017β2026 ones. The one strategy that survived this pipeline got its victory lap cancelled by a stranger with a calendar β and the record is better for it.
Round two: the same reader then challenged the flag itself β if the bar was 6%, a 5.85% is a fail, and re-deriving the bar from its rationale after seeing the number is goalpost-moving. The receipts held (the T-bills+3 derivation was registered in the same line as the number, before any run; the extension's formula bar went into the script before that run executed) β but the structural point landed: a bare number with a parenthetical rationale is two bars waiting to diverge, and the gate should never have allowed both readings. That's now a permanent rule in the gate templates: bars derived from a rationale are written as formulas, with which form governs stated up front. $0 is funded under either reading.
Round three: this note used to claim a 90/10 SPY blend "dominates the pure strategy on every tracked metric." The same reader killed that too. The 2008 column β tracked all along, just not mentioned β has pure at β1.5% and 90/10 at β5.5%: the blend triples the loss in the exact year that proves the mechanism. A finer weight grid shows 90/10 isn't even the drawdown minimum (92.5/7.5 is), and on a crash-flavored metric (worst rolling 12 months) pure wins (β9.0% vs β10.3%) β blend weight is a parameter, the frontier is an in-sample scan, and rules 7 and 8 of this page apply to their author. What survives as description: the Sharpe hump is smooth and broad (0.69 β 0.77, no knife-edge), and a token 2.5% SPY sleeve cuts years-below-T-bills from 5.1 to 3.6 β his own diagnosis ("filling the exact hole") as a number. The weight axis trades crash pain against grind pain; the frontier describes the dial and cannot pick the point. And his sharpest contribution stands unanswered by the data: the regime where both legs fail at once β entering a bear from zero rates β appears nowhere in 18 years of sample. Japan sat in it for twenty. Sizing humility has to price what the record cannot.
Round four: he proposed the test himself, so it became
Gate 26 β The Japan Test, run to his design. Borrow the regime you canβt
manufacture: Nikkei, synthetic 10-year JGB total return, gold in yen, call-rate
cash, frozen Faber rules, entry January 1990 β the month after the bubble top.
Governing window 1999β2012: both legs dead on entry. Result: survived by
0.05 percentage points β CAGR 3.17% against a formula bar of cash+3pp =
3.12%, a margin inside data noise and called that on the record β and over the
full 23 years it did exactly what he said it would: trod water (1.76% vs cash
1.27%, 15.8 years below cash). The eyeball finding is the keeper: equal-weight
buy-and-hold of the same sleeves beat the trend rule by 1.6pp/yr in ZIRP β every
whipsaw parks in an asset paying nothing β so the overlay subtracted return and
bought drawdown instead (β10.0% max vs β15.9% basket, β62.8% Nikkei). The
margin that survived came from the one sleeve that doesnβt run on interest
rates: gold in yen, up 4.9x. His kicker is now the pre-registered meaning: in
this regime the 20% sizing rule stops being humility and becomes the actual
finding. Prediction audit: the house predicted treads-water and missed by the
same five basis points; his registered calls β underperforms its own basket,
treads water against cash+3 β scored correct and 5bp-from-correct. Those calls
arrived after the numbers were already live on this page; theyβre logged with
that exact timestamp, not dressed as pre-registration, and scored anyway,
because he asked to be held to a bar in public. A pre-registration that only
counts when the clock is convenient isnβt one. His verdict
is now the meaning on record, in his words: not a return engine that
happens to survive bears β a drawdown engine that gives up return in flat
regimes to buy survival in violent ones. It survives everywhere because
it isnβt trying to win everywhere. Adopted with it: regime-transplant gates
now bar against the basketβs own buy-and-hold too, and under that bar the
overlayβs ZIRP verdict is a clean fail β the same sentence. Full receipts in
the repoβs japan_experiment/. Co-author: Zestyclose.
GATES 27β31 Β· THE COLLECTOR'S SIDE OF THE COUNTER
the standing thesis after 26 gates: our slippage is somebody's revenue. the winners collect structural payments instead of predicting price, so the last five gates went at the collecting side directly, in one 30-hour sweep
Every family measured the same thing from a different angle: the payment is always real, and it is always priced to leave nothing above cash-compensation at retail scale.
| the payment | measured | what's left at retail |
|---|---|---|
| Merger-arb spread (deal-break insurance) | 4.03% CAGR, 789 deals | bar was 4.32%; clears it only at zero cost |
| Cross-venue prediction-market gap | gaps >1Β’ in 50β95% of hours | round trip costs 2.5β4Β’; nets +2.2%/yr < T-bills |
| Maker rebates / order flow | receipted fee schedules, every venue | β0.17Β’/share; positive pay starts ~$10M/month |
The merger-arb gate is the operator-vs-machine story: the
operator registered an unhedged PASS four days before the run; the machine put it
at 65% fail. The deal set was built from SEC filings announcement-first so the
broken deals stayed in, and the famous tail turned out survivable (10% of deals
broke, mean β2.9%, book max drawdown β5.5%). What killed it was quieter: the
spread is priced almost exactly at cash+3 gross, and even 0.05%/side pushes it
under. The 2019+ holdout was never opened; the frozen protocol seals it on a
primary fail. Five validity rounds ran before any number was believed β the first
three sims printed 7%, 27%, and 17%, each destroyed by its own eyeball audit
(fictional payouts, acquirer tickers mis-cast as targets, a rebalancing rule
quietly harvesting bid-ask bounce as alpha). The full defect ledger ships in the
repo's merger_arb_experiment/DEVIATIONS.md, because a hostile reader
should get the knife with the body.
The prediction-market gate found a second body on the way down: measured raw, the arb "returns" +29%/yr, and every fat entry sits in a market's first 72 hours, where one venue's price is a bookless print that collapses 30 cents on no news. No historical order-book exists to prove any of it was ever executable. Price without a book is not a quote; that readout ships as an unverifiable upper bound, and the mature-market number, the one you could actually have, pays less than the T-bills the capital would sit in. One venue also now charges a 5% taker fee on the exact markets measured: the toll booth raised its price mid-study.
The coda is the week's real lesson. The operator asked to
beat the index by chasing momentum, so momentum got the full treatment: a
hand-rolled top-10 rotation over 1,365 stocks and 13 cross-asset ETFs in one pool
(his design: the defensive ETFs were supposed to take over in bears; in 55
months, including all of 2022, not one ETF ever out-ranked the ten hottest
stocks), then a trend overlay on the index fund that already does momentum
properly. Hand-roll: 7.14%/yr with a β41% drawdown. Professionally-run
concentrated academic momentum: 12.4%, worse than the plain index. The boring
S&P momentum index fund the whole exercise tried to beat: 19.5%/yr for a
decade, live. The overlay halved its drawdown and paid 3.4pp/yr for the
privilege against a 2pp cap written in advance. Every step toward "better" (more
concentration, faster signals, cleverer math) made it worse, monotonically.
The one keeper the hunt ever surfaced is a 13-basis-point product you buy and
leave alone, and the machine's final advice on it was about sizing, not signals.
All five gate files, predictions, and verdicts:
merger_arb_experiment/ Β· predmarket_arb_experiment/ Β·
orderflow_experiment/ Β· stock_rotation_experiment/ Β·
spmo_overlay_experiment/.
GATES 32β35 Β· THE WRAPPER ARC
an outside collaborator β a prop-firm trader β brought two strategies and a thesis: trade the firm's capital through a challenge account. all of it went through the gate
The arc is the record in miniature: his playbook died, his spec produced the record's second pass, and the wrapper turned out to be priced like everything else.
The playbook first. "Smart money concepts" on 15-minute charts, mechanized with literature-default parameters and every rule frozen in writing: 7,087 trades, five years, two currency pairs β profit factor 0.951 BEFORE costs, and statistically indistinguishable from random entries at matched frequency, risk and exits. The zones are decoration. Then his trend spec β the surprise. Faber-class cross-asset trend, vol-targeted, run through evaluation rulesets on the real 2008β2026 path: Sharpe 0.82 at honest costs, 78% funded probability with ~$19k expected value per attempt against STATIC-drawdown rules. Against trailing-drawdown-plus-consistency rulesets it fails every cell β trend profits are lumpy, and consistency rules exist to deny exactly that shape. The firm's rule sheet, not the strategy, is the constraint.
Then the transfer test, and the cleanest attribution table
in the record. Moving the spec onto FTMO's actual platform: dropping the bond leg
(no bond CFDs) cost 0.05 of Sharpe; switching to single-name oil cost nothing;
modeling the overnight financing a months-holding CFD book actually pays took it
from 0.78 to 0.11. The venue's own swap charge is the entire death β the same
kill as merger arbitrage, one venue closer to home. Swap-free static-drawdown
firms exist and restore the passing regime on their published rules (76% funded,
+$22k EV) β unless one ambiguous consistency clause in the payout policy applies,
in which case two-thirds of payout cycles get denied. As of this writing the
build waits on that sentence, in writing, from the firm. A companion rules study
priced the wrapper itself: the published rule sheet is beatable by a coin flip
(1-in-3 passes at zero edge, all the value in the funded-stage payout option),
which is exactly why challenge passes are not receipts. The standing bar for any
prop-firm claim: funded PAYOUT history over months. Files:
smc_experiment/ Β· trendprop_experiment/ Β·
propfirm_experiment/.
THE KILL-SIDE AUDIT Β· TURNING THE GAUNTLET ON OUR OWN TOMBSTONES
2026-07-29, prompted by the operator's question: "could the graveyard be artifacting kills?" β the record's skepticism had only ever been spent against passes; nobody had audited a loser tail
Criteria frozen before looking (what would count as an overturned kill), then: top-10 losers of three closed gates verified against independent price data, and every failed gate's cost ladder collated. Published whichever way it landed.
| question | answer |
|---|---|
| Any kill overturned? | No. ORB 0/10 loser flags; mean reversion 1/10; trend-follow 2/10; corrections move PF by <0.01 against 1.30 bars. |
| Any fabricated losses? | TWO β BHVN's "β91%" (Pfizer acquired Biohaven at $148.50; the ticker passed, gapless, to a ~$9 spinco) and CCXI's "β80%" (a hold jumped a 1,209-day delisting gap onto a relisted successor). Ticker splices fabricate losses exactly like they fabricate wins. |
| How cost-fragile are the kills? | Less than the auditor predicted: 10 of 18 failed gates fail even at ZERO cost. Exactly one kill is cleanly cost-conditional β pairs, which passes its bars at 0.05%/side/leg. |
The disclosure that matters, on the record: the kills share
infrastructure β one cost model (calibrated from this pipeline's measured live
slippage, always taker execution), one fill simulator, largely one data stack.
Kills that die of friction weaken TOGETHER if the cost model is wrong for your
execution; they are correlated verdicts, not 29 independent confirmations. Every
results file publishes its full cost ladder so the count you trust is the one
you recompute at your own costs. If your execution is materially better than
retail taker, the pairs gate is the tombstone to re-run first. New standing
rules from the audit: gap-scan and splice-screen every cached universe, and
eyeball the top-10 losers with the same ritual as the winners β both tails,
every gate. Full audit: kill_audit/ in the repo.
A backtest is a witness with every incentive to lie to you. Interrogate it like one.
These aren't hypotheticals. Each one is an artifact my own pipeline produced, believed for a while, and then caught. Each one became a rule in the free checklist.
The split artifact. Raw price bars plus a 1:25 reverse split inside the hold window turned a $0.17 stock into the trade of the decade. It never happened. A strategy "passed validation" for six days on trades like this before the split-adjusted rerun killed it.
rule: split-adjusted bars only, no exceptions
The fantasy fill. A backtest that fills your stop at the stop price is fiction: names that gap through a β5% stop routinely close down 10 to 15%. Half my "edges" were fill fantasies that evaporated under a checked-at-close, filled-at-close model.
rule: model the fill you actually get, not the one you want
The false-positive machine. A slot-capped portfolio simulation printed profit factor 3.96 by accidentally cherry-picking under 10% of its own trades. It did this twice, on two unrelated strategies, before the pattern was caught and banned from the pipeline.
rule: slot-capped stats are a false-positive generator
Every verdict above traces to a gate file written before its test ran. The receipts, directly:
| artifact | what it is |
|---|---|
| The full graveyard writeup | all 35 gates, every number, the three expensive lessons |
| Gates 17β25: passes, benchings, and the resumed hunt | the first PASS (and its reader-forced 2008 extension), the benched survivors, the falsified predictions (mine included), the 33-year options autopsy, and the strongest gross edge on record dying of friction |
| The artifact-hunting checklist | the free tier: every rule priced by the disaster that created it (landing in the repo this week) |
I pre-registered the launch the way I pre-register backtests: the paid tiers would ship only if people demonstrably wanted them. No fake scarcity, no countdown timer; the incentive ran the other way — packaging the kit properly took real weekends, and I'd only spend them on real demand.
Resolution (2026-07-27): the demand signal was real but modest — a four-round public audit of the flagship pass, a free-tier tester, workshop build requests — and honesty requires saying I called the gate rather than a counter tripping it. The packaging got finished, the store cleared review, and the stamps above are now live checkout links. If nobody buys, that verdict joins the ledger like every other one. The checklist tier stays free either way, and stars and issues still steer what gets built next.
Comments here are powered by GitHub: each one is filed as an issue on the graveyard repo, which means posting a comment goes on the public record. Say what you'd use, what's missing, or why the whole thing is wrong. Arguments count as interest.
The best result on this page exists because a stranger demanded it. The Gate 17 extension β the one that repriced my only pass β started as a reader's challenge: "2022 is your easy case. Extend through 2008 or don't fund it." He was right, the run happened the same day, and the record got better by getting worse. That worked well enough to become a standing offer.
Bring me a mechanism and it gets the gate treatment: pass/fail criteria pre-registered before the run, honest fills, real costs, and the result published here β pass or kill β with credit to you. Or request a testing tool: tell me which backtest lie it should catch, and it goes on the kit backlog.
Ground rules, so nobody's time gets wasted: nothing that trades a live account gets built, ever. No signals sold back to you. Requests and results are public β that's the quality control. "It backtests well" is not a mechanism; tell me who's on the other side of the trade and why they're paying you, and your idea jumps the queue. I pick what gets built and promise no deadlines. Paid data needs a budget conversation before anything runs.
Request a build β Files as a public GitHub issue β the queue is public by design.
One retail trader who spent months building a bot, watched honest testing kill every strategy it ran, and kept the machine that did the killing. The graveyard writeups converted the skeptics first, including the ones who showed up to argue. The receipts are public, the code produced them, and when my own pre-registered prediction got falsified by my own gate last week, that went in the repo too.
Questions, arguments, or a strategy you want dead? File a build request or open an issue on the repo. Arguing in public is the quality-assurance process.
Β© 2026 Β· strategygraveyard.com Β· The Falsification Kit Β· Educational software and research writeups. Not financial advice. No performance claims made or implied. Backtested results are historical simulations and do not predict future returns. The kit tests strategies; it does not provide, recommend, or execute them.
Widget not loading (some previews block it)? Post directly as a GitHub issue β