doc: FALSIFICATION_KIT.md status: pre-registered Β· checkout live 2026-07-27 performance promises: 0

case file Β· originBacktesting tools with a body count.Γ—27 FAILED

I have run 35 pre-registered backtests of retail trading strategies. The pass/fail criteria were written down before each test, and the goalposts stayed welded in place after. Twenty-seven died, among them the strongest raw edge this pipeline ever measured, killed by friction alone. Two passed and still didn't get my money, because a tiebreaker written in advance said they hadn't earned it. Every strategy died in the pipeline before it touched a live account: total tuition paid to the market, $0.

This kit is that pipeline. It will not make you money. It stops your backtest from lying to you about making money, which, if you trade, is the same thing wearing work clothes.

35 gates Β· 29 kills Β· 1 pass Β· 2 certified, then benched by their own pre-registered tiebreaker Β· 3 split verdicts (Japan + the wrapper arc) Β· $0 lost live

counted from the public gate files; every number on this page is in the repo

exhibit a Β· the killsEvery family, measured against the gate it pre-registered.

Each of these looked plausible. Several looked great, right up until honest fills, split-adjusted data, and a de-survivorship universe were applied. The picture first, then the ledger; full writeups are free in the repo.

DISTANCE FROM THE GATE β€” best honest profit factor per family

fast-strategy gate, pre-registered: PF β‰₯ 1.30 at the stated cost Β· every dot is a documented result

Not one family reached the line. The catalyst pair is the expensive lesson: a passing primary sample (hollow dot) collapsing to negative expectancy on the holdout it never touched. The pairs dot is the newest and sharpest: PF 2.23 at zero cost β€” still the strongest gross edge in all 35 gates β€” landing at 1.06 once each hedged round trip pays the spread four times. Exact figures, verdicts, and causes of death: table below.

strategy familybest honest resultverdict
Momentum scalping (213-trade window)49% win, Β±1.18%: a fair coinFAIL
Opening-range breakout, long+short ("Sharpe 2.4" paper)PF 0.82–0.85 @ 0.10%/sideFAIL
Catalyst gaps (β‰₯5% + volume)PF 1.23 primary, 0.93 holdoutFAIL
Earnings drift (real SEC 8-K dates)PF 1.11, worse than random gaps (1.29)FAIL
Mean reversion (RSI-2)PF 0.77 @ 0.25%/sideFAIL
News sentiment (19.5k articles, LLM-scored)made every config worseFAIL
Crypto funding carryreal mechanism, ~0.4% net in 2026FAIL
Cross-asset trend, monthly (Faber, zero tuned params)8.1% CAGR, βˆ’1.1% in 2022 Β· reader-forced 2008 extension: 5.9% over 18.4y, βˆ’1.5% in 2008 β€” pass holds, repricedPASS
Same rule, sampled daily7.31%; whipsaw ate itFAIL
Options premium selling (5 pro income ETFs + CBOE's own indices)every fund's Sharpe below its own underlyingFAIL
Turn-of-month window (T-4..T+3)Sharpe 0.44 vs SPY 0.73; 42% of returns in 29% of daysFAIL
Factor ETFs (MTUM/VLUE/QUAL/USMV, EW + 12-1 rotation)Sharpe 0.87 / 0.84 vs SPY 0.90FAIL
Pairs stat-arb (GGR distance, liquid, long-short)gross PF 2.23 β†’ 1.06 @ 0.10%/side/legFAIL
Insider-cluster buying (2+ insiders, $200k+, SEC Form 4)PF 1.17 @ 0.25%/side β€” vs 1.23 gross: friction wasn't the murdererFAIL
Candlestick patterns (engulfing / hammer / piercing, 43,624 signals)pooled PF 0.95 @ 0.25%/side; the hammer is a literal coin flip (1.00)FAIL
Cash-merger arbitrage (789 deals hand-built from SEC filings, breaks paid in full)4.03% CAGR vs a cash+3 bar of 4.32%: clears it only at exactly zero costFAIL
Prediction-market cross-venue arb (17 identical-resolution Kalshi/Polymarket pairs)+2.2%/yr on deployed capital, less than T-bills (3.62%)FAIL
Paid order flow / maker rebates (receipts only, no simulated fills)net βˆ’0.17 to βˆ’0.23Β’/share at retail tier; first positive payment starts at $10M/monthFAIL
Unified everything-rotation (1,365 stocks + 13 ETFs, one momentum pool)7.14% vs SPY's 11.53%; ETFs entered the top-10 in 0 of 55 monthsFAIL
Trend overlay on a momentum index fundhalved the drawdown, cost 3.4pp/yr vs a 2pp capFAIL
"Smart money concepts" M15, mechanized (7,087 forex trades)PF 0.951 before costs; inside a 20-seed random bandFAIL
Cross-asset trend inside a prop-firm eval (outside spec)Sharpe 0.82; static rules: 78% funded, +$19k EV β€” trailing+consistency rules: deadSPLIT
The same spec on FTMO's actual termsovernight financing alone: Sharpe 0.78 β†’ 0.11FAIL
The same spec, swap-free firm class76% funded, +$22k EV β€” unless the 40% consistency rule applies: 66–69% of payouts deniedSPLIT

PF = profit factor. The pattern: nearly every DIRECTIONAL price-derived signal is ~PF 1.05–1.10 gross, and retail friction eats it whole. The market-neutral exception measured 2.23 gross and died anyway: hedged trades pay friction twice on a smaller move.

GATE 20 Β· OPTIONS PREMIUM SELLING: THE LAST MAINSTREAM FAMILY

tested twice over: 5 professional covered-call/put-write ETFs, then CBOE's own benchmark indices (real traded SPX option prices back to 1990)

Every income fund had a worse Sharpe than the index it writes options on. QQQ made 21.6%/yr over QYLD's decade; QYLD kept 10% of it with 70% of the drawdown. But the 33-year index record is the real exhibit: a mechanism dying in public.

index (vs SPY, Sharpe)pre-2012post-2012
BXM covered-call+0.20 vs +0.16 Β· beats+0.61 vs +0.90 Β· lags
PUT put-write+0.24 vs +0.06 Β· beats+0.70 vs +0.90 Β· lags
BXMD 30-delta buy-write+0.50 vs +0.33 Β· beats+0.73 vs +0.90 Β· lags

The volatility premium WAS real; the academic papers were written on the pre-2012 sample. Then it got institutionalized, and the edge went to whoever collects the flow. Same life-cycle as crypto funding carry (2021: ~22% β†’ 2026: ~0.4%), two decades slower. The crash tell: in 2020 these indices fell HARDER than SPY (BXM βˆ’54.6% annualized vs βˆ’32.4%). Short vol is short the gap: insurance that pays out in drizzle and fails in the flood. Full writeup and both prediction audits in the repo.

GATE 17 Β· A READER RAN AT THE ONLY PASS

the challenge: "2022 is your easy case wearing a hard case's clothes β€” one inflation shock made trend rotate to cash exactly when cash started paying. Extend through 2008 or don't fund it."

He was right about the gap, wrong about the mechanics (the cash fix used the real monthly T-bill path, not a flat 4–5% β€” audited: it moved 2022 by +0.21pp, and legitimately). So the same frozen rules ran again, extended back to January 2008, same day. Pre-2016 prices cross-checked against the original source to 4.6bps.

regimethis strategySPY
2008 zero-rate bearβˆ’1.5% (bond sleeves carried it, not cash)βˆ’36.8%
2011–2016 ZIRP grind5.1 yrs below T-bills, 4 yrs underwater β€” the real max drawdown lives herecompounding 13–32%/yr
2022 dual bearβˆ’1.1%βˆ’18.2%
full 18.4 years5.85% CAGR Β· Sharpe 0.69 Β· maxDD βˆ’12.1%11.29% Β· 0.63 Β· βˆ’46.3%

The pass holds β€” every pre-registered readout on the extended sample cleared β€” with one flag kept on the record: against the gate's literal CAGR β‰₯ 6% bar, 5.85% is a thin fail (the bar's written rationale was T-bills + 3, which it clears at 4.37%). The price changed: no year worse than βˆ’1.6%, but you can sit five years underperforming T-bills, and the realistic failure mode is quitting in year three of the grind, not a crash. Sizing now keys off these numbers, not the friendly 2017–2026 ones. The one strategy that survived this pipeline got its victory lap cancelled by a stranger with a calendar β€” and the record is better for it.

Round two: the same reader then challenged the flag itself β€” if the bar was 6%, a 5.85% is a fail, and re-deriving the bar from its rationale after seeing the number is goalpost-moving. The receipts held (the T-bills+3 derivation was registered in the same line as the number, before any run; the extension's formula bar went into the script before that run executed) β€” but the structural point landed: a bare number with a parenthetical rationale is two bars waiting to diverge, and the gate should never have allowed both readings. That's now a permanent rule in the gate templates: bars derived from a rationale are written as formulas, with which form governs stated up front. $0 is funded under either reading.

Round three: this note used to claim a 90/10 SPY blend "dominates the pure strategy on every tracked metric." The same reader killed that too. The 2008 column β€” tracked all along, just not mentioned β€” has pure at βˆ’1.5% and 90/10 at βˆ’5.5%: the blend triples the loss in the exact year that proves the mechanism. A finer weight grid shows 90/10 isn't even the drawdown minimum (92.5/7.5 is), and on a crash-flavored metric (worst rolling 12 months) pure wins (βˆ’9.0% vs βˆ’10.3%) β€” blend weight is a parameter, the frontier is an in-sample scan, and rules 7 and 8 of this page apply to their author. What survives as description: the Sharpe hump is smooth and broad (0.69 β†’ 0.77, no knife-edge), and a token 2.5% SPY sleeve cuts years-below-T-bills from 5.1 to 3.6 β€” his own diagnosis ("filling the exact hole") as a number. The weight axis trades crash pain against grind pain; the frontier describes the dial and cannot pick the point. And his sharpest contribution stands unanswered by the data: the regime where both legs fail at once β€” entering a bear from zero rates β€” appears nowhere in 18 years of sample. Japan sat in it for twenty. Sizing humility has to price what the record cannot.

Round four: he proposed the test himself, so it became Gate 26 β€” The Japan Test, run to his design. Borrow the regime you can’t manufacture: Nikkei, synthetic 10-year JGB total return, gold in yen, call-rate cash, frozen Faber rules, entry January 1990 β€” the month after the bubble top. Governing window 1999–2012: both legs dead on entry. Result: survived by 0.05 percentage points β€” CAGR 3.17% against a formula bar of cash+3pp = 3.12%, a margin inside data noise and called that on the record β€” and over the full 23 years it did exactly what he said it would: trod water (1.76% vs cash 1.27%, 15.8 years below cash). The eyeball finding is the keeper: equal-weight buy-and-hold of the same sleeves beat the trend rule by 1.6pp/yr in ZIRP β€” every whipsaw parks in an asset paying nothing β€” so the overlay subtracted return and bought drawdown instead (βˆ’10.0% max vs βˆ’15.9% basket, βˆ’62.8% Nikkei). The margin that survived came from the one sleeve that doesn’t run on interest rates: gold in yen, up 4.9x. His kicker is now the pre-registered meaning: in this regime the 20% sizing rule stops being humility and becomes the actual finding. Prediction audit: the house predicted treads-water and missed by the same five basis points; his registered calls β€” underperforms its own basket, treads water against cash+3 β€” scored correct and 5bp-from-correct. Those calls arrived after the numbers were already live on this page; they’re logged with that exact timestamp, not dressed as pre-registration, and scored anyway, because he asked to be held to a bar in public. A pre-registration that only counts when the clock is convenient isn’t one. His verdict is now the meaning on record, in his words: not a return engine that happens to survive bears β€” a drawdown engine that gives up return in flat regimes to buy survival in violent ones. It survives everywhere because it isn’t trying to win everywhere. Adopted with it: regime-transplant gates now bar against the basket’s own buy-and-hold too, and under that bar the overlay’s ZIRP verdict is a clean fail β€” the same sentence. Full receipts in the repo’s japan_experiment/. Co-author: Zestyclose.

GATES 27–31 Β· THE COLLECTOR'S SIDE OF THE COUNTER

the standing thesis after 26 gates: our slippage is somebody's revenue. the winners collect structural payments instead of predicting price, so the last five gates went at the collecting side directly, in one 30-hour sweep

Every family measured the same thing from a different angle: the payment is always real, and it is always priced to leave nothing above cash-compensation at retail scale.

the paymentmeasuredwhat's left at retail
Merger-arb spread (deal-break insurance)4.03% CAGR, 789 dealsbar was 4.32%; clears it only at zero cost
Cross-venue prediction-market gapgaps >1Β’ in 50–95% of hoursround trip costs 2.5–4Β’; nets +2.2%/yr < T-bills
Maker rebates / order flowreceipted fee schedules, every venueβˆ’0.17Β’/share; positive pay starts ~$10M/month

The merger-arb gate is the operator-vs-machine story: the operator registered an unhedged PASS four days before the run; the machine put it at 65% fail. The deal set was built from SEC filings announcement-first so the broken deals stayed in, and the famous tail turned out survivable (10% of deals broke, mean βˆ’2.9%, book max drawdown βˆ’5.5%). What killed it was quieter: the spread is priced almost exactly at cash+3 gross, and even 0.05%/side pushes it under. The 2019+ holdout was never opened; the frozen protocol seals it on a primary fail. Five validity rounds ran before any number was believed β€” the first three sims printed 7%, 27%, and 17%, each destroyed by its own eyeball audit (fictional payouts, acquirer tickers mis-cast as targets, a rebalancing rule quietly harvesting bid-ask bounce as alpha). The full defect ledger ships in the repo's merger_arb_experiment/DEVIATIONS.md, because a hostile reader should get the knife with the body.

The prediction-market gate found a second body on the way down: measured raw, the arb "returns" +29%/yr, and every fat entry sits in a market's first 72 hours, where one venue's price is a bookless print that collapses 30 cents on no news. No historical order-book exists to prove any of it was ever executable. Price without a book is not a quote; that readout ships as an unverifiable upper bound, and the mature-market number, the one you could actually have, pays less than the T-bills the capital would sit in. One venue also now charges a 5% taker fee on the exact markets measured: the toll booth raised its price mid-study.

The coda is the week's real lesson. The operator asked to beat the index by chasing momentum, so momentum got the full treatment: a hand-rolled top-10 rotation over 1,365 stocks and 13 cross-asset ETFs in one pool (his design: the defensive ETFs were supposed to take over in bears; in 55 months, including all of 2022, not one ETF ever out-ranked the ten hottest stocks), then a trend overlay on the index fund that already does momentum properly. Hand-roll: 7.14%/yr with a βˆ’41% drawdown. Professionally-run concentrated academic momentum: 12.4%, worse than the plain index. The boring S&P momentum index fund the whole exercise tried to beat: 19.5%/yr for a decade, live. The overlay halved its drawdown and paid 3.4pp/yr for the privilege against a 2pp cap written in advance. Every step toward "better" (more concentration, faster signals, cleverer math) made it worse, monotonically. The one keeper the hunt ever surfaced is a 13-basis-point product you buy and leave alone, and the machine's final advice on it was about sizing, not signals. All five gate files, predictions, and verdicts: merger_arb_experiment/ Β· predmarket_arb_experiment/ Β· orderflow_experiment/ Β· stock_rotation_experiment/ Β· spmo_overlay_experiment/.

GATES 32–35 Β· THE WRAPPER ARC

an outside collaborator β€” a prop-firm trader β€” brought two strategies and a thesis: trade the firm's capital through a challenge account. all of it went through the gate

The arc is the record in miniature: his playbook died, his spec produced the record's second pass, and the wrapper turned out to be priced like everything else.

The playbook first. "Smart money concepts" on 15-minute charts, mechanized with literature-default parameters and every rule frozen in writing: 7,087 trades, five years, two currency pairs β€” profit factor 0.951 BEFORE costs, and statistically indistinguishable from random entries at matched frequency, risk and exits. The zones are decoration. Then his trend spec β€” the surprise. Faber-class cross-asset trend, vol-targeted, run through evaluation rulesets on the real 2008–2026 path: Sharpe 0.82 at honest costs, 78% funded probability with ~$19k expected value per attempt against STATIC-drawdown rules. Against trailing-drawdown-plus-consistency rulesets it fails every cell β€” trend profits are lumpy, and consistency rules exist to deny exactly that shape. The firm's rule sheet, not the strategy, is the constraint.

Then the transfer test, and the cleanest attribution table in the record. Moving the spec onto FTMO's actual platform: dropping the bond leg (no bond CFDs) cost 0.05 of Sharpe; switching to single-name oil cost nothing; modeling the overnight financing a months-holding CFD book actually pays took it from 0.78 to 0.11. The venue's own swap charge is the entire death β€” the same kill as merger arbitrage, one venue closer to home. Swap-free static-drawdown firms exist and restore the passing regime on their published rules (76% funded, +$22k EV) β€” unless one ambiguous consistency clause in the payout policy applies, in which case two-thirds of payout cycles get denied. As of this writing the build waits on that sentence, in writing, from the firm. A companion rules study priced the wrapper itself: the published rule sheet is beatable by a coin flip (1-in-3 passes at zero edge, all the value in the funded-stage payout option), which is exactly why challenge passes are not receipts. The standing bar for any prop-firm claim: funded PAYOUT history over months. Files: smc_experiment/ Β· trendprop_experiment/ Β· propfirm_experiment/.

THE KILL-SIDE AUDIT Β· TURNING THE GAUNTLET ON OUR OWN TOMBSTONES

2026-07-29, prompted by the operator's question: "could the graveyard be artifacting kills?" β€” the record's skepticism had only ever been spent against passes; nobody had audited a loser tail

Criteria frozen before looking (what would count as an overturned kill), then: top-10 losers of three closed gates verified against independent price data, and every failed gate's cost ladder collated. Published whichever way it landed.

questionanswer
Any kill overturned?No. ORB 0/10 loser flags; mean reversion 1/10; trend-follow 2/10; corrections move PF by <0.01 against 1.30 bars.
Any fabricated losses?TWO β€” BHVN's "βˆ’91%" (Pfizer acquired Biohaven at $148.50; the ticker passed, gapless, to a ~$9 spinco) and CCXI's "βˆ’80%" (a hold jumped a 1,209-day delisting gap onto a relisted successor). Ticker splices fabricate losses exactly like they fabricate wins.
How cost-fragile are the kills?Less than the auditor predicted: 10 of 18 failed gates fail even at ZERO cost. Exactly one kill is cleanly cost-conditional β€” pairs, which passes its bars at 0.05%/side/leg.

The disclosure that matters, on the record: the kills share infrastructure β€” one cost model (calibrated from this pipeline's measured live slippage, always taker execution), one fill simulator, largely one data stack. Kills that die of friction weaken TOGETHER if the cost model is wrong for your execution; they are correlated verdicts, not 29 independent confirmations. Every results file publishes its full cost ladder so the count you trust is the one you recompute at your own costs. If your execution is materially better than retail taker, the pairs gate is the tombstone to re-run first. New standing rules from the audit: gap-scan and splice-screen every cached universe, and eyeball the top-10 losers with the same ritual as the winners β€” both tails, every gate. Full audit: kill_audit/ in the repo.

A backtest is a witness with every incentive to lie to you. Interrogate it like one.

exhibit b Β· the artifactsThree ways a backtest lies, from this file.

These aren't hypotheticals. Each one is an artifact my own pipeline produced, believed for a while, and then caught. Each one became a rule in the free checklist.

+4,890%one fabricated "winner"

The split artifact. Raw price bars plus a 1:25 reverse split inside the hold window turned a $0.17 stock into the trade of the decade. It never happened. A strategy "passed validation" for six days on trades like this before the split-adjusted rerun killed it.

rule: split-adjusted bars only, no exceptions

βˆ’15%where a βˆ’5% "stop" really fills

The fantasy fill. A backtest that fills your stop at the stop price is fiction: names that gap through a βˆ’5% stop routinely close down 10 to 15%. Half my "edges" were fill fantasies that evaporated under a checked-at-close, filled-at-close model.

rule: model the fill you actually get, not the one you want

PF 3.96printed twice, both fake

The false-positive machine. A slot-capped portfolio simulation printed profit factor 3.96 by accidentally cherry-picking under 10% of its own trades. It did this twice, on two unrelated strategies, before the pattern was caught and banned from the pipeline.

rule: slot-capped stats are a false-positive generator

exhibit c Β· the receiptsDon't take this page's word for it.

Every verdict above traces to a gate file written before its test ran. The receipts, directly:

artifactwhat it is
The full graveyard writeupall 35 gates, every number, the three expensive lessons
Gates 17–25: passes, benchings, and the resumed huntthe first PASS (and its reader-forced 2008 extension), the benched survivors, the falsified predictions (mine included), the 33-year options autopsy, and the strongest gross edge on record dying of friction
The artifact-hunting checklistthe free tier: every rule priced by the disaster that created it (landing in the repo this week)

the instrumentThree tiers. One is free forever.

The Checklist $0 in the repo now
  • The artifact-hunting checklist: every rule priced by the disaster that created it
  • Every gate writeup, full numbers
  • Works for manual backtesters too, no code required
LIVE Β· free forever Read the graveyard β†’
The Gauntlet $54
  • Honest-fill simulator: gap-through stops, cost ladders, close-fill discipline
  • Pre-registered gate templates (the goalpost welder)
  • De-survivorship universe builder + split-adjusted bar cache
  • The ETF gate harness: the code behind the only two passes
  • SEC EDGAR earnings fetcher, funding-rate carry simulator
  • 3 runnable case studies with real output
LIVE Β· checkout via Lemon Squeezy Get The Gauntlet →
The Field Manual $139
  • Everything in The Gauntlet
  • "10 ways your backtest is lying to you" with the fix code
  • Live slippage instrumentation (signal price vs. real fill)
  • Simulator smoke-test suite: test your tester first
  • Risk-control armor modules + the data proving they're load-bearing
  • In-play relative-volume screener, with its own falsification attached
  • Decision-shadow-logger: measure a filter before you trust it
LIVE Β· checkout via Lemon Squeezy Get The Field Manual →

the protocolThis launch is gated like everything else.

I pre-registered the launch the way I pre-register backtests: the paid tiers would ship only if people demonstrably wanted them. No fake scarcity, no countdown timer; the incentive ran the other way — packaging the kit properly took real weekends, and I'd only spend them on real demand.

Resolution (2026-07-27): the demand signal was real but modest — a four-round public audit of the flagship pass, a free-tier tester, workshop build requests — and honesty requires saying I called the gate rather than a counter tripping it. The packaging got finished, the store cleared review, and the stamps above are now live checkout links. If nobody buys, that verdict joins the ledger like every other one. The checklist tier stays free either way, and stars and issues still steer what gets built next.

What's deliberately NOT in the box


the voteWould you use it? Say so below.

Comments here are powered by GitHub: each one is filed as an issue on the graveyard repo, which means posting a comment goes on the public record. Say what you'd use, what's missing, or why the whole thing is wrong. Arguments count as interest.

Widget not loading (some previews block it)? Post directly as a GitHub issue β†’


the workshopAsk me to build something.

The best result on this page exists because a stranger demanded it. The Gate 17 extension β€” the one that repriced my only pass β€” started as a reader's challenge: "2022 is your easy case. Extend through 2008 or don't fund it." He was right, the run happened the same day, and the record got better by getting worse. That worked well enough to become a standing offer.

Bring me a mechanism and it gets the gate treatment: pass/fail criteria pre-registered before the run, honest fills, real costs, and the result published here β€” pass or kill β€” with credit to you. Or request a testing tool: tell me which backtest lie it should catch, and it goes on the kit backlog.

Ground rules, so nobody's time gets wasted: nothing that trades a live account gets built, ever. No signals sold back to you. Requests and results are public β€” that's the quality control. "It backtests well" is not a mechanism; tell me who's on the other side of the trade and why they're paying you, and your idea jumps the queue. I pick what gets built and promise no deadlines. Paid data needs a budget conversation before anything runs.

Request a build β†’   Files as a public GitHub issue β€” the queue is public by design.


provenanceWho's selling this?

One retail trader who spent months building a bot, watched honest testing kill every strategy it ran, and kept the machine that did the killing. The graveyard writeups converted the skeptics first, including the ones who showed up to argue. The receipts are public, the code produced them, and when my own pre-registered prediction got falsified by my own gate last week, that went in the repo too.

Questions, arguments, or a strategy you want dead? File a build request or open an issue on the repo. Arguing in public is the quality-assurance process.

Β© 2026 Β· strategygraveyard.com Β· The Falsification Kit Β· Educational software and research writeups. Not financial advice. No performance claims made or implied. Backtested results are historical simulations and do not predict future returns. The kit tests strategies; it does not provide, recommend, or execute them.