Skip to main content

Knowing When an Edge Is Dead

Updated Jun 24, 2026 at 8:22pm

Research Draft High 1,301 words

Every trading edge is conditional and finite — it works because of some structural inefficiency (a behavioral bias, a liquidity gap, a slow-to-react participant), and edges decay as that condition fades, gets crowded out, or turns out never to have existed. "Knowing when an edge is dead" is the discipline of distinguishing a genuinely expired edge from a normal, statistically-expected drawdown — because the same losing streak can mean "keep trading, this is variance" or "stop, the inefficiency is gone," and the cost of guessing wrong in either direction is large. Quitting a live edge during routine variance forfeits the strategy's whole expected value; clinging to a dead edge bleeds capital indefinitely. There is no clean test that resolves this in real time — the core problem is that decay and bad luck look identical for a long while.

Why edges die

Practitioners and researchers converge on roughly three causes (Maven Securities; Trading Engineering Lab):

  • The edge gets crowded. As more participants trade the same signal, entries worsen, slippage rises, and the inefficiency is arbitraged toward zero. This is the dominant decay mode for any published or widely-known edge.
  • Market structure changes. Shifts in liquidity, volatility regime, tick size, participant mix, or execution venues can remove the conditions the edge depended on. An edge can be perfectly alive in one regime and dead in another without ever being "discovered."
  • The edge was never real. It was overfit to backtest noise, or a product of survivorship/data-mining bias. Here there is nothing to "die" — live trading simply reveals the truth that the backtest hid.

The single most important practical consequence: an edge that was discovered by you alone in live data dies for different reasons (regime change) than an edge that was published or is broadly known (crowding). The latter has a documented, near-inevitable decay path.

How decay is detected and measured

There is no single accepted test; the credible approaches all compare live performance against a properly-bounded expectation built before the drawdown.

  • Expected-drawdown / time-under-water bounds (Bailey & López de Prado). The most rigorous public framework. For a given Sharpe ratio you can derive the expected maximum drawdown and expected time-under-water at a chosen confidence level. Their "Triple Penance Rule" (Bailey & López de Prado, 2014) states that, under standard (iid) portfolio-theory assumptions, recovery from a drawdown takes roughly three times the period it took to form — hence "triple penance"; positive serial correlation can lengthen this further, so practitioners often cite a ~2–3× range. Crucially, they argue you should monitor recovery time, not just loss size: a small loss that fails to recover within its expected window is a stronger decay signal than a large-but-on-schedule loss.
  • Rolling performance vs in-sample baseline. Track rolling Sharpe, win rate, average R-multiple, and profit factor over a moving window and compare to the strategy's established distribution. A commonly cited robustness guideline is the "70% rule" — expect live/out-of-sample Sharpe to reach at least ~70% of in-sample, and treat persistent shortfalls below that as suspect (quantvps.com; goatfundedtrader.com). Note this is a practitioner heuristic, not a validated threshold.
  • Monte Carlo / bootstrap on trade order. Randomizing the trade sequence thousands of times yields the distribution of drawdowns the strategy can produce by chance. If the live drawdown sits inside that distribution, it is consistent with variance; if it exceeds (say) the 95th–99th percentile of simulated drawdowns, that is evidence — not proof — of breakdown.
  • Component diagnostics. Decompose the slump: did the win rate drop, the average winner shrink, or costs/slippage rise? Rising slippage on unchanged signal quality points to crowding; a collapsing win rate points to the signal itself failing. This is more diagnostic than the aggregate P&L curve.

All of these convert "I feel like it stopped working" into "live results are at the Nth percentile of what this edge should produce" — which is the only honest basis for the call.

How it's used in practice

Because the statistics are slow, most disciplined traders pre-commit to kill criteria defined before deployment, when judgment is unclouded by the open position. Typical components:

  • A maximum drawdown / time-under-water limit tied to a confidence level, beyond which the strategy is halted for review (the López de Prado stop-out logic).
  • A walk-forward / out-of-sample monitor: the strategy keeps running on a rolling basis and is re-validated; if recent out-of-sample stats fall persistently below the in-sample baseline, it is retired or down-weighted.
  • Position-size de-risking before a binary kill. Rather than a single on/off decision, many quant shops scale exposure down as confidence in the edge erodes (e.g. proportional to a decaying rolling Sharpe), which limits the cost of a wrong "it's still alive" call.

The governing principle is to make the decision rules-based and pre-registered so it isn't made emotionally mid-drawdown — which is precisely when traders both capitulate on good edges and rationalize keeping bad ones.

Standing & evidence

That known edges decay is one of the better-documented findings in finance. McLean and Pontiff (2016, Journal of Finance) studied 97 cross-sectional return predictors and found portfolio returns roughly 26% lower out-of-sample and 58% lower post-publication, consistent with sophisticated traders arbitraging the signal once it is public (they attribute the ~26% to data-mining/in-sample bias and the additional ~32% specifically to publication-informed trading). Jacobs & Müller (2020, Journal of Financial Economics), studying 241 anomalies across 39 markets, found a reliable post-publication decline only in the United States — which both corroborates the publication-arbitrage story in the most-traded market and shows anomalies persist where arbitrage is harder. So "decays substantially after publication in deep, arbitraged markets like the U.S." is well-supported; uniform global disappearance is not. The takeaway: any edge you can read about should be assumed to be in decline wherever it can be cheaply arbitraged; the uncertainty is the rate, not the direction.

Strengths & limitations

The strength of a formal decay framework is that it replaces a panic-or-hope reaction with a probability statement and a pre-committed rule. Its hard limitation is low statistical power: distinguishing a dead edge from a deep-but-normal drawdown often requires more trades than a discretionary trader will ever take, and by the time the data is conclusive, much of the damage is done. Confidence bounds also assume the past return distribution still describes the future — the very thing in question when an edge may be dying, making the test partly circular. The single most common misuse is abandoning an edge at the bottom of an ordinary drawdown — capitulating right before reversion — which is the mirror image of, and statistically just as likely as, the error of holding a dead one too long. A close second is treating the absence of a kill signal as confirmation the edge is healthy.

System relevance

This node anchors the kill-criteria end of the Edge Lifecycle & Building a Trading System branch; pair it with the sibling nodes on edge sourcing, backtesting/overfitting, and walk-forward validation. For Augustus, the operative caveat is that any setup it consumes from this corpus carries an unknown, generally-declining live edge — so it must defer to Cairn's measured live track record over any backtested or folklore base rate, and weight a setup down when Cairn's recent realized performance for that pattern sits in the lower tail of its historical distribution rather than waiting for a single hard "dead" verdict.

Sources

  • Bailey, D. & López de Prado, M. — "Stop-Outs Under Serial Correlation and the Triple Penance Rule"; and "Drawdown-Based Stop-Outs and the Triple Penance Rule" (expected max drawdown / time-under-water; recovery ≈ 2–3× formation time). davidhbailey.com; arXiv:1707.01457.
  • McLean, R.D. & Pontiff, J. (2016) — "Does Academic Research Destroy Stock Return Predictability?", Journal of Finance (26% lower out-of-sample, 58% lower post-publication). onlinelibrary.wiley.com / SSRN 2156623.
  • Jacobs, H. & Müller, S. (2020) — "Anomalies across the globe: Once public, no longer existent?", Journal of Financial Economics 135(1), 213–230 (241 anomalies, 39 markets; reliable post-publication decline only in the U.S.). ScienceDirect; SSRN 2816490.
  • Maven Securities — "Alpha Decay: what does it look like?" (causes: crowding, structure, never-real). mavensecurities.com.
  • Trading Engineering Lab — "Alpha Decay: Why Strategies Stop Working Over Time."
  • quantvps.com; goatfundedtrader.com — rolling Sharpe/drawdown monitoring, "70% rule," Monte Carlo trade-order randomization (practitioner heuristics).

Confidence: medium. Nuance: McLean-Pontiff and Jacobs-Müller agree on U.S. post-publication decay but disagree on how global/uniform it is. The "70% rule" and percentile cutoffs are practitioner conventions, not validated thresholds; the Triple Penance "3×" holds under iid assumptions and varies with serial correlation.