Skip to main content

Why Anomalies Persist (or Decay)

Updated Jun 23, 2026 at 8:47pm

Research Draft Medium 1,364 words

A documented "edge" is not a permanent fact about markets — it is a claim that lives on a spectrum from durable to already-arbitraged-away to never real in the first place. This is the meta-framework for judging where a given anomaly sits. Two opposing forces govern every published edge: forces that let real mispricings persist (limits to arbitrage, entrenched behavioral biases, genuine risk compensation, capacity/career constraints) and forces that make them decay (publication and crowding, falling trading costs, regime change) — plus the uncomfortable possibility that the "edge" was a data-snooping artifact and was never there. The core tension: the same evidence (a backtest with a strong return) is consistent with all three states, so the question is never "does the backtest look good?" but "why would this survive, and has it?"

Why anomalies persist

A real mispricing can survive for years for four distinct reasons, and the reason matters because it tells you whether to expect decay.

  • Limits to arbitrage (Shleifer & Vishny, 1997). Textbook arbitrage is riskless and costless; real arbitrage is neither. Professional arbitrage is done by a small number of specialists running other people's money on poorly diversified positions. They face noise-trader risk — the mispricing can widen before it closes — and because clients withdraw capital exactly when a position moves against the arbitrageur (performance-based withdrawals), arbitrageurs are forced out at the worst time. The result: arbitrage is least effective precisely when mispricings are largest, so prices can stay wrong. This is the foundational reason a genuine behavioral mispricing does not self-correct.
  • Persistent behavioral biases. If an anomaly is driven by hardwired investor behavior (overreaction to value/growth, anchoring, disposition effect, limited attention), and the bias does not disappear when documented, the mispricing can recur. Behavior is stickier than knowledge.
  • Risk-based explanation (the "anomaly" is a risk premium). Under an efficient-markets reading (Fama), an apparent anomaly is simply compensation for risk the asset-pricing model failed to capture. The classic tell: size- and value-sorted portfolios that show large CAPM alphas show near-zero alphas under the Fama-French model — the excess return was a factor loading, not a free lunch. A true risk premium should not decay when published, because you are being paid to bear real, undiversifiable risk. Persistence is therefore evidence for a risk-based interpretation.
  • Capacity constraints & career risk. Some edges are real but too small to scale (microcaps, illiquid names, high turnover) — big money cannot deploy size without moving the price, so the edge survives in a corner too small to be worth arbitraging. Career/agency risk compounds this: a manager who tracks a benchmark won't hold a contrarian position long enough to harvest it.

Why anomalies decay

  • Publication and crowding — McLean & Pontiff (2016). The single most important empirical result here. Reconstructing ~97 published cross-sectional predictors, they found predictor portfolio returns were about 26% lower out-of-sample (after the original study's sample period but before publication) and about 58% lower post-publication. The gap implies roughly a 32% (58% − 26%) decline attributable to publication itself — i.e. arbitrageurs learning about the mispricing and trading it away, beyond any data-mining shrinkage. Consistent with arbitrage: post-publication decline is larger for predictors with higher in-sample returns and for those easier/cheaper to trade, and post-publication periods show increased trading volume and short interest in the relevant stocks.
  • Data-snooping — it was never real. Some "anomalies" decay to zero out-of-sample because they were statistical flukes. Harvey, Liu & Zhu (2016, "…and the Cross-Section of Expected Returns") catalogued 300+ published factors and argued that, given how many were tested, the conventional t-statistic > 2.0 hurdle is far too lenient; they propose t > ~3.0 as a minimum for a new factor to be credible. Sullivan, Timmermann & White (1999) applied White's Reality Check (a multiple-testing bootstrap) to ~100 years of Dow data and found that the best technical trading rules from Brock-Lakonishok-LeBaron, while impressive in-sample even after snooping adjustment, did not outperform in the subsequent out-of-sample decade, and showed no edge on S&P 500 futures once data-snooping was accounted for. The lesson: with enough rules tried, something will look great by chance.
  • Arbitrage capital & falling trading costs. Cheaper, faster, more abundant arbitrage capital (declining commissions, sub-penny spreads, ETFs, quant funds) lowers the cost of correcting mispricings, so edges that survived on transaction-cost frictions erode as those frictions fall.
  • Structural / regime change. Rule changes (decimalization, Reg NMS), changed market microstructure, or a different macro regime can kill an edge that depended on the old structure — independent of crowding.

The decision-useful checklist (for the system / Augustus)

Treat any backtested edge as decayed-until-proven, and discount any published edge by default. Before trusting an anomaly, answer:

1. Risk-based or behavioral? If it's plausibly a risk premium, expect persistence but also expect real drawdowns (you're paid to suffer). If it's behavioral, ask whether the bias still operates and whether arbitrageurs have crowded in. 2. Published and how long ago? If published, apply a McLean-Pontiff-style haircut — assume roughly half the in-sample return is gone, more if the original return was large or the trade is cheap to execute. 3. Out-of-sample survival. Has it held up after the original sample and after publication? An edge that survives both is in a different class from a fresh backtest. 4. Multiple-testing exposure. How many variants were tried to find it? A t-stat near 2 on a heavily-searched space is probably noise (Harvey-Liu-Zhu; Sullivan-Timmermann-White). 5. Capacity & costs. Does it only work in tiny/illiquid names or at high turnover? Then it may be real but not deployable at size — and net of costs it may be gone. 6. Regime dependence. Did the structure or macro regime it relied on still hold?

If an edge cannot clear most of these, the corpus should down-weight it rather than let a stale or false anomaly contaminate live decisions.

Strengths & limitations

The framework's strength is that it protects the entire anomalies corpus from its most dangerous failure mode: trusting a number from a backtest that the market has already neutralized (or that was never real). Its limitation is that the categories overlap and are hard to disentangle in real time — persistence is genuinely ambiguous: it can mean "real risk premium" or "limits to arbitrage are still binding" or "not enough out-of-sample data yet to see the decay." Likewise, decay does not prove the edge was fake; a real edge that gets crowded also decays. The honest position is probabilistic, not binary. The single most common misuse is treating a strong in-sample backtest as a live edge without applying any out-of-sample or publication discount.

System relevance

This is a capstone/meta node for the Market Anomalies & Documented Edges section: it sets the discount rule every sibling node inherits. For Augustus, the operational rule is: when an anomaly node reports a historical effect size, do not consume it at face value — apply the publication/out-of-sample haircut from the checklist above, and prefer Cairn's measured live track record over any published or backtested number whenever the two conflict. A documented base rate is a prior to be decayed, not a fact. Cross-link: the individual anomaly nodes in this section (each should be read through this filter) and any limits-to-arbitrage or market-efficiency definition node.

Sources