Skip to main content

Alternative Data

Updated Jun 24, 2026 at 2:35pm

  • 1494f0e59c07 Satellite & Geospatial Data 1 1,256
  • 14902901af46 Credit-Card & Transaction Data 1 1,169
  • 149149130920 Web-Traffic & App-Download Data 1 1,226
  • 14959f11dcaa NLP on Filings & Transcripts 1 1,073
  • 149201905c29 Social & News Sentiment 1 1,212
  • 149384e80a67 Alt-Data Edge Decay & Crowding 1 1,312
Tree Key
Expandable — has sub-topics
475Local Id for node
a1b2c3d4Click to see full UUID
6Sub-topics
7Documents
8.6k wordsResearch depth
5Open node
Research Draft High 1,380 words

Alternative data ("alt-data") is any dataset used for investment decisions that originates outside a company's own disclosures and the traditional financial plumbing — i.e. outside SEC filings, press releases, exchange price feeds, and sell-side estimates. The defining trait of useful alt-data is that it is behavioral or physical: it measures what people and assets are actually doing — cards swiped, cars in lots, app installs, website visits, words written — rather than what a company has chosen to report. The promise is timing and ground truth: a panel can reveal that a retailer's quarter is softening weeks before the 10-Q, letting a fund "nowcast" the print and position ahead of it. The core tension of the entire domain is that this edge is sampled, expensive, perishable, and legally fraught — the gap between a raw signal and a tradeable position runs through panel bias, entity mapping, crowding, decay, and MNPI/consent risk, and that gap is where nearly all the real skill (and danger) lives.

What this section covers

This is a section node. It defines the alt-data domain and maps its sub-topics; the depth lives in the children. The branch is organized by data modality — what the raw signal is made of — plus one cross-cutting node on why edges fade:

  • Satellite & geospatial data (001) — commercial Earth-observation imagery and analytics: counting cars in parking lots, gauging oil-tank fill from roof shadows, tracking crops, shipping, and mining. The best-evidenced category, with peer-reviewed support.
  • Credit-card & transaction data (002) — anonymized, aggregated card/bank-account panels resold as a near-real-time revenue proxy. The most directly revenue-mapped category and one of the most-used.
  • Web-traffic & app-download data (003) — modeled estimates of demand for a company's digital properties (visits, sessions, installs) as a leading proxy for digital-first revenue.
  • NLP on filings & transcripts (004) — converting unstructured corporate text (10-K/Q, 8-K, earnings-call Q&A) into tone, uncertainty, topic, and change-vs-prior-period signals.
  • Social & news sentiment (005) — turning news wires and social chatter (X, Reddit/WSB, StockTwits) into tone/intensity-of-attention signals; the most behavioral and most gameable.
  • Alt-data edge decay & crowding (006) — the cross-cutting reality that any monetized, distributed, or published signal erodes as it diffuses, and that crowded signals share synchronized-unwind tail risk.

A useful split: 001–003 are physical/behavioral activity proxies (they measure real-world demand and tend to map to a revenue line); 004–005 are text/attention signals (they measure information and mood). 006 governs all of them.

How it's sourced & turned into a signal

Across modalities the pipeline is structurally similar, and the value is added far more in the processing than in the raw capture:

1. Capture / acquisition — satellites image the Earth; bank-aggregation apps and card panels harvest de-identified transactions; crawlers scrape sites and stores; vendors ingest filings and social feeds. 2. Cleaning, de-identification & normalization — the hard, skill-intensive middle: correcting panel bias and coverage drift, adjusting imagery for sun angle/weather, deduplicating, panel-normalizing same-cohort spend. 3. Entity resolution / mapping to a security — attributing noisy merchant descriptors, store footprints, or web domains to the right ticker and segment. This step is error-prone and is often the real moat. 4. Modeling the bridge — translating the panel's change into an estimate of the company's reported number, validated against the eventual print and recalibrated.

The recurring lesson: the durable edge has shifted from data exclusivity to operational capability — entity resolution, bias correction, and speed of execution — because vendor prices have collapsed as datasets commoditize (the credit-card node cites a feed that ran ~$500K/yr in 2015 trading near ~$5K by 2026; practitioner figure, not audited).

Adoption, debate & evidence

Institutional adoption is broad and growing, but the precise percentages vary widely by source, year, and respondent pool — and are vendor/survey-relayed, so they should be treated as indicative, not audited. The annual Lowenstein Sandler surveys, for example, show alt-data use among their respondents rising from ~62% (2023) to ~67% (2024) to roughly four-fifths (2025), while Coalition Greenwich's 2025 study (56 buy-side firms) found about three-quarters using non-traditional data and roughly two-thirds planning to increase spend. The directional consensus — a clear majority of professional investors now use some alt-data and most plan to spend more — is robust; the exact headline figure is not, and different surveys disagree by 10-15 points. Market-size estimates are even noisier: 2024–2025 figures cluster around the low tens of billions of dollars with double-digit-plus projected CAGRs, but forecasts diverge enormously across research firms (Grand View; IMARC) — cite the existence of growth, not a specific number.

On measured edge, the honest picture is uneven across the children. The strongest peer-reviewed support is physical: the Berkeley satellite parking-lot study (RS Metrics imagery, 44 retailers, 2011–2017) documented ~4–5% returns in the three days around earnings, and Mukherjee et al.'s "Eye in the Sky" (JFE 2021) shows satellite estimates now reduce the surprise of official macro releases. Real-time-sales nowcasting has academic support too — but the canonical Froot, Kang, Ozik & Sadka (JFE 2017) result is built on web/mobile-traffic data, not card panels (a look-alike that must not borrow card data's credibility). Text-based signals (NLP, sentiment) are widely used but have the narrowest gap between genuine edge and overfitting/crowding artifact. The unifying caveat is the 006 node's: published or widely distributed edges decay — McLean & Pontiff (2016) found predictor returns ~58% lower post-publication — so backtested alt-data Sharpe ratios should be read as upper bounds.

Strengths & limitations

Strengths. Timing (front-runs lagging official stats and quarterly disclosure), directness (it measures the actual demand driver, not a derivative of it), and — for physical data — causal grounding (cars and oil are not sentiment). A verifiable ground truth (the eventual reported number) lets models be backtested and recalibrated.

Limitations and the #1 misuse. Every alt-data read is a partial sample, not a census: card panels miss cash/online/international, satellites see only the physical slice, web panels skew demographically. The single most common misuse across the domain is reading a signal as a level and a one-click trade ("revenue will be X, so buy") rather than a direction/conviction input — ignoring that the panel-to-reported bridge drifts, that the read may already be priced by better-capitalized players, and that data + execution costs can consume a modest, decaying edge. A close second is legal/compliance complacency: the App Annie enforcement (SEC, 2021, $10M — the first action against an alt-data provider, over misrepresented sourcing) and the SEC Division of Examinations' April 2022 risk alert (Section 204A/MNPI deficiencies) put alt-data squarely within MNPI/consent liability. Vendor diligence on data provenance and aggregation is mandatory, not optional.

Sources

Dispute flags: adoption percentages and market-size figures are vendor/survey-relayed and diverge widely — treat as indicative only. The strongest measured edges are in physical (satellite) and web/mobile-traffic data; card-data-specific and text-signal alpha rest more on vendor backtests than peer review. All documented edges are subject to post-publication/crowding decay (006), so backtested figures are upper bounds.