Alternative Data
Tree Key
Alternative data ("alt-data") is any dataset used for investment decisions that originates outside a company's own disclosures and the traditional financial plumbing — i.e. outside SEC filings, press releases, exchange price feeds, and sell-side estimates. The defining trait of useful alt-data is that it is behavioral or physical: it measures what people and assets are actually doing — cards swiped, cars in lots, app installs, website visits, words written — rather than what a company has chosen to report. The promise is timing and ground truth: a panel can reveal that a retailer's quarter is softening weeks before the 10-Q, letting a fund "nowcast" the print and position ahead of it. The core tension of the entire domain is that this edge is sampled, expensive, perishable, and legally fraught — the gap between a raw signal and a tradeable position runs through panel bias, entity mapping, crowding, decay, and MNPI/consent risk, and that gap is where nearly all the real skill (and danger) lives.
What this section covers
This is a section node. It defines the alt-data domain and maps its sub-topics; the depth lives in the children. The branch is organized by data modality — what the raw signal is made of — plus one cross-cutting node on why edges fade:
- Satellite & geospatial data (
001) — commercial Earth-observation imagery and analytics: counting cars in parking lots, gauging oil-tank fill from roof shadows, tracking crops, shipping, and mining. The best-evidenced category, with peer-reviewed support. - Credit-card & transaction data (
002) — anonymized, aggregated card/bank-account panels resold as a near-real-time revenue proxy. The most directly revenue-mapped category and one of the most-used. - Web-traffic & app-download data (
003) — modeled estimates of demand for a company's digital properties (visits, sessions, installs) as a leading proxy for digital-first revenue. - NLP on filings & transcripts (
004) — converting unstructured corporate text (10-K/Q, 8-K, earnings-call Q&A) into tone, uncertainty, topic, and change-vs-prior-period signals. - Social & news sentiment (
005) — turning news wires and social chatter (X, Reddit/WSB, StockTwits) into tone/intensity-of-attention signals; the most behavioral and most gameable. - Alt-data edge decay & crowding (
006) — the cross-cutting reality that any monetized, distributed, or published signal erodes as it diffuses, and that crowded signals share synchronized-unwind tail risk.
A useful split: 001–003 are physical/behavioral activity proxies (they measure real-world demand and tend to map to a revenue line); 004–005 are text/attention signals (they measure information and mood). 006 governs all of them.
How it's sourced & turned into a signal
Across modalities the pipeline is structurally similar, and the value is added far more in the processing than in the raw capture:
1. Capture / acquisition — satellites image the Earth; bank-aggregation apps and card panels harvest de-identified transactions; crawlers scrape sites and stores; vendors ingest filings and social feeds. 2. Cleaning, de-identification & normalization — the hard, skill-intensive middle: correcting panel bias and coverage drift, adjusting imagery for sun angle/weather, deduplicating, panel-normalizing same-cohort spend. 3. Entity resolution / mapping to a security — attributing noisy merchant descriptors, store footprints, or web domains to the right ticker and segment. This step is error-prone and is often the real moat. 4. Modeling the bridge — translating the panel's change into an estimate of the company's reported number, validated against the eventual print and recalibrated.
The recurring lesson: the durable edge has shifted from data exclusivity to operational capability — entity resolution, bias correction, and speed of execution — because vendor prices have collapsed as datasets commoditize (the credit-card node cites a feed that ran ~$500K/yr in 2015 trading near ~$5K by 2026; practitioner figure, not audited).
Adoption, debate & evidence
Institutional adoption is broad and growing, but the precise percentages vary widely by source, year, and respondent pool — and are vendor/survey-relayed, so they should be treated as indicative, not audited. The annual Lowenstein Sandler surveys, for example, show alt-data use among their respondents rising from ~62% (2023) to ~67% (2024) to roughly four-fifths (2025), while Coalition Greenwich's 2025 study (56 buy-side firms) found about three-quarters using non-traditional data and roughly two-thirds planning to increase spend. The directional consensus — a clear majority of professional investors now use some alt-data and most plan to spend more — is robust; the exact headline figure is not, and different surveys disagree by 10-15 points. Market-size estimates are even noisier: 2024–2025 figures cluster around the low tens of billions of dollars with double-digit-plus projected CAGRs, but forecasts diverge enormously across research firms (Grand View; IMARC) — cite the existence of growth, not a specific number.
On measured edge, the honest picture is uneven across the children. The strongest peer-reviewed support is physical: the Berkeley satellite parking-lot study (RS Metrics imagery, 44 retailers, 2011–2017) documented ~4–5% returns in the three days around earnings, and Mukherjee et al.'s "Eye in the Sky" (JFE 2021) shows satellite estimates now reduce the surprise of official macro releases. Real-time-sales nowcasting has academic support too — but the canonical Froot, Kang, Ozik & Sadka (JFE 2017) result is built on web/mobile-traffic data, not card panels (a look-alike that must not borrow card data's credibility). Text-based signals (NLP, sentiment) are widely used but have the narrowest gap between genuine edge and overfitting/crowding artifact. The unifying caveat is the 006 node's: published or widely distributed edges decay — McLean & Pontiff (2016) found predictor returns ~58% lower post-publication — so backtested alt-data Sharpe ratios should be read as upper bounds.
Strengths & limitations
Strengths. Timing (front-runs lagging official stats and quarterly disclosure), directness (it measures the actual demand driver, not a derivative of it), and — for physical data — causal grounding (cars and oil are not sentiment). A verifiable ground truth (the eventual reported number) lets models be backtested and recalibrated.
Limitations and the #1 misuse. Every alt-data read is a partial sample, not a census: card panels miss cash/online/international, satellites see only the physical slice, web panels skew demographically. The single most common misuse across the domain is reading a signal as a level and a one-click trade ("revenue will be X, so buy") rather than a direction/conviction input — ignoring that the panel-to-reported bridge drifts, that the read may already be priced by better-capitalized players, and that data + execution costs can consume a modest, decaying edge. A close second is legal/compliance complacency: the App Annie enforcement (SEC, 2021, $10M — the first action against an alt-data provider, over misrepresented sourcing) and the SEC Division of Examinations' April 2022 risk alert (Section 204A/MNPI deficiencies) put alt-data squarely within MNPI/consent liability. Vendor diligence on data provenance and aggregation is mandatory, not optional.
Sources
- Wikipedia, Alternative data (finance) — definition, taxonomy (behavioral/physical/digital), web-scraping legal ambiguity: https://en.wikipedia.org/wiki/Alternative_data_(finance)
- Coalition Greenwich, Alternative data 2025: Fueling the AI-driven investment revolution (~three-quarters of buy-side firms use non-traditional data; 56-firm sample, vendor/survey): https://www.greenwich.com/market-structure-technology/alternative-data-2025-fueling-ai-driven-investment-revolution
- Lowenstein Sandler annual Alternative Data surveys (adoption ~62% in 2023, ~67% in 2024, ~four-fifths in 2025 among respondents — vendor/survey, indicative): https://www.lowenstein.com/news-insights/firm-news/use-of-alternative-data-in-investment-community-shows-no-signs-of-slowing-according-to-new-survey-by-lowenstein-sandler-s-investment-management-group
- Market-size estimates (divergent; cite growth not a number): Grand View Research https://www.grandviewresearch.com/industry-analysis/alternative-data-market ; IMARC https://www.imarcgroup.com/alternative-data-market
- Sidley Austin, SEC Fines App Annie $10M for Securities Fraud (MNPI/sourcing misrepresentation, 2021): https://www.sidley.com/en/insights/newsupdates/2021/09/sec-fines-app-annie-inc-10-million-for-securities-fraud
- Debevoise / NYU Compliance & Enforcement, SEC April 2022 Risk Alert on Alternative Data (MNPI, Section 204A): https://www.debevoisedatablog.com/2022/05/02/sec-risk-alert-alternative-data/
- McLean & Pontiff (2016), Does Academic Research Destroy Stock Return Predictability?, Journal of Finance 71(1):5-32 — predictor returns ~26% lower out-of-sample and ~58% lower post-publication (applied here by analogy to alt-data edge decay): https://onlinelibrary.wiley.com/doi/abs/10.1111/jofi.12365
- Child nodes (this branch) for measured base rates:
001satellite (Berkeley/Katona et al.; Mukherjee et al. JFE 2021),002card data (Froot et al. JFE 2017 — web/mobile-traffic proxy, not card panels),006edge decay & crowding.
Dispute flags: adoption percentages and market-size figures are vendor/survey-relayed and diverge widely — treat as indicative only. The strongest measured edges are in physical (satellite) and web/mobile-traffic data; card-data-specific and text-signal alpha rest more on vendor backtests than peer review. All documented edges are subject to post-publication/crowding decay (006), so backtested figures are upper bounds.