Quantitative Research
Quantitative Research
AI Summary
Not all alternative data is created equal. We evaluate the signal persistence, decay characteristics, and capacity constraints of satellite, credit card, and web-scraping datasets.
The alternative data industry has grown from a niche curiosity to a multi-billion dollar market in less than a decade. Hedge funds, asset managers, and quantitative shops now spend hundreds of millions annually on satellite imagery, credit card transaction feeds, web-scraped pricing data, social media sentiment, and dozens of other non-traditional datasets. The promise is compelling: information that is not yet reflected in prices, available before the quarterly earnings cycle, and processable at scale by machine learning systems.
The fundamental challenge with alternative data is that its alpha decays rapidly once it becomes widely adopted. A satellite imagery dataset that provided genuine edge in 2018 — when only a handful of funds had the infrastructure to process it — provides far less edge in 2026, when it is available from multiple vendors and processed by hundreds of systematic strategies. The signal has been arbitraged away.
Our evaluation framework assesses each alternative dataset on three dimensions: (1) signal persistence — how long does the predictive relationship hold after controlling for crowding? (2) decay rate — how quickly does the alpha erode as adoption increases? (3) capacity — at what AUM level does the strategy become self-defeating due to market impact? Datasets that score poorly on any of these dimensions are excluded from our production models regardless of their in-sample performance.
After rigorous out-of-sample testing, we find that three categories of alternative data retain meaningful signal in the current environment. First, proprietary web-scraped pricing data for consumer goods and services — particularly in markets where official inflation statistics lag reality — continues to provide a 2–4 week lead on CPI surprises. Second, shipping and logistics data, including port congestion metrics and container booking rates, provides early warning of supply chain disruptions that affect manufacturing PMIs. Third, job posting data from major employment platforms provides a real-time proxy for corporate hiring intentions that leads official employment data by 6–8 weeks.
By contrast, we have largely abandoned social media sentiment as a standalone signal. The noise-to-signal ratio has deteriorated significantly as bot activity, coordinated retail trading campaigns, and platform algorithm changes have made the data increasingly unreliable. It retains some value as a contrarian indicator at extremes — peak bullish sentiment has historically preceded short-term reversals — but its predictive power for fundamental outcomes is minimal.
Even high-quality alternative data is only as valuable as the infrastructure used to process it. Raw alternative datasets are typically messy, inconsistently formatted, and riddled with survivorship bias. The data engineering required to clean, normalise, and backfill these datasets is substantial — and the risk of inadvertently introducing look-ahead bias during the cleaning process is significant. We have seen multiple cases where apparently strong backtested signals evaporated entirely once the data pipeline was audited for look-ahead contamination.
Our conclusion is that alternative data is a genuine source of edge — but only for organisations with the data engineering discipline to process it correctly, the analytical rigour to test it honestly, and the patience to wait for datasets that have not yet been fully arbitraged. The bar is higher than most vendors suggest, and the graveyard of failed alternative data initiatives is larger than the industry acknowledges.
Weekly institutional research, signal updates, and market analysis. No spam, unsubscribe anytime.
Related Research
Quantitative Research
Institutional Research
Request Access