From Word Counts
to Context
Topic Models for Asset Pricing
University of California, Berkeley · *Equal contribution · †Corresponding author
Supervised by Professor Ali Kakhbod
Financial news
Contextual topics
Narrative risk
Does adding context to financial-news topic models produce more coherent narratives and change the risk factors we recover from them?
Abstract
One pipeline.
Two ways to read.
News may reveal systematic risk, but it is still unclear whether context helps us build better risk factors. We compare Latent Dirichlet Allocation, which reads word counts, with a frozen sentence transformer on the same 394,661 financial-news articles. The rest of the portfolio construction pipeline stays fixed.
The contextual branch produces higher observed topic coherence and financial performance, though the available tests do not establish outperformance. Exploratory spherical clustering and multi-horizon exposures reach an observed excess-return Sharpe of 1.03. These results are promising, but stricter point-in-time tests and broader datasets are needed.
Method
Text becomes
an exposure.
We change the text representation and keep the financial machinery fixed, isolating what a more contextual reading of news contributes.
- 01
Read the news
394,661 FNSPID articles from 2015 to 2023 are cleaned, deduplicated, and mapped to daily topic attention.
- 02
Discover narratives
LDA models word counts. The alternative embeds passages with a frozen sentence transformer and clusters them into 60 semantic topics.
- 03
Measure surprises
Unexpected deviations from a lagged trailing average become zero-mean narrative shock signals.
- 04
Price the cross-section
Rolling stock-to-shock covariances enter Sparse IPCA, which extracts three latent factors and an out-of-sample tangency portfolio.

Results
Context raises the
observed signal.
We separate the controlled comparison from the exploratory extensions. Sharpe values cover 41 out-of-sample months, from September 2020 to January 2024.
Spherical clustering, sparse-soft topic terms, and 20/60/120-day exposure horizons form the strongest observed specification.
Topic atlas
Inspect all 60 narratives.
Select a text layer to compare its topic-level coherence distribution and sample term lists. Each dot is one learned topic.
Interpretation matters. The central LDA versus Frozen ST comparison changes both the representation and how much article text each model sees. Full-sample standardization and topic construction also call for a stricter point-in-time replication. We report observed performance, not established outperformance.
Interactive demo
Watch a topic model
sort the words.
A small, seeded LDA sampler learns from 12 synthetic documents. Run a few sweeps and watch the topic assignments settle. Everything runs in your browser.
Words are colored by their current topic assignment. The example is synthetic and does not contain empirical paper weights, financial estimates, or remote requests.
Full paper
Methods, robustness,
and limitations.
The 20-page working paper documents the corpus, topic construction, narrative shocks, Sparse IPCA design, out-of-sample evaluation, and topic-count ablations.
Open PDF · 20 pages ↗to Context Topic Models for Asset Pricing UC Berkeley
Citation
Build on the work.
@article{foley2026wordcounts,
title = {From Word Counts to Context: Topic Models for Asset Pricing},
author = {Foley, Kevin and Hartadi, Jonathan and Prakash, Shivesh and Vatsal, Swapnil},
year = {2026},
note = {Working paper, University of California, Berkeley}
}