From Word Counts
to Context

Topic Models for Asset Pricing

Kevin Foley* Jonathan Hartadi* Shivesh Prakash*,† Swapnil Vatsal*

University of California, Berkeley · *Equal contribution · †Corresponding author

Supervised by Professor Ali Kakhbod

01
MARKETS · 08:42 Central bank signals a gradual path as energy prices rise

Financial news

02

Contextual topics

03

Narrative risk

Central question

Does adding context to financial-news topic models produce more coherent narratives and change the risk factors we recover from them?

Abstract

One pipeline.
Two ways to read.

News may reveal systematic risk, but it is still unclear whether context helps us build better risk factors. We compare Latent Dirichlet Allocation, which reads word counts, with a frozen sentence transformer on the same 394,661 financial-news articles. The rest of the portfolio construction pipeline stays fixed.

The contextual branch produces higher observed topic coherence and financial performance, though the available tests do not establish outperformance. Exploratory spherical clustering and multi-horizon exposures reach an observed excess-return Sharpe of 1.03. These results are promising, but stricter point-in-time tests and broader datasets are needed.

Narrative asset pricingSentence transformersLDASparse IPCA

Method

Text becomes
an exposure.

We change the text representation and keep the financial machinery fixed, isolating what a more contextual reading of news contributes.

  1. 01

    Read the news

    394,661 FNSPID articles from 2015 to 2023 are cleaned, deduplicated, and mapped to daily topic attention.

  2. 02

    Discover narratives

    LDA models word counts. The alternative embeds passages with a frozen sentence transformer and clusters them into 60 semantic topics.

  3. 03

    Measure surprises

    Unexpected deviations from a lagged trailing average become zero-mean narrative shock signals.

  4. 04

    Price the cross-section

    Rolling stock-to-shock covariances enter Sparse IPCA, which extracts three latent factors and an out-of-sample tangency portfolio.

End-to-end pipeline: financial news, semantic representation, attention shocks, stock exposures, Sparse IPCA, and systematic narrative risk.
Figure 1 End-to-end narrative asset-pricing architecture. View in paper ↗

Results

Context raises the
observed signal.

We separate the controlled comparison from the exploratory extensions. Sharpe values cover 41 out-of-sample months, from September 2020 to January 2024.

Mean topic coherence · NPMI
+21%Frozen ST vs LDA
LDA0.2875
Frozen ST0.3469
OOS Sharpe · controlled comparison
0.31Frozen ST, single horizon
LDA−0.08
Frozen ST0.31
OOS Sharpe · exploratory extension
1.03Spherical ST + multi-horizon

Spherical clustering, sparse-soft topic terms, and 20/60/120-day exposure horizons form the strongest observed specification.

Topic atlas

Inspect all 60 narratives.

Select a text layer to compare its topic-level coherence distribution and sample term lists. Each dot is one learned topic.

Mean NPMI0.2875Word-count baseline

Interpretation matters. The central LDA versus Frozen ST comparison changes both the representation and how much article text each model sees. Full-sample standardization and topic construction also call for a stricter point-in-time replication. We report observed performance, not established outperformance.

Interactive demo

Watch a topic model
sort the words.

A small, seeded LDA sampler learns from 12 synthetic documents. Run a few sweeps and watch the topic assignments settle. Everything runs in your browser.

Synthetic LDA labK = 3 · α = 0.25 · β = 0.10 · seed = 230
Example articleInitialization

Words are colored by their current topic assignment. The example is synthetic and does not contain empirical paper weights, financial estimates, or remote requests.

Full paper

Methods, robustness,
and limitations.

The 20-page working paper documents the corpus, topic construction, narrative shocks, Sparse IPCA design, out-of-sample evaluation, and topic-count ablations.

Open PDF · 20 pages ↗
Working paper · 2026 From Word Counts
to Context
Topic Models for Asset Pricing UC Berkeley

Citation

Build on the work.

@article{foley2026wordcounts,
  title   = {From Word Counts to Context: Topic Models for Asset Pricing},
  author  = {Foley, Kevin and Hartadi, Jonathan and Prakash, Shivesh and Vatsal, Swapnil},
  year    = {2026},
  note    = {Working paper, University of California, Berkeley}
}