AnyLearn
All lessons
AIadvanced

Reading the Evidence Carefully

The most prominent study in this area found a language model predicting stock reactions from headlines, and it is usually reported as proving something it explicitly does not claim. This lesson reads it precisely, follows its qualifiers to their consequences, and shows why a real statistical result and a tradable strategy are different things.

Updated · AI-authored, review-gated · how lessons are made

Not signed in: your progress and quiz score won't be saved.
Progress1 / 8

The study, stated accurately

Alejandro Lopez-Lira and Yuehua Tang published "Can ChatGPT Forecast Stock Price Movements? Return Predictability and Large Language Models" in 2023, subsequently in the Journal of Financial Economics.

The method is simple, which is part of why it is credible. Give the model a news headline about a company. Ask whether it is good, bad or irrelevant for that firm's stock price. Convert the answers into a numerical score and test the relationship with subsequent returns.

The reported findings, precisely:

Using post-knowledge-cutoff headlines, GPT-4 captures initial market responses, achieving approximately 90% portfolio-day hit rates for the non-tradable initial reaction.

The scores significantly predict the subsequent drift, especially for small stocks and negative news.

Forecasting ability increases with model size. GPT-1, GPT-2 and BERT could not accurately forecast returns, which the authors read as return predictability being an emerging capacity of more complex models.

Every qualifier in those three statements is doing work, and the popular version drops all of them.

Full lesson text

All 8 steps on one page, for reading, reference, and search.

Show

1. The study, stated accurately

Alejandro Lopez-Lira and Yuehua Tang published "Can ChatGPT Forecast Stock Price Movements? Return Predictability and Large Language Models" in 2023, subsequently in the Journal of Financial Economics.

The method is simple, which is part of why it is credible. Give the model a news headline about a company. Ask whether it is good, bad or irrelevant for that firm's stock price. Convert the answers into a numerical score and test the relationship with subsequent returns.

The reported findings, precisely:

Using post-knowledge-cutoff headlines, GPT-4 captures initial market responses, achieving approximately 90% portfolio-day hit rates for the non-tradable initial reaction.

The scores significantly predict the subsequent drift, especially for small stocks and negative news.

Forecasting ability increases with model size. GPT-1, GPT-2 and BERT could not accurately forecast returns, which the authors read as return predictability being an emerging capacity of more complex models.

Every qualifier in those three statements is doing work, and the popular version drops all of them.

2. What it did right

Before the qualifiers, the study's design deserves credit, because it addresses the previous lesson's central problem directly.

It used post-knowledge-cutoff headlines. That is the strongest available defence against contamination, and it is why the result is worth discussing at all. Most claims in this area fail at exactly this point.

It compared across model generations. Finding that GPT-1, GPT-2 and BERT cannot do this while larger models can is a meaningful control. If the effect were an artefact of the evaluation setup, it would not scale with capability in that way.

It reported where the effect concentrates rather than presenting an aggregate. Saying the drift prediction is strongest in small stocks and negative news is more informative than a single number, and it is the kind of detail that lets a reader assess tradability.

It distinguished the non-tradable initial reaction from the subsequent drift. That distinction is the most important thing in the paper and the first thing dropped in summary.

So this is careful work, and the criticism that follows is not of the study. It is of the gap between what it establishes and what it is taken to establish.

3. Non-tradable means non-tradable

The 90% figure is the number that travels, and it is attached to the non-tradable initial reaction. That phrase is not a hedge; it is a description of a specific quantity.

The initial reaction is the price move that occurs when news arrives. It happens in the seconds and minutes after publication, and by the time a headline is in a feed, retrieved, sent to a model, scored, and returned, it has already happened.

So predicting it at 90% is a statement about information content: the headline contains enough for a reader to anticipate the direction, and the model extracts it. That is a genuine and interesting finding about what language models understand.

It is not a statement about a trade, because there is nothing to trade. The move you predicted is in the past by the time the prediction exists.

The previous cursus makes this concrete. Market Design: Ticks, Venues, and the Speed Race establishes that reacting first to public information is a contest measured in microseconds, and the first lesson here put model inference five or six orders of magnitude away from that.

The tradable claim in the paper is therefore the other one: the subsequent drift. That is a smaller effect, and it is the one whose qualifiers matter most.

4. Where the qualifiers lead

Follow the tradable branch and each qualifier lands somewhere expensive.

Small stocks. The lesson Implementation Shortfall: What an Order Really Costs establishes that order size relative to average daily volume is the single best predictor of execution cost. A signal whose predictive content concentrates where books are thinnest is a signal whose costs are highest exactly where its edge is largest.

Negative news. Acting on it means selling, and for a position you do not hold that means selling short: locating a borrow, paying the fee, and accepting that in exactly the names where the signal is strongest the borrow is likely to be expensive or unavailable.

Neither observation contradicts the paper, and neither is a criticism of it. The authors reported these concentrations, which is what makes the assessment possible.

The point is that statistical significance and tradability are different properties, and the distance between them is measured in the currency of the execution cursus. A real relationship whose returns are consumed by spread, impact and borrow is a real relationship and not a strategy.

flowchart TD
A["GPT-4 scores predict returns"] --> B["90% hit rate"]
A --> C["Predicts subsequent drift"]
B --> D["On the non-tradable initial reaction"]
D --> E["Already happened when you can act"]
C --> F["Concentrated in small stocks"]
C --> G["Concentrated in negative news"]
F --> H["Thin books, high impact per unit"]
G --> I["Short selling: borrow cost and availability"]

5. The capacity ceiling

The concentration has a second consequence, separate from cost per trade and arguably more limiting.

Costs, Capacity, and a Protocol You Can Trust establishes that market impact grows with order size in a way closer to a square root than a straight line, so every strategy has a capital ceiling beyond which its own trading consumes its edge. The ceiling is set by the liquidity of the instruments it trades.

A signal whose content sits in small capitalisation names therefore has a low ceiling by construction. The names are thin, so the position that can be built without moving the price is small, and the absolute profit available is bounded regardless of how good the signal is.

That has a specific commercial consequence worth naming. A strategy with a genuine edge and a ceiling of a few tens of millions is valuable to an individual and irrelevant to a large fund, whose fixed costs exceed what the strategy can produce.

So the same result can be simultaneously true, interesting, and useless to the institution reading it.

This is not a criticism of language models. It is a general property of signals found in under-covered corners, and the reason those corners were under-covered is often that they are too small to be worth the coverage.

6. What the model-size finding suggests

The comparison across generations is the part of the study with the longest shelf life, and it points at a mechanism.

GPT-1, GPT-2 and BERT could not do this. Larger models could, with ability increasing alongside capability. The authors describe financial reasoning as an emerging capacity of complex models.

There is a reading of that which is worth taking seriously and which is less exciting than it sounds. Assessing whether a headline is good or bad for a company requires knowing what companies do, what the words mean commercially, and how the described event affects a business. That is general world knowledge and language understanding, not financial insight.

Small models lack it. Large ones have it. So the finding may be that reading comprehension at a sufficient level is what the task requires, and that scale delivered reading comprehension.

That reading is consistent with everything in the first lesson. The value is in extraction from text at scale, and headline scoring is extraction: the information is in the headline, and the model is retrieving it rather than forecasting anything.

It also predicts something checkable: the effect should be largest where the reading was not being done. Concentration in small, under-covered stocks is exactly that pattern.

7. What publication does to a signal

One further consideration applies to any published finding and applies with unusual force here.

The Same Problem in Academic Finance covers the multiple-testing problem across a profession: hundreds of published factors, tested against essentially the same history, with no individual paper accounting for the collective search. Published anomalies routinely weaken after publication.

A finding about a widely available model has that problem and a sharper one. The barrier to replication is close to zero. Reproducing a factor requires data, infrastructure and expertise; reproducing this requires an API key and the prompt, which is in the paper.

So the first lesson's argument applies directly. The information is public, the transformation is public, and the result is not private information. Whatever edge existed is competed away faster than for a conventional anomaly, because the cost of entry is a rounding error.

The conclusion is not that the study is wrong. It is that a published, easily replicated signal is close to the definition of something you should not expect to be paid for, and this is a strong instance.

Which returns to where the first lesson ended: the durable applications are the ones nobody can replicate from a paper, because they run on documents you have and others do not, or ask questions nobody thought to ask.

8. The reading, in full

ClaimStatus
An LLM extracts price-relevant information from headlinesEstablished, and the post-cutoff design supports it
Ability scales with model capabilityEstablished, with a meaningful control
It predicts the initial reaction at ~90%Established, and that reaction is non-tradable
It predicts subsequent driftEstablished, concentrated in small stocks and negative news
Therefore a profitable strategy existsNot established, and not claimed
Costs, borrow and capacity are accounted forNot addressed, and they land on the concentrations
The edge survives replicationUnlikely, given a public method

The first four rows are a real contribution. The last three are where the popular reading goes wrong, and note that the paper does not make those claims: they are supplied by readers.

The transferable skill is the one this lesson has been exercising, and it generalises past this subject. Read the qualifiers as the result. Non-tradable, small stocks, negative news, post-cutoff: each is a load-bearing constraint, and a summary that omits them describes a different and better finding than the one obtained.

Applied here, the honest summary is that language models demonstrably extract information from financial text, and that converting information extraction into net-of-cost profit is a separate problem which the evidence does not address and which the trading cursus suggests is the harder half.

That is a more useful conclusion than either enthusiasm or dismissal, and it points directly at what to build, which is the last lesson.

Check your understanding

The lesson ends with a 5-question quiz. Take it in the player above to see your score.

  1. The ~90% hit rate in Lopez-Lira and Tang applies to what, exactly?
    • The non-tradable initial reaction, which has already occurred by the time a model can score the headline
    • The subsequent drift over the following days
    • A long-short portfolio net of transaction costs
    • Directional accuracy across all stocks and horizons
  2. Why does the drift concentrating in small stocks limit tradability?
    • Small stocks have less news coverage to score
    • Order size relative to average daily volume drives execution cost, so costs are highest exactly where the edge is largest
    • Small stocks are excluded from most benchmarks
    • Model accuracy is lower on unfamiliar company names
  3. What does the finding that GPT-1, GPT-2 and BERT could not forecast returns provide?
    • Evidence that the task requires financial pretraining
    • Evidence that older models had cleaner knowledge cutoffs
    • A meaningful control: an evaluation artefact would not scale with model capability
    • Proof that scaling laws apply to financial data
  4. Why is a published LLM signal competed away faster than a conventional published anomaly?
    • Because replication needs only an API key and the prompt, which is in the paper
    • Because model providers share results between customers
    • Because language models converge on identical outputs over time
    • Because regulators require disclosure of model-based strategies
  5. What is the accurate summary of what the study establishes?
    • That LLMs can generate profitable trading strategies from public news
    • That LLMs extract price-relevant information from text, while converting that into net-of-cost profit is a separate, unaddressed problem
    • That domain-specific financial models outperform general ones
    • That headline sentiment is unrelated to subsequent returns

Related lessons

AI
intermediate

What This Teaches About Measuring Anything

The exchange is a case study with transferable rules. A conclusion resting on failures needs a failure taxonomy. Every instance must be verified solvable before anyone is scored against it. Output format is a confound whenever answers get long. And when two explanations fit the same data, the productive move is to find the prediction on which they differ, then test it.

8 steps·~12 min
AI
intermediate

The Rebuttal: Three Ways to Score Zero Without Failing

The response disputed none of the data and argued the experiment measured something other than reasoning. Models had to print move lists exceeding their output limits, and said so in the transcripts. Some instances had no solution and were scored as failures anyway. And asking for a program instead of a move list produced high accuracy on instances reported as total collapse.

8 steps·~12 min
AI
intermediate

The Experiment: Puzzles With a Difficulty Dial

Apple researchers built an evaluation designed to fix a real problem with benchmarks: puzzles where difficulty turns up smoothly while the logic stays identical, and every step can be checked. They found accuracy collapsing to zero past a threshold, and not improving when the solution algorithm was handed to the model. This lesson covers the design and why it was a good one.

8 steps·~12 min
AI
intermediate

Ten Domains, and a Profile That Is Not Flat

The framework scores ten cognitive domains at ten percent each. Running it produces something more useful than the headline totals of 27 percent for GPT-4 and around 57 for GPT-5: a jagged profile, where a model is at or near full marks on some domains and at zero on others. This lesson walks the domains, reads both profiles column by column, and shows what the jaggedness explains.

8 steps·~12 min