AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

The placebo that erases its own 1.35 Sharpe

What remains is per-tag drift worth 0.10% to 0.35% an event on the continuation side.

2026-09-08 · 9 min read · US-listed individual equities

Reviewing: Buy the Rumor, Sell the News: When Is News Priced In? · Alireza Kargarzadeh, Nariman Khaledian, Navid Parvini et al. · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Novel-Story Fundamental Drift and Reaction-Conditioned Narrative Overreaction

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (2)

This is not a replication of the paper (3)

  • The paper's NewsWitch corpus, vendor-produced five-level sentiment, article summaries, source filtering, and proprietary 17-tag labels are unavailable. Event direction, fundamental-versus-story classification, rumor flags, and story clustering must be reconstructed from the available news text/embeddings using an in-code classifier or rules; the resulting signal is not a reproduction of the paper's dataset or labels.
  • The reported 2023-2026 results rely on a much larger commercial crawl than the available news tables. A backtest would test the mechanism on available FMP/news/press-release coverage from 2020 onward, with potentially different article coverage, timestamps, duplication, and tagging quality.
  • The paper estimates event-study abnormal-return drifts rather than reporting implementable portfolio returns after trading costs, borrow constraints, and signal availability delays. These need to be explicitly modeled in a trading backtest.

The figures below measure what we could run, not the paper's own method, so they are not evidence for or against its claim.

Our own audit found this run does not follow the paper faithfully (7)

  • News corpus and sentiment provenance: The implementation uses fmp_stock_news and press_releases rather than the retained NewsWitch corpus, and therefore cannot assume the same vendor sentiment field or label distribution. (invalidates: The paper's exact article and event counts, sentiment shares, tag frequencies, source-group results, NEW shares, adjusted tag drifts, p-values, q-values, and economic-significance portfolio results.)
  • Classifier training chronology: The implementation freezes a timestamp-safe classifier before the backtest rather than applying the paper's retrospectively trained deployment classifier. (invalidates: The paper's reported corpus-wide teacher fidelity, teacher-only robustness counts, tag frequencies, and exact tag-level return estimates.)
  • Trading entry and exit timing: The strategy enters at the next session close and tests 5-, 10-, 15-, and 20-session exits instead of exclusively entering at day +5 and exiting at day +20. (invalidates: The paper's 15.8% and 15.9% annualized returns, the 1.35 gross Sharpe ratios, all reported net Sharpe ratios, and the legal-and-regulatory portfolio result in Table 6.)
  • Reaction-conditioned soft-news strategy: The implementation trades only NEW unquantified soft narratives whose publication-day abnormal return is aligned with sentiment; the paper reports event-study reversals and an unconditional launch-and-partnership fade rather than this conditioned portfolio. (invalidates: Direct comparison with the paper's product-launch, leadership, competition, and macro adjusted drifts and with the 15.8% small-cap launch-and-partnership fade.)

3 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total -18.9%Jul 2024
Sharpe
-0.23
Total Return
-18.9%
Max Drawdown
-51.9%
CAGR
-4.5%
Volatility
18.1%
Beta vs SPY
0.11
Trades
36,457

What the paper reports for its own strategy

  • Fade small-cap launch/partnership news (short after positive, buy after negative; enter close of day +5, exit close of day +20, equal-weighted, beta-hedged, 52,149 events, 2023-2026): 15.8% annualized abnormal return, gross Sharpe 1.35, Sharpe 1.20 at 5bps per side, 1.06 at 10bps, 0.77 at 20bps
  • Short any covered small cap (sentiment-ignoring benchmark, 260,472 events, same windows, 2023-2026): 15.9% annualized, gross Sharpe 1.35, 1.21 at 5bps, 1.07 at 10bps, 0.78 at 20bps
  • Follow legal and regulatory news direction (46,057 events, 2023-2026): 4.1% annualized, gross Sharpe 0.66, 0.39 at 5bps, 0.12 at 10bps, -0.43 at 20bps
  • Residual tag-map alpha cited as economically thin: earnings continuation +0.22% per event over 15 days

The paper's most useful result is 15.9%. From 2023 to 2026, shorting every small cap that appeared anywhere in the news, regardless of sentiment, produced a 15.9% annualized abnormal return. Positions ran from the close of day +5 through the close of day +20, across 260,472 events. The headline strategy uses exactly the same window to fade positive-sentiment launch and partnership coverage in small caps. It earns 15.8%, with the same gross Sharpe of 1.35, across 52,149 events. Kargarzadeh and co-authors place the rows beside each other and reach the obvious conclusion: background drift supplies the fading strategy's entire edge. News direction adds nothing.

The evaluation rule is the result.

$29 to tag 4.57 million articles

NewsWitch is the corpus, a commercial crawl operated by one of the author affiliations. It covers roughly the 3,000 most-covered US-listed stocks from 2023 to 2026. More than 70 million raw articles, drawn from over half a million sources, become 4.57 million retained articles. Each has a vendor LLM's five-level sentiment score and summary.

The authors then impose an event structure. A gpt-5-mini teacher labels 29,472 random articles using 17 event tags and five binary attributes: scheduled, forward looking, primary source, quantified, rumor. The cost is about $0.0002 an article. An 82M-parameter distilroberta student learns from those labels and tags 600,000 new articles. The 100,000 lowest-confidence cases return to the teacher, after which the student is retrained on 129,463 merged labels. On a held-out, double-labeled set of 10,000 articles, it agrees with the teacher on 87.5% of primary tags. Agreement reaches 94.0% at confidence 0.8 or above. Total teacher spend: $29. A laptop processes 125 to 160 articles per second.

Next comes story clustering. Within each stock, day and primary tag, title-and-summary embeddings are grouped at a 0.80 cosine-similarity threshold. The permitted story length varies by tag, from 90 days for M&A and legal sagas to two days for commentary. Across 22,596 stories, the mean M&A story lasts 12.9 trading days. The corpus reveals its own duplication problem here: 55% of all articles continue a story already under way.

Each event is a (stock, trading day, tag) aggregate. Articles published post-16:00 New York roll into the next session. The procedure yields 1,681,657 price-scored events. Of those, 1,317,252 are signed, with 73% positive, while 364,405 are neutral.

For returns, the paper subtracts a rolling 252-day OLS beta against SPY multiplied by the index return. The regression requires a minimum 126 observations, and missing betas are set to 1. Four windows are reported: days -5..0, day 0, +1..+5 and +6..+20. Returns receive the news sign, making a move in the reported direction positive. P-values use a 5,000-draw bootstrap that resamples trading dates.

The adjusted estimator deserves the attention. Before assigning the sign, the authors subtract the mean abnormal return for neutral events in the same dollar-volume tertile. Neutral coverage therefore acts as the placebo for appearing in the news at all.

The adjustment earns its place immediately. Following neutral coverage, stocks lag the beta benchmark by 0.92% for small caps, 0.58% for mid caps and 0.34% for large caps over the next month. Quiet stock-days, with no coverage, drift 0.74%, 0.88% and 0.59% in those buckets. Combined with a 73% positive sentiment share, the same drift generates three apparent anomalies.

Across days +6..+20, raw drift is -0.62% after good news and +0.63% after bad news. Both become 0.00% after adjustment, at p about 0.9. The raw small-cap reversal looks 2.4 times stronger than the large-cap version, -0.41% against -0.17% over a month. Adjustment removes the gradient. Legal and regulatory news supplies the only raw continuation winner, at +0.29%, p<0.001. It is also the only tag with 83% negative events, which accounts for the result. After adjustment, the estimate is -0.08%, p=0.22.

We previously covered a paper whose headline gain came from its chosen baseline (the triadic stress index). These authors applied the more revealing baseline themselves, then published the row that kills their trade.

How much does the 2.8x explain?

Across all signed events, the cumulative move in the news direction reaches +0.58% by the publication-day close and declines to +0.20% twenty days later. The ratio is 2.8. Read as "news is priced in", it carries less information than first appears. The paper acknowledges that the pooled ratio exceeds the ratio for every individual tag because the pooled after-window absorbs the background drift.

Within tags, the ratio says more about completion than size. It is 1.06 for earnings, 1.20 for analyst actions and 1.08 for guidance. By the end of publication day, the move was finished.

The pre-window is contaminated too, and the authors measure part of the problem. Limiting earnings to first reports reduces the days -5..0 result from +1.41% to +1.02%. Follow-up coverage therefore inflates apparent anticipation by about a third. Source type gives a cleaner separation. Press wires, which distribute company and regulator releases, carry -0.06% pre-publication drift. Retail analysis sites show +1.07%, portals +0.92% and mainstream media +0.86%. Commentary arrives after prices have started moving. The release coincides with the move.

The rumor result is worth keeping. Among 17,510 rumor-flagged events followed by a same-tag confirmation within 60 trading days, with a median gap of six days, the rumor day earns +0.36%. The period before confirmation returns -0.09%, and the confirmation day adds +0.01%. For the 3,335 M&A rumors, confirmation day is -0.10%, followed by -0.37% over days +6..+20.

Seventeen tags leave three survivors

On the continuation side, adjusted drift over days +6..+20 is +0.35% for capital returns (p<0.001), +0.22% for earnings (p<0.001), +0.13% for guidance (p=0.046) and +0.10% for analyst actions (p=0.012). Reversal estimates are -0.34% for macro read-throughs (p<0.001), -0.18% for product launch (p=0.022), -0.18% for leadership (p=0.016) and -0.17% for competition (p=0.026).

Because the paper tests 17 tags together, the authors apply Benjamini-Hochberg. Macro, capital returns and earnings survive at 5% FDR. Analyst, leadership and launch have q between 0.05 and 0.07. The paper describes guidance as suggestive rather than decisive.

The internal cuts weaken one of the three survivors, while another depends on the baseline design. Earnings drops from +0.22% to +0.11% when restricted to first reports, and again measures +0.11% on teacher-labeled events. Capital returns is steadier: +0.33% on first reports, +0.47% on single-tag days and +0.34% on teacher labels, compared with +0.35% across all events.

Yet the main baseline pools event tags inside each size bucket, assuming tag-independent coverage drift. Matching the baseline by tag reduces the capital-returns estimate and sends the launch estimate across zero. The authors describe magnitudes as more model-dependent than signs. That description gives the launch cell too much latitude, since crossing zero changes the sign. Decay is visible within the sample as well. Raw pooled reversal halves from -0.49% in 2024 to -0.18% in 2026, while adjusted pooled drift is zero from 2025 onward.

Costs leave little room. Earnings continuation earns +0.22% an event over 15 days on a beta-hedged leg, a result the paper itself calls thin beside plausible trading costs. Capital returns, at +0.35%, is the largest surviving estimate. Following the direction of legal news, the position suggested by the raw table, has a gross Sharpe of 0.66. It falls to 0.12 at 10bps a side and -0.43 at 20bps. Both headline books are mainly short, and the authors state that their model excludes borrow fees and locate availability.

The second moment survives the cuts. Neutral guidance events occur on days with 1.29 times the stock's normal absolute move, followed by 1.23 times over the next week. Days carrying ten or more articles move 1.36 times normal, versus 1.05 for single-article days. Post-event realized volatility is about 0.86 of the EWMA forecast. Publicity still has width after direction is exhausted. The authors reach the same conclusion: news flow offers presence and width, while direction is spent by the closing bell.

Our book measures something else

We could not reproduce the measurement. The NewsWitch corpus, its five-level vendor sentiment, its summaries and the 17-tag labels are unavailable to us. We rebuilt tags, sentiment sign, rumor flags and story clusters from FMP stock news and press releases, using our own classifier on the text we could obtain. Coverage, timestamps, duplication and tagging quality may all differ, and the size of those differences cannot be measured. The paper estimates event-study drifts rather than an implementable portfolio, requiring us to add entry delays, commissions and borrow assumptions.

Our strategy also differs from the book earning the paper's 15.8%. We trade first reports only. We follow sentiment for quantified earnings, guidance, capital-return and analyst events, while fading soft narrative events such as launch, macro, leadership and competition when publication-day abnormal return agrees with the sentiment sign. There is no small-cap restriction. We enter at the next session's close instead of the paper's day +5 and hold positions for 20 sessions. Absolute weights are equal, paired with a rolling-beta SPY hedge. Commissions are $0.004 a share with a $1 minimum, with no modeled slippage. The universe is the annual top 3,000 non-ADR US names, of which 2,989 resolved, from January 2020 through July 2024.

It lost money.

Our figures are -18.9% cumulative return, -4.55% CAGR, Sharpe -0.23, maximum drawdown -51.9%, volatility 18.1%, realized beta 0.11 and 36,457 trades. The paper reports +15.8% annualized at a gross Sharpe of 1.35 for its own overlay from 2023 to 2026. The corpora, periods and portfolios differ, so the opposite signs do not constitute a failed replication of the authors' result.

Our design choices explain much of the gap. Dropping the small-cap short tilt removed the mechanism that the paper shows behind its 15.8%. We introduced a long quantified-continuation sleeve that the authors never trade and themselves describe as thin. More than half of our window covers 2020 to 2022, when covered small caps rose sharply, reversing the sign of the drift harvested by the paper. Only about 18 months of our sample overlaps theirs.

Entry at day +1 also includes the days +1..+5 segment that the authors intentionally omit. Corpus construction and sentiment provenance differ. With per-event edges of 0.1% to 0.3%, enough classification noise can reverse the result. Our book made money on 47.25% of trades and recorded a profit factor of 0.97. Its gross edge was marginally negative before leverage.

The realized beta of 0.11 indicates that the hedge did its job, leaving the loss unexplained by a market call. Volatility of 18.1% and a -51.9% drawdown are more troubling for a market-neutral overlay. We used four times gross leverage and charged per-share commissions subject to a $1 minimum. Both likely contributed to the drawdown in a portfolio whose gross edge was already marginally negative. Even so, those differences do not comfortably explain the full drawdown, and the remaining gap cannot be closed from the evidence available to us.

The baseline is the paper's contribution

The evaluation rule matters more than the resulting tag map. Claims of news alpha should face a coverage-presence baseline like the one measured here. Without it, a strategy can rediscover background drift and label it signal. A small-cap sentiment overlay can be checked quickly: remove sentiment, retain the coverage filter and measure how much Sharpe remains. An answer of 1.35 against 1.35 means the trade has captured size drift relative to a single-beta benchmark.

I would trade the residual map, if at all, with low conviction. Three of 17 tags survive a 5% FDR. The largest earns +0.35% an event over 15 days before borrow, while adjusted pooled drift is already zero from 2025 onward. One result would change my view: the same table estimated with a multi-factor benchmark during a period outside 2023 to 2026. Capital returns and earnings would need to preserve their signs and magnitudes on live coverage rather than backfilled coverage.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each trading day t:
  1. Resolve the annual top-3,000 non-ADR stock universe using only that year's screening data.
  2. Assign each article to t using America/New_York time; roll publications after 16:00 to the next session.
  3. Assign a primary event tag and an event-sentiment sign from timestamp-safe article labels.
  4. Within each stock, day, and tag, cluster title-plus-summary embeddings at cosine similarity &gt;= 0.80.
  5. Match clusters to same-stock, same-tag open stories using the initial frozen centroid and tag-specific lifetime.
     Mark only the first cluster of a newly opened story as NEW.
  6. For a selected strategy family:
     - Quantified continuation: for NEW quantified Earnings results, Guidance and outlook,
       Capital returns, or Analyst action events, trade in the sentiment direction.
     - Reaction-conditioned reversal: for NEW Product launch, Macro through stock,
       Leadership and governance, or Competition events, require publication-day abnormal
       return AR_0 to share the sentiment sign, then trade against that sign.
     - Unconditional control: fade Product launch or Partnership and customer sentiment
       without requiring publication-day reaction alignment.
  7. Enter at the next session's observed adjusted close; skip the trade if an execution price is missing.
  8. Rebalance open stock positions to equal absolute weights. Hedge aggregate stock beta with SPY,
     where each stock beta is the 252-session OLS slope through the current date, requires 126 observations,
     and otherwise falls back to 1.
  9. Exit at the configured holding horizon using the observed adjusted close; the supplied implementation
     defaults to 20 sessions. Apply stock, hedge, and rebalance costs.

The neutral-event baseline and 5,000 date-cluster bootstrap with Benjamini-Hochberg correction are reporting tools only and do not form signals.