AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

0.55 points of volatility from workforce disruption

Zhang separates from the published labor-shortage measure on beta, while the within-firm annual result carries the weight.

2026-09-08 · 8 min read · US stocks

Reviewing: Disclosed Human-Capital Disruption and Firm-Specific Risk · Ang Zhang · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Point-in-Time Human-Capital Disruption Idiosyncratic-Risk Overlay

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (2)

This is not a replication of the paper

  • The paper's exact score depends on author-defined excerpt coding criteria, a hand-classified set of 50 excerpts, and a trained contextual language-model classifier that are not provided in the available data. A backtest can implement a reproducible approximation using transcript text/embeddings and transparent in-code labeling or NLP rules, but it tests that substitute measure rather than the paper's exact classifier and reported coefficients.

The figures below measure what we could run, not the paper's own method, so they are not evidence for or against its claim.

Our own audit found this run does not follow the paper faithfully (7)

  • Two-of-three DeBERTa-run classifier-agreement robustness variant: The executable specification reports the baseline high-precision and inclusive-F1 variants but omits the two-of-three-run variant. (invalidates: The paper's majority-rule count of 73,007 accepted excerpts, its 1.614-times-baseline count, its associated firm-year rank-stability statistics, and the majority-rule coefficients in Appendix Table 15 cannot be replicated by this specification.)
  • The executable signal is underdefined: the paper and method block require a 3,110-label RBF candidate screen at 0.65, three-sentence context, candidate-only counting, and call-level HCDisruption and HCExposure formulas per 10,000 words, but the specification provides only the 0.9208/0.6492 acceptance thresholds and a fiscal-year formula, leaving p_hat_s, HCDisruption_c, and HCExposure_c unconstructible.
  • The portfolio relies on forecast_idiosyncratic_volatility, but forecast_model specifies only a training cutoff and never implements paper equation (5)—including the logged +2-to-+43 residual-volatility target, standardized call-level disruption, matched pre-risk, coefficients, firm effects, and outcome-quarter effects—so the inverse-forecast-volatility weights cannot be computed from the specification.
  • The paper defines idiosyncratic risk using the published daily Fama–French market, SMB, and HML factors, whereas the specification replaces them with an invented 2,000-stock replica using 0.5 size and 0.3/0.7 book-to-market breakpoints and inverse price-to-book; this changes the defining residuals and detaches the paper's 0.49% 42-day, 0.005454 annual, and related risk coefficients from the implemented forecasts.

3 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 50.1%Jul 2024
Sharpe
0.56
Total Return
50.1%
Max Drawdown
-40.8%
CAGR
9.5%
Volatility
19.9%
Beta vs SPY
0.89
Trades
24,738

A one-standard-deviation rise in Zhang's earnings-call human-capital disruption score adds 0.55 percentage points of annualized idiosyncratic volatility within firm. The estimate has a t of 3.43 across 12,114 firm-years and 2,315 firms. Set beside the sample mean idiosyncratic volatility of 31.8%, the increase amounts to about 1.7% of the level. This belongs in a volatility model. Zhang's own twelve-month-ahead return coefficient is -0.0007 (t=-0.13).

We could not reproduce the score. Public materials omit the coding criteria, the 50 hand-classified excerpts that establish the economic boundary, the 729 training and 400 evaluation Claude Opus labels, and the fine-tuned DeBERTa ensemble. Nothing we ran can score a transcript as Zhang does. The backtest near the end uses a substitute measure and inverse residual-volatility sizing. It cannot test the paper's classifier or its reported coefficients.

Zhang's measure

The economic argument starts with human capital as a production input. A firm's execution becomes less certain when it struggles to hire the required skills, pays more to retain staff, absorbs unusual attrition, or works through restructuring. The expected effect lies in the firm-specific return distribution while market exposure stays unchanged. Zhang designs the tests around that split.

Human-capital disruption means a material disturbance specific to the reporting firm's workforce. The definition covers availability, skills, attrition and retention, wage and labor cost, safety and continuity, workforce restructuring, and consequential leadership transitions. Zhang excludes routine headcount disclosure, generic culture discussion, labor problems at another firm, and unsupported analyst questions.

Measurement has two stages. A workforce vocabulary identifies 1,835,450 transcript sentences. An embedding-plus-SVM screen set at 0.65 reduces them to 173,273 candidates from 36,159 calls, with precision 0.922 and recall 0.769. A three-model DeBERTa ensemble reviews each candidate alongside one sentence of context on either side. Training uses 729 Claude Opus labels. The acceptance threshold of 0.9208 was selected to deliver 90% precision on a separate 400-excerpt sample, where precision is 0.901 and recall is 0.585. Accepted excerpts are counted per 10,000 transcript words and then aggregated into Compustat fiscal years.

The source sample has 45,725 calls linked to CRSP and Compustat, covering October 2005 through May 2025 and drawn from two public transcript archives. Its annual panel spans FY2006 to FY2024, with 13,201 firm-years and 2,870 firms. The average score is 1.173 excerpts per 10,000 words, and 61.8% of firm-years have a positive score.

The annual within-firm estimates carry the most weight. Idiosyncratic volatility rises by 0.0055 (t=3.43), downside deviation by 0.0058 (t=3.92), and worst monthly return falls by -0.0046 (t=-3.75). Market beta changes by -0.0014 (t=-0.37).

At the call level, one standard deviation predicts roughly 0.50% higher log idiosyncratic volatility over trading days +2 to +43 after conditioning on matched pre-call risk. The t is 2.05 in the sample requiring that window to close before the next call. A score purged of every leadership and succession passage also predicts whether the incumbent CEO leaves before the next call. The increase is 0.391 percentage points from a 2.54% base rate (t=2.53), based on 715 exits across 27,691 intervals.

The return test lands at -0.0007, t=-0.13.

Next-year real outcomes are null as well. Employment growth has t=0.65, sales growth t=0.45, and ROA t=-1.14. Zhang gives the same reading in Section 6.6: "The current annual specifications therefore do not detect a common directional change in average real outcomes."

Separate from labor shortages?

Harford, He and Qiu, hereafter HHQ, previously published a FinBERT measure that counts labor-shortage sentences. The question is whether Zhang's wider workforce construct contributes anything beyond it. The distinction appears in the type of risk each measure tracks. Disruption loads on idiosyncratic volatility, while HHQ loads on beta.

The common annual sample contains 7,019 firm-years from 2006 to 2021. With the released HHQ measure and transcript-wide Loughran-McDonald negative and uncertainty frequencies included in the regression, disruption enters at 0.0065 (t=2.97). Removing every explicit shortage passage lifts it to 0.0068 (t=3.28). HHQ enters the same model at -0.0047 (t=-2.16). On its own, HHQ is -0.0014 with a t of -0.74, indistinguishable from zero.

Switch the outcome to market beta, using the same 7,019 observations, and the pattern reverses. Disruption is 0.0062 with a t of 1.30. HHQ is 0.0124 with a t of 2.30. Zhang treats the joint signs carefully, describing them as a decomposition of correlated text measures. His restrained conclusion is that the measures differ in their empirical relationships with systematic and firm-specific risk.

Their overlap has the same imbalance. Among 25,307 one-call firm-quarters, 62.2% of HHQ-positive observations are disruption-positive. Only 33.8% of disruption-positive observations are HHQ-positive, and the continuous correlation is 0.476.

Seven alternative classification rules probe the measurement choice. The grid changes the confidence threshold, requires agreement among independently trained classifiers, and removes shortage or leadership language. Accepted excerpt counts range from 30,294 to 78,003 around a baseline of 45,232, spanning -37% to +72%. Yet the idiosyncratic-volatility coefficient remains between 0.00503 and 0.00654, with a minimum absolute t of 3.16. Downside deviation ranges from 0.00561 to 0.00624, with minimum |t| of 3.73. Worst month runs from -0.00490 to -0.00439, with minimum |t| of 3.46. Firm-year rank correlations against the baseline range from 0.889 to 0.989.

Individual hard cases produce much weaker agreement. Across 120 boundary excerpts, Zhang and Opus agree 65.0% of the time. Cohen's kappa is 0.300, with an interval of 0.134 to 0.464. The grid matters because excerpt-level labeling noise appears to wash out before the firm-year ranking.

Timing weakens the trading case

The abstract acknowledges the timing problem and answers it through persistence. Risk is already elevated before the call, while the score predicts continuation of that firm-specific state over the next 42 trading days. Conditioning supplies the basis for the claim. Zhang writes that disruption predicts higher idiosyncratic volatility in each of the first two nonoverlapping 21-trading-day blocks after conditioning on the nearest pre-call risk realization and information available before the call.

The pre-call path fits that account. Combining the eight pre-call blocks into far, middle and near periods yields idiosyncratic-volatility coefficients of 0.63% (t=2.25), 0.69% (t=2.49), and 0.85% (t=3.36) as the call approaches. No discontinuity appears at the call date. Post-minus-pre changes have t-statistics of 0.59 for idiosyncratic volatility, -1.03 for downside deviation, and -0.36 for tail-loss magnitude.

Generic tone causes the short-horizon result to fade. Adding Loughran-McDonald negative and uncertainty frequencies at the call level lowers the 42-day coefficient from 0.00491 (t=2.09) to 0.00332 (t=1.59). Negative tone alone carries 0.02241 (t=7.36) in the annual model. Zhang reports both findings. In the 23,673 one-call HHQ quarters, he recovers a call-level estimate of 0.00653 (t=2.17). He also says plainly that the annual comparison supplies stronger evidence of incremental content. I share that ranking. The half of the paper that sounds tradable is the half I would decline to fund.

His own cross-sectional result imposes the harder limit. Section 6.3 states: "A cross-sectional specification with Fama-French 48 industry and year fixed effects, rather than firm fixed effects, yields an idiosyncratic-volatility coefficient close to zero. The annual result is concentrated in changes within a firm over time rather than in a stable cross-sectional ranking of firms." Zhang deserves credit for putting this in the body. It leaves the title and abstract with an unanswered question: why foreground a firm-specific risk measure when its only working form compares each firm with its own history?

A sort across 500 names should not be expected to produce a volatility spread from this score. The workable comparison is each firm against itself. The tests draw on a within-firm standard deviation of 1.343, versus 1.971 between firms. Subperiod estimates are 0.0094 (t=2.46) in FY2006 to FY2014 and 0.0033 (t=1.95) in FY2015 to FY2024. Two-way clustering reduces the headline t from 3.43 to 2.86.

Our substitute run

The public materials do not make the score reproducible. They leave out the coding criteria, the 50 hand-classified excerpts defining the economic boundary, the 729 training and 400 evaluation Opus labels, and the fine-tuned DeBERTa ensemble. We therefore could not score a single transcript in Zhang's manner.

Our run isolates the part of the design that can stand without text. It is a long-only book sized inversely to each stock's annualized Fama-French three-factor residual volatility. The estimation window covers the 42 trading days ending two days before rebalance. We used the 500 largest non-ADR US names and rebalanced monthly at the close. Costs were 10 bps one way plus four tenths of a cent a share. The period runs from 2020-01-01 to 2024-07-01. We omitted the disruption overlay that drops the top decile of forecast risk.

Over 2020-01-01 to 2024-07-01, our figures are total return 50.11%, Sharpe 0.56, Sortino 0.68, Calmar 0.23, volatility 19.89%, and maximum drawdown -40.76%. Read the volatility of 19.89% and the -40.76% drawdown first, since inverse-vol sizing makes a risk claim. The Sharpe of 0.56 comes from a window beginning with the March 2020 crash, monthly rebalancing, and one automated pass. It speaks only to our implementation.

Our 50.11% total return and 0.56 Sharpe have no comparable figure in Zhang's paper. The only return test I see there is the insignificant -0.0007 (t=-0.13) twelve-month-ahead coefficient, and I do not see any reported strategy performance. Because our run carries no text signal, it cannot test his claim or serve as a verdict on the authors' work.

The paper earns a place as a volatility-model term. Score each call, standardize the firm against its own recent history, and widen forecast risk by 0.55 percentage points of annualized idiosyncratic volatility for each one standard deviation increase in the score, lasting about one reporting interval. The 1.7% comparison uses the annual panel mean of 31.8%, rather than an individual stock's current level.

The missing test would change my view: a within-firm volatility forecast containing the score that sizes positions better than one excluding it. I do not see that test in the paper. Until somebody runs it, 0.55 percentage points against a 31.8% mean remains a small number on which to build a process.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

1. At each annual universe refresh, select the 500 largest eligible non-ADR US stocks
   using the point-in-time capitalization screen.
2. Construct daily MKT_RF, SMB, and HML replica factors from a separate 2,000-stock
   point-in-time universe; use DFF as the daily risk-free-rate source.
3. For each eligible stock, estimate annualized FF3 residual volatility from the
   42-trading-day lookback ending two trading days before the reference call/rebalance.
   Require complete stock and factor observations; do not impute missing inputs.
4. At executable close rebalances, assign raw_weight_i = 1 / residual_volatility_i.
   The implementation defaults to monthly rebalancing where no explicit frequency is set.
5. Cap each position at 10% of portfolio value, normalize positive weights to 100% gross
   exposure, remain long-only, and enforce the 4.0 maximum-leverage guard.
6. Submit trades at the close. Skip an order if its close price is missing; never
   synthesize an execution price.
7. Apply symmetric trading costs and report one-way traded notional relative to prior
   portfolio value.

Not executed: NLP classification, forecast idiosyncratic volatility, and exclusion of
stocks above the 90th percentile of forecast risk.