A driver that misses every inverted CDS curve should not survive a credit-model bake-off. Across the CDX high-yield panel, Baker and Capponi's subordinator reproduces 0 of 170 observed 1y-5y spread inversions in sample and 0 of 212 out of sample. It produces no false positives either. Their negative-drift driver, calibrated with year-one statics frozen, recalls 197 of 212 out of sample. The distinction can be falsified with quoted curves and a recall count over 170 and 212 inverted firm-dates. It also tests the driver class used in prior multi-credit work.
Firm i defaults when the running supremum of a latent, spectrally positive distress process first crosses an independent Exp(1) barrier. Its driver combines minus a distance-to-distress, a drift, Brownian noise and compound Poisson jumps with Exp(eta) sizes. The scalar distance-to-distress is re-fitted each date. All remaining parameters stay fixed per firm over a trailing three-month window. Baker and Capponi explain the division directly: "a single CDS curve carries eight tenors and is therefore worth roughly one effective parameter."
With mu < 0, the firm recovers between shocks. Under the classical net-profit condition mu + lambda_J E[U] < 0, its average forward hazard decays to zero. Setting sigma = 0 and mu >= 0 turns the driver into a subordinator, with accumulated hazard rising only toward a plateau. Lemma 2.8 limits any inversion to delta/T2, the buffer term divided by the longer maturity. That bound supplies the falsification.
Pricing the latent process
Theorem 2.3 provides the tractable step. With phase-type jumps, the Laplace transform of default probability becomes a finite partial-fraction sum over the roots of one polynomial. Its degree is m+2 when sigma>0 and m+1 when sigma=0. Numerical inversion uses an 18-node Talbot contour.
For portfolios, the model adds one common compound Poisson shock to every firm's driver, with arrival rate lambda_c and mean size gamma. The marginals remain phase-type at m=2, allowing the same root system to price the joint law. Simulation then runs through an event-driven Wiener-Hopf scheme at a step rate of 16 per year.
The data consists of daily Markit/S&P composites obtained through WRDS from 2021-05-27 to 2025-02-13. CDX-NAHY contributes 93,458 firm-dates, with a median 98 names per date. CDX-NAIG contributes 119,364, with a median 125. The tranche sample contains 5Y on-the-run upfronts, including 961 quotes per NAHY tranche and 962 per NAIG. Each model is fitted to the ISDA-bootstrapped forward hazard vector.
Across 14 quarterly out-of-sample folds, LBM+Exp, linear Brownian motion with exponential jumps, records the lowest mean-of-median hazard RMSE on all three boards. A board means a constituent universe here: the high-yield Series 36 cohort, the high-yield full basket or the investment-grade full basket. LBM+Exp scores 0.0098, 0.0104 and 0.0018. The subordinator reaches 0.0149, 0.0154 and 0.0052. With dependence frozen, summed absolute tranche error on the high-yield cohort drops from independence at 0.398 to 0.101. Investment grade falls from 0.421 to 0.048.
No bid-ask spreads or hedging P&L.
This remains a marking model. Its case against desk practice is a single joint law across the capital stack in place of one correlation for each attachment point.
Enough detail to replicate much of it
The paper writes out the Talbot contour, including its radius, nodes and weights. It also gives the polynomial and residues explicitly. Tranche upfront identities include the default-count thresholds, {1,22,36,51} for high-yield and {1,7,15,32} for investment grade, while the ISDA bootstrap is fully specified.
Its caching claim can be checked and has practical force. Nodes, roots and residues depend solely on maturity and the statics. Re-marking a firm's state therefore costs one vector of complex exponentials. The authors also sweep the window length. Over a common 2022-05-28 to 2023-06-30 block, the three-month trailing window is lowest or tied-lowest in 10 of 12 model-index pairs. For LBM+Exp on high-yield, the comparison is 0.0142 at 3m against 0.0150 at 12m.
Judgment enters through the optimizer. Coordinate descent minimizes panel loss by alternating a one-dimensional state update with an outer fit of the statics. The statics are warm-started from the prior fold. A bracketed one-dimensional root-find handles the state solve. We did not find printed tolerances or numeric search bounds.
Those bounds determine one benchmark result. In the one-year in-sample panel, Black-Cox on investment grade places its fitted x_0 at the search-range limit on 97% of firm-dates. Its in-sample median error is 0.0155, versus 0.0159 for the curve's own median root-mean-square. The authors identify the issue themselves and report that relaxing the limit does not alter the fit. They attribute it to monotonically rising investment-grade forward hazards, whereas Brownian first-passage hazard is unimodal. A tighter search box would give a replicator another number for that row.
The joint fit introduces a second dial that the paper does not make checkable. Its outer loop searches (gamma, lambda_c) by summed absolute tranche error. A penalty protects single-name fit quality, but the paper does not provide its form or weight. The dependence layer leaves single-name error flat when activated, which is the authors' defence. LBM+Exp records 0.0045 on the cohort with dependence frozen, matching independence, and 0.0052 after full re-marking.
Rite Aid and the inversion count
Rite Aid shows how optimizer tolerances can dominate recall. The constrained Cramér-Lundberg fit lands at mu = -8e-9, exactly at the boundary that forbids inversion. It recalls 0 of the firm's 26 inverted firm-dates. Any strictly negative drift, including -0.05, recovers all 26.
Cramér-Lundberg misses 35 of 212 observations out of sample, with Rite Aid contributing 26 of those 35. Its headline recall of 177 of 212 therefore depends on how near zero the optimizer may stop. Two of seven firms produce the entire shortfall. More broadly, the test covers 170 and 212 firm-dates across five and seven firms. Over 95% are CCC-rated, and investment grade has zero inversions in either year. The falsification applies only to high-yield and sits close to firm-specific behavior. It remains the right test, though its breadth is thinner than the five-figure firm-date panel implies.
Recall also needs the false-positive column. The control panel contains an equal-size sample of CCC firm-dates with quoted curves that are not inverted. LBM+Exp generates 32 false positives. Black-Cox produces 26 while recalling 185 of 212, despite finishing in the bottom half of the RMSE field.
Does the tranche result settle the argument?
The flat-correlation Gaussian copula supplies the sharpest benchmark. Because it accepts bootstrapped market marginals, its single-name error is zero by construction. Under the frozen protocol, it prices at 0.192 / 0.226 / 0.105 across the three boards. LBM+Exp reaches 0.101 / 0.118 / 0.048. Cramér-Lundberg reaches 0.113 / 0.117 / 0.053 and edges LBM+Exp on the full basket. Re-marking the copula's single correlation improves it only to 0.159 / 0.189 / 0.097.
Simulation noise is much smaller than those differences. Across independent seed streams, monthly-mean summed error moves by 0.010 to 0.024 in mean absolute terms, an order of magnitude below the reported separations.
Within the authors' own class, however, that same noise band leaves the three high-yield drivers indistinguishable. The paper says this directly. Under the frozen protocol, the subordinator lies within 0.03 of both alternatives on the cohort, at 0.127 against 0.101 and 0.113. The full-basket figures are 0.121 against 0.117 and 0.118. Rolling marks repeatedly anchor the model to the 5y-and-shorter portion of the curve, where a monotone driver can still approximate rising hazards.
The authors make their case against the mu >= 0 driver through the single-name results, the inversion test and investment-grade tranches. On investment grade, the subordinator with frozen dependence marks 0.447, worse than the independence level of 0.421. The objection holds in that narrower form. High-yield tranches by themselves cannot rank the three drivers.
The authors also acknowledge that the copula comparison does not represent desk practice, where convention fits a base correlation at every attachment point. Their defence is that the skew amounts to four models rather than one joint law, making flat rho the like-for-like comparison. I accept that reasoning. The practical hurdle remains the skew.
Duffie-Gârleanu carries less evidentiary weight than its table position implies. Frozen, it prices at 0.247 / 0.269 / 0.392. Re-marking all six common parameters brings little improvement. Its single-name RMSE ranges from 0.0176 to 0.0207, three to six times the running-supremum class. For investment grade, the fitted common factor reaches the edge of its admissible range, with no mean reversion and maximal volatility. The authors interpret that outcome as misspecification in the six-parameter form. A benchmark that loses on the marginals cannot decide a dependence contest.
What frozen actually means
The abstract leads with the frozen column and describes the 95% result as re-marked, an honest ordering. Frozen dependence still uses statics and dependence fitted on the preceding three months of tranche quotes. During the test quarter, only the per-firm states are re-anchored to each day's CDS curves. This is a genuine prediction over roughly a year of quotes per board and across 10 folds.
Staleness consistently hurts in one direction. Across a three-month fold, high-yield out-of-sample error rises monotonically with fit age by 11% to 26%. LBM+Exp moves from 0.0090 at 0 to 30 days to 0.0108 at 61 to 92. The investment-grade effect is roughly half as large. It disappears for Black-Cox and the subordinator.
The simulator is the claim that holds up best, and it is the result I would use first. With 95 names, the Wiener-Hopf scheme reaches a maximum absolute tranche-upfront deviation of 0.0112 in 0.76 seconds. It reaches 0.0022 in 13.79 seconds. Fine-grid Euler requires 527 seconds, 8,000 paths and dt = 1/4032 to attain 0.0032. Refining Euler's grid does not produce monotonic improvement either. At 2,000 paths, it records 0.0058 with dt = 1/252 and 0.0129 with dt = 1/1008. Wiener-Hopf accuracy depends on path count alone, allowing the coarse clock at q = 16 to beat q = 48 in wall time at equal accuracy.
We could not run any of this. Our data does not contain single-name CDS composite curves or CDX tranche upfronts, and no equity or ETF substitute can stand in for a bootstrapped forward hazard vector or a tranche attachment band.
The risk-neutral common-shock rate is the fitted quantity that most needs independent scrutiny. Its lambda_c is higher for investment grade, at 0.25 to 0.43, than for high-yield, at 0.15 to 0.16. The authors argue that a risk-neutral measure should behave this way because senior investment-grade tranches insure disasters and require substantial tail weight. The explanation is plausible. This coordinate also reveals driver misspecification most clearly: the subordinator's investment-grade fits collect in small-severity corners and mark 0.447 frozen, against an independence level of 0.421.