The objective drives the portfolio to stop reacting to the signal. Its steepest-descent direction never changes with the current policy. Both results follow exactly from the algebra, and together they sharply limit what can be built from the paper.
Aldridge begins with an identity from earlier work: expected regret equals the covariance between uncertain costs and decisions, E[R] = Cov(c, π̂(c)). The identity applies whenever E[π̂(c)] = π*. Differentiating that covariance as a functional of the policy map π̂ gives Theorem 3.1. In direction φ, the Gâteaux derivative is Cov(c, φ(c)). The current π̂ disappears.
The Riesz representer is g(c) = c − c̄. Cauchy-Schwarz therefore makes −(c − c̄) the steepest-descent direction and +(c − c̄) the ascent direction for every policy. Aldridge calls the former contrarian and the latter momentum.
For the linear policy class π̂(c) = Ac + b, the calculation becomes C[A, b] = tr(AΣc). Hence ∇_A C = Σc, ∇_b C = 0_n and, entry by entry, ∂C/∂A_ij = Cov(c_i, c_j). Since the functional is linear in A, every second derivative vanishes. There are no interior critical points, so the optimum lies at a boundary.
Under 1'w = 1, w ≥ 0 and Σc ≻ 0, Corollary 5.1 gives A = 0 and b = Σc⁻¹1/(1'Σc⁻¹1), producing C[A, b] = 0. The abstract states the result plainly, describing a zero Hessian as "implying boundary-optimal solutions such as the minimum-variance portfolio."
Everything after that develops machinery around the same endpoint. Sign-gradient duality appears in Theorem 6.1: because r = −c and Σr = Σc, regret minimization and alpha maximization use the same update, A ← A ∓ ηΣc. Algorithm 8.1 projects the gradient step as A_{k+1} = Π_Z(A_k − ηΣ̂c). The prescribed step size is η ≤ 1/‖Σc‖₂, while Theorem 8.1 claims the linear rate (1 − ηλ_min(Σc))^k. Aldridge contrasts this with the generic O(1/√T) online-convex-optimization bound and attributes the difference to zero curvature.
Theorem 9.2 sets the iteration count at K(ε) = κ ln(C[A_0]/ε), where κ = λ_max(Σc)/λ_min(Σc). For covariance estimation, Theorem 9.3 requires N ≥ Cσ⁴d ln(d/δ)/ξ² observations to obtain ‖Σ̂c − Σc‖ ≤ ξ. Its stochastic iterates retain a statistical bias floor of ξ‖A_0‖_F/∆_min. Section 10 assigns an active tilt the cost ∆C = tr(∆A Σc), then writes the regret-optimal tilt as the rank-1 update ∆A* = (w_target − w_MVP)c̄'(c̄c̄')⁻¹. The paper contains no empirical or simulated exercise; every result is analytical.
Is the descent direction a forecast?
Any direction anti-correlated with realized costs lowers the functional because the derivative is Cov(c, φ). Among them, −(c − c̄) is L2-maximal. This establishes only the sign of a contemporaneous covariance. A tradable short-horizon reversal claim concerns c_{t+1} conditional on c_t.
The derivation never dates c relative to the decision. The policy π̂ is a measurable map from a realized cost vector to a weight, making lower decisions at high costs anti-correlated with those same costs by construction. An implementer must decide whether c means yesterday's returns or next week's. The identity itself imposes no timing.
Minimum variance enters elsewhere
The regret objective sees only A.
Because ∇_b C = 0_n, the functional is blind to the intercept. The proof sketch for Corollary 5.1 first minimizes over the budget simplex and arrives at A = 0. It then chooses b by solving the standard minimum-variance quadratic program subject to 1⊤b = 1. This second calculation supplies the minimum-variance weights. The covariance gradient supplies A* = 0, meaning the portfolio should ignore c; the portfolio level comes from a variance objective absent from the regret functional.
The expected-regret identity also carries a condition at this endpoint. Setting A = 0 makes the policy the constant b, so E[π̂(c)] = w_MVP. Satisfying E[π̂(c)] = π then requires the minimum-variance portfolio to be the correct average decision already. The paper says the identity holds whenever E[π̂(c)] = π*, yet gives no conditions under which the minimum-variance portfolio would meet that requirement.
A charge for tilting, without a payoff
Remark 7.1 states that C depends on Σc and is independent of c̄, giving ∂C/∂c̄ = 0. Expected returns cannot affect the objective. Aldridge acknowledges the limitation directly: "regret is insensitive to the cost mean," while the abstract uses the zero Hessian to justify boundary solutions such as the minimum-variance portfolio.
Expected returns instead enter through constraints. Corollary 5.2 imposes a return floor with multiplier κ, producing A = κΣc⁻¹r̄c̄' and the frontier C(µ0) = κ*c̄'Σc⁻¹r̄·tr(Σc). Section 10 then charges any tilt tr(∆A Σc). Given a target, the framework can rank two tilts by regret cost.
On its own, the framework names +(c − c̄), the ascent direction, as momentum. The conclusion calls its maximizer "the maximum-leverage momentum policy," which exposes the problem. Maximizing the functional continues until the leverage boundary. No finite position size can be selected using the cost mean because that mean never enters the functional.
Several proof sketches leave algebraic gaps. Corollary 5.2 uses the stationarity condition ∇_A L = Σc − κr̄c̄' = 0, requiring a full-rank Σc to equal a rank-1 matrix. The paper does not resolve that requirement. Section 10 inverts c̄c̄', which is rank 1 and singular for d > 1, leaving only a pseudo-inverse interpretation.
There is also a direct mismatch between Corollary 3.3, reporting H_A C = Σc ⊗ I_n, and Theorem 4.1, where all second derivatives are zero. Theorems 9.2 and 9.3 are presented as proof sketches too. Since the gradient is independent of A, the update A − ηΣ̂c translates the iterate by the same matrix at every step. The sketch for 9.2 bounds ‖A_{k+1}‖ ≤ ‖A_k‖(1 − ηλ_min(Σc)), then says C "inherits the same geometric contraction." Boundedness of the feasible set carries the argument. Meanwhile, Π_Z acts on the matrix A_k even though Z was introduced as the decision set in R^n.
The paper's most appealing claim concerns estimation from costs alone: "gradient descent on regret is a zero-instrumentation algorithm." The direction requires no realized P&L. It also provides no evidence that the selected cost definition or constraint set made money.
Our run and its limits
The figures here are ours. They come from one automated pass over the point-in-time top 100 US stocks from 2015-01-01 to 2024-12-31, rebalanced weekly at the close. Across that decade, the book returned 131.06% in total. It recorded a 0.57 Sharpe, 0.84 Sortino, 35.78% max drawdown, 0.24 Calmar and 23.02% annualized volatility.
A 0.57 Sharpe paired with a 35.78% drawdown is mediocre for a long-only minimum-variance book carrying a reversal overlay.
Aldridge reports no performance figures. The paper has no dataset, simulation or backtest, so the comparison is one-sided by construction. Our 0.57 Sharpe and 131.06% total return have no author result against which they can be compared.
We estimated Σ̂ from 252 daily returns, applying 10% diagonal shrinkage and a ridge equal to 1e-6 times median variance. A long-only minimum-variance QP on that matrix provides the anchor. The tilt uses the regret-minimizing sign: minus 0.25 of an L1-normalized 20-day mean return, centered for each asset by its own 252-day mean.
Afterward came simplex projection, a 10% per-name cap, a 20% one-way weekly turnover cap and a 10% annualized ex-ante volatility target, with remaining capital held in cash. Trading costs were four tenths of a cent a share with a $1 minimum. Slippage was modeled at zero.
The 10% ex-ante volatility target failed to keep realized volatility close to 10%. The run produced 23.02% annualized volatility and a 35.78% drawdown. Across the decade, the forecast understated realized risk by better than a factor of two. Scaling against that forecast and placing the remainder in cash controls the ex-ante figure alone, and this failure weighs first against the reported performance.
Our tilt uses a 20-day mean, whereas Σ̂ comes from 1-day returns. That roughly 20x base mismatch means Corollary 3.2's ±(c − c̄) is applied to a different c from the one used to estimate the covariance. The matrix iterate also never trades. Σ̂ affects the minimum-variance anchor alone, leaving A_{k+1} = Π_Z(A_k ∓ ηΣ̂) orphaned in our implementation.
The run is a constrained minimum-variance book with a small reversal overlay, one reasonable interpretation of the paper. The gradient algorithm stayed idle. We did not separate the anchor's contribution from the tilt's, so the 131.06% cannot be attributed between them. Whatever their strength, these figures primarily describe our single automated pass rather than judge Aldridge's work.
We have made the opposite measurement error before, assigning an overlay credit for returns generated by a static anchor (the growth-defensive timer note).
A covariance-gradient tilt beating the same-constraint minimum-variance anchor plus a plain 20-day reversal overlay would change my view, provided the comparison is net of costs and uses the same universe. Until the objective can size the tilt instead of merely charging tr(∆A Σc) for it, its descent direction remains the correct derivative of a quantity that ceases to matter where the allocation choice begins.