Hugo HoennWorkX-rayContact

Cross-Asset Trend & Carry Program — and a test of ML against simple rules

Part 3 of the GS QIS product review: a systematic macro program across 21 equity, bond, currency and commodity markets, benchmarked against Goldman Sachs' Managed Futures Strategy Fund (GMSSX), with a gradient-boosted overlay tested under purged walk-forward validation. Hugo Hoenn, September 2026.

Code and data on GitHub →Public data only · every table reproducible from the repo
Goldman Sachs QIS product review, from public data: Part 1 · Absolute Return Tracker · Part 2 · ActiveBeta U.S. Large Cap · Part 3 · Cross-asset trend program vs. Managed Futures Strategy

The QIS posting describes a team that "manages exposures to global stock, bond, currency, and commodity markets to generate excess return... leveraging the latest technology, including Artificial Intelligence techniques." This project does both halves, and reports what each did.

1. The program

Markets. 21 liquid markets via total-return ETFs (excess return over T-bills ≈ a collateralised futures position, with roll yield embedded): 6 equity indices, 4 government bond buckets, 6 commodities, 5 currencies. History from 2007–08, so the sample includes the financial crisis.

Signals. Time-series momentum at 1, 3 and 12 months (return scaled by volatility, clipped, averaged), and carry where it can be measured from free data: FX carry from 3-month interbank rate differentials (FRED, lagged a month), bond carry from the US term spread. Equities and commodities are trend only. Signal = trend for equities and commodities; 0.75 trend + 0.25 carry for FX and bonds.

Sizing. Each market scaled to a volatility target, equal risk to each of the four asset classes, portfolio scaled to 10% ex-ante volatility on a 60-day covariance, gross leverage capped at 3×. Monthly rebalance. Costs 10 bps per unit of turnover, which is ETF-like and conservative: at futures-like costs of 3 bps the annualised return would be 6.6% instead of 5.1%.

Growth

2. Results, Feb 2008 – Sep 2026

Ann. return Vol Sharpe Max DD Skew Corr S&P Corr 60/40 Corr factor strat.
Trend + carry (program) 5.1% 10.2% 0.41 -20.7% 0.52 -0.15 -0.15 -0.10
Trend only 5.1% 10.3% 0.41 -16.1% 0.54 -0.18 -0.19 -0.14
Carry only 1.9% 7.9% 0.11 -34.4% -0.19 0.09 0.15 0.12
S&P 500 11.6% 15.6% 0.71 -46.3% -0.57 1.00 0.98 0.98
60/40 8.2% 10.0% 0.71 -29.7% -0.60 0.98 1.00 0.97

The program earned 5.1% a year at a 0.41 Sharpe with a -20.7% worst drawdown, positive skew, and a correlation of -0.15 to the S&P 500, -0.15 to 60/40 and -0.10 to the multi-factor equity strategy from Part 0. That combination — modest Sharpe, positive skew, negative correlation to everything a client already owns — is the entire case for trend following, and it showed up on schedule:

Period Program Trend only GMSSX S&P 500 60/40
GFC (Feb 2008–Feb 2009) +15.2% +16.2% n/a -44.9% -28.8%
Aug–Sep 2011 +1.6% +0.8% n/a -12.1% -6.4%
Q4 2018 -5.2% -4.6% +1.3% -13.5% -7.4%
COVID (Feb–Mar 2020) +15.0% +14.8% +4.9% -19.4% -11.5%
2022 (full year) +13.7% +24.6% +20.6% -18.2% -15.8%
Trend winter (2012–2019) +23.0% +22.2% +11.8% +200.9% +114.8%

The carry sleeve, as built here from free data, added nothing: on its own it earned a 0.11 Sharpe with a -34.4% drawdown, and its long-duration tilt cost the program 11 points of return in 2022 relative to trend alone. It is kept in the headline version because it was pre-specified, and reported so a reader can see it. A carry signal worth having needs futures-curve data for commodities and equities, which is the first upgrade.

Attribution

Year Program Trend only GMSSX S&P 500
2008 +18.1% +18.4% +0.0% -32.7%
2009 -4.6% -6.8% +0.0% +26.4%
2010 +6.8% +5.4% +0.0% +15.1%
2011 +9.7% +6.7% +0.0% +1.9%
2012 +0.1% -1.0% +5.8% +16.0%
2013 +3.8% +4.2% -4.4% +32.3%
2014 +3.9% +1.9% -2.6% +13.5%
2015 +8.5% +10.4% +10.4% +1.2%
2016 +1.3% -0.6% -0.6% +12.0%
2017 +15.4% +12.3% +2.7% +21.7%
2018 -3.4% +0.7% -1.7% -4.6%
2019 -7.0% -6.3% +2.4% +31.2%
2020 +11.7% +10.6% +7.0% +18.3%
2021 -0.2% +1.0% +5.0% +28.7%
2022 +13.7% +24.6% +20.6% -18.2%
2023 +0.9% -2.1% -3.7% +26.2%
2024 +6.4% +6.2% -5.1% +24.9%
2025 +16.5% +14.3% +0.6% +17.7%
2026 -2.0% -0.1% +16.3% +12.7%

3. Against Goldman's managed-futures fund

GMSSX (QIS, since Feb 2012) is the natural comparison. Common window from Mar 2012:

Ann. return Vol Sharpe Max DD Skew Corr S&P Corr 60/40 Corr factor strat.
Program (trend + carry) 4.5% 10.1% 0.34 -20.7% 0.59 -0.16 -0.16 -0.11
Program (trend only) 5.0% 10.2% 0.38 -16.1% 0.60 -0.19 -0.21 -0.14
GS Managed Futures Strategy (GMSSX) 3.4% 9.3% 0.24 -19.4% -0.22 -0.09 -0.11 -0.06
DBMF (since May 2019) 9.6% 11.2% 0.64 -17.3% 0.00 -0.15 -0.22 -0.12
KMLM (since Dec 2020) 6.5% 13.0% 0.31 -25.9% 0.10 -0.36 -0.45 -0.37
60/40 9.4% 9.1% 0.87 -20.0% -0.39 0.98 1.00 0.96

The program's correlation to GMSSX is 0.63 at 8.4% tracking error: the same family of bets, implemented differently. Over the fund's life the program earned a higher Sharpe (0.34 vs 0.24) with positive rather than negative skew; GMSSX did better in Q4 2018 and, notably, in 2022 (+20.6% vs +13.7% for the program, +24.6% for trend alone). The newer trend ETFs (DBMF, KMLM) have short records that begin after the 2012–19 "trend winter", so their Sharpe ratios are not comparable. This is a backtest against a live fund with real costs and constraints, so the gap should be read as "in the same league", not "better".

Versus funds

4. Does machine learning help? Tested honestly, no.

A gradient-boosted model (sklearn HistGradientBoostingRegressor, depth 3, 200 trees) predicts each market's next-month volatility-scaled excess return from twelve features: trend at four horizons, carry, volatility, the volatility ratio, 12-month drawdown and asset class. It is trained on a pooled panel of ~5,000 market-months and re-fitted every January on all data ending two months earlier — an expanding window with a one-month embargo, so no training target overlaps a prediction. Predictions are converted to positions and run through the identical sizing engine.

Ann. return Vol Sharpe Max DD Skew
Trend rule (same window) 5.2% 10.1% 0.42 -16.1% 0.59
ML overlay (gradient boosting) -1.3% 7.3% -0.33 -39.6% 0.17
50/50 blend 1.8% 8.8% 0.09 -17.2% 0.96

Out of sample from Mar 2010, the model's monthly rank correlation with realised returns was 0.03 (t = 1.5) against 0.06 (t = 2.4) for the plain trend rule. As a strategy it lost money. Permutation importance shows it leaning on the same trend features the rule uses, plus volatility:

Feature 2016 2020 2024
z1 0.023 0.040 0.037
z3 0.061 0.024 0.029
z6 0.068 0.048 0.029
z12 0.047 0.040 0.028
carry 0.017 0.017 0.011
vol 0.054 0.030 0.032
vol_ratio 0.039 0.031 0.041
dd12 0.036 0.028 0.022

The reading is not that ML cannot work in this setting; it is that with ~5,000 noisy observations and a signal-to-noise ratio this low, a flexible learner overfits and a rule with a 100-year literature behind it does not. The blend did not rescue it. This is the result a QIS team would expect and the one an applicant is least tempted to publish, which is why it is here.

ML vs rules

What this cannot tell you

Code

src/data.py (prices, FRED rates), src/strategy.py (signals, sizing, simulation, sleeves), src/ml.py (panel features, purged walk-forward, importance), src/analysis.py (metrics, benchmarks, crisis table, attribution, charts). Tables in output/*.csv.

Independent research for discussion. Not affiliated with or endorsed by Goldman Sachs. Not investment advice.