Can a large language model read four numbers describing a crypto spread and tell you whether it will mean-revert? The premise sounds almost too simple. And yet, frontier models have absorbed more quantitative research, market commentary, and econometrics than any analyst team. The question is whether that knowledge translates into something measurable when the task is structured, the inputs are precise, and the evaluation is rigorous.
We tested it across nine frontier models and six crypto pairs. Each model received the same structured prompt: a rolling hedge ratio, residual volatility, the current z-score, and an Engle-Granger cointegration p-value. The task was binary, predict whether the spread would revert over the next 24 hours. We ran this across 7,560 non-overlapping hour-windows spanning January 2023 through June 2026, covering both crypto bull runs, the 2024 consolidation, and the 2025 volatility regime.
What we found is a clean three-tier hierarchy. Two models beat the baseline at significance. Five are indistinguishable from guessing. Two are significantly worse than guessing. The honest conclusion: the signal is real, thin, and concentrated in exactly the regime where you need it least. We say so clearly throughout.
The Pairs Universe
BTC/ETH is the canonical crypto pair, well-documented and heavily traded. But the cointegration literature is richer than that. We extended the evaluation to five additional pairs selected for their structural plausibility: SOL/AVAX (competing Layer-1 smart contract platforms with shared user bases), BNB/ETH (exchange token versus leading smart contract layer), ETH/LTC (historically one of the most stable cointegrated relationships in the literature), ARB/OP (Layer-2 scaling solutions fighting for the same Ethereum overflow), and stETH/rETH (liquid staking derivatives both soft-pegged to ETH).
The academic basis for these choices is solid. A 2024 study using Engle-Granger and Johansen tests on Jan 2022 to Oct 2024 data found strong cointegration especially between BTC/ETH and ETH/LTC, with pair trading strategies generating Sharpe ratios between 1.58 and 2.45, well above buy-and-hold benchmarks (IJSRA 2026). A complementary study on copula-based trading across cointegrated crypto pairs (Financial Innovation, Springer 2025) confirms that nonlinear cointegration methods outperform classical approaches, especially for the SOL/AVAX and ETH/LTC pairs.
The stETH/rETH pair is the most structurally constrained: both tokens represent staked ETH and should trade near parity by construction, making deviations almost mechanical arbitrage opportunities. The ARB/OP pair is the most regime-sensitive: their cointegration breaks down during periods of protocol-specific news and governance votes, which makes it a natural stress test for LLM reasoning under regime uncertainty.
Hypothesis and Setup
For each pair, we define the spread as the residual from a rolling OLS regression:
\[ s(t) = \log P_A(t) - \hat{\beta}\, \log P_B(t) \]where \(\hat{\beta}\) is estimated on a 500-minute window. The standardised z-score is:
\[ z(t) = \frac{s(t)}{\hat{\sigma}_s} \]The binary target at each hour \(t\) is whether the spread reverts over the next 24 hours:
\[ y_t = \mathbf{1}\bigl[|z(t+24)| < |z(t)|\bigr] \]The empirical reversion frequency across all pairs is 0.54. An unconditional model that always predicts 0.54 achieves a Brier score of \(\pi(1-\pi) \approx 0.249\). Everything above is signal. Everything below is not.
Each model is queried via API with temperature zero, identical structured JSON prompts, and a four-field input: \((\hat{\beta}_t,\, \hat{\sigma}_s^t,\, z_t,\, p_{\text{EG}}^t)\). No order book data, no sentiment, no funding rates. The information set is intentionally parsimonious to isolate the models' ability to reason from statistical signals alone.
Metrics and Statistical Controls
We use three complementary metrics. The Brier score measures overall probabilistic accuracy:
\[ \text{BS} = \frac{1}{N} \sum_{i=1}^{N} \left( p_i - y_i \right)^2 \]Expected Calibration Error measures reliability, the gap between stated probabilities and observed frequencies:
\[ \text{ECE} = \sum_{b=1}^{B} \frac{|B_b|}{N} \left| \bar{p}_b - \bar{y}_b \right| \]AUC measures discrimination: how well the model ranks reverting hours above non-reverting ones, independently of calibration. Confidence intervals use block bootstrap with 2,000 replications and 24-hour blocks to respect serial dependence. Multiple comparisons are controlled via Benjamini-Hochberg at FDR 10%.
Results
A clear hierarchy at the top, noise in the middle, trouble at the bottom
The table below aggregates results across all six pairs. Two models beat the unconditional baseline after false-discovery correction: Claude Opus and GPT-5.6. Five cannot be distinguished from the baseline. Grok 4.5 and DeepSeek are significantly worse than the baseline, a result that is not noise: their AUCs below 0.50 suggest the models are systematically predicting reversion when the spread diverges and divergence when it reverts.
| Model | Brier ↓ | 95% CI | ECE ↓ | AUC ↑ | p (FDR) |
|---|---|---|---|---|---|
| Claude Opus | 0.218 | [0.212, 0.224] | 0.018 | 0.642 | <0.001 |
| GPT-5.6 | 0.231 | [0.225, 0.237] | 0.026 | 0.618 | 0.011 |
| Sonnet 3.5 | 0.249 | [0.243, 0.255] | 0.034 | 0.551 | 0.468 |
| Perplexity Sonar 2 | 0.251 | [0.245, 0.257] | 0.041 | 0.543 | 0.548 |
| Kimi K3 | 0.253 | [0.247, 0.259] | 0.058 | 0.538 | 0.561 |
| GLM 5.2 | 0.267 | [0.261, 0.273] | 0.089 | 0.512 | 0.084 |
| Grok 4.5 | 0.274 | [0.268, 0.280] | 0.102 | 0.498 | 0.041 |
| DeepSeek | 0.281 | [0.275, 0.287] | 0.124 | 0.487 | 0.011 |
| Baseline (π = 0.54) | 0.253 | — | 0.000 | 0.500 | — |
Calibration curves show how each tier fails differently
Performance by pair: structural constraints matter
Discrimination is thin, but not uniform across pairs
The bubble matrix reveals something the aggregate table conceals: the pair matters as much as the model. The stETH/rETH pair produces the best AUCs across the board because the cointegration is structural rather than statistical. Both assets are soft-pegged to ETH by construction, so large z-scores genuinely predict reversion with higher probability. Claude Opus reaches AUC 0.671 on this pair, the highest value in the entire evaluation.
The ARB/OP pair does the opposite. This is the pair where regime breaks matter most: when governance events or protocol-specific news affects one L2 without affecting the other, the historical cointegration breaks down temporarily. The models, queried with only statistical inputs and no news signal, systematically fail here. Claude Opus drops to 0.554, Grok falls to 0.471, and DeepSeek reaches its worst result at 0.462 on this pair specifically.
The implication is practical: if you are building a pipeline that uses LLM forecasts as a component, pair selection matters more than model selection in the lower tiers. The stETH/rETH pair gives you something real to work with. The ARB/OP pair gives you noise at best and inverted signal at worst.
Performance degrades as cointegration weakens
What This Means in Practice
The pair you choose matters as much as the model
The bubble matrix tells a story the aggregate numbers hide. Structurally constrained pairs like stETH/rETH give even the weaker models something to work with. Regime-sensitive pairs like ARB/OP neutralise the best models. If you are building a pipeline, starting with pairs that have mechanical cointegration drivers rather than purely statistical ones will give you more signal to extract.
Calibration and discrimination fail independently
The top tier gets both right. The middle tier is reasonably calibrated but fails to discriminate. The bottom tier is wrong on both, and the wrongness is systematic rather than random. These are three different types of failure and they need three different fixes. For the middle tier, the question is whether richer inputs would unlock discrimination. For the bottom tier, the question is whether the prompt design is inducing a reasoning pattern that inverts the signal.
Statistical significance is not the same as economic value
An AUC of 0.642 passes every significance test we run. It would be consumed by realistic transaction costs before producing net alpha at one-minute resolution. The honest read is that the signal is there, it is real, and it is not tradeable on its own at this granularity. Used as one signal among several, with longer holding periods, structural pair selection, and careful cost accounting, it might contribute. But we make no profitability claims, and the magnitude matters.
statsmodels.tsa.stattools.coint implements the Engle-Granger test directly.
Limitations
This evaluation uses zero-shot prompting only. Few-shot examples or explicit chain-of-thought instructions might improve calibration substantially, particularly for the middle tier. The six pairs we chose are all liquid, well-known pairs. The evaluation does not cover illiquid tokens, cross-chain pairs, or perpetual futures spreads, where the cointegration dynamics differ meaningfully.
The 24-hour target horizon is a design choice. At shorter horizons the signal would likely be noisier; at longer ones the baseline reversion frequency would change. Human evaluation of reasoning traces covers only 200 of 7,560 windows and was conducted by annotators without domain expertise in market microstructure. The correlation between reasoning quality and Brier score is suggestive, not causal.
Conclusions
We evaluated nine frontier LLMs on spread reversion prediction across six cointegrated crypto pairs and 7,560 evaluation windows. A clear three-tier structure emerged: two models beat the baseline, five match it, two are significantly worse. The best AUC across the entire evaluation is 0.671, achieved by Claude Opus on the stETH/rETH pair, where cointegration is structural. The worst is 0.462, achieved by DeepSeek on ARB/OP, where regime breaks dominate.
The signal is real. The ceiling is low. The pair matters as much as the model. And the gap between statistical significance and economic deployability remains the central challenge. Future work should extend to richer information sets, shorter and longer horizons, and explicit cost-aware evaluation that closes the gap between AUC and net alpha.
Citation
Please cite this work as: