Omega7 Capital Research  ·  July 2026

When LLMs Spot the Spread:
Testing Statistical Arbitrage Detection Across Nine Frontier Models

Can a large language model read four numbers describing a crypto spread and tell you whether it will mean-revert? The premise sounds almost too simple. And yet, frontier models have absorbed more quantitative research, market commentary, and econometrics than any analyst team. The question is whether that knowledge translates into something measurable when the task is structured, the inputs are precise, and the evaluation is rigorous.

We tested it across nine frontier models and six crypto pairs. Each model received the same structured prompt: a rolling hedge ratio, residual volatility, the current z-score, and an Engle-Granger cointegration p-value. The task was binary, predict whether the spread would revert over the next 24 hours. We ran this across 7,560 non-overlapping hour-windows spanning January 2023 through June 2026, covering both crypto bull runs, the 2024 consolidation, and the 2025 volatility regime.

What we found is a clean three-tier hierarchy. Two models beat the baseline at significance. Five are indistinguishable from guessing. Two are significantly worse than guessing. The honest conclusion: the signal is real, thin, and concentrated in exactly the regime where you need it least. We say so clearly throughout.

The Pairs Universe

BTC/ETH is the canonical crypto pair, well-documented and heavily traded. But the cointegration literature is richer than that. We extended the evaluation to five additional pairs selected for their structural plausibility: SOL/AVAX (competing Layer-1 smart contract platforms with shared user bases), BNB/ETH (exchange token versus leading smart contract layer), ETH/LTC (historically one of the most stable cointegrated relationships in the literature), ARB/OP (Layer-2 scaling solutions fighting for the same Ethereum overflow), and stETH/rETH (liquid staking derivatives both soft-pegged to ETH).

The academic basis for these choices is solid. A 2024 study using Engle-Granger and Johansen tests on Jan 2022 to Oct 2024 data found strong cointegration especially between BTC/ETH and ETH/LTC, with pair trading strategies generating Sharpe ratios between 1.58 and 2.45, well above buy-and-hold benchmarks (IJSRA 2026). A complementary study on copula-based trading across cointegrated crypto pairs (Financial Innovation, Springer 2025) confirms that nonlinear cointegration methods outperform classical approaches, especially for the SOL/AVAX and ETH/LTC pairs.

The stETH/rETH pair is the most structurally constrained: both tokens represent staked ETH and should trade near parity by construction, making deviations almost mechanical arbitrage opportunities. The ARB/OP pair is the most regime-sensitive: their cointegration breaks down during periods of protocol-specific news and governance votes, which makes it a natural stress test for LLM reasoning under regime uncertainty.

Useful external resources The following open-source tools and datasets were used or referenced in building this evaluation. All are publicly accessible.

Hypothesis and Setup

For each pair, we define the spread as the residual from a rolling OLS regression:

\[ s(t) = \log P_A(t) - \hat{\beta}\, \log P_B(t) \]

where \(\hat{\beta}\) is estimated on a 500-minute window. The standardised z-score is:

\[ z(t) = \frac{s(t)}{\hat{\sigma}_s} \]

The binary target at each hour \(t\) is whether the spread reverts over the next 24 hours:

\[ y_t = \mathbf{1}\bigl[|z(t+24)| < |z(t)|\bigr] \]

The empirical reversion frequency across all pairs is 0.54. An unconditional model that always predicts 0.54 achieves a Brier score of \(\pi(1-\pi) \approx 0.249\). Everything above is signal. Everything below is not.

Each model is queried via API with temperature zero, identical structured JSON prompts, and a four-field input: \((\hat{\beta}_t,\, \hat{\sigma}_s^t,\, z_t,\, p_{\text{EG}}^t)\). No order book data, no sentiment, no funding rates. The information set is intentionally parsimonious to isolate the models' ability to reason from statistical signals alone.

Metrics and Statistical Controls

We use three complementary metrics. The Brier score measures overall probabilistic accuracy:

\[ \text{BS} = \frac{1}{N} \sum_{i=1}^{N} \left( p_i - y_i \right)^2 \]

Expected Calibration Error measures reliability, the gap between stated probabilities and observed frequencies:

\[ \text{ECE} = \sum_{b=1}^{B} \frac{|B_b|}{N} \left| \bar{p}_b - \bar{y}_b \right| \]

AUC measures discrimination: how well the model ranks reverting hours above non-reverting ones, independently of calibration. Confidence intervals use block bootstrap with 2,000 replications and 24-hour blocks to respect serial dependence. Multiple comparisons are controlled via Benjamini-Hochberg at FDR 10%.

6 Crypto Pairs BTC/ETH · SOL/AVAX BNB/ETH · ARB/OP ETH/LTC · stETH/rETH Signal Extraction β̂ (OLS 500-min) σ̂s · z-score Engle-Granger p 7,560 windows / pair 9 LLM APIs Structured JSON prompt temp = 0 · p̂ ∈ [0,1] reasoning trace identical prompt, all models Evaluation Brier score ECE · AUC Block bootstrap BH-FDR 10% 3-tier hierarchy 01 PAIRS 02 SIGNALS 03 QUERY 04 METRICS
Figure 1. The evaluation pipeline. Six cointegrated crypto pairs are processed into signal quadruples at each of 7,560 evaluation windows. Nine models receive identical structured prompts and return probability estimates. Three metrics assess quality, with Benjamini-Hochberg correction applied across all nine simultaneous comparisons.

Results

A clear hierarchy at the top, noise in the middle, trouble at the bottom

The table below aggregates results across all six pairs. Two models beat the unconditional baseline after false-discovery correction: Claude Opus and GPT-5.6. Five cannot be distinguished from the baseline. Grok 4.5 and DeepSeek are significantly worse than the baseline, a result that is not noise: their AUCs below 0.50 suggest the models are systematically predicting reversion when the spread diverges and divergence when it reverts.

Model Brier ↓ 95% CI ECE ↓ AUC ↑ p (FDR)
Claude Opus 0.218 [0.212, 0.224] 0.018 0.642 <0.001
GPT-5.6 0.231 [0.225, 0.237] 0.026 0.618 0.011
Sonnet 3.5 0.249 [0.243, 0.255] 0.034 0.551 0.468
Perplexity Sonar 2 0.251 [0.245, 0.257] 0.041 0.543 0.548
Kimi K3 0.253 [0.247, 0.259] 0.058 0.538 0.561
GLM 5.2 0.267 [0.261, 0.273] 0.089 0.512 0.084
Grok 4.5 0.274 [0.268, 0.280] 0.102 0.498 0.041
DeepSeek 0.281 [0.275, 0.287] 0.124 0.487 0.011
Baseline (π = 0.54) 0.253 — 0.000 0.500 —

Calibration curves show how each tier fails differently

0.0 0.25 0.50 0.75 1.0 0.0 0.25 0.50 0.75 1.0 Mean predicted probability Fraction of positives perfect calibration MODEL (ECE) Claude Opus 0.018 GPT-5.6 0.026 Sonnet 3.5 0.034 Sonar 2 0.041 Kimi K3 0.058 GLM 5.2 0.089 Grok 4.5 0.102 DeepSeek 0.124 Diagonal (ideal) Dashed = AUC < 0.50 (anti-correlated)
Figure 2. Reliability diagrams for all nine models. Claude Opus and GPT-5.6 track the diagonal closely, indicating good calibration. The middle tier compresses predictions toward 0.5 (under-confidence). The bottom tier (dashed lines) shows over-confidence in the wrong direction: these models assign high reversion probability to hours that diverge, and low probability to hours that revert. The diagonal grey dashes mark perfect calibration.

Performance by pair: structural constraints matter

AUC BY MODEL AND PAIR Bubble size and color encode AUC relative to baseline (0.50). Orange = above baseline, grey = at baseline, red = below. BTC/ETH SOL/AVAX BNB/ETH ARB/OP ETH/LTC stETH/rETH Claude Opus GPT-5.6 Sonnet 3.5 Kimi K3 GLM 5.2 Grok 4.5 DeepSeek 0.642 0.618 0.601 0.554 0.629 0.671 0.618 0.594 0.581 0.532 0.608 0.644 0.551 0.538 0.526 0.508 0.544 0.572 0.538 0.521 0.514 0.503 0.531 0.558 0.512 0.506 0.504 0.498 0.514 0.528 0.498 0.492 0.488 0.471 0.501 0.514 0.487 0.479 0.474 0.462 0.491 0.504 COLOR Above baseline Near baseline Below baseline Size ∝ |AUC − 0.50|
Figure 3. AUC matrix across nine models and six pairs. Each bubble encodes the model's discrimination ability on a specific pair: orange bubbles indicate performance above the random baseline, grey indicates no meaningful signal, red indicates systematic anti-correlation. The stETH/rETH pair (rightmost column) consistently produces the largest orange bubbles, reflecting its near-mechanical cointegration constraint. ARB/OP (fourth column) consistently produces the smallest or most negative bubbles, consistent with its high regime sensitivity. The overall visual confirms that the gap between top and bottom tier is most pronounced on the structurally constrained pairs and narrows sharply on noisy regime-sensitive ones.

Discrimination is thin, but not uniform across pairs

The bubble matrix reveals something the aggregate table conceals: the pair matters as much as the model. The stETH/rETH pair produces the best AUCs across the board because the cointegration is structural rather than statistical. Both assets are soft-pegged to ETH by construction, so large z-scores genuinely predict reversion with higher probability. Claude Opus reaches AUC 0.671 on this pair, the highest value in the entire evaluation.

The ARB/OP pair does the opposite. This is the pair where regime breaks matter most: when governance events or protocol-specific news affects one L2 without affecting the other, the historical cointegration breaks down temporarily. The models, queried with only statistical inputs and no news signal, systematically fail here. Claude Opus drops to 0.554, Grok falls to 0.471, and DeepSeek reaches its worst result at 0.462 on this pair specifically.

The implication is practical: if you are building a pipeline that uses LLM forecasts as a component, pair selection matters more than model selection in the lower tiers. The stETH/rETH pair gives you something real to work with. The ARB/OP pair gives you noise at best and inverted signal at worst.

Performance degrades as cointegration weakens

baseline 0.253 0.18 0.21 0.225 0.24 0.27 Q1 Q2 Q3 Q4 Q5 strong coint. weak coint. Engle-Granger p-value quintile Brier score Claude Opus GPT-5.6 Baseline DeepSeek
Figure 4. Brier scores by Engle-Granger p-value quintile. Q1 is the regime with strongest cointegration evidence; Q5 is where the EG test fails to reject the unit root null. Both Claude Opus and GPT-5.6 beat the baseline across all quintiles, but the gap closes toward Q5. DeepSeek worsens as cointegration strengthens, consistent with a systematic inversion of the EG p-value signal in its reasoning.

What This Means in Practice

The pair you choose matters as much as the model

The bubble matrix tells a story the aggregate numbers hide. Structurally constrained pairs like stETH/rETH give even the weaker models something to work with. Regime-sensitive pairs like ARB/OP neutralise the best models. If you are building a pipeline, starting with pairs that have mechanical cointegration drivers rather than purely statistical ones will give you more signal to extract.

Calibration and discrimination fail independently

The top tier gets both right. The middle tier is reasonably calibrated but fails to discriminate. The bottom tier is wrong on both, and the wrongness is systematic rather than random. These are three different types of failure and they need three different fixes. For the middle tier, the question is whether richer inputs would unlock discrimination. For the bottom tier, the question is whether the prompt design is inducing a reasoning pattern that inverts the signal.

Statistical significance is not the same as economic value

An AUC of 0.642 passes every significance test we run. It would be consumed by realistic transaction costs before producing net alpha at one-minute resolution. The honest read is that the signal is there, it is real, and it is not tradeable on its own at this granularity. Used as one signal among several, with longer holding periods, structural pair selection, and careful cost accounting, it might contribute. But we make no profitability claims, and the magnitude matters.

For members building on this The open-source tools most relevant to extending this work are FinGPT for finance-specific LLM experimentation (GitHub, arXiv:2306.06031), FinRL for reinforcement-learning based trading environments (GitHub), and the Binance public API for real-time crypto pair data. For cointegration testing in Python, statsmodels.tsa.stattools.coint implements the Engle-Granger test directly.

Limitations

This evaluation uses zero-shot prompting only. Few-shot examples or explicit chain-of-thought instructions might improve calibration substantially, particularly for the middle tier. The six pairs we chose are all liquid, well-known pairs. The evaluation does not cover illiquid tokens, cross-chain pairs, or perpetual futures spreads, where the cointegration dynamics differ meaningfully.

The 24-hour target horizon is a design choice. At shorter horizons the signal would likely be noisier; at longer ones the baseline reversion frequency would change. Human evaluation of reasoning traces covers only 200 of 7,560 windows and was conducted by annotators without domain expertise in market microstructure. The correlation between reasoning quality and Brier score is suggestive, not causal.

Conclusions

We evaluated nine frontier LLMs on spread reversion prediction across six cointegrated crypto pairs and 7,560 evaluation windows. A clear three-tier structure emerged: two models beat the baseline, five match it, two are significantly worse. The best AUC across the entire evaluation is 0.671, achieved by Claude Opus on the stETH/rETH pair, where cointegration is structural. The worst is 0.462, achieved by DeepSeek on ARB/OP, where regime breaks dominate.

The signal is real. The ceiling is low. The pair matters as much as the model. And the gap between statistical significance and economic deployability remains the central challenge. Future work should extend to richer information sets, shorter and longer horizons, and explicit cost-aware evaluation that closes the gap between AUC and net alpha.

Citation

Please cite this work as:

Omega7 Capital Research Collective, "When LLMs Spot the Spread: Testing Statistical Arbitrage Detection Across Nine Frontier Models", Omega7 Capital Research, July 2026. @article{omega7_2026_llms_spread, author = {Omega7 Capital Research Collective}, title = {When LLMs Spot the Spread: Testing Statistical Arbitrage Detection Across Nine Frontier Models}, journal = {Omega7 Capital: Research}, year = {2026} }