The random walk hypothesis is among the most durable ideas in quantitative finance. In its simplest form it states that successive price changes are independent and identically distributed, so that past returns carry no usable information about future returns. If this holds, then technical prediction is futile and weak-form market efficiency is satisfied. The hypothesis is convenient, mathematically tractable, and as a first approximation, frequently defensible.
It is also not exactly true, and the interesting question is not whether it fails but where, by how much, and for whom. In this study we apply a battery of statistical tests to daily returns from 500 large-capitalisation U.S. equities over a twelve-year window. We ask a narrow, measurable question: for each stock, can we reject the null hypothesis that its returns behave like a random walk?
Our headline finding is unsurprising in direction but useful in magnitude. Roughly seventy percent of the stocks we examined are statistically indistinguishable from a random walk at conventional significance levels. The remaining thirty percent are not, and the departures are not randomly distributed across the market. They cluster in identifiable sectors, they are entangled with volatility, and they are far weaker than they first appear once trading frictions are taken seriously. We try throughout to be honest about how thin the exploitable signal really is.
Hypothesis and Motivation
Let \(p_t = \log P_t\) denote the log price of an asset at time \(t\). The random walk hypothesis with drift posits:
\[ p_t = \mu + p_{t-1} + \varepsilon_t, \qquad \varepsilon_t \overset{\text{iid}}{\sim} (0,\, \sigma^2) \]Under this model the daily log return \(r_t = p_t - p_{t-1} = \mu + \varepsilon_t\) is unpredictable from its own history. The best forecast of tomorrow's return, conditional on all information up to today, is simply the long-run drift \(\mu\):
\[ \mathbb{E}[r_{t+1} \mid \mathcal{F}_t] = \mu \]This is the weak form of the Efficient Market Hypothesis (Fama, 1970), and it makes a sharp, testable prediction. If returns are serially uncorrelated, then the autocorrelation function \(\rho(k) = \text{Corr}(r_t, r_{t-k})\) should satisfy \(\rho(k) = 0\) for all \(k \neq 0\), the variance of \(k\)-period returns should grow linearly in \(k\):
\[ \text{Var}\!\left(\sum_{s=1}^{k} r_{t+s}\right) = k\,\sigma^2 \]and runs of consecutive same-sign returns should appear with the frequency expected under independence. Each of these implications gives us a separate test, and each test probes a slightly different way the null can fail.
We emphasise what a rejection does and does not mean. Rejecting the random walk null tells us that returns contain some serial structure. It does not tell us the structure is large, stable across time, or exploitable after costs. A great deal of the empirical literature on market efficiency, including the foundational survey by Lo and MacKinlay (1988), founders precisely on this gap between statistical significance and economic significance.
Methodology
Data
Our sample consists of 500 large-capitalisation U.S. equities, selected to approximate the composition of a broad large-cap index as of January 2013. We use adjusted daily closing prices from January 2013 through December 2024, giving approximately 3,020 trading days per name. Prices are adjusted for splits and dividends, and we compute continuously compounded daily log returns \(r_t = \log(P_t/P_{t-1})\). Names with more than five percent missing observations were excluded and replaced with the next-largest eligible stock, introducing a mild survivorship bias we flag in the discussion. All returns are winsorised at the 0.1 and 99.9 percentiles to limit the influence of extreme observations associated with earnings surprises and the March 2020 dislocation.
Statistical Tests
We apply four complementary tests to each return series. The Ljung-Box test on the first ten autocorrelation lags aggregates serial correlation into a single \(\chi^2\) statistic:
\[ Q(m) = n(n+2)\sum_{k=1}^{m} \frac{\hat{\rho}(k)^2}{n-k} \;\overset{H_0}{\sim}\; \chi^2(m) \]The Lo-MacKinlay variance ratio test compares the variance of \(q\)-period returns to that implied by a random walk. Under \(H_0\), \(\text{Var}(r_t^{(q)}) = q\,\sigma^2\), so the ratio \(M(q) = \text{Var}(r_t^{(q)}) / (q\,\hat{\sigma}^2)\) should equal 1. Lo and MacKinlay derive a robust test statistic \(Z^*(q)\) that is asymptotically standard normal under heteroskedasticity. The runs test on the signs of returns is non-parametric and makes no distributional assumptions. Finally, the Jarque-Bera test probes normality via skewness \(\hat{S}\) and excess kurtosis \(\hat{K}\):
\[ \text{JB} = \frac{n}{6}\!\left(\hat{S}^2 + \frac{(\hat{K}-3)^2}{4}\right) \;\overset{H_0}{\sim}\; \chi^2(2) \]For each stock we combine the independence-oriented tests into a composite randomness score scaled from 0 to 100. A score near 100 indicates returns statistically indistinguishable from a random walk across all tests; a low score indicates at least one significant rejection. This score is a visualisation convenience and we report the underlying p-values wherever the distinction matters.
Multiple-Comparison Controls
Testing 500 stocks simultaneously invites the multiple-comparisons problem: at a five percent level we expect 25 false rejections by chance alone even if every stock were a perfect random walk. We apply Benjamini-Hochberg (1995) correction to control the false discovery rate at 5%, and report both raw and corrected rejection counts. We also conduct all tests on three non-overlapping four-year subperiods to assess persistence.
Results
Most stocks look like random walks, but a sizeable minority do not
Across the full twelve-year sample, 348 of the 500 stocks (69.6%) failed to reject the random walk null under all independence tests after false-discovery correction. The remaining 152 stocks (30.4%) showed at least one statistically significant departure. Before correction the rejection count was 184, which underscores how much of the naive signal is attributable to multiple testing rather than genuine structure. This is consistent with earlier large-scale tests of the random walk hypothesis, including Lo and MacKinlay (1988) and Fama (1991).
Departures cluster by sector
When we group the 152 non-random stocks by GICS sector, the rejections are clearly uneven. Sectors with smaller average market capitalisation, lower analyst coverage, and thinner liquidity (real estate, utilities, materials) show the highest rejection rates. Highly liquid, heavily traded sectors such as information technology and large-cap financials are closest to the random-walk benchmark. This is consistent with the limits-to-arbitrage framework (Shleifer and Vishny, 1997): predictability survives where the frictions of exploiting it are highest.
Autocorrelation is weak, short-horizon, and inconsistent in sign
For the stocks that do reject, serial correlation is concentrated at short lags, predominantly lags 1 and 2, and decays quickly thereafter. The magnitudes are small. Among rejecting stocks the median first-lag autocorrelation is approximately \(|\hat{\rho}(1)| \approx 0.06\), and even the most extreme names rarely exceed 0.12. By definition, this implies that past returns explain well under one percent of the variance of future returns:
\[ R^2_{\text{pred}} = \hat{\rho}(1)^2 \approx 0.06^2 = 0.36\% \]The sign is not uniform: some names show positive autocorrelation consistent with short-horizon momentum, others show negative first-lag autocorrelation consistent with mean reversion or bid-ask bounce. This heterogeneity itself is informative.
Volatility is highly predictable even where returns are not
The clearest and most robust departure from random-walk behaviour appears not in the returns themselves but in their squared values. Volatility clustering, the tendency of large moves to follow large moves, is present in essentially every stock regardless of whether its returns reject the independence null. Formally, while \(\hat{\rho}(k) \approx 0\) for most stocks, the autocorrelation of squared returns \(\hat{\rho}_{r^2}(k)\) is significantly positive across the universe, a signature of GARCH-type conditional heteroskedasticity (Engle, 1982):
\[ \sigma_t^2 = \omega + \sum_{i=1}^{p} \alpha_i\, \varepsilon_{t-i}^2 + \sum_{j=1}^{q} \beta_j\, \sigma_{t-j}^2 \]The Ljung-Box test applied to squared returns rejects independence for more than 95% of the universe, with overwhelming statistical significance. This is not evidence against weak-form efficiency: the level of returns can be unforecastable while their dispersion is highly forecastable.
| Test (full sample) | Stocks rejecting | Share |
|---|---|---|
| Ljung-Box on returns (lags 1–10) | 131 | 26.2% |
| Variance ratio Z*(q), Lo-MacKinlay | 118 | 23.6% |
| Runs test on return signs | 97 | 19.4% |
| Any independence test (post-FDR) | 152 | 30.4% |
| Ljung-Box on squared returns | 479 | 95.8% |
| Jarque-Bera normality | 500 | 100.0% |
Departures are not stable across time
We re-ran all independence tests on three non-overlapping four-year subperiods and asked how often a stock that rejects in one subperiod also rejects in the next. Of the stocks rejecting in the first subperiod, fewer than 40% also rejected in the second. The overlap across all three subperiods was smaller still. This temporal instability is itself consistent with market efficiency: genuine, stable inefficiencies tend to be arbitraged away, and transient departures are exactly what regime-level churn would produce.
Discussion and Implications
Taken together, our results support a nuanced rather than binary picture. The random walk hypothesis is an excellent approximation for the majority of large-cap U.S. equities, and where it fails, it fails by small and unstable margins, with three important qualifications.
The first is liquidity. The concentration of rejections in less liquid sectors fits the limits-to-arbitrage view: predictability survives precisely where the frictions of exploiting it are highest. Efficiency is better thought of as a property that holds up to transaction costs rather than as an absolute.
The second is the level-versus-magnitude distinction. With \(\hat{\rho}(1)^2 \approx 0.4\%\), the fraction of return variance explained by past returns is negligible. For risk management, option pricing, and position sizing, the random-walk-in-returns result is almost beside the point; what matters is the structure in the second moment, and that structure is strong and pervasive.
The third is the gap between statistical and economic significance. With approximately 3,000 daily observations per stock, our tests have ample power to detect autocorrelations far too small to trade against profitably. A first-lag autocorrelation of 0.06 would be swamped by realistic bid-ask spreads, commissions, and market impact. We make no profitability claims.
Conclusions
We tested 500 large-cap U.S. equities against the random walk hypothesis using four complementary statistical tests over a twelve-year window. About 70% of the stocks are statistically indistinguishable from a random walk after correcting for multiple comparisons. The 30% that reject do so by small margins (\(|\hat{\rho}(1)| \leq 0.12\)), the departures are concentrated at lags 1–2 and in less liquid sectors, and they are unstable across time. Volatility is predictable almost everywhere via GARCH-type dynamics; normality holds nowhere.
The honest summary is that weak-form efficiency is a good first approximation for this universe, the exceptions are real but economically marginal once frictions are considered, and the most reliable structure in equity returns lives in the second moment rather than the first.
Citation
Please cite this work as: