Does the opening one-minute bar know anything about the day that follows? Sixteen years of E-mini Nasdaq‑100 futures — 3,992 complete sessions — mined bar by bar: the 9:30 candle’s range, body, wicks, closing location, volume, and gap, each tested against the rest of the session, with cluster-bootstrap intervals on every number and false-discovery control across thirty-six declared tests.
We ask the oldest question on the tape — does the first candle of the day predict the rest of the day? — and answer it with sixteen years (June 2010–June 2026) of back-adjusted continuous E-mini Nasdaq-100 futures at one-minute resolution: 3,992 complete regular-hours sessions, analyzed using opening/high/low/close/volume alone, with month-block bootstrap confidence intervals on every headline statistic and Benjamini-Hochberg control within three pre-declared test families (36 associations).
(I) The candle’s anatomy is nearly information-free about itself. The median first minute spans just 8.5 bp; its close lands at 0.507 of the bar’s range (95% CI 0.496–0.518) — statistically indistinguishable from dead mid; wicks balance at 26.7%/28.1%. Its own extremes are consumed almost immediately: half of all sessions break the bar’s high and its low within two minutes of 9:31.
(II) Size grades forward — mostly because volatility persists, and measurably not only because of that. Rest-of-day range climbs monotonically from 48 bp after the quietest first minutes to 169 bp after the widest (q < 0.001), positive in every regime cut. But the obvious null is real: a trailing 21-session range alone explains R² = 0.59 of the same outcome at near-unit elasticity, so most of the raw gradient is volatility autocorrelation restated. Controlling for it — properly, with intercepts on both sides and the spread measured in scale-free units — bar0’s width retains independent information: elasticity 0.30 (Frisch–Waugh verified, p < 0.001), +3.5 points of R² on top of the prior, residualized extreme quintiles a factor 1.41 apart in rest-range, within-cut raw gradients positive 6/6. The honest headline shrinks from “a 3.5× forecast” to “a real but minor increment over a strong trailing prior”. Separately, the implementable volume measure carries no range information (p = 0.19): WP‑2026‑01’s inverse volume gradient survives only in its descriptive, future-leaking form — resolved against the folklore.
(III) Direction whispers; location is silent. Days close above the 9:31 mark 56.4% of the time after a green first candle versus 52.2% after a red one — a 4.2-point spread that survives false-discovery control (q = 0.009) but is two orders of magnitude too thin for costs. Large up-gaps (≥50 bp) continue modestly (56.9%, CI 52.7–61.1). Close location within the candle, overnight-range position, and every level-completion test we declared come back null after FDR correction. The first candle predicts how much, barely which way, and not where.
Keywords: opening range · first bar anatomy · close location value · volatility forecasting · E-mini Nasdaq-100 · false discovery rate · month-block bootstrap · OHLCV-only inference
“Does the first candle predict what the day will do?”
— the question, in the trader’s own words
At 9:30 New York time the cash equity market opens and the first one-minute candle prints on millions of screens. Within sixty seconds it has absorbed the intersection of eighteen hours of overnight repricing and the day’s first burst of institutional flow. Traders read it for shape — hammer or shooting star, wide or narrow, big volume or small — and bet the afternoon on the reading. This paper takes the reading seriously enough to grade it.
The question is deliberately naive and deliberately precise at once. “Predict” here means something a statistician can audit: does the conditional distribution of some rest-of-day outcome, given a feature of the first candle measurable at 9:31, differ from the unconditional distribution by more than sampling noise? “The first candle” means the single 9:30 one-minute bar — not the first fifteen minutes, not the opening range, not the morning. “The rest of the day” means everything printed after that bar through the 16:00 close. Every feature we study is computable at 9:31:00 from one closed bar and information already public; every response is computed strictly from later bars. Nothing leaks in either direction, by construction (§3).
The answer is not yes or no. The first candle turns out to be three different oracles wearing one body. As a geometer — its shape, its wicks, where inside itself it closed — it is nearly mute: its close sits at dead mid of its range on average, and its own highs and lows are broken within minutes. As a thermometer — how hot it ran — it is excellent: the width of one minute forecasts the width of six and a half hours with a monotonicity that survives every robustness cut we could construct. As a compass — which way it points — it murmurs: a few percentage points of directional persistence, real under audit, and useless after costs. Most of what practitioners read into the candle’s shape belongs to its size.
A deliberate constraint shapes everything: every statistic herein is computable live, bar by bar, from OHLCV and a timestamp. No order book, no tape detail, no economic calendar — where a claim would need them, we stop and say so (§9). And because this paper screens dozens of feature-response pairs, it practices on itself the suspicion its host engine applies to strategies: three test families declared in advance, cluster-bootstrap uncertainty everywhere, false-discovery rates published beside every p-value, and the complete machine-readable output committed next to the prose (§9, Appendix A).
The predecessor paper mined the window around the open: how much of the day’s volume prints in the first minutes (1.31% in the first minute, 29% in the first hour), when the daily extreme lands (77% within the hour), and how opening intensity relates to session width. Its most-discussed finding was an inversion: sessions ranked by opening relative volume showed declining day range — quiet opens preceded wide days. But its own hostile review caught a sensitivity: normalize volume differently and the gradient reversed. WP‑01 downgraded the finding to “a hypothesis to be tested with trailing-only normalization in the design window.” That hypothesis dies in §6 of this paper, quietly and by the numbers.
What WP‑01 explicitly did not do — and what this paper does:
The single bar. WP‑01’s atoms were windows (first-k-minute shares, OR15). Here the atom is one candle, decomposed into the features a chart reader actually sees: body fraction, upper and lower wicks, close location within the range, direction, absolute size, volume, and the gap it opened with.
Candle-shape lore, formalized. The practitioner literature treats a close near the high as strength and a long upper wick as rejection. Those are testable sentences. We test them as conditional-probability statements against the actual afternoon (§5, §7), with sample sizes attached.
A stricter leakage rule. WP‑01 conditioned on share-of-day volume without flagging that the denominator is unknowable intraday. We keep that statistic — it reproduces beautifully — but mark it descriptive-only wherever it appears, and run every predictive claim through measures computable at 9:31 (§3, defbox). The distinction turns out to carry the paper’s sharpest conclusion.
Multiplicity discipline. WP‑01 reported selected comparisons. This paper declares three test families before results (direction persistence, magnitude grading, location/barriers — 36 tests), applies Benjamini-Hochberg within family, and prints every p and q in Appendix A.2. Negative results are reported as findings, not discarded.
What we are not: a prediction model. There is no classifier, no training set, no learned weights. Every number below is count, divide, bootstrap — the base rates any future model would have to beat visibly.
The series. One dataset of record: the back-adjusted continuous E-mini Nasdaq-100 future at one-minute resolution (data/NQ.cont.1min.2010-2026.csv), June 2010 through June 2026 — 5,431,145 bars parsed with zero defects (no bad rows, no duplicate timestamps, no out-of-order prints), rebuilt deterministically from raw CME Globex recordings by tools/build_continuous.py. Because the series is additively back-adjusted, intraday returns and bp-denominated widths are faithful while absolute price levels are not; every quantity in this paper is therefore a ratio, never a point distance.
Sessions. Regular hours are defined in New York time via the US daylight-saving rule (second Sunday of March through first Sunday of November), timestamp-verified against zone data at six probes per day across the entire span — 35,244 checks, zero disagreements. Each CME session bucket opens 18:00 ET the prior evening; RTH is 9:30–16:00 ET. The exclusion funnel:
| Filter | Sessions | Purpose |
|---|---|---|
| Session buckets opened | 4,132 | every CME day in span |
| − buckets with no RTH print | 11 | full holidays (Globex trades, equities closed) |
| = sessions with RTH bars | 4,121 | matches WP‑01’s raw population exactly |
| − holiday half-days (span < 360 min) | 129 | early closes distort whole-day denominators |
| − broken sessions (missing/bad first or last bar) | 0 | integrity gate |
| = analysis set | 3,992 | every headline number |
Determinism. The miner is pure-stdlib Python, single-pass, fixed-seed; reruns produce byte-identical output (verified by hash). Three pinned dates were independently re-derived by hand arithmetic straight from raw CSV rows — 21 of 21 quantities matched to rounding tolerance. The machine-readable snapshot (first_candle_stats.json) sits beside this paper; every number in the text traces to a path in it, and Appendix A is a human-readable index of that file.
Why conditionals and not a model. A model would answer the trader's question with a number per morning; this paper answers it with a shape — which features move which outcome distributions, by how much, with what uncertainty. The choice is deliberate. Conditional statistics are auditable line by line: any reader can recompute a cell from the published panel file. They are also the correct substrate for the engine's downstream machinery: a strategy hypothesis that cannot be stated as "cell X of table Y beats base Z by more than costs" cannot enter the Trial protocol honestly. And they fail safely — a null conditional prints as a null, where a mis-specified model prints as a confident coefficient on noise.
Before asking what the candle predicts, measure the thing itself. Across 3,992 sessions the opening minute is very small, roughly symmetric, and — this is the sentence candle-shape lore gets wrong — its close lands almost exactly at the middle of its own range.
The size distribution is violently right-skewed: median 8.5 bp, mean 10.6, ninety-ninth percentile past 40 bp. A “big” first minute is not twice normal; it is five times normal, and it happens a few times a year. Body and wicks partition the range almost evenly on average (45/27/28), with the lower wick measurably but marginally heavier — a fingerprint of the open’s downward probe before the first auction clears, and the only anatomical asymmetry that survives scrutiny.
The volume and gap context. Bar0’s volume runs a median 1.00× its trailing twenty-session median (mean 1.07, CI 1.057–1.090) — unremarkable relative to recent opens, however loud it looks against the lunchtime tape. Its share of full-day volume centers near 1.3%, replicating the companion paper’s headline to the second decimal on an independent pipeline. The gap bar0 opens with is typically small — median 21.7 bp, mean 35.6 — but the right tail matters: ten percent of days open more than 82 bp from yesterday’s close, and 860 sessions in this sample clear that bar by half. Overnight, the market travels a median 41 bp between closes, and bar0 opens slightly above the middle of that range on average (mean position 0.53, CI 0.52–0.54) — itself a faint drift signature, since an entry uniform at random into a bounded overnight range would center exactly on 0.50.
The close-location histogram is the paper’s first small surprise. Practitioner lore reads closes near the high as “strength” and expects them to dominate. In fact closes scatter broadly — the interquartile range of clv spans 0.22 to 0.79 — but they balance: the mean is statistically identical to 0.500, and green versus red candles split nearly evenly (1,914 vs 1,978, with 100 zero-body dojis). The first minute has no persistent lean. Whatever signal it carries cannot live in its average shape; it must live in conditional structure. The rest of the paper hunts that conditional structure — and finds it mostly in one place.
Does a green first candle make an up day likelier? Yes — by four points, not fifty. The effect is real by our multiplicity standards, stable across sixteen years, and far too small to pay costs. This is the honest version of the trader’s intuition.
Baseline first: the raw probability that the day closes above the 9:31 mark is 54.4% (CI 52.9–55.9) — above half because sixteen years of NQ drifted upward. Every conditional below should be read against that base, not against 50%.
| Conditioner (at 9:31) | n | P(rest up) | 95% CI | verdict after FDR |
|---|---|---|---|---|
| none (base rate) | 3,992 | 54.4% | 52.9–55.9 | — |
| dir0 = up (green) | 1,914 | 56.4% | 54.3–58.6 | ≠ 50% (q<0.001); spread vs red q=0.009 |
| dir0 = down (red) | 1,978 | 52.2% | 50.0–54.4 | ≠ 50% marginal (q=0.067) |
| clv above mid | 2,004 | 54.9% | 52.9–56.8 | ≠ 50% (q<0.001) |
| clv below mid | 1,987 | 53.9% | 51.7–56.2 | ≠ 50% (q=0.002) |
| up-gap ≥ 50 bp | 464 | 56.9% | 52.7–61.1 | ≠ 50% (q=0.004) |
| down-gap ≤ −50 bp | 396 | 54.3% | 49.2–59.7 | indistinguishable (q=0.143) |
| green bar0 → next hour up | 1,914 | 54.8% | 52.8–56.8 | ≠ 50% (q<0.001) |
| dir0 = 0 (doji, zero-body) — unreported cell, printed for completeness | 100 | 59.0% | small-n | excluded from sign contrasts; n=100 compatible with noise |
Three readings matter. First, the green-red spread of 4.21 points survives false-discovery control and appears in every block — this is a real conditional structure, not noise. It is also two orders of magnitude below what round-turn costs consume: a 4% edge on sign pays nothing after ~1–2 bp of adverse selection per turn at any realistic holding pattern. Second, close location adds nothing beyond direction itself: clv-positive and clv-negative days differ by less than one point (q=0.55 for the difference). The wick-and-body story collapses into the sign story. Third, large up-gaps continue modestly — after gapping up ≥50 bp, the day still closes above the 9:31 mark 56.9% of the time. Read beside WP‑01’s famous “gaps follow at 51.9%, a coin flip,” this is not a contradiction but a reference-point lesson: WP‑01 measured closes against the prior close (gap-fill pulls that to even); measured against the post-open mark, up-gap momentum is faintly positive. Neither reading supports an unconditional fade or an unconditional follow.
One more directional-looking statistic belongs here because it fails instructively. Does the day break bar0’s high before its low after a green candle? Intuition says yes — the close sits nearer the high. The data say the day’s low tends to print early regardless: after green candles the low-side breaks first 51.4% of the time, after red ones 54.7% — replicating WP‑01’s 53.3% low-before-high base and showing that the open’s own direction merely modulates, weakly and not significantly after FDR (q=0.067), a downward-first drift that belongs to the auction, not to the candle’s lean.
The 54.4% base rate deserves one more paragraph, because it will be misread. It is not an edge: it is the sixteen-year drift of the instrument wearing a probability costume. A strategy that simply bought every 9:31 mark and held to the close would have captured it — along with every drawdown of sixteen years of long index exposure, at a Sharpe no better than buy-and-hold minus costs. The same caution applies to the 56.9% up-gap continuation: up-gaps cluster in bull regimes, so the statistic mixes regime selection with any genuine opening momentum. The green-minus-red contrast is the honest object — it nets the common drift out — which is exactly why Family A declares contrasts as its primary tests and treats the raw conditional levels as context.
Strip the candle of its direction and its decoration; keep its width. That single number — how far the market traveled in its first minute — orders the entire day ahead of it, monotonically, in every subsample we can construct. It speaks about volatility, never about sign. Read this section together with §6B: most of what follows is the volatility regime speaking through the open, and the control there measures exactly how much is the candle’s own.
| bar0 range quintile | cut (bp) | n | rest range (mean bp) | 95% CI | MFE | |MAE| |
|---|---|---|---|---|---|---|
| Q1 · narrowest | 0–4.9 | 798 | 48.2 | 45.7–51.1 | 23.0 | 25.2 |
| Q2 | 4.9–7.1 | 798 | 59.7 | 56.3–63.4 | 29.0 | 30.7 |
| Q3 | 7.1–10.2 | 799 | 82.2 | 76.9–87.6 | 40.2 | 42.0 |
| Q4 | 10.2–15.1 | 798 | 114.1 | 105.9–122.5 | 52.9 | 61.2 |
| Q5 · widest | 15.1+ | 799 | 169.1 | 152.9–185.9 | 81.2 | 87.9 |
Note what does not happen: no asymmetry. If wide first minutes preceded trend days, MFE would outrun |MAE| in the upper quintiles; if they preceded crash days, the reverse. Instead both scale within a few percent of each other in every row. Bar0’s width is a pure volatility forecast — it prices the day’s amplitude, and leaves its sign to §5’s four-point whisper.
Here the paper settles WP‑01’s open problem. “Loud open” mixes two different instruments. Loud in price — a wide bar0 — precedes wide days (§above, positive gradient). Loud in volume, WP‑01’s axis, precedes narrow days — their inverse gradient, which their review could not stabilize. Run both through this paper’s discipline:
| Conditioner | Q1 → Q5 rest range (bp) | gradient | Q5−Q1 p | usable at 9:31? |
|---|---|---|---|---|
| bar0 range (bp) | 48 → 169 | positive, monotone | <0.0005 | yes — the bar itself |
| |gap| (bp) | 66 → 150 | positive, monotone | <0.0005 | yes |
| share of DAY volume (descriptive) | 144 → 53 | negative, monotone | <0.0005 | NO — needs the day’s volume |
| rvol0 vs trailing 20-med | 86 → 97 | flat | 0.189 | yes — and says nothing |
The descriptive share-of-day gradient (144→53 bp) replicates WP‑01’s finding to the digit — of course it does; it is the same computation on the same data. But a quantity defined by the day’s total volume cannot be conditioned on at 9:31, which is precisely why WP‑01 fenced their finding behind “test the implementable form.” Tested: the implementable form is flat. Quiet-opening volume does not forecast wide days; wide openings forecast wide days. The folklore survives translation with its sign corrected and its mechanism renamed — it is a volatility statement, not a flow statement.
A reviewer’s question nearly killed §6, and deserved to: if volatility simply persists day to day, then a wide first minute is just today’s copy of yesterday’s news, and the 3.5× gradient is autocorrelation wearing a candle costume. This section runs that null properly. The result is neither victory nor death — it is a precise, smaller finding where a big one stood.
The confound, quantified. Define atr21 as the mean rest-of-day range of the prior 21 valid sessions (trailing; no lookahead). It forecasts today’s rest-of-day range at log-log elasticity 0.96 (cluster-bootstrap p < 0.001) with R² = 0.589 — a near-unit pass-through of the volatility regime. Yesterday’s bar0 predicts today’s bar0 at elasticity 0.69 (n=3,860 adjacent pairs). Volatility clustering on this instrument is exactly as strong as the folklore fears. Any unconditional statement of §6 must be read through that fact.
The increment, isolated. Controlling for atr21, bar0’s elasticity drops from 0.74 to 0.30 but stays decisively non-zero (cluster-bootstrap p < 0.001); the joint fit is R² = 0.624, so the first candle contributes +0.035 R² beyond the prior — a semipartial r of 0.19 against variance, or, in the residual-vs-residual terms D2’s quintile split actually uses, a partial correlation of 0.29. Two independent checks pin this down. Units: ranking sessions by feature-residual bar0 width and measuring response-residuals of log rest-range (both sides purged of atr21, consistent units), the extreme quintiles sit a factor 1.41 apart (Q5−Q1 = +0.341 log points, q < 0.001). Flatness: mean log-atr21 across those quintiles is 4.420 / 4.411 / 4.416 (Q1/Q3/Q5). Stated precisely: this flatness is guaranteed algebraically by in-sample residualization — it is an implementation verification that the residual construction did what it claims, not independent evidence; the contamination question it addresses is closed by construction, not by measurement. The apparent tension between “elasticity fell 60%” and “the raw bp spread barely moved” dissolves as geometry, not contradiction: the effect is multiplicative, so it prints additive basis points in proportion to each day’s own volatility scale — the raw-bp gap is dominated by high-vol tails where a fixed ratio is worth many points, while the elasticity and R² measure the average log-linear slope. Both statements describe the same shrunken-but-real effect; the scale-free number is the one an implementation should size against. Within every four-year block and both eras the raw gradient also stays positive: six cuts, six signs (Table A.5).
| Family-D test | estimate | p / q |
|---|---|---|
| D1 · partial slope, log range0 | log atr21 (FWL) | β=0.3015 | <0.001 / <0.001 |
| D2 · response-residual Q5−Q1 across feature-residual quintiles | +0.341 log pts (×1.41) | <0.001 / <0.001 |
| D2-diagnostic · mean log-atr across those quintiles | 4.42 / 4.41 / 4.42 | flat — identity check of the residualization, not evidence |
| D4 · log atr21 alone (the null, working) — univariate fit | β=0.961 | <0.001 / <0.001 |
| — joint model’s atr21 coefficient, for contrast | β=0.688 | not a standalone test |
| r-statistics from the two fits: semipartial r ≈ 0.19 = √0.035 (variance share); partial r ≈ 0.29 = √(0.035/0.411) (residual-vs-residual, what D2 measures) | ledger-derived | identity-checked |
| D5 · autocorrelation corroboration: yesterday’s bar0 → today’s bar0 (not a placebo — no null was tested) | β=0.689 | <0.001 / <0.001 |
What this does to the finding, stated without varnish. Most of §6’s raw gradient is volatility persistence restated: knowing what yesterday knew is worth far more than knowing today’s first minute. What remains after the control is a genuine but modest same-morning update — elasticity 0.30 over the prior’s 0.69, a factor 1.41 between residualized extremes — which is precisely the quantity any Descendant-A implementation would actually trade: the difference between sizing off the trailing prior alone and sizing off the prior corrected by this morning’s open. Pass 1’s headline (“3.5×, robust everywhere”) is hereby regraded: the multiplier is real but belongs mostly to the regime; the candle’s own contribution is the increment, and the increment survives.
If the close near bar0’s high meant accumulation, and a deep lower wick meant rejection below, then location should organize the afternoon’s races: which barrier falls first, whether prior day’s high or low breaks, how trendily the day travels. We declared twelve such tests. After false-discovery control, none survives. Location is where candle lore goes to die quietly.
The barrier race itself deserves attention before the conditionals do. Around the 9:31 mark at ±50 bp, 2,287 of 3,992 sessions resolve a winner; 1,705 — 42.7% — never travel fifty basis points either way from the morning’s settlement in six and a half hours, and ties are nil at one-minute resolution. Among resolved races the split is 47.9% upward-first (CI 46.3–49.7) — mildly bearish-tilted, the same downward-first drift §5 met from the other side. Against that base, neither where bar0 closed (quartile means 46.8/47.3/48.0/49.4%) nor where it opened in the overnight range (49.7/48.1/46.1%) separates meaningfully. Resolution itself is also location-inert: P(the race resolves at all | close-location quartile) runs 58.7 / 57.0 / 52.7 / 59.6% against a 57.3% base — non-monotone, so a location effect operating through whether the day moves at all is as invisible as one operating through direction. The apparent orderings are exactly what sampling noise produces; the slope tests say so formally (all q ≥ 0.69, Table A.4).
Prior-day levels fare no better as a pair, with one structural exception: when one of the prior day’s extremes breaks first, it is the high 58.3% of the time (n=3,579, CI 56.4–60.1). That looks like information and is mostly bookkeeping: in a sixteen-year bull drift the market simply spends more time nearer yesterday’s high, and the same drift produced §5’s 54.4% base. Treat it as a prior, not a signal.
The survival curves close the case. The first candle’s high falls as a level within minutes (median break time: one minute for both sides); whatever “support/resistance at the opening bar” means, it does not mean the printed extremes hold. Efficiency — the trendiness ratio |net move| / total range — averages 0.474 (CI 0.465–0.482) and refuses to move under any conditioner we declared: not close location, not overnight position, not bar0 direction. Days are born half-trendy on average, and nothing visible at 9:31 shifts the fraction.
If direction whispers and size shouts, does size amplify the whisper? We crossed bar0 direction with opening relative volume (terciles) and with close-location sign — the two interactions practitioner logic most often invokes (“green candle on heavy volume means conviction”).
The conviction story fails: the direction spread is not larger after relatively loud opens (tercile means +3.2/+9.1/+4.5 bp on green; −0.6/−1.3/−4.7 on red — non-monotone, no declared test, reported anyway). Crossing with close-location sign adds nothing beyond direction: green-and-upper-close (+6.4 bp, 56.7% rest-up) versus red-and-lower-close (−1.2 bp, 53.8%) is the §5 effect restated, and the off-diagonal cells shrink toward the base exactly as a single-factor world predicts.
Declared, tested, and rejected at q ≥ 0.05 — stated here so nobody mines them again unawares:
Close location as a completion predictor — twelve Family-C tests, zero survivors (§7). Overnight-range position as a barrier or efficiency predictor — null across the board. Opening volume as a range forecaster in its implementable form — flat (§6). Volume as an amplifier of directional persistence — no structure. Efficiency modulation — nothing visible at 9:31 moves how trendily the day travels. Undeclared and untested because the data ladder forbids them: order-flow imbalance, depth, scheduled-calendar effects. Where this paper is silent, it is silent on purpose.
Regime battery on the central finding. The magnitude gradient (§6) is recomputed under every cut the sample supports. It keeps its sign everywhere — including the low-volatility era, where the whole effect compresses but never flips, and excluding roll weeks, big gaps, the March–April 2020 window, and 400-bp tail days.
| Cut | n | Q1 range | Q5 range | Q5−Q1 | dir spread (pp) |
|---|---|---|---|---|---|
| 2010–2014 | 892 | 44.1 | 65.6 | +21.5 | +5.75 |
| 2014–2018 | 993 | 42.8 | 86.5 | +43.7 | +1.43 |
| 2018–2022 | 997 | 71.8 | 182.4 | +110.6 | +4.67 |
| 2022–2026 | 994 | 102.4 | 195.8 | +93.4 | +5.18 |
| low-vol era (trailing split) | 844 | 39.8 | 53.4 | +13.6 | +4.18 |
| high-vol era | 3,127 | 57.4 | 178.4 | +120.9 | +4.58 |
| exclude |gap|≥50bp | 3,132 | 45.7 | 139.7 | +94.0 | +3.36 |
| exclude roll weeks | 3,670 | 47.2 | 169.2 | +122.0 | +3.54 |
| exclude COVID window | 3,945 | 47.3 | 161.8 | +114.4 | +4.08 |
| exclude |net|>400bp days | 3,985 | 47.5 | 165.4 | +117.9 | +4.11 |
| COVID window only | 47 | 212.9 | 370.3 | +157.4 | +13.85 |
The testing ledger. Thirty-six associations were declared before any table was read: twelve in Family A (direction persistence), twelve in Family B (magnitude grading), twelve in Family C (location and barrier completion). Benjamini-Hochberg within family. Survivors at q < 0.05, counted only where the test is a contrast (the paper’s own philosophy): the green-minus-red spread (q=0.009) and low-first-after-red (q<0.001, a drift replicate). The level-vs-coin-flip cells that also clear q < 0.05 (A1/A4/A5/A7/A12) are drift-contaminated nulls by the paper’s own argument and are demoted to corroboration, not counted as independent survivors; in B, all nine range/MFE/MAE Q5−Q1 contrasts for bar0-range, |gap|, and descriptive volume share (all q<0.001) with rvol0’s three contrasts null (q≥0.22); in C, none — twelve of twelve null. A fourth family, D (four residualized magnitude tests), was added in pass 2 after pass 1 was read, prompted by external review that correctly identified the untested trailing-volatility null (§6B). It is disclosed as post-hoc everywhere it appears; its survivors confirm a shrunken effect, not a pre-declared one. Total declared associations now 40: 36 pre-registered, 4 post-hoc. Full p/q ledgers: Table A.2.
Raw-contract cross-check. The same pipeline rerun on raw single contracts NQZ5 (67 sessions) and NQH6 (50) against the continuous series restricted to the identical calendar window. Core statistics agree within small-sample noise; no stitching signature appears (Table A.6, Figure 10). Conclusions are limited to the overlap window by design.
Reading the ledger honestly. Three properties of these results deserve explicit statement. First, the survivors are coherent: direction tests survive in Family A together; magnitude tests survive in Family B as a block with rvol0's three nulls standing beside them like controls; Family C's uniform null is itself informative — had even two or three location cells squeaked under q=0.05 among twelve correlated tests, the natural reading would have been correlation, not discovery. Second, the effect sizes that survive are small by construction of the audit: FDR control within family prices in the fact that we went looking twelve times per family. Third, nothing here was selected for beauty — the paper's central figure is a monotone bar chart precisely because monotonicity was the pre-declared robustness standard, not because aesthetic search produced it.
Threats, named. No event controls: FOMC and CPI mornings inflate bar0’s range mechanically, and some of Q5’s width belongs to scheduled news — an economic calendar is not OHLCV data, the ladder forbids inventing one, and the effect is therefore carried openly rather than adjusted away. Volume semantics: Globex contract counts exclude off-exchange flow; all ratios are internal to the series. Dependence: month-block bootstraps absorb volatility clustering but are not a full GARCH. Back-adjustment: ratios only, roll weeks excluded as a cut, raw contracts checked. Volatility-clustering confound: pass 1 shipped without the obvious control; pass 2 adds it (§6B) and regrades the central finding in consequence — the audit mechanism worked. Selection: this paper screened 36 declared pairs and reports all 36; en route, dozens of undeclared exploratory statistics were computed and are preserved verbatim in the snapshot rather than laundered into the text — the file drawer is a file.
Return to the trader’s question with the evidence assembled. Does the first candle predict what the day will do? Three answers, graded by how much survives an audit.
It predicts how much — with the emphasis on how much of yesterday. The width of one minute orders the width of the day across a monotone gradient stable in every era and cut — but §6B shows most of that ordering is volatility persistence, a prior any calendar already holds. The candle’s own contribution is a real, small increment: elasticity 0.30 over the prior’s 0.69 pass-through, +3.5 R² points, residualized extremes a factor 1.41 apart. Bar0 is a same-morning update to the volatility forecast, not a source of it — a better anemometer reading than yesterday’s, not a prophecy about the storm.
It barely predicts which way. Four points of green-red persistence, a 56.9% up-gap continuation, positive in all eleven regime cuts, surviving false-discovery control — and worth approximately nothing after costs. The correct use of §5 is as a filter input and a prior, never as a thesis.
It does not predict where. Every location test — close position, wick asymmetry, overnight placement, level races — came back null under multiplicity control, and the candle’s own extremes are broken within minutes. The chart-reader’s geometry is decoration on the thermometer.
Gated descendants — what this measurement licenses next, each with a kill criterion, none yet a strategy:
Use bar0’s range (and |gap| as a second input) to scale bracket width, stop distance, and size for a separate mechanical core — never to choose direction. The §6 table is the prior; the design window is the judge.
The 4.2-point green-red spread and 56.9% up-gap continuation enter exclusively as vetoes or size tilts on some other entry — tested as interaction terms, with the prior that they contribute ≤4 pp.
Location-based entries (fade-the-wick, buy-the-strong-close) carry this paper’s null result as their burden of proof: any future test must beat Family-C’s twelve documented failures first. Recorded so the grave is marked.
Standing warning: everything above is descriptive statistics on historical futures data. No configuration herein has passed a Trial, a walk-forward, or a Final Judgment. Nothing herein is investment advice.
Conclusion. The first candle is a thermometer that traders keep consulting as a compass and a map. Sixteen years of one-minute bars say: consult it for heat — its width genuinely orders the day’s amplitude, in every regime we can construct; listen, at four-percentage-point volume, to its direction if you listen at all; and stop reading its geometry, whose completions are consumed within minutes and whose placements predict nothing that survives an honest ledger. How much, barely which way, never where. The descendants of that correction are gated above; the machinery that judges them is already running.
| Feature | mean | month-block 95% CI | median | p10 | p90 | n |
|---|---|---|---|---|---|---|
| range0 (bp) | 10.57 | 9.75–11.43 | 8.53 | 3.88 | 19.78 | 3,992 |
| |ret0| (bp) | 5.18 | 4.78–5.65 | 3.49 | 0.59 | 11.89 | 3,992 |
| body_frac | 0.452 | 0.443–0.461 | 0.450 | 0.095 | 0.808 | 3,991 |
| clv | 0.507 | 0.496–0.518 | 0.500 | 0.082 | 0.929 | 3,991 |
| uwick_frac | 0.267 | 0.261–0.274 | 0.222 | 0.039 | 0.568 | 3,991 |
| lwick_frac | 0.281 | 0.274–0.289 | 0.242 | 0.037 | 0.600 | 3,991 |
| rvol0 (×) | 1.074 | 1.057–1.090 | 1.003 | 0.689 | 1.534 | 3,972 |
| vol_share0 (% of day) | 1.306 | 1.270–1.343 | 1.201 | 0.784 | 1.928 | 3,992 |
| |gap| (bp) | 35.55 | 32.31–39.25 | 21.73 | 3.52 | 82.08 | 3,991 |
| on-range (bp) | 55.28 | 50.75–60.32 | 41.43 | 17.89 | 108.83 | 3,991 |
| Test | p | q (BH) | reading |
|---|---|---|---|
| A1 rest_up | dir0=up = 50% | <0.001 | <0.001 | level-vs-coin-flip (drift-contaminated); contrast A3 is primary |
| A2 rest_up | dir0=dn = 50% | 0.050 | 0.067 | marginal |
| A3 rest_up up−dn | 0.006 | 0.009 | +4.21 pp spread |
| A4/A5 rest_up | clv± | <0.001/0.001 | <0.001/0.002 | level-vs-coin-flip (drift-contaminated); superseded by contrast A6-null & A3 |
| A6 rest_up clv+ − clv− | 0.553 | 0.553 | null beyond dir0 |
| A7 rest_up | up-gap≥50 = 50% | 0.002 | 0.004 | level-vs-coin-flip (drift-mixed); gap-bin deconfound listed as open work |
| A8 rest_up | dn-gap≤-50 = 50% | 0.119 | 0.143 | null |
| A9/A10 hi-first | dir± | 0.232/<0.001 | 0.253/<0.001 | drift replicate |
| A11 hi-first up−dn | 0.048 | 0.067 | not significant after FDR |
| A12 next-hour-up | green | <0.001 | <0.001 | level-vs-coin-flip (demoted); no contrast declared |
| B range0 × {range,mfe,mae} | <0.001 ×3 | <0.001 | central finding |
| B abs_gap × {range,mfe,mae} | <0.001 ×3 | <0.001 | secondary |
| B vol_share0 × {range,mfe,mae} | <0.001 ×3 | <0.001 | descriptive-only axis |
| B rvol0 × {range,mfe,mae} | 0.19–0.24 | ≥0.22 | null — implementable volume says nothing |
| C clv × {race50,touchPH,eff} slopes+contrasts | 0.20–0.34 | ≥0.69 | all null |
| C on_pos × {race50,touchPH,eff} slopes+contrasts | 0.49–1.00 | ≥0.75 | all null |
| Quintile | bar0 range | |gap| | vol_share0* | rvol0 |
|---|---|---|---|---|
| Q1 | 48.2 (45.7–51.1) | 66.2 | 144.4 | 85.6 |
| Q2 | 59.7 (56.3–63.4) | 73.2 | 108.1 | 94.0 |
| Q3 | 82.2 (76.9–87.6) | 86.3 | 93.8 | 98.6 |
| Q4 | 114.1 (105.9–122.5) | 97.8 | 74.5 | 98.9 |
| Q5 | 169.1 (152.9–185.9) | 149.9 | 52.8 | 97.0 |
| Test | estimate | p | q (BH) |
|---|---|---|---|
| D1 partial slope, log range0 | log atr21 (FWL-verified) | β=0.3015 | <0.001 | <0.001 |
| D2 response-residual Q5−Q1 across feature-residual quintiles | +0.341 log pts (×1.41) | <0.001 | <0.001 |
| D2-diagnostic mean log-atr by quintile (Q1/Q3/Q5) | 4.420 / 4.411 / 4.416 — flat | — | identity check of residualization, not independent evidence |
| D4 log atr21 alone — univariate fit (the null, working) | β=0.961, R²=0.589 | <0.001 | <0.001 |
| — joint model’s atr21 coefficient (not standalone) | β=0.688 | — | context only |
| r-statistics: semipartial r ≈ 0.19 = √0.035 (variance share) · partial r ≈ 0.29 = √(0.035/0.411) (residual-vs-residual) | ledger-derived | — | identity-checked |
| D5 autocorrelation corroboration: prev bar0 → bar0 (not a placebo) | β=0.689, n=3,860 | <0.001 | <0.001 |
| Cell | n | P(+50 first | resolved) | 95% CI | efficiency |
|---|---|---|---|---|
| clv_c [−1,−.5) | 1,057 | 46.8% | 42.9–50.6 | 0.483 |
| clv_c [−.5,0) | 868 | 47.3% | 42.9–51.6 | 0.474 |
| clv_c [0,.5) | 879 | 48.0% | 43.5–52.6 | 0.463 |
| clv_c [.5,1] | 1,187 | 49.4% | 45.9–52.9 | 0.472 |
| on_pos T1 (low) | 1,329 | 49.7% | 46.5–52.9 | 0.478 |
| on_pos T2 | 1,329 | 48.1% | 44.9–51.3 | 0.471 |
| on_pos T3 (high) | 1,331 | 46.1% | 42.9–49.3 | 0.466 |
| base (resolved races) | 2,287 | 47.9% | 46.3–49.7 | 0.474 |
| Cut | point est. | note |
|---|---|---|
| Full sample | +120.9 | Q5 CI 152.9–185.9 vs Q1 CI 45.7–51.1: non-overlapping |
| Per-block / era / exclusion cuts | +13.6 … +157.4 | all positive; see §9 table and snapshot regimes block |
| Statistic | continuous (n=199) | NQZ5 raw (n=67) | NQH6 raw (n=50) |
|---|---|---|---|
| bar0 volume share, % | 1.43 | 1.22 | 1.60 |
| bar0 range, mean bp | 19.4 | 16.4 | 22.1 |
| rest-of-day range, median bp | 117.2 | 111.0 | 113.1 |
| P(rest up | green bar0) | 61.8% | 65.4% | 50.0% |
| P(+50bp first | resolved) | 51.7% | 50.8% | 47.7% |
| Q5−Q1 gradient, bp | +66.2 | +95.4 | +35.4 |
| Horizon | bar0 high unbroken | bar0 low unbroken | P(post-open extreme placed) | P(±50bp race resolved by) |
|---|---|---|---|---|
| 15 min | 21.0% | 23.2% | 32.8% | — |
| 30 min | 16.5% | 18.6% | 45.9% | — |
| 60 min | 13.3% | 14.8% | 63.1% | 28.0% |
| 120 min | 11.3% | 12.5% | 78.3% | 38.8% |
| never (full day) | 8.0% | 9.1% | 96.6%* | 57.3% |
| Artifact | Content |
|---|---|
| mine_first_candle.py | streaming miner: sessions, features, responses, bootstraps, FDR, TSVs |
| spot_check.py | independent hand-arithmetic verification, 3 pinned dates, 21/21 OK |
| mine_crosscheck_fc.py | raw-contract replication (Table A.6) |
| gen_figs.py → figs.json | deterministic SVG generation from the snapshot |
| first_candle_stats.json | full machine-readable output — every number’s home |
| sessions_panel.tsv | per-session panel: 39 columns × 3,992 rows — the paper recomputable from one file |
| data/NQ.cont.1min.2010-2026.csv | series of record; 5,431,145 bars, rebuildable from raw MDP 3 archives |
Every formula below is pinned verbatim; the miner implements exactly these strings. Definitions were fixed before results were read; any change requires a new revision line here.
| Term | Definition (exact) |
|---|---|
| RTH | ET minute-of-day [570, 960); ET derived from UTC by the US DST rule (2007+ form), timestamp-verified against zone data at six probes/day across the span (35,244 checks, zero mismatches) |
| session bucket | CME day: opens 18:00 ET prior calendar day, closes 17:00 ET |
| half_day | RTH span < 360 minutes (holiday early closes); excluded from every headline cell |
| broken | full-span session missing bar0 within 9:28–9:32 ET, or last bar outside 15:58–16:01 ET, or zero RTH volume |
| flat_open | H₀ = L₀; body/wick/clv undefined; excluded from anatomy conditionals, counted (n=1) |
| features | see defbox in §3 — functions of bar0 and pre-9:31 context only |
| responses | functions of bars [1, N) only; ref = C₀; END = final RTH close |
| barrier race X | first touch of ref·(1±X/10⁴) scanning bars [1,N); same-bar double touch = tie, excluded from denominators (zero occurred at 1-min resolution for X=50) |
| rvol0 | V₀ / median(V₀ over prior 20 valid sessions); NULL until history accrues (first 20 sessions) |
| era | low/high by trailing 21-valid-session mean day range vs expanding historical median — no lookahead |
| roll_week | ISO week containing the quarterly roll Monday (third Friday of Mar/Jun/Sep/Dec minus 4 days, build_continuous.py rule) |
| uncertainty | month-block bootstrap 95% CI, 2,000 resamples, xorshift64 seed 20260823; p-values centered two-sided cluster bootstrap |
| multiplicity | BH-FDR within Families A/B/C, declared before results; q-values published for all 36 tests |
Pass 1 (initial release, August 2026). First release of WP‑2026‑02. Families, floors, and definitions frozen pre-analysis. Companion lineage: extends WP‑2026‑01; resolves its open normalization question in §6; replicates its first-minute share (1.306% vs 1.31%) and low-before-high drift base (§5) on an independent pipeline whose session counts reconcile exactly (3,992 / 129 / 4,121).
Pass 2 (August 2026). Adds Family D — four residualized magnitude tests plus a six-cut within-block sign battery — responding to external review that identified §6’s untested trailing-volatility null. Declared post-hoc; disclosed at every appearance. Result: the raw gradient survives the control shrunken (elasticity 0.74 → 0.22 controlling atr21; +0.025 R² given the prior; residualized Q5−Q1 +123 bp; within-cut signs 6/6), and the abstract, §6’s framing, §10, and Descendant-A’s kill criterion are regraded accordingly. No pass 1 claim required retraction beyond this regrading; no other number changed (bit-stable pipeline rerun).
Pass 2b (August 2026). Two computational defects in pass 2’s Family D, caught by follow-up review reconciling D1 against D2: (i) the joint OLS was fit without an intercept, biasing all coefficients and R²s — corrected (βrange0 0.224→0.3015 with the Frisch–Waugh cross-check now exact; joint R² 0.614→0.624); (ii) D2 compared response-residual-free bp means against feature-residual ranks, mixing units — replaced by the consistent-units form (+0.341 log points, ×1.41) plus the requested diagnostic that mean trailing-atr is flat across residual quintiles (4.420/4.411/4.416), which shows the raw-bp spread was not confound-contaminated and resolves the apparent elasticity/spread tension as scale geometry of a multiplicative effect. The direction of every pass 2 conclusion survives both fixes; the magnitudes are as stated here.
Pass 3 (August 2026). Two labeling defects in the pass 2b table, caught by external review: (i) the D4 row quoted the joint model’s atr21 coefficient (0.688) under an “alone” label — the univariate estimate is β=0.961, now stored in the snapshot itself so the number can no longer live in prose alone; (ii) a semipartial r (√ΔR² ≈ 0.19, the variance-share view) was labeled and used as a partial r. The two are now stated side by side everywhere they appear, the residual-vs-residual partial r ≈ 0.29 is what D2’s quintile method measures, and Descendant A’s kill-criterion threshold is keyed to the partial (0.29), not the semipartial. Also: the within-cut gradient battery is removed from the Family-D table (it is §9’s unresidualized robustness sweep, not a net-of-regime test) and D5 is relabeled autocorrelation corroboration rather than a placebo. Structurally: every fit is now registered once in a role-tagging citation ledger (tools/stat_ledger.py) whose cite() rejects wrong-role pulls; partial/semipartial are derived by formula with identity assertions; an independent numpy/sklearn recompute of every published fit must agree to 1e-6 before release (verify_stats.py); and the rendered document passes a role-word linter before it can be written.
Pass 4 (August 2026, presentation & disclosure). No statistic changed. External review caught seven presentation-level items; six adjudicated and fixed here, one (Fig. 3 "transposition") refuted after checking figure DOM order against the caption. Fixed: the unreported doji cell printed (n=100, 59.0% rest-up, above base, within small-n noise, excluded from contrasts); Family A's survivor count reconciled with the paper's own contrast-first philosophy (A1/A4/A5/A7/A12 demoted to drift-contaminated corroboration; survivors = A3 + A10-replicate); the bootstrap censoring floor stated explicitly ("p < 0.0005" = zero exceedances in 2,000 draws); the D2 flatness diagnostic relabeled an implementation verification rather than evidence (it is algebraically guaranteed in-sample); P(race resolves | close-location) disclosed per quartile (58.7/57.0/52.7/59.6%, non-monotone). Refuted: Fig. 3 was never transposed — its bars read in block order 5.8/1.5/4.7/5.2 pp, matching the text within bar-label rounding; exact spreads are now annotated under each pair. Review items requiring NEW analysis (HAR-RV baseline, era-split residuals, gap-bin deconfound, ES replication, computed cost table) are declined for this revision and recorded as open work, not silently ignored.
Practice traditions cited, novelty not claimed over them. Selected anchors: