Skip to content
Candle Prediction

Research · measured August 2026

Candlestick patterns: a measured null.

The pattern shape adds nothing. Measured over 1.7 million bars, then attacked by three independent reviews. Every apparent edge dissolved. Nothing survives at any horizon in any asset class.

We sell a product that reads candlestick charts. This is the result that says the most famous thing about candlestick charts does not work. Both of those are true, and the second one is why the first is worth trusting.

What was measured

Eleven detectors — the five single-bar flags already in the model (engulfing ×2, doji, hammer, shooting star) plus six multi-bar ones added for this work (morning and evening star, harami ×2, three white soldiers and three black crows).

Training rows only: 1,096,620 crypto 1h rows, 273,681 4h and 44,912 daily across 22 symbols, plus 294,222 equity daily rows across 188. Never validation, test, or the final-gate window — these numbers are for readers, and computing them on evaluation data would contaminate every future evaluation.

Unconditional up-rate at ten bars: crypto 49.6%, equities 56.1%. That equity figure is drift, and it is why no conditional rate is ever shown here without its base rate beside it. A “57% up” pattern on equities isbelow baseline.

How it dissolved

The first pass controlled for the last candle’s directionand found what looked like real effects. Three reviews took it apart. The control was one variable too shallow: these patterns are defined by thesize of the last one to three bars, not just their colour.

Edge in percentage points for each pattern under three successively deeper controls, crypto 1h at one bar ahead.
crypto 1h, h1direction only+ magnitude+ close-position
bullish harami+3.49+0.08—
bearish harami−3.51+0.28—
three black crows+7.51+3.47—
hammer−1.97−1.71+0.06
shooting star+1.93+2.44+0.15
Bullish harami falls from +3.49 points to +0.08. Hammer from −1.97 to +0.06. The apparent edge was the control’s shape, not the candle’s.

The placebo settles the mechanism

“Close in the top 20% of the bar’s range” — no wick rule, no body rule, no shape at all — scores −1.45 points on 244,287 bars, against hammer’s −1.68 on 95,728. “Close in the bottom 20%” scores +1.61 against shooting star’s +2.15. A hammer’s close is mechanically near its high. The measurement was reading close-position-in-bar, which is one arithmetic line, not a candlestick.

Equities collapse on the resampling unit alone

Forward returns correlate +0.41 across symbols on the same day, and clustering the bootstrap by symbol cannot absorb that at any number of symbols — 188 does not help. Switching to 30-day calendar blocks makes every equity effect span zero. Hammer at ten bars went from [−1.38, −0.13] to [−2.35, +0.82].

Multiplicity, checked separately, was not the problem. With roughly 114 tests the expected maximum |z| under the null is 2.79, and the raw effects reached 6 to 10. They were real correlations. They were just not correlations with the shape.

What we ship instead

Counts and rates, each beside its base rate. Every edge quantity — the matched control, the bootstrap intervals, the FDR flags — is stripped before the data ever leaves the pipeline, and a test fails if one comes back. Patterns below 150 occurrences publish no rate at all, only a count and “too rare in our data to give a reliable percentage”. Crypto daily three-white-soldiers is “66.7% up” on n=12, with an interval 58 points wide. That is an anecdote with a decimal point.

Limitations

Questions this answers

Do candlestick patterns work?
Measured over 1.7 million bars across 22 crypto symbols and 188 equities, the pattern shape adds nothing. Every apparent edge dissolved once the control included the size of the bars and where they closed within their own range. Nothing survives at any horizon in any asset class.
Why do candlestick patterns appear to work in backtests?
Because the usual control is one variable too shallow. Controlling only for the last candle’s direction leaves the pattern standing in for its size and its close position. A placebo rule — “close in the top 20% of the bar’s range”, with no wick rule, no body rule and no shape at all — scores −1.45 points against hammer’s −1.68. The measurement was reading close-position-in-bar, which is one arithmetic line, not a candlestick.
Are candlestick patterns statistically significant?
The raw effects were real correlations — |z| of 6 to 10, well past the 2.79 expected as the maximum under the null across ~114 tests, so multiplicity was not the problem. They were simply not correlations with the shape. On equities they also collapse on the resampling unit alone: forward returns correlate +0.41 across symbols on the same day, and switching to 30-day calendar blocks makes every equity effect span zero.
What does Candle Prediction publish about patterns?
Counts and rates only, each beside its base rate, and nothing at all below 150 occurrences. Every edge quantity is stripped at export and a test fails if one comes back. The app never claims a pattern predicts anything, because in this data it does not.

This is the fourth time a plausible edge in this project turned out to be something else. The pattern is consistent enough to be worth naming: an effect that only exists under one specific control is a property of the control, not the market.

How confidence is measured·What Candle Prediction does