A scanner flags a pattern and shows a 90% confidence score. It looks like a verdict: nine times out of ten, the setup works. That reading is almost always wrong.
A confidence score is a number computed by a rule about how well the shape matches. It is not a measured probability that the trade will succeed, and treating it as one is a fast way to size a position on a number that means something else.
A confidence score describes how well a shape matches the scanner's rule, not how likely the outcome is. It says nothing about the actual win rate, the market context, or the sample it came from. Treating a confidence score as a probability invites oversizing and ignores everything the number does not measure.
What a Confidence Score Actually Measures
A confidence score is usually a similarity number: how closely the observed shape fits the geometric criteria of the rule. A 90% score means the shape is very close to the template the rule describes.
It does not mean the pattern succeeds 90% of the time. It is a measure of shape fit, computed by the same rule that found the pattern in the first place.
Why It Is Not a Probability
A probability would come from a measured frequency: out of N past instances, how often did the outcome follow? A confidence score has no such history behind it unless it was explicitly built and validated that way.
Even when a score is derived from historical data, it is only as good as that sample. A score trained on one market or regime does not automatically transfer to another, and a score from a small sample is mostly a guess.
| Confidence tier (scanner label) | Signals in tier | Actually followed by the predicted move | Measured win rate |
|---|---|---|---|
| 90% and above | 25 | 11 | 44.0% |
| 70%-89% | 40 | 17 | 42.5% |
| 50%-69% | 35 | 15 | 42.9% |
The measured win rate is nearly identical across all three confidence tiers, sitting between 42% and 44% — even though the score jumps from the 50s into the 90s, the actual outcome barely moves, which is exactly what you would expect if the number measures something other than true win rate.
The batch of signals labeled 90% confidence had a measured win rate of 44% — not 90%, and barely higher than the batch labeled just above 50%, which came in at 42.9%. All three tiers land on nearly the same number, which is what you would see if the 'confidence' score measures how closely the shape fits the rule's geometry, not the odds of it actually working.
| Dimension | Confidence score | Measured probability |
|---|---|---|
| What it is | Shape similarity to a rule | Observed frequency of an outcome |
| Where it comes from | The rule that found the pattern | A validated sample of past instances |
| What it tells you | How textbook the shape is | How often the outcome followed |
| Risk if misread | Oversizing on a beauty score | Acting on an unvalidated number |
The Danger of Reading It as Probability
The practical danger is sizing. A 90% score feels like a high-probability trade, so the position gets bigger and the risk per trade grows. If the number was actually a shape-fit score, the size decision was based on beauty, not evidence.
The same audit applies to a confidence score as to any flag: ask what rule produced it, whether the sample supports it, and what context the number cannot see.
The number does not upgrade the flag into a conclusion. It still needs the same rule, sensitivity, context, and manual checks as any other candidate.
How to Use a Confidence Score Honestly
Treat the score as a filter, not a verdict. Use it to rank candidates or to prefer textbook shapes, then run the same audit you would run on any flag.
If the software actually documents a measured historical win rate, check where the sample came from and how large it is before giving the number any weight. An undocumented 'confidence' is a design choice, not evidence.
Frequently Asked Questions
Can a confidence score ever be a probability?
Only if it is explicitly built from a validated sample of outcomes and documented as such. A plain similarity score is not.
Should I ignore the number entirely?
No. Use it to rank and filter candidates, just do not let it set your position size by itself.
How do I check what a score means?
Ask what rule produced it and what sample supports it. If the software cannot answer, treat the score as a design choice, not evidence.


