While building chartTrigger's strategy templates we replayed each one, unchanged, across eight markets in four asset classes at typical retail spreads. The most useful results were not about which strategy won. They were three mistakes that made nearly everything look broken for reasons that had nothing to do with the strategy.
How we tested, and the limits of it
Every template was run unchanged on FX majors, gold, the S&P 500, Bitcoin and Ethereum, on an earlier half of history and a later half, with results expressed in R-multiples so markets are comparable. Parameters were each technique's published defaults — a 200-period trend filter, a 14-period ATR — and were never swept until a backtest looked good, because that is curve-fitting and a performance claim by another route. A template only shipped if it behaved consistently across both halves.
Finding one: a spread guard in pips never passes on gold, indices or crypto
Every original template carried a spread condition like "spread under 2 pips". On EURUSD that is sensible. On gold, a normal spread expressed in the same units is around thirty-five. On an index or Bitcoin it is larger still. The result: the guard was always false, so those templates simply never traded on half the markets, and a backtest of a strategy that never fires shows nothing and looks like a bug.
The fix is to express the limit as a fraction of the market's own volatility — spread under a share of ATR(14) on closed bars. It reads as roughly a pip on EURUSD H1, under a dollar on gold, and tens of dollars on Bitcoin, which is what the guard is for: refusing a fill that costs a meaningful slice of a typical bar. It is now the default form.
Finding two: stops in pips are equally market-bound
A 12-pip stop is a sensible distance on a currency pair. On Bitcoin it is about twelve cents. Size a stop from volatility (a multiple of ATR) and it means the same thing everywhere: a stop one ATR away is "one typical bar's movement" on every instrument. If a strategy only works with a hand-tuned pip distance per market, you have not found a strategy, you have found a number.
Finding three: on M5 to M30, spread won
Every strategy in our original library on M5 to M30 lost to spread on every market we tested. This is mechanical, not mysterious. A strategy's average winner on a short timeframe is small, so a roughly fixed round-trip cost is a large fraction of it. Move the same idea to H4 or D1 and the average move grows while the cost does not, so the cost shrinks to a footnote. It is why the strategy templates now run on H4 and D1.
| Short timeframe (M5 to M30) | Longer timeframe (H4 to D1) | |
|---|---|---|
| Typical move you capture | Small | Larger |
| Spread as a share of that move | Large and visible | Small |
| Sensitivity to the spread you assume | Very high | Modest |
| Sample size per year | Many trades | Few trades |
The trade-off is real: longer timeframes give you fewer trades, so any conclusion rests on a thin sample. That cuts both ways and is the reason to hold conclusions loosely.
The spread you assume decides the answer
A backtest that charges one sampled tick as the spread is dangerous in both directions. Press Run on a weekend, at rollover or on a thin demo and you charge a spread several times the typical one, and everything loses. Use a default of "one pip" on gold and an index and costs are nearly free, which flatters the strategy as surely. So the simulator now charges the median of the broker's own historical per-bar spread on a connected account, or a stated typical retail spread on a public feed, and every result prints which source it used beside the figure.
The practical rule for any backtest you read, including your own: find the spread assumption first. If it is not stated, the result is a number with no meaning.
What to do with this
- Express spread limits as a fraction of ATR, not in pips.
- Size stops from volatility, not from a pip count that only fits one market.
- Be suspicious of any short-timeframe result that does not show its spread assumption, and re-run it with a pessimistic one.
- If a rule never fires, check the guards before you blame the idea — the Proving Ground reports which condition blocked it and how often.
- Then forward test it on paper, because history is not the last word on cost either.
Common questions
Why do scalping strategies fail in backtests?
Usually because a roughly fixed round-trip cost, the spread, is a large share of the small move a short-timeframe strategy captures. The same idea on a longer timeframe faces the same cost against a bigger move.
Why does my spread filter never let a trade through on gold or crypto?
A limit written in pips is only meaningful on the instrument it was written for. A normal gold or Bitcoin spread is dozens of pips in those units, so a currency-pair limit is always false. Express the limit relative to ATR instead.
What spread should I assume in a backtest?
A realistic typical one for the instrument and session, ideally the median of historical per-bar spread, not a single sampled tick. Whatever you use, the result is only meaningful when the assumption is stated.
Does this prove which strategy works?
No. The samples were small, and nothing here is evidence of an edge. The findings are about failure modes in testing, which apply regardless of strategy.
Backtest with the assumptions showing
Every Proving Ground run prints its spread, slippage and source next to the result, and tells you what blocked a rule that never fired.
Create a free accountRead next
Risk note. This article is educational material about how chartTrigger works. It is not investment advice, not a recommendation to trade any instrument, and nothing here forecasts results. Trading leveraged products carries a high risk of loss. Any historical simulation referred to is exactly that — a run over past bars under stated spread and slippage assumptions, not an indication of future performance.