Why Your Trading Strategy's Backtest Might Be Lying to You

ยท
Listen to this article~6 min
Why Your Trading Strategy's Backtest Might Be Lying to You

A profitable backtest doesn't prove your trading strategy adds value. Learn how fair benchmarks and a benchmark ladder can reveal whether your edge is real or just market exposure.

A trading strategy can produce an attractive equity curve and still add very little value. It may simply hold an asset that rose during the test, take more risk than the comparison portfolio, remain invested longer, or concentrate its exposure in a favorable market regime. This is why a profitable backtest is not enough. The more useful question is whether the strategy performed better than a fair alternative under comparable conditions. A benchmark gives that question structure. However, selecting the wrong benchmark can make an ordinary strategy look exceptional. Comparing a leveraged trend-following system with cash is not fair. Comparing a strategy that trades only during high-volatility periods with continuous buy-and-hold is often incomplete. Even comparing total returns can be misleading when the strategy and benchmark carry different exposure, volatility, drawdown, or trading costs. Fair benchmarking does not mean finding one universal baseline. It means building a set of increasingly demanding comparisons that isolate what the strategy is actually contributing. ### A Profitable Backtest Does Not Prove Added Value Imagine a long-only strategy tested on an equity index during a sustained bull market. The strategy earns 60%, which initially appears impressive. Over the same period, however, simply holding the index earns 85% with fewer decisions, lower turnover, and lower costs. The strategy made money, but that does not mean its entry and exit logic added value. Now suppose the strategy earned 70% while buy-and-hold earned 85%, but the strategy was invested only half of the time and experienced substantially less volatility. The conclusion becomes less obvious. Its total return was lower, yet it may have used capital more efficiently or reduced risk meaningfully. A benchmark comparison is therefore not a single return-versus-return calculation. It is an attempt to answer several separate questions: - Did the strategy outperform doing nothing? - Did it outperform passive exposure to the same market? - Did it outperform after matching time in the market? - Did it outperform after matching volatility or downside risk? - Did its complex rules beat a much simpler strategy? - Did performance remain after accounting for common market factors? - Did any advantage survive costs, unseen data, and changing regimes? Each comparison removes a different explanation for the backtest. The remaining performance becomes harder to dismiss as passive exposure, leverage, favorable timing, or unnecessary complexity. ### What Makes a Trading Benchmark Fair? A fair benchmark should represent a realistic alternative use of the same capital and should resemble the strategy along the dimensions that materially affect performance. Depending on the strategy, those dimensions may include: - Tradable instrument or investment universe - Long, short, or market-neutral exposure - Average capital employed - Time in the market - Volatility and leverage - Holding period and rebalance frequency - Transaction costs and financing costs - Liquidity and execution constraints - Market regime The goal is not to make the strategy and benchmark identical. If they were identical, there would be nothing to test. The goal is to neutralize the obvious differences so that the strategy's decisions become the main remaining variable. A useful benchmark asks: What simpler or more passive alternative could have produced a similar result with comparable exposure and risk? ### Use a Benchmark Ladder Instead of One Baseline No single benchmark can answer every question. A better approach is to use a benchmark ladder, beginning with a low hurdle and becoming progressively more demanding. | Benchmark | What It Tests | What It Can Reveal | |---|---|---| | Cash or risk-free return | Whether taking risk was rewarded | A strategy may earn a positive return without adequately compensating for risk | | Buy-and-hold | Whether active timing improved passive market exposure | The strategy may be adding value through timing or simply riding the market | | Volatility-matched index | Whether returns are adequate for the risk taken | The strategy may be taking on more risk than necessary | | Factor-based model | Whether returns are explained by common factors | The strategy may be loading on size, value, or momentum rather than skill | | Simpler strategy | Whether complexity adds value | The strategy may be over-engineered with no real edge | | Out-of-sample and after-cost | Whether the edge persists in real conditions | The strategy may be overfit or eroded by transaction costs | Each rung on the ladder strips away an excuse. If your strategy still looks good after the last comparison, you have something worth defending. ### A Simple Way to Think About It Imagine you're hiring a stock picker. You wouldn't just look at their total return. You'd ask: "Could a monkey throwing darts have done as well?" That's the benchmark ladder in a nutshell. It's not about making the comparison impossible. It's about making the hurdle honest. So next time you see a beautiful equity curve, don't get dazzled. Ask: "Compared to what?" And then build the ladder that gives you a truthful answer.