Part 1 of 2 — the intuition. Part 2 covers the methodological remedies.
Every systematic researcher has lived this moment. The baseline equity curve is decent but unexciting: Sharpe 0.8, a few uncomfortable drawdowns. Then the idea arrives. What if I filtered out entries when volatility is too high? You add the filter, rerun the backtest, and Sharpe climbs to 1.1. Pleased, you continue. What if I avoided post-FOMC days? Sharpe 1.3. What if I required the long-term trend to be aligned? Sharpe 1.5.
In a few hours you've turned a mediocre strategy into something that looks institutional. The problem is that, with high probability, you've improved nothing. You've simply memorized the past.
This process — adding conditions one at a time, keeping the ones that raise the metric and discarding the ones that lower it — is what I call sequential optimization. It's among the most widespread practices in amateur and semi-professional quantitative research, and also among the most effective at producing strategies that perform beautifully right up until the day they go into production.
Why the brain falls for it
The trap is insidious precisely because each individual step seems reasonable. You're not running a brute-force grid search over ten thousand parameter combinations — that one you recognize as overfitting. You're doing something that looks like science: you form a hypothesis ("high volatility hurts this strategy"), you test it, and you confirm it with data. The scientific method, right?
No. And the difference lies entirely in a detail that usually stays invisible: how many hypotheses you actually tested, and how many you silently discarded.
When you add a filter and keep it because "it works," you're making a choice conditioned on the outcome. But that filter didn't come from nowhere: it was one of many candidates you could have tried. The volatility filter, the time-of-day filter, the day-of-week filter, the regime filter, the correlation-with-another-instrument filter. You try five, you keep two. Those two aren't "the filters that work" — they're the two that won the noise lottery on this particular sample.
The mechanism, without formulas
Imagine your baseline strategy is actually pure chance: expected return exactly zero, no edge. On a finite sample, though, the realized return will never be exactly zero — there will be noise. Now take any binary filter, completely uncorrelated with anything real: split history into "days when the filter is active" and "days when it isn't."
By pure arithmetic, one of the two subsets will have performed better than the other. Always. Keep that one, and your "no-edge" strategy now shows a positive return. You've discovered nothing about the market: you've discovered which way the noise fell on this sample.
Repeat the game three, four, five times with different filters, keeping the good side each time, and the lucky selections compound. Each filter carves the sample keeping the portion that, by chance, was favorable. The final equity curve is a collage of lucky tails stitched together. Beautiful to look at. Devoid of predictive power.
The tell: hidden degrees of freedom
Here's the point almost everyone misses. When you report "I tested three filters," your backtest knows about far more. Every filter you could have tried and discarded after a glance at the equity curve counts as a test. Every threshold you nudged ("volatility above the 90th percentile... no, the 80th works better") is a test. Every time you looked at the result before deciding the next step, you spent a degree of freedom.
The number of hypotheses actually tested isn't the one you write in the report. It's the one that ran through your head while you stared at the charts. And that number is almost always an order of magnitude higher than you'd admit.
This is the crucial difference from grid search: in grid search the degrees of freedom are explicit and countable. In sequential optimization they are implicit and invisible, which makes them far more dangerous. You can't correct for tests you didn't record having run.
Why it survives to production
The reason sequential optimization is so poisonous — more than plain grid search — is that it produces narratively coherent strategies. At the end of the process you have not only a nice curve, but also a story: "I avoid extreme volatility, I stay out on macro days, I require trend alignment." It sounds like sensible risk management. You can present it in an interview. You can convince yourself.
Brute-force grid search produces suspicious numbers with no story, and a careful researcher eyes it with distrust. Sequential optimization produces pretty numbers with a plausible story, and it lowers every defense. It's overfitting disguised as domain understanding.
What to expect in production
The typical fate of a strategy built this way is almost a signature. In-sample, Sharpe is high and stable. In the first days or weeks of live trading, performance is mediocre but not alarming — the out-of-sample noise hasn't fully surfaced yet. Then, gradually, the strategy converges toward its true nature: the real expected return, which was zero or negative net of costs. The filters that seemed protective reveal themselves as arbitrary constraints that only reduce the number of trades, not the risk.
The diagnostic tell is brutal in its simplicity: the more your out-of-sample Sharpe falls short of your in-sample Sharpe, the more hidden degrees of freedom you spent. The gap isn't bad luck. It's the bill for the noise you memorized.
In Part 2 we go technical: how to quantify hidden degrees of freedom, why the deflated Sharpe ratio corrects for exactly this, and how to structure hold-out and nested cross-validation so that the research process itself — not just the final strategy — is subjected to honest validation.
quanthedgeai — the research arm of algosworksai. Systematic research on mid-frequency instruments.
Get the monthly Market Regime Note
Regimes, volatility and correlations across major futures markets — with the code behind the charts. Free.
Subscribe →