Glossary
Overfitting and sample size
Last reviewed: 26 September 2026·Tradelyze
Overfitting, also called curve fitting, is tuning a strategy's settings until the backtest matches the past too closely to work in the future. Say you nudge an RSI length from 14 to 17 and a stop from 1.5% to 2.1% until the Strategy Tester curve looks perfect. That can fit luck in those prices, not a repeatable edge.
In plain English
If you keep adjusting a strategy's settings until its past results look great, you can end up with settings that only fit the past. The more combinations you try, and the fewer trades the backtest contains, the more likely that becomes. Settings that still work when nudged, tested on many trades from different kinds of market, are much harder to fool yourself with. As a rough guide, a backtest with fewer than about 30 closed trades supports almost no conclusion.
New to this? Start with strategy optimization.
How many trades does a backtest need before it means anything?
Count the closed trades before you read any other number. With 20 trades, two results flipping from win to loss turn a 60% win rate into 50%, so the figure mostly measures luck. Fewer than about 30 trades supports almost no conclusion, and 100 or more trades taken in different kinds of market is a sturdier base. Neither number comes from a study; both are rules of thumb. In Tradelyze, read the Trade Count tile first: the permutation test needs at least 20 trades to report a significant result.
A closed trade is one that has been both entered and exited, so its profit or loss is final. The win-rate example is constructed arithmetic: 12 wins in 20 trades is 60%, and 10 wins in 20 is 50%. A handful of trades can swing the profit factor and the Sharpe ratio just as easily. Paying for a prop firm challenge on the strength of 20 backtested trades is paying to find out whether those trades were luck.
Tradelyze's Trade Count tile sits on the Best Metrics card and counts closed trades. Its tooltip gives the same rule of thumb: fewer than 30 trades makes the other metrics unreliable. Two checks inside the robustness score also have hard minimums. The permutation test needs at least 20 trades. It asks whether your result beats copies of the same trades with wins and losses randomly flipped. The deflated Sharpe ratio check needs at least 5. It discounts the Sharpe ratio for how many settings were tried.
Below those counts, each check refuses to run and scores zero points toward the robustness score, and its badge reads Too Few Trades. A refusal also counts as a failed check, because too few trades is a fact about the strategy, so the whole score is multiplied by 0.69. As computed with Tradelyze's scoring code, a run whose permutation test is refused, with the other three checks perfect and 400 days of data, scores 49.0, which the card shows as 49, grade D, verdict FRAGILE. Reaching a minimum only lets the check run; it does not make the sample big enough.
Every trade-count threshold in circulation, 30 included, is a convention rather than a finding. These numbers are usually quoted as though a study produced them, so the table below says where each one comes from.
| Trades | What can be claimed | Source |
|---|---|---|
| under 30 | Almost nothing. A win rate, a profit factor and a Sharpe ratio computed on 18 trades are all descriptions of 18 events. A few lucky or unlucky trades can move any of them. | No primary source. A general statistics rule of thumb, not a trading-specific finding. Where the 30 comes from is covered under Going deeper. |
| 30 | A bare minimum for a first read on out-of-sample data, meaning data the strategy was not tuned on. Enough to say a result is not obviously absurd. Not enough to rank two strategies against each other. | No primary source. Convention, widely repeated in trading education. |
| 100 | The commonly quoted threshold for reasonable confidence in a win rate or an expectancy figure (the average result per trade). | No primary source. Convention, repeated across trading forums and vendor documentation. |
| 200–500 | The standard quoted for institutional sign-off, as long as the trades span more than one kind of market. That condition is part of the threshold, not an extra. | No primary source. Convention. |
| 200 min 500 better 1,000 good |
One forum member's ladder: 200 trades as the bare minimum, 500 better, and 1,000 as the point where sample size stops being the objection. | Forex Factory, thread 297942: "200 trades are bare minimum, 500 is better, 1000 is really good." One member's opinion in a thread where others disagree, not a Forex Factory standard. |
| Trade count is necessary but not enough on its own. The trades also need to be independent of each other and to come from different kinds of market. | ||
Trades that all ride the same price move count for less than their number suggests. Why correlated trades don't count as separate evidence is covered under Going deeper.
Why is regime coverage a separate requirement?
A market regime is a stretch of prices with one broad character, such as a steady uptrend, a sustained downtrend or a quiet range. Regime coverage asks how many kinds of market the backtest saw, not how many trades it took. A backtest that never met a sustained downtrend says nothing about how the strategy behaves in one. That holds however many trades it contains. The same goes for calm and volatile periods, busy and thin markets, and, for rate-sensitive instruments, the interest-rate cycle.
To check it, look at the calendar rather than the trade list. Split the backtest by year and count the trades in each year. As a constructed example, a strategy with 400 trades, 380 of them inside 2023, is really a 2023 backtest. Both requirements have to hold at once: enough trades, and enough different conditions for those trades to come from.
How do I know if my strategy is overfit?
No single test proves a backtest is not overfit. Overfitting describes how a strategy was built, not something you can read off its results. What exists is a set of checks, listed here in order of how much each one tells you for the effort it takes.
- Parameter stability. Vary each parameter one step in each direction and look at what happens to the metric. Smooth degradation is what a real effect looks like; a collapse means the result depends on the exact value. This is the cheapest and most informative check available, and how to tell a parameter plateau from a parameter spike covers it in detail.
- Perturb the data, not the parameters. Shift the backtest start date by a month. Run the same rules on a correlated instrument. Run them on the neighboring timeframe. None of these should matter to a genuine edge, and all of them routinely destroy a fitted one. This test is stronger than parameter stability because it changes the sample rather than the model.
- Count the configurations you tried. A configuration is one complete set of input values. Count all of them, not only the ones you kept, including the manual attempts before you automated the search. That count drives how much of the result is luck of selection. It is also the input to minimum backtest length.
- Remove one rule at a time. A robust strategy degrades gracefully when a filter is deleted. An overfit one collapses, because each added condition was patching a specific historical episode rather than expressing a market behavior. If every rule is load-bearing, the rules are describing the sample.
- Check the concentration of profit. Sometimes a single trade, month or window supplies most of the net profit. Then the backtest documents one event, and the rest is padding. TradeStation's walk-forward optimizer includes a profit-concentration criterion for exactly this reason. Under TradeStation's default settings, no individual time period may contribute 50% or more of total net profit. The 50% is a configurable default rather than a fixed rule.
- Round the parameters. If a strategy needs a lookback of 47 and falls apart at 45 or 50, the 47 is fitting noise. Real parameters survive rounding to something a human would have chosen.
- Out-of-sample and walk-forward results. Results on data the strategy was not tuned on, including walk-forward analysis, are genuinely informative the first time. They get weaker every time after, for the reasons in does out-of-sample testing fix overfitting. They come last on this list rather than first, which is the reverse of how most workflows are ordered.
How do you tell a parameter plateau from a parameter spike?
A parameter plateau is a broad range of settings that all give similar, acceptable results. A parameter spike is one exact setting that scores well while its neighbors do badly. A plateau is weak evidence of a real edge, and a spike is a sign of noise. That is why practitioners pick the center of a plateau, never the single best value.
The reasoning is about what each shape implies. Sometimes the metric jumps at one value and collapses one step either side. Then the result is a property of that exact price path, not of any market behavior. When nearby values perform alike, the metric changes smoothly in that region. A whole range of settings then captures the same behavior, and that behavior is what the strategy is supposed to trade.
Choosing the center rather than the highest point follows directly. The highest point inside a plateau is where the noise happened to be most favorable. The center is furthest from the edges, so if your estimate of the right value is slightly off, you still land inside the acceptable region.
What is the difference between best parameters and recommended parameters?
In Tradelyze, Recommended Parameters are picked separately for each prop firm you select, from the trials of one parameter search. Tradelyze grades those trials against the firm's rules and ranks them the same way as the Top Trials table. The highest-ranked trial that passes every rule is recommended; if none passes, a trial is still recommended, chosen by the same ranking. So a firm's Recommended Parameters can differ from the search's overall best settings, and from another firm's. Walk-forward and the robustness checks never change them. Read them as an optimistic choice, the top of an in-sample ranking, and check the plateau around each value before trading them.
Other workflows use the two terms differently. Best parameters are the raw winner of a search: the single combination with the highest in-sample score. An in-sample score is measured on the same data the search tuned on. Some workflows also produce a separate recommended set. That set is what remains after later checks. Do neighboring values behave similarly? Did the choice hold up on unseen data? Does it survive small changes to the data?
The raw winner of a large search is where noise happened to be most favorable. That makes it the point least likely to repeat. When a separately checked set agrees with the raw winner, the search landed on a plateau and there is nothing to decide. When the two disagree, the disagreement is itself the finding: the top scorer failed a check the other set passed.
On Tradelyze's results page, the Recommended Parameters section lists the winning settings and notes how many changed from your script's defaults. It also offers a Copy JSON button. With several firms selected, that section shows one firm's settings; each prop firm card lists its own under Parameters Used for Evaluation.
Neither the robustness checks nor walk-forward analysis turns these settings into tested out-of-sample settings. Tradelyze's robustness score is measured on the search's overall winner, which can differ from the settings recommended for a particular prop firm; when it does, an amber banner on that firm's robustness card says so. Walk-forward analysis never runs these settings on data they were not chosen on. Tradelyze picks the method for each run. Re-tuned each window re-runs the search in every window, which grades the tuning method. One run, split by period, the default for eligible strategies, scores the search's winning settings on periods of the same history they were chosen on; when a firm was recommended other settings, a banner on the walk-forward card names it. How the walk-forward methods differ explains both, and why a third method, parameter stability, was retired on 26 September 2026.
The one check of these exact settings on data they were not chosen on is the held-out test on each firm's card. By default the last 25% of the history is withheld from the whole search, and each firm's Recommended Parameters are then run once over the full history. Its badge reads Confirmed, Not Confirmed, Inconclusive or Pending, or Not run when nothing was held back.
None of this replaces trade count and regime coverage. Take a recommended set drawn from 40 trades inside a single year. It describes 40 trades inside a single year, however many checks it passed. How many trades a backtest needs explains why.
How do I read the Parameter Search Space table?
Tradelyze's Parameter Search Space table lists your strategy's tunable inputs, one row each, including any left at your script's default. Click its heading to open it. A star in the Changed? column means the Best Value differs from your script's default, not that it is better.
| Column | What it shows | How to read it | Source |
|---|---|---|---|
| Parameter | The input's name, taken from its title in your script's settings where it has one. | Match it to the same input on the Inputs tab of your strategy's settings in TradingView. | Tradelyze implementation. |
| Type | "int" (a whole number), "float" (a decimal), "bool" (an on/off switch) or "categorical" (a fixed list of choices). | A switch or a list of choices can turn a whole rule on or off, which changes which trades are taken. | Tradelyze implementation. |
| Range | For numbers, the lowest and highest values searched and the step between them, written as [min..max] step=n. Switches show "true / false"; lists show their choices. | A wide range with a small step gives the search more chances to find a lucky value. | Column: Tradelyze implementation. Reading: No primary source. |
| Default | The value written in your script. | The value your strategy uses when nobody changes the input, and the starting point every change is measured from. | Tradelyze implementation. |
| Best Value | The value in the recommended settings, the same ones listed under Recommended Parameters. They are picked for a prop firm's rules, so with several firms selected the table shows one firm's settings. | A value at either end of its Range suggests the search wanted to go further than the range allowed. | Column: Tradelyze implementation. Reading: No primary source. |
| Changed? | A star, on a tinted row, when Best Value differs from Default. | The star means "different from your script", not "better". No star means the two values match, or no best value was reported. | Tradelyze implementation. |
A Best Value at the very edge of its Range suggests the search wanted to go further. That reading is methodology guidance with no primary source. Widening the range and re-running is one response. But every extra value tried is another chance to find a lucky setting, so the bar for trusting the result rises with it.
The table shows one winning value per input, not the plateau around it. To see whether neighboring values also work, re-run the strategy in TradingView with one input nudged a step either way and compare the results. An input left at your script's default carries no star. So does an input the search tuned that happened to settle on its default. A missing star alone does not tell you which of the two happened.
What does the walk-forward version of the plateau test look like?
The plateau test has a sharper form that does not require plotting anything. Run a walk-forward analysis and record which parameter value the optimizer selected in each successive run. A stable parameter clusters in a narrow band across those runs. An unstable one does not, and the formulation of the tell in Build Alpha's walk-forward optimization documentation is a parameter that jumps around like this:
12, 47, 23, 35, 55, 13
Build Alpha's reading: this is most likely curve fitting.
A sequence like that says the optimizer found a different answer every time it was shown different data. That is what happens when there is nothing stable to find. The optimizer is not malfunctioning; it is faithfully reporting the best fit to each sample, and the samples disagree. Compare a parameter that lands within a few points of the same value in every run. That is the walk-forward equivalent of a plateau.
| Signal | Parameter plateau | Parameter spike | Source |
|---|---|---|---|
| Neighboring parameter values | Perform similarly; the metric changes smoothly across the region | Degrade sharply; the metric collapses one step either side | No primary source. Practitioner convention. |
| Value selected across walk-forward runs | Clusters within a narrow band | Jumps around, as in Build Alpha's 12, 47, 23, 35, 55, 13 | Build Alpha's walk-forward optimization documentation. Vendor documentation, not a study. |
| What to select | The center of the region, not its highest point | Nothing. A spike is a coincidence, not a selection | No primary source. Practitioner convention. |
| What it is evidence of | Weak evidence of signal | Evidence of noise | No primary source. |
| Residual risk | A plateau can still be an artifact of a single regime, or a plateau over noise if the trade count is small | None left to assess — the spike is already the finding | No primary source. |
The part that is usually left out
A plateau is weak evidence, not proof. A plateau can mislead in two ways. It can be a broad region computed from too few trades, which makes it a smooth surface over noise. Or it can be a genuine plateau that exists only in one market regime, and the problem will not show until the regime changes. A plateau is a reason to keep going, not a reason to stop. Read it together with the trade count and the regime coverage described in how many trades a backtest needs.
How many parameters is too many?
Every input that is optimized rather than fixed adds another way to bend the backtest to fit the past. With enough of them, any historical price series can be fitted exactly. The result then says nothing about the future, because it was built to describe the past. Statisticians have a name for this count, explained in what degrees of freedom are under Going deeper.
The heuristic that circulates on trading forums is that more than a couple of optimized parameters puts you on thin ice. That is a heuristic and nothing more: there is no published threshold, and anyone quoting a specific maximum is quoting a preference. What can be stated exactly is how fast the number of combinations grows.
The number of combinations is the number of values tried for each input, multiplied together once per input. One worked line shows it:
| Parameters optimized | Values tried per parameter | Configurations searched |
|---|---|---|
| 1 | 10 | 10 |
| 2 | 10 | 100 |
| 3 | 10 | 1,000 |
| 4 | 10 | 10,000 |
| 5 | 10 | 100,000 |
| 6 | 10 | 1,000,000 |
| Arithmetic: multiply the number of values tried for each parameter together, once per parameter. Every configuration is another chance for noise to look like signal. | ||
Set those counts against the trade counts in how many trades a backtest needs. Four parameters at ten values each is 10,000 configurations competing to explain a backtest that may contain 100 trades. The search will find a configuration that looks excellent on those 100 trades. Finding one is not evidence, because a search over 10,000 configurations would also have found one on random data.
Two refinements that matter more than the raw count:
- Not every parameter is equally dangerous. A parameter that changes which trades are taken alters the sample itself — different entries, different exits, a different set of observations. A parameter that only scales position size leaves the sample intact and rescales the result. The first kind gives the search far more room to fit noise than the second.
- The uncounted parameters are usually the expensive ones. The instrument, the timeframe, the backtest date range and every on/off rule toggle are all choices made by searching. They are almost never counted as parameters. Take a strategy described as having "two parameters" that was arrived at after trying four instruments, three timeframes and six rule combinations. It was selected from a much larger space than two.
What is parameter gating, and why are some inputs left at their default?
Parameter gating checks whether changing an input could have changed any trade. An input that could not is left at the script's default rather than reported as a tuned recommendation. Searching such an input is not dangerous, only wasted: it multiplies the number of combinations without producing a single different backtest.
Most strategies contain on/off switches, such as use a stop loss, use a trailing stop or use breakeven. Many also have numbers the script only reads while a switch is on. With the stop loss switched off, the risk-per-trade number is never read, so every value of it produces the same backtest.
Tradelyze works this out by reading the strategy's code before the search starts. An input that only matters while a switch is on keeps your script's default on every test where that switch is off. An input whose switch the optimizer never varies is left out of the search for the whole run. It still appears in the Parameter Search Space table, so you can see that its number is yours rather than a tuned one.
When the code cannot be followed, the input is searched as normal. The bias is deliberate. Searching an input that cannot move a trade only costs trials. Freezing one that can would report a best result built on a value nothing ever varied, which is a silent and much worse failure.
Reading code is reasoning, not evidence, so Tradelyze checks it with one extra backtest. The search's overall winner is run again with every supposedly unused input moved to another value, and the trades from the two runs are compared. If not one trade changes, those inputs really were unused. If anything changes, the reading of the code was wrong, and the run says so rather than keeping the claim.
Parameter gating is Tradelyze's own engineering answer to wasted combinations. It is not a published method, and there is no citation behind it.
Does the number of parameters or the number of trials matter more?
What raises overfitting risk most is how many combinations were tried, not how many inputs a strategy has. As constructed arithmetic, two inputs tested at 50 values each is 2,500 combinations, while six inputs tested at 2 values each is only 64. Every extra try is another chance for luck to produce a winner that will not repeat.
What determines overfitting risk is the number of configurations tried, not the split between in-sample and out-of-sample data. A configuration is one complete set of input values, and a trial is one backtest of one configuration. Parameter count matters mainly because it multiplies into the number of configurations. How Tradelyze sizes its trial budget, and where a report shows how many trials ran, is covered in how many trials an optimizer should run.
The formal treatment is Bailey, Borwein, López de Prado and Zhu, Pseudo-Mathematics and Financial Charlatanism, in the May 2014 Notices of the American Mathematical Society. The paper's abstract says that high simulated performance is easily achievable after backtesting a relatively small number of alternative strategy configurations. It adds that the higher the number of configurations tried, the greater is the probability that the backtest is overfit.
A high in-sample Sharpe ratio is close to guaranteed once enough configurations have been tried, whether or not any of them has an edge. Whether the overfit winner then merely breaks even live or loses money depends on a condition the same paper spells out. Does an overfit strategy just break even, or can it lose money covers that condition.
What is minimum backtest length?
The same paper introduces minimum backtest length. It is the amount of history needed before an in-sample performance figure can be taken seriously. The more configurations were tried to obtain that figure, the more history it needs. Search harder and the evidence needed for the same reported Sharpe ratio goes up. A fixed history therefore supports only a limited amount of searching.
That inverts the usual instinct. Running more trials feels like more thorough work; in evidence terms it is spending down a fixed budget of data. Two honest responses exist: try fewer configurations, or obtain more data. Reporting the winner of a large search against a short history is not one of them. Tradelyze's robustness card includes a Min Backtest Length check that compares Required and Available Years, explained in minimum backtest length on the robustness score page. It is shown for information only: it earns no points and does not change the robustness score, grade or verdict. A different history rule does count. With less than 30 days (one month) of data, or a span that cannot be measured, the robustness score is multiplied by 0.69, so C+ and MARGINAL are the best it can show. Those days are counted over your whole upload, first bar to last, even when the last part was held back for the held-out test. The card's Enough for ROBUST (≥ 30 days) row gives that count as days of data and reads Yes or No, or NOT MEASURED when the span could not be measured.
The practical consequence is a reporting rule. Record the trial count and publish it next to the performance figure. Two identical Sharpe ratios obtained from 20 trials and from 20,000 trials are not the same evidence, and they cannot be compared without it. Tradelyze shows how many trials ran at the top of its optimization results.
A strategy summary that omits how many configurations were searched has left out the number that determines how to read every other number on it.
Does out-of-sample testing fix overfitting?
No. A holdout, meaning a slice of data kept aside for a final test, works cleanly exactly once, and almost nobody uses one only once.
The decay mechanism is straightforward and it does not require any dishonesty. Test a strategy on the holdout, get a poor result, adjust something, test again. Each rejected attempt taught you something about the holdout, and that knowledge shapes the next attempt. After a few rounds the holdout has been incorporated into the design process; it is training data that has not been labeled as such. Nothing improper happened at any step, and the guarantee is gone anyway.
The arithmetic of how quickly this happens is not subtle. Suppose each test has a 5% chance of a false positive: a pass for a strategy with no real edge. Suppose too that the tests are independent. The chance of at least one false positive across ten tests is then 1 - 0.95 to the tenth power, which is 40.1%. Read 40.1% as a floor rather than as an estimate.
Repeated holdout tests are not independent. Each attempt is shaped by what the previous one revealed about the holdout, and that dependence is exactly what makes repeated holdout use corrosive. The independence assumption is what makes the number computable, and it is the assumption that flatters the result. The corresponding point made in EliteTrader thread 314065 is that trying ten different things makes it very likely you are fooled at least once. That is the same observation without the arithmetic, and it is one trader's post rather than a study.
The order in which the work is done matters as much as the count. One trader in EliteTrader thread 290496 warned that backtesting first leaves a later walk-forward analysis massively overstated. The reason given: you're only testing stuff that you know already works. A walk-forward analysis run on a strategy developed by backtesting on the same history is not testing that strategy against unseen data. The data was seen during development, and the walk-forward is re-scoring a survivor.
Check this on your own results
Record how many times the holdout has been evaluated, and treat that count as part of the result. A holdout consulted twenty times is training data with an optimistic label. If you cannot say how many times you looked, assume it is more than you think. Assume, too, that the out-of-sample figure is overstated by an unknown amount. This applies to walk-forward analysis and walk-forward efficiency exactly as it applies to a simple holdout.
The same trap applies to re-running an optimization after each failed check. See why re-optimizing until it passes is a trap.
Out-of-sample testing does catch one specific failure: settings fitted to one stretch of data that stop working on the stretch right after it. That is a real failure and worth detecting. But it is not protection against picking a lucky winner from many trials. The probability of backtest overfitting measures that risk; a holdout cannot.
How should I read in-sample results?
Tradelyze's best metrics are mostly in-sample: they describe the recommended settings over your full history, and the search picked those settings on all of it except, by default, the last 25% it withheld for the held-out test. While a run is still in progress, or if the full-history run of those settings fails, they cover only the part the search saw, and the card says so. That makes them the most optimistic numbers on the results page. Treat them as a ceiling. Results on data the strategy has not seen will almost certainly be worse, so a best result that already looks marginal is a warning.
The Best Metrics card shows the winning settings run over the full history, including the stretch the search used to pick them. The Profit tile, for example, shows the total gain or loss in account currency and as a percentage of starting capital. A reading of +5.0% means a $100,000 account became $105,000. Read these figures against the walk-forward results rather than on their own.
In general, in-sample results are the performance of a selected configuration measured on the same data that selected it. They are not an estimate of what the strategy does; they are the maximum of a search, and a maximum is biased upward by construction. The bias grows with the number of configurations tried. That is the point of minimum backtest length: the same in-sample Sharpe ratio means less the harder it was searched for.
That makes an in-sample figure useful for two things and useless for a third. It is useful as a ceiling, since a result that is already marginal in-sample settles the question without further testing. It is useful as one side of a comparison: against the out-of-sample figure, against neighboring parameter values, and against the trial count. It is not useful as a forecast. Quoting it as one without the trial count beside it leaves out the number that says how far to discount it.
Why can a backtest that is not overfit still fail in live trading?
The backtest-to-live gap is a separate failure mode from overfitting and it is worth keeping separate, because the fixes are unrelated. A strategy can be entirely free of curve fitting, with broad plateaus, few trials and hundreds of trades across several regimes. It can still lose money live, because the simulator modeled an execution environment that does not exist.
| Backtest assumption | What actually happens live | Worst affected | Source |
|---|---|---|---|
| Tick data is real | Tick data is frequently interpolated from higher-timeframe bars, so intrabar paths are invented. A stop and a target inside the same bar are resolved by the interpolation rule rather than by what happened. | Any strategy whose stop and target can both be touched within one bar | Forex Factory thread 523334: tick data interpolated from one-minute bars. A practitioner account, not a study. |
| Historical spread is available | Historical spread is often absent from the data entirely and is replaced with a constant. Every cost estimate is then a guess that ignores the widening that happens exactly when signals fire. | News-reactive and session-open strategies | Forex Factory thread 39799: spread not deducted from results. A 2007 post disputed in its own thread. |
| Limit orders fill on touch | A limit order fills only if the queue ahead of it is consumed. Price touching the level is necessary, not sufficient, and there is no queue-position model in a typical backtester. | Passive and mean-reversion strategies that assume they are filled at the extreme | No primary source. Follows from how a limit order executes. |
| Bar-close signals fill at the close | A signal computed on the close of a bar is filled at the open of the next bar in live trading. The gap between them is a real cost that the backtest recorded as zero. | Every bar-close system; worst where the instrument gaps | No primary source. |
| Spread and swap are static | Spread moves all the time. Swap, the overnight financing charge, accrues on rules that vary by day and by position side. Backtests commonly treat both as fixed constants. | Overnight and multi-day holding strategies | Forex Factory thread 39799: spread not deducted and rollover not applied. A 2007 post disputed in its own thread. |
| Both Forex Factory threads are practitioner accounts of MetaTrader-style backtesting, not a controlled study. The limit-order row follows from how a limit order executes, not from a forum report. | |||
Scalping systems with profit targets under 20 pips are the worst case. A pip is the standard unit of price movement in a currency pair. Every item in the table of execution assumptions is an error of a fraction of a pip to a few pips. Against a 200-pip target those errors are a rounding adjustment; against a 15-pip target they are a large share of the entire edge.
A scalping backtest can be arithmetically correct and still describe a strategy that does not exist at the broker. The modeled costs and the real costs can differ by more than the target.
The separation matters for diagnosis. If a strategy passes every overfitting check on this page and still fails live, the problem is execution rather than overfitting. More parameter-stability work will not find it. The tests that do are forward testing and paper trading. Forward testing runs the frozen strategy on new prices that arrive after it was finished. Paper trading places simulated orders in live market conditions without risking money.
Where this appears in Tradelyze
In a Tradelyze report, read the Trade Count tile first. Then compare Best Metrics with Recommended Parameters and the walk-forward card to see how much of the result survived. Tradelyze re-runs an uploaded TradingView Pine Script strategy on price data you upload and checks the result against your exported trade list. It then runs parameter optimization, walk-forward analysis, a robustness score from four scored checks and prop-firm rule checks. It does not place trades, give financial advice or guarantee a challenge pass, and it is in beta.
If the walk-forward badge reads Not Confirmed or Not Consistent, or the robustness verdict reads MARGINAL or FRAGILE, do not keep changing settings until something passes. See why re-optimizing until it passes is a trap. To judge the whole report, not one tile, use the pre-trade checklist.
Create an account. Already a user? Open your strategies.
Going deeper
The sections below go deeper: the statistics behind the trade-count rules of thumb, whether an overfit strategy only breaks even or can lose money, why a t-statistic of 2 is too low a bar after a search, and how the probability of backtest overfitting is estimated. You can skip them and still read your own report.
What statistics sit behind the trade-count rules of thumb?
The 30-trade rule of thumb comes from general statistics, not from trading research. It is a textbook convention tied to the central limit theorem. That theorem says the average of many independent observations behaves more and more like a normal, bell-shaped distribution as their number grows. The convention treats about 30 observations as enough for that approximation to be workable; the theorem itself names no number.
The rule matters for two common tools. With fewer than about 30 trades, a t-test (a standard check of whether an average result differs from zero) and a confidence interval on the average trade are not dependable in practice. Exact and non-parametric tests carry no n ≥ 30 requirement, though they answer narrower questions than whether a strategy is worth trading. Tradelyze's Trade Count tooltip uses the same 30-trade rule of thumb.
Tradelyze's permutation test is a non-parametric test, and its 20-trade minimum has a different reason. It flips the sign of each trade's result at random, and n trades allow only 2 to the power n distinct sign patterns. Five trades allow just 32 patterns, too few for the p-value it would report, while 20 trades allow 1,048,576. That reasoning is documented in Tradelyze's robustness code.
Why don't correlated trades count as separate evidence?
The qualifier that outranks the trade-count table
A large number of highly correlated trades from one regime is worth less than a much smaller number of clean independent trades spanning a bull and a bear market. Trade count is a proxy for information, and the proxy breaks whenever the trades are not independent observations of the thing being measured. How much smaller is not something this page will put a number on: there is no published exchange rate between correlated and independent trades, and any specific pair of figures would be invented.
A trend-following system run through a single sustained bull year can easily produce 300 trades. They are all long, all in the same instrument, all triggered by the same signal, and all funded by the same underlying move. Statistically that is close to one observation recorded 300 times. The nominal count says 300; the effective sample size, the number of genuinely independent pieces of evidence, is far smaller, and there is no reliable way to recover it after the fact.
Three common sources of correlation between trades, all of which inflate the count without adding information:
- Overlapping positions. Trades held at the same time in the same instrument share the same price path. Two positions open across the same afternoon are not two draws from a distribution.
- One driver. A cluster of trades all responding to a single macro event, earnings release or regime shift is one bet expressed several times.
- One signal in one condition. A mean-reversion rule fired repeatedly inside a single ranging month tests the rule against one month, however many trades it generated.
What are degrees of freedom?
In statistics, degrees of freedom loosely means the number of quantities a model was free to choose when it was fitted to the data. In a strategy, every optimized input is one more. The more freely chosen quantities there are for a fixed number of trades, the easier it is to fit noise exactly, and the less the fit proves. Inputs differ in cost: one that changes which trades are taken uses up far more freedom than one that only scales position size.
Does an overfit strategy just break even, or can it lose money?
An overfit strategy does not always just break even live; in some conditions it can be expected to lose money. The second half of the finding in Bailey, Borwein, López de Prado and Zhu's May 2014 paper on backtest overfitting, Pseudo-Mathematics and Financial Charlatanism, is routinely dropped when the paper is cited: backtest overfitting does not necessarily produce out-of-sample results that average to zero.
The paper states the outcome conditionally: out-of-sample performance will be around zero if the process has no memory, but it may be significantly negative if the process has memory. Memory here means the performance series is not independent from one period to the next. Where memory is present, the selected configuration is systematically on the wrong side rather than randomly placed, because the configuration that best exploited a passing pattern in the sample is the one most exposed when that pattern reverses.
The paper names the mechanism a compensation effect: memory in the performance series raises the chance a configuration is selected in-sample precisely because some of its in-sample outcomes will be compensated for out-of-sample. In the authors' framing, the absence of memory is the only case with no reason to expect overfitting to induce negative performance. That is a conditional claim, not a general one, and this page carries the condition everywhere it carries the claim.
Check this on your own results
An overfit strategy is not reliably a neutral strategy with wasted effort attached. Bailey, Borwein, López de Prado and Zhu state the outcome conditionally — around zero if the return process has no memory, possibly significantly negative if it has memory — and the condition has to travel with the claim.
Where memory is present, expecting to lose nothing from a curve-fitted system is optimistic: the expectation is negative rather than zero, the base case for a strategy selected as the best of many thousands on a fixed history is a loss, and the size of that loss is not bounded by how mediocre the alternatives looked. The same authors put it more bluntly still: the customary disclaimer that past performance is not an indicator of future results is too optimistic in the context of backtest overfitting.
Why do I need a t-statistic higher than 2?
A t-statistic measures how far an average result sits from zero compared with how noisy the results are. A value of about 2 corresponds roughly to the conventional 5% significance level for a single, pre-specified test, meaning about a 1-in-20 chance of a result that large when there is no real effect. Neither condition holds for a trading strategy selected from a search, so the conventional threshold is the wrong one to apply.
Campbell R. Harvey, Yan Liu and Heqing Zhu argue in the Review of Financial Studies 29(1), 2016, that a newly discovered factor should clear a t-statistic of about 3.0 rather than 2.0, precisely because of how many candidate factors the literature has already tested. Their argument is a multiple-testing adjustment: when hundreds of candidates have been examined and only the successful ones published, the threshold that would be correct for one test is far too permissive for the winner of many.
Connecting that hurdle to a trade count uses one standard identity. The per-trade Sharpe ratio is the average trade result divided by the standard deviation of trade results, where standard deviation measures how widely those results vary. For n independent trades, the t-statistic of the mean is that per-trade Sharpe ratio multiplied by the square root of n. This per-trade figure is not what Tradelyze's Sharpe (Bar) or Sharpe (Daily) tiles show. Rearranged, the number of trades required to reach a given t-statistic is:
n = (t / per-trade Sharpe ratio)²
| Per-trade Sharpe | Trades for t = 2.0 | Trades for t = 3.0 | Source |
|---|---|---|---|
| 0.05 | 1,600 | 3,600 | Arithmetic, exact |
| 0.10 | 400 | 900 | Arithmetic, exact |
| 0.20 | 100 | 225 | Arithmetic, exact |
| 0.30 | 45 | 100 | Arithmetic; 45 is rounded up from 44.4 |
| 0.50 | 16 | 36 | Arithmetic, exact |
| Arithmetic from n = (t / SR)², assuming independent trades, rounded up to a whole trade. Only one cell needs the rounding: at a per-trade Sharpe of 0.30, (2 / 0.30)² = 44.4, shown as 45. Every other cell is exact. The independence assumption is generous; correlated trades require more. The t = 2.0 hurdle is the conventional single-test level; the t = 3.0 hurdle is the one Harvey, Liu and Zhu (2016) argue for. | |||
Read against the trade counts in how many trades a backtest needs, the table explains why so many backtests cannot support any claim. A per-trade Sharpe ratio of 0.10 needs 900 trades to reach a t-statistic of 3.0. A backtest with 200 trades and a per-trade Sharpe of 0.10 reaches roughly 1.4, which does not clear even the permissive single-test threshold, let alone the multiple-testing one.
The part that is usually left out
Harvey, Liu and Zhu's 3.0 was calibrated for the published asset-pricing factor literature, where the number of prior tests can be estimated from publication records. A private parameter search has a different trial count, usually unknown and frequently much larger than the number of factors in that literature. Treat 3.0 as a floor to think with rather than a constant that transfers unchanged — the correct hurdle for a search over 10,000 configurations is higher, not equal.
What is the probability of backtest overfitting (PBO)?
The probability of backtest overfitting is the probability that the configuration ranked best in-sample turns out to perform below the median of the alternatives out-of-sample. It is defined and estimated in Bailey, Borwein, López de Prado and Zhu, The Probability of Backtest Overfitting (SSRN 2326253). Tradelyze does not compute PBO. Its deflated Sharpe ratio check addresses a related question: whether a Sharpe ratio survives the number of settings tried.
The definition is worth reading closely, because it is not a statement about a strategy. It is a statement about a selection procedure: given this set of candidate configurations and this way of picking a winner from them, how often does the winner turn out to be a below-average choice on data it was not picked on.
A PBO near one half means the procedure is no better than choosing at random. A PBO near zero means the in-sample ranking carries information about the out-of-sample ranking, which is the property a strategy search is supposed to have and usually does not.
How is PBO estimated with combinatorially symmetric cross-validation?
The estimation method in the paper is combinatorially symmetric cross-validation, or CSCV. It starts from the full matrix of performance: every candidate configuration scored across every sub-period of the history, including the configurations that were rejected.
- Split the history into an even number of disjoint sub-periods of equal length.
- Form every possible combination of half of those sub-periods as a training set, with the remaining half as the matching test set.
- In each combination, find the configuration that ranks best on the training half, then record where that same configuration ranks among all configurations on the test half.
- The probability of backtest overfitting is the fraction of combinations in which the in-sample best configuration ranked below the median out-of-sample.
The method is symmetric in the sense that every sub-period serves in the training role and the test role equally often, so no single arbitrary split drives the answer. It is also deterministic given the performance matrix, which makes it reproducible in a way that a single holdout is not.
The part that is usually left out
CSCV needs the performance of every configuration tried, across every sub-period. Most retail optimization workflows discard the losing trials and keep the winner, which destroys exactly the matrix the method requires. If a probability of backtest overfitting is going to be computed at all, the full trial matrix has to be retained from the start — it cannot be reconstructed afterwards from the winning configuration.
Stage 3 · step 11 of 18. Next in the learning path: Walk-forward analysis
Frequently asked questions about overfitting and sample size
What is overfitting in a backtest?
Overfitting, also called curve fitting, is tuning a strategy's settings until the backtest matches the past too closely to work in the future. The strategy then scores well on the data used to build it and poorly on anything else, because its rules describe accidents of that one stretch of prices. A curve-fitted rule typically breaks when the start date moves, a setting is nudged or the instrument changes.
How many trades do I need for a backtest to be meaningful?
Fewer than about 30 closed trades supports almost no conclusion, and 100 or more trades taken in different kinds of market is a sturdier base. Around 200 to 500 trades spanning several market regimes is the standard quoted for institutional sign-off. None of these numbers comes from a study; all are rules of thumb. Tradelyze's permutation test needs at least 20 trades and its deflated Sharpe check at least 5.
Is 100 trades enough for a backtest?
100 trades is the commonly quoted threshold for reasonable confidence, and it is a convention with no primary source rather than a statistical result. It is also not sufficient on its own: 100 trades taken from a single trending year are far weaker evidence than 100 trades spanning a bull market and a bear market, because trade count and regime coverage are separate requirements.
How do I know if my strategy is overfit?
Check parameter stability first: vary each parameter and see whether performance degrades smoothly or collapses. Then perturb the data rather than the parameters, by shifting the start date, changing the instrument and changing the timeframe. Then count how many configurations you tried, since that number drives the risk. Out-of-sample results come last, because they decay with reuse.
What is the difference between a parameter plateau and a parameter spike?
A plateau is a broad region where neighboring parameter values perform similarly. A spike is a single value whose neighbors degrade sharply. The plateau is weak evidence of signal; the spike is evidence of noise, because a result that depends on an exact value is describing the specific price path rather than a market behavior. Select the center of the plateau, never the spike.
Are Tradelyze's Recommended Parameters tested on unseen data?
Only by the held-out test. By default Tradelyze withholds the last 25% of the history from the whole search, then runs each firm's Recommended Parameters once over the full history, and the held-out test on that firm's card shows how they did on the withheld part. With fewer than 200 bars nothing is held back. Walk-forward analysis either re-tunes each window, which tests the tuning method rather than those exact numbers, or scores the search's winning settings on periods of the history they were chosen on. Check the plateau around each value and the trade count as well.
How many parameters is too many for a trading strategy?
Each optimized parameter adds a dimension in which the equity curve can be fitted, and the forum heuristic is that more than a couple puts you on thin ice. The arithmetic is the reason: four parameters searched over ten values each is 10,000 configurations competing to explain a few hundred trades. There is no published threshold, so treat the heuristic as a heuristic.
Does the number of parameters or the number of trials matter more?
The number of configurations tried. The abstract of Bailey, Borwein, López de Prado and Zhu's May 2014 paper on backtest overfitting says high simulated performance is easily achievable after backtesting a relatively small number of alternative strategy configurations, and that the higher the number of configurations tried, the greater is the probability that the backtest is overfit. Parameter count matters only because it multiplies out into trial count.
What is the probability of backtest overfitting?
The probability of backtest overfitting, or PBO, is the probability that the configuration ranked best in-sample performs below the median of the alternatives out-of-sample. Bailey, Borwein, López de Prado and Zhu define it in The Probability of Backtest Overfitting, SSRN 2326253, and estimate it with combinatorially symmetric cross-validation. A PBO near one half means the selection procedure carries no information. Tradelyze does not compute PBO.
Does out-of-sample testing prevent overfitting?
No. A holdout works cleanly exactly once. Repeated trial and error gradually converts unseen data into known data, because each rejected attempt tells you something about the holdout that shapes the next attempt. As the point is made in EliteTrader thread 314065, trying ten different things makes it very likely you are fooled at least once. The number of holdout evaluations is part of the result.
Can a strategy pass walk-forward testing and still be overfit?
Yes. Walk-forward testing checks whether one chosen strategy, or one way of tuning it, held up on price data it was not tuned on. It cannot see how many other strategies or settings were tried before this one was picked. As it is put in EliteTrader thread 290496, if you backtest first, the performance of your walk-forward analysis will be massively overstated, since you're only testing stuff that you know already works.
Why do I need a t-statistic above 3?
Because so many candidate signals have already been tested. Harvey, Liu and Zhu argue in the Review of Financial Studies that a newly discovered factor should clear a t-statistic of about 3.0 rather than the conventional 2.0, precisely because the conventional threshold assumes a single pre-specified test. Any threshold calibrated for one test is wrong when it is applied to the winner of a search.
What is minimum backtest length?
Minimum backtest length is the amount of history required before an observed in-sample performance figure can be taken seriously, given the number of configurations that were tried to obtain it. Bailey, Borwein, López de Prado and Zhu introduce it in the same May 2014 Notices of the American Mathematical Society paper. The requirement grows with the trial count. Tradelyze's robustness card shows it as Min Backtest Length, for information only: it is not part of the robustness score, grade or verdict.
Why does my backtest not match live results?
Overfitting is one cause and the backtest-to-live gap is a separate one. Tick data is often interpolated and historical spread is often absent, limit orders are assumed to fill on touch with no queue position, bar-close signals fill at the next bar's open in live trading, and dynamic spread and swap are evaluated as static. Scalping systems with sub-20-pip targets suffer the most.
Sources
- David H. Bailey, Jonathan M. Borwein, Marcos López de Prado and Qiji Jim Zhu, Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance, Notices of the American Mathematical Society 61(5), May 2014, pp. 458–471. DOI 10.1090/noti1105; ams.org PDF (retrieved 28 July 2026). The conditional statement about out-of-sample performance is on p. 468: "performance will vary: It will be around zero if the process has no memory, but it may be significantly negative if the process has memory." The too optimistic phrase is on p. 467. The "high simulated performance is easily achievable" sentence is not in the journal article — the printed version carries no abstract, so anyone citing the sentence to a page number of the journal article is citing something that is not there. It is abstract text, and the abstract exists in two wordings: the April 2014 preprint at davidhbailey.com reads "high performance", and the abstract on the paper's institutional-repository record at scholarworks.wmich.edu reads "high simulated performance". Both retrieved 28 July 2026. That word matters: without it the sentence reads as a claim that high real returns are easy to obtain by trying configurations, which is the reverse of the paper's argument. What is easy to manufacture is the backtest.
- David H. Bailey, Jonathan M. Borwein, Marcos López de Prado and Qiji Jim Zhu, The Probability of Backtest Overfitting, SSRN 2326253 — definition of PBO and the combinatorially symmetric cross-validation procedure used to estimate it.
- Campbell R. Harvey, Yan Liu and Heqing Zhu, … and the Cross-Section of Expected Returns, Review of Financial Studies 29(1), 2016, pp. 5–68 — the argument that a newly discovered factor should clear a t-statistic of about 3.0 rather than 2.0.
- Build Alpha, What is Walk Forward Optimization?, buildalpha.com (retrieved 28 July 2026) — "a parameter that jumps from 12 to 47 to 23 to 35 to 55 to 13 is most likely curve fitting", under the heading "Optimal Strategy Parameters". Vendor documentation, not a study.
- Forex Factory threads, all retrieved 28 July 2026. Thread 297942, Manual backtest v EA testing — "200 trades are bare minimum, 500 is better, 1000 is really good." Thread 280665, How far back is ideal for backtesting — 200 trades as a minimum, offered by the poster as an opinion. Thread 523334, Best third party strategy tester for MetaTrader 4 — "MT4 takes M1 bars and interpolates ticks from that, which makes the data artificial - too 'smooth'." Thread 39799, MT4 EA - strategy tester - backtest results inaccuracies — "MT4 does NOT deduct spreads from results, nor applies daily rollover premiums", a 2007 post disputed inside its own thread. All four URLs refused automated re-checks on 16 September 2026 (HTTP 403 to every method tried), so the quotations above rest on the 28 July 2026 readings; each thread's title and topic were re-confirmed on 16 September 2026 by search. These are individual members' accounts, not Forex Factory standards and not a controlled study. No Forex Factory thread asserting the limit-order point in the backtest-to-live table could be located, so that row is reasoned from how a limit order executes. The "700 trades or more as ideal" figure that circulates alongside these could not be located on Forex Factory and is not used on this page.
- EliteTrader threads, all retrieved 28 July 2026. Thread 314065, Is walk forward / out of sample testing simply an illusion? — "if you try 10 different things even though you did not look at the out of sample data when cogitating it is very likely you are going to be fooled at least once." Thread 290496, Is backtesting necessary before walk-forward analysis? — "If you backtest first the performance of your WFA will be massively overstated, since you're only testing stuff that you know already works." The body of this page expands "WFA" to "walk-forward analysis"; the original wording is the one given here. Both are individual traders' posts in a public forum — practitioner reports, not studies.
- TradeStation, Walk-forward Test Results, help.tradestation.com (retrieved 28 July 2026) — "Using the default settings, a trading strategy passes a walk-forward analysis if … it shows an even distribution of profit, i.e. no individual time period contributed 50% or more of total net profit." The 50% is a default the user can change, and it is a different 50% from the walk-forward efficiency threshold listed beside it on the same page.
- Robert Pardo, The Evaluation and Optimization of Trading Strategies, 2nd edition, Wiley, 2008 — the walk-forward methodology that the out-of-sample discussion on this page assumes.
- Arithmetic used on this page and requiring no source: the number of configurations is the number of values tried for each parameter, multiplied together once per parameter (10 values for each of 3 inputs gives 1,000); the false-positive calculation 1 - 0.95 to the tenth power = 40.1%, which assumes independent tests and is therefore given on the page as a floor rather than an estimate, because sequential holdout tests are not independent; and n = (t / Sharpe)² for the trades-required table, every cell of which is exact except the 45, rounded up from 44.4.
- Tradelyze implementation, reviewed 14 September 2026, with the minimum trade counts and the Trade Count tooltip's wording (fewer than 30 trades makes other metrics unreliable) re-checked against the code on 15 September 2026: the Best Metrics card and its Profit and Trade Count tiles, the Recommended Parameters section, the Parameter Search Space table, parameter gating and its self-check backtest, and the minimum trade counts of the permutation test (20 trades) and the deflated Sharpe check (5 trades). Also re-checked on 15 September 2026, and again on 25 September 2026: a check refused for too few trades stays in the robustness score at full weight and scores zero, and the robustness code gives the sign-pattern count (2 to the power n, so 32 patterns at five trades) as the reason for the permutation test's 20-trade minimum. Robustness scoring, reviewed 26 September 2026: the four scored robustness checks and their 29, 29, 24 and 18 points; the minimum backtest length shown for information only; any failed scored check, a refusal for too few trades included, or less than 30 days of data or no measurable span, multiplies the points by 0.69, so the score is at most 69, C+ and MARGINAL at best; a refused check's Too Few Trades badge; the 30-day minimum history for a ROBUST verdict, counted over the whole upload; the robustness checks run on the search's overall best settings, with an amber banner on a firm's card when that firm was recommended other settings; the refused-permutation example on this page was computed with Tradelyze's scoring code and is given rounded down, as the card shows it. Recommended Parameters, reviewed 26 September 2026: they are picked separately for each prop firm from the one search's trials, graded against that firm's rules and ranked like Top Trials, so they can differ from the search's overall best; the Recommended Parameters section and the Parameter Search Space table show one firm's settings, and each prop firm card lists its own. Walk-forward and held-out test, reviewed 26 September 2026: the walk-forward method is chosen by the deployment, either re-tuned each window or one run, split by period, the default for eligible strategies; one run, split by period, measures the search's overall best settings, with a banner naming any firm recommended other settings; parameter stability was retired on 26 September 2026; the held-out test withholds the last 25% of the history from the whole search by default, runs each firm's Recommended Parameters once over the full history, which Best Metrics then describe except while a run is in progress or if that run fails, and is skipped below 200 bars. Parameter gating is not a published method and there is no citation behind it. The Best Value column is methodology guidance: reading a value at the edge of its range as a search that wanted to go further has no primary source.
- No exchange rate between correlated and independent trades is quoted on this page, because none is published. The claim that correlated trades are worth less than independent ones is qualitative on purpose.
- The 30, 100 and 200–500 trade thresholds are conventions with no primary source. They are reported here as conventions because they are what practitioners use, not because a study established them.