Glossary
In-sample vs out-of-sample testing
Last reviewed: 26 September 2026·Tradelyze
Out-of-sample testing is checking a trading strategy on price data it was never tuned on, after its settings were chosen on separate in-sample data. In-sample results show how well the settings fit the past. Out-of-sample results show whether that fit carries over to data the strategy never saw, which is the question that matters before risking money.
In plain English
Split your price history in two. Tune the strategy on the first part, then run it once, unchanged, on the second part. The second result is the honest one, and it is usually weaker.
New to this? Start with what backtesting is.
What are in-sample and out-of-sample data?
In-sample data is the stretch of price history a strategy's settings were tuned on. Out-of-sample data is a separate stretch those settings never saw, used only to score them once they are frozen.
A backtest runs trading rules over past prices. A setting, also called an input or parameter, is a number the rules depend on, such as a moving-average length or a stop distance. Tuning means trying many values and keeping the ones that scored best. Software that does the tuning automatically is called an optimizer. Any stretch of data used during that tuning is in-sample, even a stretch you only looked at while designing the rules.
Every out-of-sample test is still a backtest. The only difference is which data the reported result comes from: data that helped choose the settings, or data kept away from that choice.
Here is a dated example. Constructed illustration, not measured data. A trader tunes a moving-average strategy on 2022–2023 futures data, freezes the winning settings, and then runs them once on 2024.
| Measure | In-sample: 2022–2023 | Out-of-sample: 2024 |
|---|---|---|
| What the data was used for | Choosing the settings | Nothing until the test; run once with the settings frozen |
| Trades | 140 | 66 |
| Return | 36.8% over two years, about 18.4% a year (simple average, not compounded) | 6.2% |
| Profit factor | 1.90 | 1.18 |
In this constructed example, the tuning years returned about 18.4% a year and the unseen year returned 6.2%. Roughly a third of the tuned performance survived (6.2 ÷ 18.4 = 0.34). Profit factor, gross profit divided by gross loss, fell from 1.90 to 1.18. The 2024 figures are the better guide to what the strategy can do. The 2022–2023 figures mostly show how well the optimizer fitted those two years. Measuring that survival ratio across repeated windows is what walk-forward efficiency does.
Why does performance usually drop out of sample?
Out-of-sample performance usually drops because the in-sample result is the best of many tries. Part of any best-of-many score is luck, and luck does not repeat on new data.
An optimizer is the software that tries combinations of settings and keeps whichever scored highest on the tuning data (see strategy optimization). Some of that winning score came from a real pattern and some from noise: random moves peculiar to that stretch of prices. The optimizer cannot tell the two apart, so it rewards both. On new data the noise is different, and the part of the score it supplied disappears.
The effect is large even with few tries. The Sharpe ratio is, roughly, average return above a risk-free rate divided by how much returns swing. A Sharpe ratio of 1 can look like a genuine edge. Bailey, Borwein, López de Prado and Zhu give an example in Notices of the American Mathematical Society (May 2014, page 461). It assumes seven independent strategy configurations, none with a real edge, on a two-year backtest. By their calculation the expected best in-sample Sharpe ratio is 1, while the expected out-of-sample Sharpe ratio is 0.
Markets also change. A pattern that paid in a trending stretch can fade in a choppy one. That can happen even when the strategy was not overfit, meaning tuned to noise instead of a real pattern. That is why a moderate drop out of sample is normal, while a collapse to zero or to a loss is the warning sign. The opposite can happen too: a test stretch that happens to suit the strategy can beat the tuning stretch by luck of timing.
How much data should be held out of sample?
No universal rule sets how much data to hold out of sample. Published defaults disagree, and the useful amount depends on how many trades the test stretch will contain.
General-purpose software picks a number for convenience. According to its train/test split documentation, the scikit-learn machine-learning library tests on 25% of the data when the user does not choose a share. That is a software default for any kind of data, not a trading recommendation.
Tradelyze's held-out test keeps back the last 25% of your history by default. Its walk-forward analysis uses 70% of each window for tuning and 30% for testing by default. Trading-platform vendors and forum contributors recommend other splits, and the walk-forward train/test split comparison sets those recommendations side by side with where each comes from.
| Figure | Held out for testing | Source |
|---|---|---|
| scikit-learn train/test split when no share is chosen | 25% | scikit-learn train_test_split documentation. A software default for any kind of data, not a trading recommendation. |
| Tradelyze held-out test, default | 25% | Tradelyze implementation. The last 25% of your history is kept away from the whole search. On data under 200 bars nothing is held back. A product default, not a research finding. |
| Tradelyze walk-forward analysis, default split of each window | 30% | Tradelyze implementation. The other 70% of each window is used for tuning. A product default, not a research finding. |
| A split shown to be right for trading strategies in general | None | No primary source. |
Some researchers argue the split is not the main problem. Bailey and co-authors put it this way (Notices of the American Mathematical Society, May 2014, page 462): “Because the hold-out method does not take into account the number of trials attempted before selecting a model, it cannot assess the representativeness of a backtest.” A generous holdout does not undo a search that tried thousands of settings.
A practical way to decide: hold out enough time for the test stretch to contain a meaningful number of trades. Ideally, the test stretch also covers a different kind of market from the tuning stretch. A 70/30 split of one quiet year leaves a test stretch of under four months, which may hold only a handful of trades.
Is a single holdout as good as walk-forward testing?
A single holdout is one out-of-sample test on one stretch of history. Walk-forward testing repeats the tune-then-test step across several windows, so it shows whether the edge survived more than once.
A holdout is data set aside and tested once at the end. Its weakness is luck of timing: if the held-out year happened to suit the strategy, one good number proves little. Walk-forward analysis cuts history into windows. It tunes on the first part of each window and tests on the next part, giving one out-of-sample result per window.
Walk-forward testing is not strictly better. When the settings are re-tuned in every window, the test results describe the tuning procedure, not the one set of settings you end up trading. If you plan to freeze one set of settings, a single long holdout with those exact settings tests them more directly. Tradelyze's held-out test does this with each firm's recommended settings. The comparison of re-optimized walk-forward and one run, split by period covers that difference.
Window count matters as well. Tradelyze runs 2 rolling windows by default, so one losing window can decide the walk-forward badge, and if either window produces no usable result the badge reads Inconclusive. Rolling means every window has the same length and slides forward through time. An anchored window instead always starts at the beginning of the data (see rolling vs anchored walk-forward). In Tradelyze's implementation, each rolling window slides forward by the length of one test stretch. With the default two windows, the first window's test stretch becomes training data for the second window. The two tests are therefore not fully separate.
What quietly contaminates an out-of-sample test?
An out-of-sample test is contaminated when knowledge of the test data leaks into choosing the strategy. The most common leak is re-testing after looking at a poor result.
The sequence feels innocent: test on the held-out year, see a weak result, change a setting, test again. Each attempt taught you something about the held-out year, so after a few rounds that year is in-sample in practice. No single step was dishonest, and the protection is gone anyway.
Other common leaks:
- Picking the best of several strategies by their out-of-sample results. The test stretch becomes a selection stretch, which is what in-sample data is.
- Moving the test start or end date after seeing results. Choosing the dates that make the strategy look good is tuning by another name.
- Designing the idea on the same chart history later used as the test. A TradingView strategy adjusted by eye over five years of chart has no unseen data inside those five years.
- Rules that use information not available at the time. A signal that reads a bar's final value before that bar has closed is called look-ahead. It makes every stretch of history look better than live trading can be.
Reusing the same data both to choose a model and to judge it is called data snooping. Halbert White's paper A Reality Check for Data Snooping (Econometrica, 2000) is the standard reference on testing for it. The everyday fix is bookkeeping. Pick the test stretch before tuning, count how many times you look at it, and treat that count as part of the result.
One Tradelyze-specific point: Tradelyze's first backtest runs your script with its own default settings. That baseline is the run checked against your TradingView export, and its numbers appear on the Backtest vs TradingView card. The search does not start from those settings; its first 10 trials are random draws from your ranges. If those defaults were tuned in TradingView over the same history you upload, that baseline is already in-sample before Tradelyze does anything.
How many out-of-sample trades are enough?
No primary source sets a minimum number of out-of-sample trades. About 30 is a widely repeated rule of thumb, and more is better, especially across different market conditions.
| Figure | Trades | What it applies to | Source |
|---|---|---|---|
| Common rule of thumb | about 30 | Out-of-sample trades in a test | No primary source. |
| Tradelyze walk-forward minimum | 5 | Each window's tuning stretch, or the window is excluded; and each usable window's test stretch, or the badge reads Inconclusive | Tradelyze implementation. |
Out-of-sample stretches are short, so they hold fewer trades than the tuning stretch. At Tradelyze's default split, each test stretch is 30% of its window, under half the length of the 70% tuning stretch before it. Take a constructed 12-month window at that split. Its test stretch lasts about 3.6 months, so a strategy that trades twice a month gets about seven test trades.
Seven trades can easily be all winners or all losers by chance, so a profitable test stretch that small says little. The trade-count thresholds on the overfitting and sample size page explain where the common figures come from. They also explain why many trades from one kind of market count for less than their number suggests.
Check the test trades yourself
Tradelyze applies its five-trade minimum to both sides of a walk-forward window. A window with fewer than five trades on its tuning stretch is excluded. A usable window with fewer than five on its test stretch makes the badge Inconclusive instead of a verdict. With the default two windows, both must be usable for a verdict, so a Confirmed badge can rest on as few as ten out-of-sample trades. That is still a small sample. Check Excluded Windows, then add up the Test side of the Trade Count column in the Per-Window Results table before trusting the badge. A test window with no trades at all counts as unprofitable in Windows Profitable.
Which Tradelyze results are in-sample and which are out-of-sample?
Most Tradelyze results are in-sample. Recommended Parameters and Top Trials come from the history the search saw. By default that is all of your history except the last 25%, which is held back for the held-out test. On data under 200 bars nothing is held back and the held-out test does not run. Best Metrics and the prop firm cards run the recommended settings over your full history, most of which the search saw.
Two results measure out-of-sample performance. The held-out test on each firm's card scores that firm's recommended settings on the held-back stretch. The walk-forward card does so only when the method label beside its badge reads Re-tuned each window.
| Result shown | In-sample or out-of-sample | Why |
|---|---|---|
| Best Metrics card: Profit, Sharpe (Bar), Sharpe (Daily), Max Drawdown, Win Rate, Profit Factor, Trade Count | Mostly in-sample | The recommended settings run over your full history, most of which they were chosen on. Read them as optimistic. |
| Recommended Parameters | In-sample | Picked by the search over the history it saw. The held-out test is the one stage that runs these exact settings on data held back from the search. |
| Top Trials table | In-sample | Every trial is scored on the same history the search saw. |
| Prop firm cards: Qualifies or Not Feasible, Rule Results | Mostly in-sample | Each firm's recommended settings are picked by re-scoring the search's trials against its rules. Rule Results then describe those settings over your full history, most of which the search saw. |
| Robustness card | Not an out-of-sample test | Stress tests of the settings the search chose. None of its checks runs those settings on price data held back from the search. |
| Walk-forward card: Mean IS Sharpe, or Mean Earlier Sharpe under One run, split by period | In-sample | The average annualized Sharpe ratio on the first part of each window. Annualized means converted to a yearly figure. |
| Walk-forward card labeled Re-tuned each window: Mean OOS Sharpe, OOS Profit, Windows Profitable, WF Efficiency | Out-of-sample | Each window's own tuned settings are scored on that window's unseen test stretch. OOS Profit adds up the test windows' profit percentages rather than averaging them. |
| Walk-forward card labeled One run, split by period, or Fixed settings across periods on results from before 26 September 2026: Mean Later Sharpe, Later-Stretch Profit, Periods Profitable, Retention Ratio | Not out-of-sample | The settings were chosen on the same history the periods are cut from, every test stretch included. |
| Held-out Test card, one for each prop firm | Out-of-sample | That firm's recommended settings, run once over your full history. Only the trades opened in the held-back stretch are the test, and the search never saw that stretch. |
By default Tradelyze chooses the walk-forward method for each run automatically: One run, split by period when the strategy qualifies, and Re-tuned each window otherwise. So the label is the fact to check. Under One run, split by period, and under Fixed settings across periods on results from before 26 September 2026, a weaker later period means the later stretch of history was weaker. It is not proof that the settings were overfit.
Even a Re-tuned each window result tests the tuning procedure, not the Recommended Parameters. Each window scores the settings that window chose, and none of them is the set you are given. The held-out test is the check of the set you are given. Read the in-sample numbers as a ceiling, and the out-of-sample figures, with their trade counts, as the evidence. The explanation of the walk-forward badges covers what each badge requires.
This matters most before you pay a prop firm challenge fee or trade real money. A Qualifies verdict on a prop firm card is a mostly in-sample result: it shows the firm's rules were met over your full history, most of which the settings were tuned on. The prop firm rules and backtest metrics page explains what each rule checks. If the walk-forward badge reads Not Confirmed, Not Consistent, Inconclusive or NO VERDICT, see what to do when walk-forward fails or shows no verdict.
Where this appears in Tradelyze
In a Tradelyze report, Best Metrics are mostly in-sample. Walk-forward test windows count as unseen data only when the label beside the badge reads Re-tuned each window. The held-out test is unseen data for each firm's recommended settings. To judge the whole report, not one tile, use the pre-trade checklist.
Tradelyze re-runs an uploaded TradingView Pine Script strategy's backtest on your price data and checks it against your exported trades. It then runs parameter optimization, walk-forward analysis, a four-check robustness score and prop firm rule checks. It does not place trades, give financial advice or guarantee a challenge pass, and it is in beta.
Create an account. Already a user? Open your strategies.
Stage 3 · step 9 of 18. Next in the learning path: Strategy optimization
More questions about Tradelyze: Learn FAQ
Frequently asked questions about in-sample and out-of-sample testing
What does in-sample mean in backtesting?
In-sample means the stretch of price history a strategy's settings were tuned on. Any result measured on that same stretch shows how well the settings fit data they were chosen to fit, so in-sample results are the most optimistic figures a backtest produces. Data counts as in-sample even if you only looked at it while designing the rules.
What does out-of-sample mean?
Out-of-sample means price data that played no part in choosing a strategy's settings. The frozen settings are run on that data once, and the result estimates how the strategy behaves on prices it has never seen. Out-of-sample results are usually weaker than in-sample results, and the size of the drop is the useful information.
What is the difference between backtesting and out-of-sample testing?
Every out-of-sample test is a backtest; the difference is which data the result comes from. A plain backtest often reports performance on the same history the settings were tuned on, which is in-sample. An out-of-sample test reports performance only on history kept away from the tuning, so the strategy cannot have been fitted to it.
Why is out-of-sample performance usually worse than in-sample?
Because an optimizer keeps the best of many tries, and part of any best score is luck that does not repeat. Bailey, Borwein, López de Prado and Zhu showed in Notices of the American Mathematical Society (May 2014) that after only seven independent configurations with no real edge, the expected best in-sample Sharpe ratio is 1 on a two-year backtest, while the expected out-of-sample Sharpe ratio is 0.
What is a good split between in-sample and out-of-sample data?
No split is correct for every strategy. The scikit-learn library tests on 25% of the data when no share is chosen, and Tradelyze's walk-forward analysis tests on 30% of each window by default; both are defaults, not findings. Choose a split that leaves enough trades in the test stretch to mean something, ideally covering a different kind of market.
Can out-of-sample results be better than in-sample results?
Yes. A test stretch can happen to suit the strategy better than the tuning stretch did, for example a strongly trending year for a trend-following system. One better out-of-sample period is more likely luck of timing than a hidden edge, so read it the same way as a weak one: check how many trades it holds and whether other periods agree.
Is walk-forward analysis the same as out-of-sample testing?
Walk-forward analysis is a repeated form of out-of-sample testing. History is cut into windows; in each one the settings are tuned on the first part and scored on the next part. That gives several out-of-sample results instead of one, but when settings are re-tuned every window, the results describe the tuning procedure rather than one fixed set of settings.
What is data snooping?
Data snooping is using the same data both to choose a strategy and to judge it, so the judgment leans toward whatever happened to fit that data. Re-testing on a held-out year after every tweak is the everyday version. Halbert White's paper A Reality Check for Data Snooping (Econometrica, 2000) is the standard reference on testing for it.
Can I reuse my out-of-sample data after changing the strategy?
You can, but the data stops being out-of-sample. Once a poor test result has led you to change something, that result has shaped the strategy, just as tuning data does. After a few rounds the held-out stretch is in-sample in practice. Record how many times you have tested on it, and keep fresh data back for a final check.
How many out-of-sample trades do I need?
No primary source sets a minimum. About 30 trades is a widely repeated convention, and more is better, especially when the trades span different market conditions. A test stretch is usually much shorter than the tuning stretch, so check its trade count directly; in Tradelyze, add up the Test side of Trade Count in the Per-Window Results table.
Are Tradelyze's Best Metrics in-sample or out-of-sample?
Mostly in-sample. The Best Metrics card shows the recommended settings run over your full history. Those settings were chosen on all of it except the part held back for the held-out test, the last 25% by default, so read Profit, Sharpe (Bar), Sharpe (Daily), Max Drawdown, Win Rate, Profit Factor and Trade Count as optimistic. Compare them with the held-out test and the walk-forward card rather than reading them as a forecast.
Is Mean OOS Sharpe in Tradelyze always out-of-sample?
No. Mean OOS Sharpe is out-of-sample only when the label beside the walk-forward badge reads Re-tuned each window. When the label reads One run, split by period, the settings were chosen on the same history, every test stretch included, so the tile reads Mean Later Sharpe instead: a later stretch of history, not unseen data. Older results labeled Fixed settings across periods read the same way.
Is forward testing the same as out-of-sample testing?
Forward testing is out-of-sample testing on data that did not exist when the strategy was built: the frozen strategy runs on new prices as they arrive, usually on paper or a demo account. Historical out-of-sample tests are faster but easier to contaminate, because the test data already exists and can be looked at before the test. Neither replaces the other.
Sources
- David H. Bailey, Jonathan M. Borwein, Marcos López de Prado and Qiji Jim Zhu, Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance, Notices of the American Mathematical Society 61(5), May 2014, pages 458–471, doi:10.1090/noti1105; free full text from the American Mathematical Society at ams.org/notices/201405/rnoti-p458.pdf. The seven-configuration example is on page 461 and the hold-out quotation on page 462. DOI registration and the full-text file checked 15 September 2026.
- Halbert White, A Reality Check for Data Snooping, Econometrica 68(5), September 2000, pages 1097–1126, doi:10.1111/1468-0262.00152. DOI registration checked 15 September 2026.
- scikit-learn documentation, train_test_split: the test share is set to 0.25 when neither a test share nor a training share is given. Retrieved 15 September 2026.
- Tradelyze implementation, reviewed 26 September 2026: the held-out test keeps the last 25% of the history by default away from the search, the walk-forward and the robustness checks, then runs each firm's recommended settings once over the full history and scores the trades opened in the held-back part; it does not run on data under 200 bars; Best Metrics and the prop firm verdict describe the recommended settings over the full history; walk-forward defaults of 2 rolling windows with 70% of each window used for tuning; each rolling window slides forward by one test stretch; the walk-forward method is chosen automatically, One run, split by period when the strategy qualifies and Re-tuned each window otherwise; walk-forward test figures are out-of-sample only under the Re-tuned each window method; One run, split by period scores one fixed set of settings on periods of one continuous run of the history they were chosen on; parameter stability (Fixed settings across periods) was retired on 26 September 2026; OOS Profit is a sum across usable windows; a window with fewer than five trades on its tuning stretch is excluded, and a usable window with fewer than five trades on its test stretch makes the walk-forward badge Inconclusive; with the default two windows, both must be usable for a verdict; the first backtest runs the script's own default settings, the baseline checked against the TradingView export, and the search does not start from them.
- The 2022–2023 tuning and 2024 test example and its timeline figure, and the trade-count arithmetic for a 12-month window, are constructed illustrations, not measured data.