Glossary

Walk-forward analysis

Last reviewed: 26 September 2026·Tradelyze

Walk-forward analysis is a test of whether the settings your optimizer picked still work on price data it never saw. The history is cut into windows. In each one, the optimizer tunes the inputs on the first part, then those exact inputs are scored once on the next part. Repeating this across several windows shows whether the edge survives.

In plain English

Walk-forward analysis works like a string of practice exams. The optimizer is the tool that tries many combinations of a strategy's settings and keeps the best. It studies one stretch of past prices. Then it is graded on the stretch that came next, which it was never allowed to see. Settings that keep passing exam after exam are less likely to be a lucky fit to old prices. An edge, on this page, means a repeatable reason a strategy makes money rather than luck.

New to this? Start with in-sample vs out-of-sample testing.

This page covers the procedure

Walk-forward analysis is the process. Walk-forward efficiency is a single headline number that some implementations compute from its output. This page covers window design, split ratios, window counts and what the procedure does and does not establish. The efficiency ratio, what counts as a good value and how the ratio breaks when returns are negative are covered on the walk-forward efficiency page.

What is walk-forward analysis?

Walk-forward analysis is a procedure for testing whether an optimized trading strategy still works on data the optimizer was not allowed to see. One walk-forward window has two parts. The in-sample or training segment is the stretch of history the optimizer may study: it searches the possible settings there and picks a winner. The out-of-sample or test segment is the stretch of bars right after it. The winning settings run there exactly once, with no further tuning.

A walk-forward analysis is a sequence of those windows, each one advanced further through history than the last. You will also see the same procedure called walk-forward testing. Some writers, especially in machine learning, call it walk-forward validation.

The procedure comes from Robert Pardo. He presented walk-forward analysis in Design, Testing, and Optimization of Trading Systems (Wiley, 1992). He developed it at much greater length in the second edition, retitled The Evaluation and Optimization of Trading Strategies (Wiley, 2008), Chapter 11. That later version is the one modern implementations usually cite.

What separates walk-forward analysis from a simple holdout test is repetition. A holdout test keeps back one stretch of history and scores the strategy on it once. That gives one out-of-sample number, and the number depends heavily on the market regime that fell in that stretch. A market regime is a period when prices behave in one consistent way, such as a steady trend or a choppy range. Walk-forward analysis produces one out-of-sample result per window. So it can answer a question a single holdout cannot: was the edge present in several different stretches of history, or in only one?

The defining property

In a walk-forward analysis, the parameters that generate each out-of-sample result were chosen before that segment was visible. That is the only thing the procedure guarantees. It does not guarantee that the parameters are good, that the search was honest, or that the analysis was run only once. Any one of those can undo it.

What does walk-forward analysis look like on 24 months of data?

This worked example applies Tradelyze's walk-forward defaults to 24 months of price history. That means 2 windows, each using 70% of its length for tuning and 30% for testing, on a rolling schedule. Rolling means every tuning stretch is the same length and slides forward in time. The 24 months are the history the walk-forward receives; with the held-out test on, that leaves out the last part of your upload, 25% by default. Constructed illustration, not measured data.

In plain words, Tradelyze makes both windows the same length. It picks that length so window 1's tuning stretch plus the two test stretches exactly fill the 24 months. The arithmetic is below; the general rule is under what train/test split to use.

window length = 24 months / (2 × 0.30 + 0.70) = 24 / 1.3 ≈ 18.5 months
tuning stretch = 18.5 × 0.70 ≈ 12.9 months
test stretch = 18.5 − 12.9 ≈ 5.5 months, which is also how far window 2 starts after window 1
Constructed illustration, not measured data: 2 rolling windows at 70% tuning over 24 months of history
WindowTuning stretch (in-sample)Test stretch (out-of-sample)
Window 1About months 1–13About months 13–18
Window 2 (starts about 5.5 months later)About months 6–18Months 19–24
Tradelyze's default two rolling walk-forward windows over 24 months of history Constructed illustration, not measured data. A timeline of 24 months of price history, marked at 0, 6, 12, 18 and 24 months elapsed, showing Tradelyze's default of 2 rolling walk-forward windows with 70 percent of each window used for tuning. Each window is about 18.5 months long. Window 1 tunes on the first 12.9 months and tests on the next 5.5 months, ending about 18.5 months in. Window 2 starts about 5.5 months after window 1, tunes from about 5.5 to 18.5 months elapsed, and tests on the final 5.5 months, ending at month 24. The two test stretches do not overlap, the two tuning stretches share about 7.4 months, and the first 12.9 months are never tested on unseen data. Constructed illustration: 2 rolling windows, 70% tuning, 24 months of history Each window is about 18.5 months long and steps forward by one 5.5-month test stretch. Window 1 tune ≈ 12.9 months test ≈ 5.5 months Window 2 tune ≈ 12.9 months test ≈ 5.5 months 0 6 12 18 24 months of price history elapsed Tuning (in-sample): the optimizer picks settings here Test (out-of-sample): settings frozen, scored once
Constructed illustration, not measured data. Tradelyze's default of 2 rolling windows at 70% tuning, drawn over 24 months of history. Each window is about 18.5 months long and steps forward by one 5.5-month test stretch. Only the last 11 months or so are ever tested on data the optimizer did not see.

Three things follow from the 24-month example. Only about 11 of the 24 months are ever tested on unseen data, because the first 13 months exist only to tune window 1. The two tuning stretches share roughly months 6–13, so the two windows are not fully independent. And with two windows, one favorable test stretch is half of all the evidence. Tradelyze sizes windows in bars rather than months, so real boundaries fall on whole bars.

Why this matters for your money: when the label beside Tradelyze's walk-forward badge reads Re-tuned each window, a Confirmed badge in this example rests on roughly 11 months of unseen data. That data is split across two test stretches. When the label reads One run, split by period, the same windows are used, but none of the 24 months is unseen: the settings were chosen on all of them (see the two methods). Before you trade a strategy or pay for a prop firm challenge because of a Confirmed or Consistent badge, add history or windows. Then read the Per-Window Results table one row at a time.

Rolling vs anchored walk-forward: which should I use?

Walk-forward analysis runs on one of two schedules, rolling or anchored. The only difference is whether the training segment grows or slides forward. Neither schedule is better in general. Both are in wide use, both are defensible, and they produce different results on the same data.

An anchored walk-forward always starts training at the first bar. Each new window extends the training set further to the right, so the training set expands and no historical bar is ever discarded. This is also called an expanding window, which is the usual term in academic work.

A rolling walk-forward keeps the training set a fixed length and moves the whole thing forward. Old bars fall out of the training set as new ones enter it. The trading platforms AmiBroker and TradeStation both call this the non-anchored mode. So the same schedule appears under three names: rolling, sliding and non-anchored.

Anchored versus rolling walk-forward window schedules Two timelines, each showing three walk-forward windows drawn across the same span of price history from the oldest bar on the left to the newest bar on the right. In the anchored schedule the in-sample training segment of every window begins at the first bar and each window's training segment is longer than the last, and each is followed immediately by a fixed-length out-of-sample test segment that starts where the previous window's test segment ended. In the rolling schedule the in-sample training segment stays exactly the same length in every window and slides to the right by one out-of-sample block per window, so training segments of adjacent windows overlap, and each training segment is followed by a fixed-length out-of-sample test segment. A bracket under the rolling timeline marks the step between the start of window one and the start of window two as equal to one out-of-sample block. Anchored (expanding) — training always starts at the first bar The training set grows with every window. The oldest bars are never discarded. W1 in-sample out-of-sample W2 in-sample out-of-sample W3 in-sample out-of-sample oldest bar newest bar Rolling (sliding, non-anchored) — training stays a fixed length Each window steps forward by one out-of-sample block, so training segments overlap. W1 in-sample out-of-sample W2 in-sample out-of-sample W3 in-sample out-of-sample oldest bar newest bar step from one window to the next = one out-of-sample block In-sample: the optimizer fits here Out-of-sample: parameters frozen, scored once
The anchored schedule accumulates history and weights the distant past permanently. The rolling schedule discards it and keeps every training set the same size. Neither is correct in general; they encode different assumptions about whether market structure from years ago still applies.

The choice is an assumption about the market, not a technical detail:

  • Anchored assumes the distant past still applies. Every window trains on all of it, so a regime from years ago keeps influencing parameter selection forever. Later windows train on far more data than earlier ones, so the windows are not comparable with each other. In a five-window anchored run, window 5 can train on three times the data of window 1.
  • Rolling assumes recency matters. Every training set is the same length, so windows are comparable and the analysis adapts to structural change. The cost is that the parameter set is fitted to less data, and a rolling analysis will happily re-fit itself into every passing regime.

Tradelyze uses rolling windows by default, and its Walk-Forward & Robustness settings offer no anchored option. Tradelyze's rolling windows step forward by exactly one test segment at a time. So the test segments follow one another through the history without overlapping. When a run does use anchored windows, every window trains from the first bar and the test segments divide the rest of the data.

The Approach tile on the Tradelyze walk-forward card names the schedule a run used: Rolling (sliding window) or Anchored (expanding training). Unknown on that tile means only one window was built, which is thin evidence: the badge then always reads Inconclusive, because a verdict needs at least two usable windows.

The part that is usually left out

Tradelyze works out the Approach tile from the windows a run actually built, not from the settings that were requested. On very short data Tradelyze can build fewer windows than requested. Read the Approach and N Windows tiles on the result rather than assuming the request was honored.

What train/test split should I use?

There is no agreed train/test split for walk-forward analysis. A bigger test share gives more unseen data to judge on but less history to tune on, so no split is right on its own. The disagreement is between named, traceable sources rather than a gap in the literature. Every value in this table is a published recommendation from somebody; none of them is a research finding.

Published train/test split recommendations for walk-forward analysis, and where each comes from
Split (in : out)SourceWhat the source actually says
90:10 to 75:25 TradeStation walk-forward documentation The out-of-sample portion should be 10–25% of each run, and "research suggests a ratio of between 3 and 9 to 1" of in-sample to out-of-sample. The two statements agree: 3:1 is 75:25 and 9:1 is 90:10.
80:20 TradeStation walk-forward optimizer The shipped default, sitting in the middle of that vendor's own 10–25% band.
70:30 Tradelyze default. Also a QuantConnect community forum post — not vendor documentation Tradelyze uses 70/30 by default, which is outside TradeStation's recommended band. A 30% out-of-sample share is more than the 25% ceiling that vendor gives. The 70/30 often credited to QuantConnect comes from a community forum post; QuantConnect's own documentation gives no split.
Anything else No primary source. Splits quoted without attribution are conventions inherited from one of the other rows in this table. Treat them accordingly.

On the Tradelyze walk-forward card, the Train / Test Split tile shows how each window was divided between tuning and testing. Tradelyze uses 70% / 30% by default. The tile reports the split the windows actually got. On short data that can differ from the one requested, because Tradelyze sets minimum sizes for both segments. If the Walk-Forward & Robustness settings are enabled for your account, the Training percentage slider changes the split for future runs.

The split does not act alone. It fixes the ratio between the two segments. The segment lengths in bars come from the split, the total history and the number of windows together.

In plain words, Tradelyze gives every window the same length. It picks that length so the first training segment plus every test segment, laid end to end, fill the history. Each window then trains on its first share and tests on the rest, and the next window starts one test segment later. Written out for a rolling schedule:

window length = total bars / ( windows × (1 − training share) + training share )
training length = window length × training share
test length = window length − training length
step = test length

That is the sizing Tradelyze uses. It sets floors of 50 bars of training and 10 bars of testing, so a window can never collapse to nothing. Read it as a constraint rather than a formula to memorize. Raising the out-of-sample share shrinks the training segment, and raising the window count shrinks both. The split is one of three dials and cannot be chosen in isolation from the other two.

A defensible way to pick: start from the number of trades you need in each out-of-sample block (see how much data and how many trades). Then choose the split and window count that deliver it. Choosing a split because it is the default, then finding four trades in each test block, is the common failure.

How many walk-forward windows should I run?

Run enough walk-forward windows to smooth out luck, but not so many that each test window holds too few trades. More windows mean more separate tests of the same settings, so a verdict resting on one or two windows is thin evidence. The floor and the ceiling on window count come from different places. On a fixed amount of history they compete directly.

The floor is about noise. TradeStation's walk-forward documentation sets a minimum of 5 runs "to overcome random results". It also says that "a walk-forward analysis including at least 10 walk-forward tests approaches such reliability". With two or three windows, one unusually favorable out-of-sample segment can carry the entire aggregate. The aggregate then reports which segment happened to be lucky, not anything about the strategy.

On the Tradelyze walk-forward card, the N Windows tile shows how many train-then-test windows a run actually built. The Total Bars tile beside it counts the bars those windows covered, not the number of bars you uploaded. The walk-forward never sees the part of your history held back for the held-out test, the last 25% by default. What the default 2 windows cover on 24 months of history is drawn out in the 24-month walk-forward example.

The part that is usually left out

Tradelyze uses 2 walk-forward windows by default. Two is less than TradeStation's stated 5-run minimum. With two windows, one favorable out-of-sample segment is half the evidence. Both windows must produce a usable result before any verdict is given; otherwise the badge reads Inconclusive. Both must then make money, so one losing window turns the badge to Not Confirmed, or Not Consistent under One run, split by period.

If the Walk-Forward & Robustness settings are enabled for your account, the Walk-forward windows field sets a higher count for future runs. Otherwise that field is read-only and runs use the platform defaults. Under re-optimized walk-forward, every extra window runs another search, of 21 to 51 trials when the Walk-forward trials field is on its automatic setting, so runs take longer. Under One run, split by period, extra windows cost no extra backtests. On short data Tradelyze can build fewer windows than requested. Whatever you requested, read the N Windows tile on the result rather than assuming a count.

The ceiling is about sample size. History is fixed. Every additional window makes each out-of-sample block shorter and puts fewer trades in it. TradeStation's documentation puts roughly 30 trades on each out-of-sample run. With fewer than that, a block's win rate, return and Sharpe ratio swing so much from luck that they stop describing the strategy.

The window-count tradeoff: aggregate reliability against trades per out-of-sample block A combined bar and line chart with the number of walk-forward runs on the horizontal axis at 2, 5, 8, 10, 15 and 20 runs. The bars, read on the left axis, show how many trades land in each out-of-sample block when a fixed history of 600 trades is divided at a 20 percent out-of-sample split: 100 trades at 2 runs, 67 at 5 runs, 50 at 8 runs, 43 at 10 runs, 32 at 15 runs and 25 at 20 runs. A dashed horizontal marker at 30 trades shows the per-block floor, which the bars fall below between 15 and 20 runs. The rising line, read on the right axis, is a relative reliability index for the aggregate across runs, proportional to the square root of the number of runs and normalized to 1.0 at 20 runs: 0.32 at 2 runs, 0.50 at 5, 0.63 at 8, 0.71 at 10, 0.87 at 15 and 1.00 at 20. The region between 5 and 10 runs is shaded to mark the band TradeStation's documentation recommends. Both curves move at once and in opposite directions, so more runs buy a steadier aggregate and pay for it with a smaller sample inside every block. The bar heights are arithmetic from the split and the line is the square-root-of-n relationship, so neither is a measured result. Trades in each out-of-sample block (left axis) Reliability of the aggregate, proportional to √runs (right axis) TradeStation's recommended 5 to 10 run band 100 50 0 30 1.0 0.5 0 100 67 50 43 32 25 2 5 8 10 15 20 Number of walk-forward runs Bars: arithmetic from a fixed 600-trade history at a 20% out-of-sample split. Line: the √n relationship, not a measured result.
Both quantities move the moment you change the run count, and they move in opposite directions. The bar heights are arithmetic from the split, not an empirical finding. With 600 trades of history and a 20% out-of-sample share, each block holds 600 × 0.2 / (0.8 + 0.2n) trades. The line is the standard square-root-of-n relationship: the error luck puts into an average shrinks with the square root of the sample count. It is shown as a relative index rather than a quantity in any unit.

How to read the window-count chart: inside TradeStation's recommended 5–10 band, a 600-trade history leaves 43–67 trades in every out-of-sample block. That is comfortably more than the 30-trade floor. Push to 20 runs and each block falls to 25 trades. The aggregate looks more reliable while every input to it has become less so. That is the trap, and it is why "how many windows" has no single answer. The right count depends on how many trades your strategy generates over your history.

The honest procedure runs in the other direction. Count your total trades. Pick the out-of-sample share. Then compute how many runs leave each block at or over your per-block floor. If the answer is fewer than 5, the conclusion is that there is not enough data for a walk-forward analysis on that instrument and timeframe. That is a real result. Reducing the run count until the per-block numbers look acceptable produces an analysis that measures nothing.

Check this on your own results

Open the Per-Window Results table and look at every window's test stretch, not just the totals. Suppose the aggregate rests on one window with 90 trades and four windows with 6 trades each. That is not a five-window analysis in any useful sense.

Tradelyze excludes windows that produced no usable result and counts them on the Excluded Windows tile. A window is excluded, for example, when the settings being scored placed fewer than five trades on its tuning stretch. The badge reads Inconclusive, with no verdict, when more than half the windows are excluded or fewer than two produced a usable result.

The test stretch has a floor too. If any usable window placed fewer than five trades on its test stretch, the badge reads Inconclusive. That window still counts in the tiles, though. The Windows Profitable tile shows a count over the windows that were not excluded, such as 1 of 2, with the number the badge needs beneath it.

How much data and how many trades do I need?

The requirement is not one number, because a walk-forward analysis has to satisfy a per-block floor and a whole-analysis floor at the same time. The per-block floor is about 30 trades in each out-of-sample run, from TradeStation's walk-forward documentation. That is the binding constraint in practice, because it applies to every block individually rather than to the total. No published figure sets a floor for the training segment on its own; the in-sample number below falls out of that per-block floor and the split.

Combining the run-count guidance with the per-block floor gives a concrete total, and TradeStation publishes that total rather than leaving it to be derived. At a 20% out-of-sample split, 30 trades in each test block implies 150 trades per window, 120 of which are training. A rolling schedule that steps forward by one test block covers the first training segment once, then one test block per run:

per-window trades = 30 / 0.20 = 150 (120 in-sample + 30 out-of-sample)
10-run total = 120 + (10 × 30) = 420 trades

So a 10-run walk-forward analysis at a 20% split needs roughly 420 trades. That total is TradeStation's own published figure, and so is its split into 120 in-sample plus 300 out-of-sample. It is not an inference drawn on top of their other numbers.

The two variants either side of it are ours. Applying the same construction at a 30% out-of-sample share gives 70 + 300 = 370, and at 10% it gives 270 + 300 = 570. TradeStation publishes no total for those splits, so treat 370 and 570 as an extension of their method rather than as their recommendation.

Four hundred-odd trades is a substantial requirement. A strategy taking two trades a week produces about 100 a year, so a 10-run analysis wants roughly four years of history. That is before any allowance for four years containing only one or two genuinely distinct market regimes. This is the point at which many walk-forward analyses should be abandoned rather than shrunk to fit.

Does walk-forward analysis prevent overfitting?

No, not on its own. Overfitting means tuning a strategy until it fits past noise instead of a repeatable edge; overfitting and sample size covers it in full. Walk-forward analysis constrains one kind of overfitting: fitting parameters to the very bars you then report performance on. It leaves the more damaging kind untouched.

The mechanism is repetition. Out-of-sample data decays into in-sample data every time you look at it. You run the analysis, the out-of-sample result disappoints, you widen a parameter range or add a filter, and you run it again. The second run's out-of-sample segments are the same bars as the first run's, and they have now participated in choosing the strategy. Nothing in the procedure notices this, because each individual run is still formally clean.

An EliteTrader contributor puts the statistics plainly. If you try 10 different things, "it is very likely you are going to be fooled at least once". That is the multiple-comparisons problem in one sentence: the more versions you test, the more likely one looks good by chance. It applies to walk-forward analyses exactly as it applies to plain backtests. The out-of-sample segment protects the first run. It does not protect the tenth.

The part that is usually left out

A walk-forward result is only as clean as the number of times the analysis was run. Record how many walk-forward analyses you ran before the one you are reporting, and report that count next to the result. A first-attempt walk-forward result and a fortieth-attempt one with the same numbers are not equal evidence. No metric computed from the run can tell them apart.

There is a second limitation, and it is structural rather than procedural. Joubert, Sestovic, Barziy, Distaso and López de Prado classify walk-forward as a historical-path backtest in The Three Types of Backtests (2024, SSRN 4897573). They state the limitation directly: only a single path is tested. A historical-path backtest replays the one sequence of prices that actually happened. The strategy's behavior on paths that did not occur is not examined. It cannot be, because a walk-forward analysis has exactly one sequence of segments to work with.

Bailey, Borwein, López de Prado and Zhu made the underlying point earlier. Their paper is Pseudo-Mathematics and Financial Charlatanism (Notices of the American Mathematical Society 61(5), May 2014). It shows that, with enough trials, an impressive backtest can be produced from random data. Holding data out does not undo that when the held-out data has itself been used, indirectly and across many runs, to choose the winner.

What walk-forward analysis does contribute is real and worth keeping. Within a single run, it makes fitting parameters to the reported period impossible. It also produces several out-of-sample results instead of one, which exposes strategies that only worked in one stretch of history. It is a necessary check, not a sufficient one. The robustness score page covers the tests that address the multiple-testing side.

Which parameters do I trade after the analysis finishes?

Many walk-forward guides stop before this step, even though it is the only step that produces something you can trade. A walk-forward analysis with n windows has produced n different winning parameter sets and no instruction on what to do with them.

Two selection approaches are in use:

  • Most frequent across runs. Take the value each parameter converged on most often across the training windows. The argument: a value that keeps winning on different training data is more likely to be real than a value that won once. The cost is that it ignores recency entirely. It can also produce a combination no single window chose, because each parameter's most common value may come from a different window.
  • The current or latest run. Take the winner from the most recent training window, on the argument that it was fitted on the data most similar to what comes next. The cost is that it is a single optimization on a single segment, which is precisely the thing walk-forward analysis exists to be skeptical about.

Then there is the step both approaches usually end with: re-optimize on all available data. Once the walk-forward analysis has shown that the strategy survives out-of-sample, a final optimization is run over the entire history. The argument is that the parameters you trade should have seen every bar you have.

Tradelyze follows this pattern, with one exception. The settings on the Recommended Parameters card come from one optimization over all of your history except its last part, which is held back for the held-out test: the last 25% by default. The walk-forward analysis acts as a check on whether that optimization is trustworthy, not as the source of the settings. How that search works, and why the number of settings it tries matters, is explained on the strategy optimization page.

What the final re-optimization costs you

The parameter set you end up trading has no out-of-sample evidence behind it at all. It was fitted on all the data, including every bar the walk-forward analysis used as out-of-sample. The walk-forward result validated a procedure: "optimizing this strategy on this data tends to produce parameters that transfer". You are relying on that claim about the procedure, not on any test of the specific numbers you are about to trade. That is a defensible position, but hold it knowingly.

Tradelyze's held-out test is there to close part of that gap. The held-back stretch is kept away from the whole search, including the walk-forward and robustness checks. Each firm's recommended settings are then run once over your full history, and the trades they opened in the held-back stretch are the test. Its card reads Confirmed, Not Confirmed, Inconclusive or Pending, or Not run when the test is switched off or the data holds fewer than 200 bars; then nothing is held back and the search uses every bar. One held-back stretch is real evidence about the exact settings you are given, not proof.

What does the Per-Window Results table show?

The Per-Window Results table gives one row per walk-forward window. Each row shows how the settings used in that window performed on its tuning stretch and its test stretch, in columns marked Train and Test; click a row to see the settings themselves. Under Re-tuned each window the test stretch is unseen. Test results that stay reasonably close to the tuning results across rows are a good sign, and strong tuning results beside weak test results mean the settings were fitted to noise. Under One run, split by period the columns read Earlier and Later instead, and neither half is unseen: a weaker later half can simply mean a weaker stretch of market, not overfitting.

The table carries more information than the summary tiles, because it shows whether the result held in every window or in only one. Three things to read from it, in order:

  1. Train against test, per window. A window whose training metrics are excellent and whose test metrics are poor is an overfit window. One such window is information; a majority of them means the aggregate is not worth reading.
  2. The settings, down the table. Open each row to see the settings that window used. Under Re-tuned each window, values that stay in a narrow neighborhood across windows are the evidence behind the most-frequent selection rule. Values that jump around are the diagnostic covered under parameter sensitivity. Under One run, split by period every row holds the same set, so there is nothing to compare.
  3. Excluded windows. Tradelyze excludes a window from the summary figures when it produced no usable result. That happens when the window was too small, no settings traded enough to be scored, or the evaluation failed. Its metrics are zeroed placeholders, not results. The Excluded Windows tile counts them; when that count is high the summary is unreliable regardless of what it says.

On the Tradelyze walk-forward card, a run is summarized in a row of tiles. They are Approach, N Windows, Train / Test Split, Total Bars, WF Efficiency, Mean IS Sharpe, Mean OOS Sharpe, OOS Profit, Windows Profitable and Excluded Windows. IS means in-sample and OOS means out-of-sample. When the method label reads One run, split by period, nothing on the card is out-of-sample, so five tiles are renamed: Retention Ratio, Mean Earlier Sharpe, Mean Later Sharpe, Later-Stretch Profit and Periods Profitable. The WF Efficiency tile is explained on the walk-forward efficiency page. The Per-Window Results table sits below the tiles.

Does walk-forward test the strategy or just the parameters?

Just the parameters. The consensus in StrategyQuant's community forum is that a walk-forward analysis tests the robustness of one specific parameter set in one specific period. It does not test the strategy as an idea, or the parameter set in any other period.

This is a narrower claim than it first appears, and it has a corollary that changes how the results should be used. In a walk-forward analysis, the parameters are re-optimized at the start of every window. That is the procedure. So what the out-of-sample results describe is a strategy that gets re-optimized on the schedule the analysis used.

The corollary

Walk-forward results are only meaningful if you will actually re-optimize on the schedule you tested. An analysis with monthly test blocks measured a strategy that was re-fitted monthly. Trade one frozen parameter set for a year and you have not deployed the thing the analysis validated. You have deployed a different strategy, and nothing in the report describes its out-of-sample behavior.

Re-optimizing on the walk-forward schedule is easy to skip in practice. Skipping it opens a quiet gap between what the analysis measured and what you actually trade. If you intend to freeze parameters and leave them, the honest test is a single long holdout with frozen parameters, not a walk-forward analysis.

The narrowness also cuts in a useful direction. Because each window re-optimizes independently, the best parameters per window are evidence in their own right. A parameter that lands near the same value in every training window tells you something a jumping parameter does not, whatever the aggregate says.

What is the difference between re-optimized walk-forward and one run, split by period?

Re-optimized walk-forward and one run, split by period are the two methods Tradelyze uses to grade a walk-forward, and they answer different questions. Re-optimized walk-forward re-tunes the settings in every window. One run, split by period carries one fixed set of settings through every period of a single backtest. A Confirmed from the first is not the same claim as a Consistent from the second, so check which method produced a verdict before trusting it. A third method, parameter stability, was retired on 26 September 2026; what happened to it is covered at the end of this section.

The two methods exist because a re-optimized walk-forward result only describes a strategy you will actually re-optimize on the tested schedule. They ask different questions, they can disagree, and each supports a different way of trading.

The same window schedule, two different questions
 Re-optimized walk-forwardOne run, split by period
Label beside the badgeRe-tuned each windowOne run, split by period
What happens in each windowThe whole parameter search runs again on that window's training segmentNothing is searched. The history is backtested once, end to end, and each period is credited with what that one run did during its dates. Tradelyze checks that the periods add up exactly to the whole run
Whose parameters are scoredThe winner that window chose — a different set in each windowOne fixed set in every period: the search's overall best, which is the recommended set unless a banner on the card says otherwise
The question it answersDoes re-tuning this strategy keep working as time moves forward?Did these specific settings perform evenly across the history they were chosen on?
Out-of-sample?Yes. Each test segment was hidden from that window's searchNo. The settings were chosen on the same history, every period included
BadgesConfirmed or Not ConfirmedConsistent or Not Consistent
What a positive verdict supportsTrading the strategy and re-optimizing on the cadence testedThat the settings' results were not confined to one stretch of history. Nothing about data they never saw
CostOne search per window, each trial one backtest. The automatic budget is 21 trials when up to 5 inputs are searched, 31 for 6 to 15, 41 for 16 to 30 and 51 for more than 30No extra backtests

Both methods use the same rule. The verdict is positive when the efficiency ratio is above 0.5 and more than 60% of the usable windows made money. Under one run, split by period that ratio is the Retention Ratio: the later periods' average Sharpe ratio over the earlier periods' average, for the one fixed set. Either method can also read Inconclusive, when the stage ran but the evidence was too thin, and the card says why. NO VERDICT means nothing was asked or answered; it is not a failed check. The full rule is under what makes a walk-forward result pass or fail.

Why the method label matters

A re-optimized walk-forward verdict says nothing about the recommended settings printed beside it. No window ever ran them. That is not a criticism of the method; it is what the method is for. Reading Confirmed as approval of the recommended settings mistakes a verdict on a procedure for a verdict on a number.

A verdict from one run, split by period is the one that speaks to fixed settings. It is a weaker instrument in a different way. The settings were chosen on the whole history the periods are cut from, so Consistent shows only that their results were even across it. It cannot show that those results will carry over to data the settings never saw. For that, read the held-out test on each firm's card.

On the Tradelyze walk-forward card, a small label beside the badge names the method that produced the verdict. You do not choose the method. By default Tradelyze picks it for each run: one run, split by period when the strategy qualifies, and re-optimized walk-forward otherwise. So the label is the fact to rely on. That holds for older results too: a result keeps the badge words in use when it was produced, so an older one can read Confirmed or Not Confirmed, or PASS or FAIL, whatever its label. One run, split by period does not re-run each period, so a trade can carry across a period's start; only the profit or loss earned inside a period is credited to it.

One run, split by period measures one set of settings, the search's overall best. The settings recommended for each prop firm are picked against that firm's own rules, so they can differ. When they do, an amber banner at the top of the walk-forward card names those firms: “This check was run on the search's overall best settings. The settings recommended for <firms> are different, so for those firms this card describes other settings.” With a single firm it reads “that firm”. When a held-out test was run, it adds: “The held-out test on each firm's card uses that firm's own settings.” For a firm named there, read the walk-forward card as evidence about the strategy, not about the settings you were given.

What happened to parameter stability?

Until 26 September 2026 Tradelyze had a third method, parameter stability, labeled Fixed settings across periods. It asked the same question as one run, split by period: do fixed settings hold up across periods? Its period figures were cut from one backtest exactly as one run, split by period cuts them, but were treated as stand-ins for separate backtests of each period, so it re-ran one period as a genuine separate backtest to confirm them. That confirming re-run always landed on the first window's training part, so it never checked a testing figure, and the verdict rests on testing figures. One run, split by period answers the same question with a check that always runs, that the periods add up exactly to the whole run, at no extra cost. So parameter stability was retired on 26 September 2026.

Results produced before the retirement may still show Fixed settings across periods, with a Consistent or Not Consistent badge, or one of the older badge words noted above. Read them like one run, split by period: a verdict on fixed settings over the history they were chosen on, not a test on unseen data. The result fields that carry the fixed-settings verdict and its ratio, parameter_stability_passed and parameter_stability_efficiency, kept their names. One run, split by period now fills them, and under re-optimized walk-forward they are blank, because that method publishes its verdict and WF Efficiency in fields of its own.

Neither method replaces the other, and neither replaces a more direct measurement. If you intend to freeze parameters and leave them, a single long holdout with those frozen parameters tests them more directly than either. That is what Tradelyze's held-out test does with each firm's recommended settings, as described under which parameters to trade.

Why do walk-forward results differ from a plain backtest over the same period?

Walk-forward results differ from a plain backtest over the same dates because the two runs test different things. The differences are not rounding, and they are not a bug. Four separate mechanisms produce them, and the first is usually the largest.

  1. Different parameters are active in each segment. A plain backtest runs one parameter set across the whole range. A re-optimized walk-forward run uses a different parameter set in each test segment, because each was fitted on its own training window. The walk-forward equity curve, the running account balance, is several strategies stitched together, not a run of one. There is no reason for the two to agree. If they did agree, it would usually mean the optimizer picked the same parameters in every window.
  2. The spans are not the same. The first training segment is never scored out-of-sample by anything — it exists only to fit the first window. A walk-forward aggregate therefore covers the history minus that initial block, while the plain backtest covers all of it.
  3. Excluded windows shrink the span further. Excluded windows are dropped from the aggregate entirely. So a walk-forward result can silently be reporting on a subset of the periods a plain backtest covers.
  4. Aggregation is not compounding. Tradelyze's OOS Profit tile is the sum of the per-window percentage returns, not a compounded equity curve. Summing +1.2%, −0.5%, +0.3%, −1.8% and +0.1% gives −0.7%. An account that traded those five windows in sequence, with each return applied to the balance left by the last, would show about −0.72%. The gap is small here and grows as the returns get larger. A plain backtest reports the profit of one continuous account, so the two headline profit figures are different quantities and should not be compared directly.

The practical reading: a walk-forward result and a plain backtest over identical dates answer different questions. A gap between them is expected and is not evidence that either is wrong. The comparison worth making is a different one: the walk-forward out-of-sample result against the in-sample result of the same optimization. Walk-forward efficiency tries to summarize that comparison.

When should you not use walk-forward analysis?

Three cases, one of them stated explicitly by a vendor and the other two following from what the procedure is.

Strategies that pyramid or use variable position sizing. Pyramiding means adding to a position that is already open. TradeStation's walk-forward FAQ does not recommend walk-forward optimization for these strategies. Its reason is that profit and loss varies so much between runs that the results are skewed.

The reason is worth understanding rather than memorizing. When position size grows with account equity, a run that begins after a good stretch trades larger than one that begins after a bad stretch. The per-run results then reflect the order the windows fell in as much as the parameters. Comparing them as like-for-like is what breaks.

Strategies with no optimizable parameters. If there is nothing to fit, the training segment has nothing to do. The procedure collapses into a plain backtest chopped into pieces. That still shows how the strategy performed in each period, which is useful. But it does not test whether tuned parameters carry forward, because none were carried forward. Report it as segmented performance, not as a walk-forward result.

Histories too short to satisfy both window-count constraints at once. A walk-forward analysis needs enough runs to overcome noise, which TradeStation's documentation puts at 5 or more. It also needs enough trades in each test block, about 30 by the same documentation. When those two floors cannot both hold, running the analysis anyway produces a number with a respectable name and no content. Tradelyze applies a version of this as a hard gate. If more than half the windows are excluded, or any usable window placed fewer than five trades on its test stretch, the walk-forward badge reads Inconclusive regardless of what the surviving windows show.

One case that is sometimes listed here but should not be: strategies with few parameters. A two-parameter strategy is much less prone to overfitting than a twelve-parameter one. It can still be overfitted, though, and walk-forward analysis works on it normally. Parameter count changes how much you should worry, not whether the procedure applies.

How is walk-forward analysis different from forward testing and paper trading?

Walk-forward analysis runs on history that already exists, forward testing waits for new prices, and paper trading adds simulated live orders. The three are often confused, including by prop-firm educational content. That content sometimes presents walk-forward analysis as though it involved waiting for new data. It does not. The names are similar and the methods are not.

Walk-forward analysis, forward testing and paper trading compared
MethodWhen the test data came into existenceExecution modeledWhat it can catch that the previous method cannot
Walk-forward analysis Before the test was launched. Every out-of-sample bar already exists in the historical file. Simulated fills on historical bars Parameters fitted to the very bars being reported on
Forward testing After the strategy was frozen. The bars did not exist when the parameters were chosen. None — signals are recorded, no orders are sent Leakage from how the historical dataset was assembled, and any tuning that survived the freeze
Paper trading Live, in real time Simulated order routing with fills, slippage, partial fills and rejections Execution assumptions, latency and cost modeling

The single most useful distinction is time. Walk-forward analysis looks back and can run as soon as you have the data. Forward testing costs calendar time and cannot be sped up. That is exactly why walk-forward analysis exists, and exactly why it is weaker. Every out-of-sample bar it uses was already on disk and available to look at. The procedure relies on the honor system that you did not look.

Forward testing removes that honor-system dependency by putting the data on the other side of the freeze. Paper trading adds live order handling, which walk-forward analysis never sees. A strategy can pass a walk-forward analysis with every window profitable and still lose money live. The spread, slippage and real fills can do it, because none of them were part of what the analysis measured. The spread is the gap between the buying and selling price at any moment. Slippage is the gap between the price a backtest assumes and the price you actually get.

A note on terminology, since it causes search confusion: walk-forward analysis (WFA) and walk-forward optimization (WFO) are the same thing. Vendors and authors use the two terms interchangeably. Some writers say "optimization" when stressing the parameter search inside each window and "analysis" when stressing the whole procedure. That distinction is not consistent enough to rely on.

Going deeper

The sections below go deeper: how walk-forward analysis relates to cross-validation, and the purging, embargo and combinatorial methods built for financial data. You can skip them and still read your own report.

Is walk-forward analysis the same as cross-validation?

They are related, and a common shorthand treats walk-forward analysis as simply a type of cross-validation. That is a reasonable first approximation — both procedures hold data back, fit on the rest, and score on the held-back part — but it papers over the reason financial time series need their own machinery.

Standard k-fold cross-validation shuffles observations into folds and trains on k−1 of them while testing on the remaining one. On a time series that means training on bars that come after the bars being tested, which is lookahead by construction. Walk-forward analysis avoids this by never letting a training segment sit to the right of its own test segment.

Marcos López de Prado sets out the time-series repair in Advances in Financial Machine Learning (Wiley, 2018). Two mechanisms in Chapter 7 make cross-validation usable on financial data:

  • Purging. Training observations whose label windows overlap the test fold are removed. If a label depends on what happens over the following n bars, an observation n bars before the test fold already contains information from inside it.
  • Embargo. Training observations immediately after the test fold are also removed, because serial correlation leaks information backward across the boundary even when the label windows do not literally overlap.

Chapter 12 of the same book introduces combinatorial purged cross-validation (CPCV), and this is where the two approaches genuinely part company. Walk-forward analysis produces exactly one sequence of train/test segments: the one that history happened to lay out. CPCV forms many combinations of purged training and testing groups and reconstructs multiple out-of-sample paths from them, so the output is a distribution of out-of-sample results rather than a single figure.

Walk-forward analysis compared with cross-validation variants on time-series data
MethodRespects time orderOutputSource
Standard k-fold No — trains on bars after the test fold k out-of-sample scores from an invalid split General machine learning practice
Walk-forward analysis Yes, by construction One out-of-sample result per window along a single historical path Pardo, 1992 and 2008
Purged k-fold with embargo Yes, by removing overlapping and adjacent observations k out-of-sample scores with label leakage removed López de Prado, Advances in Financial Machine Learning, Ch. 7
CPCV Yes, with purging and embargo applied to every combination A distribution of out-of-sample paths López de Prado, Advances in Financial Machine Learning, Ch. 12

The distinction that matters: walk-forward analysis tests a single historical path, and CPCV yields a distribution of out-of-sample paths. A single path cannot tell you whether the result was typical or an outcome you were lucky to observe, because there is nothing to compare it against. That is not a flaw in how walk-forward is implemented; it is what the procedure is.

Stage 3 · step 12 of 18. Next in the learning path: Walk-forward efficiency

Check it on your own strategy

In a Tradelyze report, this is the walk-forward card: the verdict badge with the method label beside it, WF Efficiency (Retention Ratio under One run, split by period) and the Per-Window Results table. The badge reads Confirmed or Not Confirmed under Re-tuned each window, Consistent or Not Consistent under One run, split by period, and Inconclusive or NO VERDICT under either. If it reads Not Confirmed, Not Consistent, Inconclusive or NO VERDICT, see what to do when walk-forward fails or shows no verdict. To judge the whole report, not one tile, use the pre-trade checklist. Tradelyze re-runs an uploaded TradingView Pine Script strategy's backtest on your price data and checks it against your exported trades. It then runs parameter optimization, walk-forward analysis, a four-check robustness score and prop firm rule checks. It does not place trades, give financial advice or guarantee a challenge pass, and it is in beta.

Create an account. Already a user? Open your strategies.

Frequently asked questions about walk-forward analysis

What is walk-forward analysis?

Walk-forward analysis is a test of whether the settings your optimizer picked still work on price data it never saw. The history is cut into windows. In each window the optimizer tunes the inputs on the first part, and those exact inputs are then scored once on the next part. Repeating that across several windows shows whether the edge survives in more than one stretch of history.

Who invented walk-forward analysis?

Robert Pardo. He presented walk-forward analysis in Design, Testing, and Optimization of Trading Systems (Wiley, 1992) and developed it much further in the second edition of the same work, retitled The Evaluation and Optimization of Trading Strategies (Wiley, 2008), Chapter 11, which is the version most implementations cite.

What is the difference between anchored and rolling walk-forward?

Anchored walk-forward always begins training at the first bar, so the training set expands with every window and the oldest data is never dropped. Rolling walk-forward keeps the training set a fixed length and slides it forward, discarding the distant past. AmiBroker and TradeStation call rolling non-anchored; academic work usually says expanding and rolling. Tradelyze uses rolling windows by default.

What train and test split should I use for walk-forward analysis?

There is no agreed answer. TradeStation's documentation says the out-of-sample portion should be 10 to 25 percent of each run and ships a 20 percent default, noting that research suggests a ratio of between 3 and 9 to 1. Tradelyze defaults to 70/30, which is outside that band, because a 30 percent out-of-sample share exceeds TradeStation's 25 percent ceiling. No study establishes a split for trading strategies in general.

How many walk-forward windows do I need?

TradeStation's documentation sets the floor at a minimum of 5 runs to overcome random results, and says a walk-forward analysis including at least 10 walk-forward tests approaches such reliability. The ceiling is arithmetic rather than statistical: splitting fixed history into more runs shrinks every out-of-sample block, and TradeStation also says each block needs about 30 trades. Tradelyze uses 2 windows by default, below that floor.

How many trades does a walk-forward analysis need?

Enough that every individual out-of-sample block clears its own floor, not just the total. TradeStation puts roughly 30 trades on each out-of-sample run, and publishes about 420 trades as the total a 10-run analysis at a 20 percent split needs, decomposed as 120 in-sample plus 300 out-of-sample. A strategy taking two trades a week needs roughly four years of history to reach that total.

Is walk-forward analysis the same as cross-validation?

Not quite, though a common shorthand treats walk-forward as simply a type of cross-validation. Both hold data back, but standard k-fold shuffles observations and would train on bars that come after the ones it tests. Walk-forward preserves time order. Marcos López de Prado's purged k-fold with an embargo is the variant that keeps cross-validation valid on time series.

What is purged k-fold cross-validation with an embargo?

Purged k-fold cross-validation with an embargo is the time-series repair of k-fold cross-validation set out in Chapter 7 of Marcos López de Prado's Advances in Financial Machine Learning (Wiley, 2018). Purging drops training observations whose label windows overlap the test fold. The embargo additionally drops training observations immediately following the test fold, because serial correlation leaks information across the boundary.

Does walk-forward analysis prevent overfitting?

No, not on its own. Out-of-sample data decays into in-sample data through repeated iteration: every time you adjust the ranges and re-run, the held-out segments help choose the winner. As an EliteTrader contributor puts it, if you try 10 different things it is very likely you are going to be fooled at least once. Walk-forward constrains the fit; it does not undo selection bias.

Does walk-forward analysis test the strategy or only the parameters?

The consensus on the StrategyQuant forum is that walk-forward analysis tests the robustness of one specific parameter set in one specific period, not the strategy as an abstract idea. The practical corollary matters more than the semantics: walk-forward results describe a strategy that is re-optimized on the cadence you tested, so they only transfer if you actually re-optimize that often.

Which parameters should I trade after a walk-forward analysis?

Two selection rules are in use and they disagree. One takes the values that recur most often across the training runs, on the argument that a repeated winner is more likely to be signal. The other takes the most recent run's winner, on the argument that it saw the most recent regime. Most implementations then re-optimize over all available data before going live.

Why do walk-forward results differ from a plain backtest over the same dates?

Because a plain backtest runs one parameter set over the whole range while a walk-forward run uses a different parameter set in each test segment, so the two are not testing the same strategy. The reported spans also differ: the first training block is never scored out-of-sample, and windows that produced no usable optimum are dropped from the aggregate entirely.

When should I not use walk-forward analysis?

TradeStation's FAQ does not recommend walk-forward optimization for strategies that pyramid or use variable position sizing, because profit and loss then varies so much between runs that the comparison is skewed. Walk-forward analysis is also pointless for a strategy with no optimizable parameters, since there is nothing to re-fit and the procedure collapses into a segmented backtest.

Is walk-forward analysis the same as forward testing?

No. Walk-forward analysis runs entirely on historical data, because the out-of-sample bars already exist when the test is launched. Forward testing runs a frozen strategy on data that arrives after the freeze. Paper trading adds a simulated execution layer with fills and slippage. Prop-firm educational material conflates all three routinely; they catch different failures.

Sources

  • Robert Pardo, Design, Testing, and Optimization of Trading Systems, Wiley, 1992, ISBN 978-0-471-55446-2 — the earlier presentation of walk-forward analysis. The page range usually quoted for the 1992 book, 108–119, traces to Wikipedia rather than to the book, so it is not repeated on this page. The same goes for the claim that the book was the first published presentation of the technique anywhere.
  • Robert Pardo, The Evaluation and Optimization of Trading Strategies, 2nd edition, Wiley, 2008, ISBN 978-0-470-12801-5, Chapter 11 — the expanded treatment that modern implementations cite.
  • TradeStation Walk-Forward Optimizer help, retrieved 15 September 2026. Frequently Asked Questions — the 10–25% out-of-sample band, the note that "the default value of 20% is the recommended setting", "research suggests a ratio of between 3 and 9 to 1", the statement that "a walk-forward analysis including at least 10 walk-forward tests approaches such reliability", at least 30 trades per out-of-sample run, the 420-trade total for a 10-run analysis with its 120 in-sample plus 300 out-of-sample decomposition, and the recommendation against walk-forward optimization for pyramiding or variable position sizing. Perform Cluster Analysis — "a minimum of 5 Walk-forward runs are recommended to overcome random results".
  • QuantConnect community forum — the 70/30 train/test split commonly credited to QuantConnect appears in a user-submitted post still marked pending review, not in vendor documentation. QuantConnect's official documentation specifies no split ratio.
  • EliteTrader forum discussions. The thread Is Walk-Forward (out of sample) testing simply an illusion? (thread 314065) carries the argument that trying ten different things makes being fooled at least once very likely. Earlier versions of this page also carried an 85:15 split recommendation, an "at LEAST 100-150 trades in-sample" minimum and a "step size should cover at least 20 trades" floor, each attributed to EliteTrader discussions in general. None of the three could be located at source on 15 or 16 September 2026, so all three were removed on 16 September 2026 rather than left attributed to a post nobody can point at.
  • Marcos López de Prado, Advances in Financial Machine Learning, Wiley, 2018, ISBN 978-1-119-48208-6 — Chapter 7 for purged k-fold cross-validation with an embargo, Chapter 12 for combinatorial purged cross-validation.
  • Joubert, Sestovic, Barziy, Distaso and López de Prado, The Three Types of Backtests, 2024, SSRN 4897573 — walk-forward classified as a historical-path backtest that tests only a single path.
  • David H. Bailey, Jonathan M. Borwein, Marcos López de Prado and Qiji Jim Zhu, Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance, Notices of the American Mathematical Society 61(5), May 2014, beginning on page 458, doi:10.1090/noti1105.
  • StrategyQuant community forum — the position that walk-forward tests the robustness of specific parameters in a specific period.
  • Tradelyze implementation, reviewed 26 September 2026 — the default of 2 rolling windows with 70% of each window used for tuning; window sizing with 50-bar training and 10-bar testing floors; the Approach, N Windows, Train / Test Split and Total Bars tiles as measured from the windows actually built; the five-trade tuning floor and exclusion of windows with no usable result; the Inconclusive outcome when fewer than two windows are usable, more than half are excluded, a usable window placed fewer than five test trades, or a Sharpe ratio could not be measured; OOS Profit as a sum of per-window returns; the two walk-forward methods, Re-tuned each window and One run, split by period, with their badges (Confirmed or Not Confirmed, Consistent or Not Consistent, Inconclusive, NO VERDICT), their renamed tiles and table columns and their costs, including the automatic budget of 21, 31, 41 or 51 search trials per re-optimized window; the automatic choice between them; results keeping the badge words in use when they were produced; the settings banner on the walk-forward card; the held-out test of the last 25% of the history by default, not run on data under 200 bars or when switched off; and the retirement of parameter stability (Fixed settings across periods) on 26 September 2026.

Related terms

Tradelyze

Last reviewed 26 September 2026. Educational content about backtest validation methodology. Nothing here is financial advice.