Glossary

Walk-forward efficiency

Last reviewed: 26 September 2026·Tradelyze

Walk-forward efficiency is how much of a strategy's performance survives on data the optimizer never saw: the yearly return on that unseen data divided by the yearly return on the data it was tuned on. Tuned on data that returned 60% a year and tested on unseen data that returned 30% a year, it scores 50%: half the edge survived.

In plain English

Walk-forward efficiency compares a strategy on the data it was tuned on with the same strategy on data it never saw. A high value means much of the edge carried over; a low or negative value means it mostly did not. On the Tradelyze walk-forward card, WF Efficiency compares Sharpe ratios instead of returns and shows a ratio: 0.50 means half the tuned edge survived.

New to this? Start with walk-forward analysis.

What is walk-forward efficiency?

Walk-forward efficiency asks one question: after the optimizer picked its favorite settings, how much of their performance was still there on prices it never saw? In a constructed example, settings that earned 60% a year on the tuning data and 30% a year on the unseen data that followed score 30 ÷ 60 = 50%, meaning half the edge survived.

Tradelyze's walk-forward card shows the same idea as a ratio of Sharpe ratios, so 0.50 there means the same as 50% here. A Confirmed badge needs WF Efficiency above 0.5 and more than 60% of windows profitable, once the card has checked there is enough evidence for a verdict. How Pardo's book and other platforms calculate it is covered under Going deeper.

A few terms, in plain words. The optimizer is the software that tries many combinations of a strategy's input settings and keeps the best-scoring one. In-sample data is the stretch of prices it tuned on. Out-of-sample data is the stretch that comes next, scored once with the settings frozen. A window is one tuning stretch plus the test stretch after it, and a walk-forward analysis marches several windows forward through history. Both returns are annualized, meaning converted to a yearly rate, so a long tuning stretch and a short test stretch compare fairly.

walk-forward efficiency = annualized out-of-sample return / annualized in-sample return

Each window gets its own walk-forward efficiency, and the windows in one run rarely agree.

In-sample versus out-of-sample return in three walk-forward windows, with each window's efficiency Constructed illustration, not measured data. A grouped bar chart of annualized returns in three walk-forward windows, with an in-sample bar and an out-of-sample bar for each window. Window 1: in-sample 60 percent, out-of-sample 30 percent, so walk-forward efficiency is 30 divided by 60, or 50 percent, and half the edge was kept. Window 2: in-sample 45 percent, out-of-sample 36 percent, so efficiency is 36 divided by 45, or 80 percent, and most of the edge was kept. Window 3: in-sample 50 percent, out-of-sample 10 percent, so efficiency is 10 divided by 50, or 20 percent, and most of the edge was lost. The three in-sample bars are similar, while the out-of-sample bars, and so the efficiencies, differ widely from window to window. In-sample (tuning data) Out-of-sample (unseen data) Efficiency (unseen ÷ tuning) 60% 30% 0% 60% 30% 45% 36% 50% 10% Window 1 30 ÷ 60 = 50% half the edge kept Window 2 36 ÷ 45 = 80% most of the edge kept Window 3 10 ÷ 50 = 20% most of the edge lost Bars show annualized return. Constructed illustration, not measured data.
Constructed illustration, not measured data. Each window's efficiency is its out-of-sample return divided by its in-sample return. The three tuning results look alike, but the share of the edge each window kept ranges from 20% to 80%. A single headline number can hide a weak window like window 3; how platforms combine windows is covered under combining walk-forward efficiency across multiple windows.

Walk-forward efficiency matters because an optimizer always finds settings that look good on the data it tuned on. The ratio shows how much of that result held up once the data could no longer be fitted to. A strategy that keeps little of its edge on unseen data is a weak candidate for a live account. The same goes for a paid prop firm challenge, however good the backtest looks.

How do you calculate walk-forward efficiency?

Calculating walk-forward efficiency for a single window takes four steps. Optimize on the in-sample segment and record the annualized return of the winning parameter set. Freeze those parameters. Run them once on the out-of-sample segment and record that annualized return. Divide the second by the first.

in-sample: 48.0% over 24 months → 24.0% annualized
out-of-sample: 6.0% over 6 months → 12.0% annualized

walk-forward efficiency = 12.0 / 24.0 = 0.50 = 50%
Worked walk-forward efficiency calculation for one window
SegmentLengthPeriod returnAnnualized return
In-sample (optimizer fits here)24 months48.0%24.0%
Out-of-sample (parameters frozen)6 months6.0%12.0%
Walk-forward efficiency12.0 / 24.0 = 50%

This example annualizes the simple way. The 48.0% earned over 24 months is halved to 24.0% a year. The 6.0% earned over 6 months is multiplied by four to 12.0% a year. Compounding the same returns instead gives an efficiency of 57.1%, not 50%. Neither method is wrong, but a report should say which one it used. The compounding arithmetic is under technical and historical notes.

Skipping annualization gives a wrong answer. In-sample segments are usually longer than out-of-sample segments. Dividing the raw returns in this example gives 6.0 ÷ 48.0 = 12.5%. That figure is the true 50% efficiency multiplied by the length ratio of 6 months to 24 months: 0.50 × 0.25 = 0.125. It mixes efficiency and segment length together and cannot be split back apart without knowing the lengths.

A real walk-forward run has several windows, each with its own efficiency. Platforms turn those into one headline number in different ways; see combining walk-forward efficiency across multiple windows.

What is a good walk-forward efficiency?

There is no agreed answer, and the disagreement is real rather than a matter of unresolved detail. Only two thresholds in circulation have a traceable primary source, and both come from one vendor.

Walk-forward efficiency thresholds and where each one actually comes from
ValueReadingSource
≥ 100% Out-of-sample annualized return matched or beat in-sample. TradeStation's Walk-Forward Optimizer grades its walk-forward efficiency criterion pass with distinction at this level, on the efficiency figure alone. TradeStation Walk-Forward Optimizer — user-editable default
≥ 50% The strategy retained at least half its fitted edge out-of-sample. TradeStation's pass threshold for the same criterion. TradeStation Walk-Forward Optimizer — user-editable default
0.7 / 0.5 / 0.3 Commonly repeated as excellent / good / acceptable bands. No primary source. Not in Pardo, not in any vendor manual. These bands circulate between blog posts and forum answers citing each other.
< 0 The two sides had opposite signs. Usually in-sample made money and out-of-sample lost it, but the reverse also produces a negative figure. See negative returns for all four sign cases
Positive, but
both sides negative
Meaningless. The sign cancellation makes a losing strategy read as a pass. See negative returns

TradeStation's walk-forward optimizer treats 50% as the pass threshold and 100% as the pass-with-distinction grade for its walk-forward efficiency criterion. Those two numbers are vendor policy, not a statistical result. TradeStation chose them, documented them, and ships them as defaults the user can edit. Every threshold in that optimizer is a field, not a constant.

The widely repeated 0.3 / 0.5 / 0.7 bands are weaker still, because no primary source for them exists. They do not appear in Pardo's Chapter 11 and they do not appear in vendor documentation.

The practical consequence is that walk-forward efficiency should be read as a comparative measure inside one platform and one methodology, not as an absolute score. On its own, a walk-forward efficiency of 62% says little. It says far more next to the same figure for nine other candidate strategies, computed the same way.

Tradelyze's labels for its WF Efficiency tile, and the source of their 0.5 line, are under what Generalizes well, Likely overfit and Inverted mean.

What do Generalizes well, Likely overfit and Inverted mean on the WF Efficiency tile?

Generalizes well, Likely overfit and Inverted — lost out-of-sample are the labels Tradelyze attaches to its WF Efficiency number, in the tooltip on that tile. WF Efficiency is the average annualized Sharpe ratio on the unseen test windows divided by the same average on the tuning windows. A Sharpe ratio is average return divided by how much returns swing. The tile shows a ratio rather than a percentage, so 0.50 means half the tuned edge survived.

  • Generalizes well: WF Efficiency above 0.5. More than half of the tuned edge carried over to unseen data.
  • Likely overfit: WF Efficiency from 0 up to and including 0.5. At most half of the tuned edge survived, which points to settings fitted to the tuning data.
  • Inverted — lost out-of-sample: WF Efficiency under 0. The strategy made money on its tuning data and lost on the unseen data, so the edge reversed rather than faded.
  • clipped — at least this extreme: added when the ratio reaches the end of Tradelyze's scale at 2.0 or -2.0. The tile then shows ≥ 2.00 or ≤ -2.00, and the true ratio could be much larger.
  • A blank tile (--): no ratio was measured. The most common cause is an average Sharpe ratio on the tuning windows of zero or below, which leaves no edge to keep.
How five walk-forward results would read on the WF Efficiency tile. Constructed illustration, not measured data
Mean IS SharpeMean OOS SharpeTile showsLabel in the tooltip
1.200.840.70Generalizes well
1.200.540.45Likely overfit
1.20-0.30-0.25Inverted — lost out-of-sample
0.100.35≥ 2.00Generalizes well, plus clipped — at least this extreme (the raw ratio is 3.50)
-0.40-0.60--Blank: the tuning average is not positive, so no ratio is computed

Take the second row as a worked example. The tuned settings averaged an annualized Sharpe ratio of 1.20 on their tuning data. On the test windows that followed, they averaged only 0.54. Dividing 0.54 by 1.20 gives 0.45, so a little under half of the tuned edge survived and the tooltip reads Likely overfit. That run also misses the first of the two conditions for a Confirmed badge, which needs WF Efficiency above 0.5.

Why the label matters for your money. A Likely overfit or Inverted reading means much of what the optimizer found did not carry over to data it never saw. Live trading, or a paid prop firm challenge, may then look worse than the backtest. Do not treat the Recommended Parameters as validated; what to do when a strategy fails validation covers the next steps. Generalizes well is a reason to keep checking, not to trade: read Windows Profitable, Excluded Windows and the robustness card before trusting it. A clipped or blank tile needs Mean IS Sharpe read first, as explained under how Tradelyze calculates walk-forward efficiency.

These cut-offs are Tradelyze's own. The 0.5 line matches the 50% pass threshold that TradeStation's Walk-Forward Optimizer ships as a user-editable default. Tradelyze uses the same 0.5 in its walk-forward badge. The split at 0 and the label wording have no primary source.

Why is my walk-forward efficiency over 100%?

A walk-forward efficiency above 100% means the annualized out-of-sample return exceeded the annualized in-sample return. TradeStation's own walk-forward FAQ addresses this question directly, which is a fair indication of how often it comes up.

A walk-forward efficiency over 100% is possible, and it is not automatically good. Three mechanisms produce it:

  • A favorable out-of-sample regime. The test segment happened to contain a strong directional move that the training segment did not. The parameters did not become better; the market became easier.
  • A small denominator. Walk-forward efficiency of 300% on an in-sample annualized return of 3% is arithmetic, not evidence. The ratio is unstable whenever the in-sample side is close to zero. Tradelyze caps its WF Efficiency at 2.0, the ratio form of 200%. A ratio of 3.0 shows there as ≥ 2.00, with the label clipped — at least this extreme.
  • Annualizing a short segment. A modest absolute gain in a six-week out-of-sample window becomes a large annualized figure. The shorter the out-of-sample segment, the more the annualization amplifies noise.

TradeStation's guidance for results above 100% is cluster analysis. Check whether similarly high figures from neighboring parameter sets and neighboring windows surround it. If the 100%-plus result sits inside a cluster of strong results, the region of parameter space is genuinely productive. If it stands alone with mediocre or negative results on either side, it is an aberration and should be treated as one.

Why does walk-forward efficiency break when returns are negative?

Walk-forward efficiency is a ratio, and a ratio carries no information about the sign of its inputs. When the in-sample and out-of-sample figures are both negative, the two minus signs cancel. The metric then returns a flattering positive number. This is the single most consequential failure mode of the measure, and many implementations do not document how they handle it.

Here is the failure in a hypothetical implementation that divides Sharpe ratios without first checking that the in-sample figure is positive. Tradelyze does run that check, as the gate in this section explains:

mean out-of-sample Sharpe = -0.588
mean in-sample Sharpe = -0.457

walk-forward efficiency = -0.588 / -0.457 = 1.2867 → displayed as 128.7%

Treat that as an illustrative calculation rather than a derived one. The two Sharpe figures are rounded to three decimals. A platform showing exactly these two numbers can legitimately display 128.6% instead, because it divides the unrounded values. The last digit is not the point. The sign is.

On TradeStation's scale, where 100% is graded pass with distinction, this strategy reports 128.7%. It lost money on a risk-adjusted basis in every single training window, then lost more in every single test window. The out-of-sample result is not 28.7% better than in-sample. It is 28.7% worse. The metric inverted its own meaning.

How two negative Sharpe ratios produce a false walk-forward efficiency pass Three panels. Panel one shows a Sharpe ratio number line with mean out-of-sample Sharpe at minus 0.588 and mean in-sample Sharpe at minus 0.457, both in the negative half. Panel two shows the division of minus 0.588 by minus 0.457 producing plus 1.2867, with both minus signs circled to show that they cancel. Panel three shows a walk-forward efficiency scale from zero to 150 percent, with bands for below pass, pass at 50 percent and above, and pass with distinction at 100 percent and above; a marker at 128.7 percent sits inside the pass with distinction band even though every window lost money. The Sharpe ratios shown are a constructed illustration, not measured results, and they are rounded to three decimals, so the percentage is illustrative arithmetic rather than a derived figure. 1 — Both walk-forward Sharpe ratios are negative -1.0 -0.5 0 +0.5 +1.0 losing profitable mean OOS Sharpe = -0.588 mean IS Sharpe = -0.457 2 — Dividing one negative by another cancels the signs -0.588 ÷ -0.457 = +1.2867 the two minus signs cancel — the result says nothing about whether either side made money 3 — The result lands in the pass-with-distinction band 128.7% below pass pass (≥50%) pass with distinction (≥100%) 0% 50% 100% 150% Reality: every training window and every test window lost money.
Constructed illustration, not measured data, and the two Sharpe ratios are rounded to three decimals, so 128.7% is illustrative arithmetic rather than a derived figure. The ratio is computed correctly. The interpretation attached to it is what fails, because "fraction of edge retained" has no meaning when there was no edge to retain.

The sign problem is not confined to that one case. There are four sign combinations, and walk-forward efficiency was defined for only one of them. A reader who knows only the ratio cannot tell which of the four produced it.

All four sign combinations, and what walk-forward efficiency reports for each
In-sampleOut-of-sampleSign of WFEWhat the number actually means
PositivePositivePositiveThe only case the measure was defined for. The ratio is the fraction of the fitted edge that survived.
PositiveNegativeNegativeA real failure, correctly reported. The optimizer found something that did not generalize at all.
NegativePositiveNegativeReported as a failure, but the out-of-sample segments made money. The minus sign came from the denominator, not from out-of-sample performance.
NegativeNegativePositiveThe dangerous case. Both signs cancel and a strategy that lost money everywhere reads as a pass.

Two of those four rows are misreported, which is why a negative walk-forward efficiency is not self-explanatory either. It can mean the strategy failed out-of-sample, or it can mean the optimizer never found an in-sample edge for the strategy to fail at. The sign of the ratio carries no information without the sign of the denominator alongside it.

The fix is a gate, not a better formula

No rewriting of the ratio repairs this, because the problem is that the quantity being asked for does not exist. The correct handling is to refuse to compute walk-forward efficiency in the cases where it is undefined. Two rules cover it:

  1. Sign gate. If the in-sample figure is at or under zero, the window fails unconditionally and walk-forward efficiency is not computed. There was no in-sample edge, so there is no denominator with a meaningful interpretation, whatever the arithmetic returns.
  2. Magnitude floor. If the absolute value of the in-sample figure is under a small floor, the result is undefined rather than pass or fail. A near-zero denominator makes walk-forward efficiency blow up. An in-sample Sharpe of 0.01 against an out-of-sample Sharpe of 0.05 gives 500%, which neither number supports.

Tradelyze applies the sign gate: its WF Efficiency is left blank whenever the average in-sample Sharpe ratio is zero or negative. It has no larger magnitude floor, so a tiny positive in-sample average can still produce a very large ratio. That ratio is capped at 2.0; details are under how Tradelyze calculates walk-forward efficiency. One published implementation that applies both rules, with a specific floor value, is described under technical and historical notes.

Check this on your own results

If your platform reports walk-forward efficiency, find the in-sample figure it used as the denominator and confirm it is positive. If the platform does not show the denominator, treat any walk-forward efficiency above 100% as unverified. First confirm separately that the in-sample segments were profitable.

How does Tradelyze calculate walk-forward efficiency?

Tradelyze's walk-forward card links to this page from its tooltips, so this section states what the card computes and how to read it. By default Tradelyze runs 2 rolling walk-forward windows, each split into an earlier 70% and a later 30%. The method chip next to the walk-forward badge names how the run was graded. You do not choose the method. Tradelyze picks One run, split by period when the strategy qualifies, and Re-tuned each window otherwise. When the chip reads Re-tuned each window, the optimizer re-tunes on the earlier part of each window and tests on the later part. Only that method computes WF Efficiency.

How Tradelyze calculates WF Efficiency

Tradelyze divides the average annualized Sharpe ratio on the unseen test windows by the average annualized Sharpe ratio on the tuning windows. When the tuning-window average is zero or negative, there was no edge to keep. Tradelyze then leaves WF Efficiency blank instead of showing a misleading number. A blank tile means: read Mean IS Sharpe and the Per-Window Results table instead.

Three more details decide how the WF Efficiency tile should be read:

  • The tile shows a ratio, not a percentage. 0.50 on the tile is the same thing as 50% elsewhere on this page. A value of 1.40 means the unseen test windows did better than the tuning windows.
  • The scale stops at 2.0 and -2.0. A value that reaches either end shows as ≥ 2.00 or ≤ -2.00, with the label clipped — at least this extreme. Tradelyze only requires the in-sample average to be above zero. A Mean IS Sharpe very close to zero can therefore push the ratio to the cap. When the tile shows ≥ 2.00, check Mean IS Sharpe before reading the result as strong.
  • A blank tile has more than one cause. Besides a zero or negative tuning average, the tile is blank when the unseen test windows produced no usable Sharpe ratio. Runs graded with fixed settings instead of re-tuning do not compute this ratio. There the method chip next to the badge reads One run, split by period (or Fixed settings across periods, on results from before 26 September 2026), and the tile is replaced by Retention Ratio: later-period Sharpe over earlier-period Sharpe for one fixed parameter set, measured on the same history the settings were chosen on. It is not a measure of performance on unseen data.

Where this appears in Tradelyze

On the Tradelyze walk-forward card this number is the WF Efficiency tile. It sits beside Mean IS Sharpe, Mean OOS Sharpe, OOS Profit and Windows Profitable, next to a badge that reads Confirmed, Not Confirmed, Inconclusive or NO VERDICT. That badge is not the ROBUST verdict on the robustness card, which scores four separate checks; see what ROBUST, ACCEPTABLE, MARGINAL and FRAGILE mean. For the Sharpe ratio behind Mean IS Sharpe and Mean OOS Sharpe, see which Sharpe ratio Tradelyze shows.

If the badge reads Not Confirmed, Inconclusive or NO VERDICT, see what to do when walk-forward fails or shows no verdict. To judge the whole report, not one tile, use the pre-trade checklist.

What does OOS Profit mean?

The OOS Profit tile adds up the percentage return of each usable out-of-sample test window; it is a plain sum, not a compounded account balance. A negative total means the test windows' returns added up to a loss. One bad window can drag the sum negative even when most windows made money. When the method chip reads One run, split by period, the same sum appears as Later-Stretch Profit: it adds up the later part of each window, which the settings were chosen on, so it is not out-of-sample.

Because OOS Profit is a sum, it is not where an account trading those windows one after another would have finished. The gap widens as individual window returns grow. The walk-forward analysis page works through that difference under why walk-forward results differ from a plain backtest. Windows that produced no usable result are left out of the sum rather than counted as zero.

OOS Profit is also not what WF Efficiency is computed from, since WF Efficiency uses Sharpe ratios. The two can move in opposite directions on the same run without either being wrong, so read Windows Profitable alongside both.

What does Windows Profitable tell you?

The Windows Profitable tile counts the usable test windows that made money, such as 1 of 2, with the number the badge needs beneath it. Tradelyze requires more than 60% of them for a Confirmed badge, which at the default two windows means both. Results from older versions show the same figure as a percentage, % Windows Profitable. When the method chip reads One run, split by period, the count is labeled Periods Profitable, the same rule decides Consistent or Not Consistent, and the periods it counts are not unseen data. A high WF Efficiency beside few profitable windows means one window carried the run, which is weak evidence.

Windows that produced no usable result are left out of the count rather than counted as losses. With fewer than two usable windows the badge reads Inconclusive, so a single surviving window cannot carry a verdict. A test window that placed no trades at all scores zero, and zero is not a profit, so it counts as unprofitable here. It also makes the badge Inconclusive, because any usable window with fewer than 5 test trades does.

TradeStation's walk-forward optimizer treats the same quantity as a criterion in its own right. Its default pass is 50% of runs profitable, with a distinction grade at 80%. That 80% figure belongs to the profitable-runs criterion, not to the efficiency threshold it is often quoted next to.

What makes a walk-forward result pass or fail?

Tradelyze's walk-forward badge has three outcomes, and it checks the evidence before it judges the strategy. The badge reads Inconclusive when any of these is true:

  • fewer than two windows produced a usable result;
  • more than half the windows were excluded;
  • a usable window placed fewer than 5 trades after its tuning stretch; or
  • a Sharpe ratio could not be measured.

Inconclusive means the test could not tell, not that the strategy failed, and the card says which reason applied. Otherwise the badge reads Confirmed when WF Efficiency is above 0.5 and more than 60% of the test windows made money, and Not Confirmed when either one misses. The card shows that second condition as a count, such as 2 of 2 with the number needed beneath it. At the default of two windows it means both. NO VERDICT means nothing was asked or answered: there is no verdict for this check, which is not the same as a failed one.

There is no universal rule; each platform sets its own, and the thresholds are policy rather than findings. The efficiency threshold is the 50% figure TradeStation popularized; the 60% profitable-windows threshold is a Tradelyze choice with no external source behind it. Tradelyze leaves WF Efficiency blank when the in-sample average is zero or negative. So a strategy that lost money on its tuning data cannot clear the first condition.

Which question the badge answers

When the method chip next to the walk-forward badge reads Re-tuned each window, every window reruns the whole parameter search. The badge then grades the winner each window chose. That verdict is about the tuning procedure, not about the recommended parameters shown elsewhere on the results page.

When the chip reads One run, split by period, nothing is re-tuned. The history is backtested once with one frozen parameter set, and the result is cut into periods. That is a separate question, covered under re-optimized walk-forward vs one run, split by period. The same rule and the same evidence checks decide both verdicts, with Retention Ratio in place of WF Efficiency, but they test different things, so read the chip before the badge. That is also why the fixed-settings badge reads Consistent or Not Consistent rather than Confirmed or Not Confirmed. Results from before 26 September 2026 may show a third chip, Fixed settings across periods, from the retired parameter stability method; read its Consistent or Not Consistent badge the same way. A result keeps the badge words in use when it was produced, so an older one can read Confirmed or Not Confirmed, or PASS or FAIL, under any chip. The chip, not the word, says what was tested.

The 60% is weaker than it sounds. It is measured over the windows that produced a usable result, not over the windows you asked for. A run can exclude up to half of its windows and still reach a verdict, as long as at least two are usable. Read the badge together with the Excluded Windows tile. If that count is above zero, a Confirmed badge rests on fewer independent tests of the strategy than were requested.

A run in which every window was excluded reads Inconclusive, not Not Confirmed: there was no evidence either way, which is not the same as a losing strategy. NO VERDICT means nothing was asked or answered for this check.

What each walk-forward badge or band means
Badge or band shownWhat it meansReasonable next stepSource
Confirmed Re-tuned each window only. The evidence checks passed, WF Efficiency was above 0.5 and more than 60% of the usable test windows made money. Check Excluded Windows and the Per-Window Results table to see how many windows the verdict rests on, then read the robustness card and the held-out test. No primary source; Tradelyze methodology guidance.
Not Confirmed Re-tuned each window only. The evidence checks passed, but WF Efficiency was 0.5 or less, or 60% or fewer of the usable test windows made money. Do not treat the recommended parameters as validated; reduce the number of optimized inputs or add history, then re-run. No primary source; Tradelyze methodology guidance.
Consistent or Not Consistent One run, split by period, or Fixed settings across periods on results from before 26 September 2026. The same rule applied to one frozen parameter set, with Retention Ratio in place of WF Efficiency, over periods of the history the settings were chosen on. It is not a test on unseen data. Read Retention Ratio and Periods Profitable as a check on evenness across the history. For a test of the recommended settings on unseen data, read the held-out test on the firm's card. No primary source; Tradelyze methodology guidance.
Inconclusive The stage ran, but the evidence was too thin for a verdict: fewer than two usable windows, more than half excluded, a usable window with fewer than 5 test trades, or a Sharpe ratio that could not be measured. The card says which. It is not a finding about the strategy. Read the reason on the card and the Per-Window Results table. Adding history gives each window more trades. No primary source; Tradelyze methodology guidance.
NO VERDICT Nothing was asked or answered: there is no verdict for this check. It does not mean a check failed. Read the method chip next to the badge and, if it is shown, the Per-Window Results table. No primary source; Tradelyze methodology guidance.
Likely overfit WF Efficiency from 0 up to and including 0.5: at most half of the tuned edge survived on unseen data. Do not treat the recommended parameters as validated; reduce the number of optimized inputs or add history, then re-run. No primary source; Tradelyze methodology guidance.
Inverted — lost out-of-sample WF Efficiency under 0: the strategy made money on its tuning data and lost on the unseen data. Treat the tuned settings as failed; use the Per-Window Results table to see which windows lost, and simplify the strategy before re-running. No primary source; Tradelyze methodology guidance.
Blank WF Efficiency No ratio was measured: the tuning-window average Sharpe ratio was zero or negative, or the unseen windows produced no usable Sharpe ratio. A fixed-settings run shows Retention Ratio in this tile's place, and it can be blank for the same reasons. Read Mean IS Sharpe and the Per-Window Results table. A Mean IS Sharpe at or under zero means the strategy had no edge even on the data it was tuned on. No primary source; Tradelyze methodology guidance.

What are the in-sample and out-of-sample figures in a walk-forward result?

Mean IS Sharpe is the average annualized Sharpe ratio of the usable walk-forward windows on their tuning data (in-sample). Mean OOS Sharpe is the same average on their unseen test data (out-of-sample). WF Efficiency is the second divided by the first. A Mean OOS Sharpe well under a positive Mean IS Sharpe means little of the tuned edge carried over.

The Sharpe ratio is a strategy's average return divided by how much its returns swing around. Annualized means converted to a yearly figure, which puts a long tuning stretch and a shorter test stretch on the same footing. Windows that produced no usable result are left out of both averages. The two figures can then cover fewer windows than the run requested; the Excluded Windows tile, when the card shows it, gives the count.

Every walk-forward efficiency divides two summary figures like these. One comes from the segments the optimizer trained on, the other from segments it never saw. Which statistic fills those two slots is a choice each implementation makes, and establishing it is a precondition for reading anybody's walk-forward efficiency.

Statistics used as the in-sample and out-of-sample figures, and who uses which
FigureAggregated howUsed by
Annualized net profit and loss, in currencyTotals across all windows annualized, then one ratioPardo, Chapter 11 — the original definition
Annualized return, in percentVaries by vendorThe common restatement of Pardo; identical to his ratio on a constant capital base
Sharpe ratioAveraged across windows on each side, then one ratioTradelyze, using annualized Sharpe ratios; also one published gated implementation, described under technical and historical notes

One Tradelyze caveat applies to these two tiles. Sometimes the method chip next to the walk-forward badge reads One run, split by period, or Fixed settings across periods on results from before 26 September 2026. On those runs no tuning happens inside the windows, and the two tiles read Mean Earlier Sharpe and Mean Later Sharpe instead. Earlier just means the earlier part of each window. The later part was not hidden from the optimizer, and the WF Efficiency tile is replaced by Retention Ratio.

Why do my walk-forward results change when I shift the start date?

Walk-forward results change with the start date because the start date decides which bars fall into which window. With only a few windows, that assignment can carry more of the result than the strategy does.

One documented case comes from an r/algotrading post by praveenbm5 on 29 January 2017. The poster ran the same strategy and the same window schedule twice, on backtest start dates 40 days apart: 11 August 2006 and 2 July 2006. The poster describes them as start dates that "only differ by a month".

Walk-forward efficiency came out at 91% in the first run and -12% in the second. The in-sample side barely moved: 69.30% annualized in one run against 69.88% in the other. The entire difference was out-of-sample, where +63% became -8.56%.

The same strategy run at two start dates about a month apart Grouped bar chart of results reported by praveenbm5 on r/algotrading in January 2017; these are real reported figures, not constructed. Run A, the original start date, shows in-sample annual return 69.30 percent, out-of-sample return 63.00 percent and walk-forward efficiency 91 percent. Run B, with the start date shifted about 40 days earlier, shows in-sample annual return 69.88 percent, out-of-sample return negative 8.56 percent and walk-forward efficiency negative 12 percent. The in-sample bars are nearly identical while the out-of-sample and walk-forward efficiency bars flip from strongly positive to negative. In-sample annual return Out-of-sample return Walk-forward efficiency 100% 50% 0% 69.30 63.00 91 69.88 -8.56 -12 Run A original start date Run B start shifted about 40 days earlier Same strategy and same window schedule, two start dates.
Figures reported by praveenbm5 on r/algotrading, 29 January 2017 — real reported results, not constructed. The in-sample bars are effectively identical: 69.30% against 69.88%. Everything that moved was out-of-sample. A metric that swings 103 percentage points on a 40-day calendar shift is reporting window placement, not strategy quality.

Nothing about the strategy changed between the two runs. What changed was which bars landed in which segment, and therefore which market conditions the frozen parameters were tested against. A shift of about a month moved the headline metric by 103 percentage points. That run's walk-forward efficiency says almost nothing about the strategy and almost everything about where the windows fell.

Two responses are worth taking. First, re-run the analysis at several start dates and look at the spread of walk-forward efficiency, not a single value. A range of 55% to 70% across start dates tells you something about the strategy. A range of -12% to 91% tells you the sample is too small. Second, increase the number of windows, as long as each out-of-sample block still holds enough trades to measure; see how many walk-forward windows you need.

How many walk-forward windows do I need?

Tradelyze runs 2 rolling walk-forward windows by default. At that default, both windows must produce a usable result before any verdict is given, and both must make money for Confirmed, so one losing window decides the badge. Check the Excluded Windows tile and the Per-Window Results table before trusting it. With so few windows, whichever test window happened to be favorable dominates the headline ratio. In praveenbm5's r/algotrading post of January 2017, two runs with start dates 40 days apart scored 91% and -12%.

More windows is not automatically better. TradeStation's documentation puts the floor at 5 walk-forward runs and wants roughly 30 trades in each out-of-sample block. A fixed history cut into more windows leaves fewer trades in each. That 30-trade figure is TradeStation's convention for its own optimizer, not a statistical result. The full trade-off is set out under how many windows should I run.

What else has to pass besides walk-forward efficiency?

TradeStation's walk-forward optimizer applies five criteria, not one, and every threshold in the list is a user-editable default rather than a fixed rule. Quoting only the walk-forward efficiency threshold and dropping the rest removes most of the value of the framework.

  1. Total net profit greater than zero across the walk-forward runs.
  2. Walk-forward efficiency of at least 50%, graded pass with distinction at 100% or above.
  3. At least 50% of walk-forward runs profitable, graded pass with distinction at 80% or above. The 80% figure belongs to this criterion. It is not a second condition attached to the efficiency threshold, which is where it is usually misplaced.
  4. No single walk-forward run contributing 50% or more of net profit.
  5. No single walk-forward run losing more than 40% of initial capital. This is a worst-case test measured within any one run, not a drawdown on the combined equity curve. A strategy whose combined curve never draws down 40% can still fail it on one bad run.

Each criterion is graded separately and carries its own verdict. TradeStation's term for a criterion that does not clear its threshold is Failed. Pass with distinction is likewise a per-criterion grade, not a strategy-level award. One report can grade a strategy with distinction on walk-forward efficiency and Failed on profit concentration. Quoting only the first grade is the misuse to avoid.

Criterion 4 is easy to overlook. A single window supplying half the total profit means a single market regime supplied half the profit. The walk-forward efficiency figure will not reveal it, because one large win flatters the totals the ratio is built from. Criterion 3 covers the other side. A strategy can pass the efficiency threshold while a minority of windows carry the whole result.

The five criteria answer separate questions. Did it make money? Did the edge transfer? Was the edge consistent across windows, spread across regimes, and survivable? Walk-forward efficiency alone answers only the second.

How is walk-forward efficiency different from forward testing and paper trading?

Walk-forward efficiency is computed entirely from historical data: the out-of-sample bars already exist when the test runs, and real order execution is never simulated. Forward testing runs a frozen strategy on new data as it arrives. Paper trading adds simulated fills, slippage and rejections on live prices. A strategy can score well on walk-forward efficiency and still fail on execution costs. The three methods are compared in full under walk-forward analysis vs forward testing vs paper trading.

Going deeper

The sections below go deeper into the research and vendor debates. You can skip them to read your own report.

How do you combine walk-forward efficiency across multiple windows?

A walk-forward analysis with several windows produces one in-sample and one out-of-sample figure per window, but only one headline walk-forward efficiency. There is more than one way to get from many windows to one number, and the methods do not agree.

  • Average of ratios. Compute walk-forward efficiency per window, then take the arithmetic mean of those percentages. Simple, and what many implementations do.
  • Pardo's aggregate. Annualize the total out-of-sample profit across all windows, annualize the total in-sample profit across all windows, then take a single ratio. Longer and larger windows carry proportionally more weight, and one catastrophic window is not diluted into being one vote among many.
  • Ratio of averages. Average the per-window figures on each side first, then divide once. Tradelyze does this with annualized Sharpe ratios: Mean OOS Sharpe divided by Mean IS Sharpe.

The average of ratios hides the most when windows disagree. A window with a strongly negative efficiency is averaged against better windows and disappears into one mediocre-looking figure, while Pardo's aggregate, built from summed annualized profit, does not let a large loss be canceled by smaller gains in the same way. A public dispute over exactly this on one platform's forum is summarized under technical and historical notes.

The practical rule: a reported walk-forward efficiency is not comparable across platforms unless you know which aggregation produced it. If your platform publishes a per-window table, compute more than one version and look at the gap. A large gap means the windows disagree with each other, which is itself the finding.

The part that is usually left out

Pardo's definition is on annualized dollar profit and loss, not annualized return. On a constant capital base the two produce the same ratio, so the return form used throughout this page is equivalent — but an implementation that compounds equity across windows is dividing different quantities than Pardo was, and should not claim his definition for the result.

The definition being explicit has not made implementations agree. Vendors still differ on whether returns are compounded or arithmetic, on how windows of unequal length are weighted, on whether the input is profit, return or a risk-adjusted figure, and on what happens when a figure is negative.

That is why two platforms fed the same trades and the same window schedule can report different walk-forward efficiency values, and why there is no single authoritative threshold to compare against.

Does an anchored or rolling window schedule change walk-forward efficiency?

Yes, an anchored or rolling window schedule changes walk-forward efficiency. A walk-forward analysis is a sequence of train-then-test pairs marched forward through history, and two window schedules are in common use. An anchored, or expanding, schedule always starts training at the first bar, so the training set grows with every window and no old bar is ever dropped. A rolling, or sliding, schedule keeps the training set a fixed length and moves it forward, so the oldest bars fall out as new ones come in.

The two schedules train on different bars, so the optimizer can pick different settings in each window. Depending on how the history is split, they can also test on different bars. Either way they give different walk-forward efficiency numbers on the same data, which is why a report should say which schedule produced its figure.

Tradelyze uses rolling windows by default, and the Approach tile on its walk-forward card names the schedule a run used. Both schedules are drawn side by side under rolling vs anchored walk-forward, together with when each one makes sense.

What does walk-forward efficiency not fix?

Walk-forward efficiency checks one strategy on one stretch of history. It cannot see how many other settings or strategies you tried before picking this one, and the more you tried, the more likely a good score is luck. It also says nothing about the price paths that history did not happen to take.

Selection bias under multiple testing. Bailey, Borwein, López de Prado and Zhu set this out in Notices of the American Mathematical Society 61(5), May 2014, pages 458–471. The paper is Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance. Its objection to holding data out is narrower than "holdouts do not work". The holdout method, the authors write, "does not take into account the number of trials attempted".

The paper also puts a number on what that leaves out. "After trying only seven independent strategy configurations, the expected maximum SR IS is 1 for a two-year long backtest, while the expected SR OOS is 0." That is the wording on page 461 of the printed article; the preprint spells the same sentence with digits. SR IS is the in-sample Sharpe ratio and SR OOS the out-of-sample one. Seven tries on two years of data make a Sharpe ratio of 1 the expected best tuning result, even when no strategy tried has any real edge.

Their companion paper, The Probability of Backtest Overfitting (Journal of Computational Finance 20(4), 2017, pages 39–69; SSRN 2326253), gives a way to measure the effect rather than just warn about it.

Single-path testing. A 2024 paper by Joubert, Sestovic, Barziy, Distaso and López de Prado sorts backtesting into three families: "walk-forward testing, resampling, and Monte Carlo simulations". The paper is Enhanced Backtesting for Practitioners, The Journal of Portfolio Management 51(2), pages 12–27. It circulated in working-paper form as The Three Types of Backtests (SSRN 4897573). The paper makes walk-forward's limitation explicit: it tests exactly one path, the one that happened. It cannot show how a strategy behaves on paths that did not occur, which is what resampling and Monte Carlo simulation exist to address.

In practice, the size of the search changes what a score is worth. As a constructed comparison, a walk-forward efficiency of 80% from a search over 10,000 parameter combinations is usually weaker evidence than 55% from a search over 20. Record the number of trials that produced the winning parameter set. Report it next to the walk-forward efficiency, because it changes how that figure should be read.

Is walk-forward efficiency a reliable measure of robustness?

Walk-forward efficiency is not a reliable measure of robustness on its own. It answers one narrow question well: did the fitted edge survive on bars the optimizer could not see? At least one experienced platform developer reports that it does not track robustness at all. That report and its limits are summarized under technical and historical notes.

Vendor support for the metric is mixed, but it is not the clean split it is sometimes presented as. TradeStation and MultiCharts make walk-forward efficiency central to their walk-forward reporting. StrategyQuant ships the same quantity under a different name: WF Stability, a databank column computing the out-of-sample to in-sample ratio, filterable like any other column. Zorro and NinjaTrader ship walk-forward optimization with no headline efficiency metric at all.

So three vendors compute the ratio under two names, and two decline to. That is a naming difference plus two genuine abstentions. It is not enough to conclude anything about the metric from the vendor list alone.

Walk-forward efficiency says nothing about how many parameter sets were tested to find an edge. It says nothing about the risk along the path inside each segment, and nothing about whether the out-of-sample segments were typical of anything. Use it as one input among several, alongside TradeStation's other pass criteria and a check of how many trials the optimizer ran.

What are the technical and historical notes on walk-forward efficiency?

The technical and historical notes on walk-forward efficiency collect the source detail behind the metric. They cover where Pardo writes the definition down and a public dispute over how one platform combined windows. They also cover a dissenting view on what the metric measures, one published gated implementation, and a naming note.

Where Pardo defines walk-forward efficiency

Walk-forward analysis and walk-forward efficiency are developed in Robert Pardo's The Evaluation and Optimization of Trading Strategies, 2nd edition (Wiley, 2008), Chapter 11, pages 237–261. Pardo states the definition twice and works it twice. On page 256: “The numbers in the Walk-Forward Efficiency column are ratios of annualized walk-forward and optimization Net Profit & Loss.” On page 260: “Recall that WFE is the ratio of an annualized walk-forward profit and loss to an annualized optimization profit and loss.”

In the same chapter, a single-window figure appears on page 250: “Efficiency: 341 percent”. The multi-window aggregate appears on page 260. There, $11,208 of annualized walk-forward profit against $8,278 of annualized optimization profit gives 135%.

Arithmetic or compounded annualization

The annualization convention has to be stated, because it changes the answer. The one-window example under how to calculate walk-forward efficiency annualizes arithmetically: 48.0% over 24 months is halved to 24.0% a year, and 6.0% over 6 months is quadrupled to 12.0% a year. Compounding the same period returns instead gives 21.7% in-sample (1.48 to the power of one half, minus 1) and 12.4% out-of-sample (1.06 squared, minus 1), and a walk-forward efficiency of 57.1% rather than 50%.

Seven percentage points of headline metric turn on a convention that most reports never name. Neither convention is wrong; reporting a walk-forward efficiency without saying which one produced it is.

The Wealth-Lab aggregation dispute

Wealth-Lab's walk-forward implementation displayed a walk-forward efficiency of 27.45% for a run whose per-window efficiencies were 93.99%, 49.51% and -61.15%. A forum user, kribel, reverse-engineered the 27.45% and showed it was exactly the arithmetic mean of the three. That established averaging ratios as the platform's own method, not a shortcut some user had taken by hand.

The thread argued this against the aggregate Pardo works through at "255 ff", and Wealth-Lab's Eugene Cone conceded the point in February 2015. The concession was an acknowledgment in the thread; it is not evidence that the calculation was subsequently changed.

The documented Wealth-Lab case: three windows, one disputed headline number. Figures as reported in the vendor's forum thread, not constructed
WindowWalk-forward efficiency
Window 193.99%
Window 249.51%
Window 3-61.15%
Wealth-Lab's displayed figure — the arithmetic mean of the three27.45%

In the Wealth-Lab case, the average of ratios takes window 3's -61.15% and blends it with two better windows into a single mediocre-looking 27.45%. Pardo's aggregate, which works from summed annualized profit rather than summed ratios, does not let a large loss be canceled by two smaller gains in the same way.

A dissenting view from Zorro's developer

Johann Christian Lotter (JCL) is CTO of oP group Germany and lead developer of Zorro. Lotter reports that in their testing, walk-forward efficiency was not correlated with robustness. Some quite robust strategies had a low walk-forward efficiency, and some unstable ones scored high. Read "robust" in Zorro's sense: insensitive to small parameter changes. It does not mean "survived out-of-sample", which would make the claim circular, because walk-forward efficiency is itself an out-of-sample-over-in-sample ratio.

Lotter's report is also a two-sentence forum reply with no methodology attached. It is one experienced developer's account of unpublished internal testing, worth knowing about and not worth treating as a result.

A published gated implementation

Ushana Kevin Iorkumbul's MQL5 article 23250, published 9 July 2026, describes a gated walk-forward efficiency on Sharpe ratios. It applies both a sign gate and a magnitude floor to the in-sample denominator. The floor is exposed as a runtime parameter with 0.25 as its value, rather than hardcoded.

The article is user-submitted and carries MetaQuotes' standard disclaimer that they are not responsible for its accuracy. Read 0.25 as one author's choice rather than as a published standard. Gating of this kind is not something the major vendors document.

WFE, WFA and WFO, but not WFF

One naming note, since it causes confusion in search: "WFF" is not a standard abbreviation for this metric. The metric is walk-forward efficiency, abbreviated WFE. The procedure that produces it is walk-forward analysis (WFA) or walk-forward optimization (WFO).

Stage 3 · step 13 of 18. Next in the learning path: Monte Carlo simulation

Check it on your own strategy

In a Tradelyze report, this is the WF Efficiency tile on the walk-forward card. Tradelyze re-runs an uploaded TradingView Pine Script strategy's backtest on your price data and checks it against your exported trades. It then runs parameter optimization, walk-forward analysis, a four-check robustness score and prop firm rule checks. It does not place trades, give financial advice or guarantee a challenge pass, and it is in beta.

Create an account. Already a user? Open your strategies.

Frequently asked questions about walk-forward efficiency

What is walk-forward efficiency?

Walk-forward efficiency is how much of a strategy's performance survives on data the optimizer never saw: the yearly return on that unseen data divided by the yearly return on the data it was tuned on. A strategy that returned 60% a year on its tuning data and 30% a year on the unseen data that followed scores 50%, meaning half of the optimized edge survived.

What is the walk-forward efficiency formula?

Walk-forward efficiency = annualized out-of-sample performance divided by annualized in-sample performance, usually expressed as a percentage. Robert Pardo defines it on annualized dollar profit and loss in The Evaluation and Optimization of Trading Strategies, 2nd edition (Wiley, 2008), Chapter 11: walk-forward efficiency is the ratio of an annualized walk-forward profit and loss to an annualized optimization profit and loss. On a constant capital base that is the same number as a ratio of returns.

What is a good walk-forward efficiency?

TradeStation's Walk-Forward Optimizer ships 50% as the pass threshold for its walk-forward efficiency criterion and grades that criterion pass with distinction at 100% or above. Both are user-editable defaults rather than fixed rules, and they are the only widely quoted thresholds with a traceable vendor source. The 0.3, 0.5 and 0.7 bands repeated across blogs and forums have no primary source and do not appear in Pardo.

Can walk-forward efficiency be over 100%?

Yes. Walk-forward efficiency above 100% means annualized out-of-sample return exceeded annualized in-sample return. TradeStation's own walk-forward FAQ addresses this question and answers it with cluster analysis: check whether the high figure sits among similarly high neighboring results. If it stands alone it is an aberration. A small in-sample denominator also inflates the ratio arithmetically; Tradelyze caps its ratio at 2.0 and labels a capped value clipped — at least this extreme.

Can walk-forward efficiency be negative?

Yes, and in two different ways, which is why a negative figure is not self-explanatory. Walk-forward efficiency is negative when in-sample performance was positive and out-of-sample negative, meaning the optimizer found nothing that generalized. It is also negative when in-sample was negative and out-of-sample positive. And when both sides are negative the signs cancel into a positive ratio that reads as a pass.

Why is my walk-forward efficiency positive when the strategy lost money?

Because the metric is a ratio and two negatives cancel. A mean out-of-sample Sharpe of -0.588 divided by a mean in-sample Sharpe of -0.457 gives 1.2867, displayed as 128.7%, even though every window lost money. The fix is gating: if in-sample performance is at or under zero, the window fails unconditionally and walk-forward efficiency is not computed at all. Tradelyze applies that gate and leaves the figure blank.

Who invented walk-forward efficiency?

Robert Pardo. Walk-forward analysis and walk-forward efficiency are set out in The Evaluation and Optimization of Trading Strategies, 2nd edition (Wiley, 2008), Chapter 11, pages 237–261, with a single-window example on page 250 and the aggregate calculation on page 260. Pardo states the ratio explicitly, so the divergence between implementations is a vendor choice rather than a gap in the source.

Does a high walk-forward efficiency mean the strategy is not overfit?

No. Walk-forward efficiency measures out-of-sample decay for the one parameter set that was selected. It does not account for how many parameter sets were tested to find it. Bailey, Borwein, López de Prado and Zhu note in Notices of the American Mathematical Society (May 2014) that the holdout method does not take into account the number of trials attempted, and that seven independent configurations already give an expected maximum in-sample Sharpe ratio of 1 on a two-year backtest.

Why do two platforms report different walk-forward efficiency for the same strategy?

Because vendors implement Pardo's definition differently, not because the definition is missing. Some average the per-window ratios; Pardo annualizes the totals and takes one ratio. Wealth-Lab displayed 27.45% for windows of 93.99%, 49.51% and -61.15%; a forum user reverse-engineered that figure to prove it was the plain average of the three, argued it disagreed with Pardo, and the vendor conceded in February 2015.

Can a strategy with a good walk-forward efficiency still fail in forward testing?

Yes. Walk-forward efficiency is computed entirely on historical data, because the out-of-sample segments already exist when the test is run. Forward testing runs a frozen strategy on data as it arrives, and paper trading adds simulated fills and slippage on live prices. A strategy can score well on walk-forward efficiency and still fail on execution costs, or on market conditions its history never contained.

Should walk-forward efficiency use returns or Sharpe ratios?

Pardo's definition uses annualized dollar profit and loss, which gives the same ratio as annualized return on a constant capital base. Several implementations substitute a risk-adjusted figure instead; Tradelyze divides the average annualized Sharpe ratio of its unseen test windows by that of its tuning windows. Both approaches are in use and they are not interchangeable, so a reported walk-forward efficiency is uninterpretable without knowing which one was used.

Does Tradelyze's walk-forward efficiency have the negative-return problem?

No. Tradelyze's WF Efficiency divides the average annualized Sharpe ratio on the unseen test windows by the average annualized Sharpe ratio on the tuning windows, and leaves the figure blank when that tuning average is zero or negative. With no edge to measure, two losing periods cannot produce a positive score. A blank tile means reading Mean IS Sharpe and the Per-Window Results table instead.

Does an anchored or rolling window schedule change walk-forward efficiency?

Yes. An anchored, or expanding, schedule always starts training at the first bar, so the training set grows with every window. A rolling, or sliding, schedule keeps the training set a fixed length and moves it forward. The two schedules put different bars into the in-sample and out-of-sample segments, so they give different walk-forward efficiency figures on the same data. Always report which one was used. Tradelyze uses rolling windows by default.

What else should pass besides walk-forward efficiency?

TradeStation's walk-forward optimizer applies five criteria, all with user-editable default thresholds: total net profit above zero; walk-forward efficiency of at least 50%; at least 50% of walk-forward runs profitable; no single walk-forward run producing 50% or more of net profit; and no single run losing more than 40% of initial capital. Each criterion is graded on its own, so a strategy can pass one and be graded Failed on another.

Sources

  • Robert Pardo, The Evaluation and Optimization of Trading Strategies, 2nd edition, Wiley, 2008. Chapter 11, pages 237–261; walk-forward efficiency defined at pages 256 and 260, single-window example at page 250 ("Efficiency: 341 percent"), multi-window aggregate at page 260 ($11,208 / $8,278 = 135%).
  • TradeStation Walk-Forward Optimizer documentation and FAQ — the five pass criteria and their user-editable default thresholds (net profit above zero; walk-forward efficiency 50%, distinction at 100%; 50% of runs profitable, distinction at 80%; no run supplying 50% or more of net profit; no run losing more than 40% of initial capital), the 5-run minimum, the statement that a walk-forward analysis including at least 10 walk-forward tests approaches such reliability, roughly 30 trades per out-of-sample run, and the cluster-analysis guidance for results above 100%.
  • praveenbm5, Walk Forward Analysis (WFA) — Date Range Sensitivity, r/algotrading post 5qutj9, 29 January 2017 — two runs at start dates 40 days apart reporting walk-forward efficiency of 91% against -12%, in-sample 69.30% against 69.88%, and out-of-sample +63% against -8.56%. Real reported results, not constructed.
  • Wealth-Lab forum discussion of walk-forward efficiency aggregation — per-window figures of 93.99%, 49.51% and -61.15% against a platform-displayed 27.45%, shown by the user kribel to be their arithmetic mean and argued against Pardo at "255 ff"; concession by Wealth-Lab's Eugene Cone, February 2015.
  • Johann Christian Lotter (JCL), CTO of oP group Germany and lead developer of Zorro — a two-sentence forum reply reporting that some quite robust strategies had a low walk-forward efficiency and some unstable ones scored high, with no methodology given. "Robust" in Zorro's vocabulary means insensitive to small parameter changes.
  • StrategyQuant — ships the out-of-sample to in-sample ratio as the databank column WF Stability, filterable like any other databank column.
  • Ushana Kevin Iorkumbul, MQL5 article 23250, 9 July 2026 — a user-submitted article published under MetaQuotes' disclaimer that they are not responsible for its accuracy, describing a gated walk-forward efficiency with a sign gate and a magnitude floor on the in-sample Sharpe denominator, the floor exposed as a runtime parameter with 0.25 as its value.
  • David H. Bailey, Jonathan M. Borwein, Marcos López de Prado and Qiji Jim Zhu, Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance, Notices of the American Mathematical Society 61(5), May 2014, pages 458–471, DOI 10.1090/noti1105.
  • David H. Bailey, Jonathan M. Borwein, Marcos López de Prado and Qiji Jim Zhu, The Probability of Backtest Overfitting, Journal of Computational Finance 20(4), 2017, pages 39–69 (DOI 10.21314/JCF.2016.322 carries a 2016 stamp); SSRN 2326253.
  • Joubert, Sestovic, Barziy, Distaso and López de Prado, Enhanced Backtesting for Practitioners, The Journal of Portfolio Management 51(2), pages 12–27, 2024; circulated in working-paper form as The Three Types of Backtests, SSRN 4897573.
  • Tradelyze walk-forward implementation, reviewed 26 September 2026: efficiency = mean out-of-sample annualized Sharpe / mean in-sample annualized Sharpe over usable windows, undefined (tile left blank) when the in-sample mean is not positive, and clamped to the range -2.0 to 2.0; tooltip labels Generalizes well above 0.5, Likely overfit from 0 up to and including 0.5, and Inverted — lost out-of-sample below 0; Confirmed requires efficiency above 0.5 and more than 60% of usable windows profitable, shown as a count; the badge reads Inconclusive when fewer than two windows are usable, more than half are excluded, a usable window placed fewer than 5 test trades, or a Sharpe ratio could not be measured; the method is chosen automatically, One run, split by period when the strategy qualifies and Re-tuned each window otherwise; One run, split by period shows Retention Ratio, Later-Stretch Profit and Periods Profitable and reads Consistent or Not Consistent; a result keeps the badge words in use when it was produced; parameter stability (Fixed settings across periods) was retired on 26 September 2026.

Related terms

Tradelyze

Last reviewed 26 September 2026. Educational content about backtest validation methodology. Nothing here is financial advice.