Guide

What to do when a strategy fails validation

Last reviewed: 26 September 2026·Tradelyze

When a trading strategy fails validation, find the cause before changing anything: the usual causes are a data mismatch, too little evidence, too much tuning, or a position too large for the rules. Fix that cause, then run the test once more. Re-optimizing until a check passes makes overfitting worse, because every extra try is another chance for luck.

In plain English

A failed check tells you something specific. It can mean the re-run is not really your strategy, or that there is too little history to judge. It can also mean the settings were fitted to past noise, or that your position size breaks a prop firm's loss limit. Fitting settings to past noise is called overfitting: the strategy describes the past well and new prices badly. Each cause has its own fix. The move that fixes none of them is re-running the optimizer with small changes until the check turns green. By then the check has stopped measuring anything.

New to this? Start with How Tradelyze validates a strategy, which explains every card named on this page.

The order in which to work through a failed validation A left-to-right flow of four boxes joined by arrows, showing the order in which to work through failed checks. Box 1, Trade matching, asks whether the re-run is your strategy and is read on the Backtest vs TradingView card. Box 2, Enough evidence, asks whether there are enough trades and walk-forward windows to judge, read on the Trade Count tile and the N Windows tile. Box 3, Overfitting, asks whether the tuning survived data it never saw, read on the walk-forward and robustness cards. Box 4, Rules and size, asks whether the position size fits the firm's limits, read in each prop firm card's Rule Results. A dashed strip underneath all four boxes says that each fix is one new run, and that you should write down every run because no check on the report can count them for you. The diagram is conceptual and plots no data. 1 · Trade matching 2 · Enough evidence 3 · Overfitting 4 · Rules and size Is the re-run your strategy? Enough trades and windows to judge? Did the tuning survive unseen data? Does the size fit the firm's limits? Backtest vs TradingView Trade Count, N Windows Walk-forward, robustness Prop firm Rule Results Each fix is one new run. Write down every run: no check on the report can count them for you.
Work through a failed validation in this order, because each step depends on the one before it. A trade mismatch makes every later number describe a different strategy. Thin evidence makes the overfitting checks unreliable. Position size changes whether prop firm rules pass, not whether the edge is real. The order is practical guidance, not a published standard.

What does each failing result mean, and what should I do next?

Each validation check in a Tradelyze report ends in a verdict, and each failing verdict points to a different cause. The first two rows below are not verdicts at all: the run never finished, so nothing about the strategy was judged. Find your result, make the one change that addresses its cause, and run again once. The Source column separates what Tradelyze's code does from judgment calls with no published standard behind them.

What each failing result in a Tradelyze report usually means, the next step, and the move to avoid
ResultLikely causeNext stepWhat not to doSource
A card headed Optimization stopped — the backtest could not be run, or Optimization stopped — backtest failed with the name of an errorNot a verdict on the strategy. The run stopped before it measured anything, so nothing about your strategy has been judged. Without an error name, the usual cause is the backtesting service being unavailable or restarting. With one, the backtest itself failed, and the error can point at your data or script.Press Resubmit on that card and pick your firms again in the Run Optimization dialog that opens. That resubmit does not use another credit. After an outage, wait a little while first; backtests that already finished are reused if you resubmit on the same UTC day. After a named error, fix what it points at first, or the run stops the same way.Changing the ranges or the position size in response. Nothing was measured, so there is nothing to tune.Card wording and the free Resubmit: Tradelyze implementation.
A banner reading Processing is taking longer than expected, with a Cancel & Delete buttonNo result has arrived 5 hours after you submitted, while the strategy is still waiting for or going through trade matching. The banner says the processing service may be temporarily unavailable.Wait: the banner says the strategy resumes automatically once the service is back online. Cancel & Delete permanently removes the strategy and its uploaded files. The credit already spent is not returned, and a new submission costs 1 credit.Reading the delay as a verdict on the strategy.5-hour threshold, banner wording and credit cost: Tradelyze implementation.
Low Match Rate on the Backtest vs TradingView card, or a run halted on a card headed Optimization stopped with a match percentageThe re-run is not trading your TradingView strategy. Common causes: a wrong Chart Timezone, Properties or Inputs changed only in TradingView, or price data for other dates or another timeframe.Check Chart Timezone first, write your TradingView settings into the script's defaults, match the price file to the chart, then re-export and submit again. See what to fix first.Pressing Continue anyway and then reading the report as a verdict on your TradingView strategy.Card names and Continue anyway: Tradelyze implementation. Order of checks: no primary source.
Walk-forward badge reads Not Confirmed or Not ConsistentWF Efficiency, or the Retention Ratio on a run split by period, was 0.5 or less, or 60% or fewer of the usable windows made money. Not Confirmed, on a run labeled Re-tuned each window: the tuned edge faded on unseen data, or one of only a few windows lost. Not Consistent, on a run labeled One run, split by period: the search's overall best settings did worse in some periods of the history they were chosen on.Open Per-Window Results to see which window or period fell short. Add history or windows before changing the strategy; if the edge faded everywhere, search fewer inputs. See walk-forward fails.Narrowing ranges around whatever did well in the test windows and re-running until the badge turns green.Badges and pass conditions: Tradelyze implementation. Next step: no primary source.
Walk-forward badge reads Inconclusive or NO VERDICTInconclusive: the stage ran but the evidence was too thin. Fewer than 2 windows were usable, more than half were excluded, a usable window placed fewer than 5 trades in its test stretch, or a Sharpe ratio could not be measured; the card says which. NO VERDICT: no walk-forward question was asked or answered on this run.For Inconclusive, add history so each window holds more trades. For NO VERDICT, read the method label beside the badge, then the Per-Window Results table if there is one.Reading either as a pass, or as a failed check.Tradelyze implementation.
A robustness check shows NOT RUNThe check did not happen. Parameter Sensitivity is skipped automatically when fewer than 50 optimization trials ran. A check that did not run is not a failure and does not reduce the score.Set Optimization trials to Automatic or to at least 50, then run again. See NOT RUN or FRAGILE.Treating a high score as if it covered the missing check. Any NOT RUN check rules out a ROBUST verdict.Tradelyze implementation.
An amber note on the robustness card reads Score reduced, with the verdict MARGINAL or FRAGILEAt least one scored check failed, a check refused for too few trades included, or the data covers less than 30 days. The note names which. The points are multiplied by 0.69, so the score is 69 or less: C+ and MARGINAL at best.Act on the named check, as in the next row. For short data, export more history. The line under the verdict, N of 4 checks passed, shows how many passed and how many did not run.Reading the points before the reduction as the score, or tuning until the failed check only just passes.Reduction rule and card wording: Tradelyze implementation.
Robustness verdict reads FRAGILEThe robustness score is under 50. Because a failed check multiplies the points by 0.69, one badly failed check can be enough: computed with Tradelyze's scoring code, a Ruin Probability of 50% or more, with every other check perfect and 400 days of data, scores 49.0, which the card shows as 49. Which checks failed matters more than the total.Read the four scored check rows and act on the failing one: Unstable, Not Significant, Too Few Trades or a high Ruin Probability each point to a different cause. Min Backtest Length is shown for information only.Tuning until the score crosses 50 or 70.Verdict bands and scoring: Tradelyze implementation. Reading of each row: no primary source.
A prop firm card reads Not FeasibleAt least one rule in the Rule Results table failed. When a drawdown rule fails, the usual cause is a position too large for a fixed loss limit.Read the first failing row. For a drawdown rule, reduce the order size in the script's strategy() defaults, re-export and submit again. See Not Feasible.Loosening a custom rule's limit so the card reads Qualifies; the firm still applies its own limit.Badge and table: Tradelyze implementation. Sizing first: no primary source.
A Best Value sits at the edge of its Range in the Parameter Search Space tableThe search may have wanted values beyond the range it was allowed to try.Decide whether values past the edge are ones you would really trade. If so, widen that one range once and count the run. See widening a range.Trading the edge value as a tested peak; nothing beyond it was tried.No primary source.

A trial in that table is one full backtest with one set of settings. The optimizer is the part of Tradelyze that runs many trials and keeps the best. Every step in the table changes one thing and runs once. That discipline is what keeps a second run informative.

What if the verdicts disagree?

When two verdicts in a Tradelyze report disagree, let the more cautious one decide, because each check can see a problem the other cannot. A prop firm card checks the firm's rules on one backtest. The walk-forward badge asks whether the tuning held up on prices it never saw or, on a run split by period, whether the settings held up evenly across the history they were chosen on. The held-out test asks whether the exact recommended settings held up on data withheld from the search. The robustness verdict asks whether the result depended on luck. Match Rate asks whether any of those numbers describe your TradingView strategy at all.

Which verdict limits the decision when two verdicts in a Tradelyze report disagree
Verdict pairWhat it meansWhich one limits the decision
Qualifies + walk-forward Not Confirmed or Not ConsistentThe firm's rules pass on the settings recommended for that firm, graded over your full history, most or all of which the settings were chosen on. Not Confirmed says the tuning did not hold up on data it never saw. Not Consistent says the settings did worse in some periods of that same history.The walk-forward result. Don't pay the challenge fee on the Qualifies badge yet; first find out why walk-forward did not pass, and read the held-out test on the firm's card.
Walk-forward Confirmed or Consistent + Not FeasibleWalk-forward passed, but at the backtest's position size at least one of the firm's rules failed.Not Feasible, for that firm. Read the first failing Rule Results row; a drawdown or daily loss row usually points to position size. See position sizing for prop firm challenges and how Tradelyze checks a daily loss limit.
Qualifies + FRAGILEThe rules pass on this backtest, but the robustness score is under 50. The result may depend on luck, such as a lucky order of trades, or on one narrow setting.FRAGILE. A pass that rests on luck may not repeat in the challenge. Read which robustness check failed before anything else.
ROBUST + low Match RateAll four scored robustness checks passed, but the checks tested Tradelyze's re-run, which traded differently from your TradingView strategy. The scores describe a re-run that isn't your strategy.Match Rate. Fix the trade matching first, submit again, and read the scores on the new run.

A Qualifies badge means every rule on a prop firm card passed, and Not Feasible means at least one failed. FRAGILE is Tradelyze's lowest robustness verdict, given to a score under 50. A low Match Rate can sit beside a full report, because Continue anyway and Run Optimization Anyway both optimize the re-run as it is.

A constructed example, not measured data: a strategy's prop firm card reads Qualifies, with a maximum total drawdown of 7.2% against a 10% limit. Its walk-forward card, labeled Re-tuned each window, reads Not Confirmed, with WF Efficiency at 0.21 across two windows that each held more than 40 trades. WF Efficiency is roughly the share of the tuned Sharpe ratio that survived on the test windows, so only about a fifth survived. The rules passed mostly on history the settings were tuned to fit. A challenge trades new prices, so the walk-forward result is the better guide.

The badges, and what each check measures, come from Tradelyze's implementation. Which verdict should limit the decision, and the next steps in the table, are practitioner judgment with no primary source.

Why doesn't my re-run match TradingView, and what should I fix first?

A low Match Rate means Tradelyze's re-run of your script and your TradingView trade list disagree about which trades happened. Match Rate is the number of matched trades divided by the larger of the two trade counts compared, shown on the Backtest vs TradingView card. Fix a low Match Rate before reading anything else. Every later number, from Best Metrics to the prop firm cards, describes the re-run, not the strategy in your TradingView report.

Check these three causes in order, quickest to rule out first:

  1. Chart Timezone. The upload form asks for the zone shown in the bottom-right corner of your TradingView chart, because both exported files are stamped in that zone. A wrong zone shifts every trade time. Tradelyze's own code notes that the same faithful re-run can measure a 100% match against one timezone and about 16% against another. The export guide's timezone section shows which zone to pick.
  2. Script defaults. Tradelyze runs your script with its own defaults: initial capital, commission, slippage, order size, pyramiding and input values come from the strategy() declaration and the input() calls. If you changed any of these in TradingView's Properties or Inputs tabs before exporting, write the same values into the script and upload it again. TradingView strategy properties explains what each setting changes.
  3. Price data range and timeframe. The OHLCV file holds the open, high, low, close and volume of each bar. It must be for the same instrument and timeframe as the chart you exported from, and cover the same dates. A trade is compared only if it entered and exited between the file's first and last bar; any other trade is left out on both sides. If the file starts before your exported trade list does, the Match Rate drops: re-run trades inside the file but before the export's first trade count as backtest-only.

A constructed example, not measured data: a strategy with 212 TradingView trades comes back with a Match Rate of 18%. Review comparison shows the re-run took the same sequence of long and short trades, but each one entered exactly five hours after its TradingView partner. Tradelyze only pairs trades whose entries fall within five minutes of each other, so a constant offset leaves almost nothing matched. A constant offset points to the timezone, not the strategy. Correcting Chart Timezone and submitting again is the fix; retuning would change nothing, because the settings were never the problem.

Continue anyway, on the Optimization stopped card that shows a match percentage, optimizes the re-run as it is. It is reasonable only once Review comparison shows the unmatched trades are explained and minor. After a wrong timezone, changed defaults or a mismatched price file, fix the cause and submit again instead; each submission costs 1 credit. The full list of reasons trades do not match and what Continue anyway does are covered on the TradingView backtest accuracy page.

What should I do when walk-forward fails or shows no verdict?

When walk-forward does not pass, first find out whether the edge faded or the evidence was too thin, because the two need opposite responses. Walk-forward analysis tunes a strategy on one stretch of history, called in-sample, then scores it on the stretch that follows, called out-of-sample, which the tuning never saw. When the label beside the badge reads Re-tuned each window, Tradelyze does exactly this over 2 windows by default, and the badge reads Confirmed or Not Confirmed. Tradelyze picks the method for each run, not you. For strategies that qualify, the default is One run, split by period: the search's overall best settings are scored on periods of one continuous run, and the badge reads Consistent or Not Consistent. That method is not an out-of-sample test, because the settings were chosen on the same history. If a firm was recommended other settings, an amber banner on the walk-forward card names it.

Two conditions decide the pass: WF Efficiency is above 0.5, and more than 60% of the usable test windows made money. Before either is judged, Tradelyze checks the evidence. If fewer than 2 windows were usable, more than half the windows were excluded, a usable window placed fewer than 5 trades in its test stretch, or a Sharpe ratio could not be measured, the badge reads Inconclusive and the card says why. WF Efficiency is the average annualized Sharpe ratio on the test windows divided by the same average on the tuning windows. The Sharpe ratio is average return divided by how much returns swing, and annualized means converted to a yearly rate.

A two-window walk-forward result, re-tuned in each window, that reads Not Confirmed. Constructed illustration, not measured data
Figure on the cardShowsNeededMet?Source
WF Efficiency (Mean OOS Sharpe 0.81 ÷ Mean IS Sharpe 1.30)0.62Above 0.5, for ConfirmedYesTradelyze implementation
Windows Profitable (window 1 made money, window 2 lost)1 of 2More than 60%, so 2 of 2, for ConfirmedNoTradelyze implementation
Usable windows2At least 2, or the badge reads InconclusiveYesTradelyze implementation
Excluded Windows0No more than half, or the badge reads InconclusiveYesTradelyze implementation

In that constructed result, the tuned edge mostly carried over, yet the badge reads Not Confirmed because one of only two test windows lost money. One window decided the verdict. The right response is more evidence, not re-tuning. Export a longer trade list and price file for the same strategy, or raise Walk-forward windows. Tradelyze allows 2 to 6 windows when walk-forward and robustness settings are enabled for your account. More windows on the same history make each test window shorter, with fewer trades in it, as how many walk-forward windows to run explains.

A different Not Confirmed needs a different response. If WF Efficiency is at or below 0.5 across windows that each had plenty of trades, the tuned edge faded on data the optimizer never saw. The next step is a smaller search, not a bigger one. Fix the inputs you would never change, and narrow the remaining ranges to values you would really trade, using the Fixed, Min, Max and Step settings. Fewer combinations tried make any result that survives stronger evidence. What each WF Efficiency label means is covered in reading the WF Efficiency tile.

On a run labeled One run, split by period, the tiles read Retention Ratio, Mean Earlier Sharpe, Mean Later Sharpe, Later-Stretch Profit and Periods Profitable, and the pass rule is the same. Not Consistent there means the search's overall best settings did worse in some periods of the history they were chosen on. That can be a quieter market as well as overfitting, and no unseen data was tested. The check of each firm's recommended settings on data they were not chosen on is the held-out test on that firm's card: by default the last 25% of the history is withheld from the whole search. Read it before narrowing any range.

Inconclusive means the stage ran but the evidence was too thin to call; the card says which rule it hit, and the fix is more history, not different settings. NO VERDICT means no walk-forward question was asked or answered on this run. Neither is a pass or a failed check. Read the small label beside the badge first: Re-tuned each window or One run, split by period. Results from before 26 September 2026 may instead show Fixed settings across periods, from a third method, parameter stability, that has since been retired; it also reads Consistent or Not Consistent. Then read the Per-Window Results table. The methods answer different questions, set out in how the walk-forward methods differ.

Do not tune toward the test windows

Each time you change a strategy after reading a walk-forward result and run it again, the test windows help choose the strategy. After a few rounds they are no longer unseen data, and a Confirmed badge on them means much less. Change one thing for a reason you could state before seeing the result, and run once. What the walk-forward badge requires has the full gate.

What should I do when a robustness check shows NOT RUN or FRAGILE?

When a robustness check shows NOT RUN, make sure it runs next time; when the verdict is MARGINAL or FRAGILE, read which checks failed and fix that cause. The robustness score is one 0–100 number that adds up four scored stress tests of the tuned result: Monte Carlo, Permutation Test, Parameter Sensitivity and Deflated Sharpe Ratio. The card also shows Min Backtest Length, for information only; it earns no points. Tradelyze's verdict is ROBUST when all four scored checks ran and passed, the data covers at least 30 days (one month), and the score is 80 or more. Otherwise it reads ACCEPTABLE at 70 or more, MARGINAL at 50 or more, and FRAGILE below 50.

Since 26 September 2026, a failed check or a short history reduces the score. If any scored check failed, a check refused for too few trades included, or the data covers less than 30 days or has no measurable span, the points are multiplied by 0.69. Such a run scores 69 or less, so C+ and MARGINAL are the best it can show, and an amber note on the card says the score was reduced and why. The factor keeps near misses in order rather than tying them all at one number. A grade of B- or better, and so any ACCEPTABLE or ROBUST verdict, therefore means no check failed and the data covers at least 30 days. The line under the verdict reads N of 4 checks passed, plus how many did not run.

NOT RUN means a check did not happen, which is different from failing it. The common case is Parameter Sensitivity on a small search. Tradelyze's automatic sensitivity budget is zero when fewer than 50 optimization trials ran. The automatic trial budget is never fewer than 60, so the usual trigger is Optimization trials set to a custom 20, 30 or 40. If walk-forward and robustness settings are enabled for your account, set it to Automatic or to 50 or more. A NOT RUN check is left out of the score altogether, so it is not a failure and does not trigger the reduction. The verdict cannot be ROBUST without it, though: computed with Tradelyze's scoring code, a run whose other three checks are perfect, on 400 days of data, shows 100, grade A+, verdict ACCEPTABLE.

Three checks refuse to run on thin data and score zero instead. Monte Carlo needs at least 3 closed trades, the Permutation Test needs at least 20, and the Deflated Sharpe Ratio needs at least 5. A refused check's badge reads Too Few Trades. A refusal counts as a failed check, because too few trades is a fact about the strategy, so it also triggers the 0.69 reduction. Those minimums only let a check run; they do not make 20 trades enough evidence, as how many trades a backtest needs explains.

FRAGILE means the robustness score is under 50, and MARGINAL means 50 or more but under 70. The amber note names each failed check, so read those rows and act on the one that failed:

  • Parameter Sensitivity reads Unstable. Nudging the winning settings slightly cost 20% or more of the Sharpe ratio on average, or the original Sharpe ratio was not positive. The result depends on exact values: a spike, not a plateau. Search fewer inputs and prefer sensible round values; see plateau versus spike.
  • Permutation Test reads Not Significant or Too Few Trades. Not Significant: the test ran, and the trades did not clearly beat copies whose wins and losses were flipped at random, so the edge may not be there. Too Few Trades: there were fewer than 20 trades, which is a data problem.
  • Monte Carlo shows a Ruin Probability of 20% or more. Redrawn sequences of your own trades broke the firm's total drawdown limit too often. Position size is the first thing to test; see what Ruin Probability measures.
  • Deflated Sharpe Ratio reads Not Significant. The Sharpe ratio does not beat what luck would produce from the number of settings tried. Search less. With fewer than 5 trades it reads Too Few Trades instead.
  • Min Backtest Length reads Insufficient. This row is shown for information only. It earns no points and does not change the score, grade or verdict. It is still a warning sign: there may not be enough years of history for this Sharpe ratio and this many settings tried, and Required Years rises as more settings are tried. History counts for the score through one rule only: with less than 30 days (one month) of data, or a span that cannot be measured, the points are multiplied by 0.69 and the verdict is MARGINAL at best. Those days are counted over your whole upload, first bar to last. The card's Enough for ROBUST (≥ 30 days) row gives that count as days of data and shows Yes or No, or NOT MEASURED when the span could not be measured. Results produced before 26 September 2026 show this box as it was then, when Min Backtest Length was still scored.

A constructed example, not measured data: on one robustness card, Monte Carlo passed and Parameter Sensitivity reads Unstable. The Permutation Test reads Too Few Trades on 17 trades, and the DSR Value of 0.62 reads Not Significant. Three of the four scored rows point to one cause: too few trades and too much searching for the data available. The card reads 1 of 4 checks passed, and the amber note names the three failed checks. Min Backtest Length, which is not scored, points the same way: it reads Insufficient, with 4.1 Required Years against 1.5 Available Years. Raising the trial count would make that worse, because the deflated Sharpe ratio and Required Years both grow stricter as more settings are tried. Each row is explained in what each row on the robustness card means.

What should I do when a prop firm card reads Not Feasible?

When a prop firm card reads Not Feasible, read the first failing row of its Rule Results table and fix that rule first. Not Feasible means at least one rule failed on the backtest. Each row shows a Status mark, the Rule, the Actual value the backtest reached, the Limit and a Message. Tradelyze checks maximum daily drawdown, maximum total drawdown, profit target, minimum trading days, consistency and minimum trade count. It does not check evaluation time limits, news-trading rules or weekend-holding rules.

When a drawdown row fails, test a smaller position size before rewriting the strategy's logic. A drawdown is a fall in account value from a high point, and a firm's drawdown limit is a fixed amount while losses grow with size. Trading smaller also shrinks profit, so check every row after the change.

One strategy on a hypothetical rule set at two position sizes. Constructed illustration, not measured data; the 2-contract column assumes every trade's result scales with contract count
RuleLimitActual at 3 contractsActual at 2 contracts
Maximum total drawdown5%6.8% (fails)4.5% (passes)
Profit target8%11% (passes)7.3% (fails)

In that constructed case, cutting from 3 contracts to 2 scales both figures by two-thirds: 6.8% becomes about 4.5% and 11% becomes about 7.3%. The drawdown rule now passes and the profit target now fails, so the card still reads Not Feasible. Costs that do not scale with size would move the real figures further. A smaller size fixed the rule it was aimed at and exposed the next one, which is why every row needs reading after each change.

To change size, edit the order size in the script's strategy() declaration, for example default_qty_value, because Tradelyze runs the script's own defaults. Run the strategy again in TradingView so the exported trade list matches, then submit again. How to work out a size from a firm's limits is covered in position sizing for prop firm challenges. The daily and trailing rules are covered in daily loss limit and trailing drawdown.

Two further checks belong here. Tradelyze's firm presets are snapshots of each firm's rules, so compare every limit with the firm's current terms on its own site. And do not loosen a custom rule to make the card read Qualifies: the firm applies its real limit, whatever the card says. How to read the Rule Results table has the details.

When should you abandon a strategy instead of tuning it?

Abandon a strategy when the failures remain after the data problems are fixed and the evidence is as large as you can make it. No published rule says when to give up, so the signals below are practitioner judgment, not a standard:

  • The match is fixed and walk-forward still fails after you added history or windows.
  • The edge exists only at one exact setting. Parameter Sensitivity reads Unstable, and nudging an input one step in TradingView turns the result into a loss.
  • All the history you can get still gives too few trades. The Permutation Test refuses to run below 20 trades, and Min Backtest Length reads Insufficient.
  • No position size fits the firm. The only size small enough for the drawdown rules is too small to reach the profit target.
  • You have already re-run many times, and the result moves with every change instead of settling.

Dropping a strategy that fails honestly is a result, not wasted work: you found out before paying a challenge fee or risking an account. A strategy that passes is ready for more checking, not for money. The next step is forward testing on new prices, covered in backtest vs live trading, and the whole sequence is in how to validate a trading strategy. Common product questions are collected in the Learn FAQ.

Check your own report

In a Tradelyze report, each validation check ends in a verdict: Match Rate on the Backtest vs TradingView card, the walk-forward badge, the robustness verdict and each firm's Qualifies or Not Feasible badge. Tradelyze re-runs an uploaded TradingView Pine Script strategy on your price data and checks it against your exported trade list. It then runs parameter optimization, walk-forward analysis, a robustness score from four scored checks and prop-firm rule checks. It does not place trades, give financial advice or guarantee a challenge pass, and it is in beta.

To judge the whole report, not one tile, use the pre-trade checklist. When two verdicts disagree, see which one limits the decision.

Create an account

Already a user? Open your strategies.

Going deeper

The section below goes deeper: the arithmetic and research behind why re-running a strategy until a check passes weakens that check, and when widening a parameter range is fair. You can skip it and still read your own report.

Why is re-optimizing until it passes a trap?

Re-optimizing until a check passes is a trap because it turns the check into part of the search, so the eventual pass is partly luck. A validation check is only useful if the strategy could fail it. If you change settings or ranges after each failure and stop at the first pass, you have searched until the test agreed with you. Statisticians call this the multiple testing problem: the more tests you run, the more likely one passes by chance.

The arithmetic is simple. Suppose each run of a check had a 5% chance of passing a strategy with no real edge, and the runs were independent of each other:

chance of at least one false pass in 10 runs = 1 − (1 − 0.05)10 = 1 − 0.9510 ≈ 40.1%

Re-runs on the same data are not independent, so the real figure is different. But the chance of at least one lucky pass can only grow as you add runs, never shrink. The 5% is an assumption chosen for the arithmetic, not a measured property of any Tradelyze check.

Research on backtests points the same way. Bailey, Borwein, López de Prado and Zhu, in Notices of the American Mathematical Society (May 2014), argue that an impressive backtest is easy to produce by trying enough configurations. They add that the more configurations are tried, the more likely the backtest is overfit. Harvey, Liu and Zhu, in the Review of Financial Studies (2016), make a related point. So many strategies have already been tested, they argue, that a new finding should clear a t-statistic of about 3.0 rather than the conventional 2.0. A t-statistic measures how far a result stands out from noise.

The deflated Sharpe ratio, from Bailey and López de Prado (2014), lowers a Sharpe ratio to allow for how many settings were tried before it was picked. Tradelyze's DSR Value and its permutation test both correct for the distinct settings tried within the current optimization run. Neither can see runs you discarded, strategies you abandoned or ranges you changed between runs. A pass found on the sixth attempt is graded as if it were the first; see the deflated Sharpe ratio.

A constructed example, not measured data: run 1 reads Not Confirmed with WF Efficiency 0.31. You narrow two ranges around the settings that did best in the test windows. Run 2 reads Not Confirmed at 0.44, runs 3 to 5 also fall short, and run 6 reads Confirmed at 0.58. Run 6's report shows only its own trials. The five earlier runs, and everything you learned from their test windows, are invisible to every check on it.

The defense is a written record. Before you start, decide how many runs you will allow and what single change each run tests, and keep the run count next to the result you trade. Why the number of attempts matters more than the number of inputs goes further.

Is it wrong to widen a parameter range and run again?

Widening a parameter range once is not wrong when the Best Value sat at the edge of its Range and the wider values are ones you would really trade. A value at the edge suggests the search wanted to go further, a reading with no primary source. It is still another attempt. Set the new range before looking at any result beyond the old edge, run once, and accept the answer. How to read the Parameter Search Space table covers the columns.

Frequently asked questions about a strategy that fails validation

How do you fix an overfit trading strategy?

You usually fix an overfit trading strategy by searching less, not more. Fix the inputs you would never change, narrow the remaining ranges to values you would really trade, and add history so the test has more trades. Then run the test once and accept the answer. If the strategy only worked at one exact setting, there may be no edge to rescue, and dropping it is the honest fix.

Should I keep re-optimizing until my strategy passes walk-forward?

No. Each re-run after a failure lets the test windows help choose the strategy, so they stop being unseen data. As arithmetic, if each attempt had an independent 5% chance of a false pass, ten attempts would give at least one false pass 40.1% of the time. Tradelyze's checks correct only for the settings tried within one run, so they cannot see the earlier runs you discarded.

What should I do if my walk-forward efficiency is too low?

A low walk-forward efficiency means little of the tuned edge survived on data the optimizer never saw. On Tradelyze's WF Efficiency tile, a value from 0 up to 0.5 reads Likely overfit and a value below 0 reads Inverted — lost out-of-sample. On a run labeled One run, split by period, the tile reads Retention Ratio instead, and a low value means the later periods of the same history were weaker, not that unseen data was tested. Read the Per-Window Results table to see whether one window or every window faded. Then search fewer inputs or add history, and run once.

Does a walk-forward Not Confirmed mean the strategy loses money?

Not necessarily. On a run re-tuned in each window, the badge reads Not Confirmed when WF Efficiency is 0.5 or less, or when 60% or fewer of the usable windows made money. With the default two windows, one losing test window is enough, even when the other did well. Thin evidence gives Inconclusive instead: fewer than two usable windows, more than half the windows excluded, a window with fewer than 5 trades in its test stretch, or a Sharpe ratio that cannot be measured. On a run split by period, the same rule gives Not Consistent.

What should I do if a robustness test fails?

Read which of the four scored checks failed before looking at the score; an amber note on the card names each one. Unstable Parameter Sensitivity points to settings on a narrow spike. A permutation test that reads Too Few Trades points to too little data. A Ruin Probability of 20% or more points to position size. A Not Significant deflated Sharpe ratio points to too much searching. Fix that one cause. Any failed check keeps the score at 69 or below, so C+ and MARGINAL are the best such a run can show. Min Backtest Length is shown for information only and does not change the score, grade or verdict.

Why does Parameter Sensitivity show NOT RUN?

Parameter Sensitivity usually shows NOT RUN because the optimization ran fewer than 50 trials, and Tradelyze's automatic sensitivity budget is zero below that. Tradelyze's automatic trial budget is never fewer than 60, so the usual trigger is a custom Optimization trials count of 20, 30 or 40. NOT RUN is not a failure and does not reduce the score, but it keeps the robustness verdict from reaching ROBUST.

How do I fix a backtest that does not match TradingView?

Check three things in order. First, Chart Timezone must be the zone shown in the bottom-right corner of your TradingView chart. Second, any Properties or Inputs you changed in TradingView must be written into the script's own defaults. Third, the price file must cover the same instrument, timeframe and dates as the chart you exported from. Then re-export and submit again rather than retuning.

Will a smaller position size make my strategy pass a prop firm's drawdown rule?

It can, because a firm's drawdown limit is a fixed amount while losses grow with position size. It is not a guarantee. A smaller size also shrinks profit, so a strategy can move from failing the drawdown rule to failing the profit target. Change the order size in the script's strategy() defaults, re-export from TradingView, submit again, and read every row of Rule Results.

Can I loosen a custom prop firm rule to get Qualifies?

You can edit a custom rule, but loosening its limit changes only what Tradelyze checks, not what the firm enforces. A Qualifies badge earned against a limit wider than the firm's real one says nothing about the challenge you would pay for. Keep custom rules equal to the firm's current published terms, and check the presets against the firm's own site as well.

How many times can I re-run a strategy before the results stop meaning anything?

No published number answers this, because it depends on how different each run was and how much data you have. The direction is known: Bailey, Borwein, López de Prado and Zhu (2014) argue that the more configurations are tried, the more likely a backtest is overfit. Record every run and the one change it tested, and keep that count next to the result you trade.

When should I stop tuning and abandon a trading strategy?

No published rule sets this, so treat it as judgment. Strong signals are a walk-forward result that still fails after the trade match is fixed and history is added, and an edge that exists only at one exact setting. Others are too few trades even over all the history you can get, and a position size small enough for the drawdown rules but too small to reach the profit target.

Is widening a parameter range after a run the same as overfitting?

Not by itself. When a Best Value sits at the edge of its Range, the search may have wanted values it was not allowed to try. Widening that range once, to values you would really trade, is a reasonable step. It is still another attempt. Set the new range before you look at any result, run once, and count the run with the others.

What if the prop firm card says Qualifies but walk-forward fails?

Let the walk-forward result decide. Tradelyze grades a prop firm card's rules on the settings recommended for that firm, over your full history, most or all of which the settings were chosen on. So Qualifies shows little more than that the rules can pass on data the settings were fitted to. On a run re-tuned in each window, Not Confirmed says the tuning did not hold up on data it never saw. On a run split by period, Not Consistent says the settings did worse in some periods of that same history. Read the held-out test on the firm's card too, then decide before paying a challenge fee.

Does resubmitting after a backtesting service outage cost a credit?

No. A card headed Optimization stopped — the backtest could not be run means the run stopped before it measured anything, usually because the backtesting service was unavailable, so nothing was judged. That card's Resubmit button does not use another credit, and backtests that already finished are reused if you resubmit on the same UTC day. If the heading instead names the error the backtest failed with, Resubmit is still free, but fix what the error points at first. Resubmit on a Strategy Processing Failed banner is a different action: it retries that failed run with the same files and settings and costs 1 credit.

Sources

  • David H. Bailey, Jonathan M. Borwein, Marcos López de Prado and Qiji Jim Zhu, Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance, Notices of the American Mathematical Society 61(5), May 2014, pp. 458–471, DOI 10.1090/noti1105. The statement that more configurations tried raise the probability of an overfit backtest is worded most directly in the paper's abstract, which the printed journal version does not carry; the article develops the argument.
  • David H. Bailey and Marcos López de Prado, The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality, Journal of Portfolio Management 40(5), 2014, SSRN 2460551.
  • Campbell R. Harvey, Yan Liu and Heqing Zhu, … and the Cross-Section of Expected Returns, Review of Financial Studies 29(1), 2016, pp. 5–68 — the argument for a t-statistic hurdle of about 3.0 rather than 2.0.
  • Tradelyze implementation, reviewed 26 September 2026: the Optimization stopped card with Continue anyway and Review comparison, and the code note that one faithful re-run can measure 100% against one timezone and about 16% against another; the five-minute entry window for matching trades; Match Rate as matched trades over the larger of the two trade counts compared, within the uploaded price data's first and last bar, with re-run trades before the export's first trade counted as backtest-only; the walk-forward method chosen by the deployment, Re-tuned each window with Confirmed or Not Confirmed, or One run, split by period, the default for eligible strategies, with Consistent or Not Consistent, and the retirement of parameter stability on 26 September 2026; the two pass conditions, WF Efficiency above 0.5 and more than 60% of usable windows profitable; the Inconclusive badge for fewer than 2 usable windows, more than half the windows excluded, a usable window with fewer than 5 test trades or an unmeasurable Sharpe ratio; the NO VERDICT badge; the WF Efficiency labels and the tiles of each method; the held-out test on each firm's card, withholding the last 25% of the history from the whole search by default; 2 walk-forward windows by default, adjustable from 2 to 6 when the settings are enabled for an account; Optimization trials of Automatic or 20 to 300 in steps of 10, with the automatic budget never below 60; the automatic Parameter Sensitivity budget of zero below 50 optimization trials; the 3-trade minimum of Monte Carlo, the 20-trade minimum of the permutation test and the 5-trade minimum of the deflated Sharpe ratio, below which a check reads Too Few Trades; the 20% Ruin Probability limit; the four scored robustness checks and their 29, 29, 24 and 18 points; the minimum backtest length shown for information only; the 30-day minimum history for a ROBUST verdict, counted over the whole upload; the robustness verdict bands; the failure reduction, under which any failed scored check, a refusal for too few trades included, or less than 30 days of data multiplies the points by 0.69, so the score is at most 69, C+ and MARGINAL at best, while a check that did not run triggers nothing; the N of 4 checks passed line and the amber Score reduced note on the robustness card; the robustness scores on this page, computed with Tradelyze's scoring code and given rounded down, as the card shows them; one run, split by period, measuring the search's overall best settings, with an amber banner naming any firm recommended other settings; the Qualifies and Not Feasible badges, the rules checked and not checked, and the Rule Results table, with each firm's verdict graded on the trades of the settings recommended for that firm over the full history, including the held-out part when the held-out test runs; the card headed Optimization stopped — the backtest could not be run, or Optimization stopped — backtest failed with a named error, whose Resubmit opens the Run Optimization dialog, uses no further credit and, when no error is named, reuses finished backtests cached for the same UTC day, and the backend note that retrying an aborted or baseline-halted run is free; the Processing is taking longer than expected banner, shown with Cancel & Delete once 5 hours pass without a result; the 1-credit charge for a new submission and for Resubmit on a failed run, and the absence of any refund when a strategy is deleted; and the correction of the permutation test and deflated Sharpe ratio for distinct settings tried within one optimization run only.
  • Arithmetic used on this page, requiring no source: 1 − 0.9510 ≈ 0.401, which assumes independent runs and an illustrative 5% false-pass chance; 0.81 ÷ 1.30 ≈ 0.62; 6.8% × 2/3 ≈ 4.5%; 11% × 2/3 ≈ 7.3%. Every worked example on this page is a constructed illustration, not measured data.
  • The order in which to work through failures, the next steps in the results table, which verdict limits the decision when two disagree, and the signals for abandoning a strategy are practitioner judgment with no primary source.

Related terms

Tradelyze

Last reviewed 26 September 2026. Educational content about backtest validation methodology. Nothing here is financial advice.