Glossary
Robustness score
Last reviewed: 26 September 2026·Tradelyze
A robustness score is one 0–100 number that adds up several separate stress tests of a backtest. Tradelyze's score adds four tests worth 29, 29, 24 and 18 points. Each test gives partial credit, but since 26 September 2026 a failed test, or less than 30 days of data, multiplies the points by 0.69. In a constructed example on this page a strategy narrowly failing all four earns 75.6 points and scores 52.1, a D+. A fifth check, minimum backtest length, is shown for information and earns no points. Read each result before the total.
| Check and failing result | It usually means… | What to try |
|---|---|---|
| Monte Carlo: Ruin Probability of 20% or more, or Too Few Trades | Redrawn sequences of your own trades broke the firm's total drawdown limit too often. Too Few Trades means fewer than 3 closed trades, too few to redraw. | Position size is the first thing to test; see what Ruin Probability measures. For Too Few Trades, test on more history. |
| Permutation test: Not Significant, or Too Few Trades | The trades did not clearly beat copies whose wins and losses were flipped at random. A p-value is the chance of a result at least this good if the strategy had no real edge. Here it is the share of flipped copies that did as well. | Too Few Trades means fewer than 20 closed trades: a data problem, so test on more history. Not Significant means the test ran, and the edge may not be there. |
| Parameter sensitivity: Unstable | Nudging the winning settings slightly cost 20% or more of the Sharpe ratio on average, or the original Sharpe ratio was not positive. The result depends on exact values: a spike, not a plateau. | Search fewer inputs and prefer sensible round values; see plateau versus spike. |
| Deflated Sharpe ratio: Not Significant, or Too Few Trades | The Sharpe ratio does not beat what luck would produce from the number of settings tried. Too Few Trades means fewer than 5 closed trades. | Search less. For Too Few Trades, test on more history. |
| Minimum backtest length (not scored): Insufficient | There are not enough years of history for this Sharpe ratio and this many settings tried. The check is shown for information only and does not change the score, grade or verdict. | Read it as a warning sign: the Sharpe ratio may come from a wide search on too little history. Adding history, or searching less, addresses that, because Required Years rises as more settings are tried. |
| Enough for ROBUST (≥ 30 days): No | Your whole upload covers less than 30 days (one month). The points are multiplied by 0.69, so the score is 69 or below: C+ and MARGINAL at best, however well the checks did. | Test on at least 30 days of history. |
A failed check in any of the first four rows costs more than its own points: the whole score is multiplied by 0.69, so it stays at 69 or below. The Insufficient badge in the fifth row is the exception and changes nothing. A check can also show NOT RUN, which means it did not happen. That is not a failure and does not reduce the score, but it rules out ROBUST. What to do when a robustness check shows NOT RUN or FRAGILE covers that case too.
In plain English
A robustness score squeezes several different tests of a strategy into one number, so many strategies are easy to sort. A failed check, or less than a month of data, pulls the whole score down to 69 or below. So a score of 70 or more means nothing failed and there is at least a month of data, but a lower score does not say which check failed, or whether one did. Read each check on its own before you trust the total.
New to this? Start with Monte Carlo simulation.
What is a trading strategy robustness score?
A robustness score is a composite. Several separate statistical tests run on a finished backtest, which is a test of trading rules on past prices. Each test earns points, and the points combine into one 0–100 figure.
Its honest job is sorting. With a hundred candidate strategies, reading four test results for each is slow, and one number that ranks them is useful.
Its dishonest job is acting as a pass mark. Tradelyze stops a strong total from sitting on top of a failed check: since 26 September 2026 any failed check multiplies the score by 0.69. But the total still does not say which check failed. It cannot see the ideas you tried and dropped, and none of its checks looks at data the settings were not chosen on. Used alone to decide whether to risk money, the score still misleads.
Tradelyze runs five checks and scores four of them. The fifth, minimum backtest length, is shown for information only: it earns no points and does not change the grade or the verdict. Every check except parameter sensitivity works only on the list of trades the winning settings produced. Parameter sensitivity re-runs the backtest with slightly changed settings. It runs automatically only when the optimization ran at least 50 trials. A trial is one tested combination of settings.
| Check | Points | Question it answers | Passes when |
|---|---|---|---|
| Monte Carlo | 29 | Would a less lucky mix and order of the same trades have broken your drawdown limit? | Ruin Probability is below 20% |
| Permutation test | 29 | Did the long and short calls beat coin flips? | Random coin-flip versions of your trades rarely did as well (the bar gets stricter the more settings the optimizer tried). |
| Parameter sensitivity | 24 | Do slightly different settings still work? | Degradation is below 20% and the original Sharpe ratio is positive |
| Deflated Sharpe ratio | 18 | Does the Sharpe ratio survive the number of settings tried? | DSR Value is over 0.95 |
| Minimum backtest length | Not scored | Is there enough history to trust this Sharpe ratio? | Available Years is at least Required Years. Shown for information only |
The Sharpe ratio in those questions is average return divided by how much returns swing around. The weights, the pass conditions and the 0.69 reduction for a failed check are Tradelyze's own choices. No published research sets weights for a composite of these four tests. Each test has literature behind it, cited in its own section; the arithmetic that combines them has none.
Why can a composite robustness score mislead?
A composite robustness score can mislead in three ways. Partial credit can give failed checks a respectable total; Tradelyze now stops that with a reduction, but the total still does not say which check failed. A check that did not run changes what the total means. And the checks overlap, so their points are not independent evidence.
Each scored Tradelyze check earns points on a sliding scale, out of its weight of 29, 29, 24 or 18. Tradelyze then rescales the total over the checks that actually ran. If any scored check failed, or the data covers less than 30 days, it multiplies the result by 0.69. The exact point rules are under how Tradelyze turns each robustness check into points.
What stops partial credit from hiding a failed check?
None of the four scored checks scores all or nothing. A check that misses its pass condition by a hair still earns nearly its full points, and four near-misses add up to a comfortable total. Until 26 September 2026 that total was the score, so a run could fail every check and still read B and ACCEPTABLE. Now any failed check multiplies the points by 0.69.
The constructed run in the next table fails all four scored checks narrowly. For the permutation test, the bar is the P-Value it must get under. Tradelyze tightens that bar to 0.00085 when 60 distinct settings were tried, as the corrected bar explains.
| Check | Value | Pass condition | Points awarded |
|---|---|---|---|
| Monte Carlo | Ruin 20.0% | under 20%: failed | 17.4 of 29 |
| Permutation test | p = 0.0017 | under a bar of 0.00085: failed | 26.2 of 29 |
| Parameter sensitivity | Degradation 20.0% | under 20%: failed | 14.4 of 24 |
| Deflated Sharpe ratio | DSR 0.940 | over 0.95: failed | 17.6 of 18 |
| Minimum backtest length | 90% of required | information only | not scored |
| Points earned | 0 of 4 scored checks passed | 75.6 of 100 | |
| Score, grade and verdict | points × 0.69, because checks failed | 52.1 → D+ → MARGINAL |
That run failed every scored check. Partial credit still gave it 75.6 points, and the reduction turns them into a score of 52.1. The card rounds the score down to a whole number, so it shows 52, a D+ grade and a MARGINAL verdict. Under the verdict it reads 0 of 4 checks passed, and an amber note says the score was reduced from 75.6 to 52.1 and names the four failed checks. Without the reduction the same run would score 75.6, a B, and read ACCEPTABLE. The sensitivity row assumes a smooth neighborhood with no spike deduction.
Partial credit still does a job. Inside each check, points fall smoothly as the result gets worse, and the reduction is a multiplier rather than a flat cap. So runs that failed keep the order their points put them in, and a narrow miss still scores above a wide one. A flat cap would tie them all at 69. The factor is 0.69 because that keeps every run with a failed check below 70, the ACCEPTABLE line, and so below every grade from B- up.
| Ruin Probability | Points earned | Score (card) | Grade | Verdict |
|---|---|---|---|---|
| 19.9% | 88.5 | 88.5 (88) | A- | ROBUST |
| 20% | 88.4 | 61.0 (61) | C | MARGINAL |
| 30% | 82.6 | 57.0 (57) | C- | MARGINAL |
| 50% or more | 71.0 | 49.0 (49) | D | FRAGILE |
The first two rows show the cost of the rule. A tenth of a percentage point of Ruin Probability separates an A- and ROBUST from a C and MARGINAL, because 20% is where the check fails. That cliff is deliberate: past it the check failed, and the score now says so. Past the cliff, the order holds: 20% scores above 30%, and 30% above 50%. The score is published to one decimal, 61.0, and the card shows its whole-number part, 61. The amber note and Show Warnings give the same 61.0.
Tradelyze added the reduction because a high grade could sit beside a failed check. In another constructed example, a run that passed everything except a DSR Value of 0.94 earned 99.6 points and would have shown an A+. It now scores 68.7, which the card shows as 68: a C+, and MARGINAL.
How does a check that did not run change the score?
Tradelyze treats a check that did not run in two ways at once. The score leaves it out. The check earns no points and costs none, and the checks that did run are scaled up to fill 100. It is not a failure, so it does not trigger the 0.69 reduction. The verdict counts it against ROBUST. With a scored check not run, the best reachable verdict is ACCEPTABLE.
With parameter sensitivity not run, a score of 90 is 68.4 points out of the 76 the other three checks can give. It is not 90 points out of 100.
A check that ran but refused because the strategy closed too few trades is different. It stays in the score at zero, because too few trades is a fact about the strategy rather than a gap in the testing. It also counts as a failed check, so the reduction applies.
| Sensitivity check | Points earned | Score | Grade | Verdict |
|---|---|---|---|---|
| Did not run (NOT RUN) | 100 | 100 | A+ | ACCEPTABLE |
| Ran, 0% Degradation, Stable | 100 | 100 | A+ | ROBUST |
| Ran, 20% Degradation, Unstable | 90.4 | 62.4 | C | MARGINAL |
| Ran, 50% Degradation, Unstable | 76.0 | 52.4 | D+ | MARGINAL |
Check this on your own results
In the four-way example, find the strategy whose sensitivity check did not run: it scores 100. The same strategy with the check run and failed at 20% Degradation earns 90.4 points, and the reduction takes its score to 62.4, which the card shows as 62. A missing check can raise the score a long way, because a check that did not run cannot fail. What a missing check cannot do is earn ROBUST: the not-run version stops at ACCEPTABLE.
Before comparing two scores, read the line under each verdict. The first strategy shows 3 of 4 checks passed · 1 not run; the third shows 3 of 4 checks passed and an amber reduction note. Show Warnings names any missing checks and says the score rests on less evidence. A score from three checks is not comparable with one from four.
There is a third, quieter issue: the checks are not independent in the way adding their points implies. The permutation test and the deflated Sharpe ratio both rest on a Sharpe ratio from the same trade list. One unusually large winning trade moves both at once. Adding two correlated scores does not give twice the evidence. Minimum backtest length rests on the same Sharpe ratio too, but it is shown for information and adds no points.
What do the robustness grades from A+ to F mean?
The Grade on the Tradelyze robustness card is the 0–100 score relabeled as a letter. It runs from A+ at 95 or more down to F under 40. A grade of B- or better always means no scored check failed and the data covers at least 30 days. A failed check, or less than 30 days of data, multiplies the points by 0.69, and the best that can then show is 69.0, a C+. The card colors the Grade pill by letter family.
| Grade | Score | Source |
|---|---|---|
| A+ | 95 or more | Tradelyze implementation. No primary source. |
| A | 90 or more | Tradelyze implementation. No primary source. |
| A- | 85 or more | Tradelyze implementation. No primary source. |
| B+ | 80 or more | Tradelyze implementation. No primary source. |
| B | 75 or more | Tradelyze implementation. No primary source. |
| B- | 70 or more | Tradelyze implementation. No primary source. |
| C+ | 65 or more | Tradelyze implementation. No primary source. |
| C | 60 or more | Tradelyze implementation. No primary source. |
| C- | 55 or more | Tradelyze implementation. No primary source. |
| D+ | 50 or more | Tradelyze implementation. No primary source. |
| D | 45 or more | Tradelyze implementation. No primary source. |
| D- | 40 or more | Tradelyze implementation. No primary source. |
| F | under 40 | Tradelyze implementation. No primary source. |
With all four checks run, the line is sharp. The lowest score a run can reach while passing all four is about 71.6, a B-, so on that basis B- or better means everything passed, and C+ or lower means something failed or the data is short. The constructed near-miss run misses all four scored pass conditions by a hair and grades D+.
The grade still inherits the score's weaknesses. B- or better does not mean all four checks ran: a check that did not run is left out, so a run with one check not run can grade A+, as the four-way example shows. And when a check did not run, C+ or lower does not by itself mean a check failed. A constructed run with the permutation test NOT RESOLVED, the other three checks passing narrowly and the full spike deduction on parameter sensitivity scores 60.1, a C, with nothing failed. The line under the verdict and the amber reduction note tell the cases apart.
The bands also have no primary source. They are school grading intervals applied to a scale Tradelyze chose. Nothing shows that a strategy scoring 90 is meaningfully different from one scoring 85. The 0.69 factor has no primary source either; Tradelyze chose it so that every run with a failed check lands below 70.
The grade is most useful as a sort key across many strategies run through Tradelyze with the same settings, where the arbitrary parts cancel out. It is least useful as a threshold for risking money.
What do ROBUST, ACCEPTABLE, MARGINAL and FRAGILE mean?
The verdict on the Tradelyze robustness card is a one-word label: ROBUST, ACCEPTABLE, MARGINAL or FRAGILE. ROBUST is the only label that requires every scored check to have run and passed. The other three read the score alone, but since 26 September 2026 the score carries the checks: a failed check or less than 30 days of data multiplies it by 0.69, so such a run reads MARGINAL at best.
ROBUST needs all four scored checks run and passed, at least 30 days (one month) of data and a score of at least 80. ACCEPTABLE needs a score of 70 or more, which only a run with no failed check and at least 30 days of data can reach. MARGINAL needs 50 or more. Anything lower is FRAGILE.
| Verdict | Condition | What it tells you | Source |
|---|---|---|---|
| ROBUST | All four scored checks ran and passed, the data covers at least 30 days (one month), and the score is 80 or more | The only verdict that requires every scored check to run and pass. A check that did not run, could not be resolved or could not be computed rules it out. So does less than 30 days of data, or data with no measurable time span. | Tradelyze implementation. No primary source. |
| ACCEPTABLE | Not ROBUST, and the score is 70 or more | No scored check failed and the data covers at least 30 days. Either a scored check did not run, or all four passed with a score under 80. | Tradelyze implementation. No primary source. |
| MARGINAL | Score of 50 or more, under 70 | The best a run with a failed check, or under 30 days of data, can get. A run with a check not run can also land here without failing anything. | Tradelyze implementation. No primary source. |
| FRAGILE | Score under 50 | Score only. The reduction only ever lowers a score; it never lifts a run into a better verdict. | Tradelyze implementation. No primary source. |
A NOT RUN check caps the verdict at ACCEPTABLE
If any of the four scored checks did not run, the best verdict Tradelyze can show is ACCEPTABLE, however high the score. That includes a grey NOT RUN parameter sensitivity check and a NOT RESOLVED permutation test. It also includes a deflated Sharpe ratio that could not be computed. A check that did not run is not a failure, so it does not reduce the score. When the optimization ran fewer than 50 trials, parameter sensitivity does not run automatically, so ROBUST is out of reach on that run.
A failed check caps the score at 69 and the verdict at MARGINAL
If any scored check ran and failed, the points are multiplied by 0.69. The score is then 69.0 at most, the grade C+ at best and the verdict MARGINAL at best. A check refused for too few trades counts as failed. The card says so in view: an amber note gives the score before and after the reduction and names the checks that failed. The reduction is applied once, however many checks failed.
Less than 30 days of data caps the score at 69 and the verdict at MARGINAL
The 30-day rule earns no points of its own. When the data covers less than 30 days, the points are multiplied by 0.69, as for a failed check. In a constructed example, a run that earns full points on every check with only 20 days of data scores 69.0: a C+, and MARGINAL. Data with no usable dates has no measurable span and scores the same. With at least 30 days, the same run scores 100, an A+, and ROBUST. The rule reads your whole upload, first bar to last, including any part held back for the held-out test. The card shows it in the Min Backtest Length box, as Enough for ROBUST (≥ 30 days) · N days of data, where N is the length of the whole upload. No, or NOT MEASURED, means the score was reduced. That box's own Sufficient or Insufficient badge does not affect the score, grade or verdict.
One month is a Tradelyze starting value with no primary source, and it may rise later.
The labels still hide which check decided them. The constructed near-miss run fails all four scored checks and reads MARGINAL. The four-way example's run that failed only parameter sensitivity, at 50% Degradation, reads MARGINAL too, and the card shows 52 for both. Read ACCEPTABLE as "nothing failed, but not ROBUST", and read MARGINAL as a prompt to find out what failed.
The robustness verdict is also separate from the walk-forward card's badge. That badge reads Confirmed or Not Confirmed when each window is re-tuned, Consistent or Not Consistent when one run is split by period, Inconclusive when the evidence is too thin, or NO VERDICT. A strategy can read ROBUST here and Not Confirmed there, or the other way round, because the two cards test different things.
Card reads FRAGILE, or a check shows NOT RUN? See the next step for each check.
How do I read my Tradelyze robustness card?
Read the Tradelyze robustness card from the check rows up to the score, not the other way round. The score is the least informative number on the card. It adds together four answers to unrelated questions.
At the top, the card shows the score out of 100 and a Grade pill, both colored by the grade's letter family, and the verdict. The score is rounded down to a whole number, so it always sits in the same band as the grade and verdict beside it: 84.8 shows as 84, beside a B+. Under the verdict a line counts the scored checks, for example 3 of 4 checks passed, with · 1 not run added when a check did not run. When the score was reduced, an amber note says so in view: it gives the score before and after the 0.69 reduction, to one decimal, and names the checks that failed, or the short data. A failed check or under a month of data keeps a run at C+ and MARGINAL at best, and the note says that too.
Below that is one box per check: Monte Carlo, Permutation Test, Parameter Sensitivity, Deflated Sharpe Ratio and Min Backtest Length. The Min Backtest Length box is marked Not scored. When you selected a prop firm, the card names that firm, because the Monte Carlo part is judged against the firm's own drawdown limit. It can pass for one firm and fail for another, so the reduction can apply on one firm's card and not another's. If a chance of ruin cannot be worked out for that firm, the card shows Score not available instead of a number. A Show Warnings button lists anything that failed, refused or did not run, and gives the exact reason for any reduction.
One more amber banner can appear. The robustness checks run once, on the search's overall best settings, and a firm may be recommended other settings. When that happens the banner says the checks were run on the search's overall best settings, not on the settings recommended for that firm. They still describe the strategy, but not those exact settings. When a held-out test ran, the banner adds that it does use the recommended settings. The held-out test does not always run: it is skipped on short data, can be switched off and can fail to run.
The email Tradelyze sends when a run finishes shows the grade, the verdict and how many of the four checks passed, and it says when the score was reduced.
Why read the check results before the score?
Because the score does not say which check failed. In the constructed near-miss example, a run that fails all four scored checks scores 52.1. In the four-way example, a run that failed only parameter sensitivity, at 50% Degradation, scores 52.4. The card shows 52, D+ and MARGINAL for both. Only the check boxes, and the line under the verdict, tell one failed check from four.
What does each row on the robustness card mean?
| Card label | Question it answers | Pass condition in Tradelyze | What NOT RUN or NOT RESOLVED means | Source |
|---|---|---|---|---|
| Monte Carlo: MC Sharpe Mean, MC Sharpe 5th–95th, MC Sharpe Original, MC Max DD Real→P95, Ruin Probability, Ruin Check (Ruin Acceptable / Ruin Too High / Too Few Trades) | With the same trades in a less lucky mix and order, how often would the drawdown have broken the firm's limit? How deep could it have gone? | Ruin Probability below 20%: the Ruin Check badge reads Ruin Acceptable in green, and Ruin Too High in red at 20% or more. MC Max DD Real→P95 shows the drawdown that actually happened, then the 95th-percentile drawdown across the draws. | Dashes mean no figure. NOT RUN on the Ruin Check badge means the check was switched off. When the strategy closed fewer than 3 trades, the check was refused: Ruin Probability is a dash and the badge reads Too Few Trades, in red. That is not a measured ruin, but the check scores zero and counts as failed, so the score is reduced. When no firm drawdown limit could be set, the card shows Score not available. Results from before 26 September 2026 have no Ruin Check badge. | Tradelyze implementation |
| Permutation Test: P-Value, Significant / Not Significant / Too Few Trades | Were the long and short calls better than coin flips? | Significant when the P-Value is below a 0.05 bar made stricter for the number of settings tried (see the corrected bar). Needs at least 20 closed trades. | NOT RESOLVED: the flips could not reach the corrected bar, usually because the search was very wide, or the price data had no usable dates. The answer is unknown, not failed: the check leaves the score without triggering the 0.69 reduction, and ROBUST is out of reach. Price data with no usable dates still triggers the reduction, through the 30-day rule. NOT RUN: the test was switched off. Under 20 trades the check was refused: the P-Value is a dash, the badge reads Too Few Trades, in red, and the check scores zero and counts as failed, so the score is reduced. Not Significant means the test ran and missed its bar. Results from before 26 September 2026 show a refusal as Not Significant. | Tradelyze implementation |
| Parameter Sensitivity: Degradation, Mean Δ (signed), Runs / Noise, Stable / Unstable | Do slightly changed settings still work? | Stable when Degradation is below 20% and the original Sharpe ratio is positive. Unstable is a failed check. | NOT RUN: no re-tests happened, for example because the optimization ran fewer than 50 trials. The check leaves the score without triggering the reduction, and ROBUST is out of reach. | Tradelyze implementation |
| Deflated Sharpe Ratio: DSR Value, Significant / Not Significant / Too Few Trades | Is the Sharpe ratio still convincing after allowing for how many settings were tried? | Significant when the DSR Value is over 0.95. Needs at least 5 closed trades. | Under 5 trades the check was refused: DSR Value is a dash, the badge reads Too Few Trades, in red, and the check scores zero and counts as failed, so the score is reduced. A dash with a grey NOT RESOLVED badge means no deflated Sharpe could be computed. Too few distinct settings leaves it out of the score as not run, with no reduction. Price data with no usable dates also leaves it out, but triggers the reduction through the 30-day rule. Either way ROBUST is out of reach. Results from before 26 September 2026 show a refusal as Not Significant. | Tradelyze implementation |
| Min Backtest Length (Not scored): Required Years, Available Years, Sufficient / Insufficient, Enough for ROBUST (≥ 30 days) | Is there enough history to trust this Sharpe ratio, given how many settings were tried? Is there at least a month of history, so the score is not reduced? | Sufficient when Available Years is at least Required Years. Required Years is never under one year. That badge is grey: the check earns no points and does not change the score, grade or verdict. Enough for ROBUST reads Yes when your whole upload covers at least 30 days (one month). The row also gives that length, as · N days of data. No means the points were multiplied by 0.69: C+ and MARGINAL at best. Available Years is the span the checks ran on, which can be shorter than the upload when part of it was held back for the held-out test. | NOT RUN: the price data had no usable dates to measure a span from. This check is not scored, but Enough for ROBUST then reads NOT MEASURED, which counts the same as No: the points are multiplied by 0.69. Results from before 26 September 2026 show the box as it was then: no Not scored label, a colored badge and no Enough for ROBUST row. | Tradelyze implementation |
Start with the closed trade count. Tradelyze refuses to run the permutation test under 20 closed trades and the deflated Sharpe check under 5. Each refusal scores zero and counts as a failed check, so the score is multiplied by 0.69 and stays at 69 or below.
Then count how many checks ran. This is the step most readers skip. The line under the verdict counts them, for example 3 of 4 checks passed · 1 not run. Open Show Warnings for the detail. When fewer than four scored checks ran, a warning names the missing checks and says the score is out of fewer checks. It also says no strategy can be certified ROBUST without all four. A score built from three checks is not comparable with one built from four.
Then read why the score was reduced, if it was. Show Warnings gives the exact reason, for example: "Robustness score reduced from 88.4 to 61.0 (x0.69) because the Monte-Carlo shuffle check failed." For short data the reason reads "the data covers only 9.4 days (at least 30 days, one month, are needed)". The warning ends by saying that any failed check, or less than 30 days of data, keeps the score at 69 or below.
Then read the five boxes separately. They answer unrelated questions. Did your drawdown depend on a lucky sequence? Did your direction calls beat coin flips? Do your settings have to be exactly right? Does your Sharpe ratio survive the number of trials behind it? Those four are scored. The fifth box asks whether you have enough history: its Sufficient or Insufficient badge is for information only, while its Enough for ROBUST row shows the 30-day rule. A failure in any one of them is specific and actionable; the total is neither.
Then check the shape around your settings. Look for a parameter plateau rather than a spike. Use a parameter sweep, or, when the walk-forward check re-tuned every window, read each window's best settings in the walk-forward per-window table. This catches strategies that pass every aggregate check and still balance on a knife edge.
Then check results on unseen data. None of the five robustness checks looks at data the optimizer did not tune on. All of them analyze the trade list from the tuned result. That separate question is answered by the held-out test, and by a walk-forward analysis that re-tunes every window, summarized by walk-forward efficiency. The walk-forward method used by default for eligible strategies, one run split by period, is not an out-of-sample test: its settings were chosen on the same history.
Then count your attempts. Write down how many strategy ideas you tested and how many times you re-ran the pipeline. Also note how many times you changed parameter ranges. No report contains that number, and it changes what every number on the card is worth.
What does the Monte Carlo resampling check actually test?
The Monte Carlo check asks how bad your drawdown could have been with the same trades in a less lucky mix and order. A drawdown is the fall in account equity from a peak to a later low. Tradelyze redraws runs of your real trades 1,000 times by default. Some trades repeat and some are left out.
Runs of consecutive trades stay together in each draw, so a losing streak stays a streak. Tradelyze uses a fixed random seed, so the same trades give the same Monte Carlo numbers on every run. How the draws are built is covered under Going deeper.
In plain words: each draw becomes an equity curve, and its deepest fall is recorded. Ruin Probability is the share of draws whose fall broke the prop firm's limit.
The two outputs worth reading first are:
- Ruin Probability: the share of draws whose drawdown broke the total drawdown limit of the prop firm the card is shown for. Each selected firm gets its own figure, counted from the same draws. Tradelyze passes the check when Ruin Probability is under 20%, a Tradelyze choice with no primary source. The box's Ruin Check badge reads Ruin Acceptable in green when it passes and Ruin Too High in red when Ruin Probability is 20% or more. The card rounds Ruin Probability down to one decimal, so a passing 19.97% shows as 19.9%, never as 20.0%. If the strategy closed fewer than 3 trades, the check is refused and the badge reads Too Few Trades: a failed check, but not a measured ruin. Because each firm's limit is its own, the check can pass for one firm and fail for another, and the 0.69 reduction can apply on one firm's card and not another's. The other 71 points are the same for every firm. When no limit can be set for a firm, the figure is left blank rather than borrowed from another limit.
- MC Max DD Real→P95: first the drawdown of your trades in the order they happened, then the 95th-percentile drawdown across the draws. A percentile is the value a given share of draws fell under, so only 5% of draws went deeper than the 95th-percentile figure. This is the number to size an account against. It estimates how deep the drawdown could plausibly have gone with a less kind sequence of the same trades.
Ruin Probability counts breaches of the total drawdown limit only. A daily loss limit cannot be simulated this way, because resampling destroys the calendar. A strategy can pass this check and still fail a firm on one bad day. The daily limit is checked on the real backtest in that firm's Rule Results.
Resampling cannot detect anything outside the trades you got. That includes a trade you did not take, a market regime missing from your sample, or a loss larger than any in your history. The check draws from the history you have; it cannot invent history you do not.
Compare the MC Sharpe band with MC Sharpe Original, not with the headline Sharpe
Every Monte Carlo Sharpe figure spreads the drawn trades evenly across the run's bars, because a resample has no real exit times. The headline Sharpe figures on the Best Metrics card use your trades' real exit times. The two agree while trades are spread out. According to Tradelyze's tooltip for MC Sharpe Original, they can differ by tens of percent in either direction once trades cluster on the same bars.
MC Sharpe Original is the same trade list measured the same way as the band. A narrow MC Sharpe 5th–95th band means the result does not rest on a few trades; a band that sinks toward zero means it does. The Monte Carlo page explains how to read the MC Sharpe band.
What does the permutation p-value mean, and what does it not?
The P-Value on the Tradelyze robustness card is the share of coin-flip versions of your trades whose Sharpe ratio matched or beat yours. A small P-Value means coin-flip direction calls rarely did as well as yours. A large one means they often did.
A p-value, in general, is the probability of a result at least this good if a stated "nothing going on" assumption were true. Here the assumption is a strategy that entered at your exact times but picked long or short at random.
The test keeps every trade's size and entry time. It randomly flips each trade's result from profit to loss, or loss to profit, with even odds. The Sharpe ratio is recomputed on each flipped version and compared with yours. The badge reads Significant when the P-Value clears the bar, and Not Significant when it does not.
The bar starts at 0.05. Tradelyze makes it stricter the more distinct settings the optimizer tried, because the best of many tries clears a fixed bar by luck. With 60 settings tried, the bar is about 0.00085. The formula and the flip counts are under Going deeper.
What does NOT RESOLVED mean on the permutation test?
NOT RESOLVED usually means the search was so wide that even 10,000 flips could not express a P-Value small enough for the bar. No flipped version beat your real result. That is the strongest evidence this test can produce, but it still cannot say whether the bar was cleared. The badge also reads NOT RESOLVED when the price data had no usable dates: then no Sharpe ratio could be measured, the P-Value is a dash, and no comparison was made.
Tradelyze leaves a NOT RESOLVED check out of the score rather than scoring it as a failure, so it does not trigger the 0.69 reduction. Because the check did not pass, ROBUST is out of reach on that run. In a constructed example with Ruin Probability 5%, 5% Degradation and a DSR Value of 0.97, a run with a NOT RESOLVED permutation test scores 92.5, which the card shows as 92: an A, and ACCEPTABLE.
What does the P-Value not mean?
- The P-Value is not the probability that the strategy is profitable, or that it will be. It is computed under an assumption you do not believe, which is why p-values are so often misread.
- The P-Value is not "the chance the result was luck" in the everyday sense. It addresses one kind of luck: getting the direction right by accident.
- The test checks direction, not trade size or timing. Every flipped version keeps your exact trade sizes. A strategy whose profit rests on a few huge winners keeps those huge moves in every version. As Tradelyze's tooltip for the P-Value points out, that makes this test easier to pass than one that re-picks the entries.
- The P-Value knows nothing about ideas you discarded. The bar allows only for the settings tried in this optimization run. The deflated Sharpe ratio attacks the same selection problem from a different direction.
- The 0.05 starting point is a convention. A P-Value of 0.049 and one of 0.051 are not different findings. Tradelyze gives full points under the bar and partial points over it, but a P-Value that misses the bar by any margin is a failed check, and a failed check multiplies the whole score by 0.69. So the score has a cliff at the bar, just as the Significant or Not Significant badge does.
There is one harsher ending. Under 20 closed trades Tradelyze refuses to run the test. The P-Value shows a dash, the badge reads Too Few Trades, and the check scores 0 of its 29 points. The refusal counts as a failed check, so the whole score is multiplied by 0.69 as well. Five trades allow only 32 possible sign patterns, too few for a fine-grained P-Value.
What does parameter sensitivity testing detect?
The Parameter Sensitivity box on the Tradelyze robustness card shows how much the Sharpe ratio falls when the winning settings are nudged slightly. Low Degradation with a Stable badge means nearby settings still work. High Degradation with an Unstable badge means the result depends on exact values.
Tradelyze adds a small random nudge to every numeric setting, re-runs the backtest, and averages the Sharpe ratio across the re-runs. Degradation is how far that average fell, as a percentage of the original Sharpe ratio. The check reads Stable when Degradation is under 20% and the original Sharpe ratio was positive. The 20% line is a Tradelyze threshold with no primary source.
When the check shows NOT RUN
Parameter sensitivity is the only check that re-runs backtests, so it is the one most likely to be skipped. Its automatic re-test count is zero for searches below 50 trials. On a small trial budget the check shows a grey NOT RUN badge rather than a red Unstable badge, and ROBUST is out of reach.
The check also shows NOT RUN in three other cases. No numeric setting could be varied, every re-run failed, or every nudge landed back on the winning settings. A NOT RUN check is left out of the score: nothing was measured, so nothing failed, and the 0.69 reduction does not apply.
Two more labels sit in the box. Mean Δ (signed) is the same comparison with its sign kept. A positive value means the nudged settings did better on average than the winning ones, so the optimizer stopped short of a nearby peak. Runs / Noise shows how many re-runs the verdict rests on and how hard each setting was nudged.
Three scoring rules change what the box is worth. A spiky neighborhood, where re-run results scatter widely, costs up to half of this check's points. A strategy whose original Sharpe ratio is zero or negative earns none of the 24 points. And an Unstable badge is a failed check, so the whole score is multiplied by 0.69; a spike deduction on its own does not make the badge Unstable. The rest of Tradelyze's nudging rules are under Going deeper.
How do I tell a parameter plateau from a parameter spike?
Underneath the Degradation percentage is a question about shape. Plot backtest performance against one setting's value, holding the rest fixed. The shape around the chosen value tells you most of what the sensitivity number summarizes.
Stated as a rule: a setting surrounded by similar performance is weak evidence of signal, and one whose neighbors degrade sharply is noise. A spike is close to conclusive, because no plausible market mechanism works at 14 and fails at 13 and 15.
A plateau is only weak evidence in the other direction. A broad flat region can equally mean the setting barely influences the strategy. A strategy full of settings that do nothing will score beautifully on a sensitivity check.
The same test works across walk-forward runs, and it is often easier to read there. Build Alpha gives an example: a parameter that comes back as 12, then 47, 23, 35, 55 and 13 across successive runs. That pattern is most likely curve fitting. The optimizer is finding a different peak in a noisy surface every time. When the walk-forward check re-tunes every window, Tradelyze shows each window's best settings in the expanded rows of the walk-forward per-window table.
What is the deflated Sharpe ratio and why does it exist?
The DSR Value on the Tradelyze robustness card is the deflated Sharpe ratio. It is the chance, from 0 to 1, that your Sharpe ratio beats what luck alone would produce. Over 0.95 reads Significant; 0.95 or under is a failed check. A value near or under 0.5 means a search that size could have found your Sharpe ratio with no edge at all.
The deflated Sharpe ratio exists because of an unavoidable problem. If you test enough settings, the best one will show a high Sharpe ratio whether or not any edge is present. That is arithmetic, not bad practice. Draw enough samples centered on zero and the largest one is comfortably above zero. An optimizer's job is to find that largest one, so an optimizer will produce an impressive number from noise.
David H. Bailey and Marcos López de Prado formalized the correction in The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality (Journal of Portfolio Management 40(5), 2014, SSRN 2460551). First, estimate how high a Sharpe ratio the best of N trials would reach with no skill at all. Then ask how far your Sharpe ratio clears that benchmark, given how precisely the data measures it.
Read the output as a probability between 0 and 1, not as a Sharpe ratio. A DSR Value of 0.97 does not mean a Sharpe ratio of 0.97.
Two further properties decide what the number can detect:
- The trial count comes from this optimization run only. It does not include strategies you abandoned, runs you discarded, or ideas you rejected before opening the optimizer.
- Thin or incomplete data produces no number. Under 5 closed trades Tradelyze refuses: the DSR Value shows a dash, the badge reads Too Few Trades, and the check scores 0 of 18. The refusal counts as a failed check, so the whole score is multiplied by 0.69. Sometimes fewer than two distinct settings were measured, or the price data has no usable dates. Then no deflated Sharpe can be computed, the badge reads NOT RESOLVED in grey, and the check leaves the score as not run rather than failed; data with no usable dates still triggers the reduction, through the 30-day rule. Either way ROBUST is out of reach.
The formula, and how Tradelyze allows for lopsided trade results, are under Going deeper.
How much history does a backtest need to be meaningful?
Min Backtest Length on the Tradelyze robustness card compares two figures. Required Years is the history needed to trust your Sharpe ratio. Available Years is the history the checks ran on. Sufficient means you have enough; Insufficient means the Sharpe ratio may come from a wide search on too little history.
The box is marked Not scored, and its Sufficient or Insufficient badge is grey. The check earns no points and does not change the score, the grade or the verdict. It is still worth reading as a warning sign about the Sharpe ratio.
One row in the box does count: Enough for ROBUST (≥ 30 days), followed by the number of days of data. It reads Yes when your whole upload covers at least 30 days (one month), first bar to last, including any part held back for the held-out test. Available Years in the same box is the span the checks ran on, so it can be shorter than the upload. No means the points were multiplied by 0.69, so the score is 69 or below, the grade C+ at best and the verdict MARGINAL at best, however well the checks did. NOT MEASURED means the data had no usable dates, and counts the same as No. The row earns no points of its own. One month is a Tradelyze starting value with no primary source, and it may rise later.
The idea comes from Bailey, Borwein, López de Prado and Zhu in Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance (Notices of the American Mathematical Society 61(5), May 2014). The history needed to establish a given Sharpe ratio grows with the number of trials used to find it. Their headline claim is worth remembering: with enough trials, a backtest with an impressive Sharpe ratio can be built from random data.
Two rules of the arithmetic are worth knowing. Halving the Sharpe ratio your data supports quadruples Required Years. Going from 100 settings tried to 10,000 only doubles it, because the number of settings enters through a logarithm.
Required Years never falls under one year, so a backtest covering less than a year always reads Insufficient. Tradelyze added that floor because an annualized Sharpe ratio measured over a few days is an extrapolation. The same trades squeezed into a short span show a far higher yearly figure than they would over several years. Because the check is not scored, a run with between 30 days and one year of data reads Insufficient here and can still be certified ROBUST.
In practice, the check catches a common recipe for a fantasy backtest: a mediocre edge found by a wide search over a short history. It cannot detect whether your history contains more than one market regime. Ten years of a single trending market satisfies the arithmetic and tells you very little.
The formula is under Going deeper.
Why can a strategy pass every robustness test and still be overfit?
A strategy can pass every robustness test and still be overfit because all five checks run after the strategy was selected. They run on the winner's trade list, and none of them can see the candidates that lost.
Consider a realistic workflow. You test forty strategy ideas. Thirty-five look poor and you drop them. You optimize each of the remaining five, look at the robustness card, and keep the one that scored best. That strategy was selected twice: once by you across ideas, and once by the optimizer across settings. Its robustness score saw only the second selection, and only for the survivor.
The deflated Sharpe ratio is the one check aimed at selection bias, and its trial count covers the settings tried inside one optimization run. It has no way to know about the thirty-nine ideas you discarded. Nor can it know that you re-ran the pipeline eleven times with different ranges before this result.
Bailey, Borwein, López de Prado and Zhu built a tool for this in The Probability of Backtest Overfitting (SSRN 2326253). It estimates how likely the in-sample winner is to do worse than the median out-of-sample. In-sample means the data used to pick it; out-of-sample means data it never saw. That number describes the selection procedure, not the selected strategy. Tradelyze does not compute it. No robustness score computed on one finished backtest can, because the information needed is in the rejected strategies.
The one thing to record
Write down how many candidate strategies and how many pipeline runs stand behind the result you are reading. Keep that count next to the robustness score. A score of 82 from the first strategy you tested and a score of 82 from the fortieth are not the same evidence. Nothing computed from the backtest can tell them apart. Selection bias is bookkeeping, not statistics: only you have the numbers.
The same argument applies to the held-out test and the walk-forward analysis that sit next to the robustness card. Holding data out protects the first run. It does not protect the fortieth, because by then the held-out segments have helped choose the winner.
Going deeper
The sections below go deeper: the exact point rules behind the score, how many simulations each check runs, how the Monte Carlo draws are built, the corrected permutation bar, the sensitivity nudging rules, the deflated Sharpe ratio and minimum backtest length formulas, and robustness tests Tradelyze does not run. You can skip them and still read your own report.
How does Tradelyze turn each robustness check into points?
Each scored Tradelyze robustness check earns points on a sliding scale out of its weight. Tradelyze then renormalizes the total: it divides the points earned by the combined weight of the checks in the score and scales the result to 100. Finally, if any scored check failed or the data is too short, it multiplies the result by 0.69. The next table restates each rule.
| Check | Max points | Points awarded |
|---|---|---|
| Monte Carlo | 29 | 29 × (1 − Ruin Probability % / 50), never below 0. Zero if refused for fewer than 3 trades. |
| Permutation test | 29 | 29 if Significant. Otherwise 29 × log(P-Value) / log(bar), between 0 and 29, which measures how many orders of magnitude of the required evidence the P-Value reached. Zero if refused for fewer than 20 trades. |
| Parameter sensitivity | 24 | 24 × (1 − Degradation % / 50), never below 0, less up to half for a spiky neighborhood. Zero if the original Sharpe ratio is not positive. |
| Deflated Sharpe ratio | 18 | 18 × (DSR Value − 0.5) / 0.45, between 0 and 18, so 0.5 or less earns nothing and 0.95 or more earns all 18. Zero if refused for fewer than 5 trades. |
| Minimum backtest length | Not scored | Earns nothing. Shown for information only. |
| Points earned | 100 | Sum of the points above × 100 / total weight of the checks in the score |
| Robustness score | 100 | Points earned × 0.69 when any scored check ran and failed (a refusal for too few trades counts; a check that did not run does not), or when the data covers less than 30 days or has no measurable time span. Otherwise equal to points earned. |
Renormalization is why a NOT RUN check changes what the total means. With parameter sensitivity not run, a score of 90 is 68.4 points out of the 76 the other three checks can give. A NOT RESOLVED permutation test and a deflated Sharpe ratio that could not be computed are left out of the total weight too, and none of the three triggers the reduction. A check refused for too few trades stays in the total weight, scores zero and counts as failed.
The 30-day minimum earns no points either. Under 30 days, or with no measurable span, it triggers the same reduction as a failed check. The reduction is applied once, however many checks failed and whether or not the data is also short, so a reduced run shows 69.0 at most: a C+, and MARGINAL. With all four checks run, the lowest score a run can reach while passing all four is about 71.6, a B-: 17.4 points for a Ruin Probability just under 20%, 29 for a Significant permutation test, 7.2 for Degradation just under 20% with the full spike deduction, and 18 for a DSR Value over 0.95. That is why a factor of 0.69 puts every reduced run below every run that passed all four on at least 30 days of data.
Reports from before 25 September 2026 scored five checks, weighted 25, 25, 20, 15 and 15, with minimum backtest length among them. Reports from before 26 September 2026 had no failure reduction, and their cards show the minimum backtest length box without the Not scored label. Neither is directly comparable with current scores.
How many simulations each check runs
The Monte Carlo check draws 1,000 resamples by default. The permutation test runs at least 1,000 sign flips by default. It runs more, up to 10,000, when the bar it must clear needs finer resolution.
Parameter sensitivity runs no re-tests when the optimization ran fewer than 50 trials. Otherwise, by default, it runs one re-test for every five trials, at least 20 and at most 50. These counts are Tradelyze settings, not properties of your strategy. The Runs / Noise row on the card shows how many sensitivity re-tests a run used.
How does Tradelyze build the Monte Carlo draws?
The Tradelyze Monte Carlo check uses a stationary block bootstrap, not a plain reshuffle. Each draw starts at a random trade, walks forward through the trades that followed it, and now and then jumps to a new random start. Runs of consecutive trades stay together, so a losing streak stays a streak instead of being scattered through the sample.
Tradelyze picks the typical block length from how strongly your trades cluster, when there are enough trades to measure it. It uses blocks three times longer for the drawdown figures, because a deep drawdown is made of adjacent losses. A fixed random seed means the same trades give the same Monte Carlo numbers on every run.
max drawdown % = largest (peak − equity) / peak × 100 along the path
Ruin Probability = share of draws whose max drawdown broke the firm's total drawdown limit
MC Sharpe = Sharpe ratio recomputed on each draw ← a real spread, because each draw holds a different mix of trades
The Sharpe side varies because each draw holds a different mix of trades, not only a different order. That is why the MC Sharpe 5th–95th band has a real width.
The drawdown figures are floor estimates. Where the price data allows, each trade adds two steps to the equity path: the fall to its worst point, then the move to its close. A position that went deep underwater before closing green still shows a dip. A fall that opened and recovered inside a single bar stays invisible. See maximum drawdown for why the measurement convention matters.
If your script does not declare percent-of-equity position sizing, Tradelyze rebuilds the drawdowns against a fixed starting balance, the same way the backtest reports them, and a warning says the basis was assumed.
Why does the permutation test's bar get stricter as more settings are tried?
The permutation test's bar gets stricter because the settings under test are not one idea chosen in advance. They are the best of every trial the optimizer ran, and the best of many tries clears a fixed bar by luck.
In plain words: count the flipped versions that did at least as well as yours, add one to that count and one to the total number of versions, and divide. The bar starts at 0.05 and is tightened for the number of settings the optimizer tried, which is what the second line does.
bar = 1 − (1 − 0.05) ^ (1 / number of distinct settings tried)
Significant when P-Value < bar
The add-one comes from Phipson and Smyth (2010), who argue that your real result is itself one of the equally likely sign assignments, so it must be counted. As a result a permutation p-value can never be exactly zero. With 1,000 flipped versions the smallest possible value is 1 divided by 1,001, about 0.001, however good the strategy is.
Tradelyze tightens the bar with a Sidak correction, which by that formula gives 0.05 at one setting, about 0.0064 at eight and about 0.00085 at sixty. Re-testing the same combination does not raise the count, because Tradelyze counts distinct settings.
The P-Value itself is shown unchanged; only the bar moves. A P-Value that would read Significant after a small search can read Not Significant after a large one, with the underlying number exactly the same.
A tighter bar can fall under the smallest P-Value the flip count can express, and then no strategy could pass. Tradelyze avoids that by running more flips when the bar needs them, 1,170 at sixty settings by the same arithmetic, up to a ceiling of 10,000. When even 10,000 flips cannot reach the bar and no flipped version beat your result, the card shows NOT RESOLVED. Its warning gives the smallest P-Value the flips could express beside the bar it had to clear.
How does Tradelyze nudge settings in the sensitivity check?
In plain words: the Tradelyze sensitivity check nudges each numeric setting, re-runs the backtest, and compares the average re-run Sharpe ratio with the original one. Degradation reports the drop as a percentage.
Degradation % = (original Sharpe − average nudged Sharpe) / original Sharpe × 100, never below 0
Stable when Degradation < 20% and original Sharpe > 0
Tradelyze's version adjusts the basic idea in ways that change what the number means. Most adjustments stop the check from reporting stability it never tested; a few run the other way and are worth knowing for that reason:
- Nudged values are snapped back onto the search grid. A setting declared with a step of 0.5 can only take values on that ladder, so testing 2.31 would test a combination the optimizer could never have chosen.
- A nudge smaller than one rung is widened to one rung. Otherwise 5% of a value of 2.0 on a 0.5 ladder would round straight back to 2.0, and the strategy would collect full marks for a test that moved nothing.
- Settings the winning switches turn off, and settings you pinned to a fixed value, are held still. Moving a setting the strategy never reads cannot change a trade, and would make a strategy with many dead settings look steadier.
- On/off switches are not flipped. Flipping a switch changes which other settings the strategy reads at all, so it tests a different strategy rather than the stability of this one.
- Failed re-runs are excluded rather than counted as zeros. A connection failure scored as a zero Sharpe ratio would look like catastrophic fragility.
- Drop-down choices are never nudged, even when their options are coded as whole numbers. A list of options has no neighborhood, so Tradelyze holds the winning option. The effect flatters: a strategy tuned mostly through choice-type settings has less of itself tested, and the check reports no less confidence for it.
- Degradation is floored at zero. If the nudged re-runs do better on average than the winner, Degradation reads 0%. Read 0% as "did not get worse", and check Mean Δ (signed) for the direction.
- A spiky neighborhood costs points. An average cannot tell a broad plateau from a knife edge, so the score withholds up to half of this check's points when re-run results scatter widely or one re-run collapses far under the average.
- A strategy with no edge scores zero here. When the original Sharpe ratio is zero or negative there is nothing to lose, so a 0% Degradation is not a stability result, and the check earns none of its 24 points.
Two different Sharpe ratio constructions are in play across the five checks
Parameter sensitivity, the deflated Sharpe ratio and minimum backtest length read the backtest's own annualized Sharpe ratio (converted to a yearly figure), measured on the bars the trades really closed on. Monte Carlo and the permutation test cannot: both rearrange the trade list, which destroys the real exit times, so they spread the trades evenly across the run's bars. Both constructions are annualized and usually agree closely, but they can diverge when trades cluster on the same bars.
How does Tradelyze calculate the deflated Sharpe ratio?
In plain words: Tradelyze estimates the Sharpe ratio the best of your trials would reach by luck, subtracts it from your Sharpe ratio, divides by how uncertain your Sharpe ratio is, and turns the result into a probability. A standard deviation, used in the fourth line, measures how widely values spread around their average.
DSR = Φ( (SRobserved − E[max SR]) / SE )
γ = 0.5772156649 (Euler-Mascheroni) N = number of distinct settings tried
σSR = standard deviation of the Sharpe ratios across all trials, at least 0.01
SE = standard error of your Sharpe ratio, adjusted for skewed and fat-tailed trade results and for the length of the data
Φ = standard normal cumulative distribution Significant when DSR > 0.95
What Tradelyze's version includes
Tradelyze's deflated Sharpe ratio corrects for the number of settings tried, for skewness and kurtosis (lopsided and fat-tailed trade results) and for the length of the data, as the published statistic does. Two details are Tradelyze's own choices: the skewness and kurtosis are measured on trade results rather than bar-by-bar returns, and N counts distinct settings, so re-testing one combination does not count twice.
Earlier Tradelyze versions computed skewness and kurtosis without using them and had no data-length term, so DSR Values from those reports are not comparable with current ones.
The spread term does a lot of work. The luck benchmark scales with the standard deviation of the trial Sharpe ratios. A search whose trials all scored alike produces a low benchmark and a flattering DSR Value; a search that explored widely produces a demanding one. That is intended, but the value depends on the shape of your search space as well as on your strategy.
How does Tradelyze calculate the minimum backtest length?
In plain words: Tradelyze takes the low end of the Sharpe ratio your data can support, then asks how many years a Sharpe ratio that size needs before the best of N trials stops being explainable by luck.
Required Years = ( 1.645 × √ln(N) / supported Sharpe )2
N = number of distinct settings tried Required Years is at least 1 and at most 100
Sufficient when Available Years ≥ Required Years
The shape is intuitive once seen. The supported Sharpe ratio sits squared in the denominator, so halving it quadruples the requirement. The number of settings tried enters through a logarithm, so going from 100 settings to 10,000 doubles the requirement rather than multiplying it by a hundred, because the natural log of 10,000 is twice the natural log of 100.
The years figure is only as good as the Sharpe ratio under it
Required Years is computed from the backtest's own annualized Sharpe ratio, reduced to the low end of what the data supports. Earlier Tradelyze versions fed the check other Sharpe figures, one that pushed almost every run to the 100-year cap and one that let a handful of near-identical trades pass, so reports from before that change are not comparable. The output is still a rule of thumb rather than a measurement.
Which robustness tests does Tradelyze not run?
The Tradelyze robustness score is built from four scored checks, and other backtesting tools describe robustness tests that are not among them. Four of those tests are listed in the next table, each described the way its own vendor documents it. None of the four feeds the Tradelyze robustness score, so a ROBUST verdict says nothing about how a strategy would do on them. The table describes the tests; it does not rank any tool above another.
| Test | What it changes | What it detects | Described by |
|---|---|---|---|
| Noise testing | Adds or subtracts random amounts of volatility to the historical price bars, then re-trades the strategy on many altered price histories. Build Alpha suggests 1,000 or more as a start. | Noise testing detects a strategy fitted to the random wiggles of one price history, which Build Alpha reads from results that stop being profitable on the altered data. | Build Alpha robustness testing guide |
| Entry and exit delay testing | Enters or exits 1 or 2 bars later than the strategy's rules say. | Delay testing detects a strategy that only works at its exact entry and exit bars, which Build Alpha treats as a possible sign of overfitting, especially for intraday strategies. | Build Alpha robustness testing guide |
| Variance testing | Resamples the historical trades 1,000 times and keeps only the simulations whose chosen metric, such as win percentage, is worse than the backtest's by an amount the trader sets. | Variance testing detects a strategy that stops being worth trading if live results come in somewhat worse than the backtest, as in Build Alpha's example of a 61% win rate tested at 56% or lower. | Build Alpha robustness testing guide |
| Walk-forward matrix | Runs a full walk-forward optimization for every combination of several run counts and out-of-sample percentages. StrategyQuant's example uses 5 to 15 runs in steps of 2, each at 20%, 30% and 40% out of sample. | A walk-forward matrix detects a strategy that passes one walk-forward setup but fails across different re-optimization periods and training lengths. | StrategyQuant documentation: Walk-Forward Matrix |
Three of the four tests sit close to a Tradelyze check without being the same test. Tradelyze's parameter sensitivity check adds noise to the winning settings, not to the price bars, so it cannot show whether a strategy was fitted to one price history. Tradelyze's Monte Carlo check resamples your trades, as variance testing does, but it keeps every draw rather than only the draws that did worse than the backtest. The Tradelyze walk-forward card reports one walk-forward setup per run, by default 2 windows with 70% of each window used for tuning, rather than a grid of setups.
Build Alpha's guide lists more tests that are also outside the Tradelyze score, including liquidity testing and stress testing with synthetic data. A Tradelyze report is silent on all of them. Some strategies need one of these tests more than others, such as a delay test for a strategy that needs a fill on the exact bar its signal appears. That test has to be run in a tool that offers it, or by changing the strategy's rules and re-running the backtest.
Stage 3 · step 15 of 18. Next in the learning path: Prop firm rules
Where this appears in Tradelyze
In a Tradelyze report, this is the score, grade and verdict at the top of the robustness card. Tradelyze runs an uploaded TradingView Pine Script strategy as written, on a Pine Script backtesting engine and the price data you upload, and compares its trades with the trade list you exported from TradingView. It then runs parameter optimization, walk-forward analysis, a four-check robustness score and prop firm rule checks. It does not place trades, give financial advice or guarantee a challenge pass, and it is in beta.
To judge the whole report, not one tile, use the pre-trade checklist. If a check failed or did not run, see what to do when a robustness check shows NOT RUN or FRAGILE.
Create an account. Already a user? Open your strategies.
Frequently asked questions about robustness scores
What is a trading strategy robustness score?
A robustness score is one 0–100 number that adds up several separate stress tests of a backtest. Tradelyze's version scores four checks: Monte Carlo resampling, a permutation test, parameter sensitivity and a deflated Sharpe ratio. A minimum backtest length check is shown beside them for information only. The score is useful for sorting many strategies. Since 26 September 2026 a failed check keeps the score at 69 or below, but the total still does not say which check failed, so read each check before the total.
What is a good robustness score?
Tradelyze grades a score of 80 or more as B+ or better and a score under 40 as F. A grade of B- or better, a score of 70 or more, always means no scored check failed and the data covers at least 30 days. No published research sets a threshold for a composite like this, because the result depends on which tests it contains and how they are weighted. A better target than any number is four passed checks and a ROBUST verdict.
How is the robustness score calculated?
Tradelyze gives each check points on a sliding scale: up to 29 for Monte Carlo, 29 for the permutation test, 24 for parameter sensitivity and 18 for the deflated Sharpe ratio. Minimum backtest length is shown but earns no points. Checks that did not run are left out, and the points earned are rescaled so the total is still out of 100. A check refused for too few trades scores zero. If any scored check failed, a refusal included, or the data covers less than 30 days, the points are multiplied by 0.69, so the score is 69 or below.
Can a strategy fail every robustness check and still score well?
No, not since 26 September 2026. Every check still awards partial credit, so narrow misses collect most of their points: in a constructed example scored with Tradelyze's rules, a run that narrowly fails Monte Carlo, the permutation test, parameter sensitivity and the deflated Sharpe ratio earns 75.6 points. Because checks failed, those points are multiplied by 0.69, and the run scores 52.1, a D+ grade and a MARGINAL verdict. Without that rule the same run would score 75.6, a B grade and an ACCEPTABLE verdict.
What does the Monte Carlo check in a robustness score test?
The Monte Carlo check tests how deep your drawdown could have gone with a less lucky sequence of the same trades. Tradelyze redraws runs of your real trades with replacement, 1,000 times by default, so some trades repeat and some are left out. Ruin Probability is the share of draws that broke the prop firm's total drawdown limit, and it must be below 20% to pass.
Why does the Monte Carlo Sharpe not match my headline Sharpe?
The two Sharpe ratios are measured on different clocks. A resampled draw has no real exit times, so each Monte Carlo Sharpe figure spreads the drawn trades evenly across the run's bars, while the headline Sharpe uses the bars your trades really closed on. They can differ by tens of percent when trades cluster. Compare the MC Sharpe band with MC Sharpe Original, which is measured the same way.
What does a permutation test p-value mean for a trading strategy?
The permutation p-value is the share of random sign-flipped versions of your trade results whose Sharpe ratio matched or beat the real one. Each version keeps your entry times but picks long or short by coin flip, so a low value says your direction calls were unlikely to be luck. Tradelyze tightens the 0.05 bar for the number of settings the optimizer tried.
What does the permutation p-value not tell you?
The permutation p-value is not the probability that the strategy is profitable, and not the chance the result was luck in the everyday sense. The test checks one narrow assumption, random direction at your entry times, so the p-value says nothing about entry timing, trade size, execution costs, or the strategies you tried and discarded before this one.
How does the deflated Sharpe ratio affect the Tradelyze robustness score?
The deflated Sharpe ratio is worth 18 of the 100 points in the Tradelyze robustness score when all four scored checks run. A DSR Value of 0.5 or less earns none of them, 0.95 or more earns all 18, and values in between earn a share. The check needs at least 5 closed trades and scores zero below that. A DSR Value of 0.95 or less, or a refusal for too few trades, is a failed check: the whole score is then multiplied by 0.69, so it is 69 or below and ROBUST is out of reach.
What is the minimum backtest length?
Minimum backtest length is the span of history needed before a Sharpe ratio can be told apart from luck, given how many settings were tried to find it. The concept comes from Bailey, Borwein, López de Prado and Zhu in the Notices of the American Mathematical Society 61(5), May 2014. Tradelyze never asks for less than one year, and a weaker Sharpe ratio raises the requirement. Tradelyze shows the result for information only: it earns no points and does not change the grade or verdict.
What is parameter sensitivity testing?
Parameter sensitivity testing adds small random noise to the winning settings, snaps them back onto values the search allows, re-runs the backtest and measures how far the Sharpe ratio falls. Tradelyze reads Stable when that fall is below 20% on a strategy with a positive Sharpe ratio. Below 50 optimization trials the check does not run automatically, so the card shows NOT RUN.
Why is my robustness verdict ACCEPTABLE instead of ROBUST?
ROBUST needs all four scored checks to run and pass, at least 30 days (one month) of data and a score of at least 80. Since 26 September 2026 an ACCEPTABLE run has no failed check and at least 30 days of data, because either problem keeps the score at 69 or below, which is MARGINAL at best. So one of two things stopped it. Either a scored check did not run: it shows NOT RUN or NOT RESOLVED, or the deflated Sharpe ratio could not be computed. Or all four passed and the score is under 80. The line under the verdict tells you which, for example 3 of 4 checks passed · 1 not run. An Insufficient minimum backtest length is not a cause. Check parameter sensitivity first: when the optimization ran fewer than 50 trials, that check does not run automatically, and ROBUST is out of reach.
Can a strategy pass every robustness test and still be overfit?
Yes. Every test runs on the winning strategy's trade list, and none of them can see the candidates you rejected. If you picked the best of forty strategies, that choice is itself a search, and its bias is invisible to checks run afterwards. Bailey, Borwein, López de Prado and Zhu measure this in The Probability of Backtest Overfitting, which Tradelyze does not compute.
How do I test if my backtest is robust?
Start with the closed trade count: Tradelyze's permutation test needs at least 20 trades and its deflated Sharpe check at least 5. Then confirm that all four scored checks ran, which the line under the verdict counts, and read each badge separately. Look for a parameter plateau rather than a spike, confirm results on data the settings were not chosen on with the held-out test, and write down how many strategies you tried before this one.
Sources
- David H. Bailey and Marcos López de Prado, The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality, Journal of Portfolio Management 40(5), 2014, SSRN 2460551.
- David H. Bailey, Jonathan M. Borwein, Marcos López de Prado and Qiji Jim Zhu, Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance, Notices of the American Mathematical Society 61(5), May 2014: the minimum backtest length concept and the demonstration that enough trials produce an impressive backtest from random data.
- David H. Bailey, Jonathan M. Borwein, Marcos López de Prado and Qiji Jim Zhu, The Probability of Backtest Overfitting, SSRN 2326253: combinatorially symmetric cross-validation for estimating the probability that an in-sample winner underperforms out-of-sample.
- Belinda Phipson and Gordon K. Smyth, Permutation P-values Should Never Be Zero: Calculating Exact P-values When Permutations Are Randomly Drawn, Statistical Applications in Genetics and Molecular Biology 9(1), 2010: the add-one adjustment applied to the permutation p-value.
- Build Alpha, What is Walk Forward Optimization?, buildalpha.com/walk-forward-optimization, retrieved 15 September 2026: "a parameter that jumps from 12 to 47 to 23 to 35 to 55 to 13 is most likely curve fitting."
- Build Alpha, Robustness Testing Guide for Algo Trading Strategies, buildalpha.com/robustness-testing-guide, retrieved 15 September 2026: noise testing, delayed entry and exit testing, variance testing, and the further tests the guide lists, including liquidity testing and stress testing with synthetic data.
- StrategyQuant, Walk-Forward Matrix, StrategyQuant documentation, strategyquant.com/doc/strategyquant/walk-forward-matrix, retrieved 15 September 2026: the walk-forward matrix as a set of walk-forward optimizations over different run counts and out-of-sample percentages.
- Tradelyze implementation, reviewed 26 September 2026: the walk-forward defaults of 2 windows and a 70% training share, and its badge names; the four scored robustness checks and their 29, 29, 24 and 18 points; the 0.69 reduction for any failed scored check, a refusal for too few trades included, or less than 30 days of data; the minimum backtest length shown for information only; the 30-day minimum history, read over the whole upload; the scoring and renormalization rules; the grade and verdict bands; the robustness card's rounded-down headline score, check count, reduction note, Ruin Check badge, rounded-down Ruin Probability, Too Few Trades and NOT RESOLVED badges, history row and settings banner, and the completion email's robustness line; the Monte Carlo block bootstrap and 20% ruin gate; the permutation test's Sidak-corrected bar, flip counts and 20-trade minimum; the sensitivity rules, including no automatic re-tests under 50 optimization trials; the deflated Sharpe calculation and its 5-trade minimum; and the minimum backtest length formula with its one-year floor. The constructed scoring examples use Tradelyze's own scoring functions on hypothetical inputs.