Glossary
Monte Carlo simulation
Last reviewed: 26 September 2026·Tradelyze
Monte Carlo simulation replays a strategy's real trades many times in different orders and mixes to show how deep the worst losing stretch could have been. Reordering the same trades never changes total profit, but it can change the drawdown a lot. Tradelyze uses it to estimate how often a less lucky run of your trades crosses a drawdown threshold.
In plain English
A backtest is one trip through history, so its worst losing streak is simply the one that happened to show up. Monte Carlo simulation, also called Monte Carlo analysis, mixes up the same trades many times. It shows how much deeper that losing streak could reasonably have gone. That matters most when a prop firm closes the account at a fixed loss. It tells you how much room to leave for bad luck, not whether the strategy works.
New to this? A drawdown is a fall in account value from its last high. Start with maximum drawdown.
What does Monte Carlo simulation tell you about a trading strategy?
Monte Carlo simulation shows how much of a backtest's drawdown came from the order the trades happened to arrive in, not from the strategy itself. A backtest is a single run through history. Monte Carlo simulation builds many more runs out of the same trades. The run you got can then be compared with runs that were just as possible.
Those runs are usually summarized as percentiles. A percentile is the value that a given share of the runs came in under. A 95th-percentile drawdown of 20% means 95 runs in every 100 fell by 20% or less at their worst point.
A worked example shows why this matters. Take twelve trades that add up to a $1,200 gain on a $10,000 account. Every ordering of those twelve trades ends at $11,200, because adding the same numbers in a different order gives the same total. The maximum drawdown, the largest fall from a high point to a later low, changes a great deal from one ordering to the next.
In Figure 1, the order in which the constructed trades really arrived drew a 10.4% maximum drawdown. Across 200,000 random reorderings of the same twelve trades, the median drawdown was 15.4% and the 95th percentile was 23.8%. In all, 94.2% of the orderings were worse than the real one. An account sized to survive 10.4% would have been sized for a lucky draw.
Figure 1 reorders the trades without repeating any. Tradelyze's own check redraws runs of trades, repeating some and leaving others out, but the lesson about order is the same.
What Monte Carlo simulation cannot do is tell you the trades themselves were representative. Every Monte Carlo method for trading starts from material the backtest already produced, a trade list or a price series, and rearranges it. If that material came from one favorable stretch of market, resampling it gives a precise-looking answer to the wrong question. Monte Carlo simulation widens a sample into a range; it never moves the center of that range.
The one-line answer
Monte Carlo simulation answers "how bad could this have looked if the same trades had arrived in a different order or mix?" It does not answer "is this edge real?" One variant, the Monte Carlo permutation test, re-runs the strategy on rearranged prices and speaks to that question in part. It is rarely what a platform means by "Monte Carlo". The variants are compared in which of the five Monte Carlo tests does a report mean.
Why is my Monte Carlo drawdown so much worse than my backtest drawdown?
A Monte Carlo drawdown is usually worse than the backtest drawdown because maximum drawdown depends on the order of trades, while net profit does not. Net profit is a sum, and addition ignores order, so every reordering of a trade list ends at the same equity. Maximum drawdown is the deepest fall from a high point to a later low, and it moves a lot with order. Your backtest drew one order out of a huge number of possible orders, and nothing says that one was the worst.
Build Alpha publishes a case in which a strategy's backtest maximum drawdown of $1,663.90 became $5,195.17 under Monte Carlo. Two caveats come with that figure in the original, and they rarely come with the quote. First, it comes from a resample, not a reshuffle of the same trades. A resample, or bootstrap with replacement, repeats some trades and drops others; Build Alpha's own FAQ draws that distinction.
Second, the $5,195.17 figure is the worst of their resample runs, not a percentile of them. The worst of many runs keeps growing as runs are added, so it is not a sizing target. Build Alpha tells readers to size against the 90th or 95th percentile instead.
Figure 1 shows the same effect on twelve constructed trades that add up to $1,200 on a $10,000 account. The real order drew a 10.4% maximum drawdown. Across 200,000 reorderings of the same trades the 95th percentile was 23.8%, and every ordering still ended at $11,200.
When can a Monte Carlo drawdown read better than reality?
Monte Carlo simulation can also understate drawdown, and this direction gets far less attention. Four mechanisms produce it.
- The trade list is a ceiling. Reshuffling and resampling can only recombine losses that already occurred. If the future contains a loss larger than any in the backtest, no rearrangement of the backtest can show it. The worst thing a trade-order Monte Carlo can display is a reordering of things that already happened.
- Closed-trade sampling. Most implementations rebuild the equity curve from closed-trade results, so drawdown is measured only at trade exits. Dips inside open trades between exits are invisible. A drawdown that values open trades at market prices, called mark-to-market, is generally deeper.
- Streak destruction. Independent shuffling breaks the clustering of losses, which is the specific structure that produces the worst real drawdowns. This is covered in does trade shuffling understate risk.
- Fixed position size. A simulation may assume a constant size while the live system compounds, scales up after wins or holds several positions. In that case the simulated spread is narrower than the real one.
The part that is usually left out
A Monte Carlo drawdown that comes back lower than the backtest drawdown is not good news. It usually means the realized order was unusually unlucky. Or the simulation is measuring something milder than the backtest did, such as equity from closed trades only against equity that includes open trades. Check which equity series each number was computed from before treating the comparison as meaningful.
Does trade shuffling understate risk?
Trade shuffling often understates risk, and this is the strongest criticism of Monte Carlo simulation in trading. The mechanism is simple. Shuffling trades one at a time, as if trades did not affect each other, breaks up losing streaks. Those streaks are exactly what drives large real drawdowns.
Real trading losses are not independent draws. A trend-following strategy loses in chop and keeps losing until the chop ends. A mean-reversion strategy loses during a breakout and keeps losing until the breakout exhausts. Those runs are what turn a survivable decline into a terminal one. A one-at-a-time shuffle scatters those losses evenly through the sequence. The deepest simulated fall then comes out milder than the market can actually deliver. How long a losing streak to expect at a given win rate is covered in losing streaks.
Kevin Davey names the assumption, and states it more mildly than he is usually quoted. In a 2019 article on Monte Carlo probability cones (kjtradingsystems.com), he calls trade independence "generally true (or for all practical purposes true) for many trading systems". He treats serial correlation as something a trader can detect and correct for. Serial correlation means one trade's result tends to lean the same way as the one before it.
The paraphrase that serial correlation invalidates the results is stronger than what he wrote. The usual fix is a block bootstrap: it redraws blocks of consecutive trades instead of single trades, so runs of that length stay intact. Practitioners commonly use blocks of roughly 5 to 20 trades, a convention with no primary source behind it.
There is academic support for the direction of the effect, though not for trade lists specifically. Paskaramoorthy, van Zyl and Gebbie studied it in The bias of IID resampled backtests for rolling-window mean-variance portfolios (arXiv 2505.06383, 9 May 2025). The paper looks at resampling that treats returns as if they did not affect each other. They find the bias comes from breaking the link between the stretch a portfolio is tuned on and the stretch it is tested on. That link exists because returns follow on from earlier returns. The bias depends strongly on how closely each return tracks the one just before it.
Their study covers portfolio backtests, not trade-list reshuffling. They conclude that the bias "can often be tolerable" while highlighting "the need for structure-preserving resampling methods". Treat the paper as evidence about what drives the bias, not as a measurement of how large it is on your trade list.
How large the effect is depends on two things at once: how clustered the losses are, and how long a block the resampling keeps. The two interact, which is why a range of block lengths is used rather than a blanket "keep longer blocks". Figure 2 measures both on a constructed series whose clustering length is known in advance.
Figure 2 is constructed, not measured. It uses 200 synthetic trades on a $50,000 account, built so trades arrive in runs of ten, with 20,000 resamples at each block size. It shows the honest scope of the effect. A longer block raises the simulated tail only while the block is no longer than the losing runs. Past that point the tail flattens and can fall, so "longer is safer" is not the rule.
The pattern held when the construction was repeated. Across twenty independently generated series built the same way, the 95th percentile rose steeply through block size 10 every time. At block size 20 it rose further in fifteen series and fell back in five. On series with no clustering, where trades do not affect each other, the slope reverses. Across twenty such series, the 95th percentile at block size 20 was a median 0.83× the block-size-1 figure. It is lower rather than flat because at 200 trades a block size of 20 leaves only ten blocks to draw.
The sharpest number from the Figure 2 series is one the chart does not show. In the order the trades actually arrived, the series drew a 19.6% maximum drawdown. That is more than twice the 8.8% that the one-trade-at-a-time resample puts at its 95th percentile. On this series, the resample that scatters the losing runs understates a drawdown that had already happened.
One caveat about the leftmost pair of bars. A block size of 1 samples single trades with replacement, which is a plain bootstrap. That is not quite the one-at-a-time shuffle, without replacement, that this section is about.
Both break up streaks completely, so the mechanism the figure shows is unaffected. On the same series, a true shuffle reads milder still: a 95th-percentile drawdown of 8.0% against the bootstrap's 8.8%. That makes the gap against the realized 19.6% wider, not narrower.
Tradelyze's Monte Carlo check is built around this problem. It redraws runs of consecutive trades rather than single trades. When there are enough trades to measure it, the typical run length is set from how strongly your own trades cluster. For the drawdown figures the runs are kept longer still, so losing streaks are not scattered away before the drawdown is measured.
How do I read Tradelyze's Monte Carlo numbers?
Tradelyze's Monte Carlo check builds 1,000 new versions of your trade list by default. Each version is stitched together from short runs of your real trades, picked at random, so losing streaks stay together. Some trades are repeated and some are left out. Read two numbers first. Ruin Probability is the share of versions whose drawdown went past the limit, and it must stay under 20% for the check to pass. The Ruin Check badge under it reads Ruin Acceptable when the check passed, Ruin Too High when Ruin Probability was 20% or more, and Too Few Trades when there were too few trades to run the check. The 95th-percentile drawdown, the second number in MC Max DD Real→P95, is the depth that 95 of every 100 versions stayed within.
The limit behind Ruin Probability is the maximum total drawdown of the prop firm or custom rule set you selected. Each selected firm gets its own figure. A run that names no firm is checked against Tradelyze's default rule set, whose total drawdown limit is 10%. The check uses the chosen settings' trades over the history the optimizer searched. Its numbers are therefore in-sample: measured on the same data the settings were tuned on. Those settings are the search's overall best. When a firm was recommended other settings, an amber banner on that firm's robustness card says so.
The 20% pass line is Tradelyze's own choice, not a published standard. The Monte Carlo check carries 29 of the 100 points in the robustness score. A failed check costs more than those points. Since 26 September 2026, any failed scored check multiplies the robustness score by 0.69, so that firm's card shows a C+ grade and a MARGINAL verdict at best, and an amber note on the card gives the reason. Because ruin is counted at each firm's own limit, the check can fail, and the score be reduced, for one firm and not for another. If the check fails, or the robustness verdict reads MARGINAL or FRAGILE, see what to do when a robustness check shows NOT RUN or FRAGILE.
How do I read the MC Sharpe band?
The MC Sharpe band shows how much the Sharpe ratio moves when Tradelyze redraws your trades, with some repeated and some left out. A narrow band means the result does not hinge on a few trades; a band that sinks toward zero means it does. Compare the band only with MC Sharpe Original; the Sharpe on the metrics card is measured differently.
The Sharpe ratio is average return divided by the standard deviation of returns, a measure of how widely they swing. MC Sharpe Mean is the average across the draws. MC Sharpe 5th–95th gives the 5th and 95th percentiles. So 1 draw in 20 fell below the first number, and 1 in 20 rose above the second.
The gap with the Best Metrics card is about timing. A redrawn trade has no real exit date, so every Monte Carlo Sharpe, MC Sharpe Original included, spreads trades evenly across the bars. Sharpe (Bar) and Sharpe (Daily) use real trade timing.
What do Ruin Probability and MC Max DD Real→P95 mean?
Ruin Probability is Tradelyze's Monte Carlo estimate of risk of ruin: the share of simulated runs whose drawdown broke the limit. It is counted against the total drawdown limit of the prop firm or custom rule set you selected. A run that names no firm uses Tradelyze's default 10% limit. Ruin Probability must be below 20% for the Monte Carlo check to pass, and the Ruin Check badge under it shows the result. A low figure means few equally possible orderings would have closed the account.
In plain words, ruin probability counts the simulated runs that broke the limit:
Ruin Probability is not the chance of losing everything. A redrawn run has no calendar, so daily loss limits are checked on the real trade order in each firm's Rule Results table. Each selected firm gets its own figure at its own limit, so do not compare it between firms. If the chosen rule set has no total drawdown limit, the Ruin Probability figure is left blank rather than borrowed from another limit.
MC Max DD Real→P95 shows two numbers. The first is the maximum drawdown of your trades in their real order. The second is the 95th-percentile drawdown: the level that 95 of every 100 simulated runs stayed under.
Size against the right-hand number; a wide gap means the real order was kind. Both figures are rebuilt from your trades, so they differ from the Max Drawdown tile and can miss a dip inside an open trade. Treat them as a floor, not a worst case.
Can Monte Carlo tell me my odds of passing a prop firm challenge?
Monte Carlo simulation can partly answer that, and the mechanism it captures is the right one. Prop-firm evaluations fail on sequence risk, the risk that losses arrive in a damaging order. A strategy with positive expectancy, meaning it makes money per trade on average, can still breach a daily loss limit or a trailing drawdown that way. The order of trades is precisely what a trade-order reshuffle varies. That makes Monte Carlo simulation a better match for prop-firm rules than for almost anything else.
The lever that moves the answer is risk per trade, and the mechanism is arithmetic rather than something that needs a study behind it. Cutting the risk on each trade shrinks how fast a losing run eats the buffer. A longer run of losses is then needed to breach the limit, and the winning trades get more chances to arrive first.
The strategy has not improved. Only the share of orderings that end the account has fallen. That share can move a long way for a change in risk that looks small.
Three limits are worth stating before trusting a pass probability.
- Daily limits need intraday equity. A closed-trade reshuffle knows the size of each trade but not which calendar day it landed on. A daily loss limit caps a whole day's losses. Simulating it properly means grouping trades into days and reshuffling whole days, not single trades.
- Trailing drawdown is path dependent twice over. A trailing threshold moves up with new equity highs, so its breach depends on the order of wins as well as losses. The reshuffle captures this, but only on closed-trade equity. Some firms measure trailing drawdown on equity that includes open trades, which a closed-trade simulation cannot see at all.
- Streak destruction cuts the wrong way here. Shuffling trades one at a time, as if trades did not affect each other, breaks up loss runs. Loss runs are exactly what breaches a daily limit. A pass probability from a one-at-a-time shuffle is therefore optimistic in the specific way that matters most for this use.
In Tradelyze, Ruin Probability is counted separately for each selected firm, against that firm's maximum total drawdown limit. It must be below 20% for the Monte Carlo check to pass. Daily loss limits are not simulated, because a redrawn run has no calendar; Tradelyze checks them on the real trade order in each firm's Rule Results table. A strategy can clear Ruin Probability and still fail on one bad day, so read the two together.
Used with those caveats, a Monte Carlo pass rate is useful for comparison. It can rank two risk-per-trade settings, or two strategies, against the same rule set. Treated as a literal probability of passing, it is over-confident. The rule mechanics themselves are covered in prop firm rules and backtest metrics.
Should I size positions off the backtest drawdown or the Monte Carlo drawdown?
Size positions off the Monte Carlo drawdown. The backtest drawdown is one sample from a range of possible drawdowns. Sizing against it assumes the future arrives in the same order as the past.
Build Alpha publishes a worked case on $SMH showing how account size and drawdown percentage interact under Monte Carlo. The same strategy shows a 30% drawdown at the 95th percentile on a $2,500 account and 18% on a $4,000 account.
The two percentages make sense once you notice that the test holds position size at a flat 100 shares on both accounts. The dollar drawdown is about the same in both; only the account size it is divided by changes. That is the point of the example. The dollar drawdown belongs to the strategy. The percentage drawdown, which decides whether you keep trading, belongs to the account you attach it to.
Practitioners often go further than the model. A 2002 EliteTrader thread records the heuristic plainly: whatever drawdown the model predicts, double it before sizing. That is not a statistical procedure and there is no primary source establishing the factor of two. It is a safety-margin convention. It exists because several problems on this page push the estimate the same way, too low. They include broken-up streaks, closed-trade sampling and the trade list acting as a ceiling.
Sizing sequence
Take the 95th-percentile maximum drawdown from the simulation, not the backtest figure; in Tradelyze that is the right-hand number in MC Max DD Real→P95. Convert it to dollars at your intended position size. Compare it against the largest decline you can absorb without stopping. If the two are close, reduce size — the estimate is more likely to be too small than too large.
What confidence level should I use for Monte Carlo drawdown: 80%, 95% or 99%?
A StrategyQuant administrator, answering on the vendor's own support forum, says it "does not matter too much whether it is 80 or 95%". That answer is weak even coming from the vendor, and it is worth being precise about why.
The confidence level is not a statistical parameter to be chosen on statistical grounds. It is a position-sizing decision. Reading the 95th-percentile drawdown instead of the median means holding enough capital that 95% of plausible orderings would not have breached your limit. It also means accepting that 1 run in 20 would have. Whether that trade-off is right depends on what a breach costs, and the cost is wildly different across situations.
| Level | Implied tolerance | When it is defensible |
|---|---|---|
| 80% | 1 sequence in 5 breaches the sized-for drawdown | Capital you can add to, no hard external limit, many uncorrelated strategies running |
| 95% | 1 sequence in 20 breaches it | The common default. Reasonable for a single discretionary account with reserve capital |
| 99% | 1 sequence in 100 breaches it | A hard limit exists and breaching it is terminal — a prop firm account, or capital you cannot replace |
The asymmetry is the point. If a breach means adding funds, an 80% level costs you occasional inconvenience. If a breach means a closed prop-firm account and a lost fee, the 99th percentile is the only defensible read. Even that assumes the simulation captured the right risk. A reshuffle of closed trades may not, because it breaks up losing streaks and cannot see dips inside open trades.
A 99% level also needs far more simulations to estimate steadily, so the level and the simulation count have to be chosen together.
Tradelyze reports the 95th percentile, as the right-hand number in MC Max DD Real→P95. Where a breach ends the account, treat that figure as a floor rather than a comfortable estimate.
How many Monte Carlo simulations do I need?
The published answers on how many Monte Carlo simulations to run conflict, and that conflict is the honest finding rather than something to average away.
| Count | Context | Source |
|---|---|---|
| 200+ | General Monte Carlo on a trade list | StrategyQuant blog post, "Monte Carlo explained in 6 minutes" — "the magic number of simulations is 200+". Vendor blog content, not documentation and not research. |
| 500 | Product default | Adaptrade Software, Market System Analyzer default setting, documented at adaptrade.com |
| 1,000+ | General recommendation | Build Alpha published guidance. Its very next clause is "however, results are generally acceptable with 100 or more", which removes most of the apparent conflict in this table |
| Hundreds or thousands | Permutation test p-value | Timothy Masters, 2018 and 2020 books; see Sources on the often-quoted "at least 1,000" |
What actually determines the answer is not the strategy but the percentile you intend to read. A median is estimated from the whole sample and stabilizes quickly. A 95th percentile is estimated from the top 5% of runs, so 200 simulations put roughly 10 observations in the region that determines the answer.
A 99th percentile at 200 simulations rests on about 2 observations, which is not an estimate at all. The rule that follows is mechanical: the further into the tail you read, the more runs you need. A count that is enough for a median is nowhere near enough for a tail.
There is a ceiling that no simulation count fixes. Resampling cannot add information the trade list does not contain. Ten million reshuffles of 40 trades produce a very smooth distribution over 40 trades. The smoothness comes from the arithmetic, not from evidence about the strategy.
The practical test is to raise the count until the percentile you care about stops moving between runs with different random seeds. A random seed is the starting number that decides which random draws a simulation makes. If it is still moving at 1,000, it is not converged, and the answer is more runs or a less extreme percentile.
Tradelyze runs 1,000 draws by default and reads the 5th and 95th percentiles, where about 50 of the 1,000 draws sit beyond each cut-off.
How many trades do you need before Monte Carlo means anything?
An EliteTrader thread on Monte Carlo analysis gives 100 trades as a minimum and 500 as preferred. Those are practitioner conventions rather than derived thresholds, and no primary source establishes either figure. For trade counts in any backtest, not only Monte Carlo, see how many trades a backtest needs.
The sharper objection is about the mix of trades rather than the count. A market regime is a stretch in which the market behaves one way, such as a steady trend or a choppy range. Forty trades from a single regime, reshuffled ten thousand times, produce a confident-looking cloud built on a biased sample. More simulations make the output look more precise. The information in it is still fixed by how many trades there are and how varied they are. Ten thousand simulations on 40 trades produce a smooth percentile curve that still describes only 40 trades.
Check this on your own results
Count the distinct market conditions your trades came from, not just the trades. Forty trades spanning one trending year and forty trades spanning a trend, a chop and a crash are the same number and completely different evidence. If every trade in the list came from one regime, the Monte Carlo output describes that regime and says nothing about the next one.
A defensible floor combines both. Before acting on any percentile from a Monte Carlo simulation, have at least 100 trades drawn from more than one market regime. With fewer trades than that floor, the honest output is the trade count itself.
Tradelyze leaves its Monte Carlo figures blank, and scores the check at zero, when the backtest closed fewer than three trades. That is the least Tradelyze needs to produce any number at all, not a sign that three trades are enough. The refusal counts as a failed check, because too few trades is a fact about the strategy: the Ruin Check badge reads Too Few Trades, not Ruin Too High, and the robustness score is multiplied by 0.69.
Should Monte Carlo run on in-sample or out-of-sample trades?
Run Monte Carlo on out-of-sample trades, from data the optimizer never tuned on, where enough of them exist. This question gets asked explicitly and is answered almost nowhere, including by the tools that make the choice for you silently.
Running a reshuffle on in-sample trades, the ones the optimizer was fitted to, inherits the selection luck that got those parameters chosen. The winning parameter set won partly because its particular sequence of trades looked good on that particular stretch of history. Reshuffling that list builds a range of drawdowns from an already-flattering sample, so the whole range sits in the wrong place. It will still look reassuringly tight, because resampling narrows uncertainty about the sample without touching the bias in it.
Out-of-sample trades — the ones from walk-forward test segments, or from a held-out period the optimizer never touched — do not carry that selection bias. The problem is that there are usually far fewer of them. That runs straight into the trade-count floor set out in how many trades you need before Monte Carlo means anything. The two constraints genuinely conflict, and there is no arrangement that satisfies both on a short history.
The workable position is to run the simulation on out-of-sample trades when you have enough of them. When you do not, use the full trade list. In that case, say the result is optimistic rather than presenting it as a risk estimate. Whichever a platform does, it should say which, and most do not.
Tradelyze runs its Monte Carlo check on the chosen settings' trades across the same price history the optimizer searched. Those Monte Carlo numbers are therefore in-sample and lean optimistic. Tradelyze's out-of-sample results are in two places. One is the walk-forward card, but only when its label reads Re-tuned each window. A card labeled One run, split by period scores settings on the same history they were chosen on. The other is the held-out test on each firm's card: by default the last 25% of the history is withheld from the whole search, and the settings recommended for that firm are then tested on it. See in-sample vs out-of-sample testing.
Does Monte Carlo simulation detect overfitting?
Monte Carlo simulation mostly does not detect overfitting, and the reason is structural. Trade reshuffles, bootstraps and block bootstraps all begin from the realized trade list. An overfit strategy produces a trade list just as readily as a sound one, and resampling cannot tell them apart. It never sees the parameter search that produced the trade list, or a bar the strategy was not fitted on.
Only the Monte Carlo permutation test addresses overfitting, and only partly. The Monte Carlo permutation test (MCPT) re-runs the full strategy on synthetic price paths whose bar-to-bar order has been shuffled away. That lets it catch a strategy whose apparent edge came from fitting noise. If the strategy scores as well on the shuffled markets as on the real one, the p-value says so. A p-value is the probability of a result at least this good if the strategy had no edge.
The Monte Carlo permutation test leaves the market's drift in place: its overall rise or fall from the first bar to the last. So a strategy that mostly buys still earns something on every rearranged price path. How far it reaches into the parameter search depends on which of Timothy Masters' three permutation-test variants was run. Only the in-sample version re-runs the optimizer on each permuted series. None of them corrects for how many strategies you tested before keeping this one.
That is the multiple-testing problem. The top-ranked comment on an r/algotrading thread about robustness states it exactly. Every robustness test can be passed by chance if you kept the best survivor of many strategies. Run 500 variants, keep the one with the best Monte Carlo profile, and you have selected on the test rather than passed it.
The same comment points at two standard corrections. White's Reality Check judges the winner against the whole set of candidates, not the winner alone. The Bonferroni correction tightens the pass threshold by the number of tests. Neither is a Monte Carlo procedure and neither is offered by the tools that run Monte Carlo.
The literature that quantifies this is Bailey and López de Prado's, not Monte Carlo literature. The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality (SSRN 2460551) adjusts a Sharpe ratio downward for the number of trials that produced it. The Probability of Backtest Overfitting (SSRN 2326253) estimates the probability that the selected configuration is the product of the search rather than of an edge. Neither is a Monte Carlo test, and neither is substitutable for one.
Tradelyze's robustness card carries two checks aimed at this, separate from its Monte Carlo check. One is a permutation test that flips the signs of trade results at random; the other is a deflated Sharpe ratio. Both are explained on the robustness score page.
The part that is usually left out
Record how many parameter combinations were tried before the winner was chosen, and report that count next to every Monte Carlo result. A 95th-percentile drawdown from a search over 10,000 combinations is weaker evidence than the same figure from a search over 20. No amount of resampling recovers the difference. See overfitting and sample size.
Why do I get different Monte Carlo results every run?
Different Monte Carlo results on every run have two causes, and they compound: percentile instability and random seeds that are not fixed.
Tail percentiles are estimated from few effective observations. At 1,000 simulations, the 95th percentile is determined by roughly the top 50 results and the 99th by roughly the top 10. Draw a new set of 1,000 simulations with a different seed and those few extreme values change. The reported percentile changes with them, even though nothing about the strategy moved.
A post on r/quant asks about the same instability from the other end: two runs of 1,000 simulations returning −64% and −68%. Two qualifications matter before reading that as a measurement of percentile drift. Those figures are the minimum and maximum of each run, not percentiles of it. The author is also resampling with replacement rather than reshuffling. What they illustrate is how far an extreme statistic moves between seeds, which is the same mechanism amplified.
The seed is the other half. A random seed is the starting number that decides which random draws a simulation makes. A simulation with a fixed seed reproduces its output exactly, which makes results comparable between runs and between strategies but hides the instability entirely. A simulation with a random seed exposes the instability but makes two runs incomparable. Neither is wrong; what is wrong is not knowing which one your tool does.
Tradelyze fixes its seed. Tradelyze's block bootstrap and its permutation test use a fixed random seed, so the same trades give the same Monte Carlo numbers on every run. Tradelyze's permutation test flips the signs of trade results at random. It is not the price-permutation test described in what the Monte Carlo permutation test actually permutes.
That makes two Tradelyze results directly comparable with each other, and makes the reported percentile look steadier than the underlying estimate is. To see the real spread you have to resample the trade list yourself under several seeds; Tradelyze will not show it to you.
Three responses, in order of usefulness. Raise the simulation count until the percentile you care about stops moving across seeds. Read a less extreme percentile, which is estimated from more of the sample. Report a range rather than a point. A made-up example: "95th-percentile drawdown between 24% and 27% across ten seeds" is more honest than any single value from one run.
Do I need both Monte Carlo simulation and walk-forward analysis?
You need both Monte Carlo simulation and walk-forward analysis, because they answer different questions and neither substitutes for the other.
| Method | Question it answers | What it varies | What it cannot see |
|---|---|---|---|
| Walk-forward analysis | Do parameters fitted on one stretch of history still work on bars the optimizer never saw? | Which bars are used for fitting and which for scoring | Path risk within a segment. It tests one historical path — the one that happened. |
| Monte Carlo simulation | How much of the result depended on the order or mix of trades that happened to occur? | The order, and in bootstrap variants the selection, of trades | Whether the parameters generalize. It resamples the trade list a fitted strategy produced. |
A strategy can pass either one while failing the other. Parameters that transfer cleanly out-of-sample can still produce a drawdown path that ends the account before the edge pays. That is a walk-forward pass and a Monte Carlo failure. A strategy with a tight, mild range of drawdowns can be entirely fitted to its training window. That is a Monte Carlo pass and a walk-forward failure.
They also share a blind spot. Neither corrects for how many candidate strategies were tested before this one was selected, a problem set out in does Monte Carlo simulation detect overfitting. Walk-forward mechanics are covered in walk-forward analysis, and the ratio it reports in walk-forward efficiency.
Going deeper
The sections below go deeper: the five different tests that share the Monte Carlo name, and what the Monte Carlo permutation test actually permutes. You can skip them and still read your own report.
Which of the five Monte Carlo tests does a report mean?
A "Monte Carlo" report means one of several genuinely different procedures, and the report usually does not say which. This section is the technical reference for the rest of the page, and it sets out five of the most common variants. They share a name and share nothing else: they randomize different objects, they answer different questions, and each one destroys some structure in order to do its job. Reading a Monte Carlo report without knowing which one produced it is not possible.
Five is not the whole list, and no source establishes it as one. StrategyQuant's own documentation additionally offers Randomize Starting Bar, Randomize Strategy Parameters and Randomize History Data, none of which map onto any of the five in the table. Treat the table as the variants you are most likely to meet, not as a taxonomy.
| Technique | What is randomized | What it can answer | What it destroys or cannot answer |
|---|---|---|---|
| (a) Trade-order reshuffle permutation of the trade list |
The order of the realized trades. The same trades appear exactly once each, in a new sequence. | Drawdown-path risk: how deep the peak-to-trough decline could have been if the same trades had arrived in a different order. | Net profit is mathematically invariant, so it says nothing about profitability. StrategyQuant's own documentation states that randomizing trade order "doesn't change the resulting Net Profit". It also breaks up win and loss streaks. |
| (b) Bootstrap resample with replacement |
Which trades appear and how many times each one appears. Some trades are drawn twice, others not at all. | Confidence intervals on profit, maximum drawdown, profit factor and win rate, because the trade mix itself now varies. | Assumes trades do not affect each other and that the market behaves the same way throughout the sample. Both assumptions fail across a regime change. |
| (c) Execution randomization two modes, only one of which needs the price series |
The fills. Adding slippage and skipping a fraction of trades are arithmetic on the finished trade list. Moving an entry or an exit by a bar or a tick is not — that one needs the price series and a re-run. | Fragility to execution reality — whether the result depends on a handful of perfect fills that live trading will not reproduce. | Says nothing about whether the entry signal has predictive content. A strategy with no edge and cheap execution passes this test comfortably. |
| (d) Monte Carlo permutation test MCPT |
The price series, in log space. For OHLC bars Timothy Masters splits each bar into a close-to-open gap and three open-to-high, open-to-low and open-to-close offsets, then permutes the gaps and the offsets under two independent permutations, moving high, low and close together at the same index so each bar's internal geometry travels intact. A synthetic series is rebuilt and the strategy is re-run on it. | A p-value for "could a backtest this good have come from pairing this system with this market by chance?" This is the only variant that tests the signal rather than the outcome. | Destroys the bar-to-bar ordering and the autocorrelation — that is the null hypothesis, not a bug. It does not destroy the drift: the permutation pins both endpoints, so every synthetic series ends where the real one did. Costs one full backtest per permutation, so it is far slower than the others. Timothy Masters develops the method in Testing and Tuning Market Trading Systems (Apress, 2018) and Permutation and Randomization Tests for Trading System Development (2020). |
| (e) Block bootstrap fixed-length blocks |
Which contiguous blocks of trades or returns appear, and in what order. Sequence inside a block is preserved. | The same quantities as (b), but with win streaks, loss streaks and volatility clustering surviving the resampling. | Requires a block length, and the answer moves a long way with that choice — see does trade shuffling understate risk. Politis and Romano's stationary bootstrap (Journal of the American Statistical Association 89(428), 1994, pages 1303–1313) is the variant that removes the choice: its defining feature is a random, geometrically distributed block length, picked precisely so the resampled series stays stationary. A single fixed block length is the older moving-block bootstrap, not theirs. Tradelyze's Monte Carlo check uses the stationary variant. |
The deepest division in that table is not between the rows but across them, and the axis is what each technique needs as input. Techniques (a), (b) and (e) need nothing but the finished trade list, and so does half of (c) — adding slippage and dropping a fraction of trades are arithmetic on a list of numbers.
The other half of (c), and the whole of (d), need the price series: moving a fill by a bar and rebuilding a synthetic market both require bars, and both require the backtest to run again.
Only (d) re-runs the strategy on prices it has never seen, which is why it is the only one of the five that can speak to whether the entry signal predicts anything. That does not make the trade-list techniques silent about the strategy. A bootstrap confidence interval on mean trade profit that straddles zero is genuine negative evidence, and (b) is asymmetric in exactly that way: it can falsify an edge without ever being able to confirm one.
What the Monte Carlo permutation test actually permutes
The Monte Carlo permutation test is the variant most often described loosely, and three of the loose descriptions are wrong in ways that change what the p-value means. Timothy Masters' own C++ implementations, published as companion source to Testing and Tuning Market Trading Systems (Apress, 2018) and Permutation and Randomization Tests for Trading System Development (2020), settle all three.
It is not one shuffled stream of bar returns. That description fits only the single-price-series case, where a close-only series is permuted. For OHLC bars Masters decomposes each bar in log space into a close-to-open gap and three offsets measured from that bar's own open to its high, its low and its close.
He then applies two independent permutations: one to the overnight gaps, one to the intra-bar offsets, with high, low and close carried to the same destination index together. Only the gap moves on its own. Each bar's internal geometry — where the close sits inside the range, how far the high reaches above the open — survives as an intact unit, so the synthetic series still looks like the instrument it came from.
The permuted market is not structureless. Masters' permutation routine leaves the first bar and the final close alone, and its own comment says why: "the shuffled array starts and ends at their original values. Only the interior elements change." Both endpoints are pinned, so every synthetic series carries the original net drift, and a long-biased system goes on making money in the permuted market.
That is exactly why the same program separates a system's return into a trend component, a training bias and what is left over as skill, and prints all three. The null the p-value tests is correspondingly narrower than "the market had no structure": Masters labels his own output "p-value for null hypothesis that system is worthless", which is a claim about the pairing of this system with this market, not about the market on its own.
There is no single MCPT. The 2020 book publishes three, with three different nulls and three different programs. The in-sample test re-runs the whole optimizer on each permuted series, so it prices in the parameter search itself. The single out-of-sample test trains once on real data, freezes the parameters, and permutes only the out-of-sample region.
The walk-forward test permutes both regions and re-runs the entire walk-forward, re-optimizing every fold on every permutation. Collapsing them into one "MCPT" loses the only thing that distinguishes their answers, and the in-sample version is the expensive one for a reason.
Two smaller points are worth carrying across. Masters computes the p-value as the count of replications scoring at or above the original divided by the replication count, with the counter initialized to one for the unpermuted run — the conservative (r + 1) / (m + 1) form, the same add-one correction Phipson and Smyth argue for, and the reason a permutation p-value can never be exactly zero.
And the frequently repeated "Masters recommends at least 1,000 permutations" traces to a code comment in a 2006 paper of his, about a different scheme that shuffled position vectors. His 2018 program asks only for a replication count of "hundreds or thousands".
Stage 3 · step 14 of 18. Next in the learning path: Robustness score
Where this appears in Tradelyze
In a Tradelyze report, these results are the Monte Carlo rows of the robustness card: MC Sharpe band, Ruin Probability with its Ruin Check badge, and MC Max DD Real→P95. Tradelyze re-runs an uploaded TradingView Pine Script strategy on your price data and checks it against your exported trade list, then runs parameter optimization, walk-forward analysis, a robustness score from four scored checks, and prop-firm rule checks. It does not place trades, give financial advice or guarantee a challenge pass, and it is in beta.
To judge the whole report, not one tile, use the pre-trade checklist. If the Monte Carlo check fails or the robustness verdict reads MARGINAL or FRAGILE, see what to do when a robustness check shows NOT RUN or FRAGILE.
Create an account. Already a user? Open your strategies.
Frequently asked questions about Monte Carlo simulation
What is Monte Carlo simulation in trading?
Monte Carlo simulation in trading replays a strategy's trades many times, in different orders or mixes, to show the range of drawdowns the same trades could have produced. A backtest shows only one sequence. Several different tests share the name, including trade reshuffles, bootstraps, block bootstraps, execution randomization and the Monte Carlo permutation test, and each one randomizes something different.
What does a Monte Carlo simulation actually tell you about a trading strategy?
What a Monte Carlo simulation tells you depends on which variant was run. A trade-order reshuffle speaks to drawdown risk and says nothing about profitability, because reordering trades cannot change net profit. A bootstrap gives confidence intervals. A permutation test gives a p-value: the probability of a result at least this good if the strategy had no edge. A report that says only Monte Carlo has not told you what was tested.
Why is my Monte Carlo drawdown so much worse than my backtest drawdown?
Maximum drawdown is path dependent. The same set of trades in a different order produces a different peak-to-trough decline, and your backtest sampled exactly one order out of an enormous number. Build Alpha publishes a case where a backtest maximum drawdown of 1,663.90 dollars became 5,195.17 dollars, though that figure is the worst of their resample runs rather than a percentile of them.
What is block bootstrap and why does it matter for trading?
Block bootstrap resamples contiguous runs of trades or returns rather than individual ones, so win streaks, loss streaks and volatility clustering survive the resampling. It matters because clustered losses are what actually produce large drawdowns. Politis and Romano's stationary bootstrap, in the Journal of the American Statistical Association 89(428), 1994, pages 1303 to 1313, draws each block length at random instead of fixing it.
How many Monte Carlo simulations do I need?
The published answers conflict. A StrategyQuant blog post says the magic number of simulations is 200 or more, Build Alpha recommends 1,000 or more but adds that 100 or more is generally acceptable, Adaptrade's Market System Analyzer defaults to 500, and Timothy Masters says hundreds or thousands. What drives it is the percentile you want: a 99th percentile needs far more runs to stabilize than a median.
What confidence level should I use for Monte Carlo drawdown, 80%, 95% or 99%?
A StrategyQuant administrator, answering on the vendor's support forum, says it does not matter too much whether it is 80 or 95 percent. That is a weak answer, because the choice is a position-sizing decision rather than a statistical one. The confidence level sets how much capital you hold against a bad sequence, so it should be derived from what a breach would cost you, not from convention.
Should I size positions off the backtest drawdown or the Monte Carlo drawdown?
Off the Monte Carlo figure, because the backtest is one sample from the distribution. Build Alpha's SMH case shows the same strategy at a 30 percent drawdown at the 95th percentile on a 2,500 dollar account and 18 percent on a 4,000 dollar account, because position size is held flat at 100 shares and only the denominator moves. A 2002 EliteTrader heuristic goes further: double whatever drawdown the model predicts before sizing.
Does Monte Carlo simulation detect overfitting?
Mostly no. Reshuffling and resampling take the realized trade list as given, so a trade list produced by an overfit strategy is resampled just as happily as any other. Only the permutation test addresses it, and only partly. As a top-voted r/algotrading comment notes, every robustness test can be passed by chance if you kept the best survivor of many strategies.
Should Monte Carlo run on in-sample or out-of-sample trades?
Out-of-sample, where possible. Running a reshuffle on the trades you optimized over inherits the selection luck that got those parameters chosen, so the drawdown distribution is built from an already-flattering sample. The reshuffle widens that sample into a distribution but cannot recenter it. Most tools default to the full in-sample trade list, and Tradelyze's Monte Carlo check also runs on in-sample trades.
How many trades do you need before Monte Carlo means anything?
An EliteTrader thread gives 100 as a minimum and 500 as preferred. The sharper point is about composition rather than count: 40 trades drawn from one market regime, reshuffled, give you a confident-looking cloud built on a biased sample. Resampling cannot add information that the trade list does not contain, so a small or single-regime sample produces precise-looking output from thin evidence.
Why do I get different Monte Carlo results every run?
Tail percentiles are estimated from few effective observations, so they move whenever the random seed changes. A post on r/quant asks about two runs of 1,000 simulations returning minus 64 percent and minus 68 percent, although those are each run's extremes rather than percentiles. Raise the simulation count until your percentile stops moving, or report a range. Tradelyze uses a fixed random seed, so its own output does not show this movement.
Can Monte Carlo tell me my odds of passing a prop firm challenge?
Partly, and the mechanism is sequence risk: a profitable strategy can still breach a daily loss limit purely on the order in which losses arrive. Cutting risk per trade is the lever that moves the answer: smaller losses take longer to stack into a breach, so more orderings survive to the target without the strategy changing at all. Trade-close resampling cannot model intraday limits directly.
Do I need both Monte Carlo and walk-forward analysis?
Yes, because they answer different questions. Walk-forward analysis asks whether parameters fitted on one stretch of history still work on bars the optimizer never saw. Monte Carlo asks how much of the result depended on the particular order or selection of trades that happened to occur. A strategy can pass either one while failing the other.
What does Tradelyze's Monte Carlo simulation do?
Tradelyze runs a stationary block bootstrap: it redraws runs of your closed trades with replacement, 1,000 times by default, and measures how Sharpe ratios and maximum drawdowns spread across the draws. The robustness card shows MC Sharpe Mean, MC Sharpe 5th–95th, MC Sharpe Original, Ruin Probability and MC Max DD Real→P95, and a Ruin Check badge. Ruin is counted against your selected prop firm's total drawdown limit, or 10% when no firm is named, and must be below 20% to pass. A failed check keeps the robustness score at C+ and MARGINAL at best.
Sources
- Timothy Masters, Testing and Tuning Market Trading Systems: Algorithms in C++, Apress, 2018, and Permutation and Randomization Tests for Trading System Development: Algorithms in C++, 2020, together with the C++ source published as companion material to both — the OHLC decomposition into a close-to-open gap plus open-to-high, open-to-low and open-to-close offsets; the two independent permutations; the preservation of the net change from the first price to the last; the separation of in-sample, single-out-of-sample and walk-forward permutation tests; and the conservative (r+1)/(m+1) p-value. The 2018 program's own usage message asks for a replication count of "hundreds or thousands"; the widely repeated "at least 1,000" comes from a code comment in a 2006 paper of his about a different, position-vector scheme. Verified directly against
MCPT_BARS.CPPin the Apress companion source repository, which contains the bar decomposition, the two shuffle loops, the endpoint-preservation comment, the trend/bias/skill decomposition, the "hundreds or thousands" usage string and the "null hypothesis that system is worthless" p-value label. - Dimitris N. Politis and Joseph P. Romano, The Stationary Bootstrap, Journal of the American Statistical Association 89(428), 1994, pages 1303–1313 — block resampling in which the block length is drawn from a geometric distribution rather than fixed, so the resampled series remains stationary. A fixed block length, as used in this page's Figure 2, is the earlier moving-block bootstrap and not the Politis-Romano method; Tradelyze's Monte Carlo check uses the Politis-Romano stationary version.
- Andrew Paskaramoorthy, Terence van Zyl and Tim Gebbie, The bias of IID resampled backtests for rolling-window mean-variance portfolios, arXiv 2505.06383, 9 May 2025 — bias in backtests resampled as if returns did not affect each other, traced to resampling breaking the way returns follow on from earlier returns, with the authors' own caveat that the bias "can often be tolerable". Studies portfolio backtests, not trade-list reshuffling.
- David H. Bailey and Marcos López de Prado, The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality, SSRN 2460551.
- David H. Bailey, Jonathan M. Borwein, Marcos López de Prado and Qiji Jim Zhu, The Probability of Backtest Overfitting, SSRN 2326253.
- StrategyQuant — three separate kinds of source, distinguished because they carry different weight. Product documentation, Monte Carlo — trades manipulation (last updated 1 March 2019), states that randomizing trade order "doesn't change the resulting Net Profit"; that is the strongest vendor citation on this page. A company blog post, Monte Carlo explained in 6 minutes, is the source of "the magic number of simulations is 200+". A support-forum answer from a StrategyQuant administrator is the source of "does not matter too much whether it is 80 or 95%". The same documentation also describes Randomize Starting Bar, Randomize Strategy Parameters and Randomize History Data, none of which correspond to the five techniques tabulated in the section on which Monte Carlo test a report means.
- Build Alpha published Monte Carlo material, located at source on 16 September 2026 — the drawdown case and the simulation count are on buildalpha.com/monte-carlo-simulation, and the $SMH position-sizing case is on buildalpha.com/properly-funding-a-strategy-with-monte-carlo. The $1,663.90 backtest maximum drawdown against $5,195.17, which is verbatim "the worst resample drawdown from all simulations" and therefore a worst-of-N from a bootstrap with replacement, not a percentile and not a reshuffle of the same trades; the recommendation of "1,000 or more" simulations, whose own continuation reads "however, results are generally acceptable with 100 or more"; the $SMH position-sizing case showing 30% drawdown at the 95th percentile on a $2,500 account against 18% on a $4,000 account with position size held flat at 100 shares; and "all 1,000 equity curves end at the same total profit and loss amount but with wildly different paths".
- Adaptrade Software, Market System Analyzer — default of 500 simulations, documented on the Market System Analyzer Monte Carlo page.
- Kevin Davey, Monte Carlo Probability Cones, 2019, kjtradingsystems.com — trade independence is "generally true (or for all practical purposes true) for many trading systems", with serial correlation presented as detectable and correctable. The stronger paraphrase that serial correlation "invalidates the results" does not appear in the source.
- EliteTrader discussion threads — the 2002 practitioner heuristic of doubling the model's predicted drawdown before sizing, and the 100-trade minimum / 500-trade preference for Monte Carlo analysis. Practitioner conventions with no primary source.
- Practitioner community posts, cited in the body as such and without primary sources: a top-ranked r/algotrading comment on passing robustness tests by chance after keeping the best of many strategies, which also names White's Reality Check and the Bonferroni correction; and an r/quant post asking about 1,000 simulations giving −64% on one run and −68% on the next, in which those figures are each run's minimum and maximum rather than percentiles and the resampling is done with replacement.
- Figure 1 and Figure 2 are constructed illustrations, not measured data: computed demonstrations on synthetic trade series built for this page, not Tradelyze output. Figure 1 uses a 12-trade set summing to $1,200 on a $10,000 account across 200,000 reorderings. Figure 2 uses a 200-trade series on a $50,000 account in which trades arrive in runs of ten from the same regime, giving a lag-1 autocorrelation of +0.553, with 20,000 moving-block bootstrap resamples at each of block sizes 1, 5, 10 and 20. The twenty-series repeats quoted in its description and in the body text use the same construction with different generator seeds.
- Tradelyze implementation, reviewed 26 September 2026 — the stationary block bootstrap of closed-trade results with a typical block length chosen from the trade list's own clustering; 1,000 draws by default; a fixed random seed for the block bootstrap and the sign-flip permutation test; the Monte Carlo check running on the chosen settings' trades over the price history the optimizer searched; the ruin count against each selected prop firm's maximum total drawdown limit, left blank when no limit is known, and against the default rule set's 10% limit when a run names no firm; the rule that ruin must be below 20% to pass; the Ruin Check badge, reading Ruin Acceptable, Ruin Too High, or Too Few Trades when the check refused to run, for each firm; the Monte Carlo check's 29 of the 100 robustness score points; the rule that any failed scored check, a refusal for too few trades included, multiplies that firm's robustness points by 0.69, so its score is at most 69, C+ and MARGINAL at best; blank Monte Carlo figures below three closed trades; the checks running on the search's overall best settings, with an amber banner on a firm's card when that firm was recommended other settings; the evenly spaced bar grid every Monte Carlo Sharpe is measured on; the 95th-percentile maximum drawdown, shown on the robustness card as MC Max DD Real→P95; and out-of-sample results on the walk-forward card only when it re-tunes each window, and on each firm's held-out test, which by default withholds the last 25% of the history from the whole search.