The Probability of Backtest Overfitting (PBO) by David Bailey is the only statistical metric that quantifies the exact risk of backtest overfitting. Introduced in 2014 by Bailey and Marcos Lopez de Prado in the Journal of Computational Finance, this formula gives a single number between 0 and 1 representing the probability that your optimized strategy is a product of chance rather than a genuine market edge. Understanding and applying PBO has become essential before going live or submitting a strategy to a prop firm like FTMO or MyFundedFX.
The overfitting problem as measured by Bailey
Overfitting is the leading cause of failure when trading strategies move from simulation to live markets. An impressive backtest can simply reflect adaptation to past data rather than a genuine edge. Before Bailey's work, no formula allowed traders to quantify this risk precisely.
Context: the 2014 research
In 2014, David Bailey, Jonathan Borwein, Marcos Lopez de Prado, and Qiji Zhu published "Pseudo-mathematics and Financial Charlatanism" in the Notices of the American Mathematical Society. This study showed that with only 45 free parameters tested sequentially, it becomes near-certain to find a combination producing a Sharpe Ratio above 1 on entirely random data. In other words, most "good" strategies are simply the result of systematically searching a parameter space, not a true market edge.
Source: Bailey et al. (2014), Notices of the AMS
That same year, Bailey and Lopez de Prado formalized the PBO calculation using a combinatorial cross-validation method, published in the Journal of Computational Finance. Their core finding: a strategy tested across N different configurations carries a growing probability of being overfitted, regardless of how strong the in-sample backtest looks.
The PBO formula (Probability of Backtest Overfitting)
PBO is defined as the probability that the best in-sample strategy underperforms the median of all strategies on out-of-sample data:
PBO = P(SBO less than 0)
The SBO (Stochastic Backtest Optimization) is the log ratio of the best in-sample strategy's out-of-sample performance relative to the median of all N strategies. A negative SBO means the winning backtest strategy ends up below the median on unseen data: a clear sign of overfitting.
Why measure against the median?
Bailey uses the median (not absolute performance) to control selection bias. If the best in-sample strategy does not outperform half of all other strategies out-of-sample, it was optimized on historical noise, not on the real structure of the market.
How to calculate PBO
The PBO calculation relies on Combinatorial Purged Cross-Validation (CPCV). Here is the step-by-step process.
Combinatorial cross-validation method
Unlike standard k-fold cross-validation, Bailey's method tests all possible combinations of in-sample and out-of-sample partitions. For S = 6 sub-periods, this produces C(6,3) = 20 combinations, giving a statistically robust estimate of overfitting risk.
Split data into S sub-periods
Generate all combinations
Evaluate out-of-sample performance
Calculate SBO per combination
Estimate the final PBO
Source: Bailey and Lopez de Prado (2014), SSRN preprint
Interpreting the result (0 to 1)
| PBO Value | Interpretation | Recommended Action |
|---|---|---|
| PBO below 0.1 | Very robust strategy, minimal overfitting risk | Move to forward testing on demo account |
| 0.1 to 0.3 | Low risk, acceptable robustness | Validate with walk-forward before going live |
| 0.3 to 0.5 | Moderate risk, partial over-optimization likely | Reduce the number of free parameters |
| PBO above 0.5 | Overfitting probable per Bailey | Do not trade this in a live account |
| PBO above 0.8 | Overfitting near-certain | Reject the strategy and start over |
Practical thresholds for traders
Bailey himself recommends PBO = 0.5 as the decision threshold. For strategies targeting prop firms (FTMO, MyFundedFX, The5ers), professional quantitative traders aim for PBO below 0.1, even though this standard is not yet explicitly required by prop firms.
PBO is necessary but not sufficient
A low PBO validates statistical robustness on historical data, but does not guarantee future performance. Market conditions change. Combining PBO with walk-forward optimization is best practice for a complete validation.
Applying Bailey to your strategies
How many backtests are needed to interpret PBO
The more strategy configurations you test on the same data, the higher the PBO will mechanically go. Bailey demonstrates that PBO depends directly on N, the number of candidate strategies tested. With N = 50 configurations tested on the same historical period, a PBO of 0.5 or more is expected even if the winning configuration looks robust.
The practical rule: limit the number of combinations you test. If you explored 100 variants to find your optimal parameters, your PBO will be high regardless of the in-sample performance. If your strategy was built with 3 to 5 parameters and tested in fewer than 20 configurations, the PBO remains interpretable.
The role of the number of free parameters
The number of free parameters is the primary driver of overfitting. Each added parameter (moving average period, stop-loss level, volatility filter) multiplies the possible configurations and increases the risk that the optimal setup is fitted to historical noise rather than market logic.
Bailey's rule: 2k combinations for k binary parameters
With k parameters that can each be toggled on or off, the number of configurations is 2 to the power of k. For k = 10 binary parameters, that gives 1,024 configurations: PBO will almost certainly exceed 0.5 if you pick the best one. Define your market logic first, then optimize only 1 to 2 parameters.
Combining PBO and walk-forward
The combination of PBO and walk-forward optimization is the most rigorous validation method available to retail traders. PBO validates overall statistical robustness across all possible partitions of the data. Walk-forward validates parameter stability across real successive time windows.
The recommended sequence before going live:
For a deep dive into walk-forward methods, see our guide walk-forward optimization and backtesting validation.
Tools to calculate PBO
Open-source Python solutions
Marcos Lopez de Prado's mlfinlab library (Machine Learning Financial Laboratory) implements CPCV directly. For manual calculation, Python's itertools.combinations generates all sub-period pairs, and numpy handles the performance calculations. The learning curve is steep, but the method is fully documented in Bailey's original papers.
Platforms with built-in PBO
The vast majority of retail backtesting platforms (TradingView, MetaTrader 5) do not calculate PBO natively. Professional quantitative tools require coding skills, which is a real barrier for no-code traders. For an overview of available options, see our backtesting platform comparison.
Backtrex and multi-period validation
Backtrex addresses the overfitting problem through automatic multi-period validation. By testing your strategy across multiple distinct time windows (in-sample for configuration, then automatic out-of-sample validation across several periods), Backtrex generates consistency indicators in the spirit of Bailey's method: a solid strategy should perform consistently on unseen periods, not just on the optimization window.
Backtrex is accessible with a free account including a guided tour, then a lifetime Pro or Max license paid once, with 14 days to change your mind. Explore our backtesting features to learn more.
For further reading on robustness methods, see our guide stress testing and backtesting robustness or our article on backtest overfitting red flags.
Important Risk Warning
Conclusion
David Bailey's PBO formula is the most rigorous statistical tool available to determine whether a backtest reflects a genuine edge or simply adaptation to past data. With PBO below 0.1, a strategy passes the most demanding robustness test in the field. Combining this calculation with walk-forward validation and the multi-period testing available in Backtrex is the standard that separates serious traders from perfecting-the-backtest chasers.
PBO is a statistical metric developed by David Bailey and Marcos Lopez de Prado in 2014. It measures the probability that a strategy selected for its in-sample performance will achieve results below the median on out-of-sample data. The calculation uses combinatorial cross-validation across all possible partitions of the historical data. A PBO of 0 means no overfitting risk; a PBO of 1 means certain overfitting.
Bailey suggests that a PBO above 0.5 indicates probable overfitting: the strategy is statistically more likely to be a product of chance than to represent a genuine market edge. A PBO below 0.1 is considered a strong robustness signal. Between 0.1 and 0.5, additional walk-forward optimization validation is recommended before live trading.
Without coding skills, calculating the exact PBO is difficult. You can however apply Bailey's methodology in spirit by testing your strategy across multiple distinct time windows: one in-sample period for optimization, then several successive out-of-sample periods for validation. If performance drops significantly on unseen periods, overfitting is likely. Backtrex offers automatic multi-period validation with no code required.
Yes, PBO is fundamentally different from the out-of-sample Sharpe Ratio. The Sharpe measures risk-adjusted performance over a given period. PBO measures the probability that selecting the best in-sample strategy is a product of statistical chance, by testing all possible cross-validation configurations. A positive out-of-sample Sharpe is necessary but not sufficient: PBO evaluates whether that result is statistically robust across all possible data partitions.
To keep PBO below 0.3, the empirical rule is to test fewer than 20 parameter configurations in total. With 5 binary parameters, you already have 32 configurations. Bailey's recommendation: define strategy rules based on market logic first, then optimize only 1 to 2 parameters with few possible discrete values.
Yes, PBO applies to any trading strategy, including Smart Money Concepts (SMC) and ICT approaches. These strategies often contain many implicit parameters (order block levels, fair value gap size, session time filters) that increase overfitting risk. Calculating PBO on an SMC strategy backtested over several years helps distinguish a genuine edge from a configuration over-fitted to the historical optimization period.
PBO is a global statistical measure of overfitting probability, calculated via combinatorial cross-validation across all available data. Walk-forward is a sequential validation method that tests the strategy on successive unseen future periods, one after another. The two are complementary: PBO gives a global probabilistic signal, walk-forward validates temporal parameter stability. Using both together constitutes the most rigorous validation method available to retail traders.