DNS Research. Trading and Investing Blog. Free articles every day.

Sharpe Ratio, Max Drawdown, and Win Rate: Which Backtest Metrics Matter Most?

advertisement

Sharpe Ratio: The Risk-Adjusted Return Standard

The Sharpe Ratio, developed by Nobel laureate William F. Sharpe in 1966, remains the most widely cited metric in quantitative finance. It measures excess return per unit of total risk, expressed as:

Sharpe Ratio = (Rp − Rf) / σp

where Rp is the portfolio return, Rf is the risk-free rate, and σp is the standard deviation of portfolio returns. A Sharpe of 1.0 is considered acceptable, 2.0 is excellent, and anything above 3.0 in a live strategy warrants deep skepticism about overfitting or data errors.

What the Sharpe Ratio Captures Well

The Sharpe Ratio’s genius lies in its simplicity. It answers a fundamental question: how much return am I getting for each unit of volatility I endure? This makes it invaluable for comparing strategies with different return profiles. A strategy returning 40% annually with 30% volatility (Sharpe ≈ 1.3) is arguably inferior to one returning 18% with 6% volatility (Sharpe = 3.0), because the latter delivers smoother, more predictable compounding.

For portfolio construction, the Sharpe Ratio underpins mean-variance optimization and the Capital Asset Pricing Model. When allocators compare hedge funds, CTAs, or systematic strategies, Sharpe is the lingua franca.

Critical Limitations

The Sharpe Ratio assumes returns are normally distributed—a dangerous assumption in trading. Real return distributions exhibit fat tails, skewness, and kurtosis. A strategy selling far out-of-the-money options can post a Sharpe of 2.5 for years while harboring catastrophic tail risk. The Sharpe Ratio is blind to this asymmetry.

It also penalizes upside volatility identically to downside volatility. A strategy with explosive winning months gets punished despite those gains being desirable. This is why the Sortino Ratio, which divides excess return by downside deviation only, often complements Sharpe analysis.

Finally, Sharpe is time-sensitive and scale-dependent. Annualizing a daily Sharpe by multiplying by √252 assumes independent, identically distributed returns—rarely true in practice due to autocorrelation and volatility clustering. Always verify whether a reported Sharpe is computed from daily, weekly, or monthly data.

Practical Benchmarking

Strategy Type Realistic Sharpe (Net)
Long-only equity index 0.3–0.5
Global macro 0.7–1.0
Market-neutral equity 1.0–1.5
High-frequency trading 3.0–10.0+
Retail algorithmic (live) 0.5–1.2

If your backtest shows a Sharpe above 3.0 on daily data for a medium-frequency strategy, assume you have a bug before assuming you have alpha.


Maximum Drawdown: The Metric That Ends Careers

Maximum Drawdown (MDD) measures the largest peak-to-trough decline in equity before a new peak is reached:

MDD = (Trough Value − Peak Value) / Peak Value

A 40% MDD means an account must gain 66.7% just to break even. This asymmetry is why MDD is often called the “career risk” metric—it determines whether a strategy survives long enough to realize its expected return.

Why MDD Matters More Than Sharpe for Survival

Consider two strategies:

  • Strategy A: Sharpe 1.8, MDD 45%, recovery time 3 years
  • Strategy B: Sharpe 1.2, MDD 12%, recovery time 4 months

Most professional allocators choose Strategy B. The reason is behavioral and structural. Investors redeem during drawdowns. Margin calls force liquidation. Psychological capital erodes. A strategy with a deep MDD may be mathematically superior over 20 years but practically uninvestable over the 3-year window that matters to stakeholders.

Drawdown Duration: The Hidden Killer

MDD alone is insufficient. Drawdown duration—the time from peak to recovery—often inflicts more damage than depth. A 15% drawdown lasting 30 months is functionally worse than a 25% drawdown recovering in 6 weeks. Track:

  • Max Drawdown Depth: Worst peak-to-trough decline
  • Max Drawdown Duration: Longest time underwater
  • Average Drawdown: Typical pain level
  • Ulcer Index: Root-mean-square of drawdowns, weighting both depth and duration

Backtest Drawdown Underestimation

Backtested MDD is almost always optimistic. Reasons include:

  1. Survivorship bias in the tested universe
  2. Look-ahead bias in signal construction
  3. In-sample optimization that smooths the equity curve
  4. Insufficient Monte Carlo testing of trade order permutations

A robust practice: multiply your backtest MDD by 1.5–2.0x when sizing positions for live deployment. If the resulting drawdown exceeds your risk tolerance, reduce leverage before going live, not after.


Win Rate: The Most Misunderstood Metric

Win Rate = (Winning Trades / Total Trades) × 100. It is the metric retail traders obsess over and professionals treat with suspicion.

Why High Win Rates Can Destroy Accounts

A 90% win rate strategy sounds magnificent until you examine payoff asymmetry. If the average win is $100 and the average loss is $1,200, the expectancy is:

(0.90 × $100) − (0.10 × $1,200) = $90 − $120 = −$30 per trade

This is the classic “picking pennies in front of a steamroller” profile—typical of premium-selling strategies, martingale systems, and grid trading. The equity curve looks beautiful for months, then a single event wipes out years of gains.

Conversely, trend-following CTAs often win only 35–45% of trades but achieve profit factors above 2.0 because winning trades run far longer than losers.

Win Rate in Context: The Expectancy Framework

Win rate is meaningless without pairing it with average win/loss size. The core equation:

Expectancy = (Win% × Avg Win) − (Loss% × Avg Loss)

Or equivalently, using the reward-to-risk ratio (R):

Expectancy per unit risk = (Win% × R) − (Loss% × 1)

Win Rate Required R:R to Break Even
20% 4.0
33% 2.0
50% 1.0
67% 0.5
80% 0.25

This table reveals why win rate alone tells you nothing. Any strategy can be made profitable or unprofitable by adjusting the R:R ratio—until transaction costs and slippage enter the picture.

Win Rate and Psychological Sustainability

Despite its analytical weakness, win rate matters behaviorally. Traders with win rates below 40% must endure long losing streaks—10 consecutive losses is statistically expected every ~100 trades at a 40% win rate. Most humans abandon profitable systems during such streaks. A 55% win rate system with a 1.2 R:R may underperform a 40% win rate system with a 2.5 R:R on paper, but the former is more likely to be executed consistently.


Comparing the Three Metrics Side by Side

Metric Measures Strength Weakness Best Used For
Sharpe Ratio Risk-adjusted return Standardized comparison Assumes normality; ignores tails Strategy ranking, allocation
Max Drawdown Worst-case pain Captures survival risk Backward-looking; single path Position sizing, leverage limits
Win Rate Trade accuracy Intuitive; behavioral Ignores payoff size Diagnostic, not decision-making

No single metric suffices. Each answers a different question: Sharpe asks “is the return worth the volatility?”, MDD asks “can I survive the worst case?”, and win rate asks “how often am I right?”


Metrics That Belong Beside the Big Three

Sophisticated quants rarely rely on Sharpe, MDD, and win rate alone. Complementary metrics include:

  • Sortino Ratio: Excess return divided by downside deviation; corrects Sharpe’s symmetric penalty
  • Calmar Ratio: Annualized return ÷ Max Drawdown; directly ties performance to pain
  • Profit Factor: Gross profits ÷ gross losses; intuitive and robust
  • Expectancy: Average profit per trade; the engine of long-term compounding
  • Tail Ratio: 95th percentile gain ÷ 5th percentile loss; exposes hidden tail risk
  • Exposure-Adjusted Metrics: Returns per unit of time in market, not per unit of calendar time
  • Deflated Sharpe Ratio: Adjusts for multiple testing and non-normality (Bailey & López de Prado)

The Overfitting Trap: When Great Metrics Lie

A backtest with Sharpe 3.5, MDD 8%, and win rate 72% is not a dream—it is a warning. With enough parameter combinations, any dataset yields spectacular in-sample results. The probability of finding a spurious Sharpe above 2.0 grows rapidly with the number of trials.

Defenses Against Backtest Illusion

  1. Out-of-sample testing: Reserve 30–40% of data never touched during development
  2. Walk-forward analysis: Re-optimize on rolling windows, test on subsequent unseen data
  3. Monte Carlo permutation: Shuffle trade order 10,000 times to see the MDD distribution
  4. Parameter sensitivity: Profitable strategies should show a plateau, not a spike
  5. Cross-market validation: Test the same logic on correlated instruments
  6. Cost realism: Include commissions, slippage, borrow costs, and market impact
  7. Deflated Sharpe: Correct for the number of strategies tried

A strategy that survives all seven filters with a Sharpe of 1.2 is worth more than one that fails three filters with a Sharpe of 4.0.


Which Metric Matters Most? It Depends on Your Constraint

The question “which backtest metric matters most?” has no universal answer because it depends on the binding constraint:

  • If you manage outside capital: Max Drawdown dominates. Redemptions kill funds faster than poor returns.
  • If you trade your own capital with leverage: Max Drawdown still dominates—margin calls are unforgiving.
  • If you allocate across many strategies: Sharpe Ratio dominates—it standardizes comparison across volatility regimes.
  • If you trade discretionary or semi-systematic: Win Rate matters for psychological sustainability, even if analytically secondary.
  • If you sell options or run mean-reversion: Tail metrics (skew, kurtosis, tail ratio) matter more than any of the big three.
  • If you run trend-following: Profit Factor and expectancy dominate; win rate is a distraction.

A Practical Hierarchy

For most systematic traders deploying real capital, the priority order is:

  1. Max Drawdown (with duration) — determines survival
  2. Expectancy / Profit Factor — determines profitability
  3. Sharpe Ratio (or Sortino) — determines efficiency
  4. Win Rate — determines executability
  5. Tail metrics — determine robustness

Sharpe without survivable drawdown is academic. Drawdown without positive expectancy is slow death. Expectancy without tolerable win rate is abandoned mid-stream. All four must align.


Building a Composite Scorecard

Rather than fixating on one number, build a scorecard. A sample weighting for a medium-frequency systematic strategy:

  • Max Drawdown (30%): Score inversely, target < 20%
  • Sharpe or Sortino (25%): Target > 1.5 net of costs
  • Profit Factor (20%): Target > 1.5
  • Win Rate × R:R consistency (15%): Target expectancy > 0.2R
  • Tail Ratio / Skew (10%): Target > 1.0

Weight adjustments follow constraints. A fund manager facing quarterly redemption windows should shift weight toward drawdown duration. A prop trader with a hard daily loss limit should weight tail metrics more heavily.


Final Analytical Principle

Backtest metrics are diagnostic tools, not goals. Optimizing for any single metric—whether Sharpe, MDD, or win rate—invites overfitting to that metric’s blind spots. The trader who chases Sharpe builds strategies with hidden tail risk. The trader who chases low drawdown builds strategies that never take risk. The trader who chases win rate builds strategies with catastrophic payoff asymmetry.

The correct approach: define your binding constraint first (survival, capital efficiency, or psychological tolerance), select the metric that measures that constraint, then validate with at least three complementary metrics before risking capital. Metrics serve the strategy. The strategy serves the trader’s objective. Invert that hierarchy and the market will correct the error—expensively.

advertisement

latest posts

Something went wrong. Please refresh the page and/or try again.

Discover more from DNS Research

Subscribe now to keep reading and get access to the full archive.

Continue reading