Backtesting is the process of simulating a trading strategy against historical market data to evaluate its viability before risking real capital. Python has become the dominant language for this task thanks to libraries like pandas, NumPy, and vectorbt. This tutorial walks you through a complete, reproducible backtesting workflow using only open-source tools.
Step 1: Install and Import Dependencies
Start by installing the core libraries. You will need pandas for data handling, NumPy for numerical operations, yfinance for free market data, and matplotlib for visualization.
pip install pandas numpy yfinance matplotlib
Then import them in your script. Keeping imports at the top makes your backtest reproducible and easy to audit later.
import pandas as pd
import numpy as np
import yfinance as yf
import matplotlib.pyplot as plt
Step 2: Download Historical Price Data
Reliable data is the foundation of any credible backtest. For this tutorial, we will use adjusted close prices for SPY, the SPDR S&P 500 ETF, covering fifteen years of daily bars.
data = yf.download("SPY", start="2009-01-01", end="2024-01-01", auto_adjust=True)
prices = data["Close"]
Always use auto-adjusted prices. Unadjusted data includes splits and dividends that distort returns and produce false signals. Fifteen years gives you roughly 3,700 trading days, enough to cover bull markets, bear markets, and sideways chop.
Step 3: Define the Trading Strategy
We will implement a classic trend-following rule: go long when the 50-day simple moving average crosses above the 200-day simple moving average (the golden cross) and move to cash on the death cross.
df = pd.DataFrame({"price": prices})
df["sma_fast"] = df["price"].rolling(50).mean()
df["sma_slow"] = df["price"].rolling(200).mean()
df["signal"] = np.where(df["sma_fast"] > df["sma_slow"], 1.0, 0.0)
The signal column holds 1 when the strategy is invested and 0 when it is flat. Shifting this signal by one bar prevents look-ahead bias, the single most common mistake in amateur backtests.
df["position"] = df["signal"].shift(1).fillna(0)
Step 4: Calculate Strategy Returns
Compute daily percentage returns on the underlying asset, then multiply by the position to obtain strategy returns.
df["market_return"] = df["price"].pct_change().fillna(0)
df["strategy_return"] = df["position"] * df["market_return"]
df["equity_curve"] = (1 + df["strategy_return"]).cumprod()
df["benchmark_curve"] = (1 + df["market_return"]).cumprod()
The equity curve is the compounded growth of one dollar invested. Comparing it against the benchmark curve reveals whether your rules add value over simple buy-and-hold.
Step 5: Account for Transaction Costs
Ignoring costs inflates results, especially for strategies that trade frequently. A realistic assumption for SPY is five basis points per trade, covering commission, spread, and slippage.
cost_per_trade = 0.0005
df["trade"] = df["position"].diff().abs().fillna(0)
df["strategy_return_net"] = df["strategy_return"] - df["trade"] * cost_per_trade
df["equity_curve_net"] = (1 + df["strategy_return_net"]).cumprod()
Apply costs whenever the position changes, not on every bar. This single adjustment frequently turns a marginal strategy into a losing one, which is exactly why backtesting matters.
Step 6: Compute Performance Metrics
Raw equity curves are hard to compare. Calculate annualized return, volatility, Sharpe ratio, maximum drawdown, and win rate.
def metrics(returns, label):
ann_return = (1 + returns).prod() ** (252 / len(returns)) - 1
ann_vol = returns.std() * np.sqrt(252)
sharpe = ann_return / ann_vol if ann_vol else np.nan
curve = (1 + returns).cumprod()
drawdown = (curve / curve.cummax() - 1).min()
return {"Strategy": label, "CAGR": ann_return, "Vol": ann_vol,
"Sharpe": sharpe, "MaxDD": drawdown}
print(metrics(df["strategy_return_net"], "SMA Crossover"))
print(metrics(df["market_return"], "Buy & Hold"))
The Sharpe ratio divides excess return by volatility; anything above one is respectable for a daily strategy. Maximum drawdown shows the worst peak-to-trough loss, which determines whether you could actually stomach the strategy in live trading.
Step 7: Visualize the Results
Charts expose regime dependence that single numbers hide. Plot both equity curves on a log scale to compare compounding fairly.
plt.figure(figsize=(12, 6))
plt.plot(df["equity_curve_net"], label="SMA Crossover (net)")
plt.plot(df["benchmark_curve"], label="Buy & Hold", alpha=0.7)
plt.yscale("log")
plt.legend()
plt.title("SPY: Strategy vs Benchmark")
plt.show()
You will typically notice the strategy underperforms in strong uptrends but shines during prolonged declines, since the death cross moves it to cash.
Step 8: Avoid the Most Common Pitfalls
First, never use future data. Shift signals forward and avoid indicators that reference the current bar’s close when trading that same close. Second, beware overfitting: testing dozens of parameter combinations until one looks great produces curve-fitted results that fail live. Third, account for survivorship bias by using data that includes delisted assets when testing equities. Fourth, validate out-of-sample by splitting your data into an in-sample period for design and an out-of-sample period for confirmation. Fifth, remember that dividends matter; auto-adjusted prices handle this for ETFs.
Step 9: Walk-Forward Testing
A single static backtest is fragile. Walk-forward analysis repeatedly optimizes parameters on a rolling window and tests on the next unseen window, mimicking real deployment.
for start in range(0, len(df) - 1000, 250):
train = df.iloc[start:start + 750]
test = df.iloc[start + 750:start + 1000]
# optimize on train, evaluate on test
Aggregate the out-of-sample results. If performance collapses, your edge was likely an artifact of the training data.
Step 10: Extend to a Framework
For production research, hand-rolled loops become unwieldy. Vectorized backtesting with vectorbt or event-driven engines like backtrader and zipline-reloaded handle position sizing, shorting, multi-asset portfolios, and order types. Migrate your logic once the concept is validated in a simple pandas prototype, since debugging a framework is far harder than debugging twenty lines of pandas.
Move next to risk management: position sizing, stop losses, and volatility targeting typically matter more than entry signals. A mediocre signal with disciplined sizing often beats a brilliant signal with reckless leverage.







