Tier 3 · Practitioner · Module 3.4
Intro to automated backtesting concepts (Python / Pine Script)
How coded backtests work, a simple Pine Script strategy and Python example, and the biases that make automated results look far better than reality — look-ahead, overfitting, survivorship, and unrealistic costs.
Lesson 2 of 3 · 7 min read
A manual backtest of 100 trades might take a weekend. A coded backtest can test thousands of trades across years of data in seconds. That speed is powerful — and dangerous. A computer will happily produce a beautiful equity curve from a strategy that could never have worked in real life. This lesson introduces how automated backtesting works, shows two simple examples, and — most importantly — teaches you the traps to recognise, whether you write the code yourself or evaluate someone else's results.
What you'll learn
- What automated backtesting needs: rules, data, and a cost model
- A simple strategy written in TradingView's Pine Script
- How the same idea looks in Python — and why "shift" matters
- The biases that inflate backtest results
- How to judge any backtest you're shown, including a trading bot's
1. What you need
| Ingredient | What it means |
|---|---|
| Rules as code | Every decision in your written strategy expressed precisely — computers can't "use judgement" |
| Historical data | Clean price data for your market and timeframe (and volume if used) |
| A cost model | Spread, commission, slippage, and swaps — or results will be overstated |
| Metrics | Expectancy, profit factor, maximum drawdown, number of trades (see Manual backtesting on TradingView) |
The written strategy from Module 3.3 is the specification. If a rule is hard to code, it was probably ambiguous to begin with.
2. Example: Pine Script (TradingView)
Pine Script is TradingView's built-in language. A strategy() script runs your rules on the chart's history and reports results in the Strategy Tester.
This simplified example buys pullbacks to a 50 EMA in an uptrend defined by a rising 200 SMA, with an ATR-based stop and a 2R target. (For learning only — not a recommendation, and position sizing is simplified.)
//@version=5
strategy("Trend pullback (demo)", overlay=true, initial_capital=10000,
default_qty_type=strategy.percent_of_equity, default_qty_value=10,
slippage=2)
fastLen = input.int(50, "Fast EMA")
slowLen = input.int(200, "Slow SMA")
atrLen = input.int(14, "ATR length")
atrMult = input.float(1.5, "Stop (x ATR)")
rr = input.float(2.0, "Reward : risk")
fast = ta.ema(close, fastLen)
slow = ta.sma(close, slowLen)
atr = ta.atr(atrLen)
uptrend = close > slow and slow > slow[1]
pullback = low <= fast and close > fast
if uptrend and pullback and strategy.position_size == 0
stopPrice = close - atrMult * atr
targetPrice = close + rr * atrMult * atr
strategy.entry("Long", strategy.long)
strategy.exit("Exit", "Long", stop=stopPrice, limit=targetPrice)
plot(fast, "Fast EMA", color.blue)
plot(slow, "Slow SMA", color.gray)
What to notice
- Every rule is explicit: "uptrend", "pullback", stop, and target are all defined numerically.
slippage=2adds a small execution penalty; set spread and commission to match your broker in the strategy's Properties so results reflect real costs.- Signals are evaluated when a bar closes and the entry fills on the next bar — as it would live.
3. Example: Python
Python gives full control and is popular for research. Libraries such as pandas handle data, and open-source frameworks (for example backtesting.py, backtrader, or vectorbt) provide backtest engines.
The core idea — a trend filter — in a few lines of pandas:
import pandas as pd
df = pd.read_csv("eurusd_h4.csv", parse_dates=["time"], index_col="time")
df["sma200"] = df["close"].rolling(200).mean()
df["signal"] = (df["close"] > df["sma200"]).astype(int)
# Act on the NEXT bar: a signal can only use data from bars that have closed.
df["position"] = df["signal"].shift(1).fillna(0)
cost = 0.0001 # rough cost per position change, e.g. 1 pip on EUR/USD
df["strategy_ret"] = df["close"].pct_change() * df["position"]
df["strategy_ret"] -= df["position"].diff().abs() * cost / df["close"]
The shift(1) line is the most important one in the example. Without it, the code would "trade" on the same bar whose close generated the signal — using information you couldn't have had. That single missing line is one of the most common reasons homemade backtests look too good.
4. The biases that inflate results
| Bias | What happens | Defence |
|---|---|---|
| Look-ahead bias | The test uses data that wasn't available at the time of the decision | Act on the next bar; use only completed bars; check multi-timeframe data carefully |
| Overfitting (curve-fitting) | Parameters are tuned until the past looks perfect — the strategy has learned noise | Few parameters, logical rules, out-of-sample testing |
| Data snooping | Testing hundreds of variations and keeping the best — some will look great by pure chance | Decide what to test in advance; validate the winner on unseen data |
| Survivorship bias | Testing only assets that still exist today (common with stocks) | Use data that includes delisted instruments |
| Unrealistic costs | Zero spread, zero slippage, no swaps | Model realistic costs — then add a safety margin |
| Unrealistic execution | Assuming fills at exact prices during gaps or news | Add slippage; avoid strategies that rely on perfect fills |
Worked example: spotting overfitting
(Illustrative.) A trader optimises a moving-average crossover and finds that a 37/113 combination produced a 4.5 profit factor on 2019–2023 data. Neighbouring settings (35/110, 40/120) show profit factors around 1.0. On 2024 data, the 37/113 version loses money.
Diagnosis: the "best" setting was a narrow peak fitted to past noise. A robust strategy tends to show similar results across a range of nearby settings — a plateau, not a spike.
5. Judging someone else's backtest
This matters whether you're evaluating a trading course, a signal service, or a trading bot. Ask:
- How many trades, over what period and which market conditions?
- Are costs included — spread, commission, slippage, swaps?
- Was there out-of-sample or forward testing, or is it all in-sample?
- What was the maximum drawdown, and the longest losing streak?
- How many parameters, and how were they chosen?
- Is there a live or forward-tested track record to compare against the backtest?
A backtest with no drawdown information, no costs, and no out-of-sample test is marketing, not evidence.
Common beginner mistakes
- Trusting a beautiful equity curve without checking for look-ahead bias.
- Optimising dozens of parameters until the past looks perfect.
- Ignoring costs and slippage.
- Testing on too little data or only one market regime.
- Going live straight from an in-sample backtest.
Key terms
| Term | Meaning |
|---|---|
| Pine Script | TradingView's scripting language for indicators and strategies |
| Strategy Tester | TradingView's panel that reports backtest results for a strategy script |
| pandas | A Python library for working with tabular data |
| Look-ahead bias | Using information that wasn't available at decision time |
| Overfitting | Tuning a strategy to past noise so it fails on new data |
| Data snooping | Finding patterns by testing many variations until one looks good |
| Survivorship bias | Testing only on assets that survived to the present |
| Parameter plateau | A range of settings with similarly good results — a sign of robustness |
Practice
- If you use TradingView, open the Pine Editor, paste the example, and add it to a chart of your market. Set realistic commission and slippage in Properties and read the Strategy Tester summary.
- Change the fast and slow lengths to several nearby values. Are results similar (a plateau) or wildly different (a spike)?
- Write down which of the six biases your own manual backtest could be vulnerable to, and how you'll guard against each.
- Take any strategy or bot performance claim you've seen and apply the six questions in section 5.
Quick recap
- Automated backtests need rules as code, clean data, and a realistic cost model.
- Pine Script and Python can both test strategies — the logic matters more than the tool.
- Act on the next bar and use only completed data to avoid look-ahead bias.
- Beware overfitting, data snooping, survivorship bias, and unrealistic costs.
- Look for plateaus, out-of-sample results, and forward-tested track records — in your tests and in anyone else's.
Educational content only — not financial advice. Trading involves substantial risk of loss. Practise on a demo account before risking real money.
