Portfolio Backtest

Compare portfolios over the same historical window. Inspect returns, drawdowns, cash flows, and the assumptions behind them.

On this page

First run

Compare a portfolio of 60% SPY and 40% AGG with a portfolio holding only SPY. This is the historical comparison shown on the homepage.

The example runs without an account. Sign in to save it.

  1. Inspect the scenario. Both portfolios begin with $10,000 over 2005-01-03 through 2025-12-31. The example uses total return, monthly rebalancing, nominal dollars, no cash flows, and no tax-aware accounting.
  2. Read growth and loss together. Follow portfolio value through time, then inspect drawdown: the decline from a previous peak. The ending value alone does not describe the path.
  3. Change one assumption. Adjust an allocation or date window. The existing result remains tied to the last completed run until you run again.
  4. Run and compare. Check the updated assumptions, value chart, drawdown, and metrics together. A different window or cash flow schedule is a different experiment.
Worked example · historical growth and loss
PortfolioEnding valueAnnualized growthDeepest loss from peak
Classic 60/40$49,3327.90%35.56%
S&P 500 only$83,56310.64%55.19%

Annualized growth summarizes the full period. The deepest loss shows the size of a decline from a prior peak. Read both measures alongside the full path and its assumptions. These are historical results, not forecasts.

Watch a captioned walkthrough

Video walkthrough

Open the guided player with chapters →Try the example shown in this video →
Read the transcript

0:00Portfolio Backtest replays any investment mix through real market history.

0:05Compare a 60/40 portfolio with a gold-inclusive mix through 2008.

0:12Tax-aware stays off because this compares returns before taxes. Choose it for an after-tax comparison.

0:20Both start with ten thousand dollars to keep the balances on one scale. Change that amount to use a different dollar basis.

0:29The window from 2007 through 2012 fits the crash and recovery question. A different window tests different market conditions.

0:39Advanced settings define the result basis. Total return includes distributions; nominal dollars fit this replay and Holdout stays off to keep the full window. Change these for price-only returns, purchasing power, or a reserved test period.

0:57The baseline holds sixty percent SPY and forty percent AGG, rebalanced monthly to match our benchmark. Change it when you want a different comparison.

1:08A second portfolio puts both mixes on the same result page.

1:12Enter the gold mix alongside the stock-and-bond baseline.

1:18This mix holds forty five percent in SPY, twenty five percent in AGG, and thirty percent in GLD. That moves fifteen percentage points from stocks and fifteen from bonds into gold. Keeping monthly rebalancing isolates allocation. Change these weights to compare another mix.

1:39Run the comparison. Only the allocation weights differ.

1:43Predict whether the gold mix ends higher and falls less. The chart shows where their paths separate.

1:49The gold mix ends near sixteen thousand three hundred dollars. The baseline ends near twelve thousand nine hundred dollars, from the same ten thousand dollar start. That is about thirty five hundred dollars more before taxes.

2:04The baseline's deepest peak-to-trough loss was thirty five point five percent. The gold mix fell twenty eight point three percent, a shallower decline in this window.

2:15Within this fixed window, recovery tracks the climb back to a previous high. The gold mix starts twenty eight point three percent down. The baseline starts thirty five point five percent down, with more ground to regain. The curves show how that recovery unfolds.

2:32The window, basis, and rebalancing rule sit beside the result. Saving keeps these assumptions with the numbers.

2:40Save this allocation comparison with a descriptive name. Its inputs and results stay together in Workspace.

2:47Duplicate the diversified mix to test when it trades. Keep its allocation and historical window fixed.

2:54Set the copy to Deviation with a five percent threshold. This isolates when it restores the target weights.

3:02The rule waits until a position drifts five percentage points from target. Monthly rebalancing instead follows the calendar.

3:11Run all three portfolios. The two gold mixes now differ only in their rebalancing rule.

3:17The drift rule ends near sixteen thousand seven hundred dollars, about four hundred dollars above monthly. Its worst drawdown is twenty seven point two percent, versus twenty eight point three percent monthly. With other inputs fixed, the rule explains this difference in the model.

3:34Rebalancing analysis shows what each rule traded. The counts help explain the practical difference between them.

3:42Turnover measures how much portfolio value traded. Monthly rebalancing produced seventy one events and eighteen point one percent average annual turnover.

3:53The five percent drift rule produced twelve events and ten point seven percent average annual turnover.

4:00Drift traded less and finished slightly higher in this window. Another market period can change that comparison.

4:09The methodology explains the limits of this comparison.

4:14These are nominal total returns before taxes. Transaction costs and market impact are not modeled. Those costs matter when one rule trades more often. A different window can also change the ranking.

4:27Take a tour identifies the main controls. It is available whenever you return to this tool.

4:33Docs explains each setting and the calculations behind these results.

4:40Now compare both mixes over the pandemic crash. Check whether ending value, drawdown, and recovery tell the same story.

How-to

Editing a portfolio

  • Add or remove allocations at the portfolio level.
  • Edit each allocation's weight and linked strategy.
  • Enable Strategy rules, then open the strategy editor to add positions, signals, conditions, and strategy-level rebalancing.
  • Load a Library strategy or Library portfolio when you need to reuse an existing building block.

See Strategy Leaderboard and Signals for the reusable rule and indicator model used inside the editor.

Cashflows

  1. Open the portfolio you want to change and choose Add cash flow.
  2. Choose Periodic for recurring cash flows or One-time for a dated event.
  3. Enter the amount and timing. For a monthly dollar contribution, choose $ Dollar and Monthly.
  4. Run the backtest again. Inspect the cash flow events alongside the value chart.

Contributions and withdrawals are available in ordinary portfolio editing. They do not require Strategy rules.

  • Recurring cash flows: scheduled contributions or withdrawals.
  • One-time cash flows: dated events applied once inside the backtest window.
  • Sign convention: positive values contribute capital and negative values withdraw capital.

The cashflows tab in the results section shows each contribution and withdrawal event with dates, amounts, and a running total chart.

Compare a portfolio against a benchmark

  1. Configure your portfolio with the target allocations and strategies.
  2. In the benchmark section, enter a ticker such as SPY, VTI, or a custom blend.
  3. Run the backtest. The equity curve, drawdown chart, and metric table include both the portfolio and benchmark series.
  4. Review benchmark-relative metrics: excess return, tracking error, information ratio, and capture ratios appear in the metric table.

Sign in to save an analysis. If you edit inputs after a completed run, the visible result still belongs to the last completed scenario. Run again to update the evidence before comparing it with another analysis.

Save the current analysis when you need to reopen the same configuration with a result from the current engine. ArthaPilot withholds older output until you refresh or rerun the saved analysis. Create a share link when you want to send the backtest setup to someone else.

Reference

Configuration guide

Treat these fields as modeling assumptions. Change one at a time when you want to understand what drives a result.

ConfigurationWhat It MeansWhy It Matters
Portfolio allocationsThe target asset mix or strategy definition that the backtest replays.This is the portfolio hypothesis. Changing weights changes both expected return exposure and drawdown behavior.
Date rangeThe historical window used for prices, returns, and events.A short window can make a strategy look better or worse because it samples only one market regime.
Price ModeTotal return includes dividends and distributions. Raw prices keep them separate for dividend accounting.Use Total return for return comparisons. Use Raw only when you need explicit dividend tracking.
Initial value and cashflowsStarting capital plus any recurring or one-time additions and withdrawals.Cashflows make results path-dependent because deposits and withdrawals happen at specific market levels.
RebalancingThe rule that decides when the backtest trades holdings back toward target weights.More frequent rebalancing can reduce drift, but it can increase turnover, taxes, and implementation cost.
BenchmarkThe reference series used for comparison metrics.Benchmark choice changes relative metrics such as alpha, beta, tracking error, and up/down capture.
Tax-aware modeAdds account type, filing status, income, state, lot selection, and loss assumptions.After-tax results are not comparable to pre-tax results unless the tax profile is explicit.

Main inputs

  • Starting principal: initial portfolio value.
  • Date range: historical sample used in the run.
  • Price mode: total-return or raw-price return construction.
  • Allocations and strategies: the holdings and rules under test.
  • Portfolio rebalancing: how the engine maintains the top-level allocation split.
  • Cashflows: recurring or one-time contributions and withdrawals.
  • In-sample / out-of-sample holdout: optional split date used to show separate metrics for the in-sample window and the later out-of-sample window of the same realized run.

Rebalancing

Portfolio Backtest has two separate rebalance layers:

  • Strategy-level rebalancing: restores weights within the active positions of a strategy.
  • Portfolio-level rebalancing: restores the top-level split between allocations.

Both layers can use time-based or threshold-based policies. See the Rebalancing guide for the detailed mode descriptions.

Glidepath portfolios

A glidepath portfolio changes its target allocation over time. This models the common practice of shifting from growth-oriented to income-oriented holdings as the investment horizon shortens (for example, a target-date retirement fund).

  • Static vs glidepath: toggle glidepath mode under Strategy rules when a portfolio has two or more allocations.
  • Milestones: each milestone specifies a date and the target weight for every allocation. The engine uses stepwise resolution: the most recent milestone on or before the current date determines the target weights.
  • Dates must be strictly increasing. The milestone weights for each date must sum to 100%.
  • Initial allocation rule: before the first milestone, the backtest uses the initial sleeve weights. A milestone on or before the start date overrides those weights.

Example: a portfolio with two allocations (SPY, AGG) and two milestones shifts from 80/20 on 2010-01-01 to 40/60 on 2020-01-01. In a 2005-2025 backtest the initial weights would be 80/20 (first milestone applies from 2010), then 40/60 from 2020 onward.

In-sample / out-of-sample holdout

Portfolio Backtest supports an in-sample / out-of-sample holdout overlay. Choose an in-sample end date and run the backtest. The results include separate metric tables for the in-sample window and the later out-of-sample window for each portfolio.

  • The split uses the realized backtest path. It does not change trades, cashflows, tax-lot events, or rebalancing decisions.
  • The requested date snaps to the nearest prior trading day. Both sides need enough history. Otherwise, the result shows a warning instead of partial metrics.
  • Metric tables use the same metric catalog as the full-period result. Tax-aware and dividend-income metrics are not replayed per slice and are shown as unavailable with a warning.

Tax diagnostics

When you enable tax-aware mode, the tax diagnostics tab becomes the default tax surface. It summarizes total drag, current-year tax drivers, estimated embedded liquidation drag, high-impact trades, concentrated unrealized gains, and assumptions or confidence flags.

This is the fastest way to answer why a strategy is tax-inefficient before drilling into raw trade logs, lot rows, or wash-sale chains.

Dividend income

The dividend income tab shows annual dividend income as a bar chart. It also shows per-portfolio summary metrics: total income, qualified vs ordinary breakdown, and payout count. A sortable table lists individual dividend events (ex-date, ticker, shares, per-share amount, income, qualified/ordinary type).

Dividend tracking requires the raw price mode. Total return (adjusted prices) already bakes dividends into the price series, so using it with dividend tracking would double-count income.

The annual breakdown classifies each dividend as qualified or ordinary. Qualified dividends fall under the preferential long-term capital gains rates, while ordinary dividends fall under regular income rates.

Rolling metrics

The rolling metrics tab shows windowed calculations for CAGR, volatility, Sharpe, Sortino, and other risk metrics over time. This reveals whether full-period averages mask regime changes. A portfolio with a stable 10% CAGR may have experienced stretches of 20% gains and 5% losses.

The window length is configurable. Shorter windows capture more variation but have more noise. Longer windows smooth the series but can obscure transitions.

Risk vs return

The risk vs return tab plots a scatter chart of annualized return against annualized volatility for each portfolio. Point annotations show Sharpe ratio and max drawdown for comparing risk-adjusted performance across portfolios.

Seasonality

The seasonality tab shows average monthly return by calendar month. This can help identify seasonal patterns in portfolio returns. Sample sizes per month are small, so patterns may not be statistically significant over shorter backtest windows.

Correlations

The correlations tab displays a pairwise daily return correlation heatmap. This requires two or more portfolios or a benchmark, and it uses daily returns over the backtest window.

Relative analysis

The relative analysis tab shows the performance ratio of each portfolio relative to a baseline (portfolio value divided by baseline value over time). Relative metrics include tracking error, information ratio, and up/down capture ratios. This requires two or more portfolios.

Lump sum vs DCA

The lump sum vs DCA tab runs a separate dollar-cost averaging simulation using the same portfolio configuration. You can configure the DCA frequency (weekly or monthly), deployment duration, and cash yield during the deployment phase.

This compares full immediate deployment against phased entry. The result is path-dependent on the specific historical window. Lump sum wins on average but DCA reduces timing risk.

Allocations

The allocations tab shows a pie chart grid of target allocation weights for each portfolio, derived from the strategy positions and allocation weights you configured.

Tax-aware engine

Tax-aware engine runs the same portfolio settings through a lot-level tax engine. The tax profile controls filing status, income, state, account type, lot selection, and optional loss-harvesting settings.

See Tax-Aware Backtesting for the accounting assumptions and output interpretation.

Ticker modifiers

Positions can use ticker modifiers when the backtest should use a transformed series instead of the base ticker.

Regression analysis

The regression tab runs an OLS regression of each non-benchmark portfolio's returns against the benchmark's returns. The diagnostics show how much benchmark exposure accounts for a portfolio's behavior.

  • Alpha: the intercept of the OLS regression, representing return generated beyond what beta exposure explains. Reported as both a per-period value and an annualized value.
  • Beta: the slope coefficient that measures sensitivity to the benchmark. A beta of 1.0 means the portfolio moves with the benchmark. Higher values indicate more sensitivity.
  • R-squared: the fraction of portfolio return variance explained by the benchmark. Values closer to 1.0 indicate that the benchmark is a strong explanatory factor.
  • Correlation: the Pearson correlation coefficient between portfolio and benchmark returns.
  • Tracking error: the annualized standard deviation of the return difference between the portfolio and the benchmark. Lower values mean the portfolio tracks the benchmark more closely.
  • Information ratio: the annualized excess return divided by the tracking error. A higher value indicates better risk-adjusted outperformance.

Select daily or monthly frequency to change the regression interval. Monthly aggregation compounds daily returns geometrically. The scatter plot shows the fitted line. The residual chart shows deviations.

Features

  • Glidepath portfolios with time-based allocation shifts
  • Dividend tracking with configurable reinvestment
  • Tax-lot accounting with FIFO, LIFO, HIFO, or optimized selection
  • Cash flow schedules (recurring and one-time contributions and withdrawals)
  • Regression analysis against any benchmark (alpha, beta, R², tracking error)
  • Tax diagnostics for tax-aware runs, including drag drivers, high-impact events, and implementation risks
  • Side-by-side comparison of multiple portfolio configurations
  • In-sample / out-of-sample holdout split with per-slice metric tables

Glossary

Glossary

CAGR
Compound annual growth rate. The constant yearly return that would produce the same ending value as the actual path.
Volatility
Annualized standard deviation of periodic returns. A measure of how much returns swing around their average.
Drawdown
A peak-to-trough percentage decline. Max drawdown is the worst observed.
Sharpe ratio
Excess return divided by total volatility. Reward per unit of swing. Does not distinguish upside from downside.
Sortino ratio
Like Sharpe, but the denominator is downside volatility only. Penalizes losses but not gains.
Tracking error
Annualized standard deviation of the return difference between the portfolio and its benchmark.
In-sample vs out-of-sample
In-sample is the period that you used to tune a strategy. Out-of-sample is later, unseen data. Strong in-sample-only results often do not generalize.
Glidepath
A schedule changes target weights over time. For example, it can shift from stocks to bonds as a target date approaches.

Explanation

When to use it

Use this tool when the question is about how a portfolio configuration would have behaved over a historical window.

Portfolio Backtest evaluates the portfolios you specify. Use Portfolio Optimizer to search across portfolio weights.

For distributional outcomes and failure-rate analysis across many simulated paths, use Monte Carlo instead. See Historical Backtest vs Monte Carlo below for a comparison.

Portfolio model

LevelWhat It StoresWhere It Rebalances
PortfolioOne or more allocations, portfolio-level cash flows, and portfolio-level rebalancing settings.Between allocations when portfolio rebalancing is on.
AllocationOne strategy and its target portfolio weight.Inside the strategy when the strategy itself holds multiple positions.
StrategyConditions, positions, signals, and strategy-level rebalance rules.Within the active condition.

Historical backtest vs Monte Carlo

DimensionPortfolio BacktestMonte Carlo
InputOne historical return pathMany simulated return paths
OutputSingle time series, metrics, and chartsDistributional outcomes, percentile bands, and success rates
Best forExact historical behavior of a specific configurationRange of possible outcomes and failure-rate analysis
LimitationOne path is one sample. Different date ranges can give different conclusions.Simulated paths depend on the return model and its assumptions.

Use Portfolio Backtest when you need detailed historical attribution, tax-aware accounting, or strategy-level diagnostics. Use Monte Carlo when you need to evaluate withdrawal sustainability, outcome distributions, or tail risk across many scenarios.

Common pitfalls

  1. Picking too short a date range

    A short sample can miss the conditions you want to inspect. Include relevant drawdown periods when comparing historical behavior.

  2. Mixing Raw prices with dividend tracking off

    Raw keeps dividends and distributions separate, so returns read low. Choose Total return unless you specifically want to inspect dividends.

  3. Comparing a pre-tax run to an after-tax run

    Tax-aware mode adds drag and lot-level events. Either run both portfolios with tax-aware on, or both with it off.

  4. Treating one path as predictive

    A single historical replay is one sample. For probability ranges, run the same plan in Monte Carlo.