Drawdown
The fall from an earlier peak. If a mix reaches $120 and then drops to $90, its drawdown is:
Maximum drawdown is the deepest observed fall in the measured period. Future losses can be larger.
Percentages, multiplication, and a few simple comparisons explain the core ideas. Here is what the numbers mean—and what they cannot tell you.
An ETF is a fund that trades like a stock and can hold many investments. A portfolio combines ETFs in chosen percentages.
If a hypothetical mix is 60% Fund A and 40% Fund B, a simple one-period return is:
is a return. Its subscript identifies the mix, Fund A, or Fund B.
If A rises 10% and B rises 2%, the mix gains 6.8% before costs. Over several periods, weights can drift; rebalancing and costs affect the actual path.
Weights add up to 100%.
Different tickers do not always mean different risks. Funds can hold overlapping investments.
Each period builds on the value left by the previous one. CAGR is the steady annual rate that connects the starting and ending values.
is portfolio value: subscript 0 marks the start, and marks the end. is years; is the annual return as a decimal (10% = 0.10).
A 10% gain followed by a 10% loss turns $100 into $110, then $99. The two returns average to 0%, but the money falls 1%.
is CAGR as a decimal: 0.10 means 10% per year.
For $100 growing to $121 in two years, CAGR is 10%. “Net” means the metric includes the costs specified by that calculation.
Constant-rate arithmetic only. No contributions, fees, taxes, inflation, or market swings. This is not a product projection or forecast.
The fall from an earlier peak. If a mix reaches $120 and then drops to $90, its drawdown is:
Maximum drawdown is the deepest observed fall in the measured period. Future losses can be larger.
How spread out the returns are. A mix with big up-and-down moves usually has higher volatility than one with smaller moves.
It describes variability, including upward moves. It is not a complete measure of the chance of losing money.
Square each ETF weight and add the squares. Two equally weighted ETFs give:
The reciprocal, 1 ÷ 0.50 = 2, is the effective ETF count. This does not account for overlap inside the funds.
A high-ranked mix does not automatically pass every test.
The search checks return hurdles against fixed references, downside and volatility limits, turnover, liquidity, and resampling stability. A candidate can show more growth and still fail a check.
Risk limits also depend on the seed mix. These are model policies for hypothetical scenarios, not a judgment about your circumstances.
Qualified mixes come before feasible unqualified mixes. The current default ranks within each feasible group using:
is the score. is net CAGR as a decimal; multiplying by 100 gives percentage points. is risk-limit usage as a percentage: 80% usage means .
Risk-limit usage is the largest drawdown or volatility fraction of its allowed limit across the evaluated scopes and costs, expressed as a percentage.
For example, 12% CAGR with 80% usage scores −28; 9% CAGR with 70% usage scores −26. The second ranks higher within the same group.
This is a chosen scoring preference, not expected return or a probability. Effective settings are recorded per run; alternate configurations can use a different ordering.
The sequence matters: search, freeze, examine, simulate.
Fit on earlier data and evaluate in the next period. When available, a separately reserved later test examines frozen portfolios without changing their ranks.
Reuse chunks of observed history in different sequences. Compare mixes on the same sampled sequences so the comparison is fairer.
The median is the middle outcome. The 5th and 95th percentiles mark positions in the modeled results; they are not calibrated probabilities of what will happen next.
Simulations reuse the available past. They may miss events that did not occur in that history. More paths do not fix missing history, unrealistic assumptions, or model errors.
Overfitting means finding a mix that matches the luck of the past rather than a pattern that lasts. Searching many mixes makes this easier, so a high historical return alone is not enough.
When history checks pass, the program reruns the main search five times, each using only earlier data. Each search chooses a mix that meets the limits, then tests it over the next period. A different mix may be chosen each time. This helps check whether the search rules hold up beyond the data used to choose each mix, rather than only finding mixes suited to a lucky stretch of history. It tests the search process, not every final portfolio, and cannot promise future returns.
When enough history exists, a later period is set aside before any search. Final allocations and ranks are frozen before that period is evaluated. Its results cannot change qualification or rank. Shorter histories explicitly report no separate later test.
Candidate qualification requires return hurdles across three historical blocks and fixed reference mixes under configured base and stress costs, plus risk limits. Both fees currently default to zero, so they provide no distinct fee stress. Block resampling uses the same samples for each comparison to assess the advantage within that history.
When configured, repeated searches with different random seeds and a larger search budget check sensitivity to the search itself. Omitted checks are shown as not evaluated. These checks assess stability; they do not create independent market evidence.
These safeguards reduce overfitting risk; they cannot eliminate it. Resampling and simulations reuse past data, and trying more mixes can still find lucky results. Repeatedly changing inputs or choosing a mix after seeing the later test uses up its independence. Limited history, the current ETF universe, and future market changes remain limitations.
Bias can enter through which funds we include, what the search sees, and how we compare results. We apply consistent rules at each stage and disclose the limits of the evidence.
Smart pool, evolutionary search, CMA-ES, island searches, and blending explore allocations in different ways. Every evaluated candidate faces the same constraints and qualification checks. Identical allocations share one identity and cached evaluation, so discovering a mix twice does not create two pieces of evidence. Fixed budgets and reproducible random seeds make the work traceable.
Each chronological refit receives only earlier returns, prices, volume and features, excluding its entry row. Every required ETF and reference uses common actual dates; missing funds are not removed, replaced or rescaled. A reserved later test stays outside search and qualification. Matched dates, costs and sampled sequences limit unequal comparisons and look-ahead bias.
Fixed references anchor the return checks. Qualification and historical ranking have distinct recorded rules. Outer and holdout results, projections, partial disclosed-position facts and AI cannot change frozen ranks. Required training stability failures remain exploratory; missing or N/A checks cannot establish a pass.
These controls do not make the analysis unbiased. The selected ETF universe and available provider history can favor funds that survived or became popular. The chosen dates, benchmarks, thresholds, and scoring preferences shape the answer. Repeatedly tuning those choices against observed results adds selection bias, and historical tests cannot establish future performance.
Rigor means repeatable calculations, consistent comparisons, and checks that can fail. These figures describe the current methodology defaults; effective settings and completed work are recorded for each run.
With 10 ETFs, weights in 1% steps can form about 4.26 trillion mixes. This example allows zero weights and no extra limits on individual ETFs.
Weights must add up to 100%. Risk and allocation limits rule out many mixes. The optimizer searches a limited subset.
The default runs one main search, two repeats with different random seeds, and one expanded search with twice the search budget.
The five earlier-data searches are called refits. Each reruns the main search with the same search budget, using only the history available before its next test period. All five are attempted when history checks pass, even if the full searches find no qualified alternative. The actual search settings are recorded for each run.
Proposals are filtered before full testing, and repeat mixes use up attempts. Search attempts are not counts of unique tested mixes.
Each candidate faces 18 return comparisons: three historical periods × three reference mixes × two trading-cost assumptions.
Ten drawdown checks and ten volatility checks cover those periods, their combined path, and the full training period under both costs. Trading turnover and liquidity are also checked.
The test reshuffles chunks of history with average lengths of 10, 21, and 63 trading days. Each length produces 1,024 comparisons.
At each length, even the 103rd-smallest worst-case return advantage must stay above zero. This tests weaker outcomes; it is not a 90% chance of future success or an adjustment for every mix tried.
The default projection builds 5,000 possible ten-year paths. Each joins 40 sampled three-month blocks: 600,000 monthly steps per mix.
A separate check uses recent history when enough distinct months are available.
These paths reuse complete common months in the capped window after rank freeze, including any reserved later-test observations. Their ten-year horizon is separate from historical coverage. Percentiles are conditional scenarios, not independent validation, parameter uncertainty or observations of the future.
This net CAGR formula combines ETF weights, daily returns, trading costs, and compounding to measure annual growth.
is the number of ETFs; is the number of scored trading days. is ETF ’s weight before day ’s return, after any rebalance; is its daily adjusted return. is the configured one-way cost rate and is half-L1 turnover including cash. Entry turnover is 1. Weights drift between rebalances, turnover is zero without trades, and 252 is the annual trading-day convention.
The inner sum combines each scored day's ETF returns; the cost factor reduces pretrade wealth before that return. At a refit, the fit excludes the entry row, old holdings earn through its close, and transition cost belongs once to the incoming first scored return. The log-and-exp calculation annualizes compounded growth. Contributions and taxes belong to separate funded projections. More attempts, checks or paths do not guarantee better future results.
Starting values and monthly contributions affect a funded projection. They do not change the contribution-free growth calculation used to compare return paths or determine rank.
The model also applies its stated transaction costs and simplified tax assumptions. These do not account for every tax rule or individual circumstance.
Use hypothetical values. Personal financial and tax information is not required or collected.