Methodology

How Steady Otter compares ETF mixes

Steady Otter looks for a mix of exchange-traded funds (ETFs) that grew more strongly in the available market history, after estimated trading costs, while staying within limits on losses, price swings and how much money sits in each fund. An ETF is a fund traded on an exchange that holds a collection of investments.

It compares possible mixes with a starting mix and three fixed comparison portfolios: US stocks, worldwide stocks, and a mix of 60% US stocks and 40% US bonds. These are historical comparisons, not predictions.

1. Compare on the same terms

Each mix uses the same actual trading days available for every required fund and comparison portfolio, capped at the latest 2,520 returns. Younger funds can shorten this window. Missing funds are not silently removed or replaced, and weights are not rescaled. The cap is a practical policy, not a proven best amount of history. The calculation starts with money invested once, adjusts the mix back to its chosen percentages each month, and subtracts the configured trading costs. Base, stress, modeled and funded fees currently default to zero, so base and stress coincide; they do not test higher fees. Taxes, added savings and inflation are not included in the ordering of results.

Standard analysis needs at least 1,262 training days; explicitly accepted limited history needs 756. A later test uses the longest permitted period that leaves enough training plus a separate entry day. Otherwise the result says that no separate later test is available. Retry and limited rerun keep the original capture's dates and fund facts.

2. Look for stronger growth within limits

The program tries many possible mixes, though it does not try every possible combination. It checks losses from previous highs, day-to-day price swings, the amount of trading needed, and whether the funds trade in sufficient volume. A high past return alone is not enough to pass.

To earn the label Qualified, an alternative must pass those checks and grow at least half a percentage point more per year than every fixed comparison portfolio in each of three historical periods, at both trading-cost levels. For example, if a comparison grew by 8% a year, the alternative would need at least 8.5%. This rule does not require it to beat the starting mix.

A further check rearranges stretches of the recorded returns many times to see whether the advantage depends heavily on a few favorable stretches. Passing gives more support within that history; it does not tell us the chance of future success.

3. Check whether the search holds up

Repeat searches check whether different random choices or a longer search produce very different results. When the history checks pass, the program also runs five checks in time order. At each step, it reruns the main search using only earlier data. This is called a refit. It chooses a mix that meets the limits, fixes its weights, and tests it over the next period. Each step can choose a different mix.

A search can look strong because it found a mix suited to a lucky stretch of the past. These checks ask whether the same search rules hold up on data they have not yet seen. They compare later growth with the fixed comparison portfolios and check losses, price swings and trading levels. They test the search process, rather than every final mix, and cannot promise future returns.

Each search prefers qualified alternatives but can use an unqualified mix or the starting mix if it meets the limits. Its original qualification label is kept. The search leaves out the return on the day the mix enters. The new mix enters at that day's closing prices and earns returns from the next day.

The five checks also form one continuous investment history. Holdings carry from one period to the next, their values change with the market, and switching mixes includes the configured trading costs. This shows how repeated searches would have worked together over time. Starting fresh in every period could hide the cost of changing mixes.

Separate fresh-start checks are also shown, with the same entry costs for the selected mix and comparison portfolios. They do not decide whether the combined check passes. All five periods need a mix that meets the limits; otherwise the combined check is unavailable, while completed individual checks remain visible. Missing evidence cannot count as a pass.

Historical mixes, qualification labels and result order are fixed before these five checks or any reserved later period are evaluated. Failed or missing later evidence is shown alongside that comparison and cannot change its order. Projections and current disclosed holdings also cannot change it.

If required training stability checks fail or are missing, results may be marked Exploratory and left unordered. If no alternative passes all qualification checks, the page can show No qualified winner. There is no promise that a search will find a stronger mix.

4. Read the results as a comparison

Mixes that meet the limits and qualify come first. Other mixes that meet the limits follow, including the starting mix if it meets them. By default, recorded annual growth after costs is balanced against how much of the allowed historical risk budget the mix uses. Saving one percentage point of that risk budget offsets 0.5 percentage points of annual growth. For example, risk-budget usage falling from 80% to 79% offsets growth falling from 12% to 11.5%. For equal trade-off scores, lower risk comes first, followed by smaller losses and swings, less concentration, and less trading.

The results show growth, losses, risk and costs alongside the explanatory scores. Read them together; do not average the scores or treat them as a probability. Annual growth includes growth on earlier gains: $100 growing by 10% becomes $110 after one year and $121 after two. Actual yearly returns can vary widely. AI text explains the calculated results and cannot change the numbers or their order.

What the scores explain

These scores explain different parts of the comparison. Raw risk-budget usage also enters the default growth/risk ranking; its displayed log scale does not. The other scores do not decide the order, and none replaces the qualification checks. The full paper also discusses proposed ranking and stability measures; those proposals are not part of the current method.

What this cannot tell you

Future markets may behave differently. Losses can exceed those in the recorded history, and different ETFs may own many of the same investments. Results depend on the funds included, the available data, and the assumptions used. A selected mix is a focus for comparison, not an instruction to buy it or personal investment advice.

ETF HHI describes how weight is divided among funds. A separate disclosed-position HHI (partial estimate) combines the eligible disclosed holdings using the original portfolio percentages. Read it with coverage, undisclosed mass and the largest positions; the unknown tail is not rescaled. No usable holdings means unavailable, never zero. Symbols are approximate and can include funds or cash; separate share classes are not merged. Holdings dates can be unknown, and these snapshots do not show historical holdings or prove greater diversification.

Modeled paths reuse complete common months from the capped history after ranking, including any later-test months. Their ten-year horizon is separate from the available history. Percentiles describe those conditional paths, not independent validation or uncertainty in every model assumption. Repeated tuning against later tests uses up their independence. AI is attempted for each successful analysis; if unavailable, deterministic results remain complete.

Historical and contribution-free modeled costs use portfolio value times the rate times one-way turnover. Funded projections use the rate times combined purchase and sale dollars. These existing conventions differ when fees are positive; taxes and cash flows also apply only to funded projections.

Read the full version for the exact rules, formulas, score definitions, and clearly labeled proposals for future work.