How Steady Otter searches and compares ETF portfolios
Steady Otter looks for ETF allocations with stronger historical compound growth, after trading costs, within stated risk and allocation limits. It compares the alternatives with the starting portfolio and fixed benchmarks. Historical results are evidence for comparison, not a promise of future returns.
The pipeline starts with the ETF universe—the funds the search may use—and a starting portfolio. Steady Otter checks the available price, return and trading-volume history, aligns usable data with benchmarks on common trading days, and reserves a later test when enough history is available. It searches different ETF combinations and weights within the allocation rules, using only the training history and simulating monthly rebalancing under configured base and stress trading costs. It checks growth, drawdowns, volatility, turnover and liquidity; alternatives must also pass benchmark-return checks across development periods and a resampling check to qualify.
Repeat searches, a larger search budget and tests that search earlier history and evaluate the next period provide separate evidence about the search process. Training stability determines whether comparisons are ranked or exploratory; a comparison can still appear when no qualified alternative is found. The pipeline removes duplicate mixes and builds a short list alongside the starting portfolio and benchmarks. Ranked results put feasible qualified alternatives before feasible unqualified mixes; the default order within each group balances historical net growth against risk-limit usage. Any reserved later test evaluates frozen allocations. Hypothetical projections and an AI explanation, when available, complete the report. Outer tests, the reserved later test, projections, disclosed-position facts and AI cannot change the frozen historical ranking.
Reading guide. A seed is the starting portfolio; a candidate is a proposed alternative. A reference is a fixed benchmark. A session is one trading day. Development data helps find portfolios; a later test checks frozen portfolios on reserved data that did not influence the search.
Formulas support the explanations. A sum sign () means “add these values”; vertical bars mean “take the size of a difference, ignoring its sign.” Returns, risk measures and allocation weights use fractions inside formulas: 0.01 means 1%. A percentage point is the difference between percentages: 8% versus 7.5% is 0.5 points. One basis point is 0.01 points.
Method: alpha-p0-v1.2; algorithm: alpha-p0-v1.2-robust-selection-v3; result format: scenario-ranking-v2.6.
The data and portfolio rules
Every portfolio uses the same actual trading days and execution assumptions. Mixed-age ETFs start at the latest required fund/reference history; unavailable symbols block the run rather than being removed or replaced. The analyzed view uses at most the latest 2,520 common returns.1 The search sees only training; a later test is reserved before search when history permits.
Exact data, history and allocation requirements
Element
Contract
Capture
Calendar-derived 2,521 adjusted closes, including warmup; immutable returns, unadjusted close and volume; cutoff is last New York Stock Exchange session before acquisition date in New York
Input validation
Finite returns >−100%; close >0; volume ≥0; unique, ordered, gap-free common sessions; valid allocation lattice
Candidates
Long-only; total 100%; 1% ticks; ETF weights 0 or 1–100%; optional tighter bounds
One cash establishment; monthly rebalance; continuous state across development blocks; costs before returns
Analyzed history
Latest ≤2,520 common actual returns; source coverage/dates retained separately; no proxies, exclusions or renormalization
History
Standard ≥1,262 final training; accepted limited ≥756. Earliest lagged refit fits: 630 / 377; inner blocks: 126 / 75
Later-test reservation
Standard: longest of 252 or 126, needing totals 1,515 / 1,389. Limited: 126, needing 883. Totals include one execution row; otherwise valid history remains all-training
Development
training sessions: risk context, then three consecutive, nearly equal blocks. Full-training risk includes context; search/refits/liquidity use own prefixes
Standard 1,260–1,261 and limited 755 fail; training floors are structural, not guarantees of statistical adequacy. The explicit mode is not inferred from history length. Retry and limited rerun reuse the source capture's dates and frozen fund facts after request, ownership and integrity checks. Acquisition coverage, warmup, capped analysis, fit, execution and scored periods are distinct. Full capped-view finite-domain preflight still rejects invalid or overflowing projection inputs, including holdout rows; it is not a performance-selection gate.
What the main measures mean
Growth
CAGR (compound annual growth rate) is the constant annual growth rate that would produce the same final value as the recorded returns. It includes compounding and modeled trading costs. For example, a 10% CAGR means the same final value as growing by 10% each year; actual annual returns can vary considerably.
The calculation assumes 252 trading sessions per year. For daily net returns , annualized compound growth is:
Risk and concentration
Wealth starts at 1. Let be the highest wealth reached through session . Maximum drawdown is the largest loss from that peak:
Volatility measures how widely daily returns vary; higher values mean larger swings. HHI measures how concentrated the ETF weights are. Higher HHI means more weight sits in fewer ETFs; it does not measure overlap in their holdings.
Annualized volatility uses the sample standard deviation of daily returns. Allocation concentration is the sum of squared ETF weights :
A separate disclosed-position HHI (partial estimate) uses current eligible top-holdings rows after selection. If fund discloses weight in approximate symbol , preserve original portfolio weights:
The report shows whole-mix largest ten positions and their combined exposure; it does not rescale the unknown tail. Empty/unusable rows mean unavailable/null HHI. This is separate from ETF HHI and never affects selection, ranking or projections. Lower incomplete values do not establish better diversification.
The capture freezes validated Yahoo fund facts with source, retrieval date, available holdings date and cache status; unknown holdings dates remain unknown. Eligibility uses the acquisition-anchor retrieval clock: older than 30 days or more than five minutes ahead is ineligible. This does not verify that the provider's holdings themselves are current. Malformed facts are also withheld. Optional provider failure is nonfatal; verified snapshot corruption fails integrity checks. Initial/manual AI, retries and local/Batch handoff reuse the same facts; AI failure leaves deterministic concentration available. Symbols can denote funds or cash, share classes are not consolidated, and historical holdings/recursive issuer resolution are unavailable. Private per-fund provenance is retained; public analytics apply to the complete mix only.
Trading costs and turnover
Turnover measures how much the allocation changes when rebalancing. More turnover usually means more trading cost. Moving 10% from one asset to another produces 10% turnover, since that same amount is sold and bought.
Let and be target and current weights across all assets and cash. Turnover is half their total absolute difference. Trading cost uses wealth immediately before the trade:
The configured base, stress and modeled one-way cost rates currently default to zero, as does the funded-projection rate. Effective rates are recorded per run; modeled and funded settings agree. Equal zero base/stress cases provide no distinct fee stress. Establishment turnover is 1. Actual/modeled replay uses the wealth × rate × half-L1 convention above. The existing funded engine instead charges rate × (executed purchase dollars + sale dollars); these conventions differ at positive rates and are not normalized by this update.
Let be the highest rolling 252-session recurring turnover under cost case . The turnover gate uses the worst cost case:
Recurring turnover excludes establishment. “Net” deducts modeled trading costs; taxes, contributions and inflation are excluded from ranking. ETF expense drag already present in adjusted returns is retained. below is the raw base-cost CAGR of the linked development path.
How the search finds alternatives
The optimizer tries several search methods, checks whether each proposed allocation is allowed, and evaluates it on the same historical data. It does not test every possible mix. Random choices are repeatable when the captured data, settings and random seeds are unchanged.
Search methods and reproducibility
Route
Mechanism
Pool
Dirichlet weights, concentration 0.35 → legal ticks; score first unique valid proposals up to preselection limit
An alternative must respect risk, trading and liquidity limits. To qualify, it must also beat each fixed benchmark by at least 0.5 percentage points a year in each development period, under both cost assumptions, and pass a resampling check of that advantage.
Check
What passing means
Return
Beats all three references by at least 0.5 points/year in all three development blocks, at both cost rates
Drawdown
Within the smaller of 40% or the matched seed's drawdown plus 5 points
Volatility
Within the smaller of 25% or 110% of the matched seed's volatility
Turnover
Recurring turnover stays at or below 100% over a rolling 252-session period
Liquidity
Each held ETF has at least 5 million dollars in median daily dollar trading volume over the latest 60 training sessions at the search endpoint
Feasible means the allocation passes all limits except the return hurdle. Point-pass means it passes all five checks on observed returns. Qualified means a non-seed alternative also passes the resampling check. The seed can be feasible but is not labeled qualified. Qualification compares returns with the fixed references; there is no general requirement to beat seed CAGR.
The bootstrap check repeatedly samples stretches of recorded returns. Candidate and reference use the same sampled days. This asks whether an apparent advantage persists when favorable and unfavorable stretches occur in different combinations. It does not predict the probability of future success.
Exact qualification and search-fitness rules
= candidate − reference CAGR; = 0.5 pp/year; = active ETF's trailing-60 median dollar volume (close × volume). Same-scope seed caps:
= maximum member violation; passing requires :
Family
Scope
Member violation
Return
3 blocks × 3 references × 2 costs
Drawdown
Each block, linked development, full training × both costs
Volatility
Same risk scopes
Turnover
Both costs
Liquidity
Every active ETF
Legal weights required. Point-pass: all five . Feasible: all except return . Qualified: non-seed point-passer + bootstrap pass. Seed: qualification inapplicable, feasibility applies. No general seed-CAGR hurdle.
Paired circular stationary bootstrap: geometric-length blocks of computed daily candidate/reference net-return paths; shared indices (Politis–Romano).
The expected block lengths are 10, 21 and 63 sessions. The order statistic is one-based, without interpolation. Bootstrap passes if and only if .
: CAGR of resampled net returns, both paths. replicates; default 1,024 → 103rd-smallest statistic. Authorized sample-count/search-budget changes: effective run settings disclosed.
Fitness: first differing tuple field decides; smaller wins. No scalar 0–100 “Fit score.”
Feasible qualified portfolios come first, followed by feasible unqualified portfolios, including the seed when it is feasible. By default, each group ranks by CAGR/risk preference points: historical net CAGR in percentage points minus 0.5 times raw risk-limit usage. A reduction from 80% to 79% risk-limit usage offsets 0.5 percentage points/year less CAGR. The risk input is the worst drawdown/limit or volatility/limit across all historical risk scopes and both cost cases, using paired seed-relative and absolute limits.
This exchange rate expresses a policy preference, not an expected return or a validated optimum. Negative scores are normal: 12% CAGR with 80% usage scores -28; 9% CAGR with 70% usage scores -26. The logarithmic displayed risk score is not used. Growth bands remain descriptive; they do not cap the growth sacrifice.
The result retains at most 11 distinct mixes, including the seed and protected group leaders and contenders. The selected mix is a comparison focus for analysis, not an instruction to buy it. With no feasible mix, there is no selection.
Exact final ranking and retention rules
Rankable output: deduplicate evaluated searches + seed; first-source qualification retained, primary → expanded → repeats. Groups: feasible qualified ; feasible unqualified/seed . Each group's growth leader: highest . Default scoring is enabled by the server-owned JSON switch activate_cagr_bands_risk_scoring.
denotes tuple concatenation: prefix the complete search-fitness tuple with 2 for an infeasible non-seed.
: worst raw risk-limit usage over three development blocks, linked path, full-recent path and both costs. Zero risk against a zero cap contributes zero; positive risk against a zero cap is infeasible.
: preference points, not CAGR. Use raw precision; exact score ties prefer lower , then the remaining disclosed keys. Net CAGR already includes costs and compounding; the risk penalty is an additional preference.
: worst block values across costs; : all 18 reference comparisons. Modeled CAGR, later-test results and Sharpe do not rank.
Qualification bootstraps are unchanged. Separate comparison bootstraps are skipped in scoring mode.
Retain ≤11 mixes; preserve seed, group leaders/balanced contenders, primary qualified risk-lane incumbents. Select first feasible group's balanced contender for analysis; no feasible mix → no selection.
Setting the JSON switch to false restores the legacy rule: compare portfolios within 0.5 percentage points/year of each group's growth leader by stronger matched comparison-bootstrap bounds, then risk, concentration and turnover; outside the leading band, order by raw CAGR. Results record their effective switch. The v1.2/result-v2.6 cutover removes disposable older results through existing cleanup while preserving accounts. The methodology reference owns configuration, full-union retention and exact disabled-mode rules.
How the method is checked
Repeat searches test whether random choices or a larger search budget materially change the result. Separate chronological checks repeatedly search earlier data, freeze an allocation, and evaluate it on the next period. The final later test uses reserved observations and never changes the ranking.
If alternatives exist but a required training stability check fails or is missing, the output is exploratory and unranked. If every enabled search finds no qualified alternative, a feasible unqualified comparison may still be shown, labeled “No qualified winner.”
Passing history preflight → all five PRIMARY-budget searches in final 50% of training; segment fits rows before , executes at close, earns from
Prefix selection
Reuse public order: qualified first, then feasible unqualified/exempt seed; first feasible row. Keep original qualification, feasibility and eligible counts distinct; infeasible seed is descriptive
Standalone diagnostics
Each feasible fold resets selected mix/seed/references from cash with matching entry costs; preserve completed diagnostics if another fold has no feasible selection
Continuous outer pass
Requires five feasible selections; continuous slices at both costs: ≥4/5 positive minimum-reference excesses, median ≥0.5 pp/year, segment/linked risk caps; base linked recurring turnover ≤1; every bootstrap bound >0
Separate later test
Freeze final weights; entry at next session's close; score from following session; both costs, matched seed/references; rank unchanged
At segment , old holdings earn through the execution close in the preceding segment; transition cost is charged once to the incoming first scored return. Preserve drift/monthly state; continuous initial entry occurs once. Standalone resets are diagnostic and never replace the continuous gate basis. Missing a selection leaves continuous aggregates/bootstrap unavailable and cannot establish a pass; all five refits are still attempted.
Outer bootstrap: resample continuous-path slices within each segment → link five → minimum CAGR excess across three references × two costs. The denominator is all five scheduled folds. Effective PRIMARY budget, seeds, pool settings and bootstrap precision are emitted. Refits test this prefix selector, not the full primary/expanded/repeat union or every displayed mix. Freeze historical retained IDs, weights, original qualification and order before outer/holdout evaluation. Failed/missing outer evidence is shown and cannot unrank that comparison. Required training stability failure/omission with alternatives still gives exploratory output. Unanimous no-qualification across enabled searches → feasible unqualified comparison, “No qualified winner”; N/A/empty agreement does not make stability pass. Valid-domain holdout-return changes with unchanged dates cannot alter historical selection; full-view finite-domain preflight remains in force.
After freeze, projections calibrate on complete common months in the capped window, including reserved later-test rows. The ten-year projection horizon is separate from the available history. Conditional resampling percentiles are neither holdout validation nor parameter uncertainty; later-test reuse and repeated tuning limit claims of independent evidence.
Results depend on the available history, included ETFs, data-provider adjustments, trading assumptions and policy limits.
ETF HHI does not reveal underlying overlap; partial disclosed-position HHI omits the undisclosed tail and unresolved identities. Historical risk limits do not cap future losses. Resampling bounds are not future probabilities.
Searching the development data can favor patterns that fail to repeat (Cawley–Talbot). Tuning against later-test results consumes their independence.
The latest-2,520-return cap is a bounded runtime/recency convention, roughly ten trading years, not an empirically optimal history. The common window can be much shorter for younger ETFs. Recorded regimes depend on acquisition dates and the required universe; older crises and unobserved future changes may be absent. Synthetic mechanics and uncertainty experiments do not establish an optimal real-ETF horizon.
How to read portfolio rankings and evidence scores
The ranking orders portfolios using the methodology in the first paper. Five evidence scores help explain the differences. Raw risk-limit usage also enters the default CAGR/risk ranking described in the first paper; the other scores are descriptive. Read each score with its underlying return, loss or cost measure; a high number by itself is not a verdict on a portfolio.
Measure
Plain-English meaning
Direction
Risk limit used
How much of the allowed historical risk budget was used
Lower means more room below the limit
Historical growth advantage
How recorded net annual growth compared with seed
Higher means more growth relative to seed
Development hurdles passed
How many of the three development periods passed every benchmark hurdle
More periods passed
Downside severity
Largest recorded peak-to-trough loss, as a percentage
Lower means a smaller observed loss
Cost resilience
How growth differed between configured cost cases
Higher means less deterioration; equal cases supply no fee stress
Current and proposed work. P0 implements the default CAGR/risk preference score and these five evidence measures. The raw risk input is used before any logarithmic display transform. Setting activate_cagr_bands_risk_scoring to false retains the legacy growth-band/comparison-bootstrap ranking. The alternative ranking statistic , precision checks and comparison study are P1 proposals. The combined stability score is also proposed, not implemented. Current method/result identities are alpha-p0-v1.2 / scenario-ranking-v2.6.
Do not average the scores. Risk and downside overlap; persistence reuses growth evidence; trading costs already reduce net returns. These measures are not probabilities, confidence levels, or assessments of suitability.1 The existing feasibility limits and qualification groups remain unchanged.2
Reading the formulas
Use the capped common actual calendar, monthly rebalancing and the run's configured base/stress one-way costs. Both default to zero; equal cases provide no distinct fee stress. Annual returns are fractions: 0.005 is 0.5 percentage points. Drawdown is a positive loss magnitude. Round only for display.
The seed is the starting portfolio. is a portfolio's base-cost net annual compound growth on the linked development history; is the matched seed's growth. The scale means 0.5 percentage points/year.
The function turns a growth difference into a smooth score. Zero advantage sits at the midpoint. Larger advantages produce larger scores:
is 50 at zero difference and 75 at a +0.5-point annual difference. The anchor reuses a current product convention, not an empirically optimal trade-off. Never normalize against whichever candidates happen to be displayed. None of these indices is a probability, confidence level, or suitability score.
Development data helped select portfolios. Five chronological checks (outer folds) attempt all five PRIMARY-budget prefix searches when history preflight passes, selecting the first feasible mix under the public qualified-first order; original qualification is recorded separately. Each fit excludes its execution row. These folds test that prefix selector, possibly selecting different mixes each time, rather than the full multi-scope candidate union. A reserved holdout, or later test, evaluates frozen portfolios afterward and never changes rank.3 Historical IDs, weights, qualification and order are frozen before outer and holdout evaluation; missing/failed later evidence cannot change them. Training stability and its no-qualified-alternative rankability exception remain separate; N/A or empty agreement is not a pass. Replays of final weights on earlier history are descriptive. Captured history uses the latest at most 2,520 common returns, with actual history for every required ETF/reference and no exclusions/proxies/renormalization. Standard final training needs 1,262 rows (earliest prefix fit 630, inner blocks 126); accepted limited needs 756 (377/75). Reserve one execution row and the longest allowed holdout: standard 252 at total 1,515, otherwise 126 at 1,389; limited 126 at 883. Valid shorter training is historical-only. Retry/limited rerun preserve the source window and frozen metadata. These policy floors/cap do not establish statistical adequacy or an optimal history.
After rank freeze, modeled paths use complete common months in the capped window, including any later-test observations. The projection horizon is separate from historical coverage. Conditional percentiles do not measure parameter uncertainty or supply holdout validation. Resampling cannot fill missing regimes or remove current-universe/provider limitations.
Risk limit used
This measures how close historical drawdown or volatility came to its allowed limit. A value of 80 means the tightest limit was 80% used; 100 reaches it. A value above 100 means a breach. Lower means more headroom, not necessarily better returns or lower future risk.
The results page labels a rescaled version Risk score. It expands differences near the limit: 95% utilization becomes a score of 80; 100% becomes 100. A Risk score of 80 therefore does not mean 80% of the limit was used. Check the risk driver beside it for the observed drawdown or volatility and its limit. The scores summarize the existing risk checks; the rescaling does not alter their limits or portfolio order.
Risk-limit formula
Use every enforced risk scope and both cost cases. Here is a checked historical period and is the trading-cost case:
Lower means more risk-budget headroom, not greater investment quality. reaches the tightest limit; above 100 breaches it. Do not clip breaches. Zero observed risk against a zero cap contributes zero; positive risk against a zero cap is a breach, not a finite valid score.
Show the limiting metric, scope, observed value and cap. This is historical policy utilization, not total investment risk. Keep weight concentration and liquidity checks separate; ETF count and HHI do not reveal underlying overlap. The separate whole-mix disclosed-position HHI (partial estimate) matches eligible frozen top-holding symbols on original ETF weights, squares their combined exposures and leaves the unknown tail unscaled. Read it with weighted coverage, undisclosed mass and top-ten exposure. It can include funds/cash and unconsolidated share classes; it is not issuer HHI or a diversification claim. No usable positions means unavailable/null, not zero. It changes no evidence score, qualification, rank or projection; unknown holdings dates stay unknown and stale/future/malformed rows are withheld. Reports/web/AI share the same deterministic post-selection output, including when AI is unavailable. Existing results contain the necessary inputs.
How the displayed Risk score is scaled
Let be raw risk utilization on the percentage scale above. For valid utilization from 0 to 100, the display uses:
Zero utilization maps to 0, 95 maps to 80, and 100 maps to 100. This is a presentation convention, not another financial risk measure. A breach is shown as Limit breached; missing evidence is Not evaluated.
Historical growth advantage
This compares recorded net annual growth with seed. A score of 50 means equal growth; 75 corresponds to an advantage of 0.5 percentage points/year. It does not measure the chance of a future gain.
Show Historical growth advantage alongside actual net CAGR, the difference from seed, and fixed-reference comparisons.
Rising-market participation can be shown separately; higher market exposure does not establish better long-term growth. Neither receives another rank bonus.
Development hurdles passed
This counts how many of three development periods beat every fixed benchmark by at least 0.5 percentage points/year, under both cost rates. Read it as a count: 2/3 means two periods passed. It is not a probability that the next period will pass.
Let be portfolio ’s smallest annual growth advantage across all three fixed references and both cost cases in development period . The indicator contributes 1 when that period meets the hurdle, and 0 otherwise:
Show Development hurdles passed with the count, such as , as the primary display; keep the existing percentage formula for structured scores. Show the hurdle, per-period excesses and dates alongside it. The three outcomes are not independent trials or necessarily distinct regimes. All qualified portfolios pass 3/3 by definition; inventing finer distinctions would add false precision. Existing results supply the inputs.
At run level, separately report outer-fold positive-excess frequency: . Use the outer rule of strictly positive minimum reference excess, not the development +0.5-point hurdle. Show median, risk and bootstrap checks too: four positive folds alone do not establish a pass. The denominator remains all five scheduled folds. Continuous aggregates and bootstrap require five feasible selections; otherwise they are unavailable. Completed standalone cash-reset diagnostics remain visible for each feasible fold, using matched comparator entry costs. They do not supply the outer gates: those use continuous-path slices with drift and incoming transition costs. Do not copy this shared score onto each portfolio.
Downside severity
This is the largest observed loss from a previous peak across the full training history, including the initial context period. A value of 28 means a 28% loss. It describes what happened in the record, not the maximum possible future loss:
Show the full-training dates. The result contains this drawdown, but does not currently provide its peak/trough dates.
Show what happened in falling markets through a separate dated event table. Define events from a fixed reference before comparing candidates. A proposed rule is each VT running-peak-to-recovery episode whose trough falls at least 10% below its peak; mark an unfinished episode unrecovered. Include every qualifying episode within training, with matched returns, drawdown and recovery. This threshold is a display convention. Do not annualize short crisis returns or imply that a portfolio's worst loss occurred during a broad-market decline.
Stability of the search method (proposal)
This proposed run-level measure asks whether modest changes to dates, execution assumptions or random seeds materially change the selected weights and growth. The existing repeatability checks cover only part of this question.
Proposed stability formula and test variants
For each predeclared variant , rerun the same declared selection rule and compare its chosen weights with base weights . Recompute stressed net CAGR and on one identical evaluation calendar:
Here identifies an ETF; and are its variant and base weights, as fractions. Higher is better. means the worst tested change reaches a stated tolerance; the tolerances are proposed conventions. Show weight distance and growth difference separately: different allocations can behave similarly.
An initial fixed suite could include an earlier cutoff by 21 sessions, quarterly instead of monthly rebalancing, and existing alternate random seeds/expanded search. Hold the resolved universe fixed, preserve minimum-history requirements, and use only training observations. Fix the evaluation calendar to the common intersection of the valid variant training periods. For quarterly rebalancing, evaluate quarterly execution against the base monthly execution; apply each convention to its comparators too. This measures combined selection/execution sensitivity through historical replays.
Publish this once per run. If a required variant lacks valid evidence, publish the available components and mark the combined score Not evaluated. An alternative-existence disagreement is an explicit failure, not evidence to average away. No qualifying alternative is not applicable. Current results measure search repeatability/convergence; combined date/assumption sensitivity needs new work. Do not assign this shared score to each final portfolio.
Cost resilience
This measures how much annual growth falls between the configured base and stress one-way cost cases. A score of 100 means no deterioration; 50 means losing 0.5 percentage points/year. Higher means less sensitivity to trading costs. With current zero/zero defaults the paths coincide, so a score of 100 supplies no evidence about resilience to higher fees. Disclose rates beside the score. Let be the difference between base-cost and stress-cost growth:
The growth difference must be nonnegative, apart from numerical roundoff. A material increase in stressed returns is an integrity error, not a score of 100.
Label Cost resilience. Always show base/stress CAGR and stress-case excess against seed and references, plus turnover and liquidity status. A high score can coexist with no growth advantage. Name the failed comparison: No advantage over seed after costs or Does not exceed every fixed reference after costs. Cost resilience cannot offset failed implementation limits. Actual/modeled replay uses rate × pretrade wealth × half-L1 turnover including cash. Funded projection uses rate × (executed purchases + sales). Its configured rate agrees with modeled TWR, but this existing dollar-trade convention differs at positive rates; fee normalization is deferred. The recurring 252-session half-L1 limit remains 1, excluding entry. Tax, flow and expense-drag conventions are unchanged.
Ranking index (P1 proposal)
This proposal would favor growth that depends less on favorable stretches of recorded history. It would replace one comparison statistic within each group's leading growth band. It would not change feasibility or qualification rules. This separate proposal is not the implemented default CAGR/risk preference score. The proposed measure is conservative historical growth under resampling, not a forecast or future confidence bound.
Proposed ranking formulas and ordering
Calculate the experimental statistic as follows. For each cost rate , stationary-bootstrap4 block length in , and replicate , concatenate the resampled development blocks in their original block order and calculate:
Here is the daily net return. The adjusted log returns remove establishment once, at the first session of the linked path. Let be their jointly resampled observations in replicate , with development blocks kept in their original order:
Here is the configured base or stress one-way rate, currently zero for both. Remove establishment from the first session of the entire linked path, then restore it once per replicate; do not adjust later block boundaries. Recurring costs stay jointly resampled with returns. These samples are conditional statistics, not executable trading paths.
is the total development-session count. Use one primary comparison-bootstrap context, sharing indices across candidates and costs. Compute for the entire leading band of each feasible group before retention, plus seed. With 1,024 replicates, is the 103rd-smallest value, without interpolation.2
favors growth less dependent on favorable historical samples. It remains conditional historical evidence, not a future confidence bound or a correction for searching many candidates.5 Existing excess-only bounds cannot reconstruct it. Use it only near growth leaders to limit a drift toward lower market exposure.
For each feasible portfolio:
Let be 1 for qualified portfolios and 0 otherwise, including seed. Let be 1 inside the qualification group's leading 0.5-point growth band and 0 outside it.
Use the existing canonical integer band calculation. Higher is preferred:
Interval
Meaning
75–100
Qualified, in its group's leading growth band
50–75
Qualified, outside that band
25–50
Feasible unqualified, in its group's leading growth band
0–25
Feasible unqualified, outside that band
Endpoints are open for finite inputs. Infeasible portfolios and exploratory runs have no index. Preserve the existing no-alternative exception.
Sort groups and band membership first, then directly by within the band or outside it, not by rounded or transformed scores. Within-band ties use drawdown, volatility, HHI and turnover ascending, minimum/median reference excess descending, then allocation ID. Outside-band CAGR ties use ID. replaces the current comparison-bound ordering only inside the leading band.
The tiers encode policy. Crossing qualification or a band boundary can produce a large jump; a new leader can change another portfolio's band. Compare indices within one run, with rank, qualification, observed CAGR and raw . Score differences do not quantify economic-performance differences.
Comparison and validation
P1 ranking illustration with feasible, qualified, leading-band portfolios; seed , . The invented cost differences assume distinct positive base/stress rates; they are not current zero-fee output:
Measure
Scenario A
Scenario B
Observed net CAGR
14.0%
14.3%
Conditional conservative historical growth
8.0%
7.7%
Ranking index
97.5
97.2
Historical growth advantage
90
91
Risk limit used
80
95
Downside severity
20%
25%
Development hurdles passed
3/3
3/3
CAGR lost under higher costs
0.10 points
0.20 points
Cost resilience
83
71
B has 0.3 points more observed growth; A has higher conditional conservative growth, more risk headroom and less cost sensitivity.
Run results show ranking, risk and growth advantage with raw drivers; expand other evidence. All-results uses a compact comparison table, preserving selection when sorting. Shared folds and stability appear once above the explorer. Label direction, dates and counts; show missing evidence explicitly. Both pages retain separately dated holdout comparisons.
How the proposed ranking would be validated
In P1, compare one fixed challenger with the actual incumbent final ranking. At each earlier-only chronological refit, apply both rules to the same full evaluated candidate union before retention, with unchanged search policy and budgets, qualification rules and feasibility limits. Freeze selections before the following segment. Reuse outer-fold chronology and transition accounting; rerun the two declared selectors rather than reusing production prefix choices. The rules differ in benchmark relativity and worst-period versus linked-growth aggregation; compare the complete rules without attributing any benefit to one ingredient.
Check both bootstrap statistics over the complete leading bands with larger replicate counts and alternate shared streams. Within each refit, hold candidates, qualification, bands, paths, costs and tie-breakers fixed. Assess material selection changes using growth, risk, exposure, turnover and transition costs. These are proportionate offline diagnostics, not permanent runtime gates.
Before inspecting results, declare the primary cost convention, economic materiality threshold and uncertainty treatment for matched subsequent linked net compounded growth, including transitions. Count every scheduled period and preserve public fallback/failure behavior; qualified-only screening supports conditional claims. Choosing a method using these segments makes them selection evidence. Reserve an untouched period for final assessment; do not tune to current later-test results. Retain production ranking unless demonstrates a material net-growth benefit under that criterion.5
How AI uses the evidence
The AI handoff supplies calculated scores and raw drivers per scenario, with definitions once. Repeated dates and seed-comparison fields use shared references; later displays accompany their numeric cost cases. The model explains supplied evidence without calculating scores or changing rank. Page, AI and Second Opinion wording describes risk headroom and development hurdles consistently. The web projection and AI calculations share known-answer tests. Missing evidence is Not evaluated; the v1.2 cutover removes disposable older result schemas and preserves accounts. Second Opinion shares complete records in structured JSON. Expansion preserves all facts and numeric literals; source whitespace and string escaping may change.
References and notes
The five evidence measures and default CAGR/risk preference score are implemented product conventions. Raw risk-limit usage affects enabled ranking; the alternative ranking and combined stability measures are proposals. The research supports validation principles, not the particular formulas or display anchors.