← All insights

Insight · 7 min read

How accurate are our models, really?

There isn't one accuracy number on Equisect — there are four, because there are four different pricers making four different kinds of claim. Here's what each one has actually earned, in its own terms.

Two different honesty checks, not one

The three Snapshot pricer models — TA, Equisect Bayesian, and Equisect Cohort — are all fitted on the same warehouse cohort curves, with no fund-specific history to lean on beyond age, NAV, and today's numbers. For those three, the right question is "does this model beat a null strategy out-of-sample, or does it only look good because we tried enough variants until one fit the noise?" — answered by a fixed, published statistical battery, corrected for how many variants we actually searched to find them. See the methodology paper for exactly how.

Equisect Ledger, behind the Portfolio-upload and Warehouse pricers, is different in kind: it has a fund's own real dated call/distribution/NAV history to condition on, not a fitted cohort curve, so there's no comparable set of strategy variants to check for overfitting. Its honesty check instead is calibration: does its own P25–P75 band actually cover 50% of held-out funds, measured out-of-fold. Both are real accuracy claims. Neither substitutes for the other.

The three Snapshot models, held out on hundreds of fund-quarters

All three priced roughly 185 mature 2009-2016 vintage funds across ages 2-9, with no look-ahead into the future.

ModelVerdictTypical TVPI missBand covers actual (target 50%)Rank-IC
TA (age-curve baseline)Survives0.45×66%0.52
Equisect BayesianSurvives0.50×60%0.33
Equisect CohortSurvives0.49×47%0.34

All three clear the full deflation battery — see the methodology paper for what that actually requires — so none of them is "the overfit one" any more; TA stays in the product as a transparent, auditable baseline you can sanity-check by eye. Equisect Cohort's band runs a touch narrow (47% vs. the 50% target) — worth widening your own margin around it slightly, not a reason to distrust the point estimate.

Is that difference actually statistically significant?

Clearing the deflation/bootstrap battery tells you a model beats a null strategy. It doesn't tell you whether one model is measurably more accurate than another. For that we ran a two-sample Welch's t-test on every pair of models' held-out absolute pricing errors.

ModelAEPC (median)
TA29.9%
Equisect Bayesian31.6%
Equisect Cohort35.2%

AEPC (Absolute Error as % of Portfolio Commitment) normalizes the dollar forecast error against a fund's total committed capital instead of its own moving NAV — a fixed, known number, so it doesn't get distorted for funds priced late in life with a small remaining NAV base. Honest finding: on this metric TA looks best, which is a reminder that no single accuracy metric tells the whole story — this is exactly why the model with the cleanest overfitting profile (Equisect Cohort) doesn't have the lowest error on every metric.

None of the three pairwise comparisons is statistically significant at the 5% level (TA vs. Bayesian p=0.56; Bayesian vs. Cohort p=0.77; TA vs. Cohort p=0.79). Read plainly: on a per-fund, point-forecast basis, we cannot say any one of the three models is more accurate than another — despite each clearing the deflation battery above by a different margin. That battery measures something real (robustness to overfitting), but it isn't the same claim as "smaller errors," and we're not going to imply it is.

What happens at the portfolio level

A real LP doesn't hold one fund — they hold many, and independent per-fund errors partly cancel out in aggregate. We tested this directly: 300 synthetic 10-fund portfolios built from the same held-out funds, scored by summing each model's forecast and the real outcome across all ten before computing error, instead of fund by fund.

ModelPortfolio typical missPortfolio AEPC (median)
TA0.37×21.5%
Equisect Bayesian0.38×22.1%
Equisect Cohort0.33×19.4%

Every model's error shrinks substantially at the portfolio level versus per-fund (TA's typical miss drops from 0.45× to 0.37×, for instance) — independent fund-level errors genuinely do partly offset each other in aggregate, which is itself a useful thing to know if you're evaluating one fund's price versus a whole portfolio's. Rerunning the same Welch's t-test on portfolio-level errors, one comparison clears the significance bar this time: Equisect Cohort's portfolio error is statistically significantly lower than Equisect Bayesian's (p=0.03). TA vs. Cohort (p=0.10) and TA vs. Bayesian (p=0.57) are not significant. We're publishing this exactly as it came out, including the one place a per-fund non-result turned into a real portfolio-level one — not the other way around, which would have been the more flattering story to tell.

Equisect Ledger: a different kind of honesty

Trained on more than 30,000 fund-quarters across over 1,000 funds, this model's bands are calibrated out-of-fold (GroupKFold by fund, so no fund's own history calibrates its own band) against four targets: how much value is left to realize, how much of that ends up as a terminal residual, how high NAV peaks before winding down, and when. A recent hyperparameter search (per-target, GroupKFold-by-fund CV on pinball loss) replaced this model's previous single hand-picked hyperparameter set with tuned per-target ones, and the warehouse it trains on was refreshed after fixing a currency-conversion bug affecting several non-USD sources — the numbers below reflect both.

TargetBand coverage before calibrationAfter calibration (target 50%)Median bias
Net remaining value36%50%-5.4%
Terminal residual value40%50%+0.1%
Peak NAV multiple38%50%-1.3%
Years to peak44%50%-0.5%

Before calibration, every band under-covered — as low as 36% for net remaining value, meaning the raw model's range was genuinely too narrow and missed the real outcome more often than it should have. A conformal correction step widens each band until it hits its target coverage on held-out data; all four targets now land exactly on 50% (previously one, peak NAV multiple, ran wide to 60% — the tuned hyperparameters fixed that). Median bias is under 1.5% on three of four targets; net remaining value runs the widest at -5.4%, still directionally small next to the width of the band it's correcting.

How accuracy changes over a fund's full remaining life

Every number above measures a fixed decomposition window. We also re-ran the same contributions / distributions / terminal-NAV breakdown over each fund's own full remaining life instead — from its pricing age all the way to its real observed exit — and split the results by (pricing age, years remaining), so "a 5-year-old fund with 8 years left" and "a 5-year-old fund with 10 years left" get scored separately instead of blended into one number.

TypeTAEquisect BayesianEquisect CohortEquisect LedgerEquisect Bootstrap
Contributions (median APE)0.1%0.1%38.5%51.6%62.4%
Distributions (median APE)30.1%33.8%35.9%34.7%49.9%
Terminal NAV (median APE)65.4%68.4%56.0%71.6%49.4%

Contributions look almost solved for TA and Equisect Bayesian at full life — but that's structural, not a forecasting win: the two share the exact same underlying call-rate curve (only their distribution/growth assumptions differ), so by a fund's real exit, "how much gets called" mostly IS "how much commitment is left," which the curve already captures well by design. Equisect Cohort and Equisect Ledger have no equivalent structural guarantee and trail meaningfully behind — Welch's t-test confirms the gap is real (p<0.0001 against both). Distributions tell a different story: five of the six pairwise comparisons among TA, Equisect Bayesian, Equisect Cohort and Equisect Ledger are statistically indistinguishable (p>0.07) — Equisect Ledger's near-term weakness elsewhere on this page mostly washes out once errors are allowed to partly cancel over a full multi-year horizon instead of a fixed near-term window. The exception: Equisect Ledger vs. TA is significant here (p=0.019) — TA's fixed distribution curve turns out to time full-life distributions measurably better than Ledger's learned forecast, even though neither is distinguishable from Equisect Bayesian or Equisect Cohort individually. Terminal NAV is the hardest of the three for every model, and genuinely a coin flip — no pairwise comparison clears significance.

The age/remaining-life breakdown also shows accuracy isn't uniform within a model: Equisect Cohort's Contributions error ranges from as low as 10% (a 2-year-old fund with 9 years left) up to a full 100% for several close-to-exit cells (9-year-old funds with 2-4 years left) — funds nearing the end of their life have less remaining capital-call activity in absolute terms, but what's left is disproportionately hard to time. Equisect Bootstrap only covers remaining-life windows up to 8 years (its stored historical paths run that deep and no further) — with 2-4y funds now included, most of which have 9+ years left to run, roughly two-fifths of test points fall outside that horizon and are scored by every other model but not Bootstrap.

Pooled into coarser pricing-age groups instead of exact-year cells, the same pattern holds and gets easier to read at a glance. The pricing ages tested now run 2-9 (extended down from the original 4-9, once the backtest was widened to give the platform's 2-4y bucket real held-out coverage instead of an untested placeholder), so the table below now also has a 2-4y row.

Age groupTypical remaining lifeTAEquisect BayesianEquisect CohortEquisect LedgerEquisect Bootstrap
Contributions (median APE)
2-4y10y (9-11y)0.1%0.1%20.0%25.7%13.9%
4-6y8y (7-10y)0.1%0.1%29.5%41.0%32.3%
6-8y6y (5-8y)0.2%0.2%45.2%69.4%60.2%
8y+4y (3-5y)0.5%0.5%84.6%100.0%84.0%
Distributions (median APE)
2-4y10y (9-11y)25.8%32.3%34.3%30.6%30.7%
4-6y8y (7-10y)27.7%30.6%35.1%35.1%46.6%
6-8y6y (5-8y)30.1%33.8%35.2%35.3%47.4%
8y+4y (3-5y)38.4%39.1%39.3%38.3%56.0%
Terminal NAV (median APE)
2-4y10y (9-11y)71.9%77.5%57.4%76.8%41.1%
4-6y8y (7-10y)69.3%74.4%59.7%77.3%50.7%
6-8y6y (5-8y)61.6%64.0%50.6%69.9%56.9%
8y+4y (3-5y)55.4%57.2%54.9%58.0%44.6%

"Typical remaining life" is each age group's median years-to-real-exit (P25-P75 in parentheses) — younger pricing ages naturally pair with longer remaining life, so a 2-4y fund's Contributions number is being asked to forecast roughly 10 years out, an 8y+ fund's only about 4. Contributions degrades with age for every model except TA/Bayesian's structurally-near-zero curve, though not perfectly monotonically: Equisect Cohort goes from 20% (2-4y) to 30% (4-6y) to 85% (8y+), Equisect Bootstrap 14% to 84%, and Equisect Ledger the worst of the three, 26% to a full 100% median miss by 8y+ — consistent with the per-cell finding above: less absolute capital-call activity left to time, but what's left is harder to get right. Distributions moves more modestly, but not flat — every model's error rises from 2-4y to 8y+ (Equisect Cohort the least, +5pp; Equisect Bootstrap and TA the most, +25pp, +13pp). Terminal NAV is the one type that gets better with age overall for TA, Equisect Bayesian and Equisect Ledger, 17-20 points lower at 8y+ than at 2-4y — the closer a fund gets to its real exit, the less room the terminal mark has left to drift from today's NAV, though Equisect Ledger ticks up slightly from 2-4y to 4-6y before that trend takes hold. Equisect Cohort and Equisect Bootstrap don't follow the same shape: Cohort dips worse at 4-6y before recovering, and Bootstrap gets steadily worse through 6-8y (its weakest point) before improving sharply at 8y+ — neither model shows the same clean monotonic improvement with age the other three do.

Broken out further, by strategy

The age-group tables above pool every strategy together. Splitting them out by the same 6 strategies the pricer itself supports shows the blended story isn't universal — some of it holds strategy by strategy, some of it doesn't. Sample sizes vary a lot here: Buyout and Growth are the deepest (76-85 test points per age group), Secondaries the thinnest by far (exactly 10 in every age group, so read its numbers as more indicative than precise) — Equisect Bootstrap has no reportable number at all for any strategy at 2-4y (fewer than 5 points cleared the reporting bar in every one of the 6 cells, once split this finely — the pooled 2-4y Bootstrap number a few paragraphs up rests on only 9 points total), plus its usual Secondaries-at-4-6y gap.

Contributions (median APE)

StrategyAge groupTypical remaining lifeTAEquisect BayesianEquisect CohortEquisect LedgerEquisect Bootstrap
Buyout2-4y10y (10-11y)0.2%0.2%15.5%24.4%
4-6y8y (8-9y)0.3%0.3%21.7%44.4%25.2%
6-8y6y (5-7y)0.7%0.7%38.9%66.4%54.7%
8y+4y (3-5y)2.6%2.6%80.7%100.0%87.4%
Growth2-4y11y (9-12y)0.1%0.1%20.9%28.2%
4-6y9y (8-10y)0.1%0.1%34.3%42.1%38.6%
6-8y7y (5-8y)0.2%0.2%44.5%81.5%66.3%
8y+5y (3-6y)0.5%0.5%87.8%100.0%83.8%
Balanced2-4y10y (10-11y)0.0%0.0%19.7%25.0%
4-6y8y (8-9y)0.0%0.0%26.1%25.7%44.9%
6-8y6y (6-7y)0.0%0.0%40.1%53.1%67.4%
8y+4y (4-5y)0.0%0.0%80.4%100.0%71.3%
Secondaries2-4y12y (11-12y)0.0%0.0%62.1%53.0%
4-6y10y (9-10y)0.1%0.1%53.1%51.1%
6-8y8y (7-8y)0.2%0.2%59.9%47.6%42.7%
8y+6y (5-6y)0.6%0.6%70.6%89.6%59.8%
Venture Capital2-4y10y (9-11y)0.0%0.0%29.5%26.1%
4-6y8y (7-9y)0.0%0.0%54.4%65.3%58.9%
6-8y6y (5-7y)0.0%0.0%93.5%134.8%45.0%
8y+4y (3-4y)0.0%0.0%91.5%198.4%100.0%
Infrastructure2-4y10y (9-11y)0.0%0.0%14.2%6.5%
4-6y8y (7-9y)0.0%0.0%25.0%29.1%24.0%
6-8y6y (5-7y)0.0%0.0%51.3%69.0%43.4%
8y+3y (3-5y)0.1%0.1%100.0%221.6%105.0%

Distributions (median APE)

StrategyAge groupTypical remaining lifeTAEquisect BayesianEquisect CohortEquisect LedgerEquisect Bootstrap
Buyout2-4y10y (10-11y)25.9%38.6%34.3%29.7%
4-6y8y (8-9y)27.7%35.2%36.0%30.6%46.5%
6-8y6y (5-7y)31.4%34.3%41.6%34.8%48.2%
8y+4y (3-5y)36.9%38.0%38.6%44.2%55.8%
Growth2-4y11y (9-12y)24.5%25.0%35.5%29.2%
4-6y9y (8-10y)29.0%24.3%38.9%35.1%35.7%
6-8y7y (5-8y)31.7%32.8%28.0%35.3%42.2%
8y+5y (3-6y)44.3%44.2%44.9%37.8%53.7%
Balanced2-4y10y (10-11y)16.9%20.9%26.6%32.8%
4-6y8y (8-9y)20.4%19.0%19.1%37.0%48.8%
6-8y6y (6-7y)24.1%35.7%35.0%35.7%43.5%
8y+4y (4-5y)35.6%36.5%30.3%37.9%40.6%
Secondaries2-4y12y (11-12y)25.1%31.5%43.6%38.9%
4-6y10y (9-10y)27.3%26.5%27.6%30.3%
6-8y8y (7-8y)29.1%32.8%33.9%26.9%26.5%
8y+6y (5-6y)36.5%35.4%32.8%43.2%23.6%
Venture Capital2-4y10y (9-11y)51.2%60.3%26.1%31.9%
4-6y8y (7-9y)42.6%52.3%19.2%38.9%76.8%
6-8y6y (5-7y)41.8%43.6%33.9%31.9%74.6%
8y+4y (3-4y)51.8%50.8%39.2%26.9%78.0%
Infrastructure2-4y10y (9-11y)31.9%29.6%46.8%38.2%
4-6y8y (7-9y)27.1%27.7%38.2%41.9%58.5%
6-8y6y (5-7y)21.7%23.6%35.2%44.9%51.8%
8y+3y (3-5y)24.0%25.9%42.4%28.6%61.3%

Terminal NAV (median APE)

StrategyAge groupTypical remaining lifeTAEquisect BayesianEquisect CohortEquisect LedgerEquisect Bootstrap
Buyout2-4y10y (10-11y)74.8%80.3%57.0%79.2%
4-6y8y (8-9y)69.6%75.5%62.4%85.5%50.7%
6-8y6y (5-7y)62.9%66.5%55.9%74.4%58.6%
8y+4y (3-5y)55.6%59.4%61.2%65.4%45.4%
Growth2-4y11y (9-12y)82.1%85.8%63.4%81.0%
4-6y9y (8-10y)78.4%82.9%67.3%82.5%71.6%
6-8y7y (5-8y)68.5%69.3%53.9%75.2%56.8%
8y+5y (3-6y)56.3%58.9%56.4%60.9%47.3%
Balanced2-4y10y (10-11y)59.8%43.8%47.9%66.7%
4-6y8y (8-9y)54.3%39.2%46.8%72.8%119.8%
6-8y6y (6-7y)50.7%33.2%35.0%53.5%65.7%
8y+4y (4-5y)45.0%35.9%57.4%45.7%52.3%
Secondaries2-4y12y (11-12y)59.5%67.8%57.2%73.0%
4-6y10y (9-10y)59.3%64.3%57.0%61.6%
6-8y8y (7-8y)61.8%59.4%62.3%65.1%77.4%
8y+6y (5-6y)66.9%64.7%65.5%48.0%51.6%
Venture Capital2-4y10y (9-11y)63.2%69.6%55.6%64.1%
4-6y8y (7-9y)61.9%65.3%50.2%52.4%39.5%
6-8y6y (5-7y)61.3%63.6%40.0%45.7%47.6%
8y+4y (3-4y)55.4%57.8%46.5%49.5%39.1%
Infrastructure2-4y10y (9-11y)71.2%55.9%38.1%71.1%
4-6y8y (7-9y)67.4%50.5%31.9%67.4%36.3%
6-8y6y (5-7y)61.0%42.5%23.2%59.1%38.2%
8y+3y (3-5y)52.7%38.4%37.1%52.9%26.6%

Contributions' age-driven blowup is universal — every one of the 6 strategies shows Equisect Ledger's error rising sharply from 2-4y to 8y+, worst in Infrastructure (7% → 222%, +215pp) and Venture Capital (26% → 198%, +172pp), both ending well above 100% — the median forecast miss is now larger than the actual contribution figure it's missing. Terminal NAV's "gets better with age" finding from the pooled table is not universal once split by strategy: Secondaries is the one strategy where TA and Equisect Cohort both end worse at 8y+ than they started at 2-4y instead of better (TA +7.4pp, Equisect Cohort +8.3pp), while Equisect Bayesian and Equisect Ledger still improve over the same range — worth weighing against Secondaries' thin sample (n=10) before reading too much into it, but it's a real number, not smoothed away by the pooled view above.

Equisect Dynamic: blending the best of each model

No single model above wins on every component — TA/Bayesian's shared call curve is close to unbeatable on Contributions, Equisect Cohort and Equisect Bootstrap trade off the best Terminal NAV numbers, nobody dominates Distributions. That's the case for blending — not the final \$ price, but each component (Contributions, Distributions, Terminal NAV) separately, then reassembling. Equisect Dynamic weights each of the 5 models above by inverse held-out error (AEPC), separately per component, per strategy, and per pricing-age bucket (the same 2-4y/4-6y/6-8y/8y+ groups used above) — so the blend shifts as a fund ages instead of staying fixed.

Weights use a shrinkage hierarchy (per-(strategy, age) → per-strategy → global) so a thin cell like Secondaries (n=10) leans on the coarser, better-populated level instead of fitting noise to a handful of points.

The methodology point that matters most here: Dynamic's weights are GroupKFold-by-fund cross-fitted (5 folds) — a fund's own test points never influence the weights used to blend that same fund. Fitting inverse-error weights directly on the points you then "backtest" the blend against would leak, and would make the blend look artificially good on data it was tuned against — exactly the overfitting failure this whole page exists to catch. Every number below is genuinely held-out, the same standard every model above is held to.

TypeTAEquisect BayesianEquisect CohortEquisect LedgerEquisect BootstrapEquisect Dynamic
Contributions (median APE)0.1%0.1%38.5%51.6%62.4%0.6%
Distributions (median APE)30.1%33.8%35.9%34.7%49.9%30.7%
Terminal NAV (median APE)65.4%68.4%56.0%71.6%49.4%58.8%

Read this table with the same caution as every other one on this page — a lower median APE isn't automatically a real improvement, only a significantly lower held-out dollar error is. On Contributions, TA/Bayesian's 0.1% is a structural artifact (same call curve, not learned skill) that even a well-built blend can't fully match while still incorporating three much-worse models — Welch's t-test confirms TA/Bayesian both still beat the blend here (p<0.0001), though the blend significantly beats Equisect Cohort, Ledger and Bootstrap (p<0.0001 against each). On Distributions, the blend is statistically tied with TA and Equisect Bayesian (p=0.92, p=0.23 — not significant) while measurably beating Equisect Ledger (p=0.023) — the closest it comes to an outright win against the field. On Terminal NAV, none of the seven models are statistically distinguishable from each other on real dollar error (every pairwise p-value >0.08) — Equisect Bootstrap's lower headline APE doesn't hold up as a real edge once AEPC (a fixed denominator instead of the fund's own outcome) is used instead: 16.9% for Bootstrap vs. 13.8% for Dynamic, a real gap on this metric but one that doesn't clear statistical significance (p=0.08) once 2-4y funds are included in the test set.

What this means when you're actually pricing something

If you're using the Snapshot pricer, all three models — TA, Equisect Bayesian, and Equisect Cohort — clear the same deflation battery, and TA also stays useful as a transparent, auditable baseline you can sanity-check by eye. If you're uploading a fund's own history through Portfolio-upload or looking one up in the Warehouse, the calibration table above is your honesty check instead: it tells you the band is telling the truth about its own uncertainty, even though there's no "does this beat a null strategy" question to ask of a model built on one fund's own real cashflows. Equisect Dynamic is selectable on all three surfaces — Snapshot, Warehouse, and single-fund Portfolio-upload — blending Equisect Bayesian and Equisect Bootstrap (plus Equisect Ledger, whenever a real dated history is on file — Warehouse and single-fund Portfolio, not Snapshot) by the cell-and-age-specific accuracy weights this section backtests; TA and Equisect Cohort's own point forecasts are deliberately excluded from that blend (each already ships its own standalone verdict above, and blending them in was measured to make Dynamic's pricing-edge consistency worse, not better). Multi-fund Portfolio uploads still price every fund through Equisect Ledger directly, the same restriction Equisect Bootstrap already has.

Across all four, the same rule applies: read the P25–P75 band as the honest range, not the single point estimate dressed up as certainty. That's true whether the band comes from a bootstrap-tested cohort model or a calibrated Equisect Ledger forecast — it's just earned differently in each case. Portfolio-upload results also now include Equisect Bootstrap, a range built directly from real historical warehouse funds rather than any fitted curve or learned model — a fifth, structurally different kind of check, sitting alongside rather than replacing the four above.

One honesty note specific to Equisect Bayesian, Equisect Cohort and Equisect Bootstrap: when a cohort has enough history to fit a reported-vs-realized NAV bias, the fair price discounts near-term cashflows at your required return but the terminal NAV residual at the cohort's own bias-neutral rate instead — not your required return, and often meaningfully lower. Those are genuinely different quantities, so the price shown is not literally "pay this and earn your target return"; the Snapshot results page now surfaces the actual blended IRR at the fair price alongside it (previously it didn't, and the copy asserted the two were equal, which they generally aren't) — a real fix, not just an insights-page footnote, since the old wording could read as a stronger promise than the model was actually making.

Research and software, not investment advice.