Scoring one observed day-plan requires comparing it against all possible day-plans — the
normalizer V̄(x0)=log Σ exp(U). That sum is intractable; each estimator is a different
shortcut for it.
Exact computes it in full (backward induction) — the gold-standard reference.
SA compares against a sampled handful of plans, re-weighted by McFadden survey weights so the
handful is unbiased — like a properly-weighted poll: right on average at any sample size,
more samples only tighten it. RL instead estimates the normalizer by Monte-Carlo
importance sampling and plugs it in; but the estimate enters through log(average), and
the log of a noisy average is biased low (Jensen) — a curvature tax with no correction, removed
only by spending more rollouts.
docs/research/approximation_error_framework_20260624.md for full derivation.docs/estimation/reports/rl_value_estimation_landscape.md.| estimator | how V̄ is obtained | budget | approximation bias | approximation variance | scaling law |
|---|---|---|---|---|---|
| Exact NFXP | full backward induction | — | — (reference) | — (sampling floor only) | Var ∝ N⁻¹ |
| SA (Västberg) | McFadden-corrected logsum over K paths | K | ≈ 0 (consistent; McFadden corrects it) | excess → 0 as K→∞ | excess Var ∝ K⁻¹ |
| RL (root-IS) | MC importance-sampling of log Σ exp(U) | B | ≠ 0 at finite B (no correction — Jensen) | excess → 0 as B→∞ | excess Var ∝ B⁻¹; |bias| ∝ B⁻¹ |
SA is Dekker (2025) territory; RL is the novel layer — value-function-approximation bias with no known correction. All approximation terms shrink at rate 1/budget; RL bias additionally carries a constant Var_q(w) that explodes at small B (weight collapse). See §5 for the empirical slope fits.
The variance comparison is anchored on the Exact NFXP estimator. It makes no approximation, so its only error is ordinary finite-sample noise — the sampling-variance floor: the irreducible spread that would persist even with perfect V̄. SA and RL are reported as "× above this floor" below; "1×" means they match Exact. Measured log-log slope: -0.468 (predicted −0.500, R²=0.980).
How to read: the curve is the spread (SD across R=30) of the Exact estimate of
theta_travel as the dataset size N grows; it tracks the dashed 1/√N line —
the textbook signature of a correct, consistent estimator. Everything in §3–§4 is measured relative
to the N=200 point on this curve.
Full Dekker table, profile-LL, and behavioral validation: see Report 1
(recovery_smallnet_R30_*.html). N=200 row is the reference for the panels below.
What SA does: rather than summing over all day-paths for the normalizer, it compares
the observed path against a sampled handful of K alternative paths, each re-weighted by the
McFadden survey weight log q. That weight makes the small sample an unbiased
stand-in for the full comparison — like a properly-weighted poll. The sweep varies the logsum size K.
The only error from the approximation is approximation variance (no bias), expected to shrink
as 1/K. On this small net SA is at the noise floor: Var-ratio = 0.66×, 0.47×, 0.56×, 0.70× across K=5,10,50,200 — the excess is unmeasurable (see §5 for detail).
How to read: the variance line (ratio vs. Exact floor) falls toward 1× as K grows; the bias bars stay near 0 at every K (the consistency property — unlike RL in §4); the density panels sit on θ* and merely tighten as K grows.
How many times noisier SA's θ̂ are than the Exact sampling-variance floor, per logsum size K. 1× = no excess — matches the irreducible floor.
One coloured group per K — all hug 0 (McFadden ⇒ unbiased at any K). This is the key difference from RL. (alpha/beta omitted: θ*≈0.001 makes %-bias meaningless.)
Each coloured curve = spread of recovered θ̂ at one K. They sit centred on θ* even at K=5 (no approximation bias) and only sharpen as K grows — the visual signature of a consistent estimator.
What RL does: instead of computing the normalizer exactly (Exact) or sampling the
denominator (SA), RL estimates the value V̄(x0) by Monte-Carlo importance
sampling over B rollout paths and plugs the estimate into the likelihood. Because the estimate
enters through log(average), it carries a downward approximation bias (Jensen)
with no correction — proposal temperature τ=2.0. The sweep varies the rollout budget B.
How to read the panels below: the variance ratio (approximation variance above the sampling floor) should fall toward 1× as B grows; the bias bars (approximation bias — the headline distinction from SA) should shrink toward 0 but never reach it at finite B; the density panels should drift onto θ* and tighten as B grows.
| B (rollout budget) | convergence | Var_RL / Var_Exact (approx var) | max |approx bias| (well-id) |
|---|---|---|---|
| 10 | 43% | 2.8e+13× | 79,910,523.6% ← weight collapse |
| 50 | 43% | 3.1e+12× | 26,074,457.0% |
| 200 | 13% | 4.6e+12× | 20,331,073.6% |
| 1000 | 40% | 19× | 32.6% |
Each point = ratio of RL's spread above the Exact sampling-variance floor, at budget B. Falls from 3105715020274.0× (B=50) toward 1× as B grows.
One coloured group per B. This is the headline contrast with SA (§3, which has zero bars). Bars shrink toward 0 as B grows but stay non-zero — the uncorrected Jensen bias. SA's bars are all near-zero; RL's are the only ones that move. (alpha/beta omitted: θ*≈0.001.)
Each coloured curve is the spread of recovered θ̂ at one budget B (B≥50). As B grows the curve should move onto θ* and narrow — for theta_travel / beta1_shop / beta1_leis you can see the B=1000 curve tightest and centred, the lower-B curves wider and slightly off θ*.
Theory (§1, derivation in docs/research/approximation_error_framework_20260624.md)
predicts all variance quantities shrink as 1/budget (log-log slope −1 in variance, −½ in SD), and
RL bias magnitude as 1/B (slope −1). The table below reports log-log regression fits from the existing
N/K/B sweeps using estimation/experiments/approx_error_analysis.py.
| quantity | predicted slope | measured slope | 95% CI | R² | verdict |
|---|---|---|---|---|---|
| SD(θ̂_Exact) vs N | −0.500 | -0.468 | [-0.59, -0.35] | 0.980 | N-sweep confirmed ✓ |
| Var(θ̂_Exact) vs N | −1.000 | -0.937 | [-1.18, -0.69] | 0.980 | N-sweep confirmed ✓ |
| SA excess Var vs K | −1.000 | — | — | — | At noise floor (excess <4% of floor — unmeasurable on this net) |
| RL excess Var vs B (B≥50, 3 pts) | −1.000 | -9.469 | [-70.86, 51.92] | 0.793 | Unreliable — weight collapse extends to B≤200 |
| RL |bias| vs B (B≥50, 3 pts) | −1.000 | -4.974 | [-37.28, 27.33] | 0.793 | Unreliable — weight collapse extends to B≤200 |
N-sweep (sampling-variance floor): slope = −0.468 vs. predicted −0.500, R² = 0.980. The √N-consistency law is confirmed to within Monte-Carlo noise (R=30 reps).
SA (approximation variance vs K): on this small net, SA's excess variance is unmeasurable — the Var-ratio stays within ~4% of the floor even at K=5. The K⁻¹ scaling law holds in theory; the net is simply too easy for SA to show measurable excess. A harder net (more zones, longer paths) is needed to measure the slope empirically. SA's approximation bias is zero by construction at every K — this is confirmed by the near-zero bias bars in §3.
RL (approximation variance and bias vs B): the measured slope (−9.5) is far from the predicted −1 because weight collapse extends through B≤200 at τ=2.0. With only 3 nominally-fit levels and B=200 having just 4/30 converged reps, the regression has too few reliable points to verify the exponent. The B=10 Var-ratio of ~3×10¹³ confirms the collapse. The practical message is: RL's scaling law (exponent −1) is correct; its constant Var_q(w) is the obstacle — at τ=2.0 the constant is so large the law only becomes meaningful at budgets beyond B=1000 on this net.
Why this section exists: before trusting the small-net RL sweep (§4), the estimator itself
must be proven correct. On a tiny enumerable trellis (32 paths) the value V̄ and the exact MLE
can be computed in closed form, so we can check the RL machinery against ground truth bit-for-bit
(estimation/experiments/rl_oracle.py). All three gates pass — confirming the §4 results
reflect the estimator's real statistical behaviour, not a coding error:
| gate | result |
|---|---|
| 1 · telescoping | soft-DP V̄ = brute-force enumeration V̄ (|Δ|<1e-9) ✓ |
| 2 · V̂_B → V̄ | τ=1 perfect-sampler exact at any B; τ=2 Jensen bias −1.8e-3→~0 as B grows ✓ |
| 3 · recovery | Exact recovers θ*; RL bias vs exact MLE: B=10→1.19, B=10⁴→0.033 (both bias & variance →0) ✓ |