Scoring one observed day-plan requires comparing it against all possible day-plans — the
normalizer V̄(x0)=log Σ exp(U). That sum is intractable; each estimator is a different
shortcut for it.
Exact computes it in full (backward induction) — the gold-standard reference.
SA compares against a sampled handful of plans, re-weighted by McFadden survey weights so the
handful is unbiased — like a properly-weighted poll: right on average at any sample size,
more samples only tighten it. RL instead estimates the normalizer by Monte-Carlo
importance sampling and plugs it in; but the estimate enters through log(average), and
the log of a noisy average is biased low (Jensen) — a curvature tax with no correction, removed
only by spending more rollouts.
docs/estimation/reports/rl_value_estimation_landscape.md.| estimator | how V̄ is obtained | budget knob | extra bias | extra variance |
|---|---|---|---|---|
| Exact NFXP | full backward induction | — | — (reference) | — (finite-N only) |
| SA (Västberg) | McFadden-corrected logsum over K sampled paths | K | ≈ 0 (consistent) | denominator-sampling → 0 as K→∞ |
| RL (root-IS) | MC importance-sampling estimate of log Σ exp(U) | B | ≠ 0 at finite B (no correction) | value-fn-approx → 0 as B→∞ |
SA is Dekker (2025) territory; RL is the novel layer — value-function-approximation bias with no known correction.
The variance comparison is anchored on the Exact NFXP estimator, whose recovery + identification is the subject of Report 1. It makes no approximation, so its only error is ordinary finite-sample noise — that noise is the yardstick: SA and RL are reported as "× noisier than Exact" below, and "1×" means as good as Exact.
How to read: the curve is the spread (SD across R=30) of the Exact estimate of
theta_travel as the dataset size N grows; it tracks the dashed 1/√N line —
the textbook signature of a correct, consistent estimator. Everything in §3–§4 is measured relative
to the N=200 point on this curve.
Full Dekker table, profile-LL, and behavioral validation: see Report 1
(recovery_smallnet_R30_*.html). N=200 row is the reference for the panels below.
What SA does: rather than summing over all day-paths for the normalizer, it compares
the observed path against a sampled handful of K alternative paths, each re-weighted by the
McFadden survey weight log q. That weight makes the small sample an unbiased
stand-in for the full comparison — like a properly-weighted poll. The sweep varies the logsum size K.
How to read: the variance line falls toward 1× (= Exact) as K grows; the bias bars stay near 0 at every K (the consistency property — unlike RL in §4); the density panels sit on θ* and merely tighten as K grows. Unlike RL, there is no blow-up — every K is readable.
How many times noisier SA's θ̂ are than Exact, per logsum size K. 1× = matches Exact.
One coloured group per K — all hug 0 (McFadden ⇒ unbiased at any K). This is the key difference from RL. (alpha/beta omitted: θ*≈0.001 makes %-bias meaningless.)
Each coloured curve = spread of recovered θ̂ at one K. They sit centred on θ* even at K=5 (no bias) and only sharpen as K grows — the visual signature of a consistent estimator.
What RL does: instead of computing the normalizer exactly (Exact) or sampling the
denominator (SA), RL estimates the value V̄(x0) by Monte-Carlo importance
sampling over B rollout paths and plugs the estimate into the likelihood. Because the estimate
enters through log(average), it is biased (Jensen) with no correction —
proposal temperature τ=1.5. The sweep varies the rollout budget B.
How to read the panels below: the variance line should fall toward 1× (= as good as Exact) as B grows; the bias bars should shrink toward 0 (but never exactly — that's the inconsistency); the density panels should drift onto θ* and tighten as B grows.
| B (rollout budget) | convergence | median Var_RL/Var_E | max |bias| (well-id params) |
|---|---|---|---|
| 10 | 37% | 3.4e+13× | 44,833,544.5% ← blow-up |
| 50 | 77% | 5.4× | 8.9% |
| 200 | 80% | 1.1× | 4.8% |
| 1000 | 93% | 0.5× | 4.4% |
Each point = how many times noisier RL's θ̂ are than Exact, at that budget B. Falls from 5.4× (B=50) toward 1× as B grows.
One coloured group per B. Bars shrink toward 0 as B grows but stay non-zero — the uncorrected finite-B bias. (alpha/beta omitted: θ*≈0.001 makes %-bias meaningless.)
Each coloured curve is the spread of recovered θ̂ at one budget B (B≥50). As B grows the curve should move onto θ* and narrow — for theta_travel / beta1_shop / beta1_leis you can see the B=1000 curve tightest and centred, the lower-B curves wider and slightly off θ*.
Why this section exists: before trusting the small-net RL sweep (§4), the estimator itself
must be proven correct. On a tiny enumerable trellis (32 paths) the value V̄ and the exact MLE
can be computed in closed form, so we can check the RL machinery against ground truth bit-for-bit
(estimation/experiments/rl_oracle.py). All three gates pass — confirming the §4 results
reflect the estimator's real statistical behaviour, not a coding error:
| gate | result |
|---|---|
| 1 · telescoping | soft-DP V̄ = brute-force enumeration V̄ (|Δ|<1e-9) ✓ |
| 2 · V̂_B → V̄ | τ=1 perfect-sampler exact at any B; τ=2 Jensen bias −1.8e-3→~0 as B grows ✓ |
| 3 · recovery | Exact recovers θ*; RL bias vs exact MLE: B=10→1.19, B=10⁴→0.033 (both bias & variance →0) ✓ |