Approximation-Error Experiments — Exact vs Sampling-of-Alternatives vs RL

Same small synthetic network, θ*, and seeds as the recovery study (Report 1) · R=30 · N=200 · generated 2026-06-21

0 The idea, in one read

Scoring one observed day-plan requires comparing it against all possible day-plans — the normalizer V̄(x0)=log Σ exp(U). That sum is intractable; each estimator is a different shortcut for it.

Exact computes it in full (backward induction) — the gold-standard reference. SA compares against a sampled handful of plans, re-weighted by McFadden survey weights so the handful is unbiased — like a properly-weighted poll: right on average at any sample size, more samples only tighten it. RL instead estimates the normalizer by Monte-Carlo importance sampling and plugs it in; but the estimate enters through log(average), and the log of a noisy average is biased low (Jensen) — a curvature tax with no correction, removed only by spending more rollouts.

The headline contrast: SA is consistent (variance only, bias≈0 at every budget); RL is inconsistent at finite budget (extra bias and variance, both →0 only as the rollout budget B→∞). Full math + intuition: docs/estimation/reports/rl_value_estimation_landscape.md.

1 Three estimators, three error structures

estimatorhow V̄ is obtainedbudget knobextra biasextra variance
Exact NFXPfull backward induction—— (reference)— (finite-N only)
SA (Västberg)McFadden-corrected logsum over K sampled pathsK≈ 0 (consistent)denominator-sampling → 0 as K→∞
RL (root-IS)MC importance-sampling estimate of log Σ exp(U)B≠ 0 at finite B (no correction)value-fn-approx → 0 as B→∞

SA is Dekker (2025) territory; RL is the novel layer — value-function-approximation bias with no known correction.

2 The Exact reference (recap from Report 1)

The variance comparison is anchored on the Exact NFXP estimator, whose recovery + identification is the subject of Report 1. It makes no approximation, so its only error is ordinary finite-sample noise — that noise is the yardstick: SA and RL are reported as "× noisier than Exact" below, and "1×" means as good as Exact.

How to read: the curve is the spread (SD across R=30) of the Exact estimate of theta_travel as the dataset size N grows; it tracks the dashed 1/√N line — the textbook signature of a correct, consistent estimator. Everything in §3–§4 is measured relative to the N=200 point on this curve.

Full Dekker table, profile-LL, and behavioral validation: see Report 1 (recovery_smallnet_R30_*.html). N=200 row is the reference for the panels below.

3 SA — Sampling-of-Alternatives (consistent)

What SA does: rather than summing over all day-paths for the normalizer, it compares the observed path against a sampled handful of K alternative paths, each re-weighted by the McFadden survey weight log q. That weight makes the small sample an unbiased stand-in for the full comparison — like a properly-weighted poll. The sweep varies the logsum size K.

How to read: the variance line falls toward 1× (= Exact) as K grows; the bias bars stay near 0 at every K (the consistency property — unlike RL in §4); the density panels sit on θ* and merely tighten as K grows. Unlike RL, there is no blow-up — every K is readable.

Extra variance vs Exact (median, well-id params)

How many times noisier SA's θ̂ are than Exact, per logsum size K. 1× = matches Exact.

Relative bias per parameter, per K (well-identified)

One coloured group per K — all hug 0 (McFadden ⇒ unbiased at any K). This is the key difference from RL. (alpha/beta omitted: θ*≈0.001 makes %-bias meaningless.)

θ̂ distributions by K (one panel per parameter; dashed line = θ*, the truth)

Each coloured curve = spread of recovered θ̂ at one K. They sit centred on θ* even at K=5 (no bias) and only sharpen as K grows — the visual signature of a consistent estimator.

SA is consistent. Bias stays small at every K (McFadden); the density curves sit on θ* and tighten as K grows — variance is the only cost of the approximation.

4 RL — value-function estimation (inconsistent)

What RL does: instead of computing the normalizer exactly (Exact) or sampling the denominator (SA), RL estimates the value V̄(x0) by Monte-Carlo importance sampling over B rollout paths and plugs the estimate into the likelihood. Because the estimate enters through log(average), it is biased (Jensen) with no correction — proposal temperature τ=1.5. The sweep varies the rollout budget B.

How to read the panels below: the variance line should fall toward 1× (= as good as Exact) as B grows; the bias bars should shrink toward 0 (but never exactly — that's the inconsistency); the density panels should drift onto θ* and tighten as B grows.

Why the charts start at B=50, not B=10. At B=10 the value estimate is so noisy the optimizer diverges — θ̂ blows up by factors of millions, which would flatten every other bar/curve to invisibility. B=10 is kept in the table below (the "blow-up regime"); the charts show B≥50 where the trend is readable.

B (rollout budget)convergencemedian Var_RL/Var_Emax |bias| (well-id params)
1037%3.4e+13×44,833,544.5%  ← blow-up
5077%5.4×8.9%
20080%1.1×4.8%
100093%0.5×4.4%

Extra variance vs Exact (median, well-id params)

Each point = how many times noisier RL's θ̂ are than Exact, at that budget B. Falls from 5.4× (B=50) toward 1× as B grows.

Relative bias per parameter, per B (well-identified)

One coloured group per B. Bars shrink toward 0 as B grows but stay non-zero — the uncorrected finite-B bias. (alpha/beta omitted: θ*≈0.001 makes %-bias meaningless.)

θ̂ distributions by B (one panel per parameter; dashed line = θ*, the truth)

Each coloured curve is the spread of recovered θ̂ at one budget B (B≥50). As B grows the curve should move onto θ* and narrow — for theta_travel / beta1_shop / beta1_leis you can see the B=1000 curve tightest and centred, the lower-B curves wider and slightly off θ*.

RL is inconsistent at finite B. At small B the value estimate is so noisy the optimizer diverges (conv 37% at B=10); both bias and variance shrink only as B grows (conv 93% at B=1000, variance ≈ Exact). The bias has no correction — only a larger budget removes it. This is the headline contrast with SA (§3), which is unbiased at every budget.
Proposal-temperature finding: τ is critical and horizon-dependent. τ=2 (fine on the 5-step oracle, §5) makes the importance-weight variance explode on the ~48-step small-net paths → garbage at every B; this sweep uses τ=1.5. Production 96-step paths would need a lower τ / better proposal still.

5 Bit-level oracle (enumerable net)

Why this section exists: before trusting the small-net RL sweep (§4), the estimator itself must be proven correct. On a tiny enumerable trellis (32 paths) the value V̄ and the exact MLE can be computed in closed form, so we can check the RL machinery against ground truth bit-for-bit (estimation/experiments/rl_oracle.py). All three gates pass — confirming the §4 results reflect the estimator's real statistical behaviour, not a coding error:

gateresult
1 · telescopingsoft-DP V̄ = brute-force enumeration V̄ (|Δ|<1e-9) ✓
2 · V̂_B → V̄τ=1 perfect-sampler exact at any B; τ=2 Jensen bias −1.8e-3→~0 as B grows ✓
3 · recoveryExact recovers θ*; RL bias vs exact MLE: B=10→1.19, B=10⁴→0.033 (both bias & variance →0) ✓

6 Caveats

  • R=30 → the variance estimates themselves are ±~25% noisy.
  • The small net concentrates probability on a few dominant paths, so SA looks efficient even at small K; a cleaner ΔVar(K) curve would use R=50–100 and/or a harder net.
  • RL variance at B=10 is astronomical (the high-variance regime where the optimizer diverges) — read the B≥50 rows for the trend.
  • The α/β tiny-θ* params are excluded from the median variance ratio (near-zero Exact variance makes the ratio explode); see the per-parameter panels for their behavior.