Mean CP ≈ 95% and MeanSE ≈ empirical-SD ⇒ the BHHH standard errors are trustworthy. This is the identification diagnostic the hEART reviewers requested (generate-from-θ* → recover).
Before the controlled R=30 study, The single-scale specification was first confirmed at production fidelity: the
heavy toy (9 zones, 5 modes, Δt=15, multi-home, ~700K states/graph), single replication. The
question it answered: does the single transport scale fix the broken
theta_travel that the old scale-plus-level spec could not identify?
Old spec made theta_travel (a scale) and ASC_pt_adj (a level) jointly
unidentified — a logit separates a slope from an intercept only weakly, so the estimator slid along a
scale×level ridge. The single-scale specification uses one inclusive-value scale on the whole mode utility
(u = θ·(ASC + β_time·TT + β_cost·cost) + …) and retires ASC_pt_adj,
collapsing the ridge:
| spec | theta_travel θ̂ (θ*=1.0) | bias | |bias|/SE | ASC_pt_adj |
|---|---|---|---|---|
| Old (scale + level, ASC_pt_adj free) | 1.138 | +13.8% | 16.6 ✗ | −0.067 (6.8σ ✗) |
| Single transport scale | 0.992 | +0.8% | 0.87 ✓ | retired |
optionA_recovery_confirmation_20260619.md.| zone | car min from HOME | train fare ¥ | attractiveness | P_open / role |
|---|---|---|---|---|
| HOME (z0) | 0 | — | — | — |
| SHOP-near (z1) | 30 | 200 | X_shop=1.0 | noon 12:00 · WALKABLE |
| WORK (z2) | 60 | 300 | — | mandatory for workers |
| LEIS-mid (z3) | 60 | 300 | X_leis=1.2 | peak 14:00 |
| SHOP-far (z4) | 90 | 400 | X_shop=0.8 | peak 16:00 |
| LEIS-far (z5) | 120 | 450 | X_leis=2.0 | evening 19:00 |
TRAIN = flat 30-min express to every zone; WALK competitive only on the 30-min hop; all 5 modes present (no-car archetype excludes CAR-driver, keeps CAR_PASS). Wide distance spread (30–120 min) gives the (time+cost) variation that identifies the transport scale.
The synthetic population mixes the two binary characteristics that shape the choice set: car ownership (drivers can use CAR; non-owners cannot, but may ride as CAR_PASS) and worker status (workers have a mandatory WORK episode on a fixed schedule; non-workers move freely). Each archetype is ~¼ of the population; the worker archetypes are split evenly across their schedules.
| archetype | n (N=200) | modes available | WORK schedules | identifies |
|---|---|---|---|---|
| car · worker | 48 | all 5 (CAR driver + PT + walk/bike) | 08–16, 10–15, 08–13, 12–22 (×12 each) | δ, α, β + CAR-vs-TRAIN commute → θ_travel |
| car · non-worker | 50 | all 5 | — (free) | β0/β1 shop & leisure; discretionary CAR-vs-TRAIN |
| no-car · worker | 50 | PT + walk/bike + CAR_PASS (no CAR driver) | 08–16, 10–15 (×25 each) | δ, α + WALK-vs-TRAIN commute → θ_travel |
| no-car · non-worker | 50 | PT + walk/bike + CAR_PASS | — (free) | β0/β1 shop & leisure; WALK-vs-TRAIN (bulk PT signal) |
Why this mix. Workers (car + no-car) supply the on/pre/post-schedule WORK observations that identify δ/α/β — the 12:00–22:00 late shift is realistic (evening-shift workers exist in real diaries) and is what generates the post-schedule WORK steps identifying β. Non-workers' discretionary shop/leisure trips identify the activity intercepts/slopes (β0, β1). Car owners contribute CAR-vs-TRAIN choices and no-car agents the WALK-vs-TRAIN crossover — together the wide mode-choice variation that identifies the single transport scale θ_travel. (no-car excludes the CAR driver mode only; CAR_PASS stays available, matching the production pipeline.)
| schedule | type | identifies |
|---|---|---|
| 08:00–16:00 | standard | δ / α (morning arrivals) |
| 10:00–15:00 | late/short | δ / α |
| 08:00–13:00 | half-day | δ / α |
| 12:00–22:00 | LATE SHIFT (realistic evening worker) | β — post-schedule lingering |
| param | θ* | description |
|---|---|---|
| delta | +0.0300 | WORK on-schedule marginal utility (per min) |
| alpha | +0.0010 | pre-schedule earliness penalty slope |
| beta | +0.0005 | post-schedule lateness penalty slope |
| beta1_shop | +0.5000 | SHOPPING attractiveness × P_open coefficient |
| beta0_shop | -0.2200 | SHOPPING base rate (per min) |
| beta1_leis | +0.4000 | LEISURE attractiveness × P_open coefficient |
| beta0_leis | -0.4200 | LEISURE base rate (per min) |
| theta_travel | +1.0000 | transport SCALE on whole mode utility |
State \(s=(t,\text{zone},\text{prev\_act},\text{dur},\text{mode},\text{hist})\); at each \(s\) choose action \(a\) (next activity, destination, mode). Value function (Bellman logsum; terminal \(\bar V(s_T)=0\)), choice probability, and log-likelihood:
$$\bar V(s)=\log\!\!\sum_{a\in A(s)}\exp\!\big(u(s,a;\theta)+\bar V(s'(s,a))\big),\qquad P(a\mid s)=\exp\!\big(\underbrace{u(s,a)+\bar V(s')}_{Q(s,a)}-\bar V(s)\big)$$ $$\mathcal L(\theta)=\sum_{n}\sum_{t}\big[Q(s_{nt},a_{nt};\theta)-\bar V(s_{nt};\theta)\big] =\sum_n\sum_t\log P(a_{nt}\mid s_{nt})$$HOME (reference): \(u_{\text{HOME}}=\mu_{\text{home}}\,\Delta t = 0\) since \(\mu_{\text{home}}\equiv 0\).
WORK/SCHOOL (piecewise-linear around individual schedule \([t_s,t_e]\)):
$$u_{\text{WORK}}(t)=\Delta t\cdot\begin{cases} \delta-\alpha\,(t_s-t) & t\lt t_s\quad\text{(earliness)}\\ \delta & t_s\le t\le t_e\quad\text{(on-schedule)}\\ \delta-\beta\,(t-t_e) & t\gt t_e\quad\text{(lateness)}\end{cases}$$SHOPPING / LEISURE (zone \(z\), time \(t\)):
$$u_{\text{SHOP}}(z,t)=\big(\beta_1^{\text{shop}}X_{\text{shop}}[z]\,P^{\text{shop}}_{\text{open}}[z,t]+\beta_0^{\text{shop}}\big)\Delta t, \qquad u_{\text{LEIS}}\ \text{analogously}$$TRAVEL (mode \(m\), \(o\!\to\! d\), depart \(t\)) — single transport scale:
$$u_{\text{TRAVEL}}=\theta_{\text{travel}}\big(\text{ASC}[m]+\beta_{\text{time}}\,TT_{o,d,m}+\beta_{\text{cost}}\,C_{o,d,m}\big)+2\,c_{\text{change}}+t_d+t_p$$Fixed HH-MNL coefficients \(\beta_{\text{time}}=-0.039/\text{min}\), \(\beta_{\text{cost}}=-0.00088/\text{¥}\); \(\text{ASC}\): PT(TRAIN/BUS)\(=0\) reference, CAR\(=+2.756\), WALK\(=+1.948\), BICYCLE\(=-0.029\), CAR\_PASS\(=+0.111\). \(c_{\text{change}}\) is charged at departure AND arrival (\(2c_{\text{change}}\)/trip). Grid-rounding terms \(t_d,t_p\) (with \(\text{split}=(\lceil TT/\Delta t\rceil\Delta t-TT)/2\)): \(t_d=\mu(\text{prev},t,o)\tfrac{\text{split}}{\Delta t}\), \(t_p=\max_{\text{act}}\mu(\text{act},t{+}TT,d)\tfrac{\text{split}}{\Delta t}\). Small net: OD times are multiples of \(\Delta t{=}30\) so \(t_d{=}t_p{=}0\), and \(c_{\text{change}}{=}0\) — travel reduces to \(\theta_{\text{travel}}(\text{ASC}+\beta_{\text{time}}TT+\beta_{\text{cost}}C)\).
| param | description | θ* | Mean | bias | SD | MeanSE | CP% | APB | verdict |
|---|---|---|---|---|---|---|---|---|---|
| delta | WORK on-schedule marginal utility (per min) | +0.0300 | +0.0309 | +0.0009 | 0.0046 | 0.0048 | 96.7 | 12.5% | ✓ recovered |
| alpha tiny θ* | pre-schedule earliness penalty slope | +0.0010 | +0.0010 | -0.0000 | 0.0001 | 0.0001 | 93.3 | abs 0.0000 | ✓ recovered |
| beta tiny θ* | post-schedule lateness penalty slope | +0.0005 | +0.0006 | +0.0001 | 0.0004 | 0.0008 | 100.0 | abs 0.0001 | ✓ recovered |
| beta1_shop | SHOPPING attractiveness × P_open coefficient | +0.5000 | +0.5059 | +0.0059 | 0.0210 | 0.0248 | 100.0 | 3.4% | ✓ recovered |
| beta0_shop | SHOPPING base rate (per min) | -0.2200 | -0.2229 | -0.0029 | 0.0106 | 0.0131 | 100.0 | 3.8% | ✓ recovered |
| beta1_leis | LEISURE attractiveness × P_open coefficient | +0.4000 | +0.4066 | +0.0066 | 0.0239 | 0.0214 | 96.7 | 5.3% | ✓ recovered |
| beta0_leis | LEISURE base rate (per min) | -0.4200 | -0.4275 | -0.0075 | 0.0268 | 0.0242 | 93.3 | 5.6% | ✓ recovered |
| theta_travel | transport SCALE on whole mode utility | +1.0000 | +1.0051 | +0.0051 | 0.0418 | 0.0505 | 100.0 | 3.4% | ✓ recovered |
CP = % reps with θ* ∈ θ̂ ± 1.96·SE. SD = empirical std across reps; MeanSE = mean BHHH SE. ✓ = bias within ~2 SD and CP∈[85,100]. α/β have θ*≈0.001 so APB% is unstable — read the absolute bias.
theta_travel shown by default; click legend to toggle. Green = converged, red = convergence-flag false (estimate still valid).
X = replication # (each of the 30 independent datasets generated at θ*). Y = (θ̂−θ*)/SD, the estimate's error in units of its own empirical SD. Why standardize: the 8 parameters span scales 0.001–0.5; dividing by each one's SD puts them on a single comparable axis. Read it as: a band centred on 0 (no drift = unbiased) with ≈95% of points within ±2 (= consistent with the reported SD).
The Exact estimator's sampling-variance floor — the irreducible spread that remains even with a perfectly-computed value function — shown for all 8 free parameters (toggle via the legend; theta_travel shown by default). SD is normalized to its value at N=50 so every parameter starts at 1.0 and shares the single dashed 1/√N reference. Each parameter's SD tracks that line — the √N-consistency signature of a correct MLE.
Quantified slope (5 N levels, median over well-identified params, log–log): measured slope = -0.468 (predicted −½ = −0.500), R² = 0.980. Standard MLE asymptotics predict exactly −½; the recovery experiment confirms it to within the Monte-Carlo noise of R=30 reps.
X = sample size N (person-days), log scale so the levels {50,…,1000} space evenly. Y = SD(N)/SD(N=50), each parameter's spread across the 30 reps relative to its own value at the smallest N, also log. Why normalize: absolute SDs differ ~200× across params (alpha ~1e-4 vs beta0_leis ~2e-2) — dividing by SD at N=50 starts every param at 1.0 so all 8 share the single dashed 1/√N line. On log–log, SD ∝ 1/√N is a straight line of slope −½; a curve hugging it = √N-consistency.
| param | SD@50 | SD@1000 | SD ratio | CP%@1000 |
|---|---|---|---|---|
| delta | 0.0103 | 0.0022 | 4.70× | 93.3 |
| alpha | 0.0002 | 0.0000 | 5.92× | 100.0 |
| beta | 0.0007 | 0.0001 | 6.01× | 100.0 |
| beta1_shop | 0.0623 | 0.0111 | 5.60× | 93.3 |
| beta0_shop | 0.0331 | 0.0058 | 5.66× | 96.7 |
| beta1_leis | 0.0371 | 0.0105 | 3.54× | 100.0 |
| beta0_leis | 0.0419 | 0.0118 | 3.55× | 96.7 |
| theta_travel | 0.0760 | 0.0302 | 2.52× | 80.0 |
√N-ideal SD ratio = √(1000/50) = 4.5×.
| N | Mean | bias | SD | MeanSE | CP% | conv |
|---|---|---|---|---|---|---|
| 50 | 1.0265 | +0.0265 | 0.0760 | 0.1138 | 100.0 | 30/30 |
| 100 | 0.9987 | -0.0013 | 0.0621 | 0.0711 | 100.0 | 28/30 |
| 200 | 1.0051 | +0.0051 | 0.0418 | 0.0505 | 100.0 | 20/30 |
| 500 | 1.0088 | +0.0088 | 0.0328 | 0.0311 | 93.3 | 30/30 |
| 1000 | 0.9971 | -0.0029 | 0.0302 | 0.0216 | 80.0 | 30/30 |
This floor is the reference for Report 2 (Exact vs SA vs RL). The approximation variance
of SA and RL is defined as the excess above this floor — Var(θ̂approx) −
Var(θ̂Exact) — and is expected to shrink as 1/K (SA) or 1/B (RL) as the approximation
budget grows. See docs/research/approximation_error_framework_20260624.md for the
full MSE decomposition and scaling-law derivation.
1D slice of the log-likelihood in each parameter (others held at θ*). A ∩-shape peaked at θ* proves the LL responds to that parameter ⇒ identifiable. Axes normalized: x=(θ−θ*)/half-width, y=LL−max, so every identifiable curve peaks at (0,0).
| param | LL peak | θ* | LL range | curvature |
|---|---|---|---|---|
| delta | +0.0300 | +0.0300 | 103.7 | concave ✓ |
| alpha | +0.0010 | +0.0010 | 52896.3 | concave ✓ |
| beta | +0.0005 | +0.0005 | 35200.7 | concave ✓ |
| beta1_shop | +0.5000 | +0.5000 | 4836.7 | concave ✓ |
| beta0_shop | -0.2200 | -0.2200 | 46898.5 | concave ✓ |
| beta1_leis | +0.4000 | +0.4000 | 13834.3 | concave ✓ |
| beta0_leis | -0.4200 | -0.4200 | 31201.4 | concave ✓ |
| theta_travel | +1.0000 | +1.0000 | 294.8 | concave ✓ |
X = (θ−θ*)/half-width: the parameter slid off its true value, normalized so each param's scan spans [−1,1] (truth at x=0). Y = LL−max(LL): the log-likelihood with its own peak shifted to 0. Why both normalized: params have different units and the LL has different magnitudes per param — this overlays all 8 on one axis. A ∩-shape peaking at (0,0) ⇒ the LL responds to that parameter ⇒ identifiable; a flat line would mean non-identification.
| Mode share (of trips) | θ* % | θ̂ % |
|---|---|---|
| WALK | 41.3 | 40.8 |
| TRAIN | 34.0 | 33.0 |
| CAR | 22.4 | 24.0 |
| CAR_PASS | 1.6 | 1.5 |
| BICYCLE | 0.4 | 0.4 |
| BUS | 0.2 | 0.3 |
| Activity-step share | θ* % | θ̂ % |
|---|---|---|
| HOME | 43.1 | 43.0 |
| SHOP | 22.3 | 22.0 |
| LEIS | 18.5 | 18.7 |
| TRAVEL | 10.9 | 11.0 |
| WORK | 5.3 | 5.4 |
Fraction of person-steps at each activity by hour (stacked = 1.0), θ* population.
X = time of day (hour). Y = fraction of person-steps in each activity, stacked to 1.0 each hour (= 100% of agents). No normalization trick here — unlike the diagnostic plots above, the behavioral charts use natural units because the point is realism, not cross-parameter comparison: WORK fills daytime, HOME the night, shop/leisure rise toward their P_open peaks.
| sequence | count |
|---|---|
| HOME→WORK→SHOP→LEIS→HOME | 130 |
| HOME→WORK→SHOP→LEIS→SHOP→LEIS→HOME | 103 |
| HOME→SHOP→WORK→SHOP→LEIS→HOME | 76 |
| HOME→SHOP→LEIS→HOME | 64 |
| HOME→SHOP→LEIS→SHOP→LEIS→HOME | 56 |
| HOME→SHOP→LEIS→WORK→HOME | 28 |
u=θ·(ASC+β_time·t+β_cost·c) (PT=0 reference);
the separate PT level ASC_pt_adj is retired. This removes the scale-vs-level ridge that biased
theta_travel +13.8% (|bias|/SE=16.6) under the old two-parameter spec → now +0.0051
(3.4% APB).gtol (5e-4×steps) so the flag tracks the estimates.Supporting material:
docs/estimation/reports/optionA_recovery_confirmation_20260619.md — single-scale
production-fidelity (Δt=15) R=1 confirmation on the heavy toy.estimation/diagnostics/fisher_ridge_at_truth.py — scale-vs-level ridge diagnosis.estimation/diagnostics/check_schedule_variation.py — verifies the late shift
produces post-schedule (β) observations.Next: finish N-sweep curve → fork to Branch A (HH real-data estimation) and/or Branch B (Exact-vs-SA variance experiments), both de-risked by this passed gate.