Behavioral Validation — Västberg §4.4

HH θ̂ · 3120 workers · R=200 reps/person · 1,478,000 simulated paths · 2026-07-19

§0 — Goodness-of-fit headline

0.9126
McFadden ρ²
0.9125
Adjusted ρ̄²
23
Free params
3,120
Workers
200
Reps/person
Excellent fit — ρ² ≥ 0.25 (Ben-Akiva & Lerman 1985).
QuantityValueNotes
LL(θ̂)-69,967.902At estimated parameters
LL(0)-800,278.982Null: estimated params = 0 (HH MNL ASCs fixed)
ρ² = 1 − LL(θ̂)/LL(0)0.9126McFadden (1974)
ρ̄² = 1 − (LL(θ̂)−K)/LL(0)0.9125K=23 free params

LL(θ̂) and LL(0) are evaluated over all 7390 persons (26 groups, workers + non-workers), matching the mixed estimation sample, so the MLE property LL(θ̂) ≥ LL(0) holds. The null sets all estimated params (δ, α, β, β₀_shop, β₁_shop, β₀_leis, β₁_leis, θ_travel) to 0 while keeping the fixed HH MNL mode constants (ASCs, β_time, β_cost). It is not the uniform-path null.

§1 — Mode & activity shares

Blue = observed diaries (worker persons only); Orange = simulated under θ̂ (R=200 replications per worker). Mode shares = % of travel trips (per leg, count-weighted on both sides). Activity time shares exclude TRAVEL from the denominator on both sides — TRAVEL is the move-action between activity-states, not an activity itself; it is reported separately as travel-share (% of all 15-min slots).

Mode shares (% of travel trips, per leg)

ModeObsSimΔ
CAR71.8%82.9%+11.0%
TRAIN0.9%0.2%-0.7%
BUS2.2%1.7%-0.5%
WALK17.2%9.0%-8.1%
BICYCLE7.9%6.3%-1.7%

Activity time shares (TRAVEL excluded from denominator)

TRAVEL excluded from denominator on both sides — it is the move-action, not an activity-state. Reported separately as travel-share below.

ActivityObsSimΔ
HOME72.6%68.9%-3.7%
WORK21.7%21.0%-0.8%
SHOPPING1.8%3.5%+1.8%
LEISURE4.0%6.6%+2.7%

Travel time share

StatisticObsSimΔ
TRAVEL (of all 15-min slots)4.7%4.5%-0.2%

§2 — Activity start-time distributions (Västberg Fig 1)

Distribution of the start time of each activity episode, by hour of day. Orange bar = simulated count; blue line = observed (scaled to sim peak). Dashed lines mark means.

§3 — Trips & tours per person-day (Västberg Fig 2)

Distributions of trips/day, tours/day, and trips/tour. Orange bar = simulated; blue line = observed (scaled to sim peak).

§4 — Distribution of Trip Lengths by Mode (Västberg Fig 3)

Solid blue = observation; dashed dark red = simulation. Y-axis: #trips / #individuals. Method: 1-minute bins → 5-minute moving-average (Västberg 2020 §4.4). Observed: trip_data_higashihiroshima.csv (arrtime − deptime). Simulated: EndTime − StartTime of TRAVEL rows (15-min grid → step-discretised).

§5 — Aggregate statistics (Västberg Table 4)

Per-person scalar averages over the 3120-worker sample. LL(θ̂) and ρ² are computed over all 7390 persons (26 groups incl. 2 non-worker groups) — the same mixed sample used in estimation — so that the MLE property LL(θ̂) ≥ LL(0) holds. Behavioral figures (§1–§4) are worker-only.

StatisticObservedSimulatedΔ
Trips / person2.431.64-0.79
Tours / person1.070.54-0.52
Trips / tour2.283.05+0.77
Avg travel time22.02 min29.52 min+7.50 min
WORK start (h)8.42h9.13h+0.71h
TRAVEL start (h)12.90h14.82h+1.91h

Green |Δ| < 0.2 unit; red |Δ| > 1.0 unit.

§5.5 — Simulation-fidelity diagnostic (work schedule pull)

Diagnostic: does the model correctly simulate each worker doing their scheduled work, or does logit noise overwhelm the schedule pull δ=0.03814? A residual gap between sim work hours and scheduled hours after this check is F5 — genuine model misfit (weak δ or absent activity constants), not a simulation or data-loading bug.

For each worker group: scheduled hours = end−start of work window; simulated work hours = mean over R replications; any-work fraction = fraction of reps that included ≥1 WORK step. Ratio > 0.85 is green; < 0.50 is red (indicates weak schedule pull δ or excessive logit noise).

GroupPaths (N×R)Sched hSim mean h RatioAny-work %
((), (), True)218,800—0.0h—0%
((), (), False)635,200—0.0h—0%
(('WORK',), (('WORK', 480, 1200),), True)28,80012.0h12.3h1.02100%
(('WORK',), (('WORK', 480, 1200),), False)19,20012.0h12.1h1.01100%
(('WORK',), (('WORK', 630, 1200),), True)43,4009.5h9.5h1.00100%
(('WORK',), (('WORK', 630, 1200),), False)48,0009.5h9.6h1.01100%
(('WORK',), (('WORK', 480, 1080),), True)21,80010.0h10.3h1.03100%
(('WORK',), (('WORK', 480, 1080),), False)26,20010.0h10.2h1.02100%
(('WORK',), (('WORK', 420, 1080),), True)16,80011.0h11.3h1.03100%
(('WORK',), (('WORK', 420, 1080),), False)21,20011.0h11.2h1.02100%
(('WORK',), (('WORK', 450, 840),), True)4,0006.5h6.7h1.03100%
(('WORK',), (('WORK', 450, 840),), False)50,8006.5h6.7h1.03100%
(('WORK',), (('WORK', 450, 1200),), True)28,60012.5h12.8h1.02100%
(('WORK',), (('WORK', 450, 1200),), False)22,60012.5h12.7h1.02100%
(('WORK',), (('WORK', 570, 990),), True)8,6007.0h7.1h1.02100%
(('WORK',), (('WORK', 570, 990),), False)39,8007.0h7.2h1.03100%
(('WORK',), (('WORK', 540, 1080),), True)16,8009.0h9.2h1.02100%
(('WORK',), (('WORK', 540, 1080),), False)44,6009.0h9.2h1.03100%
(('WORK',), (('WORK', 570, 810),), True)7,6004.0h4.3h1.07100%
(('WORK',), (('WORK', 570, 810),), False)43,4004.0h4.4h1.10100%
(('WORK',), (('WORK', 480, 990),), True)4,6008.5h8.8h1.04100%
(('WORK',), (('WORK', 480, 990),), False)33,8008.5h8.8h1.03100%
(('WORK',), (('WORK', 450, 990),), True)9,6009.0h9.3h1.04100%
(('WORK',), (('WORK', 450, 990),), False)43,6009.0h9.2h1.03100%
(('WORK',), (('WORK', 480, 840),), True)3,0006.0h6.2h1.04100%
(('WORK',), (('WORK', 480, 840),), False)37,2006.0h6.3h1.04100%

§6 — Measurement vs Model: what was fixed vs what remains

F1–F4 (measurement asymmetries — fixed in this report)

F0 — McFadden ρ² sample discipline (fixed)

The original ρ² = −0.864 was computed on a worker-only subset (3,025p) while θ̂ was the MLE on the full mixed sample (5,218p workers+non-workers). The MLE guarantee LL(θ̂) ≥ LL(0) only holds on the estimation sample. Additionally, LL(0) was computed on a separately rebuilt graph, risking state-index misalignment. Both are fixed: LL(θ̂) and LL(0) now use all 7390 persons + the same graph-state indices (recompute_graph_utilities on the existing graph object, not rebuild).

F5 — Genuine model misfit (structural): activity shares

After all F1–F4 fixes, a gap remains: simulated WORK time share ≈21.0% vs observed 21.7% (-0.8pp gap). This is not a measurement artifact — it is a structural property of the K=10 "Option A" specification:

F5b — Genuine model misfit (structural): transit collapse

Simulated transit mode share ≈ 0% vs observed ≈ 3.4% (BUS + TRAIN). This is not a code bug or OD-data gap; it is a quantified consequence of the K=10 spec. Diagnostic run estimation/experiments/mode_collapse_diagnostic.py establishes the following:

ModeASC (raw)Effective ASC (×θ=0.408)Median TTu_travel (per trip)
CAR+2.756+1.12430 min+0.54 (positive — ASC beats time+cost)
BUS0.0000.00036 min−0.81 (negative — no ASC to offset travel penalty)
TRAIN0.0000.00050 min (0.4% feasible)−1.04 (rare service + no ASC)

The "θ_travel lifts PT share" claim in the code comment is directionally correct — at θ_travel=1.0 (full MNL), CAR ASC=2.756; at θ_travel=0.408 it is 1.124 — but the absolute PT share remains near zero because BUS still starts at 0 and has negative total utility on median trips.

Decision implication (mode): matching the observed 3.4% transit share requires at minimum a free PT constant (one extra param, "Option B" / K=11 lite) or full activity/mode ASC estimation (Västberg §4.4). See §7.

Decision implication (activity shares): to match activity-time-share marginals, the model spec needs activity-/tour-/trip-level ASCs estimated jointly with the DP parameters. That is a model-spec change; it is not in scope of this validation.

§7 — Caveats and scope