📊 Research Integration: This document demonstrates how temporal difference (TD) learning methods from reinforcement learning can be applied to estimate Dynamic Discrete Choice Models (DDCM) for activity-based travel demand, specifically addressing the computational challenges of high-dimensional state spaces in the Västberg model.
Overview
Temporal-Difference estimation applies reinforcement learning techniques to estimate structural parameters in Dynamic Discrete Choice models, particularly when state spaces are continuous or high-dimensional, avoiding the need to estimate transition densities and repeatedly solve dynamic programming problems.
Core Problem & Solution
The Challenge
Traditional CCP methods require estimating transition density and choice probabilities , then solving Bellman equations to estimate , which is computationally prohibitive with continuous states.
The TD Approach
Uses observed transitions to directly approximate value functions through two main methods:
- Semi-gradient Method: Fits linear model to TD errors (fast, closed-form)
- Approximate Value Iteration (AVI): Iteratively improves approximations using any ML method (flexible)
Mathematical Framework
DDC Model Structure
Utility Function:
Agent’s Problem:
CCP Inversion Theorem
Proposition 1 establishes that the mapping from value function differences to choice probabilities is invertible:
This allows expressing value functions without backward recursion.
TD Estimation Methods
Method 1: Linear Semi-Gradient
TD Error Definition:
Functional Approximation:
Closed-Form Solution:
Advantages:
- No transition density needed
- Simple computation (matrix inversion of dimension k × k)
- Fast closed-form solution
Method 2: Approximate Value Iteration
Iterative Formula:
Key Properties:
- Can use any ML method (random forests, neural networks, LASSO)
- Each iteration is a standard prediction problem
- Convergence after J ≈ ln(n) iterations
Locally Robust Estimation
The Correction Term
To achieve √n-consistency, add correction terms that orthogonalize the estimating equation:
This ensures small errors in estimating h have no first-order effect on estimating θ.
Application to Västberg DDCM
Model Specifications
State Space (x_t):
- Current location: 1,240 locations in Stockholm
- Current time: continuous, discretized to 10-minute intervals
- Activity history
- Mode of last trip
Action Space (a_t = (d, l, m)):
- Activity type: home, work, shopping, social, travel
- Destination location
- Travel mode: walk, bike, car, PT
Computational Advantages
High-Dimensional State Spaces:
- Basis functions automatically generalize across locations
- No need to visit every state in data
- Approximation quality depends on smoothness, not dimensionality
Continuous Time:
- Continuous basis functions naturally handle continuous time
- No discretization artifacts
- Captures smooth time-of-day variations
No Transition Density Estimation:
- Only requires observed (x, a, x’) tuples
- Avoids parametric misspecification of transitions
Basis Function Construction
Spatial Component
Temporal Component
Combined Basis
With typical dimensions, total k = 20,000, requiring only one matrix inversion.
Theoretical Guarantees
Convergence Theorem (Semi-Gradient)
Approximation Error:
Estimation Error:
Convergence Theorem (AVI)
After J iterations:
Two-Stage Estimation Procedure
Stage 1: Estimate Value Functions
- Collect data: {(a_i, x_i, a’_i, x’i)}{i=1}^n from travel diaries
- Estimate ĥ(a, x) and ĝ(a, x) using semi-gradient or AVI
- Compute v̂(a, x) = ĥ(a, x) + ĝ(a, x)
Stage 2: Estimate Structural Parameters
Maximize pseudo-likelihood:
where:
Extensions
Unobserved Heterogeneity
Combine TD with EM algorithm to handle discrete heterogeneity in preferences, different discount factors, and unobserved activity constraints.
Dynamic Games
Directly observe (a_i, a_{-i}, x, x’) in data without need for integration over other players’ actions.
Computational Scalability
- Embarrassingly parallel: each observation contributes independently
- Online learning: can update estimates as new data arrives
- Near-linear speedup with number of processors
Key Insight
The observed next state x’ already contains information about both the transition density and choice probabilities; TD methods exploit this by using x’ as a sample from K(·|a, x), avoiding explicit density estimation.