📊 Research Integration: This document demonstrates how temporal difference (TD) learning methods from reinforcement learning can be applied to estimate Dynamic Discrete Choice Models (DDCM) for activity-based travel demand, specifically addressing the computational challenges of high-dimensional state spaces in the Västberg model.


Overview

Temporal-Difference estimation applies reinforcement learning techniques to estimate structural parameters in Dynamic Discrete Choice models, particularly when state spaces are continuous or high-dimensional, avoiding the need to estimate transition densities and repeatedly solve dynamic programming problems.


Core Problem & Solution

The Challenge

Traditional CCP methods require estimating transition density and choice probabilities , then solving Bellman equations to estimate , which is computationally prohibitive with continuous states.

The TD Approach

Uses observed transitions to directly approximate value functions through two main methods:

  • Semi-gradient Method: Fits linear model to TD errors (fast, closed-form)
  • Approximate Value Iteration (AVI): Iteratively improves approximations using any ML method (flexible)

Mathematical Framework

DDC Model Structure

Utility Function:

Agent’s Problem:

CCP Inversion Theorem

Proposition 1 establishes that the mapping from value function differences to choice probabilities is invertible:

This allows expressing value functions without backward recursion.


TD Estimation Methods

Method 1: Linear Semi-Gradient

TD Error Definition:

Functional Approximation:

Closed-Form Solution:

Advantages:

  • No transition density needed
  • Simple computation (matrix inversion of dimension k × k)
  • Fast closed-form solution

Method 2: Approximate Value Iteration

Iterative Formula:

Key Properties:

  • Can use any ML method (random forests, neural networks, LASSO)
  • Each iteration is a standard prediction problem
  • Convergence after J ≈ ln(n) iterations

Locally Robust Estimation

The Correction Term

To achieve √n-consistency, add correction terms that orthogonalize the estimating equation:

This ensures small errors in estimating h have no first-order effect on estimating θ.


Application to Västberg DDCM

Model Specifications

State Space (x_t):

  • Current location: 1,240 locations in Stockholm
  • Current time: continuous, discretized to 10-minute intervals
  • Activity history
  • Mode of last trip

Action Space (a_t = (d, l, m)):

  • Activity type: home, work, shopping, social, travel
  • Destination location
  • Travel mode: walk, bike, car, PT

Computational Advantages

High-Dimensional State Spaces:

  • Basis functions automatically generalize across locations
  • No need to visit every state in data
  • Approximation quality depends on smoothness, not dimensionality

Continuous Time:

  • Continuous basis functions naturally handle continuous time
  • No discretization artifacts
  • Captures smooth time-of-day variations

No Transition Density Estimation:

  • Only requires observed (x, a, x’) tuples
  • Avoids parametric misspecification of transitions

Basis Function Construction

Spatial Component

Temporal Component

Combined Basis

With typical dimensions, total k = 20,000, requiring only one matrix inversion.


Theoretical Guarantees

Convergence Theorem (Semi-Gradient)

Approximation Error:

Estimation Error:

Convergence Theorem (AVI)

After J iterations:


Two-Stage Estimation Procedure

Stage 1: Estimate Value Functions

  1. Collect data: {(a_i, x_i, a’_i, x’i)}{i=1}^n from travel diaries
  2. Estimate ĥ(a, x) and ĝ(a, x) using semi-gradient or AVI
  3. Compute v̂(a, x) = ĥ(a, x) + ĝ(a, x)

Stage 2: Estimate Structural Parameters

Maximize pseudo-likelihood:

where:


Extensions

Unobserved Heterogeneity

Combine TD with EM algorithm to handle discrete heterogeneity in preferences, different discount factors, and unobserved activity constraints.

Dynamic Games

Directly observe (a_i, a_{-i}, x, x’) in data without need for integration over other players’ actions.

Computational Scalability

  • Embarrassingly parallel: each observation contributes independently
  • Online learning: can update estimates as new data arrives
  • Near-linear speedup with number of processors

Key Insight

The observed next state x’ already contains information about both the transition density and choice probabilities; TD methods exploit this by using x’ as a sample from K(·|a, x), avoiding explicit density estimation.