Circular Image

S. Hess

info

Please Note

3 records found

Journal article (2026) - Georges Sfeir, Gabriel Nova, Stephane Hess, Sander van Cranenburgh
Large Language Models (LLMs) are becoming widely used to support various workflows across different disciplines, yet their potential in discrete choice modelling remains relatively unexplored. This work examines the potential of LLMs as assistive agents in the specification and, where technically feasible, estimation of Multinomial Logit models. We implement a systematic experimental framework involving twelve versions of seven leading LLMs (ChatGPT, Claude, DeepSeek, Gemini, Gemma, Llama, and Mistral) evaluated under five experimental configurations. These configurations vary along three dimensions: (i) modelling goal (suggesting vs. suggesting and estimating MNL models); (ii) prompting strategy (Zero-Shot vs. Chain-of-Thoughts (CoT)); and (iii) information availability (full dataset vs. data dictionary summarising variable names and types). Each specification suggested by the LLMs is implemented, estimated, and evaluated based on goodness-of-fit metrics, behavioural plausibility, and model complexity. Our findings reveal that proprietary LLMs can generate valid and behaviourally sound utility specifications, particularly when guided by structured prompts (CoT). Open-weight models such as Llama and Gemma struggled to produce meaningful specifications. Notably, some LLMs performed better when provided with just data dictionary, suggesting that limiting raw data access may enhance internal reasoning capabilities. Among all LLMs, GPT o3, operating in an agentic setting, was uniquely capable of correctly estimating its own specifications by executing self-generated code. Overall, the results demonstrate both the promise and current limitations of LLMs as assistive agents in discrete choice modelling, not only for model specification but also for supporting modelling decision and estimation, and provide practical guidance for integrating these tools into choice modellers’ workflows. ...

Model averaging for out-of-distribution forecasting

Journal article (2026) - Stephane Hess, Sander van Cranenburgh
Travel behaviour modellers have an increasingly diverse set of models at their disposal, ranging from traditional econometric structures to models from mathematical psychology and data-driven approaches from machine learning. A key question arises as to how well these different models perform in forecasting, especially when considering trips of different characteristics from those used in estimation, i.e. out-of-distribution prediction, and whether better predictions can be obtained by combining insights from the different models. We focus on trip distance as a key example of a variable where the application context might go beyond the estimation data. Across two case studies, we show that while data-driven approaches excel in predicting mode choice for trips within the distance bands used in estimation, beyond that range, the picture is fuzzy. To leverage the relative advantages of the different model families and capitalise on the notion that multiple ‘weak’ models can result in more robust models, we put forward the use of a model averaging approach that allocates weights to different model families as a function of the distance between the characteristics of the trip for which predictions are made, and those used in model estimation. Overall, we see that the model averaging approach gives larger weight to models with stronger behavioural or econometric underpinnings the more we move outside the interval of trip distances covered in estimation. Across both case studies, we show that our model averaging approach obtains improved performance both on the estimation and test data, and crucially also when predicting mode choices for trips of distances outside the range used in estimation. While our initial proof of concept focuses on trip distance as a single trip characteristic to quantify the degree of an observation being out-of-distribution, we also provide initial insights into extending this to a multivariate context, using a Gower distance metric. ...
Journal article (2025) - Gabriel Nova, C. Angelo Guevara, Stephane Hess, Thomas O. Hancock
Discrete choice analysis aims to understand and predict decision-makers’ behaviour, a goal that is crucial across several disciplines, including transportation. This type of analysis has relied predominantly on static representations of preferences, principally through the Random Utility Maximisation (RUM) model, due to its ease of implementation, economic interpretability, and statistical formality. However, this model assumes that individuals possess complete information about all attributes of alternatives and that they can process and recall this information instantaneously, which may not align with actual human behaviour. In contrast, the Decision Field Theory (DFT) model from mathematical psychology explicitly incorporates the repeated scrutiny of attributes and recall effects within the decision-making process, which enables it to model attention weights, but lacks microeconomic interpretability and clear statistical parameter identification. This paper introduces the RUM-DFT model, which seeks to integrate strengths of both approaches. Through Monte Carlo simulations, the proposed model is shown to be able to: (i) recover parameters related to the deliberation process, (ii) replicate the dynamic behaviour of utilities during deliberation as observed in practice, (iii) maintain economic interpretability by estimating coefficients that can be used to calculate the marginal indirect utilities, and (iv) highlight the pitfalls of using a RUM model that disregards the true dynamics of data generation process. The SwissMetro case study is employed also to evaluate the RUM-DFT model using a real-world dataset, demonstrating the viability and superior goodness-of-fit of the proposed model. ...