Model-assisted multi-agent reinforcement learning for collaborative synchromodal transport re-planning under service time uncertainty

Journal Article (2026)
Author(s)

Xiaoyu Luo (Southwest Jiaotong University)

Yimeng Zhang (Southwest Jiaotong University)

Mi Gan (Southwest Jiaotong University)

Xiaobo Liu (Southwest Jiaotong University)

Bilge Atasoy (TU Delft - Mechanical Engineering)

Research Group
Transport Engineering and Logistics
DOI related publication
https://doi.org/10.1016/j.tre.2026.105051 Final published version
More Info
expand_more
Publication Year
2026
Language
English
Research Group
Transport Engineering and Logistics
Journal title
Transportation Research Part E: Logistics and Transportation Review
Volume number
214
Article number
105051
Downloads counter
8
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

Synchromodal transport operations frequently require rapid re-planning in response to uncertain terminal service times, yet such decisions are typically made by multiple carriers with heterogeneous objectives and decentralized control. This study investigates collaborative re-planning in synchromodal transport networks under service-time uncertainty from a multi-carrier perspective. We formulate the problem as a decentralized decision-making process in which carriers must coordinate re-planning actions while preserving individual rationality. To address this challenge, we propose a model-assisted multi-agent reinforcement learning (MARL) framework that combines heuristic plan generation with learning-based carrier coordination. An adaptive large neighborhood search (ALNS) heuristic is used to generate feasible candidate re-planning options, while carriers are modeled as learning agents that decide whether to accept or reject these options through a consensus mechanism. Rather than directly optimizing a centralized model, carriers learn acceptance policies based on local observations and realized outcomes, without requiring prior knowledge of service-time distributions. An event-triggered and regret-based learning scheme is introduced to align global delay mitigation with carrier-specific cost preferences under uncertainty. Extensive computational experiments on a realistic synchromodal transport network demonstrate that the proposed framework significantly improves re-planning performance compared with benchmark strategies. The results show that decentralized, heterogeneous multi-agent coordination can outperform centralized and homogeneous approaches, particularly in scenarios with high uncertainty and avoidable delays. Mechanism and ablation analyses further reveal the structural importance of multi-agent decomposition, carrier heterogeneity, and the temporal allocation of decision preferences along the transport chain. Additional analyses examine sensitivity to key reward hyperparameters and negotiation round limits, robustness under alternative disruption distributions and network-graph perturbations, and controlled extension settings with more carriers and transshipments. The proposed approach provides a practical decision-support methodology for collaborative synchromodal transport re-planning under uncertainty.

Files

Taverne
warning

File under embargo until 06-01-2027