Deep reinforcement learning for optimizing mobile depots allocation and hybrid order dispatch in last-mile delivery under multi-participants behavioral uncertainty

Journal Article (2026)
Author(s)

Shixuan Hou (Western University, IE University)

Chengming Hu (McGill University)

Annabathuni Sandeep Chowdary (Bell Canada)

Bissan Ghaddar (IE University, Western University)

Joe Naoum-Sawaya (IE University, Western University)

Jie Gao (TU Delft - Civil Engineering & Geosciences)

Research Group
Transport, Mobility and Logistics
DOI related publication
https://doi.org/10.1016/j.tre.2026.105170 Final published version
More Info
expand_more
Publication Year
2026
Language
English
Research Group
Transport, Mobility and Logistics
Journal title
Transportation Research Part E: Logistics and Transportation Review
Volume number
216
Article number
105170
Downloads counter
7
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

As last-mile delivery demand continues to grow and customer preferences become increasingly diverse, traditional single-mode logistics systems are inadequate in addressing the complexity and variability of modern urban delivery environments. In this paper, we propose an innovative hybrid last-mile delivery system that leverages mobile depots as intermediate hubs, integrating crowd-shipping, and customer self-pickup, to enhance cost-effectiveness. However multiple stakeholders in the system have conflicting interests, ignoring their behavioral uncertainties can impair operational efficiency and increase delivery costs. For example, customers and crowd-shippers may compete for the limited storage capacity of mobile depots, and crowd-shippers may be unwilling to deliver orders that customers decline to pick up. Specifically, the study captures crowd-shippers’ order-acceptance behavior and customers’ self-pickup choice through two binomial logit models, and embeds these behavioral components into a stochastic mixed-integer programming model. This model jointly optimizes mobile depots allocation and order dispatch decisions, aiming to minimize the expected total operational cost. Considering computational complexity of this model, we propose PRiME-DQN, a reinforcement learning framework that incorporates behavior-prioritized resource filtering, action biasing, and model-efficient temperature tuning to accelerate convergence and improve solution quality. A case study on Amazon’s last-mile delivery in Seattle shows that the proposed hybrid model achieves up to 49.5% cost reduction in low-shipper settings and 11.3% in higher-shipper settings compared to a crowd-shipping system, while also consistently surpassing the self-pickup baseline by 3–6%. Extensive scaling experiments show that PRiME-DQN produces high-quality feasible solutions efficiently. For the largest instances, it achieves delivery costs 1.97% lower than Gurobi’s best incumbent solution obtained within the 3,600-second time limit, while reducing the solution time to approximately 1400 seconds.