Toward Quantum Reinforcement Learning for Automated Neural Architecture Optimization in Vision-Based Structural Health Monitoring

Conference Paper (2026)
Author(s)

Álvaro Romero Mato (TU Delft - Aerospace Engineering)

Boyang Chen (TU Delft - Aerospace Engineering)

Nathan D. Eskue (TU Delft - Aerospace Engineering)

Vahid Yaghoubi (TU Delft - Aerospace Engineering)

Research Group
Group Eskue
DOI related publication
https://doi.org/10.58286/33912 Final published version
More Info
expand_more
Publication Year
2026
Language
English
Research Group
Group Eskue
Article number
1199
Publisher
NDT.net
Event
12th European Workshop on Structural Health Monitoring 2026 (2026-07-07 - 2026-07-10), Pierre Baudis Convention Centre, Toulouse, France
Page Views
38
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

Neural Architecture Search (NAS) involves exploring large and combinatorial search spaces, for which numerous approaches have been proposed, including reinforcement learning, evolutionary algorithms, Bayesian optimization, and gradient-based methods among others. Despite these advances, efficient exploration remains a key challenge, particularly as the search space scales combinatorially. Motivated by prior work based on Q-learning strategies, which are limited in large spaces, this work investigates a Quantum Reinforcement Learning (QRL) framework based on probabilistic policy optimization. We propose a QRL approach where a parameterized quantum circuit represents a policy over candidate architectures, and is trained to maximize the expected reward defined by validation accuracy. This method is compared against two baselines: Random Search and a Classical Reinforcement Learning approach using policy gradient with a Bernoulli distribution. All methods are evaluated under identical budgets using the NATS-Bench benchmark, enabling efficient and controlled comparison without full model training. The results demonstrate that both QRL and Classical RL learn meaningful search policies, consistently outperforming Random Search. QRL achieves competitive performance with respect to Classical RL, confirming its ability to effectively explore the architecture space. We further analyze the impact of key hyperparameters to identify practical trade-offs between accuracy, model complexity, and computational cost. Overall, the results demonstrate that QRL is a viable approach for discrete architecture search, capable of learning effective policies in large combinatorial spaces, while indicating clear directions for improving its efficiency and scalability.