Split CNN Inference on Networked Microcontrollers

Conference Paper (2026)
Author(s)

Junyu Lu (Student TU Delft)

Shashwath Suresh (Student TU Delft)

Hao Liu (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Qi Hong (Student TU Delft)

Qing Wang (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Research Group
Networked Systems
DOI related publication
https://doi.org/10.1109/WoWMoM69805.2026.00030 Final published version
More Info
expand_more
Publication Year
2026
Language
English
Research Group
Networked Systems
Pages (from-to)
163-172
Publisher
IEEE
ISBN (electronic)
9798331562991
Event
27th IEEE International Symposium on a World of Wireless, Mobile and Multimedia Networks, WoWMoM 2026 (2026-06-16 - 2026-06-19), Bologna, Italy
Downloads counter
24
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

Running deep neural networks on microcontroller units (MCUs) is severely constrained by limited memory resources. While TinyML techniques reduce model size and computation, they often fail in practice due to excessive peak Random Access Memory (RAM) usage during inference, dominated by intermediate activations. As a result, many models remain infeasible on standalone MCUs. In this work, we present a finegrained split inference system for networked MCUs that enables collaborative inference of Convolutional Neural Networks (CNN) models across multiple devices. Our key insight is that breaking the memory bottleneck requires splitting inference at sub-layer granularity rather than at layer boundaries. We reinterpret pretrained models to enable kernel-wise and neuron-wise partitioning, and distribute both model parameters and intermediate activations across multiple MCUs. A lightweight, resource-aware coordinator orchestrates the inference across MCU devices with heterogeneous resources. We implement the proposed system on a real testbed and evaluate it on up to 8 MCUs using MobileNetV2, a representative CNN model. Our experimental results show that CNN models infeasible on a single MCU can be executed across networked MCUs, reducing the per-MCU peak RAM usage while maintaining the practical end-to-end inference latency. All the source code of this work can be found here: https://github.com/shashsuresh/Split-Inference-on-MCUs

Files

– Personal use only – Dutch Copyright Act (Article 25fa)
warning

File under embargo until 30-01-2027