Towards modeling of Computation-in-Memory аccelerators
Latency estimation of distributed-memory architectures
D.A. Barantiev (TU Delft - Electrical Engineering, Mathematics and Computer Science)
G. Gaydadjiev – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)
M. Naderan-Tahan – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)
More Info
expand_more
Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.
Abstract
The growing data demands of industry applications have led to Von Neumann computing systems increasingly facing a performance bottleneck in the compute-memory data path because of the memory wall problem. Computation-in-Memory (CiM) is a novel type of architecture that tries to address this by combining memory storage and computing in one CiM macro device to achieve a more efficient dataflow.
However, CiM-centric systems still require theoretical analysis through the use of analytical modeling due to a broad design space and practical hardware limitations. One unexplored modeling scope is multi-macro accelerators with a distributed-memory architecture, where inter-macro data messages are necessary to combine and redistribute results. In this thesis, we present 3DCLM - a latency model
for end-to-end execution of 3D convolutional layers on multi-macro distributed-memory accelerators. In the absence of a good real-life reference for model validation, we design CIMsim-distributed, an event-based simulator of the target CiM-centric accelerator system, which accounts for local and inter-
macro data movement as well as the effects of traffic congestion. We compare predicted and simulated latencies to assess 3DCLM’s ability to capture how key parameters influence overall execution and the relative importance of compute and data movement stages. The results show that 3DCLM can predict the overall impact of convolution input channels and CiM macro columns, but often fails to represent
the importance of memory movement costs, both absolute and proportional to total execution latency. Consequently, it can be practically useful for constrained Design Space Exploration of CiM-centric systems, and future iterations should extend its model of critical data movement operations.
Files
File under embargo until 31-12-2026