FROST: Forward-Forward Round-Robin Optical-Compatible Softmax-Free Transformer
D. Krylov (TU Delft - Electrical Engineering, Mathematics and Computer Science)
S. Tan – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)
D.M.J. Tax – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)
R.A. Norte – Graduation committee member (TU Delft - Mechanical Engineering)
More Info
expand_more
Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.
Abstract
Training state-of-the-art AI models is growing far more costly than digital hardware can sustain, motivating analog photonic substrates that perform matrix multiplication directly in the propagation of light. Inference on such hardware is established, but training is not: backpropagation requires a backward pass and exact gradients that an analog chip cannot natively provide. Forward-only, backpropagation-free learning offers a way around this, yet it has been demonstrated only for multilayer perceptrons and convolutional networks, never for the Transformer, the architecture that now dominates AI and whose self-attention couples every position across the sequence, resisting purely local objectives.
This thesis asks whether a Transformer can be trained using only the forward-pass operations a photonic substrate provides, and characterizes what such training costs. The proposed method combines a layer-wise Forward-Forward prototype-based objective, directional-derivative gradient estimation (with no backward pass and no automatic differentiation), and a softmax-free Spherical attention adapted from the Kramers-Kronig kernel, with the four attention projections trained one at a time in a round-robin schedule; six training variants are compared across seven vision and sequence tasks to isolate the gradient estimator and the update schedule. The answer is affirmative: the fully forward-only variant trains a Transformer to a useful operating point on six of the seven tasks. Locality, not the forward-only gradient alone, is what makes this possible, and the depth-resilience of local learning is shown to extend, conditionally, to self-attention. The remaining gap to backpropagation has three sources, two inherent to forward-only learning (local credit assignment and gradient-estimation variance) and one architectural (the softmax-free attention cannot form the sharp selection that content-addressed retrieval requires). Because every operation reduces to a forward pass and a measured scalar loss, the method is a candidate for in-situ training on a photonic chip, the validation step this work points to.