DK

D. Krylov

info

Please Note

1 records found

Master thesis (2026) - D. Krylov, S. Tan, D.M.J. Tax, R.A. Norte
Training state-of-the-art AI models is growing far more costly than digital hardware can sustain, motivating analog photonic substrates that perform matrix multiplication directly in the propagation of light. Inference on such hardware is established, but training is not: backpropagation requires a backward pass and exact gradients that an analog chip cannot natively provide. Forward-only, backpropagation-free learning offers a way around this, yet it has been demonstrated only for multilayer perceptrons and convolutional networks, never for the Transformer, the architecture that now dominates AI and whose self-attention couples every position across the sequence, resisting purely local objectives.

This thesis asks whether a Transformer can be trained using only the forward-pass operations a photonic substrate provides, and characterizes what such training costs. The proposed method combines a layer-wise Forward-Forward prototype-based objective, directional-derivative gradient estimation (with no backward pass and no automatic differentiation), and a softmax-free Spherical attention adapted from the Kramers-Kronig kernel, with the four attention projections trained one at a time in a round-robin schedule; six training variants are compared across seven vision and sequence tasks to isolate the gradient estimator and the update schedule. The answer is affirmative: the fully forward-only variant trains a Transformer to a useful operating point on six of the seven tasks. Locality, not the forward-only gradient alone, is what makes this possible, and the depth-resilience of local learning is shown to extend, conditionally, to self-attention. The remaining gap to backpropagation has three sources, two inherent to forward-only learning (local credit assignment and gradient-estimation variance) and one architectural (the softmax-free attention cannot form the sharp selection that content-addressed retrieval requires). Because every operation reduces to a forward pass and a measured scalar loss, the method is a candidate for in-situ training on a photonic chip, the validation step this work points to. ...