Stephan Johann Lehmler; Muhammad Saif-ur-Rehman; Tobias Glasmachers; Ioannis Iossifidis
How can we tell whether a neural network is learning generalizable structure versus memorizing rare patterns? In this paper, we study distributional signatures inside networks trained under memorization-heavy regimes. Our starting point is the idea that memorization corresponds to learning “rare” input features—an effect that can be reflected in the activation probability of individual ReLU neurons.
Building on this, we derive hypotheses about distributional properties across larger network components, extending earlier Poisson-process-inspired views of activations by explicitly considering correlations between neurons. We then simulate how memorizing neurons influence activation magnitudes and weight-matrix statistics, and connect these observations to the behavior of L1/L2 regularization.
We validate the proposed indicators empirically on an MNIST classification setup and observe that activation frequency and intra-layer correlation structure can help distinguish memorizing from generalizing networks—an important step toward online memorization metrics.
Reference
In: Giuseppe Nicosia; Varun Ojha; Sven Giesselbach; M. Panos Pardalos; Renato Umeton; Emanuele La Malfa; Gabriele La Malfa (Eds.)
Machine Learning, Optimization, and Data Science, pp. 410–423, Springer Nature Switzerland, Cham, 2026.
ISBN: 978-3-032-21477-5.
Links
BibTeX
@inproceedings{lehmler2026distributional,
title={Distributional Properties of ReLU-Activations in Artificial Neural Networks That Learn by Memorization},
author={Lehmler, Stephan Johann and Saif-ur-Rehman, Muhammad and Glasmachers, Tobias and Iossifidis, Ioannis},
booktitle={Machine Learning, Optimization, and Data Science},
pages={410--423},
year={2026},
publisher={Springer Nature Switzerland}
}
