[1] Bottou, L., Curtis, F.E. and Nocedal, J. Optimization methods for large-scale machine
learning, SIAM Rev. 60(2) (2018), 223–311.
[2] Boyd, S. and Vandenberghe, L. Convex optimization, Cambridge University Press, 2004.
[3] Bubeck, S. Convex optimization: Algorithms and complexity, Found. Trends Mach. Learn.
8(3-4) (2015), 231–358.
[4] Dimitrieski, N., Cao, J. and Ebenbauer, C. Stochastic gradient descent for constrained
optimization based on adaptive relaxed barrier functions, Int. J. Robust. Nonlinear Control.
(2025).
[5] Feller, C. and Ebenbauer, C. A stabilizing iteration scheme for model predictive control
based on relaxed barrier functions, Automatica 80 (2017), 328–339.
[6] Fiacco, A.V. and McCormick, G.P. Nonlinear programming: Sequential unconstrained min-
imization techniques, John Wiley & Sons, 1968.
[7] Garrigos, G. and Gower, R.M. Handbook of convergence theorems for (stochastic) gradient
methods, arXiv preprint arXiv:2301.11235 (2023).
[8] Gower, R.M., Loizou, N., Qian, X., Sailanbayev, A., Shulgin, E. and Richtárik, P. SGD:
General analysis and improved rates, arXiv preprint arXiv:1901.09401 (2019).
[9] Hardy, G.H. Divergent series, Oxford University Press, 1949.
[10] Karandikar, R.L. and Vidyasagar, M. Convergence rates for stochastic approximation: bi-
ased noise with unbounded variance, and applications, arXiv preprint arXiv:2312.02828
(2024).
[11] Li, M., Grigas, P. and Atamtürk, A. New penalized stochastic gradient methods for linearly
constrained strongly convex optimization, SIAM J. Optim. 32(2) (2022), 859–887.
[12] Moulines, E. and Bach, F. Non-asymptotic analysis of stochastic approximation algorithms,
In Advances in Neural Information Processing Systems (NeurIPS), 451–459, 2011.
[13] Nocedal, J. and Wright, S.J. Numerical optimization, 2nd ed., Springer, 2006.
[14] Pflug, G.C. Non-asymptotic confidence bounds for stochastic approximation algorithms with
constant step size, Monatsh. Math. 110(3-4) (1990), 297–314.
[15] Polyak, B.T. and Juditsky, A.B. Acceleration of stochastic approximation by averaging,
SIAM J. Control Optim. 30(4) (1992), 838–855.
[16] Robbins, H. and Monro, S. A stochastic approximation method, Ann. Math. Stat. 22(3)
(1951), 400–407.
[17] Robbins, H. and Siegmund, D. A convergence theorem for nonnegative almost supermartin-
gales and some applications, In Optimizing Methods in Statistics, Academic Press, 1971,
233–257.
[18] Rockafellar, R.T. Convex analysis, Princeton University Press, 1970.
[19] Roy, S.K. and Harandi, M. Constrained stochastic gradient descent: The good practice, In
Int. Conf. Digit. Image Comput. Tech. Appl. (DICTA), IEEE, 1–8, 2017.
[20] Rudin, W. Principles of mathematical analysis, 3rd ed., McGraw-Hill, 1976.
[21] Wang, M. and Bertsekas, D.P. Incremental constraint projection methods for variational
inequalities, Math. Program. 150 (2015), 321–363.
[22] Wang, X., Ma, S. and Yuan, Y. Penalty methods with stochastic approximation for stochastic
nonlinear programming, arXiv preprint arXiv:1605.05609 (2016).
[23] Yan, Y. and Xu, Y. Adaptive primal-dual stochastic gradient method for expectation-
constrained convex stochastic programs, Math. Program. Comput. 14(2) (2022), 319–363.
[24] Zhang, T. Solving large scale linear prediction problems using stochastic gradient descent
algorithms, In Proc. 21st Int. Conf. Mach. Learn. (ICML), 116, 2004