[1] Bertsekas, D.P. Nonlinear programming, 2nd ed., Athena Scientific, 1999.
[2] Billingsley, P. Convergence of probability measures, John Wiley & Sons, 1968.
[3] Bottou, L., Curtis, F.E. and Nocedal, J. Optimization methods for large-scale machine
learning, SIAM Rev. 60(2) (2018), 223–311.
[4] Boyd, S. and Vandenberghe, L. Convex optimization, Cambridge University Press, 2004.
[5] Bubeck, S. Convex optimization: Algorithms and complexity, Found. Trends Mach. Learn.
8(3–4) (2015), 231–358.
[6] Dieuleveut, A., Durmus, A. and Bach, F. Bridging the gap between constant step size
stochastic gradient descent and Markov chains, arXiv preprint arXiv:1707.06386 (2018).
[7] Dimitrieski, N., Cao, J. and Ebenbauer, C. Stochastic gradient descent for constrained op-
timization based on adaptive relaxed barrier functions, Int. J. Robust. Nonlinear Control.
(2025).
[8] Feller, C. and Ebenbauer, C. A stabilizing iteration scheme for model predictive control
based on relaxed barrier functions, Automatica 80 (2017), 328–339.
[9] Fiacco, A.V. and McCormick, G.P. Nonlinear programming: Sequential unconstrained
minimization techniques, John Wiley & Sons, 1968.
[10] Garrigos, G. and Gower, R.M. Handbook of convergence theorems for (stochastic) gradient
methods, arXiv preprint arXiv:2301.11235 (2023).
[11] Gower, R.M., Loizou, N., Qian, X., Sailanbayev, A., Shulgin, E. and Richtárik, P. SGD:
General analysis and improved rates, arXiv preprint arXiv:1901.09401 (2019).
[12] Li, M., Grigas, P. and Atamtürk, A. New penalized stochastic gradient methods for linearly
constrained strongly convex optimization, SIAM J. Optim. 32(2) (2022), 859–887.
[13] Liu, J. and Yuan, Y. On almost sure convergence rates of stochastic gradient methods,
In Proceedings of the 35th Annual Conference on Learning Theory, PMLR, 178 (2022),
1–21.
[14] Nocedal, J. and Wright, S.J. Numerical optimization, 2nd ed., Springer, 2006.
[15] Polyak, B.T. and Juditsky, A.B. Acceleration of stochastic approximation by averaging,
SIAM J. Control Optim. 30(4) (1992), 838–855.
[16] Prechelt, L. Early stopping but when?, In: Orr, G.B., Müller, K.R. (eds.) Neural net-
works: Tricks of the trade, Lecture Notes in Computer Science, vol 1524, Springer, Berlin,
Heidelberg, 1998, 55–69.
[17] Robbins, H. and Monro, S. A stochastic approximation method, Ann. Math. Stat. 22(3)
(1951), 400–407.
[18] Rockafellar, R.T. Convex analysis, Princeton University Press, 1970.
[19] Roy, S.K. and Harandi, M. Constrained stochastic gradient descent: The good practice,
In Int. Conf. Digit. Image Comput. Tech. Appl. (DICTA), IEEE, 2017, 1–8.
[20] Rudin, W. Principles of mathematical analysis, 3rd ed., McGraw-Hill, 1976.
[21] Wang, M. and Bertsekas, D.P. Incremental constraint projection methods for variational
inequalities, Math. Program. 150 (2015), 321–363.
[22] Wang, X., Ma, S. and Yuan, Y. Penalty methods with stochastic approximation for
stochastic nonlinear programming, arXiv preprint arXiv:1605.05609 (2016).
[23] Yan, Y. and Xu, Y. Adaptive primal-dual stochastic gradient method for expectation-
constrained convex stochastic programs, Math. Program. Comput. 14(2) (2022), 319–363.
[24] Zhang, T. Solving large scale linear prediction problems using stochastic gradient descent
algorithms, In Proc. 21st Int. Conf. Mach. Learn. (ICML), 2004, 116