Iranian Journal of Numerical Analysis and Optimization

Iranian Journal of Numerical Analysis and Optimization

A novel three-term conjugate gradient approach for deep neural network training

Document Type : Research Article

Authors
Department of Mathematics, Faculty of Mathematics, Statistics and Computer Science, Semnan University, P.O. Box 35195-363, Semnan, Iran
Abstract
This paper presents a new three-term conjugate gradient (CG) method for large-scale unconstrained optimization, with application to training deep neural networks. The approach employs enhanced CG parameter strategies together with the strong Wolfe line search to guarantee sufficient descent and global convergence under standard smoothness assumptions. Extensive numerical experiments are conducted on both the standard CUTEr benchmark test problems and the MNIST digit classi cation task. The proposed MPRP{IFR hybrid algorithm demonstrates superior efficiency compared to traditional CG variants, achieving faster convergence and reduced computational cost on CUTEr problems as well as high classi cation accuracy and stable training dynamics for deep learning models. These findings highlight the robustness and e Effectiveness of the proposed CG scheme for large-scale optimization.
Keywords
Subjects

[1] Bottou, L. Large-scale machine learning with stochastic gradient descent. In Proc. COMP-
STAT 2010, 177–186. Springer, 2010.
[2] Dada, I.D., Akinwale, A.T., Osinuga, I.A., and Olubiyi, M.O. Optimization of deep learning
model using an improved three-term conjugate gradient algorithm. Preprint, Research Square,
2022.
[3] Dai, Y. and Yuan, Y. A nonlinear conjugate gradient method with a strong global convergence
property. SIAM J. Optim., 10(1) (1999), 177–182.
[4] Dolan, E.D. and Moré, J.J. Benchmarking optimization software with performance profiles.
Math. Program., 91(2, Ser. A) (2002), 201–213.
[5] Fletcher, R. and Reeves, C.M. Function minimization by conjugate gradients. Comput. J.,
7(2) (1964), 149–154.
[6] Goodfellow, I., Bengio, Y., and Courville, A. Deep Learning. MIT Press, 2016.
[7] Hestenes, M.R. and Stiefel, E. Methods of conjugate gradients for solving linear systems. J.
Res. Natl. Bur. Stand., 49(6) (1952), 409–436.
[8] Jiang, X. and Jian, J. Improved Fletcher–Reeves and Dai–Yuan conjugate gradient methods
with the strong Wolfe line search. J. Comput. Appl. Math., 348 (2019), 525–534.
[9] Jin, W., Ma, G., Lin, H., and Han, D. A modified conjugate gradient method for image
restoration problems. J. Math. Imaging Vision, 45(3) (2012), 327–339.
[10] LeCun, Y., Bengio, Y., and Hinton, G. Deep learning. Nature, 521(7553) (2015), 436–444.
[11] LeCun, Y. and Cortes, C. The MNIST database of handwritten digits. 1998.
[12] Lin, X. and Jiang, D. A conjugate gradient approach for signal restoration problems. Signal
Process., 104 (2014), 366–373.
[13] Ma, G., Lin, H., Jin, W., and Han, D. Two modified conjugate gradient methods for un-
constrained optimization with applications in image restoration problems. J. Appl. Math.
Comput., 68 (2022), 4733–4758.
[14] Martens, J. Deep learning via Hessian-free optimization. In Proc. 27th Int. Conf. Mach.
Learn. (ICML), 2010.
[15] Mehamdia, A. and Chaib, Y. Improved conjugate gradient methods and application to non-
parametric estimation. Appl. Math., 51(2) (2024) 147–161.
[16] Mehamdia, A. and Chaib, Y. Two modified conjugate gradient methods for unconstrained
optimization. Optim. Methods Softw., 40(2) (2025), 308–321.
[17] Mehamdia, A. and Chaib, Y. A modified conjugate gradient method for solving unconstrained
optimization with application in conditional mode regression. Int. J. Comput. Math., 102(9)
(2025), 1–13.
[18] Mehamdia, A. and Khudhur, H. A globally convergent of two conjugate gradient methods
with application to image restoration problems. Numer. Algebra Control Optim., 15(3) (2025)
645–660.
[19] Nocedal, J. and Wright, S.J. Numerical Optimization. Springer, New York, 2nd ed., 2006.
[20] Polyak, B.T. The conjugate gradient method in extremal problems. USSR Comput. Math.
Math. Phys., 9 (1969), 94–112.
[21] Yang, J., Chen, J., and Liu, Z. Adaptive conjugate gradient methods for large-scale machine
learning. J. Mach. Learn. Res., 23(1) (2022), 1–36.
[22] Zhang, J., Xie, L., Dai, B., and Yao, X. Gradient descent algorithms for deep learning: a
survey. IEEE Trans. Syst. Man Cybern. Syst., 46(3) (2016), 271–285.
Send comment about this article
Enter Name.
Enter a valid email address.
Enter a vaid affiliation.
Enter comments (At leaset 10 words)
CAPTCHA Image
Enter Security Code Correctly.