001.891.573 Improving the training efficiency of neural networks for image segmentation

Valishin A. A. (Bauman Moscow State Technical University), Zaprivoda A. V. (Bauman Moscow State Technical University), Tsukhlo S. S. (Bauman Moscow State Technical University)

CONVOLUTIONAL NEURAL NETWORKS, COMPUTER VISION, SEMANTIC SEGMENTATION, MACHINE LEARNING, ACTIVE LEARNING


doi: 10.18698/2309-3684-2025-3-103116


This paper critically analyzes modern approaches to improving the efficiency of training neural networks for image segmentation. Training efficiency refers to two interrelated aspects: computational efficiency and segmentation accuracy of the trained model. Particular attention is paid to three ways to improve training efficiency: 1. using augmentation methods. Augmentation is the process of artificially generating new data based on existing models for training. This method allows increasing the size and diversity of the dataset, which is important for improving the generalization ability of the model. In this case, augmentation included image rotations, Gaussian noise imposition, color correction, 2. optimization of neural network architectures by integrating efficient encoders based on EfficientNet, 3. using active learning methods to select the most informative training examples based on calculating the entropy of the output data. The models were trained using the Adam optimizer on the OpenEarthMap task, where the sample was 20% of the original volume, the images were downsampled to 512x512 pixels and further split into four parts of 256x256 pixels. Training was performed on 9212 images of the training set and 1536 images of the validation set for 100 training cycles. The experimental results show that augmentation increases the segmentation accuracy of the UNet model (IoU) from 36 % to 38.7 %, architectural optimization using EfficientNet-b0 and b4 increases IoU to 44.6 % and 45.3 %, respectively, and active learning based on entropy calculation shows the potential to equalize IoU across classes, although the stability of the metrics remains problematic. This work highlights the need and potential of an integrated approach to optimizing neural network models for image segmentation and points to directions for further research in the field of machine learning and improving computational efficiency.


[1] Ulku I., Akagunduz E. A survey on deep learning-based architectures for semantic segmentation on 2D images. Applied Artificial Intelligence, 2022, vol. 36, no. 1. DOI:10.1080/08839514.2022.203292
[2] Valishin A.A., Zaprivoda A.V., Klonov A.S. Mathematical modeling and comparative analysis of numerical methods for solving the problem of continuous-discrete filtering of random processes in real time. Mathematical Modeling and Computational Methods, 2024, no. 1, pp. 93-109.
[3] Dostovalova A.M. Simulation of locally homogeneous radar images using different statistical criteria. Mathematical Modeling and Computational Methods, 2021, no. 4, pp. 103-120.
[4] Kutyrkin V.A., Chaley M.B. Stochastic coding models and the distribution of structural and statistical characteristics of coding sequences. Mathematical Modeling and Computational Methods, 2017, no. 3, pp. 119-138
[5] Minaee S., Boykov Y., Porikli F., Plaza A., Kehtarnavaz N., Terzopoulos D. Image Segmentation Using Deep Learning: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021, vol. 44, no. 7, pp. 3523-3542.
[6] Valishin A.A., Zaprivoda A.V., Tsukhlo S.S. Modeling and efficiency analysis of perceptual hash functions for segmented image search. Mathematical Modeling and Computational Methods, 2024, no. 2, pp. 46-67.
[7] Ronneberger O., Fischer P., Brox T. U-Net: convolutional networks for bio-medical image segmentation. arXiv.org. [Электронный ресурс]. URL: https://arxiv.org/abs/1505.04597 (дата обращения: 05.06.2024).
[8] Tan M., Le Q. V. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. arXiv.org. [Электронный ресурс]. URL: https://arxiv.org/abs/1905.11946 (дата обращения: 23.05.2024).
[9] Yan Q., Feng Y., Zhang C., Wang P., Wu P., Dong W., Sun J., Zhang Y. You only need one color space: an efficient network for low-light image enhancement. arXiv.org. [Электронный ресурс]. URL: https://doi.org/10.48550/arXiv.2402.05809 (дата обращения: 12.05.2024).
[10] Lin T.-Yi, Goyal P., Girshick R., He K., Dollár P. Focal Loss for Dense Object Detection. Conference: 2017 IEEE International Conference on Computer Vision (ICCV), 2017. DOI:10.1109/ICCV.2017.324
[11] Kingma D.P., Ba J. Adam: A Method for Stochastic Optimization. arXiv.org. [Электронный ресурс]. URL: https://doi.org/10.48550/arXiv.1412.6980 (дата обращения: 11.06.2024).
[12] Xia J., Yokoya N., Adriano B., Broni-Bediako C. OpenEarthMap: A Benchmark Dataset for Global High-Resolution Land Cover Mapping. arXiv.org. [Электронный ресурс]. URL: https://doi.org/10.48550/arXiv.2210.10732 (дата обращения: 24.05.2024).


Валишин А.А., Запривода А.В., Цухло С.С. Повышение эффективности обучения нейронных сетей для сегментации изображений. Математическое моделирование и численные методы, 2025, № 3, с. 103–116.



Download article

Количество скачиваний: 108