doi: 10.18698/2309-3684-2025-3-85102
The article is devoted to the development of a multi-agent evacuation model that takes into account the physical characteristics of agents (age categories, speed, maneuverability), the level of panic, social interactions in groups of the “leader-follower” type, and the presence of several evacuation exits opening at a given interval (an interval of 6 seconds was considered). The Multi-Agent Proximal Policy Optimization (MAPPO) algorithm is used to train the behavior of agents. A hybrid action space is used, combining discrete output selection and continuous motion control. Training is carried out according to the curriculum learning principle: with a gradual increase in the number of agents. This allows agents to adapt to complex scenarios with high crowding and improves the generalization ability of the model for experiments with different numbers of agents. The environment is a room of given dimensions (rooms of 15×20 m were considered) with a given number of exits of a certain width (3 exits of 1.5 m each were considered). The model includes the logic of disseminating information about exits. Individual agents learn about new open exits within a radius of 5 m and transmit the signal to their neighbors. Leaders initially know about all available exits regardless of the distance. A mechanism is provided for spreading panic depending on the crowding of agents, the distance to the exit, and the time elapsed since the start of the evacuation. Specific rules of behavior for social groups are introduced: leaders make strategic decisions, and elderly followers receive a speed bonus when following the leader. In the current implementation, the choice of exit for individual agents is based on the shortest distance from the agent to it. In social groups, the decision to choose an exit is made by the leader based on the average distance of all agents. Computational experiments were conducted for 40 agents in various scenarios: with a different number of leaders (2–16) and without groups (individual evacuation). The computational experiments showed that under the considered conditions, scenarios with social groups lead to faster evacuation (the total time was reduced by about 38%). Also, during group evacuation, vulnerable agents receive the greatest advantage, in this case, the elderly. The optimal number of leaders is 4–6: further increase in their number does not provide statistically significant improvements. According to the results of the experiments, a decrease in the number of collisions and a lower level of panic with this number of leaders was recorded. The obtained results demonstrate the practical applicability of the MAPPO approach to the problems of analyzing evacuation processes in realistic conditions.
[1] Kotkova E.A., Matveev A.V., Nefedev S.A., et al. Agentnoe modelirovanie protsessa evakuatsii lyudey pri pozharakh v zdaniyakh: obzor podkhodov i issledovaniy [Agent modeling of the process of people evacuation during fire in buildings: a review of approaches and research]. Sovremennye naukoemkie tekhnologii, 2023, no. 10, pp. 55–62.
[2] Zia K., Ferscha A. An agent-based model of crowd evacuation: combining individual, social and technological aspects. Proceedings of the 2020 ACM SIGSIM conference on principles of advanced discrete simulation. New York, Association for Computing Machinery, 2020, pp. 129–140.
[3] Kotkova E.A. Perspektivy primeneniya iskusstvennykh neironnykh setey pri modelirovanii protsessa evakuatsii [Perspectives in applying artificial neural networks in modeling the evacuation process]. Pozharnaya i tekhnosfernaya bezopasnost': problemy i puti sovershenstvovaniya, 2020, no. 1(5), pp. 359–361.
[4] Sukhanov V.O., Kuzmin A.I., Skorokhodov D.V. Geoinformatsionnaya sistema podderzhki prinyatiya resheniy na evakuatsiyu naseleniya [Geoinformation system support decision-making on evacuation of the population]. Pozharnaya bezopasnost': problemy i perspektivy, 2019, vol. 1, no. 10, pp. 411–413.
[5] Sazhin I.S., Golovenko E.L., Chaniev B.Yu., et al. Intellektual'naya sistema opoveshcheniya i upravleniya evakuatsiyey lyudey na osnove informatsionnogo modelirovaniya chrezvychaynykh situatsiy v zdanii [Intelligent alerting and evacuation management system based on information emergency simulation building]. Nauka, tekhnika i obrazovanie, 2021, no. 4(79), pp. 40–44.
[6] Kotkova E.A., Matveev A.V. Metodika intellektual'nogo prognozirovaniya effektivnosti upravleniya evakuatsiey lyudey iz obshchestvennykh zdaniy [Methodology for intellectual forecastingof the efficiency of managing people evacuation of from public buildings]. Vestnik Sankt-Peterburgskogo universiteta Gosudarstvennoy protivopozharnoy sluzhby MChS Rossii, 2021, no. 4, pp. 107–120.
[7] Tsvirkun A.D., Rezchikov A.F., Samartsev A.A., et al. Integrirovannaya model' dinamiki rasprostraneniya opasnykh faktorov pozhara v pomeshcheniyakh i evakuatsii iz nikh [Integrated model of the fire dangerous factors dynamics in premises and the evacuation]. Vestnik komp'yuternykh i informatsionnykh tekhnologiy, 2019, vol. 2, no. 176, pp. 47–54.
[8] Tsvirkun A.D., Rezchikov A.F., Samartsev A.A., et al. Sistema integrirovannogo modelirovaniya rasprostraneniya opasnykh faktorov pozhara i evakuatsii lyudey iz pomeshcheniy [System of integrated simulation of spread of hazardous factors of fire and evacuation of people from indoors]. Automation and Remote Control, 2022, no 5, pp. 26–42.
[9] Samartsev A., Ivaschenko V., Rezchikov A., et al. Multiagent model of people evacuation from premises while emergency. Advances in Systems Science and Applications, 2019, vol. 19, no. 1, pp. 98–115.
[10] Gamayunova V.O., Bogomolov A.S., Kushnikov V.A., et al. Multiagentnoe modelirovanie evakuatsii iz pomeshcheniy s uchetom stolknoveniy agentov [Multi-agent modeling of evacuation from premises with consideration of agent collisions]. Izvestiya of Saratov University. Mathematics. Mechanics. Informatics, 2025, vol. 25, no 1, pp. 106–115.
[11] Rosa A.C., Falqueiro M.C., Bonacin R., et al. EvacuAI: An Analysis of Escape Routes in Indoor Environments with the Aid of Reinforcement Learning. Sensors, 2023, vol. 23, no. 21, art. no. 8892.
[12] Ünal A.E., Gezer C., Pak B.K., et al. Generating emergency evacuation route directions based on crowd simulations with reinforcement learning. 2022 Innovations in Intelligent Systems and Applications Conference ASYU. Antalya: IEEE, 2022, pp. 1–6.
[13] Xu D., Huang X., Mango J., et al. Simulating multi-exit evacuation using deep reinforcement learning. Transactions in GIS, 2021, vol. 25, no. 3, pp. 1542–1564.
[14] Malebary S.J., Basori A.H., Soliman alkayal E. Reinforcement learning for Pedestrian evacuation Simulation and Optimization during Pandemic and Panic situation. Journal of Physics: Conference Series, 2021, vol. 1817, no. 1, art. no. 012008.
[15] Komatsu H. Multi-agent reinforcement learning using echo-state network and its application to pedestrian dynamics. arXiv preprint arXiv:2312.11834, 2023.
[16] Sinpan N., Sasithong P., Chaudhary S., et al. Simulative Investigations of Crowd Evacuation by Incorporating Reinforcement Learning Scheme. ICACS '22: Proceedings of the 6th International Conference on Algorithms, Computing and Systems. New York: Association for Computing Machinery, 2022, pp. 1–5.
[17] Hassanpour S., Rassafi A.A., González V.A., et al. A hierarchical agent-based approach to simulate a dynamic decision-making process of evacuees using reinforcement learning. Journal of choice modelling, 2021, vol. 39, art. no. 100288.
[18] Schulman J., Wolski F., Dhariwal P., et al. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
[19] Yu C., Velu A., Vinitsky E., et al. The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games. arXiv preprint arXiv:2103.01955, 2022.
[20] Liu Z., Yao C., Na W., et al. MAPPO-Based Optimal Reciprocal Collision Avoidance for Autonomous Mobile Robots in Crowds. 2023 IEEE International Conference on Systems, Man, and Cybernetics SMC. Honolulu: IEEE, 2023, pp. 3907–3912.
[21] Guo Y., Liu J., Yu R., et al. MAPPO-PIS: A Multi-agent Proximal Policy Optimization Method with Prior Intent Sharing for CAVs Cooperative Decision-Making. Computer Vision – ECCV 2024 Workshops. Cham: Springer Nature Switzerland, 2025, pp. 244–263.
[22] Shixin Z., Feng P., Anni J., et al. The unmanned vehicle on-ramp merging model based on AM-MAPPO algorithm. Scientific Reports, 2024, vol. 14, no. 1, art. no. 19416.
[23] Lowe R., Wu Y., Tamar A., et al. Multi-agent actor-critic for mixed cooperative-competitive environments. Proceedings of the 31st International Conference on Neural Information Processing Systems. Red Hook: Curran Associates Inc, 2017, pp. 6382–6393.
[24] Xiong J., Wang Q., Yang Z., et al. Parametrized Deep Q-Networks Learning: Reinforcement Learning with Discrete-Continuous Hybrid Action Space. arXiv preprint arXiv:1810.06394, 2018.
[25] Srivastava N., Hinton G., Krizhevsky A., et al. Dropout: a simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 2014, vol. 15, no. 56, pp. 1929–1958.
[26] Gal Y., Ghahramani Z. Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning. Proceedings of The 33rd International Conference on Machine Learning. New York: PMLR, 2016, pp. 1050–1059.
[27] Narvekar S., Peng B., Leonetti M., et al. Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey. Journal of Machine Learning Research, 2020, vol. 21, no. 181, pp. 1–50.
[28] Sutton R.S., Barto A. Reinforcement learning: an introduction. Second edition. Cambridge, The MIT Press, 2018, 552 p.
[29] Lillicrap T.P., Hunt J.J., Pritzel A., et al. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971, 2015.
[30] Ob utverzhdenii svoda pravil SP 1.13130 «Sistemy protivopozharnoy zashchity. Evakuatsionnye puti i vykhody»: Prikaz MChS Rossii ot 19 marta 2020 g. N 194. Moskva. 2020.
[31] Trivedi A., Rao S. Agent-Based Modeling of Emergency Evacuations Considering Human Panic Behavior. IEEE Transactions on Computational Social Systems, 2018, vol. 5, no. 1, pp. 277–288.
[32] Ding N., Sun C. Experimental study of leader-and-follower behaviours dur-ing emergency evacuation. Fire Safety Journal, 2020, vol. 117, art. no. 103189.
[33] Wang L., Zheng J., Zhang X., et al. Pedestrians behavior in emergency evacuation: Modeling and simulation. Chinese Physics B, 2016, vol. 25, no. 11, art. no. 118901
Силинская А.А., Богомолов А.С., Кушников В.А. Моделирование эвакуации из помещений с учетом социальных групп и множественных выходов. Математическое моделирование и численные методы, 2025, № 3, с. 85–102.
Количество скачиваний: 125