Data Analysis of Two-Vehicle Accidents Based on Machine Learning
Dongguang Gao, Jiawei Chen, Tianyu Luo, Zijun Liu, Libo Cao, Zhongxiang Chen, Jun Wu
Road traffic accidents are the eighth leading cause of human deaths. In order to study two-vehicle accidents, this paper extracted data from 493 two-vehicle accidents from the CIDAS database from 2011 to 2022, used machine learning methods to analyze the accident data, and obtained the significance of two-vehicle accident parameters. Finally, five typical scenarios of two-vehicle accidents were obtained based on this. The results of the significance analysis show that vehicle parameters have a greater impact on occupant injury in the host vehicle; clustering results show that lighting, the number of lanes, the other vehicle’s type, and the speed of the host vehicle have a large impact on occupant injury (for example, the injury rate for the high-speed, nighttime Scenario II was 52.9%, compared to just 10.9% for the lower-speed Scenario IV). Factor analysis results show that precipitation has a large impact on occupant injury, as the frequency of injuries in rainy conditions was 13.4% higher, and the frequency of serious injuries was 7.9% higher, than in accidents without rain. This paper innovatively uses factor analysis to reduce the dimensionality of categorical variables, which provides research ideas for related research. At the same time, the clustering results obtained in this paper also provide references for the establishment of corresponding test scenarios for autonomous driving and the establishment of standards.