A Machine Learning Approach for Water Quality Assessment in the Lower Rio Grande Valley Watershed
Saika Nowshin Nowrin, Chu‐Lin Cheng, Jungseok Ho, Jinwoo An, Fatemeh Nazari
Water quality analysis plays an essential role in maintaining the health and sustainability of river ecosystems, especially in semi-arid regions like the Arroyo Colorado Watershed in South Texas. Since the river is a vital source of water supply for local communities, agriculture, and wildlife, it faces significant challenges and pollution from land use changes, climate variation, and agricultural runoff. Continuous monitoring and assessment of water quality parameters and their temporal variability are essential to ensure the drinking water supply and aquatic ecosystem health. However, comprehensive laboratory-based water quality investigations are often constrained by higher costs, logistical complexity, and limited manpower. As a result, monitoring datasets are often not available for all water quality parameters, or the datasets may be incomplete. To address such challenges, the objective of this study was to evaluate the potential of water quality index (WQI)-based assessment supported by machine learning algorithms as an alternative decision-support tool for water quality evaluation. The analysis compared four monitoring stations in the Austin and Arroyo Colorado Watersheds, with particular emphasis on one gauging station at Port Harlingen. Datasets were collected from the Texas Commission of Environmental Quality (TCEQ). A complete exploratory data analysis (EDA) was performed to understand the TCEQ water quality datasets containing sixteen parameters, and seven water quality parameters were selected based on multicollinearity checks. It was observed that seven independent water quality parameters (dissolved oxygen, ammonia, nitrate, phosphorus, temperature, fecal coliform, and residual non-filterable material concentrations) were identified as sufficient to define the WQI of the Austin monitoring stations. Moreover, U.S. Environmental Protection Agency (EPA)-based guidelines were utilized to scale individual parameters to a range of 0–100 to remove their magnitude and correlation-based bias. These parameters were further analyzed using machine learning techniques, i.e., principal component analysis, K-means, and one-class support vector machine, to compute the relative importance based on their fluctuation within the temporal dataset. Finally, the mean WQI model was developed for Port Harlingen and achieved a strong agreement with the National Sanitation Foundation (NSF) WQI (R2 = 0.91). These findings demonstrate the applicability of the proposed data-driven WQI framework for regional water quality assessment and comparative analysis across watersheds.