Physics-Informed CNN-LSTM for Street-Scale Urban Flood Prediction: Reconciling Aggregate Accuracy and Street-Level Plausibility
Luc D’Costa, Yidi Wang, Jonathan L. Goodall, Rohan Chandra
Deep learning surrogate models trained with mean-squared-error loss produce statistically accurate but physically unconstrained flood predictions: water may flow uphill, appear spontaneously, or smooth over street-level corridors. In this work, a physics-informed training framework is developed for CNN-LSTM models that predict urban flood depths at 15 min intervals over a 128×128 spatial grid. Three differentiable penalty terms are embedded directly into the loss function: (i) a gravity loss that penalizes depth increases against the water-surface-elevation gradient, (ii) a continuity loss enforcing local mass conservation with rainfall-adaptive thresholds, and (iii) a topography-aware false-alarm penalty modulated by the topographic wetness index (TWI). The framework is evaluated on the Norfolk, Virginia, flood dataset spanning two major storm events (August 2017 and September 2022) comprising 300 samples, with all variants trained on identical splits and robustness assessed over repeated random splits and leave-one-storm-out tests. A road-proximal evaluation restricted to a TWI-derived street mask quantifies street-level skill. The physics-constrained model achieves near-zero gravity violations (∼10−6) and the highest street-channel recall (0.77 ± 0.09 versus 0.44 ± 0.10 for the unconstrained baseline), the capability most relevant to downstream traffic routing, and its recall advantage more than doubles on a held-out storm, while a uniform false-alarm variant attains 16% lower mean absolute error but suppresses street recall to 0.25. The proposed TWI-modulated penalty reconciles this trade-off: it improves upon the uniform variant on every metric measured, recovering 60% higher street recall at the lowest MAE among all constrained variants and the best street-level F1 score. These results expose a fundamental tension between aggregate pixel-level error metrics and application-specific physical plausibility, and demonstrate that terrain-aware loss modulation offers a principled resolution.