September 2026

Journal

Intercomparison of flood inundation models across land use types and hydrological flood stages

By:
Nikrou, Parvaneh; Seyvani, Sadra; Cohen, Sagy; Gangrade, Sudershan ; Gutenson, Joseph; Tian, Dan; Baruah, Anupal; Follum, Michael; Kao, Shih-Chieh
Journal Name:
Journal of Hydrology
Page Number:
135410
Volume:
673
Publication Date:
September 8, 2026
View DOI Listing:
https://doi.org/10.1016/j.jhydrol.2026.135410

Abstract

Flood Inundation Mapping (FIM) model selection is a key operational decision because accurate, rapid mapping underpins early warning and resource allocation. FIM performance is context-dependent and can vary with hydrograph phase, land-use/land-cover (LULC), and the evaluation benchmark. Intercomparison studies typically assess a single near-peak snapshot against one reference dataset. Here, we provide a context-stratified intercomparison across (i) multiple hydrograph phases, (ii) LULC classes, and (iii) benchmark types, for five FIM approaches spanning a wide range of physical complexity and operational cost (TRITON, LISFLOOD-FP, HEC-RAS 2D, ARC-Curve2Flood, and OWP HAND-FIM). We use the Hurricane Matthew flood (2016) in the Neuse River Basin, North Carolina, USA, as a case study. Using high-resolution remote sensing-derived flood inundation maps, hand-labeled points, and building footprints, we assess model skill across two rising and two falling hydrograph limbs and across major LULC types. Results show that model rankings shift systematically across contexts: LISFLOOD-FP ranks highest in three of four flood phases, while TRITON leads during one rising limb phase; LISFLOOD-FP performs best in vegetated areas, whereas HEC-RAS improves relative performance in agricultural and urban areas; and benchmark choice influences conclusions, with LISFLOOD-FP performing best for flooded-building detection in the late falling limb, while TRITON ranks highest against hand-labeled points. We also report representative wall-clock runtimes for each workflow to provide use-case context for operational feasibility. Together, these results offer transferable guidance for model selection and for designing large-scale, benchmark-aware FIM intercomparison studies.