Abstract
Ensemble deep learning (EDL) methods offer an effective way to reduce model structural uncertainty and improve the reliability of river flow predictions, especially in catchments with limited data. This study evaluated the performance of a Time-Varying Dynamic Model Averaging (TV-DMA) ensemble for integrating heterogeneous machine learning and deep learning architectures, including FFNN, 1D-CNN, LSTM, CNN-LSTM and Diff-FFNN-LSTM models, for river flow prediction in two hydrologically distinct, data-scarce tropical catchments in the Upper White Nile Basin. Results show that the heterogeneous, first-order differential processing Diff-FFNN-LSTM model consistently achieved the highest individual-model river flow predictive performance (NSE ≥ 0.980), due to its ability to stabilise non-stationary river flow time series by mitigating boundary effects associated with decomposition methods, thereby improving representation of rapid flow transitions. Additionally, the TV-DMA resulted in more superior performance in both the Semliki (NSE = 0.910; PBIAS = −0.42%) and Tochi (NSE = 0.831; PBIAS = −0.34%) catchments when compared to the Bayesian Model Averaging (BMA) ensemble (Semliki: NSE = 0.848; PBIAS = −0.20% and Tochi: NSE = 0.710, PBIAS = −18.86%). BMA and TV-DMA also improved computational efficiency by 98–98.4% and 77–81%, respectively, relative to the process-based HEC-HMS model. The findings highlight the principal value of TV-DMA’s dynamic update of individual model weights based on predictive error distributions, which enhances river flow predictive performance, particularly in the rainfall-runoff-dominated Tochi catchment. This EDL approach can improve the accuracy and efficiency of river flow modelling and support the development of flood forecasting and early warning systems in data scarce regions.