AN AI-DRIVEN SMART MONITORING SYSTEM FOR ECOLOGICAL FLOW IN SMALL HYDROPOWER STATIONS
Keywords:
Ecological flow, Hydropower monitoring, Computer vision, Multimodal large modelsAbstract
Ecological flow refers to the minimum water release required to sustain downstream ecosystems. Monitoring compliance at small hydropower stations is usually challenging because outlets may vary in design, live cameras are often misaligned, or monitoring reports can be incomplete or falsified. This paper introduces a novel AI-driven monitoring system that overcomes these challenges by leveraging a multimodal large model, carefully adapted through parameter-efficient fine-tuning. Our approach not only ensures more reliable ecological flow compliance monitoring but also demonstrates how multimodal models can be applied to improve real-world environmental management. The proposed system addresses three critical questions: 1. Are cameras correctly installed? 2. Is water flowing? and 3. Does the release comply with regulatory standards? Despite being trained on a limited labelled dataset, the model achieves over 90% precision across diverse outlet types with narrow confidence intervals. The system is now deployed at nearly 5,000 stations, minimising dependence on manual inspections while enabling continuous, transparent, and scalable monitoring of ecological flow.References
[1] Yu Z, Zhang J, Zhao J, et al. A new method for calculating the downstream ecological flow of diversion-type small hydropower stations. Ecological Indicators, 2021, 125: 107530.
[2] Liu W, Chen J, Wang H, et al. Perspectives on Advancing Multimodal Learning in Environmental Science and Engineering Studies. Environmental Science & Technology, 2024. DOI: 10.1021/acs.est.4c03088.
[3] Yin S, Fu C, Zhao S, et al. A Survey on Multimodal Large Language Models. National Science Review, 2024, 11(12): nwae403. DOI: 10.1093/nsr/nwae403.
[4] Feichtenhofer C, Fan H, Malik J, et al. SlowFast Networks for Video Recognition. arXiv preprint, 2019. DOI: 10.48550/arXiv.1812.03982.
[5] Mittal A, Soundararajan R, Bovik AC. Making a "completely blind" image quality analyzer. IEEE Signal Processing Letters, 2012, 20(3): 209-212.
[6] Dosovitskiy A, Fischer P, Ilg E, et al. FlowNet: learning optical flow with convolutional networks. 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 2015: 2758-2766.
[7] Teed Z, Deng J. RAFT: recurrent all-pairs field transforms for optical flow. Computer Vision – ECCV 2020. ECCV 2020. Lecture Notes in Computer Science, 2020, 12347. DOI: 10.1007/978-3-030-58536-5_24.
[8] Elgendy H, Sharshar A, Aboeitta A, et al. GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing. arXiv preprint, 2024. DOI: 10.48550/arXiv.2410.19552.
[9] Zhang R, Isola P, Efros AA, et al. The unreasonable effectiveness of deep features as a perceptual metric. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 2018: 586-595. DOI: 10.1109/CVPR.2018.00068.
[10] GPGGWR Commission. Small hydropower stations in Guangdong Province. GPGGWR Commission, 2018.
[11] Radford A, Kim JW, Hallacy C, et al. Learning transferable visual models from natural language supervision. arXiv preprint, 2021. DOI: 10.48550/arXiv.2103.00020.
[12] Li J, Li D, Savarese S, et al. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. ICML'23: Proceedings of the 40th International Conference on Machine Learning, 2023, 814: 19730-19742.
[13] Liu H, Li C, Wu Q, et al. Visual instruction tuning. NIPS '23: Proceedings of the 37th International Conference on Neural Information Processing Systems, 2023, 1516: 34892-34916.
[14] Hu EJ, Shen Y, Wallis P, et al. LoRA: low-rank adaptation of large language models. arXiv preprint, 2021. DOI: 10.48550/arXiv.2106.09685.
[15] Han Z, Gao C, Liu J, et al. Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey. arXiv preprint, 2024. DOI: 10.48550/arXiv.2403.14608.
[16] Gao L, Wang J, Zhou M. Ecological flow management and hydropower trade-offs: A review. Renewable and Sustainable Energy Reviews, 2023, 176: 113162.
[17] Shen D, Chen L, Zhao W. Hydropower development and ecological flow requirements in China. Ecological Indicators, 2020, 117: 106603.
[18] Li Q, Huang M, Zhang L. Ecological flow policies and river biodiversity protection in China. Water Policy, 2019, 21(4): 791-805.
[19] Chen Y, Wu G, Liu Y. River health assessment and ecological flow requirements: Progress and prospects. Journal of Hydrology, 2021, 599: 126381.
[20] Xu P, Zhang L. Hydrological monitoring methods and challenges in ecological flow management. Hydrological Sciences Journal, 2021, 66(5): 747-758.
[21] Zhou W, Fang Z, Liu C. Monitoring ecological flow in small rivers: Approaches and limitations. Ecological Indicators, 2020, 110: 105948.
[22] Tran D, Bourdev L, Fergus R, et al. Learning spatiotemporal features with 3D convolutional networks. 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 2015: 4489-4497. DOI: 10.1109/ICCV.2015.510.
[23] Chen Z, Yang H, Liu W. Video-based monitoring of river flow using deep learning. Water, 2021, 13(21): 3075.
[24] Oquab M, Darcet T, Moutakanni T, et al. DINOv2: Learning Robust Visual Features without Supervision. arXiv preprint, 2023. DOI: 10.48550/arXiv.2304.07193.
[25] Liu S Y, Wang C Y, Yin H, et al. DoRA: Weight-Decomposed Low-Rank Adaptation. Proceedings of the 41st International Conference on Machine Learning (ICML), PMLR, 2024, 235: 32100-32121.
[26] Meng F, Wang Z, Zhang M. PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models. NIPS '24: Proceedings of the 38th International Conference on Neural Information Processing Systems, 2024, 3846: 121038-121072.
[27] Zhang Q, Chen M, Bukharin A, et al. AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning. International Conference on Learning Representations (ICLR), 2023.
[28] Dettmers T, Pagnoni A, Holtzman A, et al. QLoRA: efficient finetuning of quantized LLMs. NIPS '23: Proceedings of the 37th International Conference on Neural Information Processing Systems, 2023, 441: 10088-10115.
[29] Zhang R, Gui L, Sun Z, et al. Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward. arXiv preprint, 2024. DOI: 10.48550/arXiv.2404.01258.
[30] Li K, He Y, Wang Y, et al. VideoChat: Chat-Centric Video Understanding. arXiv preprint, 2023. DOI: 10.48550/arXiv.2305.06355.
[31] Chen Z, Qi Z, Cso X, et al. Class-level Structural Relation Modelling and Smoothing for Visual Representation Learning. arXiv preprint, 2023. DOI: 10.48550/arXiv.2308.04142.
[32] Houlsby N, Giurgiu A, Jastrzebski S, et al. Parameter-efficient transfer learning for NLP. arXiv preprint, 2019. DOI: 10.48550/arXiv.1902.00751.
[33] Chen Z, Wang W, Cao Y, et al. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv preprint, 2024. DOI: 10.48550/arXiv.2412.05271.