Document Type : English paper for special issue on "climate change & insurance industry"

Author

Deputy Head of Life Insurance Department, Taavon Insurance Co.

Abstract

The escalating frequency and severity of hydro-meteorological events, particularly floods driven by climate change, present a profound systemic challenge to the global insurance and reinsurance industry. Traditional actuarial pricing models, predominantly based on Generalized Linear Models (GLMs) and historical loss tables, rely on the fundamental assumption of risk stationarity. This assumption has been rendered obsolete by the rapidly changing climate, making legacy models increasingly incapable of capturing the complex, non-linear, and spatially-interconnected nature of modern flood risk. In response to these limitations, standard data-driven Deep Learning (DL) models, such as Convolutional Neural Networks (CNNs), have been proposed as superior alternatives for risk forecasting. However, despite their theoretical promise, these models suffer from two critical flaws that strictly hinder their operational adoption in real-world insurance underwriting. First, they are notoriously "data-hungry," requiring massive, high-resolution datasets of historical losses to generalize effectively. Such datasets are unavailable in most "data-scarce markets," including Iran and much of the Global South, leading to poor model generalization and high predictive uncertainty. Second, their "black-box" nature is fundamentally incompatible with the stringent regulatory and fiduciary requirements of the insurance industry, which demands model transparency, fairness, and explainability (XAI) for pricing decisions. A model that cannot causally explain why it assigns a high premium to a specific property is operationally and ethically unusable. This study addresses this critical research and implementation gap. The primary objective is to design and validate a novel conceptual framework, "PICA" (Physics-Informed and Causally-Aware), for flood risk pricing. This study aims to achieve four specific goals: (1) To formally design the PICA framework by synthesizing state-of-the-art components from diverse AI fields; (2) To specifically address the data-scarcity problem by embedding hydrodynamic physical laws directly into the model’s learning process (Physics-Informed); (3) To resolve the "black-box" problem by integrating causal inference techniques to ensure model outputs are explainable and aligned with real-world causality (Causally-Aware); and (4) To empirically validate the performance of the PICA framework against benchmark DL models, particularly under simulated data-scarce conditions. In this study we used a rigorous multi-phase, mixed-methods research design, combining a systematic review with a high-fidelity quantitative simulation. The first phase, Framework Development, utilized a systematic review following the PRISMA 2020 protocol. Data sources included Scopus, Web of Science, and IEEE Xplore (2020–2025). The analysis focused on "architectural synthesis" of 52 elite articles covering Physics-Informed Neural Networks (PINNs), Graph Neural Networks (GNNs), and Causal Machine Learning. This synthesis directly informed the novel mathematical design of the PICA framework. The second phase, Framework Validation, consisted of a quantitative simulation experiment to test the hypothesis under controlled conditions. A high-fidelity synthetic dataset was generated using the HEC-RAS 2D hydrodynamic model, simulating 50 distinct stochastic flood events over a virtual 100-square-kilometer urban watershed containing 10,000 unique property assets. The primary material was the PICA framework itself, architected as a Hybrid Graph Neural Network (GNN) to explicitly capture the network topology of flood propagation. Two benchmark models were constructed for comparison: (1) a state-of-the-art 2D-Convolutional Neural Network (CNN) and (2) a baseline Random Forest (RF) model. The experimental procedure involved training all models under two distinct scenarios: a "Data-Rich" scenario (using 100% of simulated loss data) and a "Data-Scarce" scenario (using only 20% of available data). Predictive accuracy was evaluated using Root Mean Squared Error (RMSE) and Mean Absolute Error (MAE), with statistical significance determined by the Diebold-Mariano (DM) test. Additionally, Shapley Additive Explanations (SHAP) analysis was conducted to validate model explainability. The Phase 1 review confirmed that while GNNs, PINNs, and Causal ML are individually established, no existing framework integrates all three for actuarial pricing. This led to the formalization of PICA's novel composite loss function: L_total=L_data+λ_phys*L_phys+λ_causal*L_causal. Here, L_phys penalizes violations of the Shallow Water Equations, and L_causal penalizes violations of a predefined causal graph. In Phase 2 validation, quantitative results were significant. In the "Data-Rich" scenario, the PICA framework (RMSE = 0.119) modestly outperformed the CNN (RMSE = 0.137). However, the critical finding emerged in the "Data-Scarce" scenario. The performance of the data-hungry benchmarks collapsed, with the CNN's RMSE deteriorating to 0.482. In stark contrast, the PICA framework leveraged its physics-informed component to maintain robust accuracy (RMSE = 0.165). This represents a 65.8% reduction in prediction error compared to the standard CNN in data-scarce conditions. The Diebold-Mariano test confirmed this superiority was statistically significant (p<0.001). Furthermore, SHAP analysis revealed that PICA correctly identified 'Water Depth' and 'Flow Velocity' as primary risk drivers, whereas the CNN relied on spurious spatial artifacts, confirming PICA effectively resolves the black-box problem. This research successfully validates the PICA framework, demonstrating that the strategic synthesis of physics-informed, causally-aware, and graph-based deep learning provides a superior solution for flood risk pricing in data-scarce markets. The findings provide definitive proof that "teaching" a model the underlying physics of a peril can effectively substitute for the lack of historical data.

Keywords

Letters to Editor


IJIR Journal welcomes letters to the editor for the post-publication discussions and corrections which allows debate post publication on its site, through the Letters to Editor. Letters pertaining to manuscript published in IJIR should be sent to the editorial office of IJIR within three months of either online publication or before printed publication, except for critiques of original research. Following points are to be considering before sending the letters (comments) to the editor.

[1] Letters that include statements of statistics, facts, research, or theories should include appropriate references, although more than three are discouraged.

[2] Letters that are personal attacks on an author rather than thoughtful criticism of the author’s ideas will not be considered for publication.

[3] Letters can be no more than 300 words in length.

[4] Letter writers should include a statement at the beginning of the letter stating that it is being submitted either for publication or not.

[5] Anonymous letters will not be considered.

[6] Letter writers must include their city and state of residence or work.

[7] Letters will be edited for clarity and length.
CAPTCHA Image