Document Type : Original Research Paper
Authors
1 Assistant Professor, Operations and Infromation Technology Management Department, Faculty of Management, Kharazmi University, Tehran, Iran
2 Assistant Professor, Department of Bank, Insurance and Customs, Faculty of Management, Kharazmi University
3 Department of Personal Insurance, Insurance Research Center, Tehran, Iran
Abstract
BACKGROUND AND OBJECTIVES: The dramatic increase in health insurance losses in recent years and the growth of various fraudulent behaviors by policyholders, physicians, and other healthcare providers have doubled the necessity of utilizing intelligent methods for controlling and detecting fraud. The limitations of traditional methods, the absence of labeled data, and the complexity of fraud patterns represent the primary challenges faced by insurance companies, particularly in Iran. Healthcare fraud manifests in multifaceted ways, requiring sophisticated detection mechanisms. The literature identifies practices such as incorrect coding, upcoding, submitting duplicate bills, and charging for unnecessary services. Fraud is often a collaborative effort or anomaly involving patients, doctors, and institutions. Because historical data lacks explicit fraud labels, supervised learning models are often impractical for deployments. Consequently, this study leverages unsupervised learning to identify cases that deviate from established behaviors. The theoretical foundation positions the isolation forest and autoencoder algorithms as methodologies for isolating outliers. This research was conducted with the objective of designing and evaluating an intelligent, scalable, and modular framework tailored for detecting health insurance fraud. By deeply analyzing medical prescription data and extracting diverse behavioral features, this research enables the automated and more accurate identification of various types of fraudulent activities. The absolute ultimate goal is to transition from reactive auditing to a proactive system, thereby enhancing operational efficiency and significantly reducing financial leakage for insurance providers.
METHODOLOGY: The current research focuses on developing a health insurance fraud detection framework based on the Isolation Forest (IF) algorithm, combining feature engineering, unsupervised machine learning, and the analysis of abnormal behaviors. Initially, in collaboration with insurance experts and physicians, various types of fraud and their representative features were identified. Raw data underwent cleaning and structuring within a two-stage data warehouse. Subsequently, 20 key features were extracted and utilized to train the models. These features captured dynamics such as patient-physician interaction frequency, inconsistencies of age and gender with the provided service or diagnosis, abnormal cost growth, and service rarity. The IF algorithm was deployed to cluster prescriptions and determine fraud probabilities, and the results were visualized using Principal Component Analysis (PCA). Finally, an operational software system was implemented, featuring managerial dashboards based directly on the aforementioned IF algorithm.
FINDINGS: The empirical results demonstrate that anomaly detection algorithms effectively uncover abnormal behavioral patterns with high accuracy. Key indicators such as sudden spikes in physician costs, unusual frequency of services by a single provider, inconsistencies between service types and professional specialties, and abnormal patient spending patterns proved crucial in identifying suspicious prescriptions. Principal Component Analysis (PCA) visualizations revealed that anomalies are primarily distributed at the extreme edges of the data, confirming that unsupervised models possess a robust capability to distinguish these from normal records. Comparatively, the Isolation Forest (IF) algorithm significantly outperformed the Autoencoder (AE) model in both predictive accuracy and computational efficiency. The deployed software, built with a RESTful architecture, MariaDB, and a React-based interface, incorporates 20 specific risk indicators. This system provides managerial dashboards that significantly reduce the time required for experts to analyze high-risk claims. The interactive nature of these tools allows for the ongoing guidance of training processes, parameter tuning, and detailed actor behavior analysis. Furthermore, the study emphasizes that the quality of extracted features is a decisive factor in algorithmic performance; precise selection through expert collaboration leads to significant improvements in results.
CONCLUSION: This research designed and evaluated an intelligent, modular framework for health insurance fraud detection based on unsupervised learning. Addressing the inadequacy of traditional auditing in the face of massive data volumes and complex fraud patterns, the IF-based framework offers an independent, configurable structure that adapts to the dynamic nature of non-conforming behaviors. The system’s modular architecture, comprehensive documentation, and seamless integration capabilities make it a practical and scalable solution for insurance companies. Implementing such intelligent systems significantly reduces fraud-related losses and enhances operational efficiency. By analyzing historical data and identifying deviations, these tools enable the proactive detection of suspicious cases, thereby improving decision-making and resource allocation. However, system performance requires ongoing development of features related to actors, transactions, and communication networks to capture network-based risks. Continuous model updates and real-world feedback loops are essential for maintaining adaptability. Despite these strengths, the study acknowledges limitations including data quality issues (noise and incompleteness), reliance on a single insurance company, and legal/privacy constraints that challenge real-world implementation. Future research should focus on deeper expert involvement in defining fraud types, the use of meta-heuristic algorithms for feature engineering, and the exploration of graph-based and deep learning models to identify complex collusive rings. Additionally, integrating Natural Language Processing (NLP) to analyze unstructured medical notes and expanding the framework to other domains, such as auto insurance and credit risk, represent promising paths for future development.
Keywords
Main Subjects
Letters to Editor
Send comment about this article