نوع مقاله : مقاله پژوهشی
نویسندگان
1 گروه مدیریت فنآوری اطلاعات، دانشکده مدیریت، دانشگاه خوارزمی، تهران، ایران
2 استادیار گروه بانک، بیمه و گمرک، دانشکده مدیریت، دانشگاه خوارزمی
3 گروه بیمه اشخاص، پژوهشکده بیمه، تهران، ایران
چکیده
پیشینه و اهداف: افزایش چشمگیر خسارتهای بیمه درمان در سالهای اخیر و رشد انواع رفتارهای متقلبانه از سوی بیمهشدگان، پزشکان و سایر ارائهدهندگان خدمات درمانی، ضرورت بهرهگیری از روشهای هوشمند برای کنترل و شناسایی تقلب را دوچندان کرده است. محدودیت روشهای سنتی، نبود دادههای برچسبگذاریشده، و پیچیدگی الگوهای تقلب، چالشهای اصلی شرکتهای بیمه بهویژه در ایران بهشمار میروند. پژوهش حاضر با هدف طراحی چارچوبی هوشمند، ماژولار و قابلگسترش برای کشف تقلب در بیمه درمان انجام شده است تا با تحلیل دادههای نسخههای پزشکی و استخراج ویژگیهای مختلف، امکان شناسایی دقیقتر انواع تقلب فراهم شود.
روششناسی: این پژوهش چارچوبی برای کشف تقلب در بیمه درمان بر پایه الگوریتم جنگل ایزوله، همراه با مهندسی ویژگی، یادگیریماشین بدون نظارت و تحلیل رفتارهای غیرعادی، توسعه داد. در مرحله نخست، انواع رایج تقلب و ویژگیهای نمایانگر آنها با همکاری کارشناسان بیمه و پزشکان شناسایی شد. در مرحله دوم، دادههای خام نسخهها و مطالبات در یک انبار داده دومرحلهای پاکسازی، یکپارچه و ساختاردهی شدند. در مرحله سوم، ۲۰ویژگی کلیدی شامل فراوانی تعامل بیمار و پزشک، ناسازگاری سن یا جنسیت با خدمت یا تشخیص ثبتشده، رشد غیرعادی هزینههای درمانی، نادربودن برخی خدمات، تکرار نامنظم رویههای خاص و عدم انطباق میان تخصص ارائهدهنده و خدمت ارائهشده استخراج شد. سپس الگوریتم جنگل ایزوله برای شناسایی نسخههای پرت و برآورد احتمال تقلب آنها بهکار رفت. برای افزایش تفسیرپذیری نتایج، از تحلیل مؤلفههای اصلی جهت مصورسازی استفاده شد. در نهایت، یک سامانه نرمافزاری عملیاتی همراه با داشبوردهای مدیریتی برای پایش، بررسی و تحلیل تقلب از منظر جغرافیایی، جمعیتی و زمانی پیادهسازی شد.
یافتهها: نتایج نشان داد الگوریتمهای شناسایی ناهنجاری با اثربخشی بالا قادر به شناسایی الگوهای رفتاری غیرعادی در دادههای بیمه درمان هستند. افزایش ناگهانی هزینههای مرتبط با پزشک، تکرار غیرمعمول خدمات توسط یک ارائهدهنده، ناسازگاری میان نوع خدمت و تخصص پزشکی، و الگوهای غیرعادی هزینه در بیماران، از مهمترین شاخصهای تمایز نسخههای مشکوک از رکوردهای عادی بودند. همچنین، تحلیلهای مبتنی بر مؤلفههای اصلی نشان داد که مشاهدات ناهنجار غالباً در حاشیه ساختار داده قرار میگیرند؛ موضوعی که مناسب بودن روشهای بدون نظارت را برای تفکیک موارد مشکوک از مطالبات عادی تأیید میکند. افزونبراین، الگوریتم جنگل ایزوله از نظر دقت شناسایی و کارایی محاسباتی، عملکردی بهتر از مدل خودرمزگذار داشت. نرمافزار توسعهیافته نیز امکان تحلیل چندبعدی تقلب در طول زمان، میان مناطق مختلف و در گروههای جمعیتی گوناگون را فراهم ساخت.
نتیجهگیری: چارچوب پیشنهادی نشان داد که تلفیق مهندسی ویژگی، شناسایی ناهنجاری و معماری نرمافزاری پیشرفته میتواند راهکاری عملی و مؤثر برای کشف تقلب در بیمه درمان فراهم آورد. دقت مناسب مدل و توانایی آن در عملکرد بدون دادههای برچسبگذاریشده، این چارچوب را در شرایطی که سوابق تأییدشده تقلب محدود یا در دسترس نیستند، ارزشمند میسازد. همچنین، ساختار ماژولار آن امکان بهروزرسانی مستمر، توسعه ویژگیها و انطباق با الگوهای متغیر تقلب را فراهم میکند. این سامانه میتواند به بهبود مدیریت ریسک، پایش کارآمدتر مطالبات و کاهش خسارتهای مالی در بیمه درمان کمک کند. پیشنهاد میشود مطالعات آینده با بهرهگیری از منابع داده متنوعتر، مدلهای عمیقتر و تحلیل شبکهای، عملکرد و استحکام سامانههای کشف تقلب را بیش از پیش ارتقا دهند.
کلیدواژهها
موضوعات
عنوان مقاله [English]
AI-Enabled Fraud Detection in Health Insurance: A Multi-Module Framework Based on Anomaly Analysis
نویسندگان [English]
- Mojtaba Farrokh 1
- Sirus Sharifi 2
- Nasrin Hozarmoghadam 3
1 Assistant Professor, Operations and Infromation Technology Management Department, Faculty of Management, Kharazmi University, Tehran, Iran
2 Assistant Professor, Department of Bank, Insurance and Customs, Faculty of Management, Kharazmi University
3 Department of Personal Insurance, Insurance Research Center, Tehran, Iran
چکیده [English]
BACKGROUND AND OBJECTIVES: The dramatic increase in health insurance losses in recent years and the growth of various fraudulent behaviors by policyholders, physicians, and other healthcare providers have doubled the necessity of utilizing intelligent methods for controlling and detecting fraud. The limitations of traditional methods, the absence of labeled data, and the complexity of fraud patterns represent the primary challenges faced by insurance companies, particularly in Iran. Healthcare fraud manifests in multifaceted ways, requiring sophisticated detection mechanisms. The literature identifies practices such as incorrect coding, upcoding, submitting duplicate bills, and charging for unnecessary services. Fraud is often a collaborative effort or anomaly involving patients, doctors, and institutions. Because historical data lacks explicit fraud labels, supervised learning models are often impractical for deployments. Consequently, this study leverages unsupervised learning to identify cases that deviate from established behaviors. The theoretical foundation positions the isolation forest and autoencoder algorithms as methodologies for isolating outliers. This research was conducted with the objective of designing and evaluating an intelligent, scalable, and modular framework tailored for detecting health insurance fraud. By deeply analyzing medical prescription data and extracting diverse behavioral features, this research enables the automated and more accurate identification of various types of fraudulent activities. The absolute ultimate goal is to transition from reactive auditing to a proactive system, thereby enhancing operational efficiency and significantly reducing financial leakage for insurance providers.
METHODOLOGY: The current research focuses on developing a health insurance fraud detection framework based on the Isolation Forest (IF) algorithm, combining feature engineering, unsupervised machine learning, and the analysis of abnormal behaviors. Initially, in collaboration with insurance experts and physicians, various types of fraud and their representative features were identified. Raw data underwent cleaning and structuring within a two-stage data warehouse. Subsequently, 20 key features were extracted and utilized to train the models. These features captured dynamics such as patient-physician interaction frequency, inconsistencies of age and gender with the provided service or diagnosis, abnormal cost growth, and service rarity. The IF algorithm was deployed to cluster prescriptions and determine fraud probabilities, and the results were visualized using Principal Component Analysis (PCA). Finally, an operational software system was implemented, featuring managerial dashboards based directly on the aforementioned IF algorithm.
FINDINGS: The empirical results demonstrate that anomaly detection algorithms effectively uncover abnormal behavioral patterns with high accuracy. Key indicators such as sudden spikes in physician costs, unusual frequency of services by a single provider, inconsistencies between service types and professional specialties, and abnormal patient spending patterns proved crucial in identifying suspicious prescriptions. Principal Component Analysis (PCA) visualizations revealed that anomalies are primarily distributed at the extreme edges of the data, confirming that unsupervised models possess a robust capability to distinguish these from normal records. Comparatively, the Isolation Forest (IF) algorithm significantly outperformed the Autoencoder (AE) model in both predictive accuracy and computational efficiency. The deployed software, built with a RESTful architecture, MariaDB, and a React-based interface, incorporates 20 specific risk indicators. This system provides managerial dashboards that significantly reduce the time required for experts to analyze high-risk claims. The interactive nature of these tools allows for the ongoing guidance of training processes, parameter tuning, and detailed actor behavior analysis. Furthermore, the study emphasizes that the quality of extracted features is a decisive factor in algorithmic performance; precise selection through expert collaboration leads to significant improvements in results.
CONCLUSION: This research designed and evaluated an intelligent, modular framework for health insurance fraud detection based on unsupervised learning. Addressing the inadequacy of traditional auditing in the face of massive data volumes and complex fraud patterns, the IF-based framework offers an independent, configurable structure that adapts to the dynamic nature of non-conforming behaviors. The system’s modular architecture, comprehensive documentation, and seamless integration capabilities make it a practical and scalable solution for insurance companies. Implementing such intelligent systems significantly reduces fraud-related losses and enhances operational efficiency. By analyzing historical data and identifying deviations, these tools enable the proactive detection of suspicious cases, thereby improving decision-making and resource allocation. However, system performance requires ongoing development of features related to actors, transactions, and communication networks to capture network-based risks. Continuous model updates and real-world feedback loops are essential for maintaining adaptability. Despite these strengths, the study acknowledges limitations including data quality issues (noise and incompleteness), reliance on a single insurance company, and legal/privacy constraints that challenge real-world implementation. Future research should focus on deeper expert involvement in defining fraud types, the use of meta-heuristic algorithms for feature engineering, and the exploration of graph-based and deep learning models to identify complex collusive rings. Additionally, integrating Natural Language Processing (NLP) to analyze unstructured medical notes and expanding the framework to other domains, such as auto insurance and credit risk, represent promising paths for future development.
کلیدواژهها [English]
- Anomaly-Detection
- Artificial intelligence
- Fraud-Detection
- Health insurance
- Unsupervised learning
نامه به سردبیر
سردبیر نشریه پژوهشنامه بیمه، هرگونه پیشنهاد و انتقاد دیگر نویسندگان و خوانندگان را در خصوص نقد و بررسی این مقاله مندرج در سامانه نشریه را ظرف مدت 3 ماه از تاریخ انتشار آنلاین مقاله در سامانه و قبل از انتشار چاپی نشریه، به منظور اصلاح و نظردهی امکان پذیر نموده است.، البته این نقد در مورد تحقیقات اصلی مقاله نمی باشد.
توجه به موارد ذیل پیش از ارسال نامه به سردبیر لازم است در نظر گرفته شود:
[1] نامه هایی که شامل گزارش آماری، واقعیت ها، تحقیقات یا نظریه پردازی ها هستند، لازم است همراه با منابع معتبر و مناسب همراه باشد، اگرچه ارسال بیش از زمان 3 نامه توصیه نمی گردد.
[2] نامه هایی که بجای انتقاد سازنده به ایده های تحقیق، مشتمل بر حملات شخصی به نویسنده باشند، توجه و چاپ نمی شود.
[3] نامه ها نباید بیش از 300 کلمه باشد.
[4] نویسندگان نامه لازم است در ابتدای نامه تمایل یا عدم تمایل خود را نسبت به چاپ نظریه ارسالی نسبت به یک مقاله خاص اعلام نمایند.
[5] به نامه های ناشناس ترتیب اثر داده نمی شود.
[6] شهر، کشور و محل سکونت نویسندگان نامه باید در نامه مشخص باشد.
[7] به منظور شفافیت بیشتر و محدودیت حجم نامه، ویرایش بر روی آن انجام می پذیرد.
ارسال نظر در مورد این مقاله