فصلنامه علمی

نوع مقاله : مقاله پژوهشی

نویسندگان

1 گروه مدیریت فن‌آوری اطلاعات، دانشکده مدیریت، دانشگاه خوارزمی، تهران، ایران

2 استادیار گروه بانک، بیمه و گمرک، دانشکده مدیریت، دانشگاه خوارزمی

3 گروه بیمه اشخاص، پژوهشکده بیمه، تهران، ایران

چکیده

پیشینه و اهداف: افزایش چشمگیر خسارت‌های بیمه درمان در سال‌های اخیر و رشد انواع رفتارهای متقلبانه از سوی بیمه‌شدگان، پزشکان و سایر ارائه‌دهندگان خدمات درمانی، ضرورت بهره‌گیری از روش‌های هوشمند برای کنترل و شناسایی تقلب را دوچندان کرده است. محدودیت روش‌های سنتی، نبود داده‌های برچسب‌گذاری‌شده، و پیچیدگی الگوهای تقلب، چالش‌های اصلی شرکت‌های بیمه به‌ویژه در ایران به‌شمار می‌روند. پژوهش حاضر با هدف طراحی چارچوبی هوشمند، ماژولار و قابل‌گسترش برای کشف تقلب در بیمه درمان انجام شده است تا با تحلیل داده‌های نسخه‌های پزشکی و استخراج ویژگی‌های مختلف، امکان شناسایی دقیق‌تر انواع تقلب فراهم شود.
روش‌شناسی: این پژوهش چارچوبی برای کشف تقلب در بیمه درمان بر پایه الگوریتم جنگل ایزوله، همراه با مهندسی ویژگی، یادگیری‌ماشین بدون نظارت و تحلیل رفتارهای غیرعادی، توسعه داد. در مرحله نخست، انواع رایج تقلب و ویژگی‌های نمایانگر آن‌ها با همکاری کارشناسان بیمه و پزشکان شناسایی شد. در مرحله دوم، داده‌های خام نسخه‌ها و مطالبات در یک انبار داده دومرحله‌ای پاک‌سازی، یکپارچه و ساختاردهی شدند. در مرحله سوم، ۲۰ویژگی کلیدی شامل فراوانی تعامل بیمار و پزشک، ناسازگاری سن یا جنسیت با خدمت یا تشخیص ثبت‌شده، رشد غیرعادی هزینه‌های درمانی، نادربودن برخی خدمات، تکرار نامنظم رویه‌های خاص و عدم انطباق میان تخصص ارائه‌دهنده و خدمت ارائه‌شده استخراج شد. سپس الگوریتم جنگل ایزوله برای شناسایی نسخه‌های پرت و برآورد احتمال تقلب آن‌ها به‌کار رفت. برای افزایش تفسیرپذیری نتایج، از تحلیل مؤلفه‌های اصلی جهت مصورسازی استفاده شد. در نهایت، یک سامانه نرم‌افزاری عملیاتی همراه با داشبوردهای مدیریتی برای پایش، بررسی و تحلیل تقلب از منظر جغرافیایی، جمعیتی و زمانی پیاده‌سازی شد.
یافته‌ها: نتایج نشان داد الگوریتم‌های شناسایی ناهنجاری با اثربخشی بالا قادر به شناسایی الگوهای رفتاری غیرعادی در داده‌های بیمه درمان هستند. افزایش ناگهانی هزینه‌های مرتبط با پزشک، تکرار غیرمعمول خدمات توسط یک ارائه‌دهنده، ناسازگاری میان نوع خدمت و تخصص پزشکی، و الگوهای غیرعادی هزینه در بیماران، از مهم‌ترین شاخص‌های تمایز نسخه‌های مشکوک از رکوردهای عادی بودند. همچنین، تحلیل‌های مبتنی بر مؤلفه‌های اصلی نشان داد که مشاهدات ناهنجار غالباً در حاشیه ساختار داده قرار می‌گیرند؛ موضوعی که مناسب بودن روش‌های بدون نظارت را برای تفکیک موارد مشکوک از مطالبات عادی تأیید می‌کند. افزونبر‌این، الگوریتم جنگل ایزوله از نظر دقت شناسایی و کارایی محاسباتی، عملکردی بهتر از مدل خودرمزگذار داشت. نرم‌افزار توسعه‌یافته نیز امکان تحلیل چندبعدی تقلب در طول زمان، میان مناطق مختلف و در گروه‌های جمعیتی گوناگون را فراهم ساخت.
نتیجه‌گیری: چارچوب پیشنهادی نشان داد که تلفیق مهندسی ویژگی، شناسایی ناهنجاری و معماری نرم‌افزاری پیشرفته می‌تواند راهکاری عملی و مؤثر برای کشف تقلب در بیمه درمان فراهم آورد. دقت مناسب مدل و توانایی آن در عملکرد بدون داده‌های برچسب‌گذاری‌شده، این چارچوب را در شرایطی که سوابق تأییدشده تقلب محدود یا در دسترس نیستند، ارزشمند می‌سازد. همچنین، ساختار ماژولار آن امکان به‌روزرسانی مستمر، توسعه ویژگی‌ها و انطباق با الگوهای متغیر تقلب را فراهم می‌کند. این سامانه می‌تواند به بهبود مدیریت ریسک، پایش کارآمدتر مطالبات و کاهش خسارت‌های مالی در بیمه درمان کمک کند. پیشنهاد می‌شود مطالعات آینده با بهره‌گیری از منابع داده متنوع‌تر، مدل‌های عمیق‌تر و تحلیل شبکه‌ای، عملکرد و استحکام سامانه‌های کشف تقلب را بیش از پیش ارتقا دهند.

کلیدواژه‌ها

موضوعات

عنوان مقاله [English]

AI-Enabled Fraud Detection in Health Insurance: A Multi-Module Framework Based on Anomaly Analysis

نویسندگان [English]

  • Mojtaba Farrokh 1
  • Sirus Sharifi 2
  • Nasrin Hozarmoghadam 3

1 Assistant Professor, Operations and Infromation Technology Management Department, Faculty of Management, Kharazmi University, Tehran, Iran

2 Assistant Professor, Department of Bank, Insurance and Customs, Faculty of Management, Kharazmi University

3 Department of Personal Insurance, Insurance Research Center, Tehran, Iran

چکیده [English]

BACKGROUND AND OBJECTIVES: The dramatic increase in health insurance losses in recent years and the growth of various fraudulent behaviors by policyholders, physicians, and other healthcare providers have doubled the necessity of utilizing intelligent methods for controlling and detecting fraud. The limitations of traditional methods, the absence of labeled data, and the complexity of fraud patterns represent the primary challenges faced by insurance companies, particularly in Iran. Healthcare fraud manifests in multifaceted ways, requiring sophisticated detection mechanisms. The literature identifies practices such as incorrect coding, upcoding, submitting duplicate bills, and charging for unnecessary services. Fraud is often a collaborative effort or anomaly involving patients, doctors, and institutions. Because historical data lacks explicit fraud labels, supervised learning models are often impractical for deployments. Consequently, this study leverages unsupervised learning to identify cases that deviate from established behaviors. The theoretical foundation positions the isolation forest and autoencoder algorithms as methodologies for isolating outliers. This research was conducted with the objective of designing and evaluating an intelligent, scalable, and modular framework tailored for detecting health insurance fraud. By deeply analyzing medical prescription data and extracting diverse behavioral features, this research enables the automated and more accurate identification of various types of fraudulent activities. The absolute ultimate goal is to transition from reactive auditing to a proactive system, thereby enhancing operational efficiency and significantly reducing financial leakage for insurance providers.

METHODOLOGY: The current research focuses on developing a health insurance fraud detection framework based on the Isolation Forest (IF) algorithm, combining feature engineering, unsupervised machine learning, and the analysis of abnormal behaviors. Initially, in collaboration with insurance experts and physicians, various types of fraud and their representative features were identified. Raw data underwent cleaning and structuring within a two-stage data warehouse. Subsequently, 20 key features were extracted and utilized to train the models. These features captured dynamics such as patient-physician interaction frequency, inconsistencies of age and gender with the provided service or diagnosis, abnormal cost growth, and service rarity. The IF algorithm was deployed to cluster prescriptions and determine fraud probabilities, and the results were visualized using Principal Component Analysis (PCA). Finally, an operational software system was implemented, featuring managerial dashboards based directly on the aforementioned IF algorithm.

FINDINGS: The empirical results demonstrate that anomaly detection algorithms effectively uncover abnormal behavioral patterns with high accuracy. Key indicators such as sudden spikes in physician costs, unusual frequency of services by a single provider, inconsistencies between service types and professional specialties, and abnormal patient spending patterns proved crucial in identifying suspicious prescriptions. Principal Component Analysis (PCA) visualizations revealed that anomalies are primarily distributed at the extreme edges of the data, confirming that unsupervised models possess a robust capability to distinguish these from normal records. Comparatively, the Isolation Forest (IF) algorithm significantly outperformed the Autoencoder (AE) model in both predictive accuracy and computational efficiency. The deployed software, built with a RESTful architecture, MariaDB, and a React-based interface, incorporates 20 specific risk indicators. This system provides managerial dashboards that significantly reduce the time required for experts to analyze high-risk claims. The interactive nature of these tools allows for the ongoing guidance of training processes, parameter tuning, and detailed actor behavior analysis. Furthermore, the study emphasizes that the quality of extracted features is a decisive factor in algorithmic performance; precise selection through expert collaboration leads to significant improvements in results.

CONCLUSION: This research designed and evaluated an intelligent, modular framework for health insurance fraud detection based on unsupervised learning. Addressing the inadequacy of traditional auditing in the face of massive data volumes and complex fraud patterns, the IF-based framework offers an independent, configurable structure that adapts to the dynamic nature of non-conforming behaviors. The system’s modular architecture, comprehensive documentation, and seamless integration capabilities make it a practical and scalable solution for insurance companies. Implementing such intelligent systems significantly reduces fraud-related losses and enhances operational efficiency. By analyzing historical data and identifying deviations, these tools enable the proactive detection of suspicious cases, thereby improving decision-making and resource allocation. However, system performance requires ongoing development of features related to actors, transactions, and communication networks to capture network-based risks. Continuous model updates and real-world feedback loops are essential for maintaining adaptability. Despite these strengths, the study acknowledges limitations including data quality issues (noise and incompleteness), reliance on a single insurance company, and legal/privacy constraints that challenge real-world implementation. Future research should focus on deeper expert involvement in defining fraud types, the use of meta-heuristic algorithms for feature engineering, and the exploration of graph-based and deep learning models to identify complex collusive rings. Additionally, integrating Natural Language Processing (NLP) to analyze unstructured medical notes and expanding the framework to other domains, such as auto insurance and credit risk, represent promising paths for future development.

کلیدواژه‌ها [English]

  • Anomaly-Detection
  • Artificial intelligence
  • Fraud-Detection
  • Health insurance
  • Unsupervised learning

نامه به سردبیر


سردبیر نشریه پژوهشنامه بیمه، هرگونه پیشنهاد و انتقاد دیگر نویسندگان و خوانندگان را در خصوص نقد و بررسی این مقاله مندرج در سامانه نشریه را ظرف مدت 3 ماه از تاریخ انتشار آنلاین مقاله در سامانه و قبل از انتشار چاپی نشریه، به منظور اصلاح و نظردهی امکان پذیر نموده است.، البته این نقد در مورد تحقیقات اصلی مقاله نمی باشد.
توجه به موارد ذیل پیش از ارسال نامه به سردبیر لازم است در نظر گرفته شود:
[1] نامه هایی که شامل گزارش آماری، واقعیت ها، تحقیقات یا نظریه پردازی ها هستند، لازم است همراه با منابع معتبر و مناسب همراه باشد، اگرچه ارسال بیش از زمان 3 نامه توصیه نمی گردد.
[2] نامه هایی که بجای انتقاد سازنده به ایده های تحقیق، مشتمل بر حملات شخصی به نویسنده باشند، توجه و چاپ نمی شود.
[3] نامه ها نباید بیش از 300 کلمه باشد.
[4] نویسندگان نامه لازم است در ابتدای نامه تمایل یا عدم تمایل خود را نسبت به چاپ نظریه ارسالی نسبت به یک مقاله خاص اعلام نمایند.
[5] به نامه های ناشناس ترتیب اثر داده نمی شود.
[6] شهر، کشور و محل سکونت نویسندگان نامه باید در نامه مشخص باشد.
[7] به منظور شفافیت بیشتر و محدودیت حجم نامه، ویرایش بر روی آن انجام می پذیرد.


 

CAPTCHA Image