Contents
Download PDF
pdf Download XML
66 Views
40 Downloads
Share this article
Research Article | Volume 15 Issue 12 (Dec, 2025) | Pages 1383 - 1387
A Cross-Sectional Analysis of AI-Generated Educational Material for Medical Professionals Compared to UpToDate on Cardiac Catheterization
 ,
 ,
1
Medical Officer, Department of Prosthodontics, Bangabandhu Sheikh Mujib Medical University (BSMMU), Dhaka, Bangladesh
2
Associate Professor, Department of Prosthodontics, Bangladesh Dental College, Dhaka, Bangladesh.
3
Rama Medical College Hospital and Research Centre Rama City, NH-9, Delhi Meerut Expressway, Near Mother Dairy, Pilkhuwa, Hapur (U.P.) - 245304. India.
Under a Creative Commons license
Open Access
Received
Nov. 1, 2025
Revised
Nov. 15, 2025
Accepted
Dec. 25, 2025
Published
Dec. 28, 2025
Abstract

Background: For healthcare professionals, having access to precise and up-to-date educational materials is essential to deliver effective patient care, especially during intricate and high-risk procedures like cardiac catheterization. In response to the increasing demand for current educational resources, the rapid progress of artificial intelligence (AI) in the healthcare industry has brought attention to tools such as ChatGPT. Objectives: The objective of the study is to compare educational content produced by ChatGPT with established medical reference resource, UpToDate on Cardiac Catheterization.  Methodology: A Cross-Sectional study on Cardiac Catheterization was conducted, comparing educational material from UpToDate, with ChatGPT. Tools used for readability are Flesch Reading Ease (FRE), Flesch-Kincaid Grade Level (FKGL), and SMOG Index (Simple Measure of Gobbledygook). Statistical analysis performed using IBM SPSS and R. Wilcoxon signed rank test used, and p<0.05 was significant. Results: ChatGPT and UpToDate were compared for readability of educational content on cardiac catheterization. UpToDate generated more extensive content, demonstrating significant differences in word count (p = 0.0051), sentence count (p = 0.0064), and word-to-sentence ratio (p = 0.0202). It also had higher Flesch-Kincaid Grade Level (p = 0.045), Simple Measure of Gobbledygook index values (p = 0.045), and difficult word count (p = 0.0082). Conversely, ChatGPT had a higher proportion of difficult words (p = 0.0306). Conclusions: In comparison, ChatGPT generated more concise content, required slightly lower reading grade levels, yet utilized denser vocabulary, whereas UpToDate provided more detailed, relatively complex Material.

Keywords
INTRODUCTION

Cardiac catheterization is a cornerstone diagnostic and interventional procedure in cardiovascular medicine. The procedure involves placing a catheter into the heart chambers and coronary vessels, typically through a radial or femoral artery, to measure intracardiac pressures, oxygen saturation, and visualize coronary anatomy. This procedure is essential for diagnosing and treating coronary artery disease, valvular disorders, congenital heart defects, and cardiomyopathies. [1, 2]. Diagnostic catheterization provides hemodynamic data, while interventional techniques such as percutaneous coronary intervention (PCI) allow therapeutic management of obstructive lesions. Advancements such as intravascular ultrasound (IVUS), optical coherence tomography (OCT), and fractional flow reserve (FFR) have improved diagnostic accuracy and supported individualized patient care. [3]. Despite its widespread use, cardiac catheterization carries risks, including bleeding, arrhythmias, coronary dissection, vascular injury, and contrast-induced nephropathy [4]. Therefore, thorough knowledge of indications, procedural steps, and complication management is essential.

 

Continuous medical education in this field ensures clinicians remain updated with evolving procedural standards and evidence-based practices. Resources like UpToDate provide peer-reviewed, continually updated clinical summaries that integrate international guidelines, expert consensus, and the latest research findings [1, 2]. By consolidating vast medical evidence into accessible clinical guidance, UpToDate enhances decision-making accuracy, reduces diagnostic errors, and improves patient outcomes [5]. Its educational value is particularly significant for trainees and practicing clinicians seeking to maintain procedural competency and align with current best practices.

 

Artificial intelligence (AI) technologies, especially large language models (LLMs) such as ChatGPT, are transforming the dissemination and accessibility of medical information. These models use deep learning and natural language processing methods to generate coherent and contextually relevant text from large-scale training datasets. In medical education, AI is increasingly used for summarizing complex topics, generating clinical explanations, and simplifying evidence-based content for diverse learning levels [6]. The primary advantage lies in their speed, accessibility, and adaptability clinicians and students can obtain instant summaries tailored to their specific queries.

 

However, several concerns persist regarding their reliability. AI models may lack transparency about data sources, produce inaccurate or outdated information, or omit critical clinical nuances. In invasive procedures like cardiac catheterization where precision in describing normal hemodynamic, pressure gradients, or contraindications is crucial these inaccuracies could have significant educational and safety implications [6]. ChatGPT, developed by OpenAI, represents a widely used LLM capable of generating detailed explanations on cardiovascular procedures. While it demonstrates high linguistic fluency and readability, its factual accuracy and clinical reliability require systematic evaluation before integration into professional learning frameworks [5, 6].

 

This research aims to assess the readability and educational value of AI-generated content from ChatGPT by comparing it with evidence-based medical information from UpToDate on the subject of cardiac catheterization. The focus is on determining whether AI-generated material achieves greater readability without compromising medical accuracy. Readability metrics such as Flesch Reading Ease (FRE), Flesch–Kincaid Grade Level (FKGL) and Gunning Fog Index (GFI) will be employed to objectively evaluate text complexity, while content validity will be assessed by expert clinicians. This comparative analysis will identify whether ChatGPT-generated explanations maintain adequate clinical depth and conceptual clarity when compared to expert-authored reference material. The outcome will help establish whether AI tools can supplement traditional resources in professional education and whether they are capable of supporting lifelong learning in cardiovascular medicine without risking misinterpretation or misinformation [4–6]. The primary objective of this study is to compare the readability of educational content on cardiac catheterization generated by ChatGPT, an LLM, with content provided by UpToDate, a widely used evidence-based clinical reference. By employing validated readability assessment tools, the objective is to assess the clarity, complexity, and potential educational utility of AI-generated content within the context of contemporary medical education.

 

MATERIALS AND METHODS

This original research, designed as a cross-sectional study, was carried out over the course of one month, from September 1 to September 30 2025. Since the study did not include human participants, identifiable data, or any form of intervention, approval from an institutional ethics committee was not required [7,8].

The topic selected for analysis was Cardiac Catheterization. Educational content aimed at medical professionals was generated using ChatGPT 5 (assessed on September 11, 2025), a large language model developed specifically for scientific and medical domains. The AI tool was prompted with the instruction:

 

“Write an educational guide for medical professionals on cardiac catheterization, including definition, clinical features, diagnosis, and treatment options.” The generated output was compiled into a Microsoft Word document for evaluation.

For comparison, standard reference content on cardiac catheterization was retrieved from UpToDate (accessed on September 11, 2025), a widely used, evidence-based clinical decision support resource. To maintain consistency, only the main disease summary was included, while tables, references, and figure legends were omitted.

 

he readability of both texts was assessed using the Flesch Reading Ease (FRE) score and the Flesch–Kincaid Grade Level (FKGL), which were calculated using an online Flesch–Kincaid readability calculator. Parameters assessed included word count, sentence count, and average words per sentence, FRE Score, FKGL, and the proportion of difficult words. These metrics allowed for a structured comparison of the accessibility of the educational content for its intended audience i.e., medical professionals.

 

The data were gathered in Microsoft Excel, and statistical analyses were performed using IBM SPSS (version 25) [9] and R software [10] (version 4.3.2, R Core Team, 2023). The Wilcoxon signed-rank test was applied to compare the readability scores of ChatGPT and UpToDate content. A p-value below 0.05 was regarded as statistically significant.

RESULTS

ChatGPT and UpToDate were utilized to generate educational material on cardiac catheterization aimed at healthcare professionals. Content from both sources was analyzed and assessed for readability using parameters such as word count, sentence count, word-to-sentence ratio, Flesch Reading Ease (FRE), Flesch-Kincaid Grade Level (FKGL), SMOG Index, number of difficult words, and percentage of difficult words. The objective was to determine whether AI-generated content demonstrates readability and complexity comparable to a standard clinical reference source.

 

Table 1 presents the comparison of readability characteristics between ChatGPT and UpToDate for content on cardiac catheterization. Significant differences were observed in word count (p = 0.0051), sentence count (p = 0.0064), and word-to-sentence ratio (p = 0.0202), with UpToDate producing considerably longer content. UpToDate also showed higher FKGL (p = 0.045) and SMOG index (p = 0.045) values, reflecting a slightly higher reading grade level. Difficult word count was significantly higher for UpToDate (p = 0.0082), while the proportion of difficult words was greater for ChatGPT (p = 0.0306). No statistically significant difference was found in FRE scores (p = 0.229), suggesting comparable overall readability between the two sources.

 

Figure 1 illustrates the topic-wise comparison of FRE, FKGL, SMOG index, and difficult word percentage between ChatGPT and UpToDate across six subtopics under cardiac catheterization. FRE scores were generally higher for UpToDate, indicating greater ease of readability, with the largest difference observed in Topic 3 (UpToDate: 20.1 vs ChatGPT: 1.6). In contrast, ChatGPT consistently showed lower FKGL and SMOG index values, suggesting that its content required a slightly lower reading grade level. Difficult word percentage was higher in ChatGPT content for all six topics, with the greatest difference seen in Topic 2 (ChatGPT: 40.52% vs UpToDate: 28.45%), indicating proportionally denser use of complex vocabulary. These visual trends are consistent with the statistical results, underscoring that ChatGPT generates shorter, lower-grade-level content but with a higher density of difficult words, whereas UpToDate produces longer, more detailed material at a slightly higher reading level.

 

Tables and Figures

 

Table 1. Readability characteristics of educational content on cardiac catheterization   generated by UpToDate and ChatGPT

 Variable Median (IQR) Wilcoxon Statistic P value+  
UpToDate ChatGPT  
Word Count 4600.5 (3634.0-5041.0) 613.0 (575.0-926.0) 57 0.0051*
Sentence Count 194.0 (140.0-250.0) 76.0 (73.0-84.0) 56.5 0.0064*
Word/Sentence Count 21.1 (19.9-24.1) 9.5 (7.9-15.7) 54 0.0202*
FRE 20.1 (18.4-25.7) 16.4 (5.1-19.6) 47 0.229
FKGL 15.3 (15.2-15.5) 14.5 (13.4-15.0) 52 0.045*
SMOG Index 13.3 (13.1-13.8) 10.7 (10.4-12.3) 52 0.045*
Difficult Word Count 1082.5 (736.0-1349.0) 200.0 (152.0-245.0) 56 0.0082*
Difficult Word Percentage 27.6 (26.6-28.7) 31.6 (29.2-39.8) 25 0.0306*

 

+ Wilcoxon signed-rank Test. P-values <0.05 are regarded as statistically significant

Values are presented as Median (Interquartile Range, IQR) unless otherwise specified. Statistical comparison between groups was performed using the Wilcoxon signed-rank test. P-values <0.05 were considered statistically significant. Abbreviations: FRE – Flesch Reading Ease; FKGL – Flesch–Kincaid Grade Level; SMOG – Simple Measure of Gobbledygook Index.

Figure 1. Comparison of readability indices for educational content on cardiac catheterization   generated by UpToDate and ChatGPT

The figure illustrates the comparison of FRE, FKGL, SMOG Index, and difficult word percentage for educational content related to cardiac catheterization   generated by ChatGPT and UpToDate. Statistical comparison was performed using the Wilcoxon signed-rank test, and P-values <0.05 were considered statistically significant. Abbreviations: FRE – Flesch Reading Ease; FKGL – Flesch–Kincaid Grade Level; SMOG – Simple Measure of Gobbledygook Index.

DISCUSSION

This cross-sectional study compared educational material on cardiac catheterization generated by ChatGPT with the expert-curated resource, UpToDate. Our analysis revealed that ChatGPT produced significantly more concise responses, with lower word count, sentence count, and word-to-sentence ratio (Table 1). Paradoxically, while its content scored at a lower reading grade level (FKGL and SMOG Index), it included a higher percentage of difficult words (Figure 1).

 

Artificial intelligence (AI) tools are being increasingly incorporated into medical education for generating brief summaries that support just-in-time learning. These models allow clinicians to retrieve information more efficiently than traditional resources, serving as complementary tools to established platforms like UpToDate. This is particularly useful for early-career professionals and trainees in time-pressured environments, offering rapid information synthesis and point-of-care decision support [11, 12]. 

Tools like the Flesch Reading Ease (FRE), Flesch-Kincaid Grade Level (FKGL), and the SMOG index are employed to evaluate how easily a text can be understood. Higher FRE values reflect easier readability, while higher FKGL and SMOG scores suggest greater complexity. In medical education, better readability is vital for quick clinical decision-making [13]. In our study, ChatGPT had lower FKGL and SMOG values than UpToDate, indicating simpler sentence structure. However, this was contrasted with a significantly higher difficult word percentage and a non-significant FRE difference, revealing a reliance on more challenging vocabulary despite simpler syntax.

 

Our findings are consistent with prior evaluations of LLMs in medical contexts. For example, Lazris et al. (2024) found that AI-generated summaries were notably shorter than traditional resources [14]. Similarly, Patel et al. (2025) noted that AI can effectively simplify dense medical topics while still delivering accurate and relevant information [15]. This is further supported by Olszewski et al. (2024) showing that AI content can achieve low readability scores (FKGL and SMOG) while presenting advanced academic concepts [16]. This pattern suggests a common characteristic of current AI models: they excel at condensing information, reinforcing their potential as a supplementary tool in medical education. 

 

Conversely, our findings diverge from studies that characterize AI-generated content as overly simplified or lacking in clinical depth. Li et al. (2025) reported that ChatGPT often gave verbose and generic answers [17]. Cao et al. (2025) mentioned that the output omitted important technical nuances [18]. In contrast, our analysis found the information to be highly specific, albeit dense. This difference is likely attributable to the prompt design and intended audience. Furthermore, the reliability of LLMs remains a valid concern due to the potential for factual inaccuracies from outdated training data [11, 19]. These models may lack knowledge of the most recent clinical guidelines or medical discoveries, which can compromise their clinical utility. The variability in AI performance underscores the critical need for standardized assessments of its generated material [13].

 

Limitations: This research faces a few limitations. The study primarily focused on a single AI tool and a specific clinical topic, which may limit the generalizability of the findings to other platforms or areas of medicine. Furthermore, the version of ChatGPT used for this study is static, and its performance may not be representative of newer, more advanced iterations of the model. Furthermore, our study focused exclusively on quantitative textual metrics and did not include a qualitative assessment by subject matter experts on the content's clinical accuracy, evidence-based alignment, or educational utility. Future research should address these gaps through broader studies involving multiple AI tools, diverse clinical topics, and qualitative reviews by experienced clinicians to provide a more holistic understanding of the role of LLMs in medical education.

CONCLUSION

The present study systematically compared educational content produced by ChatGPT and UpToDate on cardiac catheterization targeted toward medical professionals. Statistically significant differences emerged in multiple readability parameters and linguistic complexity. ChatGPT generated shorter, more succinct materials with a lower word and sentence count, a reduced word-to-sentence ratio, and a slightly lower required reading grade level. Despite this, ChatGPT’s responses featured a higher proportion of difficult words, indicating denser vocabulary use. UpToDate, meanwhile, produced more elaborate content with increased detail and marginally greater reading complexity. These findings suggest that while artificial intelligence tools like ChatGPT hold potential as rapid and accessible resources for medical education, their integration should be approached judiciously. Comprehensive evaluations of their clarity, accuracy, and content depth are necessary. Future research is warranted to extend these assessments across a broader range of AI platforms and clinical topics to better define their role and optimize their utility in educational and clinical contexts

REFERENCES

1.      Manda YR, Baradhi KM. Cardiac catheterization risks and complications [Internet]. Treasure Island (FL): StatPearls Publishing; 2025 [updated 2023 Jun 5; cited 2026 Jun 22]. Available from: https://www.ncbi.nlm.nih.gov/books/NBK531461

2.      Naidu SS, Abbott JD, Bagai J, Blankenship J, Garcia S, Iqbal SN, et al. SCAI expert consensus update on best practices in the cardiac catheterization laboratory: this statement was endorsed by the American College of Cardiology (ACC), the American Heart Association (AHA), and the Heart Rhythm Society (HRS) in April 2021. Catheter Cardiovasc Interv. 2021;98(2):255-276. doi:10.1002/ccd.29744.

3.      Ojong M, Weinhold C, Salako A, Andres J. Cardiac catheterization: a review for pharmacists. US Pharm. 2014;39:13-16.

4.      Nishimura RA, Carabello BA. Hemodynamics in the cardiac catheterization laboratory of the 21st century. Circulation. 2012;125(17):2138-2150. doi:10.1161/CIRCULATIONAHA.111.060319.

5.      Tavakol M, Ashraf S, Brener SJ. Risks and complications of coronary angiography: a comprehensive review. Glob J Health Sci. 2012;4(1):65-93. doi:10.5539/gjhs.v4n1p65.

6.      Bangalore S, Barsness GW, Dangas GD, Kern MJ, Rao SV, Shore-Lesserson L, et al. Evidence-based practices in the cardiac catheterization laboratory: a scientific statement from the American Heart Association. Circulation. 2021;144(5):e107-e119. doi:10.1161/CIR.0000000000000996.

7.      Morrissey M, McLaughlin J, O’Connor Y. Reporting of ethical considerations in qualitative research using social media data. J Med Internet Res. 2024;26:e51496. doi:10.2196/51496.

8.      Deoghare S. Virtual research designs and IRB requirements: clarifying what truly needs ethics approval. Minerva Med. 2026. Epub ahead of print. doi:10.23736/S0026-4806.26.09862-9.

9.      IBM Corp. IBM SPSS Statistics for Windows, version 25.0 [computer software]. Armonk (NY): IBM Corp.; 2017.

10.   R Core Team. R: a language and environment for statistical computing, version 4.3.2 [computer software]. Vienna: R Foundation for Statistical Computing; 2023 [cited 2026 Jun 22]. Available from: https://www.r-project.org/

11.   Yu E, Chu X, Zhang W, Meng X, Yang Y, Ji X, et al. Large language models in medicine: applications, challenges, and future directions. Int J Med Sci. 2025;22(11):2792-2801. doi:10.7150/ijms.111780.

12.   Harrington J, Booth RG, Jackson KT. Large language models in nursing education: concept analysis. JMIR Nurs. 2025;8:e77948. doi:10.2196/77948.

13.   DeTemple DE, Meine TC. Comparison of the readability of ChatGPT and Bard in medical communication: a meta-analysis. BMC Med Inform Decis Mak. 2025;25(1):325. doi:10.1186/s12911-025-03035-2.

14.   Lazris D, Schenker Y, Thomas TH. AI-generated content in cancer symptom management: a comparative analysis between ChatGPT and NCCN. J Pain Symptom Manage. 2024;68(4):e303-e311. doi:10.1016/j.jpainsymman.2024.06.019.

15.   Patel EA, Herrmann PT, Fleischer L, Filip P, Joe S, Kshirsagar RS, et al. Comparative analysis of AI-generated study guides in otolaryngology education. Am J Otolaryngol. 2025;46(5):104693. doi:10.1016/j.amjoto.2025.104693.

16.   Olszewski R, Watros K, Mańczak M, Owoc J, Jeziorski K, Brzeziński J. Assessing the response quality and readability of chatbots in cardiovascular health, oncology, and psoriasis: a comparative study. Int J Med Inform. 2024;190:105562. doi:10.1016/j.ijmedinf.2024.105562.

17.   Li Y, Li J, Li M, Yu E, Rhee D, Amith M, et al. VaxBot-HPV: a GPT-based chatbot for answering HPV vaccine-related questions. JAMIA Open. 2025;8(1):ooaf005. doi:10.1093/jamiaopen/ooaf005.

18.   Cao H, Hao C, Zhang T, Zheng X, Gao Z, Wu J, et al. Battle of the artificial intelligence: a comprehensive comparative analysis of DeepSeek and ChatGPT for urinary incontinence-related questions. Front Public Health. 2025;13:1605908. doi:10.3389/fpubh.2025.1605908.

19.  Xu X, Liu S, Zhu L, Long Y, Zeng Y, Lu X, et al. Development and evaluation of a retrieval-augmented large language model framework for enhancing endodontic education. Int J Med Inform. 2025;203:106006. doi:10.1016/j.ijmedinf.2025.106006.

Recommended Articles
Research Article
Clinical Profile and Risk Factors of Upper Gastrointestinal Bleeding in a Tertiary Care Hospital
...
Published: 25/06/2026
Download PDF
Research Article
Relationship Between Syntax Score And Global Longitudinal Strain In Ischemic Heart Disease Evaluation.
...
Published: 24/08/2026
Download PDF
Research Article
Cardiac Autonomic Dysfunction and Subclinical Myocardial Deformation in Patients with Heart Disease: Evidence for Neuro-Mechanical Coupling Despite Preserved Ejection Fraction.
Published: 25/03/2026
Download PDF
Research Article
Burden and Echocardiographic Patterns of Congenital Heart Disease in the High-Altitude Sikkim State of India: A Five-Year Tertiary Care Hospital–Based Study
...
Published: 29/08/2025
Download PDF
Chat on WhatsApp
Copyright © EJCM Publisher. All Rights Reserved.