<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article">
 <front>
  <journal-meta>
   <journal-id journal-id-type="publisher-id">
    jbm
   </journal-id>
   <journal-title-group>
    <journal-title>
     Journal of Biosciences and Medicines
    </journal-title>
   </journal-title-group>
   <issn pub-type="epub">
    2327-5081
   </issn>
   <issn publication-format="print">
    2327-509X
   </issn>
   <publisher>
    <publisher-name>
     Scientific Research Publishing
    </publisher-name>
   </publisher>
  </journal-meta>
  <article-meta>
   <article-id pub-id-type="doi">
    10.4236/jbm.2024.1211021
   </article-id>
   <article-id pub-id-type="publisher-id">
    jbm-137387
   </article-id>
   <article-categories>
    <subj-group subj-group-type="heading">
     <subject>
      Articles
     </subject>
    </subj-group>
    <subj-group subj-group-type="Discipline-v2">
     <subject>
      Biomedical 
     </subject>
     <subject>
       Life Sciences
     </subject>
    </subj-group>
   </article-categories>
   <title-group>
    A Qualitative Assessment of Medical Diagnosis Capabilities of Three Artificial Intelligence Models: ChatGPT-4o, CodyMD, and Dr. Gupta
   </title-group>
   <contrib-group>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Yordanka
      </surname>
      <given-names>
       Eneva
      </given-names>
     </name>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Bora
      </surname>
      <given-names>
       Dogan
      </given-names>
     </name>
    </contrib>
   </contrib-group> 
   <aff id="affnull">
    <addr-line>
     aDepartment of Physics and Biophysics, Faculty of Pharmacy, Medical University of Varna “Prof. Dr. Paraskev Stoyanov”, Varna, Bulgaria
    </addr-line> 
   </aff> 
   <pub-date pub-type="epub">
    <day>
     30
    </day> 
    <month>
     10
    </month>
    <year>
     2024
    </year>
   </pub-date> 
   <volume>
    12
   </volume> 
   <issue>
    11
   </issue>
   <fpage>
    243
   </fpage>
   <lpage>
    254
   </lpage>
   <history>
    <date date-type="received">
     <day>
      22,
     </day>
     <month>
      September
     </month>
     <year>
      2024
     </year>
    </date>
    <date date-type="published">
     <day>
      12,
     </day>
     <month>
      September
     </month>
     <year>
      2024
     </year> 
    </date> 
    <date date-type="accepted">
     <day>
      12,
     </day>
     <month>
      November
     </month>
     <year>
      2024
     </year> 
    </date>
   </history>
   <permissions>
    <copyright-statement>
     © Copyright 2014 by authors and Scientific Research Publishing Inc. 
    </copyright-statement>
    <copyright-year>
     2014
    </copyright-year>
    <license>
     <license-p>
      This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/
     </license-p>
    </license>
   </permissions>
   <abstract>
    <b>Background:</b> Artificial intelligence (AI) has the potential to transform medical diagnostics by enhancing the accuracy and efficiency of diagnostic processes. Its application in clinical practice can greatly support medical professionals by offering improved tools for faster and more precise diagnoses. Understanding AI’s capabilities is essential for its successful integration into medical diagnostics. In this context, evaluating the performance of different AI models in the diagnostic process becomes particularly important. The objective of this study is to qualitatively evaluate the diagnostic performance of three AI models—ChatGPT-4o, CodyMD, and Dr. Gupta—based on patient-reported symptoms. 
    <b>Objectives:</b> The aim of the study is to compare the three AI models in terms of diagnostic accuracy, the level of detail in the provided information, the interaction between the models and patients, and the number of differential diagnoses offered. 
    <b>Results:</b> ChatGPT-4o achieved the highest accuracy, correctly diagnosing 90% of the cases. The model provides basic information and focuses on a single most likely diagnosis. CodyMD and Dr. Gupta achieved 50% accuracy, with CodyMD using an interactive approach and offering differential diagnoses with probability percentages for each. Dr. Gupta provided educational medical information and differential diagnoses without probability estimates. 
    <b>Conclusions:</b> The AI models assessed can assist medical professionals in the diagnostic process, but they require further refinement and optimization. ChatGPT-4o stands out for its high accuracy, though increased patient interaction is needed. CodyMD excels in offering an interactive approach and more detailed responses, but requires improved accuracy. Dr. Gupta provided differential diagnoses, but the information provided is suitable for educational purposes. While these models show potential for clinical use, further research is needed to optimize and validate them in real-world settings.
   </abstract>
   <kwd-group> 
    <kwd>
     Artificial Intelligence
    </kwd> 
    <kwd>
      Medical Diagnostics
    </kwd> 
    <kwd>
      ChatGPT-4o
    </kwd> 
    <kwd>
      CodyMD
    </kwd> 
    <kwd>
      Dr. Gupta
    </kwd> 
    <kwd>
      Qualitative Assessment
    </kwd>
   </kwd-group>
  </article-meta>
 </front>
 <body>
  <sec id="s1">
   <title>1. Introduction</title>
   <p>The rapid development of technologies related to artificial intelligence (AI) offers new opportunities for its implementation in clinical practice. The vast availability of diverse patient data, such as medical images <xref ref-type="bibr" rid="scirp.137387-1">
     [1]
    </xref>, text, and electronic health records <xref ref-type="bibr" rid="scirp.137387-2">
     [2]
    </xref>, combined with numerous studies evaluating the potential of AI in areas such as prevention <xref ref-type="bibr" rid="scirp.137387-3">
     [3]
    </xref>, screening and diagnosis <xref ref-type="bibr" rid="scirp.137387-4">
     [4]
    </xref>, disease progression prediction <xref ref-type="bibr" rid="scirp.137387-5">
     [5]
    </xref>, and clinical decision-making and treatment selection <xref ref-type="bibr" rid="scirp.137387-6">
     [6]
    </xref>, provides the foundation for developing AI solutions that can greatly enhance the diagnostic process.</p>
   <p>There are studies evaluating the diagnostic capabilities of various AI models <xref ref-type="bibr" rid="scirp.137387-7">
     [7]
    </xref>, but important questions arise regarding their reliability, accuracy, and ability to interact with medical professionals and patients. This study focuses on three AI models—ChatGPT-4o, CodyMD, and Dr. Gupta—with the aim of assessing their diagnostic capabilities based on patient-reported symptoms. Key aspects such as diagnostic accuracy, level of interaction, and the provided differential diagnoses are evaluated. Understanding these factors is critically important, as AI has the potential to assist medical professionals in decision-making, but there are also risks of errors, especially in more complex cases. To implement AI effectively in medical practice, research must focus on evaluating the diagnostic accuracy of various AI models. This is crucial for their development, improvement, and training of Natural Language Processing (NLP) models. Such research is vital for advancing medical care and enhancing public health, as it enriches the datasets required for model training and provides key insights for optimizing and adapting algorithms to meet real clinical needs. Therefore, this type of research is particularly relevant and valuable.</p>
   <p>This study aims to qualitatively evaluate the diagnostic capabilities of three different AI models—ChatGPT-4o, CodyMD, and Dr. Gupta—in terms of accuracy, user approach, level of detail, user interaction, and the number of differential diagnoses generated, based solely on patient-reported symptoms. ChatGPT-4o is one of the most widely used AI models, while CodyMD and Dr. Gupta are specifically designed for medical purposes. The qualitative evaluation of their diagnostic methods in this study provides valuable insights for both users and developers.</p>
  </sec><sec id="s2">
   <title>2. Materials and Method</title>
   <p>
    <xref ref-type="bibr" rid="scirp.137387-"></xref>Based on the results of our previous research <xref ref-type="bibr" rid="scirp.137387-7">
     [7]
    </xref> on the diagnostic capabilities of ChatGPT-3.5, where the AI demonstrated over 70% diagnostic accuracy, we decided to extend our research and investigate the diagnostic capabilities of ChatGPT-4o as well as other active medical AIs. The methodological approach (<xref ref-type="fig" rid="fig1">
     Figure 1
    </xref>) follows a structured process that includes selecting the AI models, choosing clinical cases, and analyzing the diagnostic results from the AI models.</p>
   <fig id="fig1" position="float">
    <label>Figure 1</label>
    <caption>
     <title>Figure 1. Flowchart of the research methodology.</title>
    </caption>
    <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/2152828-rId12.jpeg?20241115014645" />
   </fig>
   <sec id="s2_1">
    <title>2.1. AI Model Selection</title>
    <p>The selection of AI models for this study was carefully considered to ensure a diverse and relevant evaluation. ChatGPT-4o was chosen due to its broad usage by the general population for a wide range of informational purposes, including medical diagnostics. In contrast, CodyMD and Dr. Gupta were selected specifically for their focused applications in the medical field, offering specialized diagnostic capabilities that align with the study’s objectives. During the research phase, other AI models like Docus and Doctronic also emerged through a search using the keyword “AI doctors”. However, CodyMD and Dr. Gupta were ultimately selected for their established reputation and strong performance in medical diagnosis. CodyMD is known for its integration with clinical decision-making, and Dr. Gupta has a track record of being used in real-world healthcare scenarios, making them more suitable for this comparison.</p>
    <p>CodyMD was created by David Sanders and Albert DiPiero. It is designed to work alongside real doctors and offers several features, including “Medical Diagnosis”, “Specialists”, “Treat”, “Ask”, “Interpret”, and “Talk”. It is clear from the website that this is a sophisticated product with a wide range of capabilities. In the present study, we utilize the “Medical Diagnosis” feature.</p>
    <p>Dr. Gupta is another AI specifically developed for medical purposes by DL Software Inc. It utilizes advanced natural language processing and machine learning tools to interpret user queries and provide accurate medical information and advice. Dr. Gupta allows users to input data such as age, weight, symptoms, allergies, medications, vitals (temperature, heart rate, respiratory rate, oxygen saturation, waist circumference, hip circumference, systolic blood pressure, and diastolic blood pressure), and lab test results. While this offers a more comprehensive view of the patient’s condition, our goal is to explore diagnostic models based solely on patient-reported symptoms. Therefore, we limit our input to age and symptom data for each patient. This ensures that all three AIs receive the same information and operate under identical conditions for this study.</p>
   </sec>
   <sec id="s2_2">
    <title>2.2. Literature Review</title>
    <p>The literature review was conducted through an extensive search of databases including PubMed, Google Scholar, Web of Science, and Scopus (<xref ref-type="fig" rid="fig1">
      Figure 1
     </xref>). Our focus was on articles published in English that relate to medical diagnoses made by AI. For our search, we used a combination of keywords: “medical diagnosis” and “artificial intelligence”. However, due to the large number of articles, we narrowed our focus to “ChatGPT and medical diagnosis”. A literature search using the keywords “Dr. Gupta and medical diagnosis” and “CodyMD and medical diagnosis” yielded no results. The absence of studies involving both models makes this study unique.</p>
   </sec>
   <sec id="s2_3">
    <title>2.3. Selection of Clinical Cases to Evaluate the Diagnostic Capabilities of Models</title>
    <p>For an effective evaluation of the AI models’ diagnostic performance, it was essential to select clinical cases representing a diverse range of medical conditions with varying levels in complexity and severity. This approach ensures that the AI models are challenged across various diagnostic scenarios, mimicking the daily challenges faced by a general practitioner and allowing for a better assessment of their diagnostic capabilities. For this purpose, we conducted an extensive search in PubMed (<xref ref-type="fig" rid="fig1">
      Figure 1
     </xref>) and selected cases that cover a spectrum of diseases, ranging from rare to common, with varying levels of complexity and clearly defined diagnoses confirmed by specialists. The case selection criteria were as follows:</p>
    <p>Based on these criteria, ten clinical cases were selected (<xref ref-type="table" rid="table1">
      Table 1
     </xref>). Each case was assigned a diagnostic difficulty rating on a scale of one to five stars, determined by the following factors:</p>
    <p>For each case, a clinical vignette was created that included the patient’s age and disease symptoms (<xref ref-type="table" rid="table1">
      Table 1
     </xref>). These vignettes were presented to the three AI models in separate, isolated chat sessions, with each model being asked, “What is the most likely diagnosis?” This process ensured that none of the models were influenced by previous interactions. Each model received the same patient-reported symptoms and medical history presented consistently, and without any additional cues or guidance, allowing independent analysis and response to the clinical data. The diagnoses generated by the models are presented in <xref ref-type="table" rid="table2">
      Table 2
     </xref>.</p>
    <table-wrap id="table1">
     <label>
      <xref ref-type="table" rid="table1">
       Table 1
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.137387-"></xref>Table 1. Reported clinical symptoms, corresponding diagnoses, and diagnostic difficulties.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="12.80%"><p style="text-align:center">Case Report</p></td> 
       <td class="custom-bottom-td acenter" width="49.84%"><p style="text-align:center">Complaints of the patient</p></td> 
       <td class="custom-bottom-td acenter" width="18.50%"><p style="text-align:center">Correct diagnosis</p></td> 
       <td class="custom-bottom-td acenter" width="18.86%"><p style="text-align:center">Diagnostic difficulty</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="12.80%"><p style="text-align:center">1</p></td> 
       <td class="custom-top-td aleft" width="49.84%"><p style="text-align:left">An 18-year-old female patient presents the following symptoms: </p><p style="text-align:left">- Partially blocked left nostril along with bilateral nasal itchy feeling</p><p style="text-align:left">- Sneezing for up to 1h and 80 to 100 sneezes every day usually in the morning time</p><p style="text-align:left">- Watery discharge from the nose (rhinorrhea)</p><p style="text-align:left">- Heaviness in the head region</p><p style="text-align:left">- Loss of concentration</p><p style="text-align:left">- Weakness <xref ref-type="bibr" rid="scirp.137387-8">
          [8]
         </xref></p></td> 
       <td class="custom-top-td acenter" width="18.50%"><p style="text-align:center">Allergic Rhinitis</p></td> 
       <td class="custom-top-td acenter" width="18.86%"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="12.80%"><p style="text-align:center">2</p></td> 
       <td class="aleft" width="49.84%"><p style="text-align:left">A 27-year-old woman has progressively deteriorating abdominal symptoms over the previous 5 years. Complaints include:</p><p style="text-align:left">- Diarrhoea alternating with constipation, and on occasions, episodes of faecal incontinence</p><p style="text-align:left">- Colicky abdominal pain on a daily basis accompanied by abdominal distension</p><p style="text-align:left">- Lethargy</p><p style="text-align:left">- Low back pain</p><p style="text-align:left">- Nausea</p><p style="text-align:left">- Bladder symptoms consistent with a diagnosis of irritable bladder <xref ref-type="bibr" rid="scirp.137387-9">
          [9]
         </xref></p></td> 
       <td class="acenter" width="18.50%"><p style="text-align:center">Irritable Bowel Syndrome</p></td> 
       <td class="acenter" width="18.86%"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="12.80%"><p style="text-align:center">3</p></td> 
       <td class="aleft" width="49.84%"><p style="text-align:left">A 7-year-old boy presents the following symptoms:</p><p style="text-align:left">- Intermittent fevers</p><p style="text-align:left">- Lower quadrant abdominal pain</p><p style="text-align:left">- Vomiting, without bilious and bloody emesis <xref ref-type="bibr" rid="scirp.137387-10">
          [10]
         </xref></p></td> 
       <td class="acenter" width="18.50%"><p style="text-align:center">Appendicitis</p></td> 
       <td class="acenter" width="18.86%"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="12.80%"><p style="text-align:center">4</p></td> 
       <td class="aleft" width="49.84%"><p style="text-align:left">A 36-year-old woman presents the following symptoms: - Cyclic pain on the C-section scar</p><p style="text-align:left">- Moderate to severe dysmenorrhoea and dyspareunia</p><p style="text-align:left">- Painful, palpable, small firm mass of approximately 3 cm in the lower abdomen wall, at the site of the caesarean section scar <xref ref-type="bibr" rid="scirp.137387-11">
          [11]
         </xref></p></td> 
       <td class="acenter" width="18.50%"><p style="text-align:center">(Abdominal Wall) Endometriosis</p></td> 
       <td class="acenter" width="18.86%"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="12.80%"><p style="text-align:center">5</p></td> 
       <td class="aleft" width="49.84%"><p style="text-align:left">An 18-year-old boy presents the following symptoms:</p><p style="text-align:left">- Acute onset of fever</p><p style="text-align:left">- Rhinitis</p><p style="text-align:left">- Myalgia</p><p style="text-align:left">- Headache</p><p style="text-align:left">- Decreased taste and smell sensation <xref ref-type="bibr" rid="scirp.137387-12">
          [12]
         </xref></p></td> 
       <td class="acenter" width="18.50%"><p style="text-align:center">COVID-19</p></td> 
       <td class="acenter" width="18.86%"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="12.80%"><p style="text-align:center">6</p></td> 
       <td class="aleft" width="49.84%"><p style="text-align:left">A 41-year-old woman has hypertension, hypothyroidism, and asthma. She presents the following symptoms:</p><p style="text-align:left">–1-month history of fever associated with chills and rigors</p><p style="text-align:left">–Pleuritic chest pain</p><p style="text-align:left">–Pain in the small joints of the hand</p><p style="text-align:left">–Cold in the extremities</p><p style="text-align:left">–Photosensitivity</p><p style="text-align:left">–1-year history of progressive fatigue, arthralgia, 20 kg weight loss, and intermittent low- and high-grade fever <xref ref-type="bibr" rid="scirp.137387-13">
          [13]
         </xref></p></td> 
       <td class="acenter" width="18.50%"><p style="text-align:center">Systemic Lupus Erythematosus</p></td> 
       <td class="acenter" width="18.86%"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="12.80%"><p style="text-align:center">7</p></td> 
       <td class="aleft" width="49.84%"><p style="text-align:left">A 35-year-old man has a medical history of hyperglycemia, hyperhidrosis, and high blood pressure. He presents the following symptoms:</p><p style="text-align:left">–Fatigue and cough</p><p style="text-align:left">–Dyspnea, accompanied by chest tightness, and inability of lying supine at night</p><p style="text-align:left">–Bout of cold followed by general malaise <xref ref-type="bibr" rid="scirp.137387-14">
          [14]
         </xref></p></td> 
       <td class="acenter" width="18.50%"><p style="text-align:center">Heart Failure</p></td> 
       <td class="acenter" width="18.86%"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="12.80%"><p style="text-align:center">8</p></td> 
       <td class="aleft" width="49.84%"><p style="text-align:left">A 6-year-old girl, who is a known asthmatic, presented the following symptoms:</p><p style="text-align:left">–Generalized bruises</p><p style="text-align:left">–Fever</p><p style="text-align:left">–Eczematous rashes</p><p style="text-align:left">–Six or seven episodes of loose stools per day for 3 months accompanied by loss of appetite <xref ref-type="bibr" rid="scirp.137387-15">
          [15]
         </xref></p></td> 
       <td class="acenter" width="18.50%"><p style="text-align:center">Celiac Disease</p></td> 
       <td class="acenter" width="18.86%"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="12.80%"><p style="text-align:center">9</p></td> 
       <td class="aleft" width="49.84%"><p style="text-align:left">A 30-year-old white male presents the following symptoms:</p><p style="text-align:left">–Decrease of visual acuity</p><p style="text-align:left">–Intermittent diplopia</p><p style="text-align:left">–Photophobia in both eyes</p><p style="text-align:left">–Paresthesia of left hand <xref ref-type="bibr" rid="scirp.137387-16">
          [16]
         </xref></p></td> 
       <td class="acenter" width="18.50%"><p style="text-align:center">Multiple Sclerosis</p></td> 
       <td class="acenter" width="18.86%"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="12.80%"><p style="text-align:center">10</p></td> 
       <td class="aleft" width="49.84%"><p style="text-align:left">A woman in her late 70s presents the following symptoms:</p><p style="text-align:left">–Significant short-term memory impairment</p><p style="text-align:left">–Episodes of confusion</p><p style="text-align:left">–Difficulty with language skills</p><p style="text-align:left">–Reclusive and disengaged from her previous social networks</p><p style="text-align:left">–Disorientation during seasonal changes that leads to periods of wandering and becoming lost <xref ref-type="bibr" rid="scirp.137387-17">
          [17]
         </xref></p></td> 
       <td class="acenter" width="18.50%"><p style="text-align:center">Alzheimer’s disease</p></td> 
       <td class="acenter" width="18.86%"><p style="text-align:center"></p></td> 
      </tr> 
     </table>
    </table-wrap>
    <table-wrap id="table2">
     <label>
      <xref ref-type="table" rid="table2">
       Table 2
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.137387-"></xref>Table 2. Diagnoses offered by the three AI models.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="acenter" width="13.73%"><p style="text-align:center">Case Report</p></td> 
       <td class="acenter" width="16.15%"><p style="text-align:center">ChatGPT-4o</p></td> 
       <td class="acenter" width="30.62%"><p style="text-align:center">CodyMD</p></td> 
       <td class="acenter" width="21.40%"><p style="text-align:center">Dr. Gupta</p></td> 
       <td class="acenter" width="18.10%"><p style="text-align:center">Correct Diagnosis</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="13.73%"><p style="text-align:center">1</p></td> 
       <td class="custom-top-td acenter" width="16.15%"><p style="text-align:center">Allergic Rhinitis</p></td> 
       <td class="custom-top-td aleft" width="30.62%"><p style="text-align:left">1. Allergic Rhinitis (60% probability);</p><p style="text-align:left">2. Vasomotor Rhinitis (25% probability);</p><p style="text-align:left">3. Common Cold (15% probability)</p></td> 
       <td class="custom-top-td aleft" width="21.40%"><p style="text-align:left">Allergic Rhinitis</p></td> 
       <td class="custom-top-td acenter" width="18.10%"><p style="text-align:center">Allergic Rhinitis</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.73%"><p style="text-align:center">2</p></td> 
       <td class="acenter" width="16.15%"><p style="text-align:center">Irritable Bowel Syndrome</p></td> 
       <td class="aleft" width="30.62%"><p style="text-align:left">1. Irritable Bowel Syndrome (IBS) (60% probability);</p><p style="text-align:left">2. Inflammatory Bowel Disease (IBD) (30% probability);</p><p style="text-align:left">3. Functional Dyspepsia (10% probability)</p></td> 
       <td class="aleft" width="21.40%"><p style="text-align:left">1. Irritable Bowel Syndrome (IBS);</p><p style="text-align:left">2. Inflammatory Bowel Disease;</p><p style="text-align:left">3. Gastrointestinal infection</p></td> 
       <td class="acenter" width="18.10%"><p style="text-align:center">Irritable Bowel Syndrome</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.73%"><p style="text-align:center">3</p></td> 
       <td class="acenter" width="16.15%"><p style="text-align:center">Appendicitis</p></td> 
       <td class="aleft" width="30.62%"><p style="text-align:left">1. Gastroenteritis (40% probability);</p><p style="text-align:left">2. Urinary tract infection (25% probability);</p><p style="text-align:left">3. Appendicitis (20% probability)</p></td> 
       <td class="aleft" width="21.40%"><p style="text-align:left">1. Viral or bacterial gastroenteritis;</p><p style="text-align:left">2. Appendicitis;</p><p style="text-align:left">3. Urinary tract infection</p></td> 
       <td class="acenter" width="18.10%"><p style="text-align:center">Appendicitis</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.73%"><p style="text-align:center">4</p></td> 
       <td class="acenter" width="16.15%"><p style="text-align:center">Cesarean Scar Endometriosis</p></td> 
       <td class="aleft" width="30.62%"><p style="text-align:left">1. Cesarean Scar Endometriosis (70% probability);</p><p style="text-align:left">2. Scar Tissue Adhesion (20% probability);</p><p style="text-align:left">3. Incisional Hernia (10% probability);</p></td> 
       <td class="aleft" width="21.40%"><p style="text-align:left">Incisional Hernia</p></td> 
       <td class="acenter" width="18.10%"><p style="text-align:center">(Abdominal Wall) Endometriosis</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.73%"><p style="text-align:center">5</p></td> 
       <td class="acenter" width="16.15%"><p style="text-align:center">COVID-19</p></td> 
       <td class="aleft" width="30.62%"><p style="text-align:left">1. Influenza (40% probability);</p><p style="text-align:left">2. Common Cold (30% probability);</p><p style="text-align:left">3. COVID-19 (30% probability)</p></td> 
       <td class="aleft" width="21.40%"><p style="text-align:left">1. Common cold or the flu;</p><p style="text-align:left">2. Sinus infection</p></td> 
       <td class="acenter" width="18.10%"><p style="text-align:center">COVID-19</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.73%"><p style="text-align:center">6</p></td> 
       <td class="acenter" width="16.15%"><p style="text-align:center">Systemic Lupus Erythematosus (SLE)</p></td> 
       <td class="aleft" width="30.62%"><p style="text-align:left">1. Rheumatoid arthritis (40% probability);</p><p style="text-align:left">2. Systemic lupus erythematosus (30% probability);</p><p style="text-align:left">3. Tuberculosis (20% probability)</p></td> 
       <td class="aleft" width="21.40%"><p style="text-align:left">1. Systemic Lupus Erythematosus (SLE);</p><p style="text-align:left">2. Tuberculosis (TB);</p><p style="text-align:left">3. Rheumatoid Arthritis or other autoimmune conditions</p></td> 
       <td class="acenter" width="18.10%"><p style="text-align:center">Systemic Lupus Erythematosus</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.73%"><p style="text-align:center">7</p></td> 
       <td class="acenter" width="16.15%"><p style="text-align:center">Congestive Heart Failure (CHF)</p></td> 
       <td class="aleft" width="30.62%"><p style="text-align:left">1. Chronic obstructive pulmonary disease (COPD) (40% probability);</p><p style="text-align:left">2. Asthma (30% probability);</p><p style="text-align:left">3. Gastroesophageal reflux disease (GERD) (20% probability)</p></td> 
       <td class="aleft" width="21.40%"><p style="text-align:left">1. Respiratory infection (bronchitis or pneumonia);</p><p style="text-align:left">2. Allergies, asthma, anxiety</p></td> 
       <td class="acenter" width="18.10%"><p style="text-align:center">Heart Failure</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.73%"><p style="text-align:center">8</p></td> 
       <td class="acenter" width="16.15%"><p style="text-align:center">Wiskott-Aldrich syndrome (WAS)</p></td> 
       <td class="aleft" width="30.62%"><p style="text-align:left">1. Inflammatory Bowel Disease (30% probability);</p><p style="text-align:left">2. Eczema Herpeticum (25% probability);</p><p style="text-align:left">3. Food Allergy (20% probability)</p></td> 
       <td class="aleft" width="21.40%"><p style="text-align:left">1. Viral Gastroenteritis;</p><p style="text-align:left">2. Inflammatory Bowel Disease;</p><p style="text-align:left">3. Blood disorder or a bleeding disorder</p></td> 
       <td class="acenter" width="18.10%"><p style="text-align:center">Celiac Disease</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.73%"><p style="text-align:center">9</p></td> 
       <td class="acenter" width="16.15%"><p style="text-align:center">Multiple </p><p style="text-align:center">Sclerosis (MP)</p></td> 
       <td class="aleft" width="30.62%"><p style="text-align:left">1. Multiple Sclerosis (30% probability);</p><p style="text-align:left">2. Migraine (25% probability);</p><p style="text-align:left">3. Optic Neuritis (20% probability)</p></td> 
       <td class="aleft" width="21.40%"><p style="text-align:left">1. Multiple Sclerosis (MS);</p><p style="text-align:left">2. Asthma;</p><p style="text-align:left">3. Uveitis or optic neuritis</p></td> 
       <td class="acenter" width="18.10%"><p style="text-align:center">Multiple Sclerosis</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.73%"><p style="text-align:center">10</p></td> 
       <td class="acenter" width="16.15%"><p style="text-align:center">Alzheimer’s disease</p></td> 
       <td class="aleft" width="30.62%"><p style="text-align:left">1. Alzheimer’s disease (40% probability);</p><p style="text-align:left">2. Vascular dementia (30% probability);</p><p style="text-align:left">3. Mild cognitive impairment (20% probability)</p></td> 
       <td class="aleft" width="21.40%"><p style="text-align:left">Alzheimer’s disease or another form of dementia</p></td> 
       <td class="acenter" width="18.10%"><p style="text-align:center">Alzheimer’s disease</p></td> 
      </tr> 
     </table>
    </table-wrap>
   </sec>
   <sec id="s2_4">
    <title>2.4. Selection of Criteria for Evaluating the Diagnostic Capabilities of Models</title>
    <p>The diagnostic results provided by each AI model were evaluated based on the following qualitative criteria:</p>
   </sec>
  </sec><sec id="s3">
   <title>3. Results</title>
   <p>The AI diagnoses (<xref ref-type="table" rid="table2">
     Table 2
    </xref>) were compared with those published in the literature and evaluated based on the criteria outlined above. The results are presented in <xref ref-type="table" rid="table3">
     Table 3
    </xref>. The table shows that out of ten clinical vignettes, ChatGPT made correct and accurate diagnoses in nine of them. In contrast, CodyMD and Dr. Gupta provided differential diagnoses, with the correct diagnosis being the first one about half the time. This can be viewed as a positive feature of these models, as they enable physicians to explore multiple potential diagnoses instead of being restricted to just one. By promoting a broader differential diagnosis, these models help prevent premature conclusions and foster a more thorough assessment, ultimately leading to more accurate and informed decision-making in patient care.</p>
   <p>For case 3, CodyMD assigned a probability of 20% to the correct diagnosis, and for cases 5 and 6, –30%. Dr. Gupta provided potential differential diagnoses but did not assign any probability percentages. In case 3, the correct diagnosis was listed second, while for cases 4, 5, 7, and 8, Dr. Gupta did not provide a correct diagnosis. These two AIs are designed to provide medical information, but symptoms alone may not be sufficient for accurate diagnoses due to the numerous diseases with overlapping symptoms. As a result, they may require additional information about the patient’s physiological state, including blood tests, imaging studies (e.g., X-rays, MRI, or CT scans), and other diagnostic tools. These tools can be electrocardiograms (ECGs) to assess heart function, pulmonary function tests to evaluate lung capacity, biopsies for tissue analysis, genetic testing to identify hereditary diseases and urinalysis to detect abnormalities in kidney function.</p>
   <p>ChatGPT-4o provides only one differential diagnosis based on the patient’s symptoms, using an informative approach without engaging in direct dialogue with the patient. It communicates in accessible language for the general public and provides a brief rationale for its response. While it suggests possible conditions, it typically emphasizes the need for further investigations, tests, and a physical examination by a medical professional. Although it mentions potential causes and provides context for the symptoms, it does not offer more than one differential diagnosis compared to other AI. It often focuses on a single condition, which can be limiting given the variety of presenting symptoms. The level of detail in its responses is moderate, providing plausible explanations in accessible language. However, it lacks personalized interaction, making it less engaging for the patient.</p>
   <p>CodyMD, on the other hand, provides three differential diagnoses for all cases and assigns a percentage probability to each potential diagnosis. This AI uses an interactive approach, creating a conversational format that resembles a doctor-patient dialogue. It asks detailed questions to gather additional information about symptoms and pain levels, aiming to refine the diagnosis. CodyMD provides a list of three potential diagnoses and offers a detailed treatment plan for one of them. The level of detail is high, including comprehensive explanations, self-care tips, and potential lifestyle changes. CodyMD is very interactive, friendly, and polite, frequently using encouraging phrases such as “Thanks for the confirmation!” and “Thanks for sharing with me!” It maintains constant communication with the patient by asking questions like “What should I call you?” and “Would you like to see a treatment plan?” This fosters a collaborative atmosphere, and strengthens the physician-patient connection, and mimics a sense of empathy for the patient.</p>
   <p>Dr. Gupta, like CodyMD, provides differential diagnoses in most cases (see <xref ref-type="table" rid="table2">
     Table 2
    </xref>) but does not state the percent likelihood for each. It begins by stating the most likely diagnosis and offers a brief explanation of the condition, along with potential causes of the symptoms. Dr. Gupta employs an educational approach, making it suitable for medical students during their training. The AI explains the disease and the causes of specific symptoms, emphasizing that a formal diagnosis can only be made by a medical professional through evaluation. It suggests appropriate diagnostic steps (such as physical exams and imaging) and recommends treatment options (including pain control, hormone therapy, and surgery). While it provides clear and precise explanations, it does not delve deeply into the specifics of the disease and lacks the interactivity and personalization found in CodyMD’s responses.</p>
   <table-wrap id="table3">
    <label>
     <xref ref-type="table" rid="table3">
      Table 3
     </xref></label>
    <caption>
     <title>
      <xref ref-type="bibr" rid="scirp.137387-"></xref>Table 3. Comparative analysis of the three AI models.</title>
    </caption>
    <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
     <tr> 
      <td class="custom-bottom-td custom-top-td acenter" width="24.48%"><p style="text-align:center">Criteria</p></td> 
      <td class="custom-bottom-td custom-top-td acenter" width="24.07%"><p style="text-align:center">ChatGPT-4o</p></td> 
      <td class="custom-bottom-td custom-top-td acenter" width="26.76%"><p style="text-align:center">CodyMD</p></td> 
      <td class="custom-bottom-td custom-top-td acenter" width="24.69%"><p style="text-align:center">Dr. Gupta</p></td> 
     </tr> 
     <tr> 
      <td class="custom-top-td acenter" width="24.48%"><p style="text-align:center">Accuracy</p></td> 
      <td class="custom-top-td acenter" width="24.07%"><p style="text-align:center">90%</p></td> 
      <td class="custom-top-td acenter" width="26.76%"><p style="text-align:center">50%</p></td> 
      <td class="custom-top-td acenter" width="24.69%"><p style="text-align:center">50%</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.48%"><p style="text-align:center">Approach</p></td> 
      <td class="acenter" width="24.07%"><p style="text-align:center">informative</p></td> 
      <td class="acenter" width="26.76%"><p style="text-align:center">interactive</p></td> 
      <td class="acenter" width="24.69%"><p style="text-align:center">educational</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.48%"><p style="text-align:center">Level of Detail</p></td> 
      <td class="acenter" width="24.07%"><p style="text-align:center">moderate</p></td> 
      <td class="acenter" width="26.76%"><p style="text-align:center">high</p></td> 
      <td class="acenter" width="24.69%"><p style="text-align:center">high</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.48%"><p style="text-align:center">User Interaction</p></td> 
      <td class="acenter" width="24.07%"><p style="text-align:center">Does not engage in dialog</p></td> 
      <td class="aleft" width="26.76%"><p style="text-align:left">Engages in dialog and seeks more information from the patient</p></td> 
      <td class="acenter" width="24.69%"><p style="text-align:center">Does not engage in dialog</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="24.48%"><p style="text-align:center">Number of possible diagnoses provided</p></td> 
      <td class="acenter" width="24.07%"><p style="text-align:center">1</p></td> 
      <td class="acenter" width="26.76%"><p style="text-align:center">3</p></td> 
      <td class="acenter" width="24.69%"><p style="text-align:center">1, 2 or 3</p></td> 
     </tr> 
    </table>
   </table-wrap>
  </sec><sec id="s4">
   <title>4. Conclusion</title>
   <p>
    <xref ref-type="bibr" rid="scirp.137387-"></xref>The results of the present study indicate that the three AIs investigated have the potential to assist healthcare professionals in diagnosing complex cases with multiple symptoms. They can also be used as virtual assistants for initial patient consultations, referring patients to appropriate tests or specialists based on their reported symptoms. In both cases, however, optimization and refinement are needed to improve their accuracy and reliability in clinical settings. The interactivity demonstrated by Cody is particularly useful for acquiring more accurate and complete information about the disease and its history. There is a need to advance more personalized AI decision-making based on patient-supplied symptoms and the analysis of genetic and clinical data to enable more accurate diagnoses and individualized treatment plans. Future studies will need to focus not only on refining the algorithms, but also on integrating these systems into real clinical scenarios to ensure their effectiveness and practical applicability.</p>
  </sec>
 </body><back>
  <ref-list>
   <title>References</title>
   <ref id="scirp.137387-ref1">
    <label>1</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Asha, P., Srivani, P., iqbaldoewes, R., Al Ayub Ahmed, A., Kolhe, A. and Nomani, M.Z.M. (2022) Artificial Intelligence in Medical Imaging: An Analysis of Innovative Technique and Its Future Promise. Materials Today: Proceedings, 56, 2236-2239. &gt;https://doi.org/10.1016/j.matpr.2021.11.558 
    </mixed-citation>
   </ref>
   <ref id="scirp.137387-ref2">
    <label>2</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Yang, X., Chen, A., PourNejatian, N., Shin, H.C., Smith, K.E., Parisien, C., et al. (2022) A Large Language Model for Electronic Health Records. npj Digital Medicine, 5, Article No. 194. &gt;https://doi.org/10.1038/s41746-022-00742-2 
    </mixed-citation>
   </ref>
   <ref id="scirp.137387-ref3">
    <label>3</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Patel, V. and Shah, M. (2021) A Comprehensive Study on Artificial Intelligence and Machine Learning in Drug Discovery and Drug Development. Intelligent Medicine, 2, 134-140.
    </mixed-citation>
   </ref>
   <ref id="scirp.137387-ref4">
    <label>4</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Nakamura, T. and Sasano, T. (2022) Artificial Intelligence and Cardiology: Current Status and Perspective. Journal of Cardiology, 79, 326-333. &gt;https://doi.org/10.1016/j.jjcc.2021.11.017 
    </mixed-citation>
   </ref>
   <ref id="scirp.137387-ref5">
    <label>5</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Muthalaly, R.G. and Evans, R.M. (2020) Applications of Machine Learning in Cardiac Electrophysiology. Arrhythmia&amp;Electrophysiology Review, 9, 71-77. &gt;https://doi.org/10.15420/aer.2019.19 
    </mixed-citation>
   </ref>
   <ref id="scirp.137387-ref6">
    <label>6</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Van den Eynde, J., Lachmann, M., Laugwitz, K., Manlhiot, C. and Kutty, S. (2023) Successfully Implemented Artificial Intelligence and Machine Learning Applications in Cardiology: State-of-the-Art Review. Trends in Cardiovascular Medicine, 33, 265-271. &gt;https://doi.org/10.1016/j.tcm.2022.01.010 
    </mixed-citation>
   </ref>
   <ref id="scirp.137387-ref7">
    <label>7</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Eneva, Y. and Dogan, B. (2024) Evaluation of Medical Diagnosis Capabilities of Three Artificial Intelligence Models—ChatGPT-3.5, Google Gemini, Microsoft Copilot. International Journal of Advanced Natural Sciences and Engineering Researches, 8, 102-108.
    </mixed-citation>
   </ref>
   <ref id="scirp.137387-ref8">
    <label>8</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Sharma, R. and Bhat, P. (2023) Management of Allergic Rhinitis with Rajanyadi Churna and Guduchi Kwatha—A Case Report. Journal of Ayurveda and Integrative Medicine, 14, Article 100740. &gt;https://doi.org/10.1016/j.jaim.2023.100740 
    </mixed-citation>
   </ref>
   <ref id="scirp.137387-ref9">
    <label>9</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Pearson, J.S., Niven, R.M., Meng, J., Atarodi, S. and Whorwell, P.J. (2015) Immunoglobulin E in Irritable Bowel Syndrome: Another Target for Treatment? A Case Report and Literature Review. Therapeutic Advances in Gastroenterology, 8, 270-277. &gt;https://doi.org/10.1177/1756283x15588875 
    </mixed-citation>
   </ref>
   <ref id="scirp.137387-ref10">
    <label>10</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Wang, Z., Ye, J., Wang, Y. and Liu, Y. (2019) Diagnostic Accuracy of Pediatric Atypical Appendicitis. Medicine, 98, e15006. &gt;https://doi.org/10.1097/md.0000000000015006 
    </mixed-citation>
   </ref>
   <ref id="scirp.137387-ref11">
    <label>11</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Doroftei, B., Armeanu, T., Maftei, R., Ilie, O., Dabuleanu, A. and Condac, C. (2020) Abdominal Wall Endometriosis: Two Case Reports and Literature Review. Medicina, 56, Article 727. &gt;https://doi.org/10.3390/medicina56120727 
    </mixed-citation>
   </ref>
   <ref id="scirp.137387-ref12">
    <label>12</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Devassikutty, F.M., Jain, A., Edavazhippurath, A., Joseph, M.C., Peedikayil, M.M.T., Scaria, V., et al. (2021) X-Linked Agammaglobulinemia and COVID-19: Two Case Reports and Review of Literature. Pediatric Allergy, Immunology, and Pulmonology, 34, 115-118. &gt;https://doi.org/10.1089/ped.2021.0002 
    </mixed-citation>
   </ref>
   <ref id="scirp.137387-ref13">
    <label>13</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Al-Nokhatha, S.A., Khogali, H.I., Al Shehhi, M.A. and Jassim, I.T. (2019) Myocarditis as a Lupus Challenge: Two Case Reports. Journal of Medical Case Reports, 13, Article No. 343. &gt;https://doi.org/10.1186/s13256-019-2242-1 
    </mixed-citation>
   </ref>
   <ref id="scirp.137387-ref14">
    <label>14</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Zhang, J., Cao, L., Yan, L., Jin, C. and Zhang, D. (2022) A Young Patient with Heart Failure Was Diagnosed with Extra-Adrenal Paraganglioma: A Case Report. BMC Cardiovascular Disorders, 22, Article No. 574. &gt;https://doi.org/10.1186/s12872-022-03026-5 
    </mixed-citation>
   </ref>
   <ref id="scirp.137387-ref15">
    <label>15</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Irfan, O., Mahmood, S., Nand, H. and Billoo, G. (2018) Celiac Disease Associated with Aplastic Anemia in a 6-Year-Old Girl: A Case Report and Review of the Literature. Journal of Medical Case Reports, 12, Article No. 16. &gt;https://doi.org/10.1186/s13256-017-1527-5 
    </mixed-citation>
   </ref>
   <ref id="scirp.137387-ref16">
    <label>16</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Costin, D., Pînzaru, G.M., Pătraşcu, A.M., Moţoc, A. and Moraru, A.D. (2018) Multiple Sclerosis with Ophthalmologic Onset—Case Report. Romanian Journal of Ophthalmology, 62, 78-82. &gt;https://doi.org/10.22336/rjo.2018.11 
    </mixed-citation>
   </ref>
   <ref id="scirp.137387-ref17">
    <label>17</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Quail, Z., Carter, M.M., Wei, A. and Li, X. (2020) Management of Cognitive Decline in Alzheimer’s Disease Using a Non-Pharmacological Intervention Program. Medicine, 99, e20128. &gt;https://doi.org/10.1097/md.0000000000020128
    </mixed-citation>
   </ref>
  </ref-list>
 </back>
</article>