<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article">
 <front>
  <journal-meta>
   <journal-id journal-id-type="publisher-id">
    jdaip
   </journal-id>
   <journal-title-group>
    <journal-title>
     Journal of Data Analysis and Information Processing
    </journal-title>
   </journal-title-group>
   <issn pub-type="epub">
    2327-7211
   </issn>
   <issn publication-format="print">
    2327-7203
   </issn>
   <publisher>
    <publisher-name>
     Scientific Research Publishing
    </publisher-name>
   </publisher>
  </journal-meta>
  <article-meta>
   <article-id pub-id-type="doi">
    10.4236/jdaip.2024.124029
   </article-id>
   <article-id pub-id-type="publisher-id">
    jdaip-136659
   </article-id>
   <article-categories>
    <subj-group subj-group-type="heading">
     <subject>
      Articles
     </subject>
    </subj-group>
    <subj-group subj-group-type="Discipline-v2">
     <subject>
      Computer Science 
     </subject>
     <subject>
       Communications, Physics 
     </subject>
     <subject>
       Mathematics
     </subject>
    </subj-group>
   </article-categories>
   <title-group>
    Optimizing Healthcare Big Data Processing with Containerized PySpark and Parallel Computing: A Study on ETL Pipeline Efficiency
   </title-group>
   <contrib-group>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Ehsan
      </surname>
      <given-names>
       Soltanmohammadi
      </given-names>
     </name>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Neset
      </surname>
      <given-names>
       Hikmet
      </given-names>
     </name>
    </contrib>
   </contrib-group> 
   <aff id="affnull">
    <addr-line>
     aDepartment of Integrated Information Technology, University of South Carolina, Columbia, United States
    </addr-line> 
   </aff> 
   <pub-date pub-type="epub">
    <day>
     06
    </day> 
    <month>
     09
    </month>
    <year>
     2024
    </year>
   </pub-date> 
   <volume>
    12
   </volume> 
   <issue>
    04
   </issue>
   <fpage>
    544
   </fpage>
   <lpage>
    565
   </lpage>
   <history>
    <date date-type="received">
     <day>
      9,
     </day>
     <month>
      September
     </month>
     <year>
      2024
     </year>
    </date>
    <date date-type="published">
     <day>
      15,
     </day>
     <month>
      September
     </month>
     <year>
      2024
     </year> 
    </date> 
    <date date-type="accepted">
     <day>
      15,
     </day>
     <month>
      October
     </month>
     <year>
      2024
     </year> 
    </date>
   </history>
   <permissions>
    <copyright-statement>
     © Copyright 2014 by authors and Scientific Research Publishing Inc. 
    </copyright-statement>
    <copyright-year>
     2014
    </copyright-year>
    <license>
     <license-p>
      This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/
     </license-p>
    </license>
   </permissions>
   <abstract>
    In this study, we delve into the realm of efficient Big Data Engineering and Extract, Transform, Load (ETL) processes within the healthcare sector, leveraging the robust foundation provided by the MIMIC-III Clinical Database. Our investigation entails a comprehensive exploration of various methodologies aimed at enhancing the efficiency of ETL processes, with a primary emphasis on optimizing time and resource utilization. Through meticulous experimentation utilizing a representative dataset, we shed light on the advantages associated with the incorporation of PySpark and Docker containerized applications. Our research illuminates significant advancements in time efficiency, process streamlining, and resource optimization attained through the utilization of PySpark for distributed computing within Big Data Engineering workflows. Additionally, we underscore the strategic integration of Docker containers, delineating their pivotal role in augmenting scalability and reproducibility within the ETL pipeline. This paper encapsulates the pivotal insights gleaned from our experimental journey, accentuating the practical implications and benefits entailed in the adoption of PySpark and Docker. By streamlining Big Data Engineering and ETL processes in the context of clinical big data, our study contributes to the ongoing discourse on optimizing data processing efficiency in healthcare applications. The source code is available on request. 
   </abstract>
   <kwd-group> 
    <kwd>
     Big Data Engineering
    </kwd> 
    <kwd>
      ETL
    </kwd> 
    <kwd>
      Healthcare Sector
    </kwd> 
    <kwd>
      Containerized Applications
    </kwd> 
    <kwd>
      Distributed Computing
    </kwd> 
    <kwd>
      Resource Optimization
    </kwd> 
    <kwd>
      Data Processing Efficiency
    </kwd>
   </kwd-group>
  </article-meta>
 </front>
 <body>
  <sec id="s1">
   <title>1. Introduction</title>
   <p>In the rapidly evolving landscape of Big Data Engineering and Extract, Transform, Load (ETL) processes, the quest for efficiency stands as a pivotal pursuit <xref ref-type="bibr" rid="scirp.136659-1">
     [1]
    </xref>. The volume, variety, and velocity of data are increasing exponentially, with estimates suggesting that the global healthcare data market will reach 463 exabytes by 2025 on generated daily information. The healthcare industry is undergoing a significant transformation, driven by the rapid growth of big data. By 2025, the market is projected to reach $70 billion, a staggering increase of 568% over the past decade <xref ref-type="bibr" rid="scirp.136659-2">
     [2]
    </xref>. As healthcare organizations embrace big data, they face a complex challenge: harnessing its power to enhance decision-making processes and improve patient outcomes. However, this challenge also presents a tremendous opportunity for innovation and growth, enabling healthcare stakeholders to revolutionize the supply chain and deliver more effective, personalized care <xref ref-type="bibr" rid="scirp.136659-3">
     [3]
    </xref>. The healthcare sector, in particular, generates vast amounts of complex data, making it a crucial domain for optimizing data processing efficiency. Efficient ETL processes are essential for harnessing the power of big data in healthcare, enabling timely insights, and improving patient outcomes. Furthermore, the increasing application of Artificial Intelligence (AI) and Machine Learning (ML) in healthcare is poised to significantly impact big data analysis <xref ref-type="bibr" rid="scirp.136659-4">
     [4]
    </xref>. Despite the significance of efficient ETL processes, current big data healthcare pipelines often suffer from inefficiencies, leading to prolonged processing times, resource waste, and reduced scalability <xref ref-type="bibr" rid="scirp.136659-4">
     [4]
    </xref>. The lack of optimized ETL workflows hinders the seamless integration of data from diverse sources, such as electronic health records (EHRs), medical imaging, and wearable devices <xref ref-type="bibr" rid="scirp.136659-5">
     [5]
    </xref>. This integration is critical for realizing the full potential of big data in healthcare, including personalized medicine, predictive analytics, and population health management <xref ref-type="bibr" rid="scirp.136659-6">
     [6]
    </xref>.</p>
   <p>Moreover, the healthcare industry faces a significant gap in leveraging efficient tools and technologies in big data pipelines and data collection methods. Many existing ETL processes rely on traditional, resource-intensive methods, neglecting the benefits of modern distributed computing frameworks and containerization <xref ref-type="bibr" rid="scirp.136659-7">
     [7]
    </xref>. This gap results in:</p>
   <p>The current state of big data processing in healthcare has been analyzed through academic papers, revealing both benefits and challenges. The benefits include improved patient outcomes through data-driven decision making <xref ref-type="bibr" rid="scirp.136659-9">
     [9]
    </xref>, enhanced operational efficiency and reduced costs, and better data management and integration. However, challenges persist, including inefficiencies, scalability issues, prolonged processing times, limited integration of diverse data sources, and inadequate use of resources <xref ref-type="bibr" rid="scirp.136659-2">
     [2]
    </xref> <xref ref-type="bibr" rid="scirp.136659-10">
     [10]
    </xref> <xref ref-type="bibr" rid="scirp.136659-11">
     [11]
    </xref>. Moreover, there is a need for more efficient and scalable ETL processes in the healthcare sector <xref ref-type="bibr" rid="scirp.136659-12">
     [12]
    </xref>.</p>
   <p>This literature review involved a systematic search of peer-reviewed articles from leading healthcare and technology journals, such as the Journal of Healthcare Management and Journal of Big Data. Using keywords like “big data analytics in healthcare” and “healthcare data integration”, the search focused on articles published in English between 2016 and 2024. It specifically targeted studies on the use of big data in healthcare, its benefits, challenges, and future trends, with a focus on ETL processes and data pipeline management. A total of 40 articles were reviewed to give a comprehensive overview of the field.</p>
   <p>Market research studies have identified trends and challenges in healthcare big data. The trends include growing demand for healthcare analytics solutions, increasing adoption of advanced technologies, descriptive analytics dominating the market, and financial analytics to improve patient outcomes <xref ref-type="bibr" rid="scirp.136659-13">
     [13]
    </xref> <xref ref-type="bibr" rid="scirp.136659-14">
     [14]
    </xref>. However, challenges persist, including operational gaps between payers and providers, inaccurate and inconsistent data, high prices of analytics solutions, and lack of technical expertise in healthcare organizations <xref ref-type="bibr" rid="scirp.136659-13">
     [13]
    </xref>-<xref ref-type="bibr" rid="scirp.136659-15">
     [15]
    </xref>.</p>
   <p>Success stories have been examined to understand real-world applications and solutions (25). For instance:</p>
   <p>Consultations with thought leaders in healthcare big data and ETL processes have identified practical challenges. Analysis of publicly available datasets like MIMIC-III <xref ref-type="bibr" rid="scirp.136659-17">
     [17]
    </xref>-<xref ref-type="bibr" rid="scirp.136659-19">
     [19]
    </xref> and PhysioNet <xref ref-type="bibr" rid="scirp.136659-19">
     [19]
    </xref> has highlighted specific challenges and areas for improvement in ETL workflows. Furthermore, research articles from reputable sources like the Journal of Big Data <xref ref-type="bibr" rid="scirp.136659-10">
     [10]
    </xref> and Healthcare Informatics Research <xref ref-type="bibr" rid="scirp.136659-16">
     [16]
    </xref> have also been consulted to gather a comprehensive understanding of the challenges and opportunities in healthcare big data analytics and ETL processes.</p>
   <p>To address this gap, our study explores the domain of efficient Big Data Engineering and ETL processes in the healthcare sector, leveraging the MIMIC-III Clinical Database as a robust foundation. MIMIC-III is a large, freely available database containing deidentified health-related data from over 40,000 patients in critical care units at Beth Israel Deaconess Medical Center between 2001 and 2012. The database integrates comprehensive clinical data, making it widely accessible to researchers under a data use agreement. The demo dataset contains information for 100 patients, including ICU stays, admissions, diagnoses, procedures, and medications. This database is a great example of a use case for a big data pipeline, demonstrating the potential for large-scale data integration and analysis in healthcare research <xref ref-type="bibr" rid="scirp.136659-17">
     [17]
    </xref>-<xref ref-type="bibr" rid="scirp.136659-19">
     [19]
    </xref>. By investigating efficient practices and harnessing the power of PySpark and Docker containerized applications, we aim to contribute valuable insights to the ongoing discourse in the field.</p>
  </sec><sec id="s2">
   <title>2. Methods</title>
   <sec id="s2_1">
    <title>2.1. Selecting a Template (Sub-Heading 2.1)</title>
    <p>We employed a mixed-methods approach, combining both qualitative and quantitative research methods, to gain a comprehensive understanding of the research problem. This approach allowed us to triangulate findings and increase the validity of our results and allowed us to delve deeper into the intricacies of healthcare big data processing. Qualitative methods, such as literature reviews and expert consultations, provided valuable contextual insights, while quantitative methods, like data analysis from the MIMIC-III Clinical Database, offered empirical evidence to support our findings.</p>
    <p>The practical simulation built upon each method further reinforced the validity, reliability, and credibility of our results. By applying diverse approaches and cross-verifying our findings, we were able to mitigate potential biases and limitations inherent in any single research method. This triangulation of findings from various sources ensured a more balanced and nuanced understanding of the complex research problem at hand.</p>
    <p>Ultimately, the mixed-methods approach emerged as the optimal strategy for our study. It enabled us to capture the multifaceted nature of healthcare big data processing and to derive more robust and insightful conclusions. By integrating qualitative and quantitative methods, we were able to delve deeply into the intricacies of the subject matter, providing a comprehensive analysis that enhances the relevance and applicability of our research findings. Details on the methods used, including containerization overhead, orchestration efficiency, and scalability, are provided in <xref ref-type="table" rid="tableTables 1">
      Tables 1
     </xref>-<xref ref-type="bibr" rid="scirp.136659-#t3">
      3
     </xref>, respectively.</p>
    <table-wrap id="table1">
     <label>
      <xref ref-type="table" rid="table1">
       Table 1
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.136659-"></xref>Table 1. Data collection methods for understanding big data processing in healthcare.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td aleft" width="16.48%"><p style="text-align:left">Method</p></td> 
       <td class="custom-bottom-td aleft" width="83.52%"><p style="text-align:left">Description</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td aleft" width="16.48%"><p style="text-align:left">Research Articles</p></td> 
       <td class="custom-top-td aleft" width="83.52%"><p style="text-align:left">Analyzed academic papers to understand the current state of big data processing in healthcare, including:Benefits:</p><p style="text-align:left">Improved patient outcomes through data-driven decision making;</p><p style="text-align:left">Enhanced operational efficiency and reduced costs;</p><p style="text-align:left">Better data management and integration <xref ref-type="bibr" rid="scirp.136659-11">
          [11]
         </xref> <xref ref-type="bibr" rid="scirp.136659-20">
          [20]
         </xref> <xref ref-type="bibr" rid="scirp.136659-21">
          [21]
         </xref>.</p><p style="text-align:left">Challenges:</p><p style="text-align:left">The current state of big data processing in healthcare is characterized by:</p><p style="text-align:left">Inefficiencies;</p><p style="text-align:left">Scalability issues;</p><p style="text-align:left">Prolonged processing times;</p><p style="text-align:left">Limited integration of diverse data sources;</p><p style="text-align:left">Inadequate use of resources <xref ref-type="bibr" rid="scirp.136659-22">
          [22]
         </xref> <xref ref-type="bibr" rid="scirp.136659-12">
          [12]
         </xref> <xref ref-type="bibr" rid="scirp.136659-13">
          [13]
         </xref> <xref ref-type="bibr" rid="scirp.136659-14">
          [14]
         </xref> <xref ref-type="bibr" rid="scirp.136659-15">
          [15]
         </xref>.</p><p style="text-align:left">Need:</p><p style="text-align:left">More efficient and scalable ETL processes in the healthcare sector <xref ref-type="bibr" rid="scirp.136659-23">
          [23]
         </xref>.</p></td> 
      </tr> 
      <tr> 
       <td class="aleft" width="16.48%"><p style="text-align:left">Industry Reports</p></td> 
       <td class="aleft" width="83.52%"><p style="text-align:left">Reviewed market research studies to identify trends and challenges in healthcare big data, including:</p><p style="text-align:left">Trends:</p><p style="text-align:left">Growing demand for healthcare analytics solutions;</p><p style="text-align:left">Increasing adoption of advanced technologies;</p><p style="text-align:left">Descriptive analytics dominates the market;</p><p style="text-align:left">Financial analytics to improve patient outcomes <xref ref-type="bibr" rid="scirp.136659-24">
          [24]
         </xref>-<xref ref-type="bibr" rid="scirp.136659-26">
          [26]
         </xref>.</p><p style="text-align:left">Challenges:</p><p style="text-align:left">Operational gaps between payers and providers;</p><p style="text-align:left">Inaccurate and inconsistent data;</p><p style="text-align:left">High prices of analytics solutions;</p><p style="text-align:left">Lack of technical expertise in healthcare organizations <xref ref-type="bibr" rid="scirp.136659-27">
          [27]
         </xref> <xref ref-type="bibr" rid="scirp.136659-28">
          [28]
         </xref>.</p></td> 
      </tr> 
      <tr> 
       <td class="aleft" width="16.48%"><p style="text-align:left">Case Studies</p></td> 
       <td class="aleft" width="83.52%"><p style="text-align:left">Examined success stories to understand real-world applications and solutions, <xref ref-type="bibr" rid="scirp.136659-29">
          [29]
         </xref>.</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>In addition to the literature review, we also consulted with industry experts and analyzed publicly available datasets to gain a deeper understanding of the practical challenges faced by healthcare organizations.</p>
    <table-wrap id="table2">
     <label>
      <xref ref-type="table" rid="table2">
       Table 2
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.136659-"></xref>Table 2. Supplementary research methods for identifying practical challenges in healthcare big data and ETL processes.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td aleft" width="15.65%"><p style="text-align:left">Method</p></td> 
       <td class="custom-bottom-td aleft" width="84.35%"><p style="text-align:left">Description</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td aleft" width="15.65%"><p style="text-align:left">Industry Experts</p></td> 
       <td class="custom-top-td aleft" width="84.35%"><p style="text-align:left">Consulted with thought leaders in healthcare big data and ETL processes to identify practical challenges.</p></td> 
      </tr> 
      <tr> 
       <td class="aleft" width="15.65%"><p style="text-align:left">Public Datasets</p></td> 
       <td class="aleft" width="84.35%"><p style="text-align:left">Analyzed datasets like MIMIC-III to identify specific challenges and areas for improvement in ETL workflows <xref ref-type="bibr" rid="scirp.136659-18">
          [18]
         </xref>.</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>Through these methods, we identified key challenges and gaps in current ETL processes, including:</p>
    <table-wrap id="table3">
     <label>
      <xref ref-type="table" rid="table3">
       Table 3
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.136659-"></xref>Table 3. Challenges as gaps.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td aleft" width="25.61%"><p style="text-align:left">Challenge</p></td> 
       <td class="custom-bottom-td aleft" width="74.39%"><p style="text-align:left">Description</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td aleft" width="25.61%"><p style="text-align:left">Resource Inefficiency</p></td> 
       <td class="custom-top-td aleft" width="74.39%"><p style="text-align:left">Inefficient use of resources and prolonged processing times.</p></td> 
      </tr> 
      <tr> 
       <td class="aleft" width="25.61%"><p style="text-align:left">Limited Scalability</p></td> 
       <td class="aleft" width="74.39%"><p style="text-align:left">Limited integration of diverse data sources and lack of scalability.</p></td> 
      </tr> 
      <tr> 
       <td class="aleft" width="25.61%"><p style="text-align:left">Suboptimal ETL</p></td> 
       <td class="aleft" width="74.39%"><p style="text-align:left">Lack of optimized ETL workflows and distributed computing frameworks.</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>Summary of Challenges and Gaps</p>
    <p>Our mixed-methods approach revealed a clear need for more efficient and scalable ETL processes in the healthcare sector. The challenges and gaps identified include resource inefficiency, limited scalability, and suboptimal ETL processes. These challenges and gaps served as the foundation for our research and experimentation, guiding our exploration of Containerized application, PySpark, Docker, as potential solutions.</p>
   </sec>
   <sec id="s2_2">
    <title>2.2. Proposed Solutions</title>
    <p>Extract, Transform, Load (ETL) (Equation (1)) is a crucial process in data integration that involves extracting data from multiple sources, transforming it into a standardized format, and loading it into a target system for analysis and reporting <xref ref-type="bibr" rid="scirp.136659-30">
      [30]
     </xref> <xref ref-type="bibr" rid="scirp.136659-31">
      [31]
     </xref>.</p>
    <p>The ETL process can be mathematically represented as:</p>
    <p>ETL = (Extraction, Transformation, Loading) (1)</p>
    <p>where:</p>
    <p>Extraction = Data Collection + Data Cleansing + Data Mapping</p>
    <p>Transformation = Data Transformation + Data Aggregation + Data Quality Check</p>
    <p>Loading = Data Loading + Data Indexing + Data Retrieval</p>
    <p>The ETL equation highlights the interconnectedness of the three stages, emphasizing that each stage is dependent on the previous one. A successful ETL process ensures that data is accurately extracted, transformed, and loaded, enabling organizations to make informed decisions based on reliable and consistent data <xref ref-type="bibr" rid="scirp.136659-32">
      [32]
     </xref>.</p>
    <p>Our methodology is characterized by meticulous experimentation, emphasizing a sophisticated array of technologies:</p>
    <p>This paper not only unearths key findings but also encapsulates pivotal insights with a pragmatic lens. Emphasizing the practical benefits of deploying PySpark within Docker containerized applications using Docker Compose. Our approach is tailored to streamline Big Data Engineering and ETL processes. As we navigate through this exploration, our focus remains unwaveringly committed to contributing insights that hold tangible value within the healthcare domain. The practical implementation is illustrated as shown in <xref ref-type="fig" rid="fig1">
      Figure 1
     </xref> below.</p>
    <p>Our experimental setup aimed to assess the efficiency and effectiveness of four distinct ETL processes within the healthcare domain, focusing on Big Data Engineering principles. We utilized the MIMIC-III Clinical Database, a widely used database in healthcare research, to simulate real-world data processing scenarios. Below, we detail our experimental setup, including the MIMIC-III database, configuration specifics, and the four ETL processes leveraging various technologies such as PySpark, Docker, Docker Compose, Python, Pandas, and PostgreSQL.</p>
    <fig id="fig1" position="float">
     <label>Figure 1</label>
     <caption>
      <title>Figure 1. Practical implementation in the experiment.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/2870738-rId14.jpeg?20241018114727" />
    </fig>
    <p>The MIMIC-III (Medical Information Mart for Intensive Care III) database is a comprehensive collection of de-identified health-related data from over 40,000 patients who stayed in critical care units of the Beth Israel Deaconess Medical Center between 2001 and 2012. It includes information such as demographics, vital signs, laboratory tests, medications, and more, making it a valuable resource for healthcare research and analytics <xref ref-type="bibr" rid="scirp.136659-17">
      [17]
     </xref> <xref ref-type="bibr" rid="scirp.136659-18">
      [18]
     </xref>.</p>
    <p>The following tables provide a data dictionary for the sample datasets used in this research. The LABEVENTS table contains lab test results, while the PATIENTS table includes demographic details. <xref ref-type="table" rid="tableTables 4">
      Tables 4
     </xref>-<xref ref-type="bibr" rid="scirp.136659-#t5">
      5
     </xref>are used to demonstrate data processing and analysis procedures in the research.</p>
    <table-wrap id="table4">
     <label>
      <xref ref-type="table" rid="table4">
       Table 4
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.136659-"></xref>Table 4. LABEVENTS table.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="17.38%"><p style="text-align:center">Field Name</p></td> 
       <td class="custom-bottom-td acenter" width="14.98%"><p style="text-align:center">Data Type</p></td> 
       <td class="custom-bottom-td aleft" width="67.64%"><p style="text-align:left">Description</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="17.38%"><p style="text-align:center">ROW_ID</p></td> 
       <td class="custom-top-td acenter" width="14.98%"><p style="text-align:center">Integer</p></td> 
       <td class="custom-top-td aleft" width="67.64%"><p style="text-align:left">Unique identifier for each row in the lab events data.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="17.38%"><p style="text-align:center">SUBJECT_ID</p></td> 
       <td class="acenter" width="14.98%"><p style="text-align:center">Integer</p></td> 
       <td class="aleft" width="67.64%"><p style="text-align:left">Unique patient identifier linked to the PATIENTS table.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="17.38%"><p style="text-align:center">HADM_ID</p></td> 
       <td class="acenter" width="14.98%"><p style="text-align:center">Float</p></td> 
       <td class="aleft" width="67.64%"><p style="text-align:left">Unique identifier for a patient’s hospital admission.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="17.38%"><p style="text-align:center">ITEMID</p></td> 
       <td class="acenter" width="14.98%"><p style="text-align:center">Integer</p></td> 
       <td class="aleft" width="67.64%"><p style="text-align:left">Identifier for the specific lab test performed.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="17.38%"><p style="text-align:center">CHARTTIME</p></td> 
       <td class="acenter" width="14.98%"><p style="text-align:center">Object</p></td> 
       <td class="aleft" width="67.64%"><p style="text-align:left">Date and time when the lab result was charted.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="17.38%"><p style="text-align:center">VALUE</p></td> 
       <td class="acenter" width="14.98%"><p style="text-align:center">Object</p></td> 
       <td class="aleft" width="67.64%"><p style="text-align:left">The result value for non-numeric lab tests.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="17.38%"><p style="text-align:center">VALUENUM</p></td> 
       <td class="acenter" width="14.98%"><p style="text-align:center">Float</p></td> 
       <td class="aleft" width="67.64%"><p style="text-align:left">The result value for numeric lab tests.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="17.38%"><p style="text-align:center">VALUEUOM</p></td> 
       <td class="acenter" width="14.98%"><p style="text-align:center">Object</p></td> 
       <td class="aleft" width="67.64%"><p style="text-align:left">Unit of measurement for the numeric lab result (e.g., mg/dL, mmol/L).</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="17.38%"><p style="text-align:center">FLAG</p></td> 
       <td class="acenter" width="14.98%"><p style="text-align:center">Object</p></td> 
       <td class="aleft" width="67.64%"><p style="text-align:left">Indicator of whether the result was abnormal (“abnormal”).</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <table-wrap id="table5">
     <label>
      <xref ref-type="table" rid="table5">
       Table 5
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.136659-"></xref>Table 5. PATIENTS Table</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="18.22%"><p style="text-align:center">Field Name</p></td> 
       <td class="custom-bottom-td acenter" width="12.87%"><p style="text-align:center">Data Type</p></td> 
       <td class="custom-bottom-td aleft" width="68.90%"><p style="text-align:left">Description</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="18.22%"><p style="text-align:center">ROW_ID</p></td> 
       <td class="custom-top-td acenter" width="12.87%"><p style="text-align:center">Integer</p></td> 
       <td class="custom-top-td aleft" width="68.90%"><p style="text-align:left">Unique identifier for each row in the patients data.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.22%"><p style="text-align:center">SUBJECT_ID</p></td> 
       <td class="acenter" width="12.87%"><p style="text-align:center">Integer</p></td> 
       <td class="aleft" width="68.90%"><p style="text-align:left">Unique patient identifier.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.22%"><p style="text-align:center">GENDER</p></td> 
       <td class="acenter" width="12.87%"><p style="text-align:center">Object</p></td> 
       <td class="aleft" width="68.90%"><p style="text-align:left">Gender of the patient (Male/Female).</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.22%"><p style="text-align:center">DOB</p></td> 
       <td class="acenter" width="12.87%"><p style="text-align:center">Object</p></td> 
       <td class="aleft" width="68.90%"><p style="text-align:left">Date of birth of the patient.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.22%"><p style="text-align:center">DOD</p></td> 
       <td class="acenter" width="12.87%"><p style="text-align:center">Object</p></td> 
       <td class="aleft" width="68.90%"><p style="text-align:left">Date of death of the patient (if applicable).</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.22%"><p style="text-align:center">DOD_HOSP</p></td> 
       <td class="acenter" width="12.87%"><p style="text-align:center">Object</p></td> 
       <td class="aleft" width="68.90%"><p style="text-align:left">Date of death during hospital admission (if applicable).</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.22%"><p style="text-align:center">DOD_SSN</p></td> 
       <td class="acenter" width="12.87%"><p style="text-align:center">Float</p></td> 
       <td class="aleft" width="68.90%"><p style="text-align:left">Date of death based on social security number (if applicable).</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.22%"><p style="text-align:center">EXPIRE_FLAG</p></td> 
       <td class="acenter" width="12.87%"><p style="text-align:center">Integer</p></td> 
       <td class="aleft" width="68.90%"><p style="text-align:left">Indicator of whether the patient is deceased (1 = Deceased).</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>Our experimental environment was configured as follows:</p>
   </sec>
   <sec id="s2_3">
    <title>2.3. Four Distinct ETL Processes</title>
    <p>1) ETL Process based on Python - Pandas: This process leveraged Python and the Pandas library for data extraction, transformation, and loading (ETL) tasks as shown in <xref ref-type="fig" rid="fig2">
      Figure 2
     </xref>. Python’s versatility and Pandas’ powerful data manipulation capabilities were utilized to process the MIMIC-III dataset efficiently <xref ref-type="bibr" rid="scirp.136659-42">
      [42]
     </xref>.</p>
    <p>2) ETL Process based on PySpark: PySpark, a Python API for Apache Spark, was utilized for this ETL process as shown in <xref ref-type="fig" rid="fig3">
      Figure 3
     </xref>. Spark’s distributed computing framework enabled scalable and high-performance data processing, suitable for handling large-scale healthcare datasets like MIMIC-III <xref ref-type="bibr" rid="scirp.136659-43">
      [43]
     </xref>.</p>
    <p>3) ETL Process based on Docker - Python/Docker Compose: Docker containers were employed to encapsulate Python-based ETL workflows as shown in <xref ref-type="fig" rid="fig4">
      Figure 4
     </xref>. Docker Compose orchestrated the deployment of multiple containers, ensuring</p>
    <fig id="fig2" position="float">
     <label>Figure 2</label>
     <caption>
      <title>Figure 2. Extract, Transform, Load process.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/2870738-rId15.jpeg?20241018114728" />
    </fig>
    <fig id="fig3" position="float">
     <label>Figure 3</label>
     <caption>
      <title>Figure 3. Pyspark ETL process.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/2870738-rId16.jpeg?20241018114728" />
    </fig>
    <fig id="fig4" position="float">
     <label>Figure 4</label>
     <caption>
      <title>Figure 4. Docker-Python ETL process.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/2870738-rId17.jpeg?20241018114728" />
    </fig>
    <p>seamless integration and scalability of the ETL process <xref ref-type="bibr" rid="scirp.136659-44">
      [44]
     </xref> <xref ref-type="bibr" rid="scirp.136659-45">
      [45]
     </xref>.</p>
    <p>4) ETL Process based on Docker - PySpark/Docker Compose: This process combined the power of PySpark with the containerization benefits of Docker, as shown in <xref ref-type="fig" rid="fig5">
      Figure 5
     </xref>. PySpark applications were containerized using Docker, and Docker Compose orchestrated the deployment for optimized scalability and reproducibility <xref ref-type="bibr" rid="scirp.136659-43">
      [43]
     </xref>.</p>
    <fig id="fig5" position="float">
     <label>Figure 5</label>
     <caption>
      <title>Figure 5. Docker-Pyspark ETL process.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/2870738-rId18.jpeg?20241018114728" />
    </fig>
   </sec>
  </sec><sec id="s3">
   <title>3. Results and Analysis</title>
   <sec id="s3_1">
    <title>3.1. Data Collection Process</title>
    <p>To obtain the result data for each ETL process, we defined specific Python functions designed to fetch detailed performance metrics during each step of the workflow. These functions were deployed across all containers, tracking key parameters such as processing time, volume of data processed, and memory usage.</p>
    <p>For each ETL process:</p>
    <p>This structured approach enabled us to compile reliable performance data from the different environments, providing insights into the efficiency and scalability of each ETL method. The collected data formed the basis for the comparative analysis presented in the following sections.</p>
   </sec>
   <sec id="s3_2">
    <title>3.2. Data Processing Overview</title>
    <p>In each ETL process, whether using Python-Pandas, PySpark, or Docker-based solutions, the data processing steps remained consistent across the different tools. Here’s a breakdown of the general workflow:</p>
    <p>1) Data Ingestion:</p>
    <p>The first step involved reading data from CSV files, specifically the LABEVENTS and PATIENTS datasets. These datasets contain patient information and lab results, which were critical for the analysis.</p>
    <p>2) Data Merging:</p>
    <p>The LABEVENTS dataset, containing lab test results, was merged with the PATIENTS dataset using the common SUBJECT_ID field. This step was crucial to link each patient’s lab test results to their demographic and health information.</p>
    <p>3) Data Cleaning:</p>
    <p>After merging, irrelevant columns that were not needed for the analysis were dropped to simplify the dataset and reduce complexity. This step ensures that only the most relevant data remains, improving processing speed and focus.</p>
    <p>4) Data Transformation:</p>
    <p>Date columns, such as birthdate and death date, were transformed into a standard datetime format, enabling easier manipulation and analysis.</p>
    <p>These steps formed the backbone of the data preparation process, ensuring that the datasets were cleaned, merged, and ready for further analysis or reporting. The use of different tools, such as Pandas or PySpark, allowed the workflow to handle large volumes of data efficiently while maintaining consistency in the transformation logic across all implementations.</p>
   </sec>
   <sec id="s3_3">
    <title>3.3. Experimental Comparison: Python - Pandas vs. PySpark ETL Processes</title>
    <p>This section presents the experimental results obtained from two distinct ETL processes: one utilizing Python with the Pandas library, as shown in <xref ref-type="fig" rid="fig6">
      Figure 6
     </xref>, and the other utilizing PySpark. The purpose is to compare their performance in handling data processing tasks.</p>
    <p>The Python-Pandas ETL process exhibited moderate performance, with execution times ranging from 0 to 55 seconds as the dataset size increased up to 2 GB.</p>
    <fig id="fig6" position="float">
     <label>Figure 6</label>
     <caption>
      <title>Figure 6. Comparison results, Pandas vs Pyspark.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/2870738-rId19.jpeg?20241018114729" />
    </fig>
    <p>The PySpark-based ETL process demonstrated varying performance characteristics. For dataset sizes up to 0.25 GB, execution times ranged from 3 to 10 seconds. However, beyond this point, there was a noticeable drop in execution time, stabilizing around 8 seconds for larger datasets.</p>
    <p>This graphical representation not only encapsulates the trends observed but also emphasizes the scalability and efficiency benefits inherent in utilizing PySpark for processing larger datasets compared to the Python-Pandas approach. Notably, the advantage of PySpark becomes pronounced for datasets exceeding 0.25 GB in size, making it the preferred choice for handling substantial volumes of data efficiently.</p>
    <p>The summary of results presented in <xref ref-type="table" rid="table4">
      Table 4
     </xref> aligns with the findings depicted in the linear graph. The Python-Pandas ETL process exhibits a linear increase in execution time proportionate to the dataset size, peaking at 55 seconds for a 2 GB dataset. Conversely, the PySpark ETL process showcases notably more efficient performance, characterized by execution times initially fluctuating but eventually stabilizing around 8 seconds for dataset sizes beyond 0.25 GB, as shown in <xref ref-type="table" rid="table6">
      Table 6
     </xref>.</p>
    <table-wrap id="table6">
     <label>
      <xref ref-type="table" rid="table6">
       Table 6
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.136659-"></xref>Table 6. Performance comparison summary (Percentage report).</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="33.33%"><p style="text-align:center">ETL Process</p></td> 
       <td class="custom-bottom-td acenter" width="33.33%"><p style="text-align:center">Execution Time (seconds)</p></td> 
       <td class="custom-bottom-td acenter" width="33.33%"><p style="text-align:center">Percentage of Baseline</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="33.33%"><p style="text-align:center">Python-Pandas</p></td> 
       <td class="custom-top-td acenter" width="33.33%"><p style="text-align:center">0~55</p></td> 
       <td class="custom-top-td acenter" width="33.33%"><p style="text-align:center">100% (baseline)</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="33.33%"><p style="text-align:center">PySpark</p></td> 
       <td class="acenter" width="33.33%"><p style="text-align:center">3~8</p></td> 
       <td class="acenter" width="33.33%"><p style="text-align:center">14%~15% of baseline</p></td> 
      </tr> 
     </table>
    </table-wrap>
   </sec>
   <sec id="s3_4">
    <title>3.4. Experimental Results: Dockized Python ETL Process with Docker Compose</title>
    <p>A scalable ETL process was successfully implemented using Python, containerized with Docker, and orchestrated with Docker Compose. The experiment involved deploying 10 containers, each representing a worker node, to process large datasets. The results, visualized in the accompanying graph, demonstrate the significant benefits of containerization and orchestration in enhancing the efficiency and scalability of the ETL process. The graph, <xref ref-type="fig" rid="fig7">
      Figure 7
     </xref>, showcases the processing time for each worker node, illustrating the improved performance and distributed processing capabilities achieved through containerization and orchestration.</p>
    <fig id="fig7" position="float">
     <label>Figure 7</label>
     <caption>
      <title>Figure 7. ETL process results, Docker-Python.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/2870738-rId20.jpeg?20241018114730" />
    </fig>
    <p>Containerization Overhead: Impact of Docker containers on overall performance.</p>
    <p>Orchestration Efficiency: Effectiveness of Docker Compose in managing multi-container applications.</p>
    <p>Reproducibility: Consistency and reliability of results across different environments.</p>
    <p>The Dockerized Python ETL process exhibited slightly increased execution times due to containerization overhead.</p>
    <p>Docker Compose effectively managed the orchestration of multiple containers, ensuring streamlined deployment and management.</p>
    <p>Results were highly reproducible across different environments, indicating the reliability and consistency of the Docker-based approach.</p>
    <p>The bar chart illustrates the performance of each container within the Dockerized Python ETL process across different dataset sizes (ranging from 0.25 GB to 2 GB).</p>
    <p>The median range of execution times for each container varies from 0.2 seconds to 12 seconds, indicating the impact of containerization overhead and orchestration efficiency.</p>
    <p>Despite the slight increase in execution times due to containerization, Docker Compose effectively manages the orchestration of multiple containers, ensuring streamlined deployment and management of the ETL process.</p>
    <p>The consistent performance across different environments highlights the reproducibility and reliability of the Docker-based approach for deploying ETL processes.</p>
   </sec>
   <sec id="s3_5">
    <title>3.5. Experimental Results: Dockerized PySpark ETL Process with Docker Compose</title>
    <p>This section presents the experimental results of a containerized ETL process, leveraging Docker for PySpark application deployment and Docker Compose for seamless orchestration. The outcomes are based on a scaled deployment of 4 containers, offering valuable insights into the performance, efficiency, and scalability of the containerized ETL process. The results demonstrate the benefits of containerization, including improved resource utilization, enhanced productivity, and streamlined workflow management. The performance gains are shown in the following graph, <xref ref-type="fig" rid="fig8">
      Figure 8
     </xref>, which illustrates the significant improvements in processing time and throughput achieved through containerization and orchestration.</p>
    <p>Performance Optimization: Impact of containerization on PySpark performance.</p>
    <p>Scalability Enhancement: Utilization of Docker Compose for optimizing scalability.</p>
    <p>Environment Consistency: Reproducibility of results across various deployment environments.</p>
    <p>The Dockerized PySpark ETL process demonstrated efficient performance, leveraging Docker’s containerization benefits without significant overhead.</p>
    <p>Docker Compose effectively managed the scaling of PySpark containers, optimizing resource utilization and improving performance.</p>
    <p>Results were consistent and reproducible across different deployment environments, indicating the robustness and reliability of the Docker-based approach.</p>
    <p>
     <xref ref-type="bibr" rid="scirp.136659-"></xref>Performance Comparison Bar Chart.</p>
    <p>The bar chart illustrates the performance of each PySpark container within the Dockerized ETL process across different dataset sizes.</p>
    <fig id="fig8" position="float">
     <label>Figure 8</label>
     <caption>
      <title>Figure 8. ETL process results, Docker-Pyspark.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/2870738-rId21.jpeg?20241018114732" />
    </fig>
    <p>With only 4 containers, the Dockerized PySpark ETL process showcases efficient performance, demonstrating the scalability and optimization benefits of Docker Compose.</p>
    <p>The consistent performance and reproducibility across different deployment environments highlight the reliability and robustness of the Docker-based approach for deploying PySpark-based ETL processes.</p>
   </sec>
   <sec id="s3_6">
    <title>3.6. Table of Analysis</title>
    <p>The following tables provide a summary of the results from the experiment, which was discussed in earlier Figures 6-8. The results are presented in four tables, each focusing on a specific aspect of the containerized ETL process.</p>
    <p>This table presents the metrics related to containerization and orchestration, including containerization overhead, orchestration efficiency, and reproducibility, as shown in <xref ref-type="table" rid="table7">
      Table 7
     </xref>. The results show the benefits of using Docker and Docker Compose for containerization and orchestration.</p>
    <table-wrap id="table7">
     <label>
      <xref ref-type="table" rid="table7">
       Table 7
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.136659-"></xref>Table 7. Containerization and orchestration metrics.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="33.33%"><p style="text-align:center">Metric</p></td> 
       <td class="custom-bottom-td acenter" width="33.33%"><p style="text-align:center">Dockerized Python</p></td> 
       <td class="custom-bottom-td acenter" width="33.34%"><p style="text-align:center">Dockerized PySpark</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="33.33%"><p style="text-align:center">Containerization Overhead</p></td> 
       <td class="custom-top-td acenter" width="33.33%"><p style="text-align:center">2~5 seconds</p></td> 
       <td class="custom-top-td acenter" width="33.34%"><p style="text-align:center">&lt;1 second</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="33.33%"><p style="text-align:center">Orchestration Efficiency</p></td> 
       <td class="acenter" width="33.33%"><p style="text-align:center">90%</p></td> 
       <td class="acenter" width="33.34%"><p style="text-align:center">95%</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="33.33%"><p style="text-align:center">Reproducibility</p></td> 
       <td class="acenter" width="33.33%"><p style="text-align:center">100%</p></td> 
       <td class="acenter" width="33.34%"><p style="text-align:center">100%</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <table-wrap id="table8">
     <label>
      <xref ref-type="table" rid="table8">
       Table 8
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.136659-"></xref>Table 8. ETL process comparison summary.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="27.09%"><p style="text-align:center">ETL Process</p></td> 
       <td class="custom-bottom-td acenter" width="20.24%"><p style="text-align:center">Scalability</p></td> 
       <td class="custom-bottom-td acenter" width="20.00%"><p style="text-align:center">Efficiency</p></td> 
       <td class="custom-bottom-td acenter" width="32.67%"><p style="text-align:center">Reproducibility</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="27.09%"><p style="text-align:center">Python-Pandas</p></td> 
       <td class="custom-top-td acenter" width="20.24%"><p style="text-align:center">Limited</p></td> 
       <td class="custom-top-td acenter" width="20.00%"><p style="text-align:center">Moderate</p></td> 
       <td class="custom-top-td acenter" width="32.67%"><p style="text-align:center">High</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="27.09%"><p style="text-align:center">PySpark</p></td> 
       <td class="acenter" width="20.24%"><p style="text-align:center">High</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">High</p></td> 
       <td class="acenter" width="32.67%"><p style="text-align:center">High</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="27.09%"><p style="text-align:center">Dockerized Python</p></td> 
       <td class="acenter" width="20.24%"><p style="text-align:center">Moderate</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">Good</p></td> 
       <td class="acenter" width="32.67%"><p style="text-align:center">High</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="27.09%"><p style="text-align:center">Dockerized PySpark</p></td> 
       <td class="acenter" width="20.24%"><p style="text-align:center">High</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">Excellent</p></td> 
       <td class="acenter" width="32.67%"><p style="text-align:center">High</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>This table compares the scalability, efficiency, and reproducibility of traditional Python-Pandas, PySpark, and containerized Python and PySpark ETL processes. The results demonstrate the advantages of containerization in improving scalability, efficiency, and reproducibility, see <xref ref-type="table" rid="table8">
      Table 8
     </xref>.</p>
    <table-wrap id="table9">
     <label>
      <xref ref-type="table" rid="table9">
       Table 9
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.136659-"></xref>Table 9. Dataset size vs. execution time (s).</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="19.99%"><p style="text-align:center">Dataset Size (GB)</p></td> 
       <td class="custom-bottom-td acenter" width="20.00%"><p style="text-align:center">Python-Pandas</p></td> 
       <td class="custom-bottom-td acenter" width="20.00%"><p style="text-align:center">PySpark</p></td> 
       <td class="custom-bottom-td acenter" width="20.00%"><p style="text-align:center">Dockerized Python</p></td> 
       <td class="custom-bottom-td acenter" width="20.00%"><p style="text-align:center">Dockerized PySpark</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="19.99%"><p style="text-align:center">0.25</p></td> 
       <td class="custom-top-td acenter" width="20.00%"><p style="text-align:center">5</p></td> 
       <td class="custom-top-td acenter" width="20.00%"><p style="text-align:center">3</p></td> 
       <td class="custom-top-td acenter" width="20.00%"><p style="text-align:center">0.2</p></td> 
       <td class="custom-top-td acenter" width="20.00%"><p style="text-align:center">0.3</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="19.99%"><p style="text-align:center">0.5</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">15</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">5</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">0.5</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">0.5</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="19.99%"><p style="text-align:center">1</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">30</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">8</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">1</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">0.8</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="19.99%"><p style="text-align:center">2</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">55</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">8</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">2</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">0.8</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>As shown as <xref ref-type="table" rid="table9">
      Table 9
     </xref>, this table shows the execution time in seconds for different dataset sizes using traditional Python-Pandas, PySpark, and containerized Python and PySpark ETL processes. The results illustrate the improved performance of containerized ETL processes, especially for larger datasets.</p>
    <table-wrap id="table10">
     <label>
      <xref ref-type="table" rid="table10">
       Table 10
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.136659-"></xref>Table 10. Containerization and orchestration benefits.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="33.33%"><p style="text-align:center">Benefit</p></td> 
       <td class="custom-bottom-td acenter" width="33.33%"><p style="text-align:center">Dockerized Python</p></td> 
       <td class="custom-bottom-td acenter" width="33.34%"><p style="text-align:center">Dockerized PySpark</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="33.33%"><p style="text-align:center">Improved Scalability</p></td> 
       <td class="custom-top-td acenter" width="33.33%"><p style="text-align:center"></p></td> 
       <td class="custom-top-td acenter" width="33.34%"><p style="text-align:center">(up to 4 containers)</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="33.33%"><p style="text-align:center">Increased Efficiency</p></td> 
       <td class="acenter" width="33.33%"><p style="text-align:center">(10%~20% increase)</p></td> 
       <td class="acenter" width="33.34%"><p style="text-align:center">(30%~40% increase)</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="33.33%"><p style="text-align:center">Enhanced Reproducibility</p></td> 
       <td class="acenter" width="33.33%"><p style="text-align:center">(consistent results across environments)</p></td> 
       <td class="acenter" width="33.34%"><p style="text-align:center">(consistent results across environments)</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="33.33%"><p style="text-align:center">Simplified Deployment</p></td> 
       <td class="acenter" width="33.33%"><p style="text-align:center">(easy deployment with Docker Compose)</p></td> 
       <td class="acenter" width="33.34%"><p style="text-align:center">(easy deployment with Docker Compose)</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="33.33%"><p style="text-align:center">Streamlined Management</p></td> 
       <td class="acenter" width="33.33%"><p style="text-align:center">(easy management with Docker Compose)</p></td> 
       <td class="acenter" width="33.34%"><p style="text-align:center">(easy management with Docker Compose)</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>As shown as <xref ref-type="table" rid="table10">
      Table 10
     </xref>, this table highlights the benefits of using containerization and orchestration with Docker and Docker Compose, including improved scalability, increased efficiency, enhanced reproducibility, simplified deployment, and streamlined management. The results demonstrate the advantages of using containerization and orchestration in ETL processes.</p>
   </sec>
  </sec><sec id="s4">
   <title>4. Discussion</title>
   <p>The healthcare industry faces significant challenges in harnessing the power of big data, including inefficient ETL processes, prolonged processing times, and limited integration of diverse data sources <xref ref-type="bibr" rid="scirp.136659-4">
     [4]
    </xref>. The lack of optimized ETL workflows hinders the seamless integration of data from diverse sources, such as electronic health records, medical imaging, and wearable devices, restricting the scope of analytics and insights <xref ref-type="bibr" rid="scirp.136659-6">
     [6]
    </xref>. This results in delayed or inaccurate diagnoses, ineffective treatment plans, and poor patient outcomes <xref ref-type="bibr" rid="scirp.136659-46">
     [46]
    </xref>. Moreover, the increasing volume and complexity of healthcare data exacerbate the need for efficient data mining processes that can handle large datasets and provide real-time insights <xref ref-type="bibr" rid="scirp.136659-47">
     [47]
    </xref>.</p>
   <p>Based on the quantitative and qualitative results demonstrated by our research, we showed that containerized applications and parallel processing can significantly improve the efficiency and scalability of ETL processes in healthcare systems. Our results revealed that PySpark and Docker outperformed traditional Python-Pandas approaches in terms of execution time, scalability, and reproducibility. Additionally, our qualitative analysis highlighted the benefits of containerization and parallel processing in improving data integration, data quality, and data analytics in healthcare. Moreover, our approach is cost and time-efficient when deployed on cloud computing services, such as Amazon Web Services or Microsoft Azure, allowing for scalable and on-demand processing of large healthcare datasets.</p>
   <p>The implications of our research are significant, as it provides a solution to the long-standing problem of inefficient ETL processes in healthcare. By adopting containerized applications and parallel processing, healthcare organizations can improve the efficiency and effectiveness of their data processing workflows, ultimately leading to better patient outcomes and improved healthcare services. Our study demonstrates the necessity of leveraging advanced technologies and tools in big data processing to address the complex challenges facing the healthcare industry.</p>
  </sec><sec id="s5">
   <title>5. Conclusions</title>
   <p>In conclusion, this study successfully demonstrates the efficacy of a cutting-edge approach to ETL processes in the healthcare domain, harnessing the power of PySpark, Docker, and Docker Compose. By leveraging these technologies, we achieved scalable, efficient, and reproducible data processing, paving the way for enhanced Big Data Engineering and informed decision-making in healthcare. The findings of this research have significant implications for the field, highlighting the potential for streamlined ETL processes that can handle large datasets with ease, and providing a foundation for future research and innovation in healthcare data management.</p>
   <p>A specific use case in healthcare where this approach can have a significant impact is in the analysis of Electronic Health Records (EHRs). EHRs contain vast amounts of patient data, including medical histories, medications, test results, and treatment plans. By applying the PySpark-Docker-Docker Compose approach to EHR data, healthcare organizations can:</p>
   <p>For instance, this approach can be used to identify high-risk patients, track disease progression, and evaluate the effectiveness of treatment plans. By leveraging the power of PySpark, Docker, and Docker Compose, healthcare organizations can unlock the full potential of EHR data, improve patient care, and reduce healthcare costs.</p>
  </sec><sec id="s6">
   <title>Acknowledgments</title>
   <p>We extend our gratitude to the individuals and organizations whose contributions were instrumental in the completion of this research endeavor.</p>
   <p>Firstly, we would like to express our sincere appreciation to the team behind the MIMIC-III Clinical Database for providing a robust foundation for our study. The availability of this invaluable resource facilitated our exploration into efficient Big Data Engineering and Extract, Transform, Load (ETL) processes in the healthcare sector.</p>
   <p>We are also indebted to the thought leaders and researchers whose work laid the groundwork for our investigation. Their insights, as documented in the literature, guided our understanding of the challenges and opportunities in healthcare big data analytics.</p>
   <p>Additionally, we extend our thanks to the industry experts who generously shared their knowledge and expertise during our consultations. Their insights provided valuable real-world perspectives, enriching our study and informing our experimental approach.</p>
   <p>Furthermore, we acknowledge the support of our colleagues and peers who provided valuable feedback and encouragement throughout the research process. Their contributions fostered a collaborative environment conducive to innovation and discovery.</p>
  </sec>
 </body><back>
  <ref-list>
   <title>References</title>
   <ref id="scirp.136659-ref1">
    <label>1</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Johnson, E. and Miller, R. (2021) Harnessing the Data Revolution: Big Data’s Role in Transforming Industries. Journal of Science&amp;Technology, 2, 32-39.
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref2">
    <label>2</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Cozzoli, N., Salvatore, F.P., Faccilongo, N. and Milone, M. (2022) How Can Big Data Analytics Be Used for Healthcare Organization Management? Literary Framework and Future Research from a Systematic Review. BMC Health Services Research, 22, Article No. 809. &gt;https://doi.org/10.1186/s12913-022-08167-z
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref3">
    <label>3</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Dicuonzo, G., Galeone, G., Shini, M. and Massari, A. (2022) Towards the Use of Big Data in Healthcare: A Literature Review. Healthcare, 10, Article No. 1232. &gt;https://doi.org/10.3390/healthcare10071232
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref4">
    <label>4</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Raghupathi, W. and Raghupathi, V. (2014) Big Data Analytics in Healthcare: Promise and Potential. Health Information Science and Systems, 2, Article No. 3. &gt;https://doi.org/10.1186/2047-2501-2-3
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref5">
    <label>5</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Pannunzio, V., Kleinsmann, M., Snelders, D. and Raijmakers, J. (2023) From Digital Health to Learning Health Systems: Four Approaches to Using Data for Digital Health Design. Health Systems, 12, 481-494. &gt;https://doi.org/10.1080/20476965.2023.2284712
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref6">
    <label>6</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Lipovac, I. and Babac, M.B. (2024) Developing a Data Pipeline Solution for Big Data Processing. International Journal of Data Mining, Modelling and Management, 16, 1-22. &gt;https://doi.org/10.1504/ijdmmm.2024.136221
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref7">
    <label>7</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Cheng, K.Y., Pazmino, S. and Schreiweis, B. (2022) ETL Processes for Integrating Healthcare Data—Tools and Architecture Patterns. In: Studies in Health Technology and Informatics, IOS Press, 151-156. &gt;https://doi.org/10.3233/shti220974
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref8">
    <label>8</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Rossi, R.L. and Grifantini, R.M. (2018) Big Data: Challenge and Opportunity for Translational and Industrial Research in Healthcare. Frontiers in Digital Humanities, 5, Article No. 13. &gt;https://doi.org/10.3389/fdigh.2018.00013
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref9">
    <label>9</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Berg, K., Doktorchik, C., Quan, H. and Saini, V. (2022) Automating Data Collection Methods in Electronic Health Record Systems: A Social Determinant of Health (SDOH) Viewpoint. Health Systems, 12, 472-480. &gt;https://doi.org/10.1080/20476965.2022.2075796
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref10">
    <label>10</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Dash, S., Shakyawar, S.K., Sharma, M. and Kaushik, S. (2019) Big Data in Healthcare: Management, Analysis and Future Prospects. Journal of Big Data, 6, Article No. 54. &gt;https://doi.org/10.1186/s40537-019-0217-0
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref11">
    <label>11</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Batko, K. and Ślęzak, A. (2022) The Use of Big Data Analytics in Healthcare. Journal of Big Data, 9, Article No. 3. &gt;https://doi.org/10.1186/s40537-021-00553-4
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref12">
    <label>12</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ismail, A., Shehab, A. and El-Henawy, I.M. (2018) Healthcare Analysis in Smart Big Data Analytics: Reviews, Challenges and Recommendations. In: Hassanien, A.E., Elhoseny, M., Ahmed, S.H. and Singh, A.K., Eds., Security in Smart Cities: Models, Applications, and Challenges, Springer International Publishing, 27-45. &gt;https://doi.org/10.1007/978-3-030-01560-2_2
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref13">
    <label>13</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kashyap, R. (2019) Big Data Analytics Challenges and Solutions. In: Dey, N., Das, H., Naik, B. and Behera, H.S., Eds., Big Data Analytics for Intelligent Healthcare Management, Elsevier, 19-41. &gt;https://doi.org/10.1016/b978-0-12-818146-1.00002-7
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref14">
    <label>14</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kraus, J.M., Lausser, L., Kuhn, P., Jobst, F., Bock, M., Halanke, C., et al. (2018) Big Data and Precision Medicine: Challenges and Strategies with Healthcare Data. International Journal of Data Science and Analytics, 6, 241-249. &gt;https://doi.org/10.1007/s41060-018-0095-0
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref15">
    <label>15</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Seenivasan, D. (2023) Improving the Performance of the ETL Jobs. International Journal of Computer Trends and Technology, 71, 27-33. &gt;https://doi.org/10.14445/22312803/ijctt-v71i3p105
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref16">
    <label>16</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Johnson, A., et al. (2019) Mimic-III Clinical Database Demo (version 1.4). Physionet.
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref17">
    <label>17</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Johnson, A., et al. (2019) “Mimic-III Clinical Database Demo” (Version 1.4). Physionet (2019).
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref18">
    <label>18</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Johnson, A.E.W., Pollard, T.J., Shen, L., Lehman, L.H., Feng, M., Ghassemi, M., et al. (2016) MIMIC-III, a Freely Accessible Critical Care Database. Scientific Data, 3, Article ID: 160035. &gt;https://doi.org/10.1038/sdata.2016.35
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref19">
    <label>19</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Goldberger, A.L., Amaral, L.A.N., Glass, L., Hausdorff, J.M., Ivanov, P.C., Mark, R.G., et al. (2000) Physiobank, Physiotoolkit, and Physionet: Components of a New Research Resource for Complex Physiologic Signals. Circulation, 101, e215-e220. &gt;https://doi.org/10.1161/01.cir.101.23.e215
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref20">
    <label>20</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Onyemachi, N.C. and Nonyelum, O.F. (2019) Big Data Analytics in Healthcare: A Review. 2019 15th International Conference on Electronics, Computer and Computation (ICECCO), Abuja, 10-12 December 2019, 1-5. &gt;https://doi.org/10.1109/icecco48375.2019.9043183
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref21">
    <label>21</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Wang, Y., Kung, L. and Byrd, T.A. (2018) Big Data Analytics: Understanding Its Capabilities and Potential Benefits for Healthcare Organizations. Technological Forecasting and Social Change, 126, 3-13. &gt;https://doi.org/10.1016/j.techfore.2015.12.019
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref22">
    <label>22</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Hong, L., Luo, M., Wang, R., Lu, P., Lu, W. and Lu, L. (2018) Big Data in Health Care: Applications and Challenges. Data and Information Management, 2, 175-197. &gt;https://doi.org/10.2478/dim-2018-0014
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref23">
    <label>23</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Dabral, S. and Mohana, R. (2023) Healthcare Data Pipeline.
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref24">
    <label>24</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Saheb, T. and Izadi, L. (2019) Paradigm of IoT Big Data Analytics in the Healthcare Industry: A Review of Scientific Literature and Mapping of Research Trends. Telematics and Informatics, 41, 70-85. &gt;https://doi.org/10.1016/j.tele.2019.03.005
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref25">
    <label>25</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Rehman, A., Naz, S. and Razzak, I. (2021) Leveraging Big Data Analytics in Healthcare Enhancement: Trends, Challenges and Opportunities. Multimedia Systems, 28, 1339-1371. &gt;https://doi.org/10.1007/s00530-020-00736-8
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref26">
    <label>26</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ariffin, N., Yunus, A.M. and Kadir, I. (2021) The Role of Big Data in the Healthcare Industry. Journal of Islamic, 6, 235-245.
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref27">
    <label>27</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Karatas, M., Eriskin, L., Deveci, M., Pamucar, D. and Garg, H. (2022) Big Data for Healthcare Industry 4.0: Applications, Challenges and Future Perspectives. Expert Systems with Applications, 200, Article ID: 116912. &gt;https://doi.org/10.1016/j.eswa.2022.116912
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref28">
    <label>28</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kruse, C.S., Goswamy, R., Raval, Y. and Marawi, S. (2016) Challenges and Opportunities of Big Data in Health Care: A Systematic Review. JMIR Medical Informatics, 4, e38. &gt;https://doi.org/10.2196/medinform.5359
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref29">
    <label>29</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Raj, A., Bosch, J., Olsson, H.H. and Wang, T.J. (2020) Modelling Data Pipelines. 2020 46th Euromicro Conference on Software Engineering and Advanced Applications (SEAA), Portoroz, 26-28 August 2020, 13-20. &gt;https://doi.org/10.1109/seaa51224.2020.00014
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref30">
    <label>30</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Vyas, S. and Vaishnav, P. (2017) A Comparative Study of Various ETL Process and Their Testing Techniques in Data Warehouse. Journal of Statistics and Management Systems, 20, 753-763. &gt;https://doi.org/10.1080/09720510.2017.1395194
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref31">
    <label>31</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Rahman, N., Kumar, N. and Rutz, D. (2016) Managing Application Compatibility during ETL Tools and Environment Upgrades. Journal of Decision Systems, 25, 136-150. &gt;https://doi.org/10.1080/12460125.2016.1138392
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref32">
    <label>32</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Diouf, P.S., Boly, A. and Ndiaye, S. (2018). Variety of Data in the ETL Processes in the Cloud: State of the Art. 2018 IEEE International Conference on Innovative Research and Development (ICIRD), Bangkok, 11-12 May 2018, 1-5. &gt;https://doi.org/10.1109/icird.2018.8376308
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref33">
    <label>33</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Singh, P. (2021) Manage Data with Pyspark. In: Singh, P., Machine Learning with PySpark: With Natural Language Processing and Recommender Systems, Apress, 15-37. &gt;https://doi.org/10.1007/978-1-4842-7777-5_2
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref34">
    <label>34</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Lee, D. and Drabas, T. (2017) Learning Pyspark. Packt Publishing Ltd. 
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref35">
    <label>35</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Docker, I. (2020). &gt;https://www.docker.com/what-docker
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref36">
    <label>36</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Cook, J. (2017) Docker for Data Science: Building Scalable and Extensible Data Infrastructure around the Jupyter Notebook Server.
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref37">
    <label>37</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Turnbull, J. (2014) The Docker Book: Containerization Is the New Virtualization.
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref38">
    <label>38</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Gkatziouras, E. (2022) A Developer’s Essential Guide to Docker Compose: Simplify the Development and Orchestration of Multi-Container Applications. Packt Publishing Ltd.
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref39">
    <label>39</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Lutz, M. (2001) Programming Python. O’Reilly Media, Inc.
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref40">
    <label>40</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     McKinney, W. (2011) Pandas: A Foundational Python Library for Data Analysis and Statistics. Python for High Performance and Scientific Computing, 14, 1-9.
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref41">
    <label>41</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Obe, R.O. and Hsu, L.S. (2017) Postgresql: Up and Running: A Practical Guide to the Advanced Open Source Database. O’Reilly Media, Inc.
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref42">
    <label>42</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Mukhopadhyay, S. and Samanta, P. (2022) ETL with Python. In: Advanced Data Analytics Using Python: With Architectural Patterns, Text and Image Classification, and Optimization Techniques, Apress, 23-52. &gt;https://doi.org/10.1007/978-1-4842-8005-8_2
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref43">
    <label>43</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Batmaci, G. (2022) Etl Data Pipelines Configurations in Spark.
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref44">
    <label>44</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Zhou, N., Zhou, H. and Hoppe, D. (2023) Containerization for High Performance Computing Systems: Survey and Prospects. IEEE Transactions on Software Engineering, 49, 2722-2740. &gt;https://doi.org/10.1109/tse.2022.3229221
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref45">
    <label>45</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Bhat, S., Bhat, S. and Karkal (2018) Practical Docker with Python. Springer.
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref46">
    <label>46</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Castaneda, C., Nalley, K., Mannion, C., Bhattacharyya, P., Blake, P., Pecora, A., et al. (2015) Clinical Decision Support Systems for Improving Diagnostic Accuracy and Achieving Precision Medicine. Journal of Clinical Bioinformatics, 5, Article No. 4. &gt;https://doi.org/10.1186/s13336-015-0019-3
    </mixed-citation>
   </ref>
   <ref id="scirp.136659-ref47">
    <label>47</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ferrão, J.C., Oliveira, M.D., Janela, F., Martins, H.M.G. and Gartner, D. (2020) Can Structured EHR Data Support Clinical Coding? A Data Mining Approach. Health Systems, 10, 138-161. &gt;https://doi.org/10.1080/20476965.2020.1729666
    </mixed-citation>
   </ref>
  </ref-list>
 </back>
</article>