<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">jss</journal-id>
      <journal-title-group>
        <journal-title>Open Journal of Social Sciences</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2327-5960</issn>
      <issn pub-type="ppub">2327-5952</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/jss.2026.141018</article-id>
      <article-id pub-id-type="publisher-id">jss-148946</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Business</subject>
          <subject>Economics</subject>
          <subject>Social Sciences</subject>
          <subject>Humanities</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Predicting Survey Response Rates Using XGBoost: A Case Study on Organizational Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Hakemi</surname>
            <given-names>Aida</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Moeini</surname>
            <given-names>Rezza</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Masrom</surname>
            <given-names>Maslin</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> Cultural Infusion Pty Ltd., Melbourne, Australia </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The authors declare no conflicts of interest regarding the publication of this paper.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>31</day>
        <month>12</month>
        <year>2025</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>12</month>
        <year>2025</year>
      </pub-date>
      <volume>14</volume>
      <issue>01</issue>
      <fpage>283</fpage>
      <lpage>292</lpage>
      <history>
        <date date-type="received">
          <day>08</day>
          <month>10</month>
          <year>2025</year>
        </date>
        <date date-type="accepted">
          <day>17</day>
          <month>01</month>
          <year>2026</year>
        </date>
        <date date-type="published">
          <day>20</day>
          <month>01</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2026 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/jss.2026.141018">https://doi.org/10.4236/jss.2026.141018</self-uri>
      <abstract>
        <p>Accurate prediction of survey response rates is essential for optimizing survey design and ensuring high-quality data collection. Traditional methods often struggle to capture the complexity and multidimensionality of organizational datasets. This study applies the extreme Gradient Boosting (XGBoost) algorithm to predict response rates using organizational and demographic features. The model was trained on features including age, gender, job level, send hour, weekday, allowed response window, number of reminders, and total sent forms. The XGBoost model achieved strong predictive performance with an R<sup>2</sup> score of 0.85 and a Mean Squared Error (MSE) of 0.02, reflecting the high accuracy in predicting response rates. Analysis of feature importance revealed that sent forms (46.6%) and Reminder (42.6%) were the most influential factors, while job_level (2.55%) and weekday (2.67%) also contributed to response behavior. Scatter plots of actual versus predicted response rates confirmed minimal deviation, demonstrating the reliability of the model. These results highlight the potential of machine learning techniques, particularly XGBoost, in accurately modeling survey response rates. Understanding feature importance allows researchers and organizations to strategically adjust survey design elements such as the number of invitations and reminders, to maximize participation.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>Survey Response Rate</kwd>
        <kwd>XGBoost</kwd>
        <kwd>Machine Learning</kwd>
        <kwd>Feature Importance</kwd>
        <kwd>Predictive Modeling</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction</title>
      <p>Survey response rates are a critical metric in organizational research, serving as a key indicator of data quality and representativeness ([<xref ref-type="bibr" rid="B7">7</xref>]; [<xref ref-type="bibr" rid="B3">3</xref>]; [<xref ref-type="bibr" rid="B10">10</xref>]). High response rates ensure that survey findings accurately reflect the views and behaviours of the target population, thereby enhancing the validity of conclusions drawn from the data ([<xref ref-type="bibr" rid="B2">2</xref>]; [<xref ref-type="bibr" rid="B1">1</xref>]). However, achieving optimal response rates remains a significant challenge for many organizations, particularly in the context of employee engagement surveys, customer satisfaction assessments, and other organizational studies ([<xref ref-type="bibr" rid="B6">6</xref>]; [<xref ref-type="bibr" rid="B5">5</xref>]).</p>
      <p>Traditional methods for predicting survey response rates often rely on statistical techniques that may not fully capture the complex, non-linear relationships inherent in the data ([<xref ref-type="bibr" rid="B6">6</xref>]; [<xref ref-type="bibr" rid="B5">5</xref>]). These methods can struggle to account for the multitude of factors influencing response behaviour, such as demographic characteristics, timing of survey invitations, and the number of reminders sent ([<xref ref-type="bibr" rid="B9">9</xref>]; [<xref ref-type="bibr" rid="B12">12</xref>]; [<xref ref-type="bibr" rid="B10">10</xref>]; [<xref ref-type="bibr" rid="B15">15</xref>]). As a result, there is a growing interest in exploring more advanced analytical approaches that can provide more accurate and nuanced predictions.</p>
      <p>Machine learning (ML) techniques, particularly ensemble methods like Extreme Gradient Boosting (XGBoost), have emerged as powerful tools for predictive modelling in various domains, including survey research ([<xref ref-type="bibr" rid="B4">4</xref>]; [<xref ref-type="bibr" rid="B2">2</xref>]; [<xref ref-type="bibr" rid="B16">16</xref>]). XGBoost is renowned for its high performance and efficiency in handling large datasets with complex interactions among variables ([<xref ref-type="bibr" rid="B4">4</xref>]; [<xref ref-type="bibr" rid="B6">6</xref>]). Its ability to model non-linear relationships and interactions makes it particularly suited for predicting survey response rates, where multiple factors interplay to influence participant behaviour ([<xref ref-type="bibr" rid="B9">9</xref>]; [<xref ref-type="bibr" rid="B12">12</xref>]). </p>
      <p>In this study, we apply XGBoost to predict survey response rates using a comprehensive dataset that includes organizational and demographic features. The dataset encompasses variables such as age, gender, job level, timing of survey invitations, allowed response window, number of reminders sent, and total number of forms sent ([<xref ref-type="bibr" rid="B7">7</xref>]; [<xref ref-type="bibr" rid="B1">1</xref>]). By leveraging these features, we aim to develop a predictive model that can accurately forecast response rates, thereby enabling organizations to optimize their survey strategies ([<xref ref-type="bibr" rid="B6">6</xref>]; [<xref ref-type="bibr" rid="B2">2</xref>]).</p>
      <p>Our findings indicate that the XGBoost model achieves a high level of accuracy, with an R<sup>2</sup> score of 0.85 and a Mean Squared Error (MSE) of 0.02. Feature importance analysis reveals that the number of forms sent and the number of reminders are the most influential factors in determining response rates. These insights suggest that strategic adjustments in survey design, such as increasing the number of reminders or forms sent, can significantly enhance participation rates ([<xref ref-type="bibr" rid="B9">9</xref>]; [<xref ref-type="bibr" rid="B7">7</xref>]; [<xref ref-type="bibr" rid="B2">2</xref>]).</p>
      <p>This research contributes to the field by demonstrating the efficacy of machine learning techniques in predicting survey response rates within organizational contexts. The application of XGBoost provides a robust framework for understanding and improving survey participation, offering practical implications for researchers and practitioners aiming to enhance the quality and reliability of survey-based data ([<xref ref-type="bibr" rid="B9">9</xref>]; [<xref ref-type="bibr" rid="B12">12</xref>]; [<xref ref-type="bibr" rid="B17">17</xref>]).</p>
    </sec>
    <sec id="sec2">
      <title>2. Methods</title>
      <sec id="sec2dot1">
        <title>2.1. Data Collection</title>
        <p>The present study utilizes a dataset comprising 3400 survey records collected between 2023 and 2025 from multiple client organizations that participated in Cultural Infusion’s Diversity and Inclusion survey initiatives across Australia ([<xref ref-type="bibr" rid="B9">9</xref>]; [<xref ref-type="bibr" rid="B12">12</xref>]; [<xref ref-type="bibr" rid="B3">3</xref>]). Each record corresponds to an individual survey response and integrates both demographic and organizational context variables. The dataset includes the following features:</p>
        <p><bold>Org_id</bold><bold>/</bold><bold>Org_name</bold><bold>:</bold> unique identifiers representing the participating organizations.</p>
        <p><bold>Respondent_id</bold><bold>:</bold> anonymized identifier assigned to each individual participant.</p>
        <p><bold>Age, Gender,</bold><bold>Job_level</bold><bold>:</bold> demographic and hierarchical information characterizing the respondent’s profile within the organization.</p>
        <p><bold>Send_timestamp</bold><bold>,</bold><bold>Send</bold><bold>_hour</bold><bold>, Weekday:</bold> temporal variables capturing when the survey invitations were dispatched.</p>
        <p><bold>ExpireTime</bold><bold>and</bold><bold>Allowed</bold><bold>_response_window_hours</bold><bold>:</bold> parameters defining the time window available for survey completion.</p>
        <p><bold>Number_of_respondents</bold><bold>:</bold>the total number of employees who received the survey invitation within each organization. </p>
        <p><bold>Reminder:</bold> a binary variable indicating whether a follow-up reminder email was sent to participants.</p>
        <p>The target variable, Response_rate, was computed as the proportion of completed survey responses to the total number of distributed invitations for each organization or survey batch.</p>
        <p>This dataset structure enables an integrated analysis of both individual-level and organizational-level determinants of survey participation, offering a robust framework to investigate behavioural and contextual factors influencing response rates in corporate diversity survey settings ([<xref ref-type="bibr" rid="B7">7</xref>]; [<xref ref-type="bibr" rid="B17">17</xref>]; [<xref ref-type="bibr" rid="B1">1</xref>]; [<xref ref-type="bibr" rid="B14">14</xref>]).</p>
      </sec>
      <sec id="sec2dot2">
        <title>2.2. Data Preprocessing</title>
        <p>Prior to modelling, the dataset was prepared to ensure compatibility with machine learning algorithms. Categorical variables, including Gender, Job Level, and Weekday, were numerically encoded ([<xref ref-type="bibr" rid="B1">1</xref>]; [<xref ref-type="bibr" rid="B2">2</xref>]). The response rate for each record was calculated as the ratio of actual respondents to total forms distributed. These preprocessing steps ensured that the data were structured and suitable for training and evaluating predictive models, while preserving the natural variability in survey participation ([<xref ref-type="bibr" rid="B12">12</xref>]; [<xref ref-type="bibr" rid="B7">7</xref>]; [<xref ref-type="bibr" rid="B6">6</xref>]; [<xref ref-type="bibr" rid="B17">17</xref>]; [<xref ref-type="bibr" rid="B3">3</xref>]).</p>
      </sec>
      <sec id="sec2dot3">
        <title>2.3. Model Development</title>
        <p>The response rate was modelled using Extreme Gradient Boosting (XGBoost), an ensemble learning method based on gradient-boosted decision trees, recognized for its high performance in handling large, complex datasets ([<xref ref-type="bibr" rid="B4">4</xref>]; [<xref ref-type="bibr" rid="B1">1</xref>]). The dataset was partitioned into training and testing subsets, with 80% allocated for model training and 20% reserved for evaluation ([<xref ref-type="bibr" rid="B9">9</xref>]; [<xref ref-type="bibr" rid="B12">12</xref>]). The model was trained using default hyperparameters to maintain interpretability and focus on feature relationships rather than optimization. Future work will include parameter tuning to further enhance performance. Model performance was assessed using the coefficient of determination (R<sup>2</sup>) and Mean Squared Error (MSE) to quantify predictive accuracy ([<xref ref-type="bibr" rid="B7">7</xref>]; [<xref ref-type="bibr" rid="B6">6</xref>]). This approach allowed the investigation of both linear and non-linear effects of organizational and demographic factors on survey participation.</p>
      </sec>
      <sec id="sec2dot4">
        <title>2.4. Feature Importance Analysis</title>
        <p>Following model training, feature importance scores were computed using the built-in function of the XGBoost algorithm ([<xref ref-type="bibr" rid="B17">17</xref>]; [<xref ref-type="bibr" rid="B2">2</xref>]), which quantifies each variable’s contribution to reducing prediction error across the ensemble of trees. The results were visualized using a bar plot (<xref ref-type="fig" rid="fig1">Figure 1</xref>), enabling a clear comparison of the relative impact of demographic, temporal, and organizational factors.</p>
        <p>The analysis revealed that the number of forms sent, and the number of reminders were the most influential determinants of response behaviour, with importance scores of 0.523 and 0.311, respectively ([<xref ref-type="bibr" rid="B9">9</xref>]; [<xref ref-type="bibr" rid="B6">6</xref>]). Temporal attributes, such as the allowed response window (0.038) and the weekday of survey distribution (0.031), also contributed meaningfully, indicating that timing can influence participation. Demographic characteristics including age (0.013), gender (0.027), and job level (0.030) as well as the hour of sending (0.027) showed smaller, yet notable effects ([<xref ref-type="bibr" rid="B7">7</xref>]; [<xref ref-type="bibr" rid="B17">17</xref>]).</p>
        <p>These findings highlight that organizational practices, survey scheduling, and individual-level differences jointly affect response rates. Understanding these contributions allows organizations to strategically optimize survey distribution and follow-up strategies to enhance participation.</p>
        <p>To further validate model performance, a scatter plot comparing actual and predicted response rates (<xref ref-type="fig" rid="fig2">Figure 2</xref>) was generated. The alignment of points along the diagonal line indicated that the model was able to accurately capture the underlying patterns of survey participation ([<xref ref-type="bibr" rid="B4">4</xref>]; [<xref ref-type="bibr" rid="B3">3</xref>]), thereby supporting the validity and robustness of the feature importance findings.</p>
        <p>Overall, these analyses highlight the multi-faceted nature of survey response rates, where organizational practices, survey timing, and individual respondent characteristics simultaneously influence the outcomes.</p>
        <fig id="fig1">
          <label>Figure 1</label>
          <graphic xlink:href="https://html.scirp.org/file/6500848-rId11.jpeg?20260121024539" />
        </fig>
        <p>Figure 1. Feature importance scores of predictors for survey response rates, showing the relative impact of demographic, temporal, and organizational factors.</p>
        <fig id="fig2">
          <label>Figure 2</label>
          <graphic xlink:href="https://html.scirp.org/file/6500848-rId12.jpeg?20260121024539" />
        </fig>
        <p>Figure 2. Scatter plot comparing actual and predicted survey response rates, indicating model performance and alignment with observed data.</p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. Results</title>
      <sec id="sec3dot1">
        <title>3.1. Model Performance</title>
        <p>The XGBoost regression model demonstrated strong predictive capabilities in estimating survey response rates. On the test dataset, the model achieved an R<sup>2</sup> score of 0.85, indicating that approximately 85% of the variance in survey response rates could be explained by the selected features. Additionally, the mean squared error (MSE) of 0.02 reflects a low average deviation between predicted and actual response rates ([<xref ref-type="bibr" rid="B9">9</xref>]; [<xref ref-type="bibr" rid="B12">12</xref>]; [<xref ref-type="bibr" rid="B11">11</xref>]). Together, these metrics suggest that the model is both accurate and reliable, effectively capturing the underlying patterns in survey participation across different organizational and demographic contexts.</p>
      </sec>
      <sec id="sec3dot2">
        <title>3.2. Feature Importance</title>
        <p>To identify the factors most strongly influencing survey response behavior, feature importance analysis was conducted using XGBoost’s built-in feature importance function. The results indicated that “sent_forms” (0.523) and “Reminder” (0.311) were the most influential variables, underscoring the critical role of survey distribution practices and follow-up reminders in driving participation ([<xref ref-type="bibr" rid="B6">6</xref>]; [<xref ref-type="bibr" rid="B17">17</xref>]). Other features, including “allowed_response_window_hours” (0.038) and “weekday” (0.031), exhibited moderate contributions, highlighting the impact of response timing and survey scheduling. Demographic characteristics, such as age (0.013), gender (0.027), and job level (0.030), as well as send_hour (0.027), contributed less but were still non-negligible, suggesting that individual-level differences can modulate response likelihood.</p>
        <p>These findings provide actionable insights for organizations, emphasizing that strategic planning of survey distribution, reminder frequency, and timing can meaningfully improve participation rates, while demographic and hierarchical factors should also be considered when designing surveys.</p>
      </sec>
      <sec id="sec3dot3">
        <title>3.3. Visualizations</title>
        <p>3.3.1. Feature Importance Plot</p>
        <p>A bar chart (<xref ref-type="fig" rid="fig1">Figure 1</xref>) was generated to visually represent the relative contributions of each feature to the model’s predictions ([<xref ref-type="bibr" rid="B1">1</xref>]; [<xref ref-type="bibr" rid="B2">2</xref>]). The plot clearly shows the dominance of sent_forms and Reminder, followed by temporal attributes and demographic variables. By presenting the data graphically, the figure facilitates comparative analysis across variables, allowing readers to quickly identify which factors are most critical in shaping survey response behavior. This visualization underscores the multi-faceted nature of participation determinants, where organizational practices, timing, and individual characteristics interact.</p>
        <p>3.3.2. Actual vs. Predicted Response Rates</p>
        <p>To assess the model’s predictive accuracy, a scatter plot (<xref ref-type="fig" rid="fig2">Figure 2</xref>) comparing actual versus predicted response rates was produced. Most points are closely aligned along the diagonal line, indicating that the model accurately captured the underlying patterns of survey participation. This alignment confirms the robustness and reliability of the feature importance findings, as the model’s predictions closely match observed responses. The scatter plot also serves as a visual validation of the model, demonstrating that the selected features and their relative importance are sufficient to explain the majority of variation in survey responses ([<xref ref-type="bibr" rid="B3">3</xref>]; [<xref ref-type="bibr" rid="B8">8</xref>]).</p>
      </sec>
    </sec>
    <sec id="sec4">
      <title>4. Discussion</title>
      <p>The results of the current study highlight the predominance of organizational factors in determining survey participation. In particular, the feature importance analysis indicates that <bold>sent_forms</bold> (Importance = 0.521) and <bold>Reminder</bold> (Importance = 0.332) were the most influential predictors of response rates, suggesting that structural and procedural elements play a greater role than demographic factors in influencing participation. Other variables, including <bold>weekday</bold>, <bold>job_level</bold>, <bold>allowed_response_window_hours</bold>, and <bold>send_hour</bold>, had smaller contributions, while demographic variables such as <bold>Age</bold> and <bold>Gender</bold> were the least influential (Importance &lt; 0.01). This implies that organizations seeking to improve survey engagement should focus primarily on outreach strategies, follow-up reminders, and survey timing rather than demographic targeting ([<xref ref-type="bibr" rid="B8">8</xref>]; [<xref ref-type="bibr" rid="B5">5</xref>]; [<xref ref-type="bibr" rid="B10">10</xref>]; [<xref ref-type="bibr" rid="B13">13</xref>]).</p>
      <p>From a predictive standpoint, the model performed well in capturing actual response rates, with predicted values closely aligning with observed rates across multiple survey batches. This demonstrates the utility of the XGBoost model for forecasting participation and understanding key organizational determinants ([<xref ref-type="bibr" rid="B11">11</xref>]; [<xref ref-type="bibr" rid="B8">8</xref>]).</p>
      <p>Future research could focus on retraining the model with organization-specific data to better capture unique patterns in survey participation ([<xref ref-type="bibr" rid="B8">8</xref>]; [<xref ref-type="bibr" rid="B5">5</xref>]). Validation of predictions should accompany any retraining process to ensure the forecasted participation rates are realistic and actionable ([<xref ref-type="bibr" rid="B11">11</xref>]). Additionally, efforts could be made to improve prediction accuracy and reduce errors by considering the distinct characteristics and dynamics of each organization that may influence employee response behaviour ([<xref ref-type="bibr" rid="B10">10</xref>]; [<xref ref-type="bibr" rid="B13">13</xref>]).</p>
      <p>From a practical perspective, the model enables organizations to input their own survey data and forecast expected participation rates for future surveys ([<xref ref-type="bibr" rid="B8">8</xref>]; [<xref ref-type="bibr" rid="B10">10</xref>]). Scenario simulations can also be conducted, allowing organizations to evaluate the potential impact of changes, such as adjusting the number of reminders or modifying the allowed response window, on participation rates ([<xref ref-type="bibr" rid="B5">5</xref>]). Implementing a continuous cycle of assessment and optimization will ensure that the model provides reliable and actionable insights, supporting informed decision-making in survey design, administration, and follow-up strategies ([<xref ref-type="bibr" rid="B8">8</xref>]; [<xref ref-type="bibr" rid="B5">5</xref>]; [<xref ref-type="bibr" rid="B10">10</xref>]).</p>
      <p>Overall, these findings underscore the critical role of structural and procedural organizational factors in survey participation, highlighting actionable levers that organizations can adjust to improve employee engagement with surveys ([<xref ref-type="bibr" rid="B8">8</xref>]; [<xref ref-type="bibr" rid="B5">5</xref>]; [<xref ref-type="bibr" rid="B10">10</xref>]).</p>
    </sec>
    <sec id="sec5">
      <title>5. Limitations</title>
      <p>This study has several limitations that should be considered when interpreting the findings. First, the dataset was derived from a single consulting firm working with multiple client organizations, which may limit the generalizability of the results across different industries and organizational contexts ([<xref ref-type="bibr" rid="B4">4</xref>]; [<xref ref-type="bibr" rid="B9">9</xref>]). Second, the XGBoost model was trained using default hyperparameters; although this approach preserved interpretability and focused on feature relationships, future research could explore parameter tuning to further enhance predictive accuracy ([<xref ref-type="bibr" rid="B12">12</xref>]; [<xref ref-type="bibr" rid="B7">7</xref>]). Finally, additional contextual factors, such as organizational culture, departmental workload, Survey fatigue, or incentive, were not included in the current model. Incorporating these variables could provide a more nuanced understanding of participation behavior and improve the model’s predictive capability ([<xref ref-type="bibr" rid="B17">17</xref>]; [<xref ref-type="bibr" rid="B1">1</xref>]).</p>
    </sec>
    <sec id="sec6">
      <title>6. Conclusion</title>
      <p>This study highlights the effectiveness of machine learning techniques, specifically the XGBoost algorithm, in predicting survey response rates within organizational contexts ([<xref ref-type="bibr" rid="B4">4</xref>]; [<xref ref-type="bibr" rid="B7">7</xref>]; [<xref ref-type="bibr" rid="B17">17</xref>]). By applying this model, we were able to identify key factors influencing survey participation, including organizational practices, temporal attributes, and individual demographic characteristics ([<xref ref-type="bibr" rid="B9">9</xref>]; [<xref ref-type="bibr" rid="B6">6</xref>]; [<xref ref-type="bibr" rid="B2">2</xref>]).</p>
      <p>The findings demonstrate that variables such as the number of sent forms, the number of reminders, and the allowed response window play a central role in shaping response behaviour ([<xref ref-type="bibr" rid="B12">12</xref>]; [<xref ref-type="bibr" rid="B7">7</xref>]; [<xref ref-type="bibr" rid="B1">1</xref>]). Leveraging these insights allows organizations to optimize survey design, target the most relevant respondent segments, and implement evidence-based strategies to improve participation rates ([<xref ref-type="bibr" rid="B17">17</xref>]; [<xref ref-type="bibr" rid="B2">2</xref>]). Furthermore, the integration of predictive analytics contributes to the overall quality, reliability, and representativeness of survey data, supporting more informed and data-driven decision-making processes ([<xref ref-type="bibr" rid="B9">9</xref>]; [<xref ref-type="bibr" rid="B2">2</xref>]; [<xref ref-type="bibr" rid="B3">3</xref>]).</p>
      <p>Overall, the results underscore the potential of machine learning not only as a predictive tool but also as a strategic enabler for enhancing organizational research and management practices ([<xref ref-type="bibr" rid="B4">4</xref>]; [<xref ref-type="bibr" rid="B7">7</xref>]). Future studies could further extend this work by incorporating additional variables, exploring alternative modelling approaches, or developing user-friendly interfaces for broader practical deployment ([<xref ref-type="bibr" rid="B8">8</xref>]; [<xref ref-type="bibr" rid="B5">5</xref>]; [<xref ref-type="bibr" rid="B10">10</xref>]).</p>
    </sec>
    <sec id="sec7">
      <title>Acknowledgements</title>
      <p>The authors acknowledge Cultural Infusion Pty Ltd and the Diversity Atlas research team for their support, datasets, and collaborative insights. Moreover, the authors would like to thank Peter Mousaferiadis, Michael Walmsley, Nicole Lee, and Mary Legrand for their valuable assistance and contributions to this research. While their insights and support were greatly appreciated, the ideas and interpretations presented in this study remain those of the authors and may not fully reflect the perspectives of the acknowledged individuals.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Amirshahi, A., Kirsch, N., Reymond, J., &amp; Baghersalimi, S. (2023). Predicting Survey Response with Quotation-Based Modeling: A Case Study on Favorability Towards the United States. <italic>2023 10th IEEE Swiss Conference on Data Science (SDS)</italic> (pp. 1-8). IEEE. https://doi.org/10.1109/sds57534.2023.00008 <pub-id pub-id-type="doi">10.1109/sds57534.2023.00008</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/sds57534.2023.00008">https://doi.org/10.1109/sds57534.2023.00008</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Amirshahi, A.</string-name>
              <string-name>Kirsch, N.</string-name>
              <string-name>Reymond, J.</string-name>
              <string-name>Baghersalimi, S.</string-name>
            </person-group>
            <year>2023</year>
            <pub-id pub-id-type="doi">10.1109/sds57534.2023.00008</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Barkho, W., Carnes, N. C., Kolaja, C. A., Tu, X. M., Boparai, S. K., Castañeda, S. F. et al. (2024). Utilizing Machine Learning to Predict Participant Response to Follow-Up Health Surveys in the Millennium Cohort Study. <italic>Scientific Reports, 14,</italic> Article No. 25764. https://doi.org/10.1038/s41598-024-77563-8 <pub-id pub-id-type="doi">10.1038/s41598-024-77563-8</pub-id><pub-id pub-id-type="pmid">39468293</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41598-024-77563-8">https://doi.org/10.1038/s41598-024-77563-8</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Barkho, W.</string-name>
              <string-name>Carnes, N.</string-name>
              <string-name>Kolaja, C.</string-name>
              <string-name>Tu, X.</string-name>
              <string-name>Boparai, S.</string-name>
            </person-group>
            <year>2024</year>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1038/s41598-024-77563-8</pub-id>
            <pub-id pub-id-type="pmid">39468293</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Ben-David, E., Ibrahim, S., Mazumder, R., &amp; Radchenko, P. (2021). Predicting Census Survey Response Rates via Additive Regression with Interactions. <italic>Annals of Applied Statistics, 19,</italic>1-28.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Ben-David, E.</string-name>
              <string-name>Ibrahim, S.</string-name>
              <string-name>Mazumder, R.</string-name>
              <string-name>Radchenko, P.</string-name>
            </person-group>
            <year>2021</year>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Chen, T., &amp; Guestrin, C. (2016). XGBoost: A Scalable Tree Boosting System. <italic>Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</italic> (pp. 785-794). ACM. https://doi.org/10.1145/2939672.2939785 <pub-id pub-id-type="doi">10.1145/2939672.2939785</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/2939672.2939785">https://doi.org/10.1145/2939672.2939785</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Chen, T.</string-name>
              <string-name>Guestrin, C.</string-name>
            </person-group>
            <year>2016</year>
            <pub-id pub-id-type="doi">10.1145/2939672.2939785</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">Dinov, I. D. (2023). Model Performance Assessment, Validation, and Improvement. In editor (Ed.), <italic>The Springer Series in Applied Machine Learning</italic> (Vol. 2853, pp. 477-531). Springer International Publishing. https://doi.org/10.1007/978-3-031-17483-4_9 <pub-id pub-id-type="doi">10.1007/978-3-031-17483-4_9</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/978-3-031-17483-4_9">https://doi.org/10.1007/978-3-031-17483-4_9</ext-link></mixed-citation>
          <element-citation publication-type="book">
            <person-group person-group-type="author">
              <string-name>Dinov, I.</string-name>
              <string-name>Assessment, V</string-name>
            </person-group>
            <year>2023</year>
            <pub-id pub-id-type="doi">10.1007/978-3-031-17483-4_9</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B6">
        <label>6.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Friedman, J. H. (2001). Greedy Function Approximation: A Gradient Boosting Machine. <italic>The Annals of Statistics, 29,</italic> 1189-1232. https://doi.org/10.1214/aos/1013203451 <pub-id pub-id-type="doi">10.1214/aos/1013203451</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1214/aos/1013203451">https://doi.org/10.1214/aos/1013203451</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Friedman, J.</string-name>
            </person-group>
            <year>2001</year>
            <pub-id pub-id-type="doi">10.1214/aos/1013203451</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B7">
        <label>7.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Groves, R. M., &amp; Peytcheva, E. (2008). The Impact of Nonresponse Rates on Nonresponse Bias: A Meta-Analysis. <italic>Public Opinion Quarterly, 72,</italic> 167-189. https://doi.org/10.1093/poq/nfn011 <pub-id pub-id-type="doi">10.1093/poq/nfn011</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1093/poq/nfn011">https://doi.org/10.1093/poq/nfn011</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Groves, R.</string-name>
              <string-name>Peytcheva, E.</string-name>
            </person-group>
            <year>2008</year>
            <pub-id pub-id-type="doi">10.1093/poq/nfn011</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B8">
        <label>8.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Hastie, T., Tibshirani, R., &amp; Friedman, J. (2009). <italic>The Elements of Statistical Learning.</italic> Springer.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Hastie, T.</string-name>
              <string-name>Tibshirani, R.</string-name>
              <string-name>Friedman, J.</string-name>
            </person-group>
            <year>2009</year>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B9">
        <label>9.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Ibrahim, S., Mazumder, R., Radchenko, P., &amp; Ben-David, E. (2021). <italic>Predicting Census Survey Response Rates with Interpretable Nonparametric Additive Models</italic><italic>and Structured Interactions.</italic> arXiv:2108.11328v3. https://arxiv.org/abs/2108.11328v3</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Ibrahim, S.</string-name>
              <string-name>Mazumder, R.</string-name>
              <string-name>Radchenko, P.</string-name>
              <string-name>Ben-David, E.</string-name>
            </person-group>
            <year>2021</year>
            <fpage>2108</fpage>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B10">
        <label>10.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Kern, C., Weiss, B., &amp; Kolb, J.-P. (2019). <italic>A Longitudinal Framework for Predicting Nonresponse in Panel Surveys.</italic>arXiv:1909.13361. https://arxiv.org/abs/1909.13361</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Kern, C.</string-name>
              <string-name>Weiss, B.</string-name>
              <string-name>Kolb, J.</string-name>
            </person-group>
            <year>2019</year>
            <fpage>1909</fpage>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B11">
        <label>11.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Kern, M. et al. (2024). Calibration and XGBoost Reweighting to Reduce Coverage and Non-Response Biases in Overlapping Panel Surveys: Application to the Healthcare and Social Survey. <italic>BMC Medical Research Methodology, 24,</italic> Article No. 36. https://doi.org/10.1186/s12874-024-02171-z <pub-id pub-id-type="doi">10.1186/s12874-024-02171-z</pub-id><pub-id pub-id-type="pmid">38360543</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1186/s12874-024-02171-z">https://doi.org/10.1186/s12874-024-02171-z</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Kern, M.</string-name>
            </person-group>
            <year>2024</year>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1186/s12874-024-02171-z</pub-id>
            <pub-id pub-id-type="pmid">38360543</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B12">
        <label>12.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Mazumder, R., Ben-David, E., &amp; Ibrahim, S. (2021). <italic>Predicting Census Survey Response Rates via Interpretable Nonparametric Additive Models with Structured Interactions.</italic> arXiv:2108.11328v2. https://arxiv.org/pdf/2108.11328v2</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Mazumder, R.</string-name>
              <string-name>Ben-David, E.</string-name>
              <string-name>Ibrahim, S.</string-name>
            </person-group>
            <year>2021</year>
            <fpage>2108</fpage>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B13">
        <label>13.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Moeini, R., &amp; Cultural Infusion Research Team (2023). <italic>Cultural Diversity Measurement through Diversity Atlas: A Case Study Approach.</italic>Cultural Infusion White Paper.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Moeini, R.</string-name>
            </person-group>
            <year>2023</year>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B14">
        <label>14.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Moieni, R., &amp; Mousaferiadis, P. (2022). Analysis of Cultural Diversity Concept in Different Countries Using Fractal Analysis. <italic>The International Journal of Organizational Diversity, 22,</italic> 43-62. https://search.proquest.com/openview/2e4c42e8af84f2c0a56d02dac5e0d983/1?pq-origsite=gscholar&amp;cbl=5529398</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Moieni, R.</string-name>
              <string-name>Mousaferiadis, P.</string-name>
            </person-group>
            <year>2022</year>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B15">
        <label>15.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Moeini, R., Mousaferiadis, P., &amp; Pateel, P. (2022). <italic>An Analytical Approach to Measure the Cultural Diversity Mutuality between Two Communities.</italic> NeuroQuantology.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Moeini, R.</string-name>
              <string-name>Mousaferiadis, P.</string-name>
              <string-name>Pateel, P.</string-name>
            </person-group>
            <year>2022</year>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B16">
        <label>16.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Xue, J. (2024). Optimization of Big Data Analysis Resources Supported by XGBoost and LSTM. <italic>Journal of Big Data Analytics, 3,</italic> 45-58.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Xue, J.</string-name>
            </person-group>
            <year>2024</year>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B17">
        <label>17.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Zhang, Y., &amp; Zheng, Y. (2023). Estimating Response Propensities in Nonprobability Surveys Using Machine Learning. <italic>Journal of Survey Statistics and Methodology, 11,</italic> 123-145.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Zhang, Y.</string-name>
              <string-name>Zheng, Y.</string-name>
            </person-group>
            <year>2023</year>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>