<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">ce</journal-id>
      <journal-title-group>
        <journal-title>Creative Education</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2151-4771</issn>
      <issn pub-type="ppub">2151-4755</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/ce.2026.178091</article-id>
      <article-id pub-id-type="publisher-id">ce-153395</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Social Sciences</subject>
          <subject>Humanities</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>AI-DRS: A Dialogic Generative-AI Instructional Model for Scientific Reasoning in Primary Education</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes">
          <contrib-id contrib-id-type="orcid">0000-0002-9124-2245</contrib-id>
          <name name-style="western">
            <surname>Kalogiannakis</surname>
            <given-names>Michail</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <contrib-id contrib-id-type="orcid">0009-0004-9960-9488</contrib-id>
          <name name-style="western">
            <surname>Spasopoulos</surname>
            <given-names>Theodoros</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <contrib-id contrib-id-type="orcid">0009-0006-0266-0098</contrib-id>
          <name name-style="western">
            <surname>Papakonstantinou</surname>
            <given-names>Nikos</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <contrib-id contrib-id-type="orcid">0000-0002-7300-7380</contrib-id>
          <name name-style="western">
            <surname>Xenakis</surname>
            <given-names>Apostolos</given-names>
          </name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> Department of Special Education, University of Thessaly, Volos, Greece </aff>
      <aff id="aff2"><label>2</label> Department of Digital Systems, University of Thessaly, Larissa, Greece </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The authors declare no conflicts of interest regarding the publication of this paper.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>06</day>
        <month>08</month>
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>08</month>
        <year>2026</year>
      </pub-date>
      <volume>17</volume>
      <issue>08</issue>
      <fpage>1556</fpage>
      <lpage>1586</lpage>
      <history>
        <date date-type="received">
          <day>24</day>
          <month>06</month>
          <year>2026</year>
        </date>
        <date date-type="accepted">
          <day>22</day>
          <month>08</month>
          <year>2026</year>
        </date>
        <date date-type="published">
          <day>25</day>
          <month>08</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2026 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/ce.2026.178091">https://doi.org/10.4236/ce.2026.178091</self-uri>
      <abstract>
        <p>Cultivating scientific reasoning in primary education involves not just making measurements, but also turning them into evidence and explanations. Physical-computing activities often do not support this process, and using generative artificial intelligence (GenAI) as a tool for answers can skip this step. This article suggests a teaching method for primary education that limits GenAI’s role so it can support pupils’ reasoning and then step back. In our research-and-development design, we created the AI-DRS (AI-supported Dialogic Reasoning Scaffolding) Physical Inquiry Model. This model pairs a Nezha-micro:bit soil-moisture investigation with sensor data generated by pupils. It follows a Claim-Evidence-Reasoning framework and includes a dialogic control layer based on three principles: grounding in data, ensuring proper epistemic oversight, and being responsive to pupils. The model incorporates fading based on established criteria, teacher-led interventions, and specific protections for each child. The article specifies the scientific target of the investigation, the sensor-calibration and data-quality rules, the assignment and balancing of inquiry variables, the technical configuration and rule precedence of the system, a teacher escalation protocol with explicit stop rules, and a cluster-aware analysis plan. Next steps include expert validation and a controlled pilot study. The model has a seven-phase process, an architecture that responds to pupil input, a scoring system, and a coding scheme that is sensitive to how pupils engage. It separates the tool’s dialogic moves from pupil responses, allowing for individual analysis. This proposal is theoretical, with no empirical results presented. It establishes a structured approach where GenAI can ask follow-up questions while the teacher maintains control over the knowledge. Validating the model’s effectiveness and understanding how it works is the essential next step.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>Generative Artificial Intelligence</kwd>
        <kwd>Micro:bit</kwd>
        <kwd>Claim-Evidence-Reasoning</kwd>
        <kwd>Scientific Reasoning</kwd>
        <kwd>Primary Education</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction</title>
      <sec id="sec1dot1">
        <title>1.1. Background of the Study</title>
        <p>Scientific literacy in primary education rests on a skill that is simple to describe but hard to teach: reasoning from evidence. Children must distinguish observation from interpretation, make testable claims, and select data that support them. The most demanding step is explaining why that data supports the claim through a scientific mechanism. In primary classrooms, pupils often reach the conclusion but struggle with the explicit mechanistic link between evidence and claim ([<xref ref-type="bibr" rid="B9">9</xref>]; [<xref ref-type="bibr" rid="B19">19</xref>]).</p>
        <p>Physical computing eases this challenge. Sensors let children build systems that interact with the physical world and return continuous real-time measurements, turning abstract phenomena into observable data ([<xref ref-type="bibr" rid="B25">25</xref>]; [<xref ref-type="bibr" rid="B28">28</xref>]; [<xref ref-type="bibr" rid="B7">7</xref>]). Reliable data, however, do not produce interpretation on their own: the critical step is moving from recording a value to reading a pattern as evidence of a mechanism, and without support, pupils find this hard ([<xref ref-type="bibr" rid="B16">16</xref>]). GenAI promises personalized support, timely feedback, and new forms of assessment ([<xref ref-type="bibr" rid="B42">42</xref>]), but used as an answer provider it performs the reasoning the learner should do, encouraging over-trust and weakening independent thinking ([<xref ref-type="bibr" rid="B8">8</xref>]; [<xref ref-type="bibr" rid="B15">15</xref>]). The question for primary science is therefore not whether GenAI will be present, but which instructional design can make it support reasoning rather than replace it.</p>
      </sec>
      <sec id="sec1dot2">
        <title>1.2. Problem of the Study</title>
        <p>Research on the micro:bit has focused mainly on technical features, motivational benefits, and approaches to teaching programming, and less on whether such activities support scientific thinking or conceptual understanding ([<xref ref-type="bibr" rid="B30">30</xref>]; [<xref ref-type="bibr" rid="B12">12</xref>]). Completing a physical-computing task does not demonstrate understanding of the phenomenon, the variables, or the link between measurement and scientific model: a device that works is not a pupil who thinks. Similarly, many educational chatbots ask generic questions and give feedback without requiring pupils to draw conclusions from their own data, and without criteria for when to advance, persist, hint, or withdraw.</p>
        <p>Two gaps emerge. Physical computing generates rich data but does not ensure scientific reasoning; GenAI generates rich dialogue but can easily detach from the learner’s data and substitute for the learner’s own thinking. This study treats the two gaps as complementary and addressable through a single intervention, provided the dialogue stays tied to pupils’ own measurements and follows explicit epistemic and developmental rules.</p>
      </sec>
      <sec id="sec1dot3">
        <title>1.3. State of the Art</title>
        <p>Dialogic teaching views learning as a process in which ideas are shared, examined, and progressively refined through mutual interaction ([<xref ref-type="bibr" rid="B1">1</xref>]; [<xref ref-type="bibr" rid="B23">23</xref>]; [<xref ref-type="bibr" rid="B35">35</xref>]). Unlike static prompts, generative artificial intelligence (GenAI) dialogue systems can adapt to how learners are thinking, which makes them a promising support tool when they advance inquiry through questions rather than immediate answers ([<xref ref-type="bibr" rid="B14">14</xref>]; [<xref ref-type="bibr" rid="B31">31</xref>]; [<xref ref-type="bibr" rid="B32">32</xref>]; [<xref ref-type="bibr" rid="B33">33</xref>]). LLM-based platforms have also been developed to help educators design computational-thinking, AI, and STEM activities, illustrating the wider potential of generative systems as pedagogical infrastructures ([<xref ref-type="bibr" rid="B40">40</xref>]).</p>
        <p>Customized chatbots can support reasoning and argumentation in secondary science ([<xref ref-type="bibr" rid="B33">33</xref>]), and recent theoretical work casts GenAI as a conversational partner rather than a source of information ([<xref ref-type="bibr" rid="B37">37</xref>]). Neither specifies an implementation for primary-aged children: a developmentally appropriate method grounded in children’s own sensor data, with rules for when the system should question, prompt, or support, and when that support should be withdrawn. Scaffolding theory supplies the missing element. A scaffold is the support that enables a learner to solve a problem beyond their current ability ([<xref ref-type="bibr" rid="B38">38</xref>]), and its defining features are contingency, fading, and the transfer of responsibility to the learner ([<xref ref-type="bibr" rid="B34">34</xref>]); permanent mediation is not effective support, even when the immediate answer is correct. Work on scientific explanation stresses reducing written aids as skill develops ([<xref ref-type="bibr" rid="B17">17</xref>]; [<xref ref-type="bibr" rid="B19">19</xref>]), and studies of AI-supported scaffolding emphasize continuous diagnosis and progressive withdrawal rather than answer provision ([<xref ref-type="bibr" rid="B3">3</xref>]; [<xref ref-type="bibr" rid="B5">5</xref>]; [<xref ref-type="bibr" rid="B36">36</xref>]).</p>
        <p>Three further findings clarify the developmental stakes. Young children may place considerable and largely uncritical trust in conversational agents ([<xref ref-type="bibr" rid="B22">22</xref>]). AI literacy must be deliberately cultivated rather than assumed to emerge from technology use ([<xref ref-type="bibr" rid="B24">24</xref>]). And chatbot effects are, meta-analytically, smaller for younger learners, with novelty-related gains that may fade ([<xref ref-type="bibr" rid="B39">39</xref>]). Syntheses of GenAI-supported learning agree: the technology’s educational potential cannot be separated from its effects on cognition, metacognition, and learner agency ([<xref ref-type="bibr" rid="B42">42</xref>]).</p>
        <p>Consequently, general-purpose chatbot designs cannot simply be transferred to primary education; they require developmentally appropriate constraints, scaffolding, and safeguards.</p>
      </sec>
      <sec id="sec1dot4">
        <title>1.4. Research Gap and Objective</title>
        <p>The closest existing demonstration, a customized chatbot supporting reasoning and argumentation in secondary science ([<xref ref-type="bibr" rid="B33">33</xref>]), leaves three issues unaddressed for primary education: dialogue grounded in pupils’ own sensor data rather than curricular text, explicit criteria for fading support, and safeguards for children who may over-trust conversational agents. Effectiveness is defined here not by whether a pupil completes a conversation but by whether the pupil can subsequently produce an independent Claim-Evidence-Reasoning (CER) explanation ([<xref ref-type="bibr" rid="B18">18</xref>]) without AI. The aim is to develop the AI-DRS (AI-supported Dialogic Reasoning Scaffolding) Physical Inquiry Model and prepare it for validation, treating the design itself as the research contribution.</p>
        <p>The study is guided by a single design question: which design principles and architecture can connect GenAI dialogue to primary pupils’ own sensor data so that it supports, and then deliberately fades, independent evidence-based reasoning? The evaluative questions that follow concern whether the model improves independent CER performance relative to an equivalent static scaffold, how it affects pupils’ use of their own data and the trajectory of scaffolding, whether pupils articulate synthesis themselves, whether reasoning transfers without AI, and how perceived usefulness and calibrated trust relate to reasoning quality. These form the protocol of the planned pilot and are not addressed empirically here.</p>
      </sec>
    </sec>
    <sec id="sec2">
      <title>2. Method</title>
      <sec id="sec2dot1">
        <title>2.1. Type and Design</title>
        <p>The study uses a Research and Development (R&amp;D) design with two sequential parts: (a) developing the intervention model and (b) planning its expert validation and subsequent pilot evaluation. This article reports only the first part; the second is set out as a protocol for future work. The study is deliberately not presented as an efficacy trial: the planned pilot examines feasibility, implementation fidelity, preliminary indications of effect, and above all, the links among scaffolding, data use, and reasoning.</p>
        <p>Development is organized into seven R&amp;D stages: </p>
        <p>1) needs analysis through curriculum documents, research literature, and the authors’ teaching experience, with no participant data; </p>
        <p>2) mapping of curriculum, target concepts, and likely misconceptions; </p>
        <p>3) development of the inquiry task, the AI-DRS prompt, the worksheet, and the teacher guide; </p>
        <p>4) expert validation;</p>
        <p>5) a small-group technical and usability test; </p>
        <p>6) revision of the package; and </p>
        <p>7) pilot implementation with process evaluation. </p>
        <p>The first three stages are complete, and their output, the model in Section 3, constitutes this article’s contribution; the remainder are reported as protocol, and Sections 2.2 - 2.4 set it out in full for replicability, describing planned rather than collected data. Because the intervention involves GenAI use with children, it assumes a school-managed environment, data minimization, parental consent and pupil assent, pseudonymized chat logs, defined retention rules, and continuous teacher oversight. The system states that it can be wrong, asks pupils to check its suggestions against their own observations, and routes uncertain or problematic cases to the teacher. The planned pilot will be submitted to the Research Ethics Committee of the University of Thessaly before any data collection ([<xref ref-type="bibr" rid="B27">27</xref>]).</p>
        <p>The design must also address how children may lawfully interact with GenAI. Mainstream consumer chatbots are not permitted for this age group: OpenAI’s terms require users to be at least 13, with parental consent for minors ([<xref ref-type="bibr" rid="B26">26</xref>]), and Anthropic’s require users to be at least 18 ([<xref ref-type="bibr" rid="B2">2</xref>]). The model, therefore, assumes mediated access. Pupils hold no accounts and use no consumer interface; the dialogue runs through a school-controlled application operating under the teacher’s institutional account and the selected provider’s application programming interface (API). That application logs pseudonymized interaction data inside the school environment and transmits only data-minimized, non-identifying inputs through the API, under approved data-processing, security, and retention arrangements.</p>
      </sec>
      <sec id="sec2dot2">
        <title>2.2. Data and Data Sources</title>
        <p>A sample of about 40 to 60 Grade 5 or Grade 6 pupils is proposed, with 20 to 30 per condition. Because pupils build and measure in groups of three to four and groups are the unit of assignment, this corresponds to approximately five to ten groups per condition, a figure that governs the analysis plan in Section 2.4. Pupils work in groups during construction, but the pre-test, post-test, and transfer measures are administered individually. The planned data sources are the individual pre- and post-CER tasks, an individual transfer task completed without AI, pseudonymized chat logs, teacher observation logs, technical and fidelity records, a short perceived-usefulness and calibrated-trust questionnaire, and the expert-validation ratings collected before the pilot.</p>
        <p>The intervention is built on the Nezha Inventor’s Kit for micro:bit (<xref ref-type="fig" rid="fig1">Figure 1</xref><bold>)</bold>. Because the assessed comparison requires both soil conditions to be set up and measured simultaneously within each group (Section 2.2.3), every group is equipped with two complete measurement channels, one per container: two BBC micro:bit units, two soil-moisture probes, and LED indicators. This exceeds the standard kit, which contains a single soil-moisture sensor, so a second kit, or a second micro:bit and probe, is required per group; this is stated explicitly as a resource condition of the design rather than assumed. </p>
        <p>The materials support a Smart Plant Investigation in which pupils build a soil-moisture monitoring system and examine how soil type affects the change in soil moisture over time. Every group investigates the same focal comparison under the same controlled conditions, so that datasets and reasoning demands are equivalent across groups and conditions. After calibration and a stability check on each channel, pupils record the container, elapsed time, calibrated readings, quality flags, and their observations.</p>
        <fig id="fig1">
          <label>Figure 1</label>
          <graphic xlink:href="https://html.scirp.org/file/6309660-rId16.jpeg?20260825041905" />
        </fig>
        <p><bold>Figure 1.</bold>Components of the Nezha Inventor’s Kit used in the proposed Smart Plant Investigation. Source: Authors’ photograph.</p>
        <p>The soil-moisture probe is the kit’s own sensor, which keeps the build within the kit’s wiring conventions and MakeCode blocks; components such as a BME280 sensor or an OLED display are not part of the standard configuration ([<xref ref-type="bibr" rid="B10">10</xref>]). Duplicating the micro:bit and the probe, as described above, is the only departure from the standard configuration.</p>
        <p>2.2.1. Inquiry Question, Target Mechanism, and Permissible Causal Claim</p>
        <p>Claim-Evidence-Reasoning (CER) responses can be scored consistently only if the investigation has one question, one expected mechanism, and one clearly bounded causal claim that applies to every group. The Smart Plant Investigation is therefore built around a single focal comparison rather than group-selected variables. The operational inquiry question is: how does soil type (sandy soil versus potting compost) affect the rate at which soil-moisture readings decrease over 40 minutes, when the initial water volume, container, soil mass, probe placement, bench position, light exposure, and room temperature are held constant?</p>
        <p>The expected empirical pattern is a steeper decline of the calibrated Moisture Index in sandy soil than in compost. The target mechanism, expressed at a level appropriate for Grade 5-6, is that sandy soil consists of larger particles with larger pore spaces and little organic matter, so water drains through it faster and is exposed to the air across a larger internal surface, whereas the finer particles and organic matter in compost hold water against drainage and evaporation. A full-credit reasoning statement links a stated numerical difference in the rate of decline to this particle-size and pore-space account; a partial-credit statement identifies the difference but restates the data instead of explaining it, as in “it dried faster because the numbers went down faster”.</p>
        <p>The permissible causal claim is deliberately bounded. Because soil type was manipulated while the remaining conditions were held constant, pupils may claim that, in this investigation and within the 40-minute window, the type of soil caused the difference in the rate of moisture loss. They may not claim which soil is better for growing plants, extend the claim to plant health or growth, read the index as an absolute quantity of water, or make causal claims about light, temperature, or container size. <bold>Table 1</bold> states this target explicitly so that raters, the teacher, and the AI-DRS system apply the same scientific standard: it is supplied to the system as a fixed reference in every dialogue turn (Section 3.4) and it governs the Mechanism dimension of the rubric (Section 3.7).</p>
        <p><bold>Table 1.</bold> Scientific target of the smart plant investigation.</p>
        <table-wrap id="tbl1">
          <label>Table 1</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Element</bold>
                </td>
                <td>
                  <bold>Specification</bold>
                </td>
              </tr>
              <tr>
                <td>Operational inquiry question</td>
                <td>How does soil type (sandy soil vs. potting compost) affect the rate at which soil-moisture readings decrease over 40 minutes, when all other conditions are held constant?</td>
              </tr>
              <tr>
                <td>Manipulated variable</td>
                <td>Soil type, at two levels: sandy soil and potting compost. Both levels are set up and measured simultaneously within every group, on two separate probes.</td>
              </tr>
              <tr>
                <td>Outcome measure</td>
                <td>Calibrated Moisture Index (MI, nominally 0 - 100; the index is unbounded and is logged unclipped, and only the plotted value is clipped to −5 to 105) recorded every 5 minutes for 40 minutes on each of the two containers; derived measures are the total decline (MI at t0 minus MI at t40) and the mean rate of decline (MI points per 10 minutes).</td>
              </tr>
              <tr>
                <td>Controlled variables</td>
                <td>Container type and volume, dry soil mass (±5 g), initial water volume (100 ml), probe depth and position, bench position and light exposure, room temperature, and time of day.</td>
              </tr>
              <tr>
                <td>Expected pattern</td>
                <td>A steeper decline in sandy soil than in compost across the measurement window.</td>
              </tr>
              <tr>
                <td>Target mechanism (pupil level)</td>
                <td>Sandy soil has larger particles and larger pore spaces and little organic matter, so water drains faster, and more of it is exposed to the air; the finer particles and organic matter in compost hold water against drainage and evaporation.</td>
              </tr>
              <tr>
                <td>Permissible causal claim</td>
                <td>“In this test, the type of soil caused the difference in how fast the soil lost water over 40 minutes”, supported by at least two specific readings from each container.</td>
              </tr>
              <tr>
                <td>Claims outside scope</td>
                <td>Which soil is better for plants; effects on growth or plant health; absolute water content; behavior beyond 40 minutes or after re-watering; causal claims about light, temperature, or container size.</td>
              </tr>
              <tr>
                <td>Anticipated misconceptions</td>
                <td>Reading the sensor value as an absolute amount of water; attributing the decline to the plant “drinking” it; treating a single reading as a trend; equating a faster decline with a “worse” soil.</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>2.2.2. Sensor Calibration, Measurement Schedule, and Data-Quality Rules</p>
        <p>The soil-moisture sensor returns an uncalibrated analog value, so the same physical condition can produce different readings across sensors, sessions, and probe placements. Interpretation is therefore preceded by a fixed two-point calibration and a stability check, and the dataset is screened with rules defined in advance rather than negotiated during the dialogue.</p>
        <p>Calibration is carried out for each of a group’s two probes at the start of every session and repeated whenever a probe is removed and re-inserted: a dry reference in air and a wet reference at the marked depth in tap water, each the mean of five readings after a settling period (<bold>Table 2</bold>). Readings are converted in MakeCode to a Moisture Index, MI = 100 × (R − R_dry)/(R_wet − R_dry). The index is unbounded: the unclipped value is logged and screened for quality, and only the plotted value is clipped to −5 to 105, so that an out-of-range reading remains detectable by the range rule. Each channel yields its own pair of calibration constants; both pairs are recorded on the group worksheet and stored with the dataset, labeled by container.</p>
        <p>The acceptance criteria and the action taken when they are not met are set out in <bold>Table 2</bold>: a reference is accepted only when the five readings span no more than 2% of full scale, and a probe that fails three attempts, or whose calibration span falls below 150 raw units, is replaced before the investigation proceeds. After the standardised watering at t = 0, MI is recorded on both channels every 5 minutes for 40 minutes, giving nine time points per container; each value is the median of three consecutive readings taken 5 seconds apart, which suppresses single-sample noise without requiring statistical treatment by pupils. Probes remain at a fixed depth of 5 cm and at least 2 cm from the container wall, marked with tape; any displacement is recorded as an event.</p>
        <p>The investigation is timetabled across three 45-minute lessons, set out in the final row of <bold>Table 2</bold>. Lesson 2 is the tightest, since a 40-minute window leaves no margin in a 45-minute slot; it is therefore timetabled as a double period where the school allows one, and otherwise calibration is completed at the end of lesson 1 so that lesson 2 begins with watering at t = 0. If neither is workable, the schedule is shortened to 30 minutes and seven time points, but only as a study-level decision taken for the whole pilot before data collection, never per group, so that all groups produce datasets of identical structure; the adopted schedule and its dependent thresholds are then reported. The 40-minute, nine-point schedule is the reference version used throughout this article. Lesson 2 deliberately does not carry the dialogue: about four minutes between readings allows graphing and the claim and data-grounding moves but not the full contingent exchange, so conceptual bridging, fading, and transfer occur in lesson 3 from a closed dataset.</p>
        <p><bold>Table 2.</bold> Calibration and measurement procedure.</p>
        <table-wrap id="tbl2">
          <label>Table 2</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Step</bold>
                </td>
                <td>
                  <bold>Procedure</bold>
                </td>
                <td>
                  <bold>Acceptance criterion</bold>
                </td>
                <td>
                  <bold>Action if not met</bold>
                </td>
              </tr>
              <tr>
                <td>Dry reference (R_dry)</td>
                <td>Each probe clean, dry, in air; 60 s settling; mean of five readings at 10 s intervals.</td>
                <td>Spread of the five readings ≤ 20 raw units (2% of full scale).</td>
                <td>Wait 60 s and repeat, up to three attempts; then replace the sensor.</td>
              </tr>
              <tr>
                <td>Wet reference (R_wet)</td>
                <td>Each probe inserted to the marked depth in tap water, electronics above the waterline; 60 s settling; mean of five readings.</td>
                <td>Spread ≤ 20 raw units.</td>
                <td>As above; a probe failing three attempts is withdrawn.</td>
              </tr>
              <tr>
                <td>Span check</td>
                <td>Compute |R_wet − R_dry| and record it on the worksheet.</td>
                <td>Span ≥ 150 raw units.</td>
                <td>Re-seat the probe and check the connector; if unchanged, replace the unit.</td>
              </tr>
              <tr>
                <td>Conversion</td>
                <td>MI = 100 × (R − R_dry) / (R_wet − R_dry), computed on each micro:bit. The unbounded index is logged; only the plotted value is clipped to −5 to 105.</td>
                <td>Both pairs of calibration constants logged with the dataset and labeled by container.</td>
                <td>The dataset is not interpreted until the constants are recorded.</td>
              </tr>
              <tr>
                <td>Measurement schedule</td>
                <td>Reading on both channels every 5 minutes for 40 minutes after watering at t = 0; each value is the median of three readings taken 5 s apart.</td>
                <td>Nine time points per container under the reference schedule, with no missing points.</td>
                <td>A missing point is re-taken at the next interval and marked as delayed.</td>
              </tr>
              <tr>
                <td>Probe placement</td>
                <td>Fixed depth of 5 cm, at least 2 cm from the container wall, position marked with tape.</td>
                <td>Probe not moved between t = 0 and t = 40.</td>
                <td>Displacement is logged as an event; subsequent readings are flagged.</td>
              </tr>
              <tr>
                <td>Session scheduling</td>
                <td>Three 45-minute lessons: (1) orientation, prediction, fair test, construction, programming, calibration; (2) re-calibration, watering at t = 0, the 40-minute window; (3) CER dialogue, fading, transfer task.</td>
                <td>Lesson 2 starts from completed calibration and ends with nine time points per container.</td>
                <td>A double period where available; a shortened 30-minute, seven-point schedule may be adopted for the whole pilot before data collection, never for single groups. An unfinished run is not interpreted; lesson 3 uses the class reference dataset.</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Five rules identify unreliable readings before any interpretation begins (<bold>Table 3</bold>). A single flagged point is re-taken once and, if the flag recurs, is excluded from the evidence base while remaining in the log. A run that fails the drift rule, or a container with more than two flagged points out of nine, is not interpreted at all: the system halts interpretation and refers the group to the teacher, which is the operational trigger for rule R8 in <bold>Table 8</bold>. The group then completes the CER task using a validated class reference dataset, so that a hardware fault does not cost the pupils the reasoning task. These definitions give the system’s unreliable-data category a concrete referent instead of leaving it to real-time judgment by a generative model.</p>
        <p><bold>Table 3.</bold> Rules for identifying unreliable readings.</p>
        <table-wrap id="tbl3">
          <label>Table 3</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Rule</bold>
                </td>
                <td>
                  <bold>Trigger</bold>
                </td>
                <td>
                  <bold>Interpretation</bold>
                </td>
                <td>
                  <bold>Consequence</bold>
                </td>
              </tr>
              <tr>
                <td>Range</td>
                <td>The unbounded index falls outside −5 to 105 MI, screened before the plotted value is clipped.</td>
                <td>Calibration or contact failure.</td>
                <td>Point discarded and re-taken once; a recurrence flags the container.</td>
              </tr>
              <tr>
                <td>Jump</td>
                <td>A change greater than 15 MI points between consecutive 5-minute points after t = 5, with no logged event.</td>
                <td>Probe movement or intermittent contact. The first interval (t = 0 to t = 5) is exempt because rapid drainage after watering is the phenomenon under study; the threshold is set per soil condition from the class reference dataset.</td>
                <td>Point re-taken once; if it recurs, the point is excluded from the evidence base.</td>
              </tr>
              <tr>
                <td>Flatline</td>
                <td>An identical raw value for four or more consecutive readings.</td>
                <td>Disconnection or a frozen input.</td>
                <td>Hardware check by the teacher before measurement resumes.</td>
              </tr>
              <tr>
                <td>Drift</td>
                <td>A post-session dry check differing from R_dry by more than 25 raw units, or by more than 5% of the calibration span, whichever is larger.</td>
                <td>Sensor drift across the session, judged against a floor that exceeds the calibration noise accepted above.</td>
                <td>The whole run is flagged unreliable and is not interpreted.</td>
              </tr>
              <tr>
                <td>Run level</td>
                <td>More than two flagged points in a container, out of the nine recorded under the reference schedule.</td>
                <td>The dataset cannot support an evidence claim.</td>
                <td>Interpretation halted, teacher check, and the class reference dataset is used for the CER task.</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>The class reference dataset is not improvised during the lesson. It is collected by the research team in a pre-pilot, with at least six complete runs per soil condition, using the same kit, soils, containers, calibration procedure, and 40-minute schedule as the pupils, and screened against the same five rules. Its pooled runs fix the per-condition jump thresholds used in <bold>Table 3</bold>, provide the substitute dataset for a run that cannot be interpreted, and establish that the two soils are distinguishable within the measurement window. Because the substitute data follow the same protocol, a group working from them faces the same CER task; every substitution is recorded as a fidelity deviation and reported with the results.</p>
        <p>2.2.3. Assignment and Balancing of Inquiry Variables</p>
        <p>Allowing each group to select its own variable would produce datasets of different shapes and reasoning tasks of different difficulty, confounding any comparison of CER performance between conditions. Here the focal variable is identical for all groups, and both of its levels are investigated within each group: every group sets up one sandy-soil container and one compost container and measures them simultaneously on its two channels. Every group therefore produces a dataset of the same structure, two conditions by nine time points, and faces the same reasoning demand. Light exposure, initial water amount, and container surface area are retained only as optional extensions after the assessed CER task, recorded as process variables and never used in a between-condition comparison.</p>
        <p>Assignment proceeds in three steps. First, pupils are ranked within class on the pre-test CER total and allocated to mixed-attainment groups of three to four, so that group mean prior attainment is comparable. Second, groups rather than pupils are randomly assigned to the AI-DRS or the static-scaffold condition, using block randomization with class as the blocking factor and a computer-generated sequence prepared by a researcher not involved in delivery; allocation is concealed until the pre-test has been scored. </p>
        <p>The group is thus both the unit of assignment and of delivery, which is the basis of the analysis in Section 2.4. Third, equipment is counterbalanced: probes, micro:bit units, and bench positions are rotated across conditions and across the two soils within each group, so that no device or location is systematically associated with one condition or one soil.</p>
        <p>Standardization of the materials is specified in <bold>Table 4</bold>. Balance is verified before analysis by reporting, separately by condition, the distribution of pre-test scores, soil batches, sensor units, bench positions, ambient temperature and relative humidity, and prior micro:bit and GenAI experience. Any imbalance is reported descriptively rather than adjusted away: with only five to ten groups per condition, randomization alone cannot be relied upon to produce equivalence, and the comparisons are interpreted accordingly.</p>
        <p><bold>Table 4.</bold> Standardization and assignment of the inquiry variables.</p>
        <table-wrap id="tbl4">
          <label>Table 4</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Variable</bold>
                </td>
                <td>
                  <bold>Status in the design</bold>
                </td>
                <td>
                  <bold>Specification</bold>
                </td>
                <td>
                  <bold>Balancing procedure</bold>
                </td>
              </tr>
              <tr>
                <td>Soil type</td>
                <td>Manipulated, within group</td>
                <td>Sandy soil and potting compost, equal dry mass (±5 g).</td>
                <td>Both levels in every group; one batch per soil type for the whole pilot.</td>
              </tr>
              <tr>
                <td>Initial water volume</td>
                <td>Controlled</td>
                <td>100 ml delivered to each container with the same syringe at t = 0.</td>
                <td>Identical instruction and equipment; volume recorded on the worksheet.</td>
              </tr>
              <tr>
                <td>Container</td>
                <td>Controlled</td>
                <td>Identical pots: same volume, surface area, and drainage.</td>
                <td>Supplied pre-assembled and checked before the session.</td>
              </tr>
              <tr>
                <td>Probe depth and position</td>
                <td>Controlled</td>
                <td>5 cm depth, at least 2 cm from the wall, marked with tape.</td>
                <td>Checked by the teacher before t = 0.</td>
              </tr>
              <tr>
                <td>Bench position and light</td>
                <td>Controlled in the assessed task</td>
                <td>Same room and bench, no direct sunlight.</td>
                <td>Positions rotated across conditions.</td>
              </tr>
              <tr>
                <td>Temperature and humidity</td>
                <td>Recorded covariate</td>
                <td>Recorded once per session.</td>
                <td>Both conditions run in the same room within the same two-hour window; cross-condition exposure is logged and reported (Section 2.2.4).</td>
              </tr>
              <tr>
                <td>Probes and micro:bit units</td>
                <td>Recorded and counterbalanced</td>
                <td>Two probes and two micro:bit units per group, one per container; unit identifier logged with each container.</td>
                <td>Units rotated across conditions and across the two soils within each group.</td>
              </tr>
              <tr>
                <td>Group composition</td>
                <td>Assigned</td>
                <td>Mixed-attainment groups of three to four, stratified on pre-test CER.</td>
                <td>Block randomization of groups to condition, with class as the block.</td>
              </tr>
              <tr>
                <td>Light, water amount, surface area</td>
                <td>Manipulated only in the optional extension</td>
                <td>Held constant throughout the assessed CER task; varied only in the optional extension that follows.</td>
                <td>Excluded from all between-condition comparisons.</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>2.2.4. Comparison Conditions</p>
        <p>The two conditions are identical in every respect except the form of analysis support: the experimental groups receive the adaptive AI-DRS dialogue, the active control groups an equivalent static worksheet with the same questions in a fixed order. Hardware, phenomenon, dataset, time on task, and the individual final assessment are the same in both (<bold>Table 5</bold>). One source of contamination cannot be removed at this scale. Because both conditions run in the same room within the same two-hour window (<bold>Table 4</bold>), control groups can see and overhear the AI-DRS groups, and any diffusion of dialogic moves would attenuate the observed difference. Separating the conditions by room or session would confound condition with setting, so cross-condition exposure is instead logged and reported as a limitation.</p>
        <p><bold>Table 5.</bold> Comparison conditions.</p>
        <table-wrap id="tbl5">
          <label>Table 5</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Element</bold>
                </td>
                <td>
                  <bold>Experimental group</bold>
                </td>
                <td>
                  <bold>Active control</bold>
                </td>
              </tr>
              <tr>
                <td>Hardware and phenomenon</td>
                <td>Nezha-micro:bit Smart Plant inquiry</td>
                <td>Identical</td>
              </tr>
              <tr>
                <td>Dataset</td>
                <td>Pupil-generated sensor data</td>
                <td>Own group-generated dataset, identical protocol</td>
              </tr>
              <tr>
                <td>CER questions</td>
                <td>Adaptive AI-DRS dialogue</td>
                <td>Equivalent static worksheet</td>
              </tr>
              <tr>
                <td>Time on task</td>
                <td>Equivalent</td>
                <td>Equivalent</td>
              </tr>
              <tr>
                <td>Teacher support</td>
                <td>Recorded</td>
                <td>Recorded</td>
              </tr>
              <tr>
                <td>Final assessment</td>
                <td>Individual, without AI</td>
                <td>Individual, without AI</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
      </sec>
      <sec id="sec2dot3">
        <title>2.3. Data Collection Technique</title>
        <p>Data will be collected across the seven-phase learning sequence detailed in Section 3. The system captures pupils’ interactions as pseudonymized chat logs, their measurements as datasets and graphs, and their explanations as written CER responses, while the teacher completes a structured observation log throughout. Equivalence and fidelity indicators are recorded: prior science knowledge, baseline CER performance, prior micro:bit and GenAI experience, time on task, absences, technical problems, teacher interventions, and number of AI turns. A subset of lessons is reviewed against a fidelity checklist by a second observer.</p>
      </sec>
      <sec id="sec2dot4">
        <title>2.4. Data Analysis</title>
        <p>In the planned pilot, two independent raters will score at least 25% of the responses. Weighted Cohen’s kappa will be computed for the ordinal rubric dimensions and an intraclass correlation coefficient for the total CER score, with disagreements resolved through documented consensus. Chat logs will be examined using a scheme that classifies the system’s moves as Socratic questioning, data grounding, conceptual bridging, guided choice, explicit correction, or metacognitive prompting, and pupil responses as full, partial, none, or not reached. Synthesis is recognized only when the pupil articulates it; accepting an AI-generated summary does not count as independent reasoning.</p>
        <p>The analysis plan follows from the unit of assignment. Because groups of three to four pupils are randomized and taught as a unit while outcomes are measured individually, pupils are nested within groups, and the planned sample corresponds to only five to ten groups per condition. An individual-level ANCOVA treating pupils as independent observations is therefore not an appropriate primary analysis, where the intraclass correlation is positive, standard errors are understated, and the associated p-values are not interpretable, and reporting an intraclass correlation alongside an uncorrected model does not remedy this ([<xref ref-type="bibr" rid="B11">11</xref>]). Multilevel modeling is not a workable alternative at this scale either, because variance components estimated from fewer than roughly ten clusters per arm are unstable and their standard errors are biased downwards ([<xref ref-type="bibr" rid="B20">20</xref>]).</p>
        <p>The primary analysis is therefore specified at cluster level and is descriptive rather than confirmatory. For each group, the mean post-test and mean pre-test CER score are computed, and the between-condition contrast is the difference in group mean post-test scores adjusted for group mean pre-test scores. Every group mean is displayed individually, so that the small number of clusters is visible rather than hidden behind an aggregate. The effect size is a cluster-adjusted standardized mean difference using the total between-plus-within variance in its denominator, reported with the intraclass correlation on which it depends ([<xref ref-type="bibr" rid="B11">11</xref>]). The primary inferential procedure is an exact permutation test on the groups’ condition labels, because with five to ten clusters per arm it is the only exact procedure available. At the lower end of that range, five groups per arm, the test cannot return a two-sided p-value below about .008 and has negligible power against any plausible effect; it is therefore reported as a descriptive statement of how extreme the observed contrast is among the possible label assignments, not as a test of efficacy. Interval estimates use a wild cluster bootstrap-t, developed precisely because pairs-resampling and percentile-type bootstraps behave poorly with few clusters; a bias-corrected and accelerated bootstrap resampling of whole groups is deliberately not used, for the same reason ([<xref ref-type="bibr" rid="B6">6</xref>]). Both procedures are conservative at this number of clusters and the resulting intervals will be wide, which is accepted as the cost of correct inference rather than concealed behind a nominally narrower pupil-level interval.</p>
        <p>Pupil-level ANCOVA is retained only as a secondary, illustrative analysis, with cluster-robust (CR2) standard errors and Satterthwaite degrees of freedom, designed for this situation of few clusters ([<xref ref-type="bibr" rid="B4">4</xref>]); no pupil-level p-value is treated as evidence of effectiveness. The intraclass correlation is reported as an interval, not a point estimate, since with at most ten clusters per arm its confidence interval is necessarily wide, and it is accompanied by a sensitivity table giving the design effect, Deff = 1 + (m − 1) × ICC, and the effective sample size, N/Deff, across a plausible range: with groups of three to four, an intraclass correlation of .05, .10, or .20 reduces an effective sample of 60 pupils to about 53, 48, or 40. This bounds the number of clusters a subsequent cluster-randomized trial would require, where the number of groups rather than of pupils determines power; it is a planning range, not an estimate a pilot of this size could deliver. The analysis plan, including the stopping rules for the process measures, will be pre-registered. <bold>Table 6</bold> summarizes the unit of analysis, estimand, and inference for each outcome.</p>
        <p><bold>Table 6.</bold> Analysis plan: outcomes, units of analysis, estimands, and inference.</p>
        <table-wrap id="tbl6">
          <label>Table 6</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Outcome</bold>
                </td>
                <td>
                  <bold>Unit of analysis</bold>
                </td>
                <td>
                  <bold>Estimand</bold>
                </td>
                <td>
                  <bold>Inference</bold>
                </td>
                <td>
                  <bold>Status</bold>
                </td>
              </tr>
              <tr>
                <td>Post-test CER</td>
                <td>Group (5 - 10 per condition)</td>
                <td>Difference in group mean post-test scores, adjusted for group mean pre-test scores.</td>
                <td>Exact permutation test on group labels (primary); wild cluster bootstrap-t for interval estimates.</td>
                <td>Primary, exploratory</td>
              </tr>
              <tr>
                <td>Transfer task without AI</td>
                <td>Group</td>
                <td>Difference in group means on the individually administered transfer task.</td>
                <td>As for the primary outcome.</td>
                <td>Key outcome, exploratory</td>
              </tr>
              <tr>
                <td>Post-test CER</td>
                <td>Pupil</td>
                <td>ANCOVA-adjusted difference with pre-test as covariate.</td>
                <td>CR2 cluster-robust standard errors with Satterthwaite degrees of freedom; p-values not reported as evidence.</td>
                <td>Secondary, illustrative</td>
              </tr>
              <tr>
                <td>Clustering parameters</td>
                <td>Group</td>
                <td>Intraclass correlation as an interval; design effect and effective sample size across a sensitivity range (0.05, 0.10, 0.20).</td>
                <td>Interval estimate with an explicit sensitivity table; no point estimate is relied upon.</td>
                <td>Planning range for a future trial</td>
              </tr>
              <tr>
                <td>Chat-log process codes</td>
                <td>Turn and pupil</td>
                <td>Frequencies of system moves and of pupil uptake.</td>
                <td>Descriptive statistics and joint displays.</td>
                <td>Descriptive</td>
              </tr>
              <tr>
                <td>Rater agreement</td>
                <td>Response</td>
                <td>Weighted Cohen’s kappa and the intraclass correlation for totals.</td>
                <td>Point estimates with confidence intervals.</td>
                <td>Quality control</td>
              </tr>
              <tr>
                <td>Fidelity and feasibility</td>
                <td>Session and group</td>
                <td>Completion, time on task, technical failures, teacher interventions, flags raised.</td>
                <td>Descriptive statistics.</td>
                <td>Feasibility</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>The qualitative analysis will focus on patterns of CER error and will use thematic analysis for explanations, integrating turn-level chat-log coding with comparative cases of full, partial, and zero uptake in joint displays that connect CER gain, data-grounding frequency, bridge uptake, engagement, and transfer.</p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. The AI-DRS Physical Inquiry Model</title>
      <p>The development strand produces the AI-DRS Physical Inquiry Model, a proposed intervention in which GenAI acts as a dialogic scaffold that helps primary pupils move from observation to evidence-based scientific reasoning and then steps back. The sections below set out its reasoning progression, control principles, seven-phase sequence, response-contingency architecture, implementation specification, escalation protocol, fading logic, assessment instruments, and alignment with the Technological Pedagogical Content Knowledge (TPACK) framework. <xref ref-type="fig" rid="fig2">Figure 2</xref> illustrates its overall logic.</p>
      <sec id="sec3dot1">
        <title>3.1. From Measurement to Scientific Explanation</title>
        <p>Sensors convert physical states into numerical values, but their scientific meaning comes from the relations among values, conditions, and theoretical ideas. Pupils must decide which data are relevant, whether a measurement is trustworthy, which patterns are present, and what mechanism might explain them. The model foregrounds four practices: generating data, checking data quality, selecting evidence, and explaining mechanistically. It organizes them through a CER progression ([<xref ref-type="bibr" rid="B18">18</xref>]) adapted for the primary years: Observe/Describe, Predict, Claim, Evidence, Reasoning, and Reflect and Transfer.</p>
        <p>At the heart of the model is a dialogic control layer that responds to the pupil’s input. The system advances the pupil’s thinking through targeted questions instead of presenting a ready-made explanation: the interaction starts with observation, asks for comparisons, requires a claim, returns the pupil to the actual measurements, and only then prompts for mechanistic reasoning. Its value depends on the quality of the pupil’s responses, not the number of exchanges. Three principles govern the interaction. Data grounding requires pupils to cite specific values, changes, or patterns from their own data and rejects vague statements such as “the soil dried out more”. Epistemic gatekeeping makes the move from claim to evidence conditional on a testable claim, the move from evidence to reasoning conditional on relevant data, and completion conditional on a mechanism stated in the pupil’s own words.</p>
        <p>Summaries produced by the system are never counted as the pupil’s own synthesis. Contingent responsiveness determines the next move from the quality of the response, classified as adequate and independent, partially adequate, off-target, or scientifically inaccurate, for instructional routing only. <xref ref-type="fig" rid="fig2">Figure 2</xref> shows these principles as a layer beneath the reasoning progression, with AI support decreasing as reasoning becomes more independent.</p>
        <fig id="fig2">
          <label>Figure 2</label>
          <graphic xlink:href="https://html.scirp.org/file/6309660-rId17.jpeg?20260825041906" />
        </fig>
        <p><bold>Figure 2.</bold> The AI-DRS reasoning progression and its dialogic control layer.</p>
      </sec>
      <sec id="sec3dot2">
        <title>3.2. The Seven-Phase Learning Sequence</title>
        <p>The intervention is enacted through seven phases, each pairing a pupil activity with a defined role for the AI-DRS system or the teacher and yielding a distinct piece of evidence for analysis (<bold>Table 7</bold>). The system does not answer during orientation, offers no correction during prediction, and enforces fair-test criteria before construction begins.</p>
        <p><bold>Table 7.</bold> The seven phases of the AI-DRS Physical Inquiry Model.</p>
        <table-wrap id="tbl7">
          <label>Table 7</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Phase</bold>
                </td>
                <td>
                  <bold>Pupil activity</bold>
                </td>
                <td>
                  <bold>AI-DRS/teacher role</bold>
                </td>
                <td>
                  <bold>Evidence produced</bold>
                </td>
              </tr>
              <tr>
                <td>1. Orientation</td>
                <td>Confront a real watering problem</td>
                <td>Teacher activates prior ideas; no AI answer</td>
                <td>Initial conceptions</td>
              </tr>
              <tr>
                <td>2. Prediction</td>
                <td>Predict the pattern and give an initial rationale</td>
                <td>AI requests a clear prediction, not a correction</td>
                <td>Prediction statement</td>
              </tr>
              <tr>
                <td>3. Experimental design</td>
                <td>Apply the fair-test protocol: identify the manipulated variable and the controls</td>
                <td>Gatekeeping for a fair test</td>
                <td>Group protocol</td>
              </tr>
              <tr>
                <td>4. Construction and calibration</td>
                <td>Connect, program, and calibrate</td>
                <td>Technical hints: teacher escalation for hardware</td>
                <td>Working system</td>
              </tr>
              <tr>
                <td>5. Data collection</td>
                <td>Take repeated measurements and graph them</td>
                <td>Quality check of unusual values</td>
                <td>Dataset and graph</td>
              </tr>
              <tr>
                <td>6. CER dialogue</td>
                <td>State claim, evidence, and reasoning</td>
                <td>Data grounding, gatekeeping, conceptual bridging</td>
                <td>AI-supported CER</td>
              </tr>
              <tr>
                <td>7. Fading and transfer</td>
                <td>Solve a new individual problem without full help</td>
                <td>Reduced or absent prompts</td>
                <td>Independent CER transfer</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
      </sec>
      <sec id="sec3dot3">
        <title>3.3. The Response-Contingency Architecture</title>
        <p>The model’s adaptivity is defined by an explicit rule set linking the quality of a pupil’s response to the system’s next move, so that support remains just sufficient to keep the reasoning on track (<bold>Table 8</bold>). </p>
        <p>A general statement without data triggers a data-grounding request; data without a mechanism triggers a conceptual bridge; a scientifically inaccurate response triggers explicit correction and a teacher notification. Rule R9 sits outside the instructional sequence: if the exchange touches on a pupil’s welfare, causes distress, or moves off task, the dialogue stops and the teacher is alerted.</p>
        <p><bold>Table 8.</bold>Response-contingency rules.</p>
        <table-wrap id="tbl8">
          <label>Table 8</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Rule</bold>
                </td>
                <td>
                  <bold>Pupil response or system state</bold>
                </td>
                <td>
                  <bold>Next move</bold>
                </td>
                <td>
                  <bold>Criterion</bold>
                </td>
              </tr>
              <tr>
                <td>R1</td>
                <td>Complete and independent</td>
                <td>Acknowledge, advance, fade</td>
                <td>Clear epistemic practice without a prompt</td>
              </tr>
              <tr>
                <td>R2</td>
                <td>Partially adequate: the prerequisite for the next step is missing</td>
                <td>Focused prompt for the missing element</td>
                <td>Has a claim or data, but not the full link</td>
              </tr>
              <tr>
                <td>R3</td>
                <td>General, without data</td>
                <td>Data-grounding prompt</td>
                <td>Requires specific values</td>
              </tr>
              <tr>
                <td>R4</td>
                <td>Data without a mechanism</td>
                <td>Conceptual bridge</td>
                <td>Why/how connection</td>
              </tr>
              <tr>
                <td>R5</td>
                <td>First unsuccessful attempt</td>
                <td>Reformulation</td>
                <td>Same goal, simpler language</td>
              </tr>
              <tr>
                <td>R6</td>
                <td>Second unsuccessful attempt</td>
                <td>Guided choice or bounded hint</td>
                <td>Avoiding an impasse</td>
              </tr>
              <tr>
                <td>R7</td>
                <td>Scientifically unsafe misconception</td>
                <td>Explicit correction and teacher flag</td>
                <td>Risk of consolidating an error</td>
              </tr>
              <tr>
                <td>R8</td>
                <td>
                  Technically unreliable data (
                  <bold>Table 3</bold>
                  )
                </td>
                <td>Halt interpretation, teacher check</td>
                <td>Defective measurements are not explained</td>
              </tr>
              <tr>
                <td>R9</td>
                <td>Safeguarding-relevant, distressing, or off-task content</td>
                <td>Dialogue stops; neutral holding message; teacher alerted</td>
                <td>Welfare and safety take precedence over the task</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>The architecture rests on a technical assumption. Although the contingency rules are explicitly defined, they are implemented by a stochastic generative model, so the real-time classification of pupils’ responses may not always be reliable.</p>
        <p>Misrouting is possible: an adequate answer may attract an unnecessary prompt, or an unsafe one may pass. Three design decisions bound the consequences. The classification is used only for low-stakes instructional routing, never for assessment; the default action under uncertainty is the mildest one, a reformulation, with a teacher flag raised if it recurs; and every interaction is logged, so that routing accuracy can be measured and becomes part of the validation criteria in Section 3.9.</p>
        <p>Dialogic moves are delivered as short, age-appropriate prompts, for example “Which two or three readings from your own table best support that claim?” for data grounding, “How might soil structure, drainage, or evaporation explain the different pattern?” for the conceptual bridge, and “Check my suggestion against your graph. What evidence would show that it is wrong?” for calibrated trust. Appendix A gives the full register and move specifications.</p>
      </sec>
      <sec id="sec3dot4">
        <title>3.4. Implementation Specification of the AI-DRS System</title>
        <p>A dialogic model implemented on a generative system is reproducible only if the configuration that produced the dialogue is stated. The reference implementation is therefore specified at the level of provider, model snapshot, prompt version, decoding settings, input format, and rule precedence.</p>
        <p>Pupils hold no accounts and use no consumer interface. A school-controlled web application delivers the dialogue; a group signs in with a class code and a pseudonymous group label, and the application calls the provider’s application programming interface server-side, under the institution’s account and key. The reference configuration uses a single pinned, dated model snapshot, logged with every turn. The snapshot named here, the OpenAI API with gpt-4o-2024-11-20, fixes the reference build rather than prescribing the model to be used; the snapshot in force at the pilot is reported verbatim, and any change triggers a re-run of the validation suite before use with pupils.</p>
        <p>The system prompt is a fixed, version-controlled artefact (AI-DRS-SP v1.0) rather than an ad hoc instruction, and is reproduced verbatim in Appendix A. Its twelve blocks specify role and audience; the register constraint of one question per turn, at most 45 words, and no technical vocabulary not yet introduced in class; the never-answer rule, under which the system does not state the claim, the evidence, or the mechanism on the pupil’s behalf and does not confirm an unjustified claim; data grounding; epistemic gatekeeping; the target-explanation reference from <bold>Table 1</bold>; the contingency rules of <bold>Table 8</bold> and their precedence; the fading levels; the calibrated-trust statement; the refusal and safeguarding rules; and the output schema. The four response categories named in Section 3.1 are the coarse form of the eight operational classification values defined in Block 6 of Appendix A, which separate the unsupported, unreliable, and safeguarding cases that the four categories collapse together.</p>
        <p>Decoding settings, output schema, failure fallback, and the exact input payload are specified in <bold>Table 9</bold>. Three of these are load-bearing. The 160-token limit keeps replies short enough for the target age group. A generation that fails schema validation is never shown to pupils as free text, but is retried once and then replaced by the deterministic reformulation move with a teacher flag. And nothing is transmitted beyond the group’s own calibrated dataset, the pupil’s current response, the phase and fade level, the two preceding turns, and the fixed target explanation, with class-list names removed by a client-side filter.</p>
        <p><bold>Table 9.</bold>Reference implementation specification of the AI-DRS system.</p>
        <table-wrap id="tbl9">
          <label>Table 9</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Parameter</bold>
                </td>
                <td>
                  <bold>Specification</bold>
                </td>
              </tr>
              <tr>
                <td>Access architecture</td>
                <td>School-controlled web application; no pupil accounts; class code and pseudonymous group label; server-side API calls under the institutional account.</td>
              </tr>
              <tr>
                <td>Provider and model</td>
                <td>A single pinned, dated snapshot, logged with every turn. The snapshot named here (OpenAI API, gpt-4o-2024-11-20) is indicative of the reference build; the snapshot current at the pilot is reported verbatim, and any change triggers re-validation.</td>
              </tr>
              <tr>
                <td>Provider data settings</td>
                <td>No training on submitted data; shortest available retention; only data-minimized, non-identifying content transmitted.</td>
              </tr>
              <tr>
                <td>System prompt</td>
                <td>AI-DRS-SP v1.0, fixed and version-controlled, reproduced verbatim in Appendix A: role, register, never-answer rule, data grounding, gatekeeping, target explanation, contingency rules and precedence, fading levels, calibrated trust, safeguarding, output schema.</td>
              </tr>
              <tr>
                <td>Register constraints</td>
                <td>One question per turn; maximum 45 words; no un-introduced technical vocabulary; no multi-part questions.</td>
              </tr>
              <tr>
                <td>Decoding settings</td>
                <td>Temperature 0.2; top-p 1.0; maximum 160 output tokens; no frequency or presence penalties; fixed seed where supported.</td>
              </tr>
              <tr>
                <td>Output format</td>
                <td>Schema-validated JSON with the fields classification, rule_fired, move_type, message, fade_level, and flag.</td>
              </tr>
              <tr>
                <td>Failure fallback</td>
                <td>One retry on schema failure, then the deterministic reformulation move with a teacher flag; generated free text is never shown unvalidated.</td>
              </tr>
              <tr>
                <td>Input payload</td>
                <td>Calibrated dataset (≤ 2 containers × 9 points), the pupil’s current response, phase and fade level, the two preceding turns, and the fixed target explanation.</td>
              </tr>
              <tr>
                <td>Data minimization</td>
                <td>Client-side removal of class-list names; no pupil identifiers, no teacher notes, no prior sessions transmitted.</td>
              </tr>
              <tr>
                <td>Logging and audit</td>
                <td>Timestamp, snapshot, prompt version, hash of the transmitted input, classification, rule fired, move issued, and any flag.</td>
              </tr>
              <tr>
                <td>Pre-session validation</td>
                <td>
                  A fixed regression suite of at least 36 synthetic pupil responses covering every category in
                  <bold>Table 8</bold>
                  , with no fewer than five items each for R7 and R9. Required before use with pupils: at least 80% overall, at least 70% in every category, and correct routing on every R7 and R9 item. This is a pre-session check on curated items, distinct from and weaker than the small-group usability criterion of Section 3.9.
                </td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Because more than one contingency rule can match a single response, the rules are ordered, and the system applies only the highest-priority matching rule. <bold>Table 10</bold> sets out that hierarchy and maps every priority level onto the corresponding rule identifier in <bold>Table 8</bold>. Safeguarding overrides everything; a data-quality halt precedes any interpretive move, since interpreting unreliable measurements is worse than not interpreting them at all; scientific correction precedes gatekeeping; gatekeeping precedes data grounding and conceptual bridging; and fading is applied last, only when no higher rule has fired. When classification confidence is low, or two rules of equal priority match, the system takes the mildest available action, a reformulation, and raises a teacher flag if the difficulty recurs, so that the classifier’s failure mode is conservative by design.</p>
        <p><bold>Table 10.</bold>Rule precedence in the AI-DRS dialogue.</p>
        <table-wrap id="tbl10">
          <label>Table 10</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Priority</bold>
                </td>
                <td>
                  <bold>Condition detected</bold>
                </td>
                <td>
                  <bold>System action</bold>
                </td>
                <td>
                  <bold>Escalation</bold>
                </td>
                <td>
                  <bold>Table 8</bold>
                  <bold>rule</bold>
                </td>
              </tr>
              <tr>
                <td>P1 Safeguarding</td>
                <td>Personal disclosure, distress, unsafe or off-task content.</td>
                <td>Dialogue stops immediately; a neutral holding message is shown.</td>
                <td>Red flag; teacher within 2 minutes.</td>
                <td>R9</td>
              </tr>
              <tr>
                <td>P2 Data quality</td>
                <td>
                  Run-level or drift flag from
                  <bold>Table 3</bold>
                  .
                </td>
                <td>Interpretation halted; no claim, evidence, or reasoning move is issued.</td>
                <td>Amber flag; teacher within the lesson.</td>
                <td>R8</td>
              </tr>
              <tr>
                <td>P3 Scientific accuracy</td>
                <td>A scientifically unsafe misconception in the pupil response.</td>
                <td>Brief explicit correction followed by a re-check question.</td>
                <td>Red flag; teacher notified.</td>
                <td>R7</td>
              </tr>
              <tr>
                <td>P4 Gatekeeping</td>
                <td>The prerequisite for the next step is missing.</td>
                <td>The pupil is held at the current step with a focused prompt.</td>
                <td>Logged only.</td>
                <td>R2</td>
              </tr>
              <tr>
                <td>P5 Data grounding</td>
                <td>A general statement with no specific values.</td>
                <td>Request for two or three specific readings from the pupil’s own table.</td>
                <td>Logged only.</td>
                <td>R3</td>
              </tr>
              <tr>
                <td>P6 Conceptual bridging</td>
                <td>Data cited but no mechanism offered.</td>
                <td>A why-or-how prompt directed at the target mechanism.</td>
                <td>Logged only.</td>
                <td>R4</td>
              </tr>
              <tr>
                <td>P7 Fading</td>
                <td>The epistemic practice is performed independently.</td>
                <td>Acknowledge, advance, and reduce the level of prompting.</td>
                <td>Logged only.</td>
                <td>R1</td>
              </tr>
              <tr>
                <td>P8 Default</td>
                <td>Low classification confidence, or two rules of equal priority.</td>
                <td>Reformulation in simpler language; after two failures, a bounded guided choice.</td>
                <td>Amber flag after the second failure.</td>
                <td>R5, then R6</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
      </sec>
      <sec id="sec3dot5">
        <title>3.5. Teacher Escalation, Review, and Stop Rules</title>
        <p>Flagging is useful only if it is attached to a named person, a defined response time, and a defined endpoint. The class teacher is the first responder and is present throughout; a live dashboard shows one row per group with the current phase, the fade level, and any open flag. A designated science lead, or a member of the research team, is the second reviewer and examines every flagged interaction within 24 hours, recording the outcome in a review log. Where a flag concerns a child’s welfare rather than the science, the school’s designated safeguarding lead takes over under the school’s existing child-protection procedure.</p>
        <p>Flags are graded (<bold>Table 11</bold>). A red flag pauses the dialogue automatically and alerts the teacher, who is expected to reach the group within two minutes; triggers include a safeguarding disclosure, signs of distress, content the filters should have blocked, explicit correction of a scientifically unsafe misconception, and any turn in which the system appears to have performed the pupils’ reasoning. An amber flag does not pause the dialogue but requires a teacher visit within the same lesson. Routine events such as a change of fade level are logged as green entries and reviewed after the lesson.</p>
        <p>The corrective sequence is fixed: the dialogue pauses; the teacher reads the last three turns on the dashboard, speaks with the group, and establishes whether the difficulty is scientific, technical, or affective; the teacher then resumes with corrected input, transfers the group to the static worksheet scaffold for the remainder of the session, or ends that group’s AI dialogue. The second and third outcomes are recorded as fidelity deviations.</p>
        <p>The dialogue must stop, and may not resume within that session, if the system produces scientifically incorrect content the teacher cannot correct on the spot; if it supplies the claim, evidence, or mechanism the pupils were asked to produce; if a pupil discloses safeguarding-relevant information or shows distress; if the system is used off task after one warning; if two red flags occur in one group within a session; or if a data-quality halt cannot be resolved. The group then completes the sequence with the static scaffold, so that no pupil loses the science lesson because of a system failure. The incident is logged and reported to the research team within 24 hours and, where reportable, to the ethics committee. Recurrent prompt failures identified in the weekly review trigger a prompt revision, a version increment, and a re-run of the regression suite.</p>
        <p><bold>Table 11.</bold> Teacher escalation protocol.</p>
        <table-wrap id="tbl11">
          <label>Table 11</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Level</bold>
                </td>
                <td>
                  <bold>Triggers</bold>
                </td>
                <td>
                  <bold>Who acts, and when</bold>
                </td>
                <td>
                  <bold>Action</bold>
                </td>
                <td>
                  <bold>Review</bold>
                </td>
              </tr>
              <tr>
                <td>Red</td>
                <td>Safeguarding disclosure; distress; blocked content that appeared; explicit correction of an unsafe misconception; the system reasoning on the pupils’ behalf.</td>
                <td>Class teacher, within 2 minutes; dialogue pauses automatically.</td>
                <td>Read the last three turns, speak with the group, then resume, switch to the static scaffold, or stop.</td>
                <td>Second reviewer within 24 hours; safeguarding lead where welfare is involved.</td>
              </tr>
              <tr>
                <td>Amber</td>
                <td>Two unsuccessful reformulations; data-quality halt; hardware fault; latency above 30 s.</td>
                <td>Class teacher, within the same lesson; dialogue continues.</td>
                <td>Technical or procedural fix; the affected data are marked.</td>
                <td>Second reviewer within 24 hours.</td>
              </tr>
              <tr>
                <td>Green</td>
                <td>Routine events such as a change of fade level or a completed phase.</td>
                <td>No action during the lesson.</td>
                <td>Logged for process analysis.</td>
                <td>Weekly review of aggregated logs.</td>
              </tr>
              <tr>
                <td>Hard stop</td>
                <td>Uncorrectable scientific error; the system supplies claim, evidence, or reasoning; disclosure or distress; off-task use after one warning; two red flags in one group; an unresolved data-quality halt.</td>
                <td>Class teacher, immediately.</td>
                <td>AI dialogue ends for that group; the sequence is completed with the static scaffold.</td>
                <td>Reported to the research team within 24 hours, and to the ethics committee where reportable; prompt revision and re-validation follow.</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
      </sec>
      <sec id="sec3dot6">
        <title>3.6. Fading, Uptake and Engagement</title>
        <p>The presence of a scaffold does not guarantee uptake: a conceptual bridge may lead to full articulation, simple acceptance, or no response at all. The model therefore separates the system’s moves from the quality of the pupil’s uptake and treats engagement as a function of the substance of the response, data use, revision, and persistence rather than the number of turns. Fading occurs when the pupil performs the required practice independently and proceeds through four levels: a direct prompt naming the step, a delayed prompt issued only after a turn’s pause, a minimal prompt that acknowledges the response without asking anything further, and no prompt at all. Support is reduced by one level at a time and may return temporarily if performance declines. Block 9 of Appendix A states the corresponding fade_level values.</p>
      </sec>
      <sec id="sec3dot7">
        <title>3.7. Assessment Instruments</title>
        <p>Two instruments make reasoning visible. The CER rubric (<bold>Table 12</bold>) scores six dimensions to a maximum of twelve points. Its Mechanism dimension is scored against the target mechanism of <bold>Table 1</bold> rather than a general standard: a score of 2, partial mechanism, requires a correct but incomplete appeal to particle size, pore space, or organic matter, for example naming faster drainage without relating it to soil structure, whereas a score of 3, full and consistent, requires the pupil to link a stated numerical difference in the rate of decline to the particle-size and pore-space account of <bold>Table 1</bold> while making no claim outside the bounds listed there. Raters use the same <bold>Table 1</bold> reference that is supplied to the AI-DRS system, so that the human and machine standards coincide. The chat-log coding scheme classifies the system’s moves and the pupil’s responses so that only pupil-expressed synthesis counts as independent reasoning.</p>
        <p><bold>Table 12.</bold>Claim-Evidence-Reasoning (CER) rubric.</p>
        <table-wrap id="tbl12">
          <label>Table 12</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Dimension</bold>
                </td>
                <td>
                  <bold>0</bold>
                </td>
                <td>
                  <bold>1</bold>
                </td>
                <td>
                  <bold>2</bold>
                </td>
                <td>
                  <bold>3</bold>
                </td>
              </tr>
              <tr>
                <td>Claim</td>
                <td>Absent/incorrect</td>
                <td>Vague</td>
                <td>Clear and testable</td>
                <td>-</td>
              </tr>
              <tr>
                <td>Evidence relevance</td>
                <td>No reference</td>
                <td>Partially relevant</td>
                <td>Relevant and sufficient</td>
                <td>-</td>
              </tr>
              <tr>
                <td>Data specificity</td>
                <td>No data</td>
                <td>General trend</td>
                <td>Specific values/changes</td>
                <td>-</td>
              </tr>
              <tr>
                <td>Evidence-claim link</td>
                <td>Absent</td>
                <td>Implicit connection</td>
                <td>Explicit connection</td>
                <td>-</td>
              </tr>
              <tr>
                <td>
                  Mechanism(against
                  <bold>Table 1</bold>
                  )
                </td>
                <td>Absent/wrong</td>
                <td>Descriptive</td>
                <td>Partial mechanism: correct but incomplete appeal to particle size, pore space, or organic matter</td>
                <td>
                  Full and consistent: numerical rate difference linked to the particle-size and pore-space account of
                  <bold>Table 1</bold>
                </td>
              </tr>
              <tr>
                <td>Limitations</td>
                <td>Not recognized</td>
                <td>Recognizes a limit/alternative</td>
                <td>-</td>
                <td>-</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
      </sec>
      <sec id="sec3dot8">
        <title>3.8. TPACK Alignment</title>
        <p>The model is designed so that technology is not treated separately from the science content and the pedagogical goals it supports ([<xref ref-type="bibr" rid="B21">21</xref>]). This view is consistent with computational-pedagogy frameworks for STEAM education, which emphasize the coordinated interaction of content, pedagogical practice, computational methods, and technological tools ([<xref ref-type="bibr" rid="B29">29</xref>]), and with work integrating AI-based assessment models and authentic STEAM engineering tasks in contexts such as precision agriculture ([<xref ref-type="bibr" rid="B40">40</xref>]). Each phase involves a specific combination of content, pedagogical, and technological knowledge (<bold>Table 13</bold>); the analysis and explanation phases represent the fullest TPACK intersection, where data logging and the AI-DRS dialogue work with CER scaffolding to promote mechanistic reasoning.</p>
        <p><bold>Table 13.</bold>TPACK alignment of the intervention phases.</p>
        <table-wrap id="tbl13">
          <label>Table 13</label>
          <table>
            <tbody>
              <tr>
                <td>
                  <bold>Phase</bold>
                </td>
                <td>
                  <bold>Content knowledge</bold>
                </td>
                <td>
                  <bold>Pedagogical knowledge</bold>
                </td>
                <td>
                  <bold>Technology</bold>
                </td>
                <td>
                  <bold>Intersection</bold>
                </td>
              </tr>
              <tr>
                <td>1. Orientation</td>
                <td>Moisture, plants, change</td>
                <td>Elicitation</td>
                <td>-</td>
                <td>PCK</td>
              </tr>
              <tr>
                <td>2. Prediction</td>
                <td>Variables and mechanisms</td>
                <td>Inquiry learning</td>
                <td>-</td>
                <td>PCK</td>
              </tr>
              <tr>
                <td>3. Experimental design</td>
                <td>Fair testing</td>
                <td>Collaborative inquiry</td>
                <td>micro:bit design</td>
                <td>TPACK</td>
              </tr>
              <tr>
                <td>4. Construction and calibration</td>
                <td>Measurement and calibration</td>
                <td>Learning by making</td>
                <td>Nezha, sensor, MakeCode</td>
                <td>TCK/TPK</td>
              </tr>
              <tr>
                <td>5. Data collection</td>
                <td>Patterns and graphs</td>
                <td>CER scaffolding</td>
                <td>Data logging</td>
                <td>TPACK</td>
              </tr>
              <tr>
                <td>6. CER dialogue</td>
                <td>Evidence-mechanism</td>
                <td>Dialogic scaffolding</td>
                <td>AI-DRS</td>
                <td>TPACK</td>
              </tr>
              <tr>
                <td>7. Fading and transfer</td>
                <td>Applying the model</td>
                <td>Fading/metacognition</td>
                <td>Reduced AI support</td>
                <td>TPACK</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
      </sec>
      <sec id="sec3dot9">
        <title>3.9. Readiness for Validation</title>
        <p>Finally, the development strand sets out the criteria that must be met before the pilot can proceed. A panel of five to seven experts in science education, primary teaching methods, educational technology, and research ethics will rate each element for relevance, clarity, feasibility, developmental appropriateness, scientific accuracy, and safeguarding, and will record item- and scale-level content-validity indices (I-CVI ≥ .78 and S-CVI/Ave ≥ .90).</p>
        <p>The minimum requirements are: no unresolved objection to scientific accuracy; a successful review of the prompts against misconception, privacy, and refusal cases; routing accuracy of at least 80% in the small-group usability test, measured as the percentage of authentic pupil responses that the system categorizes as an independent human coder does, with no category of <bold>Table 8</bold> below 70% and with correct routing on every safeguarding (R9) and scientific-accuracy (R7) item, since an aggregate threshold can be satisfied while the two safety-critical rules fail; the pre-session regression suite of <bold>Table 9</bold>is scored against the same per-category floors but is a separate and weaker check, because its items are curated and easier than authentic responses, and meeting it does not discharge the usability criterion; stable hardware operation and logging in at least 90% of trials; pupil comprehension of instructions without extended explanation; teacher feasibility within the timetabled slots; a documented dry run of the escalation protocol including at least one simulated red flag; and a pilot scoring agreement of κ ≥ .70.</p>
      </sec>
    </sec>
    <sec id="sec4">
      <title>4. Discussion</title>
      <p>The model can be situated against three areas of prior research. First, GenAI in science often assumes the role of a single source of knowledge ([<xref ref-type="bibr" rid="B8">8</xref>]), which, combined with young children’s tendency to over-trust conversational agents ([<xref ref-type="bibr" rid="B22">22</xref>]), explains why adding an answer-provider role to reasoning instruction would backfire. AI-DRS reverses that role: the technology is designed to ask rather than to answer, and its data-grounding and gatekeeping rules keep pupils’ own measurements at the center of every explanation. It thereby extends the finding that a customized chatbot can support reasoning and argumentation in secondary science ([<xref ref-type="bibr" rid="B33">33</xref>]) to primary education, adding pupil-generated sensor data, explicit gatekeeping, and structured fading. This positioning is consistent with earlier evidence that digital technologies in young children’s science learning are most productive when they complement hands-on inquiry and teacher mediation rather than replace them ([<xref ref-type="bibr" rid="B13">13</xref>]).</p>
      <p>Second, the model sharpens the literature that reads AI through a scaffolding lens. Reviews there discuss scaffolding as personalized learning paths, tailored support, and immediate feedback, but are largely based in higher education and tend to treat scaffolding as adaptive delivery ([<xref ref-type="bibr" rid="B3">3</xref>]; [<xref ref-type="bibr" rid="B5">5</xref>]; [<xref ref-type="bibr" rid="B36">36</xref>]). Chatbot effects appear weaker for younger learners and to diminish over time ([<xref ref-type="bibr" rid="B39">39</xref>]), and the benefits of GenAI are entangled with risks to cognition and agency ([<xref ref-type="bibr" rid="B39">39</xref>]; [<xref ref-type="bibr" rid="B43">43</xref>]). This is not an argument against GenAI in primary education; it is what makes the model’s developmental and safety features load-bearing, since attunement, fading, and teacher involvement are necessary conditions at this age rather than optional extras. The model also foregrounds a dimension that delivery-centered accounts underplay: the dialogic and argumentative quality of the interaction ([<xref ref-type="bibr" rid="B1">1</xref>]; [<xref ref-type="bibr" rid="B23">23</xref>]; [<xref ref-type="bibr" rid="B35">35</xref>]; [<xref ref-type="bibr" rid="B37">37</xref>]).</p>
      <p>Third, the model must navigate tensions rather than remove them. The first concerns epistemic agency: heavy reliance on GenAI is associated with reduced critical thinking and weaker self-regulation ([<xref ref-type="bibr" rid="B15">15</xref>]; [<xref ref-type="bibr" rid="B43">43</xref>]). AI-DRS responds with fading criteria, a stance that withholds answers, and assessment only of knowledge pupils express themselves, though whether these safeguards work remains open. The second concerns trust and misinformation: young children are credulous users, so a system that occasionally errs could instill misconceptions, which is why teacher verification and AI literacy are built in from the outset ([<xref ref-type="bibr" rid="B22">22</xref>]; [<xref ref-type="bibr" rid="B24">24</xref>]). The third concerns the teacher, since the model assumes the technological, pedagogical, and content knowledge that makes professional development a precondition ([<xref ref-type="bibr" rid="B21">21</xref>]). The fourth concerns equity, since unequal access to devices and connectivity could widen gaps unless inclusion is addressed explicitly.</p>
      <p>The main theoretical contribution of this article is a sharper account of scaffolding when the scaffold is generative and conversational. Classically, a scaffold is defined by contingency, fading, and the transfer of responsibility to the learner ([<xref ref-type="bibr" rid="B34">34</xref>]; [<xref ref-type="bibr" rid="B38">38</xref>]). A generative system satisfies the first well, assessing a response and adjusting its next move, but it does not deliver the other two: it has no intrinsic mechanism for reducing support, and children who trust these systems are unlikely to question or refuse it ([<xref ref-type="bibr" rid="B22">22</xref>]). Left unregulated, the result is a scaffold that never diminishes, which in classical terms is a permanent aid rather than a scaffold. Better prompts cannot resolve this; it reflects a difference in kind between human scaffolding, which rests on pedagogical judgment, and automated scaffolding, whose responses are generated probabilistically within predefined rules.</p>
      <p>AI-DRS builds the consequences of that difference into its design. If fading will not emerge from the technology, it must be written into the interaction, with explicit criteria for when support is reduced and who decides. Success must therefore be measured by the independent explanation that follows the conversation, not by the completed exchange. This is why the model credits only pupil-generated synthesis, treats the AI-free transfer task as the key outcome, and specifies fading as a requirement rather than a hope. Its value is greatest where a teacher cannot supply tailored questioning to every group: the system increases the available dialogue while the teacher retains responsibility for scientific accuracy, safety, and the decision to use it at all.</p>
    </sec>
    <sec id="sec5">
      <title>5. Conclusions</title>
      <p>This study developed the AI-DRS Physical Inquiry Model as a structured, safety-oriented approach to supporting scientific reasoning in primary education. The model connects pupil-generated sensor data, contingent dialogue, CER scaffolding, and systematic fading within a Nezha-micro:bit soil-moisture investigation. Its central premise is that GenAI has educational value only when it helps pupils turn measurements into evidence and evidence into mechanistic explanations, while progressively transferring responsibility back to the learner.</p>
      <p>The study integrates elements previously examined largely in isolation: pupil-generated sensor data, dialogic GenAI support, CER scaffolding, and the systematic reduction of technological assistance. Physical-computing activities let pupils generate authentic measurements but rarely support the move from measurement to evidence and mechanism ([<xref ref-type="bibr" rid="B28">28</xref>]; [<xref ref-type="bibr" rid="B12">12</xref>]; [<xref ref-type="bibr" rid="B16">16</xref>]), whereas GenAI systems sustain personalized dialogue but may detach from learners’ observations or reason on their behalf ([<xref ref-type="bibr" rid="B8">8</xref>]; [<xref ref-type="bibr" rid="B42">42</xref>]). AI-DRS addresses this dual gap by keeping every substantive dialogue move grounded in pupils’ own measurements and positioning GenAI as a temporary reasoning scaffold rather than an answer-generating authority. Its first contribution is therefore a data-grounded bridge between physical inquiry and dialogic artificial intelligence in primary science.</p>
      <p>A second contribution is the explicit response-contingency and fading architecture. Rather than a general prompt or an unrestricted chatbot conversation, AI-DRS specifies how the system responds to complete, partial, unsupported, scientifically inaccurate, and technically unreliable reasoning; introduces gatekeeping criteria that block progress until defined requirements are met; and states when support is reformulated, intensified, reduced, temporarily restored, or transferred to the teacher. Fading thereby becomes an operational component of the design rather than an assumed by-product of repeated AI use, extending conventional accounts of scaffolding into a generative conversational environment ([<xref ref-type="bibr" rid="B38">38</xref>]; [<xref ref-type="bibr" rid="B34">34</xref>]). </p>
      <p>A third contribution concerns how learning and learner agency are evaluated. The model distinguishes the dialogic move produced by the AI from the pupil’s uptake of it, avoiding the assumption that exposure to an AI-generated explanation is equivalent to learning. Only explanations pupils articulate themselves count as evidence of reasoning, and the central outcome is an independent CER explanation produced after AI support has been removed. With teacher oversight, calibrated-trust prompts, age-appropriate safeguarding, and mediated access through a school-controlled environment, this emphasis on post-scaffold transfer offers a framework for addressing over-trust, cognitive dependence, and epistemic agency in GenAI-supported learning ([<xref ref-type="bibr" rid="B22">22</xref>]; [<xref ref-type="bibr" rid="B43">43</xref>]).</p>
      <p>The contribution is an intervention ready for validation rather than proof of long-term effectiveness, and its limitations follow. The model has been designed but not tested in authentic classrooms; the real-time classification performed by a stochastic generative model may be unreliable; a pilot cannot demonstrate long-term effectiveness; the number of randomized groups is small, so the pilot is designed to bound design parameters such as the intraclass correlation and the design effect rather than estimate them precisely or test efficacy; both conditions are taught in the same room within the same session, so contamination between conditions cannot be excluded and would attenuate any observed difference; and novelty, group-work effects, model variability, sensor error, and teacher influence may all complicate results. The next steps are expert validation and the pilot with an active control condition, assessing feasibility and the proposed mechanism: whether data-grounded dialogue helps pupils construct their own scientific explanations without AI. If it does, GenAI should function not as an answer provider but as a temporary dialogic scaffold that elicits, probes, and progressively transfers reasoning responsibility to the learner.</p>
    </sec>
    <sec id="sec6">
      <title>Appendix A: AI-DRS-SP v1.0: System Prompt, Verbatim</title>
      <p>The text below is the complete system prompt used in the reference implementation described in Section 3.4. It is reproduced without modification so that the dialogue behavior reported in any subsequent empirical study can be replicated and audited. The bracketed headings mark the twelve blocks and are part of the prompt; the target explanation and the dataset named in Blocks 4 and 7 are inserted at run time from the group’s own record.</p>
      <p>[BLOCK 1 - ROLE AND AUDIENCE]</p>
      <p>You are a science-inquiry partner for pupils aged 10 to 12 working in a small group in a primary school classroom. You are not a teacher, not an assessor, and not an encyclopedia. Your only purpose is to help this group turn their own soil-moisture measurements into a claim, evidence, and reasoning. A human teacher is present in the room at all times and can see every turn of this conversation.</p>
      <p>[BLOCK 2 - REGISTER]</p>
      <p>Write one question per turn. Never more than one. Use no more than 45 words. Use short sentences and everyday words. Do not use any technical term that does not already appear in the pupils’ worksheet or in the target explanation given in Block 7. Never ask a multi-part question. Never use bullet points, headings, or emoji.</p>
      <p>[BLOCK 3 - NEVER-ANSWER RULE]</p>
      <p>You must never state the claim, the evidence, or the reasoning on the pupils’ behalf, in whole or in part, even if the pupils ask you to, even if they say they are stuck, and even if they have already given a nearly complete answer. You must never confirm a claim that the pupils have not yet justified with their own data. You may not supply the mechanism. You may not summarize the pupils’ explanation back to them as if it were finished. If the pupils ask you directly for the answer, reply with a question that returns them to their own data.</p>
      <p>[BLOCK 4 - DATA GROUNDING]</p>
      <p>Every substantive move you make must refer to the group’s own dataset, which is supplied to you in the input payload. Do not accept a general statement such as “it dried out more” or “sand is drier”. Require at least two specific readings, with their times and containers, before you allow the pupils to move from evidence to reasoning. If the pupils cite a reading that does not appear in their dataset, ask them to check their table; do not correct the number for them.</p>
      <p>[BLOCK 5 - EPISTEMIC GATEKEEPING]</p>
      <p>Hold the pupils at the current step until its requirement is met.</p>
      <p>Evidence claim: the pupils must have stated a claim that is testable with their dataset and that names the comparison, not just an outcome.</p>
      <p>Evidence to reasoning: the pupils must have cited at least two specific readings from each container that are relevant to the claim.</p>
      <p>Reasoning to completion: the pupils must have connected the stated difference in the rate of decline to a physical mechanism in their own words.</p>
      <p>Do not advance the phase yourself. Advancing is signaled by the application, not by you.</p>
      <p>[BLOCK 6 - CLASSIFICATION]</p>
      <p>Classify the pupils’ most recent response as exactly one of: complete_independent, partially_adequate, general_no_data, data_no_mechanism, unsuccessful_attempt, unsafe_misconception, unreliable_data, safeguarding_or_off_task. This classification is used only to choose your next instructional move. It is never an assessment of the pupils and is never shown to them. If you are not confident, classify as unsuccessful_attempt and reformulate; the application raises an amber flag after a second consecutive unsuccessful attempt, as set out in Block 8, P8.</p>
      <p>[BLOCK 7 - TARGET EXPLANATION]</p>
      <p>The target explanation for this investigation is fixed and is supplied in the input payload as target_explanation. You must not accept an explanation that contradicts it, and you must not extend the pupils’ claim beyond the bounds it states. The current target is: sandy soil has larger particles and larger pore spaces and little organic matter, so water drains through it faster and more of it is exposed to the air; the finer particles and the organic matter in compost hold water against drainage and evaporation. Out-of-scope claims, which you must not endorse or encourage, are: which soil is better for plants; effects on plant growth or health; the index as an absolute amount of water; behavior beyond 40 minutes or after re-watering; and causal claims about light, temperature, or container size.</p>
      <p>[BLOCK 8 - CONTINGENCY RULES AND PRECEDENCE]</p>
      <p>Exactly one rule fires per turn. Evaluate in this order and stop at the first match.</p>
      <p>P1 safeguarding (R9): if the response contains a personal disclosure, signs of distress, or unsafe or off-task content, stop the dialogue, return the neutral holding message “Thank you. Let’s pause here and ask your teacher to come over.”, set flag to “red”, and issue no further question.</p>
      <p>P2 data quality (R8): if the payload marks the run as drift-flagged or run-level-flagged, halt interpretation, issue no claim, evidence, or reasoning move, set flag to “amber”, and tell the pupils that the teacher will check the equipment.</p>
      <p>P3 scientific accuracy (R7): if the response contains a scientifically unsafe misconception, give one brief explicit correction of the misconception only, follow it with one re-check question, and set flag to “red”.</p>
      <p>P4 gatekeeping (R2): if the prerequisite for the next step is missing, hold the pupils at the current step with one focused prompt for the missing element.</p>
      <p>P5 data grounding (R3): if the response is general with no specific values, ask for two or three specific readings from the pupils’ own table.</p>
      <p>P6 conceptual bridging (R4): if data are cited but no mechanism is offered, ask one why-or-how question directed at the target mechanism, without naming the mechanism.</p>
      <p>P7 fading (R1): if the epistemic practice was performed independently, acknowledge briefly, and reduce the prompt level by one.</p>
      <p>P8 default (R5, then R6): if classification confidence is low, or two rules of equal priority match, reformulate the previous question in simpler language with the same goal. On a second consecutive unsuccessful attempt, offer a bounded guided choice of two options, neither of which states the mechanism, and set flag to “amber”.</p>
      <p>[BLOCK 9 - FADING LEVELS]</p>
      <p>fade_level 3, direct prompt: ask the full question, naming the step.</p>
      <p>fade_level 2, delayed prompt: wait one turn; ask only “What would you add to that?”.</p>
      <p>fade_level 1, minimal prompt: acknowledge only; ask nothing.</p>
      <p>fade_level 0, no prompt: return an empty message.</p>
      <p>Reduce fade_level by one when rule R1 fires. Increase it by one, to a maximum of 3, if the pupils’ next response falls below the level of the previous one. Never reduce fade_level by more than one step in a turn.</p>
      <p>[BLOCK 10 - UNCERTAINTY AND CALIBRATED TRUST]</p>
      <p>You can be wrong. At least once per phase, ask the pupils to test what you have said against their own graph, for example: “Check my suggestion against your graph. What evidence would show that it is wrong?” Never claim certainty about the pupils’ data. If you cannot tell what the data show, say so and ask the pupils to look.</p>
      <p>[BLOCK 11 - REFUSAL AND SAFEGUARDING]</p>
      <p>Do not discuss anything other than this investigation. Do not answer questions about yourself, about other pupils, about health, medicine, family, or personal circumstances, or about topics outside the lesson. Do not give advice of any kind. If any such content appears, apply P1 immediately: stop, return the holding message, set flag to “red”, and end the exchange. Do not repeat back any name, address, school, or other identifying detail that appears in a pupil’s text.</p>
      <p>[BLOCK 12 - OUTPUT FORMAT]</p>
      <p>Return a single JSON object and nothing else, conforming to this schema:</p>
      <p>{"classification": string, "rule_fired": string, "move_type": string, "message": string, "fade_level": integer, "flag": string}</p>
      <p>classification is one of the eight values in Block 6. rule_fired is one of R1-R9. move_type is one of acknowledge, focused_prompt, data_grounding, conceptual_bridge, reformulation, guided_choice, correction, halt, stop. message is the text shown to the pupils and obeys Block 2. fade_level is 0 to 3. flag is one of green, amber, red. Return nothing outside this object. If the object you return fails schema validation, the application discards it and substitutes the deterministic fallback move server-side; no unvalidated generation is ever shown to a pupil.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Alexander, R. (2018). Developing Dialogic Teaching: Genesis, Process, Trial. <italic>Research</italic><italic>Papers</italic><italic>in</italic><italic>Education,</italic><italic>33,</italic> 561-598. https://doi.org/10.1080/02671522.2018.1481140 <pub-id pub-id-type="doi">10.1080/02671522.2018.1481140</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1080/02671522.2018.1481140">https://doi.org/10.1080/02671522.2018.1481140</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Alexander, R.</string-name>
              <string-name>Genesis, P</string-name>
            </person-group>
            <year>2018</year>
            <pub-id pub-id-type="doi">10.1080/02671522.2018.1481140</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="web">Anthropic (2025). <italic>Consumer Terms of Service</italic>. https://www.anthropic.com/legal/terms</mixed-citation>
          <element-citation publication-type="web">
            <year>2025</year>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Bai, S. T., Yeung, S. S., &amp; Lo, C. K. (2026). Enhancing the Effect of AI-Assisted Learning: The Use of Scaffolding Strategies to Develop Students’ Prompt Engineering Skills. <italic>Interactive</italic><italic>Learning</italic><italic>Environments,</italic> 1-22. https://doi.org/10.1080/10494820.2026.2649547 <pub-id pub-id-type="doi">10.1080/10494820.2026.2649547</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1080/10494820.2026.2649547">https://doi.org/10.1080/10494820.2026.2649547</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Bai, S.</string-name>
              <string-name>Yeung, S.</string-name>
              <string-name>Lo, C.</string-name>
            </person-group>
            <year>2026</year>
            <pub-id pub-id-type="doi">10.1080/10494820.2026.2649547</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Bell, R. M., &amp; McCaffrey, D. F. (2002). Bias Reduction in Standard Errors for Linear Regression with Multi-Stage Samples. <italic>Survey Methodology, 28,</italic> 169-181.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Bell, R.</string-name>
              <string-name>McCaffrey, D.</string-name>
            </person-group>
            <year>2002</year>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Cai, L., Msafiri, M. M., &amp; Kangwa, D. (2025). Exploring the Impact of Integrating AI Tools in Higher Education Using the Zone of Proximal Development. <italic>Education</italic><italic>and</italic><italic>Information</italic><italic>Technologies,</italic><italic>30,</italic> 7191-7264. https://doi.org/10.1007/s10639-024-13112-0 <pub-id pub-id-type="doi">10.1007/s10639-024-13112-0</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s10639-024-13112-0">https://doi.org/10.1007/s10639-024-13112-0</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Cai, L.</string-name>
              <string-name>Msafiri, M.</string-name>
              <string-name>Kangwa, D.</string-name>
            </person-group>
            <year>2025</year>
            <pub-id pub-id-type="doi">10.1007/s10639-024-13112-0</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B6">
        <label>6.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Cameron, A. C., Gelbach, J. B., &amp; Miller, D. L. (2008). Bootstrap-Based Improvements for Inference with Clustered Errors. <italic>Review</italic><italic>of</italic><italic>Economics</italic><italic>and</italic><italic>Statistics,</italic><italic>90,</italic> 414-427. https://doi.org/10.1162/rest.90.3.414 <pub-id pub-id-type="doi">10.1162/rest.90.3.414</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1162/rest.90.3.414">https://doi.org/10.1162/rest.90.3.414</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Cameron, A.</string-name>
              <string-name>Gelbach, J.</string-name>
              <string-name>Miller, D.</string-name>
            </person-group>
            <year>2008</year>
            <pub-id pub-id-type="doi">10.1162/rest.90.3.414</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B7">
        <label>7.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Cao, J., Chan, S. W. T., Garbett, D. L., Denny, P., Nassani, A., Scholl, P. M. et al. (2021). Sensor-Based Interactive Worksheets to Support Guided Scientific Inquiry. In <italic>Proceedings of the 20th Annual ACM Interaction Design and Children Conference</italic> (pp. 1-7). ACM. https://doi.org/10.1145/3459990.3460716 <pub-id pub-id-type="doi">10.1145/3459990.3460716</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3459990.3460716">https://doi.org/10.1145/3459990.3460716</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Cao, J.</string-name>
              <string-name>Chan, S.</string-name>
              <string-name>Garbett, D.</string-name>
              <string-name>Denny, P.</string-name>
              <string-name>Nassani, A.</string-name>
              <string-name>Scholl, P.</string-name>
            </person-group>
            <year>2021</year>
            <pub-id pub-id-type="doi">10.1145/3459990.3460716</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B8">
        <label>8.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Cooper, G. (2023). Examining Science Education in ChatGPT: An Exploratory Study of Generative Artificial Intelligence. <italic>Journal</italic><italic>of</italic><italic>Science</italic><italic>Education</italic><italic>and</italic><italic>Technology,</italic><italic>32,</italic> 444-452. https://doi.org/10.1007/s10956-023-10039-y <pub-id pub-id-type="doi">10.1007/s10956-023-10039-y</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s10956-023-10039-y">https://doi.org/10.1007/s10956-023-10039-y</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Cooper, G.</string-name>
            </person-group>
            <year>2023</year>
            <pub-id pub-id-type="doi">10.1007/s10956-023-10039-y</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B9">
        <label>9.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Diola, W. Y., Jalon Jr., J. B., &amp; Prudente, M. S. (2025). Exploring the Use of Claim-Evidence-Reasoning in Promoting Scientific Reasoning Skills of Elementary School Students. <italic>Anatolian</italic><italic>Journal</italic><italic>of</italic><italic>Education,</italic><italic>10,</italic> 203-214. https://doi.org/10.29333/aje.2025.10115a <pub-id pub-id-type="doi">10.29333/aje.2025.10115a</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.29333/aje.2025.10115a">https://doi.org/10.29333/aje.2025.10115a</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Diola, W.</string-name>
              <string-name>Prudente, M.</string-name>
            </person-group>
            <year>2025</year>
            <pub-id pub-id-type="doi">10.29333/aje.2025.10115a</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B10">
        <label>10.</label>
        <mixed-citation publication-type="web">ELECFREAKS (n.d.). <italic>Nezha Inventor’s Kit for Micro:Bit: Product Documentation and Learning Cases</italic>. https://wiki.elecfreaks.com/en/microbit/building-blocks/nezha-inventors-kit/</mixed-citation>
      </ref>
      <ref id="B11">
        <label>11.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Hedges, L. V. (2007). Effect Sizes in Cluster-Randomized Designs. <italic>Journal</italic><italic>of</italic><italic>Educational</italic><italic>and</italic><italic>Behavioral</italic><italic>Statistics,</italic><italic>32,</italic> 341-370. https://doi.org/10.3102/1076998606298043 <pub-id pub-id-type="doi">10.3102/1076998606298043</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3102/1076998606298043">https://doi.org/10.3102/1076998606298043</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Hedges, L.</string-name>
            </person-group>
            <year>2007</year>
            <pub-id pub-id-type="doi">10.3102/1076998606298043</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B12">
        <label>12.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Kalelioğlu, F., &amp; Sentance, S. (2020). Teaching with Physical Computing in School: The Case of the Micro:Bit. <italic>Education</italic><italic>and</italic><italic>Information</italic><italic>Technologies,</italic><italic>25,</italic> 2577-2603. https://doi.org/10.1007/s10639-019-10080-8 <pub-id pub-id-type="doi">10.1007/s10639-019-10080-8</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s10639-019-10080-8">https://doi.org/10.1007/s10639-019-10080-8</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Sentance, S.</string-name>
            </person-group>
            <year>2020</year>
            <pub-id pub-id-type="doi">10.1007/s10639-019-10080-8</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B13">
        <label>13.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Kalogiannakis, M., Ampartzaki, M., Papadakis, S., &amp; Skaraki, E. (2018). Teaching Natural Science Concepts to Young Children with Mobile Devices and Hands-On Activities. A Case Study. <italic>International</italic><italic>Journal</italic><italic>of</italic><italic>Teaching</italic><italic>and</italic><italic>Case</italic><italic>Studies,</italic><italic>9,</italic> 171-183. https://doi.org/10.1504/ijtcs.2018.090965 <pub-id pub-id-type="doi">10.1504/ijtcs.2018.090965</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1504/ijtcs.2018.090965">https://doi.org/10.1504/ijtcs.2018.090965</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Kalogiannakis, M.</string-name>
              <string-name>Ampartzaki, M.</string-name>
              <string-name>Papadakis, S.</string-name>
              <string-name>Skaraki, E.</string-name>
            </person-group>
            <year>2018</year>
            <pub-id pub-id-type="doi">10.1504/ijtcs.2018.090965</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B14">
        <label>14.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Kalogiannakis, M., Papakonstantinou, N., &amp; Sotiropoulos, D. (2025). From Support Tool to Learning Partner: A Systematic Review of GenAI Integration in University Science Labs. <italic>Creative</italic><italic>Education,</italic><italic>16,</italic> 1364-1401. https://doi.org/10.4236/ce.2025.169083 <pub-id pub-id-type="doi">10.4236/ce.2025.169083</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.4236/ce.2025.169083">https://doi.org/10.4236/ce.2025.169083</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Kalogiannakis, M.</string-name>
              <string-name>Papakonstantinou, N.</string-name>
              <string-name>Sotiropoulos, D.</string-name>
            </person-group>
            <year>2025</year>
            <pub-id pub-id-type="doi">10.4236/ce.2025.169083</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B15">
        <label>15.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Lee, H.-P., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R. et al. (2025). The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects from a Survey of Knowledge Workers. In Association for Computing Machinery (Ed.), <italic>Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems</italic> (pp. 1-22). ACM. https://doi.org/10.1145/3706598.3713778 <pub-id pub-id-type="doi">10.1145/3706598.3713778</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3706598.3713778">https://doi.org/10.1145/3706598.3713778</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Lee, H.</string-name>
              <string-name>Sarkar, A.</string-name>
              <string-name>Tankelevitch, L.</string-name>
              <string-name>Drosos, I.</string-name>
              <string-name>Rintel, S.</string-name>
              <string-name>Banks, R.</string-name>
            </person-group>
            <year>2025</year>
            <pub-id pub-id-type="doi">10.1145/3706598.3713778</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B16">
        <label>16.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Lin, X.-F., Hwang, G.-J., Wang, J., Zhou, Y., Li, W., Liu, J., &amp; Liang, Z.-M. (2023). Effects of a Contextualised Reflective Mechanism-Based Augmented Reality Learning Model on Students’ Scientific Inquiry Learning Performances, Behavioural Patterns, and Higher Order Thinking. <italic>Interactive</italic><italic>Learning</italic><italic>Environments,</italic><italic>31,</italic> 6931-6951. https://doi.org/10.1080/10494820.2022.2057546 <pub-id pub-id-type="doi">10.1080/10494820.2022.2057546</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1080/10494820.2022.2057546">https://doi.org/10.1080/10494820.2022.2057546</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Lin, X.</string-name>
              <string-name>Hwang, G.</string-name>
              <string-name>Wang, J.</string-name>
              <string-name>Zhou, Y.</string-name>
              <string-name>Li, W.</string-name>
              <string-name>Liu, J.</string-name>
              <string-name>Liang, Z.</string-name>
              <string-name>Performances, B</string-name>
            </person-group>
            <year>2023</year>
            <pub-id pub-id-type="doi">10.1080/10494820.2022.2057546</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B17">
        <label>17.</label>
        <citation-alternatives>
          <mixed-citation publication-type="thesis">Masters, H., &amp; Docktor, J. (2022). Preservice Teachers’ Abilities and Confidence with Constructing Scientific Explanations as Scaffolds Are Faded in a Physics Course for Educators. <italic>Journal</italic><italic>of</italic><italic>Science</italic><italic>Teacher</italic><italic>Education,</italic><italic>33,</italic> 786-813. https://doi.org/10.1080/1046560x.2021.2004641 <pub-id pub-id-type="doi">10.1080/1046560x.2021.2004641</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1080/1046560x.2021.2004641">https://doi.org/10.1080/1046560x.2021.2004641</ext-link></mixed-citation>
          <element-citation publication-type="thesis">
            <person-group person-group-type="author">
              <string-name>Masters, H.</string-name>
              <string-name>Docktor, J.</string-name>
            </person-group>
            <year>2022</year>
            <pub-id pub-id-type="doi">10.1080/1046560x.2021.2004641</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B18">
        <label>18.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">McNeill, K. L., &amp; Krajcik, J. S. (2012). <italic>Supporting Grade 5-8 Students in Constructing Explanations in Science: The Claim, Evidence, and Reasoning Framework for Talk and Writing</italic>. Pearson.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>McNeill, K.</string-name>
              <string-name>Krajcik, J.</string-name>
              <string-name>Claim, E</string-name>
            </person-group>
            <year>2012</year>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B19">
        <label>19.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">McNeill, K. L., Lizotte, D. J., Krajcik, J., &amp; Marx, R. W. (2006). Supporting Students’ Construction of Scientific Explanations by Fading Scaffolds in Instructional Materials. <italic>Journal</italic><italic>of</italic><italic>the</italic><italic>Learning</italic><italic>Sciences,</italic><italic>15,</italic> 153-191. https://doi.org/10.1207/s15327809jls1502_1 <pub-id pub-id-type="doi">10.1207/s15327809jls1502_1</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1207/s15327809jls1502_1">https://doi.org/10.1207/s15327809jls1502_1</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>McNeill, K.</string-name>
              <string-name>Lizotte, D.</string-name>
              <string-name>Krajcik, J.</string-name>
              <string-name>Marx, R.</string-name>
            </person-group>
            <year>2006</year>
            <pub-id pub-id-type="doi">10.1207/s15327809jls1502_1</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B20">
        <label>20.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">McNeish, D. M., &amp; Stapleton, L. M. (2016). The Effect of Small Sample Size on Two-Level Model Estimates: A Review and Illustration. <italic>Educational</italic><italic>Psychology</italic><italic>Review,</italic><italic>28,</italic> 295-314. https://doi.org/10.1007/s10648-014-9287-x <pub-id pub-id-type="doi">10.1007/s10648-014-9287-x</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s10648-014-9287-x">https://doi.org/10.1007/s10648-014-9287-x</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>McNeish, D.</string-name>
              <string-name>Stapleton, L.</string-name>
            </person-group>
            <year>2016</year>
            <pub-id pub-id-type="doi">10.1007/s10648-014-9287-x</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B21">
        <label>21.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Mishra, P., Warr, M., &amp; Islam, R. (2023). TPACK in the Age of ChatGPT and Generative AI. <italic>Journal</italic><italic>of</italic><italic>Digital</italic><italic>Learning</italic><italic>in</italic><italic>Teacher</italic><italic>Education,</italic><italic>39,</italic> 235-251. https://doi.org/10.1080/21532974.2023.2247480 <pub-id pub-id-type="doi">10.1080/21532974.2023.2247480</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1080/21532974.2023.2247480">https://doi.org/10.1080/21532974.2023.2247480</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Mishra, P.</string-name>
              <string-name>Warr, M.</string-name>
              <string-name>Islam, R.</string-name>
            </person-group>
            <year>2023</year>
            <pub-id pub-id-type="doi">10.1080/21532974.2023.2247480</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B22">
        <label>22.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Movahed, S. V., &amp; Martin, F. G. (2025). Ask Me Anything: Exploring Children’s Attitudes toward an Age-Tailored AI-Powered Chatbot. <italic>International</italic><italic>Journal</italic><italic>of</italic><italic>Artificial</italic><italic>Intelligence</italic><italic>in</italic><italic>Education,</italic><italic>35,</italic> 3979-4001. https://doi.org/10.1007/s40593-025-00523-4 <pub-id pub-id-type="doi">10.1007/s40593-025-00523-4</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s40593-025-00523-4">https://doi.org/10.1007/s40593-025-00523-4</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Movahed, S.</string-name>
              <string-name>Martin, F.</string-name>
            </person-group>
            <year>2025</year>
            <pub-id pub-id-type="doi">10.1007/s40593-025-00523-4</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B23">
        <label>23.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Muhonen, H., Rasku-Puttonen, H., Pakarinen, E., Poikkeus, A., &amp; Lerkkanen, M. (2016). Scaffolding through Dialogic Teaching in Early School Classrooms. <italic>Teaching</italic><italic>and</italic><italic>Teacher</italic><italic>Education,</italic><italic>55,</italic> 143-154. https://doi.org/10.1016/j.tate.2016.01.007 <pub-id pub-id-type="doi">10.1016/j.tate.2016.01.007</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.tate.2016.01.007">https://doi.org/10.1016/j.tate.2016.01.007</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Muhonen, H.</string-name>
              <string-name>Rasku-Puttonen, H.</string-name>
              <string-name>Pakarinen, E.</string-name>
              <string-name>Poikkeus, A.</string-name>
              <string-name>Lerkkanen, M.</string-name>
            </person-group>
            <year>2016</year>
            <pub-id pub-id-type="doi">10.1016/j.tate.2016.01.007</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B24">
        <label>24.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Ng, D. T. K., Su, J., Leung, J. K. L., &amp; Chu, S. K. W. (2024). Artificial Intelligence (AI) Literacy Education in Secondary Schools: A Review. <italic>Interactive Learning Environments, 32,</italic> 6204-6224. https://doi.org/10.1080/10494820.2023.2255228 <pub-id pub-id-type="doi">10.1080/10494820.2023.2255228</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1080/10494820.2023.2255228">https://doi.org/10.1080/10494820.2023.2255228</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Ng, D.</string-name>
              <string-name>Su, J.</string-name>
              <string-name>Leung, J.</string-name>
              <string-name>Chu, S.</string-name>
            </person-group>
            <year>2024</year>
            <pub-id pub-id-type="doi">10.1080/10494820.2023.2255228</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B25">
        <label>25.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Ntourou, V., Kalogiannakis, M., &amp; Psycharis, S. (2021). A Study of the Impact of Arduino and Visual Programming in Self-Efficacy, Motivation, Computational Thinking and 5th Grade Students’ Perceptions on Electricity. <italic>Eurasia</italic><italic>Journal</italic><italic>of</italic><italic>Mathematics,</italic><italic>Science</italic><italic>and</italic><italic>Technology</italic><italic>Education,</italic><italic>17,</italic> em1960. https://doi.org/10.29333/ejmste/10842 <pub-id pub-id-type="doi">10.29333/ejmste/10842</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.29333/ejmste/10842">https://doi.org/10.29333/ejmste/10842</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Ntourou, V.</string-name>
              <string-name>Kalogiannakis, M.</string-name>
              <string-name>Psycharis, S.</string-name>
              <string-name>Self-Efficacy, M</string-name>
              <string-name>Mathematics, S</string-name>
            </person-group>
            <year>2021</year>
            <pub-id pub-id-type="doi">10.29333/ejmste/10842</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B26">
        <label>26.</label>
        <citation-alternatives>
          <mixed-citation publication-type="web">OpenAI (2025). <italic>Terms of Use</italic>. https://openai.com/policies/row-terms-of-use/</mixed-citation>
          <element-citation publication-type="web">
            <year>2025</year>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B27">
        <label>27.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Petousi, V., &amp; Sifaki, E. (2020). Contextualising Harm in the Framework of Research Misconduct. Findings from Discourse Analysis of Scientific Publications. <italic>International</italic><italic>Journal</italic><italic>of</italic><italic>Sustainable</italic><italic>Development,</italic><italic>23,</italic> 149-174. https://doi.org/10.1504/ijsd.2020.115206 <pub-id pub-id-type="doi">10.1504/ijsd.2020.115206</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1504/ijsd.2020.115206">https://doi.org/10.1504/ijsd.2020.115206</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Petousi, V.</string-name>
              <string-name>Sifaki, E.</string-name>
            </person-group>
            <year>2020</year>
            <pub-id pub-id-type="doi">10.1504/ijsd.2020.115206</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B28">
        <label>28.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Przybylla, M., &amp; Romeike, R. (2014). Physical Computing and Its Scope—Towards a Constructionist Computer Science Curriculum with Physical Computing. <italic>Informatics</italic><italic>in</italic><italic>Education,</italic><italic>13,</italic> 225-240. https://doi.org/10.15388/infedu.2014.14 <pub-id pub-id-type="doi">10.15388/infedu.2014.14</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.15388/infedu.2014.14">https://doi.org/10.15388/infedu.2014.14</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Przybylla, M.</string-name>
              <string-name>Romeike, R.</string-name>
            </person-group>
            <year>2014</year>
            <pub-id pub-id-type="doi">10.15388/infedu.2014.14</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B29">
        <label>29.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Psycharis, S., Kalovrektis, K., &amp; Xenakis, A. (2020). A Conceptual Framework for Computational Pedagogy in STEAM Education: Determinants and Perspectives. <italic>Hellenic</italic><italic>Jour</italic><italic>nal</italic><italic>of</italic><italic>STEM</italic><italic>Education,</italic><italic>1,</italic> 17-32. https://doi.org/10.51724/hjstemed.v1i1.4 <pub-id pub-id-type="doi">10.51724/hjstemed.v1i1.4</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.51724/hjstemed.v1i1.4">https://doi.org/10.51724/hjstemed.v1i1.4</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Psycharis, S.</string-name>
              <string-name>Kalovrektis, K.</string-name>
              <string-name>Xenakis, A.</string-name>
            </person-group>
            <year>2020</year>
            <pub-id pub-id-type="doi">10.51724/hjstemed.v1i1.4</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B30">
        <label>30.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Sentance, S., Waite, J., Hodges, S., MacLeod, E., &amp; Yeomans, L. (2017). “Creating Cool Stuff”: Pupils’ Experience of the BBC Micro:Bit. In <italic>Proceedings of the 2017 ACM</italic><italic>SIGCSE Technical Symposium on Computer Science Education</italic> (pp. 531-536). Association for Computing Machinery. https://doi.org/10.1145/3017680.3017749 <pub-id pub-id-type="doi">10.1145/3017680.3017749</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3017680.3017749">https://doi.org/10.1145/3017680.3017749</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Sentance, S.</string-name>
              <string-name>Waite, J.</string-name>
              <string-name>Hodges, S.</string-name>
              <string-name>MacLeod, E.</string-name>
              <string-name>Yeomans, L.</string-name>
            </person-group>
            <year>2017</year>
            <pub-id pub-id-type="doi">10.1145/3017680.3017749</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B31">
        <label>31.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Sotiropoulos, D., Xenakis, A., Kalogiannakis, M., &amp; Taşar, M. F. (2026). The Interactive Design Process Framework (IDPF): Utilizing GenAI as a Collaborative Agent for Creating STEAM Projects. <italic>Hellenic</italic><italic>Journal</italic><italic>of</italic><italic>STEM</italic><italic>Education,</italic><italic>5,</italic> 1-18. https://doi.org/10.51724/hjstemed.v5i1.76 <pub-id pub-id-type="doi">10.51724/hjstemed.v5i1.76</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.51724/hjstemed.v5i1.76">https://doi.org/10.51724/hjstemed.v5i1.76</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Sotiropoulos, D.</string-name>
              <string-name>Xenakis, A.</string-name>
              <string-name>Kalogiannakis, M.</string-name>
            </person-group>
            <year>2026</year>
            <pub-id pub-id-type="doi">10.51724/hjstemed.v5i1.76</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B32">
        <label>32.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Spasopoulos, T., Sotiropoulos, D., &amp; Kalogiannakis, M. (2025). Generative AI in Pre-Service Science Teacher Education: A Systematic Review. <italic>Advances</italic><italic>in</italic><italic>Mobile</italic><italic>Learning</italic><italic>Educational</italic><italic>Research,</italic><italic>5,</italic> 1501-1523. https://doi.org/10.25082/amler.2025.02.007 <pub-id pub-id-type="doi">10.25082/amler.2025.02.007</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.25082/amler.2025.02.007">https://doi.org/10.25082/amler.2025.02.007</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Spasopoulos, T.</string-name>
              <string-name>Sotiropoulos, D.</string-name>
              <string-name>Kalogiannakis, M.</string-name>
            </person-group>
            <year>2025</year>
            <pub-id pub-id-type="doi">10.25082/amler.2025.02.007</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B33">
        <label>33.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Tang, K.-S., &amp; Putra, G. B. S. (2026). Generative AI as a Dialogic Partner: Enhancing Multiple Perspectives, Reasoning, and Argumentation in Science Education with Customized Chatbots. <italic>Journal</italic><italic>of</italic><italic>Science</italic><italic>Education</italic><italic>and</italic><italic>Technology,</italic><italic>35,</italic> 128-140. https://doi.org/10.1007/s10956-025-10240-1 <pub-id pub-id-type="doi">10.1007/s10956-025-10240-1</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s10956-025-10240-1">https://doi.org/10.1007/s10956-025-10240-1</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Tang, K.</string-name>
              <string-name>Putra, G.</string-name>
              <string-name>Perspectives, R</string-name>
            </person-group>
            <year>2026</year>
            <pub-id pub-id-type="doi">10.1007/s10956-025-10240-1</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B34">
        <label>34.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">van de Pol, J., Volman, M., &amp; Beishuizen, J. (2010). Scaffolding in Teacher-Student Interaction: A Decade of Research. <italic>Educational</italic><italic>Psychology</italic><italic>Review,</italic><italic>22,</italic> 271-296. https://doi.org/10.1007/s10648-010-9127-6 <pub-id pub-id-type="doi">10.1007/s10648-010-9127-6</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s10648-010-9127-6">https://doi.org/10.1007/s10648-010-9127-6</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Pol, J.</string-name>
              <string-name>Volman, M.</string-name>
              <string-name>Beishuizen, J.</string-name>
            </person-group>
            <year>2010</year>
            <pub-id pub-id-type="doi">10.1007/s10648-010-9127-6</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B35">
        <label>35.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Vrikki, M., Wheatley, L., Howe, C., Hennessy, S., &amp; Mercer, N. (2019). Dialogic Practices in Primary School Classrooms. <italic>Language</italic><italic>and</italic><italic>Education,</italic><italic>33,</italic> 85-100. https://doi.org/10.1080/09500782.2018.1509988 <pub-id pub-id-type="doi">10.1080/09500782.2018.1509988</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1080/09500782.2018.1509988">https://doi.org/10.1080/09500782.2018.1509988</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Vrikki, M.</string-name>
              <string-name>Wheatley, L.</string-name>
              <string-name>Howe, C.</string-name>
              <string-name>Hennessy, S.</string-name>
              <string-name>Mercer, N.</string-name>
            </person-group>
            <year>2019</year>
            <pub-id pub-id-type="doi">10.1080/09500782.2018.1509988</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B36">
        <label>36.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Wan Hamedi, W. H., Awang Ali, F. D., Abdullah, W. Y., Ab Hamid, H., Mohammad Shuhaimi, N. I., &amp; Mohamad Amir, M. (2025). AI as a Digital Scaffold: An Integrative Review of Vygotsky’s Zone of Proximal Development in Modern Education. <italic>International</italic><italic>Journal</italic><italic>of</italic><italic>Modern</italic><italic>Education,</italic><italic>7,</italic> 579-589. https://doi.org/10.35631/ijmoe.726038 <pub-id pub-id-type="doi">10.35631/ijmoe.726038</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.35631/ijmoe.726038">https://doi.org/10.35631/ijmoe.726038</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Hamedi, W.</string-name>
              <string-name>Ali, F.</string-name>
              <string-name>Abdullah, W.</string-name>
              <string-name>Hamid, H.</string-name>
              <string-name>Shuhaimi, N.</string-name>
              <string-name>Amir, M.</string-name>
            </person-group>
            <year>2025</year>
            <pub-id pub-id-type="doi">10.35631/ijmoe.726038</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B37">
        <label>37.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Wegerif, R., &amp; Casebourne, I. (2026). A Dialogic Theoretical Foundation for Integrating Generative AI into Pedagogical Design. <italic>British</italic><italic>Journal</italic><italic>of</italic><italic>Educational</italic><italic>Technology,</italic><italic>57,</italic> 639-654. https://doi.org/10.1111/bjet.70026 <pub-id pub-id-type="doi">10.1111/bjet.70026</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1111/bjet.70026">https://doi.org/10.1111/bjet.70026</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Wegerif, R.</string-name>
              <string-name>Casebourne, I.</string-name>
            </person-group>
            <year>2026</year>
            <pub-id pub-id-type="doi">10.1111/bjet.70026</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B38">
        <label>38.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Wood, D., Bruner, J. S., &amp; Ross, G. (1976). The Role of Tutoring in Problem Solving. <italic>Journal</italic><italic>of</italic><italic>Child</italic><italic>Psychology</italic><italic>and</italic><italic>Psychiatry,</italic><italic>17,</italic> 89-100. https://doi.org/10.1111/j.1469-7610.1976.tb00381.x <pub-id pub-id-type="doi">10.1111/j.1469-7610.1976.tb00381.x</pub-id><pub-id pub-id-type="pmid">932126</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1111/j.1469-7610.1976.tb00381.x">https://doi.org/10.1111/j.1469-7610.1976.tb00381.x</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Wood, D.</string-name>
              <string-name>Bruner, J.</string-name>
              <string-name>Ross, G.</string-name>
            </person-group>
            <year>1976</year>
            <pub-id pub-id-type="doi">10.1111/j.1469-7610.1976.tb00381.x</pub-id>
            <pub-id pub-id-type="pmid">932126</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B39">
        <label>39.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Wu, R., &amp; Yu, Z. (2024). Do AI Chatbots Improve Students Learning Outcomes? Evidence from a Meta-Analysis. <italic>British Journal of Educational Technology, 55,</italic> 10-33. https://doi.org/10.1111/bjet.13334 <pub-id pub-id-type="doi">10.1111/bjet.13334</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1111/bjet.13334">https://doi.org/10.1111/bjet.13334</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Wu, R.</string-name>
              <string-name>Yu, Z.</string-name>
            </person-group>
            <year>2024</year>
            <pub-id pub-id-type="doi">10.1111/bjet.13334</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B40">
        <label>40.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Xenakis, A. (2025). AI-Based Models for Assessing STEAM Engineering Literacy for University Students: The Case of Digital Systems for Precision Agriculture. <italic>Hellenic Journal of STEM Education, 4,</italic> 1-9. https://doi.org/10.51724/hjstemed.v4i1.37 <pub-id pub-id-type="doi">10.51724/hjstemed.v4i1.37</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.51724/hjstemed.v4i1.37">https://doi.org/10.51724/hjstemed.v4i1.37</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Xenakis, A.</string-name>
            </person-group>
            <year>2025</year>
            <pub-id pub-id-type="doi">10.51724/hjstemed.v4i1.37</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B41">
        <label>41.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Xenakis, A., Dimos, I., Feidakis, M., Sotiropoulos, D., Kalovrektis, K., &amp; Nikolaou, G. (2025). An LLM-Based Smart Repository Platform to Support Educators with Computational Thinking, AI, and STEM Activities. In S. Papadakis, &amp; M. Kalogiannakis (Eds.), <italic>Empowering STEM Educators with Digital Tools</italic> (pp. 107-136). IGI Global. https://doi.org/10.4018/979-8-3693-9806-7.ch005 <pub-id pub-id-type="doi">10.4018/979-8-3693-9806-7.ch005</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.4018/979-8-3693-9806-7.ch005">https://doi.org/10.4018/979-8-3693-9806-7.ch005</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Xenakis, A.</string-name>
              <string-name>Dimos, I.</string-name>
              <string-name>Feidakis, M.</string-name>
              <string-name>Sotiropoulos, D.</string-name>
              <string-name>Kalovrektis, K.</string-name>
              <string-name>Nikolaou, G.</string-name>
              <string-name>Thinking, A</string-name>
            </person-group>
            <year>2025</year>
            <pub-id pub-id-type="doi">10.4018/979-8-3693-9806-7.ch005</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B42">
        <label>42.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Yan, L., Greiff, S., Teuber, Z., &amp; Gašević, D. (2024). Promises and Challenges of Generative Artificial Intelligence for Human Learning. <italic>Nature Human Behaviour, 8,</italic> 1839-1850. https://doi.org/10.1038/s41562-024-02004-5 <pub-id pub-id-type="doi">10.1038/s41562-024-02004-5</pub-id><pub-id pub-id-type="pmid">39438686</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41562-024-02004-5">https://doi.org/10.1038/s41562-024-02004-5</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Yan, L.</string-name>
              <string-name>Greiff, S.</string-name>
              <string-name>Teuber, Z.</string-name>
            </person-group>
            <year>2024</year>
            <pub-id pub-id-type="doi">10.1038/s41562-024-02004-5</pub-id>
            <pub-id pub-id-type="pmid">39438686</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B43">
        <label>43.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Yan, L., Pammer-Schindler, V., Mills, C., Nguyen, A., &amp; Gašević, D. (2025). Beyond Efficiency: Empirical Insights on Generative AI’s Impact on Cognition, Metacognition and Epistemic Agency in Learning. <italic>British Journal of Educational Technology, 56,</italic> 1675-1685. https://doi.org/10.1111/bjet.70000 <pub-id pub-id-type="doi">10.1111/bjet.70000</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1111/bjet.70000">https://doi.org/10.1111/bjet.70000</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Yan, L.</string-name>
              <string-name>Pammer-Schindler, V.</string-name>
              <string-name>Mills, C.</string-name>
              <string-name>Nguyen, A.</string-name>
              <string-name>Cognition, M</string-name>
            </person-group>
            <year>2025</year>
            <pub-id pub-id-type="doi">10.1111/bjet.70000</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>