<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">jcc</journal-id>
      <journal-title-group>
        <journal-title>Journal of Computer and Communications</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2327-5227</issn>
      <issn pub-type="ppub">2327-5219</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/jcc.2026.149001</article-id>
      <article-id pub-id-type="publisher-id">jcc-153887</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Computer Science</subject>
          <subject>Communications</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Development of a System for the Recognition of Isolated Signs in Text of the Niger Sign Language (LSNi) Based on Neural Networks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Idi</surname>
            <given-names>Bachir Moussa</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Kadri</surname>
            <given-names>Chaibou</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Alley</surname>
            <given-names>Ibrahim Bouwey</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Ganda</surname>
            <given-names>Yahaya Morou</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Naroua</surname>
            <given-names>Harouna</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> Département de Mathématiques et Informatique, Faculté des Sciences et Techniques, Université Abdou Moumouni, Niamey, Niger </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The authors declare no conflicts of interest regarding the publication of this paper.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>16</day>
        <month>09</month>
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>09</month>
        <year>2026</year>
      </pub-date>
      <volume>14</volume>
      <issue>09</issue>
      <fpage>1</fpage>
      <lpage>9</lpage>
      <history>
        <date date-type="received">
          <day>06</day>
          <month>08</month>
          <year>2026</year>
        </date>
        <date date-type="accepted">
          <day>13</day>
          <month>09</month>
          <year>2026</year>
        </date>
        <date date-type="published">
          <day>16</day>
          <month>09</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2026 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/jcc.2026.149001">https://doi.org/10.4236/jcc.2026.149001</self-uri>
      <abstract>
        <p>This article contributes to improving communication between the deaf community and the rest of the population through the use of digital technology and artificial intelligence techniques, notably Deep Learning, by developing a system capable of recognizing isolated signs in text of Niger Sign Language (NiSL) gestures (words) in real time using a webcam. For this purpose, a video dataset specific to Niger Sign Language was created. Landmarks of articulations were extracted using MediaPipe Holistic [<xref ref-type="bibr" rid="B1">1</xref>] and used to train a Long Short-Term Memory (LSTM) classification model to perform this recognition of isolated signs in text. The results obtained show a good performance of the model, with a recognition rate of 62.22% when only hand characteristics are used, which is comparable to results of previous works in the literature on more documented sign languages.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>Niger Sign Language</kwd>
        <kwd>Automatic Translation</kwd>
        <kwd>Recurrent Neural Networks</kwd>
        <kwd>LSTM</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction</title>
      <p>Communication is a fundamental process in human interaction, enabling the exchange of ideas, information and emotions [<xref ref-type="bibr" rid="B2">2</xref>]. For the deaf community, often deprived of hearing or speech, non-verbal communication, and more specifically sign language, becomes the main vector of interaction.</p>
      <p>In Niger, disabled people represent 4.2% of the population, approximately 715,497 individuals, among whom there are nearly 24,000 deaf people [<xref ref-type="bibr" rid="B3">3</xref>]. The latter use Niger Sign Language (NiSL) to communicate. Sign language is a natural language in its own right, which is based on a visual-gestural modality involving the hands, facial expressions and body movements [<xref ref-type="bibr" rid="B4">4</xref>]. However, in an environment where the spoken language predominates, this community faces social isolation which hinders its integration.</p>
      <p>Technological advances, particularly in artificial intelligence (AI), offer promising solutions to overcome this isolation. Systems based on digital gloves (data gloves) [<xref ref-type="bibr" rid="B5">5</xref>] or computer vision [<xref ref-type="bibr" rid="B6">6</xref>] have been developed for other sign languages. However, NiSL remains underrepresented in academic work and technological applications, suffering from a critical lack of digital resources.</p>
      <p>This work aims to address this challenge by developing a system capable of recognizing NiSL gestures at word level in real time and translating them into text, using a simple webcam.</p>
    </sec>
    <sec id="sec2">
      <title>2. Materials and Methods</title>
      <p>To design this translation system, we adopted a multi-step methodology, ranging from the linguistic study of the NiSL to the validation of the AI model.</p>
      <sec id="sec2dot1">
        <title>2.1. Linguistic Aspects of Sign Language</title>
        <p>Contrary to popular belief, sign languages are not simple gestural transcriptions of spoken languages. They are complex languages with their own phonology, morphology, and syntax. The basic unit, analogous to the phoneme, is the sign, which is composed of several essential parameters. William Stokoe [<xref ref-type="bibr" rid="B4">4</xref>] was the first to identify three of such parameters for American Sign Language (ASL):</p>
        <p>Hand configuration: The shape the hand takes (e.g., closed fist, open hand). Movement: The movement of the hand in space (direction, speed). Location: The place where the sign is produced on the body or in space.</p>
        <p>Other researchers have complemented this model by adding palm orientation and non-manual elements (facial expressions, body movements), which are crucial for nuanced meaning [<xref ref-type="bibr" rid="B7">7</xref>]. Our modeling approach aims to capture this multimodal richness.</p>
      </sec>
      <sec id="sec2dot2">
        <title>2.2. Construction of a Data Set for NiSL</title>
        <p>One of the major challenges of this project was the lack of a digital corpus for the NiSL. To address this, we built our own dataset:</p>
        <p><bold>1)</bold><bold>Source:</bold> We based our work on the Dictionary of Signs in use in Niger, a paper document used in schools for the deaf in Niamey. The corpus was recorded by a single experienced deaf signer, a deaf person from the Niamey School for the Deaf/Hassane Banâ Bâ, a public institution in Niamey. This signer provided all nine original recordings. Using a deaf signer instead of directly reproducing the dictionary illustrations is a methodological choice: the source only has static drawings, and only a skilled signer can convey the movement, rhythm, and articulation that a drawing can’t capture.</p>
        <p><bold>2)</bold><bold>Acquisition:</bold> We recorded videos of 9 distinct signs related to the family estate. Each video was captured in high definition with a Canon LEGRIA HF R806 camera to ensure accurate gesture analysis.</p>
        <p><bold>3)</bold><bold>Data</bold><bold>enrichment</bold>: To increase the robustness and diversity of our initial corpus of 9 videos, we applied data augmentation techniques:</p>
        <p><bold>Horizontal</bold><bold>Flip</bold> [<xref ref-type="bibr" rid="B8">8</xref>]: To simulate the variability between right-handed and left-handed signers, each video was duplicated in a mirror version, bringing the corpus to 18 videos.<bold>Temporal</bold><bold>Variation:</bold> For each video, we generated 15 new instances with variable execution speeds (slowed down and accelerated).</p>
        <p>In the end, the dataset consists of <bold>270</bold><bold>videos</bold> (9 signs × 2 hand variations × 15 speed variation) (see <xref ref-type="fig" rid="fig1">Figure 1</xref>).</p>
        <fig id="fig1">
          <label>Figure 1</label>
          <graphic xlink:href="https://html.scirp.org/file/1733639-rId13.jpeg?20260916024440" />
        </fig>
        <p><bold>Figure 1.</bold> The 9 dictionary signs used.</p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. Modeling and Design</title>
      <p>The heart of this system is a Deep Learning model capable of processing sequential and multimodal data.</p>
      <sec id="sec3dot1">
        <title>3.1. Feature Extraction with MediaPipe Holistic</title>
        <p>To transform the raw video data into a format usable by a neural network, we used Google’s MediaPipe Holistic solution [<xref ref-type="bibr" rid="B1">1</xref>]. This powerful tool allows real time and simultaneous extraction of key points (landmarks) of the body, face and hands from each frame of the video. </p>
        <p>Hands: 21 landmarks per hand (see <xref ref-type="fig" rid="fig2">Figure 2</xref>)Face: 468 landmarks (see <xref ref-type="fig" rid="fig3">Figure 3</xref>)Pose: 33 landmarks (see <xref ref-type="fig" rid="fig4">Figure 4</xref>)</p>
        <p>For each video frame, the feature vectors for the right hand <inline-formula><mml:math display="inline"><mml:mrow><mml:mi> v </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mi> M </mml:mi><mml:mi> d </mml:mi></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> , the left hand <inline-formula><mml:math display="inline"><mml:mrow><mml:mi> v </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mi> M </mml:mi><mml:mi> g </mml:mi></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> , the face <inline-formula><mml:math display="inline"><mml:mrow><mml:mi> v </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mi> F </mml:mi><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> and the pose <inline-formula><mml:math display="inline"><mml:mrow><mml:mi> v </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mi> P </mml:mi><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> are extracted. These vectors are then concatenated to form a single feature vector <inline-formula><mml:math display="inline"><mml:mrow><mml:mi> v </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mi> I </mml:mi><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> for frame.</p>
        <disp-formula id="FD1">
          <label>(1)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:mi>v</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>I</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>=</mml:mo>
              <mml:mi>v</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>M</mml:mi>
                    <mml:mi>d</mml:mi>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>⊕</mml:mo>
              <mml:mi>v</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>M</mml:mi>
                    <mml:mi>g</mml:mi>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>⊕</mml:mo>
              <mml:mi>v</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>F</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>⊕</mml:mo>
              <mml:mi>v</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mi>P</mml:mi>
                <mml:mo>)</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <inline-formula><mml:math display="inline"><mml:mo> ⊕ </mml:mo></mml:math></inline-formula> represents the concatenation operation. A video is thus represented by a sequence of these vectors <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> V </mml:mi><mml:mrow><mml:mtext> sign </mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula></p>
        <disp-formula id="FD2">
          <label>(2)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:msub>
                <mml:mi>V</mml:mi>
                <mml:mrow>
                  <mml:mtext>sign</mml:mtext>
                </mml:mrow>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:mi>v</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>I</mml:mi>
                        <mml:mn>1</mml:mn>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mo>,</mml:mo>
                  <mml:mi>v</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>I</mml:mi>
                        <mml:mn>2</mml:mn>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mo>,</mml:mo>
                  <mml:mo>⋯</mml:mo>
                  <mml:mo>,</mml:mo>
                  <mml:mi>v</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>I</mml:mi>
                        <mml:mi>N</mml:mi>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <italic>N</italic> denotes the total number of frames in the video.</p>
        <fig id="fig2">
          <label>Figure 2</label>
          <graphic xlink:href="https://html.scirp.org/file/1733639-rId32.jpeg?20260916024441" />
        </fig>
        <p><bold>Figure 2</bold><bold>.</bold> Landmarks of a hand [<xref ref-type="bibr" rid="B9">9</xref>].</p>
        <fig id="fig3">
          <label>Figure 3</label>
          <graphic xlink:href="https://html.scirp.org/file/1733639-rId33.jpeg?20260916024441" />
        </fig>
        <p><bold>Figure 3</bold><bold>.</bold> Landmarks of a face [<xref ref-type="bibr" rid="B10">10</xref>].</p>
        <fig id="fig4">
          <label>Figure 4</label>
          <graphic xlink:href="https://html.scirp.org/file/1733639-rId34.jpeg?20260916024441" />
        </fig>
        <p><bold>Figure 4</bold><bold>.</bold> Landmarks of a pose [<xref ref-type="bibr" rid="B9">9</xref>].</p>
      </sec>
      <sec id="sec3dot2">
        <title>3.2. LSTM Model Architecture</title>
        <p>For the task of classifying gesture sequences, we opted for a recurrent neural network, and more specifically an LSTM (Long Short-Term Memory) architecture [<xref ref-type="bibr" rid="B11">11</xref>]. LSTMs are particularly effective at learning long-term dependencies in sequential data, which is essential for understanding the dynamic movements of signs.</p>
        <p>The architecture of our model is as follows: (see <xref ref-type="fig" rid="fig5">Figure 5</xref>)</p>
        <p>1) LSTM Layers: Two overlapping LSTM layers (64 and 128 neurons) to process the sequence of landmark vectors and learn temporal patterns.</p>
        <p>2) Dense Layers (MLP): Two fully connected layers (64 neurons each) to interpret the features learned by the LSTMs.</p>
        <p>3) Dropout Layer: A regularization layer to prevent overfitting.</p>
        <p>4) Output Layer: A dense layer with a softmax activation function that produces a probability distribution over the nine (9) classes (the classes of the nine signs).</p>
        <fig id="fig5">
          <label>Figure 5</label>
          <graphic xlink:href="https://html.scirp.org/file/1733639-rId35.jpeg?20260916024441" />
        </fig>
        <p><bold>Figure 5</bold><bold>.</bold> LSTM neural network architecture.</p>
      </sec>
    </sec>
    <sec id="sec4">
      <title>4. Results and Discussion</title>
      <p>For the evaluation, we didn’t use part of these 270 sequences. We did a new, independent real-time recording: each sign was performed 10 times, making 90 test trials, with no overlap with the training data. So, the reported performance is calculated exclusively on these 90 independent trials. We trained and tested the model in four different configurations to assess the relative importance of each modality (hands, face, pose). The model was tested in real time via a webcam, performing 10 attempts for each of the 9 signs. </p>
      <p>The average recognition rates obtained are as follows: (see <bold>Table 1</bold>, <xref ref-type="fig" rid="fig6">Figure 6</xref> and <xref ref-type="fig" rid="fig7">Figure 7</xref>)</p>
      <p>Hands only: 62.22%Hands + Face: 27.77%Hands + Pose: 35.55%Hands + Face + Pose: 37.77% </p>
      <p><bold>Table 1</bold><bold>.</bold> Overall recognition rate based on features used.</p>
      <table-wrap id="tbl1">
        <label>Table 1</label>
        <table>
          <tbody>
            <tr>
              <td>
                <bold>Criteria</bold>
              </td>
              <td>
                <bold>Hands</bold>
                <bold>only</bold>
              </td>
              <td>
                <bold>Hand</bold>
                <bold>+</bold>
                <bold>Pose</bold>
              </td>
              <td>
                <bold>Hand</bold>
                <bold>+</bold>
                <bold>Face</bold>
              </td>
              <td>
                <bold>Hand</bold>
                <bold>+</bold>
                <bold>Face</bold>
                <bold>+</bold>
                <bold>Pose</bold>
              </td>
            </tr>
            <tr>
              <td>
                <bold>Number</bold>
                <bold>of</bold>
                <bold>signs</bold>
                <bold>tested</bold>
              </td>
              <td>9</td>
              <td>9</td>
              <td>9</td>
              <td>9</td>
            </tr>
            <tr>
              <td>
                <bold>Repetitions</bold>
                <bold>by</bold>
                <bold>sign</bold>
              </td>
              <td>10</td>
              <td>10</td>
              <td>10</td>
              <td>10</td>
            </tr>
            <tr>
              <td>
                <bold>Correct</bold>
                <bold>predictions</bold>
              </td>
              <td>56</td>
              <td>32</td>
              <td>25</td>
              <td>34</td>
            </tr>
            <tr>
              <td>
                <bold>Incorrect</bold>
                <bold>predictions</bold>
              </td>
              <td>34</td>
              <td>58</td>
              <td>65</td>
              <td>56</td>
            </tr>
            <tr>
              <td>
                <bold>Total</bold>
                <bold>number</bold>
                <bold>of</bold>
                <bold>predictions</bold>
              </td>
              <td>90</td>
              <td>90</td>
              <td>90</td>
              <td>90</td>
            </tr>
            <tr>
              <td>
                <bold>Accuracy</bold>
                <bold>rate</bold>
              </td>
              <td>
                <bold>62</bold>
                <bold>.</bold>
                <bold>22%</bold>
              </td>
              <td>
                <bold>35</bold>
                <bold>.</bold>
                <bold>55%</bold>
              </td>
              <td>
                <bold>27</bold>
                <bold>.</bold>
                <bold>77%</bold>
              </td>
              <td>
                <bold>37</bold>
                <bold>.</bold>
                <bold>77%</bold>
              </td>
            </tr>
          </tbody>
        </table>
      </table-wrap>
      <fig id="fig6">
        <label>Figure 6</label>
        <graphic xlink:href="https://html.scirp.org/file/1733639-rId36.jpeg?20260916024442" />
      </fig>
      <p><bold>Figure 6</bold><bold>.</bold> Comparison of the overall accuracy rates of the four configurations.</p>
      <fig id="fig7">
        <label>Figure 7</label>
        <graphic xlink:href="https://html.scirp.org/file/1733639-rId37.jpeg?20260916024442" />
      </fig>
      <p><bold>Figure 7</bold><bold>.</bold> Accuracy by sign according to feature configuration.</p>
      <p>These results suggest that using only hand features achieves the best recognition rate. This score reflects the central role of manual gestures, which convey the majority of meaning in sign language. The lower complexity of this configuration (fewer key points to process) allows the model to efficiently focus on the most relevant information.</p>
      <p>On the other hand, adding facial features significantly reduces accuracy. This can be explained by the high variability of facial expressions and the large number of extracted key points (468), which may introduce noise and distract the model from more discriminative hand features.</p>
      <p>Including pose features slightly improves accuracy compared to using facial features alone, likely because body posture provides useful additional spatial context. However, its contribution remains secondary compared to hand gestures.</p>
      <p>Finally, combining all features only marginally improves performance over face and pose alone, while remaining well below the hand-only configuration. The increased data dimensionality and the absence of a mechanism to weight the relative importance of each modality appear to dilute key signals, reducing the overall effectiveness of the model.</p>
      <p>Our best result (62.22%) is encouraging and comparable to those reported in the literature for more extensively studied sign languages. For instance, studies conducted on the WLASL dataset for American Sign Language (ASL) reported accuracies of 62.63% [<xref ref-type="bibr" rid="B12">12</xref>], although it should be noted that experimental conditions differ significantly in terms of vocabulary size and dataset scale.</p>
      <p>It should also be noted that, due to the binary nature of the collected data (correct/incorrect prediction), only accuracy and recall could be computed in this study. Metrics such as precision, F1-score, and confusion matrix require per-prediction class labels, which constitute a limitation of the current evaluation and a direction for future work. </p>
    </sec>
    <sec id="sec5">
      <title>5. Conclusions</title>
      <p>This work has enabled the development of a functional prototype of a real-time translation system from Niger Sign Language to text, using Deep Learning techniques. The main contribution lies in the creation of a first video corpus for NiSL and in the demonstration that an LSTM model, trained on hand movements alone, can achieve a promising performance of 62.22%.</p>
      <p>Our results emphasize that hands convey the essence of the message and that the addition of other modalities can, paradoxically, degrade performance due to the complexity and noise added.</p>
      <p>To improve this system, future perspectives include the expansion of the corpus. The integration of attention mechanisms, such as those proposed in Transformers models [<xref ref-type="bibr" rid="B13">13</xref>], could also help to better prioritize information from different modalities. Ultimately, such a system could greatly promote the social and professional inclusion of deaf people in Niger.</p>
    </sec>
    <sec id="sec6">
      <title>Author Contributions</title>
      <p>Conceptualization, Moussa Idi Bachir and Bouwey Alley Ibrahim; methodology, Moussa Idi Bachir; software, Bouwey Alley Ibrahim; validation, Naroua Harouna, Kadri Chaibou, and Moussa Idi Bachir; resources, Bouwey Alley Ibrahim; data curation, Bouwey Alley Ibrahim; writing—original draft preparation, Moussa Idi Bachir and Kadri Chaibou; writing—review and editing, Kadri Chaibou and Morou Ganda Yahaya; supervision, Naroua Harouna.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Grishchenko, I. and Bazarevsky, V. (2020) MediaPipe Holistic—Simultaneous Face, Hand and Pose Prediction, on Device. Google Research Blog.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Grishchenko, I.</string-name>
              <string-name>Bazarevsky, V.</string-name>
              <string-name>Face, H</string-name>
            </person-group>
            <year>2020</year>
            <article-title>MediaPipe Holistic—Simultaneous Face, Hand and Pose Prediction, on Device</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">Littlejohn, S.W. and Foss, K.A. (2010) Theories of Human Communication. 10th Edition, Waveland Press.</mixed-citation>
          <element-citation publication-type="book">
            <person-group person-group-type="author">
              <string-name>Littlejohn, S.W.</string-name>
              <string-name>Foss, K.A.</string-name>
              <string-name>Edition, W</string-name>
            </person-group>
            <year>2010</year>
            <article-title>Theories of Human Communication</article-title>
            <source>10th Edition</source>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Studio Kalangou (2023) La langue des sourds: Un apport à l’éducation des personnes déficientes auditives et malentendantes au Niger.</mixed-citation>
          <element-citation publication-type="other">
            <year>2023</year>
            <article-title>La langue des sourds: Un apport à l’éducation des personnes déficientes auditives et malentendantes au Niger</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Stokoe, W.C. (1960) Sign Language Structure: An Outline of the Visual Communication Systems of the American Deaf. Studies in Linguistics: Occasional Papers 8, University of Buffalo.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Stokoe, W.C.</string-name>
            </person-group>
            <year>1960</year>
            <article-title>Sign Language Structure: An Outline of the Visual Communication Systems of the American Deaf</article-title>
            <source>Studies in Linguistics: Occasional Papers 8</source>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Ahmed, M.A., Zaidan, B.B., Zaidan, A.A., Salih, M.M. and Lakulu, M.M.b. (2018) A Review on Systems-Based Sensory Gloves for Sign Language Recognition State of the Art between 2007 and 2017. <italic>Sensors</italic>, 18, Article 2208. https://doi.org/10.3390/s18072208 <pub-id pub-id-type="doi">10.3390/s18072208</pub-id><pub-id pub-id-type="pmid">29987266</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/s18072208">https://doi.org/10.3390/s18072208</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Ahmed, M.A.</string-name>
              <string-name>Zaidan, B.B.</string-name>
              <string-name>Zaidan, A.A.</string-name>
              <string-name>Salih, M.M.</string-name>
              <string-name>Lakulu, M.M.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>A Review on Systems-Based Sensory Gloves for Sign Language Recognition State of the Art between 2007 and 2017</article-title>
            <source>Sensors</source>
            <volume>18</volume>
            <elocation-id>2208</elocation-id>
            <pub-id pub-id-type="doi">10.3390/s18072208</pub-id>
            <pub-id pub-id-type="pmid">29987266</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B6">
        <label>6.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Mitra, S. and Acharya, T. (2007) Gesture Recognition: A Survey. <italic>IEEE</italic><italic>Transactions</italic><italic>on</italic><italic>Systems</italic>, <italic>Man</italic><italic>and</italic><italic>Cybernetics</italic>, <italic>Part</italic><italic>C</italic> ( <italic>Applications</italic><italic>and</italic><italic>Reviews</italic>), 37, 311-324. https://doi.org/10.1109/tsmcc.2007.893280 <pub-id pub-id-type="doi">10.1109/tsmcc.2007.893280</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/tsmcc.2007.893280">https://doi.org/10.1109/tsmcc.2007.893280</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Mitra, S.</string-name>
              <string-name>Acharya, T.</string-name>
              <string-name>Systems, M</string-name>
              <string-name>Cybernetics, P</string-name>
            </person-group>
            <year>2007</year>
            <article-title>Gesture Recognition: A Survey</article-title>
            <source>IEEE Transactions on Systems</source>
            <volume>37</volume>
            <pub-id pub-id-type="doi">10.1109/tsmcc.2007.893280</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B7">
        <label>7.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">Crasborn, O.A. (2006) Nonmanual Structures in Sign Language. In: Brown, K., Ed., <italic>Encyclopedia</italic><italic>of</italic><italic>Language</italic><italic>&amp;</italic><italic>Linguistics</italic>, Elsevier, 668-672. https://doi.org/10.1016/b0-08-044854-2/04216-4 <pub-id pub-id-type="doi">10.1016/b0-08-044854-2/04216-4</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/b0-08-044854-2/04216-4">https://doi.org/10.1016/b0-08-044854-2/04216-4</ext-link></mixed-citation>
          <element-citation publication-type="book">
            <person-group person-group-type="author">
              <string-name>Crasborn, O.A.</string-name>
              <string-name>Brown, K.</string-name>
              <string-name>Linguistics, E</string-name>
            </person-group>
            <year>2006</year>
            <article-title>Nonmanual Structures in Sign Language</article-title>
            <source>In: Brown</source>
            <volume>668</volume>
            <pub-id pub-id-type="doi">10.1016/b0-08-044854-2/04216-4</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B8">
        <label>8.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Shorten, C. and Khoshgoftaar, T.M. (2019) A Survey on Image Data Augmentation for Deep Learning. <italic>Journal</italic><italic>of</italic><italic>Big</italic><italic>Data</italic>, 6, Article No. 60. https://doi.org/10.1186/s40537-019-0197-0 <pub-id pub-id-type="doi">10.1186/s40537-019-0197-0</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1186/s40537-019-0197-0">https://doi.org/10.1186/s40537-019-0197-0</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Shorten, C.</string-name>
              <string-name>Khoshgoftaar, T.M.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>A Survey on Image Data Augmentation for Deep Learning</article-title>
            <source>Journal of Big Data</source>
            <volume>6</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1186/s40537-019-0197-0</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B9">
        <label>9.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">GeeksforGeeks (2023) Python-Facial and Hand Recognition Using MediaPipe Holistic. GeeksforGeeks.</mixed-citation>
          <element-citation publication-type="other">
            <year>2023</year>
            <article-title>Python-Facial and Hand Recognition Using MediaPipe Holistic</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B10">
        <label>10.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Albadawi, Y., AlRedhaei, A. and Takruri, M. (2023) Real-Time Machine Learning-Based Driver Drowsiness Detection Using Visual Features. <italic>Journal</italic><italic>of</italic><italic>Imaging</italic>, 9, Article 91. https://doi.org/10.3390/jimaging9050091 <pub-id pub-id-type="doi">10.3390/jimaging9050091</pub-id><pub-id pub-id-type="pmid">37233309</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/jimaging9050091">https://doi.org/10.3390/jimaging9050091</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Albadawi, Y.</string-name>
              <string-name>AlRedhaei, A.</string-name>
              <string-name>Takruri, M.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Real-Time Machine Learning-Based Driver Drowsiness Detection Using Visual Features</article-title>
            <source>Journal of Imaging</source>
            <volume>9</volume>
            <elocation-id>91</elocation-id>
            <pub-id pub-id-type="doi">10.3390/jimaging9050091</pub-id>
            <pub-id pub-id-type="pmid">37233309</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B11">
        <label>11.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Hochreiter, S. and Schmidhuber, J. (1997) Long Short-Term Memory. <italic>Neural</italic><italic>Computation</italic>, 9, 1735-1780. https://doi.org/10.1162/neco.1997.9.8.1735 <pub-id pub-id-type="doi">10.1162/neco.1997.9.8.1735</pub-id><pub-id pub-id-type="pmid">9377276</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1162/neco.1997.9.8.1735">https://doi.org/10.1162/neco.1997.9.8.1735</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Hochreiter, S.</string-name>
              <string-name>Schmidhuber, J.</string-name>
            </person-group>
            <year>1997</year>
            <article-title>Long Short-Term Memory</article-title>
            <source>Neural Computation</source>
            <volume>9</volume>
            <pub-id pub-id-type="doi">10.1162/neco.1997.9.8.1735</pub-id>
            <pub-id pub-id-type="pmid">9377276</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B12">
        <label>12.</label>
        <citation-alternatives>
          <mixed-citation publication-type="web">Li, D.X., <italic>et al</italic>. (2020) WLASL: A Large-Scale Word-Level American Sign Language Dataset for Deep Learning-Based Recognition. https://dxli94.github.io/WLASL/</mixed-citation>
          <element-citation publication-type="web">
            <person-group person-group-type="author">
              <string-name>Li, D.X.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>WLASL: A Large-Scale Word-Level American Sign Language Dataset for Deep Learning-Based Recognition</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B13">
        <label>13.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L. and Polosukhin, I. (2017) Attention Is All You Need. 31 <italic>st Conference on Neural Information Processing Systems</italic> ( <italic>NIPS</italic>2017), Long Beach, CA, 1-11. https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf</mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Vaswani, A.</string-name>
              <string-name>Shazeer, N.</string-name>
              <string-name>Parmar, N.</string-name>
              <string-name>Uszkoreit, J.</string-name>
              <string-name>Jones, L.</string-name>
              <string-name>Gomez, A.N.</string-name>
              <string-name>Kaiser, L.</string-name>
              <string-name>Polosukhin, I.</string-name>
              <string-name>Beach, C</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Attention Is All You Need</article-title>
            <source>31st Conference on Neural Information Processing Systems (NIPS 2017)</source>
            <volume>1</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>