<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article">
 <front>
  <journal-meta>
   <journal-id journal-id-type="publisher-id">
    ojapps
   </journal-id>
   <journal-title-group>
    <journal-title>
     Open Journal of Applied Sciences
    </journal-title>
   </journal-title-group>
   <issn pub-type="epub">
    2165-3917
   </issn>
   <issn publication-format="print">
    2165-3925
   </issn>
   <publisher>
    <publisher-name>
     Scientific Research Publishing
    </publisher-name>
   </publisher>
  </journal-meta>
  <article-meta>
   <article-id pub-id-type="doi">
    10.4236/ojapps.2025.158160
   </article-id>
   <article-id pub-id-type="publisher-id">
    ojapps-144967
   </article-id>
   <article-categories>
    <subj-group subj-group-type="heading">
     <subject>
      Articles
     </subject>
    </subj-group>
    <subj-group subj-group-type="Discipline-v2">
     <subject>
      Biomedical 
     </subject>
     <subject>
       Life Sciences, Chemistry 
     </subject>
     <subject>
       Materials Science, Computer Science 
     </subject>
     <subject>
       Communications, Engineering, Physics 
     </subject>
     <subject>
       Mathematics
     </subject>
    </subj-group>
   </article-categories>
   <title-group>
    Addressing Hearing Impairments through Machine Learning: A Review of Sound Detection and Assistive Technologies
   </title-group>
   <contrib-group>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Nada
      </surname>
      <given-names>
       Barnawi
      </given-names>
     </name>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Mohammed
      </surname>
      <given-names>
       Alnuem
      </given-names>
     </name>
    </contrib>
   </contrib-group> 
   <aff id="affnull">
    <addr-line>
     aInformation Systems Department, King Saud University, Riyadh, KSA
    </addr-line> 
   </aff> 
   <pub-date pub-type="epub">
    <day>
     01
    </day> 
    <month>
     08
    </month>
    <year>
     2025
    </year>
   </pub-date> 
   <volume>
    15
   </volume> 
   <issue>
    08
   </issue>
   <fpage>
    2383
   </fpage>
   <lpage>
    2407
   </lpage>
   <history>
    <date date-type="received">
     <day>
      16,
     </day>
     <month>
      July
     </month>
     <year>
      2025
     </year>
    </date>
    <date date-type="published">
     <day>
      18,
     </day>
     <month>
      July
     </month>
     <year>
      2025
     </year> 
    </date> 
    <date date-type="accepted">
     <day>
      18,
     </day>
     <month>
      August
     </month>
     <year>
      2025
     </year> 
    </date>
   </history>
   <permissions>
    <copyright-statement>
     © Copyright 2014 by authors and Scientific Research Publishing Inc. 
    </copyright-statement>
    <copyright-year>
     2014
    </copyright-year>
    <license>
     <license-p>
      This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/
     </license-p>
    </license>
   </permissions>
   <abstract>
    This paper studies recent assistive technologies and AI sound detection systems that have been developed to support both the safety and communication of individuals who are deaf. It highlights how modern sound detection systems effectively address challenges such as real-time processing and polyphonic audio environments, while integrating speech recognition to improve situational awareness and interaction. The findings confirm that these technologies not only increase auditory accessibility but also empower greater independence and security for deaf and hard-of-hearing users in everyday environments.
   </abstract>
   <kwd-group> 
    <kwd>
     Machine Learning-Based
    </kwd> 
    <kwd>
      Deaf and Hard of Hearing
    </kwd> 
    <kwd>
      Sound Detection
    </kwd> 
    <kwd>
      Assistive Technologies
    </kwd> 
    <kwd>
      Speech Recognition
    </kwd> 
    <kwd>
      Realtime
    </kwd>
   </kwd-group>
  </article-meta>
 </front>
 <body>
  <sec id="s1">
   <title>1. Introduction</title>
   <p>In daily life, humans face several dangers that often require constant vigilance and caution to deal with surrounding risks in general. This necessitates attention, identification of potential hazards, determination of the type of risk, and subsequent action accordingly. Risks vary in terms of response speed and the level of danger they pose. Some, like responding to fire alarms in a building, demand swift movement and evacuation, potentially even saving lives.</p>
   <p>One of the most significant obstacles that may impede a swift response is the loss of hearing. Deaf or hard-of-hearing individuals face additional risks, such as an inability to respond to potential warnings and alerts. They heavily rely on others in times of danger, leading to a loss in response speed and the ability to promptly handle similar risks. Furthermore, beyond difficulty in hearing and identifying risks, they struggle with oral communication in their daily lives, especially in times of potential danger.</p>
   <p>Recent data from the World Health Organization (WHO) highlights a growing global trend in hearing impairment. Currently, over 430 million people, more than 5% of the global population, suffer from hearing loss, whether congenital or acquired. By 2050, this number is expected to exceed 700 million, meaning one in ten people will experience disabling hearing loss <xref ref-type="bibr" rid="scirp.144967-1">
     [1]
    </xref>. This group is identified as the second-largest among individuals with disabilities in the General Census of Population and Housing <xref ref-type="bibr" rid="scirp.144967-2">
     [2]
    </xref>.</p>
   <p>In response to these growing needs, rapid advancements in computer science—particularly in artificial intelligence (AI)—have introduced powerful tools to support the hearing-impaired community. Technologies such as sound recognition and speech-to-text systems have significantly improved their ability to detect surrounding hazards and communicate in daily life. While early assistive solutions focused on limited alerts (like vibrations or flashing lights), recent AI-based systems use machine learning to analyze and classify environmental sounds in real time. These modern systems represent a leap forward, offering both greater functionality and adaptability.</p>
   <p>So, in this paper, it will serve as the literature review, elucidating the key concepts central to understanding how machine learning can contribute and utilize to enhance comfort and safety environment for deaf and hard of hearing people. It commences with a concise overview of various assistive technologies. Also, provides an exploration of sound detection technologies and comprehensively covers sound detection within machine learning, encompassing preprocessing, feature extraction, model training and classification, real-time aspects, and associated challenges and constraints. In addition, a grasp of speech recognition and automatic speech recognition is established. To conclude, this comprehensive review presents an examination of research papers and applications offering insightful reviews.</p>
  </sec><sec id="s2">
   <title>2. Assistive Technologies to Hearing Disabilities</title>
   <p>Assistive technologies (AT) have been among the most significant advancements in the past 20 years for helping individuals with disabilities, including those with hearing impairments. According to ISO 9999:2016 and UNE-ISO 9999:2017, AT encompasses any product—whether a device, piece of equipment, instrument, or software—specifically designed to improve the engagement and functionality of individuals with disabilities. These technologies assist, support, train, measure, or substitute for bodily functions, aiming to prevent impairments, activity limitations, or participation restrictions <xref ref-type="bibr" rid="scirp.144967-3">
     [3]
    </xref>.</p>
   <p>For those with hearing impairments, AT includes hearing aids, communication systems, low-tech devices, cochlear implants, and specialized software and hardware that enhance hearing and communication abilities. These tools not only promote independence and well-being but also prevent secondary health issues and provide socioeconomic benefits by reducing healthcare costs and stimulating economic growth <xref ref-type="bibr" rid="scirp.144967-3">
     [3]
    </xref>.</p>
   <p>In communication for the deaf and hard of hearing, significant technological progress has been made. Text telephones (TTY) and telecommunications devices have been crucial in enabling telephone communication for this community, with communications assistants helping to bridge the gap between hearing-impaired and standard telephone users by relaying messages . Recent innovations include sign recognition systems for interpreting sign language gestures and Personalized Emergency Response Systems that enhance safety through sensor-based alarms <xref ref-type="bibr" rid="scirp.144967-4">
     [4]
    </xref> <xref ref-type="bibr" rid="scirp.144967-5">
     [5]
    </xref>. Mobile applications have also become a breakthrough, empowering the hearing-impaired to live more autonomously and communicate effectively within their communities . These technologies cater to the diverse needs of individuals with hearing impairments across various areas of life.</p>
  </sec><sec id="s3">
   <title>3. Sound Detection Technologies</title>
   <p>Sound detection technology has seen growing interest due to recent advancements that have made it more reliable and precise. These developments enable machines to mimic human hearing and understand context, allowing for applications such as smartphone alerts <xref ref-type="bibr" rid="scirp.144967-6">
     [6]
    </xref>, diagnosing conditions like coughing <xref ref-type="bibr" rid="scirp.144967-7">
     [7]
    </xref>, improving security by identifying threats, aiding in cataloging audio archives <xref ref-type="bibr" rid="scirp.144967-8">
     [8]
    </xref>, and enhancing safety monitoring on construction sites <xref ref-type="bibr" rid="scirp.144967-9">
     [9]
    </xref>. These technologies provide critical situational awareness, especially in environments beyond visual or attentive reach.</p>
   <p>Sound detection involves identifying sound events in a continuous audio signal and analyzing them using various methodologies to extract relevant information, depending on the intended application. This process typically involves multiple techniques related to audio signal processing or machine learning. For example, traditional computational analysis systems extract specific acoustic features from an input signal, which can then be categorized and detected using supervised classifiers like neural networks. Developing sound detection applications requires defining several factors, such as the nature of the application, technological constraints, desired complexity and precision, and data availability <xref ref-type="bibr" rid="scirp.144967-10">
     [10]
    </xref>.</p>
   <p>Sound event detection systems are usually customized for specific tasks and environments, requiring a combination of different techniques and processes. The implementation of sound detection technologies involves integrating various methods tailored to specific purposes. The following section will explore the core processes and key techniques used in sound detection technologies, particularly those involving machine learning, in different environments.</p>
  </sec><sec id="s4">
   <title>4. Machine Learning-Based Sound Detection</title>
   <p>Machine learning has significantly advanced numerous domains, as highlighted by Rebala et al. They have succinctly defined the core processes of sound detection within the broader machine learning context. They characterize machine learning as a computer science discipline dedicated to automating solutions for intricate problems that defy conventional programming methods. Traditional programming necessitates meticulous design and code implementation, posing challenges for tasks such as character recognition or sound event detection. In contrast, machine learning algorithms acquire knowledge from labeled data, bypassing the need for explicit rules. These algorithms excel at solving complex problems with greater accuracy and objectivity compared to rules crafted by humans. Their approach involves creating a model from a dataset and subsequently predicting labels for new data points <xref ref-type="bibr" rid="scirp.144967-11">
     [11]
    </xref>.</p>
   <p>Therefore, in order to improve the detection of sound events across various applications, it is contemporary to employ machine learning methods and techniques, whether in real-time or not. It is also essential to grasp the key challenges and constraints associated with machine learning-based sound detection. So, the following sections will discuss more about these topics.</p>
   <sec id="s4_1">
    <title>4.1. Processes and Technics for Sound Detection</title>
    <p>To gain a deeper understanding of the fundamental processes and methodologies utilized in machine learning for sound detection, the two referenced papers <xref ref-type="bibr" rid="scirp.144967-12">
      [12]
     </xref> <xref ref-type="bibr" rid="scirp.144967-13">
      [13]
     </xref> provide insights into the core stages of sound detection technologies or sound event detection systems. These stages comprise three key components: 1) Preprocessing, 2) Feature Extraction, and 3) Model Training and Classification. Various papers will demonstrate a range of machine-learning techniques for sound detection, with supervised learning emerging as the dominant approach in tackling the sound event detection task.</p>
    <p>The initial stage in sound event detection entails the application of various techniques to enhance the quality of the audio data before feature extraction. This step is necessary because raw audio data cannot be directly employed as input for machine learning-based classification. The rationale behind this necessity lies in the presence of signal redundancy, which must be addressed. Preprocessing activities typically encompass tasks such as noise reduction, equalization, low-pass filtering, and segmenting the original audio signal into audio and silent events to facilitate subsequent feature extraction <xref ref-type="bibr" rid="scirp.144967-12">
      [12]
     </xref> <xref ref-type="bibr" rid="scirp.144967-13">
      [13]
     </xref>.</p>
    <p>Data Collection: Effective sound detection systems hinge on understanding sound data and context, especially in supervised learning, where the training data must closely resemble the intended application scenario. As sound event detection can cover a broad range of sound classes and environments, no single dataset or acoustic model fits all scenarios. Instead, multiple datasets are curated to address specific challenges, with dataset size often reflecting the complexity of the labels. Access to diverse datasets is crucial for training models that can adapt to different environments and sound events <xref ref-type="bibr" rid="scirp.144967-13">
      [13]
     </xref>.</p>
    <p>Three research papers explore gunshot-related sound detection systems using various datasets, achieving different accuracy levels. The first study uses a sound event recognition model trained with 692 samples from sources like Freesound.org and YouTube, achieving an accuracy of 77.32% in indoor event classification <xref ref-type="bibr" rid="scirp.144967-14">
      [14]
     </xref>. The second paper introduces a hybrid algorithm for detecting gunshots in indoor settings, with data collected from online sources and real-world locations like shopping malls and universities, achieving accuracy between 91.65% and 94.97%, depending on the classifier <xref ref-type="bibr" rid="scirp.144967-15">
      [15]
     </xref>. The third paper presents a novel approach for recognizing environmental sounds across various settings, using a dataset of 1000 sounds across 10 categories, including gunshots, and achieving up to 92.22% classification accuracy <xref ref-type="bibr" rid="scirp.144967-16">
      [16]
     </xref>.</p>
    <p>Audio Signal Preprocessing is pivotal for the effectiveness of machine learning algorithms, particularly in creating generalized predictive models for classification tasks <xref ref-type="bibr" rid="scirp.144967-17">
      [17]
     </xref>. This phase involves a variety of techniques to prepare the audio signal for feature extraction, with the choice of methods depending on the sensitivity and accuracy requirements of the specific application or system.</p>
    <p>Normalization is an essential technique in data analysis, especially when handling data from different sources with varying measurement scales. Dalwinder and Birmohan define data normalization as a process that resizes or transforms raw data so that each feature contributes uniformly. For example, when data includes both age (in years) and height (in centimeters), normalization ensures consistency in the magnitude of these parameters <xref ref-type="bibr" rid="scirp.144967-17">
      [17]
     </xref> <xref ref-type="bibr" rid="scirp.144967-18">
      [18]
     </xref>.</p>
    <p>In signal processing, Noise Reduction and Filtering are two fundamental methods. Noise refers to unwanted disturbances within a specific frequency range, such as electric waves and random variations <xref ref-type="bibr" rid="scirp.144967-19">
      [19]
     </xref>. Noise reduction is crucial for cleaning data that is susceptible to noise interference <xref ref-type="bibr" rid="scirp.144967-18">
      [18]
     </xref>. Filtering is closely related and focuses on eliminating noise, often using low-pass or high-pass filters to remove high-frequency noise or highlight specific data features <xref ref-type="bibr" rid="scirp.144967-20">
      [20]
     </xref>. These methods are particularly useful in medical fields like Electrocardiogram (ECG) signal analysis. ECG signals, which record heart rates, are vital for investigating abnormal heart functions, such as arrhythmias and conduction disturbances. Filters like low-pass, high-pass, and Butterworth filters are employed to preprocess these signals by effectively removing high-frequency noise, with Butterworth filters being particularly effective <xref ref-type="bibr" rid="scirp.144967-21">
      [21]
     </xref>.</p>
    <p>Another important preprocessing step is Silence Removal, which addresses the presence of complete silence at the beginning, end, or within audio signals. This process applies a specific threshold to remove unvoiced portions that lack relevant data, retaining only the voiced sections <xref ref-type="bibr" rid="scirp.144967-22">
      [22]
     </xref>. In studies related to cough sound detection, silence removal is crucial as it allows the focus to remain on sound events, thereby improving detection accuracy <xref ref-type="bibr" rid="scirp.144967-23">
      [23]
     </xref>-<xref ref-type="bibr" rid="scirp.144967-25">
      [25]
     </xref>.</p>
    <p>In practice, sounds often overlap or occur sequentially, requiring effective segmentation for classification. Segmentation prepares audio signals for feature extraction by dividing them into distinct segments based on temporal proximity and setting thresholds to determine the relevance of sound segments <xref ref-type="bibr" rid="scirp.144967-26">
      [26]
     </xref>.</p>
    <p>The final preprocessing method is Feature Windowing, which treats non-stationary signals as quasi-stationary by sliding a window over the entire signal for comprehensive analysis . Unlike traditional acoustic analysis systems that divide sound recordings into fixed-sized windows, contemporary methods adapt the window size to the signal’s characteristics <xref ref-type="bibr" rid="scirp.144967-27">
      [27]
     </xref> <xref ref-type="bibr" rid="scirp.144967-28">
      [28]
     </xref>.</p>
    <p>Segmentation and feature windowing are closely related but serve different purposes. While segmentation focuses on creating meaningful signal segments, feature windowing concentrates on extracting features within these windows. For instance, Baughman et al. used a peak detection method to identify specific sounds in a tennis match recording, isolating the sound of interest within a single window. This technique is valuable for classifying acoustic events using machine learning. However, if the sound of interest spans two different windows, classification accuracy may be reduced <xref ref-type="bibr" rid="scirp.144967-28">
      [28]
     </xref>.</p>
    <p>Ultimately, these preprocessing methods lead to feature extraction, where the extracted features are stored in a database for training the classifier, a topic to be discussed in the next step.</p>
    <p>Feature extraction is vital in audio content analysis as it involves creating a numerical representation, or feature vector, that captures key acoustic characteristics of audio segments <xref ref-type="bibr" rid="scirp.144967-12">
      [12]
     </xref> <xref ref-type="bibr" rid="scirp.144967-13">
      [13]
     </xref>. This vector is foundational for various audio analysis and information extraction algorithms <xref ref-type="bibr" rid="scirp.144967-29">
      [29]
     </xref>, as it condenses extensive data into a more manageable format while retaining essential information <xref ref-type="bibr" rid="scirp.144967-30">
      [30]
     </xref>.</p>
    <p>Time-frequency analysis techniques are commonly used in feature extraction to focus on the signal’s frequency domain, breaking the signal into overlapping frames to track frequency distribution changes over time <xref ref-type="bibr" rid="scirp.144967-30">
      [30]
     </xref>. These techniques, such as Mel frequency cepstral coefficients (MFCCs), Log-Mel energies, and Spectrograms, divide the signal into smaller segments and calculate the frequency content for each. The resulting magnitude spectrum shows energy distribution over frequency for each segment, allowing for the computation of a concise set of features that capture fundamental spectral characteristics. This compact feature set is preferred for machine learning algorithms, as it maintains informativeness while reducing complexity, making it widely applicable in signal processing tasks <xref ref-type="bibr" rid="scirp.144967-29">
      [29]
     </xref>.</p>
    <p>Different sound extraction features are used depending on the system and its performance needs. Common methods include log-mel energies, MFCCs, spectrograms, and constant-Q filterbank-based features. For instance, Jain et al. developed ProtoSound, an interactive system that enhances sound awareness for deaf or hard-of-hearing individuals. ProtoSound personalizes sound recognition models using user recordings and log-mel spectrogram features, significantly improving performance when integrated with deep convolutional neural networks (CNN) <xref ref-type="bibr" rid="scirp.144967-31">
      [31]
     </xref>.</p>
    <p>Deep Neural Networks (DNNs) have been used to classify cough sounds by extracting MFCC features. Liu et al. <xref ref-type="bibr" rid="scirp.144967-32">
      [32]
     </xref> reported a DNN with MFCC features achieving 90.1% accuracy for positive cases and 85% for negative cases, while Amoh and Odame <xref ref-type="bibr" rid="scirp.144967-33">
      [33]
     </xref> attained 86.8% accuracy for positive cases and 92.7% for negative cases using a similar approach.</p>
    <p>In speech-based emotion recognition, extensive reviews have compared different approaches. Studies using the RAVDESS, Emo-DB, and IEMOCAP datasets found that Log-Mel spectrogram features outperform MFCCs, challenging their prevalent use in this field <xref ref-type="bibr" rid="scirp.144967-34">
      [34]
     </xref> <xref ref-type="bibr" rid="scirp.144967-35">
      [35]
     </xref>. Other feature types are also effective for sound event recognition. For example, Sing et al. analyzed constant-Q filter bank-based time-frequency representations, offering superior frequency resolution at low frequencies compared to MFSC. Wang et al. used a model based on discrete Fourier parameters, where frequency harmonics and their derivatives effectively distinguished emotion classes. Additionally, Badshah et al. introduced a method combining spectrograms and CNN for sound event recognition, achieving promising results in emotion prediction <xref ref-type="bibr" rid="scirp.144967-36">
      [36]
     </xref>-<xref ref-type="bibr" rid="scirp.144967-38">
      [38]
     </xref>.</p>
    <p>During this phase, the system learns to correlate extracted audio signal features with specific class labels to develop a model for categorizing audio recordings into predefined classes. For instance, a sound scene classification system might categorize recordings as “home,” “street,” or “office” <xref ref-type="bibr" rid="scirp.144967-12">
      [12]
     </xref> <xref ref-type="bibr" rid="scirp.144967-13">
      [13]
     </xref>.</p>
    <p>Classification, a machine learning method, assigns input patterns to predefined categories using a classifier. This process involves two main phases: training the classifier with samples representing each class and then categorizing unknown inputs into these classes. Different classification techniques use various algorithms and rules, which can impact accuracy based on the specific application <xref ref-type="bibr" rid="scirp.144967-39">
      [39]
     </xref>. Four notable classification algorithms for sound detection are:</p>
    <p>1) Convolutional Neural Networks (CNNs): CNNs, commonly used in deep learning, excel in image classification but require a specialized approach for sound. Sound is converted into spectrogram images, which CNNs can then analyze. CNNs consist of layers that process sound data by generating feature maps to identify patterns <xref ref-type="bibr" rid="scirp.144967-40">
      [40]
     </xref> <xref ref-type="bibr" rid="scirp.144967-41">
      [41]
     </xref>. They are effective but require longer training times <xref ref-type="bibr" rid="scirp.144967-8">
      [8]
     </xref>. For example, CNNs have been used to classify bird sounds in normal and threatened conditions by analyzing spectrograms .</p>
    <p>2) Deep Neural Networks (DNNs): Unlike standard neural networks with a single hidden layer, DNNs have multiple hidden layers, mimicking the human brain’s visual recognition model. This depth allows DNNs to achieve high precision by progressively recognizing complex information <xref ref-type="bibr" rid="scirp.144967-29">
      [29]
     </xref>. Research by Li et al. on DNN hyperparameters for speech recognition and audio analysis demonstrates their adaptability and high performance <xref ref-type="bibr" rid="scirp.144967-8">
      [8]
     </xref>.</p>
    <p>3) Recurrent Neural Networks (RNNs): RNNs are designed for sequence modeling, capturing temporal dependencies by considering both current input and previous hidden states. This allows RNNs to handle sequences of varying lengths and contexts. Arsenali et al. utilized RNNs for sound event classification, achieving high accuracy and sensitivity with their optimized model <xref ref-type="bibr" rid="scirp.144967-42">
      [42]
     </xref> <xref ref-type="bibr" rid="scirp.144967-43">
      [43]
     </xref>.</p>
    <p>4) Decision Trees: Decision trees, a non-linear classification method, use a hierarchical structure to eliminate classes sequentially until the correct class is reached. They are efficient for problems with many classes. The Ordinary Binary Decision Tree is a common variant, and research by Saifan et al. explored its application in sound engine classification <xref ref-type="bibr" rid="scirp.144967-39">
      [39]
     </xref>.</p>
    <p>Based on the different stages of sound detection technology, the entire process can be categorized and summarized into three primary stages: Preprocessing, Feature Extraction, and Model Training and Classification. These stages encompass the core methodologies and techniques employed in sound detection systems. <xref ref-type="table" rid="table1">
      Table 1
     </xref> provides a detailed overview of these processes, highlighting the key steps, methods, and relevant research associated with each stage. This structured approach allows for a clearer understanding of how sound detection technology operates and the advancements made at each stage.</p>
    <table-wrap id="table1">
     <label>
      <xref ref-type="table" rid="table1">
       Table 1
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.144967-"></xref>Table 1. Summary of key processes in sound detection technology.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td aleft" width="20.93%"><p style="text-align:left">Process Stage</p></td> 
       <td class="custom-bottom-td aleft" width="30.85%"><p style="text-align:left">Description</p></td> 
       <td class="custom-bottom-td aleft" width="17.36%"><p style="text-align:left">Key References</p></td> 
       <td class="custom-bottom-td aleft" width="30.86%"><p style="text-align:left">Key Techniques/Methods</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="20.93%"><p style="text-align:center">Preprocessing</p></td> 
       <td class="custom-top-td acenter" width="30.85%"><p style="text-align:center">Enhances audio data quality before feature extraction.</p></td> 
       <td class="custom-top-td acenter" width="17.36%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.144967-12">
          [12]
         </xref>-<xref ref-type="bibr" rid="scirp.144967-28">
          [28]
         </xref></p></td> 
       <td class="custom-top-td acenter" width="30.86%"><p style="text-align:center">Noise Reduction, Equalization, Filtering, Silence Removal, Segmentation, Feature Windowing</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="20.93%"><p style="text-align:center">Feature Extraction</p></td> 
       <td class="acenter" width="30.85%"><p style="text-align:center">Converts audio data into a numerical representation (feature vector) that captures key acoustic characteristics.</p></td> 
       <td class="acenter" width="17.36%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.144967-12">
          [12]
         </xref> <xref ref-type="bibr" rid="scirp.144967-13">
          [13]
         </xref> <xref ref-type="bibr" rid="scirp.144967-29">
          [29]
         </xref>-<xref ref-type="bibr" rid="scirp.144967-38">
          [38]
         </xref></p></td> 
       <td class="acenter" width="30.86%"><p style="text-align:center">MFCCs, Log-Mel Energies, Spectrograms, Time-Frequency Analysis</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="20.93%"><p style="text-align:center">Model Training &amp; Classification</p></td> 
       <td class="acenter" width="30.85%"><p style="text-align:center">Develops models to classify audio recordings into predefined classes based on extracted features.</p></td> 
       <td class="acenter" width="17.36%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.144967-12">
          [12]
         </xref> <xref ref-type="bibr" rid="scirp.144967-13">
          [13]
         </xref> <xref ref-type="bibr" rid="scirp.144967-39">
          [39]
         </xref>-<xref ref-type="bibr" rid="scirp.144967-43">
          [43]
         </xref></p></td> 
       <td class="acenter" width="30.86%"><p style="text-align:center">CNNs, DNNs, RNNs, Decision Trees</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>The table categorizes sound detection technology into three main stages, highlighting key processes and methods employed in each stage, alongside relevant research, to present a clear overview of advancements in sound detection technology</p>
   </sec>
   <sec id="s4_2">
    <title>4.2. Real-Time Sound Detection</title>
    <p>Real-time processing refers to the immediate handling of data as it becomes available, with two main requirements: the processing must be completed faster than the data’s duration, and delays should be minimized to ideally zero. In computer operating systems, “real-time” also involves precise scheduling to handle events within a set timeframe <xref ref-type="bibr" rid="scirp.144967-44">
      [44]
     </xref>.</p>
    <p>Signal classification can be divided into three categories:</p>
    <p>1) Analog Signal Processing: Deals with signals that have not been digitized, such as those from radios or older televisions.</p>
    <p>2) Continuous Signal Processing: Focuses on signals with continuous amplitude variations, including modeling continuous systems, system function adjustments, and time-based filtering.</p>
    <p>3) Discrete Signal Processing: Handles signals sampled and quantized at specific intervals, represented as a sequence of numbers. Discrete-time signals are crucial for real-time applications like sound event detection and speech recognition due to their ability to capture and process temporal dynamics <xref ref-type="bibr" rid="scirp.144967-45">
      [45]
     </xref>.</p>
    <p>Non-real-time signal processing involves manipulating pre-gathered and digitized signals without real-time constraints, while real-time processing demands precise timing from both hardware and software <xref ref-type="bibr" rid="scirp.144967-46">
      [46]
     </xref>. Key strategies for real-time sound detection include:</p>
    <p>1) On-line Processing: Manages live data streams, computing results as the data is recorded or transmitted, typically using short buffers.</p>
    <p>2) Incremental Processing: Optimizes on-line processing by reducing the delay between data input and analytical results.</p>
    <p>Digital Signal Processing (DSP) significantly enhances real-time sound detection by digitally representing and analyzing signals. DSP involves breaking down signals, applying mathematical operations like filtering and Fourier transforms, and integrating results for effective analysis. Recent advancements in DSP technology have enabled real-time applications where analog methods are impractical <xref ref-type="bibr" rid="scirp.144967-46">
      [46]
     </xref>.</p>
    <p>An example of DSP application is in real-time arrhythmia classification, involving three stages: preprocessing to reduce noise, feature extraction using techniques like Wavelet Transform, and classification with algorithms such as probabilistic neural networks. DSP is essential in minimizing noise, extracting features, and classifying arrhythmias <xref ref-type="bibr" rid="scirp.144967-47">
      [47]
     </xref>.</p>
   </sec>
   <sec id="s4_3">
    <title>4.3. Security Aspect</title>
    <p>Machine learning (ML) is becoming an essential part of modern cybersecurity, helping systems detect threats, analyze behavior, and respond automatically. However, as ML becomes more deeply integrated into security infrastructure, it also introduces new vulnerabilities. These risks largely stem from ML’s heavy reliance on large datasets and complex algorithms, which can compromise data confidentiality, integrity, and availability if not properly secured <xref ref-type="bibr" rid="scirp.144967-48">
      [48]
     </xref>.</p>
    <p>To mitigate these vulnerabilities—especially in the face of quantum computing threats—post-quantum cryptographic (PQC) schemes such as Kyber and NTRU have gained attention. These lattice-based algorithms offer quantum-resistant key encapsulation mechanisms, ensuring secure communication and data protection even under quantum-enabled attacks <xref ref-type="bibr" rid="scirp.144967-49">
      [49]
     </xref>.</p>
    <p>Kyber, in particular, can be embedded into ML pipelines to secure the transmission and storage of sensitive training data, model parameters, and updates. It protects against unauthorized access during remote deployment and communication, making it a strong candidate for ML-based systems operating in distributed environments <xref ref-type="bibr" rid="scirp.144967-50">
      [50]
     </xref>.</p>
    <p>Similarly, NTRU complements Kyber in offering low-latency encryption suitable for IoT and edge devices, which often rely on ML for local inference and security decisions. Both schemes have demonstrated resistance to side-channel attacks and ML-based cryptanalysis <xref ref-type="bibr" rid="scirp.144967-49">
      [49]
     </xref>.</p>
    <p>Many PQC schemes—including Kyber and NTRU—depend on efficient Number Theoretic Transform (NTT) implementations to accelerate polynomial multiplication, a core operation in lattice-based cryptography. High-performance NTT architectures such as pipelined R2MDC improve encryption and decryption speeds, enabling real-time operation in ML-driven security systems, especially those deployed in resource-constrained environments. Efficient NTT integration ensures that cryptographic operations do not become performance bottlenecks in ML workflows, allowing seamless real-time key exchange and digital signature generation—key elements in secure ML inference engines and autonomous systems <xref ref-type="bibr" rid="scirp.144967-51">
      [51]
     </xref>.</p>
    <p>Beyond encryption, ML can enhance system resilience by enabling adaptive responses to hardware faults or attacks. ML models can monitor cryptographic systems, analyze error patterns, and automatically trigger mitigation protocols. This approach is particularly effective for securing lightweight block ciphers like LED and HIGHT, which are often deployed in embedded systems with limited computational capacity. These adaptive techniques allow for the development of self-healing security architectures, which maintain cryptographic integrity despite environmental faults or malicious interference <xref ref-type="bibr" rid="scirp.144967-52">
      [52]
     </xref>.</p>
    <p>The importance of adopting quantum-resistant encryption was emphasized during the 2023 NIST Post-Quantum Cryptography Standardization process. Algorithms such as ML-KEM (based on Kyber) and ML-DSA (based on Dilithium) were selected as the future standards for secure communication. These algorithms are now being integrated into a wide range of ML-based systems, including smart devices, autonomous vehicles, and cloud services. As machine learning becomes more central to automation and critical infrastructure, combining it with PQC is key to building systems that are not only intelligent, but also secure and future-proof <xref ref-type="bibr" rid="scirp.144967-53">
      [53]
     </xref>.</p>
   </sec>
   <sec id="s4_4">
    <title>4.4. Challenges and Limitations</title>
    <p>Creating automatic systems for sound event detection is a complex task with several challenges, particularly related to sound characteristics, data collection, and annotation. Addressing these challenges is crucial for the effectiveness of machine learning techniques <xref ref-type="bibr" rid="scirp.144967-13">
      [13]
     </xref>. Recent advancements offer potential solutions to these issues, which are detailed below.</p>
    <p>A major challenge in sound event detection is handling overlapping sound events, a task known as polyphonic sound event detection. This involves identifying all coinciding sounds simultaneously <xref ref-type="bibr" rid="scirp.144967-54">
      [54]
     </xref>. To tackle this, supervised classification methods like RNNs <xref ref-type="bibr" rid="scirp.144967-55">
      [55]
     </xref> <xref ref-type="bibr" rid="scirp.144967-56">
      [56]
     </xref> and CNNs <xref ref-type="bibr" rid="scirp.144967-57">
      [57]
     </xref> are commonly used. These methods predict the presence of each sound event on a frame-by-frame basis, helping to manage the complexity of real-world sound environments.</p>
    <p>Data challenges, such as insufficient samples and class imbalance, also impact classifier performance. Data augmentation, which involves synthetically increasing the data through techniques like pitch shifting, noise removal, compression, and time stretching, is a key strategy to improve the system’s performance by enhancing data representation <xref ref-type="bibr" rid="scirp.144967-58">
      [58]
     </xref>.</p>
    <p>Traditional audio processing often separates feature representation from classifier design, which can lead to suboptimal features. Deep Neural Networks (DNNs) address this by integrating feature extraction and classification, optimizing both processes simultaneously. For example, in speech recognition, lower DNN layers adapt to speaker characteristics while upper layers focus on class discrimination <xref ref-type="bibr" rid="scirp.144967-42">
      [42]
     </xref>.</p>
    <p>Another challenge is weakly labeled data, where only the presence or absence of events is known, but not their exact timings. Multiple Instance Learning (MIL) is a useful approach for dealing with such data. MIL treats the entire audio clip as a “bag” for classification when annotations are partial, allowing for effective sound event detection despite weak labels <xref ref-type="bibr" rid="scirp.144967-59">
      [59]
     </xref> <xref ref-type="bibr" rid="scirp.144967-60">
      [60]
     </xref>.</p>
    <p>As machine learning becomes increasingly integrated into critical sectors such as healthcare, finance, and cybersecurity, new security concerns and vulnerabilities have emerged. This research <xref ref-type="bibr" rid="scirp.144967-48">
      [48]
     </xref> highlights several key challenges that affect the secure deployment and operation of ML systems:</p>
    <p>1) Emerging Vulnerabilities and Attack Surfaces: the widespread use of ML has introduced new entry points for malicious attacks. These vulnerabilities can be exploited to disrupt the integrity, availability, or confidentiality of ML systems. For example, attackers may take advantage of flaws in data preprocessing or the model architecture to alter outputs or extract sensitive training information.</p>
    <p>2) Privacy vs. Accuracy Trade-offs: a core dilemma in ML security lies in balancing data privacy with model performance. Techniques such as differential privacy aim to protect individual data by introducing noise, but this often comes at the cost of reduced accuracy. As the number of queries increases, the risk of privacy leaks or performance degradation also grows.</p>
    <p>3) Confidentiality and Intellectual Property Concerns: in applications involving sensitive or proprietary information, such as patient records or financial data, maintaining the confidentiality of ML models and their underlying data is critical. If an attacker gains access to model parameters, they may be able to reconstruct proprietary algorithms or extract confidential data, resulting in serious privacy breaches and intellectual property theft.</p>
    <p>The growing deployment of ML also exposes it to a variety of sophisticated attacks, as discussed in <xref ref-type="bibr" rid="scirp.144967-52">
      [52]
     </xref>:</p>
    <p>1) Adversarial Attacks: these attacks involve subtly altering input data to fool the model into making incorrect predictions. Such manipulations can undermine the reliability of ML-based security systems by causing false positives or negatives.</p>
    <p>2) Model Inversion Attacks: in this case, attackers analyze model outputs to reconstruct the original training data, potentially revealing personal or confidential information that was assumed to be protected.</p>
    <p>3) Data Poisoning Attacks: by injecting malicious or misleading data into the training set, attackers can corrupt the learning process. This compromises the integrity of the model and can be exploited to influence its behavior in harmful ways.</p>
    <p>4) Evasion Attacks: these attacks involve modifying malicious behavior or input data in a way that avoids detection by ML-based security systems. For example, attackers might disguise malware traffic to bypass an intrusion detection system trained on known patterns.</p>
    <p>5) Denial of Service (DoS) Attacks: ML systems can also be overwhelmed by large volumes of requests or data, causing them to slow down or crash. In critical environments, such as automated surveillance or fraud detection, such disruptions can have severe consequences.</p>
   </sec>
  </sec><sec id="s5">
   <title>5. Speech Recognition</title>
   <p>Recent advancements in speech recognition have enhanced communication across languages and interactions with various devices. These improvements stem from three main factors: increasing computational power from multi-core processors and GPU clusters, access to extensive datasets, and the rise of mobile, wearable, and smart home technologies <xref ref-type="bibr" rid="scirp.144967-61">
     [61]
    </xref>.</p>
   <p>Speech technology impacts both human-to-human (HHC) and human-to-machine communication (HMC). In HHC, research has focused on assisting individuals with speech impairments through applications like Google’s Speech API, which converts speech to text with high accuracy <xref ref-type="bibr" rid="scirp.144967-62">
     [62]
    </xref>. Studies also address transcription for the deaf and hard of hearing and explore deep learning models for language translation <xref ref-type="bibr" rid="scirp.144967-63">
     [63]
    </xref> <xref ref-type="bibr" rid="scirp.144967-64">
     [64]
    </xref>.</p>
   <p>In HMC, advancements include voice-activated smart home systems, though these often overlook less common languages like Romanian. To address this, new acoustic and grammar models for Romanian have been developed, alongside remote speaker recognition techniques <xref ref-type="bibr" rid="scirp.144967-65">
     [65]
    </xref> <xref ref-type="bibr" rid="scirp.144967-66">
     [66]
    </xref>. Additionally, the field covers voice search and interaction with mobile devices and entertainment systems <xref ref-type="bibr" rid="scirp.144967-67">
     [67]
    </xref>.</p>
   <p>This paper emphasizes automatic speech recognition as a critical component in spoken language technology, focusing on its role in real-time audio transcription.</p>
   <sec id="s5_1">
    <title>Automatic Speech Recognition</title>
    <p>In this book <xref ref-type="bibr" rid="scirp.144967-61">
      [61]
     </xref>, the architecture of automatic speech recognition (ASR) is elucidated, comprising four primary components: Signal processing and feature extraction, acoustic model (AM), language model (LM), and hypothesis search. Signal processing readies audio input by converting it into feature vectors. The AM evaluates the likelihood of feature sequences, while the LM estimates word sequence probabilities. The hypothesis search combines AM and LM scores to yield the recognition result. The AM must address challenges such as variable-length feature vectors and acoustic variability, stemming from factors like speaker characteristics, speech style, noise, and accents. Real-world ASR systems encounter further complexities, including extensive vocabularies, spontaneous speech, and multilingual contexts. While traditional ASR systems employed features like MFCC and RASTA-PLP with GMM-HMM models trained via maximum likelihood criteria, recent advances include discriminative training methods such as MCE and MPE in the 2000s and the adoption of DNNs and discriminative hierarchical models like CD-DNN-HMM, substantially improving accuracy due to enhanced computational capabilities and more extensive training data.</p>
   </sec>
  </sec><sec id="s6">
   <title>6. Related Applications for Deaf and Hard-of-Hearing Individuals</title>
   <p>This section will present research papers and applications aimed at improving the lives of deaf and hard-of-hearing individuals. Numerous studies focus on leveraging recent technologies, such as sound detection and speech recognition, to enhance their safety and overall quality of life.</p>
   <sec id="s6_1">
    <title>6.1. Related Research Papers</title>
    <p>The Enssat application <xref ref-type="bibr" rid="scirp.144967-68">
      [68]
     </xref> utilizes Google Glass as a wearable device to support individuals who are deaf or hard of hearing in both Arabic and English. This application primarily focuses on functions such as sound detection and speech recognition. The sound detection feature of Enssat utilizes the microphone of either a mobile phone or Google Glass to identify ambient sounds. When a sound is detected, a snippet of it is recorded and sent to another thread for identification. The system then compares this recorded sound with a set of stored sounds using the MusicG library to determine the degree of similarity. Regarding speech recognition, Enssat offers real-time transcription of speech. It employs Google’s Speech-to-Text service to convert audio files into text. The challenge lies in effectively processing continuous speech and segmenting it into snippets for accurate transcription. Additionally, the Enssat application provides translation capabilities, enabling the translation of both spoken words and text captured in images. To achieve this, the system utilizes Google’s Translation API to deliver real-time translation services.</p>
    <p>A different document explores the development of a specialized virtual assistant catering to individuals with hearing impairments, highlighting its proficiency in identifying and categorizing various sounds. The proposed remedy encompasses a sound classification module, a gesture recognition module, and a multilingual translation module. The sound classification module is engineered to recognize and categorize diverse sounds, such as those produced by vehicles, to notify users of potential hazards. It leverages audio data from the UrbanSound8K dataset and employs a deep neural network for sound recognition. The gesture recognition module translates gestures from Indian Sign Language into text and audio, facilitating communication between non-deaf individuals and those with hearing impairments. The multilingual translation module converts the generated text into various regional languages, offering translation services for hearing-impaired individuals in India. This solution is seamlessly integrated into an Android application and has undergone a comparative analysis with existing apps, assessing factors such as response time, accuracy, output predictions, and alert systems <xref ref-type="bibr" rid="scirp.144967-6">
      [6]
     </xref>.</p>
    <p>Saifan presents the Deaf Assistant Digital System <xref ref-type="bibr" rid="scirp.144967-39">
      [39]
     </xref>, a solution utilizing smartphones to provide alerts for individuals with hearing impairments across various situations. Employing vibrations and visual notifications on the smartphone screen, the system ensures effective communication with the user. The paper delves into the technical intricacies of the speech and sound recognition engines, detailing the utilization of deep auto-encoder-based low-dimensional feature extraction from FFT spectral envelopes. This approach enables the identification of diverse sounds and words. Additionally, the paper highlights the application of Praat, a computer program for speech analysis, synthesis, and manipulation, in extracting sound features. The system’s matching engines play a crucial role in comparing recognized words and sounds with predefined cautionary words and sound alerts. Upon identifying a match, these engines trigger appropriate actions such as vibration and visual effects to alert individuals with hearing impairments.</p>
    <p>In this paper <xref ref-type="bibr" rid="scirp.144967-69">
      [69]
     </xref>, the significance of sound detection is examined, and diverse technologies and systems designed for this objective are explored. The paper emphasizes the growing presence of comprehensive sound detection systems in the market. In the realm of sound detection technology, a proof-of-concept for a sound detection algorithm based on Gaussian Mixture Model (GMM) is discussed. Additionally, the paper briefly touches upon the application of Gaussian Mixture Models for speaker identification and verification in the context of speech recognition.</p>
    <p>The iHelp application <xref ref-type="bibr" rid="scirp.144967-70">
      [70]
     </xref> introduces a real-time mobile emergency assistance system designed to aid deaf-mute individuals or elderly individuals living alone in promptly and effectively reporting emergencies. This system employs mobile application software installed on smartphones, enabling users to report emergencies through SMS, even in the absence of internet access. By doing so, it optimizes the dispatching of rescue units and enhances the overall success rate of emergency rescue operations. The system is comprised of three key components: the report subsystem, dispatch system, and rescue subsystem.</p>
   </sec>
   <sec id="s6_2">
    <title>6.2. Related Mobile Applications</title>
    <p>The Sound Alert App functions as a tool for capturing and informing users of significant environmental and household occurrences. It can identify various sounds like doorbells, phone rings, microwave beeps, alarms, and intercoms without the need for pre-recording. What distinguishes this solution is its seamless integration with existing building infrastructure and alarm systems, providing a cost-effective alternative to flashy lights and expensive hardware setups. This app proves especially beneficial for individuals who are hard-of-hearing, deaf, elderly, or heavy sleepers. By activating “Detection Mode,” the app’s intelligent algorithm continuously monitors the environment through the smartphone’s microphone. Additionally, it can sync with Pebble Watch for extra notification options, including vibration, flashing lights, and on-screen icons with event names <xref ref-type="bibr" rid="scirp.144967-71">
      [71]
     </xref>. Another similar app, The Deaf and Hearing Impaired (APK), is designed to assist deaf or hearing-impaired individuals by alerting them through vibration and flashlight signals when a loud sound occurs nearby <xref ref-type="bibr" rid="scirp.144967-72">
      [72]
     </xref>.</p>
    <p>The Android app Live Transcribe &amp; Sound Notifications enhances accessibility for individuals with hearing impairments. It enables real-time transcription of spoken words in over 80 languages and dialects, allowing users to customize word additions. The app also notifies users of various sounds, including potentially risky situations. Users have the flexibility to adjust settings, save transcriptions for three days, and search within saved transcriptions. Developed in collaboration with Gallaudet University, a leading institution for the deaf and hard of hearing, the app is compatible with Android 6.0 and newer devices. It incorporates features such as vibrating when the user’s name is spoken and supports external microphones for improved audio reception <xref ref-type="bibr" rid="scirp.144967-73">
      [73]
     </xref>.</p>
    <p>Rogervoice, a revolutionary call transcription application, has transformed phone communication for individuals who are deaf or hard of hearing. By offering real-time call subtitles in over 80 languages, it enables users to independently connect with family, friends, medical professionals, and customer service helplines. The app is designed to be user-friendly, allowing calls to be initiated either from contacts or by entering numbers manually. Conversations are transcribed instantly, and users have the option to respond through speech or typing, with a voice synthesizer delivering text messages. It’s important to note that Rogervoice does not support emergency calls or premium-rate numbers. Subscriptions are required for calls to individuals who do not use the application, and pricing details can be found on the website <xref ref-type="bibr" rid="scirp.144967-74">
      [74]
     </xref>.</p>
    <p>In the pursuit of modernizing communication within the realm of security and improving the efficiency of security personnel in managing emergency reports, the General Directorate of Public Security has introduced the “Kulluna Amn” mobile application. This application is designed to actively involve citizens and residents in the security framework and has been launched with the direct support and guidance of HRH Prince Mohammed bin Nayef, Deputy Prime Minister and Minister of Interior. The app empowers users to report unusual incidents, which are then transmitted to the thirty-nine operation rooms situated across the kingdom. Users can furnish details about incidents, including photos and GPS location, choose the incident category, and even pinpoint the nearest police or traffic department based on their current geographical coordinates. “Kulluna Amn” signifies a noteworthy stride in enhancing emergency response and encouraging public participation in upholding security <xref ref-type="bibr" rid="scirp.144967-75">
      [75]
     </xref>.</p>
    <p>The TapSOS application serves as a crucial tool for reaching Emergency Services in situations where verbal communication is challenging or unsafe, especially for individuals with hearing impairments. Users establish profiles containing essential information, which is communicated to Emergency Call Handlers through visual icons. A medical profile provides valuable information for First Responders. The GPS feature automatically identifies the user’s location, allowing manual adjustments for precision. By responding to a series of questions aligned with Emergency Services protocols, users trigger alerts that are directly transmitted to the UK’s 999 Emergency Call Handlers. This makes TapSOS an indispensable tool for non-verbal emergency communication <xref ref-type="bibr" rid="scirp.144967-76">
      [76]
     </xref>.</p>
   </sec>
   <sec id="s6_3">
    <title>6.3. Comprehensive Comparsion</title>
    <p>This section explores a variety of research papers and mobile applications designed to improve the lives of deaf and hard-of-hearing individuals. These solutions leverage technologies such as sound detection, speech recognition, and real-time transcription to provide timely alerts, enhance communication, and increase situational awareness. Applications like Sound Alert App <xref ref-type="bibr" rid="scirp.144967-71">
      [71]
     </xref>, Live Transcribe &amp; Sound Notifications <xref ref-type="bibr" rid="scirp.144967-73">
      [73]
     </xref>, and Rogervoice <xref ref-type="bibr" rid="scirp.144967-74">
      [74]
     </xref> offer real-time features powered by cloud-based services, while others like iHelp <xref ref-type="bibr" rid="scirp.144967-70">
      [70]
     </xref> operate offline to ensure accessibility in emergencies. Solutions such as Saifan’s Deaf Assistant <xref ref-type="bibr" rid="scirp.144967-39">
      [39]
     </xref> and Enssat <xref ref-type="bibr" rid="scirp.144967-68">
      [68]
     </xref> integrate deep learning for accurate sound classification and multilingual support, including Arabic. Meanwhile, research initiatives like the Virtual Assistant with gesture recognition <xref ref-type="bibr" rid="scirp.144967-6">
      [6]
     </xref> and YAMNet-based firearm detection system <xref ref-type="bibr" rid="scirp.144967-69">
      [69]
     </xref> demonstrate the effectiveness of deep learning in both general and specialized sound classification tasks. Other apps, such as Deaf and Hearing Impaired (APK) <xref ref-type="bibr" rid="scirp.144967-72">
      [72]
     </xref>, Kulluna Amn <xref ref-type="bibr" rid="scirp.144967-75">
      [75]
     </xref>, and TapSOS <xref ref-type="bibr" rid="scirp.144967-76">
      [76]
     </xref>, emphasize safety, emergency reporting, and user-friendly design for accessible communication. Collectively, these solutions highlight the diversity of assistive technologies, balancing real-time performance, offline functionality, localization, and usability to support users with hearing impairments. <xref ref-type="table" rid="table2">
      Table 2
     </xref> presents a summary and comparison of these applications and studies. <xref ref-type="table" rid="table3">
      Table 3
     </xref> classifies them based on core assistive features such as sound detection, live transcription, Arabic support, emergency services, user interface design, and alert notifications.</p>
    <table-wrap id="table2">
     <label>
      <xref ref-type="table" rid="table2">
       Table 2
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.144967-"></xref>Table 2. Summary of related applications and research papers.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td aleft" width="14.71%"><p style="text-align:left">Category</p></td> 
       <td class="custom-bottom-td aleft" width="20.59%" colspan="2"><p style="text-align:left">Application/Research Paper</p></td> 
       <td class="custom-bottom-td aleft" width="52.03%"><p style="text-align:left">Description</p></td> 
       <td class="custom-bottom-td aleft" width="12.67%"><p style="text-align:left">Reference</p></td> 
      </tr> 
      <tr> 
       <td rowspan="4" class="custom-top-td tbtextacenter" width="14.71%"><p style="text-align:center">Related Research Papers</p></td> 
       <td class="custom-top-td acenter" width="20.59%" colspan="2"><p style="text-align:center">Enssat Application</p></td> 
       <td class="custom-top-td acenter" width="52.03%"><p style="text-align:center">Utilizes Google Glass for sound detection and speech recognition, offers real-time transcription and translation.</p></td> 
       <td class="custom-top-td acenter" width="12.67%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.144967-68">
          [68]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="20.59%" colspan="2"><p style="text-align:center">Specialized Virtual Assistant</p></td> 
       <td class="acenter" width="52.03%"><p style="text-align:center">Features sound classification, gesture recognition, and multilingual translation for individuals with hearing impairments.</p></td> 
       <td class="acenter" width="12.67%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.144967-6">
          [6]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="20.59%" colspan="2"><p style="text-align:center">Deaf Assistant Digital System</p></td> 
       <td class="acenter" width="52.03%"><p style="text-align:center">Uses smartphones to provide alerts via vibrations and visual notifications, with sound and speech recognition.</p></td> 
       <td class="acenter" width="12.67%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.144967-39">
          [39]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="20.59%" colspan="2"><p style="text-align:center">Sound Detection Technology</p></td> 
       <td class="acenter" width="52.03%"><p style="text-align:center">Discusses a proof-of-concept for sound detection using Gaussian Mixture Models (GMM).</p></td> 
       <td class="acenter" width="12.67%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.144967-69">
          [69]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td tbtextacenter" width="14.71%"><p style="text-align:center"></p></td> 
       <td class="custom-bottom-td acenter" width="19.27%"><p style="text-align:center">iHelp Application</p></td> 
       <td class="custom-bottom-td acenter" width="53.35%" colspan="2"><p style="text-align:center">Mobile emergency assistance system for deaf-mute or elderly individuals, includes reporting emergencies via SMS.</p></td> 
       <td class="custom-bottom-td acenter" width="12.67%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.144967-70">
          [70]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td rowspan="6" class="custom-top-td tbtextacenter" width="14.71%"><p style="text-align:center">Related Mobile Applications</p></td> 
       <td class="custom-top-td acenter" width="19.27%"><p style="text-align:center">Sound Alert App</p></td> 
       <td class="custom-top-td acenter" width="53.35%" colspan="2"><p style="text-align:center">Identifies and alerts users to significant sounds, integrates with building infrastructure, and supports Pebble Watch.</p></td> 
       <td class="custom-top-td acenter" width="12.67%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.144967-71">
          [71]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="19.27%"><p style="text-align:center">The Deaf and Hearing Impaired (APK)</p></td> 
       <td class="acenter" width="53.35%" colspan="2"><p style="text-align:center">Alerts users to loud sounds nearby through vibrations and flashlight signals.</p></td> 
       <td class="acenter" width="12.67%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.144967-72">
          [72]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="19.27%"><p style="text-align:center">Live Transcribe &amp; Sound Notifications</p></td> 
       <td class="acenter" width="53.35%" colspan="2"><p style="text-align:center">Provides real-time transcription of spoken words and notifications of various sounds in over 80 languages.</p></td> 
       <td class="acenter" width="12.67%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.144967-73">
          [73]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="19.27%"><p style="text-align:center">Rogervoice</p></td> 
       <td class="acenter" width="53.35%" colspan="2"><p style="text-align:center">Real-time call transcription application with subtitles in over 80 languages, enabling phone communication.</p></td> 
       <td class="acenter" width="12.67%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.144967-74">
          [74]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="19.27%"><p style="text-align:center">Kulluna Amn</p></td> 
       <td class="acenter" width="53.35%" colspan="2"><p style="text-align:center">Mobile app for reporting incidents to security forces, includes GPS location and incident details.</p></td> 
       <td class="acenter" width="12.67%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.144967-75">
          [75]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="19.27%"><p style="text-align:center">TapSOS</p></td> 
       <td class="acenter" width="53.35%" colspan="2"><p style="text-align:center">Emergency communication app for non-verbal communication, provides visual icons and GPS location for emergency services.</p></td> 
       <td class="acenter" width="12.67%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.144967-76">
          [76]
         </xref></p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p><sup>a</sup>This table reviews technologies and research designed to improve the lives of deaf and hard-of-hearing individuals, focusing on sound detection and speech recognition.</p>
    <table-wrap id="table3">
     <label>
      <xref ref-type="table" rid="table3">
       Table 3
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.144967-"></xref>Table 3. Feature classification of applications and research.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td aleft" width="16.19%"><p style="text-align:left">Features</p></td> 
       <td class="custom-bottom-td aleft" width="10.29%"><p style="text-align:left">Sound detection</p></td> 
       <td class="custom-bottom-td aleft" width="16.18%"><p style="text-align:left">Live transcriptions during the call</p></td> 
       <td class="custom-bottom-td aleft" width="11.77%"><p style="text-align:left">Support Arabic language interface</p></td> 
       <td class="custom-bottom-td aleft" width="11.77%"><p style="text-align:left">Emergency call</p></td> 
       <td class="custom-bottom-td aleft" width="10.29%"><p style="text-align:left">Easy to use &amp; attractive interface</p></td> 
       <td class="custom-bottom-td aleft" width="11.77%"><p style="text-align:left">Android Application</p></td> 
       <td class="custom-bottom-td aleft" width="11.75%"><p style="text-align:left">Alert notification</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="16.19%"><p style="text-align:center">Enssat</p></td> 
       <td class="custom-top-td acenter" width="10.29%"><p style="text-align:center">√</p></td> 
       <td class="custom-top-td acenter" width="16.18%"><p style="text-align:center"></p></td> 
       <td class="custom-top-td acenter" width="11.77%"><p style="text-align:center">√</p></td> 
       <td class="custom-top-td acenter" width="11.77%"><p style="text-align:center"></p></td> 
       <td class="custom-top-td acenter" width="10.29%"><p style="text-align:center"></p></td> 
       <td class="custom-top-td acenter" width="11.77%"><p style="text-align:center"></p></td> 
       <td class="custom-top-td acenter" width="11.75%"><p style="text-align:center">√</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="16.19%"><p style="text-align:center">Virtual Assistant for Hearing Impaired</p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="16.18%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="11.75%"><p style="text-align:center">√</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="16.19%"><p style="text-align:center">Deaf Assistant Digital System</p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="16.18%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="11.75%"><p style="text-align:center">√</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="16.19%"><p style="text-align:center">Sound Detection Technology</p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="16.18%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="11.75%"><p style="text-align:center">√</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="16.19%"><p style="text-align:center">iHelp</p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="16.18%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.75%"><p style="text-align:center">√</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="16.19%"><p style="text-align:center">Sound Alert</p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="16.18%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.75%"><p style="text-align:center">√</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="16.19%"><p style="text-align:center">The Deaf and Hearing Impaired (APK)</p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="16.18%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="11.75%"><p style="text-align:center">√</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="16.19%"><p style="text-align:center">Google live transcribe</p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="16.18%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="11.75%"><p style="text-align:center">√</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="16.19%"><p style="text-align:center">Rogervoice</p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="16.18%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="11.75%"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="16.19%"><p style="text-align:center">Kulluna Amn</p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="16.18%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="11.75%"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="16.19%"><p style="text-align:center">TapSOS</p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="16.18%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="10.29%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="11.77%"><p style="text-align:center">√</p></td> 
       <td class="acenter" width="11.75%"><p style="text-align:center"></p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p><sup>a</sup>This table compares the related studies and applications, based on functionality and features.</p>
   </sec>
  </sec><sec id="s7">
   <title>7. Discussions</title>
   <p>Based on the discussion presented in this review paper, it can be concluded that artificial intelligence plays a crucial role in enabling real-time alarm sound detection applications designed to support individuals who are deaf or hard of hearing and speak Arabic. Such applications can significantly facilitate communication with government agencies in Saudi Arabia during emergency situations. In other words, an effective approach would be to design an integrated application that combines key elements from existing solutions. The application should:</p>
   <p>Incorporate Sound Detection Technology: Utilize algorithms similar to those in the Sound Alert App and The Deaf and Hearing-Impaired app to detect alarm sounds. This will enable the app to identify and alert users about important environmental and home events, specifically tailored to the needs of deaf and hard-of-hearing individuals.</p>
   <p>Implement Speech Recognition and Transcription Features: Drawing from Live Transcribe &amp; Sound Notifications and Rogervoice, integrate real-time transcription of spoken words in Arabic. This feature will aid in communication during emergency situations, enabling individuals to understand spoken information.</p>
   <p>Facilitate Communication with Government Agencies: Inspired by Kulluna Amn &amp; TapSOS, the application should allow users to communicate with the relevant authorities, specifically in Saudi Arabia. This ensures a direct link to the appropriate agencies during emergencies.</p>
   <p>Combining these elements into a single application tailored to the Arabic-speaking population, particularly in Saudi Arabia, would address the core challenges of sound detection, communication during emergencies, and interaction with relevant government agencies for the deaf and hard-of-hearing community.</p>
  </sec><sec id="s8">
   <title>8. Conclusions</title>
   <p>By utilizing breakthroughs in artificial intelligence and machine learning, modern assistive hearing technologies have made significant progress over conventional auditory aids. These modern systems have better accuracy, more contextual understanding, and better security measures. Today’s remedies, unlike earlier standalone devices, are connected, flexible, and significantly more sophisticated, signaling a radical change in the way hearing aid is provided.</p>
   <p>In summary, incorporating advanced technologies especially in the areas of machine learning and artificial intelligence shows a lot of potentials to improve the safety and the quality of life of hearing-impaired individuals. This paper has analyzed assistive technologies, sound recognition, and sonification for the enhancement of the functionality of patients suffering from hearing loss. Hearing aids, cochlear implants, and mobile apps breakdown communication barriers and promote self-sufficiency, effectively changing the status quo of everyday routines. On the other hand, machine learning-centered technologies for sound detection provide advanced means of real-time situational awareness for safety, security, and environmental monitoring.</p>
   <p>The key findings of this review highlight the transformative impact of these technologies. Augmented hearing aids and cochlear implants are essential; both assist in increasing overall auditory fulfillment, while empowering a sense of independence. Sound detection systems driven by machine learning based models, generate alerts in time and also provide contextual information when required that specifically addresses the safety problems encountered to people with hearing impairments. Further, the evolution of speech recognition technologies enables communication across disparate contexts enabling greater context-aware support for integration and cross-interaction within different environments.</p>
   <p>There is still a lot of room for innovation by using cutting-edge ML approaches, even with these advancements. Future research should concentrate on a few key areas in order to overcome existing constraints and increase the functionality of assistive audio devices:</p>
   <p>By integrating these machine learning techniques, future assistive hearing devices can offer more individualized, safe, and dependable assistance, enabling people with hearing loss to live more independently and safely in their surroundings.</p>
  </sec>
 </body><back>
  <ref-list>
   <title>References</title>
   <ref id="scirp.144967-ref1">
    <label>1</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     World Health Organization (2023) Deafness and Hearing Loss. &gt;https://www.who.int/news-room/fact-sheets/detail/deafness-and-hearing-loss 
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref2">
    <label>2</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Abdallah, E.E. and Fayyoumi, E. (2016) Assistive Technology for Deaf People Based on Android Platform. Procedia Computer Science, 94, 295-301. &gt;https://doi.org/10.1016/j.procs.2016.08.044 
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref3">
    <label>3</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     World Health Organization (2018) Improving Access to Assistive Technology. &gt;http://apps.who.int/gb/ebwha/pdf_files/WHA71/A71_21-en.pdf 
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref4">
    <label>4</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kbar, G., Bhatia, A., Abidi, M.H. and Alsharawy, I. (2016) Assistive Technologies for Hearing, and Speaking Impaired People: A Survey. Disability and Rehabilitation: Assistive Technology, 12, 3-20. &gt;https://doi.org/10.3109/17483107.2015.1129456
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref5">
    <label>5</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Advanced Solutions International, Inc. (2012) Educational Technology. The Delta Kappa Gamma Bulletin. &gt;https://www.media.mit.edu/~mres/papers/educational-technology-2012.pdf 
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref6">
    <label>6</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ozarkar, S., Chetwani, R., Devare, S., Haryani, S. and Giri, N. (2020) AI for Accessibility: Virtual Assistant for Hearing Impaired. 2020 11th International Conference on Computing, Communication and Networking Technologies (ICCCNT), Kharagpur, 1-3 July 2020, 1-7. &gt;https://doi.org/10.1109/icccnt49239.2020.9225392
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref7">
    <label>7</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Alqudaihi, K.S., Aslam, N., Khan, I.U., Almuhaideb, A.M., Alsunaidi, S.J., Ibrahim, N.M.A.R., et al. (2021) Cough Sound Detection and Diagnosis Using Artificial Intelligence Techniques: Challenges and Opportunities. IEEE Access, 9, 102327-102344. &gt;https://doi.org/10.1109/access.2021.3097559
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref8">
    <label>8</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Li, J., Dai, W., Metze, F., Qu, S. and Das, S. (2017) A Comparison of Deep Learning Methods for Environmental Sound Detection. 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, 5-9 March 2017, 126-130. &gt;https://doi.org/10.1109/icassp.2017.7952131
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref9">
    <label>9</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Lee, Y., Shariatfar, M., Rashidi, A. and Lee, H.W. (2020) Evidence-Driven Sound Detection for Prenotification and Identification of Construction Safety Hazards and Accidents. Automation in Construction, 113, Article 103127. &gt;https://doi.org/10.1016/j.autcon.2020.103127
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref10">
    <label>10</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Virtanen, T., Plumbley, M.D. and Ellis, D. (2017) Introduction to Sound Scene and Event Analysis. In: Virtanen, T., Plumbley, M. and Ellis, D., Eds., Computational Analysis of Sound Scenes and Events, Springer International Publishing, 3-12. &gt;https://doi.org/10.1007/978-3-319-63450-0_1
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref11">
    <label>11</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Rebala, G., Ravi, A. and Churiwala, S. (2019) Machine Learning Definition and Basics. In: Rebala, G., Ravi, A. and Churiwala, S., An Introduction to Machine Learning, Springer International Publishing, 1-17. &gt;https://doi.org/10.1007/978-3-030-15729-6_1
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref12">
    <label>12</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Babaee, E., Anuar, N.B., Abdul Wahab, A.W., Shamshirband, S. and Chronopoulos, A.T. (2017) An Overview of Audio Event Detection Methods from Feature Extraction to Classification. Applied Artificial Intelligence, 31, 661-714. &gt;https://doi.org/10.1080/08839514.2018.1430469
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref13">
    <label>13</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Mesaros, A., Heittola, T., Virtanen, T. and Plumbley, M.D. (2021) Sound Event Detection: A Tutorial. IEEE Signal Processing Magazine, 38, 67-83. &gt;https://doi.org/10.1109/msp.2021.3090678
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref14">
    <label>14</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Min, K., Jung, M., Kim, J. and Chi, S. (2018) Sound Event Recognition-Based Classification Model for Automated Emergency Detection in Indoor Environment. In: Mutis, I. and Hartmann, T., Eds., Advances in Informatics and Computing in Civil and Construction Engineering, Springer International Publishing, 529-535. &gt;https://doi.org/10.1007/978-3-030-00220-6_63
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref15">
    <label>15</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Rahman, S.U., Khan, A., Abbas, S., Alam, F. and Rashid, N. (2020) Hybrid System for Automatic Detection of Gunshots in Indoor Environment. Multimedia Tools and Applications, 80, 4143-4153. &gt;https://doi.org/10.1007/s11042-020-09936-w
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref16">
    <label>16</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Souli, S. and Lachiri, Z. (2018) Audio Sounds Classification Using Scattering Features and Support Vectors Machines for Medical Surveillance. Applied Acoustics, 130, 270-282. &gt;https://doi.org/10.1016/j.apacoust.2017.08.002
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref17">
    <label>17</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Singh, D. and Singh, B. (2020) Investigating the Impact of Data Normalization on Classification Performance. Applied Soft Computing, 97, Article 105524. &gt;https://doi.org/10.1016/j.asoc.2019.105524
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref18">
    <label>18</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Caixinha, M. and Nunes, S. (2016) Machine Learning Techniques in Clinical Vision Sciences. Current Eye Research, 42, 1-15. &gt;https://doi.org/10.1080/02713683.2016.1175019
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref19">
    <label>19</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Fink, D. (2019) A New Definition of Noise: Noise Is Unwanted and/or Harmful Sound. Noise Is the New ‘Secondhand Smoke’. Proceedings of Meetings on Acoustics, 39, Article 050002. &gt;https://doi.org/10.1121/2.0001186
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref20">
    <label>20</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Cadena, L., et al. (2017) Noise Reduction Techniques for Processing of Medical Images. Noise Reduction Techniques for Processing of Medical Images. &gt;http://www.iaeng.org/publication/WCE2017/WCE2017_pp496-500.pdf 
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref21">
    <label>21</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Celin, S. and Vasanth, K. (2018) ECG Signal Classification Using Various Machine Learning Techniques. Journal of Medical Systems, 42, Article No. 241. &gt;https://doi.org/10.1007/s10916-018-1083-6
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref22">
    <label>22</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Deshmukh, G., Gaonkar, A., Golwalkar, G. and Kulkarni, S. (2019) Speech Based Emotion Recognition Using Machine Learning. 2019 3rd International Conference on Computing Methodologies and Communication (ICCMC), Erode, 27-29 March 2019, 812-817. &gt;https://doi.org/10.1109/iccmc.2019.8819858
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref23">
    <label>23</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Adhi Pramono, R.X., Anas Imtiaz, S. and Rodriguez-Villegas, E. (2019) Automatic Identification of Cough Events from Acoustic Signals. 2019 41st Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), Berlin, 23-27 July 2019, 217-220. &gt;https://doi.org/10.1109/embc.2019.8856420
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref24">
    <label>24</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Cohen-McFarlane, M., Goubran, R. and Knoefel, F. (2019) Comparison of Silence Removal Methods for the Identification of Audio Cough Events. 2019 41st Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), Berlin, 23-27 July 2019, 1263-1268. &gt;https://doi.org/10.1109/embc.2019.8857889
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref25">
    <label>25</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Hall, J.I., Lozano, M., Estrada-Petrocelli, L., Birring, S. and Turner, R. (2020) The Present and Future of Cough Counting Tools. Journal of Thoracic Disease, 12, 5207-5223. &gt;https://doi.org/10.21037/jtd-2020-icc-003
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref26">
    <label>26</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Tsalera, E., Papadakis, A., Samarakou, M. and Voyiatzis, I. (2022) CNN-Based Segmentation and Classification of Sound Streams under Realistic Conditions. Proceedings of the 26th Pan-Hellenic Conference on Informatics, Athens, 25-27 November 2022, 373-378. &gt;https://doi.org/10.1145/3575879.3576020
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref27">
    <label>27</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Sharma, G., Umapathy, K. and Krishnan, S. (2020) Trends in Audio Signal Feature Extraction Methods. Applied Acoustics, 158, Article 107020. &gt;https://doi.org/10.1016/j.apacoust.2019.107020
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref28">
    <label>28</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Baughman, A., Morales, E., Reiss, G., Greco, N., Hammer, S. and Wang, S. (2019) Detection of Tennis Events from Acoustic Data. Proceedings Proceedings of the 2nd International Workshop on Multimedia Content Analysis in Sports, Nice, 25 October 2019, 91-99. &gt;https://doi.org/10.1145/3347318.3355520
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref29">
    <label>29</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Tzanetakis, G. (2004) Manipulation, Analysis and Retrieval Systems for Audio Signals. ProQuest. 
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref30">
    <label>30</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ramadan, M., Bani Younes, A. and Moheidat, J. (2019) Robust Sound Detection&amp;Localization Algorithms for Robotics Applications. AIAA Scitech 2019 Forum. &gt;https://doi.org/10.2514/6.2019-2046
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref31">
    <label>31</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Jain, D., Huynh Anh Nguyen, K., M. Goodman, S., Grossman-Kahn, R., Ngo, H., Kusupati, A., et al. (2022) ProtoSound: A Personalized and Scalable Sound Recognition System for Deaf and Hard-of-Hearing Users. CHI Conference on Human Factors in Computing Systems, New Orleans, 9 April 2022-5 May 2022, 1-16. &gt;https://doi.org/10.1145/3491102.3502020
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref32">
    <label>32</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Liu, J., You, M., Wang, Z., Li, G., Xu, X. and Qiu, Z. (2014) Cough Detection Using Deep Neural Networks. 2014 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), Belfast, 2-5 November 2014, 560-563. &gt;https://doi.org/10.1109/bibm.2014.6999220
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref33">
    <label>33</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Amoh, J. and Odame, K. (2016) Deep Neural Networks for Identifying Cough Sounds. IEEE Transactions on Biomedical Circuits and Systems, 10, 1003-1011. &gt;https://doi.org/10.1109/tbcas.2016.2598794
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref34">
    <label>34</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Venkataramanan, K. and Rajamohan, H.R. (2019) Emotion Recognition from Speech. arXiv:1912.10458&gt;https://arxiv.org/abs/1912.10458 
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref35">
    <label>35</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Pandey, S.K., Shekhawat, H.S. and Prasanna, S.R. (2019) Deep Learning Techniques for Speech Emotion Recognition: A Review. 2019 29th International Conference Radioelektronika (RADIOELEKTRONIKA). [Preprint]
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref36">
    <label>36</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Singh, D., et al. (2020) Deep Learning Techniques for Sound Classification. IEEE Transactions on Audio, Speech, and Language Processing, 28, 286-297.
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref37">
    <label>37</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Wang, K., An, N., Li, B.N., Zhang, Y. and Li, L. (2015) Speech Emotion Recognition Using Fourier Parameters. IEEE Transactions on Affective Computing, 6, 69-75. &gt;https://doi.org/10.1109/taffc.2015.2392101
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref38">
    <label>38</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Badshah, A.M., Ahmad, J., Rahim, N. and Baik, S.W. (2017) Speech Emotion Recognition from Spectrograms with Deep Convolutional Neural Network. 2017 International Conference on Platform Technology and Service (PlatCon), Busan, 13-15 February 2017, 1-5. &gt;https://doi.org/10.1109/platcon.2017.7883728
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref39">
    <label>39</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Saifan, R.R., Dweik, W. and Abdel‐Majeed, M. (2018) A Machine Learning Based Deaf Assistance Digital System. Computer Applications in Engineering Education, 26, 1008-1019. &gt;https://doi.org/10.1002/cae.21952
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref40">
    <label>40</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Permana, S.D.H., Saputra, G., Arifitama, B., Yaddarabullah, Caesarendra, W. and Rahim, R. (2022) Classification of Bird Sounds as an Early Warning Method of Forest Fires Using Convolutional Neural Network (CNN) Algorithm. Journal of King Saud University-Computer and Information Sciences, 34, 4345-4357. &gt;https://doi.org/10.1016/j.jksuci.2021.04.013
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref41">
    <label>41</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kattenborn, T., Leitloff, J., Schiefer, F. and Hinz, S. (2021) Review on Convolutional Neural Networks (CNN) in Vegetation Remote Sensing. ISPRS Journal of Photogrammetry and Remote Sensing, 173, 24-49. &gt;https://doi.org/10.1016/j.isprsjprs.2020.12.010
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref42">
    <label>42</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Purwins, H., Li, B., Virtanen, T., Schluter, J., Chang, S. and Sainath, T. (2019) Deep Learning for Audio Signal Processing. IEEE Journal of Selected Topics in Signal Processing, 13, 206-219. &gt;https://doi.org/10.1109/jstsp.2019.2908700
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref43">
    <label>43</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Arsenali, B., van Dijk, J., Ouweltjes, O., den Brinker, B., Pevernagie, D., Krijn, R., et al. (2018) Recurrent Neural Network for Classification of Snoring and Non-Snoring Sound Events. 2018 40th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), Honolulu, 18-21 July 2018, 328-331. &gt;https://doi.org/10.1109/embc.2018.8512251
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref44">
    <label>44</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Eyben, F. (2018) Real-Time Speech and Music Classification by Large Audio Feature Space Extraction. Springer International Publishing.
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref45">
    <label>45</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Dastres, R. and Soori, M. (2021) A review in Advanced Digital Signal Processing Systems. International Journal of Electrical and Computer Engineering, 15, 122-127. &gt;https://www.researchgate.net/publication/350449625_A_Review_in_Advanced_Digital_Signal_Processing_Systems 
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref46">
    <label>46</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kuo, S.M., Lee, B.H. and Tian, W. (2017) Real-Time Digital Signal Processing: Fundamentals, Implementations and Applications. John Wiley&amp;Sons.
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref47">
    <label>47</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Gutiérrez-Gnecchi, J.A., Morfin-Magaña, R., Lorias-Espinoza, D., Tellez-Anguiano, A.d.C., Reyes-Archundia, E., Méndez-Patiño, A., et al. (2017) Dsp-Based Arrhythmia Classification Using Wavelet Transform and Probabilistic Neural Network. Biomedical Signal Processing and Control, 32, 44-56. &gt;https://doi.org/10.1016/j.bspc.2016.10.005
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref48">
    <label>48</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Papernot, N., McDaniel, P., Sinha, A. and Wellman, M.P. (2018) SoK: Security and Privacy in Machine Learning. 2018 IEEE European Symposium on Security and Privacy (EuroS&amp;P), London, 24-26 April 2018, 399-414. &gt;https://doi.org/10.1109/eurosp.2018.00035
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref49">
    <label>49</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ehsan, M.A., Alayed, W., Rehman, A.U., ul Hassan, W. and Zeeshan, A. (2025) Post-quantum KEMs for IoT: A Study of Kyber and NTRU. Symmetry, 17, Article 881. &gt;https://doi.org/10.3390/sym17060881
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref50">
    <label>50</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Sanal, P., Karagoz, E., Seo, H., Azarderakhsh, R. and Mozaffari-Kermani, M. (2021) Kyber on ARM64: Compact Implementations of Kyber on 64-Bit ARM Cortex-A Processors. In: Garcia-Alfaro, J., Li, S., Poovendran, R., Debar, H. and Yung, M., Eds., Lecture Notes of the Institute for Computer Sciences, Social Informatics and Telecommunications Engineering, Springer International Publishing, 424-440. &gt;https://doi.org/10.1007/978-3-030-90022-9_23
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref51">
    <label>51</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kundi, D.E.S., Mera, J.M.B., Strub, P. and Hutter, M. (2024) High-Performance NTT Hardware Accelerator to Support ML-KEM and ML-DSA. Proceedings of the 2024 Workshop on Attacks and Solutions in Hardware Security, Salt Lake City, 14-18 October 2024, 100-105. &gt;https://doi.org/10.1145/3689939.3695785
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref52">
    <label>52</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Subramanian, S., Mozaffari-Kermani, M., Azarderakhsh, R. and Nojoumian, M. (2017) Reliable Hardware Architectures for Cryptographic Block Ciphers LED and Hight. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, 36, 1750-1758. &gt;https://doi.org/10.1109/tcad.2017.2661811
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref53">
    <label>53</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Darzi, S. and Yavuz, A.A. (2024) PQC Meets ML or AI: Exploring the Synergy of Machine Learning and Post-Quantum Cryptography. TechRxiv, IEEE Security and Privacy Magazine. &gt;https://www.techrxiv.org/users/711847/articles/699030-pqc-meets-ml-or-ai-exploring-the-synergy-of-machine-learning-and-post-quantum-cryptography 
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref54">
    <label>54</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Adavanne, S., Politis, A., Nikunen, J. and Virtanen, T. (2019) Sound Event Localization and Detection of Overlapping Sources Using Convolutional Recurrent Neural Networks. IEEE Journal of Selected Topics in Signal Processing, 13, 34-48. &gt;https://doi.org/10.1109/jstsp.2018.2885636
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref55">
    <label>55</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Parascandolo, G., Huttunen, H. and Virtanen, T. (2016) Recurrent Neural Networks for Polyphonic Sound Event Detection in Real Life Recordings. 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, 20-25 March 2016, 6440-6444. &gt;https://doi.org/10.1109/icassp.2016.7472917
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref56">
    <label>56</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Adavanne, S. et al. (2017) Sound Event Detection in Multichannel Audio Using Spatial and Harmonic Features. arXiv:1706.02293 &gt;https://arxiv.org/abs/1706.02293 
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref57">
    <label>57</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Phan, H., Hertel, L., Maass, M. and Mertins, A. (2016) Robust Audio Event Recognition with 1-Max Pooling Convolutional Neural Networks. INTERSPEECH 2016, San Francisco, 8-12 September 2016, 3653-3657. &gt;https://doi.org/10.21437/interspeech.2016-123
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref58">
    <label>58</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Abayomi-Alli, O.O., Damaševičius, R., Qazi, A., Adedoyin-Olowe, M. and Misra, S. (2022) Data Augmentation and Deep Learning Methods in Sound Classification: A Systematic Review. Electronics, 11, Article 3795. &gt;https://doi.org/10.3390/electronics11223795
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref59">
    <label>59</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kong, Q., Xu, Y., Wang, W. and Plumbley, M.D. (2020) Sound Event Detection of Weakly Labelled Data with CNN-Transformer and Automatic Threshold Optimization. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 28, 2450-2460. &gt;https://doi.org/10.1109/taslp.2020.3014737
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref60">
    <label>60</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Carbonneau, M., Cheplygina, V., Granger, E. and Gagnon, G. (2018) Multiple Instance Learning: A Survey of Problem Characteristics and Applications. Pattern Recognition, 77, 329-353. &gt;https://doi.org/10.1016/j.patcog.2017.10.009
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref61">
    <label>61</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Yu, D. and Deng, L. (2016) Automatic Speech Recognition: A Deep Learning Approach. Springer. &gt;https://doi.org/10.1145/3689939.3695785
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref62">
    <label>62</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Anggraini, N., Kurniawan, A., Wardhani, L.K. and Hakiem, N. (2018) Speech Recognition Application for the Speech Impaired Using the Android-Based Google Cloud Speech Api. TELKOMNIKA (Telecommunication Computing Electronics and Control), 16, Article 2733. &gt;https://doi.org/10.12928/telkomnika.v16i6.9638
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref63">
    <label>63</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Gaur, Y. (2015) The Effects of Automatic Speech Recognition Quality on Human Transcription Latency. Proceedings of the 17th International ACM SIGACCESS Conference on Computers&amp;Accessibility, Lisbon, 26-28 October 2015, 367-368.
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref64">
    <label>64</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Gibadullin, R.F., Perukhin, M.Y. and Ilin, A.V. (2021) Speech Recognition and Machine Translation Using Neural Networks. 2021 International Conference on Industrial Engineering, Applications and Manufacturing (ICIEAM), Sochi, 17-21 May 2021, 398-403. &gt;https://doi.org/10.1109/icieam51226.2021.9446474
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref65">
    <label>65</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Caranica, A., Cucu, H., Burileanu, C., Portet, F. and Vacher, M. (2017) Speech Recognition Results for Voice-Controlled Assistive Applications. 2017 International Conference on Speech Technology and Human-Computer Dialogue (SPED), Bucharest, 6-9 July 2017, 1-8. &gt;https://doi.org/10.1109/sped.2017.7990438
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref66">
    <label>66</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Peng, S., Lv, T., Han, X., Wu, S., Yan, C. and Zhang, H. (2019) Remote Speaker Recognition Based on the Enhanced LDV-Captured Speech. Applied Acoustics, 143, 165-170. &gt;https://doi.org/10.1016/j.apacoust.2018.08.007
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref67">
    <label>67</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     S, K. and E, C. (2016) A Review on Automatic Speech Recognition Architecture and Approaches. International Journal of Signal Processing, Image Processing and Pattern Recognition, 9, 393-404. &gt;https://doi.org/10.14257/ijsip.2016.9.4.34
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref68">
    <label>68</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Alkhalifa, S. and Al-Razgan, M. (2018) Enssat: Wearable Technology Application for the Deaf and Hard of Hearing. Multimedia Tools and Applications, 77, 22007-22031. &gt;https://doi.org/10.1007/s11042-018-5860-5
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref69">
    <label>69</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Bragg, D., Huynh, N. and Ladner, R.E. (2016) A Personalizable Mobile Sound Detector App Design for Deaf and Hard-of-Hearing Users. Proceedings of the 18th International ACM SIGACCESS Conference on Computers and Accessibility, Reno, 23-26 October 2016, 3-13. 
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref70">
    <label>70</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Chen, L., Tsai, C., Chang, W., Cheng, Y. and Li, K.S. (2016) A Real-Time Mobile Emergency Assistance System for Helping Deaf-Mute People/Elderly Singletons. 2016 IEEE International Conference on Consumer Electronics (ICCE), Las Vegas, 7-11 January 2016, 45-46. &gt;https://doi.org/10.1109/icce.2016.7430516
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref71">
    <label>71</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Erts to Sounds Around You (No Date) Sound Alert|Convert Sounds into Visual and Sensory Notifications. &gt;http://www.soundalert.co/ 
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref72">
    <label>72</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     APKPure.Com. (2015) The Deaf and Hearing Impaired APK for Android Download. &gt;https://apkpure.com/the-deaf-and-hearing-impaired/kr.ac.kaist.isilab.doyeob.humanityProject2 
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref73">
    <label>73</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Live Transcribe&amp;Notification—Apps on Google Play (No Date). Google. &gt;https://play.google.com/store/apps/details?id=com.google.audio.hearing.visualization.accessibility.scribe 
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref74">
    <label>74</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Caption All Your Phone Calls Instantly! (No Date). Rogervoice. &gt;https://rogervoice.com/en/ 
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref75">
    <label>75</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kamn كلنا أمن-التطبيقات على google play (No Date). Google. &gt;https://play.google.com/store/apps/details?id=sa.gov.moi.securityinform&amp;hl=ar 
    </mixed-citation>
   </ref>
   <ref id="scirp.144967-ref76">
    <label>76</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Home (2023) TapSOS. &gt;https://tapsos.com/
    </mixed-citation>
   </ref>
  </ref-list>
 </back>
</article>