<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">jcc</journal-id>
      <journal-title-group>
        <journal-title>Journal of Computer and Communications</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2327-5227</issn>
      <issn pub-type="ppub">2327-5219</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/jcc.2026.144005</article-id>
      <article-id pub-id-type="publisher-id">jcc-150860</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Computer Science</subject>
          <subject>Communications</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>A Review of Pedestrian Recognition Based on Millimeter-Wave Radar and Video Fusion</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Li</surname>
            <given-names>Hongtao</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Lu</surname>
            <given-names>Qimeng</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Gao</surname>
            <given-names>Huijia</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Xu</surname>
            <given-names>Beibei</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Chen</surname>
            <given-names>Zijia</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author" corresp="yes">
          <name name-style="western">
            <surname>Wang</surname>
            <given-names>Zhengjie</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> College of Electronic and Information Engineering, Shandong University of Science and Technology, Qingdao, China </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The authors declare no conflicts of interest regarding the publication of this paper.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>16</day>
        <month>04</month>
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>04</month>
        <year>2026</year>
      </pub-date>
      <volume>14</volume>
      <issue>04</issue>
      <fpage>103</fpage>
      <lpage>118</lpage>
      <history>
        <date date-type="received">
          <day>05</day>
          <month>04</month>
          <year>2026</year>
        </date>
        <date date-type="accepted">
          <day>19</day>
          <month>04</month>
          <year>2026</year>
        </date>
        <date date-type="published">
          <day>22</day>
          <month>04</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2026 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/jcc.2026.144005">https://doi.org/10.4236/jcc.2026.144005</self-uri>
      <abstract>
        <p>Person identification serves as a core supporting technology in public security, intelligent home applications, and person tracking in different environments. Achieving stable identity matching across environments and locations has become a critical research requirement. Single-modal identification technologies suffer from inherent limitations. The cross-modal fusion method based on millimeter-wave radar and video realizes the complementary advantages of the two modalities. It can effectively support person identification across scenarios, becoming a research hotspot in complex open scenarios. This paper systematically reviews the research progress and core technical systems in this field. Firstly, it sorts out the information extraction methods of millimeter-wave radar features and video visual features, summarizing the extraction ideas of key features such as range-Doppler, micro-Doppler, point cloud, and skeleton. Secondly, it summarizes three fusion strategies (data-level, feature-level, decision-level) and two core alignment methods (spatio-temporal alignment, feature distribution alignment). Then, it introduces typical applications, including intelligent security monitoring, human-computer interaction authentication, and multi-target tracking. It analyzes current challenges such as complex environmental interference, cross-location spatio-temporal alignment, feature distribution shift, and scarcity of large-scale data. Finally, it discusses future research directions, including scene-invariant feature learning, cross-domain alignment optimization, unified feature space construction, and few-shot generalization, providing a comprehensive reference for person identification research.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>Information Feature Extraction</kwd>
        <kwd>Feature Fusion and Alignment</kwd>
        <kwd>Cross-Modal Recognition and Matching</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction</title>
      <p>With the rapid development of intelligent security, intelligent transportation, and public service technologies, person identification has become a core supporting technology in various fields [<xref ref-type="bibr" rid="B1">1</xref>]. Systems capable of stably completing person identification in different environments and locations are increasingly important. Traditional identification mostly relies on single-modal sensing in fixed scenarios and devices, which makes it difficult to meet the identity verification and target retrieval requirements across locations, devices, and environments in squares, stations, communities, buildings, etc. Person identification systems based on different sensing modalities have been continuously developed and applied. At present, mainstream identification systems can be mainly divided into: wearable sensor-based systems, computer vision-based systems, millimeter-wave radar-based systems, and millimeter-wave radar and video fusion-based systems. </p>
      <p>Wearable sensor-based identification systems require users to wear devices integrated with sensors, such as fingerprints and heart rate, to collect biometric features [<xref ref-type="bibr" rid="B2">2</xref>]. However, wearable devices are inconvenient to use, high in maintenance costs, and cannot achieve non-cooperative cross-scenario identification, limiting their wide application in public areas. Unlike the cooperative identification that relies on active user cooperation in wearable sensor systems, the millimeter-wave radar and video fusion system features inherent non-cooperative sensing, with no need for on-body devices or active cooperation from targets. This characteristic is critical for identity verification in unconstrained public security scenarios such as squares, railway stations, and communities. </p>
      <p>Computer vision-based identification systems complete identity verification by collecting visual features such as faces and gaits through cameras [<xref ref-type="bibr" rid="B3">3</xref>][<xref ref-type="bibr" rid="B4">4</xref>]. RGB cameras can capture facial textures, and RGB-D sensors can obtain 3D structures [<xref ref-type="bibr" rid="B1">1</xref>], with high accuracy in ideal environments. However, visual methods have weak cross-scenario generalization ability, are easily affected by illumination, angle, occlusion, and clothing changes, and their performance decreases significantly in open environments such as squares and stations [<xref ref-type="bibr" rid="B5">5</xref>], making it difficult to support cross-location person identity matching and identification. </p>
      <p>Millimeter-wave radar-based identification systems realize identity perception by analyzing radar echo signals caused by human movements, with advantages of non-cooperation, occlusion resistance, illumination independence, and strong privacy protection. Moreover, identity features such as gaits and micro-movements are more consistent across scenarios. However, a single radar modality has limited feature dimensions, making it difficult to achieve high-precision identity discrimination in long-distance and multi-target scenarios. </p>
      <p>Millimeter-wave radar and video fusion-based person identification systems fully combine the complementary advantages of millimeter-wave radar and video. Millimeter-wave radar can penetrate obstacles, is not limited by illumination, and can capture 3D information with strong cross-scenario consistency, such as distance, speed, and gait trajectories, while protecting privacy [<xref ref-type="bibr" rid="B5">5</xref>][<xref ref-type="bibr" rid="B6">6</xref>]; video can provide rich visual details such as faces, contours, and clothing [<xref ref-type="bibr" rid="B7">7</xref>]. Through dual-modal fusion, the system not only retains the high discrimination of visual features but also obtains the strong robustness of radar features, significantly improving the stability and accuracy of cross-environment, cross-location, and cross-device person identification. It is an ideal solution to realize cross-scenario retrieval and identity confirmation of personnel in open areas [<xref ref-type="bibr" rid="B8">8</xref>]. </p>
      <p>To date, remarkable progress has been made in the research on millimeter-wave radar and video fusion for person identification, but there is still a lack of comprehensive reviews specifically targeting this fusion technology. This paper systematically summarizes the latest research results on person identification methods based on millimeter-wave radar and video fusion. Firstly, it introduces the overall framework of the fusion identification system, including information feature extraction, feature fusion and alignment, and cross-modal recognition and matching algorithms. Then, it analyzes and summarizes typical application scenarios of the fusion system. Next, it discusses the current limitations and unsolved problems of fusion-based person identification and proposes potential future research directions. Finally, it concludes the full text to provide a reference for subsequent research in this field. </p>
      <p>This paper is divided into five parts: The first part describes the development and application of person identification technologies based on different systems, emphasizing the advantages of millimeter-wave radar and video fusion systems. The second part details the overall framework and key technologies of the fusion identification system, including information feature extraction, feature fusion and alignment, and cross-modal recognition and matching algorithms. The third part introduces typical application scenarios of the fusion system and analyzes their performance characteristics. The fourth part points out the limitations and future research directions of fusion-based person identification. The fifth part presents the conclusions. </p>
    </sec>
    <sec id="sec2">
      <title>2. Architecture of Millimeter-Wave Radar and Video Fusion Person Identification System</title>
      <p>The basic principle of person identification based on millimeter-wave radar and video fusion is as follows. Millimeter-wave radar and video sensors respectively, collect relevant information about target persons. The algorithms preprocess raw data to extract effective features with scene invariance, eliminate modal differences and scenario shifts through feature fusion and alignment, and finally complete identity verification through recognition and matching algorithms [<xref ref-type="bibr" rid="B1">1</xref>]. The general framework of the fusion identification system mainly includes three key stages, as shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>. </p>
      <p>Information feature extraction: In this stage, raw data of target persons are collected through millimeter-wave radar and video sensors. The radar transmits frequency-modulated continuous waves and receives echo signals to extract 3D features such as distance, speed, and angle. Video sensors capture image</p>
      <fig id="fig1">
        <label>Figure 1</label>
        <graphic xlink:href="https://html.scirp.org/file/1733518-rId13.jpeg?20260422044526" />
      </fig>
      <p><bold>Figure 1.</bold>Architecture of person identification based on millimeter-wave radar and video fusion.</p>
      <p>sequences to extract visual features such as facial textures, gait contours, and skeletal structures. </p>
      <p>Feature fusion and alignment: to remove noise and redundant information, the preprocessing of the extracted heterogeneous features is applied. This method gives the complementary advantages of the two modalities through fusion strategies such as data-level, feature-level, and decision-level. Meanwhile, alignment methods to reduce modal differences and map features to a unified feature space are implemented. Cross-modal recognition and matching: the system feeds the fused features into classification or matching algorithms to complete identity verification. A variety of machine learning and deep learning algorithms can be used, and the output result is the identity label of the target person. </p>
      <p>These three stages constitute the general framework of person identification based on millimeter-wave radar and video fusion, providing support for the development of robust and reliable identification systems.</p>
      <sec id="sec2dot1">
        <title>2.1. Feature Extraction</title>
        <p>Feature extraction is the basis of fusion identification, aiming to fully capture identity-related features from radar and video data to lay the foundation for subsequent fusion and matching, as shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>.</p>
        <fig id="fig2">
          <label>Figure 2</label>
          <graphic xlink:href="https://html.scirp.org/file/1733518-rId14.jpeg?20260422044526" />
        </fig>
        <p><bold>Figure 2.</bold>Feature extraction processes of the two modalities.</p>
        <p>2.1.1. Radar Feature Extraction</p>
        <p>Millimeter-wave radar extracts features based on echo signals, mainly including range-Doppler features, spatial structure features, and motion trajectory features [<xref ref-type="bibr" rid="B6">6</xref>]. </p>
        <p>The radar data collection usually requires the target to move naturally within the sensing range [<xref ref-type="bibr" rid="B2">2</xref>] to obtain sufficient echo information. Typical radar feature extraction applications involve various feature types, data forms, and sample sizes, as shown in <bold>Table 1</bold>. </p>
        <p><bold>Table 1.</bold> Typical applications of radar data acquisition and feature extraction. </p>
        <table-wrap id="tbl1">
          <label>Table 1</label>
          <table>
            <tbody>
              <tr>
                <td>Feature Type</td>
                <td>Data Form</td>
                <td>Sample Size</td>
                <td>Radar Parameters</td>
                <td>Reference</td>
              </tr>
              <tr>
                <td>Contour features, motion trajectory</td>
                <td>Voxelized point cloud</td>
                <td>5 action categories, 3,200 samples</td>
                <td>77 GHz FMCW radar</td>
                <td>
                  [
                  <xref ref-type="bibr" rid="B9">9</xref>
                  ][
                  <xref ref-type="bibr" rid="B10">10</xref>
                  ]
                </td>
              </tr>
              <tr>
                <td>MicroDoppler features, range-angle features</td>
                <td>Timefrequency spectrogram</td>
                <td>Multiple users, 1000+ samples</td>
                <td>24 GHz SIMO radar</td>
                <td>
                  [
                  <xref ref-type="bibr" rid="B11">11</xref>
                  ]
                </td>
              </tr>
              <tr>
                <td>Vital sign features</td>
                <td>Time series signal</td>
                <td>20 volunteers</td>
                <td>76 - 81 GHz FMCW radar</td>
                <td>
                  [
                  <xref ref-type="bibr" rid="B2">2</xref>
                  ]
                </td>
              </tr>
              <tr>
                <td>Radar point cloud features, 3D structural features</td>
                <td>Coordinate matrix</td>
                <td>58 volunteers</td>
                <td>IWR6843ISKODS radar</td>
                <td>
                  [
                  <xref ref-type="bibr" rid="B7">7</xref>
                  ]
                </td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Common radar feature extraction techniques include two-dimensional Fast Fourier Transform (2D-FFT), Multiple Signal Classification (MUSIC) algorithm, and point cloud voxelization. 2D-FFT converts time-domain signals into Range-Doppler Maps (RDM) to extract distance and speed features [<xref ref-type="bibr" rid="B9">9</xref>]. For example, Cao<italic>et al</italic>. [<xref ref-type="bibr" rid="B3">3</xref>] used 2D-FFT to process radar echo signals to obtain range-Doppler map sequences containing gait motion information; the MUSIC algorithm is used for high-precision angle estimation to generate Angle-Time Maps (ATM) to supplement spatial orientation features; point cloud voxelization converts sparse point clouds into regular voxel grids to facilitate feature extraction by deep learning models. Luo<italic>et al</italic>. [<xref ref-type="bibr" rid="B6">6</xref>] adopted a sliding window-based method to convert single-frame point clouds into multi-frame time-series data, retaining the temporal motion features of skeletal postures. </p>
        <p>2.1.2. Video Feature Extraction</p>
        <p>Video feature extraction focuses on identity-related visual information, such as facial features, gait features, and skeletal features [<xref ref-type="bibr" rid="B7">7</xref>]. Facial features include texture, contour, and key point information, extracted through face detection and feature encoding methods [<xref ref-type="bibr" rid="B12">12</xref>]; gait features originate from motion patterns such as walking posture and step frequency [<xref ref-type="bibr" rid="B10">10</xref>], usually obtained by analyzing image sequences or contour sequences [<xref ref-type="bibr" rid="B3">3</xref>][<xref ref-type="bibr" rid="B4">4</xref>]; skeletal features extract 3D coordinates of human joints through pose estimation algorithms [<xref ref-type="bibr" rid="B6">6</xref>][<xref ref-type="bibr" rid="B13">13</xref>]. </p>
        <p>Typical video feature extraction applications involve various feature types and processing methods, as shown in <bold>Table 2</bold>. </p>
        <p>Common data forms include Range-Doppler Spectrograms (RDM), 4D point clouds (x, y, z, v), time-frequency spectrograms, etc. Extraction technologies mainly include the following three categories: </p>
        <p><bold>Table 2.</bold>Typical applications of video data acquisition and feature extraction.</p>
        <table-wrap id="tbl2">
          <label>Table 2</label>
          <table>
            <tbody>
              <tr>
                <td>Feature Type</td>
                <td>Processing Method</td>
                <td>Data Source</td>
                <td>Sample Size</td>
                <td>Reference</td>
              </tr>
              <tr>
                <td>RGB texture features, facial keypoint features</td>
                <td>CNN-based feature extraction</td>
                <td>RGB camera</td>
                <td>58 volunteers × multiple images</td>
                <td>
                  [
                  <xref ref-type="bibr" rid="B7">7</xref>
                  ]
                </td>
              </tr>
              <tr>
                <td>Gait Energy Image (GEI), motion contour features</td>
                <td>Spatial attention module</td>
                <td>RGB camera</td>
                <td>121 subjects × 8 views</td>
                <td>
                  [
                  <xref ref-type="bibr" rid="B4">4</xref>
                  ]
                </td>
              </tr>
              <tr>
                <td>Motion sequence features, visual semantic features</td>
                <td>Transformer-based temporal modeling</td>
                <td>RGB camera</td>
                <td>Multiple sequences × 30+ frames</td>
                <td>
                  [
                  <xref ref-type="bibr" rid="B8">8</xref>
                  ]
                </td>
              </tr>
              <tr>
                <td>Facial identity features, gait motion features</td>
                <td>Knowledge distillation encoding</td>
                <td>RGB camera</td>
                <td>36 subjects × 5000 + face-gait pairs</td>
                <td>
                  [
                  <xref ref-type="bibr" rid="B12">12</xref>
                  ]
                </td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Signal transformation technology converts time-domain signals into range-Doppler spectrograms through 2D-FFT to capture target motion features [<xref ref-type="bibr" rid="B14">14</xref>] and generates time-frequency spectrograms using STFT to extract micro-Doppler features [<xref ref-type="bibr" rid="B15">15</xref>][<xref ref-type="bibr" rid="B16">16</xref>]. Spatial feature extraction processes sparse point clouds using networks such as PointNet and DGCNN to capture 3D human structure and contour topology [<xref ref-type="bibr" rid="B7">7</xref>]. It also achieves high-precision angle estimation through the MUSIC algorithm to generate ATM [<xref ref-type="bibr" rid="B17">17</xref>]. Temporal feature mining converts single-frame point clouds into time series using a sliding window mechanism to capture dynamic patterns of gaits and actions [<xref ref-type="bibr" rid="B18">18</xref>] and optimizes trajectory features through Kalman filtering to improve stability [<xref ref-type="bibr" rid="B19">19</xref>]. </p>
      </sec>
      <sec id="sec2dot2">
        <title>2.2. Feature Fusion and Alignment</title>
        <p>Feature fusion and alignment are key elements for solving the heterogeneity of millimeter-wave radar and video data, directly affecting the performance of the fusion identification system [<xref ref-type="bibr" rid="B20">20</xref>]. The specific implementation is shown in <xref ref-type="fig" rid="fig3">Figure 3</xref>. </p>
        <fig id="fig3">
          <label>Figure 3</label>
          <graphic xlink:href="https://html.scirp.org/file/1733518-rId15.jpeg?20260422044527" />
        </fig>
        <p><bold>Figure 3.</bold>Implementation flow of spatio-temporal alignment, feature distribution alignment, and fusion strategies.</p>
        <p>2.2.1. Fusion Strategies</p>
        <p>According to different fusion stages, fusion strategies can be divided into data-level fusion, feature-level fusion, and decision-level fusion [<xref ref-type="bibr" rid="B5">5</xref>]. </p>
        <p>Data-level fusion strategy directly integrates raw radar and video data before feature extraction. For example, Zong <italic>et al</italic>. [<xref ref-type="bibr" rid="B19">19</xref>] spliced radar range-angle information with video pixel data to form a multimodal data matrix, which is input into the fusion detection network. This strategy retains the most original information but has high requirements for data synchronization and noise reduction. </p>
        <p>The feature-level fusion strategy integrates extracted radar and video features in a unified feature space, which is the most widely used fusion strategy in current research. For example, Liu <italic>et al</italic>. [<xref ref-type="bibr" rid="B7">7</xref>] designed a cross-modal similarity estimation method to fuse radar point cloud features and video 2D image features. Shi <italic>et al</italic>. [<xref ref-type="bibr" rid="B4">4</xref>] fused radar micro-motion features and video gait appearance features in a multi-scale feature space. Chen <italic>et al</italic>. [<xref ref-type="bibr" rid="B8">8</xref>] adopted a dual contrast learning framework to fuse radar Doppler features and video motion features. Feature-level fusion balances information retention and computational efficiency, and can adaptively adjust feature weights according to scenario changes. </p>
        <p>Decision-level fusion strategy combines the recognition results of radar and video single-modal systems to obtain the final identity judgment [<xref ref-type="bibr" rid="B15">15</xref>]. It has strong robustness to single-modal failure but relies on the accuracy of single-modal decision results. </p>
        <p>The three fusion strategies present significant trade-offs between computational cost and system latency. Data-level fusion entails processing raw heterogeneous data, yielding high computational complexity and maximum system latency. Feature-level fusion only integrates the extracted feature vectors, with moderate computational cost and latency, balancing information integrity and real-time performance. Decision-level fusion merely conducts posterior fusion on the recognition outcomes of single-modal systems and delivers minimal computational overhead and the lowest system latency, thus being better suited for application scenarios with stringent real-time requirements. </p>
        <p>2.2.2. Alignment Methods</p>
        <p>Alignment methods aim to eliminate spatio-temporal asynchrony and feature distribution differences between radar and video modalities [<xref ref-type="bibr" rid="B21">21</xref>], including spatio-temporal alignment and feature distribution alignment. </p>
        <p>Spatio-temporal alignment ensures that the features of the two modalities correspond to the same target at the same time [<xref ref-type="bibr" rid="B6">6</xref>]. Time alignment is usually realized through sensor timestamp synchronization [<xref ref-type="bibr" rid="B7">7</xref>]. Spatial alignment uses calibration methods to map radar spatial coordinates (distance, angle) to video image coordinates to achieve spatial registration of targets. For example, Yang <italic>et al</italic>. proposed the BiCro method [<xref ref-type="bibr" rid="B22">22</xref>], which corrects noisy correspondences through bidirectional cross-modal similarity consistency to improve alignment accuracy. </p>
        <p>Feature distribution alignment reduces the distribution gap between radar and video features. Metric learning is a common method. Wang <italic>et al</italic>. [<xref ref-type="bibr" rid="B20">20</xref>] learned a coupled feature space to map heterogeneous features to a unified space. Wei<italic>et al</italic>. [<xref ref-type="bibr" rid="B21">21</xref>] proposed a universal weighted metric learning method to adaptively adjust feature weights. Liong<italic>et al</italic>. [<xref ref-type="bibr" rid="B23">23</xref>] adopted deep coupled metric learning to realize nonlinear mapping of features. In addition, contrastive learning and knowledge distillation are also widely used in alignment tasks [<xref ref-type="bibr" rid="B12">12</xref>]. Chen<italic>et al</italic>. [<xref ref-type="bibr" rid="B8">8</xref>] transferred video facial identity knowledge to radar gait features through knowledge distillation to achieve semantic alignment. </p>
      </sec>
      <sec id="sec2dot3">
        <title>2.3. Cross-Modal Recognition and Matching Algorithms</title>
        <p>Cross-modal matching algorithms [<xref ref-type="bibr" rid="B24">24</xref>] and recognition [<xref ref-type="bibr" rid="B25">25</xref>] complete identity verification based on fused and aligned features, which can be divided into machine learning algorithms and deep learning algorithms. Traditional machine learning relies on manually designed features, with lightweight models, fast training speed, good compatibility with small sample data, and strong interpretability. However, it has limited feature expression ability in complex scenarios and vulnerable generalization performance, and is suitable for classification and regression tasks with small data volume and clear features. Deep learning can automatically learn deep abstract features from massive data, with higher accuracy in complex data processing such as images and radar signals, but has large model parameters, a long training time, requires a large amount of labeled data, poor interpretability, and is prone to overfitting. Overall, machine learning is more advantageous in small-sample and low-dimensional scenarios; deep learning has significantly better feature extraction and recognition performance in high-dimensional and nonlinear data tasks such as millimeter-wave radar and video fusion. </p>
        <p>2.3.1. Machine Learning Algorithms</p>
        <p>Machine learning algorithms mainly use traditional classification and matching methods to process fused features [<xref ref-type="bibr" rid="B20">20</xref>]. Common algorithms include Dynamic Time Warping (DTW), K-Nearest Neighbor (KNN), Hidden Markov Model (HMM), and Random Forests. </p>
        <p>DTW is used for matching time-series features such as gaits and motion trajectories, calculating the similarity between test sequences and reference templates through dynamic programming, suitable for radar motion feature matching. KNN performs identity classification based on feature similarity [<xref ref-type="bibr" rid="B9">9</xref>]. For example, Guo <italic>et al</italic>. [<xref ref-type="bibr" rid="B2">2</xref>] used KNN to classify fused vital sign features with an accuracy of over 85%. HMM models the temporal dependence of features, suitable for dynamic gesture and gait-based recognition [<xref ref-type="bibr" rid="B10">10</xref>]. Random Forest performs classification by integrating multiple decision trees and can handle high-dimensional fused features [<xref ref-type="bibr" rid="B11">11</xref>]. However, machine learning algorithms have a limited ability to process complex heterogeneous features and perform worse than deep learning algorithms on large-scale datasets. </p>
        <p>2.3.2. Deep Learning Algorithms</p>
        <p>Deep learning algorithms have powerful feature learning and pattern recognition capabilities, making them the mainstream algorithms in current fusion identification [<xref ref-type="bibr" rid="B7">7</xref>][<xref ref-type="bibr" rid="B12">12</xref>]. Common algorithms include Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM), and Transformer. </p>
        <p>CNN is used to extract spatial features from voxelized point clouds and multimodal feature maps [<xref ref-type="bibr" rid="B14">14</xref>]. For example, Shi<italic>et al</italic>. [<xref ref-type="bibr" rid="B4">4</xref>] used a deep CNN to fuse radar time-Doppler spectrograms and video gait energy images. Luo<italic>et al</italic>. [<xref ref-type="bibr" rid="B6">6</xref>] used CNN to extract spatial features of radar point clouds. LSTM is suitable for processing time-series features such as gait sequences and radar signal sequences. Shan<italic>et al</italic>. [<xref ref-type="bibr" rid="B12">12</xref>] used LSTM to model the temporal dependence of gait features to improve recognition robustness. Transformer captures global feature dependencies through a self-attention mechanism, suitable for cross-modal feature matching [<xref ref-type="bibr" rid="B18">18</xref>]. Chen <italic>et al</italic>. [<xref ref-type="bibr" rid="B8">8</xref>] used Transformer-based networks for temporal modeling and instance discrimination to enhance the discriminative ability of fused features. </p>
        <p>Some innovative algorithms were proposed for these applications. Yao<italic>et al</italic>. [<xref ref-type="bibr" rid="B25">25</xref>] proposed the semi-supervised recognition system mmSignature, achieving 96.3% accuracy with a small amount of labeled data. Shao <italic>et al</italic>. [<xref ref-type="bibr" rid="B18">18</xref>] realized fast action recognition based on point cloud sequence learning with inference time less than 0.2 ms. Deep learning algorithms can also be combined with multi-task learning, transfer learning, and other technologies to improve performance. For example, Shan <italic>et al</italic>. [<xref ref-type="bibr" rid="B12">12</xref>] used knowledge distillation for transfer learning, and Deng <italic>et al</italic>. [<xref ref-type="bibr" rid="B26">26</xref>] used synthetic data to enhance model generalization ability. </p>
        <p>In summary, deep learning delivers superior performance in high-dimensional, heterogeneous cross-modal feature learning and complex-scenario recognition. By contrast, traditional machine learning models still present more prominent practical advantages for resource-constrained edge computing devices, few-shot annotation scenarios, and low-latency lightweight deployment, owing to their lightweight architecture, low computational cost, and fast inference speed. </p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. Typical Applications</title>
      <p>This section introduces several typical application systems of person identification based on millimeter-wave radar and video fusion, and analyzes their performance characteristics from multiple perspectives. These studies are divided into the following categories according to application scenarios: intelligent security monitoring, human-computer interaction identity authentication, and multi-target tracking identification. </p>
      <sec id="sec3dot1">
        <title>3.1. Intelligent Security Monitoring</title>
        <p>This application aims to achieve real-time person identification and pedestrian re-identification in public places such as shopping malls, stations, and communities [<xref ref-type="bibr" rid="B27">27</xref>], ensure public safety, and provide cross-environment identity confirmation and trajectory tracking of personnel [<xref ref-type="bibr" rid="B6">6</xref>]. The core requirements are high precision and strong robustness to complex environments. </p>
        <p>Cao<italic>et al</italic>. [<xref ref-type="bibr" rid="B3">3</xref>] proposed a cross vision-RF gait re-identification system based on low-cost RGB-D cameras and millimeter-wave radar, fusing gait motion features extracted by radar and visual gait features obtained by cameras, achieving about 92.5% top-1 accuracy and 97.5% top-5 accuracy among 56 volunteers, and stably identifying targets even in multi-person scenarios. Liu <italic>et al</italic>. [<xref ref-type="bibr" rid="B7">7</xref>] designed the Mission system, which detects targets through radar in camera-restricted areas and identifies person images from the camera network, achieving 85% top-1 accuracy and 90% top-5 accuracy among 58 volunteers, suitable for security monitoring scenarios. </p>
        <p>The core challenge of intelligent security monitoring is to cope with environmental interference such as crowd density, occlusion, and illumination changes. Future research needs to focus on improving the real-time performance and multi-target discrimination ability of the system. </p>
      </sec>
      <sec id="sec3dot2">
        <title>3.2. Human-Computer Interaction Identity Authentication</title>
        <p>This application combines person identification with human-computer interaction, realizing identity authentication while completing gesture and action control, and is applied to smart homes, smart vehicles, and other scenarios. The core requirements are fast response and non-contact operation. Janakaraj<italic>et al</italic>. [<xref ref-type="bibr" rid="B28">28</xref>] proposed the STAR system, realizing simultaneous tracking and recognition of persons using millimeter-wave radar and deep learning, fusing radar tracking information and video identity features, achieving real-time recognition in human-computer interaction scenarios with high stability. Huang<italic>et al</italic>. [<xref ref-type="bibr" rid="B11">11</xref>] proposed the RPCRS system that extracted human activity characteristics from millimeter-wave radar signals and combined them with video identity verification to achieve safe human-computer interaction, with recognition accuracy over 90%. The MatchAnything method, proposed by He <italic>et al</italic>. [<xref ref-type="bibr" rid="B29">29</xref>], is capable of matching diverse image data, making it applicable for cross-modal behavior recognition, where it can also be utilized to process radar signals to enhance recognition precision. Yao<italic>et al</italic>.’s mmSignature system [<xref ref-type="bibr" rid="B30">30</xref>] realizes identity verification for interactive scenarios such as device unlocking and personalized services through semi-supervised learning, with an accuracy of 96.3%. </p>
        <p>The core challenge of human-computer interaction identity authentication is to balance real-time performance and accuracy. Future research needs to design lightweight fusion models to reduce computational delays. </p>
      </sec>
      <sec id="sec3dot3">
        <title>3.3. Multi-Target Tracking Identification</title>
        <p>This application aims to continuously maintain identity and stably identify multiple moving personnel in complex and dense scenarios such as stations, squares, and parks. The core requirement is to distinguish different personnel identities and maintain recognition consistency even under multi-person interference, occlusion, and staggered movement. Luo <italic>et al</italic>. [<xref ref-type="bibr" rid="B6">6</xref>] adopted the CNN-BiGRU architecture to fuse radar point clouds and video joint information, realizing multi-target skeletal positioning and identity discrimination, with a multi-target identity identification accuracy of 89%. Cao<italic>et al</italic>. [<xref ref-type="bibr" rid="B3">3</xref>] fused radar gait features and video visual features, and can stably match target identities even in multi-person parallel movement scenarios, effectively avoiding target confusion and identity jumps. Shao <italic>et al</italic>. [<xref ref-type="bibr" rid="B18">18</xref>] realized real-time multi-target identity tracking and discrimination based on rapid modeling of radar point cloud sequences, maintaining high frame rate and identity consistency in complex crowd scenarios. </p>
        <p>The core challenge of multi-target tracking identity identification is feature confusion and identity association errors caused by occlusion interference and target staggering. Future optimization of the multi-target association mechanism and cross-modal feature matching strategy is needed to improve the stability of identity identification in dense scenarios. </p>
      </sec>
      <sec id="sec3dot4">
        <title>3.4. Summary</title>
        <p><bold>Table 3</bold> briefly summarizes several mainstream applications, listing the goals and key elements of four typical application categories.</p>
        <p><bold>Table 3.</bold> Key features of four typical application categories.</p>
        <table-wrap id="tbl3">
          <label>Table 3</label>
          <table>
            <tbody>
              <tr>
                <td>Application Type</td>
                <td>Goal</td>
                <td>Highlights</td>
                <td>Reference</td>
              </tr>
              <tr>
                <td>Intelligent security monitoring</td>
                <td>Real-time identification in public places</td>
                <td>High precision, anti-interference</td>
                <td>
                  [
                  <xref ref-type="bibr" rid="B1">1</xref>
                  ][
                  <xref ref-type="bibr" rid="B3">3</xref>
                  ][
                  <xref ref-type="bibr" rid="B7">7</xref>
                  ]
                </td>
              </tr>
              <tr>
                <td>Human-computer interaction authentication</td>
                <td>Authentication during interaction</td>
                <td>Fast response, non-contact</td>
                <td>
                  [
                  <xref ref-type="bibr" rid="B11">11</xref>
                  ][
                  <xref ref-type="bibr" rid="B28">28</xref>
                  ]
                </td>
              </tr>
              <tr>
                <td>Multi-target tracking identification</td>
                <td>Simultaneous identification of multiple targets</td>
                <td>Occlusion resistance, consistent tracking</td>
                <td>
                  [
                  <xref ref-type="bibr" rid="B6">6</xref>
                  ][
                  <xref ref-type="bibr" rid="B18">18</xref>
                  ]
                </td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
      </sec>
    </sec>
    <sec id="sec4">
      <title>4. Current Challenges</title>
      <sec id="sec4dot1">
        <title>4.1. Complex Environmental Interference</title>
        <p>Most current research experiments are conducted in laboratory environments with less interference [<xref ref-type="bibr" rid="B2">2</xref>][<xref ref-type="bibr" rid="B7">7</xref>]. However, in practical applications, the system is often affected by factors such as crowd density, occlusion, illumination changes, and electromagnetic interference [<xref ref-type="bibr" rid="B1">1</xref>]. For example, in dense crowd scenarios, radar signals are easily mixed with echoes from multiple targets, and video images are severely occluded, leading to a significant decrease in recognition accuracy [<xref ref-type="bibr" rid="B6">6</xref>]. In addition, severe weather such as rain and fog will attenuate radar signals and blur video images, further affecting system performance [<xref ref-type="bibr" rid="B5">5</xref>]. </p>
      </sec>
      <sec id="sec4dot2">
        <title>4.2. Limited Modal Alignment Accuracy</title>
        <p>Although existing alignment methods have made progress, there are still gaps in spatio-temporal synchronization and feature distribution consistency [<xref ref-type="bibr" rid="B20">20</xref>][<xref ref-type="bibr" rid="B23">23</xref>]. Spatio-temporal asynchrony between radar and video sensors may lead to feature misalignment of the same target [<xref ref-type="bibr" rid="B6">6</xref>]. Feature distribution differences caused by modal heterogeneity [<xref ref-type="bibr" rid="B29">29</xref>] make it difficult to fully integrate complementary information, limiting the improvement of recognition accuracy [<xref ref-type="bibr" rid="B12">12</xref>]. Especially in dynamic scenarios, target movement will exacerbate alignment difficulties [<xref ref-type="bibr" rid="B8">8</xref>]. </p>
      </sec>
      <sec id="sec4dot3">
        <title>4.3. Scarcity of Large-Scale Data</title>
        <p>The collection and annotation of millimeter-wave radar and video fusion data are time-consuming and labor-intensive. Cross-scenario data collection and annotation are costly and privacy-sensitive, making it difficult to construct large-scale standardized datasets [<xref ref-type="bibr" rid="B26">26</xref>]. Currently, publicly available millimeter-wave radar and video fusion datasets for person identification are extremely scarce. Two representative datasets are as follows. The mmBody dataset provides synchronized millimeter-wave radar point clouds and RGBD video data, covering 20 subjects, 100 motions, and 7 scenes, serving as a commonly used multimodal benchmark for human perception [<xref ref-type="bibr" rid="B31">31</xref>]. Such public datasets generally suffer from limitations, including a small number of subjects, single-scene settings, and insufficient sample sizes, which still cannot meet the requirements of large-scale cross-scenario model training. Radar data require professional equipment and technicians for collection, and video data involve privacy issues, making it difficult to build large-scale annotated datasets [<xref ref-type="bibr" rid="B16">16</xref>]. Although some studies use synthetic data to supplement training samples [<xref ref-type="bibr" rid="B26">26</xref>], there is a distribution gap between synthetic data and real data, affecting model generalization ability [<xref ref-type="bibr" rid="B30">30</xref>]. </p>
      </sec>
      <sec id="sec4dot4">
        <title>4.4. Privacy and Ethical Risks</title>
        <p>Although millimeter-wave radar data is privacy-preserving and avoids direct collection of sensitive human visual information, fusion of radar data and video data forms a multimodal data linkage that can precisely associate individual identities, thereby introducing new privacy leakage and ethical risks. Cross-modal data association and matching may break through the privacy protection boundaries of a single modality. Without sufficient authorization, this can easily lead to the abuse of identity information and continuous tracking, violating the principles of data minimization and privacy compliance. This represents a critical ethical and compliance challenge that must be addressed in the practical deployment of fusion technologies. </p>
      </sec>
    </sec>
    <sec id="sec5">
      <title>5. Future Research Directions</title>
      <sec id="sec5dot1">
        <title>5.1. Enhance Environmental Adaptability and Learn Scene-Invariant Features</title>
        <p>To improve anti-interference ability, it is important to optimize feature extraction algorithms [<xref ref-type="bibr" rid="B3">3</xref>]. For example, we may use attention mechanisms to focus on extracting “illumination-independent, angle-independent, occlusion-robust” identity features, and use attention mechanisms, domain adaptation, and data enhancement to improve cross-environment generalization ability and suppress environmental noise [<xref ref-type="bibr" rid="B4">4</xref>][<xref ref-type="bibr" rid="B8">8</xref>]. We also design a multi-view sensor layout to reduce occlusion impact [<xref ref-type="bibr" rid="B32">32</xref>]. In addition, we may enhance model robustness through domain adaptation and data enhancement technologies [<xref ref-type="bibr" rid="B26">26</xref>]. </p>
      </sec>
      <sec id="sec5dot2">
        <title>5.2. Improve Modal Alignment Accuracy</title>
        <p>We may develop more accurate spatio-temporal calibration methods to achieve millisecond-level time synchronization and pixel-level spatial registration [<xref ref-type="bibr" rid="B6">6</xref>]. At the same time, we can explore advanced feature alignment technologies such as cross-modal contrastive learning and generative adversarial networks to reduce feature distribution differences [<xref ref-type="bibr" rid="B12">12</xref>]. For dynamic scenarios, it is important to design adaptive alignment strategies that can be adjusted in real time according to target movement [<xref ref-type="bibr" rid="B8">8</xref>]. </p>
      </sec>
      <sec id="sec5dot3">
        <title>5.3. Expand Data Resources and Cross-Scenario Generalization</title>
        <p>High-quality synthetic data may be generated to supplement real data [<xref ref-type="bibr" rid="B26">26</xref>] to reduce dependence on large-scale annotated data and achieve rapid adaptation to new locations and environments with a small number of samples. For example, [<xref ref-type="bibr" rid="B33">33</xref>] optimizes the video-to-radar conversion model to improve the authenticity of synthetic radar data [<xref ref-type="bibr" rid="B16">16</xref>]. We may use federated learning to achieve data sharing under privacy protection, expand the training data scale, and explore semi-supervised and unsupervised learning algorithms to reduce dependence on annotated data [<xref ref-type="bibr" rid="B29">29</xref>]. </p>
      </sec>
      <sec id="sec5dot4">
        <title>5.4. Unified Feature Space Construction</title>
        <p>Unified feature space construction is a crucial research direction, as it can map data from multiple modalities into a common feature space. As a result, it reduces the discrepancies among different modalities, enables more diverse applications, and allows the use of more network models. Unified feature space construction can rely on modern architectures, including cross-modal Transformers, dual-branch contrastive learning frameworks, and deep cross-modal metric learning networks. It realizes global cross-modal feature interaction via self-attention mechanisms and narrows the feature distance of homologous targets through contrastive learning, ultimately achieving precise alignment and efficient representation of heterogeneous radar and video features in a unified latent space. This serves as the core technical pathway to address modal heterogeneity. </p>
      </sec>
    </sec>
    <sec id="sec6">
      <title>6. Conclusions</title>
      <p>This paper comprehensively reviews the research status of person identification methods based on millimeter-wave radar and video fusion. Firstly, it reviews common person identification systems, dividing them into wearable sensor-based, computer vision-based, millimeter-wave radar-based, and millimeter-wave radar and video fusion-based systems, highlighting the advantages of dual-modal fusion. Then, it introduces the overall framework of the fusion identification system, including information feature extraction, feature fusion and alignment, and cross-modal recognition and matching algorithms, and elaborates on key technologies and implementation methods in detail. Next, it reviews the latest research on fusion-based person identification, dividing it into three typical application scenarios: intelligent security monitoring, human-computer interaction identity authentication, and multi-target tracking identification, and analyzes the implementation process and performance characteristics of each system. Finally, it discusses the limitations and existing problems of current research and proposes future research directions combined with research trends. </p>
      <p>Person identification based on millimeter-wave radar and video fusion combines the advantages of the two modalities, overcomes the limitations of single-modal identification, and has broad application prospects in intelligent security, human-computer interaction, and other fields. With the continuous advancement of sensor technology and artificial intelligence algorithms, future fusion identification systems will achieve higher accuracy, greater robustness, and better real-time performance, providing more reliable technical support for various intelligent applications. This paper can provide a reference for researchers’ follow-up research and help develop more efficient and practical person identification systems.</p>
    </sec>
    <sec id="sec7">
      <title>Funding</title>
      <p>The work is funded by the Foundation of the Innovation and Entrepreneurship Training Program for College Students (X202510424033). </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Yang, Q., Quan, Z., Li, J., Jiang, T., Deng, Z., Qiu, X., <italic>et al</italic>. (2025) A Survey of Pedestrian Re-Identification Based on Millimeter Wave Radar and Vision Fusion. <italic>Journal of Computer and Communications</italic>, 13, 64-80. https://doi.org/10.4236/jcc.2025.136005 <pub-id pub-id-type="doi">10.4236/jcc.2025.136005</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.4236/jcc.2025.136005">https://doi.org/10.4236/jcc.2025.136005</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Yang, Q.</string-name>
              <string-name>Quan, Z.</string-name>
              <string-name>Li, J.</string-name>
              <string-name>Jiang, T.</string-name>
              <string-name>Deng, Z.</string-name>
              <string-name>Qiu, X.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>A Survey of Pedestrian Re-Identification Based on Millimeter Wave Radar and Vision Fusion</article-title>
            <source>Journal of Computer and Communications</source>
            <volume>13</volume>
            <pub-id pub-id-type="doi">10.4236/jcc.2025.136005</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Guo, J., Wei, J., Xiang, Y. and Han, C. (2024) Millimeter-Wave Radar-Based Identity Recognition Algorithm Built on Multimodal Fusion. <italic>Sensors</italic>, 24, Article 4051. https://doi.org/10.3390/s24134051 <pub-id pub-id-type="doi">10.3390/s24134051</pub-id><pub-id pub-id-type="pmid">39000830</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/s24134051">https://doi.org/10.3390/s24134051</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Guo, J.</string-name>
              <string-name>Wei, J.</string-name>
              <string-name>Xiang, Y.</string-name>
              <string-name>Han, C.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Millimeter-Wave Radar-Based Identity Recognition Algorithm Built on Multimodal Fusion</article-title>
            <source>Sensors</source>
            <volume>24</volume>
            <elocation-id>4051</elocation-id>
            <pub-id pub-id-type="doi">10.3390/s24134051</pub-id>
            <pub-id pub-id-type="pmid">39000830</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Cao, D., Liu, R., Li, H., Wang, S., Jiang, W. and Lu, C.X. (2022) Cross Vision-RF Gait Re-Identification with Low-Cost RGB-D Cameras and Mmwave Radars. <italic>Proceedings of the ACM on Interactive</italic>, <italic>Mobile</italic>, <italic>Wearable and Ubiquitous Technologies</italic>, 6, 1-25. https://doi.org/10.1145/3550325 <pub-id pub-id-type="doi">10.1145/3550325</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3550325">https://doi.org/10.1145/3550325</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Cao, D.</string-name>
              <string-name>Liu, R.</string-name>
              <string-name>Li, H.</string-name>
              <string-name>Wang, S.</string-name>
              <string-name>Jiang, W.</string-name>
              <string-name>Lu, C.X.</string-name>
              <string-name>Interactive, M</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Cross Vision-RF Gait Re-Identification with Low-Cost RGB-D Cameras and Mmwave Radars</article-title>
            <source>Proceedings of the ACM on Interactive</source>
            <volume>6</volume>
            <pub-id pub-id-type="doi">10.1145/3550325</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Shi, Y., Du, L., Chen, X., Liao, X., Yu, Z., Li, Z., <italic>et al</italic>. (2023) Robust Gait Recognition Based on Deep CNNs with Camera and Radar Sensor Fusion. <italic>IEEE Internet of Things Journal</italic>, 10, 10817-10832. https://doi.org/10.1109/jiot.2023.3242417 <pub-id pub-id-type="doi">10.1109/jiot.2023.3242417</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/jiot.2023.3242417">https://doi.org/10.1109/jiot.2023.3242417</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Shi, Y.</string-name>
              <string-name>Du, L.</string-name>
              <string-name>Chen, X.</string-name>
              <string-name>Liao, X.</string-name>
              <string-name>Yu, Z.</string-name>
              <string-name>Li, Z.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Robust Gait Recognition Based on Deep CNNs with Camera and Radar Sensor Fusion</article-title>
            <source>IEEE Internet of Things Journal</source>
            <volume>10</volume>
            <pub-id pub-id-type="doi">10.1109/jiot.2023.3242417</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Wang, S., Mei, L., Liu, R., Jiang, W., Yin, Z., Deng, X., <italic>et al</italic>. (2025) Multi-Modal Fusion Sensing: A Comprehensive Review of Millimeter-Wave Radar and Its Integration with Other Modalities. <italic>IEEE Communications Surveys &amp; Tutorials</italic>, 27, 322-352. https://doi.org/10.1109/comst.2024.3398004 <pub-id pub-id-type="doi">10.1109/comst.2024.3398004</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/comst.2024.3398004">https://doi.org/10.1109/comst.2024.3398004</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Wang, S.</string-name>
              <string-name>Mei, L.</string-name>
              <string-name>Liu, R.</string-name>
              <string-name>Jiang, W.</string-name>
              <string-name>Yin, Z.</string-name>
              <string-name>Deng, X.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>Multi-Modal Fusion Sensing: A Comprehensive Review of Millimeter-Wave Radar and Its Integration with Other Modalities</article-title>
            <source>IEEE Communications Surveys &amp; Tutorials</source>
            <volume>27</volume>
            <pub-id pub-id-type="doi">10.1109/comst.2024.3398004</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B6">
        <label>6.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Luo, Y., He, Y., Li, Y., Liu, H., Wang, J. and Gao, F. (2025) A Sliding Window-Based CNN-BiGRU Approach for Human Skeletal Pose Estimation Using Mmwave Radar. <italic>Sensors</italic>, 25, Article 1070. https://doi.org/10.3390/s25041070 <pub-id pub-id-type="doi">10.3390/s25041070</pub-id><pub-id pub-id-type="pmid">40006298</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/s25041070">https://doi.org/10.3390/s25041070</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Luo, Y.</string-name>
              <string-name>He, Y.</string-name>
              <string-name>Li, Y.</string-name>
              <string-name>Liu, H.</string-name>
              <string-name>Wang, J.</string-name>
              <string-name>Gao, F.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>A Sliding Window-Based CNN-BiGRU Approach for Human Skeletal Pose Estimation Using Mmwave Radar</article-title>
            <source>Sensors</source>
            <volume>25</volume>
            <elocation-id>1070</elocation-id>
            <pub-id pub-id-type="doi">10.3390/s25041070</pub-id>
            <pub-id pub-id-type="pmid">40006298</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B7">
        <label>7.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Liu, R., Yao, T., Shi, R., Mei, L., Wang, S., Yin, Z., <italic>et al</italic>. (2024) Mission: mmWave Radar Person Identification with RGB Cameras. <italic>Proceedings of the</italic>22 <italic>nd ACM Conference on Embedded Networked Sensor Systems</italic>, Hangzhou, 4-7 November 2024, 309-321. https://doi.org/10.1145/3666025.3699340 <pub-id pub-id-type="doi">10.1145/3666025.3699340</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3666025.3699340">https://doi.org/10.1145/3666025.3699340</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Liu, R.</string-name>
              <string-name>Yao, T.</string-name>
              <string-name>Shi, R.</string-name>
              <string-name>Mei, L.</string-name>
              <string-name>Wang, S.</string-name>
              <string-name>Yin, Z.</string-name>
              <string-name>Systems, H</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Mission: mmWave Radar Person Identification with RGB Cameras</article-title>
            <source>Proceedings of the 22nd ACM Conference on Embedded Networked Sensor Systems</source>
            <volume>4</volume>
            <pub-id pub-id-type="doi">10.1145/3666025.3699340</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B8">
        <label>8.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Chen, Y. and Cheng, K. (2024) BiCLR: Radar-Camera-Based Cross-Modal Bi-Contrastive Learning for Human Motion Recognition. <italic>IEEE Sensors Journal</italic>, 24, 4102-4119. https://doi.org/10.1109/jsen.2023.3344789 <pub-id pub-id-type="doi">10.1109/jsen.2023.3344789</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/jsen.2023.3344789">https://doi.org/10.1109/jsen.2023.3344789</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Chen, Y.</string-name>
              <string-name>Cheng, K.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>BiCLR: Radar-Camera-Based Cross-Modal Bi-Contrastive Learning for Human Motion Recognition</article-title>
            <source>IEEE Sensors Journal</source>
            <volume>24</volume>
            <pub-id pub-id-type="doi">10.1109/jsen.2023.3344789</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B9">
        <label>9.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Singh, A.D., Sandha, S.S., Garcia, L. and Srivastava, M. (2019) RadHAR: Human Activity Recognition from Point Clouds Generated through a Millimeter-Wave Radar. <italic>Proceedings of the</italic>3 <italic>rd ACM Workshop on Millimeter</italic>- <italic>Wave Networks and Sensing Systems</italic>, Los Cabos, 25 October 2019, 51-56. https://doi.org/10.1145/3349624.3356768 <pub-id pub-id-type="doi">10.1145/3349624.3356768</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3349624.3356768">https://doi.org/10.1145/3349624.3356768</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Singh, A.D.</string-name>
              <string-name>Sandha, S.S.</string-name>
              <string-name>Garcia, L.</string-name>
              <string-name>Srivastava, M.</string-name>
              <string-name>Systems, L</string-name>
            </person-group>
            <year>2019</year>
            <article-title>RadHAR: Human Activity Recognition from Point Clouds Generated through a Millimeter-Wave Radar</article-title>
            <source>Proceedings of the 3rd ACM Workshop on Millimeter-Wave Networks and Sensing Systems</source>
            <volume>25</volume>
            <pub-id pub-id-type="doi">10.1145/3349624.3356768</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B10">
        <label>10.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Li, Z., Le Kernec, J., Abbasi, Q., Fioranelli, F., Yang, S. and Romain, O. (2023) Radar-based Human Activity Recognition with Adaptive Thresholding towards Resource Constrained Platforms. <italic>Scientific Reports</italic>, 13, Article No. 3473. https://doi.org/10.1038/s41598-023-30631-x <pub-id pub-id-type="doi">10.1038/s41598-023-30631-x</pub-id><pub-id pub-id-type="pmid">36859571</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41598-023-30631-x">https://doi.org/10.1038/s41598-023-30631-x</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Li, Z.</string-name>
              <string-name>Kernec, J.</string-name>
              <string-name>Abbasi, Q.</string-name>
              <string-name>Fioranelli, F.</string-name>
              <string-name>Yang, S.</string-name>
              <string-name>Romain, O.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Radar-based Human Activity Recognition with Adaptive Thresholding towards Resource Constrained Platforms</article-title>
            <source>Scientific Reports</source>
            <volume>13</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1038/s41598-023-30631-x</pub-id>
            <pub-id pub-id-type="pmid">36859571</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B11">
        <label>11.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Huang, T., Liu, G., Li, S. and Liu, J. (2023) RPCRS: Human Activity Recognition Using Millimeter Wave Radar. 2022 <italic>IEEE</italic>28 <italic>th International Conference on Parallel and Distributed Systems</italic>( <italic>ICPADS</italic>), Nanjing, 10-12 January 2023, 122-129. https://doi.org/10.1109/icpads56603.2022.00024 <pub-id pub-id-type="doi">10.1109/icpads56603.2022.00024</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/icpads56603.2022.00024">https://doi.org/10.1109/icpads56603.2022.00024</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Huang, T.</string-name>
              <string-name>Liu, G.</string-name>
              <string-name>Li, S.</string-name>
              <string-name>Liu, J.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>RPCRS: Human Activity Recognition Using Millimeter Wave Radar</article-title>
            <source>2022 IEEE 28th International Conference on Parallel and Distributed Systems (ICPADS)</source>
            <volume>10</volume>
            <pub-id pub-id-type="doi">10.1109/icpads56603.2022.00024</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B12">
        <label>12.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Shan, L., Zhang, R., Chilukoti, S.V., Zhang, X., Lee, I. and Hei, X. (2024) IdentityKD: Identity-Wise Cross-Modal Knowledge Distillation for Person Recognition via Mmwave Radar Sensors. <italic>Proceedings of the</italic>6 <italic>th ACM International Conference on Multimedia in Asia</italic>, Auckland, 3-6 December 2024, 1-7. https://doi.org/10.1145/3696409.3700254 <pub-id pub-id-type="doi">10.1145/3696409.3700254</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3696409.3700254">https://doi.org/10.1145/3696409.3700254</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Shan, L.</string-name>
              <string-name>Zhang, R.</string-name>
              <string-name>Chilukoti, S.V.</string-name>
              <string-name>Zhang, X.</string-name>
              <string-name>Lee, I.</string-name>
              <string-name>Hei, X.</string-name>
              <string-name>Asia, A</string-name>
            </person-group>
            <year>2024</year>
            <article-title>IdentityKD: Identity-Wise Cross-Modal Knowledge Distillation for Person Recognition via Mmwave Radar Sensors</article-title>
            <source>Proceedings of the 6th ACM International Conference on Multimedia in Asia</source>
            <volume>3</volume>
            <pub-id pub-id-type="doi">10.1145/3696409.3700254</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B13">
        <label>13.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Yang, Z. and Yang, Y. (2025) Wisk: A Lightweight Multimodal Human Activity Recognition Method Based on WiFi Channel State Information and Video Skeleton Images. <italic>Concurrency and Computation</italic>: <italic>Practice and Experience</italic>, 37, e70383. https://doi.org/10.1002/cpe.70383 <pub-id pub-id-type="doi">10.1002/cpe.70383</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1002/cpe.70383">https://doi.org/10.1002/cpe.70383</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Yang, Z.</string-name>
              <string-name>Yang, Y.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>Wisk: A Lightweight Multimodal Human Activity Recognition Method Based on WiFi Channel State Information and Video Skeleton Images</article-title>
            <source>Concurrency and Computation: Practice and Experience</source>
            <volume>37</volume>
            <pub-id pub-id-type="doi">10.1002/cpe.70383</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B14">
        <label>14.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Cao, P., Xia, W., Ye, M., Zhang, J. and Zhou, J. (2018) Radar‐ID: Human Identification Based on Radar Micro‐Doppler Signatures Using Deep Convolutional Neural Networks. <italic>IET Radar</italic>, <italic>Sonar &amp; Navigation</italic>, 12, 729-734. https://doi.org/10.1049/iet-rsn.2017.0511 <pub-id pub-id-type="doi">10.1049/iet-rsn.2017.0511</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1049/iet-rsn.2017.0511">https://doi.org/10.1049/iet-rsn.2017.0511</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Cao, P.</string-name>
              <string-name>Xia, W.</string-name>
              <string-name>Ye, M.</string-name>
              <string-name>Zhang, J.</string-name>
              <string-name>Zhou, J.</string-name>
              <string-name>Radar, S</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Radar‐ID: Human Identification Based on Radar Micro‐Doppler Signatures Using Deep Convolutional Neural Networks</article-title>
            <source>IET Radar</source>
            <volume>12</volume>
            <pub-id pub-id-type="doi">10.1049/iet-rsn.2017.0511</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B15">
        <label>15.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Hafner, F.M., Bhuiyan, A., Kooij, J.F.P. and Granger, E. (2019) RGB-Depth Cross-Modal Person Re-Identification. 2019 16 <italic>th IEEE International Conference on Advanced Video and Signal Based Surveillance</italic> ( <italic>AVSS</italic>), 18-21 September 2019, 1-8. https://doi.org/10.1109/avss.2019.8909838 <pub-id pub-id-type="doi">10.1109/avss.2019.8909838</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/avss.2019.8909838">https://doi.org/10.1109/avss.2019.8909838</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Hafner, F.M.</string-name>
              <string-name>Bhuiyan, A.</string-name>
              <string-name>Kooij, J.F.P.</string-name>
              <string-name>Granger, E.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>RGB-Depth Cross-Modal Person Re-Identification</article-title>
            <source>2019 16th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS)</source>
            <volume>18</volume>
            <pub-id pub-id-type="doi">10.1109/avss.2019.8909838</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B16">
        <label>16.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Ahuja, K., Jiang, Y., Goel, M. and Harrison, C. (2021) Vid2Doppler: Synthesizing Doppler Radar Data from Videos for Training Privacy-Preserving Activity Recognition. <italic>Proceedings of the</italic>2021 <italic>CHI Conference on Human Factors in Computing Systems</italic>, Yokohama, 8-13 May 2021, 1-10. https://doi.org/10.1145/3411764.3445138 <pub-id pub-id-type="doi">10.1145/3411764.3445138</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3411764.3445138">https://doi.org/10.1145/3411764.3445138</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Ahuja, K.</string-name>
              <string-name>Jiang, Y.</string-name>
              <string-name>Goel, M.</string-name>
              <string-name>Harrison, C.</string-name>
              <string-name>Systems, Y</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Vid2Doppler: Synthesizing Doppler Radar Data from Videos for Training Privacy-Preserving Activity Recognition</article-title>
            <source>Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems</source>
            <volume>8</volume>
            <pub-id pub-id-type="doi">10.1145/3411764.3445138</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B17">
        <label>17.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Xie, S., Wang, C., Yang, X., Wan, Y., Zeng, T. and Liu, Z. (2022) Millimeter-Wave Radar Target Detection Based on Inter-Frame DBSCAN Clustering. 2022 <italic>IEEE</italic>22 <italic>nd International Conference on Communication Technology</italic> ( <italic>ICCT</italic>), Nanjing, 11-14 November 2022, 1703-1708. https://doi.org/10.1109/icct56141.2022.10072664 <pub-id pub-id-type="doi">10.1109/icct56141.2022.10072664</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/icct56141.2022.10072664">https://doi.org/10.1109/icct56141.2022.10072664</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Xie, S.</string-name>
              <string-name>Wang, C.</string-name>
              <string-name>Yang, X.</string-name>
              <string-name>Wan, Y.</string-name>
              <string-name>Zeng, T.</string-name>
              <string-name>Liu, Z.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Millimeter-Wave Radar Target Detection Based on Inter-Frame DBSCAN Clustering</article-title>
            <source>2022 IEEE 22nd International Conference on Communication Technology (ICCT)</source>
            <volume>11</volume>
            <pub-id pub-id-type="doi">10.1109/icct56141.2022.10072664</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B18">
        <label>18.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Shao, T., Du, Z., Li, C., Wu, T. and Wang, M. (2024) Fast Human Action Recognition via Millimeter Wave Radar Point Cloud Sequences Learning. <italic>Proceedings of the</italic>33 <italic>rd ACM International Conference on Information and Knowledge Management</italic>, Boise, 21-25 October 2024, 2024-2033. https://doi.org/10.1145/3627673.3679787 <pub-id pub-id-type="doi">10.1145/3627673.3679787</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3627673.3679787">https://doi.org/10.1145/3627673.3679787</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Shao, T.</string-name>
              <string-name>Du, Z.</string-name>
              <string-name>Li, C.</string-name>
              <string-name>Wu, T.</string-name>
              <string-name>Wang, M.</string-name>
              <string-name>Management, B</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Fast Human Action Recognition via Millimeter Wave Radar Point Cloud Sequences Learning</article-title>
            <source>Proceedings of the 33rd ACM International Conference on Information and Knowledge Management</source>
            <volume>21</volume>
            <pub-id pub-id-type="doi">10.1145/3627673.3679787</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B19">
        <label>19.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Zong, M., Wu, J., Zhu, Z. and Ni, J. (2024) A Method for Target Detection Based on Mmw Radar and Vision Fusion. arXiv:2403.16476.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Zong, M.</string-name>
              <string-name>Wu, J.</string-name>
              <string-name>Zhu, Z.</string-name>
              <string-name>Ni, J.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>A Method for Target Detection Based on Mmw Radar and Vision Fusion</article-title>
            <fpage>2403</fpage>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B20">
        <label>20.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Wang, K., He, R., Wang, W., Wang, L. and Tan, T. (2013) Learning Coupled Feature Spaces for Cross-Modal Matching. 2013 <italic>IEEE International Conference on Computer Vision</italic>, Sydney, 1-8 December 2013, 2088-2095. https://doi.org/10.1109/iccv.2013.261 <pub-id pub-id-type="doi">10.1109/iccv.2013.261</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/iccv.2013.261">https://doi.org/10.1109/iccv.2013.261</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Wang, K.</string-name>
              <string-name>He, R.</string-name>
              <string-name>Wang, W.</string-name>
              <string-name>Wang, L.</string-name>
              <string-name>Tan, T.</string-name>
              <string-name>Vision, S</string-name>
            </person-group>
            <year>2013</year>
            <article-title>Learning Coupled Feature Spaces for Cross-Modal Matching</article-title>
            <source>2013 IEEE International Conference on Computer Vision</source>
            <volume>1</volume>
            <pub-id pub-id-type="doi">10.1109/iccv.2013.261</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B21">
        <label>21.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Wei, J., Xu, X., Yang, Y., Ji, Y., Wang, Z. and Shen, H.T. (2020) Universal Weighting Metric Learning for Cross-Modal Matching. 2020 <italic>IEEE</italic>/ <italic>CVF Conference on Computer Vision and Pattern Recognition</italic> ( <italic>CVPR</italic>), Seattle, 13-19 June 2020, 6420-6429. https://doi.org/10.1109/cvpr42600.2020.01302 <pub-id pub-id-type="doi">10.1109/cvpr42600.2020.01302</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/cvpr42600.2020.01302">https://doi.org/10.1109/cvpr42600.2020.01302</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Wei, J.</string-name>
              <string-name>Xu, X.</string-name>
              <string-name>Yang, Y.</string-name>
              <string-name>Ji, Y.</string-name>
              <string-name>Wang, Z.</string-name>
              <string-name>Shen, H.T.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Universal Weighting Metric Learning for Cross-Modal Matching</article-title>
            <source>2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>
            <volume>13</volume>
            <pub-id pub-id-type="doi">10.1109/cvpr42600.2020.01302</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B22">
        <label>22.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Yang, S., Xu, Z., Wang, K., You, Y., Yao, H., Liu, T., <italic>et al</italic>. (2023) BiCro: Noisy Correspondence Rectification for Multi-Modality Data via Bi-Directional Cross-Modal Similarity Consistency. 2023 <italic>IEEE</italic>/ <italic>CVF Conference on Computer Vision and Pattern Recognition</italic> ( <italic>CVPR</italic>), Vancouver, 17-24 June 2023, 19883-19892. https://doi.org/10.1109/cvpr52729.2023.01904 <pub-id pub-id-type="doi">10.1109/cvpr52729.2023.01904</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/cvpr52729.2023.01904">https://doi.org/10.1109/cvpr52729.2023.01904</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Yang, S.</string-name>
              <string-name>Xu, Z.</string-name>
              <string-name>Wang, K.</string-name>
              <string-name>You, Y.</string-name>
              <string-name>Yao, H.</string-name>
              <string-name>Liu, T.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>BiCro: Noisy Correspondence Rectification for Multi-Modality Data via Bi-Directional Cross-Modal Similarity Consistency</article-title>
            <source>2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>
            <volume>17</volume>
            <pub-id pub-id-type="doi">10.1109/cvpr52729.2023.01904</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B23">
        <label>23.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Liong, V.E., Lu, J., Tan, Y. and Zhou, J. (2017) Deep Coupled Metric Learning for Cross-Modal Matching. <italic>IEEE Transactions on Multimedia</italic>, 19, 1234-1244. https://doi.org/10.1109/tmm.2016.2646180 <pub-id pub-id-type="doi">10.1109/tmm.2016.2646180</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/tmm.2016.2646180">https://doi.org/10.1109/tmm.2016.2646180</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Liong, V.E.</string-name>
              <string-name>Lu, J.</string-name>
              <string-name>Tan, Y.</string-name>
              <string-name>Zhou, J.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Deep Coupled Metric Learning for Cross-Modal Matching</article-title>
            <source>IEEE Transactions on Multimedia</source>
            <volume>19</volume>
            <pub-id pub-id-type="doi">10.1109/tmm.2016.2646180</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B24">
        <label>24.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Wei, J., Xu, X., Wang, Z. and Wang, G. (2021) Meta Self-Paced Learning for Cross-Modal Matching. <italic>Proceedings of the</italic>29 <italic>th ACM International Conference on Multimedia</italic>, Virtual Event, 20-24 October 2021, 1613-1621. https://doi.org/10.1145/3474085.3475451 <pub-id pub-id-type="doi">10.1145/3474085.3475451</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3474085.3475451">https://doi.org/10.1145/3474085.3475451</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Wei, J.</string-name>
              <string-name>Xu, X.</string-name>
              <string-name>Wang, Z.</string-name>
              <string-name>Wang, G.</string-name>
              <string-name>Multimedia, V</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Meta Self-Paced Learning for Cross-Modal Matching</article-title>
            <source>Proceedings of the 29th ACM International Conference on Multimedia</source>
            <volume>20</volume>
            <pub-id pub-id-type="doi">10.1145/3474085.3475451</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B25">
        <label>25.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Yao, Y., Zhang, H., Xia, P., Liu, C., Geng, F., Bai, Z., <italic>et al</italic>. (2023) Mmsignature: Semi-Supervised Human Identification System Based on Millimeter Wave Radar. <italic>Engineering Applications of Artificial Intelligence</italic>, 126, Article 106939. https://doi.org/10.1016/j.engappai.2023.106939 <pub-id pub-id-type="doi">10.1016/j.engappai.2023.106939</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.engappai.2023.106939">https://doi.org/10.1016/j.engappai.2023.106939</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Yao, Y.</string-name>
              <string-name>Zhang, H.</string-name>
              <string-name>Xia, P.</string-name>
              <string-name>Liu, C.</string-name>
              <string-name>Geng, F.</string-name>
              <string-name>Bai, Z.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Mmsignature: Semi-Supervised Human Identification System Based on Millimeter Wave Radar</article-title>
            <source>Engineering Applications of Artificial Intelligence</source>
            <volume>126</volume>
            <elocation-id>106939</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.engappai.2023.106939</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B26">
        <label>26.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Deng, K., Zhao, D., Han, Q., Zhang, Z., Wang, S., Zhou, A., <italic>et al</italic>. (2023) Midas: Generating mmWave Radar Data from Videos for Training Pervasive and Privacy-Preserving Human Sensing Tasks. <italic>Proceedings of the ACM on Interactive</italic>, <italic>Mobile</italic>, <italic>Wearable</italic><italic>and Ubiquitous Technologies</italic>, 7, 1-26. https://doi.org/10.1145/3580872 <pub-id pub-id-type="doi">10.1145/3580872</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3580872">https://doi.org/10.1145/3580872</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Deng, K.</string-name>
              <string-name>Zhao, D.</string-name>
              <string-name>Han, Q.</string-name>
              <string-name>Zhang, Z.</string-name>
              <string-name>Wang, S.</string-name>
              <string-name>Zhou, A.</string-name>
              <string-name>Interactive, M</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Midas: Generating mmWave Radar Data from Videos for Training Pervasive and Privacy-Preserving Human Sensing Tasks</article-title>
            <source>Proceedings of the ACM on Interactive</source>
            <volume>7</volume>
            <pub-id pub-id-type="doi">10.1145/3580872</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B27">
        <label>27.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Hafner, F.M., Bhuyian, A., Kooij, J.F.P. and Granger, E. (2022) Cross-Modal Distillation for RGB-Depth Person Re-Identification. <italic>Computer Vision and Image Understanding</italic>, 216, Article 103352. https://doi.org/10.1016/j.cviu.2021.103352 <pub-id pub-id-type="doi">10.1016/j.cviu.2021.103352</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.cviu.2021.103352">https://doi.org/10.1016/j.cviu.2021.103352</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Hafner, F.M.</string-name>
              <string-name>Bhuyian, A.</string-name>
              <string-name>Kooij, J.F.P.</string-name>
              <string-name>Granger, E.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Cross-Modal Distillation for RGB-Depth Person Re-Identification</article-title>
            <source>Computer Vision and Image Understanding</source>
            <volume>216</volume>
            <elocation-id>103352</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.cviu.2021.103352</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B28">
        <label>28.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Janakaraj, P., Jakkala, K., Bhuyan, A., Sun, Z., Wang, P. and Lee, M. (2019) STAR: Simultaneous Tracking and Recognition through Millimeter Waves and Deep Learning. 2019 12 <italic>th IFIP Wireless and Mobile Networking Conference</italic> ( <italic>WMNC</italic>), Paris, 11-13 September 2019, 211-218. https://doi.org/10.23919/wmnc.2019.8881354 <pub-id pub-id-type="doi">10.23919/wmnc.2019.8881354</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.23919/wmnc.2019.8881354">https://doi.org/10.23919/wmnc.2019.8881354</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Janakaraj, P.</string-name>
              <string-name>Jakkala, K.</string-name>
              <string-name>Bhuyan, A.</string-name>
              <string-name>Sun, Z.</string-name>
              <string-name>Wang, P.</string-name>
              <string-name>Lee, M.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>STAR: Simultaneous Tracking and Recognition through Millimeter Waves and Deep Learning</article-title>
            <source>2019 12th IFIP Wireless and Mobile Networking Conference (WMNC)</source>
            <volume>11</volume>
            <pub-id pub-id-type="doi">10.23919/wmnc.2019.8881354</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B29">
        <label>29.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">He, X.Y., Yu, H., Peng, S.D., Tan, D.L., Shen, Z.H., Bao, H.J. and Zhou, X.W. (2025) MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training. arXiv:2501.07556v1.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>He, X.Y.</string-name>
              <string-name>Yu, H.</string-name>
              <string-name>Peng, S.D.</string-name>
              <string-name>Tan, D.L.</string-name>
              <string-name>Shen, Z.H.</string-name>
              <string-name>Bao, H.J.</string-name>
              <string-name>Zhou, X.W.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training</article-title>
            <fpage>2501</fpage>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B30">
        <label>30.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Vishwakarma, S., Li, W., Tang, C., Woodbridge, K., Adve, R. and Chetty, K. (2023) SimHumalator: An Open Source End-to-End Radar Simulator for Human Activity Recognition. <italic>IEEE Aerospace and Electronic Systems Magazine</italic>, 37, 6-22. https://doi.org/10.1109/MAES.2021.3138948 <pub-id pub-id-type="doi">10.1109/MAES.2021.3138948</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/MAES.2021.3138948">https://doi.org/10.1109/MAES.2021.3138948</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Vishwakarma, S.</string-name>
              <string-name>Li, W.</string-name>
              <string-name>Tang, C.</string-name>
              <string-name>Woodbridge, K.</string-name>
              <string-name>Adve, R.</string-name>
              <string-name>Chetty, K.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>SimHumalator: An Open Source End-to-End Radar Simulator for Human Activity Recognition</article-title>
            <source>IEEE Aerospace and Electronic Systems Magazine</source>
            <volume>37</volume>
            <pub-id pub-id-type="doi">10.1109/MAES.2021.3138948</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B31">
        <label>31.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Chen, A., Wang, X., Shi, K., Huo, Y., Chen, J. and Ye, Q. (2025) Toward Weather-Robust 3D Human Body Reconstruction: Millimeter-Wave Radar-Based Dataset, Benchmark, and Multi-Modal Fusion. <italic>IEEE Transactions on Circuits and Systems for Video Technology</italic>, 35, 273-286. https://doi.org/10.1109/tcsvt.2024.3461960 <pub-id pub-id-type="doi">10.1109/tcsvt.2024.3461960</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/tcsvt.2024.3461960">https://doi.org/10.1109/tcsvt.2024.3461960</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Chen, A.</string-name>
              <string-name>Wang, X.</string-name>
              <string-name>Shi, K.</string-name>
              <string-name>Huo, Y.</string-name>
              <string-name>Chen, J.</string-name>
              <string-name>Ye, Q.</string-name>
              <string-name>Dataset, B</string-name>
            </person-group>
            <year>2025</year>
            <article-title>Toward Weather-Robust 3D Human Body Reconstruction: Millimeter-Wave Radar-Based Dataset, Benchmark, and Multi-Modal Fusion</article-title>
            <source>IEEE Transactions on Circuits and Systems for Video Technology</source>
            <volume>35</volume>
            <pub-id pub-id-type="doi">10.1109/tcsvt.2024.3461960</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B32">
        <label>32.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Choi, J., Hor, S., Yang, S. and Arbabian, A. (2025) Mvdoppler-Pose: Multi-Modal Multi-View Mmwave Sensing for Long-Distance Self-Occluded Human Walking Pose Estimation. 2025 <italic>IEEE</italic>/ <italic>CVF Conference on Computer Vision and Pattern Recognition</italic> ( <italic>CVPR</italic>), Nashville, 10-17 June 2025, 1-8. https://doi.org/10.1109/cvpr52734.2025.02584 <pub-id pub-id-type="doi">10.1109/cvpr52734.2025.02584</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/cvpr52734.2025.02584">https://doi.org/10.1109/cvpr52734.2025.02584</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Choi, J.</string-name>
              <string-name>Hor, S.</string-name>
              <string-name>Yang, S.</string-name>
              <string-name>Arbabian, A.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>Mvdoppler-Pose: Multi-Modal Multi-View Mmwave Sensing for Long-Distance Self-Occluded Human Walking Pose Estimation</article-title>
            <source>2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>
            <volume>10</volume>
            <pub-id pub-id-type="doi">10.1109/cvpr52734.2025.02584</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B33">
        <label>33.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Huan, S., Wang, Z., Wang, X., Wu, L., Yang, X., Huang, H., <italic>et al</italic>. (2023) A Lightweight Hybrid Vision Transformer Network for Radar-Based Human Activity Recognition. <italic>Scientific</italic><italic>Reports</italic>, 13, Article No. 17996. https://doi.org/10.1038/s41598-023-45149-5 <pub-id pub-id-type="doi">10.1038/s41598-023-45149-5</pub-id><pub-id pub-id-type="pmid">37865672</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41598-023-45149-5">https://doi.org/10.1038/s41598-023-45149-5</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Huan, S.</string-name>
              <string-name>Wang, Z.</string-name>
              <string-name>Wang, X.</string-name>
              <string-name>Wu, L.</string-name>
              <string-name>Yang, X.</string-name>
              <string-name>Huang, H.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>A Lightweight Hybrid Vision Transformer Network for Radar-Based Human Activity Recognition</article-title>
            <source>Scientific Reports</source>
            <volume>13</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1038/s41598-023-45149-5</pub-id>
            <pub-id pub-id-type="pmid">37865672</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>