<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article">
 <front>
  <journal-meta>
   <journal-id journal-id-type="publisher-id">
    jcc
   </journal-id>
   <journal-title-group>
    <journal-title>
     Journal of Computer and Communications
    </journal-title>
   </journal-title-group>
   <issn pub-type="epub">
    2327-5219
   </issn>
   <issn publication-format="print">
    2327-5227
   </issn>
   <publisher>
    <publisher-name>
     Scientific Research Publishing
    </publisher-name>
   </publisher>
  </journal-meta>
  <article-meta>
   <article-id pub-id-type="doi">
    10.4236/jcc.2025.136005
   </article-id>
   <article-id pub-id-type="publisher-id">
    jcc-143380
   </article-id>
   <article-categories>
    <subj-group subj-group-type="heading">
     <subject>
      Articles
     </subject>
    </subj-group>
    <subj-group subj-group-type="Discipline-v2">
     <subject>
      Computer Science 
     </subject>
     <subject>
       Communications
     </subject>
    </subj-group>
   </article-categories>
   <title-group>
    A Survey of Pedestrian Re-Identification Based on Millimeter Wave Radar and Vision Fusion
   </title-group>
   <contrib-group>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Qingyuan
      </surname>
      <given-names>
       Yang
      </given-names>
     </name>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Zhipeng
      </surname>
      <given-names>
       Quan
      </given-names>
     </name>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Jingxuan
      </surname>
      <given-names>
       Li
      </given-names>
     </name>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Tingyv
      </surname>
      <given-names>
       Jiang
      </given-names>
     </name>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Zhihao
      </surname>
      <given-names>
       Deng
      </given-names>
     </name>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Xinran
      </surname>
      <given-names>
       Qiu
      </given-names>
     </name>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Zhengjie
      </surname>
      <given-names>
       Wang
      </given-names>
     </name>
    </contrib>
   </contrib-group> 
   <aff id="affnull">
    <addr-line>
     aCollege of Electronic and Information Engineering, Shandong University of Science and Technology, Qingdao, China
    </addr-line> 
   </aff> 
   <pub-date pub-type="epub">
    <day>
     11
    </day> 
    <month>
     06
    </month>
    <year>
     2025
    </year>
   </pub-date> 
   <volume>
    13
   </volume> 
   <issue>
    06
   </issue>
   <fpage>
    64
   </fpage>
   <lpage>
    80
   </lpage>
   <history>
    <date date-type="received">
     <day>
      30,
     </day>
     <month>
      April
     </month>
     <year>
      2025
     </year>
    </date>
    <date date-type="published">
     <day>
      16,
     </day>
     <month>
      April
     </month>
     <year>
      2025
     </year> 
    </date> 
    <date date-type="accepted">
     <day>
      16,
     </day>
     <month>
      June
     </month>
     <year>
      2025
     </year> 
    </date>
   </history>
   <permissions>
    <copyright-statement>
     © Copyright 2014 by authors and Scientific Research Publishing Inc. 
    </copyright-statement>
    <copyright-year>
     2014
    </copyright-year>
    <license>
     <license-p>
      This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/
     </license-p>
    </license>
   </permissions>
   <abstract>
    With the advancement of technology and the growth of human demand, pedestrian re-identification is a key technology of intelligent systems and plays an important role in daily life. Traditional vision methods have certain limitations, but the millimeter-wave-vision fusion system, which complements the advantages of cameras and millimeter-wave radars, plays a greater role and is attracting widespread attention. This paper first introduces the effects of vision camera and millimeter-wave radar camera on human re-identification in a single mode, and then discusses the detailed processing required to fuse the data of the two modes, including the key technologies of sensor settings, data synchronization and sensor calibration. We also review the classification and evolution of millimeter-wave radar and visual fusion pedestrian re-identification methods, including data-level, feature-level and decision-level methods, and review the previous research methods, which will provide great inspiration for future research. Finally, this paper discusses the typical applications of millimeter-wave vision fusion systems, as well as the key technologies and potential challenges of fusing millimeter-wave radar and visual data, and looks forward to future research directions.
   </abstract>
   <kwd-group> 
    <kwd>
     Pedestrian Re-Identification
    </kwd> 
    <kwd>
      RGB Cameras
    </kwd> 
    <kwd>
      mmWave Radar
    </kwd> 
    <kwd>
      Feature Fusion
    </kwd>
   </kwd-group>
  </article-meta>
 </front>
 <body>
  <sec id="s1">
   <title>1. Introduction</title>
   <p>Pedestrian re-identification (ReID) is a fundamental computer vision task that involves matching images or video sequences of the same person captured by different cameras at different times and locations <xref ref-type="bibr" rid="scirp.143380-1">
     [1]
    </xref>. With a wide range of promising prospects, it plays a crucial role in various applications, including intelligent surveillance systems for public safety, human-robot interaction, autonomous driving perception, and retail analytics <xref ref-type="bibr" rid="scirp.143380-2">
     [2]
    </xref>.</p>
   <p>Traditionally, ReID research has focused on visual data captured by RGB cameras <xref ref-type="bibr" rid="scirp.143380-3">
     [3]
    </xref>. Significant advancements have been made leveraging deep learning techniques, leading to impressive performance on benchmark datasets under ideal conditions <xref ref-type="bibr" rid="scirp.143380-4">
     [4]
    </xref>. However, visual ReID methods face inherent limitations in real-world scenarios. Their performance drastically degrades under challenging environmental conditions, such as: illumination variations, occlusion, viewpoint changes, low resolution, adverse weather, etc.</p>
   <p>To overcome these limitations, researchers have increasingly explored multi-modal approaches, integrating information from sensors beyond the visible spectrum <xref ref-type="bibr" rid="scirp.143380-5">
     [5]
    </xref>. Among various sensor modalities, millimeter-wave (mmWave) radar has emerged as a particularly promising candidate for enhancing pedestrian ReID robustness <xref ref-type="bibr" rid="scirp.143380-6">
     [6]
    </xref>. mmWave radar operates in the 30 - 300 GHz frequency range, offering unique advantages: robustness to environmental conditions, range and velocity information, penetration capability, privacy preservation, etc.</p>
   <p>However, mmWave radar also has limitations, primarily its low spatial resolution compared to cameras and the sparsity of its point cloud data, which lacks rich texture and color information essential for appearance-based matching. This inherent complementarity between visual cameras and mmWave radar motivates their fusion for pedestrian ReID. By combining the strengths of both modalities, fused systems aim to achieve more reliable and robust performance across a wider range of operating conditions than is possible with either sensor alone <xref ref-type="bibr" rid="scirp.143380-7">
     [7]
    </xref>. <xref ref-type="fig" rid="fig1">
     Figure 1
    </xref> illustrates the conceptual framework of mmWave-visual fusion for enhanced pedestrian ReID.</p>
   <fig id="fig1" position="float">
    <label>Figure 1</label>
    <caption>
     <title>Figure 1. Conceptual diagram of mmWave radar and visual fusion for pedestrian re-identification.</title>
    </caption>
    <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1733173-rId16.jpeg?20250619013725" />
   </fig>
   <p>The system integrates data from cameras (capturing appearance) and mmWave radar (capturing range, velocity, and operating in adverse conditions) to achieve more robust ReID performance compared to single-modality systems.</p>
   <p>This review provides a comprehensive survey of the research landscape in mmWave-visual fused pedestrian ReID. The main contributions are:</p>
   <p>1) A clear explanation of the fundamental concepts and the motivation behind fusing mmWave radar and visual data for ReID.</p>
   <p>2) A detailed discussion and comparison of techniques for data acquisition, synchronization, calibration, and pre-processing specific to this multi-modal setup.</p>
   <p>3) A systematic classification and critical analysis of existing fusion methods, highlighting their evolution and comparative performance.</p>
   <p>4) An identification of key challenges and a discussion of promising future research directions.</p>
   <p>This review aims to serve as a valuable resource for researchers entering the field and to stimulate further advancements in robust multi-modal perception systems. The subsequent sections delve into the fundamental theories (Section 2), data acquisition and processing (Section 3), fusion methodologies (Section 4), typical applications (Section 5), challenges and future directions (Section 6), and finally conclude the review (Section 7).</p>
  </sec><sec id="s2">
   <title>2. Fundamental Theories about Pedestrian Re-Identification</title>
   <p>This section establishes the theoretical foundation necessary to understand mmWave-visual pedestrian ReID. We define pedestrian ReID and multi-modal ReID, followed by an analysis of the characteristics of visual and mmWave modalities and the rationale underpinning their fusion.</p>
   <sec id="s2_1">
    <title>
     <xref ref-type="bibr" rid="scirp.143380-"></xref>2.1. Pedestrian Re-Identification (ReID)</title>
    <p>Pedestrian Re-identification is the task of associating observations of the same individual across a network of non-overlapping camera views <xref ref-type="bibr" rid="scirp.143380-1">
      [1]
     </xref>. Given a query image of a person of interest captured in one camera view, the objective is to retrieve all instances of the same person from a gallery set containing images from other camera views. It is fundamentally a matching problem, aiming to determine if two observations correspond to the same physical person despite variations in viewpoint, pose, illumination, occlusion, and background clutter. Early methods relied on hand-crafted features <xref ref-type="bibr" rid="scirp.143380-8">
      [8]
     </xref>, while modern approaches predominantly utilize deep learning to learn discriminative feature representations directly from data <xref ref-type="bibr" rid="scirp.143380-3">
      [3]
     </xref> <xref ref-type="bibr" rid="scirp.143380-4">
      [4]
     </xref>.</p>
   </sec>
   <sec id="s2_2">
    <title>2.2. Multi-Modal Pedestrian ReID</title>
    <p>Multi-Modal Pedestrian ReID extends the traditional ReID paradigm by leveraging information from multiple sensing modalities to improve matching accuracy and robustness <xref ref-type="bibr" rid="scirp.143380-5">
      [5]
     </xref>. While visual data (RGB) is the most common modality, it can be limited under challenging conditions. Integrating complementary sensor data, such as depth maps (RGB-D), thermal imagery (RGB-T), or radio frequency signals (like mmWave radar), can provide additional cues to overcome the shortcomings of visual sensors alone. The core idea is that different modalities capture distinct aspects of a person or the environment, and their fusion can lead to a more comprehensive and resilient representation for matching <xref ref-type="bibr" rid="scirp.143380-9">
      [9]
     </xref>.</p>
   </sec>
   <sec id="s2_3">
    <title>2.3. Characteristics of Visual and mmWave Modalities</title>
    <p>Understanding the individual strengths and weaknesses of visual cameras and mmWave radar is crucial for effective fusion.</p>
   </sec>
   <sec id="s2_4">
    <title>2.4. Necessity for Fusion</title>
    <p>The complementary nature of visual and mmWave data forms the core rationale for their fusion in pedestrian ReID. Cameras excel at capturing appearance details crucial for identification in good conditions, while mmWave radar, which persists even when visual data is degraded, provides robust geometric and velocity information.</p>
    <p>Fusion aims to leverage these complementary strengths:</p>
    <p>Compared with systems that rely on a single modality, this combination ensures that the application system is more robust and reliable in complex real-world environments <xref ref-type="bibr" rid="scirp.143380-10">
      [10]
     </xref>.</p>
   </sec>
  </sec><sec id="s3">
   <title>3. Data Acquisition and Processing</title>
   <p>Successfully fusing mmWave radar and visual data for pedestrian ReID depends on data acquisition and processing. This section discusses the typical sensor setups, critical techniques for data synchronization and sensor calibration, and essential pre-processing steps for both modalities.</p>
   <sec id="s3_1">
    <title>3.1. Sensor Setup and Hardware Configuration</title>
    <p>A typical mmWave-visual ReID system involves one or more cameras and mmWave radar sensors deployed to cover the area of interest.</p>
    <p>This configuration facilitates calibration and ensures that both sensors observe the same scene region, simplifying data association. The diagram should also present the subsequent processing steps.</p>
   </sec>
   <sec id="s3_2">
    <title>3.2. Data Synchronization</title>
    <p>Temporal alignment of data streams from cameras and radars is critical for accurate fusion. Mismatched timestamps can lead to incorrect associations between visual features and radar points corresponding to the same pedestrian at a given moment. Common synchronization techniques include:</p>
    <fig id="fig2" position="float">
     <label>Figure 2</label>
     <caption>
      <title>Figure 2. Example of a co-located mmWave radar and camera sensor setup.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1733173-rId17.jpeg?20250619013727" />
    </fig>
   </sec>
   <sec id="s3_3">
    <title>3.3. Sensor Calibration</title>
    <p>Spatial alignment, or calibration, is necessary to establish the geometric relationship between the coordinate systems of the camera(s) and the radar(s). This allows projecting radar points onto the image plane or transforming visual features into the radar’s coordinate system, enabling data association and fusion.</p>
    <p>In multi-sensor systems, the collaboration between millimeter-wave radar and RGB cameras faces three core challenges: calibration, synchronization, and data association. Hardware-level synchronization can achieve microsecond-level precision but is costly, while software synchronization (NTP/PTP) requires motion compensation algorithms to correct the data misalignment caused by millisecond-level delays. In terms of external parameter calibration, traditional checkerboard calibration has high accuracy but is limited to static environments, whereas target-less calibration is flexible but relies on rich natural feature matching. In data association, the modal differences between the radar’s sparse point cloud and the camera’s dense image can easily lead to target mismatches, necessitating the fusion of probabilistic models and deep learning features to enhance consistency. In practical use, it is essential to balance cost and performance: first, use a calibration board for precise calibration and automatically adjust during operation; prioritize hardware synchronization, but if the budget is insufficient, resort to software synchronization with algorithm compensation. Ultimately, success depends on the collaboration between hardware and algorithms to adapt to environmental changes in real-time.</p>
   </sec>
   <sec id="s3_4">
    <title>3.4. Data Pre-Processing</title>
    <p>Raw data from cameras and radar require significant pre-processing before fusion.</p>
    <p>Visual Data Pre-processing includes the following context:</p>
    <p>mmWave Radar Data Pre-processing including the following context:</p>
    <p>Effective pre-processing cleans the raw data, extracts relevant information (e.g. pedestrian bounding boxes, clustered radar points associated with pedestrians), and prepares it for subsequent fusion stages.</p>
   </sec>
  </sec><sec id="s4">
   <title>4. Classification and Evolution of Pedestrian Re-Identification Methods Using Millimeter-Wave Radar and Visual Fusion</title>
   <p>Pedestrian re-identification (ReID) is a critical task in computer vision and sensor fusion, aiming to recognize and track the same individual across different views or time instances. This section classifies fusion methods into data-level, feature-level, and decision-level approaches, reviews representative studies, and discusses their evolution, supported by comparative tables.</p>
   <sec id="s4_1">
    <title>4.1. Data-Level Fusion</title>
    <p>Data-level fusion involves merging raw or minimally processed data from mmWave radar and visual sensors before detection or identification. This approach leverages radar’s precise localization to generate regions of interest (ROI) in images, which are then processed for re-identification. Early methods relied on straightforward data integration, but recent advancements incorporate deep learning for enhanced performance.</p>
    <p>One of the earliest works, Milch and Behrens <xref ref-type="bibr" rid="scirp.143380-18">
      [18]
     </xref>, proposed a method that uses a radar-generated target list to define the ROIs, and then verifies the pedestrian through visual profile analysis. This approach was computationally efficient but limited by the simplicity of visual features. Guo et al. <xref ref-type="bibr" rid="scirp.143380-19">
      [19]
     </xref> advanced this by introducing intra-frame clustering and inter-frame tracking, using radar data to filter noise and guide visual confirmation. Their method improved robustness in noisy environments.</p>
    <p>More recently, Wang et al. <xref ref-type="bibr" rid="scirp.143380-20">
      [20]
     </xref> developed a three-layer fusion model integrating radar and monocular vision, using a visual attention mechanism to prioritize ROIs. This method enhanced detection in dynamic scenes. Similarly, Streubel and Yang <xref ref-type="bibr" rid="scirp.143380-21">
      [21]
     </xref> explored data-level fusion for indoor pedestrian tracking, projecting radar points onto stereo camera images. Their approach achieved high localization accuracy but required precise sensor calibration.</p>
    <p>Data-level fusion excels in leveraging raw data complementarity but faces challenges in computational complexity and sensor synchronization. Recent studies, such as Plascencia et al. <xref ref-type="bibr" rid="scirp.143380-22">
      [22]
     </xref>, have begun integrating deep learning into process fused data, improving scalability. <xref ref-type="table" rid="table1">
      Table 1
     </xref> compares key data-level fusion methods.</p>
    <table-wrap id="table1">
     <label>
      <xref ref-type="table" rid="table1">
       Table 1
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.143380-"></xref>Table 1. Comparison of data-level fusion methods.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="16.10%"><p style="text-align:center">Study</p></td> 
       <td class="custom-bottom-td acenter" width="9.17%"><p style="text-align:center">Year</p></td> 
       <td class="custom-bottom-td acenter" width="20.62%"><p style="text-align:center">Technique</p></td> 
       <td class="custom-bottom-td acenter" width="14.37%"><p style="text-align:center">Application</p></td> 
       <td class="custom-bottom-td acenter" width="15.75%"><p style="text-align:center">Strengths</p></td> 
       <td class="custom-bottom-td acenter" width="15.92%"><p style="text-align:center">Weaknesses</p></td> 
       <td class="custom-bottom-td acenter" width="8.06%"><p style="text-align:center">Reference</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="16.10%"><p style="text-align:center">Milch and Behrens</p></td> 
       <td class="custom-top-td acenter" width="9.17%"><p style="text-align:center">2001</p></td> 
       <td class="custom-top-td acenter" width="20.62%"><p style="text-align:center">Radar target lists, visual contour</p></td> 
       <td class="custom-top-td acenter" width="14.37%"><p style="text-align:center">Detection</p></td> 
       <td class="custom-top-td acenter" width="15.75%"><p style="text-align:center">Simple, low computation</p></td> 
       <td class="custom-top-td acenter" width="15.92%"><p style="text-align:center">Limited feature complexity</p></td> 
       <td class="custom-top-td acenter" width="8.06%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.143380-18">
          [18]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="16.10%"><p style="text-align:center">Guo et al.</p></td> 
       <td class="acenter" width="9.17%"><p style="text-align:center">2018</p></td> 
       <td class="acenter" width="20.62%"><p style="text-align:center">Clustering, visual confirmation</p></td> 
       <td class="acenter" width="14.37%"><p style="text-align:center">Detection/Tracking</p></td> 
       <td class="acenter" width="15.75%"><p style="text-align:center">Noise reduction</p></td> 
       <td class="acenter" width="15.92%"><p style="text-align:center">Calibration sensitivity</p></td> 
       <td class="acenter" width="8.06%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.143380-19">
          [19]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="16.10%"><p style="text-align:center">Wang et al.</p></td> 
       <td class="acenter" width="9.17%"><p style="text-align:center">2011</p></td> 
       <td class="acenter" width="20.62%"><p style="text-align:center">Three-layer fusion, attention</p></td> 
       <td class="acenter" width="14.37%"><p style="text-align:center">Detection</p></td> 
       <td class="acenter" width="15.75%"><p style="text-align:center">Dynamic scene handling</p></td> 
       <td class="acenter" width="15.92%"><p style="text-align:center">High computation</p></td> 
       <td class="acenter" width="8.06%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.143380-20">
          [20]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="16.10%"><p style="text-align:center">Streubel and Yang</p></td> 
       <td class="acenter" width="9.17%"><p style="text-align:center">2016</p></td> 
       <td class="acenter" width="20.62%"><p style="text-align:center">Radar projection, stereo vision</p></td> 
       <td class="acenter" width="14.37%"><p style="text-align:center">Tracking</p></td> 
       <td class="acenter" width="15.75%"><p style="text-align:center">High accuracy</p></td> 
       <td class="acenter" width="15.92%"><p style="text-align:center">Calibration complexity</p></td> 
       <td class="acenter" width="8.06%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.143380-21">
          [21]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="16.10%"><p style="text-align:center">Plascencia et al.</p></td> 
       <td class="acenter" width="9.17%"><p style="text-align:center">2023</p></td> 
       <td class="acenter" width="20.62%"><p style="text-align:center">Deep learning fusion</p></td> 
       <td class="acenter" width="14.37%"><p style="text-align:center">Detection</p></td> 
       <td class="acenter" width="15.75%"><p style="text-align:center">Scalable</p></td> 
       <td class="acenter" width="15.92%"><p style="text-align:center">Data requirements</p></td> 
       <td class="acenter" width="8.06%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.143380-22">
          [22]
         </xref></p></td> 
      </tr> 
     </table>
    </table-wrap>
   </sec>
   <sec id="s4_2">
    <title>4.2. Feature-Level Fusion</title>
    <p>Feature-level fusion extracts features from radar and visual data separately, and then combines them for re-identification. This approach has gained prominence with deep learning, enabling models to learn complex feature representations and fusion strategies dynamically.</p>
    <p>Nobis et al. <xref ref-type="bibr" rid="scirp.143380-23">
      [23]
     </xref> proposed a deep learning architecture, which fuses projected radar data with camera images within network layers. This method improved 2D detection accuracy by learning optimal fusion levels. Plascencia et al. <xref ref-type="bibr" rid="scirp.143380-22">
      [22]
     </xref> extended this concept, transforming radar and lidar data into 2D grayscale images and fusing them with RGB images using a SegNet-based network. Their approach enhanced pedestrian detection in cluttered environments.</p>
    <p>In addition to that, attention mechanisms have further refined feature-level fusion. Li et al. <xref ref-type="bibr" rid="scirp.143380-24">
      [24]
     </xref> introduced an attention-based network for pedestrian liveness detection, combining radar cross-section (RCS) features with visual data to distinguish real pedestrians from static images. Similarly, Yu et al. <xref ref-type="bibr" rid="scirp.143380-25">
      [25]
     </xref> proposed a dual cross-attention module (DCAM) for feature fusion, initially for vehicle detection but adaptable to pedestrians. Liu et al. <xref ref-type="bibr" rid="scirp.143380-26">
      [26]
     </xref> developed a multi-modal network integrating radar gait features with visual appearance, achieving robust re-identification in occluded scenes.</p>
    <p>Feature-level fusion benefits from deep learning’s ability to model complex relationships but requires substantial training data and computational resources. <xref ref-type="table" rid="table2">
      Table 2
     </xref> summarizes key methods.</p>
    <table-wrap id="table2">
     <label>
      <xref ref-type="table" rid="table2">
       Table 2
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.143380-"></xref>Table 2. Comparison of feature-level fusion methods.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="13.89%"><p style="text-align:center">Study</p></td> 
       <td class="custom-bottom-td acenter" width="7.69%"><p style="text-align:center">Year</p></td> 
       <td class="custom-bottom-td acenter" width="20.09%"><p style="text-align:center">Technique</p></td> 
       <td class="custom-bottom-td acenter" width="14.63%"><p style="text-align:center">Application</p></td> 
       <td class="custom-bottom-td acenter" width="15.64%"><p style="text-align:center">Strengths</p></td> 
       <td class="custom-bottom-td acenter" width="20.28%"><p style="text-align:center">Weaknesses</p></td> 
       <td class="custom-bottom-td acenter" width="7.78%"><p style="text-align:center">Reference</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="13.89%"><p style="text-align:center">Nobis et al.</p></td> 
       <td class="custom-top-td acenter" width="7.69%"><p style="text-align:center">2019</p></td> 
       <td class="custom-top-td acenter" width="20.09%"><p style="text-align:center">CRF-Net, deep fusion</p></td> 
       <td class="custom-top-td acenter" width="14.63%"><p style="text-align:center">Detection</p></td> 
       <td class="custom-top-td acenter" width="15.64%"><p style="text-align:center">Adaptive fusion</p></td> 
       <td class="custom-top-td acenter" width="20.28%"><p style="text-align:center">Data-intensive</p></td> 
       <td class="custom-top-td acenter" width="7.78%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.143380-23">
          [23]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.89%"><p style="text-align:center">Plascencia et al.</p></td> 
       <td class="acenter" width="7.69%"><p style="text-align:center">2023</p></td> 
       <td class="acenter" width="20.09%"><p style="text-align:center">SegNet, grayscale fusion</p></td> 
       <td class="acenter" width="14.63%"><p style="text-align:center">Detection</p></td> 
       <td class="acenter" width="15.64%"><p style="text-align:center">Clutter robustness</p></td> 
       <td class="acenter" width="20.28%"><p style="text-align:center">Computation-heavy</p></td> 
       <td class="acenter" width="7.78%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.143380-22">
          [22]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.89%"><p style="text-align:center">Li et al.</p></td> 
       <td class="acenter" width="7.69%"><p style="text-align:center">2022</p></td> 
       <td class="acenter" width="20.09%"><p style="text-align:center">Attention, RCS features</p></td> 
       <td class="acenter" width="14.63%"><p style="text-align:center">Liveness Detection</p></td> 
       <td class="acenter" width="15.64%"><p style="text-align:center">High specificity</p></td> 
       <td class="acenter" width="20.28%"><p style="text-align:center">Training complexity</p></td> 
       <td class="acenter" width="7.78%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.143380-24">
          [24]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.89%"><p style="text-align:center">Yu et al.</p></td> 
       <td class="acenter" width="7.69%"><p style="text-align:center">2025</p></td> 
       <td class="acenter" width="20.09%"><p style="text-align:center">DCAM fusion</p></td> 
       <td class="acenter" width="14.63%"><p style="text-align:center">Detection</p></td> 
       <td class="acenter" width="15.64%"><p style="text-align:center">Flexible</p></td> 
       <td class="acenter" width="20.28%"><p style="text-align:center">Limited pedestrian focus</p></td> 
       <td class="acenter" width="7.78%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.143380-25">
          [25]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.89%"><p style="text-align:center">Liu et al.</p></td> 
       <td class="acenter" width="7.69%"><p style="text-align:center">2024</p></td> 
       <td class="acenter" width="20.09%"><p style="text-align:center">Gait and visual fusion</p></td> 
       <td class="acenter" width="14.63%"><p style="text-align:center">ReID</p></td> 
       <td class="acenter" width="15.64%"><p style="text-align:center">Occlusion handling</p></td> 
       <td class="acenter" width="20.28%"><p style="text-align:center">Data requirements</p></td> 
       <td class="acenter" width="7.78%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.143380-26">
          [26]
         </xref></p></td> 
      </tr> 
     </table>
    </table-wrap>
   </sec>
   <sec id="s4_3">
    <title>4.3. Decision-Level Fusion</title>
    <p>Decision-level fusion involves independent detection or identification by each sensor, followed by result integration. This modular approach is robust to sensor failures and suitable for real-time applications.</p>
    <p>Yang et al. <xref ref-type="bibr" rid="scirp.143380-27">
      [27]
     </xref> proposed a decision-level fusion method using an unscented Kalman filter (UKF) for radar data and YOLOv5 with DeepSORT for visual tracking, matching targets in polar coordinates. This method achieved high tracking precision. An earlier study <xref ref-type="bibr" rid="scirp.143380-28">
      [28]
     </xref> introduced a multi-sensor tracking algorithm with back-projection and multi-hypothesis association, enhancing trajectory accuracy. Zhao et al. <xref ref-type="bibr" rid="scirp.143380-29">
      [29]
     </xref> focused on nighttime detection, combining infrared vision and radar with an improved YOLOv5 and extended Kalman filter.</p>
    <p>Additional studies include Graves et al. <xref ref-type="bibr" rid="scirp.143380-30">
      [30]
     </xref>, who used decision-level fusion for pedestrian collision warning, integrating radar localization with visual classification, and Zhu et al. <xref ref-type="bibr" rid="scirp.143380-31">
      [31]
     </xref>, who developed a track-to-track fusion method for multi-pedestrian tracking. These methods prioritize simplicity and robustness but may miss early-stage data synergies. <xref ref-type="table" rid="table3">
      Table 3
     </xref> compares key approaches.</p>
    <table-wrap id="table3">
     <label>
      <xref ref-type="table" rid="table3">
       Table 3
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.143380-"></xref>Table 3. Comparison of decision-level fusion methods.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="12.13%"><p style="text-align:center">Study</p></td> 
       <td class="custom-bottom-td acenter" width="6.39%"><p style="text-align:center">Year</p></td> 
       <td class="custom-bottom-td acenter" width="19.97%"><p style="text-align:center">Technique</p></td> 
       <td class="custom-bottom-td acenter" width="17.67%"><p style="text-align:center">Application</p></td> 
       <td class="custom-bottom-td acenter" width="17.67%"><p style="text-align:center">Strengths</p></td> 
       <td class="custom-bottom-td acenter" width="17.67%"><p style="text-align:center">Weaknesses</p></td> 
       <td class="custom-bottom-td acenter" width="8.51%"><p style="text-align:center">Reference</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="12.13%"><p style="text-align:center">Yang et al.</p></td> 
       <td class="custom-top-td acenter" width="6.39%"><p style="text-align:center">2023</p></td> 
       <td class="custom-top-td acenter" width="19.97%"><p style="text-align:center">UKF, YOLOv5/DeepSORT</p></td> 
       <td class="custom-top-td acenter" width="17.67%"><p style="text-align:center">Tracking</p></td> 
       <td class="custom-top-td acenter" width="17.67%"><p style="text-align:center">High precision</p></td> 
       <td class="custom-top-td acenter" width="17.67%"><p style="text-align:center">Limited synergy</p></td> 
       <td class="custom-top-td acenter" width="8.51%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.143380-27">
          [27]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="12.13%"><p style="text-align:center">Cui et al.</p></td> 
       <td class="acenter" width="6.39%"><p style="text-align:center">2021</p></td> 
       <td class="acenter" width="19.97%"><p style="text-align:center">Back-projection, association</p></td> 
       <td class="acenter" width="17.67%"><p style="text-align:center">Tracking</p></td> 
       <td class="acenter" width="17.67%"><p style="text-align:center">Trajectory accuracy</p></td> 
       <td class="acenter" width="17.67%"><p style="text-align:center">Complexity</p></td> 
       <td class="acenter" width="8.51%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.143380-28">
          [28]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="12.13%"><p style="text-align:center">Zhao et al.</p></td> 
       <td class="acenter" width="6.39%"><p style="text-align:center">2023</p></td> 
       <td class="acenter" width="19.97%"><p style="text-align:center">YOLOv5, Kalman filter</p></td> 
       <td class="acenter" width="17.67%"><p style="text-align:center">Nighttime Detection</p></td> 
       <td class="acenter" width="17.67%"><p style="text-align:center">Low-light robustness</p></td> 
       <td class="acenter" width="17.67%"><p style="text-align:center">Data association</p></td> 
       <td class="acenter" width="8.51%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.143380-29">
          [29]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="12.13%"><p style="text-align:center">Graves et al.</p></td> 
       <td class="acenter" width="6.39%"><p style="text-align:center">2022</p></td> 
       <td class="acenter" width="19.97%"><p style="text-align:center">Radar-visual matching</p></td> 
       <td class="acenter" width="17.67%"><p style="text-align:center">Collision Warning</p></td> 
       <td class="acenter" width="17.67%"><p style="text-align:center">Simplicity</p></td> 
       <td class="acenter" width="17.67%"><p style="text-align:center">Limited integration</p></td> 
       <td class="acenter" width="8.51%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.143380-30">
          [30]
         </xref></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="12.13%"><p style="text-align:center">Zhu et al.</p></td> 
       <td class="acenter" width="6.39%"><p style="text-align:center">2022</p></td> 
       <td class="acenter" width="19.97%"><p style="text-align:center">Track-to-track fusion</p></td> 
       <td class="acenter" width="17.67%"><p style="text-align:center">Tracking</p></td> 
       <td class="acenter" width="17.67%"><p style="text-align:center">Multi-target handling</p></td> 
       <td class="acenter" width="17.67%"><p style="text-align:center">Synchronization</p></td> 
       <td class="acenter" width="8.51%"><p style="text-align:center">
         <xref ref-type="bibr" rid="scirp.143380-31">
          [31]
         </xref></p></td> 
      </tr> 
     </table>
    </table-wrap>
   </sec>
   <sec id="s4_4">
    <title>
     <xref ref-type="bibr" rid="scirp.143380-"></xref>4.4. Evolution of Methods</title>
    <p>The evolution of mmWave radar and visual fusion for pedestrian ReID reflects advancements in sensor technology and algorithms. Early methods (2000s) focused on data-level fusion, using radar to guide visual processing in simple scenarios. The 2010s saw decision-level fusion gain traction for its robustness, as seen in collision warning systems. Since 2015, feature-level fusion has dominated, driven by deep learning’s ability to model complex multi-modal relationships. Future directions may involve end-to-end deep learning models and efficient handling of sparse radar data, addressing real-time constraints in autonomous systems.</p>
   </sec>
   <sec id="s4_5">
    <title>4.5. Typical System Analysis</title>
    <p>To better illustrate the fusion approach, we present a detailed explanation of several representative re-identification methods. This will facilitate a more comprehensive understanding of the advantages offered by millimeter-wave radar and vision fusion for re-identification tasks.</p>
    <p>Zheng et al. <xref ref-type="bibr" rid="scirp.143380-1">
      [1]
     </xref> proposed the pedestrian alignment network (PAN) that utilizes two convolutional branches (the base branch and the alignment branch) and an affine estimation branch to simultaneously address pedestrian alignment and recognition issues. The base branch deploys a pre-trained ResNet-50 model on ImageNet and removes the final fully connected (FC) layer. The alignment branch consists of 3 ResBlocks and 1 average pooling layer, also adding an FC layer to predict multiclass probabilities. The affine estimation branch receives two activated input tensors from the base branch. The Res4 Feature Maps contain shallow feature maps of the original image, reflecting local pattern information; the Res2 Feature Maps are closer to the classification layer and encode attention and semantic cues for pedestrian recognition.</p>
    <p>Zheng et al. <xref ref-type="bibr" rid="scirp.143380-2">
      [2]
     </xref> proposed a deep learning model that combines the advantages of verification models and recognition models. Establishing relationships through pairwise comparisons, such as partial matching and contrastive loss, is performed. Contrastive loss directly computes the Euclidean distance between two embeddings. In the recognition model, there exists an implicit relationship between the learned embeddings constructed using cross-entropy loss. The cross-entropy loss can be used. When the directions of the embedding vectors are similar, the network converges, thereby maintaining the similarity of the embeddings. The proposed model simultaneously utilizes both types of loss functions and benefits from pre-training on ImageNet, thereby overcoming the limitations of a single model.</p>
    <p>In the performance evaluation of data-level fusion methods, computational complexity and real-time indicators have significant advantages. Regarding the issue of computational complexity, existing research mainly focuses on two dimensions: noise suppression and accuracy optimization. For example, Yu et al. <xref ref-type="bibr" rid="scirp.143380-25">
      [25]
     </xref> utilized modules to reduce technical complexity, providing a reference technical path for complexity control. In terms of real-time performance, most studies present a processing delay of less than 50 ms, which is largely attributed to the inherent characteristics of data-level fusion due to operating directly at the raw data layer, thus avoiding the time overhead associated with higher-level processing, such as feature extraction and decision reasoning.</p>
    <p>Zheng et al. <xref ref-type="bibr" rid="scirp.143380-1">
      [1]
     </xref> adopted a multi-branch architecture (base branch, aligned branch, affine estimated branch), which has a higher complexity, resulting in a surge in memory and computation, and its multi-task joint optimization (recognition, alignment, and feature learning) further increases the training complexity. In addition, due to high-resolution feature map processing and multi-branch parallel computing, PAN has higher hardware requirements, resulting in large inference delays. The model proposed by Zheng et al. <xref ref-type="bibr" rid="scirp.143380-2">
      [2]
     </xref> fused the comparison loss of the verification model (based on Euclidean distance) and the cross-entropy loss of the recognition model, although it is manifested in the complexity of dual-objective optimization and high-dimensional embedding calculation. The transfer learning of the ImageNet pre-trained model (such as ResNet) reduces the training cost. In general, the model proposed by Zheng et al. <xref ref-type="bibr" rid="scirp.143380-2">
      [2]
     </xref> has made a breakthrough in complexity and real-time, which can be used as a good reference for similar processing of complexity and real-time in this project to achieve more accurate pedestrian re-identification.</p>
    <p>Although the constructed models of the two are different, both have improved the accuracy of pedestrian re-identification, providing certain insights and guidance for the fusion re-identification system.</p>
   </sec>
  </sec><sec id="s5">
   <title>5. Typical Applications of Pedestrian Re-Identification Methods Using Millimeter-Wave Radar and Visual Fusion</title>
   <p>Multi-modal perception technology utilizing millimeter-wave (mmWave) radar has increasingly become a significant focus and an important field of study <xref ref-type="bibr" rid="scirp.143380-32">
     [32]
    </xref>, which supports critical applications in autonomous driving, smart cities, and surveillance. Object detection and tracking based on radar-camera fusion have also gained growing attention <xref ref-type="bibr" rid="scirp.143380-33">
     [33]
    </xref>. We investigate these references, hoping to make a greater contribution to research on fusion-based pedestrian re-identification.</p>
   <sec id="s5_1">
    <title>5.1. Nighttime Pedestrian Detection</title>
    <p>Nighttime or low-light conditions challenge visual sensors, but radar’s penetration ability ensures reliability. Zhao et al. <xref ref-type="bibr" rid="scirp.143380-29">
      [29]
     </xref> proposed a decision-level fusion framework using infrared vision and radar, achieving high accuracy in dark environments. Similarly, Zhang et al. <xref ref-type="bibr" rid="scirp.143380-11">
      [11]
     </xref> combined thermal imaging with radar for nighttime ReID, improving robustness.</p>
   </sec>
   <sec id="s5_2">
    <title>5.2. Pedestrian Tracking</title>
    <p>Real-time pedestrian tracking supports path planning and collision avoidance. Yang et al. <xref ref-type="bibr" rid="scirp.143380-27">
      [27]
     </xref> developed a decision-level fusion method using UKF and DeepSORT, ensuring precise trajectory tracking. Zhu et al. <xref ref-type="bibr" rid="scirp.143380-31">
      [31]
     </xref> introduced track-to-track fusion for multi-pedestrian scenarios, handling occlusions effectively. These methods enable vehicles to anticipate pedestrian movements.</p>
   </sec>
   <sec id="s5_3">
    <title>5.3. Impact and Future Directions</title>
    <p>Fusion-based ReID methods significantly enhance safety and reliability in autonomous systems. Studies like <xref ref-type="bibr" rid="scirp.143380-28">
      [28]
     </xref> demonstrate reduced false positives compared to single-sensor approaches. Future advancements may leverage datasets and focus on real-time processing and adverse weather performance.</p>
   </sec>
  </sec><sec id="s6">
   <title>6. Key Technical Challenges and Future Research Directions</title>
   <p>Despite the promising potential of fusing mmWave radar and visual data for robust pedestrian ReID, several significant technical challenges hinder its widespread adoption and performance optimization. Addressing these challenges constitutes key future research directions.</p>
   <sec id="s6_1">
    <title>6.1. Key Technical Challenges</title>
    <p>1) Lack of Dedicated Benchmarks: The absence of large-scale, diverse, and publicly available datasets specifically designed for mmWave-visual pedestrian ReID (with ground-truth IDs across non-overlapping views under various conditions) is arguably the biggest obstacle. This impedes standardized evaluation, fair comparison of methods, and the training of data-hungry deep learning models.</p>
    <p>2) Data Heterogeneity and Representation: Effectively fusing the sparse, geometric, and point cloud data from radar with the dense, semantic, and appearance-rich pixel data from cameras remains fundamentally challenging. Finding optimal representations for radar data that facilitate effective fusion with visual features is crucial.</p>
    <p>3) Radar Data Quality and Interpretation: mmWave radar data can be noisy, suffer from multipath reflections (ghost targets), and have low angular resolution compared to cameras. Sparsity makes it difficult to infer detailed shapes or associate points reliably to specific body parts for fine-grained gait analysis or ReID, especially in crowds. Extracting discriminative features solely from sparse radar points is non-trivial.</p>
    <p>4) Complexity vs. Real-Time Performance: Sophisticated hybrid fusion models (e.g. using transformers or complex attention mechanisms) often achieve better performance but come with high computational costs, making real-time deployment on resource-constrained platforms (like robots or edge devices) challenging.</p>
    <p>5) Generalization and Domain Adaptation: Models trained on data from one specific sensor setup, environment, or weather condition may not generalize well to others (domain shift). Variability in radar hardware, camera types, environmental clutter, and pedestrian densities poses significant generalization challenges.</p>
   </sec>
   <sec id="s6_2">
    <title>6.2. Future Research Directions</title>
    <p>1) Benchmark Dataset Development: Creating large-scale, diverse mmWave-visual pedestrian ReID datasets covering various environments (indoor/outdoor), weather conditions, pedestrian densities, and sensor viewpoints is paramount. Including annotations for persistent IDs across non-overlapping views is essential.</p>
    <p>2) Advanced Fusion Architectures: Exploring novel deep learning architectures tailored for heterogeneous sensor fusion. This includes investigating more advanced transformer variants, graph neural networks specifically designed for radar point cloud structure and radar-visual interaction, and perhaps integrating neural rendering techniques to bridge the modality gap.</p>
    <p>3) Exploiting Richer Radar Information: Moving beyond basic point clouds (x, y, z, v). Research into effectively incorporating radar micro-Doppler signatures for gait recognition utilizing the full Range-Azimuth-Elevation-Doppler radar tensor, or learning discriminative features directly from raw radar ADC data holds significant promise.</p>
    <p>4) Real-Time Optimization: Studying model compression, quantization, knowledge distillation, and efficient network architectures to enable real-time execution of complex fusion models on edge devices.</p>
    <p>5) Cross-Modal Adaptation: Developing geometry-aware fusion models with 3D pose estimation to overcome the core challenge of view-invariant matching in non-overlapping sensor configurations, while addressing domain shift through adversarial feature representation.</p>
    <p>Overcoming these challenges and pursuing these research directions will be crucial for unlocking the full potential of mmWave-visual fusion and realizing truly robust and reliable pedestrian ReID systems for real-world applications.</p>
   </sec>
  </sec><sec id="s7">
   <title>7. Conclusions</title>
   <p>Pedestrian re-identification is a critical technology for intelligent systems, but traditional visual methods struggle in challenging real-world conditions. This review has surveyed the emerging field of mmWave-visual fusion for pedestrian ReID, motivated by the complementary strengths of cameras and mmWave radar.</p>
   <p>We began by outlining the fundamental concepts of ReID, multi-modal ReID, and the characteristics of the visual and mmWave modalities, establishing the strong rationale for their fusion. We then discussed the essential practical aspects of data acquisition, including sensor setup, synchronization, calibration, and pre-processing techniques crucial for successful integration. The core of the review provided a systematic classification and analysis of fusion methodologies, tracing their evolution from early and late fusion approaches to the currently dominant intermediate/hybrid fusion strategies, particularly those leveraging deep learning, attention mechanisms, and graph networks. We highlighted the importance of experimental validation and explained the lack of dedicated ReID benchmarks. Finally, we identified key technical challenges, including data scarcity, heterogeneity, radar limitations, complexity, and generalization issues. Based on these challenges, we proposed promising future research directions, emphasizing benchmark creation, advanced fusion architectures, richer radar feature utilization, self-supervised learning, and real-time optimization.</p>
   <p>In conclusion, fusing mmWave radar and visual data offers a compelling pathway towards achieving robust, all-weather, and reliable pedestrian re-identification systems. While significant challenges remain, the ongoing advancements in sensor technology, deep learning, and multi-modal fusion techniques promise substantial progress in this important research area, paving the way for more capable perception systems in surveillance, robotics, and autonomous driving.</p>
  </sec><sec id="s8">
   <title>Acknowledgements</title>
   <p>The work is funded by the foundation of the Innovation and Entrepreneurship Training Program for College Students (202410424057).</p>
  </sec>
 </body><back>
  <ref-list>
   <title>References</title>
   <ref id="scirp.143380-ref1">
    <label>1</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Zheng, Z., Zheng, L. and Yang, Y. (2019) Pedestrian Alignment Network for Large-Scale Person Re-Identification. IEEE Transactions on Circuits and Systems for Video Technology, 29, 3037-3045. &gt;https://doi.org/10.1109/tcsvt.2018.2873599
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref2">
    <label>2</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Zheng, Z., Zheng, L. and Yang, Y. (2017) A Discriminatively Learned CNN Embedding for Person Reidentification. ACM Transactions on Multimedia Computing, Communications, and Applications, 14, 1-20. &gt;https://doi.org/10.1145/3159171
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref3">
    <label>3</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ye, M., Shen, J., Lin, G., Xiang, T., Shao, L. and Hoi, S.C.H. (2022) Deep Learning for Person Re-Identification: A Survey and Outlook. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44, 2872-2893. &gt;https://doi.org/10.1109/tpami.2021.3054775
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref4">
    <label>4</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Luo, H., Jiang, W., Gu, Y., Liu, F., Liao, X., Lai, S., et al. (2020) A Strong Baseline and Batch Normalization Neck for Deep Person Re-Identification. IEEE Transactions on Multimedia, 22, 2597-2609. &gt;https://doi.org/10.1109/tmm.2019.2958756
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref5">
    <label>5</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Uddin, M.K., Bhuiyan, A., Bappee, F.K., Islam, M.M. and Hasan, M. (2023) Person Re-Identification with RGB-D and RGB-IR Sensors: A Comprehensive Survey. Sensors, 23, Article No. 1504. &gt;https://doi.org/10.3390/s23031504
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref6">
    <label>6</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Bartsch, A., Fitzek, F. and Rasshofer, R.H. (2012) Pedestrian Recognition Using Automotive Radar Sensors. Advances in Radio Science, 10, 45-55. &gt;https://doi.org/10.5194/ars-10-45-2012
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref7">
    <label>7</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Cui, F., Zhang, Q., Wu, J., Song, Y., Xie, Z., Song, C., et al. (2023) Online Multipedestrian Tracking Based on Fused Detections of Millimeter Wave Radar and Vision. IEEE Sensors Journal, 23, 15702-15712. &gt;https://doi.org/10.1109/jsen.2023.3255924 
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref8">
    <label>8</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Gray, D. and Tao, H. (2008) Viewpoint Invariant Pedestrian Recognition with an Ensemble of Localized Features. 10th European Conference on Computer Vision, Marseille, 12-18 October 2008, 262-275. &gt;https://doi.org/10.1007/978-3-540-88682-2_21
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref9">
    <label>9</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ye, M., Wang, Z., Lan, X. and Yuen, P.C. (2018) Visible Thermal Person Re-Identification via Dual-Constrained Top-Ranking. Proceedings of the 27th International Joint Conference on Artificial Intelligence, Stockholm, 13-19 July 2018, 1092-1099. &gt;https://doi.org/10.24963/ijcai.2018/152
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref10">
    <label>10</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Yao, S., Guan, R., Huang, X., Li, Z., Sha, X., Yue, Y., et al. (2024) Radar-Camera Fusion for Object Detection and Semantic Segmentation in Autonomous Driving: A Comprehensive Review. IEEE Transactions on Intelligent Vehicles, 9, 2094-2128. &gt;https://doi.org/10.1109/tiv.2023.3307157
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref11">
    <label>11</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Zhang, Z. (2000) A Flexible New Technique for Camera Calibration. IEEE Transactions on Pattern Analysis and Machine Intelligence, 22, 1330-1334. &gt;https://doi.org/10.1109/34.888718
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref12">
    <label>12</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Oh, J., Kim, K., Park, M. and Kim, S. (2018) A Comparative Study on Camera-Radar Calibration Methods. 2018 15th International Conference on Control, Automation, Robotics and Vision (ICARCV), Singapore, 18-21 November 2018, 1057-1062. &gt;https://doi.org/10.1109/icarcv.2018.8581329
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref13">
    <label>13</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Liu, X., Deng, Z. and Zhang, G. (2025) Targetless Radar-Camera Extrinsic Parameter Calibration Using Track-to-Track Association. Sensors, 25, Article No. 949. &gt;https://doi.org/10.3390/s25030949
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref14">
    <label>14</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Redmon, J. and Farhadi, A. (2018) YOLOv3: An Incremental Improvement. arXiv: 1804.02767. &gt;https://arxiv.org/abs/1804.02767
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref15">
    <label>15</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ren, S., He, K., Girshick, R. and Sun, J. (2017) Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39, 1137-1149. &gt;https://doi.org/10.1109/tpami.2016.2577031
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref16">
    <label>16</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ester, M., Kriegel, H.-P., Sander, J. and Xu, X. (1996) A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise. Proceedings of the Second International Conference on Knowledge Discovery and Data Mining, Portland, 2-4 August 1996, 226-231. &gt;https://dl.acm.org/doi/10.5555/3001460.3001507
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref17">
    <label>17</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kalman, R.E. (1960) A New Approach to Linear Filtering and Prediction Problems. Journal of Basic Engineering, 82, 35-45. &gt;https://doi.org/10.1115/1.3662552
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref18">
    <label>18</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Milch, S. and Behrens, M. (2001) Pedestrian Detection with Radar and Computer Vision. Proceedings of PAL 2001-Progress in Automobile Lighting, Held Laboratory of Lighting Technology, Vol. 9, 657-664.
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref19">
    <label>19</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Guo, X., Du, J., Gao, J. and Wang, W. (2018) Pedestrian Detection Based on Fusion of Millimeter Wave Radar and Vision. Proceedings of the 2018 International Conference on Artificial Intelligence and Pattern Recognition, Beijing, 18-20 August 2018, 38-42. &gt;https://doi.org/10.1145/3268866.3268868
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref20">
    <label>20</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Wang, T., Zheng, N., Xin, J. and Ma, Z. (2011) Integrating Millimeter Wave Radar with a Monocular Vision Sensor for On-Road Obstacle Detection Applications. Sensors, 11, 8992-9008. &gt;https://doi.org/10.3390/s110908992
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref21">
    <label>21</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Streubel, R. and Yang, B. (2016) Fusion of Stereo Camera and MIMO-FMCW Radar for Pedestrian Tracking in Indoor Environments. 2016 19th International Conference on Information Fusion (FUSION), Heidelberg, 5-8 July 2016, 565-572. &gt;https://ieeexplore.ieee.org/document/7527938
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref22">
    <label>22</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Plascencia, A.C., García-Gómez, P., Perez, E.B., DeMas-Giménez, G., Casas, J.R. and Royo, S. (2023) A Preliminary Study of Deep Learning Sensor Fusion for Pedestrian Detection. Sensors, 23, Article No. 4167. &gt;https://doi.org/10.3390/s23084167
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref23">
    <label>23</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Nobis, F., Geisslinger, M., Weber, M., Betz, J. and Lienkamp, M. (2019) A Deep Learning-Based Radar and Camera Sensor Fusion Architecture for Object Detection. 2019 Sensor Data Fusion: Trends, Solutions, Applications (SDF), Bonn, 15-17 October 2019, 1-7. &gt;https://doi.org/10.1109/sdf.2019.8916629
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref24">
    <label>24</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Li, H., Liu, R., Wang, S., Jiang, W. and Lu, C.X. (2022) Pedestrian Liveness Detection Based on mmWave Radar and Camera Fusion. 2022 19th Annual IEEE International Conference on Sensing, Communication, and Networking (SECON), Stockholm, 20-23 September 2022, 262-270. &gt;https://doi.org/10.1109/secon55815.2022.9918553
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref25">
    <label>25</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Yu, X., Hu, T. and Zhu, H. (2025) Roadside Perception Applications Based on DCAM Fusion and Lightweight Millimeter-Wave Radar-Vision Integration. Electronics, 14, Article No. 1576. &gt;https://doi.org/10.3390/electronics14081576
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref26">
    <label>26</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Liu, R., Yao, T., Shi, R., Mei, L., Wang, S., Yin, Z., et al. (2024) Mission: mmWave Radar Person Identification with RGB Cameras. Proceedings of the 22nd ACM Conference on Embedded Networked Sensor Systems, Hangzhou, 4-7 November 2024, 309-321. &gt;https://doi.org/10.1145/3666025.3699340
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref27">
    <label>27</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Yang, C., Huan, S., Wu, L., Weng, Q. and Xiong, W. (2023) Fusion of Millimeter-Wave Radar and Camera Vision for Pedestrian Tracking. 2023 5th International Conference on Communications, Information System and Computer Engineering (CISCE), Guangzhou, 14-16 April 2023, 317-321. &gt;https://doi.org/10.1109/cisce58541.2023.10142444
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref28">
    <label>28</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Cui, F., Song, Y., Wu, J., Xie, Z., Song, C., Xu, Z., et al. (2021) Online Multi-Target Tracking for Pedestrian by Fusion of Millimeter Wave Radar and Vision. 2021 IEEE Radar Conference (RadarConf21), Atlanta, 7-14 May 2021, 1-6. &gt;https://doi.org/10.1109/radarconf2147009.2021.9455185
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref29">
    <label>29</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Zhao, W., Wang, T., Tan, A. and Ren, C. (2023) Nighttime Pedestrian Detection Based on a Fusion of Visual Information and Millimeter-Wave Radar. IEEE Access, 11, 68439-68451. &gt;https://doi.org/10.1109/access.2023.3291398
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref30">
    <label>30</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Graves, K., Kanwal, M., Yu, X. and Saniie, J. (2022) Design Flow of mmWave Radar and Machine Vision Fusion for Pedestrian Collision Warning. 2022 IEEE International Conference on Electro Information Technology (eIT), Mankato, 19-21 May 2022, 176-181. &gt;https://doi.org/10.1109/eit53891.2022.9813942
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref31">
    <label>31</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Zhu, Y., Wang, T. and Zhu, S. (2022) Adaptive Multi-Pedestrian Tracking by Multi-Sensor: Track-to-Track Fusion Using Monocular 3D Detection and MMW Radar. Remote Sensing, 14, Article No. 1837. &gt;https://doi.org/10.3390/rs14081837
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref32">
    <label>32</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Wang, S., Mei, L., Liu, R., Jiang, W., Yin, Z., Deng, X., et al. (2025) Multi-Modal Fusion Sensing: A Comprehensive Review of Millimeter-Wave Radar and Its Integration with Other Modalities. IEEE Communications Surveys&amp;Tutorials, 27, 322-352. &gt;https://doi.org/10.1109/comst.2024.3398004
    </mixed-citation>
   </ref>
   <ref id="scirp.143380-ref33">
    <label>33</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Shi, K., et al. (2024) Radar and Camera Fusion for Object Detection and Tracking: A Comprehensive Survey. arXiv: 2410.19872. &gt;https://doi.org/10.48550/arXiv.2410.19872
    </mixed-citation>
   </ref>
  </ref-list>
 </back>
</article>