<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article">
 <front>
  <journal-meta>
   <journal-id journal-id-type="publisher-id">
    ojapps
   </journal-id>
   <journal-title-group>
    <journal-title>
     Open Journal of Applied Sciences
    </journal-title>
   </journal-title-group>
   <issn pub-type="epub">
    2165-3917
   </issn>
   <issn publication-format="print">
    2165-3925
   </issn>
   <publisher>
    <publisher-name>
     Scientific Research Publishing
    </publisher-name>
   </publisher>
  </journal-meta>
  <article-meta>
   <article-id pub-id-type="doi">
    10.4236/ojapps.2024.1412234
   </article-id>
   <article-id pub-id-type="publisher-id">
    ojapps-138267
   </article-id>
   <article-categories>
    <subj-group subj-group-type="heading">
     <subject>
      Articles
     </subject>
    </subj-group>
    <subj-group subj-group-type="Discipline-v2">
     <subject>
      Biomedical 
     </subject>
     <subject>
       Life Sciences, Chemistry 
     </subject>
     <subject>
       Materials Science, Computer Science 
     </subject>
     <subject>
       Communications, Engineering, Physics 
     </subject>
     <subject>
       Mathematics
     </subject>
    </subj-group>
   </article-categories>
   <title-group>
    Research on Vehicle Tracking Method Based on YOLOv8 and Adaptive Kalman Filtering: Integrating SVM Dynamic Selection and Error Feedback Mechanism
   </title-group>
   <contrib-group>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Liping
      </surname>
      <given-names>
       Zheng
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff1"> 
      <sup>1</sup>
     </xref>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Hao
      </surname>
      <given-names>
       Gou
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff2"> 
      <sup>2</sup>
     </xref>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Kaiwen
      </surname>
      <given-names>
       Xiao
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff1"> 
      <sup>1</sup>
     </xref>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Moran
      </surname>
      <given-names>
       Qiu
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff1"> 
      <sup>1</sup>
     </xref>
    </contrib>
   </contrib-group> 
   <aff id="aff1">
    <addr-line>
     aSchool of Computer Science, Sichuan University Jinjiang College, Meishan, China
    </addr-line> 
   </aff> 
   <aff id="aff2">
    <addr-line>
     aSchool of Automotive and Transportation, Xihua University, Yibin, China
    </addr-line> 
   </aff> 
   <pub-date pub-type="epub">
    <day>
     03
    </day> 
    <month>
     12
    </month>
    <year>
     2024
    </year>
   </pub-date> 
   <volume>
    14
   </volume> 
   <issue>
    12
   </issue>
   <fpage>
    3569
   </fpage>
   <lpage>
    3588
   </lpage>
   <history>
    <date date-type="received">
     <day>
      26,
     </day>
     <month>
      November
     </month>
     <year>
      2024
     </year>
    </date>
    <date date-type="published">
     <day>
      16,
     </day>
     <month>
      November
     </month>
     <year>
      2024
     </year> 
    </date> 
    <date date-type="accepted">
     <day>
      16,
     </day>
     <month>
      December
     </month>
     <year>
      2024
     </year> 
    </date>
   </history>
   <permissions>
    <copyright-statement>
     © Copyright 2014 by authors and Scientific Research Publishing Inc. 
    </copyright-statement>
    <copyright-year>
     2014
    </copyright-year>
    <license>
     <license-p>
      This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/
     </license-p>
    </license>
   </permissions>
   <abstract>
    Vehicle tracking plays a crucial role in intelligent transportation, autonomous driving, and video surveillance. However, challenges such as occlusion, multi-target interference, and nonlinear motion in dynamic scenarios make tracking accuracy and stability a focus of ongoing research. This paper proposes an integrated method combining YOLOv8 object detection with adaptive Kalman filtering. The approach employs a support vector machine (SVM) to dynamically select the optimal filter (including standard Kalman filter, extended Kalman filter, and unscented Kalman filter), enhancing the system’s adaptability to different motion patterns. Additionally, an error feedback mechanism is incorporated to dynamically adjust filter parameters, further improving responsiveness to sudden events. Experimental results on the KITTI and UA-DETRAC datasets demonstrate that the proposed method significantly improves detection accuracy (mAP@0.5 increased by approximately 3%), tracking accuracy (MOTA improved by 5%), and system robustness, providing an efficient solution for vehicle tracking in complex environments.
   </abstract>
   <kwd-group> 
    <kwd>
     Multi-Target Tracking
    </kwd> 
    <kwd>
      YOLOv8-Based Detection
    </kwd> 
    <kwd>
      Adaptive Filtering
    </kwd> 
    <kwd>
      Support Vector Machine
    </kwd> 
    <kwd>
      Error Feedback Mechanism
    </kwd>
   </kwd-group>
  </article-meta>
 </front>
 <body>
  <sec id="s1">
   <title>1. Introduction</title>
   <p>Vehicle tracking technology finds extensive applications in intelligent transportation systems, autonomous driving, and video surveillance <xref ref-type="bibr" rid="scirp.138267-1">
     [1]
    </xref>. Its core task is to precisely localize and continuously track target vehicles, supporting key functions such as path planning, traffic flow analysis, and behavior prediction <xref ref-type="bibr" rid="scirp.138267-2">
     [2]
    </xref>.</p>
   <p>In autonomous driving, accurate vehicle tracking enables vehicles to identify dynamic obstacles in complex environments, facilitating safe and efficient decision-making. In traffic surveillance, it aids in detecting traffic violations, monitoring congestion, and improving road efficiency.</p>
   <p>However, the highly dynamic and unpredictable nature of real-world environments presents numerous challenges for vehicle tracking:</p>
   <p>1) Occlusion Problem</p>
   <p>Occlusion is a common issue in vehicle tracking within multi-vehicle scenarios <xref ref-type="bibr" rid="scirp.138267-3">
     [3]
    </xref>. For instance, when a vehicle is partially or fully obscured by other vehicles, pedestrians, or static objects, the continuity of detection and tracking may be severely affected.</p>
   <p>2) Dynamic Scenarios and Complex Backgrounds</p>
   <p>In densely populated dynamic traffic scenes, vehicle motion can exhibit nonlinear behaviors such as acceleration, deceleration, and turning. Additionally, complex backgrounds—including multi-target interference, strong lighting, or shadows—can significantly reduce the accuracy of detection and tracking <xref ref-type="bibr" rid="scirp.138267-4">
     [4]
    </xref>.</p>
   <p>3) Error Accumulation</p>
   <p>In traditional tracking frameworks, discrepancies may arise between detection results and trajectory predictions by filters. When tracking persists over time or the dynamics of the scene intensify, these errors can accumulate, ultimately leading to tracking failures.</p>
   <p>To overcome these challenges, it is crucial to develop vehicle tracking methods that offer high accuracy, strong robustness, and dynamic adaptability. This paper proposes an integrated approach combining object detection with trajectory prediction. The method utilizes the deep learning-based object detection model YOLOv8 <xref ref-type="bibr" rid="scirp.138267-5">
     [5]
    </xref> and an adaptive filter selector to achieve efficient and robust vehicle tracking.</p>
   <p>To address the limitations of existing methods, this paper introduces two key innovations:</p>
   <p>1) SVM-Based Adaptive Filter Selection Mechanism</p>
   <p>Vehicle motion characteristics (e.g., constant velocity, acceleration, nonlinear movement) vary across different scenarios, making it challenging for a single filter to effectively handle diverse conditions. To address this, an adaptive filter selection mechanism based on Support Vector Machine (SVM) is designed. This mechanism dynamically selects the optimal filter—standard Kalman Filter (KF), Extended Kalman Filter (EKF), or Unscented Kalman Filter (UKF)—based on the current motion characteristics of the vehicle. This approach significantly enhances tracking accuracy and robustness in complex and dynamic environments.</p>
   <p>2) Integration of Detection and Tracking Error Feedback Mechanism</p>
   <p>To minimize discrepancies between detection and prediction, an error feedback mechanism is proposed. This mechanism dynamically feeds the error between YOLOv8 detection results and filter predictions back to the filter parameter adjustment module. By adaptively tuning the process noise covariance and measurement noise covariance, the filter can better handle sudden events (e.g., abrupt acceleration or sharp turns), thereby improving the consistency and stability of tracking.</p>
  </sec><sec id="s2">
   <title>2. Related Work</title>
   <p>Vehicle tracking technology is a key research area in intelligent transportation, autonomous driving, and security fields. Its main objective is to accurately locate the target vehicle and predict its motion trajectory. Currently, vehicle tracking methods are mainly divided into two categories: deep learning-based object detection methods <xref ref-type="bibr" rid="scirp.138267-6">
     [6]
    </xref> and classical filtering-based trajectory prediction methods <xref ref-type="bibr" rid="scirp.138267-7">
     [7]
    </xref>. This section reviews these two approaches and discusses their advantages and limitations based on the latest research.</p>
   <sec id="s2_1">
    <title>2.1. Deep Learning-Based Object Detection Methods</title>
    <p>With the development of deep learning, vehicle tracking methods based on object detection have gradually become mainstream. Object detectors generate candidate bounding boxes by detecting each video frame, and vehicle trajectories are formed using subsequent association strategies.</p>
    <p>1) YOLO Series Models</p>
    <p>The YOLO (You Only Look Once) model, proposed by Redmon et al. <xref ref-type="bibr" rid="scirp.138267-8">
      [8]
     </xref>, pioneered single-stage object detection. The current popular version, YOLOv8, optimizes the network structure and feature fusion, demonstrating excellent accuracy and speed in vehicle detection tasks <xref ref-type="bibr" rid="scirp.138267-9">
      [9]
     </xref> <xref ref-type="bibr" rid="scirp.138267-10">
      [10]
     </xref>.</p>
    <p>The advantage of the YOLOv8 model is its ability to achieve real-time vehicle detection with a high frame rate, performing well even in occlusion and low-light conditions. However, its detection results are highly dependent on the diversity of the training data, and there is limited coordination with subsequent tracking modules.</p>
    <p>2) Faster R-CNN</p>
    <p>Faster R-CNN, introduced by Ren et al. <xref ref-type="bibr" rid="scirp.138267-11">
      [11]
     </xref>, significantly improves detection accuracy by incorporating a Region Proposal Network (RPN). Zhang Ying et al. <xref ref-type="bibr" rid="scirp.138267-12">
      [12]
     </xref> designed a vehicle detection solution based on the Faster R-CNN model for unmanned aerial vehicle (UAV) platforms. Ouyang Bo et al. <xref ref-type="bibr" rid="scirp.138267-13">
      [13]
     </xref> proposed a lightweight tracking model, FA-SORT, which achieved a tracking speed of 29.93 FPS when tested on the UAVDT dataset.</p>
    <p>The advantage of Faster R-CNN lies in its high accuracy, making it particularly suitable for complex backgrounds and dense target scenarios. However, its computational complexity is relatively high, making it difficult to meet real-time processing requirements.</p>
   </sec>
   <sec id="s2_2">
    <title>2.2. Kalman Filtering and Its Variants</title>
    <p>Classic filtering methods, represented by the Kalman filter, estimate and update vehicle trajectories by combining detection results with predicted states. These methods are typically used in conjunction with object detectors for multi-target tracking tasks.</p>
    <p>1) Kalman Filter (KF)</p>
    <p>The Kalman Filter (KF) is an efficient linear state estimation method widely used for vehicle tracking in scenarios involving constant or linear motion. Gordon et al. <xref ref-type="bibr" rid="scirp.138267-14">
      [14]
     </xref> applied KF in real-time traffic flow monitoring, significantly improving the accuracy of vehicle trajectory predictions. However, KF performs poorly when dealing with nonlinear or complex motion.</p>
    <p>2) Extended Kalman Filter (EKF)</p>
    <p>The Extended Kalman Filter (EKF) extends the KF to nonlinear systems by performing a first-order linearization, making it suitable for a wider range of applications. Reid et al. <xref ref-type="bibr" rid="scirp.138267-15">
      [15]
     </xref> applied EKF to multi-target tracking, achieving high accuracy in vehicle scenarios involving rapid acceleration and sharp turns. However, its performance in highly nonlinear scenarios is still limited by linearization errors.</p>
    <p>3) Unscented Kalman Filter (UKF)</p>
    <p>The Unscented Kalman Filter (UKF) improves trajectory prediction accuracy in nonlinear systems by using sigma-point sampling, eliminating the need for linearization. The UKF algorithm, proposed by van der Merwe et al. <xref ref-type="bibr" rid="scirp.138267-16">
      [16]
     </xref>, demonstrates excellent performance in dynamic target tracking. Farag et al. <xref ref-type="bibr" rid="scirp.138267-17">
      [17]
     </xref> combined UKF with deep detectors to perform multi-target tracking in complex traffic environments.</p>
   </sec>
   <sec id="s2_3">
    <title>2.3. Integrated Approaches Combining Deep Learning and Classical Filtering</title>
    <p>In recent years, researchers have explored the deep integration of object detection and filtering techniques to optimize vehicle tracking performance. Bui T et al. <xref ref-type="bibr" rid="scirp.138267-18">
      [18]
     </xref> proposed a tracking framework that combines YOLOv5 with an optimized Kalman filter (KF) algorithm in DeepSORT, achieving a 15.5% improvement in mAP and a 14.2% increase in MOTA. Gao J et al. <xref ref-type="bibr" rid="scirp.138267-19">
      [19]
     </xref> reduced the impact of observation noise on detection accuracy during nonlinear motion by incorporating a lightweight channel block attention mechanism (LCBAM) and noise-adaptive Extended Kalman Filter (NSA-EKF). These methods demonstrate the complementarity of deep detectors and classical filtering techniques in vehicle tracking tasks, making them a current research hotspot.</p>
    <p>Deep learning-based object detection methods offer efficient target recognition capabilities, while classical filtering methods provide irreplaceable advantages in state estimation and trajectory prediction. Integrating the strengths of both approaches into a unified tracking framework is an effective solution for vehicle tracking in complex dynamic environments.</p>
   </sec>
  </sec><sec id="s3">
   <title>3. Methodology</title>
   <sec id="s3_1">
    <title>3.1. YOLOv8 Detection Module</title>
    <p>In this study, YOLOv8 is responsible for the initial vehicle detection, providing high-quality target information for the subsequent tracking module. This information is then used to enhance the overall tracking performance through Kalman filtering and error feedback mechanisms. The following section outlines the configuration and structural optimizations of the YOLOv8 model used in this research.</p>
    <p>1) Overall Framework Configuration</p>
    <p>In this study, YOLOv8 still uses CSPDarknet as the backbone network, incorporating a bottleneck structure (Bottleneck CSP) to reduce redundant computations and lower model complexity, thereby improving detection speed while maintaining accuracy. Additionally, multi-scale feature fusion modules (FPN and PAN) are introduced in the structure to enhance the model’s performance in detecting vehicles of varying sizes. Furthermore, the application of depthwise separable convolutions significantly reduces the computational cost of YOLOv8, enabling real-time processing. Finally, to improve detection accuracy, a reinforced data augmentation strategy is employed during YOLOv8 training, including random scaling, translation, flipping, and color adjustment, which enhances the model’s robustness and generalization ability. The overall framework of YOLOv8 is shown in <xref ref-type="fig" rid="fig1">
      Figure 1
     </xref>.</p>
    <fig id="fig1" position="float">
     <label>Figure 1</label>
     <caption>
      <title>Figure 1. YOLOv8 structural frame diagram.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/2312887-rId16.jpeg?20241219021041" />
    </fig>
    <p>2) Input and Output Formats</p>
    <p>The input format for the YOLOv8 detection module is a standard RGB image, which is normalized and resized to a specific dimension (e.g., 640 × 640) before being fed into the model. This ensures consistent detection accuracy across images with different resolutions. During the input process, YOLOv8 also incorporates anchor box adaptive adjustment techniques, allowing the model to more flexibly accommodate vehicles of various sizes and perspectives.</p>
    <p>The output format of YOLOv8 includes the coordinates of the predicted bounding boxes, target categories, and confidence scores. For vehicle detection tasks, the model outputs the vehicle’s bounding box positions (x, y, width, height) along with the probability distribution for the corresponding category (e.g., car, truck, motorcycle), where the confidence score is used to filter out low-confidence targets. After post-processing with Non-Maximum Suppression (NMS), overlapping redundant boxes are eliminated to ensure the accuracy of the output results. In this study, the detection boxes output by YOLOv8 are directly fed into the Kalman filter module for further trajectory prediction and state estimation.</p>
   </sec>
   <sec id="s3_2">
    <title>3.2. Kalman Filtering and Its Improvements</title>
    <p>In vehicle tracking tasks, the Kalman filter is a commonly used state estimation method that combines observation and prediction to estimate the object’s position and motion state. While the standard Kalman filter (KF) performs excellently in linear systems with Gaussian noise, it has limitations in dynamic, nonlinear, and non-Gaussian environments. Therefore, this study introduces several variants of the Kalman filter, including the Extended Kalman Filter (EKF), Unscented Kalman Filter (UKF), and Interactive Multiple Model (IMM) filter. Additionally, an adaptive filter selection mechanism based on Support Vector Machine (SVM) is proposed to dynamically select the optimal filter in different scenarios, thus optimizing detection and tracking performance.</p>
    <p>1) Standard Kalman Filter (KF)</p>
    <p>The standard Kalman filter assumes that both the system state model and the measurement model are linear and that the noise follows a Gaussian distribution. The state estimation process consists of two stages: prediction and update.</p>
    <p>The state prediction equations are given by the following Equation (1) and (2).</p>
    <p>
     <xref ref-type="bibr" rid="scirp.138267-"></xref> 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mover accent="true"> 
         <mi>
           x 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           | 
         </mo> 
         <mi>
           k 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </msub> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mi>
          F 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <msub> 
        <mover accent="true"> 
         <mi>
           x 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
         <mo>
           | 
         </mo> 
         <mi>
           k 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </msub> 
       <mo>
         + 
       </mo> 
       <msub> 
        <mi>
          B 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <msub> 
        <mi>
          u 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math> (1)</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          P 
        </mi> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           | 
         </mo> 
         <mi>
           k 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </msub> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mi>
          F 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <msub> 
        <mi>
          P 
        </mi> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
         <mo>
           | 
         </mo> 
         <mi>
           k 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </msub> 
       <msubsup> 
        <mi>
          F 
        </mi> 
        <mi>
          k 
        </mi> 
        <mtext>
          T 
        </mtext> 
       </msubsup> 
       <mo>
         + 
       </mo> 
       <msub> 
        <mi>
          Q 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math> (2)</p>
    <p>The state update equations are given by the following Equation (3), (4) and (5).</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          K 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mi>
          P 
        </mi> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           | 
         </mo> 
         <mi>
           k 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </msub> 
       <msubsup> 
        <mi>
          H 
        </mi> 
        <mi>
          k 
        </mi> 
        <mtext>
          T 
        </mtext> 
       </msubsup> 
       <msup> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              H 
            </mi> 
            <mi>
              k 
            </mi> 
           </msub> 
           <msub> 
            <mi>
              P 
            </mi> 
            <mrow> 
             <mi>
               k 
             </mi> 
             <mo>
               | 
             </mo> 
             <mi>
               k 
             </mi> 
             <mo>
               − 
             </mo> 
             <mn>
               1 
             </mn> 
            </mrow> 
           </msub> 
           <msubsup> 
            <mi>
              H 
            </mi> 
            <mi>
              k 
            </mi> 
            <mtext>
              T 
            </mtext> 
           </msubsup> 
           <mo>
             + 
           </mo> 
           <msub> 
            <mi>
              R 
            </mi> 
            <mi>
              k 
            </mi> 
           </msub> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mrow> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </msup> 
      </mrow> 
     </math> (3)</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mover accent="true"> 
         <mi>
           x 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           | 
         </mo> 
         <mi>
           k 
         </mi> 
        </mrow> 
       </msub> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mover accent="true"> 
         <mi>
           x 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           | 
         </mo> 
         <mi>
           k 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </msub> 
       <mo>
         + 
       </mo> 
       <msub> 
        <mi>
          K 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            z 
          </mi> 
          <mi>
            k 
          </mi> 
         </msub> 
         <mo>
           − 
         </mo> 
         <msub> 
          <mi>
            H 
          </mi> 
          <mi>
            k 
          </mi> 
         </msub> 
         <msub> 
          <mover accent="true"> 
           <mi>
             x 
           </mi> 
           <mo>
             ^ 
           </mo> 
          </mover> 
          <mrow> 
           <mi>
             k 
           </mi> 
           <mo>
             | 
           </mo> 
           <mi>
             k 
           </mi> 
           <mo>
             − 
           </mo> 
           <mn>
             1 
           </mn> 
          </mrow> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> (4)</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          P 
        </mi> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           | 
         </mo> 
         <mi>
           k 
         </mi> 
        </mrow> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           I 
         </mi> 
         <mo>
           − 
         </mo> 
         <msub> 
          <mi>
            K 
          </mi> 
          <mi>
            k 
          </mi> 
         </msub> 
         <msub> 
          <mi>
            H 
          </mi> 
          <mi>
            k 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <msub> 
        <mi>
          P 
        </mi> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           | 
         </mo> 
         <mi>
           k 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </msub> 
      </mrow> 
     </math> (5)</p>
    <p>Wherein, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          F 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math> represents the state transition matrix, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          H 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math> is the measurement matrix, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          Q 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          R 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math> denote the process noise and measurement noise covariance matrices, respectively, and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          P 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math> is the estimation error covariance matrix.</p>
    <p>2) Extended Kalman Filter (EKF)</p>
    <p>The Extended Kalman Filter (EKF) extends the linear assumption of the standard Kalman Filter (KF) to accommodate mildly nonlinear systems. The EKF achieves approximate linearization by performing a first-order Taylor expansion on the nonlinear state transition function and observation function.</p>
    <p>The equations for the state prediction phase (nonlinear) are given by Equations (6) and (7) as follows.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mover accent="true"> 
         <mi>
           x 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           | 
         </mo> 
         <mi>
           k 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mover accent="true"> 
           <mi>
             x 
           </mi> 
           <mo>
             ^ 
           </mo> 
          </mover> 
          <mrow> 
           <mi>
             k 
           </mi> 
           <mo>
             − 
           </mo> 
           <mn>
             1 
           </mn> 
           <mo>
             | 
           </mo> 
           <mi>
             k 
           </mi> 
           <mo>
             − 
           </mo> 
           <mn>
             1 
           </mn> 
          </mrow> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         + 
       </mo> 
       <msub> 
        <mi>
          B 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <msub> 
        <mi>
          u 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math> (6)</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          P 
        </mi> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           | 
         </mo> 
         <mi>
           k 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </msub> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mi>
          F 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <msub> 
        <mi>
          P 
        </mi> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
         <mo>
           | 
         </mo> 
         <mi>
           k 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </msub> 
       <msubsup> 
        <mi>
          F 
        </mi> 
        <mi>
          k 
        </mi> 
        <mtext>
          T 
        </mtext> 
       </msubsup> 
       <mo>
         + 
       </mo> 
       <msub> 
        <mi>
          Q 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math> (7)</p>
    <p>The state update equations are given by the following Equation (8), (9) and (10).</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          K 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mi>
          P 
        </mi> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           | 
         </mo> 
         <mi>
           k 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </msub> 
       <msubsup> 
        <mi>
          H 
        </mi> 
        <mi>
          k 
        </mi> 
        <mtext>
          T 
        </mtext> 
       </msubsup> 
       <msup> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              H 
            </mi> 
            <mi>
              k 
            </mi> 
           </msub> 
           <msub> 
            <mi>
              P 
            </mi> 
            <mrow> 
             <mi>
               k 
             </mi> 
             <mo>
               | 
             </mo> 
             <mi>
               k 
             </mi> 
             <mo>
               − 
             </mo> 
             <mn>
               1 
             </mn> 
            </mrow> 
           </msub> 
           <msubsup> 
            <mi>
              H 
            </mi> 
            <mi>
              k 
            </mi> 
            <mtext>
              T 
            </mtext> 
           </msubsup> 
           <mo>
             + 
           </mo> 
           <msub> 
            <mi>
              R 
            </mi> 
            <mi>
              k 
            </mi> 
           </msub> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mrow> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </msup> 
      </mrow> 
     </math> (8)</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mover accent="true"> 
         <mi>
           x 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           | 
         </mo> 
         <mi>
           k 
         </mi> 
        </mrow> 
       </msub> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mover accent="true"> 
         <mi>
           x 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           | 
         </mo> 
         <mi>
           k 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </msub> 
       <mo>
         + 
       </mo> 
       <msub> 
        <mi>
          K 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            z 
          </mi> 
          <mi>
            k 
          </mi> 
         </msub> 
         <mo>
           − 
         </mo> 
         <mi>
           h 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mover accent="true"> 
             <mi>
               x 
             </mi> 
             <mo>
               ^ 
             </mo> 
            </mover> 
            <mrow> 
             <mi>
               k 
             </mi> 
             <mo>
               | 
             </mo> 
             <mi>
               k 
             </mi> 
             <mo>
               − 
             </mo> 
             <mn>
               1 
             </mn> 
            </mrow> 
           </msub> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> (9)</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          P 
        </mi> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           | 
         </mo> 
         <mi>
           k 
         </mi> 
        </mrow> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           I 
         </mi> 
         <mo>
           − 
         </mo> 
         <msub> 
          <mi>
            K 
          </mi> 
          <mi>
            k 
          </mi> 
         </msub> 
         <msub> 
          <mi>
            H 
          </mi> 
          <mi>
            k 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <msub> 
        <mi>
          P 
        </mi> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           | 
         </mo> 
         <mi>
           k 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </msub> 
      </mrow> 
     </math> (10)</p>
    <p>In this context, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mo>
          ⋅ 
        </mo> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         h 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mo>
          ⋅ 
        </mo> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> represent the nonlinear functions for state transition and observation, respectively.</p>
    <p>3) Unscented Kalman Filter (UKF)</p>
    <p>The Unscented Kalman Filter (UKF) addresses the state estimation problem in nonlinear systems through a set of weighted sample points known as “sigma points.” Without the need for linearization, the UKF achieves higher-precision state estimation under stronger nonlinear conditions.</p>
    <p>Sigma Point Generation and Propagation: A set of weighted sigma points is generated, and the propagation results of each sigma point are used to update the mean and covariance, thereby calculating the estimated state distribution.</p>
    <p>In complex vehicle tracking scenarios, a single filter struggles to cope with all situations due to the presence of different motion patterns (such as uniform speed, acceleration, nonlinear trajectories, etc.). Therefore, an adaptive filter selection mechanism based on Support Vector Machines (SVMs) is proposed. This mechanism utilizes SVMs to classify motion patterns in the input feature space, thereby selecting the most suitable Kalman filter variant according to the characteristics of the scene. This dynamic selection can significantly enhance the adaptability of the filter and ensure the accuracy and robustness of the tracking process.</p>
    <p>The design of the SVM-based adaptive filter selection mechanism can be divided into three main steps: feature extraction, SVM classifier training, and dynamic filter selection.</p>
    <p>1) Feature Extraction</p>
    <p>To achieve accurate filter selection, the system needs to extract features that can effectively distinguish between motion patterns. These features include the target’s velocity ( 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        v 
      </mi> 
     </math>), acceleration ( 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        a 
      </mi> 
     </math>), and the rate of change in its motion trajectory (such as curvature changes). The distributions of these features under different motion patterns are typically distinct, serving as effective bases for classification. Therefore, we define the feature vector as shown in Equation (11):</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mstyle mathvariant="bold" mathsize="normal"> 
        <mi>
          X 
        </mi> 
       </mstyle> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mo>
          [ 
        </mo> 
        <mrow> 
         <mi>
           v 
         </mi> 
         <mo>
           , 
         </mo> 
         <mi>
           a 
         </mi> 
         <mo>
           , 
         </mo> 
         <mi>
           θ 
         </mi> 
        </mrow> 
        <mo>
          ] 
        </mo> 
       </mrow> 
      </mrow> 
     </math> (11)</p>
    <p>Wherein, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        v 
      </mi> 
     </math> represents the instantaneous velocity of the target at the current moment; 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        a 
      </mi> 
     </math> denotes the acceleration of the target; and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        θ 
      </mi> 
     </math> indicates the change angle of the movement direction (i.e., steering angle or trajectory curvature).</p>
    <p>2) SVM Classifier</p>
    <p>Within the feature space, the feature vector X is input into the Support Vector Machine (SVM) classifier. The SVM utilizes the trained classification boundary to separate different motion patterns. Specifically, we train a multi-class SVM classifier that categorizes vehicle motion into three classes: uniform speed, acceleration, and complex nonlinear motion.</p>
    <p>Assuming the training data is denoted as 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         D 
       </mi> 
       <mo>
         = 
       </mo> 
       <msubsup> 
        <mrow> 
         <mrow> 
          <mo>
            { 
          </mo> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <msub> 
              <mstyle mathvariant="bold" mathsize="normal"> 
               <mi>
                 X 
               </mi> 
              </mstyle> 
              <mi>
                i 
              </mi> 
             </msub> 
             <mo>
               , 
             </mo> 
             <msub> 
              <mi>
                y 
              </mi> 
              <mi>
                i 
              </mi> 
             </msub> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mo>
            } 
          </mo> 
         </mrow> 
        </mrow> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mo>
           = 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mi>
          N 
        </mi> 
       </msubsup> 
      </mrow> 
     </math>, where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mstyle mathvariant="bold" mathsize="normal"> 
         <mi>
           X 
         </mi> 
        </mstyle> 
        <mi>
          i 
        </mi> 
       </msub> 
      </mrow> 
     </math> represents the input features and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          y 
        </mi> 
        <mi>
          i 
        </mi> 
       </msub> 
      </mrow> 
     </math> corresponds to the motion pattern labels (1 for uniform motion, 2 for accelerated motion, and 3 for nonlinear motion). The objective is to find an optimal classification boundary by solving the following optimization problem using SVM, as shown in Equations (12) and (13):</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <munder> 
        <mrow> 
         <mi>
           min 
         </mi> 
        </mrow> 
        <mrow> 
         <mstyle mathvariant="bold" mathsize="normal"> 
          <mi>
            w 
          </mi> 
         </mstyle> 
         <mo>
           , 
         </mo> 
         <mi>
           b 
         </mi> 
        </mrow> 
       </munder> 
       <mfrac> 
        <mn>
          1 
        </mn> 
        <mn>
          2 
        </mn> 
       </mfrac> 
       <msup> 
        <mrow> 
         <mrow> 
          <mo>
            ‖ 
          </mo> 
          <mstyle mathvariant="bold" mathsize="normal"> 
           <mi>
             w 
           </mi> 
          </mstyle> 
          <mo>
            ‖ 
          </mo> 
         </mrow> 
        </mrow> 
        <mn>
          2 
        </mn> 
       </msup> 
       <mo>
         + 
       </mo> 
       <mi>
         C 
       </mi> 
       <mstyle displaystyle="true"> 
        <msubsup> 
         <mo>
           ∑ 
         </mo> 
         <mrow> 
          <mi>
            i 
          </mi> 
          <mo>
            = 
          </mo> 
          <mn>
            1 
          </mn> 
         </mrow> 
         <mi>
           N 
         </mi> 
        </msubsup> 
        <mrow> 
         <msub> 
          <mi>
            ε 
          </mi> 
          <mi>
            i 
          </mi> 
         </msub> 
        </mrow> 
       </mstyle> 
      </mrow> 
     </math> (12)</p>
    <p>subject to 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          y 
        </mi> 
        <mi>
          i 
        </mi> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mstyle mathvariant="bold" mathsize="normal"> 
          <mi>
            w 
          </mi> 
         </mstyle> 
         <mo>
           ⋅ 
         </mo> 
         <mo>
           ∅ 
         </mo> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mstyle mathvariant="bold" mathsize="normal"> 
             <mi>
               X 
             </mi> 
            </mstyle> 
            <mi>
              i 
            </mi> 
           </msub> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           + 
         </mo> 
         <mi>
           b 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ≥ 
       </mo> 
       <mn>
         1 
       </mn> 
       <mo>
         − 
       </mo> 
       <msub> 
        <mi>
          ε 
        </mi> 
        <mi>
          i 
        </mi> 
       </msub> 
       <mo>
         , 
       </mo> 
       <mtext>
           
       </mtext> 
       <msub> 
        <mi>
          ε 
        </mi> 
        <mi>
          i 
        </mi> 
       </msub> 
       <mo>
         ≥ 
       </mo> 
       <mn>
         0 
       </mn> 
       <mo>
         , 
       </mo> 
       <mtext>
           
       </mtext> 
       <mi>
         i 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         1 
       </mn> 
       <mo>
         , 
       </mo> 
       <mo>
         ⋯ 
       </mo> 
       <mo>
         , 
       </mo> 
       <mi>
         N 
       </mi> 
      </mrow> 
     </math> (13)</p>
    <p>Wherein, W represents the classification weight vector, B is the bias, CC stands for the slack variable, C is the penalty parameter, and QQ denotes the kernel function used to map features into a high-dimensional space. By solving the aforementioned optimization problem, the SVM obtains the optimal separating hyperplane, enabling classification of newly input data points.</p>
    <p>3) Filter Selection Strategy</p>
    <p>Once the current motion pattern is determined, the system dynamically selects the corresponding Kalman filter variant based on the SVM classification result. The specific selection strategy is as follows:</p>
    <p>Therefore, given the input features X at the current time, the SVM outputs a label 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         y 
       </mi> 
       <mo>
         ∈ 
       </mo> 
       <mrow> 
        <mo>
          { 
        </mo> 
        <mrow> 
         <mn>
           1 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           2 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           3 
         </mn> 
        </mrow> 
        <mo>
          } 
        </mo> 
       </mrow> 
      </mrow> 
     </math>, and the corresponding filter is selected based on this label to enhance vehicle tracking performance.</p>
    <p>Compared to a fixed filter, the SVM-based adaptive selection mechanism overcomes the limitations of a single filter in complex scenarios by dynamically switching filters, thereby improving tracking accuracy and stability.</p>
    <p>
     <xref ref-type="fig" rid="fig2">
      Figure 2
     </xref> illustrates the flowchart of the SVM-based adaptive filter selection mechanism.</p>
    <p>Firstly, data preprocessing and feature extraction are conducted to compute the current speed, acceleration, and direction change angle from the target’s historical trajectory, forming the feature vector X.</p>
    <p>Subsequently, the feature vector X is input into an SVM classifier, which outputs the current motion mode label y (in a multi-class SVM, based on the trained classification boundaries, the SVM can effectively distinguish between uniform, accelerated, and complex nonlinear motions).</p>
    <fig id="fig2" position="float">
     <label>Figure 2</label>
     <caption>
      <title>Figure 2. Flowchart of the adaptive filter selection mechanism.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/2312887-rId75.jpeg?20241219021041" />
    </fig>
    <p>Concurrently, based on the classification result y, the corresponding Kalman filter variant is selected. For instance, when y = 1, the system employs the standard Kalman Filter (KF); when y = 2, the Extended Kalman Filter (EKF) is used; and when y = 3, the Unscented Kalman Filter (UKF) is chosen.</p>
    <p>Finally, with the selected filter, the YOLOv8 detection results are input as observations to the filter, and the updated target state is calculated to obtain more precise location information.</p>
   </sec>
   <sec id="s3_3">
    <title>3.3. Error Feedback Mechanism</title>
    <p>Apart from selecting the appropriate Kalman filter, dynamic changes or unexpected events in the environment (such as sudden acceleration, deceleration, occlusion, or unexpected direction changes of the target) can also lead to increased tracking errors. To enhance the system’s responsiveness to these unexpected events, this study proposes an error feedback mechanism that dynamically adjusts the parameters of the Kalman filter by feeding back detection errors. This mechanism not only compensates for errors generated during detection and tracking but also adapts when the filter’s response is poor, thereby improving the system’s stability and accuracy in complex environments.</p>
    <p>The design of the error feedback mechanism is based on dynamic feedback adjustment of detection errors and mainly consists of the following steps:</p>
    <p>1) Error Calculation</p>
    <p>In each frame, the detection error 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          e 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math> is calculated based on the difference between the target position detected by YOLOv8 (detection result) and the predicted position by the Kalman filter (prediction result).</p>
    <p>Assuming the detection result is 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          z 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math> (observation) and the filter’s prediction is 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mover accent="true"> 
         <mi>
           x 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           | 
         </mo> 
         <mi>
           k 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </msub> 
      </mrow> 
     </math>, the error is defined as Equation (14):</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          e 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mi>
          z 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <mo>
         − 
       </mo> 
       <msub> 
        <mover accent="true"> 
         <mi>
           x 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           | 
         </mo> 
         <mi>
           k 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </msub> 
      </mrow> 
     </math> (14)</p>
    <p>Wherein, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          e 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math> denotes the detection error vector at time 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        k 
      </mi> 
     </math>, encompassing information on the differences in position and velocity.</p>
    <p>2) Error Weight Calculation</p>
    <p>Different feedback intensities should be applied to detection errors in various scenarios. By introducing an error weight 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          W 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math>, the errors are weighted accordingly. The error weight can be dynamically adjusted based on the magnitude of the detection error and the frequency of occurrence of unexpected events. For instance, a larger feedback weight is assigned to significantly increased errors.</p>
    <p>The feedback weight can be defined by Equation (15):</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          W 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mi>
         α 
       </mi> 
       <mo>
         ⋅ 
       </mo> 
       <mrow> 
        <mo>
          ‖ 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            e 
          </mi> 
          <mi>
            k 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          ‖ 
        </mo> 
       </mrow> 
       <mo>
         + 
       </mo> 
       <mi>
         β 
       </mi> 
      </mrow> 
     </math> (15)</p>
    <p>Wherein, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        α 
      </mi> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        β 
      </mi> 
     </math> are adjustable parameters, and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mrow> 
        <mo>
          ‖ 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            e 
          </mi> 
          <mi>
            k 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          ‖ 
        </mo> 
       </mrow> 
      </mrow> 
     </math> represents the magnitude of the error. By adjusting the feedback weight, the filter’s responsiveness to unexpected events can be enhanced.</p>
    <p>3) Filter Gain Adjustment</p>
    <p>The error is fed back to the process noise covariance 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          Q 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math> and measurement noise covariance 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          R 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math> of the Kalman filter. By dynamically adjusting 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          Q 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          R 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math>, the filter can enhance its adaptability to different error scenarios, thereby improving tracking accuracy.</p>
    <p>The adjustment model for the feedback mechanism can be expressed as Equations (16) and (17):</p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          Q 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mi>
          Q 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <mo>
         + 
       </mo> 
       <msub> 
        <mi>
          W 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <mo>
         ⋅ 
       </mo> 
       <mo stretchy="false">
         ( 
       </mo> 
       <msub> 
        <mi>
          e 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <msubsup> 
        <mi>
          e 
        </mi> 
        <mi>
          k 
        </mi> 
        <mi>
          T 
        </mi> 
       </msubsup> 
       <mo stretchy="false">
         ) 
       </mo> 
      </mrow> 
     </math> (16)</p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          R 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mi>
          R 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <mo>
         + 
       </mo> 
       <msub> 
        <mi>
          W 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <mo>
         ⋅ 
       </mo> 
       <mo stretchy="false">
         ( 
       </mo> 
       <msub> 
        <mi>
          e 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
       <msubsup> 
        <mi>
          e 
        </mi> 
        <mi>
          k 
        </mi> 
        <mi>
          T 
        </mi> 
       </msubsup> 
       <mo stretchy="false">
         ) 
       </mo> 
      </mrow> 
     </math> (17)</p>
    <p>The error feedback adjusts the process noise 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          Q 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math> and the measurement noise covariance matrix 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          R 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math>, enabling the filter to adaptively adjust its gain matrix based on the current error.</p>
    <p>
     <xref ref-type="fig" rid="fig3">
      Figure 3
     </xref> illustrates the complete overall process of adaptive tracking.</p>
    <fig id="fig3" position="float">
     <label>Figure 3</label>
     <caption>
      <title>Figure 3. Overall flowchart of adaptive tracking.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/2312887-rId114.jpeg?20241219021041" />
    </fig>
   </sec>
  </sec><sec id="s4">
   <title>4. Experimental Setup</title>
   <sec id="s4_1">
    <title>4.1. Datasets and Evaluation Metrics</title>
    <p>To comprehensively evaluate the performance of the vehicle detection and tracking system proposed in this study, two widely used standard datasets, KITTI <xref ref-type="bibr" rid="scirp.138267-20">
      [20]
     </xref> and UA-DETRAC <xref ref-type="bibr" rid="scirp.138267-21">
      [21]
     </xref>, are selected. These datasets cover a variety of scenarios (such as urban roads, highways, and adverse weather conditions) and different dynamic conditions (such as occlusion and low-light environments). Additionally, to quantify the model’s performance, mean Average Precision (mAP), Multiple Object Tracking Accuracy (MOTA), and Frames Per Second (FPS) for real-time performance are adopted as the core evaluation metrics.</p>
    <p>1) KITTI Dataset</p>
    <p>The KITTI dataset is one of the standard datasets for autonomous driving research. It primarily consists of images from common traffic scenarios such as urban roads and highways, including 7481 training images and 7518 test images.</p>
    <p>2) UA-DETRAC Dataset</p>
    <p>The UA-DETRAC dataset focuses on vehicle detection and tracking tasks. It covers high-density traffic areas such as urban roads and intersections. It provides 10 hours of video data, comprising 1210 video clips and 140,000 annotated frames.</p>
    <p>
     <xref ref-type="table" rid="table1">
      Table 1
     </xref> summarizes the main characteristics of the KITTI and UA-DETRAC datasets:</p>
    <table-wrap id="table1">
     <label>
      <xref ref-type="table" rid="table1">
       Table 1
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.138267-"></xref>Table 1. Table type styles (Table caption is indispensable).</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="25.64%"><p style="text-align:center">Dataset</p></td> 
       <td class="custom-bottom-td acenter" width="25.64%"><p style="text-align:center">Task Type</p></td> 
       <td class="custom-bottom-td acenter" width="25.64%"><p style="text-align:center">Data Volume</p></td> 
       <td class="custom-bottom-td acenter" width="25.64%"><p style="text-align:center">Scenario Coverage</p></td> 
       <td class="custom-bottom-td acenter" width="25.64%"><p style="text-align:center">Challenges</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="25.64%"><p style="text-align:center">KITTI</p></td> 
       <td class="custom-top-td acenter" width="25.64%"><p style="text-align:center">Detection &amp; Tracking</p></td> 
       <td class="custom-top-td acenter" width="25.64%"><p style="text-align:center">7481 Training Images</p></td> 
       <td class="custom-top-td acenter" width="25.64%"><p style="text-align:center">Urban, Highways</p></td> 
       <td class="custom-top-td acenter" width="25.64%"><p style="text-align:center">Occlusion, Multi-target, Illumination Changes</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="25.64%"><p style="text-align:center">UA-DETRAC</p></td> 
       <td class="acenter" width="25.64%"><p style="text-align:center">Detection &amp; Tracking</p></td> 
       <td class="acenter" width="25.64%"><p style="text-align:center">1210 Video Clips</p></td> 
       <td class="acenter" width="25.64%"><p style="text-align:center">Urban, Traffic Intersections</p></td> 
       <td class="acenter" width="25.64%"><p style="text-align:center">Complex Background, Dense Targets, Low-light Scenarios</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>This paper quantitatively evaluates the system performance from both detection and tracking perspectives, mainly using the following metrics:</p>
    <p>1) Mean Average Precision (mAP@0.5)</p>
    <p>mAP is a core metric for object detection, used to measure the overall performance of a detection model across all categories. mAP@0.5 indicates the mean average precision calculated with an Intersection over Union (IoU) threshold of 0.5. Its calculation is shown in Equation (18):</p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         mAP 
       </mtext> 
       <mo>
         @ 
       </mo> 
       <mn>
         0.5 
       </mn> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mn>
          1 
        </mn> 
        <mi>
          N 
        </mi> 
       </mfrac> 
       <mstyle displaystyle="true"> 
        <msubsup> 
         <mo>
           ∑ 
         </mo> 
         <mrow> 
          <mi>
            i 
          </mi> 
          <mo>
            = 
          </mo> 
          <mn>
            1 
          </mn> 
         </mrow> 
         <mi>
           N 
         </mi> 
        </msubsup> 
        <mrow> 
         <mi>
           A 
         </mi> 
         <msub> 
          <mi>
            P 
          </mi> 
          <mi>
            i 
          </mi> 
         </msub> 
        </mrow> 
       </mstyle> 
      </mrow> 
     </math> (18)</p>
    <p>where, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        N 
      </mi> 
     </math> is the number of categories, and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         A 
       </mi> 
       <msub> 
        <mi>
          P 
        </mi> 
        <mi>
          i 
        </mi> 
       </msub> 
      </mrow> 
     </math> is the average precision for the i-th category.</p>
    <p>2) Multiple Object Tracking Accuracy (MOTA)</p>
    <p>MOTA is used to quantify the overall performance of a tracking task, taking into account object losses, mis-tracking, and ID switches. Its calculation is shown in Equation (19):</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         MOTA 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mn>
         1 
       </mn> 
       <mo>
         − 
       </mo> 
       <mfrac> 
        <mrow> 
         <msub> 
          <mstyle mathsize="140%" displaystyle="true"> 
           <mo>
             ∑ 
           </mo> 
          </mstyle> 
          <mi>
            t 
          </mi> 
         </msub> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mi>
             F 
           </mi> 
           <msub> 
            <mi>
              N 
            </mi> 
            <mi>
              t 
            </mi> 
           </msub> 
           <mo>
             + 
           </mo> 
           <mi>
             F 
           </mi> 
           <msub> 
            <mi>
              P 
            </mi> 
            <mi>
              t 
            </mi> 
           </msub> 
           <mo>
             + 
           </mo> 
           <mi>
             I 
           </mi> 
           <mi>
             D 
           </mi> 
           <msub> 
            <mi>
              S 
            </mi> 
            <mi>
              t 
            </mi> 
           </msub> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mrow> 
         <msub> 
          <mstyle mathsize="140%" displaystyle="true"> 
           <mo>
             ∑ 
           </mo> 
          </mstyle> 
          <mi>
            t 
          </mi> 
         </msub> 
         <mtext>
             
         </mtext> 
         <mi>
           G 
         </mi> 
         <msub> 
          <mi>
            T 
          </mi> 
          <mi>
            t 
          </mi> 
         </msub> 
        </mrow> 
       </mfrac> 
      </mrow> 
     </math> (19)</p>
    <p>where, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         F 
       </mi> 
       <msub> 
        <mi>
          N 
        </mi> 
        <mi>
          t 
        </mi> 
       </msub> 
      </mrow> 
     </math>, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         F 
       </mi> 
       <msub> 
        <mi>
          P 
        </mi> 
        <mi>
          t 
        </mi> 
       </msub> 
      </mrow> 
     </math>, and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         I 
       </mi> 
       <mi>
         D 
       </mi> 
       <msub> 
        <mi>
          S 
        </mi> 
        <mi>
          t 
        </mi> 
       </msub> 
      </mrow> 
     </math> represent the number of false negatives, false positives, and ID switches at frame 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        t 
      </mi> 
     </math>, respectively, and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         G 
       </mi> 
       <msub> 
        <mi>
          T 
        </mi> 
        <mi>
          t 
        </mi> 
       </msub> 
      </mrow> 
     </math> represents the total number of ground truth objects at frame 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        t 
      </mi> 
     </math>.</p>
    <p>3) Real-time Performance (FPS)</p>
    <p>Frames Per Second (FPS) is used to measure the speed performance of the model and is an important indicator of real-time performance. A high FPS value indicates that the system is suitable for practical applications.</p>
   </sec>
   <sec id="s4_2">
    <title>4.2. Experimental Parameters and Environment</title>
    <p>The experiments in this study were conducted in a high-performance computing environment to ensure the real-time and reliable performance of vehicle detection and tracking tasks. The following details the software and hardware configurations of the experimental environment, as well as the specific hyperparameter settings for the YOLOv8 detection module and Kalman filter.</p>
    <p>The software and hardware configurations used in the experiments are shown in <xref ref-type="table" rid="table2">
      Table 2
     </xref>:</p>
    <table-wrap id="table2">
     <label>
      <xref ref-type="table" rid="table2">
       Table 2
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.138267-"></xref>Table 2. Software and hardware configuration table.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="25.64%"><p style="text-align:center">Software/Hardware Component</p></td> 
       <td class="custom-bottom-td acenter" width="25.64%"><p style="text-align:center">Configuration Details</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="25.64%"><p style="text-align:center">Operating System</p></td> 
       <td class="custom-top-td acenter" width="25.64%"><p style="text-align:center">Windows 11</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="25.64%"><p style="text-align:center">Graphics Card</p></td> 
       <td class="acenter" width="25.64%"><p style="text-align:center">NVIDIA RTX 3060 12GB</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="25.64%"><p style="text-align:center">Memory</p></td> 
       <td class="acenter" width="25.64%"><p style="text-align:center">128GB DDR4</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="25.64%"><p style="text-align:center">Python</p></td> 
       <td class="acenter" width="25.64%"><p style="text-align:center">3.9</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="25.64%"><p style="text-align:center">PyTorch</p></td> 
       <td class="acenter" width="25.64%"><p style="text-align:center">1.12.1</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="25.64%"><p style="text-align:center">CUDA</p></td> 
       <td class="acenter" width="25.64%"><p style="text-align:center">11.6</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>The GPU (RTX 3060) supports rapid training and inference for YOLOv8, and the sufficient memory ensures the loading and processing of the dataset. The detection part of the experiment is implemented based on the PyTorch framework, while the Kalman filter and data processing modules utilize NumPy for matrix operations.</p>
    <p>The hyperparameters of the YOLOv8 model have been fine-tuned to ensure a balance between detection accuracy and speed. The specific configurations are shown in <xref ref-type="table" rid="table3">
      Table 3
     </xref>:</p>
    <table-wrap id="table3">
     <label>
      <xref ref-type="table" rid="table3">
       Table 3
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.138267-"></xref>Table 3. YOLOv8 hyperparameter configuration table.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="33.34%"><p style="text-align:center">Parameter Names</p></td> 
       <td class="custom-bottom-td acenter" width="16.04%"><p style="text-align:center">Setting</p></td> 
       <td class="custom-bottom-td acenter" width="50.62%"><p style="text-align:center">Instructions</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="33.34%"><p style="text-align:center">Image Size</p></td> 
       <td class="custom-top-td acenter" width="16.04%"><p style="text-align:center">640 × 640</p></td> 
       <td class="custom-top-td acenter" width="50.62%"><p style="text-align:center">Adjust all input images to a uniform size.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="33.34%"><p style="text-align:center">Confidence</p></td> 
       <td class="acenter" width="16.04%"><p style="text-align:center">0.5</p></td> 
       <td class="acenter" width="50.62%"><p style="text-align:center">Filter out detection bounding boxes with low confidence scores.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="33.34%"><p style="text-align:center">NMS (Non-Maximum Suppression)</p></td> 
       <td class="acenter" width="16.04%"><p style="text-align:center">0.45</p></td> 
       <td class="acenter" width="50.62%"><p style="text-align:center">Remove excessively overlapping detection bounding boxes.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="33.34%"><p style="text-align:center">Epochs</p></td> 
       <td class="acenter" width="16.04%"><p style="text-align:center">50</p></td> 
       <td class="acenter" width="50.62%"><p style="text-align:center">Number of iterations for model training.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="33.34%"><p style="text-align:center">Batch Size</p></td> 
       <td class="acenter" width="16.04%"><p style="text-align:center">16</p></td> 
       <td class="acenter" width="50.62%"><p style="text-align:center">Number of images input to the model per batch.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="33.34%"><p style="text-align:center">Learning Rate (Initial Value)</p></td> 
       <td class="acenter" width="16.04%"><p style="text-align:center">0.01</p></td> 
       <td class="acenter" width="50.62%"><p style="text-align:center">Dynamically adjust the learning rate using a cosine annealing strategy.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="33.34%"><p style="text-align:center">Optimizer</p></td> 
       <td class="acenter" width="16.04%"><p style="text-align:center">SGD</p></td> 
       <td class="acenter" width="50.62%"><p style="text-align:center">Optimize model parameters using stochastic gradient descent.</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="33.34%"><p style="text-align:center">Data Augmentation</p></td> 
       <td class="acenter" width="16.04%"><p style="text-align:center">Open</p></td> 
       <td class="acenter" width="50.62%"><p style="text-align:center">Data augmentation techniques include random cropping, rotation, scaling, and color jittering.</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>These hyperparameter settings have been optimized on the KITTI and UA-DETRAC datasets, effectively enhancing the detection performance of YOLOv8 in various scenarios.</p>
    <p>To adapt to dynamic scenarios in vehicle tracking, the initialization parameters for various Kalman filters are shown in <xref ref-type="table" rid="table4">
      Table 4
     </xref>:</p>
    <table-wrap id="table4">
     <label>
      <xref ref-type="table" rid="table4">
       Table 4
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.138267-"></xref>Table 4. Kalman filter hyperparameter configuration table.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="25.64%"><p style="text-align:center">Filter Type</p></td> 
       <td class="custom-bottom-td acenter" width="25.64%"><p style="text-align:center">State Transition Covariance (Q)</p></td> 
       <td class="custom-bottom-td acenter" width="25.64%"><p style="text-align:center">Measurement Noise Covariance(R)</p></td> 
       <td class="custom-bottom-td acenter" width="25.64%"><p style="text-align:center">Initial Error Covariance(P)</p></td> 
       <td class="custom-bottom-td acenter" width="25.64%"><p style="text-align:center">State Dimension</p></td> 
       <td class="custom-bottom-td acenter" width="25.64%"><p style="text-align:center">Observation Dimension</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="25.64%"><p style="text-align:center">KF</p></td> 
       <td class="custom-top-td acenter" width="25.64%"><p style="text-align:center">0.1I</p></td> 
       <td class="custom-top-td acenter" width="25.64%"><p style="text-align:center">0.05I</p></td> 
       <td class="custom-top-td acenter" width="25.64%"><p style="text-align:center">1.0I</p></td> 
       <td class="custom-top-td acenter" width="25.64%"><p style="text-align:center">4</p></td> 
       <td class="custom-top-td acenter" width="25.64%"><p style="text-align:center">2</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="25.64%"><p style="text-align:center">EKF</p></td> 
       <td class="acenter" width="25.64%"><p style="text-align:center">0.15I</p></td> 
       <td class="acenter" width="25.64%"><p style="text-align:center">0.05I</p></td> 
       <td class="acenter" width="25.64%"><p style="text-align:center">1.0I</p></td> 
       <td class="acenter" width="25.64%"><p style="text-align:center">4</p></td> 
       <td class="acenter" width="25.64%"><p style="text-align:center">2</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="25.64%"><p style="text-align:center">UKF</p></td> 
       <td class="acenter" width="25.64%"><p style="text-align:center">0.2I</p></td> 
       <td class="acenter" width="25.64%"><p style="text-align:center">0.1I</p></td> 
       <td class="acenter" width="25.64%"><p style="text-align:center">1.0I</p></td> 
       <td class="acenter" width="25.64%"><p style="text-align:center">4</p></td> 
       <td class="acenter" width="25.64%"><p style="text-align:center">2</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>Where Q, R, and P represent the process noise covariance matrix, measurement noise covariance matrix, and initial error covariance matrix, respectively; I denotes the identity matrix.</p>
    <p>The hardware and software environments of the experiment provide powerful computational support for large-scale detection and tracking tasks. The optimized hyperparameters of YOLOv8 enable it to be efficient and robust in multiple scenarios, while the reasonable initialization and dynamic adjustment strategies of the Kalman filter ensure the accuracy and real-time performance of the tracking module.</p>
   </sec>
  </sec><sec id="s5">
   <title>5. Experimental Results and Analysis</title>
   <sec id="s5_1">
    <title>5.1. Comparative Analysis of Detection and Tracking</title>
    <p>Vehicle detection serves as the foundational step in achieving multi-object tracking, with its detection accuracy directly impacting the predictive performance of subsequent filters. To evaluate the performance of the YOLOv8 detection model used in this paper, comprehensive comparative experiments were conducted with mainstream object detection models (such as YOLOv5, Faster R-CNN, and RetinaNet) across multiple typical scenarios (including daytime, nighttime, occlusion, low-light, etc.). <xref ref-type="table" rid="table5">
      Table 5
     </xref> summarizes the performance of different detection models in these typical scenarios.</p>
    <table-wrap id="table5">
     <label>
      <xref ref-type="table" rid="table5">
       Table 5
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.138267-"></xref>Table 5. Comparison of performance among different detection models.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="18.99%"><p style="text-align:center">Models</p></td> 
       <td class="custom-bottom-td acenter" width="14.71%"><p style="text-align:center">Daytime</p><p style="text-align:center">mAP@0.5</p></td> 
       <td class="custom-bottom-td acenter" width="14.73%"><p style="text-align:center">Nighttime</p><p style="text-align:center">mAP@0.5</p></td> 
       <td class="custom-bottom-td acenter" width="14.71%"><p style="text-align:center">Occlusion</p><p style="text-align:center">mAP@0.5</p></td> 
       <td class="custom-bottom-td acenter" width="14.73%"><p style="text-align:center">low-light</p></td> 
       <td class="custom-bottom-td acenter" width="11.06%"><p style="text-align:center">FPS</p></td> 
       <td class="custom-bottom-td acenter" width="11.06%"><p style="text-align:center">Loss</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="18.99%"><p style="text-align:center">YOLOv8</p></td> 
       <td class="custom-top-td acenter" width="14.71%"><p style="text-align:center">94.3%</p></td> 
       <td class="custom-top-td acenter" width="14.73%"><p style="text-align:center">90.1%</p></td> 
       <td class="custom-top-td acenter" width="14.71%"><p style="text-align:center">87.8%</p></td> 
       <td class="custom-top-td acenter" width="14.73%"><p style="text-align:center">84.5%</p></td> 
       <td class="custom-top-td acenter" width="11.06%"><p style="text-align:center">45</p></td> 
       <td class="custom-top-td acenter" width="11.06%"><p style="text-align:center">0.43</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.99%"><p style="text-align:center">YOLOv5</p></td> 
       <td class="acenter" width="14.71%"><p style="text-align:center">91.2%</p></td> 
       <td class="acenter" width="14.73%"><p style="text-align:center">85.4%</p></td> 
       <td class="acenter" width="14.71%"><p style="text-align:center">83.1%</p></td> 
       <td class="acenter" width="14.73%"><p style="text-align:center">79.8%</p></td> 
       <td class="acenter" width="11.06%"><p style="text-align:center">35</p></td> 
       <td class="acenter" width="11.06%"><p style="text-align:center">0.57</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.99%"><p style="text-align:center">Faster R-CNN</p></td> 
       <td class="acenter" width="14.71%"><p style="text-align:center">88.5%</p></td> 
       <td class="acenter" width="14.73%"><p style="text-align:center">83.2%</p></td> 
       <td class="acenter" width="14.71%"><p style="text-align:center">79.5%</p></td> 
       <td class="acenter" width="14.73%"><p style="text-align:center">75.4%</p></td> 
       <td class="acenter" width="11.06%"><p style="text-align:center">15</p></td> 
       <td class="acenter" width="11.06%"><p style="text-align:center">0.62</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.99%"><p style="text-align:center">RetinaNet</p></td> 
       <td class="acenter" width="14.71%"><p style="text-align:center">89.1%</p></td> 
       <td class="acenter" width="14.73%"><p style="text-align:center">82.8%</p></td> 
       <td class="acenter" width="14.71%"><p style="text-align:center">80.3%</p></td> 
       <td class="acenter" width="14.73%"><p style="text-align:center">77.0%</p></td> 
       <td class="acenter" width="11.06%"><p style="text-align:center">20</p></td> 
       <td class="acenter" width="11.06%"><p style="text-align:center">0.59</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>From the comparison of mean Average Precision (mAP) at IoU threshold of 0.5, it can be seen that YOLOv8 outperforms other models in all scenarios, particularly in nighttime and occlusion scenarios, achieving mAP values of 90.1% and 87.8%, respectively.</p>
    <p>The final loss value of YOLOv8 is 0.43, which is approximately 25% lower than that of YOLOv5 and RetinaNet, and approximately 30% lower than that of Faster R-CNN. This verifies the rationality of selecting YOLOv8 as the detection module in this paper, providing robust support for subsequent tracking tasks.</p>
    <p>
     <xref ref-type="table" rid="table6">
      Table 6
     </xref> presents the performance of the SVM-based adaptive filter selection mechanism across different scenarios. The experimental results indicate that the adaptive selection mechanism significantly outperforms fixed filter schemes in terms of tracking accuracy and robustness in complex scenarios such as accelerated motion and nonlinear motion.</p>
    <table-wrap id="table6">
     <label>
      <xref ref-type="table" rid="table6">
       Table 6
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.138267-"></xref>Table 6. Performance of the SVM-based adaptive filter selection mechanism.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="24.27%"><p style="text-align:center">Scene Types</p></td> 
       <td class="custom-bottom-td acenter" width="24.27%"><p style="text-align:center">Filter Selection</p></td> 
       <td class="custom-bottom-td acenter" width="17.15%"><p style="text-align:center">Tracking Accuracy(MOTA)</p></td> 
       <td class="custom-bottom-td acenter" width="17.15%"><p style="text-align:center">Mean Squared Error (MSE)</p></td> 
       <td class="custom-bottom-td acenter" width="17.15%"><p style="text-align:center">Computational Latency(ms)</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="24.27%"><p style="text-align:center">Uniform Motion</p></td> 
       <td class="custom-top-td acenter" width="24.27%"><p style="text-align:center">KF (Kalman Filter)</p></td> 
       <td class="custom-top-td acenter" width="17.15%"><p style="text-align:center">89%</p></td> 
       <td class="custom-top-td acenter" width="17.15%"><p style="text-align:center">1.5</p></td> 
       <td class="custom-top-td acenter" width="17.15%"><p style="text-align:center">12</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="24.27%"><p style="text-align:center">Accelerated Motion</p></td> 
       <td class="acenter" width="24.27%"><p style="text-align:center">EKF (Extended Kalman Filter)</p></td> 
       <td class="acenter" width="17.15%"><p style="text-align:center">92%</p></td> 
       <td class="acenter" width="17.15%"><p style="text-align:center">1.2</p></td> 
       <td class="acenter" width="17.15%"><p style="text-align:center">14</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="24.27%"><p style="text-align:center">Complex Nonlinear Motion</p></td> 
       <td class="acenter" width="24.27%"><p style="text-align:center">UKF (Unscented Kalman Filter)</p></td> 
       <td class="acenter" width="17.15%"><p style="text-align:center">94%</p></td> 
       <td class="acenter" width="17.15%"><p style="text-align:center">0.9</p></td> 
       <td class="acenter" width="17.15%"><p style="text-align:center">20</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="24.27%"><p style="text-align:center">SVM-Based Adaptive Selection</p></td> 
       <td class="acenter" width="24.27%"><p style="text-align:center">Dynamic Switching</p></td> 
       <td class="acenter" width="17.15%"><p style="text-align:center">95%</p></td> 
       <td class="acenter" width="17.15%"><p style="text-align:center">0.8</p></td> 
       <td class="acenter" width="17.15%"><p style="text-align:center">15</p></td> 
      </tr> 
     </table>
    </table-wrap>
   </sec>
   <sec id="s5_2">
    <title>5.2. Effectiveness of Error Feedback Mechanism</title>
    <p>In scenarios involving occlusion, sudden motion, and multi-object environments, the error feedback + SVM dynamic selection mechanism significantly improves tracking accuracy compared to other methods. Below is a comparison of the Mean Squared Error (MSE) for each method across different scenarios.</p>
    <table-wrap id="table7">
     <label>
      <xref ref-type="table" rid="table7">
       Table 7
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.138267-"></xref>Table 7. Performance of the error feedback + SVM dynamic selection mechanism.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="18.57%"><p style="text-align:center">Scene Types</p></td> 
       <td class="custom-bottom-td acenter" width="36.30%"><p style="text-align:center">Methods</p></td> 
       <td class="custom-bottom-td acenter" width="10.34%"><p style="text-align:center">MSE</p></td> 
       <td class="custom-bottom-td acenter" width="17.40%"><p style="text-align:center">Trajectory Smoothness</p></td> 
       <td class="custom-bottom-td acenter" width="17.40%"><p style="text-align:center">Computational Latency (ms)</p></td> 
      </tr> 
      <tr> 
       <td rowspan="5" class="custom-top-td acenter" width="18.57%"><p style="text-align:center">Occlusion Scenarios</p></td> 
       <td class="custom-top-td acenter" width="36.30%"><p style="text-align:center">Single KF</p></td> 
       <td class="custom-top-td acenter" width="10.34%"><p style="text-align:center">15.2</p></td> 
       <td class="custom-top-td acenter" width="17.40%"><p style="text-align:center">Low</p></td> 
       <td class="custom-top-td acenter" width="17.40%"><p style="text-align:center">10</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="36.30%"><p style="text-align:center">Single EKF</p></td> 
       <td class="acenter" width="10.34%"><p style="text-align:center">12.4</p></td> 
       <td class="acenter" width="17.40%"><p style="text-align:center">Medium</p></td> 
       <td class="acenter" width="17.40%"><p style="text-align:center">12</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="36.30%"><p style="text-align:center">Single UKF</p></td> 
       <td class="acenter" width="10.34%"><p style="text-align:center">11.5</p></td> 
       <td class="acenter" width="17.40%"><p style="text-align:center">Medium</p></td> 
       <td class="acenter" width="17.40%"><p style="text-align:center">20</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="36.30%"><p style="text-align:center">SVM Dynamic Selection Mechanism</p></td> 
       <td class="acenter" width="10.34%"><p style="text-align:center">9.7</p></td> 
       <td class="acenter" width="17.40%"><p style="text-align:center">High</p></td> 
       <td class="acenter" width="17.40%"><p style="text-align:center">18</p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td acenter" width="36.30%"><p style="text-align:center">Error Feedback + SVM Dynamic Selection Mechanism</p></td> 
       <td class="custom-bottom-td acenter" width="10.34%"><p style="text-align:center">7.1</p></td> 
       <td class="custom-bottom-td acenter" width="17.40%"><p style="text-align:center">High</p></td> 
       <td class="custom-bottom-td acenter" width="17.40%"><p style="text-align:center">20</p></td> 
      </tr> 
      <tr> 
       <td rowspan="5" class="custom-top-td acenter" width="18.57%"><p style="text-align:center">Sudden Motion Scenarios</p></td> 
       <td class="custom-top-td acenter" width="36.30%"><p style="text-align:center">Single KF</p></td> 
       <td class="custom-top-td acenter" width="10.34%"><p style="text-align:center">18.7</p></td> 
       <td class="custom-top-td acenter" width="17.40%"><p style="text-align:center">Low</p></td> 
       <td class="custom-top-td acenter" width="17.40%"><p style="text-align:center">10</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="36.30%"><p style="text-align:center">Single EKF</p></td> 
       <td class="acenter" width="10.34%"><p style="text-align:center">13.2</p></td> 
       <td class="acenter" width="17.40%"><p style="text-align:center">Medium</p></td> 
       <td class="acenter" width="17.40%"><p style="text-align:center">15</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="36.30%"><p style="text-align:center">Single UKF</p></td> 
       <td class="acenter" width="10.34%"><p style="text-align:center">11.0</p></td> 
       <td class="acenter" width="17.40%"><p style="text-align:center">Medium</p></td> 
       <td class="acenter" width="17.40%"><p style="text-align:center">20</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="36.30%"><p style="text-align:center">SVM Dynamic Selection Mechanism</p></td> 
       <td class="acenter" width="10.34%"><p style="text-align:center">9.2</p></td> 
       <td class="acenter" width="17.40%"><p style="text-align:center">High</p></td> 
       <td class="acenter" width="17.40%"><p style="text-align:center">18</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="36.30%"><p style="text-align:center">Error Feedback + SVM Dynamic Selection Mechanism</p></td> 
       <td class="acenter" width="10.34%"><p style="text-align:center">6.9</p></td> 
       <td class="acenter" width="17.40%"><p style="text-align:center">High</p></td> 
       <td class="acenter" width="17.40%"><p style="text-align:center">20</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>As shown in <xref ref-type="table" rid="table7">
      Table 7
     </xref>, the error feedback + SVM dynamic selection mechanism significantly reduces MSE in occlusion and sudden motion scenarios, with superior smoothness and real-time performance indicators compared to single filters and the standalone SVM dynamic selection mechanism.</p>
    <p>As shown in <xref ref-type="table" rid="table8">
      Table 8
     </xref>, a further evaluation is conducted to assess the difference in effectiveness between the standalone SVM dynamic selection mechanism and the SVM dynamic selection mechanism combined with error feedback. Two different error feedback mechanisms under distinct strategies are employed and compared across multiple scenarios.</p>
    <p>1) Error Feedback Standalone Adjustment Strategy: Error feedback only adjusts the parameters of the currently selected filter without triggering a re-selection.</p>
    <p>2) Error Feedback + Filter Re-selection Strategy: Error feedback not only adjusts the parameters but also triggers a re-selection of the filter.</p>
    <table-wrap id="table8">
     <label>
      <xref ref-type="table" rid="table8">
       Table 8
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.138267-"></xref>Table 8. Comparative experiments of different strategies.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="24.89%"><p style="text-align:center">Scene Types</p></td> 
       <td class="custom-bottom-td acenter" width="39.27%"><p style="text-align:center">Strategies</p></td> 
       <td class="custom-bottom-td acenter" width="10.34%"><p style="text-align:center">MSE</p></td> 
       <td class="custom-bottom-td acenter" width="25.50%"><p style="text-align:center">Switching Frequency (times/second)</p></td> 
      </tr> 
      <tr> 
       <td rowspan="2" class="custom-top-td acenter" width="24.89%"><p style="text-align:center">Occlusion Scenarios</p></td> 
       <td class="custom-top-td acenter" width="39.27%"><p style="text-align:center">Standalone Adjustment Strategy</p></td> 
       <td class="custom-top-td acenter" width="10.34%"><p style="text-align:center">8.8</p></td> 
       <td class="custom-top-td acenter" width="25.50%"><p style="text-align:center">2</p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td acenter" width="39.27%"><p style="text-align:center">Filter Re-selection Strategy</p></td> 
       <td class="custom-bottom-td acenter" width="10.34%"><p style="text-align:center">7.1</p></td> 
       <td class="custom-bottom-td acenter" width="25.50%"><p style="text-align:center">3</p></td> 
      </tr> 
      <tr> 
       <td rowspan="2" class="custom-top-td acenter" width="24.89%"><p style="text-align:center">Sudden Motion Scenarios</p></td> 
       <td class="custom-top-td acenter" width="39.27%"><p style="text-align:center">Standalone Adjustment Strategy</p></td> 
       <td class="custom-top-td acenter" width="10.34%"><p style="text-align:center">8.5</p></td> 
       <td class="custom-top-td acenter" width="25.50%"><p style="text-align:center">3</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="39.27%"><p style="text-align:center">Filter Re-selection Strategy</p></td> 
       <td class="acenter" width="10.34%"><p style="text-align:center">6.9</p></td> 
       <td class="acenter" width="25.50%"><p style="text-align:center">4</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>The performance advantages of SVM in dynamic filter selection primarily stem from the following aspects:</p>
    <p>a) Discriminative power of motion features: Experiments have demonstrated that velocity, acceleration, and direction change angles exhibit significant distribution differences across different motion patterns, providing a clear basis for filter classification.</p>
    <p>b) Nonlinear classification capability: Compared to traditional linear classifiers, SVM effectively handles nonlinear patterns through kernel function mapping, enabling more precise selection among KF, EKF, and UKF.</p>
    <p>c) Matching of errors with scenario characteristics: By incorporating an error feedback mechanism, SVM dynamically adjusts selection strategies based on real-time error variations, further enhancing the system’s adaptability.</p>
   </sec>
   <sec id="s5_3">
    <title>5.3. Robustness of Adaptive Kalman Filtering</title>
    <p>To validate the robustness of the SVM-based adaptive Kalman filtering, this paper analyzes the visualization results of tracking and recognition in various complex environments, as shown in <xref ref-type="fig" rid="fig4">
      Figure 4
     </xref>.</p>
    <p>(a) The four frames in row (a) illustrate the tracking and recognition performance of the model for a target vehicle with ID 17 in a turning scenario. It can be observed that the model demonstrates good performance.</p>
    <p>(b) The four frames in row (b) show the tracking and recognition performance of the model for a target vehicle with ID 1 at different angles during nighttime. It is evident that the model maintains good locking on the target vehicle, even during the process of being overtaken.</p>
    <p>(c) The four frames in row (c) present the tracking and recognition performance of the model for a target vehicle with ID 1 across multiple scenarios on a road segment. It is clear that the model can effectively identify and track the target vehicle even in multi-target and complex scenarios.</p>
    <p>From this, it can be concluded that the SVM-based adaptive Kalman filtering exhibits strong robustness in complex environments. It can provide decision support for intelligent transportation, autonomous driving, security, and other fields.</p>
    <fig id="fig4" position="float">
     <label>Figure 4</label>
     <caption>
      <title>Figure 4. Visualization of inference results.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/2312887-rId135.jpeg?20241219021044" />
    </fig>
   </sec>
   <sec id="s5_4">
    <title>5.4. Comparison with Existing Methods</title>
    <p>Through experimental comparisons with current mainstream vehicle tracking methods such as DeepSORT and FAIR MOT, the proposed research method in this study demonstrates advantages in detection accuracy, tracking accuracy, and system stability.</p>
    <p>As shown in <xref ref-type="table" rid="table9">
      Table 9
     </xref>, compared to DeepSORT’s 91.5% mAP on the KITTI dataset and FAIR MOT’s 89.2% mAP, the YOLOv8 detection module combined with a dynamic filter selection mechanism proposed in this paper achieves a 94.3% mAP, exhibiting stronger adaptability to complex scenarios. Additionally, compared to DeepSORT’s 35 FPS and FAIR MOT’s 30 FPS, the real-time performance of the proposed method (45 FPS) significantly enhances processing speed while meeting accuracy requirements, making it suitable for practical deployment.</p>
    <table-wrap id="table9">
     <label>
      <xref ref-type="table" rid="table9">
       Table 9
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.138267-"></xref>Table 9. Comparison of experimental results table.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="19.99%"><p style="text-align:center">Datasets</p></td> 
       <td class="custom-bottom-td acenter" width="20.00%"><p style="text-align:center">Methods</p></td> 
       <td class="custom-bottom-td acenter" width="20.00%"><p style="text-align:center">mAP</p></td> 
       <td class="custom-bottom-td acenter" width="20.00%"><p style="text-align:center">MOTA</p></td> 
       <td class="custom-bottom-td acenter" width="20.00%"><p style="text-align:center">FPS</p></td> 
      </tr> 
      <tr> 
       <td rowspan="3" class="custom-top-td acenter" width="19.99%"><p style="text-align:center">KITTI</p></td> 
       <td class="custom-top-td acenter" width="20.00%"><p style="text-align:center">DeepSORT</p></td> 
       <td class="custom-top-td acenter" width="20.00%"><p style="text-align:center">91.5%</p></td> 
       <td class="custom-top-td acenter" width="20.00%"><p style="text-align:center">87.5%</p></td> 
       <td class="custom-top-td acenter" width="20.00%"><p style="text-align:center">35</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="20.00%"><p style="text-align:center">FAIR MOT</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">89.2%</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">88.3%</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">30</p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td acenter" width="20.00%"><p style="text-align:center">Ours</p></td> 
       <td class="custom-bottom-td acenter" width="20.00%"><p style="text-align:center">94.3%</p></td> 
       <td class="custom-bottom-td acenter" width="20.00%"><p style="text-align:center">90.6%</p></td> 
       <td class="custom-bottom-td acenter" width="20.00%"><p style="text-align:center">45</p></td> 
      </tr> 
      <tr> 
       <td rowspan="3" class="custom-top-td acenter" width="19.99%"><p style="text-align:center">UA-DETRAC</p></td> 
       <td class="custom-top-td acenter" width="20.00%"><p style="text-align:center">DeepSORT</p></td> 
       <td class="custom-top-td acenter" width="20.00%"><p style="text-align:center">90.2%</p></td> 
       <td class="custom-top-td acenter" width="20.00%"><p style="text-align:center">85.4%</p></td> 
       <td class="custom-top-td acenter" width="20.00%"><p style="text-align:center">33</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="20.00%"><p style="text-align:center">FAIR MOT</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">88.1%</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">86.1%</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">28</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="20.00%"><p style="text-align:center">Ours</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">92.1%</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">89.7%</p></td> 
       <td class="acenter" width="20.00%"><p style="text-align:center">40</p></td> 
      </tr> 
     </table>
    </table-wrap>
   </sec>
  </sec><sec id="s6">
   <title>6. Conclusions and Future Work</title>
   <p>This paper proposes an SVM-based adaptive filter selection mechanism and an error feedback mechanism, addressing the challenges of vehicle tracking in dynamic and complex scenarios with impressive results.</p>
   <p>Although the integration of the components (YOLOv8, KF variants, SVM dynamic selection, and error feedback mechanism) represents incremental innovation, it yields the following significant system-level advantages:</p>
   <p>1) Enhanced adaptability in complex dynamic scenarios, reducing accuracy degradation due to changing motion patterns through dynamic filter selection.</p>
   <p>2) Improved coordination between detection and prediction, with the error feedback mechanism ensuring rapid system response to unexpected events.</p>
   <p>3) Simplified complexity in multi-module development through deep integration, while optimizing both real-time performance and resource utilization.</p>
   <p>Future research will focus on multimodal data fusion, incorporating sensor data from LiDAR and mmWave radars to enhance system applicability in adverse weather conditions, and large-scale application deployment, exploring distributed computing methods to meet real-time demands of large-scale traffic monitoring systems.</p>
  </sec>
 </body><back>
  <ref-list>
   <title>References</title>
   <ref id="scirp.138267-ref1">
    <label>1</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Tong, R.H., Li, J., Chen, H.S., et al. (2002) A Solution for Vehicle Tracking System Based on GPS and GIS Integration. Computer Engineering and Science, 1, 29-32.
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref2">
    <label>2</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Brown, M., Funke, J., Erlien, S. and Gerdes, J.C. (2017) Safe Driving Envelopes for Path Tracking in Autonomous Vehicles. Control Engineering Practice, 61, 307-316. &gt;https://doi.org/10.1016/j.conengprac.2016.04.013 
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref3">
    <label>3</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Liu, H.M., Guan, H., Yu, M.H., et al. (2020) Research and Implementation of a Vehicle Tracking Algorithm Based on Multi-feature Fusion. Journal of Chinese Computer Systems, 41, 1258-1262.
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref4">
    <label>4</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Wang, H.Y. (2022) Research on Vehicle Tracking Method Based on Intelligent Reflecting Surfaces in Complex Environments. Jilin University.
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref5">
    <label>5</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Varghese, R. and M., S. (2024). Yolov8: A Novel Object Detection Algorithm with Enhanced Performance and Robustness. 2024 International Conference on Advances in Data Engineering and Intelligent Computing Systems, Chennai, 18-19 April 2024, 1-6. &gt;https://doi.org/10.1109/adics58448.2024.10533619 
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref6">
    <label>6</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kim, H. (2019) Multiple Vehicle Tracking and Classification System with a Convolutional Neural Network. Journal of Ambient Intelligence and Humanized Computing, 13, 1603-1614. &gt;https://doi.org/10.1007/s12652-019-01429-5 
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref7">
    <label>7</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Fu, C., Lu, K., Zheng, G., Ye, J., Cao, Z., Li, B., et al. (2023) Siamese Object Tracking for Unmanned Aerial Vehicle: A Review and Comprehensive Analysis. Artificial Intelligence Review, 56, 1417-1477. &gt;https://doi.org/10.1007/s10462-023-10558-5 
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref8">
    <label>8</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Redmon, J., Divvala, S., Girshick, R. and Farhadi, A. (2016) You Only Look Once: Unified, Real-Time Object Detection. 2016 IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, 27-30 June 2016, 779-788. &gt;https://doi.org/10.1109/cvpr.2016.91 
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref9">
    <label>9</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Azimjonov, J. and Özmen, A. (2021) A Real-Time Vehicle Detection and a Novel Vehicle Tracking Systems for Estimating and Monitoring Traffic Flow on Highways. Advanced Engineering Informatics, 50, Article 101393. &gt;https://doi.org/10.1016/j.aei.2021.101393 
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref10">
    <label>10</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Bakirci, M. (2024) Enhancing Vehicle Detection in Intelligent Transportation Systems via Autonomous UAV Platform and Yolov8 Integration. Applied Soft Computing, 164, Article 112015. &gt;https://doi.org/10.1016/j.asoc.2024.112015 
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref11">
    <label>11</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ren, S., He, K., Girshick, R. and Sun, J. (2017) Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39, 1137-1149. &gt;https://doi.org/10.1109/tpami.2016.2577031 
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref12">
    <label>12</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Zhang, Y., Liu, Z.L. and Wan, W. (2021) UAV-Based Vehicle Target Detection Using Faster R-CNN. Electronic Science and Technology, 34, 11-20. &gt;https://doi.org/10.16180/j.cnki.issn1007-7820.2021.11.002 
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref13">
    <label>13</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ouyang, B., Zhu, Y.J., Yang, L.K., et al. (2024) FA-SORT: A Lightweight Multi-Vehicle Tracking Algorithm. Computer Engineering and Applications, 60, 122-134.
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref14">
    <label>14</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Gordon, N.J., Salmond, D.J. and Smith, A.F.M. (1993) Novel Approach to Nonlinear/non-Gaussian Bayesian State Estimation. IEE Proceedings F Radar and Signal Processing, 140, 107-113. &gt;https://doi.org/10.1049/ip-f-2.1993.0015 
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref15">
    <label>15</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Reid, D. (1979) An Algorithm for Tracking Multiple Targets. IEEE Transactions on Automatic Control, 24, 843-854. &gt;https://doi.org/10.1109/tac.1979.1102177 
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref16">
    <label>16</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Wan, E.A. and Van Der Merwe, R. (2000) The Unscented Kalman Filter for Nonlinear Estimation. Proceedings of the IEEE 2000 Adaptive Systems for Signal Processing, Communications, and Control Symposium, Lake Louise, 4 October 2000, 153-158. &gt;https://doi.org/10.1109/asspcc.2000.882463 
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref17">
    <label>17</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Farag, W. (2022) Multiple Road-Objects Detection and Tracking for Autonomous Driving. Journal of Engineering Research, 10, 237-262.
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref18">
    <label>18</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Bui, T., Wang, G., Wei, G. and Zeng, Q. (2024) Vehicle Multi-Object Detection and Tracking Algorithm Based on Improved You Only Look Once 5s Version and Deepsort. Applied Sciences, 14, Article 2690. &gt;https://doi.org/10.3390/app14072690 
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref19">
    <label>19</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Gao, J., Han, G., Zhu, H. and Liao, L. (2024) Multiple Moving Vehicles Tracking Algorithm with Attention Mechanism and Motion Model. Electronics, 13, Article 242. &gt;https://doi.org/10.3390/electronics13010242 
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref20">
    <label>20</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Geiger, A., Lenz, P., Stiller, C. and Urtasun, R. (2013) Vision Meets Robotics: The KITTI Dataset. The International Journal of Robotics Research, 32, 1231-1237. &gt;https://doi.org/10.1177/0278364913491297 
    </mixed-citation>
   </ref>
   <ref id="scirp.138267-ref21">
    <label>21</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Wen, L., Du, D., Cai, Z., Lei, Z., Chang, M., Qi, H., et al. (2020) UA-DETRAC: A New Benchmark and Protocol for Multi-Object Detection and Tracking. Computer Vision and Image Understanding, 193, Article 102907. &gt;https://doi.org/10.1016/j.cviu.2020.102907
    </mixed-citation>
   </ref>
  </ref-list>
 </back>
</article>