<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">ojapps</journal-id>
      <journal-title-group>
        <journal-title>Open Journal of Applied Sciences</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2165-3925</issn>
      <issn pub-type="ppub">2165-3917</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/ojapps.2026.164080</article-id>
      <article-id pub-id-type="publisher-id">ojapps-151115</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Biomedical</subject>
          <subject>Life Sciences</subject>
          <subject>Chemistry</subject>
          <subject>Materials Science</subject>
          <subject>Computer Science</subject>
          <subject>Communications</subject>
          <subject>Engineering</subject>
          <subject>Physics</subject>
          <subject>Mathematics</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Research on Sheep Face Recognition Based on Deep Learning and YOLOv3</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Zhen</surname>
            <given-names>Xu</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author" corresp="yes">
          <name name-style="western">
            <surname>Chen</surname>
            <given-names>Guoqing</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Wen</surname>
            <given-names>Lerong</given-names>
          </name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Jia</surname>
            <given-names>Haifeng</given-names>
          </name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> Inner Mongolia Key Laboratory of Aeolian Physics and Desertification Engineering, Hohhot, China </aff>
      <aff id="aff2"><label>2</label> College of Desert Control Science and Engineering, Inner Mongolia Agricultural University, Hohhot, China </aff>
      <aff id="aff3"><label>3</label> Inner Mongolia Hangjin Desert Ecological Position Research Station, Ordos, China </aff>
      <aff id="aff4"><label>4</label> Inner Mongolia Caodu Grass and Livestock Ecological Technology Co., Ltd., Hohhot, China </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The authors declare no conflicts of interest regarding the publication of this paper.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>02</day>
        <month>04</month>
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>04</month>
        <year>2026</year>
      </pub-date>
      <volume>16</volume>
      <issue>04</issue>
      <fpage>1386</fpage>
      <lpage>1399</lpage>
      <history>
        <date date-type="received">
          <day>22</day>
          <month>10</month>
          <year>2026</year>
        </date>
        <date date-type="accepted">
          <day>27</day>
          <month>04</month>
          <year>2026</year>
        </date>
        <date date-type="published">
          <day>30</day>
          <month>04</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2026 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/ojapps.2026.164080">https://doi.org/10.4236/ojapps.2026.164080</self-uri>
      <abstract>
        <p>To address the issues of low management efficiency, imprecise data tracking, and cumbersome rapid decision-making in traditional grassland animal husbandry, this study proposes a livestock monitoring model based on biometric technology. The study integrates sheep face recognition and in-depth data analysis, and combines them with deep learning algorithms to achieve non-contact, high-precision individual identification of livestock. A total of 400 sheep were tested in typical grassland areas of Inner Mongolia. The results indicated that the system achieved a comprehensive recognition accuracy of 94.1% and a recall rate of 81%, which effectively resolved the problems of traditional ear tags being prone to loss and damage, while also facilitating the monitoring of livestock diseases. This technology enables the full-lifecycle tracking of livestock, encompassing the intelligent monitoring of health status, the precise management of pastures, and the scientific prediction of breeding cycles. The research provides technical support for the digital transformation of grassland animal husbandry, and particularly offers practical assistance for the sustainable development of ecological animal husbandry in degraded grassland regions.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>YOLOv3</kwd>
        <kwd>Deep Learning</kwd>
        <kwd>Biometric Technology</kwd>
        <kwd>Grassland Livestock Husbandry</kwd>
        <kwd>Sustainable Development</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction</title>
      <p>Dairy sheep and meat sheep in Inner Mongolia are vital components of China’s agricultural sector [<xref ref-type="bibr" rid="B1">1</xref>], and smart animal husbandry serves as a key approach to boosting livestock production efficiency [<xref ref-type="bibr" rid="B2">2</xref>]. Currently, the development of smart pastures requires the implementation of intelligent management, meat product traceability, and epidemic monitoring. However, traditional ear-tag identification suffers from issues such as easy detachment and high labor costs; additionally, individual sheep exhibit high similarity, leading to low accuracy in long-distance identification, which fails to meet practical demands [<xref ref-type="bibr" rid="B3">3</xref>][<xref ref-type="bibr" rid="B4">4</xref>]. Meanwhile, livestock diseases severely compromise the quality and safety of livestock products [<xref ref-type="bibr" rid="B5">5</xref>][<xref ref-type="bibr" rid="B6">6</xref>]. Against this backdrop, biometric detection and recognition solutions based on computer vision technology have become critical. These solutions not only can replace ear tags as essential pasture infrastructure but also record livestock growth data to ensure animal health, facilitate rapid monitoring of herd health status and female livestock reproduction, and provide support for modern livestock production [<xref ref-type="bibr" rid="B4">4</xref>]. Therefore, individual identification is the core focus of this study. </p>
      <p>Driven by deep learning, biometric technology has achieved breakthroughs in multiple fields, evolving from traditional manual feature extraction to an efficient recognition system centered on convolutional neural networks (CNN) [<xref ref-type="bibr" rid="B7">7</xref>][<xref ref-type="bibr" rid="B8">8</xref>] and transfer learning. It has demonstrated remarkable effectiveness in fields such as face recognition, iris recognition, and gait recognition [<xref ref-type="bibr" rid="B9">9</xref>]. Specifically, CNN enhance the adaptability of iris recognition to open scenarios [<xref ref-type="bibr" rid="B10">10</xref>], while multi-network fusion improves the performance of gait recognition under low-resolution conditions. In the agricultural sector, this technology offers a new pathway for grassland herd management: by combining YOLOv3 with DeepSORT tracking technology [<xref ref-type="bibr" rid="B11">11</xref>]-[<xref ref-type="bibr" rid="B13">13</xref>], real-time sheep detection and individual identity binding are realized, addressing the problems of traditional ear-tag identification (e.g., easy detachment, high costs, and poor real-time monitoring). Nevertheless, sheep face recognition still has room for optimization, including issues like adaptability to complex grassland environments, insufficient multi-modal feature fusion, and hardware compatibility. In terms of technology integration, multi-modal biometric recognition has emerged as a research hotspot—for instance, multi-modal fusion of hand features enhances recognition robustness [<xref ref-type="bibr" rid="B10">10</xref>], and deep learning-based analysis of biometric information improves recognition accuracy. </p>
      <p>Looking ahead, deep learning will continue to drive the development of biometric technology [<xref ref-type="bibr" rid="B14">14</xref>], providing new ideas for various fields. With technological advancements and expanded application scenarios, biometric technology will achieve higher levels of intelligence and precision, contributing to the development of multiple industries. Modern grassland animal husbandry requires real-time monitoring of individual livestock to update their health status, thereby establishing a real-time IoT-based monitoring big data set. This big data set enables early warning of abnormal herd conditions, prediction of herd productivity, and data-supported health management. However, accurate individual identification of herds remains the most fundamental task to be accomplished [<xref ref-type="bibr" rid="B15">15</xref>]. Consequently, this study explores the individual identification of grassland herds based on image recognition technology.</p>
    </sec>
    <sec id="sec2">
      <title>2. Basic Principles</title>
      <p>Deep learning is a key AI technology with outstanding performance in image classification, speech recognition and other fields but complex to achieve. Built on four advanced technologies, it is implemented in many fields using open-source databases and PaddlePaddle. Convolutional neural networks are important deep learning models, with core modules including convolution and pooling. The convolutional layer extracts image features via operations, serving as feature extraction core; the pooling layer divides images into non-overlapping blocks and samples elements. Additionally, the model relies on ReLU activation function, batch normalization and dropout to jointly ensure recognition performance.</p>
      <sec id="sec2dot1">
        <title>2.1. Object Detection and Recognition</title>
        <p>Using convolutional neural networks, two detection algorithms can be chosen for image object detection and classification: one stage detection with fast speed and low accuracy, and two stage detection with slow speed but high accuracy. This time, a two-stage Retina Net algorithm is used to recognize the target image. This algorithm adopts a fusion algorithm of ResNet, FPN, and FCN, and uses Focal loss as the loss function. Focal loss mainly solves the problem of target detection loss being affected by negative samples due to the imbalance of positive and negative sample areas during the object detection process. Firstly, classify the image (<xref ref-type="fig" rid="fig1">Figure 1(a)</xref>) to identify it as a picture of a sheep. Then, perform object detection (<xref ref-type="fig" rid="fig1">Figure 1(b)</xref>) to recognize it as a photo of a sheep and mark the position of each sheep in the image.</p>
        <fig id="fig1">
          <label>Figure 1</label>
          <graphic xlink:href="https://html.scirp.org/file/2313475-rId15.jpeg?20260430040550" />
        </fig>
        <p><bold>Figure 1</bold><bold>.</bold> Image classification (a) and object detection (b).</p>
      </sec>
      <sec id="sec2dot2">
        <title>2.2. Detection Object Target Box</title>
        <p>In convolutional neural networks, the preferred region needs to be determined, and we use exhaustive search to generate candidate regions. The A pixel on the image and the B pixel in the lower right corner of A can determine a rectangular box, denoted as AB. A is located in the upper left corner of the image, and B traverses all positions except A to generate rectangular boxes A1 B1, ..., A1 Bn (<xref ref-type="fig" rid="fig2">Figure 2(a)</xref>). After obtaining a certain position of A in the middle of the image, B traverses all positions in the lower right corner of A and generates rectangular boxes Ak B1, ..., Ak Bn (<xref ref-type="fig" rid="fig2">Figure 2(b)</xref>). When A traverses all the pixels on the image, B traverses all the pixels in the lower right corner, and finally generates a set of rectangular boxes {AiBj}, which includes all the cocoa selection areas on the image.</p>
        <fig id="fig2">
          <label>Figure 2</label>
          <graphic xlink:href="https://html.scirp.org/file/2313475-rId16.jpeg?20260430040550" />
        </fig>
        <p><bold>Figure 2</bold><bold>.</bold> Candidate area confirmation.</p>
        <p>As long as the classification of each candidate region is accurate enough, it is certain that a region that is close enough to the actual object can be found. The exhaustive search method may yield accurate prediction results, but its computational complexity is also enormous. Assuming H = W = 100, the total number will reach 2.5 × 107 points. Assuming the classification is fine enough, exhaustive search can theoretically complete the detection task, but it requires designing bounding boxes, anchor boxes, and intersection to union ratios for object detection to accurately provide candidate regions.</p>
        <disp-formula id="FD1">
          <label>(1)</label>
          <mml:math>
            <mml:mrow>
              <mml:mfrac>
                <mml:mrow>
                  <mml:msup>
                    <mml:mtext>W</mml:mtext>
                    <mml:mtext>2</mml:mtext>
                  </mml:msup>
                  <mml:msup>
                    <mml:mi>H</mml:mi>
                    <mml:mn>2</mml:mn>
                  </mml:msup>
                </mml:mrow>
                <mml:mtext>4</mml:mtext>
              </mml:mfrac>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>Boundary box: Boundary boxes are often used to represent the specific position of objects and can be inserted into rectangular boxes of objects. It can be inferred that boundary boxes are a detection task that simultaneously predicts and detects the position and category of objects. There are usually two forms of bounding boxes: xyxy (x<sub>1</sub>, y<sub>1</sub>, x<sub>2</sub>, y<sub>2</sub>) and xywh (x, y, w, h).</p>
        <p>In detection tasks, bounding boxes are called real boxes because the labels on the training dataset provide the coordinates of the real bounding box of the target object, which are (x<sub>1</sub>, y<sub>1</sub>, x<sub>2</sub>, y<sub>2</sub>). The model accurately predicts the possible positions of the target object, and the bounding boxes predicted by the model are called prediction boxes [<xref ref-type="bibr" rid="B16">16</xref>].</p>
        <p>Anchor box: Anchor box is a type of target box imagined by people. Draw an anchor box of a specific size, find a center point, and draw a rectangular box.</p>
        <p>In the detection task, an anchor box is formed on the image, and it is determined whether the target object is included before proceeding to the next step of object detection. Due to the mismatch between the object and the anchor box, it is necessary to make fine adjustments to the original anchor box to form an anchor box that can accurately obtain the position of the object. The model needs to predict the magnitude of the fine adjustment, and different models have different methods for generating anchor boxes and different micro adjustment amplitudes, because the anchor box position is fixed at the beginning.</p>
        <p>Intersection to union ratio: Intersection to union ratio is a measurement indicator in object detection. In object detection tasks, the intersection to union ratio is used as a measurement indicator in object detection, and the specific calculation formula is as follows:</p>
        <disp-formula id="FD2">
          <label>(2)</label>
          <mml:math>
            <mml:mrow>
              <mml:msub>
                <mml:mi>I</mml:mi>
                <mml:mi>O</mml:mi>
              </mml:msub>
              <mml:mi>U</mml:mi>
              <mml:mo>=</mml:mo>
              <mml:mfrac>
                <mml:mrow>
                  <mml:mi>A</mml:mi>
                  <mml:mo>∩</mml:mo>
                  <mml:mi>B</mml:mi>
                </mml:mrow>
                <mml:mrow>
                  <mml:mi>A</mml:mi>
                  <mml:mo>∪</mml:mo>
                  <mml:mi>B</mml:mi>
                </mml:mrow>
              </mml:mfrac>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>The green area in the “Intersection” section of the figure represents the overlapping area of two boxes, while the green area in the “Merge” section represents the merged area of two boxes. Dividing these two areas yields the intersection to union ratio between them, which is also known as the intersection to union ratio (<xref ref-type="fig" rid="fig3">Figure 3</xref>).</p>
        <fig id="fig3">
          <label>Figure 3</label>
          <graphic xlink:href="https://html.scirp.org/file/2313475-rId21.jpeg?20260430040550" />
        </fig>
        <p><bold>Figure 3</bold><bold>.</bold> Comparison of Intersection and Intersection Ratio Operations.</p>
        <p>If the positions of these two rectangular boxes A and B are A (x<sub>a1</sub><sub>’</sub>y<sub>a1</sub>, x<sub>a2</sub><sub>’</sub>y<sub>a2</sub>) and B (x<sub>b1</sub><sub>’</sub>y<sub>b1</sub>, x<sub>b2</sub><sub>’</sub>y<sub>b2</sub>), respectively, as shown in <xref ref-type="fig" rid="fig3">Figure 3(d)</xref>: If there is an intersection between the two, the coordinates of the upper left corner of the intersection are:</p>
        <disp-formula id="FD3">
          <label>(3)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:msub>
                <mml:mtext>x</mml:mtext>
                <mml:mn>1</mml:mn>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:mi>max</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mtext>x</mml:mtext>
                    <mml:mrow>
                      <mml:mtext>a</mml:mtext>
                      <mml:mn>1</mml:mn>
                    </mml:mrow>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mtext>x</mml:mtext>
                    <mml:mrow>
                      <mml:mtext>b</mml:mtext>
                      <mml:mn>1</mml:mn>
                    </mml:mrow>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>,</mml:mo>
              <mml:msub>
                <mml:mtext>y</mml:mtext>
                <mml:mn>2</mml:mn>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:mi>max</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mtext>y</mml:mtext>
                    <mml:mrow>
                      <mml:mtext>a</mml:mtext>
                      <mml:mn>1</mml:mn>
                    </mml:mrow>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mtext>y</mml:mtext>
                    <mml:mrow>
                      <mml:mtext>b</mml:mtext>
                      <mml:mn>1</mml:mn>
                    </mml:mrow>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>The coordinates of the lower right corner of the intersection are:</p>
        <disp-formula id="FD4">
          <label>(4)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:msub>
                <mml:mtext>x</mml:mtext>
                <mml:mn>2</mml:mn>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:mi>min</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mtext>x</mml:mtext>
                    <mml:mrow>
                      <mml:mtext>a</mml:mtext>
                      <mml:mn>2</mml:mn>
                    </mml:mrow>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mtext>x</mml:mtext>
                    <mml:mrow>
                      <mml:mtext>b</mml:mtext>
                      <mml:mn>2</mml:mn>
                    </mml:mrow>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>,</mml:mo>
              <mml:msub>
                <mml:mtext>y</mml:mtext>
                <mml:mn>2</mml:mn>
              </mml:msub>
              <mml:mo>=</mml:mo>
              <mml:mi>min</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mtext>y</mml:mtext>
                    <mml:mrow>
                      <mml:mtext>a</mml:mtext>
                      <mml:mn>2</mml:mn>
                    </mml:mrow>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mtext>y</mml:mtext>
                    <mml:mrow>
                      <mml:mtext>b</mml:mtext>
                      <mml:mn>2</mml:mn>
                    </mml:mrow>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>Calculate the area of the first part to be delivered:</p>
        <disp-formula id="FD5">
          <label>(5)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:mtext>interection</mml:mtext>
              <mml:mo>=</mml:mo>
              <mml:mi>max</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mtext>x</mml:mtext>
                    <mml:mn>2</mml:mn>
                  </mml:msub>
                  <mml:mo>−</mml:mo>
                  <mml:msub>
                    <mml:mtext>x</mml:mtext>
                    <mml:mn>1</mml:mn>
                  </mml:msub>
                  <mml:mo>+</mml:mo>
                  <mml:mn>1.0</mml:mn>
                  <mml:mo>,</mml:mo>
                  <mml:mn>0</mml:mn>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>⋅</mml:mo>
              <mml:mi>max</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mtext>y</mml:mtext>
                    <mml:mn>2</mml:mn>
                  </mml:msub>
                  <mml:mo>−</mml:mo>
                  <mml:msub>
                    <mml:mtext>y</mml:mtext>
                    <mml:mn>1</mml:mn>
                  </mml:msub>
                  <mml:mo>+</mml:mo>
                  <mml:mn>1.0</mml:mn>
                  <mml:mo>,</mml:mo>
                  <mml:mn>0</mml:mn>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>The area of rectangular boxes A and B is:</p>
        <disp-formula id="FD6">
          <label>(6)</label>
          <mml:math display="inline">
            <mml:mtable columnalign="left">
              <mml:mtr>
                <mml:mtd>
                  <mml:msub>
                    <mml:mtext>s</mml:mtext>
                    <mml:mi>A</mml:mi>
                  </mml:msub>
                  <mml:mo>=</mml:mo>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mtext>x</mml:mtext>
                        <mml:mrow>
                          <mml:mtext>a</mml:mtext>
                          <mml:mn>2</mml:mn>
                        </mml:mrow>
                      </mml:msub>
                      <mml:mo>−</mml:mo>
                      <mml:msub>
                        <mml:mtext>x</mml:mtext>
                        <mml:mrow>
                          <mml:mtext>a</mml:mtext>
                          <mml:mn>1</mml:mn>
                        </mml:mrow>
                      </mml:msub>
                      <mml:mo>+</mml:mo>
                      <mml:mn>1.0</mml:mn>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mo>⋅</mml:mo>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mtext>y</mml:mtext>
                        <mml:mrow>
                          <mml:mtext>a</mml:mtext>
                          <mml:mn>1</mml:mn>
                        </mml:mrow>
                      </mml:msub>
                      <mml:mo>−</mml:mo>
                      <mml:msub>
                        <mml:mtext>y</mml:mtext>
                        <mml:mrow>
                          <mml:mtext>a</mml:mtext>
                          <mml:mn>2</mml:mn>
                        </mml:mrow>
                      </mml:msub>
                      <mml:mo>+</mml:mo>
                      <mml:mn>1.0</mml:mn>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mtd>
              </mml:mtr>
              <mml:mtr>
                <mml:mtd>
                  <mml:msub>
                    <mml:mtext>s</mml:mtext>
                    <mml:mi>B</mml:mi>
                  </mml:msub>
                  <mml:mo>=</mml:mo>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mtext>x</mml:mtext>
                        <mml:mrow>
                          <mml:mtext>b</mml:mtext>
                          <mml:mn>2</mml:mn>
                        </mml:mrow>
                      </mml:msub>
                      <mml:mo>−</mml:mo>
                      <mml:msub>
                        <mml:mtext>x</mml:mtext>
                        <mml:mrow>
                          <mml:mtext>b</mml:mtext>
                          <mml:mn>1</mml:mn>
                        </mml:mrow>
                      </mml:msub>
                      <mml:mo>+</mml:mo>
                      <mml:mn>1.0</mml:mn>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mo>⋅</mml:mo>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mtext>y</mml:mtext>
                        <mml:mrow>
                          <mml:mtext>b</mml:mtext>
                          <mml:mn>1</mml:mn>
                        </mml:mrow>
                      </mml:msub>
                      <mml:mo>−</mml:mo>
                      <mml:msub>
                        <mml:mtext>y</mml:mtext>
                        <mml:mrow>
                          <mml:mtext>b</mml:mtext>
                          <mml:mn>2</mml:mn>
                        </mml:mrow>
                      </mml:msub>
                      <mml:mo>+</mml:mo>
                      <mml:mn>1.0</mml:mn>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mtd>
              </mml:mtr>
            </mml:mtable>
          </mml:math>
        </disp-formula>
        <p>Calculate the combined area:</p>
        <disp-formula id="FD7">
          <label>(7)</label>
          <mml:math>
            <mml:mrow>
              <mml:mtext>union</mml:mtext>
              <mml:mo>=</mml:mo>
              <mml:msub>
                <mml:mtext>s</mml:mtext>
                <mml:mi>A</mml:mi>
              </mml:msub>
              <mml:mo>+</mml:mo>
              <mml:msub>
                <mml:mtext>s</mml:mtext>
                <mml:mi>B</mml:mi>
              </mml:msub>
              <mml:mo>−</mml:mo>
              <mml:mtext>intersection</mml:mtext>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>Calculate the intersection to union ratio:</p>
        <disp-formula id="FD8">
          <label>(8)</label>
          <mml:math>
            <mml:mrow>
              <mml:msub>
                <mml:mi>I</mml:mi>
                <mml:mi>O</mml:mi>
              </mml:msub>
              <mml:mi>U</mml:mi>
              <mml:mo>=</mml:mo>
              <mml:mfrac>
                <mml:mrow>
                  <mml:mtext>intersection</mml:mtext>
                </mml:mrow>
                <mml:mrow>
                  <mml:mtext>union</mml:mtext>
                </mml:mrow>
              </mml:mfrac>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>In order to clearly demonstrate the relationship between the size of the intersection ratio and the degree of overlap, <xref ref-type="fig" rid="fig4">Figure 4</xref> shows the corresponding positional relationship between two target boxes under different intersection ratios, ranging from <italic>I</italic><italic><sub>o</sub></italic><italic>U</italic> = 0.95 to <italic>I</italic><italic><sub>o</sub></italic><italic>U</italic> = 0.00.</p>
        <p>R-CNN generates candidate regions through selective search, segments the image, and merges similar regions. The candidate boxes are uniformly adjusted to 227 × 227 and input into a CNN (such as AlexNet) to extract 4096 dimensional features. SVM is used for classification and non-maximum suppression is used to remove overlapping detection boxes, but the process is complex and time-consuming. YOLO divides the image into N × N grids, and directly predicts the bounding box position and category probability for each grid, achieving end-to-end detection and greatly improving speed. SSD is improved based on VGG16 by adding multi-scale convolutional layers, using default boxes for detection on different feature layers, and combining multiple aspect ratio presets to achieve adaptability to targets of different sizes. The three represent the technological evolution paths from regional proposal, single-stage detection to multi-scale prediction.</p>
        <fig id="fig4">
          <label>Figure 4</label>
          <graphic xlink:href="https://html.scirp.org/file/2313475-rId34.jpeg?20260430040550" />
        </fig>
        <p><bold>Figure 4</bold><bold>.</bold> Relative position diagram between two boxes under different intersection and merger ratios.</p>
      </sec>
      <sec id="sec2dot3">
        <title>2.3. Target Detection Evaluation</title>
        <p>In order to further evaluate the performance of the sheep detection model and select evaluation indicators for the experimental results, including recall rate <italic>R</italic>, accuracy rate <italic>P</italic>, mean average precision mAP, and average precision <italic>AP</italic>. The recall rate <italic>R</italic> is an evaluation criterion that reflects the incompleteness of the target detected by a model; Accuracy <italic>P</italic> is an evaluation criterion that reflects the accuracy of a model’s predictions; <italic>AP</italic> is the area under the Precision call curve, and generally speaking, the better the classifier, the higher the <italic>AP</italic> value. MAP is the average of multiple categories of <italic>AP</italic>. Mean is the average of the <italic>AP</italic> for each class, resulting in mAP. The size of mAP must be in the range of [0, 1], the larger the better. The calculation formulas for <italic>P</italic>, <italic>R</italic>, <italic>AP</italic>, and mAP are shown in equations (9), (10), (11), and (12), respectively.</p>
        <disp-formula id="FD9">
          <label>(9)</label>
          <mml:math>
            <mml:mrow>
              <mml:mi>P</mml:mi>
              <mml:mo>=</mml:mo>
              <mml:mfrac>
                <mml:mrow>
                  <mml:mi>T</mml:mi>
                  <mml:mi>P</mml:mi>
                </mml:mrow>
                <mml:mrow>
                  <mml:mi>T</mml:mi>
                  <mml:mi>P</mml:mi>
                  <mml:mo>+</mml:mo>
                  <mml:mi>F</mml:mi>
                  <mml:mi>P</mml:mi>
                </mml:mrow>
              </mml:mfrac>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <disp-formula id="FD10">
          <label>(10)</label>
          <mml:math>
            <mml:mrow>
              <mml:mi>R</mml:mi>
              <mml:mo>=</mml:mo>
              <mml:mfrac>
                <mml:mrow>
                  <mml:mi>T</mml:mi>
                  <mml:mi>P</mml:mi>
                </mml:mrow>
                <mml:mrow>
                  <mml:mi>T</mml:mi>
                  <mml:mi>P</mml:mi>
                  <mml:mo>+</mml:mo>
                  <mml:mi>F</mml:mi>
                  <mml:mi>N</mml:mi>
                </mml:mrow>
              </mml:mfrac>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <disp-formula id="FD11">
          <label>(11)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:mi>A</mml:mi>
              <mml:mi>P</mml:mi>
              <mml:mo>=</mml:mo>
              <mml:mstyle displaystyle="true">
                <mml:mrow>
                  <mml:msubsup>
                    <mml:mo>∫</mml:mo>
                    <mml:mn>0</mml:mn>
                    <mml:mn>1</mml:mn>
                  </mml:msubsup>
                  <mml:mrow>
                    <mml:mi>P</mml:mi>
                    <mml:mrow>
                      <mml:mo>(</mml:mo>
                      <mml:mi>R</mml:mi>
                      <mml:mo>)</mml:mo>
                    </mml:mrow>
                    <mml:mi>d</mml:mi>
                    <mml:mrow>
                      <mml:mo>(</mml:mo>
                      <mml:mi>R</mml:mi>
                      <mml:mo>)</mml:mo>
                    </mml:mrow>
                  </mml:mrow>
                </mml:mrow>
              </mml:mstyle>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <disp-formula id="FD12">
          <label>(12)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:mtext>mAP</mml:mtext>
              <mml:mo>=</mml:mo>
              <mml:mfrac>
                <mml:mn>1</mml:mn>
                <mml:mi>m</mml:mi>
              </mml:mfrac>
              <mml:mstyle displaystyle="true">
                <mml:msubsup>
                  <mml:mo>∑</mml:mo>
                  <mml:mrow>
                    <mml:mi>i</mml:mi>
                    <mml:mo>=</mml:mo>
                    <mml:mn>1</mml:mn>
                  </mml:mrow>
                  <mml:mi>m</mml:mi>
                </mml:msubsup>
                <mml:mrow>
                  <mml:mi>A</mml:mi>
                  <mml:msub>
                    <mml:mi>P</mml:mi>
                    <mml:mi>i</mml:mi>
                  </mml:msub>
                  <mml:malignmark>
                  </mml:malignmark>
                </mml:mrow>
              </mml:mstyle>
            </mml:mrow>
          </mml:math>
        </disp-formula>
      </sec>
      <sec id="sec2dot4">
        <title>2.4. Object Detection YOLOv3 Algorithm</title>
        <p>Data preprocessing is necessary before the YOLOv3 algorithm, and it is a crucial and primary step in training convolutional neural network structures. Firstly, selecting appropriate data preprocessing methods can prevent overfitting, and then implementing data reading and preprocessing can help accelerate processing. The biometric technology based on YOLOv3 first normalizes the input biometric image with a size of 416 × 416, and then extracts multi-scale features through the Darknet-53 backbone network, generating feature maps at three scales: 13 × 13, 26 × 26, and 52 × 52. Using Feature Pyramid Network (FPN) to fuse deep and shallow features, three prior boxes are pre-set at each grid point to predict boundary coordinates and category probabilities. Finally, the non maximum suppression (NMS) algorithm is used to optimize the detection results, eliminate redundant boxes, and retain the optimal prediction, forming an end-to-end recognition pipeline from feature extraction to target localization, providing reliable technical support for individual identity recognition.</p>
        <p>Data retrieval is the process of storing all descriptive information of an image in records, where each element contains a description of the image. The subsequent program demonstrates how to retrieve the image and annotate it based on the description in the records. Data preprocessing is the random processing of images, but with minimal changes. The main purpose is to suppress overfitting and improve the generalization ability of the model by increasing the training dataset. Common methods include randomly changing brightness, contrast, and color, random filling, random cropping, random scaling, random flipping, random scrambling of the real box arrangement order, and using numpy to implement these data augmentation methods.</p>
        <p>This article standardizes the animal and plant dataset based on the Pascal VOC 2007 set format and uses the LableImg tool for animal and plant image annotation. Use the VOC_ label. py script to navigate this path to your own dataset, with the class selected as “face”. Finally, running the script will generate files such as train.txt, val.txt, test.txt, etc., and generate a labels file in devkit. Change the model parameters of YOLOv3 and adjust the batch. When batch equals 64, it’s enough, then adjust other parameters to end. The parameter settings are shown in <bold>Table 1</bold>. If the average loss of the result is less than 0.0607, the operation will be stopped at this time.</p>
        <p><bold>Table 1</bold><bold>.</bold> Parameter settings for YOLOv3.</p>
        <table-wrap id="tbl1">
          <label>Table 1</label>
          <table>
            <tbody>
              <tr>
                <td>Parameter</td>
                <td>Value</td>
                <td>Parameter</td>
                <td>Value</td>
              </tr>
              <tr>
                <td>Batch</td>
                <td>64</td>
                <td>Exposure</td>
                <td>1.5</td>
              </tr>
              <tr>
                <td>Subdivisions</td>
                <td>2</td>
                <td>Hue</td>
                <td>0.1</td>
              </tr>
              <tr>
                <td>Width</td>
                <td>416</td>
                <td>learning_rate</td>
                <td>0.001</td>
              </tr>
              <tr>
                <td>Height</td>
                <td>416</td>
                <td>Burn_in</td>
                <td>400</td>
              </tr>
              <tr>
                <td>Channels</td>
                <td>3</td>
                <td>Max_batches</td>
                <td>5200</td>
              </tr>
              <tr>
                <td>Momentum</td>
                <td>0.9</td>
                <td>Policy</td>
                <td>Steps</td>
              </tr>
              <tr>
                <td>Decay</td>
                <td>0.0005</td>
                <td>Steps</td>
                <td>3800</td>
              </tr>
              <tr>
                <td>Angle</td>
                <td>0</td>
                <td>Scales</td>
                <td>0.1</td>
              </tr>
              <tr>
                <td>Saturation</td>
                <td>1.5</td>
                <td>
                </td>
                <td>
                </td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Adopting the Triplet loss proposed in FaceNet as the loss function. The Triplet Loss function consists of three images: the reference image (anchor), the positive example image (positive), and the negative example image (negative). Through continuous training, there is a small gap between the positive example image and the negative example image, but a large gap between them. The selected VGGFace16 and resNet50 represent two types of convolutional network structures with different network layers (16 layers and 50 layers), respectively, combined with two embeddings (2096 and 128) to construct a sheep face recognition model.</p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. Result Analysis</title>
      <sec id="sec3dot1">
        <title>3.1. Dataset Preprocessing Results</title>
        <p>3.1.1. Individual Labeling of Livestock Herds </p>
        <p>The dataset of this article mainly comes from photos of livestock herds in Yangchang Village, Zhaojun Town, Dalate Banner, Ordos. LabelImg software was used to label the above data as sheep, including pictures of sheep in various scenarios, totaling 400 images. Divide the annotated dataset into training, testing, and validation sets in an 8:1:1 ratio, resulting in 320 training sets, 40 testing sets, and 40 validation sets, respectively. Finally, the dataset was organized to obtain its label situation as shown in <xref ref-type="fig" rid="fig5">Figure 5</xref><bold>:</bold>(a) The y-axis of the figure represents the number of labels, and the x-axis represents the type of labels. (b) In the figure, the horizontal axis represents the ratio of label width to image width, and the vertical axis represents the ratio of label height to image height.</p>
        <fig id="fig5">
          <label>Figure 5</label>
          <graphic xlink:href="https://html.scirp.org/file/2313475-rId43.jpeg?20260430040552" />
        </fig>
        <p><bold>Figure 5</bold><bold>.</bold> Distribution of sheep face labeling labels (a) and aspect ratio (b).</p>
        <p>3.1.2. Individual Tracking of Livestock Herds</p>
        <p>The Deepsort object tracking algorithm is based on detectors, such as the YOLO object detector. Firstly, the object detector detects sheep in the image and passes the detection coordinates of the sheep to the Deepsort algorithm. The Deepsort algorithm predicts the target position of the sheep at the next moment based on the Kalman filter, and finally completes the target tracking based on the Hungarian algorithm (<xref ref-type="fig" rid="fig6">Figure 6</xref>), and assigns a unique ID number to each target.</p>
        <fig id="fig6">
          <label>Figure 6</label>
          <graphic xlink:href="https://html.scirp.org/file/2313475-rId44.jpeg?20260430040553" />
        </fig>
        <p><bold>Figure 6</bold><bold>.</bold> Individual ID tracking of livestock herd.</p>
      </sec>
      <sec id="sec3dot2">
        <title>3.2. Sheep Face Recognition Evaluation</title>
        <p>3.2.1. Target Detection Model Training</p>
        <p>The experiment used a self-built sheep dataset image, with sheep as the sample. The YOLOv3 network was trained using end-to-end stochastic gradient descent method, and the input image size was 640 × 640. The parameter values are set as follows: batch size is 64, momentum value is 0.937, decay value is 0.0005, total rounds are 100, and initial learning rate is 0.01. When one round of training is completed, the model training situation is validated using the validation set. The loss curves of the training set and validation set (<xref ref-type="fig" rid="fig7">Figure 7</xref>) show that from the 0th to the 50th batch on the training set, the loss of the training set decreases during training, and from the 50th to the 100th batch, the loss of the training set tends to stabilize, and the model gradually converges. Around the 25th batch on the validation set, the model gradually converges, and around the 50th batch, the Object loss shows an upward trend, indicating that the model may be overfitting.</p>
        <fig id="fig7">
          <label>Figure 7</label>
          <graphic xlink:href="https://html.scirp.org/file/2313475-rId45.jpeg?20260430040554" />
        </fig>
        <fig id="fig8">
          <label>Figure 8</label>
          <graphic xlink:href="https://html.scirp.org/file/2313475-rId46.jpeg?20260430040554" />
        </fig>
        <fig id="fig9">
          <label>Figure 9</label>
          <graphic xlink:href="https://html.scirp.org/file/2313475-rId47.jpeg?20260430040553" />
        </fig>
        <fig id="fig10">
          <label>Figure 10</label>
          <graphic xlink:href="https://html.scirp.org/file/2313475-rId48.jpeg?20260430040554" />
        </fig>
        <fig id="fig11">
          <label>Figure 11</label>
          <graphic xlink:href="https://html.scirp.org/file/2313475-rId49.jpeg?20260430040553" />
        </fig>
        <fig id="fig12">
          <label>Figure 12</label>
          <graphic xlink:href="https://html.scirp.org/file/2313475-rId50.jpeg?20260430040554" />
        </fig>
        <p><bold>Figure 7</bold><bold>.</bold> Loss curves of training and validation sets.</p>
        <p>The performance curve of the validation set shows that mAP is the area enclosed by the precision Pr and recall Re plotted as two axes, and mAP and mAP are the average detection accuracy under general threshold and high threshold, respectively (<xref ref-type="fig" rid="fig8">Figure 8</xref>). When the values corresponding to the curve gradually stabilize, the optimal training model can be determined. In <xref ref-type="fig" rid="fig8">Figure 8</xref>, the top and bottom are the loss curves of the training set and validation set, respectively. It can be seen that the three types of loss functions have a significant decrease in the model training from 0 to 50 times, and remain stable at 50 to 100 times, indicating that training improves the detection performance of the model.</p>
        <p>The mAP curve fluctuates significantly from 0 to 50 times at the beginning of training, indicating that the convergence speed of the model in the early stage of training is fast and meets the requirements of model training. After 50 iterations, the model remained stable with minimal changes, indicating good training and no overfitting. The curve around 90 to 100 tends to stabilize, indicating that the training of the sheep detection model is basically completed at this time (<xref ref-type="fig" rid="fig8">Figure 8</xref>). The selection of the optimal model is calculated based on mAP-U accounting for 10% and mAP-H accounting for 90%. The accuracy and recall were 94.1% and 81%, respectively, and the mAP reached 63.3%.</p>
        <fig id="fig13">
          <label>Figure 13</label>
          <graphic xlink:href="https://html.scirp.org/file/2313475-rId51.jpeg?20260430040554" />
        </fig>
        <fig id="fig14">
          <label>Figure 14</label>
          <graphic xlink:href="https://html.scirp.org/file/2313475-rId52.jpeg?20260430040553" />
        </fig>
        <fig id="fig15">
          <label>Figure 15</label>
          <graphic xlink:href="https://html.scirp.org/file/2313475-rId53.jpeg?20260430040553" />
        </fig>
        <fig id="fig16">
          <label>Figure 16</label>
          <graphic xlink:href="https://html.scirp.org/file/2313475-rId54.jpeg?20260430040553" />
        </fig>
        <fig id="fig17">
          <label>Figure 17</label>
          <graphic xlink:href="https://html.scirp.org/file/2313475-rId55.jpeg?20260430040553" />
        </fig>
        <p><bold>Figure 8</bold><bold>.</bold> Verification set performance curve.</p>
        <p>3.2.2. Accuracy Evaluation of Sheep Face Recognition</p>
        <p>After training the sheep face detection model with YOLO v3 object detection, computer vision technology is used to read the video stream, perform sheep face recognition and detection, and save the sheep face information in the detection box. After a series of tests, the accuracy is 94.1% and the recall rate is 81%. Real time detection of sheep face data has been achieved. The training data for sheep face detection includes sheep with different postures (<xref ref-type="fig" rid="fig9">Figure 9</xref>), different lighting conditions, and camera shake, indicating that the model can accurately detect sheep under different movements and angles, meeting the needs of sheep face detection.</p>
        <fig id="fig18">
          <label>Figure 18</label>
          <graphic xlink:href="https://html.scirp.org/file/2313475-rId56.jpeg?20260430040554" />
        </fig>
        <p><bold>Figure 9</bold><bold>.</bold> Sheep face detection results ((a) front view; (b) side view).</p>
        <p>Divide the sheep image dataset of different scenes into training set, testing set, and validation set in an 8:1:1 ratio, and conduct model training based on this division. When the performance test results of the model using VGGFace neural network and embedding dimension 2096 are obtained, the performance of this model is relatively high. Using a specific model to recognize sheep faces in different situations within the same sheepfold. The use of transfer method to train models with a small number of samples can achieve high accuracy and recall, and can better complete the task of recognizing frontal sheep faces in a small dataset.</p>
      </sec>
    </sec>
    <sec id="sec4">
      <title>4. Discussion</title>
      <p>In this study on sheep face recognition models, nose stripe recognition and comparative analysis between the two should be included in order to better and more accurately recognize sheep, improve accuracy, and achieve the ideal state. In addition to methods such as data collection, data cleaning, model construction, model training, and deployment, methods to improve recognition accuracy should also consider lighting conditions to make recognition more accurate. Generally speaking, conventional recognition systems require uniform illumination of light in the recognition area, without scattered light such as shadows and flashes. Some high-end products have reduced lighting requirements, but their costs are relatively high. Therefore, in order to improve the accuracy of sheep face recognition, it is necessary to add supplementary lighting equipment in situations with poor lighting conditions. There is also a hardware factor, which refers to the performance of the camera and control motherboard in the recognition system. The commonly used recognition camera pixels are between 2 million and 4 million pixels, not necessarily the higher the pixel, the better.</p>
    </sec>
    <sec id="sec5">
      <title>5. Conclusion</title>
      <p>The sheep face detection dataset for training scenarios is divided into training set, testing set, and validation set in an 8:1:1 ratio, and the results meet the requirements. YOLOv3 accuracy and recall display can achieve real-time monitoring of sheep faces under natural light. In smaller datasets and with lower threshold for devices (hardware camera pixels of 2 - 4 million), it saves hardware configuration costs and meets the requirements for use in pastoral areas.</p>
    </sec>
    <sec id="sec6">
      <title>Acknowledgments</title>
      <p>We are grateful to the reviewers and editors for their valuable comments and suggestions, which have significantly improved the quality of this manuscript. This work was supported by the Science and Technology Plan Project of Inner Mongolia Autonomous Region (Grant No. 2022YFYZ0009 and 2023YFDZ0078).</p>
    </sec>
    <sec id="sec7">
      <title>Authors’ Contributions</title>
      <p>Guoqing Chen, Lerong Wen, and Xu Zheng designed the experiments and interpreted the results; Xu Zheng wrote the draft; Zheng Xu and Guoqing Chen prepared the figures and reviewed the manuscript; all authors revised the manuscript.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Zhu, M.F. and Cheng, G.Q. (2025) The Realistic Challenges, Key Issues, and Promotion Strategies for the High-Quality Development of China’s Herbivorous Animal Husbandry Industry. <italic>Journal of Social Sciences</italic>, No. 6, 164-175. (In Chinese)</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Zhu, M.F.</string-name>
              <string-name>Cheng, G.Q.</string-name>
              <string-name>Challenges, K</string-name>
              <string-name>Sciences, N</string-name>
            </person-group>
            <year>2025</year>
            <article-title>The Realistic Challenges, Key Issues, and Promotion Strategies for the High-Quality Development of China’s Herbivorous Animal Husbandry Industry</article-title>
            <source>Journal of Social Sciences</source>
            <volume>164</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Yang, C.D., Qi, J.L., Yang, T.H. and Zhang, L.J. (2025) The Spatial Correlation Network and Driving Mechanism for the Coupling and Coordination of Digital Rural Construction and Green Transformation of Animal Husbandry in China. <italic>Research on Agricultural Modernization</italic>, 47, 14-25. (In Chinese)</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Yang, C.D.</string-name>
              <string-name>Qi, J.L.</string-name>
              <string-name>Yang, T.H.</string-name>
              <string-name>Zhang, L.J.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>The Spatial Correlation Network and Driving Mechanism for the Coupling and Coordination of Digital Rural Construction and Green Transformation of Animal Husbandry in China</article-title>
            <source>Research on Agricultural Modernization</source>
            <volume>47</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="thesis">Zhang, W.Y. (2020) Research and Implementation of Smart Ranch Management System. Master’s Thesis, Harbin University of Science and Technology.</mixed-citation>
          <element-citation publication-type="thesis">
            <person-group person-group-type="author">
              <string-name>Zhang, W.Y.</string-name>
              <string-name>Thesis, H</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Research and Implementation of Smart Ranch Management System</article-title>
            <source>Master’s Thesis</source>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Alomair, R., Al-Amoudi, A., Javaid, A., Alnaser, M. and Al binali, S. (2024) Enhancing Precision Livestock Farming Management with AI-Driven Ear Tag Detection and OCR Recognition. 2024 <italic>IEEE International Conference on Technology Management</italic>, <italic>Operations and Decisions</italic> ( <italic>ICTMOD</italic>), Sharjah, 4-6 November 2024, 1-6. https://doi.org/10.1109/ictmod63116.2024.10878196 <pub-id pub-id-type="doi">10.1109/ictmod63116.2024.10878196</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/ictmod63116.2024.10878196">https://doi.org/10.1109/ictmod63116.2024.10878196</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Alomair, R.</string-name>
              <string-name>Al-Amoudi, A.</string-name>
              <string-name>Javaid, A.</string-name>
              <string-name>Alnaser, M.</string-name>
              <string-name>Management, O</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Enhancing Precision Livestock Farming Management with AI-Driven Ear Tag Detection and OCR Recognition</article-title>
            <source>2024 IEEE International Conference on Technology Management</source>
            <volume>4</volume>
            <pub-id pub-id-type="doi">10.1109/ictmod63116.2024.10878196</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Tomley, F.M. and Shirley, M.W. (2009) Livestock Infectious Diseases and Zoonoses. <italic>Philosophical</italic><italic>Transactions</italic><italic>of</italic><italic>the</italic><italic>Royal</italic><italic>Society</italic><italic>B</italic>: <italic>Biological</italic><italic>Sciences</italic>, 364, 2637-2642. https://doi.org/10.1098/rstb.2009.0133 <pub-id pub-id-type="doi">10.1098/rstb.2009.0133</pub-id><pub-id pub-id-type="pmid">19687034</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1098/rstb.2009.0133">https://doi.org/10.1098/rstb.2009.0133</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Tomley, F.M.</string-name>
              <string-name>Shirley, M.W.</string-name>
            </person-group>
            <year>2009</year>
            <article-title>Livestock Infectious Diseases and Zoonoses</article-title>
            <source>Philosophical Transactions of the Royal Society B: Biological Sciences</source>
            <volume>364</volume>
            <pub-id pub-id-type="doi">10.1098/rstb.2009.0133</pub-id>
            <pub-id pub-id-type="pmid">19687034</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B6">
        <label>6.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Ouali, M., Belhouadjeb, F.A., Soufan, W. and Rihan, H.Z. (2023) Sustainability Evaluation of Pastoral Livestock Systems. <italic>Animals</italic>, 13, Article 1335. https://doi.org/10.3390/ani13081335 <pub-id pub-id-type="doi">10.3390/ani13081335</pub-id><pub-id pub-id-type="pmid">37106899</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/ani13081335">https://doi.org/10.3390/ani13081335</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Ouali, M.</string-name>
              <string-name>Belhouadjeb, F.A.</string-name>
              <string-name>Soufan, W.</string-name>
              <string-name>Rihan, H.Z.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Sustainability Evaluation of Pastoral Livestock Systems</article-title>
            <source>Animals</source>
            <volume>13</volume>
            <elocation-id>1335</elocation-id>
            <pub-id pub-id-type="doi">10.3390/ani13081335</pub-id>
            <pub-id pub-id-type="pmid">37106899</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B7">
        <label>7.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Girshick, R. (2015) Fast R-CNN. 2015 <italic>IEEE International Conference on Computer Vision</italic> ( <italic>ICCV</italic>), Santiago, 7-13 December 2015, 1440-1448. https://doi.org/10.1109/iccv.2015.169 <pub-id pub-id-type="doi">10.1109/iccv.2015.169</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/iccv.2015.169">https://doi.org/10.1109/iccv.2015.169</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Girshick, R.</string-name>
            </person-group>
            <year>2015</year>
            <article-title>Fast R-CNN</article-title>
            <source>2015 IEEE International Conference on Computer Vision (ICCV)</source>
            <volume>7</volume>
            <pub-id pub-id-type="doi">10.1109/iccv.2015.169</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B8">
        <label>8.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">He, K., Gkioxari, G., Dollar, P. and Girshick, R. (2017) Mask R-CNN. 2017 <italic>IEEE International Conference on Computer Vision</italic> ( <italic>ICCV</italic>), Venice, 22-29 October 2017, 2980-2988. https://doi.org/10.1109/iccv.2017.322 <pub-id pub-id-type="doi">10.1109/iccv.2017.322</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/iccv.2017.322">https://doi.org/10.1109/iccv.2017.322</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>He, K.</string-name>
              <string-name>Gkioxari, G.</string-name>
              <string-name>Dollar, P.</string-name>
              <string-name>Girshick, R.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Mask R-CNN</article-title>
            <source>2017 IEEE International Conference on Computer Vision (ICCV)</source>
            <volume>22</volume>
            <pub-id pub-id-type="doi">10.1109/iccv.2017.322</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B9">
        <label>9.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Sun, Z., Tian, L., Du, Q., Bhutto, J.A. and Wang, Z. (2022) Feature Learning via Multi-Action Forms Supervising Force for Face Recognition. <italic>Neural</italic><italic>Computing</italic><italic>and</italic><italic>Applications</italic>, 34, 4425-4436. https://doi.org/10.1007/s00521-021-06598-z <pub-id pub-id-type="doi">10.1007/s00521-021-06598-z</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s00521-021-06598-z">https://doi.org/10.1007/s00521-021-06598-z</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Sun, Z.</string-name>
              <string-name>Tian, L.</string-name>
              <string-name>Du, Q.</string-name>
              <string-name>Bhutto, J.A.</string-name>
              <string-name>Wang, Z.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Feature Learning via Multi-Action Forms Supervising Force for Face Recognition</article-title>
            <source>Neural Computing and Applications</source>
            <volume>34</volume>
            <pub-id pub-id-type="doi">10.1007/s00521-021-06598-z</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B10">
        <label>10.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Zhang, X., Song, Y., McCurdy, W., Wang, X. and Zuo, F. (2024) Revising the Problem of Partial Labels from the Perspective of CNNs’ Robustness. 2024 <italic>IEEE</italic>/ <italic>ACIS</italic> 22 <italic>nd International Conference on Software Engineering Research</italic>, <italic>Management and Applications</italic> ( <italic>SERA</italic>), Honolulu, 30 May-1 June 2024, 88-93. https://doi.org/10.1109/sera61261.2024.10685603 <pub-id pub-id-type="doi">10.1109/sera61261.2024.10685603</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1109/sera61261.2024.10685603">https://doi.org/10.1109/sera61261.2024.10685603</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Zhang, X.</string-name>
              <string-name>Song, Y.</string-name>
              <string-name>McCurdy, W.</string-name>
              <string-name>Wang, X.</string-name>
              <string-name>Zuo, F.</string-name>
              <string-name>Research, M</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Revising the Problem of Partial Labels from the Perspective of CNNs’ Robustness</article-title>
            <source>2024 IEEE/ACIS 22nd International Conference on Software Engineering Research</source>
            <volume>30</volume>
            <pub-id pub-id-type="doi">10.1109/sera61261.2024.10685603</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B11">
        <label>11.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Zou, X., Yin, Z., Li, Y., Gong, F., Bai, Y., Zhao, Z., <italic>et al</italic>. (2023) Novel Multiple Object Tracking Method for Yellow Feather Broilers in a Flat Breeding Chamber Based on Improved YOLOv3 and Deep SORT. <italic>International</italic><italic>Journal</italic><italic>of</italic><italic>Agricultural</italic><italic>and</italic><italic>Biological</italic><italic>Engineering</italic>, 16, 44-55. https://doi.org/10.25165/j.ijabe.20231605.7836 <pub-id pub-id-type="doi">10.25165/j.ijabe.20231605.7836</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.25165/j.ijabe.20231605.7836">https://doi.org/10.25165/j.ijabe.20231605.7836</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Zou, X.</string-name>
              <string-name>Yin, Z.</string-name>
              <string-name>Li, Y.</string-name>
              <string-name>Gong, F.</string-name>
              <string-name>Bai, Y.</string-name>
              <string-name>Zhao, Z.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Novel Multiple Object Tracking Method for Yellow Feather Broilers in a Flat Breeding Chamber Based on Improved YOLOv3 and Deep SORT</article-title>
            <source>International Journal of Agricultural and Biological Engineering</source>
            <volume>16</volume>
            <pub-id pub-id-type="doi">10.25165/j.ijabe.20231605.7836</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B12">
        <label>12.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Zhang, C., Liu, X., Shang, Q., Wu, C., Wang, Y., Wang, L., <italic>et al</italic>. (2025) Personnel Target Detection in Infrared Environment Based on YOLOv3-Tinier Network and Its FPGA Implementation. <italic>Infrared</italic><italic>Physics</italic><italic>&amp;</italic><italic>Technology</italic>, 150, Article ID: 106015. https://doi.org/10.1016/j.infrared.2025.106015 <pub-id pub-id-type="doi">10.1016/j.infrared.2025.106015</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.infrared.2025.106015">https://doi.org/10.1016/j.infrared.2025.106015</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Zhang, C.</string-name>
              <string-name>Liu, X.</string-name>
              <string-name>Shang, Q.</string-name>
              <string-name>Wu, C.</string-name>
              <string-name>Wang, Y.</string-name>
              <string-name>Wang, L.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>Personnel Target Detection in Infrared Environment Based on YOLOv3-Tinier Network and Its FPGA Implementation</article-title>
            <source>Infrared Physics &amp; Technology</source>
            <volume>150</volume>
            <fpage>106015</fpage>
            <elocation-id>ID</elocation-id>
            <pub-id pub-id-type="doi">10.1016/j.infrared.2025.106015</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B13">
        <label>13.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Nigam, N., Singh, D.P. and Choudhary, J. (2025) Real-Time Traffic Spatial Occupancy Calculation with Modified YOLOv3 in Complex Environments. <italic>Procedia</italic><italic>Computer</italic><italic>Science</italic>, 260, 717-724. https://doi.org/10.1016/j.procs.2025.03.251 <pub-id pub-id-type="doi">10.1016/j.procs.2025.03.251</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.procs.2025.03.251">https://doi.org/10.1016/j.procs.2025.03.251</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Nigam, N.</string-name>
              <string-name>Singh, D.P.</string-name>
              <string-name>Choudhary, J.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>Real-Time Traffic Spatial Occupancy Calculation with Modified YOLOv3 in Complex Environments</article-title>
            <source>Procedia Computer Science</source>
            <volume>260</volume>
            <pub-id pub-id-type="doi">10.1016/j.procs.2025.03.251</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B14">
        <label>14.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Chuong, V.H., Cuong, V.H., Dat, V.N., Thanh, N.T.C., Thanh, P.T. and Le Quan, N. (2024) Facial Expression Recognition: A Lite Deep Learning-Based Approach. In: Yang, XS., Sherratt, S., Dey, N. and Joshi, A., Eds., <italic>Proceedings of Ninth International Congress on Information and Communication Technology</italic>, Springer, 125-135. https://doi.org/10.1007/978-981-97-3559-4_10 <pub-id pub-id-type="doi">10.1007/978-981-97-3559-4_10</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/978-981-97-3559-4_10">https://doi.org/10.1007/978-981-97-3559-4_10</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Chuong, V.H.</string-name>
              <string-name>Cuong, V.H.</string-name>
              <string-name>Dat, V.N.</string-name>
              <string-name>Thanh, N.T.C.</string-name>
              <string-name>Thanh, P.T.</string-name>
              <string-name>Quan, N.</string-name>
              <string-name>Yang, X</string-name>
              <string-name>Sherratt, S.</string-name>
              <string-name>Dey, N.</string-name>
              <string-name>Joshi, A.</string-name>
              <string-name>Technology, S</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Facial Expression Recognition: A Lite Deep Learning-Based Approach</article-title>
            <source>In: Yang</source>
            <volume>125</volume>
            <pub-id pub-id-type="doi">10.1007/978-981-97-3559-4_10</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B15">
        <label>15.</label>
        <citation-alternatives>
          <mixed-citation publication-type="thesis">Wei, B. (2020) Sheep Face Detection and Recognition Based on Deep Learning. Master’s Thesis, North-West A&amp;F University.</mixed-citation>
          <element-citation publication-type="thesis">
            <person-group person-group-type="author">
              <string-name>Wei, B.</string-name>
              <string-name>Thesis, N</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Sheep Face Detection and Recognition Based on Deep Learning</article-title>
            <source>Master’s Thesis</source>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B16">
        <label>16.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Kolsrud, D. (2015) A Time-Simultaneous Prediction Box for a Multivariate Time Series. <italic>Journal</italic><italic>of</italic><italic>Forecasting</italic>, 34, 675-693. https://doi.org/10.1002/for.2366 <pub-id pub-id-type="doi">10.1002/for.2366</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1002/for.2366">https://doi.org/10.1002/for.2366</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Kolsrud, D.</string-name>
            </person-group>
            <year>2015</year>
            <article-title>A Time-Simultaneous Prediction Box for a Multivariate Time Series</article-title>
            <source>Journal of Forecasting</source>
            <volume>34</volume>
            <pub-id pub-id-type="doi">10.1002/for.2366</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>