<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JCC</journal-id><journal-title-group><journal-title>Journal of Computer and Communications</journal-title></journal-title-group><issn pub-type="epub">2327-5219</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jcc.2025.137018</article-id><article-id pub-id-type="publisher-id">JCC-144438</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Computer Science&amp;Communications</subject></subj-group></article-categories><title-group><article-title>
 
 
  Plunger Pump Fault Diagnosis Method Based on Wavelet Convolution and Multi-Head Self-Attention Mechanism for Multi-Channel Feature Fusion
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Pu</surname><given-names>Zhou</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Zhichun</surname><given-names>Qian</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Dongliang</surname><given-names>Fu</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Yi</surname><given-names>Zhang</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>Shanghai Marine Equipment Research Institute, Shanghai, China</addr-line></aff><pub-date pub-type="epub"><day>02</day><month>07</month><year>2025</year></pub-date><volume>13</volume><issue>07</issue><fpage>345</fpage><lpage>356</lpage><history><date date-type="received"><day>18,</day>	<month>June</month>	<year>2025</year></date><date date-type="rev-recd"><day>27,</day>	<month>July</month>	<year>2025</year>	</date><date date-type="accepted"><day>30,</day>	<month>July</month>	<year>2025</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  To address the issues of low accuracy, high dependence on prior knowledge, and poor adaptability in fusing multi-channel features in existing plunger pump fault diagnosis methods, a new method based on single-channel wavelet convolution and multi-head self-attention mechanism is proposed. This method first applies wavelet decomposition and 2D convolution to extract local features of each channel signal individually, and then utilizes the multi-head self-attention mechanism to enhance inter-channel mutual information perception, enabling intelligent diagnosis of plunger pump conditions. Experimental results show that this method can effectively diagnose plunger pump faults with an accuracy of 99.54%, outperforming other deep learning models in terms of training efficiency and diagnostic accuracy.
 
</p></abstract><kwd-group><kwd>Plunger Pump Fault Diagnosis</kwd><kwd> Wavelet Decomposition</kwd><kwd> Multi-Head  Self-Attention Mechanism</kwd><kwd> Multi-Channel Fusion</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>As the “power heart” of hydraulic transmission systems, plunger pumps play a crucial role in marine propulsion systems [<xref ref-type="bibr" rid="scirp.144438-ref1">1</xref>]. However, due to the harsh operating environments, internal components of the plunger pump are prone to damage [<xref ref-type="bibr" rid="scirp.144438-ref2">2</xref>]. Coupled with the complexity of the equipment structure, maintenance personnel often struggle to accurately and promptly locate fault positions and analyze their causes. Therefore, during operation, plunger pumps are at risk of efficiency degradation, shutdown, or even safety accidents due to faults, which seriously threatens the system’s stability and personnel safety [<xref ref-type="bibr" rid="scirp.144438-ref3">3</xref>]. To enhance operational reliability and performance, and to reduce economic losses caused by excessive maintenance or faults, research on fault diagnosis methods for plunger pumps has significant practical and application value.</p><p>Traditional fault diagnosis methods for plunger pumps are mainly based on signal processing techniques, which help extract features from vibration, acoustic, and other signals while filtering irrelevant noise [<xref ref-type="bibr" rid="scirp.144438-ref4">4</xref>]. Common methods include wavelet-based techniques such as Empirical Wavelet Transform (EWT) [<xref ref-type="bibr" rid="scirp.144438-ref5">5</xref>] and Wavelet Packet Decomposition (WPD) [<xref ref-type="bibr" rid="scirp.144438-ref6">6</xref>], mode decomposition techniques such as Empirical Mode Decomposition (EMD) [<xref ref-type="bibr" rid="scirp.144438-ref7">7</xref>] and Ensemble Empirical Mode Decomposition (EEMD) [<xref ref-type="bibr" rid="scirp.144438-ref8">8</xref>], as well as frequency-domain analysis methods like Short-Time Fourier Transform (STFT) [<xref ref-type="bibr" rid="scirp.144438-ref9">9</xref>] and Hilbert Transform (DCT) [<xref ref-type="bibr" rid="scirp.144438-ref10">10</xref>]. These methods have shown strong capabilities in feature extraction and classification.</p><p>With the rise of deep learning, which offers good adaptability and generalization, researchers have introduced convolutional neural networks (CNNs) such as LeNet-5 [<xref ref-type="bibr" rid="scirp.144438-ref11">11</xref>] and AlexNet [<xref ref-type="bibr" rid="scirp.144438-ref12">12</xref>], as well as recurrent neural networks (RNNs) like LSTM [<xref ref-type="bibr" rid="scirp.144438-ref13">13</xref>], into feature extraction tasks. Furthermore, since fault data of plunger pumps are often derived from multiple sensors, multi-source fusion methods like Bayesian networks [<xref ref-type="bibr" rid="scirp.144438-ref14">14</xref>] and Dempster-Shafer (D-S) theory [<xref ref-type="bibr" rid="scirp.144438-ref15">15</xref>] have been applied to integrate features from different signals, thereby improving fault characterization and diagnosis accuracy.</p><p>However, existing studies still face three main limitations:</p><p>・ Traditional signal processing methods struggle to extract deep features and rely heavily on expert knowledge, which limits the generalization ability of diagnosis models;</p><p>・ Classical neural network architectures have inherent shortcomings: CNNs are less effective in capturing global patterns, while RNNs tend to lose temporal details and local information during iteration;</p><p>・ Existing multi-channel fusion methods often rely on manually set rules or prior knowledge, lacking the ability to adaptively adjust fusion strategies based on data, thus failing to achieve true adaptive information fusion.</p><p>To address the above issues, this paper proposes a fault diagnosis method based on wavelet convolution and a multi-head self-attention mechanism for multi-channel feature fusion. First, wavelet convolution operators are employed to extract local features from each signal channel across different frequency domains, adapting to the non-stationary nature of fault signals. Then, multi-head self-attention is used to adaptively fuse features across channels and enhance the expression of key fault-related information. Experimental results demonstrate that the proposed method can capture critical features within and across channels, improving feature representation and fault classification accuracy.</p></sec><sec id="s2"><title>2. Method</title><sec id="s2_1"><title>2.1. Overview</title><p>As illustrated in <xref ref-type="fig" rid="fig1">Figure 1</xref>, the proposed model mainly consists of three parts:</p><p>First, the multi-sensor input signals are fed into a multi-channel Haar wavelet and 2D convolution module for initial feature extraction from the vibration signals.</p><p>Second, the extracted features are passed into a multi-head self-attention module to establish global dependencies in the time dimension.</p><p>Finally, a fully connected layer serves as the classifier to output the plunger pump fault diagnosis results.</p></sec><sec id="s2_2"><title>2.2. Network Structure</title><sec id="s2_2_1"><title>2.2.1. Haar Wavelet Convolution Module</title><p>Inspired by the Haar wavelet downsampling method proposed by Xu G. et al. [<xref ref-type="bibr" rid="scirp.144438-ref16">16</xref>], this module introduces the Haar wavelet basis for each channel, performing a first-order Haar wavelet transform to downsample the signal while preserving local features of the original signal. The result of the Haar wavelet transform can be expressed as follows:</p><p>a j [ k ] = 1 2 ( a j − 1 [ 2 k ] + a j − 1 [ 2 k + 1 ] ) (1)</p><p>d j [ k ] = 1 2 ( a j − 1 [ 2 k ] − a j − 1 [ 2 k + 1 ] ) (2)</p><p>In Equations (1) and (2), a j [ k ] represents the low-frequency wavelet coefficients, while d j [ k ] represents the high-frequency wavelet coefficients.</p><p>The obtained wavelet coefficients are concatenated along the convolution channel dimension and then sequentially passed through two 2D convolution modules. The first convolution module uses 16 kernels along the time dimension to extract features from both low- and high-frequency wavelet coefficients, expanding the signal into 16 channels. Batch normalization and ReLU activation are then applied to ensure training stability and enhance the model's nonlinear representation capacity. The second convolution module is similar to the first, except that the number of convolution channels remains unchanged. The parameters of the two convolutional modules are summarized in <xref ref-type="table" rid="table1">Table 1</xref>.</p><p>After passing through these two modules, the wavelet coefficients are transformed into the final output of the single-channel wavelet convolution module.</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Configuration of the convolution modules</title></caption><table><tbody><thead><tr><th align="center" valign="middle"  rowspan="2"  >Layer</th><th align="center" valign="middle"  colspan="4"  >Layer parameters</th></tr></thead><tr><td align="center" valign="middle" >Input Channels</td><td align="center" valign="middle" >Output Channels</td><td align="center" valign="middle" >Subhead</td><td align="center" valign="middle" >Stride</td></tr><tr><td align="center" valign="middle" >Conv2D</td><td align="center" valign="middle" >2</td><td align="center" valign="middle" >16</td><td align="center" valign="middle" >(1,16)</td><td align="center" valign="middle" >8</td></tr><tr><td align="center" valign="middle" >BatchNorm2D</td><td align="center" valign="middle" >16</td><td align="center" valign="middle" >16</td><td align="center" valign="middle" >―</td><td align="center" valign="middle" >―</td></tr><tr><td align="center" valign="middle" >ReLU</td><td align="center" valign="middle" >―</td><td align="center" valign="middle" >―</td><td align="center" valign="middle" >―</td><td align="center" valign="middle" >―</td></tr><tr><td align="center" valign="middle" >Conv2D</td><td align="center" valign="middle" >16</td><td align="center" valign="middle" >16</td><td align="center" valign="middle" >(1,16)</td><td align="center" valign="middle" >8</td></tr><tr><td align="center" valign="middle" >BatchNorm2D</td><td align="center" valign="middle" >16</td><td align="center" valign="middle" >16</td><td align="center" valign="middle" >―</td><td align="center" valign="middle" >―</td></tr><tr><td align="center" valign="middle" >ReLU</td><td align="center" valign="middle" >―</td><td align="center" valign="middle" >―</td><td align="center" valign="middle" >―</td><td align="center" valign="middle" >―</td></tr></tbody></table></table-wrap><p>Finally, the wavelet convolution outputs from all channels are concatenated along the signal channel dimension and then merged with the convolutional channels.</p><p>The merged dimensions are subsequently transposed with the time dimension, resulting in a matrix of shape Seq &#215; Token, where Seq represents the length of the time sequence and Token represents the feature vector at each time step.</p><p>This feature representation is then fed into the self-attention module to enable more expressive and high-level representation learning via the subsequent self-attention structure.</p></sec><sec id="s2_2_2"><title>2.2.2. Multi-Head Self-Attention Module</title><p>Due to the limited receptive field of convolution kernels, the aforementioned modules can only extract features within local regions of the signal. To capture the global features of plunger pump fault signals, a multi-head self-attention mechanism is introduced in this study.</p><p>First, the Token dimension in the Seq &#215; Token matrix is evenly divided into h parts, resulting in vectors x i ∈ ℝ S e q ∗ T o k e n / h . Then, each x i is linearly projected using learned weights W q i , W k i , W v i , producing the query, key, and value matrices: Q i , K i , V i , respectively.</p><p>For each head, the attention output is computed as:</p><p>A t t e n t i o n i = s o f t m a x ( Q i K i T / √ T o k e n / h ) V i (3)</p><p>Finally, the outputs from all attention heads are concatenated to obtain the final output of the multi-head self-attention module:</p><p>M H A = [ A t t e n t i o n 1 , A t t e n t i o n 2 , … , A t t e n t i o n h ] (4)</p></sec><sec id="s2_2_3"><title>2.2.3. Classifier</title><p>The output of the multi-head self-attention module retains the shape Seq &#215; Token. To further compress the information along the time dimension and extract more discriminative features, a classifier primarily based on fully connected layers is designed.</p><p>This classifier first reduces the temporal dimension and then compresses the feature dimension to match the number of target classes num_classes. The specific architecture is shown in <xref ref-type="table" rid="table2">Table 2</xref> below:</p><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Structure of the classifier</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Layer</th><th align="center" valign="middle" >Input Size</th><th align="center" valign="middle" >Output Size</th></tr></thead><tr><td align="center" valign="middle" >Fully Connected</td><td align="center" valign="middle" >Seq * Token</td><td align="center" valign="middle" >16 * Token</td></tr><tr><td align="center" valign="middle" >ReLU + Dropout</td><td align="center" valign="middle" >16 * Token</td><td align="center" valign="middle" >16 * Token</td></tr><tr><td align="center" valign="middle" >Fully Connected</td><td align="center" valign="middle" >16 * Token</td><td align="center" valign="middle" >1 * Token</td></tr><tr><td align="center" valign="middle" >Squeeze</td><td align="center" valign="middle" >1 * Token</td><td align="center" valign="middle" >Token</td></tr><tr><td align="center" valign="middle" >Fully Connected</td><td align="center" valign="middle" >Token</td><td align="center" valign="middle" >128</td></tr><tr><td align="center" valign="middle" >ReLU + Dropout</td><td align="center" valign="middle" >128</td><td align="center" valign="middle" >128</td></tr><tr><td align="center" valign="middle" >Fully Connected</td><td align="center" valign="middle" >128</td><td align="center" valign="middle" >num_classes</td></tr></tbody></table></table-wrap></sec></sec></sec><sec id="s3"><title>3. Experiments and Analysis</title><sec id="s3_1"><title>3.1. Data Acquisition and Preprocessing</title><p><xref ref-type="fig" rid="fig2">Figure 2</xref> shows the plunger pump fault simulation test bench used in this study. Fourteen vibration measurement channels were arranged as listed in <xref ref-type="table" rid="table3">Table 3</xref>, using Br&#252;el &amp; Kj&#230;r 4514-B accelerometers, with a sampling frequency of 25,600 Hz for all channels.</p><p>To obtain fault data of the plunger pump, 12 typical fault types were injected into the test bench as listed in <xref ref-type="table" rid="table4">Table 4</xref>. For each fault type, data were collected before and after fault injection to obtain both normal and fault-state samples.</p><p>To increase the dataset size and improve classification accuracy, each continuous data segment was first split into training and testing sets at a 7:3 ratio over time, and then sliced into 1-second segments with a 0.5-second overlap. Each data slice has a shape of 14 channels &#215; 25,600 samples. The number of samples for each category is summarized in <xref ref-type="table" rid="table5">Table 5</xref>.</p><table-wrap id="table3" ><label><xref ref-type="table" rid="table3">Table 3</xref></label><caption><title> Sensor placement information</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Channel No.</th><th align="center" valign="middle" >Measurement Location</th><th align="center" valign="middle" >Type of Measurement</th><th align="center" valign="middle" >Sensor Axis Direction</th></tr></thead><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle"  rowspan="3"  >Plunger pump</td><td align="center" valign="middle"  rowspan="3"  >Triaxial acceleration</td><td align="center" valign="middle" >X-axis</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >Y-axis</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >Z-axis</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle"  rowspan="3"  >Gear pump</td><td align="center" valign="middle"  rowspan="3"  >Triaxial acceleration</td><td align="center" valign="middle" >X-axis</td></tr><tr><td align="center" valign="middle" >5</td><td align="center" valign="middle" >Y-axis</td></tr><tr><td align="center" valign="middle" >6</td><td align="center" valign="middle" >Z-axis</td></tr><tr><td align="center" valign="middle" >7</td><td align="center" valign="middle"  rowspan="3"  >Motor front end (near coupling)</td><td align="center" valign="middle"  rowspan="3"  >Triaxial acceleration</td><td align="center" valign="middle" >X-axis</td></tr><tr><td align="center" valign="middle" >8</td><td align="center" valign="middle" >Y-axis</td></tr><tr><td align="center" valign="middle" >9</td><td align="center" valign="middle" >Z-axis</td></tr><tr><td align="center" valign="middle" >10</td><td align="center" valign="middle"  rowspan="3"  >Motor front end (far coupling)</td><td align="center" valign="middle"  rowspan="3"  >Triaxial acceleration</td><td align="center" valign="middle" >X-axis</td></tr><tr><td align="center" valign="middle" >11</td><td align="center" valign="middle" >Y-axis</td></tr><tr><td align="center" valign="middle" >12</td><td align="center" valign="middle" >Z-axis</td></tr><tr><td align="center" valign="middle" >13</td><td align="center" valign="middle" >Motor base</td><td align="center" valign="middle" >Uniaxial acceleration</td><td align="center" valign="middle" >Z-axis</td></tr><tr><td align="center" valign="middle" >14</td><td align="center" valign="middle" >Pump base</td><td align="center" valign="middle" >Uniaxial acceleration</td><td align="center" valign="middle" >Z-axis</td></tr></tbody></table></table-wrap><table-wrap id="table4" ><label><xref ref-type="table" rid="table4">Table 4</xref></label><caption><title> Injected fault types and labels for plunger pump</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Label</th><th align="center" valign="middle" >Fault Type</th><th align="center" valign="middle" >Label</th><th align="center" valign="middle" >Fault Type</th></tr></thead><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >Loose fasteners</td><td align="center" valign="middle" >7</td><td align="center" valign="middle" >Poor suction in gear pump</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >Deformed fasteners</td><td align="center" valign="middle" >8</td><td align="center" valign="middle" >Gear pump wear</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >Damaged elastomer</td><td align="center" valign="middle" >9</td><td align="center" valign="middle" >Plunger pair wear</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >Overflow regulation failure</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >Inner ring bearing fault</td></tr><tr><td align="center" valign="middle" >5</td><td align="center" valign="middle" >Loose slipper</td><td align="center" valign="middle" >11</td><td align="center" valign="middle" >Rolling element bearing fault</td></tr><tr><td align="center" valign="middle" >6</td><td align="center" valign="middle" >Valve plate wear</td><td align="center" valign="middle" >12</td><td align="center" valign="middle" >Outer ring bearing fault</td></tr></tbody></table></table-wrap><table-wrap id="table5" ><label><xref ref-type="table" rid="table5">Table 5</xref></label><caption><title> Number of samples</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Type</th><th align="center" valign="middle" >Training Set</th><th align="center" valign="middle" >Test Set</th><th align="center" valign="middle" >Total</th></tr></thead><tr><td align="center" valign="middle" >Normal</td><td align="center" valign="middle" >84</td><td align="center" valign="middle" >36</td><td align="center" valign="middle" >120</td></tr><tr><td align="center" valign="middle" >Faulty (12 types)</td><td align="center" valign="middle" >12 &#215; 35</td><td align="center" valign="middle" >12 &#215; 15</td><td align="center" valign="middle" >12 &#215; 50 = 600</td></tr><tr><td align="center" valign="middle" >Total</td><td align="center" valign="middle" >504</td><td align="center" valign="middle" >216</td><td align="center" valign="middle" >720</td></tr></tbody></table></table-wrap></sec><sec id="s3_2"><title>3.2. Model Training</title><p>All neural network models in this section were implemented using Python 3.10.11 and the deep learning framework PyTorch 2.12, and executed on a GeForce RTX 4070 GPU using the CUDA 12.1.66 parallel computing architecture.</p><p>After experimental tuning, the training configuration was finalized as follows:</p><p>・ Learning rate: 1e−4</p><p>・ Batch size: 64</p><p>・ Loss function: Cross-entropy loss</p><p>・ Optimizer: Adam optimizer (with β<sub>1</sub> = 0.9 and β<sub>2</sub> = 0.999)</p><p>・ Dropout rate: 0.1</p><p>During training, early stopping was introduced to prevent overfitting. Specifically, training was terminated early if the validation loss did not decrease within 30 consecutive epochs, and the model was restored to the state with the best validation performance.</p><p>As shown in <xref ref-type="fig" rid="fig3">Figure 3</xref> and <xref ref-type="fig" rid="fig4">Figure 4</xref>, the model achieved a 99.54% accuracy on the test set at the 28<sup>th</sup> epoch, with a total training time of 24.07 seconds. At this point, the confusion matrix of the test set is shown in <xref ref-type="fig" rid="fig5">Figure 5</xref>, where only one normal sample was misclassified as an overflow regulation failure.</p></sec><sec id="s3_3"><title>3.3. Ablation Study</title><p>To verify the effectiveness of the key structural components in the proposed model, two comparative models were designed for ablation analysis against the original architecture:</p><p>・ Model A: The Haar wavelet downsampling method is replaced with a simple interpolation method, which only performs shape transformation on the raw input signal.</p><p>・ Model B: The multi-head self-attention module in the proposed model is removed. Instead, features extracted by wavelet convolution from each channel are directly pooled and classified, ignoring inter-channel correlations.</p><p>All other parameters remain unchanged. The training results are shown in <xref ref-type="table" rid="table6">Table 6</xref>.</p><p>As shown in <xref ref-type="fig" rid="fig6">Figure 6</xref> and <xref ref-type="fig" rid="fig7">Figure 7</xref>, Model A fails to converge effectively, while Model B demonstrates a certain level of classification ability but still falls significantly short of the complete model. These results indicate that simple waveform reshaping through interpolation is insufficient for capturing the temporal and spatial features of the signal. Moreover, removing the attention mechanism and ignoring the relationships among channels leads to performance degradation.</p><table-wrap id="table6" ><label><xref ref-type="table" rid="table6">Table 6</xref></label><caption><title> Ablation study results</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Model</th><th align="center" valign="middle" >Training Steps</th><th align="center" valign="middle" >Training Time (s)</th><th align="center" valign="middle" >Accuracy</th></tr></thead><tr><td align="center" valign="middle" >Model A</td><td align="center" valign="middle" >151</td><td align="center" valign="middle" >121.64</td><td align="center" valign="middle" >66.67%</td></tr><tr><td align="center" valign="middle" >Model B</td><td align="center" valign="middle" >151</td><td align="center" valign="middle" >74.42</td><td align="center" valign="middle" >93.98%</td></tr><tr><td align="center" valign="middle" >Proposed Model</td><td align="center" valign="middle" >28</td><td align="center" valign="middle" >24.07</td><td align="center" valign="middle" >99.54%</td></tr></tbody></table></table-wrap><p>Therefore, the proposed combination of wavelet convolution and multi-head self-attention not only achieves superior accuracy but also improves training efficiency, demonstrating the effectiveness and rationality of the model’s structural design.</p></sec><sec id="s3_4"><title>3.4. Comparative Experiments</title><p>To further validate the effectiveness of the proposed model in processing plunger pump fault signals, this section compares it with several classical convolutional neural network (CNN) models. Since these baseline models require inputs of specific dimensions, the original input signals are first expanded into a three-channel format via 2D convolution. The last two dimensions are then adjusted using reshaping and interpolation techniques to meet the input size requirements.</p><p>All models are trained using the same parameter settings and dataset partitions to ensure a fair comparison in terms of classification accuracy and training efficiency. The experimental results are presented in <xref ref-type="table" rid="table7">Table 7</xref>.</p><table-wrap id="table7" ><label><xref ref-type="table" rid="table7">Table 7</xref></label><caption><title> Comparative experiment results</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Model</th><th align="center" valign="middle" >Training Steps</th><th align="center" valign="middle" >Training Time (s)</th><th align="center" valign="middle" >Accuracy</th></tr></thead><tr><td align="center" valign="middle" >AlexNet</td><td align="center" valign="middle" >71</td><td align="center" valign="middle" >53.25</td><td align="center" valign="middle" >99.54%</td></tr><tr><td align="center" valign="middle" >GoogLeNet</td><td align="center" valign="middle" >75</td><td align="center" valign="middle" >86.54</td><td align="center" valign="middle" >90.74%</td></tr><tr><td align="center" valign="middle" >LeNet-5</td><td align="center" valign="middle" >123</td><td align="center" valign="middle" >56.64</td><td align="center" valign="middle" >44.91%</td></tr><tr><td align="center" valign="middle" >ResNet34</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >15.38</td><td align="center" valign="middle" >92.13%</td></tr><tr><td align="center" valign="middle" >Proposed Model</td><td align="center" valign="middle" >28</td><td align="center" valign="middle" >24.07</td><td align="center" valign="middle" >99.54%</td></tr></tbody></table></table-wrap><p>The results show that although classical CNN models perform well in image recognition tasks, they face challenges such as low training efficiency and suboptimal accuracy when applied to the fault diagnosis of plunger pumps. In contrast, the proposed model achieves comparable or even higher classification accuracy while maintaining a compact architecture and superior training efficiency. These findings demonstrate the proposed model’s effectiveness and practicality in multi-channel temporal signal classification tasks.</p></sec></sec><sec id="s4"><title>4. Conclusions</title><p>This paper proposes a fault diagnosis method for plunger pumps based on wavelet convolution and multi-head self-attention mechanisms. The main conclusions are as follows:</p><p>・ The proposed model effectively extracts instantaneous feature variations across different frequency bands within individual channels through the wavelet convolution module. Meanwhile, the multi-head self-attention mechanism enhances the model’s ability to adaptively capture key inter-channel features, which helps suppress redundant information and contributes to improved detection accuracy.</p><p>・ Experiments conducted on real-world datasets show that the proposed model achieves a classification accuracy of 99.54% on multi-channel vibration signals, with a training time of 24.07 seconds on an RTX 4070 GPU. Compared with classical convolutional models, the proposed approach significantly improves both training efficiency and diagnostic accuracy, demonstrating clear advantages.</p><p>This study focuses on multi-channel vibration signals of plunger pumps and achieves effective fault type classification based on the signal characteristics. Future work will explore the integration of real-time signal processing and edge computing to facilitate on-site deployment and real-time application of the proposed method in industrial environments.</p></sec><sec id="s5"><title>Conflicts of Interest</title><p>The authors declare no conflicts of interest regarding the publication of this paper.</p></sec></body><back><ref-list><title>References</title><ref id="scirp.144438-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Zhu, Y., Li, G., Wang, R., Tang, S., Su, H. and Cao, K. (2023) Intelligent Fault Diagnosis Methods for Hydraulic Piston Pumps: A Review. Journal of Marine Science and Engineering, 11, 1609. https://doi.org/10.3390/jmse11081609</mixed-citation></ref><ref id="scirp.144438-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Kumar, S., Bergada, J.M. and Watton, J. (2009) Axial Piston Pump Grooved Slipper Analysis by CFD Simulation of Three Dimensional NVS Equation in Cylindrical Coordinates. Computers &amp; Fluids, 38, 648-663.  
https://doi.org/10.1016/j.compfluid.2008.06.007</mixed-citation></ref><ref id="scirp.144438-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Khan, K., Sohaib, M., Rashid, A., Ali, S., Akbar, H., Basit, A. and Ahmad, T. (2021) Recent Trends and Challenges in Predictive Maintenance of Aircraft’s Engine and Hydraulic System. Journal of the Brazilian Society of Mechanical Sciences and Engineering, 43, 403. https://doi.org/10.1007/s40430-021-03121-2</mixed-citation></ref><ref id="scirp.144438-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Yang, Y., Ding, L., Xiao, J., Fang, G. and Li, J. (2022) Current Status and Applications for Hydraulic Pump Fault Diagnosis: A Review. Sensors, 22, 9714.  
https://doi.org/10.3390/s22249714</mixed-citation></ref><ref id="scirp.144438-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Yu, H., Li, H., Li, Y. and Li, Y. (2019) A Novel Improved Full Vector Spectrum Algorithm and Its Application in Multi-Sensor Data Fusion for Hydraulic Pumps. Measurement, 133, 145-161. https://doi.org/10.1016/j.measurement.2018.10.011</mixed-citation></ref><ref id="scirp.144438-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Hou, W., Lu, C., Liu, H. and Lu, C. (2009) Fault Diagnosis Based on Wavelet Package for Hydraulic Pump: Proceedings of the 8th International Conference on Reliability, Maintainability and Safety, Chengdu, China, 20-24 July 2009, 831-835. 
https://doi.org/10.1109/ICRMS.2009.5269950</mixed-citation></ref><ref id="scirp.144438-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Meng, L., Dong, C., Zhou, J., Lu, X., Wang, Y., Deng, M., Huang, C. and Yuan, B. (2021) Typical Fault Simulation and On-Line Monitoring for Aviation Hydraulic Pump. Machinery Tool and Hydraulics, 49, 170-174.</mixed-citation></ref><ref id="scirp.144438-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Jiang, W.L., Zhang, P.Y., Li, M. and Zhang, S.Q. (2021) Axial Piston Pump Fault Diagnosis Method Based on Symmetrical Polar Coordinate Image and Fuzzy C-Means Clustering Algorithm. Shock and Vibration, 2021, 6681751.  
https://doi.org/10.1155/2021/6681751</mixed-citation></ref><ref id="scirp.144438-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Przystupa, K., Ambrozkiewicz, B. and Litak, G. (2020) Diagnostics of Transient States in Hydraulic Pump System with Short Time Fourier Transform. Advances in Science and Technology Research Journal, 178-183.  
https://doi.org/10.12913/22998624/116971</mixed-citation></ref><ref id="scirp.144438-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Yu, H., Li, H. and Li, Y. (2020) Vibration Signal Fusion Using Improved Empirical Wavelet Transform and Variance Contribution Rate for Weak Fault Detection of Hydraulic Pumps. ISA Transactions, 107, 385-401. 
https://doi.org/10.1016/j.isatra.2020.07.025</mixed-citation></ref><ref id="scirp.144438-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Zhu, Y., Li, G., Wang, R., Tang, S., Su, H. and Cao, K. (2021) Intelligent Fault Diagnosis of Hydraulic Piston Pump Combining Improved LeNet-5 and PSO Hyperparameter Optimization. Applied Acoustics, 183, 108336. 
https://doi.org/10.1016/j.apacoust.2021.108336</mixed-citation></ref><ref id="scirp.144438-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Zhu, Y., Li, G., Wang, R., Tang, S., Su, H. and Cao, K. (2021) Intelligent Fault Diagnosis of Hydraulic Piston Pump Based on Wavelet Analysis and Improved AlexNet. Sensors, 21, 549. https://doi.org/10.3390/s21020549</mixed-citation></ref><ref id="scirp.144438-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Zhu, Y., Su, H., Tang, S., Zhang, S. and Wang, J. (2023) A Novel Fault Diagnosis Method Based on SWT and VGG-LSTM Model for Hydraulic Axial Piston Pump. Journal of Marine Science and Engineering, 11, 594.  
https://doi.org/10.3390/jmse11030594</mixed-citation></ref><ref id="scirp.144438-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Liu, Z., Liu, Y., Shan, H., Cai, B. and Huang, Q. (2015) A Fault Diagnosis Method-ology for Gear Pump Based on EEMD and Bayesian Network. PLoS ONE, 10, e0125703.  
https://doi.org/10.1371/journal.pone.0125703</mixed-citation></ref><ref id="scirp.144438-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Lu, C., Wang, S. and Wang, X. (2017) A Multi-Source Information Fusion Fault Diagnosis for Aviation Hydraulic Pump Based on the New Evidence Similarity Distance. Aerospace Science and Technology, 71, 392-401. 
https://doi.org/10.1016/j.ast.2017.09.040</mixed-citation></ref><ref id="scirp.144438-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Xu, G., Liao, W., Zhang, X., Li, C., He, X. and Wu, X. (2023) Haar Wavelet Downsampling: A Simple but Effective Downsampling Module for Semantic Segmentation. Pattern Recognition, 143, 109819. https://doi.org/10.1016/j.patcog.2023.109819</mixed-citation></ref></ref-list></back></article>