<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JCC</journal-id><journal-title-group><journal-title>Journal of Computer and Communications</journal-title></journal-title-group><issn pub-type="epub">2327-5219</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jcc.2021.911007</article-id><article-id pub-id-type="publisher-id">JCC-113211</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Computer Science&amp;Communications</subject></subj-group></article-categories><title-group><article-title>
 
 
  A Hybrid K-Means-GRA-SVR Model Based on Feature Selection for Day-Ahead Prediction of Photovoltaic Power Generation
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Jiemin</surname><given-names>Lin</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Haiming</surname><given-names>Li</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>School of Computer Science and Technology, Shanghai University of Electric Power, Shanghai, China</addr-line></aff><pub-date pub-type="epub"><day>05</day><month>11</month><year>2021</year></pub-date><volume>09</volume><issue>11</issue><fpage>91</fpage><lpage>111</lpage><history><date date-type="received"><day>13,</day>	<month>April</month>	<year>2021</year></date><date date-type="rev-recd"><day>14,</day>	<month>November</month>	<year>2021</year>	</date><date date-type="accepted"><day>17,</day>	<month>November</month>	<year>2021</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  In order to ensure that the large-scale application of photovoltaic power generation does not affect the stability of the grid, accurate photovoltaic (PV) power generation forecast is essential. A short-term PV power generation forecast method using the combination of K-means++, grey relational analysis (GRA) and support vector regression (SVR) based on feature selection (Hybrid Kmeans-GRA-SVR, HKGSVR) was proposed. The historical power data were clustered through the multi-index K-means++ algorithm and divided into ideal and non-ideal weather. The GRA algorithm was used to match the similar day and the nearest neighbor similar day of the prediction day. And selected appropriate input features for different weather types to train the SVR model. Under ideal weather, the average values of MAE, RMSE and R2 were 0.8101, 0.9608 kW and 99.66%, respectively. And this method reduced the average training time by 77.27% compared with the standard SVR model. Under non-ideal weather conditions, the average values of MAE, RMSE and R2 were 1.8337, 2.1379 kW and 98.47%, respectively. And this method reduced the average training time of the standard SVR model by 98.07%. The experimental results show that the prediction accuracy of the proposed model is significantly improved compared to the other five models, which verify the effectiveness of the method.
 
</p></abstract><kwd-group><kwd>Feature Selection</kwd><kwd> Grey Relational Analysis</kwd><kwd> K-Means++</kwd><kwd> Nearest Neighbor Similar Day</kwd><kwd> Photovoltaic Power</kwd><kwd> Support Vector Regression</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>In the face of limited fossil energy and the need to adjust the energy structure, the exploration of renewable energy power generation technology is of great significance [<xref ref-type="bibr" rid="scirp.113211-ref1">1</xref>]. A study shows that the earth receives about 1.8 &#215; 10<sup>11</sup> MW of power per second from solar radiation [<xref ref-type="bibr" rid="scirp.113211-ref2">2</xref>]. Photovoltaic power generation is one of the most promising solar power technologies [<xref ref-type="bibr" rid="scirp.113211-ref3">3</xref>]. Photovoltaic energy has the advantages of cleanliness, wide distribution and abundant reserves, and has become the best substitute for industrial and residential power generation [<xref ref-type="bibr" rid="scirp.113211-ref4">4</xref>]. According to the 2020 report of the International Renewable Energy Agency, in the past 8 years, the global photovoltaic power generation cost has dropped by more than 70%, and the global installed capacity has reached 578.553 GW [<xref ref-type="bibr" rid="scirp.113211-ref5">5</xref>].</p><p>However, due to the chaotic nature of the weather system, the production of photovoltaic energy is highly random, volatile and intermittent, which may lead to grid power and voltage imbalances, and also greatly increase the difficulty of large-scale photovoltaic energy applications [<xref ref-type="bibr" rid="scirp.113211-ref6">6</xref>] [<xref ref-type="bibr" rid="scirp.113211-ref7">7</xref>]. In order to improve the power system’s ability to consume photovoltaic energy, many solutions have been proposed, including energy storage optimization [<xref ref-type="bibr" rid="scirp.113211-ref8">8</xref>], demand response strategy [<xref ref-type="bibr" rid="scirp.113211-ref9">9</xref>] [<xref ref-type="bibr" rid="scirp.113211-ref10">10</xref>], power flow optimization [<xref ref-type="bibr" rid="scirp.113211-ref11">11</xref>], stand-alone microgrid [<xref ref-type="bibr" rid="scirp.113211-ref12">12</xref>], and PV power forecasting [<xref ref-type="bibr" rid="scirp.113211-ref13">13</xref>]. Considering economy and feasibility comprehensively, photovoltaic power generation forecast is one of the most promising solutions to the impact of large-scale photovoltaic energy application on the grid [<xref ref-type="bibr" rid="scirp.113211-ref14">14</xref>] [<xref ref-type="bibr" rid="scirp.113211-ref15">15</xref>].</p><p>The current photovoltaic power generation forecasting technologies have three main directions: physical methods, time series statistical methods and ensemble methods [<xref ref-type="bibr" rid="scirp.113211-ref14">14</xref>]. [<xref ref-type="bibr" rid="scirp.113211-ref16">16</xref>] proposed a partial function linear regression model to forecast the day-ahead photovoltaic power generation. The regression method has a low amount of calculation, but the prediction accuracy is relatively low. [<xref ref-type="bibr" rid="scirp.113211-ref17">17</xref>] proposed an ANN model based on an extreme learning machine algorithm to predict photovoltaic power generation. Artificial neural network can handle nonlinear problems and has excellent self-learning ability, so it has high prediction accuracy. However, the ANN multi-layer network structure greatly increases the complexity of the model, which makes training and optimizing the model consume a lot of computing resources and longer training time. In [<xref ref-type="bibr" rid="scirp.113211-ref18">18</xref>] [<xref ref-type="bibr" rid="scirp.113211-ref19">19</xref>] [<xref ref-type="bibr" rid="scirp.113211-ref20">20</xref>], the support vector machine (SVM) is used for short-term photovoltaic output forecasting. SVM can also handle non-linear problems, has excellent learning ability and does not rely heavily on prior knowledge. The training speed is fast and has the ability to prevent overfitting, with good generalization.</p><p>The ensemble method solves the limitations of a single model by mixing different models with unique functions, thereby improving the prediction performance [<xref ref-type="bibr" rid="scirp.113211-ref21">21</xref>]. For the prediction of photovoltaic power generation, the ensemble method that mixes various effective methods is more effective and accurate [<xref ref-type="bibr" rid="scirp.113211-ref22">22</xref>]. For example, the hybrid GA-SVM model [<xref ref-type="bibr" rid="scirp.113211-ref20">20</xref>], which performed better than the SVM model. In [<xref ref-type="bibr" rid="scirp.113211-ref23">23</xref>], a hybrid Kmeans-GRA-Elman model was proposed, the performance of Kmeans-GRA-Elman was better than BP neural network, Elman, GRA-BPNN and GRA-Elman.</p><p>Photovoltaic power generation has obvious seasonal and weather characteristics [<xref ref-type="bibr" rid="scirp.113211-ref24">24</xref>]. Weather conditions can be roughly divided into two types: the ideal weather type (sunny day), and non-ideal weather types [<xref ref-type="bibr" rid="scirp.113211-ref25">25</xref>]. For ideal weather, the prediction accuracy of many prediction methods is high enough [<xref ref-type="bibr" rid="scirp.113211-ref26">26</xref>]. It can be seen from [<xref ref-type="bibr" rid="scirp.113211-ref27">27</xref>] [<xref ref-type="bibr" rid="scirp.113211-ref28">28</xref>] that the prediction accuracy of these methods for non-ideal weather was much lower than that of ideal weather. In order to improve the prediction performance under non-ideal weather, similar algorithms have been used in many studies to extract output features under similar weather. For example, [<xref ref-type="bibr" rid="scirp.113211-ref29">29</xref>] proposed a prediction method based on similar days and improved BP neural network. The similarity algorithm can effectively extract the output characteristics of different weather types. Moreover, compared to directly using a large amount of historical data to train the model, the use of similar days not only saves a lot of computing resources, but also improves the prediction accuracy of the model. However, if the time interval between the similar day and the forecast day is too long, the characteristics of the photovoltaic array (surface cleanliness, module aging, conversion efficiency, etc.) have changed a lot, which will cause a large error between the predicted result and the actual value [<xref ref-type="bibr" rid="scirp.113211-ref25">25</xref>].</p><p>A short-term photovoltaic power generation forecast method using the combination of K-means++, grey relational analysis (GRA) and support vector regression (SVR) based on feature selection is proposed in this paper. The proposed HKGSVR (hybrid Kmeans-GRA-Support Vector Regression) forecasting model is compared with SVR, HKGLSTM (hybrid Kmeans-GRA-LSTM), HKGBP (hybrid Kmeans-GRA-Back Propagation Neural Network), HKGLR (hybrid Kmeans-GRA-Linear Regression), HKGARIMA (hybrid Kmeans-GRA-Auto-regressive Integrated Moving Average), respectively, to demonstrate its superiority in predictive performance. The main contributions of this paper include:</p><p>1) A novel day-ahead PV power forecasting method utilizes SVR, clustering and similarity algorithms is proposed.</p><p>2) Clustering historical power data through multi-index K-means++ to obtain power generation modes of different weather types. Overcome the limitation of directly categorizing according to weather tags. According to the average power of each cluster, it is divided into ideal weather cluster and non-ideal weather cluster.</p><p>3) The GRA algorithm is used to match the nearest neighbor similar day, and the error caused by the long time interval between the similar day and the forecast day is reduced by using the information of the nearest neighbor similar day.</p><p>4) By analyzing the correlation between photovoltaic output power and various meteorological factors, 10 feature combinations are proposed. Select appropriate input features for ideal and non-ideal weather to further improve the prediction accuracy of photovoltaic power generation.</p><p>The remainder of this paper is organized as follows. Section 2 describes the hybrid Kmeans-GRA-SVR model. Section 3 illustrates clustering and model evaluation metrics. Section 4 introduces the experiments and result analysis. Finally, conclusions are given in Section 5.</p></sec><sec id="s2"><title>2. Hybrid K-Means-GRA-SVR Model</title><sec id="s2_1"><title>2.1. K-Means++ Clustering Algorithm</title><p>K-means++ clustering algorithm is an improved version of K-means algorithm. This algorithm separates the K initial cluster centers more from each other. In this work, which is selected as the classifier due to its higher efficiency and improved robustness compared with others (e.g., standard K-means, K-medoids, Gaussian mixture models, etc.) [<xref ref-type="bibr" rid="scirp.113211-ref30">30</xref>]. The running process of K-means++ is as follows:</p><p>Step 1: Randomly select a sample as the first cluster center C<sub>1</sub>;</p><p>Step 2: Calculate the probability of each sample being selected as the next cluster center:</p><p>D ( x ) 2 ∑ x ∈ X D ( x ) 2 (1)</p><p>where, D(x) represents the distance between the sample and the nearest cluster center.</p><p>Then use the roulette method to select the next cluster center;</p><p>Step 3: Repeat step 2 until K cluster centers are selected;</p><p>Step 4: For each sample x<sub>i</sub> in the datasets, calculate its distance to K cluster centers, and then put it into the class corresponding to the smallest distance cluster center;</p><p>Step 5: For each cluster, recalculate its cluster center C<sub>i</sub>:</p><p>c i = 1 c i ∑ x ∈ c i x (2)</p><p>Step 6: Repeat steps 4 and 5 until the position of the cluster center does not change.</p><p>In this part, the historical power data is directly clustered by season to obtain different power generation modes due to the diversity of weather. Moreover, the aging of the equipment itself and its own parameters will be different under different weather, it is difficult for us to accurately measure these changes. The characteristics of historical power data will integrate these changes into it. After clustering the historical power data, the centroid value of each cluster is calculated by the minimum, average and maximum of global horizontal irradiance (GHI), diffuse horizontal irradiance (DHI), relative humidity (RH) and temperature (T)(12 meteorological factor eigenvalues).</p></sec><sec id="s2_2"><title>2.2. Grey Relational Analysis Algorithm</title><p>The basic idea of the grey relational analysis algorithm is to judge the correlation degree by comparing the geometric similarity between the reference sequence and several data columns. Generally, the more consistent the change tendency of the reference sequence and the comparison sequence, the higher the degree of correlation between the two variables. The flow of the GRA algorithm is as follows:</p><p>Step 1: Determine the reference sequence y and the comparison sequence x<sub>i</sub>:</p><p>y = { y ( k ) | k = 1 , 2 , ⋯ , n } (3)</p><p>x i = { x i ( k ) | k = 1 , 2 , ⋯ , n } , i = 1 , 2 , ⋯ , m (4)</p><p>where, n and m represent the dimension of the eigenvalues and the number of comparison sequence, respectively.</p><p>Step 2: Non-dimensionalization of variables:</p><p>d j * ( k ) = D j ( k ) − D a v ( k ) D max ( k ) − D min ( k ) , k = 1 , 2 , ⋯ , n ;   i = 0 , 1 , 2 , ⋯ , m ;   j = 1 , 2 , ⋯ , m + 1 (5)</p><p>where, D<sub>j</sub>(k) contains reference sequence and comparison sequence, D<sub>av</sub>(k), D<sub>min</sub>(k) and D<sub>max</sub>(k) are the average, minimum and maximum values of each column, j represents sum of the number of reference sequence and comparison sequence.</p><p>Non-dimensionalization is used to solve the problem that the columns cannot be compared due to the different dimensions.</p><p>Step 3: Calculate correlation coefficient ξ<sub>i</sub>(k):</p><p>ξ i ( k ) = min i min k | y ( k ) − x i ( k ) | + ρ max i max k | y ( k ) − x i ( k ) | | y ( k ) − x i ( k ) | + ρ max i max k | y ( k ) − x i ( k ) | (6)</p><p>where, ρ is called the resolution coefficient, here, ρ is 0.5.</p><p>Step 4: Calculate correlation degree.</p><p>Calculate the average value of the correlation coefficient at each moment (that is, each point in the curve) r<sub>i</sub>:</p><p>r i = 1 n ∑ k = 1 n ξ i ( k ) , k = 1 , 2 , ⋯ , n (7)</p><p>Step 5: Sort correlation degree.</p><p>After determining the cluster to which the prediction day belongs, the correlation between the prediction day and each sample in the cluster is calculated by GRA based on 12 meteorological factor eigenvalues, and the date with the correlation degree greater than the threshold (an appropriate correlation value that takes into account the similarity and the number of samples) is regarded as the similar days. Based on GRA global matching results: for the ideal weather, the sample with the highest correlation in the 7 days before the forecast date is set as the nearest neighbor similar day; for non-ideal weather, the sample with the highest correlation in the 30 days before the prediction is set as the nearest neighbor similar day.</p></sec><sec id="s2_3"><title>2.3. Support Vector Regression</title><p>Based on the structural risk minimization theory, the support vector machine constructs a hyperplane in the feature space, thereby overcoming the local optimal problem and requiring fewer training samples [<xref ref-type="bibr" rid="scirp.113211-ref30">30</xref>]. When the data type is complex, support vector regression is used. For a set of data { ( X i , Y i ) , i = 1 , 2 , ⋯ , n } , X<sub>i</sub> is the input variable of the sample, and Y<sub>i</sub> is the target value. The support vector machine equation based on Vapnik theory is as follows [<xref ref-type="bibr" rid="scirp.113211-ref31">31</xref>]:</p><p>f ( x ) = ω T ϕ ( x ) + b (8)</p><p>where, ω is a vector of weight coefficients, Φ(x) is the nonlinear mapping function and b denotes a bias constant.</p><p>ω and b can be obtained by the following formula:</p><p>minimize : 1 2 ‖ ω ‖ 2 + C ∑ i = 1 n ξ i + ξ i * (9)</p><p>subject   to { y i − 〈 ω , ϕ ( x i ) 〉 − b ≤ ε + ξ i 〈 ω , ϕ ( x i ) 〉 + b − y i ≤ ε + ξ i * ξ i ≥ 0 , ξ i * ≥ 0 (10)</p><p>where, ξ i and ξ i * are slack variables, and C denotes the penalty variable, ε is the insensitive loss function.</p><p>By introducing Lagrangian multipliers and optimal constraints, (8) can be transformed into:</p><p>f ( x , a i a i * ) = ∑ i = 1 n ( a i − a i * ) K ( x , x i ) + b (11)</p><p>where, K(x, x<sub>i</sub>) = Φ(x<sub>i</sub>)Φ(x<sub>j</sub>) is the kernel function.</p><p>In this paper, the radial basis function (RBF) kernel is applied to construct the SVR model. The RBF kernel is presented as:</p><p>K ( x i , x i ) = exp ( − γ ‖ x i − x i ‖ 2 ) (12)</p><p>where, γ is the kernel parameter.</p></sec><sec id="s2_4"><title>2.4. The HKGSVR Model Workflow</title><p>The flow chart of the hybrid K means-GRA-LSTM model is shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>, and the workflow is as follows:</p><p>Step 1: Obtain historical photovoltaic output power and meteorological factor data, and deal with the missing and abnormal data in the data set.</p><p>Step 2: Use the multi-index K-means++ algorithm to cluster historical photovoltaic power data by season, and calculate the 12 meteorological factor eigenvalues as the central value of each cluster. According to the average power of each cluster, it is divided into ideal weather cluster and non-ideal weather cluster.</p><p>Step 3: The Euclidean distance, Pearson correlation coefficient and GRA correlation between the 12 meteorological eigenvalues of the forecast day and the centroid value of each cluster are calculated to determine the cluster to which the forecast day belongs.</p><p>Step 4: Calculate the correlation between the predicted day and each sample in the matched cluster through GRA to obtain the similar days (as the training set) and the nearest neighbor similar day (as the validation set) and normalize them.</p><p>Step 5: Select appropriate input features and similar days are used to train SVR. Determine the C and γ of SVR through grid search and cross-validation, and use the nearest neighbor similar day to test.</p><p>Step 6: Use the trained model to predict the prediction day.</p></sec></sec><sec id="s3"><title>3. Evaluation Metrics</title><sec id="s3_1"><title>3.1. Clustering Evaluation Metrics</title><p>If the ground truth labels are not known, evaluation must be performed using the model itself. The Silhouette Coefficient is an example of such an evaluation, the score is higher when clusters are dense and well separated. Silhouette Coefficient S(i) is defined as follows [<xref ref-type="bibr" rid="scirp.113211-ref32">32</xref>]:</p><p>S ( i ) = b ( i ) − a ( i ) max { a ( i ) , b ( i ) } (13)</p><p>where, a(i) is the mean distance between a sample and all other points in the same cluster, b(i) is the mean distance between a sample and all other points in the next nearest cluster. Average the Silhouette Coefficient of all points, which is the total Silhouette Coefficient of the clustering result.</p><p>Davies-Bouldin index is defined as follows [<xref ref-type="bibr" rid="scirp.113211-ref33">33</xref>]:</p><p>DBI = 1 n ∑ i = 1 n max j ≠ i ( S i &#175; + S j &#175; ‖ ω i − ω j ‖ 2 ) (14)</p><p>where, S i &#175; is the average distance from the points in the cluster to the cluster centroid, ‖ ω i − ω j ‖ 2 is the distance between the centroid of cluster i and j.</p><p>The Davies-Bouldin index is lower if the model clusters have better separation.</p><p>SSE is also an effective metric, that is, the sum of squared errors of the distance between the centroid of each cluster and the points in the cluster. SSE is defined as follows:</p><p>SSE = ∑ i = 1 K ∑ d i s t ( x , c i ) 2 (15)</p></sec><sec id="s3_2"><title>3.2. Metrics of Photovoltaic Power Forecasting Techniques</title><p>In order to evaluate the performance of the proposed method HKGSVR for photovoltaic power generation forecasting, the root mean square error (RMSE), average absolute error (MAE) and coefficient of determination (R<sup>2</sup>) indicators were calculated. The mean absolute error can better reflect the difference between the predicted value and the true value. RMSE is used to measure the deviation between the predicted value and the actual value, so it is more sensitive to outliers (that is, if the predicted value of a point is very different from the true value, the RMSE of the curve will be very large). R<sup>2</sup> is used to test the fit of the predicted value to the true value, and is generally used to evaluate the prediction performance of the model. They are defined as follows [<xref ref-type="bibr" rid="scirp.113211-ref14">14</xref>].</p><p>1) The RMSE is defined as:</p><p>RMSE = 1 N ∑ i = 1 N P f i − P a i (16)</p><p>where, P<sub>ai</sub> and P<sub>fi</sub> are the actual and predicted value at i hour. N refers to the number of hours a sample contains.</p><p>2) The MAE is expressed as:</p><p>MAE = 1 N ∑ i = 1 N | P f i − P a i | (17)</p><p>3) The R<sup>2</sup> is given as:</p><p>R 2 = ( N ∑ i = 1 N P f i P a i − ∑ i = 1 N P f i ∑ i = 1 N P a i ) 2 ( N ∑ i = 1 N P f i 2 − ( ∑ i = 1 N P f i ) 2 ) ( N ∑ i = 1 N P a i 2 − ( ∑ i = 1 N P a i ) 2 ) (18)</p></sec></sec><sec id="s4"><title>4. Experimental Analysis</title><sec id="s4_1"><title>4.1. Data</title><p>In this paper, the general datasets on the DKASC (Desert Knowledge Australia Solar Center) website are used for related experiments. The photovoltaic array is composed of 22 polycrystalline silicon photovoltaic panels with a rated power of 265 W, whose total rated power is 5.83 kW. The photovoltaic array is located at the Desert Knowledge Precinct in Alice Springs, a town in the Northern Territory that enjoys one of the country’s highest solar resources in an arid desert environment. The configuration information of the photovoltaic array is shown in <xref ref-type="table" rid="table1">Table 1</xref>. Meteorology (global horizontal irradiance, diffuse horizontal irradiance, relative humidity and temperature) and historical power data of PV arrays from March 1, 2018 to February 29, 2020 were used in the experiment. The experiment uses data with an interval of 1 hour from 7:00 to 18:00 every day.</p></sec><sec id="s4_2"><title>4.2. Number of Clusters and Weather Division</title><p>In order to obtain the appropriate number of clusters for each season, SSE, DBI and Silhouette Coefficient (S) are used for evaluation. Taking autumn as an example, the experimental results are shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>.</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> The configuration information of photovoltaic array</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Item</th><th align="center" valign="middle" >Information</th></tr></thead><tr><td align="center" valign="middle" >Array Rating</td><td align="center" valign="middle" >5.83 kW</td></tr><tr><td align="center" valign="middle" >Panel Rating</td><td align="center" valign="middle" >265 W</td></tr><tr><td align="center" valign="middle" >Number Of Panels</td><td align="center" valign="middle" >22</td></tr><tr><td align="center" valign="middle" >Panel Type</td><td align="center" valign="middle" >HSL 60 S</td></tr><tr><td align="center" valign="middle" >Array Area</td><td align="center" valign="middle" >36.74</td></tr><tr><td align="center" valign="middle" >Inverter Size/Type</td><td align="center" valign="middle" >SMA SMC 6000A</td></tr><tr><td align="center" valign="middle" >Array Tilt/Azimuth</td><td align="center" valign="middle" >Tilt = 20, Azimuth = 0 (Solar North)</td></tr><tr><td align="center" valign="middle" >Nominal working temperature</td><td align="center" valign="middle" >45 &#177; 3 Celsius</td></tr><tr><td align="center" valign="middle" >Temperature coefficient of power</td><td align="center" valign="middle" >−0.41%/Celsius</td></tr></tbody></table></table-wrap><p>It can be seen from <xref ref-type="fig" rid="fig2">Figure 2</xref> that SSE decreases as the number of clusters increases. When the number of clusters is 3, the downward trend begins to slow down. DBI has the best performance when the value of K is 3. When the value of K is 2 and 3, the value of S is 0.71 and 0.64, respectively. Then, as the value of K increases, the value of S drops sharply. So the value of K is chosen between 2 and 3. When K = 2, the blue cluster and the red cluster merge into one cluster. However, the blue clusters are mostly smooth arcs, while the red clusters are mostly polylines. Therefore, the value of K is chosen to be 3. The blue cluster is selected as the ideal weather cluster (most of the curves are smooth and the average power is larger in the cluster), and the green and red clusters are non-ideal weather clusters (most of them are broken lines in the clusters, and the average power is small, the average power of the green cluster is 123.09 kW, and the average power of the red cluster is 301.90 kW).</p><p>The evaluation of clustering results in each season is shown in <xref ref-type="table" rid="table2">Table 2</xref>. In order to prevent local optima or other abnormal situations, 100 rounds of experiments were carried out. Considering all indicators and clustering results comprehensively, the number of clusters in spring is 3, the number of clusters in summer is 2, the number of clusters in autumn is 3, and the number of clusters in winter is 3.</p><p>The clustering results of each season are divided into ideal weather clusters and non-ideal weather clusters by comparing the average power and geometric shape (arc and polyline) of each cluster. The average power of each cluster in the four seasons is shown in <xref ref-type="table" rid="table3">Table 3</xref>. The ideal weather clusters are mostly smooth</p><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Cluster evaluation metrics for each season</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Metrics</th><th align="center" valign="middle" >Spring</th><th align="center" valign="middle" >Summer</th><th align="center" valign="middle" >Autumn</th><th align="center" valign="middle" >Winter</th></tr></thead><tr><td align="center" valign="middle" >SSE K = 2</td><td align="center" valign="middle" >44,578.41</td><td align="center" valign="middle" >20,278.67</td><td align="center" valign="middle" >26,813.26</td><td align="center" valign="middle" >4218.80</td></tr><tr><td align="center" valign="middle" >K = 3</td><td align="center" valign="middle" >32,045.12</td><td align="center" valign="middle" >14,899.18</td><td align="center" valign="middle" >15,383.50</td><td align="center" valign="middle" >2942.09</td></tr><tr><td align="center" valign="middle" >K = 4</td><td align="center" valign="middle" >24,354.12</td><td align="center" valign="middle" >11,971.90</td><td align="center" valign="middle" >11,657.92</td><td align="center" valign="middle" >2473.00</td></tr><tr><td align="center" valign="middle" >DBI K = 2</td><td align="center" valign="middle" >0.6723</td><td align="center" valign="middle" >0.9740</td><td align="center" valign="middle" >0.8402</td><td align="center" valign="middle" >0.7100</td></tr><tr><td align="center" valign="middle" >K = 3</td><td align="center" valign="middle" >0.7846</td><td align="center" valign="middle" >0.9440</td><td align="center" valign="middle" >0.7098</td><td align="center" valign="middle" >0.6375</td></tr><tr><td align="center" valign="middle" >K = 4</td><td align="center" valign="middle" >1.0187</td><td align="center" valign="middle" >1.0247</td><td align="center" valign="middle" >0.7962</td><td align="center" valign="middle" >0.9049</td></tr><tr><td align="center" valign="middle" >S K = 2</td><td align="center" valign="middle" >0.6383</td><td align="center" valign="middle" >0.6156</td><td align="center" valign="middle" >0.7119</td><td align="center" valign="middle" >0.4861</td></tr><tr><td align="center" valign="middle" >K = 3</td><td align="center" valign="middle" >0.5893</td><td align="center" valign="middle" >0.5866</td><td align="center" valign="middle" >0.6390</td><td align="center" valign="middle" >0.5074</td></tr><tr><td align="center" valign="middle" >K = 4</td><td align="center" valign="middle" >0.5231</td><td align="center" valign="middle" >0.5705</td><td align="center" valign="middle" >0.4073</td><td align="center" valign="middle" >0.5018</td></tr></tbody></table></table-wrap><table-wrap id="table3" ><label><xref ref-type="table" rid="table3">Table 3</xref></label><caption><title> Average power of each cluster in four seasons (kW)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Cluster number</th><th align="center" valign="middle" >Spring</th><th align="center" valign="middle" >Summer</th><th align="center" valign="middle" >Autumn</th><th align="center" valign="middle" >Winter</th></tr></thead><tr><td align="center" valign="middle" >Cluster1</td><td align="center" valign="middle" >437.41</td><td align="center" valign="middle" >440.65</td><td align="center" valign="middle" >423.34</td><td align="center" valign="middle" >339.07</td></tr><tr><td align="center" valign="middle" >Cluster2</td><td align="center" valign="middle" >281.24</td><td align="center" valign="middle" >331.46</td><td align="center" valign="middle" >123.09</td><td align="center" valign="middle" >376.46</td></tr><tr><td align="center" valign="middle" >Cluster3</td><td align="center" valign="middle" >111.08</td><td align="center" valign="middle" ></td><td align="center" valign="middle" >301.90</td><td align="center" valign="middle" >419.81</td></tr></tbody></table></table-wrap><p>arcs, and the average power is relatively large. The non-ideal weather clusters are mostly broken lines, and the average power is small. Therefore, spring cluster 1, summer cluster 1, autumn cluster 1 and winter cluster 2 and 3 are divided into ideal weather clusters, and the rest are non-ideal weather clusters.</p></sec><sec id="s4_3"><title>4.3. Selection of Similar Day Threshold and Nearest Similar Day</title><p>The similar days are obtained by calculating the GRA correlation between the predicted days and the samples in the matching clusters. In order to improve the prediction accuracy while reducing the computational cost and speeding up the training speed of the model, it is necessary to select an appropriate correlation threshold. A higher correlation threshold can improve the prediction accuracy, but too few training samples may cause overfitting. After comprehensive consideration, the similar day correlation threshold of each forecast day, the nearest neighbor similar day and its correlation degree are shown in <xref ref-type="table" rid="table4">Table 4</xref> and <xref ref-type="table" rid="table5">Table 5</xref>. It can be seen that the nearest neighbor similar days of ideal weather are mostly adjacent days, while the time intervals of nearest neighbor similar days of non-ideal weather are relatively long.</p></sec><sec id="s4_4"><title>4.4. Design of SVR Model</title><p>This part is mainly to explore the optimal C and γ of SVR, which are usually related to the characteristics of power generation in different seasons. Grid search and cross-validation are used to find the optimal number of C and γ for SVR. This experiment uses PyCharm (python3.6) to train and optimize the SVR model on a Win 7 System personal computer with Intel core i5-3230CPU, 2.60 GHz processor and 4 GB RAM.</p><table-wrap id="table4" ><label><xref ref-type="table" rid="table4">Table 4</xref></label><caption><title> Forecasting day, similar day correlation threshold, nearest neighbor similarity day and correlation under ideal weather</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Item</th><th align="center" valign="middle" >Spring</th><th align="center" valign="middle" >Summer</th><th align="center" valign="middle" >Autumn</th><th align="center" valign="middle" >Winter</th></tr></thead><tr><td align="center" valign="middle" >Forecasting day</td><td align="center" valign="middle" >09/15/2019</td><td align="center" valign="middle" >02/09/2020</td><td align="center" valign="middle" >04/18/2019</td><td align="center" valign="middle" >08/17/2019</td></tr><tr><td align="center" valign="middle" >similar days threshold</td><td align="center" valign="middle" >0.90</td><td align="center" valign="middle" >0.88</td><td align="center" valign="middle" >0.92</td><td align="center" valign="middle" >0.86</td></tr><tr><td align="center" valign="middle" >nearest neighbor similarity day</td><td align="center" valign="middle" >09/14/2019</td><td align="center" valign="middle" >02/08/2020</td><td align="center" valign="middle" >04/15/2019</td><td align="center" valign="middle" >08/16/2019</td></tr><tr><td align="center" valign="middle" >Correlation degree</td><td align="center" valign="middle" >0.9295</td><td align="center" valign="middle" >0.9452</td><td align="center" valign="middle" >0.9913</td><td align="center" valign="middle" >0.9579</td></tr></tbody></table></table-wrap><table-wrap id="table5" ><label><xref ref-type="table" rid="table5">Table 5</xref></label><caption><title> Forecasting day, similar day correlation threshold, nearest neighbor similarity day and correlation under non-ideal weather</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Item</th><th align="center" valign="middle" >Spring</th><th align="center" valign="middle" >Summer</th><th align="center" valign="middle" >Autumn</th><th align="center" valign="middle" >Winter</th></tr></thead><tr><td align="center" valign="middle" >Forecasting day</td><td align="center" valign="middle" >11/02/2019</td><td align="center" valign="middle" >02/25/2020</td><td align="center" valign="middle" >04/28/2019</td><td align="center" valign="middle" >08/07/2019</td></tr><tr><td align="center" valign="middle" >similar days threshold</td><td align="center" valign="middle" >0.92</td><td align="center" valign="middle" >0.87</td><td align="center" valign="middle" >0.91</td><td align="center" valign="middle" >0.91</td></tr><tr><td align="center" valign="middle" >nearest neighbor similarity day</td><td align="center" valign="middle" >10/18/2019</td><td align="center" valign="middle" >02/05/2020</td><td align="center" valign="middle" >04/17/2019</td><td align="center" valign="middle" >07/12/2019</td></tr><tr><td align="center" valign="middle" >Correlation degree</td><td align="center" valign="middle" >0.9295</td><td align="center" valign="middle" >0.9452</td><td align="center" valign="middle" >0.9913</td><td align="center" valign="middle" >0.9579</td></tr></tbody></table></table-wrap><p>It can be seen from <xref ref-type="table" rid="table6">Table 6</xref> and <xref ref-type="table" rid="table7">Table 7</xref> that the optimal training time of the ideal weather model for each season is 2.0342, 1.9506, 2.3272, 0.6826 s, and the average time is 1.74865 s. And the optimal training time of the model for each season of non-ideal weather is 0.2490, 0.2400, 0.2240, 0.2060 s, and the average time is 0.22975 s. The number of similar days matched has a greater impact on the training optimization time. Comparing <xref ref-type="table" rid="table6">Table 6</xref> and <xref ref-type="table" rid="table7">Table 7</xref>, it can be found that because the data complexity of non-ideal weather is higher than that of ideal weather, the C of non-ideal weather is generally larger than that of ideal weather, and the γ of non-ideal weather is generally smaller than that of ideal weather.</p></sec><sec id="s4_5"><title>4.5. Feature Selection</title><p>In order to select the appropriate input feature, GRA and Pearson correlation analysis is performed between the power generation and various meteorological factors. The historical power and meteorological data for the year from March 1, 2018 to February 28, 2019 are used for analysis. The result is shown in <xref ref-type="fig" rid="fig3">Figure 3</xref>. The definition of Pearson correlation coefficient is as follows:</p><table-wrap id="table6" ><label><xref ref-type="table" rid="table6">Table 6</xref></label><caption><title> The optimal SVR structure for each season under ideal weather</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Structure</th><th align="center" valign="middle" >Spring</th><th align="center" valign="middle" >Summer</th><th align="center" valign="middle" >Autumn</th><th align="center" valign="middle" >Winter</th></tr></thead><tr><td align="center" valign="middle" >Training data</td><td align="center" valign="middle" >(564, 2)</td><td align="center" valign="middle" >(408, 2)</td><td align="center" valign="middle" >(612, 2)</td><td align="center" valign="middle" >(156, 2)</td></tr><tr><td align="center" valign="middle" >C</td><td align="center" valign="middle" >100,000.00</td><td align="center" valign="middle" >100,000.0</td><td align="center" valign="middle" >100,000.00</td><td align="center" valign="middle" >1.00</td></tr><tr><td align="center" valign="middle" >γ</td><td align="center" valign="middle" >0.0001</td><td align="center" valign="middle" >0.0001</td><td align="center" valign="middle" >0.002</td><td align="center" valign="middle" >1.00</td></tr><tr><td align="center" valign="middle" >Time (s)</td><td align="center" valign="middle" >2.0342</td><td align="center" valign="middle" >1.9506</td><td align="center" valign="middle" >2.3272</td><td align="center" valign="middle" >0.6826</td></tr></tbody></table></table-wrap><table-wrap id="table7" ><label><xref ref-type="table" rid="table7">Table 7</xref></label><caption><title> The optimal SVR structure for each season under non-ideal weather</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Structure</th><th align="center" valign="middle" >Spring</th><th align="center" valign="middle" >Summer</th><th align="center" valign="middle" >Autumn</th><th align="center" valign="middle" >Winter</th></tr></thead><tr><td align="center" valign="middle" >Training data</td><td align="center" valign="middle" >(108, 3)</td><td align="center" valign="middle" >(180, 3)</td><td align="center" valign="middle" >(144, 3)</td><td align="center" valign="middle" >(60, 3)</td></tr><tr><td align="center" valign="middle" >C</td><td align="center" valign="middle" >1,000,000.0</td><td align="center" valign="middle" >10,000.0</td><td align="center" valign="middle" >1,000,000.0</td><td align="center" valign="middle" >1,000,000.0</td></tr><tr><td align="center" valign="middle" >γ</td><td align="center" valign="middle" >0.000001</td><td align="center" valign="middle" >0.0001</td><td align="center" valign="middle" >0.00001</td><td align="center" valign="middle" >0.000001</td></tr><tr><td align="center" valign="middle" >Time (s)</td><td align="center" valign="middle" >0.2490</td><td align="center" valign="middle" >0.2400</td><td align="center" valign="middle" >0.2240</td><td align="center" valign="middle" >0.2060</td></tr></tbody></table></table-wrap><p>ρ X , Y = N ∑ X Y − ∑ X ∑ Y N ∑ X 2 − ( ∑ X ) 2 N ∑ Y 2 − ( ∑ Y ) 2 (19)</p><p>where, X and Y are meteorological factors and photovoltaic output power respectively, and N is the number of sampling points per day.</p><p>It can be seen from <xref ref-type="fig" rid="fig5">Figure 5</xref> that the Pearson correlation coefficients between photovoltaic power generation and T, RH, GHI, and DHI are 0.35, −0.41, 0.97, and 0.35, respectively. GHI has the greatest impact on photovoltaic output, and there is a negative correlation between relative humidity and photovoltaic power. The GRA correlations between photovoltaic power generation and T, RH, GHI, and DHI are 0.68, 0.60, 0.88 and 0.67, respectively. GHI still has the largest impact on photovoltaic output.</p><p>Based on the above analysis, the paper proposes 10 feature combinations. The prefix N represents the nearest neighbor similar day, P, G, and M respectively represent Power, GHI and meteorological factor eigenvalues. For example, NG_MG represents the nearest neighbor day GHI and predicted day meteorological factor eigenvalues and GHI.</p><p>For ideal weather, due to its high prediction accuracy, the main consideration for the selection of its input features is to select a feature combination that is easier to obtain and requires less data accuracy while ensuring sufficient prediction accuracy. Therefore, the input features of the ideal weather are power of the nearest neighbor similar day and 12 meteorological factor eigenvalues of the forecast day (NP_M). For non-ideal weather, the main goal of feature selection is to improve the prediction accuracy. Tables 8-10 show the evaluation of 10 feature combinations of non-ideal weather in each season. The best performance of the evaluation indicators in the table is bolded.</p><p>It can be seen from Tables 8-10 that the MAE of NG_MG feature combination in each season is 1.3733, 2.0817, 1.6475, and 2.2323 kW, respectively. And</p><table-wrap id="table8" ><label><xref ref-type="table" rid="table8">Table 8</xref></label><caption><title> MAE (in kW) evaluation of 10 feature combinations</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Feature</th><th align="center" valign="middle" >Spring</th><th align="center" valign="middle" >Summer</th><th align="center" valign="middle" >Autumn</th><th align="center" valign="middle" >Winter</th><th align="center" valign="middle" >Average</th></tr></thead><tr><td align="center" valign="middle" >NG_MG</td><td align="center" valign="middle" >1.3733</td><td align="center" valign="middle" >2.0817</td><td align="center" valign="middle" >1.6475</td><td align="center" valign="middle" >2.2323</td><td align="center" valign="middle" >1.8337</td></tr><tr><td align="center" valign="middle" >NMP_M</td><td align="center" valign="middle" >6.9802</td><td align="center" valign="middle" >5.8165</td><td align="center" valign="middle" >9.2623</td><td align="center" valign="middle" >3.1488</td><td align="center" valign="middle" >6.3020</td></tr><tr><td align="center" valign="middle" >NP_M</td><td align="center" valign="middle" >5.2198</td><td align="center" valign="middle" >6.7131</td><td align="center" valign="middle" >10.2075</td><td align="center" valign="middle" >2.7698</td><td align="center" valign="middle" >6.2275</td></tr><tr><td align="center" valign="middle" >NG_M</td><td align="center" valign="middle" >3.6054</td><td align="center" valign="middle" >6.3498</td><td align="center" valign="middle" >8.3381</td><td align="center" valign="middle" >3.6163</td><td align="center" valign="middle" >5.4774</td></tr><tr><td align="center" valign="middle" >NM_M</td><td align="center" valign="middle" >12.1144</td><td align="center" valign="middle" >12.9347</td><td align="center" valign="middle" >15.7812</td><td align="center" valign="middle" >16.1616</td><td align="center" valign="middle" >14.2480</td></tr><tr><td align="center" valign="middle" >NMPG_MG</td><td align="center" valign="middle" >2.7202</td><td align="center" valign="middle" >2.0359</td><td align="center" valign="middle" >3.9729</td><td align="center" valign="middle" >2.0185</td><td align="center" valign="middle" >2.6869</td></tr><tr><td align="center" valign="middle" >NPG_MG</td><td align="center" valign="middle" >2.9866</td><td align="center" valign="middle" >2.3195</td><td align="center" valign="middle" >4.2225</td><td align="center" valign="middle" >2.1003</td><td align="center" valign="middle" >2.9072</td></tr><tr><td align="center" valign="middle" >NP_MG</td><td align="center" valign="middle" >2.5947</td><td align="center" valign="middle" >2.2328</td><td align="center" valign="middle" >4.1497</td><td align="center" valign="middle" >2.4596</td><td align="center" valign="middle" >2.8592</td></tr><tr><td align="center" valign="middle" >NG_G</td><td align="center" valign="middle" >2.8937</td><td align="center" valign="middle" >2.2657</td><td align="center" valign="middle" >3.3403</td><td align="center" valign="middle" >1.6866</td><td align="center" valign="middle" >2.5466</td></tr><tr><td align="center" valign="middle" >MG</td><td align="center" valign="middle" >1.6100</td><td align="center" valign="middle" >3.3220</td><td align="center" valign="middle" >3.7172</td><td align="center" valign="middle" >1.6253</td><td align="center" valign="middle" >2.5686</td></tr></tbody></table></table-wrap><table-wrap id="table9" ><label><xref ref-type="table" rid="table9">Table 9</xref></label><caption><title> RMSE (in kW) evaluation of 10 feature combinations</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Feature</th><th align="center" valign="middle" >Spring</th><th align="center" valign="middle" >Summer</th><th align="center" valign="middle" >Autumn</th><th align="center" valign="middle" >Winter</th><th align="center" valign="middle" >Average</th></tr></thead><tr><td align="center" valign="middle" >NG_MG</td><td align="center" valign="middle" >1.4699</td><td align="center" valign="middle" >2.6625</td><td align="center" valign="middle" >1.8700</td><td align="center" valign="middle" >2.5492</td><td align="center" valign="middle" >2.1379</td></tr><tr><td align="center" valign="middle" >NMP_M</td><td align="center" valign="middle" >7.2800</td><td align="center" valign="middle" >6.3683</td><td align="center" valign="middle" >10.2603</td><td align="center" valign="middle" >3.8242</td><td align="center" valign="middle" >6.9332</td></tr><tr><td align="center" valign="middle" >NP_M</td><td align="center" valign="middle" >6.0625</td><td align="center" valign="middle" >7.8041</td><td align="center" valign="middle" >10.6872</td><td align="center" valign="middle" >3.2925</td><td align="center" valign="middle" >6.9616</td></tr><tr><td align="center" valign="middle" >NG_M</td><td align="center" valign="middle" >4.2990</td><td align="center" valign="middle" >8.1207</td><td align="center" valign="middle" >10.2338</td><td align="center" valign="middle" >4.5935</td><td align="center" valign="middle" >6.8118</td></tr><tr><td align="center" valign="middle" >NW_M</td><td align="center" valign="middle" >12.8652</td><td align="center" valign="middle" >14.9801</td><td align="center" valign="middle" >17.5154</td><td align="center" valign="middle" >18.6947</td><td align="center" valign="middle" >16.0138</td></tr><tr><td align="center" valign="middle" >NMPG_MG</td><td align="center" valign="middle" >2.9276</td><td align="center" valign="middle" >2.4981</td><td align="center" valign="middle" >4.3010</td><td align="center" valign="middle" >2.3739</td><td align="center" valign="middle" >3.0251</td></tr><tr><td align="center" valign="middle" >NPG_MG</td><td align="center" valign="middle" >3.1398</td><td align="center" valign="middle" >2.7112</td><td align="center" valign="middle" >4.4120</td><td align="center" valign="middle" >2.4219</td><td align="center" valign="middle" >3.1712</td></tr><tr><td align="center" valign="middle" >NP_MG</td><td align="center" valign="middle" >3.0082</td><td align="center" valign="middle" >2.6862</td><td align="center" valign="middle" >4.2862</td><td align="center" valign="middle" >2.8474</td><td align="center" valign="middle" >3.2070</td></tr><tr><td align="center" valign="middle" >NG_G</td><td align="center" valign="middle" >3.1922</td><td align="center" valign="middle" >2.7433</td><td align="center" valign="middle" >3.4749</td><td align="center" valign="middle" >1.9734</td><td align="center" valign="middle" >2.8459</td></tr><tr><td align="center" valign="middle" >MG</td><td align="center" valign="middle" >1.7974</td><td align="center" valign="middle" >3.7430</td><td align="center" valign="middle" >3.9093</td><td align="center" valign="middle" >2.0774</td><td align="center" valign="middle" >2.8818</td></tr></tbody></table></table-wrap><table-wrap id="table10" ><label><xref ref-type="table" rid="table1">Table 1</xref>0</label><caption><title> R<sup>2</sup> evaluation of 10 feature combinations</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Feature</th><th align="center" valign="middle" >Spring</th><th align="center" valign="middle" >Summer</th><th align="center" valign="middle" >Autumn</th><th align="center" valign="middle" >Winter</th><th align="center" valign="middle" >Average</th></tr></thead><tr><td align="center" valign="middle" >NG_MG</td><td align="center" valign="middle" >0.9918</td><td align="center" valign="middle" >0.9731</td><td align="center" valign="middle" >0.9926</td><td align="center" valign="middle" >0.9814</td><td align="center" valign="middle" >0.9847</td></tr><tr><td align="center" valign="middle" >NMP_M</td><td align="center" valign="middle" >0.7987</td><td align="center" valign="middle" >0.8464</td><td align="center" valign="middle" >0.7770</td><td align="center" valign="middle" >0.9582</td><td align="center" valign="middle" >0.8451</td></tr><tr><td align="center" valign="middle" >NP_M</td><td align="center" valign="middle" >0.8604</td><td align="center" valign="middle" >0.7693</td><td align="center" valign="middle" >0.7580</td><td align="center" valign="middle" >0.9690</td><td align="center" valign="middle" >0.8392</td></tr><tr><td align="center" valign="middle" >NG_M</td><td align="center" valign="middle" >0.9298</td><td align="center" valign="middle" >0.7502</td><td align="center" valign="middle" >0.7781</td><td align="center" valign="middle" >0.9396</td><td align="center" valign="middle" >0.8494</td></tr><tr><td align="center" valign="middle" >NM_M</td><td align="center" valign="middle" >0.3714</td><td align="center" valign="middle" >0.1499</td><td align="center" valign="middle" >0.3501</td><td align="center" valign="middle" >−1e−5</td><td align="center" valign="middle" >0.2178</td></tr><tr><td align="center" valign="middle" >NMPG_MG</td><td align="center" valign="middle" >0.9675</td><td align="center" valign="middle" >0.9764</td><td align="center" valign="middle" >0.9608</td><td align="center" valign="middle" >0.9839</td><td align="center" valign="middle" >0.9721</td></tr><tr><td align="center" valign="middle" >NPG_MG</td><td align="center" valign="middle" >0.9626</td><td align="center" valign="middle" >0.9722</td><td align="center" valign="middle" >0.9588</td><td align="center" valign="middle" >0.9832</td><td align="center" valign="middle" >0.9692</td></tr><tr><td align="center" valign="middle" >NP_MG</td><td align="center" valign="middle" >0.9656</td><td align="center" valign="middle" >0.9727</td><td align="center" valign="middle" >0.9611</td><td align="center" valign="middle" >0.9768</td><td align="center" valign="middle" >0.9690</td></tr><tr><td align="center" valign="middle" >NG_G</td><td align="center" valign="middle" >0.9613</td><td align="center" valign="middle" >0.9715</td><td align="center" valign="middle" >0.9744</td><td align="center" valign="middle" >0.9889</td><td align="center" valign="middle" >0.9740</td></tr><tr><td align="center" valign="middle" >MG</td><td align="center" valign="middle" >0.9877</td><td align="center" valign="middle" >0.9469</td><td align="center" valign="middle" >0.9676</td><td align="center" valign="middle" >0.9877</td><td align="center" valign="middle" >0.9725</td></tr></tbody></table></table-wrap><p>the average MAE is 1.8337 kW, which is the smallest among all feature combinations. The RMSE of the NG_MG feature combination in each season was 1.4699, 2.6625, 1.8700, 2.5492 kW. And the average RMSE was 2.1379 kW, the best performance among all the feature combinations. The R<sup>2</sup> of each season of NG_MG feature combination is 99.18%, 97.31%, 99.26% and 98.14%. And the average R<sup>2</sup> is 98.47%, which is the highest degree of fit among all feature combinations. From the perspective of comprehensive performance, NG_MG feature combination has higher prediction accuracy and robustness, so NG_MG is selected as the input feature of non-ideal weather.</p></sec><sec id="s4_6"><title>4.6. Forecasting Results and Discussion</title><p><xref ref-type="fig" rid="fig4">Figure 4</xref> shows the prediction results of the models in each season under ideal</p><p>weather. The feature combination is NP_M. The average value of R<sup>2</sup> is 0.9966. The MAE are 1.4521, 1.4661, 0.7120, and 0.2132 kW respectively. The average value of RMSE is 0.9608 kW. The proposed model’s MAE enhancement with respect to the SVR model is 45.51%, 25.37%, 81.21%, 91.72%, respectively. The presented model’s RMSE improvement relative to the SVR model is 41.63%, 38.80%, 77.97%, 90.37%, respectively. The average R<sup>2</sup> of the proposed model is also better than the SVR model. And this method reduced the average training time by 77.27% compared with the standard SVR model.</p><p>Under non-ideal weather, the forecast results of the HKGSVR model and the other five forecast models for the four seasons are shown in Figures 5-8. It can be seen from the figure that the HKGSVR model has the highest degree of fit in each season.</p><p>From the spring forecast results in <xref ref-type="fig" rid="fig5">Figure 5</xref>, it can be seen that HKGSVR has the highest degree of fit, HKGLSTM is the second, and the SVR trend is more consistent with the predicted day. From the summer forecast results in <xref ref-type="fig" rid="fig6">Figure 6</xref>, it can be seen that each point of HKGSVR has a high degree of fit, and HKGLSTM has a good performance except for one point that has a lower degree of fit. Through the observation of the autumn forecast results in <xref ref-type="fig" rid="fig7">Figure 7</xref>, HKGSVR still performs best, and the trends of SVR and HKGARIMA are more consistent with the forecasted day. From the winter forecast results in <xref ref-type="fig" rid="fig8">Figure 8</xref>, it is found that both HKGSVR and SVR perform better, and HKGLSTM and HKGBP have poor performance due to less training data.</p><p>According to the MAE value of each model given in <xref ref-type="fig" rid="fig9">Figure 9</xref>, the HKGSVR's average value of MAE is 1.8337 kW, which is the minimum value of all models. The presented model’s average MAE improvement relative to the compared five models (SVR, HKGLSTM, HKGBP, HKGLR, HKGARIMA) are 26.61%, 41.80%, 58.18%, 57.36%, 52.98%, respectively.</p><p>Observation in <xref ref-type="fig" rid="fig1">Figure 1</xref>0 finds that the average RMSE of the HKGSVR model is 2.1379 kW, which is the best value among all models. The proposed model’s average RMSE enhancement with respect to the compared five models is 24.13%, 52.12%, 59.87%, 61.37%, 52.66%, respectively.</p><p>Comparing the R<sup>2</sup> values of the models shown in <xref ref-type="table" rid="table1">Table 1</xref>1 shows that the average R<sup>2</sup> of the proposed model is 0.9847, which is better than other models.</p><table-wrap id="table11" ><label><xref ref-type="table" rid="table1">Table 1</xref>1</label><caption><title> Daily R<sup>2</sup> comparison results of 6 models</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Models</th><th align="center" valign="middle" >Spring</th><th align="center" valign="middle" >Summer</th><th align="center" valign="middle" >Autumn</th><th align="center" valign="middle" >Winter</th><th align="center" valign="middle" >Average</th></tr></thead><tr><td align="center" valign="middle" >HKGSVR</td><td align="center" valign="middle" >0.9918</td><td align="center" valign="middle" >0.9731</td><td align="center" valign="middle" >0.9926</td><td align="center" valign="middle" >0.9814</td><td align="center" valign="middle" >0.9847</td></tr><tr><td align="center" valign="middle" >SVR</td><td align="center" valign="middle" >0.9586</td><td align="center" valign="middle" >0.9650</td><td align="center" valign="middle" >0.9886</td><td align="center" valign="middle" >0.9804</td><td align="center" valign="middle" >0.9732</td></tr><tr><td align="center" valign="middle" >HKGLSTM</td><td align="center" valign="middle" >0.9406</td><td align="center" valign="middle" >0.9468</td><td align="center" valign="middle" >0.9486</td><td align="center" valign="middle" >0.9369</td><td align="center" valign="middle" >0.9432</td></tr><tr><td align="center" valign="middle" >HKGBP</td><td align="center" valign="middle" >0.9242</td><td align="center" valign="middle" >0.9109</td><td align="center" valign="middle" >0.9001</td><td align="center" valign="middle" >0.9431</td><td align="center" valign="middle" >0.9196</td></tr><tr><td align="center" valign="middle" >HKGARIMA</td><td align="center" valign="middle" >0.9424</td><td align="center" valign="middle" >0.9308</td><td align="center" valign="middle" >0.9400</td><td align="center" valign="middle" >0.9544</td><td align="center" valign="middle" >0.9419</td></tr><tr><td align="center" valign="middle" >HKGLR</td><td align="center" valign="middle" >0.9374</td><td align="center" valign="middle" >0.8685</td><td align="center" valign="middle" >0.9323</td><td align="center" valign="middle" >0.9038</td><td align="center" valign="middle" >0.9105</td></tr></tbody></table></table-wrap><p>Under non-ideal weather, compared to the standard SVR (the average training optimization time is 11.5389 s), the average training optimization time of the HKGSVR model is 0.2225 s, which is 98.07% less than the standard SVR.</p></sec></sec><sec id="s5"><title>5. Conclusion</title><p>A hybrid day-ahead photovoltaic power generation prediction model based on K-means++, GRA and SVR is proposed. The historical power data are clustered by multi-index K-means++, and divided into ideal weather clusters and non-ideal weather clusters according to the average power of each cluster. And it chooses the appropriate feature combination for different weather, different feature combination which has a greater impact on the model performance. It also uses GRA to match the similar day and the nearest neighbor similar day of the prediction day to improve the prediction accuracy and reduce the training optimization time of the model. Compared with the standard SVR model under ideal weather, the HKGSVR model not only improves the prediction accuracy but also greatly shortens the training time. Under non-ideal weather, the average MAE, RMSE and R<sup>2</sup> of the proposed model are 1.8337, 2.1379 kW and 98.47%, respectively, which have better performance than the other five models. And the training time is 0.2225 s, which is 98.07% less than the standard SVR. When there are more accurate forecasting weather information and less training data, the HKGSVR model has higher forecast accuracy. In general, HKGSVR has higher accuracy, shorter training time and better generalization performance. Therefore, the model can be used to predict the daily power generation of photovoltaic power plants. However, in terms of model structure optimization, this paper uses grid search, so there is room for improvement in the optimization speed and search range, which will be the direction of further research.</p></sec><sec id="s6"><title>Conflicts of Interest</title><p>The authors declare no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s7"><title>Cite this paper</title><p>Lin, J.M. and Li, H.M. (2021) A Hybrid K-Means-GRA-SVR Model Based on Feature Selection for Day-Ahead Prediction of Photovoltaic Power Generation. Journal of Computer and Communications, 9, 91-111. https://doi.org/10.4236/jcc.2021.911007</p></sec></body><back><ref-list><title>References</title><ref id="scirp.113211-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Tsai, Y.C., Chan, Y.K., Ko, Y.K., et al. (2018) Integrated Operation of Renewable Energy Sources and Water Resources. Energy Conversion and Management, 160, 439-454. https://doi.org/10.1016/j.enconman.2018.01.062</mixed-citation></ref><ref id="scirp.113211-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Shah, A.S.B.M., Yokoyama, H. and Kakimoto, N. (2015) High-Precision Forecasting Model of Solar Irradiance Based on Grid Point Value Data Analysis for an Efficient Photovoltaic System. IEEE Transactions on Sustainable Energy, 6, 474-481. https://doi.org/10.1109/TSTE.2014.2383398</mixed-citation></ref><ref id="scirp.113211-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Wang, F., Zhen, Z., Liu, C., et al. (2018) Image Phase Shift Invariance Based Cloud Motion Displacement Vector Calculation Method for Ultra-Short-Term Solar PV Power Forecasting. Energy Conversion and Management, 157, 123-135. https://doi.org/10.1016/j.enconman.2017.11.080</mixed-citation></ref><ref id="scirp.113211-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Koo, C., Hong, T., Jeong, K., et al. (2017) Development of the Smart Photovoltaic System Blind and Its Impact on Net-Zero Energy Solar Buildings Using Technical-Economic-Political Analyses. Energy, 124, 382-396. https://doi.org/10.1016/j.energy.2017.02.088</mixed-citation></ref><ref id="scirp.113211-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">IRENA (2020) Renewable Energy Statistics 2020. The International Renewable Energy Agency, Abu Dhabi.</mixed-citation></ref><ref id="scirp.113211-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Jiang, H. and Dong, Y. (2017) Forecast of Hourly Global Horizontal Irradiance Based on Structured Kernel Support Vector Machine: A Case Study of Tibet Area in China. Energy Conversion and Management, 142, 307-321. https://doi.org/10.1016/j.enconman.2017.03.054</mixed-citation></ref><ref id="scirp.113211-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">K&amp;ouml;hler, C., Steiner, A., Saint-Drenan, et al. (2017) Critical Weather Situations for Renewable Energies—Part B: Low Stratus Risk for Solar Power. Renewable Energy, 101, 794-803. https://doi.org/10.1016/j.renene.2016.09.002</mixed-citation></ref><ref id="scirp.113211-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Wang, F., Zhou, L., Ren, H., et al. (2018) Multi-Objective Optimization Model of Source-Load-Storage Synergetic Dispatch for a Building Energy Management System Based on TOU Price Demand Response. IEEE Transactions on Industry Applications, 54, 1017-1028. https://doi.org/10.1109/TIA.2017.2781639</mixed-citation></ref><ref id="scirp.113211-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Talari, S., Shafiekhah, M., Wang, F., et al. (2019) Optimal Scheduling of Demand Response in Pre-Emptive Markets Based on Stochastic Bilevel Programming Method. IEEE Transactions on Industrial Electronics, 66, 1453-1464. https://doi.org/10.1109/TIE.2017.2786288</mixed-citation></ref><ref id="scirp.113211-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Chen, Q., Wang, F., Hodge, B.M., et al. (2017) Dynamic Price Vector Formation Model-Based Automatic Demand Response Strategy for PV-Assisted EV Charging Stations. IEEE Transactions on Smart Grid, 8, 2903-2915. https://doi.org/10.1109/TSG.2017.2693121</mixed-citation></ref><ref id="scirp.113211-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Biswas, P.P., Suganthan, P.N. and Amaratunga, G.A.J. (2017) Optimal Power Flow Solutions Incorporating Stochastic Wind and Solar Power. Energy Conversion and Management, 148, 1194-1207. https://doi.org/10.1016/j.enconman.2017.06.071</mixed-citation></ref><ref id="scirp.113211-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Wang, F., Zhou, L., Wang, B., et al. (2017) Modified Chaos Particle Swarm Optimization-Based Optimized Operation Model for Stand-Alone CCHP Microgrid. Applied Sciences, 7, 754. https://doi.org/10.3390/app7080754</mixed-citation></ref><ref id="scirp.113211-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Zhen, Z., Xuan, Z., Wang, F., et al. (2019) Image Phase Shift Invariance Based Multi-Transform-Fusion Method for Cloud Motion Displacement Calculation Using Sky Images. Energy Conversion and Management, 197, 11853. https://doi.org/10.1016/j.enconman.2019.111853</mixed-citation></ref><ref id="scirp.113211-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Sobri, S., Koohi-Kamali, S. and Rahim, N.A. (2018) Solar Photovoltaic Generation Forecasting Methods: A Review. Energy Conversion and Management, 156, 459-497. https://doi.org/10.1016/j.enconman.2017.11.019</mixed-citation></ref><ref id="scirp.113211-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Wang, F., Li, K., Liu, C., et al. (2018) Synchronous Pattern Matching Principle-Based Residential Demand Response Baseline Estimation: Mechanism Analysis and Approach Description. IEEE Transactions on Smart Grid, 9, 6972-6985. https://doi.org/10.1109/TSG.2018.2824842</mixed-citation></ref><ref id="scirp.113211-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Wang, G., Su, Y. and Shu, L. (2016) One-Day-Ahead Daily Power Forecasting of Photovoltaic Systems Based on Partial Functional Linear Regression Models. Renewable Energy, 96, 469-478. https://doi.org/10.1016/j.renene.2016.04.089</mixed-citation></ref><ref id="scirp.113211-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Teo, T.T., Logenthiran, T., Woo, W.L., et al. (2016) Forecasting of Photovoltaic Power Using Extreme Learning Machine. IEEE Region 10 Conference, Singapore, 455-458. https://doi.org/10.1109/TENCON.2016.7848040</mixed-citation></ref><ref id="scirp.113211-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Felice, M.D., Petitta, M. and Ruti, P.M. (2015) Short Term Predictability of Photovoltaic Production over Italy. Renewable Energy, 80, 197-204. https://doi.org/10.1016/j.renene.2015.02.010</mixed-citation></ref><ref id="scirp.113211-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Zeng, J.W. and Qiao, W. (2013) Short-Term Solar Power Prediction Using a Support Vector Machine. Renewable Energy, 52, 118-127. https://doi.org/10.1016/j.renene.2012.10.009</mixed-citation></ref><ref id="scirp.113211-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Wang, J., Ran, R., Song, Z. and Sun, J. (2017) Short-Term Photovoltaic Power Generation Forecasting Based on Environmental Factors and GA-SVM. Journal of Electrical Engineering &amp; Technology, 12, 64-71. https://doi.org/10.5370/JEET.2017.12.1.064</mixed-citation></ref><ref id="scirp.113211-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Leva, S., Dolara, A., et al. (2017) Analysis and Validation of 24 Hours Ahead Neural Network Forecasting of Photovoltaic Output Power. Mathematics and Computers in Simulation, 131, 88-100. https://doi.org/10.1016/j.matcom.2015.05.010</mixed-citation></ref><ref id="scirp.113211-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">Du, P., Wang, J., Yang, W., et al. (2018) Multi-Step Ahead Forecasting in Electrical Power System Using a Hybrid Forecasting System. Renewable Energy, 122, 533-550. https://doi.org/10.1016/j.renene.2018.01.113</mixed-citation></ref><ref id="scirp.113211-ref23"><label>23</label><mixed-citation publication-type="other" xlink:type="simple">Lin, P., Peng, Z., Lai, Y., et al. (2018) Short-Term Power Prediction for Photovoltaic Power Plants Using a Hybrid Improved Kmeans-GRA-Elman Model Based on Multivariate Meteorological Factors and Historical Power Datasets. Energy Conversion and Management, 177, 704-717. https://doi.org/10.1016/j.enconman.2018.10.015</mixed-citation></ref><ref id="scirp.113211-ref24"><label>24</label><mixed-citation publication-type="other" xlink:type="simple">Han, Y.T., Wang, N.B., Ma, M., et al. (2019) A PV Power Interval Forecasting Based on Seasonal Model and Nonparametric Estimation Algorithm. Solar Energy, 184, 515-526. https://doi.org/10.1016/j.solener.2019.04.025</mixed-citation></ref><ref id="scirp.113211-ref25"><label>25</label><mixed-citation publication-type="other" xlink:type="simple">Gao, M., Li, J., Hong, F. and Long, D. (2019) Day-Ahead Power Forecasting in a Large-Scale Photovoltaic Plant Based on Weather Classification Using LSTM. Energy, 187, Article ID: 115838. https://doi.org/10.1016/j.energy.2019.07.168</mixed-citation></ref><ref id="scirp.113211-ref26"><label>26</label><mixed-citation publication-type="other" xlink:type="simple">Benali, L., Notton, G. and Fouilloy, A. (2019) Solar Radiation Forecasting Using Artificial Neural Network and Random Forest Methods: Application to Normal Beam, Horizontal Diffuse and Global Components. Renewable Energy, 132, 871-884. https://doi.org/10.1016/j.renene.2018.08.044</mixed-citation></ref><ref id="scirp.113211-ref27"><label>27</label><mixed-citation publication-type="other" xlink:type="simple">Ding, M., Wang, L. and Bi, R. (2011) An ANN-Based Approach for Forecasting the Power Output of Photovoltaic System. Procedia Environmental Sciences, 11, 1308-1315. https://doi.org/10.1016/j.proenv.2011.12.196</mixed-citation></ref><ref id="scirp.113211-ref28"><label>28</label><mixed-citation publication-type="other" xlink:type="simple">Shi, J., Lee, W.J., Liu, Y., et al. (2015) Forecasting Power Output of Photovoltaic Systems Based on Weather Classification and Support Vector Machines. IEEE Transactions on Industry Applications, 48, 1064-1069. https://doi.org/10.1109/TIA.2012.2190816</mixed-citation></ref><ref id="scirp.113211-ref29"><label>29</label><mixed-citation publication-type="other" xlink:type="simple">Li, P., Zang, C. and Wang, K. (2013) Photovoltaic Generation Prediction Based on Similar Days and Neural Network. Renewable Energy Resources, 31, 1-9.</mixed-citation></ref><ref id="scirp.113211-ref30"><label>30</label><mixed-citation publication-type="other" xlink:type="simple">Luo, X., Zhu, X. and Lim, E.G. (2019) A Parametric Bootstrap Algorithm for Cluster Number Determination of Load Pattern Categorization. Energy, 180, 50-60. https://doi.org/10.1016/j.energy.2019.04.089</mixed-citation></ref><ref id="scirp.113211-ref31"><label>31</label><mixed-citation publication-type="other" xlink:type="simple">Vapnik, V.N. (1998) Statistical Learning Theory. Wiley-Interscience, New York.</mixed-citation></ref><ref id="scirp.113211-ref32"><label>32</label><mixed-citation publication-type="other" xlink:type="simple">Rousseeuw, P.J. (1987) Silhouettes: A Graphical Aid to the Interpretation and Validation of Cluster Analysis. Computational and Applied Mathematics, 20, 53-65. https://doi.org/10.1016/0377-0427(87)90125-7</mixed-citation></ref><ref id="scirp.113211-ref33"><label>33</label><mixed-citation publication-type="other" xlink:type="simple">Davies, D.L. and Bouldin, D.W. (1979) A Cluster Separation Measure. IEEE Transactions on Pattern Analysis and Machine Intelligence, 1, 224-227. https://doi.org/10.1109/TPAMI.1979.4766909</mixed-citation></ref></ref-list></back></article>