<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OJS</journal-id><journal-title-group><journal-title>Open Journal of Statistics</journal-title></journal-title-group><issn pub-type="epub">2161-718X</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/ojs.2017.72015</article-id><article-id pub-id-type="publisher-id">OJS-75523</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Application of SVR Models in Stock Index Forecast Based on Different Parameter Search Methods
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Jiechao</surname><given-names>Chen</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Huazhou</surname><given-names>Chen</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Yajuan</surname><given-names>Huo</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Wanting</surname><given-names>Gao</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>College of Science, Guilin University of Technology, Guilin, China</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>hzchengut@foxmail.com(HC)</email>;</corresp></author-notes><pub-date pub-type="epub"><day>20</day><month>04</month><year>2017</year></pub-date><volume>07</volume><issue>02</issue><fpage>194</fpage><lpage>202</lpage><history><date date-type="received"><day>15,</day>	<month>March</month>	<year>2017</year></date><date date-type="rev-recd"><day>17,</day>	<month>April</month>	<year>2017</year>	</date><date date-type="accepted"><day>20,</day>	<month>April</month>	<year>2017</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Stock index forecast is regarded as a challenging task of financial time-series prediction. In this paper, the non-linear support vector regression (SVR) method was optimized for the application in stock index prediction. The parameters (C, σ) of SVR models were selected by three different methods of grid search (GRID), particle swarm optimization (PSO) and genetic algorithm (GA).The optimized parameters were used to predict the opening price of the test samples. The predictive results shown that the SVR model with GRID (GRID-SVR), the SVR model with PSO (PSO-SVR) and the SVR model with GA (GA-SVR) were capable to fully demonstrate the time-dependent trend of stock index and had the significant prediction accuracy. The minimum root mean square error (RMSE) of the GA-SVR model was 15.630, the minimum mean absolute percentage error (MAPE) equaled to 0.39% and the correspondent optimal parameters (C, σ) were identified as (45.422, 0.012). The appreciated modeling results provided theoretical and technical reference for investors to make a better trading strategy.
 
</p></abstract><kwd-group><kwd>CSI 300 Index</kwd><kwd> Support Vector Regression</kwd><kwd> Grid Search</kwd><kwd> Particle Swarm Optimization</kwd><kwd> Genetic Algorithm</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Stock index forecast is a non-linear dynamic system. There are many factors affecting the stock index, which goes with the complex fluctuation [<xref ref-type="bibr" rid="scirp.75523-ref1">1</xref>] . It had become a popular and interesting research issue to calculate the stock index to avoid the investment risk [<xref ref-type="bibr" rid="scirp.75523-ref2">2</xref>] . K-curve analysis was successfully applied to predict the trend of stock prices [<xref ref-type="bibr" rid="scirp.75523-ref3">3</xref>] , but it couldn’t accomplish the quantitative calculated. To quantitatively forecast the stock index price, traditional time-series models were introduced, such as autoregressive moving average model, which still failed in non-linear and non-stationary prediction [<xref ref-type="bibr" rid="scirp.75523-ref4">4</xref>] . At present, the research methods for stock index prediction vary from time series to artificial intelligence.</p><p>A variety of machine learning methods had been applied to stock index forecasting. They perform excellently with their merits of self-organization, self- learning and nonlinear approximation [<xref ref-type="bibr" rid="scirp.75523-ref5">5</xref>] . Support vector regression (SVR) is a kind of the simple, global optimization machine learning method, mainly described as nonlinear mapping transforming the low-dimensional data into a high-dimensional space, so that the data can be explained by a set of linear functions [<xref ref-type="bibr" rid="scirp.75523-ref6">6</xref>] . Improvement of SVR models depends on parameter optimization (regularization parameter C) and the selection of kernel functions [<xref ref-type="bibr" rid="scirp.75523-ref7">7</xref>] [<xref ref-type="bibr" rid="scirp.75523-ref8">8</xref>] . In this paper, the SVR models were trained by optimizing C and the kernel function. The radial basis function (RBF) was selected as the kernel. An over-large value of C possibly reduces the prediction ability of SVR models, and RBF kernel width (σ) commonly influences the model complexity [<xref ref-type="bibr" rid="scirp.75523-ref9">9</xref>] [<xref ref-type="bibr" rid="scirp.75523-ref10">10</xref>] .</p><p>CSI 300 index is a capitalization-weighted stock market index designed to replicate the performance of 300 stocks traded in the Shanghai and Shenzhen stock exchanges. We established SVR calibration models to predict the opening price of CSI 300 index. In order to find out the optimizational modeling parameter combination of (C, σ), we tried to respectively utilize grid search (GRID) method, particle swarm optimization (PSO) and genetic algorithm (GA) for parameter selection. Furthermore, the opening price of CSI 300 index was predicted using the best parameters of the SVR model with GRID (GRID-SVR), the SVR model with PSO (PSO-SVR) and the SVR model with GA (GA-SVR). The rest of the paper is organized as follows. The methodology demonstrated in Section 1, Section 2 illustrated the SVR modeling process and prediction results. Section 3 contains the conclusions.</p></sec><sec id="s2"><title>2. Experiments</title><sec id="s2_1"><title>2.1. Data Acquisition and Pretreatment</title><p>Daily trading data of CSI 300 index was scraped from the Wind Financial Terminal One-Stop Platform. The daily trading data (January 4, 2013 to November 30, 2016) were selected as sample set (a total of 949-day data). The data originally includes eight variables, which are opening price, ceiling price, the lowest price, closing price, charge rate, volume, turnover and the margin balance of China stock markets. We further calculated the 5-day average charge rate, the 20-day average charge rate, the 5-day average volume and the 20-day average volume as four new indicators to identify the index-changing trend. SVR models were established by the 949 samples with 12 variables to predict the opening price of the next day. The CSI 300 index daily opening price is shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>. The sample set was normalized to the [0, 1] before establishing calibration models and the SVR prediction results were renormalized.</p></sec><sec id="s2_2"><title>2.2. Model Evaluation Indices</title><p>Four fifths of the total samples were selected for training, and rest one fifth for testing. The calibration models were practiced using the training samples and the parameters are optimized. Then, the training models with their parameters were applied to predict the opening price of the test samples. The prediction per- formance was evaluated using the root mean square error (RMSE) and mean absolute percentage error (MAPE). The formulas are as follows:</p><p>RMSE = ∑ t = 1 n ( y t − y ^ t ) 2 n − 1 , MAPE = 1 n ∑ t = 1 n | y t − y ^ t | | y t |</p><p>where y t represents the real opening price, y ^ t represents the predictive value of the opening price and n is total number of sample.</p></sec><sec id="s2_3"><title>2.3. SVR Modeling and Parameter Search Method</title><p>Stock index prediction requires establishing an optimal prediction function based on stock history data and other interference to calculate the stock index price and to reveal the trend of the stock index. The function is defined as follows:</p><p>y t + 1 = f ( x 1 , x 2 , ⋯ , x t ) ,</p><p>where y<sub>t</sub><sub>+1</sub> represents the next-day opening price and x 1 , x 2 , ⋯ , x t are input samples.</p><p>The SVR algorithm is used to estimate the function. The input data is mapped onto a high-dimensional feature space using RBF kernel. SVR is formulated as minimization of the following optimization problem,</p><p>min ω , ξ i , ξ i * 1 2 ‖ ω ‖ 2 + C ∑ i = 1 l [ ( ξ i ) + ( ξ i * ) ] ,</p><p>where ω is the vector-form coefficients, ξ i and ξ i * represent the relaxation factors. This optimization formulation can be transformed into the dual problem, and its solution is given by</p><p>f ( x ) = ∑ i = 1 l ( α i * − α i ) K ( x i , x * ) + b * ,</p><p>St. K ( x , x * ) = exp ( − ‖ x − x * ‖ 2 / 2 σ 2 ) ,</p><p>where Lagrange multipliers ( α i , α i * ) are controlled by C, K(x, x<sup>*</sup>) is a RBF kernel, σ represents the kernel width. C and σ are tuned respectively in the GRID-SVR, PSO-SVR and GA-SVR models respectively, to test the predictive capabilities.</p><p>The best parameter combination of (C, σ) is determined according to the mi- nimum mean square error (MSE) under the 5-fold cross validation. The MSE is defined as follows,</p><p>MSE = 1 n ∑ t = 1 l ( y t − y ^ t ) 2 .</p><p>In the Grid search process, C and σ were pre-set in the tuning range of { 2 − 8 , 2 − 7.5 ⋯ 2 7.5 , 2 8 } . They constitute a two-dimensional dynamic network. The optimal parameters (C, σ) are determined by searching the minimum MSE in this dynamic network.</p><p>In the PSO search process, a group of particles (random solutions of C and σ) were randomly initialized [<xref ref-type="bibr" rid="scirp.75523-ref11">11</xref>] [<xref ref-type="bibr" rid="scirp.75523-ref12">12</xref>] . And the optimal solution was found by iterations, with the training result (i.e. the minimum MSE) as the fitness value. In each iteration, the velocity and position of the particle swarm were globally and individually updated by searching the minimum MSE’s. They were renewed by the following iterative equation,</p><p>v d i ( t + 1 ) = ω ⋅ v d i ( t ) + c 1 ⋅ ( p best i ( t ) − x d i ( t ) ) + c 2 ⋅ ( g best i ( t ) − x d i ( t ) ) ,</p><p>x d i ( t + 1 ) = x d i ( t ) + v d i ( t + 1 ) ,</p><p>where x d <sub> </sub>represents the position of the particle, v d represents the velocity of the particle, t is the number of iteration, i is the number of particles and ω is the velocity weight, c<sub>1</sub> and c<sub>2</sub> are learning factors. In this study, c<sub>1</sub> and c<sub>2</sub> were both valued 2, ω valued 0.5. The globally-optimal and individually-optimal velocity and position of the particle swarm were found by 100 iterative computations, so that the correspondent optimized parameters (C, σ) are determined.</p><p>In the GA search process, a population of chromosomes was randomly initialized [<xref ref-type="bibr" rid="scirp.75523-ref13">13</xref>] [<xref ref-type="bibr" rid="scirp.75523-ref14">14</xref>] [<xref ref-type="bibr" rid="scirp.75523-ref15">15</xref>] . And a new population was generated by selection, crossover, and mutation. The candidates were evolved toward better solutions with the MSE selected as the fitness value for iterative calculation. The evolution terminates when the maximum number of generations reached 100. The best SVR training model was obtained with the optimal parameters (C, σ).</p><p>The GRID, PSO and GA methods were respectively applied for SVR parameters optimization. The flow charts of the experimental algorithms are shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>. This figure depicted the entire modeling process.</p></sec></sec><sec id="s3"><title>3. Results and Discussion</title><p>The trading data of CSI 300 index from January 4, 2013 to November 30, 2016</p><p>was prepared for establishing the SVR calibration models. The records of previous consecutive 759 days were used as training samples and the remaining 190-dayrecords as the test samples. Twelve variables of each sample were input to the SVR training models and a series of the next-day opening price were predicted. The parameters (C, σ) of SVR models were optimized respectively by GRID, PSO and GA.</p><p>During the GRID-SVR modeling process, the parameters (C, σ) were tuned for searching the minimum MSE. A larger value of C and a smaller value of σ generated a smaller MSE. The MSE contours are shown in <xref ref-type="fig" rid="fig3">Figure 3</xref>. The minimum MSE was found equaling to 7.967 &#215; 10 − 5 , when C reaches 2<sup>8</sup> and σ reaches 2<sup>−7</sup> (the solid point in <xref ref-type="fig" rid="fig3">Figure 3</xref>).</p><p>For the PSO-SVR models, an initial group of particles was randomly generated and then the positions and velocities of particles were globally and individually updated by 100 iterative computations. The MSE convergence process is shown in <xref ref-type="fig" rid="fig4">Figure 4</xref>. The MSE went rapidly to the minimum value (the solid point, equaling to 8.043 &#215; 10 − 5 ) at the 24th iteration, and it kept skipping reiteratively over the optimal value, the correspondent parameters (C, σ) are (60.576, 0.010).</p><p>In the GA-SVR parametric tuning, a group of candidate solutions was randomly initialized, and reiteratively evolved to a more appreciate alternative group of solutions by genetic selection, crossover, and mutation. <xref ref-type="fig" rid="fig5">Figure 5</xref> showed the MSE for each step of iteration. After 99 iterations, the minimum MSE found as 8.072 &#215; 10 − 5 and the parameter(C, σ) are identified as (45.422, 0.012).</p><p>In summary, the optimally selected parameters (C, σ) of the SVR models were determined respectively for GRID, PSO and GA optimalization. The GRID-SVR, PSO-SVR and GA-SVR calibration models with their corresponding optimal values of (C, σ) were applied to predict the validation samples. The parameters and prediction results are both presented in <xref ref-type="table" rid="table1">Table 1</xref>. As is shown in <xref ref-type="table" rid="table1">Table 1</xref>, the</p><p>obtained low values of RMSE and MAPE indicate that GRID, PSO and GA optimalizing methods were acceptable for parametric optimization of SVR models, while the GA-SVR model was best validated. It provided the lowest RMSE of 15.630 and lowest MAPE of 0.39%. The comparison between the real and the GA-SVR predictive opening price is depicted in <xref ref-type="fig" rid="fig6">Figure 6</xref>. Therefore, the GRID-SVR, PSO- SVR and GA-SVR calibration models were feasible to accurately predict the short- term trend of opening price, and the GA-SVR had the highest prediction accuracy.</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Comparison of the predictive results for GRID-SVR, PSO-SVR and GA-SVR models</title></caption><table><tbody><thead><tr><th align="center" valign="middle"  rowspan="2"  >Model Type</th><th align="center" valign="middle"  colspan="2"  >Model Parameters</th><th align="center" valign="middle"  colspan="2"  >Predictive Results</th></tr></thead><tr><td align="center" valign="middle" >C</td><td align="center" valign="middle" >σ</td><td align="center" valign="middle" >RMSE</td><td align="center" valign="middle" >MAPE (%)</td></tr><tr><td align="center" valign="middle" >GRID Optimalization</td><td align="center" valign="middle" >256</td><td align="center" valign="middle" >0.008</td><td align="center" valign="middle" >17.954</td><td align="center" valign="middle" >0.47</td></tr><tr><td align="center" valign="middle" >PSO Optimalization</td><td align="center" valign="middle" >60.576</td><td align="center" valign="middle" >0.010</td><td align="center" valign="middle" >16.179</td><td align="center" valign="middle" >0.41</td></tr><tr><td align="center" valign="middle" >GA Optimalization</td><td align="center" valign="middle" >45.422</td><td align="center" valign="middle" >0.012</td><td align="center" valign="middle" >15.630</td><td align="center" valign="middle" >0.39</td></tr></tbody></table></table-wrap></sec><sec id="s4"><title>4. Conclusion</title><p>In this study, the SVR models with GRID, PSO and GA parametric optimization were applied to predict the opening price of CSI 300 index. The optimal parameters (C, σ) were selected as (256, 0.008), (60.576, 0.010) and (45.422, 0.012) for GRID-SVR, PSO-SVR and GA-SVR models, respectively. The optimized SVR models were applied to the validation samples, obtaining the predictive RMSE’s in the range of (15.63, 17.96), and the MAPE’s ranged from 0.39% to 0.47%. The results showed that the GRID-SVR, PSO-SVR and GA-SVR calibration models were feasible to predict the short-term trend of opening price, and the GA-SVR had the most accurate prediction. The modeling performance provided theoretical and technical reference for investors to make a better trading strategy.</p></sec><sec id="s5"><title>Acknowledgements</title><p>This work was supported by the National Natural Scientific Foundation of China (61505037), the Natural Scientific Foundation of Guangxi (2016GXNSFBA38- 0077, 2015GXNSFBA139259).</p></sec><sec id="s6"><title>Cite this paper</title><p>Chen, J.C., Chen, H.Z., Huo, Y.J. and Gao, W.T. (2017) Application of SVR Models in Stock Index Forecast Based on Different Parameter Search Methods. Open Journal of Statistics, 7, 194-202. https://doi.org/10.4236/ojs.2017.72015</p></sec></body><back><ref-list><title>References</title><ref id="scirp.75523-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Cao, Y., Liu, S. and Qiu, W. (2006) Research on Determinants of Intraday Price Movement in Shanghai Security Market. System Engineering Theory and Practice, 26, 77-85.</mixed-citation></ref><ref id="scirp.75523-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Zheng, X. and Zhang, H. (2014) Application of Combination Forecasting Model in Stock Price Prediction. Industrial Control Computer, 27, 121-122.</mixed-citation></ref><ref id="scirp.75523-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Wang, W., Yuan, Z., Xie, W. and Yang, J. (2009) Association Rule Study on Combination of Stock K Line. Journal of Chengdu University (Natural Science Edition), 28, 268-271.</mixed-citation></ref><ref id="scirp.75523-ref4"><label>4</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Zheng</surname><given-names> W. </given-names></name>,<etal>et al</etal>. (<year>2014</year>)<article-title>Short-Term Forecast of Stock Price of Shanghai Composite Index Based on ARIMA Model</article-title><source> Economic Research Guide</source><volume> 234</volume>,<fpage>136</fpage>-<lpage>137</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.75523-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Tan, P., Steinbach, M. and Kumar, V. (2011) Introduction to Data Mining. Posts &amp; Telecom Press, Beijing.</mixed-citation></ref><ref id="scirp.75523-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Vapnik, V.N. (1995) The Nature of Statistical Learning Theory. Springer Science + Business Media, New York. https://doi.org/10.1007/978-1-4757-2440-0</mixed-citation></ref><ref id="scirp.75523-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Gao, Z. and Yang, J. (2014) Financial Time Series Forecasting with Grouped Predictors Using Hierarchical Clustering and Support Vector Regression. International Journal of Grid &amp; Distributed Computing, 7, 53-64.  
https://doi.org/10.14257/ijgdc.2014.7.5.05</mixed-citation></ref><ref id="scirp.75523-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Cherkassky, V. and Ma, Y. (2004) Practical Selection of SVM Parameters and Noise Estimation for SVM Regression. Neural Networks, 17, 113-126.</mixed-citation></ref><ref id="scirp.75523-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Zhu, Y. and Zhang, Y. (2003) The Study on Some Problems of Support Vector Classifier. Computer Engineering and Applications, 39, 36-38.</mixed-citation></ref><ref id="scirp.75523-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Xiong, W. and Xu, B. (2006) Study on Optimization of SVR Parameters Selection Based on PSO. Journal of System Simulation, 18, 2442-2445.</mixed-citation></ref><ref id="scirp.75523-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Chen, C., Tian, Y. and Bie, R. (2008) Research of SVR Optimized by PSO Compared with BP Network Trained by PSO. Journal of Beijing Normal University (Natural Science), 44, 449-453.</mixed-citation></ref><ref id="scirp.75523-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Gu, W., Chai, B. and Teng, Y. (2014) Research on Support Vector Machine Based on Particle Swarm Optimization. Transactions of Beijing Institute of Technology, 34, 705-709.</mixed-citation></ref><ref id="scirp.75523-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Tang, K., Hu, G., Che, X. and Hu, L. (2010) Grid Host Load Prediction Model of Support Vector Regression Optimized by Genetic Algorithm. Journal of Jilin University (Science Edition), 48, 251-255.</mixed-citation></ref><ref id="scirp.75523-ref14"><label>14</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Dai</surname><given-names> H. </given-names></name>,<etal>et al</etal>. (<year>2008</year>)<article-title>Forecasting Population Based on Support Vector Regression with Intelligent Genetic Algorithms</article-title><source> Computer Engineering and Applications</source><volume> 44</volume>,<fpage> 9</fpage>-<lpage>11</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.75523-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Jiang, W., Wei, H., Qu, T. and Zhu, F. (2011) Predication of the Calorific Value for Fuel Coal Based on the Support Vector Regression Machine with Parameters Optimized by Genetic Algorithm. Thermal Power Generation, 40, 14-19.</mixed-citation></ref></ref-list></back></article>