<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JMF</journal-id><journal-title-group><journal-title>Journal of Mathematical Finance</journal-title></journal-title-group><issn pub-type="epub">2162-2434</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jmf.2015.54033</article-id><article-id pub-id-type="publisher-id">JMF-61461</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Business&amp;Economics</subject><subject> Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Predictive Analytics on CSI 300 Index Based on ARIMA and RBF-ANN Combined Model
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>yuxun</surname><given-names>Yang</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Xi</surname><given-names>Cheng</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>The School of Statistics and Mathematics, Zhongnan University of Economics and Law, Wuhan, China</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>yanglvxun@gmail.com(YY)</email>;<email>851998781@qq.com(XC)</email>;</corresp></author-notes><pub-date pub-type="epub"><day>05</day><month>11</month><year>2015</year></pub-date><volume>05</volume><issue>04</issue><fpage>393</fpage><lpage>400</lpage><history><date date-type="received"><day>20</day>	<month>October</month>	<year>2015</year></date><date date-type="rev-recd"><day>accepted</day>	<month>22</month>	<year>November</year>	</date><date date-type="accepted"><day>25</day>	<month>November</month>	<year>2015</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
   The time series of share prices is a highly noised, non-stationary chaotic system which possesses both linear and non-linear characteristics. The alternative of either linear or non-linear prediction models is of its inherent limitation. The paper establishes an ARIMA and RBF-ANN combined model and makes a short-term prediction on the time series of CSI 300 index by choosing various typical input variables. Results show that the combined model with multiple input indicators, compared with single ARIMA model, single RBF-ANN model, or models with single input variable, is of higher precision. 
 
</p></abstract><kwd-group><kwd>ARIMA</kwd><kwd> RBF-ANN</kwd><kwd> CSI 300 Index</kwd><kwd> Prediction Model</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Prediction model of share price or index is of both practical and theoretical significance. Y. Bai [<xref ref-type="bibr" rid="scirp.61461-ref1">1</xref>] has made a forecast on Shanghai securities composite index by using ARIMA model. J. Ouyang [<xref ref-type="bibr" rid="scirp.61461-ref2">2</xref>] and X. Yang [<xref ref-type="bibr" rid="scirp.61461-ref3">3</xref>] have made a forecast on the rise and fall of share prices by using different improved methods of BP neural network models, which has made some effects. However, share price is influenced by complicated factors such as economy, politics and society. Its mathematic model is a highly noised, non-stationary chaotic system [<xref ref-type="bibr" rid="scirp.61461-ref4">4</xref>] with both linear and non-linear characteristics. So, to individually make predictions with linear and non-linear models has some certain limitation. In addition, since the change of share price or index is related to lots of factors, it is not proper to only select historical value sequence as the input variable.</p><p>Autoregressive Integrated Moving Average (ARIMA) Model is a notable model in time series data prediction [<xref ref-type="bibr" rid="scirp.61461-ref5">5</xref>] and is widely used in econometrics study. It fits the linear characteristics of non-stationary time series to some extent. With a good ability of functional approximation, Artificial Neural Network (ANN) can explore the non-linear law of time series to a large extent. However, the conventional Back Propagation (BP) neural network has many defects, such as difficulty in determining hidden layer unit numbers, slowing rate of convergence and tendency to local minimum. Its modeling process can even be called “the process of artistic creation” [<xref ref-type="bibr" rid="scirp.61461-ref6">6</xref>] . The neural network of Radial Basis Function (RBF) can effectively improve the above-mentioned defects and has been successfully used in different fields such as none-linear functional approximation, data classification and picture processing with its fast learning and good fitting sufficiency. This paper proposes a method to combine the ARIMA model with RBF neural network model, which makes full use of the characteristics of both to predict linear and non-linear rules so as to select various indicators such as the stock price index, its amplitude and transaction volume as the combined input variables of the neural network, thus promoting the preciseness of index prediction models.</p></sec><sec id="s2"><title>2. Establishment of Prediction Models</title><sec id="s2_1"><title>2.1. ARIMA Model</title><p>ARIMA model is to turn integrated time series into stationary time series, and then recovering its lagged value of dependent variables, present values and lagged values of stochastic error terms.</p><p>If the sequence <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x6.png" xlink:type="simple"/></inline-formula> can be turned into steady sequence <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x7.png" xlink:type="simple"/></inline-formula> via d-order difference, to introduce the delay operator B, namely:</p><disp-formula id="scirp.61461-formula1671"><label>. (1)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/7-1490387x8.png"  xlink:type="simple"/></disp-formula><p>We can establish ARMA (p, q) Model, which is:</p><disp-formula id="scirp.61461-formula1672"><label>. (2)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/7-1490387x9.png"  xlink:type="simple"/></disp-formula><p>In this Equation, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x10.png" xlink:type="simple"/></inline-formula>is the stochastic error term. After introducing the delay operator B, ARMA (p, q) Model can be expressed as:</p><disp-formula id="scirp.61461-formula1673"><label>. (3)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/7-1490387x11.png"  xlink:type="simple"/></disp-formula><p>From the view of operation, the ARIMA modeling idea of Box and Jenkins can be summarized into 6 steps:</p><p>1) To carry out a stationary test on the sequences (such as the ADF unit root test). When the sequences fail to meet the condition of smoothness, turn the sequence into stationary sequences via differential transform (or logarithm differential transform);</p><p>2) To determine the order of ARIMA model by generally using Auto Correlation Function (ACF) and Partial Auto Correlation Function (PACF) for preliminary order determination, as well as Akaike Information Criterion (AIC) and Schwarz Criterion (SC) for quantization order determination;</p><p>3) To estimate the model parameters and test the significance of parameters so as to evaluate their rationality and adjust the model;</p><p>4) To carry out hypothesis test so as to check whether the diagnosis of the residual sequence is white noise. However, for the combined model illustrated below, since the residual of ARIMA model shall be fitted by other models, it shall still work when this hypothesis test is not passed;</p><p>5) To carry out diagnostic analysis so as to confirm that the obtained model is in accordance with the observed data characteristics;</p><p>6) To use the adopted model to predict the analysis.</p></sec><sec id="s2_2"><title>2.2. RBF Neural Network Model</title><sec id="s2_2_1"><title>2.2.1. RBF Neural Network Structure</title><p>RBF Neural Network is a novel, effective feed forward artificial neural network with faster learning convergence rate. Theories show that RBF Neural Network can approximate any continuous rational function and possess the Best Approximation ability which the BP neural network doesn’t possess [<xref ref-type="bibr" rid="scirp.61461-ref7">7</xref>] [<xref ref-type="bibr" rid="scirp.61461-ref8">8</xref>] . The RBF Neural Network structure includes the input layer, hidden layer and output layer. As is shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>, the input layer maps the input signals to the hidden space. The hidden layer has several hidden unit nodes and carries out non-linear mapping on input vectors via mapping function RBF. The output layer applies the mapping mode of</p><fig id="fig1"  position="float"><label><xref ref-type="fig" rid="fig1">Figure 1</xref></label><caption><title> RBF neural network structure</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/7-1490387x12.png"/></fig><p>linear weighting to the output signals of the hidden layer.</p><p>RBF is usually defined as a monotone function of Euclidean distance between two random points in space. It is a non-negative, non-linear function with radial symmetry and bi-directional attenuation. The most common is Gaussian Function:</p><disp-formula id="scirp.61461-formula1674"><label>(4)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/7-1490387x13.png"  xlink:type="simple"/></disp-formula><p>In the Equation, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x14.png" xlink:type="simple"/></inline-formula>is the i-th input sample. <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x15.png" xlink:type="simple"/></inline-formula>are respectively the center and variance of the j-th node in the hidden layer. <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x16.png" xlink:type="simple"/></inline-formula>are respectively the amount of nodes in the input and hidden layers. <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x17.png" xlink:type="simple"/></inline-formula>is Euclidean norm. Now, the linear function of the output layer is:</p><disp-formula id="scirp.61461-formula1675"><label>. (5)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/7-1490387x18.png"  xlink:type="simple"/></disp-formula><p>In the Equation, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x19.png" xlink:type="simple"/></inline-formula>is the k-th output signal. <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x20.png" xlink:type="simple"/></inline-formula>is the hyperlink weight from the j-th node in the hidden layer to the k-th output in the k-th output layer. <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x20.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x21.png" xlink:type="simple"/></inline-formula>is Gaussian Function.</p></sec><sec id="s2_2_2"><title>2.2.2. The Learning Process of RBF Neural Network</title><p>The learning process of RBF Neural Network can be classified into unsupervised and supervised learning. The unsupervised learning process is used to learn how to solve the center and variance of a primary function. The common K-means clustering algorithm is to cluster training samples and solve center of clustering<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x22.png" xlink:type="simple"/></inline-formula>, then solve the variance via the following Equation:</p><disp-formula id="scirp.61461-formula1676"><label>. (6)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/7-1490387x23.png"  xlink:type="simple"/></disp-formula><p>In the Equation: <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x24.png" xlink:type="simple"/></inline-formula>is the maximum range from the j-th data center to any other one. n is the amount of the nodes in the hidden layer.</p><p>Supervision of the learning stage aims to determine the link weight<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x25.png" xlink:type="simple"/></inline-formula>. The least square method of repeated iteration is mainly used, namely to continuously calculate the output values and compare the calculated value with the actual value and to adjust the relevant <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x25.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x26.png" xlink:type="simple"/></inline-formula> according to the OLS principle while carrying out repeated correction until the output error of mean square has reached the required preciseness. After it’s done, all the parameters can be determined, namely to be used to predict the unknown output value of some group of known input volume.</p></sec></sec><sec id="s2_3"><title>2.3. Combined Model of ARIMA and RBF-ANN</title><p>The time series <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x27.png" xlink:type="simple"/></inline-formula> of share price can be divided into three parts:</p><disp-formula id="scirp.61461-formula1677"><label>(7)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/7-1490387x28.png"  xlink:type="simple"/></disp-formula><p><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x29.png" xlink:type="simple"/></inline-formula>is a linear part to meet the conditions of autoregressive integration and moving average rules, which can be predicted with ARIMA model; <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x29.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x30.png" xlink:type="simple"/></inline-formula>is a predictable part in the non-linear part which doesn’t meet the general regularity but can be predicted with RBF Neural Network Model; <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x29.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x30.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x31.png" xlink:type="simple"/></inline-formula>is an inconsistent, unpredictable pure white noise range.</p><p>First, we need to make prediction on <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x32.png" xlink:type="simple"/></inline-formula> with ARIMA model, namely to regress Equation (2) and extract the residual sequence<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x32.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x33.png" xlink:type="simple"/></inline-formula>. This sequence can be divided into:</p><disp-formula id="scirp.61461-formula1678"><label>. (8)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/7-1490387x34.png"  xlink:type="simple"/></disp-formula><p>Next, the prediction on Sequence <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x35.png" xlink:type="simple"/></inline-formula> with the RBF Neural Network Model which has a very strong predictive ability of the non-linear rule is made. Based on the economic theory that there are some connections among the closing price and the prices before several days, as well as the ups and downs in a day (changes of opening quotation and closing quotation), amplitude (the highest and lowest changes) and volume of transaction, The paper chooses the opening price<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x35.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x36.png" xlink:type="simple"/></inline-formula>, the closing price<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x35.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x36.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x37.png" xlink:type="simple"/></inline-formula>, the maximum price<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x35.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x36.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x37.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x38.png" xlink:type="simple"/></inline-formula>, the minimum price<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x35.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x36.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x37.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x38.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x39.png" xlink:type="simple"/></inline-formula> and the volume of transaction <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x35.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x36.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x37.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x38.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x39.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x40.png" xlink:type="simple"/></inline-formula> as the combined input variables. If the lag phase is “s”, the input matrix shall be:</p><disp-formula id="scirp.61461-formula1679"><label>. (9)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/7-1490387x41.png"  xlink:type="simple"/></disp-formula><p>Correspondingly, the output matrix is:<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x42.png" xlink:type="simple"/></inline-formula>. Each training shall use the i-th line <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x42.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x43.png" xlink:type="simple"/></inline-formula> of the P matrix as its input layer and the i-th line of the T matrix as its target output value. After the training, RBF Neural Network Model which is be used to predict Sequence <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x42.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x43.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x44.png" xlink:type="simple"/></inline-formula> can be obtained. Then, the predicted value of the combined model shall be:</p><disp-formula id="scirp.61461-formula1680"><label>. (10)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/7-1490387x45.png"  xlink:type="simple"/></disp-formula><p>In the Equation: B is the lag operator; <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x46.png" xlink:type="simple"/></inline-formula>and <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x46.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x47.png" xlink:type="simple"/></inline-formula> are the parameters of ARIMA model; <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x46.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x47.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x48.png" xlink:type="simple"/></inline-formula>is the well trained neural network total mapping function.</p></sec></sec><sec id="s3"><title>3. Empirical Analysis</title><sec id="s3_1"><title>3.1. Model Effectiveness Test</title><sec id="s3_1_1"><title>3.1.1. Test Design</title><p>To test the effectiveness of the combined model, this paper designs to use the previous historical data to predict the subsequent historical data, so that we can compare the prediction with the real data. Since it is considered that the change of Chinese stock market around the end of year 2014 to the beginning of year 2015 is especially striking and hard to forecast, this paper chooses data around this period to test the model, which are, the data of 190 trading days from July 1<sup>st</sup>, 2014, to April 10<sup>th</sup>, 2015.Considering that this model is aimed at predicting only one future closing price of the day after the last day of input data, but one-day prediction can hardly test the effectiveness of the model, this test carries out a dynamic prediction. For example, using data of the 1<sup>st</sup> to the 126<sup>th</sup> day to predict closing price of the 127<sup>th</sup> day, and using data of 2<sup>nd</sup> to the 127<sup>th</sup> day to predict closing price of the 128<sup>th</sup> day, and so on. By this method, we predict 64 days’ closing price, and compare them with the real data. The data source is the Wind database.</p></sec><sec id="s3_1_2"><title>3.1.2. To Carry out Parameter Estimation on the ARIMA Section</title><p>First, we need to determine the order of ARIMA (p, d, q) Model. From the view of the sequence chart drawn with Eviews 7.2, it’s clear that the stock data don’t conform to the characteristics of zero-mean equal variance. And the probability value of the ADF unit root test is 0.9999, obviously larger than 0.05, which means the sequence is not stable. According to the proportional change characteristics of price indexes, logarithm taking shall be done on data first and then first difference shall be done on the processed data. Now the probability value of the ADF unit root test is less than 0.0001. This means that the sequence passes the stationary test. So it’s determined that d = 1. <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x49.png" xlink:type="simple"/></inline-formula>is the sequence after the difference.</p><disp-formula id="scirp.61461-formula1681"><label>. (11)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/7-1490387x50.png"  xlink:type="simple"/></disp-formula><p>Then, it requires inspecting of the multiple-order ACF and PACF images of Sequence <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x51.png" xlink:type="simple"/></inline-formula> before determining p and q. At most 10-order is considered here. As is shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>.</p><p>According to the preliminary judgment from the picture, 2- and 4-order truncations appear on the ACF and PACF pictures. So four combinations, namely p = 2, 4 and q = 2, 4 are tried. According to the calculation of AIC and SC values, it is found that AIC = −6.06868 and SC = −5.92467 when p = 4, q = 2, being the smallest. So ARIMA (4, 1, 2) Model is used. Now, the t test probability value of all regression terms are less than 0.05, the critical value. So, ARIMA model, all of whose items are obviously effective, is obtained. By the LM order autocorrelation test of the residual error, it is found that the probability value (0.9999) is obviously greater than 0.05, showing that there’s no self-correlation element in the residual sequence. Via the inverse operation of first difference and logarithm taking, the forecasting sequence <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x52.png" xlink:type="simple"/></inline-formula> can be obtained. The ARIMA model shall be:</p><disp-formula id="scirp.61461-formula1682"><label>. (12)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/7-1490387x53.png"  xlink:type="simple"/></disp-formula></sec><sec id="s3_1_3"><title>3.1.3. To Use the ARIMA and RBF-ANN Combined Model for Prediction</title><p>Remove the predicted <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x54.png" xlink:type="simple"/></inline-formula> from the original sequence <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x54.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x55.png" xlink:type="simple"/></inline-formula> and obtain the sequence<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x54.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x55.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x56.png" xlink:type="simple"/></inline-formula>. Due to four lagged items of ARIMA model, there are only 121 items of available <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x54.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x55.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x56.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x57.png" xlink:type="simple"/></inline-formula> sequences to go with them. <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x54.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x55.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x56.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x57.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x58.png" xlink:type="simple"/></inline-formula>is selected. Matrix P is input and Matrix T is output in accordance with RBF-ANN learning method structure. RBF Neural Network is established with MATLAB R2010b software and training shall be carried out with normalized P and T until the RBF Neural Network that meets the requirements of preciseness is obtained.</p><p>Now we can predict the data behind with this model. When predicting the t-th day, data of the (t − 1)th day shall be the training set of the neural network. The predicted value can be treated in an anti-normalization way to obtain<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x59.png" xlink:type="simple"/></inline-formula>. It totals up with ARIMA forecasting sequence <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x59.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x60.png" xlink:type="simple"/></inline-formula> to obtain the final forecasting sequence<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x59.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x60.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x61.png" xlink:type="simple"/></inline-formula>. The predicting outcomes are shown in <xref ref-type="fig" rid="fig3">Figure 3</xref>. The full lines in the picture are the real sequence <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x59.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x60.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x61.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x62.png" xlink:type="simple"/></inline-formula> and the dotted lines are ARIMA model forecasting sequence<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x59.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x60.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x61.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x62.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x63.png" xlink:type="simple"/></inline-formula>. Dots are forecasting sequence <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x59.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x60.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x61.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x62.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x63.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x64.png" xlink:type="simple"/></inline-formula> of ARIMA and RBF-ANN combined model. It can be seen that the forecasting sequence of combined model is highly matched with the original values and superior to ARIMA model predicted value. A white noise test is carried out on the residual sequence <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x59.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x60.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x61.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x62.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x63.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x64.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x65.png" xlink:type="simple"/></inline-formula> and the outcome is obvious. Meanwhile, the directional predictive effects of the model on the ups and downs are inspected. The accuracy rate of the prediction on the ups and downs is calculated with the following Equation.</p><disp-formula id="scirp.61461-formula1683"><label>(13)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/7-1490387x66.png"  xlink:type="simple"/></disp-formula><p>In the Equation, # is the symbolic function, which means 1 for the positive number and 0 for the rest. “n” is total predicted number. <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/7-1490387x67.png" xlink:type="simple"/></inline-formula>is calculated. It is seen that the qualitative forecasting of the model on the</p><fig id="fig2"  position="float"><label><xref ref-type="fig" rid="fig2">Figure 2</xref></label><caption><title> Self-correlation and partial self-correlation pictures of at most 10 orders in the steady sequence</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/7-1490387x68.png"/></fig><p>ups and downs is relatively accurate.</p></sec><sec id="s3_1_4"><title>3.1.4. Contrastive Analysis</title><p>To inspect the superiority of the combined model, we also fit the price sequence respectively with ARIMA model and RBF neural network model. Relative errors of the predicting outcomes of the three models over the true values are in <xref ref-type="fig" rid="fig4">Figure 4</xref>. The relative error of ARIMA and RBF-ANN combined model indicated by full lines; relative error of ARIMA model indicated by dotted lines; relative error of RBF neural network model indicated by dots. Mean Absolute Percent Error (MAPE) of the predicted value of three models is indicated as <xref ref-type="table" rid="table1">Table 1</xref> shows.</p><p>Clearly, the error fluctuation degree of ARIMA and RBF-ANN combined model proposed by this paper reaches the least and the prediction preciseness reaches the highest with the superiority.</p><p>Furthermore, the effects of further prediction on models are inspected. For the prediction of the data on the t-th day, the paper respectively uses the first t − 1, t − 2, t − 3 and t − 4 data as the training set for prediction. The obtained MAPE is shown as <xref ref-type="table" rid="table2">Table 2</xref>.</p><fig id="fig3"  position="float"><label><xref ref-type="fig" rid="fig3">Figure 3</xref></label><caption><title> The true value curve, ARIMA model prediction curve and the combined model prediction curve</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/7-1490387x69.png"/></fig><fig id="fig4"  position="float"><label><xref ref-type="fig" rid="fig4">Figure 4</xref></label><caption><title> The relative error picture of two single models and combined model</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/7-1490387x70.png"/></fig><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> The respective MAPE values of two single models and combined model</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Model</th><th align="center" valign="middle" >ARIMA</th><th align="center" valign="middle" >RBF-ANN</th><th align="center" valign="middle" >Combined Model</th></tr></thead><tr><td align="center" valign="middle" >MAPE</td><td align="center" valign="middle" >1.5691%</td><td align="center" valign="middle" >1.7772%</td><td align="center" valign="middle" >0.9071%</td></tr></tbody></table></table-wrap><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Prediction preciseness of the combined model of different lag-phase training set that is used</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Training set</th><th align="center" valign="middle" >The first t − 1</th><th align="center" valign="middle" >The first t − 2</th><th align="center" valign="middle" >The first t − 3</th><th align="center" valign="middle" >The first t − 4</th></tr></thead><tr><td align="center" valign="middle" >ARIMA Model MAPE</td><td align="center" valign="middle" >1.5691%</td><td align="center" valign="middle" >2.0427%</td><td align="center" valign="middle" >2.4952%</td><td align="center" valign="middle" >2.6728%</td></tr><tr><td align="center" valign="middle" >Combined Model MAPE</td><td align="center" valign="middle" >0.9071%</td><td align="center" valign="middle" >1.6206%</td><td align="center" valign="middle" >3.1094%</td><td align="center" valign="middle" >3.5059%</td></tr></tbody></table></table-wrap><p>It can be seen that when the training set is lagged for above 2 phases, the forecast errors are obviously higher and the errors of the combined model are more than those of the ARIMA model. It fully verifies that the chaos characteristics of stock data increase dramatically as the forecast period goes up. The neural network model is more applicable to the capture of its short-term change rules. Therefore, it is right for the paper to choose to make short-term prediction on the future phase 1.</p><p>Furthermore, to inspect the effects of this paper to use multiple index combination input volumes such as the opening price, the highest and lowest price and transaction volume to enhance the prediction preciseness when constructing RBF neural network, the paper also carries out check experiments: only using the pure closing price sequence as the input volume to make combined model prediction under the same condition without considering other factors. Now, the predicted MAPE = 1.2624%, higher than the MAPE = 0.9071% predicted by the model in this paper. This means that the combination input added by this paper has effectively elevated the prediction accuracy and is totally necessary.</p></sec></sec><sec id="s3_2"><title>3.2. Prediction for Real Future</title><p>To make it meaningful, the paper also uses the combined model to predict the real future data, which is the closing price of November 2<sup>nd</sup>, 2015. The input data are the daily opening price, closing price, highest price, lowest price and transaction volume of 202 trading days from January 1<sup>st</sup>, 2015 to October 31<sup>st</sup>, 2015. The output price of the combined model is 3504.32. To apply this model for more future days, it has to take the dynamic prediction method illustrated in the part of model effectiveness test to predict future closing prices day by day.</p></sec></sec><sec id="s4"><title>4. Conclusion</title><p>Regarding the chaos and complexity of share price fluctuation, which possesses both linear and non-linear characteristics, it’s difficult to fit all the features of the information with single linear or non-linear prediction model. The paper proposes an ARIMA and RBF-ANN combined prediction model in which several prices and transaction volume indicators are chosen as combined input variables, hence endowing the model with high precision and good operability. An empirical study on short-term prediction on CSI 300 index is implemented to verify the superiority of the combined model over the other two, the precision of this combined model has obviously exceeded the single linear or non-linear model, as well as the prediction pattern with single closing price sequence as input variables. However the short-term prediction effects of the model are obviously better than long-term prediction effects, so relations between the share price and different factors, as well as the prediction on weekly and monthly data, have yet to be studied.</p></sec><sec id="s5"><title>Cite this paper</title><p>LyuxunYang,XiCheng, (2015) Predictive Analytics on CSI 300 Index Based on ARIMA and RBF-ANN Combined Model. Journal of Mathematical Finance,05,393-400. doi: 10.4236/jmf.2015.54033</p></sec></body><back><ref-list><title>References</title><ref id="scirp.61461-ref1"><label>1</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Bai</surname><given-names> Y. </given-names></name>,<etal>et al</etal>. (<year>2009</year>)<article-title>Forecast and Analysis of Shanghai Stock Index Based upon ARIMA model</article-title><source> Science Technology and Engineering</source><volume> 9</volume>,<fpage> 4885</fpage>-<lpage>4888</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.61461-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Ouyang, J. and Lu, L. (2011) Application of Integrative Improved BP Neural Network Algorithm in Stock Price Forecast. Computer and Digital Engineering, 39, 57-59.</mixed-citation></ref><ref id="scirp.61461-ref3"><label>3</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Yang</surname><given-names> X. </given-names></name>,<etal>et al</etal>. (<year>2014</year>)<article-title>Analysis on the Share Price Prediction Based on the Main Component and BP Neural Network</article-title><source> Statistics and Decision</source><volume> 12</volume>,<fpage> 42</fpage>-<lpage>43</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.61461-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Hann, T.H. and Steurer, E. (1996) Much Ado about Nothing? Exchange Rates Forecasting: Neural Networks vs. Linear Models Using Monthly and Weekly Data. Neurocomputing, 10, 323-339. &lt;/br&gt;http://dx.doi.org/10.1016/0925-2312(95)00137-9</mixed-citation></ref><ref id="scirp.61461-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Ediger, V., Akar, S. and Ugurlu, B. (2006) Forecasting Production of Fossil Fuel Sources in Turkey Using a Comparative Regression and ARIMA Model. Energy Policy, 18, 3836-3846. &lt;/br&gt;http://dx.doi.org/10.1016/j.enpol.2005.08.023</mixed-citation></ref><ref id="scirp.61461-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Zhang, L., Cao, J. and Jiang, S. (2008) Practical Course of Neural Network. Beijing Mechanical Industry Press, Beijing.</mixed-citation></ref><ref id="scirp.61461-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Wei, M. and Yu, L. (2012) A RBF Neural Network with Optimum Learning Rates and its Application. Journal of Management Sciences in China, 15, 50-57.</mixed-citation></ref><ref id="scirp.61461-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Piggio, T. and Girosi, F. (1989) A Theory of Networks for Approximation and Learning. A.I. Memo No. 1140, Artificial Intelligence Laboratory, Massachusetts Institute of Technology, Cambridge.</mixed-citation></ref></ref-list></back></article>