<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JDAIP</journal-id><journal-title-group><journal-title>Journal of Data Analysis and Information Processing</journal-title></journal-title-group><issn pub-type="epub">2327-7211</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jdaip.2018.63006</article-id><article-id pub-id-type="publisher-id">JDAIP-85799</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Computer Science&amp;Communications</subject><subject> Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  The Study on China’s Flu Prediction Model Based on Web Search Data
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Yan</surname><given-names>Bu</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Jinhong</surname><given-names>Bai</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Zhuo</surname><given-names>Chen</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Mingjing</surname><given-names>Guo</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Fan</surname><given-names>Yang</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>School of Economy and Management, China University of Geosciences, Wuhan, China</addr-line></aff><aff id="aff2"><addr-line>School of Business, Central South University, Changsha, China</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>buyan10@126.com(MG)</email>;</corresp></author-notes><pub-date pub-type="epub"><day>03</day><month>07</month><year>2018</year></pub-date><volume>06</volume><issue>03</issue><fpage>79</fpage><lpage>92</lpage><history><date date-type="received"><day>12,</day>	<month>June</month>	<year>2018</year></date><date date-type="rev-recd"><day>1,</day>	<month>July</month>	<year>2018</year>	</date><date date-type="accepted"><day>4,</day>	<month>July</month>	<year>2018</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Influenza is a kind of infectious disease, which spreads quickly and widely. The outbreak of influenza has brought huge losses to society. In this paper, four major categories of flu keywords, “prevention phase”, “symptom phase”, “treatment phase”, and “commonly-used phrase” were set. Python web crawler was used to obtain relevant influenza data from the National Influenza Center’s influenza surveillance weekly report and Baidu Index. The establishment of support vector regression (SVR), least absolute shrinkage and selection operator (LASSO), convolutional neural networks (CNN) prediction models through machine learning, took into account the seasonal characteristics of the influenza, also established the time series model (ARMA). The results show that, it is feasible to predict influenza based on web search data. Machine learning shows a certain forecast effect in the prediction of influenza based on web search data. In the future, it will have certain reference value in influenza prediction. The ARMA(3,0) model predicts better results and has greater generalization. Finally, the lack of research in this paper and future research directions are given.
 
</p></abstract><kwd-group><kwd>Data Mining</kwd><kwd> Web Search</kwd><kwd> Machine Learning</kwd><kwd> Baidu Index</kwd><kwd> Influenza Prediction</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Influenza, referred to as the flu, is an acute respiratory infectious disease caused by influenza virus that cannot be completely controlled until now [<xref ref-type="bibr" rid="scirp.85799-ref1">1</xref>] . According to the WHO (World Health Organization) study of seasonal influenza, seasonal influenza causes about 3 to 5 million serious diseases each year, resulting in approximately 250,000 to 500,000 deaths [<xref ref-type="bibr" rid="scirp.85799-ref2">2</xref>] . From the Spanish flu (H1N1) in 1918, the Asian flu (H2N2) in 1957, the Hong Kong flu (H3N2) in 1968, and the Russian flu (H1N1) in 1977 to April 2009, the outbreak of H1N1 has caused a huge loss of human society for every outbreak of flu [<xref ref-type="bibr" rid="scirp.85799-ref3">3</xref>] [<xref ref-type="bibr" rid="scirp.85799-ref4">4</xref>] [<xref ref-type="bibr" rid="scirp.85799-ref5">5</xref>] . For all countries in the world, the prevention and control of influenza has always been a serious problem.</p><p>First of all, in order to control the spread of influenza virus and reduce the losses caused by influenza, it is necessary to use reasonable methods to predict the trend of influenza activity. However, the influenza virus has the characteristics of strong infectiousness, rapid propagation, wide spread, and antigen variability [<xref ref-type="bibr" rid="scirp.85799-ref6">6</xref>] , which brings great difficulties to prevention and monitoring. As a result, researchers in various countries are focusing more on improving the timeliness of forecasting the flu epidemic. Second, the use of more timely and accurate data sources is the main means of improving timeliness. In order to obtain influenza case data, most national influenza surveillance agencies generally conduct surveys on suspected influenza cases in hospitals. However, this method requires the collection of national influenza case data. There are complex data processing processes, heavy workload, and monitoring data lag about influenza development and other issues. Finally, in order to obtain more data on the flu cases, flu monitoring agencies used data such as telephone consultations on influenza, sales of flu-type non-prescription drugs, and page views of relevant websites to predict the incidence of influenza [<xref ref-type="bibr" rid="scirp.85799-ref7">7</xref>] . To a certain extent, it improves the accuracy and timeliness of short-term forecasting.</p><p>Nowadays, search engines are increasingly becoming the main method for people to obtain information. Web search data has become an ideal data source for influenza surveillance. In the United States, about 90 million adults annually search the Internet for health information such as disease and medicine [<xref ref-type="bibr" rid="scirp.85799-ref8">8</xref>] . Compared with other data sources, web search data has a stronger tendency and immediacy, and search keywords can directly reflects the intent of the inquirer, and the search data can be collected in a timely manner to maintain complete synchronization with the development of the flu epidemic. In addition, the search data has a wider range of survey populations. It can show the attention of all Internet users in a certain area to the flu, and the data is closer to the true whole. Using web search data to monitor epidemic disease is a faster, more accurate and low-cost way. It can be used as an auxiliary measure of traditional investigation methods to provide early warning of disease and is important for the prevention and control of infectious diseases in China and beyond.</p></sec><sec id="s2"><title>2. Literature Review</title><p>Influenza has caused great difficulties in prevention and monitoring due to its rapid mutation rate. Therefore, the most important task in influenza epidemic surveillance research is to improve the timeliness of predictions. The use of more immediate and accurate data sources is the main reason for improving timeliness [<xref ref-type="bibr" rid="scirp.85799-ref7">7</xref>] . In the era of big data, web search data has become an ideal data source for influenza surveillance. The flu monitoring application based on web search data mainly includes the following aspects.</p><sec id="s2_1"><title>2.1. Using Search Engines for Influenza Surveillance</title><p>In 2008, Polgreen et al. [<xref ref-type="bibr" rid="scirp.85799-ref9">9</xref>] used the Web search data for the first time. They used the search volume of influenza-related search terms on the Yahoo! search engine in the United States to verify the correlation between search volume and influenza mortality. Jeremy G et al. [<xref ref-type="bibr" rid="scirp.85799-ref10">10</xref>] published a flu trend monitoring research based on Google search data in Nature, which laid the theoretical foundation for the Google Flu Trends (GFT) launched by Google later [<xref ref-type="bibr" rid="scirp.85799-ref11">11</xref>] . GFT is an online flu trend online warning system based on its own search data released by Google. It provides flu trend predictions in 28 countries around the world. After the GFT was released, it was applied to influenza surveillance activities in different countries. In addition to Google’s search engine, search data can also be obtained through other methods, such as China’s Baidu Index, Weibo Micro Index, etc. Q. Yuan et al. [<xref ref-type="bibr" rid="scirp.85799-ref12">12</xref>] studied the relationship between search terms and flu trends through Baidu Index and fitted a multiple regression monitoring model. Lu Li et al. [<xref ref-type="bibr" rid="scirp.85799-ref13">13</xref>] compared and analyzed the role of Baidu Index and Sina Weibo micro index in the monitoring of influenza in China and found that the Baidu Index was more relevant to the flu epidemic. Search engine-based influenza surveillance estimates the incidence of influenza due to the search frequency of keywords alone. This can easily lead to over-sensitivity of the model, causing “overestimation” of the epidemic, as well as seasonal and geographical impacts. After it is still insufficient, it needs to improve.</p></sec><sec id="s2_2"><title>2.2. Using Social Networks for Influenza Surveillance</title><p>The prediction of events through social networks is a hot topic of big data research. In foreign countries, there are many researchers who use the social platform Twitter to do data analysis, including flu trend monitoring. Nigel Collier et al. [<xref ref-type="bibr" rid="scirp.85799-ref14">14</xref>] used SVM algorithm to analyze the epidemic situation by studying user behavior information in the information posted by users on Twitter, and compared the results with that of the CDC (United States Centers for Disease Control and Prevention), and found that it had a very strong relationship with that. Lampos, V. et al. [<xref ref-type="bibr" rid="scirp.85799-ref15">15</xref>] observed and tracked the Twitter information published by users in the UK’s most popular 49 regions. Using the flu keyword weighted filtering method, it was found that the flu episode showed strong linear correlation with the HPA’s influenza-like illness (ILI) data. Similar examples of flu predictions based on social platforms are numerous. Chen et al. [<xref ref-type="bibr" rid="scirp.85799-ref16">16</xref>] used Facebook, micro-blog, and Instagram as research data to filter textual data for flu symptoms keywords, to obtain suspected influenza users, and to associate GPS information on Instagram to geographically monitor the flu. Also as a social media in recent years, Weibo has been popular among Chinese citizens. At present, there are many researchers who are doing data mining based on Weibo, such as: analysis of social relations based on Weibo, public opinion analysis based on Weibo, and outbreak analysis based on Weibo [<xref ref-type="bibr" rid="scirp.85799-ref17">17</xref>] . However, the data did not make significant progress in the study of seasonal influenza surveillance based on Weibo.</p></sec><sec id="s2_3"><title>2.3. Using Existing Disease Surveillance Platforms for Influenza Surveillance</title><p>At present, the most representative foreign influenza surveillance platform is Flu Near You. Flu Near You is a flu monitoring and visualization system that can be intuitively displayed on maps. It is also participatory for the general public. Users can submit the relevant information about flu symptoms every week. These data for researchers better understand the spread of the flu, while ordinary citizens can also watch the surrounding communities where they live and the spread of national flu [<xref ref-type="bibr" rid="scirp.85799-ref18">18</xref>] . In China, Baidu and the Chinese Center for Disease Control and Prevention launched its disease prediction platform. The Baidu Disease Forecasting Platform provides an online map tool to show people how active certain diseases are in each region, and to make predictions about disease changes in the past 30 days and the next seven days.</p><p>Nowadays, there is no standardized flu prediction model in China and there are not many researches on the use of Web search data to study flu prediction models. This study establishes some prediction models by using Python to crawl relevant flu data together with machine learning. Considering the seasonality of influenza, a time series model has also been established, which has certain reference value for the monitoring and prevention of influenza.</p></sec></sec><sec id="s3"><title>3. Sources of Data</title><sec id="s3_1"><title>3.1. Influenza-Like Illness</title><p>A major indicator of influenza surveillance at home and abroad is the proportion of influenza-like illness (ILI). It refers to fever (body temperature ≥ 38˚C) and cough in all outpatient clinics at sentinel hospitals. Sore throat is one of the cases of acute respiratory infection [<xref ref-type="bibr" rid="scirp.85799-ref19">19</xref>] . The flu epidemic data used in this paper comes from the weekly influenza surveillance report (http://www.cnic.org.cn/) published by the China National Influenza Center website. The sample period is from the 16th week of 2016 (2016/16, starting on April 25, 2016) to the 16<sup>th</sup> week of 2018 (2018/16, April 23, 2018). The data collected in this paper are mainly the proportion of influenza-like cases in the country (The proportion of flu-like cases is the total number of ILI patients divided by the number of outpatients, expressed as ILI%).</p></sec><sec id="s3_2"><title>3.2. Web Data</title><p>In order to obtain more time-sensitive data, this paper uses Python to write a crawler program and uses Baidu’s webpage data as crawling objects. Selected more primitive search terms such as “flu vaccine”, “cold”, “flu treatment”, “flu medicine” and “H7N9”, and summarized the keywords selected by other related studies and keyword recommendations of search engines. Expand the number of words for each keyword type, sum up all valid keywords, and form an initial vocabulary. As shown in <xref ref-type="table" rid="table1">Table 1</xref>, the data were crawled according to the initial vocabulary.</p></sec></sec><sec id="s4"><title>4. Model Introduction</title><sec id="s4_1"><title>4.1. Feature Selection</title><p>At the beginning of the model establishment of data mining and machine learning algorithms, in order to minimize the problem of model deviation due to the lack of important variables, we usually choose as many independent variables as possible. However, during the actual modeling process, it is usually necessary to find the subset of independent variables that have the ability to interpret the response variables to improve the model’s ability to interpret and predict. This process is called feature selection.</p><p>Principal component analysis (PCA) is a method of dimension reduction for unsupervised learning. It requires only eigenvalue decomposition to compress and denoise the data. Therefore, this paper uses PCA algorithm to extract features of influenza keywords.</p><p>Algorithm flow: Input: n-dimensional sample set D = ( x ( 1 ) , x ( 2 ) , ⋯ , x ( m ) ) , to reduce dimension to n ′ dimension. Output: Sample set D ′ after dimension reduction.</p><p>1) Centralize all samples: x ( i ) = x ( i ) − 1 m ∑ j = 1 m x ( j ) ,</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Key words and extended primaries</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Classification</th><th align="center" valign="middle" >Key words</th><th align="center" valign="middle" >Extended words</th></tr></thead><tr><td align="center" valign="middle"  rowspan="2"  >Prevention stage</td><td align="center" valign="middle" >Flu vaccine</td><td align="center" valign="middle" >Flu vaccine side effects, Influenza vaccine necessary to fight it, Bird flu vaccine</td></tr><tr><td align="center" valign="middle" >Flu prevention</td><td align="center" valign="middle" >How to prevent flu, Prevent influenza, Influenza outbreak, Prevent influenza A</td></tr><tr><td align="center" valign="middle"  rowspan="3"  >Stage of symptoms</td><td align="center" valign="middle" >Cold</td><td align="center" valign="middle" >Gastro-intestinal flu, Influenza, Cold symptoms, Viral influenza</td></tr><tr><td align="center" valign="middle" >Respiratory infection</td><td align="center" valign="middle" >Upper respiratory tract infection, Nasal congestion, Cough, Bronchitis, Sore throat, Runny nose, Rhinitis, Pharyngitis</td></tr><tr><td align="center" valign="middle" >Fever</td><td align="center" valign="middle" >Fever, High fever, Headache, Dizziness, Fatigue, Fever, Chills</td></tr><tr><td align="center" valign="middle"  rowspan="2"  >Treatment stage</td><td align="center" valign="middle" >Flu treatment</td><td align="center" valign="middle" >A stream of treatment, What medicine to eat flu</td></tr><tr><td align="center" valign="middle" >Cold medicine</td><td align="center" valign="middle" >Baijiahei, Contac, Tylenol, Gankang, Amoxicillin, Cough, Antipyretics, Cephalosporins, Oseltamivir</td></tr><tr><td align="center" valign="middle"  rowspan="3"  >Commonly used words</td><td align="center" valign="middle" >H1N1</td><td align="center" valign="middle" >H1N1 flu, H7N9, H7N9 flu</td></tr><tr><td align="center" valign="middle" >Influenza-A</td><td align="center" valign="middle" >Influenza A, Type A H1N1 flu, Type A flu symptoms, What is Influenza A</td></tr><tr><td align="center" valign="middle" >Influenza</td><td align="center" valign="middle" >Flu virus, Swine flu, Bird flu</td></tr></tbody></table></table-wrap><p>2) Calculate the sample’s covariance matrix XX<sup>T</sup> ,</p><p>3) Perform eigenvalue decomposition on the matrix XX<sup>T</sup> , and take out the eigenvector corresponding to the largest n ′ eigenvalue ( w 1 , w 2 , ⋯ , w n ′ ) , After all the eigenvectors are normalized, they form a matrix of eigenvector W,</p><p>4) Transform each sample x ( i ) in the sample set into a new sample z ( i ) = W T x ( i ) ,</p><p>5) Get output sample set D ′ = ( z ( 1 ) , z ( 2 ) , ⋯ , z ( m ) ) .</p></sec><sec id="s4_2"><title>4.2. Model Introduction</title><sec id="s4_2_1"><title>4.2.1. Support Vector Regression</title><p>Provided that the training sample is ( x i , y i ) and ( i = 1 , 2 , ⋯ , l ) , the simplest support vector regression (SVR) uses a linear function f ( x , ω ) = ( ω ⋅ x ) + b to model the sample points Together, where w and b are the normal vector and the offset of the linear regression function respectively. Assume that all training data are fitted with a linear function without errors under ε . Solve the following optimization problem:</p><p>min α Φ ( w ) = 1 2 ‖ w ‖ 2 (1)</p><p>s .t . { ( w ⋅ x i + b ) − y i ≤ ε , i = 1 , 2 , ⋯ , l y i − ( w ⋅ x i + b ) ≤ ε , i = 1 , 2 , ⋯ , l (2)</p><p>When we cannot fully satisfy the above two-condition constraint, we introduce the slack variables ξ i , ξ i * and the penalty parameter C to “soften” the same as the linear inseparable support vector classification. The original optimization problem becomes:</p><p>min a , ξ i , ξ i * , b 1 2 ‖ w ‖ 2 + C ∗ 1 l ∑ i = 1 l ( ξ i + ξ i * ) (3)</p><p>s .t . { ( w ⋅ x i + b ) − y i ≤ ε + ξ i , i = 1 , 2 , ⋯ , l y i − ( w ⋅ x i + b ) ≤ ε + ξ i * , i = 1 , 2 , ⋯ , l ξ i ≥ 0 , ξ i * ≥ 0 , i = 1 , 2 , ⋯ , l (4)</p><p>To solve the problem, you can get the normal vector and the regression function of the regression function:</p><p>w = ∑ i = 1 l ( α i * − α i ) x i (5)</p><p>f ( x ) = ∑ i = 1 l ( α i * − α i ) ( x i ⋅ x ) + b (6)</p><p>Here, ( x i ⋅ x ) is the inner product of the vector x i and the vector x .</p></sec><sec id="s4_2_2"><title>4.2.2. Least Absolute Shrinkage and Selection Operator</title><p>Least Absolute Shrinkage and Selection Operator (LASSO), also known as linear regression L1 regularity, is a kind of compression estimation. It obtains a refined model by constructing a penalty function, making it compress some coefficients and setting some coefficients to zero. Therefore, the advantage of subset shrinkage is preserved, which is a kind of biased estimation of multiple colinearity data. The objective function is:</p><p>J ( w ) = min m { 1 2 N ‖ X T w − y ‖ 2 2 + α ‖ w ‖ 1 } (7)</p><p>Among them, y is the proportion of influenza-like cases, X is the independent variable that affects influenza cases, N is the number of data groups, α = 0.001, and w is the regression coefficient of the influenza model.</p></sec><sec id="s4_2_3"><title>4.2.3. Convolutional Neural Networks</title><p>Convolutional Neural Networks (CNN) is a deep neural network model containing convolutional layers. It has become a hot topic in the field of speech analysis and image recognition. Since CNN’s feature detection layer learns through training data, when CNN is used, explicit feature extraction is avoided, and learning is implicitly performed from training data. Furthermore, because the neuron weights on the same feature map are the same, the network can learn in parallel. Therefore, this paper selected CNN to establish influenza prediction model.</p><p>Several important levels of convolutional neural networks:</p><p>1) Convolution layer: Each neuron is seen as a filter, which calculates the local data. Take a data window, this data window slides continuously until all samples are covered.</p><p>2) Pooled layer: The pooled layer is sandwiched between successive convolution layers to compress the amount of data and parameters and reduce overfitting.</p><p>3) Excitation layer: The excitation layer has an excitation function that performs non-linear mapping of the convolutional output.</p><p>4) Fully connected layer: In the fully connected layer, all neurons between the two layers have the right to reconnect. Usually the fully connected layer is at the tail of the convolutional neural network because the amount of information at the tail does not begin to be as large.</p><p>In this paper, CNN is divided into six layers: input layer, first convolution layer, pooled layer, second convolution layer, fully connected layer, and output layer. Here, the convolutional layer excitation layer adds the excitation function ReLU to each convolution process. In addition, the droupout layer was also added to the fully connected layer, and the inactivation ratio was 0.3, which means that 70% of the neurons were retained and the overfitting phenomenon was reduced.</p><p>Enter a size of 1*16 for each training matrix. Before the first convolutional layer, change the matrix size to 4*4 and use a convolution kernel of 2*2*32. The horizontal step is 1 and the vertical step is 1, the result is 4*4*64. Enter the pooling layer to get a 2*2*32 matrix. The function used by the pooling layer is MaxPool. Then enter the next layer of convolution layer, enter 2*2*32, use the convolution kernel as 2*2*64, get 2*2*64, horizontal step is 1, vertical step is 1. Finally enter the fully connected layer, the learning efficiency is 0.01, finding the best value of the mean-square error (MSE) function by using the stochastic gradient method, the results obtained before reduce the dimension, stretched into a 512*1 matrix, and set the deactivation rate. The output to the output layer completes a training. The CNN training was completed after 500 training steps.</p></sec><sec id="s4_2_4"><title>4.2.4. Time Series Model</title><p>Taking into account the seasonal characteristics of influenza, this article considers the establishment of a time series model. The time series modeling refers to the model established by using only its past values and random disturbance terms. Its general form is:</p><p>Y t = F ( Y t − 1 , Y t − 2 , ⋯ , u t ) (8)</p><p>At present, there are two types of time-series models. One is the ARMA (Auto Regression Moving Average) model, which is an autoregressive moving average model; the other is the ARIMA (Auto Regression Integrated Moving Average) model, which is an autoregression integral moving average model. The ARMA model is suitable for stationary time series data, and the ARIMA model is suitable for non-stationary time series data.</p></sec></sec></sec><sec id="s5"><title>5. Results Analysis</title><p>A total of 47 indicators were crawled in this study from the 16<sup>th</sup> week of 2016 (started on April 25, 2016) to the 16<sup>th</sup> week of 2018 (April 23, 2018). Firstly, after PCA dimensionality reduction, there are 16 main components remaining, and the 16 main components after dimensionality reduction are included in SVR, LASSO and CNN for modeling respectively.</p><p>In this paper, a total of 105 sets of data were randomly selected from the 105 sets of data to perform tests on 10 groups. SVR, LASSO and CNN were all using the same 10 groups for testing, and the remaining 95 groups were trained.</p><p>The fitting results of the SVR, LASSO and CNN models are shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>. The training results (TR) of the three models fitted with the trends of the flu.</p><p>In the SVR model, the ploynomial kernel was used for the kernel function, C = 9.1896, gamma = 0.0474, training RMSE (Root Mean Square Error) = 0.1027, and test RMSE = 6.4906.</p><p>The LASSO model uses the penalty function L1, α = 0.001, the training RMSE = 3.9954, and the test RMSE = 2.2268.</p><p>The learning efficiency of the CNN model is 0.01. In order to prevent over-fitting, the penalty function increases the Dropout layer. Some neurons are randomly deactivated at a ratio of 0.7. The training RMSE = 1.8670 and the test RMSE = 9.6885.</p><p>Due to the seasonal features of influenza, the time series model was considered in this paper. Since the time series model requires consistency and completeness of time series data, the first 95 groups were used as training data and the last 10 groups were taken as Test Data. The unit root test results show ADF = −3.6991, p = 0.0041, indicating that the time series is a stationary time series and can be modeled with time series. The AIC rule of ARMA model is used to determine the order, and the minimum AIC value p = 3 and q = 0 are calculated. The ARMA(3,0) model is selected. The result of ARMA(3,0) fitting is shown in <xref ref-type="fig" rid="fig2">Figure 2</xref>. The training RMSE = 1.7123 and the test RMSE = 1.4333.</p><p>From the training and predictive results of the SVR, LASSO, CNN and ARMA models, it is feasible to predict the proportion of influenza-like illnesses through the Web search data. Each model shows a certain predictive result, as shown in <xref ref-type="fig" rid="fig3">Figure 3</xref> and <xref ref-type="fig" rid="fig4">Figure 4</xref>. <xref ref-type="fig" rid="fig5">Figure 5</xref> shows the accumulation absolute error of the SVR, LASSO and CNN models (SVR-AE, LASSO-AE, CNN-AE). The LASSO model has the smallest absolute error. At the same time point (2016/52, 2017/10, 2017/30, 2018/8) almost all of the three models exhibited relatively large absolute errors. Explain that the three models have poor predictability for certain periods</p><p>of influenza. The absolute error of ARMA(3,0) is smaller and the error range is (0, 2.5).</p><p>From the training RMSE of the model (in <xref ref-type="table" rid="table2">Table 2</xref>): LASSO &gt; CNN &gt; ARMA(3,0) &gt; SVR, from the perspective of the test RMSE of the model: CNN &gt; SVR &gt; LASSO &gt; ARMA(3,0). By comparison, the ARMA(3,0) model predicts better results and has greater generalization. This reflects the preference for time-series models in predicting the number of influenza cases. The LASSO model also shows a good prediction effect. SVR model performance is poor. The CNN model has the worst prediction effect, which may be due to the small amount of data, resulting in unsatisfactory learning results.</p></sec><sec id="s6"><title>6. Conclusions and Prospects</title><sec id="s6_1"><title>6.1. Conclusions</title><p>The use of web search data to predict flu epidemics is a popular research in developed countries in recent years. Baidu Index was selected as the web search data source, and the feature extraction of influenza keywords was performed through PCA algorithm. Four influenza prediction models were established and compared to explore the use of web search data to assist in the application of influenza surveillance. The conclusions are as follows:</p><p>1) It is feasible to predict the proportion of influenza-like cases by web search data.</p><p>2) Machine learning shows a certain predictive effect in the prediction of influenza based on web search data, and it has certain reference value in the future of influenza prediction.</p><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Training and prediction results based on SVR, LASSO, CNN and ARMA models</title></caption><table><tbody><thead><tr><th align="center" valign="middle"  rowspan="2"  >Date</th><th align="center" valign="middle"  rowspan="2"  >Real value (RV)</th><th align="center" valign="middle"  colspan="3"  >Predictive value (PV)</th><th align="center" valign="middle"  rowspan="2"  >Date</th><th align="center" valign="middle"  colspan="2"  >ARMA</th></tr></thead><tr><td align="center" valign="middle" >SVR-PV</td><td align="center" valign="middle" >LASSO-PV</td><td align="center" valign="middle" >CNN-PV</td><td align="center" valign="middle" >ARMA-RV</td><td align="center" valign="middle" >ARMA-PV</td></tr><tr><td align="center" valign="middle" >2016/22</td><td align="center" valign="middle" >3.3</td><td align="center" valign="middle" >2.190572</td><td align="center" valign="middle" >3.82588</td><td align="center" valign="middle" >5.128867</td><td align="center" valign="middle" >2018/7</td><td align="center" valign="middle" >33</td><td align="center" valign="middle" >35.43721</td></tr><tr><td align="center" valign="middle" >2016/32</td><td align="center" valign="middle" >2.7</td><td align="center" valign="middle" >2.759205</td><td align="center" valign="middle" >3.552288</td><td align="center" valign="middle" >4.206986</td><td align="center" valign="middle" >2018/8</td><td align="center" valign="middle" >31.4</td><td align="center" valign="middle" >30.83961</td></tr><tr><td align="center" valign="middle" >2016/42</td><td align="center" valign="middle" >8.3</td><td align="center" valign="middle" >8.047085</td><td align="center" valign="middle" >7.270786</td><td align="center" valign="middle" >11.35356</td><td align="center" valign="middle" >2018/9</td><td align="center" valign="middle" >24.9</td><td align="center" valign="middle" >25.3212</td></tr><tr><td align="center" valign="middle" >2016/52</td><td align="center" valign="middle" >23.3</td><td align="center" valign="middle" >16.69593</td><td align="center" valign="middle" >20.84384</td><td align="center" valign="middle" >41.21413</td><td align="center" valign="middle" >2018/10</td><td align="center" valign="middle" >19.8</td><td align="center" valign="middle" >20.00432</td></tr><tr><td align="center" valign="middle" >2017/10</td><td align="center" valign="middle" >16.7</td><td align="center" valign="middle" >22.42617</td><td align="center" valign="middle" >19.99406</td><td align="center" valign="middle" >6.872914</td><td align="center" valign="middle" >2018/11</td><td align="center" valign="middle" >15.6</td><td align="center" valign="middle" >14.70278</td></tr><tr><td align="center" valign="middle" >2017/20</td><td align="center" valign="middle" >6.1</td><td align="center" valign="middle" >6.652845</td><td align="center" valign="middle" >4.370603</td><td align="center" valign="middle" >3.13537</td><td align="center" valign="middle" >2018/12</td><td align="center" valign="middle" >11.8</td><td align="center" valign="middle" >9.889074</td></tr><tr><td align="center" valign="middle" >2017/30</td><td align="center" valign="middle" >18.1</td><td align="center" valign="middle" >30.91959</td><td align="center" valign="middle" >18.31482</td><td align="center" valign="middle" >26.71198</td><td align="center" valign="middle" >2018/13</td><td align="center" valign="middle" >8.1</td><td align="center" valign="middle" >6.629661</td></tr><tr><td align="center" valign="middle" >2017/40</td><td align="center" valign="middle" >6.8</td><td align="center" valign="middle" >1.038435</td><td align="center" valign="middle" >6.23335</td><td align="center" valign="middle" >4.948258</td><td align="center" valign="middle" >2018/14</td><td align="center" valign="middle" >6.6</td><td align="center" valign="middle" >5.14302</td></tr><tr><td align="center" valign="middle" >2017/50</td><td align="center" valign="middle" >40.2</td><td align="center" valign="middle" >44.69814</td><td align="center" valign="middle" >43.72965</td><td align="center" valign="middle" >44.19198</td><td align="center" valign="middle" >2018/15</td><td align="center" valign="middle" >5</td><td align="center" valign="middle" >3.50494</td></tr><tr><td align="center" valign="middle" >2018/8</td><td align="center" valign="middle" >24.9</td><td align="center" valign="middle" >13.69753</td><td align="center" valign="middle" >21.04979</td><td align="center" valign="middle" >4.802393</td><td align="center" valign="middle" >2018/16</td><td align="center" valign="middle" >4</td><td align="center" valign="middle" >2.24089</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >Training RMSE</td><td align="center" valign="middle" >0.1027</td><td align="center" valign="middle" >3.9954</td><td align="center" valign="middle" >1.8670</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" >1.7123</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >Test RMSE</td><td align="center" valign="middle" >6.4906</td><td align="center" valign="middle" >2.2268</td><td align="center" valign="middle" >9.6885</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" >1.4333</td></tr></tbody></table></table-wrap><p>3) The ARMA(3,0) model has a better predictive result and is more generalized. It also reflects that seasonal characteristics should be taken into account when predicting the proportion of influenza-like cases.</p></sec><sec id="s6_2"><title>6.2. Prospects</title><p>The outbreak and epidemic of influenza are affected by a variety of factors, including meteorological factors, virus activity intensity, and air pollution, as well as the combined effects of various factors such as the level of antibody in the population and behavioral patterns. In this study, we only studied flu prediction models by using web search data and influenza history data. Although the use of web search data for influenza surveillance has improved real-time performance, there is still a lack of accuracy, especially at the peak season of the flu season.</p><p>Future study directions for this topic include:</p><p>1) From the aspect of data sources, on the one hand, we can consider integrating the original search data of multiple search engines to reflect the search behavior of Internet users as fully as possible. In addition, we can obtain interactive behaviors through social networks, professional medical information portals, etc. and browsing behaviors to get more information on influenza concerns; on the other hand, we can collect other metrics that reflect the outbreak and epidemic of flu as a part of the predictive model input.</p><p>2) With regard to the scope of research, the scope of the study can be narrowed down to the scope of cities and counties. Based on a regional influenza prediction study, the impact of regional differences can be filtered out, and meteorological factors and other measurement indicators can be introduced more easily.</p><p>3) In the aspect of model optimization, more forecasting models can be used for weighted combinatorial optimization, and other better combinatorial optimization methods can also be used. The next optimization goal is to improve the early warning capability and achieve prediction in advance for a period of time.</p><p>4) For predictive visualization, some data visualization software can be combined to display the predictive analysis results by using charts and other methods. Displaying the real-time changes of various indicators can help users quickly obtain relevant information and respond quickly.</p></sec></sec><sec id="s7"><title>Acknowledgements</title><p>This project was supported by the Fundamental Research funds for Central Universities, China University of Geosciences (Wuhan) (1810491T09) and Laboratory Research Funds, China University of Geosciences (Wuhan) (SKJ2018240).</p></sec><sec id="s8"><title>Cite this paper</title><p>Bu, Y., Bai, J.H., Chen, Z., Guo, M.J. and Yang, F. (2018) The Study on China’s Flu Prediction Model Based on Web Search Data. Journal of Data Analysis and Information Processing, 6, 79-92. https://doi.org/10.4236/jdaip.2018.63006</p></sec></body><back><ref-list><title>References</title><ref id="scirp.85799-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Huang, L.R., Yuan, L., Li, R.C., et al. (2011) Research on Safety and Immunogenicity of Domestic Influenza Virus Split Vaccine. Fifth National Symposium on Immunodiagnosis and Vaccine, Yinchuan, 1 August 2011, 298-301. (In Chinese)</mixed-citation></ref><ref id="scirp.85799-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">WHO (2014) Seasonal Influenza. http://www.who.int/mediacentre/factsheets/fs211/en/</mixed-citation></ref><ref id="scirp.85799-ref3"><label>3</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Brady</surname><given-names> R.C. </given-names></name>,<etal>et al</etal>. (<year>2010</year>)<article-title>Influenza</article-title><source> Adolescent Medicine State of the Art Reviews</source><volume> 21</volume>,<fpage> 236</fpage>-<lpage>250</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.85799-ref4"><label>4</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Shimao</surname><given-names> T. </given-names></name>,<etal>et al</etal>. (<year>2009</year>)<article-title>Spanish Flu Related Data</article-title><source> Kekkaku</source><volume> 84</volume>,<fpage> 685</fpage>-<lpage>689</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.85799-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Centers for Disease Control and Prevention (CDC) (2010) Influenza Activity-United States and Worldwide, June 13-September 25, 2010. Morbidity and Mortality Weekly Report, 59, 1270-1273.</mixed-citation></ref><ref id="scirp.85799-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Dijk, A.V., Aramini, J., Edge, G. and Moore, K.M. (2009) Real-Time Surveillance for Respiratory Disease Outbreaks, Ontario, Canada. Emerging Infectious Diseases, 15, 799-801. https://doi.org/10.3201/eid1505.081174</mixed-citation></ref><ref id="scirp.85799-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Li, X.T., Liu, F., Dong, J.C.H., et al. (2013) Chinese Influenza Surveillance Based on Internet Search Data. Systems Engineering Theory and Practice, 33, 3028-3034. (In Chinese)</mixed-citation></ref><ref id="scirp.85799-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">FOXS (2016) Online Health Search 2006. http://www.Pewinternet.Org/2006/10/29/online-health-search-2006/</mixed-citation></ref><ref id="scirp.85799-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Polgreen, P., Chen, Y.D. and Nelson, F. (2008) Using Internet Searches for Influenza Surveillance. Clinical Infectious Diseases, 47, 1443-1448. https://doi.org/10.1086/593098</mixed-citation></ref><ref id="scirp.85799-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Ginsberg, J., Mohebbi, M.H., Patel, R.S., Brammer, L., Smolinski, M.S. and Brilliant, L. (2009) Detecting Influenza Epidemics Using Search Engine Query Data. Nature, 457, 1012. https://doi.org/10.1038/nature07634</mixed-citation></ref><ref id="scirp.85799-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Wang, R.J. (2016) Comparison and Optimization of Influenza Alerting Models Based on Internet Search Data. Doctoral Dissertation, Nankai University, Tianjing. (In Chinese)</mixed-citation></ref><ref id="scirp.85799-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Yuan, Q., Nsoesie, E.O., Lv, B., et al. (2013) Monitoring Influenza Epidemics in China with Search Query from Baidu. PLoS ONE, 8, e64323. https://doi.org/10.1371/journal.pone.0064323</mixed-citation></ref><ref id="scirp.85799-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Lu, L., Zou, Y.Q., Peng, Y.S., et al. (2016) Comparative Analysis of Baidu Index and Micro Index in Chinese Influenza Surveillance. Computer Applied Research, 33, 392-395. (In Chinese)</mixed-citation></ref><ref id="scirp.85799-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Collier, N. (2011) Omg u Got Flu? Analysis of Shared Health Messages for Bio-Surveillance. Journal of Biomedical Semantics, 2, S9. https://doi.org/10.1186/2041-1480-2-S5-S9</mixed-citation></ref><ref id="scirp.85799-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Lampos, V., Bie, T.D. and Cristianini, N. (2010) Flu Detector-Tracking Epidemics on Twitter. In: European Conference on Machine Learning and Knowledge Discovery in Databases, Vol. 6323, Springer-Verlag, Berlin, 599-602. https://doi.org/10.1007/978-3-642-15939-8_42</mixed-citation></ref><ref id="scirp.85799-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Xie, Y., Chen, Z., Cheng, Y., Zhang, K., Agrawal, A., Liao, W.K., et al. (2013) Detecting and Tracking Disease Outbreaks by Mining Social Media Data. Proceedings of the 23rd International Joint Conference on Artificial Intelligence, Beijing, 3-9 August 2013, 2958-2960.</mixed-citation></ref><ref id="scirp.85799-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Huang, J., Zhao, H. and Zhang, J. (2013) Detecting Flu Transmission by Social Sensor in China. IEEE International Conference on Green Computing and Communications and IEEE Internet of Things and IEEE Cyber, Physical and Social Computing, Beijing, 20-23 August 2013, 1242-1247. https://doi.org/10.1109/GreenCom-iThings-CPSCom.2013.216</mixed-citation></ref><ref id="scirp.85799-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Lu, L. (2015) Prediction of Chinese Flu Trends Based on Internet Data. Doctoral Dissertation, Hunan University, Changsha. (In Chinese)</mixed-citation></ref><ref id="scirp.85799-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Chinese Center for Disease Control and Prevention (2005) National Implementation Plan for Influenza/Human Bird Flu Surveillance. (In Chinese)</mixed-citation></ref></ref-list></back></article>