<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">APM</journal-id><journal-title-group><journal-title>Advances in Pure Mathematics</journal-title></journal-title-group><issn pub-type="epub">2160-0368</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/apm.2018.810049</article-id><article-id pub-id-type="publisher-id">APM-88053</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Comparison of Different Regularized and Shrinkage Regression Methods to Predict Daily Tropospheric Ozone Concentration in the Grand Casablanca Area
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Halima</surname><given-names>Oufdou</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Lise</surname><given-names>Bellanger</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Amal</surname><given-names>Bergam</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Angélina</surname><given-names>El Ghaziri</given-names></name><xref ref-type="aff" rid="aff3"><sup>3</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Kenza</surname><given-names>Khomsi</given-names></name><xref ref-type="aff" rid="aff4"><sup>4</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>El</surname><given-names>Mostafa Qannari</given-names></name><xref ref-type="aff" rid="aff5"><sup>5</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>Laboratory MAE2D, University of Abdelmalek Essaadi, Larache, Morocco</addr-line></aff><aff id="aff4"><addr-line>National Meteorological Office, Casablanca, Morocco</addr-line></aff><aff id="aff5"><addr-line>StatSC, ONIRIS, INRA, Nantes, France</addr-line></aff><aff id="aff3"><addr-line>Maison des Sciences de l’Homme Ange Guépin USR 3491, University of Nantes, Nantes, France</addr-line></aff><aff id="aff2"><addr-line>Laboratory of Mathematics Jean Leray UMR CNRS 6629, University of Nantes, Nantes, France</addr-line></aff><pub-date pub-type="epub"><day>26</day><month>10</month><year>2018</year></pub-date><volume>08</volume><issue>10</issue><fpage>793</fpage><lpage>812</lpage><history><date date-type="received"><day>21,</day>	<month>July</month>	<year>2018</year></date><date date-type="rev-recd"><day>23,</day>	<month>October</month>	<year>2018</year>	</date><date date-type="accepted"><day>26,</day>	<month>October</month>	<year>2018</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Tropospheric ozone (O3) is one of the pollutants that have a significant impact on human health. It can increase the rate of asthma crises, cause permanent lung infections and death. Predicting its concentration levels is therefore important for planning atmospheric protection strategies. The aim of this study is to predict the daily mean O3 concentration one day ahead in the Grand Casablanca area of Morocco using primary pollutants and meteorological variables. Since the available explanatory variables are multicollinear, multiple linear regressions are likely to lead to unstable models.
   
  To counteract the multicollinearity problem, we compared several alternative regression methods: 
  1
  ) Continuum Regression
  ;
   
  2
  ) Ridge &amp; Lasso Regressions
  ;
   
  3
  ) Principal component regression (PCR)
  ;
   
  4
  ) Partial least Square regression &amp; sparse PLS and
  ;
   
  5
  ) Biased Power Regression. The aim is to set up a good prediction model of the daily ozone in the Grand Casablanca area. These models are fitted on a training data set (from the years 2013 and 2014), tested on a data set (from 2015) and validated on yet another data set data (from 2015). The Lasso model showed a better performance for the prediction of ozone concentrations compared to multiple linear regression and its other alternative methods.
 
</p></abstract><kwd-group><kwd>Multiple Linear Regression</kwd><kwd> Multicollinearity</kwd><kwd> Penalized Regression</kwd><kwd> Statistical Forecasting</kwd><kwd> Tropospheric Ozone</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Tropospheric ozone (O3) is a dangerous air pollutant that threatens the human health [<xref ref-type="bibr" rid="scirp.88053-ref1">1</xref>]. Indeed, epidemiologic studies have shown that current ambient exposures are associated with reduced baseline lung function, exacerbation of asthma and premature mortality [<xref ref-type="bibr" rid="scirp.88053-ref2">2</xref>]. It is a secondary trace gas in the atmosphere, not directly emitted from any natural or anthropogenic source, but rather formed through a complex set of several chemical reactions in presence of sunlight [<xref ref-type="bibr" rid="scirp.88053-ref3">3</xref>].</p><p>As all the large cities in the world, Casablanca has a serious photochemical tropospheric ozone (O3) air pollution problem. The urban emission pattern of O3-forming pollutants is caused by meteorological factors: exposure to sunshine, temperature and wind speed and also by a series of atmospheric reactions involving precursor pollutants caused by car and industry emissions.</p><p>Various statistical methods are available to predict daily O3 [<xref ref-type="bibr" rid="scirp.88053-ref4">4</xref>] [<xref ref-type="bibr" rid="scirp.88053-ref5">5</xref>] [<xref ref-type="bibr" rid="scirp.88053-ref6">6</xref>] [<xref ref-type="bibr" rid="scirp.88053-ref7">7</xref>] Multiple linear regression (MLR) is frequently used by several environmental protection agencies involved in air quality monitoring (e.g. [<xref ref-type="bibr" rid="scirp.88053-ref8">8</xref>] [<xref ref-type="bibr" rid="scirp.88053-ref9">9</xref>] [<xref ref-type="bibr" rid="scirp.88053-ref10">10</xref>] etc.). The prediction ability of this type of models is generally satisfactory, notwithstanding the fact that, very often, the predictor variables are highly collinear. In the following, MLR will stand as the standard method to which alternative methods will be compared. In order to tackle the multicollinearity issue, various methods are proposed in the literature [<xref ref-type="bibr" rid="scirp.88053-ref11">11</xref>]. Ridge Regression [<xref ref-type="bibr" rid="scirp.88053-ref12">12</xref>] was certainly the first method proposed in this context. This method of analysis is based on a regularization strategy which aims at constraining the length (as measured by the L2 norm) of the vector of regression coefficients to be relatively small. Similarly, Lasso regression [<xref ref-type="bibr" rid="scirp.88053-ref13">13</xref>] follows the same principle as Ridge Regression, but, this time, the length of the regression coefficients is measured by L1 norm. Other alternative methods to MLR encompass Principal Component Regression [<xref ref-type="bibr" rid="scirp.88053-ref14">14</xref>] and Partial Least Squares regression [<xref ref-type="bibr" rid="scirp.88053-ref15">15</xref>] [<xref ref-type="bibr" rid="scirp.88053-ref16">16</xref>] [<xref ref-type="bibr" rid="scirp.88053-ref17">17</xref>]. These latter techniques were combined into a single approach, Continuum Regression (CR), proposed by Stone and Brooks [<xref ref-type="bibr" rid="scirp.88053-ref18">18</xref>]. Sundberg [<xref ref-type="bibr" rid="scirp.88053-ref19">19</xref>] shows that CR is also related to Ridge regression. Recently, a new biased regression strategy consisting in gradually shedding off the correlations among the independent variables was proposed by Qannari and El Ghaziri [<xref ref-type="bibr" rid="scirp.88053-ref20">20</xref>].</p><p>In this study, we compare different regression models to predict the daily mean O3 concentration in the Grand Casablanca area using O3 persistence and meteorological variables. We follow two successive stages. In the first stage, we fit statistical models using two years (2013-2014, calibration sets) of pollutants and observed meteorological data. In the second stage, in order to choose the best predictive model, we compare the prediction abilities using the observed dataset of 2015 (test set) and the meteorological forecasting dataset of 2015 (prediction dataset test). The aim of the study is to select the best model in terms of prediction ability.</p></sec><sec id="s2"><title>2. Material and Methods</title><sec id="s2_1"><title>2.1. Data</title><p>The three years datasets used in this study are provided by the National Meteorological Office of Morocco (DMN). Their collection ranged from January 2013 to December 2015. A detailed description of the 25 available variables is given in Appendix A. The data consist of daily O3 pollutant concentrations, observed at “Jahid” monitoring site, located in the western center of the Casablanca city which is the most important industrial area (<xref ref-type="fig" rid="fig1">Figure 1</xref>). Following the DMN recommendation, we use in this study 23 meteorological variables such as temperature, humidity, duration of sunshine, wind direction, wind speed, precipitation, pressure, etc. also measured at the center of Casablanca. After a step of pre-processing of these data which involved in particular the imputation of missing values, we dispose of a dataset containing for each day the observed meteorological data, forecasted meteorological data acquired from the numeric model ALADIN-Maroc and the measured O3 concentrations. The period of the study is limited to the hot and sunny season (April-September) when ozone concentrations are at their maximum [<xref ref-type="bibr" rid="scirp.88053-ref21">21</xref>].</p></sec><sec id="s2_2"><title>2.2. Exploratory Data Analysis</title><p>Inevitably, the collected data contain missing values and it is important to tackle</p><p>this problem before further analyses are performed. The reasons for which a value can be missing are numerous. For instance, in air quality applications, data can be missing due to a dysfunction of the equipment or an insufficient resolution of a sensor device. Therefore, it is necessary to identify missing values and choose an appropriate imputation technique in order to keep as much data as possible e.g. [<xref ref-type="bibr" rid="scirp.88053-ref22">22</xref>]. Various imputation procedures are used in practice. We can cite for instance regression imputation, nearest-neighbour imputation, random hot-deck imputation [<xref ref-type="bibr" rid="scirp.88053-ref23">23</xref>] [<xref ref-type="bibr" rid="scirp.88053-ref24">24</xref>] [<xref ref-type="bibr" rid="scirp.88053-ref25">25</xref>]. We choose the K-nearest neighbors (KNN) strategy because it is simple and efficient. With this technique, the imputation is based on the neighboring observations to each missing value [<xref ref-type="bibr" rid="scirp.88053-ref26">26</xref>] [<xref ref-type="bibr" rid="scirp.88053-ref27">27</xref>]. More precisely, missing values are replaced by values extracted from cases that are similar to the recipient with respect to the observed (i.e. non missing) characteristics.</p><p>Once the data were imputed, we applied a standardized Principal Components Analysis (PCA) to investigate the relationships between the variables and assess the degree of collinearity among the predictor variables [<xref ref-type="bibr" rid="scirp.88053-ref28">28</xref>] [<xref ref-type="bibr" rid="scirp.88053-ref29">29</xref>] [<xref ref-type="bibr" rid="scirp.88053-ref30">30</xref>].</p></sec><sec id="s2_3"><title>2.3. Linear Modelling Approach</title><p>Several statistical models are available to predict tropospheric ozone concentration. Since the available explanatory variables are potentially highly correlated, we investigate alternative methods to the classical Multiple Linear Regression (MLR) that circumvent the problem of multicollinearity which is likely to lead to unstable models. The emphasis is put on: 1) classical regularized regression methods: Principal Components Regression [<xref ref-type="bibr" rid="scirp.88053-ref14">14</xref>] , Partial Least Squares Regression [<xref ref-type="bibr" rid="scirp.88053-ref16">16</xref>] , Sparse PLS [<xref ref-type="bibr" rid="scirp.88053-ref31">31</xref>] [<xref ref-type="bibr" rid="scirp.88053-ref32">32</xref>] and Continuum Regression introduced by Stone and Brooks in 1990; 2) Penalized regression methods: Ridge [<xref ref-type="bibr" rid="scirp.88053-ref12">12</xref>] and Lasso developed by Tibshirani [<xref ref-type="bibr" rid="scirp.88053-ref13">13</xref>] and finally; 3) Biased Power Regression recently introduced by Qannari and El Ghaziri [<xref ref-type="bibr" rid="scirp.88053-ref20">20</xref>].</p><sec id="s2_3_1"><title>2.3.1. Multiple Linear Regression (MLR)</title><p>We assume the MLR model using Equation (1):</p><p>O 3 i = β 0 + β 1 O 3 i − 1 + ∑ j = 2 p β j v a r m e t e o i j + e i (1)</p><p>where O 3 i : ozone concentration at day i; O 3 i − 1 : ozone concentration at day i − 1 (i.e. the persistence); v a r m e t e o i j : Meteorological variable j observed on day i.</p><p>Equation (1) can be written using usual matrix format after centring of the response variable:</p><p>y = X β + e (2)</p><p>where y is an ( n &#215; 1 ) vector of centered dependant variable (O3 concentrations at day i), X is a ( n &#215; p ) matrix of standardized predictors (Observed meteorological variables and O3 concentrations at day i − 1), β is an ( p &#215; 1 ) vector of unknown regression coefficients and e is an ( n &#215; 1 ) vector of random errors. Classically, the distribution of e is assumed to be normal with mean equal to 0 and a variance covariance matrix equal to σ 2 I ; where I is the identity matrix.</p><p>The usual unbiased Ordinary Least Squares (OLS) estimator is expressed by (3) [<xref ref-type="bibr" rid="scirp.88053-ref33">33</xref>] :</p><p>β ^ OLS = ( X T X ) − 1 X T y (3)</p><p>The prediction of y using OLS, y ^ OLS , is given by y ^ OLS = X β ^ OLS . It is well known that this estimator is likely to lead to an unstable model and poor predictions in presence of quasi-collinearity among the predictors or in the case of the small sample and high dimensional setting.</p></sec><sec id="s2_3_2"><title>2.3.2. Principal Component Regression (PCR)</title><p>The principal components regression (PCR) approach involves running PCA on the predictor variables and, thereafter, using the first m principal components (PC) with 1 ≤ m ≤ p , F 1 , ⋯ , F m , as the predictors in a linear regression model [<xref ref-type="bibr" rid="scirp.88053-ref14">14</xref>] [<xref ref-type="bibr" rid="scirp.88053-ref34">34</xref>].</p><p>The appropriate number, m, of first principal components to be introduced in the model can be determined in practice by a validation technique such as Leave One Out cross validation (LOO). The PCR model can be written as in Equation (4):</p><p>y = F ( m ) β PCR ( m ) + e (4)</p><p>where F ( m ) is a matrix with m columns containing the m first PCs. The Ordinary Least Squares (OLS) estimator of β PCR is given by: β ^ PCR = ( F ( m ) T F ( m ) ) − 1 F ( m ) T y .</p><p>It is easy to express this model in terms of the original variables by remarking that F ( m ) = X U ( m ) , where U ( m ) ( U T U = I p ) is the matrix containing the m dominant normalised eigenvectors of X T X . It follows that β ^ PCR = U ( m ) T β ^ OLS .</p><p>PCR gives a biased estimate of the regression coefficients. If all the PCs are included in the model, we retrieve the usual MLR estimator, β ^ OLS .</p></sec><sec id="s2_3_3"><title>2.3.3. Partial Least Squares Regression (PLS)</title><p>As with PCR, PLS regression, introduced by [<xref ref-type="bibr" rid="scirp.88053-ref35">35</xref>] , consists in regressing y on components, also called latent variables, which are linear combinations of the p predictor variables.</p><p>The major difference between PCR and PLS regression is that, whereas PCR uses only X to construct the components to be used as regressors, PLS regression uses both X and y to determine such components. More precisely, the PLS components are determined sequentially and, at each step, we seek to determine a new component, constrained to be orthogonal to the components determined at the previous stages, so as to maximize the covariance between this component and the independent variable.</p><p>Suppose that m PLS components are determined. Again, the number m of latent variables to be introduced in the model can be selected by means of LOO cross validation technique in practice. These latent variables can be stacked into a matrix T ( m ) , T ( m ) = X W ( m ) where W ( m ) = ( w 1 , ⋯ , w m ) is the X-weight matrix. Equation (5) gives the vector of fitted values obtained by regressing y on T ( m ) , namely:</p><p>y ^ PLS m = T ( m ) ( T ( m ) T T ( m ) ) − 1 T ( m ) T y (5)</p><p>We can show in Equation (6) the expression of PLS regression coefficients as:</p><p>β ^ PLS = W ( m ) ( P ( m ) T W ( m ) ) − 1 ( T ( m ) T T ( m ) ) − 1 T ( m ) T y (6)</p><p>where P ( m ) = X T T ( m ) ( T ( m ) T T ( m ) ) − 1 and W ( m ) = X T y / ‖ X T y ‖ , where ‖   ‖ is the L<sup>2</sup> norm.</p><p>PLS regression is often helpful to reduce the number of predictors to a small number of latent variables constructed by linear combinations of the columns of original predictors. It yields a biased estimate of the regression coefficients.</p></sec><sec id="s2_3_4"><title>2.3.4. Sparse PLS Regression (SPLS)</title><p>The Sparse PLS method defined by [<xref ref-type="bibr" rid="scirp.88053-ref32">32</xref>] is a direct adaptation of the PLS regression method. It allows us to operate a dimensionality reduction using regression PLS.</p><p>In SPLS regression, w the first vector of loadings is sought as an optimal solution to:</p><p>max w ( w T M w ) subject to ‖ w T w ‖ = 1 , ‖ w ‖ 1 ≤ η ,</p><p>where M = X T y y T X , ‖ w ‖ 1 is the L<sup>1</sup>-norm of vector w and η &gt; 0 is a scalar which controls the degree of sparsity.</p><p>The regression coefficients estimation of y on X is calculated in the following way: The coefficients of the non-selected variables are set to 0, and the coefficients of the selected variables are those obtained by means of the “standard” PLS regression. We can also give an expression of the SPLS regression coefficients defined by (7) [<xref ref-type="bibr" rid="scirp.88053-ref32">32</xref>]</p><p>( β ^ SPLS ) j = { ( β ^ PLS ) j ,     if   w j ≠ 0     and     j = 1 , ⋯ , m 0                                       otherwise (7)</p><p>The interest of the SPLS is two folds. On the one hand, thanks to the sparsity, it yields an easy to interpret model and, on the other hand, it prevents the problem of multicollinearity by using the PLS framework. SPLS estimator is biased comparing to OLS estimator. Moreover, SPLS is computationally efficient with a tunable sparsity parameter to select the important variables.</p></sec><sec id="s2_3_5"><title>2.3.5. Continuum Regression (CR)</title><p>The CR prediction model is chosen from a continuum of candidates among which we find methods of analysis related to OLS estimation, PCR, PLSR. [<xref ref-type="bibr" rid="scirp.88053-ref19">19</xref>] gives a general overview of the continuum approach regression and shows how different methods relate to “least squares ridge regression”. As with PCR and PLS regression, CR consists in a regression upon latent variables (i.e. optimal linear combinations of the independent variables). More precisely, these latent variables are determined in a sequential way where at each stage, a latent variable is defined so as to realize a balance between stability (as assessed by the variance of the latent variable) and prediction ability (as assessed by the correlation between the latent variable and the dependent variable y ). The prediction of y using CR, y ^ CR , is given by y ^ CR = X β ^ CR . The CR achieves a reduction of the variance of estimator at the cost of introducing a small bias [<xref ref-type="bibr" rid="scirp.88053-ref18">18</xref>].</p><p>CR aims at transforming the explanatory variables into new latent predictors which are orthogonal to each other and constructed as linear combinations of the original predictors. It makes it possible to circumvent the problem of multicollinearity between predictors. However, the CR regression does not specifically aim at selecting a subset of variables [<xref ref-type="bibr" rid="scirp.88053-ref18">18</xref>].</p></sec><sec id="s2_3_6"><title>2.3.6. Penalized Regression</title><p>Another general strategy to circumvent the problem of multicollinearity consists in imposing a constraint on the vector of regression coefficients. The two most popular methods in this context are Ridge and Lasso regressions.</p><p>&#216; Ridge regression</p><p>Ridge regression is the first regularization procedure that was proposed to cope with the multicollinearity problem [<xref ref-type="bibr" rid="scirp.88053-ref12">12</xref>]. The Ridge estimator is given by (8):</p><p>β ^ R = ( X T X + k I ) − 1 X T y (8)</p><p>where k ≥ 0 is a constant to be selected. Note that if k = 0 , the Ridge estimator amounts to the least-squares estimator.</p><p>Ridge estimator is obtained as a solution to the following least squares problem defined by (9):</p><p>β ^ R = arg min β ∈ ℝ p , ‖ β ‖ 2 ≤ δ ( ‖ y − X β ‖ 2 )     where   δ ≥ 0 (9)</p><p>There is a one to one correspondence between the Ridge parameter k and the upper bound, δ, imposed on the vector of regression coefficients, β . From a practical point of view, these parameters can be selected by means of a cross-validation technique.</p><p>The Ridge regression shrinks the OLS estimators towards 0. It yields a biased estimator, but with a smaller variance than that of OLS estimator.</p><p>&#216; Lasso regression</p><p>The Least Absolute Shrinkage and Selection Operator, or Lasso [<xref ref-type="bibr" rid="scirp.88053-ref13">13</xref>] is another penalized regression where L<sup>2</sup> penalty of ridge regression is replaced by an L<sup>1</sup> penalty: ‖ β ‖ 1 = ∑ j = 1 p | β j | . This is a subtle change that has important consequences. Indeed, this constraint entails that some of the regression coefficients are shrunk exactly to zero. This means that this regression strategy operates de facto a selection of variables since the unimportant variables are discarded, their regression coefficients being equal to zero. Formally, the lasso estimator is given as a solution to the following optimization problem by (10):</p><p>β ^ Lasso = arg min β ∈ ℝ p , ‖ β ‖ 1 ≤ δ ( ‖ y − X β ‖ 2 ) , (10)</p><p>where δ ≥ 0 .</p><p>The parameter δ controls the degree of sparsity and, in practice; it is determined by Leave One Out (LOO) cross validation procedure. The smaller this parameter is, the larger is the number of discarded variable. Contrariwise, if δ is larger than δ 0 = ∑ j = 1 p | β ^ j | (where β ^ j are the OLS estimators) then β ^ Lasso = β ^ OLS . Lasso regression has the double effect of shrinking the β coefficients, allowing to decrease the variance of the regression coefficients as with Ridge regression, and, more importantly, to operate an automatic selection of the variables, by cancelling out some β j coefficients.</p></sec><sec id="s2_3_7"><title>2.3.7. Biased Power Regression (BPR)</title><p>Recently, a new biased regression called Biased Power regression (BPR) strategy was proposed [<xref ref-type="bibr" rid="scirp.88053-ref20">20</xref>]. It consists in gradually shedding off the correlation among the independent variables by means of a tuning parameter α. More precisely, the BPR estimator of β is given by (11):</p><p>β ^ BP = ( X T X ) α − 1 X T y (11)</p><p>where α is a tuning parameter which ranges between 0 and 1.</p><p>In practice, α is selected using a cross validation procedure.</p><p>Clearly, when α = 0 , we retrieve the OLS estimator and as α increases, the variance-covariance matrix of the predictor variables is shrunk to the identity matrix. The prediction of y using BPR, y ^ BP is given by y ^ BP = X β ^ BP .</p><p>BP-regression shares the same properties as Ridge regression (see Section 2.3.4) and thus can highlight those variables whose coefficients become very small. However, it was not designed to select a subset of variables [<xref ref-type="bibr" rid="scirp.88053-ref20">20</xref>].</p></sec></sec><sec id="s2_4"><title>2.4. Evaluation of the Methods</title><p>To assess the prediction ability of the various models listed above on the Grand Casablanca O3 data, we performed a cross validation technique on a training set to determine the appropriate parameters (number of components, Ridge or lasso constant…) to be used in the prediction models. Using these parameters, the performance of the different models is assessed on the basis of a fresh data set. More precisely, we partitioned the available data into two complementary datasets: 1) summer period of 2013 and 2014 (called the training set) used to adjust the models; and 2) summer period of 2015 (called the validation set or testing set) used to “test” the models obtained in the training phase. The models are fitted on the training set used to predict the ozone responses for a) the observed meteorological data on 2015 of the validation set (obstest) and b) the forecasted meteorological data on 2015 for real validation (prevtest).</p><p>The performance of the models is measured with standard indicators defined by Equations (12)-(14) generally used to compare statistical models [<xref ref-type="bibr" rid="scirp.88053-ref36">36</xref>].</p><p>In a first stage, an internal validation (2013 and 2014 datasets) is performed on the basis of the following criteria in order to assess the quality of the model adjustment:</p><p>The multiple correlation coefficient R<sup>2</sup> allows us to assess the quality of the</p><p>adjustment based on the training set: R 2 = ∑ i = 1 n train ( y i − y ^ i ) 2 ∑ i = 1 n train ( y i − y &#175; ) 2 , where n train is the</p><p>size of training sample.</p><p>The Root Mean Squared Errors (RMSE): This is computed according to the following expression:</p><p>RMSE = 1 n train ∑ i = 1 n train ( y i − y ^ i ) 2 (12)</p><p>The smallest value of this criterion corresponds to the best adjustment of the model.</p><p>For the external validation (on summer 2015 observed dataset), the following criterion is used to assess the prediction ability of the models [<xref ref-type="bibr" rid="scirp.88053-ref37">37</xref>] :</p><p>The Root Mean Squared Errors of Prediction (RMSEP). This criterion is similar to RMSE but, this time, the validation data set is used instead of the training data set.</p><p>RMSEP obs = 1 n obstest ∑ i = 1 n obstest ( y i − y ^ i ) 2 (13)</p><p>where n obstest is the size of the observed the validation set (obstest).</p><p>Obviously, the best predictive model corresponds to the smallest RMSEP.</p><p>The following criterion is used to assess the performance of the models with observed meteorological data (obstest) and real meteorological forecast data (prevtest) for summer period of 2015.</p><p>In the same way, we define the RMSEP of prevision based on the forecasted dataset as:</p><p>RMSEP prev = 1 n prevtest ∑ i = 1 n prevtest ( y i − y ^ i ) 2 (14)</p><p>where n prevtest is the size of the sample size of the forecasted data (prevtest).</p></sec></sec><sec id="s3"><title>3. Results and Discussion</title><p>Experiments were run on an Intel(R) Core(TM) i7-6600U CPU computer with 2.60 GHz, 8 Go in RAM, Windows 10 Professional 64 bits.</p><p>All the statistical analyses were performed using the free software R. (http://www.rproject.org/).</p><sec id="s3_1"><title>3.1. Data Description</title><p>In this study, the dataset is composed of 25 explanatory variables. Appendix A gives the abbreviation of these variables.</p><p><xref ref-type="table" rid="table1"><xref ref-type="table" rid="table">Table </xref>1</xref> provides descriptive statistics of the meteorological variables and tropospheric ozone concentrations calculated on the data with missing values.</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1"><xref ref-type="table" rid="table">Table </xref>1</xref></label><caption><title> Statistics of measured variables at Grand Casablanca area from 01 April 2013 to 30 September 2014</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Variable</th><th align="center" valign="middle" >Min</th><th align="center" valign="middle" >Max</th><th align="center" valign="middle" >Mean</th><th align="center" valign="middle" >St.Dev</th><th align="center" valign="middle" >NA</th></tr></thead><tr><td align="center" valign="middle" >TMPMAX</td><td align="center" valign="middle" >16.2</td><td align="center" valign="middle" >37.5</td><td align="center" valign="middle" >24.5</td><td align="center" valign="middle" >3.09</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >TMPMIN</td><td align="center" valign="middle" >8.20</td><td align="center" valign="middle" >23.50</td><td align="center" valign="middle" >18.35</td><td align="center" valign="middle" >3.02</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >TMPMOY</td><td align="center" valign="middle" >12.40</td><td align="center" valign="middle" >29.90</td><td align="center" valign="middle" >21.45</td><td align="center" valign="middle" >2.88</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >RRQUOT</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >19.30</td><td align="center" valign="middle" >0.39</td><td align="center" valign="middle" >1.98</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >DRINSQ</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >13.30</td><td align="center" valign="middle" >9.72</td><td align="center" valign="middle" >2.79</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >HUMREL06h</td><td align="center" valign="middle" >50.00</td><td align="center" valign="middle" >100.0</td><td align="center" valign="middle" >87.42</td><td align="center" valign="middle" >8.00</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >HUMREL12h</td><td align="center" valign="middle" >34.00</td><td align="center" valign="middle" >95.00</td><td align="center" valign="middle" >68.32</td><td align="center" valign="middle" >8.78</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >HUMREL18h</td><td align="center" valign="middle" >28.00</td><td align="center" valign="middle" >97.00</td><td align="center" valign="middle" >75.66</td><td align="center" valign="middle" >9.66</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >PRESTN06h</td><td align="center" valign="middle" >9997.7</td><td align="center" valign="middle" >1017.3</td><td align="center" valign="middle" >1008.2</td><td align="center" valign="middle" >2.97</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >PRESTN12h</td><td align="center" valign="middle" >997.7</td><td align="center" valign="middle" >1016.5</td><td align="center" valign="middle" >1008.9</td><td align="center" valign="middle" >2.91</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >PRESTN18h</td><td align="center" valign="middle" >999</td><td align="center" valign="middle" >1016</td><td align="center" valign="middle" >1008</td><td align="center" valign="middle" >2.88</td><td align="center" valign="middle" >0</td></tr><tr><td align="center" valign="middle" >FFVM06h</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >4.00</td><td align="center" valign="middle" >1.55</td><td align="center" valign="middle" >0.80</td><td align="center" valign="middle" >3</td></tr><tr><td align="center" valign="middle" >FFVM12h</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >6.00</td><td align="center" valign="middle" >3.58</td><td align="center" valign="middle" >0.98</td><td align="center" valign="middle" >4</td></tr><tr><td align="center" valign="middle" >FFVM18h</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >7.00</td><td align="center" valign="middle" >3.46</td><td align="center" valign="middle" >1.04</td><td align="center" valign="middle" >4</td></tr><tr><td align="center" valign="middle" >DDVM06degre</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >360.0</td><td align="center" valign="middle" >176.4</td><td align="center" valign="middle" >117.87</td><td align="center" valign="middle" >3</td></tr><tr><td align="center" valign="middle" >DDVM12hDEG</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >360.0</td><td align="center" valign="middle" >227.3</td><td align="center" valign="middle" >141.63</td><td align="center" valign="middle" >4</td></tr><tr><td align="center" valign="middle" >DDVM18hDEG</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >360.0</td><td align="center" valign="middle" >189.2</td><td align="center" valign="middle" >152.21</td><td align="center" valign="middle" >4</td></tr><tr><td align="center" valign="middle" >Vx06</td><td align="center" valign="middle" >−2.95</td><td align="center" valign="middle" >3.46</td><td align="center" valign="middle" >−0.05</td><td align="center" valign="middle" >1.06</td><td align="center" valign="middle" >3</td></tr><tr><td align="center" valign="middle" >Vx12</td><td align="center" valign="middle" >−5.91</td><td align="center" valign="middle" >3.94</td><td align="center" valign="middle" >−0.59</td><td align="center" valign="middle" >1.98</td><td align="center" valign="middle" >4</td></tr><tr><td align="center" valign="middle" >Vx18</td><td align="center" valign="middle" >−5.91</td><td align="center" valign="middle" >4.50</td><td align="center" valign="middle" >−0.10</td><td align="center" valign="middle" >1.84</td><td align="center" valign="middle" >4</td></tr><tr><td align="center" valign="middle" >Vy06</td><td align="center" valign="middle" >−4.00</td><td align="center" valign="middle" >4.00</td><td align="center" valign="middle" >0.08</td><td align="center" valign="middle" >1.38</td><td align="center" valign="middle" >3</td></tr><tr><td align="center" valign="middle" >Vy12</td><td align="center" valign="middle" >−3.06</td><td align="center" valign="middle" >6.00</td><td align="center" valign="middle" >2.75</td><td align="center" valign="middle" >1.39</td><td align="center" valign="middle" >4</td></tr><tr><td align="center" valign="middle" >Vy18</td><td align="center" valign="middle" >−5.36</td><td align="center" valign="middle" >6.00</td><td align="center" valign="middle" >2.79</td><td align="center" valign="middle" >1.36</td><td align="center" valign="middle" >4</td></tr><tr><td align="center" valign="middle" >O3veilleJahid</td><td align="center" valign="middle" >10.00</td><td align="center" valign="middle" >130.0</td><td align="center" valign="middle" >52.83</td><td align="center" valign="middle" >25.66</td><td align="center" valign="middle" >23</td></tr><tr><td align="center" valign="middle" >O3Jahid</td><td align="center" valign="middle" >10.00</td><td align="center" valign="middle" >130.0</td><td align="center" valign="middle" >52.84</td><td align="center" valign="middle" >25.62</td><td align="center" valign="middle" >23</td></tr></tbody></table></table-wrap><p>Minimum, maximum, mean and standard deviation statistics are provided to describe the characteristics of the data set.</p><p>The 2013 and 2014 studied periods are characterized by high temperatures. We can notice that, in the Grand Casablanca Area, the maximal temperature (TMPMAX) can go up to 37.5˚C and the minimal temperature is 16.2˚C. The maximal daily total sunshine duration is of 13.3 hours. We can also notice that there is almost no rain is these periods (RRQUOT). The Wind strength average is relatively high at 18 hours (FFVM18h = 7 m/s). The O3 concentrations are between 10 and 130 &#181;g/m<sup>3</sup>.</p><p>There are in total 90 missing values for the 366 recording days, distributed on 14 variables. This represents around 2% of missing values to be imputed before the prediction models are performed.</p><sec id="s3_1_1"><title>3.1.1. Missing Values Imputation</title><p>As mentioned above, a strategy based on the K-nearest neighbors was performed. Different K values were used in the literature and the choice of K = 10 led to the best results [<xref ref-type="bibr" rid="scirp.88053-ref38">38</xref>].</p></sec><sec id="s3_1_2"><title>3.1.2. Multivariate Analysis</title><p><xref ref-type="fig" rid="fig2">Figure 2</xref> shows the scatter plots associated with pairs of the available variables. It highlights the pairwise correlations between these variables. We also indicate in <xref ref-type="fig" rid="fig2">Figure 2</xref>, the histograms associated with each variable and the correlation coefficients between each pair of variables. For instance, we can see that there is a high correlation between O3veilleJahid and O3 Jahid (the two last columns) with a correlation coefficient equal to 0.92. We can also observe large correlations (around 0.94) between the first three explanatory variables (TMPMAX and TMPMOY, TMPMIN and TMPMOY).</p><p>The diagonal entries show the histograms associated with the various variables and the upper entries indicate the coefficients of correlations between pairs of variables.</p><p>We also performed a PCA on the imputed dataset. PCA is run on the complete data (2013 and 2014) after imputation and standardization of the variables. The data is composed of 366 days (from April to September of 2013 and 2014) and 24 variables. The first five principal components recover up to 65% of the total variance (<xref ref-type="table" rid="table2"><xref ref-type="table" rid="table">Table </xref>2</xref>). In the following, only the results related to the first two principal components which recover around 40% of the total inertia are shown.</p><p><xref ref-type="fig" rid="fig3">Figure 3</xref> shows the correlations of the explanatory variables with the first two principal components. The variable O3 Jahid is superimposed as an illustrative variable with a blue arrow to depict its relationships with the explanatory variables. This figure highlights the strong correlations among the variables, which may be harmful for the prediction models.</p><p>The first PC is linked to wind direction (Vx06, Vx12, Vy12 and Vx18) and pressure variables (PRESTN at 06 h, 12 h and 18 h). We can notice, for example, that variables TMPMAX, TMPMIN and TMPMOY as well as PRESTN06h, PRESTN12h and PRESTN18h are strongly correlated. A strong correlation also exists between variables Vx06, Vy06, Vx18 and Vx12. O3veilleJahid variable is very correlated to O3 Jahid but it is not very well represented in the plan (PC1-PC2).</p></sec></sec><sec id="s3_2"><title>3.2. Prediction Models</title><p>In this section, we compare the results obtained from the different regression models described in section 2.3, namely: 1) the Multiple Linear Regression (MLR) model applied to all the variables of the dataset (24 variables), 2) the Reduced MLR with seven variables selected by means of Akaike Information Criterion (AIC) [<xref ref-type="bibr" rid="scirp.88053-ref39">39</xref>] [<xref ref-type="bibr" rid="scirp.88053-ref40">40</xref>] , 3) The Principal Component Regression (PCR) model, 4) PLS and sparse PLS models, 5) Continuum Regression (CR), 6) Ridge and Lasso</p><table-wrap id="table2" ><label><xref ref-type="table" rid="table2"><xref ref-type="table" rid="table">Table </xref>2</xref></label><caption><title> Percentage of total variance recovered by the principal components</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Component</th><th align="center" valign="middle" >Eigenvalue</th><th align="center" valign="middle" >Percentage of variance</th><th align="center" valign="middle" >Cumulative percentage of variance</th></tr></thead><tr><td align="center" valign="middle" >comp 1</td><td align="center" valign="middle" >4.83</td><td align="center" valign="middle" >23.02</td><td align="center" valign="middle" >23.02</td></tr><tr><td align="center" valign="middle" >comp 2</td><td align="center" valign="middle" >3.67</td><td align="center" valign="middle" >17.47</td><td align="center" valign="middle" >40.49</td></tr><tr><td align="center" valign="middle" >comp 3</td><td align="center" valign="middle" >1.91</td><td align="center" valign="middle" >9.11</td><td align="center" valign="middle" >49.60</td></tr><tr><td align="center" valign="middle" >comp 4</td><td align="center" valign="middle" >1.77</td><td align="center" valign="middle" >8.43</td><td align="center" valign="middle" >58.04</td></tr><tr><td align="center" valign="middle" >comp 5</td><td align="center" valign="middle" >1.52</td><td align="center" valign="middle" >7.27</td><td align="center" valign="middle" >65.31</td></tr><tr><td align="center" valign="middle" >comp 6</td><td align="center" valign="middle" >1.30</td><td align="center" valign="middle" >6.19</td><td align="center" valign="middle" >71.51</td></tr><tr><td align="center" valign="middle" >comp 7</td><td align="center" valign="middle" >1.01</td><td align="center" valign="middle" >4.82</td><td align="center" valign="middle" >76.33</td></tr><tr><td align="center" valign="middle" >comp 8</td><td align="center" valign="middle" >0.99</td><td align="center" valign="middle" >4.73</td><td align="center" valign="middle" >81.06</td></tr><tr><td align="center" valign="middle" >comp 9</td><td align="center" valign="middle" >0.83</td><td align="center" valign="middle" >3.93</td><td align="center" valign="middle" >84.99</td></tr><tr><td align="center" valign="middle" >comp 10</td><td align="center" valign="middle" >0.64</td><td align="center" valign="middle" >3.05</td><td align="center" valign="middle" >88.05</td></tr></tbody></table></table-wrap><p>regressions, 7) Biased Power regression (BP-regression).</p><p>A cross validation procedure (LOO) is applied on the data collected during the period extending from April 1<sup>st</sup> to September 30<sup>th</sup> in 2013 and 2014 (training data) to determine for each model the parameters (number of components, Ridge and Lasso parameters…) leading to the minimum of the Root Mean Squared Error (RMSE). Then, the prediction ability according to the Root Mean Squared Error Predicted (RMSEP) of the various models is assessed on the basis of: 1) observed data (test data), RMSEP<sub>obs</sub>, and 2) forecasted data, RMSEP<sub>prev</sub> from the summer period of 2015.</p><p><xref ref-type="table" rid="table3"><xref ref-type="table" rid="table">Table </xref>3</xref> shows the results of the various methods according to several criteria (RMSE, R<sup>2</sup>, RMSEPobs and RMSEPprev).</p><p>Concerning the adjustment of the model on training data (internal validation), not surprisingly, the MLR model leads to the lowest RMSE (9.503), but the other models lead to close values and take into account multicollinearity problem. However, if the goal is to get the best predictive model, the RMSE alone is unsufficient and we need to analyze the RMSEP to assess the predictive quality of each model.</p><p>As for the criterion RMSEP<sub>obs</sub> in external validation, Lasso, Ridge, PLS, BP</p><table-wrap id="table3" ><label><xref ref-type="table" rid="table3"><xref ref-type="table" rid="table">Table </xref>3</xref></label><caption><title> Comparison of different models according RMSE, R<sup>2</sup> and RMSEP criteria</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Model</th><th align="center" valign="middle" >MLR</th><th align="center" valign="middle" >Reduced MLR</th><th align="center" valign="middle" >PCR</th><th align="center" valign="middle" >PLS</th><th align="center" valign="middle" >Sparse PLS</th><th align="center" valign="middle" >CR</th><th align="center" valign="middle" >Ridge</th><th align="center" valign="middle" >Lasso</th><th align="center" valign="middle" >BP Reg</th></tr></thead><tr><td align="center" valign="middle" >Nb var</td><td align="center" valign="middle" >24</td><td align="center" valign="middle" >7</td><td align="center" valign="middle" >24</td><td align="center" valign="middle" >24</td><td align="center" valign="middle" >14</td><td align="center" valign="middle" >24</td><td align="center" valign="middle" >24</td><td align="center" valign="middle" >11</td><td align="center" valign="middle" >24</td></tr><tr><td align="center" valign="middle" >Parameter</td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" >ncp = 10</td><td align="center" valign="middle" >ncp = 5</td><td align="center" valign="middle" >ncp = 3 η = 0.56</td><td align="center" valign="middle" >ncp = 1 α = 0.1</td><td align="center" valign="middle" >λopt = 9.322</td><td align="center" valign="middle" >Fract = 0.2</td><td align="center" valign="middle" >α = 0.01</td></tr><tr><td align="center" valign="middle" >RMSE</td><td align="center" valign="middle" >9.503</td><td align="center" valign="middle" >9.587</td><td align="center" valign="middle" >10.521</td><td align="center" valign="middle" >9.59</td><td align="center" valign="middle" >9.703</td><td align="center" valign="middle" >10.11</td><td align="center" valign="middle" >9.537</td><td align="center" valign="middle" >9.676</td><td align="center" valign="middle" >9.535</td></tr><tr><td align="center" valign="middle" >R<sup>2</sup></td><td align="center" valign="middle" >0.862</td><td align="center" valign="middle" >0.859</td><td align="center" valign="middle" >0.831</td><td align="center" valign="middle" >0.859</td><td align="center" valign="middle" >0.858</td><td align="center" valign="middle" >0.872</td><td align="center" valign="middle" >0.818</td><td align="center" valign="middle" >0.829</td><td align="center" valign="middle" >0.834</td></tr><tr><td align="center" valign="middle" >RMSEPobs</td><td align="center" valign="middle" >11.84</td><td align="center" valign="middle" >11.68</td><td align="center" valign="middle" >13.45</td><td align="center" valign="middle" >11.74</td><td align="center" valign="middle" >12.24</td><td align="center" valign="middle" >11.83</td><td align="center" valign="middle" >11.73</td><td align="center" valign="middle" >11.58</td><td align="center" valign="middle" >11.74</td></tr><tr><td align="center" valign="middle" >RMSEPprev</td><td align="center" valign="middle" >15.40</td><td align="center" valign="middle" >14.49</td><td align="center" valign="middle" >15.80</td><td align="center" valign="middle" >13.49</td><td align="center" valign="middle" >14.92</td><td align="center" valign="middle" >15.21</td><td align="center" valign="middle" >14.98</td><td align="center" valign="middle" >12.74</td><td align="center" valign="middle" >14.39</td></tr></tbody></table></table-wrap><p>regressions and CR outperform the other methods. Lasso shows the best predictive ability since it has the smallest RMSEPobs. Moreover, this method of analysis has yet another advantage since the model is based on fewer predictive variables (11 variables such as TMPMAX, TMPMIN, DRINSQ, HUMREL12h, PRESTN06h, FFVM06h, FFVM18h, DDVM12h, Vx06h, Vy06h and O3veilleJahid) than the other models with the exception of the reduced model (7 predictive variables such as TMPMIN, TMPMOY, DRINSQ, PRESTN06, Vx06, Vx12, O3veilleJahid). Among these significant variables, TMPMIN and TMPMOY are strongly correlated so the reduced model does not solve the multicollinearity problem by comparing it to the Lasso model.</p><p>Most important are the results of the RMSEP<sub>prev</sub> based on the forecasted meteorological data for 2015. We recall that the forecasted meteorological data will be the data used on daily basis to predict the O3 concentration as obtained by Aladin-Maroc numerical forecasted model. It turns out that Lasso has by far the best RMSEP_prev (equal to 12.74), a value close to its RMSEP_obs (11.58) followed by the PLS and BP regression model. However, these last two models keep all the predictive variables, unlike the Lasso model, which keeps fewer variables, thus obtaining a model that is simple and easy to interpret.</p><p><xref ref-type="fig" rid="fig4">Figure 4</xref> shows a good correlation (around 0.723) between observed O3 and forecasted O3 data one day ahead obtained with the Lasso regression model only in 2015.</p><p>The most important finding (<xref ref-type="table" rid="table3"><xref ref-type="table" rid="table">Table </xref>3</xref>, <xref ref-type="fig" rid="fig4">Figure 4</xref>) is that the Lasso regression model has the best performance in predicting O3 concentrations in Jahid compared to the other models. Moreover, it clearly gives stable regression coefficients compared to Reduced MLR model. <xref ref-type="table" rid="table">Table </xref>A1 of explanatory variables and <xref ref-type="table" rid="table">Table </xref>B1 of the coefficients estimated by the models used in this study shows that the explanatory variables most retained by the models are: TMPMAX, TMPMIN, DRINSQ, PRESTNO6h, Vx06 and O3veilleJahid. Indeed, the formation of ozone in the Grand Casablanca area is related more particularly to: 1) the maximum and minimum daily temperature; 2) the period of intense sunshine; 3) the weak wind that accumulates the massive concentration of ozone and; 4) the previous day’s concentration, which, to a large extent, determines the next day’s</p><p>ozone concentration.</p></sec></sec><sec id="s4"><title>4. Conclusions</title><p>Starting with a multiple linear regression model, which is plagued by multicollinearity among the predictor variables, we have considered nine more or less recent alternative methods to relate meteorological and pollution variables. The emphasis was put on the prediction ability of the daily tropospheric ozone of these models in the Grand Casablanca area as the first comparative study of its type in such region.</p><p>We proposed the selected Lasso model based on a comparison of several linear forecasting methods to reduce the multicollinearity problem. The results obtained over two years of training data (2013 and 2014), verified on observed data (2015) and validated on forecast data (2015) show that the Lasso model has the best predictive capacity O3 for the Jahid station located in Grand Casablanca area. Moreover, using the dataset of 2015, Lasso model still gives the best predictive ability for O3 in Jahid station. The Lasso model presents the interest of being relatively simple and easily interpretable. The choice of this model is explained by the fact that it yields the best criteria in comparison to the alternative models discussed in this paper. These criteria include R&#178;, RMSE, RMSEPobs and RMSEPprev. Furthermore, besides yielding a more stable model than multiple linear regression, Lasso is based on a relatively small number of explanatory variables. This feature presents a significant advantage for the daily prediction of the ozone concentration in the Grand Casablanca.</p><p>This contribution proposes the first linear model of daily O3 concentration forecast in Morocco and more particularly in the Grand Casablanca area.</p><p>In perspective, we plan to widen our study by comparing the performances of the Lasso model with those of other non-parametric models and we will add more data (2017-2018) to ensure model validation. The most appropriate forecast model will be routinely implemented by the National Meteorological Office of Morocco (DMN).</p></sec><sec id="s5"><title>Acknowledgements</title><p>The National Meteorological Office of Morocco (Direction de la M&#233;t&#233;orologie Nationale DMN) is gratefully acknowledged for providing the necessary data for undertaking the present study.</p></sec><sec id="s6"><title>Conflicts of Interest</title><p>The authors declare no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s7"><title>Cite this paper</title><p>Oufdou, H., Bellanger, L., Bergam, A., El Ghaziri, A., Khomsi, K. and Qannari, E.M. (2018) Comparison of Different Regularized and Shrinkage Regression Methods to Predict Daily Tropospheric Ozone Concentration in the Grand Casablanca Area. Advances in Pure Mathematics, 8, 793-812. https://doi.org/10.4236/apm.2018.810049</p></sec><sec id="s8"><title>Appendix A</title><table-wrap id="table4" ><label><xref ref-type="table" rid="table">Table </xref>A1</label><caption><title> Variables abbreviation and units of measurement</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Abbreviation</th><th align="center" valign="middle" >Variable</th><th align="center" valign="middle" >Unit</th></tr></thead><tr><td align="center" valign="middle" >TMPMAX</td><td align="center" valign="middle" >Maximal temperature</td><td align="center" valign="middle" >˚C</td></tr><tr><td align="center" valign="middle" >TMPMIN</td><td align="center" valign="middle" >Minimal temperature</td><td align="center" valign="middle" >˚C</td></tr><tr><td align="center" valign="middle" >TMPMOY</td><td align="center" valign="middle" >Average temperature</td><td align="center" valign="middle" >˚C</td></tr><tr><td align="center" valign="middle" >RRQUOT</td><td align="center" valign="middle" >Total precipitation</td><td align="center" valign="middle" >mm</td></tr><tr><td align="center" valign="middle" >DRINSQ</td><td align="center" valign="middle" >Sunshine duration</td><td align="center" valign="middle" >heure</td></tr><tr><td align="center" valign="middle" >HUMREL06h</td><td align="center" valign="middle" >Relative humidity at 06 h</td><td align="center" valign="middle" >%</td></tr><tr><td align="center" valign="middle" >HUMREL12h</td><td align="center" valign="middle" >Relative humidity at 12 h</td><td align="center" valign="middle" >%</td></tr><tr><td align="center" valign="middle" >HUMREL18h</td><td align="center" valign="middle" >Relative humidity at 18 h</td><td align="center" valign="middle" >%</td></tr><tr><td align="center" valign="middle" >PRESTN06h</td><td align="center" valign="middle" >Pressure at the station level at 06 h</td><td align="center" valign="middle" >hpa</td></tr><tr><td align="center" valign="middle" >PRESTN12h</td><td align="center" valign="middle" >Pressure at the station level at 12 h</td><td align="center" valign="middle" >hpa</td></tr><tr><td align="center" valign="middle" >PRESTN18h</td><td align="center" valign="middle" >Pressure at the station level at 18 h</td><td align="center" valign="middle" >hpa</td></tr><tr><td align="center" valign="middle" >FFVM06h</td><td align="center" valign="middle" >Wind force at 06 h</td><td align="center" valign="middle" >m/s</td></tr><tr><td align="center" valign="middle" >FFVM12h</td><td align="center" valign="middle" >Wind force at 12 h</td><td align="center" valign="middle" >m/s</td></tr><tr><td align="center" valign="middle" >FFVM18h</td><td align="center" valign="middle" >Wind force at 18 h</td><td align="center" valign="middle" >m/s</td></tr><tr><td align="center" valign="middle" >DDVM06h</td><td align="center" valign="middle" >Wind direction at 06 h</td><td align="center" valign="middle" >degree</td></tr><tr><td align="center" valign="middle" >DDVM12h</td><td align="center" valign="middle" >Wind direction at 12 h</td><td align="center" valign="middle" >degree</td></tr><tr><td align="center" valign="middle" >DDVM18h</td><td align="center" valign="middle" >Wind direction at 18 h</td><td align="center" valign="middle" >degree</td></tr><tr><td align="center" valign="middle" >Vx06</td><td align="center" valign="middle" >Horizontal wind at 06 h</td><td align="center" valign="middle" >m/s</td></tr><tr><td align="center" valign="middle" >Vx12</td><td align="center" valign="middle" >Horizontal wind at 12 h</td><td align="center" valign="middle" >m/s</td></tr><tr><td align="center" valign="middle" >Vx18</td><td align="center" valign="middle" >Horizontal wind at 18 h</td><td align="center" valign="middle" >m/s</td></tr><tr><td align="center" valign="middle" >Vy06</td><td align="center" valign="middle" >Vertical wind at 06 h</td><td align="center" valign="middle" >m/s</td></tr><tr><td align="center" valign="middle" >Vy12</td><td align="center" valign="middle" >Vertical wind at 12 h</td><td align="center" valign="middle" >m/s</td></tr><tr><td align="center" valign="middle" >Vy18</td><td align="center" valign="middle" >Vertical wind at 18 h</td><td align="center" valign="middle" >m/s</td></tr><tr><td align="center" valign="middle" >O3veilleJahid</td><td align="center" valign="middle" >Ozone concentrations of the day before</td><td align="center" valign="middle" >&#181;g/m<sup>3</sup></td></tr><tr><td align="center" valign="middle" >O3veille</td><td align="center" valign="middle" >Ozone concentrations</td><td align="center" valign="middle" >&#181;g/m<sup>3</sup></td></tr></tbody></table></table-wrap></sec><sec id="s9"><title>Appendix B</title><table-wrap id="table5" ><label><xref ref-type="table" rid="table">Table </xref>B1</label><caption><title> Comparison of regression coefficients estimated by the different models</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Variables</th><th align="center" valign="middle" >Complete Reg</th><th align="center" valign="middle" >Reduced Reg</th><th align="center" valign="middle" >PCR</th><th align="center" valign="middle" >PLS</th><th align="center" valign="middle" >SPLS</th><th align="center" valign="middle" >CR</th><th align="center" valign="middle" >Ridge</th><th align="center" valign="middle" >Lasso</th><th align="center" valign="middle" >BP Reg</th></tr></thead><tr><td align="center" valign="middle" >TMPMAX</td><td align="center" valign="middle" >24.55</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.03</td><td align="center" valign="middle" >−1.06</td><td align="center" valign="middle" >−1.41</td><td align="center" valign="middle" >−1.74</td><td align="center" valign="middle" >−2.52</td><td align="center" valign="middle" >−0.55</td><td align="center" valign="middle" >21.62</td></tr><tr><td align="center" valign="middle" >TMPMIN</td><td align="center" valign="middle" >30.45</td><td align="center" valign="middle" >6.99</td><td align="center" valign="middle" >1.13</td><td align="center" valign="middle" >1.09</td><td align="center" valign="middle" >1.60</td><td align="center" valign="middle" >2.16</td><td align="center" valign="middle" >2.85</td><td align="center" valign="middle" >0.79</td><td align="center" valign="middle" >27.38</td></tr><tr><td align="center" valign="middle" >TMPMOY</td><td align="center" valign="middle" >−51.62</td><td align="center" valign="middle" >−6.55</td><td align="center" valign="middle" >0.61</td><td align="center" valign="middle" >−0.001</td><td align="center" valign="middle" >0.08</td><td align="center" valign="middle" >0.17</td><td align="center" valign="middle" >−0.003</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >−45.91</td></tr><tr><td align="center" valign="middle" >RRQUOT</td><td align="center" valign="middle" >−0.08</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.84</td><td align="center" valign="middle" >0.27</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >−0.03</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >−0.08</td></tr><tr><td align="center" valign="middle" >DRINSQ</td><td align="center" valign="middle" >2.15</td><td align="center" valign="middle" >2.11</td><td align="center" valign="middle" >−1.47</td><td align="center" valign="middle" >0.66</td><td align="center" valign="middle" >0.92</td><td align="center" valign="middle" >1.66</td><td align="center" valign="middle" >1.97</td><td align="center" valign="middle" >1.02</td><td align="center" valign="middle" >2.08</td></tr><tr><td align="center" valign="middle" >HUMREL06h</td><td align="center" valign="middle" >0.51</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >−0.40</td><td align="center" valign="middle" >−0.43</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.10</td><td align="center" valign="middle" >0.36</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.48</td></tr><tr><td align="center" valign="middle" >HUMREL12h</td><td align="center" valign="middle" >0.27</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >1.37</td><td align="center" valign="middle" >0.84</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.53</td><td align="center" valign="middle" >0.33</td><td align="center" valign="middle" >0.14</td><td align="center" valign="middle" >0.28</td></tr><tr><td align="center" valign="middle" >HUMREL18h</td><td align="center" valign="middle" >−0.62</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >−1.33</td><td align="center" valign="middle" >−0.36</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >−0.34</td><td align="center" valign="middle" >−0.48</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >−0.59</td></tr><tr><td align="center" valign="middle" >PRESTN06h</td><td align="center" valign="middle" >−1.46</td><td align="center" valign="middle" >−1.64</td><td align="center" valign="middle" >−0.51</td><td align="center" valign="middle" >−1.04</td><td align="center" valign="middle" >−0.65</td><td align="center" valign="middle" >−0.97</td><td align="center" valign="middle" >−1.15</td><td align="center" valign="middle" >−0.87</td><td align="center" valign="middle" >−1.42</td></tr><tr><td align="center" valign="middle" >PRESTN12h</td><td align="center" valign="middle" >−0.02</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >−0.46</td><td align="center" valign="middle" >−0.89</td><td align="center" valign="middle" >−0.41</td><td align="center" valign="middle" >−0.43</td><td align="center" valign="middle" >−0.28</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >−0.05</td></tr><tr><td align="center" valign="middle" >PRESTN18h</td><td align="center" valign="middle" >0.08</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >−0.54</td><td align="center" valign="middle" >−0.85</td><td align="center" valign="middle" >−0.39</td><td align="center" valign="middle" >−0.15</td><td align="center" valign="middle" >0.06</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.06</td></tr><tr><td align="center" valign="middle" >FFVM06h</td><td align="center" valign="middle" >0.27</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >−0.19</td><td align="center" valign="middle" >0.64</td><td align="center" valign="middle" >0.45</td><td align="center" valign="middle" >0.41</td><td align="center" valign="middle" >0.31</td><td align="center" valign="middle" >0.26</td><td align="center" valign="middle" >0.28</td></tr><tr><td align="center" valign="middle" >FFVM12h</td><td align="center" valign="middle" >0.54</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.37</td><td align="center" valign="middle" >−0.03</td><td align="center" valign="middle" >0.69</td><td align="center" valign="middle" >0.31</td><td align="center" valign="middle" >0.42</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.52</td></tr><tr><td align="center" valign="middle" >FFVM18h</td><td align="center" valign="middle" >0.41</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.35</td><td align="center" valign="middle" >0.47</td><td align="center" valign="middle" >1.18</td><td align="center" valign="middle" >0.43</td><td align="center" valign="middle" >0.41</td><td align="center" valign="middle" >0.52</td><td align="center" valign="middle" >0.41</td></tr><tr><td align="center" valign="middle" >DDVM06deg</td><td align="center" valign="middle" >0.05</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.66</td><td align="center" valign="middle" >0.45</td><td align="center" valign="middle" >0.56</td><td align="center" valign="middle" >0.22</td><td align="center" valign="middle" >0.11</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.07</td></tr><tr><td align="center" valign="middle" >DDVM12hDEG</td><td align="center" valign="middle" >−0.11</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >−0.87</td><td align="center" valign="middle" >−0.60</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >−0.36</td><td align="center" valign="middle" >−0.19</td><td align="center" valign="middle" >−0.11</td><td align="center" valign="middle" >−0.11</td></tr><tr><td align="center" valign="middle" >DDVM18hDEG</td><td align="center" valign="middle" >−0.58</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >−0.18</td><td align="center" valign="middle" >−0.09</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >−0.39</td><td align="center" valign="middle" >−0.54</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >−0.56</td></tr><tr><td align="center" valign="middle" >Vx06</td><td align="center" valign="middle" >−1.33</td><td align="center" valign="middle" >−1.35</td><td align="center" valign="middle" >−1.16</td><td align="center" valign="middle" >−1.23</td><td align="center" valign="middle" >−0.94</td><td align="center" valign="middle" >−1.34</td><td align="center" valign="middle" >−1.31</td><td align="center" valign="middle" >−0.82</td><td align="center" valign="middle" >−1.31</td></tr><tr><td align="center" valign="middle" >Vx12</td><td align="center" valign="middle" >1.28</td><td align="center" valign="middle" >1.10</td><td align="center" valign="middle" >1.03</td><td align="center" valign="middle" >0.39</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.62</td><td align="center" valign="middle" >0.98</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >1.22</td></tr><tr><td align="center" valign="middle" >Vx18</td><td align="center" valign="middle" >−0.70</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.38</td><td align="center" valign="middle" >−0.24</td><td align="center" valign="middle" >0.51</td><td align="center" valign="middle" >−0.61</td><td align="center" valign="middle" >−0.71</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >−0.68</td></tr><tr><td align="center" valign="middle" >Vy06</td><td align="center" valign="middle" >0.69</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >−2.20</td><td align="center" valign="middle" >0.26</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.79</td><td align="center" valign="middle" >0.73</td><td align="center" valign="middle" >0.53</td><td align="center" valign="middle" >0.69</td></tr><tr><td align="center" valign="middle" >Vy12</td><td align="center" valign="middle" >−0.97</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >1.86</td><td align="center" valign="middle" >0.41</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >−0.19</td><td align="center" valign="middle" >−0.69</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.91</td></tr><tr><td align="center" valign="middle" >Vy18</td><td align="center" valign="middle" >0.24</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >1.75</td><td align="center" valign="middle" >1.18</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.67</td><td align="center" valign="middle" >0.36</td><td align="center" valign="middle" >0.00</td><td align="center" valign="middle" >0.27</td></tr><tr><td align="center" valign="middle" >O3veilleJahid</td><td align="center" valign="middle" >23.36</td><td align="center" valign="middle" >23.26</td><td align="center" valign="middle" >22.34</td><td align="center" valign="middle" >23.03</td><td align="center" valign="middle" >23.31</td><td align="center" valign="middle" >23.21</td><td align="center" valign="middle" >22.75</td><td align="center" valign="middle" >23.14</td><td align="center" valign="middle" >22.97</td></tr></tbody></table></table-wrap></sec></body><back><ref-list><title>References</title><ref id="scirp.88053-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Anenberg, S.C., Horowitz, L.W., Tong, D.Q. and West, J.J. (2010) An Estimate of the Global Burden of Anthropogenic Ozone and Fine Particulate Matter on Premature Human Mortality Using Atmospheric Modeling. Environmental Health Perspectives, 118, 1189-1195. https://doi.org/10.1289/ehp.0901220</mixed-citation></ref><ref id="scirp.88053-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Malig, B.J., Pearson, D.L., Chang, Y.B., Broadwin, R., Basu, R., Green, R.S. and Ostro, B. (2016) A Time-Stratified Case-Crossover Study of Ambient Ozone Exposure and Emergency Department Visits for Specific Respiratory Diagnoses in California (2005-2008). Environmental Health Perspectives, 124, 745-753. https://doi.org/10.1289/ehp.1409495</mixed-citation></ref><ref id="scirp.88053-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Russell, B., Cooley, D., Porter, W. and Heald, C. (2016) Modeling the Spatial Behavior of the Meteorological Driver’s Effects on Extreme Ozone. Environmetrics, 27, 334-344. https://doi.org/10.1002/env.2406</mixed-citation></ref><ref id="scirp.88053-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Bel, L., Bellanger, L., Bobbia, M., Ciuperca, G., Dacunha-Castelle, D., Gilibert, E., Jackubowicz, P., Oppenheim, G. and Tomassone, R. (1998) On Forecasting Ozone Episodes in the Paris Area. Listy Biometryczne-Biometrical Letters, 35, 37-66.</mixed-citation></ref><ref id="scirp.88053-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Bel, L., Bellanger, L., Bonneau, V., Ciuperca, G., Dacunha-Castelle, D., Deniau, C., Ghattas, B., Misiti, M., Misiti, Y., Oppenheim, G., Poggi, J.M. and Tomassone, R. (1999) Eléments de comparaison de prévisions statistiques des pics d’ozone. Revue de Statistique Appliquée, XLVII, 7-25.</mixed-citation></ref><ref id="scirp.88053-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Besse, P., Milhem, H., Mestre, O., Dufour, A. and Peuch, V.H. (2007) Comparaison de techniques de &lt;Data Mining&gt; pour l’adaptation statistique des prévisions d’ozone du modèle de chimie-transport MOCAGE. Pollution atmosphérique, 195, 285-292. https://doi.org/10.4267/pollution-atmospherique.1442</mixed-citation></ref><ref id="scirp.88053-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Tamas, W. (2015) Prévision statistique de la qualité de l’air et d’épisodes de pollution atmosphérique en Corse: Génie des procédés. Ph.D. Thesis, Université de Corse Pascale Paoli, Francais, 254 p.</mixed-citation></ref><ref id="scirp.88053-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Lethrosne, M. (2008) Adaptation statistique des prévisions d’ozone issues du système Esméralda. Rapport de stage Master 2 Ingénierie Statistique à Airparif. Université de Rennes.</mixed-citation></ref><ref id="scirp.88053-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Oufdou, H. (2010) Réalisation de modèles de prévision par adaptation statistique. Rapport de stage Master 2 Ingénierie Statistique à AIRAQ. Université Bordeaux2.</mixed-citation></ref><ref id="scirp.88053-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Eljohra, B. (2005) Prévisibilité à 24 heures des concentrations en ozone troposphérique à Casablanca. Centre National de Recherches Météorologiques, Service Météorologie Sectorielle. DMN Maroc.</mixed-citation></ref><ref id="scirp.88053-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Hastie, T., Tibshirani, R. and Friedman, J. (2009) The Element of Statistical Learning: Data Mining, Inference, and Prediction. Springer, Berlin. https://doi.org/10.1007/978-0-387-84858-7</mixed-citation></ref><ref id="scirp.88053-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Hoerl, A. and Kennard, R. (1970) Ridge Regression: Biased Estimation for Non Orthogonal Problems. Technometrics, 12, 55-67. https://doi.org/10.1080/00401706.1970.10488634</mixed-citation></ref><ref id="scirp.88053-ref13"><label>13</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Tibshirani</surname><given-names> R. </given-names></name>,<etal>et al</etal>. (<year>1996</year>)<article-title>Regression Shrinkage and Selection via the Lasso</article-title><source> Journal of the Royal Statistical Society: Series B</source><volume> 58</volume>,<fpage> 267</fpage>-<lpage>288</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.88053-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Joliffe, I.T. (1982) A Note on the Use of Principal Components in Regression. Journal of the Royal Statistical Society: Series C, 31, 300-303. https://doi.org/10.2307/2348005</mixed-citation></ref><ref id="scirp.88053-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Hoskuldsson, A. (1988) PLS Regression Methods. Journal of Chemometrics, 2, 211-228. https://doi.org/10.1002/cem.1180020306</mixed-citation></ref><ref id="scirp.88053-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Tenenhaus, M., Gauchi, J.P. and Ménardo, C. (1995) Régression PLS et applications. Revue de Statistique Appliquée, 43, 7-63.</mixed-citation></ref><ref id="scirp.88053-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Tenenhaus, M. (1998) La régression PLS, théorie et pratique. Technip, Paris.</mixed-citation></ref><ref id="scirp.88053-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Stone, M. and Brooks, R.J. (1990) Continuum Regression: Cross-Validated Sequentially Constructed Prediction Embracing Ordinary Least Squares, Partial Least Squares and Principal Components Regression. Journal of the Royal Statistical Society: Series B, 52, 237-269.</mixed-citation></ref><ref id="scirp.88053-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Sundberg, R. (1993) Continuum Regression and Ridge Regression. Journal of the Royal Statistical Society: Series B: Methodological, 55, 653-659. http://www.jstor.org/stable/2345877</mixed-citation></ref><ref id="scirp.88053-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Qannari, E.M., El Ghaziri, A. and Hanafi, M. (2017) Biased Power Regression: A New Biased Estimation Procedure in Linear Regression. Electronic Journal of Applied Statistical Analysis, 10, 160-179.</mixed-citation></ref><ref id="scirp.88053-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Houze, M.L. (1999) Concentrations en ozone dans les agglomérations dijonnaise et chalonnaise et conditions météorologiques (avril-aout 1998), Mémoire de maitrise de géographie. Université de Bourgogne, 67 p.</mixed-citation></ref><ref id="scirp.88053-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">Little, R.J.A. and Rubin, D.B. (2002, 2014) Statistical Analysis with Missing Data. John Wiley &amp; Sons, Hoboken.</mixed-citation></ref><ref id="scirp.88053-ref23"><label>23</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Sande</surname><given-names> I.G. </given-names></name>,<etal>et al</etal>. (<year>1983</year>)<article-title>Hot-Deck Imputation Procedures</article-title><source> Incomplete Data in Sample Surveys</source><volume> 3</volume>,<fpage> 334</fpage>-<lpage>350</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.88053-ref24"><label>24</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Ford</surname><given-names> B.L. </given-names></name>,<etal>et al</etal>. (<year>1983</year>)<article-title>An Overview of Hot-Deck Procedures</article-title><source> Incomplete Data in Sample Surveys</source><volume> 2</volume>,<fpage> 185</fpage>-<lpage>207</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.88053-ref25"><label>25</label><mixed-citation publication-type="other" xlink:type="simple">Fuller, W.A. and Kim, J.K. (2005) Hot Deck Imputation for the Response Model. Survey Methodology, 31, 139.</mixed-citation></ref><ref id="scirp.88053-ref26"><label>26</label><mixed-citation publication-type="other" xlink:type="simple">Rubin, D.B. (1987) Multiple Imputation for Non Response in Surveys. Wiley, New York. https://doi.org/10.1002/9780470316696</mixed-citation></ref><ref id="scirp.88053-ref27"><label>27</label><mixed-citation publication-type="other" xlink:type="simple">Beretta, L. and Santaniello, A. (2016) Nearest Neighbor Imputation Algorithms: A Critical Evaluation. BMC Medical Informatics and Decision Making, 16, 74. https://doi.org/10.1186/s12911-016-0318-z</mixed-citation></ref><ref id="scirp.88053-ref28"><label>28</label><mixed-citation publication-type="other" xlink:type="simple">Jolliffe, I.T. (2002) Principal Component Analysis. Second Edition, Chapter 7.</mixed-citation></ref><ref id="scirp.88053-ref29"><label>29</label><mixed-citation publication-type="other" xlink:type="simple">Husson, F., Le, S. and Pages, J. (2010) Exploratory Multivariate Analysis by Example Using R. Chapman and Hall, London. https://doi.org/10.1201/b10345</mixed-citation></ref><ref id="scirp.88053-ref30"><label>30</label><mixed-citation publication-type="other" xlink:type="simple">Bellanger, L. and Tomassone, R. (2014) Exploration de données et méthodes statistiques: Data analysis &amp; Data mining avec R. Collection Références Sci, Editions Ellipses, Paris, 480 p.</mixed-citation></ref><ref id="scirp.88053-ref31"><label>31</label><mixed-citation publication-type="other" xlink:type="simple">Le Cao, K.A., Rossouw, D., Robert-Granié, C. and Besse, P. (2008) A Sparse PLS for Variable Selection When Integrating Omics Data. Statistical Applications in Genetics and Molecular Biology, 7, 32. https://doi.org/10.2202/1544-6115.1390</mixed-citation></ref><ref id="scirp.88053-ref32"><label>32</label><mixed-citation publication-type="other" xlink:type="simple">Chun, H. and Keles, S. (2010) Sparse Partial Least Squares Regression for Simultaneous Dimension Reduction and Variable Selection. Journal of the Royal Statistical Society: Series B, Statistical Methodology, 72, 3-25. https://doi.org/10.1111/j.1467-9868.2009.00723.x</mixed-citation></ref><ref id="scirp.88053-ref33"><label>33</label><mixed-citation publication-type="other" xlink:type="simple">Draper, N.R. and Smith, H. (1998) Selecting the “Best” Regression Equation. In: Applied Regression Analysis, 3rd Edition, John Wiley &amp; Sons, Inc., Hoboken, 473-504.</mixed-citation></ref><ref id="scirp.88053-ref34"><label>34</label><mixed-citation publication-type="other" xlink:type="simple">James, G., Witten, D., Hastie, T. and Tibshirani, R. (2013) An Introduction to Statistical Learning with Applications in R. Springer Science + Business Media, New York.</mixed-citation></ref><ref id="scirp.88053-ref35"><label>35</label><mixed-citation publication-type="book" xlink:type="simple">Wold, H. (1966) Estimation of Principal Components and Related Models by Iterative Least Squares. In: Krishnaiah, P.R., Ed., Multivariate Analysis, Academic Press, New York, 391-420.</mixed-citation></ref><ref id="scirp.88053-ref36"><label>36</label><mixed-citation publication-type="other" xlink:type="simple">Abudu, S., Cui, C.L., King, J.P. and Abudukadeer, K. (2010) Comparison of Performance of Statistical Models in Forecasting Monthly Streamflow of Kizil River, China. Water Science and Engineering, 3, 269-281.</mixed-citation></ref><ref id="scirp.88053-ref37"><label>37</label><mixed-citation publication-type="other" xlink:type="simple">Sayegh, A., Munir, S. and Habeebullah, T.M. (2014) Comparing the Performance of Statistical Models for Predicting PM10 Concentrations. Aerosol and Air Quality Research, 14, 653-665. https://doi.org/10.4209/aaqr.2013.07.0259</mixed-citation></ref><ref id="scirp.88053-ref38"><label>38</label><mixed-citation publication-type="other" xlink:type="simple">Batista, G.E.A.P.A. and Monard, M.C. (2002) K-Nearest Neighbour as Imputation Method: Experimental Results. Technical Report, ICMC-USP.</mixed-citation></ref><ref id="scirp.88053-ref39"><label>39</label><mixed-citation publication-type="other" xlink:type="simple">Akaike, H. (1974) A New Look at the Statistical Model Identification. IEEE Transactions on Automatic Control, 19, 716-723. https://doi.org/10.1109/TAC.1974.1100705</mixed-citation></ref><ref id="scirp.88053-ref40"><label>40</label><mixed-citation publication-type="other" xlink:type="simple">Zhang, Z. (2016) Variable Selection with Stepwise and Best Subset Approaches. Annals of Translational Medicine, 4, 136. https://doi.org/10.21037/atm.2016.03.35</mixed-citation></ref></ref-list></back></article>