<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OJS</journal-id><journal-title-group><journal-title>Open Journal of Statistics</journal-title></journal-title-group><issn pub-type="epub">2161-718X</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/ojs.2017.74039</article-id><article-id pub-id-type="publisher-id">OJS-78088</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Sparse Additive Gaussian Process with Soft Interactions
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Garret</surname><given-names>Vo</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Debdeep</surname><given-names>Pati</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>Department of Industrial and Manufacturing Engineering, Florida State University, Tallahassee, FL, USA</addr-line></aff><aff id="aff2"><addr-line>Department of Statistics, Texas A and M University, College Station, TX, USA</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>garretvo19@gmail.com(GV)</email>;</corresp></author-notes><pub-date pub-type="epub"><day>21</day><month>07</month><year>2017</year></pub-date><volume>07</volume><issue>04</issue><fpage>567</fpage><lpage>588</lpage><history><date date-type="received"><day>13,</day>	<month>June</month>	<year>2017</year></date><date date-type="rev-recd"><day>28,</day>	<month>July</month>	<year>2017</year>	</date><date date-type="accepted"><day>31,</day>	<month>July</month>	<year>2017</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  This paper presents a novel variable selection method in additive nonparametric regression model. This work is motivated by the need to select the number of nonparametric components and number of variables within each nonparametric component. The proposed method uses a combination of hard and soft shrinkages to separately control the number of additive components and the variables within each component. An efficient algorithm is developed to select the importance of variables and estimate the interaction network. Excellent performance is obtained in simulated and real data examples.
 
</p></abstract><kwd-group><kwd>Additive</kwd><kwd> Gaussian Process</kwd><kwd> Interaction</kwd><kwd> Lasso</kwd><kwd> Sparsity</kwd><kwd> Variable Selection</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Variable selection has played a pivotal role in scientific and engineering applications, such as biochemical analysis [<xref ref-type="bibr" rid="scirp.78088-ref1">1</xref>] , bioinformatics [<xref ref-type="bibr" rid="scirp.78088-ref2">2</xref>] and text mining [<xref ref-type="bibr" rid="scirp.78088-ref3">3</xref>] , among other areas. A significant portion of existing variable selection methods are only applicable to linear parametric models. Despite the linearity and additivity assumption, variable selection in linear regression models has been popular since 1970, referring to Akaike information criterion (AIC; [<xref ref-type="bibr" rid="scirp.78088-ref4">4</xref>] ); Bayesian information criterion (BIC; [<xref ref-type="bibr" rid="scirp.78088-ref5">5</xref>] ) and Risk inflation criterion (RIC; [<xref ref-type="bibr" rid="scirp.78088-ref6">6</xref>] ).</p><p>Popular classical sparse-regression methods such as Least absolute shrinkage operator (LASSO [<xref ref-type="bibr" rid="scirp.78088-ref7">7</xref>] [<xref ref-type="bibr" rid="scirp.78088-ref8">8</xref>] ), and related penalization methods [<xref ref-type="bibr" rid="scirp.78088-ref9">9</xref>] [<xref ref-type="bibr" rid="scirp.78088-ref10">10</xref>] [<xref ref-type="bibr" rid="scirp.78088-ref11">11</xref>] [<xref ref-type="bibr" rid="scirp.78088-ref12">12</xref>] have gained popularity over the last decade due to their simplicity, computational scalability and efficiency in prediction when the underlying relation between the response and the predictors can be adequately described by parametric models. Bayesian methods [<xref ref-type="bibr" rid="scirp.78088-ref13">13</xref>] [<xref ref-type="bibr" rid="scirp.78088-ref14">14</xref>] [<xref ref-type="bibr" rid="scirp.78088-ref15">15</xref>] with sparsity inducing priors offer greater applicability beyond parametric models and are a convenient alternative when the underlying goal is in inference and uncertainty quantification. However, there is still a limited amount of literature which seriously considers relaxing the linearity assumption, particularly when the dimension of the predictors is high. Moreover, when the focus is on learning the interactions between the variables, parametric models are often restrictive since they require very many parameters to capture the higher-order interaction terms.</p><p>Smoothing based non-additive nonparametric regression methods [<xref ref-type="bibr" rid="scirp.78088-ref16">16</xref>] [<xref ref-type="bibr" rid="scirp.78088-ref17">17</xref>] [<xref ref-type="bibr" rid="scirp.78088-ref18">18</xref>] [<xref ref-type="bibr" rid="scirp.78088-ref19">19</xref>] can accommodate a wide range of relationships between predictors and response leading to excellent predictive performance. Such methods have been adapted for different methods of functional component selection with non- linear interaction terms: component selection and smoothing operator (COSSO; [<xref ref-type="bibr" rid="scirp.78088-ref20">20</xref>] ), sparse addictive model (SAMS; [<xref ref-type="bibr" rid="scirp.78088-ref21">21</xref>] ) and variable selection using adaptive nonlinear interaction structure in high dimensions (VANISH; [<xref ref-type="bibr" rid="scirp.78088-ref22">22</xref>] ). However, when the number of variables is large and their interaction network is complex, modeling each functional component is highly expensive.</p><p>Nonparametric variable selection based on kernel methods is increasingly becoming popular over the last few years. Liu et al. [<xref ref-type="bibr" rid="scirp.78088-ref23">23</xref>] provided a connection between the least square kernel machine (LKM) and the linear mixed models. Zou et al. [<xref ref-type="bibr" rid="scirp.78088-ref24">24</xref>] , Savitsky et al. [<xref ref-type="bibr" rid="scirp.78088-ref25">25</xref>] introduced Gaussian process with dimension- specific scalings for simultaneous variable selection and prediction. Yang et al. [<xref ref-type="bibr" rid="scirp.78088-ref26">26</xref>] argued that a single Gaussian process with variable bandwidths can achieve the optimal rate in estimation when the true number of covariates s ≍ O ( l o g n ) . However, when the true number of covariates is relatively high, the suitability of using a single Gaussian process is questionable. Moreover, such an approach is not convenient to recover the interaction among variables. Fang et al. [<xref ref-type="bibr" rid="scirp.78088-ref27">27</xref>] used the nonnegative Garotte kernel to select variables and capture interaction. Though these methods can successfully perform variable selection and capture the interaction, non-additive nonparametric models are not sufficiently scalable when the dimension of the relevant predictors is even moderately high. [<xref ref-type="bibr" rid="scirp.78088-ref27">27</xref>] claimed that extensions to additive models may cause over-fitting issues in capturing the interaction between variables (i.e. capture more interacting variables than the ones which are influential).</p><p>To circumvent this bottleneck, Yang et al. [<xref ref-type="bibr" rid="scirp.78088-ref26">26</xref>] , Qamar and Tokdar [<xref ref-type="bibr" rid="scirp.78088-ref28">28</xref>] introduced the additive Gaussian process with sparsity inducing priors for both the number of components and variables within each component. The additive Gaussian process captures interactions among variables, can scale up to moderately high dimensions and are suitable for low sparse regression functions. However, the use of two component sparsity inducing prior forced them to develop a tedious Markov chain Monte Carlo algorithm to sample from the posterior distribution.</p><p>To overcome the computational challenge facing in Yang et al. [<xref ref-type="bibr" rid="scirp.78088-ref26">26</xref>] , Qamar and Tokdar [<xref ref-type="bibr" rid="scirp.78088-ref28">28</xref>] , we propose a novel method, called the additive Gaussian process with soft interactions. More specifically, we decompose the unknown regression function F into k components, such as F = ϕ 1 f 1 + ϕ 2 f 2 + ⋯ + ϕ k f k , for k hard shrinkage parameters ϕ l , l = 1 , ⋯ , k , k ≥ 1 . Each component f l is independent. Each of them is assigned to a Gaussian process prior. To induce sparsity within each Gaussian process, we introduce an additional level of soft shrinkage parameters. The combination of hard and soft shrinkage priors makes our approach very straightforward to implement and computationally efficient, while retaining all the advantages of the additive Gaussian process proposed by Qamar and Tokdar [<xref ref-type="bibr" rid="scirp.78088-ref28">28</xref>] . We propose a combination of Markov chain Monte Carlo (MCMC) and the Least Angle Regression algorithm (LARS) to select the Gaussian process components and variables within each component.</p><p>The rest of the paper is organized as follows. Section 2 presents the additive Gaussian process model. Section 3 describes the two-level regularization and the prior specifications. The posterior computation is detailed in Section 4 and the variable selection and interaction recovery approach are presented in Section 5. The simulation study results are presented in Section 6. A couple of real data examples are considered in Section 7. We conclude with a discussion in Section 8.</p></sec><sec id="s2"><title>2. Additive Gaussian Process</title><p>For observed predictor-response pairs ( x i , y i ) ∈ ℝ p &#215; ℝ , where i = 1 , 2 , ⋯ , n (i.e. n is the sample size and p is the dimension of the predictors), an additive nonparametric regression model can be expressed as</p><p>y i = F ( x i ) + ϵ i ,   ϵ i ~ Ν ( 0 , σ 2 ) F ( x i ) = ϕ 1 f 1 ( x i ) + ϕ 2 f 2 ( x i ) + ⋯ + ϕ k f k ( x i ) . (1)</p><p>The regression function F in (1) is a sum of k regression functions, with the relative importance of each function controlled by the set of non-negative parameters ϕ = ( ϕ 1 , ϕ 2 , ⋯ , ϕ k ) T . Typically the unknown parameter ϕ is assumed to be sparse to prevent F from over-fitting the data.</p><p>Gaussian process (GP) [<xref ref-type="bibr" rid="scirp.78088-ref29">29</xref>] provides a flexible prior for each of the component functions in { f l , l = 1 , ⋯ , k } . GP defines a prior on the space of all continuous functions, denoted f ~ GP ( μ , c ) for a fixed function μ : ℝ p → ℝ and a positive definite function c defined on ℝ p &#215; ℝ p such that for any finite collection of points { x i , i = 1 , ⋯ , L } , the distribution of { f ( x 1 ) , ⋯ , f ( x L ) } is multivariate Gaussian with mean { μ ( x 1 ) , ⋯ , μ ( x L ) } and variance-covariance matrix Σ = { c ( x i , x i ′ ) } 1 ≤ , i , i ′ ≤ L . The choice of the covariance kernel is crucial to ensure the sample path realizations of the Gaussian process are appropriately smooth. A squared exponential covariance kernel c ( x , x ′ ) = e x p ( − κ ‖ x − x ′ ‖ 2 ) with an Gamma hyperprior assigned to the inverse-bandwidth parameter κ ensures optimal estimation of an isotropic regression function [<xref ref-type="bibr" rid="scirp.78088-ref30">30</xref>] even when a single component function is used ( k = 1 ). When the dimension of the covariates is high, it is natural to assume that the underlying regression function is not isotropic. In that case, Bhattacharya et al. [<xref ref-type="bibr" rid="scirp.78088-ref31">31</xref>] showed that a single bandwidth parameter might be inadequate and dimension specific scalings with appropriate shrinkage priors are required to ensure that the posterior distribution can adapt to the unknown dimensionality. However, Yang et al. [<xref ref-type="bibr" rid="scirp.78088-ref26">26</xref>] showed that single Gaussian process might not be appropriate to capture interacting variables and also does not scale well with the true dimension of the predictor space. In that case, an additive Gaussian process is a more effective alternative which also leads to interaction recovery as a bi-product. In this article, we work with the additive representation in (1) with dimension specific scalings (inverse-bandwidth parameters) κ l j along dimension j for the lth Gaussian process component, j = 1 , ⋯ , p and l = 1 , ⋯ , k .</p><p>We assume that the response vector y = ( y 1 , y 2 , ⋯ , y n ) in (1) is centered and scaled. Let f l ~ GP ( 0, c l ) with</p><p>c l ( x , x ′ ) = exp { − ∑ j = 1 p       κ l j ( x j − x ′ j ) 2 } . (2)</p><p>In the next section, we discuss appropriate regularization on ϕ and { κ l j , l = 1 , ⋯ , k ; j = 1 , ⋯ , p } . A shrinkage prior on the { κ l j , j = 1 , ⋯ , p } facilitates the selection of variables within component l and allows adaptive local smoothing. An appropriate regularization on ϕ allows F to adapt to the degree of additivity in the data without over-fitting.</p></sec><sec id="s3"><title>3. Regularization</title><p>A full Bayesian specification will require placing prior distribution on both ϕ and κ . However, such a specification requires tedious posterior sampling algorithms to sample from the posterior distribution as seen in [<xref ref-type="bibr" rid="scirp.78088-ref28">28</xref>] . Moreover, it is difficult to identify the role of ϕ l and κ j l , j = 1 , ⋯ , p since one can remove the effect of the lth component by either setting ϕ l to zero or by having κ l j = 0 , j = 1 , ⋯ , p . This ambiguous representation causes mixing issues in a full-blown MCMC. To facilitate computation, we adopt a hybrid approach between frequentist and Bayesian to regularize ϕ and κ l j , respectively. The hybrid-algorithm is a combination of i) MCMC, to sample κ conditional on ϕ ii) and optimization to estimate ϕ conditional on κ . With this viewpoint, we propose the following regularization on κ and ϕ . With the parameter γ , each component controls the selection of variables and interaction among them. In addition to γ , the parameter Γ allows the model (1) to select significant components, which includes interested variables and interaction network. Together, Γ and γ are the global-local shrinkage on F.</p><sec id="s3_1"><title>3.1. L<sub>1</sub> Regularization for f</title><p>Conditional on f 1 , ⋯ , f k , (1) with ϕ l , and ϕ l &gt; 0 . Hence we impose L 1 regularization on ϕ l , which is as following</p><p>1 n ∑ i = 1 N { y ( x i ) − ∑ l = 1 k     ϕ l f l ( x i ) } + λ ∑ l = 1 k     ϕ l (3)</p><p>In the algorithm, ϕ l is updated using least absolute shrinkage and selection operator (LASSO) [<xref ref-type="bibr" rid="scirp.78088-ref7">7</xref>] [<xref ref-type="bibr" rid="scirp.78088-ref32">32</xref>] [<xref ref-type="bibr" rid="scirp.78088-ref33">33</xref>] . The L 1 regularization enforces sparsity on ϕ at each stage of the algorithm, thereby pruning the unnecessary Gaussian process components in F. The parameter λ in (3) is selected using five fold cross validation.</p></sec><sec id="s3_2"><title>3.2. Choice of k Components</title><p>The proposed model has the number of components, k, which determines how many components to fit the data and build the prediction. We propose using LASSO to choose k. First, we start with a large k value. As ϕ j is updated with the LASSO algorithm, the LASSO algorithm prunes unnecessary Gaussian process f l . Therefore, the value of k is updated, which is equal to the number of components which are not pruned.</p></sec><sec id="s3_3"><title>3.3. Global-local shrinkage for κ l j</title><p>The parameters κ l j controls the effective number of variables within each component. For each l, { κ l j , j = 1 , ⋯ , p } are assumed to be sparse. As opposed to the two component mixture prior on κ l j in [<xref ref-type="bibr" rid="scirp.78088-ref28">28</xref>] , we enforce weak-sparsity using a global-local continuous shrinkage prior which potentially have substantial computational advantages over mixture priors. Many continuous shrinkage priors have been proposed recently [<xref ref-type="bibr" rid="scirp.78088-ref34">34</xref>] - [<xref ref-type="bibr" rid="scirp.78088-ref39">39</xref>] . These priors can be unified through a global-local (GL) scale mixture representation of [<xref ref-type="bibr" rid="scirp.78088-ref40">40</xref>] below,</p><p>κ l j ~ N ( 0, ψ l j τ l ) ,   τ l ~ f g ,   ψ l j ~ f l , (4)</p><p>for each fixed l, where f g and f l are densities on the positive real line. In (4), τ l controls global shrinkage towards the origin while the local parameters { ψ l j , j = 1 , ⋯ , p } allow local deviations in the degree of shrinkage for each predictor. Special cases include Bayesian lasso [<xref ref-type="bibr" rid="scirp.78088-ref34">34</xref>] , relevance vector machine [<xref ref-type="bibr" rid="scirp.78088-ref35">35</xref>] , normal-gamma mixtures [<xref ref-type="bibr" rid="scirp.78088-ref36">36</xref>] and the horseshoe [<xref ref-type="bibr" rid="scirp.78088-ref37">37</xref>] [<xref ref-type="bibr" rid="scirp.78088-ref38">38</xref>] among others. Motivated by the remarkable performance of horseshoe, we assume both f g and f l to be square-root of half-Cauchy distributions. Both τ l and ψ l j will be updated using the MCMC algorithm.</p></sec></sec><sec id="s4"><title>4. Hybrid Algorithm for Prediction, Selection and Interaction Recovery</title><p>In this section, we develop a fast algorithm which is a combination of L 1 optimization and conditional MCMC to estimate the parameters ϕ l , ψ l j , and τ l for l = 1 , ⋯ , k and j = 1 , ⋯ , p . Conditional on κ l j , (1) is linear in ϕ l and hence we resort to the least angle regression procedure [<xref ref-type="bibr" rid="scirp.78088-ref8">8</xref>] with five fold cross validation to estimate ϕ l , l = 1 , ⋯ , k . The computation of the lasso solutions is a quadratic programming problem, and can be tackled by standard numerical analysis algorithms.</p><p>The least angle regression procedure better approach which exploits the special structure of the lasso problem, and provides an efficient way to compute the solutions. Next, we describe the conditional MCMC to sample from κ l j and F ( x * ) at a new point x * conditional on the parameters ϕ l . For two collection of vectors X v and Y v of size m 1 and m 2 respectively, denote by c ( X v , Y v )</p><p>the m 1 &#215; m 2 matrix { c ( x , y ) } x ∈ X v , y ∈ Y v . Let X = { x 1 , x 2 , ⋯ , x n } and define</p><p>c ( X , X ) , c ( x * , X ) , c ( X , x * ) and c ( x * , x * ) denote the corresponding matrices. For a random variable q, we denote by q | − the conditional distribution of q given the remaining random variables.</p><p>Observe that the algorithm does not necessarily produce samples which are approximately distributed as the true posterior distribution. The combination of optimization and conditional sampling is similar to stochastic EM [<xref ref-type="bibr" rid="scirp.78088-ref41">41</xref>] [<xref ref-type="bibr" rid="scirp.78088-ref42">42</xref>] which is employed to avoid computing costly integrals required to find maximum likelihood in latent variable models. Conditional on ϕ l , l = 1 , ⋯ , k , the MCMC algorithm to update ψ l j , τ l , and ϕ l is as following:</p><p>1) Compute the kernel k ( x , x ) , k ( x , x * ) , k ( x * , x ) , k ( x * , x * ) with the kernel formula k ( x , x ′ ) = exp ( − γ d j ‖ x − x ′ ‖ 2 ) .</p><p>2) Compute f l − ( x i ) = ∑ j ≠ l     ϕ j f j ( x i ) . Compute the predictive mean</p><p>μ l * = k ( x * , x ) [ c ( X , X ) + σ 2 I ] − 1 ( y − f l − ) (5)</p><p>3) Compute the predictive variance</p><p>Σ l * = c ( x * , x * ) − c ( x * , X ) [ c ( X , X ) + σ 2 ] − 1 c ( X , x * ) . (6)</p><p>4) Sample f l | − , y ~ N ( μ l * , Σ l * ) .</p><p>5) Compute the predictive</p><p>F ( x * ) = ϕ 1 f 1 * + ϕ 2 f 2 * + ⋯ + ϕ k f k * . (7)</p><p>6) Update ψ l j by sampling from the following posterior distribution, p ( ψ l j | − , y )</p><p>p ( ψ l j | − , y ) ∝ e x p { − 1 2 y T [ c ( X , X ) + σ 2 I ] − 1 y } | c ( X , X ) + σ 2 I | p ( ψ l j ) . (8)</p><p>7) Update τ l , j = 1 , ⋯ , k by sampling from the following posterior distribution, p ( τ l | − , y )</p><p>p ( τ l | − , y ) ∝ e x p { − 1 2 y T [ c ( X , X ) + σ 2 I ] − 1 y } | c ( X , X ) + σ 2 I | p ( τ l ) . (9)</p><p>8) Update γ d j by using the formula γ d j = τ j Ψ d j .</p><p>9) Update the vector Γ with the LASSO estimation.</p><p>10) Update κ l j by sampling</p><p>κ l j ~ N ( 0, ψ l j τ l ) (10)</p><p>11) Update ϕ j and prune unnecessary f j where j ≠ l with the LASSO algorithm.</p><p>The MCMC algorithm above is illustrated with the following flow-chart.</p><disp-formula id="scirp.78088-formula14"><graphic  xlink:href="//html.scirp.org/file/3-1240904x129.png"  xlink:type="simple"/></disp-formula><p>In the MCMC algorithm above, the conditional distributions of τ j and ψ l j are not available in closed form. Therefore, we sample them using Metropolis- Hastings algorithm [<xref ref-type="bibr" rid="scirp.78088-ref43">43</xref>] . In this paper, we give the algorithm for updating τ l only, as the steps for ψ l j are similar. Assuming that the chain is currently at the iteration t, the Metropolis-Hastings algorithm to sample τ l t + 1 independently for l = 1 , ⋯ , k proceeds as following:</p><p>1) Propose log τ l * ~ N ( log τ l t , σ τ 2 ) .</p><p>2) Compute the Metropolis ratio:</p><p>p = min [ p ( τ l * | − ) p ( τ l t | − ) , 1 ] (11)</p><p>3) Sample u ~ U ( 0,1 ) . If u &lt; p then log τ l t + 1 = log τ l * , else log τ l t + 1 = log τ l t .</p><p>The flowchart for the above Metropolis-Hastings algorithm is as following:</p><disp-formula id="scirp.78088-formula15"><graphic  xlink:href="//html.scirp.org/file/3-1240904x142.png"  xlink:type="simple"/></disp-formula><p>The proposal variance σ τ 2 is tuned to ensure that the acceptance probability is between 20% - 40%. We also propose a similar Metropolis-Hastings algorithm to sample from the conditional distribution of ψ l j | − .</p></sec><sec id="s5"><title>5. Variable Selection and Interaction Recovery for Selected Variables</title><p>In this section, we first state a generic algorithm to select important variables based on the samples of the parameter vector γ . This algorithm is independent of the prior for γ and unlike other variable selection algorithms, it requires few tuning parameters making it suitable for practical purposes. The idea is based on finding the most probable set of variables in the median of the γ samples. Since the distribution for the number of important variables is more stable and largely unaffected by the Metropolis-Hastings algorithm, we find the mode H of the distribution for the number of important variables. Then, we select the H largest coefficients from the posterior mean of γ .</p><p>In this algorithm, we use k-means algorithm [<xref ref-type="bibr" rid="scirp.78088-ref44">44</xref>] [<xref ref-type="bibr" rid="scirp.78088-ref45">45</xref>] with k = 2 at each iteration to form two clusters, corresponding to signal and noise variables respectively. One cluster contains values concentrating around zero, corresponding to the noise variables. The other cluster contains values concentrating away from zeros, corresponding to the signals. At the t<sup>th</sup> iteration, the number of non-zero signals h ( t ) is estimated by the smaller cluster size out of the two clusters. We take the mode over all the iterations to obtain the final estimate H for the number of non-zero signals i.e. H = mode ( h ( t ) ) . The H largest entries of the posterior median of | γ | are identified as the non-zero signals.</p><p>We run the algorithm for 5000 iterations with a burn-in of 2000 to ensure convergence. Based on the remaining iterates, we apply the algorithm to κ j l for each component f l to select important variables within each f l for l = 1 , ⋯ , k . Using this approach, we select the important variables within each function. We define the inclusion score of a variable as the proportion of functions (out of k) which contains that variable. Next, we apply the algorithm to ϕ and select the important functions. Let us denote by A f the set of active functions, obtained from the LASSO algorithm as discussed in Section 3.2. The interaction score between a pair of selected variables is defined as the proportion of functions within A f in which the selected pair appears together. Using these interaction scores, we can find the interaction between important variables with optimal number of active components. Observe that the inclusion and interaction scores are not a functional of the posterior distribution and is purely a property of the additive representation. Hence, we do not require the sampler to converge to the posterior distribution. As illustrated in Section 6, these inclusion and the interaction scores provide an excellent representation of a variable or an interaction being present or absent in the model. An illustratfor both variable selection and interaction will be displayed in Section 6.</p></sec><sec id="s6"><title>6. Simulation Examples</title><p>In this section, we consider eight different simulation settings with 50 replicated datasets each and test the performance of our algorithm with respect to variable selection, interaction recovery, and prediction. To generate the simulated data, we draw x i j ~ Unif ( 0,1 ) , and y i ~ N ( f ( x i ) , σ 2 ) , where 1 ≤ i ≤ n , 1 ≤ j ≤ p and σ 2 = 0.02 . <xref ref-type="table" rid="table1">Table 1</xref> and <xref ref-type="table" rid="table2">Table 2</xref> summary the result and signal to noise ratio (SNR) for the eight different datasets with different combinations of p and n for both non-interaction and interaction cases, respectively.</p><sec id="s6_1"><title>6.1. Variable Selection</title><p>We compute the Inclusion score for each variable in each simulated dataset, then provide the bar plots as in Figures 1-4 below.</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Non-interaction simulated datasets</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" ></th><th align="center" valign="middle" ></th><th align="center" valign="middle"  colspan="2"  >Equation for the Dataset</th></tr></thead><tr><td align="center" valign="middle" >Simulated Dataset</td><td align="center" valign="middle" >n</td><td align="center" valign="middle" >p</td><td align="center" valign="middle" >Non-interaction Data</td><td align="center" valign="middle" >SNR</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/3-1240904x167.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >37.3274</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/3-1240904x168.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >36.9188</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >20</td><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/3-1240904x169.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >41.1118</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/3-1240904x170.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >41.6303</td></tr></tbody></table></table-wrap><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Interaction simulated datasets</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" ></th><th align="center" valign="middle" ></th><th align="center" valign="middle"  colspan="2"  >Equation for the Dataset</th></tr></thead><tr><td align="center" valign="middle" >Simulated Dataset</td><td align="center" valign="middle" >n</td><td align="center" valign="middle" >p</td><td align="center" valign="middle" >Interaction Data</td><td align="center" valign="middle" >SNR</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/3-1240904x171.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >41.9095</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/3-1240904x172.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >42.1258</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >20</td><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/3-1240904x173.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >43.0888</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/3-1240904x174.png" xlink:type="simple"/></inline-formula></td><td align="center" valign="middle" >44.4024</td></tr></tbody></table></table-wrap><p>From these histograms, we rank the Inclusion score value. Based on our ranking, we select a threshold value to identify the signal based on the top Inclusion score values. From our ranking, the selected threshold value is 0.1. The ranking of variables selection has been mentioned in Guyon and Elisseeff [<xref ref-type="bibr" rid="scirp.78088-ref46">46</xref>] , Forman [<xref ref-type="bibr" rid="scirp.78088-ref47">47</xref>] , Stoppiglia et al. [<xref ref-type="bibr" rid="scirp.78088-ref48">48</xref>] . The choice of the threshold variable based upon the data has been mentioned in Genuer et al. [<xref ref-type="bibr" rid="scirp.78088-ref49">49</xref>] . When we obtain selected variables, we compute the false positive rate (FPR), which is the proportion of true signals not detected by our algorithm, and false negative rate (FNR), which is the proportion of false signals detected by our algorithm. Both values are recorded in <xref ref-type="table" rid="table3">Table 3</xref> to assess the quantitative performance of our algorithm.</p><p>Based on the results in <xref ref-type="table" rid="table3">Table 3</xref>, it is immediate that the algorithm is very successful in delivering accurate variable selection for both non-interaction and interaction cases.</p></sec><sec id="s6_2"><title>6.2. Interaction Recovery</title><p>In order to capture the interaction network, we compute the probability of interaction between two variables by calculating the proportion of functions in which both the variables jointly appear. Since we are interested in capturing the interaction between selected variables, we plot interaction heat map for selected variables with their probability of interaction values, for each dataset for both the non-interaction and interaction cases.</p><p>Based on Figures 5-8, it is evident that the estimated interaction probabilities for the non-interacting variables are less than the corresponding number for interacting variables. With these heat map values, we plot the interaction</p><table-wrap id="table3" ><label><xref ref-type="table" rid="table3">Table 3</xref></label><caption><title> The average false positive (FPR) and false negative (FNR) for replicated datasets</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle"  colspan="2"  >Non-interaction Dataset</th><th align="center" valign="middle"  colspan="2"  >Interaction Dataset</th></tr></thead><tr><td align="center" valign="middle" >Dataset</td><td align="center" valign="middle" >FPR</td><td align="center" valign="middle" >FNR</td><td align="center" valign="middle" >FPR</td><td align="center" valign="middle" >FNR</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.05</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.01</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.01</td></tr></tbody></table></table-wrap><p>network in <xref ref-type="fig" rid="fig9">Figure 9</xref> &amp; <xref ref-type="fig" rid="fig10">Figure 10</xref> for selected variables.</p><p>Based on the interaction network in <xref ref-type="fig" rid="fig9">Figure 9</xref> &amp; <xref ref-type="fig" rid="fig10">Figure 10</xref>, we observe that edges for interaction cases are thicker than edges for non-interaction cases. In interaction cases, interacted variables are connected in the network, while every variables is connected in non-interaction cases. Therefore, our algorithm successfully captures the interaction network in all the datasets for selected variables according to the Inclusion score.</p></sec><sec id="s6_3"><title>6.3. Predictive Performance</title><p>We randomly partition each dataset into training (50%) and test (50%) observations. We apply our algorithm on the training data and compare the performance on the test dataset. For the sake of brevity we plot the predicted vs. the observed test observations only for a few cases in <xref ref-type="fig" rid="fig11">Figure 11</xref>.</p><p>From <xref ref-type="fig" rid="fig11">Figure 11</xref>, the predicted observations and the true observations fall very closely along the y = x line demonstrating a good predictive performance. We compare our results with [<xref ref-type="bibr" rid="scirp.78088-ref27">27</xref>] . However, their additive model was not able to capture higher order interaction and thus have a poor predictive performance compared to our method.</p></sec><sec id="s6_4"><title>6.4. Comparison with BART</title><p>Bayesian Additive Regression Tree (BART; [<xref ref-type="bibr" rid="scirp.78088-ref50">50</xref>] ) is a state of the art method for variable selection in nonparametric regression problems. BART is a Bayesian “sum of tree” framework which fits and infers the data through an iterative back-fitting MCMC algorithm to generate samples from a posterior. Each tree in BART [<xref ref-type="bibr" rid="scirp.78088-ref50">50</xref>] is constrained by a regularization prior. Hence BART is similar to our method which also resorts to back-fitting MCMC to generate samples from a posterior.</p><p>Since BART is well-known to deliver excellent prediction results, its performance in terms of variable selection and interaction recovery in high- dimensional setting is worth investigating. In this section, we compare our method with BART in all the three aspects: variable selection, interaction recovery and predictive performance. For comparison, with BART, we used the same simulation settings as in <xref ref-type="table" rid="table1">Table 1</xref> with all combinations of (n, p), where n = 100 and p = 10 , 20 , 100 , 150 , 200 .</p><p>We used 50 replicated datasets and compute average inclusion probabilities for each variable. Similar to &#167;6.1, we ranked the Inclusion score, and chose the threshold value equal to 0.1 in order to find selected variables. Then, we computed the false positive and false negative rates for both algorithms as in <xref ref-type="table" rid="table3">Table 3</xref>. These values are recorded in <xref ref-type="table" rid="table4">Table 4</xref>.</p><p>In <xref ref-type="table" rid="table4">Table 4</xref>, the first column indicates which equations are used to generate the data with the respective p and n values in the second and third column for both non-interaction and interaction cases. For example, if the dataset is 1, the equations to generate the data is x 1 + x 2 2 + x 3 + ϵ and x 1 + x 2 2 + x 3 + x 1 x 2 + x 2 x 3 + x 3 x 1 + ϵ for non-interaction and interaction case, respectively. NA value means that the algorithm cannot run at all for that particular combination of p and n values.</p><p>According to <xref ref-type="table" rid="table4">Table 4</xref>, BART performs similar to our algorithm when p = 10 and n = 100 . However, as p increases, BART fails to perform adequately while our algorithm still performs well even when p is larger than n. When p is twice</p><table-wrap id="table4" ><label><xref ref-type="table" rid="table4">Table 4</xref></label><caption><title> Comparison between our algorithm and BART for variable selection</title></caption><table><tbody><thead><tr><th align="center" valign="middle" ></th><th align="center" valign="middle" ></th><th align="center" valign="middle" ></th><th align="center" valign="middle"  colspan="4"  >Our Algorithm</th><th align="center" valign="middle"  colspan="4"  >BART</th></tr></thead><tr><td align="center" valign="middle" >Dataset</td><td align="center" valign="middle" >p</td><td align="center" valign="middle" >n</td><td align="center" valign="middle" >Non-interaction</td><td align="center" valign="middle" ></td><td align="center" valign="middle" >Interaction</td><td align="center" valign="middle" ></td><td align="center" valign="middle" >Non-interaction</td><td align="center" valign="middle" ></td><td align="center" valign="middle" >Interaction</td><td align="center" valign="middle" ></td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" ></td><td align="center" valign="middle" >FPR</td><td align="center" valign="middle" >FNR</td><td align="center" valign="middle" >FPR</td><td align="center" valign="middle" >FNR</td><td align="center" valign="middle" >FPR</td><td align="center" valign="middle" >FNR</td><td align="center" valign="middle" >FPR</td><td align="center" valign="middle" >FNR</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >1.0</td><td align="center" valign="middle" >1.0</td><td align="center" valign="middle" >1.0</td><td align="center" valign="middle" >1.0</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >20</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.05</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.05</td><td align="center" valign="middle" >1.0</td><td align="center" valign="middle" >1.0</td><td align="center" valign="middle" >1.0</td><td align="center" valign="middle" >1.0</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.01</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.01</td><td align="center" valign="middle" >1.0</td><td align="center" valign="middle" >1.0</td><td align="center" valign="middle" >1.0</td><td align="center" valign="middle" >1.0</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >150</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >1.0</td><td align="center" valign="middle" >1.0</td><td align="center" valign="middle" >1.0</td><td align="center" valign="middle" >1.0</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >150</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >1.0</td><td align="center" valign="middle" >1.0</td><td align="center" valign="middle" >1.0</td><td align="center" valign="middle" >1.0</td></tr><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >200</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >0.01</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >NA</td><td align="center" valign="middle" >NA</td><td align="center" valign="middle" >NA</td><td align="center" valign="middle" >NA</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >200</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >0.0</td><td align="center" valign="middle" >NA</td><td align="center" valign="middle" >NA</td><td align="center" valign="middle" >NA</td><td align="center" valign="middle" >NA</td></tr></tbody></table></table-wrap><p>as n, BART fails to run while our algorithm provides excellent results in variable selection. Overall, our algorithm performs significantly better than BART in terms of variable selection.</p></sec></sec><sec id="s7"><title>7. Real Data Analysis</title><p>In this section, we demonstrate the performance of our method in two real data sets. We use the Boston housing data and concrete slump test datasets obtained from UCI machine learning repository. Both data have been used extensively in the literature.</p><sec id="s7_1"><title>7.1. Boston Housing Data</title><p>In this section, we used the Boston housing data to compare the performance between BART and our algorithm. The Boston housing data [<xref ref-type="bibr" rid="scirp.78088-ref51">51</xref>] contains information collected by the United States Census Service on the median value of owner occupied homes in Boston, Massachusetts. The data has 506 number of instances with thirteen continuous variables and one binary variable. The data is split into 451 training and 51 test observations. The description for each variable is summarized in <xref ref-type="table" rid="table5">Table 5</xref>.</p><p>MEDV is chosen as the response and the remaining variables are included as predictors. We ran our algorithm for 5000 iterations and the prediction result for both algorithms is shown in <xref ref-type="fig" rid="fig12">Figure 12</xref>.</p><p>Although our algorithm has a comparable prediction error with BART, we argue below that we have a more convincing result in terms of variable selection. We displayed the Inclusion score barplot in <xref ref-type="fig" rid="fig13">Figure 13</xref>.</p><table-wrap id="table5" ><label><xref ref-type="table" rid="table5">Table 5</xref></label><caption><title> Boston housing Dataset variable</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Variables</th><th align="center" valign="middle" >Abbreviation</th><th align="center" valign="middle" >Description</th></tr></thead><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >CRIM</td><td align="center" valign="middle" >Per capita crime rate</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >ZN</td><td align="center" valign="middle" >Proportion of residential land zoned for lots over 25,000 squared feet</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >INDUS</td><td align="center" valign="middle" >Proportion of non-retail business acres per town</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >CHAS</td><td align="center" valign="middle" >Charles River dummy variable (= 1 if tract bounds river; 0 otherwise)</td></tr><tr><td align="center" valign="middle" >5</td><td align="center" valign="middle" >NOX</td><td align="center" valign="middle" >Nitric oxides concentration (parts per 10 million)</td></tr><tr><td align="center" valign="middle" >6</td><td align="center" valign="middle" >RM</td><td align="center" valign="middle" >Average number of rooms per dwelling</td></tr><tr><td align="center" valign="middle" >7</td><td align="center" valign="middle" >AGE</td><td align="center" valign="middle" >Proportion of owner-occupied units built prior to 1940</td></tr><tr><td align="center" valign="middle" >8</td><td align="center" valign="middle" >DIS</td><td align="center" valign="middle" >Weighted distances to five Boston employment centers</td></tr><tr><td align="center" valign="middle" >9</td><td align="center" valign="middle" >RAD</td><td align="center" valign="middle" >Index of accessibility to radial highways</td></tr><tr><td align="center" valign="middle" >10</td><td align="center" valign="middle" >TAX</td><td align="center" valign="middle" >Full-value property-tax rate per $10,000</td></tr><tr><td align="center" valign="middle" >11</td><td align="center" valign="middle" >PTRATIO</td><td align="center" valign="middle" >Pupil-teacher ratio by town</td></tr><tr><td align="center" valign="middle" >12</td><td align="center" valign="middle" >B</td><td align="center" valign="middle" ><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/3-1240904x193.png" xlink:type="simple"/></inline-formula>where <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/3-1240904x194.png" xlink:type="simple"/></inline-formula> is the proportion of blacks by town</td></tr><tr><td align="center" valign="middle" >13</td><td align="center" valign="middle" >LSTAT</td><td align="center" valign="middle" >Percentage of lower status of the population</td></tr><tr><td align="center" valign="middle" >14</td><td align="center" valign="middle" >MEDV</td><td align="center" valign="middle" >Median value of owner-occupied homes in $1000’s</td></tr></tbody></table></table-wrap><p>Based on the histograms, we chose the threshold value equal to 0.1 for easily comparing BART and our algorithm. From the ranking and the chosen threshold value, BART only selected NOX and RM, while our algorithm selected CRIM, ZN, NOX, DIS, B and LSTAT. In order to compare the performance, we looked at Savitsky et al. [<xref ref-type="bibr" rid="scirp.78088-ref25">25</xref>] , which previously analyzed this dataset and selected variables RM, DIS and LSTAT. Clearly, the set of selected variables from our method has more common elements with that of Savitsky et al. [<xref ref-type="bibr" rid="scirp.78088-ref25">25</xref>] .</p></sec><sec id="s7_2"><title>7.2. Concrete Slump Test</title><p>In this section we consider an engineering application to compare our algorithm against BART. The concrete slump test dataset records the test results of two executed tests on concrete to study its behavior [<xref ref-type="bibr" rid="scirp.78088-ref52">52</xref>] [<xref ref-type="bibr" rid="scirp.78088-ref53">53</xref>] .</p><p>The first test is the concrete-slump test, which measures concrete’s plasticity. Since concrete is a composite material with mixture of water, sand, rocks and cement, the first test determines whether the change in ingredients of concrete is consistent. The first test records the change in the slump height and the flow of water. If there is a change in a slump height, the flow must be adjusted to keep the ingredients in concrete homogeneous to satisfy the structure ingenuity. The second test is the “Compressive Strength Test”, which measures the capacity of a concrete to withstand axially directed pushing forces. The second test records the compressive pressure on the concrete.</p><p>The concrete slump test dataset has 103 instances. The data is split into 53 instances for training and 50 instances for testing. There are seven continuous input variables, which are seven ingredients to make concrete, and three outputs, which are slump height, flow height and compressive pressure. Here we only consider the slump height as the output. The description for each variable and output is summarized in <xref ref-type="table" rid="table6">Table 6</xref>.</p><p>The predictive performance is illustrated in <xref ref-type="fig" rid="fig14">Figure 14</xref>.</p><p>Similar to the Boston housing dataset, our algorithm performs closely to BART in prediction. Next, we investigated the performances in terms of variable selection. We plotted the bar-plot of the Inclusion score for each variable in <xref ref-type="fig" rid="fig15">Figure 15</xref>.</p><p>Yurugi et al. [<xref ref-type="bibr" rid="scirp.78088-ref54">54</xref>] determined that coarse aggregation has a significant impact on the plasticity of a concrete. Since the difference in slump’s height is to</p><table-wrap id="table6" ><label><xref ref-type="table" rid="table6">Table 6</xref></label><caption><title> Concrete Slump Test dataset</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Variables</th><th align="center" valign="middle" >Ingredients</th><th align="center" valign="middle" >Unit</th></tr></thead><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >Cement</td><td align="center" valign="middle" >kg</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >Slag</td><td align="center" valign="middle" >kg</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >Fly ash</td><td align="center" valign="middle" >kg</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >Water</td><td align="center" valign="middle" >kg</td></tr><tr><td align="center" valign="middle" >5</td><td align="center" valign="middle" >Super-plasticizer (SP)</td><td align="center" valign="middle" >kg</td></tr><tr><td align="center" valign="middle" >6</td><td align="center" valign="middle" >Coarse Aggregation</td><td align="center" valign="middle" >kg</td></tr><tr><td align="center" valign="middle" >7</td><td align="center" valign="middle" >Fine Aggregation</td><td align="center" valign="middle" >kg</td></tr><tr><td align="center" valign="middle" >8</td><td align="center" valign="middle" >Slump</td><td align="center" valign="middle" >cm</td></tr><tr><td align="center" valign="middle" >9</td><td align="center" valign="middle" >Flow</td><td align="center" valign="middle" >cm</td></tr><tr><td align="center" valign="middle" >10</td><td align="center" valign="middle" >28-day Compressive Strength</td><td align="center" valign="middle" >Mpa</td></tr></tbody></table></table-wrap><p>measure the plasticity of a concrete, coarse aggregation is a critical variable in the concrete slump test. According to <xref ref-type="fig" rid="fig15">Figure 15</xref>, our algorithm selects coarse aggregation as the most important variable unlike BART, which clearly demonstrates the efficacy of our algorithm compared to BART.</p></sec><sec id="s7_3"><title>7.3. Community and Crime Dataset</title><p>In this section we consider a dataset, which has more than 100 predictors to compare our algorithm against BART. Therefore, we chose the Community and Crime dataset. This dataset describes the socio-economic, law enforcement, and crime data in communities of the United States in 1995 [<xref ref-type="bibr" rid="scirp.78088-ref55">55</xref>] [<xref ref-type="bibr" rid="scirp.78088-ref56">56</xref>] .</p><p>In this data, there are about 124 predictors, 5 non-predictors, and 18 response values. The details for each response value can be found at University of California, Irvine (UCI) Machine Learning Database [<xref ref-type="bibr" rid="scirp.78088-ref57">57</xref>] . Since this data has missing values and non-predictors, we preprocessed the data before applying our algorithm and BART on it. After the preprocessing, the number of observations n goes from 2215 to 114 observations, and the number of predictors p becomes 123. Therefore, in this example, we have a case that the number of predictors p is larger than the number of observation n. We split the data into 79 instances for training and 35 instances for testing. We investigated both algorithms’ performance in variable selection. We plotted the histogram of the Inclusion score for each variable in <xref ref-type="fig" rid="fig16">Figure 16</xref>.</p><p>Since BART and our algorithm has different Inclusion score values, we cannot pick the threshold values to identify variables for comparison. Since our algorithm only selects 10 predictors, we decided to rank predictors in BART based on their Inclusion score. Then, we chose BART’s top 10 predictors with highest Inclusion score to compare with ours. <xref ref-type="table" rid="table7">Table 7</xref> lists selected factors affecting violent crime rate based on our algorithm and BART.</p><p>According to Blumstein and Rosenfeld [<xref ref-type="bibr" rid="scirp.78088-ref58">58</xref>] , the crime trend in the United States is contributed by the following factors 1) Economic condition 2) Policing 3) Control of firearms 4) Drugs markets 5) Gangs 6) Socialization and social service 7) Incarceration percent 8) Demographic change. Our selected variables can be grouped into three categories based on above factors. The first category</p><table-wrap id="table7" ><label><xref ref-type="table" rid="table7">Table 7</xref></label><caption><title> Selected variables between our algorithm and BART</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Variables</th><th align="center" valign="middle" >Our Algorithm</th><th align="center" valign="middle" >BART</th></tr></thead><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >% household with social security income</td><td align="center" valign="middle" >% African-American</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >% Mom and kids under labor force</td><td align="center" valign="middle" >income per capita for Asian heritage</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >% immigrants in the last 8 years</td><td align="center" valign="middle" >% employed in manufacturing</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >% immigrants in the last 3 years</td><td align="center" valign="middle" >% kids in two parents family</td></tr><tr><td align="center" valign="middle" >5</td><td align="center" valign="middle" >% housing occupied</td><td align="center" valign="middle" >% of working mom</td></tr><tr><td align="center" valign="middle" >6</td><td align="center" valign="middle" >% vacant housing more than 6 months</td><td align="center" valign="middle" >% kids in unmarried families</td></tr><tr><td align="center" valign="middle" >7</td><td align="center" valign="middle" >Number of housing occupied in upper quantile</td><td align="center" valign="middle" >% immigrants in the last 8 years</td></tr><tr><td align="center" valign="middle" >8</td><td align="center" valign="middle" >Number of sworn full time police officer</td><td align="center" valign="middle" >number of unit house built</td></tr><tr><td align="center" valign="middle" >9</td><td align="center" valign="middle" >Number of sworn police officer in operation</td><td align="center" valign="middle" >number of housing without plumbing facilities</td></tr><tr><td align="center" valign="middle" >10</td><td align="center" valign="middle" >Total request for police per police officer</td><td align="center" valign="middle" >% people living in the same city since 1985</td></tr></tbody></table></table-wrap><p>is economic condition: variables 1, 2, 5, 6, and 7. The second category is demographic change: variables 3 and 4. The third category is policing: variables 8, 9, and 10. Similarly, selected variables in BART can be grouped into three categories. The first category is economic condition: variables 3, 2, 5 and 8. The second category is demographic change: variables 1 and 7. The third category is socialization and social service: variable 6. Based on these grouping, one can see that our selected variables is more agreeable to the study of Blumstein and Rosenfeld [<xref ref-type="bibr" rid="scirp.78088-ref58">58</xref>] than BART.</p></sec></sec><sec id="s8"><title>8. Conclusion</title><p>In this paper, we propose a novel Bayesian nonparametric approach for variable selection and interaction recovery with excellent performance in selection and interaction recovery in both simulated and real datasets. Our method obviates the computation bottleneck in recent unpublished work [<xref ref-type="bibr" rid="scirp.78088-ref28">28</xref>] by proposing a simpler regularization involving a combination of hard and soft shrinkage parameters.</p><p>Although such sparse additive models are well known to adapt to the underlying true dimension of the covariates [<xref ref-type="bibr" rid="scirp.78088-ref26">26</xref>] , literature on consistent selection and interaction recovery in the context of nonparametric regression models is missing. As a future work, we propose to investigate consistency of the variable selection and interaction of our method.</p></sec><sec id="s9"><title>Cite this paper</title><p>Vo, G. and Pati, D. (2017) Sparse Additive Gaussian Process with Soft Interactions. Open Journal of Sta- tistics, 7, 567-588. https://doi.org/10.4236/ojs.2017.74039</p></sec></body><back><ref-list><title>References</title><ref id="scirp.78088-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Manavalan, P. and Johnson, W.C. (1987) Variable Selection Method Improves the Prediction of Protein Secondary Structure from Circular Dichroism Spectra. Analytical Biochemistry, 167, 76-85.</mixed-citation></ref><ref id="scirp.78088-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Saeys, Y., Inza, I. and Larranaga, P. (2007) A Review of Feature Selection Techniques in Bioinformatics. Bioinformatics, 23, 2507-2517.</mixed-citation></ref><ref id="scirp.78088-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Kwon, O.-W., Chan, K., Hao, J. and Lee, T.-W. (2003) Emotion Recognition by Speech Signals. 8th European Conference on Speech Communication and Technology, Geneva, 1-4 September 2003.</mixed-citation></ref><ref id="scirp.78088-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Akaike, H. (1973) Maximum Likelihood Identification of Gaussian Autoregressive Moving Average Models. Biometrika, 60, 255-265.  
https://doi.org/10.1093/biomet/60.2.255</mixed-citation></ref><ref id="scirp.78088-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Schwarz, G., et al. (1978) Estimating the Dimension of a Model. The Annals of Statistics, 6, 461-464. https://doi.org/10.1214/aos/1176344136</mixed-citation></ref><ref id="scirp.78088-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Foster, D.P. and George, E.I. (1994) The Risk Ination Criterion for Multiple Regression. The Annals of Statistics, 22, 1947-1975.</mixed-citation></ref><ref id="scirp.78088-ref7"><label>7</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Tibshirani</surname><given-names> R. </given-names></name>,<etal>et al</etal>. (<year>1996</year>)<article-title>Regression Shrinkage and Selection via the Lasso. Journal of the Royal Statistical Society</article-title><source> Series B (Methodological)</source><volume> 58</volume>,<fpage> 267</fpage>-<lpage>288</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.78088-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Efron, B., Hastie, T., Johnstone, I., Tibshirani, R., et al. (2004) Least Angle Regression. The Annals of Statistics, 32, 407-499.  
https://doi.org/10.1214/009053604000000067</mixed-citation></ref><ref id="scirp.78088-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Fan, J. and Li, R. (2001) Variable Selection via Nonconcave Penalized Likelihood and Its Oracle Properties. JASA, 96, 1348-1360.  
https://doi.org/10.1198/016214501753382273</mixed-citation></ref><ref id="scirp.78088-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Zou, H. and Hastie, T. (2005) Regularization and Variable Selection via the Elastic Net. Journal of the Royal Statistical Society, Series B, 67, 301-320.  
https://doi.org/10.1111/j.1467-9868.2005.00503.x</mixed-citation></ref><ref id="scirp.78088-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Zou, H. (2006) The Adaptive Lasso and Its Oracle Properties. JASA, 101, 1418-1429. https://doi.org/10.1198/016214506000000735</mixed-citation></ref><ref id="scirp.78088-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Zhang, C.-H. (2010) Nearly Unbiased Variable Selection under Minimax Concave Penalty. The Annals of Statistics, 38, 894-942. https://doi.org/10.1214/09-AOS729</mixed-citation></ref><ref id="scirp.78088-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Mitchell, T.J. and Beauchamp, J.J. (1988) Bayesian Variable Selection in Linear Regression. JASA, 83, 1023-1032. https://doi.org/10.1080/01621459.1988.10478694</mixed-citation></ref><ref id="scirp.78088-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">George, E.I. and McCulloch, R.E. (1993) Variable Selection via Gibbs Sampling. Journal of the American Statistical Association, 88, 881-889.  
https://doi.org/10.1080/01621459.1993.10476353</mixed-citation></ref><ref id="scirp.78088-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">George, E.I. and McCulloch, R.E. (1997) Approaches for Bayesian Variable Selection. Statistica sinica, 7, 339-373.</mixed-citation></ref><ref id="scirp.78088-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Laerty, J. and Wasserman, L. (2008) Rodeo: Sparse, Greedy Nonparametric Regression. The Annals of Statistics, 36, 28-63.</mixed-citation></ref><ref id="scirp.78088-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Wahba, G. (1990) Spline Models for Observational Data. Vol. 59, Siam.  
https://doi.org/10.1137/1.9781611970128</mixed-citation></ref><ref id="scirp.78088-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Green, P.J. and Silverman, B.W. (1993) Nonparametric Regression and Generalized Linear Models: A Roughness Penalty Approach. CRC Press, Boca Raton.</mixed-citation></ref><ref id="scirp.78088-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Hastie, T.J. and Tibshirani, R.J. (1990) Generalized Additive Models. Vol. 43, CRC Press, Boca Raton.</mixed-citation></ref><ref id="scirp.78088-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Lin, Y., Zhang, H.H., et al. (2006) Component Selection and Smoothing in Multivariate Nonparametric Regression. The Annals of Statistics, 34, 2272-2297.  
https://doi.org/10.1214/009053606000000722</mixed-citation></ref><ref id="scirp.78088-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Ravikumar, P., Laerty, J., Liu, H. and Wasserman, L. (2009) Sparse Additive Models. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 71, 1009-1030. https://doi.org/10.1111/j.1467-9868.2009.00718.x</mixed-citation></ref><ref id="scirp.78088-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">Radchenko, P. and James, G.M. (2010) Variable Selection Using Adaptive Nonlinear Interaction Structures in High Dimensions. Journal of the American Statistical Association, 105, 1541-1553. https://doi.org/10.1198/jasa.2010.tm10130</mixed-citation></ref><ref id="scirp.78088-ref23"><label>23</label><mixed-citation publication-type="other" xlink:type="simple">Liu, D., Lin, X. and Ghosh, D. (2007) Semiparametric Regression of Multidimensional Genetic Pathway Data: Least-Squares Kernel Machines and Linear Mixed Models. Biometrics, 63, 1079-1088.  
https://doi.org/10.1111/j.1541-0420.2007.00799.x</mixed-citation></ref><ref id="scirp.78088-ref24"><label>24</label><mixed-citation publication-type="other" xlink:type="simple">Zou, F., Huang, H., Lee, S. and Hoeschele, I. (2010) Nonparametric Bayesian Variable Selection with Applications to Multiple Quantitative Trait Loci Mapping with Epistasis and Gene-Environment Interaction. Genetics, 186, 385-394.  
https://doi.org/10.1534/genetics.109.113688</mixed-citation></ref><ref id="scirp.78088-ref25"><label>25</label><mixed-citation publication-type="other" xlink:type="simple">Savitsky, T., Vannucci, M. and Sha, N. (2011) Variable Selection for Nonparametric Gaussian Process Priors: Models and Computational Strategies. Statistical Science: A Review Journal of the Institute of Mathematical Statistics, 26, 130.</mixed-citation></ref><ref id="scirp.78088-ref26"><label>26</label><mixed-citation publication-type="other" xlink:type="simple">Yang, Y., Tokdar, S.T., et al. (2015) Minimax-Optimal Nonparametric Regression in High Dimensions. The Annals of Statistics, 43, 652-674.  
https://doi.org/10.1214/14-AOS1289</mixed-citation></ref><ref id="scirp.78088-ref27"><label>27</label><mixed-citation publication-type="other" xlink:type="simple">Fang, Z., Kim, I. and Schaumont, P. (2012) Flexible Variable Selection for Recovering Sparsity in Nonadditive Nonparametric Models.</mixed-citation></ref><ref id="scirp.78088-ref28"><label>28</label><mixed-citation publication-type="other" xlink:type="simple">Qamar, S. and Tokdar, S.T. (2014) Additive Gaussian Process Regression.</mixed-citation></ref><ref id="scirp.78088-ref29"><label>29</label><mixed-citation publication-type="other" xlink:type="simple">Rasmussen, C.E. and Williams, C.K.I. (2006) Gaussian Processes for Machine Learning.</mixed-citation></ref><ref id="scirp.78088-ref30"><label>30</label><mixed-citation publication-type="other" xlink:type="simple">Van der Vaart, A.W. and van Zanten, J.H. (2009) Adaptive Bayesian Estimation Using a Gaussian Random Field with Inverse Gamma Bandwidth. The Annals of Statistics, 37, 2655-2675.</mixed-citation></ref><ref id="scirp.78088-ref31"><label>31</label><mixed-citation publication-type="other" xlink:type="simple">Bhattacharya, A., Pati, D. and Dunson, D.B. (2014) Anisotropic Function Estimation Using Multi-Bandwidth Gaussian Processes. The Annals of Statistics, 42, 352-381. https://doi.org/10.1214/13-AOS1192</mixed-citation></ref><ref id="scirp.78088-ref32"><label>32</label><mixed-citation publication-type="other" xlink:type="simple">Tibshirani, R., et al. (1997) The Lasso Method for Variable Selection in the Cox Model. Statistics in Medicine, 16, 385-395.  
https://doi.org/10.1002/(SICI)1097-0258(19970228)16:4&lt;385::AID-SIM380&gt;3.0.CO;2-3</mixed-citation></ref><ref id="scirp.78088-ref33"><label>33</label><mixed-citation publication-type="other" xlink:type="simple">Hastie, T., Tibshirani, R., Friedman, J. and Franklin, J. (2005) The Elements of Statistical Learning: Data Mining, Inference and Prediction. The Mathematical Intelligencer, 27, 83-85. https://doi.org/10.1007/BF02985802</mixed-citation></ref><ref id="scirp.78088-ref34"><label>34</label><mixed-citation publication-type="other" xlink:type="simple">Park, T. and Casella, G. (2008) The Bayesian Lasso. Journal of the American Statistical Association, 103, 681-686. https://doi.org/10.1198/016214508000000337</mixed-citation></ref><ref id="scirp.78088-ref35"><label>35</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Tipping</surname><given-names> M.E. </given-names></name>,<etal>et al</etal>. (<year>2001</year>)<article-title>Sparse Bayesian Learning and the Relevance Vector Machine</article-title><source> The Journal of Machine Learning Research</source><volume> 1</volume>,<fpage> 211</fpage>-<lpage>244</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.78088-ref36"><label>36</label><mixed-citation publication-type="other" xlink:type="simple">Grin, J.E. and Brown, P.J. (2010) Inference with Normal-Gamma Prior Distributions in Regression Problems. Bayesian Analysis, 5, 171-188.  
https://doi.org/10.1214/10-BA507</mixed-citation></ref><ref id="scirp.78088-ref37"><label>37</label><mixed-citation publication-type="other" xlink:type="simple">Carvalho, C.M., Polson, N.G. and Scott, J.G. (2010) The Horseshoe Estimator for Sparse Signals. Biometrika, 97, 465-480. https://doi.org/10.1093/biomet/asq017</mixed-citation></ref><ref id="scirp.78088-ref38"><label>38</label><mixed-citation publication-type="other" xlink:type="simple">Carvalho, C.M., Polson, N.G. and Scott, J.G. (2009) Handling Sparsity via the Horseshoe. International Conference on Artificial Intelligence and Statistics, Clearwater, 16-18 April 2009, 73-80.</mixed-citation></ref><ref id="scirp.78088-ref39"><label>39</label><mixed-citation publication-type="other" xlink:type="simple">Bhattacharya, A., Pati, D., Pillai, N.S. and Dunson, D.B. (2014) Dirichlet-Laplace Priors for Optimal Shrinkage. Journal of the American Statistical Association, 110, 1479-1490.</mixed-citation></ref><ref id="scirp.78088-ref40"><label>40</label><mixed-citation publication-type="other" xlink:type="simple">Polson, N.G. and Scott, J.G. (2010) Shrink Globally, Act Locally: Sparse Bayesian Regularization and Prediction. Bayesian Statistics, 9, 501-538.</mixed-citation></ref><ref id="scirp.78088-ref41"><label>41</label><mixed-citation publication-type="other" xlink:type="simple">Diebolt, J., Ip, E. and Olkin, I. (1994) A Stochastic EM Algorithm for Approximating the Maximum Likelihood Estimate. Technical Report 301, Department of Statistics, Stanford University, Stanford.</mixed-citation></ref><ref id="scirp.78088-ref42"><label>42</label><mixed-citation publication-type="other" xlink:type="simple">Meng, X.-L. and Rubin, D.B. (1994) On the Global and Component Wise Rates of Convergence of the EM Algorithm. Linear Algebra and Its Applications, 199, 413-425.</mixed-citation></ref><ref id="scirp.78088-ref43"><label>43</label><mixed-citation publication-type="other" xlink:type="simple">Hastings, W. (1970) Monte Carlo Sampling Methods Using Markov Chains and Their Applications. Biometrika, 57, 97-109. https://doi.org/10.1093/biomet/57.1.97</mixed-citation></ref><ref id="scirp.78088-ref44"><label>44</label><mixed-citation publication-type="other" xlink:type="simple">Bishop, C.M. (2006) Pattern Recognition and Machine Learning. Springer, Berlin.</mixed-citation></ref><ref id="scirp.78088-ref45"><label>45</label><mixed-citation publication-type="other" xlink:type="simple">Han, J., Kamber, M. and Pei, J. (2011) Data Mining: Concepts and Techniques: Concepts and Techniques. Elsevier, Amsterdam.</mixed-citation></ref><ref id="scirp.78088-ref46"><label>46</label><mixed-citation publication-type="other" xlink:type="simple">Guyon, I. and Elisseeff, A. (2003) An Introduction to Variable and Feature Selection. Journal of Machine Learning Research, 3, 1157-1182.</mixed-citation></ref><ref id="scirp.78088-ref47"><label>47</label><mixed-citation publication-type="other" xlink:type="simple">George Forman (2003) An Extensive Empirical Study of Feature Selection Metrics for Text Classification. Journal of Machine Learning Research, 3, 1289-1305.</mixed-citation></ref><ref id="scirp.78088-ref48"><label>48</label><mixed-citation publication-type="other" xlink:type="simple">Stoppiglia, H., Dreyfus, G., Dubois, R. and Oussar, Y. (2003) Ranking a Random Feature for Variable and Feature Selection. Journal of Machine Learning Research, 3, 1399-1414.</mixed-citation></ref><ref id="scirp.78088-ref49"><label>49</label><mixed-citation publication-type="other" xlink:type="simple">Genuer, R., Poggi, J.M. and Tuleau-Malot, C. (2010) Variable Selection Using Random Forests. Pattern Recognition Letters, 31, 2225-2236.</mixed-citation></ref><ref id="scirp.78088-ref50"><label>50</label><mixed-citation publication-type="other" xlink:type="simple">Chipman, H.A., George, E.I. and McCulloch, R.E. (2010) Bart: Bayesian Additive Regression Trees. The Annals of Applied Statistics, 4, 266-298.</mixed-citation></ref><ref id="scirp.78088-ref51"><label>51</label><mixed-citation publication-type="other" xlink:type="simple">Harrison, D. and Rubinfeld, D.L. (1978) Hedonic Housing Prices and the Demand for Clean Air. Journal of Environmental Economics and Management, 5, 81-102.</mixed-citation></ref><ref id="scirp.78088-ref52"><label>52</label><mixed-citation publication-type="other" xlink:type="simple">Yeh, I., et al. (2008) Modeling Slump of Concrete with Yash and Superplasticizer. Computers and Concrete, 5, 559-572. https://doi.org/10.12989/cac.2008.5.6.559</mixed-citation></ref><ref id="scirp.78088-ref53"><label>53</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Yeh</surname><given-names> I.C. </given-names></name>,<etal>et al</etal>. (<year>2007</year>)<article-title>Modeling Slumpow of Concrete Using Second-Order Regressions and Artificial Neural Networks</article-title><source> Cement and Concrete Composites</source><volume> 29</volume>,<fpage> 474</fpage>-<lpage>480</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.78088-ref54"><label>54</label><mixed-citation publication-type="other" xlink:type="simple">Yurugi, M., Sakata, N., Iwai, M. and Sakai, G. (1993) Mix Proportion for Highly Workable Concrete. Proceedings of Concrete, Dundee, 7-9 September 1993, 579-589.</mixed-citation></ref><ref id="scirp.78088-ref55"><label>55</label><mixed-citation publication-type="other" xlink:type="simple">Redmond, M.A. and Highley, T. (2010) Empirical Analysis of Caseediting Approaches for Numeric Prediction. In: Innovations in Computing Sciences and Software Engineering, Springer, Berlin, 79-84.</mixed-citation></ref><ref id="scirp.78088-ref56"><label>56</label><mixed-citation publication-type="other" xlink:type="simple">Buczak, A.L. and Gifford, C.M. (2010) Fuzzy Association Rule Mining for Community Crime Pattern Discovery. In: ACM SIGKDD Workshop on Intelligence and Security Informatics, ACM, New York, 2.</mixed-citation></ref><ref id="scirp.78088-ref57"><label>57</label><mixed-citation publication-type="other" xlink:type="simple">Blake, C. and Merz, C.J. (1998) Repository of Machine Learning Databases.</mixed-citation></ref><ref id="scirp.78088-ref58"><label>58</label><mixed-citation publication-type="other" xlink:type="simple">Blumstein, A. and Rosenfeld, R. (2008) Factors Contributing to Us Crime Trends. In: Understanding Crime Trends: Workshop Report, The National Academies Press, Washington DC, 13-43.</mixed-citation></ref></ref-list></back></article>