<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OJS</journal-id><journal-title-group><journal-title>Open Journal of Statistics</journal-title></journal-title-group><issn pub-type="epub">2161-718X</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/ojs.2015.55040</article-id><article-id pub-id-type="publisher-id">OJS-58573</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Small Sample Behaviors of the Delete-&lt;i&gt;d&lt;/i&gt; Cross Validation Statistic
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>ude</surname><given-names>H. Kastens</given-names></name><xref ref-type="aff" rid="aff1"><sub>1</sub></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><label>1</label><addr-line>Kansas Biological Survey (KBS), University of Kansas, Higuchi Hall, Lawrence, USA</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>kastens@ku.edu</email></corresp></author-notes><pub-date pub-type="epub"><day>23</day><month>07</month><year>2015</year></pub-date><volume>05</volume><issue>05</issue><fpage>382</fpage><lpage>392</lpage><history><date date-type="received"><day>31</day>	<month>March</month>	<year>2015</year></date><date date-type="rev-recd"><day>accepted</day>	<month>2</month>	<year>August</year>	</date><date date-type="accepted"><day>5</day>	<month>August</month>	<year>2015</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Built upon an iterative process of resampling without replacement and out-of-sample prediction, the delete-d cross validation statistic CV(
  <em>d</em>) provides a robust estimate of forecast error variance. To compute CV(
  <em>d</em>), a dataset consisting of n observations of predictor and response values is systematically and repeatedly partitioned (split) into subsets of size 
  <em>n</em> – 
  <em>d</em> (used for model training) and 
  <em>d</em> (used for model testing). Two aspects of CV(
  <em>d</em>) are explored in this paper. First, estimates for the unknown expected value E[CV(
  <em>d</em>)] are simulated in an OLS linear regression setting. Results suggest general formulas for E[CV(
  <em>d</em>)] dependent on σ
  <sup>2</sup> (“true” model error variance), 
  <em>n</em> – 
  <em>d</em> (training set size), and 
  <em>p</em> (number of predictors in the model). The conjectured E[CV(
  <em>d</em>)] formulas are connected back to theory and generalized. The formulas break down at the two largest allowable 
  <em>d</em> values (
  <em>d</em> = 
  <em>n</em> – 
  <em>p</em> – 1 and 
  <em>d</em> = 
  <em>n</em> – 
  <em>p</em>, the 1 and 0 degrees of freedom cases), and numerical instabilities are observed at these points. An explanation for this distinct behavior remains an open question. For the second analysis, simulation is used to demonstrate how the previously established asymptotic conditions {
  <em>d</em>/
  <em>n</em> → 1 and 
  <em>n</em> – 
  <em>d</em> → ∞ as 
  <em>n</em> → ∞} required for optimal linear model selection using CV(
  <em>d</em>) for model ranking are manifested in the smallest sample setting, using either independent or correlated candidate predictors.
 
</p></abstract><kwd-group><kwd>Expected Value</kwd><kwd> Forecast Error Variance</kwd><kwd> Linear Regression</kwd><kwd> Model Selection</kwd><kwd> Simulation</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Cross validation (CV) is a model evaluation technique that utilizes data splitting. To describe CV, suppose that each data observation consists of a response value (the dependent variable) and corresponding predictor values (the independent variables) that will be used in some specified model form for the response. The data observations are split (partitioned) into two subsets. One subset (the training set) is used for model parameter estimation. Using these parameter values, the model is then applied to the other subset (the testing set). The model predictions determined for the testing set observations are compared to their corresponding actual response values, and an “out-of-sample” mean squared error is computed. For delete-d cross validation, all possible data splits with testing sets that contain d observations are evaluated. The aggregated mean squared error statistic that results from this computational effort is denoted CV(d).</p><p>To define the CV(d) statistic used in ordinary least squares (OLS) linear regression, let p &lt; n be positive integers, and let I<sub>k</sub> denote the k-by-k identity matrix. Let<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x5.png" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x6.png" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x7.png" xlink:type="simple"/></inline-formula>, and<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x8.png" xlink:type="simple"/></inline-formula>, where β and ε are unknown and<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x9.png" xlink:type="simple"/></inline-formula>. Also, let β and ε be such that Y = Xβ + ε is the “true” optimal (minimum σ<sup>2</sup>) linear statistical model for predicting Y using X. Each row of the matrix [X Y] corresponds to a data observation (p predictors and one response), and each column of X corresponds to a particular predictor. Assume that each p-row submatrix of X has full rank, a necessary condition for CV(d) to be computable for all <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x10.png" xlink:type="simple"/></inline-formula>. This is a reasonable assumption when the data observations are independent. Let S be an arbitrary subset of<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x11.png" xlink:type="simple"/></inline-formula>, and let S<sup>C</sup> = N\S. Let X<sub>S</sub> denote the row subset of X indexed by S, and define</p><p><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x12.png" xlink:type="simple"/></inline-formula>to be the OLS parameter vector estimated using X<sub>S</sub>. Define<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x13.png" xlink:type="simple"/></inline-formula>.</p><p>Let <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x14.png" xlink:type="simple"/></inline-formula> denote the l<sup>2</sup> norm and let <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x15.png" xlink:type="simple"/></inline-formula> denote set cardinality. Then the delete-d cross validation statistic is given by</p><disp-formula id="scirp.58573-formula430"><label>(1)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-1240497x16.png"  xlink:type="simple"/></disp-formula><p>This equation can be found in [<xref ref-type="bibr" rid="scirp.58573-ref1">1</xref>] , [<xref ref-type="bibr" rid="scirp.58573-ref2">2</xref>] (p. 255), and [<xref ref-type="bibr" rid="scirp.58573-ref3">3</xref>] (p. 405). The form of (1) is that of a mean squared error, because there are d observations in each testing set S and there are n-choose-d unique subsets S such that<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x17.png" xlink:type="simple"/></inline-formula>.</p></sec><sec id="s2"><title>2. Background for CV(d)</title><p>Three popular papers provided some of the early groundwork for cross validation. Allen [<xref ref-type="bibr" rid="scirp.58573-ref4">4</xref>] introduced the prediction sum of squares (PRESS) statistic, which involves sequential prediction of single observations using models estimated from the full data absent the data point to be predicted. Stone [<xref ref-type="bibr" rid="scirp.58573-ref5">5</xref>] examined the use of delete-1 cross validation methods for regression coefficient “shrinker” estimation. Geisser [<xref ref-type="bibr" rid="scirp.58573-ref6">6</xref>] presented one of the first introductions of a multiple observation holdout sample reuse method similar to delete-d cross validation. One of the first major practical implementations of CV appeared in [<xref ref-type="bibr" rid="scirp.58573-ref7">7</xref>] , where “V-fold cross validation” is offered as a way to estimate model accuracy during optimization of classification and regression tree models.</p><p>Numerous authors have discussed and examined the properties of CV(1) specifically in the context of model selection (e.g., [<xref ref-type="bibr" rid="scirp.58573-ref2">2</xref>] [<xref ref-type="bibr" rid="scirp.58573-ref8">8</xref>] [<xref ref-type="bibr" rid="scirp.58573-ref9">9</xref>] ). These studies and others have established that in spite of the merits of using CV(1), this method does not always perform well in optimal model identification studies when compared to other direct methods such as information criteria (e.g., [<xref ref-type="bibr" rid="scirp.58573-ref2">2</xref>] ). The consensus is that CV(1) has a tendency in many situations to select overly complex models; i.e., it does not sufficiently penalize for over fitting [<xref ref-type="bibr" rid="scirp.58573-ref10">10</xref>] (p.303). Many researchers have examined CV(d) for one or more d values for actual and simulated case studies involving model selection (e.g., [<xref ref-type="bibr" rid="scirp.58573-ref1">1</xref>] [<xref ref-type="bibr" rid="scirp.58573-ref2">2</xref>] [<xref ref-type="bibr" rid="scirp.58573-ref11">11</xref>] ), but not to the extent of exposing any general, finite-sample statistical tendencies of CV(d) as a function of d.</p><p>Asymptotic equivalence of CV(1) to the delete-1 jackknife, the standard bootstrap, and other model selection statistics such as Mallows’s C<sub>p</sub> [<xref ref-type="bibr" rid="scirp.58573-ref12">12</xref>] and the Akaike information criteria (AIC) [<xref ref-type="bibr" rid="scirp.58573-ref13">13</xref>] has been established (see [<xref ref-type="bibr" rid="scirp.58573-ref11">11</xref>] &amp; [<xref ref-type="bibr" rid="scirp.58573-ref14">14</xref>] and references therein). By defining d to increase at a rate d/n → a &lt; 1, Zhang [<xref ref-type="bibr" rid="scirp.58573-ref1">1</xref>] shows that CV(d) and a particular form of the mean squared prediction error ( [<xref ref-type="bibr" rid="scirp.58573-ref15">15</xref>] ; this is a generalization of the “final prediction error” of [<xref ref-type="bibr" rid="scirp.58573-ref16">16</xref>] ) are asymptotically equivalent under certain constraints. Arguably the most compelling result can be found in [<xref ref-type="bibr" rid="scirp.58573-ref11">11</xref>] , where it is shown that if d is selected such that d/n → 1 and n ? d → ∞ as n → ∞, these conditions are necessary and sufficient to ensure consistency of using CV(d) for optimal linear model selection under certain general conditions. The proof depends on a formula for E[CV(d)] that applies to the X-fixed case, and the author (Shao) briefly describes the additional constraints (which are unrelated to the E[CV(d)] formula) necessary for almost sure application of the result to the X-random case.</p><p>Unfortunately, none of these asymptotic theoretical developments provide practitioners with specific guidance or information helpful for making a judicious choice for d in an arbitrary small sample setting, for either forecast error variance estimation or model ranking for model selection. For this study, two questions are addressed regarding CV(d) that are relevant to the small sample setting. First, expressions are developed for E[CV(d)] for the X-random case using simulation, which are linked back to theory and generalized. Second, a model selection simulation is used to illustrate how Shao’s conditions {d/n → 1 and n ? d → ∞ as n → ∞} are manifested in the smallest sample setting.</p></sec><sec id="s3"><title>3. Expected Value for CV(d)</title><sec id="s3_1"><title>3.1. Problem Background</title><p>For the case with X-fixed, Shao &amp; Tu [<xref ref-type="bibr" rid="scirp.58573-ref17">17</xref>] (p. 309) indicate that</p><disp-formula id="scirp.58573-formula431"><label>(2)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-1240497x18.png"  xlink:type="simple"/></disp-formula><p>This expression implies that CV(d) provides an estimate for<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x19.png" xlink:type="simple"/></inline-formula>, where <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x20.png" xlink:type="simple"/></inline-formula> refers to the squared prediction error when making a prediction for a future observation at a design point (row of predictor matrix X) and X contains n?d independent observations (rows).</p><p>A theorem introduced in this section establishes that (2) applies to the case of the mean (intercept) model. As previously noted, we are interested in the X-random case. Based on an extremely tight correspondence with simulation results, expressions are conjectured for E[CV(d)] when the linear model contains at least one random valued predictor, for cases with and without an intercept. Building from work described in Miller [<xref ref-type="bibr" rid="scirp.58573-ref9">9</xref>] , the conjectured formulas are linked back to theory, allowing them to be generalized beyond the constraints of the simulation.</p></sec><sec id="s3_2"><title>3.2. Results</title><p>Begin with the simplest case, where the only predictor is the intercept. Suppose X = 1<sup>n</sup><sup>&#215;</sup><sup>1</sup> (an n-vector of ones), so that the linear regression model under investigation is the mean model (so called because<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x21.png" xlink:type="simple"/></inline-formula>). The expected value for (1) under the mean model is given in</p><p>THEOREM 1: Suppose X = 1<sup>n</sup><sup>&#215;1</sup>, and let <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x22.png" xlink:type="simple"/></inline-formula> be such that the<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x22.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x23.png" xlink:type="simple"/></inline-formula>. Then, for<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x22.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x23.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x24.png" xlink:type="simple"/></inline-formula>, the expected value for CV(d) is given by</p><disp-formula id="scirp.58573-formula432"><label>(3)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-1240497x25.png"  xlink:type="simple"/></disp-formula><p>Proof: Define<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x26.png" xlink:type="simple"/></inline-formula>, where<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x26.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x27.png" xlink:type="simple"/></inline-formula>. Then, for<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x26.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x27.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x28.png" xlink:type="simple"/></inline-formula>, we will show that the expected value for a single summand term of (1) is given by</p><p><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x29.png" xlink:type="simple"/></inline-formula>.</p><p>The d-by-1 vector <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x30.png" xlink:type="simple"/></inline-formula> has components of the form<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x30.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x31.png" xlink:type="simple"/></inline-formula>, where y<sub>j</sub> is a “deleted” observation (entry in Y<sub>S</sub>) and <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x30.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x31.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x32.png" xlink:type="simple"/></inline-formula> is the sample mean of (n ? d) Y-values in <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x30.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x31.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x32.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x33.png" xlink:type="simple"/></inline-formula> that were not deleted. Since y<sub>j</sub> and <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x30.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x31.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x32.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x33.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x34.png" xlink:type="simple"/></inline-formula> are statistically independent and have the same expected value (m), we have</p><p><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x35.png" xlink:type="simple"/></inline-formula>.</p><p>Since y<sub>j</sub> is an arbitrary element of holdout set Y<sub>S</sub>, we have</p><p><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x36.png" xlink:type="simple"/></inline-formula>.</p><p>Because of the linearity of<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x37.png" xlink:type="simple"/></inline-formula>, (3) immediately follows from this derivation, which applies to an arbitrary split of the dataset. QED</p><p>Different results appear when simulating models that include at least one random valued predictor. Using the random number generator in MATLAB&#174; to simulate data sets {X,Y}, values for CV(d) were simulated for numerous cases with <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x38.png" xlink:type="simple"/></inline-formula> and<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x38.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x39.png" xlink:type="simple"/></inline-formula>. Y-values were simulated using Y<sub>j</sub> = X<sub>j</sub>β + ε<sub>j</sub>, with error<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x38.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x39.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x40.png" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x38.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x39.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x40.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x41.png" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x38.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x39.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x40.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x41.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x42.png" xlink:type="simple"/></inline-formula>, and<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x38.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x39.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x40.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x41.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x42.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x43.png" xlink:type="simple"/></inline-formula>. For each simulated data set, for<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x38.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x39.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x40.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x41.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x42.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x43.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x44.png" xlink:type="simple"/></inline-formula>, coefficient values <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x38.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x39.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x40.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x41.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x42.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x43.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x44.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x45.png" xlink:type="simple"/></inline-formula> were IID Bernoulli with<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x38.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x39.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x40.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x41.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x42.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x43.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x44.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x45.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x46.png" xlink:type="simple"/></inline-formula>. For each simulated data set, for<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x38.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x39.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x40.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x41.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x42.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x43.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x44.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x45.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x46.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x47.png" xlink:type="simple"/></inline-formula>, deviation values σ = 10<sup>α</sup> and <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x38.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x39.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x40.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x41.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x42.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x43.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x44.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x45.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x46.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x47.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x48.png" xlink:type="simple"/></inline-formula> were IID such that α, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x38.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x39.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x40.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x41.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x42.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x43.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x44.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x45.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x46.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x47.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x48.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x49.png" xlink:type="simple"/></inline-formula>(to help limit rounding errors, α and α<sub>k</sub> values outside the interval [?3,3] were snapped to the appropriate interval endpoint). Simulations using an intercept were also examined, in which case the first column of X was populated with ones rather than random values.<sub> </sub></p><p>For a particular (n,p), after simulating at least 20,000 CV(d) values for each possible<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x50.png" xlink:type="simple"/></inline-formula>, mean simulated CV(d) values were computed to provide simulated E[CV(d)]. Upon inspection, simulated E[CV(d)]/σ<sup>2</sup> were found to follow rational number sequences clear enough to conjecture general formulas for E[CV(d)] dependent on n ? d, p, and σ<sup>2</sup>.</p><p>An apparently related outcome is the identification of a two-point region of numerical instability of the E[CV(d)] error curve for any tested model that includes a random valued predictor. Specifically, simulation results reveal two points of increasing instability at <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x51.png" xlink:type="simple"/></inline-formula> and E[CV(d<sub>max</sub>)], where d<sub>max</sub> = n ? p. The term “increasing instability” is apt because the coefficient of variation (=standard deviation/mean) calculated for the simulated CV(d) values is stable for d &lt; d<sub>max</sub> ? 1, but increasingly blows up (along with E[CV(d)]) at d = d<sub>max</sub> ? 1 and d = d<sub>max</sub> (results not shown). The reason for this phenomenon is, at present, an interesting open question. This exceptional behavior is not incompatible with the conjectured formulas for E[CV(d)] because the formulas break down at the two largest allowable d values.</p><p>To gauge the accuracy of the conjectured formulas for E[CV(d)], the author used an absolute percent error statistic defined by</p><disp-formula id="scirp.58573-formula433"><label>(4)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-1240497x52.png"  xlink:type="simple"/></disp-formula><p>To provide a simple gauge for rounding error magnitude, values were simulated for the expected error of regression, given by</p><disp-formula id="scirp.58573-formula434"><label>(5)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-1240497x53.png"  xlink:type="simple"/></disp-formula><p>REG has the property that E[REG] = σ<sup>2</sup>.</p><p>Two findings are notable: (a) distinct but related patterns for E[CV(d)] emerge when considering linear models consisting entirely of random valued predictors and those that use an intercept; and (b) two points of increasing instability appear at <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x54.png" xlink:type="simple"/></inline-formula> and E[CV(d<sub>max</sub>)]. The existence of the two-point instability appears to be robust to increasing dimensionality. Result (a) is expressed in Conjectures 1 and 2, which do not conflict with the exceptional behavior noted in (b). For Conjectures 1 and 2, suppose that<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x54.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x55.png" xlink:type="simple"/></inline-formula>, with β the assumed “true” linear model coefficient vector.</p><p>CONJECTURE 1: Let X be the n-by-p design matrix where p &lt; n ? 2 and the predictors in X are multivariate normal. Then, for<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x56.png" xlink:type="simple"/></inline-formula>, the expected value for CV(d) is given by</p><disp-formula id="scirp.58573-formula435"><label>. (6)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-1240497x57.png"  xlink:type="simple"/></disp-formula><p>In the search for this equation, the author scrutinized simulated values for E[CV(d)], examining a variety of cases. Using the approximation in (2) as a starting point for exploring possible forms for the RHS of (6), the author eventually arrived at (6) through trial and error.</p><p>At <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x58.png" xlink:type="simple"/></inline-formula> (the first point of the two-point instability in the E[CV(d)] error curve), (6) has a singularity. At d = d<sub>max</sub>, (6) takes the nonsensical value of<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x58.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x59.png" xlink:type="simple"/></inline-formula>. Though (6) and (2) are similar, the inclusion of “?p ? 1” in the denominator of the dilation factor in (6) presents an obvious disagreement that becomes increasingly substantial as p increases. For example, the largest value that (2) can achieve is 2σ<sup>2</sup>, realized at d = d<sub>max</sub>. Compare this to the maximum E[CV(d)] value<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x58.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x59.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x60.png" xlink:type="simple"/></inline-formula>, realized by (6) at<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x58.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x59.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x60.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x61.png" xlink:type="simple"/></inline-formula>.</p><p><xref ref-type="fig" rid="fig1">Figure 1</xref> shows results for the case<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x62.png" xlink:type="simple"/></inline-formula>, with a single random valued predictor comprising the design matrix X. The simulated E[CV(d)] error curve is displayed along with corresponding predicted E[CV(d)] error curves obtained using (6) and (2), so that all three error curves can be examined simultaneously. Note the congruity between simulated E[CV(d)] and predicted E[CV(d)] from Conjecture 1, and the widening (with d)</p><fig id="fig1"  position="float"><label><xref ref-type="fig" rid="fig1">Figure 1</xref></label><caption><title> A comparison between simulated and predicted E[CV(d)] error curves using a linear model with a single random valued predictor, for sample size n = 10. Note the good correspondence between the simulated E[CV(d)] values and the predicted values from Conjecture 1. Also visible is the two-point instability, with simulated E[CV(d)] blowing up at d = d<sub>max</sub> ? 1 = 8 and d = d<sub>max</sub> = 9</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/4-1240497x63.png"/></fig><p>disparity between simulated E[CV(d)] and predicted E[CV(d)] from the approximation provided in (2). Also note the blowup in simulated E[CV(d)] at the two largest d values, reflecting the previously described two-point instability of the E[CV(d)] error curve when at least one random valued predictor is used in the model.</p><p><xref ref-type="fig" rid="fig2">Figure 2</xref>(a) shows results for the case<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x64.png" xlink:type="simple"/></inline-formula>, using a model with two random valued predictors. <xref ref-type="fig" rid="fig3">Figure 3</xref>(a) shows results for the case with<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x64.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x65.png" xlink:type="simple"/></inline-formula>, using a model with eight random valued predictors. The same observations noted above for <xref ref-type="fig" rid="fig1">Figure 1</xref> apply to <xref ref-type="fig" rid="fig2">Figure 2</xref>(a) and <xref ref-type="fig" rid="fig3">Figure 3</xref>(a). Now consider the case where an intercept is included in the linear model.</p><p>CONJECTURE 2: Let X be the n-by-p design matrix where p &lt; n ? 2, the first column of X is an intercept, and the other predictors in X are multivariate normal. Then, for<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x66.png" xlink:type="simple"/></inline-formula>, the expected value for CV(d) is given by</p><disp-formula id="scirp.58573-formula436"><label>. (7)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-1240497x67.png"  xlink:type="simple"/></disp-formula><p>In the search for this equation, the author once again scrutinized simulated values for E[CV(d)], examining a variety of cases. This time, (6) was used as a starting point for exploring possible forms for the RHS of (7). Specifically, the author reasoned that substitution of an intercept for a random valued predictor reduces model complexity, suggesting that the E[CV(d)] expression for models that include an intercept might take the form of a dampened version of (6). Indeed, after much trial and error, this was found to be the case once the RHS of (7) was “discovered”.</p><p>Like Equation (6), at <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x68.png" xlink:type="simple"/></inline-formula> and<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x68.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x69.png" xlink:type="simple"/></inline-formula>, (7) has a singularity and a non-positive (and thus nonsensical) value, respectively. Note that (7) constitutes a downward adjustment of (6). Apparently, the 1/(n?d) term provides an adjustment for the reduced model complexity when substituting an intercept for a random-valued predictor. The 2/p term dampens the adjustment as p gets larger and the general effect of this substitution on the model becomes less pronounced.</p><p><xref ref-type="fig" rid="fig2">Figure 2</xref>(b) shows results for the case<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x70.png" xlink:type="simple"/></inline-formula>, using a model with an intercept and one random valued predictor. <xref ref-type="fig" rid="fig3">Figure 3</xref>(b) shows results for the case<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x70.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x71.png" xlink:type="simple"/></inline-formula>, using a model with an intercept and seven random valued predictors. The same general observations noted above for <xref ref-type="fig" rid="fig1">Figure 1</xref> apply to <xref ref-type="fig" rid="fig2">Figure 2</xref>(b) and <xref ref-type="fig" rid="fig3">Figure 3</xref>(b).</p><fig id="fig2"  position="float"><label><xref ref-type="fig" rid="fig2">Figure 2</xref></label><caption><title> A comparison between simulated and predicted E[CV(d)] error curves using a linear model with (a) two random valued predictors, and (b) an intercept and one random valued predictor, for sample size n = 10. Note the good correspondence between the simulated E[CV(d)] values and the predicted values from Conjectures 1 and 2. Also visible is the two-point instability, with simulated E[CV(d)] blowing up at d = d<sub>max</sub> ? 1 = 7 and d = d<sub>max</sub> = 8</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/4-1240497x72.png"/></fig><fig id="fig3"  position="float"><label><xref ref-type="fig" rid="fig3">Figure 3</xref></label><caption><title> A comparison between simulated and predicted E[CV(d)] error curves using the linear model with (a) eight random valued predictors, and (b) an intercept and seven random valued predictors, for sample size n = 20. Note the good correspondence between the simulated E[CV(d)] values and the predicted values from Conjectures 1 and 2. Also visible is the two-point instability, with simulated E[CV(d)] blowing up at d = d<sub>max</sub> ? 1 = 11 and d = d<sub>max</sub> = 12</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/4-1240497x73.png"/></fig><p>These simulation results using independent normal predictors and errors provide strong evidence for the validity of Conjectures 1 and 2. Graphical evidence for this assertion can be seen in Figures 1-3. APE values (4) computed comparing (6) and (7) to corresponding simulated E[CV(d)] were generally O(10<sup>−2</sup>) to O(10<sup>−1</sup>). To provide a gauge for these error magnitudes, E[REG] values (5) were simulated and compared to the known value of σ<sup>2</sup>. APE values from this comparison also were generally O(10<sup>−2</sup>) to O(10<sup>−1</sup>), indicating that rounding error was solely responsible for the slight differences observed between simulated E[CV(d)] and predicted E[CV(d)] from (6) and (7). To justify the more general assumptions in the conjectures using multivariate normal predictors and second-order errors, we use the following theoretical connection.</p></sec><sec id="s3_3"><title>3.3. Connecting Simulation Results Back to Theory</title><p>Equation (2) was examined because it was the only explicitly stated estimate for E[CV(d)] found in the literature. This expression gives the expected mean squared error of prediction (MSEP) for using a linear regression model to make a prediction for some future observation at a design point. However, (2) provides an inaccurate characterization for CV(d) in any arbitrary small sample setting where there are substantially more possible design point values than observations. In this situation, the random subset design used for making “out-of-sample” predictions when computing the CV(d) statistic more logically is associated with the expected MSEP for using a linear regression model to make a prediction for some future observation at a random X value.</p><p>In Miller [<xref ref-type="bibr" rid="scirp.58573-ref9">9</xref>] (pp. 132-133), an expression is derived for E[MSEP] in the random X case, using a model with an intercept and predictor variables independently sampled from some fixed multivariate normal distribution and general second order errors with mean 0 and variance σ<sup>2</sup>. Miller credits this result to [<xref ref-type="bibr" rid="scirp.58573-ref18">18</xref>] , but uses a derivation from [<xref ref-type="bibr" rid="scirp.58573-ref19">19</xref>] . The equation that Miller derives is</p><disp-formula id="scirp.58573-formula437"><label>. (8)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-1240497x74.png"  xlink:type="simple"/></disp-formula><p>The “1/n” term in the dilation factor accounts for the variance of the intercept parameter estimated in the model. The other term in the dilation factor is developed using a Hotelling T<sup>2</sup>-statistic, which is a generalization of Student’s t-statistic that is used in multivariate hypothesis testing. If we replace n ? d with n in (7), then it is easy to show that (8) and (7) are equivalent.</p><p>Following Miller’s derivation for the case using a model with an intercept, we can also derive an expression equivalent to (6) for the “no intercept” case. The “1/n” term in (8) is not needed because all predictor and response variables are distributed with 0 mean, and no intercept is used in the model. It is a straightforward exercise to show that the other term in the dilation factor becomes<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x75.png" xlink:type="simple"/></inline-formula>, giving</p><disp-formula id="scirp.58573-formula438"><label>. (9)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/4-1240497x76.png"  xlink:type="simple"/></disp-formula><p>(9) is identical to (6) if we substitute n for n ? d in (6).</p><p>In support of the error generalization, simulations using <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x77.png" xlink:type="simple"/></inline-formula> (where a is randomly drawn from the interval <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x77.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x78.png" xlink:type="simple"/></inline-formula> for each simulated data set) were found to produce APE values on the same order as those observed using normally distributed ε. Regarding the normality constraint on X, if we independently sample pre-</p><p>dictor values from<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x79.png" xlink:type="simple"/></inline-formula>, then simulated E[CV(d)] values are less than the conjectured values.</p><p>For example, with <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x80.png" xlink:type="simple"/></inline-formula> and no intercept, E[CV(d)] values ranged from 2.8% - 9.6% smaller than the conjectured formula in (6) as d increased from 1 to (d<sub>max</sub> ? 2), but simulated E[REG] values were unchanged (as expected, since E[REG] is independent of predictor distribution). Therefore, unlike some of the more general properties for OLS linear regression, E[CV(d)] appears to depend on predictor distribution. It is worth noting that the two-point instability phenomenon persisted in the uniformly distributed X case and thus appears to be robust to predictor distribution.</p></sec><sec id="s3_4"><title>3.4. Simulating CV(d) for Model Selection in a Small Sample Setting</title><p>Recall the asymptotic model selection conditions from [<xref ref-type="bibr" rid="scirp.58573-ref11">11</xref>] requiring that d/n → 1and n ? d → ∞ as n → ∞. The conclusion to be drawn from these constraints is that when using CV(d) for model ranking in model selection, a value for d that is an appreciable fraction of sample size n is preferred. However, no specific guidance is provided, as the finite sample situation is inconsequential to the asymptotic result. For example, setting<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x81.png" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x81.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x82.png" xlink:type="simple"/></inline-formula>, satisfies Shao’s two conditions, yet imposes no certain constraint on what values for d are desirable.</p><p>To use CV(d) for model selection in a manner that is consistent with Shao’s setup, one begins with a pool of candidate predictors and evaluates all possible linear models defined by non-empty subsets of this predictor pool for the purpose of estimating some response. The optimal model contains only and all of the predictors that contribute to the response. CV(d) is evaluated for each candidate model, and the model exhibiting the smallest CV(d) value is selected. Define the optimal d value (d<sub>opt</sub>) to correspond with the CV(d) that exhibits the highest rate of selecting the optimal model. If Shao’s result has relevance in the small sample setting, one would expect d<sub>opt</sub> to generally be among the larger allowable d values. Further, we would expect d<sub>opt</sub> to increase nearly at the same rate as n, while at the same time also observing growth in n ? d<sub>opt</sub>.</p><p>For this simulation, define predictor pool<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x83.png" xlink:type="simple"/></inline-formula>, which consists of an intercept and two random variables that are<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x83.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x84.png" xlink:type="simple"/></inline-formula>, where n is sample size. From this three-element predictor pool, there are seven candidate models defined by the non-empty subsets. Values for response variable Y were simulated using <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x83.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x84.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x85.png" xlink:type="simple"/></inline-formula>, where error<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x83.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x84.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x85.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x86.png" xlink:type="simple"/></inline-formula>. The optimal model is defined by the choice for</p><p><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x87.png" xlink:type="simple"/></inline-formula>used to construct Y, for which five values were examined representing the unique cases:</p><p><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x88.png" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x88.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x89.png" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x88.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x89.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x90.png" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x88.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x89.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x90.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x91.png" xlink:type="simple"/></inline-formula>, and<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x88.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x89.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x90.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x91.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x92.png" xlink:type="simple"/></inline-formula>. For each sample size n = 4:20, at least 12,000 iterations (i.e., model selection opportunities) were simulated for each choice for β. For each iteration, CV(d) values were computed for <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x88.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x89.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x90.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x91.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x92.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x93.png" xlink:type="simple"/></inline-formula> for each of the seven candidate models, with the upper bound for d (called d<sub>ub</sub>) determined by the largest allowable d for the full model (which uses all three predictors). For each d, the model with the smallest CV(d) value was identified, allowing for a rate of optimal model selection to be estimated across the iterations. For comparison, REG values also were computed and evaluated for optimal model selection rate.</p><p>Optimal model selection rates using model selectors CV(d) and REG (which is plotted at d = 0 for convenience) are shown in Figures 4(a)-(e) for the five unique optimal model cases, followed by the average optimal model selection rate in <xref ref-type="fig" rid="fig4">Figure 4</xref>(f) reflecting the case where the optimal model can be any one of the seven candidate models. To compute the average, results from optimal model cases <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x94.png" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x94.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x95.png" xlink:type="simple"/></inline-formula> were doubly weighted to account for the redundant cases <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x94.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x95.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x96.png" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x94.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x95.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x96.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x97.png" xlink:type="simple"/></inline-formula> that were not separately simulated.</p><p>At the extreme cases <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x98.png" xlink:type="simple"/></inline-formula> (<xref ref-type="fig" rid="fig4">Figure 4</xref>(a); the intercept model) and <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x98.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x99.png" xlink:type="simple"/></inline-formula> (<xref ref-type="fig" rid="fig4">Figure 4</xref>(e); the full model), we see d<sub>opt</sub> = d<sub>ub</sub> and d<sub>opt</sub> = 1, respectively, for all examined n. For cases in between (Figures 4(b)-(d)), d<sub>opt</sub> varies with n in a generally logical progression. Ultimately we are interested in the behavior of d<sub>opt</sub> in the arbitrary case, where the optimal model can be any one of the seven candidate models. This situation is depicted in <xref ref-type="fig" rid="fig4">Figure 4</xref>(f), which shows the average behavior of CV(d) and REG for model selection. Indeed, d<sub>opt</sub> does appear to exhibit behavior not unlike that implied by Shao’s conditions, suggesting that the essence of Shao’s asymptotic result may well have applicability in this most elementary of model selection scenarios, and perhaps the small sample model selection setting in general.</p><p>Additional simulation results (<xref ref-type="fig" rid="fig5">Figure 5</xref>) using 80% shared variance (correlation) between X<sub>1</sub> and X<sub>2</sub> exhibit similar behavior but with lower and flatter CV(d) rate curves (i.e., reduced capability for optimal model selection and less distinction for d<sub>opt</sub>) and attenuated growth in d<sub>opt</sub>. Though this situation does not precisely conform to Shao’s setup, it is valuable none the less because model selection situations using real data frequently involve correlated predictors.</p></sec></sec><sec id="s4"><title>4. Conclusions</title><p>The first objective of this research was to examine values of E[CV(d)] in a small sample setting using simulation. This effort resulted in Conjectures 1 and 2, which constitute the first explicitly stated, generally applicable formulas for E[CV(d)]. The link established between (6) and (7) and the random-X MSEP described in [<xref ref-type="bibr" rid="scirp.58573-ref9">9</xref>] allowed for the generalization of Conjectures 1 and 2 beyond the limited scope of the simulation, to multivariate normal X and<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x100.png" xlink:type="simple"/></inline-formula>. Simulation using uniformly distributed predictors suggested that the E[CV(d)] statistic does depend on predictor distribution.</p><p>Revelation of the two-point numerical instability at the end of the E[CV(d)] error curve, which was not incompatible with the conjectured formulas for E[CV(d)] because of their breakdown at these values, was an</p><fig id="fig4"  position="float"><label><xref ref-type="fig" rid="fig4">Figure 4</xref></label><caption><title> Optimal model selection rates are shown for the model selection simulation using a candidate predictor pool {1, X<sub>1</sub>, X<sub>2</sub><sub>&#173;</sub>}, for sample sizes n = 4:20. Optimal d values (d<sub>opt</sub>) are circled. Results for particular optimal models (predictor subsets) are shown in (a) {1}; (b) {X<sub>1</sub>} (or {X<sub>2</sub>}); (c) {1,X<sub>1</sub>} (or {1,X<sub>2</sub>}); (d) {X<sub>1</sub>,X<sub>2</sub>}; and (e) {1,X<sub>1</sub>,X<sub>2</sub>}. β values correspond to notation used in the text. (f) shows the average model selection rate across (a)-(e), with (b) and (c) double counted in the average for completeness</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/4-1240497x101.png"/></fig><p>unexpected outcome. This phenomenon, which did not appear to depend on predictor distribution, suggests the curious result that OLS linear regression models fit using just 1 or 0 degrees of freedom must be unique in some way compared to models fit using 2 or more degrees of freedom. Theoretical investigation of this exceptional behavior might best begin by examining the development of the Hotelling T<sup>2</sup>-statistic for the multivariate normal X case that provides the basis for the X-random MSEP in (8).</p><p>For the second objective, an elementary model selection simulation with candidate predictors<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x102.png" xlink:type="simple"/></inline-formula>, where X<sub>1</sub> and X<sub>2</sub> were<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x102.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/4-1240497x103.png" xlink:type="simple"/></inline-formula>, was used to show how the asymptotic model selection conditions {d/n → 1and n ? d → ∞ as n → ∞} of [<xref ref-type="bibr" rid="scirp.58573-ref11">11</xref>] are manifested in the smallest sample setting. When the optimal model was the full model (i.e., the most complex model), then CV(1) was the best model ranking statistic (d<sub>opt</sub> = 1). When the optimal model was the mean model (i.e., the simplest model), then CV(d<sub>ub</sub>) was the best model ranking statistic (d<sub>opt</sub> = d<sub>ub</sub>). For cases in between, d<sub>opt</sub> was observed to vary with n in a generally logical progression dependent on optimal model complexity. Ultimately we are interested in the arbitrary optimal model case, which was simulated by averaging all of the specific optimal model selection rates. For the arbitrary optimal model case, d<sub>opt</sub> and optimal model selection rate demonstrated behavior reflective of the conditions prescribed in [<xref ref-type="bibr" rid="scirp.58573-ref11">11</xref>] , whereby (i) the optimal model selection rate at d = d<sub>opt</sub> increased as n increased, (ii) d<sub>opt</sub> was generally among the larger allowable d values, (iii) d<sub>opt</sub> increased at nearly the same rate as n, and (iv) growth occurred in n ? d<sub>opt</sub>. These behaviors persisted in a dampened fashion when correlated predictors were used, with faster growth observed in n ? d<sub>opt</sub>.</p><fig id="fig5"  position="float"><label><xref ref-type="fig" rid="fig5">Figure 5</xref></label><caption><title> Same as <xref ref-type="fig" rid="fig4">Figure 4</xref>(f), but with X<sub>1</sub> and X<sub>2</sub> values simulated such that Corr(X<sub>1</sub>,X<sub>2</sub>) ≈ 0.8. Optimal d values (d<sub>opt</sub>) are indicated by dots. Results from <xref ref-type="fig" rid="fig4">Figure 4</xref>(f) are shown in the background for comparison</title></caption><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/4-1240497x104.png"/></fig><p>For practitioners, the analyses presented in this paper shed new light on computed CV(d) values, especially for small sample model selection and forecast error variance estimation problems where little is known about the behavior of CV(d). With the conjectured formulas for E[CV(d)] (which appear to be exact for the multivariate normal case), theoreticians can formulate more precise series expansions for CV(d) to facilitate the furthering of our mathematical understanding of this interesting and useful statistic.</p></sec><sec id="s5"><title>Cite this paper</title><p>Jude H.Kastens, (2015) Small Sample Behaviors of the Delete-d Cross Validation Statistic. Open Journal of Statistics,05,382-392. doi: 10.4236/ojs.2015.55040</p></sec></body><back><ref-list><title>References</title><ref id="scirp.58573-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Zhang, P. (1993) Model Selection via Multifold Cross Validation. The Annals of Statistics, 21, 299-313.  
http://dx.doi.org/10.1214/aos/1176349027</mixed-citation></ref><ref id="scirp.58573-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">McQuarrie, A.D.R. and Tsai, C. (1998) Regression and Time Series Model Selection. World Scientific Publishing Co. Pte. Ltd., River Edge, NJ.</mixed-citation></ref><ref id="scirp.58573-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Seber, G.A.F. and Lee, A.J. (2003) Linear Regression Analysis, Second Edition. John Wiley &amp; Sons, Inc., Hoboken, NJ. http://dx.doi.org/10.1002/9780471722199</mixed-citation></ref><ref id="scirp.58573-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Allen, D.M. (1974) The Relationship between Variable Selection and Data Augmentation and a Method for Prediction. Technometrics, 16, 125-127. http://dx.doi.org/10.1080/00401706.1974.10489157</mixed-citation></ref><ref id="scirp.58573-ref5"><label>5</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Stone</surname><given-names> M. </given-names></name>,<etal>et al</etal>. (<year>1974</year>)<article-title>Cross-Validatory Choice and Assessment of Statistical Prediction (with Discussion)</article-title><source> Journal of the Royal Statistical Society (Series B)</source><volume> 36</volume>,<fpage> 111</fpage>-<lpage>147</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.58573-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Geisser, S. (1975) The Predictive Sample Reuse Method with Applications. Journal of the American Statistical Association, 70, 320-328. http://dx.doi.org/10.1080/01621459.1975.10479865</mixed-citation></ref><ref id="scirp.58573-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Breiman, L., Friedman, J.H., Olshen, R.A. and Stone, C.J. (1984) Classification and Regression Trees. Wadsworth, Belmont, CA.</mixed-citation></ref><ref id="scirp.58573-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Hjorth, J.S.U. (1994) Computer Intensive Statistical Methods. Chapman &amp; Hall/CRC, New York.</mixed-citation></ref><ref id="scirp.58573-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Miller, A. (2002) Subset Selection in Regression. 2nd Edition, Chapman &amp; Hall/CRC, New York.  
http://dx.doi.org/10.1201/9781420035933</mixed-citation></ref><ref id="scirp.58573-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Davison, A.C. and Hinkley, D.V. (1997) Bootstrap Methods and their Application. Cambridge University Press, New York. http://dx.doi.org/10.1017/CBO9780511802843</mixed-citation></ref><ref id="scirp.58573-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Shao, J. (1993) Linear Model Selection by Cross-Validation. Journal of the American Statistical Association, 88, 486-494. http://dx.doi.org/10.1080/01621459.1993.10476299</mixed-citation></ref><ref id="scirp.58573-ref12"><label>12</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Mallows</surname><given-names> C.L. </given-names></name>,<etal>et al</etal>. (<year>1973</year>)<article-title>Some Comments on Cp</article-title><source> Technometrics</source><volume> 15</volume>,<fpage> 661</fpage>-<lpage>675</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.58573-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Akaike, H. (1973) Information Theory and an Extension of the Maximum Likelihood Principle. Proceedings of 2nd International Symposium on Information Theory, Budapest, 267-281.</mixed-citation></ref><ref id="scirp.58573-ref14"><label>14</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Shao</surname><given-names> J. </given-names></name>,<etal>et al</etal>. (<year>1997</year>)<article-title>An Asymptotic Theory for Linear Model Selection</article-title><source> Statistica Sinica</source><volume> 7</volume>,<fpage> 221</fpage>-<lpage>264</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.58573-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Shibata, R. (1984) Approximate Efficiency of a Selection Procedure for the Number of Regression Variables. Biometrika, 71, 43-49. http://dx.doi.org/10.1093/biomet/71.1.43</mixed-citation></ref><ref id="scirp.58573-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Akaike, H. (1970) Statistical Predictor Identification. Annals of the Institute of Statistical Mathematics, 22, 203-217.  
http://dx.doi.org/10.1007/BF02506337</mixed-citation></ref><ref id="scirp.58573-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Shao, J. and Tu, D. (1995) The Jackknife and Bootstrap. Springer-Verlag, Inc., New York.  
http://dx.doi.org/10.1007/978-1-4612-0795-5</mixed-citation></ref><ref id="scirp.58573-ref18"><label>18</label><mixed-citation publication-type="book" xlink:type="simple">Stein, C. (1960) Multiple Regression. In: Olkin, I., et al., Eds., Contributions to Probability and Statistics, Stanford University Press, Stanford, CA, 424-443.</mixed-citation></ref><ref id="scirp.58573-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Bendel, R.B. (1973) Stopping Rules in Forward Stepwise-Regression. Ph.D. Dissertation, Univ. of California at Los Angeles.</mixed-citation></ref></ref-list></back></article>