<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OJS</journal-id><journal-title-group><journal-title>Open Journal of Statistics</journal-title></journal-title-group><issn pub-type="epub">2161-718X</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/ojs.2013.31002</article-id><article-id pub-id-type="publisher-id">OJS-27914</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Detecting Global Influential Observations in Liu Regression Model
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>boobacker</surname><given-names>Jahufer</given-names></name><xref ref-type="aff" rid="aff1"><sub>1</sub></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><label>1</label><addr-line>Department of Mathematical Sciences, Faculty of Applied Sciences, South Eastern University of Sri Lanka, Sammanthurai, Sri Lanka</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>jahufer@yahoo.com</email></corresp></author-notes><pub-date pub-type="epub"><day>20</day><month>02</month><year>2013</year></pub-date><volume>03</volume><issue>01</issue><fpage>5</fpage><lpage>11</lpage><history><date date-type="received"><day>May</day>	<month>14,</month>	<year>2012</year></date><date date-type="rev-recd"><day>June</day>	<month>30,</month>	<year>2012</year>	</date><date date-type="accepted"><day>July</day>	<month>10,</month>	<year>2012</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
   In linear regression analysis, detecting anomalous observations is an important step for model building process. Various influential measures based on different motivational arguments and designed to measure the influence of observations on different aspects of various regression results are elucidated and critiqued. The presence of influential observations in the data is complicated by the presence of multicollinearity. In this paper, when Liu estimator is used to mitigate the effects of multicollinearity the influence of some observations can be drastically modified. Approximate deletion formulas for the detection of influential points are proposed for Liu estimator. Two real macroeconomic data sets are used to illustrate the methodologies proposed in this paper.
     
 
</p></abstract><kwd-group><kwd>Liu Estimator; Global Influential Observations; Diagnostics; Multicollinearity; Case Deletion; Approximate Deletion Formulas</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>The presence of multicollinearity in the regressors seriously affects the parameter estimation and prediction. Therefore, mixed estimation and ridge type regression are suggested to mitigate this effect. In addition to multicollinearity, the presence of influential observations in the observed data is seriously affected by the estimators as those estimators are not unbiased [1,2].</p><p>In literature, many authors [1-3] noted that the influential observations on ridge type estimators are different from the corresponding least squares estimate and that multicollinearity can even disguise anomalous data. Belsley [<xref ref-type="bibr" rid="scirp.27914-ref3">3</xref>] investigated the leverage in ridge regression. Walker and Bitch [<xref ref-type="bibr" rid="scirp.27914-ref2">2</xref>] studied the influence of observations in ordinary ridge regression estimator (ORRE) based on case deletion method. Shi [<xref ref-type="bibr" rid="scirp.27914-ref4">4</xref>] proposed the local influence in principal component analysis by defining a generalized Cook statistic and showed that his method is equivalent to Cook’s approach under the likelihood framework. Shi and Wang [<xref ref-type="bibr" rid="scirp.27914-ref2">2</xref>] analyzed the influential cases in ORRE using local influence method. Jahufer and Chen [<xref ref-type="bibr" rid="scirp.27914-ref5">5</xref>] studied global influential observations in modified ridge regression estimator (MRRE) moreover, Jahufer and Chen [<xref ref-type="bibr" rid="scirp.27914-ref6">6</xref>] studied local influential observations in MRRE. Besides, Jahufer and Chen [<xref ref-type="bibr" rid="scirp.27914-ref7">7</xref>] analyzed local influential observations in Liu estimator.</p><p>The main aim of this paper is therefore to assess the global influence of observations in the linear Liu estimator using the method of case deletion. This method has been extensively studied and it is very powerful for detecting influential cases because of its intuitive appeal and its direct connection to the sample influence curves. Also, it is widely accepted as the foundation of many other statistical methods approached. The methodology proposed in this paper is illustrated using two real macroeconomic data sets. The first Data set is macro impact of foreign direct investment in Sri Lanka. This data set contains four regressors and a response variable with 27 observations. The second data set is Longley [<xref ref-type="bibr" rid="scirp.27914-ref8">8</xref>] data set. It consists of six regressors and a response variable with 16 observations. This paper is composed of six sections: Section 2 gives the background and definition of influential measures in least squares; Section 3 derives the influence measures in Liu estimator; Section 4 describes approximate deletion formulas for Liu estimator; Section 5 reports the examples using two real data sets. Discussion is given in the last section.</p></sec><sec id="s2"><title>2. Background and Definition</title><sec id="s2_1"><title>2.1. Background</title><p>A matrix multiple linear regression model can be written as</p><disp-formula id="scirp.27914-formula50468"><label>(1)</label><graphic position="anchor" xlink:href="2-1240103\1002a85a-e4ba-4078-b225-8128679e3f04.jpg"  xlink:type="simple"/></disp-formula><p>where y is an n &#215; 1 response vector, X is an n &#215; p centered and standardized known matrix (i.e. the length of the column of X is standardized to one), <img src="2-1240103\908e0d9e-70ae-4fa0-84ed-4f750e853d74.jpg" />is a p &#215; 1 vector of an unknown parameter, e is an n &#215; 1 error vector with <img src="2-1240103\4e3ab347-08fa-4f6b-9783-3485d196e6be.jpg" /> and <img src="2-1240103\3508dcd5-e7fa-45b9-93c4-56ebed7aa195.jpg" /> and <img src="2-1240103\137431fb-76d8-497b-a83f-5d8851532856.jpg" /> is an identity matrix of order n. Then the ordinary least squares estimator (OLSE) of <img src="2-1240103\447ededb-17b4-4335-bfc3-81ffb7a63770.jpg" /> is <img src="2-1240103\cf6a6398-e6b2-471e-a25d-2d334b31785c.jpg" /> The estimator of <img src="2-1240103\beaa2d2a-404c-4ae0-9697-03134f19b17e.jpg" /> is <img src="2-1240103\b639de64-14d1-4fe3-af21-c466eb710e46.jpg" /> where residual vector <img src="2-1240103\2ebe4e02-7249-4772-977d-52ff63656ca7.jpg" /></p><p>This article assumes that the reader is acquainted with the basic ideas of leverage and influence in regression analysis, as presented, for instance, in the works of [<xref ref-type="bibr" rid="scirp.27914-ref1">1</xref>] and [<xref ref-type="bibr" rid="scirp.27914-ref9">9</xref>].</p></sec><sec id="s2_2"><title>2.2. Definition of Influential Measures in Least Squares</title><p>The general purpose of influential analysis is to measure the changes induced in a given aspect of the analysis when the data set is perturbed. A particularly appealing perturbation scheme is case deletion. Note that this scheme is used throughout this article.</p><p>In general, the influence of a case can be viewed as the product of two factors: the first a function of the residual and the second a function of the position of the point in the X space. The position or leverage of the i-th point is measured by h<sub>i</sub>, which is the i-th diagonal element of the “hat” matrix<img src="2-1240103\915f4a2a-b46c-443a-a6d4-8ff30081569c.jpg" />.</p><p>Among the most popular single-case influential measure is the difference in fit standardized DFFITS [<xref ref-type="bibr" rid="scirp.27914-ref1">1</xref>], which evaluated at the i-th case is given by</p><disp-formula id="scirp.27914-formula50469"><label>(2)</label><graphic position="anchor" xlink:href="2-1240103\4c9caaef-2ed0-4367-9980-360b37674750.jpg"  xlink:type="simple"/></disp-formula><p>where <img src="2-1240103\1e8381a9-c5cf-4d47-bf2b-26dc04d2a26f.jpg" /> is the least squares estimator of <img src="2-1240103\5502406f-f730-4c38-932c-5f139d560075.jpg" /> without the i-th case and <img src="2-1240103\7d5e80b7-05e6-4607-99c0-fe194b257c5e.jpg" /> is an estimator of the standard error (SE) of the fitted values. DFFITS is the standardized change in the fitted value of a case when it is deleted. Thus it can be considered a measure of influence on individual fitted values.</p><p>Another useful measure of influence is Cook’s D [<xref ref-type="bibr" rid="scirp.27914-ref9">9</xref>], which evaluated at the i-th case is given by</p><disp-formula id="scirp.27914-formula50470"><label>(3)</label><graphic position="anchor" xlink:href="2-1240103\5d4433df-2c5a-44ed-a30d-b519ca069b8a.jpg"  xlink:type="simple"/></disp-formula><p>where <img src="2-1240103\28c99795-bab8-4ecd-98bb-0e3a93e587b2.jpg" /> is a measure of the change in all of the fitted values when a case is deleted. Even though <img src="2-1240103\5d9caad0-618a-435c-a874-d9e74269799b.jpg" /> is based on different theoretical consideration, it is closely related to DFFITS.</p><p>Points with large values of <img src="2-1240103\b04e589f-5e0e-4df1-8b16-c9c87d9c171b.jpg" /> have considerable influence on the least squares estimate<img src="2-1240103\39befb5b-b818-4194-a67b-976ced30a429.jpg" />. In general, points for which <img src="2-1240103\50fc6484-b4cd-4eb7-8251-281cbb3b901a.jpg" /> to be influential. In DFFITS measures any observation for which</p><p><img src="2-1240103\72bd387d-541d-4925-add5-d92852d330d8.jpg" />warrants attention.</p><p>It is important to mention that these measures are useful for detecting single cases with an unduly high influence. For generalizations of Equations (2) and (3) for detecting influential sets (see [1,9]).</p></sec></sec><sec id="s3"><title>3. Influence Measures in Liu Estimator</title><sec id="s3_1"><title>3.1. Liu Estimator</title><p>The Liu estimator <img src="2-1240103\cd9f4c3a-ec43-46f3-83ff-534506c74c35.jpg" /> was introduced by Liu [<xref ref-type="bibr" rid="scirp.27914-ref10">10</xref>] and is defined as</p><disp-formula id="scirp.27914-formula50471"><label>(4)</label><graphic position="anchor" xlink:href="2-1240103\be53c045-8888-4b5e-bb00-871d5c4d9641.jpg"  xlink:type="simple"/></disp-formula><p>where I is an identity matrix, <img src="2-1240103\c2e8c6dc-5be9-4d9c-a4cf-7048dd9fe847.jpg" />is OLSE and d is Liu estimator biasing parameter and it is<img src="2-1240103\31db124e-9d3c-49ba-9e61-90c8ba10434d.jpg" />. The Liu estimator combines the ORRE [11,12] estimator. The ORRE is effective in practice, but it is complicated function of its biasing parameter. Thus we often meet some complicated equations when we use some popular methods, such as ([<xref ref-type="bibr" rid="scirp.27914-ref13">13</xref>], C<sub>k</sub> criterion [<xref ref-type="bibr" rid="scirp.27914-ref14">14</xref>], GCV criterion [<xref ref-type="bibr" rid="scirp.27914-ref15">15</xref>]) and etc. to choose ridge regression biasing parameter k. The advantage of Liu estimator over ORRE is that Liu estimator is a linear function of its biasing parameter d. Therefore, it is convenient to choose Liu estimator biasing parameter d.</p><p>The Liu estimator is very useful to mitigate the effect of near multicollinearity. Also, the recent literature, particularly in the area of econometrics, engineering and other statistical areas, the Liu estimator has produced a number of new techniques and ideas for example [5-7, 16-20].</p></sec><sec id="s3_2"><title>3.2. Leverage and Residual Measures in Liu Estimator</title><p>Using Equation (4), the vector of fitted value of Liu estimator is</p><p><img src="2-1240103\6dda8543-b6bc-4052-8627-22748ad26e9a.jpg" /></p><p>where <img src="2-1240103\76da8797-1edd-45b0-9ac6-b32803b15fc9.jpg" /> is Liu estimator “hat matrix” (see [2,10]) and plays the same role as the hat matrix in OLSE. It is important to note that the matrix H<sub>d</sub> is not a projection matrix because it is not idempotent and H<sub>d</sub> is called quasi-projection matrix (see [<xref ref-type="bibr" rid="scirp.27914-ref2">2</xref>]). The i-th fitted value can be written in terms of elements of H<sub>d</sub> as <img src="2-1240103\5c35936c-36d1-400c-9014-dc1e3c800520.jpg" /> consequently,</p><p><img src="2-1240103\e8b58fdc-1c99-4bdb-82dd-631360b5ab08.jpg" />The “Liu hat diagonals” h<sub>di</sub> can be interpreted as leverage in the same sense as the hat diagonals in OLSE. It is important to note that the H<sub>d</sub> is not idempotent and it is called a quasi-projection matrix (see [<xref ref-type="bibr" rid="scirp.27914-ref2">2</xref>]).</p><p>The single value decomposition (SVD) (see [<xref ref-type="bibr" rid="scirp.27914-ref21">21</xref>]) allows X to be decomposed as X = UDV', where D is a p &#180;</p><p>p diagonal matrix with i-th diagonal elements <img src="2-1240103\1dc7268e-77d8-430e-b250-9dfbc69e8de5.jpg" /> (<img src="2-1240103\9eb793c5-4b91-4263-86b2-541ac53a95f9.jpg" /></p><p>is the i-th eigenvalue of X'X), the column of V are the eigenvectors of X'X. The (ij)-th element of the n &#180; p matrix <img src="2-1240103\24bc5723-3719-438e-bf83-df5737b86051.jpg" /> is such that <img src="2-1240103\2a296f9c-17c3-4485-a5fd-54bb4b63feed.jpg" /> is the projection of the i-th row, x<sub>i</sub>, onto the j-th principal axis (eigenvector) of X. Using the SVD, the Liu estimator leverage of the i-th point can be written as <img src="2-1240103\cd6702f3-ca89-47de-81d7-174fb7c1385a.jpg" /> (see [<xref ref-type="bibr" rid="scirp.27914-ref22">22</xref>]).</p><p>Several important facts can be deduced from the preceding expression. First, for d &gt; 0, <img src="2-1240103\36de34be-25a8-4afe-a5e2-7e8af8c49f49.jpg" />for<img src="2-1240103\7a7df7e5-83d0-4655-9474-2bd6b618d2fd.jpg" />; that is, for every observation the Liu estimator leverage is smaller than the corresponding OLSE leverage. It can be confirmed from the above equation that</p><p><img src="2-1240103\3f7e3b7a-d4ab-4001-8f30-b049bcd69bea.jpg" /></p><p>Second, the leverage decreases monotonically as d increases. Finally, the rate of decrement of leverage depends on the position of the particular row of X along the principal axes. More specifically, the leverage of a row that lies in the direction of principal axes associated with large eigenvalues will be reduced less than the leverage of a row that lies in the direction of principal axes associated with small eigenvalues (see [<xref ref-type="bibr" rid="scirp.27914-ref2">2</xref>]).</p><p>The influence can be differentially affected as d increases. Remember, that influence is not a function of leverage but also of the residual. Although the leverage of every point decreases monotonically as d increases, the effect of this increment on the residuals is far less clear.</p><p>The i-th Liu estimator residual is defined as</p><p><img src="2-1240103\b2f98247-31f4-4906-bee3-22de8cdfdcf8.jpg" /></p></sec><sec id="s3_3"><title>3.3. DFFITS and Cook’s Measures in Liu Estimator</title><p>The DFFITS for Liu estimator can be written as</p><disp-formula id="scirp.27914-formula50472"><label>(5)</label><graphic position="anchor" xlink:href="2-1240103\a8f76cd1-3dfe-49b4-8223-331178646291.jpg"  xlink:type="simple"/></disp-formula><p>where <img src="2-1240103\84aaed65-eb84-425c-8de0-cdee64c51c35.jpg" /> is the Liu estimator in (4) computed with the i-th case deleted and the denominator is an estimator of the standard error of the Liu estimator fitted value. If Liu estimator biasing parameter d is assumed non-stochastic, then</p><p><img src="2-1240103\253a2bd0-b65f-4174-9a28-f0d760110958.jpg" /></p><p>Hence the mean squared error is a function of the fitted values and the response, neither of which depends on individual eigenvalues of <img src="2-1240103\d83a53ae-6941-4cfd-9ad2-307b8483281f.jpg" /> it is not affected by multicollinearity. For this reason, the OLSE of <img src="2-1240103\7aa80f6b-10bc-46f1-b61b-fc509688c53c.jpg" /> will be used as measures of scale.</p><p>At least two versions of Cook’s D<sub>i</sub> can be constructed for Liu estimator, they are</p><disp-formula id="scirp.27914-formula50473"><label>(6)</label><graphic position="anchor" xlink:href="2-1240103\cf84274f-044b-4eb7-adc6-c63d87b6561f.jpg"  xlink:type="simple"/></disp-formula><p>and</p><disp-formula id="scirp.27914-formula50474"><label>(7)</label><graphic position="anchor" xlink:href="2-1240103\5f001efc-c644-4260-a076-0484c4f4ea01.jpg"  xlink:type="simple"/></disp-formula><p>where <img src="2-1240103\ce605654-b6a3-430f-b37b-3d63c1090e8a.jpg" /> is the direct generalization of Cook’s D in (3) and <img src="2-1240103\d296374a-c5f9-4fe1-a019-df0a7d58dfa8.jpg" /> is based on the fact that</p><p><img src="2-1240103\1891a629-2b32-41db-ac57-87b8cbeb8782.jpg" /></p><p>Note that both <img src="2-1240103\2c6374c8-1a93-49ad-87f3-051f1f744384.jpg" /> and <img src="2-1240103\98207010-9add-45b4-b24c-d5eb3d27493d.jpg" /> simplify to D<sub>i</sub> in (3) when d = 1.</p><p>It would be desirable to be able to write these measures as functions of leverage and residual, as was done in (2) and (3). This is not possible, however, because of the scale dependency of the Liu estimator. Since the Liu estimator is not scale invariant, <img src="2-1240103\3115addc-5e5c-481f-82c9-fea6bc7feb18.jpg" />(the X matrix with the i-th row deleted) has to be rescaled to unit-column length before computing <img src="2-1240103\6bac85f6-f06a-43b6-9ef4-d43c7be94191.jpg" /> In the following section some approximate deletion formulas are proposed.</p></sec></sec><sec id="s4"><title>4. Deletion Formulas for Liu Estimator</title><p>In the analysis of influential observations to quantify the impact of the i-th case, the most common approach is to compute single-case diagnostics with the i-th case deleted. In Liu regression, it is impossible to derive an exact formula using case deletion because of the scale dependency of the Liu estimator. Hence, approximate deletion formulas are derived for influential measures DFFITS and two versions of Cook’s statistics.</p><p>The scale dependence of Liu estimator precludes the development of deletion formulas of the types (2) and (3). The main problem resides in the computation of <img src="2-1240103\082539ef-e1f6-4bb2-ab99-7574826755de.jpg" /> because the matrix <img src="2-1240103\aca9e92a-a9fb-454a-a358-fa8f03af1356.jpg" /> has to be standardized same as in Section 2.1, for the small values of d and/or cases with low leverage. However, approximate deletion formulas can be obtained using the Sherman-MorrisonWoodbury (SMW) theorem (see [<xref ref-type="bibr" rid="scirp.27914-ref1">1</xref>]).</p><p>When i-th row is deleted from <img src="2-1240103\2c4c3324-a953-4ba5-b0cf-d2d74e866b02.jpg" /> then <img src="2-1240103\2ec1d797-9829-493a-bf45-8d7c754b57df.jpg" /> can be written as</p><p><img src="2-1240103\381926a8-cc5d-4339-9aea-dbd2c4dcfebe.jpg" /></p><p>where <img src="2-1240103\09b066e9-bdf8-41a9-b67c-1a286b7cf47a.jpg" /> is the matrix X without the i-th row and <img src="2-1240103\a8b316d6-c61b-45f5-871c-5b7fcd1358c1.jpg" /> is the vector of response without the i-th entry. We assume that <img src="2-1240103\94f33c87-dcd1-4e2f-a3d6-3289f2716d2b.jpg" />is centered and scaled so that <img src="2-1240103\762b2bbb-e69d-4f4d-adcc-1e1c70ab2c3e.jpg" />is in correlation form.</p><p>If after deletion the i-th row X is not recentered and rescaled. Then, <img src="2-1240103\3e5dd544-7271-4389-a2ef-adca10b6c97c.jpg" />will not be in exact correlation form. Thus,</p><p><img src="2-1240103\405d52c2-8d66-4136-8175-987335e95a75.jpg" /></p><p>which uses the SMW theorem with <img src="2-1240103\2d9541ab-640c-4bd8-8818-f6eae9149dcb.jpg" /> (see the Appendix), <img src="2-1240103\602a6d4c-e6ac-417f-b9cf-d151763e97b0.jpg" />can be approximated as:</p><p><img src="2-1240103\a09313e4-4c70-4521-a7b3-2c96751847c6.jpg" /></p><p><img src="2-1240103\ceb62f3b-9af3-43df-8305-71f2752b2304.jpg" /></p><p>Therefore,<img src="2-1240103\c2e54c75-fc7d-4b63-9b5c-1d6d87443c40.jpg" />.</p><p>Based on the above result, approximate version of (5), (6) and (7) can written as</p><disp-formula id="scirp.27914-formula50475"><label>, (8)</label><graphic position="anchor" xlink:href="2-1240103\0fad7df7-5e55-4696-a413-5fe352675cc5.jpg"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.27914-formula50476"><label>(9)</label><graphic position="anchor" xlink:href="2-1240103\951f731d-36b6-4fe5-b1c7-9183d02694c9.jpg"  xlink:type="simple"/></disp-formula><disp-formula id="scirp.27914-formula50477"><label>(10)</label><graphic position="anchor" xlink:href="2-1240103\1062d6a8-0d8e-4ebb-829a-5e7dff594f95.jpg"  xlink:type="simple"/></disp-formula></sec><sec id="s5"><title>5. Examples</title><sec id="s5_1"><title>5.1. Example 1: Macroeconomic Impact of Foreign Direct Investment (MIFDI) Data</title><p>Sun [<xref ref-type="bibr" rid="scirp.27914-ref23">23</xref>] studied MIFDI in China 1979-1996. Based on his theory, the MIFDI data were collected in Sri Lanka form 1978 to 2004 to illustrate the methodologies proposed in this paper. The data set consists of four regressors (Foreign Direct Investment, Gross Domestic Product Per Capita, Exchange Rate and Interest Rate) and one response variable (Total Domestic Investment) with 27 observations. The selected variables were tested for statistical conditions: Integration and Multicollinearity. The test results showed that variables are integrated with a same order of integration I(1) at 1% level of significance. The scaled condition number of this data set is 31,244, this large value suggests the presence of an unusually high level of multicollinearity among the regressors (the proposed cutoff is 30; see [<xref ref-type="bibr" rid="scirp.27914-ref1">1</xref>]). The Liu estimator biasing parameter for this data set is estimated d = 0.692.</p><p>Global influence measures Leverage, Residual, DFFITS and two versions of Cook’s D<sub>i</sub> were computed and the results for the most seven influential cases are given in <xref ref-type="table" rid="table1">Table 1</xref>. The influential cases detected by these methods are same except case 26 in leverage and case 14 in residual measures, but only the order of magnitude is changed.</p><p>Using the approximate case deletion formulas in Equations (8)-(10) the influential cases are estimated and it is given in <xref ref-type="table" rid="table2">Table 2</xref>. From this table, it can be seen that the influential cases detected by approximate case deletion formulas DFFITS and two versions of Cook’s D<sub>i</sub> are same but only the order of magnitude is changed. Moreover, influential cases detected by case deletion one by</p><p><xref ref-type="table" rid="table1">Table 1</xref>. The most seven influential observations using leverage, residual, DFFITS and two versions of Cook’s.</p><p><img src="2-1240103\ebda583a-d42f-4fd3-a49b-8ae5385108c8.jpg" /></p><p>one formulas in Equations (5)-(7) and approximate deletion formulas in (8)-(10) are exactly same but only the order of magnitude is changed.</p><p>For verifying these results, it is plane to contribute <xref ref-type="table" rid="table3">Table 3</xref> of Liu estimates for the full data and the data without some influential cases. In this table, the parenthesis value indicates the percentage of change in the parameter value. The result reveals that case 3 is the most influential case while case 22 is the seventh influential point among the detected cases. It is also clear from this table, that omission of single influential cases 3, 23, 1, 27, 2, 18 and 22 contribute the substantial change in the Liu estimator. Among all of these 3, 23, 1, 2 and 18 have a remarkable influence while cases 27 and 22 have a little influence.</p></sec><sec id="s5_2"><title>5.2. Example 2: Longley Data</title><p>The Longley data set [<xref ref-type="bibr" rid="scirp.27914-ref8">8</xref>] has also been used to explain the effect of extreme multicollinearity on the OLSE. The scaled condition number (see [<xref ref-type="bibr" rid="scirp.27914-ref1">1</xref>]) of this data set is 43,275. This large value suggests the presence of high level of multicollinearity among regressors. Cook [<xref ref-type="bibr" rid="scirp.27914-ref24">24</xref>] used this data to identify the influential observations in OLSE using the method of Cook’s D<sub>i</sub> and found that cases 5, 16, 4, 10 and 15 (in this order) were the most influential observations (see <xref ref-type="table" rid="table4">Table 4</xref>). Walker and Birch [<xref ref-type="bibr" rid="scirp.27914-ref2">2</xref>] analyzed the same data to detect anomalous observations in ORRE using case deletion method. They found that cases 16, 10, 4, 15 and 5 (in this order) were most influential observations (see <xref ref-type="table" rid="table4">Table 4</xref>). Shi and Wang [<xref ref-type="bibr" rid="scirp.27914-ref25">25</xref>] also analyzed the same data to detect influential observations on the ridge regression estimator using local influence method. They detected cases 10, 4, 15, 16 and 1 (in this order) were the most five anomalous observations.</p><p>In this paper, I used the same data set to assess the influential observations in Liu estimator using global influence methods: Cook’s D<sub>i</sub>, DFFITS, Leverage and Residual. The estimated results, the most five influential cases and the corresponding values are given in <xref ref-type="table" rid="table4">Table 4</xref>.</p><p>The approximate deletion formulas in Equations (8)- (10) are used to detect influential cases for Longley data and detected influential cases are given in <xref ref-type="table" rid="table5">Table 5</xref>. According to this table, it can be confirmed that identified influential cases using approximate deletion formulas DFFITS and two versions of Cook’s D<sub>i</sub> are exactly same. Besides, influential cases identified by case deletion one by one formulas in Equations (5)-(7) and approximate deletion formulas in (8)-(10) are precisely same for Longly data.</p><p>The influential cases in MIFDI and Longly data were identified using one by one deletion formulas in Equations (5)-(7) and approximate deletion formulas in Equations (8)-(10) respectively. The identified influential cases for MIFDI and Longly data are same for both measures but only the order of magnitude is changed for MIFDI data. Hence, instead of using the one by one case deletion formulas to identify influential cases in Liu estimator the approximate case deletion formulas are more suitable and appropriate to identify influential cases in Liu estimator.</p><p><xref ref-type="table" rid="table2">Table 2</xref>. The most seven influence observations according to approximate case deletion formulas.</p><p><img src="2-1240103\3f3efdf2-ed13-4685-95b4-f06a8ac58734.jpg" /></p><p><xref ref-type="table" rid="table3">Table 3</xref>. Impact of influential cases on Liu estimator parameter.</p><p><img src="2-1240103\63deb5ee-5ee9-4226-b65d-48714f7918c6.jpg" /></p><p><xref ref-type="table" rid="table4">Table 4</xref>. The five most influence observations in Longley data.</p><p><img src="2-1240103\14d90677-9d17-4ab2-876e-3af614ab51c4.jpg" /></p><p><xref ref-type="table" rid="table5">Table 5</xref>. The most five influential cases for Longly data using deletion formulas.</p><p><img src="2-1240103\933e39b2-74b1-4f09-b292-6bca23f30c71.jpg" /></p></sec></sec><sec id="s6"><title>6. Discussion</title><p>In this article, I show that the Liu estimator user not rely on influence measures obtained for OLSE. Once the value of d is determined, influence measures should be computed for that d. If, after analyzing these indexes, it is decided to delete one or more cases from the analysis, the whole situation has to be reassessed in terms of both influence and multicollinearity.</p><p>In this research study, the Liu estimator shrinkage parameter d is estimated first and for that d value the Liu estimator co-efficients are estimated. Using these parameter quantities the influential observations are identified. But, the value of shrinkage parameter d depends on the every observation. Hence, for every influential case the value of d should be estimated. This is very difficult task so this issue will be studied in future research study.</p><p>The main advantage of the deletion formulas in Section 4 is that, as in least squares, the estimator does not have to be computed every time a case is deleted. For a value of d all of the elements in (8), (9) and (10) are readily available from a single run of Liu estimator. Moreover, these measures, based on deletion formulas are particularly helpful for large data sets. Furthermore, the deletion formulas provide computationally inexpensive approximate influence measures for Liu estimator.</p><p>Although no conventional cutoff points are introduced or developed for the Liu estimator global influence diagnostics: Cook’s measures, DFFITS, Leverage and Liu Estimator standardized residual, it seems that index plot is an optimistic and conventional procedure to disclose influential cases. It is a bottleneck for cutoff values for the influence method. These are additional active issues for future research study.</p></sec><sec id="s7"><title>REFERENCES</title></sec><sec id="s8"><title>Appendix</title>Sherman Morrison-Woodbury (SMW) Theorem<p>Consider the p &#215; p matrix <img src="2-1240103\ee108776-c212-403c-8031-554df095bfe1.jpg" /> and let <img src="2-1240103\026c151e-0396-4d44-831a-894d3770a135.jpg" /> be the i-th row of <img src="2-1240103\495006e6-84bc-4921-b637-b3dcc7704163.jpg" />Note that <img src="2-1240103\b84d009e-e6c9-488e-ae60-7e4dc9da26d8.jpg" /> is the <img src="2-1240103\b1dd074f-e7c7-4e19-93e7-37ccd9155b20.jpg" /> matrix with the i-th row removed. The inverse matrix of <img src="2-1240103\fa51ca80-27db-4bcb-8c2e-94cf24e9513c.jpg" /> is:</p><p><img src="2-1240103\0099f402-7c29-4833-bef0-f57d36ec395f.jpg" />.</p></sec></body><back><ref-list><title>References</title><ref id="scirp.27914-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">D. A. Belsley, E. Kuh and R. E. Welsch, “Regression Diagnostics: Identifying Influence Data and Source of Collinearity,” Wiley, New York, 1980. 
doi:10.1002/0471725153</mixed-citation></ref><ref id="scirp.27914-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">E. Walker and J. B. Birch, “Influence Measures in Ridge Regression,” Technometrics, Vol. 30, No. 2, 1988, pp. 221- 227. doi:10.1080/00401706.1988.10488370</mixed-citation></ref><ref id="scirp.27914-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">D. A. Belsley, “Conditioning Diagnostics: Collinearity and Weak Data in Regression,” Wiley, New York, 1991.</mixed-citation></ref><ref id="scirp.27914-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">L. Shi, “Local Influence in Principal Component Analysis,” Biometrika, Vol. 84, No. 1, 1997, pp. 175-186. 
doi:10.1093/biomet/84.1.175</mixed-citation></ref><ref id="scirp.27914-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">A. Jahufer and J. B. Chen, “Assessing Global Influential Observations in Modified Ridge Regression,” Statistics and Probability Letters, Vol. 79, No. 4, 2009, pp. 513- 518. doi:10.1016/j.spl.2008.09.019</mixed-citation></ref><ref id="scirp.27914-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">A. Jahufer and J. Chen, “Measuring Local Influential Observations in Modified Ridge Regression,” Journal of Data Science, Vol. 9, No. 3, 2011, pp. 359-372.</mixed-citation></ref><ref id="scirp.27914-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">A. Jahufer and J. B. Chen, “Identifying Local Influential Observations in Liu Estimator,” Journal of Metrika, Vol. 75, No. 3, 2012, pp. 425-438. 
doi:10.1007/s00184-010-0334-4</mixed-citation></ref><ref id="scirp.27914-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">J. W. Longley, “An Appraisal of Least Squares Programs for Electronic Computer for the Point of View of the User,” Journal of American Statistical Association, Vol. 62, No. 319, 1967, pp. 819-841. 
doi:10.1080/01621459.1967.10500896</mixed-citation></ref><ref id="scirp.27914-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">R. D. Cook and S. Weisberg, “Residuals and Influence in Regression,” Chapman &amp; Hall, London, 1982.</mixed-citation></ref><ref id="scirp.27914-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">K. Liu, “A New Class of Biased Estimate in Linear Regression,” Communications in Statistics—Theory and Methods, Vol. 22, No. 2, 1993, pp. 393-402.</mixed-citation></ref><ref id="scirp.27914-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">A. E. Hoerl and R. W. Kennard, “Ridge Regression: Biased Estimation for Nonorthogonal Problems,” Technometrics, Vol. 12, No. 1, 1970, pp. 55-67. 
doi:10.1080/00401706.1970.10488634</mixed-citation></ref><ref id="scirp.27914-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">C. Stein, “Inadmissibility of the Usual Estimator for the Mean of a Multivariate Normal Distribution,” Proceeding of the third Berkeley Symposium on Mathematical Statistics and Probability, Berkeley, December 1954 and July-August 1955, pp. 197-206.</mixed-citation></ref><ref id="scirp.27914-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">G. C. Mcdonald and D. I. Galarneau, “A Monte Carlo Evaluation of Some Ridge-Type Estimators,” Journal of American Statistical Association, Vol. 70, No. 350, 1975, pp. 407-416. doi:10.1080/01621459.1975.10479882</mixed-citation></ref><ref id="scirp.27914-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">C. L. Mallows, “Some comments on Cp,” Technometrics, Vol. 15, No. 4, 1973, pp. 661-675.</mixed-citation></ref><ref id="scirp.27914-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">G. Wahba, G. H. Golub and C. G. Heath, “Generalized Cross-Validation as a Method for Choosing a Good Ridge Parameter,” Technometrics, Vol. 24, No. 2, 1979, pp. 215-223. doi:10.1080/00401706.1979.10489751</mixed-citation></ref><ref id="scirp.27914-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">F. Akdeniz and S. Ka?iranlar, “More on the New Biased Estimator in Linear Regression,” The Indian Journal of Statistics, Vol. 63, No. 3, 2001, pp. 321-325. </mixed-citation></ref><ref id="scirp.27914-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">S, Ka?iranlar and S. Sakallio?in, “Combining the Liu Estimator and the Principal Component Regression Estimator,” Communications in Statistics—Theory and Methods, Vol. 30, No. 12, 2001, pp. 2699-2705. </mixed-citation></ref><ref id="scirp.27914-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">S, Ka?iranlar, G. P. H. Styan and H. J. Werner, “A New Biased Estimator In Linear Regression and a Detailed Analysis of the Widely Analyzed Dataset on Portland Cement,” The Indian Journal of Statistics, Vol. 61, No. B3, 1999, pp. 443-459. </mixed-citation></ref><ref id="scirp.27914-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">M. H. Hubert and P. Wijekoon, “Improvement of the Liu Estimator in Linear Regression Model,” Journal of Statistical Papers, Vol. 47, No. 3, 2006, pp. 471-479.  
doi:10.1007/s00362-006-0300-4</mixed-citation></ref><ref id="scirp.27914-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">N. Torigoe and K. Ujiie, “On the Restricted Liu Estimator in the Gauss-Markov Model,” Communications in Statistics—Theory and Methods, Vol. 35, No. 9, 2006, pp. 1713-1722.</mixed-citation></ref><ref id="scirp.27914-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">J. Mandel, “Use of the Singular Value Decomposition in Regression Analysis,” The American Statistician, Vol. 36, No. 1, 1982, pp. 15-24.</mixed-citation></ref><ref id="scirp.27914-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">A. S. Top?uba?i and N. Billor, “A Class of Biased Estimators and Their Diagnostic Measures,” 2001. 
http://idari.cu.edu.tr /sempozyum/bil26.htm </mixed-citation></ref><ref id="scirp.27914-ref23"><label>23</label><mixed-citation publication-type="other" xlink:type="simple">H. Sun, “Macroeconomic Impact of Direct Foreign Investment in China: 1979-1996,” Blackwell Publishers Ltd., Oxford, 1998.</mixed-citation></ref><ref id="scirp.27914-ref24"><label>24</label><mixed-citation publication-type="other" xlink:type="simple">R. D. Cook, “Detection of Influential Observations in Linear Regression,” Technometrics, Vol. 19, No. 1, 1977, pp. 15-18. doi:10.2307/1268249</mixed-citation></ref><ref id="scirp.27914-ref25"><label>25</label><mixed-citation publication-type="other" xlink:type="simple">L. Shi and X. Wang, “Local Influence in Ridge Regression,” Computational Statistics &amp; Data Analysis, Vol. 31, No. 3, 1999, pp. 341-353. 
doi:10.1016/S0167-9473(99)00019-5</mixed-citation></ref></ref-list></back></article>