<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OJS</journal-id><journal-title-group><journal-title>Open Journal of Statistics</journal-title></journal-title-group><issn pub-type="epub">2161-718X</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/ojs.2021.113023</article-id><article-id pub-id-type="publisher-id">OJS-109743</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Pivot Points in Bivariate Linear Regression
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Carl</surname><given-names>V. Lutzer</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>David</surname><given-names>L. Farnsworth</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>School of Mathematical Sciences, Rochester Institute of Technology, Rochester, New York, USA</addr-line></aff><pub-date pub-type="epub"><day>08</day><month>05</month><year>2021</year></pub-date><volume>11</volume><issue>03</issue><fpage>393</fpage><lpage>399</lpage><history><date date-type="received"><day>11,</day>	<month>May</month>	<year>2021</year></date><date date-type="rev-recd"><day>6,</day>	<month>June</month>	<year>2021</year>	</date><date date-type="accepted"><day>9,</day>	<month>June</month>	<year>2021</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  There are little-noticed points in the plane, which are artifacts of linear regression. The points, which are called pivot points, are the intersections of sets of regression lines. We derive the coordinates of the pivot point and explain its sources. We show how a pivot point arises in a certain notable data set, which has been analyzed often for points of high leverage. We obtain the application of pivot points that shortens calculations when updating a set of bivariate observations by adding a new point.
 
</p></abstract><kwd-group><kwd>Augmented Data Set</kwd><kwd> Bilinear Regression</kwd><kwd> Influence</kwd><kwd> Leverage</kwd><kwd> Pivot Point</kwd><kwd> Updating a Regression Line</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>It is common to produce many lines to fit bivariate data as the observations are being altered in some way. For example, in order to determine a particular data point’s influence on the best fit, the point may be moved by changing its y-coordinate and a new line created. Some diagnostic tests are based on this. A point, which is called the pivot point, is the intersection of certain lines that are often used for examining influence.</p><p>An example of a pivot point is presented in Section 2. In Section 3, we derive the coordinates of the pivot point. We show that a pivot point can be created in two ways. One way is augmenting an original set of bivariate observations with an additional point, which can have arbitrary multiplicity. Another way is altering an existing observation’s y-coordinate as described above. Section 4 presents the benefit of the pivot point in that it can be useful to shorten calculations when adding a new observation.</p></sec><sec id="s2"><title>2. Illustrative Example</title><p>Consider the data in <xref ref-type="table" rid="table1">Table 1</xref> [<xref ref-type="bibr" rid="scirp.109743-ref1">1</xref>]. The predictor variable (x) is the age in months at which a child says their first word, and the response variable (y) is the child’s Gesell Adaptive Score from an aptitude test. These data have been analyzed many times for influential and outlying observations [<xref ref-type="bibr" rid="scirp.109743-ref2">2</xref>] - [<xref ref-type="bibr" rid="scirp.109743-ref7">7</xref>]. Using various criteria, Cases 2, 18, and 19 have been identified as significant. For illustrative purposes, we focus on Case 18.</p><p>When examining an individual observation’s influence on a bivariate least-squares linear regression, it is common to generate a sequence of regression lines. These lines fit the same set of observations, except that the y-coordinate is made to vary while its x-coordinate is unchanged for the specified data point of interest. The influence of Case 18 on the least-squares regression line is examined by keeping its x-coordinate of 42 and giving its y-coordinate the values 57, 77, 97, 117, and 137. This produces the five regression lines in <xref ref-type="fig" rid="fig1">Figure 1</xref>. Clearly, Case 18 could have a large influence on the regression line. Some authors have illustrated and evaluated leverage in this way [<xref ref-type="bibr" rid="scirp.109743-ref8">8</xref>] [<xref ref-type="bibr" rid="scirp.109743-ref9">9</xref>] [<xref ref-type="bibr" rid="scirp.109743-ref10">10</xref>] [<xref ref-type="bibr" rid="scirp.109743-ref11">11</xref>]. All these regression lines pass though a common point, called the pivot point [<xref ref-type="bibr" rid="scirp.109743-ref12">12</xref>]. In <xref ref-type="fig" rid="fig1">Figure 1</xref>, the pivot point (12.3, 96.1) is shared by the five lines, and its location is indicated by the symbol D.</p></sec><sec id="s3"><title>3. Derivation of the Pivot Point</title><p>We derive the formula for the coordinates of the pivot point. The pivot point can be created by augmenting an original set of bivariate observations with an additional point, which can have arbitrary multiplicity, which is another method to diagnose influence on the line [<xref ref-type="bibr" rid="scirp.109743-ref5">5</xref>] [<xref ref-type="bibr" rid="scirp.109743-ref13">13</xref>] [<xref ref-type="bibr" rid="scirp.109743-ref14">14</xref>] [<xref ref-type="bibr" rid="scirp.109743-ref15">15</xref>]. We show that formulation to be equivalent to varying the location of a single point, while keeping the same first coordinate, as is done in <xref ref-type="fig" rid="fig1">Figure 1</xref>.</p><p>Consider the bivariate data set S 0 = { ( x i , y i ) : i = 1 , 2 , ⋯ , n } . For simplicity, assume that coordinates are selected so that ( ∑ x / n , ∑ y / n ) = ( 0 , 0 ) . Unindexed summations are over the elements of S<sub>0</sub>. Define V = ∑ x 2 / n . Introduce m copies of the new point R(u,v). If R is a point in S<sub>0</sub>, these are additional copies. The aggregate of S<sub>0</sub> and m &gt; 0 copies of R is denoted S<sub>m</sub>.</p><p>For m = 0, the least-squares regression line of S<sub>0</sub> is</p><p>y = a 0 + b 0 x = ( ∑ x y / ∑ x 2 ) x .</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Age at First Word (x) and Gesell Adaptive Score (y)</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >Case</th><th align="center" valign="middle" >x</th><th align="center" valign="middle" >y</th><th align="center" valign="middle" >Case</th><th align="center" valign="middle" >x</th><th align="center" valign="middle" >y</th><th align="center" valign="middle" >Case</th><th align="center" valign="middle" >x</th><th align="center" valign="middle" >y</th></tr></thead><tr><td align="center" valign="middle" >1</td><td align="center" valign="middle" >15</td><td align="center" valign="middle" >95</td><td align="center" valign="middle" >8</td><td align="center" valign="middle" >11</td><td align="center" valign="middle" >100</td><td align="center" valign="middle" >15</td><td align="center" valign="middle" >11</td><td align="center" valign="middle" >102</td></tr><tr><td align="center" valign="middle" >2</td><td align="center" valign="middle" >26</td><td align="center" valign="middle" >71</td><td align="center" valign="middle" >9</td><td align="center" valign="middle" >8</td><td align="center" valign="middle" >104</td><td align="center" valign="middle" >16</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >100</td></tr><tr><td align="center" valign="middle" >3</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >83</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >20</td><td align="center" valign="middle" >94</td><td align="center" valign="middle" >17</td><td align="center" valign="middle" >12</td><td align="center" valign="middle" >105</td></tr><tr><td align="center" valign="middle" >4</td><td align="center" valign="middle" >9</td><td align="center" valign="middle" >91</td><td align="center" valign="middle" >11</td><td align="center" valign="middle" >7</td><td align="center" valign="middle" >113</td><td align="center" valign="middle" >18</td><td align="center" valign="middle" >42</td><td align="center" valign="middle" >57</td></tr><tr><td align="center" valign="middle" >5</td><td align="center" valign="middle" >15</td><td align="center" valign="middle" >102</td><td align="center" valign="middle" >12</td><td align="center" valign="middle" >9</td><td align="center" valign="middle" >96</td><td align="center" valign="middle" >19</td><td align="center" valign="middle" >17</td><td align="center" valign="middle" >121</td></tr><tr><td align="center" valign="middle" >6</td><td align="center" valign="middle" >20</td><td align="center" valign="middle" >87</td><td align="center" valign="middle" >13</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >83</td><td align="center" valign="middle" >20</td><td align="center" valign="middle" >11</td><td align="center" valign="middle" >86</td></tr><tr><td align="center" valign="middle" >7</td><td align="center" valign="middle" >18</td><td align="center" valign="middle" >93</td><td align="center" valign="middle" >14</td><td align="center" valign="middle" >11</td><td align="center" valign="middle" >84</td><td align="center" valign="middle" >21</td><td align="center" valign="middle" >10</td><td align="center" valign="middle" >100</td></tr></tbody></table></table-wrap><p>For any integer m ≥ 0, the least-squares regression line of S<sub>m</sub> is</p><p>y = a m + b m x = m V ( v − b 0 u ) ( m + n ) V + m u 2 + ( m + n ) V b 0 + m u v ( m + n ) V + m u 2 x , (1)</p><p>and the point of means is</p><p>M m = ( m m + n u , m m + n v ) , (2)</p><p>which is on line (1) for S<sub>m</sub>.</p><p>When m &gt; 0 and u ≠ 0, the pivot point</p><p>P = ( − V u , − V b 0 u ) (3)</p><p>is on the least-squares line for all setsS<sub>m</sub>. This can be seen by substituting point (3) into the equation of the line (1), that is,</p><p>a m + b m ( − V / u ) = − V b 0 / u .</p><p>Point P on (3) is called the pivot point ofR with respect toS<sub>0</sub>, because P is on all regression lines for S<sub>m</sub>, which have different slopes. Because the y-coordinate v of R is absent from the coordinates of P, it is also called the pivot point ofu with respect to S<sub>0</sub>. The set of regression lines that is created by adding copies of R, is called a pencil of lines or fan of lines throughP.</p><p>When u = 0, the best-fit line (1) translates in the y-direction as m increases, and the pivot point is said to be at infinity. The pivot point is solely an artifact of the least-squares regression equations. Initially, it was found and explained in a linear-algebraic setting [<xref ref-type="bibr" rid="scirp.109743-ref12">12</xref>].</p><p>The regression lines in a fan, which is formed by vertically moving one point in the data set, intersect at the pivot point. In particular, the regression line formed by addingm copies of the pointR(u,v) toS<sub>0</sub> is equivalent to the line formed by adding a single point (u,v<sub>m</sub>) with</p><p>v m = n ( 1 − m ) V b 0 u ( m + n ) V + m u 2 + m ( ( 1 + n ) V + u 2 ) ( m + n ) V + m u 2 v ,</p><p>which can be seen algebraically by setting m = 1 and v = v m in line (1), which yields (1).</p><p>Pivot points occur when the data are not centered at the origin. All best-fit lines can be rigidly translated, so that the new center is ( x &#175; , y &#175; ) . The slope of each line can be found from</p><p>∑ ( x − x &#175; ) ( y − y &#175; ) ∑ ( x − x &#175; ) 2 ,</p><p>which shows the dependence solely on the differences of each coordinate from its mean. The observations in <xref ref-type="fig" rid="fig1">Figure 1</xref> are centered at the data set’s mean point ( x &#175; , y &#175; ) .</p></sec><sec id="s4"><title>4. Computational Shortcuts When Augmenting a Bivariate Set</title><p>The pivot point offers two shortcuts for computing equations for regression lines. This is analogous to adding the n + 1<sup>st</sup> value a to the data set { x i : i = 1 , 2 , ⋯ , n } , whose mean is x &#175; . The new mean can be calculated using ( n x &#175; + a ) / ( n + 1 ) , which requires considerably less computation than not using x &#175; [<xref ref-type="bibr" rid="scirp.109743-ref11">11</xref>].</p><p>One shortcut is, given setS<sub>0</sub>, the regression line forS<sub>m</sub> can be computed as the line containing the point of means (2) and the pivot point (3). Recall that in (4), V and b<sub>0</sub> are based only on the unaugmented data set.</p><p>The second shortcut involves the line obtained when multiplicitym becomes very large, then the line (1) approaches the line</p><p>y = a ∞ + b ∞ x = V ( v − b 0 u ) V + u 2 + V b 0 + u v V + u 2 x , (4)</p><p>which contains the new pointR and the pivot pointP. The coefficients in (4) provide the tool for rapid computation for the line (1) for any m, including m = 1 for a single additional point. In (1), a<sub>m</sub> is a weighted average of a<sub>0</sub> and a ∞ , andb<sub>m</sub> is a weighted average ofb<sub>0</sub> and b ∞ with the same weights, in particular,</p><p>a m = w a 0 + ( 1 − w ) a ∞ and b m = w b 0 + ( 1 − w ) b ∞ , (5)</p><p>where</p><p>w = n V ( m + n ) v + m u 2 , (6)</p><p>Equations (5) are seen by substituting a<sub>0</sub> and b<sub>0</sub> from (1), a ∞ and b ∞ from (4), and w from (6) into the right-hand sides of (5), which yields a<sub>m</sub> and b<sub>m</sub> in (1).</p></sec><sec id="s5"><title>5. Conclusion</title><p>Pivot points are omnipresent in applications of bivariate linear regression. In particular, they are points through which new lines pass when a data point is altered. One important purpose of altering a point is to determine its influence. We have displayed this phenomenon with the well-known data set of ages at first word versus Gesell scores, which has been analyzed by many authors from many points of view. A pivot point is a handy and efficient tool for shortening calculations when new data arises.</p></sec><sec id="s6"><title>Acknowledgements</title><p>We are grateful to many of our colleagues who have frequently and freely shared their knowledge about regression and computational statistics.</p></sec><sec id="s7"><title>Conflicts of Interest</title><p>The authors declare no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s8"><title>Cite this paper</title><p>Lutzer, C.V. and Farnsworth, D.L. (2021) Pivot Points in Bivariate Linear Regression. Open Journal of Statistics, 11, 393-399. https://doi.org/10.4236/ojs.2021.113023</p></sec></body><back><ref-list><title>References</title><ref id="scirp.109743-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Mickey, R.M., Dunn, O.J. and Clark, V. (1967) Note on the Use of Stepwise Regression in Detecting Outliers. Computer and Biomedical Research, 1, 105-111. 
https://doi.org/10.1016/0010-4809(67)90009-2</mixed-citation></ref><ref id="scirp.109743-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Andrews, D.F. and Pregibon, D. (1978) Finding the Outliers that Matter. Journal of the Royal Statistical Society, Series B (Methodological), 40, 85-93. 
https://doi.org/10.1111/j.2517-6161.1978.tb01652.x</mixed-citation></ref><ref id="scirp.109743-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Dempster, A.P. and Gasko-Green, M. (1981) New Tools for Residual Analysis. Annals of Statistics, 9, 945-959. https://doi.org/10.1214/aos/1176345575</mixed-citation></ref><ref id="scirp.109743-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Draper, N.R. and John, J.A. (1981) Influential Observations and Outliers in Regression. Technometrics, 23, 21-26. https://doi.org/10.1080/00401706.1981.10486232</mixed-citation></ref><ref id="scirp.109743-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Moore, D.S., Notz, W.I. and Fligner, M.A. (2017) The Basic Practice of Statistics. 8th Edition, Freeman, New York.</mixed-citation></ref><ref id="scirp.109743-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Paul, S.R. (1983) Sequential Detection of Unusual Points in Regression. Journal of the Royal Statistical Society, Series D (The Statistician), 32, 417-424. 
https://doi.org/10.2307/2987543</mixed-citation></ref><ref id="scirp.109743-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Rousseeuw, P.J. and Leroy, A.M. (1987) Robust Regression and Outlier Detection. Wiley, New York. https://doi.org/10.1002/0471725382</mixed-citation></ref><ref id="scirp.109743-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Chatterjee, S. and Hadi, A.S. (1986) Influential Observations, High Leverage Points, and Outliers in Linear Regression. Statistical Science, 1, 379-393. 
https://doi.org/10.1214/ss/1177013630</mixed-citation></ref><ref id="scirp.109743-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Hoaglin, D.C. (1988) Using Leverage and Influence to Introduce Regression Diagnostics. College Mathematics Journal, 19, 387-416. 
https://doi.org/10.1080/07468342.1988.11973146</mixed-citation></ref><ref id="scirp.109743-ref10"><label>10</label><mixed-citation publication-type="book" xlink:type="simple">Hoaglin, D.C. (1992) Diagnostics. In: Hoaglin, D.C. and Moore, D.S., Eds., Perspectives on Contemporary Statistics, Mathematical Association of America, Washington, 123-144.</mixed-citation></ref><ref id="scirp.109743-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Montgomery, D.C., Runger, G.C. and Hubele, N.F. (2011) Engineering Statistics. 5th Edition, Wiley, New York.</mixed-citation></ref><ref id="scirp.109743-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Lutzer, C.V. (2017) A Curious Feature of Regression. College Mathematics Journal, 48, 189-198. https://doi.org/10.4169/college.math.j.48.3.189</mixed-citation></ref><ref id="scirp.109743-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Brase, C.H. and Brase, C.P. (2017) Understandable Statistics: Concepts and Methods. 12th Edition, Cengage Learning, Boston.</mixed-citation></ref><ref id="scirp.109743-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Larose, D.T. (2015) Discovering Statistics. 3rd Edition, Freeman, New York.</mixed-citation></ref><ref id="scirp.109743-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Triola, M.F. (2017) Elementary Statistics. 13th Edition, Pearson, Boston.</mixed-citation></ref></ref-list></back></article>