<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OJS</journal-id><journal-title-group><journal-title>Open Journal of Statistics</journal-title></journal-title-group><issn pub-type="epub">2161-718X</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/ojs.2013.32017</article-id><article-id pub-id-type="publisher-id">OJS-30719</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  The Statistical Analysis of Interval-Censored Failure Time Data with Applications
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>adhey</surname><given-names>S. Singh</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Dishna</surname><given-names>P. Totawattage</given-names></name><xref ref-type="aff" rid="aff2"><sup>2</sup></xref></contrib></contrib-group><aff id="aff2"><addr-line>Formerly at University of Guelph, Guelph, Canada</addr-line></aff><aff id="aff1"><addr-line>University of Waterloo, Waterloo, Canada</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>rssingh@uoguelph.ca(ASS)</email>;</corresp></author-notes><pub-date pub-type="epub"><day>16</day><month>04</month><year>2013</year></pub-date><volume>03</volume><issue>02</issue><fpage>155</fpage><lpage>166</lpage><history><date date-type="received"><day>June</day>	<month>14,</month>	<year>2012</year></date><date date-type="rev-recd"><day>July</day>	<month>17,</month>	<year>2012</year>	</date><date date-type="accepted"><day>July</day>	<month>31,</month>	<year>2012</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
   The analysis of survival data is a major focus of statistics. Interval censored data reflect uncertainty as to the exact times the units failed within an interval. This type of data frequently comes from tests or situations where the objects of interest are not constantly monitored. Thus events are known only to have occurred between the two observation periods. Interval censoring has become increasingly common in the areas that produce failure time data. This paper explores the statistical analysis of interval-censored failure time data with applications. Three different data sets, namely Breast Cancer, Hemophilia, and AIDS data were used to illustrate the methods during this study. Both parametric and nonparametric methods of analysis are carried out in this study. Theory and methodology of fitted models for the interval-censored data are described. Fitting of parametric and non-parametric models to three real data sets are considered. Results derived from different methods are presented and also compared. 
 
</p></abstract><kwd-group><kwd>Interval Cens oring; Survival Analysis; Parametric;Non-Parametric; Semi-Parametric; Survival Functions; Survival Curves; Kaplan-Meier Estimate; Turnbull Estimator; Logspline Estimation</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>A great many studies in statistics deal with deaths or failures of components: they involve the numbers of deaths, the timing of death, or the risks of death to which different classes of individuals are exposed. The analysis of survival data is a major focus of statistics.</p><p>In standard time-to-event analysis, the time to a particular event of interest is observed exactly or right-censored. Numerous methods are available for estimating the survival curve and also for estimation of the effects of covariates for these cases. In certain situations, the times of the event of interest may not be exactly known. This means that it may have occurred within particular time duration. In clinical trials, patients are often seen at pre-scheduled visits but the event of interest may have occurred in between visits. These types of data are known as interval-censored data.</p><p>Right-censored data can be considered as a special case of interval-censored data. Some of the inference approaches for right-censored data can be directly, or with minor modifications, used to analyze interval-censored data. However, most of the inference approaches for right-censored data are not appropriate for interval-censored data due to fundamental differences between these two types of censoring. The censoring approach behind interval censoring is more complicated than that of right censoring. For right-censored failure time data, substantial advances in the theory and development of modern statistical methods are based on the counting processes theory, which is not applicable to interval-censored data. Due to the complexity and special structure of interval censoring, the same theory is not applicable to interval-censored data.</p><p>Interval censoring has become increasingly common in the areas that produce failure time data. Over the past two decades, a lot of literature on the statistical analysis of interval-censored failure time data has appeared.</p><p>Lindsay and Ryan [<xref ref-type="bibr" rid="scirp.30719-ref1">1</xref>] provided a tutorial on Biostatistical methods for interval-censored data. This paper illustrated and compared available methods which correctly treated the data as being interval-censored. This paper did not provide a full review of all existing methods. However, all approaches were illustrated on two data sets and compared with methods which ignore the interval-censored nature of the data. In this paper, we have used some of the methodologies, notations and equations used by Lindsay and Ryan [<xref ref-type="bibr" rid="scirp.30719-ref1">1</xref>].</p><p>Lindsay [<xref ref-type="bibr" rid="scirp.30719-ref2">2</xref>] showed that parametric models for interval censored data can now easily be fitted with minimal programming in certain standard statistical software packages. Regression equations were introduced and finite mixture models were also fitted. Models based on nine different distributions were compared for three examples of heavily censored data as well as a set of simulated data. It has been found that interval censoring can be ignored for parametric models. Parametric models are remarkably robust with changing distributional assumptions and more informative than the corresponding non-parametric models for heavily interval censored data.</p><p>Finkelstein and Wolfe [<xref ref-type="bibr" rid="scirp.30719-ref3">3</xref>] provided a method for regression analysis to accommodate interval-censored data. Finkelstein [<xref ref-type="bibr" rid="scirp.30719-ref4">4</xref>] develops a method for fitting proportional hazards regression model when the data contain intervalcensored observations. The method described in this paper is used to analyze data from an animal study and also a clinical trial.</p><p>Peto [<xref ref-type="bibr" rid="scirp.30719-ref5">5</xref>] provided a method of calculating an estimate of the cumulative distribution function from intervalcensored data, which was similar to the life-table technique.</p><p>Rosenberg [<xref ref-type="bibr" rid="scirp.30719-ref6">6</xref>] presented a flexible parametric procedure to model the hazard function as a linear combination of cubic B-Splines and derived maximum likelihood estimates from censored survival data. This provided smooth estimates of the hazard and survivorship functions that are intermediate between parametric and non-parametric models. HIV infections data that were intervalcensored were used to illustrate the methods.</p><p>Odell et al. [<xref ref-type="bibr" rid="scirp.30719-ref7">7</xref>] studied the use of a Weibull-based accelerated failure time regression model when intervalcensored data were observed. They have used two alternative methods to analyze the data. Turnbull [<xref ref-type="bibr" rid="scirp.30719-ref8">8</xref>] has used non-parametric estimation of a distribution function for censored data. A simple algorithm using self-consistency as a basis was used to get maximum likelihood estimates.</p><p>Farrington [<xref ref-type="bibr" rid="scirp.30719-ref9">9</xref>] provided a method for weak parametric modeling of interval-censored data using generalized linear models. Three types of models, namely, additive, multiplicative and proportional hazard model with discrete baseline survival function were considered. Goetghebeur and Ryan [<xref ref-type="bibr" rid="scirp.30719-ref10">10</xref>] introduced semi-parametric regression analysis of interval censored data. A semi-parametric approach to the proportional hazards regression analysis of interval-censored data was proposed in this paper. The method was illustrated on data from the breast cancer cosmetics trial, previously analyzed by Finkelstein [<xref ref-type="bibr" rid="scirp.30719-ref4">4</xref>].</p><p>Lawless [<xref ref-type="bibr" rid="scirp.30719-ref11">11</xref>] provides a unified treatment of models and statistical methods used in the analysis of lifetime or response time data (Chapter 3, Section 3.5.3, p. 124). Numerical illustrations and examples involving real data demonstrate the application of each method to problems in areas such as reliability, product performance evaluation, clinical trials, and experimentation in the biomedical sciences. Collet [<xref ref-type="bibr" rid="scirp.30719-ref12">12</xref>] describes and illustrates the modeling approach to the analysis of survival data. Some methods for analyzing interval-censored data are described and illustrated. This begins with an introduction to survival analysis and a description of four studies in which survival data was obtained. These and other data sets then illustrate the techniques presented, including the Cox and Weibull proportional hazards models, accelerated failure time models, models with time-dependent variables, interval-censored survival data and model checking.</p><p>Sun [<xref ref-type="bibr" rid="scirp.30719-ref13">13</xref>] has recently presented statistical models and methods specifically developed for the analysis of interval-censored failure time data. This book collects and unifies statistical models and methods that have been proposed for analyzing interval-censored failure time data. It provides the first comprehensive coverage of the topic of interval-censored data. This focuses on non-parametric and semi-parametric inferences, but it also describes parametric and imputation approaches. This paper reviews the substantial body of recent work in this field and also provides some applications.</p></sec><sec id="s2"><title>2. Statistical Methodology</title><sec id="s2_1"><title>2.1. Parametric Methods</title><p>The straightforward procedure to analyze censored data is to assume a parametric model for the failure times. It is possible to fit Accelerated failure time (AFT) models for a variety of distributions to interval censored data. In the statistical area of survival analysis, an accelerated failure time model is a parametric model that provides an alternative to the commonly-used proportional hazards models. A proportional hazards model assumes that the effect of a covariate is to multiply the hazard by some constant; an AFT model assumes that the effect of a covariate is to multiply the predicted event time by some constant. AFT models can therefore be framed as linear models for the logarithm of the survival time.</p><p>The results of AFT models are easily interpreted. For example, the results of a clinical trial with mortality as the endpoint could be interpreted as a certain percentage increase in future life expectancy on the new treatment compared to the control. So a patient could be informed that he would be expected to live (say) 15% longer if he took the new treatment. Hazard ratios are harder to explain in layman’s terms. More probability distributions can be used in AFT models than parametric proportional hazards models. A distribution must have a parameterization that includes a scale parameter to be used in an AFT model. The logarithm of the scale parameter is then modeled as a linear function of the covariates.</p><p>The Weibull distribution (including the exponential distribution as a special case) can be parameterized as either a proportional hazards model or an AFT model, and is the only family of distributions to have this property. The results of fitting a Weibull model can therefore be interpreted in either framework. Unlike the Weibull distribution, log-logistic distribution can exhibit a non-monotonic hazard function which increases at early times and will decrease at later times.&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;&#160;</p><p>Other distributions suitable for AFT models include the log-normal and log-gamma distributions, although they are less popular than the log-logistic, partly as their cumulative distribution functions do not have a closed form.</p><p>The SAS procedure LIFEREG provides a way of fitting accelerated failure time models for a variety of distributions to interval censored data. The AFT model is defined by the transformation</p><disp-formula id="scirp.30719-formula16695"><label>, (2.1)</label><graphic position="anchor" xlink:href="10-1240113\47ba10dc-b73b-432d-a466-af1d2f231572.jpg"  xlink:type="simple"/></disp-formula><p>where <img src="10-1240113\86beaba2-109a-489e-b58d-b9bf6d7fcef0.jpg" /> is the failure time random variable for an individual with covariate z and <img src="10-1240113\86590254-e5c8-49a5-a000-92418ddfe3e5.jpg" /> is the failure time that the individual would have if they had covariate value 0. The effect of changing covariates is to shrink or stretch the time to event. If <img src="10-1240113\0a203844-3cfd-4ed0-a2a7-5b1b7162911d.jpg" /> is negative, then the covariate has the effect of “speeding up time” so that individuals with larger values of z have higher failure rates and hence shorter survival times. The survival function can be written as</p><disp-formula id="scirp.30719-formula16696"><label>, (2.2)</label><graphic position="anchor" xlink:href="10-1240113\3fa4fc32-d3ca-4cee-baf9-1f38b08edc65.jpg"  xlink:type="simple"/></disp-formula><p>where <img src="10-1240113\c0fabe4a-415b-414f-a592-91dafc58e94b.jpg" /> is the survival function for an individual with covariate value 0. Taking natural logarithm, the AFT model can be expressed as</p><disp-formula id="scirp.30719-formula16697"><label>. (2.3)</label><graphic position="anchor" xlink:href="10-1240113\9ad0b433-9625-4788-9c79-92f99da4ff0f.jpg"  xlink:type="simple"/></disp-formula><p>If we assume that <img src="10-1240113\80b92d55-fb51-4ad1-ba09-bbf2d63d5689.jpg" /> can be expressed as<img src="10-1240113\bc77b49c-3745-4e0c-970b-b732c4ef9919.jpg" />, where W is a random variable, then the model can be written in a linear model-like form:</p><disp-formula id="scirp.30719-formula16698"><label>. (2.4)</label><graphic position="anchor" xlink:href="10-1240113\bbdede16-bc20-4a1d-a088-423aeb84a4a5.jpg"  xlink:type="simple"/></disp-formula><p>The PROC LIFEREG module of SAS fits this model, except that the sign is changed on the regression coefficients. That is, SAS fits</p><disp-formula id="scirp.30719-formula16699"><label>. (2.5)</label><graphic position="anchor" xlink:href="10-1240113\9b72f255-fc80-4838-98bf-abddc8fe4286.jpg"  xlink:type="simple"/></disp-formula><p>It is possible to include a variety of distributions to be placed on the error term W with SAS, including the log of the exponential, log-normal and log-gamma distributions. The intercept parameter <img src="10-1240113\0755b2c5-3296-4a56-8d53-707a8c1b326d.jpg" /> and the scale parameter <img src="10-1240113\8cbc5184-d836-4d8d-bade-def854b28d9f.jpg" /> are usually not of direct interest, although for some distributions, there is a relationship between the AFT model and a proportional hazards model through the scale parameter. For example, if W is an extreme value distribution (log of a unit exponential), then T has a Weibull distribution. Note that because of the change in sign implicit in the AFT formulation, the direction of covariate effects will be opposite to those fit with a Cox proportional hazards model.</p></sec><sec id="s2_2"><title>2.2. Non-Parametric Estimation of Survival Curve</title><sec id="s2_2_1"><title>2.2.1. Kaplan-Meier Estimator</title><p>The Kaplan-Meier estimator estimates the survival function for life-time data. In medical research, it might be used to measure the fraction of patients living for a certain amount of time after treatment. An economist might measure the length of time people remain unemployed after a job loss. An engineer might measure the time until failure of machine parts.</p><p>A plot of the Kaplan-Meier estimate of the survival function is a series of horizontal steps of declining magnitude which, when a large enough sample is taken, approaches the true survival function for that population. The value of the survival function between successive distinct sampled observations is assumed to be constant.</p><p>An important advantage of the Kaplan-Meier curve is that the method can take censored data into account, for instance, if a patient withdraws from a study. When no truncation or censoring occurs, the Kaplan-Meier curve is equivalent to the empirical distribution.</p><p>Let <img src="10-1240113\a232dc97-4b59-40ca-8516-218b23a999b9.jpg" /> be the probability that an item from a given population will have a lifetime exceeding t. For a sample from this population of size N let the observed times until death of N sample members be</p><disp-formula id="scirp.30719-formula16700"><label>. (2.6)</label><graphic position="anchor" xlink:href="10-1240113\3f06d14c-5e4d-4fc2-a026-616af97f47f4.jpg"  xlink:type="simple"/></disp-formula><p>Corresponding to each <img src="10-1240113\4c906ca4-cf64-4067-b4a0-d260194f728b.jpg" /> is<img src="10-1240113\6c553708-fb61-4148-b415-cb31df370269.jpg" />, the number “at risk” just prior to time<img src="10-1240113\ab5948b5-04ef-4dcc-8bbd-1e783fa6ea94.jpg" />, and<img src="10-1240113\6abe4e53-e455-4113-a372-5afa898ed786.jpg" />, the number of deaths at time<img src="10-1240113\4915ca9e-81c7-4648-8f7a-847d4901ef5e.jpg" />.</p><p>Note that the intervals between each time typically will not be uniform. For example, a small data set might begin with 10 cases, have a death at Day 3, a loss (censored case) at Day 9, and another death at Day 11. Then we have<img src="10-1240113\e3ed8914-fc71-4969-9c95-cf5938aa3967.jpg" />, <img src="10-1240113\d016f5f8-ee37-48c3-858f-317539859c2e.jpg" />, and<img src="10-1240113\23546d4f-d1e3-49e5-a6d0-8d772385d3dc.jpg" />.</p><p>The Kaplan-Meier estimator is the nonparametric maximum likelihood estimate of<img src="10-1240113\570a5447-9633-4555-9ace-6dd04adf4fec.jpg" />. It is a product of the form</p><disp-formula id="scirp.30719-formula16701"><label>. (2.7)</label><graphic position="anchor" xlink:href="10-1240113\34cc153c-2b30-4525-8788-75741d81dfbf.jpg"  xlink:type="simple"/></disp-formula><p>When there is no censoring, <img src="10-1240113\f32d7a76-1ce6-4007-b4d1-a45a3d9d1234.jpg" />is just the number of survivors just prior to time<img src="10-1240113\bac42751-94fc-4317-8c7e-5b11f55cd0a0.jpg" />. With censoring, <img src="10-1240113\e2431173-7cf5-44c0-a2c0-6197227edc87.jpg" />is the number of survivors less the number of losses (censored cases). It is only those surviving cases that are still being observed (have not yet been censored) that are “at risk” of an (observed) death.</p><p>Let T be the random variable that measures the time of failure and let <img src="10-1240113\04f7495f-e515-43ac-b59a-3a71a062beec.jpg" /> be its cumulative distribution function. Note that</p><disp-formula id="scirp.30719-formula16702"><label>(2.8)</label><graphic position="anchor" xlink:href="10-1240113\6ffdb159-a9c6-40a2-9429-051ce73c08af.jpg"  xlink:type="simple"/></disp-formula><p>Consequently, the right-continuous definition of <img src="10-1240113\8cd62dd7-3fec-428d-8705-354be1e7b007.jpg" /> may be preferred in order to make the estimate compatible with a right-continuous estimate of<img src="10-1240113\574b33b9-f162-407c-942f-fc6a26f7b066.jpg" />.</p><p>With right-censored data, Kaplan and Meier [<xref ref-type="bibr" rid="scirp.30719-ref14">14</xref>] showed that the closed form product limit estimator is the generalized maximum likelihood estimate. This curve jumps at each observed event time.</p></sec><sec id="s2_2_2"><title>2.2.2. Turnbull Estimator</title><p>In most applications, the data may be interval-censored. By interval-censored data, a random variable of interest is known only to lie in an interval, instead of being observed exactly. In such cases, the only information available for each individual is that their event time falls in an interval, but the exact time is unknown. A nonparametric estimate of the survival function can also be found in such interval censored situations. The survival function is perhaps the most important function in medical and health studies. In this section, the iterative procedure proposed by Turnbull [<xref ref-type="bibr" rid="scirp.30719-ref8">8</xref>] to estimate such function is described and illustrated.</p><p>Situations where the observed response for each individual under study is either an exact survival time or a censoring time are common in practice. Other situations, however, can occur, and amongst them we find the longitudinal studies, where the individuals are followed for a pre-fixed time period or visited periodically for a fixed number of times. In this context, the time <img src="10-1240113\ab0e70e2-f5d0-4bbc-b29f-79033973b6c6.jpg" /> until the occurrence of the event of interest for each individual is only known (whenever it occurs) to be within the interval between visits, i.e., between the visit in time <img src="10-1240113\ea1db413-f003-4784-a763-088ffb1771b7.jpg" /> and the visit in time<img src="10-1240113\eea24f70-bca7-40cf-947b-30a33c5563f2.jpg" />. Note that in such studies, the survival times <img src="10-1240113\4038a168-0ae9-43df-9bd4-cd09de0733de.jpg" /> are no longer known exactly. It is only known that the event of interest has occurred within the interval <img src="10-1240113\4eac9cbe-2b8a-4613-8f18-d4e64cab1aa4.jpg" /> with<img src="10-1240113\61cb47ed-bb86-49da-b367-ca76a6222a67.jpg" />. Furthermore, note that if the event occurs exactly at the moment of a visit, which is very improbable but can happen, then we have an exact survival time. In this case it is assumed that<img src="10-1240113\5e31bf90-520c-4c89-a10e-b612d1217a4f.jpg" />.</p><p>On the other hand, it is known for the individuals with right censoring that the event of interest did not occur until the last visit but it can happen at any time from that moment on. We therefore assumed in this case that <img src="10-1240113\0940f41d-3908-45d6-9bd1-f42e501d9482.jpg" /> can occur within the interval <img src="10-1240113\56379eb3-0afc-4c18-827d-9bf9e3197f71.jpg" /> &#160;with <img src="10-1240113\9168903c-f2e7-41e3-b803-9db79cb70f76.jpg" /> being equal to the period of time from the beginning of the study until the last visit and<img src="10-1240113\569d1cfd-8898-4204-a958-f8192591da66.jpg" />.</p><p>Similarly, it is known for the individuals that are left censored, that the event of interest has occurred before the first visit and, hence, we assume that <img src="10-1240113\01db5e5e-dab2-4ddd-931f-78f72b12c718.jpg" /> falls in the interval <img src="10-1240113\12cd7b2e-2104-4f37-a8f8-cf43a8263ea9.jpg" /> with <img src="10-1240113\37e93607-93e0-4d4c-b1ca-7772517f7664.jpg" /> representing the beginning of the study and <img src="10-1240113\397e0308-9ce9-4cff-873b-9e2a275b33ac.jpg" /> is the period of time from the beginning of the study until the first visit.</p><p>Note from what we have presented so far that exact survival times as well as right and left censored data, are all special cases of interval survival data with <img src="10-1240113\3c1741e4-2d0f-4690-a851-6ee08bcc13d1.jpg" /> for exact times, <img src="10-1240113\fbc48a1c-bfa1-4206-9878-3cc9f50b5ad7.jpg" />for right censoring and <img src="10-1240113\6e64a2f6-2726-4b22-ae15-1924c5a61236.jpg" /> &#160;for left censoring. We can therefore state that interval survival data generalize any situation with combinations of survival times (exact or interval) and right and left censoring that can occur in survival studies.</p><p>As usual in the analysis of non-interval survival data, it is also of interest to estimate the survival function <img src="10-1240113\4f3ad796-d2e3-44d2-b253-7b3218835bfc.jpg" /> and to assess the importance of potential prognostic factors. Few statistical software allow for such data, and for this reason a common practice amongst data analysts is to assume that the event occurring within the interval <img src="10-1240113\37b31139-3ffb-4ea6-9c1c-78ee6165bcd1.jpg" /> has occurred either at the upper/lower limit of the interval or, at the middle point of each interval. Rucker and Messerer [<xref ref-type="bibr" rid="scirp.30719-ref15">15</xref>], Odell et al. [<xref ref-type="bibr" rid="scirp.30719-ref7">7</xref>] and Dorey et al. [<xref ref-type="bibr" rid="scirp.30719-ref16">16</xref>] stated that assuming interval survival times as exact times can lead to biased estimates as well as results and conclusions that were not fully reliable. In this work we describe a nonparametric procedure for estimation of the survival function for interval survival data.</p><p>Peto [<xref ref-type="bibr" rid="scirp.30719-ref5">5</xref>] was the first to propose a non-parametric method for estimating the survival distribution based on interval-censored data. Turnbull [<xref ref-type="bibr" rid="scirp.30719-ref8">8</xref>] derived the same estimator, but used a different approach in estimation. Suppose<img src="10-1240113\60c2752a-5e6b-4133-a953-856b7b550b3a.jpg" />, the survival times for n patients, are independent random variables with right continuous survival function<img src="10-1240113\1280d0f8-12aa-461c-a5c0-027daf4b1594.jpg" />. If <img src="10-1240113\2885573c-0a1e-4aa4-9c0e-c8fbbfbabc05.jpg" /> are not observed directly, but instead are known to lie in the interval<img src="10-1240113\8ddfd546-537a-4c2d-a69f-4930a735a59b.jpg" />, then the likelihood for the n observations is,</p><disp-formula id="scirp.30719-formula16703"><label>. (2.9)</label><graphic position="anchor" xlink:href="10-1240113\88bbc082-e330-42ac-ae80-d3eac4267a91.jpg"  xlink:type="simple"/></disp-formula><p>By<img src="10-1240113\fe5c70b3-704a-46a6-962d-d00ddead08f5.jpg" />, we mean</p><disp-formula id="scirp.30719-formula16704"><label>(2.10)</label><graphic position="anchor" xlink:href="10-1240113\d81e7ae4-a097-4a80-8ea4-9a4834a26258.jpg"  xlink:type="simple"/></disp-formula><p>which may be different from<img src="10-1240113\9c9f40df-a876-4314-9e0b-e0d8c56c8f47.jpg" />, since <img src="10-1240113\b82f0edb-3710-4762-8c72-f055e7aa3887.jpg" /> is left continuous. It is important to note that different authors vary in their conventions regarding definition of the censoring interval. The Convention of Peto [<xref ref-type="bibr" rid="scirp.30719-ref5">5</xref>] and Turnbull [<xref ref-type="bibr" rid="scirp.30719-ref8">8</xref>] who assumed a closed interval, <img src="10-1240113\ca61a382-3e30-4192-99e2-c91fe581b34f.jpg" />, was followed. This definition facilitates the accommodation of observations that are known exactly, that is, <img src="10-1240113\643735ff-f670-41e3-89b3-bebee62d5684.jpg" />, but necessitates the use of the <img src="10-1240113\8482a473-6907-4585-a8e3-76871a833319.jpg" /> notation in above Equation (2.9) to allow a non-zero contribution to the likelihood for these observations. Finkelstein [<xref ref-type="bibr" rid="scirp.30719-ref4">4</xref>] assumed semiclosed censoring intervals, which need to add the convention that the likelihood contribution for any observation with an exact failure time, <img src="10-1240113\5384cce0-7caa-4d80-a709-0c8bc6ab2d93.jpg" />, is<img src="10-1240113\ef4ca947-117d-4a6b-a267-78401bf213fb.jpg" />. Good arguments against almost any convention can be made for defining the censoring intervals. In practice, the choice will have little impact and all reasonable conventions can be adopted.</p><p>Turnbull [<xref ref-type="bibr" rid="scirp.30719-ref8">8</xref>] derived the same estimator using an iterative self-consistency algorithm, described below. Gentleman and Geyer [<xref ref-type="bibr" rid="scirp.30719-ref17">17</xref>] showed that this self-consistent estimator is not always the maximum likelihood estimator (MLE), and that the MLE is not necessarily unique and discuss conditions under which this can be determined.</p><p>Since the observed event times are known to occur only within potentially overlapping intervals, the survival curve can only jump within so-called equivalence sets<img src="10-1240113\5585f97b-71ec-4f57-9d89-700bdb8ad4d8.jpg" />, <img src="10-1240113\201ec3ed-afad-4c09-aedc-49eafa58959e.jpg" />, where<img src="10-1240113\c2494939-6ce4-4df3-9d86-d41c3493677b.jpg" />. The curve between <img src="10-1240113\ddee049f-ec3c-452f-98eb-ea966d4476f3.jpg" /> and <img src="10-1240113\6a48732f-c183-4866-a166-9179a376a923.jpg" /> is flat. The estimate of <img src="10-1240113\cdff3c0e-f4ff-4135-be36-6d4033376476.jpg" /> is unique only up to these equivalence classes; any function that jumps the appropriate amount within the equivalence class will yield the same likelihood.</p><p>An analog of the Product-Limit estimator of the survival function for interval-censored data is presented in this section. This estimator, which has no closed form, is based on an iterative procedure and has been suggested by Turnbull [<xref ref-type="bibr" rid="scirp.30719-ref8">8</xref>].</p><p>To construct the estimator, let <img src="10-1240113\4bc49a43-8f7f-4d7b-8d0a-76b624959f57.jpg" /> be a grid of time which includes all the points <img src="10-1240113\d3997fe5-acc2-44ba-a799-fdd78f1d1763.jpg" /> and <img src="10-1240113\db3dcb57-84b8-4838-8bb7-df9858b1394c.jpg" /> for<img src="10-1240113\a50c3bec-c80e-4ce1-8d51-73eec3ae5ee1.jpg" />. For the <img src="10-1240113\7b5a98ac-1c87-4abe-8980-cfe749ea31ef.jpg" /> observation, define a weight <img src="10-1240113\fc5af972-d12d-4151-b222-30a453bf45b2.jpg" /> to be 1 if the interval <img src="10-1240113\37698d6c-a578-4515-ab3c-cd97b8aaa193.jpg" /> is contained in the interval <img src="10-1240113\3f778a70-4aae-4b67-a30d-f32dfb8cf65c.jpg" /> and 0, otherwise. The weight <img src="10-1240113\6dcfb899-953f-4ce6-a07a-b08de73f29ec.jpg" /> indicates whether the event which occurs in the interval <img src="10-1240113\a89420e4-f8ae-4c12-92b0-405f1133a26d.jpg" /> could have occurred at<img src="10-1240113\4326fa3a-6850-4978-8064-50b52d074b3f.jpg" />. An initial guess at <img src="10-1240113\5094e668-183e-416b-bde4-8ba921fbb943.jpg" /> is made and Turnbull’s algorithm is as follows:</p><p>Step 1: Compute the probability of an event occurring at time <img src="10-1240113\d47a0559-2c16-4992-94cb-75f88cd639d8.jpg" /> by</p><disp-formula id="scirp.30719-formula16705"><label>(2.11)</label><graphic position="anchor" xlink:href="10-1240113\f694484a-1f20-44aa-911f-d2e6ab02b2c8.jpg"  xlink:type="simple"/></disp-formula><p>Step 2: Estimate the number of events which occurred at <img src="10-1240113\8a1d007f-efaa-4279-8c2e-4dada2cc4534.jpg" /> by</p><disp-formula id="scirp.30719-formula16706"><label>(2.12)</label><graphic position="anchor" xlink:href="10-1240113\f5b5ed6f-3910-47ac-9512-88b13f07ba67.jpg"  xlink:type="simple"/></disp-formula><p>Step 3: Compute the estimated number at risk at time</p><p><img src="10-1240113\074ee09f-ef0d-45f4-8d17-3db6d64db6e8.jpg" />by <img src="10-1240113\c7739e8b-fe13-413e-a186-a9b8a24d62b1.jpg" /></p><p>Step 4: Compute the updated Product-Limit estimator using the pseudo data found in Steps 2 and 3. If the updated estimate of S is close to the old version of S for all<img src="10-1240113\d145dc1e-7f18-49ea-b3c4-a17952ec8307.jpg" />’s, stop the iterative process, otherwise repeat Steps 1- 3, using the updated estimate of S.</p></sec><sec id="s2_2_3"><title>2.2.3. Logspline Estimation of the Survival Curve</title><p>Kooperberg and Stone [<xref ref-type="bibr" rid="scirp.30719-ref18">18</xref>] have introduced the Logspline density estimation. They have developed a system for data that may be right censored, left censored, or interval censored. A fully automatic method was used to determine the estimate, which involved the maximum likelihood method and may involve stepwise knot deletion and either the Akaike information criterion (AIC) or Bayesian information criterion (BIC), was used to determine the estimate.</p><p>Kooperberg and Stone [<xref ref-type="bibr" rid="scirp.30719-ref18">18</xref>] provided software (logspline. fit, available through Statlib for S-plus2) which can be used to obtain smoothed estimates of the survival function based on interval censored data using splines. Smooth functions were fitted to the log-density function of the failure times within subsets of the time axis defined by the “knots”, and constrained to be continuous at those points. This provides a loosely parametric framework for finding estimates of the survival and hazard functions which can be useful for exploratory data analysis. Their approach is related to that of Rosenberg [<xref ref-type="bibr" rid="scirp.30719-ref6">6</xref>] who uses splines to model the hazard function.</p></sec></sec></sec><sec id="s3"><title>3. Applications</title><p>There are essentially three approaches to fit survival models. The first straightforward method is the parametric approach, where a specific functional form for the baseline hazard <img src="10-1240113\f9816fb3-744b-4293-95c4-32dcd18a9e0d.jpg" /> is assumed. Examples are: models based on the exponential, Weibull, gamma and generalized F distributions. A second approach might be called a flexible or semi-parametric strategy, where mild assumptions are made about the baseline hazard<img src="10-1240113\d6e71a1c-aaae-431c-a238-4e488f5f37db.jpg" />. Specifically, time is subdivided into reasonably small intervals and it is assumed that the baseline hazard is constant in each interval leading to a piecewise exponential model. The third approach is a non-parametric strategy that focuses on estimation of the parameters leaving the baseline hazard <img src="10-1240113\e2d73d30-1673-4894-ac48-f70903ed5aeb.jpg" /> completely unspecified. This approach relies on a partial likelihood function proposed by Cox [<xref ref-type="bibr" rid="scirp.30719-ref19">19</xref>].</p><p>Five ways of estimating the time to event ignoring the effects of covariates are considered initially. Standard Kaplan-Meier estimator is used first. It is assumed that exact times of event are known and this is done either by assuming the event occurred at the left interval, or at the right interval. These two extreme cases should roughly bracket the estimates derived using the interval-censoring methods. A second approach is using a Weibull model, where the survivals function is modeled using the estimates from the SAS Proc LIFEREG. A Third procedure is to model the interval-censored nature of the data using the techniques proposed by Turnbull. The fourth is to use splines models proposed by Kooperberg and Stone [<xref ref-type="bibr" rid="scirp.30719-ref18">18</xref>]. Finally, the survival function is estimated using the piecewise exponential model.</p><sec id="s3_1"><title>3.1. Breast Cancer Data</title><p>This data set is from a retrospective study of patients with breast cancer designed to compare radiation therapy alone versus in combination with chemotherapy with respect to the time to cosmetic deterioration. This data set has been analyzed by several authors to illustrate various methods for interval censored data. Patients were seen initially every 4 to 6 months, with decreasing frequency over time. If deterioration was seen, it was known only to have occurred between two visits. Deterioration was not observed in all patients during the course of the trial, so some data were right-censored.</p><p>The breast cancer data set is described in detail in Finkelstein and Wolfe [<xref ref-type="bibr" rid="scirp.30719-ref3">3</xref>] and it consists of a total of 94 observations from a retrospective study looking at the time to cosmetic deterioration. Information is available on one covariate, type of therapy, either radiation alone (coded 0), or in combination with chemotherapy (coded 1). Of the 94 observations, 56 are interval-censored and 38 are right-censored. All estimated curves for the Breast Cancer data are presented in <xref ref-type="fig" rid="fig1">Figure 1</xref>. KM for R and KM for L represent Kaplan Meier estimates for right censored and left censored data respectively in Figures 1-3.</p></sec><sec id="s3_2"><title>3.2. AIDS Data</title><p>This data set focused on the development of drug resistance (measured using a plaque reduction assay) to zidovudine in patients enrolled in four clinical trials for the treatment of AIDS. Samples were collected on the patients at a subset of the scheduled visit times dictated by the four protocols. Since the resistance assays were very expensive, there were few assessments on each patient, resulting in very wide intervals, <img src="10-1240113\89949f6f-dbea-4342-bde4-058dd94e253e.jpg" />, if resistance was seen to have occurred, and a high proportion of right-censored observations. Because of the sparseness of these data, this is a challenging data set to analyze. The variables of interest were the effects of stage of disease,</p><p>dose of zidovudine and CD4 lymphocyte counts at time of randomization on the time to development of resistance. All estimated curves for the AIDS data are presented in <xref ref-type="fig" rid="fig2">Figure 2</xref>.</p><p>Lindsay and Ryan [<xref ref-type="bibr" rid="scirp.30719-ref1">1</xref>] have presented both Breast Cancer and AIDS data sets. More information about these two data sets could be found there. Some of the analyticcal methods, notations and results presented in this paper are similar to their paper.</p></sec><sec id="s3_3"><title>3.3. Hemophilia Data</title><p>In 1978, 262 persons with Type A or B hemophilia have been treated at Hospital Kremlin Bicetre and Hospital Coeur des Yvelines in France. Twenty-five of the hemophiliacs were found to be infected with HIV on their first test for infection. By August 1988, 197 had become infected and 43 of these had developed clinical symptoms (AIDS, lymphadenopathy, or leukopenia) relating to their HIV infection. All of the infected persons were believed to have become infected by contaminated blood factor received for their hemophilia. The observations for the 262 patients were based on a discretization of the time axis into 6-month intervals. Here time is measured in 6-month intervals, with L = 1 denoting July 1, 1978, and Z denoting chronologic time of first clinical symptom. The 25 hemophiliacs infected at entry are assigned L = 1. Victor and Stephen [<xref ref-type="bibr" rid="scirp.30719-ref20">20</xref>] have presented this data in their paper and more information could be found there. They have initially analyzed Hemophilia data considering this as a Doubly-Censored Survival Data. However, we have analyzed this data set taking left censoring, right censoring and interval censoring into consideration. All estimated curves for the Hemophilia data are presented in <xref ref-type="fig" rid="fig3">Figure 3</xref>.</p><p>For the breast cancer example, the Kaplan-Meier estimates, bracket the Turnbull estimate. The Turnbull curve lies very close to both Weibull and logspline curves. At the same time, estimates from the Weibull, and logspline estimates are quite close to each other. The piecewise curve does not properly fall within the Kaplan-Meier estimates.</p><p>The estimated survival curve for the AIDS data took very few steps in the non-parametric models, which reflected the high degree of censoring in this small data set. The Kaplan-Meier estimates no longer bracketed the Turnbull estimate, mainly because the Turnbull estimate had very few jumps due to the particular configuration of this data set. The logspline estimate also tracked the parametric models closely. The non-parametric methods were not very helpful in understanding the AIDS data.</p><p>For the Hemophilia data, the results were quite similar to breast cancer data except the logspline model. Results derived for the piecewise exponential are not accurate for all three data sets.</p></sec><sec id="s3_4"><title>3.4. Covariate Effects on Time to Event</title><p>To compare the two treatments, for Breast cancer data a retrospective study of 46 radiation only and 48 radiation plus chemotherapy patients was conducted. Using Turnbull’s algorithm the estimated survival functions were obtained for radiotherapy only and radiation plus chemotherapy groups respectively, which are shown in <xref ref-type="fig" rid="fig4">Figure 4</xref>. Note that the estimated survival curves did not show striking differences from 0 to 18 months. From 18 onwards, however, a fast decay of the curve is seen for patients given radiotherapy plus chemotherapy. Note, for instance, that only 11.06% of the patients in the radiotherapy plus chemotherapy group were estimated to be free of any evidence of breast retraction at time t = 40 months against 47.37% in the radiotherapy group.</p><p>Using the midpoint of each interval, is a common practice among analysts due to the lack of well-known statistical methodology and available software. Then applying the Kaplan-Meier method, we obtained the estimated survival curves presented in <xref ref-type="fig" rid="fig5">Figure 5</xref>. The curves estimated previously are also shown in the graph. Comparing the curves we can see that the estimates obtained using both, the midpoints and the intervals, are very similar to each other at several times but they tend to be under or over estimated at others. Although not shown here, under or over estimation became more evident if it is assumed that the event occurred to the end or at the beginning of each interval instead of at the midpoint. The range of each interval also contributes for the magnitude of these differences. They are more accentuated as the range of each interval increases.</p><p>Positive parameter estimates in Cox regression indicate higher failure rates for individuals with larger values of the covariate. The exponential model parameter should be of comparable magnitude to the Cox model, but with the sign reversed. The Weibull is the only family of models that is both proportional hazards and AFT. It can easily be shown that the estimated regression coefficient should be comparable to the coefficients from the Cox.</p><p>Results using the Cox regression models assuming exact event times (taken to be the left, midpoint and right extremes of the interval), and based on the exponential, Weibull and log-normal models for the breast cancer data is shown in <xref ref-type="table" rid="table1">Table 1</xref>. Parameter estimates of each of the above models considered, standard errors and P-values obtained for these different models are presented in this Table. All four analyses give similar results about the treatment comparison and suggest that the adjuvant chemotherapy significantly increases the risk of breast retraction. Note that the Cox analysis has the minimal impact from the differing assumptions about timing of events.</p><p>Cox models are fitted with left end point, middle point and right end point. Results obtained for fitting these models for the AIDS data with all four covariates, stage, dose, CD4<sub>1</sub> and CD4<sub>2</sub> are shown in <xref ref-type="table" rid="table2">Table 2</xref>. The results differed from those obtained from the Breast Cancer data, although none of the estimated covariate effects are significant except in two instances. Possible explanations are that the sample size is quite small, and the observed information is very limited due to interval-censoring. Another reason could be that the covariates are correlated. The last two covariates, the indicators of CD4 count, CD4<sub>1</sub> and CD4<sub>2</sub> are correlated and both are also correlated with the other covariates the stage of the disease. Therefore, the last two covariates CD4<sub>1</sub> and CD4<sub>2</sub> are removed, and the new analysis results are presented in <xref ref-type="table" rid="table3">Table 3</xref> along with the other results for exponential, Weibull, Log-Normal and piecewise exponential.</p><p>Now we take a look at the individual effects of stage and dose on the time to development of resistance. For stage, all methods indicate an increased risk of developing resistance for the patients in a later stage of disease. The non-parametric methods do not perform as well as the parametric methods. Unlike the breast cancer data, changing assumptions about when events are assumed to occur has a big impact on the Cox analysis. The strength of the significance is also affected when the midpoint or right extreme of the interval is used as the exact event time. With such large effects, using a method which accounts for the interval-censored nature of the data is preferable, but with so few steps in the survival curve using the non-parametric methods, a parametric analysis is the best choice. Similar trends are seen in fitting the effect of dose. The Cox model results are highly dependent on the assumptions about when the event oc-</p><p><xref ref-type="table" rid="table1">Table 1</xref>. Breast cancer data: Effect of therapy on time to event.</p><p><img src="10-1240113\4f383bf8-c583-407f-b34e-43a1792b307f.jpg" /></p><p><xref ref-type="table" rid="table2">Table 2</xref>. Aids data: Effect of stage of disease, dose of zidovudine, CD4<sub>1</sub> and CD4<sub>2</sub>.<sub></sub></p><disp-formula id="scirp.30719-formula16707"><graphic  xlink:href="10-1240113\5ce9c484-5862-446d-bb92-a03a52eeff51.jpg"  xlink:type="simple"/></disp-formula><p><xref ref-type="table" rid="table3">Table 3</xref>. Aids data: Effect of stage of disease and dose of zidovudine.</p><p><img src="10-1240113\a2643eb2-7585-415e-a604-456b66175ce8.jpg" /></p><p>curred.</p><p>Interval-censored data often occur in medical applications. As seen in the AIDS data set, when data are heavily censored, making assumptions about when events occurred and using techniques such as Cox regression can lead to inaccurate conclusions. It can also result in unstable estimation in the non-parametric methods. The parametric methods available in SAS, S-Plus and R are the most readily available alternatives. As seen with the examples presented in this paper, these parametric approaches can be highly satisfactory in their performance. This is especially so if one chooses the Weibull or lognormal family that allows a reasonably wide range of distributional shapes.</p><p>Results using the Cox regression models, the exponential, Weibull and log-normal models for the Hemophilia data are shown in <xref ref-type="table" rid="table4">Table 4</xref>. Parameter estimates of each of the above models considered, standard errors and Pvalues obtained for these different models are presented in this Table. All of these four analyses show significant evidence of treatment effects.</p></sec></sec><sec id="s4"><title>4. Conclusions and Further Work</title><p>Interval censoring has become increasingly common in the areas that produce failure time data. This type of data frequently comes from tests or situations where the objects of interests are not constantly monitored. Thus events are known only to have occurred within particular time durations. The purpose of this study was to illustrate available parametric and non-parametric methods that consider the data as being interval censored.</p><p>The time to event ignoring the effects of covariates has been considered using five different techniques. The Kaplan-Meier estimator is used, assuming the event occurred at the left interval, or at the right interval. A Weibull model, where the survival function is modeled using the estimates from SAS PROC LIFEREG is carried out. A Third approach is accomplished by modeling the interval-censored nature of the data using the methods proposed by Turnbull [<xref ref-type="bibr" rid="scirp.30719-ref8">8</xref>]. Splines models presented by Kooperberg and Stone [<xref ref-type="bibr" rid="scirp.30719-ref18">18</xref>] are implemented. Finally, the survival function is estimated using the piecewise exponential model. Parametric models for interval-censored data can follow a number of distributions such as generalized gamma, the log-normal, the Weibull and the exponential distribution. Different independent covariates or categorical variables have also been included in the model to study their effect on the response variable.</p><p>Parametric and non-parametric methods of analysis are two different types of techniques in general for the analysis of censored data. However, for the analysis of interval-censored data, the terminology behind them is the same. An important advantage of parametric inference approaches is that their implementation is quite straightforward in principle and the theory of standard maximum likelihood can be applied. A primary disadvantage of these methods is that there often does not exist enough prior information or data to verify a parametric model. The major advantage of non-parametric methods such as Kaplan-Meier and Turnbull approach is that, one can avoid complicated interval censoring issues and make use of the existing inference procedures for the rightcensored data. It is also assumed that the censoring mechanism or variables are independent of the survival variables of interest.</p><p>For the Breast Cancer example, the Kaplan-Meier estimates, bracket the Turnbull estimate. The Turnbull curve lies very close to both Weibull and logspline curves. At the same time, Weibull estimates and logspline estimates are quite close to each other. Cox regression models and parametric models with covariates using Exponential, Weibull and Lognormal were fitted. Four analyses produced similar results qualitatively and all showed an increased hazard for group on radiation and chemotherapy which was statistically significant.</p><p>Unlike the results obtained for the Breast Cancer data, the estimated survival curve for AIDS data took very few steps in the non-parametric models. Possible explanations could be that the sample size is small, and the observed information was very limited due to interval cen</p><p><xref ref-type="table" rid="table4">Table 4</xref>. Hemophilia data: Effect of treatment on time to event.</p><p><img src="10-1240113\9c6c70c1-e90a-43c2-94f1-761e05d5109e.jpg" /></p><p>soring. The Kaplan-Meier estimates no longer bracketed the Turnbull estimate. The logspline estimate tracks the parametric models closely. The non-parametric methods were not very helpful in understanding the AIDS data. For the covariate effects of time to development of resistance, stage and dose were taken into account. All methods indicated an increased risk of developing resistance for the patients in a later stage of disease. The Cox model results were highly dependent on the assumptions about when the event occurred. No methods showed a significant effect of dose on the time to development of resistance.</p><p>The estimated survival curve for the AIDS data took very few steps in the non-parametric models, which reflected the high degree of censoring in this small data set. For the AIDS data, we took a look at the individual effects of stage and dose on the time to development of resistance. For these data, the non-parametric methods were not very helpful in understanding this data. For the breast cancer data, the four analyses gave similar results qualitatively. All showed an increased hazard for the group on radiation and chemotherapy which was statistically significant. Kaplan-Meier estimates, as expected, bracket all other survival curves.</p><p>For the Hemophilia data, the results derived were very much similar to those of Breast Cancer and AIDS data sets except the logspline model. Cox regression models and parametric models with covariates using Exponential, Weibull and Lognormal were fitted. Four analyses produced similar results qualitatively and all have shown an increased hazard for the group on radiation and chemotherapy, which is statistically significant. Results derived for the piecewise exponential are not accurate for all three data sets.</p><p>Major statistical packages such as SAS, S-Plus and R have procedures for analyzing interval-censored data using parametric models. Some non-parametric methods are easily programmed. In particular, the Turnbull [<xref ref-type="bibr" rid="scirp.30719-ref8">8</xref>] method for non-parametric estimation of the survival distribution, the Kooperburg and Stone [<xref ref-type="bibr" rid="scirp.30719-ref18">18</xref>] logspline estimates of the survival function and Finkelstein’s (1986) test for covariates are recommended. As seen in the AIDS data set, when data are heavily censored, making assumptions about when events occurred and using techniques such as Cox regression can lead to inaccurate conclusions. It can also result in unstable estimation in the non-parametric methods.</p><p>As examples presented in this paper show, parametric approaches can be highly satisfactory in their performance. Especially when the Weibull or log-normal family is chosen, it allows a reasonably wide range of distributional shapes. To allow more flexible modeling with weak parametric assumptions, we suggest the use of a piecewise constant hazards model. Finkelstein and Wolfe [<xref ref-type="bibr" rid="scirp.30719-ref3">3</xref>], Self and Grossman [<xref ref-type="bibr" rid="scirp.30719-ref21">21</xref>], Miller [<xref ref-type="bibr" rid="scirp.30719-ref22">22</xref>] and Buckley and James [<xref ref-type="bibr" rid="scirp.30719-ref23">23</xref>] have all proposed tests for assessing the covariate effects.</p></sec><sec id="s5"><title>REFERENCES</title></sec><sec id="s6"><title>NOTES</title></sec></body><back><ref-list><title>References</title><ref id="scirp.30719-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">J. C. Lindsey and L. M. Ryan, “Tutorial in Biostatistics Methods for Interval-Censored Data,” Statistics in Medicine, Vol. 17, No. 2, 1998, pp. 219-238. 
doi:10.1002/(SICI)1097-0258(19980130)17:2&lt;219::AID-SIM735&gt;3.0.CO;2-O</mixed-citation></ref><ref id="scirp.30719-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">J. K. Lindsey, “A Study of Interval Censoring in Parametric Regression Models,” Life Time Data Analysis, Vol. 4, No. 4, 1998, pp. 329-354. 
doi:10.1023/A:1009681919084</mixed-citation></ref><ref id="scirp.30719-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">D. M. Finkelstein and R. A. Wolfe, “A Semi-Parametric Model for Regression Analysis of Interval Censored Failure Time Data,” Biometrics, Vol. 41, No. 4, 1985, pp. 933-945. doi:10.2307/2530965</mixed-citation></ref><ref id="scirp.30719-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">D. M. Finkelstein, “A Proportional Hazards Model for Interval-Censored Failure Time Data”, Biometrics, Vol. 42, No. 4, 1986, pp. 845-854. doi:10.2307/2530698</mixed-citation></ref><ref id="scirp.30719-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">R. Peto, “Experimental Survival Curves for Interval-Censored Data,” Applied Statistics, Vol. 22, No. 1, 1973, pp. 86-91. doi:10.2307/2346307</mixed-citation></ref><ref id="scirp.30719-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">P. S. Rosenberg, “Hazard Function Estimation Using B-Splines,” Biometrics, Vol. 51, No. 3, 1995, pp. 874-887. 
doi:10.2307/2532989</mixed-citation></ref><ref id="scirp.30719-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">P. M. Odell, K. M. Anderson and R. B. Agostino, “Maximum Likelihood Estimation for Interval-Censored Data Using a Weibull-Based Accelerated Failure Time Model,” Biometrics, Vol. 48, No. 3, 1992, pp. 951-959. 
doi:10.2307/2532360</mixed-citation></ref><ref id="scirp.30719-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">B. W. Turnbull, “The Empirical Distribution Function with Arbitrarily Grouped, Censored and Truncated Data”, Journal of the Royal Statistical Society, Series B, Vol. 38, No. 3, 1976, pp. 290-295.</mixed-citation></ref><ref id="scirp.30719-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">C. P. Farrington, “Interval Censored Survival Data: A Generalized Linear Modeling Approach,” Statistics in Medicine, Vol. 15, No. 3, 1996, pp. 283-292. 
doi:10.1002/(SICI)1097-0258(19960215)15:3&lt;283::AID-SIM171&gt;3.0.CO;2-T</mixed-citation></ref><ref id="scirp.30719-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">E. Goetghebeur and L. Ryan, “Semi-Parametric Regression Analysis of Interval-Censored Data,” Biometrics, Vol. 56, No. 4, 2000, pp. 1139-44.</mixed-citation></ref><ref id="scirp.30719-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">J. F. Lawless, “Statistical Models and Methods for Lifetime Data,” Wiley, 2003.</mixed-citation></ref><ref id="scirp.30719-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">D. Collett, “Modelling Survival Data in Medical Research,” Chapman and Hall, London, 2003.</mixed-citation></ref><ref id="scirp.30719-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">J. Sun “The Statistical Analysis of Interval-Censored Failure Time Data,” Springer, New York/Heidelberg, 2006.</mixed-citation></ref><ref id="scirp.30719-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">E. L. Kaplan and P. Meier, “Nonparametric Estimation from Incomplete Observations,” Journal of the American Statistical Association, Vol. 53, No. 282, 1958, pp. 457-481. doi:10.1080/01621459.1958.10501452</mixed-citation></ref><ref id="scirp.30719-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">G. Rucker and D. Messerer, “Remission Duration: An Example of Interval-Censored Observations,” Statistics in Medicine, Vol. 7, No. 11, 1988, pp. 1139-1145. 
doi:10.1002/sim.4780071106</mixed-citation></ref><ref id="scirp.30719-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">F. J. Dorey, R. J. Little and N. Schenker, “Multiple Imputation for Threshold-Crossing Data with Interval Censoring”, Statistics in Medicine, Vol. 12, No. 17, 1993, pp. 1589-1603. doi:10.1002/sim.4780121706</mixed-citation></ref><ref id="scirp.30719-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">R. Gentleman and C. J. Geyer, “Maximum Likelihood for Interval Censored Data: Consistency and Computation,” Biometrika, Vol. 81, No. 3, 1994, pp. 618-623. 
doi:10.1093/biomet/81.3.618</mixed-citation></ref><ref id="scirp.30719-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">C. Kooperberg and C. J. Stone, “Logspline Density Estimation for Censored Data,” Journal of Computational and Graphical Statistics, Vol. 1, No. 4, 1992, pp. 301-328.</mixed-citation></ref><ref id="scirp.30719-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">D. R. Cox, “Regression Models and Life Tables (with Discussion),” Journal of the Royal Statistical Society, Series B, Vol. 34, No. 2, 1972, pp. 187-220.</mixed-citation></ref><ref id="scirp.30719-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">V. De Gruttola and S. W. Lagakos, “Analysis of Doubly-Censored Survival Data, with Application to AIDS,” Biometrics, Vol. 45, No. 1, 1989, pp. 1-11. 
doi:10.2307/2532030</mixed-citation></ref><ref id="scirp.30719-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">S. G. Self and E. A. Grossman, “Linear Rank Tests for Interval-Censored Data with Application to PCB levels in Adipose Tissue of Transformer Repair Workers,” Biometrics, Vol. 42, No. 3, 1996, pp. 521-530. 
doi:10.2307/2531202</mixed-citation></ref><ref id="scirp.30719-ref22"><label>22</label><mixed-citation publication-type="other" xlink:type="simple">R. G. Miller, “Least Squares Regression with Censored Data,” Biometrika, Vol. 63, No. 3, 1976, pp. 447-464. 
doi:10.1093/biomet/63.3.449</mixed-citation></ref><ref id="scirp.30719-ref23"><label>23</label><mixed-citation publication-type="other" xlink:type="simple">J. Buckley and I. James, “Linear Regression with Censored Data,” Biometrika, Vol. 66, No. 3, 1979, pp. 429-436. 
doi:10.1093/biomet/66.3.429</mixed-citation></ref></ref-list></back></article>