<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">OJS</journal-id><journal-title-group><journal-title>Open Journal of Statistics</journal-title></journal-title-group><issn pub-type="epub">2161-718X</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/ojs.2021.111012</article-id><article-id pub-id-type="publisher-id">OJS-107493</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Maximum Likelihood Estimation for the Pooled Repeated Partly Interval-Censored Observations Logistic Regression Model
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Naghmeh</surname><given-names>Daneshi</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref></contrib><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Jong</surname><given-names>Sung Kim</given-names></name><xref ref-type="aff" rid="aff1"><sup>1</sup></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><addr-line>Fariborz Maseeh Department of Mathematics and Statistics, Portland State University, Portland, OR, USA</addr-line></aff><pub-date pub-type="epub"><day>04</day><month>01</month><year>2021</year></pub-date><volume>11</volume><issue>01</issue><fpage>230</fpage><lpage>242</lpage><history><date date-type="received"><day>22,</day>	<month>December</month>	<year>2020</year></date><date date-type="rev-recd"><day>23,</day>	<month>February</month>	<year>2021</year>	</date><date date-type="accepted"><day>26,</day>	<month>February</month>	<year>2021</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
  Often in longitudinal studies, some subjects complete their follow-up visits, but others miss their visits due to various reasons. For those who miss follow-up visits, some of them might learn that the event of interest has already happened when they come back. In this case, not only are their event times interval-censored, but also their time-dependent measurements are incomplete. This problem was motivated by a national longitudinal survey of youth data. Maximum likelihood estimation (MLE) method based on expectation-maximization (EM) algorithm is used for parameter estimation. Then missing information principle is applied to estimate the variance-covariance matrix of the MLEs. Simulation studies demonstrate that the proposed method works well in terms of bias, standard error, and power for samples of moderate size. The national longitudinal survey of youth 1997 (NLSY97) data is analyzed for illustration.
 
</p></abstract><kwd-group><kwd>EM Algorithm</kwd><kwd> Longitudinal Studies</kwd><kwd> Louis’ Method</kwd><kwd> Partly Interval-Censored Failure Time Data</kwd><kwd> Pooled Repeated Observations</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>In longitudinal studies, subjects who are likely to progress to a new state during the study are monitored over time. For example, in clinical trials, subjects who are at high risk of a certain disease are monitored and have follow-up visits. Some subjects complete all of their follow-up visits and their failure times are recorded. However, others miss their follow-up visits, and they may learn that the event of interest had already occurred when they came back. The event times for these patients are censored within the corresponding person-specific time intervals. Although there are multiple follow-up visiting intervals for each subject, researchers often use one particular interval that contains the true unknown failure time unless they had accurately determined the failure time. This is known as “partly interval-censored failure time data”. There are quite a few research works based on partly interval-censored data such as [<xref ref-type="bibr" rid="scirp.107493-ref1">1</xref>] [<xref ref-type="bibr" rid="scirp.107493-ref2">2</xref>] [<xref ref-type="bibr" rid="scirp.107493-ref3">3</xref>] and [<xref ref-type="bibr" rid="scirp.107493-ref4">4</xref>] among others.</p><p>Another commonly available data type in longitudinal studies is called pooled repeated observations. Subjects have multiple follow-up visits as usual. From every visit, a subject obtains a binary outcome for the event of interest. All those repeated binary outcomes are pooled together to develop a model to analyze the effects of time-dependent covariates on the occurrence of the event. [<xref ref-type="bibr" rid="scirp.107493-ref5">5</xref>] and [<xref ref-type="bibr" rid="scirp.107493-ref6">6</xref>] pooled such repeated observations with binary outcomes for the event of interest into a single sample. Then they used logistic regression model to estimate the effects of the risk factors on the occurrence of the event. Each observation interval is considered a mini follow-up study in which the current risk factors are updated to predict events in the interval. Once an individual has an event in a particular interval, all subsequent intervals from that individual are excluded from the analysis.</p><p>Now, we define pooled repeated partly interval-censored data. We have pooled repeated observations, but some binary outcomes and covariates are incomplete. They can only be determined with certain unknown probabilities within the corresponding specific follow-up visits. In this case, the analysis of such data requires a new method that combines a model that handles pooled repeated observations without censoring and a method that deals with partly or completely interval-censored data.</p><p>The main goal of this study is to estimate the effects of the time-dependent covariates on the occurrence of the event of interest (e.g., progression to a disease, becoming a frequent smoker, etc.). We extend the work of [<xref ref-type="bibr" rid="scirp.107493-ref7">7</xref>], who employed conditional expected score test (CEST) to determine the presence of association of a longitudinal marker and an event with missing binary outcomes to the estimation problem when the event of interest has a single progression state and the response is pooled, repeated, and partly interval-censored. We assume that the missing data is missing at random (MAR). In MAR data, there might be systematic differences between the observed and missing data, but the differences can be explained by the observed data. EM algorithm was originally developed to handle MAR data.</p><p>The organization of this paper is as follows. In Section 2, we present a logistic regression model for pooled repeated partly interval-censored data. In Section 3, we provide the details of computation of the MLEs of the regression parameter via EM algorithm and the variance estimation through the missing information principle. Section 4 displays the simulation study results. Section 5 illustrates an application to a real data set. Finally, Section 6 briefly summarizes what we have achieved and also discusses potential extensions of our work.</p></sec><sec id="s2"><title>2. Model</title><p>We consider a case of longitudinal studies, where subjects are at risk of an event of interest and have follow-up visits. Some subjects make complete follow-up visits, but others miss some of their follow-up appointments and come back after the event of interest has occurred. Whenever they miss a visit, both their binary outcome of the event of the interest and covariates are missing. Our proposed model estimates the effects of time-dependent covariates on the event of interest.</p><p>Let T i be the time subject i experiences the event of interest, i = 1, ⋯ , n . At the beginning of the study, every subject is assigned to the same follow-up visits, t j , j = 1, ⋯ , M . Let y i j be the indicator of whether or not subject i has had the event of interest in the jth interval given a subject was event-free through t j − 1 and x i j , the covariate at time t j − 1 . Since we are interested in modeling a binary outcome, we use a logit link to model the probability of event as in [<xref ref-type="bibr" rid="scirp.107493-ref7">7</xref>].</p><p>logit ( p i j ) = log ( p i j / ( 1 − p i j ) ) = α + β ′ x i j , (1)</p><p>where</p><p>p i j = P ( y i j = 1 | x i j , T i &gt; t j − 1 ) . (2)</p><p>We construct the full (complete) log-likelihood, assuming as if there were no missing visits while subjects are in the study.</p><p>l = ∑ i = 1 n   ∑ j = 1 M i [ − log ( 1 + exp ( α + β ′ x i j ) ) + y i j ( α + β ′ x i j ) ] , (3)</p><p>where M i is the index of the last time subject i was in the study.</p></sec><sec id="s3"><title>3. Methods</title><sec id="s3_1"><title>3.1. Parameter Estimation</title><p>Assume that the i<sup>th</sup> subject missed visits after time t L i and came back at t R i . L i is the index of the last time subject i made the visit and was event-free. R i is the index of the first time subject i was observed with the event of interest. Then y i L i = 0 , y i R i = 1 , and y i j is missing for L i + 1 ≤ j ≤ R i − 1 . For the subjects who do not miss visits, L i + 1 = R i . Whenever subjects miss visits, their covariate value, x i j , is also missing. We use the EM algorithm ( [<xref ref-type="bibr" rid="scirp.107493-ref8">8</xref>] ) to estimate the parameters.</p><p>E-step: For individuals whose failure times are interval-censored, we need to estimate both y i j and x i j in the expression (3) for j ∈ { L i + 1, ⋯ , R i − 1 } .</p><p>x i j could be continuous or categorical ( [<xref ref-type="bibr" rid="scirp.107493-ref9">9</xref>] ). We assume that x i j has a linear growth curve with fixed effects to incorporate a real data, NLSY97. That is,</p><p>x i j = θ 0 i + θ 1 i t j − 1 + ε i j , (4)</p><p>where ε i j ∼ N ( 0, σ ε 2 ) , c o v ( ε i j , ε i j ′ ) = 0, j ≠ j ′ . We estimate x i j by x ^ i j = θ ^ 0 i + θ ^ 1 i t j − 1 for L i + 1 ≤ j ≤ R i − 1 , where θ ^ 0 i and θ ^ 1 i are least squares estimators.</p><p>If x i j is ordinal, we assign numbers to corresponding categories. Then we again assume linear growth curve with fixed effects to estimate the missing x i j ’s. Let n c be the number of categories for this ordinal variable. For each individuali, the observed x i j ’s are used in model (4) to compute θ ^ 0 i and θ ^ 1 i . Then we compute x ^ i j = θ ^ 0 i + θ ^ 1 i t j − 1 as usual.</p><p>Next, we create n c − 1 thresholds in order to uniquely assign x ^ i j into one of the n c categories. Note that x ^ i j ∼ N ( θ 0 i + θ 1 i t j − 1 , σ x ^ i 2 ) and x ^ i = { x ^ i , L i + 1 , ⋯ , x ^ i , R i − 1 , x i , R i } . We use the quantiles of this normal distribution to define the thresholds. Since we need to compute σ ^ x ^ i 2 to define thresholds, we need at least three distinct observed covariate values, x i j ’s for each subject, otherwise, σ ^ x ^ i 2 would be undefined due to the zero degrees of freedom.</p><p>The observed ordinal covariates for some subjects do not include the entire ordinal categories. Therefore, the ordinal logistic regression model does not work for estimating ordinal covariates. In Appendix 2, we provide a detailed rationale for choosing fixed effects model, its extension in a general setting, and challenges with random effects model.</p><p>y ^ i j = E [ y i j | Y i , x ^ i , α , β , T i &gt; t j − 1 ] = P [ T i = t j | t L i &lt; T i ≤ t R i , Y i , x ^ i , α , β ] = ( p ^ i j p ^ i , L i + 1 + ∑ k = L i + 2 R i − 1 [ p ^ i k ∏ o = L i + 1 k − 1 ( 1 − p ^ i o ) ] + ∏ o = L i + 1 R i − 1 ( 1 − p ^ i o )     if   j = L i + 1 , p ^ i j ∏ o = L i + 1 j − 1 ( 1 − p ^ i o ) p ^ i , L i + 1 + ∑ k = L i + 2 R i − 1 [ p ^ i k ∏ o = L i + 1 k − 1 ( 1 − p ^ i o ) ] + ∏ o = L i + 1 R i − 1 ( 1 − p ^ i o )     if   j ∈ { L i + 2 , ⋯ , R i − 1 } , ∏ o = L i + 1 j − 1 ( 1 − p ^ i o ) p ^ i , L i + 1 + ∑ k = L i + 2 R i − 1 [ p ^ i k ∏ o = L i + 1 k − 1 ( 1 − p ^ i o ) ] + ∏ o = L i + 1 R i − 1 ( 1 − p ^ i o )     if   j = R i , (5)</p><p>where</p><p>p ^ i j = exp ( α + β ′ x ^ i j ) 1 + exp ( α + β ′ x ^ i j ) (6)</p><p>This is an extension of a geometric-type experiment, where the probability of success (progression) changes at each follow-up visit, t j , L i + 1 ≤ j ≤ R i .</p><p>M-step: We find the values of α and β that maximize the expected value of log-likelihood in Equation (3), conditioned on the observed data. Therefore, we have</p><p>( α ^ , β ^ ) = arg max l α , β | y ^ i j , x ^ i j , (7)</p><p>where y ^ i j = y i j , x ^ i j = x i j , if uncensored.</p><p>Expressions (5)-(7) are repeated until convergence. As there are no closed forms for α ^ and β ^ , we used an optimization package optim in R to obtain ( α ^ , β ^ ) .</p></sec><sec id="s3_2"><title>3.2. Variance Estimation</title><p>We apply Louis’ method for variance estimation using the notation in [<xref ref-type="bibr" rid="scirp.107493-ref10">10</xref>]. Following the missing information principle, we compute the observed information by subtracting the missing information from the complete information.</p><p>− ∂ 2 log P ( θ | W ) ∂ θ 2 = − ∫ v ∂ 2 log P ( θ | W , V ) ∂ θ 2 P ( V | θ , W ) d V       − V a r ( − ∂ log P ( θ | W , V ) ∂ θ ) , (8)</p><p>where W is observed data, i.e., partly interval-censored pooled repeated observations. V is latent data, the true unknown counterpart of the interval-censored portion of W. θ | W is the observed posterior and θ | W , V is the augmented posterior.</p><p>The details of the expression (8) are provided in Appendix 3.</p></sec></sec><sec id="s4"><title>4. Simulation Study</title><sec id="s4_1"><title>4.1. Data Simulation</title><p>We considered n = 300 subjects who have M = 7 follow-up visits each. We generated covariates as follows:</p><p>x 1 i j ∼ N ( 5.8 + 0.3 t j − 1 ,0.1 ) .</p><p>x 2 i j ∼   N ( 0.4 + 0.15 t j − 1 ,0.1 ) .</p><p>x 1 i j represents a continuous covariate with larger values and faster growth rate over time, while x 2 i j represents one with smaller values and slower growth rate over time.</p><p>First, we generate n = 300 subjects who have complete follow-up visits. This makes the original complete data (OC), pooled repeated data. We randomly choose n 1 subjects out of these. This makes the exact data (E), a proper subset of the OC. For the remaining n 2 = n − n 1 subjects, we randomly designate some of their follow-up visits missing. This makes the pooled repeated interval-censored observations. The observed data (O), which is the pooled repeated partly interval-censored data, is the mix of pooled repeated data (E) and pooled repeated interval-censored data. We considered several values for n 1 and n 2 to cover different proportions of exact data.</p><p>We randomly sampled L i and R i for each patient. Note that for the exact data, we have R i = L i + 1 and for the pooled repeated interval-censored data, R i ≥ L i + 2 . Then for j = 1, ⋯ , L i , we have y i j = 0 and for j = R i , ⋯ , M , we have y i j = 1 . y i j is missing for j = L i + 1, ⋯ , R i − 1 in the pooled repeated interval-censored data. y i j is 1 when the i<sup>th</sup> subject at risk at the j<sup>th</sup> visit experiences the event of interest in thej<sup>th</sup> interval.</p><p>We computed the bias and variance for original complete data, exact data, and observed data based on B = 1500 replications. In addition, we investigated the power of our test.</p></sec><sec id="s4_2"><title>4.2. Results</title><p>We first considered the case where there was only one attribute in the model. The EM algorithm (Section 3.1) was used for the parameter estimation. The variance of the parameter estimator was calculated using Louis’ method (Section 3.2).</p><p>The results are shown in <xref ref-type="table" rid="table1">Table 1</xref>. For all the different combinations of n 1 and n 2 , the proposed estimator based on the observed data produces a smaller bias and a smaller variance than that based on the exact data alone. In particular, for the case of (250, 50), containing E 84% (250) and only 16% (50) pooled-repeated interval-censored data, the proposed estimator produces a smaller bias and a smaller variance than that based on E alone. We also notice that the more exact data we have, the smaller bias and variance we get. These results have a quite similar pattern to those in [<xref ref-type="bibr" rid="scirp.107493-ref3">3</xref>], who employed a proportional hazards model with partly interval-censored data. [<xref ref-type="bibr" rid="scirp.107493-ref6">6</xref>] notes that pooled repeated observations logistic regression is close to the time-dependent covariate Cox regression analysis. Therefore, this simulation result coincides with what we expected. In order to see if bootstrap would be of help, we also ran simulations with various pairs of n 1 and n 2 to compare the bootstrap variance with the variances for the O, E, and OC. We considered two covariates; one is continuous and the other is ordinal. <xref ref-type="table" rid="table2">Table 2</xref> shows the results. For all pairs of n 1 and n 2 , the bootstrap variance for the O is smaller than that for OC, which is supposed to be the smallest. This is, bootstrapping suffers from substantial underestimation. Therefore, we do not recommend it for this setting. Another issue is that it is time-consuming.</p><p>Next, we computed the power of the test H 0 : β = β 0 vs. H 1 : β ≠ β 0 . We considered both one-dimensional covariate and two-dimensional covariates. We considered 3 sample sizes (100, 200, and 300), and for each of these sample sizes we ran B = 1000 replications of the test. The power was calculated as the proportion of times H 0 was rejected at 5% level of significance. Both <xref ref-type="fig" rid="fig1">Figure 1</xref> and <xref ref-type="fig" rid="fig2">Figure 2</xref> show the powers for different values of β 0 and different sample sizes. The power curves are symmetric for all the different sample sizes. As a sample size increases or the parameter values are farther apart from the true parameter value (i.e., an effect size increases), the corresponding power increases. From <xref ref-type="fig" rid="fig1">Figure 1</xref>, with a sample of size n = 300 , one can achieve 80% power for the effect size of 0.45. Moreover, for the effect size of 0.55, a sample of size n = 200 is enough to achieve 80% power. [<xref ref-type="bibr" rid="scirp.107493-ref11">11</xref>] achieved approximately 80% power in</p><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Results for 1-dimensional β , β t r u e = 3.6 , B: Bias, σ 2 : variance</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >( n 1 , n 2 )</th><th align="center" valign="middle" >B E</th><th align="center" valign="middle" >B O</th><th align="center" valign="middle" >B O C</th><th align="center" valign="middle" >σ E 2</th><th align="center" valign="middle" >σ O 2</th><th align="center" valign="middle" >σ O C 2</th></tr></thead><tr><td align="center" valign="middle" >(250, 50)</td><td align="center" valign="middle" >0.559</td><td align="center" valign="middle" >0.241</td><td align="center" valign="middle" >0.021</td><td align="center" valign="middle" >0.043</td><td align="center" valign="middle" >0.028</td><td align="center" valign="middle" >0.017</td></tr><tr><td align="center" valign="middle" >(200, 100)</td><td align="center" valign="middle" >0.624</td><td align="center" valign="middle" >0.326</td><td align="center" valign="middle" >0.023</td><td align="center" valign="middle" >0.056</td><td align="center" valign="middle" >0.031</td><td align="center" valign="middle" >0.022</td></tr><tr><td align="center" valign="middle" >(150, 150)</td><td align="center" valign="middle" >0.769</td><td align="center" valign="middle" >0.457</td><td align="center" valign="middle" >0.025</td><td align="center" valign="middle" >0.059</td><td align="center" valign="middle" >0.034</td><td align="center" valign="middle" >0.022</td></tr><tr><td align="center" valign="middle" >(100, 200)</td><td align="center" valign="middle" >0.812</td><td align="center" valign="middle" >0.608</td><td align="center" valign="middle" >0.022</td><td align="center" valign="middle" >0.065</td><td align="center" valign="middle" >0.038</td><td align="center" valign="middle" >0.023</td></tr><tr><td align="center" valign="middle" >(50, 250)</td><td align="center" valign="middle" >0.838</td><td align="center" valign="middle" >0.809</td><td align="center" valign="middle" >0.023</td><td align="center" valign="middle" >0.078</td><td align="center" valign="middle" >0.044</td><td align="center" valign="middle" >0.026</td></tr></tbody></table></table-wrap><table-wrap id="table2" ><label><xref ref-type="table" rid="table2">Table 2</xref></label><caption><title> Estimated variance, boot: bootstrap</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >( n 1 , n 2 )</th><th align="center" valign="middle" >Parameter</th><th align="center" valign="middle" >σ O 2</th><th align="center" valign="middle" >σ O C 2</th><th align="center" valign="middle" >σ E 2</th><th align="center" valign="middle" >σ B o o t 2</th></tr></thead><tr><td align="center" valign="middle" >(50, 250)</td><td align="center" valign="middle" >β 1</td><td align="center" valign="middle" >0.7312</td><td align="center" valign="middle" >0.1514</td><td align="center" valign="middle" >1.2726</td><td align="center" valign="middle" >0.1244</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >β 2</td><td align="center" valign="middle" >0.0152</td><td align="center" valign="middle" >0.0049</td><td align="center" valign="middle" >0.0287</td><td align="center" valign="middle" >0.0035</td></tr><tr><td align="center" valign="middle" >(100 ,200)</td><td align="center" valign="middle" >β 1</td><td align="center" valign="middle" >0.4723</td><td align="center" valign="middle" >0.1617</td><td align="center" valign="middle" >0.5491</td><td align="center" valign="middle" >0.1157</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >β 2</td><td align="center" valign="middle" >0.0108</td><td align="center" valign="middle" >0.0050</td><td align="center" valign="middle" >0.0160</td><td align="center" valign="middle" >0.0022</td></tr><tr><td align="center" valign="middle" >(150, 150)</td><td align="center" valign="middle" >β 1</td><td align="center" valign="middle" >0.2765</td><td align="center" valign="middle" >0.1463</td><td align="center" valign="middle" >0.3130</td><td align="center" valign="middle" >0.1175</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >β 2</td><td align="center" valign="middle" >0.0076</td><td align="center" valign="middle" >0.0043</td><td align="center" valign="middle" >0.0087</td><td align="center" valign="middle" >0.0025</td></tr><tr><td align="center" valign="middle" >(200, 100)</td><td align="center" valign="middle" >β 1</td><td align="center" valign="middle" >0.1848</td><td align="center" valign="middle" >0.1442</td><td align="center" valign="middle" >0.2153</td><td align="center" valign="middle" >0.1391</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >β 2</td><td align="center" valign="middle" >0.0064</td><td align="center" valign="middle" >0.0049</td><td align="center" valign="middle" >0.0077</td><td align="center" valign="middle" >0.0028</td></tr><tr><td align="center" valign="middle" >(250, 50)</td><td align="center" valign="middle" >β 1</td><td align="center" valign="middle" >0.1522</td><td align="center" valign="middle" >0.1342</td><td align="center" valign="middle" >0.1604</td><td align="center" valign="middle" >0.1274</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >β 2</td><td align="center" valign="middle" >0.0053</td><td align="center" valign="middle" >0.0050</td><td align="center" valign="middle" >0.0059</td><td align="center" valign="middle" >0.0039</td></tr><tr><td align="center" valign="middle" >(270, 30)</td><td align="center" valign="middle" >β 1</td><td align="center" valign="middle" >0.1677</td><td align="center" valign="middle" >0.1621</td><td align="center" valign="middle" >0.1696</td><td align="center" valign="middle" >0.1497</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >β 2</td><td align="center" valign="middle" >0.0052</td><td align="center" valign="middle" >0.0049</td><td align="center" valign="middle" >0.0055</td><td align="center" valign="middle" >0.0026</td></tr><tr><td align="center" valign="middle" >(290, 10)</td><td align="center" valign="middle" >β 1</td><td align="center" valign="middle" >0.1483</td><td align="center" valign="middle" >0.1461</td><td align="center" valign="middle" >0.1509</td><td align="center" valign="middle" >0.1321</td></tr><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >β 2</td><td align="center" valign="middle" >0.0047</td><td align="center" valign="middle" >0.0046</td><td align="center" valign="middle" >0.0048</td><td align="center" valign="middle" >0.0029</td></tr></tbody></table></table-wrap><p>detecting the effect size of 0.75 for the proportional hazards model with a sample of size 300 using current status data. Considering that pooled repeated partly interval-censored data has more information than current status data, we fully</p><table-wrap id="table3" ><label><xref ref-type="table" rid="table3">Table 3</xref></label><caption><title> The 95% coverage probabilities</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >( n 1 , n 2 )</th><th align="center" valign="middle" >β 1</th><th align="center" valign="middle" >β 2</th><th align="center" valign="middle" >Joint</th></tr></thead><tr><td align="center" valign="middle" >(50, 250)</td><td align="center" valign="middle" >0.878</td><td align="center" valign="middle" >0.853</td><td align="center" valign="middle" >0.835</td></tr><tr><td align="center" valign="middle" >(100, 200)</td><td align="center" valign="middle" >0.883</td><td align="center" valign="middle" >0.874</td><td align="center" valign="middle" >0.866</td></tr><tr><td align="center" valign="middle" >(150, 150)</td><td align="center" valign="middle" >0.899</td><td align="center" valign="middle" >0.886</td><td align="center" valign="middle" >0.881</td></tr><tr><td align="center" valign="middle" >(200, 100)</td><td align="center" valign="middle" >0.910</td><td align="center" valign="middle" >0.906</td><td align="center" valign="middle" >0.903</td></tr><tr><td align="center" valign="middle" >(250, 50)</td><td align="center" valign="middle" >0.947</td><td align="center" valign="middle" >0.931</td><td align="center" valign="middle" >0.936</td></tr></tbody></table></table-wrap><p>agree with this better power result. The 95% coverage probabilities for different proportions of pooled repeated partly interval-censored data are shown in <xref ref-type="table" rid="table3">Table 3</xref>.</p><p>In summary, even a small amount of pooled repeated interval-censored data within O does make our statistical inference more accurate and more powerful.</p></sec></sec><sec id="s5"><title>5. Analysis of NLSY97 Data</title><p>For more than 4 decades, the National Longitudinal Surveys (NLS) data have served as an important tool for economists, sociologists, and other researchers. The NLSY97 is a nationally representative sample of approximately 9000 youths who were 12 to 18 years old as of December 31, 1996. The NLSY97 is designed to document the transition from school to work and into adulthood. It collects extensive information about youths’ labor market behavior and educational experiences over time. In addition to educational and labor market experiences, the NLSY97 contains detailed information on many other topics. Some of the areas included in the data are criminal behavior, alcohol, and drug use. For the purpose of illustration of our methods, we use the NLSY97 data from 1997 to 2013 ( [<xref ref-type="bibr" rid="scirp.107493-ref12">12</xref>] ). We illustrate how to analyze the effects of covariates that may affect an adolescent’s smoking behavior.</p><p>There are 8984 subjects in the data set. We analyze the 1822 subjects who did not smoke at the beginning of the study in 1997, but by the end of 2013 became frequent smokers (smoking for more than 10 days in a month). That is O. The response variable is defined as</p><p>y i j = ( 1, a frequent smoker 0, not a frequent smoker . (9)</p><p>Exact observations (E) are available in approximately 87.5% of those analyzed. The 1<sup>st</sup> covariate, x 1 i j , is the number of days an individual drank alcohol in the last 30 days. The 2<sup>nd</sup> covariate, x 2 i j , is an individual’s self-evaluation of “general state of health”. x 2 i j is defined as: 1 = excellent, 2 = very good, 3 = good, 4 = fair, and 5 = poor. The covariate effects are estimated by the EM algorithm in Section 0. The standard errors of these estimators are computed by Louis’ method in Section 0. The results are shown in <xref ref-type="table" rid="table4">Table 4</xref>. Fixing an individual’s self-evaluated health level as the subject drinks alcohol one more day during the past 30 days, the log of odds of becoming a frequent smoker increases by 0.1 (s.e. = 0.002). Furthermore, by fixing an individual’s amount of drink as the subject’s health level rises (i.e., gets worse) by one unit, the log of odds of becoming a frequent smoker increases by 0.19 (s.e. = 0.015).</p><p>Additionally, we analyzed only E from O in order to see how much smaller the pooled repeated interval-censored data can help make the analysis more accurate. Another rationale for this is some practitioners often analyze only E due to the unavailability of software. The results are shown in <xref ref-type="table" rid="table5">Table 5</xref>. The parameter estimates are very close to those from O. However, the estimated standard errors are much larger than those from O. This is consistent with the simulation results in Section 3.2. The Wald test statistic for testing ( β 1 , β 2 ) = ( 0,0 ) is quite large for both E alone and O. Therefore, the p-values are nearly 0. Though both tests tell us that the covariates have a statistically significant effect on adolescent’s smoking behavior, O provides us with much stronger evidence for the</p><table-wrap id="table4" ><label><xref ref-type="table" rid="table4">Table 4</xref></label><caption><title> The results of NLSY97 analysis using the observed data</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >α ^</th><th align="center" valign="middle" >s e ( α ^ )</th><th align="center" valign="middle" >β ^ 1</th><th align="center" valign="middle" >s e ( β ^ 1 )</th><th align="center" valign="middle" >β ^ 2</th><th align="center" valign="middle" >s e ( β ^ 2 )</th></tr></thead><tr><td align="center" valign="middle" >−2.36</td><td align="center" valign="middle" >0.041</td><td align="center" valign="middle" >0.103</td><td align="center" valign="middle" >0.002</td><td align="center" valign="middle" >0.19</td><td align="center" valign="middle" >0.015</td></tr></tbody></table></table-wrap><table-wrap id="table5" ><label><xref ref-type="table" rid="table5">Table 5</xref></label><caption><title> The results of NLSY97 analysis using only the exact data</title></caption><table><tbody><thead><tr><th align="center" valign="middle" >α ^</th><th align="center" valign="middle" >s e ( α ^ )</th><th align="center" valign="middle" >β ^ 1</th><th align="center" valign="middle" >s e ( β ^ 1 )</th><th align="center" valign="middle" >β ^ 2</th><th align="center" valign="middle" >s e ( β ^ 2 )</th></tr></thead><tr><td align="center" valign="middle" >−2.35</td><td align="center" valign="middle" >0.067</td><td align="center" valign="middle" >0.102</td><td align="center" valign="middle" >0.004</td><td align="center" valign="middle" >0.18</td><td align="center" valign="middle" >0.028</td></tr></tbody></table></table-wrap><p>effect. Therefore, this data analysis reaffirms that even a small amount of pooled repeated interval-censored portion of O increases the sensitivity of the test.</p></sec><sec id="s6"><title>6. Discussion</title><p>We focused on developing a method to estimate the regression parameters and the variance-covariance matrix of those estimators for the pooled repeated partly interval-censored data logistic regression model. We employed the EM algorithm to estimate the parameters and missing information principle to estimate the variance-covariance matrix of those estimators.</p><p>Monte Carlo simulation demonstrates acceptable levels of bias, standard error, and power. To our knowledge, this is the first extensive power study for the pooled repeated partly interval-censored data logistic regression model. The simulation results suggest that in practice, one needs a sample of size around 300 to achieve an 80% power of the test to detect a very small effect size (0.45) for the regression parameter of interest. However, one needs a much smaller size, only around 200, for a bit larger effect size (0.55).</p><p>There are several potential extensions of our methods. Our methods can also be used when the predetermined follow-up visits were person-dependent. Our methods can be extended to handle correlated covariates by employing a ridge regression model ( [<xref ref-type="bibr" rid="scirp.107493-ref13">13</xref>] ), variable selections by lasso regression ( [<xref ref-type="bibr" rid="scirp.107493-ref14">14</xref>] ), and multiple progression states due to the fact that the likelihood factors into a distinct term for each interval ( [<xref ref-type="bibr" rid="scirp.107493-ref15">15</xref>] ).</p><p>Last but not least, we note that there are challenges in including either left-censoring or right-censoring. Refer to Appendix 1 for details.</p></sec><sec id="s7"><title>Acknowledgements</title><p>The authors appreciate Dr. Alexis Dinno for introducing the data.</p></sec><sec id="s8"><title>Conflicts of Interest</title><p>The authors declare no conflicts of interest regarding the publication of this paper.</p></sec><sec id="s9"><title>Cite this paper</title><p>Daneshi, N. and Kim, J.S. (2021) Maximum Likelihood Estimation for the Pooled Repeated Partly Interval-Censored Observations Logistic Regression Model. Open Journal of Statistics, 11, 230-242. https://doi.org/10.4236/ojs.2021.111012</p></sec><sec id="s10"><title>Appendix 1. Right and Left Censoring in the Model</title><p>In some special cases, the visiting time of some subjects in the data may have either right or left censoring. If a subject has not failed at the last visit ( y i L i = 0 ) and does not come back for the proceeding interview visits, then the subject’s time to the event of interest is right-censored. In this case L i = M i and R i = M . As NLSY predetermined M for all subjects, M plays the role of ∞ .</p><p>One may want to impute the covariate, x i j and reponse, y i j according to the procedures in Section 3.1. Unfortunately, extrapolating the covariates x i j for j &gt; L i using the linear growth curve in Section 3.1 may well increase bias and variance.</p><p>If a subject’s first visit is at time k and the subject shows the symptoms of the event of interest, then both y i j and x i j are missing for j = 1, ⋯ , k − 1 , and y i k = 1 . Therefore, the covariate, x i j and response, y i j should be estimated for j ≤ k − 1 at E-step. We merely have L i = 0 , R i = k , and two observed covariate values x i 0 and x i k . Therefore, we cannot fit the subject-dependent growth curve to estimate the covariates at the missed visits.</p><p>In summary, there is no merit to include individuals whose event-times are either left-censored or right-censored when fitting a logistic regression model with pooled repeated observations.</p></sec><sec id="s11"><title>Appendix 2. Imputation of Covariates</title><p>In Section 3.1, we assumed that covariates have a linear growth curve with fixed effects. This was motivated by NLSY97 data. In NLSY97, follow-up interviews were relatively far apart (1 year). Additionally, some individuals had no change in their covariate values, e.g., some individuals had no drinking throughout the study. This motivated us to assume that for a given individual, the covariate values are uncorrelated at different follow-up visits, i.e., c o v ( ε i j , ε i j ′ ) = 0, j ≠ j ′ .</p><p>If the follow-up time intervals are relatively short and there are no constant covariate values for any individual over time, one may adopt a linear growth curve with fixed effects and autocorrelated errors. That is,</p><p>x i j = θ 0 i + θ 1 i t j − 1 + ε i j , (10)</p><p>where ε i j is an autoregressive process with lag 1, AR (1), ε i j ∼ N ( 0, σ ε 2 ) , c o v ( ε i , j , ε i , j + 1 ) = ρ , ρ ≠ 0 .</p><p>[<xref ref-type="bibr" rid="scirp.107493-ref7">7</xref>] and [<xref ref-type="bibr" rid="scirp.107493-ref16">16</xref>] assumed random effects. In a linear growth curve with random effects, all subjects have the same growth curve distribution, which depends on time points and it is correlated within the same subject. The least squares estimators for this model are the same for all subjects. For example, assume that we get θ ^ 0 and θ ^ 1 for a random effect model. If two subjects i * and i * * have missing covariates at a given time point j, then they will have the same estimated covariate x ^ i j = θ ^ 0 + θ ^ 1 t j for i = i * , i * * . This may cause a substantial amount of bias.</p></sec><sec id="s12"><title>Appendix 3. Formulas for Computing the Variance in Section 3.2</title><p>The variance estimation is based on Y i , the observed binary outcomes for subject i and the variability of missing response, y i j conditioned on x i , where x i = { x i 1 , x i 2 , ⋯ , x i , L i , x ^ i , L i + 1 , ⋯ , x ^ i , R i − 1 , x i , R i } . Let θ = ( α , β → ) , and z i = ( 1, x i ) . Then the complete information matrix in (8) can be computed by</p><p>∑ i     z i z i T exp ( θ T z i ) ( 1 + exp ( θ T z i ) ) 2 . (11)</p><p>The missing information in (8) is computed by Monte Carlo simulations.</p><p>V a r ( − ∂ log P ( θ | W , V ) ∂ θ ) = E ( − ∂ log P ( θ | W , V ) ∂ θ ) 2 − [ E ( − ∂ log P ( θ | W , V ) ∂ θ ) ] 2 , (12)</p><p>where</p><p>E ( − ∂ log P ( θ | W , V ) ∂ θ ) ≈ 1 B ∑ b = 1 B [ − ∂ log P ( θ | W , V b ) ∂ θ ]</p><p>and</p><p>E ( − ∂ log P ( θ | W , V ) ∂ θ ) 2 ≈ 1 B ∑ b = 1 B [ − ∂ log P ( θ | W , V b ) ∂ θ ] 2 .</p></sec></body><back><ref-list><title>References</title><ref id="scirp.107493-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Wulfsohn, M.S. and Tsiatis, A. (1997) A Joint Model for Survival and Longitudinal Data Measured with Error. Biometrics, 53, 330-339.  
https://doi.org/10.2307/2533118</mixed-citation></ref><ref id="scirp.107493-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Allison, P.D. (2010) Survival Analysis Using SAS: A Practical Guide. SAS Institute, Cary.</mixed-citation></ref><ref id="scirp.107493-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Tibshirani, R. (1996) Regression Shrinkage and Selection via the Lasso. Journal of the Royal Statistical Society: Series B (Methodological), 58, 267-288.  
https://doi.org/10.1111/j.2517-6161.1996.tb02080.x</mixed-citation></ref><ref id="scirp.107493-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Hoerl, A., Kennard, R. and Baldwin, K. (1975) Ridge Regression: Some Simulations. Communications in Statistics, 4, 105-123.</mixed-citation></ref><ref id="scirp.107493-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Bureau of Labor Statistics, U.S. Department of Labor (2015) National Longitudinal Survey of Youth 1997 Cohort, 1997-2013 (Rounds 1-16) Produced by the National Opinion Research Center, the University of Chicago and Distributed by the Center for Human Resource Research, The Ohio State University. Columbus.</mixed-citation></ref><ref id="scirp.107493-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Mongoué-Tchokoté, S. and Kim, J.S. (2008) New Statistical Software for the Proportional Hazards Model with Current Status Data. Computational Statistics and Data Analysis, 52, 4272-4286. https://doi.org/10.1016/j.csda.2008.02.007</mixed-citation></ref><ref id="scirp.107493-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Tanner, M.A. (1996) Tools for Statistical Inference. 3rd Edition, Springer-Verlag, New York. https://doi.org/10.1007/978-1-4612-4024-2</mixed-citation></ref><ref id="scirp.107493-ref8"><label>8</label><mixed-citation publication-type="book" xlink:type="simple">Masyn, K.E., Petras, H. and Liu, W. (2014) Growth Curve Models with Categorical Outcomes. In: Bruinsma, G. and Weisburd, D., Eds., Encyclopedia of Criminology and Criminal Justice, Springer, New York, 2013-2025. 
https://doi.org/10.1007/978-1-4614-5690-2_404</mixed-citation></ref><ref id="scirp.107493-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Dempster, A.P., Laird, N.M. and Rubin, D.B. (1997) Maximum Likelihood from Incomplete Data via the EM Algorithm. Journal of the Royal Statistical Society: Series B (Methodologica), 39, 1-22.  
https://doi.org/10.1111/j.2517-6161.1977.tb01600.x</mixed-citation></ref><ref id="scirp.107493-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Finkelstein, D.M., Wang, R., Ficociello, L.H. and Schoenfeld, D.A. (2010) A Score Test for Association of a Longitudinal Marker and an Event with Missing Data. Biometrics, 66, 726-732. https://doi.org/10.1111/j.1541-0420.2009.01326.x</mixed-citation></ref><ref id="scirp.107493-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">D’Agostino, R., Lee, M.L., Belanger, A., Cupples, L.A., Anderson, K. and Kannel, W.B. (1990) Relation of Pooled Logistic Regression to Time Dependent Cox Regression Analysis: The Framingham Heart Study. Statistics in Medicine, 9, 1501-1515.  
https://doi.org/10.1002/sim.4780091214</mixed-citation></ref><ref id="scirp.107493-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Adrienne Cupples, L., D’Agostino, R.B., Anderson, K. and Kannel, W.B. (1988) Comparison of Baseline and Repeated Measure Covariate Techniques in the Framingham Heart Study. Statistics in Medicine, 7, 205-218.  
https://doi.org/10.1002/sim.4780070122</mixed-citation></ref><ref id="scirp.107493-ref13"><label>13</label><mixed-citation publication-type="journal" xlink:type="simple"><name name-style="western"><surname>Huang</surname><given-names> J. </given-names></name>,<etal>et al</etal>. (<year>1999</year>)<article-title>Asymptotic Properties of Nonparametric Estimation Based on Partly Interval-Censored Data</article-title><source> Statistica Sinica</source><volume> 9</volume>,<fpage> 501</fpage>-<lpage>519</lpage>.<pub-id pub-id-type="doi"></pub-id></mixed-citation></ref><ref id="scirp.107493-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Kim, J.S. (2003) Maximum Likelihood Estimation for the Proportional Hazards Model with Partly Interval-Censored Data. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 65, 489-502. 
https://doi.org/10.1111/1467-9868.00398</mixed-citation></ref><ref id="scirp.107493-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Zhao, X.Q., Zhao, Q., Sun, J.G. and Kim, J.S. (2008) Generalized Log-Rank Tests for Partly Interval-Censored Failure Time Data. Biometrical Journal, 50, 375-385. 
https://doi.org/10.1002/bimj.200710419</mixed-citation></ref><ref id="scirp.107493-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Gao, F., Zeng, D.L. and Lin, D.Y. (2017) Semiparametric Estimation of the Accelerated Failure Time Model with Partly Interval-Censored Data. Biometrics, 73, 1161-1168. https://doi.org/10.1111/biom.12700</mixed-citation></ref></ref-list></back></article>