Prevalence, Domain Structure, and Correlates of Multidimensional Male Partner Involvement in Maternal and Newborn Health in Sierra Leone ()
1. Introduction
Few countries carry a maternal and newborn mortality burden as heavy as Sierra Leone’s. The most recent United Nations inter-agency modelled estimate places the maternal mortality ratio at 354 deaths per 100,000 livebirths, which is more than five times the Sustainable Development Goal 3.1 target of fewer than 70 [1] [2]. A rather different figure emerges from the 2019 Demographic and Health Survey, whose sisterhood method captures deaths occurring in the community as well as in facilities and produces a population-based estimate of approximately 717 per 100,000 live births [3]. Facility-based surveillance gives a lower figure still, 183.8 per 100,000 livebirths for January to November 2024, although this figure by construction omits the deaths that never reach a facility [4]. The neonatal mortality rate stands at 34 per 1000 livebirths, with prematurity, birth asphyxia, and neonatal sepsis together accounting for four-fifths of neonatal deaths [5].
What makes these figures difficult to explain is that they have persisted through a period of substantial institutional reform. The Free Health Care Initiative, launched in April 2010, abolished user fees for pregnant women, lactating mothers, and children under five, and facility-based deliveries subsequently rose from 25% in 2008 to 83% in 2019 [6] [7]. The reform worked, in the narrow sense that it removed the price at the facility door. The mortality burden nevertheless remains. Part of the explanation lies in the clinical character of the deaths themselves. Obstetric haemorrhage, hypertensive disorders, obstructed labour and sepsis account for the large majority of direct maternal deaths in Sierra Leone, and all four share a property that makes them unusually unforgiving: their lethality depends on time [5] [8]. Postpartum haemorrhage can progress from onset to death within two hours. The Three Delays framework formalises the consequence, which is that hours consumed at the household decision-making stage can render the clinical capacity of a referral centre irrelevant [9] [10]. More recent theoretical work has extended the framework to draw attention to the connection points between phases, where continuity of care is most often lost [11]. Regional review evidence indicates further that facility-based maternal deaths across West Africa remain systematically under-investigated relative to deaths occurring in the community, so the true distribution of delay is likely to be understated in routine data [12].
Where household decision-making is controlled by men, the male partner becomes the effective gatekeeper of that first delay. This mechanism accounts for much of the global policy interest in male partner involvement. Systematic reviews and meta-analyses have reported associations between male engagement and higher antenatal attendance, skilled birth attendance, and postnatal care utilisation in low-income and middle-income settings [13]-[15], and a 2026 meta-analysis reported a pooled odds ratio of 2.05 (95% CI: 1.63 - 2.59) for institutional delivery among women whose partners accompanied them to antenatal care [16].
Three features of that evidence limit its usefulness in Sierra Leone. The first is geographic. Reviews of determinants across sub-Saharan Africa draw predominantly on Ethiopian, Kenyan, Tanzanian, and Zimbabwean samples, and West African settings are comparatively thinly represented [17] [18]. Ethiopian evidence, in particular, now dominates the meta-analytic literature on delivery care [19], raising a question of transferability that has rarely been tested directly. The second limitation is measurement. Exposure is often reduced to a single binary item that records whether a man attended at least one antenatal visit. A natural-language-processing review of 282 studies found that antenatal attendance was the single most commonly used indicator, appearing in 40% of studies, and that one third of all studies relied on a single indicator alone [20]; an international Delphi process reached the same conclusion about the inadequacy of such measures [21]. Critical scholarship has gone further, arguing that accompaniment-based benchmarks import a companionate-marriage model that may simply not describe how support is enacted elsewhere [22] [23]. Qualitative work from Ghana has shown that men in patriarchal West African settings recognise the value of skilled care yet rarely involve themselves unless complications arise, with fewer than a quarter of male participants having ever accompanied a partner to antenatal or postnatal care [24]; the same body of work has documented that some women actively resist male accompaniment because it curtails their own social space within the clinic [25]. Ethnographic work in Sierra Leone specifically has shown that men frame their obligations in material and logistical terms, and that clinic presence is neither expected of them nor normatively rewarded [26].
The third limitation is the reliance on a single reporter. Because almost all studies ask either the woman or the man but not both, the systematic divergence between their accounts of the same behaviour remains largely unquantified [27]. The problem is not new; it was identified in the family planning literature three decades ago [28] and has since been documented repeatedly in South Asian settings, where spousal reports of contraceptive communication and use diverge substantially [29]. It has, however, seldom been examined in African maternal health, and never at the national scale in Sierra Leone. A recent scoping review of fatherhood and men’s participation across rural sub-Saharan Africa reached the same conclusion, noting that men’s participation is constituted relationally and cannot be adequately captured from one side of the relationship [30].
This study was designed to address all three limitations. It reports the first two objectives of a national dyadic mixed-methods investigation: to assess male partner involvement across the six domains of a purpose-built index and along the continuum of care, and to analyse the factors associated with men’s participation in maternal and newborn health. Because the design is cross-sectional, the analysis identifies correlates rather than causes, and the language of this paper has been kept to association throughout.
2. Research Methodology
2.1. Research Design
A descriptive cross-sectional survey design with a dyadic (paired-respondent) structure was employed, forming the quantitative strand of a convergent parallel mixed-methods study. The design was appropriate because it permitted simultaneous measurement of a behavioural exposure and its correlates in a large sample at a single point in time, while the paired structure allowed the same behaviour to be measured independently from two accounts. Inferential techniques, including Pearson correlation, independent-samples t-tests, one-way analysis of variance, and ordinary least squares multiple regression, were used to identify factors independently associated with involvement. Reporting of the quantitative strand follows the STROBE statement for cross-sectional studies [31]; reporting of the qualitative strand is described in Section 2.5.
2.2. Description of Study Area
The study was conducted in 15 of Sierra Leone’s 16 administrative districts, covering all five provinces/regions, namely Eastern, Northern, North West, Southern, and Western Area, between 20 January and 24 June 2026. Karene District in the North West Province was not included for logistical reasons related to the seasonal inaccessibility of its northern chiefdoms during the fieldwork window. This national scope was deliberate. It was necessary in order to capture the full structural range of access to maternity care, which extends from the tertiary referral capacity concentrated in Western Area Urban to the pronounced geographic isolation of districts such as Falaba, Koinadugu, and Bonthe. Households in this setting navigate a plural health system in which a formal referral pyramid of Peripheral Health Units (PHUs) and hospitals coexists with a traditional sector comprising traditional birth attendants and the Sande female society, an institution whose authority over the birth space is considerable in the southern and eastern provinces.
2.3. Sampling and Sampling Techniques
The target population comprised paired respondents. Eligible women were of reproductive age (15 - 49 years), had delivered a livebirth or stillbirth preceding data collection, and were usual residents of a selected cluster. The acknowledged husband or partner associated with the index pregnancy was enrolled as the paired respondent. Women without a contactable partner were not enrolled, since the dyadic design requires both accounts.
The interval between reported delivery and interview is reported on the 449 dyads (89.8%) whose recorded delivery date preceded the interview date. Among these, the median interval was 19 months (interquartile range, 11 to 28), with a maximum of 52 months; 30.5% of births occurred within 12 months of the interview, and 64.8% within 24 months. The remaining 51 records (10.2%) have delivery dates between one and nine months after the interview date and are described in Section 5; they are excluded from the recall statistics but retained in all analyses that do not depend on the delivery date.
A four-stage stratified cluster sampling procedure was adopted, using PHUs as sampling hubs. In the first stage, the 15 participating districts were treated as explicit strata to preserve sub-national coverage and capture the travel-time gradients central to any assessment of Phase II delays. At the second stage, PHUs (Maternal and Child Health Posts, Community Health Posts, or Community Health Centres) were selected at random within each district. At the third stage, each selected PHU identified its catchment communities, and enumerators were directed to recruit from both the PHU village, where access is comparatively good, and distal villages more than 30 minutes away, to achieve maximum variation in geographic access. At the fourth stage, cluster listing and a systematic random-walk procedure were used within villages to identify and enroll dyads.
Three features of this procedure bear directly on how the estimates below should be read and are stated explicitly in response to review. First, selection probabilities were not proportional to size at any stage. Districts were allocated on operational rather than demographic grounds, and the achieved allocation ranged from 11 dyads in Koinadugu to 53 in Western Area Urban. No sampling frame of recently delivered women with an identifiable partner exists for Sierra Leone at the district level, so design weights could not be constructed. All estimates reported here are consequently unweighted, and prevalences should be read as describing the sampled dyads rather than as design-weighted national parameters. Second, the deliberate inclusion of distal villages was a purposive maximum-variation device rather than a probability mechanism. It was adopted so that the travel-time gradient would be represented across its full range, but it means that the achieved sample over-represents geographically isolated households relative to a self-weighting design, and that the sample mean travel time of 53.6 minutes is not an unbiased national estimate. Third, the number of PHUs and catchment communities visited is recorded in the field logs but was not carried into the analytic file, which retains only district, a sequential field-cluster identifier for the 50 groups of ten dyads, and interviewer identifiers (12 female and 8 male enumerators). Because no primary sampling unit identifier is available, a fully design-based variance estimator cannot be fitted. Section 2.6 sets out the design-adjusted approach adopted instead, and Section 5 records this as a limitation.
Sample size was determined using Cochran’s formula for estimating a proportion in a large population [32]:
n0 = Z2·p(1 − p)/e2(1)
where n0 is the size required for an effectively infinite population, Z is the standard normal deviate for the chosen confidence level (1.96 for 95% confidence), p is the expected prevalence of the key indicator, and e is the margin of error, set at 0.05. Because no prior Sierra Leonean estimate of the prevalence of multidimensional male involvement existed, p was set conservatively at 0.50, the value that maximises p(1 − p) and therefore yields the largest and most cautious requirement. Substituting these values returned a base requirement of approximately 385 respondents. Two adjustments followed. A design effect of 2.0, a conventional working assumption for maternal health cluster surveys, raised the requirement to approximately 770; inflation for non-response and incomplete pairing at 10% produced a planning target of approximately 855 women.
The executed dataset comprised 500 mother-partner couples. Within these pairs, 490 of the women’s interviews (98.0%) and 483 of the men’s (96.6%) were recorded as fully completed, the remainder as partial. Recruitment-level figures, namely the number of households approached and the number of partners who could not be contacted, are held in the field logs and are not reproduced in the analytic file; incomplete pairing rather than refusal by women was the dominant reason for the shortfall against the planning target, a point returned to in the limitations. Table 1 shows the distribution of the achieved sample across provinces.
2.4. Sources of Data Collection
Primary data were collected using two parallel structured instruments. Instrument A, administered to women, captured sociodemographic characteristics, household assets, women’s autonomy, obstetric history, and the woman’s report of her partner’s involvement. Instrument B, administered to male partners, captured parallel demographics, self-reported involvement, perceived barriers, and attitudinal items. Secondary data were drawn from Ministry of Health and Sanitation reports, Demographic and Health Survey publications, District Health Information System aggregates, and peer-reviewed literature.
Table 1. Distribution of the achieved sample by province and facility access.
Province |
Distribution of the Achieved Sample by Facility Access |
Districts Covered |
Dyads (n) |
% of Sample |
Mean Travel Time to Facility (Min) |
Eastern |
3 of 3 |
128 |
25.6 |
55.3 |
Northern |
4 of 4 |
109 |
21.8 |
74.5 |
Southern |
4 of 4 |
105 |
21.0 |
48.0 |
Western Area |
2 of 2 |
87 |
17.4 |
20.3 |
North West |
2 of 3 |
71 |
14.2 |
67.2 |
Total |
15 of 16 |
500 |
100.0 |
53.6 |
a. Travel time is the self-reported time to the nearest delivery facility. The national mean was 53.6 minutes (SD = 65.3), but the provincial range is wide: 20.3 minutes in the Western Area versus 74.5 minutes in the Northern Province. Karene District (North West Province) was not covered. b. Allocation across districts was non-proportional and reflects operational access rather than population size; the figures are therefore descriptive of the achieved sample and are not design-weighted national estimates.
The Male Involvement Index is a purpose-built 30-item instrument organised into six five-item domains. Cognitive engagement covers knowledge of danger signs, the antenatal schedule, and birth planning. Financial support covers money for antenatal care, delivery, transport, emergency reserves, and postnatal needs. Logistical support covers transport and place-of-delivery arrangements, complication planning, workload relief, and nutrition. Psychosocial support covers encouragement, couple communication, and reassurance. Decision-making partnership records whether care decisions are made jointly. Clinical accompaniment records attendance at antenatal, delivery, and postnatal contacts where culturally acceptable. Frequency and intensity items were scored on a five-point scale from 0 to 4, while event, knowledge, and decision-maker items retained categorical form. Domain construction was informed by earlier multidimensional instruments, notably the correlates framework developed for peri-urban Myanmar [33] and the ten-item scale validated for prevention of mother-to-child transmission settings in Kenya, which recovered a two-factor structure separating encouragement and reminders from active participation [34].
The scoring of the decision-making domain warrants particular comment, because it is the point at which most existing instruments become vulnerable. For each item asking who decided, a jointly made decision received full credit, a decision made by the woman alone received partial credit on the grounds that her autonomy was preserved and no male override occurred, and a decision made by the man alone or by a third party received no credit at all. A high decision-making score is therefore attainable only through collaboration and cannot be produced by control. This design choice follows the conceptualisation of empowerment as the acquisition of a capacity to choose that was previously denied [35], and it responds directly to evidence from Bangladesh that joint decision-making is associated with better reproductive outcomes than sole decision-making by either partner [36].
Each domain score is the sum of valid item points divided by the maximum attainable from those same valid items, expressed on a 0 - 100 scale, with not-applicable and missing items excluded from both numerator and denominator. The composite is the unweighted mean of the six domain scores. Three continuum aggregates were then constructed from these domain scores: an antenatal aggregate, the mean of the cognitive, financial, and logistical domains; an intrapartum aggregate, consisting of the clinical accompaniment domain alone; and a postnatal aggregate, the mean of the psychosocial and decision-making domains. The allocation of domains to phases follows the phase structure set out in the WHO antenatal care recommendations [37]. These are constructed aggregates of index domains and not phase-specific measurements, a distinction with direct consequences for how Section 3.3 may be read; it is set out there in full.
Internal consistency of the full 30-item composite was good in both instruments (Cronbach’s alpha 0.855 for Instrument A and 0.782 for Instrument B). Domain-level alpha was high for financial support (0.805 and 0.728), logistical support (0.798 and 0.705), and psychosocial support (0.815 and 0.745), but low for cognitive engagement (0.319 and 0.234), clinical accompaniment (0.207 and 0.073), and decision-making partnership (0.063 and 0.038). This pattern is expected rather than anomalous. The three high-alpha domains are built from effect indicators, in which items reflect a common underlying disposition and should therefore correlate. The three low-alpha domains are built from causal indicators, in which the items jointly constitute the construct without needing to covary: a man may accompany his partner to antenatal care and not to delivery, and he may know the danger signs of haemorrhage without knowing the antenatal schedule. Coefficient alpha is not an appropriate reliability statistic for indices of this second type [38] [39], and the low values should not be read as evidence that those domains are poorly constructed. The instruments were reviewed by maternal health specialists and by the supervisory team to establish content and face validity, and were pre-tested in a community outside the study area. Household wealth quintiles were constructed via principal components analysis of household assets, following the standard approach for settings without expenditure data [40] [41]. Women’s autonomy was measured by a five-item composite covering final say over her own health care, large household purchases, daily purchases, visits to family, and the use of her own earnings, with higher scores indicating greater autonomy (Cronbach’s alpha = 0.688).
Data were collected on tablets using Kobo Toolbox, with near-real-time encrypted synchronisation to a secure server that permitted daily supervisory monitoring of completeness and inter-interviewer variation. Twenty final-year BSc Nursing students on national public health postings served as lead enumerators; their clinical training enabled accurate identification of the morbidity proxies in the instruments, and their existing deployment across the participating districts provided a national fieldwork infrastructure. Interviews were sex-matched and conducted separately, with female enumerators interviewing women and male team members interviewing partners. At no point were partners interviewed in each other’s presence. This separation was the principal structural safeguard for women’s candour and made it possible to document controlling forms of involvement that joint interviewing would have concealed.
2.5. The Qualitative Strand
Findings from the concurrent qualitative strand are used at three points below to interpret quantitative patterns, and the methods are therefore summarised here. The strand was developed within a constructivist and phenomenological framework and addresses the fifth objective of the parent study, on men’s perceived roles, motivations, and obstacles.
Purposive maximum-variation sampling was used to span rural and urban settings, ethnic and religious groupings, and levels of health system access. Forty-three sessions were conducted across ten districts: eight focus group discussions with male partners of women who had delivered within the recall period (six to eight men per group, 54 participants in total); twenty key informant interviews with maternal health providers (midwives, community health officers, state enrolled community health nurses, maternal and child health aides and facility in-charges); ten key informant interviews with community gatekeepers (traditional leaders, imams, pastors, Sande society elders and women’s leaders); and five key informant interviews with policy and programme actors (District Health Management Team officers, the national reproductive, maternal, newborn, child and adolescent health directorate, and non-governmental programme managers). Four semi-structured guides were used, one per informant category. All sessions were audio-recorded with consent, transcribed, translated into English, and back-checked against the recordings.
Transcripts were analysed using Braun and Clarke’s six-phase reflexive thematic analysis [42], managed in NVivo. Coding was primarily inductive, with community-generated categories drawn from free-listing and spontaneous language used to build the codebook, which resolved into a 57-node code tree spanning emic and etic domains. Integration with the survey followed a joint-display matrix in which quantitative gradients were set alongside the qualitative mechanisms proposed to account for them. Convergence, divergence, and expansion were recorded for each pairing. Where qualitative material appears in Sections 3.2, 3.3, and 3.6 below, it is presented as an interpretive proposition arising from that matrix, not as corroborating evidence, and it is labelled as such at each point of use. Reporting of this strand follows the COREQ criteria [43]. Full thematic results are reported separately.
2.6. Data Analysis
Analyses were performed in R version 4.3.0 and in Python using the statsmodels and SciPy libraries. Continuous variables are summarised as means with standard deviations and categorical variables as frequencies and percentages. Bivariate associations with the MII composite were assessed using Pearson correlation for continuous predictors, independent samples t-tests for dichotomous predictors, and one-way analysis of variance for multilevel categorical predictors, with eta squared reported as the effect size against the conventional benchmarks [44]. Cohen’s d for dichotomous contrasts is computed from the difference in group means divided by the pooled standard deviation. Inter-reporter agreement was assessed in three ways: by Pearson correlation between paired composite scores, by a two-way mixed-effects intraclass correlation coefficient for absolute agreement, and by cross-classifying couples into categories of each reporter’s distribution, using both a median split and a tertile split, with couples falling into different categories classified as discordant. Both cut-points are reported throughout, because the two yield materially different discordance rates and reporting either alone would misrepresent the finding.
Factors independently associated with involvement were identified using ordinary least squares multiple linear regression of the women’s-report composite on seven pre-specified covariates: education level, household wealth quintile, travel time, maternal age, parity, Muslim religion with Christian faith as reference, and Western Province residence with all other provinces as reference. Model assumptions were examined through inspection of residual plots, and multicollinearity was assessed by variance inflation factors, which ranged from 1.03 to 1.72 and indicated no material collinearity. No predictor contained missing values, and no case was excluded, so all 500 observations were retained in every model; the residual degrees of freedom of 492 follow from 500 cases and eight estimated parameters. Statistical significance was set at a two-sided alpha of 0.05.
Women’s autonomy was not among the pre-specified covariates for conceptual and empirical reasons, and both are reported so that the exclusion can be judged. Conceptually, autonomy is not an antecedent of involvement in the sense that education, wealth, and distance are; it is a relational property of the same household bargaining process that produces involvement and is as plausibly consequent upon it as before it, so entering it as a covariate would treat a co-determined quantity as exogenous. Empirically, as Section 3.5 reports, autonomy is strongly collinear with education. A sensitivity model adding autonomy to the seven pre-specified covariates was nevertheless fitted and is reported in full in Section 3.5 so that readers may weigh the exclusion for themselves.
Because respondents were recruited through a clustered design, intra-cluster correlation coefficients were estimated by one-way analysis of variance with district as the clustering unit, and design effects were derived as 1 + (m − 1) ρ, where m is the average number of dyads per district. Design-adjusted confidence intervals for principal descriptive estimates were obtained by inflating the simple-random-sampling standard error by the square root of the design effect, and the corresponding effective sample size is reported. The regression model was refitted with cluster-robust (sandwich) standard errors, again with district as the clustering unit. District is a coarser unit than the primary sampling unit actually used, so these adjustments are conservative with respect to between-district variation but cannot capture residual clustering within districts; the qualification is recorded in Section 5.
2.7. Ethical Consideration
Ethical clearance was obtained from the Njala University Institutional Review Board and the Sierra Leone Ethics and Scientific Review Committee. Written authorisation was granted by the Ministry of Health and Sanitation and by all participating District Health Management Teams, and community assent from paramount chiefs preceded recruitment in every chiefdom. Written informed consent was obtained from literate participants, and witnessed oral consent with thumbprint attestation, verified by an impartial witness independent of the research team, was obtained from participants with limited literacy. Participation was voluntary throughout, and refusal carried no consequence for access to health services. A trained intimate partner violence referral protocol operated at every site, with linkage to district gender-based violence services. Identifiers were stored separately from responses, and all electronic data was encrypted to the AES-256 standard.
3. Results and Discussion
3.1. Sociodemographic Characteristics of the Respondents
Table 2 presents the sociodemographic profile of the 500 dyads. The sample was young. More than a quarter of women (27.0%) were under 20 years of age, and the mean age was 25.5 years (SD = 8.4) for women and 30.1 years (SD = 8.6) for men, a spousal age gap of 4.5 years. Educational disadvantage was pronounced, with 154 women (30.8%) and 95 men (19.0%) reporting no formal schooling. This is broadly consistent with the profile recorded in district-level maternal health surveys elsewhere in the country, where median maternal age was 25 years and educational attainment was similarly low [45]. Muslim respondents formed the majority (63.0%), followed by Christians (33.8%). More than half of the women (53.4%) were primigravid, a group that providers in the qualitative strand independently identified as the least prepared for obstetric emergency. Mean distance to the nearest delivery facility was 9.6 km (SD = 11.5) and mean travel time 53.6 minutes (SD = 65.3).
Two sociodemographic variables were significantly associated with facility-based delivery: education level (χ2 = 13.971, df = 4, p = 0.007) and province of residence (χ2 = 10.558, df = 4, p = 0.032). Religion was also significantly associated, though weakly (χ2 = 4.146, df = 1, p = 0.042, Yates continuity correction, restricted to the 484 Muslim and Christian respondents), an association that did not survive adjustment. The wealth gradient was steeper than either, with facility delivery rising from 34.0% in the poorest quintile through 55.0%, 80.0%, and 87.0% to 91.0% in the wealthiest (χ2 = 110.621, df = 4, p < 0.001). National survey analyses reach the same conclusion from a different direction: only a small minority of Sierra Leonean women complete the full continuum of maternal and newborn care [46], and facility childbirth tracks women’s empowerment indices closely [47]. These associations identify the confounding structure that the multivariable model was designed to address.
Table 2. Sociodemographic characteristics of the respondents (N = 500 dyads).
Characteristic |
Sociodemographic Characteristics of the Respondents |
Category |
n (Women/Men) |
% Women |
% Men |
χ2 with Facility Delivery |
Age (Years) |
<20 |
135 |
27.0 |
— |
— |
20 - 24 |
145 |
29.0 |
— |
— |
25 - 29 |
83 |
16.6 |
— |
— |
30 - 34 |
54 |
10.8 |
— |
— |
≥35 |
83 |
16.6 |
— |
— |
Mean (SD) |
— |
25.5 (8.4) |
30.1 (8.6) |
— |
Educational Level |
No Formal Education |
154/95 |
30.8 |
19.0 |
χ2 = 13.971, df = 4, p = 0.007 |
Primary (Incomplete) |
88/98 |
17.6 |
19.6 |
|
Primary (Complete) |
78/98 |
15.6 |
19.6 |
|
Secondary (BECE/SSCE) |
143/160 |
28.6 |
32.0 |
|
Post-Secondary/Tertiary |
37/49 |
7.4 |
9.8 |
|
Religion |
Muslim |
315 |
63.0 |
— |
χ2 = 4.146, df = 1, p = 0.042 |
Christian |
169 |
33.8 |
— |
|
Other/Traditional |
16 |
3.2 |
— |
|
Parity |
1 (Primigravid) |
267 |
53.4 |
— |
— |
2 |
81 |
16.2 |
— |
— |
3 |
62 |
12.4 |
— |
— |
≥4 |
90 |
18.0 |
— |
— |
Wealth Quintile |
Q1 (Poorest) to Q5 (Wealthiest) |
100 Each |
20.0 Each |
— |
χ2 = 110.621, df = 4, p < 0.001 |
Province |
Eastern |
128 |
25.6 |
— |
χ2 = 10.558, df = 4, p = 0.032 |
Northern |
109 |
21.8 |
— |
|
Southern |
105 |
21.0 |
— |
|
Western Area |
87 |
17.4 |
— |
|
North West |
71 |
14.2 |
— |
|
Facility Access |
Distance (km), Mean (SD) |
|
9.6 (11.5) |
— |
r with MII = −0.193, p < 0.001 |
Travel Time (min), Mean (SD) |
|
53.6 (65.3) |
— |
r with MII = −0.181, p < 0.001 |
a. Religion, parity, and ethnicity were not collected in the male sociodemographic module. Wealth quintiles were constructed via principal components analysis of household assets, with equal allocation (n = 100 per quintile); the χ2 shown corresponds to the facility-delivery gradient of 34.0%, 55.0%, 80.0%, 87.0%, and 91.0% across Q1 to Q5. MII = Male Involvement Index. b. The religion χ2 is a 2 × 2 test with Yates continuity correction comparing Muslim (n = 315) and Christian (n = 169) respondents only, and therefore has an analytic sample of 484; the 16 respondents of other or traditional faith are excluded from that test alone.
3.2. Level and Domain Structure of Male Partner Involvement
Mean MII composite was 27.54 out of 100 (SD = 15.52) by women’s report and 19.54 (SD = 11.33) by men’s self-report (Table 3, Figure 1). Both fall well below the scale midpoint of 50, and so does every one of the six constituent domains in both instruments. Read plainly, a composite of 27.54 means that roughly three-quarters of the supportive behaviour the index was built to capture was simply absent. Precision on this estimate is materially lower than a simple random sample of 500 would imply because involvement clusters geographically: the intra-cluster correlation at the district level is 0.100, yielding a design effect of 4.21 and an effective sample size of approximately 119. The design-adjusted 95% confidence interval for the composite is accordingly 24.75 to 30.33, against 26.18 to 28.90 unadjusted. The figure is difficult to compare directly with the literature because most studies report a binary prevalence rather than a graded score, but it sits at the low end of what binary measures have found: 21% of men present during childbirth in Winneba Municipality, Ghana [48], 32.95% involved in postnatal care in Wolaita Sodo, southern Ethiopia [49], 20.8% in Motta District, northwest Ethiopia [50], and roughly a third across urban Ghanaian communities [51].
![]()
Figure 1. Male Involvement Index domain scores by reporter (N = 500 dyads). Error bars show one standard error of the mean. Every domain in both instruments falls below the scale midpoint of 50.
The internal profile of the index is more interesting than the headline figure. Clinical accompaniment recorded the lowest score in both instruments (19.68 and 13.72), while decision-making partnership recorded the highest (43.66 and 35.40). This inverts the assumption embedded in most measurement practice, where accompaniment serves as the single-item proxy for involvement as a whole. In this population, the behaviour most commonly used to represent involvement is the behaviour least often performed. The finding is consistent with the ethnographic argument that clinic presence is not a locally salient expression of male support in Sierra Leone [26], with Ghanaian evidence that men do not attend unless complications arise [24], and with Nigerian survey work reporting that men’s involvement concentrates on financial and logistical rather than physical presence [52]. It aligns with the conclusion of the Delphi consensus process that accompaniment-based measures capture only a fraction of the construct [21].
The high decision-making score requires careful handling and should not be read as evidence of egalitarian households. The joint-display integration described in Section 2.5 places this quantitative result alongside a qualitative pattern in which unilateral patriarchal decision-making was coded approximately twice as often as joint decision-making, and the proposition arising from that pairing is that the domain is inflated by gatekeeping rather than by partnership. The proposition is interpretive; it is not established by the survey data, and the two strands diverge rather than converge at this point. It nevertheless carries a construct validity warning of direct relevance to the wider field. An index that records whether a man was involved in a decision, rather than whether the decision was made jointly, will systematically score control as though it were support. Comparable warnings have been raised in Bangladeshi work, where women’s perceptions of male involvement diverged sharply from men’s reported participation [53], and in the Kenyan scale development study that separated encouragement from active participation as distinct factors [34]. The scoring protocol adopted here, in which joint decisions receive full credit and male-alone decisions receive none, is one available correction and is recommended as a design standard for comparable settings.
Women reported systematically higher involvement than men reported of themselves, by a mean of 8.00 points (SD = 15.73). Classifying couples by the direction of that gap, women reported higher by more than five points in 261 couples (52.2%), men reported higher by more than five points in 86 (17.2%), and the remaining 153 (30.6%) fell within five points of one another. Only 15 couples (3.0%) produced exactly identical composites. The paired correlation between reporters was moderate but far from interchangeable (r = 0.347, p < 0.001; Figure 2). Because Pearson correlation measures consistency rather than agreement, and a systematic eight-point level difference separates the two instruments, the intraclass correlation coefficient for absolute agreement is the more appropriate statistic; it was lower still, at 0.282.
Discordance depends on the cut-point, and both are reported. Classifying each couple by whether the two reporters fell on the same side of their respective sample medians, 184 of 500 couples (36.8%) were discordant. Cross-classifying instead by tertile, a stricter criterion because it creates three categories rather than two, and therefore more opportunities to disagree, 295 couples (59.0%) were placed in different categories depending on which partner was asked. Neither figure is the discordance rate; each is the discordance rate under a stated cut-point, and the submitted version of this paper conflated them by attaching the median-split figure of 36.8% to a tertile classification. Both are reported here wherever discordance is quantified, including in the abstract and the conclusion.
![]()
Figure 2. Agreement between women’s and men’s reports of the same partner’s involvement (N = 500 couples). Points falling above the dashed line of exact agreement are couples in which the woman reported higher involvement than the man reported for himself.
Comparable discordance has been documented in contraceptive decision-making research in India [27] [29] and in family planning studies from Nepal, where concordance on communication was only moderate, even though reported joint decision-making was high [54]. It has seldom been quantified in African maternal health. The implication is methodological and consequential: if between a third and three-fifths of couples are classified differently depending on who is asked, single-reporter studies will be subject to substantial exposure misclassification. Where such misclassification is non-differential with respect to the outcome, the expected consequence for a binary exposure is attenuation of effect estimates towards the null, although the present cross-sectional design does not quantify that attenuation, and the claim is offered as an inference from measurement theory rather than as a result of this study. Some of the heterogeneity that characterises this literature may therefore reflect measurement rather than genuine variation between settings.
Table 3. Male Involvement Index composite and domain scores by reporter, with a couple of concordance statistics.
Domain/Composite |
Male Involvement Index Composite and Domain Scores by Reporter with a Couple of Concordance Statistics |
Women’s Report Mean (SD) |
Men’s Self-Report Mean (SD) |
Discrepancy (A − B) |
Paired r |
p |
Cognitive Engagement |
25.68 (22.57) |
19.16 (19.41) |
+6.52 |
0.122 |
0.006 |
Financial Support |
23.35 (19.26) |
16.44 (15.39) |
+6.91 |
0.315 |
<0.001 |
Logistical Support |
24.39 (18.98) |
13.80 (13.38) |
+10.59 |
0.322 |
<0.001 |
Psychosocial Support |
28.48 (20.74) |
18.70 (16.09) |
+9.78 |
0.368 |
<0.001 |
Decision-Making Partnership |
43.66 (20.90) |
35.40 (18.35) |
+8.26 |
0.092 |
0.040 |
Clinical Accompaniment |
19.68 (19.32) |
13.72 (15.71) |
+5.96 |
0.115 |
0.010 |
MII Composite |
27.54 (15.52) |
19.54 (11.33) |
+8.00 (15.73) |
0.347 |
<0.001 |
Intraclass Correlation (Absolute Agreement) |
0.282 |
|
|
|
|
Tertile Distribution (Women’S Report) |
T1 = 176; T2 = 160; T3 = 164 |
|
|
|
|
Discordance, Median Split |
184 (36.8%) |
|
|
|
|
Discordance, Tertile Split |
295 (59.0%) |
|
|
|
|
a. MII scored 0 - 100 with a scale midpoint of 50. Discrepancy is women’s report minus men’s self-report; positive values indicate women reported higher involvement than men reported themselves. Paired r is the Pearson correlation between the two instruments for the same couple. b. Median-split discordance counts couples in which one partner scored at or above his or her own instrument’s sample median and the other below it. Tertile discordance counts couples placed in different tertiles of their respective distributions. Both cut-points are reported because they yield materially different rates, and neither alone represents the discordance of the sample. c. An analysis of variance comparing domain scores across tertiles derived from the composite of those same domains would be circular and is therefore not reported.
3.3. Involvement across Constructed Continuum Aggregates
The three quantities compared in this section are aggregates of index domains assigned to phases of care, not measurements taken at those phases, and they are not constituted alike. The antenatal aggregate is the mean of three domains, the postnatal aggregate is the mean of two, and the intrapartum aggregate consists of the clinical accompaniment domain alone. A comparison across them therefore conflates two things: any genuine variation in involvement by phase, and the differing composition and reliability of the aggregates themselves. Clinical accompaniment is also the narrowest of the six domains in behavioural range and, at alpha 0.207, among the least internally consistent. The comparison is presented for what it is, a descriptive ordering of constructed aggregates, and the term “nadir” used in the submitted version has been removed because it implies a phase-specific measurement that the design does not deliver.
Disaggregating the index in this way (Table 4, Figure 3), the antenatal aggregate was 24.47 (SD = 17.53), the intrapartum aggregate 19.68 (SD = 19.32), and the postnatal aggregate 36.07 (SD = 17.53). All three fall below the midpoint, and the proportion of respondents scoring below the midpoint was 90.4%, 92.4%, and 76.4%, respectively. The relative elevation of the postnatal aggregate is attributable almost entirely to the decision-making domain, with the interpretive caution already noted in Section 3.2; without that domain, the postnatal aggregate would be the psychosocial score of 28.48.
Figure 3. Constructed continuum aggregates of the Male Involvement Index, women’s report (N = 500). Error bars show one standard error of the mean. The three aggregates are composed of three, one, and two domains, respectively, and are therefore not directly comparable measures of phase-specific involvement.
Subject to that qualification, the low standing of clinical accompaniment retains clinical interest, because the behaviour it measures is the one required at the point where postpartum haemorrhage, the leading cause of maternal death in this country, can kill within two hours [5] [8]. The Three Delays framework and its recent extensions predict that this is where male non-involvement translates most directly into mortality [9]-[11]. The qualitative strand proposes three converging structural mechanisms for the pattern rather than an attitudinal one: Sande society rules that prohibit male presence in the birth space, facility environments that provide no seating or sanitation for male companions, and a peer-ridicule norm that penalises men who attend. These are propositions from the joint-display matrix and are not tested by the survey.
Each of those three mechanisms has a counterpart elsewhere in the region. Malawian work found that men who wished to be involved were deterred by the physical arrangements of the facility and by the reactions of staff [55]; policymaker and practitioner interviews across the Pacific identified the same combination of institutional and normative exclusion [56]; and Ghanaian opinion leaders described male attendance at delivery as socially anomalous rather than individually undesirable [57]. What these accounts share with ours is that the obstacle is not located inside the individual man. Each is modifiable. None is modifiable by communication campaigns addressed to individual men.
Table 4. Male Involvement Index constructed continuum aggregates (women’s report, N = 500).
Continuum Aggregate |
Male Involvement Index Constructed Continuum Aggregates |
Domains Included (Number) |
Mean (SD) |
% of Respondents Scoring below Midpoint |
Interpretation |
Antenatal |
Cognitive + Financial + Logistical (3) |
24.47 (17.53) |
90.4 |
Material provision dominant |
Intrapartum |
Clinical Accompaniment
Only (1) |
19.68 (19.32) |
92.4 |
Lowest aggregate; single narrow domain |
Postnatal |
Psychosocial +
Decision-Making (2) |
36.07 (17.53) |
76.4 |
Elevated by the decision-making domain (43.66) |
Composite (All Domains) |
All Six (6) |
27.54 (15.52) |
89.2 |
Systemically below the midpoint throughout |
a. Scale midpoint is 50 out of 100. Allocation of domains to phases follows the WHO antenatal care framework [37]. b. The aggregates comprise three, one, and two domains respectively and are not equivalently constituted; differences between them reflect both any genuine variation in involvement by phase and the differing composition and reliability of the aggregates. The intrapartum aggregate is the clinical accompaniment domain alone (Cronbach’s alpha 0.207). c. The final column reports the percentage of respondents whose score fell below 50. The corresponding column in the submitted version was mislabelled: the values given there (approximately 75, 80, 60, and 73) were 100 minus the mean score, not the percentage of respondents below the midpoint, and are corrected here.
3.4. Factors Associated with Male Partner Participation: Bivariate Analysis
Table 5 presents the bivariate associations. Women’s autonomy emerged as the strongest single correlate of involvement (r = 0.362, p < 0.001), exceeding household wealth (r = 0.238, p < 0.001) and both measures of geographic access. This is a finding with considerable programmatic weight, because it indicates that male involvement and women’s decision-making power move together rather than trading off against one another. It is consistent with multilevel Demographic and Health Survey evidence linking married women’s autonomy to health care utilisation across high-fertility African settings [58], and with multi-country analyses showing that women scoring higher on validated empowerment indices report fewer barriers to obtaining permission to seek care [59]. Its relationship with education, and the consequences for the multivariable model, are addressed in Section 3.5.
Read alongside the critical literature warning that male engagement can shade into male control [22] [23] [60], the implication is that programming which raises involvement without simultaneously protecting women’s decision-making authority risks reinforcing the very gatekeeping it sets out to overcome. This is not a hypothetical concern. Critical reflections from gender-transformative programming have documented how interventions framed around male responsibility can, without deliberate design safeguards, restore rather than redistribute male authority [61], and evaluations in rural South Africa found that shifts in gender ideology were uneven and required sustained group work to hold [62]. The one large randomised trial in this space, the Bandebereho couples’ intervention in Rwanda, achieved gains in male accompaniment and reductions in intimate partner violence precisely because it was designed around relationship power rather than around male attendance alone [63].
Geographic isolation was significantly and inversely associated with involvement, for both distance (r = −0.193) and travel time (r = −0.181), each at p < 0.001. Residence in Western Area Urban was associated with higher involvement than elsewhere (32.50 versus 26.95; t = 2.473, df = 498, p = 0.014; d = 0.36). Neither maternal age nor parity reached significance in bivariate analysis, and Muslim and Christian respondents did not differ significantly (26.79 versus 28.78; t = −1.415, df = 482, p = 0.158; d = 0.13). One-way analysis of variance confirmed significant variation by education (F = 25.411, df = 4495, p < 0.001, η2 = 0.17, a large effect by conventional criteria), wealth (F = 9.138, df = 4495, p < 0.001, η2 = 0.07), and province (F = 7.450, df = 4495, p < 0.001, η2 = 0.06). Two features of these gradients deserve note. The wealth gradient is not strictly monotonic: mean involvement runs 20.70, 25.40, 30.67, 28.98, and 31.95 across quintiles one to five, so the fourth quintile falls below the third. And the lowest-scoring province is North West at 22.86 rather than Northern at 23.64, with Western Area highest at 32.26 and Southern close behind at 31.49.
Table 5. Bivariate associations between predictor variables and the MII composite score (N = 500).
Predictor |
Bivariate Associations between Predictor Variables and the MII Composite Score |
Test Statistic |
p |
Effect Size |
Direction/Group Means |
Women’s Autonomy Scale |
r = 0.362 |
<0.001 |
Moderate |
Higher autonomy, higher involvement |
Household Wealth Quintile |
r = 0.238 |
<0.001 |
Small to Moderate |
Q1 = 20.70 to Q5 = 31.95 |
Distance to Facility (km) |
r = −0.193 |
<0.001 |
Small |
Inverse |
Travel Time (min) |
r = −0.181 |
<0.001 |
Small |
Inverse |
Maternal Age (years) |
r = 0.065 |
0.148 |
Negligible |
Not significant |
Parity |
r = −0.033 |
0.457 |
Negligible |
Not significant |
Western Area Urban vs Other |
t = 2.473, df = 498 |
0.014 |
d = 0.36 |
32.50 vs 26.95 |
Muslim vs Christian |
t = −1.415, df = 482 |
0.158 |
d = 0.13 |
26.79 vs 28.78 |
Education Level (5 groups) |
F = 25.411, df = 4495 |
<0.001 |
η2 = 0.17 (large) |
None = 19.86 to post-secondary = 35.68 |
Wealth Quintile (5 groups) |
F = 9.138, df = 4495 |
<0.001 |
η2 = 0.07 |
20.70, 25.40, 30.67, 28.98, 31.95 |
Province (5 groups) |
F = 7.450, df = 4495 |
<0.001 |
η2 = 0.06 |
Western = 32.26; North West = 22.86 |
a. MII = Male Involvement Index composite (women’s report, 0 - 100). Effect size conventions follow Cohen: r small = 0.10, medium = 0.30; Cohen’s d small = 0.20, medium = 0.50; η2 small = 0.01, medium = 0.06, large = 0.14. b. Degrees of freedom corrected. The education, wealth, and province analyses of variance each use all 500 cases and therefore have df = 4,495; the submitted version reported df = 4,493 for education in error. The Muslim versus Christian t-test compares 315 Muslim and 169 Christian respondents, giving an analytic sample of 484 and df = 482; the submitted version reported df = 483, which is not attainable for an independent-samples test on those group sizes. c. Cohen’s d corrected. The values now reported are computed from the difference in group means divided by the pooled standard deviation. The submitted version converted the point-biserial correlation to d using a formula that assumes equal group sizes, which understates d where groups are as unequal as 53 versus 447. d. The women’s autonomy scale is a five-item composite from Instrument A with a Cronbach’s alpha of 0.688. The submitted version gave 0.855 at this point, which is the alpha of the 30-item MII composite, not of the autonomy scale.
3.5. Multivariable Analysis of Factors Associated with Involvement
The regression model was statistically significant (F = 17.708, df = 7,492, p < 0.001) and accounted for 20.1% of the variance in the composite (R2 = 0.201; adjusted R2 = 0.190). Four covariates retained independent significance (Table 6). Education was by a clear margin the strongest correlate (B = 4.021, 95% CI: 3.068 - 4.974, p < 0.001), and it was also the strongest single correlate of involvement in bivariate terms (r = 0.405, p < 0.001). Because the variable is coded across five ordered levels, the coefficient implies a fitted difference of approximately 16 MII points between a man whose partner has no formal schooling and one whose partner has post-secondary education, which is the largest structural gradient documented anywhere in the study. Household wealth quintile (B = 1.162, p = 0.014), travel time (B = −0.025, p = 0.013), and maternal age (B = 0.194, p = 0.047) were also independently associated. Parity, Muslim religion, and Western Province residence were not significant after adjustment, which is worth emphasising: the apparent provincial and religious differences observed at the bivariate level are largely explained by the education, wealth, and access differences that accompany them.
Refitting the model with cluster-robust standard errors at the district level left every substantive conclusion intact and, in two respects, strengthened it (Table 6, final columns). Education (p < 0.001), wealth (p = 0.027), travel time (p = 0.006), and maternal age (p = 0.015) all retained significance, and parity, which was marginal under ordinary least squares (p = 0.072), reached significance under robust estimation (B = −0.898, p = 0.024). Religion and province remained non-significant. Because between-district variation is being modelled explicitly, robust standard errors are smaller than model-based ones for several covariates; this is a legitimate consequence of the design rather than an artifact, but the parity result is reported as suggestive only, since it rests on 15 clusters and does not replicate under model-based estimation.
Women’s autonomy is the strongest bivariate correlate of involvement, and its absence from the multivariable model requires justification. Two considerations bear on it. The first is conceptual and was set out in Section 2.6: autonomy is a relational property of the same household bargaining process that produces involvement, as plausibly consequent upon it as prior to it, and entering it as a covariate would treat a co-determined quantity as exogenous. The second is empirical and is decisive. Autonomy correlates 0.807 with education level in this sample. Adding it to the seven pre-specified covariates raises R2 from 0.2012 to 0.2017, a change of 0.0005 that does not approach significance (partial F = 0.303, p = 0.582); its own coefficient is small and non-significant (B = 0.188, 95% CI: −0.483 to 0.859, p = 0.582); and its inclusion inflates the variance inflation factor for education from 1.12 to 2.96 and returns 3.32 for autonomy itself, while leaving education significant at B = 3.679 (p < 0.001). The two variables are, in these data, largely measuring the same underlying resource. Table 7 reports the sensitivity model in full so that the exclusion can be judged rather than taken on assertion.
This pattern of correlates aligns with, and sharpens, the regional evidence. Education has emerged as a dominant predictor in Ethiopian meta-analytic work and in multi-country Demographic and Health Survey analyses across West Africa [19] [64], and it predicts completion of the recommended antenatal schedule across 36 sub-Saharan African countries [65] and correlates with antenatal utilisation across 31 [66]. Community studies reach the same conclusion on a smaller scale: educational attainment predicted male antenatal involvement in southern Ethiopia [67], in central Tanzania [68], and in a refugee settlement in northern Uganda, where distance to the facility and long clinic waiting times also emerged as independent constraints [69]. Theoretically, the results sit comfortably within the Andersen behavioural model, in which enabling resources such as education, wealth, and geographic access dominate predisposing characteristics such as religion and parity in shaping health-related behaviour [70].
Table 6. Ordinary least squares regression of factors independently associated with the MII composite score (n = 500), with model-based and cluster-robust inference.
Predictor |
Ordinary Least Squares Regression of Factors Independently Associated with the MII Composite Score |
B |
SE |
95% CI |
t |
p |
Robust SE |
Robust p |
Intercept |
15.927 |
2.769 |
10.487 to 21.368 |
5.75 |
<0.001 |
— |
— |
Education Level |
4.021 |
0.485 |
3.068 to 4.974 |
8.29 |
<0.001 |
0.631 |
<0.001 |
Household Wealth Quintile |
1.162 |
0.472 |
0.234 to 2.089 |
2.46 |
0.014 |
0.525 |
0.027 |
Travel Time (min) |
−0.025 |
0.010 |
−0.045 to −0.005 |
−2.49 |
0.013 |
0.009 |
0.006 |
Maternal Age (years) |
0.194 |
0.097 |
0.002 to 0.385 |
1.99 |
0.047 |
0.080 |
0.015 |
Parity |
−0.898 |
0.499 |
−1.878 to 0.081 |
−1.80 |
0.072 |
0.399 |
0.024 |
Muslim Religion |
−0.493 |
1.310 |
−3.067 to 2.081 |
−0.38 |
0.707 |
1.187 |
0.678 |
Western Province |
0.893 |
1.751 |
−2.547 to 4.332 |
0.51 |
0.610 |
1.539 |
0.562 |
Model Fit |
R2 = 0.201 |
Adj. R2 = 0.190 |
F = 17.708, df = 7492 |
|
<0.001 |
|
|
a. Dependent variable is the MII composite (women’s report, 0 - 100). B is the unstandardised coefficient, expressed as a change in MII points per one-unit increase in the predictor. Religion is dummy coded (Muslim = 1, Christian = 0); province is dummy coded (Western = 1, other = 0). b. Analytic sample corrected: no case was excluded, and no predictor contained missing values, so all 500 observations were retained. The footnote to this table in the submitted version stated that two cases were excluded for missing predictor data, which was incorrect and is inconsistent with the residual degrees of freedom of 492 reported in the same table. c. Robust standard errors and p-values derive from cluster-robust (sandwich) variance estimation with district (15 clusters) as the clustering unit. The coefficients themselves are unchanged. d. Confidence intervals shown are model-based.
3.6. Barriers Reported by Male Partners
Men’s self-reported barriers reproduced the structural hierarchy visible in the regression results (Table 8, Figure 4). Work and livelihood obligations were reported by 272 of 500 men (54.4%), geographic distance by 196 (39.2%), and financial constraints by 176 (35.2%). Cultural or normative barriers were reported by 150 (30.0%), feeling unwelcome at the facility by 73 (14.6%), embarrassment
Table 7. Sensitivity model adding women’s autonomy to the pre-specified covariates (n = 500).
Predictor |
Sensitivity Model Adding Women’s Autonomy to the Pre-Specified Covariates |
B (Main Model) |
B (with Autonomy) |
95% CI (with Autonomy) |
p (with Autonomy) |
VIF (with Autonomy) |
Education Level |
4.021 |
3.679 |
2.131 to 5.228 |
<0.001 |
2.96 |
Household Wealth Quintile |
1.162 |
1.066 |
0.076 to 2.055 |
0.035 |
1.30 |
Travel Time (min) |
−0.025 |
−0.026 |
−0.045 to −0.006 |
0.012 |
1.11 |
Maternal Age (years) |
0.194 |
0.190 |
−0.002 to 0.382 |
0.052 |
1.73 |
Parity |
−0.898 |
−0.892 |
−1.873 to 0.089 |
0.075 |
1.72 |
Muslim Religion |
−0.493 |
−0.470 |
−3.047 to 2.108 |
0.721 |
1.03 |
Western Province |
0.893 |
0.853 |
−2.592 to 4.298 |
0.627 |
1.13 |
Women’s Autonomy |
not entered |
0.188 |
−0.483 to 0.859 |
0.582 |
3.32 |
Model Fit |
R2 = 0.2012 |
R2 = 0.2017 |
ΔR2 = 0.0005 |
Partial F = 0.303, p = 0.582 |
|
a. Added in response to review. The women’s autonomy scale correlates 0.807 with education level and 0.432 with wealth quintile in this sample. Adding it changes R2 by 0.0005, does not approach significance on a partial F test, and raises the variance inflation factor for education from 1.12 to 2.96. Education remains significant with the autonomy term in the model. Variance inflation factors in the main model ranged from 1.03 to 1.72. b. The main model coefficients in the second column are reproduced from Table 6 for comparison.
by 39 (7.8%) and partner or family refusal by only 10 (2.0%). The ordering matters. Behaviour-change communication addressed to individual men presumes that the binding constraint is belief. These data indicate that for the majority, it is income. The three commonest male occupations in the study communities, namely subsistence farming, daily-wage motorcycle taxi riding, and petty trading, share the property that a day away from work is a day without earnings. Antenatal clinic hours are 07:30 to 12:00; therefore, they do not represent an inconvenience for such men; they represent a structural impossibility. Nigerian analyses of antenatal non-use reach a parallel conclusion, identifying cost and distance rather than attitude as the principal deterrents [71], and Ghanaian women themselves identify men’s work commitments and the absence of male-accommodating space as the leading obstacles to accompaniment [72]. The finding reinforces a conclusion increasingly evident in recent scoping reviews of enablers and barriers across sub-Saharan Africa, namely that the binding constraints on male involvement in low-income settings are principally material rather than attitudinal [18] [30].
That said, the cultural and facility-level barriers reported by the minority should not be dismissed as residual. Qualitative work in Limpopo Province, South Africa, documented that cultural prohibitions can operate absolutely rather than probabilistically, foreclosing attendance entirely for the men subject to them [73], and Sesotho-speaking communities in the Free State described a normative frame in which male attendance signals suspicion of the partner rather than support for her [74]. A barrier endorsed by 30% of men in aggregate may be decisive for those 30%. Ugandan and Malawian work similarly found that men who did attend often did so despite, rather than in the absence of, normative disapproval [75] [76], and systematic review evidence from prevention of mother-to-child transmission programs across sub-Saharan Africa shows that facility-level unfriendliness is one of the most consistently reported deterrents across the region [77].
![]()
Figure 4. Barriers to involvement reported by male partners (N = 500, multiple responses permitted). Bars are shaded by barrier class. The three most frequently reported barriers are all structural.
Table 8. Barriers to involvement reported by male partners (N = 500, multiple responses permitted).
Barrier |
Barriers to Involvement Reported by Male Partners |
n |
% |
Barrier Class |
Principal Mechanism |
Work/Livelihood Obligations |
272 |
54.4 |
Structural |
Daily-wage income loss; clinic hours 07:30 - 12:00 |
Geographic Distance |
196 |
39.2 |
Structural |
Mean travel time 53.6 min; Phase II delay enabler |
Financial Constraints |
176 |
35.2 |
Structural |
Indirect costs persist despite fee abolition |
Cultural/Normative Barriers |
150 |
30.0 |
Cultural |
Sande birth-space rules; peer ridicule |
Feeling Unwelcome at Facility |
73 |
14.6 |
Supply-Side |
No seating or sanitation; staff dismissiveness |
Embarrassment at the Clinic |
39 |
7.8 |
Social |
Normative disapproval of attending men |
Partner or Family Refusal |
10 |
2.0 |
Relational |
Rare; within-household objection |
a. Percentages exceed 100 because respondents could select all applicable barriers. b. Chi-square tests of each barrier against facility-based delivery were non-significant, with all p-values at or above 0.23 (smallest p = 0.233, feeling unwelcome at the facility); the submitted version stated all p > 0.28. This is consistent with barriers acting on outcomes indirectly through involvement.
None of the individual barriers was significantly associated with facility delivery in direct chi-square testing (all p ≥ 0.23; the smallest p-value, 0.233, was for feeling unwelcome at the facility). This is consistent with barriers operating on outcomes indirectly, through their association with involvement, rather than directly, and it argues against treating barrier endorsement as a proxy for outcome risk.
4. Conclusions and Recommendations
This study provides the first nationally distributed dyadic measurement of multidimensional male partner involvement in Sierra Leone. Four conclusions follow. Involvement is low: the composite score was 27.54 out of 100, and every domain fell below the scale midpoint. It is lowest on the behaviour required where obstetric risk is highest, clinical accompaniment standing at 19.68 during the phase when the leading cause of maternal death can kill within two hours, although the aggregate representing that phase is constituted by a single domain and is not comparable like-for-like with the multi-domain antenatal and postnatal aggregates. Involvement is patterned structurally: education, household wealth, travel time, and maternal age were each independently associated with it, while religion and parity were not, and work, distance, and cost outranked cultural norms in men’s own accounts of what prevents them. The fourth conclusion is methodological. Single-reporter measurement classified 36.8% of couples differently under a median split and 59.0% under a tertile split, so the discordance rate is a function of the cut-point and should always be reported with it. Dyadic measurement should be a design requirement rather than a refinement. Because the design is cross-sectional, none of these associations establishes causation, and reverse and reciprocal pathways remain open.
Three recommendations follow, framed as associations rather than demonstrated effects. First, male engagement programming should address the structural barriers that men themselves report, principally clinic scheduling incompatible with daily-wage work, the cost and time of reaching a facility, and facility environments that make no provision for male companions, rather than relying on behaviour change communication alone; the implementation literature reaches the same conclusion across African settings [78]. Second, programming should be designed to protect and extend women’s decision-making authority at the same time as it raises male participation, in line with World Health Organization guidance on health promotion interventions for maternal and newborn health [79] and with the gender-transformative trial evidence cited above [63]. The current legislative environment in Sierra Leone makes this an opportune moment for such design [80]. Third, involvement should be measured multidimensionally and dyadically, with the postnatal period treated as a distinct measurement occasion rather than inferred from antenatal behaviour.
5. Limitations of the Study
Six limitations qualify these findings. First, the cross-sectional design precludes causal inference. Associations between involvement and its correlates may be bidirectional, and educated men may both be more involved and more likely to partner with educated women. Formal mediation analysis requires longitudinal data or bootstrapped approaches and is reserved for a later paper.
Second, the sampling design constrains the generalisability of the prevalence estimates. Selection probabilities were not proportional to size at any stage; distal villages were included purposively to maximise variation in access, and no district-level sampling frame of recently delivered women with a contactable partner existed from which design weights could be constructed. The estimates are therefore unweighted and describe the sampled dyads rather than the national population. Clustering was addressed using intra-cluster correlations, design effects, and cluster-robust standard errors, but the district is a coarser unit than the primary sampling unit actually used because the number of Peripheral Health Units and catchment communities visited was recorded in the field logs rather than in the analytic file. A fully design-based variance estimator could not, therefore, be fitted, and the adjustments reported here should be read as conservative with respect to between-district variation and silent with respect to residual clustering within districts. Reporting the primary sampling unit identifier in the analytic file is a straightforward remedy for future rounds.
Third, involvement was measured retrospectively for the index pregnancy and is therefore subject to recall bias. The dyadic design allows reporting divergence to be characterised, but the systematically lower male self-report is open to more than one reading, of which under-claiming of behaviours that carry no masculine credit is only one.
Fourth, the Male Involvement Index is new. The full composite’s internal consistency was good, but three of the six domains returned low alpha coefficients. This is a property of causal-indicator domains rather than a defect [38] [39], but it means those domains should not be treated as unidimensional scales, and it bears directly on the constructed continuum aggregates in Section 3.3, one of which rests on a single low-alpha domain. Confirmatory factor analysis and test-retest reliability are required before the index is recommended for routine use in West Africa.
Fifth, the achieved sample of 500 dyads fell short of the planning target of approximately 855, principally because eligible women lacked a contactable or consenting partner. Power was sufficient for the primary comparisons but limited for district-level estimates, and the smallest district cells should be interpreted with caution. The paired design also excluded women without contactable partners, which removes the most extreme cases of non-involvement and may bias the composite upwards. Karene District was not covered, so the results describe 15 of Sierra Leone’s 16 districts.
Sixth, a data quality issue is disclosed. Fifty-one records (10.2%) carry a recorded delivery date between one and nine months after the interview date. These are most plausibly expected delivery dates entered in place of actual ones, or keying errors. They are excluded from the recall-interval statistics reported in Section 2.3 and retained in all analyses that do not depend on the delivery date, none of which is affected. The discrepancy is being reconciled against the original field submissions, and the affected records are flagged in the archived dataset.
Acknowledgements
The authors thank the 500 women and 500 men who gave their time to this study; the District Health Management Teams of all participating districts and the paramount chiefs who granted community assent; the peripheral health unit in-charges and maternal and child health aides who facilitated recruitment; and the 20 final-year BSc Nursing students of Njala University who served as enumerators.
Author Contributions
Conceptualisation: E.J.S. and S.J.B.; Methodology: E.J.S.; Software: E.J.S.; Validation: E.J.S., S.J.B., and P.T.M.; Formal analysis: E.J.S.; Investigation: E.J.S.; Resources: E.J.S.; Data curation: E.J.S.; Writing—original draft preparation: E.J.S.; Writing—review and editing: E.J.S., S.J.B., P.T.M., and J.J.M.; Visualisation: E.J.S.; Supervision: S.J.B. and P.T.M.; Project administration: E.J.S. All authors have read and agreed to the published version of the manuscript.