MARATTO

article · Journal of Statistical Sciences and Computational Intelligence

Regression estimator of population mean with random missing values in stratified two stage sampling using double sampling for auxiliary information

2025Open accessBenue State University

In plain language

Missing data frequently impairs the reliability and efficiency of statistical estimates drawn from surveys. To address this, an evaluation was conducted on a regression-type estimator designed to calculate population means within a stratified two-stage sampling structure containing incomplete observations. The assessment utilised field survey records consisting of school attendance records as auxiliary data and mathematics test scores as the primary variable. Within this setup, schools formed primary sampling units while students represented secondary units. Missing values were simulated by setting twenty percent of test scores as missing completely at random, which were subsequently replaced using ratio and regression imputation methods. Tested across multiple sample sizes against existing ratio and difference estimators, the regression estimator consistently yielded lower variances and coefficients of variation. Its relative efficiency rose with larger sample sizes, particularly when paired with regression imputation.

Key takeaways

  • A regression-type estimator was tested for estimating population means under stratified two-stage sampling with missing data.
  • The estimator demonstrated lower variances and coefficients of variation compared to existing ratio and difference estimators.
  • Performance and efficiency gains increased as sample sizes expanded from 25 to 100.
  • Superior accuracy was maintained in the presence of random missing values, especially when using regression imputation.

Why it matters

Real-world surveys routinely suffer from incomplete responses, which can distort results and weaken statistical confidence. Demonstrating that a regression estimator handles missing values effectively provides survey practitioners with a more dependable method for calculating population averages from multi-stage field data, such as educational testing records.

Commercialisation angle

The method represents early-stage analytical research that could be integrated into survey processing software, educational assessment platforms, or statistical toolkits used by government statistical agencies and polling organisations. While demonstrated on empirical school survey data, the abstract does not indicate that software packages or direct commercial tools have yet been built.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

Missing data is a recurring challenge in survey sampling, often reducing the efficiency and reliability of estimators. This study proposed and investigated the regression-type estimator of the population mean under a stratified two-stage sampling design in the presence of missing values. The population of the study comprised of Field survey data on students’ school attendance (auxiliary variable) and mathematics test scores (study variable). A stratified two-stage design was adopted, with schools as primary sampling units and students as secondary units. To reflect item nonresponse, 20% of the study variable was declared missing completely at random (MCAR) and handled through both regression and ratio imputation. The performance of the proposed regression estimator was compared with the Bahl-Saini (2011) ratio and difference estimator across sample sizes of 25, 40, 70, and 100, using coefficient of variation (CV), and confidence intervals as evaluation criteria. Results showed that the regression estimator consistently achieved lower variances and CVs than the existing estimators, with efficiency improving as sample size increased. Even in the presence of missing data, the regression estimator maintained superior performance, particularly under regression imputation. The study demonstrated the efficiency of the regression estimator in handling incomplete data and highlights its practical significance for reliable estimation in complex survey designs.

Research topics

  • Survey Sampling and Estimation Techniques
  • HIV, Drug Use, Sexual Risk
  • Survey Methodology and Nonresponse

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.64497/jssci.128

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.