Non-sampling error

Indian Economy glossary

Topic: Economic Data: Census, NSS, Surveys and Statistical Tools · NCERT: Class 11, Ch 2 "Collection of Data"; Class 11, Ch 8 "Use of Statistical Tools"

Meaning

Non-sampling error is any error in survey or census data that does not come from studying only a part of the population. It comes from the plan, the list used to pick the sample, the people who answer, and the way answers are recorded.

It matters because NCERT calls it more serious than sampling error. A bigger sample does not remove it. Even a census (complete enumeration, which counts every unit) has it. Official numbers on poverty, jobs and population can be biased even when the survey is very large.

Explanation

Where it fits: the two families of error

  • A parameter is the true value for the whole population, for example the true mean income of all farmers in Manipur.
  • An estimate is the value we calculate from our data.
  • An error is any gap between the estimate and the parameter. There are two kinds:
  • Sampling error comes from studying only a part. Sampling error = population parameter − sample estimate. It is a matter of chance, not anyone's mistake.
  • Non-sampling error covers everything else: a poor design, an incomplete list, people who refuse, and wrong recording.

  • NCERT worked example (sampling error, for contrast):

  • Incomes of 5 farmers: 500, 550, 600, 650, 700. Population mean = 3000 ÷ 5 = 600.
  • A sample of (500, 600) has a mean of 550, so the sampling error is 600 − 550 = 50.
  • A sample of four has a mean of 575, so the error is 25. A census of all five gives an error of 0.
  • Lesson: sampling error shrinks as the sample grows. Non-sampling error does not follow this rule.

The three forms (NCERT)

  • (a) Sampling bias
  • The sampling plan cannot reach part of the population, so that part is never in the sampling frame (the list or map from which the sample is drawn).
  • Examples: phone or online surveys miss households with no phone or internet, which are often poorer and rural. Stale frames miss new colonies and new urban areas.
  • Effect: the error is one-sided. If poor households are left out, average income comes out too high.

  • (b) Non-response error

  • People in the sample cannot be contacted or refuse to answer.
  • The harm depends on who drops out. If the people who refuse (for example, the rich) are different from those who answer, the estimate becomes biased.

  • (c) Error in data acquisition (measurement error)

  • The right person was reached, but a wrong answer gets recorded.
  • NCERT examples:
    • Instrument differences: students' measuring tapes differ, so they get different lengths for the same table.
    • Price variation: orange prices differ by shop, market and quality, so only averages can be recorded.
    • Recall lapses: people forget past spending or events.
    • Transcription slips: writing 13 for 31.

Why it is the bigger danger

  • A bigger sample does not cure it.
  • More interviews → more chances of refusals and recording slips → the error can even grow.

  • A census has it too.

  • Counting everyone removes sampling error.
  • But it does not fix bad questions, missed houses or wrong answers.

  • It is hard to measure.

  • Sampling error is random and can be measured with statistics.
  • Non-sampling error is often a bias (it pushes the result in one direction), so it does not cancel out.

  • What reduces it: better frames, better questionnaires, trained field staff, CAPI (Computer Assisted Personal Interviewing, where answers go straight into a tablet) and careful checking of data.

Worked example: measuring non-sampling error in a census

  • The Post Enumeration Survey (PES) checks how many people a census missed or counted twice.
  • Net coverage error = undercount (omissions) − overcount (duplications).
  • Suppose a PES finds 30 omissions and 5 duplications per 1,000 true persons:
  • Net undercount = 30 − 5 = 25 per 1,000.
  • Net omission rate = 25 ÷ 1,000 = 2.5%.

  • The census has zero sampling error, yet it has a 2.5% error here. All of that error is non-sampling error.

In India

Institutions that fight it

  • The National Statistics Office (NSO) under MoSPI now collects primary data through CAPI or web applications [2].
  • These tools have built-in validation checks. For example, the tablet flags an age of 13 for a head of household who is "married with 4 children" [2].
  • This cuts data-entry and transcription errors at the moment of collection [2].

  • Other NSO safeguards [2]:

  • questionnaires designed with experts and field staff
  • multi-layer training of field officials on concepts, coverage and handling respondent behaviour, which helps against non-response
  • multi-level data scrutiny and follow-up visits.

  • The National Quality Assurance Framework (NQAF) of MoSPI sets the wider quality standard.

Census coverage checks: non-sampling error inside a census

  • The Registrar General of India (RGI) runs the Post Enumeration Survey (PES) after each census.
  • It measures coverage errors (people missed or counted twice) and content errors (wrong details recorded).
  • An independent P-sample (population sample) is interviewed again, and each person is matched against the census records [5].

  • Omissions usually exceed duplications, so most censuses show a net undercount [4].

  • India's net omission rate rose slowly, from 0.68% (1961) to 2.33% (2001) [4].
  • Omissions are highest among children under 5 and young adults (PES 1981, 1991, 2001) [4].

The rich don't answer (non-response and under-reporting)

  • Rich households refuse the survey or report less spending than they really have.
  • Result: NSS consumption falls short of Private Final Consumption Expenditure (PFCE), the national-accounts estimate of what all households spend.
  • The NSS-NAS gap (the gap between survey consumption and national-accounts consumption) has widened over time.
  • top-end spending is missed → average MPCE (Monthly Per Capita Consumption Expenditure) comes out too low → inequality looks smaller than it is, and poverty estimates can shift.

Recall period changes the answer (data-acquisition error)

  • A recall period is how far back a person is asked to remember. People report more spending on frequent items with a 7-day recall than with a 30-day recall, because they forget small purchases over a month.
  • Uniform Reference Period (URP): a 30-day recall for all items.
  • Modified Mixed Reference Period (MMRP) uses three recall periods [3]:
  • last 7 days: edible oil, egg, fish, meat, milk and milk products, vegetables, fruits, spices, beverages, processed food, pan and tobacco
  • last 365 days: clothing, bedding, footwear, education, institutional medical care and durable goods
  • last 30 days: all other items.

  • HCES 2022-23 (Household Consumption Expenditure Survey, fieldwork August 2022 – July 2023) used MMRP and covered 2,61,746 households (1,55,014 rural and 1,06,732 urban) [3].

  • Average MPCE without imputation (no money value added for free items): ₹3,773 rural and ₹6,459 urban (2022-23), against ₹1,430 and ₹2,630 in 2011-12 (NSS 68th round) [3].
  • With imputed values of free items such as NFSA/PMGKAY grain, laptops, bicycles and school uniforms: ₹3,860 rural and ₹6,521 urban (2022-23) [3].
  • Free health care under PM-JAY and free education were not imputed, because HCES is not a record-based survey and their value cannot be checked [3].
  • Lesson: the choice of recall period and imputation changes the headline number, and neither is a sampling error.

Aging frames (sampling bias)

  • NSS and PLFS (Periodic Labour Force Survey) draw their village and urban-block frames from Census 2011.
  • The census due in 2021 was postponed and is now awaited as Census 2027, so the frames have grown stale.
  • new towns, peri-urban growth and migrant colonies are under-covered → estimates of urbanisation, jobs and consumption are biased.

Don't confuse with

  • Sampling error: it comes only from studying a part. It shrinks with a bigger sample and is zero in a census. Non-sampling error does not shrink with sample size and is present in a census.
  • Sampling bias vs non-response error: in sampling bias, part of the population is never in the frame, for example households without phones in a phone survey. In non-response error, the unit was selected but refused or could not be contacted.
  • Data-acquisition error vs non-response: in data-acquisition error the person did answer, but a wrong value was recorded, for example 13 written for 31 or a forgotten purchase. In non-response, no answer was collected.
  • Coverage error vs content error (PES terms): coverage error means people were missed or counted twice. Content error means people were counted but their details were recorded wrongly. Both are non-sampling errors.

Prelims Hooks

  • NCERT's three forms of non-sampling error are sampling bias, non-response error and error in data acquisition.
  • Increasing the sample size reduces sampling error only. It does not reduce non-sampling error and may even increase it.
  • A census has zero sampling error but does have non-sampling error. Any statement that a census is "error-free" is false.
  • Recording 13 as 31 is a transcription error, which is a data-acquisition (non-sampling) error. It is not a sampling error.
  • The Post Enumeration Survey is run by the RGI to measure census coverage and content error. India's net omission rate was 2.33% in 2001 [4].
  • CAPI with built-in validation checks is used by NSO (MoSPI) to cut data-entry errors [2]. MMRP uses 7-day / 30-day / 365-day recall periods by item group [3].

Mains Points

  • Quality over quantity in official surveys.
  • A very large but badly run survey can be less accurate than a small, careful one, because non-sampling error does not fall as the sample grows.
  • Spending on better frames, training and CAPI often buys more accuracy than adding households. Even a full census left a 2.33% net undercount (2001) [4].

  • Survey bias and welfare policy (GS-III).

  • Rich households refuse or under-report, and recall-period and imputation choices shift MPCE (for example ₹3,773 vs ₹3,860 rural in 2022-23, without vs with imputation) [3].
  • So poverty lines, inequality measures and CPI weights built on HCES can be biased. This supports checking surveys against national accounts and administrative data (tax, GST, UPI).

  • Census delay as a data-governance issue (GS-II).

  • Survey frames still date from Census 2011 while Census 2027 is awaited. Stale frames bias PLFS, HCES and ASUSE (Annual Survey of Unincorporated Sector Enterprises) estimates, especially for urban and migrant populations.
  • A timely census is basic statistical infrastructure. Publishing standard errors (for sampling error) and non-response rates (for non-sampling error), along with NSO's safeguards [2] and the NQAF, would build trust in official data.

Related concepts

Read more

Sources

  1. 1Class 11, Ch 2 "Collection of Data"; Class 11, Ch 8 "Use of Statistical Tools" (primary)
  2. 2PIB — National Statistics Office (NSO) under MoSPI is committed to ensuring accurate and reliable data while minimizing non-sampling errorspib.gov.in · tier 1
  3. 3MoSPI/NSSO — Press Release: Results of Household Consumption Expenditure Survey 2022-23 (24 February 2024)mospi.gov.in · tier 1
  4. 4UN DESA Population Division — Census counts, undercounts and population estimates (Technical Paper 2020)un.org · tier 2
  5. 5UN Statistics Division — Post Enumeration Surveys (handbook)unstats.un.org · tier 2