Sampling and non-sampling errors

Economic Data: Census, NSS, Surveys and Statistical Tools · section 4 of 12

In this note
  1. Detail
  2. Prelims Hooks
  3. Mains Points

Detail

1. Why errors arise at all

  • A population (or universe) is the whole group we want to study, for example all farmers in Manipur.
  • A parameter is a true value for the whole population, such as its true mean income.
  • A sample is a small part of the population, picked to stand for the whole.
  • An estimate (or statistic) is the value we calculate from the sample.
  • An error is any gap between the estimate and the true parameter. There are two families of error:
  • Sampling error comes from studying only a part.
  • Non-sampling error comes from everything else: the plan, the list, the people answering, and the recording.

2. Sampling error

  • Definition: Sampling error = population parameter − sample estimate.
  • It happens because a sample never matches the population exactly. It is a matter of chance, not a mistake by anyone.
  • NCERT worked example (Manipur farmers):
  • Incomes of 5 farmers: 500, 550, 600, 650, 700.
  • Population mean = 3000 ÷ 5 = 600.
  • Sample of two farmers (500, 600): sample mean = 1100 ÷ 2 = 550.
  • Sampling error = 600 − 550 = 50.

  • Extension (same data):

  • A sample of four (500, 550, 600, 650) has a mean of 2300 ÷ 4 = 575, so the error is 600 − 575 = 25.
  • All five together form a census. The mean is 600, so the error is 0.
  • Lesson: as the sample grows, sampling error shrinks, and it disappears in a census.

  • Key property: sampling error exists only in sample surveys. A census (complete enumeration, which covers every unit) has no sampling error.

  • It can be measured with statistics and made smaller by:
  • taking a larger sample
  • using a better design, such as stratified random sampling, where the population is split into groups and each group is sampled.

3. Non-sampling error: the bigger danger

  • Definition: Non-sampling error is any error that does not come from picking only a part of the population.
  • NCERT calls it more serious than sampling error for two reasons:
  • A bigger sample does not cure it. More interviews bring more chances for refusals and recording slips, so it can even grow.
  • A census has it too. Covering everyone does not remove bad questions, missed houses or wrong answers.

  • It takes three forms.

(a) Sampling bias

  • Definition: the sampling plan cannot reach part of the target population, so that part is never in the frame. A sampling frame is the list or map from which the sample is drawn.
  • Examples:
  • Phone or online surveys miss households without a phone or internet connection. These are often poorer and rural households.
  • Stale frames miss new settlements, new urban areas and new colonies that grew after the list was made.

  • Effect: the error is one-sided. For example, if poor households are left out, average income comes out too high.

(b) Non-response error

  • Definition: sampled people cannot be contacted or refuse to answer, so the sample that actually answers no longer represents the population.
  • The harm depends on who drops out. If refusers differ from those who answer (for example, the rich), the estimate becomes biased.

(c) Error in data acquisition (measurement error)

  • Definition: incorrect responses get recorded, even when the right person was reached.
  • NCERT examples:
  • Instrument differences: students' measuring tapes differ, so they get different lengths for the same table.
  • Price variation: orange prices differ by shop, market and quality, so only averages can be recorded.
  • Recall lapses: informants forget past spending or events.
  • Transcription slips: writing 13 for 31.

4. Comparison table

Sampling error Non-sampling error
Source Studying a part, not the whole Design, frame, response, recording
Larger sample Reduces it Does not reduce it (may even raise it)
In a census Absent Present
Nature Random; can be measured statistically Often one-sided (bias); hard to measure
Cure Bigger or better-designed sample Better frames, questionnaires, training, CAPI, scrutiny

5. Indian illustrations

The rich don't answer (non-response and under-reporting)

  • Affluent households refuse the survey or under-report their spending.
  • Result: NSS consumption falls short of Private Final Consumption Expenditure (PFCE). PFCE is the national-accounts estimate of what all households spend.
  • The NSS-NAS gap (the gap between survey consumption and national-accounts consumption) has widened over time.
  • Chain of effect:
  • top-end spending is missed
  • average MPCE (Monthly Per Capita Consumption Expenditure) comes out too low
  • inequality is understated, and poverty estimates can shift.

Recall period changes the answer (data-acquisition error)

  • Recall period means how far back the respondent is asked to remember.
  • The same household reports more spending on frequent items with a 7-day recall than with a 30-day recall, because people forget small purchases over a month.
  • Uniform Reference Period (URP): a 30-day recall for all items.
  • Modified Mixed Reference Period (MMRP) splits items into three recall periods [3]:
  • last 7 days: edible oil, egg, fish, meat, milk and milk products, vegetables, fruits, spices, beverages, processed food, pan and tobacco
  • last 365 days: clothing, bedding, footwear, education, institutional medical care and durable goods
  • last 30 days: all other items.

  • HCES 2022-23 (Household Consumption Expenditure Survey, fieldwork August 2022 – July 2023) used MMRP [3]:

  • It covered 2,61,746 households: 1,55,014 rural and 1,06,732 urban [3].
  • Average MPCE at current prices (without imputation, which means without adding a money value for free items): ₹3,773 rural and ₹6,459 urban (2022-23), against ₹1,430 and ₹2,630 (2011-12, NSS 68th round) [3].
  • With imputed values of free items from welfare schemes (NFSA/PMGKAY grain, laptops, bicycles, school uniforms and so on): ₹3,860 rural and ₹6,521 urban (2022-23) [3].
  • Free health care under PM-JAY and free education were not imputed. The reason: HCES is not a record-based survey, so their value cannot be checked [3].
  • Lesson: the choices of recall period and imputation change the headline number. Neither is a sampling error.

Coverage checks on the Census (non-sampling error in a census)

  • The Post Enumeration Survey (PES) is run by the Registrar General of India (RGI) after the census. It measures the census's own coverage errors (people missed or counted twice) and content errors (wrong details recorded).
  • Method: an independent P-sample (population sample) is interviewed again. Each person is then matched against the census records [5].
  • Net coverage error = undercount (omissions) − overcount (duplications). Omissions usually exceed duplications, so most censuses show a net undercount [4].
  • Worked example: suppose a PES finds 30 omissions and 5 duplications per 1,000 true persons. The net undercount is 30 − 5 = 25 per 1,000, so the net omission rate is 2.5%.

  • India's record: PES results show the net omission rate rising slowly, from a low of 0.68% (1961) to 2.33% (2001) [4].

  • Omissions are highest among children under 5 and young adults (PES 1981, 1991, 2001) [4].

  • Lesson: this is direct proof that a census, with zero sampling error, still carries non-sampling error.

Aging frames (sampling bias)

  • NSS and PLFS (Periodic Labour Force Survey) frames of villages and urban blocks come from Census 2011.
  • The census due in 2021 was postponed and is now awaited as Census 2027. Meanwhile, the frames have grown stale.
  • Effect:
  • new towns, peri-urban growth and migrant colonies are under-covered
  • estimates of urbanisation, employment and consumption carry bias.

Technology helps (reducing acquisition error)

  • The National Statistics Office (NSO) under MoSPI now collects primary data through CAPI (Computer Assisted Personal Interviewing, where the interviewer enters answers on a tablet) or web applications [2].
  • These tools have built-in validation checks. For example, the tablet flags an age of 13 for a head of household who is "married with 4 children" [2].
  • This cuts data-entry and transcription errors at the moment of collection [2].

  • Other NSO safeguards against non-sampling error [2]:

  • questionnaires designed with experts and field staff
  • multi-layer training of field officials on concepts, definitions, coverage and handling respondent behaviour, which helps against non-response
  • multi-level data scrutiny and follow-up visits.

  • The National Quality Assurance Framework (NQAF) of MoSPI sets the wider quality standard.

Prelims Hooks

  • Sampling error = parameter − estimate. In NCERT's Manipur example it is 600 − 550 = 50.
  • A census has zero sampling error but does have non-sampling error. Statements claiming a census is "error-free" are false.
  • Increasing sample size reduces sampling error only. It does not reduce non-sampling error, and it may even increase it.
  • NCERT's three forms of non-sampling error: sampling bias, non-response error, error in data acquisition.
  • Sampling bias = part of the population is unreachable (phone or online surveys, stale frames). Non-response = sampled units refuse or cannot be contacted. Do not confuse the two.
  • Recording 13 as 31 is a transcription error, a data-acquisition (non-sampling) error. It is not a sampling error.
  • MMRP uses 7-day / 30-day / 365-day recall periods by item group. HCES 2022-23 average MPCE was ₹3,773 rural and ₹6,459 urban (without imputation) [3].
  • The Post Enumeration Survey is run by the RGI. It measures census coverage and content error. India's net omission rate was 2.33% in 2001 [4].
  • CAPI reduces data-entry errors through built-in validation. It is used by NSO (MoSPI) [2].

Mains Points

  • Quality over quantity in surveys. A very large but badly designed survey can be less accurate than a small, careful one.
  • Money spent on frames, training and CAPI often buys more accuracy than simply adding households.
  • The census and PES show that complete coverage still leaves a 2.33% net undercount (2001) [4].

  • The NSS-PFCE gap and policy. Rich households refuse or under-report, and MMRP and imputation choices shift MPCE [3].

  • So poverty lines, inequality measures and CPI weights built on HCES can be biased.
  • This argues for reconciling surveys with national accounts and administrative data (tax, GST, UPI).

  • Census delay as a data-governance issue. The frames still date from Census 2011 while Census 2027 is awaited.

  • Stale frames bias PLFS, HCES and ASUSE (Annual Survey of Unincorporated Sector Enterprises) estimates, especially for urban and migrant populations.
  • A regular, on-time census is statistical infrastructure, not just a headcount (GS-II governance, GS-III planning).

  • Institutional safeguards: NSO's CAPI validation, multi-layer training and multi-level scrutiny [2], plus the NQAF, show how India can reduce non-sampling error. Publishing standard errors (for sampling error) and non-response rates (for non-sampling error) would further build trust in official data.

Sources

  1. 1Class 11, Ch 2 "Collection of Data"; Class 11, Ch 1 "Introduction (Statistics for Economics)"; Class 11, Ch 3 "Organisation of Data"; Class 11, Ch 4 "Presentation of Data"; Class 11, Ch 5 "Measures of Central Tendency"; Class 11, Ch 6 "Correlation"; Class 11, Ch 8 "Use of Statistical Tools" (primary)
  2. 2PIB — National Statistics Office (NSO) under MoSPI is committed to ensuring accurate and reliable data while minimizing non-sampling errorspib.gov.in · tier 1
  3. 3MoSPI/NSSO — Press Release: Results of Household Consumption Expenditure Survey 2022-23 (24 February 2024)mospi.gov.in · tier 1
  4. 4UN DESA Population Division — Census counts, undercounts and population estimates (Technical Paper 2020)un.org · tier 2
  5. 5UN Statistics Division — Post Enumeration Surveys (handbook)unstats.un.org · tier 2