Sampling error
Topic: Economic Data: Census, NSS, Surveys and Statistical Tools · NCERT: Class 11, Ch 2 "Collection of Data"; Class 11, Ch 8 "Use of Statistical Tools"
Meaning
Sampling error is the gap between the true value for the whole population (the parameter) and the value we calculate from a sample (the estimate). It happens only because we studied a part of the population instead of all of it.
Sampling error = population parameter − sample estimate
It matters because almost all of India's big economic data, such as NSS, PLFS and HCES, comes from sample surveys. Every number they publish carries some sampling error. To read official data properly, an aspirant must know how big this error can be and how it can be made smaller.
Explanation
Basic terms
- Population (universe): the whole group we want to study. Example: all farmers in Manipur.
- Parameter: the true value for the whole population, such as its true mean (average) income.
- Sample: a small part of the population, chosen to stand for the whole.
- Estimate (statistic): the value we calculate from the sample.
- Error: any gap between the estimate and the parameter. There are two kinds:
- Sampling error comes from studying only a part.
- Non-sampling error comes from everything else: the survey plan, the list used to pick the sample, the people who answer, and the way answers are recorded.
How sampling error works: NCERT's Manipur example
- Incomes of 5 farmers: 500, 550, 600, 650, 700.
- Population mean = 3000 ÷ 5 = 600. This is the parameter.
- Sample of 2 (500, 600): mean = 1100 ÷ 2 = 550.
-
Sampling error = 600 − 550 = 50.
-
Sample of 4 (500, 550, 600, 650): mean = 2300 ÷ 4 = 575.
-
Sampling error = 600 − 575 = 25.
-
All 5 farmers (this is a census): mean = 600.
-
Sampling error = 600 − 600 = 0.
-
Lesson: as the sample grows, sampling error gets smaller. In a census it becomes zero.
Key features
- It is a matter of chance, not a mistake. No sample matches the population exactly. Nobody has done anything wrong.
- It exists only in sample surveys. A census (complete enumeration, which covers every unit) has no sampling error.
- It is random, not one-sided. One sample may give an estimate that is too high and another may give one that is too low.
- It can be measured with statistics. Surveys report it as a standard error (a number that shows how far a sample estimate is likely to be from the true value).
What makes it rise or fall
- Sample size:
- bigger sample → estimate closer to the parameter → smaller sampling error
-
census → sampling error is zero.
-
Sample design:
- Stratified random sampling means the population is split into groups (strata), such as rural and urban, or small and large farmers, and a sample is drawn from each group.
-
Every group is then represented, so the estimate is more accurate even when the sample is the same size.
-
What does NOT help: better questionnaires or better training of interviewers do not change sampling error. Those steps cut non-sampling error.
In India
- Who runs the sample surveys: the National Statistics Office (NSO) under MoSPI conducts NSS rounds, the PLFS (Periodic Labour Force Survey) and the HCES (Household Consumption Expenditure Survey). All of them are sample surveys, so their estimates carry sampling error.
- HCES 2022-23 (a large sample):
- It covered 2,61,746 households: 1,55,014 rural and 1,06,732 urban [3].
- Average MPCE (Monthly Per Capita Consumption Expenditure) without imputation, meaning without adding a money value for free items, was ₹3,773 rural and ₹6,459 urban [3].
-
A sample this large keeps sampling error low. It does not protect the numbers from non-sampling error.
-
The Census is the zero-sampling-error benchmark:
- The Census, run by the Registrar General of India (RGI), counts every person, so it has no sampling error.
-
It still has non-sampling error. The Post Enumeration Survey (PES) found a net omission rate that rose from 0.68% (1961) to 2.33% (2001) [4].
-
The frame problem: NSS and PLFS draw their samples from Census 2011 lists of villages and urban blocks. Census 2021 was postponed and Census 2027 is awaited. This stale frame causes sampling bias, which is a non-sampling error. A bigger sample cannot fix it.
- Transparency: publishing standard errors alongside survey estimates tells users how large the sampling error is. This builds trust in official data.
Don't confuse with
- Non-sampling error: it comes from the survey plan, the sampling frame, the people answering or the recording, not from picking a part. A larger sample does not reduce it and may even increase it. It is present even in a census.
- Sampling bias: part of the population is never in the sampling frame (the list or map used to draw the sample). Phone surveys that miss poor households are an example. It is one-sided and is a non-sampling error, even though its name contains "sampling".
- Error in data acquisition: a wrong answer gets recorded, for example writing 13 for 31 (a transcription slip) or a respondent forgetting past spending (recall lapse). This is a non-sampling error, not a sampling error.
- Non-response error: people in the sample refuse to answer or cannot be contacted. Example: rich households skip the survey. This is a non-sampling error.
Prelims Hooks
- Sampling error = population parameter − sample estimate. In NCERT's Manipur example it is 600 − 550 = 50. With a sample of 4 it falls to 25, and in a census it is 0.
- A census has zero sampling error but still has non-sampling error. A statement that a census is "free of all errors" is false. PES showed a 2.33% net omission rate in 2001 [4].
- Increasing the sample size reduces only sampling error. It does not reduce non-sampling error and can even increase it.
- Sampling error is random and statistically measurable. It is reduced by a larger sample or a better design, such as stratified random sampling.
- Trap: recording 13 as 31, recall lapses, refusals and phone-only surveys are all non-sampling errors, not sampling errors.
- NCERT calls non-sampling error more serious than sampling error.
Mains Points
- Size versus quality in surveys. A larger sample cuts sampling error, but beyond a point extra households add cost and more chances for refusals and recording slips.
- Money spent on fresh frames, interviewer training and CAPI (Computer Assisted Personal Interviewing, where answers are entered on a tablet with built-in checks) often improves accuracy more than a bigger sample [2].
-
A very large but badly designed survey can be less accurate than a small, careful one.
-
Honest reporting of error builds trust in official data. Surveys such as PLFS and HCES guide poverty lines, CPI weights and welfare targeting.
- Publishing standard errors (for sampling error) and non-response rates (for non-sampling error) lets users judge how reliable each number is.
-
This links to MoSPI's National Quality Assurance Framework (NQAF) and to data governance (GS-II and GS-III).
-
A timely census makes sample surveys work. A census has no sampling error, and it also supplies the frame for every sample survey.
- The Census 2011 frame is stale while Census 2027 is awaited, so new towns and migrant colonies are under-covered.
- A larger sample cannot correct this bias. Only a fresh frame can. That is why the census is basic statistical infrastructure and not just a headcount.
Related concepts
Read more
Sources
- 1Class 11, Ch 2 "Collection of Data"; Class 11, Ch 8 "Use of Statistical Tools" (primary)
- 2PIB — National Statistics Office (NSO) under MoSPI is committed to ensuring accurate and reliable data while minimizing non-sampling errorspib.gov.in · tier 1
- 3MoSPI/NSSO — Press Release: Results of Household Consumption Expenditure Survey 2022-23 (24 February 2024)mospi.gov.in · tier 1
- 4UN DESA Population Division — Census counts, undercounts and population estimates (Technical Paper 2020)un.org · tier 2