Where data come from: sources, questionnaires and survey modes
Economic Data: Census, NSS, Surveys and Statistical Tools · section 2 of 12
In this note
Detail
1. Basic terms: variable and observation
- Variable: a quantity whose value changes from one situation to another. It is usually written as X, Y or Z.
-
Examples: price, income, foodgrain output, age.
-
Observation: one single value taken by a variable.
- NCERT example. Variable = India's foodgrain output. Each year's figure is one observation.
| Year | Foodgrain output |
|---|---|
| 1970-71 | 108 MT |
| 1978-79 | 132 MT |
| 1990-91 | 176 MT |
| 2015-16 | 252 MT |
| 2016-17 | 272 MT (NCERT figure) |
| 2024-25 | 357.73 MT, a record (final estimates) [7] |
- Update: Output in 2024-25 was 357.73 MT, against 251.54 MT in 2015-16. That is about 106 MT more [7]. (NCERT: 272 MT in 2016-17.)
- Worked example (growth of a variable):
- Change from 2015-16 to 2024-25 = 357.73 − 251.54 = 106.19 MT
-
% change = (106.19 ÷ 251.54) × 100 ≈ 42.2%
-
Record crops in 2024-25: rice at 1,501.84 lakh tonnes and wheat at 1,179.45 lakh tonnes [7].
2. Primary data vs secondary data
- Primary data: data you collect yourself, first-hand, through your own enquiry (a survey or an interview).
- Secondary data: data that some other agency has already collected and processed. "Processed" means the data were checked (scrutinised) and put into tables (tabulated).
| Primary data | Secondary data | |
|---|---|---|
| Who collects | The researcher, first-hand, through an enquiry | Another agency has already collected and processed it |
| Source | Survey, interview | Government reports, newspapers, books, websites |
| Cost | High and slow | Saves time and cost |
| Example | Your own survey of a film star's popularity | Someone else reusing your published report |
- Key idea: primary and secondary depend on who is using the data.
- Data are primary for the agency that collects them first.
- The same data are secondary for everyone who uses them later.
-
Example: NSS unit-level data are primary for the NSO (National Statistical Office), which collects them. For a researcher who downloads them, they are secondary.
-
Care with secondary data: always check who collected the data, what the definitions were, what the reference year was and how the data were collected. A different definition can make two data sets impossible to compare.
3. Survey, instrument and the people involved
- Survey: a method of collecting information from individuals.
-
Examples: a firm testing a new product, or a political party testing how popular a candidate is.
-
Questionnaire: the list of questions. The respondent can fill it in alone (self-administered).
- Interview schedule: the same kind of question list, but an enumerator fills it in.
-
Enumerator: a trained person who asks the questions and records the answers. NSS enumerators "canvass" (ask and fill in) schedules in households.
-
Informant (respondent): the person or unit that gives the information, for example a household, a firm or a farmer.
- Legal backing in India:
- Census Act, 1948 (Act No. 37 of 1948, amended in 1994). This is the legal basis of the Census [9].
- Collection of Statistics Act, 2008. Parliament passed it on 7 January 2009 and it came into force on 11 June 2010. It replaced (repealed) the Collection of Statistics Act, 1953 [8].
- Confidentiality: Under the 2008 Act, information collected cannot be disclosed or used as evidence in any court proceeding. Data about a person cannot be released unless the details that identify the person are removed [8].
- Why this matters: informants trust that their answers stay secret → they answer honestly → the data are more accurate.
4. Rules for writing good questions (NCERT's poor/good pairs)
| Rule | Poor question | Better approach |
|---|---|---|
| Short, simple words | Long question with difficult or ambiguous words | Short, plain words |
| General → specific | Asking "Is the tariff hike justified?" first | Ask "Is supply regular?" first, then the tariff question |
| Precise | Asking about clothing spending "in order to look presentable" | Drop the extra phrase |
| Unambiguous | "Do you spend a lot of money on books?" | Give ranges: below Rs 200, Rs 200-300 … |
| No double negatives | "Don't you think smoking should be prohibited?" | "Do you think smoking should be prohibited?" |
| No leading question | "How do you like the flavour of this high-quality tea?" | "How do you like the flavour of this tea?" |
| No suggested alternatives | "Would you like a job after college or be a housewife?" | "What would you like to do after college?" |
- Leading question: a question that hints at the "right" answer. The word "high-quality" pushes the respondent to answer positively. This produces bias, meaning answers that are wrong in one direction.
- Ranges fix ambiguity: "a lot" means Rs 100 to one person and Rs 1,000 to another. Ranges give everyone the same scale.
5. Types of questions
- Closed-ended (structured) question: the respondent chooses from answers that are already given.
- Two-way question: only two options (yes/no).
- Multiple-choice question: several options, usually ending with "Any other (specify)" so that no true answer is left out.
- Pros: easy to score, code and tabulate. Good for large surveys.
-
Cons: hard to write well. The given options may not include the respondent's true answer.
-
Open-ended question: the respondent answers in their own words, e.g. "What is your view on globalisation?"
- Pro: rich, detailed answers.
-
Con: hard to interpret, compare and score.
-
Pilot survey (pre-testing): a small try-out of the questionnaire before the main survey. It checks:
- whether the questions are clear
- whether the instructions are clear
- how well the enumerators perform
- how much the full survey will cost and how long it will take
6. Survey modes compared
| Mode | Advantages | Disadvantages |
|---|---|---|
| Personal interview (face-to-face) | Highest response; suits all question types; best for open questions; doubts can be cleared; reactions can be seen | Most expensive; slowest; the interviewer can influence answers |
| Mailed questionnaire survey | Cheapest; reaches remote areas; no interviewer influence; keeps anonymity; best for sensitive questions | Low response; useless for people who cannot read; long response time; no chance to clarify; reactions unseen |
| Telephone interview | Fairly cheap and quick; doubts can be cleared; useful when people avoid face-to-face meetings | Only reaches people with phones; reactions unseen; some interviewer influence possible |
| Online survey / SMS | Fast and cheap at large scale | Misses people who are not connected (sampling bias, §4) |
- Response rate: the share of people contacted who actually reply.
- Formula: Response rate = (Completed responses ÷ Questionnaires sent) × 100
-
Worked example (illustrative): 1,000 questionnaires are mailed and 250 come back → response rate = 25%. If 900 of 1,000 face-to-face interviews are completed → 90%. This is why mailed surveys are called "low response".
-
Sampling bias in online surveys: people without internet are left out → the sample is richer, younger and more urban than the population → the results are misleading for the whole country.
- How to choose a mode (NCERT): it depends on
- the objective of the survey
- the literacy of respondents
- how easy they are to reach
7. Modern practice: from paper to digital
- CAPI (Computer Assisted Personal Interviewing): the enumerator records answers directly on a tablet instead of paper.
- MoSPI developed CAPI software for data collection in the PLFS (Periodic Labour Force Survey) [4].
- PLFS was launched in April 2017. It is the main official source of data on the labour force, employment and unemployment [4].
- NSS surveys now run on CAPI through the e-SIGMA platform. It has built-in validation checks (automatic error checks), real-time data submission, screens in many languages and AI chatbot support [5].
-
Why it helps: errors are caught at the doorstep → there is less cleaning later → results come out faster.
-
PLFS reform (January 2025): the methodology was revised so that PLFS now gives monthly national estimates of key labour indicators [5][6].
- Census 2027: India's first digital census
- It is the 16th census in the series and the 8th since independence [2].
- Self-enumeration (SE) is offered for the first time: households can fill in and submit their own details online at se.census.gov.in [2][3].
- An optional 15-day self-enumeration window comes before the house-to-house visit [2].
- Enumerators use a secure mobile app that works offline. Only people registered on the CMMS portal can use it. Data go straight from the field to the server, with no paper forms [2].
- The app and the portal are available in 16 languages, including Hindi and English [2].
-
Phase I, Houselisting and Housing Census (HLO): April to September 2026, with a 30-day window in each State/UT [2].
-
Link to NCERT modes: self-enumeration works like a modern "mailed/online questionnaire" (cheap, private, needs literacy). The enumerator's app visit is a digital "personal interview". Census 2027 combines both, so people who cannot or do not self-enumerate are still covered.
Prelims Hooks
- The same data set is primary for the agency that collects it (e.g. NSO for NSS) and secondary for a researcher who reuses it.
- Informant = the person who gives the information. Enumerator = the trained person who collects it. An interview schedule is filled in by the enumerator, not the informant.
- "How do you like the flavour of this high-quality tea?" is a leading question. It is not a double negative.
- A pilot survey checks question clarity, enumerator performance, cost and time. It does not produce final estimates.
- Mailed questionnaire: cheapest and best for sensitive questions, but lowest response and useless for people who cannot read. Personal interview: highest response and most expensive.
- Collection of Statistics Act, 2008: in force from 11 June 2010; repealed the 1953 Act; bars disclosure of data that identify a person [8].
- PLFS was launched in April 2017. It has given monthly national estimates since January 2025 [4][5].
- Census 2027: 16th census, 8th since independence; first digital census; first self-enumeration facility; 16 languages [2].
- Foodgrain output reached a record 357.73 MT in 2024-25 (NCERT: 272 MT in 2016-17) [7].
Mains Points
- Choosing a survey mode is a trade-off between cost, coverage and quality. Face-to-face CAPI costs more but reaches people who cannot read and people without internet. Online and self-enumeration modes are cheap but can leave out the digitally excluded. This is why Census 2027 keeps the enumerator visit after the self-enumeration window [2].
- Data quality starts with how questions are designed. Leading, vague or double-negative questions create bias. When bias enters labour or consumption data, it can distort policy targets such as poverty lines, welfare targeting and employment schemes. Pilot surveys and in-built CAPI validation checks [5] are the protection against this.
- Confidentiality supports trust, and trust supports accuracy. The Collection of Statistics Act, 2008 and the Census Act, 1948 give legal protection to informants [8][9]. Digital collection increases the need for data security, so the statistical law may need to be read together with data-protection rules.
- Faster data helps policy. Monthly PLFS estimates (from January 2025) and field-to-server census data [2][5] shorten the time between collecting data and making policy. This is part of the wider reform of India's statistical system (GS-III: growth and employment; GS-II: governance).
Sources
- 1Class 11, Ch 2 "Collection of Data"; Class 11, Ch 1 "Introduction (Statistics for Economics)"; Class 11, Ch 3 "Organisation of Data"; Class 11, Ch 4 "Presentation of Data"; Class 11, Ch 5 "Measures of Central Tendency"; Class 11, Ch 6 "Correlation"; Class 11, Ch 8 "Use of Statistical Tools" (primary)
- 2Census 2027: India's First Digital Enumeration Exercise (PIB)pib.gov.in · tier 1
- 3For the First Time, Census 2027 to Enable Digital Data Collection and Self-Enumeration (PIB)pib.gov.in · tier 1
- 4PLFS Instruction Manual Vol. I (MoSPI)mospi.gov.in · tier 1
- 5Counting What Counts: Strengthening India's National Accounts and Core Statistics (PIB, Jan 2026)static.pib.gov.in · tier 1
- 6MoSPI Annual Report 2025-26mospi.gov.in · tier 1
- 7Record foodgrain output breaks all previous highs (PIB)pib.gov.in · tier 1
- 8The Collection of Statistics Act, 2008 (MoSPI)mospi.gov.in · tier 1
- 9The Census Act, 1948 (Act No. 37 of 1948), as amended in 1994 (India Code) — )_doc.pdfindiacode.nic.in · tier 1