Where data come from: sources, questionnaires and survey modes

Economic Data: Census, NSS, Surveys and Statistical Tools · section 2 of 12

In this note
  1. Detail
  2. Prelims Hooks
  3. Mains Points

Detail

1. Basic terms: variable and observation

  • Variable: a quantity whose value changes from one situation to another. It is usually written as X, Y or Z.
  • Examples: price, income, foodgrain output, age.

  • Observation: one single value taken by a variable.

  • NCERT example. Variable = India's foodgrain output. Each year's figure is one observation.
Year Foodgrain output
1970-71 108 MT
1978-79 132 MT
1990-91 176 MT
2015-16 252 MT
2016-17 272 MT (NCERT figure)
2024-25 357.73 MT, a record (final estimates) [7]
  • Update: Output in 2024-25 was 357.73 MT, against 251.54 MT in 2015-16. That is about 106 MT more [7]. (NCERT: 272 MT in 2016-17.)
  • Worked example (growth of a variable):
  • Change from 2015-16 to 2024-25 = 357.73 − 251.54 = 106.19 MT
  • % change = (106.19 ÷ 251.54) × 100 ≈ 42.2%

  • Record crops in 2024-25: rice at 1,501.84 lakh tonnes and wheat at 1,179.45 lakh tonnes [7].

2. Primary data vs secondary data

  • Primary data: data you collect yourself, first-hand, through your own enquiry (a survey or an interview).
  • Secondary data: data that some other agency has already collected and processed. "Processed" means the data were checked (scrutinised) and put into tables (tabulated).
Primary data Secondary data
Who collects The researcher, first-hand, through an enquiry Another agency has already collected and processed it
Source Survey, interview Government reports, newspapers, books, websites
Cost High and slow Saves time and cost
Example Your own survey of a film star's popularity Someone else reusing your published report
  • Key idea: primary and secondary depend on who is using the data.
  • Data are primary for the agency that collects them first.
  • The same data are secondary for everyone who uses them later.
  • Example: NSS unit-level data are primary for the NSO (National Statistical Office), which collects them. For a researcher who downloads them, they are secondary.

  • Care with secondary data: always check who collected the data, what the definitions were, what the reference year was and how the data were collected. A different definition can make two data sets impossible to compare.

3. Survey, instrument and the people involved

  • Survey: a method of collecting information from individuals.
  • Examples: a firm testing a new product, or a political party testing how popular a candidate is.

  • Questionnaire: the list of questions. The respondent can fill it in alone (self-administered).

  • Interview schedule: the same kind of question list, but an enumerator fills it in.
  • Enumerator: a trained person who asks the questions and records the answers. NSS enumerators "canvass" (ask and fill in) schedules in households.

  • Informant (respondent): the person or unit that gives the information, for example a household, a firm or a farmer.

  • Legal backing in India:
  • Census Act, 1948 (Act No. 37 of 1948, amended in 1994). This is the legal basis of the Census [9].
  • Collection of Statistics Act, 2008. Parliament passed it on 7 January 2009 and it came into force on 11 June 2010. It replaced (repealed) the Collection of Statistics Act, 1953 [8].
  • Confidentiality: Under the 2008 Act, information collected cannot be disclosed or used as evidence in any court proceeding. Data about a person cannot be released unless the details that identify the person are removed [8].
  • Why this matters: informants trust that their answers stay secret → they answer honestly → the data are more accurate.

4. Rules for writing good questions (NCERT's poor/good pairs)

Rule Poor question Better approach
Short, simple words Long question with difficult or ambiguous words Short, plain words
General → specific Asking "Is the tariff hike justified?" first Ask "Is supply regular?" first, then the tariff question
Precise Asking about clothing spending "in order to look presentable" Drop the extra phrase
Unambiguous "Do you spend a lot of money on books?" Give ranges: below Rs 200, Rs 200-300 …
No double negatives "Don't you think smoking should be prohibited?" "Do you think smoking should be prohibited?"
No leading question "How do you like the flavour of this high-quality tea?" "How do you like the flavour of this tea?"
No suggested alternatives "Would you like a job after college or be a housewife?" "What would you like to do after college?"
  • Leading question: a question that hints at the "right" answer. The word "high-quality" pushes the respondent to answer positively. This produces bias, meaning answers that are wrong in one direction.
  • Ranges fix ambiguity: "a lot" means Rs 100 to one person and Rs 1,000 to another. Ranges give everyone the same scale.

5. Types of questions

  • Closed-ended (structured) question: the respondent chooses from answers that are already given.
  • Two-way question: only two options (yes/no).
  • Multiple-choice question: several options, usually ending with "Any other (specify)" so that no true answer is left out.
  • Pros: easy to score, code and tabulate. Good for large surveys.
  • Cons: hard to write well. The given options may not include the respondent's true answer.

  • Open-ended question: the respondent answers in their own words, e.g. "What is your view on globalisation?"

  • Pro: rich, detailed answers.
  • Con: hard to interpret, compare and score.

  • Pilot survey (pre-testing): a small try-out of the questionnaire before the main survey. It checks:

  • whether the questions are clear
  • whether the instructions are clear
  • how well the enumerators perform
  • how much the full survey will cost and how long it will take

6. Survey modes compared

Mode Advantages Disadvantages
Personal interview (face-to-face) Highest response; suits all question types; best for open questions; doubts can be cleared; reactions can be seen Most expensive; slowest; the interviewer can influence answers
Mailed questionnaire survey Cheapest; reaches remote areas; no interviewer influence; keeps anonymity; best for sensitive questions Low response; useless for people who cannot read; long response time; no chance to clarify; reactions unseen
Telephone interview Fairly cheap and quick; doubts can be cleared; useful when people avoid face-to-face meetings Only reaches people with phones; reactions unseen; some interviewer influence possible
Online survey / SMS Fast and cheap at large scale Misses people who are not connected (sampling bias, §4)
  • Response rate: the share of people contacted who actually reply.
  • Formula: Response rate = (Completed responses ÷ Questionnaires sent) × 100
  • Worked example (illustrative): 1,000 questionnaires are mailed and 250 come back → response rate = 25%. If 900 of 1,000 face-to-face interviews are completed → 90%. This is why mailed surveys are called "low response".

  • Sampling bias in online surveys: people without internet are left out → the sample is richer, younger and more urban than the population → the results are misleading for the whole country.

  • How to choose a mode (NCERT): it depends on
  • the objective of the survey
  • the literacy of respondents
  • how easy they are to reach

7. Modern practice: from paper to digital

  • CAPI (Computer Assisted Personal Interviewing): the enumerator records answers directly on a tablet instead of paper.
  • MoSPI developed CAPI software for data collection in the PLFS (Periodic Labour Force Survey) [4].
  • PLFS was launched in April 2017. It is the main official source of data on the labour force, employment and unemployment [4].
  • NSS surveys now run on CAPI through the e-SIGMA platform. It has built-in validation checks (automatic error checks), real-time data submission, screens in many languages and AI chatbot support [5].
  • Why it helps: errors are caught at the doorstep → there is less cleaning later → results come out faster.

  • PLFS reform (January 2025): the methodology was revised so that PLFS now gives monthly national estimates of key labour indicators [5][6].

  • Census 2027: India's first digital census
  • It is the 16th census in the series and the 8th since independence [2].
  • Self-enumeration (SE) is offered for the first time: households can fill in and submit their own details online at se.census.gov.in [2][3].
  • An optional 15-day self-enumeration window comes before the house-to-house visit [2].
  • Enumerators use a secure mobile app that works offline. Only people registered on the CMMS portal can use it. Data go straight from the field to the server, with no paper forms [2].
  • The app and the portal are available in 16 languages, including Hindi and English [2].
  • Phase I, Houselisting and Housing Census (HLO): April to September 2026, with a 30-day window in each State/UT [2].

  • Link to NCERT modes: self-enumeration works like a modern "mailed/online questionnaire" (cheap, private, needs literacy). The enumerator's app visit is a digital "personal interview". Census 2027 combines both, so people who cannot or do not self-enumerate are still covered.

Prelims Hooks

  • The same data set is primary for the agency that collects it (e.g. NSO for NSS) and secondary for a researcher who reuses it.
  • Informant = the person who gives the information. Enumerator = the trained person who collects it. An interview schedule is filled in by the enumerator, not the informant.
  • "How do you like the flavour of this high-quality tea?" is a leading question. It is not a double negative.
  • A pilot survey checks question clarity, enumerator performance, cost and time. It does not produce final estimates.
  • Mailed questionnaire: cheapest and best for sensitive questions, but lowest response and useless for people who cannot read. Personal interview: highest response and most expensive.
  • Collection of Statistics Act, 2008: in force from 11 June 2010; repealed the 1953 Act; bars disclosure of data that identify a person [8].
  • PLFS was launched in April 2017. It has given monthly national estimates since January 2025 [4][5].
  • Census 2027: 16th census, 8th since independence; first digital census; first self-enumeration facility; 16 languages [2].
  • Foodgrain output reached a record 357.73 MT in 2024-25 (NCERT: 272 MT in 2016-17) [7].

Mains Points

  • Choosing a survey mode is a trade-off between cost, coverage and quality. Face-to-face CAPI costs more but reaches people who cannot read and people without internet. Online and self-enumeration modes are cheap but can leave out the digitally excluded. This is why Census 2027 keeps the enumerator visit after the self-enumeration window [2].
  • Data quality starts with how questions are designed. Leading, vague or double-negative questions create bias. When bias enters labour or consumption data, it can distort policy targets such as poverty lines, welfare targeting and employment schemes. Pilot surveys and in-built CAPI validation checks [5] are the protection against this.
  • Confidentiality supports trust, and trust supports accuracy. The Collection of Statistics Act, 2008 and the Census Act, 1948 give legal protection to informants [8][9]. Digital collection increases the need for data security, so the statistical law may need to be read together with data-protection rules.
  • Faster data helps policy. Monthly PLFS estimates (from January 2025) and field-to-server census data [2][5] shorten the time between collecting data and making policy. This is part of the wider reform of India's statistical system (GS-III: growth and employment; GS-II: governance).

Sources

  1. 1Class 11, Ch 2 "Collection of Data"; Class 11, Ch 1 "Introduction (Statistics for Economics)"; Class 11, Ch 3 "Organisation of Data"; Class 11, Ch 4 "Presentation of Data"; Class 11, Ch 5 "Measures of Central Tendency"; Class 11, Ch 6 "Correlation"; Class 11, Ch 8 "Use of Statistical Tools" (primary)
  2. 2Census 2027: India's First Digital Enumeration Exercise (PIB)pib.gov.in · tier 1
  3. 3For the First Time, Census 2027 to Enable Digital Data Collection and Self-Enumeration (PIB)pib.gov.in · tier 1
  4. 4PLFS Instruction Manual Vol. I (MoSPI)mospi.gov.in · tier 1
  5. 5Counting What Counts: Strengthening India's National Accounts and Core Statistics (PIB, Jan 2026)static.pib.gov.in · tier 1
  6. 6MoSPI Annual Report 2025-26mospi.gov.in · tier 1
  7. 7Record foodgrain output breaks all previous highs (PIB)pib.gov.in · tier 1
  8. 8The Collection of Statistics Act, 2008 (MoSPI)mospi.gov.in · tier 1
  9. 9The Census Act, 1948 (Act No. 37 of 1948), as amended in 1994 (India Code) — )_doc.pdfindiacode.nic.in · tier 1