Economic Data: Census, NSS, Surveys and Statistical Tools
In this note
- Why statistics in economics: from data to policy
- Where data come from: sources, questionnaires and survey modes
- Census or sample: populations, samples and sampling methods
- Sampling and non-sampling errors
- Organising and presenting data (compact)
- Summarising and relating data: averages, spread and correlation (compact)
- Reading economic data: time series, growth rates and high-frequency indicators
- The Census of India: 1881 to the 2027 round
- Consumption surveys: HCES, MPCE and reference periods
- Employment and enterprise surveys: PLFS, ASI, ASUSE and the Economic Census
- Household-welfare and farm surveys: AIDIS, SAS, NFHS, time use and crop cutting
- India's statistical system: institutions, reforms and credibility
- Exam angles
1. Why statistics in economics: from data to policy
What economics needs facts about
- Consumption, production and distribution are the three core areas.
- Economics also needs facts on special problems: poverty, disparity, unemployment, and the cost of disasters such as tsunamis, earthquakes and bird flu.
- Economic data are economic facts expressed in numbers. They are collected to understand a problem and the causes behind it.
The chain from data to policy
- Data: the facts are collected first.
- Economic analysis: the problem is explained by its causes. For example, poverty is explained by unemployment, low productivity and backward technology.
- Economic policy: a measure is chosen to solve the problem.
- Evaluation: statistics then checks whether the policy worked. Class 11, Introduction asks whether family planning has checked population growth.
- Without data there is no analysis, and without analysis there is no policy.
What "statistics" means
- Statistics has two meanings. It means the data themselves (plural). It also means the method: collection → organisation and presentation → analysis → interpretation.
- Quantitative data can be measured in numbers. NCERT's example: rice output rose from 39.58 million tonnes (1974-75) to 106.5 MT (2013-14). (NCERT: 106.5 MT; now about 150 MT in 2024-25, verify current.)
- Qualitative data record attributes that cannot be measured, such as gender. They are sometimes recorded in degrees, for example unskilled / skilled / highly skilled, or sick / healthy.
What statistics does
- Precision: "310 people died in the Kashmir earthquake" is a statistical fact. "Hundreds died" is not.
- Condensing: a mass of data is reduced to a few summary measures (mean, variance). You cannot remember thousands of incomes, but you can remember one average income.
- Finding and testing relationships: price and demand, income and consumption, government spending and the price level.
- Statistical prediction: statistical tools are applied to past or survey data to estimate future values for planning. NCERT's examples are deciding in 2017 how much to produce in 2020, and how much oil India should import in 2025.
The limit: statistical methods are no substitute for common sense
- A family of four knew the river's average depth. It was less than the family's average height, so they crossed, and the children drowned.
- The average was misused. It hid the spread. Section 6 comes back to this story.
The project cycle (Class 11, Use of Statistical Tools)
- A project report follows fixed steps: identify the problem → choose the target group → collect primary and/or secondary data → organise and present → analyse (averages, dispersion, correlation) → conclude, with predictions and suggestions → bibliography.
- The target group is the set of people the study focuses on. It follows from the objective: middle and high-income groups for a study on cars; all rural and urban consumers for soap; rural and urban people for safe drinking water.
- Worked example, the toothpaste project: a sample of 100 households (67% urban). Mean monthly family income was Rs 18,000 (SD Rs 9,000). Mean spending on toothpaste was Rs 104 per household per month (SD Rs 35.60). Pepsodent, Colgate and Close-up were the top brands, and television was the main media influence.
- This note follows the same cycle: collection (§2-4) → organisation (§5) → analysis (§6-7) → India's data system (§8-12).
- Cross-ref: Marshall's definition, scarcity and choice → economic-problem-systems.
2. Where data come from: sources, questionnaires and survey modes
Basic terms
- A variable is a quantity that takes different values in different situations. It is usually written X, Y or Z.
- An observation is each single value of a variable.
- NCERT example: foodgrain output was 108 MT (1970-71), 132 MT (1978-79), 176 MT (1990-91), 252 MT (2015-16) and 272 MT (2016-17). (NCERT: 272 MT; now a record above 350 MT in 2024-25, verify current.)
Primary vs secondary data
| Primary data | Secondary data | |
|---|---|---|
| Who collects | The researcher, first-hand, through an enquiry | Another agency has already collected and processed (scrutinised, tabulated) it |
| Source | Survey, interview | Government reports, newspapers, books, websites |
| Cost | High, slow | Saves time and cost |
| Example | Your survey of a film star's popularity | Someone else reusing your published report |
- Data are primary to the agency that first collects them and secondary to every later user. NSS unit-level data are primary for NSO and secondary for a researcher who downloads them.
Survey and instrument
- A survey is the method of gathering information from individuals, for example a firm testing a product or a party testing a candidate.
- A questionnaire is the list of questions. It can be self-administered by the respondent. An interview schedule is filled in by an enumerator, a trained person who collects data. NSS enumerators canvass schedules.
- The informant is the person or unit who gives the information.
Rules for good questions (NCERT's poor/good pairs)
- Keep it short and easy, with no ambiguous or difficult words.
- Move from general to specific: ask "Is supply regular?" before "Is the tariff hike justified?"
- Be precise: drop "in order to look presentable" from the clothing question.
- Be unambiguous: replace "a lot of money on books?" with ranges (below Rs 200, Rs 200-300 …).
- No double negatives: not "Don't you think smoking should be prohibited?"
- No leading question, i.e. one that hints at the answer, such as "How do you like the flavour of this high-quality tea?"
- Do not suggest alternatives: not "a job after college or be a housewife?"
Question types
- A closed-ended question (structured) offers answers to choose from. It can be two-way (yes/no) or multiple choice with an "Any other (specify)" option. It is easy to score and code, but hard to write and may restrict the true answer.
- An open-ended question allows individual answers ("Your view on globalisation?"). It is rich, but hard to interpret and score.
- A pilot survey (pre-testing) is a try-out on a small group. It checks the questions, the clarity of instructions, the enumerators' performance, and the cost and time of the full survey.
Survey modes compared
| Mode | Advantages | Disadvantages |
|---|---|---|
| Personal interview (face-to-face) | Highest response; all question types; best for open questions; clarifies doubts; reactions can be watched | Most expensive; slowest; interviewer can influence answers |
| Mailed questionnaire survey | Least expensive; reaches remote areas; no interviewer influence; keeps anonymity; best for sensitive questions | Low response; useless for illiterates; long response time; no clarification; reactions unseen |
| Telephone interview | Relatively cheap and quick; can clarify; good when respondents avoid face-to-face | Limited to those with phones; reactions unseen; some influence possible |
| Online survey / SMS | Fast and cheap at scale | Misses the unconnected (sampling bias, §4) |
- Beyond NCERT: NSO surveys now use CAPI (Computer Assisted Personal Interviewing) on tablets. Census 2027 uses mobile-app enumeration with a self-enumeration window (verify current).
- The choice of mode depends on the objective, the literacy of respondents and how easy they are to reach (Class 11, Collection of Data).
3. Census or sample: populations, samples and sampling methods
Population and census
- The population (universe) is every unit that has the characteristic under study. The purpose of the study defines it.
- A census (complete enumeration) collects data from every unit. India's decennial census is a house-to-house count of all households, run by the Registrar General of India.
Sample survey
- A sample is the part of the population from which information is actually taken.
- A sample survey studies a sample and uses the result for the whole population.
- A representative sample mirrors the population's make-up. It gives reasonably accurate results at lower cost and in less time.
- The sample estimate stands in for the population parameter, the true value in the population (for example, true average income).
- NCERT example: to study agricultural labourers in Churachandpur (Manipur), the population is all agricultural labourers in the district and the sample is 10% of them.
Why most surveys are samples
- They are cheaper and quicker.
- They allow more detail through intensive enquiry.
- They need a smaller team, which is easier to train and supervise.
Why a census is still needed
- Some tasks need every unit's data: electoral rolls and delimitation, small-area (village/ward) and caste counts.
- The census also provides the sampling frames that surveys themselves draw from.
Random sampling
- In random sampling, every unit has an equal chance of selection. Random does not mean haphazard.
- It first needs a sampling frame, the full list of units. NCERT's example is all 300 households in a locality, studied for the effect of a petrol-price rise.
- The lottery method: write all 300 names on slips, mix them and draw 30 one by one.
- A random number table (such as Tippett's, printed in Class 11, Use of Statistical Tools) does the same job with numbered units. Today, computer generators are used.
- NCERT exercise: choosing 3 of 10 students gives ¹⁰C₃ = 10!/(3!·7!) = 120 possible samples.
Non-random sampling
- In non-random sampling, units do not have equal chances. They are picked by the investigator's judgement, purpose, convenience or quota, so bias creeps in.
- NCERT example: picking 10 of 100 households that are nearby or known to you.
Exit polls
- An exit poll asks a random sample of voters leaving polling booths whom they voted for, to predict the result.
- Exit polls misfire through non-response, sampling bias and false answers (§4).
- Section 126A of the Representation of the People Act 1951 bars conducting or publishing exit polls from the start of polling in the first phase until half an hour after polling closes in the last phase.
Beyond NCERT: how NSS samples
- The NSS uses stratified multi-stage sampling. First-stage units are census villages (rural) and Urban Frame Survey (UFS) blocks (urban). Households are selected at the second stage.
- Systematic sampling (every k-th unit from a list) is another common method.
4. Sampling and non-sampling errors
Sampling error
- Sampling error = population parameter − sample estimate.
- NCERT example: five Manipur farmers earn 500, 550, 600, 650 and 700. The population mean is 3000/5 = 600. A sample of (500, 600) gives a mean of 550. The sampling error is 600 − 550 = 50.
- It shrinks as the sample grows. It exists only in samples, so a census has none.
Non-sampling error
- Non-sampling error is more serious. A bigger sample does not cure it, and even a census has it. It takes three forms.
- Sampling bias: the sampling plan cannot reach part of the target population.
- Phone or online surveys miss people without connections.
-
Stale frames miss new settlements and new urban areas.
-
Non-response error: sampled people cannot be contacted or refuse, so the sample stops being representative.
- Error in data acquisition: incorrect responses get recorded.
- Students' measuring tapes differ when they measure the same table.
- Orange prices vary by shop, market and quality, so only averages can be recorded.
- Informants forget (recall lapses).
- Transcription slips happen, such as writing 13 for 31.
| Sampling error | Non-sampling error | |
|---|---|---|
| Source | Studying a part, not the whole | Design, frame, response, recording |
| Larger sample | Reduces it | Does not reduce it (may even raise it) |
| In a census | Absent | Present |
Indian illustrations (carried into §9-12)
- The rich don't answer. Affluent households refuse or under-report. So NSS consumption falls short of national-accounts Private Final Consumption Expenditure (PFCE), and the NSS-NAS gap widens.
- Recall period changes the answer. The same household reports more spending with a 7-day recall than a 30-day one (URP vs MMRP, §9).
- Coverage checks. RGI's Post Enumeration Survey (PES) measures the census's own coverage and content errors.
- Aging frames. Survey frames based on Census 2011 have grown stale while Census 2027 is awaited.
- Technology helps. CAPI cuts transcription error through built-in validation checks.
5. Organising and presenting data (compact)
Classification
- Raw data are unorganised and bulky. NCERT's examples are the maths marks of 100 students and the monthly food spending of 50 households (Rs 1,007 to Rs 5,090). Drawing conclusions from them is tedious.
- Classification arranges data into groups by some criterion to bring order. NCERT compares it to how a kabadiwallah sorts junk. There are four kinds:
- Chronological classification (by time) gives a time series. India's population: 35.7 crore (1951) → 43.8 → 54.6 → 68.4 → 81.8 → 102.7 → 121.0 crore (2011).
- Spatial classification (by place). Wheat yield in 2013, kg/ha: India 3,154; France 7,254; Germany 7,998; China 5,055; Canada 3,594; Pakistan 2,787 (Agricultural Statistics at a Glance 2015).
- Qualitative classification is by attribute, a quality that cannot be measured (gender, religion, literacy). It works by presence or absence: first male/female, then married/unmarried within each.
- Quantitative classification is by measurable size (marks, income, height), grouped into classes.
Variables
- A continuous variable can take any value, including fractions: height, weight, time.
- A discrete variable moves in jumps: number of students. It may still take fractions, e.g. 1/8, 1/16 …, as long as it takes no value in between.
Frequency distribution
- Frequency is how often a value occurs.
- A frequency distribution shows the classes of a variable with their class frequency (the count in each class). NCERT's marks example: 0-10: 1, 10-20: 8, 20-30: 6, 30-40: 7, 40-50: 21, 50-60: 23, 60-70: 19, 70-80: 6, 80-90: 5, 90-100: 4; total 100.
- Class limits are the lower and upper ends of a class. The class interval (width) = UL − LL.
- The class mark = (UL + LL)/2. It represents every value in the class.
- The range is largest minus smallest value. It sets the number of classes, usually 6-15: number of classes = range ÷ class width.
- Unequal class intervals suit two cases: a huge range (income, from near zero to crores) or values clustered in a small part of the range. In NCERT's example, the 40-70 band is split into width-5 classes, and age-at-death tables use very short early age intervals.
Inclusive vs exclusive classes
- In the inclusive method, both limits belong to the class (0-10, 11-20 …).
- In the exclusive method, the upper limit goes to the next class (10-20, 20-30), so the classes are continuous.
- The adjustment in class interval restores continuity in inclusive classes. Take half the gap (900 − 899 = 1, so 0.5), subtract it from each lower limit and add it to each upper limit. So 800-899 becomes 799.5-899.5.
- NCERT error: Class 11, Organisation of Data says inclusive intervals are "used very often" for continuous variables. In fact continuity needs exclusive or adjusted classes.
- An open-ended class ("70 and over", "less than 10") is undesirable. It also blocks calculation of the mean.
Other points
- Loss of information: once data are grouped, every value is treated as the class mark. The values 20, 22, 25, 25, 25, 28 all become 25.
- Tally marking counts class frequencies in groups of five.
- A frequency array classifies a discrete variable value by value, e.g. household size 1, 2, 3 … with the number of households at each.
- Relative frequency = class frequency ÷ total. The 40-70 marks band holds 63%.
- A univariate distribution covers one variable.
- A bivariate frequency distribution covers two variables, with joint frequencies in each cell. NCERT's example is sales vs advertising spending for 20 firms, e.g. 3 firms with sales of Rs 135-145 lakh and advertising of Rs 64-66 thousand. This feeds correlation (§6).
Presentation
- Textual presentation puts data in sentences. It suits small data but must be read in full. NCERT's example is Census 2001: 102 crore people, 40 crore workers and 62 crore non-workers.
- Tabular presentation uses rows and columns. It suits further statistical work and can be one-, two- or three-way.
- Parts of a statistical table: table number, title, caption (column heading), stub (row heading; the left column is the stub column), body, unit of measurement, source, note.
- Example, a 3×3 table of Census 2011 literacy for ages 7+ (%):
| Rural | Urban | Total | |
|---|---|---|---|
| Male | 79 | 90 | 82 |
| Female | 59 | 80 | 65 |
| Total | 68 | 84 | 74 |
- Diagrammatic presentation is the quickest to read but less exact than a table.
Geometric diagrams
- Bar diagram: one-dimensional. The bars are equally spaced and equally wide, and only height matters. It suits discrete data, attributes and non-frequency data. NCERT's example is male literacy by state, 2011.
- Multiple bar diagram: compares sets side by side. NCERT's example is female literacy in 2001 vs 2011. Bihar (33.1 → 53.3), Jharkhand and UP rose most. Kerala was highest at 92.0 and Rajasthan lowest at 52.7 in 2011. India went from 53.7 to 65.5.
- Component bar diagram: shows the parts of a total. In a Bihar district, 41.4% of girls aged 6-14 were out of school (8.5% of boys).
- Pie diagram: 360°/100, so 1% = 3.6°. Working status in 2011 was main workers 29.8% (107°), marginal workers 9.9% (36°), non-workers 60.3% (217°).
- NCERT error: Table 4.8 prints the total as 102 crore, but the components (12 + 36 + 73) add to about 121 crore, the 2011 population.
Frequency diagrams
- Histogram: two-dimensional. Adjacent rectangles have area proportional to frequency.
- With unequal classes, height = frequency density = class frequency ÷ class width.
- It is drawn only for continuous data. It locates the mode graphically.
-
NCERT's example is daily wages of 85 earners.
-
Frequency polygon: joins the midpoints of the histogram tops and closes at the zero-frequency classes at each end. It is best for comparing distributions.
- Frequency curve: a smooth freehand curve through the polygon points.
- Ogive: plots cumulative frequency.
- The "less than" ogive plots against upper limits and never falls.
- The "more than" ogive plots against lower limits and never rises.
- The two intersect at the median.
Arithmetic line graph
- An arithmetic line graph (time-series graph) puts time on the X-axis. It shows trend and periodicity.
- DGCI&S data, 1993-94 to 2013-14 (Rs 100 crore): exports rose from 698 to 19,050 and imports from 731 to 27,154. Imports stayed higher all through, and the gap widened after 2001-02.
- Export shares in 2013-14 (total US$ 314.4 bn): Other Asia 29.4%, GCC 15.3%, USA 12.5%, Other EU 10.9%.
- (NCERT: 2013-14 data; now total goods-and-services exports were about US$ 825 bn in 2024-25, verify current.)
6. Summarising and relating data: averages, spread and correlation (compact)
Measures of central tendency reduce data to one typical value. NCERT's case is Baiju, who owns 1 acre in Balapur (Buxar, Bihar). Is he above the mean, the median or the mode of the 50 farmers' holdings?
Arithmetic mean
- Arithmetic mean: X̄ = ΣX/N. For grouped data, X̄ = Σfm/Σf, where m is the class mark.
- Example: six family incomes average Rs 1,547.
- The assumed-mean shortcut is X̄ = A + Σd/N. Step deviation divides d by a common factor c. In NCERT's example, weekly incomes of 10 families average Rs 1,116.
- Properties:
- Deviations from the mean sum to zero: Σ(X − X̄) = 0.
- It uses every value.
- An outlier (extreme value) drags it. NCERT's series 1, 2, 3000 shows this.
- It cannot be computed with open-ended classes.
Weighted arithmetic mean
- Weighted arithmetic mean = Σwx/Σw. In NCERT, budget shares are used as weights on mango and potato prices, which is the basis of price indices.
- Beyond NCERT: trade-weighted vs simple average tariffs, the weighted average lending rate, and the revenue-neutral GST rate.
- Index construction → inflation-price-indices.
Median
- The median is the positional middle value: size of the (N+1)/2-th item.
- For continuous data: Median = L + [(N/2 − c.f.)/f] × h. NCERT's example gives a median daily wage of Rs 35.83.
- It can be read from the intersection of the ogives.
- It is robust to outliers, so it is the right average for skewed income or MPCE.
Partition values
- Quartiles: Q1 = size of the (N+1)/4-th item, Q3 = size of the 3(N+1)/4-th item. In NCERT's marks example, Q1 = 13.5.
- A decile divides the data into 10 equal parts. Percentiles divide it into 100.
- Q2 = D5 = P50 = median.
- Scoring the 82nd percentile means 18% scored higher.
- HCES reports MPCE by fractile classes built this way (§9).
Mode
- The mode is the most frequent value. It suits qualitative data, such as the shoe size in most demand.
- A distribution can be unimodal, bimodal or multimodal, or have no mode (1, 1, 2, 2, 3, 3, 4, 4).
- It lies in the modal class: Mo = L + [D₁/(D₁ + D₂)] × h. This needs equal, exclusive classes. NCERT's example gives about Rs 27,273.
- It can be read from a histogram.
Relative position of mean, median and mode
- In a symmetric distribution, all three coincide.
- In a moderately skewed distribution, Mode ≈ 3 Median − 2 Mean (Karl Pearson).
- NCERT error: Class 11, Measures of Central Tendency says the median is "always" between the mean and the mode. That holds only for moderately skewed unimodal data.
- Income is right-skewed, so mean > median > mode. That is why per capita figures overstate the typical person.
Geometric and harmonic mean
- Geometric mean and harmonic mean: GM = ⁿ√(x₁…xₙ), for growth rates and ratios (CAGR, §7). HM = n/Σ(1/x), for rates such as speed.
Spread
- Measures of dispersion:
- Range.
- Mean deviation: the average absolute deviation from the mean or median.
- Variance: σ² = Σ(X − X̄)²/N.
-
Standard deviation: σ = √variance, in the same units as the data. The toothpaste income SD was Rs 9,000.
-
An average means little without dispersion, as the river story shows.
Correlation
- Correlation measures the direction and intensity of covariation between variables.
- Correlation vs causation: correlation shows that two variables move together. It never proves that one causes the other.
- Positive correlation: variables move together (income-consumption, temperature-ice-cream).
- Negative correlation: they move oppositely (apple price-demand, vegetable arrivals-price).
- Perfect correlation: all points lie on a line, r = ±1.
- A scatter diagram plots the pairs of values. It shows the relationship's form visually, including non-linear forms.
- A linear relationship is one a straight line can represent.
Karl Pearson's coefficient of correlation
- Covariance = Σ(X − X̄)(Y − Ȳ)/N. Its sign sets the sign of r.
- r = Cov(X,Y)/(σx·σy) = Σxy/(N·σx·σy).
- Properties:
- It is unit-free.
- −1 ≤ r ≤ +1.
-
It measures linear relations only. r = 0 means no linear relation, not independence (NCERT: X = −3…3, Y = X² gives r = 0).
-
Worked example: farmers' years of schooling vs yield per acre gives r = 42/(√112·√38) = 0.644. Price index vs money supply gives r = 0.98.
- Change of origin and scale: with U = (X − A)/B and V = (Y − C)/D, where B and D have the same sign, r_uv = r_xy. This is the basis of the step-deviation shortcut.
Spearman's rank correlation
- Spearman's rank correlation: rs = 1 − 6ΣD²/(n³ − n), where D is the difference in ranks.
- Use it for attributes (beauty, honesty), for unmeasurable variables (heights in a village without a measuring rod), and for data with outliers.
- It is generally no more than r for precisely measured data.
- Beauty-contest example: judges A-B = 0.3, A-C = 0.5, B-C = 0.9.
- Tied ranks get the mean rank plus the correction (m³ − m)/12.
Spurious correlation
- Spurious correlation can arise three ways:
- Coincidence: migratory birds and the local birth rate.
- A third variable: ice-cream sales and drownings are both driven by temperature.
-
Timing and confounding: more doctors are sent to epidemic villages, and deaths rise. The cases were terminal, the benefit comes later, and other shocks intervene.
-
Beyond NCERT: policy needs causal evidence. That comes from RCTs (Nobel 2019: Banerjee, Duflo, Kremer) and natural experiments (Nobel 2021: Card; Angrist and Imbens).
7. Reading economic data: time series, growth rates and high-frequency indicators
Time series and its parts
- A time series is data in chronological order, such as population by census year or trade by year.
- Trend is the long-run movement, seen on a line graph. Examples: rising exports and imports, 1993-2014.
- Seasonality is a regular pattern within a year: kharif/rabi arrivals, festive demand, monsoon vegetable prices.
- This is why CPI and IIP are quoted year-on-year (same month last year).
-
It is also why seasonally adjusted month-on-month figures exist.
-
Cyclicity is a pattern longer than a year, i.e. business cycles. Phases → growth-theories-business-cycles.
- Irregular movements are one-off shocks, such as COVID in 2020-21.
Growth arithmetic
- y-o-y vs m-o-m: m-o-m is noisy and seasonal; y-o-y compares like with like.
- Base effect: a low base last year inflates this year's y-o-y rate, and a high base deflates it. Growth after the 2020-21 slump looks high partly for this reason.
- Percentage vs percentage points: an unemployment rate falling from 6.1% to 3.2% is a fall of 2.9 percentage points, but a ~48% fall (2.9/6.1).
- Compound annual growth rate: CAGR = (end/start)^(1/n) − 1. It is a geometric mean.
- Decadal population growth of 17.7% in 2001-11 → 1.177^(1/10) − 1 ≈ 1.64% a year, matching NCERT.
- Nominal vs real → national-income-accounting / inflation-price-indices.
Purchasing Managers' Index (PMI)
- The Purchasing Managers' Index is a monthly, survey-based business indicator.
- In India: HSBC India Manufacturing, Services and Composite PMIs, compiled by S&P Global from panels of about 400 firms each.
- It is a diffusion index of month-on-month change. >50 = expansion, <50 = contraction, 50 = no change. It measures the direction of change, not the level of activity.
- Manufacturing PMI weights:
| Component | Weight |
|---|---|
| New orders | 30% |
| Output | 25% |
| Employment | 20% |
| Suppliers' delivery times (inverted) | 15% |
| Stocks of purchases | 10% |
- A flash estimate is released mid/late month and the final print early the next month.
Indicators by timing
- Leading indicators turn before the economy does. They are used to spot turning points: PMI new orders, stock prices, the slope of the yield curve, consumer confidence.
- Coincident indicators move with the economy: IIP, the eight core industries, GST and e-way bills, power demand.
- Lagging indicators turn after: unemployment, CPI inflation, bank NPAs.
RBI's forward-looking surveys (read by the MPC)
- OBICUS: order books, inventories and capacity utilisation.
- Industrial Outlook Survey.
- Consumer Confidence Survey.
- Inflation Expectations Survey of Households.
- Survey of Professional Forecasters.
8. The Census of India: 1881 to the 2027 round
History and law
- 1872: a non-synchronous count was completed under Lord Mayo.
- 1881: the first synchronous census, under W.C. Plowden. It has been decennial since. 1951 was the first after Independence.
- The decennial population census has run every ten years from 1881 to 2011. The 2021 round slipped to 2027.
- Legal basis: the Census Act 1948. Census is Entry 69 of the Union List.
- It is run by the Office of the Registrar General & Census Commissioner, India (MHA, set up 1949).
- NCERT lists "Census of India" and "RGI" as if they were separate. They are one office.
- The RGI also runs:
- SRS (Sample Registration System): IMR, TFR and MMR between censuses.
- CRS (Civil Registration System): under the RBD Act 1969, amended 2023.
- NPR (National Population Register).
What it collects
- Demographic data: size, density, sex ratio, literacy, migration, rural-urban split, birth and death rates, life expectancy, and workers (main/marginal).
- The houselisting phase adds housing and amenities.
Census 2011 (15th overall, 7th after Independence)
- Population 121.09 crore (NCERT's 2001 figure: 102.87 crore).
- Sex ratio 943.
- Literacy 74.04%.
- Density 382/km².
- Decadal growth 17.7% (17.64%), the lowest of any decade since Independence.
Census 2027
- Timeline:
- The 2021 round was postponed because of COVID.
- It was notified in June 2025.
-
Reference date: 00:00 hrs, 1 March 2027. For Ladakh and snow-bound parts of J&K, HP and Uttarakhand: 1 October 2026.
-
Two phases: houselisting in April-September 2026, then population enumeration in February 2027.
- Firsts:
- The first digital census, with mobile apps and self-enumeration.
- The first caste enumeration since 1931, after a CCPA decision in April 2025.
- SECC 2011 was a separate exercise, and its caste data remain unreleased.
-
(verify current: field progress, budget.)
-
Stakes:
- The delimitation freeze (Arts 82 and 170) lasts until the first census after 2026.
- Women's reservation (106th Amendment, Art 334A) starts after that census and the delimitation based on it.
- NFSA coverage and survey frames still run on 2011 figures.
Reading the population series (crore)
| 1901 | 1951 | 1961 | 1971 | 1981 | 1991 | 2001 | 2011 |
|---|---|---|---|---|---|---|---|
| 23.83 | 35.7 | 43.8 | 54.6 | 68.4 | 81.8 | 102.7 | 121.0 |
- The population rose by over 97 crore in 110 years.
- Average annual growth fell: 2.2% (1971-81) → 1.97% (1991-2001) → 1.64% (2001-11).
- Female literacy rose from 53.7% to 65.5% (2001-11).
- India became the world's most populous country in 2023 (UN estimate).
- Cross-ref: the 1881 census and the 1921 "great divide" → colonial-economy-1947.
9. Consumption surveys: HCES, MPCE and reference periods
History
- NSS consumer expenditure surveys began with the 1st round (1950-51).
- "Thick" quinquennial rounds ran from the 27th (1972-73) to the 68th (2011-12), which is NCERT's example (NCERT outdated).
- The 75th round (2017-18) was withheld in November 2019, citing "data quality", after a leak showed real MPCE falling.
HCES 2022-23 and 2023-24
- Periods: August 2022-July 2023 and August 2023-July 2024.
- Three questionnaires: food; consumables and services; durables. They were canvassed over three monthly visits using CAPI.
- Monthly per capita consumption expenditure (MPCE) = household monthly consumption ÷ household size.
- MPCE is reported with and without the imputed value of free items (PDS grain, laptops, bicycles and similar).
Results
| 2022-23 | 2023-24 | |
|---|---|---|
| Rural MPCE | Rs 3,773 | Rs 4,122 |
| Urban MPCE | Rs 6,459 | Rs 6,996 |
| Urban-rural gap | 71% (84% in 2011-12) | 70% |
| Gini: rural / urban | 0.266 / 0.314 | 0.237 / 0.284 |
- The food share is below 50%: about 47% rural and about 40% urban (2023-24).
- Beverages and processed food are the largest food item.
- Fractile classes run from the bottom 5% to the top 5% (§6 percentiles).
Reference periods (recall periods)
| Method | Recall used |
|---|---|
| Uniform reference period (URP) | 30 days for all items |
| Mixed reference period (MRP) | 365 days for clothing, bedding, footwear, durables, education, institutional medical care; 30 days for all else |
| Modified mixed reference period (MMRP) (from 2009-10) | 7 days for perishables (edible oil, egg-fish-meat, vegetables, fruits, spices, beverages, processed food, pan/tobacco/intoxicants); 365 days as in MRP; 30 days for the rest |
Why recall matters
- A shorter recall captures more spending, because people forget less.
- So MPCE: MMRP > MRP > URP, and measured poverty is lower under MMRP.
- This is non-sampling error (§4) shaping a headline number.
One survey, three uses
- Poverty lines: Lakdawala used URP, Tendulkar MRP, Rangarajan MMRP. Analysis → poverty-inequality.
- CPI weights: 2011-12 MMRP data gave the CPI (base 2012) weights. HCES 2023-24 feeds the new CPI (base 2024) (verify current). Construction → inflation-price-indices.
- Cross-check on PFCE: survey consumption is compared with national-accounts PFCE. This exposes the NSS-NAS gap, partly caused by non-response and under-reporting by the rich.
10. Employment and enterprise surveys: PLFS, ASI, ASUSE and the Economic Census
Periodic Labour Force Survey (PLFS)
- Launched by NSO in April 2017, designed on the Amitabh Kundu committee's recommendations.
- It replaced the quinquennial NSS employment-unemployment rounds (last in 2011-12) and the Labour Bureau's annual EUS.
- Two ways of measuring activity:
- Usual status (ps+ss) looks back 365 days: principal plus subsidiary activity.
-
Current weekly status (CWS) looks back 7 days.
-
Urban areas use a rotational panel, so the same households are revisited.
- The indicators are LFPR, WPR and UR. Definitions and analysis → employment-informal-sector.
- 2025 revamp (verify current):
- Monthly CWS estimates for rural + urban.
- Quarterly estimates for rural and urban (earlier urban only).
- About 2.7 lakh households a quarter.
-
Calendar-year annual reports.
-
2023-24 (July-June), usual status, 15+:
- Unemployment rate (UR) 3.2%.
- Labour force participation rate (LFPR) 60.1%.
- Worker population ratio (WPR) 58.2%.
-
Female LFPR 41.7%.
-
Flashpoints:
- The 2017-18 leak ("6.1%, a 45-year high"), followed by two NSC members resigning in January 2019.
- Comparisons with CMIE's CPHS.
- EPFO payroll data used as a proxy for formal jobs.
Enterprise surveys
- ASI (Annual Survey of Industries)
- Runs since 1960, now under the Collection of Statistics Act 2008.
- Covers factories under Sections 2(m)(i)/(ii) of the Factories Act 1948: 10+ workers with power, 20+ without.
- Also covers bidi and cigar units and some electricity undertakings.
-
Gives organised manufacturing's GVA, employment, wages and capital.
-
ASUSE (Annual Survey of Unincorporated Sector Enterprises)
- Covers unincorporated non-agricultural (informal) enterprises annually from 2021-22.
-
Quarterly bulletins from 2025 (verify current).
-
Economic Census
- A complete count of all establishments.
-
1st in 1977 … 7th in 2019-21, conducted with CSC e-Governance.
-
Labour Bureau
- Runs quarterly establishment surveys (QES under AQEES) for jobs in selected sectors.
Why this matters
- ASI captures the organised sector well. The informal sector needs ASUSE, which only recently became annual.
- This organised-unorganised measurement gap also affects GDP estimates (→ national-income-accounting).
11. Household-welfare and farm surveys: AIDIS, SAS, NFHS, time use and crop cutting
Household-welfare surveys
- AIDIS (All India Debt and Investment Survey)
- Decennial since 1961-62. RBI ran the rural credit surveys from 1951-52, and the NSS took over in 1992 (48th round).
- 77th round (January-December 2019): assets, liabilities, and institutional vs non-institutional credit.
- Incidence of indebtedness was 35.0% rural and 22.4% urban.
-
Credit analysis → financial-inclusion-rural-credit.
-
SAS (Situation Assessment Survey of agricultural households), 77th round
- Average monthly income per agricultural household: Rs 10,218 (agricultural year 2018-19).
- About half of households were indebted.
-
It feeds the doubling-farm-income debate.
-
NFHS (National Family Health Survey)
- Run by MoHFW with IIPS Mumbai as nodal agency. It is India's version of the DHS.
- Rounds: NFHS-1 (1992-93) … NFHS-5 (2019-21), with district-level estimates.
- NFHS-5: TFR 2.0; 1,020 women per 1,000 men; child stunting 35.5%; anaemia in women (15-49) 57%.
-
NFHS-6 fieldwork was in 2023-24 (verify current: results).
-
NSS thematic range: NCERT lists NSS estimates on literacy, school enrolment, morbidity, maternity, child care and PDS use. Example: the 60th round (January-June 2004) on morbidity and healthcare.
Time use
- Time Use Survey (NSO, 2019 and 2024).
- In 2024, women (15-59) spent 289 minutes a day on unpaid domestic and care work, against 88 for men.
- It is the basis for valuing unpaid work.
Crop estimation
- A crop cutting experiment (CCE) harvests crops from randomly selected plots to estimate average yield.
- CCEs are run by states under the General Crop Estimation Surveys (GCES), with technical guidance and supervision from NSO's Field Operations Division.
- Yield × area = production. This feeds the DES advance estimates.
- The same yields settle PMFBY insurance claims.
- Technology is coming in: the CCE-Agri app, YES-TECH (satellite and remote-sensing yield estimation), and the Digital Crop Survey (verify current).
Complete enumerations alongside the samples
- Agriculture Census (11th, 2021-22, MoA&FW).
- Livestock Census (21st, 2024-25, DAHD).
12. India's statistical system: institutions, reforms and credibility
Constitutional basis
- Union List Entry 94: statistics for Union matters.
- Concurrent List Entry 45: statistics for concurrent matters.
- Union List Entry 69: census.
- The system is decentralised. Ministries and states produce their own statistics.
Centre
- MoSPI was set up in 1999.
- NSO was formed in May 2019 by merging:
- CSO, set up in 1951, which produced national accounts, IIP and CPI.
-
NSSO, whose NSS began in 1950 on P.C. Mahalanobis's initiative.
-
(NCERT outdated: Class 11, Collection of Data still lists CSO and NSS separately.)
- NSO produces GDP, IIP, CPI (Rural/Urban/Combined), ASI, the NSS surveys and the Economic Census.
Oversight and law
- NSC (National Statistical Commission) was set up in 2005-06 on the Rangarajan Commission's 2001 advice. It is non-statutory.
- The Collection of Statistics Act 2008 is the legal backbone.
- ISI Kolkata was founded in 1931 and became an Institute of National Importance by the ISI Act 1959.
- Statistics Day is 29 June, Mahalanobis's birthday.
Line agencies
| Agency | Parent | Output |
|---|---|---|
| RGI | MHA | Census, SRS, CRS, NPR |
| DGCI&S (Kolkata, 1862) | Commerce | Foreign trade data (source of NCERT's trade table) |
| Labour Bureau (1946, Shimla/Chandigarh) | Labour & Employment | CPI-IW/AL/RL, establishment surveys |
| DES | MoA&FW | Crop estimates, Agricultural Statistics at a Glance |
| Office of the Economic Adviser | DPIIT | WPI |
| RBI | — | Handbook of Statistics, DBIE, forward-looking surveys |
| State DESs | States | State income, district data |
Secondary-data publications
- Sarvekshana, the NSS journal (NCERT calls it quarterly).
- Statistical Year Book; Women and Men in India.
- Economic Survey:
- Prepared by DEA, Ministry of Finance, under the Chief Economic Adviser.
- Tabled just before the Union Budget.
- The first was in 1950-51. It was delinked from the Budget in 1964.
- Its statistical appendix is a standard data source.
- Its policy role → economic-problem-systems.
Credibility flashpoints
- GDP series: the 2015 series (base 2011-12, MCA-21 corporate data) and the 2018 back-series row.
- Overestimation claim: Arvind Subramanian (2019) argued GDP growth was overstated.
- PLFS leak: the 2017-18 report was leaked, and NSC members resigned (January 2019).
- Junked CES 2017-18: this left no official poverty estimate for a decade.
- Census delay: frames went stale, welfare coverage was outdated, and delimitation was stalled.
- IMF grade: the IMF's 2025 Article IV graded India's national accounts "C" (verify current).
Reforms and responses
- Demands for a statutory NSC and advance release calendars.
- The Standing Committee on Statistics (2023; later restructured, verify current).
- The eSankhyiki portal (2024).
- New base years: GDP 2022-23, CPI 2024, IIP 2022-23, rolled out in 2026 (verify current).
- Monthly PLFS.
- Administrative and big data (GSTN, EPFO, UPI, satellite imagery), with privacy governed by the DPDP Act 2023.
The thread of this note
- Every number carries its design choices: frame, sample, recall period, response.
- Credibility needs those choices to be transparent and the institutions to be independent.
- The Economic Survey 2018-19 framed this as data as a public good.
Exam angles
Prelims — high-yield facts and traps
- Agency-product matching (the classic trap):
- RGI (MHA): Census, SRS, CRS, NPR.
- NSO (MoSPI): HCES, PLFS, AIDIS, SAS, TUS, ASUSE, ASI, IIP, CPI-Combined, GDP, Economic Census.
- Labour Bureau (MoLE): CPI-IW/AL/RL.
- OEA, DPIIT: WPI.
- DGCI&S (Kolkata, Commerce): trade statistics.
- MoHFW via IIPS Mumbai: NFHS.
- MoA&FW: Agriculture Census and crop estimates. DAHD: Livestock Census.
- DEA, Ministry of Finance: Economic Survey. It is NOT prepared by RBI or NITI Aayog.
- S&P Global for HSBC: PMI. It is NOT from RBI or MoSPI.
-
RBI: OBICUS, Consumer Confidence and Inflation Expectations surveys.
-
Chronology: 1872/1881 census → 1950 NSS → 1951 CSO → 1999 MoSPI → 2005-06 NSC → 2008 Collection of Statistics Act → 2017 PLFS → May 2019 NSO → 2022-24 HCES → 2025 monthly PLFS → Census reference date 1 March 2027 (1 October 2026 for snow-bound areas).
- Formulas:
- Class mark = (UL + LL)/2.
- X̄ = ΣX/N; weighted mean = Σwx/Σw.
- Median = L + [(N/2 − c.f.)/f]·h.
- Mode = L + [D₁/(D₁ + D₂)]·h.
- Mode ≈ 3 Median − 2 Mean.
- r = Cov(X,Y)/(σxσy).
- rs = 1 − 6ΣD²/(n³ − n).
- CAGR = (end/start)^(1/n) − 1.
- Pie: 1% = 3.6°.
-
¹⁰C₃ = 120.
-
Statement traps:
- "A larger sample reduces non-sampling error": FALSE. It reduces only sampling error.
- "A census is free of all error": FALSE. It has no sampling error but does have non-sampling error.
- "Random sampling means haphazard selection": FALSE. It means an equal chance for every unit.
- "Exit polls use non-random samples": FALSE as NCERT describes them. They are random samples of voters.
- "A histogram is one-dimensional": FALSE. It is 2-D and for continuous data only. A bar diagram is 1-D.
- "The median is read from a histogram": FALSE. The histogram gives the mode; the ogives' intersection gives the median.
- "The mean is unaffected by outliers": FALSE. The median is robust. The mode suits qualitative data.
- "r = 0 means the variables are independent": FALSE. It means only no linear relation.
- "r has the units of X × Y": FALSE. It is unit-free, between −1 and +1.
- "Rank correlation is more accurate than r for precisely measured data": FALSE. Spearman's is for ranks, attributes and data with outliers.
- "PMI 52 means output is at 52% of capacity": FALSE. It means m-o-m expansion.
- Seasonality is under a year; cyclicity is over a year.
-
"Census and RGI are separate bodies": FALSE. They are one office.
-
Survey facts to know:
- Recall periods: URP 30 days; MRP 365/30; MMRP 7/30/365.
- HCES MPCE rural/urban: Rs 3,773/6,459 (2022-23) and Rs 4,122/6,996 (2023-24).
- PLFS: CWS is 7 days, usual status 365 days; monthly estimates since 2025. 2023-24 UR 3.2%.
- AIDIS: decennial; 77th round in 2019 (indebtedness rural 35%, urban 22.4%).
- NFHS-5: TFR 2.0.
- Census 2011: 121.09 crore, sex ratio 943, literacy 74.04%, decadal growth 17.7%.
- ASI coverage thresholds: 10 workers with power, 20 without.
- Census is Union List Entry 69; statistics are Union List Entry 94 and Concurrent List Entry 45.
Mains — GS-III themes
- Official statistics as the basis of evidence-based policy (GS-III, GS-II governance). They drive welfare targeting (PDS and NFSA coverage frozen on 2011 figures), poverty estimation, and CPI and GDP weights. The census delay has real costs: stale frames, outdated coverage, and delimitation and women's reservation left waiting.
- Caste enumeration in Census 2027 (GS-I/GS-II). Data for social justice and sub-categorisation vs identity politics and reservation pressures. Also: SECC 2011's unreleased caste data, and state caste surveys (Bihar 2023).
- Independence and credibility of India's statistical system. A statutory NSC, release calendars, the GDP methodology controversy, the junked CES 2017-18, the PLFS leak, and the IMF "C" grade. The reform agenda: new base years, high-frequency data, administrative and big data, and privacy under the DPDP Act.
-
Methodology shapes findings: - Recall periods change measured poverty. - Non-response by the rich widens the NSS-NAS gap. - CMIE and PLFS give different unemployment figures. - Usual status and CWS give different pictures of work. - Old sampling frames bias surveys.
-
Reading data critically: - Averages hide distribution: mean vs median income, per capita income vs Gini. - Correlation is not causation in policy evaluation; RCTs and natural experiments give causal evidence. - Base effects and percentage points vs per cent can mislead. - PMI and other high-frequency indicators help spot turning points.
Current-affairs hooks
- Census 2027: houselisting (April-September 2026), population enumeration (February 2027), caste questions, self-enumeration and digital tools, and the delimitation debate over North-South seat shares.
- MoSPI releases:
- HCES fact sheets (February 2024, December 2024) and any next round.
- New series in 2026 (verify current): CPI (base 2024), GDP (base 2022-23), IIP (base 2022-23).
- PLFS monthly/quarterly bulletins and annual reports; ASI and ASUSE results.
-
Time Use Survey 2024; NFHS-6 results; SRS statistical reports.
-
Economic Survey: released in late January, before the Budget.
- Statistics days: Statistics Day on 29 June (the NSS completed 75 years in 2025); World Statistics Day on 20 October (every five years; 2025).
- Monthly and meeting-linked releases: HSBC PMI flash and final prints each month; RBI survey releases alongside MPC meetings.
- Recurring news triggers:
- Exit-poll misses after elections and the Section 126A embargo.
- UN World Population Prospects on India's rank.
- The IMF Article IV data-adequacy rating.
- World Bank poverty estimates built on HCES.
Detailed notes
- Why statistics in economics: from data to policy
- Where data come from: sources, questionnaires and survey modes
- Census or sample: populations, samples and sampling methods
- Sampling and non-sampling errors
- Organising and presenting data (compact)
- Summarising and relating data: averages, spread and correlation (compact)
- Reading economic data: time series, growth rates and high-frequency indicators
- The Census of India: 1881 to the 2027 round
- Consumption surveys: HCES, MPCE and reference periods
- Employment and enterprise surveys: PLFS, ASI, ASUSE and the Economic Census
- Household-welfare and farm surveys: AIDIS, SAS, NFHS, time use and crop cutting
- India's statistical system: institutions, reforms and credibility