Summarising and relating data: averages, spread and correlation (compact)
Economic Data: Census, NSS, Surveys and Statistical Tools · section 6 of 12
In this note
Detail
1. Why summarise data: measures of central tendency
- A measure of central tendency is one number that stands for a whole set of data. It gives the "typical" value.
- NCERT case: Baiju owns 1 acre of land in Balapur village (Buxar, Bihar). The village has 50 farmers. Is Baiju above the typical farmer?
- The answer depends on which average you use: the mean, the median or the mode.
-
Each average can give a different answer for the same data.
-
The three main averages are the mean, the median and the mode. The geometric mean and harmonic mean are used in special cases.
2. Arithmetic mean
- Arithmetic mean (the sum of all values divided by the number of values):
- Individual series: X̄ = ΣX / N
-
Grouped data: X̄ = Σfm / Σf. Here f is the frequency (how many items fall in a class), and m is the class mark (the mid-point of the class, (lower limit + upper limit)/2).
-
NCERT example: the incomes of six families have a mean of Rs 1,547.
- Assumed-mean (short-cut) method: X̄ = A + Σd / N
- A is any value you pick as the "assumed mean".
- d = X − A (the deviation of each value from A).
-
Worked example: take the values 40, 50, 55, 78, 58 and pick A = 55. Then d = −15, −5, 0, +23, +3, so Σd = 6. X̄ = 55 + 6/5 = 56.2. The direct method gives the same answer: 281/5 = 56.2.
-
Step-deviation method: divide each d by a common factor c. Then X̄ = A + (Σd′/N) × c, where d′ = d/c. This keeps the numbers small.
-
NCERT example: the weekly incomes of 10 families have a mean of Rs 1,116.
-
Properties:
- Deviations from the mean always sum to zero: Σ(X − X̄) = 0.
- It uses every value in the data.
- An outlier (an extreme value far from the rest) pulls the mean towards itself.
- NCERT's series is 1, 2, 3000. The mean is 1001, which describes none of the three values. The median is 2.
- It cannot be calculated for open-ended classes (for example "above 500"), because such a class has no mid-point.
3. Weighted arithmetic mean
- Weighted arithmetic mean = Σwx / Σw. Each value x is multiplied by its importance, the weight w.
- NCERT use: a family's budget shares are used as weights on the prices of mangoes and potatoes. The same idea is the basis of price indices such as the CPI and WPI.
- Worked example: potatoes cost Rs 20/kg (weight 3) and mangoes cost Rs 50/kg (weight 1).
- Weighted mean = (20×3 + 50×1)/(3+1) = 110/4 = Rs 27.5.
-
The simple mean is Rs 35. It overstates the cost of living because the family buys far more potatoes than mangoes.
-
Beyond NCERT, the same logic in policy:
- Simple vs trade-weighted average tariff. A simple average treats every tariff line equally. A trade-weighted average weights each line by how much is actually imported, so it is usually different.
- Weighted average lending rate (WALR). Banks' loan interest rates are weighted by the amount lent at each rate.
-
Revenue-neutral GST rate. This is the single rate that would raise the same revenue as the old taxes. It is a weighted average over the whole tax base.
-
How indices are built → see the inflation-price-indices note.
4. Median
- The median is the middle value after the data are arranged in order. It is an average by position.
- Median = the size of the (N+1)/2-th item.
- Example: in 5, 7, 9, 12, 20 (N = 5), the 3rd item is the median, which is 9.
-
Example: in 5, 7, 9, 12 (N = 4), you need the 2.5th item, so the median is (7+9)/2 = 8.
-
Continuous (grouped) data: Median = L + [(N/2 − c.f.)/f] × h
- L = lower limit of the median class. c.f. = cumulative frequency of the class before the median class. f = frequency of the median class. h = class width.
- Worked example: take the classes 0–10 (f 5), 10–20 (f 8) and 20–30 (f 7), so N = 20 and N/2 = 10. The median class is 10–20, and the c.f. before it is 5. Median = 10 + [(10 − 5)/8] × 10 = 16.25.
-
NCERT example: the median daily wage is Rs 35.83.
-
Graphically: the median is where the "less than" ogive and the "more than" ogive cross. (An ogive is a cumulative frequency curve.)
- Robust to outliers: a few very rich households do not move it. So it is the right average for skewed data such as income or MPCE (monthly per capita consumption expenditure).
5. Partition values: quartiles, deciles, percentiles
- Quartiles cut ordered data into 4 equal parts.
- Q1 = size of the (N+1)/4-th item. Q3 = size of the 3(N+1)/4-th item.
- Example: in 3, 5, 7, 9, 11, 13, 15 (N = 7), Q1 is the 2nd item = 5 and Q3 is the 6th item = 13.
-
NCERT marks example: Q1 = 13.5.
-
A decile cuts the data into 10 equal parts (D1…D9). A percentile cuts it into 100 equal parts (P1…P99).
- Q2 = D5 = P50 = median. This is a common MCQ.
- Scoring the 82nd percentile means 82% scored at or below you and 18% scored higher. It does not mean 82% marks.
- In official data: the Household Consumption Expenditure Survey (HCES) reports average MPCE for each fractile class (for example the bottom 5%, 5–10%, and so on up to the top 5%). These classes are built from percentiles.
- HCES 2023-24: the rise in average MPCE over 2022-23 was largest for the bottom 5–10% of the population, in both rural and urban areas [2][3].
- The gap in per capita calorie intake between the bottom 5% and the top 5% fractile classes narrowed significantly in 2023-24, in both sectors [2][3].
6. Mode
- The mode is the value that occurs most often.
- It is the only average that suits qualitative data (data that describe a quality, not a number). Example: the shoe size that shoppers demand most.
- Types:
- Unimodal: one mode.
- Bimodal or multimodal: two or more modes.
-
No mode: every value repeats equally. Example: 1, 1, 2, 2, 3, 3, 4, 4.
-
Grouped data: first find the modal class (the class with the highest frequency). Then:
- Mo = L + [D₁/(D₁ + D₂)] × h
- D₁ = f(modal) − f(previous class). D₂ = f(modal) − f(next class).
- The classes must be equal in width and exclusive (the upper limit of one class is not counted in it but in the next).
- Worked example: the modal class is 20–30 with f = 12. The class before has f = 8 and the class after has f = 6. So D₁ = 4 and D₂ = 6. Mo = 20 + (4/10) × 10 = 24.
-
NCERT example: the mode is about Rs 27,273.
-
Graphically: the mode is read from a histogram. Draw the crossing lines from the tallest bar to its neighbours.
7. Relative position of mean, median and mode
- Symmetric distribution: mean = median = mode.
- Moderately skewed distribution (Karl Pearson): Mode ≈ 3 Median − 2 Mean.
- Example: mean = 50 and median = 45, so mode ≈ 135 − 100 = 35.
-
Here mean > median > mode, so the data are right-skewed.
-
NCERT error: the Class 11 chapter Measures of Central Tendency says the median "always" lies between the mean and the mode.
- That is true only for moderately skewed, unimodal data.
-
It can fail for strongly skewed or multimodal data.
-
Income is right-skewed. A few people earn very high incomes, which gives the distribution a long right tail. So mean > median > mode.
-
This is why per capita (mean) income or consumption overstates what the typical person has.
-
Official illustration:
- Average MPCE in 2023-24 was Rs 4,122 in rural India and Rs 6,996 in urban India (up from Rs 1,430 and Rs 2,630 in 2011-12) [2][3].
- So urban MPCE is about 70% above rural MPCE (6,996 ÷ 4,122 ≈ 1.70).
- Because consumption is right-skewed, these means are above the median household's spending.
8. Geometric mean and harmonic mean
- Geometric mean (GM) = ⁿ√(x₁ × x₂ × … × xₙ). Use it for growth rates and ratios, for example the CAGR (compound annual growth rate).
-
Example: growth is 10% in year 1 and 20% in year 2. GM = √(1.10 × 1.20) = √1.32 ≈ 1.149, so the true average growth is about 14.9%, not 15%.
-
Harmonic mean (HM) = n / Σ(1/x). Use it for rates, such as speed over equal distances.
-
Example: you drive at 60 km/h going and 40 km/h returning. HM = 2/(1/60 + 1/40) = 48 km/h, not 50 km/h.
-
For positive data that are not all equal: AM > GM > HM.
9. Spread: measures of dispersion
- Dispersion is how widely values are scattered around the average.
- Why it matters (NCERT's river story): a man hears that a river's average depth is low and tries to walk across. He drowns in the deep middle. An average means little unless you also know the spread.
- Measures:
- Range = largest value − smallest value. It is simple, but it depends only on the two extreme values.
- Mean deviation = the average of the absolute deviations (signs ignored), taken from the mean or the median.
- Variance σ² = Σ(X − X̄)² / N, the average of the squared deviations.
-
Standard deviation (SD) σ = √variance. It is in the same units as the data.
- NCERT toothpaste case: the SD of consumer income was Rs 9,000.
-
Worked example: data 2, 4, 4, 4, 5, 5, 7, 9 (N = 8, mean = 5).
- Range = 9 − 2 = 7.
- Absolute deviations: 3, 1, 1, 1, 0, 0, 2, 4, which sum to 12. Mean deviation = 12/8 = 1.5.
-
Squared deviations: 9, 1, 1, 1, 0, 0, 4, 16, which sum to 32. Variance = 32/8 = 4, so SD = 2.
-
Beyond NCERT: the Gini coefficient is a relative measure of spread used for inequality. It runs from 0 (everyone has the same) to 1 (one person has everything).
- Consumption Gini, rural: 0.266 (2022-23) → 0.237 (2023-24) [3].
- Consumption Gini, urban: 0.314 (2022-23) → 0.284 (2023-24) [3].
- It fell in almost all major states, in both sectors [3].
10. Correlation: basic ideas
- Correlation measures the direction and strength of how two variables move together.
- Correlation vs causation: correlation shows that two variables move together. It never proves that one causes the other.
- Types:
- Positive correlation: both move in the same direction. Examples: income and consumption; temperature and ice-cream sales.
- Negative correlation: they move in opposite directions. Examples: the price of apples and the demand for apples; vegetable arrivals in the market and vegetable prices.
-
Perfect correlation: all points lie exactly on a straight line, so r = +1 or −1.
-
Scatter diagram: a graph that plots each (X, Y) pair as a dot.
- It shows the form of the relationship visually, including non-linear (curved) forms.
-
It gives no number.
-
Linear relationship: a relationship that a straight line can represent.
11. Karl Pearson's coefficient of correlation (r)
- Covariance = Σ(X − X̄)(Y − Ȳ) / N. It measures how X and Y vary together, and its sign sets the sign of r.
- Formula: r = Cov(X,Y) / (σx · σy) = Σxy / (N · σx · σy), where x = X − X̄ and y = Y − Ȳ.
- Mini example: X = 1, 2, 3 and Y = 2, 4, 6.
- The deviations are x = −1, 0, 1 and y = −2, 0, 2, so Σxy = 4 and Cov = 4/3.
- σx = √(2/3) and σy = √(8/3), so σx·σy = 4/3.
-
r = (4/3)/(4/3) = +1 (perfect positive correlation).
-
Properties:
- It is unit-free: a pure number, not in rupees or kg.
- −1 ≤ r ≤ +1.
-
It measures linear relations only. r = 0 means no linear relation. It does not mean the variables are independent.
- NCERT example: X = −3 … 3 and Y = X². Y depends fully on X, yet r = 0, because the curve is U-shaped.
-
NCERT worked examples:
- Farmers' years of schooling vs yield per acre: r = 42/(√112 × √38) = 42/65.24 = 0.644 (moderate positive correlation).
-
Price index vs money supply: r = 0.98 (very high positive correlation).
-
Change of origin and scale: let U = (X − A)/B and V = (Y − C)/D, where B and D have the same sign.
- Then r_uv = r_xy. Shifting the origin or rescaling the data does not change r.
- If B and D have opposite signs, only the sign of r flips.
- This is why the step-deviation shortcut works for r.
12. Spearman's rank correlation (rs)
- Formula: rs = 1 − 6ΣD² / (n³ − n), where D is the difference between the two ranks of each item.
- Worked example: n = 5 and D = 1, −1, 0, 2, −2, so ΣD² = 10. rs = 1 − 60/120 = 0.5.
- When to use it:
- Attributes that cannot be measured, only ranked (beauty, honesty).
- Variables that cannot be measured in practice. Example: heights in a village without a measuring rod, where people can still be lined up by height.
-
Data with outliers, because ranks reduce the effect of extreme values.
-
For precisely measured data, rs is generally no more than r, because ranking throws away some information.
- NCERT beauty-contest example: judges A-B 0.3, A-C 0.5, B-C 0.9. So judges B and C have the most similar taste.
- Tied ranks:
- Give each tied item the mean of the ranks they share. Example: two items tied for 2nd and 3rd each get 2.5.
- Add the correction (m³ − m)/12 to ΣD² for each tie, where m is the number of tied items. For m = 2, the correction is (8 − 2)/12 = 0.5.
13. Spurious correlation and causal evidence
- Spurious correlation: a correlation that is real in the numbers but has no real cause-and-effect link.
- It can arise three ways:
- Coincidence. Example: the arrival of migratory birds and the local birth rate.
- A third variable. Example: ice-cream sales and drownings.
- Hot weather → more ice-cream is sold.
- Hot weather → more people swim → more people drown.
- Temperature drives both, so ice-cream does not cause drowning.
-
Timing and confounding (a hidden factor mixes up the result). Example: more doctors are sent to villages hit by an epidemic, and deaths rise.
- The worst, terminal cases were already there.
- The doctors' benefit shows only later.
- Other shocks were also at work.
- So "doctors cause deaths" is a false reading.
-
Beyond NCERT: how policy gets causal evidence.
- RCTs (randomised controlled trials): people are randomly split into a group that gets a programme and a group that does not, then outcomes are compared.
- 2019 Nobel (Economics): Abhijit Banerjee, Esther Duflo and Michael Kremer, "for an experimental approach to alleviating global poverty".
- They used field experiments in education, health, credit and technology adoption, over more than two decades [4].
- Natural experiments: real-life events or rules happen to split people "as if" at random.
- 2021 Nobel: half went to David Card for empirical contributions to labour economics. The other half went to Joshua Angrist and Guido Imbens for methodological contributions to the analysis of causal relationships [5].
- Angrist and Imbens developed the LATE (Local Average Treatment Effect) framework for drawing valid causal conclusions from natural experiments [5].
Prelims Hooks
- Q2 = D5 = P50 = median. Scoring the 82nd percentile means 18% scored higher. It does not mean 82% marks.
- Σ(X − X̄) = 0: deviations from the arithmetic mean always sum to zero.
- The mean cannot be calculated for open-ended classes. The median and mode can.
- The mode is the only average for qualitative data. The grouped mode formula needs equal, exclusive classes.
- Pearson's empirical relation: Mode ≈ 3 Median − 2 Mean, valid only for moderately skewed data. In right-skewed income data, mean > median > mode.
- GM suits growth rates (CAGR). HM suits speeds and rates. For unequal positive values, AM > GM > HM.
- −1 ≤ r ≤ +1. r is unit-free and unaffected by change of origin and scale.
- r = 0 means no linear relation, not independence. Trap: Y = X² on symmetric X gives r = 0.
- Spearman: rs = 1 − 6ΣD²/(n³ − n). Use it for ranked attributes such as beauty or honesty.
- HCES 2023-24 consumption Gini: rural 0.237 and urban 0.284, both lower than in 2022-23 (0.266 and 0.314) [3]. Nobel 2019: RCTs for poverty (Banerjee–Duflo–Kremer) [4].
Mains Points
- Mean vs median in welfare measurement:
- Per capita income and MPCE are means, so a small rich group pulls them up.
- Policy should also track the median and fractile-wise data. HCES 2023-24 fractiles showed the fastest MPCE growth at the bottom 5–10% [2][3].
- It should also track dispersion measures such as the Gini (rural 0.237 and urban 0.284 in 2023-24) [3].
-
Otherwise "average growth" can hide who actually gains.
-
Weighting drives official numbers:
- The CPI and WPI, trade-weighted tariffs, the WALR and the revenue-neutral GST rate are all weighted means.
-
Old weights (an outdated consumption basket) bias inflation and policy signals. This is why HCES-based weight revisions matter for the CPI.
-
Correlation is not causation in policy evaluation:
- Spurious links (third variables, timing) can make a good scheme look bad, as in the epidemic-doctors case, or a bad one look good.
-
Evidence-based policy needs RCTs and natural experiments [4][5]. NITI Aayog-style evaluation of DBT or nutrition schemes should use comparison groups, not before-after correlations.
-
Averages without spread mislead (the river story):
- Regional averages hide backward districts.
- This supports district-level indicators, as in the Aspirational Districts Programme, and state-wise Gini tracking.
Sources
- 1Class 11, Ch 2 "Collection of Data"; Class 11, Ch 1 "Introduction (Statistics for Economics)"; Class 11, Ch 3 "Organisation of Data"; Class 11, Ch 4 "Presentation of Data"; Class 11, Ch 5 "Measures of Central Tendency"; Class 11, Ch 6 "Correlation"; Class 11, Ch 8 "Use of Statistical Tools" (primary)
- 2Household Consumption Expenditure Survey: 2023-24 — Press Note, MoSPImospi.gov.in · tier 1
- 3Household Consumption Expenditure Survey: 2023-24 — PIBpib.gov.in · tier 1
- 4Michael Kremer | Biography & Facts (also Abhijit Banerjee, Esther Duflo entries) — Britannicabritannica.com · tier 3
- 5Joshua Angrist | Biography, Nobel Prize (also David Card, Guido Imbens entries) — Britannicabritannica.com · tier 3