Organising and presenting data (compact)
Economic Data: Census, NSS, Surveys and Statistical Tools · section 5 of 12
In this note
Detail
1. Why organise data at all
- Raw data means data collected but not yet arranged. It is unorganised and bulky.
- NCERT gives two examples:
- the maths marks of 100 students
-
the monthly food spending of 50 households, from Rs 1,007 to Rs 5,090
-
Reading 100 loose numbers does not show any pattern, so drawing conclusions is slow and tiring.
- Classification arranges data into groups using one common criterion (a rule for sorting).
- NCERT's comparison is a kabadiwallah, who sorts junk into paper, plastic, glass and metal so that each pile can be priced and sold.
-
Classification does the same for numbers. Like items go together, so we can compare and analyse them.
-
Official statistics go through the same step. MoSPI publishes survey results as ready tables in PDF, CSV and XLSX formats, not as loose numbers [6].
- MoSPI also runs a DataViz platform. It turns survey data (for example from PLFS) into charts [7].
2. Four kinds of classification
| Kind | Basis | Output | NCERT example |
|---|---|---|---|
| Chronological | Time | Time series | India's population in crore |
| Spatial | Place | Cross-country or cross-state comparison | Wheat yield, 2013 |
| Qualitative | Attribute (a quality that cannot be measured) | Groups by presence or absence | Gender, religion, literacy |
| Quantitative | Measurable size | Classes such as 0-10, 10-20 | Marks, income, height |
- Chronological (time series): India's population in crore:
- 35.7 (1951) → 43.8 (1961) → 54.6 (1971) → 68.4 (1981) → 81.8 (1991) → 102.7 (2001) → 121.0 (2011)
-
Over 60 years the population became about 3.4 times larger.
-
Spatial: wheat yield in 2013, kg per hectare (Agricultural Statistics at a Glance 2015):
- Germany 7,998 · France 7,254 · China 5,055 · Canada 3,594 · India 3,154 · Pakistan 2,787
-
India's yield was less than half of Germany's.
-
Qualitative: works in layers by presence or absence of an attribute.
- First split: male or female.
-
Then, inside each group: married or unmarried.
-
Quantitative: the variable is measured and grouped into classes (ranges of values).
3. Variables
- A variable is a quantity that changes from one unit to another, such as marks from student to student.
- Continuous variable: can take any value within a range, including fractions. Examples: height (152.4 cm), weight, time.
- Discrete variable: moves in jumps. Example: the number of students (30 or 31, never 30.5).
- Trap: a discrete variable can still take fractions (1/2, 1/4, 1/8, 1/16 …). What makes it discrete is that it takes no value in between those fixed points.
4. Frequency distribution
- Frequency: how many times a value (or a value inside a class) occurs.
- Frequency distribution: a table showing each class of a variable and its class frequency (the count of observations in that class).
- NCERT's marks of 100 students, worked through below:
| Class (marks) | Frequency (f) | Class mark (m) | f × m | Cumulative f ("less than" UL) |
|---|---|---|---|---|
| 0-10 | 1 | 5 | 5 | 1 |
| 10-20 | 8 | 15 | 120 | 9 |
| 20-30 | 6 | 25 | 150 | 15 |
| 30-40 | 7 | 35 | 245 | 22 |
| 40-50 | 21 | 45 | 945 | 43 |
| 50-60 | 23 | 55 | 1,265 | 66 |
| 60-70 | 19 | 65 | 1,235 | 85 |
| 70-80 | 6 | 75 | 450 | 91 |
| 80-90 | 5 | 85 | 425 | 96 |
| 90-100 | 4 | 95 | 380 | 100 |
| Total | 100 | 5,220 |
Key terms and formulas
- Class limits: the two ends of a class. For 40-50, the lower limit (LL) = 40 and the upper limit (UL) = 50.
- Class interval (class width) = UL − LL. For 40-50: 50 − 40 = 10.
- Class mark (mid-value) = (UL + LL) / 2. For 40-50: (50 + 40)/2 = 45.
-
The class mark stands for every value in the class in later calculations.
-
Range = largest value − smallest value.
- Number of classes = Range ÷ Class width. Usually 6 to 15 classes are used.
- Example: range 0-100 = 100, width 10 → 100 ÷ 10 = 10 classes.
-
Too few classes hide the pattern. Too many make the table as messy as the raw data.
-
Relative frequency = class frequency ÷ total frequency.
-
Example: the 40-70 band = (21 + 23 + 19)/100 = 63/100 = 63% of students.
-
Mean from grouped data = Σfm ÷ Σf = 5,220 ÷ 100 = 52.2 marks.
- Median (read graphically from the ogives, see §9):
- N/2 = 50. The cumulative frequency passes 50 in the 50-60 class, so that is the median class.
- Formula: Median = L + [(N/2 − cf)/f] × h = 50 + [(50 − 43)/23] × 10 ≈ 53.0.
-
Here L = lower limit of the median class, cf = cumulative frequency of the class before it, f = its frequency, h = its width.
-
Modal class (the class with the highest frequency) = 50-60, with f = 23.
Unequal class intervals
-
Equal widths are the default. Unequal widths suit two cases: 1. A huge range. Income runs from almost zero to crores, so equal classes would leave most classes empty. 2. Values packed into a small part of the range. In NCERT's marks example, 63% of students fall in 40-70, so that band is split into narrower width-5 classes (40-45, 45-50 …) to show detail.
-
Age-at-death tables use very short intervals for the early ages (0-1, 1-4 years), because infant and child deaths are frequent and matter for policy.
5. Inclusive and exclusive classes
| Exclusive method | Inclusive method | |
|---|---|---|
| Form | 10-20, 20-30 | 0-10, 11-20, 21-30 |
| Upper limit | Goes to the next class (20 is counted in 20-30) | Belongs to the same class |
| Continuity | Continuous, with no gaps | Gaps between classes (10 → 11) |
| Best for | Continuous variables | Discrete variables |
- Adjustment in class interval removes the gaps in inclusive classes: 1. Find the gap = next class's LL − this class's UL. Example: 900 − 899 = 1. 2. Take half the gap: 1/2 = 0.5. 3. Subtract 0.5 from every lower limit and add 0.5 to every upper limit. 4. So 800-899 → 799.5-899.5 and 900-999 → 899.5-999.5. The classes now join up.
-
The class width stays the same (100) and the class mark stays the same (849.5).
-
NCERT error: Class 11, Organisation of Data, says inclusive intervals are "used very often" for continuous variables. That is wrong. A continuous variable needs exclusive or adjusted classes, because a value such as 10.5 would otherwise have no class.
- Open-ended class: a class with no lower or no upper limit, such as "less than 10" or "70 and over".
- It is undesirable because it has no class mark, so the mean cannot be calculated.
- Official income or age tables sometimes still use one (for example "60+ years").
6. Other points on organisation
- Loss of information: once data are grouped, each value is treated as the class mark.
- Example: 20, 22, 25, 25, 25 and 28 in class 20-30 all become 25.
-
This is the price paid for a compact summary. The grouped mean (52.2 above) can differ a little from the true mean of the raw marks.
-
Tally marking: counting class frequencies with strokes. The fifth stroke crosses the first four (𝍸), so counting goes in blocks of five.
- Frequency array: the version of a frequency distribution used for a discrete variable. Each value gets its own row instead of a class.
-
Example: household size 1, 2, 3, 4 … with the number of households of each size.
-
Univariate distribution: covers one variable (for example, marks only).
- Bivariate frequency distribution: covers two variables together. Each cell holds a joint frequency (units that fall in class i of X and class j of Y at the same time).
- NCERT example: sales vs advertising spending for 20 firms. For instance, 3 firms have sales of Rs 135-145 lakh and advertising of Rs 64-66 thousand.
- This table is the starting point for correlation (§6 of the parent note), because it shows whether the two variables move together.
7. Three ways to present data
| Mode | What it is | Strength | Weakness |
|---|---|---|---|
| Textual (descriptive) | Data written in sentences | Good for small data and for stressing a point | Must be read in full; hard to compare |
| Tabular | Rows and columns | Exact; ready for further statistical work | Slower to read than a picture |
| Diagrammatic | Charts and graphs | Quickest to read and remember | Less exact than a table |
- Textual example (NCERT, Census 2001): India had 102 crore people, with 40 crore workers and 62 crore non-workers.
- Tabular presentation can be:
- one-way: one characteristic, such as literacy by sex
- two-way: two characteristics, such as sex × rural/urban
- three-way: three characteristics, such as sex × rural/urban × state
8. Parts of a statistical table
- Table number: for reference.
- Title: says what, where and when, in short.
- Caption: the column headings.
- Stub: the row headings. The left-most column is the stub column.
- Body: the actual numbers.
- Unit of measurement: %, Rs crore, kg/ha, and so on.
- Source: where the data came from, such as Census or NSS.
- Note: special points, such as definitions or estimates.
- Worked table (a two-way 3×3 table): Census 2011 literacy rate, age 7 and above (%):
| Stub ↓ / Caption → | Rural | Urban | Total |
|---|---|---|---|
| Male | 79 | 90 | 82 |
| Female | 59 | 80 | 65 |
| Total | 68 | 84 | 74 |
- India's overall literacy rose from 64.83% (2001) to 74.04% (2011), a gain of 9.21 percentage points [5]. This matches the "Total-Total" cell (74).
- Reading the table:
- The widest gap is between rural females (59) and urban males (90), a gap of 31 points.
- The gender gap is smaller in urban areas (10 points) than in rural areas (20 points).
9. Diagrammatic presentation
A. Geometric diagrams
- Bar diagram: a one-dimensional diagram. Only the height of each bar carries meaning.
- Bars are equally wide and equally spaced.
-
It suits discrete data, attributes and non-frequency data (for example, male literacy by state in 2011).
-
Multiple bar diagram: two or more bars placed side by side for each item, to compare sets.
-
NCERT example: female literacy, 2001 vs 2011.
- Bihar rose from 33.1 to 53.3, and Jharkhand and UP also rose sharply.
- Kerala was highest at 92.0 and Rajasthan lowest at 52.7 (2011).
- India: 53.7 → 65.5.
-
Component (sub-divided) bar diagram: one bar split into parts to show how a total is made up.
-
NCERT example: in a Bihar district, 41.4% of girls aged 6-14 were out of school, against 8.5% of boys.
-
Pie diagram: a circle divided into sectors in proportion to their shares.
- A full circle is 360° and represents 100%, so 1% = 360/100 = 3.6°.
- Worked example, working status in Census 2011:
- Main workers 29.8% × 3.6 = 107.3° ≈ 107°
- Marginal workers 9.9% × 3.6 = 35.6° ≈ 36°
- Non-workers 60.3% × 3.6 = 217.1° ≈ 217°
- Check: 107 + 36 + 217 = 360°
- NCERT error: Table 4.8 prints the total as 102 crore. But the components (12 + 36 + 73 crore) add up to about 121 crore, which is the 2011 population. So 102 crore is the 2001 figure printed by mistake.
B. Frequency diagrams
- Histogram: a two-dimensional diagram of adjacent rectangles with no gaps. The area of each rectangle (not its height) is proportional to the frequency.
- It is drawn only for continuous data in exclusive classes.
- With unequal classes, the height = frequency density = class frequency ÷ class width.
- Example: class 0-10 with f = 20 → height 20/10 = 2. Class 10-30 with f = 30 → height 30/20 = 1.5.
- Using raw frequency (30) as the height would make the wider class look falsely large.
- It locates the mode graphically. Draw two crossing lines from the tallest bar's corners to the corners of the bars next to it. A vertical line dropped from where they cross gives the mode.
-
NCERT example: daily wages of 85 earners.
-
Frequency polygon: join the midpoints of the tops of the histogram bars with straight lines.
- Close it at both ends at the mid-points of imaginary zero-frequency classes, so its area equals the histogram's.
-
It is the best diagram for comparing two or more distributions on one graph.
-
Frequency curve: a smooth freehand curve drawn through the polygon points.
- Ogive (cumulative frequency curve):
- "Less than" ogive: cumulative frequency plotted against upper limits. It is rising (never falls).
- In the marks table: (10, 1), (20, 9) … (100, 100).
- "More than" ogive: plotted against lower limits. It is falling (never rises).
- (0, 100), (10, 99), (20, 91) …
- The two ogives cross at the median. For the marks data this is ≈ 53 on the X-axis, with Y = N/2 = 50.
C. Arithmetic line graph (time-series graph)
- Time on the X-axis and the value on the Y-axis, with points joined by lines.
- It shows trend (the long-run direction) and periodicity (a pattern that repeats).
- NCERT example, DGCI&S data in Rs 100 crore, 1993-94 to 2013-14:
- Exports rose from 698 to 19,050.
- Imports rose from 731 to 27,154.
- Imports stayed above exports all through, so there was a trade deficit every year.
-
The gap widened after 2001-02.
-
Export shares by destination, 2013-14 (total US$ 314.4 bn): Other Asia 29.4%, GCC 15.3%, USA 12.5%, Other EU 10.9%. This kind of data suits a pie diagram.
- Update: India's total exports of goods and services reached a record about US$ 825 bn in 2024-25, about 6% growth over US$ 778.1 bn in 2023-24 [2][3].
- Merchandise exports: US$ 437.42 bn (2024-25).
- Services exports: US$ 387.54 bn (2024-25).
-
Latest: total exports were estimated at US$ 860.09 bn in 2025-26, up 4.22% from US$ 825.26 bn in 2024-25 [4].
- (NCERT: 2013-14 merchandise exports of US$ 314.4 bn, a narrower measure than goods plus services.)
10. Link to official surveys
- NSS and PLFS reports reach the public as the tables and diagrams described above.
- From 2025, the PLFS annual report follows the calendar year (January-December) instead of July-June [8].
- Any time-series line graph of PLFS indicators must note this change in reference period, so that readers do not compare mismatched periods.
Prelims Hooks
- Class mark = (UL + LL)/2. Class interval = UL − LL. Number of classes = Range ÷ class width, usually 6-15.
- Adjustment of inclusive classes: half the gap (0.5) is subtracted from the LL and added to the UL. So 800-899 → 799.5-899.5.
- Open-ended classes block calculation of the mean, because they have no class mark.
- Pie diagram: 1% = 3.6°. A 25% share = 90°.
- Histogram: area, not height, is proportional to frequency. Unequal classes use frequency density = f ÷ width. It is drawn only for continuous data and locates the mode.
- Bar diagram is one-dimensional, with gaps between bars. Histogram is two-dimensional, with no gaps. This is a common "which of the following" trap.
- The "less than" and "more than" ogives intersect at the median (not the mean or the mode).
- Stub = row headings. Caption = column headings.
- A discrete variable can take fractional values (1/8, 1/16). It just cannot take the values in between.
- India's literacy rate: 64.83% (2001) → 74.04% (2011) [5]. Female literacy was 65.5% (2011).
Mains Points
- Organising data has a cost. Grouping, averages and pie charts compress data but lose detail through the class-mark assumption.
- Policy built only on headline averages (for example, state-average literacy) can hide gaps such as rural female literacy (59%) versus urban male literacy (90%).
-
So we need disaggregated tables (data broken down by sex, rural/urban, social group and district) for targeted schemes.
-
How data are presented shapes how people understand them. Unequal class widths shown without frequency density, pie charts with wrong totals (NCERT Table 4.8), or line graphs across changed survey periods (PLFS moving to the calendar year in 2025 [8]) can mislead.
-
Openly shared metadata and machine-readable releases (MoSPI's CSV/XLSX tables and DataViz [6][7]) make statistics more credible and let others check them.
-
Trade data as a time series. In 1993-94 to 2013-14, imports ran above exports every year and the gap widened after 2001-02.
- In 2024-25 services exports grew 13.6% (to US$ 387.54 bn) while merchandise stayed flat at US$ 437.42 bn [2][3].
- This supports a GS-III answer on India's structural trade deficit and the growing role of services exports.
Sources
- 1Class 11, Ch 2 "Collection of Data"; Class 11, Ch 1 "Introduction (Statistics for Economics)"; Class 11, Ch 3 "Organisation of Data"; Class 11, Ch 4 "Presentation of Data"; Class 11, Ch 5 "Measures of Central Tendency"; Class 11, Ch 6 "Correlation"; Class 11, Ch 8 "Use of Statistical Tools" (primary)
- 2India's Total Exports Grow by 6.01% to Reach Record $824.9 Billion in 2024–25: RBI Report (PIB)pib.gov.in · tier 1
- 3Cumulative exports (merchandise & services) during FY 2024-25 (PIB, Ministry of Commerce)pib.gov.in · tier 1
- 4Cumulative exports during FY 2025-26 estimated at US$ 860.09 Billion (PIB)pib.gov.in · tier 1
- 5Country's Population Reaches 1210 Million as Per Census 2011 (PIB)pib.gov.in · tier 1
- 6MoSPI — Download Tables/Datamospi.gov.in · tier 1
- 7MoSPI — DataViz indexmospi.gov.in · tier 1
- 8Periodic Labour Force Survey (PLFS) Annual Report, 2025 [January–December 2025] (PIB)pib.gov.in · tier 1