Data, in raw form, is just a pile of numbers. Before you can extract any meaning from it — before you can calculate averages, spot trends, or compare groups — you have to collect it systematically, sort it into categories, and present it in a form that the human eye can parse quickly. That three-step flow — collect, classify, present — is what this entire chapter is about.
Think of it like a government census. Field officers collect raw responses from households (collection). The data is then sorted by state, age group, income bracket, and so on (classification). Finally, the Planning Commission publishes tables and charts in annual reports (presentation). The same logic applies to any dataset you encounter in an exam question.
Data collection refers to the process of gathering observations or measurements. Primary data is collected directly by the researcher through surveys, experiments, or interviews. Secondary data is borrowed from existing sources — published reports, databases, previous studies. In exam questions, the distinction rarely matters beyond knowing which is which.
Classification is the process of arranging raw data into groups or classes that share a common characteristic. Quantitative classification uses numerical ranges (class intervals like 10–20, 20–30). Qualitative classification uses categories (gender, district, occupation). The output of classification is a frequency distribution table — probably the most tested object in this chapter.
Presentation converts the classified data into visual form: tables, bar diagrams, histograms, frequency polygons, ogives, and pie charts. Each diagram has a specific purpose. Confusing them — especially histogram vs. bar chart, or ogive vs. frequency polygon — is the single most common error in this chapter.
The analogy that sticks: raw data is a box of unsorted exam answer sheets. Classification is sorting them by score range. Presentation is the bar chart on the principal's notice board. Each step has rules, and the exam tests whether you know those rules precisely.
Given a dataset, the first task is to build a frequency distribution. Here are the key terms you must know cold:
(Lower boundary + Upper boundary) / 2 = 25 for 20–30.f / N, where N is total observations. Multiply by 100 for percentage.Number of classes formula: Number of classes = Range / Class width, where Range = Highest value − Lowest value. If the result is not a whole number, round up — never round down, or some observations will fall outside the last class.
Cumulative frequency is constructed by adding each class frequency to the sum of all previous class frequencies. For a "less than" ogive, you accumulate from the top; for a "more than" ogive, you accumulate from the bottom.
Histogram: Bars are drawn on a continuous scale with no gaps between them, because the x-axis represents a continuous variable (class intervals). The area of each bar is proportional to the frequency of that class. When class widths are unequal, you must plot frequency density (= frequency / class width) on the y-axis, not raw frequency. This is a detail many candidates miss.
Bar diagram: Bars have gaps between them, because the x-axis represents discrete or categorical data. Heights represent frequencies or values. Histogram ≠ bar diagram — the gap is the giveaway.
Frequency polygon: Connect the midpoints (class marks) of the tops of histogram bars with straight lines. Extend the polygon to the x-axis by adding hypothetical classes of zero frequency at each end. It shows the shape of the distribution.
Ogive (cumulative frequency curve): Plot cumulative frequency against class boundaries.
Pie chart: Represents data as sectors of a circle. Sector angle for a category = (Category frequency / Total frequency) × 360°. Purpose: showing part-to-whole relationships. It cannot show trends over time; it cannot compare absolute values meaningfully across multiple datasets.
When data is classified by two attributes simultaneously, you get a contingency table (also called a cross-tabulation or two-way table). If attribute A has r categories and attribute B has c categories, the inner grid is r × c cells. Adding a column for row totals gives r × (c+1) cells in the data rows, and adding a row for column totals (with a grand total cell) adds (c+1) more cells. Total cells = (r+1)(c+1).
Geographical: Data classified by region (state, district). Chronological (temporal): Classified by time period (monthly, yearly). Qualitative: Classified by attribute (gender, religion, occupation). Quantitative: Classified by measurable variable (age, income, marks).
Cumulative frequency questions almost always give you a small table (3–5 rows) and ask for CF at a specific boundary. Don't look for a formula. Just add downward: CF at any boundary = sum of all frequencies from the first class through that class. With a 3-row table, this takes 5 seconds. The trap is candidates subtract instead of add — keep the direction clear (always accumulating downward). Standard method: write each CF separately (15s). Add-down: running mental total (5s).
When asked for relative frequency %, first mentally check: total frequency should divide the class frequency cleanly or near-cleanly. If total = 45 and class f = 18, notice 18/45 = 2/5 = 40%. Convert to a simple fraction first, then to %. The options will include distractors at 36% (18/50) and 30% (trying 18/60). Using the fraction route kills the arithmetic in 8 seconds vs. long division in 25 seconds.
Formula: Number of classes = Range / Class width. If the answer is a decimal like 5.9, always round up to 6 — never round down to 5, because a 5-class system would leave the value 79 outside the last class boundary. Memorize the direction: decimals go up, never down. This single rule eliminates the "5 vs 6" trap in every number-of-classes question. One decision, saves 20 seconds of re-checking.
Four diagrams appear in almost every options set. Assign each a one-word purpose: Histogram → shape/spread, Frequency polygon → trend/comparison, Ogive → cumulative/median/quartiles, Pie chart → parts-of-whole. When the question says "composition" or "parts of a total" or "budget allocation", the answer is always pie chart. When it says "median graphically", the answer is ogive. This two-second label lookup eliminates 3 wrong options without calculation.
Instead of counting cells manually, use the formula (r+1)(c+1) where r = number of row categories, c = number of column categories. The +1 accounts for the totals row/column. For r=3, c=4: (3+1)(4+1) = 4×5 = 20. Manual counting on a drawn grid takes 30+ seconds. The formula takes 5 seconds.
When you see a data presentation question in the exam hall, run through this decision path:
Is it a cumulative frequency question? → Add all class frequencies from the start through the specified boundary. Done.
Is it a relative frequency / percentage question? → Sum all frequencies for total N. Divide class frequency by N. Convert to fraction first, then multiply by 100.
Is it a number-of-classes question? → Compute Range = Highest − Lowest boundary. Divide by class width. Round up if decimal.
Is it asking which diagram to use? → Apply the one-word purpose labels: composition = pie, cumulative = ogive, frequency distribution = histogram, shape over classes = frequency polygon.
Is it a contingency table structure question? → Use (r+1)(c+1) for total cells including marginals.
Is it an ogive-reading question? → Identify whether "less than" (upper boundaries, rising curve) or "more than" (lower boundaries, falling curve). Read the y-axis value at the required x-coordinate.
If none of the above, read the definitions precisely — most remaining questions are conceptual definition matches.
Why this question: Tests the most fundamental skill in this chapter — computing cumulative frequency from a two-row table. If you miss this, you will miss every ogive question too.
Solving path: CF up to 40 = frequency of (20–30) + frequency of (30–40) = 15 + 25 = 40. No formula required — straight addition. Distractor "35" traps candidates who subtract instead of add.
Why this question: Tests diagram-purpose knowledge — the most recycled conceptual question in this chapter.
Solving path: Apply the one-word label system. "Parts and whole" → pie chart. Eliminate histogram (frequency shape), frequency polygon (trend), ogive (cumulative). One decision, zero calculation.
Why this question: Tests relative frequency calculation with a clean fraction — a standard numerical question that appears across tiers.
Solving path: Total N = 5+12+18+10 = 45. Class 30–40 frequency = 18. Relative frequency = 18/45 = 2/5 = 40%. Distractor 36% comes from using 50 as total (wrong). Distractor 30% from using 60 as total (wrong). Always verify your total before dividing.
Why this question: Tests the number-of-classes formula and the rounding rule — a calculation question that catches candidates who round down.
Solving path: Range = 79 − 20 = 59. Classes = 59/10 = 5.9 → round up to 6. The classes are 20–30, 30–40, 40–50, 50–60, 60–70, 70–80, which covers 79. Rounding down to 5 would leave the last class at 60–70, missing observations from 70–79. Always round up.
Why this question: Tests ogive structure — what each axis represents — a definition question that trips candidates who confuse ogive with frequency polygon.
Solving path: An ogive is specifically a cumulative frequency curve. The y-axis always carries cumulative frequency. The x-axis carries class boundaries (upper boundaries for "less than" ogive). Relative frequency is a ratio (not cumulative), class frequency is the raw count, class mark is the midpoint — none of these are on the y-axis of an ogive.
Why this question: Tests multi-step CF calculation for a "less than" ogive — requires adding three frequencies without error.
Solving path: CF up to 20 = 8. CF up to 30 = 8+15 = 23. CF up to 40 = 23+22 = 45. The third class interval is 30–40, upper boundary is 40, so the answer is 45. Distractor 37 = 8+15+14 (partial count of third class). Distractor 52 = 8+15+22+7 (partial overshoot). Add cleanly in sequence.
Why this question: Tests contingency table structure — a higher-order question that requires knowing the marginals formula, not just the inner grid.
Solving path: Inner grid = 3 rows × 4 columns = 12 cells. Including row totals column: 3 × 5 = 15. Including column totals row + grand total: +5 more cells. Total = 20. Shortcut: (r+1)(c+1) = (3+1)(4+1) = 4 × 5 = 20. Distractor 12 ignores all marginals. Distractor 16 counts only partial marginals.
Rounding down in number-of-classes problems. If Range / Class width = 5.9, candidates write 5. The correct answer is 6 — always round up, or you'll leave data outside the last class boundary.
Confusing histogram with bar diagram. A histogram has no gaps between bars (continuous data). A bar diagram has gaps (discrete or categorical data). When an exam question asks which diagram "has adjacent bars touching", the answer is histogram, not bar diagram.
Misidentifying what the y-axis of an ogive carries. It is always cumulative frequency, not relative frequency and not raw class frequency. This distinction appears directly as a question option.
Using the wrong total in relative frequency. Always recompute the total from the given frequencies — do not assume a round number like 50 or 100 unless the table explicitly states it. The table in the exam will often have a total like 45 that catches candidates who assume 50.
Confusing "less than" and "more than" ogive direction. A "less than" ogive is a rising S-curve plotted against upper class boundaries. A "more than" ogive is a falling S-curve plotted against lower class boundaries. Their intersection gives the median — not their sum, not their difference.
Forgetting marginal totals in contingency table cell counts. The inner grid has r × c cells, but the complete table including all row totals, column totals, and the grand total has (r+1)(c+1) cells. Stopping at r × c = 12 instead of (r+1)(c+1) = 20 is a consistent error in this question type.