Data Collection, Classification and Presentation for SSC CGL Statistics

beginner 18 min read

Concept

Data, in raw form, is just a pile of numbers. Before you can extract any meaning from it — before you can calculate averages, spot trends, or compare groups — you have to collect it systematically, sort it into categories, and present it in a form that the human eye can parse quickly. That three-step flow — collect, classify, present — is what this entire chapter is about.

Think of it like a government census. Field officers collect raw responses from households (collection). The data is then sorted by state, age group, income bracket, and so on (classification). Finally, the Planning Commission publishes tables and charts in annual reports (presentation). The same logic applies to any dataset you encounter in an exam question.

Data collection refers to the process of gathering observations or measurements. Primary data is collected directly by the researcher through surveys, experiments, or interviews. Secondary data is borrowed from existing sources — published reports, databases, previous studies. In exam questions, the distinction rarely matters beyond knowing which is which.

Classification is the process of arranging raw data into groups or classes that share a common characteristic. Quantitative classification uses numerical ranges (class intervals like 10–20, 20–30). Qualitative classification uses categories (gender, district, occupation). The output of classification is a frequency distribution table — probably the most tested object in this chapter.

Presentation converts the classified data into visual form: tables, bar diagrams, histograms, frequency polygons, ogives, and pie charts. Each diagram has a specific purpose. Confusing them — especially histogram vs. bar chart, or ogive vs. frequency polygon — is the single most common error in this chapter.

The analogy that sticks: raw data is a box of unsorted exam answer sheets. Classification is sorting them by score range. Presentation is the bar chart on the principal's notice board. Each step has rules, and the exam tests whether you know those rules precisely.


Deep Dive

Frequency Distribution Tables

Given a dataset, the first task is to build a frequency distribution. Here are the key terms you must know cold:

Number of classes formula: Number of classes = Range / Class width, where Range = Highest value − Lowest value. If the result is not a whole number, round up — never round down, or some observations will fall outside the last class.

Cumulative frequency is constructed by adding each class frequency to the sum of all previous class frequencies. For a "less than" ogive, you accumulate from the top; for a "more than" ogive, you accumulate from the bottom.

Types of Graphical Presentation

Histogram: Bars are drawn on a continuous scale with no gaps between them, because the x-axis represents a continuous variable (class intervals). The area of each bar is proportional to the frequency of that class. When class widths are unequal, you must plot frequency density (= frequency / class width) on the y-axis, not raw frequency. This is a detail many candidates miss.

Bar diagram: Bars have gaps between them, because the x-axis represents discrete or categorical data. Heights represent frequencies or values. Histogram ≠ bar diagram — the gap is the giveaway.

Frequency polygon: Connect the midpoints (class marks) of the tops of histogram bars with straight lines. Extend the polygon to the x-axis by adding hypothetical classes of zero frequency at each end. It shows the shape of the distribution.

Ogive (cumulative frequency curve): Plot cumulative frequency against class boundaries.

Pie chart: Represents data as sectors of a circle. Sector angle for a category = (Category frequency / Total frequency) × 360°. Purpose: showing part-to-whole relationships. It cannot show trends over time; it cannot compare absolute values meaningfully across multiple datasets.

Bivariate / Contingency Tables

When data is classified by two attributes simultaneously, you get a contingency table (also called a cross-tabulation or two-way table). If attribute A has r categories and attribute B has c categories, the inner grid is r × c cells. Adding a column for row totals gives r × (c+1) cells in the data rows, and adding a row for column totals (with a grand total cell) adds (c+1) more cells. Total cells = (r+1)(c+1).

Classification of Data

Geographical: Data classified by region (state, district). Chronological (temporal): Classified by time period (monthly, yearly). Qualitative: Classified by attribute (gender, religion, occupation). Quantitative: Classified by measurable variable (age, income, marks).


Memory Tricks & Shortcuts

patternCF Add-Down Rule

Cumulative frequency questions almost always give you a small table (3–5 rows) and ask for CF at a specific boundary. Don't look for a formula. Just add downward: CF at any boundary = sum of all frequencies from the first class through that class. With a 3-row table, this takes 5 seconds. The trap is candidates subtract instead of add — keep the direction clear (always accumulating downward). Standard method: write each CF separately (15s). Add-down: running mental total (5s).

estimationRelative Frequency Percentage Check

When asked for relative frequency %, first mentally check: total frequency should divide the class frequency cleanly or near-cleanly. If total = 45 and class f = 18, notice 18/45 = 2/5 = 40%. Convert to a simple fraction first, then to %. The options will include distractors at 36% (18/50) and 30% (trying 18/60). Using the fraction route kills the arithmetic in 8 seconds vs. long division in 25 seconds.

patternNumber of Classes — Always Round Up

Formula: Number of classes = Range / Class width. If the answer is a decimal like 5.9, always round up to 6 — never round down to 5, because a 5-class system would leave the value 79 outside the last class boundary. Memorize the direction: decimals go up, never down. This single rule eliminates the "5 vs 6" trap in every number-of-classes question. One decision, saves 20 seconds of re-checking.

eliminationDiagram Selector — One-Question Elimination

Four diagrams appear in almost every options set. Assign each a one-word purpose: Histogram → shape/spread, Frequency polygon → trend/comparison, Ogive → cumulative/median/quartiles, Pie chart → parts-of-whole. When the question says "composition" or "parts of a total" or "budget allocation", the answer is always pie chart. When it says "median graphically", the answer is ogive. This two-second label lookup eliminates 3 wrong options without calculation.

patternContingency Table Cell Count — (r+1)(c+1)

Instead of counting cells manually, use the formula (r+1)(c+1) where r = number of row categories, c = number of column categories. The +1 accounts for the totals row/column. For r=3, c=4: (3+1)(4+1) = 4×5 = 20. Manual counting on a drawn grid takes 30+ seconds. The formula takes 5 seconds.


Fast-Solving Framework

When you see a data presentation question in the exam hall, run through this decision path:

  1. Is it a cumulative frequency question? → Add all class frequencies from the start through the specified boundary. Done.

  2. Is it a relative frequency / percentage question? → Sum all frequencies for total N. Divide class frequency by N. Convert to fraction first, then multiply by 100.

  3. Is it a number-of-classes question? → Compute Range = Highest − Lowest boundary. Divide by class width. Round up if decimal.

  4. Is it asking which diagram to use? → Apply the one-word purpose labels: composition = pie, cumulative = ogive, frequency distribution = histogram, shape over classes = frequency polygon.

  5. Is it a contingency table structure question? → Use (r+1)(c+1) for total cells including marginals.

  6. Is it an ogive-reading question? → Identify whether "less than" (upper boundaries, rising curve) or "more than" (lower boundaries, falling curve). Read the y-axis value at the required x-coordinate.

If none of the above, read the definitions precisely — most remaining questions are conceptual definition matches.


Solved PYQs

Why this question: Tests the most fundamental skill in this chapter — computing cumulative frequency from a two-row table. If you miss this, you will miss every ogive question too.

Previous Year Questionपिछले वर्ष का प्रश्न
A frequency distribution table shows that the class interval 20–30 has a frequency of 15 and the class interval 30–40 has a frequency of 25. What is the cumulative frequency up to 40?
एक बारंबारता वितरण तालिका में वर्ग-अंतराल 20–30 की बारंबारता 15 है और 30–40 की बारंबारता 25 है। 40 तक की संचयी बारंबारता क्या होगी?
  1. 35
  2. 40
  3. 15
  4. 25
  1. 35
  2. 40
  3. 15
  4. 25
Solutionसमाधान
Cumulative frequency is the running total of frequencies up to a given class boundary. Adding the frequencies of the first two classes: 15 (for 20–30) + 25 (for 30–40) = 40. So the cumulative frequency up to 40 is 40.
संचयी बारंबारता एक निश्चित वर्ग-सीमा तक सभी बारंबारताओं का कुल योग होती है। 20–30 की बारंबारता (15) और 30–40 की बारंबारता (25) को जोड़ने पर: 15 + 25 = 40। अतः 40 तक की संचयी बारंबारता 40 है।

Solving path: CF up to 40 = frequency of (20–30) + frequency of (30–40) = 15 + 25 = 40. No formula required — straight addition. Distractor "35" traps candidates who subtract instead of add.


Why this question: Tests diagram-purpose knowledge — the most recycled conceptual question in this chapter.

Previous Year Questionपिछले वर्ष का प्रश्न
Which type of diagram is most suitable for showing the relationship between the parts and the whole of a dataset?
किसी डेटा के भागों और समग्र के बीच संबंध दर्शाने के लिए कौन-सा आरेख सबसे उपयुक्त है?
  1. Histogram
  2. Frequency polygon
  3. Ogive
  4. Pie chart
  1. हिस्टोग्राम
  2. बारंबारता बहुभुज
  3. ओजाइव
  4. पाई चार्ट
Solutionसमाधान
A pie chart represents data as slices of a circle, where each slice shows the proportion of each category relative to the total. It is best suited for displaying part-to-whole relationships. A histogram shows frequency distribution, a frequency polygon connects midpoints of class intervals, and an ogive shows cumulative frequency.
पाई चार्ट डेटा को एक वृत्त के टुकड़ों के रूप में दर्शाता है, जहाँ प्रत्येक टुकड़ा कुल के सापेक्ष प्रत्येक श्रेणी का अनुपात दिखाता है। यह भाग-से-समग्र संबंध दर्शाने के लिए सबसे उपयुक्त है। हिस्टोग्राम बारंबारता वितरण दिखाता है, बारंबारता बहुभुज वर्ग-मध्यबिंदुओं को जोड़ता है, और ओजाइव संचयी बारंबारता दर्शाता है।

Solving path: Apply the one-word label system. "Parts and whole" → pie chart. Eliminate histogram (frequency shape), frequency polygon (trend), ogive (cumulative). One decision, zero calculation.


Why this question: Tests relative frequency calculation with a clean fraction — a standard numerical question that appears across tiers.

Previous Year Questionपिछले वर्ष का प्रश्न
A frequency distribution table shows the following data: Class: 10-20, 20-30, 30-40, 40-50 Frequency: 5, 12, 18, 10 What is the relative frequency (in %) of the class 30-40?
एक बारंबारता वितरण तालिका निम्नलिखित डेटा दिखाती है: वर्ग: 10-20, 20-30, 30-40, 40-50 बारंबारता: 5, 12, 18, 10 वर्ग 30-40 की सापेक्ष बारंबारता (% में) क्या है?
  1. 45%
  2. 30%
  3. 40%
  4. 36%
  1. 45%
  2. 30%
  3. 40%
  4. 36%
Solutionसमाधान
Total frequency = 5 + 12 + 18 + 10 = 45. Relative frequency of class 30-40 = (18/45) × 100 = 40%. This represents the proportion of observations falling in the class 30-40.
कुल बारंबारता = 5 + 12 + 18 + 10 = 45। वर्ग 30-40 की सापेक्ष बारंबारता = (18/45) × 100 = 40%। यह 30-40 वर्ग में आने वाले प्रेक्षणों का अनुपात दर्शाती है।

Solving path: Total N = 5+12+18+10 = 45. Class 30–40 frequency = 18. Relative frequency = 18/45 = 2/5 = 40%. Distractor 36% comes from using 50 as total (wrong). Distractor 30% from using 60 as total (wrong). Always verify your total before dividing.


Why this question: Tests the number-of-classes formula and the rounding rule — a calculation question that catches candidates who round down.

Previous Year Questionपिछले वर्ष का प्रश्न
The following data represents marks of 40 students. If the class width (class size) is 10 and the lowest class boundary is 20, how many classes will be formed if the highest mark is 79? (Use the formula: Number of classes = Range / Class width)
40 छात्रों के अंकों का डेटा दिया गया है। यदि वर्ग चौड़ाई 10 है और सबसे निचली वर्ग सीमा 20 है, तो यदि सबसे अधिक अंक 79 हैं, तो कितने वर्ग बनेंगे? (सूत्र: वर्गों की संख्या = परिसर / वर्ग चौड़ाई)
  1. 7
  2. 5
  3. 6
  4. 8
  1. 7
  2. 5
  3. 6
  4. 8
Solutionसमाधान
Range = Highest value – Lowest class boundary = 79 – 20 = 59. Number of classes = Range / Class width = 59/10 = 5.9, which is rounded up to 6. So 6 classes are formed: 20-30, 30-40, 40-50, 50-60, 60-70, 70-80.
परिसर = अधिकतम मान – न्यूनतम वर्ग सीमा = 79 – 20 = 59। वर्गों की संख्या = 59/10 = 5.9, जिसे पूर्णांकित करके 6 किया जाता है। अतः 6 वर्ग बनते हैं: 20-30, 30-40, 40-50, 50-60, 60-70, 70-80।

Solving path: Range = 79 − 20 = 59. Classes = 59/10 = 5.9 → round up to 6. The classes are 20–30, 30–40, 40–50, 50–60, 60–70, 70–80, which covers 79. Rounding down to 5 would leave the last class at 60–70, missing observations from 70–79. Always round up.


Why this question: Tests ogive structure — what each axis represents — a definition question that trips candidates who confuse ogive with frequency polygon.

Previous Year Questionपिछले वर्ष का प्रश्न
In a 'less than' ogive, the y-axis represents:
'से कम' ओजाइव में y-अक्ष क्या दर्शाता है?
  1. Cumulative frequency
  2. Relative frequency
  3. Class frequency
  4. Class mark
  1. संचयी बारंबारता
  2. सापेक्ष बारंबारता
  3. वर्ग बारंबारता
  4. वर्ग मध्य बिंदु
Solutionसमाधान
An ogive (cumulative frequency curve) plots cumulative frequency on the y-axis against upper class boundaries (for 'less than' ogive) on the x-axis. It is used to determine the median, quartiles, and percentiles graphically.
ओजाइव (संचयी बारंबारता वक्र) में y-अक्ष पर संचयी बारंबारता और x-अक्ष पर ऊपरी वर्ग सीमाएँ ('से कम' ओजाइव के लिए) दर्शाई जाती हैं। इसका उपयोग मध्यिका, चतुर्थक और शतमक ज्ञात करने के लिए होता है।

Solving path: An ogive is specifically a cumulative frequency curve. The y-axis always carries cumulative frequency. The x-axis carries class boundaries (upper boundaries for "less than" ogive). Relative frequency is a ratio (not cumulative), class frequency is the raw count, class mark is the midpoint — none of these are on the y-axis of an ogive.


Why this question: Tests multi-step CF calculation for a "less than" ogive — requires adding three frequencies without error.

Previous Year Questionपिछले वर्ष का प्रश्न
A frequency distribution has class intervals 10–20, 20–30, 30–40, 40–50, and 50–60 with frequencies 8, 15, 22, 12, and 3 respectively. If the data is represented as a 'less than' ogive, what is the cumulative frequency corresponding to the upper boundary of the third class interval?
एक आवृत्ति वितरण में वर्ग-अंतराल 10–20, 20–30, 30–40, 40–50, और 50–60 हैं और उनकी आवृत्तियाँ क्रमशः 8, 15, 22, 12, और 3 हैं। यदि डेटा को 'से कम' ओजाइव के रूप में प्रस्तुत किया जाए, तो तीसरे वर्ग-अंतराल की ऊपरी सीमा के संगत संचयी आवृत्ति क्या होगी?
  1. 52
  2. 37
  3. 22
  4. 45
  1. 52
  2. 37
  3. 22
  4. 45
Solutionसमाधान
The cumulative frequency for the 'less than' ogive is calculated by adding frequencies progressively: up to 20 = 8, up to 30 = 8+15 = 23, up to 40 = 23+22 = 45. The third class interval is 30–40, and its upper boundary is 40, so the cumulative frequency is 45.
'से कम' ओजाइव के लिए संचयी आवृत्ति क्रमिक रूप से जोड़कर निकाली जाती है: 20 तक = 8, 30 तक = 8+15 = 23, 40 तक = 23+22 = 45। तीसरा वर्ग-अंतराल 30–40 है और इसकी ऊपरी सीमा 40 है, अतः संचयी आवृत्ति 45 है।

Solving path: CF up to 20 = 8. CF up to 30 = 8+15 = 23. CF up to 40 = 23+22 = 45. The third class interval is 30–40, upper boundary is 40, so the answer is 45. Distractor 37 = 8+15+14 (partial count of third class). Distractor 52 = 8+15+22+7 (partial overshoot). Add cleanly in sequence.


Why this question: Tests contingency table structure — a higher-order question that requires knowing the marginals formula, not just the inner grid.

Previous Year Questionपिछले वर्ष का प्रश्न
In a bivariate data set, 200 observations are classified by two attributes A and B. Attribute A has 3 categories and attribute B has 4 categories. If a complete two-way contingency table (including marginal totals) is constructed, what is the total number of cells in the table, including all row totals, column totals, and the grand total?
एक द्विचर (bivariate) डेटा सेट में 200 प्रेक्षणों को दो गुणों A और B के आधार पर वर्गीकृत किया गया है। गुण A की 3 श्रेणियाँ और गुण B की 4 श्रेणियाँ हैं। यदि एक पूर्ण द्वि-दिशा सारणी (पंक्ति योग, स्तंभ योग और महायोग सहित) बनाई जाए, तो सारणी में कुल कितने कोष्ठक (cells) होंगे?
  1. 15
  2. 12
  3. 20
  4. 16
  1. 15
  2. 12
  3. 20
  4. 16
Solutionसमाधान
A two-way contingency table for A (3 categories) and B (4 categories) has 3 rows × 4 columns = 12 inner cells. Adding 1 column for row totals gives 3×5 = 15 cells in data rows. Adding 1 row for column totals (4 cells + 1 grand total cell) gives 5 more cells. Total = 15 + 5 = 20 cells.
A (3 श्रेणियाँ) और B (4 श्रेणियाँ) के लिए द्वि-दिशा सारणी में 3×4 = 12 आंतरिक कोष्ठक होते हैं। पंक्ति योग के लिए 1 अतिरिक्त स्तंभ जोड़ने पर 3×5 = 15 कोष्ठक बनते हैं। स्तंभ योग की 1 अतिरिक्त पंक्ति (4 + 1 महायोग) = 5 कोष्ठक और जोड़ने पर कुल 15+5 = 20 कोष्ठक होते हैं।

Solving path: Inner grid = 3 rows × 4 columns = 12 cells. Including row totals column: 3 × 5 = 15. Including column totals row + grand total: +5 more cells. Total = 20. Shortcut: (r+1)(c+1) = (3+1)(4+1) = 4 × 5 = 20. Distractor 12 ignores all marginals. Distractor 16 counts only partial marginals.


Common Mistakes


Related Topics


Practice on SarkariRise

Sign up + get 3 free mocks →