Why Every Frequency Polygon Needs Two Extra Zeros
Tell raw, arrayed and grouped data apart, build a frequency table with tally marks, convert discontinuous class intervals into continuous ones, and draw a frequency polygon that closes properly.
What is the difference between raw, arrayed and grouped data?
Raw data is the list as collected, arrayed data is the same list sorted, and grouped data is that list summarised into classes with counts.
Suppose you note the marks of students out of in the order they hand in their papers:
That is raw data — accurate, complete and almost unreadable. Sorting it into ascending order gives arrayed data, which immediately shows the smallest value , the largest , and therefore the range:
Counting how many marks fall in each ten-mark band gives grouped data, and only then can you see the shape of the class's performance.
Each step trades detail for clarity. Grouping loses the individual marks forever — you can no longer recover that someone scored exactly — but it lets you see at a glance where most of the class sits. That trade is the whole subject of statistics, and every technique in this chapter is a choice about how much detail to give up.
The variable matters too. Marks out of can only be whole numbers, so that variable is discrete. The heights of the same students can take any value in between — cm is as real as cm — so height is continuous. Numbers of members in a household, runs scored and vehicles counted at a crossing are discrete; height, mass, temperature and time taken are continuous.
This page covers the ICSE Class 9 Mathematics chapter on statistics: types of data and variables, frequency tables with tally marks, class limits and boundaries, and the frequency polygon.
Suppose you note the marks of students out of in the order they hand in their papers:
That is raw data — accurate, complete and almost unreadable. Sorting it into ascending order gives arrayed data, which immediately shows the smallest value , the largest , and therefore the range:
Counting how many marks fall in each ten-mark band gives grouped data, and only then can you see the shape of the class's performance.
Each step trades detail for clarity. Grouping loses the individual marks forever — you can no longer recover that someone scored exactly — but it lets you see at a glance where most of the class sits. That trade is the whole subject of statistics, and every technique in this chapter is a choice about how much detail to give up.
The variable matters too. Marks out of can only be whole numbers, so that variable is discrete. The heights of the same students can take any value in between — cm is as real as cm — so height is continuous. Numbers of members in a household, runs scored and vehicles counted at a crossing are discrete; height, mass, temperature and time taken are continuous.
This page covers the ICSE Class 9 Mathematics chapter on statistics: types of data and variables, frequency tables with tally marks, class limits and boundaries, and the frequency polygon.
How do you build a frequency distribution table with tally marks?
Find the range, choose a class size, list the classes, then tally each value once into exactly one class.
Step one — the number of classes. Divide the range by the class size you want:
Always round up, since a part-class still needs a class of its own. Between five and ten classes is usual: too few hides the shape, too many leaves most classes almost empty.
Step two — the classes. With a class size of starting at , the classes are -, -, - and -.
Step three — the convention. These are exclusive classes: the upper limit belongs to the next class. So goes into -, not into -. Decide this once and apply it to every value, or some values get counted twice and others not at all.
Step four — tally and count. Working through the raw list in order and putting one stroke per value, grouping them in fives:
- -: the values give a frequency of
- -: the values give a frequency of
- -: the values give a frequency of
- -: the values give a frequency of
Step five — the check that must never be skipped. Add the frequencies:
and that must equal the number of observations you started with. If the total is wrong, the table is wrong, and no later calculation can rescue it.
Why tally marks rather than counting? Because you pass through the raw list once, in its original order, placing each value as you meet it. Counting class by class means reading the whole list four times and it is where values get missed. The fifth stroke drawn across the other four makes the groups readable at a glance, so the final count is a multiplication rather than a recount.
Step one — the number of classes. Divide the range by the class size you want:
Always round up, since a part-class still needs a class of its own. Between five and ten classes is usual: too few hides the shape, too many leaves most classes almost empty.
Step two — the classes. With a class size of starting at , the classes are -, -, - and -.
Step three — the convention. These are exclusive classes: the upper limit belongs to the next class. So goes into -, not into -. Decide this once and apply it to every value, or some values get counted twice and others not at all.
Step four — tally and count. Working through the raw list in order and putting one stroke per value, grouping them in fives:
- -: the values give a frequency of
- -: the values give a frequency of
- -: the values give a frequency of
- -: the values give a frequency of
Step five — the check that must never be skipped. Add the frequencies:
and that must equal the number of observations you started with. If the total is wrong, the table is wrong, and no later calculation can rescue it.
Why tally marks rather than counting? Because you pass through the raw list once, in its original order, placing each value as you meet it. Counting class by class means reading the whole list four times and it is where values get missed. The fifth stroke drawn across the other four makes the groups readable at a glance, so the final count is a multiplication rather than a recount.
What is the difference between class limits and class boundaries?
Class limits are the numbers written in the table; class boundaries are the true dividing points, and they differ whenever there is a gap between one class and the next.
Marks are whole numbers, so a table is often written in the inclusive form, where both end values belong to the class:
- -
- -
- -
- -
There is a gap here. Nothing in the table covers the space between and . For a discrete variable that is harmless, but a histogram or a polygon needs bars that touch, so the classes must be made continuous.
The adjustment factor is half the gap.
Subtract it from every lower limit and add it to every upper limit:
- - becomes -
- - becomes -
- - becomes -
- - becomes -
Now the classes touch, and the frequencies have not changed at all — no value was moved, only the labels on the walls.
Worked example — the four quantities. For the class -:
- class limits: lower , upper
- class boundaries: lower , upper
- class size:
- class mark (the mid-value):
Check the class mark a second way: the average of the two class limits is as well. The class mark is the same whichever pair you average, because the adjustment moves both ends by the same amount.
**Notice that the class size is , not .** Reading it as is the standard error, and it matters because class size enters the histogram's bar width and, in Class 10, the mean and median formulas for grouped data.
And here is why the difference is worth this much care. If a class is written - and you use and as if they were boundaries, every class mark and every bar width is slightly wrong, and a mean computed from them drifts by half a unit. Convert to boundaries first, then calculate — always in that order.
Marks are whole numbers, so a table is often written in the inclusive form, where both end values belong to the class:
- -
- -
- -
- -
There is a gap here. Nothing in the table covers the space between and . For a discrete variable that is harmless, but a histogram or a polygon needs bars that touch, so the classes must be made continuous.
The adjustment factor is half the gap.
Subtract it from every lower limit and add it to every upper limit:
- - becomes -
- - becomes -
- - becomes -
- - becomes -
Now the classes touch, and the frequencies have not changed at all — no value was moved, only the labels on the walls.
Worked example — the four quantities. For the class -:
- class limits: lower , upper
- class boundaries: lower , upper
- class size:
- class mark (the mid-value):
Check the class mark a second way: the average of the two class limits is as well. The class mark is the same whichever pair you average, because the adjustment moves both ends by the same amount.
**Notice that the class size is , not .** Reading it as is the standard error, and it matters because class size enters the histogram's bar width and, in Class 10, the mean and median formulas for grouped data.
And here is why the difference is worth this much care. If a class is written - and you use and as if they were boundaries, every class mark and every bar width is slightly wrong, and a mean computed from them drifts by half a unit. Convert to boundaries first, then calculate — always in that order.
How do you draw a frequency polygon, with or without a histogram?
Plot the class mark against the frequency for each class, join the points with straight lines, and close the figure by adding a class of zero frequency at each end.
Worked example. Draw a frequency polygon for this distribution:
- -: frequency
- -: frequency
- -: frequency
- -: frequency
- -: frequency
First the total, as a check: observations.
Now the class marks, each the average of the two boundaries:
So the points to plot are , , , and .
The two extra zeros. Add an imaginary class before the first, from to with class mark , and one after the last, from to with class mark , each with frequency . That gives two more points, and , and joining all seven in order closes the figure onto the horizontal axis.
With a histogram: draw the bars first, mark the mid-point of the top of each bar, and join those mid-points. Without a histogram: plot the class marks directly against the frequencies. The polygon is identical either way.
Why the zeros are not optional. A polygon is a closed figure, and without the two end points the line stops in mid-air. More importantly, closing it makes the area under the polygon equal to the area of the histogram: each little triangle the straight line cuts off the top of a bar is exactly matched by a triangle it adds to the neighbouring bar. The polygon redistributes area without changing the total, which is why it represents the same data faithfully.
Check that claim on the example. Between class marks and the line falls from to , crossing the class boundary at a height of . The triangle it cuts off the -bar and the triangle it adds to the -bar are congruent — both have base and height — so the areas cancel. Equal class widths are what make that cancellation exact, and it is the reason a frequency polygon should not be drawn for classes of unequal width.
The practical advantage over a histogram. Two polygons can be drawn on the same pair of axes and compared directly — two sections' marks, or the same test before and after revision — because they are lines rather than solid bars. A histogram cannot be overlaid on another histogram without hiding one of them, and that is the main reason this diagram exists.
Worked example. Draw a frequency polygon for this distribution:
- -: frequency
- -: frequency
- -: frequency
- -: frequency
- -: frequency
First the total, as a check: observations.
Now the class marks, each the average of the two boundaries:
So the points to plot are , , , and .
The two extra zeros. Add an imaginary class before the first, from to with class mark , and one after the last, from to with class mark , each with frequency . That gives two more points, and , and joining all seven in order closes the figure onto the horizontal axis.
With a histogram: draw the bars first, mark the mid-point of the top of each bar, and join those mid-points. Without a histogram: plot the class marks directly against the frequencies. The polygon is identical either way.
Why the zeros are not optional. A polygon is a closed figure, and without the two end points the line stops in mid-air. More importantly, closing it makes the area under the polygon equal to the area of the histogram: each little triangle the straight line cuts off the top of a bar is exactly matched by a triangle it adds to the neighbouring bar. The polygon redistributes area without changing the total, which is why it represents the same data faithfully.
Check that claim on the example. Between class marks and the line falls from to , crossing the class boundary at a height of . The triangle it cuts off the -bar and the triangle it adds to the -bar are congruent — both have base and height — so the areas cancel. Equal class widths are what make that cancellation exact, and it is the reason a frequency polygon should not be drawn for classes of unequal width.
The practical advantage over a histogram. Two polygons can be drawn on the same pair of axes and compared directly — two sections' marks, or the same test before and after revision — because they are lines rather than solid bars. A histogram cannot be overlaid on another histogram without hiding one of them, and that is the main reason this diagram exists.
Exam tip
What layout mistakes cost marks in a statistics question?
Write the class size, state the convention you are using, and total the frequency column. The marks in this chapter sit in the table's structure rather than in any calculation.
- State the range and the class size as separate steps: *range ; with a class size of we need classes.* This shows the choice was reasoned, not guessed
- Say which convention you are using — exclusive (-, upper limit in the next class) or inclusive (-). Mixing them in one table loses several marks
- Show the tally column, in groups of five. It is the evidence that you passed through the data once and it is separately creditable
- **Total the frequencies and compare with . Write the total in the table; an unchecked table is usually a wrong table
- Convert to boundaries before drawing anything**, and write the adjustment factor explicitly: *gap , so adjustment *
- Label both axes with what they measure, and mark the scale. For a polygon, the horizontal axis carries class marks, not class limits
- Add the two zero-frequency classes and show them on the graph. A polygon that does not touch the axis at both ends is incomplete
- Use a kink or a break symbol on the axis if your scale does not start at zero, so the reader is not misled
The confusion to clear up. A histogram has bars whose widths are class sizes and whose areas represent frequency; a bar graph has separated bars for categories and means something different. Bars that touch say the variable is continuous, and drawing gaps in a histogram claims something about the data that is not true.
- State the range and the class size as separate steps: *range ; with a class size of we need classes.* This shows the choice was reasoned, not guessed
- Say which convention you are using — exclusive (-, upper limit in the next class) or inclusive (-). Mixing them in one table loses several marks
- Show the tally column, in groups of five. It is the evidence that you passed through the data once and it is separately creditable
- **Total the frequencies and compare with . Write the total in the table; an unchecked table is usually a wrong table
- Convert to boundaries before drawing anything**, and write the adjustment factor explicitly: *gap , so adjustment *
- Label both axes with what they measure, and mark the scale. For a polygon, the horizontal axis carries class marks, not class limits
- Add the two zero-frequency classes and show them on the graph. A polygon that does not touch the axis at both ends is incomplete
- Use a kink or a break symbol on the axis if your scale does not start at zero, so the reader is not misled
The confusion to clear up. A histogram has bars whose widths are class sizes and whose areas represent frequency; a bar graph has separated bars for categories and means something different. Bars that touch say the variable is continuous, and drawing gaps in a histogram claims something about the data that is not true.
Did you know
What happens to a frequency polygon when the classes get narrower?
Take the same data and redraw the polygon with narrower classes — widths of instead of , then , then . Each time you get more points and the line becomes less jagged.
Keep going, with a large amount of data and very narrow classes, and the polygon stops looking like a polygon at all. It settles into a smooth curve — a frequency curve — and for many natural measurements that curve takes on a particular shape: a single hump, roughly symmetric, tailing away on both sides.
That shape turns up in the heights of students in a large school, in the masses of packets filled by the same machine, and in the errors of repeated measurements of the same quantity. It has a name you will meet in Class 11 and 12: the normal distribution.
The reason the smoothing works is the area property from the previous section. Because the polygon preserves the histogram's total area, and the histogram's area is the total frequency, the curve you end up with also has a fixed total area — and that is what makes it possible to treat areas under the curve as proportions of the data. In Class 12 those areas become probabilities.
You can see the beginning of it in your own table. The five frequencies already rise to a peak and fall away. With ten times as much data and half the class width, the rise and fall would be smoother but the shape would be recognisably the same.
And that is why the humble class mark matters. Grouping throws away the individual values, but it keeps the shape — and the shape is what lets you say something about students, packets or measurements you have not seen. A frequency polygon is the first step from describing a set of numbers to predicting the next one.
Keep going, with a large amount of data and very narrow classes, and the polygon stops looking like a polygon at all. It settles into a smooth curve — a frequency curve — and for many natural measurements that curve takes on a particular shape: a single hump, roughly symmetric, tailing away on both sides.
That shape turns up in the heights of students in a large school, in the masses of packets filled by the same machine, and in the errors of repeated measurements of the same quantity. It has a name you will meet in Class 11 and 12: the normal distribution.
The reason the smoothing works is the area property from the previous section. Because the polygon preserves the histogram's total area, and the histogram's area is the total frequency, the curve you end up with also has a fixed total area — and that is what makes it possible to treat areas under the curve as proportions of the data. In Class 12 those areas become probabilities.
You can see the beginning of it in your own table. The five frequencies already rise to a peak and fall away. With ten times as much data and half the class width, the rise and fall would be smoother but the shape would be recognisably the same.
And that is why the humble class mark matters. Grouping throws away the individual values, but it keeps the shape — and the shape is what lets you say something about students, packets or measurements you have not seen. A frequency polygon is the first step from describing a set of numbers to predicting the next one.
Exam relevance
How does this statistics chapter feed into JEE and NEET preparation?
This is foundation work, and it is the chapter whose conventions cause the most avoidable errors later.
Where it leads. In Class 10 the same grouped tables are used to compute the mean, median and mode of grouped data and to draw ogives — and every one of those formulas needs class boundaries and class marks, exactly as defined here. In Class 11 Statistics the same table carries through to variance and standard deviation, which is a JEE Main topic, usually set as a direct numerical from a frequency distribution.
Where it appears outside Mathematics. Class 11 and 12 Physics practical work uses frequency distributions of repeated readings to discuss random error, and Biology for NEET uses them in variation and inheritance data. In both, the skill being tested is reading the table correctly rather than any calculation.
Question types to expect. At this level: build a table, convert to continuous classes, find class marks, draw the polygon. In competitive papers: a mean, variance or standard deviation computed from a grouped table, and assertion-reason items about whether a variable is discrete or continuous.
The single trap that costs marks. Using class limits where boundaries are needed. In a Class 11 standard-deviation question on an inclusive table like -, -, the class mark must come from -, and taking it from - shifts every value by the same amount. The mean shifts with it and the whole answer is wrong while every line of working looks correct. Convert first, always.
A second trap worth knowing now. Class size must be constant for a frequency polygon and for the standard grouped-data formulas. A table with one wider class has to be handled separately, and in a histogram that class gets a shorter bar so that its area stays proportional to its frequency. In a histogram the area carries the frequency, not the height — a fact examined in assertion-reason form and forgotten by most candidates.
Board versus competitive emphasis. ICSE marks the table, the tally and the labelled diagram; a competitive paper marks one number from the same table. The conventions are the part that transfers, so getting them exactly right now saves marks for years.
Where it leads. In Class 10 the same grouped tables are used to compute the mean, median and mode of grouped data and to draw ogives — and every one of those formulas needs class boundaries and class marks, exactly as defined here. In Class 11 Statistics the same table carries through to variance and standard deviation, which is a JEE Main topic, usually set as a direct numerical from a frequency distribution.
Where it appears outside Mathematics. Class 11 and 12 Physics practical work uses frequency distributions of repeated readings to discuss random error, and Biology for NEET uses them in variation and inheritance data. In both, the skill being tested is reading the table correctly rather than any calculation.
Question types to expect. At this level: build a table, convert to continuous classes, find class marks, draw the polygon. In competitive papers: a mean, variance or standard deviation computed from a grouped table, and assertion-reason items about whether a variable is discrete or continuous.
The single trap that costs marks. Using class limits where boundaries are needed. In a Class 11 standard-deviation question on an inclusive table like -, -, the class mark must come from -, and taking it from - shifts every value by the same amount. The mean shifts with it and the whole answer is wrong while every line of working looks correct. Convert first, always.
A second trap worth knowing now. Class size must be constant for a frequency polygon and for the standard grouped-data formulas. A table with one wider class has to be handled separately, and in a histogram that class gets a shorter bar so that its area stays proportional to its frequency. In a histogram the area carries the frequency, not the height — a fact examined in assertion-reason form and forgotten by most candidates.
Board versus competitive emphasis. ICSE marks the table, the tally and the labelled diagram; a competitive paper marks one number from the same table. The conventions are the part that transfers, so getting them exactly right now saves marks for years.
Key takeaways
What should you be able to do with data before the mean and median?
This chapter is about organising data correctly, so that every later calculation starts from a sound table.
- Raw data is as collected, arrayed is sorted, grouped is summarised into classes with frequencies
- Range largest smallest, and number of classes range class size, rounded up
- Discrete variables take separated values (marks, runs, members in a household); continuous variables take any value in a range (height, mass, time)
- Exclusive classes put the upper limit in the next class; inclusive classes include both ends — pick one convention and state it
- Tally in groups of five, passing through the raw data once, then total the frequencies and check against
- Adjustment factor half the gap between consecutive classes; subtract from lower limits and add to upper limits to get class boundaries
- Class size upper boundary lower boundary, and class mark the average of either pair of ends
- A frequency polygon plots class marks against frequencies and is closed with a zero-frequency class at each end, which keeps its area equal to the histogram's
The quickest way to check this has landed is to take the inclusive table -, -, -, write down the boundaries, the class size and the class marks from memory — and then confirm that the class size is rather than .
- Raw data is as collected, arrayed is sorted, grouped is summarised into classes with frequencies
- Range largest smallest, and number of classes range class size, rounded up
- Discrete variables take separated values (marks, runs, members in a household); continuous variables take any value in a range (height, mass, time)
- Exclusive classes put the upper limit in the next class; inclusive classes include both ends — pick one convention and state it
- Tally in groups of five, passing through the raw data once, then total the frequencies and check against
- Adjustment factor half the gap between consecutive classes; subtract from lower limits and add to upper limits to get class boundaries
- Class size upper boundary lower boundary, and class mark the average of either pair of ends
- A frequency polygon plots class marks against frequencies and is closed with a zero-frequency class at each end, which keeps its area equal to the histogram's
The quickest way to check this has landed is to take the inclusive table -, -, -, write down the boundaries, the class size and the class marks from memory — and then confirm that the class size is rather than .