Free Mathematics Class 8 ICSE notes · practise this chapter with an AI quiz

← All study notes

Three Averages, One Data Set, Three Different Answers

Learn the difference between raw and arrayed data, find the range, build tally and grouped frequency tables with class marks, and calculate the mean, median and mode.

Which average should you actually use?

It depends on the question — and the three averages can disagree sharply on the same data.

Take the marks of students out of :



The mean works out to , the median to and the mode to . The mean is nearly two marks higher than the other two, and the reason is the handful of high scores at and pulling it upwards.

So the average mark was 16.8 and the average mark was 15 are both true statements about this class, and they would leave a reader with different impressions. Knowing which average a figure refers to matters as much as being able to calculate it.

This page covers the first part of the ICSE Class 8 Mathematics chapter on data handling, and works with that one data set throughout so the ideas can be compared directly.

What is the difference between raw data and arrayed data?

Raw data is recorded in the order it was collected; arrayed data has been sorted. Sorting changes nothing about the information and everything about how easily you can read it.

The marks above are raw — that is the order the papers came off the pile. Sorted into ascending order they become arrayed:



Now several things are visible at a glance: the lowest mark is , the highest is , and occurs more often than anything else. None of that was apparent in the raw list.

The range is the simplest measure of spread:



The range is a single number, not a pair. Writing *the range is to * loses the mark; the range is . And note the range has the same unit as the data — here marks — because it is a difference, not a count.

What the range hides. It uses only two values and ignores every other observation, so two very different classes can share a range. A class with marks and one with both have a range of , though the first is clustered at the bottom and the second is spread out. The range is a quick check on spread, not a description of it — which is why the averages, and later the graphs, are needed as well.

Sort before doing anything else. The median cannot be found from raw data, and the mode is easy to miscount in an unsorted list. Arraying the data first is the step that makes the rest of this page routine, and it costs one line.

How do you make a frequency table with tally marks?

Write each distinct value once, then walk the raw list marking a tally against the value you meet.

Tally marks are grouped in fives — four upright strokes with the fifth drawn diagonally across them — because a human eye can count groups of five far more reliably than a row of nineteen strokes.

Worked example — the marks data. Going through the raw list once:

- — frequency
- — frequency
- — frequency
- — frequency
- — frequency
- — frequency

Always total the frequency column.



which matches the students. A frequency total that does not match the number of observations means a value was missed or double-counted, and that check is the entire reason tally marks exist — you are tallying precisely so you can prove nothing was lost.

Walk the raw list, not the sorted one. The point of a tally is that you can build the table directly from the data as recorded, in one pass, without sorting first. Going back and forth hunting for all the s is exactly what causes the miscount.

The table has not lost any information. From the frequency table you could write the whole data set back out — five s, six s and so on. That is worth noticing, because the grouped table in the next section genuinely does lose information, and the difference between the two matters.

When to use an ungrouped table. This works well because the data has only six distinct values. If the students had scored different marks, a frequency table with rows each of frequency would tell you nothing at all — and that is the situation grouping was invented for.

How do you build a grouped frequency distribution?

Choose class intervals of equal width covering all the data, then count how many observations fall in each.

Worked example — the same marks, grouped in fives. Using the exclusive form, where the upper limit belongs to the next class:

- : the five s, so frequency
- : six s and three s, so frequency
- : three s and one , so frequency
- : two s, so frequency



In the exclusive form, a value equal to a class boundary goes into the upper class. So the s went into , not , and the s into . Deciding that rule once and applying it consistently is the whole of the difficulty — a value placed in the wrong class shifts two frequencies at once.

**The vocabulary, applied to the class :

-
Lower class limit**:
- Upper class limit:
- Class size (or width):
- Class mark (the midpoint):

The class marks of the four classes are , , and — each apart, as they must be when the classes have equal width.

Worked example — the inclusive form. The same data can be grouped as , , , , where both limits belong to the class:

- : frequency
- : frequency
- : frequency
- : frequency

The frequencies are identical, but the arithmetic of the class details is not:



**In the inclusive form the class size needs the **, because is five values. Using is the standard error, and it also throws off the class mark.

Choosing the class size. Aim for somewhere between five and ten classes: too few and the shape of the data is flattened away, too many and the table is as unhelpful as the raw list. With a range of , classes of width gave four classes, which is reasonable for observations.

Grouping loses information, deliberately. Once you know that nine students scored between and , you can no longer recover that six scored exactly . That is the price paid for seeing the shape of the distribution, and it is why calculations from grouped data use the class mark as a stand-in for every value in the class.

How do you calculate the mean, median and mode?

The mean is the total shared out, the median is the middle value of sorted data, and the mode is the most frequent value.



Worked example 1 — the mean of the marks data. Using the frequency table rather than adding twenty numbers one at a time:







Multiplying each value by its frequency is the same as adding them all, and far less error-prone — which is a second reason to build the frequency table first.

Worked example 2 — the median. With observations the middle falls between the th and th values in the arrayed list. Both are , so



For an even number of observations the median is the average of the two middle values; for an odd number it is the single middle value. With observations the median position is , so gives position — halfway between the tenth and eleventh, exactly as found.

Worked example 3 — the mode. The highest frequency in the table is , against the value , so



The mode is the value, not the frequency. Answering is the classic slip — is how often the mode occurs.

Worked example 4 — a smaller set, all three averages. For :



Worked example 5 — an even count. For :



There is no mode, since every value appears once. A data set may have no mode, one mode, or several — unlike the mean and median, which always exist and are always unique.

Worked example 6 — a missing observation. The mean of five numbers is , and four of them are . Find the fifth.

If the mean is , the total must be :



Check: . Correct.

Sort before finding the median. Taking the middle of the raw list gives here by luck — the th and th raw values happen to be and — but that is a coincidence, and on almost any other data set it would be wrong. The median is defined on arrayed data, and the sorting step is not optional.

Why the mean sits above the other two here. The mean uses every value, so the two scores of and the one of each pull it upward. The median only cares about position, so an extreme value moves it hardly at all. That is exactly when the median is the more honest summary — and why it is preferred for figures like household income, where a few very large values would otherwise dominate.
Exam tip

Exam tip: sort first, total the frequencies, and state the value not the count

Array the data before anything else. The median is defined on sorted data, and the mode is easy to miscount otherwise.

Range is one number: , not * to *.

Total your frequency column and match it to the number of observations — that is what tally marks are for. Build the table in one pass through the raw list.

In the exclusive form (, ), a value on a boundary goes into the upper class. In the inclusive form (, ), class size is — **do not forget the .

Class mark is the midpoint**: in exclusive form, in inclusive. Aim for five to ten classes.

Find the mean with : . Far safer than adding twenty numbers.

The median position is , so gives the average of the th and th values.

The mode is the value, not its frequency, not . And a set may have no mode or several.

For a missing observation, multiply the mean by the count to get the total, then subtract.

And state the units on every answer.
Did you know

Why the middle value is sometimes the fairer summary

Imagine nine people in a small shop, each earning ₹ a month. The mean monthly earning is ₹, the median is ₹, and both describe the room accurately.

Now one more person walks in, earning ₹ a month. The mean leaps to



while the median moves to the average of the fifth and sixth values, which are still ₹ each — so the median stays at ₹.

The mean now says the average person in the room earns ₹ a month, which describes nobody in it. Nine of the ten earn far less and one earns far more. The median's ₹ describes nine of the ten exactly.

Neither figure is wrong. The mean is the correct answer to if the total were shared equally, what would each get? The median answers what is a typical person here? — and when a few values are extreme, those two questions have very different answers.

This is why figures about income, house prices and land holdings are usually quoted as medians, while figures about rainfall or examination marks are usually means. It is also the reason the marks data in this page had a mean of and a median of : a few high scores, and the two averages parted company. Asking which average a number is is often more useful than asking how it was calculated.
Key takeaways

Data, frequency tables and averages: quick revision

- Raw data is in collection order; arrayed data is sorted. Always array first.
- Range highest lowest — a single number with the data's unit. It uses only two values, so it hides the shape.
- Tally marks in groups of five let you build a frequency table in one pass through the raw data.
- For the marks data: occurs times, six times, three, three, once, twice — totalling . Always total the frequency column.
- An ungrouped table loses no information; a grouped one does, deliberately.
- Exclusive form , , , gives frequencies . A boundary value goes to the upper class.
- For the class : lower limit , upper limit , class size , class mark . The class marks are .
- Inclusive form , , , gives the same frequencies, but class size is and the class mark is . **The matters.
- Aim for
five to ten classes of equal width.
-
Mean** . Here , from .
- Median position is . With that is , so average the th and th values: .
- Mode is the most frequent value, not its frequency of .
- gives mean , median , mode . And gives mean , median , and no mode.
- A set may have no mode or several; the mean and median are always unique.
- Missing observation: mean over five numbers means a total of , so the fifth is .
- The mean is pulled by extreme values and the median is not — which is why the marks data gave against , and why incomes are usually reported as medians.

Take the marks data, add one score of , and recalculate all three averages — watching which of them move and by how much is the fastest way to understand what each one measures.

Ready to put this into practice?

Create a personalized quiz on this exact topic — free to start.

Create your own quiz on Data Handling — Part 1Create a free account
← Back to all articles