One Unusual Value Can Ruin an Average
Learn to tell a statistical question from an ordinary one, organise data into a frequency table, calculate the mean and the median, and choose the value that represents the data honestly.
Can one unusual value make an average misleading?
Easily. If four friends get ₹100 to ₹130 pocket money a month and a fifth gets ₹900, the mean climbs to ₹270 — a figure none of the five actually receives.
The median stays at ₹120 and describes the group far better. This page covers everything in the CBSE Class 7 Mathematics chapter's first half: recognising a statistical question, collecting and organising data, calculating the mean and the median, and deciding which one represents the data honestly.
The median stays at ₹120 and describes the group far better. This page covers everything in the CBSE Class 7 Mathematics chapter's first half: recognising a statistical question, collecting and organising data, calculating the mean and the median, and deciding which one represents the data honestly.
What makes a question a statistical question?
A statistical question is one whose answer varies across the data you collect, so you need many responses rather than one.
"How tall is Meera?" has a single answer, so it is not statistical. "How tall are the students in my class?" produces a different answer for each student, so it is. The variation is the whole point.
- Statistical: How many hours a day do students in Class 7 read?
- Not statistical: How many students are in Class 7?
- Not a question at all: Class 7 students read a lot. That is a statement, and statements need data to support them.
A school canteen survey shows why it matters. Asking "what does the canteen sell?" gets one fixed answer, while "what do students buy most often?" gets answers that differ from person to person — and only the second can be analysed.
So before collecting anything, check that your question will produce varied answers. A question with one fixed answer needs a fact, not a survey, and no amount of data collection will make it statistical.
"How tall is Meera?" has a single answer, so it is not statistical. "How tall are the students in my class?" produces a different answer for each student, so it is. The variation is the whole point.
- Statistical: How many hours a day do students in Class 7 read?
- Not statistical: How many students are in Class 7?
- Not a question at all: Class 7 students read a lot. That is a statement, and statements need data to support them.
A school canteen survey shows why it matters. Asking "what does the canteen sell?" gets one fixed answer, while "what do students buy most often?" gets answers that differ from person to person — and only the second can be analysed.
So before collecting anything, check that your question will produce varied answers. A question with one fixed answer needs a fact, not a survey, and no amount of data collection will make it statistical.
How do you organise collected data into a frequency table?
Group identical values together and count how many times each appears — that count is its frequency.
Suppose 20 students report the number of glasses of water they drink in a day: 4, 6, 5, 4, 7, 5, 6, 4, 5, 8, 6, 5, 4, 6, 7, 5, 4, 6, 5, 6.
Work through the list once, putting a tally mark against each value as you meet it, then total the tallies:
- 4 glasses — frequency 5
- 5 glasses — frequency 6
- 6 glasses — frequency 6
- 7 glasses — frequency 2
- 8 glasses — frequency 1
The frequencies must add to the number of observations: . That check catches a missed or double-counted entry straight away, and is worth doing every time.
The organised table already answers questions the raw list hid — most students drink 5 or 6 glasses, and only one drinks 8.
The habit to build is tallying in a single pass through the data. Scanning the list again for each separate value is slow and is where values get counted twice.
Suppose 20 students report the number of glasses of water they drink in a day: 4, 6, 5, 4, 7, 5, 6, 4, 5, 8, 6, 5, 4, 6, 7, 5, 4, 6, 5, 6.
Work through the list once, putting a tally mark against each value as you meet it, then total the tallies:
- 4 glasses — frequency 5
- 5 glasses — frequency 6
- 6 glasses — frequency 6
- 7 glasses — frequency 2
- 8 glasses — frequency 1
The frequencies must add to the number of observations: . That check catches a missed or double-counted entry straight away, and is worth doing every time.
The organised table already answers questions the raw list hid — most students drink 5 or 6 glasses, and only one drinks 8.
The habit to build is tallying in a single pass through the data. Scanning the list again for each separate value is slow and is where values get counted twice.
Formula
How do you calculate the mean and what does it represent?
The mean, or arithmetic average, shares the total equally among all the members:
Worked example. Five friends receive monthly pocket money of ₹100, ₹100, ₹120, ₹130 and ₹150.
The fair share reading is what makes the mean meaningful: if the five pooled all their money and split it equally, each would get ₹120.
Run it backwards when a question gives you the mean. If 6 students scored a mean of 15 marks, the total must be marks — useful whenever a new value is added to a set.
Two features follow from the definition. The mean always lies between the smallest and largest observations, so a mean of 160 for the list above would be impossible. And the mean need not be one of the data values — a mean of 5.6 glasses of water is perfectly correct even though nobody drank 5.6 glasses.
Worked example. Five friends receive monthly pocket money of ₹100, ₹100, ₹120, ₹130 and ₹150.
The fair share reading is what makes the mean meaningful: if the five pooled all their money and split it equally, each would get ₹120.
Run it backwards when a question gives you the mean. If 6 students scored a mean of 15 marks, the total must be marks — useful whenever a new value is added to a set.
Two features follow from the definition. The mean always lies between the smallest and largest observations, so a mean of 160 for the list above would be impossible. And the mean need not be one of the data values — a mean of 5.6 glasses of water is perfectly correct even though nobody drank 5.6 glasses.
How does an outlier affect the mean more than the median?
The median is the middle value once the data is arranged in order, and because it depends only on position, an extreme value barely moves it.
Arrange the pocket money in order: 100, 100, 120, 130, 150. The middle value is the third, so the median is ₹120 — the same as the mean here.
Now replace ₹150 with ₹900, an outlier far from the rest:
The median is still the third value, so it stays at ₹120. The mean more than doubled; the median did not budge.
The mean uses every value's size, so one huge number drags it upward. The median uses only the order, so it ignores how extreme the extreme is.
For an even number of observations there is no single middle, so take the mean of the two middle values: for 12, 15, 18, 20 the median is .
This is why house prices and incomes are usually reported as medians. When a few values are far above the rest, the median is the value that describes a typical member.
Arrange the pocket money in order: 100, 100, 120, 130, 150. The middle value is the third, so the median is ₹120 — the same as the mean here.
Now replace ₹150 with ₹900, an outlier far from the rest:
The median is still the third value, so it stays at ₹120. The mean more than doubled; the median did not budge.
The mean uses every value's size, so one huge number drags it upward. The median uses only the order, so it ignores how extreme the extreme is.
For an even number of observations there is no single middle, so take the mean of the two middle values: for 12, 15, 18, 20 the median is .
This is why house prices and incomes are usually reported as medians. When a few values are far above the rest, the median is the value that describes a typical member.
Exam tip
Exam tip: ordering the data before you find the median
The commonest error in this chapter is reading the middle of the unordered list. Arranging the data in increasing order is step one of every median question, and skipping it makes the answer wrong even when the method is understood.
Write the ordered list out, then count in from both ends together until you meet. With an even count you will land on two values, and the median is their mean.
For the mean, show the fraction with the sum on top and the count below before dividing, so the method earns marks even if the division slips. And count the observations carefully — dividing 600 by 4 instead of 5 is a silent error the answer alone will not reveal.
When a question asks which average is "more suitable", name the outlier explicitly and say the median is less affected by it. That comparison is what the mark is for.
Write the ordered list out, then count in from both ends together until you meet. With an even count you will land on two values, and the median is their mean.
For the mean, show the fraction with the sum on top and the count below before dividing, so the method earns marks even if the division slips. And count the observations carefully — dividing 600 by 4 instead of 5 is a silent error the answer alone will not reveal.
When a question asks which average is "more suitable", name the outlier explicitly and say the median is less affected by it. That comparison is what the mark is for.
Did you know
Why does the median ignore how extreme an outlier is?
Because it only asks where a value sits in the order, never how far away it is.
Change the largest value in a list from ₹150 to ₹900 or to ₹9000 and the median is unchanged either time — the number is still simply the last one in the queue.
The mean, by contrast, adds every value in full, so it feels the whole distance. Two averages built from the same data can therefore tell quite different stories, which is why a report that quotes only one of them is worth questioning.
Change the largest value in a list from ₹150 to ₹900 or to ₹9000 and the median is unchanged either time — the number is still simply the last one in the queue.
The mean, by contrast, adds every value in full, so it feels the whole distance. Two averages built from the same data can therefore tell quite different stories, which is why a report that quotes only one of them is worth questioning.
Key takeaways
Mean, median and data handling: quick revision
- A statistical question has answers that vary across the data collected; a question with one fixed answer is not statistical.
- Build a frequency table by tallying in a single pass, then check that the frequencies add to the number of observations.
- , and it represents the fair share if the total were split equally.
- The mean lies between the smallest and largest values but need not be one of them; total when working backwards.
- The median is the middle value of the ordered data, or the mean of the two middle values when the count is even.
- An outlier shifts the mean sharply and leaves the median almost untouched, so the median is the better representative for skewed data.
You will remember all of this far better after answering five questions on it than after reading it twice.
- Build a frequency table by tallying in a single pass, then check that the frequencies add to the number of observations.
- , and it represents the fair share if the total were split equally.
- The mean lies between the smallest and largest values but need not be one of them; total when working backwards.
- The median is the middle value of the ordered data, or the mean of the two middle values when the count is even.
- An outlier shifts the mean sharply and leaves the median almost untouched, so the median is the better representative for skewed data.
You will remember all of this far better after answering five questions on it than after reading it twice.