Free Mathematics Class 10 CBSE notes · practise this chapter with an AI quiz

← All study notes

The Middle Value and the Most Common Value Can Sit in Different Places

Build a cumulative frequency table, locate the median class and compute the median with its formula, identify the modal class and compute the mode, recover a missing frequency from a given median, and decide which average describes a real data set best.

Why does one data set need three different kinds of average?

Suppose a small workshop pays nine workers ₹ a month each and one owner ₹. **The mean salary is ₹ — a figure that describes nobody in the building. Nine of the ten people earn less than a third of it.

The median, the middle value when the salaries are lined up, is ₹. So is the mode, the most frequent value. Those two describe the workshop honestly and the mean does not, and the reason is that the mean is dragged upward by one extreme value while the other two are not.

So the three averages answer three different questions:

-
The mean — if the total were shared equally, how much would each get?
-
The median — what is the value in the middle, with half the data on either side?
-
The mode — which value occurs most often?

None of them is the right answer in general. The mean uses every observation and is best when the data has no extreme outliers. The median ignores how far out the extremes are and so survives them. The mode is the only one that makes sense for data that is not numerical at all, such as the most common shoe size sold.

Part 1 dealt with the mean. This part gives you the other two for grouped data, and both need a new idea first: the cumulative frequency, which tells you how many observations fall below a given point. You cannot find a median without it, because the median is defined by a position rather than by a value.

And then one relation ties all three together approximately:**



which is useful as a rough check and as a way of estimating the third when two are known.

This page covers the second part of the CBSE Class 10 Maths chapter on statistics: cumulative frequency, the median and mode of grouped data, missing frequencies, and comparing the three averages.

How do you build a cumulative frequency table?

Add each frequency to the running total of all the frequencies before it. The result tells you how many observations are less than the upper limit of each class.

Worked example — the distribution used throughout this page. A test was taken by students and their marks recorded as follows.

- ****: frequency
- ****: frequency
- ****: frequency
- ****: frequency
- ****: frequency

The less-than cumulative frequencies, each obtained by adding the next frequency to the previous total:

- **Less than **:
- **Less than **:
- **Less than **:
- **Less than **:
- **Less than **:

The final cumulative frequency must equal the total number of observations, and here it does: . That is the compulsory check on any cumulative column — if the last entry is not , a frequency has been added wrongly or skipped.

Reading the table. "Less than is " means ** of the students scored under marks.** It says nothing about how those are distributed inside the classes below ; it only counts them.

There is also a more-than type, built by adding from the bottom up:

- **More than or equal to **:
- **More than or equal to **:
- **More than or equal to **:
- **More than or equal to **:
- **More than or equal to **:

Check that the two types agree. At any boundary the two must add to : less than is , more than or equal to is , and . Correct.

Why the median needs this table and the mean did not. The mean was a total divided by a count, so the order of the data was irrelevant. The median is defined by position — the value with half the data below it — so you must be able to say how many observations lie below any point. That is exactly what a cumulative frequency is.

One boundary case about the two types. A question may give you a cumulative table and ask you to recover the ordinary frequencies. Subtract each entry from the next one: from , , , , you get , , , , — the original frequencies. Converting in both directions is examinable, and the give-away that a table is cumulative is that its entries never decrease.
Formula

What is the median formula for grouped data and how do you use it?

Find the class in which the middle observation falls, then interpolate inside it.



where

- ** is the lower limit of the median class
-
is the total number of observations
-
is the cumulative frequency of the class immediately before the median class
-
is the frequency of the median class itself
-
is the class width

How to find the median class.** Compute and look down the cumulative frequency column for the first entry that is greater than or equal to it. That class is the median class.

Worked example. Find the median of the distribution of marks tabulated above.

Step 1 — the position.



Step 2 — the median class. The cumulative frequencies are , , , , . The first one to reach is , belonging to the class **.

Step 3 — read off the four quantities.**

- , the lower limit of
- , the cumulative frequency of the previous class
- , the frequency of itself
- , the class width

Step 4 — substitute.



Check that the answer lies inside the median class. It must, by construction, and does lie between and . An answer outside the median class means one of the four quantities was misread — most often , by taking the median class's own cumulative frequency instead of the previous one.

What the formula is actually doing. There are students below marks and we want the th, so we need more students. The class contains students spread across marks, so ** students' worth of it is of the way across**, which is marks. **Starting at and moving gives . The formula is nothing more than that sentence written in symbols, and reconstructing it that way is safer than remembering it.

Two quantities that are easy to confuse.

-
is the cumulative frequency of the class before the median class**, not of the median class. Here it is , not
- ** is the plain frequency of the median class**, not a cumulative one. Here it is , not

**Using for would give** , which falls outside the median class — and that is exactly why the plausibility check catches this particular error every time.

**One note on .** For grouped data the position used is whether is odd or even. There is no averaging of two middle values as there is for ungrouped data, because the data is continuous and the formula interpolates rather than picking an observation.

How do you find the mode of grouped data?

Find the class with the highest frequency, then interpolate inside it using how much taller it is than its two neighbours.



where

- ** is the lower limit of the modal class
-
is the frequency of the modal class
-
is the frequency of the class before it
-
is the frequency of the class after it
-
is the class width

Worked example.** Find the mode of the same distribution of marks.

Step 1 — the modal class. The frequencies are , , , , , and the largest is . **The modal class is .

Step 2 — read off the quantities.** , , , , .

Step 3 — substitute.




Check that the mode lies inside the modal class: is between and . Correct.

What the formula is doing. The modal class is taller than the class before it by and taller than the class after it by . The mode is pulled toward the taller neighbour, and the fraction is the share of the class width you move across. Written that way the denominator makes sense, because



So the denominator is the sum of the two differences, and computing it that way is much less error-prone than computing directly.

A useful consequence to sanity-check with. If the two neighbours have equal frequencies, the two differences are equal, the fraction is , and the mode sits exactly at the middle of the modal class. The mode leans toward whichever neighbour is taller, and if it comes out on the wrong side, and have been swapped.

Worked example 2 — the modal class at an end. Suppose the frequencies were , , , , instead, so the modal class is the first one, .

There is no class before it, so :



Taking the missing neighbour's frequency as zero is the convention, and it is what a question with a modal first or last class expects.

When the mode is not well defined. If two classes share the highest frequency, the distribution has two modal classes and the formula cannot be applied to either without ambiguity. Such data is called bimodal, and a question will not ask for a single mode from it — but recognising the situation is worth a mark if asked.

Putting all three averages together for this data. The mean of the same marks, by the direct method with class marks , , , , :



**So the three averages are: mean , median , mode — all close together, which tells you the distribution is fairly symmetric and has no extreme outliers. Compare that with the workshop salaries of the opening section**, where the mean was ₹ and the median ₹: a large gap between the mean and the median is the signature of a skewed data set.

And the empirical relation, checked on this data.




Close but not exactly equal, which is why the relation is written with an approximation sign. It is a rule of thumb for estimating a third average from two, not an identity, and quoting it as an exact equation is a marked error.

How do you find missing frequencies when the median is given?

Use the total frequency to get one equation, and the median formula to get another. Two unknowns need two conditions, exactly as in the mean chapter.

Worked example. The median of the following distribution is and the total frequency is . Find and .

- ****:
- ****:
- ****:
- ****:
- ****:
- ****:
- ****:
- ****:
- ****:
- ****:

Step 1 — the total frequency gives the first equation.




Step 2 — locate the median class from the given median. The median is , which lies in the class ****. So , , , and .

Step 3 — the cumulative frequency before the median class, in terms of :



Step 4 — substitute into the median formula.





Step 5 — the second unknown, from the first equation:



Check by recomputing the median. With the cumulative frequency before the median class is , so



Exactly the given median, and the frequencies now total . Both conditions are satisfied.

**Also confirm that really is the median class with these values.** The cumulative frequencies become , , , , , , , , , . The first to reach is , belonging to . Consistent — and that check matters, because a value of that shifted the median class elsewhere would have invalidated the whole calculation.

The structure of every question in this family.

- One unknown frequency — the median or the mean alone is enough
- Two unknown frequencies — you need the total frequency as well, and the total gives the easier equation, so use it first
- Always identify the median class from the given median, not by guessing, since the median lies inside its own class by construction

Worked example 2 — one unknown, from the mode instead. Suppose the frequencies of , , , are , , , , the modal class is , and the mode is . Find .

With , , and :




Trying gives . **So **, which is the distribution of the previous section — as it should be, since that is where the mode came from.

**Note that a mode equation is not linear in **, because appears in the denominator too. That is why questions prefer the median or the mean for missing-frequency problems, and why a mode version usually supplies enough information to check a candidate value rather than to solve from scratch.
Exam tip

What does a full-mark median or mode answer contain?

Show the cumulative frequency column, name the median or modal class explicitly, list the four quantities with their values, and then substitute. Each of those is a separate mark.

- Build the cumulative column and check the last entry equals
- **Compute and state it before looking for the median class
-
Name the median class**: "the first cumulative frequency to reach is , so the median class is "
- **Write , , and as a list**, and take from the class before the median class
- For the mode, name the modal class as the one with the highest frequency, and list , , , and
- **Compute the mode's denominator as ** rather than as
- Take a missing neighbour's frequency as zero when the modal class is first or last
- Check the answer lies inside its own class — the median inside the median class, the mode inside the modal class
- Use the total frequency first when two frequencies are unknown
- Re-verify a found frequency by recomputing the median and confirming the median class has not moved
- Quote the empirical relation with an approximation sign, never as an equality

The misconception to name. The median class and the modal class are not always the same class, and when they coincide it is a property of the data rather than a rule. The median class is found from the cumulative frequencies; the modal class from the plain frequencies — and a distribution with a long tail can easily have them in different places. Using one class for both calculations without checking is a real error even when it happens to give the right answer here.

A second trap. Taking as the median class's own cumulative frequency. It must be the previous class's, and the error is self-announcing: it makes negative, so the median comes out below the lower limit of its own class. In this chapter's data it gives for a median class of , which is impossible.
Did you know

Why does one very large salary move the mean but not the median?

Go back to the workshop: nine workers on ₹ and one owner on ₹. **Now suppose the owner's pay doubles to ₹.

The mean jumps** from ₹ to ₹, because the total has risen by ₹ and been shared among ten people.

The median does not move at all. It is still ₹, because the middle of the list is still a worker's salary — and it would stay ₹ if the owner earned ten crore.

That difference has a clean explanation. The mean uses the value of every observation, so a value far from the rest pulls it. The median uses only the order, so moving the largest observation further out changes nothing except how far out it is. The median is said to be resistant to outliers, and the mean is not.

Which is why the choice of average is a real decision and not a matter of taste.

- To describe a typical income, rent or house price, the median is used, because a few very large values would otherwise distort the picture
- To compute a total from an average, the mean is the only one that works — the mean times the count gives the total, and the median does not
- To decide which size of shirt to stock most of, the mode is the only sensible answer, since you cannot stock a mean size

The mode has one advantage the others lack entirely. It applies to data that is not numbers at all. The most common blood group in a sample, the most common mode of travel to school, the most frequently sold flavour — none of these has a mean or a median, because you cannot add or order them meaningfully. The mode is defined for any kind of data.

And a small warning about the mode of grouped data. The formula returns a value inside the modal class, but the modal class itself depends on how you chose the classes. Regroup the same observations into classes of width instead of and the modal class may change, and the mode with it. The mean and the median are far less sensitive to the grouping — which is a genuine limitation of the mode and a good answer to "which average is least reliable for grouped data?"

One last observation about the empirical relation. holds well for moderately skewed data and poorly for badly skewed data. Check it on the workshop: the mean is ₹, the median and mode are both ₹, so the left side is ₹ and the right side is ₹. Nowhere near. The relation is a rule of thumb whose failure is itself informative — when it fails badly, the data has extreme outliers, which is exactly the case where the mean should not have been used to describe the data in the first place.
Exam relevance

How are median and mode used in JEE and NEET?

This is foundation work for Class 11 Statistics, a JEE Main topic, and the data-description ideas here are used directly in NEET Biology and in laboratory work.

Where the median leads. Class 11 uses it to define the mean deviation about the median, and proves the useful fact that the mean deviation is least when taken about the median rather than about any other point. The positional definition you learn here is what makes that true, and JEE Main asks for mean deviation about the median as a direct numerical.

Where the cumulative frequency leads. It becomes the basis of quartiles, deciles and percentiles, all located by the same interpolation formula with or in place of . The formula does not change at all — only the fraction of you look for in the cumulative column. Recognising that means the median formula is really one formula for a whole family.

Where the comparison of the three averages leads. Class 11 introduces the coefficient of variation to compare the consistency of two data sets, and the discussion of skewness begins exactly where this chapter ends: a gap between the mean and the median signals skewness. NEET questions on interpreting biological data rely on being able to say which average describes a sample fairly.

Where the outlier-resistance idea leads. It matters in every practical measurement. A single bad reading shifts the mean of a set of observations but not their median, which is why experimental work often reports both, and why a discordant reading must be identified rather than silently averaged in. That reasoning appears in JEE and NEET Physics error-analysis questions.

Where the mode leads. It is the average used for categorical data, which is the only kind available in much of Biology — blood groups, phenotypes, the most frequent allele. NEET Biology uses modal reasoning without naming it, and in genetics the most probable phenotype of a cross is a modal statement.

Question types to expect. At this level: cumulative frequency tables, median and mode calculations, missing frequencies from a given median, and choosing the appropriate average. In competitive papers: mean deviation about the mean and the median, variance and standard deviation, coefficient of variation, and interpretation questions.

The single trap that costs marks. Using the median class's own cumulative frequency as . It must be the previous class's, and the error shows up as a median lying below the median class — for a class of in this chapter's data. A plausibility check catches it in one second.

A second trap. Quoting as an exact identity. It is approximate, holding well only for moderately skewed data — on this chapter's distribution it gives against , and on the workshop salaries it fails completely. Write it with an approximation sign and say what it is for.

Board versus competitive emphasis. The CBSE paper marks the cumulative column, the named class, the listed quantities, the substitution and the conclusion; a competitive paper marks a single number, often a quartile or a standard deviation. The transferable habit is locating a value by its position in the cumulative column — because that one technique gives you the median, every quartile and every percentile, and it is the whole of the interpolation method you will use next year.
Key takeaways

What must you be able to do from this part?

One cumulative column, two formulas and one approximate relation.

- The mean shares the total equally, the median is the middle value, the mode is the most frequent
- The median resists outliers; the mean does not. A single very large value moves the mean and leaves the median where it was
- Cumulative frequency is the running total, and the last entry must equal
- **Less-than and more-than entries at the same boundary add to **, so
- Recover plain frequencies from a cumulative table by subtracting consecutive entries
- Median , with from the class before the median class
- Find the median class as the first class whose cumulative frequency reaches
- **For the -mark distribution**: , median class , , , , , giving a median of
- **Using would give , outside the median class — which is how that error announces itself
-
Mode** , and the denominator is
- Same distribution: modal class , , , , giving
- Take a missing neighbour as zero when the modal class is first or last
- Equal neighbours put the mode at the middle of the modal class; otherwise it leans toward the taller one
- **The mean of the same data is , so all three averages are close and the distribution is nearly symmetric
-
A large gap between the mean and the median signals a skewed data set
-
** — here against , so approximate and not exact
- Two missing frequencies need two conditions: a median of with a total of gives and
- Re-verify that the median class has not moved after substituting a found frequency
- The median class and the modal class need not be the same class
- The mode is the only average defined for non-numerical data, and the only one sensitive to how the classes were chosen

The sharpest self-test is one table, three averages. Take the -mark distribution, compute the mean, the median and the mode yourself, and then test the empirical relation — and decide, from how close the three come out, whether you would be comfortable describing this class's performance by the mean alone.

Ready to put this into practice?

Create a personalized quiz on this exact topic — free to start.

Create your own quiz on Statistics — Part 2Create a free account
← Back to all articles