Three Different Methods Must Give the Same Average, and One of Them Is Far Quicker
Find the mean of grouped data by the direct method using class marks, shift the numbers with an assumed mean, divide them down with the step-deviation method, decide which method the data suits, and work backwards to a missing frequency.
How do you average data when you no longer know the individual numbers?
Suppose a survey records the ages of people, but instead of the ages it reports only how many fall in each band: two people in the band –, three in –, seven in –, and so on. The individual ages are gone. You cannot add them up, because you do not know them.
So you do the next best thing: assume everyone in a band sits at its middle. The two people in – are treated as being each; the seven in – as each. That middle value is the class mark:
Then the ordinary averaging method works again, because you have a single number for each group and a count of how many share it.
The answer is an estimate, and it is worth being honest about that. Some people in a band are above its middle and some below; the assumption is that those cancel out across the band. For a large data set they very nearly do, which is why this is the standard method and not a compromise.
From there the chapter gives you three ways to do the arithmetic, and all three must give the same answer:
- The direct method — multiply each class mark by its frequency and divide
- The assumed mean method — shift all the numbers down by a convenient amount first
- The step-deviation method — shift them and then divide them down by the class width
They are not three different means. They are three different routes to one number, and the reason for having three is purely arithmetic: **the third one replaces multiplications like with multiplications like .** On a distribution with equal class widths it is dramatically faster, and that is why it exists.
This page covers the first part of the CBSE Class 10 Maths chapter on statistics: the mean of grouped data by all three methods, choosing between them, and finding a missing frequency.
So you do the next best thing: assume everyone in a band sits at its middle. The two people in – are treated as being each; the seven in – as each. That middle value is the class mark:
Then the ordinary averaging method works again, because you have a single number for each group and a count of how many share it.
The answer is an estimate, and it is worth being honest about that. Some people in a band are above its middle and some below; the assumption is that those cancel out across the band. For a large data set they very nearly do, which is why this is the standard method and not a compromise.
From there the chapter gives you three ways to do the arithmetic, and all three must give the same answer:
- The direct method — multiply each class mark by its frequency and divide
- The assumed mean method — shift all the numbers down by a convenient amount first
- The step-deviation method — shift them and then divide them down by the class width
They are not three different means. They are three different routes to one number, and the reason for having three is purely arithmetic: **the third one replaces multiplications like with multiplications like .** On a distribution with equal class widths it is dramatically faster, and that is why it exists.
This page covers the first part of the CBSE Class 10 Maths chapter on statistics: the mean of grouped data by all three methods, choosing between them, and finding a missing frequency.
Formula
What are the three formulas for the mean of grouped data?
All three compute the same mean; they differ only in what they multiply the frequencies by.
The direct method, using the class marks :
The assumed mean method, where is any convenient value — usually a class mark near the middle — and :
The step-deviation method, where is the common class width and :
Why the second one works. Subtracting the same from every value lowers the mean by exactly , so adding back at the end restores it. The mean of the shifted data plus the shift equals the mean of the original data — which is a statement about averages that holds for any data at all.
Why the third one works. Dividing every deviation by divides the mean of the deviations by , so multiplying by at the end restores it. The step-deviation method is the assumed mean method with a second, undone, simplification.
Notice the structural similarity. Each formula is "something a correction", and the correction is a weighted average of the deviations. **In the direct method the something is and the deviations are the values themselves.
The condition on the third method, which examiners ask about. The step-deviation method needs all class widths to be equal**, because a single has to divide every deviation. If the classes have unequal widths, use the direct or the assumed mean method — and saying so is itself an examinable point.
A convention that must be stated for the first method to be well defined. Class marks are computed from the class limits as given. If the classes are written as –, – and so on, they are already continuous and the marks are , and so on. **If they are written as –, –, with gaps, they must first be made continuous** by taking the class boundaries halfway between the limits — and a question that gives gapped classes is testing exactly that step.
The direct method, using the class marks :
The assumed mean method, where is any convenient value — usually a class mark near the middle — and :
The step-deviation method, where is the common class width and :
Why the second one works. Subtracting the same from every value lowers the mean by exactly , so adding back at the end restores it. The mean of the shifted data plus the shift equals the mean of the original data — which is a statement about averages that holds for any data at all.
Why the third one works. Dividing every deviation by divides the mean of the deviations by , so multiplying by at the end restores it. The step-deviation method is the assumed mean method with a second, undone, simplification.
Notice the structural similarity. Each formula is "something a correction", and the correction is a weighted average of the deviations. **In the direct method the something is and the deviations are the values themselves.
The condition on the third method, which examiners ask about. The step-deviation method needs all class widths to be equal**, because a single has to divide every deviation. If the classes have unequal widths, use the direct or the assumed mean method — and saying so is itself an examinable point.
A convention that must be stated for the first method to be well defined. Class marks are computed from the class limits as given. If the classes are written as –, – and so on, they are already continuous and the marks are , and so on. **If they are written as –, –, with gaps, they must first be made continuous** by taking the class boundaries halfway between the limits — and a question that gives gapped classes is testing exactly that step.
How do you find the mean by the direct method?
Write the class mark for every class, multiply each by its frequency, add both columns, and divide.
Worked example — the distribution used throughout this page. Find the mean of the following distribution of observations.
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
The two totals:
The mean:
Check the answer for plausibility. The class marks run from to , so the mean must lie somewhere between them — and does. More usefully, the frequencies are heavier in the upper classes, so the mean should sit above the middle of the range, which is . **It does, at . That two-second check catches a misplaced decimal point or a column added wrongly.
How the class marks were found.** Each is the average of the two limits:
**And they rise by each time, which is the class width. That regular step is what makes the step-deviation method possible, and noticing it early tells you which method to reach for.
Where the direct method costs you.** The products , , and all involve multiplying a two-digit frequency by a number with a decimal. On a distribution with larger class marks — thousands of rupees, say — this becomes genuinely slow and error-prone, and that is the practical reason the other two methods exist.
When the direct method is nevertheless the right choice. Use it when:
- The class marks are small whole numbers, so the products are easy
- The class widths are unequal, which rules the step-deviation method out
- There are only two or three classes, so no method saves much
One boundary case about class marks. A class written as – and the next as – appear to share the value . **The convention is that an observation of exactly is counted in the class –, so each class includes its lower limit and excludes its upper. The totals are unaffected either way**, but the convention has to be stated when a question asks how a boundary value is treated.
Worked example — the distribution used throughout this page. Find the mean of the following distribution of observations.
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
The two totals:
The mean:
Check the answer for plausibility. The class marks run from to , so the mean must lie somewhere between them — and does. More usefully, the frequencies are heavier in the upper classes, so the mean should sit above the middle of the range, which is . **It does, at . That two-second check catches a misplaced decimal point or a column added wrongly.
How the class marks were found.** Each is the average of the two limits:
**And they rise by each time, which is the class width. That regular step is what makes the step-deviation method possible, and noticing it early tells you which method to reach for.
Where the direct method costs you.** The products , , and all involve multiplying a two-digit frequency by a number with a decimal. On a distribution with larger class marks — thousands of rupees, say — this becomes genuinely slow and error-prone, and that is the practical reason the other two methods exist.
When the direct method is nevertheless the right choice. Use it when:
- The class marks are small whole numbers, so the products are easy
- The class widths are unequal, which rules the step-deviation method out
- There are only two or three classes, so no method saves much
One boundary case about class marks. A class written as – and the next as – appear to share the value . **The convention is that an observation of exactly is counted in the class –, so each class includes its lower limit and excludes its upper. The totals are unaffected either way**, but the convention has to be stated when a question asks how a boundary value is treated.
How do the assumed mean and step-deviation methods speed this up?
Subtract a convenient value from every class mark, and then divide by the class width. Both operations are undone at the end, so the mean is unchanged.
Worked example 1 — the assumed mean method on the same data. Take , the class mark of the middle-ish class, and compute .
- **–**: , , , product
- **–**: , , , product
- **–**: , , , product
- **–**: , , , product
- **–**: , , , product
- **–**: , , , product
**The same mean of , reached with much smaller numbers. And notice what happened to the class containing : its deviation is zero, so it contributes nothing at all.** Choosing to be the class mark of the class with the largest frequency kills the biggest product outright, which is the whole reason for choosing from the middle of the data rather than from an end.
Worked example 2 — the step-deviation method on the same data. Every deviation above is a multiple of , which is the class width. Divide by it: .
- **–**: , , product
- **–**: , , product
- **–**: , , product
- **–**: , , product
- **–**: , , product
- **–**: , , product
**The same again, and this time every number in the working column was a single digit. Compare the three sums you had to compute: , then , then . That is the whole case for the step-deviation method.
Which method is most efficient for this data, and why. The step-deviation method, because:
- All six class widths are equal**, at , so a single divides every deviation
- The class marks are large and carry decimals, which makes the direct products awkward
- **The deviations are all exact multiples of **, which they always are when the widths are equal and is a class mark
That last point deserves emphasis. If you choose to be one of the class marks, then every comes out as a whole number — and if you choose to be something else, they do not. **So choose from the class marks, ideally the one with the highest frequency.
Worked example 3 — a distribution where the step-deviation method is not available.** Suppose the classes were –, –, – and –, with widths , , and .
The widths are unequal, so no single divides all the deviations into whole numbers, and the step-deviation method cannot be used. Use the assumed mean method instead, which still shrinks the numbers without needing equal widths. The direct method also works, as it always does.
A comparison worth remembering. The three methods relate like this: the direct method has no simplification, the assumed mean method shifts, and the step-deviation method shifts and scales. Each undoes its own simplification at the end, which is why they cannot disagree — and if yours do disagree, one of the columns has an arithmetic error rather than a method error.
Worked example 1 — the assumed mean method on the same data. Take , the class mark of the middle-ish class, and compute .
- **–**: , , , product
- **–**: , , , product
- **–**: , , , product
- **–**: , , , product
- **–**: , , , product
- **–**: , , , product
**The same mean of , reached with much smaller numbers. And notice what happened to the class containing : its deviation is zero, so it contributes nothing at all.** Choosing to be the class mark of the class with the largest frequency kills the biggest product outright, which is the whole reason for choosing from the middle of the data rather than from an end.
Worked example 2 — the step-deviation method on the same data. Every deviation above is a multiple of , which is the class width. Divide by it: .
- **–**: , , product
- **–**: , , product
- **–**: , , product
- **–**: , , product
- **–**: , , product
- **–**: , , product
**The same again, and this time every number in the working column was a single digit. Compare the three sums you had to compute: , then , then . That is the whole case for the step-deviation method.
Which method is most efficient for this data, and why. The step-deviation method, because:
- All six class widths are equal**, at , so a single divides every deviation
- The class marks are large and carry decimals, which makes the direct products awkward
- **The deviations are all exact multiples of **, which they always are when the widths are equal and is a class mark
That last point deserves emphasis. If you choose to be one of the class marks, then every comes out as a whole number — and if you choose to be something else, they do not. **So choose from the class marks, ideally the one with the highest frequency.
Worked example 3 — a distribution where the step-deviation method is not available.** Suppose the classes were –, –, – and –, with widths , , and .
The widths are unequal, so no single divides all the deviations into whole numbers, and the step-deviation method cannot be used. Use the assumed mean method instead, which still shrinks the numbers without needing equal widths. The direct method also works, as it always does.
A comparison worth remembering. The three methods relate like this: the direct method has no simplification, the assumed mean method shifts, and the step-deviation method shifts and scales. Each undoes its own simplification at the end, which is why they cannot disagree — and if yours do disagree, one of the columns has an arithmetic error rather than a method error.
How do you find a missing frequency when the mean is given?
**Call the missing frequency , write both totals as expressions in , set the mean formula equal to the given mean, and solve the linear equation.
Worked example 1.** The mean of the following distribution is . Find the missing frequency .
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
**The two totals as expressions in :**
**Set the mean equal to :**
**Check by computing the mean with :**
Exactly the given mean. That check is compulsory, because the algebra is easy to get right in a way that still leaves a wrong answer if a product was mis-multiplied.
Notice why the equation came out linear. The unknown appears once in the numerator and once in the denominator, and cross-multiplying leaves single powers of on both sides. Every missing-frequency question is a one-step linear equation, and if yours turns out quadratic, an has been squared by mistake.
Worked example 2 — with the step-deviation method, which is faster here. Solve the same problem taking and , so that .
- ****: ,
- ****: ,
- ****: ,
- ****: ,
- ****: ,
- ****: ,
- ****: ,
Now the mean formula:
so the correction term must be zero:
One line. Because the assumed mean was chosen to equal the given mean, the whole correction had to vanish — so the equation became and nothing else. That is the fastest possible route to a missing frequency, and it works whenever the given mean happens to be one of the class marks.
Worked example 3 — two missing frequencies. If two frequencies are unknown, one equation is not enough. You need a second piece of information, and it is always supplied — usually the total frequency. With given as well, you get two linear equations in two unknowns and solve them by elimination, exactly as in the linear-equations chapter.
The general shape of every question in this family. One unknown needs one condition; two unknowns need two. Count your unknowns and count your conditions before you start, and if they do not match, re-read the question — a total frequency or a second mean is hiding in it somewhere.
Worked example 1.** The mean of the following distribution is . Find the missing frequency .
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
- **–**: frequency , class mark , product
**The two totals as expressions in :**
**Set the mean equal to :**
**Check by computing the mean with :**
Exactly the given mean. That check is compulsory, because the algebra is easy to get right in a way that still leaves a wrong answer if a product was mis-multiplied.
Notice why the equation came out linear. The unknown appears once in the numerator and once in the denominator, and cross-multiplying leaves single powers of on both sides. Every missing-frequency question is a one-step linear equation, and if yours turns out quadratic, an has been squared by mistake.
Worked example 2 — with the step-deviation method, which is faster here. Solve the same problem taking and , so that .
- ****: ,
- ****: ,
- ****: ,
- ****: ,
- ****: ,
- ****: ,
- ****: ,
Now the mean formula:
so the correction term must be zero:
One line. Because the assumed mean was chosen to equal the given mean, the whole correction had to vanish — so the equation became and nothing else. That is the fastest possible route to a missing frequency, and it works whenever the given mean happens to be one of the class marks.
Worked example 3 — two missing frequencies. If two frequencies are unknown, one equation is not enough. You need a second piece of information, and it is always supplied — usually the total frequency. With given as well, you get two linear equations in two unknowns and solve them by elimination, exactly as in the linear-equations chapter.
The general shape of every question in this family. One unknown needs one condition; two unknowns need two. Count your unknowns and count your conditions before you start, and if they do not match, re-read the question — a total frequency or a second mean is hiding in it somewhere.
Exam tip
What does a full-mark mean calculation look like on paper?
Set the work out as a table with a column for every quantity you compute, and total the columns you need. An examiner marks the columns, so working done in your head earns nothing even when the answer is right.
- Make the classes continuous first if they are given with gaps, such as – and –
- Write the class mark column as , and check the marks rise by the class width
- **Choose to be a class mark**, preferably the one with the largest frequency, so that every is a whole number
- Use the step-deviation method only when all class widths are equal, and say so
- **Total the column and the column, and show both totals
- Add back and multiply by ** in the right order: , with the multiplying only the fraction
- Keep negative deviations negative and add the column carefully — most errors in this chapter are sign errors in the column
- Check the mean lies between the smallest and largest class marks, and on the side the frequencies lean toward
- Verify a missing frequency by recomputing the mean with it substituted
- Name the method you used in your answer
The misconception to name. The mean of grouped data is an estimate, not the exact mean of the original observations. It assumes every value in a class sits at the class mark, which is almost never literally true. The estimate is close because the errors within a class largely cancel, and a question asking "why is this only an estimate?" wants exactly that sentence.
A second trap. Multiplying by in the wrong place. **The formula is **, so the multiplies the fraction and not the whole expression. Writing scales the assumed mean too, and in this chapter's example would have given instead of — an answer so far out that the plausibility check catches it at once.
- Make the classes continuous first if they are given with gaps, such as – and –
- Write the class mark column as , and check the marks rise by the class width
- **Choose to be a class mark**, preferably the one with the largest frequency, so that every is a whole number
- Use the step-deviation method only when all class widths are equal, and say so
- **Total the column and the column, and show both totals
- Add back and multiply by ** in the right order: , with the multiplying only the fraction
- Keep negative deviations negative and add the column carefully — most errors in this chapter are sign errors in the column
- Check the mean lies between the smallest and largest class marks, and on the side the frequencies lean toward
- Verify a missing frequency by recomputing the mean with it substituted
- Name the method you used in your answer
The misconception to name. The mean of grouped data is an estimate, not the exact mean of the original observations. It assumes every value in a class sits at the class mark, which is almost never literally true. The estimate is close because the errors within a class largely cancel, and a question asking "why is this only an estimate?" wants exactly that sentence.
A second trap. Multiplying by in the wrong place. **The formula is **, so the multiplies the fraction and not the whole expression. Writing scales the assumed mean too, and in this chapter's example would have given instead of — an answer so far out that the plausibility check catches it at once.
Did you know
Why does shifting all the data leave the average shifted by exactly the same amount?
Take any five numbers — say , , , , — whose mean is . Now subtract from each: you get , , , , . **Their mean is , since they add to zero.
That is not a property of those particular numbers. Subtracting the mean from every observation always gives a set whose mean is zero, because the amounts above the mean and the amounts below it are equal by definition — which is what being the mean means.
So the mean is a balance point. Imagine the number line as a see-saw with a unit weight placed at each observation. The mean is the point where it balances**, and the deviations are the distances of the weights from the pivot. The balance condition is exactly .
And that is why the assumed mean method works. You pivot the see-saw at instead of at the true mean, and measure how far out of balance it is. **The quantity is exactly how far you must move the pivot to restore the balance** — so adding it to lands you on the mean.
Which explains two things you may have noticed.
- **If is chosen below the true mean, comes out positive**, and the correction pushes the answer up. In this chapter's example gave , and the mean was indeed above
- **If you happened to choose equal to the true mean, the sum would be exactly zero — which is precisely the shortcut that solved the missing-frequency question in one line
The scaling in the step-deviation method has the same kind of explanation.** Measuring distances in units of instead of divides every deviation by , so the out-of-balance amount comes out in those units too. **Multiplying by converts it back, exactly as converting a distance from metres to centimetres would.
One consequence that matters far beyond this chapter. Because the deviations from the mean always sum to zero, they are useless for measuring spread — you cannot average them to see how scattered the data is, since the answer is always zero. That is why spread is measured by squaring the deviations first, and the quantity that results is the variance. The fact that makes the assumed mean method work is the same fact that forces the definition of variance**, and in Class 11 the two appear a page apart.
That is not a property of those particular numbers. Subtracting the mean from every observation always gives a set whose mean is zero, because the amounts above the mean and the amounts below it are equal by definition — which is what being the mean means.
So the mean is a balance point. Imagine the number line as a see-saw with a unit weight placed at each observation. The mean is the point where it balances**, and the deviations are the distances of the weights from the pivot. The balance condition is exactly .
And that is why the assumed mean method works. You pivot the see-saw at instead of at the true mean, and measure how far out of balance it is. **The quantity is exactly how far you must move the pivot to restore the balance** — so adding it to lands you on the mean.
Which explains two things you may have noticed.
- **If is chosen below the true mean, comes out positive**, and the correction pushes the answer up. In this chapter's example gave , and the mean was indeed above
- **If you happened to choose equal to the true mean, the sum would be exactly zero — which is precisely the shortcut that solved the missing-frequency question in one line
The scaling in the step-deviation method has the same kind of explanation.** Measuring distances in units of instead of divides every deviation by , so the out-of-balance amount comes out in those units too. **Multiplying by converts it back, exactly as converting a distance from metres to centimetres would.
One consequence that matters far beyond this chapter. Because the deviations from the mean always sum to zero, they are useless for measuring spread — you cannot average them to see how scattered the data is, since the answer is always zero. That is why spread is measured by squaring the deviations first, and the quantity that results is the variance. The fact that makes the assumed mean method work is the same fact that forces the definition of variance**, and in Class 11 the two appear a page apart.
Exam relevance
How is the mean of grouped data used in JEE and NEET?
This is foundation work for Class 11 Statistics, which is a JEE Main topic, and the summation notation here is used throughout Physics and Chemistry.
Where the mean formulas lead. Class 11 keeps unchanged and adds the mean deviation, the variance and the standard deviation, all built on the same grouped table with one extra column. JEE Main asks for the variance or the standard deviation of a grouped distribution as a direct numerical, and the table you learn to lay out here is the table those questions need.
Where the step-deviation idea leads. Class 11 uses exactly the same shift-and-scale trick for variance, with the result that the variance of times gives the variance of . **The shift leaves the variance unchanged and the scale multiplies it by — which is the natural next question after seeing the shift and scale behave as they do for the mean, and it is a favourite JEE Main one-liner.
Where the balance-point picture leads. It becomes the centre of mass** in Class 11 Physics, where replaces with masses in place of frequencies. The formula is identical, and the assumed-mean trick becomes the standard practice of choosing a convenient origin. Both JEE and NEET Physics rely on that choice.
Where the deviations-sum-to-zero fact leads. It is why variance uses squares, and it recurs in Chemistry when averaging bond enthalpies and in error analysis when combining measurements.
Where the missing-frequency technique leads. Competitive papers set problems giving a mean and one unknown observation, or a combined mean of two groups from which one group's mean must be recovered. The method is the same: write the totals as expressions in the unknown and solve.
Question types to expect. At this level: mean by each of the three methods, justifying the choice of method, and a missing frequency. In competitive papers: variance and standard deviation of grouped data, the effect of shifting and scaling on the mean and the variance, and combined means.
The single trap that costs marks. Misplacing the in the step-deviation formula. It multiplies the fraction only, and the plausibility check — does the mean lie between the extreme class marks? — catches the error immediately. In Class 11 the corresponding slip is multiplying the variance by instead of .
A second trap. Using the step-deviation method on unequal class widths. **A single cannot divide every deviation into whole numbers**, so the method does not apply, and a question with widths , , and is testing whether you notice. Say which method you are using and why.
Board versus competitive emphasis. The CBSE paper marks the table, each column, the two totals, the substitution and the named method; a competitive paper marks a single value, often a variance rather than a mean. The transferable habit is choosing the origin to make the arithmetic easy — an assumed mean in statistics, a convenient origin in mechanics, a reference level for potential energy. It is the same decision every time, and it always saves work.
Where the mean formulas lead. Class 11 keeps unchanged and adds the mean deviation, the variance and the standard deviation, all built on the same grouped table with one extra column. JEE Main asks for the variance or the standard deviation of a grouped distribution as a direct numerical, and the table you learn to lay out here is the table those questions need.
Where the step-deviation idea leads. Class 11 uses exactly the same shift-and-scale trick for variance, with the result that the variance of times gives the variance of . **The shift leaves the variance unchanged and the scale multiplies it by — which is the natural next question after seeing the shift and scale behave as they do for the mean, and it is a favourite JEE Main one-liner.
Where the balance-point picture leads. It becomes the centre of mass** in Class 11 Physics, where replaces with masses in place of frequencies. The formula is identical, and the assumed-mean trick becomes the standard practice of choosing a convenient origin. Both JEE and NEET Physics rely on that choice.
Where the deviations-sum-to-zero fact leads. It is why variance uses squares, and it recurs in Chemistry when averaging bond enthalpies and in error analysis when combining measurements.
Where the missing-frequency technique leads. Competitive papers set problems giving a mean and one unknown observation, or a combined mean of two groups from which one group's mean must be recovered. The method is the same: write the totals as expressions in the unknown and solve.
Question types to expect. At this level: mean by each of the three methods, justifying the choice of method, and a missing frequency. In competitive papers: variance and standard deviation of grouped data, the effect of shifting and scaling on the mean and the variance, and combined means.
The single trap that costs marks. Misplacing the in the step-deviation formula. It multiplies the fraction only, and the plausibility check — does the mean lie between the extreme class marks? — catches the error immediately. In Class 11 the corresponding slip is multiplying the variance by instead of .
A second trap. Using the step-deviation method on unequal class widths. **A single cannot divide every deviation into whole numbers**, so the method does not apply, and a question with widths , , and is testing whether you notice. Say which method you are using and why.
Board versus competitive emphasis. The CBSE paper marks the table, each column, the two totals, the substitution and the named method; a competitive paper marks a single value, often a variance rather than a mean. The transferable habit is choosing the origin to make the arithmetic easy — an assumed mean in statistics, a convenient origin in mechanics, a reference level for potential energy. It is the same decision every time, and it always saves work.
Key takeaways
What must you be able to do from this part?
Three methods, one answer and one condition on the fastest of them.
- A class mark is , and every observation in a class is treated as sitting there
- The mean of grouped data is therefore an estimate, close because errors within a class largely cancel
- Make classes continuous first if they are given with gaps
- Direct method:
- Assumed mean method: , with
- Step-deviation method: , with
- **The multiplies the fraction only, not the whole expression
- All three give the same mean.** For the six-class distribution of observations they give , then , then — all equal to
- **The sums shrink from to to , which is the whole case for the step-deviation method
- Choose to be a class mark**, ideally the one with the largest frequency, so every is a whole number and one product vanishes
- The step-deviation method needs equal class widths — with widths , , , it cannot be used, and the assumed mean method should be
- For a missing frequency, write both totals in terms of and set the mean equal to the given value — a mean of gives and so
- Verify: with ,
- **If the given mean equals a class mark, take it as ** and the whole correction must vanish, giving in one line
- Two unknown frequencies need two conditions, usually the total frequency as well
- Check the mean lies between the extreme class marks and leans the way the frequencies do
- The deviations from the mean always sum to zero, which is why the mean is the balance point of the data
The most convincing self-test is one table done twice. Take any grouped distribution, find its mean by the direct method and again by the step-deviation method, and see whether the two land on the same number — and notice how much smaller the second set of products was.
- A class mark is , and every observation in a class is treated as sitting there
- The mean of grouped data is therefore an estimate, close because errors within a class largely cancel
- Make classes continuous first if they are given with gaps
- Direct method:
- Assumed mean method: , with
- Step-deviation method: , with
- **The multiplies the fraction only, not the whole expression
- All three give the same mean.** For the six-class distribution of observations they give , then , then — all equal to
- **The sums shrink from to to , which is the whole case for the step-deviation method
- Choose to be a class mark**, ideally the one with the largest frequency, so every is a whole number and one product vanishes
- The step-deviation method needs equal class widths — with widths , , , it cannot be used, and the assumed mean method should be
- For a missing frequency, write both totals in terms of and set the mean equal to the given value — a mean of gives and so
- Verify: with ,
- **If the given mean equals a class mark, take it as ** and the whole correction must vanish, giving in one line
- Two unknown frequencies need two conditions, usually the total frequency as well
- Check the mean lies between the extreme class marks and leans the way the frequencies do
- The deviations from the mean always sum to zero, which is why the mean is the balance point of the data
The most convincing self-test is one table done twice. Take any grouped distribution, find its mean by the direct method and again by the step-deviation method, and see whether the two land on the same number — and notice how much smaller the second set of products was.