PrepShorts · Study sheet · Class 10 Mathematics · Chapter 13, Statistics
Chapter 13 · Statistics
Dividing through by the class width to shrink the arithmetic further
This video could not be loaded. Reload the page to try again.
Sign in with Google15 min.
Keep your place in this chapter — sign in, it’s free.Sign in
The step in the step-deviation method is usually introduced as the class width, and on a table of six classes with three different widths that description stops making sense. The method runs anyway, at a step of 20 that matches none of them, and lands on exactly the number the direct method lands on - because the algebra never asked h to be the class width. It only ever asked h not to be nought.
The idea
The second reduction scales where the first one shifted, and scaling survives averaging for the same reason shifting did — divide every value by h and the average divides by h too. So multiplying back by h at the end recovers the mean exactly, and the two-parameter formula follows from the same pair of facts about sums. The surprise is how little h is required to be. The chapter introduces it as the class size, but the algebra only ever needs it to be non-zero, and Example 3 spends that freedom openly: it runs the method across six classes whose widths take three different values, using an h equal to none of the three, and accepts a deviation that will not divide evenly. The step is a convenience the method can lose without losing its correctness.
What you should be able to do
- Build the u column as (x − a) ÷ h for a given assumed mean and step
- Compute a grouped mean as a + h × (Σf·u ÷ Σf)
- Derive that formula from the definition of u, naming the property used at each step
- Choose h as a common divisor of the deviations rather than assuming it must be the class width
- Apply the method to a distribution whose classes have unequal widths, and justify why this is legitimate
- Verify a step-deviation answer against the direct method on the same data
- Decide which of the three methods a given table calls for, and defend the choice
Words to know
| Term | Definition in one line | First introduced |
|---|---|---|
| step-deviation method | computing the mean from deviations that have been divided through by a chosen h | printed in this chapter, §13.2, p. 177 |
| class size | the width of a class, and the chapter's first choice for h | printed in this chapter, §13.2, p. 176 |
| assumed mean | the value a that the class marks are measured from | printed in this chapter, §13.2, p. 174 |
| deviation | the signed gap x − a before any division | printed in this chapter, §13.2, p. 174 |
| common factor | a number dividing every entry of the deviation column | printed in this chapter, §13.2, p. 177 |
| divisor | what the chapter asks h to be with respect to the deviations, when class sizes differ | printed in this chapter, §13.2, p. 179 |
| Direct Method | the mean straight from class marks, with no shift and no scaling | printed in this chapter, §13.2, p. 174 |
| Assumed Mean Method | the shift alone, with no scaling | printed in this chapter, §13.2, p. 176 |
| step | one unit of the u column, worth h of the original scale | printed in this chapter inside the method's name, §13.2, p. 177; the standalone gloss given here is added here |
| two-parameter family | the single formula that a and h specialise into all three methods | an added term |
Where people slip up
- "h must equal the class size." It must divide the deviations usefully. Example 3 uses 20 against widths of 40, 50 and 100, and the answer is right. The class size is merely the choice that works when all classes share one.
- "Unequal classes rule the method out." They do not rule anything out. What the chapter says twice — in the mode formula's symbol list on p. 184 and the median's on p. 193 — is that those printed formulas are written assuming the classes are equal in size. It then says twice more, in the Remark on p. 186 and Remark 2 on p. 197, that both measures can be found for unequal classes and that it will not go into how. The median formula in fact runs unchanged with h read as the median class's own width. The mean has no such caveat at all, and Example 3 is the chapter demonstrating it.
- "−3.75 in the u column means I picked the wrong h." It means 75 is not a multiple of 20. The arithmetic proceeds unharmed; only the tidiness suffers. Choosing h = 5 removes the fraction and enlarges everything else.
- "The three methods are three formulas to memorise." They are one formula at three settings of a and h. A student who sees that has two fewer things to remember and a check they can run for free.
- "Multiply by h at the start, when building u." You divide by h to build u and multiply by h at the end. Reversing this is the single commonest slip, and it shows up as an answer wrong by a factor of h².
- "Forgetting h just gives a slightly wrong answer." Dropping the final multiplication on Example 3 gives 200 − 2.36, about 197.6, against 152.89 — not slightly wrong. The check against the direct method catches it instantly.
- "The mean must land inside the busiest class." Example 3's mean, 152.89, falls in 150–250, while the busiest class is 100–150. The mean is not a location of frequency.
Ask your teacher a person
Your teacher reads this and writes back, usually within a day. For an instant answer, use Ask the video in the sidebar.
Your class sees the question and the answer. Only your teacher sees that it was you.
No questions on this topic yet.
Worked answers: Exercise 13.1 · Exercise 13.2 · Exercise 13.3 · this video explains Exercise 13.1 Q8
Transcript2,014 words
Here is the deviation column from last time. Thirty students, six classes, and a flag planted at forty-seven and a half. Minus thirty, minus fifteen, nought, fifteen, thirty, forty-five. That was already a large improvement on the class marks. But look at those six numbers again. Every one of them is a multiple of fifteen. Which is not a coincidence. The classes are fifteen wide, so consecutive class marks are fifteen apart, and every deviation from one of them is a whole number of fifteens.
There is a factor sitting in every entry of the column, and you are carrying it through every single multiplication. So take it out. Divide the whole column through by fifteen. Minus thirty becomes minus two. Minus fifteen becomes minus one. Nought stays nought. Then one, two, three. Six entries, none of them wider than a single digit. That column has a name. It is called u, and each entry is the deviation divided by the number you divided through by, which is written h.
And it is worth asking what those new numbers are actually counting. Minus two does not mean minus two marks. It means two classes below the flag. The class marks are spaced fifteen apart, so one unit of u is one whole class — one step along the row. That is where the method gets its name. Step deviation. The deviations, counted in steps rather than in marks. So rebuild the table with that column in it.
Multiply each u by its frequency. Minus two times two is minus four. Minus one times three is minus three. Nought. Then one times six is six, two times six is twelve, three times six is eighteen. Total the column: twenty-nine. Compare that with the four hundred and thirty-five you were totalling before, and the one thousand eight hundred and sixty before that. Every entry now fits in two digits.
Twenty-nine shared among thirty students is twenty-nine thirtieths. Now — that is not the answer, and it is not even the shortfall. It is the shortfall measured in steps. So turn it back into marks by multiplying by fifteen. Fifteen times twenty-nine thirtieths is fourteen and a half. And add the flag back. Forty-seven and a half plus fourteen and a half is sixty-two. Sixty-two again. Which raises exactly the question the last video raised, one level up.
The flag was free — any centre gave the same answer. Is the step free too? And there is a real worry here that there was not before. Subtracting a number and adding it back is obviously reversible. But dividing a whole column and then multiplying one number by the same amount — that is doing two different-looking things at two different places. The claim is that dividing every value by h divides their average by h, and nothing else happens.
If that is true, multiplying back at the end recovers the mean exactly. If it is only nearly true, the method is worthless. It is true, and it takes the same two facts about sums that the last derivation used. Start from the average of the u column. Sum of f times u, over sum of f. Every u is x minus a, all over h. Write that in. Now h is a constant. It is the same for every class, so it comes out of the sum — exactly as a did last time, and for exactly the same reason.
That leaves one over h, times the sum of f times x minus a, over the sum of f. And that second piece is a quantity you already know. It is the average of the deviations, which the last video proved is the mean minus a. So the average of the u column is the mean, minus a, all divided by h. Nothing has been approximated. The h that went into every entry came out of the average untouched, because a constant factor survives averaging.
Which tells you exactly what to do to get back. Multiply both sides by h, and you have h times the u-average equals the mean minus a. Add a to both sides, and the method is finished. The mean is a, plus h, times the sum of f u over the sum of f. Two parameters now instead of one. You choose a, you choose h, and the formula returns the same number whatever you chose.
And notice the order, because reversing it is the commonest slip in the whole topic. You DIVIDE by h when you build the column, and you MULTIPLY by h at the very end. Do it the other way round and the answer is wrong by a factor of h squared. On this table that turns sixty-two into three thousand three hundred and ten. Here is the picture that makes the word step mean something.
Draw the original scale along the top: seventeen and a half, thirty-two and a half, forty-seven and a half, and so on. Now draw a second axis underneath, ticked every fifteen units, starting from the flag. Number those ticks nought, one, two, and minus one, minus two going the other way. That second axis is the u scale, and every class mark drops straight onto one of its ticks. Ninety-two and a half lands on three. Seventeen and a half lands on minus two.
So u is not some algebraic convenience. It is a reading off a second ruler, and h is how far apart that ruler's marks are. The mean comes out as a position on that ruler, and multiplying by h is simply converting it back to the scale you started on. Now for the table that shows how little h is really required to be. Forty-five bowlers, and how many wickets each has taken across a career. Six classes.
Twenty to sixty. Sixty to a hundred. A hundred to a hundred and fifty. A hundred and fifty to two fifty. Two fifty to three fifty. Three fifty to four fifty. Look at the widths. Forty. Forty. Fifty. A hundred. A hundred. A hundred. Three different class widths in one table, which the previous examples never had. And the step used here is twenty. Twenty is not the class width. It is not any of the three class widths. There is no class in this table twenty wide.
Draw a ruler ticked every twenty units, lay it under those six classes, and it fits none of them. And it does not matter in the slightest, because h was never required to be the class width. Look back at the derivation: h appears in a division and in a multiplication that undo each other, and that is all. Take the centre at two hundred and the deviations come out minus a hundred and sixty, minus a hundred and twenty, minus seventy-five, nought, a hundred, two hundred.
Divide those by twenty. Minus eight. Minus six. Then minus three point seven five. And there it is — a decimal, sitting in the middle of a column that was supposed to be small whole numbers. That is not a mistake, and it is not a sign you chose h badly. It means seventy-five is not a multiple of twenty, and that is the entire explanation. The arithmetic proceeds completely unharmed. Multiply through: minus fifty-six, minus thirty, minus sixty, nought, ten, thirty.
Total, minus one hundred and six. Over forty-five bowlers, times twenty, plus two hundred. One hundred and fifty-two point eight nine. And here is the check that settles it. Run the direct method on the same table — every class mark times its frequency — and you get six thousand eight hundred and eighty over forty-five. The same number, to the last digit. The unequal widths and the decimal did no damage whatever.
If the decimal bothers you, you can get rid of it. There is a step that divides every one of those six deviations exactly, and it is five. At h equals five the u column is minus thirty-two, minus twenty-four, minus fifteen, nought, twenty, forty. Every entry a whole number. The products come to minus four hundred and twenty-four, and the answer is one hundred and fifty-two point eight nine again.
Two different steps, two completely different columns, one answer. So which should you have used? Count the digits your hand forms. At a step of twenty, the product column and its total cost fourteen characters. At a step of five, eighteen. The column with the decimal in it is the CHEAPER of the two. Which is worth sitting with, because a column of whole numbers looks like less work and here it is more.
Making h smaller removes fractions and makes every entry bigger. Making h larger shrinks the entries and risks a fraction. A judgement, then, not a rule. So what does h actually have to satisfy? Only one thing. It may not be nought — because you divide by it, and you cannot divide by nothing. That is the whole restriction. Everything else is convenience. But convenience is worth having, so take the deviations and find their greatest common divisor.
On the first table that is fifteen, which is also the class width — which is why the class width is the obvious choice when every class shares one. On the bowlers it is five. On a table of heartbeats it is three, and on one of household spending, fifty. And on a table of days absent from school it is one. Seven classes, of widths six, four, four, six, eight, ten and two — and the deviations share no factor beyond one. There is no useful step for that table at all.
So the advice runs out, and the honest thing is to stop trying. Compute it directly: four hundred and ninety-nine over forty, twelve point four seven five days. Not every table wants this method. Knowing which ones do is part of the skill. Now put the two dials side by side, because there are not three methods here. Take the first table. Set a to forty-seven and a half and h to fifteen, and the u column runs minus two to three, totals twenty-nine, and gives sixty-two. That is the step deviation method.
Now turn h down to one. Dividing by one changes nothing, so the u column IS the deviation column. It totals four hundred and thirty-five, and gives sixty-two. That is the assumed mean method. Now turn a down to nought as well. Subtracting nothing changes nothing, so the column is the class marks themselves. It totals one thousand eight hundred and sixty, and gives sixty-two. That is the direct method.
One formula. Two dials. Three settings, three completely different columns, and the same answer every time. Which means these are not three things to memorise. There is one thing, and the other two are what it does when you leave a dial alone. So how do you choose, given that all three are guaranteed to agree? On cost, and nothing else. Count the digit characters. On that first table the three columns cost twenty-two, sixteen and ten. Each move takes about a third off the last, and every one returns sixty-two.
So look at the table before you start. If the marks and the frequencies are small, go direct — there is nothing to gain. If the marks are large, subtract one of them and work from there. And if the deviations you get all share a factor, divide it out and count in steps instead. One more thing this buys you, for free. Because the three methods cannot disagree, running any two of them is a complete check on your own arithmetic.
Drop the final multiplication by h on the bowlers and you get one hundred and ninety-seven point six four instead of one hundred and fifty-two point eight nine. That is an error of more than a quarter of the answer, and the direct method catches it in one line. The step is a convenience. The answer never was.
Where this fits
Taken from the notes each video was made from, not from the reading order — these are the ideas this one rests on and the ones that later rest on it.
Builds on
- Shifting the origin: why guessing a centre cannot change the answerClass 10 · Ch 13, Statistics
- The direct method, and where it becomes unwieldyClass 10 · Ch 13, Statistics
Comes up again in
- Which of the three averages a given question actually wantsClass 10 · Ch 13, Statistics
Either side of this one
- Finding the busiest class, then placing the mode inside itClass 10 · Ch 13, Statistics