PrepShorts · Study sheet · Class 10 Mathematics · Chapter 13, Statistics
This video could not be loaded. Reload the page to try again.
Sign in with Google15 min.
Keep your place in this chapter — sign in, it’s free.Sign in
The direct method is not a method. Write the formula for the mean of a table of values above the formula for the mean of a table of classes and they are the same characters in the same order - only the x column has changed hands. That is why this one route to a grouped mean needs no proof, and why the two that save you arithmetic do.
The idea
The direct method is not really a method — it is the definition of the mean, applied without alteration to a table whose values have been replaced by class marks. That is why it alone needs no justification, and it is the reason the other two methods do. Its weakness is arithmetic bulk, never correctness: the products pile up as the class marks and frequencies grow. So the chapter's fix changes the numbers rather than the formula, and the moment it does, it takes on a debt — a rescaling that alters every entry has to be proved to leave the answer alone, and that proof is the whole content of the two methods that follow.
What you should be able to do
- State the direct-method formula and identify each symbol in a given table
- Lay out a grouped-mean calculation as a table with an f·x column and a totals row
- Compute the mean of a grouped distribution by the direct method and round the result sensibly
- Judge, before starting, whether the direct method or a reduced-arithmetic method suits a given table, and give the reason in terms of the size of the class marks and frequencies
- Explain why the direct method requires no proof of validity while the assumed mean and step-deviation methods do
- Show that the mean is unaffected by converting inclusive classes into continuous ones, and say why
Words to know
| Term | Definition in one line | First introduced |
|---|---|---|
| Direct Method | computing the grouped mean straight from the class marks and frequencies | printed in this chapter, §13.2, p. 174 |
| class mark | the mid-point standing in for every observation in a class | printed in this chapter, §13.2, p. 173 |
| class size | the width of a class, its upper limit less its lower limit | printed in this chapter, §13.2, p. 176 |
| frequency | the count of observations in a class | printed in this chapter, §13.2, p. 171 |
| class interval | one of the ranges the data has been cut into | printed in this chapter, §13.2, p. 173 |
| continuous classes | classes whose upper and lower limits meet, leaving no gap | printed in this chapter, in the closing note on p. 201 and in the hint to Exercise 13.3 Q4, p. 199 |
| assumed mean method | the first of the two reduced-arithmetic routes | printed in this chapter, §13.2, p. 176 |
| step-deviation method | the second, which also divides through by a chosen size | printed in this chapter, §13.2, p. 177 |
| product column | the working column holding each f·x | an added label for a column the chapter builds but does not name |
| arithmetic bulk | how large the intermediate numbers get, as distinct from how hard the idea is | an added phrasing |
Where people slip up
- "The direct method is the crude one and the others are better." All three return the same number. The later two are labour-saving devices, not improvements in accuracy — a point the chapter states flatly after Example 2.
- "Choose the method by the size of the answer." Choose it by the size of the intermediate products. Q7's answer is about 0.1 and its arithmetic is still fiddly; Q1's answer is 8.1 and its arithmetic is trivial.
- "Σf·x is the total of the data." It is the total of the reconstructed data. In Example 1 the real total was 1779 and the reconstructed total 1860.
- "You may simplify the frequencies too." You may not. Frequencies are the observed counts; changing one changes the data. Only the class marks are relabelled, and only because the relabelling is later undone.
- "Continuity correction changes every calculation." It changes the class limits, and therefore the mode and median, which are computed from limits. The class mark is the average of the two limits and both move by the same half unit in opposite directions, so the mean does not budge.
- "Round every intermediate value to two places." Example 2's mean is 39.714…, rounded once at the end. Rounding the products first will move the last digit.
Ask your teacher a person
Your teacher reads this and writes back, usually within a day. For an instant answer, use Ask the video in the sidebar.
Your class sees the question and the answer. Only your teacher sees that it was you.
No questions on this topic yet.
Worked answers: Exercise 13.1 · Exercise 13.2 · Exercise 13.3 · this video explains Exercise 13.1 Q1, Exercise 13.1 Q2, Exercise 13.1 Q3, Exercise 13.1 Q4, Exercise 13.1 Q5, Exercise 13.1 Q6, Exercise 13.1 Q7, Exercise 13.1 Q9
Transcript1,858 words
Here is the formula for the mean of a set of values, each one appearing some number of times. Add up every value times how often it occurs, and divide by how many there are altogether. And here is the formula for the mean of a table whose data has been sorted into classes. Look at them. They are the same formula. Character for character. The only thing that changed is what sits in the x column: real values in the first, class marks in the second.
That is why this one is called the direct method. It is not a method at all. It is the definition of an average, applied without alteration to a table whose values have been replaced by stand-ins. That is worth stopping on, because it has a consequence you will need twice more. There are two other ways to compute a grouped mean, and both of them start by changing the numbers.
One subtracts a chosen value from every class mark. The other divides through as well. Both of them then have to show you something: that the answer came back unharmed. That is a debt. It is not difficult to pay, but it has to be paid. The direct method borrows nothing, so it owes nothing. It changes no number at all. Which means it can never be wrong, and it is the standard the other two are checked against.
Its only weakness is labour, and labour is a measurable thing. So what does the labour actually consist of? Four physical actions, in order, on any table you are ever given. First, write down the class mark of every class: the average of its two limits. Second, multiply each mark by its frequency. That is a column of products, one per class. Third, total two columns: the frequencies, and the products.
Fourth, divide one total by the other. Once. That is the whole procedure, and notice how the four are not equally heavy. The marks are averages of two small numbers. The division happens a single time. Everything expensive is in step two. Let us run it on something real. Here is the percentage of women among primary school teachers in rural areas, recorded for thirty-five states and territories. The classes run from 15 to 25 up to 75 to 85, and the counts are 6, 11, 7, 4, 4, 2 and 1.
Class marks first: 20, 30, 40, 50, 60, 70, 80. Now the product column. 6 times 20 is 120. 11 times 30 is 330. And so on down the seven. The frequencies total 35, which is not a coincidence about classes at all. It is the number of territories. The products total 1390. One division: 1390 over 35. That comes to 39.714 and so on, which we report as 39.71.
Rounded once, at the end. Never before. Now look back at what that cost. Seven products, none of them wider than three digits. The biggest was 330. That is a table you can do on paper without complaint. Here is another one with exactly the same shape. The lengths of 400 boxes, in five classes, with 135 of them sitting against a class mark of 57. 135 times 57 is 7695.
Same method. Same four steps. One product with four digits in it instead of three, and five such products in a row. Nothing became harder to understand. Something became harder to do. The usual advice at this point is to escape to a method with smaller numbers, and that advice is sound. But it is worth making it precise, because the version students remember is wrong. Unwieldy sounds like a feeling. It is not. It is a number, and here is how to get it.
Write every product out in full and count its digits, before the point and after it. Add those counts up, and add the digits of the total they go into. That is the pen-work in a table: how many figures your hand has to form. The teacher percentages come to 24. The box lengths come to 24 as well. Equal in total, and not equal in feel, because the boxes concentrate theirs into wider single products.
A table of plants per house comes to 14. A table of daily pay comes to 25. And a table of air pollution readings, with class marks like 0.02 and 0.06, comes to 20. Five tables, one method, and a real spread in what it costs. Now put those two lists side by side, and something breaks. Order the five tables by the size of their answers. The pollution table is smallest, at about 0.0987. Then plants at 8.1, then teachers at 39.71, then boxes at 57.19, then pay at 545.2.
Now order the same five by pen-work. Plants is cheapest at 14. Then pollution at 20, then teachers and boxes at 24 each, then pay at 25. The first two have swapped. The table with the smallest answer is not the table with the least arithmetic. Its answer is about a tenth. Its class marks are 0.02, 0.06, 0.10, and every product is a decimal you have to line up.
So the rule is not 'look at the answer'. You cannot see the answer yet. The rule is: look at the class marks and the frequencies, and ask how wide their products will be. Which brings us to the escape, and to a line you must not cross while taking it. A grouped table has two columns of numbers, and they are not the same kind of thing. The frequencies are the data. Somebody counted them. Change one and you are no longer describing the same thirty-five territories.
The class marks are labels. Nobody scored a class mark. They were appointed, and they can be appointed differently. Test that, rather than believing it. Take every one of the five tables, subtract some chosen number from every class mark, divide every result by some chosen size, find the mean of that, then multiply back and add on. Twenty-five different combinations of shift and size, on each of five tables. A hundred and twenty-five in all.
Every single one came back to exactly the number the table started with. Now do the opposite experiment. Bump one frequency by one, in every class of every table. Thirty-three attempts, and the mean moved on thirty-two of them. The one that did not move is worth a moment. It was a table whose middle class sat exactly on its own mean, so adding one more observation there changed nothing.
One exception, and it is an exception you can point at and explain. That is the difference between a rule and a coincidence. But notice how much had to be checked to say that relabelling is safe. And notice that the relabelling has to be a genuine one. Square every class mark instead. That is also a change to the label column, and it is also reversible. Run the same hundred and twenty-five checks with squares in place of shifts, and not one of them comes back.
So 'you may alter the class marks' is not licence. It is a specific permission for shifting and scaling, and it is exactly what the next two methods spend their proofs on. The direct method has no such paragraph, because it never asked for one. Here are four tables side by side, worked the same way, so you can see the shape stay still while the numbers grow. Plants per house, over 20 houses. Marks 1, 3, 5 up to 13. Products 1, 6, 5, 35, 54, 22, 39. Total 162. Mean 8.1 plants.
Air pollution over 30 localities. Marks 0.02 up to 0.22. The products total 2.96, and the mean is 0.0987, or 0.099 to three places. Daily pay for 50 workers. Marks 510 up to 590. Products in the thousands. Total 27260, and the mean is 545.2 exactly. Box lengths over 400 boxes. Marks 51 up to 63. Total 22875, and the mean is 57.1875 exactly. Four answers of wildly different size. One column of products, two totals, one division, every time.
That is the direct method's real virtue. There is nothing to remember and nothing to get wrong except arithmetic. There is one way to get that arithmetic wrong that has nothing to do with size, and it is a tempting one. Rounding as you go. The pollution table's class marks are 0.02, 0.06, 0.10, 0.14, 0.18 and 0.22. Round them to one decimal place first, to make the multiplying easier.
You get 0.0, 0.1, 0.1, 0.1, 0.2, 0.2, and the work really does get easier. The answer comes out at 0.1067. The true answer is 0.0987. That is an error of about eight per cent of the answer itself. In a table whose whole point is a reading to three decimal places. Round once, at the end, when there is nothing left to multiply. One more thing about this method, and it is the thing most often got wrong.
The box table's classes were written 50 to 52, then 53 to 55, then 56 to 58, and so on. They do not touch. There is a gap of one between each pair. Some calculations need classes that meet exactly, so the usual correction is to lower every lower limit by a half and raise every upper limit by a half. 50 to 52 becomes 49.5 to 52.5. The gaps close.
Every limit in that table has moved. Every class width has gone from 2 to 3. And the mean has not moved at all. Watch why. The class mark is the average of the two limits. One went down by a half and the other went up by a half. Their average is exactly where it was. 51 before, 51 after, in every class. So the mean, which is built out of class marks, does not notice.
Anything built out of limits does. Lower only the lower limits and every mark drops a quarter, and so does the mean. That is the whole reason the correction matters for some quantities and not for this one. Last, an equation is readable in both directions, and this one is no exception. Here is a table of pocket money with one frequency missing, and the mean given as 18. The classes are 11 to 13 up to 23 to 25, marks 12 up to 24, with counts 7, 6, 9, 13, then a gap, then 5 and 4.
You could solve for it algebraically. Or you could simply try. Offer the table every whole number from 0 to 200 in the empty slot, and ask each time where it balances. Exactly one of them balances at 18, and it is 20. Ask instead for a mean of 17, and no whole number works at all. Not every target is reachable. With 20 in place, the classes contribute 84, 84, 144, 234, 400, 110 and 96. That totals 1152, over 64 children, which is 18.
One product column. Two totals. One division. It is the definition, and it never needed to be anything else.
Where this fits
Taken from the notes each video was made from, not from the reading order — these are the ideas this one rests on and the ones that later rest on it.
Builds on
- Grouping loses the raw values, so we stand the class mark in for themClass 10 · Ch 13, Statistics
Comes up again in
- Shifting the origin: why guessing a centre cannot change the answerClass 10 · Ch 13, Statistics
- Dividing through by the class width to shrink the arithmetic furtherClass 10 · Ch 13, Statistics