PrepShorts · Study sheet · Class 10 Mathematics · Chapter 13, Statistics
Chapter 13 · Statistics
Grouping loses the raw values, so we stand the class mark in for them
This video could not be loaded. Reload the page to try again.
Sign in with Google13 min.
Keep your place in this chapter — sign in, it’s free.Sign in
Thirty students, one paper. Work out the average and you get 59.3. Sort the same thirty marks into six classes, work out the average from that table, and you get 62. Neither number is a mistake, and the 2.7 between them is exactly the amount by which the classes fail to balance about their middles.
The idea
Once data is grouped, no observation's own value survives anywhere in the table — only a count per class does. So a mean can no longer be computed; it has to be reconstructed, by appointing one stand-in value for every observation in a class. The class mark is that stand-in, and choosing the mid-point is a bet that the observations inside each class balance around it. The chapter proves the bet is not always won on its own data: the same thirty marks give 59.3 before grouping and 62 after. The gap is not an error in the arithmetic — it is the price of grouping, and it is exactly the amount by which the classes fail to balance.
What you should be able to do
- Compute the mean of an ungrouped frequency distribution as Σf·x ÷ Σf, laying the products out in a column
- Regroup a set of raw values into classes of a stated width, applying the convention that a value sitting exactly on a boundary is counted in the class above
- Compute the class mark of a class as the average of its two limits
- Explain why a grouped table cannot yield the exact mean, and name what has been discarded
- Compute the mean of a grouped distribution using class marks, and compare it with the ungrouped mean of the same data
- State the assumption the class mark encodes, and say when it is nearly true and when it fails
- Identify, for a given class, whether its observations sit above or below its mid-point, and predict the direction in which the grouped mean will err
Words to know
| Term | Definition in one line | First introduced |
|---|---|---|
| ungrouped data | observations still carrying their own individual values | printed in this chapter, §13.1, p. 171 |
| grouped data | observations condensed into classes, so only counts per class remain | printed in this chapter, §13.1, p. 171 |
| measures of central tendency | the numerical stand-ins for a whole distribution — mean, median, mode | printed in this chapter, §13.1, p. 171 |
| frequency | how many observations a value or a class holds | printed in this chapter, §13.2, p. 171 |
| class interval | one of the ranges the data is cut into | printed in this chapter, §13.2, p. 173 (Table 13.2) |
| class mark | a class's mid-point, adopted as the value of every observation it holds | printed in this chapter, §13.2, p. 173 |
| mid-point | the same quantity, named for its position rather than its role | printed in this chapter, §13.2, p. 173 |
| upper class limit | the top end of a class | printed in this chapter, §13.2, p. 173 |
| lower class limit | the bottom end of a class | printed in this chapter, §13.2, p. 173 |
| summation | the operation the capital sigma sign stands for | printed in this chapter, §13.2, p. 171 |
| exact mean | the mean computed before any grouping | printed in this chapter, §13.2, p. 174 |
| approximate mean | the mean computed from class marks | printed in this chapter, §13.2, p. 174 |
| balance point of a class | whether a class's observations sit above or below its mid-point | an added phrasing; the chapter reasons about this without naming it |
Where people slip up
- "One of 59.3 and 62 is a mistake." Neither is. They are answers to two different questions, because after grouping the marks 10 and 20 no longer exist in the table — only "two students somewhere in 10–25" does.
- "Grouping always pushes the mean up." It pushed it up here because most classes happened to be bottom-heavy. Class 25–40 in this very data leans the other way. The direction is a fact about the particular data, not about grouping.
- "The class mark is the average of the observations in the class." It is the average of the class's two limits. Whether it equals the average of the observations is exactly the thing being assumed, and in Example 1 it does not.
- "A mark of 40 could go in either 25–40 or 40–55." Not once the convention is fixed. Every boundary value goes to the class above, or the same student would be counted twice and the frequencies would not total 30.
- "Wider classes are simpler, so use them." Wider classes discard more position information, so the stand-in has further to stretch. Width buys tidiness with accuracy.
- "Σf·x ÷ Σf is a new formula for grouped data." It is the same formula as for ungrouped data, run on a table whose x column has been replaced by stand-ins. Nothing about the mean changed; the data did.
Ask your teacher a person
Your teacher reads this and writes back, usually within a day. For an instant answer, use Ask the video in the sidebar.
Your class sees the question and the answer. Only your teacher sees that it was you.
No questions on this topic yet.
Worked answers: Exercise 13.1 · Exercise 13.2 · Exercise 13.3
Transcript1,734 words
Thirty students sat a paper. Here are their marks. Work out the average and you get 59.3. Now here are the same thirty marks, sorted into six bands. Work out the average from this table and you get 62. Same students. Same paper. Nobody changed a single mark. So one of these two numbers is a mistake. No. Neither of them is. They are answers to two different questions, and the 2.7 between them is not an error. It is the price of the second table.
By the end of this you will be able to say what that price buys, and predict which way it will fall. Start with the marks as they were recorded. Thirteen different numbers appear, from 10 up to 95. But thirty students sat the paper, and a number's frequency is how many of them scored it. Four students got 40. Three got 92. One got 95. That distinction is the whole of this table: thirteen values, thirty observations.
And here is what an average is, before any formula. Lay the thirty marks along a line and hang an equal weight at each one. The average is the point the line balances about. Slide the pivot until everything above it exactly cancels everything below. It settles at 59.3, and it settles there exactly. Not close to it. Every student's own mark took part in finding that point. Thirty marks fit on a page. Thirty thousand do not.
Real data arrives in quantities nobody can look at, so it gets condensed: cut the range into bands and report how many landed in each. You lose detail and you gain a shape you can see in one glance. That is a good trade, and it is made everywhere. But be exact about what is being given up. Once the marks are in bands, the table says how many students were in each band.
It does not say what any of them scored. Those individual numbers are not hidden somewhere in the table. They are gone. And the average we just found was built out of nothing else. Cut these thirty into bands 15 wide. 10 to 25, 25 to 40, 40 to 55, 55 to 70, 70 to 85, 85 to 100. Now let every mark fall into the band it belongs to.
The 10 and the 20 go into the first band. Two students. The three on 36 go into the second. Three students. Keep going, and the six bands hold 2, 3, 7, 6, 6 and 6. Those add to 30, which they had better, because every student went somewhere and nobody went twice. That is the condensed table. Six numbers instead of thirty. These bands have a proper name: class intervals. Each one has a lower limit and an upper limit, and the distance between them is its width.
Except that one thing in that sorting was not automatic. Four students scored exactly 40, and 40 is where two bands meet. It is the top of 25 to 40 and the bottom of 40 to 55. So which band are those four in? The answer is not discovered. It is decided, once, and then applied to everybody. The rule taken here is that a mark landing on a boundary goes to the band above. So the four on 40 go up.
Watch what happens if the rule is turned over instead. The four on 40 drop into the second band, the four on 70 drop into the third, and the lowest mark of all, the 10, now belongs to no band at all, because nothing sits below the very first boundary. A completely different table, from the same thirty marks. And with no rule at all, the students on 40 and on 70 land in two bands each. The counts come to 38 for a class of 30.
So the convention is not decoration. It is what makes the table a table. Now the real problem appears. Ask the condensed table for the average and it cannot answer. Look at the third band: seven students, somewhere between 40 and 55. Somewhere. The table does not know where, and neither do you. To balance a line you need to know where to hang each weight, and this table gives you a count without a position.
The average cannot be computed from it. It has to be rebuilt. And rebuilding means putting a number back where a number used to be, for all thirty students, out of a table that no longer holds one. Here is the move. Appoint one stand-in value for every student in a band. The obvious candidate is the middle of the band, and it has a name: the class mark. For the band 10 to 25, that is 10 plus 25, halved. 17.5.
The six class marks are 17.5, 32.5, 47.5, 62.5, 77.5 and 92.5. Read the definition carefully, because it is where this whole video turns. The class mark is the average of the band's two edges. It is not the average of the students in the band. Nobody has checked that, and nobody can, because those marks are gone. Appointing the mid-point is a bet: that inside each band the students sit evenly enough around the middle to cancel out.
It is a reasonable bet. It is still a bet, and the rest of this is about how to price it. So build a new set of thirty numbers. Every student in the first band becomes a 17.5. Every student in the second becomes a 32.5. And so on down the six. Notice that not one student's own mark survives that. All thirty were replaced. In fact nobody in this room scored 17.5, or 32.5, or any of the six. Those values were never on anybody's paper.
Now ask the identical question we asked at the start. Hang thirty equal weights at these thirty positions, and find the point they balance about. Nothing about the method has changed. Only the data did. It balances at 62. Exactly 62. So here they are together. 59.3, from thirty marks each standing where its student put it. 62, from thirty stand-ins each standing where its band's middle is. Neither calculation is wrong. They are answers about two different sets of numbers.
One of those sets is the class. The other is what is left of the class after condensing. There are two useful words for this. Call 59.3 the exact mean, and 62 the approximate one. But approximate is doing a lot of quiet work in that sentence, and it is worth making it say something. The gap is 2.7 marks. Where did 2.7 come from? Take the bands one at a time, and ask two questions of each.
What did this band's students actually contribute? And what does the stand-in credit them with? First band: the 10 and the 20 really total 30. Two stand-ins at 17.5 credit them with 35. Over-credited by 5. Third band: four students on 40 and three on 50 total 310. Seven stand-ins at 47.5 credit them with 332.5. Over by 22.5. Do all six and you get five over-credits, totalling 91.5, and one band that goes the other way, by 10.5.
Net: 81 marks handed out that nobody earned. Share 81 among thirty students, and that is 2.7 each. So the gap is not an arithmetic slip anywhere. It is the exact amount by which the bands failed to balance around their middles, measured band by band and added up. Which brings us to the band that went the other way. The second band, 25 to 40, holds three students, and all three scored 36.
Their own average is 36. Their band's middle is 32.5. The stand-in is below them, so that band is under-credited, and its contribution to the gap is negative. That single band is enough to kill a very natural idea: that grouping pushes an average up. It went up here because most of these bands happened to be bottom-heavy. That is a fact about these marks, not about grouping. Take two hundred different sets of thirty and band them all the same way.
The average went up in 170 of them. It went down in 26. And in 4 it landed exactly right. Down and exact are both rarer, and both real. So the honest statement is that a band leans whichever way its students happen to sit, and the whole gap is those leans added together. Then the question a student should ask next: how bad can this get? There is a clean answer, and it comes straight from what a stand-in is.
Every student is replaced by the middle of their own band, so nobody is moved by more than half a band width. Thirty numbers each moved by under half a width cannot move their balance point by more than half a width either. In a thousand groupings that bound was never once broken, and the worst miss seen was 4.6 against a bound of 12.5. And the bound is not generous. Put every student on their band's lower edge and the average is out by exactly half the width, every time.
Which answers the other tempting idea: that wider bands are simpler, so use wider bands. Wider bands give the stand-in further to stretch. Across those same datasets the average miss climbed at every step, from about half a mark at width 5 to about 1.2 at width 25. Tidiness is bought with accuracy, at a rate you can measure. So the mid-point is not a formula. It is an assumption, and now you can say exactly what it assumes.
That inside each band, the students balance about the middle. When they do, grouping costs nothing at all. Line ten observations up so each one sits on its own band's mark, group them, and the two averages are the same number. When they do not, the gap is the sum of the leans, and it points whichever way the leans do. Nothing about the average changed here. The same balancing, on the same kind of line.
What changed was the data underneath it: thirty real marks became thirty stand-ins. So when you are handed a banded table and asked for an average, you can give one. Just know that you are answering for stand-ins, and that nobody scored the class mark.
Where this fits
Taken from the notes each video was made from, not from the reading order — these are the ideas this one rests on and the ones that later rest on it.
Comes up again in
- The direct method, and where it becomes unwieldyClass 10 · Ch 13, Statistics
- Shifting the origin: why guessing a centre cannot change the answerClass 10 · Ch 13, Statistics
- Finding the busiest class, then placing the mode inside itClass 10 · Ch 13, Statistics
- Running totals, and converting between the two cumulative tablesClass 10 · Ch 13, Statistics
- Locating the middle class and interpolating across itClass 10 · Ch 13, Statistics
Either side of this one
- When a piece has been scooped out: apparent capacity against actualClass 10 · Ch 12, Surface Areas and Volumes