PrepShorts · Study sheet · Class 10 Mathematics · Chapter 13, Statistics
This video could not be loaded. Reload the page to try again.
Sign in with Google15 min.
Keep your place in this chapter — sign in, it’s free.Sign in
Thirty scripts give a mean of 62, a median of 62.5 and a mode of 52, and nobody has made an arithmetic slip. The three are not three attempts at one quantity - they are exact answers to three different questions, and the whole difference between them is how much of the table each one actually reads.
The idea
"Which average is the right one?" has no answer, because the three are answers to three different questions — and which question you are asking is decided by what the number will be used for, not by the data. The differences all follow from one thing: how much of the data each measure actually consults. The mean reads every observation, which is why it compares distributions well and why a single far-out value can drag it. The median reads only positions, so a value's size cannot reach it once it is past the middle. The mode reads three adjacent counts and nothing else. Lopsidedness is therefore one thing that drives them apart, but not the only one: on the chapter's own opening dataset the mean and median sit half a mark apart while the mode is ten marks away, because the busiest class is simply not where the weight of the distribution lies. That is also why the chapter's own thumb rule linking the three, which works to within a rounding error on some of its exercises, misses by six per cent on the very dataset the chapter opened with.
What you should be able to do
- Compute all three measures for one grouped distribution and set them side by side
- Explain which observations each measure depends on, and which it ignores
- Given a stated purpose, choose the appropriate measure and justify the choice
- Predict whether the three will be close or far apart from the shape of the distribution
- Show, on a given table, that removing an extreme group moves the mean but not the mode
- Apply the empirical relation between the three, and state the accuracy it can and cannot be trusted to
- Recognise which distributions the chapter's formulas are not equipped for
Words to know
| Term | Definition in one line | First introduced |
|---|---|---|
| mean | the total of the observations shared equally among them | printed in this chapter, §13.1, p. 171 |
| median | the value of the middle observation once the data is ordered | printed in this chapter, §13.1, p. 171 |
| mode | the value occurring most often | printed in this chapter, §13.3, p. 183 |
| measures of central tendency | the collective name for the three | printed in this chapter, §13.1, p. 171 |
| extreme values | observations far from the bulk, which pull the mean | printed in this chapter, §13.4, p. 197 |
| empirical relationship | the approximate rule tying the three together | printed in this chapter, §13.4, p. 197 |
| modal class | the busiest class, all the mode ever looks at | printed in this chapter, §13.3, p. 184 |
| median class | the class the middle observation falls in | printed in this chapter, §13.4, p. 193 |
| skew | a distribution's lopsidedness, which is what drives the three apart | an added term; the chapter describes lopsided data without naming the property |
| resistance | a measure's tendency not to move when a far-out group is added or removed | an added term |
Where people slip up
- "One of the three is the real average and the others are approximations." All three are exact answers to their own questions. Nothing in the chapter ranks them.
- "The mean is always the best because it uses all the data." Using all the data is precisely why one far-out group can move it, as Exercise 13.2 Q4 shows. The chapter says outright that the mean can stop representing the data in such cases.
- "Mode is always less than median is always less than mean." Exercise 13.3 Q6 has that order; Exercise 13.2 Q1 has the reverse; and Example 1's data has the mean below the median with the mode far off to one side. Three of the chapter's own datasets, three different orderings.
- "Use 3 × median = mode + 2 × mean to find whichever average is missing." On Example 1's data that rule is out by 11.5 on a quantity of about 180. It is a sanity check, not a formula, and the chapter calls it empirical for a reason.
- "The median ignores extreme values, so extreme values do not matter." They do not move the median much, but they are still part of the data and may be the most important part of it. Choosing the median is a decision to set them aside, and that decision should be conscious.
- "The mode tells you where most of the data is." It tells you where the single busiest class is. In Example 1's data the busiest class is 40–55 while eighteen of the thirty students score above 55.
- "These formulas work on any table." The mode formula needs equal class widths and the chapter declines unequal ones; the median formula is stated for equal widths too; and both need continuous classes. The mean alone is free of all three restrictions, which is why Example 3 could use unequal widths.
Ask your teacher a person
Your teacher reads this and writes back, usually within a day. For an instant answer, use Ask the video in the sidebar.
Your class sees the question and the answer. Only your teacher sees that it was you.
No questions on this topic yet.
Worked answers: Exercise 13.1 · Exercise 13.2 · Exercise 13.3 · this video explains Exercise 13.3 Q1
Transcript2,036 words
Thirty scripts, marks out of a hundred, sorted into six classes. Two students in the lowest class, three in the next, then seven, six, six and six. One table. Now put all three averages on it. The mean comes out at sixty-two exactly. The median comes out at sixty-two point five. And the mode comes out at fifty-two. Fifty-two, sixty-two, sixty-two point five. Three numbers, one table, and nobody has made an arithmetic slip.
So which one is the average of this data? That question has no answer, and this video is about why. The three are not three attempts at one quantity. They are exact answers to three different questions, and which question you are asking is decided by what the number is for. Everything that follows comes from one difference between them: how much of the table each one actually reads. Start with the mean, because it is the greediest of the three.
The mean reads every observation, without exception. Every count in the table enters the total, and every class mark is multiplied by its own count. There is a way to test that rather than assert it. Take a table, add one observation to a class, and see whether the answer moves. Do that for every class in turn. On a table like this one, every single bump moves the answer, because every count is standing in the arithmetic.
Nothing in the table is ignored, and nothing in it is rounded away. That is the mean's whole character, and it is the reason for both of its properties. It compares two distributions well, because it has looked at all of both of them. And it can be dragged, because one far-out group is part of all of them too. Here is the drag, on a real table. Thirty-five regions, counted by students per teacher, in classes five wide.
Three regions, then eight, nine, ten, three - and then two empty classes, and two regions right out at the end. With all thirty-five regions the mean is twenty-nine point two one. Now set those two far-out regions aside and work with the remaining thirty-three. The mean falls to twenty-seven point eight. One and four tenths, moved by two regions out of thirty-five. That is not a mistake in the arithmetic. It is the arithmetic working exactly as designed.
The mean was asked to represent every region, and it did. Whether that is a virtue depends entirely on what you wanted the number for, which is the whole point. The median reads something else. It reads positions, and never sizes. Once an observation is past the middle, the median knows it is there and does not know how big it is. You can test that too, and the test is sharper than the last one.
Take the same table and push the far-out group further out - three whole class widths further out - without changing a single count. Nobody has arrived and nobody has left. The data has the same shape and the same total. Only the size of those far-out observations has changed. Run all five of our tables through that, and the mean moves every time. The median does not move at all. Not on one of the five.
That is not resistance to being dragged a little. That is the median being structurally unable to see the change. Positions moved nowhere, so the middle stayed exactly where it was. Now, is that steadiness good? It depends, and the honest answer has two halves. When your far-out values are errors, or one-offs, or the sort of thing that will not happen again, ignoring their size is exactly what you want.
A typical wage, a typical house price, a typical time taken - these are all median questions, because a handful of enormous values would otherwise decide the answer for everybody. But the cost is real. Those far-out values did not go away. They are still in the data, and sometimes they are the most important part of it. Choosing the median is a decision to set them aside. It is a perfectly good decision. It should just be a conscious one, rather than something that happened because a formula was convenient.
The median ignoring extreme values does not mean extreme values do not matter. The mode is the narrowest reader of the three, and by a long way. It looks at the busiest class and at its two immediate neighbours. Three adjacent counts. Everything else in the table is invisible to it. Test that the same way. Change a count anywhere outside those three bars, by as much as you like, and as long as the busiest bar is still the busiest, the mode does not move.
Which explains something on the regions table. We set two far-out regions aside and the mean shifted by one and four tenths. The mode was thirty point six two five before, and thirty point six two five after. Not approximately. Exactly, to the last digit. The busiest class was thirty to thirty-five, its neighbours held nine and three, and none of those three numbers was touched. The mode had nothing to notice.
So here is the map, and it is a map from purposes, not from data. If the number will be added, or shared out, or compared with another distribution, you want the mean, because it is the only one built out of the whole of the data. Total pay divided among the staff. Average rainfall across regions. Anything you intend to multiply back up by the count. If the number has to survive a few extreme values, you want the median.
A typical income, a typical waiting time, a typical price. And if the question is which single group is busiest - which shoe size to stock, which class to timetable a room for, which band most customers fall into - you want the mode. Notice that none of those three sentences mentions the data at all. They are all about the job the number has to do. That is the answer to which average is right: name the job first, and the measure follows.
Sometimes the choice barely matters, and it is worth knowing when. Sixty-eight households, counted by electricity used, in classes twenty wide. Four, five, thirteen, twenty, fourteen, eight, four. That shape rises to a peak and comes down again at about the same rate. It is close to balanced about its busiest class. Its mode is a hundred and thirty-five point seven seven. Its median is a hundred and thirty-seven, exactly.
Its mean is a hundred and thirty-seven point zero six. The whole spread of the three is one point two nine. On numbers of that size, that is a rounding error. For a balanced distribution the three questions nearly coincide, so choosing between them is nearly free. Which is a good habit to have: look at the shape before you agonise over the choice. Now back to the thirty scripts, where they do not coincide at all.
Fifty-two, sixty-two, sixty-two point five. The mean and the median are half a mark apart. The mode is ten marks below the mean, and ten and a half below the median. Why is the mode so far out on its own? Because of what it reads. The busiest class here is forty to fifty-five, holding seven. And the busiest class is simply not where the weight of this distribution sits.
Count the students above fifty-five - above the top of the busiest class. Eighteen of them. Against twelve below. The mode is answering the question it was asked, faithfully: which single class is busiest. Forty to fifty-five is. It was never asked where the bulk of the data is, and it does not know. That is not lopsidedness dragging the mean. It is a peak in one place and a mass of data in another.
There is a rule students memorise, and it is worth destroying properly. The rule says mode below median below mean. Here is a table where that is exactly right. A hundred family names, counted by how many letters they have, in classes three wide. Its mode is seven point eight eight, its median eight point zero five, its mean eight point three two. Mode, median, mean. Just as promised. And here is a table where the order is precisely reversed. Eighty patients by age, in classes ten wide.
Mean thirty-five point three seven five, median thirty-five point eight seven, mode thirty-six point eight two. Mean, median, mode. The exact opposite order. And the thirty scripts we started with run a third way again: mode, then mean, then median. Three tables, three orderings, and none of them a special case invented to embarrass you. There is no ordering to memorise. There is only what each measure reads. There is also a thumb rule that ties the three together. Three times the median is roughly the mode plus twice the mean.
Roughly is doing a great deal of work in that sentence, so let us find out how much. Run it on all four tables and measure the miss. On the patients, three times the median is a hundred and seven point six one, against a hundred and seven point five seven. Out by four hundredths. That is nought point nought four per cent. On the households, out by one point one one. Nought point three per cent.
On the family names, out by nought point three seven. One point five per cent. And on the thirty scripts, three times the median is a hundred and eighty-seven point five, against a hundred and seventy-six. Out by eleven and a half. Six per cent. Look at where the rule works and where it fails. It is excellent on the three tables whose measures were already close together, and poor on the one where they were far apart.
Which is to say it is least reliable exactly where you would most want to lean on it. One caution, because it would be easy to overclaim. Across those first three, the size of the miss does not follow the spread in order. The safe statement is that the rule falls apart once the three separate, not that you can predict the error from how far apart they are. Use it to check an answer. Never use it to compute a missing one.
Last, the conditions these formulas come with, and they are not arbitrary. The mode formula and the median formula both want classes of equal width, and both want classes that meet. Meeting matters because both of them place an answer inside a class by assuming the observations are spread evenly across it. A gap between classes is a stretch of the number line the assumption says nothing about. The mean is free of all three restrictions.
Give it classes of ten, ten and thirty and it answers perfectly happily, because it never places anything inside a class. It only ever uses the class marks. The step-deviation shortcut does need equal widths, but that is the shortcut's requirement and not the mean's. And one honest refusal. If two classes tie for busiest, the mode construction has no answer, and neither does the formula. The construction joins corners of the busiest bar to its neighbours and crosses them. With a flat top, those two lines are parallel and never meet.
A refusal is the right answer there. A number would be worse. One table, three numbers, and the whole difference between them is what each one reads. The mean reads every observation, which is why it compares distributions and why it can be dragged. The median reads positions and never sizes, which is why a far-out value can be pushed three class widths further out without moving it a hair.
The mode reads three adjacent counts, which is why it did not flinch when two regions left the table. None of the three is the real average with the other two as approximations. So do not ask which average is right. Ask what the number is for. Then the measure is not a choice at all. It is a consequence.
Where this fits
Taken from the notes each video was made from, not from the reading order — these are the ideas this one rests on and the ones that later rest on it.
Builds on
- Dividing through by the class width to shrink the arithmetic furtherClass 10 · Ch 13, Statistics
- Finding the busiest class, then placing the mode inside itClass 10 · Ch 13, Statistics
- Locating the middle class and interpolating across itClass 10 · Ch 13, Statistics
Either side of this one
- Equally likely outcomes, and the everyday cases where that failsClass 10 · Ch 14, Probability