PrepShorts · Study sheet · Class 11 Mathematics · Chapter 13, Statistics
Chapter 13 · Statistics
Working it out about the mean and about the median, for a plain list
This video could not be loaded. Reload the page to try again.
Sign in with Google14 min.
Keep your place in this chapter — sign in, it’s free.Sign in
Slide a reference point along a data set, and the total distance to every reading falls only while more readings lie ahead of it than behind - which is why nothing beats the median, not even the mean.
The idea
§13.4.1 prints four steps, and they are one idea carried out twice: fix the centre, then average the distances to it. What the four steps hide is that the first of them is a decision with a consequence. Choosing the median rather than the mean never raises the answer and usually lowers it, because a median is a point at which the total distance is already as small as it can be — and the reason is a counting argument, not a computation. The chapter's own three worked examples cannot show this, because in two of them the mean and the median happen to coincide; the third is where the two centres separate.
What you should be able to do
- Carry out the four steps of §13.4.1 for the mean deviation about the mean of a given list
- Carry out the same four steps about the median, including the sorting step the mean does not require
- Apply the odd-count and even-count median rules correctly
- Compute the mean deviation for each of the four ungrouped data sets set in Exercise 13.1
- Explain, by a counting argument, why shifting the reference point away from a median cannot decrease the total distance
- Recognise data sets where the mean and the median coincide, and say why those cannot illustrate the difference between the two mean deviations
- Round a recurring quotient to two decimal places as the chapter's examples do
Words to know
| Term | Definition in one line | First introduced |
|---|---|---|
| ungrouped data | observations listed one by one, with no frequency column | printed in this chapter (§13.4.1 heading, p. 260) |
| mean deviation about the mean | the average distance of the observations from their own mean | printed in this chapter (§13.4.1, p. 260) |
| mean deviation about the median | the average distance of the observations from their median | printed in this chapter (§13.4.1, p. 260) |
| median | the middle entry of the sorted list, or the average of the two middle entries | printed in this chapter (§13.1, p. 258) |
| ascending order | the sorted arrangement the median step requires | printed in this chapter (§13.4.1, p. 262) |
| absolute value | the size of a difference with its sign removed | printed in this chapter (§13.4, p. 260) |
| M | the chapter's symbol for the median throughout this chapter | printed in this chapter (§13.4.1 Note, p. 261) |
| minimising point | the reference value that makes the total distance smallest | an added term; the chapter states the resulting inequality in §13.4.3 but gives it no name |
Where people slip up
- "The mean deviation about the mean and about the median are the same thing computed twice." They are different numbers for most data. On the Example 3 list they are about 5.39 and about 5.27.
- "The median version is bigger because the median is smaller." The direction has nothing to do with which centre is numerically smaller. The median version is never the larger of the two, whichever way round the two centres sit.
- "You can find the median without sorting." Every median step in these examples begins by arranging the data in order, and the chapter's Example 5 points out when the given data is already ordered so that step can be skipped.
- "The median of a twelve-value list is the sixth value." It is the average of the sixth and the seventh. Exercise 13.1 items 3 and 4 are set precisely to catch this.
- "5.27 is the exact answer." It is 58/11 rounded. Keep the fraction in view and round only at the end, or a student who rounds early will report 5.3 and believe the difference from 5.27 is the book's error.
- "Dropping the sign is the same as ignoring the negative observations." The distances of the below-centre observations are kept in full; only the direction is discarded. In Example 1 the largest single distance, 5, comes from an observation below the mean.
Ask your teacher a person
Your teacher reads this and writes back, usually within a day. For an instant answer, use Ask the video in the sidebar.
Your class sees the question and the answer. Only your teacher sees that it was you.
No questions on this topic yet.
Worked answers: Exercise 13.1 · Exercise 13.2 · Miscellaneous Exercise · this video explains Exercise 13.1 Q1, Exercise 13.1 Q2, Exercise 13.1 Q3, Exercise 13.1 Q4
Transcript2,019 words
The recipe for a mean deviation is four steps long, and it is worth writing them out once. One: fix a centre. Two: subtract that centre from every reading. Three: throw away the signs. Four: average what is left. That is it, and steps two, three and four are simply the definition of a mean deviation, carried out in order. The recipe adds no mathematics at all. It only tells you what order to do the work in.
Which means the only interesting thing in the whole procedure is step one. Step one says fix a centre, and it does not say which. In practice the two that get used are the mean and the median. That looks like a matter of taste, and it is presented as one. It is not. Choosing the median rather than the mean changes the answer, and it changes it in a direction that is fixed in advance.
By the end of this video we will have measured that direction on four hundred records of odd length and four hundred of even, none of them chosen by anybody, and we will know why it cannot come out the other way. But we should start by running the recipe properly, because the first two lists anyone reaches for cannot show any of this. Eight readings: six, seven, ten, twelve, thirteen, four, eight and twelve.
Step one, the mean. They total seventy-two, there are eight of them, so the centre is nine. Step two, subtract nine from each, and step three, drop the signs. The eight distances are three, two, one, three, four, five, one and three. Notice the largest single distance, the five, comes from the reading of four, which sits below the centre. Dropping the sign is not the same as ignoring the readings below the middle; their distances are kept in full.
Step four. Those eight distances total twenty-two, and twenty-two over eight is two point seven five. That is the mean deviation about the mean. Now do it again about the median, and watch what happens. Sorting the eight readings gives four, six, seven, eight, ten, twelve, twelve, thirteen. Eight is an even count, so the median is the average of the fourth and fifth entries, which are eight and ten.
The median is nine. The same nine the mean gave. So step two onwards is word for word the work we already did, and the answer is two point seven five again. This list cannot tell you anything about the difference between the two centres, because on this list there is no difference to see. A demonstration run on it demonstrates nothing. The obvious next move is a bigger list, and it does not help.
Twenty readings, totalling two hundred, so the mean is a comfortable ten. Their twenty distances from ten add to a hundred and twenty-four, and dividing by twenty gives six point two. That list is genuinely useful for one thing: twenty distances is where a student's arithmetic starts to slip, and a round centre is what keeps it survivable. But sort those twenty and the tenth and eleventh entries are nine and eleven, so the median is ten.
The mean again. Two lists, two chances to see the two centres come apart, and both of them wasted. So here is a list where they do come apart. Eleven readings: three, nine, five, three, twelve, ten, eighteen, four, seven, nineteen and twenty-one. The median step always begins by sorting, and sorted they run three, three, four, five, seven, nine, ten, twelve, eighteen, nineteen, twenty-one. Eleven is odd, so there is a genuine middle entry rather than a pair to average, and the sixth one along is nine.
Now the distances from nine, taken in the order the readings were given: six, nought, four, six, three, one, nine, five, two, ten and twelve. One of them is nought, because nine is itself one of the readings. They total fifty-eight, and fifty-eight over eleven is the answer. As a decimal that is five point two seven. Now the same eleven readings about their mean instead. They total a hundred and eleven, and there are eleven of them, so the mean is a hundred and eleven over eleven.
That is ten and a bit, and it is not a whole number, which means every one of the eleven distances is going to be a fraction. Written over eleven they are seventy-eight, twelve, fifty-six, seventy-eight, twenty-one, one, eighty-seven, sixty-seven, thirty-four, ninety-eight and a hundred and twenty. Those numerators add to six hundred and fifty-two, so the total distance is six hundred and fifty-two elevenths. Divide by the eleven readings and the mean deviation is six hundred and fifty-two over a hundred and twenty-one.
Turning that fraction into a decimal is where a careful student and a careless one part company. Six hundred and fifty-two over a hundred and twenty-one is five point three eight eight and so on. Cut it off after two places and you get five point three eight. Round it to two places and you get five point three nine. Those are different numbers, and only one of them is right.
The median's answer happened not to care: fifty-eight over eleven reads five point two seven whether you round it or cut it. So the habit that got you through the first answer will quietly fail you on the second. Keep the fraction on the page and round once, at the very end. Now put the two answers side by side. About the median, five point two seven. About the mean, five point three nine.
The mean's answer is the larger of the two, by fourteen over a hundred and twenty-one, which is about nought point one two. A small difference, and on one list a small difference proves nothing at all. The question worth asking is whether the direction was an accident. It was not, and the reason has nothing to do with arithmetic. It is a counting argument. Forget both centres for a moment and think of the reference point as something you can slide along the line.
Put it somewhere, add up all eleven distances, and then nudge it one step to the right. Every reading that is behind you now sits one further away. Every reading still ahead of you sits one closer. So the total distance changes by the step, times the number behind, minus the number ahead. Try it at seven, where four readings are behind and six ahead. The total falls by one.
Try it further out, at four, with two behind and eight ahead: the total falls by five. And on the other side, at twelve, with seven behind and three ahead, the total rises by five. The sliding is always worth doing towards the side with more readings on it. That is not a description of what happened, it is an identity, and it can be checked. Take a hundred and twenty records nobody chose, every whole reference point from nought to fifty-seven, and strides of one, two and three.
For every step that jumps over no reading, compare what the total distance actually did against what the counting predicted. The number of disagreements is nought. Not nearly nought. So the rule holds wherever you put the reference point, on any list at all, and we can use it to find where the total stops falling. The total keeps falling as long as more readings lie ahead than behind. It stops falling exactly when those two counts balance.
And a point with as many readings below it as above it is precisely what a median is. That is the whole argument. Not a computation, a count. On the eleven readings the balance sits at nine, with five below and five above, and a step to the right from there raises the total by one rather than lowering it. Search every whole number from nought to sixty and the number of them that tie nine is exactly one.
Across four hundred odd-length records, the number with exactly one whole point that ties is all four hundred of them. With an even count something different happens, and it is the key to everything left. Take ten readings whose two middle entries are forty-six and forty-nine. The median is the average of those, forty-seven and a half. Now measure the total distance at forty-six: seventy. At forty-seven and a half: seventy.
At forty-nine: seventy. The total does not fall anywhere between those two middle entries, because between them the counts already balance. Step outside and it rises at once: seventy-four at forty-five, seventy-two at fifty. So with an even count the best reference point is not a point at all. It is a whole stretch, and the median is just one value in it. That stretch is what decides whether the two answers come out equal.
If the mean happens to land inside it, the mean is doing exactly as well as the median and the two answers are identical. Check it: across four hundred even-length records, the number where the two answers agree is a hundred and seventy-five. The number whose mean lands inside the stretch is a hundred and seventy-five. And the number belonging to one of those groups but not the other is nought.
That is not a correlation, it is the same condition counted twice. With an odd count the stretch shrinks to a single reading, so the mean can only land in it by being equal to the median. Across four hundred odd-length records the two answers agree three times, and all three of those have a mean equal to their median. Which finally explains the first two lists properly. They did not fail because their mean equals their median; that is just the odd way of saying it.
They failed because the mean landed in the flat stretch. Now the direction, measured rather than argued. Across four hundred odd-length records and four hundred even-length ones, the number where the median's answer is larger than the mean's is nought, and nought. Stronger: search every whole reference point from nought to sixty on every one of those eight hundred records, and the number of records where anything at all beats the median is nought.
The median is not merely a good choice; nothing is better. And the reason is not that the median is the smaller number, which is the explanation students usually reach for. Of those four hundred odd-length records, two hundred and four have a median below their mean and a hundred and ninety-three have one above it, with three where the two are equal. In the group where the median is the larger number, the count whose median answer comes out larger is nought.
In the group where it is the smaller number, also nought. Which side the median sits on has nothing to do with it. Four more lists to finish, two asked about the mean and two about the median. The first has a mean of ten and an answer of three. Its median is nine and a half, which is not the same number at all and yet gives the same answer of three, because ten sits inside its flat stretch.
The second has a mean of fifty and an answer of eight point four; its median of forty-seven would have given eight, which is smaller, exactly as promised. The third asks for the median and has twelve readings, so the even rule runs: the median is thirteen and a half. That is not a whole number and it is not one of the readings, and both of those facts are the point of setting it.
Its distances total twenty-eight, so the answer is twenty-eight over twelve, about two point three three. The fourth has ten readings and a median of forty-seven and a half, again neither whole nor one of the readings, with distances totalling seventy and an answer of exactly seven. Its mean of fifty would have given seven point two. Four lists, and the median never once did worse.
Where this fits
Taken from the notes each video was made from, not from the reading order — these are the ideas this one rests on and the ones that later rest on it.
Builds on
- Signed deviations cancel to nothing, so the sign has to be discardedClass 11 · Ch 13, Statistics
Comes up again in
- The same procedure once the data arrive already groupedClass 11 · Ch 13, Statistics
- Where this measure breaks down, and why another was neededClass 11 · Ch 13, Statistics