PrepShorts · Study sheet · Class 7 Mathematics · Chapter 5, Connecting the Dots...
Chapter 5 · Connecting the Dots...
Outliers, and why the median survives them
This video could not be loaded. Reload the page to try again.
Sign in with Google10 min.
Keep your place in this chapter — sign in, it’s free.Sign in
Two families, eleven people, every height measured. The average says the second family is shorter. Then go and look at them.
The idea
The mean is built out of the total, so every value votes with a weight equal to its own size; one value far from the rest can therefore drag the mean away from where almost all the data actually sits. The median is built out of position, so a far-out value votes once no matter how far out it is. That single structural difference is the whole lesson — and it pays a bonus: comparing the two numbers tells you which way the data leans. Mean below median means something is pulling from the low end; mean above means something is pulling from the high end.
What you should be able to do
- Find the median of a sorted list, for an odd and for an even number of values
- Explain why an even-sized list forces you to average the two values that share the middle
- Identify an outlier in a small data set and say which end it lies at
- Explain, from how each is computed, why an outlier moves the mean far more than the median
- Predict the direction of the shift: a low outlier pushes the mean below the median, a high one above it
- Recompute both summaries with the outlier removed and describe what changed
- Decide, for a given data set, which of the two better represents it, and justify the choice
- Use the phrase measures of central tendency for the pair
Words to know
| Term | Definition in one line | First introduced |
|---|---|---|
| median | the middle value of a sorted list, or the halfway point between the two that share the middle when the count is even | printed in bold in §5.2, Part II, p.106 |
| outlier | a value sitting far away from the rest of the set | printed in bold in §5.2, Part II, p.107 |
| sorted data | the same values rearranged from smallest to largest | printed in §5.2, Part II, p.106 and in the SUMMARY, Part II, p.134 |
| central tendency | the pull of a set of values towards one central value | printed in §5.2, Part II, p.109 |
| measures of central tendency | the collective name this chapter gives to the mean and the median | printed in §5.2, Part II, p.109 |
| mean | the total of the values divided by how many there are | printed in bold in §5.2, Part II, p.100 |
| average | the chapter's everyday word for the mean | printed in bold in §5.2, Part II, p.100 |
| deviate | to sit away from the rest — the chapter's verb for what an outlier does | printed in §5.2, Part II, p.107 |
| resistant | said of a summary an outlier cannot move much | an added term; this chapter shows the property on p.108 without giving it a name |
| skew | the lean a distribution takes when one tail is longer | an added term, not printed in this chapter, which describes the same lean as the mean shifting towards the outlier |
Where people slip up
- "The median is unaffected by outliers." Too strong, and the chapter is careful not to say it — its own wording is that outliers did not move it much. Deleting the 118 from Poovizhi's family moves the median from 170 to 171.5, because removing a value also changes which position is the middle. The right statement is comparative: the median moved by 1.5 cm and the mean by 10.55 cm on the same edit.
- "An outlier is a mistake in the data." Sometimes it is; here it is a real child. The chapter's outliers — a young sibling, a student who read forty stories — are all genuine values. The problem is with the summary, not the observation.
- "Take the middle value of the list as written." The list has to be sorted first, and the chapter's two examples both come pre-scrambled precisely so that the sorting step cannot be skipped.
- "With an even count you pick either middle value." You average them, and the reason is worth a sentence: the median is defined so that as many values sit below it as above, and with an even count no listed value does that job.
- "The median is always the better summary." Not the claim being made. The chapter's own newspaper example has no serious outlier and the two summaries agree; there the mean is fine, and it uses the whole data set rather than just the middle of it. Which to prefer is a judgement about the data.
- "Mean below median means the data is small." It means something is pulling from the low end. Direction of pull, not size — this is the specific confusion the three-case comparison on p.108 exists to break.
- "Any value at the end of the range is an outlier." Every set has a largest and a smallest value. What makes the 118 and the 40 outliers is the gap between them and everything else, which is exactly what a dot plot shows and a sorted list does not.
Ask your teacher a person
Your teacher reads this and writes back, usually within a day. For an instant answer, use Ask the video in the sidebar.
Your class sees the question and the answer. Only your teacher sees that it was you.
No questions on this topic yet.
Worked answers to this chapter’s exercises · this video explains Figure it Out · 4 Q5
Transcript1,447 words
Two families, and everybody in each of them measured. The first family is six people: a hundred and sixty nine centimetres, a hundred and seventy three, a hundred and fifty five, a hundred and sixty five, a hundred and sixty, and a hundred and sixty four. The second is five: a hundred and seventy, a hundred and seventy three, a hundred and sixty five, a hundred and eighteen, and a hundred and seventy five.
Average them out. The first family comes to a hundred and sixty four point three. The second comes to a hundred and sixty point two. So the second family is the shorter one. That is the arithmetic, and the arithmetic is correct. Now go and look at the second family again. A hundred and sixty five. A hundred and seventy. A hundred and seventy three. A hundred and seventy five.
Every single one of those is taller than a hundred and sixty point two. Four of the five people are taller than the number that is meant to describe them. That is not a rounding problem. It is a summary sitting below four fifths of what it summarises. And the calculation behind it is still perfectly right. So the fault is not in the arithmetic. It is that this particular number is the wrong tool for this particular job.
So try a different tool, a very old one. Write the five heights out in order, smallest to largest. A hundred and eighteen. A hundred and sixty five. A hundred and seventy. A hundred and seventy three. A hundred and seventy five. Now walk in from both ends at the same time, and see where you meet. You meet at a hundred and seventy. Two people below it, two above it. That is the middle value, and it is called the median.
And notice what it never asked. It did not ask how far away the shortest person was. Only that they were the shortest. Do the same for the first family and you hit a snag. Six people. In order: a hundred and fifty five, a hundred and sixty, a hundred and sixty four, a hundred and sixty five, a hundred and sixty nine, a hundred and seventy three. Walk in from both ends and you do not land on anybody. You land between two people.
Check all six if you like. Not one of them has as many below it as above it, and with an even count none of them ever can. So the median is the halfway point between the two that share the middle. A hundred and sixty four point five. Which is nobody's height, and does not need to be. So the medians are a hundred and sixty four point five, and a hundred and seventy. The median says the second family is the taller one.
Two summaries. The same eleven people. Opposite answers. And you can see the reason sitting on its own, down at a hundred and eighteen. That is a real person. A much younger child. Nobody wrote the number down wrong. A value sitting that far from the rest has a name. It is called an outlier. But be careful about what earns that name. It is not being the smallest. Every set has a smallest.
It is the gap. From a hundred and eighteen up to the next person is forty seven centimetres, and the whole family only spans fifty seven. One gap is over eighty percent of the entire range. In the first family, the largest gap between neighbours is five. Now the part that matters. Why does one summary move so much more than the other? It is how each one is built.
The mean is the total, shared out. Every value goes into that total at its own size, so a value far away pulls hard. The median is a position. Every value counts once, wherever it happens to be standing. Watch what that means. Take the youngest child and imagine them shorter still. A hundred. Then fifty. Then zero. The mean drops to a hundred and fifty six point six. Then a hundred and forty six point six. Then a hundred and thirty six point six. Twenty three and a half centimetres of drift, and still going.
The median does not move at all. A hundred and seventy, a hundred and seventy, a hundred and seventy, a hundred and seventy. One value slid a hundred and eighteen centimetres and the median did not shift by a millimetre. It was the smallest before and it is the smallest now, and that is the only question the median asked. Outliers come from the other end too. Fifteen students, asked how many short stories they read last year.
Six, three, zero, eight, two, five, seven, fifteen, twelve, ten, forty, five, zero, one, and eight. There is the forty. And there is the gap that makes it one: twenty five clear of the next student along. Sort them and the middle value is six. Add them up and divide, and the average is eight point one. So this time the average sits above the middle, not below it. The pull is upwards, because the far-out value is upwards.
And here is a set with nothing much going on at all. A newspaper, seven days running: sixteen pages, eighteen, twenty, twenty two, twenty six, sixteen, and ten. The middle value is eighteen. The average is eighteen point three. Three tenths of a page apart. For any practical purpose they are the same number. The largest gap between neighbours here is six, in a range of sixteen. Nothing sits on its own, so nothing is pulling, and the two summaries have no reason to disagree.
Three sets, three relationships, and together they give you something you can use. When the average sits below the middle value, something is pulling from the low end. When it sits above, something is pulling from the high end. When the two are nearly equal, nothing much is pulling either way. So you can be handed those two numbers, with none of the data behind them, and still say which way the set leans.
And notice what that is not saying. The average being below the middle does not mean the values are small. It means the shape is lopsided, and it tells you which side is doing it. Now the tempting sentence, and it is wrong. People say the median is unaffected by outliers. It is not. Take the hundred and eighteen out of that family altogether, and work both summaries again. Four heights left. The average leaps from a hundred and sixty point two to a hundred and seventy point seven five.
That is a jump of ten and a half centimetres. And the median goes from a hundred and seventy to a hundred and seventy one point five. It moved as well. By one and a half. It moved because taking somebody out changes how many people there are, and that changes where the middle sits. So the honest sentence is a comparison. On the very same edit, the average travelled seven times as far. The median is hard to push, not impossible.
The two of them share a name. They are both attempts to say where a set of values is centred, so they are called measures of central tendency. And neither one is the right answer in general. For the newspaper the average is fine, and it has a real advantage: it uses every day of the week, not just the middle one. For that family the median is plainly better, and the average is actively misleading.
Which to reach for is a judgement about the data in front of you, and you now know what to look at before making it. One last case, and it is the awkward one. What happens when a set has an outlier at each end? Here is one built on purpose. Zero, forty eight, forty nine, fifty, fifty one, fifty two, and a hundred. The middle value is fifty. The average is fifty.
They agree exactly. Now throw both far-out values away and do it again. The average is fifty. The middle value is fifty. Neither of them has moved by anything at all. And yet the set has gone from spanning a hundred to spanning four. So two summaries agreeing is not proof that a set is well behaved. The two pulls simply cancelled, and both numbers are blind to the values that make this set what it is.
Which is a different question entirely. Not where the values are centred, but how far apart they are.
Where this fits
Taken from the notes each video was made from, not from the reading order — these are the ideas this one rests on and the ones that later rest on it.
Builds on
- The arithmetic mean as fair-shareClass 7 · Ch 5, Connecting the Dots...
- Dot plots: seeing spread and clustering at a glanceClass 7 · Ch 5, Connecting the Dots...
Comes up again in
- Why one number is never enough to describe a data setClass 7 · Ch 5, Connecting the Dots...