PrepShorts · Study sheet · Class 11 Mathematics · Chapter 13, StatisticsPrepShorts

Chapter 13 · Statistics

The cheapest measure of spread, and everything it fails to notice

Why a single central value is not enough15 min

This video could not be loaded. Reload the page to try again.

Sign in with Google

15 min.

Two batsmen average the same fifty-three runs and share the same middle score - one subtraction, largest innings less smallest, finally tells them apart: 117 against 14.

The idea

Subtract the smallest observation from the largest and you have a number that does what no measure of central tendency could: on the two batsmen it returns 117 against 14 and separates them instantly. That is a real success and it is why §13.3 comes first. But the subtraction consults exactly two of the ten innings and discards the other eight, and it never once refers to a centre — so it cannot distinguish a set that crowds its middle from a set that piles up at its two ends. Every measure in the rest of the chapter is an answer to that second failure, and §13.3 says so in its closing paragraph.

What you should be able to do

  • List the four measures of dispersion §13.2 names, and state which one this chapter sets aside
  • Compute the range of a data set as the difference between its largest and smallest observations
  • Compute and compare the ranges of the chapter's two batting records
  • Explain why a range built from two observations cannot report how the remaining observations are arranged
  • Construct two data sets with the same range and visibly different scatter, and compute a distance-based figure that separates them
  • State the requirement §13.3 arrives at: a measure of dispersion must be built from deviations about a central value
  • Say why the range still gets computed in practice despite these limitations

Words to know

TermDefinition in one lineFirst introduced
rangethe largest observation minus the smallestprinted in this chapter (§13.3 heading, p. 259)
quartile deviationa measure of dispersion the chapter names in its list and then excludes from studyprinted in this chapter (§13.2, p. 259)
mean deviationthe average of the distances of the observations from a chosen central valueprinted in this chapter (§13.2 and §13.4, p. 259)
standard deviationthe measure of dispersion §13.5 builds from squared deviationsprinted in this chapter (§13.2, p. 259)
deviationthe difference between an observation and a fixed valueprinted in this chapter (§13.4, p. 259)
maximum valuethe largest observation in the setprinted in this chapter (§13.3, p. 259)
minimum valuethe smallest observation in the setprinted in this chapter (§13.3, p. 259)
seriesthe chapter's word for one data set when its range is being takenprinted in this chapter (§13.3, p. 259)
extreme observationeither of the two values a range is computed froman added compound; the chapter describes these values without giving them a joint name

Where people slip up

  • "The range is a bad measure." It is a narrow one, and the chapter says as much rather than dismissing it — it uses the range to separate the two batsmen before anything better exists. The right verdict is that it answers a smaller question than the one being asked.
  • "A bigger range always means more scatter." The manufactured record in section 5 has exactly B's range and more than twice his average distance from the centre. Equal ranges do not imply equal scatter; the range only bounds it.
  • "Quartile deviation is not in the syllabus, so it does not exist." §13.2 names it as one of four and then declines it. Say that it exists and is not developed here, rather than quietly leaving it off the list.
  • "Range needs the data sorted." It needs the two extreme entries, which can be found in one pass. Sorting is the median's requirement, not the range's.
  • "Every measure of dispersion is a distance from the centre." The range is not — it is a distance between two observations, and never refers to a centre at all. That is precisely the gap §13.4 opens by.
  • "Range is useless for grouped data." It is computable from the class limits, but it then reports the width of the reported classes rather than of the observations. The chapter does not raise this; do not build on it.
Transcript2,085 words

Two players, ten games each. The same two records as before. The first scored thirty, ninety-one, nought, sixty-four, forty-two, eighty, thirty, five, a hundred and seventeen, seventy-one. The second scored fifty-three, forty-six, forty-eight, fifty, fifty-three, fifty-three, fifty-eight, sixty, fifty-seven, fifty-two. Their averages are both fifty-three. Their middle values are both fifty-three. Of those comparisons, the number that come out different is nought. Now do something almost embarrassingly cheap. Take the largest score in each record and subtract the smallest.

The first player: a hundred and seventeen, take away nought. A hundred and seventeen. The second player: sixty, take away forty-six. Fourteen. A hundred and seventeen against fourteen. That is not a close call, and it is the first summary that has told these two records apart at all. That number has a name. It is called the range, and this video is about how much it gets right and exactly where it stops.

Look at what the range costs to compute. You walk the record once, keeping the biggest thing you have seen and the smallest thing you have seen. At the end you subtract. There is no sorting. Sorting is what the middle value needs, because the middle value is a position and positions only exist once things are in order. The range is not a position. It is a gap, and a gap needs only its two ends.

One walk, one subtraction. Nothing in the whole of this subject is cheaper, and it just separated two records that the centres could not. So the question is not whether the range works. It is what question it is answering, because it is plainly not answering nothing. Before going further, the shape of the subject. A number that reports how widely a set of observations is strewn is called a measure of dispersion, and there is more than one of them.

The range is the one being built here. The mean deviation and the standard deviation come later in this series, and each gets a video of its own. There is also the quartile deviation. It is real, it is used, and this series does not build it. It is being named rather than quietly left off the list, so that meeting the phrase later is not a surprise. The range comes first for a reason. It is the cheapest, and its failure is the clearest, and understanding its failure is most of the argument for everything that replaces it.

Here is the second player's ten games on a line, with the two ends lit up. Forty-six and sixty. Those two scores went into the subtraction. The other eight - forty-eight, fifty, fifty-two, fifty-three, fifty-three, fifty-three, fifty-seven and fifty-eight - did not. Not that they mattered less. They did not enter the arithmetic at all. That claim can be tested rather than asserted. Take each position of the record in turn and replace whatever is sitting there with every whole number strictly between forty-six and sixty.

There are thirteen such numbers, and there are eight positions where no replacement can shift the range. Eight positions, thirteen values, a hundred and four rewrites. Of those hundred and four rewrites, the number that move the range is nought. The number that move the average is ninety-six, and the number that move the average distance from fifty-three is ninety-one. So eight of the ten games could have gone almost any other way and the range would have come out fourteen regardless. Whatever it is measuring, it is not measuring them.

Take that seriously for a moment and build a record out of it. Keep the second player's two extreme scores, forty-six and sixty. Use nothing else. Five games of forty-six, and five games of sixty. A player who either had a bad day or had a good day and never once had an ordinary one. Nothing about that record was chosen freely. Both of its values came off the record it is being compared with.

Its range is sixty take away forty-six. Fourteen. The same fourteen. And now the interesting part, because the range is not the only thing it kept. Add the ten scores up. Five forty-sixes is two hundred and thirty, five sixties is three hundred, and the total is five hundred and thirty. Five hundred and thirty is exactly what the real record totalled, so the average is fifty-three. The same fifty-three.

Sort the manufactured record and look at the middle. The fifth entry is forty-six and the sixth is sixty, and halfway between them is fifty-three. Which is worth stopping on. This record contains fifty-three exactly nought times, and its middle value is fifty-three. So here are two records that agree on their average, agree on their middle value, and now agree on their range as well. The range has just joined the summaries that cannot tell these two apart. It bought us the first separation and it cannot buy us this one.

They are obviously different records. So separate them. Take each score, ask how far it is from fifty-three, and average those distances. For the manufactured record every single score is seven away from fifty-three, so the average distance is seven exactly. For the real record it is three point two. Seven against three point two. One is two point one eight times the other. Square the distances first and then average them, and it is forty-nine against seventeen point four - two point eight one times.

Same two ends. Same range. More than twice the strewing, whichever way the distances are handled. That seven is not a coincidence, and this is the part worth remembering. Seven is half of fourteen. Half of the range. For any record at all, the average distance from its own average can never exceed half its range. It is not allowed to. Over four hundred records nobody chose, the number whose average distance comes out above half their range is nought.

The number that reach exactly half is also nought - among records made at random, nothing gets that far. But the manufactured record does reach it. Its average distance is exactly half its range, and the real record's is not. So the range is not a measurement of strewing that happens to be inaccurate. It is a ceiling. It tells you how strewn a record could possibly be. It says nothing whatever about how strewn it actually is.

The two records are one example, and one example is never the argument. So build every record in between. Start from the tightest thing those two ends allow: one score at forty-six, one at sixty, and the other eight sitting on fifty-three. Then move one of the middle scores down to forty-six and another up to sixty, and keep doing that until nothing is left in the middle. That gives five records, ending at the five-and-five one we already built.

Every one of them has a range of fourteen. The number with any other range is nought. Every one of them averages fifty-three and has a middle value of fifty-three. The number with any other centre is nought. And their average distances from fifty-three run one point four, two point eight, four point two, five point six, seven. Five records, five different values. Count what each summary can do with that ladder.

There are ten pairs of records in it. The number of those pairs the range can tell apart is nought. The number the average distance can tell apart is ten. All of them. That is the whole failure, stated as a count rather than as an opinion. And it is not an artefact of a ladder built on purpose. Take four hundred records nobody chose, every one of them carrying nought and a hundred and twenty, so every one of them has a range of a hundred and twenty.

The number with any other range is nought. The number of different average distances among them is two hundred and eighty-four. They run from twenty point four four to forty-eight point seven. Same range, right across that band. Why is the range blind to this, when it was so decisive a few minutes ago? Because of what the subtraction refers to. It refers to two observations, and to nothing else.

It never mentions a centre. There is no average in it, no middle value, no fixed point of any kind - just one score take away another score. Watch what that buys and what it costs. Move a whole record bodily along the line, adding thirty-seven to every score. Over four hundred records, the number whose range changes is nought, and the number whose average distance from their own average changes is also nought. Both are blind to where the record sits, which is correct, because both are measuring spread.

Now go back to the interior rewrites. Nought of them moved the range. Ninety-six moved the average, and ninety-one moved the average distance. That is the gap, in one comparison. Both quantities ignore where a record sits. Only one of them notices how its observations are arranged inside their own two ends. None of this makes the range a bad measure, and it is worth being precise about that, because it is easy to leave with the wrong verdict.

The range is the exact answer to a question. Just not the question we have been asking. The question it answers exactly is: what is the largest difference between any two observations in this record? Test that. Walk every pair of scores in a record, take the size of each difference, and keep the largest - a routine that never asks for a biggest or a smallest at all. Over four hundred records, the number of times that disagrees with the range is nought. Not approximately. Exactly, every time.

And to show that this comparison is capable of saying no, run the obvious near miss through it: the two ends added instead of subtracted. That gets thirty-three of the four hundred right - which is exactly the number of records that happen to contain a nought. Thirty-three accidental agreements out of four hundred. One case coming out right proves nothing, and now you can see why. One more thing the range does that is worth knowing before you use it.

Take the first player's record and remove one game from each end - the nought and the hundred and seventeen. One unusually bad day, one unusually good one. What is left runs from five to ninety-one. A range of eighty-six. The range fell by thirty-one. The average moved by one and three eighths. Two games out of ten, and one summary barely flinched while the other lost thirty-one points. Remove only the highest game and nine are left, so the middle rule takes a single entry rather than averaging two. The range reads ninety-one, and the middle value drops from fifty-three to forty-two.

Push the other way instead. Take that highest score and move it further out by every amount from one to fifty. All fifty times, the range grows by exactly the amount you moved it. All fifty times, the average grows by exactly a tenth of that amount - a tenth, because there are ten games and the move is shared among all of them. The range takes the whole of an unusual observation. Every centre divides it up. That is a real property and sometimes it is the one you want.

So where does that leave us. The range answers its own question exactly, and answers it from two observations, and never once refers to a centre. Because it never refers to a centre, it cannot report how the observations are arranged around one. It bounds that arrangement from above and stops. Which tells you precisely what a replacement has to be made of. It has to pick a central value. It has to ask how far each observation lies from that value - every observation, not two of them.

And it has to combine all of those distances into a single number. Average the distances as they come and you get one such measure. Square them first and then average and you get another. Those are the two that get built next, and everything they can do that the range cannot comes from that one change: every observation, measured from a centre. The range gave us the first honest separation we have had. It is just not able to give us the second.

Where this fits

Taken from the notes each video was made from, not from the reading order — these are the ideas this one rests on and the ones that later rest on it.

Builds on

Comes up again in

The book

Open in a new tab