PrepShorts · Study sheet · Class 11 Mathematics · Chapter 13, StatisticsPrepShorts

Chapter 13 · Statistics

Taking the root to get the spread back into the units of the data

Squaring instead of discarding the sign13 min

This video could not be loaded. Reload the page to try again.

Sign in with Google

13 min.

A batting record averaging fifty-three runs, best score a hundred and seventeen, carries a variance of 1300.6 - bigger than the record's entire span.

The idea

The variance ranks data sets correctly and describes none of them, because squaring changed the units: a variance of runs is in runs-squared, and there is no axis on which it can be laid beside the mean. §13.5.1 fixes exactly that and nothing else. Taking the non-negative square root returns the measure to the units the observations were recorded in, and because the square root is increasing it cannot disturb any ordering the variance established. The standard deviation therefore carries no information the variance did not; what it carries is the ability to be read on the same scale as the data, which is why it is the number that gets reported.

What you should be able to do

  • State why the variance cannot be compared with the mean or plotted on the observations' own axis
  • Define the standard deviation as the non-negative square root of the variance, and say why the non-negative root is chosen
  • Explain why taking the root cannot change which of two data sets is judged more dispersed
  • Compute the variance of an ungrouped data set with the mean found by the step-deviation method
  • Compute the corresponding standard deviation and round it as the chapter does
  • Compare the standard deviation, the mean deviation and the range of the same data set and say what each is telling you
  • Identify a printed slip in the chapter's own working table and say what the column should have read

Words to know

TermDefinition in one lineFirst introduced
standard deviationthe non-negative square root of the varianceprinted in this chapter from §13.2's list, p. 259, and again in §13.4.3, p. 271; defined at the §13.5.1 heading, p. 274
variancethe mean of the squared deviations about the meanprinted in this chapter (§13.5, p. 273)
square-rootthe operation §13.5.1 applies to the varianceprinted in this chapter (§13.5.1, p. 274)
unitswhat the observations are measured in, and what the standard deviation is restored toprinted in this chapter (§13.5.1, p. 274)
assumed meanthe convenient value the deviations are measured from in the workingprinted in this chapter (§13.4.2, p. 266; used again in §13.5.1, p. 274)
step-deviation methodthe route to the mean used in the Example 8 tableprinted in this chapter (§13.4.2, p. 266; named again on p. 274)
ungrouped dataobservations listed one by one, with no frequency columnprinted in this chapter (§13.4.1, p. 260; §13.5.1, p. 274)
root mean squarethe general name for the operation that produces a standard deviation from a list of deviationsan added term; not printed in this chapter

Where people slip up

  • "Variance and standard deviation are two different measures." They are the same measure in two unit systems, and either determines the other. Choose by what you need to do with it: report the standard deviation, compute with the variance.
  • "The square root can be negative." By convention the root written here is the non-negative one, because a standard deviation is a distance. The chapter says so in the sentence that introduces it.
  • "Taking the root might reverse which data set is more spread." It cannot. The square root is increasing on the non-negative numbers.
  • "The step-deviation columns can also be used for the squares." They cannot. The step-deviations in Table 13.7 are measured from 14 and shrunk by 2; the squared column is measured from the true mean, 15, at full scale. What happens to a variance under shifting and scaling is the subject of §13.5.4, not of a shortcut taken here.
  • "5.74 is exact." It is the square root of 33 rounded to two places. Keep the surd until the last line.
  • "A standard deviation of 5.74 means most observations are within 5.74 of the mean." Nothing in this chapter says that. It is a typical squared-weighted distance, not a guarantee about how many observations fall inside it. For this data four of the ten lie further out than 5.74.
  • "Both mean deviation and standard deviation should give the same number." For the same data they generally differ, and the standard deviation is the larger unless every distance from the mean is identical.
Transcript1,763 words

The variance works. It ranks two records the right way round, it is nought exactly when there is no scatter, and it can be built up from parts. And you cannot put it on the board next to the data. Here is a batting record of ten innings. Its average is fifty-three runs. Its variance is one thousand three hundred point six. The best innings in the whole record was one hundred and seventeen, so the entire span of the data is one hundred and seventeen runs - and the variance is bigger than that.

A measure of spread that is larger than the whole span is not wrong. It is in different units. That sentence gets said a lot, and on its own it is a slogan. Here is the version you can check. The units of a quantity are visible in what happens when you change the unit the observations are recorded in. Multiply every observation by some factor, and see what each quantity does.

The average multiplies by the factor. The distances from the average multiply by the factor. The variance multiplies by the factor SQUARED. Three hundred records at four different factors, twelve hundred trials, nought exceptions - and nought trials in which the variance moved by the factor itself rather than by its square. That is not a slogan. That is the exponent, measured. So undo it. Take the square root of the variance.

That is the standard deviation, and it is written with a plain sigma - the same letter as the variance, without the square. Two things are worth saying out loud before we use it. The root is taken to be the non-negative one, by choice, because this number is going to be read as a distance and a distance is not negative. And the root always exists, because a variance is an average of squares and can never be negative in the first place.

Over a pool of three hundred and eighty records - ordinary ones, two-value ones, and flat ones with no spread at all - there is not a single negative variance to take the root of. The ones that come out nought are exactly the twenty flat records, and no others. Now the honest question. Is the standard deviation a better measure than the variance? No. It cannot be, and that is the whole point of it.

The square root is increasing: if one non-negative number is bigger than another, so is its root. So any ranking the variance gives you, the root gives you back unchanged. Take sixty records, and rank every pair of them by variance, and then rank every pair again by the root. One thousand seven hundred and seventy pairs, nought disagreements - and every single one of those pairs has two different variances, so not one of them is a tie that could have hidden a disagreement.

The standard deviation carries no information the variance did not carry. What it carries is the ability to be read on the same axis as the data, and that is why it is the one that gets reported. Let us do one properly, from the beginning. Ten observations: six, eight, ten, twelve, fourteen, sixteen, eighteen, twenty, twenty-two, twenty-four. Evenly spaced, two apart, all the way along. We want the variance and then the standard deviation, and the first job is the centre.

You could add ten numbers. There is a cheaper route and it is worth taking, because the cheap route is where the interesting mistake lives. Pick a convenient value near the middle and measure from it. Fourteen will do. The distances from fourteen are all even, so divide each of them by two as well. That gives a short column: minus four, minus three, minus two, minus one, nought, one, two, three, four, five.

It totals five. So the average of that column is five over ten, which is a half. Multiply back up by the two we divided out, and add the fourteen we measured from. Fourteen plus a half times two. Fifteen. Adding the ten numbers the long way gives the same fifteen, so the shortcut is sound for what it was built to do. Now the deviations, from the real centre, at full size.

Every observation is even and fifteen is odd, so fifteen is not one of the ten, and not one deviation is nought. They are minus nine, minus seven, minus five, minus three, minus one, one, three, five, seven, nine. The odd numbers, one to nine, once on each side. That is not a coincidence of this data. Ten values evenly spaced sit symmetrically about their own centre, so the distances pair off.

Square them. Eighty-one, forty-nine, twenty-five, nine, one - and then the same five again on the other side. The total is three hundred and thirty, which is twice one hundred and sixty-five. And one hundred and sixty-five is the first five odd squares added up. There is a closed form for a run of odd squares, and it agrees with adding them for every count up to sixty: sixty-one counts tried, nought disagreements.

Divide the three hundred and thirty by ten observations and the variance is thirty-three. And thirty-three is not an accident of these ten numbers either. For n values spaced h apart, the variance is h squared times n squared minus one, over twelve. Ten values two apart: four times ninety-nine over twelve, which is thirty-three. Checked against the long way over a hundred different combinations of count and spacing, nought disagreements.

Now the mistake, because it is the one everybody makes. We already have a short column of numbers on the board: minus four through five, the convenient ones. Why not square those? Their squares total eighty-five. Divide by ten and you get eight point five. That is measured in units shrunk by two, so multiply back by two squared, which is four. Thirty-four. The answer is thirty-three. Not approximately thirty-three.

Thirty-three, and the shortcut says thirty-four. The gap is one. And one is the square of the distance between fourteen, the value we measured from, and fifteen, the value we should have measured from. That is exact, and it is not about this data. Take any record and any centre at all. The average of the squared distances from that centre is the variance, plus the square of how far the centre is from the true mean.

Three hundred records against a hundred and forty-one different centres - forty-two thousand three hundred trials, nought exceptions. Two things fall straight out of it. No centre ever beats the mean, because you are always adding a square to the variance - nought of those forty-two thousand came out below. And the ties are exactly the thirty-two cases where the centre we tried happened to BE that record's mean, which is the same thirty-two counted a completely different way.

So: variance thirty-three. Take the root. Thirty-three is not a perfect square, so the root is not going to be a tidy number, and what gets written down is a reading of it. To two places it is five point seven four. Five point seven four squared is thirty-two point nine four seven six, which is not thirty-three. Keep the root itself in the working and round once, at the end.

Here, as it happens, rounding and truncating agree at two places - but they are different operations and on the batting record they part company, where one says thirty-six point one and the other says thirty-six point nought. Now put the ten observations on a line and mark what we have. The centre is at fifteen. The whole span, largest less smallest, is eighteen. The average distance from the centre - the older measure, distances without squares - is fifty over ten, which is five.

And the standard deviation is five point seven four. All three of those can be drawn as a length on this axis, beside a centre of fifteen. The variance, thirty-three, cannot: it is longer than the entire span of the data. The span reports everything, the average distance reports a typical distance, and the standard deviation reports that same typical distance with the far-out observations counting for more. Notice that the standard deviation came out larger than the average distance.

Five point seven four against five. That is not this data being awkward. For any list of distances, the root of the average of the squares is at least the average itself. Over the same three hundred and eighty records, there is not one where it falls below. But it is not always strictly larger, and the pool is built so you can see when. The two agree on exactly eighty records - and those eighty are exactly the eighty whose distances from the centre are all the same size.

Nought records belong to one of those groups and not the other. A pool of ordinary records would have shown you no equality at all, and you would have learnt the wrong rule. For our ten, the distances are nine, seven, five, three and one - not all the same size - so the inequality is strict. One thing the number does not say. A standard deviation of five point seven four does not promise that most of the observations are within five point seven four of the centre.

Nothing here says that. It is a typical distance with the far ones weighted up, not a count of anything. Check it against these ten: four of the ten lie further from fifteen than five point seven four, and six lie closer. Six out of ten is not most, and it is not a rule either - it is just what this record happens to do. So where does that leave the two of them?

They are not two measures. They are one measure in two unit systems, and either one hands you the other. Square the standard deviation and you have the variance; root the variance and you have the standard deviation back. Use the variance when you are going to calculate with it, because it is the one that adds and expands. Report the standard deviation, because it is the one that can be laid beside the data.

One record of innings, spread thirty-six point one runs about an average of fifty-three. Ten observations, spread five point seven four about a centre of fifteen. Both of those are sentences you can say. Neither of them could be said with a variance.

Where this fits

Taken from the notes each video was made from, not from the reading order — these are the ideas this one rests on and the ones that later rest on it.

Builds on

Comes up again in

The book

Open in a new tab