PrepShorts · Study sheet · Class 11 Mathematics · Chapter 13, Statistics
This video could not be loaded. Reload the page to try again.
Sign in with Google13 min.
Keep your place in this chapter — sign in, it’s free.Sign in
Six readings spread from 5 to 55, and thirty-one packed from 15 to 45 - and the plain total of their squared deviations ranks the packed set as more scattered.
The idea
Squaring solves the sign problem the modulus solved, and it does it with an operation that can be expanded — which is the whole reason for the change. But the first quantity you can build from squares, the plain total of the squared deviations, is not a measure of dispersion at all, and §13.5 proves it rather than asserting it: two data sets centred on the same value are put up, one with six observations reaching 25 either side and one with thirty-one reaching only 15, and the totals come out 1750 and 2480 — the wrong way round. The total is counting observations as much as it is measuring spread. Dividing by the number of observations removes that, and the quotient is the variance.
What you should be able to do
- State why squaring removes the sign problem without introducing a modulus
- Show that the total of the squared deviations about the mean is zero exactly when every observation equals the mean
- Compute the total of the squared deviations for each of the chapter's two sets
- Use the sum-of-squares formula for the first n natural numbers to evaluate the larger of those totals without adding thirty-one terms
- Explain why that total ranks the two sets in the opposite order to their actual spread
- Divide by the number of observations and show that the ranking corrects
- Define the variance and give its symbol
- Quantify how much more weight squaring gives to a distant observation than averaging distances does
Words to know
| Term | Definition in one line | First introduced |
|---|---|---|
| variance | the mean of the squared deviations of the observations from their mean | printed in this chapter (§13.5, p. 273) |
| sigma square | the chapter's reading of the symbol it writes the variance with | printed in this chapter (§13.5, p. 273) |
| squares of deviations | the non-negative quantities the section is built from | printed in this chapter (§13.5, pp. 271–273) |
| degree of dispersion | the chapter's phrase for how much scatter a data set shows | printed in this chapter (§13.5, p. 272) |
| mean | the centre all the deviations in this section are taken about | printed throughout this chapter (§13.1, p. 257) |
| observation | one recorded value in the data set | printed throughout this chapter (§13.1, p. 258) |
| geometrical representation | the chapter's name for the dot diagrams that confirm the corrected ranking | printed in this chapter (§13.5, p. 273) |
| leverage | the extra weight squaring gives to an observation far from the centre | an added term; the chapter demonstrates the effect and does not name it |
Where people slip up
- "Squaring is just another way to drop the minus sign, so it must give the same answer." It gives a different answer and a different ranking of magnitudes, because it stretches large deviations more than small ones. Set A's outermost pair carries 56 per cent of the mean deviation and 71 per cent of the variance.
- "A bigger total of squares means more scatter." Not across data sets of different sizes. That is exactly what §13.5's two sets are printed to refute.
- "So the total of squares is useless." It is not — it is zero precisely when there is no scatter at all, and it is the numerator of everything that follows. It is unusable as a comparator, which is a narrower complaint.
- "Dividing by six and by thirty-one is unfair to set B." It is the only way the two are comparable at all. Dividing by the count is what turns a total into a per-observation figure.
- "The variance of set A is bigger, so set A's observations are bigger." Both sets have mean 30. Variance says nothing about where the data sits; it was built from deviations, which are indifferent to the location of the centre.
- "291.67 is exact." It is 1750/6 rounded to two places. Keep the fraction in the working.
- "Set B is more spread because it covers thirty-one values." It covers a narrower interval, 15 to 45, against set A's 5 to 55. The number of observations is not the width.
Ask your teacher a person
Your teacher reads this and writes back, usually within a day. For an instant answer, use Ask the video in the sidebar.
Your class sees the question and the answer. Only your teacher sees that it was you.
No questions on this topic yet.
Worked answers: Exercise 13.1 · Exercise 13.2 · Miscellaneous Exercise
Transcript1,717 words
The plan has not changed. Take each observation's deviation from the centre, remove the sign, average the results. Exactly one step is different: the sign comes off by squaring instead of by taking a size. And it has to be a square rather than any other even power, because a square is the one that expands cheaply into terms you can total separately - which is what everything after this depends on.
A fourth power would remove the sign just as well and be useless afterwards. So: square the deviations. What can we build out of them? First, notice what squaring buys immediately. A deviation of minus five and a deviation of plus five are different numbers. Squared, they are the same number, twenty-five. Both sides of the centre land on the same non-negative value, and nothing on one side can cancel anything on the other.
That was the whole problem the signs caused. It is now gone, and gone without a modulus. The simplest thing to build is the total. Square every deviation, add them all up, and call that the answer. It has one property already worth having. Every term in that sum is a square, so no term can be negative, so the total can only be nought if every single term is nought.
And a squared deviation is nought only when the observation sits exactly on the mean. So the total is nought exactly when every observation is the mean - when there is no scatter at all. That is not nothing. It is a real property, and it is the reason squares are worth trying. That last statement is an if-and-only-if, and those are easy to believe and hard to test, because on ordinary data neither side of it ever happens.
Take two hundred and sixty records, and deliberately build sixty of them out of a single repeated value. Sixty have a total of nought. Sixty are constant. And the number that belongs to one of those groups but not the other is nought. Had I used only ordinary records, both counts would have been nought and the check would have proved nothing whatsoever. Now the question that decides everything: is that total a measure of spread?
Here is a set of six observations - five, fifteen, twenty-five, thirty-five, forty-five and fifty-five. They add to a hundred and eighty, so the mean is thirty. The deviations are minus twenty-five, minus fifteen, minus five, five, fifteen and twenty-five. Squared: six hundred and twenty-five, two hundred and twenty-five, twenty-five, twenty-five, two hundred and twenty-five, six hundred and twenty-five. Add them and you get one thousand seven hundred and fifty.
Notice that this set reaches twenty-five either side of its centre. Here is a second set: every whole number from fifteen to forty-five. Thirty-one observations, and by symmetry the mean is thirty again. Its deviations run from minus fifteen up through nought to fifteen. And they pair off - for every deviation of a given size on the left there is one of exactly that size on the right. Listed by size, they are every whole number from one to fifteen, twice over, with a single nought in the middle.
So the total of their squares is twice the total of the squares of one to fifteen. Notice: this set reaches only fifteen either side of its centre. Half as far as the first set. Rather than write out thirty-one squares, use the formula for the squares of the first whole numbers: n times n plus one, times twice n plus one, all over six. That formula deserves a moment of suspicion, so I checked it against actually adding the squares, for every count from nought to sixty.
Sixty-one values tried; nought disagreements. At fifteen it gives fifteen times sixteen times thirty-one over six, which is one thousand two hundred and forty. Double it, because the deviations pair off, and the second set's total is two thousand four hundred and eighty. Now put the two side by side, and something has gone wrong. The first set reaches twenty-five either side of its centre; the second reaches fifteen. The first set spans fifty, from five to fifty-five; the second spans thirty, from fifteen to forty-five.
By any reading of the word, the first set is the more spread out. But its total is one thousand seven hundred and fifty, and the second set's is two thousand four hundred and eighty. The total says the second set is the more dispersed. It has the order exactly backwards. The pictures make it plain. Draw both on the same axis, with the same centre marked at thirty. The first set is six dots stretched right across, from five out to fifty-five.
The second is a solid unbroken block of thirty-one dots crowded between fifteen and forty-five. One of these is spread out and one of them is packed together, and there is no argument about which. The total of squares disagrees with the picture. So what went wrong? Nothing subtle. The first set has six observations and the second has thirty-one - twenty-five more. Every one of those twenty-five extra observations contributes another square, and a square is never negative.
The total was measuring spread, and it was also counting observations, and it added the two together. That can be made exact. Append one more observation to any record, and the total goes up by exactly n over n plus one, times the square of that new observation's distance from the old mean. I tried that three thousand three hundred times. The identity failed nought times, and the total fell nought times.
It can only climb. Five of those three thousand three hundred appends changed the total by nothing at all - and all five were the case where the new observation landed exactly on the old mean, adding no distance to anything. But here is the cleanest demonstration that the total is counting. Take any record and write every observation down three times. Nothing about its spread has changed. Same mean, same deviations, same reach from the centre, same everything.
The total is three times as big. Over three hundred records, at three different multiples, that is nine hundred trials, and it failed nought times. A quantity that triples when you change nothing is not measuring what you thought it was measuring. The repair is the obvious one. If the total is partly a count of observations, divide by the count of observations. One thousand seven hundred and fifty over six is eight hundred and seventy-five thirds - two hundred and ninety-one point six seven, rounded.
Keep the fraction in the working, though: two hundred and ninety-one point six seven is a rounding, and truncating instead would print two hundred and ninety-one point six six. Two thousand four hundred and eighty over thirty-one is exactly eighty. Two hundred and ninety-one and change against eighty. The order is back, and it agrees with the picture. That quotient - the average of the squared deviations - is the variance.
It is written with a sigma squared, and read as sigma square. That is the whole definition: take the deviations from the mean, square them, average them. The first set's variance is eight hundred and seventy-five thirds and the second set's is eighty. The first is about three point six five times the second. Before we move on, one honest check. It is no use asking whether the variance ranks two records the way the variance ranks them.
So I scored both measures against something built from neither: how far a record reaches from its own mean. Of three hundred pairs, a hundred and two had the shorter record reaching further than the longer one. The total of squares gets sixty-seven of those hundred and two backwards. The variance gets eighteen of them backwards. Not nought - eighteen. The variance is a far better comparator, and it is not a perfect one, and a video that told you it was would be lying to you.
There is one more thing squaring did, and it was not free. Look again at the first set's total. The two outermost observations contribute six hundred and twenty-five each - one thousand two hundred and fifty of the one thousand seven hundred and fifty. That is seventy-one point four per cent of the whole answer, from two of the six observations. Now measure the same set the older way, by averaging distances rather than squares.
The same two observations carry fifty-five point six per cent. And if you had used fourth powers, they would carry eighty-eight point four per cent. Squaring did not merely remove the sign. It moved a large share of the measure onto the observations furthest from the centre. That is a design decision, not an accident, and it is worth knowing it holds generally. Over three hundred records nobody chose, the outer pair's share under squares came out below its share under distances nought times, and strictly above it three hundred times.
Every single record. But it is not a law about every possible record either. Take nought, twenty, nought, twenty. Its mean is ten and all four deviations have the same size, ten. So the outer pair carries exactly half of the answer no matter which power you use - a half under distances, a half under squares, a half under fourth powers. When every deviation is the same size, squaring has nothing to move the weight towards.
One last thing about what the variance does not tell you. It says nothing at all about where the data sits. Add the same amount to every observation - shift the whole record up the line - and the mean moves by exactly that amount while the variance does not move at all. Three hundred records at three different shifts, nine hundred trials, nought failures. Our two sets are a case of it: both of them sit on thirty.
They are told apart entirely by their spread, which is exactly what a measure of dispersion should do. So: square the deviations, average them, and you have the variance. The squares are still there, though, and their units are the square of whatever you measured - and that is the next thing to fix.
Where this fits
Taken from the notes each video was made from, not from the reading order — these are the ideas this one rests on and the ones that later rest on it.
Builds on
- Where this measure breaks down, and why another was neededClass 11 · Ch 13, Statistics
Comes up again in
- Taking the root to get the spread back into the units of the dataClass 11 · Ch 13, Statistics