PrepShorts · Study sheet · Class 11 Mathematics · Chapter 13, Statistics
This video could not be loaded. Reload the page to try again.
Sign in with Google19 min.
Keep your place in this chapter — sign in, it’s free.Sign in
Slide every observation along by the same amount and not one gap between any two of them changes, so the variance never notices a shift at all.
The idea
The substitution §13.5.4 makes is two transformations wearing one name, and they do different things to the spread. Subtracting an assumed value slides every observation the same distance, so no gap between any two of them changes — and a variance is built only out of gaps, so a shift cannot touch it. Dividing by a positive common factor shrinks every gap by that factor, so the standard deviation shrinks by it and the variance by its square. Which is why the mean has to be shifted back and scaled back, while the spread only ever has to be scaled back. Getting that asymmetry right is the whole content of the section, and two of the Miscellaneous Examples, on pp. 282–284, state the two halves of it separately, as results in their own right.
What you should be able to do
- Write the step-deviation substitution and its inverse
- Prove that adding a constant to every observation leaves the variance unchanged
- Prove that multiplying every observation by a positive constant multiplies the standard deviation by that constant and the variance by its square
- Recover the mean of the original data from the mean of the transformed data
- Recover the standard deviation of the original data from that of the transformed data, and say why only one of the two operations has to be undone
- Apply the assembled shortcut formula to a continuous distribution and check it against the direct computation
- State what the transformation results assume about the sign of the scale factor
Words to know
| Term | Definition in one line | First introduced |
|---|---|---|
| shortcut method | the chapter's name for computing a mean and variance through a shifted, scaled variable | printed in this chapter (§13.4.2, p. 266; §13.5.4 heading, p. 279) |
| step-deviation | the transformed observation, a deviation from the assumed value divided by the scale | printed in this chapter (§13.4.2, p. 267) |
| assumed mean | the value the observations are measured from before scaling | printed in this chapter (§13.4.2, p. 266) |
| common factor | the number divided out of all the deviations | printed in this chapter (§13.4.2, p. 267) |
| class-intervals | the bands whose common width supplies the scale factor here | printed in this chapter (§13.4.2, p. 265; §13.5.4, p. 279) |
| change of scale | the chapter's phrase for dividing the deviations by the common factor | printed in this chapter (§13.4.2, p. 267) |
| shifting of origin | the chapter's phrase for measuring from the assumed mean instead of zero | printed in this chapter (§13.4.2, p. 267) |
| variance | the mean of the squared deviations about the mean | printed in this chapter (§13.5, p. 273) |
| translation invariance | the property that a shift leaves the variance alone | an added term; the chapter proves the property and gives it no name |
Where people slip up
- "If the mean shifts by the assumed value, so does the spread." It does not. This is the single error the section exists to prevent, and the cancellation in line (4) is the place to show why.
- "Both the shift and the scale have to be undone in the standard deviation." Only the scale. Undoing the shift as well — adding the assumed value back into the spread — is the commonest wrong answer on this material.
- "Dividing by the class width makes the answer smaller, so it is an approximation." It is exact. The division is undone by the multiplication in line (4).
- "The variance is multiplied by the scale factor." By its square. The standard deviation gets the plain factor; the variance gets the square. Example 12 has a factor of 10 and a factor of 100 in front of its bracket.
- "The assumed value must be near the mean for the answer to be right." It must be near for the arithmetic to be small. The result is correct for any choice; a badly chosen one simply loses the benefit.
- "The relation between the two standard deviations holds for any scale factor." As printed it is stated with the factor multiplying directly, which reads correctly because a class width is positive. A negative scale factor would reverse the sign of the transformed observations and the relation would need the factor's size rather than the factor itself, since a standard deviation is never negative. The chapter never uses a negative factor; say the condition rather than leaving it implicit.
- "Example 12 is a new distribution." It is Example 10's distribution recomputed. The agreement of the two answers is the point.
Ask your teacher a person
Your teacher reads this and writes back, usually within a day. For an instant answer, use Ask the video in the sidebar.
Your class sees the question and the answer. Only your teacher sees that it was you.
No questions on this topic yet.
Worked answers: Exercise 13.1 · Exercise 13.2 · Miscellaneous Exercise · this video explains Exercise 13.2 Q6, Exercise 13.2 Q9, Miscellaneous Exercise Q3, Miscellaneous Exercise Q4
Transcript2,660 words
There is a trick for making the arithmetic of a variance small, and it is usually taught as one move. Take each observation, subtract some convenient value, and divide by some convenient factor. One line, one name, one recipe. But look at what is written there. A subtraction, and then a division. Those are two different operations, and the whole difficulty of this topic is that they do two different things to the spread.
One of them is free. The other one costs you something, and you have to pay it back at the end. So before touching a single table, we are going to take them apart and ask what each one does on its own. Here is a row of observations on an axis. Subtracting the same number from every one of them slides the whole row bodily along, without changing its shape.
Now measure something. Not the observations - the distance between two of them. That distance was three before the slide, and it is three after it. And that is not a lucky pair. Every difference between two observations is unchanged, because both ends moved by the same amount and a difference cannot tell. Written out for three observations at two, five and eleven, the gaps are three, nine and six, and after any slide you like they are still three, nine and six.
Three observations make three gaps, five make ten, and eight make twenty-eight, and none of them notices. Now here is why that settles the variance. A variance is built out of squared deviations, and a deviation is a difference: the observation, less the mean. When every observation slides by the same amount, the mean slides by exactly that amount too. It has to - the mean is an average of things that all moved together.
So inside the bracket you have the observation plus the shift, less the mean plus the shift. The shift cancels. Not approximately, not nearly - it is the same symbol added and subtracted, and it is gone. Every deviation is the deviation it was, every square is the square it was, and the variance is the number it was. Over sixteen hundred and eighty trials of records and shifts, the variance failed to survive a shift zero times.
The mean, on the same sixteen hundred and eighty trials, moved by exactly the shift every single time. So the shift is doing something. It is just not doing it to the spread. Now the other operation. Multiply every observation by a common factor - or divide, which is multiplying by a fraction. The row does not slide this time. It stretches, or it shrinks, about the origin. And now measure that same distance again.
It was three; multiply everything by two and it is six. Every gap is multiplied by the factor, because both ends were multiplied by the factor and a difference of two multiples is a multiple of the difference. Over two thousand two hundred and forty trials the gap list came out as the old gaps times the factor, and it failed zero times. So this operation is not free. It changes the very thing a variance is built from.
So the spread changes - but by how much? This is the single place people go wrong, so let us be slow. Every deviation is multiplied by the factor. The variance averages the SQUARES of the deviations. A square of a multiple is the multiple of the square - by the factor SQUARED. So the variance is multiplied by the factor squared, and the standard deviation, which is a root, gets the plain factor back.
Over the same two thousand two hundred and forty trials, that held with zero exceptions. And giving the variance the plain factor instead was right on eight hundred of those trials - which sounds like a lot until you ask which eight hundred. They are exactly the trials where the factor was one, or the factor was nothing at all, or the record had no spread to multiply. In other words, the only places where a number and its square are the same number.
Everywhere else - on more than half the pool - the plain factor is simply wrong. There is a condition on that last statement that almost never gets said out loud. The standard deviation gets the plain factor - but what if the factor is negative? Multiply every observation by minus four. The record flips over to the other side of the origin, and it stretches by four. Its spread is still a positive number - spread has no direction.
But minus four times the old standard deviation is negative, and no standard deviation is ever negative. Over seven hundred and twenty trials with a negative factor and a record that actually had some spread, the scaled record's spread came out positive seven hundred and twenty times out of seven hundred and twenty. So the relation needs the factor's SIZE, not the factor. And here is the interesting part: the variance relation is fine.
On eight hundred and forty negative-factor trials the variance was exactly the factor squared times the old one, because squaring throws the sign away before you ever see it. Squaring is what hides the condition. It is safe to write it without the modulus only because the factor we are about to use is a class width, and a class width is positive. Right. We now do the two operations together, on purpose, to make the numbers small - and then we have to get the real answers back.
Start with the mean. The transformed observations were made by subtracting, then dividing. To undo that you multiply, then add. Multiply the transformed mean by the factor, and add the value you subtracted. Both operations get undone, because the mean is a position and the shift moved every position. Over eight hundred table trials that recovered mean was the table's own mean, with zero disagreements. Leave out the addition and you get the wrong answer.
Leave out the multiplication and you get the wrong answer. The mean wants both. Now the spread, and this is the sentence the whole topic exists to deliver. The spread only wants ONE of them back. Multiply by the factor - and stop. Do not add the subtracted value back in. It was never taken out. Look again at what happened when we slid the row: the gaps did not move, so there was nothing to restore.
Over the same eight hundred trials, undoing only the scale gave the table's own variance, with zero disagreements. And adding the value back as well was wrong on six hundred of those eight hundred - which is exactly the six hundred where a value was actually subtracted. On the table we are about to work, the right answer is two hundred and one and the wrong route gives two hundred and sixty-six.
Out by sixty-five, which is precisely the number that was subtracted and had no business coming back. So here is the whole thing assembled. The mean is the assumed value, plus the factor times the average of the transformed column. The variance is the factor squared, times the two column totals form you already have - the total of f y squared over N, less the square of the total of f y over N.
Notice what is in that second line and what is not. The factor is in it, squared. The assumed value is nowhere in it at all. That absence is not a simplification anybody made. It is the shift cancelling, three scenes ago, showing up as a symbol that never got written down. And the second line is not a new formula either - it is the same two column totals form, with a factor squared out in front.
Let us actually use it. Seven classes running from thirty to a hundred, with frequencies three, seven, twelve, fifteen, eight, three and two. Fifty observations in all. The mid-points are thirty-five, forty-five, and so on up to ninety-five. Take sixty-five as the assumed value, and take the class width, ten, as the factor. Now watch what those mid-points turn into. Minus three, minus two, minus one, nought, one, two, three.
The whole column collapsed to seven small whole numbers, and the middle one is nothing at all. The f y column totals minus fifteen. The f y squared column totals a hundred and five. Every number I just read out is one you could add up in your head. Now put both answers back. The mean: minus fifteen over fifty is minus three tenths, times ten is minus three, added to sixty-five gives sixty-two.
The variance: a hundred over two thousand five hundred, times fifty times a hundred and five less two hundred and twenty-five. That is a hundred over two thousand five hundred, times five thousand and twenty-five, which is two hundred and one. And nowhere in that second calculation did sixty-five appear. The standard deviation is the root of two hundred and one. Two hundred and one is not a perfect square, so that root does not stop - it is a reading, not a number.
Rounded to two places it reads fourteen point one eight; truncated at two places it reads fourteen point one seven. Those are two different operations on the same root, and it is worth knowing which one produced the digits in front of you. So what did all that substituting actually buy? The claim is that the arithmetic got smaller, and a claim about size can be measured. The biggest number anywhere on the transformed table is twenty-eight, and every entry in it is a whole number.
Now do the same table without transforming anything. The deviation column, measured from sixty-two, carries entries up to two thousand one hundred and eighty-seven. The squares column carries entries up to sixty-three thousand three hundred and seventy-five. So the biggest thing a deviation column would have made you write is seventy-eight times over what the transformed table asks for, and the squares column more than two thousand two hundred times it.
That is the whole payment, and it is a payment in ink and in mistakes, not in accuracy. Nothing here is an approximation. And there is a way to be certain of that, which is better than being told. Those seven classes, and those seven frequencies, are the same seven classes and the same seven frequencies as a table we worked the long way before - written out one for one, they match.
So the two routes are answering about the same fifty observations. The long route builds a mean, then a deviation column, then a column of squares in the thousands, and lands on sixty-two and two hundred and one. The short route never builds a deviation at all, keeps every entry under thirty, and lands on sixty-two and two hundred and one. Same standard deviation, fourteen point one eight. That agreement is the evidence that the transformation was undone correctly, and it is worth something only because the two routes are genuinely different - one writes a column the other never writes.
One more thing about the assumed value, because it is a common worry. Does it have to be near the mean? Take the same table and try five different assumed values: sixty-five, thirty-five, ninety-five, nought, and minus forty. Every single one returns a mean of sixty-two and a variance of two hundred and one. Five tried, five right, none wrong. So the choice cannot make the answer wrong. What it makes is work.
The biggest entry the chosen value asks you to write is twenty-eight. For the others it is a hundred and thirty-five, a hundred and ninety-two, and then two entries that are not whole numbers at all - they land on quarters. The worst of the five asks for numbers fifty-nine times bigger than the best. So a badly chosen value is not an error. It is a wasted opportunity, which is a different thing and worth saying differently.
Both halves of this are worth stating on their own, without any table around them, because you will meet them as questions. First half. Take twenty observations whose variance is five. Their squared deviations total a hundred. Now add seventeen to every single one of them. The mean goes up by seventeen. The deviations - every one of the twenty, term by term - are identical to what they were.
The squared deviations still total a hundred, and the variance is still five. Subtract seventeen instead and exactly the same thing happens. So if you are told a data set has some variance and then told every value was increased by a constant, you already have the answer, and the answer is that nothing happened. Second half. Same twenty observations, variance five, squared deviations totalling a hundred. Now double every one of them.
The squared deviations total four hundred, and the variance is twenty. Twenty is four times five, not twice five - the factor arrived squared, exactly as promised. Here is the same thing in a standard deviation. Six observations with a mean of eight and a standard deviation of four. Multiply every one by three. The mean becomes twenty-four, the variance becomes a hundred and forty-four, and its root is exactly twelve.
The mean took the plain factor and the spread took the plain factor, because a standard deviation is a root - but underneath, the variance took nine. And one edge case, since people ask about it: multiplying by nothing is allowed, and it does not break the rule. Every observation collapses to nought, the mean is nought times the old mean, and the variance is nought squared times the old variance.
The rule survives; it is the data that does not. Two more tables, to show the two things that can happen. First: nine values, sixty through sixty-eight, standing for a hundred observations. Take sixty-four as the assumed value and take the factor as one, since the values are already one apart. The f y column totals nothing at all - zero. So the mean is sixty-four exactly, with nothing to add back.
The f y squared column totals two hundred and eighty-six, so the variance is two point eight six exactly, and the standard deviation reads one point six nine. Second: nine classes of heights in centimetres, sixty children. Take ninety-two point five as the assumed value and five as the factor. This time the f y column totals six, not nought, so the mean is ninety-two point five plus a tenth of five - ninety-three.
The f y squared column totals two hundred and fifty-four, and the variance works out at a hundred and five point five eight - a reading, because that one does not stop either, it is a hundred and five point five with a three running on for ever. The standard deviation reads ten point two eight. So here is the whole topic in one picture. Two operations went out. A slide, which moved every position and moved no gap.
And a stretch, which moved every gap by its factor. The mean is a position, so it feels both, and both have to be undone in it. The spread is built out of gaps, so it never felt the slide, and only the stretch has to be undone in it. That asymmetry is not a rule to memorise. It is a consequence of what each of the two things is made of, and if you can see that, you will never add the assumed value back into a standard deviation again.
One last caution, because it is the one that usually goes unsaid: the spread takes the factor's size, not the factor. With a positive width nobody notices the difference. With a negative one, everybody would.
Where this fits
Taken from the notes each video was made from, not from the reading order — these are the ideas this one rests on and the ones that later rest on it.
Builds on
- Carrying the frequencies through, for discrete and for grouped dataClass 11 · Ch 13, Statistics
- The same procedure once the data arrive already groupedClass 11 · Ch 13, Statistics
Either side of this one
- Every description of a happening picks out a part of the outcome listClass 11 · Ch 14, Probability