PrepShorts · Study sheet · Class 11 Mathematics · Chapter 13, StatisticsPrepShorts

Chapter 13 · Statistics

Carrying the frequencies through, for discrete and for grouped data

Squaring instead of discarding the sign16 min

This video could not be loaded. Reload the page to try again.

Sign in with Google

16 min.

A deviation column cannot begin until the mean has already been written out in full, and nothing guarantees that mean will be a tidy number.

The idea

Frequencies enter the variance exactly as they entered the mean deviation, as repetition counts used as weights, and for the two worked examples that costs nothing extra. What grouping does force is an ordering problem: the mean must be computed and written down before a single squared deviation can be, and nothing guarantees it will be a whole number — the chapter's own Example 11 supplies one that is not. §13.5.3 removes that dependence algebraically. Expanding the square splits the sum into three, the middle one collapses because the weighted total of the observations is the count times the mean, and what survives is the mean of the squares less the square of the mean — an expression in two column totals, with no deviation column anywhere in it. This is the payoff §13.4.3 promised when it complained that a modulus could not be manipulated.

What you should be able to do

  • Write, in terms of the frequencies, the variance of a discrete frequency distribution and its standard deviation
  • Compute both for a discrete distribution using a five-column working table
  • Replace the classes of a continuous distribution by mid-points and compute the variance
  • Derive the second form of the variance by expanding the square and using the weighted total of the observations
  • State the identity that the variance is the mean of the squares less the square of the mean, and read off the inequality it forces between those two quantities
  • Compute a standard deviation from the two column totals alone, without a deviation column
  • Decide, for a given distribution, which of the two routes will be less work

Words to know

TermDefinition in one lineFirst introduced
discrete frequency distributiona table of distinct values with the number of times each occursprinted in this chapter (§13.4.2, p. 262; §13.5.2 heading, p. 275)
continuous frequency distributiona table of class intervals with no gaps, and their frequenciesprinted in this chapter (§13.4.2, p. 265; §13.5.3 heading, p. 276)
mid-pointthe value standing for a whole class in every computationprinted in this chapter (§13.4.2, p. 265; §13.5.3, p. 276)
variancethe weighted mean of the squared deviations about the meanprinted in this chapter (§13.5, p. 273)
standard deviationthe non-negative square root of the varianceprinted in this chapter (§13.5.1, p. 274)
frequencythe number of times a value or class occursprinted throughout this chapter (§13.4.2, p. 262)
Nthe chapter's symbol for the total of the frequenciesprinted in this chapter (§13.4.2, p. 263)
another formula for standard deviationthe chapter's own heading for the expression in the two column totalsprinted in this chapter (§13.5.3, p. 276)
deviation-free formthe property of the second formula that no deviation column is neededan added label; the chapter derives the form and gives it no name

Where people slip up

  • "The two formulas give different variances." They are the same quantity, and Example 10 and Example 12 work the same distribution by two different routes and reach 201 both times. Be exact about which two routes those are: Example 10 goes by the definition and Example 12 by the step-deviation shortcut of §13.5.4, so the pair evidences definition-against-shortcut. The one worked example that applies the alternative formula itself is Example 11. If a student's two routes disagree, the arithmetic is wrong, not the theory.
  • "The second formula is an approximation." It is an identity, derived in four algebraic steps with nothing dropped.
  • "You still need the mean for formula (3)." You do not. The mean was substituted out; the formula needs only the two column totals and N. That is precisely why it exists.
  • "Squaring the observations is the same as squaring the deviations." It is not, and the identity says by exactly how much they differ — the square of the mean.
  • "N is the number of rows in the table." N is the total of the frequency column: 30 in Example 9, 50 in Example 10, 48 in Example 11.
  • "Since the mid-points stand in for the classes, the answer is approximate." The computation is exact for the reconstructed data; the approximation lives in the reconstruction, not in the arithmetic. Say which is which.
  • "The observation with the largest frequency contributes most." Not necessarily. In Example 9 the value 32 occurs once and contributes 324 of the 1374, more than the value 11 contributes with a frequency of nine.
Transcript2,217 words

Here is a column of numbers with a second column beside it, and the second column is not a measurement. It is a count of repeats. The value four appears three times, the value eight appears five times, and writing them out that way would take thirty lines. The table takes seven. So the first thing to be clear about is what N means. N is the total of the frequency column, not the number of rows.

This table has seven rows and stands for thirty observations. And the formula that follows is not a new formula at all: it is the definition you already have, with each squared deviation counted as many times as its value occurs. That sentence is easy to say and easy to believe without checking, so it is worth checking. Take a table, write out the list it stands for, and run the plain-list machinery on the list.

Then run the weighted machinery on the table. If a frequency really is a repeat count, the two must agree everywhere. Over two hundred and eighty tables, most of them chosen by nobody, the two routes disagree zero times, on the total, on the total of the squares, on the mean and on the variance. So nothing here is new mathematics. It is a shorter way of writing the same sum, and everything already proved about the sum still holds.

Take the table properly. Seven values: four, eight, eleven, seventeen, twenty, twenty-four and thirty-two, with frequencies three, five, nine, five, four, three and one. The frequencies total thirty, so N is thirty. Multiply each value by its frequency and add: the products come to four hundred and twenty, so the mean is fourteen. Now the deviations from fourteen: minus ten, minus six, minus three, three, six, ten and eighteen. Square each, multiply by the frequency, and add.

That total is one thousand three hundred and seventy-four. Divide by thirty and the variance is forty-five point eight, and the standard deviation reads six point seven seven. Before moving on, look at the column you just added. The seven contributions are three hundred, one hundred and eighty, eighty-one, forty-five, one hundred and forty-four, three hundred, and three hundred and twenty-four. The largest of those is the last one. It comes from the value thirty-two, which occurs once.

One observation out of thirty carries three hundred and twenty-four of the one thousand three hundred and seventy-four -- twenty-three point five eight per cent of the whole measure. Meanwhile the value eleven, which occurs nine times, carries eighty-one between all nine of them. The row with the largest frequency does not carry the largest contribution, and it is not close. That is squaring doing what squaring does, and a frequency column hides it, because the tall row and the heavy row are not the same row.

Now the harder kind of table, where the left column is not values but intervals. Thirty to forty, forty to fifty, and so on up to ninety to a hundred. You do not know where inside its class any observation actually sat. So each class is replaced by its mid-point, and from there the method is exactly the one you just used. Thirty to forty becomes thirty-five, forty to fifty becomes forty-five, and the seven mid-points run from thirty-five to ninety-five in tens.

No new idea arrives at this step. What arrives is a cost, and it is worth being exact about what the cost is and what it is not. Here is the exact statement. The answer this table gives is the exact answer for one particular record: the record with every observation sitting precisely on its class mid-point. For that record the arithmetic is not an approximation at all. But other records produce the same table.

Build one that pushes each class's observations out towards the two ends instead of the middle. Same classes, same seven frequencies, same table on the page. Its own variance is two hundred and twenty-eight point three eight, not two hundred and one. So the table is exactly right about one of those records and wrong about the other, and there is nothing in the table to say which one you are holding.

The approximation lives in the reconstruction, not in the sum. With that said, work the grouped table. The frequencies are three, seven, twelve, fifteen, eight, three and two, totalling fifty. Mid-point times frequency, added down the column, comes to three thousand one hundred, so the mean is sixty-two. The squared deviations from sixty-two are seven hundred and twenty-nine, two hundred and eighty-nine, forty-nine, nine, one hundred and sixty-nine, five hundred and twenty-nine, and one thousand and eighty-nine.

Weighted, they total ten thousand and fifty. Divide by fifty: the variance is two hundred and one, and the standard deviation reads fourteen point one eight. Two tables done, and both of them went smoothly. Now look at what the smoothness was hiding. Go back to the order you did the work in. You filled the frequency column, you filled the products column, you added it, you divided by N -- and only then could you start writing deviations.

The third column cannot begin until the mean exists in writing. In both of those tables the mean came out a whole number, fourteen and sixty-two, so nobody noticed the dependence. Nothing guarantees that. Here is a table where it does not happen. Five values -- three, eight, thirteen, eighteen and twenty-three -- with frequencies seven, ten, fifteen, ten and six. N is forty-eight and the products total six hundred and fourteen.

So the mean is six hundred and fourteen over forty-eight, which is three hundred and seven over twenty-four. Twenty-four carries a factor of three, and a fraction whose denominator carries anything other than twos and fives has a decimal that never terminates. This one runs twelve point seven nine one six six six, and keeps going. Every entry in a deviation column built on that is a rounding, and every square of a rounding is a bigger rounding.

It is worth knowing how unusual that is here. Of the ten tables and lists checked behind this video, exactly one has a mean that does not terminate, and it is this one. That is not luck. The data was chosen so that the next idea would visibly earn its keep. So: get rid of the dependence. The quantity we want is the mean of f times the square of x minus the mean.

The trouble is entirely in that bracket, so open it. The square of x minus m is x squared, plus m squared, minus twice m times x. Nothing subtle has happened. That is one line of algebra you have known for years, and it is the line the modulus would never have allowed. This is what squaring bought. Now put the frequencies back and split the sum. A sum of three things, added up across the rows, is three sums added together.

The first is the total of f times x squared. The second is m squared times the total of the frequencies, because m squared is the same in every row and comes outside. The third is twice m, also the same in every row, times the total of f times x. Look at what those three are made of. The first is a column total. The second contains N. And the third contains the other column total, the one you added to get the mean in the first place.

Now the substitution that does the work. The total of f times x is N times the mean. That is not a new fact; it is the definition of the mean, rearranged. Put it into the second term: m squared times N. Put it into the third: twice m, times N times m, which is twice N times m squared. So the second term is N m squared and the third is two N m squared, and they are being subtracted.

One cancels the other and leaves minus N m squared. The middle of the expression has emptied itself out. Divide the whole thing by N and see what is left. The variance is the total of f times x squared, over N, minus the square of the mean. In words: the mean of the squares, less the square of the mean. Two quantities, and neither of them is a deviation.

Checked over the same two hundred and eighty tables, that holds with zero exceptions. And if it feels familiar, it should. The mean squared distance from any centre you like is the variance plus the square of the distance from that centre to the mean. Read that at a centre of nought and you get exactly this. Nothing new was proved here. An old statement was read at a new place.

Before using it, read the identity backwards, because it says something the working never mentions. A variance is a mean of squares, so it cannot be negative. The mean of the squares, less the square of the mean, cannot be negative either. So for any data whatever, the mean of the squares is at least the square of the mean. Over two hundred and eighty tables it falls below zero times.

And equality is not impossible: it happens exactly when every observation is the same value. That claim would be worth nothing on a pool where equality never happens, so sixty constant tables are mixed in on purpose. Sixty tables meet the equality, sixty tables have one value, and zero belong to one group without the other. There is still a mean in the expression, and it can go too. The mean is the first column total divided by N, so write it that way and clear the denominators.

The standard deviation is one over N, times the square root of N times the second column total, minus the square of the first column total. Read the ingredients. The total of f x, the total of f x squared, and N. Three numbers, two of them column totals you were going to add anyway. No mean, no deviation column, and no waiting. You can fill the whole table left to right without stopping in the middle to divide.

Take the table whose mean would not terminate. The frequency column totals forty-eight. The f x column totals six hundred and fourteen. For the last column, square each value first -- nine, sixty-four, a hundred and sixty-nine, three hundred and twenty-four, five hundred and twenty-nine -- then weight and add, giving nine thousand six hundred and fifty-two. Now: forty-eight times nine thousand six hundred and fifty-two is four hundred and sixty-three thousand two hundred and ninety-six.

Six hundred and fourteen squared is three hundred and seventy-six thousand nine hundred and ninety-six. Subtract: eighty-six thousand three hundred. Its square root reads two hundred and ninety-three point seven seven, and a forty-eighth of that is six point one two. Every number in that working is a whole number, and the mean that would not terminate never appeared. It would be dishonest to stop there, because that route costs something too.

Look at the subtraction again. Four hundred and sixty-three thousand, less three hundred and seventy-six thousand, leaving eighty-six thousand. Eighty-one point three seven per cent of the larger number is thrown away by that one subtraction, and what survives is the answer. That makes the answer sensitive to the two totals in a way the deviation column is not. How sensitive is exact. An error of one in the second column total moves the variance by one over N.

An error of one in the first column total moves it by twice that total, plus one, over N squared. The ratio between them is twice the mean, plus one over N. For this table that is twenty-five point six. A slip of one in the column you add first costs twenty-five times what the same slip costs in the column you add second. Both routes are exact; they are not equally forgiving.

So the choice between them is a real choice, and it is decided before you start. If the mean is going to come out clean, the deviation column is short, honest and easy to check. A list of eight with mean nine and variance nine point two five, or a table with mean nineteen and variance forty-three point four, wants that route. If the mean is going to be a fraction that will not stop, the two totals win, and they win by more the longer the table is.

One more thing worth noticing, about tables that arrive with gaps in them. Some come with gaps between the classes -- thirty-three to thirty-six, then thirty-seven to forty -- and the standard instruction is to close those gaps by shifting every limit half a unit. Do it, because a continuous table is what the method assumes. But it changes nothing. Move both limits of a class by the same amount in opposite directions and the mid-point does not move at all -- checked across thirty-six different gapped tables, not one mid-point shifted.

Same mid-points, same mean, same variance, before and after. The instruction tidies the table. It does not touch the answer.

Where this fits

Taken from the notes each video was made from, not from the reading order — these are the ideas this one rests on and the ones that later rest on it.

Builds on

Comes up again in

The book

Open in a new tab