PrepShorts · Study sheet · Class 11 Mathematics · Chapter 13, Statistics
This video could not be loaded. Reload the page to try again.
Sign in with Google21 min.
Keep your place in this chapter — sign in, it’s free.Sign in
A table of bands, unlike a table of values, loses the readings themselves. Every mean deviation computed from it afterwards assumes each reading sat at its band's exact midpoint.
The idea
Grouping does not change the idea; it changes the bookkeeping and, in one case, the truth. A frequency is a repetition count, so every total in §13.4.1 becomes a weighted total and the count n becomes the frequency total N — that is all that happens to a discrete distribution, and no information is lost. A continuous distribution is different: the individual observations are gone, and the method survives only by pretending every value in a class sits at the class mid-point. What §13.4.2 then computes is the mean deviation of a reconstructed data set, not of the one that was measured. The assumed mean and the step-deviations are arithmetic relief and nothing more, and the Note on p. 268 is careful to say how narrowly that relief applies. The median formula is not relief at all: it is a second assumption of the same family as the mid-point one, spreading a class's frequency evenly across its width so that a position inside it can be interpolated.
What you should be able to do
- Compute the mean deviation about the mean of a discrete frequency distribution, using a worked table with the products written out
- Locate the median of a discrete frequency distribution from its cumulative frequencies
- Replace each class of a continuous distribution by its mid-point and say precisely what assumption that makes
- Compute the mean of a continuous distribution by the step-deviation method, and state what the Note on p. 268 restricts that method to
- Locate the median class of a continuous distribution and apply the interpolation formula, explaining what the formula assumes about the interior of that class
- Compute the mean deviation about the median for a continuous distribution
- Convert a distribution with gaps between its classes into a continuous one, as Exercise 13.1 item 12 requires
Words to know
| Term | Definition in one line | First introduced |
|---|---|---|
| discrete frequency distribution | a table of distinct values with the number of times each occurs | printed in this chapter (§13.4.2, p. 262) |
| continuous frequency distribution | a table of class intervals, with no gaps between them, and their frequencies | printed in this chapter (§13.4.2, p. 265) |
| frequency | the number of times a value or a class occurs | printed throughout this chapter (§13.4.2, p. 262) |
| cumulative frequency | the running total of the frequencies down the table | printed in this chapter (§13.4.2, p. 263) |
| class interval | one of the bands the observations are sorted into | printed in this chapter (§13.4.2, p. 265) |
| mid-point | the value standing for a whole class in every computation | printed in this chapter (§13.4.2, p. 265) |
| median class | the class whose cumulative frequency first reaches half the frequency total | printed in this chapter from §13.4.2, p. 268, where the process is set up; the formula and its four symbols follow on p. 269 |
| assumed mean | a convenient value the deviations are measured from instead of the true mean | printed in this chapter (§13.4.2, p. 266) |
| step-deviation | a deviation from the assumed mean divided by the common factor | printed in this chapter (§13.4.2, p. 267) |
| common factor | the number all the deviations share, divided out to shrink them | printed in this chapter (§13.4.2, p. 267) |
| shortcut method | the chapter's name for the step-deviation route to the mean | printed in this chapter (§13.4.2, p. 266) |
| N | the chapter's symbol for the total of all the frequencies | printed in this chapter (§13.4.2, p. 263) |
| mid-point assumption | the claim that a class's whole frequency may be treated as sitting at its mid-point | an added label; the chapter makes the assumption in words on p. 265 and gives it no name |
Where people slip up
- "The grouped formula is a new definition." It is the same average of distances, with repeated distances collected. Write out one small frequency longhand once and the weighting stops looking like a rule.
- "N is the number of classes." N is the total of the frequencies — 40 in Example 4 and Example 6, 30 in Example 5, 50 in Example 7. The number of classes is the number of rows and appears only as the upper limit of the summation.
- "The mid-point is where the observations actually are." It is where they are assumed to be. In Example 6 the answer 10 is the mean deviation of forty values placed exactly at 15, 25, 35 and so on, and no real measurement guarantees that.
- "Half of N will be one of the cumulative frequencies." Usually it will not. The rule is to take the first cumulative frequency that reaches or passes it; in Example 5 half of 30 is 15 and the cumulative frequencies step from 14 straight to 18.
- "The median is the mid-point of the median class." In Example 7 the median class is 20–30, whose mid-point is 25, and the median is 28. The formula exists precisely because the median is generally not the mid-point.
- "The assumed mean has to be the true mean." It is chosen for convenience — a value near the middle of the table. The formula corrects for whatever was assumed. Table 13.5 happens to assume the true mean and so corrects by zero.
- "The step-deviation table also gives the mean deviation." The Note on p. 268 rules this out. The step-deviation columns feed the mean only.
- "Class intervals like 16–20 and 21–25 are continuous." They have a gap between 20 and 21, and Exercise 13.1 item 12 has to close it before the median formula can be applied.
Ask your teacher a person
Your teacher reads this and writes back, usually within a day. For an instant answer, use Ask the video in the sidebar.
Your class sees the question and the answer. Only your teacher sees that it was you.
No questions on this topic yet.
Worked answers: Exercise 13.1 · Exercise 13.2 · Miscellaneous Exercise · this video explains Exercise 13.1 Q5, Exercise 13.1 Q6, Exercise 13.1 Q7, Exercise 13.1 Q8, Exercise 13.1 Q9, Exercise 13.1 Q10, Exercise 13.1 Q11, Exercise 13.1 Q12
Transcript2,996 words
Sometimes the readings do not arrive as a list. They arrive already counted, in a table, and a table comes in two shapes. The first is a column of values with a column saying how many times each one occurred. The second is a column of bands - nought to ten, ten to twenty, and so on - with a column saying how many readings fell in each band. The difference between those two looks small and it is not.
The first shape loses nothing at all. The second one loses the readings themselves, and everything awkward about this topic comes from that. Start with the shape that loses nothing. A frequency is a repetition count and nothing more. If a value of six occurred ten times, then ten of the distances in step three are the same distance, ten times over. Adding it once and multiplying by ten is faster than writing it out ten times, and that is the entire content of the weighted formula.
So every total becomes a weighted total, and the count becomes the frequency total. Nothing has been defined twice. It is worth saying out loud, because the weighted formula is usually met as a second rule to memorise rather than as the first one written more compactly. Here is such a table. Six values - two, five, six, eight, ten and twelve - occurring two, eight, ten, seven, eight and five times.
Six rows, and forty observations. Multiply each value by its frequency and the products total three hundred, so the mean is three hundred over forty, which is seven point five. Notice what that is: seven point five is not one of the six values in the table. It is a point in the middle of the readings, not a member of them. Now the distances from seven point five. Five point five, two point five, one point five, nought point five, two point five and four point five.
Weighted by their frequencies they total ninety-two, and ninety-two over forty is two point three. The one place this goes wrong is the divisor. The table has six rows and forty observations, and it is forty you divide by. Divide by six instead and the answer is forty-six thirds, which is fifteen point three three and so on - twenty thirds of the size it should be. And the check that this is really the ungrouped procedure with the repeats collected is available, so take it: write the table out longhand as forty separate readings, run the plain four steps on those, and the mean and the answer come out identical.
That divisor deserves its own moment, because the letter N is doing double duty in most students' heads. N is the total of the frequencies. It is not the number of rows. On the four tables this video works, the row counts are six, eight, seven and six. The frequency totals are forty, thirty, forty and fifty. The number of rows appears in the working exactly once, as how far down the column you have to add, and it never appears as a divisor.
Now the other centre, on a table of the same shape. The median is the value with as many readings below it as above it, so you need to know how many readings you have passed by the time you reach each row. That is a running total down the frequency column, and it has a name: the cumulative frequency. Then the rule is: take the value at the first row whose running total reaches half of N.
Reaches, or passes. Not lands on. Here is a table of eight values with frequencies three, four, five, two, four, five, four and three. The running totals are three, seven, twelve, fourteen, eighteen, twenty-three, twenty-seven and thirty. So N is thirty and half of N is fifteen. Look at the column: it steps from fourteen straight to eighteen. Fifteen is not on it, and that is normal rather than unlucky.
The first running total to reach fifteen is eighteen, on the row whose value is thirteen, so the median is thirteen. And unlike the seven point five from the first table, thirteen IS one of the listed values. The distances from it are ten, seven, four, one, nought, two, eight and nine; weighted they total a hundred and forty-nine; and the answer is a hundred and forty-nine over thirty. That is recurring, so the two printing habits part company on it again: rounded to two places it reads four point nine seven, cut off it reads four point nine six.
Written out longhand, those thirty readings have exactly that median and exactly that answer. Two small tables to pin down what that rule actually says, because both of the tables above sidestep it. Take three values occurring two, two and four times. N is eight, half of N is four, and the running totals are two, four and eight - so here half of N IS on the column, exactly.
The rule says the first row to REACH it, which is the second row, whose value is two. A rule that said the first row to PASS it would take the third row instead, and get a different median. Second table: frequencies four, six and two. N is twelve and half of N is six. Six is one of the frequencies, sitting right there in the column you were just reading.
It is not one of the running totals, which are four, ten and twelve. The test has to be looking at the right column, and the only way to know it is is to hand it a table where the two columns disagree. Now the second shape of table, where the rows are bands rather than values. Every reading that fell between ten and twenty is on one row, and which of them was eleven and which was nineteen is not recorded anywhere.
The readings are gone. So the procedure cannot be run at all - unless you decide where inside the band each reading is going to be taken to have been. And the decision the method makes is: all of them, at the middle of the band. That is not a technicality to be waved through. It is an assumption, it is the only thing that makes the next four columns possible, and it is worth about ninety seconds of your attention before we use it.
Here is what the assumption costs, measured rather than described. Take a table of seven bands ten wide, from ten to eighty, with frequencies two, three, eight, fourteen, eight, three and two. Now build two real records that both produce exactly that table. In the first, every reading sits on its band's mid-point. In the second, each band's readings are pushed outwards instead, to sit just inside its two ends.
Sort either of them into those seven bands and you get back exactly those seven frequencies, forty readings each time. The table cannot tell them apart. But the first record's own mean is forty-five and its own answer is ten. The second record's mean is forty-five point two and its answer is eleven point four. The table, worked by the method, says ten. It is exactly right about the first record and wrong about the second, and it has no way of knowing which one it was handed.
With that said plainly, the working itself is short. The mid-points of those seven bands are fifteen, twenty-five, thirty-five, forty-five, fifty-five, sixty-five and seventy-five. Weighted by the frequencies they total eighteen hundred, and eighteen hundred over forty is forty-five. The distances from forty-five are thirty, twenty, ten, nought, ten, twenty and thirty. Weighted they total four hundred, and the answer is four hundred over forty, which is ten. Every number in that came out whole, which is not luck either: the frequencies read the same forwards and backwards, so the table is symmetric about its middle band.
The mid-points there were small and friendly. They are often not. So there is a piece of arithmetic relief, and it is worth being precise about what it relieves and what it does not. Pick any convenient value near the middle of the table and call it the assumed mean. Measure every mid-point from that instead of from nought. On a number line this is one move: slide the origin to sit under your assumed value, and relabel the same ticks.
The points have not moved. Only the numbers written under them have. Then do it once more, to the scale rather than the origin. If the bands are ten wide, every one of those deviations is a multiple of ten, so divide them all by ten and work with the small numbers. Those are the step-deviations. On the table we just did, with the assumed mean at forty-five and a common factor of ten, the step-deviations are minus three, minus two, minus one, nought, one, two and three.
Weighted by the frequencies they total nought. So the mean comes back as forty-five plus ten times nought over forty, which is forty-five. The same answer the long way round gave. That nought in the middle is worth stopping on, because it looks like a convenience and it is a consequence. The deviations about the TRUE mean always add to nothing. Forty-five is the true mean of that table, so of course they did.
Assume something else and the nought does not appear, and the formula corrects for whatever was assumed. Assume thirty-five, still with a factor of ten: the weighted step total is forty, and thirty-five plus ten times forty over forty is forty-five again. Keep forty-five but halve the factor to five: the weighted total is nought and the mean is forty-five. Assume sixty with a factor of twenty: the weighted total is minus thirty, and the mean is forty-five.
Four different assumptions, four different columns, one answer. The assumed mean does not have to be right, it has to be convenient. And now the trap, which is the reason this shortcut needs a warning attached rather than just a demonstration. The step-deviation column gets you the mean. It does not get you the answer. A student who keeps going down that same column, takes the sizes and averages them, gets one.
One, when the answer is ten. It is out by the common factor, which is exactly what dividing by the common factor would do. Multiply by ten and you land on ten, which makes it look like a rule. It is not a rule. Do the same thing from an assumed mean of thirty-five and the step column gives one point three five; times ten that is thirteen point five, and the answer is still ten.
The repair only worked the first time because the assumed mean happened to be the true one. Build the distance column from the real mid-points and the real mean, every time. The other centre on a banded table is a different kind of problem altogether. Here is a table of six bands ten wide, from nought to sixty, with frequencies six, seven, fifteen, sixteen, four and two. The running totals are six, thirteen, twenty-eight, forty-four, forty-eight and fifty.
N is fifty and half of N is twenty-five. The first running total to reach twenty-five is twenty-eight, so the median is somewhere inside the band from twenty to thirty. Somewhere inside it - and now you have to say where. Four numbers do that. The band's lower limit, twenty. The running total before it, thirteen. The band's own frequency, fifteen. And the band's width, ten. Start at twenty, and move in by the shortfall - twenty-five less thirteen, which is twelve - as a fraction of the band's own fifteen, scaled by the width.
Twelve fifteenths of ten is eight, so the median is twenty-eight. And look where that is not: the mid-point of that band is twenty-five. The median is three above it. That last step is a second assumption, of exactly the same family as the mid-point one, and it is almost never named. Moving in by twelve fifteenths of the width only makes sense if the fifteen readings in that band are spread evenly across it - so that the twelfth of them is twelve fifteenths of the way along.
You can build that record and check. Lay each band's readings at equal steps across its own width, all six bands, and you get fifty readings whose twenty-fifth, in order, is exactly twenty-eight. The formula is describing that record. Now build a second one, identical except that the fifteen readings in the middle band are bunched near its top instead of spread. Same fifty readings, same six bands, and sorted into them it produces the same table - the formula still says twenty-eight.
The twenty-fifth reading of that record is twenty-nine. So the formula is a straight-line guess about the inside of one band, and it is the right guess to make when you have nothing else, and it is still a guess. The rest is the ordinary working. The mid-points of the six bands are five, fifteen, twenty-five, thirty-five, forty-five and fifty-five. Their distances from twenty-eight are twenty-three, thirteen, three, seven, seventeen and twenty-seven.
Weighted by the frequencies they total five hundred and eight, and five hundred and eight over fifty is ten point one six. Which is exact to two places, so nothing has to be rounded at all. One last table, because it breaks a rule the others quietly obeyed. Its bands run sixteen to twenty, then twenty-one to twenty-five, then twenty-six to thirty, and so on. There is a gap of one between every band and the next.
The formula reads a lower limit and a width off the median band, and a band whose lower limit does not touch the band below it is describing a stretch of the line where nothing was measured and nothing is claimed. So close the gaps first: half a unit off every lower limit and half a unit onto every upper one. Sixteen to twenty becomes fifteen point five to twenty point five, and the gaps are gone.
Here is the surprise, though. On this particular table it changes the answer not at all. The median band is the same one either way, and the median is thirty-eight either way. The mid-points do not move at all - taking a half unit off the bottom and putting a half unit on the top leaves the middle exactly where it was. And the median agrees only because of a coincidence in the numbers: the shortfall here is thirteen and the median band's frequency is twenty-six, so the shortfall is exactly half of it.
Change one frequency - drop the fourth band from fourteen readings to ten - and the shortfall over the frequency becomes fifteen twenty-sixths, and now the two routes give four hundred and ninety-eight thirteenths and four hundred and ninety-nine thirteenths, a thirteenth apart. So close the gaps because the formula is not entitled to that lower limit otherwise, not because the number always moves. Eight questions to finish, and they are worth sorting by what they ask for rather than by their order.
Two are tables of values asked about the mean. The first has twenty-five readings, a mean of fourteen and an answer of a hundred and fifty-eight over twenty-five, which is six point three two. The second has eighty readings, a mean of fifty, and an answer of exactly sixteen. Two are tables of values asked about the median. The third has running totals eight, fourteen, sixteen, eighteen, twenty and twenty-six, so N is twenty-six, half of N is thirteen, the first total to reach it is fourteen, and the median is seven; the answer is three point two three.
The fourth has running totals three, eight, fourteen, twenty-one and twenty-nine, so half of N is fourteen and a half, the first total to reach it is twenty-one, and the median is thirty; the answer is five point one nought. Neither of those halves is on its column either. Two are banded tables asked about the mean. One has fifty readings, mid-point products totalling seventeen thousand nine hundred, a mean of three hundred and fifty-eight, and an answer of a hundred and fifty-seven point nine two.
The other has a hundred readings, products totalling twelve thousand five hundred and thirty, a mean of a hundred and twenty-five point three, and an answer of eleven point two eight eight. One is a banded table asked about the median: its median band is twenty to thirty, its median is a hundred and ninety-five sevenths - about twenty-seven point eight six - and its answer is about ten point three four.
And the eighth is the one whose bands had gaps, which after closing them has a median of thirty-eight and an answer of seven point three five. So: what did grouping actually change? For a table of values, nothing but the bookkeeping. Every total became a weighted total, the count became the frequency total, and the answer is the same answer the longhand list gives, which you can check on any of them in a minute.
For a table of bands, something real. The readings are gone, and both centres are now computed about a reconstructed set of readings rather than about the ones that were measured. The mid-point assumption puts every reading at the middle of its band. The median formula spreads a band's readings evenly across it. Neither is a rule of arithmetic; both are decisions about data you no longer have. They are the right decisions - there is nothing better to do with a table - but a student who knows that will read a grouped answer with exactly the confidence it deserves, and one who does not will read it as a measurement.
Where this fits
Taken from the notes each video was made from, not from the reading order — these are the ideas this one rests on and the ones that later rest on it.
Builds on
- Working it out about the mean and about the median, for a plain listClass 11 · Ch 13, Statistics
Comes up again in
- Where this measure breaks down, and why another was neededClass 11 · Ch 13, Statistics
- Taking the root to get the spread back into the units of the dataClass 11 · Ch 13, Statistics
- Carrying the frequencies through, for discrete and for grouped dataClass 11 · Ch 13, Statistics
- Shifting and scaling the observations to make the arithmetic smallClass 11 · Ch 13, Statistics