PrepShorts · Study sheet · Class 10 Mathematics · Chapter 13, StatisticsPrepShorts

Chapter 13 · Statistics

Running totals, and converting between the two cumulative tables

Median of grouped data15 min

This video could not be loaded. Reload the page to try again.

Sign in with Google

15 min.

A frequency table answers 'how many are here?' and cannot answer 'which class is the 27th script in?'. A running total answers the second, and it can be run in either direction - so the same distribution produces two different-looking columns whose entries add to the total at every shared boundary. That is not a rule of arithmetic; it is what a clean split means.

The idea

A frequency table answers "how many are in this class?". A positional average asks a different question — "which observation is the middle one?" — and that question cannot be answered from per-class counts until they are accumulated. Cumulative frequency is that accumulation, and it is what makes a median computable at all. The chapter builds two accumulations, upward and downward, and they look like two separate facts; they are one. At every shared boundary the pair adds to the total, so each table is the other subtracted from n. And the operation reverses: differencing a cumulative column gives the frequencies back, which is the only reason a distribution handed over in purely cumulative form can be used at all.

What you should be able to do

  • Build a cumulative frequency column from a frequency column, upward and downward
  • State what a given cumulative entry counts, in words, for a named boundary
  • Show that the two cumulative forms sum to n at every shared boundary, and use that to convert one into the other
  • Recover a class-frequency column by differencing a cumulative column
  • Construct the class intervals implied by a table given only in cumulative form
  • Identify which limits — upper or lower — each cumulative form is indexed by
  • Explain why the two cumulative tables cover different boundary sets

Words to know

TermDefinition in one lineFirst introduced
cumulative frequencythe running total of frequencies up to a stated boundaryprinted in §13.1, p. 171, where the chapter announces what it will cover; the column itself is built and named in §13.4, p. 189
Cumulative Frequency Tablethe two columns pairing each boundary with its running totalprinted in this chapter, §13.4, p. 190
medianthe value of the middle observation once the data is orderedprinted in this chapter, §13.1, p. 171
upper limitthe top end of a class, which indexes the upward accumulationprinted in this chapter, §13.4, p. 191
lower limitthe bottom end of a class, which indexes the downward oneprinted in this chapter, §13.4, p. 192
ogivethe graph of a cumulative frequency distributionprinted in this chapter, §13.1, p. 171, and again in the closing note, p. 201
cumulative frequency curvethe chapter's other name for the same graphprinted in this chapter, §13.1, p. 171
frequencythe count of observations in one classprinted in this chapter, §13.2, p. 171
shared boundarya value appearing in both cumulative tables — every class edge except the two extremesan added term
differencingsubtracting consecutive cumulative entries to recover frequenciesan added term for the reverse operation

Where people slip up

  • "Cumulative frequency is just another column to fill in." It changes what the table can answer. Frequencies answer "how many here?"; running totals answer "how many so far?", and only the second lets you locate a position.
  • "The two cumulative tables are two different distributions." They are one distribution counted from opposite ends. Either can be produced from the other by subtracting from n, without going back to the frequencies at all.
  • "Less than 20 means the class 10–20." It means every class below 20 — here 0–10 and 10–20 together, so 8 and not 3. This is the single commonest error in building the column.
  • "The last cumulative entry can be anything." It must equal n. If it does not, the column is wrong, and that is a free self-check on every question.
  • "Both cumulative tables are indexed by the same numbers." The upward one is indexed by upper limits, the downward one by lower limits. They differ by one class width all the way along.
  • "A table given cumulatively is ready to use." It is not — the mode and median formulas need class frequencies, so the column must be differenced first. Example 7 and Exercise 13.3 Q3 both open with that step.
  • "An open first class is a misprint." Example 7's lowest class genuinely has no printed lower limit. Sometimes context supplies one, as the age-18 floor does in Exercise 13.3 Q3, and sometimes nothing does.
Transcript1,920 words

Here is a frequency table. Fifty-three scripts, marks out of a hundred, sorted into classes of ten. Five scored under ten. Three between ten and twenty. Four between twenty and thirty. And so on, up to eight in the top class. Ask it how many scored between forty and fifty and it answers instantly. Three. Now ask it a different question. Which script is the middle one? Fifty-three scripts, so the middle one is the twenty-seventh. Which class is the twenty-seventh script in?

The table cannot tell you. Every row says how many are in one class, and none of them says how many came before it. That is the difference between a question about quantity and a question about position, and a frequency table answers only the first. So how would you find the twenty-seventh by hand? You would line the scripts up in order and count along. Five, and the first class is done. Eight, the second. Twelve, the third.

Notice what you are doing: at every boundary you carry a total forward — not the count in that class, but everything so far. That is a running total, and it is the whole idea: walk along it and stop when it passes the position you want. It is called the cumulative frequency — a long name for the total up to a stated boundary. It is not just another column to fill in. It answers a question the frequency column cannot.

Frequencies say how many are here. Running totals say how many so far — and only the second can locate a position. Build it on the fifty-three scripts, one row at a time. Five in the first class, so at ten the total is five. Add three, and at twenty it is eight. Add four, and at thirty it is twelve. Then fifteen, eighteen, twenty-two. Then the classes fill up: seven takes it to twenty-nine, nine to thirty-eight, seven more to forty-five.

And the last eight take it to fifty-three. That final entry is worth more than it looks. It has to be fifty-three, because by the top of the table you have counted everybody. So the last entry always equals the number of observations. If it does not, the column is wrong, and you have just checked your own arithmetic for free. And the opening question now has an answer: the twenty-seventh script falls between twenty-two and twenty-nine, so it is in the class sixty to seventy.

One entry deserves a closer look, because this is where the commonest mistake lives. The entry at twenty is eight. What does eight count? It is tempting to say the class ending at twenty, which holds three. That is wrong, and it quietly ruins the whole column. Eight counts every script that scored less than twenty. That is two classes, not one — the five under ten and the three between ten and twenty.

Five and three is eight. So a boundary gathers everything beneath it. Not the nearest class. Everything. Say it out loud and the error disappears: eight scripts scored under twenty. Every entry in the column reads the same way, and reading them properly is most of the skill. Now do the same thing from the other end. Instead of asking how many are below a boundary, ask how many are at it or above it.

At nought, everybody, because every script scored at least nothing. Fifty-three. At ten, everybody except the five in the bottom class. Forty-eight. At twenty, drop the next three. Forty-five. Then forty-one, thirty-eight, thirty-five, thirty-one, twenty-four. At eighty, only the last two classes are left. Fifteen. And at ninety, only the top class. Eight. This is a second cumulative table, accumulated downward — built the way the first one was, by counting from a different end.

Nothing here was subtracted from anything. The two columns are separate counts of the same fifty-three scripts. Which matters, because of what happens when you set them side by side. Take one boundary and read both entries. At twenty, the upward table says eight and the downward table says forty-five. Eight and forty-five is fifty-three. That is not a coincidence, and the reason is almost too simple to notice. Take any script. Either it scored under twenty, or it scored twenty or more. It cannot do both, and it cannot do neither.

So the boundary splits the fifty-three into two groups with nothing left over, and two groups holding everybody between them must add to everybody. That is the whole proof. The two cumulative tables are one distribution counted from opposite ends, and at every shared boundary they add to the total. And it depends on the split being clean. Let two classes overlap, so a script could be counted on both sides, and the totals stop adding up — the identity is a fact about a partition, not a rule of arithmetic.

One boundary is a demonstration. Do all of them. At ten, five and forty-eight. Fifty-three. At twenty, eight and forty-five. Fifty-three. At thirty, twelve and forty-one. At forty, fifteen and thirty-eight. At fifty, eighteen and thirty-five. At sixty, twenty-two and thirty-one. At seventy, twenty-nine and twenty-four. At eighty, thirty-eight and fifteen. At ninety, forty-five and eight. Nine boundaries. Nine pairs. Nine totals of fifty-three, and not one exception. Watch the two numbers as the sweep moves right. One climbs, the other falls, and they trade exactly.

Which has a practical consequence. Given one of these columns you never have to go back to the frequencies for the other: subtract each entry from the total. One thing about the pair catches people out. The two tables are not indexed by the same numbers. The upward one is indexed by upper limits — ten, twenty, thirty, up to a hundred. Ten of them. The downward one is indexed by lower limits — nought, ten, twenty, up to ninety. Also ten.

Ten rows each, but they overlap in only nine values, because the two lists are offset by exactly one class width. Draw all eleven class edges on one line, the upward index points above it and the downward ones below, and the offset is there to see. The two ends that do not match carry no information anyway: everybody scored under a hundred, and everybody scored at least nought, so both read fifty-three.

So the nine shared boundaries are the ones that say anything, which is why the sweep had nine pairs and not ten. Everything so far has gone one way: frequencies in, running totals out. Now run it backwards, because you will need to. If each entry is everything up to that boundary, the difference between consecutive entries is exactly the class between them. Twelve minus eight is four, and four is the class twenty to thirty. Fifteen minus twelve is three. Twenty-two minus eighteen is four.

The first entry needs no subtraction, because there is nothing beneath it. It is already the first class. Do that all the way down and the frequency column comes back whole: five, three, four, three, three, four, seven, nine, seven, eight. And there is the free check again: add the recovered frequencies, and if they do not total the last cumulative entry the differencing went wrong. Subtracting to reverse an addition is no surprise. What matters is remembering you can, because sometimes the running totals are all you are given.

Here is exactly that. Fifty-one girls, measured by height, and the data arrives already accumulated. Under a hundred and forty centimetres, four. Under a hundred and forty-five, eleven. Then twenty-nine, forty, forty-six, and fifty-one at a hundred and sixty-five. This table is not ready to use: almost anything you compute from it needs class frequencies, and there are none. So difference it. Four stays four. Eleven minus four is seven. Twenty-nine minus eleven is eighteen. Forty minus twenty-nine is eleven. Forty-six minus forty is six. Fifty-one minus forty-six is five.

Four, seven, eighteen, eleven, six, five. Add them: fifty-one. The check passes. And the classes come straight off the boundaries: a hundred and forty to a hundred and forty-five, then to a hundred and fifty, and so on. With one oddity at the bottom. The lowest class is written only as under a hundred and forty: no stated lower limit, open below, and not a misprint. Four girls were shorter than that, and nothing says how much shorter.

Sometimes, though, the missing limit is supplied by the question rather than by the table. A hundred policy holders, by age, given the same way. Under twenty, two. Under twenty-five, six. Then twenty-four, forty-five, seventy-eight, eighty-nine, ninety-two, ninety-eight, and a hundred at sixty. Difference it and you get two, four, eighteen, twenty-one, thirty-three, eleven, three, six and two. Totalling one hundred, so the arithmetic is sound. Now the trap. What is the lowest class?

Reading the table alone, you would say everything below twenty. But the wording says policies are only issued from age eighteen. So the lowest class is not open below at all. It is eighteen to twenty, and those two policy holders sit in a two-year window, not an open one. That matters: every other class here is five years wide and this one is two, so anything assuming equal widths has a problem.

The lesson costs nothing: read the sentence above the table before you decide what the table means. Now the thing a running total does that is easy to miss. Here is a simpler set: a hundred students, marks out of fifty, listed as individual values rather than classes. Twenty, twenty-five, twenty-eight, twenty-nine, thirty-three, thirty-eight, forty-two, forty-three — with six, twenty, twenty-four, twenty-eight, fifteen, four, two and one students scoring them.

The running totals are six, twenty-six, fifty, seventy-eight, ninety-three, ninety-seven, ninety-nine and a hundred. A hundred students, so the middle is between the fiftieth and the fifty-first. The fiftieth student scored twenty-eight and the fifty-first scored twenty-nine. Where does that come from? Read it off the running totals. By the mark twenty-five the total is twenty-six, so positions one to twenty-six are settled. The twenty-four students on twenty-eight come next, occupying positions twenty-seven through fifty — the total lands exactly on fifty.

So the next block begins at fifty-one: the twenty-eight students who scored twenty-nine, running to seventy-eight. So a running total does not merely count. It tells you which positions belong to which value — and that is what makes a middle value findable at all. One last thing: these two columns are asking to be drawn. Plot the upward table: boundary along the bottom, running total up the side. Ten and five, twenty and eight, thirty and twelve, and so on to a hundred and fifty-three.

Join them: a curve that starts low and climbs to the total. It can never fall: a running total can only stay the same or grow, and it stays the same exactly where a class is empty. Now plot the downward table on the same axes. It starts at fifty-three and falls to eight, and it is the first curve turned upside down. The two cross somewhere in the middle, and where they cross, the number below equals the number above.

The name for either curve is an ogive. But the arithmetic is the part that matters. A frequency table counts. A cumulative one counts so far. One column becomes the other by subtracting from the total, and either becomes the frequencies again by differencing. Two columns, two directions, one distribution.

Where this fits

Taken from the notes each video was made from, not from the reading order — these are the ideas this one rests on and the ones that later rest on it.

Builds on

Comes up again in

Either side of this one

The book

Open in a new tab