PrepShorts · Study sheet · Class 10 Mathematics · Chapter 13, Statistics
This video could not be loaded. Reload the page to try again.
Sign in with Google15 min.
Keep your place in this chapter — sign in, it’s free.Sign in
Grouping throws the actual values away, so the middle observation still exists but the table cannot name it. The median formula is not five symbols to memorise - it is one sentence: find the class the middle falls in, assume its observations are evenly spread, go as far into that class as the middle stands through its count, and measure that in the data's own units.
The idea
The median formula looks like five symbols to memorise and is really one sentence of straight-line reasoning. Running totals tell you which class the middle observation falls in, but not where inside it — so we assume the class's observations are spread evenly across its width. Then the middle one stands as far along the class as its position stands through the class's count, and multiplying that fraction by the class width converts it into a number on the measurement scale. Each symbol in the formula is one clause of that sentence, which is why the formula also reads backwards: Example 8 fixes the median and solves for a missing frequency instead.
What you should be able to do
- Compute n ÷ 2 and identify the median class from a cumulative frequency column
- Explain what each of l, n, cf, f and h refers to in a specific table
- Derive the median formula as linear interpolation rather than quoting it
- Compute the median of a grouped distribution and state its meaning in context
- Compute a median from a table supplied only in cumulative form
- Use the formula in reverse to find one or two missing frequencies given the median and the total
- Apply the continuity correction before computing a median, and say what difference it makes
Words to know
| Term | Definition in one line | First introduced |
|---|---|---|
| median | the value of the middle observation once the data is ordered | printed in this chapter, §13.1, p. 171 |
| median class | the class in which the middle observation falls | printed in this chapter, §13.4, p. 193 |
| cumulative frequency | the running total of frequencies up to a boundary | printed in §13.1, p. 171, where the chapter announces what it will cover; the column itself is built and named in §13.4, p. 189 |
| lower limit | the bottom end of the median class, from which the interpolation starts | printed in this chapter, §13.4, p. 193 |
| class size | the width of the median class, written h | printed in this chapter, §13.2, p. 176 |
| continuous | describing classes whose limits meet, required before this formula applies | printed in this chapter, in the closing note, p. 201 |
| interpolation | estimating a value between two known ones by assuming a straight line between them | an added term; the chapter performs the interpolation without naming it |
| even spreading | the assumption that a class's observations sit at equal intervals across it | an added phrasing for the assumption the formula encodes |
Where people slip up
- "The median is the class mark of the median class." In the chapter's worked case the class mark of 60–70 is 65 and the median is 66.43. They agree only when the middle position happens to sit centrally in the class.
- "cf is the cumulative frequency of the median class." It is the running total for the class before it. Using the median class's own value is the most frequent wrong answer on this topic, and it is self-diagnosing every single time: the median class is by definition the first whose running total has reached n ÷ 2, so its own total is never less than n ÷ 2 and the substitution always leaves a numerator of zero or below. A negative or vanishing numerator means this mistake and no other.
- "Use (n + 1) ÷ 2 as you did for ungrouped data." Grouped questions use n ÷ 2. On the chapter's own data the two give 66.43 and 67.14 — a visible gap.
- "Pick whichever class has a running total closest to n ÷ 2." Pick the first one that has reached it. A nearer value on the low side has not yet reached the middle observation, so the middle cannot be in it.
- "The median needs the frequencies, so a cumulative table is unusable." Difference it first. Example 7 and Exercise 13.3 Q3 both begin that way.
- "The continuity correction is bookkeeping that does not change anything." On Exercise 13.3 Q4 it moves the median from 147 mm to 146.75 mm. It changes the lower limit you start from and the width you multiply by.
- "Even spreading inside the class is a fact about the data." It is an assumption, and it is the same species of assumption as the class mark for the mean — a stand-in adopted because the real positions were discarded.
- "If the median formula gives a value outside the median class, that's fine." It cannot, if the working is right. The fraction is between 0 and 1, so the answer always lands between l and l + h — a free sanity check.
Ask your teacher a person
Your teacher reads this and writes back, usually within a day. For an instant answer, use Ask the video in the sidebar.
Your class sees the question and the answer. Only your teacher sees that it was you.
No questions on this topic yet.
Worked answers: Exercise 13.1 · Exercise 13.2 · Exercise 13.3 · this video explains Exercise 13.3 Q2, Exercise 13.3 Q4, Exercise 13.3 Q5, Exercise 13.3 Q6, Exercise 13.3 Q7
Transcript2,091 words
The median is the middle observation. With a list of values in front of you that is a rule you can follow: put them in order, count to the middle, read it off. Grouping throws the values away. Fifty-three scripts, marks out of a hundred, sorted into classes of ten. Nobody's actual mark survives that. So the middle observation is still there - it exists, somebody wrote it - but the table cannot tell you what it is.
What the table can still do is narrow it down. It can tell you which class the middle observation fell into. And then it can estimate where inside that class it sits. That estimate is what the median formula computes, and the rest of this video is the reasoning behind it. It looks like five symbols to memorise. It is really one sentence. First, which class. Half of fifty-three is twenty-six point five, so the middle of the data sits at position twenty-six and a half.
Now walk down the running totals: five, eight, twelve, fifteen, eighteen, twenty-two, twenty-nine. Twenty-two is not enough. Twenty-nine is. So the middle observation is somewhere in the class sixty to seventy, and that class has a name: the median class. The rule is the first class whose running total has reached half of n. First - not nearest. Nearest is a tempting rule and it is wrong. A running total on the low side has not yet reached the middle observation, so the middle cannot be in that class, however close the number looks.
Reaching it is what matters. One thing changes quietly here, and it is worth a moment. For a plain list of values with an odd count, the middle sits at position n plus one, over two. Fifty-three values, and the middle one is the twenty-seventh. For grouped data the rule is half of n, flat. Twenty-six and a half. Both of those fall inside sixty to seventy, so the median class is the same either way.
But the answers are not the same. Half of n gives sixty-six point four three. The other rule gives sixty-seven point one four. Five sevenths of a mark apart. Small, and still a different answer. For grouped data, use half of n. The reason is that grouping has turned fifty-three separate things into a quantity spread continuously along a scale, and half of a quantity is half of it. So the middle observation is in the class sixty to seventy. Now where inside it?
Here is the honest answer: the table does not know. Seven scripts scored somewhere between sixty and seventy. The table records that there were seven. It records nothing about where. They could all be bunched at sixty-one. They could all be at sixty-nine. They could be scattered evenly. Cover the neighbouring classes and stare at the number seven, and there is no way to get a position out of it.
So we cannot deduce the answer. We have to assume something. And it is worth saying plainly that what follows is an assumption, not a fact about the data. It is the same species of assumption as letting a class mark stand for a whole class when you compute a mean. The assumption is the simplest one available: even spreading. Take the seven scripts in the class and place them at equal intervals across its width.
That is all of it. No cleverness. And it buys everything, because now a position turns into a place. The first of the seven stands a seventh of the way through the class. The fourth stands four sevenths through. The seventh stands at the far end. Position becomes distance, and distance is a number on the marks scale. If the real data were not evenly spread, this estimate would be off - and there is no way to tell from the table whether it was.
That is the price of grouping, and it was paid the moment the values were thrown away. Now the counting. The middle sits at position twenty-six and a half of the whole data set. Twenty-two observations are already accounted for below sixty. So inside the class, the middle is the four-and-a-half-th of the seven. Twenty-six and a half minus twenty-two is four and a half. That is the count still owed.
Seven is the count available. Four and a half over seven is nine fourteenths, which is about nought point six four three. So the middle stands roughly sixty-four per cent of the way through the class. Still owed, over available. That fraction is the heart of the whole thing, and it is always between nought and one, because the class is the one the middle observation is in. A fraction is not an answer yet. Sixty-four per cent of the way through what?
Through a class ten marks wide. So multiply. Nine fourteenths of ten marks is forty-five sevenths, which is about six point four three marks. That is a distance, in the same units as the data. Start at the lower limit of the class, sixty, and walk six point four three marks up it. Sixty-six point four three. And watch how the units behave. A fraction has no units. A class width is in marks. Their product is in marks, and it is added to a lower limit in marks.
If the data were heights in centimetres, the width would be in centimetres and so would the answer. The machine does not care what is being measured. Now write down what we just did, and the formula appears on its own. Start at l, the lower limit of the median class. Sixty. Add h, the class width, times a fraction. Ten. The fraction is n over two, minus cf, all over f.
n over two is the middle position. Twenty-six and a half. cf is the running total of everything below the median class. Twenty-two. And cf is the class before it, never the median class's own total. That one matters, and we will come back to it. f is the count inside the median class. Seven. Median equals l, plus h times, n over two minus cf, over f. Five symbols, and every one of them is a clause of the sentence we just spoke.
Start at the bottom of the class. Go as far through it as the middle stands through the count. Measure that in the data's own units. There is nothing to memorise if you can rebuild it. Run it on the fifty-three scripts. l is sixty. h is ten. cf is twenty-two. f is seven. And n over two is twenty-six and a half. Sixty, plus ten times four and a half over seven.
Ten times four and a half is forty-five. Forty-five over seven is six point four three. Sixty-six point four three - or sixty-six point four, to one place. Now say what it means, because a number with no sentence attached is not an answer. Roughly half the class scored below sixty-six point four, and roughly half above. Not exactly half. That is what makes it an estimate. But that is what the number is claiming.
And notice that the class mark of sixty to seventy is sixty-five. The median is not the class mark. The two agree only when the middle happens to sit dead centre in the class, and here it does not. The same machine, on a table that arrives the wrong way round. Fifty-one girls, measured by height, and the data comes already accumulated. Under a hundred and forty centimetres, four. Under a hundred and forty-five, eleven. Then twenty-nine, forty, forty-six, and fifty-one.
There are no class frequencies anywhere in that, and the formula needs one. So the first move is to difference it: four, seven, eighteen, eleven, six, five. They total fifty-one, which is the check that the differencing was done right. Half of fifty-one is twenty-five and a half. The running total passes it at twenty-nine. So the median class is a hundred and forty-five to a hundred and fifty. l is a hundred and forty-five, cf is eleven, f is eighteen, and h is five.
A hundred and forty-five, plus five times fourteen and a half over eighteen. Seventy-two and a half over eighteen is four point zero three. So a hundred and forty-nine point zero three centimetres. Half the girls are shorter than that and half are taller. Same machine, different units. Now run the formula backwards, which is the strongest evidence there is that it is reasoning and not ritual. Ten classes, nought to a hundred, up to nine hundred to a thousand.
The frequencies are two, five, then an unknown x, then twelve, seventeen, twenty, then another unknown y, then nine, seven, four. You are told the total is a hundred, and the median is five hundred and twenty-five. The total gives one relation at once. Everything known adds to seventy-six, so x plus y is twenty-four. Five hundred and twenty-five sits in the class five hundred to six hundred, so that is the median class.
l is five hundred, h is a hundred, f is twenty, and cf - everything below five hundred - is thirty-six plus x. So five hundred and twenty-five equals five hundred, plus a hundred times, fifty minus thirty-six minus x, over twenty. Twenty-five equals five times fourteen minus x. So fourteen minus x is five, and x is nine. And then y is fifteen. Now close the loop, which is the step almost everyone skips.
With x equal to nine the running totals are two, seven, sixteen, twenty-eight, forty-five, sixty-five. Fifty falls between forty-five and sixty-five, so five hundred to six hundred really is the median class. The assumption we started from holds. There is one thing the formula demands, and it is easy to miss. It needs the classes to meet. l plus h has to be the next class's lower limit, or the arithmetic is describing a scale with holes in it.
Whole-number data often comes with gaps between the classes. A hundred and eighteen to a hundred and twenty-six, then a hundred and twenty-seven to a hundred and thirty-five, and so on. That is not one continuous scale. There is a gap between a hundred and twenty-six and a hundred and twenty-seven, and the formula has nothing to say about it. The repair is to move each boundary to the middle of the gap.
A hundred and seventeen point five to a hundred and twenty-six point five, then to a hundred and thirty-five point five, and on up. Every class is now nine wide instead of eight, and they meet. Does it matter? On forty leaves measured that way, the corrected median is a hundred and forty-six point seven five millimetres. Skip the correction, treat the median class as a hundred and forty-five to a hundred and fifty-three, eight wide, and you get a hundred and forty-seven.
A quarter of a millimetre. Small - but it is not bookkeeping. The correction changes the lower limit you start from and the width you multiply by, and both of those are in the answer. Three checks come free with this method, and between them they catch most of what goes wrong. First: the answer has to land inside the median class. The fraction is between nought and one, so l plus h times it can only land between l and l plus h. Sixty-six point four three is between sixty and seventy.
If your answer is outside the class, the working is wrong, and you know that before you check anything else. Second: cf is the running total of the class before the median class, and never its own. That mistake is self-diagnosing, which is a rare and lovely thing. The median class is by definition the first whose running total has reached half of n. So its own total is never below half of n.
Use it as cf and the numerator comes out nought or negative. Every single time. A negative numerator on this topic means that mistake and no other. Third: the last cumulative entry has to equal n. If it does not, the column is wrong before you start. And the whole method, in one sentence. Find the class the middle falls in, assume its observations are evenly spread, go as far into that class as the middle stands through its count, and measure it in the data's own units.
Five symbols. One sentence. Rebuildable from the picture every time.
Where this fits
Taken from the notes each video was made from, not from the reading order — these are the ideas this one rests on and the ones that later rest on it.
Builds on
- Running totals, and converting between the two cumulative tablesClass 10 · Ch 13, Statistics
- Grouping loses the raw values, so we stand the class mark in for themClass 10 · Ch 13, Statistics
Comes up again in
- Which of the three averages a given question actually wantsClass 10 · Ch 13, Statistics