PrepShorts · Study sheet · Class 9 Mathematics · Chapter 7, The Mathematics of Maybe: Introduction to Probability
Chapter 7 · The Mathematics of Maybe: Introduction to Probability
Estimating from statistical data, and scaling the estimate up
This video could not be loaded. Reload the page to try again.
Sign in with Google11 min.
Keep your place in this chapter — sign in, it’s free.Sign in
Turning a survey into a claim about a whole school is one multiplication. Everything difficult about it is in the assumption underneath.
The idea
Turning a sample into a statement about a whole population is arithmetically trivial and epistemically expensive: it is one multiplication, resting on the assumption that the part you looked at resembles the whole you did not. The chapter builds its example so that the assumption is impossible to miss — fifty students from one class standing in for fifteen hundred across a school — and then names its own remedy in a way that hides two quite different repairs inside one sentence. Asking more students makes the estimate steadier; asking students from other classes and grades makes it about the school at all. A larger sample drawn from the same class fixes nothing, and that is the distinction this topic exists to draw.
What you should be able to do
- Compute a probability estimate from survey counts as a relative frequency
- Explain why a randomly chosen respondent is what makes that relative frequency a probability
- Scale a sample proportion up to a population and state the estimate as a count
- Show that the scaled estimates across all categories must total the population, and use that as a check
- Distinguish increasing a sample's size from making it representative, and say what each fixes
- Use the words population, sample and sampling correctly for a described survey
- Explain why scaling with the exact fraction is better than scaling with a rounded decimal, with a worked case
- Read probabilities directly off a table of a thousand recorded cases
- State what a scaled estimate is not — a count, a guarantee, or a measurement of the population
Words to know
| Term | Definition in one line | First introduced |
|---|---|---|
| statistical data | records of what has already happened, used as evidence for a probability | printed in the heading of §7.2.3 (p. 162) and in §7.2's first route (p. 159) |
| population | the whole group the estimate is meant to be about | printed in bold in §7.2.3 (p. 162) |
| sample | the part of the population you actually collected data from | printed in bold in §7.2.3 (p. 162) |
| sampling | the practice of choosing which part of the population to look at | printed in bold at the end of §7.2.3 (p. 163) |
| representative | of a sample that resembles the population in the ways that matter | printed in §7.2.3 (p. 163) and again in the LEARN MORE ABOUT SAMPLING box (p. 163) |
| biased | of a sample that systematically over- or under-represents part of the population | printed in the LEARN MORE ABOUT SAMPLING box (p. 163) |
| sample size | how many members of the population were surveyed | printed in the LEARN MORE ABOUT SAMPLING box (p. 163); defined for a sample space in §7.3.1 (p. 166) |
| relative frequency | a count divided by the total number collected | printed in bold in §7.2.1 (p. 160), and the quantity §7.2.3 computes |
| estimate | a value worked out from partial evidence rather than measured | printed in §7.2.3 (p. 162) |
| forecasting | predicting future quantities, given as one of this method's uses | printed in §7.2.3 (p. 162) |
| scaling up | multiplying a sample proportion by the population size | an added phrasing; the chapter performs the step and does not name it |
| category total check | adding the scaled estimates to confirm they account for the whole population | an added term; the chapter never performs this check |
Where people slip up
- "The sample proportion is the population proportion." It is an estimate of it. The chapter's word for 600 mangoes is approximately, and the word is carrying the whole argument.
- "A bigger sample fixes a biased one." It does not. Ask the same class two hundred times over and you learn about that class with great precision. Size and representativeness are separate repairs, and the chapter's own box names them separately.
- "600 is how many mangoes the school will want." It is what the estimate suggests. A real school of 1500 will not divide 40 : 30 : 20 : 10.
- "Probability from data is a different kind of probability." It is a relative frequency, exactly as in §7.2.1, with respondents in place of trials.
- "Round the decimal, then multiply." 0.23 × 600 = 138 where 7/30 × 600 = 140. Cancel first.
- "Just survey everyone." The chapter's own reason for sampling is that collecting from the whole population is usually impractical. That is a fact about cost and access, not laziness.
- "Anonymous means random." The fruit survey is anonymous, which protects the answers. What makes the estimate a probability is that the student is picked at random, and that is a different property.
Ask your teacher a person
Your teacher reads this and writes back, usually within a day. For an instant answer, use Ask the video in the sidebar.
Your class sees the question and the answer. Only your teacher sees that it was you.
No questions on this topic yet.
Worked answers to this chapter’s exercises · this video explains Exercise Set 7.2 Q1, Exercise Set 7.2 Q2, End-of-Chapter Exercises Q7
Transcript1,444 words
Not every probability comes from an experiment you run. Most of the ones that matter come from records of what has already happened. A shop deciding what to stock, a company forecasting sales, an insurer setting a price, a researcher reading a survey: all reasoning from data somebody already collected. What is new is the last step, and it is one multiplication. You take a proportion measured on the people you asked and apply it to a much larger group you did not ask.
That step is trivial to perform and expensive to justify, and this video is about the difference. Here is the case to work with. Fifty students in one class are asked, anonymously, for their favourite fruit. Twenty say mango, fifteen say apples, ten say bananas, five say grapes. First thing to check, before anything else: those four counts add to fifty. Nobody was left out and nobody counted twice, and that matters more than it looks.
Now the probability that a student picked from this class prefers mango. Twenty of the fifty, which is two fifths, or nought point four, or forty per cent. And the four proportions add to exactly one, because every student gave exactly one answer. Pause on the word probability, because it has been smuggled in. A count over a total is just a proportion of the class. It becomes a probability only when you say how the student will be chosen.
Choose at random and every one of the fifty carries the same weight, one fiftieth. Add up the weights of the mango answers and you get two fifths back. That is why the proportion is the probability: the picking rule put equal weight everywhere. Now pick the student sitting nearest the fruit bowl instead. That rule lands on the same person every time, so the answer is not two fifths.
It is nought or one, depending who that is, and the same counts support no claim at all. Anonymous protects the answers; random is what makes the number mean something. Now the school, which has fifteen hundred students. If two fifths of them prefer mango, that is two fifths of fifteen hundred. Cancel first: fifteen hundred divided by fifty is thirty, times twenty is six hundred. Six hundred students, from a survey of fifty.
That is the whole of the arithmetic. It assumes the fifty resemble the fifteen hundred in the one respect being measured. Not that they are the same people, only that mango is as popular among those you did not ask. Everything hangs on that sentence, and nothing in the survey can check it. Do the other three the same way. Apples, fifteen fiftieths, gives four hundred and fifty. Bananas, ten fiftieths, gives three hundred.
Grapes, five fiftieths, gives one hundred and fifty. Add them: six hundred, four hundred and fifty, three hundred, one hundred and fifty. Fifteen hundred exactly, which is the whole school. That is a real check, the only one available on a scale-up, but it works only if you divide by the number you actually asked. Lose one respondent — twenty, fifteen, ten and four, all over fifty — and the estimates come to one thousand four hundred and seventy.
The school does not add up, and that tells you a count went missing. So what is the six hundred? It is not a measurement of the school; nobody measured the school. It is not a guarantee, and it is not how many mangoes will be wanted. It is one number, two fifths, applied to fifteen hundred people, and only as good as the two fifths. Change one student's answer, so twenty-one say mango instead of twenty.
The estimate moves from six hundred to six hundred and thirty. One person in that classroom carries thirty students of the school, because fifteen hundred divided by fifty is thirty. That is not an error in the method; it is the method. The obvious fix is to ask more people, and the standard advice is to take a larger sample and make it more representative. That is one sentence holding two different repairs, and only one of them fixes this.
Watch what happens on a school small enough to check every possibility. Twelve students in three classes of four. In the surveyed class three of the four like mango; in each of the others, only one. So the school is five out of twelve; the surveyed class is three out of four. Now take every possible sample from that class. Average over every one of them, at every size, and it is three quarters every time.
Exactly three quarters, including when you ask the whole class. The gap to the school's five twelfths is a third, and asking more of the same class never closes it at all. Now draw the sample from the whole school instead. Run the same exhaustive count: every possible sample of every possible size. The average answer is five twelfths, the school's own figure, at every size from one to twelve.
And now something happens that did not happen before. Measure how far a typical sample strays from it, and that falls at every step as the sample grows. By six of the twelve it has more than halved; by twelve it is nothing. A bigger sample makes the answer steadier. A fairer sample makes it an answer about the right group. A bigger sample of the wrong group is a more precise measurement of the wrong group, and precision was not the problem.
Three words are worth keeping apart, because they get used interchangeably. The population is the whole group the estimate is about: the fifteen hundred. The sample is the part you collected data from: the fifty. Sampling is the practice of choosing which part, and it is where all the difficulty lives. The sample sits inside the population, and it is one thirtieth of it. So the estimate speaks for one thousand four hundred and fifty people nobody asked.
That is not an objection to sampling; asking everybody is usually impossible. It is a reminder of how much weight the sampling step is carrying. One more thing, and it is pure arithmetic. A bag of sweets is mixed and thirty taken at random: ten red, eight green, seven yellow, five blue. The bag holds six hundred. How many are yellow? Seven thirtieths of six hundred. Cancel first: six hundred over thirty is twenty, times seven is one hundred and forty, exactly.
Now do it the other way. Seven divided by thirty rounds to nought point two three, and that times six hundred is one hundred and thirty-eight. Two sweets have vanished into the rounding. It goes the other way too: green is exactly one hundred and sixty, but rounding gives one hundred and sixty-two. And the damage is not rare. Count every way thirty answers could split four ways, round each share to two places, and more than a quarter no longer fill the bag.
Cancel first, divide last. Notice something about that sweet bag that the fruit survey did not have. The thirty sweets were drawn out of the very bag being estimated. The sample is part of the thing being described, and the total is known. Compare a different survey: forty students, a random sample of a whole school of eight hundred. Fourteen say science club, eleven arts, nine sports, six debate. Arts is eleven fortieths, nought point two seven five, and sports scales to one hundred and eighty.
That scale-up stands on firmer ground, and the difference is one phrase: the forty came from the school, not from one class. Same arithmetic, better foundation. The multiplication never tells you which situation you are in. Finally, the case where none of this applies. A tyre company records how far a thousand tyres ran before replacement. Under four thousand kilometres, twenty of them. Four to nine thousand, two hundred and ten.
Nine to fourteen thousand, three hundred and twenty-five. Over fourteen thousand, four hundred and forty-five. Those add to a thousand, so the records are the whole population and no scaling happens at all. The relative frequencies are the answers: nought point nought two, and nought point four four five. For four to fourteen thousand you join two columns, two hundred and ten plus three hundred and twenty-five, which is nought point five three five.
And the three answers add to one. Read those boundaries as continuous ranges, not whole kilometres, or four thousand exactly belongs to no column. When your data already covers everybody you are not estimating but counting, and the hard part of all this does not arise.
Where this fits
Taken from the notes each video was made from, not from the reading order — these are the ideas this one rests on and the ones that later rest on it.
Builds on
- Experimental probability: relative frequency over many trialsClass 9 · Ch 7, The Mathematics of Maybe: Introduction to Probability
- What "rational" means, and why the denominator cannot be zeroClass 9 · Ch 3, The World of Numbers
Either side of this one
- Theoretical probability: counting favourable outcomes when all are equally likelyClass 9 · Ch 7, The Mathematics of Maybe: Introduction to Probability
- Fair, unbiased, and memoryless: the gambler's fallacyClass 9 · Ch 7, The Mathematics of Maybe: Introduction to Probability