PrepShorts · Study sheet · Class 10 Mathematics · Chapter 13, Statistics
Chapter 13 · Statistics
Shifting the origin: why guessing a centre cannot change the answer
This video could not be loaded. Reload the page to try again.
Sign in with Google14 min.
Keep your place in this chapter — sign in, it’s free.Sign in
Subtract any number you like from every class mark, average what is left, and the result falls short of the true mean by exactly that number. Never by more, never by less. So the centre you measure from is a genuinely free choice, and a bad guess costs you arithmetic — not a single decimal of accuracy.
The idea
Subtract any number you like from every class mark, average what is left, and the result falls short of the true mean by exactly that number — never by more, never by less. That is not a coincidence to be checked case by case; it follows in one line from a single fact, that averaging a constant gives the constant back. So the choice of centre is genuinely free, and the name "assumed mean" is misleading: nothing is being assumed about where the middle lies, and a poor guess costs extra arithmetic but not a decimal of accuracy. Activity 1 is the chapter handing the student six chances to break this and watching all six fail.
What you should be able to do
- Choose an assumed mean from a table and justify the choice on arithmetic grounds
- Build the deviation column d = x − a and the product column f·d for a given distribution
- Compute a grouped mean as a + Σf·d ÷ Σf
- Derive that the mean of the deviations equals the true mean less a, naming the property of sums used at each step
- Predict, before computing, that changing the assumed mean will not change the final answer, and verify it on a second choice
- Explain what a badly chosen assumed mean actually costs
- Recognise the direct method as the special case in which nothing is subtracted
Words to know
| Term | Definition in one line | First introduced |
|---|---|---|
| assumed mean | the number chosen to measure the class marks from, written a | printed in this chapter, §13.2, p. 174 |
| Assumed Mean Method | computing the grouped mean via deviations from a chosen a | printed in this chapter, §13.2, p. 176 |
| deviation | the signed gap between a class mark and the assumed mean | printed in this chapter, §13.2, p. 174 |
| mean of the deviations | the average of the d column, written with a bar over d | printed in this chapter, §13.2, p. 175 |
| class mark | the mid-point standing in for a whole class | printed in this chapter, §13.2, p. 173 |
| Direct Method | the same mean computed straight from the class marks | printed in this chapter, §13.2, p. 174 |
| shifting the origin | relabelling every value by measuring from a new zero | an added phrasing; the chapter performs the shift without naming it |
| free choice | a parameter the answer provably does not depend on | an added term |
Where people slip up
- "The assumed mean is a guess at the answer, so a good guess gives a better answer." It is not a guess at anything. Every choice returns the identical mean; the derivation leaves no room for the choice to matter.
- "a must be one of the class marks." It is convenient — one deviation becomes 0 and one product vanishes — but it is not required. A centre of 100 works on this data and lies above every class mark.
- "a must lie inside the range of the data." No. See the a = 100 case.
- "You add a back because it seems fair." You add it back because the algebra produces exactly a and nothing else.
- "If the deviations total zero the method has failed." A total of zero means the chosen centre happened to be the mean. That is the method succeeding perfectly, not failing — try a = 62 on this data.
- "Negative deviations should be dropped or made positive." The signs are the whole mechanism. Making them positive computes a different quantity entirely and destroys the cancellation.
- "Different a, different answer — I got 62 and my neighbour got 62.5." One of you made an arithmetic slip. Because the answer cannot depend on a, disagreement between two choices is a reliable error detector — a genuinely useful exam habit.
Ask your teacher a person
Your teacher reads this and writes back, usually within a day. For an instant answer, use Ask the video in the sidebar.
Your class sees the question and the answer. Only your teacher sees that it was you.
No questions on this topic yet.
Worked answers: Exercise 13.1 · Exercise 13.2 · Exercise 13.3
Transcript1,975 words
Here is a table of thirty students' marks, sorted into six classes. To find the mean by the direct method you multiply each class mark by its frequency, total the column, and divide by thirty. Nothing about that is hard. Look at the column it makes. Six products, most of them three figures wide, and two of them carrying a half. The complaint is not that the method is wrong. The complaint is the size of what your hand has to write.
So somebody had an idea. The class marks are large because they are being measured from nought, and nought is nowhere near this data. Measure them from somewhere closer, and they get small. The whole of this video is about whether you are allowed to do that. Put the six class marks on a line, fifteen apart, starting at seventeen and a half. Now plant a flag at one of them — say forty-seven and a half — and measure everything from the flag instead of from nought.
The third mark is at the flag, so it becomes zero. The two marks below it become minus fifteen and minus thirty. The three above become fifteen, thirty and forty-five. That is the deviation column, and it is what the method is named after: each value's signed gap from a chosen number. The chosen number is called the assumed mean, and it is written a. Hold on to that word, assumed. By the end of this you are going to want it back.
Look at what has happened to the numbers. The largest entry was ninety-two and a half. The largest entry now is forty-five. And the halves are gone, because every class mark ends in a half and so does the flag, so they cancel in the subtraction. And two of the six entries are negative. That is not damage. That is the whole mechanism. The marks below the flag pull the answer down and the marks above pull it up, and the signs are what let them cancel.
Make the deviations positive because negatives look untidy and you have not simplified anything. You have computed a different quantity, and we will see exactly what it comes to. Two negative, one nought, three positive. Keep the signs. Now do the same work as before, on the new column. Multiply each deviation by its frequency. Minus sixty. Minus forty-five. Then nought times seven, which is nothing at all — that whole row costs you no arithmetic.
Then ninety, one hundred and eighty, two hundred and seventy. Total the column, minding the signs: four hundred and thirty-five. Share that among the thirty students and you get fourteen and a half. Fourteen and a half is not the mean. It is the average of the deviations — the average distance from the flag, counted with sign. So add the flag back. Forty-seven and a half plus fourteen and a half.
Sixty-two. Which is exactly what the direct method gives, off a column of three-figure numbers. Now, that could be a coincidence. One table, one guess, one lucky agreement. So put it on trial. Move the flag. Take a as sixty-two and a half instead — one class mark further along. Every deviation changes, and so does every product. And the total changes sign: minus fifteen. Minus fifteen shared among thirty is minus a half.
And sixty-two and a half, minus a half, is sixty-two. The columns share not one entry with the columns before. The answer is the same number. So either that is a second coincidence, or the shortfall is doing something exact. It is doing something exact, and it takes four lines to see. Start with the thing we actually computed: the average of the deviation column. Sum of f times d, over sum of f.
Line two. Every d is just x minus a, so write that in. Line three. A sum of differences is the difference of the sums, so split it in two. Line four. Split the division across the subtraction, and you have two separate fractions. The first is sum of f x over sum of f. That is the mean — the thing we were trying to find in the first place.
The second is sum of f times a, over sum of f. And a is a constant: it does not vary from class to class, so it comes out of the sum. Leaving a times sum of f, over sum of f. Which is the only interesting step in the argument. A times sum of f, over sum of f. Sum of f is the number of observations. So you have taken one number, written it down once for every student in the table, added them all up, and divided by how many students there are.
You get a back. The average of a constant is that constant. Say it out loud and it sounds like nothing. It is the load-bearing line of the method. The sums cancel — not approximately, exactly, because they are the same number top and bottom. So the second fraction is a. Not roughly a. A. And the four lines have come out to this: the average of the deviations is the mean, minus a.
The column does not fall short by some amount that depends on how good your guess was. It falls short by the guess. Add a to both sides and the method is finished: the mean is a, plus the average of the deviations. So try every class mark in the table as the flag, and watch the columns. Six completely different columns, and six completely different totals. One thousand three hundred and thirty-five. Eight hundred and eighty-five. Four hundred and thirty-five. Minus fifteen. Minus four hundred and sixty-five. Minus nine hundred and fifteen.
Now look at that as a sequence rather than a list. Each one is four hundred and fifty below the one before it. Every single time. And four hundred and fifty is thirty times fifteen — thirty students, and fifteen is how far the flag moved each step. The shortfall does not drift with the guess. It tracks it exactly. That is the theorem, written as six integers. And all six answers are sixty-two.
If the choice really is free, then it is free, and the flag does not have to be sensible. Put it at one hundred. That is above every class mark in the table — there is no student anywhere near it. Now every deviation is negative, all the way down to minus seven and a half. The column totals minus one thousand one hundred and forty, which shared among thirty is minus thirty-eight.
One hundred minus thirty-eight is sixty-two. So the flag does not have to sit inside the data. Try minus forty. Sixty-two. Try a thousand. Sixty-two. Try one third — not a class mark, not a whole number, not near anything at all. Sixty-two. Twenty-one centres were tried on five tables — a hundred and five runs — and every one returned the answer that table already had. One more. Put the flag at sixty-two: at the answer itself.
Then the deviation column totals nought, which students read as the method failing. It is the opposite. You guessed the mean exactly, so the deviations balance, so there is nothing to add on, and the answer is the flag. Since nothing you do to the flag can break this, it is worth knowing what can. Four things. One: make the deviations positive. At forty-seven and a half that gives sixty-nine.
And notice where the trap hides. If your flag sits at or below every mark, nothing is negative, so there is nothing to make positive, and you get the right answer anyway. Across the hundred and five runs it agreed forty-seven times — and every one of those forty-seven was a flag with nothing below it. That is how a wrong method survives. Two: never add the flag back. That gives fourteen and a half — short by forty-seven and a half, which is the flag, which is exactly what the derivation said.
Three: divide by the number of classes rather than the number of students. Six instead of thirty gives one hundred and twenty, for a table whose largest mark is ninety-two and a half. Four: subtract the flag from the frequency column. The frequencies are the data, and nobody may relabel those. It returns a negative answer from a table with no negative number in it. So if accuracy is not at stake, what is?
Arithmetic. And you can count it: write out every entry of the product column and its total, and count the digit characters your hand actually forms. For the six choices of flag that comes to nineteen, eighteen, sixteen, fifteen, seventeen and eighteen. The cheapest is fifteen, at sixty-two and a half, the class mark nearest the answer. So the rule of thumb is a real effect, not folklore. The dearest is nineteen, at the mark furthest away.
The direct method — which, remember, is the flag at nought — costs twenty-two here. And the flag at one hundred costs twenty-three. So a bad enough guess costs more than not guessing at all, and still returns sixty-two. Every half unit from nought to a hundred was tried, and the cheapest of the two hundred and one is sixty-one and a half, at fourteen — neither a class mark nor the mean.
Which tells you what kind of quantity this is. Pen-work is worth minimising, and it has nothing to do with whether the answer is right. Here is a table where the saving is worth having. Thirty-five territories, class marks running from twenty to eighty in tens. The direct method's column runs up to three hundred and thirty and totals one thousand three hundred and ninety. Twenty-four digit characters. Take the flag at fifty and the deviations are minus thirty up to thirty, in tens.
The products total minus three hundred and sixty. Eighteen digit characters. Six characters cheaper, and every number in it small enough to hold in your head. Minus three hundred and sixty over thirty-five is minus ten point two eight, and fifty minus that is thirty-nine point seven one — which is what one thousand three hundred and ninety over thirty-five gives, to the digit. And that hands you something useful. If two people pick different flags and get different answers, one of them has slipped, because the method cannot produce that.
Thirty-six single-entry slips were made in one column here, and all thirty-six showed up as a different answer. Two flags is a free check on your own arithmetic. One choice left, and it is the interesting one. Put the flag at nought. Then d is x minus nought, which is x. The deviation column IS the class mark column. The product column is the f times x column, and it totals one thousand eight hundred and sixty.
Share that among thirty, and add the flag — which is nought, so add nothing. Sixty-two. That is the direct method. Not something like it — it is this method with a chosen to be nought, and it always was. So there are not two methods here. There is one method with a parameter, and the direct method is the case where nobody moved the origin. Which brings us back to the word. Assumed mean.
Nothing has been assumed. You have not guessed at the answer and you are not hoping to be close. The derivation left no room for your choice to matter: it produced a, exactly, and cancelled it. It is not an assumed mean. It is a chosen origin. And what a good choice buys you is smaller numbers to write down — nothing more, and nothing less.
Where this fits
Taken from the notes each video was made from, not from the reading order — these are the ideas this one rests on and the ones that later rest on it.
Builds on
- The direct method, and where it becomes unwieldyClass 10 · Ch 13, Statistics
- Grouping loses the raw values, so we stand the class mark in for themClass 10 · Ch 13, Statistics
Comes up again in
- Dividing through by the class width to shrink the arithmetic furtherClass 10 · Ch 13, Statistics