PrepShorts · Study sheet · Class 8 Mathematics · Chapter 5, Tales by Dots and Lines
Chapter 5 · Tales by Dots and Lines
Mean and median from a frequency table, by hand and in a spreadsheet
This video could not be loaded. Reload the page to try again.
Sign in with Google9 min.
Keep your place in this chapter — sign in, it’s free.Sign in
A frequency table is an ordinary list with the repetitions folded up. Every rule for computing with one falls out of unfolding it.
The idea
A frequency table is an ordinary data list with the repetitions folded up, and every rule for computing with it follows from unfolding it in your head. That is why averaging the distinct values is wrong — it gives a family size that one student reported the same weight as one that eleven students reported — and why the correct mean is a weighted total divided by the total frequency, not by the number of rows. The same folding is what makes the median cheap: running the frequencies up from the smallest value tells you which value sits in any given position, so you never write the list out. A spreadsheet then adds nothing mathematical at all. It only lets you name a block of cells instead of a list of numbers, which is what makes the same two operations survive a table of a hundred and thirty-two marks.
What you should be able to do
- Explain what a frequency table records, and reconstruct the underlying list from one
- Say why averaging the distinct values of a frequency table is wrong, and what quantity that computation actually gives
- Compute a mean from a frequency table as a weighted total over the total frequency
- Identify the correct divisor for such a mean, and say why it is not the number of rows
- Locate the median from a frequency table by accumulating frequencies, and state which positions a given value occupies
- Report a mean that does not come out exactly, and say what has been rounded
- Name a spreadsheet cell from its column letter and row number, and read a value out of a named cell
- Write a range as a start cell and an end cell, and say which values it collects
- Write a formula that totals a row and one that averages part of a row, and predict its result before the sheet computes it
- Compute the mean, median, smallest and largest value of a data set given as a frequency table
Words to know
| Term | Definition in one line | First introduced |
|---|---|---|
| frequency | how many times a particular value occurs in the data | printed as a column heading and used throughout Part II p.110 |
| frequency table | a two-column record pairing each value with how often it occurs | the chapter prints the table and the subsection heading "Mean and Median with Frequencies" (Part II p.110); this compound is the explanation's shorthand |
| mean | the total of all the values divided by how many values there are | printed throughout Part II §5.1 |
| median | the middle value of the sorted data, or the average of the two middle values | printed throughout Part II §5.1; located from frequencies on Part II p.110 |
| sorted | in order of size — the state the data has to be in before positions mean anything | printed on Part II p.110 in exactly this role |
| spreadsheet | an application laid out as a grid of cells that can hold text, numbers or formulae | printed in Part II §5.1, the unnumbered subsection "Spreadsheets" (Part II pp.111–112) |
| cell | one box of the grid, named by its column letter and row number | printed in bold at its first use (Part II p.112) |
| formulae | expressions typed into a cell that compute from other cells | printed in Part II p.112 |
| SUM, AVERAGE | the two spreadsheet functions the chapter uses, written with a range in brackets | printed in Part II p.112 and again in the formula bar on Part II p.113 |
| range | a block of cells written as its first cell, a colon, and its last cell | the chapter writes the pattern as Start:End (Part II p.112); range is an added label for it and is not printed in this chapter |
| weighted total | the sum in which each value is multiplied by its frequency before being added | an added term; the chapter performs exactly this computation on Part II p.110 and gives it no name |
| cumulative frequency | the running total of frequencies from the smallest value upwards | an added term; the chapter accumulates the frequencies on Part II p.110 without naming the quantity |
Where people slip up
- "Average the numbers in the left-hand column." This is printed as the tempting answer for a reason. It computes the average of the reported values, which is a different question, and it is wrong by more than a rounding: 6.5 against 5.22.
- "Divide by the number of rows." Divide by the total frequency. Eight rows, thirty-six students. The dart table makes it starker: ten rows, sixty-two students, two of the rows contributing nobody at all.
- "A frequency of 1 and a frequency of 11 are both just one row, so they count the same." That is the whole error, stated plainly. Show the table unfolded into thirty-six tick marks once and the objection disappears.
- "You must write out all the values to find the median." The chapter asks this question directly and answers it: accumulate the frequencies instead. For thirty-six values writing them out is merely tedious; for the sixty-two dart throwers it is a real obstacle.
- "The median is the middle row of the table." The middle row here is between 5 and 6, and the median is 5. Rows are values, not positions.
- "5.22 is the family size." No family has 5.22 members. It is the balance point of thirty-six family sizes, rounded to two places from 5.2222…
- "A spreadsheet knows statistics." It knows how to add up a named block. Every formula in this section is one of the two operations already done by hand, addressed differently.
- "B7 is a column." The chapter's own question says column and means cell. Use the slip: a letter alone names a column, a number alone names a row, and it takes both to name a cell.
- "Above 30 includes 30." Ashwin's Social Science mark is exactly 30. Strict comparisons are half the work in reading a table, and this item is built on one.
Ask your teacher a person
Your teacher reads this and writes back, usually within a day. For an instant answer, use Ask the video in the sidebar.
Your class sees the question and the answer. Only your teacher sees that it was you.
No questions on this topic yet.
Worked answers to this chapter’s exercises · this video explains Figure it Out · 5.1 Q10, Figure it Out · 5.1 Q11
Transcript1,274 words
A class is asked a simple question. How big is the average family in this room? Everybody says a number, and instead of writing thirty-six numbers on the board, somebody tallies them. Three students said three. Eleven said four. Nine said five, seven said six, three said seven, and one each said eight, nine and ten. Eight rows. Thirty-six students. That table is not a summary of the data. It IS the data, with the repetitions folded up - and every rule for computing with one comes straight out of unfolding it again.
Here is the answer somebody always gives first. Add up the left-hand column - three, four, five, six, seven, eight, nine, ten - and you get fifty-two. Divide by eight rows. Six point five. It is wrong, and it is worth more than being told so. That computation is not meaningless. It is the average of the eight family sizes that were REPORTED - one vote each, whether one student said it or eleven did.
A real quantity. Just not the one anybody asked for. So unfold the table and look at what the data actually is. The row that says three, three does not mean one three. It means three threes. The row that says four, eleven means eleven fours. Write them all out and there are thirty-six numbers on the board, most of them fours and fives. Now the objection dies on its own. A row with a frequency of one and a row with a frequency of eleven are not two equal things. One of them is eleven pieces of data.
You do not have to write them out, though. You only have to count each value as often as it occurs. Three threes are nine. Eleven fours are forty-four. Nine fives are forty-five. Seven sixes are forty-two. Then twenty-one, eight, nine and ten from the last four rows. Add those eight products and the whole class's families come to a hundred and eighty-eight people. That is the same total you would have got by adding thirty-six numbers. It is just the folded way of getting it.
Now what goes underneath. The total is a hundred and eighty-eight. Divided by what? Not by eight. Eight is how many rows the table has, and a row is not a student. A hundred and eighty-eight over eight is twenty-three and a half, which is not the average of anything in this room. Divide by thirty-six - the total frequency, which is how many students there actually are. That is the single most useful sentence in this topic. The divisor is the total frequency, and never the number of rows.
A hundred and eighty-eight over thirty-six. That does not come out exactly. It is five point two two two two, going on for ever. Reported to two places it is five point two two - and two different things are worth saying out loud about that number. The first: no family has five point two two people in it. It is the balance point of thirty-six family sizes, not a family.
The second: five point two two is rounded. The exact answer is not a decimal that stops. A class should always know which of those two it is looking at. Now the middle. Thirty-six values, so the median sits between the eighteenth and the nineteenth of them, in order. You could write out all thirty-six and count along. Do not. Run the frequencies up from the smallest value instead. Three. Then three plus eleven is fourteen. Then plus nine is twenty-three. Then thirty, thirty-three, thirty-four, thirty-five, thirty-six.
And here is what those running totals ARE. Each one is the last position that value occupies in the sorted list. Read them that way and the table answers a question about any position at all. The fourteenth value is a four, because the fours run out at fourteen. The twenty-third is a five, because the fives run out at twenty-three. So every position from the fifteenth to the twenty-third holds a five. That is the whole run.
Eighteen and nineteen are both inside it. Both are fives, so the median is five. Not five point two two, and not the middle row of the table either - the middle row sits between six and seven. Rows are values. Positions are data. Thirty-six numbers can be done by hand. A hundred and thirty-two cannot, comfortably - and that is where a sheet earns its place. Put a table of marks into one: a column of names, then a column for each subject.
Every box has an address made of two parts. A letter names the column, a number names the row, and it takes both to name a box. So E five is the fifth row's mark in column E - that student's maths mark, forty-two. And B seven is not a column, whatever it looks like. B is a column. Seven is a row. B seven is one box, and it holds twenty-seven.
The second idea is the one that does the work. You can name a whole block of boxes instead of listing numbers. Write the first box, a colon, and the last box. B three colon G three is one student's six marks, all the way across. Then ask for the total of that block, and the sheet answers two hundred and fifty. Ask for the average of B seven colon D seven - three of another student's marks - and it answers thirty exactly.
Notice what the sheet did not do. It did not know any statistics. It added up a block, and it added up a block and divided by three. The same two operations you have been doing by hand, with a shorter way of saying which numbers. Two tables to finish on, and a warning in the middle of them. Forty-two students recorded how many bicycle rides they took in a week, and the counts are drawn as a stack of dots over each number.
Same object as a frequency table. Forty-two students, a hundred and ninety-three rides, so the mean is four point six oh - rounded - and the median, read off the running totals, is four. Now the warning. A week has seven days. Anybody who rode more than seven times must have ridden twice on some day, and five students are above seven. But that is AT LEAST five, not exactly five. The plot records weekly totals, so a student with four rides could have taken two of them on one day and nothing on the board rules it out.
The data forces five and permits more. Knowing the difference between those two is the point. The last table is the one that makes the rule unmissable. Sixty-two students threw darts at a target, and the table records how many throws each needed. One student needed one throw. Nobody needed two, and nobody needed three. Ten rows. Two of them empty. So the divisor is not ten. Four hundred and seventy-three throws over sixty-two students is seven point six three, rounded.
And when you run the frequencies up, the empty rows do not move the running total at all. It stands still at one for three rows together. Which is exactly right, because a value nobody reported occupies no positions. The thirty-first and thirty-second are both eight, so the median is eight. The smallest anybody needed was one, and the largest was ten. A frequency table is a list with the repetitions folded up. Weight to average it, accumulate to find the middle, and never once divide by the number of rows.
Where this fits
Taken from the notes each video was made from, not from the reading order — these are the ideas this one rests on and the ones that later rest on it.
Builds on
- The mean as the point where the distances balanceClass 8 · Ch 5, Tales by Dots and Lines
- Whether adding a value raises or lowers the medianClass 8 · Ch 5, Tales by Dots and Lines
Comes up again in
- Telling a story with data, and letting it raise the next questionClass 8 · Ch 5, Tales by Dots and Lines
Either side of this one
- Working backwards from an average to a missing valueClass 8 · Ch 5, Tales by Dots and Lines
- Line graphs, and what change over time looks likeClass 8 · Ch 5, Tales by Dots and Lines