PrepShorts · Study sheet · Class 8 Mathematics · Chapter 5, Tales by Dots and LinesPrepShorts

Chapter 5 · Tales by Dots and Lines

Telling a story with data, and letting it raise the next question

Reading and making data graphics10 min

This video could not be loaded. Reload the page to try again.

Sign in with Google

10 min.

A data story does not end in a conclusion. It ends in a better question — one that names the data it would need.

The idea

A data story does not end in a conclusion; it ends in a better question. The chapter's own example is a curve of how long Indians sleep at each age, and the curve does something no single number could — it shows a shape, falling through the teens, flattening near eight hours in middle life and lifting again after fifty. But every "why" that shape provokes needs data the graph does not contain, because each point is an average over everyone of that age. That is not a defect to apologise for. It is how the work proceeds: a visualisation earns its place by making a pattern unmissable and by making visible exactly which question you must now go and collect data for.

What you should be able to do

  • Describe the shape of a curve in words, naming where it falls, flattens and rises
  • Read stated values off a curve at given ages, to the precision the axis supports
  • Explain why hundreds of closely spaced points look like a smooth curve, and why a column graph of the same data would be unreadable
  • Compare a full-scale panel with a zoomed panel of the same data, and say what each is good for
  • Separate the claims a curve of averages supports from the claims it cannot
  • Turn a "why" question raised by a figure into a statement of what data would answer it
  • Design a small collection: what to record, from whom, for how long, and with what definition of the quantity
  • Summarise collected data with a mean and a median for each of several groups, and say what each summary adds
  • Say what the next figure should show, given what the last one raised

Words to know

TermDefinition in one lineFirst introduced
data storya piece of writing built around a data set, which lays out what the data shows and what it asks nextprinted as the subsection heading "Data Story: Sleepy-Deepy" (Part II p.126)
smooth curvewhat a line graph looks like when its points are close enough together to hide the segmentsprinted in the discussion of the figure (Part II p.127)
data pointsthe individual plotted values a line graph joinsprinted in the same paragraph (Part II p.127)
column graphthe bar-style alternative the chapter rejects for this dataprinted in the same paragraph (Part II p.127)
sleep durationhow long someone sleeps in a day, the quantity plottedprinted in the paragraph introducing the figure (Part II p.126)
National Time Use Surveythe survey the figure's source line credits, dated 2024printed as the source line beneath the two panels (Part II p.126), read on the printed page because it does not extract
zoomed panela second copy of a figure drawn over a narrower range of the vertical axisthe chapter calls the second picture a zoomed-in version of the first (Part II p.126); the compound is added here
average across agesa value that summarises everyone of one age, and therefore describes no individualan added phrase; it is what every point on this curve is, and the chapter does not say so

Where people slip up

  • "A data story should end with an answer." This one ends with three questions, and the SUMMARY endorses that. Ending on a question is not a failure to conclude; it is what a finding looks like before the next collection.
  • "A smooth curve means the quantity varies smoothly." It means the points are close together. The data is still one value per age; the smoothness is a drawing effect, and the chapter says so.
  • "The zoomed panel exaggerates." Both panels are the same data. The full-scale panel understates the dip as surely as the zoomed one dramatises it, and the reason to print both is that neither is sufficient alone.
  • "The curve says people sleep less as they get older." It says people of different ages reported different amounts at one time. Following one person for sixty years is a different study, and it might not give this shape.
  • "So the average person sleeps eight hours — that is how much sleep I need." The text's seven-to-nine band is about the range across individuals; the curve's eight hours is a middle of that range. Neither is a prescription for anybody, and the paragraph before the figure lists the reasons an individual's need differs.
  • "The graph covers all ages." It runs from 6 to 75. Newborns, whom the closing question is about, are outside it, which is exactly why the question has to be asked rather than read off.
  • "Collecting data is the easy part." The group project makes the difficulty concrete: two groups that disagree about whether naps count have collected two different quantities and cannot pool them.
  • "An interesting question is a good question." A good question here is one that names the data that would settle it. Section 11 should convert each of the three into a sentence of the form "we would need to record …".
Transcript1,345 words

Some animals sleep two hours a day. Some sleep twenty. That is a fact you can put in one sentence, and it is the end of it. Now ask it about us. How long do people sleep, and does it change as they get older? That question does not have a one-sentence answer. It has a SHAPE, and a shape needs a picture. So here is a data story. And a data story does not end in a conclusion. It ends in a better question.

Before a single point is plotted, notice what has already been said. That people typically sleep seven to nine hours. That how much somebody needs depends on age, on where they live, on what they eat, on what they do all day. That babies sleep longer than adults. Every one of those is a claim. Not one of them is going to be tested by the graph. That is not a criticism either. Listing what you already believe, before you look, is what makes this a story and not just a figure.

Because at the end you will want to know which of those beliefs the data touched. The answer, mostly, is going to be none of them. The data is one number for every age, from six to seventy-five. That is seventy ages, and seventy points. Watch what happens as you plot them. Seven points, joined: you can see every segment. Twenty points: the corners are getting small. Seventy points: it is a smooth curve.

And here is the thing to be careful about. The quantity did not become smooth. The POINTS became crowded. There is still exactly one value per age. The smoothness is a fact about the drawing. Draw the same seventy numbers as columns and you get seventy columns, heavy and cluttered, and the pattern disappears behind them. That is why the line was chosen. Now read the shape, in words. It starts at nine point six hours at age six. It falls. Through the teens it falls fastest of all - eight tenths of an hour lost between twelve and nineteen.

By thirty it is down to eight point two. It bottoms out at forty, at eight point nought five. And then it turns. It lifts, slowly, all the way to eight point seven at seventy-five. Down, level, up. That is the finding, and no single number could have carried it. One more measurement before we go on. From the highest point to the lowest is one point five five hours. The whole story fits inside an hour and a half.

Which is a problem for the picture. Here is the curve on an axis running from nought to twelve hours - the honest, full scale, with nothing left out. The whole variation occupies somewhere between a ninth and a seventh of the height. It looks like a flat line with a wobble in it. So here is the same data again, on an axis from seven to twelve, ruled every half hour.

Now the dip through the teens is obvious. Now the lift after forty is obvious. Same numbers. Same curve. Different axis. And this is where somebody says: the second one is cheating. It is not, and it is worth being exact about why. The variation is drawn two point four times larger, because a twelve-hour axis became a five-hour axis. That is the whole of the effect, and you can state it.

And the second panel is not even full. There are two point four hours of empty axis above the curve, and just over an hour below it. The variation takes up between three tenths and a third of the height. Neither panel lies. The full-scale one understates the dip exactly as much as the zoomed one dramatises it. The fair thing is to print both, and to say the zoom factor out loud. A zoomed axis you have named is not a trick. A zoomed axis you have not named is.

So what does this figure actually support? That average sleep falls from childhood into middle life. That the fall is steepest in the teens. That the middle-life level is close to eight hours. That the trend turns upward after forty. And that the whole range is well under two hours. Five claims, all read straight off the picture, all safe. Now the other column. Every point on this curve is an average over everybody of that age.

So it cannot tell you how long any particular person sleeps. Not you, not anybody. And it cannot settle the seven-to-nine-hour claim from the beginning, because that claim is about the spread across individuals, and this curve threw the spread away when it took the average. Four people averaging eight hours might all sleep eight hours, or two might sleep six and two might sleep ten. The average is the same. The curve cannot tell them apart.

The figure is about averages. It is not about anybody. And there is a second limit, harder to see and more important. This data was collected at one moment. The point at age twenty and the point at age sixty are DIFFERENT PEOPLE. So the curve cannot tell you that getting older changes how you sleep. It tells you that people of different ages, asked at the same time, reported different amounts.

Following one person for sixty years is a completely different study, and it might not give this shape at all. One more thing it cannot do: the figure starts at six. It has nothing whatsoever to say about newborns. Which is exactly why the story ends where it does. On three questions. Why do newborns sleep so long? Do people in different places sleep differently? Do other animals show the same shape across their own ages?

An interesting question is easy. A GOOD question is one that names the data which would settle it. So: to answer the first, record sleep across a whole day, for babies under two, in families with a baby, for a fortnight. For the second, record sleep across a whole day and where the sleeper lives, from people of the same ages in several places, for a week. That is what turns a question into work. Who, what, and for how long.

Now try collecting some yourself, and meet the difficulty immediately. Track everybody in a family for a week. Here are three children, seven days each. One group counts night sleep only, and gets an average of nine point six eight hours. The other group counts naps as well, and gets ten point nought eight. Exactly twenty-four minutes apart. And both are right, because they are not measuring the same thing.

So their numbers cannot be pooled. Put them together and you get a figure sitting halfway between two quantities, describing neither. Which makes the most important sentence in the whole project the one that says what counts as sleep. Agree the definition first, or you have collected nothing you can add up. Agree it, then. Naps count. Now pool three age bands. Children, just over ten hours. Adults, seven point eight exactly. Elders, just under seven and a half.

And take the median as well as the mean, because they are not the same number, and the gap between them is telling you something about the shape of the group. Here is a cleaner one to finish on. How long is a school day? Five schools. Three of the five run exactly seven hours. One runs six and three quarters, one runs six and a half. The median is seven hours. The mean is six point eight five. The median says what a typical school does; the mean has been pulled down by the two short ones. Report both.

And notice what has happened. We started with a picture, and we have ended with a collection. That is the loop. Looking hard at data throws up fresh questions and fresh lines worth following. That is not the figure failing to conclude. That is the figure working.

Where this fits

Taken from the notes each video was made from, not from the reading order — these are the ideas this one rests on and the ones that later rest on it.

Builds on

Either side of this one

The book

Open in a new tab