PrepShorts · Study sheet · Class 7 Mathematics · Chapter 5, Connecting the Dots...PrepShorts

Chapter 5 · Connecting the Dots...

Clustered bar graphs: comparing across categories and across time

यह वीडियो हिंदी में भी · Watch in Hindi

Visualising data10 min

This video could not be loaded. Reload the page to try again.

Sign in with Google

10 min.

Also recorded in Hindi.Englishहिन्दी

A year of onion prices in two towns, drawn as two graphs on identical scales. Nothing is wrong with either one.

The idea

Two bar graphs printed one above the other cannot really be compared: the eye has to carry a height across a page gap and it is bad at that. Interleaving the bars, so that every category owns a little cluster of its own, replaces the memory task with a direct side-by-side look — and that is the only thing the clustered graph does. What it cannot do is tell you whether the left-to-right order of the categories carries meaning. Months are a sequence and shuffling them destroys information; organisations are a list and shuffling them costs nothing. The chapter puts those two graphs a page apart on purpose, and the reader has to supply the difference.

What you should be able to do

  • Read a value off a clustered column graph by referring to the markings on the vertical line
  • State the scale of a graph, and use it to estimate a bar that falls between two markings
  • Explain what the clustered form buys over two separate graphs of the same data
  • Say when reordering the categories changes the meaning of a graph and when it does not
  • Apply the chapter's two-step routine — first identify what is given, then infer from it
  • Distinguish a value read off the graph from an estimate, and say which you are giving
  • Choose a scale and draw a clustered graph from a table
  • Explain why a pattern fill as well as a colour is used to tell two series apart

Words to know

TermDefinition in one lineFirst introduced
data visualisationshowing data as a picture so it can be understood fasterprinted in §5.3, Part II, p.114
column grapha graph whose values are shown as the heights of upright barsprinted in §5.3, Part II, p.114
clustered column grapha graph with one small group of bars for each categoryprinted in bold in §5.3, Part II, p.115
double column graphthe same picture when each cluster holds exactly two barsprinted in bold in §5.3, Part II, p.115
double bar graphthe chapter's alternative name for the two-bar formprinted in §5.3, Part II, pp.117, 122
clustered bar graphthe name the SUMMARY settles on for the familyprinted in §5.3, Part II, p.119 and in the SUMMARY, Part II, p.134
scalehow many units of the quantity one step on the graph is worthprinted in §5.3, Part II, p.115
greyscaleprinting in black, white and grey only, with no colourprinted in §5.3, Part II, p.115
infographica picture that carries data alongside pictures and textprinted in §5.3, Part II, p.122
guidelinesthe faint ruled lines that help you read a bar's valueprinted in §5.3, Part II, p.116
categoryone of the things being compared, which gets one cluster of barsprinted as category in §5.3, Part II, p.117 and as categories in the SUMMARY, Part II, p.134
seriesone of the repeated bars inside every cluster, such as a single yearprinted in §5.2, Part II, p.98, but only in the cricket sense of a run of matches; the graph sense is an added term

Where people slip up

  • "Any two graphs of the same thing can be compared." Only if they share a scale, and even then the eye is unreliable across a page break. The chapter's p.114 pair is drawn on identical scales and is still hard to use, which is the honest version of the argument.
  • "A cluster of bars adds up to something." It does not. Two bars in one cluster are two separate measurements of the same category, not parts of a whole. This is the error a stacked bar chart invites and a clustered one must not.
  • "Reading a value off a graph gives you the value." It gives an estimate whose precision is set by the markings. The rocket graph's "perhaps 61" is the chapter modelling exactly this.
  • "You can always reorder the bars to make a graph tidier." Sorting the months of the year by price would destroy the seasonal pattern that is the whole reason for plotting them. Sorting the organisations loses nothing. The test is whether the category order is itself data.
  • "Colour is enough to tell two series apart." The chapter adds hatching and dots and says why: some readers cannot separate the colours, and greyscale printing removes them for everyone.
  • "A graph shows the numbers in the table." The daylight graph does not — it shows monthly totals divided by the days in the month. Always ask what quantity a bar's height actually is.
  • "If a graph has no numbers printed on it, it is a bad graph." The rocket graph is deliberately number-free and the chapter uses it to teach estimating. What makes a graph unreadable is a missing scale, not missing labels.
Transcript1,408 words

Here is a year of onion prices, in two towns, one price for every month. Draw the northern town as a column graph. Twelve bars, January to December, in rupees a kilogram. Now draw the southern town underneath it. The same twelve months, and the same scale, zero to sixty. Two honest graphs. There is nothing wrong with either of them. So answer one question. In May, which town was dearer, and by how much?

You just did something difficult, and you probably did it badly. To answer that, your eye had to find May in the top graph, measure the bar, hold the number, travel down the page, find May again, and only then compare. The measuring is not the hard part. The holding is. And the answer was worth having. In May the gap is eight rupees, and that is the widest gap between the two towns in the whole year.

The biggest thing in the picture, and it is nearly invisible, because the two bars that carry it are a whole graph apart. So do the obvious thing. Take January's two bars and slide them together until they touch. Then February's. Then March's. Twelve little groups of two, standing side by side along the bottom. Every month owns a cluster of its own now, and the comparison you were making from memory you can make by looking.

That is a clustered column graph. When every cluster holds exactly two bars, it also gets called a double graph. And that is the whole of what it does. Not one number has changed. One detail that looks decorative and is not. The two bars in each cluster are different colours. They also carry different fills. One is hatched with diagonal lines, the other is covered in dots. That is deliberate. It is belt and braces.

Some readers cannot separate the two colours at all. And any graph can end up photocopied, or come out in black and white, and then the colour is simply gone for everybody. So take the colour out and look again. The hatching and the dots still tell you which bar is which. A graph that only works in colour is a graph that stops working. Now read a value off it.

There is no number written on any bar. What there is, is a line up the side with markings on it, and that line is the only reason the picture means anything. Labels every ten, and a faint rule every five in between. That is the scale. October in the south. The bar reaches the sixty mark exactly. Sixty rupees, and it is the tallest bar on the graph. April in the north. It stops somewhere between the twenty five rule and the thirty label, nearer to thirty. About twenty eight.

Sixty is a reading. Twenty eight is an estimate. Always say which one you are giving. Here is the mistake this picture invites, and it is worth naming. February. Twenty four in the north, seventeen in the south. Two bars, side by side, touching. It is very tempting to see them as two parts of one thing. They are not. They are two separate measurements of the same month, and adding them together means nothing.

Twenty four plus seventeen is forty one, and forty one is not a price anybody paid, in either town, in any month of that year. A cluster is not a total. It is a comparison. Now the same idea, turned on its side. Ten launch providers, three years each. Thirty bars, running left to right instead of upward. These counts are drawn for this video, not measured. Everything you are about to do with them is real.

One cluster per provider, three bars in every cluster, one for each year. The line of values runs along the bottom this time, from zero to a hundred. With markings only every twenty. That last detail is about to matter more than anything else on the graph. Before you conclude anything, identify what you have actually been given. The scale first. One step along the bottom is twenty launches, and nothing finer is marked.

Second, the bottom row is not a provider at all. It says Other, and it is everybody too small to draw separately. It is a bin. Third, what is missing. There are no faint rules between the markings, so any bar that does not land near a multiple of twenty has to be guessed. Only now, infer. The top row climbs every year, and its middle bar is about twice its first. Perhaps sixty one, that one.

One row falls every year, and ends at half of what it was. And the bin at the bottom finishes somewhere around twenty five. Perhaps, about and around are doing real work in those sentences. And here is a question you should refuse to answer. Over the three years, did the Indian provider launch more than the second Chinese one? Look at the two clusters. They look identical. The true totals are fourteen and thirteen. One launch apart.

On a line whose nearest markings are twenty apart, one launch is a twentieth of the gap between two rules. That is not a thin difference. It is an invisible one. So the honest answer is that this graph does not know, and neither do you. But notice you can still see something. In the first year the Chinese provider was ahead. In the third, the Indian one was. The shape survives even when the total does not.

Now the thing neither picture can tell you. Both of them have categories along one edge. Months on one. Providers on the other. You could sort either of them by height, biggest first, and the graph would look tidier. Do that to the providers and you lose nothing at all. Every row carries its own name, so the order was never holding any information. Do it to the months and you destroy the graph.

In calendar order, nine of the eleven steps from one month to the next are a rise. Prices climb almost all year and then drop. Shuffle the months at random and you would expect five and a half rises. Sort them by height and you get eleven, every single step a rise, and it tells you nothing whatsoever. The order of the categories can be data. The graph will not warn you either way.

Two more cities, and a graph that does not plot the numbers it was handed. Here are the hours of daylight each one gets, month by month, across a year. Two rows of twelve totals. You cannot put those straight onto a graph, because February is three days shorter than October. A short month looks dark whether it is dark or not. So divide each month's total by the number of days in that month, and plot the hours per day instead.

Watch what that fixes. In the second city, October carries twenty five more hours of daylight than February does. And February's days are the longer ones, by half an hour each. Now the picture. The first city runs from six hours a day in December up to eighteen point eight in June. Exactly a quarter of the day, and very nearly three quarters. The second city does the opposite. Its best month is December and its worst is June.

One of them is far north. The other is far south. Each one's summer is the other's winter. One last graph. Twenty overs of a cricket match, two teams, one cluster per over. Drawn for this video again. Scale up the side, one step to five runs. And a small circle above a bar means a wicket fell in that over. That circle is a second kind of information riding on the same picture without needing a second scale, and it costs nothing.

Here are questions it answers instantly. In over twelve, thirteen runs and eleven. The second team's leanest over is the sixteenth, with two. And here is one it cannot answer at all. What was the target? To get that you would have to add twenty bars together. The graph gives you every single over and no total, because a cluster is not a total, and neither is a graph full of them.

It answers what it was drawn to answer. The habit worth keeping is asking what that was.

Where this fits

Taken from the notes each video was made from, not from the reading order — these are the ideas this one rests on and the ones that later rest on it.

Builds on

Comes up again in

Either side of this one

The book

Open in a new tab