Adding Context with Small Multiples and Reference Displays

Statistical analysis rarely asks only what values were observed. It usually asks how those observations compare with something else.

The comparison might involve:

  • another group;
  • another time period;
  • a historical average;
  • an expected range;
  • a target;
  • a control group;
  • a scientific benchmark;
  • a previous version of a system.

A graph without context may show what happened while making it difficult to determine whether the result is ordinary, unusual, improving, or deteriorating.

Data become informative through comparison.

Why Context Matters

Suppose a city recorded an average temperature of \(15^\circ\text{C}\). That number alone is difficult to interpret.

Relevant questions include:

  • Which month was measured?
  • Is this warmer than the historical average?
  • How variable are temperatures during that month?
  • How does it compare with nearby cities?
  • Is it part of a long-term trend?
  • Is \(15^\circ\text{C}\) inside the usual range?

Similarly, an unemployment rate of 7% has little meaning without information about:

  • the normal range;
  • previous years;
  • other regions;
  • economic conditions;
  • whether the rate is rising or falling.

Context converts an isolated value into an interpretable comparison.

Forms of Graphical Context

Context can be added to a graph in several ways:

  • reference lines;
  • reference bands;
  • historical ranges;
  • comparison groups;
  • targets;
  • annotations;
  • benchmarks;
  • repeated panels;
  • uncertainty intervals;
  • previous-period values.

Two particularly useful techniques are:

  1. small multiples;
  2. explicit reference regions.

What Are Small Multiples?

Small multiples are a collection of similar graphs arranged in a grid. Each panel displays a different group, category, location, or time period while using a consistent visual design.

For example:

January February March April
[boxplot] [boxplot] [boxplot] [boxplot]
May June July August
[boxplot] [boxplot] [boxplot] [boxplot]
September October November December
[boxplot] [boxplot] [boxplot] [boxplot]

Each panel uses the same:

  • type of graph;
  • variables;
  • visual encoding;
  • units;
  • axis meaning;
  • formatting.

Only the group or subset changes.

This makes small multiples a visual equivalent of a table: viewers can compare corresponding positions and patterns across panels.

Why Small Multiples Work

Small multiples take advantage of the human ability to notice:

  • repeated shapes;
  • changes in position;
  • differences in spread;
  • recurring peaks;
  • unusual deviations;
  • synchronized movements;
  • gradual trends.

Because every panel follows the same design, the viewer learns how to read one panel and can apply that interpretation to all the others.

The repeated structure reduces cognitive effort. Instead of decoding a new chart for every group, the viewer focuses on the differences between groups.

Consistency makes comparison possible; repetition makes patterns visible.

Example: Monthly Temperature Boxplots

Suppose a weather station has recorded monthly temperatures over 50 years. For each month, a boxplot summarizes the 50 annual observations for that month.

The January boxplot summarizes:\[ T_{\text{Jan},1}, T_{\text{Jan},2}, \ldots, T_{\text{Jan},50}. \]

The February boxplot summarizes:\[ T_{\text{Feb},1}, T_{\text{Feb},2}, \ldots, T_{\text{Feb},50}. \]

This continues through December.

Placed side by side, the twelve boxplots reveal the annual temperature cycle.

What the Monthly Boxplots Show

For each month, a boxplot displays information such as:

  • median temperature;
  • first and third quartiles;
  • interquartile range;
  • whiskers;
  • potential outliers.

The display supports several comparisons.

Seasonal center

The median line shows how typical temperature changes throughout the year.

A sequence such as\[ \text{Jan}<\text{Feb}<\text{Mar}<\cdots<\text{Jul} \]

may reveal warming into summer, followed by cooling toward winter.

Seasonal variability

The height or width of each box shows the IQR:\[ \operatorname{IQR}_m = Q_{3,m}-Q_{1,m}, \]

where \(m\) represents the month.

Some months may be much more variable than others. Transitional seasons may have wider boxes than months with more stable weather.

Extreme observations

Points beyond the whiskers identify years in which a month was unusually hot or cold relative to that month’s historical distribution.

Distribution asymmetry

A median positioned near one side of the box or one whisker extending farther may indicate skewness.

Why One Boxplot Is Not Enough

A single boxplot for all temperatures would combine winter and summer observations.

That combined distribution might have:

  • a wide range;
  • several peaks;
  • a median that represents no actual season particularly well;
  • apparent variability caused mainly by seasonal differences.

Separating the observations by month preserves the seasonal structure.

This illustrates an important analytical principle:

Variation within groups and variation between groups should not be confused.

Monthly boxplots distinguish:

  • variability from year to year within a month;
  • systematic differences between months.

Shared Scales Are Essential

Small multiples are easiest to compare when their axes use the same scale.

Suppose one panel’s vertical axis runs from 0% to 5%, while another runs from 0% to 15%. Identical line heights would represent different values.

This can create false impressions:

  • a small fluctuation may look dramatic;
  • a large fluctuation may look modest;
  • two panels may appear similar when they are not.

For direct magnitude comparisons, use:

  • identical axis limits;
  • identical tick spacing;
  • identical units;
  • identical aspect ratios where practical.

Free scales can be useful when the goal is to compare shapes rather than magnitudes, but this choice should be stated clearly.

Small Multiples for Time Series

A time-series plot places time on the horizontal axis and a measured value on the vertical axis.

For region \(r\), a time series can be written as:\[ y_{r,t}, \]

where:

  • \(r\) identifies the region;
  • \(t\) identifies time;
  • \(y_{r,t}\) is the observed value.

A small-multiple display creates one time-series panel per region:

Region A Region B Region C
/\ ___ /\/\
__/ \__ __/ \__ __/ \_
Region D Region E Region F
___/\___ _/\/\___ _____/\/

This arrangement makes it possible to identify both shared patterns and regional differences.

Example: State Unemployment Rates

Consider monthly unemployment rates for every US state from January 1976 through April 2009.

Each panel contains:

  • time on the horizontal axis;
  • unemployment rate on the vertical axis;
  • a line showing changes over time;
  • a reference band representing a selected comparison range.

The panel for state \(s\) shows:\[ u_{s,t}, \]

the unemployment rate for that state at month \(t\).

Viewing every state in a repeated grid makes it easier to detect:

  • national recessions affecting many states;
  • states with consistently higher unemployment;
  • states with relatively stable rates;
  • different recovery speeds;
  • regional economic shocks;
  • unusual local patterns.

Reference Bands

Suppose a shaded band covers unemployment rates from 4% to 6%.

Mathematically, the reference region is:\[ 4\%\leq u_{s,t}\leq6\%. \]

The band gives viewers an immediate visual benchmark:

Unemployment rate
10% | /\
8% | __/ \__
6% |============== ===== Upper reference boundary
| Reference range
4% |=========================== Lower reference boundary
2% |
+-------------------------------- Time

A viewer can quickly see:

  • when the series enters the band;
  • how long it remains there;
  • when it rises above it;
  • how far it departs from the reference;
  • whether it later returns.

A Reference Band Must Be Defined Carefully

Calling 4–6% a “normal” range requires justification. It might represent:

  • a historical central range;
  • a policy target;
  • an economic convention;
  • an organization’s operational threshold;
  • a range selected for illustration.

These interpretations are not equivalent.

A graph should state what the band means and how it was chosen. Otherwise, viewers may interpret a descriptive historical range as an official target or scientific boundary.

A better label might be:

Historical comparison range: 4–6%

rather than simply:

Normal

unless “normal” has a precise definition.

Detecting Common Shocks

When many panels show a rise at the same time, the cause may be a shared external event.

In unemployment data, a financial crisis may produce simultaneous increases across numerous states. Small multiples reveal this common timing while preserving differences in magnitude.

For example:

  • California and Michigan may show sharp increases;
  • Montana and South Dakota may show smaller increases;
  • some states may begin rising earlier;
  • some may recover more slowly.

A single national average would hide much of this variation.

Why Not Put Every Line on One Graph?

An alternative would be to draw all state unemployment rates in one panel.

With many lines, this often produces a spaghetti plot:

Rate
| /\/\_/\/\_/\/\_/\
|/_/\_/\/\__/\/\_/\
|\/\_/\/\_/\/\__/\/
+---------------------- Time

Problems include:

  • overlapping lines;
  • difficulty following one state;
  • confusing color legends;
  • poor accessibility;
  • hidden regional patterns;
  • excessive visual clutter.

Small multiples trade compactness for clarity. Each series gets its own panel, while alignment preserves comparison.

When a Single Combined Plot May Be Better

Small multiples are not always the best solution.

A single combined plot may be preferable when:

  • only two or three series are compared;
  • precise differences at the same time are important;
  • the series rarely overlap;
  • direct interaction is available;
  • one series serves as a benchmark.

For example, comparing one state with the national average may be clearer in a single panel containing two well-labeled lines.

The choice depends on the task:

  • use small multiples to compare patterns across many groups;
  • use a shared panel to compare a small number of values directly.

Ordering the Panels

Panel order influences what patterns are visible.

Possible arrangements include:

  • alphabetical order;
  • geographic order;
  • descending average value;
  • peak value;
  • trend similarity;
  • region;
  • cluster membership;
  • final-period value.

For state data, a geographic arrangement may reveal regional patterns. Sorting by average unemployment may reveal persistent differences. Clustering by time-series shape may reveal states that respond similarly to economic events.

There is no universally correct ordering. The arrangement should support the analytical question.

Adding a Shared Benchmark

Each panel can include a common reference series, such as a national average.

A useful design is:

  • state series in a strong color;
  • national series in a light gray;
  • identical axes across panels.

This allows the viewer to compare each state against the same benchmark without overcrowding one chart.

The deviation can also be plotted directly:\[ d_{s,t} = u_{s,t}-u_{\text{national},t}. \]

Then:

  • \(d_{s,t}>0\) means the state rate is above the national rate;
  • \(d_{s,t}<0\) means it is below;
  • \(d_{s,t}=0\) means they are equal.

Sometimes plotting deviations communicates the comparison more directly than plotting raw values.

Visual Hierarchy in Small Multiples

Repeated panels can become cluttered if every visual element is equally prominent.

A useful hierarchy is:

  1. the data series;
  2. important reference lines or bands;
  3. panel titles;
  4. subtle axes and gridlines;
  5. supporting annotations.

Good design choices include:

  • thin, light gridlines;
  • a restrained reference-band color;
  • consistent line widths;
  • short panel labels;
  • shared outer-axis labels;
  • minimal repeated legends.

The panels should look like coordinated parts of one display rather than dozens of unrelated charts.

Annotation Should Explain, Not Obstruct

Annotations can identify important events:

  • the beginning of a recession;
  • a policy change;
  • an extreme weather year;
  • a measurement-system change;
  • a major local event.

When the same event affects every panel, a shared vertical reference line may be more effective than repeating a label in every panel.

For an event at time \(t_0\):\[ t=t_0 \]

can be marked consistently across the grid.

This helps viewers examine whether changes occur before, during, or after the event.

Annotations should not imply causation unless the analysis supports that conclusion. A line marking an event shows temporal correspondence, not proof that the event caused the observed change.

Small Multiples and Boxplots Serve Different Comparisons

Both techniques provide context, but in different ways.

Repeated boxplots

Useful for comparing distributions across groups:

  • center;
  • spread;
  • skewness;
  • potential outliers.

Repeated time-series plots

Useful for comparing change:

  • trends;
  • timing;
  • volatility;
  • peaks;
  • recovery;
  • synchronized events.

Repeated histograms

Useful for comparing distribution shapes:

  • skewness;
  • multiple modes;
  • tails;
  • concentration.

The repeated structure is the common principle. The panel type should match the analytical question.

Small Multiples in Machine Learning

The same approach is valuable when evaluating models.

Performance by subgroup

Create one panel for each:

  • demographic group;
  • device type;
  • geographic region;
  • product category;
  • lighting condition;
  • noise level.

Learning curves

Use one panel per model to show:

  • training loss;
  • validation loss;
  • overfitting;
  • convergence speed.

Confusion patterns

Use comparable matrices for different models or data subsets.

Time-dependent performance

Display one time-series panel per model, region, or service to detect performance drift.

A single overall accuracy value may hide large differences across important subgroups.

Common Small-Multiple Mistakes

Using inconsistent scales without warning

This can make small and large changes look identical.

Repeating too much decoration

Full legends, dense tick labels, and heavy borders in every panel produce clutter.

Making panels too small

If trends or labels cannot be read, the repeated structure loses its value.

Using an arbitrary panel order

A thoughtful ordering can reveal patterns that a random order conceals.

Encoding groups only by color

Each panel should have a direct title so that interpretation does not depend on color matching.

Omitting the reference definition

A shaded region or benchmark must be explained.

Showing too many panels without structure

Large collections should be grouped, sorted, filtered, or made interactive.

A Practical Design Procedure

Step 1: Define the comparison

Examples:

  • How do temperatures differ by month?
  • Which regions reacted most strongly to a recession?
  • Which model performs poorly under low light?

Step 2: Select the panel unit

Choose one panel per:

  • month;
  • state;
  • model;
  • category;
  • subgroup;
  • sensor.

Step 3: Choose the appropriate chart

Use:

  • boxplots for distributions;
  • line charts for time series;
  • histograms for distribution shape;
  • scatterplots for relationships.

Step 4: Standardize the visual structure

Keep consistent:

  • axes;
  • units;
  • bin boundaries;
  • colors;
  • symbols;
  • reference elements.

Step 5: Add a meaningful reference

Possible references include:

  • historical median;
  • acceptable range;
  • baseline model;
  • control group;
  • national average;
  • previous year.

Step 6: Order panels intentionally

Arrange them to support comparison.

Step 7: Highlight the important pattern

Use restrained annotation or color to guide attention without obscuring the data.

Step 8: Verify interpretation

Confirm that viewers can identify both the common pattern and the important differences.

Key Takeaway

Statistical graphics become more informative when observations are compared with meaningful references. Small multiples present the same kind of graph repeatedly across groups, making differences in center, spread, shape, and change easy to detect. Shared scales, consistent design, intentional ordering, and clearly defined reference bands are essential. The purpose is not merely to show many graphs—it is to create a structured visual comparison that reveals patterns an isolated value or combined plot might hide.

Similar Posts

Questions, corrections, or additional insights?