Skip to main content

Chapter 12 Data Visualization For Analysis Tableau I: Chapter 12 Data Visualization For Analysis Tableau I

Chapter 12 Data Visualization For Analysis Tableau I
Chapter 12 Data Visualization For Analysis Tableau I
  • Show the following:

    Annotations
    Resources
  • Adjust appearance:

    Font
    Font style
    Color Scheme
    Light
    Dark
    Annotation contrast
    Low
    High
    Margins
  • Search within:
    • My Notes + Comments
    • Notifications
    • Privacy
  • Project HomeAPPLIED BUSINESS ANALYTICS FOR MARKETING DECISION-MAKING
  • Projects
  • Learn more about Manifold

Notes

table of contents
  1. Chapter Information
  2. Chapter Learning Objectives
  3. 12.1 Marketing Decision Context: Seeing the Expansion Story
    1. 12.1.1 Opening Case Questions
    2. 12.2 Visualization's Two Jobs
    3. 12.3 The Grammar of a Chart
    4. 12.4 The Perceptual Hierarchy
    5. 12.5 Choosing Charts by Analytic Purpose
    6. 12.6 Scales and Honest Axes
    7. 12.7 How Charts Mislead
    8. 12.8 Color as an Encoding System
    9. 12.9 Reference Points, Baselines, and Annotation
    10. 12.10 The Tableau Data Model
    11. 12.11 From Prep to First Charts
    12. 12.12 AI as a Chart-Drafting Assistant
    13. 12.13 Hands-On Application in Tableau
      1. 12.13.1 Lab 12.1, Part A: The Prep Flow and the Two Grains
      2. 12.13.2 Lab 12.1, Part B: Six Purposes, Six Views
      3. 12.13.3 Lab 12.1, Part C: Small Multiples and the Store-Age Artifact
      4. 12.13.4 Lab 12.2, Part A: The AI-Drafted Chart, Predicted and Diagnosed
      5. 12.13.5 Lab 12.2, Part B: The Repair, and What Changed in the Reading
    14. 12.14 Marketing Interpretation and Managerial Insight
    15. 12.15 Business Analytics in Practice
    16. 12.16 Ethics, Misleading Mechanics as an Integrity Obligation
    17. 12.17 Chapter Summary
    18. 12.18 Exercises for Practice and Homework
      1. 12.18.1 Concept Check (Required Practice)
      2. 12.18.2 Predict the Encoding (Required Practice)
      3. 12.18.3 Purpose First (Required Practice)
      4. 12.18.4 The Grain Trap (Required Practice)
      5. 12.18.5 The Visual Audit Deliverable (Homework Submission)
      6. 12.18.6 AI Chart Audit (Homework Submission)
      7. 12.18.7 Spot the Mechanic (In-Class Discussion)
      8. 12.18.8 The Colleague's Deck (In-Class Discussion)
    19. 12.19 Glossary of Terms
    20. 12.20 Further Readings
    21. 12.21 References

Data Visualization for Analysis (Tableau I)

Seeing Before Persuading

Dr. Jose Mendoza, Academic Director and Clinical Associate Professor

Version 1.1 · July 2026

Except where otherwise noted, this chapter is licensed under CC BY 4.0.

Chapter Information

ABSTRACT

This chapter builds the analytical half of data visualization, using StyleCraft's store-expansion question as its running material and Tableau as its instrument. It establishes the grammar of a chart — marks, encodings, scales, aggregation, and sort — and the perceptual hierarchy that orders encoding channels by decoding accuracy, converting chart selection from taste into a prediction. Charts are chosen by analytic purpose rather than from a gallery, scales and baselines are assigned by what a mark encodes, and five standing misleading mechanics are demonstrated on StyleCraft's numbers and repaired. Color is treated as an encoding system in which accessibility is a correctness requirement, and reference lines supply Chapter 2's baseline discipline in visual form. A Tableau Prep flow certifies the extract at two grains, six views build the expansion story, and small multiples test whether the newest stores' apparent weakness is largely an artifact of store age.

KEYWORDS

data visualization; perceptual hierarchy; chart selection; visual encoding; axis truncation; color accessibility; reference lines; Tableau; aggregation; small multiples

VERSION AND DATE

Version 1.1 · July 2026 · Language: English (United States)

SUGGESTED CITATION

Mendoza, J. (2026). Data visualization for analysis (Tableau I). In Applied business analytics for marketing decision-making: Business analytics and data visualization (Chapter 12, Version 1.1) [Open educational resource]. CC BY 4.0.

LICENSE AND RIGHTS

Copyright © 2026 Jose Mendoza. Except where otherwise noted, this work is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). You may share and adapt this material for any purpose, provided appropriate credit is given. Third-party trademarks, screenshots, figures, and other materials remain subject to their respective rights and licenses.

Tableau, Tableau Desktop, and Tableau Prep are trademarks of Salesforce, Inc. Google Colab is a product of Google LLC. ChatGPT is a product of OpenAI. Claude is a product of Anthropic. Gemini is a product of Google LLC. GitHub Copilot is a product of GitHub, Inc. NotebookLM is a product of Google LLC. Product names are used for identification only and do not imply endorsement. StyleCraft Collective is a fictional company created for instruction.

COMPANION REPOSITORY

Datasets, notebooks, and figure sources for this chapter: https://github.com/jrmst102/businessanalytics


GENERATIVE AI USE

Generative artificial intelligence and other AI-assisted tools were used in the research, writing, revision, and production of this chapter, including literature discovery, source organization, outlining, preliminary drafts, prose revision, support for code and analytical examples, and document formatting. These tools were used under the author's direction and are not credited as authors, researchers, or sources. The author determined the chapter's scope, learning objectives, methods, interpretations, and recommendations, and reviewed and approved all AI-assisted material: factual claims and citations were checked against the underlying sources rather than accepted from AI-generated summaries, and code and analytical outputs were tested or otherwise reviewed for accuracy. Responsibility for the accuracy, originality, and final form of this chapter rests entirely with the author. A fuller statement appears in the front matter of the complete guide.

Chapter Learning Objectives

By the end of this chapter, students should be able to:

  1. Distinguish visualization's analytical job from its communicative job, state which job a given chart was built for, and explain why a chart that succeeds at one can fail at the other.
  2. Decompose any chart into its grammar — the marks it draws, the data encoded to each visual channel, and the axes and scales those channels are read against — and describe the chart types used since Chapter 4 in that vocabulary.
  3. State the perceptual hierarchy of encoding channels, explain the evidence behind it, and predict which of two candidate encodings a reader will decode more accurately for a stated comparison.
  4. Select a chart form from an analytic purpose — distribution, comparison, trend, relationship, composition, or geographic — using a purpose-first selection table rather than a chart-type gallery, and justify the selection in terms of the comparison the reader must make.
  5. Choose between linear and logarithmic scales for a stated marketing question, and apply the zero-baseline obligation by chart type rather than as a blanket rule.
  6. Recognize, demonstrate, and repair the five standing misleading-chart mechanics — axis truncation, dual axes, cherry-picked windows, area distortion, and three-dimensional effects — and state what each does to a reader's estimate of the underlying quantity.
  7. Match a color palette type — sequential, diverging, or categorical — to a variable's role and level of measurement per Section 3.4, and evaluate a palette for contrast and color-vision deficiency as a condition of correctness rather than of politeness.
  8. Add reference lines, baselines, and annotation that supply the analytic context a chart cannot otherwise carry, explaining this as the visual form of the baseline discipline introduced in Section 2.7.
  9. Map Tableau's dimensions and measures onto the variable roles of Section 3.5, distinguish discrete from continuous fields, and explain how aggregation in the view can silently answer a question at a grain other than the one asked, per Section 3.3.
  10. Build a Tableau Prep flow that produces a certified extract at a stated grain, with a row-count and total-revenue reconciliation that verifies the flow did what was intended.
  11. Construct core chart forms in Tableau, including small multiples and faceted views, and state what a facet reveals that an aggregate cannot.
  12. Audit an AI-drafted chart by predicting its faults from Sections 12.3 through 12.9 before looking at it, correcting it against the Appendix E checklist, and documenting the exchange per Appendix D.
  13. Assemble a visual audit of a marketing story that a colleague can reproduce: the question each view answers, the encoding chosen and why, the reference context supplied, the grain the view aggregates at, and the charts considered and rejected.

Chapter 11 closed Part II with an interval and a rule for acting on it. Eleven chapters of machinery now sit behind the analyst as a shelf of finished evidence: a cleaned transaction file with a documented verification log, a decomposition of the suburban average-order-value gap, four recovered customer segments, a driver model with coefficients in dollars, a scored churn list with a threshold derived from retention economics, a forecast with a prediction interval that survives the store-opening inflection, and an experimentally measured lift with a break-even line beneath it. Every one of those artifacts is currently a table, a metric, or a paragraph — which is to say that every one of them currently exists in a form that only the analyst who built it can read. That is the problem Part III takes up, and it is not a cosmetic one. A finding that cannot be seen is a finding that will be re-litigated by whoever has the most confident slide, and the great majority of analytical work that dies in organizations dies in the last two feet between the notebook and the meeting. Part III therefore treats visualization not as the decoration applied to finished analysis but as an instrument with its own failure modes, its own verification discipline, and its own capacity to produce confident wrong answers. This chapter owns the first of visualization's two jobs — the analytical one, in which a chart is built so that the analyst can find out what the numbers contain, before anyone has decided what the numbers should say. Chapter 13 will own the second. The distinction sounds academic and is entirely practical: the chart that helps an analyst discover the expansion story and the chart that helps a chief executive act on it are different objects, built to different standards, and confusing them is how a room ends up looking at a beautiful picture of nothing.

Concept

What This Chapter Is Really About

A chart is a machine for making a comparison easy. That is the whole of it, and nearly every practical rule in this chapter is a consequence of taking the sentence literally. If the comparison a reader must make is "which of these seventeen stores earned the most," the machine that makes it easy is a sorted bar chart, because human beings compare lengths against a common baseline with high accuracy and compare wedge angles poorly. If the comparison is "did the daily series step up when the store opened," the machine is a line with a dated reference mark, because the eye follows connected position over time and cannot hold nine columns of a table in memory long enough to notice a break. And if a chart is built so that the comparison the reader most needs is the one the encoding makes hardest — categories in a pie, magnitudes in a bubble's area, two unrelated series forced onto twin axes — then the chart has failed at its only job, regardless of how good it looks or how correct its underlying query was. This is why the chapter spends its first half on perception rather than on software. Tableau will draw whatever it is asked to draw, in seconds, beautifully, and an AI assistant will now ask it on the analyst's behalf; neither has any opinion about whether the resulting picture supports the comparison anyone actually needs. The judgment that remains the analyst's is therefore the judgment about what to encode where, which is where every chart is won or lost, and it is a judgment that can be checked in advance. Predict what the reader will see, build it, then look and grade the prediction. That is predict-then-verify, arriving in Part III as the discipline that turns chart-making from a matter of taste into a matter that can be right or wrong.

12.1 Marketing Decision Context: Seeing the Expansion Story

Project #1 was presented last Tuesday. The evidence is assembled, defended, and — for the first time since Chapter 5 — nobody is asking the analyst for another number.

What the VP of Marketing is asking for instead is a pre-read. On December 2, StyleCraft's chief executive convenes a long-scheduled review: whether to commit to eight suburban and resort sites under letter of intent, or to stop the expansion and redirect the capital into the urban and digital business. Flagship candidates remain in the pipeline but are not yet under letter of intent and require a separate economic benchmark; the real-estate committee needs a decision before the leases price, and the decision will be made in one meeting, from one packet, by people who have not read a line of the analysis behind it. Chapter 13 will build the artifact that meeting runs on. This week's job comes first: the VP wants a visual read of the expansion story — everything Part II established about the suburban and resort cohort, rendered so that a person can see it — circulated to the review's four participants by Monday.

The commission is unusual in one respect worth naming immediately, because it changes what "good" means. Nobody is being persuaded this week. The pre-read's four readers are the analyst's own colleagues — the VP, the finance partner from Chapter 11, the head of merchandising, and the director of real estate — and their job on receipt is to disagree, to notice what is missing, and to arrive on December 2 having already had the argument. A pre-read that makes the expansion look good is a failed pre-read, and so is one that makes it look bad. What is wanted is a set of views in which the story is visible enough to be argued with — a materially different design brief from the one Chapter 13 will answer, and the subject of Section 12.2.

There is also a complication, and it arrived Thursday afternoon in an email with six charts attached. The merchandising analyst — capable, junior, and under the same deadline — asked an AI assistant to "make some charts showing how the suburban stores are doing," pasted the resulting specifications into Tableau, and circulated the result in about ninety minutes. The deck is attractive. It uses the brand's colors, it is titled in complete sentences, and it has already been forwarded to the director of real estate with the note "this is great, let's use these." It also contains, on close reading, four separate errors of the kind this chapter exists to catch, and a fifth that lives in the covering email rather than in any image: the email describes the loyalty-tier chart as showing "average order value," while the worksheet behind it computes AVG(line_revenue) on a line-grain extract. The lead chart shows quarterly revenue in the suburban and resort stores rising from $241,000 to $252,000 — a real increase of 4.6 percent — drawn on a vertical axis that begins at $240,000 and ends at $255,000, so that the second bar stands roughly twelve times the height of the first. The second chart puts suburban revenue and cumulative store count on twin vertical axes whose ranges have been chosen, by the software's defaults rather than by anyone's intent, so that the two lines lie almost exactly on top of each other, inviting the reading that each opening produced the revenue that followed it. The third is a pie chart of seven product categories, from which no reader can rank the middle four. The fourth encodes StyleCraft's four loyalty tiers — None, Insider, VIP, and Backstage — in a red-to-green palette that implies an ordering the tiers do have and a valence they do not, and that renders two of the four tiers indistinguishable to the roughly one in twelve male readers with a color-vision deficiency. A distribution list of this size could easily include such a reader, and the design obligation does not depend on knowing whether it does.

None of that was done in bad faith, and the distinction matters for how the analyst handles it. The merchandising analyst did not truncate an axis to mislead the real-estate director; the assistant proposed a chart specification, the software rendered it with a default axis range chosen to fill the available space, and nobody in the chain had a reason to look at the axis. This is the modal way misleading charts enter marketing organizations in 2026. They are not forged. They are defaults, accepted quickly, by people with a deadline and no framework for knowing which defaults are dangerous. Section 12.7 supplies the framework, Section 12.12 supplies the audit that catches it before circulation, and Lab 12.2 makes the student produce and repair the failure herself.

So the commission has two halves, and they are the chapter's two labs. The first is constructive: build the expansion story out of transactions_enriched, in six views chosen by analytic purpose, at grains that survive inspection, with the reference context that lets each view be read against something. The second is corrective: take the circulating deck apart against a checklist, repair each chart, and state — in a sentence per chart, for the merchandising analyst rather than for the file — what changed in the reading. And the standard for both is the one the pre-read's purpose implies. Every view must be a view the analyst would be willing to have argued with, which means every view must be built so that a reader who disagrees can see exactly what she is disagreeing with.

12.1.1 Opening Case Questions

Keep these questions in mind while reading, and return to them after completing the labs.

  • The lead chart in the circulating deck is arithmetically correct: suburban and resort revenue really did rise from $241,000 to $252,000. Write one sentence naming what the chart nonetheless communicates falsely, and one sentence explaining why the phrase "but the numbers are right there on the axis" is not a defense.
  • The pre-read is meant to be argued with rather than agreed with. List two design choices you would make differently for a chart built to be disagreed with than for a chart built to be accepted, and predict which section of this chapter each choice comes from.
  • The dual-axis chart invites the reading that store openings produced the revenue that followed them. Name the chapter and section of Part II that already supplies the correct objection to that reading, and state what the chart would have to show for the objection to be answered.
  • Before Section 12.4 tells you, rank these four encodings by how accurately you think a reader can compare two values through them: length of bars from a common baseline, area of circles, angle of pie wedges, position of points along a common axis. Write your ranking down and grade it later.

12.2 Visualization's Two Jobs

The pre-read due Monday and the dashboard due in December are both collections of charts about the same company, built from the same file, in the same software, by the same analyst. They are not the same kind of object, and the reason they are not is the distinction this section installs — a distinction that has been implicit since Chapter 4, when the first histogram was drawn to find out whether the order-value distribution was skewed, and which becomes load-bearing now that Part III has to build both kinds deliberately.

A chart has a dominant job, and it is one of two. In its analytical job, the chart is an instrument of discovery: it is drawn by the analyst, for the analyst, to find out what the data contains, and its success criterion is whether it reveals structure that the analyst did not already know — a break in a series, a bimodal distribution, an outlier store, a relationship that reverses inside a subgroup. In its communicative job, the chart is an instrument of transfer: it is drawn by the analyst, for someone else, to convey a finding already established, and its success criterion is whether a reader who has never seen the data arrives at the intended understanding in the time available. Gelman and Unwin (2013) put the tension precisely, noting that the graphics statisticians build to interrogate data and the graphics designers build to communicate results are optimized against different objectives and routinely disappoint each other's audiences. Tukey (1977) had staked out the first job a generation earlier under the name exploratory data analysis, whose governing image — the analyst as detective, drawing quickly and disposably in order to notice — remains the best description of what an analytical chart is for.

Concept

Analysis Charts and Communication Charts Are Different Objects

An analytical chart is disposable, dense, unpolished, and drawn to answer a question the analyst has not yet answered; twenty of them may be made in an hour and nineteen thrown away. Its labeling is spare rather than absent: a private, disposable view can be terse, but any view that will be revisited, reused, or reproduced still needs units, the filters applied, the grain one mark represents, and field labels a reader can interpret. The difference between the two kinds of chart is density and polish, not whether the chart can be understood later. A communication chart is durable, sparse, finished, and drawn to convey a finding the analyst has already established; there may be four of them in a deliverable, each surviving several rounds of revision, each titled with the sentence it wants the reader to leave holding. The properties that make each good actively harm the other. Density that lets an analyst scan seventeen stores at once overwhelms an executive with ninety seconds; simplification that lets an executive grasp one comparison instantly hides the second comparison the analyst needed to see. Hence the working rule this chapter and the next enforce together: decide which job a chart has before drawing it, and never let a chart drift from the first job to the second without being rebuilt. The two jobs are dominant purposes rather than exclusive categories — some artifacts, the pre-read among them, serve both audiences at once, and naming which job dominates is what makes the design brief decidable. The most common form of that drift is the exploratory view — made in five minutes, full of defaults, correct for its purpose — that gets pasted into a deck because it was already there. This chapter owns analysis charts. Chapter 13 owns communication charts, and its first act will be to rebuild rather than to reuse.

Figure 12.1

Visualization's Two Jobs

Figure 12.1 sets the two objects beside each other. In other words, the pre-read is an analytical artifact serving a semi-public purpose, which is exactly the awkward middle case that makes it a good teaching problem. Its four readers are colleagues who will interrogate it, so it must carry the density and the honesty of an analytical view; but it leaves the analyst's screen, so it must be legible to someone who did not build it. The resolution the chapter adopts is to hold the analytical standard and add only what legibility requires — real axis labels, units, a note of the grain each view aggregates at, and a reference line where a number needs something to be read against. What it does not add is the assertion titling, the aggressive decluttering, and the narrative sequencing that Chapter 13 will teach, because those techniques direct a reader toward a conclusion, and directing this week's readers toward a conclusion is precisely what the VP does not want.

One consequence of the division deserves stating now, because it governs the rest of the chapter. Analytical charts are where an analyst's own errors are caught, which means the analytical stage is the one that has to be built for self-suspicion rather than for speed. Chapter 4's exploratory charts were read for anomalies; Chapter 7's scatterplots enforced the always-plot discipline that a correlation coefficient alone will not; Chapter 10's decomposition plot made a structural break visible that no summary statistic named. Every one of those was a chart doing the analytical job, and in every case the value came from the chart's willingness to show something inconvenient. A chart built to make a point will not do that, because a chart built to make a point has already decided what it is going to find. The order matters, and it is the order Part III follows: see first, then persuade — and only about what you saw.

12.3 The Grammar of a Chart

Chapter 4 introduced four chart types by name and used them: the histogram for shape, the box plot for spread and outliers, the bar for comparison across categories, and the scatterplot for relationship. They were introduced as objects to be read, and the reading was enough for the job. This section supplies the theory that was deferred — the observation that those four are not four unrelated inventions but four points in a small space of choices, and that once the space is named, a chart stops being something an analyst selects from a gallery and becomes something she specifies.

The idea is due to Bertin (1983), who cataloged the visual variables available for encoding information on a plane, and to Wilkinson (2005), who formalized the result as a grammar in which a chart is generated by combining a small number of independent components rather than chosen from a fixed list; Wickham (2010) rendered the grammar in the layered form that most modern software implements, Tableau included. The practical payoff for a marketing analyst is not the formalism. It is that a chart can be specified in a sentence, audited component by component, and altered one component at a time.

Definition

Mark and Encoding

A mark is the geometric object a chart draws to represent data — a point, a line, a bar, a rectangle, a circle, an arc, or a filled area — and a chart's mark type determines what kind of thing the reader is looking at. An encoding is the assignment of a data field to a visual channel of those marks: position along an axis, length, angle, area or size, color hue, color intensity, shape, or texture. A chart is fully specified by its marks, its encodings, and the axes and scales against which the positional channels are read; two charts with the same underlying query and different encodings are different instruments answering different questions.

Source: Adapted from Bertin (1983) and Wilkinson (2005).

In other words, when the merchandising analyst's fourth chart maps loyalty tier to color hue, that is an encoding decision, separable from the decision to use bar marks and from the decision to put revenue on the vertical axis, and it can be criticized and replaced without touching either. Take the four chart types of Chapter 4 and read them in this vocabulary. A histogram is a bar mark whose horizontal position encodes a binned continuous variable and whose height encodes a count; the binning is a transformation applied before encoding, which is why bin width changes the picture without changing the data. A box plot is a composite mark whose position encodes distributional summaries — the quartiles, the whiskers, the outlying points — against a single continuous axis. A bar chart of revenue by store is a bar mark whose length encodes revenue and whose categorical position encodes store, with sort order as a design choice that carries real information and is almost always left at the software's default. And a scatterplot is a point mark with two positional encodings, which is why it is the most information-efficient chart in common use and why Chapter 7 insisted on it: both of its channels are the ones readers decode best.

Three components sit underneath the marks and are easy to overlook precisely because software supplies them automatically. The first is the axis and its scale, which determines what a unit of visual distance means and which Section 12.6 takes up in full. The second is the aggregation applied before the mark is drawn: a bar showing "revenue by metro" is showing a sum, and one showing "order value by metro" is almost certainly showing an average, and the two are different claims about the world made with identical geometry — the point Section 12.10 will make operational when Tableau's aggregation in the view is introduced, and the point Section 3.3 established when it observed that changing grain changes the claim. The third is the sort order, which is not decoration. A bar chart of seventeen stores sorted alphabetically and the same chart sorted by revenue contain the same numbers and support different reasoning, because sorting by the encoded value turns the comparison "which is largest" from a search problem into a reading problem.

The reason to hold the grammar in mind rather than a gallery of chart types is that the gallery does not scale and the grammar does. There are perhaps forty named chart types in circulation and an analyst will meet a new one every few months, usually with a vendor's name attached; there are half a dozen mark types and eight or nine encoding channels, and every new chart type is a recombination of them. A radial bar chart is a bar chart with position mapped to angle instead of to a linear axis, which tells you immediately what it costs (Section 12.4 will say why) without requiring any prior opinion about radial bar charts. This is also, and not incidentally, why the grammar is what makes an AI-drafted chart auditable: the assistant returns a specification, the specification decomposes into marks, encodings, scale, aggregation, and sort, and each component can be checked against the question the chart was supposed to answer.

12.4 The Perceptual Hierarchy

Section 12.3 established that a chart's designer chooses which channel carries which field. This section supplies the fact that makes the choice consequential rather than aesthetic: the channels are not equally good, the ordering among them has been measured, and the ordering is stable enough to be used as a design rule and as a prediction.

The foundational work is Cleveland and McGill (1984), who ran controlled experiments in which subjects judged the ratio between two encoded values and whose accuracy was recorded by channel. Important parts of the resulting ranking have been replicated at scale on crowdsourced platforms with modern stimuli and have held up well (Heer & Bostock, 2010). What that evidence supports is a robust average ordering for elementary quantitative judgments rather than a universal ranking: measured performance varies with the task, the display, the reader, and the context. The ordering is short enough to memorize and stable enough to design against, which is all a working rule needs to be.

Definition

Perceptual Hierarchy

The perceptual hierarchy is the empirically measured ordering of visual encoding channels by the accuracy with which readers extract quantitative values from them. Judgments of position along a common scale are most accurate, followed by position along identical non-aligned scales, then length, then angle and slope, then area, then volume, and finally color intensity and saturation. Because accuracy degrades substantially down the ordering, a quantitative field that must be compared precisely belongs on a high-ranking channel, while low-ranking channels are better used for fields that only need to be grouped, ordered coarsely, or distinguished. Two conditions belong to the definition rather than being caveats appended to it. The ordering ranks decoding accuracy alone, and accuracy is not the only design objective, so a low-ranking channel can be the correct choice for a field that need not be decoded precisely. And the ordering governs the comparison the reader actually makes, which the chart's layout determines as much as its channels do. The hierarchy is therefore a robust average ordering for elementary quantitative judgments, not a universal ranking that overrides a stated analytical purpose.

Source: Adapted from Cleveland and McGill (1984) and Heer and Bostock (2010).

In other words, the eye is a good ruler, a mediocre protractor, and a poor scale for judging area, and a chart's design should spend its best channel on its most important comparison. The consequences are immediate and they explain most of the practical advice in circulation that otherwise sounds like taste. A sorted bar chart beats a pie chart for comparing category magnitudes because bars use length from a common baseline and pies use angle and arc. Bubble maps understate large differences because readers judge circles by something between radius and area and therefore compress the range — which is why a store with four times another store's revenue does not look four times anything. Encoding a quantity by color intensity alone, as a heat map does, is acceptable when the reader needs to spot hot regions and unacceptable when she needs to know whether Rye beat Saddle River. And the reason Chapter 7's scatterplots carried so much weight is now visible: a scatterplot spends both of the top two channels on the two variables whose relationship is the entire question.

Both conditions in the definition deserve a worked instance, because they are what keep the hierarchy from being applied mechanically. On the first, grouping, memorability, and the ability to show many series at once are real objectives, and a channel low in the ordering is the right choice for a field that does not need precise decoding — StyleCraft's four metros carried by color hue on a scatterplot, where the reader needs to tell the groups apart, not to rank them. On the second, bars in a single sorted column are compared along a common baseline; the same bars split into four side-by-side panels are compared across identical non-aligned scales, one rank lower, which is the small price small multiples pay and Section 12.11 will weigh against what they buy. Figure 12.2 puts the ordering on one set of values so that it can be experienced rather than accepted. Before looking at it, fix the task: in each of the four panels, estimate how many times larger the largest value is than the smallest, and write the four estimates down before reading on. Zacks and Tversky (1999) added a related caution worth carrying: readers infer meaning from mark type itself, reading bars as discrete comparisons and lines as continuous trends, so a line drawn across categorical positions implies a continuity that does not exist, and a bar chart of a time series discards the trend reading the eye is best at.

Figure 12.2

The Perceptual Hierarchy in Practice

The four panels encode the same four values — 8, 14, 23, and 46, so the largest is 5.75 times the smallest. Compare that against the four estimates. Most readers land close on the aligned points and the bars, drift on the wedges, and compress the ratio badly on the circles, which is the ordering the hierarchy predicts and the reason the next section spends its best channel on its most important comparison.

The reason to teach the hierarchy in a chapter about analysis rather than communication is that it converts chart criticism into a prediction, which is what makes it verifiable. Before building a view, the analyst can state what comparison the reader must make, name the channel that comparison will be made through, and predict how accurately it will be made — and then look at the built chart and grade the prediction. That routine is the chapter's verification theme in miniature, it is what Section 12.12 asks of an AI-drafted chart before the chart is even displayed, and it is the first exercise in Section 12.18.

12.5 Choosing Charts by Analytic Purpose

Section 12.4 established which channels carry a comparison well. This section answers the question that actually confronts an analyst at the keyboard, which is not "which channel is most accurate" in the abstract but "what am I trying to find out, and what does that make the right picture." The move is to select from purposes rather than from chart types, and the reason is that a gallery of chart types answers a question nobody has — what charts exist — while a purpose list answers the question everybody has, which is what do I do next.

Definition

Analytic Purpose

An analytic purpose is the class of question a view is built to answer, stated as the comparison or structure the reader must extract rather than as the chart form that displays it. Six purposes cover nearly all marketing analysis: distribution (how are values of one variable spread), comparison (how do values differ across categories), trend (how does a value move over time), relationship (how do two variables move together), composition (how does a total divide into parts), and geographic (how does a value vary across places). Purpose determines which encoding channels must carry which fields, and therefore determines the chart form; selecting the form first inverts the dependency and produces charts that display data without answering anything.

Source: Adapted from Shneiderman (1996) and Munzner (2014).

In other words, an analyst who opens Tableau thinking "I'll make a bar chart of this" has already skipped the only decision that matters. Amar et al. (2005) cataloged the low-level tasks people actually perform with visualizations — retrieve a value, find an extremum, sort, characterize a distribution, find anomalies, correlate — and the catalog reads as a decomposition of the six purposes into the operations each supports. The practical form of the catalog is Table 12.1, which is written to be used at the moment of building rather than read once. Its rightmost column matters as much as the others: knowing which chart form is wrong for a purpose is what stops an analyst from producing the pie chart in the circulating deck.

Table 12.1

Choosing a Chart by Analytic Purpose

Purpose

The question it answers

A StyleCraft instance from Chapters 4–10

Default form

Common wrong form and why

Distribution

How are the values of one variable spread — center, spread, shape, outliers?

Are suburban order values simply higher, or higher and wider than urban ones (Chapters 4 and 5)?

Histogram; box plot for comparison across a few groups

A bar of the mean, which discards the shape that was the question and hides the bimodality Chapter 5 found

Comparison

How do values differ across categories, and which are largest?

In which store types do returning customers account for the largest share of orders (Chapter 5)?

Sorted horizontal bar

A pie, which replaces length judgments with angle judgments; an unsorted bar, which turns reading into searching

Trend

How does a value move over time, and does the movement break?

Does daily revenue step up at the Greenwich opening (Chapter 10)?

Line, with time on the horizontal axis

A bar chart of a time series, which discards the continuity the eye reads best; a line across categories, which invents continuity that does not exist

Relationship

How do two variables move together, and for which units does the pattern break?

Does app engagement accompany higher spend, and does the pattern hold in both designed segments (Chapter 7)?

Scatterplot, with a fit line where a model exists

A dual-axis time chart, which shows two series against time and is routinely misread as showing them against each other

Composition

How does a total divide into parts, and how does the division change?

How does category revenue split across metros, and is occasionwear's share really the suburban difference (Chapter 5)?

Stacked bar for a few parts; 100 percent stacked bar when shares are the question; treemap for many parts

A pie with more than about five slices, which nobody can rank; a stacked bar used to compare the middle segments, whose baselines float

Geographic

How does a value vary across places, and is location itself the explanation?

Where are the Wave 4 stores, and does the ramp pattern follow geography or store age (Chapters 5 and 9)?

Symbol map sized by value; filled map for rates over defined areas

A filled map of raw counts, which mostly draws a population map; a bubble map used to rank magnitudes precisely

Two rules of use travel with the table. The first is that the purpose is a property of the question, not of the data, and the same fields serve different purposes in different views. StyleCraft's line_revenue appears in this chapter as a trend (daily revenue across the openings), a comparison (revenue per store), a composition (category share by metro), and a geographic view (revenue by location), and all four are legitimate because all four answer different questions. What is not legitimate is producing all four and then deciding what the question was, which is the visual form of the multiple-comparisons problem Section 11.8 named and which Section 12.14 will meet again in a different costume.

The second rule is that a purpose can require more than one view, and analysis is where the extra views are cheap. Chapter 5's finding that suburban baskets are larger rather than more expensive is a composition question and a distribution question at once — the mix of occasionwear must be seen and so must the spread of units per order — and no single chart carries both. The analytical stage's advantage over the communication stage is precisely that it is allowed to answer a question with four pictures. Chapter 13 will have to pick one, and picking one is easier for an analyst who has already seen all four than for one who built only the chart she expected to need.

12.6 Scales and Honest Axes

Marks are drawn against axes, and an axis is a claim about what visual distance means. This section takes up the two decisions that claim involves — whether the scale is linear or logarithmic, and where the axis begins — and it treats them as analytical decisions with correct and incorrect answers for a stated question, rather than as the matters of preference they are usually treated as. Section 12.7 will then show what happens when the same decisions are made badly, which is why the honest treatment comes first: an analyst who cannot say why an axis is where it is has no basis for objecting to anyone else's.

Begin with the scale. On a linear axis, equal vertical distances represent equal differences, so the question the axis answers is "how much more." On a logarithmic axis, equal vertical distances represent equal ratios, so the question it answers is "how many times more" — or, for a series over time, "at what rate is this growing." Neither is more honest than the other; each answers a different question, and the error is using one to answer the other's question. One condition is not a matter of judgment: a logarithmic scale is available only when every plotted value is positive. Zero and negative values cannot be placed on a log axis in Tableau, so a series containing them requires a different transformation or a different chart (Tableau, 2026e).

StyleCraft supplies the case that makes the choice concrete. Plot weekly revenue for the SoHo store, mature and running near $7,900 a week, and for the Washington store, opened four months ago and running near $1,800 a week, on a single linear axis. SoHo's line dominates, Washington's line is a nearly flat trace along the bottom, and the fact that matters most to the December 2 decision — that Washington's weekly revenue has doubled since opening while SoHo's has moved 3 percent — is invisible, not because it is small but because the axis was built to answer a different question. Redraw the same two series on a logarithmic vertical axis and the comparison becomes possible: Washington's line rises visibly steeper than SoHo's, because on a logarithmic vertical scale against linear time the slope of a line represents proportional growth — a straight line indicates approximately constant exponential growth — and two series with very different levels can be compared on the quantity that is actually in dispute. The cost is real and must be stated. Log axes are unfamiliar to most executive readers, gridlines fall at 1, 2, 5, 10 rather than at even intervals, and a reader who does not notice the scale will badly underestimate the level difference. This is why the log axis belongs in an analytical chapter: it is often the right instrument for finding something out and frequently the wrong instrument for telling somebody, and a communication artifact built for a reader with ten minutes and no prior exposure will usually make the opposite call about the same pair of series, for good reasons.

Now the baseline, which is where more damage is done. The rule usually taught — "always start the y-axis at zero" — is wrong as a blanket rule and right for a specific and important class of charts, and the specific class is defined by the perceptual hierarchy rather than by convention.

Concept

The Zero-Baseline Obligation Follows the Encoding, Not the Chart

A bar encodes its value in the length of the mark, and length is read as a proportion: a bar twice as long is read as twice as much. That reading is only true if the bar's origin is zero, so a truncated bar chart does not merely exaggerate — it makes a false statement in the reader's most trusted channel. A line or a point encodes its value in position against a labeled axis, and position is read as a location rather than as a magnitude, so a line chart's axis may legitimately be restricted to the range where the variation lives, because the question a line answers is usually about movement rather than about level. The obligation, then, attaches to the encoding. An ordinary bar whose full length encodes magnitude must begin at zero. A floating or interval mark — a Gantt bar, a range bar, a waterfall component — measures its length from an explicitly defined reference rather than from zero, and must begin at that reference and say so, because the quantity it encodes is a duration, an interval, or a change rather than a level. For area marks the rule is stated on the area rather than on an axis that may not exist: zero must map to zero area. And any mark whose position carries the value requires a labeled axis whose range was chosen for the question and stated where a reader might be misled. And the analyst's duty in the second case is not a smaller version of the first — a line chart on a restricted axis is honest only if the restriction is visible and defensible, and "the software chose it" is not a defense, as Section 12.7 will show with the deck already circulating.

Table 12.2 turns the principle into the check that runs in Lab 12.2 and in Appendix E.

Table 12.2

Scale and Baseline Obligations by Chart Form

Form

What carries the value

Baseline obligation

Log scale appropriate?

What to state on the chart

Bar or column (magnitude)

Full length from a common origin

Zero baseline required

No — length from zero is meaningless on a log axis

Units and aggregation ("sum of revenue, $000")

Floating, interval, Gantt, or waterfall bar

Length measured from a defined reference

Must begin at the explicitly defined reference, which must be stated

No

The reference the length is measured from

Stacked bar or area

Length of each segment; total height

Zero baseline required

No

Whether segments are values or shares

Line

Position against the axis

Not required; range chosen for the question

Yes, when growth rates are the question

The axis range, and the scale type if not linear

Scatterplot

Position on two axes

Not required on either axis

Yes on either axis, when the relationship is multiplicative

Units on both axes; what one point is

Dot plot

Position against a category axis

Not required; label the value axis clearly

Yes

The value axis range

Bubble or symbol map

Area of the mark

Area proportional to value, with zero mapped to zero area

Rarely; only with a stated legend

A size legend with at least two labeled reference sizes

Index or percent-change view

Position relative to a base period

The base (100 or 0 percent) must be drawn as a reference line

Sometimes, for compounding growth

The base period, explicitly

One further point closes the section, because it is the one most often missed in Tableau specifically. Software chooses axis ranges to fill the available space, which means the default behavior of nearly every charting tool is to truncate. Tableau's default for a continuous measure on a bar chart is to include zero, which is correct; its default for a line chart is a range fitted to the data, which is defensible; and the moment a filter is applied, a measure is switched, or a quick table calculation is added, the range recomputes silently. An axis is therefore not a thing an analyst sets once. It is a thing an analyst checks every time the view changes, and the check takes two seconds — read the bottom number on the axis and ask whether a reader would guess it.

12.7 How Charts Mislead

Everything in Sections 12.3 through 12.6 was stated positively. This section states the same material as a catalog of failures, because the failures are what an analyst actually encounters — in a vendor's case study, in another team's deck, in an AI assistant's output, and, most instructively, in her own first drafts. Five mechanics account for most of the damage. Each is a specific misuse of a specific component of the grammar, each leaves a recognizable signature, and each has a repair. The mechanics are owned here; the professional obligation they create — that an analyst who can recognize them is the last line before a distorted chart ships — belongs to Section 12.16 and is not argued in this section.

The first mechanic is the one in the circulating deck's lead chart.

Definition

Axis Truncation

Axis truncation is the omission of the zero origin from an axis whose marks encode value by length or area, so that the visual proportion between marks no longer corresponds to the numerical proportion between the values. Because readers decode bar length as a ratio, truncation inflates apparent differences by a factor that depends on how far the axis floor sits above zero and that can be arbitrarily large; the distortion is invisible to a reader who does not inspect the axis, and it survives the presence of correct numerical labels. Truncation is a legitimate choice for position-encoded marks such as lines and points, where it restricts the view rather than misstating the proportion.

Source: Adapted from Pandey et al. (2015) and Correll et al. (2020).

Work the deck's chart arithmetically, because the size of the effect is the argument. Quarterly revenue in the suburban and resort stores moved from $241,000 to $252,000, an increase of 4.6 percent. Drawn on an axis running from $240,000 to $255,000, the first bar occupies about 7 percent of the plot's height and the second about 80 percent. State the effect in the two quantities that are actually comparable, and avoid any single "exaggeration factor" whose numerator and denominator are undeclared: the second bar is drawn roughly twelve times as tall as the first even though the underlying value is only 1.046 times as large, and a reader who converts the drawing into percentage growth sees an apparent increase of roughly 1,100 percent against an actual increase of about 4.6 percent. The misstatement is made in the channel the perceptual hierarchy says readers trust most, while the axis labels displayed are entirely correct. Pandey et al. (2015) measured exactly this effect experimentally and found that truncated axes substantially shifted readers' judgments of magnitude relative to identical data drawn from zero. Correll et al. (2020) added the finding that matters for practice: the distortion persists even when readers are shown the axis clearly and even when they are warned, which is why "the numbers are right there" is not a defense and why the repair has to be structural. The repair is to draw the bars from zero — at which point, as Figure 12.3 shows, the two bars look nearly identical, which is the honest picture of a 4.6 percent change — and, if the small change is genuinely the finding, to switch the encoding rather than the axis: a line chart of the eight quarters on a restricted, labeled axis, or a bar chart of quarter-over-quarter percentage change drawn from a zero reference line, both of which put the small movement on a positional channel where restricting the range is legitimate.

Figure 12.3

Axis Truncation Changes the Reading

The second mechanic is the deck's second chart, and it is subtler because it usually has no dishonest intention behind it at all.

Definition

Dual Axis

A dual-axis chart plots two measures against a shared horizontal axis using two independent vertical scales, one on each side. Because each scale's range can be set independently and the relative position of the two series depends entirely on that arbitrary choice, the apparent alignment, divergence, or crossing of the series is a property of the scaling rather than of the data; the same two series can be made to appear tightly coupled or unrelated by adjusting either range. Dual axes are defensible only when the two measures are genuinely related quantities whose scales are fixed by definition, and are otherwise best replaced by two vertically stacked panels sharing the horizontal axis, or by indexing both series to a common base.

Source: Adapted from Cairo (2019) and Few (2012).

The deck's chart puts monthly suburban revenue on the left axis and cumulative store count on the right, with both ranges chosen by the software to fill the frame, and the two lines run nearly on top of each other. The reader's inference — each opening produced the revenue that followed — is invited by the geometry and supported by nothing. Two objections are available and the analyst should have both ready. The first is Section 12.6's: the coincidence is manufactured by two independent range choices, and shifting the right axis by two stores separates the lines entirely without changing a single data point. The second is Part II's, and it is the more important one: even a real coincidence between the two series would be a correlation in observational data with a season and a national brand trend running underneath it, which is the confounding of Section 7.5 and the reason Chapter 11 exists. Figure 12.4 makes the dependence on the axis setting visible and shows the repaired form. The repair is to stack two panels sharing a time axis so that each series is read on its own honest scale, to add the urban series as the comparison group whose movement over the same period is what the suburban series must be judged against, and to leave the causal question to the design Section 11.9 sketched rather than to the drawing.

Figure 12.4

Dual Axes Manufacture Relationships

The third mechanic requires no misuse of any axis at all.

Cherry-picking is the selection of the displayed window — of time, of categories, or of units — so that the visible pattern differs from the pattern in the full data. The deck does not contain this one, but the version the analyst will meet before December certainly does: an eight-week window beginning at the Greenwich opening, over which suburban revenue rose 31 percent. Every number in that chart is correct. What the window omits is that the eight weeks run from late October into the holiday peak that Section 10.3's decomposition attributes to seasonality, and that suburban revenue had also been rising through the six months before the opening. Extend the window to the full twenty-four months and the step at the opening is still there — Chapter 10 found it and priced it — but it is a modest step on a rising, strongly seasonal series rather than the takeoff the short window shows. The signature of the mechanic is a window whose start date has no stated justification, and the repair is a standing rule rather than a case-by-case judgment: state why the window starts where it starts, and show at least one full seasonal cycle whenever the series has a season, per Section 10.3.

The fourth mechanic is a failure of the encoding rather than of the scale.

Definition

Area Distortion

Area distortion is the misencoding of a quantity in the size of a two-dimensional mark, most commonly by scaling the mark's radius or side length in proportion to the value so that its area grows with the square of the value. A circle drawn with twice the radius covers four times the area and is read as far more than twice as much; the resulting exaggeration compounds with the general tendency of readers to underestimate area differences, so that size-encoded quantities are simultaneously overstated by the construction and imprecisely decoded by the eye. Correct practice scales area in proportion to value and supplies a size legend with labeled reference marks; precise magnitude comparisons should be moved to a length- or position-encoded view.

Source: Adapted from Cleveland and McGill (1984) and Tufte (2001).

StyleCraft's store map is where this arrives. Trailing-twelve-month revenue runs from roughly $412,000 at SoHo to roughly $26,000 at Washington, which has been open four months and has ramped from under $1,000 a week to $1,800, a ratio of about 16 to 1. Encoded correctly by area, SoHo's circle has about four times Washington's diameter, which already produces a mark large enough to cover several neighboring stores; encoded by radius, it would have 16 times the diameter and roughly 250 times the area, which is not a chart but a blot. The deeper point is that neither version supports the comparison anyway, because Section 12.4 ranks area low. The repair is a division of labor between views: use the map for what maps are uniquely good at, which is showing where things are and whether spatial adjacency explains anything, and put the magnitude comparison in a sorted bar chart beside it. An analyst who wants both from one picture will get neither.

The fifth mechanic can be dealt with briefly, because it has no legitimate use in analysis. Three-dimensional effects applied to two-dimensional data — extruded bars, tilted pies, perspective ribbons — add a depth channel that carries no information and actively corrupts the channels that do: perspective makes near marks larger than far marks of equal value, occlusion hides small marks behind large ones, and the reader has no reliable way to determine where a bar's top intersects the axis. Wainer (1984), writing a deliberately inverted guide to displaying data badly, cataloged the family; nothing since has rehabilitated it. The repair is to remove the effect, and the practice lesson is that a chart type offered by the software is not thereby endorsed by anyone.

Table 12.3 collects the five for use rather than for reading. As with Table 11.4, the symptom column is the one that gets used most, because in practice the analyst is looking at somebody else's finished chart with no access to the underlying view.

Table 12.3

The Five Standing Ways Charts Mislead

Mechanic

What is misused

Symptom in a finished chart

Repair

Axis truncation

The baseline of a length-encoded mark

Bars whose height ratio is wildly larger than their labels' ratio; an axis whose lowest gridline is a suspiciously round non-zero number

Draw length-encoded marks from zero; move a genuinely small difference to a line or a percent-change view with a labeled reference

Dual axes

Two independent scale ranges on one frame

Two series that track each other implausibly well; a right-hand axis whose range has no natural relationship to the left

Stack two panels on a shared time axis, or index both series to a common base; add the comparison series the causal reading requires

Cherry-picked window

The extent of the displayed data

A window with no stated start rationale; a period shorter than one seasonal cycle in a seasonal business

State the window's justification; show at least one full cycle; display the pre-period whenever an event is being credited

Area distortion

The mapping from value to mark size

Size differences that dwarf the value differences; a size legend with one entry or none

Scale area, not radius, in proportion to value; supply a two-point size legend; move precise comparisons to a bar chart

Three-dimensional effects

Depth as a decorative channel

Extruded bars, tilted pies, perspective; marks occluding one another

Remove the effect; there is no analytical case for it

One closing observation belongs here rather than in Section 12.16, because it is a fact about how these charts get made rather than a claim about anyone's character. Four of the five mechanics are available as software defaults or one-click options, and the fifth — the cherry-picked window — is what a date filter does when nobody chooses a date. That is the reason this catalog is taught to analysts rather than to executives. The people in a position to catch these are the people whose hands are on the tool, and catching them requires knowing what to look at, which is a checklist, which is Appendix E.

12.8 Color as an Encoding System

Section 12.7 cataloged the failures of the positional channels and the marks drawn against them; this section takes up the channel that catalog barely touched, and it is the one where the most damage is done quietly. Color is the channel most used and least designed. It arrives in every chart whether or not anyone chose it, it is the first thing a reader notices and the last thing a builder inspects, and it is the only encoding channel in common use that can render a chart unreadable for a substantial share of the audience without anyone in the room being aware of it. This section treats color as what Section 12.3's grammar says it is — a set of visual channels to which fields are assigned — and derives the palette rules from the kind of field being assigned, which means from the levels of measurement introduced in Section 3.4.

Definition

Sequential, Diverging, and Categorical Palettes

A sequential palette varies a single hue through a monotonic progression of lightness or saturation and encodes a quantity with a meaningful low-to-high ordering, such as revenue or conversion rate. A diverging palette runs two contrasting hues outward from a neutral midpoint and encodes a quantity with a meaningful center — a target, a zero, a chain average — where distance and direction from that center are both the point. A categorical (or qualitative) palette assigns distinguishable hues of roughly equal lightness to unordered groups, and carries no magnitude information at all; beyond roughly seven or eight categories color alone usually becomes an unreliable identifier and should be supplemented with direct labels or replaced by a different channel — a practical working limit rather than a fixed perceptual boundary. Matching palette type to the field's level of measurement is a correctness requirement: a categorical field in a sequential palette asserts an order the data does not have, and a quantitative field in a categorical palette discards the order it does.

Source: Adapted from Harrower and Brewer (2003) and Munzner (2014).

In other words, the palette is chosen by asking what kind of variable is being colored, which is a question Chapter 3 already taught students to answer. Table 12.4 completes the mapping.

Table 12.4

Palette Type by Variable Role and Level of Measurement

Field being colored

Level of measurement (Section 3.4)

Palette

StyleCraft example

Failure if mismatched

Unordered groups

Nominal

Categorical, equal lightness, no more than 7–8 hues

home_metro; acquisition_channel; channel

A sequential ramp implies NYC Urban is "less" than Resort

Ordered groups

Ordinal

Sequential, light-to-dark within one hue

loyalty_tier (None → Insider → VIP → Backstage); age_band

A categorical palette discards the tier order the business cares about

Quantity with no meaningful center

Ratio or interval

Sequential

Revenue per store; units per order

A diverging palette invents a midpoint and splits the range into two false classes

Quantity measured against a target or a zero

Ratio or interval, with a reference

Diverging, midpoint set at the reference

Contribution margin versus the chain average; forecast error versus zero

A sequential ramp buries the sign change that is the whole finding

A binary flag used for emphasis

Nominal, two values

Two colors: one accent, one neutral gray

is_occasionwear; treatment versus control from Chapter 11

Two saturated hues make both groups shout and neither stand out

Figure 12.5

Color Is an Encoding System

Figure 12.5 shows the three palette types applied correctly and summarizes the accessibility checks each visualization must undergo. Three rules of construction do most of the practical work, and all three are violated by the fourth chart in the circulating deck.

The first is that color must not be the only channel carrying a distinction the reader needs. Roughly 8 percent of men and 0.5 percent of women of Northern European descent have some form of red–green color-vision deficiency, and the deficiency is invisible to everyone else in the meeting, including the person who chose the palette (Wong, 2011). The deck's red-to-green loyalty ramp fails twice over: it uses the two hues that some readers confuse most, and it uses hue alone to separate categories that a light-to-dark progression within a single hue would separate for everyone. The general repair is redundant encoding — carry the distinction in position, in ordering, in direct labels, in shape or line pattern, or in lightness as well as hue — and direct labels, position, and ordering are better safeguards than any choice of hue. The test runs in three parts, and the first part alone is a screening check rather than a certification. Convert the chart to grayscale: if the categories collapse, the chart depends on hue alone and must be rebuilt, and if they survive, the chart has passed one test and not all of them, because grayscale exposes reliance on hue and weak lightness contrast but does not simulate every form of color-vision deficiency. Then inspect the chart through a color-vision-deficiency simulator. Then confirm that color is not the only means by which any essential meaning is carried, which is the requirement accessibility guidance actually imposes, along with adequate non-text contrast for graphical elements. Contrast obeys the same logic. Light gray text on white, and pale categorical hues on a light background, fail for readers with ordinary vision on a projector in a lit room, which describes the December 2 meeting exactly.

The second rule is that color should be spent, not spread. A chart in which every category is a different saturated hue has used its most attention-grabbing channel to encode the least interesting fact, which is that the categories are different from one another — something the axis already said. The high-value use of color in an analytical chart is emphasis: one accent hue for the series under examination, gray for everything else, so that the eye goes where the question is. StyleCraft's seventeen-store weekly revenue chart is unreadable in seventeen hues and immediately readable with the three suburban stores in one accent color and the other fourteen in gray, and no information is lost, because the store identities live in the labels where they belong.

The third rule is that a palette applied to a quantity must be interpretable as a scale, which means it needs a legend with real values and a progression the eye reads monotonically. The classic violation is the rainbow ramp, still the default in some tools and still popular because it is colorful: it is not perceptually uniform, so equal steps in the data produce visually unequal steps; it introduces sharp apparent boundaries at yellow and cyan where the data has none; and it is unreadable in grayscale and under most color-vision deficiencies (Borland & Taylor, 2007). A single-hue sequential ramp answers every question the rainbow was supposed to answer and answers them correctly.

Finally, a note on brand palettes, because it is the question students ask first and the one that causes the most trouble in practice. StyleCraft has brand colors, and marketing charts are frequently expected to use them. Brand palettes are designed for identity rather than for encoding, so they typically contain three or four hues of similar lightness and no sequential ramp at all, which makes them adequate for a two-or-three-category chart and unusable for anything quantitative. The professional resolution is not to fight the brand and not to bend the data: use the brand hues for accent and identity, derive a sequential ramp by varying lightness within a single brand hue, and keep a neutral gray as the workhorse. Where a chart's encoding requirement and the brand guideline genuinely conflict, the encoding wins in an analytical chart and the conversation gets escalated for a communication one — a boundary an analyst has to be able to defend in front of an actual audience. What no palette can supply, however sensibly chosen, is something to read the values against, which is the subject of Section 12.9.

12.9 Reference Points, Baselines, and Annotation

A chart with a correct encoding, an honest axis, and a defensible palette can still be uninterpretable, and the reason has nothing to do with any of them. A number displayed alone answers "how much" and leaves the reader with the question she actually has, which is "compared to what." This section supplies the answer on the canvas, and it is the direct visual descendant of a discipline this guide installed in Chapter 2: the baseline as the thing a result must be read against, introduced in Section 2.7 as an interpretive obligation, operationalized in Section 8.7 as the model a new model must beat, extended in Section 10.5 to the naive and seasonal-naive forecasts, and rendered here in ink.

Definition

Reference Line

A reference line is a line, band, or marker drawn at a separately defined benchmark — external or data-derived — against which the plotted marks are interpreted. The benchmark may be a target, an average, a median, a prior period, a break-even level, a threshold, or a dated event. Its function is to supply the comparison the reader would otherwise have to supply from memory, converting an absolute display into a relative judgment; a reference band performs the same role for a range rather than a point. References come in two kinds, and the discipline differs between them. An externally supplied decision threshold — a target, a break-even level, a committed plan number — must be declared before the results are inspected, because a threshold chosen after the fact is not a benchmark but a description. A descriptive reference such as an average, a median, or a prior period may legitimately be computed from the data being plotted, provided it is labeled as such so that a reader knows the line is a summary of the marks rather than an independent standard. In both cases the reference must be named on the chart, because the choice of reference determines what the chart appears to say.

Source: Adapted from Few (2012) and Munzner (2014).

StyleCraft's expansion story needs four of them, and naming them is a useful demonstration of how much interpretive work the ink does. A bar chart of average order value by home metro needs a reference line at the chain average, because "$156 for NYC Suburban customers" means nothing until the reader knows the chain runs near $108. A line chart of daily revenue needs a dated vertical reference at the Greenwich opening, because the analytical question is whether the series stepped at that date and the eye cannot locate a date on a two-year axis unaided — this is the structural break of Section 10.4, drawn. A bar chart of contribution by store type needs a reference at the level that covers a store's fixed footprint, because the question the real-estate committee is actually asking is whether these stores clear their own cost. And a chart of repeat-purchase rate by store cohort needs a reference at the rate the urban cohort reached at the same age, which is the comparison-group logic of Section 11.9 doing its work in a picture rather than in an estimator.

Concept

A Chart Without a Reference Is a Number Without a Baseline

Section 2.7 established that a result quoted without a baseline is not evidence but a quantity, and every chapter since has honored the rule in prose: the model beat a stated baseline, the forecast beat seasonal-naive, the treatment arm beat a control that represented the real alternative. The visual form of the same rule is that a chart which displays values and supplies nothing to read them against has offloaded the interpretive work onto a reader who will do it from memory, from expectation, or from whatever the last chart showed — and will do it badly and invisibly. The working rule this chapter adopts, and Appendix E checks: before a view leaves the analyst's screen, name the comparison the reader must make and confirm that the thing being compared against is on the chart. If the comparison is against a target, draw the target. If it is against last year, plot last year. If it is against the other stores, show the other stores or the average of them. And if the honest answer is that there is nothing to compare against — which happens, and is worth saying — then the chart is a description and its title should not contain a word like strong, weak, up, or behind, because all four are comparative claims that the chart has not earned.

Annotation is the second instrument, and it does a different job.

Definition

Annotation

An annotation is explanatory text placed on the chart itself — a label on a specific mark, a note on a region, a caption on an axis, or a statement of the view's grain, filters, and source — that supplies context the encoding cannot carry. Annotation is how a chart records what its builder knows and its reader does not: why a point is anomalous, what happened on a marked date, which units are excluded, and at what grain the values are aggregated. Because it travels with the image, annotation is also the mechanism by which an analytical chart remains interpretable after it is separated from the analyst who built it.

Source: Adapted from Tufte (2001) and Few (2012).

Two habits make annotation useful rather than decorative. The first is to annotate the exception rather than the rule: the reader can see that the series rises, and cannot see that the November spike sits inside a promotional window recorded in the campaigns table, or that the small step in a March is the near-duplicate re-export that Chapter 4's defect catalog documents and that this extract removed. Those are exactly the facts that belong in ink on the chart, and they are exactly the facts that live only in the analyst's head. The second is the footer discipline, which is unglamorous and is the single highest-return habit in this chapter: every analytical view carries, in small type, the source table, the date range, the filters applied, the grain of aggregation, and the date the extract was refreshed. Five lines of text, added once to a template, prevent the most common failure of the pre-read genre — a chart that circulates for three weeks while nobody can determine whether it includes online orders.

One boundary should be marked before the section closes, since both this chapter and the next put words on charts: annotation supplies the context a reader needs in order to interpret correctly, while the titling discipline Chapter 13 owns supplies the conclusion the analyst wants her to take away, and the pre-read due Monday carries the first and deliberately withholds the second, for the reason Section 12.2 gave. With the canvas now carrying its own context, the chapter can turn to the tool that will draw it.

12.10 The Tableau Data Model

The chapter has so far described charts in a vocabulary that belongs to no particular software. This section takes up Tableau specifically, and the argument it makes is that Tableau's central abstraction is closely analogous to a distinction students already own from Chapter 3 — which is why an analyst who understood Section 3.5 will find Tableau's behavior largely predictable, and one who did not will find it arbitrary. The analogy is close but not an identity: dimensions and measures are software roles governing how Tableau builds a query, not the analytic variable roles of Section 3.5, and the two can come apart. A field can be converted from measure to dimension or back, and role is orthogonal to whether a field is discrete or continuous, so any field may occupy one of four states — discrete dimension, continuous dimension, discrete measure, or continuous measure (Tableau, 2026a).

Definition

Dimension and Measure

In Tableau's data model, a dimension is a field used to partition, group, or slice the data — it defines the categories, dates, and identifiers at which values are separated — while a measure is a field that is aggregated within those partitions to produce a number. The distinction is closely analogous to the variable roles of Section 3.5: dimensions are typically the categorical and identifying fields that define the units and groups of an analysis, and measures are typically the quantitative fields from which metrics are computed. It remains a software role rather than an exact equivalent — dimensions affect the level of detail of the view, measures are aggregated by default, and a field's role can be changed without changing what the field means (Tableau, 2026a). Because Tableau aggregates every measure to the level defined by the dimensions present in the view, the dimensions in a view determine what each mark represents, and adding or removing one changes the question the chart answers.

Source: Adapted from Murray (2016) and the variable-role framework of Section 3.5.

Two further distinctions govern how fields behave once placed, and they are the ones that confuse new users because they are orthogonal to the first.

Definition

Discrete and Continuous Fields

A discrete field takes a finite set of separate values and produces headers — labeled panels, categories, or bins along an axis — while a continuous field takes values along an unbroken range and produces an axis. In Tableau the distinction is displayed by color, with discrete fields shown in blue and continuous fields in green, and it is independent of whether a field is a dimension or a measure: a date can be treated discretely as a set of month labels or continuously as a timeline, and a measure can be made discrete to produce labeled panels rather than an axis. The choice determines whether the chart's marks are separated into panels or positioned along a scale, which is a decision about the chart's grammar rather than about the data.

Source: Adapted from Murray (2016).

Definition

Aggregation in the View

Aggregation in the view is Tableau's behavior of computing each measure at the level of detail defined by the dimensions currently placed in the view, so that the same underlying field yields different numbers as dimensions are added or removed. Because the aggregation is performed silently and recomputed on every change, a view's values are a property of its dimension set rather than of the measure alone; a metric whose definition requires a specific grain — an average per order, per customer, or per store — will be computed at the wrong grain, without warning or error, whenever the view's dimensions do not match that grain.

Source: Adapted from Murray (2016) and the grain discipline of Section 3.3.

The last of those three is where the chapter's most expensive error lives, and StyleCraft's own headline finding demonstrates it with unusual clarity. transactions_enriched is a flat file at order-line grain: one row per line of one order. Drag line_revenue into a view, set the aggregation to Average, split by home_metro, and Tableau returns the average line value — roughly $49.60 for NYC Urban customers and $51.10 for NYC Suburban, a difference of about 3 percent. That number is computed correctly, and it is not average order value. Average order value requires the order grain, which in this file means the sum of revenue divided by the count of distinct orders: SUM([line_revenue]) / COUNTD([order_id]), which returns roughly $94 for urban and $156 for suburban, a difference of 66 percent. The entire suburban finding — the one that motivates the December 2 meeting, that Chapter 5 decomposed and Chapter 7 modeled — vanishes at the wrong grain and reappears at the right one, and nothing in the software distinguishes the two cases. There is no error message, no warning, and no visual difference between the correct chart and the incorrect one. The only defense is the one Section 3.3 installed: state the grain the metric requires before building the view, and check that the view's dimensions and aggregation produce it.

The practical shape of that defense in Tableau is worth stating as a routine, because it is short. Before building any view, write down what one mark is supposed to represent — one store, one month, one customer, one order. Then look at what the view's dimensions actually make one mark represent. If the two differ, either change the dimensions or move the aggregation upstream into Tableau Prep, which is what Lab 12.1 does and which is generally the better answer for a metric that will be used repeatedly. Calculated fields can express the correction inside the workbook, as the average-order-value expression above does, and this chapter uses them only at that introductory level; the row-level-versus-aggregate distinction that governs when such expressions succeed, and the level-of-detail expressions that solve the general version of the grain problem, belong to the advanced visual-analytics work that follows this guide.

12.11 From Prep to First Charts

The concepts are in place and the tool's vocabulary is mapped. This section describes the path from the certified file to the first views — what Tableau Prep is for, what belongs in it rather than in the workbook, how the core chart forms are constructed, and what faceting buys — and it is deliberately short on click paths, because the click paths belong in the labs of Section 12.13 where a student can follow them at the keyboard.

Tableau Prep's role in this course is narrow and specific: it is where the data is brought to the grain and shape the analysis requires, and where the decisions that would otherwise be repeated in every workbook are made once. transactions_enriched arrives already denormalized — the joins to customers, products, and stores that Chapter 4 taught have been performed, which is precisely what makes it a Tableau file rather than a modeling file — so Prep is not being used here for joining. It is being used for three things. First, inspection: Prep's profile pane displays, for every field, the distribution of values and the count of nulls, which is Chapter 4's inspection routine rendered as an interface and which should be read the same way, looking for the categorical variants, the impossible dates, and the nulls that the defect catalog taught students to expect. Second, derivation: fields the analysis needs repeatedly — a month or week truncation of order_date with the week start declared explicitly, a store wave taken from the store roster rather than grouped by hand, a store-age-in-weeks field that Lab 12.1 Part C will need — belong here rather than in each workbook, because a definition that lives in one place can be corrected in one place. Each of those derivations carries its own inspection: check for negative store ages, for stores with a missing opening_date, for transactions dated before their store opened, and for how the online location is treated, since ONLINE has no opening date and therefore no store age at all. Third, and most importantly, aggregation to a stated grain, which is the answer to Section 12.10's trap: an aggregate step that rolls order lines up to orders produces a file on which average order value is simply an average, and on which no future analyst can compute it wrongly.

That last move deserves a caution, because it is a real trade-off rather than a free improvement. Aggregating to order grain discards product_id, category, and is_occasionwear, which are line-level facts, so the composition question of Section 12.5 cannot be answered on the aggregated file. The professional answer is not to choose but to output both: a line-grain extract for composition and product questions, and an order-grain extract for order-level metrics, with a documented reconciliation between them. The reconciliation is the verification, and it is exactly one number — total revenue must agree to the cent across the two outputs, because aggregation redistributes rows and must not create or destroy dollars.

Construction of the core forms in the workbook then follows the grammar directly, and the mapping is worth stating once so that the labs can move quickly. A field on the Columns shelf and a field on Rows establish the positional encodings; the Marks card sets the mark type and receives the non-positional encodings, with fields dropped onto Color, Size, Label, Detail, and Tooltip; and the aggregation of any measure is set on the field itself. A sorted bar chart is a dimension on Rows, a measure on Columns, and a sort applied to the measure. A trend line is a continuous date on Columns, a measure on Rows, and a Line mark. A scatterplot is two measures on the two shelves with the unit of analysis on Detail — and that last element is the one beginners omit, which is why their scatterplot shows one point: without a dimension on Detail, Tableau aggregates everything into a single mark, which is Section 12.10's lesson arriving in its most visible form.

The last construction the chapter owns is the one that will do the most analytical work in Lab 12.1.

Definition

Small Multiples and Faceting

Small multiples are a series of small charts sharing identical axes, scales, and encodings, each showing the same measures for a different subset of the data, arranged in a grid so that the subsets can be compared by scanning. Faceting is the operation that produces them: partitioning the data by a dimension and drawing one panel per value of that dimension. Because every panel shares a scale, differences between panels are read as differences in the data rather than as differences in the drawing; because the panels are separate, the comparison is made across identical non-aligned scales rather than against a common baseline, which costs some precision in exchange for the ability to compare many subsets at once.

Source: Adapted from Tufte (2001) and Munzner (2014).

In other words, faceting trades one rung of the perceptual hierarchy for the ability to see seventeen things at the same time, and in analysis that trade is very often correct. StyleCraft's case makes the argument by itself. Weekly revenue for all seventeen stores on one set of axes is a thicket in which nothing can be traced; the same data as seventeen panels sorted by opening date is a picture in which each store's ramp shape can be traced, and in which the apparent underperformance of the newest stores can be tested against the possibility that they are simply four months old. Putting that question where it can be answered is the analytical payoff of the whole chapter, it is what the December 2 meeting most needs to have seen before it starts, and no aggregate chart produces it — which is the general case for small multiples, and the reason they are the last thing this chapter builds before turning the tool over to the assistant.

12.12 AI as a Chart-Drafting Assistant

The division of labor this guide has refined since Chapter 7 — the assistant drafts mechanics, the analyst audits meaning — meets a new complication in Part III, and the complication is that a chart's mechanics and its meaning are produced by the same act. In Chapter 8 an assistant could write a model-fitting cell whose correctness was separable from the question of whether the model answered anything; the code either ran or did not, and the meaning was audited afterward against outputs the analyst could read. A chart has no such separation. The specification that says bar mark, revenue on length, quarter on the category axis, vertical range 240,000 to 255,000 is simultaneously the mechanics and the claim, and the assistant that produced it made a claim about what the reader should conclude in the same breath as it chose a mark type. There is no later stage at which the meaning gets audited unless the analyst inserts one, which is what this section installs.

The chapter's verification theme is that insertion, and it is unusually cheap to run. The assistant returns a chart specification or a set of build instructions; the analyst reads the specification, predicts against Sections 12.3 through 12.9 what the resulting picture will show and where it will mislead, then builds it and grades the prediction. Predicting first is what makes the audit possible, because a chart that has already been rendered is persuasive in a way its specification is not — it looks finished, it looks designed, and the human tendency is to evaluate whether it is attractive rather than whether it is right.

The failure modes are specific enough to be listed for the audit, and they follow directly from what an assistant is and is not. The truncated axis to order is the first and most consequential: an assistant asked to "make the difference clear" or "show the growth" will do exactly that, and the most reliable way to make a 4.6 percent difference clear is to restrict the axis, which the assistant will do without comment because the instruction was followed. The grain slip is the second and the costliest, and it is Section 11.11's failure arriving in visual form: asked for "average order value by metro" against a line-grain file, an assistant will very often return an average of line_revenue, which is a correct average of the wrong thing, and the resulting chart shows a 3 percent gap where a 66 percent gap exists. The default pie is the third: composition is the one purpose for which the assistant's training data is saturated with a bad answer, and a seven-category pie will be proposed for StyleCraft's product mix unless the prompt forbids it. The palette mismatch is the fourth, in both directions — an ordered field such as loyalty_tier given a categorical palette that discards the order, or an unordered field such as home_metro given a sequential ramp that invents one — and it is accompanied by the accessibility omission, since nothing in the request will have mentioned color-vision deficiency and nothing in the assistant's default behavior will raise it. The invented field is the fifth and the easiest to catch: assistants routinely reference columns that would be natural for the question and do not exist in the file, an aov or a customer_ltv or a region where the schema has home_region, and the failure surfaces as an error rather than as a wrong chart only if the analyst runs the build rather than trusting it. The wrong-purpose chart is the sixth: asked whether suburban stores outperform urban ones, an assistant will frequently return a time series, which is a trend answer to a comparison question and which will look entirely reasonable. The undisclosed filter is the seventh and the most dangerous in a pre-read, because it leaves no trace in the image: nulls dropped, ONLINE excluded as "not a store," a date range quietly restricted to complete months. The default sort is the eighth and the most trivially fixable, and it is left in place constantly. And the overclaiming title is the ninth: asked to title a chart, an assistant produces a fluent sentence containing a verb the view has not earned, and the verbs are the ones Section 11.10's ladder governs — drives, causes, proves — attached to a description that licenses none of them.

Where the assistant genuinely helps, the list is long, and it is longer here than in any previous chapter, because more of visualization's labor is mechanical than practitioners like to admit. Enumerating candidate forms for a stated purpose is a good use: asked for five ways to display the composition question with their trade-offs, an assistant produces a better slate than most analysts generate unaided, and the analyst's judgment is then spent on choosing rather than on remembering. Prep step logic is boilerplate, as are calculated-field expressions of the kind Section 12.10 used, and the assistant's syntax is generally correct where its grain judgment is not — which is a useful asymmetry to exploit deliberately, by specifying the grain in the prompt and delegating only the expression. Generating a color-blind-safe categorical palette as an explicit list of hex values, with the grayscale check described, is a task the assistant does quickly and well. Drafting the footer discipline of Section 12.9 — source, window, filters, grain, refresh date — is pure formatting. Writing alternative text for a chart, which almost nobody does and which accessibility increasingly requires, is a natural fit. And the adversarial use is once again where the assistant earns most: handed a description of a finished view and asked to list every way it could mislead a reader, it produces a checklist that will contain items the builder missed, at a cost of thirty seconds.

The governing instrument extends the audit lineage one more time. Table 8.5's five points do not all transfer, and saying which is part of the audit: frame becomes the question the view was built to answer, baseline becomes the reference the reader is meant to read the values against, and error in decision units becomes whether the view's quantities are in units a decision-maker can act on, while leakage and split hygiene have no visual analogue and are marked not applicable rather than skipped silently. Table 12.5 adds the three checks visualization makes necessary, and it is ordered so that the cheapest and most frequently violated check runs first.

Table 12.5

The Visualization Supplement to the Five-Point Audit (run with Table 8.5)

Supplemental point

The check

Fails when

V1. Grain and aggregation

State what one mark represents before looking at the view, then confirm the view's dimensions and the measure's aggregation produce exactly that; verify any per-order, per-customer, or per-store metric against a hand computation on one group

A per-order metric is computed from line-grain rows; a scatterplot shows one mark because no dimension is on Detail; totals change when an unrelated dimension is added

V2. Scale and baseline honesty

Read the axis floor on every length-encoded mark and confirm it is zero; read the range on every position-encoded axis and confirm the range was chosen rather than inherited; confirm any dual axis is justified or replaced; confirm the displayed window covers at least one seasonal cycle and that its start is defensible

A bar axis begins at a non-zero round number; two series track implausibly well on twin axes; a window begins at an event with no pre-period shown

V3. Encoding fits the purpose

Name the comparison the reader must make, name the channel carrying it, and confirm the channel's rank in Section 12.4 is appropriate; confirm palette type matches the field's level of measurement and survives a grayscale check; confirm the sort order was chosen; confirm the reference the reader needs is on the chart

Magnitudes are carried by angle or area; a categorical field is in a sequential ramp; bars are sorted alphabetically; a value is displayed with nothing to read it against

AI in Practice

Predict the Faults Before You Look

The chapter's verification theme, run as a five-step routine that takes about ten minutes and should be run on every AI-drafted view before it is shown to anyone. Step one is the request, written to expose rather than to conceal the assistant's assumptions: "I have transactions_enriched at order-line grain, with these columns [paste the schema]. I want to answer this question: do suburban and resort stores convert first-time buyers into repeat buyers at a lower rate than urban stores? Propose three candidate views. For each, state the mark type, every field-to-channel encoding, the aggregation and the grain one mark represents, the axis ranges, the sort order, and the palette. Do not use any column not in the schema I pasted; if the question requires a field I have not given you, say so instead of assuming it. Do not render or describe conclusions." The final two clauses do most of the work: the first defeats the invented field, the second stops the assistant from telling you what you are about to see. Step two is the prediction, made from the specification alone and written down before anything is built: for each candidate, name the comparison the reader will make, the channel it will be made through, the rank of that channel in Section 12.4, and the two faults you expect — then commit to which of the three candidates you expect to win and why. Step three is the build, in Tableau, following the specification exactly as written, including the faults; a specification corrected silently during construction cannot be audited. Step four is the audit and the grade: run the three transferable points of Table 8.5 and Table 12.5's three, line by line, on the built view, then compare the faults you found against the faults you predicted, and record both the misses and the false alarms, because a fault you predicted and did not find is as informative about your own model of the assistant as one you missed. Step five is the repair against Appendix E, one mechanic at a time, with a one-sentence note per repair recording what changed in the reading — not what changed in the chart. File both exchanges per Appendix D. And observe the rule that gives the routine its teeth: the assistant may propose any view it likes, but the analyst who circulates it owns every encoding in it, which is the visual form of signing the recommendation as the analyst of record.

12.13 Hands-On Application in Tableau

The preceding sections built the vocabulary; this section applies it to StyleCraft's own file in two labs that follow the pre-read's chronology rather than the chapter's section order. Lab 12.1 constructs: a Tableau Prep flow that certifies transactions_enriched at the two grains the analysis needs, six views built from the six analytic purposes of Section 12.5, and a faceted view that resolves the newest stores' apparent underperformance into a fact about store age. Lab 12.2 corrects: the AI round trip of Section 12.12, run on the deck circulating since Thursday, with each mechanic diagnosed against Table 12.3 and repaired against Appendix E.

One convention differs from Chapters 4 through 11 and is worth stating before the first box. Those chapters' labs presented runnable code cells; a Tableau lab's equivalent is a build path, so the boxed element here is labeled Build and contains the sequence of actions that produces the view, followed by the same plain-language account of input, transformation, and output that a code cell receives. Appendix A covers Colab mechanics and does not cover Tableau; students who have not opened Tableau Desktop or Tableau Prep before should work through the vendor's own Getting Started tour before this lab, and the Prep and Desktop versions assumed are those specified in the course technology requirements. AI assistants may draft any build path in this section, and every output is predicted before it is produced.

The labs use the certified transactions_enriched file, which is the denormalized join of transactions to customers, products, and stores described in the course data dictionary: one row per order line, carrying order_id, order_line_id, customer_id, product_id, store_id, channel, order_date, quantity, unit_price, discount_pct, line_revenue, line_cost, and campaign_id from the fact table, the customer attributes signup_date, acquisition_channel, home_metro, home_region, age_band, generation, gender, email_opt_in, sms_opt_in, app_user, and loyalty_tier, the product attributes category, subcategory, launch_date, and is_occasionwear, and the store attributes store_name, city, state, metro, store_type, opening_date, latitude, and longitude.

12.13.1 Lab 12.1, Part A: The Prep Flow and the Two Grains

Everything in the two labs depends on a fact established in Section 12.10 and easy to forget under deadline: transactions_enriched is at order-line grain, and half the questions the pre-read must answer are at order grain. This part builds the flow that produces both, and it verifies the flow with a single number.

Build

The certified flow in Tableau Prep

1. Open Tableau Prep Builder. Connect to Data → Text file → select transactions_enriched.csv. The input step appears; in the input pane, confirm the field types, and set order_date, signup_date, opening_date, and launch_date to Date.

2. Click the + on the input step → Clean Step. In the profile pane, read every field: the value distribution, the number of distinct values, and the null count. Record the row count shown in the step header.

3. Branch the flow before any date filter is applied, because one derived field must be computed from the full history. Click + on the clean step → Aggregate. Drag customer_id to Grouped Fields and order_date to Aggregated Fields, set to Minimum, and rename the result First Order Date. This branch reads every row in the file, including the partial months trimmed below and any purchase that precedes the displayed analysis window. Confirm that the step returns exactly one row per customer.

4. Return to the clean step and take the second branch, which is the analysis window. Right-click order_date → Filter → Range of Dates, and set the range to begin on the first day of the first complete month and end on the last day of the last complete month in the file. Record the new row count and the number of rows removed. The partial months at each end are trimmed because a monthly view built on them shows two artificial dips that are properties of the extract rather than of the business.

5. Still on the filtered branch, create the derived fields the analysis needs repeatedly. Create Calculated Field → name it Order Month, formula DATETRUNC('month', [order_date]). Create Order Week with DATETRUNC('week', [order_date], 'monday'). Create Store Age Weeks with DATEDIFF('week', [opening_date], [order_date], 'monday'). The week-start argument is not optional in this course: omitted, Tableau uses the data source's configured start of week, and two students on differently configured sources will produce different weekly bins and different store ages from the same file (Tableau, 2026c, 2026h). Both DATETRUNC and DATEDIFF accept the optional start_of_week argument in Prep as well as in Desktop.

6. Create Store Wave from the canonical store_wave field on the store roster. If your roster does not carry one, create it as an explicit calculated field that lists the store identifiers assigned to each wave — IF [store_id] IN ('SOHO','FLAT','BKLN') THEN 1 ELSEIF … END — rather than as a manual grouping on opening_date. A manual grouping cannot be audited from a screenshot and is easily reconstructed wrongly by the next person to open the flow.

7. Join the full-history branch to the filtered branch. Click + on the filtered clean step → Join; set the join type to Left, with the filtered line-level data on the left and the First Order Date table on the right, joined on customer_id. Read the join profile before continuing: the left row count must be unchanged, every filtered transaction must have matched exactly one customer, and no row may show a First Order Date later than its own order_date. Prep's aggregate step changes grain according to the fields you group by, and its join profile reports row inclusion for exactly this kind of check (Tableau, 2026b).

8. On the joined step create Is Repeat Order with IF [order_date] > [First Order Date] THEN 1 ELSE 0 END. Computing it here rather than in the workbook is the general principle of this step: a field whose correct value depends on a grain other than the view's belongs upstream. Note what the rule does with same-day orders. Because order_date carries a date rather than a timestamp, every order placed on a customer's first purchase date is classified as a first-day order rather than as a repeat. That is a defensible convention and it is declared here; it is not the same thing as identifying the first individual order, and any statement built on this field should say so.

9. Click the + on the joined step → Output. Name it lines_certified, set the output type to Tableau Hyper Extract, and note the path. This is the line-grain file.

10. Return to the joined step, click + → Aggregate. Drag to Grouped Fields: order_id, customer_id, order_date, Order Month, Order Week, Is Repeat Order, store_id, store_name, city, state, latitude, longitude, metro, store_type, channel, home_metro, generation, loyalty_tier, app_user, Store Age Weeks, Store Wave, First Order Date. Drag to Aggregated Fields: line_revenue (Sum), line_cost (Sum), quantity (Sum), and order_line_id (Count, renamed Lines In Order).

11. Click + on the aggregate step → Output. Name it orders_certified, again as a Hyper extract.

12. Run Flow. Record the row count of each output.

The input step reads the flat file; the first aggregate step computes each customer's first order date from the full history, before any window is imposed; the clean step performs the window filter and the derivations that would otherwise be repeated in every workbook; the join attaches the first-order date without changing the line count; and the second aggregate step collapses order lines into orders by grouping on every field that is constant within an order and summing the fields that are not. The order of those steps is the substantive lesson of the flow. Computing a customer's first purchase after trimming partial months would classify as a first purchase any order that merely happens to be the earliest one still visible, which misclassifies every customer whose true first purchase fell in a trimmed month or before the analysis window opens. The output is two certified extracts describing the same business at two grains — one on which product and category questions are answerable, one on which order-level metrics are simply averages. Note what the grouped-field list encodes: every field placed there is an assertion that the field does not vary within an order, and an assertion that turns out to be false will silently split one order into several rows, which is why the verification below counts orders as well as dollars. The store coordinates and the customer's first-order date are carried through the grouping for the same reason — the map of View F and the repeat-order metric of View B both need them at order grain, and a field that the workbook would have to reconstruct is a field the workbook can reconstruct wrongly.

Verification Check

Predict every number below before running the flow, then reconcile rather than spot-check. The row counts and totals quoted here — 97,912 input lines, 96,844 after the complete-months filter, 45,206 orders, and $4,882,247.60 of revenue — are the values the certified StyleCraft extract is built to return; confirm them against the answer key published in the companion repository before treating any of them as a target and report a mismatch as a finding rather than adjusting the flow until the numbers appear. The line-to-order ratio should come to 2.14 — predict it first from what Chapter 5 established about basket size and note that a ratio near 1.0 would mean the grouping keys split orders and a ratio far above 3 would mean orders were merged. Revenue reconciliation is necessary and not sufficient. Confirm all of the following across the two outputs: the row count of orders_certified equals COUNTD(order_id) in lines_certified; total revenue, total cost, and total quantity agree to the cent or to the unit; distinct customers and distinct stores agree; and the minimum and maximum order_date agree. Confirm separately that the join in step 7 preserved the line count exactly. Total revenue must agree to the cent because an aggregation redistributes rows and must never create or destroy dollars. Finally, compute average order value two ways on the two files — SUM([line_revenue]) / COUNTD([order_id]) on the line-grain extract, and a simple AVG([line_revenue]) on the order-grain extract — and confirm they return the same $108.00, which is the chain average this chapter's reference lines will use. Then compute AVG([line_revenue]) on the line-grain file, observe that it returns $50.41, and write down what a reader would conclude if that number were labeled "average order value." That single wrong label is the most common error in this chapter and the hardest to see, because the chart it produces is beautiful.

12.13.2 Lab 12.1, Part B: Six Purposes, Six Views

With both grains certified, the pre-read can be built. The discipline of this part is that each view is specified before it is constructed — purpose, comparison, encoding, grain, reference — and that the specification is written down where it can be graded. Table 12.6 is that specification for all six, and filling it in is the assignment; the build paths below construct three of the six, and the remaining three are the student's.

Table 12.6

The Expansion Story in Six Views

View

Purpose (Section 12.5)

The comparison the reader must make

Grain of one mark

Reference the view must carry

A. Order value distribution by metro

Distribution

Are suburban orders uniformly larger, or larger and more dispersed?

One order

The chain median, drawn

B. Repeat-order share by store type

Comparison

In which store types is the largest share of orders placed by customers who had already bought before?

One store type

The chain share for the same store population, drawn

C. Weekly revenue with the openings marked

Trend

Does the series step at the Greenwich opening, and does it step again later?

One week

Dated vertical marks at each opening; one full seasonal cycle visible

D. Items per order against occasionwear share, by store

Relationship

Do the stores with big baskets also sell the occasionwear mix, or are these two different stories?

One store

The chain averages on both axes, as crossed reference lines

E. Category revenue share by metro

Composition

Does the suburban category mix differ from the urban one, and by how much?

One category within one metro

None required; shares are self-referencing

F. Store locations with trailing revenue

Geographic

Where are the newest stores, and does geography explain the ramp?

One store

A labeled size legend with two reference sizes

Build

View B — repeat-order share by store type (comparison)

1. Connect a new workbook to orders_certified.

2. Confirm that Is Repeat Order arrived from the Prep flow of Part A. It is already correct at order grain, which is why it was built there.

3. Create Repeat-Order Share as SUM([Is Repeat Order]) / COUNTD([order_id]). Format it as a percentage with one decimal. Name it precisely, because the obvious name is wrong: this is the share of orders placed after the customer's first purchase date, subject to the same-day convention declared in Part A, and it is not the share of customers who returned. A view labeled "returning-customer share" invites a customer-level reading the metric does not support.

4. Drag store_type to Rows and Repeat-Order Share to Columns. The Marks card should read Bar; if it does not, set it. Exclude store_id = ONLINE on the Filters shelf, since the comparison is across physical store types.

5. Click the sort icon on the Repeat-Order Share axis to sort descending. Confirm the sort is applied and not merely apparent.

6. Right-click the horizontal axis → Edit Axis → confirm Include zero is checked and the range starts at 0.

7. Supply the chain reference deliberately, because a constant line requires a value and does not compute the pooled rate for you. Tableau's reference lines can be based on a constant or on a computed value, but the computation and its scope must be specified (Tableau, 2026d). Choose one of three routes and record which you used: compute the verified chain-wide share separately and enter it as a constant; precompute it in Prep and carry it into the extract; or create a parameter populated from the certified answer key. Analytics pane → drag Constant Line to the view → enter that value → label it "Chain average (physical stores)."

7a. Do not use the average of the four displayed store-type rates. The chain rate is computed on the pooled denominator and differs from the unweighted mean of the four rates, for the reason Chapter 5's denominator discussion gives. And because step 4 filters ONLINE out, the reference must be the pooled rate for the same physical-store population; if a comparison against the whole business including online is intended, draw a second, separately labeled line rather than redefining the first.

8. Add the footer: Worksheet → Show Title, and in the caption record the source extract, the date window, the filters, the grain ("one bar = one store type; share = orders placed after the customer's first purchase date ÷ distinct orders"), and the refresh date.

Input is the order-grain extract; the transformation is a ratio of two aggregations computed within each store type; the output is a sorted bar chart with a reference line. Three of the eight steps are encoding discipline rather than construction — the sort, the zero baseline, and the reference line — and they are the three that a five-minute version of this chart omits.

Build

View C — weekly revenue with the openings marked (trend)

1. New worksheet on orders_certified. Drag Order Week to Columns; right-click the pill and confirm it is continuous (green), not discrete.

2. Drag line_revenue to Rows with aggregation Sum. Set the Marks card to Line.

3. Drag metro to Color. State the emphasis hierarchy before choosing hues: NYC Suburban is the primary subject of the December 2 question and Resort is secondary. Assign neutral gray to the two metros not under examination, one saturated accent hue to NYC Suburban, and a muted variant of that accent — same hue, lower saturation — to Resort, so the pairing reads as primary and secondary rather than as two competing signals.

4. Analytics pane → Reference Line → Table → set Value to a constant at each opening date, or add an annotation: right-click the point on the axis at the Greenwich opening date → Annotate → Point, and label it "Greenwich opens."

5. Repeat the annotation for the Wave 4 opening dates, keeping the labels short.

6. Right-click the vertical axis → Edit Axis. Because the marks are lines and encode value by position, the axis range may be restricted; set it deliberately and record the choice in the caption rather than accepting the default.

7. Confirm the horizontal range spans the full twenty-four months, so that at least one complete seasonal cycle is visible per Section 10.3.

8. Add the footer as in View B.

Input is the order-grain extract; the transformation is a sum of revenue within week and metro; the output is a multi-series line chart with dated annotations. The color choice in step 3 is the emphasis discipline of Section 12.8 rather than a decoration: four saturated hues would make the chart about the existence of four metros, while two neutral grays plus one accent and its muted variant make it about the suburban series first and the resort series second, which is the hierarchy the question implies.

Build

View F — store locations with trailing revenue (geographic)

1. New worksheet on orders_certified. Confirm latitude and longitude carry geographic roles; if not, right-click each → Geographic Role → Latitude / Longitude. Either use Tableau's generated Latitude and Longitude fields or use the store roster's own coordinates.

2. Place the coordinates as continuous measures, which is how Tableau treats latitude and longitude in the standard proportional-symbol workflow and where its generated fields appear (Tableau, 2026f, 2026g): AVG(Longitude) on Columns and AVG(Latitude) on Rows. Drag store_name to Detail so that the average is taken within one store and one mark is one store.

3. Set the Marks card to Circle. Drag line_revenue to Size. Filter to the trailing twelve months using the anchored window defined below rather than a relative-date filter.

4. Open the quantitative Size legend → Edit Sizes → set Sizes vary to From zero, which is the mapping that makes circle area proportional to the value from an origin of zero (Tableau, 2026f). Retain the continuous size legend, and add a written note on the view stating what circle area represents. If the pre-read requires two fixed labeled reference circles, build them as a small separate legend worksheet; a standard quantitative size legend does not reliably provide two custom reference symbols.

5. Drag Store Wave to Color. Waves 1 through 4 are ordered, so a categorical palette would contradict Table 12.4. Use a sequential palette running from earlier to later waves, or — better for this question, which is about the newest stores rather than about the distinction among all four waves — neutral gray for Waves 1 through 3 and one accent hue for Wave 4. Then run the three-part accessibility check of Section 12.8.

6. Do not attempt to read magnitudes from this view. Place View B's bar chart beside it in the pre-read, per Section 12.7's division of labor.

Input is the order-grain extract with a trailing-twelve-month filter; the transformation is a sum of revenue per store; the output is a symbol map.

The trailing window is anchored to the dataset rather than to the clock, and this is a course-wide convention rather than a preference. A relative-date filter evaluates against the current system date, so a workbook built this term will silently return a different window next term and eventually no records at all from a static teaching extract. Define the end of the window as the maximum complete date in the certified extract and derive the start from it:

Analysis End Date = #2026-06-30#

Trailing Start Date = DATEADD('day', 1,

DATEADD('year', -1, [Analysis End Date]))

Two details in that expression are the point of it. The date is written as a Tableau date literal between hash marks, or supplied as a Date parameter, because a bare 2026-06-30 is not valid syntax in a calculated field. And the start is derived by stepping back one year and forward one day rather than back twelve months, because an inclusive filter anchored twelve months back would begin on June 30, 2025 and return twelve months plus a day. The window is therefore July 1, 2025 through June 30, 2026, and it remains reproducible in 2027 and beyond, so that any student who reruns the workbook gets the figures the answer key documents. Step 6 is the substantive instruction: the map's job is to answer whether location explains the ramp, and the honest answer will come from the faceted view of Part C rather than from the circles.

Verification Check

Predict each view's shape before building it, in one sentence apiece, and predict the specific numbers where Part II already supplies them. View A should show suburban orders both higher and substantially wider than urban ones, with the suburban distribution's right tail carrying the occasionwear baskets; predict the two medians before you see them. View B's repeat-order share should return roughly 61 percent for Urban, 33 percent for Flagship, 24 percent for Suburban, and 19 percent for Resort — predict the ordering first, from what Chapter 5 and Chapter 9 established. Then predict the chain reference line, and notice that it is not the average of those four numbers: the chain-wide share sits near 47 percent because urban stores carry most of the chain's orders, well above the unweighted mean of 34 percent. A student who predicts 34 and finds 47 has not found an error in the extract; she has found the difference between an average of rates and a rate computed on the pooled denominator, which Chapter 5's denominator discussion already named. View C should show the Greenwich step Chapter 10 already found, and it should show it as a modest step on a rising, strongly seasonal series rather than as a takeoff; if the step looks dramatic, check the axis range you set in step 6 before believing it. View D should show the two axes are correlated but not redundant, with at least two stores off the diagonal, and those stores are worth naming in the pre-read. View E should show occasionwear's share roughly three times higher in suburban than in urban metros while average line value differs by about 3 percent, which is the designed finding of Chapter 5 restated visually and the single most important slide in the packet. Predict that View F will show no geographic pattern that survives inspection — and hold the prediction as a prediction. If the certified data does show a spatial pattern, that is a result to investigate against store age and store type, not a defect in the view. Then run V1 from Table 12.5 across all six: for each view, state what one mark represents, and confirm it against the dimensions actually in the view.

12.13.3 Lab 12.1, Part C: Small Multiples and the Store-Age Artifact

The pre-read's hardest question is the one the real-estate director will ask first: the newest stores look weak, so why open eight more? Every aggregate view in Part B says the Wave 4 stores underperform, and every aggregate view is telling the truth about a comparison that should never have been made. This part builds the view that can settle it, and it is the chapter's demonstration that faceting is an analytical instrument rather than a layout preference.

Build

Seventeen panels, one scale

1. New worksheet on orders_certified. Drag Store Age Weeks to Columns (continuous) and line_revenue to Rows (Sum). Set the Marks card to Line.

2. Drag store_name to Rows, to the left of the measure. Tableau now draws one row of panels per store.

3. Right-click store_name → Sort → Field, sort by MIN(opening_date) ascending, so the panels run oldest to newest. Sort by the field rather than manually, so that the panel order is reproducible by anyone who opens the workbook.

4. Analysis menu → Table Layout → confirm the axes are not independent: every panel must share one vertical scale, or the comparison the small multiple exists to support is destroyed.

5. Set the panel size small enough that all seventeen are visible without scrolling; use Fit → Entire View and reduce the row height.

5a. Before the benchmark can be computed, build a complete store-week scaffold, because the order-grain extract contains no row for a week in which a store took no orders. Generate one row per physical store for every store-age week from 0 through 25, left-join weekly revenue onto it, and set revenue to zero for any store-week in which the store was open and observable but recorded no orders. Leave a store-week null — rather than zero — where the store was not yet open or the window had not yet closed, so that an absent observation is never read as a zero.

5b. Carry both horizontal variables on the scaffold, because the two panels of this lab differ only in which one is placed on Columns. Store Age Weeks is the scaffold's own key; derive the matching calendar week from it:

Order Week = DATEADD('week', [Store Age Weeks],

DATETRUNC('week', [opening_date], 'monday'))

Both displays are then built from exactly the same scaffold rows — store-age weeks 0 through 25 for every physical store — with only the field on Columns changing. That is not a convenience. The exhibit's whole claim is that nothing changes except the horizontal encoding, and it cannot be made honestly if the calendar-time panel silently draws on the full transaction history while the store-age panel draws on the first twenty-six weeks.

6. Add a reference line at the chain-average weekly revenue for stores in their first twenty-six weeks, so each new store's ramp is read against the ramp its predecessors had. Compute it on the scaffold, not on the view, and define it as

SUM(revenue for store-age weeks 0-25)

/ COUNT(scaffold store-weeks 0-25 that were open and observable)

so that every eligible store-week counts exactly once. Two failures are avoided here. Computing the average from the marks in the view would weight stores by how many order lines they generated rather than by the store-weeks the benchmark describes. And omitting the zero-order weeks — which is what the raw extract does silently — drops the weakest observations from the denominator and biases the benchmark upward, which is the direction that would make a new store's ramp look worse than it is.

6a. Build the small multiples themselves on the same scaffold. Without it, Tableau draws a line straight across an absent week, so a store that sold nothing for a fortnight is rendered as though the fortnight never happened.

7. Duplicate the sheet, and on the copy replace Store Age Weeks with Order Week on Columns. Change nothing else — same scaffold, same rows, same measure, same vertical scale, same panel order, same styling.

Input is the order-grain extract; the transformation is a sum of revenue by store and by store age; the output is seventeen small charts on one scale. Step 7 is the crux of the lab, and Figure 12.6 is the result: the same data, faceted the same way, plotted against two different horizontal variables. On calendar time, the Wave 4 stores are eighteen months of white space followed by a short low line, and they look like failures. On store age, they are laid over their predecessors' first months, and the comparison becomes the one the question actually requires — whatever that comparison then turns out to show.

Figure 12.6

Calendar Time Versus Store Age

Verification Check

Predict the two pictures before building either or predict them precisely enough to be wrong. On the calendar-time facet, the Wave 4 panels will be empty for most of the horizontal range and will end with a short line at a level far below SoHo's; predict what a reader who saw only this panel would recommend. On the store-age facet, the designed expectation is that most Wave 4 ramps will fall inside the band traced by the older stores' first twenty-six weeks — predict which one or two will not, because the honest version of this finding is that store age explains most of the gap rather than all of it. Treat that expectation as something to verify rather than as the answer: a different result is a finding to investigate and report, not an error to edit away. Then run the comparison the whole part exists for: name, in one sentence, what changed between the two views, and confirm that nothing changed except the horizontal encoding. No filter was applied, no row was removed, no measure was redefined. The recommendation a reader would make is flipped because of a choice about what one axis means, which is the most important sentence in this chapter and the one Section 12.16 will hold the analyst responsible for. Finally, connect whatever you find to what the course data dictionary already documents about storage and marketing exposure — the confound Chapter 9 met as a fairness problem in a model, and which this view meets as an encoding problem in a chart. If the facet confirms the expectation, it is the same artifact caught by a different instrument; if it does not, you have found something the data dictionary does not yet describe, and that is the more valuable result.

12.13.4 Lab 12.2, Part A: The AI-Drafted Chart, Predicted and Diagnosed

The constructive lab is complete; the corrective one begins with the deck that has been circulating since Thursday. Run the five-step routine of Section 12.12's AI in Practice box end to end, documenting each move per Appendix D.

Begin with the request, and make it the bad request on purpose, because reproducing the failure is the point: ask an assistant to "make a chart showing how much the suburban stores have grown this quarter and make the growth clear." Do not paste the schema. Do not state the grain. Ask for a build path for Tableau.

Then stop and predict. Before reading what came back, write down five predictions: which mechanic from Table 12.3 the assistant will use to satisfy "make the growth clear"; whether it will compute a per-order or a per-line quantity; what it will do about the sort; whether it will name a column the file does not contain; and what verb its suggested title will use. Then read the response and grade the predictions.

Build

Reproducing the circulating chart, faults included

1. New worksheet on lines_certified. Drag order_date to Columns and set it to discrete Quarter. Filter to the two most recent quarters.

2. Filter store_type to Suburban and Resort.

3. Drag line_revenue to Rows with aggregation Sum. Marks card: Bar.

4. Right-click the vertical axis → Edit Axis → uncheck Include zero → set the range from 240,000 to 255,000.

5. Add a second sheet: Order Month on Columns, line_revenue (Sum) on Rows as a line, then drag a running count of open stores to Rows and right-click → Dual Axis; do not synchronize the axes.

6. Add a third sheet: category on Color, line_revenue on Angle, Marks card set to Pie.

7. Add a fourth sheet: loyalty_tier on Color with the red-to-green diverging palette, line_revenue on Rows as bars.

8. Do not repair anything yet. Screenshot all four.

Input is the line-grain extract; the transformation is a set of sums; the output is the deck as circulated, rebuilt so that its mechanics can be inspected from the inside rather than argued about from the image. Building the faults deliberately is what makes them legible: an analyst who has set an axis floor to 240,000 by hand will never again fail to notice one.

Verification Check

Diagnose all four against Table 12.3 and Table 12.5 before reading further and quantify each rather than merely naming it. For the first chart, compute the ratio of the two bars' drawn heights — with an axis from 240,000 to 255,000, the first bar occupies about 7 percent of the plot height and the second about 80 percent, a drawn ratio near twelve to one against a true value ratio of 1.046 to one — and then state the effect the way Section 12.7 states it, as a drawing that suggests roughly 1,100 percent growth where the data shows about 4.6 percent. Do not report a single exaggeration factor unless you also declare which two quantities you divided, because three different ratios are available here and they differ by more than an order of magnitude. For the second, change the right axis's range by two stores and observe that the coupling disappears; then state the second, more serious objection, which is the one from Section 7.5 and has nothing to do with drawing. For the third, cover the labels and try to rank the middle four categories; then rebuild the same data as a sorted bar and try again. For the fourth, export the image to grayscale and count how many of the four loyalty tiers remain distinguishable. Then run V1 across all four sheets and find the fifth fault the deck's readers have not noticed. Line grain is not itself an error: three of the four charts display sums of revenue, for which the line-grain extract is entirely valid. The fault is the mismatch between the extract and a claim made about it — the covering email calls the loyalty-tier chart "average order value," and the worksheet computes AVG(line_revenue), which is average line value. State the rule that follows, because it is the one students most often get backward: a chart built on line-grain data is wrong only when the chart or the text around it makes a per-order, per-customer, or per-store claim that the grain does not support.

12.13.5 Lab 12.2, Part B: The Repair, and What Changed in the Reading

The last part rebuilds the four charts against the Appendix E checklist, and it is graded on the notes rather than on the charts.

Build

The four repairs

1. Chart one: re-point the sheet at orders_certified. Right-click the axis → Edit Axis → check Include zero. Observe that the two bars are now nearly identical. Then decide whether the finding survives its honest rendering: if the 4.6 percent change is genuinely the point, duplicate the sheet, extend the filter to all eight quarters, switch the Marks card to Line, and add a reference line at the prior-year same-quarter value.

2. Chart two: delete the dual axis. Build two sheets — suburban and resort revenue by month; cumulative store count by month — and place them in a vertical arrangement sharing the horizontal axis. Add urban revenue to the first sheet in gray as the comparison series.

3. Chart three: change the Marks card from Pie to Bar, move category from Color to Rows, move line_revenue from Angle to Columns, sort descending, and confirm the axis includes zero. Then duplicate and convert to a 100 percent stacked bar split by metro, which is the view that answers the composition question the pie was pretending to answer.

4. Chart four: replace the red-to-green palette with a single-hue sequential ramp ordered None → Insider → VIP → Backstage. Run the grayscale check. Add the footer discipline of Section 12.9 to all four sheets.

Input is the two certified extracts; the transformation is a sequence of encoding changes with no change to any underlying query; the output is four charts that say what the data says. Note that step 1 contains the chapter's most uncomfortable moment, and it should not be skipped: the honest chart shows almost nothing, and the analyst must then decide whether there is a finding at all. That decision is analysis, and it is what the truncated axis was substituting for.

Verification Check

For each repair, write one sentence recording what changed in the reading rather than what changed in the chart — "a reader now sees a quarter that moved slightly, rather than a quarter that appeared to jump twelvefold" is the form. Then write the note that actually gets sent. The merchandising analyst who built the original deck is a colleague under the same deadline who used the tool the way it presented itself, and the memo that lands well is the one that names the mechanic rather than the person, shows the before and after side by side, and states the general rule so that the next chart is built correctly rather than corrected later. Draft it in four sentences and confirm before sending that it contains no sentence beginning with "You should have." Finally, count what the whole lab cost: four repairs, roughly thirty minutes, against a deck that had already reached the director of real estate. That ratio — minutes of audit against the credibility of a December meeting — is the argument for making the audit routine, and it is the argument Section 12.15 will find running through every organization that has ever adopted a chart standard.

12.14 Marketing Interpretation and Managerial Insight

The labs produced two certified extracts, six views, a faceted resolution of the store-age question, and four repaired charts. The pre-read circulates Monday, and what circulates is not the charts, but the sentences readers will form while looking at them. This section translates, and as in every chapter of Part II it does so partly by exhibiting the wrong readings, because analytical charts attract their own family of misinterpretation — one that differs from the misreadings of models and forecasts in a specific way. A wrong reading of a regression coefficient requires the reader to know what a coefficient is. A wrong reading of a chart requires nothing at all, arrives instantly, and is held with the confidence of something seen rather than something inferred.

The first wrong reading is the one this chapter's outline predicted, and it will be spoken by whoever opens the circulating deck first: "Suburban revenue is surging — look at the jump last quarter." The reader has decoded a length ratio of roughly twelve to one, which is what the chart drew, and has translated it into the word surging, which is what a twelve-to-one ratio would justify. The underlying movement is 4.6 percent, from $241,000 to $252,000, which on a quarterly series with StyleCraft's seasonality is inside the range the last eight quarters have wandered through without anyone using a verb. The correction is not a hedge and should not be delivered as one: suburban and resort revenue rose 4.6 percent quarter over quarter, which is within the ordinary quarter-to-quarter variation of the series and is smaller than the holiday effect Section 10.3 already decomposed, and the chart that made it look otherwise had an axis floor of $240,000. Two features of that sentence are worth copying. It gives the number and the comparison in the same breath, and it names the mechanic rather than merely asserting that the impression was wrong — because the reader who was misled will believe a correction that shows her how she was misled and will resist one that simply contradicts what she saw.

The second wrong reading is the expensive one, and it does not come from a broken chart at all. "The Wave 4 stores are underperforming; we should stop the expansion." Every aggregate view supports it. Trailing-twelve-month revenue by store puts the eight newest at the bottom; the calendar-time facet shows them as short low lines; the map shows small circles in Dallas, Scottsdale, and Washington. And the comparison being made — eight stores that have been open four months against nine stores that have been open one to two years — is one that no analyst would make in a table and that the chart makes irresistible, because the eye compares whatever is placed side by side on a shared scale and does not ask how old anything is. The correction is the store-age facet of Lab 12.1 Part C, and the way to deliver it in a meeting is to show both pictures in sequence and let the room test its own conclusion against the second one. The general lesson is worth naming because it will recur in every visual comparison anyone builds: a chart's implicit claim is that the things it places on a common scale are comparable, and that claim is made by the layout rather than by the data. An analyst who would never run a regression without asking what the units have in common should not build a bar chart without asking the same question.

The third wrong reading belongs to the map, and it is the most seductive because geography feels explanatory. "The West Coast stores are our growth engine — look at Beverly Hills and Scottsdale." Two errors are stacked. The first is Section 12.7's: the reader is ranking magnitudes by circle area, a channel the perceptual hierarchy places near the bottom, and a difference she reads as large may be a difference of thirty percent or of three hundred. The second is subtler and is the reason maps are so often the wrong instrument in retail analytics: a map of a per-store quantity mostly draws where the stores are, and StyleCraft's stores are where they are because of a real-estate strategy rather than because of customer demand. The map answers "where" and cannot answer "why," and every use of it to answer "why" is the ecological reasoning Section 3.3.1 warned about arriving in cartographic form. The correction is the division of labor Lab 12.1 built into the pre-read: the map establishes location and adjacency, and the sorted bar beside it carries every magnitude claim anyone intends to make.

The pre-read that survives all three has a fixed anatomy, and its ordering is an argument in the same way Chapter 11's deliverable ordering was. Page one is the reading guide: one paragraph naming the question each of the six views answers, the grain of one mark in each, and the window all six share, so that a reader who disagrees knows exactly what she is disagreeing with. Page two is the composition finding, because it is the one that changes minds — occasionwear share roughly three times higher in suburban metros, with average line value differing by about 3 percent, which together establish that StyleCraft's suburban advantage is basket construction rather than pricing power and therefore that the advantage is a merchandising fact rather than a market fact. Page three is the distribution and comparison pair: suburban orders larger and wider, suburban repeat-order share less than half the urban share, both against drawn chain averages, which is the tension the December meeting exists to resolve. Page four is the trend with the openings annotated and the full seasonal cycle visible, which prevents the cherry-picked reading before anyone constructs it. Page five is the pair of faceted views from Lab 12.1 Part C, presented together on one page rather than separated, with the calendar-time facet on the left and the store-age facet on the right and a single line between them stating that nothing changed except the horizontal encoding; presenting the pair rather than the conclusion is what keeps the page an analytical exhibit, and the question of how a finding should be ordered for a reader who has ten minutes rather than a week belongs to Chapter 13. Page six is the map, captioned to say what it does not establish. And an appendix carries the charts that were built and rejected, with a line apiece on why, which is the visual counterpart of Chapter 11's exploratory section and serves the same function: it makes the analyst's search visible, so that nobody has to wonder whether these six views are the six that worked.

One translation deserves its own paragraph, because it is what makes this chapter's deliverable different in kind from Part II's. Every artifact Chapters 8 through 11 produced was valuable for the number inside it. This chapter's artifact is valuable for a decision it makes possible in the other direction — the same asymmetry Chapter 11 found in experiments, arriving here in a cheaper form. The pre-read's highest-value output is not the six views that shipped. It is the four charts that were repaired before they reached the director of real estate, and the one conclusion that reversed when an axis changed meaning. Managers judge visualization work by the picture that made something clear. Analysts should judge it by the picture that stopped something wrong, and should say so, because the second kind is invisible by construction and therefore never gets credited unless the analyst names it.

12.15 Business Analytics in Practice

This section turns from the fictional case to how visualization standards, governed data sources, and dashboard portfolios operate inside professional organizations — where the governance of charts has quietly become an infrastructure problem, where a single axis has cost real credibility, and where the enormous investment in self-service dashboards has produced a graveyard that most analytics leaders can describe from memory.

The first vignette is the rise of the visualization standard, which in most large firms arrived not as a design initiative but as damage control. It is a composite drawn from recurring organizational practice rather than a report on one named firm, and the quantities in it are illustrative of the pattern rather than measured. The pattern is consistent enough to be a genre. A company adopts a self-service BI platform; within eighteen months it has several thousand workbooks, several dozen definitions of revenue, and a weekly meeting in which two directors present contradictory charts of the same metric and neither can reconstruct the other's filters. The response, at firms from national retailers to global banks, has been to build two things at once. The first is a style guide — a documented chart standard specifying the corporate palette with its sequential and diverging ramps, the mandatory zero baseline on bar charts, the required footer carrying source, window, filters, and grain, and an explicit list of prohibited forms that almost always names the three-dimensional chart and the multi-slice pie. The second, and the more consequential, is the certified data source: a governed, versioned dataset published by a data team with its joins, grain, and metric definitions fixed, which analysts connect to instead of building their own extract from the warehouse. The certification does for a chart what the data dictionary of Section 3.8 does for a metric — it moves a definition out of individual judgment and into shared infrastructure — and it directly prevents the failure Lab 12.1 Part A was built around, because a certified source publishes its grain and a workbook built on it cannot quietly compute a per-order metric from line-grain rows. The practice lesson for a marketing analyst is about where the argument goes. In an organization with a standard, the argument about whether a bar chart may start at 4,000 happens once, in a document, among people thinking clearly and without a deadline; in an organization without one, it happens every week, in a meeting, about a specific chart whose author is present and defensive. The standard's value is not aesthetic consistency. It is that it converts a recurring interpersonal negotiation into a settled rule.

The second vignette is the cost of a single axis, and it is worth telling carefully because the mechanism is more interesting than the scandal. Truncated axes appear regularly in investor presentations, product-launch decks, and public-health graphics, and they are almost never fabrications: the underlying numbers are correct, the axis is labeled, and the presenter can point to both. What the research shows is that none of that helps. Pandey et al. (2015) presented readers with matched pairs of charts — identical data, one drawn from zero and one truncated — and measured substantial shifts in readers' judgments of how large the differences were. Correll et al. (2020) pushed the finding further, testing whether making the truncation obvious repaired the damage, and found that readers' magnitude judgments remained distorted even when the axis break was clearly signaled. The practical consequence, seen repeatedly when a truncated chart from a corporate deck is picked up and redrawn by a journalist or a competitor, is that the organization loses a specific and hard-to-recover asset: the presumption that its charts can be read at face value. Once an audience has learned that this company's axes require checking, every subsequent chart is read with suspicion, including the honest ones, and the analyst who spent a year building careful views pays the cost of a colleague's ninety-minute deck. The practice lesson is that chart credibility is organizational rather than individual, and it is consumed by whoever spends it.

The third vignette is the dashboard graveyard, and it is the vignette most likely to describe a student's first job. It is likewise a composite based on recurring organizational practice rather than a citation of published usage statistics. Organizations that deployed self-service BI at scale generally built far more dashboards than anyone uses, and the ones that later instrumented their platforms — most BI tools log which workbooks are opened, by whom, how often, and for how long — commonly report distributions that are uncomfortably skewed: a small minority of dashboards carrying nearly all of the viewing, a long tail opened once by their author and never again, and a substantial population of workbooks whose only regular visitor is the automated refresh job. The diagnosis, in the retrospectives that followed, was rarely that the charts were badly drawn. It was that they had been built without a decision attached — commissioned as "visibility into X," delivered as twelve views of X, and abandoned because nobody's Tuesday actually required any of them. The teams that responded well did two things that connect directly to this guide's spine. They began requiring a named decision and a named decision-maker before a dashboard was built, which is the analytic specification of Section 2.5 applied to a visual deliverable. And they began treating usage telemetry as a feedback loop, retiring unused workbooks on a schedule and interviewing the users of the surviving ones, which turned dashboard-building from a request-fulfillment queue into something closer to product work. The practice lesson is that the failure rate of visualization work in industry is dominated not by encoding errors but by unattached artifacts — beautiful charts nobody needed — and that the defense is upstream of anything in this chapter.

In your first analyst job, these three compress into one expectation, and it is this chapter's closing thread: analysts are judged as much on the charts they refuse to make as on the ones they build. Some of the refusals are the obvious ones — the pie with eleven slices, the truncated bar, the dual axis requested specifically because the two lines "look better together." Most are quieter. The refusal to show a metric without the reference it must be read against. The refusal to build the fourteenth view of a dashboard nobody opens. The refusal to put a magnitude comparison on a map because the map is what the executive asked for. The refusal to let a five-minute exploratory chart, correct for its purpose, be pasted into a deck where its defaults become claims. None of these produce a deliverable, and all of them are visible to the people who matter, because a colleague who says "I can build that, and here is what it would make people believe that is not true, so here is what I would build instead" is doing something a chart-generating tool cannot do and an assistant will not do unasked. That sentence is the job. The tool draws; the assistant drafts; the analyst decides what should be seen, and takes responsibility for what the seeing produces.

12.16 Ethics, Misleading Mechanics as an Integrity Obligation

Section 12.15 ended on the charts an analyst declines to build. This section examines why declining is an obligation rather than a preference, and it extends the guide's running ethics discussion — data use (Section 1.13), problem framing (Section 2.12), measurement design (Section 3.13), cleaning as editorial power (Section 4.14), honest summarization (Section 5.14), differential treatment of segments (Section 6.16), causal language (Section 7.15), acting on predictions about people (Section 8.15), fairness across groups (Section 9.15), forecast accountability (Section 10.15), and experimentation on customers (Section 11.15) — into the one medium where a false impression can be produced without a single false statement.

That is the feature which makes visualization ethics distinctive, and it should be stated at full strength. Every prior chapter's ethical failures involved something an analyst wrote, chose, or withheld that could in principle be quoted back: a metric defined to flatter, an exclusion left undocumented, a causal verb attached to correlational evidence, a forecast stripped of its interval. The mechanics of Section 12.7 involve none of that. The deck circulating since Thursday contains no false number, no misdefined metric, no overclaiming sentence, and no undocumented exclusion. It contains an axis that starts at $240,000. Every element of it would survive a fact-check, and its lead chart nonetheless caused a director of real estate to believe something about StyleCraft's suburban business that is not true. An ethics built entirely on the accuracy of statements does not reach that case, which is why the profession needs one built on the accuracy of impressions.

Concept

The Default Is Not a Defense

Four of the five mechanics in Table 12.3 are available as software defaults or single clicks, and this fact is usually offered in mitigation. It is the opposite. A defense of the form "the tool chose the axis" concedes precisely the thing that matters, which is that nobody looked — and once an analyst has read this chapter, "nobody looked" is a description of the analyst rather than of the tool. The obligation this creates is small and absolute: before any view leaves your screen, read its axis floor, name its grain, and state the comparison a reader will make. That is under a minute of work, it catches the great majority of what Section 12.7 catalogs, and it converts the standard from did I intend to mislead — a question about the analyst's interior state, which nobody else can check and which is therefore useless as a professional standard — to did I check, which leaves evidence and can be audited. The distinction matters most in the case that actually occurs. Almost nobody in marketing sets out to build a deceptive chart. A great many people ship one, quickly, under deadline, because the default filled the frame and the deadline filled the day, and the injury to the reader is identical either way.

Three obligations follow, and they run from the narrow to the broad.

The first is the duty to inspect one's own output, which is what the Appendix E checklist exists to make routine. An obligation that depends on remembering to care will be honored on quiet weeks and abandoned on the weeks that matter, which are exactly the weeks when charts reach decision-makers. Making it a checklist — read the axis, name the grain, name the comparison, run the grayscale check, write the footer — removes it from the domain of virtue and puts it in the domain of process, which is where obligations survive.

The second is the duty to correct other people's charts, and it is harder because it is social. The merchandising analyst who built the circulating deck is a colleague, junior, and under the same deadline; the director of real estate has already praised it; and the analyst who raises the axis question is choosing a small, immediate, personal cost in order to prevent a larger, deferred, institutional one. This is the ordinary shape of professional integrity and it is worth naming as such, because students consistently underestimate how often analytical ethics presents itself as an awkward conversation rather than as a dilemma. Two things make the conversation survivable, and both were built into Lab 12.2 Part B. Name the mechanic rather than the person, since "this axis starts at 240,000, which makes a 4.6 percent change read as a twelvefold jump" is a fact about a chart that anyone can verify, while "your chart is misleading" is a claim about a colleague that invites defense. And bring the repair, because a correction that arrives with a rebuilt chart attached is a contribution and one that arrives alone is a complaint.

The third is the duty that attaches to the audience rather than to the chart, and it is the one that will grow over these students' careers. Readers of a marketing chart are almost never in a position to check it. They do not have the extract, they do not know the grain, they cannot tell whether the window was chosen or inherited, and they are reading in a meeting with eleven other items on the agenda. Every one of those asymmetries transfers responsibility to the person who built the view, and the transfer is total in the case that Chapter 13 will inherit: a chart that travels beyond the room where its author can be questioned. The working standard this guide adopts is therefore a disclosure standard rather than an intention standard. Build the view so that the choices which could have gone another way are visible on it — the axis range, the window, the grain, the filters, the reference — and an audience that cannot check your work can at least see what you decided. That is what the footer discipline of Section 12.9 is actually for, and it is why five lines of small type at the bottom of a chart are an ethical instrument rather than a formatting convention.

One boundary closes the section, because students reasonably ask where the line sits and this chapter is not the place where it is drawn. Emphasis is legitimate. Choosing to color the suburban series and gray the others, choosing to sort by the value under discussion, choosing to annotate the opening date — all of these direct a reader's attention, all of them are decisions the analyst makes on the reader's behalf, and none of them misstates a quantity. Distortion is different in kind: it changes what the reader believes the numbers are. This chapter owns the mechanics by which the second happens and asserts the obligation to prevent them. Where exactly emphasis becomes distortion, in a chart built to persuade an executive who has ten minutes and a decision to make, is the harder question, and it belongs to Chapter 13, which will have to answer it in front of an audience rather than in a lab.

12.17 Chapter Summary

This chapter opened Part III by treating visualization as an instrument rather than as a finish, and by claiming the first of its two jobs. A chart in its analytical role is built for discovery — dense, disposable, drawn so that the analyst can find out what the data contains — and a chart in its communicative role is built for transfer, and the properties that make each good actively damage the other, which is why the pre-read due Monday and the dashboard due in December are different objects rather than different drafts. Underneath both sits a grammar: every chart is a set of marks, a set of assignments from data fields to visual channels, and the axes, aggregations, and sort orders those channels are read against, so that a chart can be specified in a sentence, audited component by component, and altered one component at a time. The grammar becomes consequential through the perceptual hierarchy, which is the chapter's central empirically measured ordering — position read most accurately, then length, then angle, then area, then color intensity — and which converts chart selection from taste into a prediction about how accurately a reader will decode the comparison she is being asked to make. From there the chapter made the choices operational. Charts were selected by analytic purpose rather than from a gallery, with distribution, comparison, trend, relationship, composition, and geographic mapped to the StyleCraft questions Chapters 4 through 10 had already asked. Scales were examined as claims about what visual distance means, with the logarithmic axis recovered as the correct instrument for comparing growth rates across very different levels and the zero-baseline obligation reassigned from blanket rule to encoding-specific duty: an ordinary bar whose full length encodes magnitude must start at zero, a floating or interval mark must start at the reference its length is measured from and say so, an area mark must map zero to zero area, and any mark whose position carries the value needs a range that was chosen and stated rather than inherited from the software. The five standing mechanics were then cataloged and repaired — axis truncation, which on StyleCraft's own quarterly numbers drew the second bar twelve times as tall as the first for a value only 1.046 times as large, and which experimental work shows survives both correct labeling and explicit warning; dual axes, whose apparent coupling is manufactured by two arbitrary range choices and which invites a causal reading that Section 7.5 already disposed of; the cherry-picked window, defeated by showing a full seasonal cycle and stating why the window starts where it starts; area distortion, defeated by scaling area rather than radius and by moving magnitude comparisons off the map entirely; and the three-dimensional effect, which has no analytical use. Color completed the encoding system as a system: palette type chosen from the field's level of measurement per Section 3.4, color spent on emphasis rather than spread across every category, accessibility treated as a condition of correctness, screened by a grayscale conversion and then confirmed with a color-vision-deficiency simulation and a check that no essential meaning rests on color alone. Reference lines and annotation then supplied what no encoding can, which is something to read the values against — the visual form of the baseline discipline of Section 2.7 — along with the footer that records source, window, filters, and grain and thereby keeps a chart interpretable after it separates from its author. The tool arrived last and mapped onto what students already owned: Tableau's dimensions and measures closely parallel the variable roles of Section 3.5 while remaining software roles that govern query construction, discrete and continuous determine whether a field produces headers or an axis, and aggregation in the view is where the grain discipline of Section 3.3 becomes invisible — which the labs demonstrated by showing StyleCraft's entire suburban finding disappear into a 3 percent difference when average order value is computed from line-grain rows and reappear as a 66 percent difference when it is computed at order grain, with no error message in either case. The labs built the expansion story six ways from two certified extracts that reconcile to the cent, and then produced the chapter's most useful single exhibit: seventeen small multiples on one shared scale, plotted first against calendar time, where the newest stores look like failures, and then against store age, where they are read against the band their predecessors traced — the same data, the same facets, one changed axis, and a recommendation that may not survive the change. The AI section installed the chapter's verification theme as a routine that predicts an AI-drafted chart's faults from its specification before the chart is rendered, on the grounds that a picture is persuasive in a way a specification is not, and named the nine failure modes an assistant reliably produces, chief among them the axis truncated on request and the per-order metric computed from per-line rows. And the ethics section drew the distinction that makes this chapter's obligations unusual: every mechanic in it produces a false impression without a false statement, so the professional standard cannot be did I intend to mislead but must be did I check, which is a minute of work and leaves evidence.

Looking ahead, the pre-read now exists, and its readers will spend the next week disagreeing with it, which is what it was built for. On December 2 the disagreement ends and a decision begins, and the artifact that meeting runs on cannot be this one. The next chapter turns to data visualization for communication and decision-making, where the reader is a chief executive with ten minutes rather than a colleague with a week, where the three-second test replaces the visual audit as the governing constraint, and where every technique this chapter deliberately withheld — assertion titling, aggressive decluttering, narrative sequencing, interactivity that guides rather than explores — becomes the subject. The mechanics of Section 12.7 travel forward unchanged and are not re-argued there; what Chapter 13 must add is the harder line this chapter declined to draw, between the emphasis that helps a decision-maker see what matters and the concealment that helps her agree. Between here and there sits Project #2, whose brief the next chapter releases and whose deliverable is the dashboard Chapter 13 teaches. The exercises below prepare the certified extracts and exploratory views that students may reuse once that brief arrives.

12.18 Exercises for Practice and Homework

The following exercises practice the chapter's main habits: name the analytic purpose before choosing a form, state what one mark represents before building a view, read the axis floor on anything encoded by length, spend color on emphasis rather than on decoration, put the reference the reader needs on the chart itself, and predict what a chart will show before rendering it. Not everything below is required. Core practice, which every student should complete: 12.18.1, 12.18.2, 12.18.3, and 12.18.4. Homework options, from which your instructor will assign a subset: 12.18.5 and 12.18.6. Instructor-led discussion, prepared for class rather than submitted: 12.18.7 and 12.18.8. Each exercise also carries its label individually. No new project deliverable is due this week; the two Homework Submission items are assignment options from which your instructor will select, and any of them that is assigned is graded. Project #2's brief is released in Chapter 13, and Section 12.18.5 produces the certified extracts and the first views that project will build on, so students who complete it will start Project #2 a week ahead.

12.18.1 Concept Check (Required Practice)

Answer each in two or three sentences, in your own words.

  • Distinguish visualization's analytical job from its communicative job, and name one property that improves a chart for the first job while damaging it for the second.
  • Define a mark and an encoding, and specify a sorted bar chart of revenue by store in that vocabulary, naming every component including the ones the software supplies by default.
  • State the perceptual hierarchy in order, and explain why a pie chart of seven categories is a worse instrument than a sorted bar chart of the same seven.
  • Explain why ordinary magnitude bars must begin at zero while position-encoded lines need not, in terms of what each mark encodes rather than in terms of convention, and state what a floating or interval bar must begin at instead.
  • A colleague defends a truncated bar chart on the grounds that the axis is clearly labeled. State the empirical finding that answers him, and cite the study.
  • Explain what makes a dual-axis chart's apparent coupling arbitrary, and name a replacement that answers the same question honestly.
  • State the rule for matching a palette type to a field, and give one StyleCraft example of each of the three palette types.
  • Explain what a reference line does that no encoding can do, and connect it to the discipline introduced in Section 2.7.
  • Define aggregation in the view, and explain how it can produce a correct number that answers the wrong question with no error message.
  • Explain what small multiples buy and what they cost, in terms of the perceptual hierarchy, and give the StyleCraft case in which the purchase was clearly worth the price.

12.18.2 Predict the Encoding (Required Practice)

For each of the following comparisons, state (a) the analytic purpose, (b) the encoding channel you would use to carry the comparison, (c) that channel's rank in the perceptual hierarchy, and (d) one encoding you considered and rejected, with the reason. Do not build anything.

  • Which of StyleCraft's seventeen stores produced the most trailing-twelve-month revenue, and by how much more than the second.
  • Whether the share of revenue coming from occasionwear differs between NYC Urban and NYC Suburban customers.
  • Whether weekly revenue at the Washington store is growing faster than weekly revenue at SoHo, given that SoHo runs at roughly $7,900 a week and Washington at roughly $1,800.
  • Whether customers with the app spend more than customers without it, across the full range of spend rather than on average.
  • Whether the eight Wave 4 stores are clustered in any particular part of the country.

Then answer the question the exercise is actually about. For the third item, state which scale you chose and write the one-sentence caption that would prevent a reader from misreading the level difference. For the fifth, state what your chosen view can and cannot establish about why those stores were opened where they were.

12.18.3 Purpose First (Required Practice)

Return to Table 12.6 and complete it for the three views the labs did not build — A, D, and E — filling every column: purpose, the comparison the reader must make, the grain of one mark, and the reference the view must carry. Then, for each of the three, write a full specification in the form Section 12.12's request prompt asks an assistant to produce: mark type, every field-to-channel encoding, the aggregation and grain, the axis ranges with a justification for each, the sort order, and the palette with its type. Finally, for View E, name two chart forms that could answer the composition question, state which you would build for the pre-read and which you would build for yourself while exploring, and explain the difference in terms of Section 12.2.

12.18.4 The Grain Trap (Required Practice)

Using transactions_enriched in Tableau or in Colab, whichever you prefer:

Compute average order value for NYC Urban and NYC Suburban customers two ways — as AVG([line_revenue]) on the line-grain file, and as SUM([line_revenue]) / COUNTD([order_id]) — and report all four numbers. State which pair is average order value and which pair is something else, and name the something else precisely. Then compute the difference between the metros under each definition, in dollars and as a percentage, and write the one sentence a reader would take away from each, so that the two sentences sit side by side on the page.

Next, decompose the correct gap. Compute items per order and average line value for each metro, and show that the gap equals the basket-size component plus the price-and-mix component, with the arithmetic displayed. State in one sentence what the decomposition establishes about whether StyleCraft charges suburban customers more, and connect the finding to what the course data dictionary documents about the dataset's design.

Finally, list three other metrics used anywhere in Chapters 5 through 11 that require a grain other than order-line, name the grain each requires, and state where in a Tableau workflow you would enforce it — in Prep, in a calculated field, or in the view's dimensions — with one sentence of justification.

12.18.5 The Visual Audit Deliverable (Homework Submission)

Build the pre-read described in Section 12.14, as a Tableau workbook plus a two-page written guide.

The workbook must contain, in order: the two certified extracts from Lab 12.1 Part A with the revenue reconciliation documented; the six views of Table 12.6, each carrying its reference and its footer; and the two faceted views of Lab 12.1 Part C, on calendar time and on store age. Every view must state, in its caption, what one mark represents.

The written guide must contain the reading guide paragraph, one paragraph per view naming the question it answers and the finding it supports, the two-picture argument about store age with an explicit statement of what changed between the pictures, a paragraph on what the map does not establish, and an appendix listing at least three charts you built and rejected with one sentence apiece on why. Include, as the guide's final paragraph, the single sentence you would say first if you had thirty seconds with the director of real estate — and note that Chapter 13 will ask you to rebuild the whole thing around that sentence.

12.18.6 AI Chart Audit (Homework Submission)

Run the five-step routine of Section 12.12's AI in Practice box on a question of your choosing from Table 12.6, and document it per Appendix D. Submit: (a) your request prompt and the assistant's three candidate specifications, with the schema you pasted and any column the assistant invented despite it; (b) your written predictions, made from the specifications before building anything, naming for each candidate the comparison, the channel, its hierarchy rank, and the two faults you expected; (c) the built views, faults included, with screenshots; (d) the completed audit against the three transferable points of Table 8.5 and all three of Table 12.5, with each fault named, quoted, and quantified where quantification is possible; (e) your grade of your own predictions, including both the faults you missed and the ones you predicted that did not occur, with a paragraph on what the pattern reveals about your model of the assistant; and (f) the repaired views with a one-sentence note per repair recording what changed in the reading.

Your grade rests on the audit and the prediction, not on the assistant's output. An exchange in which the assistant produced four faults and you predicted three of them is a better submission than one in which it produced a clean chart you did not have to inspect.

12.18.7 Spot the Mechanic (In-Class Discussion)

Each of the following describes a chart you will meet in commercial practice. Name the mechanic from Table 12.3 or the audit point from Table 12.5 that it violates, state in one sentence what a reader will believe that is not true, and state the single change that would repair it.

  • A quarterly revenue bar chart whose vertical axis runs from $4.2 million to $4.6 million, with both values labeled on the bars.
  • A line chart showing email sends and revenue on twin vertical axes, presented with the caption "email drives revenue."
  • A map of the United States shaded by total orders, presented as evidence that the Northeast is StyleCraft's strongest market.
  • A bar chart of average order value by loyalty tier, sorted alphabetically as None, Insider, VIP, Backstage, in four saturated hues.
  • A chart of campaign performance covering the six weeks following a campaign launch, in a business whose fourth quarter carries a documented holiday lift.
  • A dashboard tile reading "Average order value: $50.41," computed on a file with one row per order line.

For the last item, state what the correct number is and estimate what the error would do to a revenue projection built on it.

12.18.8 The Colleague's Deck (In-Class Discussion)

You are the analyst in Section 12.1. The merchandising analyst's deck has reached the director of real estate, who has replied "this is great, let's use these." You have four working days before the pre-read circulates, and the December 2 meeting is three weeks out. The merchandising analyst reports to a different manager, is two years into her first job, and built the deck in ninety minutes using an AI assistant because she was asked to produce something by Friday.

Prepare a position, and address each of the following explicitly. What do you send, to whom, and in what order — the merchandising analyst first, her manager first, the director first, or all three at once? What exactly does the message contain, given Section 12.16's instruction to name the mechanic rather than the person and to bring the repair? How do you handle the fact that the director has already endorsed the charts publicly? Which of the four faults would you raise and which, if any, would you let stand, and on what principle do you distinguish them? What would you do differently if the deck had already gone to the chief executive rather than to a peer of yours? And finally: the merchandising analyst did nothing that this course had taught her not to do, and produced in ninety minutes something that took you two labs to check. State what organizational change — not what individual virtue — would have prevented the situation, and name the section of this chapter that describes it.

12.19 Glossary of Terms

This glossary includes only the terms introduced in this chapter. Each definition is tied to the sources used in the chapter rather than added for decoration.

Aggregation in the view. Tableau's behavior of computing each measure at the level of detail defined by the dimensions currently in the view, recomputed silently on every change, so that a metric requiring a specific grain will be computed at the wrong grain without warning whenever the view's dimensions do not match it (adapted from Murray, 2016, and the grain discipline of Section 3.3).

Analytic purpose. The class of question a view is built to answer, stated as the comparison or structure the reader must extract rather than as the chart form that displays it; the six purposes covering nearly all marketing analysis are distribution, comparison, trend, relationship, composition, and geographic (adapted from Shneiderman, 1996; Munzner, 2014).

Analytical visualization. A chart whose dominant job is discovery — built by the analyst, for the analyst, to find out what the data contains, and judged by whether it reveals structure not already known; typically dense, disposable, and labeled enough to be reproduced rather than polished for an outside reader (Section 12.2; Tukey, 1977; Gelman & Unwin, 2013).

Annotation. Explanatory text placed on a chart itself — a label on a mark, a note on a region, or a footer recording source, window, filters, and grain — supplying context the encoding cannot carry and keeping the chart interpretable after it separates from its author (adapted from Tufte, 2001; Few, 2012).

Area distortion. Misencoding of a quantity in the size of a two-dimensional mark, most often by scaling radius or side length in proportion to value so that area grows with the square of the value; compounded by readers' general underestimation of area differences, and repaired by scaling area to value, labeling a size legend, and moving precise comparisons to a length- or position-encoded view (adapted from Cleveland & McGill, 1984; Tufte, 2001).

Axis truncation. Omission of the zero origin from an axis whose marks encode value by length or area, so that visual proportion no longer corresponds to numerical proportion; the distortion survives correct axis labels and explicit warnings, and is legitimate only for position-encoded marks, where it restricts the view rather than misstating the proportion (adapted from Pandey et al., 2015; Correll et al., 2020).

Cherry-picked window. Selection of the displayed extent of data — of time, of categories, or of units — so that the visible pattern differs from the pattern in the full data; every displayed number may be correct. The signature is a window whose start has no stated justification, and the repair is to state why the window begins where it begins and to show at least one full seasonal cycle in a seasonal business (Section 12.7; Section 10.3).

Communication visualization. A chart whose dominant job is transfer — built by the analyst, for someone else, to convey a finding already established, and judged by whether a reader who has never seen the data arrives at the intended understanding in the time available; typically sparse, durable, and finished, and the subject of Chapter 13 rather than of this chapter (Section 12.2; Gelman & Unwin, 2013).

Dimension. In Tableau's data model, a field used to partition, group, or slice the data, defining the categories, dates, and identifiers at which values are separated and thereby setting the level of detail of a view; closely parallel to the categorical and identifying variable roles of Section 3.5, but a software role that can be changed without changing what the field means (adapted from Murray, 2016; Tableau, 2026a).

Discrete and continuous fields. A discrete field takes separate values and produces headers or panels; a continuous field takes values along an unbroken range and produces an axis. The distinction is independent of whether a field is a dimension or a measure, and determines whether marks are separated into panels or positioned along a scale (adapted from Murray, 2016).

Dual axis. A chart plotting two measures against a shared horizontal axis using two independently scaled vertical axes, so that the series' apparent alignment or divergence is a property of the arbitrary scaling rather than of the data; defensible only when the two scales are fixed by definition, and otherwise replaced by stacked panels or by indexing both series to a common base (adapted from Cairo, 2019; Few, 2012).

Encoding. The assignment of a data field to a visual channel of a chart's marks — position, length, angle, area or size, color hue, color intensity, shape, or texture — which together with the mark type and the scales fully specifies a chart (adapted from Bertin, 1983; Wilkinson, 2005).

Faceting. Partitioning data by a dimension and drawing one panel per value of that dimension, producing small multiples; the operation that trades common-baseline comparison for the ability to compare many subsets at once (adapted from Tufte, 2001; Munzner, 2014).

Linear scale. An axis on which equal visual distances represent equal differences, so that the question the axis answers is "how much more"; the default for nearly every chart form and the required scale wherever a mark's length or area encodes magnitude (Section 12.6).

Logarithmic scale. An axis on which equal visual distances represent equal ratios, so that the question the axis answers is "how many times more" and, for a series against linear time, at what proportional rate a quantity is growing; a straight line indicates approximately constant exponential growth. Available only when every plotted value is positive, since zero and negative values cannot be placed on a log axis; inappropriate for length-encoded marks, and unfamiliar enough to most executive readers that its use requires an explicit note (Section 12.6; Tableau, 2026e).

Mark. The geometric object a chart draws to represent data — point, line, bar, rectangle, circle, arc, or filled area — whose type determines what kind of object the reader is looking at and which visual channels are available to encode fields (adapted from Bertin, 1983; Wilkinson, 2005).

Measure. In Tableau's data model, a quantitative field aggregated by default within the partitions defined by the dimensions present in a view; closely parallel to the metric and quantitative variable roles of Section 3.5, but a software role, and one a field can be converted out of (adapted from Murray, 2016; Tableau, 2026a).

Perceptual hierarchy. The empirically measured average ordering of visual encoding channels by the accuracy with which readers extract quantitative values from them — position on a common scale, then position on identical non-aligned scales, then length, angle and slope, area, volume, and color intensity — used to assign the most important comparison to the highest-ranking available channel. A robust average ordering for elementary quantitative judgments rather than a universal ranking: measured performance varies with task, display, reader, and context, and the ordering governs the comparison the layout actually asks the reader to make (adapted from Cleveland & McGill, 1984; Heer & Bostock, 2010).

Reference line. A line, band, or marker drawn at a separately defined benchmark — external or data-derived — against which the plotted marks are interpreted, supplying the comparison the reader would otherwise have to provide from memory. Externally supplied decision thresholds should be declared before results are inspected; descriptive references such as an average or a median may be computed from the plotted data but must be labeled as such. The visual form of the baseline discipline of Section 2.7 (adapted from Few, 2012; Munzner, 2014; Tableau, 2026d).

Sequential, diverging, and categorical palettes. Sequential palettes vary lightness within one hue and encode ordered quantities; diverging palettes run two hues outward from a meaningful midpoint and encode distance and direction from a reference; categorical palettes assign distinguishable hues of roughly equal lightness to unordered groups and carry no magnitude information. Matching palette type to the field's level of measurement is a correctness requirement rather than a stylistic one (adapted from Harrower & Brewer, 2003; Munzner, 2014).

Small multiples. A series of small charts sharing identical axes, scales, and encodings, each showing the same measures for a different subset, arranged so the subsets can be compared by scanning; because the scale is shared, differences between panels are read as differences in the data (adapted from Tufte, 2001; Munzner, 2014).

12.20 Further Readings

Students who want additional background may begin with the following readings. The design treatments are listed first, the perceptual evidence second, and the tool reference last.

  • Cairo (2019) for the most readable modern treatment of how charts mislead, written for a general audience by a journalist and academic who has spent a career on the problem; its chapters on scale and on charts that appear to show causation are the direct expansion of Sections 12.6 and 12.7.
  • Few (2012) for the practitioner's standard on business charts and tables — the source of much of the reference-line, annotation, and emphasis discipline used in Sections 12.8 and 12.9, and the most immediately applicable book on this list for a marketing analyst.
  • Munzner (2014) for the systematic treatment of visualization as a design problem, organized around what task a visualization must support and which channels can support it; the level above this chapter's Sections 12.3 through 12.5, and the best available bridge from these ideas to the more advanced visualization work that follows this guide.
  • Cleveland and McGill (1984) for the original experiments behind the perceptual hierarchy — short, readable, and worth reading in the original, because seeing how the ordering was actually measured is what makes it a fact rather than an opinion to be argued with.
  • Correll et al. (2020) for the modern experimental treatment of axis truncation, including the uncomfortable finding that warning readers does not repair the distortion; the paper behind Section 12.7's insistence that the repair must be structural.
  • Murray (2016) together with Tableau Help (2026a through 2026h) for the conceptual and the current-interface treatments of the tool. Read Murray for the account of the data model, dimensions and measures, and the discrete-continuous distinction that Section 12.10 maps onto Chapter 3's variable roles; read the documentation for the authoritative current build paths behind every instruction in Section 12.13 — field roles, Prep aggregation and joins, Prep functions, reference lines, axis configuration, geographic roles, proportional-symbol maps, and quantitative mark sizing. Where a printed source and the vendor's documentation disagree about an interface, the documentation governs, and it is the reference an analyst should learn to check first.

12.21 References

Amar, R., Eagan, J., & Stasko, J. (2005). Low-level components of analytic activity in information visualization. In Proceedings of the 2005 IEEE Symposium on Information Visualization (pp. 111–117). IEEE. https://doi.org/10.1109/INFVIS.2005.1532136

Bertin, J. (1983). Semiology of graphics: Diagrams, networks, maps (W. J. Berg, Trans.). University of Wisconsin Press. (Original work published 1967)

Borland, D., & Taylor, R. M. (2007). Rainbow color map (still) considered harmful. IEEE Computer Graphics and Applications, 27(2), 14–17. https://doi.org/10.1109/MCG.2007.323435

Cairo, A. (2019). How charts lie: Getting smarter about visual information. W. W. Norton.

Cleveland, W. S., & McGill, R. (1984). Graphical perception: Theory, experimentation, and application to the development of graphical methods. Journal of the American Statistical Association, 79(387), 531–554. https://doi.org/10.1080/01621459.1984.10478080

Correll, M., Bertini, E., & Franconeri, S. (2020). Truncating the y-axis: Threat or menace? In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (pp. 1–12). Association for Computing Machinery. https://doi.org/10.1145/3313831.3376222

Few, S. (2012). Show me the numbers: Designing tables and graphs to enlighten (2nd ed.). Analytics Press.

Gelman, A., & Unwin, A. (2013). Infovis and statistical graphics: Different goals, different looks. Journal of Computational and Graphical Statistics, 22(1), 2–28. https://doi.org/10.1080/10618600.2012.761137

Harrower, M., & Brewer, C. A. (2003). ColorBrewer.org: An online tool for selecting colour schemes for maps. The Cartographic Journal, 40(1), 27–37. https://doi.org/10.1179/000870403235002042

Heer, J., & Bostock, M. (2010). Crowdsourcing graphical perception: Using Mechanical Turk to assess visualization design. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (pp. 203–212). Association for Computing Machinery. https://doi.org/10.1145/1753326.1753357

Munzner, T. (2014). Visualization analysis and design. CRC Press.

Murray, D. G. (2016). Tableau your data! Fast and easy visual analysis with Tableau Software (2nd ed.). Wiley.

Pandey, A. V., Rall, K., Satterthwaite, M. L., Nov, O., & Bertini, E. (2015). How deceptive are deceptive visualizations? An empirical analysis of common distortion techniques. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems (pp. 1469–1478). Association for Computing Machinery. https://doi.org/10.1145/2702123.2702608

Shneiderman, B. (1996). The eyes have it: A task by data type taxonomy for information visualizations. In Proceedings of the 1996 IEEE Symposium on Visual Languages (pp. 336–343). IEEE. https://doi.org/10.1109/VL.1996.545307

Tableau. (2026a). Dimensions and measures, blue and green. Tableau Help. Retrieved July 30, 2026, from https://help.tableau.com/current/pro/desktop/en-us/datafields_typesandroles.htm

Tableau. (2026b). Aggregate, join, or union data. Tableau Prep Help. Retrieved July 30, 2026, from https://help.tableau.com/current/prep/en-us/prep_combine.htm

Tableau. (2026c). Date properties for a data source. Tableau Help. Retrieved July 30, 2026, from https://help.tableau.com/current/pro/desktop/en-us/date_properties.htm

Tableau. (2026d). Reference lines, bands, distributions, and boxes. Tableau Help. Retrieved July 30, 2026, from https://help.tableau.com/current/pro/desktop/en-us/reference_lines.htm

Tableau. (2026e). Edit axes. Tableau Help. Retrieved July 30, 2026, from https://help.tableau.com/current/pro/desktop/en-us/formatting_editaxes.htm

Tableau. (2026f). Create maps that show quantitative values in Tableau. Tableau Help. Retrieved July 30, 2026, from https://help.tableau.com/current/pro/desktop/en-us/maps_howto_symbol.htm

Tableau. (2026g). Format geographic fields in Tableau. Tableau Help. Retrieved July 30, 2026, from https://help.tableau.com/current/pro/desktop/en-us/maps_geographicroles.htm

Tableau. (2026h). Tableau Prep functions reference. Tableau Prep Help. Retrieved July 30, 2026, from https://help.tableau.com/current/prep/en-us/prep_function_reference.htm

Tufte, E. R. (2001). The visual display of quantitative information (2nd ed.). Graphics Press.

Tukey, J. W. (1977). Exploratory data analysis. Addison-Wesley.

Wainer, H. (1984). How to display data badly. The American Statistician, 38(2), 137–147. https://doi.org/10.1080/00031305.1984.10483186

Wickham, H. (2010). A layered grammar of graphics. Journal of Computational and Graphical Statistics, 19(1), 3–28. https://doi.org/10.1198/jcgs.2009.07098

Wilkinson, L. (2005). The grammar of graphics (2nd ed.). Springer.

Wong, B. (2011). Points of view: Color blindness. Nature Methods, 8(6), 441. https://doi.org/10.1038/nmeth.1618

Zacks, J., & Tversky, B. (1999). Bars and lines: A study of graphic communication. Memory & Cognition, 27(6), 1073–1079. https://doi.org/10.3758/BF03201236

Annotate

Book Chapters
Powered by Manifold Scholarship. Learn more at
Opens in new tab or windowmanifoldapp.org