Notes
Chapter 07
Relationships, Drivers, and Regression
From Whom to What Moves Them: Measuring, Modeling, and Refusing to Overclaim the Drivers of Customer Spend
Dr. Jose Mendoza, Academic Director and Clinical Associate Professor
Version 1.0 · July 2026
Except where otherwise noted, this chapter is licensed under CC BY 4.0.
Chapter Information
ABSTRACT
This chapter develops the analysis of relationships among marketing variables: how association is measured, modeled, and kept distinct from causation. It builds linear association through covariance and the Pearson correlation coefficient, then catalogs what correlation hides — nonlinearity, outlier sensitivity, and mixed populations — and develops confounding, defined as a pre-existing common cause, on marketing’s most reliable confounder, the calendar. Linear regression is built from the simple case through multiple regression, dummy variables, and interactions, with coefficients read as model-adjusted comparisons rather than levers and interactions read conditionally. Statistical evidence is developed at working level: heteroskedasticity-robust standard errors, confidence intervals, p-values, and variance inflation factors. A section on AI assistance installs the four-point audit of signs, magnitudes, units, and diagnostics. The labs estimate spend an average-order-value drivers on the Chapter 6 feature table, recover the chapter’s designed interaction, and verify an AI-fitted model against certified descriptive findings.
KEYWORDS
correlation; covariance; regression; drivers; dummy variables; interaction terms; multicollinearity; confounding; omitted variable bias; causal overclaiming
VERSION AND DATE
Version 1.0 · July 2026 · Language: English (United States)
SUGGESTED CITATION
Mendoza, J. (2026). Relationships, drivers, and regression. In Applied business analytics for marketing decision-making: Business analytics and data visualization (Chapter 7, Version 1.0) [Open educational resource]. CC BY 4.0.
LICENSE AND RIGHTS
Copyright © 2026 Jose Mendoza. Except where otherwise noted, this work is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). You may share and adapt this material for any purpose, provided appropriate credit is given. Third-party trademarks, screenshots, figures, and other materials remain subject to their respective rights and licenses.
Google Colab is a product of Google LLC. "Python" and the Python logos are trademarks or registered trademarks of the Python Software Foundation. pandas, NumPy, and Matplotlib are sponsored projects of NumFOCUS, a 501(c)(3) nonprofit charity in the United States. statsmodels is a community-developed project distributed under the modified BSD license. ChatGPT is a product of OpenAI, Claude of Anthropic, Gemini and NotebookLM of Google LLC, and GitHub Copilot of GitHub, Inc. Product names are used for identification only and do not imply endorsement. StyleCraft Collective is a fictional company created for instruction.
COMPANION REPOSITORY
Datasets, notebooks, and figure sources for this chapter: https://github.com/jrmst102/businessanalytics
GENERATIVE AI USE
Generative artificial intelligence and other AI-assisted tools were used in the research, writing, revision, and production of this chapter, including literature discovery, source organization, outlining, preliminary drafts, prose revision, support for code and analytical examples, and document formatting. These tools were used under the author’s direction and are not credited as authors, researchers, or sources. The author determined the chapter’s scope, learning objectives, methods, interpretations, and recommendations, and reviewed and approved all AI-assisted material: factual claims and citations were checked against the underlying sources rather than accepted from AI-generated summaries, and code and analytical outputs were tested or otherwise reviewed for accuracy. Responsibility for the accuracy, originality, and final form of this chapter rests entirely with the author. A fuller statement appears in the front matter of the complete guide.
Chapter Learning Objectives
- By the end of this chapter, students should be able to:
- Distinguish association from causation as a standing discipline and explain why “driver” is a causal word that correlational evidence places on probation.
- Compute and interpret sample covariance and the sample Pearson correlation coefficient among marketing variables, stating the direction, the strength, the unit-free character of r, and the difference between the population quantity and the sample estimate the software reports.
- Identify what correlation misses — nonlinearity, outlier sensitivity, and mixed populations — and apply the always-plot discipline before trusting any correlation.
- Read a correlation matrix without committing its characteristic misreadings and explain why a matrix of pairwise correlations is not a map of drivers.
- Define a confounder as a pre-existing common cause, distinguish it from a variable that is merely correlated with both members of a pair, and recognize seasonality as marketing’s standing confounder.
- Construct and interpret simple linear regression models: the fitted line, the intercept and slope in native units, least squares as a criterion, and residuals as diagnostic evidence.
- Interpret R² as shared-variation fit under its stated conditions, and explain why fit and usefulness are different judgments, anchored to the baseline discipline of Section 2.7.
- Construct and interpret multiple regression models, read a coefficient as a model-adjusted comparison rather than a lever, and reason about omitted variable bias and its direction, including the qualification that applies once a model holds several predictors.
- Encode categorical variables with dummy variables and reference categories, interpret segment coefficients against the reference, and specify and interpret an interaction term conditionally.
- Assess coefficient evidence using heteroskedasticity-robust standard errors, confidence intervals, p-values, and variance inflation factors at working level, and distinguish prediction from explanation as modeling goals.
- Verify AI-generated regression output against fundamentals — signs, magnitudes, units, fit, and diagnostics — and against certified descriptive findings.
Chapter 6 ended with a map and a question the map could not answer. The map was the treatment map: four segments, estimated, validated, named by rule, and paired with journeys on the fall calendar. The question came from the sign-off meeting, and it was the natural next one — what moves these customers? Which levers, at what strength, with what confidence? Segments describe who; this chapter is about what moves with their behavior, and about the analyst’s principal instrument for that question: regression. The route runs from the raw material of relationships — association measured honestly by correlation, and the many ways correlation deceives — through the regression model itself, simple and then multiple, with the categorical machinery that lets StyleCraft’s estimated segments enter the model as evidence. Because every step of this route is now one prompt away from an AI assistant, the chapter also installs its own verification discipline — the audit of signs, magnitudes, units, and diagnostics — and its own ethical guardrail, because the characteristic failure of this chapter’s tools is not a wrong number but a wrong verb: an association, fluently narrated, wearing the language of cause.
Concept
What This Chapter Is Really About
Chapter 6 argued that grouping is the third analysis, and that an algorithmic grouping is a hypothesis to be validated rather than a fact to be admired. This chapter takes the next step in the same argument: relating is the fourth analysis, and it is the first whose output looks like an answer to the question executives ask — what should we do? A correlation of 0.6, a coefficient of $41 per order, a fitted line through a cloud of customers: each is arithmetic about the world as it is, and each will be read, the moment it reaches a meeting, as a claim about what would happen if the world were changed. The entire discipline of this chapter lives in that gap.
The methods are not hard; a modern AI assistant fits a defensible regression in seconds. What cannot be delegated is the reading: whether the signs and magnitudes survive contact with certified facts, whether the model’s assumptions survive contact with its residuals, and whether the sentence written about the coefficient claims only what the evidence can carry. An analyst who master’s this chapter does not become the person who finds drivers. They become the person who knows exactly how much of “driver” the data has earned.
Source: Course concept developed for this guide, informed by Freedman (1991) and Shmueli (2010).
7.1 Marketing Decision Context: Three Levers and One Budget
Chapter 6 left the CRM manager in possession of a signed treatment map: four validated segments — the Urban Loyal Core, the Suburban Occasion shoppers, the Discount-Dependent, and the New/At-Risk — each attached to one of the four journeys the platform can sustain this fall. The map answered who. The sign-off meeting immediately generated the next decision, and this one has a budget attached. Beyond the fixed cost of building the four journeys, StyleCraft’s fall retention program carries an incremental budget — the money that funds what the journeys do — and the VP of Marketing must allocate it across three candidate levers before the pre-holiday budget lock in twenty-one days. The three levers on the table: deepen the discount ladder (richer offers, triggered more often), accelerate app adoption (an install campaign plus app-exclusive perks), and increase email frequency for enrolled customers. The CFO will attend the allocation meeting. After the margin conversation that opened this guide in Chapter 1, every dollar of discount is presumed guilty until proven incremental.
Three claims are already circulating in the building, one per lever, and each arrives dressed in numbers. The app claim comes from the platform team’s deck: “App users spend 2.4 times more than non-app customers. App adoption is our highest-leverage investment — every install is worth hundreds of dollars in annual spend.” The discount claim comes from a finance analysis: “Customers with high discount exposure show below-average total spend. Discounts are training our customers to pay less — cut the ladder and spend will recover.” The email claim comes from the CRM team itself: “Email-engaged customers buy twice as often as the unengaged. Doubling send frequency doubles touches, and touches drive orders.” Each claim cites a real pattern; each pattern will reappear, verifiable, in this chapter’s labs. And each claim commits, in its second sentence, an inference its first sentence cannot support. The VP has asked the analyst — you — for something specific by the budget lock: an honest drivers analysis of customer spend. Which variables move with spend, how strongly, in what units, holding what constant — and, for each of the three claims, a verdict: what the data supports, what it cannot support, and what evidence would settle the difference.
The deliverable has a shape this chapter will fill in. First, the relationships themselves, measured honestly: which of StyleCraft’s behavioral and engagement variables are associated with spend, at what strength, and with what caveats about shape, outliers, and the mixed population Chapter 6 showed the base to be. Second, a model: spend regressed on the candidate drivers together, so that each claim can be examined holding the others constant — because the app claim’s “2.4 times” compares customers who differ in far more than app ownership, and the meeting deserves to know how much of the gap survives when tenure, recency, engagement, and segment membership are held level. Third, the segments as evidence: the four groups from Chapter 6 enter the model as dummy variables, both because segment membership is itself a candidate driver and because Chapter 5’s certified findings about the segments — the suburban average-order-value gap and its decomposition — are the known truths against which any fitted model must be checked. Fourth, and running through everything, the language: the deliverable’s sentences must survive a CFO who has learned to ask, of every “driver,” whether anyone changed anything and watched what happened.
The stakes of the language are not academic. If the app coefficient is read causally and the install campaign is funded on it, the budget buys the customers StyleCraft already has — the young, urban, loyalty-enrolled core who were always going to install the app — dressed up as growth. If the discount association is read causally and the ladder is cut, the Discount-Dependent segment, one quarter of the treatment map, loses the behavioral lever with which its purchases are most visibly associated, three weeks before the holiday, on evidence that has never tested whether the segment responds to that lever at all. The chapter builds the machinery in order. Sections 7.2 through 7.5 discipline the raw material: association and causation, correlation and its failure modes, confounding and the calendar. Sections 7.6 through 7.10 build the model: simple regression, fit, multiple regression, the categorical machinery, and the reading of statistical evidence. Section 7.11 turns to the AI assistant, which will fit all this fluently and narrate it causally by default. The labs in Section 7.12 estimate the drivers of spend an average order value on the feature table Chapter 6 built — where, by design, the true coefficient structure is known — and Sections 7.13 through 7.15 rehearse the meeting: the managerial translation, the industry mirror, and the ethics of the verb “drive.”
7.2 Association, Causation, and the Meaning of “Driver”
Everything in this chapter rests on one distinction, and the guide has been preparing it since Chapter 2 named the anatomy of a decision. Two variables are associated when they move together: when knowing one tells you something about the likely value of the other. One variable cause another when intervening on the first changes the second — when the world after the intervention differs from the world without it. Association is a fact about observed data; causation is a claim about what would happen under an action not yet taken. Marketing decisions are actions, so marketing questions are, in their bones, causal — “should we deepen the discount ladder?” is a question about a world that does not yet exist. Marketing data, almost always, records associations. The gap between the question’s nature and the data’s nature is the standing condition of the marketing analyst’s work, and this chapter’s tools live entirely on the association side of the gap.
Definition
Association and Causation
Two variables are associated when their values move together in observed data, so that knowing one improves a guess about the other. A variable causes another when changing the first, with all else undisturbed, would change the second. Association is symmetric and observational; causation is directional and interventional. Observed association is consistent with causation in either direction, with a third variable influencing both, or with coincidence — and cannot, by itself, distinguish among them.
Source: Adapted from Simon (1954) and Freedman (1991).
In other words: the data can show that app users spend more; it cannot, alone, show that the app made them do it. The definition’s last sentence enumerates the rivals every association carries. The app–spend association could exist because the app increases spending (the platform team’s reading); because spending increases app adoption — frequent customers install the app to track orders they were already going to place (reverse causation); because a third variable, engagement with the brand, influences both installs and orders (confounding, Section 7.5’s subject); or, in small samples, because of nothing at all. Nothing in the number 2.4 favors any of these. Ruling rivals out requires either design — the experiments of Chapter 11, where random assignment severs every backdoor — or assumptions strong enough to substitute for design, which multiple regression will attempt in Section 7.8 and only partly achieve.
This is why the word “driver,” the word the VP used and the word this chapter’s title uses, deserves formal probation. In business language, a driver analysis is understood loosely: which variables move with the outcome strongly enough to matter. In causal language, a driver is a lever: push it and the outcome moves. The looseness is not harmless, because the loose sense is what the analysis delivers and the strict sense is what the budget meeting hears. This guide’s convention, which the deliverable in Section 7.13 will follow, is to use “driver” only with its probation officer present: a driver, in this chapter, is a variable that predicts the outcome in a specified model, and every sentence that promotes it to a lever must say what additional evidence performed the promotion.
None of this makes association worthless — the opposite. Association is evidence: it nominates candidate levers, sizes relationships, falsifies confident stories (a lever uncorrelated with the outcome has a steep hill to climb), and disciplines the prioritization of the expensive causal tests to come. Chapter 2’s evidence discipline applies: the analyst’s job is not to refuse conclusions but to match the confidence of the sentence to the strength of the evidence. The rest of this chapter builds the machinery for measuring association well; the standing distinction rides along in every section, and the ethics section will show it is the chapter’s real subject.
7.3 Covariance and Correlation: Putting a Number on Linear Association
Measurement starts, as it did in Chapter 3, with a definition decision: what number should summarize “moving together”? The natural starting point is covariance. Take two variables at the customer grain — frequency and monetary value from the feature table Chapter 6 built. For each customer, measure how far each variable sits from its own mean, and multiply the two deviations. A customer above average on both, or below average on both, contributes a positive product; a customer above on one and below on the other contributes a negative product. Covariance summarizes these products across customers: positive when the variables tend to sit on the same side of their means, negative when they sit on opposite sides, near zero when there is no consistent pattern.
One detail in that summary decides whether the printed number will match the definition, and this guide fixes it now rather than letting the software fix it silently. There are two covariances, and they differ in their denominator. The population covariance is the mean product of paired deviations across a full population: sum the products, divide by the number of members. The sample covariance estimates that quantity from a sample by dividing the sum of products by one fewer than the number of observations. The distinction is not decorative here, because the labs run on pandas, and pandas reports the sample version by default — so a chapter that defines covariance as “the average product” and then prints a number computed with the smaller denominator has taught two different quantities under one name.
Definition
Covariance
The population covariance of two variables is the mean product of their paired deviations from their respective means, positive when the variables tend to deviate in the same direction and negative when they tend to deviate in opposite directions. The sample covariance estimates it from n observations by dividing the sum of the paired deviation products by n − 1. Its magnitude depends on the units of both variables, which makes it a direction detector rather than a strength measure. In this chapter, pandas report the sample covariance and the sample correlation unless stated otherwise.
Source: Adapted from Rodgers and Nicewander (1988).
Covariance has the right sign and the wrong scale, and the reason echoes Section 6.7’s tyranny of units: the products are denominated in the units of both variables at once — order-dollars, in the frequency–monetary case — so the number’s size says as much about measurement units as about the relationship. Chapter 6 solved this problem for distance by standardization, and correlation is the same solution applied to covariance: divide by the spread of each variable. The result is the Pearson correlation coefficient, written r, and the division buys three properties that made it the lingua franca of association. It is unit-free: r between frequency and monetary is the same whether monetary is in dollars or cents. It is bounded: r always lies between −1 and +1, with the endpoints reached only when the points fall exactly on a line. And it is symmetric: r of X with Y equals r of Y with X — a property to remember, because “drivers” language is directional and r is not.
Definition
Pearson Correlation Coefficient (r)
The sample Pearson correlation coefficient is the sample covariance of two variables divided by the product of their two sample standard deviations. Because the same n − 1 denominator appears in the covariance and in each standard deviation, the correlation is unaffected by the choice of denominator even though the covariance is not. It measures the direction and strength of linear association on a fixed, unit-free scale from −1 (perfect negative linear association) through 0 (no linear association) to +1 (perfect positive linear association).
Source: Adapted from Pearson (1895) and Rodgers and Nicewander (1988).
That last property is worth one further sentence, because it explains why the sample-versus-population distinction matters for covariance and not for r. Dividing a covariance by two standard deviations cancels the denominator: whichever convention is used, provided it is used consistently in all three quantities, the ratio is identical. So, the covariance a student computes by hand can disagree with pandas in the third digit while the correlation agrees exactly — which is itself a small, memorable lesson about where conventions bite and where they do not, and Lab 7.1 arranges for the student to meet it.
Reading r takes calibration, and Table 7.1 supplies the working scale this guide uses — with the warning that context outranks any fixed ladder. In customer-level marketing data, where individual behavior is noisy by nature, correlations above 0.5 between distinct behaviors are strong findings and correlations near 0.9 should trigger suspicion rather than celebration: behaviorally distinct quantities rarely track each other that tightly, and near-perfect r usually means the two columns are the same fact wearing two names — monetary and its own components, a metric and its numerator, the arithmetic relatives Section 7.10 will flag for the model. The table is introduced here and consulted throughout the labs.
Table 7.1
Reading r in customer-level marketing data
Range of |r| | Working label | Reading in this guide | Standing caution |
|---|---|---|---|
0.0–0.1 | Negligible | No linear association worth a sentence | Could still hide a nonlinear pattern (Section 7.4) |
0.1–0.3 | Weak | Detectable; matters mainly in aggregate | Easily produced by a shared confounder |
0.3–0.5 | Moderate | A real behavioral relationship | Plot before believing; check outlier influence |
0.5–0.8 | Strong | Unusually tight for distinct behaviors | Ask what shared machinery produces it |
0.8–1.0 | Suspect | Rare between genuinely distinct quantities | Often two names for one fact; check derivations |
Three boundary facts complete the working picture. First, r is undefined when either variable has zero variance: the formula divides by both standard deviations, so a variable that never moves has no correlation with anything, and software will report NaN rather than zero. The practical form of this rule is a warning about near-misses — inside a narrow subgroup, a variable with almost no spread yields an r that is numerically defined and substantively meaningless, which Lab 7.1’s within-segment exercise is built to expose. Second, r is a customer-grain statistic here, and the grain declaration discipline of Section 3.3 applies with full force: the correlation between frequency and monetary across customers is a different fact from the correlation between daily orders and daily revenue across days, and Section 7.5 will show that the daily version imports the calendar as a hidden third variable. Third, r is a two-variable statistic, and modern practice rarely stops at one pair: the correlation matrix — every pairwise r among a set of variables, arranged in a grid — is the standard first exhibit of a drivers analysis, and it is also the exhibit most reliably misread. The matrix’s seductions are specific enough to deserve their own treatment, and they open the next section.
7.4 What Correlation Misses: The Always-Plot Discipline
The correlation coefficient compresses a full scatter of customers into one number, and compression loses information by design. This section catalogs what is lost — and installs the discipline that recovers it, which costs one line of code and has been part of this guide’s toolkit since Chapter 4’s exploratory charts: plot the data before, or at latest alongside, any correlation you intend to repeat in a meeting.
The canonical demonstration is Anscombe’s quartet: four small datasets constructed so that the means, variances, correlation (r = 0.816), and fitted regression line are essentially identical across all four — while the scatterplots show four utterly different worlds: a well-behaved linear cloud; a clean curve that is not a line at all; a perfect line with one outlier that has dragged the fitted slope; and a pillar of identical x-values propped into a “relationship” by a single distant point (Anscombe, 1973). Anscombe built the quartet to argue against the habit of computing without graphing, and half a century later the argument has sharpened, because the computing is now delegated: an AI assistant asked for a correlation matrix returns the numbers, not the shapes, and narrates the numbers fluently. The quartet’s lesson, restated for this guide: identical statistics do not mean identical evidence, and the scatterplot is the cheapest audit in analytics.
Three specific losses matter most on StyleCraft’s data, and each is a designed feature of the base the labs will meet. Nonlinearity first: r measures linear association only, and marketing relationships are routinely curved. Recency and future spending, engagement and response, tenure and frequency — each tends to rise or fall steeply in one region and flatten in another, and a straight-line summary of a curve understates the relationship where it is strong and overstates it where it is flat. An r near zero does not mean “no relationship”; it means “no linear one,” and the plot decides which.
Outlier sensitivity second: r is built from products of deviations, so a single customer far from both means contributes an enormous product — Chapter 5’s concentration analysis showed that StyleCraft’s top decile holds an outsized share of revenue, and those whales can single-handedly manufacture, inflate, or reverse a correlation among spend variables. The lab’s routine, executed in Code 7.5: compute r with and without the top spend decile, and if the story changes, the story was about the whales.
Third, and most important on this dataset: mixed populations. Chapter 6 established that the base contains four segments with different behavioral regimes. A correlation computed across the pooled base averages over regimes that may individually show different relationships — frequency and basket size can be positively related within every segment and negatively related across the pool, because the pool’s variation is dominated by the between-segment contrast of frequent small-basket urban customers against infrequent large-basket suburban ones. This is the correlational face of the aggregation lesson of Section 3.3.1 — the claim changes with the grain, and here, with the population sliced — and the working rule is the one Chapter 5 used for every summary: compute the pooled number, then compute it within segments, and treat disagreement between the two as a finding about structure, not a nuisance.
The correlation matrix concentrates all these hazards in one exhibit, plus two of its own. Table 7.2 collects the misreadings this guide’s students most reliably commit on their first matrix, with the corrections the sections of this chapter supply. The matrix remains the right first exhibit — it is the fastest wide-angle view of a feature table — but it is a screening instrument, and Table 7.2’s right-hand column is the protocol for everything it flags.
Table 7.2
Correlation-matrix misreadings and their corrections
Misreading | Why it is wrong | Correction |
|---|---|---|
“The biggest r is the biggest driver” | r is symmetric association, not directed influence; and the biggest r is often an arithmetic relative | Rank candidates only after Section 7.8’s model adjusts for the others; check derivations first |
“r ≈ 0, so this variable doesn’t matter” | Zero linear association can hide a strong curve or offsetting subgroup patterns | Plot it; compute within segments (Section 7.4) |
“These two drivers both correlate with spend, so both add value” | They may carry the same information — multicollinearity in waiting | Check their correlation with each other and their variance inflation factors; Section 7.10 |
“r = 0.9 — excellent, strong finding” | Near-perfect r between distinct behaviors usually flags a definitional overlap | Trace both columns’ derivation rules (Table 6.5) before celebrating |
“The matrix shows what to change to move spend” | Every cell is observational; no cell survives Section 7.2’s rivals | Treat the matrix as a nomination list for modeling and, later, testing |
One habit closes the section, borrowed whole from Chapter 1’s workflow: predict before you compute. Before running the matrix in Lab 7.1, the analyst writes the expected sign of every correlation with spend — recency negative, engagement positive, and so on — from marketing logic and from Chapter 5’s certified findings. The matrix then grades the predictions, and every surprise is an assignment: a plot, a derivation check, or a within-segment recomputation. Surprises that survive the assignment are findings. This is predict-then-verify (Section 1.7) applied to association, and it is the small-scale rehearsal of the audit this chapter’s AI section will formalize.
7.5 Confounding and Spurious Correlation
Section 7.2 listed the rivals behind every association; this section develops the one that does the most damage in marketing, because marketing’s operating environment manufactures it on a schedule. An association between X and Y is confounded when a third variable Z, present before either, influences both — creating or distorting a correlation between X and Y that partly or wholly reflects Z’s double influence rather than any relationship between X and Y themselves. When the observed association has no direct component at all — when it exists only through Z or through coincidence — it is called spurious: real as a number, empty as evidence (Simon, 1954).
The definition must be tighter than the one student usually carries, and the looseness is worth naming because it produces real errors. “A variable related to both” is too broad a test. A variable that sits on the pathway between X and Y — a mediator — is related to both, and adjusting for it removes exactly the part of the relationship the analysis was trying to see. A variable that is a common consequence of both — a collider — is also related to both, and adjusting for it can manufacture an association where none existed. This chapter does not develop causal diagrams; at working level the repair is to define the confounder by its position rather than by its correlations, and to treat a double correlation as a flag that starts an inquiry rather than as a verdict that ends one.
Definition
Confounding and Spurious Correlation
A confounder is a pre-treatment variable that influences, or helps determine, both the predictor or exposure and the outcome, creating or distorting their observed association. Being correlated with both variables is a warning sign, not by itself proof that a variable is a confounder: a variable on the pathway between them, or a common consequence of them, is also correlated with both and calls for different handling. An association is spurious when it reflects only the operation of confounders or chance, so that neither variable responds to the other. Detecting confounding requires variables and reasoning beyond the correlated pair itself; nothing in the pair’s own correlation reveals it.
Source: Adapted from Simon (1954) and Freedman (1991).
The definition’s closing sentence is the operational point: the correlation between X and Y, examined ever more closely, will never confess that Z exists. Confounders are found by thinking about the world, not by staring at the pair — which is why domain knowledge is a technical input to this chapter, not decoration around it. Public collections of absurd examples make the logic memorable — time-series pairs like per-capita cheese consumption tracking engineering doctorates, correlated at 0.9 and connected by nothing but parallel drift (Vigen, 2015) — but the absurd cases are the safe ones, because no one budgets on cheese. The dangerous confounders are the plausible ones, and marketing’s most reliable confounder is the calendar.
Here is the seasonality example this guide will reuse for the rest of Part II. Take StyleCraft’s daily grain: marketing activity by day and revenue by day, across the 24-month window. Email volume and revenue will correlate strongly and positively across days. The mechanism is only partly the one the email claim wants: StyleCraft, like every retailer, concentrates its marketing where demand already concentrates — November and December, the designed holiday lift of the dataset’s own revenue curve, plus the back-to-school swell. The calendar raises sends (campaign planners schedule into the season) and raises revenue (customers buy for the holidays regardless of email), and the daily correlation gratefully records the two effects as one. A naive read — “days with more email have more revenue; email drives revenue; send more” — is the confounded read, and its recommendation self-destructs in February. The correction pattern generalizes far beyond email: compare within season rather than across it (does extra email correlate with extra revenue among December days?), bring the confounder into the analysis as a variable — which is precisely the job Section 7.8’s multiple regression takes up — or change the design so the confounder cannot operate, which is Chapter 11’s business. Note what the example does to the email claim from the opening case: the claim’s customer-grain version (engaged customers buy more often) and its day-grain version (heavy-send days sell more) are different claims with different confounders, and neither, unexamined, supports the recommendation to double frequency.
Confounding also explains the finance claim’s trap, and Chapter 6 already built the machinery to see it. Discount exposure and total spend correlate negatively across the pooled base — the pattern is real and the labs will confirm it. But segment membership is associated with both discount exposure and total spend: the Discount-Dependent segment concentrates its purchases on promotion and spends less in total than the Urban Loyal Core, while the Core buys at full price and spends the most. The pooled correlation records that composition as if it were a consequence — the same between-group mechanism as Section 7.4’s mixed-population warning. The segment label earns its place here as a descriptive stratifier, because it exposes the customer composition sitting behind the pooled association. It does not automatically earn the stronger title. Chapter 6 estimated that label from same-window behavioral features that include monetary value and discount share — the very outcome and the very exposure this example is examining — so the label is not an independently measured pre-treatment variable. It is partly constructed from the behaviors whose relationship is under study and treating it as a clean confounder or a causal control would be reasoning in a circle: controlling for a grouping that was defined, in part, by the exposure and the outcome themselves. This is worth stating plainly because the reflex it corrects is common. An analyst who has just built a segmentation naturally reaches for it as the obvious control, and the reach is right about the arithmetic and wrong about the warrant.
What the segment label does support is a specific and useful list: exposing the composition behind a pooled number, describing heterogeneity, testing whether pooled and within-group patterns disagree, reproducing certified descriptive structure with a new instrument, and improving an explanatory model’s descriptive fit. Whether discounts reduce, leave unchanged, or increase any individual customer’s spending is simply not answered by the pooled r, and it is not answered by the stratified version either. The within-segment recomputation diagnoses how composition shapes the pooled association; the fitted model with segment dummies (Section 7.9) is a second descriptive adjustment, useful for examining heterogeneity and insufficient for causal identification. A genuine pre-treatment confounder — the calendar, tenure at the time of exposure, a demographic measured at signup — is a different kind of variable, and Section 7.8 returns to what controlling for one can and cannot buy. The finance claim may still be right. The point of this section is that its current evidence cannot tell.
What correlation cannot do, then, is settle any of the opening case’s three claims. What it has done, by the end of Lab 7.1, is discipline them: signs and strengths measured, shapes inspected, whales counted, segments separated, and the calendar named as a standing suspect. The next step is the instrument that examines relationships jointly instead of pairwise — and whose output format, the coefficient, is the one the meeting will quote.
7.6 Simple Linear Regression: The Line and Its Coefficients
Correlation answers “do these move together, and how tightly?” The budget meeting needs a different quantity: “if customers differ by one unit of this, how much do they differ in spend?” That question asks for a rate of exchange between variables, denominated in their native units — dollars per order, dollars per day of recency — and the instrument that estimates it is linear regression. The name and the idea are old: Galton (1886), studying heredity, noticed that extreme parents had fewer extreme children — values “regressing” toward the mean — and the line he drew through that cloud of points became the general tool for summarizing how one variable runs with another. This guide develops it as working machinery, one step above intuition and well short of formalism, per the program’s no-prerequisite design.
Simple linear regression models one outcome variable Y — the response — as a straight-line function of one predictor X, plus an error term that absorbs everything the line does not capture. The fitted model is an equation of the form Y = b₀ + b₁X, and its two coefficients are the deliverable. The intercept b₀ is the model’s estimate of Y when X is zero — sometimes meaningful (predicted spend at zero discount exposure), sometimes a mathematical anchor outside the data’s range (predicted spend at zero days of tenure describes a customer who signed up this morning), and the analyst states which. The slope b₁ is the working number: the estimated difference in Y associated with a one-unit difference in X. Its units are always outcome-units per predictor-unit, and this guide’s standing rule is that a coefficient is not read until its units are said aloud. A slope of −0.85 in a regression of monetary value on recency_days is not “negative 0.85”; it is “eighty-five cents less total spend per additional day since last purchase, across the observed range” — a sentence with dollars, a direction, a denominator, and a scope.
Definition
Simple Linear Regression
Simple linear regression models a numeric response variable as a linear function of one predictor: an intercept (the estimated response when the predictor is zero) plus a slope (the estimated difference in response per one-unit difference in the predictor), with an error term absorbing what the line does not capture. Coefficients are estimated from data and read in the native units of the variables.
Source: Adapted from James et al. (2021).
Where does the line come from? Any line through the scatter leaves gaps between itself and the actual customers, and each gap has a name: the residual, the customer’s actual Y minus the value the line predicts for their X. Positive residual, the model underestimated them; negative, it overestimated. The least squares criterion chooses the one line that makes the sum of squared residuals as small as possible — squared for the same reason inertia squared distances in Section 6.8: large misses are penalized disproportionately, and the arithmetic is computable. Uniqueness carries a condition worth attaching now, because Sections 7.9 and 7.10 will meet the exception. When the model’s predictor matrix has full column rank — no predictor is an exact linear combination of the others — least squares yield a unique coefficient vector. When predictors are perfectly redundant, as in the dummy-variable trap, software may still return a fitted result by numerical convention, but the individual coefficients are not uniquely identified and should not be interpreted. The intuition worth keeping is the balance-point image: the least squares line is the scatter’s center of gravity stretched into a slope, it passes through the point of means, and — like every mean since Chapter 5 — it is sensitive to extreme points, which is Anscombe’s third panel and the whales again. The formalism beyond this is machinery the software owns; what the analyst owns is the criterion and its sensitivities.
Definition
Residual
A residual is the difference between an observation’s actual response value and the value the fitted model predicts for it: actual minus predicted. Residuals are the model’s itemized account of what it missed, and their patterns — trends, curves, funnels, clusters — are diagnostic evidence about where the model’s form is wrong.
Source: Adapted from James et al. (2021).
Residuals deserve the second definition box because they are this chapter’s diagnostic instrument, the regression counterpart of Chapter 4’s verification log. Software can almost always return a fitted result; whether its coefficients are uniquely identified and substantively meaningful depends on the model’s design, and whether the line is an honest summary shows up in what it leaves behind. The residual plot — residuals against predicted values, or against the predictor — should look like weather: patternless scatter around zero. A curve in the residuals means the relationship was not a line (the model form is wrong, not the world); a funnel that widens with predicted spend means the model misses big spenders by more than small ones, a near-universal feature of revenue data that Chapter 5’s skew already predicted, and one with a direct consequence for the uncertainty columns that Section 7.10 will take up; a cluster of large residuals that turns out to be one segment means the single line is averaging over populations that want different lines — the mixed-population lesson again, now visible as diagnostic evidence and fixable with Section 7.9’s machinery. Reading residual plots is a chart-reading skill in exactly Chapter 4’s sense — the objects were introduced there; this chapter supplies what to look for when the chart is about a model.
Regression and correlation are close arithmetic relatives — for one predictor, the slope is r rescaled by the ratio of standard deviations, so they always agree in sign — but they answer different questions and fail differently. r is unit-free and symmetric: it grades the tightness of the cloud. The slope is unit-bearing and directional in form: it prices the relationship in decision units. That pricing is why the meeting quotes slopes, and why the analysts caution transfers whole: the regression of Y on X and the regression of X on Y are different lines answering different questions, and nothing in either line knows which variable is in charge — the model’s direction is the analyst’s framing, not a discovery. Every rival explanation from Section 7.2 survives the fitting of a line untouched.
7.7 How Good Is the Line? R², Fit, and Usefulness
A fitted line and honest units are not yet a report card. The standard single-number grade is R², the coefficient of determination: the share of the response variable’s variation that the model accounts for. The construction is worth seeing once in words. Total variation is the sum of squared deviations of Y around its own mean — how much there is to explain, the baseline ignorance of a model that predicts the mean for everyone. Residual variation is the sum of squared residuals — the ignorance that remains after the model. R² is the fraction of the original variation the model removed: R² = 1 minus (residual variation ÷ total variation).
Three properties are usually stated as though they were laws of arithmetic, and they are better stated with their conditions attached, because students meet software that reports R² under other conventions. For in-sample ordinary least squares models fitted with an intercept to the same observations, R² lies between 0 and 1 and cannot decrease when predictors are added to a nested model; and in a simple one-predictor model with an intercept, R² equals r². Remove the intercept and the arithmetic changes: statsmodels, for example, computes a centered R² when a constant is present and an uncentered version when it is absent, so the two are not comparable and the uncentered figure is routinely and misleadingly larger (Seabold & Perktold, 2010). Change the observations or evaluate the model on data it was not fitted to, and the guarantee that R² stays inside the unit interval disappears with it — a fact Chapter 8 will need immediately. The conditions are not fussiness. They are the difference between a number that means what the sentence says and a number that merely looks familiar.
Definition
R² (Coefficient of Determination)
R² is the proportion of a response variable’s total variation, measured around its mean, that a fitted model accounts for: one minus the ratio of residual variation to total variation. It compares the model to the baseline of predicting the mean for every observation. For in-sample ordinary least squares models fitted with an intercept to the same observations, R² lies between 0 and 1 and cannot decrease when predictors are added; in a simple one-predictor model with an intercept, it equals r². Because it cannot decrease, a rising R² is not, by itself, evidence of a better model.
Source: Adapted from Wooldridge (2020) and James et al. (2021); software conventions per Seabold and Perktold (2010).
Notice what R² secretly is: a baseline comparison, in exactly Section 2.7’s sense. “The model accounts for 34 percent of spend variation” means “against the interpretive baseline of guessing the mean, residual error shrank by a third.” This guide flags the connection because baselines are a running thread — Chapter 8 will operationalize them for prediction and Chapter 10 will extend them to time — and because remembering what R² a comparison is against inoculates against its two standing misreadings. The first misreading is that high R² certifies a good model. It does not: R² measures fit to this data, and Anscombe’s quartet holds four identical R² values over four different truths; a curve, an outlier, or a leaking variable can each buy fit without buying validity. In customer-level data the suspicion runs mostly the other way — an R² above roughly 0.9 on individual behavior should trigger the Table 7.1 reflex, because individual humans are not that predictable, and near-perfect fit usually means the model was handed the answer, the trap Code 7.18 demonstrates on purpose.
The second misreading is that low R² certifies a useless model. It does not: R² grades total variance accounted for, while decisions often turn on the slope. A model accounting for 12 percent of spend variation can still estimate the recency slope precisely enough to price a win-back program — noisy individuals, reliable rates. In other words: fit is a property of the model against this data; usefulness is a property of the model against a decision, and the two judgments can disagree in both directions.
Concept
Fit Is Not Usefulness
R² answers one narrow question — how much of the outcome’s variation this model absorbs in this sample. Decisions rarely ask that question. They ask whether a specific coefficient is estimated well enough to act on, whether the model’s form survives its residuals, and whether the relationship is stable enough to hold next quarter. A high-R² model can fail all three; a low-R² model can pass all three. Report R² — it is the honest disclosure of how much is not accounted for — but never let it stand alone as the verdict and treat “too good” with the same suspicion as “too poor.” The number that flatters the model is the number to audit first.
Source: Course concept developed for this guide, informed by Shmueli (2010).
7.8 Multiple Regression and the Ceteris Paribus Promise
The opening case’s claims all founder on the same rock: each compares customers who differ in many things at once and attributes the difference to one of them. The instrument built for exactly this problem is multiple regression: one response, several predictors, fitted jointly. The model form extends naturally — Y = b₀ + b₁X₁ + b₂X₂ + … — and least squares still choose the coefficients that minimize squared residuals. What changes is the meaning of each coefficient, and the change is the entire reason the method matters: in a multiple regression, b₁ estimates the difference in Y per one-unit difference in X₁ among observations that share the same values of every other predictor in the model. The Latin tag is ceteris paribus — other things equal — and the operation is statistical control (Wooldridge, 2020).
One word must be handled carefully here, because the natural English for statistical control is misleading. Textbook prose and classroom shorthand both reach for “matched” or “comparable customers,” and ordinary least squares does not match anybody. It fits a surface to the whole cloud and reports the surface’s slope in one direction. The estimate is therefore a model-based comparison — an adjusted difference computed under the model’s assumed functional form — and it can be produced even where few or no real customers share the relevant combination of covariate values. Where that happens, the model is extrapolating across a region the data do not populate rather than comparing genuinely similar customers, and nothing in the printed coefficient distinguishes the two cases. The honest phrasing, used throughout this chapter, is that a coefficient compares customers statistically, holding the included predictors constant within the model’s assumed form.
Definition
Multiple Regression
Multiple regression models a numeric response as a linear function of two or more predictors fitted jointly, so that each coefficient estimates the difference in the response per one-unit difference in its predictor, holding the other included predictors constant within the model’s assumed functional form. It implements statistical control by adjustment rather than by matching: the comparison is model-based, extends only to the variables the model was given, and may extrapolate where the data contain few observations sharing the relevant combination of values.
Source: Adapted from Wooldridge (2020).
Applied to the app claim, the machinery earns its keep immediately. The raw comparison — app users versus non-users, a 2.4× spend ratio — mixes the app difference with every difference that travels with it: app users are younger, more urban, more loyalty-enrolled, longer-engaged, more email-active. Put spend on the left and app ownership on the right alongside recency, tenure, email opt-in, and segment membership, and the app coefficient now estimates the spend difference between app owners and non-owners adjusted for all of those. On StyleCraft’s data the labs will find that this adjusted difference is far smaller than the raw ratio implied — most of “2.4×” was the Urban Loyal Core’s profile traveling with the app, not the app itself. That shrinkage from raw gap to adjusted gap is the most useful single exhibit a drivers analysis produces, and the meeting-ready sentence practically writes itself: “customers with the app spend more; most of that gap reflects who installs the app rather than anything the model can attribute to installing it.”
The promise, however, has strict boundaries, and the guide states them as sharply as the promise itself. Ceteris paribus means other included things equal. The model holds constant what it was given and nothing else; every relevant variable not in the model remains free to vary between the compared customers and to smuggle its influence into the coefficients. This failure has a name — omitted variable bias — and a usable direction logic, provided the logic is stated with its scope. In the simplest linear case, one included predictor and one omitted variable, the direction of the bias follows the product of two relationships: the omitted variable’s effect on the response, times its association with the included predictor. In a multiple regression the relevant association is narrower: it is the part of the omitted variable’s association with the included predictor that remains after the model’s other predictors are considered. That qualification matters in practice, because a variable that looks strongly related to the predictor in a raw cross-tabulation can be nearly unrelated to it once the model’s other controls are in place, and the bias reasoning must follow the adjusted relationship rather than the raw one.
Definition
Omitted Variable Bias
Omitted variable bias is the distortion of a fitted coefficient that arises when a variable related to both the response and an included predictor is left out of the model: the included predictor absorbs credit or blame belonging to the omitted one. In the simplest linear case, the direction of the bias follows the product of the omitted variable’s effect on the response and its association with the included predictor. In a multiple regression, the relevant association is the part that remains after the model’s other predictors are considered.
Source: Adapted from Wooldridge (2020).
If underlying brand engagement increases spend and is positively associated with email activity, then omitting it would tend to bias the email coefficient upward: under those assumptions the fitted model attributes to email part of an association that belongs to enthusiasm. Note the shape of that sentence. The direction follows from stated assumptions about the omitted variable, not from a causal fact the data have established — which is exactly what makes it usable, since the reasoning can be run before the missing variable is ever measured. The direction logic converts anxiety into analysis: for every coefficient the meeting will quote, the analyst lists the plausible omitted variables, reasons the sign of each bias, and reports which way the estimate leans. Statistical control is not a causal machine; it is a better class of comparison, and the honest sentence names both halves.
Two practical notes complete the working level. First, model inputs are governed by measurement, exactly per Section 3.4: numeric predictors enter as they are; ordinal predictors (loyalty tier) enter either as ordered scores — asserting equal steps, an assertion to declare — or as categories per Section 7.9; nominal predictors (segment, metro, channel) must enter as dummy variables, because arithmetic on category codes is the level-of-measurement error Chapter 3 warned against. Second, adding predictors is not free: each addition changes every other coefficient’s meaning (“holding constant” now includes the newcomer), R² rises mechanically whether the addition carries information, and predictors that share information destabilize one another — the multicollinearity of Section 7.10. The model is a specified instrument, not a bucket; every variable in it should be there because the analysis’s question, written per Section 2.5’s specification discipline, put it there.
7.9 Dummy Variables, Reference Categories, and Interaction Terms
The drivers question at StyleCraft is inseparable from the segments, and the segments are categories: four names, no natural ordering, no units. Regression’s arithmetic needs numbers, and the encoding that supplies them honestly is the dummy variable: an indicator column equal to 1 when the observation belongs to a category and 0 otherwise. A categorical variable with c categories enters the model as c − 1 dummies; the category left without a dummy is the reference category, and it is not omitted from the analysis — it becomes the analysis’s baseline. Every dummy coefficient is then read as an adjusted difference against the reference: with Urban Loyal Core as the reference, the Suburban Occasion dummy’s coefficient is the estimated difference in the response between Suburban Occasion customers and Urban Loyal Core customers, holding the model’s other variables constant. One category must always be left out — including all c dummies alongside an intercept gives the model two ways to say the same thing, a perfect redundancy that breaks the fitting (the dummy variable trap, multicollinearity in its terminal form).
What the intercept means in such a model need stating precisely, because the convenient shorthand is wrong as soon as the model contains anything else. It is tempting to say that the reference category’s average outcome is carried by the intercept. That is true only in the special case of a model containing nothing but that one set of dummies. Add a numeric predictor and the intercept become the model’s expected response for the reference category when that numeric predictor equals zero — which may describe no customer who exists. Add a second categorical predictor and the intercept refer to the reference level of both. The reference groups observed mean and the model’s intercept are different quantities and confusing them is how an analyst ends up reporting a negative baseline spend with a straight face.
Definition
Dummy Variable and Reference Category
A dummy variable encodes category membership as an indicator: 1 if the observation is in the category, 0 otherwise. A categorical predictor with c categories enters a regression as c − 1 dummies, and the reference category is represented when all dummies for that predictor equal zero. With an intercept in the model, each dummy coefficient is an adjusted difference from that reference category. The intercept represents the model’s expected response for the reference category when all other predictors are at their coded zero or reference values — which is the reference group’s observed mean only in the special case of a model containing no other predictors.
Source: Adapted from Wooldridge (2020).
The choice of reference category changes no substance — every model with any reference fits identically and tells the same story — but it changes every printed number and therefore every sentence, so it is chosen for readability and declared. The guide’s convention: the reference is the largest or most decision-central category, so that coefficients read as “compared to our core.” Table 7.3 shows the encoding for StyleCraft’s four segments with Urban Loyal Core as reference, and the reading column is the deliverable skill: students who can compute dummies and cannot read them against the reference produce the most common wrong sentence in applied regression — “the model shows Suburban Occasion customers spend $X” instead of “$X more or less than the Core, adjusted for the model’s controls.”
Table 7.3
Dummy encoding of the four segments, with Urban Loyal Core as the reference category
Segment (from Chapter 6) | seg_ | seg_ | seg_new | How the coefficient reads |
|---|---|---|---|---|
Urban Loyal Core | 0 | 0 | 0 | Reference: no coefficient of its own. The model’s expected response for this group appears only when every other predictor is also at zero or at its own reference level. |
Suburban Occasion | 1 | 0 | 0 | Adjusted difference from the Urban Loyal Core, holding the model’s other predictors constant |
Discount-Dependent | 0 | 1 | 0 | Adjusted difference from the Urban Loyal Core, holding the model’s other predictors constant |
New / At-Risk | 0 | 0 | 1 | Adjusted difference from the Urban Loyal Core, holding the model’s other predictors constant |
Dummies let group membership shift the model’s level. The next construction lets it shift the model’s slopes — because the sharpest drivers questions are not “does this lever matter?” but “for whom does the relationship differ?” An interaction term is a new predictor built as the product of two others, and its coefficient estimates how the association between one variable and the response differs across values of the other.
Interactions must be read conditionally, and the reading is mechanical enough to state as a rule. Suppose the model contains app use, loyalty enrollment, and their product, both coded 0 and 1. Then the app coefficient is the app–spend association among customers whose loyalty enrollment equals zero — the non-enrolled. The loyalty coefficient is the loyalty–spend association among customers whose app use equals zero. And the interaction coefficient is the difference between the app association among the enrolled and the app association among the non-enrolled: the app association for enrolled customers is the sum of the app coefficient and the interaction coefficient. Nothing in the model reports an “overall” app association once the interaction is present, and an analyst who quotes the app coefficient as though it were one has quoted a conditional number as an unconditional claim.
Definition
Interaction Term
An interaction term is a predictor formed as the product of two others, whose coefficient estimates how the association between one predictor and the response differs across values of the other. A model with an interaction no longer has one slope per variable; it has slopes that depend on context. Each main-effect coefficient is the association for its variable when the interacting partner equals zero, and the association at any other value of the partner is obtained by adding the interaction coefficient the appropriate number of times.
Source: Adapted from Wooldridge (2020) and James et al. (2021).
The registry example is the one the budget meeting cares about, and this guide names it one way and holds to it: loyalty enrollment × app-use status, with app use serving as the selected channel-engagement indicator. The dataset’s specification plants a loyalty-by-channel structure, and app use is the channel variable StyleCraft’s schema carries at the customer grain, so the model estimates the planted structure through that indicator rather than through a channel construct the table does not contain. The question the interaction answers is whether the association between app engagement and spend differs by loyalty enrollment — and the honest report of the answer is two conditional associations in dollars, not one significance verdict.
Two standing cautions travel with every interaction. The main-effect coefficients change meaning the moment their interaction enters, which is the conditional reading just given, and it is the single most common source of misquotation in applied regression. And interactions multiply the model’s appetite for data: estimating a separate association per group requires enough customers in each combination of the two variables, a substantiality check Chapter 6’s segment sizes already supplied and Code 7.15 re-checks by printing the cell counts before the model is fitted. Where an interaction is materially nonzero, the finding to write is descriptive and conditional — “the association between discount exposure and spend differs by loyalty status” — rather than the causal restatement that a slide will reach for, and the exercises test the difference.
7.10 Reading Regression Output as Evidence: p-Values, Confidence Intervals, and Multicollinearity
A fitted model arrives as a table: one row per coefficient, and columns the software prints whether anyone can read them. This section supplies the working-level reading of the columns that matter after the estimate itself — the uncertainty columns — and the two structural judgments that govern how much any of it means: whether the predictors undermined each other, and what job the model was built to do. Everything here is calibrated to the analyst’s seat: enough to weigh evidence and defend a table in a meeting, on the explicit understanding that the full statistical machinery belongs to later coursework.
Every coefficient is an estimate from one sample of customers and would come out somewhat differently on another sample — the fitted $41 per order is the center of a range of values the data cannot distinguish, not a constant of nature. The standard error measures that sampling wobble, and the two working instruments are built from it. The 95 percent confidence interval is the honest way to report a coefficient: the range of values compatible with the data at conventional confidence — “$34 to $48 per order” — whose width is the finding about precision, and whose reading follows Chapter 5’s spread discipline: a wide interval around a large estimate says, “big but poorly pinned down,” and a narrow interval around a small one says, “small and we know it.”
Before either instrument can be trusted, one assumption behind the printed standard errors must be checked, and Section 7.6 already predicted its failure. Ordinary least squares compute standard errors, by default, on the assumption that the residual variance is the same across the range of fitted values. Revenue data rarely obliges: the residual plot funnels, because the model misses big spenders by more than small ones. That is heteroskedasticity, and its consequence is specific — it does not bias the fitted coefficients, and it does distort the standard errors, so the intervals and p-values built from them are wrong in a direction the default output does not disclose. The repair is a different covariance estimator, requested at fit time. Statsmodels defaults to a non-robust estimator and supports several robust alternatives; this guide uses HC3, which performs well in samples of the size marketing analysts work with (MacKinnon & White, 1985; Seabold & Perktold, 2010). A chapter that teaches students to expect a residual funnel and then quotes non-robust intervals two pages later would be teaching the diagnosis and skipping the treatment, so every model in Lab 7.2 is fitted with robust standard errors and Code 7.9 states plainly what changes and what does not: the coefficients and the predictions are identical, and the standard errors, confidence intervals, and p-values are not.
The coefficient’s p-value answers a narrower question than either instrument above, or it is worth stating exactly, because the loose version students carry — “the probability of getting an estimate this far from zero by chance” — omits the machinery that produces it. The software computes a test statistic from the coefficient and its estimated standard error and refers that statistic to a distribution implied by the null hypothesis and the model’s assumptions. Small p-values say the data are hard to reconcile with “no association, given this model”; they say nothing about size, importance, or cause. The American Statistical Association’s consensus statement, issued after decades of misuse, compresses to three warnings this guide adopts: a p-value is not the probability that a hypothesis is true, statistical significance is not practical significance, and no single p-value should by itself determine a scientific, business, or policy conclusion (Wasserstein & Lazar, 2016). On eight thousand customers, near-trivial associations clear 0.05 easily — at this sample size, significance is table stakes, and the analyst’s questions are the interval’s location and width in decision units. The asterisk column is where reading starts, not where it ends.
Definition
Coefficient p-Value and Confidence Interval
A p-value is the probability, assuming the stated null hypothesis and statistical model are correct, of obtaining a test statistic at least as extreme as the one observed. For a regression coefficient, the null hypothesis is that the coefficient is zero given the model’s other variables, and the test statistic is computed from the estimate and its standard error — so the p-value inherits every assumption behind that standard error, including the assumption about residual variance that robust estimators relax. A 95 percent confidence interval is the range of coefficient values compatible with the data at conventional confidence; its center is the estimate, and its width is the precision. Together they grade the strength of evidence for a relationship, never the relationship’s causal status.
Source: Adapted from Wasserstein and Lazar (2016) and Wooldridge (2020).
The first structural judgment is multicollinearity: predictors that carry overlapping information. When two variables move together tightly — frequency and monetary; email opt-in and app use; tenure and recency in a base that grew in waves — the model struggles to apportion the shared movement between them, and the symptoms are characteristic: coefficients with large standard errors and unstable values that swing when a variable enters or leaves; individually “insignificant” predictors in a model whose joint fit is strong; occasionally a sign that flips against all descriptive evidence.
The predictor correlation matrix is the first diagnostic, read with Table 7.1’s “suspect” row, and it is not sufficient by itself. Pairwise correlations detect only pairwise overlap, and the damaging case is often multi-variable: one predictor is closely approximated by a combination of several others while no single pairwise correlation looks extreme. The instrument that sees this is the variance inflation factor, which measures how much a coefficient’s estimated variance is inflated by the predictor’s relationships with all the other predictors together (Seabold & Perktold, 2010). This guide declines to publish a universal cutoff, because none survives contact with the literature. VIF is a diagnostic rather than a verdict: large values indicate that the model has difficulty separating a predictor’s coefficient from the information carried by the others, and what counts as large depends on how precisely the coefficient needs to be estimated for the decision at hand. The paired instrument is the instability test — refit with and without the twin and watch the survivor’s coefficient — which Code 7.17 runs as a specification-sensitivity check.
Definition
Multicollinearity
Multicollinearity is strong overlap in the information carried by a model’s predictors — whether between a pair or between one predictor and a combination of several others — which leaves their joint explanatory contribution largely intact but makes the division of that contribution among individual coefficients unstable and imprecise: inflated standard errors, estimate swings across specifications, and occasionally reversed signs. It is diagnosed from predictor correlations, variance inflation factors, and specification-sensitivity checks, and addressed by combining, dropping, or jointly reporting the entangled predictors. Multicollinearity may leave in-sample fit largely intact while destabilizing coefficients; it can also weaken out-of-sample prediction when the relationship among predictors shifts.
Source: Adapted from Wooldridge (2020); variance inflation factor per Seabold and Perktold (2010).
The response to a multicollinearity finding is a decision, not a formula: combine the twins into one variable, drop one and say so, or keep both and report the pair’s joint story while refusing to interpret their individual split. And one special case is a defect rather than a judgment call: a predictor that is arithmetically derived from the response produces a near-perfect fit that measures the derivation, not behavior. Regressing monetary value on the product of average order value and frequency is the pure form of the error, because that product is monetary by construction (Table 6.5), and Code 7.18 demonstrates it deliberately so that students meet the sensation of a model that is too good. Chapter 8 will name the general disease leakage; here it is caught by the derivation-tracing habit Table 7.2 installed.
The second structural judgment sorts everything this chapter has built, and it is the distinction the rest of Part II stands on: prediction and explanation are different modeling goals, and the same equation serves them differently (Shmueli, 2010). A predictive model has one obligation — accurate outputs on customers it has not seen — and it tolerates much that an explanatory model cannot: entangled predictors, uninterpretable coefficients, variables that merely correlate with the future. An explanatory model has a different obligation — coefficients an analyst can defend as honest structure — and it tolerates lost accuracy in exchange: variables chosen by the specification’s question, collinear twins separated or set aside, every retained coefficient readable in units with its bias direction reasoned. The drivers analysis of this chapter is explanatory work; the spend model of Chapter 8 will be predictive work; and most applied failures of regression trace to answering one goal’s question with the other goal’s model — quoting individual coefficients out of a kitchen-sink predictive fit, or crippling a forecasting model to keep its coefficients tidy. Table 7.4 sets the comparison.
Table 7.4
Prediction and explanation as distinct modeling goals
Dimension | Explanation (this chapter) | Prediction (Chapter 8) |
|---|---|---|
Question | How do these variables relate to the outcome? | What will the outcome be for new cases? |
Success measure | Defensible coefficients: signs, units, bias reasoning | Out-of-sample accuracy against a baseline |
Variable selection | By the specification’s question | By predictive contribution |
Multicollinearity | A serious problem — it muddles attribution among coefficients | Often tolerable in-sample, but a risk when predictor relationships shift out of sample |
Coefficient reading | The deliverable | Optional; often declined |
Characteristic failure | Causal overclaiming (Section 7.15) | Leakage and overfitting (Chapter 8) |
7.11 AI as a Regression Assistant
Chapter 6 marked the point where AI assistance began delivering whole analytic acts; this chapter is where the acts start arriving in the language of decisions. Prompted with the feature table and “what drives spend?”, a current assistant will select variables, fit models, format the coefficient table, and — this is the part that changes the analyst’s job — narrate the results in fluent managerial prose, complete with recommendations. The narration is the hazard. Code errors announce themselves; a regression that runs, fits, and is narrated with confident causal language fails silently, and it fails in the exact register the meeting will quote. The division of labor therefore follows the pattern set in Section 6.12, with this chapter’s twist: the assistant may draft the machinery, and the analyst audits not only the numbers but the verbs.
Where the assistant genuinely helps, use it. Drafting the modeling code — formula syntax, dummy encoding, the residual-plot boilerplate, the tidy coefficient table with intervals — is legitimate delegation of mechanics this chapter has specified, provided the variable list comes from your written specification rather than the assistant’s initiative. Generating the audit exhibits themselves — predicted-sign checklists, with-and-without-whales correlation runs, specification-sensitivity refits — is exactly the tedious-but-mechanical work assistants accelerate. Translating a finished, verified coefficient table into plain-language draft sentences is useful, provided every sentence is then checked against the verb discipline of Section 7.2. And the assistant is a genuinely good brainstorming partner for candidate confounders — “what variables might influence both discount exposure and total spend?” — because enumerating rivals is a breadth task, and breadth is what it has.
What must never be delegated are the judgments, and the failure modes are characteristic enough to name. The confident narration: asked for drivers, assistants return causes — “app adoption drives a $212 spend lift” — because fluent causal prose saturates their training data; the arithmetic may be flawless while the verb is unearned. The silent specification: the assistant chose which variables entered, which were dropped, what the reference category was, and whether an interaction was included — choices that change every number — and reported none of them as choices. The plausible-but-impossible coefficient: a sign or magnitude that contradicts certified facts (a positive recency coefficient; an intercept implying negative spend; a segment gap triple anything Chapter 5 measured), delivered with the same serene formatting as a correct one. The units slip: narrating a coefficient without units, or with the wrong ones — “recency matters most because its coefficient is largest” when recency is in days and frequency is in orders, the standardization confusion of Section 6.7 reborn in a coefficient table. The default covariance estimator: a model fitted on visibly funneling residuals and reported with non-robust intervals, because nothing in the default workflow connects the diagnostic to the inference. The diagnostics skip: no residual plot, no mention of the funnel or the curve, because the default workflow fits and reports but does not examine. The unconditional reading of a conditional coefficient: a main effect quoted as an overall association while an interaction sits in the same model. And the significance theater: p-values below 0.05 presented as importance, exactly the misreading this section retired.
The audit routine that answers them is this chapter’s verification theme, compact enough to run on every AI-fitted model and every colleague’s model — four points, in order, before any narration is accepted. Table 7.5 states it as a working checklist; the AI in Practice box operationalizes it as a prompt-and-audit pattern.
Table 7.5
The four-point audit of a fitted regression
Point | The check | Fails when |
|---|---|---|
1. Signs | Every coefficient’s direction against written predictions and certified descriptives (Chapter 5) | A sign contradicts known structure with no confounding story offered |
2. Magnitudes | Every coefficient’s size against hand-computable anchors: group means, Chapter 5 gaps, plausible ranges | A coefficient implies differences no descriptive table shows, in either direction, with no explanation grounded in the model’s variables |
3. Units | Every coefficient restated aloud in outcome-units per predictor-unit; comparisons only on common scales; conditional coefficients read conditionally | A narration compares raw coefficients across different units, or quotes a main effect from a model containing its interaction as though it were unconditional |
4. Diagnostics | Residual plots inspected; the covariance estimator matched to what the residuals show; variance inflation factors read; R² read against Section 7.7; specification sensitivity spot-checked | Fit is reported but residuals were never looked at, or intervals are quoted from a non-robust fit on funneling residuals |
AI in Practice
The Regression That Must Survive Its Audit
A two-prompt pattern with the audit between. Prompt one, the specification: paste the variable list with derivation rules (Table 6.5), the declared response, the reference category, the covariance estimator, and the exact model formula from your written spec, then ask for mechanics only — “Fit this model as specified with HC3 standard errors; report the coefficient table with 95 percent intervals in native units, the residual-versus-fitted plot, variance inflation factors, and R²; change nothing in the specification; narrate nothing yet.”
Then run the four-point audit of Table 7.5 yourself: check every sign against the predictions you wrote before fitting; check magnitudes against Chapter 5’s certified gaps; say every coefficient in its units and read any coefficient whose interaction is in the model as the conditional quantity it is; read the residual plot and confirm the covariance estimator matches what it shows.
Prompt two, only after the model survives: “Draft three sentences reporting these estimates as associations, using 'is associated with’ or 'predicts,' naming units and the held-constant list, and making no causal claim.” Audit the verbs in the reply as strictly as the numbers — assistants reintroduce causal language at the narration step even when the numbers are clean.
Source: Course concept developed for this guide, informed by Wasserstein and Lazar (2016) and Shmueli (2010).
The deeper continuity: in Chapter 6 the assistant proposed structure and the analyst tested whether the structure was real; here the assistant proposes relationships and the analyst tests whether the relationships are honestly stated. The labs that follow are built so the audit has teeth. The dataset’s specification planted the coefficient structure the labs recover — which variables carry spend, at what strengths, with one designed interaction — so an AI-fitted model’s claims can be graded against a known answer, and Exercise 7.6 does exactly that with a deliberately underspecified prompt.
7.12 Marketing Interpretation and Managerial Insight
The lab’s outputs are coefficient tables; the budget lock runs on sentences. This section translates — and, per this guide’s standing practice, it does so partly by exhibiting a wrong managerial reading and correcting it, because this chapter’s characteristic misreading is the most expensive one in the guide so far: a misread summary misinforms a meeting, a misread segmentation runs for a season, but a misread driver reallocates a budget and then defends the reallocation with the very analysis that caused it.
The wrong reading will be offered in the allocation meeting by someone reasonable, probably while pointing at your own Code 7.12 table: “The model says app users spend $140 more than comparable non-users. We have several thousand customers without the app. Fund the install campaign — at $140 a head, it pays for itself many times over.” Every number in the sentence can be exactly right, and the conclusion does not follow. The coefficient is a model-adjusted comparison between existing app users and non-users; it is not a forecast of what happens when a non-user is persuaded to install. The customers who chose the app differ from those who did not in everything the model cannot see — enthusiasm for the brand, planned future purchases, the underlying engagement that Section 7.8’s omitted-variable logic says is inflating the coefficient right now. Installing the app does not install the enthusiasm.
The corrected reading, which the deliverable should print under the table: “Customers with the app spend $140 more than otherwise similar customers without it, adjusting for recency, tenure, email engagement, discount exposure, and segment. Some of that difference may reflect what the app enables; some may reflect selection into app use, and this analysis cannot split the two. The raw gap the platform deck quoted was several times larger, and the model attributes most of it to profile rather than to the product. Across four defensible specifications the estimate ranges from $X to $Y, which is the range the decision should carry. If the install campaign is funded, it should be funded as a test — Chapter 11’s instrument, not this chapter’s — with a holdout that would tell us, one quarter from now, what a caused install is worth.” That paragraph — association stated, rivals named, sensitivity disclosed, decisive evidence specified — is what an honest drivers deliverable sounds like, and it is a stronger meeting position than the overclaim, because it cannot be dismantled by the CFO’s first question.
The same discipline settles the other two claims, and the deliverable should give each its verdict paragraph. The discount claim dissolves under the lab’s within-segment exhibit from Code 7.7 and the faceted plot from Code 7.6: the pooled negative association reflects which customers use discounts more than what discounts do, and discount exposure is more strongly associated with the Discount-Dependent segment’s purchasing pattern than with any change in it. Cutting the ladder on this evidence would withdraw the lever most visibly associated with a quarter of the treatment map’s purchasing, on a correlation that composition substantially accounts for, and without any test of how that quarter responds. The recommendation that survives hold the ladder, and test depth variations within the Discount-Dependent journey, where the lever’s marginal effect is the live question.
The email claim earns the mildest verdict, and it is now a verdict the chapter has earned rather than assumed. Code 7.16 estimated the email association separately in each segment. Report what it found in the deliverable’s own words: email engagement is positively associated with spend in every segment, in the specified model, with the four conditional estimates falling in a band the deliverable should print; the joint test of whether those associations differ is reported alongside them; and a large p-value on that test does not establish that the association is identical across segments, because failing to detect a difference is not evidence of sameness, especially where the smaller segments carry the widest intervals. What the chapter cannot supply from any of this is a price for additional sends — the day-grain version of the claim was disqualified by the calendar in Section 7.5, and no model here varies send frequency at all. Doubling frequency on this evidence would be paying twice for the same association; the honest sentence funds the content of the engaged journeys, not the volume, pending a frequency test.
The average-order-value finding gets a fourth paragraph and its own careful verb. After adjusting for units per order and occasionwear share, the estimated suburban difference becomes much smaller — these two variables statistically account for most of the observed premium in this dataset. Say it that way. The tempting version, that the premium runs through composition, asserts a pathway, and a pathway is a causal structure that two fitted models cannot establish. The adjustment claim is exactly as useful to the meeting, agrees with Chapter 5’s certified decomposition, and does not require a retraction when someone asks how it was tested.
Two further translation disciplines round out the deliverable. First, magnitudes travel in decision units and with their intervals: “$0.85 less per day of recency, interval $0.70 to $1.00” prices a win-back trigger’s timing; “recency was highly significant” prices nothing. Every quoted coefficient carries units, direction, the held-constant list in one clause, and its interval — the four-element metric discipline of Section 3.5.1, reborn for model outputs — and where the interval was computed with robust standard errors, which in this chapter is everywhere, the deliverable says so once in a methods line rather than in every sentence.
Second, the model’s silences are reported alongside its statements: the drivers model says nothing about levers not in the data (creative quality, price perception, service), nothing about individuals (these are adjusted differences averaged over customers), nothing about combinations of predictor values the base barely contains, and nothing about the future. The one-page deliverable for the VP therefore has a fixed anatomy: the four verdict paragraphs; the coefficient table with intervals, in units; the composition exhibits that discipline the discount claim; the two conditional app associations from the interaction, in dollars, because that is the finding the four journeys can act on this fall; the specification-sensitivity range for any coefficient the recommendation leans on; and a closing line that names the next instrument — the questions this analysis cannot answer are precisely the ones an experiment can, and the cheapest of those experiments should be in the field before the next budget cycle. Drivers analysis, honestly delivered, does not end arguments; it upgrades them.
7.13 Business Analytics in Practice
This section turns from the fictional case to three ways regression is used — and misread — in professional marketing analytics, where it is at once the most institutionalized method in the field and the one whose outputs are most routinely pushed past what they can bear. The second and third vignettes are composites of recurring professional patterns rather than reports about a single named organization.
7.13.1 Marketing Mix Modeling
Marketing mix modeling is regression’s flagship deployment and one of the oldest continuous analytics practices in the industry. It descends directly from the marketing-mix concept (Borden, 1964): regress a brand’s sales, usually weekly and by market, on its marketing inputs — media spend by channel, price, promotion, distribution — plus the confounders this chapter taught you to name first: seasonality, holidays, competitor activity, macro conditions. The fitted coefficients become the currency of budget season, translated into contribution charts and response curves that answer, channel by channel, what a marginal dollar returned. For decades the method lived mostly at large consumer-goods firms with the data and patience for it; the privacy era revived it broadly, because as cookie deprecation, platform walled gardens, and tracking limits degraded user-level attribution, an approach that needs no individual tracking — aggregate spend in, aggregate sales out — became newly attractive.
Practitioners at the center of that revival have been candid about the method’s fragilities, and the list reads like this chapter’s table of contents at industrial scale: limited variation in the spend data, strong correlation among channels that are planned together, selection effects when spend chases demand, and results acutely sensitive to specification (Chan & Perry, 2017). Three of those deserve unpacking, because they are the difference between a model that is fitted and a model that is governed.
The first is that the specification is governed rather than merely chosen. In a mature practice the variable list, the functional forms for diminishing returns and carryover, the treatment of price and distribution, and the set of control variables are agreed in advance, documented, and changed through a review process — because a coefficient that moves when a control is added is a coefficient whose value is partly an artifact of a modeling choice, and an ungoverned specification lets that choice be made, unrecorded, by whoever ran the model last. The second is the specific fragility of the time controls. Seasonality, holiday timing, and the promotion calendar are the confounders that matter most, and they are also the ones most entangled with the media plan, since campaigns are scheduled into exactly the weeks when demand is highest; a model with too coarse a time control credits the season to the media, and a model with too fine a time control absorbs the media effect into the calendar. Neither error announces itself in the fit statistics. The third is collinearity among channels: media budgets move together, up in the fourth quarter and down in the first, so the model is asked to separate contributions that the data present as a bundle — Section 7.10’s problem, with a budget attached to its answer.
One distinction governs how the output may be used, and it is this chapter’s distinction wearing an industrial suit. A marketing mix model used for historical attribution answers a descriptive question: given what we spent and what happened, how does the model apportion the sales we observed? A claim that moving next year’s budget from one channel to another will cause a particular return is a different claim, resting on the assumption that the fitted relationship would survive an intervention that changes the very spend pattern the model was estimated on. Careful teams state which of the two they are doing, and increasingly they pay for the upgrade the same way this chapter recommends: deliberate spend experiments — geographic holdouts, staged rollouts — whose results calibrate and correct the regression. The practice lesson: the world’s most consequential marketing regressions survive not because the fitting is clever but because the checking is institutional.
7.13.2 Confounding in the Wild
The second vignette is a composite of a pattern retail analytics teams meet regularly. A national retailer’s team estimated discount elasticity from two years of transaction data: how much does volume respond when price is cut? The first model said, absurdly, that discounts barely mattered — and a specification review found the reason. Promotions clustered in December, when volume would have surged anyway; and the deepest discounts were routinely applied to the slowest merchandise, so discount depth was correlated with weak demand by managerial construction. The calendar pushed the estimate up; the clearance policy pushed it down; the two biases nearly canceled into a number that looked precise and meant nothing.
The repaired analysis compared like with like — within-season, within-category, promotion versus non-promotion weeks for the same items — and the elasticity that emerged was several times larger and finally decision-grade. The lesson is Section 7.5 operating at line-item scale: before trusting any observational coefficient, ask who chose the variable’s values and why — because in business data the treatment is almost never assigned at random; it is assigned by someone’s policy, and the policy is the confounder. It is also a lesson about the danger of two biases in opposite directions, since a badly confounded estimate does not always look wrong. Sometimes it looks reasonable, which is worse.
7.13.3 Key Driver Analysis
The third vignette is key driver analysis, the regression genre most students will meet first: customer-satisfaction and brand trackers, where survey attributes — ease of use, value for money, feels like a brand for me — are regressed on overall satisfaction or intent, and the coefficients are relabeled impact scores and sorted into a bar chart. The genre is useful, because it prioritizes attention across dozens of attributes that no one can weigh by intuition. It is also where every pathology this chapter names tends to appear at once, which makes a driver deck a good exam for a new analyst.
The attributes are strongly intercorrelated, because a satisfied customer rates everything higher, so the individual impact scores are multicollinearity’s unstable division of shared credit, per Section 7.10 — and the ranking that the bar chart implies is precisely the quantity that instability makes unreliable. The causal direction is asserted rather than established: satisfied customers may rate value higher because they are satisfied, which is the reverse-causation rival from Section 7.2 sitting in the survey instrument itself. And the standard chart’s language — improving this attribute by one point will lift satisfaction by X — is Section 7.8’s coefficient-as-lever fallacy printed as a recommendation. Careful practitioners keep the genre honest by reporting attribute groups rather than pretending to separate entangled twins, by triangulating stated importance against modeled association, and by writing the deck’s verbs in association language. Which mostly means that the analyst in the room did to the vendor’s regression exactly what Table 7.5 does to an assistant’s.
7.13.4 In Your First Analyst Job
In your first analyst job, these vignettes compress into one expectation: you will spend more meeting-room minutes defending regressions from misreading than building them. The mix-model budget chart, the elasticity estimate, the driver bar chart — each will arrive in your inbox with causal language already attached, and the marginal value you add is rarely a better fit; it is the audit of signs, magnitudes, units, and diagnostics, the confounder named before the decision, the specification recorded so that next quarter’s model is comparable to this one, and the verb corrected before the slide ships. The statistics of this chapter are a solved problem in every software stack on earth. The judgment — what the coefficient can and cannot carry — is the job.
7.14 Ethics, Causal Overclaiming in Stakeholder Language
The Business Analytics in Practice section described analysts defending regressions from misreading. This section examines the defense as an ethical obligation, extending the guide’s running discussions — data use (Section 1.13), problem framing (Section 2.12), measurement design (Section 3.13), cleaning as editorial power (Section 4.14), honest summarization (Section 5.14), and differential treatment (Section 6.16) — to the act this chapter adds: stating what moves what. The chapter’s ethics is not a caution about a side effect. Overclaiming is the technique’s characteristic temptation, because association analysis produces artifacts that look exactly like causal answers, and the analyst is the only person in the chain positioned to know the difference.
Begin with how overclaiming happens, because the mechanism is rarely a lie. It is drift. The analyst writes “app ownership is associated with $140 higher spend, holding recency, tenure, engagement, and segment constant.” The deck version becomes “app users spend $140 more.” The executive summary becomes “the app drives $140 per customer.” The board slide becomes “app adoption is a $1M growth lever.” Each rewrite is shorter, each was made in good faith by someone optimizing for clarity, and the final sentence claims something the analysis never measured. No one decided to deceive; the qualifications were simply the least memorable part of every handoff, and compression removed them in order. The drift is worst under exactly two conditions this chapter creates: when the finding is quantified, because numbers survive compression and caveats do not, and when the finding flatters an existing plan, because the app team was always going to propose the install campaign. AI assistance intensifies both — Section 7.11 showed that assistants narrate causally by default, which means the first draft of the deck sentence now often arrives pre-drifted, and the analyst’s role shifts from author to editor of claims.
Why does it matter, beyond precision for its own sake? Because overclaimed drivers move real money onto weak evidence, and the costs land asymmetrically — the asymmetric-cost lens of Section 2.7, applied to language. Fund the install campaign on the drifted claim, and the cost of being wrong is a wasted budget and a metric that quietly fails to move, discovered quarters later and attributed to execution. Cut the discount ladder on the drifted claim, and the cost lands on the Discount-Dependent segment — a quarter of the treatment map loses the lever most associated with its purchasing, on evidence that never tested how it responds — and on the P&L in the holiday quarter. And there is a second-order cost, paid by the discipline itself: every quantified overclaim that later fails burns credibility that the next honest analysis must spend to be believed. The analyst’s overclaim is borrowed against colleagues’ future evidence.
The discipline this guide installs has three parts, all rehearsed in this chapter and none of them new machinery. First, the verb registry of Section 7.2, applied at the point of writing: “is associated with,” “predicts,” and “differs by” for this chapter’s findings; “drives,” “causes,” and “will lift” reserved for evidence that earned them, with the earning stated. The registry is a writing rule, which means it is auditable — a reviewer can grade a deliverable’s verbs against its methods in five minutes, and Exercise 7.7 does. Second, the drift defense: the analyst controls the source document, so the source document carries the qualification inside the sentence, not in a footnote — “app users spend $140 more than otherwise similar non-users, a difference this analysis cannot attribute to the app itself” survives one more round of compression than the same content footnoted. Anticipating compression is part of writing the sentence. Third, the escalation duty: when a drifted claim surfaces in a meeting — someone else’s slide, quoting your number with an upgraded verb — the analyst of record corrects it, with the same professional obligation as correcting a wrong number. The correction is rarely comfortable and is the whole point of having an analyst of record: Chapter 1 defined the role as accountability for what the analysis establishes, and the verb is part of what it establishes.
One boundary keeps the discipline honest in the other direction, because refusing all causal language is its own failure — an analyst who answers every question with “correlation is not causation” has abandoned the field to whoever will overclaim. The evidence for a relationship comes in grades: raw association; association with confounders named; association robust to controls and to specification, which is what Code 7.17 measures; association consistent with a designed mechanism; and, above them all, experimental evidence. Honest stakeholder language locates the finding on that ladder and states what would move it up a rung: “this is a controlled association that holds across four specifications; the December test would make it causal evidence.” The analyst’s job is not to withhold judgment forever; it is to price the judgment correctly and name the purchase that would upgrade it. That framing — evidence has a price list, and experiments are how you pay — is the bridge on which this chapter ends.
7.15 Chapter Summary
This chapter took up the question Chapter 6’s treatment map could not answer — what moves these customers? — and equipped the budget-lock decision with an honest drivers analysis. The main point is the standing distinction the chapter opened with and never set down: association is what marketing data records, causation is what marketing decisions require, and the analyst’s value lies in measuring the first rigorously while refusing to let it impersonate the second. “Driver” entered the chapter as a word on probation, and it leaves the same way — usable, but only with its evidence named.
The chapter’s machinery arrived in two movements. The measurement movement-built covariance — with the population quantity and its sample estimate distinguished, because the software reports one of them and the definition had better name which — and then the Pearson correlation coefficient, whose division by two standard deviations makes it indifferent to the convention that the covariance is not. It then cataloged what the correlation cannot buy: nonlinearity, outlier sensitivity, and mixed populations, each answered by the always-plot discipline of Anscombe’s quartet, and each demonstrated on StyleCraft’s base, where a hand-computed r of 0.20 between frequency and monetary value taught the mixed-population lesson on ten rows. Confounding gave the rivals of every association their mechanism, with the confounder defined by position — a pre-existing common cause — rather than by the loose test of being correlated with both, and with the calendar installed as marketing’s standing confounder. Segment composition entered on different terms: an estimated segment label built from same-window behavior is a descriptive stratifier that exposes what a pooled number is averaging over, not a clean pre-treatment control, and the chapter was careful to claim only the first.
The modeling movement-built regression from the fitted line upward: coefficients read aloud in native units, least squares as a criterion with a mean’s sensitivities, residuals as the model’s confession, and R² as a baseline comparison whose familiar properties hold under stated conditions — in-sample, with an intercept, on the same observations, in nested models. Multiple regression delivered the ceteris paribus reading and its strict boundaries: other included things equal, an adjustment computed from a fitted surface rather than a matching of real customers, and an estimate that can extrapolate where few customers share the relevant combination of values. Omitted variable bias supplied the direction logic for reasoning about everything excluded, with the qualification that in a multiple regression the relevant association is the one that survives the model’s other predictors. Dummy variables and reference categories brought the four segments into the model as evidence, with the intercept correctly described as the expected response at every predictor’s zero rather than as the reference group’s mean. The interaction term let one association differ across groups, read conditionally and reported as two-dollar figures rather than one significance verdict.
The evidence-reading section calibrated the uncertainty columns at working level and connected them to the diagnostic that precedes them: a funneling residual plot is heteroskedasticity, heteroskedasticity distorts standard errors without biasing coefficients, and the repair is a robust covariance estimator requested at fit time — so every model in the labs was fitted with HC3, and the labs printed the nonrobust and robust standard errors side by side to show exactly what changes and what does not. The p-value was defined through the test statistic it is computed from, under the null hypothesis and the model’s assumptions, and read with the American Statistical Association’s three warnings attached. Multicollinearity was diagnosed with variance inflation factors as well as pairwise correlations, deliberately without a universal cutoff, and described honestly: it may leave in-sample fit largely intact while destabilizing coefficients, and it can weaken out-of-sample prediction when the relationship among predictors shifts. Prediction and explanation were separated as modeling goals with different tolerances.
The AI section named the assistant-era failure modes — the confident causal narration, the silent specification, the plausible-but-impossible coefficient, the default covariance estimator on funneling residuals, the unconditional reading of a conditional coefficient — and installed the four-point audit of signs, magnitudes, units, and diagnostics that the labs then rehearsed. Eighteen numbered cells carried the argument into executable form: the two covariances computed side by side, the labeled miniature scatter with its mean lines, a correlation matrix preceded by the completeness check that makes it one exhibit rather than several, whale sensitivity, faceted and within-segment views of the discount claim, the composition features built and reconciled against certified totals, robust and nonrobust inference contrasted, the funnel measured rather than eyeballed, the specified model with its encoding asserted, a coefficient table carrying units, variance inflation factors, Chapter 5’s decomposition remeasured, the designed interaction reported as two conditional associations, the email claim tested rather than assumed, a specification sweep, and a leakage demonstration in which a model handed the outcome under another name returns an R² of one. The interpretation, practice, and ethics sections carried the findings into the meeting room, where the verbs proved to be part of the finding and the analyst of record proved to be their editor.
Looking ahead, the drivers model was built to explain: its coefficients were the deliverable, and their honesty was the discipline. The next question is the one the model was not built for — how well do these patterns hold on customers the model has never seen? The moment a model’s outputs, rather than its coefficients, become the product, the obligations change: targets and horizons must be defined, data must be split, baselines must be beaten, and a new class of silent failure — the model that grades itself on information it should never have had, of which Code 7.18 was a deliberately obvious instance — becomes the central danger. Chapter 8 takes up predictive modeling for marketing decisions, where Part II’s machinery turns from explaining the past toward scoring the future, and where the too-good result replaces the wrong verb as the red flag an analyst is paid to notice. Further out, Chapter 11 supplies the instrument this chapter kept naming and could not use: experimental design, where random assignment severs the backdoors that every observational coefficient in this chapter had to reason around, and where the verbs this chapter rationed can finally be earned.
7.16 Exercises for Practice
The following exercises practice the chapter’s main habits: predict signs before computing, plot before believing, read every coefficient in its units against its reference, read conditional coefficients conditionally, reason omitted-variable-bias directions, match the covariance estimator to the residuals, run the four-point audit on every fitted model, and grade every verb against its evidence. They are organized into two groups. Core chapter practice is the required path and should be completed by every student.
7.16.1 Core Chapter PracticeExercise 7.1 Concept Check
Answer each in two or three sentences, in your own words.
- State the difference between association and causation and give the opening case’s app claim as an example of the gap between them.
- Why is covariance a direction detector but not a strength measure, and what operation converts it into one? State also which covariance pandas reports and explain why the same choice does not change the correlation.
- An analyst reports r = 0.02 between tenure and average order value and concludes the two are unrelated. Name two distinct ways this conclusion could be wrong, citing the section that covers each.
- Define a confounder by its position rather than by its correlations. Then explain why “related to both variables” is too broad a test, naming the two other kinds of variable that pass it and stating why adjusting for each causes a different problem.
- State the ceteris paribus reading of a multiple regression coefficient and then state three distinct boundaries on it — one about the variable list, one about what the adjustment is, and one about the regions of the data where it may be extrapolating.
- Why must one category be excluded when a categorical variable enters a regression as dummies, and what does the intercept represent in a model that also contains a numeric predictor?
- A model contains app_user, loyalty_enrolled, and their product. State what each of the three coefficients means and write the expression for the app association among loyalty-enrolled customers.
- State the three familiar properties of R² together with the conditions under which they hold and give one situation in which a reported R² would not satisfy them.
- A coefficient has p = 0.001 and a 95 percent confidence interval of $0.40 to $0.62 per unit. Write the honest one-sentence report, state precisely what quantity the p-value is the probability of and name the misreading the p-value alone invites.
- A residual plot funnels outward as fitted spend rises. Say what that is called, what it does and does not distort, and what the analyst should change about the fit. Then state what stays the same after the change.
- What does a variance inflation factor measure that a pairwise correlation matrix cannot, and why does this guide decline to publish a cutoff for it?
- State the difference between a model built to explain and a model built to predict, and name one practice that is a failure in one and tolerable in the other.
Exercise 7.2 Predict the Signs
Before running Lab 7.1 Part B, write the predicted sign (positive, negative, or near zero) of the correlation between monetary and each of: recency_days, frequency, aov, discount_share, store_share, orders_per_month, revenue_per_month, app_user, and email_opt_in — with one sentence of reasoning each, citing Chapter 5’s certified findings or Chapter 6’s segment profiles where they apply. Mark in advance which two of the nine you expect to be arithmetic relatives of monetary rather than behavioral findings and say how you would tell the difference from the matrix alone. After running Code 7.4, grade every prediction and write one paragraph on your largest miss: what did the data know that your reasoning did not, or what did the pooled computation conceal that Section 7.4’s tools would reveal?
Exercise 7.3 Read the Coefficient Aloud
For each fitted result below, write the single honest sentence that reports it — units, direction, held-constant clause, and a verb that stops at the evidence. Then write the overclaimed version a deck would produce and underline the exact words that changed the claim.
- In the Code 7.11 spend model: recency_days coefficient −0.85, HC3 interval −1.00 to −0.70.
- In the Code 7.11 spend model: seg_suburban coefficient −180, HC3 interval −220 to −140, with Urban Loyal Core as the reference.
- In the Code 7.11 spend model: discount_share coefficient −310, HC3 interval −390 to −230. Be careful with the units on this one; state what a one-unit difference in discount_share is.
- In the Code 7.14 controlled AOV model: seg_suburban coefficient 4, HC3 interval −2 to 10.
- In the Code 7.15 interaction model: app_user coefficient 45, app_user:loyalty_enrolled coefficient 60, both intervals excluding zero. Report the two conditional associations rather than the three coefficients and explain why the deck should never quote the 45 on its own.
Exercise 7.4 Hand Arithmetic on the Miniature
Using only the ten-customer miniature of Code 7.1 and no code: verify the sample covariance of 26.73 between frequency and monetary by computing the ten deviation products and summing them, and state what the population covariance would be and why the two differ. Explain in one sentence why C002 and C010 contribute exactly zero to the sum despite being the two largest spenders. Recompute r with C006 removed and state what the change reveals about single-observation influence. Sketch the scatterplot by hand with both mean lines drawn, label the quadrant each customer falls in, and mark the region that would mislead a straight-line summary. Close with one sentence connecting the exercise to Anscombe’s argument.
Exercise 7.5 The Full Drivers Deliverable
Complete Labs 7.1 and 7.2 on the certified files and assemble the VP’s one-page deliverable per Section 7.13: the four verdict paragraphs on the app, discount, email, and average-order-value findings; the coefficient table with HC3 intervals and a units column; the pooled-versus-within-segment discount exhibit with its faceted plot; the two conditional app associations from the designed interaction, in dollars; the specification-sensitivity range for any coefficient your recommendation leans on; and the closing line naming the experiment that would upgrade the weakest verdict. State once, in a methods line, that all intervals are heteroskedasticity-robust, and say why. Submissions are graded on the verdict paragraphs’ verbs and the audit evidence as heavily as on the code.
Exercise 7.6 AI Regression Audit
Give an AI assistant the customer_features table and this deliberately loose prompt, verbatim: “Find what drives customer spend and tell me what to do about it.” Then audit the response with Table 7.5’s four points. Check every reported sign against your Exercise 7.2 predictions. Check every magnitude against a hand-computed anchor — group means, Chapter 5 gaps — and note whether any adjusted figure exceeds its descriptive counterpart, which is permitted but requires an explanation grounded in the model’s variables. Restate every quoted coefficient in its units, flag any cross-unit comparison, and flag any main effect quoted unconditionally from a model that contains its interaction. Determine whether the assistant examined residuals, which covariance estimator it used, and whether it connected the two. Separately, list every causal verb in the reply and grade each against Section 7.2’s registry. Report what specification the assistant chose — variables, reference category, interactions — and whether it disclosed choosing.
7.17 Glossary of Terms
This glossary includes only the terms introduced in this chapter. Each definition is tied to the sources used in the chapter rather than added for decoration.
Association. The tendency of two variables to move together in observed data, so that knowing one improves a guess about the other; symmetric, observational, and consistent with multiple causal structures (adapted from Simon, 1954; Freedman, 1991).
Causation. The relation in which changing one variable, with all else undisturbed, would change another; directional and interventional, and never established by association alone (adapted from Simon, 1954; Freedman, 1991).
Ceteris paribus. The “other things equal” reading of a multiple regression coefficient: the estimated difference in the response per unit difference in a predictor among observations alike on the model’s other included variables, computed as an adjustment under the model’s assumed functional form rather than by matching observations (adapted from Wooldridge, 2020).
Confidence interval (coefficient). The range of coefficient values compatible with the data at a stated confidence level, centered on the estimate; its width is the finding about precision, and its validity depends on the standard error it was built from (adapted from Wasserstein & Lazar, 2016; Wooldridge, 2020).
Confounder. A pre-treatment variable that influences, or helps determine, both the predictor or exposure and the outcome, creating or distorting their observed association. Correlation with both variables is a warning sign rather than proof: a variable on the pathway between them, or a common consequence of them, is also correlated with both and requires different handling (adapted from Simon, 1954; Freedman, 1991).
Correlation matrix. The grid of pairwise Pearson correlations among a set of variables; a screening exhibit who’s characteristic misreadings — ranking drivers, celebrating near-perfect cells, treating cells as levers — this chapter catalogs, and whose cells are computed pair by pair over the observations complete for that pair, so incomplete columns silently change the population behind each figure (behavior per pandas development team, 2026).
Covariance. The mean product of two variables’ paired deviations from their means in a population; estimated from a sample by dividing the sum of the paired deviation products by n − 1. Positive for same-direction movement, negative for opposite, with a magnitude that depends on both variables’ units (adapted from Rodgers & Nicewander, 1988).
Dummy variable. An indicator predictor equal to 1 for membership in a category and 0 otherwise, by which categorical variables enter regression; c categories require c − 1 dummies (adapted from Wooldridge, 2020).
Heteroskedasticity. Residual variance that differs systematically across the range of fitted values, visible as a funnel in a residual plot. It does not bias the fitted coefficients, and it does distort the standard errors computed under the default assumption, so the intervals and p-values built from them are unreliable until a robust covariance estimator is used (adapted from Wooldridge, 2020; MacKinnon & White, 1985).
Interaction term. A predictor formed as the product of two others, whose coefficient estimates how the association between one predictor and the response differs across values of the other; once it is present, each main-effect coefficient is the association for its variable when the interacting partner equals zero (adapted from Wooldridge, 2020; James et al., 2021).
Least squares. The estimation criterion that chooses regression coefficients to minimize the sum of squared residuals. The solution is unique when the predictor matrix has full column rank; under perfect redundancy among predictors, software may still return a fitted result whose individual coefficients are not identified. Like every mean, it is sensitive to extreme observations (adapted from James et al., 2021).
Multicollinearity. Strong overlap in the information carried by a model’s predictors, whether pairwise or between one predictor and a combination of others, which leaves joint explanatory power largely intact while making individual coefficients imprecise and unstable across specifications (adapted from Wooldridge, 2020).
Multiple regression. The modeling of a numeric response as a linear function of two or more predictors fitted jointly, so that each coefficient carries a ceteris paribus reading with respect to the model’s other variables (adapted from Wooldridge, 2020).
Omitted variable bias. The distortion of an included predictor’s coefficient when a variable related to both the response and that predictor is left out of the model. In the simplest linear case, its direction follows the product of the omitted variable’s effect on the response and its association with the included predictor; in a multiple regression the relevant association is the part remaining after the model’s other predictors are considered (adapted from Wooldridge, 2020).
p-value (coefficient). The probability, assuming the stated null hypothesis and statistical model are correct, of obtaining a test statistic at least as extreme as the one observed; for a regression coefficient, the null hypothesis is that the coefficient is zero given the model’s other variables. It is a measure of surprise under that hypothesis, not of size, importance, or truth (adapted from Wasserstein & Lazar, 2016).
Pearson correlation coefficient (r). The sample covariance of two variables divided by the product of their sample standard deviations: a unit-free, symmetric measure of linear association ranging from −1 to +1, unaffected by the denominator convention that changes the covariance (adapted from Pearson, 1895; Rodgers & Nicewander, 1988).
R² (coefficient of determination). The proportion of a response’s variation around its mean that a fitted model accounts for. For in-sample ordinary least squares fitted with an intercept to the same observations, it lies between 0 and 1, cannot decrease when predictors are added to a nested model, and equals r² in the simple one-predictor case; it grades fit to the sample and never, by itself, usefulness for a decision (adapted from Wooldridge, 2020; James et al., 2021).
Reference category. The category of a categorical predictor represented when all its dummy variables equal zero, against which every dummy coefficient is read as an adjusted difference. With an intercept in the model, the intercept is the expected response for the reference category when all other predictors are at their coded zero or reference values — the reference group’s observed mean only when the model contains nothing else (adapted from Wooldridge, 2020).
Residual. An observation’s actual response minus the model’s predicted value; the itemized record of what the model missed, whose patterns are the primary diagnostic evidence about model form (adapted from James et al., 2021).
Robust standard error (heteroskedasticity-consistent). A standard error computed under an estimator that does not assume constant residual variance, requested at fit time and used throughout this chapter’s labs in its HC3 form. It changes the standard errors, confidence intervals, and p-values, and leaves the fitted coefficients and predictions unchanged (adapted from MacKinnon & White, 1985; Seabold & Perktold, 2010).
Simple linear regression. The modeling of a numeric response as an intercept plus a slope times one predictor, with coefficients estimated by least squares and read in the native units of the variables (adapted from James et al., 2021).
Spurious correlation. An observed association that reflects only the operation of confounders or chance, so that neither variable responds to the other (adapted from Simon, 1954).
Standard error (coefficient). The estimated sampling variability of a fitted coefficient — how much the estimate would move across repeated samples from the same population — and the quantity from which confidence intervals and test statistics are built. It describes the precision of the estimate, not the spread of the underlying variable (adapted from Wooldridge, 2020).
Variance inflation factor. A diagnostic measuring how much a coefficient’s estimated variance is inflated by the predictor’s relationships with all the other predictors together, and therefore able to detect overlap that not pairwise correlation reveals. It is a diagnostic rather than a verdict, and this guide publishes no universal cutoff for it (adapted from Seabold & Perktold, 2010; Wooldridge, 2020).
7.18 References
Anscombe, F. J. (1973). Graphs in statistical analysis. The American Statistician, 27(1), 17–21. https://doi.org/10.1080/00031305.1973.10478966
Borden, N. H. (1964). The concept of the marketing mix. Journal of Advertising Research, 4(2), 2–7.
Chan, D., & Perry, M. (2017). Challenges and opportunities in media mix modeling. Google Inc. https://research.google/pubs/challenges-and-opportunities-in-media-mix-modeling/
Freedman, D. A. (1991). Statistical models and shoe leather. Sociological Methodology, 21, 291–313. https://doi.org/10.2307/270939
Galton, F. (1886). Regression towards mediocrity in hereditary stature. Journal of the Anthropological Institute of Great Britain and Ireland, 15, 246–263. https://doi.org/10.2307/2841583
James, G., Witten, D., Hastie, T., & Tibshirani, R. (2021). An introduction to statistical learning: With applications in R (2nd ed.). Springer. https://doi.org/10.1007/978-1-0716-1418-1
MacKinnon, J. G., & White, H. (1985). Some heteroskedasticity-consistent covariance matrix estimators with improved finite sample properties. Journal of Econometrics, 29(3), 305–325. https://doi.org/10.1016/0304-4076(85)90158-7
pandas development team. (2026). pandas.DataFrame.corr [Software documentation]. https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.corr.html
Pearson, K. (1895). Note on regression and inheritance in the case of two parents. Proceedings of the Royal Society of London, 58, 240–242. https://doi.org/10.1098/rspl.1895.0041
Rodgers, J. L., & Nicewander, W. A. (1988). Thirteen ways to look at the correlation coefficient. The American Statistician, 42(1), 59–66. https://doi.org/10.1080/00031305.1988.10475524
Seabold, S., & Perktold, J. (2010). Statsmodels: Econometric and statistical modeling with Python. In Proceedings of the 9th Python in Science Conference (pp. 92–96). https://doi.org/10.25080/Majora-92bf1922-011
Shmueli, G. (2010). To explain or to predict? Statistical Science, 25(3), 289–310. https://doi.org/10.1214/10-STS330
Simon, H. A. (1954). Spurious correlation: A causal interpretation. Journal of the American Statistical Association, 49(267), 467–479. https://doi.org/10.2307/2281124
Vigen, T. (2015). Spurious correlations. Hachette Books.
Wasserstein, R. L., & Lazar, N. A. (2016). The ASA’s statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129–133. https://doi.org/10.1080/00031305.2016.1154108
Wooldridge, J. M. (2020). Introductory econometrics: A modern approach (7th ed.). Cengage Learning.