1. What makes observational studies different?
- In an experiment, the researcher assigns treatment.
- In an observational (nonexperimental) study, the researcher just observes who got which “treatment” (coaching vs no coaching, program vs no program, drug vs no drug), but didn’t control it.
- The data often look like experimental data (treatment + outcome for each unit), so people analyze them like experiments.
- But: in observational data, the treated and control groups can be very different to begin with, especially in unmeasured ways. That’s the core problem.
Example from the text:
- SAT coaching = experimental if the school randomized students to coaching vs no coaching.
- SAT coaching = observational if students chose coaching themselves.
- In the second case, coached students might already be more motivated → bias.
2. When can observational studies still support causal inference?
They list 5 conditions that make an observational study more credible:
- Good control for background differences: we have enough covariates to explain why some got treated and others didn’t.
- Enough units in each treatment group: so posteriors aren’t too wide.
- Design not tailored to the desired answer: the way data were collected isn’t “influenced” by the outcome we want.
- Missingness handled: dropout, attrition, etc. are either small or modeled.
- Analysis matches design: you actually condition on the variables that determined treatment.
If these aren’t met, causal conclusions get very sensitive.
3. Why sensitivity is worse than in experiments
- In randomized experiments, treated and control groups are balanced in expectation.
- In observational studies, treated and control groups can be unbalanced — e.g. only richer people take the training.
- When there is a lot of imbalance, then adjusting with a regression model $y \mid x, \text{treatment}$ becomes model-dependent:
- with one functional form, you see a positive effect,
- with another plausible functional form, the effect disappears.

- That’s exactly what their above image is showing:
- (a) balanced → treatment effect is stable across models.
- (b), (c) unbalanced → treatment effect flips depending on whether you fit linear or mildly nonlinear.
So: lack of overlap ⇒ high modeling sensitivity.
4. Role of propensity scores
- Propensity score = $e(x) = \Pr(I_i = 1 \mid X_i = x)$: probability of getting the treatment, given covariates.
- In a bad observational study, treated units have propensity scores near 1, controls near 0 → no overlap.
- If there is little or no overlap in propensity scores, then causal estimates are fragile — exactly like in their pictures.
- Using propensity scores (matching, trimming, subclassification) can:
- restrict analysis to the region where treated and control look similar,
- and therefore reduce sensitivity to the exact outcome model.
So in Bayesian observational analysis, estimating the propensity score is often a diagnostic step: it tells you how bad the imbalance is.
5. Bayesian view for observational studies
What do we do in Bayesian analysis?
- Collect a lot of covariates so that treatment assignment is (approximately) ignorable given those covariates.
- “Ignorable” here = $\Pr(I \mid y, x, \phi)$ depends only on $x$, not on unobserved potential outcomes.
- If we accept ignorability, we can analyze the observational data as if it were an experiment with non-equal assignment probabilities:
- $p(\theta \mid x, y_{\text{obs}}) \propto p(\theta)\,p(y_{\text{obs}} \mid x, \theta)$
- Then we can do the usual Bayesian things: regression adjustment, poststratification, averaging over covariate distribution.
But: if we cannot make treatment “almost ignorable” with available covariates, then
- inferences become very sensitive,
- and sometimes the data basically can’t answer the causal question.
That’s their warning sentence: “without enough covariates to make the design approximately ignorable, sensitivity … can be so great that the observed data may provide essentially no information.”
6. Tools to reduce sensitivity in observational studies
They list standard design/analysis tricks that all have the same philosophy: constrain the x-space so you don’t have to extrapolate much.
- Matched sampling: pair treated and control units with similar x.
- Subclassification / stratification / blocking: group by x (or by propensity score) and compare within groups.
- Poststratification (very natural in Bayesian models): model within cells and average over population cell proportions.
- Analysis of covariance (regression adjustment): model y∣x,treatmenty \mid x, \text{treatment}y∣x,treatment and then average.
All of these try to make it so that the model for $y \mid x$ is used locally, not to extrapolate from rich to poor or from old to young when only one group got treated.
7. Two main difficulties they highlight
- Unbalanced treatments
- If good students choose coaching and weak students don’t, and you don’t have enough x to fully adjust, then your result depends a lot on assumptions (e.g. linearity, additivity).
- Sometimes that makes the dataset basically unusable for the causal question.
- Unknown / unmeasured drivers of treatment
- In real observational data, people take treatment for reasons we don’t fully observe (motivation, health, mood).
- Then treatment is not ignorable, and your posterior for θ will depend on how you model that nonignorable mechanism.
So again: data collection quality matters a lot more than in randomized experiments.
8. Principal stratification (intermediate variables on the causal path)
Now they shift to a special but very common situation: randomized experiment with noncompliance (or, more generally, an intermediate post-treatment variable).
Set-up:
- $I$: assignment (1 = assigned treatment, 0 = assigned control)
- $C$: intermediate post-assignment variable (e.g. did the person actually take the vitamin?)
- Final outcome: $Y$
Naive mistake: treat observed $C_{\text{obs}}$ as if it were a baseline covariate and stratify on it. That’s wrong because $sC_{\text{obs}}$ depends on assignment.
Correct idea: define principal strata by the pair of potential intermediate outcomes $(C(1), C(0))$:
- “Compliers”: $C(1)=1, C(0)=0$
- “Never-takers”: $C(1)=0, C(0)=0$
- (In other cases you can also have always-takers or defiers, depending on the study.)
These principal strata are pre-treatment characteristics conceptually — they don’t change with assignment — so it’s legitimate to stratify on them.
That’s what “principal stratification” means.
9. Example: vitamin A experiment with noncompliance
- Vitamin A was available only to those assigned to treatment.
- So: control group can’t get it → everyone has $C(0)=0$.
- Among those assigned treatment, some actually took it (compliers), some didn’t (never-takers).
- So there are two relevant strata: compliers and never-takers.
- The table they show is just counts by (assignment, exposure, outcome).
Goal: estimate the causal effect of actually taking vitamin A for compliers. This is called:
- CACE = Complier Average Causal Effect.
- Also called LATE (Local Average Treatment Effect) in econometrics.
10. Link to instrumental variables
In this noncompliance setting:
- Assignment III is randomized → it’s a valid instrument.
- Vitamin actually taken is the treatment actually received.
- Outcome is infant survival.
They give the key identity:
$\text{ITT} = Y^1 – Y^0 = p_c \cdot \text{CACE} + (1 – p_c)\cdot \text{NACE}$
where
- ITT (intention-to-treat effect) = effect of being assigned treatment,
- $p_c$ = proportion of compliers,
- CACE = effect among compliers,
- NACE = effect among never-takers.
If we make the exclusion restriction:
- assignment has no effect on never-takers (NACE = 0),
then
$\text{CACE} = \frac{\text{ITT}}{\hat p_c}$
That’s the instrumental variables (IV) estimate: divide the ITT by the compliance rate.
11. Bayesian causal inference with noncompliance
They say: the IV formula is nice and simple, but it’s a method-of-moments style estimate.
The Bayesian way:
- treat the compliance status of the control group as missing data,
- write a full model for outcomes, assignment, and compliance,
- sample from the posterior using MCMC,
- and you can relax exclusion restriction if you want.
This is a good example of their overall message: when design is not perfectly clean (observational, noncompliance, intermediate variables), Bayesian modeling lets you write down the full data-generating story — but your results will be as sensitive as your assumptions.
12. Core message of this section
- Observational studies can be analyzed “like experiments” only if you make treatment assignment (approximately) ignorable by conditioning on rich covariates.
- Even then, poor covariate overlap ⇒ high sensitivity to functional-form assumptions.
- Propensity scores and design-like devices (matching, stratifying, blocking) are ways to reduce that sensitivity by restricting to comparable units.
- Principal stratification is the right way to handle intermediate post-treatment variables (like compliance).
- In noncompliance settings, the CACE / IV logic gives you the causal effect for people who actually take treatment when offered.
