The hierarchical SAT coaching model uses the data model
which is reasonable because each school’s estimate is based on a large sample.
The population model
is less obviously justified, although model checking indicates that it provides adequate posterior intervals. Posterior inferences can nevertheless be sensitive to modeling choices even when the model fits the observed data well.
To examine robustness, an alternative population distribution is introduced:
Let
denote the posterior distribution under this population model, and let
denote the posterior under the normal model used earlier.
1. Robust Inference Under a $t_4$ Population Model
Long-tailed population distributions reduce shrinkage for extreme schools and reduce the influence of the largest observed effect on the estimate of . To evaluate this possibility, the population model is replaced with a distribution, while leaving the likelihood
unchanged and retaining the hyperprior
Gibbs Sampling
Sampling proceeds using the mixture representation for the t distribution described earlier. With , Gibbs sampling yields 2500 posterior draws after combining the second halves of five chains. Table reports posterior quantiles. The results are nearly identical to those under the normal population model, with only slightly reduced shrinkage for extreme schools such as school A.

2. Computing the Robust Posterior via Importance Resampling
The robust posterior under can also be approximated without running a new Gibbs sampler.
Step 1: Draw from the original posterior
Generate 5000 draws
Step 2: Compute the importance ratio
For each draw, compute:
The likelihood and hyperprior terms cancel, leaving
Step 3: Importance resampling
Select 500 draws without replacement, with probabilities proportional to the ratios above.
This provides an approximate sample from
The long right tail of the importance ratios indicates potential instability, suggesting limits on accuracy for this approximation.
3. Sensitivity Analysis Across Different Degrees of Freedom $\nu$
To examine sensitivity to the normality assumption, posterior simulation is repeated for
in addition to the normal () and robust () models.
Results are summarized by the posterior means and standard deviations of the eight ’s as functions of . This reparameterization maps the normal model to and produces a finite interval [0,1] spanning all distributions from normal to Cauchy.
The figure shows variability arising mostly from simulation noise, with no systematic dependence of posterior inferences on .
4. Treating $\nu$ as an Unknown Parameter
A full Bayesian treatment places a prior on and integrates it out.
Prior for
A uniform prior is assigned to over the interval , producing equal prior weight on the range from the normal model to the Cauchy distribution. The conditional prior on is set to
The scale parameter does not directly correspond to a variance when , but this parameterization remains workable for posterior computation.
Posterior simulation
The Gibbs sampler is expanded with a Metropolis step to update . Posterior draws of are shown in Figure; the distribution is diffuse but does not materially alter the inferences for the school effects.

5. Discussion
Posterior medians and central posterior intervals (50% and 95%) for the eight ’s exhibit little sensitivity to the choice of normal vs. population distributions. Extreme tail quantities (such as 99.9% intervals) could show stronger dependence on , although such extremes are not substantively important in this application. Overall, the educational testing example demonstrates relative insensitivity of the inferences to the normality assumption for the population distribution of the school-level effects.
