A key element in hierarchical modeling is the choice of prior distribution for the group-level variance (scale) parameter $\tau$.
In previous analyses, we have used a uniform prior on $\tau$, but many other so-called noninformative priors have been proposed in the Bayesian literature.
In practice, the choice of such a “noninformative” prior can have a large effect on inference—especially when the number of groups $J$ is small or the true group-level variation $\tau$ is small.
Although this discussion uses the normal hierarchical model for illustration, the same ideas apply generally to variance parameters in multilevel models.
Conceptual Background
1. Improper priors as limits of proper priors
Improper prior densities can—but do not necessarily—lead to proper posterior distributions.
To reason clearly, it helps to treat improper distributions as limits of proper ones.
Two commonly considered examples for the group-level variance are:
- $\tau \sim \text{Uniform}(0, A)$ with $A \to \infty$
- $\tau^2 \sim \text{Inverse-Gamma}(\epsilon, \epsilon)$ with $\epsilon \to 0$
For the normal hierarchical model, the uniform prior on $\tau$ yields a proper posterior as long as $J \ge 3$.
Hence, for a large enough finite $A$, inference is not sensitive to the exact upper limit.
In contrast, the $\text{Inverse-Gamma}(\epsilon, \epsilon)$ prior does not have a proper limiting posterior as $\epsilon \to 0$; thus, inferences become highly sensitive to the chosen $\epsilon$.
2. Calibration of the posterior mean
Posterior calibration is the Bayesian analogue of classical bias analysis.
Let the posterior mean be $\hat{\theta} = E(\theta \mid y).$
The miscalibration of this estimator is defined as $E(\theta \mid \hat{\theta}) – \hat{\theta}.$
If the prior distribution used in the model matches the true data-generating prior, this quantity should be zero for all $\hat{\theta}$.
However, when using an improper prior, this calibration can fail.
For instance, if the true prior for $\tau$ were $\text{Uniform}(0, A)$ but we used $\text{Uniform}(0, \infty)$, the posterior would integrate over unrealistically large $\tau$ values.
This yields positive miscalibration—that is, a tendency to overestimate $\tau$ on average.
Classes of Noninformative or Weakly Informative Priors
(a) Uniform Priors
Several versions of “uniform” priors are possible, depending on which scale is used.
- Uniform on $\log \tau$:
Although this may seem natural since $\tau > 0$, it leads to an improper posterior in hierarchical models.
The problem arises because $p(y \mid \tau)$ approaches a nonzero limit as $\tau \to 0$, so the posterior accumulates infinite mass near zero. - Uniform on $\tau$:
The prior $p(\tau) \propto 1$ is often used in practice.
It avoids divergence at $\tau = 0$ and works reasonably well when $J \ge 3$.
However, it has infinite mass in the upper tail ($\tau \to \infty$), producing a mild positive bias (overestimation of $\tau$).
With very small $J$, such as 1 or 2, the posterior becomes improper, and for $J = 4$ or $5$, the right tail is excessively heavy, resulting in under-pooling of the group-level effects $\theta_j$. - Uniform on $\tau^2$:
This choice further exaggerates the overestimation problem, and a proper posterior requires $J \ge 4$.
Hence, it is not generally recommended.
Mathematically, the uniform-on-$\log\tau$ prior corresponds to $p(\tau) \propto \frac{1}{\tau} \quad \text{or} \quad p(\tau^2) \propto \frac{1}{\tau^2},$
which can be viewed as the limit of an inverse-$\chi^2$ prior with 0 degrees of freedom.
Similarly, $p(\tau) \propto 1$ corresponds to $τp(\tau^2) \propto \frac{1}{\tau}$, an inverse-$\chi^2$ prior with −1 degree of freedom, which can also be regarded as a limiting case of the half-$t$ family with infinite scale.
(b) Inverse-Gamma $(\epsilon, \epsilon)$ Priors
The inverse-gamma family is conditionally conjugate for variance parameters.
That is, if $\tau^2 \sim \text{Inv-Gamma}(\alpha, \beta)$, then the conditional posterior $p(\tau^2 \mid \theta, \mu, y)$ is also inverse-gamma.
Using the parameterization $s^2 = \frac{\beta}{\alpha}, \quad \nu = 2\alpha,$
this corresponds to an inverse-$\chi^2(\nu, s^2)$ prior.
The $\text{Inv-Gamma}(\epsilon, \epsilon)$ prior, with small $\epsilon$ (e.g., 1, 0.01, or 0.001), is often mistakenly considered “noninformative.”
However, as $\epsilon \to 0$, the posterior becomes improper, and for realistic data with small possible $\tau$, inference is highly sensitive to $\epsilon$.
Thus, this prior is not genuinely noninformative.
(c) Half-ttt and Half-Cauchy Priors
A practical and flexible alternative is the half-$t$ family, in particular the half-Cauchy distribution: $\tau \sim \text{half-Cauchy}(0, A),$
where $A$ is a scale parameter chosen to be large enough to be weakly informative.
This prior has a broad peak at zero, a slow-decaying tail, and a single interpretable scale parameter.
In the limit $A \to \infty$, it approaches the uniform prior on $\tau$.
For large but finite $A$, it acts as a weakly informative prior: it constrains only implausibly large $\tau$ values while letting the data dominate over most of the plausible range.
Half-Cauchy priors are particularly useful when $J$ is small, because the uniform prior is often too weak to produce a realistic posterior.
Example: The Eight-Schools Problem
In the eight-schools example, the parameters $\theta_1, \dots, \theta_8$ represent the treatment effects (in SAT-V score points) for each school, and $\tau$ represents the between-school standard deviation.
Given the test score range (200–800), the plausible maximum difference is about 300 points, so a realistic upper bound for $\tau$ might be around 100.

Figure in the text compares three “noninformative” priors for this model:
- Uniform on $\tau$
The posterior for $\tau$ is mainly below 20, with a mild right tail—reasonable since $J=8$ provides enough data to constrain the parameter. - Inverse-Gamma(1, 1) on $\tau^2$
The posterior becomes much tighter, concentrated around $\tau \in [0.5, 5]$, and closely mirrors the prior.
Shrinkage of the school effects $\theta_j$ increases, implying the prior is overly restrictive. - Inverse-Gamma(0.001, 0.001) on $\tau^2$
The posterior collapses near zero, further distorting inference and yielding unrealistically strong pooling.
Hence, the “noninformative” inverse-gamma priors actually impose strong information, favoring $\tau \approx 0$.
In contrast, the simple uniform prior on $\tau$ performs well, giving a reasonable posterior without imposing unrealistic constraints.
Note that these behaviors depend strongly on the plotting scale.
When plotted on the log scale, the inverse-gamma(0.001, 0.001) prior looks flattest, but this is misleading because the posterior remains extremely sensitive to $\epsilon$.
Example: The Three-Schools Problem
The uniform prior on $\tau$ works well when $J=8$, but fails when $J$ is much smaller.
With $J=3$, the data provide little information about between-group variability, and the uniform prior produces a posterior with an unrealistically long right tail, implying implausibly large $\tau$ values.
Such a posterior leads to under-pooling of the $\theta_j$ estimates.
A better choice here is a half-Cauchy prior with a weakly informative scale, for example:
$\tau \sim \text{half-Cauchy}(0, 25)$
In this educational testing context, a scale of 25 is large enough to allow realistic variation (plausible $\tau < 50$) but still prevents the tail from extending to absurdly high values.
This prior reflects a weak expectation that large $\tau$ values are unlikely, while still letting the likelihood dominate within the plausible range.
Summary
- The choice of prior for the variance parameter $\tau$ is not trivial; so-called “noninformative” priors can behave quite differently.
- The uniform prior on $\log\tau$ produces an improper posterior (infinite mass at zero).
- The uniform prior on $\tau$ is often reasonable for moderate $J$ (≥ 3), but slightly overestimates $\tau$ and leads to heavy right tails.
- The inverse-gamma$(\epsilon, \epsilon)$ priors, though popular, are not truly noninformative and often produce misleading shrinkage toward zero.
- The half-Cauchy prior (or more generally, a half-$t$ prior with a large scale) is a robust weakly informative choice, especially when $J$ is small.
It constrains the posterior gently without dominating the data.
In short, weakly informative priors—like the half-Cauchy—provide a principled compromise between overly vague and overly restrictive priors for variance parameters in hierarchical Bayesian models.
