Posterior
1. Definition
- The posterior is the updated probability distribution of a parameter (or hypothesis) after observing data.
- It represents the new belief after combining:
- the prior belief (what we thought before), and
- the likelihood (evidence from data).
In short:
Posterior = Prior updated with Evidence.
2. Formula (Bayes’ Theorem)
For a parameter $\theta$ and data $D$:
$P(\theta \mid D) = \frac{P(D \mid \theta) \cdot P(\theta)}{P(D)}$
Where:
- $P(\theta)$ = prior (belief before data)
- $P(D \mid \theta)$ = likelihood (probability of data given parameter)
- $P(D)$ = marginal likelihood (normalizing constant)
- $P(\theta \mid D)$ = posterior (updated belief)
3. Interpretation
- The posterior distribution shows which parameter values are most plausible after observing the data.
- It allows statements like:
- “There’s a 95% probability that the conversion rate is between 4% and 6%.”
- This is different from frequentist confidence intervals, which don’t allow direct probability statements about parameters.
4. Example – Coin Toss
We want to estimate the probability of heads (ppp).
- Prior: Uniform prior = Beta(1,1) → all values of $p \in [0,1]$ equally likely.
- Data: 10 tosses, 7 heads.
- Likelihood: Binomial(10, p).
- Posterior:
$P(p \mid \text{data}) \propto p^7 (1-p)^3 \cdot \text{Prior}(p)$
Since prior = Beta(1,1), posterior = Beta(1+7, 1+3) = Beta(8,4).
Posterior distribution: mean ≈ 0.67 → belief updated toward p ~ 0.67 after data.
5. Posterior vs Posterior Probability
- Posterior distribution: Full probability distribution of parameter after data.
- Posterior probability: Probability that a hypothesis/parameter lies in a certain region (e.g., P(p > 0.5 | data) = 0.9).
6. Applications
- A/B testing: posterior probability that variant B > A.
- Medicine: probability that treatment works after observing trial data.
- Machine learning: Bayesian neural networks, Naive Bayes.
- Forecasting: update predictions with new data.
7. Key Takeaways
- Posterior = updated belief about parameters/hypotheses after seeing data.
- Formula: Posterior ∝ Prior × Likelihood.
- More data → posterior is dominated by likelihood (prior matters less).
- Core of Bayesian inference → all decisions are based on posterior.
In short:
The posterior is the probability distribution of a parameter after observing data. It combines prior belief with the likelihood of observed evidence, giving the updated belief used in Bayesian inference.
Discover more from Insightful Data Lab
Subscribe to get the latest posts sent to your email.
