Bayesian Inference
This note is from a statistics point of view and the textbook source is "All of Statistics", chapter 11
Bayesian Philosophy
Postulates:
- Probability describes degree of belief, not limiting frequency
- We can make probability statements about parameters, even though they are fixed constants
- we make inferences about a parameter
by producing a probability distribution for . Inferences, such as point estimates and interval estimates, may then be extracted from this distribution.
First postulates means that for example it is “correct” to say that “the probability that Alber Einstein drank a cup of tea on August 1, 1948” is 0.35. This does not refer to frequency but the my strenght of belief that the proposition is true.
The Bayesian Method
Bayesian Inference steps:
- We choose a probability density
) called the prior distribution that expresses our belief about a parameter before we see any data. - We choose a statistical mdoel
that reflects our belief about given . - After observing data
we update our belief and calculate the posterio distribution .
To see how the third step is carried out, first suppose that
Since we are treating
- This is the Bayes Theorem (Statistics)
Considering
The version for continuos variables is obtained by using probability density functions:
replace
is the likelihood
Then:
notation for ) (and also is a normalization constant gived by:
Posterior is proportional to Likelihood times Prior or in symbols:
The constant isn’t considered and can be recovered later.
With the posterio distribution, we can get a point estimate by summarizing the center of the posterior (typically mean or mode of the posterior)
The posterior Covariance, Variance and Mean|mean is:
We can also obtain a Bayesian interval estimate. We find
Let
So
Example with Bernoulli
Let
By Bayes’ Theorem, the posterior has the form:
where
Recall that a random variable has a Beta distribution with parameters
We see that the posterior for
That we write as:
Notice that we have figured out the normalizing constant without actually doing the integral
It is instructive to rewrite the estimator as:
is the MLE is the prior mean
A 95 percent posterior interval can be obtained by numerically finding
Now, suppose that instead of a uniform prior, we use the prior
The flat prior is just the special case with
The posterior mean is:
is the prior mean.
In the previous example, the prior was a Beta distribution and the posterior was a Beta distribution. When the prior and the posterior are in the same family, we say that the prior is conjugate with respect to the model.
Example with Gaussian
Let
Suppose we take as prior
i.e the standard error of the MLE .
This is another example of a conjugate prior.
Since
Now, say we want to find the interval
We choose
We want to find
We know that
implying that
So a 95 percent Bayesian interval is
Large Sample Properties of Bayes Procedures
We saw in the previous two examples that posterio mean was close to MLE. This is true in greater generality
Theorem Let
There is also a Bayesian delta method. Let
- where
and
Flat Priors, Improper Priors and Noninformative Priors
An important question is: where does one get the prior
Moreover, injecting subjective opinion into the analysis is contrary to the goal of making scientific inference as objective as possible.
Noninformative prior
An alternative is try to define some sort of noninformative prior, one obvious candidate is the flat prior
Improper Priors
Let
But this is not a problem because applying the Bayes theorem and computing the posterior density by multiplting the prior and the likelihood we obtain:
Flat priors are not invariant. Let
which is not flat.
But if we are ignorant about
Jeffreys’s Prior: it’s a rule for creating priors:
Example
Consider the Bernoulli
Jeffrey’s rule says to use the prior:
In a multiparameter problem, the Jeffrey’s prior is defined to be