Prior probability
Prior probability represents belief before new evidence.
A prior probability distribution represents the assumed probability of an uncertain quantity before any specific evidence is considered. For instance, it could model the expected proportions of voters supporting a candidate in a future election. The unknown quantity might be a model parameter or a latent variable, not necessarily an observable one. In Bayesian statistics, Bayes' rule dictates how this prior is updated with new data to yield a posterior probability distribution. Historically, priors were often chosen from a conjugate family relative to a given likelihood function, ensuring the posterior belonged to the same family and remained mathematically tractable. The advent of Markov chain Monte Carlo methods has largely alleviated this constraint.
Priors can be constructed in several ways. They may be derived from past information, such as previous experiments, or elicited from the subjective judgment of an experienced expert. When no information is available, an uninformative prior may be adopted, justified by the principle of indifference. In modern practice, priors are also selected for their mechanical properties, such as regularization or feature selection. Priors themselves often depend on hyperparameters, and uncertainty about these can be expressed through hyperprior distributions, forming hierarchical priors with multiple conditional levels.
Informative priors convey specific, definite information. For example, a prior for tomorrow’s noon temperature might be a normal distribution centered on today’s noon temperature, with variance reflecting day-to-day atmospheric variability. A strong prior is a type of informative prior where the prior information dominates the data, so the posterior changes little from the prior. A weakly informative prior provides partial information, steering analysis toward reasonable solutions without overly constraining results—for instance, using a normal distribution with a wide standard deviation to loosely bound a temperature estimate. Uninformative or diffuse priors express vague information, such as “the variable is positive,” and often yield results similar to conventional statistical analysis. The search for objectively required priors remains philosophically debated, with objective Bayesians, following Edwin T. Jaynes, arguing for priors based on symmetries and maximum entropy, while subjective Bayesians vi
- definition
- Assumed probability distribution before evidence
- role
- Foundation for Bayesian updating
- types
- Informative, weakly informative, uninformative
- construction_methods
- Past information, expert elicitation, principle of indifference, mechanical properties
- key_concept
- Hyperparameters and hierarchical priors
- controversy
- Objective vs. subjective Bayesianism
Lore & Background
A prior probability distribution represents the assumed probability of an uncertain quantity before any evidence is considered. Its appearance is that of a probability distribution, which can take various mathematical forms, such as a normal or beta distribution, depending on the nature of the unknown quantity. The range of a prior is defined by the possible values of that quantity; for instance, a prior for a proportion is confined between zero and one, while a prior for a temperature might span a wide range of degrees. The habitat of a prior is within Bayesian statistical analysis, where it serves as the starting point for updating with new data via Bayes' rule to produce a posterior distribution. A defining characteristic is that priors can be informative, weakly informative, or uninformative. An informative prior expresses specific, definite information, such as using today’s temperature as the expected value for tomorrow’s. A weakly informative prior provides partial information for regularization, loosely constraining estimates without dominating the data. An uninformative prior, often justified by the principle of indifference, assigns equal probabilities to all possibilities and expresses vague information like "the variable is positive." Historically, priors were often chosen from a conjugate family to ensure a tractable posterior, but modern Markov chain Monte Carlo methods have relaxed this constraint. Priors can also be hierarchical, with their own parameters—called hyperparameters—which may themselves have hyperpriors. The posterior from one problem can become the prior for another, and as evidence accumulates, the posterior is largely determined by the data rather than the original assumption, provided that assumption allowed for the evidence.
Reader's Guide
The concept of prior probability is central to Bayesian statistics, providing a formal mechanism to incorporate existing knowledge into statistical analysis. Its significance lies in enabling the updating of beliefs via Bayes' rule, producing posterior distributions that combine prior information with new data. The historical constraint of conjugate priors for tractability has been largely overcome by computational methods like Markov chain Monte Carlo, broadening the applicability of Bayesian methods. The construction of priors ranges from informative (based on past data or expert opinion) to weakly informative (for regularization) to uninformative (based on principles like indifference or invariance). The philosophical debate between objective and subjective Bayesianism highlights the tension between seeking logically required priors and acknowledging the subjective nature of prior beliefs. This debate, exemplified by arguments from Edwin T. Jaynes and others, underscores that priors represent a state of knowledge rather than an observer-independent feature of the world. The legacy of prior probability is its foundational role in Bayesian inference, enabling principled uncertainty quantification across scientific disciplines.
Did You Know?
- A prior can be determined from past information, such as previous experiments.
- The Haldane prior is an improper prior distribution with infinite mass.
- A strong prior is a type of informative prior where the prior dominates the data.
- The principle of indifference assigns equal probabilities to all possibilities.
More in Probability & Statistics 1-24
Spotted an error? Know more?
This is a living reference — every entry is fact-audited, and reader corrections feed straight into our audit queue. Suggest an edit · See this site's audit record
