P-value
Probability of extreme results under the null hypothesis.
The p-value is a statistical concept used in null-hypothesis significance testing. It quantifies the probability of obtaining test results at least as extreme as the observed result, assuming the null hypothesis is true. Despite its widespread use in academic publications across many quantitative fields, the p-value is frequently misinterpreted and misused, a topic of ongoing discussion in mathematics and metascience.
In null-hypothesis significance testing, the p-value is the probability, under the assumption that the null hypothesis is correct, of obtaining a test statistic at least as extreme as the one actually observed. A very small p-value indicates that such an extreme outcome would be very unlikely if the null hypothesis were true. The null hypothesis itself is a default conjecture that a specified property, such as a correlation or difference between means, does not exist in the population of interest. The p-value is used to quantify the statistical significance of a result; lower p-values are generally taken as stronger evidence against the null hypothesis, and a result is deemed statistically significant if it leads to the rejection of that hypothesis. However, rejecting the null hypothesis does not specify which alternative values are most plausible, nor does it indicate the real-world importance of any deviation. The precision of the test increases with more independent observations, but this also heightens the need to evaluate the scientific relevance of any detected effect. The p-value is a random variable; if the null hypothesis is true and the test statistic’s distribution is continuous, the p-value is uniformly distributed between zero and one. Different p-values from independent data sets can be combined, for instance using Fisher’s combined probability test. The threshold for significance, commonly set at 0.05, was originally proposed by Ronald Fisher in 1925. The American Statistical Association has formally stated that p-values do not measure the probability that a studied hypothesis is true, nor the size or importance of an effect, and that they require context and other evidence to be meaningful. Nevertheless, a later task force concluded that p-values and significance tests, when properly applied and interpreted, increase the rigor of conclusions drawn from data.
- field
- Statistics
- known_for
- Quantifying statistical significance in null-hypothesis testing
- associated_organization
- American Statistical Association (ASA)
- related_concept
- Null hypothesis
Lore & Background
In null-hypothesis significance testing, the p-value is defined as the probability, under the null hypothesis, of obtaining a real-valued test statistic at least as extreme as the one observed. For a one-sided right-tail test, this is Pr(T ≥ t | H₀); for a left-tail test, Pr(T ≤ t | H₀); and for a two-sided test, often 2 min{Pr(T ≥ t | H₀), Pr(T ≤ t | H₀)}. The p-value is a random variable, as it depends on the chosen test statistic; when the null hypothesis is true and its distribution is continuous, the p-value is uniformly distributed between 0 and 1. A very small p-value indicates that the observed result would be very unlikely under the null hypothesis, leading to its rejection if the p-value falls below a predefined significance level, commonly set at 0.05. This threshold was originally proposed by Ronald Fisher. Despite widespread use in academic publications across quantitative fields, p-values are frequently misinterpreted and misused, a major topic in mathematics and metascience. The American Statistical Association formally stated that p-values do not measure the probability that the studied hypothesis is true, nor the size or importance of an effect, and do not alone provide good evidence for a model without context. However, when properly applied and interpreted, p-values and significance tests increase the rigor of conclusions. Different p-values from independent data sets can be combined, for instance using Fisher's combined probability test.
Reader's Guide
The p-value is a random variable that depends on the chosen test statistic; if the null hypothesis is true and its distribution is continuous, the p-value is uniformly distributed between zero and one. In practice, only a single p-value is typically observed for a given hypothesis, and no effort is made to estimate its underlying distribution. The p-value is used in null-hypothesis significance testing to quantify statistical significance: the lower the p-value, the lower the probability of obtaining the observed result if the null hypothesis were true. A result is deemed statistically significant when it allows rejection of the null hypothesis, with smaller p-values generally considered stronger evidence against it. However, rejecting the null hypothesis does not specify which alternative is most plausible, nor does it indicate the size or practical relevance of an effect. The significance level α, commonly set at 0.05, is a threshold chosen by the researcher before examining the data, not derived from it. Misinterpretation and misuse of p-values are widespread in quantitative fields, prompting the American Statistical Association to formally state that p-values do not measure the probability that a hypothesis is true, the probability that data arose from random chance alone, the size of an effect, or the importance of a result without additional context. Despite these cautions, a later task force concluded that p-values and significance tests, when properly applied and interpreted, increase the rigor of conclusions drawn from data.
Did You Know?
- The p-value is the probability of obtaining test results at least as extreme as the observed result, assuming the null hypothesis is true.
- Different p-values based on independent data sets can be combined using Fisher's combined probability test.
More in Probability & Statistics 1-24
Spotted an error? Know more?
This is a living reference — every entry is fact-audited, and reader corrections feed straight into our audit queue. Suggest an edit · See this site's audit record
