Covariance
Measure of joint variability between two random variables.
Covariance is a foundational concept in probability theory and statistics that measures how two random variables vary together. It is defined for two jointly distributed real-valued random variables with finite second moments as the expected value of the product of their deviations from their respective means. This can be expressed as the expected value of the product of the variables minus the product of their expected values, an identity useful for mathematical derivations but prone to numerical instability due to catastrophic cancellation. The sign of the covariance indicates the tendency of the linear relationship: a positive value suggests the variables tend to increase or decrease together, while a negative value suggests they tend to move in opposite directions. The magnitude of the covariance represents the geometric mean of the variances shared by the two variables, with a larger magnitude implying a stronger dependence. However, covariance is not dimensionless; its value is expressed in the product of the units of the two variables, so changing units (e.g., from meters to millimeters) scales the covariance proportionally. This dependence on units makes it difficult to assess the strength of a relationship from the covariance alone, especially when comparing different pairs of variables with different units. To address this, the correlation coefficient is used, which normalizes the covariance by dividing by the geometric mean of the total variances (the product of the standard deviations), yielding a dimensionless value between -1 and 1. A distinction is made between the population covariance, a parameter of the joint probability distribution, and the sample covariance, which describes a sample and serves as an estimate of the population parameter. For complex random variables, the definition includes complex conjugation of the second factor, and a related pseudo-covariance exists. For discrete random variables, the covariance can be computed from the joint probabilities and the means, or equivalently without directly referencing the means. A useful identity for computing covariance is Hoeffding’s covariance identity, which involves the joint cumulative distribution function and the marginal distributions. Random variables with zero covariance are termed uncorrelated, but uncorrelatedness does not imply independence, as demonstrated by a nonlinear r
- field
- Probability theory and statistics
- known_for
- Measuring joint variability of two random variables; basis for correlation coefficient
Lore & Background
Covariance is a statistical measure that captures the joint variability of two random variables, indicating the direction of their linear relationship. The sign of the covariance reveals the tendency in this relationship: a positive value suggests the variables tend to exhibit similar behavior, while a negative value indicates they tend to show opposite behavior. The magnitude of the covariance represents the geometric mean of the variances shared between the two variables, with a larger magnitude signifying a stronger mutual dependence. However, because covariance carries the units of measurement of the variables (for instance, the product of meters and grams), its magnitude is directly affected by those units. Changing the units, such as converting from meters to millimeters, proportionally alters the covariance value, making it difficult to assess the strength of the relationship from the covariance alone. To compare the strength of association between different pairs of random variables with potentially different units, the correlation coefficient is used, which normalizes the covariance to a dimensionless value between -1 and 1 by dividing by the product of the standard deviations. A key distinction exists between the covariance of two random variables as a population parameter, which is a property of their joint probability distribution, and the sample covariance, which describes a sample and serves as an estimate of the population parameter. Random variables with a covariance of zero are termed uncorrelated, though uncorrelatedness does not generally imply independence; for example, a non-linear relationship can yield zero covariance despite dependence. However, if two variables are jointly normally distributed, uncorrelatedness does imply independence.
Reader's Guide
Covariance serves as a fundamental concept in statistics, providing a measure of how two random variables change together. Its sign indicates the tendency of the linear relationship: positive when variables show similar behavior, negative when they show opposite behavior. However, because covariance has units and its magnitude is affected by those units, it is difficult to assess the strength of the relationship from covariance alone. To compare the strength of joint association between different pairs of random variables with possibly different units, the correlation coefficient is used, which normalizes covariance to a value between -1 and 1 by dividing by the product of the standard deviations. A distinction exists between the population covariance, a parameter of the joint probability distribution, and the sample covariance, which serves as both a sample descriptor and an estimate of the population parameter. The identity cov(X,Y) = E[XY] - E[X]E[Y] is useful for mathematical derivations but is susceptible to catastrophic cancellation in numerical computation.
Did You Know?
- Covariance is defined as the expected value of the product of deviations from individual expected values.
- The sign of covariance shows the tendency in the linear relationship between variables.
- Changing the units of measurement changes the covariance value proportionally.
- The correlation coefficient normalizes covariance to a value between -1 and 1.
More in Probability And Stochastic Processes 1-21
Spotted an error? Know more?
This is a living reference — every entry is fact-audited, and reader corrections feed straight into our audit queue. Suggest an edit · See this site's audit record
