Pearson product moment correlation coefficient¶
Correlation is a commonly used method to examine the relationship between quantitative variables. The most commonly used statistic is the linear correlation coefficient, $r$, which is also known as the Pearson product moment correlation coefficient in honor of its developer, Karl Pearson. It is given by
$$r = \frac{\sum_{i=1}^n(x_i- \bar x)(y_i - \bar y)}{\sqrt{\sum_{i=1}^n(x_i- \bar x)^2}\sqrt{\sum_{i=1}^n(y_i- \bar y)^2}}=\frac{s_{xy}}{s_x s_y}\text{,}$$
where $s_{xy}$ is the covariance of $x$ and $y$, $s_x$ and $s_y$ are the standard deviations of $x$ and $y$, respectively. By dividing by the sample standard deviations, $s_x$ and $s_y$, the linear correlation coefficient, $r$, becomes scale independent and takes values between $-1$ and $1$.
Important note: Arithmetic mean and standard deviation are exclusive parameter of the normal distribution adn only meaningful on open interval scale. Consequently, the correlation coefficient inherits these two limitations. Analyzing real world data with positive measures only, a log-transformation is mandatory to fulfill the requirement of (additive) interval scale. In case of double constraint feature space, a logistic transformation is a feasible way to obtain an unlimited scale with meaningful additive measures. Subordinated transformation are probably necessary to satisfy the need for normality.
The linear correlation coefficient measures the strength of the linear relationship between two variables. If $r$ is close to $\pm 1$, the two variables are highly correlated, and when plotted on a scatter plot, the data points cluster around a line. If $r$ is far from $\pm 1$, the data points are more widely scattered. If $r$ is near $0$, the data points are essentially scattered around a line indicating that there is almost no linear relationship between the variables.
An interesting property of $r$ is that its sign reflects the slope of the linear relationship between two variables. A positive value of $r$ suggests that the variables are positively linearly correlated, indicating that $y$ tends to increase linearly as $x$ increases.
A negative value of $r$ suggests that the variables are negatively linearly correlated, indicating that $y$ tends to decrease linearly as $x$ increases.
There is no unambiguous classification rule for the quantity of a linear relationship between two variables. However, the following table may serve an as rule of thumb how to address the numerical values of the Pearson product moment correlation coefficient:
$$ \begin{array}{lc} \hline \ \text{Strong linear relationship} & r > 0.9 \\ \ \text{Medium linear relationship} & 0.7 < r \le 0.9\\ \ \text{Weak linear relationship} & 0.5 < r \le 0.7 \\ \ \text{No or doubtful linear relationship} & 0 < r \le 0.5 \\ \hline \end{array} $$
Pearson's correlation assumes the variables to be roughly normally distributed, and it is not robust in the presence of outliers.
In a later section on linear regression we discuss the coefficient of determination, $R^2$, a descriptive measure for the quality of linear models. There is a close relation between $R^2$ and the linear correlation coefficient, $r$. The coefficient of determination, $R^2$, equals the square of the linear correlation coefficient, $r$:
$$\text{coefficient of determination }(R^2) =r^2 $$
Pearson correlation coefficient: An example¶
For this example, we will use weather data from the DWD weather station Berlin-Dahlem (FU), downloaded from Climate Data Center. For the purpose of this tutorial the dataset is provided here.
Here you can find a documentation of the data set,
| STATIONS_ID | MESS_DATUM | QN_3 | FX | FM | QN_4 | RSK | RSKF | SDK | SHK_TAG | NM | VPM | PM | TMK | UPM | TXK | TNK | TGK | eor | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 403 | 19500101 | NaN | NaN | NaN | 5 | 2.2 | 7 | NaN | 0.0 | 5.0 | 4.0 | 1025.60 | -3.2 | 83.00 | -1.1 | -4.9 | -6.3 | eor |
| 1 | 403 | 19500102 | NaN | NaN | NaN | 5 | 12.6 | 8 | NaN | 0.0 | 8.0 | 6.1 | 1005.60 | 1.0 | 95.00 | 2.2 | -3.7 | -5.3 | eor |
| 2 | 403 | 19500103 | NaN | NaN | NaN | 5 | 0.5 | 1 | NaN | 0.0 | 5.0 | 6.5 | 996.60 | 2.8 | 86.00 | 3.9 | 1.7 | -1.4 | eor |
| 3 | 403 | 19500104 | NaN | NaN | NaN | 5 | 0.5 | 7 | NaN | 0.0 | 7.7 | 5.2 | 999.50 | -0.1 | 85.00 | 2.1 | -0.9 | -2.3 | eor |
| 4 | 403 | 19500105 | NaN | NaN | NaN | 5 | 10.3 | 7 | NaN | 0.0 | 8.0 | 4.0 | 1001.10 | -2.8 | 79.00 | -0.9 | -3.3 | -5.2 | eor |
| ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... |
| 26293 | 403 | 20211227 | NaN | NaN | NaN | 3 | 0.0 | 8 | 0.183 | 0.0 | 5.9 | 3.8 | 998.13 | -3.7 | 79.67 | -0.7 | -7.9 | -9.9 | eor |
| 26294 | 403 | 20211228 | NaN | NaN | NaN | 3 | 1.5 | 6 | 0.000 | 0.0 | 6.4 | 5.3 | 990.17 | -0.5 | 88.46 | 2.7 | -3.9 | -5.1 | eor |
| 26295 | 403 | 20211229 | NaN | NaN | NaN | 3 | 0.3 | 6 | 0.000 | 0.0 | 7.5 | 8.2 | 994.40 | 4.0 | 100.00 | 5.6 | 1.8 | 0.0 | eor |
| 26296 | 403 | 20211230 | NaN | NaN | NaN | 3 | 3.2 | 6 | 0.000 | 0.0 | 7.9 | 11.5 | 1001.70 | 9.0 | 98.54 | 12.7 | 4.6 | 2.3 | eor |
| 26297 | 403 | 20211231 | NaN | NaN | NaN | 3 | 5.5 | 6 | 0.000 | 0.0 | 7.7 | 12.5 | 1004.72 | 12.8 | 84.96 | 14.0 | 11.5 | 10.7 | eor |
26298 rows × 19 columns
For our demonstration purpose we consider a sample of the data set considering the following features:
- TNK: Minimum of the day of the air temperature in 2 m height
- TXK: Maximum of the day of the air temperature in 2 m height
The scatter plot indicates that there exists a linear relationship between the two variables under consideration.
For the sake of this exercise we calculate the linear correlation coefficient by hand at first and, then we apply the np.corrcoef() function from the NumPy package in Python.
Recall the equation from above:
$$r = \frac{\sum_{i=1}^n(x_i- \bar x)(y_i - \bar y)}{\sqrt{\sum_{i=1}^n(x_i- \bar x)^2}\sqrt{\sum_{i=1}^n(y_i- \bar y)^2}}=\frac{s_{xy}}{s_x s_y}$$
np.float64(0.916)
np.float64(0.916)
Recall the conditions on calculating the mean and the standard deviations. Both measures are included in the definition of the correlation coefficient!
In order to achieve reliable correlations, we have to transform the data to normal distributions with data in the whole real space!
Step 1: Min-Max scaling Apply the min-max transformation to transform the data into the open interval (0,1). Using the classical scaling
$$x'=\frac{x-x_{min}}{x_{max}-x_{min}}$$
transforms the data into the range $[0,1]$ that includes "0" and "1". We want to avoid zero-valued values, because in the second step we apply the logarithm. Therefore, we use slightly larger boundary values $x_{min}-1$ and $x_{max}+1$. Thus
$$x'=\frac{x-x_{min}+1}{x_{max}-x_{min}+2}\, .$$
Let us check how the histograms look like:
We can recognize that the min-max transformation only changed the scale. The data is now in the open interval between zero and one. But it does not change the distribution!
Step2: Logistic transformation
As a second step, we use the logistic transformation $$x''=\log\frac{x'}{1-x'}$$
and also compare the distributions of the temperature data:
So now our data is symmetrically distributed around zero and mapped on the whole real space $\mathbb{R}$.
Now, we can calculate the correlation coefficient with the above formula that contains the arithmetic mean and the standard deviation.
Let us first calculate the correlation coefficient "by hand":
Finally, we apply the in-build np.corrcoef() function:
r = 0.916
Perfect. The three calculations yield the nearly same result! The linear correlation coefficient evaluates to $r = 0.9$. Thus, we may conclude that there is a strong linear correlation between the maximal and minimal temperature and the deviation from the normal distribution is subsidiary in our example!
Of course a correlation analysis is not restricted to two variables. We present multiple and partial correlation analysis in our chapter about Multiple linear Regression MLR
But for explorative statistics, we are able to conduct a pairwise correlation analysis for more than two variables.
The pd.corr() comes in handy for correlation multiple columns of a pandas DataFrame.
Note: Correlation functions implemented in Python, such as
np.corrcoef()orpd.corr(), include different types of correlation coefficients such as Pearson's, Spearman's and Kendall's correlation coefficients. To pick one particular formula you add the argumentmethodto the function call. The Pearson's correlation coefficient is the default setting.
We will have a closer look at Spearman's rank correlation coefficient on the following page.
Note: It is always recommended advancing with a statistical test, in order to assess whether the result is statistically significant or whether the variation is just due to chance. Check out the sections on Hypothesis Testing for further information!
Spurious correlations¶
A spurious correlation occurs, when there are initial dependencies between two variables, which are not founded by the measured values. In other words: When the observed correlation is rooted in common dependencies of both variables or complete coincidence, rather than a relation in causality.
A good example to emphasize this is the occurrence of childbirths and stork sightings. Both the stork population and the birth rates of humans peak in spring to summer and therefore correlate in timing. Whatsoever, both variables have no effect on each other of course Sappsford & Jupp.
A spurious correlation can occur for a number of reasons:
For once, the stork example shows, that the similar timing of unrelated events can cause a spurious correlation. Here, the reoccurring seasons are the common, governing factor for both variables.
Secondly, a constant sum causes spurious correlations. This is commonly seen in concentrations, like oxygen or carbon-dioxide concentrations in air. If the concentration of one gas in the air mixture is reduced, all other gases will rise in concentration, as the volume will stay constant at $100$%. This is, of course not, for the absolute mass.
A third, common reason is the skewness of the variables. Positive skewness will lead to exaggerated correlations, due to the distributions shape. The correlation itself is calculated by the arithmetic mean, which is very sensitive to positive outliers.
Remember: Arithmetic mean and standard deviation are exclusive parameters of the normal distribution and, thus both variables have to be normally distributed for a reliable estimation of correlation.
Let us have a closer look on the geometric meaning of the correlation coefficient:
Correlation, cosine and scalar product¶
We may consider two variable $x$ and $y$, both on a metric scale, as vectors $\vec x = (x_i - \bar x; i=1,...,n )$ and $\vec y = (y_i - \bar y; i=1,...,n )$ in $\mathbb{R}^n$ .
In this scenario, the correlation can be interpreted as cosine of the angle between $\vec x$ and $\vec y$:
$$\cos\!\big(\vec{x},\vec{y}\big)= \frac{\langle \vec{x},\vec{y}\rangle}{\|\vec{x}\|\|\vec{y}\|}= \frac{\sum_{i=1}^n x_i y_i}{\sqrt{\sum_{i=1}^n x_i^2}\sqrt{\sum_{i=1}^n y_i^2}}= \frac{\sum_{i=1}^n (x_i-\bar{x})(y_i-\bar{y})}{\sqrt{\sum_{i=1}^n (x_i-\bar{x})^2}\sqrt{\sum_{i=1}^n (y_i-\bar{y})^2}}= r$$
By this relationship, we can examine three common cases for alternating $\vec y$:
- $r=0$ : $\cos\!\big(\vec{x},\vec{y}\big)=0 \rightarrow\angle(\vec{x},\vec{y}\big)=90^\circ$
Both variable vectors are rectangular to each other!
- $r = 1$ : $\cos\!\big(\vec{x},\vec{y}\big)=1\rightarrow \angle(\vec{x},\vec{y}\big)=0^\circ$
Both variable vectors point to the same direction!
- $r=-1$ : $\cos\!\big(\vec{x},\vec{y}\big)=-1\rightarrow \angle(\vec{x},\vec{y}\big)=180^\circ$
Both variable vectors point in the exact opposite direction of each other!
Considering the meaning of $r=(-1,0,1)$ this context seems logical!
Here we see, that the correlation coefficient $r$ is only meaningful in a data space with an orthonormal base (Euclidean data space)
Under these requirements the correlation coefficient $r$, derived by the $cos$ gives the x-axis extent to which the two variable vectors $\overrightarrow{x}$ and $\overrightarrow{y}$ are pointing towards the same direction:
Since the correlation coefficient is the normalized covariance:
$$\cos^2 + \sin^2 = r^2 +( 1-r)^2 = 1 $$
Thus, the coefficient of determination (COD) also has a geometric meaning and we can define the squared sine as the Coefficient of non-determination (COnD) and split the variance meaningful into a "common portion" and a "unique portion"!
However, covariance, correlations, and COD/COnD together with their geometric implications are only meaningfull in a reference system with an orthogonal base (= linear independent variable space)!
In cases of inherent dependencies through constant sum constraints, compositional dependencies or external common relations (e.g. seasonal or spatial dependencies resp. temporal/spatial autocorrelation), the reference system becomes oblique!
If two variables are spurious correlated (linear dependent), it means that their bases become oblique to each other through a common perspective (variable). Imagine: it is impossible to project a three-axis rectangular object onto a 2D plane preserving right angles.
This leads to the $\cos$, the correlation coefficient $r$ and the coefficient of determination $R^2$ to become meaningless:
Important conclusion:\
- Spurious correlation indicates an oblique projection of our data
- Spurious correlations obscure the message in our data
- Spurious correlations produce artificial patterns
- Correlation analysis needs interval scales!
- Correlation on ratio scales is meaningless without proper transformation!
- Correlation of variables in a constrained data space is nonsense!
- Variables with units of %, mg/g, ppm, etc. have to be transformed before applying correlation/covariance analyses (spurious correlation by constant sum constraint).
Citation
The E-Learning project SOGA-Py was developed at the Department of Earth Sciences by Annette Rudolph, Joachim Krois and Kai Hartmann. You can reach us via mail by soga[at]zedat.fu-berlin.de.

You may use this project freely under the Creative Commons Attribution-ShareAlike 4.0 International License.
Please cite as follow: Rudolph, A., Krois, J., Hartmann, K. (2023): Statistics and Geodata Analysis using Python (SOGA-Py). Department of Earth Sciences, Freie Universitaet Berlin.