Skip to content

Correlation and covariance calculator

Calculate Pearson correlation and sample or population covariance for two paired lists, with a scatter plot.

Paired data

Processed in your browser.

2–200 values. Separate with spaces, commas or semicolons. Decimal point: 1.5.

2–200 values. Separate with spaces, commas or semicolons. Decimal point: 1.5.

Result

Enter two lists, then calculate.

How to use

  1. Enter the two lists in matching order: the first number in one list belongs with the first number in the other. Keep each list in one consistent unit and remove headings or unit symbols.
  2. Use spaces, commas or semicolons between numbers, and a dot for decimals, such as 1.5. A comma always starts another value, even when your display language uses decimal commas. Scientific notation such as 2e-3 is accepted.
  3. Results and charts update as soon as both lists are valid. Incomplete or mismatched input leaves the results empty. Press Calculate if you want to see an explanation of an input error.
  4. Copy results copies the numerical summary as text. Clear empties the inputs and results. The chart is an on-screen explanation; this tool does not export an image or data file.

Example

Do X and Y move together? Each dot is a pair. Here Y rises with X: the correlation is positive.

YX0214263

cov(X,Y) = Σ[(X − X̄)(Y − Ȳ)] / (n − 1)

X: 1 2 3 · Y: 2 4 6

r = 1 · Sample covariance = 2

The input fields start with this example. Edit them to use your own data, then calculate.

What the results mean

Use this calculator to describe how two measured quantities vary together. It is useful for checking classroom calculations, exploring paired measurements, or comparing the direction of association before fitting a model. Each dot represents one original pair (X, Y), not a separate summary or a randomly sampled point.

Let dx = X − mean(X) and dy = Y − mean(Y). Pearson r = Σ(dx × dy) / √[Σ(dx²) × Σ(dy²)]. Sample covariance divides Σ(dx × dy) by n − 1; population covariance divides it by n. Both conventions are shown so you can choose the one that matches your task.

Pearson r lies between −1 and 1 when both lists vary. Positive values describe an upward linear tendency; negative values describe a downward tendency. Values near zero indicate little linear association, but can hide a curved relationship. Read the scatter plot before attaching a verbal strength label to a rounded coefficient.

Covariance has the product of the original units. Changing meters to centimeters changes its magnitude, even though the underlying relationship is unchanged. Correlation is unitless and is unchanged by positive rescaling. Use population covariance for the complete set you intend to describe, and sample covariance when treating the observations as a sample.

Worked examples

X = 1, 2, 3 and Y = 2, 4, 6 give n = 3, r = 1, sample covariance = 2 and population covariance ≈ 1.33. Every dot lies on an upward line. The two covariance results differ only because their divisors are 2 and 3.

X = 1, 2, 3 and Y = 3, 2, 1 give r = −1, sample covariance = −1 and population covariance ≈ −0.67. Replacing Y with 5, 5, 5 instead gives covariance 0 and undefined r: no Y variation remains to normalize.

Input, precision and output

For meaningful comparison, preserve which measurements belong together. Sorting X and Y independently creates new pairs and can manufacture a relationship. A large coefficient can also be driven by a single unusual observation or by combining different groups. This summary does not establish causation, independence, statistical significance or a useful prediction model.

Both lists need the same 2–200 values. Missing entries, text, NaN and infinity are rejected. Zero, negative numbers and repeated values are allowed. Remove a missing observation from both lists, rather than shifting only one list.

Nonzero input magnitudes must be between 1e-100 and 1e100. Calculations use browser floating-point numbers, with about 15–16 significant digits. Distinct entries that collapse to the same stored number are rejected; reduce an offset or choose more suitable units before retrying.

Most displayed results use at most two decimal places. Small nonzero values use scientific notation, and calculations keep unrounded intermediate values. Offset/scaled axes explicitly show how to recover the original values, so tiny differences remain visible beside a large baseline.

Frequently asked questions

Why is correlation undefined for a constant list?

Its sum of squared deviations is zero, so the denominator of r is zero. Zero covariance still makes sense here; reporting r = 0 would incorrectly imply a measured lack of linear association.

Can r = 0 still hide a relationship?

Yes. For example, values arranged in a U-shaped pattern can have little linear correlation. The scatter plot is essential for seeing shape, clusters and outliers that a single coefficient cannot describe.

Why do the sample and population values differ?

They divide the same centered cross-product sum by n − 1 and n, respectively. Choose the convention from the purpose of the analysis; the calculator cannot determine whether your data represent a full population.

Does a correlation near 1 prove cause and effect?

No. A common cause, selection of observations, or coincidence can produce association. With only two varying pairs, the coefficient is always +1 or −1, so additional evidence matters.