Standard Deviation in R: Calculate It Correctly

R code calculating standard deviation for a numeric vector

Use sd() to calculate standard deviation in R for a numeric vector or a data-frame column. For standard deviation in R, the main interpretation decision is whether your data represent a sample or an entire population. R’s default result is the sample standard deviation.

The function also returns NA when missing values are present unless you explicitly remove them during the calculation.

Standard deviation in R with sd()

Pass a numeric vector to sd():

x <- c(12, 15, 14, 10, 9)

sd(x)

This returns approximately 2.54951. The values have a mean of 12, and R divides the sum of squared deviations by 4, which is the sample denominator: the number of observations minus one.

You can apply the same standard deviation function to a numeric column in a data frame:

scores <- data.frame(score = c(12, 15, 14, 10, 9))

sd(scores$score)

This also returns approximately 2.54951. Use scores[[“score”]] as an alternative to scores$score. Both expressions select the underlying numeric vector. By contrast, scores[“score”] returns a one-column data frame, which is not the intended input for this calculation.

The standard deviation function in R handles missing values

By default, sd() does not ignore missing observations:

measurements <- c(4, 7, NA, 10)

sd(measurements)

The result is NA because the missing value propagates through the calculation. Add na.rm = TRUE to exclude missing values:

sd(measurements, na.rm = TRUE)

This calculates the sample standard deviation of 4, 7, and 10. The argument removes only the NA values; it does not replace them or estimate what they might have been.

Use the same argument with a data-frame column:

sd(scores$score, na.rm = TRUE)

If no nonmissing values remain, or only one nonmissing value remains, a standard deviation cannot be estimated and R returns NA.

Sample versus population standard deviation

sd() uses the sample standard deviation formula:

sqrt(sum((x – mean(x))^2) / (n – 1))

Here, n is the number of observed values. The n – 1 denominator estimates the variability of a larger population from a sample, so do not label R’s default result as a population standard deviation.

Use a population standard deviation when your vector contains every member of the population being measured. The population formula divides by n:

sqrt(sum((x – mean(x))^2) / n)

In R, calculate it directly with:

population <- c(12, 15, 14, 10, 9)

sqrt(mean((population – mean(population))^2))

If missing values are possible, remove them first:

complete <- population[!is.na(population)]

sqrt(mean((complete – mean(complete))^2))

You can also convert a sample result to a population result with sd(complete) * sqrt((n – 1) / n), where n <- length(complete).

Check R standard deviation with a worked vector

This vector makes the denominator difference easy to verify:

values <- c(2, 4, 4, 4, 5, 5, 7, 9)

mean(values)

sd(values)

The mean is 5. The sum of squared deviations from that mean is 32, and there are eight observations. Therefore:

  • Sample standard deviation: sqrt(32 / 7), which is approximately 2.13809 and matches sd(values).
  • Population standard deviation: sqrt(32 / 8), which equals 2.

Use sd() for the usual sample estimate, add na.rm = TRUE when missing values should be excluded, and use the denominator-n formula when the data represent the complete population.