Category: R Analysis

  • Rename Function in R: How to Rename Columns

    Rename Function in R: How to Rename Columns

    The rename function in R is commonly dplyr::rename(). It changes column labels without changing rows, values, or columns you do not select. For direct, readable mappings, use new_name = old_name. For full-vector control, use base R’s names().

    The examples below assume an existing data frame named df. Both approaches can rename columns in R while preserving the data itself.

    How does the rename function in R work?

    dplyr::rename() returns a modified data frame and uses a deliberately explicit mapping:

    • The new column name goes on the left.
    • The existing column name goes on the right.
    • Columns not listed in the call keep their names and positions.

    Therefore, full_name = first_name changes first_name to full_name. Reversing the order produces the wrong result or an error if the proposed old name does not exist.

    How do you rename columns in R with dplyr?

    Use rename() for one column or several named changes. Refer to the package explicitly or load it with library(dplyr).

    One column: df2 <- dplyr::rename(df, full_name = first_name)

    This creates df2 with first_name renamed to full_name. The other columns remain unchanged. Add comma-separated mappings for multiple columns:

    Several columns: df2 <- dplyr::rename(df, full_name = first_name, test_score = score)

    Use rename_with() when a rule should transform selected names rather than mapping each name manually. This is useful for capitalization, prefixes, suffixes, or consistent formatting:

    All names: df2 <- dplyr::rename_with(df, toupper)

    Selected names: df2 <- dplyr::rename_with(df, ~ paste0(“score_”, .x), .cols = dplyr::starts_with(“score”))

    The function receives the selected names as .x. In the second example, only names beginning with score receive the score_ prefix; other column names are untouched.

    How can you rename in R with base R?

    Base R stores a data frame’s column labels in its names vector. To change one column by its existing name, assign through a logical match:

    One column: names(df)[names(df) == “first_name”] <- “full_name”

    This method changes every matching name and leaves all other names intact. You can also rename by position, but position-based assignments are more fragile if the data-frame layout changes.

    To replace the complete name vector, assign one new name for every column:

    Full vector: names(df) <- c(“full_name”, “test_score”, “status”)

    Full-vector assignment is appropriate when you know the exact column order. It changes every label, so it can accidentally rename columns you intended to preserve.

    How do you check renamed columns and preserve the rest?

    Inspect the resulting labels with names() immediately after either method:

    names(df2)

    For an automated check, verify the new name exists and the old name does not:

    stopifnot(“full_name” %in% names(df2), !”first_name” %in% names(df2))

    To confirm that untouched columns survived a dplyr rename, save the original names first:

    old_names <- names(df)
    df2 <- dplyr::rename(df, full_name = first_name)
    stopifnot(all(setdiff(old_names, “first_name”) %in% names(df2)))

    Use rename() for explicit old-to-new mappings, rename_with() for repeatable naming rules, and names() when you need direct control of one label or the entire name vector.

  • Filter in R: How to Filter Data Frame Rows

    Filter in R: How to Filter Data Frame Rows

    To filter in R, use dplyr::filter() to keep data-frame rows that satisfy one or more conditions. It handles numeric comparisons, text matches, missing values, and combinations of conditions in a readable way.

    To filter data in R, remember that filtering changes which rows remain; it does not choose which columns are displayed. The R filter workflow below uses a small data frame and then shows the equivalent base R approach.

    How do you filter in R with dplyr::filter(), not stats::filter()?

    Start with a data frame containing numeric, text, and missing values:

    sales <- data.frame(product = c(“A”, “B”, “A”, “C”), region = c(“East”, “West”, NA, “East”), units = c(12, 7, NA, 20))

    Use dplyr::filter() with a comparison such as greater than or equal to:

    large_sales <- dplyr::filter(sales, units >= 10)

    This keeps rows where units is at least 10. Common comparison operators are == for equal to, != for not equal to, >, <, >=, and <=. The function returns all columns for the matching rows. In contrast, dplyr::select(sales, product, units) selects columns rather than filtering rows.

    Use dplyr::filter() for row operations. The separate stats::filter() function is designed for time-series and other filtering operations on vectors, so it is not the row-filtering function used here.

    How do you combine conditions with AND, OR, and negation?

    Use & for AND when every condition must be true:

    east_large <- dplyr::filter(sales, region == “East” & units >= 10)

    Use | for OR when either condition can be true:

    east_or_west <- dplyr::filter(sales, region == “East” | region == “West”)

    Use ! for negation. For example, this keeps rows whose region is not West:

    not_west <- dplyr::filter(sales, !(region == “West”))

    Parentheses make compound logic easier to read and prevent ambiguity. Use & and | for row-by-row conditions; && and || are scalar operators and are generally inappropriate for filtering a full column.

    How do you filter text and missing values safely?

    Match text by comparing a character column with a quoted value:

    east_sales <- dplyr::filter(sales, region == “East”)

    R represents a missing value as NA. A comparison such as region == “East” produces NA when region is missing, not TRUE or FALSE. dplyr::filter() keeps only rows where the condition is TRUE, so those uncertain rows are excluded.

    Test missing values explicitly with is.na() or its negation:

    missing_region <- dplyr::filter(sales, is.na(region))

    known_region <- dplyr::filter(sales, !is.na(region))

    Combine the missing-value check with another condition when needed:

    known_large <- dplyr::filter(sales, !is.na(units) & units >= 10)

    How does an R filter work with base R logical indexing?

    Base R filters rows by placing a logical condition before the comma inside square brackets. Include an explicit missing-value check so an NA does not create an unintended missing row in the result:

    large_sales_base <- sales[!is.na(sales$units) & sales$units >= 10, , drop = FALSE]

    The expression before the comma chooses rows; the blank expression after the comma keeps all columns. A text condition with AND works the same way:

    east_base <- sales[!is.na(sales$region) & sales$region == “East”, , drop = FALSE]

    For OR, use parentheses around the alternatives:

    east_or_west_base <- sales[!is.na(sales$region) & (sales$region == “East” | sales$region == “West”), , drop = FALSE]

    To select columns instead of rows, place column names after the comma: sales[, c(“product”, “units”), drop = FALSE]. That distinction keeps base R logical indexing focused on rows while column selection remains a separate operation.

  • How to Clear the Global Environment in R

    How to Clear the Global Environment in R

    Use rm() to remove selected objects or clear the objects listed in R’s global environment. A targeted command is safer when you need to preserve some work; a full rm(list = ls()) clears the ordinary, visible objects in the current environment.

    How to clear the global environment in R with rm(list = ls())

    Run this command at the Global Environment prompt:

    rm(list = ls())

    It removes every object returned by ls() in that environment. Check the contents before and after to confirm the result:

    ls()
    [1] “data” “model” “results”

    rm(list = ls())

    ls()
    character(0)

    The result character(0) means that no visible objects remain. The command is evaluated in the current environment, so run it from the Global Environment if that is the environment you intend to clear.

    By default, ls() does not list hidden names beginning with a period. To include those names in a more comprehensive cleanup, use:

    rm(list = ls(all.names = TRUE))

    Use the extended form only when you also intend to remove hidden objects managed or created in that environment.

    Clear environment in R selectively with rm(object) and rm(list = c(‘data’, ‘model’))

    Remove one object by passing its name to rm():

    object <- 42
    rm(object)

    Verify that the object is gone:

    ls()
    character(0)

    For a real workspace, inspect the names first and remove only the objects you no longer need:

    ls()
    [1] “data” “model” “results”

    rm(list = c(‘data’, ‘model’))

    ls()
    [1] “results”

    The names supplied to rm(list = c(…)) must be character strings. If a name does not exist, R can report an error. Check names with ls() first, or use rm(data, model) when the objects are known to exist.

    Clear data in R: what rm() removes and what it does not

    In R, a data frame, vector, list, function, or fitted model is an object. Removing it deletes its binding from the selected environment:

    rm(data)

    This does not delete an original CSV file, database table, or other external source used to create the object. It also does not remove individual columns from a data frame. To change the data frame itself, assign a revised object or remove a column with a separate operation.

    Removing a large object makes it eligible for memory cleanup, but it does not restart R or necessarily return memory to the operating system immediately. Use targeted removal when you want to keep analysis results, imported data, or model objects.

    Verify with ls() or objects(): what packages, plots, and session state remain?

    objects() is an alternative to ls() for listing objects. Use either command after removal:

    objects()
    character(0)

    Clearing objects does not detach loaded packages. Functions from attached packages remain available, and package namespaces are not unloaded. It also does not close graphics devices: a displayed plot remains on its device until you close it with a graphics command such as dev.off(). A plot stored as an object, however, is removed if its name is included in the rm() call.

    Working-directory settings, options, open connections, and the R session itself also remain. Clearing the global environment is therefore not a complete R session restart. Restarting R is a separate action that ends and starts the session again.

  • Standard Deviation in R: Calculate It Correctly

    Standard Deviation in R: Calculate It Correctly

    Use sd() to calculate standard deviation in R for a numeric vector or a data-frame column. For standard deviation in R, the main interpretation decision is whether your data represent a sample or an entire population. R’s default result is the sample standard deviation.

    The function also returns NA when missing values are present unless you explicitly remove them during the calculation.

    Standard deviation in R with sd()

    Pass a numeric vector to sd():

    x <- c(12, 15, 14, 10, 9)

    sd(x)

    This returns approximately 2.54951. The values have a mean of 12, and R divides the sum of squared deviations by 4, which is the sample denominator: the number of observations minus one.

    You can apply the same standard deviation function to a numeric column in a data frame:

    scores <- data.frame(score = c(12, 15, 14, 10, 9))

    sd(scores$score)

    This also returns approximately 2.54951. Use scores[[“score”]] as an alternative to scores$score. Both expressions select the underlying numeric vector. By contrast, scores[“score”] returns a one-column data frame, which is not the intended input for this calculation.

    The standard deviation function in R handles missing values

    By default, sd() does not ignore missing observations:

    measurements <- c(4, 7, NA, 10)

    sd(measurements)

    The result is NA because the missing value propagates through the calculation. Add na.rm = TRUE to exclude missing values:

    sd(measurements, na.rm = TRUE)

    This calculates the sample standard deviation of 4, 7, and 10. The argument removes only the NA values; it does not replace them or estimate what they might have been.

    Use the same argument with a data-frame column:

    sd(scores$score, na.rm = TRUE)

    If no nonmissing values remain, or only one nonmissing value remains, a standard deviation cannot be estimated and R returns NA.

    Sample versus population standard deviation

    sd() uses the sample standard deviation formula:

    sqrt(sum((x – mean(x))^2) / (n – 1))

    Here, n is the number of observed values. The n – 1 denominator estimates the variability of a larger population from a sample, so do not label R’s default result as a population standard deviation.

    Use a population standard deviation when your vector contains every member of the population being measured. The population formula divides by n:

    sqrt(sum((x – mean(x))^2) / n)

    In R, calculate it directly with:

    population <- c(12, 15, 14, 10, 9)

    sqrt(mean((population – mean(population))^2))

    If missing values are possible, remove them first:

    complete <- population[!is.na(population)]

    sqrt(mean((complete – mean(complete))^2))

    You can also convert a sample result to a population result with sd(complete) * sqrt((n – 1) / n), where n <- length(complete).

    Check R standard deviation with a worked vector

    This vector makes the denominator difference easy to verify:

    values <- c(2, 4, 4, 4, 5, 5, 7, 9)

    mean(values)

    sd(values)

    The mean is 5. The sum of squared deviations from that mean is 32, and there are eight observations. Therefore:

    • Sample standard deviation: sqrt(32 / 7), which is approximately 2.13809 and matches sd(values).
    • Population standard deviation: sqrt(32 / 8), which equals 2.

    Use sd() for the usual sample estimate, add na.rm = TRUE when missing values should be excluded, and use the denominator-n formula when the data represent the complete population.