To filter in R, use dplyr::filter() to keep data-frame rows that satisfy one or more conditions. It handles numeric comparisons, text matches, missing values, and combinations of conditions in a readable way.
To filter data in R, remember that filtering changes which rows remain; it does not choose which columns are displayed. The R filter workflow below uses a small data frame and then shows the equivalent base R approach.
How do you filter in R with dplyr::filter(), not stats::filter()?
Start with a data frame containing numeric, text, and missing values:
sales <- data.frame(product = c(“A”, “B”, “A”, “C”), region = c(“East”, “West”, NA, “East”), units = c(12, 7, NA, 20))
Use dplyr::filter() with a comparison such as greater than or equal to:
large_sales <- dplyr::filter(sales, units >= 10)
This keeps rows where units is at least 10. Common comparison operators are == for equal to, != for not equal to, >, <, >=, and <=. The function returns all columns for the matching rows. In contrast, dplyr::select(sales, product, units) selects columns rather than filtering rows.
Use dplyr::filter() for row operations. The separate stats::filter() function is designed for time-series and other filtering operations on vectors, so it is not the row-filtering function used here.
How do you combine conditions with AND, OR, and negation?
Use & for AND when every condition must be true:
east_large <- dplyr::filter(sales, region == “East” & units >= 10)
Use | for OR when either condition can be true:
east_or_west <- dplyr::filter(sales, region == “East” | region == “West”)
Use ! for negation. For example, this keeps rows whose region is not West:
not_west <- dplyr::filter(sales, !(region == “West”))
Parentheses make compound logic easier to read and prevent ambiguity. Use & and | for row-by-row conditions; && and || are scalar operators and are generally inappropriate for filtering a full column.
How do you filter text and missing values safely?
Match text by comparing a character column with a quoted value:
east_sales <- dplyr::filter(sales, region == “East”)
R represents a missing value as NA. A comparison such as region == “East” produces NA when region is missing, not TRUE or FALSE. dplyr::filter() keeps only rows where the condition is TRUE, so those uncertain rows are excluded.
Test missing values explicitly with is.na() or its negation:
missing_region <- dplyr::filter(sales, is.na(region))
known_region <- dplyr::filter(sales, !is.na(region))
Combine the missing-value check with another condition when needed:
known_large <- dplyr::filter(sales, !is.na(units) & units >= 10)
How does an R filter work with base R logical indexing?
Base R filters rows by placing a logical condition before the comma inside square brackets. Include an explicit missing-value check so an NA does not create an unintended missing row in the result:
large_sales_base <- sales[!is.na(sales$units) & sales$units >= 10, , drop = FALSE]
The expression before the comma chooses rows; the blank expression after the comma keeps all columns. A text condition with AND works the same way:
east_base <- sales[!is.na(sales$region) & sales$region == “East”, , drop = FALSE]
For OR, use parentheses around the alternatives:
east_or_west_base <- sales[!is.na(sales$region) & (sales$region == “East” | sales$region == “West”), , drop = FALSE]
To select columns instead of rows, place column names after the comma: sales[, c(“product”, “units”), drop = FALSE]. That distinction keeps base R logical indexing focused on rows while column selection remains a separate operation.
