Category: Python Data

  • df.rename: How to rename a pandas column in a DataFrame

    df.rename: How to rename a pandas column in a DataFrame

    Use df.rename() to rename a pandas column, several columns, or index labels without rebuilding the DataFrame. The method accepts explicit old-to-new mappings for targeted edits and callable functions for systematic cleanup.

    By default, rename() returns a new DataFrame and leaves the original unchanged. That behavior makes it suitable for a clear, assignable pandas DataFrame’s rename method.

    How to rename a DataFrame column or several columns with df.rename()

    Use columns={old: new} when you rename a pandas DataFrame

    Pass a dictionary to columns. Each key is the existing column label, and each value is its replacement. The mapping always goes from old name to new name:

    renamed = df.rename(columns={“Name”: “name”})

    This creates a new DataFrame with Name changed to name. The original df still has its original labels. To rename several columns, include multiple entries:

    renamed = df.rename(columns={
        “First Name”: “first_name”,
        “Order Total”: “order_total”
    })

    Columns not listed in the dictionary remain unchanged. This is the direct way to rename a DataFrame column when you know the exact labels. Do not reverse the dictionary: {“name”: “Name”} searches for a column already named name and changes it to Name.

    Should you return a new DataFrame or change a DataFrame column name in place?

    Return a new DataFrame by default and assign it

    The default operation returns a renamed copy-like DataFrame object:

    df2 = df.rename(columns={“old_name”: “new_name”})

    Use assignment when you want to preserve the original DataFrame, keep a clearly named transformed object, or chain the result with other operations. You can also replace the original variable:

    df = df.rename(columns={“old_name”: “new_name”})

    Use inplace=True when you do not need a returned object

    Set inplace=True to update df directly:

    df.rename(columns={“old_name”: “new_name”}, inplace=True)

    This operation returns None, so do not assign its result back to df. Choose the default returned-DataFrame behavior when you want easier debugging, reversible steps, or a transformation pipeline. Choose in-place behavior when direct mutation is intentional and you do not need the method’s return value.

    Choose errors=’raise’ or ‘ignore’ for missing labels

    By default, errors=”ignore” leaves a mapping unchanged if its old label is absent. To catch misspelled or unexpected labels, use:

    df.rename(columns={“Oder Total”: “order_total”}, errors=”raise”)

    A missing source label then raises a KeyError. Use errors=”ignore” when a mapping may apply only to some DataFrames, and errors=”raise” when every requested rename must succeed.

    How to transform every column label with a callable

    Use a callable to clean labels and display the final columns

    Pass a function to columns when every label needs the same transformation. The callable receives one label at a time and must return its replacement. This example strips surrounding whitespace, converts labels to lowercase, and replaces spaces with underscores:

    clean = df.rename(
        columns=lambda column: column.strip().lower().replace(” “, “_”)
    )

    For input labels Customer ID, Order Total, and Order Date, display the result with:

    print(clean.columns.tolist())

    [‘customer_id’, ‘order_total’, ‘order_date’]

    A callable is useful for consistent label cleanup, while a dictionary is safer when only selected columns should change. If labels are not all strings, account for their types before calling string methods.

    How to rename index labels or a MultiIndex level

    Rename ordinary index labels

    Use index instead of columns to rename row labels:

    df2 = df.rename(index={0: “first”, 1: “second”})

    This changes labels, not row positions or the values stored in the DataFrame.

    Rename one level of a MultiIndex

    For hierarchical columns or rows, supply the level name or number. To rename labels in the second level of a MultiIndex column:

    df2 = df.rename(
        columns={“Q1”: “first_quarter”},
        level=1
    )

    Use index= for a MultiIndex on rows. The level argument limits the mapping to one hierarchy level, leaving labels in the other levels unchanged.

  • Append to pandas DataFrame with pd.concat

    Append to pandas DataFrame with pd.concat

    Use pd.concat to append to pandas DataFrame objects in current pandas versions. It replaces the removed df.append pattern and can combine DataFrames, dictionaries converted to rows, or batches of records. For most row additions, use ignore_index=True so the result receives one continuous index.

    pd.concat returns a new DataFrame; it does not modify either input in place. The basic pattern is pd.concat([df, new_data], ignore_index=True).

    Replace df.append with pd.concat: side-by-side legacy and current code with ignore_index

    The pandas DataFrame append method was removed in pandas 2.0. If older code contains an append call, replace it with a list passed to pd.concat.

    1. Legacy code, no longer available: df.append(new_df, ignore_index=True)
    2. Current code: pd.concat([df, new_df], ignore_index=True)

    The replacement is usually a direct change. Both operations place the rows from new_df below the rows in df. Assign the result if you need the combined object:

    result = pd.concat([df, new_df], ignore_index=True)

    Use ignore_index=True when the original index labels are not meaningful and the output should run from 0 through the number of rows minus 1. Without it, pandas preserves the existing labels. That can produce duplicate index values, especially when both inputs use the default range index.

    For example, concatenating two three-row DataFrames without ignore_index can create an index of 0, 1, 2, 0, 1, 2. This is valid, but label-based selection and joins may become ambiguous. Use result.reset_index(drop=True) later if resetting the index is more convenient.

    Append DataFrame objects with pd.concat: column alignment, missing values, and duplicate indexes

    To append DataFrame objects, pass them in order inside a list. By default, pd.concat combines rows with axis=0 and aligns columns by column name:

    result = pd.concat([sales_january, sales_february], ignore_index=True)

    Column order in the result generally follows the first DataFrame, followed by columns introduced by later DataFrames. If the inputs do not contain the same columns, pandas uses the union of their column names. A missing value appears as NaN or another appropriate missing-value marker.

    • If the first DataFrame has customer and total, while the second has customer and currency, the result contains all three columns.
    • January rows have missing values in currency.
    • February rows have missing values in total.

    This name-based alignment prevents values from being assigned to the wrong field when column order differs. If you want only columns shared by every input, use join=”inner”:

    result = pd.concat([df_a, df_b], join=”inner”, ignore_index=True)

    Use the default join=”outer” when preserving every column matters. Check the resulting dtypes after concatenation if one input contains numbers and another contains strings or missing values; pandas may broaden a column’s dtype to accommodate both.

    Duplicate index labels are retained unless you explicitly request a new index. They are not automatically errors. Preserve them when the labels identify source records, or set ignore_index=True when the combined DataFrame represents a new sequential collection.

    Add a dictionary or one row when appending to a pandas DataFrame

    A dictionary represents one row when each key is a column name and each value is that row’s value. Wrap it in a one-item list before creating a DataFrame:

    row = {“product”: “Notebook”, “quantity”: 3, “price”: 4.50}
    result = pd.concat([df, pd.DataFrame([row])], ignore_index=True)

    The list is important because pd.DataFrame([row]) creates one record. Calling pd.DataFrame(row) with scalar values does not provide enough information for pandas to determine the row index.

    Keys are matched to existing columns by name. If the dictionary omits a column, the new row receives a missing value there. If it introduces a new key, pd.concat adds a new column and fills that column with missing values for earlier rows:

    new_row = {“product”: “Pen”, “quantity”: 10}
    df = pd.concat([df, pd.DataFrame([new_row])], ignore_index=True)

    For one-row additions inside a controlled, interactive workflow, direct assignment such as df.loc[len(df)] = row can be concise. For a consistent append DataFrame workflow, converting the record and using pd.concat makes column alignment explicit.

    Build many rows efficiently with pandas concat for rows

    Do not repeatedly concatenate a growing DataFrame inside a loop. Each operation can allocate and copy the accumulated data, making a long sequence much slower than one final combination. The efficient pandas concat rows approach is to collect DataFrames or records first.

    Collect records as dictionaries, convert the complete batch once, and concatenate once:

    records = []
    for item in source:
        records.append({“id”: item.id, “status”: item.status})
    new_rows = pd.DataFrame(records)
    result = pd.concat([df, new_rows], ignore_index=True)

    This works well when each iteration produces a dictionary. If each iteration already produces a DataFrame, store those frames in a list instead:

    frames = [df]
    for batch in batches:
        frames.append(batch)
    result = pd.concat(frames, ignore_index=True)

    For an empty batch, decide whether to return the original DataFrame or concatenate it with an explicitly shaped empty DataFrame. When column types matter, define the expected columns and dtypes before processing so missing or empty inputs do not unexpectedly change the result.