Ravindra BagaleCourses & study guides

9. Data Cleaning A–Z in pandas

9.4 Outliers

# fictional: delivery mins > 120 treated as bad sensor / typo
df.loc[df["mins"] > 120, "mins"] = np.nan
df["mins"] = df["mins"].fillna(df["mins"].median())

Cleaning checklist

  1. Normalise column names.
  2. Strip / title-case categories (city, status).
  3. to_numeric for amounts; coerce junk to NaN.
  4. Decide fill vs drop for missing cities.
  5. drop_duplicates on business key (order_id).
  6. Cap or null out impossible mins / amounts.
  7. df.to_csv("data/orders_clean.csv", index=False).
  8. Re-run info() and compare row counts before/after.

What you should see. One row for BLK-2; Amount 1299 as number; Mins 999 becomes median; City null → Unknown or dropped per your rule — document the rule.

Ravindra Bagale's Tip

Khup students dropna() full DataFrame var sodtat aani arda data gamatat. Column-wise vichara: city missing ≠ amount missing. subset= vapra. Bilkul visru naka.

Practice task

Clean the sample so only real Maharashtra course cities remain (drop Unknown if you filled it). Export orders_clean.csv and print df["city"].value_counts().

Thodkyaat sangaycha tar (quick recap)

  • Rename → standardise text → numeric coerce → missing strategy → duplicates → outliers → export
  • Never silent-drop without a count log

Samajla ka? Aata pudhe jaauya dates, times (IST) and text.