9. Data Cleaning A–Z in pandas
9.4 Outliers
# fictional: delivery mins > 120 treated as bad sensor / typo
df.loc[df["mins"] > 120, "mins"] = np.nan
df["mins"] = df["mins"].fillna(df["mins"].median())
Cleaning checklist
- Normalise column names.
- Strip / title-case categories (city, status).
to_numericfor amounts; coerce junk to NaN.- Decide fill vs drop for missing cities.
drop_duplicateson business key (order_id).- Cap or null out impossible mins / amounts.
df.to_csv("data/orders_clean.csv", index=False).- Re-run
info()and compare row counts before/after.
What you should see. One row for BLK-2; Amount 1299 as number; Mins 999 becomes median; City null → Unknown or dropped per your rule — document the rule.
Ravindra Bagale's Tip
Khup students dropna() full DataFrame var sodtat aani arda data gamatat. Column-wise vichara: city missing ≠ amount missing. subset= vapra. Bilkul visru naka.
Practice task
Clean the sample so only real Maharashtra course cities remain (drop Unknown if you filled it). Export orders_clean.csv and print df["city"].value_counts().
Thodkyaat sangaycha tar (quick recap)
- Rename → standardise text → numeric coerce → missing strategy → duplicates → outliers → export
- Never silent-drop without a count log
Samajla ka? Aata pudhe jaauya dates, times (IST) and text.