Data Science · मराठी आवृत्ती
9.4 Outliers: वेगळ्या दिसणाऱ्या values तपासा
या काल्पनिक lab मध्ये 120 पेक्षा जास्त delivery minutes sensor error किंवा typo मानले आहेत. त्या values NaN करून उरलेल्या minutes च्या median ने भरतो:
# fictional: delivery mins > 120 treated as bad sensor / typo
df.loc[df["mins"] > 120, "mins"] = np.nan
df["mins"] = df["mins"].fillna(df["mins"].median())
Cleaning checklist
- Column names एकसारखे करा.
- City आणि status चा text strip/title-case करा.
- Amounts साठी to_numeric; invalid text NaN मध्ये बदला.
- Missing city भरायची की row काढायची ते ठरवा.
- योग्य business key वर duplicate काढा — या छोट्या sample मध्ये order_id.
- अशक्य minutes/amounts तपासून ठरलेल्या नियमानुसार cap किंवा null करा.
df.to_csv("data/orders_clean.csv", index=False)run करा.- पुन्हा
info()आणि आधी/नंतरचे row counts तपासा.
BLK-2 ची एक row, numeric amount 1299, Mins 999 च्या जागी median आणि missing city च्या जागी Unknown दिसेल. वेगळा नियम वापरला तर तो नोंदवा.
Practice
फक्त course मधल्या महाराष्ट्रातील cities ठेवा. Unknown भरला असेल तर त्या rows काढा. orders_clean.csv export करा आणि df["city"].value_counts() print करा.
Recap
Rename → text standardise → numeric conversion → missing values चा नियम → duplicates → outliers → export. Rows काढल्यावर count ची नोंद ठेवा. आता dates, IST आणि text पाहू.