# 12.2 पहिल्या तपासणीची checklist

Source: https://ravindrabagale.com/mr/datascience/ch12-exploratory-data-analysis-eda/12-2-first-pass-checklist.html
Language: mr (Marathi with English technical terms)

Jupyter मध्ये EDA steps

df.shape, df.head(), df.tail().

df.info() आणि df.isna().sum().

df["order_id"].duplicated().sum().

df.describe(include="all"); किंवा numeric/text वेगळं तपासा.

df["city"].value_counts(); वाटा पाहण्यासाठी normalize=True वापरा.

df["status"].value_counts().

pd.crosstab(df["city"], df["status"], margins=True).

df.groupby("festival")["amount"].agg(["sum", "mean", "count"]).

Markdown cell मध्ये तीन findings लिहा.

या काल्पनिक sample मध्ये Pune सर्वाधिक वेळा दिसतं. Diwali rows च्या amounts जास्त आहेत. Kolhapur मध्ये एक Cancelled आणि Pune मध्ये एक Returned row आहे.

रवींद्र बागले यांची tip

Count आणि percentage दोन्ही पाहा. एखाद्या city मध्ये rows कमी असतील तर एक cancellation मुळे percentage मोठी दिसू शकते.
