Ravindra BagaleCourses & study guides

12. Exploratory Data Analysis (EDA)

12.2 First Pass Checklist

EDA steps in Jupyter

  1. df.shape, df.head(), df.tail().
  2. df.info() and df.isna().sum().
  3. df["order_id"].duplicated().sum().
  4. df.describe(include="all") — or split numeric / object.
  5. df["city"].value_counts() and normalize=True.
  6. df["status"].value_counts().
  7. pd.crosstab(df["city"], df["status"], margins=True).
  8. df.groupby("festival")["amount"].agg(["sum", "mean", "count"]).
  9. Write 3 bullet findings in a markdown cell.

What you should see. Pune appears most; Diwali rows show higher amounts in this fictional sample; one Cancelled (Kolhapur) and one Returned (Pune).