12. Exploratory Data Analysis (EDA)
12.1 Question-Driven EDA
Do not plot everything. Start with questions:
- How many rows, columns, nulls, duplicate keys?
- What is the grain?
- Which cities / statuses dominate?
- Are festival days different?
- Any impossible values (mins, amounts)?
import pandas as pd
df = pd.DataFrame({
"order_id": [f"BLK-{i}" for i in range(1, 13)],
"city": [
"Pune", "Pune", "Nashik", "Nagpur", "Solapur", "Kolhapur",
"Sambhaji Nagar", "Pune", "Nashik", "Pune", "Nagpur", "Pune",
],
"amount": [64, 900, 450, 240, 1299, 58, 195, 320, 880, 110, 240, 75],
"status": [
"Delivered", "Delivered", "Delivered", "Delivered", "Delivered", "Cancelled",
"Delivered", "Delivered", "Delivered", "Returned", "Delivered", "Delivered",
],
"festival": [
"None", "Diwali", "Diwali", "None", "None", "None",
"Diwali", "Ganeshotsav", "Diwali", "None", "None", "None",
],
"delivery_mins": [9, 18, 16, 11, 18, 14, 10, 12, 17, 13, 11, 8],
"customer": [
"Ruhi Bagale", "Amir", "Salman", "Zoya", "Raja", "Rani",
"Shahrukh", "Ravina", "Shraddha Bagale", "Ruhi Bagale", "Amir", "Salman",
],
})