# 12.1 प्रश्नांपासून EDA सुरू करा

Source: https://ravindrabagale.com/mr/datascience/ch12-exploratory-data-analysis-eda/12-1-question-driven-eda.html
Language: mr (Marathi with English technical terms)

प्रत्येक column चा chart बनवत बसू नका. आधी हे प्रश्न विचारा:

किती rows, columns, nulls आणि duplicate keys आहेत?

प्रत्येक row चा grain काय आहे?

कोणत्या cities किंवा statuses जास्त दिसतात?

Festival tags असलेल्या rows वेगळ्या दिसतात का?

Minutes किंवा amounts मध्ये अशक्य values आहेत का?

पुढच्या steps साठी हा काल्पनिक sample तयार करा:

import pandas as pd

df = pd.DataFrame({
    "order_id": [f"BLK-{i}" for i in range(1, 13)],
    "city": [
        "Pune", "Pune", "Nashik", "Nagpur", "Solapur", "Kolhapur",
        "Sambhaji Nagar", "Pune", "Nashik", "Pune", "Nagpur", "Pune",
    ],
    "amount": [64, 900, 450, 240, 1299, 58, 195, 320, 880, 110, 240, 75],
    "status": [
        "Delivered", "Delivered", "Delivered", "Delivered", "Delivered", "Cancelled",
        "Delivered", "Delivered", "Delivered", "Returned", "Delivered", "Delivered",
    ],
    "festival": [
        "None", "Diwali", "Diwali", "None", "None", "None",
        "Diwali", "Ganeshotsav", "Diwali", "None", "None", "None",
    ],
    "delivery_mins": [9, 18, 16, 11, 18, 14, 10, 12, 17, 13, 11, 8],
    "customer": [
        "Ruhi Bagale", "Amir", "Salman", "Zoya", "Raja", "Rani",
        "Shahrukh", "Ravina", "Shraddha Bagale", "Ruhi Bagale", "Amir", "Salman",
    ],
})

रवींद्र बागले यांची tip

EDA म्हणजे data मध्ये काय मिळेल ते बघत राहणं एवढंच नाही. प्रत्येक check कोणत्या प्रश्नाचं उत्तर देतो ते स्पष्ट ठेवा.
