7. Data Cleaning A–Z in Power Query
7.1 A Cleaning Workflow You Can Repeat
Clean in the same order every time. This avoids surprises. Udaharan mhanje, if you remove duplicates before you trim spaces, then "Pune" and "Pune·" are treated as two different values.
| # | Stage | What you do | Section |
|---|---|---|---|
| 1 | Profile | Turn on the column quality, distribution and profile tools; profile the entire dataset | 7.2 |
| 2 | Shape | Remove title rows, promote headers, remove footer/total rows, remove unwanted columns | 7.3 |
| 3 | Text hygiene | Trim, Clean, fix case, standardise spellings | 7.6, 7.9 |
| 4 | Types | Set data types (use Using Locale for Indian dates); convert ₹ text to numbers | 7.7, 7.8 |
| 5 | Blanks & errors | Replace, fill or remove nulls; handle errors | 7.5, 7.15 |
| 6 | Duplicates | Remove full-row duplicates or duplicates by key | 7.4 |
| 7 | Validity | Flag or filter invalid values and outliers | 7.13, 7.16 |
| 8 | Restructure | Split, merge, extract, unpivot, pivot | 7.10–7.12, 7.17 |
| 9 | Combine | Append, merge, folder combine, group | 7.20–7.23 |
| 10 | Document | Rename steps, check query dependencies, turn off load for helper queries | Module 6 |
Ravindra Bagale's Tip
Mitrano, sagalyat jast disnari chuk mhanje cleaning in a random order, udaharan mhanje changing types before removing ₹ symbols or before promoting headers. Follow the same order every time: remove junk rows → promote headers → trim and clean text → fix values → set types → remove duplicates. A fixed order prevents most conversion errors. Samjla ka?