Exploratory Data Analysis (EDA)
Before you engineer a single feature or train a single model, you should actually look at your data. Exploratory Data Analysis (EDA) is the habit of examining a dataset — its shape, its distributions, its oddities — before you start building anything on top of it. Skipping this step is one of the most common ways teams waste days debugging a model that was never the actual problem; the data was.
A reasonable EDA checklist:
- Shape and types. How many rows and columns? What type is each column — numeric, categorical, date, text? Does that match what you'd expect?
- Missingness. Which columns have missing values, and how much? See the Handling Missing Values resource for what to do once you know.
- Distributions. Plot histograms for numeric columns and bar charts for categorical ones. Is anything heavily skewed? Are there unexpected spikes (a suspicious number of exact zeros, or a value like 9999 that's probably a placeholder for missing data)?
- Outliers. Box plots and scatter plots will often surface them immediately — see Handling Outliers.
- Relationships between features. A correlation matrix or pairwise scatter plots for numeric features; group-by summaries for categorical ones against your target. This is often where you first spot candidates for feature selection or engineering.
- Target variable. What does its distribution look like? A heavily imbalanced classification target changes your whole approach — see the Handling Class Imbalance resource.
- Duplicates and inconsistencies. Repeated rows, inconsistent category spellings ("NY" vs "New York" vs "new york"), and mismatched units are extremely common and easy to miss without deliberately checking for them.
Why is this important? Every downstream step — which features to engineer, which encoding or scaling to use, which model family makes sense — is a decision you're making blind if you skip this. A five-minute look at a histogram can save hours spent debugging a model that was actually just being fed garbage.
Tips:
- Do EDA on the training set only, or a held-out exploration sample — not the full dataset including your test set, for the same reasons covered in Data Leakage.
- Write down what you notice as you go, even briefly. It's easy to spot something odd, move on, and forget it by the time it matters.
- Don't treat EDA as a one-off step in week one. It's worth revisiting whenever you add new data or notice the model behaving strangely.
Where to go deeper: the Python libraries pandas-profiling (now ydata-profiling) and Sweetviz can auto-generate a full EDA report (distributions, correlations, missingness) from a single line of code, which is a great way to get an initial overview before diving in manually.