Dimensionality Reduction (PCA and friends)
When you have a large number of features, models can become slow to train, harder to visualise, and prone to the "curse of dimensionality" (in high-dimensional spaces, data points become sparse and distance-based methods struggle to find meaningful patterns). Dimensionality reduction techniques compress your features into a smaller set that still captures most of the useful information.
Principal Component Analysis (PCA) is the classic starting point. It finds new axes (principal components) that are combinations of your original features, ordered so that the first component captures the most variance in the data, the second captures the next most (while being uncorrelated with the first), and so on. Keeping only the top few components lets you represent most of the dataset's variation in far fewer dimensions.
What PCA is good for:
- Speeding up training and reducing memory use when you have hundreds or thousands of features.
- Removing redundancy between highly correlated features.
- Visualising high-dimensional data by reducing it to 2 or 3 components you can actually plot.
- Reducing noise, since lower-variance components often capture more noise than signal.
What it's not good for: PCA components are linear combinations of the original features and are often hard to interpret directly — "component 3" doesn't map cleanly onto a real-world concept the way an original feature like "age" does. This makes PCA a poor fit when interpretability matters (e.g. explaining a decision to a stakeholder or regulator). It's also a linear technique, so it can miss more complex, non-linear structure in the data — for that, non-linear alternatives like t-SNE or UMAP are commonly used, particularly for visualisation.
Pitfalls to avoid:
- Scale your features first (see Feature Scaling and Normalization) — PCA is sensitive to the units of your features, and an unscaled feature with a huge numeric range can dominate the components for no meaningful reason.
- Fit PCA on the training data only, then apply the same transformation to the test data — fitting it on everything first is another route to data leakage.
- Check how much variance is actually explained by the components you keep (most libraries report this directly). Reducing to 2 dimensions for a plot is fine for visualisation, but throwing away too much variance before modelling can quietly hurt performance.
- t-SNE and UMAP are excellent for visualisation but are generally not meant to be used as preprocessing input to a downstream model in the same way PCA is — their distances and axes aren't as directly meaningful for that purpose.
Where to go deeper: scikit-learn's decomposition module documentation covers PCA with a worked example, and the UMAP documentation is a good next step for non-linear dimensionality reduction, especially for visualisation.