← Back to Resources

Dimensionality Reduction (PCA and friends)

Practical Resources - AI Engineering

When you have a large number of features, models can become slow to train, harder to visualise, and prone to the "curse of dimensionality" (in high-dimensional spaces, data points become sparse and distance-based methods struggle to find meaningful patterns). Dimensionality reduction techniques compress your features into a smaller set that still captures most of the useful information.


Principal Component Analysis (PCA) is the classic starting point. It finds new axes (principal components) that are combinations of your original features, ordered so that the first component captures the most variance in the data, the second captures the next most (while being uncorrelated with the first), and so on. Keeping only the top few components lets you represent most of the dataset's variation in far fewer dimensions.

What PCA is good for:

What it's not good for: PCA components are linear combinations of the original features and are often hard to interpret directly — "component 3" doesn't map cleanly onto a real-world concept the way an original feature like "age" does. This makes PCA a poor fit when interpretability matters (e.g. explaining a decision to a stakeholder or regulator). It's also a linear technique, so it can miss more complex, non-linear structure in the data — for that, non-linear alternatives like t-SNE or UMAP are commonly used, particularly for visualisation.

Pitfalls to avoid:

Where to go deeper: scikit-learn's decomposition module documentation covers PCA with a worked example, and the UMAP documentation is a good next step for non-linear dimensionality reduction, especially for visualisation.