Feature Selection Techniques
More features aren't automatically better. Irrelevant or redundant features can slow training down, make models harder to interpret, and in some cases actively hurt performance by giving the model more opportunities to fit noise instead of signal. Feature selection is the process of deciding which features actually earn their place in the model.
Three broad approaches:
- Filter methods — score each feature independently of any model, using statistics like correlation with the target, chi-squared tests, or mutual information, and keep the top-scoring ones. Fast and simple, but ignores interactions between features (a feature that's useless alone but powerful in combination with another can get filtered out by mistake).
- Wrapper methods — actually train models on different subsets of features and see which subset performs best, using strategies like recursive feature elimination (repeatedly drop the least useful feature and retrain) or forward/backward selection. More accurate than filter methods since they account for interactions, but far more computationally expensive.
- Embedded methods — feature selection happens as part of training the model itself. Lasso (L1) regression naturally shrinks unhelpful feature coefficients to exactly zero; tree-based models like random forests and gradient boosting produce feature importance scores as a natural byproduct of training. Often the best balance of accuracy and cost.
Why is this important? Beyond the performance argument, fewer, well-chosen features usually mean a model that's faster to train, cheaper to run in production, and much easier to explain to a stakeholder or debug when something goes wrong. A model with 200 barely-relevant features is a much harder thing to reason about than one with 15 features you can each justify.
Pitfalls to avoid:
- Do feature selection inside your cross-validation loop, not once on the full dataset beforehand — selecting features using information from your test set is a subtle but common form of data leakage, and will make your evaluation metrics look better than they really are.
- A feature being weakly correlated with the target on its own doesn't mean it's useless — some features are only informative in combination with others. Filter methods alone can miss this.
- Feature importance from tree-based models can be biased towards high-cardinality features; don't treat it as an absolute ranking without a bit of skepticism.
Where to go deeper: scikit-learn's feature selection documentation implements filter methods, recursive feature elimination, and model-based selection with ready-to-use code.