Encoding Categorical Variables
Most machine learning algorithms only understand numbers, but a huge amount of real data comes as categories — colours, cities, product types, yes/no answers. Encoding is how you turn those categories into numbers without accidentally telling the model something untrue, like implying "blue" is somehow greater than "red."
The main techniques:
- One-hot encoding — create a new binary column for each category. The safe default for categories with no natural order (nominal data) and a small-to-moderate number of unique values. Downside: it can create a huge number of columns when a feature has thousands of unique categories (high cardinality), like zip codes or user IDs.
- Label / ordinal encoding — map each category to a single integer. Appropriate when the categories genuinely have an order (e.g. "low," "medium," "high"). Using it on unordered categories is a common mistake, since it invents a fake ranking the model will try to learn from.
- Target encoding — replace each category with a statistic of the target variable for that category (e.g. the average outcome for that group). Keeps the column count low even for high-cardinality features, but is prone to overfitting and target leakage if not done carefully (usually mitigated with cross-validation or smoothing).
- Frequency / count encoding — replace each category with how often it appears in the data. Simple, avoids target leakage entirely, and can still carry useful signal (rare categories often behave differently from common ones).
- Hashing / binary encoding — more advanced options for very high-cardinality features, trading off a small amount of information loss (hash collisions) for a fixed, manageable number of columns.
Which one should you use? As a rough default: one-hot for low-cardinality nominal features, ordinal encoding only when a real order exists, and target or frequency encoding once one-hot would create an unreasonable number of columns. Tree-based models tend to handle label-encoded or target-encoded high-cardinality features gracefully; linear models and neural networks are more sensitive to fake ordinal relationships, so lean towards one-hot or target encoding for those.
Pitfalls to avoid:
- If you're using target encoding, compute the target statistics using only the training set (ideally with cross-validation folds), never the full dataset — otherwise you're leaking the label into your features. See the Data Leakage resource.
- Decide up front how to handle a category at inference time that never appeared in training (an "unknown" category). Most encoders let you configure a fallback value; forgetting to handle this is a common cause of production errors.
- Watch out for high cardinality quietly bloating your dataset with one-hot encoding — check the number of unique values in a categorical column before blindly one-hot encoding it.
Where to go deeper: the category_encoders library implements most of the techniques above (including target, frequency, hashing, and binary encoding) with a consistent scikit-learn-style API, which is a good way to experiment with several approaches quickly.