← Back to Resources

Encoding Categorical Variables

Practical Resources - AI Engineering

Most machine learning algorithms only understand numbers, but a huge amount of real data comes as categories — colours, cities, product types, yes/no answers. Encoding is how you turn those categories into numbers without accidentally telling the model something untrue, like implying "blue" is somehow greater than "red."


The main techniques:

Which one should you use? As a rough default: one-hot for low-cardinality nominal features, ordinal encoding only when a real order exists, and target or frequency encoding once one-hot would create an unreasonable number of columns. Tree-based models tend to handle label-encoded or target-encoded high-cardinality features gracefully; linear models and neural networks are more sensitive to fake ordinal relationships, so lean towards one-hot or target encoding for those.

Pitfalls to avoid:

Where to go deeper: the category_encoders library implements most of the techniques above (including target, frequency, hashing, and binary encoding) with a consistent scikit-learn-style API, which is a good way to experiment with several approaches quickly.