← Back to Resources

Tips on Data Augmentation

Practical Resources - AI Engineering

Data augmentation is the practice of artificially expanding your training set by creating modified copies of existing data, rather than collecting entirely new samples. It's one of the cheapest ways to improve a model when you don't have enough labelled data to begin with, or when your model is overfitting a small dataset.


Why it matters: models generalise better when they've seen more variation. If your dataset only shows a cat facing forward in bright light, your model may struggle with a cat facing sideways in dim light — even though it's obviously still a cat. Augmentation exposes the model to that kind of variation without the cost of collecting and labelling thousands more real examples.

Common techniques by data type:

Things to watch out for:

Where to go deeper: most modern deep learning frameworks have augmentation built in and ready to use — for images, look at torchvision.transforms (PyTorch), tf.keras.layers augmentation layers (TensorFlow/Keras), or the standalone Albumentations library, which is fast and has a huge range of transforms. For tabular data, the imbalanced-learn library implements SMOTE and its variants.