Text and NLP Feature Engineering
Raw text — a sentence, a review, a support ticket — isn't something a model can use directly. Text feature engineering is the process of turning words into numeric representations a model can actually learn from, ranging from very simple counting methods to modern learned embeddings.
Cleaning and preprocessing, before feature extraction:
- Lowercasing, removing punctuation/special characters — reduces unnecessary variation ("Dog" and "dog" shouldn't usually be treated as different tokens).
- Tokenization — splitting text into words or sub-word units.
- Stopword removal — dropping very common, low-information words ("the," "is," "and"), though this is worth skipping for tasks where those words carry meaning, like some sentiment or authorship tasks.
- Stemming / lemmatization — reducing words to a common root or dictionary form ("running" → "run"), so the model doesn't have to treat every inflected form as a separate word.
From text to numbers, roughly from simplest to most powerful:
- Bag of Words — count how many times each word appears in a document. Simple, interpretable, and a reasonable baseline, but ignores word order and context entirely, and treats every word as equally distinct.
- TF-IDF (Term Frequency–Inverse Document Frequency) — like Bag of Words, but down-weights words that appear in almost every document (less informative) and up-weights words that are distinctive to a particular document. A very common, strong, and cheap baseline for text classification.
- N-grams — count sequences of 2 or 3 consecutive words instead of (or alongside) single words, capturing a bit of local word order ("not good" carries different meaning than "not" and "good" separately).
- Word embeddings (word2vec, GloVe) — represent each word as a dense vector learned from large text corpora, such that words with similar meaning end up close together in the vector space. A big step up from Bag of Words/TF-IDF for capturing meaning, though a given word still gets one fixed vector regardless of context.
- Contextual embeddings (BERT and similar transformer models) — represent a word's meaning differently depending on its surrounding sentence, which handles cases like "bank" (river vs. finance) that fixed embeddings can't distinguish. State of the art for most modern NLP tasks, at the cost of more compute.
Which should you use? For a small student project or a quick baseline, TF-IDF with a simple classifier (logistic regression, linear SVM) is fast, interpretable, and often surprisingly competitive. Reach for pretrained embeddings or transformer models once you need to capture meaning and context more precisely, and you have the compute budget (or can use a pretrained model off the shelf) to support it.
Pitfalls to avoid:
- Fit your TF-IDF vocabulary (or any text vectorizer) on the training set only, then apply it to the test set — the same leakage principle as everywhere else in this section applies here too.
- Don't over-clean. Aggressive preprocessing (like removing all punctuation) can strip out meaningful signal for some tasks — e.g. exclamation marks and capitalization often matter for sentiment analysis.
- Domain-specific text (medical notes, legal documents, code) often benefits from a domain-tuned tokenizer or embedding model rather than a general-purpose one trained on generic web text.
Where to go deeper: scikit-learn's text feature extraction documentation covers Bag of Words and TF-IDF with working examples; the Hugging Face Transformers documentation is the standard starting point for pretrained contextual embeddings like BERT.