Baseline Models: Why You Should Always Build One First
Before building anything sophisticated, build the simplest possible model that could plausibly work. This feels like a step you can skip to save time; in practice, skipping it usually costs you more time than it saves.
What counts as a baseline:
- For classification: predict the majority class every time, or a random guess weighted by class frequency.
- For regression: predict the mean (or median) of the target every time.
- A simple rule-based heuristic, if you have domain knowledge — e.g. "flag a transaction as fraud if it's over $1,000 and from a new device."
- A simple, well-understood model with default settings and minimal feature engineering — logistic/linear regression is a common choice, since it trains in seconds and gives you a real number to compare against.
Why is this important?
- It tells you what "good" actually means for your problem. 85% accuracy sounds impressive until you learn the majority-class baseline already gets 90%. Without a baseline, you have no way to know if your fancy model is actually adding value.
- It catches pipeline bugs early. A baseline model runs end-to-end through your whole pipeline — data loading, splitting, evaluation — and that plumbing needs to work correctly before a more complex model's results mean anything. Better to find a broken evaluation script when your baseline gives a nonsensical number than after a week spent tuning a neural network on top of the same broken pipeline.
- It sets realistic expectations. If even a strong classical model can't do much better than the baseline, that's useful information about the problem itself — maybe the features don't carry much signal, or the task is genuinely close to its ceiling of predictability.
- It gives you an early, working system. Especially relevant if you're following the "build a minimal working version first" approach described in the Managing Overwhelming Projects resource — a baseline model, however crude, is a real end-to-end deliverable you can build on and demo early.
How to use it going forward: every time you try a new model, feature, or technique, compare it explicitly against the baseline (and against your previous best result), not in isolation. A model that improves on the baseline by 1% might not be worth the added complexity and training cost it introduces — that's a judgement call the baseline lets you actually make, instead of guessing.
Where to go deeper: scikit-learn's DummyClassifier and DummyRegressor implement common baseline strategies (majority class, stratified random, mean/median prediction) with one line of code, making it easy to set a baseline before writing anything more complex.