Supervised versus unsupervised is usually taught as a one-line definition: one has labels, the other does not. That is true and nearly useless when you face a real dataset, because the useful question is not which category a method belongs to but what you are asking the data to tell you, what it costs to get the answer key, and how you will know whether the result is any good.
This article explains the split from first principles. It starts from what each family optimises, walks through the core algorithms on each side, spends a full section on the problem unsupervised learning never escapes (evaluation without an answer key), and then works through a realistic project that uses both: clustering fifty thousand unlabelled support tickets to discover categories, then training a classifier on a small labelled sample. It ends with the self-supervised methods that blur the line in modern deep learning, and a checklist for choosing on your own data.
The split is about the objective
Every learning method minimises some objective over data. The difference between the two families is what that objective refers to.
Supervised learning receives pairs of an input x and a target y and learns a function that predicts y from x. Probabilistically, it models the conditional distribution p(y | x). Classification predicts a category, regression a number, ranking an order. The objective is a loss that compares predictions with the given targets, such as cross-entropy or squared error, and the goal is that the loss stays low on new inputs drawn from the same distribution.
Unsupervised learning receives inputs only and learns something about their structure: the distribution p(x) itself, or a simpler description of it. Clustering finds groups of similar points. Dimensionality reduction finds a few directions that explain most of the variation. Density estimation and anomaly detection find which points are typical and which are rare. The objective is defined entirely by the inputs, for example the total distance from points to their cluster centres, or the reconstruction error after compressing and decompressing.
One consequence follows directly and shapes everything else: in supervised learning the answer you want is written into the objective, so a good loss on held-out data means a good model. In unsupervised learning the objective is a proxy for what you actually want, so a good objective value does not guarantee a useful result.
Supervised learning from first principles
A supervised model is a parameterised function f with parameters theta. Training picks the theta that minimises the average loss over the training pairs, usually by gradient descent or a closed-form solution for simple models. Logistic regression, gradient-boosted trees and neural networks differ in the shape of f and how they search, not in this basic contract.
The point of training is generalisation, not fitting. A model that memorises its training pairs scores perfectly on them and badly on anything new. That is why supervised projects hold out data the model never trains on and report loss or accuracy there, and why the way you split data matters as much as the model; see train, validation and test splits. With a clean held-out set, supervised learning has the most honest feedback loop in machine learning: you can measure exactly how often the model is right on the thing you care about.
The price is labels. Someone has to produce them, either people annotating examples or outcomes recorded later, such as whether a loan defaulted. Labels are slow, expensive, sometimes inconsistent between annotators, and specific to one task: labels for spam detection do nothing for topic routing. They also encode a fixed taxonomy decided before training, which is a problem when you do not yet know what the categories should be.
Unsupervised learning from first principles
k-means clustering is the clearest example. Choose a number of clusters k. The objective, called inertia, is the sum of squared distances from each point to the centre of its cluster. Lloyd's algorithm alternates two steps: assign each point to its nearest centre, then move each centre to the mean of its points. Each step can only lower the objective, so it converges, but to a local optimum that depends on the starting centres, which is why implementations run several initialisations and keep the best.
# Lloyd's algorithm for k-means, the core of most clustering pipelines
def kmeans(X, k, iters=100):
centroids = random_sample(X, k) # k-means++ seeding is better
for _ in range(iters):
# assignment step: each point joins its nearest centroid
labels = [argmin_j(dist2(x, centroids[j])) for x in X]
# update step: each centroid moves to the mean of its points
new = [mean(X[labels == j]) for j in range(k)]
if new == centroids: # converged: objective cannot drop further
break
centroids = new
inertia = sum(dist2(x, centroids[l]) for x, l in zip(X, labels))
return labels, centroids, inertia # local optimum: rerun with several seedsNotice what the objective assumes. Squared Euclidean distance means features with larger numeric ranges dominate, so unscaled data clusters on whichever column has the biggest numbers. And minimising distance to a centre favours compact, roughly round clusters of similar size. Long, curved or very unequal groups get split or merged. Density-based methods such as DBSCAN and HDBSCAN, and Gaussian mixture models, make different assumptions and suit different shapes.
Principal component analysis finds orthogonal directions of maximum variance; projecting onto the first few keeps as much variation as possible in fewer dimensions. It is computed with a singular value decomposition of the centred data. It is used to compress features, to visualise data, and as a preprocessing step that removes noise before clustering. Its assumption is that variance means information, which fails when the signal you care about is small and a nuisance factor, such as image brightness, varies a lot.
Anomaly detection models what is typical and scores how unusual each point is. Isolation forests, for instance, count how few random splits it takes to isolate a point; rare points isolate quickly. These methods find rarity, which is not the same as badness, a distinction that matters in fraud and security work.
Evaluating without an answer key
Because an unsupervised objective is a proxy, you need other evidence that the structure is real and useful. No single number settles it, so combine several checks.
- Internal metrics. Inertia always falls as k grows, so look for where the improvement flattens rather than its minimum. The silhouette score compares each point's distance to its own cluster with its distance to the nearest other cluster, ranging from -1 to 1. Both reward compact, separated clusters and inherit k-means' shape assumptions.
- Stability. Rerun with different seeds or on random subsamples and measure agreement with the adjusted Rand index. Clusters that reshuffle every run are artefacts of initialisation, not structure.
- A small labelled probe. Label a few hundred examples by hand and check how well clusters line up with the categories that matter to you. This borrows a little supervision to validate a lot of unsupervised work.
- Human inspection. Read twenty examples from each cluster and the terms or features that distinguish them. If you cannot name a cluster, it probably is not one.
- Downstream usefulness. The final test is whether the structure improves something measurable: routing accuracy, a classifier trained on cluster-derived features, or an investigation queue with a higher hit rate.
Worked example: discover, then classify
A support team has 50,000 historical tickets and wants automatic routing, but the existing categories are a free-text field that agents filled in inconsistently. Labelling all tickets against a taxonomy nobody has agreed on would waste weeks. The plan uses each family for what it does well.
Phase 1, discovery. Represent each ticket as a TF-IDF vector reduced to 100 dimensions with truncated SVD and normalised to unit length, so Euclidean distance behaves like cosine similarity. Run k-means for several values of k and record inertia and silhouette on a sample. Suppose silhouette peaks weakly around k of 10 and two seeds agree with an adjusted Rand index of 0.8: the structure is reasonably stable. Reading samples from each cluster yields names like billing disputes, password resets, shipping delays and API errors, plus one cluster that turns out to be auto-replies, which should be filtered out rather than routed.
Phase 2, labelling. The team merges two clusters that are really one topic, splits one that mixes refunds with invoices, and agrees nine categories. Annotators label 150 tickets per discovered cluster, about 1,500 in total; sampling per cluster ensures rare categories are covered, which random sampling would miss.
Phase 3, supervised model. A logistic regression on TF-IDF features, trained on 75% of the labelled set and tested on the rest, gives per-category precision and recall that can be checked against the routing requirement. Unlike the clusters, this model has an honest error rate, and it predicts the agreed categories rather than whatever k-means happened to find.
import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.decomposition import TruncatedSVD
from sklearn.preprocessing import Normalizer
from sklearn.pipeline import make_pipeline
from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score, adjusted_rand_score
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import classification_report
texts = load_tickets() # 50,000 unlabelled support tickets
# --- Phase 1: unsupervised discovery -------------------------------------
embed = make_pipeline(TfidfVectorizer(min_df=5, ngram_range=(1, 2), sublinear_tf=True),
TruncatedSVD(n_components=100, random_state=0),
Normalizer(copy=False)) # unit length: Euclidean ~ cosine
Z = embed.fit_transform(texts)
sample = np.random.default_rng(0).choice(len(Z), 5000, replace=False)
for k in (6, 8, 10, 12, 15):
km = KMeans(n_clusters=k, n_init="auto", random_state=0).fit(Z)
print(k, round(km.inertia_), round(silhouette_score(Z[sample], km.labels_[sample]), 3))
# stability: do two seeds agree on the same k?
a = KMeans(n_clusters=10, n_init="auto", random_state=1).fit_predict(Z)
b = KMeans(n_clusters=10, n_init="auto", random_state=2).fit_predict(Z)
print("seed agreement (ARI):", round(adjusted_rand_score(a, b), 3))
# --- Phase 2: humans name the clusters and label a stratified sample ---------
labelled = label_sample(texts, a, per_cluster=150) # [(text, category), ...]
X_txt, y = zip(*labelled)
# --- Phase 3: supervised classifier on the agreed taxonomy ------------------
X_tr, X_te, y_tr, y_te = train_test_split(X_txt, y, test_size=0.25,
stratify=y, random_state=0)
clf = make_pipeline(TfidfVectorizer(min_df=2, ngram_range=(1, 2), sublinear_tf=True),
LogisticRegression(max_iter=2000, class_weight="balanced"))
clf.fit(X_tr, y_tr)
print(classification_report(y_te, clf.predict(X_te)))The clustering never became the product. It made the labelling cheaper and the taxonomy better, and the supervised model, which can be measured, carries the production load. The same pattern works for transactions, logs, images and documents. When budgets are tight, choosing which tickets to label next with active learning cuts the labelling cost further.
The bridge: self-supervised and semi-supervised learning
Modern deep learning blurs the old line. Self-supervised learning creates labels from the data itself: a language model predicts the next token, a vision model predicts a masked patch or learns that two crops of one image belong together. The training loop is supervised, with a loss against a target, but no human provided the target, so it scales to unlabelled data. The result is a representation that a small supervised head can then adapt to a task, which is why pretrained embeddings would replace TF-IDF in a modern version of the ticket example.
Semi-supervised learning trains on a few labels and many unlabelled examples together, for instance by pseudo-labelling confident predictions or enforcing consistent predictions under perturbation. It is covered in semi-supervised learning. In practice most production systems are a mix: unsupervised or self-supervised representations, a modest labelled set and a supervised model on top.
Choosing for your problem
| Question | Points to supervised | Points to unsupervised |
|---|---|---|
| Do you know the categories or target? | Yes, and they are stable | No, you want to discover them |
| Can you get labels? | Affordable, consistent, available | Scarce, slow or ambiguous |
| How will results be judged? | Accuracy on a defined target | Insight, exploration, triage |
| Is the target rare and changing? | Enough labelled examples exist | Novel cases matter, like new fraud |
| Does output drive automated decisions? | Yes: needs a measurable error rate | Only with a human or a validated proxy |
Failure modes
- Clustering on scale, not meaning. Unscaled numeric features let the largest column decide the clusters. Standardise or normalise first.
- Treating clusters as ground truth. k-means always returns k groups, even for structureless data. Validate with stability and a labelled probe before acting on them.
- Rare equals bad. Anomaly detectors flag unusual but legitimate behaviour and miss attacks that look normal. Measure the hit rate of flagged items with reviewers.
- Label noise and leakage. Inconsistent annotation caps supervised accuracy, and features that encode the label, such as a resolution code, inflate test scores. Audit both.
- Stale taxonomy. A supervised model can only predict categories it was trained on. Periodically re-cluster recent unrouted or low-confidence items to find new ones.
- Distribution shift. Both families assume future data looks like training data. Monitor input statistics and prediction confidence in production.
What to do next
- Write down the question you want answered and whether a correct answer can be defined; if it can, plan a supervised model and a labelled test set.
- If the categories are unknown, embed a sample, run k-means for several k with scaled or normalised features, and check silhouette, seed stability and human-readable cluster names.
- Label a stratified sample per discovered cluster and fix the taxonomy with domain owners before labelling more.
- Train a simple supervised baseline, such as logistic regression, and report per-class precision and recall on held-out data.
- Try pretrained embeddings in place of TF-IDF and keep them only if the held-out metrics improve.
- Schedule periodic re-clustering of low-confidence predictions to catch new categories. Read overfitting before trusting a high test score.