Skip to content

An independent study reference written by Dr Phuc V. Nguyen. It is not official subject material — for assessment requirements always follow your subject outline and vUWS.

Supervised and unsupervised learning

The split turns on one question: do you already have the answer for past cases? Supervised learning starts from examples carrying a known outcome, called a label, and learns a rule mapping inputs to that outcome so it can be applied to new cases. Unsupervised learning has no labels. It searches the inputs themselves for structure, grouping similar records or compressing many variables into a few. The practical difference is how you check the work. Supervised results can be scored against withheld labels, while unsupervised results have no answer sheet and must be judged on stability and usefulness.

Try it yourself

Unsupervised: k-means on unlabelled customers

The same 120 customers appear in both modes. Here nothing is labelled, so there is no answer sheet. k-means splits the cloud into however many groups you ask for, and the only honest questions are whether the split is stable and whether anyone can use it. Set k, then walk the loop one step at a time.

WCSS at convergence, k = 3
20,467
squared index points
Refit agreement, adjusted Rand index
0.98
1.00 is the same split, 0.00 is chance
00252550507575100100monthly spend index (0 to 100)visits index (0 to 100)123

Both features are synthetic index numbers, built to run 0 to 100 with comparable spread, about 23 index points each across this cloud, and with no deliberate difference in importance. That is why a step of one carries about the same weight on either axis here and this cloud needs no rescaling. Shared bounds on their own would not buy that. Two indices can both run 0 to 100 and still have spreads of 2 and 30, or be meant to matter unequally. In real work the scaling decision turns on units, spread, outliers and what the business thinks each feature is worth, and getting it wrong lets one feature decide every distance on its own.

Converged WCSS at every k, and why WCSS cannot choose k
57,242k = 220,467k = 3on screen17,543k = 414,548k = 512,173k = 6

The best attainable WCSS cannot rise as k rises, because splitting one group in two always matches the old total and usually beats it. It falls at every step here, so picking the lowest bar would pick k = 6 every time. That is why the lowest WCSS is not an answer to the question of how many groups there are.

WCSS at this step 20,467Groups 3Smallest group 36 customers
Number of groups, k3 groups
Step 7 of 7. The assignment repeated itself, so the loop has stopped.
Three tests the concept page sets for a clustering result
  1. Do the groups differ in ways a marketing team can act on. That is a business judgement and this widget does not make it. The group table below gives the raw material.
  2. Do they survive refitting on a different half of the customers. Measured here. An adjusted Rand index of 0.98 at k = 3, against 1.00 for the same split and 0.00 for chance.
  3. Can a person describe each group in a sentence. Also a person's job, from the centres and sizes in the table.
The 3 groups at step 7 of 7. Centres are index points on the same 0 to 100 scale as the axes. Sizes read an em dash until the first assign step, because before it no customer belongs to a group.
GroupCustomersCentre spendCentre visitsMarker
Group 14426.634.0circle
Group 23669.127.1square
Group 34070.274.7triangle
At k = 3 the loop has stopped, and the within-cluster sum of squares stands at 20,467 squared index points. Group sizes are 44, 36, 40. WCSS measures fit and nothing else. The refit test is the only number here that speaks to whether the same split comes back from a different half of the file, and even that is reproducibility rather than proof that the groups are real. No clustering run can supply that proof. The adjusted Rand index between this split and the refits is 0.98, 1.00 on the even-indexed half and 0.95 on the odd-indexed half. It reads both partitions, so a pair this k separated and a refit merges counts against the score and not only a pair that was broken, and it is corrected for the agreement two unrelated splits reach by chance, which is why it can be compared across different k with more confidence than a raw count of pairs held together. Across k = 2 to 6 the highest agreement here is at k = 3 on 0.98, and the lowest WCSS is at k = 6. The two criteria point at different k here, and only one of them is about whether the split reproduces at all.
Stylised cloud, not observed customers, generated from three seeded blobs. The generator's own blob membership is never shown, because unsupervised work has no such key. The halves are the even-indexed and odd-indexed records, and the blobs were interleaved when the cloud was built so that split is not a split on blob. The run on screen is the best of 6 seeded k-means++ starts for this k, judged on final WCSS, so the same k always gives the same answer. A group that ends an assign step with no customers keeps its centre where it is and is reported as empty. That happens at no step of any run shown here, and the rule is in the code so the recentre step can never divide by zero.

Why it matters

Supervised learning is a student with a marked practice exam. Every question comes with the right answer, so the student can test the rule they invented and correct it. Unsupervised learning is being handed a box of unlabelled photographs and asked to sort them into piles. There is no answer sheet. Someone has to look at the piles afterwards and say whether they mean anything.

Before you read on — recall

A fraud team holds 2,000 confirmed fraudulent transactions among 4 million records and wants to flag new ones. The strongest objection to reporting accuracy as the headline measure is

Formulas

What supervised learning is fitting
f^=arg⁡min⁡f  1n∑i=1nL(yi, f(xi))\hat{f} = \arg\min_{f} \; \frac{1}{n} \sum_{i=1}^{n} L\big(y_i,\, f(x_i)\big)
Each training record has inputs xix_i and a known outcome yiy_i. The loss LL scores how wrong a prediction is, using squared error for a number or a penalty for a wrong class. The algorithm searches for the rule ff with the lowest average loss. That average is measured on the training records, which is exactly why performance must then be rechecked on data the rule never saw.
What k-means clustering is minimising
J=∑k=1K∑i∈Ck∥xi−μk∥2J = \sum_{k=1}^{K} \sum_{i \in C_k} \lVert x_i - \mu_k \rVert^2
A common unsupervised method. Pick KK centres, assign every record to its nearest centre, move each centre to the mean of the records assigned to it, and repeat until nothing moves. The quantity JJ is the total squared distance from records to their own centre. Its best attainable value cannot increase as KK rises, so it cannot tell you the right number of groups, and that choice stays a judgement.

Worked examples

Scenario

A telecommunications retailer wants to reduce churn. It holds 24 months of account records and knows which customers left. Which kind of learning applies, and what is the first trap?

Solution

The known outcome makes this supervised. Define the label precisely first: left within 90 days of the observation date, not left at any point ever, or the rule will quietly learn from the future. Then split the data, fit on one part and score on the withheld part. The trap is accuracy. If 3 in every 100 customers churn, a rule predicting that nobody churns scores 97 per cent and is worthless. Judge it instead on how many genuine leavers land in the top slice the retention team can afford to contact.

Scenario

The same retailer asks for customer segments to guide its marketing. No list of correct segments exists anywhere in the business.

Solution

This is unsupervised. A clustering run will always return groups, however many you ask for, and the algorithm cannot tell you whether they are real. Judge them on use instead. Do the groups differ in ways the marketing team can act on, do they survive refitting on a different half of the customers, and can a person describe each one in a sentence. Groups that fail those three tests are arithmetic rather than segments, and campaigns built on them are aimed at nobody in particular.

Common mistakes

  • ✗Unsupervised learning is easier because there is no label to prepare. It is harder to trust. Without labels there is no independent score, so the only available checks are stability under refitting, interpretability, and whether acting on the output works.
  • ✗A high accuracy figure means a good model. Accuracy misleads whenever the classes are unbalanced. Predicting the majority class every time scores well and helps nobody, so report measures that show how well the rare class is found.
  • ✗The label is an objective fact about the customer. Someone defined it. Whether a customer counts as churned after 30, 90 or 180 days of inactivity is a business decision, and changing it changes what the model learns and how useful it is.
  • ✗Clustering discovers the true segments hidden in the data. It partitions whatever you give it into however many groups you request. The structure returned depends on the variables included, how they were scaled, and the number of clusters chosen.

Revision bullets

  • •Supervised needs labelled outcomes, unsupervised has none
  • •Supervised splits into classification (a category) and regression (a number)
  • •Unsupervised covers clustering, association and dimension reduction
  • •Only supervised work can be scored against withheld labels
  • •Accuracy misleads under class imbalance, so check the rare class
  • •The label definition is a business choice that shapes everything downstream

Quick check

A fraud team holds 2,000 confirmed fraudulent transactions among 4 million records and wants to flag new ones. The strongest objection to reporting accuracy as the headline measure is

A clustering run on customer data returns five neat groups. Before presenting them, the most informative check is

Connected topics

More in Decisions and Models

Sources

  1. James, G., Witten, D., Hastie, T., & Tibshirani, R. An Introduction to Statistical Learning. 2nd ed., Springer, 2021.
    Introductory treatment of the supervised and unsupervised split, of splitting data into training and test sets, and of why fit to training data overstates real performance.
  2. Hastie, Tibshirani & Friedman (2009)
    Hastie, T., Tibshirani, R., & Friedman, J. The Elements of Statistical Learning. 2nd ed., Springer, 2009.
    Fuller technical account, including the point that unsupervised results lack any external criterion for correctness.
How to cite this page
Dr. Phil's Quant Lab. (2026). Supervised and unsupervised learning. Business Analytics Atlas. https://phucnguyenvan.com/analytics_atlas/concept/ba-supervised-unsupervised
Next concept
Models as deliberate abstraction
Built by Dr. Phuc V. Nguyen ·Follow on LinkedInWork with PhilEmail