Classification — predicting categories
Describe classification problems and the difference between scores, thresholds, and hard labels.
Sign in to track progress on this lesson.
01 — Main lesson
Full walk-through. · 7.4 MB
Spoken script — useful when names or terms sound ambiguous.
Welcome back. Last time you learned to predict a number, like how many minutes late a delivery will arrive. Today we take the next step. Instead of predicting a number, we predict a category. This is called classification, and it is probably the most common type of machine learning problem in the real world.
Let me start with a simple shift in thinking. In regression, your model outputs a continuous value. Twelve minutes late, four minutes late, thirty minutes late. In classification, your model outputs a label. Spam or not spam. Fraud or not fraud. Cat, dog, or bird. The answer is a category, not a number on a sliding scale.
Now, there are two flavours of classification you need to know. The first is binary classification. That means there are exactly two possible answers. Yes or no. Positive or negative. Fraud or legitimate. Most business problems start here. The second is multi-class classification. That means there are three or more possible answers, but still only one correct answer per example. Think of classifying an email as personal, work, promotional, or social. Or classifying a tumour image as benign, malignant, or uncertain. The key point is that in multi-class, each example belongs to exactly one category. You are not allowed to pick two.
There is actually a third flavour called multi-label, where an example can belong to several categories at once, like tagging a photo with both beach and sunset. But that is a topic for another walk. For today, let us stay with binary and multi-class.
Here is something that surprises people new to classification. Most classification models do not actually output a label directly. They output a score. A number. Often that number looks like a probability, somewhere between zero and one. For example, your fraud model might look at a transaction and output zero point zero three. That means it thinks there is a three percent chance this transaction is fraudulent. Another transaction might get zero point nine one. Ninety-one percent likely to be fraud.
But your business does not want a probability. Your business wants a decision. Do we block this transaction or let it through? Do we flag this email as spam or put it in the inbox? To turn that score into a decision, you need a threshold. A threshold is just a cutoff. If the score is above the threshold, you say yes. If it is below, you say no.
This is one of the most important ideas in classification, so let me say it clearly. The model gives you a score. You choose the threshold. The threshold is where your judgement enters the system, just like the loss function did in regression.
Let me make this concrete. Imagine you work for a bank and you have built a fraud detection model. For every card transaction, the model outputs a fraud score between zero and one. You need to decide what threshold to use. If you set the threshold at zero point five, you only block transactions where the model is more than fifty percent sure it is fraud. If you set it at zero point two, you block anything where the model is even slightly suspicious. If you set it at zero point nine, you only block the most obvious fraud.
Each choice has consequences. A low threshold means you catch more fraud, but you also block more legitimate transactions. Angry customers. A high threshold means you rarely bother legitimate customers, but you miss more fraud. Lost money. There is no universally correct threshold. It depends on the cost of each kind of mistake.
Let us talk about those mistakes, because this is where classification gets genuinely interesting and where a lot of people go wrong.
When your model makes a prediction, there are four possible outcomes for a binary problem. Let us use the fraud example. The model says fraud or not fraud. The truth is fraud or not fraud. That gives us four combinations.
The first is the model says fraud and it is fraud. That is a true positive. You caught the bad guy. Good. The second is the model says not fraud and it is not fraud. That is a true negative. You let an honest customer through. Also good. The third is the model says fraud but it is actually legitimate. That is a false positive. You blocked an honest customer. They are annoyed. The fourth is the model says not fraud but it is actually fraud. That is a false negative. You let a criminal through. You lost money.
This two-by-two grid is called a confusion matrix. It sounds technical, but it is just a way of counting your hits and your misses. Let me tell it as a story. You ran your fraud model on one thousand transactions last week. Seven hundred were correctly flagged as legitimate. Two hundred were correctly flagged as fraud. Fifty were legitimate but got blocked. Fifty were fraud but got through. That is your confusion matrix. Four numbers. Two good outcomes, two bad outcomes. Once you can tell that story, you understand the confusion matrix.
Now, the simplest metric people reach for is accuracy. Accuracy is the fraction of predictions that were correct. In that example, seven hundred plus two hundred is nine hundred correct out of one thousand. So accuracy is ninety percent. That sounds great. But here is the trap. Imagine only ten of those one thousand transactions were actually fraud. Your model could have achieved ninety percent accuracy by simply saying not fraud every single time. It would catch zero fraud and still look fantastic on paper.
This is the imbalance problem. When one category is much rarer than the other, accuracy becomes misleading. Fraud is rare. Disease is rare. Customer churn is rare. In all these cases, a model that always says no can look brilliant while being completely useless.
So we need better metrics. Two of the most important are precision and recall. Let me explain them in plain English.
Recall asks, of all the actual fraud cases, how many did we catch? If there were two hundred and fifty real fraud cases and we caught two hundred, our recall is two hundred out of two hundred and fifty, which is eighty percent. Recall is about not missing things. It is the true positives divided by all the actual positives.
Precision asks, of all the cases we flagged as fraud, how many were actually fraud? If we flagged two hundred and fifty transactions as fraud and two hundred were real fraud, our precision is two hundred out of two hundred and fifty, which is also eighty percent. Precision is about not crying wolf. It is the true positives divided by all the predicted positives.
These two metrics pull against each other. If you want to catch every single fraud case, you lower your threshold and flag more transactions. Your recall goes up. But you will also flag more legitimate transactions, so your precision goes down. If you want to be very sure that every flagged transaction is really fraud, you raise your threshold. Your precision goes up. But you will miss more fraud, so your recall goes down.
This trade-off is the heart of classification. You cannot maximise both at the same time. You have to choose based on what matters more in your situation.
Let me give you two contrasting examples so you can feel this in your bones.
First, imagine you are building a spam filter. If you flag a legitimate email as spam, the person might miss an important message from their boss or a job offer. That is costly. If you let a spam email through, it is annoying but usually not catastrophic. So for spam, you want high precision. You would rather let some spam through than risk hiding a real email. You set your threshold high.
Second, imagine you are building a screening model for a serious disease. If you miss a sick patient, they might not get treatment in time. That could be fatal. If you flag a healthy patient as possibly sick, they get a follow-up test. It is stressful and costs money, but it is not catastrophic. So for medical screening, you want high recall. You would rather have false alarms than miss a real case. You set your threshold low.
Do you see the difference? Same model structure, same metrics, completely different threshold choices. The threshold is a business decision, not a technical one.
Let me bring this together with a practical scenario. You are hired by an insurance company to build a model that flags potentially fraudulent claims. The company processes ten thousand claims a month. Historically, about two percent are fraudulent. You build a model that outputs a fraud score for each claim. You need to decide on a threshold.
You start by looking at what happens at different thresholds. At a threshold of zero point five, the model flags about one hundred claims a month. Of those, sixty are actually fraudulent. Precision is sixty percent. But there are two hundred real fraud cases, so recall is only thirty percent. You are missing one hundred and forty fraudulent claims a month. The finance team is not happy.
You lower the threshold to zero point two. Now the model flags four hundred claims. Of those, one hundred and twenty are actually fraudulent. Precision drops to thirty percent. But recall rises to sixty percent. You are catching more fraud, but your investigators are now looking at two hundred and eighty legitimate claims that got flagged. They are overwhelmed.
You try zero point three. The model flags two hundred claims. One hundred are fraudulent. Precision is fifty percent. Recall is fifty percent. You catch half the fraud and half your flagged claims are real. Is that good enough? That depends on the cost of an investigator's time versus the cost of a missed fraudulent claim.
This is the conversation you need to have with the business. You say, we can catch more fraud, but it will cost more in investigation time and we will upset more honest customers. We can be more conservative, but we will let more fraud through. Where do you want the dial set? That is the threshold conversation.
Notice what you did not do. You did not just pick zero point five because it sounds nice. You did not report accuracy and call it a day. You showed the trade-off in terms the business understands. Cost of fraud, cost of investigation, cost of customer friction. That is what good classification practice looks like.
Let me give you one more mental model before we wrap up. Think of the threshold as a dial. Turn it one way and you become more aggressive. You catch more of the thing you are looking for, but you also make more false accusations. Turn it the other way and you become more conservative. You make fewer false accusations, but you miss more of the real cases. Your job is not to find the magic number. Your job is to understand the costs on both sides and help the business choose wisely.
As you walk, try this. Think of a yes-or-no decision in your own work or life where you could use a model. Maybe it is whether to follow up on a sales lead. Maybe it is whether to recommend a customer for a credit upgrade. Ask yourself, what is the cost of a false positive? What is the cost of a false negative? Which one would hurt more? If you can answer that, you already know which way to turn the dial.
Next time, we will look at how to measure a classifier across all possible thresholds at once, so you can compare models before you ever pick a cutoff. For now, enjoy the rest of your walk.
Let me start with a simple shift in thinking. In regression, your model outputs a continuous value. Twelve minutes late, four minutes late, thirty minutes late. In classification, your model outputs a label. Spam or not spam. Fraud or not fraud. Cat, dog, or bird. The answer is a category, not a number on a sliding scale.
Now, there are two flavours of classification you need to know. The first is binary classification. That means there are exactly two possible answers. Yes or no. Positive or negative. Fraud or legitimate. Most business problems start here. The second is multi-class classification. That means there are three or more possible answers, but still only one correct answer per example. Think of classifying an email as personal, work, promotional, or social. Or classifying a tumour image as benign, malignant, or uncertain. The key point is that in multi-class, each example belongs to exactly one category. You are not allowed to pick two.
There is actually a third flavour called multi-label, where an example can belong to several categories at once, like tagging a photo with both beach and sunset. But that is a topic for another walk. For today, let us stay with binary and multi-class.
Here is something that surprises people new to classification. Most classification models do not actually output a label directly. They output a score. A number. Often that number looks like a probability, somewhere between zero and one. For example, your fraud model might look at a transaction and output zero point zero three. That means it thinks there is a three percent chance this transaction is fraudulent. Another transaction might get zero point nine one. Ninety-one percent likely to be fraud.
But your business does not want a probability. Your business wants a decision. Do we block this transaction or let it through? Do we flag this email as spam or put it in the inbox? To turn that score into a decision, you need a threshold. A threshold is just a cutoff. If the score is above the threshold, you say yes. If it is below, you say no.
This is one of the most important ideas in classification, so let me say it clearly. The model gives you a score. You choose the threshold. The threshold is where your judgement enters the system, just like the loss function did in regression.
Let me make this concrete. Imagine you work for a bank and you have built a fraud detection model. For every card transaction, the model outputs a fraud score between zero and one. You need to decide what threshold to use. If you set the threshold at zero point five, you only block transactions where the model is more than fifty percent sure it is fraud. If you set it at zero point two, you block anything where the model is even slightly suspicious. If you set it at zero point nine, you only block the most obvious fraud.
Each choice has consequences. A low threshold means you catch more fraud, but you also block more legitimate transactions. Angry customers. A high threshold means you rarely bother legitimate customers, but you miss more fraud. Lost money. There is no universally correct threshold. It depends on the cost of each kind of mistake.
Let us talk about those mistakes, because this is where classification gets genuinely interesting and where a lot of people go wrong.
When your model makes a prediction, there are four possible outcomes for a binary problem. Let us use the fraud example. The model says fraud or not fraud. The truth is fraud or not fraud. That gives us four combinations.
The first is the model says fraud and it is fraud. That is a true positive. You caught the bad guy. Good. The second is the model says not fraud and it is not fraud. That is a true negative. You let an honest customer through. Also good. The third is the model says fraud but it is actually legitimate. That is a false positive. You blocked an honest customer. They are annoyed. The fourth is the model says not fraud but it is actually fraud. That is a false negative. You let a criminal through. You lost money.
This two-by-two grid is called a confusion matrix. It sounds technical, but it is just a way of counting your hits and your misses. Let me tell it as a story. You ran your fraud model on one thousand transactions last week. Seven hundred were correctly flagged as legitimate. Two hundred were correctly flagged as fraud. Fifty were legitimate but got blocked. Fifty were fraud but got through. That is your confusion matrix. Four numbers. Two good outcomes, two bad outcomes. Once you can tell that story, you understand the confusion matrix.
Now, the simplest metric people reach for is accuracy. Accuracy is the fraction of predictions that were correct. In that example, seven hundred plus two hundred is nine hundred correct out of one thousand. So accuracy is ninety percent. That sounds great. But here is the trap. Imagine only ten of those one thousand transactions were actually fraud. Your model could have achieved ninety percent accuracy by simply saying not fraud every single time. It would catch zero fraud and still look fantastic on paper.
This is the imbalance problem. When one category is much rarer than the other, accuracy becomes misleading. Fraud is rare. Disease is rare. Customer churn is rare. In all these cases, a model that always says no can look brilliant while being completely useless.
So we need better metrics. Two of the most important are precision and recall. Let me explain them in plain English.
Recall asks, of all the actual fraud cases, how many did we catch? If there were two hundred and fifty real fraud cases and we caught two hundred, our recall is two hundred out of two hundred and fifty, which is eighty percent. Recall is about not missing things. It is the true positives divided by all the actual positives.
Precision asks, of all the cases we flagged as fraud, how many were actually fraud? If we flagged two hundred and fifty transactions as fraud and two hundred were real fraud, our precision is two hundred out of two hundred and fifty, which is also eighty percent. Precision is about not crying wolf. It is the true positives divided by all the predicted positives.
These two metrics pull against each other. If you want to catch every single fraud case, you lower your threshold and flag more transactions. Your recall goes up. But you will also flag more legitimate transactions, so your precision goes down. If you want to be very sure that every flagged transaction is really fraud, you raise your threshold. Your precision goes up. But you will miss more fraud, so your recall goes down.
This trade-off is the heart of classification. You cannot maximise both at the same time. You have to choose based on what matters more in your situation.
Let me give you two contrasting examples so you can feel this in your bones.
First, imagine you are building a spam filter. If you flag a legitimate email as spam, the person might miss an important message from their boss or a job offer. That is costly. If you let a spam email through, it is annoying but usually not catastrophic. So for spam, you want high precision. You would rather let some spam through than risk hiding a real email. You set your threshold high.
Second, imagine you are building a screening model for a serious disease. If you miss a sick patient, they might not get treatment in time. That could be fatal. If you flag a healthy patient as possibly sick, they get a follow-up test. It is stressful and costs money, but it is not catastrophic. So for medical screening, you want high recall. You would rather have false alarms than miss a real case. You set your threshold low.
Do you see the difference? Same model structure, same metrics, completely different threshold choices. The threshold is a business decision, not a technical one.
Let me bring this together with a practical scenario. You are hired by an insurance company to build a model that flags potentially fraudulent claims. The company processes ten thousand claims a month. Historically, about two percent are fraudulent. You build a model that outputs a fraud score for each claim. You need to decide on a threshold.
You start by looking at what happens at different thresholds. At a threshold of zero point five, the model flags about one hundred claims a month. Of those, sixty are actually fraudulent. Precision is sixty percent. But there are two hundred real fraud cases, so recall is only thirty percent. You are missing one hundred and forty fraudulent claims a month. The finance team is not happy.
You lower the threshold to zero point two. Now the model flags four hundred claims. Of those, one hundred and twenty are actually fraudulent. Precision drops to thirty percent. But recall rises to sixty percent. You are catching more fraud, but your investigators are now looking at two hundred and eighty legitimate claims that got flagged. They are overwhelmed.
You try zero point three. The model flags two hundred claims. One hundred are fraudulent. Precision is fifty percent. Recall is fifty percent. You catch half the fraud and half your flagged claims are real. Is that good enough? That depends on the cost of an investigator's time versus the cost of a missed fraudulent claim.
This is the conversation you need to have with the business. You say, we can catch more fraud, but it will cost more in investigation time and we will upset more honest customers. We can be more conservative, but we will let more fraud through. Where do you want the dial set? That is the threshold conversation.
Notice what you did not do. You did not just pick zero point five because it sounds nice. You did not report accuracy and call it a day. You showed the trade-off in terms the business understands. Cost of fraud, cost of investigation, cost of customer friction. That is what good classification practice looks like.
Let me give you one more mental model before we wrap up. Think of the threshold as a dial. Turn it one way and you become more aggressive. You catch more of the thing you are looking for, but you also make more false accusations. Turn it the other way and you become more conservative. You make fewer false accusations, but you miss more of the real cases. Your job is not to find the magic number. Your job is to understand the costs on both sides and help the business choose wisely.
As you walk, try this. Think of a yes-or-no decision in your own work or life where you could use a model. Maybe it is whether to follow up on a sales lead. Maybe it is whether to recommend a customer for a credit upgrade. Ask yourself, what is the cost of a false positive? What is the cost of a false negative? Which one would hurt more? If you can answer that, you already know which way to turn the dial.
Next time, we will look at how to measure a classifier across all possible thresholds at once, so you can compare models before you ever pick a cutoff. For now, enjoy the rest of your walk.
02 — Refresh
Short recap. · 795 KB
Spoken script — useful when names or terms sound ambiguous.
Quick recap. Classification is about predicting categories, not numbers. Binary classification has two outcomes, like fraud or not fraud. Multi-class has three or more, like cat, dog, or bird.
Most models output a score, often a probability between zero and one. You choose a threshold to turn that score into a yes-or-no decision. The threshold is your judgement call.
Accuracy can mislead when classes are imbalanced. If fraud is rare, a model that always says not fraud can look accurate while being useless.
Recall measures how many real cases you caught. Precision measures how many of your flagged cases were actually real. They pull against each other. Lower the threshold and recall rises but precision falls. Raise it and precision rises but recall falls.
A confusion matrix is just four counts. True positives, true negatives, false positives, and false negatives. It tells the story of your hits and misses.
Choose your threshold based on costs. For spam filtering, you want high precision to avoid hiding real emails. For disease screening, you want high recall to avoid missing sick patients. The threshold is a business decision, not a technical one.
Most models output a score, often a probability between zero and one. You choose a threshold to turn that score into a yes-or-no decision. The threshold is your judgement call.
Accuracy can mislead when classes are imbalanced. If fraud is rare, a model that always says not fraud can look accurate while being useless.
Recall measures how many real cases you caught. Precision measures how many of your flagged cases were actually real. They pull against each other. Lower the threshold and recall rises but precision falls. Raise it and precision rises but recall falls.
A confusion matrix is just four counts. True positives, true negatives, false positives, and false negatives. It tells the story of your hits and misses.
Choose your threshold based on costs. For spam filtering, you want high precision to avoid hiding real emails. For disease screening, you want high recall to avoid missing sick patients. The threshold is a business decision, not a technical one.