Build a KNN Breast Tumor Classifier
Build and evaluate a leakage-safe KNN tumor classifier with scikit-learn.
Introduction
30 Second Summary
A single score can make a prediction system look dependable while hiding the cases it gets wrong. In health-related data, those hidden mistakes deserve a closer look.
In this project, you will build a command-line breast tumor classifier with K-nearest neighbors in scikit-learn using a real Kaggle dataset. The evaluation shows how each diagnosis class performs on held-out samples.
What You'll Build
Run your finished classifier to see its benign and malignant prediction errors exposed in a terminal report before a labeled confusion matrix opens in a desktop plot window.
By the end of this project, you'll have:
- A dataset check that loads 569 real tumor observations. You can confirm that the file contains 30 usable measurements.
- A model comparison that prints unscaled KNN accuracy beside the accuracy from a scaled pipeline on the same held-out observations.
- A class-level evaluation that prints precision, recall, F1-score, and support for Benign and Malignant predictions. A labeled confusion matrix then makes false predictions visible.
- Secret Mission: Use cross-validation to choose the number of neighbors without tuning against the held-out test set.
Are there any prerequisites?
You need basic Python syntax plus theoretical familiarity with supervised machine learning.
You also need Python 3.11 or newer, Visual Studio Code, and a Kaggle account.
Before We Start
Before any hands-on work begins, this checkpoint helps you define the purpose of your KNN classifier. You are building an educational learning exercise that stays separate from medical decision-making.
Set Up the Windows ML Workspace
A machine learning project can behave differently when its packages share space with unrelated work. This step prevents those conflicts before they reach your classifier.
An isolated virtual environment keeps this project's libraries separate. You will confirm a compatible Python runtime in Visual Studio Code before installing the pinned versions.
In this step, get ready to:
- Create the knn-breast-tumor workspace on your Desktop.
- Confirm that Python 3.11 or newer is available.
- Install the pinned libraries inside an activated .venv.
Prepare the project workspace
Windows File Explorer gives the project a predictable home on your Desktop. That location keeps the folder easy to find in later steps.
- Press the Windows key to open the search bar.
- Type File Explorer into the search field.
- Press Enter to open File Explorer.
- Select Desktop in the left navigation area.
- Press Ctrl+Shift+N to create a new folder.
- Type knn-breast-tumor as the folder name.
- Press Enter to save the name.
You will see the empty knn-breast-tumor folder on your Desktop.
- Press the Windows key to open the search bar.
- Type Visual Studio Code into the search field.
- Press Enter to open Visual Studio Code.
- Click File in the menu bar.
- Click Open Folder....
- Select the knn-breast-tumor folder on your Desktop.
- Confirm your folder selection in the dialog.
- Confirm that you trust the folder if the Workspace Trust dialog appears.
You will see KNN-BREAST-TUMOR at the top of the Explorer sidebar. This confirms that Visual Studio Code is using the correct workspace.
- Click Terminal in the menu bar.
- Click New Terminal.
The integrated terminal opens at the bottom of Visual Studio Code. Its profile should be PowerShell.
Terminal opened a different shell?
The shell selector is easy to miss. Set PowerShell as the default before continuing.
- Press Ctrl+Shift+P to open the Command Palette.
- Type Terminal: Select Default Profile into the Command Palette.
- Press Enter.
- Select PowerShell.
- Click Terminal in the menu bar.
- Click New Terminal.
Help me switch the Visual Studio Code terminal to PowerShell on Windows.
- Check the installed Python version by running this command:
python --version
What does this command do?
The --version option prints the Python version used by this terminal. The pinned libraries require Python 3.11 or newer.
✔️ I see version 3.11 or higher
Great, your Python runtime meets the package requirements. You can build the isolated environment with this interpreter.
- Continue to the next substep with the current PowerShell terminal.
ⓧ I see an older version
The installed version cannot run every pinned library in this project. Installing a newer Python version leaves your existing project files in place.
- Open the official Python downloads page in your browser.
- Use the Python Install Manager to install Python 3.14.
- Complete the installation prompts.
- Press Ctrl+Shift+` in Visual Studio Code to create a fresh terminal.
- Confirm the installed version by running this command:
python --version
What should I see?
The fresh terminal detects the newly installed interpreter. You should now see Python 3.14 in the version output.
ⓧ Command not found
PowerShell cannot currently find a Python installation. The Python Install Manager adds the runtime needed for this project.
- Open the official Python downloads page in your browser.
- Use the Python Install Manager to install Python 3.14.
- Complete the installation prompts.
- Press Ctrl+Shift+` in Visual Studio Code to create a fresh terminal.
- Confirm that PowerShell can now find Python by running this command:
python --version
What should I see?
The command should now print Python 3.14. This confirms that the new terminal can access the installed runtime.
Create and activate the virtual environment
The .venv folder stores a private Python environment inside knn-breast-tumor. Package installations made after activation stay tied to this project.
- Create .venv inside knn-breast-tumor by running this command:
python -m venv .venv
What does this command do?
The venv module creates an isolated Python environment. The environment lives in the local .venv folder.
- Activate the new environment by running this command:
.\.venv\Scripts\Activate.ps1
What does activation change?
Activation points this terminal at the Python interpreter inside .venv. Package commands now update the project environment.
Your PowerShell prompt now includes .venv near the beginning. That small prompt change confirms that the isolated environment is active.
PowerShell blocked activation?
PowerShell may block local activation scripts under its current execution policy. The fallback below changes the policy for your Windows user.
Help me understand why PowerShell blocked my virtual environment activation.
- If activation was blocked, allow signed local scripts before retrying with these commands:
- Approve the policy change if PowerShell displays a confirmation prompt.
Set-ExecutionPolicy -ExecutionPolicy RemoteSigned -Scope CurrentUser
.\.venv\Scripts\Activate.ps1
What changes here?
The first command applies the RemoteSigned policy to your current Windows user. The second command retries activation.
The policy command only needs to run once for this user account. Successful activation adds .venv to the PowerShell prompt.
Your project now has its own active Python environment. Package changes in the next substep remain isolated from your other Python projects.
Install and verify the pinned packages
scikit-learn provides the KNN model used by the classifier. pandas loads the tumor records from a CSV file.
Matplotlib displays the confusion matrix later in the project. Pinning each version makes the environment repeatable.
Package installation can pause while files download. A quiet terminal for a short stretch is normal.
- Install the three pinned packages into the active environment by running:
python -m pip install scikit-learn==1.9.1 pandas==3.0.6 matplotlib==3.11.2
What is being installed?
- scikit-learn 1.9.1 supplies the classifier plus its preprocessing tools.
- pandas 3.0.6 supplies the tabular data structures used for the CSV.
- Matplotlib 3.11.2 supplies the desktop plotting window.
- Exact version pins prevent an unexpected package update from changing the project behavior.
When the PowerShell prompt returns without an installation error, the three libraries are available inside .venv.
Package installation failed?
Check that the terminal prompt includes .venv. A missing environment marker means the packages may be installing through a different Python interpreter.
Help me debug the pinned package installation in my Windows virtual environment.
Before you run the final check, which three pinned package versions do you expect to find in the list?
- Verify the active environment by running this command:
python -m pip freeze
What does this command prove?
The freeze command lists the exact package versions installed through the active Python interpreter. Finding all three pins confirms that the project environment is ready.
You will see scikit-learn==1.9.1 in the output.
You will also see pandas==3.0.6 in the output.
The list will include matplotlib==3.11.2 as well.
Your Windows ML workspace is now isolated and reproducible. The classifier has a dependable foundation for every step that follows.
Missing a pinned package?
- Confirm that the PowerShell prompt includes .venv.
- Run the package installation command above again.
- Repeat the final package check.
Help me find a missing package in my pip freeze output.
Your isolated workspace is ready. Next, you will bring the Kaggle tumor dataset into the project and inspect its contents.
Download and Inspect the Kaggle Dataset
Your isolated workspace is ready for the first real input to your KNN project. A model can only learn responsibly when you know what its data contains.
A filename cannot confirm the target labels or the usable measurements inside a dataset. In this step, you will download a real Kaggle CSV into your project.
Your script will inspect the diagnosis labels. It will also report the feature count and measurement ranges before any model is fitted.
In this step, get ready to:
- Place the Kaggle dataset at data/breast-cancer.csv.
- Separate the diagnosis target from the usable features.
- Print the dataset checks from knn_breast_tumor.py.
Download and organize the dataset
The script needs a predictable file path so every later run reads the same dataset. The downloaded archive can be a little fiddly because its CSV keeps a long original name.
- Open the Kaggle breast cancer dataset page in your browser.
- Sign in with your Kaggle account.
- Use the page's dataset download control to download the archive.
- Extract the downloaded archive in Windows File Explorer.
You should see breast-cancer-wisconsin-data_data.csv inside the extracted folder. This confirms that you downloaded the expected dataset file.
- Switch back to the knn-breast-tumor workspace in Visual Studio Code.
- Use the new-folder control in the Explorer sidebar to create data inside knn-breast-tumor.
- Move breast-cancer-wisconsin-data_data.csv from File Explorer into the new data folder.
The Explorer sidebar should now show the original CSV inside data.
- Rename the CSV to breast-cancer.csv.
Good progress. Your dataset now has the stable path data/breast-cancer.csv that the script expects.
Load and separate the dataset
The script begins with the complete import set used throughout the project. Running these imports now confirms that the activated environment can find every installed package.
- Use the new-file control in the Explorer sidebar to create knn_breast_tumor.py inside knn-breast-tumor.
You should see an empty knn_breast_tumor.py tab open in the editor.
- Add the project imports to knn_breast_tumor.py by pasting this code:
import matplotlib.pyplot as plt
import pandas as pd
from sklearn.metrics import ConfusionMatrixDisplay, classification_report
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
What Do These Imports Provide?
- Matplotlib handles the plot window used for visual evaluation.
- pandas reads the CSV into a table that Python can inspect.
- scikit-learn provides the preprocessing and modeling components used by the classifier.
- Save knn_breast_tumor.py.
- Confirm that every import loads by running this command in the active PowerShell terminal:
python knn_breast_tumor.py
What Should You See?
PowerShell should return to its prompt without an import error. This proves that .venv can access the pinned packages.
Imports Not Loading?
Confirm that the terminal prompt still shows .venv as active. Switch back to the activated PowerShell terminal from the previous step if needed.
Check each import against the code block for spelling differences. Package imports are case-sensitive.
Help me diagnose the import problem in my Python script.
A supervised learning dataset separates the value to predict from the measurements used to make that prediction. Here, y holds the target while X holds the features.
- Add the data-loading section below from sklearn.preprocessing import StandardScaler by pasting this code:
DATA_PATH = "data/breast-cancer.csv"
# Load and inspect the Kaggle CSV.
df = pd.read_csv(DATA_PATH)
y = df["diagnosis"]
X = df.drop(columns=["id", "diagnosis"]).dropna(axis="columns", how="all")
print(f"Rows: {len(df)}")
print("Diagnosis counts:")
print(y.value_counts())
print(f"Usable features: {X.shape[1]}")
What Does This Code Do?
- DATA_PATH stores the stable location of the renamed CSV.
- df holds every row loaded from the dataset.
- y selects the diagnosis column as the prediction target.
- X removes id because it does not describe a tumor measurement.
- dropna(axis="columns", how="all") removes any feature column that is completely empty.
- The print statements expose the row total and diagnosis counts. They also report the number of usable features.
- Save knn_breast_tumor.py.
- Run the first dataset inspection in the active PowerShell terminal with this command:
python knn_breast_tumor.py
What Should You See?
You should see Rows: 569 followed by diagnosis counts for B and M.
You should also see Usable features: 30. These results prove that the script loaded the expected records and removed the non-feature columns.
Dataset Not Loading?
Confirm that the Explorer sidebar shows data/breast-cancer.csv inside knn-breast-tumor. A different folder or filename prevents DATA_PATH from finding the CSV.
Confirm that you extracted the archive before moving the CSV. The script cannot read the compressed archive as breast-cancer.csv.
Help me fix the dataset loading problem.
Inspect the feature ranges
KNN calculates distances between observations. Printing two measurement ranges reveals whether those distances combine features with very different numeric scales.
- Add the feature-range checks below print(f"Usable features: {X.shape[1]}") by pasting this code:
print(
"smoothness_mean range:",
min(X["smoothness_mean"]),
"to",
max(X["smoothness_mean"]),
)
print(
"area_mean range:",
min(X["area_mean"]),
"to",
max(X["area_mean"]),
)
What Do These Checks Measure?
Each print statement finds the smallest value and the largest value for one feature. Comparing those spans exposes the scale difference that the model must eventually handle.
- Save knn_breast_tumor.py.
Before you run the finished script, which feature do you expect to span the larger numeric range?
- Run the complete dataset inspection in the active PowerShell terminal with this command:
python knn_breast_tumor.py
What Should You See?
You should see Rows: 569 and Usable features: 30. The diagnosis section should include counts for both B and M.
The final lines should show minimum and maximum values for smoothness_mean and area_mean. Their different ranges make the scale gap visible.
That is your first real data checkpoint complete. You have proved that the classifier reads the expected observations and separates 30 usable measurements from the diagnosis target.
Missing a Feature Range?
Check that the column names are exactly smoothness_mean and area_mean. A spelling difference prevents X from selecting the requested feature.
Confirm that X comes from the same df loaded through DATA_PATH.
Help me debug the missing feature-range output.
✔️ Awesome, I've got everything!
Your script now validates the dataset structure and exposes two feature ranges. Make sure knn_breast_tumor.py is saved.
ⓧ I'd like to double check the full code
Compare your saved knn_breast_tumor.py with this complete version.
import matplotlib.pyplot as plt
import pandas as pd
from sklearn.metrics import ConfusionMatrixDisplay, classification_report
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
DATA_PATH = "data/breast-cancer.csv"
# Load and inspect the Kaggle CSV.
df = pd.read_csv(DATA_PATH)
y = df["diagnosis"]
X = df.drop(columns=["id", "diagnosis"]).dropna(axis="columns", how="all")
print(f"Rows: {len(df)}")
print("Diagnosis counts:")
print(y.value_counts())
print(f"Usable features: {X.shape[1]}")
print(
"smoothness_mean range:",
min(X["smoothness_mean"]),
"to",
max(X["smoothness_mean"]),
)
print(
"area_mean range:",
min(X["area_mean"]),
"to",
max(X["area_mean"]),
)
Your real dataset is organized and validated. Next, you will split these observations and train an unscaled KNN baseline.
Build the Unscaled KNN Baseline
Your dataset inspection confirmed 569 rows. You can now create a repeatable baseline from those records.
A K-nearest neighbors model classifies each sample using the labels of nearby training samples. This step trains a raw KNN model so you have a result to compare against later.
Your existing range output shows very different numeric scales. Large-number measurements can dominate the distance calculation that defines which samples count as neighbors.
In this step, get ready to:
- Create a stratified split for training data and held-out test data.
- Train a KNN classifier with five neighbors on the unscaled features.
- Print the unscaled test accuracy beneath the contrasting feature ranges.
Create the stratified split
A held-out test set checks the model on records it did not use for training. Stratification keeps the diagnosis proportions stable across both sets.
- Place your cursor below the closing parenthesis of the area_mean range printout in knn_breast_tumor.py.
- Create the four split variables by pasting this code:
# Keep the class proportions stable in the held-out test set.
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42,
stratify=y,
)
What Does This Code Do?
- The test_size=0.2 setting reserves 20 percent of the records for testing.
- The random_state=42 setting produces the same split every time the script runs.
- The stratify=y setting preserves the benign and malignant proportions in both sets.
- The four output variables keep the training features separate from the test features and their matching labels.
- Save knn_breast_tumor.py.
- Confirm the split leaves the script runnable by running this command:
python knn_breast_tumor.py
What Does This Check Prove?
The command runs the updated script inside your active environment. A successful run proves that all four split variables are created without interrupting the inspection logic.
You should still see Rows: 569 plus Usable features: 30. Both feature range lines should follow.
Script No Longer Runs?
- Check that the split block sits after the complete area_mean printout.
- Match the parentheses and commas in the split block with the reference above.
Help me fix my stratified train and test split in knn_breast_tumor.py.
Train the raw KNN baseline
Your first model intentionally uses the raw measurements. Its test score becomes the comparison point for the next step.
- Place your cursor below the closing parenthesis of the train_test_split block.
- Build the unscaled baseline by pasting this code:
# Establish a deliberately unscaled baseline.
unscaled_model = KNeighborsClassifier(n_neighbors=5)
unscaled_model.fit(X_train, y_train)
unscaled_accuracy = unscaled_model.score(X_test, y_test)
print(f"Unscaled KNN accuracy: {unscaled_accuracy:.3f}")
How Does the Baseline Work?
- The KNeighborsClassifier(n_neighbors=5) constructor tells KNN to use five nearby training samples for each classification.
- The fit call learns from the unscaled training features and their diagnosis labels.
- The score call measures accuracy on the held-out test records.
- The :.3f format displays the resulting accuracy with three decimal places.
- Save knn_breast_tumor.py.
Before you run the baseline, do you expect measurements on larger numeric scales to exert more influence on KNN's distance calculations?
- Run the completed baseline with this command:
python knn_breast_tumor.py
What Should You See?
You should see an Unscaled KNN accuracy value after the two feature range lines. This confirms that the raw classifier trained successfully on the fixed split.
You now have a repeatable raw-model score. The much wider area_mean range gives that measurement more numerical influence than smoothness_mean during neighbor calculations.
Accuracy Line Missing?
- Check that the baseline block sits below the complete train_test_split block.
- Confirm that unscaled_model.fit(X_train, y_train) appears before the score call.
- Confirm that the final print line uses unscaled_accuracy.
Help me debug why the unscaled KNN accuracy is missing.
✔️ Awesome, I've got everything!
Your saved script now creates a fixed stratified split. It also trains and scores the unscaled KNN baseline.
ⓧ I'd like to double check the full code
- Compare the full reference below with knn_breast_tumor.py.
import matplotlib.pyplot as plt
import pandas as pd
from sklearn.metrics import ConfusionMatrixDisplay, classification_report
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
DATA_PATH = "data/breast-cancer.csv"
# Load and inspect the Kaggle CSV.
df = pd.read_csv(DATA_PATH)
y = df["diagnosis"]
X = df.drop(columns=["id", "diagnosis"]).dropna(axis="columns", how="all")
print(f"Rows: {len(df)}")
print("Diagnosis counts:")
print(y.value_counts())
print(f"Usable features: {X.shape[1]}")
print(
"smoothness_mean range:",
min(X["smoothness_mean"]),
"to",
max(X["smoothness_mean"]),
)
print(
"area_mean range:",
min(X["area_mean"]),
"to",
max(X["area_mean"]),
)
# Keep the class proportions stable in the held-out test set.
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42,
stratify=y,
)
# Establish a deliberately unscaled baseline.
unscaled_model = KNeighborsClassifier(n_neighbors=5)
unscaled_model.fit(X_train, y_train)
unscaled_accuracy = unscaled_model.score(X_test, y_test)
print(f"Unscaled KNN accuracy: {unscaled_accuracy:.3f}")
Your raw KNN baseline now exposes the distance problem in a working model. Next, you will keep feature scaling with KNN inside a pipeline so the preprocessing learns only from the training data.
Fix the Distance Problem with a Pipeline
Your unscaled KNN baseline now runs on a fixed split. Its feature-range output shows how large numeric measurements can steer the distance calculation.
Feature scaling puts the measurements on comparable scales. A scikit-learn Pipeline keeps scaling inside the training workflow. This prevents data leakage from the held-out test set.
In this step, get ready to:
- Combine feature scaling with the KNN classifier in one pipeline.
- Fit the pipeline on the existing training split.
- Compare the unscaled accuracy with the scaled pipeline accuracy.
Build the scaled pipeline
The pipeline treats preprocessing plus prediction as one model. Its scaler learns from X_train before KNN calculates any distances.
- In knn_breast_tumor.py, locate the final print(f"Unscaled KNN accuracy: {unscaled_accuracy:.3f}") line.
- Replace that final line with the code below.
# Keep scaling and KNN together so preprocessing is learned from training data.
model = Pipeline(
steps=[
("scaler", StandardScaler()),
("knn", KNeighborsClassifier(n_neighbors=5)),
]
)
model.fit(X_train, y_train)
scaled_accuracy = model.score(X_test, y_test)
print(f"Unscaled KNN accuracy: {unscaled_accuracy:.3f}")
print(f"Scaled pipeline accuracy: {scaled_accuracy:.3f}")
What does this code do?
- Pipeline runs its named steps in order.
- StandardScaler() learns the scale of each feature from the training data.
- model.fit(X_train, y_train) fits the scaler before training the five-neighbor classifier.
- model.score(X_test, y_test) applies the learned transformation to the held-out features before calculating accuracy.
- scaled_accuracy stores the pipeline's score for the side-by-side comparison.
- Save knn_breast_tumor.py.
Run the side-by-side comparison
Before you run the script, do you expect scaling to change the accuracy score?
- Switch back to the PowerShell terminal from earlier.
- Run the updated classifier by entering this command:
python knn_breast_tumor.py
What should I see?
The terminal prints a line beginning with Unscaled KNN accuracy. It also prints a line beginning with Scaled pipeline accuracy.
The scaled score can be higher or lower for this split. The pipeline guarantees that feature scaling learns only from the training data.
Both scores come from the same held-out observations. This keeps the comparison consistent.
You now have a direct comparison between raw distances and scaled distances. Your final model also protects the test set from influencing preprocessing.
Missing the scaled accuracy?
- Confirm that you saved knn_breast_tumor.py before running the command.
- Check that the pipeline block appears below the line that calculates unscaled_accuracy.
- Check that the PowerShell prompt still shows the active .venv environment.
Help me debug why my scaled KNN pipeline is not printing its accuracy.
✔️ Awesome, I've got everything!
Strong work. Your saved classifier now keeps scaling with KNN throughout training plus evaluation.
ⓧ I'd like to double check the full code
- Compare your complete knn_breast_tumor.py file with this reference.
import matplotlib.pyplot as plt
import pandas as pd
from sklearn.metrics import ConfusionMatrixDisplay, classification_report
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
DATA_PATH = "data/breast-cancer.csv"
# Load and inspect the Kaggle CSV.
df = pd.read_csv(DATA_PATH)
y = df["diagnosis"]
X = df.drop(columns=["id", "diagnosis"]).dropna(axis="columns", how="all")
print(f"Rows: {len(df)}")
print("Diagnosis counts:")
print(y.value_counts())
print(f"Usable features: {X.shape[1]}")
print(
"smoothness_mean range:",
min(X["smoothness_mean"]),
"to",
max(X["smoothness_mean"]),
)
print(
"area_mean range:",
min(X["area_mean"]),
"to",
max(X["area_mean"]),
)
# Keep the class proportions stable in the held-out test set.
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42,
stratify=y,
)
# Establish a deliberately unscaled baseline.
unscaled_model = KNeighborsClassifier(n_neighbors=5)
unscaled_model.fit(X_train, y_train)
unscaled_accuracy = unscaled_model.score(X_test, y_test)
# Keep scaling and KNN together so preprocessing is learned from training data.
model = Pipeline(
steps=[
("scaler", StandardScaler()),
("knn", KNeighborsClassifier(n_neighbors=5)),
]
)
model.fit(X_train, y_train)
scaled_accuracy = model.score(X_test, y_test)
print(f"Unscaled KNN accuracy: {unscaled_accuracy:.3f}")
print(f"Scaled pipeline accuracy: {scaled_accuracy:.3f}")
What does this reference confirm?
This file keeps the original inspection logic plus the fixed stratified split. It adds the scaled pipeline before printing both accuracy values.
Your classifier now uses leakage-safe scaling for its distance calculations. Next, you will inspect how well it predicts each diagnosis class.
Evaluate the Tumor Classifier
Your leakage-safe KNN pipeline now predicts held-out tumor records. Its single accuracy score still compresses every prediction into one number.
Class-level classification metrics reveal how the model behaves for each diagnosis. A confusion matrix makes the correct predictions visible beside the false predictions.
In this step, get ready to:
- Generate held-out predictions from the fitted pipeline.
- Print class-level metrics for benign records and malignant records.
- Open a labeled confusion matrix to inspect false predictions.
Print class-level metrics
Accuracy reports the share of correct predictions across the full test set. A classification report separates that result into measurements for Benign predictions and Malignant predictions.
- In Visual Studio Code, return to knn_breast_tumor.py from earlier.
- Find this block at the bottom of the file:
scaled_accuracy = model.score(X_test, y_test)
print(f"Unscaled KNN accuracy: {unscaled_accuracy:.3f}")
print(f"Scaled pipeline accuracy: {scaled_accuracy:.3f}")
Where Does This Edit Begin?
This block currently calculates the scaled score. It then prints both accuracy results.
- Replace the block you found with this evaluation section:
scaled_accuracy = model.score(X_test, y_test)
y_pred = model.predict(X_test)
print(f"Unscaled KNN accuracy: {unscaled_accuracy:.3f}")
print(f"Scaled pipeline accuracy: {scaled_accuracy:.3f}")
print("\nClassification report:")
print(
classification_report(
y_test,
y_pred,
labels=["B", "M"],
target_names=["Benign", "Malignant"],
)
)
What Does This Code Do?
- The model.predict(X_test) call classifies every held-out record with the fitted pipeline.
- The y_pred variable holds the resulting diagnosis predictions.
- The classification_report function compares those predictions with the true labels in y_test.
- The labels argument keeps B before M in the report.
- The target_names argument displays those classes as Benign and Malignant.
- Save knn_breast_tumor.py.
Before you run the script, make a mental prediction about whether one accuracy score can reveal which diagnosis class has lower recall.
- Run the updated classifier in the activated PowerShell terminal by running this command:
python knn_breast_tumor.py
What Should You See?
The terminal prints Classification report: below the two accuracy values. You should see separate rows for Benign and Malignant.
- Precision shows how often predictions for a displayed class are correct.
- Recall shows how many true records from a displayed class the model finds.
- F1-score balances precision with recall.
- Support counts the true records from each class in the test set.
Classification Report Missing?
- Check that y_pred = model.predict(X_test) appears below the scaled_accuracy calculation.
- Check that the closing parentheses around classification_report match the replacement block.
- Save knn_breast_tumor.py before rerunning the script.
Help me debug the missing classification report.
Open the labeled confusion matrix
A confusion matrix compares each true diagnosis with the model's predicted diagnosis. Fixed display labels make the two classes easy to distinguish in the plot.
- Scroll below the classification_report block in knn_breast_tumor.py.
- Add this plotting block at the end of the file:
ConfusionMatrixDisplay.from_predictions(
y_test,
y_pred,
labels=["B", "M"],
display_labels=["Benign", "Malignant"],
)
plt.show()
What Does This Code Do?
- The ConfusionMatrixDisplay.from_predictions method compares y_test with y_pred.
- The labels argument preserves the same B followed by M class order used in the report.
- The display_labels argument places Benign and Malignant on the plot axes.
- The plt.show() call displays the completed figure in a desktop window.
- Save knn_breast_tumor.py.
Before you run the final check, predict where the false predictions will appear in the matrix.
- Run the completed classifier in the activated PowerShell terminal by running this command:
python knn_breast_tumor.py
What Should You See?
The terminal shows precision, recall, F1-score, and support for Benign and Malignant. A labeled Matplotlib confusion matrix opens in a desktop plot window.
- Inspect the cells outside the main diagonal of the confusion matrix.
The diagonal cells contain correct predictions. The off-diagonal cells contain false predictions where the predicted diagnosis differs from the true diagnosis.
Confusion Matrix Window Missing?
- Confirm that plt.show() remains the final line in knn_breast_tumor.py.
- Check the terminal for an earlier Python error that stopped the script before the plot block.
- Rerun the script after saving the file.
Help me debug the missing confusion matrix window.
Keep This Experiment Educational
This matrix describes model behavior on held-out Kaggle records. It does not establish clinical safety.
Keep this classifier within an educational experiment. The classifier has no role in medical decisions.
That is the full evaluation loop working. Your terminal now exposes class-level behavior while your plot makes each false prediction visible.
✔️ Awesome, I've got everything!
Your saved classifier now prints class-level metrics and opens a labeled confusion matrix.
ⓧ I'd like to double check the full code
- Compare your saved knn_breast_tumor.py file with the complete version below.
import matplotlib.pyplot as plt
import pandas as pd
from sklearn.metrics import ConfusionMatrixDisplay, classification_report
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
DATA_PATH = "data/breast-cancer.csv"
# Load and inspect the Kaggle CSV.
df = pd.read_csv(DATA_PATH)
y = df["diagnosis"]
X = df.drop(columns=["id", "diagnosis"]).dropna(axis="columns", how="all")
print(f"Rows: {len(df)}")
print("Diagnosis counts:")
print(y.value_counts())
print(f"Usable features: {X.shape[1]}")
print(
"smoothness_mean range:",
min(X["smoothness_mean"]),
"to",
max(X["smoothness_mean"]),
)
print(
"area_mean range:",
min(X["area_mean"]),
"to",
max(X["area_mean"]),
)
# Keep the class proportions stable in the held-out test set.
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42,
stratify=y,
)
# Establish a deliberately unscaled baseline.
unscaled_model = KNeighborsClassifier(n_neighbors=5)
unscaled_model.fit(X_train, y_train)
unscaled_accuracy = unscaled_model.score(X_test, y_test)
# Keep scaling and KNN together so preprocessing is learned from training data.
model = Pipeline(
steps=[
("scaler", StandardScaler()),
("knn", KNeighborsClassifier(n_neighbors=5)),
]
)
model.fit(X_train, y_train)
scaled_accuracy = model.score(X_test, y_test)
y_pred = model.predict(X_test)
print(f"Unscaled KNN accuracy: {unscaled_accuracy:.3f}")
print(f"Scaled pipeline accuracy: {scaled_accuracy:.3f}")
print("\nClassification report:")
print(
classification_report(
y_test,
y_pred,
labels=["B", "M"],
target_names=["Benign", "Malignant"],
)
)
ConfusionMatrixDisplay.from_predictions(
y_test,
y_pred,
labels=["B", "M"],
display_labels=["Benign", "Malignant"],
)
plt.show()
Secret mission
Tune the Number of Neighbors
Five neighbors was a sensible starting choice. Build a cross-validated search that compares several neighbor counts before one final evaluation on the held-out test set.
Clean Up Your Resources
Clean Up Your Resources
All project work stays inside the local knn-breast-tumor folder. There are no cloud resources or ongoing costs.
Decide whether to keep the workspace ready. You can also pause your session or delete the folder entirely.
Resources you used:
- The local knn-breast-tumor folder containing every project file.
Keep everything running
No action is needed. Choose this if you plan to keep testing neighbor values or evaluation settings.
- Keep the knn-breast-tumor folder in its current location.
- Return to the existing Visual Studio Code workspace when you want to rerun the classifier.
Pause - I'll come back to this later
End the active local session while preserving every project file. This option keeps the workspace ready for your next experiment.
- Close the confusion matrix window if it is still open.
- Close the open Visual Studio Code window.
The terminal closes with its active .venv session. The complete knn-breast-tumor folder remains on your computer.
Delete - I don't want to use this again
Deleting the folder permanently removes every project file. Copy it somewhere else first if you want to preserve your results.
- Close the confusion matrix window if it is still open.
- Close the open Visual Studio Code window.
- Press the Windows key to open system search.
- Type knn-breast-tumor into the search field.
- Select the knn-breast-tumor folder in the results.
- Press Shift+Delete to delete the selected folder permanently.
- Confirm the permanent deletion in the prompt.
A final search confirms that the local resource is gone.
- Press the Windows key to reopen system search.
- Type knn-breast-tumor into the search field again.
You should see no matching project folder in the results.
Nice Work!
Nice Work!
You did it! You’ve built an educational KNN classifier with scikit-learn. It turns a real Kaggle dataset into inspectable predictions for benign and malignant tumor records.
You’ve learned how to:
- Load real CSV data into a clean feature table. Confirm 569 rows plus 30 usable measurements before training.
- Compare an unscaled KNN baseline against a scaled model on the same held-out observations. Keep feature scaling inside a pipeline to prevent data leakage.
- Inspect class-level metrics for benign and malignant predictions. Display a labeled confusion matrix that makes off-diagonal mistakes visible.
- Secret Mission: Use cross-validation to compare several neighbor counts using only the training data. Select one neighbor count before calculating the final held-out test score.
Ready to quiz yourself?