Build a Wine Cultivar Classifier

Train and evaluate a multiclass wine classifier in a Jupyter notebook.

Introduction

30 Second Summary

Wine labels tell you which grape produced a bottle. The chemistry inside the wine carries its own measurable clues.

In this project, you will build a wine cultivar classifier in a Jupyter Notebook with Python. Its output traces the path from a naive baseline to an explained prediction from a scaled logistic regression pipeline.

What You'll Build

The finished notebook shows a weak one-class pattern give way to a model that identifies a held-out wine with probabilities for all three cultivars.

By the end of this project, you'll have:

  • A Wine dataset summary showing 178 samples across 13 chemical features in three classes.
  • A baseline-to-model comparison that pairs each accuracy result with a confusion matrix.
  • A held-out prediction that names one wine's predicted cultivar alongside probabilities for all three classes.
  • Secret Mission: Use five-fold cross-validation to check whether one train/test result was unusually lucky or unlucky.

Are there any prerequisites?

Basic Python knowledge plus basic statistics are enough to begin. Your Windows computer needs Jupyter Notebook plus internet access for the package installation in Step 1.

Before We Start

This is your moment to commit to building a supervised multiclass classifier that predicts a wine cultivar from 13 chemical measurements. A most-frequent-class baseline gives you the evidence needed to judge whether the trained model improves on guesswork.

Set Up the Notebook Environment

A fair model comparison depends on every notebook cell running in the same verified environment. Your local Python installation needs a compatibility check before you can trust the results.

A Jupyter Notebook keeps your code beside its results. This step pins the required packages before you load the Wine dataset.

In this step, get ready to:
  • Confirm Python 3.11 or newer is available.
  • Install pinned releases of scikit-learn plus Matplotlib.
  • Create the notebook that loads the Wine dataset.
Prepare a compatible Python environment

The pinned scikit-learn release requires Python 3.11 or newer. Checking first prevents package compatibility problems during installation.

  • Press the Windows key to open Windows search.
  • Type PowerShell into the search field.
  • Press Enter to open Windows PowerShell.
  • Check the installed Python version by running this command:
py --version

What does this command check?

The Windows py launcher reports the Python version available to your commands. The result determines which path you follow below.

✔️ I see version 3.11 or newer

Your Python runtime meets the package requirements. Continue with the pinned package installation below.

ⓧ I see an older version

The reported Python version is below 3.11. Install the current 64-bit release before continuing.

  • Visit the official Python downloads for Windows page.
  • Select Download Windows installer (64-bit) for the current stable release.
  • Open the downloaded installer.
  • Select Install Now.

The installer updates the Python version available through the Windows launcher.

  • Close the existing PowerShell window.
  • Press the Windows key to return to Windows search.
  • Type PowerShell into the search field.
  • Press Enter to reopen Windows PowerShell.
  • Repeat the version check by running this command:
py --version

What should the new check show?

The command now reports Python 3.11 or newer. That confirms the updated runtime is available to the package installer.

ⓧ Command not found

The Python launcher is unavailable in PowerShell. Install the current 64-bit Python release to add it.

  • Visit the official Python downloads for Windows page.
  • Select Download Windows installer (64-bit) for the current stable release.
  • Open the downloaded installer.
  • Select Install Now.

The installer adds Python to the Windows launcher.

  • Close the existing PowerShell window.
  • Press the Windows key to return to Windows search.
  • Type PowerShell into the search field.
  • Press Enter to reopen Windows PowerShell.
  • Verify the installation by running this command:
py --version

What should the version check show?

The command now reports Python 3.11 or newer. The Windows launcher is ready for the installation command.

Pinned versions give every later cell the same library behavior. You will use scikit-learn 1.9.1, Matplotlib 3.11.2, plus Notebook 7.6.3.

This download includes several packages. Expect the installation to take a few minutes while PowerShell retrieves them.

  • Install the pinned environment by running this command:
py -m pip install "scikit-learn==1.9.1" "matplotlib==3.11.2" "notebook==7.6.3"

What does this command install?

  • The py -m pip portion runs the package installer through the compatible Python runtime.
  • The scikit-learn==1.9.1 pin provides the dataset loader plus the machine learning tools used later.
  • The matplotlib==3.11.2 pin provides the plotting tools for the confusion matrices.
  • The notebook==7.6.3 pin provides the browser-based notebook application.

Good progress. PowerShell finishes with an installation summary for the pinned packages.

Did the package installation stop?

  • Confirm the version check reports Python 3.11 or newer.
  • Check that your computer still has an internet connection.
  • Run the pinned installation command again after correcting the failed condition.

Help me troubleshoot the pinned Python package installation.

Create the Jupyter notebook

Jupyter runs a local notebook server from PowerShell. The server opens its browser interface while the PowerShell process stays active.

  • Start the notebook server by running this command:
jupyter notebook

What does this command do?

The jupyter notebook command starts the local server. It also opens the notebook dashboard in your default browser.

Keep this PowerShell window running. Your browser shows the Jupyter dashboard when the server is ready.

Didn't the dashboard open?

  • Return to the PowerShell window from earlier.
  • Check that the notebook server is still running without a failure.
  • Run the launch command again if the process stopped.

Help me open the local Jupyter Notebook dashboard.

  • Click New Notebook on the Jupyter dashboard.

Jupyter asks which available Python kernel should run the notebook.

  • Choose Python when Jupyter asks for a kernel.

You now see an empty notebook editor in the browser.

  • Click the notebook title at the top of the editor.

Jupyter opens the notebook naming field.

  • Enter wine_classifier.ipynb as the notebook name.
  • Press Enter to confirm the name.

The editor now shows wine_classifier.ipynb at the top. This is the notebook you build throughout the project.

Load and verify the Wine dataset

The Wine dataset is bundled with scikit-learn. Its shape gives you a quick compatibility check because the data should contain 178 samples with 13 chemical features.

Before you run the first check, what two dimensions do you expect the dataset to report?

  • Verify the dataset loader in the first notebook cell by entering this code:
from sklearn.datasets import load_wine; print(load_wine().data.shape)

What should you see?

Press Shift+Enter to execute the cell. You will see (178, 13) beneath it.

The first number counts the wine samples. The second number counts the chemical measurements available for each sample.

Didn't the shape appear?

  • Confirm the notebook is using the Python kernel selected earlier.
  • Compare the package name in the import with sklearn.
  • Return to PowerShell to confirm the notebook server is still running.

Help me run the Wine dataset verification cell.

The shape confirms that the package installation can load the complete dataset. You can now turn the verification cell into the reusable imports and variables for the classifier.

  • Select all code in the first cell of wine_classifier.ipynb.
  • Replace the selected code with this complete first cell:
# Cell 1: imports and data
import matplotlib.pyplot as plt
from sklearn.datasets import load_wine
from sklearn.dummy import DummyClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, ConfusionMatrixDisplay
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

wine = load_wine()
X = wine.data
y = wine.target

What does this cell prepare?

  • The import section makes the plotting, baseline, evaluation, splitting, scaling, pipeline, plus classification tools available to later cells.
  • The wine variable holds the loaded dataset with its feature names plus class names.
  • The X variable holds the 13 chemical measurements for all 178 samples.
  • The y variable holds the known cultivar class for each sample.
  • Save wine_classifier.ipynb using the notebook's save control.
  • Execute the completed first cell by pressing Shift+Enter.

The cell finishes without an error. Its execution count advances to confirm that every import succeeded.

Did the completed cell fail?

  • Compare each import with the full-code reference below.
  • Confirm PowerShell completed the pinned package installation.
  • Check that the notebook is still using the Python kernel selected earlier.

Help me fix the first cell in my wine classifier notebook.

✔️ Awesome, I've got everything!

Your first cell has run successfully. Double check that wine_classifier.ipynb is saved before continuing.

ⓧ I'd like to double check the full code

# Cell 1: imports and data
import matplotlib.pyplot as plt
from sklearn.datasets import load_wine
from sklearn.dummy import DummyClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, ConfusionMatrixDisplay
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

wine = load_wine()
X = wine.data
y = wine.target

That's the environment ready. Next, you will train a baseline that deliberately ignores all 13 measurements so you can see exactly where guesswork falls short.

Build a Naive Baseline

Your Jupyter Notebook already loads the Wine dataset. You now need a fair reference point for every model you test.

A baseline gives you that reference point by always predicting the most frequent class. It ignores all 13 chemical measurements.

In this step, get ready to:
  • Create a reproducible stratified train and test split.
  • Train a classifier that always predicts the most frequent class.
  • Expose the baseline's limitation with an accuracy score plus a confusion matrix.
Create a reproducible held-out split

A held-out test set gives you samples that the classifier does not use for training. Stratification keeps the class proportions similar across both parts of the dataset.

  • Switch back to wine_classifier.ipynb from earlier.
  • Click the empty cell directly below Cell 1.
  • Create the reproducible split by pasting this code into the empty cell:
# Cell 2: one reproducible held-out split
X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.25,
    random_state=42,
    stratify=y,
)

What does this code do?

  • The train_test_split function separates the measurements plus their labels into training and test sets.
  • The test_size=0.25 setting reserves one quarter of the samples for testing.
  • The random_state=42 setting produces the same split each time you run the notebook.
  • The stratify=y setting preserves the class proportions in both sets.
  • Save wine_classifier.ipynb using the save icon in the notebook toolbar.
  • Run Cell 2 by pressing Shift+Enter.

You'll see an execution number appear beside Cell 2. That number confirms the split ran successfully.

Seeing an error in Cell 2?

  • Run Cell 1 again if X or y is unavailable.
  • Check that the split defines X_train, X_test, y_train, plus y_test.

Help me fix my train and test split.

Run the naive baseline

The DummyClassifier class from scikit-learn provides a simple reference model. Its most_frequent strategy makes predictions without using the chemical measurements.

  • Click the empty cell directly below Cell 2.
  • Create the baseline evaluation by pasting this code into the empty cell:
# Cell 3: naive baseline
baseline = DummyClassifier(strategy="most_frequent")
baseline.fit(X_train, y_train)
baseline_predictions = baseline.predict(X_test)
baseline_accuracy = accuracy_score(y_test, baseline_predictions)

print(f"Baseline accuracy: {baseline_accuracy:.3f}")
ConfusionMatrixDisplay.from_predictions(
    y_test,
    baseline_predictions,
    display_labels=wine.target_names,
)
plt.show()

What does this code do?

  • The fit call learns which class appears most often in the training labels.
  • The predict call produces one prediction for every held-out sample.
  • The accuracy_score function calculates the fraction of test labels predicted correctly.
  • The ConfusionMatrixDisplay.from_predictions call shows how the predicted classes compare with the known classes.
  • Save wine_classifier.ipynb using the save icon in the notebook toolbar.

Before you run Cell 3, where do you think its predictions will collect in the confusion matrix? Your prediction gives you something concrete to test.

  • Run Cell 3 by pressing Shift+Enter.

There it is. You'll see a baseline accuracy followed by a three-class confusion matrix.

Every prediction appears in one class column. This intended shortfall proves that the baseline cannot respond to any of the 13 measurements.

Missing the accuracy or confusion matrix?

  • Run Cells 1 through 3 in order if the notebook reports that a required name is unavailable.
  • Check that plt.show() is the final line in Cell 3 if the confusion matrix does not appear.
  • Check that Cell 3 uses X_test for predictions plus y_test for the accuracy and confusion matrix.

Help me debug my baseline output.

✔️ Awesome, I've got everything!

Great. Your complete baseline workflow is saved in wine_classifier.ipynb.

ⓧ I'd like to double check the full code

# Cell 1: imports and data
import matplotlib.pyplot as plt
from sklearn.datasets import load_wine
from sklearn.dummy import DummyClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, ConfusionMatrixDisplay
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

wine = load_wine()
X = wine.data
y = wine.target

# Cell 2: one reproducible held-out split
X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.25,
    random_state=42,
    stratify=y,
)

# Cell 3: naive baseline
baseline = DummyClassifier(strategy="most_frequent")
baseline.fit(X_train, y_train)
baseline_predictions = baseline.predict(X_test)
baseline_accuracy = accuracy_score(y_test, baseline_predictions)

print(f"Baseline accuracy: {baseline_accuracy:.3f}")
ConfusionMatrixDisplay.from_predictions(
    y_test,
    baseline_predictions,
    display_labels=wine.target_names,
)
plt.show()

Your naive baseline now exposes the cost of ignoring the measurements. Next, you'll inspect the evidence that this classifier missed.

Inspect the Wine Data

Your baseline proved that an accuracy score can exist even when a model ignores every measurement. Its confusion matrix packed every prediction into one class column.

In this step, you'll expose the class balance in the dataset. You'll compare all 13 feature ranges before choosing feature scaling.

In this step, get ready to:
  • Print the sample count and feature count.
  • Count the wines in each cultivar class.
  • Compare the minimum and maximum of every feature.
Print the dataset evidence

Each row in X represents one wine. Each column stores one chemical measurement.

The labels in y identify the cultivar class for each wine. Counting those labels reveals how the three classes are distributed.

  • Click the empty notebook cell directly below the baseline output.
  • Fill the selected cell with the dataset inspection by pasting this code:
# Cell 4: inspect classes and feature scales
print(f"Samples: {X.shape[0]}")
print(f"Features: {X.shape[1]}")
print("\nClass distribution:")
for class_id, class_name in enumerate(wine.target_names):
    class_count = (y == class_id).sum()
    print(f"{class_name}: {class_count}")

print("\nFeature ranges:")
for feature_name, minimum, maximum in zip(
    wine.feature_names,
    X.min(axis=0),
    X.max(axis=0),
):
    print(f"{feature_name:30} {minimum:8.2f} to {maximum:8.2f}")

What does this code inspect?

  • The expression X.shape[0] reports the number of wine samples.
  • The expression X.shape[1] reports the number of chemical features.
  • The function enumerate(wine.target_names) pairs each class name with its numeric label.
  • The expression (y == class_id).sum() counts the labels belonging to one class.
  • The function zip() pairs every feature name with its minimum and maximum.
  • The argument axis=0 calculates those endpoints separately for each feature column.

Before you run this cell, do you expect the 13 measurements to occupy similar numeric ranges?

  • Execute the selected cell by pressing Shift+Enter.
  • Save wine_classifier.ipynb by pressing Ctrl+S.

You'll see Samples: 178 plus Features: 13 at the top of the output. Below them, you'll see three named class rows with counts of 59, 71, and 48.

The Feature ranges: section should contain one minimum-to-maximum range for each of the 13 features.

Missing Part of the Inspection Output?

  • Check that the selected cell begins with # Cell 4: inspect classes and feature scales.
  • Check that both loops are indented with four spaces.
  • Run the cell again with Shift+Enter.

Help me fix my Cell 4 dataset inspection output.

✔️ Awesome, I've got everything!

That inspection is working. Cell 4 now exposes the class balance and feature scales that the baseline ignored.

ⓧ I'd like to double check the full code

# Cell 1: imports and data
import matplotlib.pyplot as plt
from sklearn.datasets import load_wine
from sklearn.dummy import DummyClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, ConfusionMatrixDisplay
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

wine = load_wine()
X = wine.data
y = wine.target

# Cell 2: one reproducible held-out split
X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.25,
    random_state=42,
    stratify=y,
)

# Cell 3: naive baseline
baseline = DummyClassifier(strategy="most_frequent")
baseline.fit(X_train, y_train)
baseline_predictions = baseline.predict(X_test)
baseline_accuracy = accuracy_score(y_test, baseline_predictions)

print(f"Baseline accuracy: {baseline_accuracy:.3f}")
ConfusionMatrixDisplay.from_predictions(
    y_test,
    baseline_predictions,
    display_labels=wine.target_names,
)
plt.show()

# Cell 4: inspect classes and feature scales
print(f"Samples: {X.shape[0]}")
print(f"Features: {X.shape[1]}")
print("\nClass distribution:")
for class_id, class_name in enumerate(wine.target_names):
    class_count = (y == class_id).sum()
    print(f"{class_name}: {class_count}")

print("\nFeature ranges:")
for feature_name, minimum, maximum in zip(
    wine.feature_names,
    X.min(axis=0),
    X.max(axis=0),
):
    print(f"{feature_name:30} {minimum:8.2f} to {maximum:8.2f}")
Interpret the feature ranges

A model receives all 13 measurements together. Large differences between their raw ranges can cause bigger-valued features to dominate the calculation.

Before you check, which feature do you think has the widest printed range?

  • Count the rows under Feature ranges:.
  • Compare the minimum with the maximum for each feature.
  • Identify the feature with the widest printed range.

You should count 13 feature rows. Their minimum and maximum values occupy noticeably different numeric scales.

This confirms why the next model needs a consistent preprocessing step. The evidence now supports scaling the features before classification.

You have shown exactly what the baseline ignored: three classes supported by 13 differently scaled measurements. Next, you'll turn that evidence into a scaled classification pipeline.

Train a Scaled Pipeline

Your notebook has exposed the evidence that the baseline ignored. The 13 chemical measurements use very different numeric ranges.

A pipeline applies feature scaling using the training data. It passes the scaled measurements to logistic regression without leaking test-set statistics into training.

In this step, get ready to:
  • Combine feature scaling with logistic regression in one pipeline.
  • Fit the pipeline using the training samples.
  • Compare its held-out results with the baseline.
Build the scaled pipeline

The scaler learns how to transform each feature from the training set. The classifier then learns from those transformed measurements.

  • Switch back to wine_classifier.ipynb in Jupyter Notebook.
  • Add a new code cell directly below Cell 4 using the notebook's cell controls.
  • Build the pipeline in the new cell by copying this code:
# Cell 5: scaled logistic regression pipeline
model = make_pipeline(
    StandardScaler(),
    LogisticRegression(max_iter=1000),
)
model.fit(X_train, y_train)
model_predictions = model.predict(X_test)
model_accuracy = accuracy_score(y_test, model_predictions)

print(f"Baseline accuracy: {baseline_accuracy:.3f}")
print(f"Pipeline accuracy: {model_accuracy:.3f}")
ConfusionMatrixDisplay.from_predictions(
    y_test,
    model_predictions,
    display_labels=wine.target_names,
)
plt.show()

What Does This Code Do?

  • The make_pipeline call keeps scaling and classification in one ordered workflow.
  • The StandardScaler step learns its statistics when model.fit receives X_train.
  • The LogisticRegression step uses the scaled training measurements to learn the three cultivar classes.
  • The same y_test labels measure both models. This keeps the accuracy comparison fair.
  • The final display maps every held-out prediction against its known class.
  • Save wine_classifier.ipynb using the notebook's save control.
Compare the pipeline with the baseline

Before you run the cell, do you expect a model using all 13 measurements to outperform one that always chooses the most frequent class?

  • Run the new cell by pressing Shift-Enter.

You'll see values beside Baseline accuracy: and Pipeline accuracy:. A three-class confusion matrix appears below them.

  • Compare the two printed accuracy values.
  • Trace the matrix columns to confirm the pipeline predicts all three cultivar classes.
  • Check the main diagonal to find the correctly classified held-out samples.

You should see a higher pipeline accuracy than the baseline accuracy. You should also see predictions across all three class columns instead of the baseline's single column.

That is the key improvement working. Your classifier now responds to the chemical evidence in each held-out wine sample.

Cell Does Not Run?

  • Run Cells 1 through 4 again if the notebook kernel restarted.
  • Check that Cell 1 defines StandardScaler before Cell 5 uses it.
  • Check that Cell 2 defines X_train and X_test before the pipeline runs.

Help me debug my scaled pipeline cell.

✔️ Awesome, I've got everything!

Your wine_classifier.ipynb notebook now contains the fitted pipeline. It also shows both accuracy values and the new confusion matrix.

ⓧ I'd like to double check the full code

# Cell 1: imports and data
import matplotlib.pyplot as plt
from sklearn.datasets import load_wine
from sklearn.dummy import DummyClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, ConfusionMatrixDisplay
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

wine = load_wine()
X = wine.data
y = wine.target

# Cell 2: one reproducible held-out split
X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.25,
    random_state=42,
    stratify=y,
)

# Cell 3: naive baseline
baseline = DummyClassifier(strategy="most_frequent")
baseline.fit(X_train, y_train)
baseline_predictions = baseline.predict(X_test)
baseline_accuracy = accuracy_score(y_test, baseline_predictions)

print(f"Baseline accuracy: {baseline_accuracy:.3f}")
ConfusionMatrixDisplay.from_predictions(
    y_test,
    baseline_predictions,
    display_labels=wine.target_names,
)
plt.show()

# Cell 4: inspect classes and feature scales
print(f"Samples: {X.shape[0]}")
print(f"Features: {X.shape[1]}")
print("\nClass distribution:")
for class_id, class_name in enumerate(wine.target_names):
    class_count = (y == class_id).sum()
    print(f"{class_name}: {class_count}")

print("\nFeature ranges:")
for feature_name, minimum, maximum in zip(
    wine.feature_names,
    X.min(axis=0),
    X.max(axis=0),
):
    print(f"{feature_name:30} {minimum:8.2f} to {maximum:8.2f}")

# Cell 5: scaled logistic regression pipeline
model = make_pipeline(
    StandardScaler(),
    LogisticRegression(max_iter=1000),
)
model.fit(X_train, y_train)
model_predictions = model.predict(X_test)
model_accuracy = accuracy_score(y_test, model_predictions)

print(f"Baseline accuracy: {baseline_accuracy:.3f}")
print(f"Pipeline accuracy: {model_accuracy:.3f}")
ConfusionMatrixDisplay.from_predictions(
    y_test,
    model_predictions,
    display_labels=wine.target_names,
)
plt.show()

Your scaled pipeline now beats the evidence-blind baseline on the same held-out samples. Next up, you'll explain one prediction using its class probabilities.

Explain an Unseen Prediction

Your fitted scikit-learn pipeline now uses all 13 measurements to classify held-out wines.

An aggregate accuracy score cannot explain what happened to a specific wine. A useful final demo needs one concrete prediction with its known answer.

You will inspect the model's probability estimates for all three cultivar classes. This shows how the model distributes its confidence for one sample it never saw during training.

In this step, get ready to:
  • Select one held-out wine for a fresh prediction.
  • Compare the predicted cultivar with the known cultivar.
  • Display the probability estimate for every class.
Select one held-out sample

Your test set contains wines that were excluded from model training. The first row gives you a concrete example for checking the pipeline's decision.

  • Add a code cell directly below Cell 5 from the Jupyter Notebook toolbar.
  • Paste this code into the new cell:
# Cell 6: explain one held-out prediction
sample = X_test[[0]]
predicted_class = model.predict(sample)[0]
actual_class = y_test[0]
probabilities = model.predict_proba(sample)[0]

print(f"Predicted class: {wine.target_names[predicted_class]}")
print(f"Actual class:    {wine.target_names[actual_class]}")
print("Class probabilities:")
for class_name, probability in zip(wine.target_names, probabilities):
    print(f"  {class_name}: {probability:.3f}")

What does this cell do?

  • The X_test[[0]] expression selects the first held-out wine while keeping the row shape expected by the pipeline.
  • The model.predict(sample) expression chooses the most likely numeric class for that wine.
  • The y_test[0] expression retrieves the known class for the same held-out wine.
  • The model.predict_proba(sample) expression calculates one probability estimate for each cultivar class.
  • The final loop pairs each value with its matching name from wine.target_names.
Run and interpret the prediction

The predicted class shows the pipeline's decision. The actual class reveals whether that decision matches the known label.

  • Save wine_classifier.ipynb.

Before you run the cell, do you expect the predicted class to match the actual class?

  • Execute Cell 6 by pressing Shift+Enter.

You'll see the predicted cultivar followed by the actual cultivar. You'll also see one probability line for each of the three classes.

You have turned a headline model score into an explainable result for one unseen wine.

Don't see all three probabilities?

Run Cell 5 again to restore the fitted model before executing Cell 6.

Check that the sample uses X_test[[0]] with two pairs of square brackets. Confirm that the probability line uses model.predict_proba(sample)[0].

Help me debug my held-out prediction cell.

✔️ Awesome, I've got everything!

Your saved notebook now explains one held-out prediction with its known label and three class probability estimates.

ⓧ I'd like to double check the full code

# Cell 1: imports and data
import matplotlib.pyplot as plt
from sklearn.datasets import load_wine
from sklearn.dummy import DummyClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, ConfusionMatrixDisplay
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

wine = load_wine()
X = wine.data
y = wine.target

# Cell 2: one reproducible held-out split
X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.25,
    random_state=42,
    stratify=y,
)

# Cell 3: naive baseline
baseline = DummyClassifier(strategy="most_frequent")
baseline.fit(X_train, y_train)
baseline_predictions = baseline.predict(X_test)
baseline_accuracy = accuracy_score(y_test, baseline_predictions)

print(f"Baseline accuracy: {baseline_accuracy:.3f}")
ConfusionMatrixDisplay.from_predictions(
    y_test,
    baseline_predictions,
    display_labels=wine.target_names,
)
plt.show()

# Cell 4: inspect classes and feature scales
print(f"Samples: {X.shape[0]}")
print(f"Features: {X.shape[1]}")
print("\nClass distribution:")
for class_id, class_name in enumerate(wine.target_names):
    class_count = (y == class_id).sum()
    print(f"{class_name}: {class_count}")

print("\nFeature ranges:")
for feature_name, minimum, maximum in zip(
    wine.feature_names,
    X.min(axis=0),
    X.max(axis=0),
):
    print(f"{feature_name:30} {minimum:8.2f} to {maximum:8.2f}")

# Cell 5: scaled logistic regression pipeline
model = make_pipeline(
    StandardScaler(),
    LogisticRegression(max_iter=1000),
)
model.fit(X_train, y_train)
model_predictions = model.predict(X_test)
model_accuracy = accuracy_score(y_test, model_predictions)

print(f"Baseline accuracy: {baseline_accuracy:.3f}")
print(f"Pipeline accuracy: {model_accuracy:.3f}")
ConfusionMatrixDisplay.from_predictions(
    y_test,
    model_predictions,
    display_labels=wine.target_names,
)
plt.show()

# Cell 6: explain one held-out prediction
sample = X_test[[0]]
predicted_class = model.predict(sample)[0]
actual_class = y_test[0]
probabilities = model.predict_proba(sample)[0]

print(f"Predicted class: {wine.target_names[predicted_class]}")
print(f"Actual class:    {wine.target_names[actual_class]}")
print("Class probabilities:")
for class_name, probability in zip(wine.target_names, probabilities):
    print(f"  {class_name}: {probability:.3f}")

Your notebook now demonstrates the full path from a weak baseline to an evidence-based prediction for an unseen wine.

Secret mission

Check Whether One Split Was Lucky

Use five-fold cross-validation to evaluate the complete pipeline across five partitions. Review the five scores. Use their mean to summarize performance while their standard deviation shows the spread.

Clean Up Your Resources

Clean Up Your Resources

Everything in this project runs locally, so there are no ongoing costs or cloud resources to manage. Choose whether to keep your work available, pause the running server, or delete the notebook entirely.

Resources you used:

  • The running Jupyter Notebook server in PowerShell.
  • The wine_classifier.ipynb file containing your classifier and saved results.

Keep everything running

No action needed. Choose this if you want to keep experimenting with your wine classifier.

  • Leave the Jupyter Notebook server running in PowerShell.
  • Keep wine_classifier.ipynb in its current folder.
  • Leave the pinned Python packages installed for future notebooks.
  • Continue rerunning cells or testing new wine samples whenever you are ready.

Pause - I'll come back to this later

Shut down the notebook server to free its running process. Your notebook stays on your computer with its saved cells and outputs.

  • Save the open notebook by pressing Ctrl+S.
  • Switch back to the PowerShell window running Jupyter Notebook.
  • Press Ctrl+C.
  • Follow the shutdown prompt if PowerShell asks for confirmation.

That frees the running process while preserving every cell you built. Your wine_classifier.ipynb file remains ready for your return.

  • Start Jupyter Notebook again from PowerShell with the launch command you used in Step 1.

Delete - I don't want to use this again

Remove the listed project resources if you want to start fresh. This permanently deletes only wine_classifier.ipynb. Your shared Python installation and pinned packages remain available.

  • Switch back to the PowerShell window running Jupyter Notebook.
  • Press Ctrl+C.
  • Follow the shutdown prompt if PowerShell asks for confirmation.

The local notebook server is now stopped. You can remove its saved project file next.

  • Press the Windows key to open the Windows search bar.
  • Type File Explorer into the search field.
  • Press Enter.
  • Navigate to the folder where you created wine_classifier.ipynb in Step 1.
  • Select wine_classifier.ipynb.
  • Press Shift+Delete to delete the file permanently.
  • Approve the permanent deletion if Windows asks for confirmation.

That completes the local cleanup. You should no longer see wine_classifier.ipynb in the folder.

Nice Work!

Nice Work!

You did it! Your local Jupyter Notebook now demonstrates a complete multiclass classification workflow from a deliberately weak baseline to an explained wine cultivar prediction.

You've learned how to:

  • Established a most-frequent baseline that exposed why ignoring all 13 chemical measurements sends every prediction into one cultivar class.
  • Built a scikit-learn pipeline that scales every measurement before fitting logistic regression.
  • Evaluated unseen wines with accuracy plus a three-class confusion matrix. Explained one held-out prediction with class probabilities for every cultivar.
  • Secret Mission: Reported five fold scores with their mean plus standard deviation to test whether one split was unusually lucky using five-fold cross-validation.

Ready to quiz yourself?