Build an AI Ticket Triage Dashboard

Build a local dashboard that classifies support tickets with a trained ML model.

Introduction

30 Second Summary

Two customers can describe the same problem with completely different words. A person sees the connection immediately, while a rigid word list can send one message to the wrong place.

In this project, you will build a local support dashboard that routes customer messages into billing, account, or technical categories. The page compares a keyword baseline with a machine learning classifier trained on labeled examples.

What You'll Build

Your finished dashboard lets you paste a ticket into the browser to see a weak keyword result beside the trained model's category, confidence, probabilities, and held-out evidence.

By the end of this project, you'll have:

  • One-click ticket routing in your browser for billing, account, or technical requests.
  • A side-by-side comparison that shows the keyword baseline returning Unknown while the trained model still supplies a category.
  • Evaluation evidence you can inspect through held-out accuracy, individual test predictions, and the full probability breakdown.
  • Secret Mission: A human-review route that catches predictions below 60 percent confidence instead of forcing a category.

Are there any prerequisites?

You need a Mac with Visual Studio Code plus Python 3.11 or newer. Basic Python skills are enough because you do not need prior machine learning experience, an API key, or a cloud account.

Before We Start

This opening step locks in the portfolio story that will guide every later choice. You are building a local AI dashboard that classifies support messages as billing, account, or technical. It compares brittle keyword rules with a trained model to show how labeled examples improve routing.

Prepare a Reproducible Python Workspace

Your support ticket dashboard depends on a predictable Python runtime. Package conflicts can turn a strong demo into a setup problem for anyone reviewing it.

You will verify Python compatibility before creating an isolated virtual environment in Visual Studio Code. A pinned requirements.txt file will reproduce the same package setup.

In this step, get ready to:
  • Verify that your Mac runs Python 3.11 or newer.
  • Create the ai-ticket-triage folder with an isolated environment.
  • Install the pinned dependencies from requirements.txt.
  • Prove the installation by launching the official Hello app.
Check your Python version

The classifier package requires Python 3.11 or newer. Checking first prevents an incompatible interpreter from reaching the package installation.

  • Press Cmd+Space to open Spotlight.
  • Type Terminal into Spotlight.
  • Press Enter to open Terminal.
  • Check the installed Python version by running this command:
python3 --version

What does this command do?

The python3 command selects the Python 3 interpreter available in your terminal. The --version option asks it to print its version.

Your terminal output places your Mac in one of these three states.

✔️ I see version 3.11 or newer

Good start. Your Python installation meets the project requirement.

ⓧ I see an older version

Your current interpreter is too old for scikit-learn 1.9.1. The latest stable macOS release is Python 3.14.8.

  • Open the official Python downloads page in your browser.
  • Download the signed universal2 macOS installer for Python 3.14.8.
  • Double-click the downloaded .pkg file.
  • Complete the standard installation steps.

The installer adds the newer interpreter to your shell path. Reopening Terminal allows that updated path to take effect.

  • Quit Terminal.
  • Press Cmd+Space to open Spotlight.
  • Type Terminal into Spotlight.
  • Press Enter to reopen Terminal.
  • Run the version check command above again.

macOS can keep resolving its older system interpreter after a new installation. The Python installer includes a shell-profile updater for this situation.

  • Open Finder through Spotlight if the version remains older than 3.11.
  • Open the /Applications/Python 3.14/ folder.
  • Double-click Update Shell Profile.command.
  • Reopen Terminal through Spotlight.
  • Run the version check command above again.

ⓧ Command not found

Your shell cannot find a Python 3 installation. The official macOS installer adds one to your Mac.

  • Open the official Python downloads page in your browser.
  • Download the signed universal2 macOS installer for Python 3.14.8.
  • Double-click the downloaded .pkg file.
  • Complete the standard installation steps.

A fresh terminal session loads the path created by the installer.

  • Quit Terminal.
  • Press Cmd+Space to open Spotlight.
  • Type Terminal into Spotlight.
  • Press Enter to reopen Terminal.
  • Run the version check command above again.
Create the project workspace

A dedicated folder keeps every project file in one place. Visual Studio Code can use that folder as a workspace with its terminal pointed at the correct location.

  • Move to your Desktop by running this command:
cd ~/Desktop

Why start on the Desktop?

The cd command changes the terminal location to your Desktop. This gives the new project folder a predictable location.

  • Create the project folder by running this command:
mkdir ai-ticket-triage

What does this command do?

The mkdir command creates the ai-ticket-triage folder on your Desktop. This folder becomes the home for the dashboard.

  • Confirm the folder exists by running this command:
ls

What should I see?

The ls command lists the contents of your Desktop. You should see ai-ticket-triage in the output.

  • Press Cmd+Space to open Spotlight.
  • Type Visual Studio Code into Spotlight.
  • Press Enter to open Visual Studio Code.
  • Click File in the menu bar.
  • Click Open Folder....
  • Select the ai-ticket-triage folder on your Desktop.
  • Click Open.
  • Approve the Workspace Trust prompt if it appears.

You will see ai-ticket-triage at the top of the Explorer view. That confirms Visual Studio Code has opened the correct workspace.

  • Press Ctrl+` to open the integrated terminal.
  • Create the isolated environment inside ai-ticket-triage by running this command:
python3 -m venv .venv

What does this command do?

The venv module creates a private Python environment inside .venv. Packages installed there stay separate from your other Python projects.

The .venv folder now appears in the Explorer view. That visible folder confirms the environment was created.

  • Activate the environment by running this command:
source .venv/bin/activate

What does activation change?

The source command updates your current shell session. From this point, python points to the interpreter inside .venv.

Before you check, predict whether the active environment will report a compatible Python version.

  • Verify the active environment by running this command:
python --version

What does this check prove?

This command checks the interpreter selected by the active environment. A version of 3.11 or newer proves that the project environment uses a compatible Python release.

Your output shows Python 3.11 or newer. That is the isolated runtime your dashboard will use.

Install and test the pinned packages

A requirements file records the exact package versions for a project. Version pins make a future installation follow the same setup.

  • Click the new-file icon beside ai-ticket-triage in the Explorer view.
  • Type requirements.txt as the file name.
  • Press Enter to create the file.

The empty requirements.txt file opens in the editor. Its name also appears in the Explorer view.

  • Add the verified package pins to requirements.txt by pasting this content:
scikit-learn==1.9.1
streamlit==1.65.0

What do these pins install?

  • The scikit-learn==1.9.1 line installs the machine learning package used by the classifier.
  • The streamlit==1.65.0 line installs Streamlit for the browser-based dashboard.
  • The double equals signs pin each package to one exact version.
  • Save requirements.txt by pressing Cmd+S on macOS or Ctrl+S on Windows.

The saved file shows exactly two package lines. Each line contains its verified version pin.

✔️ Awesome, I've got everything!

Great. Double check that requirements.txt is saved before installing the packages.

ⓧ I'd like to double check the full code

Compare your complete requirements.txt file with this version line by line.

scikit-learn==1.9.1
streamlit==1.65.0

How should I compare this file?

Check both package names for spelling. Confirm that each version number matches exactly.

  • Install the pinned packages into the active environment by running this command:
python -m pip install -r requirements.txt

What does this command do?

The python -m pip portion runs the package installer from your active environment. The -r option tells it to read every pin from requirements.txt.

The packages are installed inside .venv. Your system-wide Python packages remain separate.

The command finishes with both package names in the completed installation output. Your terminal then returns to the active environment prompt.

Package installation failed?

  • Confirm that the environment name still appears at the start of your terminal prompt.
  • Run the activation command above again if the environment is inactive.
  • Check that requirements.txt contains the two package pins exactly as shown.

Still stuck? Help me diagnose why my pinned Python packages will not install inside this virtual environment.

The Hello command starts a local server in the integrated terminal. You can return to the prompt with Ctrl+C after capturing the checkpoint.

Before you run this, predict whether the installation test will reach your browser.

  • Launch the official installation test by running this command:
python -m streamlit hello

What does this command prove?

The command runs Streamlit through the Python interpreter inside .venv. Reaching the Hello page proves the package can start a local web app.

Your browser opens the Streamlit Hello app. You can see the interactive example page served from your Mac.

Strong start. Your isolated workspace can now run the same framework that will power the ticket dashboard.

Hello page did not open?

  • Check the terminal for a local address if the browser stayed closed.
  • Open that local address in your browser.
  • Confirm that the active environment name still appears in the terminal prompt before rerunning the command.

Need another pair of eyes? Help me work out why the Streamlit Hello app is not opening from my active virtual environment.

  • Return to the integrated terminal after capturing your screenshot.
  • Press Ctrl+C to stop the Hello server.

Your Python workspace is isolated, version-checked, and ready to run Streamlit. Next up, you will build a keyword-based triage app that exposes the limits of fixed rules.

Build a Rule-Based Triage App

Your reproducible Python workspace is ready. Now you can turn it into the first working slice of your support triage product.

A rule-based baseline scans customer messages for fixed keywords. Paraphrased language tests the central weakness of this approach.

Streamlit gives the baseline a local browser interface. The visible result makes its behavior easy to inspect.

In this step, get ready to:
  • Define a baseline that routes literal keyword matches.
  • Build a Streamlit interface for ticket classification.
  • Test the default paraphrased ticket in the browser.
Create the keyword baseline

The KEYWORDS mapping stores literal terms for each support category. The keyword_baseline() function uses those terms to choose a route.

The DEMO_TICKET value gives the dashboard one consistent customer message to test. Its wording describes a loading problem without naming the problem directly.

  • Select the Explorer view in Visual Studio Code's left Activity Bar.
  • Select the New File... button in the Explorer view.
  • Enter triage.py as the file name.
  • Press Enter to create the file.
  • Define the baseline module by pasting the code below into triage.py:
KEYWORDS = {
    "billing": ["charge", "charged", "invoice", "payment", "refund", "price"],
    "account": ["account", "password", "locked", "sign in", "verification"],
    "technical": ["crash", "error", "upload", "slow", "blank", "freeze"],
}

DEMO_TICKET = "The dashboard keeps spinning and never finishes loading."


def keyword_baseline(ticket):
    normalized = ticket.lower()
    for category, words in KEYWORDS.items():
        if any(word in normalized for word in words):
            return category
    return "Unknown"

What does this code do?

  • The KEYWORDS dictionary groups literal words under the billing, account, or technical category.
  • The DEMO_TICKET variable holds the default message that appears in the dashboard.
  • The ticket.lower() call normalizes capitalization before matching begins.
  • The loop checks whether any category keyword appears inside the normalized ticket.
  • The function returns Unknown when no keyword matches.
  • Save triage.py by pressing Cmd+S.
  • Confirm triage.py now appears in the Explorer view beside requirements.txt.

Baseline file missing or incomplete?

  • Confirm triage.py is saved inside the ai-ticket-triage folder.
  • Compare every quote with the reference below.
  • Compare every bracket with the reference below.

Still stuck? Help me compare my triage.py file with the expected keyword baseline.

✔️ Awesome, I've got everything!

Your baseline module is saved in the project folder.

ⓧ I'd like to double check the full code

This is the complete triage.py file for this step.

KEYWORDS = {
    "billing": ["charge", "charged", "invoice", "payment", "refund", "price"],
    "account": ["account", "password", "locked", "sign in", "verification"],
    "technical": ["crash", "error", "upload", "slow", "blank", "freeze"],
}

DEMO_TICKET = "The dashboard keeps spinning and never finishes loading."


def keyword_baseline(ticket):
    normalized = ticket.lower()
    for category, words in KEYWORDS.items():
        if any(word in normalized for word in words):
            return category
    return "Unknown"
Build the Streamlit interface

The dashboard collects a support ticket through a text area. A button passes that message to keyword_baseline() and displays the returned route.

  • Select the New File... button in the Explorer view.
  • Enter app.py as the file name.
  • Press Enter to create the file.
  • Build the dashboard by pasting the code below into app.py:
import streamlit as st

from triage import DEMO_TICKET, keyword_baseline

st.title("AI Support Ticket Triage")
st.write("Compare a brittle keyword baseline with a trained text classifier.")

ticket = st.text_area("Support ticket", value=DEMO_TICKET)

if st.button("Classify ticket", type="primary"):
    if ticket.strip():
        st.write("Keyword baseline:", keyword_baseline(ticket))
    else:
        st.write("Enter a support ticket before classifying it.")

What does this code do?

  • The import statements load Streamlit plus the two project values from triage.py.
  • The st.title() call names the dashboard.
  • The st.write() call explains what the comparison demonstrates.
  • The st.text_area() call preloads the default customer message.
  • The st.button() call classifies the current ticket after a click.
  • The inner condition asks for a support ticket when the text area is empty.
  • Save app.py by pressing Cmd+S.
  • Confirm app.py now appears in the Explorer view beside triage.py.

Dashboard code showing a warning?

  • Confirm the import uses the exact file name triage.
  • Confirm DEMO_TICKET matches the capitalization in triage.py.
  • Confirm keyword_baseline matches the function name in triage.py.

Need another pair of eyes? Help me find the mismatch between app.py and triage.py.

✔️ Awesome, I've got everything!

Your dashboard file is saved beside the baseline module.

ⓧ I'd like to double check the full code

This is the complete app.py file for this step.

import streamlit as st

from triage import DEMO_TICKET, keyword_baseline

st.title("AI Support Ticket Triage")
st.write("Compare a brittle keyword baseline with a trained text classifier.")

ticket = st.text_area("Support ticket", value=DEMO_TICKET)

if st.button("Classify ticket", type="primary"):
    if ticket.strip():
        st.write("Keyword baseline:", keyword_baseline(ticket))
    else:
        st.write("Enter a support ticket before classifying it.")
Test the paraphrased ticket

The Streamlit runner serves app.py as a local web app. The terminal remains occupied while the dashboard is running.

Before you run the app, consider whether fixed vocabulary can route a sentence about endless spinning.

  • Launch the dashboard from the activated terminal by running this command:
python -m streamlit run app.py

What does this command do?

This command starts the Streamlit server with app.py as its entry point. Your browser opens the local dashboard served by that process.

  • Wait for the browser to show the AI Support Ticket Triage heading.
  • Leave the default support ticket unchanged.
  • Click Classify ticket.

You'll see Keyword baseline: Unknown beneath the button.

The interface works correctly. The Unknown result is the intended shortfall because none of the technical keywords appears in the paraphrased ticket.

Dashboard not loading?

  • Confirm the .venv environment remains activated in the terminal.
  • Confirm the terminal is still inside the ai-ticket-triage folder.
  • Open the local address printed in the terminal if the browser did not open automatically.

Still blocked? Help me debug why my Streamlit dashboard is not loading or showing Unknown.

Your first dashboard is live. It has exposed the exact limit of literal keyword matching.

Your dashboard now gives you a concrete baseline to improve. Next, you'll train a text classifier that learns patterns from labeled support tickets.

Train and Evaluate the Text Classifier

Your working dashboard exposed the keyword baseline's central weakness. It only recognizes the fixed terms stored in KEYWORDS.

You will now use scikit-learn to learn patterns from labeled training data. A reproducible train/test split measures performance on tickets that the model did not use for training.

In this step, get ready to:
  • Add 36 balanced support tickets with known categories.
  • Build a two-stage pipeline from TF-IDF features plus logistic regression.
  • Evaluate the classifier on nine held-out tickets.
Add balanced labeled tickets

Each training example pairs a customer message with its correct category. Equal numbers of billing, account, and technical examples keep the small dataset balanced.

  • In triage.py, place the model imports above KEYWORDS by copying this block:
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline

What do these imports provide?

  • TfidfVectorizer converts ticket language into weighted numeric features.
  • LogisticRegression learns how those features relate to the three categories.
  • train_test_split reserves examples for evaluation.
  • Pipeline keeps feature extraction and classification in one reusable object.
  • Save triage.py.
  • Check that the imports load by running this command in the terminal from earlier:
python triage.py

What does this check prove?

The terminal returns to its prompt without a traceback. That confirms the active environment can import the machine learning tools.

Does Python report a missing package?

Confirm that the terminal prompt still shows the active .venv environment. The imports depend on the scikit-learn installation inside that environment.

If the environment is active but the import still fails, help me diagnose the missing scikit-learn import.

  • In triage.py, place TRAINING_DATA between KEYWORDS and DEMO_TICKET by copying these first two category groups:
TRAINING_DATA = [
    ("I was charged twice for this month", "billing"),
    ("Why did my subscription price increase", "billing"),
    ("I need a refund for my last payment", "billing"),
    ("My card was declined during checkout", "billing"),
    ("The invoice total is wrong", "billing"),
    ("I do not recognize this charge", "billing"),
    ("Please update the card on my subscription", "billing"),
    ("I was billed after cancelling", "billing"),
    ("Where can I download my receipt", "billing"),
    ("My payment is still pending", "billing"),
    ("The discount was not applied to my order", "billing"),
    ("Can I change from annual to monthly billing", "billing"),
    ("I forgot my password", "account"),
    ("Please help me reset my password", "account"),
    ("My account is locked", "account"),
    ("I cannot sign in with my email", "account"),
    ("I need to change the email on my account", "account"),
    ("How do I delete my account", "account"),
    ("The verification code never arrived", "account"),
    ("I think someone accessed my account", "account"),
    ("Two factor authentication is not working", "account"),
    ("My profile name is incorrect", "account"),
    ("I want to reactivate my account", "account"),
    ("I cannot complete account verification", "account"),
]

How are the examples labeled?

Each tuple contains one ticket followed by its expected category. These first 24 rows provide 12 billing examples plus 12 account examples.

The wording varies within each category. That variation gives the model more than one phrase to learn from.

  • Save triage.py.

You should now see TRAINING_DATA between the keyword mapping and the demo ticket.

  • In TRAINING_DATA, find the closing square bracket.
  • Insert the technical examples immediately above that bracket by copying this block:
    ("The app crashes when I open it", "technical"),
    ("The page keeps loading after I log in", "technical"),
    ("The dashboard spinner never stops", "technical"),
    ("The upload button does nothing", "technical"),
    ("I see an error when I save my work", "technical"),
    ("The mobile app freezes on launch", "technical"),
    ("The report will not load", "technical"),
    ("Notifications are not appearing", "technical"),
    ("The screen is blank after the update", "technical"),
    ("The search feature returns no results", "technical"),
    ("The export gets stuck halfway through", "technical"),
    ("The website is extremely slow today", "technical"),

Why add several phrasings?

The technical examples describe crashes, loading failures, blank screens, and stuck operations. Their shared context gives the model several signals for the same category.

  • Save triage.py.
  • Confirm that all 36 rows form valid Python by running this command:
python triage.py

What does this data check prove?

The terminal returns to its prompt without a traceback. Python can now parse the complete balanced dataset.

Seeing a syntax problem in the dataset?

Check that every ticket row ends with a comma. Confirm that all technical examples sit above the final square bracket.

If the line number is unclear, help me find the syntax problem in TRAINING_DATA.

Build the training and evaluation pipeline

A TF-IDF vectorizer gives more weight to useful words or short phrases. Logistic regression uses those weighted features to estimate a category.

Evaluation uses 25 percent of the examples as held-out data. The model trains on the other 75 percent before it predicts the nine reserved tickets.

  • In triage.py, add build_pipeline() below keyword_baseline() by copying this block:
def build_pipeline():
    return Pipeline(
        [
            ("tfidf", TfidfVectorizer(ngram_range=(1, 2))),
            (
                "classifier",
                LogisticRegression(l1_ratio=0.0, max_iter=1000),
            ),
        ]
    )

What does this pipeline do?

  • TfidfVectorizer(ngram_range=(1, 2)) learns from individual words plus two-word phrases.
  • LogisticRegression(l1_ratio=0.0, max_iter=1000) performs the three-category classification.
  • Pipeline applies both stages in the same order during training and prediction.
  • Save triage.py.
  • Check the pipeline definition by running this command:
python triage.py

What does this pipeline check prove?

The terminal returns to its prompt without a traceback. The pipeline constructor is valid Python.

Does the pipeline fail to parse?

Check that the vectorizer tuple closes before the classifier tuple begins. Confirm that the final parentheses close both Pipeline and return.

If the parentheses are difficult to match, help me compare my build_pipeline function with the expected structure.

  • Add the first part of train_and_evaluate() directly below build_pipeline() by copying this block:
def train_and_evaluate():
    texts = [text for text, _ in TRAINING_DATA]
    labels = [label for _, label in TRAINING_DATA]

    train_texts, test_texts, train_labels, test_labels = train_test_split(
        texts,
        labels,
        test_size=0.25,
        random_state=42,
        stratify=labels,
    )

    evaluation_model = build_pipeline()
    evaluation_model.fit(train_texts, train_labels)
    test_predictions = evaluation_model.predict(test_texts)
    accuracy = float(accuracy_score(test_labels, test_predictions))

How does the held-out test work?

  • texts and labels separate each message from its known category.
  • test_size=0.25 reserves nine of the 36 examples for evaluation.
  • random_state=42 makes the same examples land in the test set on repeated runs.
  • stratify=labels preserves the category balance across the split.
  • accuracy_score calculates the fraction of held-out predictions that match their labels.
  • Save triage.py.
  • Check the first evaluation stage by running this command:
python triage.py

What does this evaluation check prove?

The terminal returns to its prompt without a traceback. Python can parse the split, training, prediction, and accuracy logic.

  • In train_and_evaluate(), find the line that calculates accuracy.
  • Add the evaluation rows plus final training stage directly below that line:
    evaluation_rows = [
        {
            "ticket": text,
            "actual": actual,
            "predicted": str(predicted),
        }
        for text, actual, predicted in zip(
            test_texts,
            test_labels,
            test_predictions,
        )
    ]

    final_model = build_pipeline()
    final_model.fit(texts, labels)
    return final_model, accuracy, evaluation_rows

Why train a second model?

evaluation_rows preserves each held-out ticket beside its actual label and predicted label. This makes individual successes or mistakes inspectable.

final_model trains on all 36 examples after evaluation is complete. The returned accuracy still comes from the separate held-out test.

  • Save triage.py.
  • Confirm that the completed training function parses by running this command:
python triage.py

What does the completed check prove?

The terminal returns to its prompt without a traceback. The complete training function is ready to be called by the command-line demo.

Does the training function report a syntax problem?

Keep every line from evaluation_rows through the return statement indented inside train_and_evaluate().

If Python points at the list comprehension, help me check the evaluation_rows indentation and brackets.

Classify the demo ticket and inspect the results

The fitted pipeline exposes a predicted category plus one probability for each class. The highest probability becomes the model's confidence for its selected category.

  • Add classify_ticket() below train_and_evaluate() by copying this block:
def classify_ticket(model, ticket):
    prediction = str(model.predict([ticket])[0])
    probabilities = model.predict_proba([ticket])[0]
    scores = {
        str(label): round(float(probability), 3)
        for label, probability in zip(model.classes_, probabilities)
    }
    return prediction, scores[prediction], scores

How does classification work?

  • model.predict([ticket]) selects one learned category.
  • model.predict_proba([ticket]) returns the probability assigned to every category.
  • model.classes_ supplies the labels in the same order as those probabilities.
  • scores[prediction] retrieves the probability for the winning category.
  • Add the reproducible command-line demo below classify_ticket() by copying this block:
if __name__ == "__main__":
    trained_model, held_out_accuracy, held_out_rows = train_and_evaluate()
    predicted_category, confidence, class_scores = classify_ticket(
        trained_model,
        DEMO_TICKET,
    )

    print("Keyword baseline:", keyword_baseline(DEMO_TICKET))
    print("ML prediction:", predicted_category)
    print("Class probabilities:", class_scores)
    print("Held-out accuracy:", held_out_accuracy)
    print("Held-out predictions:", held_out_rows)

What does the demo reveal?

The main block trains the model before classifying DEMO_TICKET. It prints the weak baseline beside the learned prediction.

The final lines expose all three probabilities plus the held-out evidence. You can inspect the model's behavior instead of accepting one label blindly.

  • Save triage.py.

Before you run the finished script, which result do you expect to reveal more about the model: its winning category or the complete held-out predictions?

  • Run the finished classifier from the activated terminal with this command:
python triage.py

What should you see?

The output starts with Keyword baseline: Unknown. The next lines show an ML prediction plus three class probabilities.

You will also see a held-out accuracy value. The final list contains nine dictionaries with ticket, actual, and predicted fields.

That is the key leap complete. Your script now learns from examples and produces inspectable evidence about its performance.

Does the final script fail?

Use the final traceback line to identify which function needs attention. A problem before any printed results usually points to training or evaluation.

Confirm that classify_ticket() appears above the main block. Confirm that the main block remains flush with the left edge.

If the traceback is still unclear, help me debug the final triage.py run.

✔️ Awesome, I've got everything!

Your terminal now shows the failed keyword baseline beside the trained model's prediction. Double-check that triage.py is saved.

ⓧ I'd like to double check the full code

Compare these three project files with yours. Only triage.py changed during this step.

scikit-learn==1.9.1
streamlit==1.65.0

What should remain in requirements.txt?

The file still contains the two verified package pins from your workspace setup. This step does not add another dependency.

import streamlit as st

from triage import DEMO_TICKET, keyword_baseline

st.title("AI Support Ticket Triage")
st.write("Compare a brittle keyword baseline with a trained text classifier.")

ticket = st.text_area("Support ticket", value=DEMO_TICKET)

if st.button("Classify ticket", type="primary"):
    if ticket.strip():
        st.write("Keyword baseline:", keyword_baseline(ticket))
    else:
        st.write("Enter a support ticket before classifying it.")

Why is app.py unchanged?

The dashboard still runs the keyword baseline from the previous step. This step builds and verifies the model layer in triage.py.

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline

KEYWORDS = {
    "billing": ["charge", "charged", "invoice", "payment", "refund", "price"],
    "account": ["account", "password", "locked", "sign in", "verification"],
    "technical": ["crash", "error", "upload", "slow", "blank", "freeze"],
}

TRAINING_DATA = [
    ("I was charged twice for this month", "billing"),
    ("Why did my subscription price increase", "billing"),
    ("I need a refund for my last payment", "billing"),
    ("My card was declined during checkout", "billing"),
    ("The invoice total is wrong", "billing"),
    ("I do not recognize this charge", "billing"),
    ("Please update the card on my subscription", "billing"),
    ("I was billed after cancelling", "billing"),
    ("Where can I download my receipt", "billing"),
    ("My payment is still pending", "billing"),
    ("The discount was not applied to my order", "billing"),
    ("Can I change from annual to monthly billing", "billing"),
    ("I forgot my password", "account"),
    ("Please help me reset my password", "account"),
    ("My account is locked", "account"),
    ("I cannot sign in with my email", "account"),
    ("I need to change the email on my account", "account"),
    ("How do I delete my account", "account"),
    ("The verification code never arrived", "account"),
    ("I think someone accessed my account", "account"),
    ("Two factor authentication is not working", "account"),
    ("My profile name is incorrect", "account"),
    ("I want to reactivate my account", "account"),
    ("I cannot complete account verification", "account"),
    ("The app crashes when I open it", "technical"),
    ("The page keeps loading after I log in", "technical"),
    ("The dashboard spinner never stops", "technical"),
    ("The upload button does nothing", "technical"),
    ("I see an error when I save my work", "technical"),
    ("The mobile app freezes on launch", "technical"),
    ("The report will not load", "technical"),
    ("Notifications are not appearing", "technical"),
    ("The screen is blank after the update", "technical"),
    ("The search feature returns no results", "technical"),
    ("The export gets stuck halfway through", "technical"),
    ("The website is extremely slow today", "technical"),
]

DEMO_TICKET = "The dashboard keeps spinning and never finishes loading."


def keyword_baseline(ticket):
    normalized = ticket.lower()
    for category, words in KEYWORDS.items():
        if any(word in normalized for word in words):
            return category
    return "Unknown"


def build_pipeline():
    return Pipeline(
        [
            ("tfidf", TfidfVectorizer(ngram_range=(1, 2))),
            (
                "classifier",
                LogisticRegression(l1_ratio=0.0, max_iter=1000),
            ),
        ]
    )


def train_and_evaluate():
    texts = [text for text, _ in TRAINING_DATA]
    labels = [label for _, label in TRAINING_DATA]

    train_texts, test_texts, train_labels, test_labels = train_test_split(
        texts,
        labels,
        test_size=0.25,
        random_state=42,
        stratify=labels,
    )

    evaluation_model = build_pipeline()
    evaluation_model.fit(train_texts, train_labels)
    test_predictions = evaluation_model.predict(test_texts)
    accuracy = float(accuracy_score(test_labels, test_predictions))

    evaluation_rows = [
        {
            "ticket": text,
            "actual": actual,
            "predicted": str(predicted),
        }
        for text, actual, predicted in zip(
            test_texts,
            test_labels,
            test_predictions,
        )
    ]

    final_model = build_pipeline()
    final_model.fit(texts, labels)
    return final_model, accuracy, evaluation_rows


def classify_ticket(model, ticket):
    prediction = str(model.predict([ticket])[0])
    probabilities = model.predict_proba([ticket])[0]
    scores = {
        str(label): round(float(probability), 3)
        for label, probability in zip(model.classes_, probabilities)
    }
    return prediction, scores[prediction], scores


if __name__ == "__main__":
    trained_model, held_out_accuracy, held_out_rows = train_and_evaluate()
    predicted_category, confidence, class_scores = classify_ticket(
        trained_model,
        DEMO_TICKET,
    )

    print("Keyword baseline:", keyword_baseline(DEMO_TICKET))
    print("ML prediction:", predicted_category)
    print("Class probabilities:", class_scores)
    print("Held-out accuracy:", held_out_accuracy)
    print("Held-out predictions:", held_out_rows)

How should you use this reference?

Compare the imports, data order, function order, indentation, and final print statements line by line. The complete file matches the script that produced your terminal results.

Your classifier now turns labeled tickets into a prediction plus measurable evidence. Next, you will connect those results to the dashboard so someone can compare both routing systems in one page.

Put the Model Behind the Dashboard

Your command-line demo proved that the trained classifier can route the default ticket. It also exposed the confidence score plus the complete probability breakdown.

Support teams need those results inside an interface they can use. This step connects the classifier to your Streamlit dashboard while keeping the keyword baseline visible for comparison.

In this step, get ready to:
  • Train the classifier when the dashboard starts.
  • Compare the keyword result with the model prediction.
  • Display the model confidence plus every class probability.
Train the model at dashboard startup

The dashboard needs a trained model before it can classify a submitted ticket. Loading the model during startup makes it available to every button click.

  • Select app.py in the Visual Studio Code Explorer sidebar.
  • Select everything from import streamlit as st through the current st.write("Compare a brittle keyword baseline with a trained text classifier.") line.
  • Replace the selected section with this dashboard startup code:
import streamlit as st

from triage import (
    DEMO_TICKET,
    classify_ticket,
    keyword_baseline,
    train_and_evaluate,
)

model, accuracy, evaluation_rows = train_and_evaluate()

st.title("AI Support Ticket Triage")
st.write(
    "Compare a brittle keyword baseline with a trained text classifier."
)

What does this code do?

  • The expanded import makes classify_ticket() available to the dashboard.
  • The call to train_and_evaluate() trains the final classifier on all labeled examples.
  • The model variable holds the trained pipeline used for live classification.
  • The accuracy variable plus evaluation_rows preserve the held-out evidence returned during training.
  • Save app.py.

Before you start the dashboard, do you expect the classifier to train without interrupting the page load?

  • Start the dashboard from the activated terminal by running this command:
python -m streamlit run app.py

What does this command do?

This command runs app.py through the active virtual environment. Streamlit serves the dashboard in your browser after the model finishes training.

Your browser opens the support ticket dashboard with the default ticket already filled in. Good progress. The dashboard now trains its classifier during startup.

Dashboard not loading?

Confirm that the terminal still uses the active .venv environment. Check that every imported name matches the code above.

A startup error may also point to an unsaved change in triage.py. Help me debug the dashboard startup problem.

Show the model prediction and confidence

A predicted category tells the viewer where the model would route the ticket. The confidence percentage shows the strength of that choice.

  • In app.py, locate the if ticket.strip(): block.
  • Select the indented line that displays the keyword result.
  • Replace that line with this comparison block:
        baseline_result = keyword_baseline(ticket)
        prediction, confidence, scores = classify_ticket(model, ticket)

        st.write("Keyword baseline:", baseline_result)
        st.metric("ML prediction", prediction.title())
        st.metric("Model confidence", f"{confidence:.1%}")

How does the comparison work?

  • The baseline_result variable stores the route produced by the fixed keyword mapping.
  • The prediction variable stores the category selected by the trained classifier.
  • The confidence variable stores the probability assigned to that category.
  • The scores dictionary keeps the probability assigned to every category.
  • Save app.py.
  • Refresh the running dashboard in your browser.
  • Click Classify ticket.

You see Unknown for the keyword baseline. You also see an ML category plus its confidence percentage.

The comparison is now live. Your dashboard shows how the trained classifier handles language that defeats the fixed vocabulary.

Model metrics missing?

Check that the replacement block remains inside if ticket.strip():. Every new line needs the same base indentation as the original keyword result.

Confirm that you saved app.py before refreshing the browser. Help me fix the missing model metrics.

Render every class probability

The confidence metric shows only the winning category. The complete probability distribution reveals how the model divided its confidence across billing, account, and technical tickets.

  • In app.py, locate the metric labeled Model confidence.
  • Place your cursor directly below that metric.
  • Add the complete probability display with this code:
        st.write("Class probabilities")
        st.json(
            {
                label: f"{score:.1%}"
                for label, score in scores.items()
            }
        )

Why show every probability?

The dictionary comprehension formats each score as a percentage. The st.json() display keeps the three labeled values together.

Viewers can now compare the winning probability with both alternatives. This makes uncertainty visible inside the product.

  • Save app.py.

Before you test the completed dashboard, do you expect the baseline result to match the trained model route for the default ticket?

  • Refresh the running dashboard in your browser.
  • Keep the prefilled support ticket unchanged.
  • Click Classify ticket.

You see Unknown for the keyword baseline. The page also shows an ML category, a confidence percentage, and three class probabilities.

That is the full dashboard integration working. Your classifier now gives viewers a route plus the evidence behind its choice.

Probability breakdown missing?

Confirm that the st.json() block sits inside if ticket.strip():. Align it with the prediction metrics above.

Check that the dictionary comprehension reads from scores.items(). Help me debug the probability display.

✔️ Awesome, I've got everything!

Your app.py now compares both routing systems. It also exposes the model confidence plus every category probability.

ⓧ I'd like to double check the full code

  • Compare your complete app.py with this version line by line.
import streamlit as st

from triage import (
    DEMO_TICKET,
    classify_ticket,
    keyword_baseline,
    train_and_evaluate,
)

model, accuracy, evaluation_rows = train_and_evaluate()

st.title("AI Support Ticket Triage")
st.write(
    "Compare a brittle keyword baseline with a trained text classifier."
)

ticket = st.text_area("Support ticket", value=DEMO_TICKET)

if st.button("Classify ticket", type="primary"):
    if ticket.strip():
        baseline_result = keyword_baseline(ticket)
        prediction, confidence, scores = classify_ticket(model, ticket)

        st.write("Keyword baseline:", baseline_result)
        st.metric("ML prediction", prediction.title())
        st.metric("Model confidence", f"{confidence:.1%}")
        st.write("Class probabilities")
        st.json(
            {
                label: f"{score:.1%}"
                for label, score in scores.items()
            }
        )
    else:
        st.write("Enter a support ticket before classifying it.")

What should you compare?

Check the imported names plus the model training assignment. Confirm that every result display remains inside the non-empty ticket branch.

Your trained model now works through a browser-based product. Next up, you will expose the held-out evaluation so viewers can inspect model quality.

Make Model Quality Visible

Your Streamlit dashboard already compares the keyword baseline with a trained model. It also shows the probability assigned to each category.

A convincing AI demo needs visible evidence about model behavior. This step adds a held-out evaluation so viewers can assess accuracy before trusting a live prediction.

In this step, get ready to:
  • Add a held-out accuracy metric to the dashboard.
  • Display all nine evaluation tickets with their actual categories plus predicted categories.
  • Test an original ticket using model confidence plus the evaluation evidence.
Display held-out evaluation evidence

The app already receives accuracy plus evaluation_rows from train_and_evaluate(). Displaying both values turns the evaluation into evidence that a viewer can inspect.

  • Select app.py in the project sidebar.
  • Scroll to the final line of app.py.
  • Add the evaluation output at the left edge of the file by copying this code below:
st.write("Held-out evaluation")
st.metric("Held-out accuracy", f"{accuracy:.1%}")
st.write(evaluation_rows)

What does this code show?

  • The first st.write() call separates the evaluation evidence from the live classifier.
  • The st.metric() call formats accuracy as a percentage with one decimal place.
  • The final st.write() call renders every dictionary in evaluation_rows for inspection.
  • Save app.py.

Before you refresh, what evidence do you expect to appear below the live classifier?

  • Return to the running app in your browser.
  • Refresh the page with your browser's refresh control.

You'll see a Held-out evaluation heading followed by an accuracy card. Below it, you'll see nine rows containing ticket text plus actual and predicted categories.

That is the evidence layer complete. Your dashboard now exposes the model's results on examples withheld from evaluation training.

Evaluation section missing?

  • Check that the three new lines sit at the left edge of app.py. Indented lines may place the evaluation inside the button logic.
  • Save app.py before refreshing the browser.
  • Use the launch command from the previous step if the local app has stopped.

Still missing the evaluation? Help me find why my Streamlit evaluation output is not appearing.

✔️ Awesome, I've got everything!

Great. Double check that app.py is saved. Confirm that the held-out evaluation appears below the classifier.

ⓧ I'd like to double check the full code

Here is the complete app.py file for comparison. Check each line against your version.

import streamlit as st

from triage import (
    DEMO_TICKET,
    classify_ticket,
    keyword_baseline,
    train_and_evaluate,
)

model, accuracy, evaluation_rows = train_and_evaluate()

st.title("AI Support Ticket Triage")
st.write(
    "Compare a brittle keyword baseline with a trained text classifier."
)

ticket = st.text_area("Support ticket", value=DEMO_TICKET)

if st.button("Classify ticket", type="primary"):
    if ticket.strip():
        baseline_result = keyword_baseline(ticket)
        prediction, confidence, scores = classify_ticket(model, ticket)

        st.write("Keyword baseline:", baseline_result)
        st.metric("ML prediction", prediction.title())
        st.metric("Model confidence", f"{confidence:.1%}")
        st.write("Class probabilities")
        st.json(
            {
                label: f"{score:.1%}"
                for label, score in scores.items()
            }
        )
    else:
        st.write("Enter a support ticket before classifying it.")

st.write("Held-out evaluation")
st.metric("Held-out accuracy", f"{accuracy:.1%}")
st.write(evaluation_rows)
Inspect each evaluation prediction

The accuracy card summarizes performance across the held-out examples. The individual rows reveal which tickets produced that score.

  • Read the value shown in the Held-out accuracy card.
  • Count the tickets listed beneath the accuracy card.
  • Compare the actual value with the predicted value in each row.
  • Identify any row where the two category values differ.

How to read the evidence

The held-out score summarizes predictions for nine examples separated from the model's evaluation training data. Each new route still needs judgment.

The row-level results make mistakes visible. A viewer can inspect the exact ticket behind every correct or incorrect prediction.

Judge an original ticket route

Confidence shows how strongly the trained model favors its winning category. The held-out results provide wider context for deciding whether that route deserves trust.

  • Replace the text in the Support ticket field with one original support request of your own.

Before you classify it, which category do you expect the model to choose?

  • Click Classify ticket.

You'll see the baseline route followed by the model prediction. You'll also see its confidence plus all three class probabilities.

  • Compare the model prediction with the category you intended.
  • Read the percentage shown under Model confidence.
  • Use the held-out accuracy to judge the route.
  • Use the individual evaluation rows to support your decision.

That closes the loop. Every live route now sits beside visible model evidence.

Secret mission

Route Uncertain Tickets to Human Review

Your classifier currently sends every ticket to one of three categories, even when its confidence is weak. Add a review threshold that gives uncertain tickets a safe route to a human reviewer.

Clean Up Your Resources

Clean Up Your Resources

Your Streamlit dashboard runs entirely on your Mac. It has no ongoing costs.

Decide whether to keep your resources running, pause them to come back later, or delete them entirely.

Resources you used:

  • The running Streamlit dashboard process.
  • The local ai-ticket-triage folder. This includes the .venv environment.

Keep everything running

No action needed. Choose this if you want to keep testing support tickets in your local dashboard.

  • Leave the dashboard process running while you test more tickets.
  • Keep the ai-ticket-triage folder in its current location.
  • Keep the .venv environment inside the project folder.

Pause - I'll come back to this later

Shut down the dashboard to free up your terminal. Your code stays ready for your next session.

  • Return to the terminal panel running the dashboard.
  • Press Ctrl+C to stop the local dashboard.
  • Close the browser tab containing the dashboard.
  • Leave the ai-ticket-triage folder unchanged.

Delete - I don't want to use this again

Deleting this folder is permanent. The command below removes only ai-ticket-triage. This includes its isolated environment.

  • Copy any files you want to keep to a location outside ai-ticket-triage.
  • Return to the terminal panel from earlier.
  • Press Ctrl+C if the local dashboard is still running.
  • Remove the complete project folder by running these commands:
cd ..
rm -rf ai-ticket-triage
ls

What Does This Cleanup Do?

  • The first command moves the terminal to the folder containing your project.
  • The second command permanently removes the project folder.
  • The final command lists the remaining folders so you can verify the deletion.

The final terminal output should no longer include ai-ticket-triage. Your project files and isolated packages are now removed.

Still See the Project Folder?

  • Check that you returned to the folder containing ai-ticket-triage before removing it.
  • Check that the folder name uses the exact spelling ai-ticket-triage.

Help me safely remove my local project folder.

Nice Work!

Nice Work!

You did it! Your local AI support ticket triage dashboard now classifies customer messages while keeping its evidence visible.

You've learned how to:

  • Build a browser-based triage dashboard in Python with Streamlit.
  • Expose the limits of a keyword baseline with a paraphrased technical ticket. Train a logistic regression classifier on TF-IDF features from labeled examples.
  • Evaluate the classifier against held-out tickets. Display its accuracy alongside each evaluation prediction. Keep confidence scores plus class probabilities visible for live tickets.
  • Secret Mission: Add a human-review route that catches uncertain predictions before they receive a category.

Ready to quiz yourself?