Build a Leakage-Aware Topic Classifier

Build a local text classifier that reveals and fixes metadata leakage.

Introduction

30 Second Summary

A high score can make a text-sorting tool look trustworthy. Hidden clues in forum headers and quoted replies can make that score misleading.

In this project, you will build a browser-based topic classifier powered by TF-IDF and Complement Naive Bayes. You will expose metadata leakage by comparing a baseline score with a cleaner held-out evaluation.

What You'll Build

Your finished Streamlit app turns a pasted forum-style post into atheism discussion, religion discussion, computer graphics, or space science.

By the end of this project, you'll have:

  • Run a repeatable held-out benchmark that prints baseline accuracy beside cleaned accuracy.
  • Classify pasted text as one of four readable topics through a local browser app.
  • Explain why a cleaned score can be more trustworthy after headers, footers, and quoted replies are removed.
  • Secret Mission: Rank all four model-reported probabilities from highest to lowest. Flag a prediction for human review when its top probability falls below the 60% threshold.

Are there any prerequisites?

You need basic Python syntax plus internet access for the dataset's first download. This project assumes Python and Visual Studio Code are available on Windows.

Before We Start

You are committing to a four-topic classifier that labels forum-style text as atheism discussion, religion discussion, computer graphics, or space science. Honest evaluation matters because metadata leakage can inflate apparent performance.

Prepare a Reproducible Windows Environment

Your classifier needs a compatible Python version before it can train or launch in a browser. A project-specific virtual environment keeps its libraries isolated from the rest of your machine.

Pinned dependency versions make the setup repeatable on another computer. Ignoring generated files keeps future commits focused on the project code.

In this step, get ready to:
  • Verify Python 3.11 through 3.14 in the project terminal.
  • Create the pinned dependency file plus the Git ignore rules.
  • Build an isolated environment with scikit-learn 1.9.1 plus Streamlit 1.65.0.
Verify Python inside your project folder

An opened folder gives Visual Studio Code one workspace for your files plus its terminal. The integrated Windows PowerShell terminal starts inside that folder.

  • Press the Windows key to open Windows search.
  • Type Visual Studio Code into the search bar.
  • Press Enter to open Visual Studio Code.
  • Select File from the top menu.
  • Select Open Folder....
  • Select your Desktop in the folder dialog.
  • Select New Folder.
  • Name the folder your project folder name.

Why open a folder first?

Visual Studio Code treats the opened folder as a workspace. Commands in its integrated terminal start from the same location as the files shown in the Explorer sidebar.

  • Select the new folder in the folder dialog.
  • Select Select Folder to open it as your workspace.
  • Confirm that you trust the folder if the Workspace Trust dialog appears.
  • Select View from the top menu.
  • Select Terminal to open the integrated terminal.
  • Select PowerShell from the terminal profile dropdown if another shell opens.
  • Check the installed Python version by running this command:
python --version

What does this command check?

The command prints the Python interpreter currently available to this PowerShell session. This project supports Python 3.11 through 3.14.

✔️ I see Python 3.11 through 3.14

Your Python version is compatible with both pinned libraries. You can continue with the interpreter already available in this terminal.

ⓧ I see a version outside 3.11 to 3.14

This project needs a Python version supported by both pinned libraries. Install Python 3.14.8 before creating the virtual environment.

  • Open the official Windows downloads page.
  • Download the Windows installer for Python 3.14.8.
  • Run the downloaded installer.
  • Complete the installer prompts.
  • Close Visual Studio Code after the installation finishes.
  • Return to your project folder name in Visual Studio Code.
  • Open the integrated PowerShell terminal from the View menu.
  • Check the new Python version by running this command:
python --version

What should this confirm?

PowerShell should now report Python 3.14.8. The new interpreter can create the project environment.

Still seeing the previous version?

Close every Visual Studio Code window before reopening the project folder. This gives the new terminal a fresh view of the installed Python interpreter.

If the previous version remains active, help me troubleshoot which Python installation PowerShell is using.

ⓧ Command not found

PowerShell cannot currently find a Python installation. Install Python 3.14.8 before continuing.

  • Open the official Windows downloads page.
  • Download the Windows installer for Python 3.14.8.
  • Run the downloaded installer.
  • Complete the installer prompts.
  • Close Visual Studio Code after the installation finishes.
  • Return to your project folder name in Visual Studio Code.
  • Open the integrated PowerShell terminal from the View menu.
  • Check that PowerShell can find Python by running this command:
python --version

What should this confirm?

PowerShell should now print Python 3.14.8. That output confirms the command is available in the project terminal.

Still cannot find Python?

Close every Visual Studio Code window before reopening the project folder. A terminal that stayed open during installation may still use the previous system configuration.

If the command is still unavailable, help me make Python available in Windows PowerShell.

Create the pinned project files

The requirements.txt file records the exact library versions used by this project. The .gitignore file keeps generated environment files plus Python cache files out of future commits.

  • Select New File in the Visual Studio Code Explorer sidebar.
  • Enter requirements.txt as the file name.
  • Add the pinned dependencies by copying this content into requirements.txt:
scikit-learn==1.9.1
streamlit==1.65.0

What does this file control?

  • The first line pins scikit-learn to version 1.9.1 for model training plus evaluation.
  • The second line pins Streamlit to version 1.65.0 for the browser interface.
  • Save requirements.txt by pressing Ctrl+S.
  • Confirm that requirements.txt appears in the Explorer sidebar with both pinned versions.

Does the dependency file look different?

Check that each package is on its own line. Remove any extra spaces around the double equals signs.

If the file still looks wrong, help me compare my requirements file with the expected two-line version.

  • Select New File in the Explorer sidebar.
  • Enter .gitignore as the file name.
  • Add the generated-file exclusions by copying this content into .gitignore:
.venv/
__pycache__/

What do these rules exclude?

  • The .venv/ rule excludes the project-specific environment.
  • The __pycache__/ rule excludes cache folders created when Python runs project code.
  • Save .gitignore by pressing Ctrl+S.
  • Confirm that .gitignore appears beside requirements.txt in the Explorer sidebar.

Is the ignore file missing?

Check that the file name begins with a dot. Make sure the editor did not add a text-file extension.

If the file remains hidden or renamed, help me create a dotfile in the Visual Studio Code Explorer.

Use the comparison below to confirm both project files before creating the environment.

✔️ Awesome, I've got everything!

Both files match the project configuration. Your dependencies are pinned plus generated files are excluded.

ⓧ I'd like to double check the full code

scikit-learn==1.9.1
streamlit==1.65.0

What should match?

Your requirements.txt file should contain these two package pins in this order.

.venv/
__pycache__/

What should match?

Your .gitignore file should contain these two directory patterns in this order.

The project files now describe the environment. The next command creates that isolated environment inside the same folder.

  • Create .venv in the project folder by running this command:
python -m venv .venv

What does this command create?

Python creates an isolated environment named .venv inside the project folder. Its own interpreter plus package directory keep this project's libraries separate.

  • Confirm that .venv appears in the Explorer sidebar after the command finishes.

Was the environment not created?

Check that the terminal still points to the folder opened in Visual Studio Code. Confirm that the earlier version command reported a supported Python version.

If creation still fails, help me diagnose why Python cannot create .venv in my project folder.

  • Activate the new environment in the current PowerShell session by running this command:
.venv\Scripts\Activate.ps1

What does activation change?

Activation makes this terminal use the Python interpreter plus package directory inside .venv. Packages installed next stay inside this project environment.

  • Confirm that the environment name appears at the start of the PowerShell prompt.

Did PowerShell block activation?

Confirm that .venv exists in the Explorer sidebar. Make sure the activation command was run from the project folder.

If PowerShell still refuses to activate the environment, help me diagnose my virtual environment activation problem.

Install and verify the pinned tools

The activated environment is an empty container for project packages. Installing from requirements.txt recreates the exact dependency set recorded in the project.

  • Install the pinned dependencies into the active environment by running this command:
pip install -r requirements.txt

What does this command install?

The package installer reads both pinned lines from requirements.txt. It installs those versions plus the supporting packages they need into .venv.

  • Wait for PowerShell to return to the active environment prompt.
  • Confirm that the installation finishes without an error.

Did the installation fail?

Check that the environment name is visible in the terminal prompt. Confirm that both package pins in requirements.txt match the comparison above.

If the installation still fails, help me troubleshoot the dependency installation output.

  • Verify the scikit-learn installation by running this command:
python -c "import sklearn; sklearn.show_versions()"

What does this verification prove?

Python imports scikit-learn from the active environment. The library then prints its installed version plus information about its supporting packages.

You should see scikit-learn version details in PowerShell. That output proves the model library can load successfully.

Did the import check fail?

Confirm that the environment name remains visible in the prompt. Run the dependency installation command again if PowerShell was reopened after activation.

If the import still fails, help me trace which Python environment contains scikit-learn.

The final check starts a local example server. PowerShell stays occupied until you stop the server.

Before you run this, what do you expect a successful Streamlit installation to open?

  • Launch the Streamlit verification app by running this command:
python -m streamlit hello

What does this command verify?

Python starts Streamlit from the active virtual environment. Streamlit serves its example app locally plus asks your browser to display it.

Your browser should open the Streamlit Hello app. That is the full environment working from package import to browser output.

  • Return to the Visual Studio Code PowerShell terminal.
  • Press Ctrl+C to stop the Streamlit example server.
  • Confirm that the PowerShell prompt returns with the environment still active.

Did the browser stay closed?

Check the PowerShell output for the local app address. Open that address in your browser if Streamlit started without opening a tab.

If the server did not start, help me troubleshoot the Streamlit Hello command in my virtual environment.

Your isolated Windows environment is ready to build the classifier. Next, you will train the first four-topic model plus see its baseline held-out accuracy.

Train a Baseline Topic Classifier

Your isolated Windows environment is ready. Your next goal is a four-topic model built with supervised classification.

A working first model gives you an early result to inspect. It also preserves message metadata, creating a tempting baseline that you will challenge after you see it run.

In this step, get ready to:
  • Load four documented topics into separate training and test subsets.
  • Train a TF-IDF and Complement Naive Bayes pipeline.
  • Measure held-out accuracy on the test subset.
Load the four topics

The 20 Newsgroups dataset contains forum-style posts grouped by topic. You will use scikit-learn to load its predefined train and test data as separate subsets.

  • Create model.py inside the project folder with the file sidebar's new-file control in Visual Studio Code.
  • Add the dataset imports plus the first runnable main() function by pasting this code:
from sklearn.datasets import fetch_20newsgroups
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics import accuracy_score
from sklearn.naive_bayes import ComplementNB
from sklearn.pipeline import Pipeline

CATEGORIES = [
    "alt.atheism",
    "talk.religion.misc",
    "comp.graphics",
    "sci.space",
]


def main():
    data_train = fetch_20newsgroups(
        subset="train",
        categories=CATEGORIES,
        shuffle=True,
        random_state=42,
    )

    print("Topics:", ", ".join(data_train.target_names))


if __name__ == "__main__":
    main()

What does this code do?

  • The imports make the dataset loader plus the feature extraction plus the classification plus the evaluation tools available to this file.
  • CATEGORIES limits the experiment to four topic identifiers.
  • fetch_20newsgroups loads the predefined training subset for those topics.
  • random_state=42 keeps the shuffled ordering reproducible.
  • The final print statement exposes the dataset's topic identifiers as your first visible result.
  • Save model.py.

Before you run this first build, consider whether the terminal will list all four topic identifiers.

  • Run the dataset loader in the activated PowerShell terminal with this command:
python model.py

What should you see?

The first run downloads the required dataset into scikit-learn's default local cache. The terminal then prints a Topics: line containing all four identifiers.

That first data loop is working. Your script can now fetch the training documents for the exact topics this classifier needs.

Dataset not loading?

Confirm that your internet connection is available for the first dataset download. Check that the terminal prompt still shows the activated .venv environment.

Compare every category string in CATEGORIES with the code block above. A spelling mismatch prevents the requested topic from loading.

Help me fix this dataset loading problem.

Build and train the baseline pipeline

Raw documents need numeric features before a classifier can learn from them. A machine learning pipeline connects TF-IDF feature extraction to a Complement Naive Bayes classifier.

  • Locate the existing main() function in model.py.
  • Compare the function with this reference before replacing it:
def main():
    data_train = fetch_20newsgroups(
        subset="train",
        categories=CATEGORIES,
        shuffle=True,
        random_state=42,
    )

    print("Topics:", ", ".join(data_train.target_names))

What are you replacing?

This first version only loads the training subset. The replacement keeps that loader while adding the held-out test subset plus the complete training and evaluation flow.

  • Replace the current main() function with this version:
def main():
    data_train = fetch_20newsgroups(
        subset="train",
        categories=CATEGORIES,
        shuffle=True,
        random_state=42,
    )
    data_test = fetch_20newsgroups(
        subset="test",
        categories=CATEGORIES,
        shuffle=True,
        random_state=42,
    )

    model = Pipeline(
        [
            ("tfidf", TfidfVectorizer()),
            ("classifier", ComplementNB()),
        ]
    )
    model.fit(data_train.data, data_train.target)
    predictions = model.predict(data_test.data)
    accuracy = accuracy_score(data_test.target, predictions)

    print("Topics:", ", ".join(data_train.target_names))
    print(f"Baseline accuracy: {accuracy:.3f}")

How does the baseline work?

  • data_train contains the documents and labels used for learning.
  • data_test contains separate held-out documents used for evaluation.
  • TfidfVectorizer() converts each document into weighted text features.
  • ComplementNB() learns how those features differ across the four topics.
  • model.fit() trains the complete pipeline on the training subset.
  • accuracy_score calculates the fraction of held-out predictions that match their labels.
  • Save the updated model.py.

Before you run this version, consider how the fitted model changes the terminal output.

  • Train the pipeline and evaluate its predictions by running this command:
python model.py

What should you see now?

The terminal prints the four topic identifiers again. It also prints Baseline accuracy: followed by a score rounded to three decimal places.

Your first classifier now completes the full learning loop. It trains on one subset before measuring its predictions on separate held-out documents.

Baseline run failing?

Check that the updated main() function remains above the final script entry block. Confirm that both pipeline step names and their surrounding quotes match the code block.

Check the indentation inside main() if Python points to a syntax problem. Every statement in the function begins with four spaces.

Help me debug my baseline classifier.

✔️ Awesome, I've got everything!

Your model.py is complete for this baseline.

ⓧ I'd like to double check the full code

from sklearn.datasets import fetch_20newsgroups
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics import accuracy_score
from sklearn.naive_bayes import ComplementNB
from sklearn.pipeline import Pipeline

CATEGORIES = [
    "alt.atheism",
    "talk.religion.misc",
    "comp.graphics",
    "sci.space",
]


def main():
    data_train = fetch_20newsgroups(
        subset="train",
        categories=CATEGORIES,
        shuffle=True,
        random_state=42,
    )
    data_test = fetch_20newsgroups(
        subset="test",
        categories=CATEGORIES,
        shuffle=True,
        random_state=42,
    )

    model = Pipeline(
        [
            ("tfidf", TfidfVectorizer()),
            ("classifier", ComplementNB()),
        ]
    )
    model.fit(data_train.data, data_train.target)
    predictions = model.predict(data_test.data)
    accuracy = accuracy_score(data_test.target, predictions)

    print("Topics:", ", ".join(data_train.target_names))
    print(f"Baseline accuracy: {accuracy:.3f}")


if __name__ == "__main__":
    main()
Challenge the baseline score

The score looks like a clean measure of topic recognition. However, each document still contains message headers plus signature footers plus quoted replies.

Those fields can reveal author or thread clues that make the prediction task easier. This is metadata leakage because the model can exploit shortcuts beyond the message content you want it to classify.

Before the final run, consider which two output lines prove that the dataset and evaluation loop are both working.

  • Perform the final baseline check by running this command:
python model.py

What does this result prove?

You should see a Topics: line containing all four identifiers. You should also see a Baseline accuracy: line containing the held-out score.

This is the intended shortfall in the first model. The evaluation uses separate test documents, but their retained metadata can make the score more optimistic than a content-focused evaluation.

You now have a reproducible baseline plus a reason to question its score. Next, you will repeat the experiment without the message metadata shortcuts.

Expose and Fix Metadata Leakage

Your baseline classifier already produces a held-out score from separate train and test data. This gives you a working benchmark.

However, the messages still include author details plus thread clues that can cause metadata leakage. This step repeats the same learning process after removing those shortcuts.

In this step, get ready to:
  • Extract reusable training logic.
  • Evaluate intact messages against metadata-stripped messages.
  • Return the cleaned model artifact for the browser app.
Extract the reusable pipeline

A reusable machine learning pipeline keeps feature extraction plus classification consistent across both evaluations. Each evaluation receives a fresh model with the same structure.

  • In model.py, place your cursor on the blank line below the closing ] for CATEGORIES.
  • Add the shared definitions by pasting this code:
METADATA_FIELDS = ("headers", "footers", "quotes")


def make_pipeline():
    return Pipeline(
        [
            ("tfidf", TfidfVectorizer()),
            ("classifier", ComplementNB()),
        ]
    )

What does this code do?

  • The METADATA_FIELDS tuple records the message sections that the cleaned evaluation removes.
  • The make_pipeline() function creates a fresh Pipeline for each training run.
  • The tfidf step converts message text into TF-IDF features.
  • The classifier step uses ComplementNB to assign a topic.
  • Inside main(), find the Pipeline(...) expression shown below.
Pipeline(
        [
            ("tfidf", TfidfVectorizer()),
            ("classifier", ComplementNB()),
        ]
    )

What are you replacing?

This expression constructs the classifier directly inside main(). The new factory gives every evaluation the same model structure.

  • Replace the assignment containing that expression with this factory call:
    model = make_pipeline()

Why use the factory?

The factory keeps the model definition in one place. This prevents the baseline evaluation from drifting away from the cleaned evaluation.

  • Save model.py.

Before you run this check, predict whether the baseline entry point still produces its original output.

  • Run the refactored baseline by executing this command:
python model.py

What does this check prove?

This command executes the current main() function. A successful run proves that the pipeline factory creates a trainable classifier.

You will still see the four topic identifiers plus the baseline held-out accuracy. The reusable factory now produces the same working result.

Baseline no longer running?

Check that make_pipeline() sits above main(). Confirm that the replacement line remains indented inside main().

Compare the parentheses in the factory with the snippet above. One missing parenthesis prevents Python from parsing the file.

Help me debug the pipeline factory refactor.

Make evaluation repeatable

The remove parameter controls which message sections the dataset loader excludes. An empty tuple keeps every section.

  • Add the reusable evaluation function above main() by pasting this code:
def train_and_evaluate(remove=()):
    data_train = fetch_20newsgroups(
        subset="train",
        categories=CATEGORIES,
        shuffle=True,
        random_state=42,
        remove=remove,
    )
    data_test = fetch_20newsgroups(
        subset="test",
        categories=CATEGORIES,
        shuffle=True,
        random_state=42,
        remove=remove,
    )

    model = make_pipeline()
    model.fit(data_train.data, data_train.target)
    predictions = model.predict(data_test.data)
    accuracy = accuracy_score(data_test.target, predictions)

    return model, data_train.target_names, accuracy

How does evaluation stay fair?

  • The train subset supplies examples for fitting the model.
  • The separate test subset measures performance on held-out messages.
  • The remove value is applied to both subsets.
  • The returned tuple keeps the fitted model plus its readable target names plus its held-out accuracy.
  • Add train_comparison() between train_and_evaluate() and main() by pasting this code:
def train_comparison():
    baseline_result = train_and_evaluate()
    clean_model, clean_target_names, clean_accuracy = train_and_evaluate(
        remove=METADATA_FIELDS
    )

    return {
        "model": clean_model,
        "target_names": clean_target_names,
        "baseline_accuracy": baseline_result[2],
        "clean_accuracy": clean_accuracy,
    }

What does the comparison return?

  • The first evaluation uses the default remove=() value to preserve the baseline condition.
  • The second evaluation uses METADATA_FIELDS to strip headers plus footers plus quoted replies.
  • The returned model is the cleaned classifier that the browser app can use.
  • The returned accuracy values make the two evaluation conditions easy to compare.
  • Save model.py.

Before you run this check, predict whether Python can parse both new functions while the existing entry point remains in place.

  • Check the refactored file by running this command:
python model.py

What does this intermediate run prove?

Python parses every function before it calls the current entry point. Seeing the baseline output confirms that the reusable evaluation functions contain valid syntax.

Good progress. The baseline output appears again after the reusable comparison logic has been added.

File failing after the new functions?

Confirm that both function definitions begin at the left edge of model.py. Check that each function body uses four spaces of indentation.

Make sure train_comparison() appears after train_and_evaluate(). Python must define the evaluation helper before the comparison calls it.

Help me fix the reusable evaluation functions.

Run the honest comparison

The current entry point still runs only the original baseline flow. The final entry point calls train_comparison() so both data conditions are evaluated in one repeatable run.

  • In model.py, select everything from def main(): through the bottom of the file.
  • Replace the selected code with this final entry point:
def main():
    result = train_comparison()
    print("Topics:", ", ".join(result["target_names"]))
    print(f"Baseline accuracy: {result['baseline_accuracy']:.3f}")
    print(f"Cleaned accuracy: {result['clean_accuracy']:.3f}")
    print("Removed metadata:", METADATA_FIELDS)


if __name__ == "__main__":
    main()

What changes in the final entry point?

  • The result dictionary holds the cleaned model artifact plus both evaluation scores.
  • The baseline line reports performance when metadata remains available.
  • The cleaned line reports performance after the three metadata fields are removed.
  • The final line prints the removal tuple used for the cleaned evaluation.
  • Save model.py.

Before you run the final check, predict whether the cleaned score will match the baseline score.

  • Run both held-out evaluations by executing this command:
python model.py

What does the final run test?

The first evaluation leaves metadata intact. The second evaluation repeats the same training process with the removal tuple applied.

A lower cleaned score can provide the more trustworthy estimate because the model has fewer author plus thread shortcuts available.

You will see Baseline accuracy plus Cleaned accuracy in the terminal. The final line will show ('headers', 'footers', 'quotes') as the removed metadata tuple.

Missing one of the comparison lines?

Confirm that main() calls train_comparison(). Check that the four dictionary keys match the returned dictionary exactly.

If the run stops during data loading, confirm that the activated .venv still contains scikit-learn. The dataset should load from the existing local cache.

Help me debug the final leakage comparison.

✔️ Awesome, I've got everything!

Your final model.py now compares the leaky baseline with the metadata-stripped model. Make sure the file is saved.

ⓧ I'd like to double check the full code

from sklearn.datasets import fetch_20newsgroups
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics import accuracy_score
from sklearn.naive_bayes import ComplementNB
from sklearn.pipeline import Pipeline

CATEGORIES = [
    "alt.atheism",
    "talk.religion.misc",
    "comp.graphics",
    "sci.space",
]

METADATA_FIELDS = ("headers", "footers", "quotes")


def make_pipeline():
    return Pipeline(
        [
            ("tfidf", TfidfVectorizer()),
            ("classifier", ComplementNB()),
        ]
    )


def train_and_evaluate(remove=()):
    data_train = fetch_20newsgroups(
        subset="train",
        categories=CATEGORIES,
        shuffle=True,
        random_state=42,
        remove=remove,
    )
    data_test = fetch_20newsgroups(
        subset="test",
        categories=CATEGORIES,
        shuffle=True,
        random_state=42,
        remove=remove,
    )

    model = make_pipeline()
    model.fit(data_train.data, data_train.target)
    predictions = model.predict(data_test.data)
    accuracy = accuracy_score(data_test.target, predictions)

    return model, data_train.target_names, accuracy


def train_comparison():
    baseline_result = train_and_evaluate()
    clean_model, clean_target_names, clean_accuracy = train_and_evaluate(
        remove=METADATA_FIELDS
    )

    return {
        "model": clean_model,
        "target_names": clean_target_names,
        "baseline_accuracy": baseline_result[2],
        "clean_accuracy": clean_accuracy,
    }


def main():
    result = train_comparison()
    print("Topics:", ", ".join(result["target_names"]))
    print(f"Baseline accuracy: {result['baseline_accuracy']:.3f}")
    print(f"Cleaned accuracy: {result['clean_accuracy']:.3f}")
    print("Removed metadata:", METADATA_FIELDS)


if __name__ == "__main__":
    main()

What should the complete file contain?

The complete file has one pipeline factory plus one reusable evaluation function plus one comparison function. Its entry point prints both scores plus the exact metadata removal tuple.

That is the central lesson proved: your classifier now reports an evaluation that accounts for metadata leakage. Next, you will put the cleaned model behind an interactive browser interface.

Build the Interactive App

Your cleaned classifier now reports a baseline score beside its metadata-stripped score. That terminal result proves your evaluation works.

A terminal metric is awkward for another person to test. In this step, you'll use Streamlit to turn the cleaned pipeline into a browser app.

In this step, get ready to:
  • Connect the app to the cleaned model artifact.
  • Build the browser interface for pasted forum text.
  • Add submission handling for warnings or readable predictions.
Connect the app to the cleaned model

The app needs the same evaluated code path that produced your cleaned score. A readable label map turns each dataset identifier into a topic that another person can understand.

  • In the VS Code file sidebar, select the new-file control beside the project folder name.
  • Name the new file app.py.
  • Paste the following imports and topic label map into app.py:
import streamlit as st

from model import train_comparison

DISPLAY_NAMES = {
    "alt.atheism": "Atheism discussion",
    "talk.religion.misc": "Religion discussion",
    "comp.graphics": "Computer graphics",
    "sci.space": "Space science",
}

What does this code do?

  • The Streamlit import provides the browser interface components used throughout the app.
  • The train_comparison() import reuses the evaluated training process from model.py.
  • The DISPLAY_NAMES map converts dataset identifiers into labels that are easier to read.
  • Save app.py.
  • Confirm app.py appears beside model.py in the VS Code file sidebar.

Can't see app.py?

Check that you created the file in the same folder as model.py. Confirm that the filename ends with .py.

Ask for help if the file is missing from your project: Help me create app.py beside model.py in VS Code.

Training the comparison produces a reusable model artifact. st.cache_resource keeps that artifact available while the app handles browser interactions.

  • Place the following code below the closing brace of DISPLAY_NAMES in app.py:
@st.cache_resource
def load_artifact():
    return train_comparison()


artifact = load_artifact()

How does the model reach the app?

  • The load_artifact() function calls the existing comparison workflow.
  • The cache decorator preserves the returned model artifact for reuse.
  • The artifact dictionary gives the app access to the cleaned model.
  • Save app.py.

Before you launch the connection check, consider whether the model artifact can load without changing model.py.

  • Launch the connection check from the active PowerShell terminal by running this command:
python -m streamlit run app.py

What does this command do?

This starts Streamlit with the active virtual environment. It serves app.py as a local browser app.

You'll see a browser page served by Streamlit without an application error. That result confirms the cached artifact finished training inside the app process.

  • Return to the VS Code PowerShell terminal.
  • Stop the connection check with Ctrl+C.

Did the connection fail?

Confirm that app.py sits beside model.py. Check that train_comparison matches the function name in model.py.

Use this prompt for help: Help me debug why my Streamlit app cannot import train_comparison from model.py.

Build the classifier interface

The interface needs enough context for another person to understand the model's purpose. It also needs a place where they can paste the text they want to classify.

  • Place the following interface code below artifact = load_artifact() in app.py:
st.title("Leakage-Aware Tech Topic Classifier")
st.write(
    "Paste a forum-style post and classify it with a model trained after "
    "removing message headers, footers, and quoted replies."
)
st.write(f"Cleaned test accuracy: {artifact['clean_accuracy']:.1%}")

post = st.text_area("Forum post", height=180)

What's happening here?

  • The title names the product that someone is testing.
  • The description states that the model was trained after metadata removal.
  • The accuracy line exposes the cleaned held-out score inside the app.
  • The text area stores the pasted forum text in post.
  • Save app.py.

Before you check the interface, predict which details will help a visitor understand how the classifier was evaluated.

  • Launch the interface from the active PowerShell terminal by running this command:
python -m streamlit run app.py

What does this command check?

This serves the updated interface from app.py. The browser reflects the title plus the cleaned evaluation result.

You'll see the classifier title above its explanation. You'll also see the cleaned test accuracy above a large forum post text area.

  • Return to the VS Code PowerShell terminal.
  • Stop the interface check with Ctrl+C.

Is the interface incomplete?

Confirm that the interface code sits below artifact = load_artifact(). Check each closing parenthesis in the multi-line description.

Use this prompt for help: Help me find the syntax problem in my Streamlit title, description, accuracy, or text area code.

Classify submitted text

A useful classifier must handle an empty submission safely. Valid text should reach the cleaned model before the numeric class becomes a readable topic.

  • Place the following submission logic below the post text area in app.py:
if st.button("Classify", type="primary"):
    if not post.strip():
        st.warning("Enter some text before classifying.")
    else:
        prediction = int(artifact["model"].predict([post])[0])
        target_name = artifact["target_names"][prediction]
        st.success(f"Predicted topic: {DISPLAY_NAMES.get(target_name, target_name)}")

How does classification work?

  • The button starts classification during the interaction that selects it.
  • The empty-input check removes whitespace before deciding whether text is available.
  • The cleaned model predicts a numeric class for the submitted text.
  • The target list resolves that number to a dataset identifier.
  • The label map presents the final topic in readable language.
  • Save app.py.

✔️ Awesome, I've got everything!

Great. Confirm that app.py is saved before launching the finished classifier.

ⓧ I'd like to double check the full code

import streamlit as st

from model import train_comparison

DISPLAY_NAMES = {
    "alt.atheism": "Atheism discussion",
    "talk.religion.misc": "Religion discussion",
    "comp.graphics": "Computer graphics",
    "sci.space": "Space science",
}


@st.cache_resource
def load_artifact():
    return train_comparison()


artifact = load_artifact()

st.title("Leakage-Aware Tech Topic Classifier")
st.write(
    "Paste a forum-style post and classify it with a model trained after "
    "removing message headers, footers, and quoted replies."
)
st.write(f"Cleaned test accuracy: {artifact['clean_accuracy']:.1%}")

post = st.text_area("Forum post", height=180)

if st.button("Classify", type="primary"):
    if not post.strip():
        st.warning("Enter some text before classifying.")
    else:
        prediction = int(artifact["model"].predict([post])[0])
        target_name = artifact["target_names"][prediction]
        st.success(f"Predicted topic: {DISPLAY_NAMES.get(target_name, target_name)}")

How is the file organized?

  • The first section imports the app library plus the evaluated model workflow.
  • The next section caches the cleaned artifact returned by train_comparison().
  • The interface section displays the evaluation context plus the text input.
  • The final section validates the submission before displaying a readable prediction.

Before you launch the finished app, predict what an empty submission will trigger.

  • Launch the completed classifier from the active PowerShell terminal by running this command:
python -m streamlit run app.py

What does this launch prove?

This serves the completed app.py file with the cleaned model artifact. The running browser app now supports input validation plus classification.

  • Leave the Forum post text area empty.
  • Select Classify.

You'll see Enter some text before classifying. in a warning message. That confirms empty input cannot reach the model.

  • Enter NASA launched a spacecraft into orbit to study planets and stars. in the Forum post text area.
  • Select Classify.

You'll see Predicted topic: followed by one of the four readable topic labels. The result comes from the metadata-stripped model returned by train_comparison().

No prediction appearing?

Confirm that the text area contains visible characters before selecting Classify. Check that the prediction uses artifact["model"].

Use this prompt for help: Help me debug why my Streamlit classifier does not display a readable prediction.

You did it. Your classifier now turns pasted forum text into a readable topic from the cleaned model.

Secret mission

Add Probability-Aware Predictions

A predicted label can look decisive even when the model barely prefers it. Your extension ranks all four model-reported probabilities. A project-defined threshold sends weak results to human review.

Clean Up Your Resources

Clean Up Your Resources

This local Windows project has no ongoing cloud costs. Decide whether to keep its resources, pause the environment, or delete the local files.

Resources you used:

  • The local project folder containing .venv, .gitignore, requirements.txt, model.py, and app.py.
  • The default scikit-learn cache folder named scikit_learn_data.
  • The local Streamlit server process running in PowerShell.

Keep everything running

No deletion is needed. Choose this option if you want to keep refining the classifier.

  • Switch back to the VS Code PowerShell terminal.
  • Press Ctrl+C to stop the Streamlit server.
  • Keep the project folder in its current location.
  • Keep scikit_learn_data in its current location.

Your classifier code, virtual environment, and downloaded dataset remain ready for another session.

Pause - I'll come back to this later

This option stops the active processes while preserving every local file for later.

  • Switch back to the VS Code PowerShell terminal.
  • Press Ctrl+C to stop the Streamlit server.
  • Deactivate .venv by running this command:
deactivate

What Does Deactivate Do?

The deactivate command exits the project-specific virtual environment in your current terminal session.

  • Leave the project folder in its current location.
  • Leave scikit_learn_data in its current location.
  • Return to the same project folder in VS Code when you want to continue.
  • Reactivate .venv from its PowerShell terminal by running this command:
.venv\Scripts\Activate.ps1

How Does Reactivation Work?

The activation script reconnects the terminal to the existing .venv. Your pinned packages are already installed there.

Delete - I don't want to use this again

Deleting these local resources is permanent once they leave Windows recovery locations. Your online accounts remain untouched.

  • Switch back to the VS Code PowerShell terminal.
  • Press Ctrl+C to stop the Streamlit server.
  • Press the Windows key to open search.
  • Type File Explorer into the search field.
  • Press Enter to open File Explorer.
  • Navigate to the parent location of the project folder currently open in VS Code.
  • Select the project folder.
  • Press Delete to remove it.

What Gets Removed?

Deleting the project folder removes .venv plus every project file. The dataset cache remains separate.

  • Search your Windows user files for scikit_learn_data in File Explorer.
  • Select the scikit_learn_data folder shown in the results.
  • Press Delete to remove the cached dataset.
  • Confirm that the project folder no longer appears in its parent location.
  • Confirm that scikit_learn_data no longer appears in the File Explorer search results.

Nice Work!

Nice Work!

Excellent work! Your leakage-aware topic classifier turns pasted forum posts into readable predictions backed by an honest held-out evaluation.

You've learned how to:

  • Built a reproducible four-class natural language classifier with TF-IDF feeding a Complement Naive Bayes model.
  • Exposed metadata leakage through a held-out comparison of the baseline model against the cleaned model.
  • Created an interactive browser app with Streamlit so another person can paste text to see a readable prediction plus the cleaned test accuracy.
  • Secret Mission: Ranked every model-reported class probability from highest to lowest. Added a human-review warning below the project-defined threshold.

Ready to quiz yourself?