Predict Titanic Survival

Train a logistic regression model to predict Titanic passenger survival.

Introduction

30 Second Summary

A passenger manifest records one life per row. The patterns behind survival only become visible when those rows are compared.

In this project, you will build a local Python program that uses logistic regression to predict whether a Titanic passenger survived. A leakage-safe pipeline keeps the test rows unseen until the final accuracy check.

What You'll Build

Your finished script prints a clear verdict on its own performance, including a majority baseline, held-out model accuracy, and ten actual-versus-predicted outcomes.

By the end of this project, you'll have:

  • A runnable Titanic binary classifier that reads train.csv to predict passenger survival.
  • A leakage-safe scikit-learn preprocessing pipeline that handles blank ages plus text categories without a traceback.
  • A terminal accuracy report that shows whether the trained model beats a majority-class baseline. A ten-row prediction table connects that score to actual passenger outcomes.
  • Secret Mission: Measure the model with five-fold cross-validation to print five accuracy scores plus their mean.

Are there any prerequisites?

You need basic Python syntax plus a downloaded copy of the Kaggle Titanic train.csv file. The setup step checks Python on your Windows computer before creating an isolated environment in Visual Studio Code.

Before We Start

Before the hands-on work begins, this moment locks in your goal. You are predicting Titanic survival while measuring accuracy on passenger rows the model never trained on.

Set Up the Local ML Project

A model becomes hard to reproduce when its data moves between folders or its package versions drift. This setup gives every later result a stable starting point.

You'll use Visual Studio Code to keep the project files together. A virtual environment isolates the pinned pandas and scikit-learn packages from the rest of your computer.

In this step, get ready to:
  • Prepare the project folder with train.csv and requirements.txt.
  • Confirm a supported Python version in the Visual Studio Code terminal.
  • Create an isolated environment with the verified package versions.
Prepare the project folder

The project folder gives your dataset a fixed location. The dependency file records the package versions that the model expects.

  • Press the Windows key to open the search bar.
  • Type File Explorer into the search bar.

You'll see File Explorer in the search results.

  • Press Enter to open File Explorer.

You'll see a File Explorer window with your folders in the left sidebar.

  • Select Desktop in the left sidebar.

You'll see the files and folders stored on your Desktop.

  • Press Ctrl+Shift+N to create a folder.

You'll see a new folder with its name ready for editing.

  • Type titanic-logistic-regression as the folder name.
  • Press Enter to save the folder name.

You'll see the titanic-logistic-regression folder on your Desktop.

  • Return to the folder containing the downloaded train.csv file.
  • Select train.csv.

The selected train.csv file is highlighted.

  • Press Ctrl+C to copy train.csv.
  • Return to the titanic-logistic-regression folder on your Desktop.

You'll see the new project folder ready to receive the dataset.

  • Press Ctrl+V to paste the copied file.

You'll see train.csv inside titanic-logistic-regression.

  • Press the Windows key to return to the search bar.
  • Type Visual Studio Code into the search bar.

You'll see Visual Studio Code in the search results.

  • Press Enter to open Visual Studio Code.

You'll see the Visual Studio Code window.

  • Click File in the top menu.
  • Click Open Folder.

You'll see a folder picker.

  • Select the titanic-logistic-regression folder on your Desktop.
  • Click Select Folder.

You'll see train.csv in the Explorer sidebar.

Seeing a folder trust prompt?

  • Confirm the trust prompt because you created this folder yourself.

Visual Studio Code then opens the project files for editing.

The package pins belong in requirements.txt. This file makes the same dependency versions available whenever you recreate the environment.

  • Click the New File icon at the top of the Explorer sidebar.

You'll see a file name field inside the Explorer sidebar.

  • Type requirements.txt into the file name field.
  • Press Enter to create the file.

You'll see a blank requirements.txt editor tab.

  • Add the pinned dependencies to requirements.txt by pasting this code:
pandas==3.0.6
scikit-learn==1.9.1

What does this file control?

  • The first line pins pandas to 3.0.6.
  • The second line pins scikit-learn to 1.9.1.
  • Exact pins prevent a future package update from changing this project's behavior.
  • Save requirements.txt by pressing Ctrl+S.
  • Confirm requirements.txt appears beside train.csv in the Explorer sidebar.

Your project folder now contains the downloaded data plus its reproducible dependency list.

Don't see both files?

  • Confirm the Explorer heading shows TITANIC-LOGISTIC-REGRESSION.
  • Check that the dependency file is named exactly requirements.txt.
  • Return to File Explorer if train.csv was pasted into a different folder.

Help me check my project folder structure.

✔️ Awesome, I've got everything

Your dataset and dependency file are saved in the project folder.

ⓧ I'd like to double check the full code

The complete dependency file contains these two pins.

pandas==3.0.6
scikit-learn==1.9.1
Prepare the Python environment

Both pinned packages require Python 3.11 or newer. Checking first prevents the installation from failing against an unsupported interpreter.

  • Click Terminal in the Visual Studio Code top menu.
  • Click New Terminal.

You'll see an integrated terminal at the bottom of the project window.

  • Check the installed Python 3 version by running this command:
py -3 --version

What does this check prove?

The command asks the Windows Python launcher for its installed Python 3 version. The result determines whether the pinned packages can run in this project.

✔️ Python is supported

If the result shows Python 3.11 or newer, your interpreter is compatible with both pinned packages.

ⓧ Python is too old

An older interpreter cannot run the pinned package versions. Install Python 3.15.0 before creating the environment.

You'll see the Python installer in your browser's downloads.

  • Install Python 3.15.0 with the Windows installer.
  • Restart Visual Studio Code after the installation completes.

The restarted editor can now detect the new Python installation.

  • Return to the titanic-logistic-regression folder from earlier.
  • Open a new integrated terminal from the top menu.
  • Confirm the upgraded Python version by running:
py -3 --version

What should the new check confirm?

The terminal should now report Python 3.15.0. That version satisfies the package requirements.

ⓧ Command not found

Windows cannot currently find a Python 3 installation through its launcher. Install Python 3.15.0 before continuing.

You'll see the Python installer in your browser's downloads.

  • Install Python 3.15.0 with the Windows installer.
  • Restart Visual Studio Code after the installation completes.

The restarted editor can now detect Python through the Windows launcher.

  • Return to the titanic-logistic-regression folder from earlier.
  • Open a new integrated terminal from the top menu.
  • Confirm Python is available by running:
py -3 --version

What should the new check confirm?

The terminal should now report Python 3.15.0. The project can now create its isolated environment.

A supported interpreter can now create the project's isolated environment. The .venv folder keeps these packages separate from other Python projects.

  • Create the .venv environment by running this command:
python -m venv .venv

What does this command create?

  • Python creates an isolated environment inside the .venv folder.
  • The environment contains its own Python executable plus a separate package location.
  • Confirm the .venv folder appears in the Explorer sidebar.

The new folder proves that the isolated environment exists inside titanic-logistic-regression.

Don't see the environment folder?

  • Confirm the terminal is open inside the titanic-logistic-regression project.
  • Check the terminal output for a message mentioning a failure to create the environment.

Help me create the missing virtual environment.

  • Activate the new environment by running this command:
.\.venv\Scripts\activate

What does activation change?

Activation points the current terminal session at the Python executable inside .venv. Package installations now stay inside this project.

You'll see .venv at the start of the terminal prompt. That label confirms the isolated environment is active.

Environment not activating?

  • Confirm the Explorer sidebar contains the .venv folder.
  • Check that the terminal is using its PowerShell profile.
  • Retype the activation command with both backslashes before running it again.

Help me activate my Windows virtual environment.

Install the pinned packages

The active environment is empty until it reads requirements.txt. Installing from that file gives the classifier the exact library versions used throughout this project.

The first installation can take a minute while the packages download. A short pause in the terminal is expected.

  • Install the project dependencies by running this command in the activated terminal:
python -m pip install -r requirements.txt

What does this installation use?

  • The command reads each pinned package from requirements.txt.
  • The active environment receives pandas 3.0.6 plus scikit-learn 1.9.1.
  • The rest of your computer keeps its existing Python packages unchanged.

The terminal returns to its prompt after the dependencies finish installing.

Package installation failing?

  • Confirm .venv appears at the start of the terminal prompt.
  • Confirm requirements.txt is visible in the Explorer sidebar.
  • Check that your Python version is 3.11 or newer.

Help me diagnose the dependency installation.

Before you check, which two version numbers do you expect the activated environment to print?

  • Verify both installed package versions by running this command:
python -c "import pandas, sklearn; print(pandas.__version__, sklearn.__version__)"

What does this verification test?

  • Python imports both installed libraries from the active environment.
  • Each library reports its installed version.
  • A successful import proves the packages are available to the model script you build next.

You'll see 3.0.6 1.9.1 in the terminal. You cleared the setup hurdle with an isolated terminal running the exact package versions this classifier expects.

Versions missing or different?

  • Confirm .venv still appears at the start of the terminal prompt.
  • Compare requirements.txt with the full-code tab from earlier.
  • Repeat the package installation inside the activated environment if the versions differ.

Help me fix the package version check.

Your local machine learning workspace now has its data plus a reproducible Python environment. Next, you'll inspect the raw Titanic rows before testing what the model can handle.

Expose the Raw Data Problem

Your pinned local environment is ready. The next question is whether the Titanic passenger rows can move directly into a classifier.

The table combines numeric values with text categories. Several selected columns also contain missing values.

You will inspect the data before passing it directly to a logistic regression classifier from scikit-learn. The classifier's response tells you whether preprocessing is needed.

In this step, get ready to:
  • Inspect the downloaded passenger table.
  • Create a stratified held-out split.
  • Test raw passenger values with Logistic Regression.
Inspect the passenger data

Pandas loads the CSV file into a table called a DataFrame. Its shape reveals the row count plus the column count.

A focused preview confirms that the expected feature columns exist. Missing-value counts reveal which selected columns have incomplete rows.

  • In VS Code's file sidebar, click the page-shaped icon with a plus sign beside titanic-logistic-regression.
  • Type titanic_model.py as the file name.
  • Press Enter to create the file.
  • Add the dataset inspection code below:
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

FEATURES = ["Pclass", "Sex", "Age", "SibSp", "Parch", "Fare", "Embarked"]
TARGET = "Survived"


data = pd.read_csv("train.csv")
print(f"Rows and columns: {data.shape}")
print(data[FEATURES + [TARGET]].head())
print("\nMissing values in selected columns:")
print(data[FEATURES + [TARGET]].isna().sum())

What Does This Inspection Code Do?

  • The imports make pandas available as pd. They also prepare the classifier plus the splitting function used later.
  • FEATURES identifies the seven passenger columns used as model inputs.
  • TARGET identifies Survived as the label the model tries to predict.
  • pd.read_csv("train.csv") loads the downloaded file into data.
  • data.shape reports the number of rows plus the number of columns.
  • head() displays the first five rows from the selected columns.
  • isna().sum() counts missing values in each selected column.
  • Save titanic_model.py by pressing Ctrl+S.
  • Run the inspection in the activated VS Code terminal by using this command:
python titanic_model.py

What Does This Command Do?

The python command runs titanic_model.py with the interpreter from your active .venv environment.

You'll see Rows and columns: followed by the dataset shape. A five-row preview appears below it.

You'll also see missing-value counts for the selected columns. That is your first data checkpoint: pandas can read the expected Titanic file.

Can't See the Dataset Output?

  • Confirm that train.csv sits beside titanic_model.py inside titanic-logistic-regression.
  • Check that the data file is named exactly train.csv.
  • Confirm that the VS Code terminal still shows the active .venv environment.

Help me troubleshoot why my Titanic script cannot load or display train.csv.

Test a held-out raw-data split

A train-test split reserves some labeled rows for evaluation. Stratification keeps the survival class proportions similar across both groups.

A fixed random state makes the row assignment reproducible. This experiment sends the training portion directly into the classifier so its response can reveal whether the raw feature table is ready.

  • In titanic_model.py, place your cursor below print(data[FEATURES + [TARGET]].isna().sum()).
  • Add the held-out split plus the raw model attempt below:
X = data[FEATURES]
y = data[TARGET]

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.2,
    random_state=42,
    stratify=y,
)

model = LogisticRegression(max_iter=1000)
model.fit(X_train, y_train)

What Does This Experiment Do?

  • X contains the passenger feature table. y contains the survival labels.
  • test_size=0.2 reserves 20 percent of the labeled rows for testing.
  • stratify=y keeps the survival class proportions similar across the training set plus the test set.
  • random_state=42 makes the split reproducible.
  • LogisticRegression(max_iter=1000) creates the classifier with a higher iteration limit.
  • model.fit(X_train, y_train) asks the classifier to learn directly from the raw training rows.
  • Save titanic_model.py by pressing Ctrl+S.

Before you run the script, do you think the classifier can learn directly from every value in X_train?

  • Test the raw model attempt in the activated VS Code terminal by running:
python titanic_model.py

The Failure Is the Result

You'll see the dataset shape first. The selected-column preview follows.

The missing-value counts appear next. The traceback begins during model.fit(X_train, y_train).

Sex plus Embarked contain text categories. The selected feature table also includes missing values.

You found the exact raw-data boundary the model cannot cross. That planned failure gives the next preprocessing work a clear purpose.

No Traceback After the Preview?

  • Confirm that the final two lines create model before calling model.fit(X_train, y_train).
  • Check that FEATURES still includes Sex plus Embarked.
  • Save titanic_model.py before running it again.

Help me understand why my raw Titanic model does not produce the planned preprocessing traceback.

✔️ Awesome, I've got everything!

Your script prints the dataset inspection before reaching the planned raw-data failure.

ⓧ I'd like to double check the full code

The complete titanic_model.py file should match this reference.

import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

FEATURES = ["Pclass", "Sex", "Age", "SibSp", "Parch", "Fare", "Embarked"]
TARGET = "Survived"


data = pd.read_csv("train.csv")
print(f"Rows and columns: {data.shape}")
print(data[FEATURES + [TARGET]].head())
print("\nMissing values in selected columns:")
print(data[FEATURES + [TARGET]].isna().sum())

X = data[FEATURES]
y = data[TARGET]

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.2,
    random_state=42,
    stratify=y,
)

model = LogisticRegression(max_iter=1000)
model.fit(X_train, y_train)

How to Use This Reference

This block shows the cumulative file at the end of this step. Its final line deliberately sends the raw training features into the classifier.

Your raw-data experiment has exposed the model's input requirements. Next, you will place imputation, encoding, plus scaling inside a leakage-safe pipeline.

Build a Leakage-Safe Pipeline

Your raw training attempt exposed the project’s key barrier. Logistic Regression cannot train directly on Titanic’s text categories.

A preprocessing pipeline turns each feature group into numeric values. It also fills missing entries consistently.

The complete workflow fits on X_train before it touches X_test. This boundary prevents data leakage from making your accuracy look better than it really is.

In this step, get ready to:
  • Build the numeric preprocessing path.
  • Build the categorical preprocessing path.
  • Evaluate the combined classifier on held-out rows.
Clear the raw fit attempt

The final two lines still send unprocessed passenger values straight into the classifier. Removing that attempt gives you a clean place to build the working pipeline.

  • Delete both final raw-fit lines from titanic_model.py.
  • Save titanic_model.py.
  • Check the cleaned script by running this command in the activated terminal:
python titanic_model.py

What does this check prove?

The command runs the script after the failing fit attempt has been removed. The file should now finish after printing the data inspection.

You’ll see the dataset shape followed by the preview and missing-value counts. Good recovery. Your script now completes without the raw-model traceback.

Still seeing the preprocessing failure?

  • Return to the bottom of titanic_model.py.
  • Remove the standalone classifier assignment containing LogisticRegression(max_iter=1000).
  • Remove the standalone model.fit(X_train, y_train) call.

Help me remove the raw fitting attempt from my Titanic script.

Build the feature preprocessing paths

Numeric columns need missing-value imputation followed by feature scaling. Categorical columns need imputation followed by one-hot encoding.

Why a pipeline instead of manual preprocessing?

Manual preprocessing can accidentally learn from the test rows. A pipeline learns its imputation values and scaling rules only while fitting the training rows.

The same fitted rules transform every later row. That keeps training and prediction consistent.

  • Replace every line from import pandas as pd through TARGET = "Survived" with this updated header:
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler

FEATURES = ["Pclass", "Sex", "Age", "SibSp", "Parch", "Fare", "Embarked"]
NUMERIC_FEATURES = ["Age", "SibSp", "Parch", "Fare"]
CATEGORICAL_FEATURES = ["Pclass", "Sex", "Embarked"]
TARGET = "Survived"

What does this header prepare?

  • The new imports provide the transformers needed for missing values and feature conversion.
  • The Pipeline import lets several transformations run as one fitted workflow.
  • The NUMERIC_FEATURES constant identifies the columns that contain measured values.
  • The CATEGORICAL_FEATURES constant identifies the columns that represent groups.
  • Save titanic_model.py.
  • Check the updated header by running this command in the activated terminal:
python titanic_model.py

What does this run check?

This run confirms that Python can load every new import. It also confirms that both feature lists use valid names.

You’ll see the complete dataset inspection without a traceback. The new pipeline tools are ready to use.

Seeing an import or name problem?

  • Confirm the activated terminal still shows .venv in its prompt.
  • Compare every import in titanic_model.py with the header above.
  • Check the capitalization of NUMERIC_FEATURES.

Help me fix the updated imports or feature constants.

The numeric path fills each missing value with a median learned from the training rows. It then puts the numeric columns onto comparable scales.

  • Add this numeric pipeline below the closing parenthesis of the train_test_split block:
numeric_pipeline = Pipeline(
    steps=[
        ("imputer", SimpleImputer(strategy="median")),
        ("scaler", StandardScaler()),
    ]
)

How does the numeric path work?

  • The imputer step learns the median for each numeric feature.
  • The scaler step learns the center and spread of each imputed feature.
  • The ordered steps list guarantees that imputation happens before scaling.
  • Save titanic_model.py.
  • Check the numeric pipeline by running this command in the activated terminal:
python titanic_model.py

What does this run confirm?

The command verifies that the numeric pipeline is valid Python. The dataset inspection should still complete normally.

You’ll see the dataset output without a traceback. That proves the numeric transformation path is assembled correctly.

Numeric pipeline not running?

  • Check that numeric_pipeline sits below the complete train-test split.
  • Confirm that every opening parenthesis has a matching closing parenthesis.
  • Compare the spelling of SimpleImputer with its import.

Help me repair the numeric preprocessing pipeline.

The categorical path replaces missing categories with the most frequent training value. It then converts every category into numeric indicator columns.

  • Add this categorical pipeline directly below numeric_pipeline in titanic_model.py.
categorical_pipeline = Pipeline(
    steps=[
        ("imputer", SimpleImputer(strategy="most_frequent")),
        ("onehot", OneHotEncoder(handle_unknown="ignore")),
    ]
)

How does the categorical path work?

  • The categorical imputer learns the most frequent value for each selected column.
  • The onehot step converts text categories into binary columns.
  • The handle_unknown="ignore" setting lets later rows contain categories that were absent during fitting.
  • Save titanic_model.py.
  • Check the categorical pipeline by running this command in the activated terminal:
python titanic_model.py

What does this run confirm?

This run verifies the categorical pipeline’s structure. The script should complete its dataset inspection without reaching a fitting error.

You’ll see the familiar dataset output without a traceback. Both feature-specific preprocessing paths now load successfully.

Categorical pipeline not running?

  • Check that categorical_pipeline appears after the complete numeric pipeline.
  • Confirm that most_frequent uses an underscore.
  • Compare the spelling of OneHotEncoder with its import.

Help me repair the categorical preprocessing pipeline.

Assemble and evaluate the pipeline

The two preprocessing paths need a router that sends each column to the correct transformer. A ColumnTransformer provides that routing.

  • Add this preprocessor directly below categorical_pipeline in titanic_model.py.
preprocessor = ColumnTransformer(
    transformers=[
        ("numeric", numeric_pipeline, NUMERIC_FEATURES),
        ("categorical", categorical_pipeline, CATEGORICAL_FEATURES),
    ]
)

How are the columns routed?

  • The numeric route sends NUMERIC_FEATURES through numeric_pipeline.
  • The categorical route sends CATEGORICAL_FEATURES through categorical_pipeline.
  • The transformer combines both numeric outputs into one model-ready feature table.
  • Save titanic_model.py.
  • Check the column routing by running this command in the activated terminal:
python titanic_model.py

What does this run verify?

The command checks that both feature lists connect to valid pipelines. A successful run proves that the transformer definition is ready for model fitting.

You’ll see the dataset inspection complete without a traceback. The numeric and categorical routes now belong to one preprocessor.

Column routing not working?

  • Confirm that both transformer entries sit inside the transformers list.
  • Check that NUMERIC_FEATURES uses numeric_pipeline.
  • Check that CATEGORICAL_FEATURES uses categorical_pipeline.

Help me correct the ColumnTransformer routes.

The outer pipeline treats preprocessing and classification as one estimator. Calling fit on this object learns every transformation from the same training split.

  • Add this model pipeline directly below preprocessor in titanic_model.py.
model = Pipeline(
    steps=[
        ("preprocessor", preprocessor),
        ("classifier", LogisticRegression(max_iter=1000)),
    ]
)

How does the model pipeline work?

  • The preprocessor step converts the raw passenger columns into model-ready values.
  • The classifier step trains Logistic Regression on those transformed values.
  • The ordered steps keep preprocessing inside every future fit or prediction.
  • Save titanic_model.py.
  • Check the assembled model pipeline by running this command in the activated terminal:
python titanic_model.py

What does this run verify?

The command confirms that the outer pipeline accepts the complete preprocessor. The model definition should load without starting training yet.

You’ll see the dataset inspection finish normally. Your leakage-safe estimator is assembled and ready to fit.

Model pipeline not loading?

  • Confirm that model appears after the complete preprocessor definition.
  • Check that the first model step references preprocessor.
  • Check that the classifier uses LogisticRegression(max_iter=1000).

Help me fix the outer model pipeline.

Held-out evaluation trains the complete workflow on X_train. It measures prediction accuracy only against y_test.

  • Add this evaluation code directly below the model pipeline:
model.fit(X_train, y_train)
model_predictions = model.predict(X_test)
model_accuracy = accuracy_score(y_test, model_predictions)

print(f"Model accuracy: {model_accuracy:.3f}")

How is held-out accuracy measured?

  • The fit call learns preprocessing rules from X_train.
  • The predict call applies those fitted rules to X_test.
  • The accuracy_score call calculates the fraction of correct held-out predictions.
  • The final line prints the result with three decimal places.
  • Save titanic_model.py.

Before you run the completed script, where do you expect its accuracy to fall between 0 and 1?

  • Test the completed workflow by running this command in the activated terminal:
python titanic_model.py

What does the final run test?

The command loads the passenger data before fitting the complete pipeline. It then evaluates predictions made for rows held out from training.

You’ll see the dataset inspection followed by Model accuracy: and a value between 0 and 1. That’s the failed raw model transformed into a working held-out evaluation.

Still seeing a traceback?

  • Confirm that model.fit(X_train, y_train) appears after the complete model pipeline.
  • Check that model.predict(X_test) uses the test features.
  • Check that accuracy_score(y_test, model_predictions) uses the test labels.

Help me debug the completed leakage-safe pipeline.

✔️ Awesome, I've got everything!

Great. Your saved script now trains the leakage-safe pipeline and prints held-out model accuracy.

ⓧ I'd like to double check the full code

Compare your complete titanic_model.py file with this version:

import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler

FEATURES = ["Pclass", "Sex", "Age", "SibSp", "Parch", "Fare", "Embarked"]
NUMERIC_FEATURES = ["Age", "SibSp", "Parch", "Fare"]
CATEGORICAL_FEATURES = ["Pclass", "Sex", "Embarked"]
TARGET = "Survived"


data = pd.read_csv("train.csv")
print(f"Rows and columns: {data.shape}")
print(data[FEATURES + [TARGET]].head())
print("\nMissing values in selected columns:")
print(data[FEATURES + [TARGET]].isna().sum())

X = data[FEATURES]
y = data[TARGET]

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.2,
    random_state=42,
    stratify=y,
)

numeric_pipeline = Pipeline(
    steps=[
        ("imputer", SimpleImputer(strategy="median")),
        ("scaler", StandardScaler()),
    ]
)

categorical_pipeline = Pipeline(
    steps=[
        ("imputer", SimpleImputer(strategy="most_frequent")),
        ("onehot", OneHotEncoder(handle_unknown="ignore")),
    ]
)

preprocessor = ColumnTransformer(
    transformers=[
        ("numeric", numeric_pipeline, NUMERIC_FEATURES),
        ("categorical", categorical_pipeline, CATEGORICAL_FEATURES),
    ]
)

model = Pipeline(
    steps=[
        ("preprocessor", preprocessor),
        ("classifier", LogisticRegression(max_iter=1000)),
    ]
)

model.fit(X_train, y_train)
model_predictions = model.predict(X_test)
model_accuracy = accuracy_score(y_test, model_predictions)

print(f"Model accuracy: {model_accuracy:.3f}")

How should you compare the file?

Check each section from top to bottom. Pay particular attention to indentation inside every pipeline and transformer.

Your classifier now handles mixed passenger data without leaking test information into training. Next, you’ll give its accuracy a simple reference point.

Compare Model Predictions

Your leakage-safe pipeline now trains on the Titanic rows without a traceback. It reports held-out accuracy on data it did not train on.

An accuracy number becomes useful when it has a simple reference. You will compare your logistic regression model with a majority-class baseline before connecting its score to ten concrete predictions.

In this step, get ready to:
  • Benchmark the logistic regression model against a majority-class baseline.
  • Build a ten-row table of actual and predicted survival labels.
  • Run the finished script to inspect the complete evaluation report.
Add a majority-class baseline

A baseline shows the accuracy available from a simple strategy. DummyClassifier creates that reference by predicting the most frequent training label for every test row.

  • In titanic_model.py, locate the import from sklearn.compose.
  • Add the new classifier import directly below it by copying this line:
from sklearn.dummy import DummyClassifier

What does this import do?

This import makes DummyClassifier available to your script. The classifier gives your trained model a simple benchmark to beat.

  • Find the line that starts with model.fit(X_train, y_train).
  • Insert the baseline evaluation directly above that line by copying this code:
baseline = DummyClassifier(strategy="most_frequent")
baseline.fit(X_train, y_train)
baseline_predictions = baseline.predict(X_test)
baseline_accuracy = accuracy_score(y_test, baseline_predictions)

How does the baseline work?

  • The most_frequent strategy always chooses the most common label learned from y_train.
  • The fit call finds that majority label without using any passenger features.
  • The predict call applies the majority label to every row in X_test.
  • The accuracy_score call measures the baseline against the same y_test labels used for your model.
  • Scroll to the final line in titanic_model.py.
  • Confirm the current reporting line looks like this:
print(f"Model accuracy: {model_accuracy:.3f}")

What does the current report show?

The current report only prints the trained model's accuracy. It does not provide a reference for deciding whether that result is useful.

  • Replace the current reporting line with these two lines:
print(f"\nBaseline accuracy: {baseline_accuracy:.3f}")
print(f"Model accuracy: {model_accuracy:.3f}")

What changes in the report?

  • The first line prints the accuracy from always choosing the majority class.
  • The second line prints the accuracy from your trained logistic regression pipeline.
  • Save titanic_model.py by pressing Ctrl+S.

Before you run this, predict which accuracy you expect to be higher.

  • Return to the activated VS Code terminal from the previous step.
  • Run the updated classifier by entering this command:
python titanic_model.py

What should you see?

You should see one line for Baseline accuracy: followed by one line for Model accuracy:. Both values use the same held-out test labels.

Good progress. Your model accuracy now has a fair reference point.

Does the baseline report fail?

  • Confirm the DummyClassifier import sits directly below the ColumnTransformer import.
  • Confirm the baseline block appears above model.fit(X_train, y_train).
  • Confirm the terminal still shows (.venv) before its prompt.

Help me debug my baseline comparison.

Inspect ten predictions

A score summarizes every test prediction into one number. A pandas table lets you inspect where the predicted survival labels agree with the actual labels.

  • In titanic_model.py, locate the line that calculates model_accuracy.
  • Insert the prediction comparison directly below that line by copying this code:
comparison = pd.DataFrame(
    {
        "Actual": y_test.to_numpy(),
        "Predicted": model_predictions,
    }
).head(10)

How is the comparison built?

  • The to_numpy() call converts the held-out labels into values for the table.
  • The Actual column holds the labels from y_test.
  • The Predicted column holds the labels produced by the model.
  • The head(10) call limits the report to ten rows.
  • Scroll to the two reporting lines at the bottom of titanic_model.py.
  • Confirm they currently look like this:
print(f"\nBaseline accuracy: {baseline_accuracy:.3f}")
print(f"Model accuracy: {model_accuracy:.3f}")

What is still missing?

These lines report both accuracy values. The new comparison table still needs its own output lines.

  • Replace the two reporting lines with the complete output block below:
print(f"\nBaseline accuracy: {baseline_accuracy:.3f}")
print(f"Model accuracy: {model_accuracy:.3f}")
print("\nSample predictions:")
print(comparison.to_string(index=False))

What does the final output include?

  • The first two lines print the baseline accuracy followed by the model accuracy.
  • The third line labels the sample prediction section.
  • The final line prints the table without a pandas row index.
  • Save titanic_model.py by pressing Ctrl+S.

Before you run this, predict what agreement and disagreement will look like across the ten rows.

  • Return to the activated VS Code terminal from earlier.
  • Run the completed report by entering this command:
python titanic_model.py

What should the final report show?

  • You should see the dataset shape followed by its preview and missing-value counts.
  • You should see a value after Baseline accuracy:.
  • You should see a value after Model accuracy:.
  • You should see ten rows beneath Sample predictions: with Actual and Predicted columns.

Strong finish. Your classifier now reports performance against a baseline and connects that score to passenger-level predictions.

Missing the prediction table?

  • Confirm the comparison block appears after model_accuracy is calculated.
  • Confirm the final line uses comparison.to_string(index=False).
  • Confirm the closing parenthesis after head(10) remains in place.

Help me fix my missing prediction table.

✔️ My report matches

Your completed script is saved. It now prints both accuracy values and ten sample predictions.

ⓧ I'd like to double check the full code

  • Compare your titanic_model.py with the complete reference below.
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.dummy import DummyClassifier
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler

FEATURES = ["Pclass", "Sex", "Age", "SibSp", "Parch", "Fare", "Embarked"]
NUMERIC_FEATURES = ["Age", "SibSp", "Parch", "Fare"]
CATEGORICAL_FEATURES = ["Pclass", "Sex", "Embarked"]
TARGET = "Survived"


data = pd.read_csv("train.csv")
print(f"Rows and columns: {data.shape}")
print(data[FEATURES + [TARGET]].head())
print("\nMissing values in selected columns:")
print(data[FEATURES + [TARGET]].isna().sum())

X = data[FEATURES]
y = data[TARGET]

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.2,
    random_state=42,
    stratify=y,
)

numeric_pipeline = Pipeline(
    steps=[
        ("imputer", SimpleImputer(strategy="median")),
        ("scaler", StandardScaler()),
    ]
)

categorical_pipeline = Pipeline(
    steps=[
        ("imputer", SimpleImputer(strategy="most_frequent")),
        ("onehot", OneHotEncoder(handle_unknown="ignore")),
    ]
)

preprocessor = ColumnTransformer(
    transformers=[
        ("numeric", numeric_pipeline, NUMERIC_FEATURES),
        ("categorical", categorical_pipeline, CATEGORICAL_FEATURES),
    ]
)

model = Pipeline(
    steps=[
        ("preprocessor", preprocessor),
        ("classifier", LogisticRegression(max_iter=1000)),
    ]
)

baseline = DummyClassifier(strategy="most_frequent")
baseline.fit(X_train, y_train)
baseline_predictions = baseline.predict(X_test)
baseline_accuracy = accuracy_score(y_test, baseline_predictions)

model.fit(X_train, y_train)
model_predictions = model.predict(X_test)
model_accuracy = accuracy_score(y_test, model_predictions)

comparison = pd.DataFrame(
    {
        "Actual": y_test.to_numpy(),
        "Predicted": model_predictions,
    }
).head(10)

print(f"\nBaseline accuracy: {baseline_accuracy:.3f}")
print(f"Model accuracy: {model_accuracy:.3f}")
print("\nSample predictions:")
print(comparison.to_string(index=False))

What should match?

Check the import order first. Then check that the baseline runs before the model while the comparison table is built after both model metrics exist.

Secret mission

Measure Accuracy Across Five Folds

One held-out split can make your classifier look unusually strong or weak. Measure the complete pipeline across five folds to get a more stable view of its accuracy.

Clean Up Your Resources

Clean Up Your Resources

Your classifier runs locally in Visual Studio Code with no ongoing cloud costs. Decide whether to keep your resources available, pause your work, or delete the project entirely.

Resources you used:

  • The local titanic-logistic-regression folder containing titanic_model.py, requirements.txt, and the copied train.csv file.
  • The .venv virtual environment containing pandas 3.0.6 and scikit-learn 1.9.1.

Keep everything running

No action needed. Choose this if you want to rerun the classifier or experiment with its features.

Your script, dataset copy, dependencies, held-out results, and five-fold evaluation remain available without creating ongoing costs.

Pause - I'll come back to this later

Leave the active virtual environment to pause your work while keeping every project file on your machine.

  • Return to the VS Code terminal from earlier.
  • Leave the virtual environment by running this command:
deactivate

What does this command do?

This removes the virtual environment from your current terminal session. The .venv folder stays ready for your next session.

You can close Visual Studio Code after the environment name disappears from the terminal prompt.

Delete - I don't want to use this again

Remove the local project resources if you want to start again from an empty folder.

  • Return to the VS Code terminal from earlier.
  • Leave the active virtual environment by running this command:
deactivate

What does this command do?

This releases the virtual environment from the current terminal session. Your files remain in place until the next command removes their folder.

This permanently removes only the local titanic-logistic-regression folder.

  • Delete the project folder by running these PowerShell commands:
cd ..
Remove-Item -Recurse -Force titanic-logistic-regression

What do these commands remove?

  • The first command moves the terminal to the folder containing your project folder.
  • The second command removes the project files, the copied dataset, and the entire virtual environment.
  • Check the VS Code Explorer sidebar. You should no longer see files inside titanic-logistic-regression.

Nice Work!

Nice Work!

You did it! You built a local logistic regression classifier that predicts survival from Titanic passenger data. Your evaluation now shows how the model performs beyond the rows used for training.

You've learned how to:

  • Build a leakage-safe pipeline in scikit-learn. It imputes missing values. It scales numeric columns. It one-hot encodes categorical columns.
  • Measure held-out accuracy on unseen rows. Compare it with a majority-class baseline to judge the model's value.
  • Inspect the ten-row comparison report to connect each prediction with its actual passenger outcome.
  • Complete the Secret Mission by measuring accuracy with five-fold cross-validation. Keep preprocessing inside the model for every fold.

Ready to quiz yourself?