Build a Medical Cost Predictor
Build a scikit-learn pipeline to predict medical insurance costs from CSV data.
Introduction
30 Second Summary
Medical bills can vary wildly between people who seem similar on paper. A prediction only earns trust when its mistakes are visible.
In this project, you will build a command-line medical insurance cost predictor with scikit-learn. You will use supervised regression to compare three modeling approaches against unseen test data.
What You'll Build
The finished Python script names the best model from unseen test rows while a saved four-panel figure reveals where its predictions miss.
By the end of this project, you'll have:
- A data-quality check that confirms the CSV has the expected seven columns before any model trains.
- A fair model scoreboard that compares the Mean baseline, Linear regression, and Random forest against the same unseen test rows.
- A saved four-panel diagnostic figure at outputs/model_diagnostics.png that exposes the prediction errors of both trained models.
- Secret Mission: A five-fold cross-validation report that reveals whether the Random forest performs consistently across different held-out subsets.
Are there any prerequisites?
You need a Windows machine with Python 3.11 or newer plus Visual Studio Code. A Kaggle account handles the manual dataset download. Comfort with Python and pandas lets you focus on modeling.
Before We Start
You are building a model that predicts medical insurance charges from six personal and demographic features. Before any hands-on work, lock in why held-out test rows must stay unseen during training so the final evaluation stays fair.
Set Up the Local Project
Reliable regression starts with a stable dataset path. Exact package versions keep the experiment consistent across machines.
This step builds that foundation with a real CSV from Kaggle inside Visual Studio Code. A dedicated virtual environment keeps the project's scikit-learn tools isolated.
In this step, get ready to:
- Place the Kaggle CSV at data/medical-charges.csv.
- Pin the project packages in requirements.txt.
- Prepare a Python 3.11 or newer virtual environment with the pinned packages.
Download and organize the real dataset
The Medical Insurance Cost Prediction dataset contains the six input features used by this project. Its exact filename lets the Python script find the data without a custom path.
- Open your web browser.
- Visit Kaggle's Medical Insurance Cost Prediction dataset page.
- Sign in to your Kaggle account.
- Locate medical-charges.csv on the dataset page.
- Use the browser download control beside medical-charges.csv.
Your browser places the downloaded CSV or its archive in your Downloads folder.
- Press Windows+E to open File Explorer.
- Select Downloads in the left navigation pane.
- Right-click the downloaded package if its filename ends in .zip.
- Select Extract All... if the package is zipped.
- Follow the extraction instructions.
The extracted folder contains medical-charges.csv. Next, give the dataset a predictable home on your Desktop.
- Select Desktop in the File Explorer navigation pane.
- Select New.
- Select Folder.
- Type insurance-regression as the folder name.
- Press Enter.
- Double-click insurance-regression.
The empty insurance-regression folder is your project root. Two child folders keep the input CSV separate from the diagnostic images you create later.
- Select New.
- Select Folder.
- Type data as the folder name.
- Press Enter.
- Select New again.
- Select Folder.
- Type outputs as the folder name.
- Press Enter.
Your project now has a place for its source data plus an empty destination for generated plots.
- Return to Downloads in File Explorer.
- Select medical-charges.csv.
- Select Cut.
- Return to the data folder inside insurance-regression.
- Select Paste.
The dataset now lives at data/medical-charges.csv. That exact relative path is what the modeling script uses.
- Press the Windows key to open Windows search.
- Type Visual Studio Code and press Enter to open it.
- Click File in the top menu bar.
- Click Open Folder....
- Select the insurance-regression folder on your Desktop.
- Click Select Folder.
Seeing Workspace Trust?
This folder contains the files you just created. Select Yes, I trust the authors to enable the editor and terminal features used in this project.
- Expand data in the Explorer sidebar.
- Confirm that medical-charges.csv appears inside data.
- Expand outputs in the Explorer sidebar.
- Confirm that outputs contains no files.
Good progress. VS Code now shows the real dataset in the exact location your regression workflow expects.
Dataset file missing?
Check that the filename is exactly medical-charges.csv. Make sure the file is inside data instead of beside that folder.
If you downloaded an archive, confirm that you moved the extracted CSV. The modeling script cannot read the compressed package directly.
help me place my Kaggle CSV in the correct project folder
Pin the project dependencies
A dependency pin records the exact package release used by the project. Your requirements.txt file makes the environment reproducible.
- Hover over the INSURANCE-REGRESSION heading in the Explorer sidebar.
- Click the New File icon.
- Type requirements.txt and press Enter.
- Add the pinned dependencies by copying this file content:
matplotlib==3.11.2
pandas==3.0.6
scikit-learn==1.9.1
What does this file control?
- The matplotlib==3.11.2 line pins Matplotlib for the diagnostic plots.
- The pandas==3.0.6 line pins pandas for loading the CSV.
- The scikit-learn==1.9.1 line pins the modeling library used throughout the experiment.
- Save requirements.txt by pressing Ctrl+S.
- Confirm that requirements.txt appears beside the data folder in the Explorer sidebar.
✔️ Awesome, I've got everything!
Your dependency pins are saved. Keep requirements.txt at the top level of insurance-regression.
ⓧ I'd like to double check the full code
Compare your complete requirements.txt file with this reference:
matplotlib==3.11.2
pandas==3.0.6
scikit-learn==1.9.1
What should match?
Each package name uses lowercase text. Each line uses two equals signs before its pinned version.
These package releases require Python 3.11 or newer. Check the interpreter before creating the environment so incompatible packages do not interrupt the installation.
- Click Terminal in the VS Code menu bar.
- Click New Terminal.
- Select the terminal profile dropdown in the terminal panel.
- Select Command Prompt.
The Command Prompt terminal starts inside insurance-regression because that folder is open in VS Code.
- Check the Python interpreter available to this terminal by running:
python --version
What does this command check?
The command reports the Python interpreter selected by this Command Prompt session. The project needs version 3.11 or newer for all three pinned packages.
✔️ I see version 3.11 or higher
Your Python interpreter meets the package requirement. You can create the isolated environment with this interpreter.
ⓧ I see an older version
The listed interpreter is below the project's minimum version. Update Python before creating the environment.
- Visit the official Python downloads page for Windows.
- Download a current 64-bit Python 3 release.
- Complete the Python installer.
- Restart Visual Studio Code after the installation finishes.
- Repeat the version check above in a new Command Prompt terminal.
ⓧ Command not found
Windows cannot currently find a Python interpreter from this terminal. Install Python before continuing.
- Visit the official Python downloads page for Windows.
- Download a current 64-bit Python 3 release.
- Complete the Python installer.
- Restart Visual Studio Code after the installation finishes.
- Repeat the version check above in a new Command Prompt terminal.
A virtual environment gives this project its own package space. That boundary protects the pinned releases from packages used by unrelated Python projects.
- Create the sklearn-env virtual environment by running:
python -m venv sklearn-env
What does this command create?
Python creates a sklearn-env directory inside insurance-regression. The directory contains the isolated interpreter plus its package installation location.
You should see sklearn-env appear in the VS Code Explorer after the command finishes.
- Activate the new environment in Command Prompt by running:
sklearn-env\Scripts\activate
What does activation change?
Activation points python plus pip at the environment inside sklearn-env. Packages installed next stay attached to this project.
Your Command Prompt now includes sklearn-env at the start of its prompt. That visible label confirms the environment is active.
Environment did not activate?
Confirm that the terminal profile says Command Prompt. The activation command in this project is written for that shell.
Check that sklearn-env appears at the top level of the VS Code Explorer. Create the environment again if the folder is missing.
help me activate my Windows virtual environment
Install and verify the environment
The active environment is ready to receive the versions recorded in requirements.txt. The installation connects the reproducible file to the interpreter that runs the project.
The first installation can take a few minutes while the packages download. A stream of package names in the terminal means the process is moving forward.
- Install the pinned dependencies into sklearn-env by running:
python -m pip install -r requirements.txt
What does this command install?
Python runs the pip installer from the active environment. The installer reads each exact package pin from requirements.txt.
The command finishes after installing Matplotlib, pandas, scikit-learn, plus the supporting packages they require.
The completed command returns control to your sklearn-env prompt. Your isolated machine-learning environment is now populated.
Package installation failed?
Confirm that the terminal prompt starts with sklearn-env. An inactive environment sends the installation to a different Python location.
Run the Python version check again if the output mentions package compatibility. The pinned releases require Python 3.11 or newer.
help me diagnose my pinned package installation
- List the packages installed in the active environment by running:
python -m pip freeze
What does this command confirm?
The command prints the installed package names plus their resolved versions. This gives you direct evidence that the active environment contains the requested releases.
You should find matplotlib==3.11.2, pandas==3.0.6, plus scikit-learn==1.9.1 in the output.
Before the final check, which scikit-learn version do you expect the environment report to show?
- Display the complete scikit-learn environment report by running:
python -c "import sklearn; sklearn.show_versions()"
What does this verification prove?
Python imports scikit-learn from the active environment. The version report identifies the installed scikit-learn release plus its runtime dependencies.
You should see scikit-learn 1.9.1 in the report. This confirms that the pinned modeling library can be imported successfully.
- Return to the Explorer sidebar in VS Code.
- Confirm that data/medical-charges.csv is present.
- Confirm that outputs/ is empty.
- Confirm that requirements.txt contains the three pinned dependencies.
- Keep the activated Command Prompt terminal open for the remaining steps.
That's the setup complete. Your dataset path is fixed and your isolated environment can import the required scikit-learn version.
Version report missing scikit-learn 1.9.1?
Check that the terminal prompt still includes sklearn-env. Activate the environment again if the label is missing.
Compare requirements.txt with the full-file reference above. A missing equals sign can install an unintended release.
help me verify my scikit-learn environment
Your local project is ready for modeling. Next, you will load the CSV, protect an unseen test set, and establish the mean baseline every real model must beat.
Load the Data and Beat the Mean
Your verified environment is ready. The downloaded CSV is waiting inside the existing data folder.
Before a model can earn trust, you need to confirm that the dataset has the expected structure. You also need a protected train/test split that keeps evaluation rows away from training.
A mean-only baseline model gives you a visible first benchmark. Running it exposes how little a prediction can achieve when it ignores every feature.
In this step, get ready to:
- Load the medical charges CSV with pandas.
- Protect an unseen test set with a reproducible split.
- Measure a mean-only predictor with regression metrics.
Load and inspect the CSV
A reliable data-loading function checks the file before reading it. This gives failures a clear location inside regression_project.py.
- Create regression_project.py inside the existing insurance-regression folder using the file creation control in the Visual Studio Code file sidebar.
- Build the first runnable data-loading slice by adding this code to regression_project.py:
from pathlib import Path
import pandas as pd
from sklearn.dummy import DummyRegressor
from sklearn.metrics import mean_absolute_error, r2_score, root_mean_squared_error
from sklearn.model_selection import train_test_split
DATA_PATH = Path("data") / "medical-charges.csv"
def load_and_check_data(path):
if not path.exists():
raise FileNotFoundError(
f"Could not find {path}. Place medical-charges.csv inside the data folder."
)
data = pd.read_csv(path)
print(f"Dataset shape: {data.shape}")
print("\nFirst five rows:")
print(data.head())
return data
What does this code do?
- Path builds a Windows-compatible route to data/medical-charges.csv.
- load_and_check_data() stops with a focused message when the CSV is missing.
- pd.read_csv() loads the CSV into a DataFrame.
- data.head() provides a quick view of the first five records.
- Add the first version of the script entry point beneath load_and_check_data() using this code:
def main():
data = load_and_check_data(DATA_PATH)
if __name__ == "__main__":
main()
How does the script start?
main() calls the loader with DATA_PATH. The final condition runs main() when you launch this file as a script.
- Save regression_project.py.
- Predict which dataset facts the first run can reveal.
- Return to the activated sklearn-env terminal from the previous step.
- Run the first data-loading slice with this command:
python regression_project.py
What should I see?
You should see Dataset shape: followed by the dataset dimensions. The first five rows should appear underneath.
Your first data read works. The script can now find the downloaded CSV from the project root.
Can't find the CSV?
Confirm that the activated terminal is still at the insurance-regression project root. The script resolves data/medical-charges.csv from that location.
Check that the filename is exactly medical-charges.csv. Windows may hide a second file extension.
Help me diagnose the missing CSV.
Reading rows proves that the path works. A schema check goes further by confirming that the CSV contains exactly the seven columns your model expects.
- Add the remaining project constants directly below DATA_PATH using this code:
OUTPUT_PATH = Path("outputs") / "model_diagnostics.png"
TARGET = "charges"
RANDOM_STATE = 42
EXPECTED_COLUMNS = [
"age",
"sex",
"bmi",
"children",
"smoker",
"region",
"charges",
]
What do these constants control?
- OUTPUT_PATH reserves the location used for diagnostic images later in the project.
- TARGET identifies charges as the value the model predicts.
- RANDOM_STATE stores the fixed value used for reproducible splitting.
- EXPECTED_COLUMNS records the complete dataset schema.
- Replace the current load_and_check_data() function with this complete version:
def load_and_check_data(path):
if not path.exists():
raise FileNotFoundError(
f"Could not find {path}. Place medical-charges.csv inside the data folder."
)
data = pd.read_csv(path)
missing_columns = set(EXPECTED_COLUMNS) - set(data.columns)
unexpected_columns = set(data.columns) - set(EXPECTED_COLUMNS)
if missing_columns or unexpected_columns:
raise ValueError(
"Dataset columns do not match the expected schema. "
f"Missing: {sorted(missing_columns)}. "
f"Unexpected: {sorted(unexpected_columns)}."
)
print(f"Dataset shape: {data.shape}")
print("\nFirst five rows:")
print(data.head())
print("\nData types:")
print(data.dtypes)
print("\nMissing values by column:")
print(data.isna().sum())
print(f"\nDuplicate rows: {data.duplicated().sum()}")
return data
How does the full check protect the project?
- missing_columns finds required names that are absent from the CSV.
- unexpected_columns finds names outside the expected schema.
- data.dtypes shows how pandas interpreted each column.
- data.isna().sum() counts missing values by column.
- data.duplicated().sum() counts repeated rows.
- Save regression_project.py.
- Predict whether the schema check accepts the downloaded CSV.
- Rerun the expanded data check with this command:
python regression_project.py
What should the data check show?
You should see 1,338 rows across seven columns. The output should list data types plus missing-value counts for every expected column.
The missing-value counts should all be zero. The duplicate-row count appears at the end of the report.
Seeing a schema error?
Read the missing plus unexpected column lists in the error. They identify which CSV headers differ from EXPECTED_COLUMNS.
Confirm that you downloaded the medical-charges.csv file from the Medical Insurance Cost Prediction dataset.
Help me compare my CSV schema.
Protect the test set
The six input columns belong in X. The medical charges target belongs in y.
The split reserves 20 percent of the rows for testing. A fixed random state recreates the same row assignment on every run.
- Add split_data() between load_and_check_data() and main() using this code:
def split_data(data):
X = data.drop(columns=[TARGET])
y = data[TARGET]
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=RANDOM_STATE,
)
print(f"\nTraining rows: {len(X_train)}")
print(f"Test rows: {len(X_test)}")
return X_train, X_test, y_train, y_test
What does the split do?
- data.drop(columns=[TARGET]) creates the feature table without charges.
- data[TARGET] creates the prediction target.
- test_size=0.2 reserves one fifth of the rows for final checks.
- random_state=RANDOM_STATE keeps the split reproducible.
- Replace the current main() function with this version:
def main():
data = load_and_check_data(DATA_PATH)
X_train, X_test, y_train, y_test = split_data(data)
How does the split enter the workflow?
main() passes the validated DataFrame into split_data(). The four returned objects keep features plus targets separated across training rows plus test rows.
- Save regression_project.py.
- Predict how 1,338 rows divide under the 80/20 split.
- Run the script to reveal the row counts:
python regression_project.py
What should the split show?
You should see 1,070 training rows. You should also see 268 test rows.
Those counts prove that the held-out rows exist before any model is fitted.
Row counts missing?
Confirm that main() calls split_data(data). The row-count messages come from that function.
Check the indentation inside split_data(). Each split statement must remain inside the function body.
Help me debug the train/test split.
Beat the mean
A DummyRegressor with a mean strategy learns one value from y_train. It uses that same constant for every prediction.
Your evaluation function measures average error through MAE. It also reports RMSE plus R2 for a broader view of model performance.
- Add evaluate_model() between split_data() and main() using this code:
def evaluate_model(name, model, X_train, X_test, y_train, y_test):
model.fit(X_train, y_train)
train_predictions = model.predict(X_train)
test_predictions = model.predict(X_test)
metrics = {
"Model": name,
"Train MAE": mean_absolute_error(y_train, train_predictions),
"Test MAE": mean_absolute_error(y_test, test_predictions),
"Test RMSE": root_mean_squared_error(y_test, test_predictions),
"Train R2": r2_score(y_train, train_predictions),
"Test R2": r2_score(y_test, test_predictions),
}
return metrics, test_predictions
What do the metrics reveal?
- Train MAE measures the model's average miss on rows used for fitting.
- Test MAE measures the average miss on unseen rows in charge units.
- Test RMSE gives larger misses more influence.
- Train R2 plus Test R2 compare the model with constant mean prediction.
- Replace the current main() function with this complete baseline workflow:
def main():
data = load_and_check_data(DATA_PATH)
X_train, X_test, y_train, y_test = split_data(data)
metrics, _ = evaluate_model(
"Mean baseline",
DummyRegressor(strategy="mean"),
X_train,
X_test,
y_train,
y_test,
)
print("\nMean baseline:")
print(pd.DataFrame([metrics]).round(2).to_string(index=False))
How is the baseline trained?
evaluate_model() fits the mean predictor using only the training rows. It then evaluates predictions separately on the training set plus the protected test set.
The final DataFrame turns the metric dictionary into one rounded comparison row. This row becomes the benchmark for every later model.
- Save regression_project.py.
- Predict whether one constant prediction can explain differences among people.
- Run the completed baseline workflow with this command:
python regression_project.py
What should the completed run show?
You should see the dataset shape, expected data checks, training row count, plus test row count. A Mean baseline metrics row should follow.
The row includes train MAE, test MAE, test RMSE, train R2, plus test R2. Train R2 should be zero because the model predicts the training-target mean.
The run succeeds, yet every person receives the same predicted charge. Age, BMI, smoking status, plus every other feature have no effect on this baseline.
You now have your first working benchmark. Every real model must earn its place by improving on this visible result.
Baseline metrics missing?
Confirm that DummyRegressor(strategy="mean") is passed into evaluate_model() inside main().
Check that the metric names inside the dictionary match the names used by the final DataFrame. A spelling mismatch can interrupt the printed row.
Help me debug the mean baseline.
✔️ Awesome, I've got everything!
Your saved regression_project.py now loads the CSV, validates its schema, protects a test set, plus reports the mean baseline.
ⓧ I'd like to double check the full code
Compare your saved regression_project.py with this complete version:
from pathlib import Path
import pandas as pd
from sklearn.dummy import DummyRegressor
from sklearn.metrics import mean_absolute_error, r2_score, root_mean_squared_error
from sklearn.model_selection import train_test_split
DATA_PATH = Path("data") / "medical-charges.csv"
OUTPUT_PATH = Path("outputs") / "model_diagnostics.png"
TARGET = "charges"
RANDOM_STATE = 42
EXPECTED_COLUMNS = [
"age",
"sex",
"bmi",
"children",
"smoker",
"region",
"charges",
]
def load_and_check_data(path):
if not path.exists():
raise FileNotFoundError(
f"Could not find {path}. Place medical-charges.csv inside the data folder."
)
data = pd.read_csv(path)
missing_columns = set(EXPECTED_COLUMNS) - set(data.columns)
unexpected_columns = set(data.columns) - set(EXPECTED_COLUMNS)
if missing_columns or unexpected_columns:
raise ValueError(
"Dataset columns do not match the expected schema. "
f"Missing: {sorted(missing_columns)}. "
f"Unexpected: {sorted(unexpected_columns)}."
)
print(f"Dataset shape: {data.shape}")
print("\nFirst five rows:")
print(data.head())
print("\nData types:")
print(data.dtypes)
print("\nMissing values by column:")
print(data.isna().sum())
print(f"\nDuplicate rows: {data.duplicated().sum()}")
return data
def split_data(data):
X = data.drop(columns=[TARGET])
y = data[TARGET]
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=RANDOM_STATE,
)
print(f"\nTraining rows: {len(X_train)}")
print(f"Test rows: {len(X_test)}")
return X_train, X_test, y_train, y_test
def evaluate_model(name, model, X_train, X_test, y_train, y_test):
model.fit(X_train, y_train)
train_predictions = model.predict(X_train)
test_predictions = model.predict(X_test)
metrics = {
"Model": name,
"Train MAE": mean_absolute_error(y_train, train_predictions),
"Test MAE": mean_absolute_error(y_test, test_predictions),
"Test RMSE": root_mean_squared_error(y_test, test_predictions),
"Train R2": r2_score(y_train, train_predictions),
"Test R2": r2_score(y_test, test_predictions),
}
return metrics, test_predictions
def main():
data = load_and_check_data(DATA_PATH)
X_train, X_test, y_train, y_test = split_data(data)
metrics, _ = evaluate_model(
"Mean baseline",
DummyRegressor(strategy="mean"),
X_train,
X_test,
y_train,
y_test,
)
print("\nMean baseline:")
print(pd.DataFrame([metrics]).round(2).to_string(index=False))
if __name__ == "__main__":
main()
How to use this reference
Compare each import, constant, function, plus the final entry point with your saved file. Keep the order unchanged so later additions fit into the same workflow.
Your protected split plus mean benchmark are ready. Next, you will give a linear model access to the numeric plus categorical features without leaking information from the test rows.
Build a Leakage-Safe Linear Pipeline
Your mean baseline now provides a score that every useful model must beat. Its constant predictions also expose the limitation of ignoring all six input features.
In this step, a scikit-learn Pipeline prepares mixed numeric and categorical data for linear regression. Keeping every learned transformation inside the pipeline prevents data leakage from the test rows.
In this step, get ready to:
- Apply numeric scaling and categorical encoding with a column-specific preprocessor.
- Join the preprocessor to a Linear regression model inside a pipeline.
- Compare the Linear regression model with the Mean baseline on the same held-out rows.
Define column-specific preprocessing
The numeric columns need comparable scales. The text columns need numeric representations that the model can use.
A ColumnTransformer applies the appropriate transformation to each named group. It learns every scaling value and category from the data passed to fit.
- Return to regression_project.py in VS Code.
- Select the import section at the top of the file.
- Replace the selected imports with this updated group:
from pathlib import Path
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.dummy import DummyRegressor
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error, r2_score, root_mean_squared_error
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
What do these imports add?
- The ColumnTransformer import lets one preprocessor handle separate column groups.
- The StandardScaler import provides scaling for the numeric features.
- The OneHotEncoder import converts categorical values into numeric columns.
- The Pipeline and LinearRegression imports provide the complete trainable model.
- Find the end of the split_data function.
- Add the preprocessing function below split_data:
def build_preprocessor():
numeric_features = ["age", "bmi", "children"]
categorical_features = ["sex", "smoker", "region"]
return ColumnTransformer(
transformers=[
("numeric", StandardScaler(), numeric_features),
(
"categorical",
OneHotEncoder(handle_unknown="ignore"),
categorical_features,
),
],
sparse_threshold=0,
)
What does this preprocessor do?
- The numeric_features list sends age, bmi, and children through StandardScaler.
- The categorical_features list sends sex, smoker, and region through OneHotEncoder.
- The handle_unknown="ignore" setting lets the pipeline transform a future category that was absent from its training rows.
- The sparse_threshold=0 setting makes the combined transformed output dense.
- Save regression_project.py.
- Confirm that Python can import and parse the new preprocessor by running:
python regression_project.py
What should this check show?
You should see the dataset checks and the Mean baseline metrics complete without a traceback. This proves the new imports and build_preprocessor definition are valid Python.
Does the script stop before the baseline?
Check that every new import appears before the constants. Confirm that build_preprocessor starts at the left edge of the file.
Compare the parentheses around ColumnTransformer with the code above. A missing closing parenthesis prevents the script from loading.
help me debug my preprocessing function.
Join preprocessing and training
The preprocessor becomes leakage-safe when the model fits it only through the training pipeline. The held-out test rows reach the pipeline later through prediction.
A model dictionary keeps the baseline and linear pipeline under clear names. The next comparison can evaluate both models through the same function.
- Find the end of build_preprocessor in regression_project.py.
- Add the model-building function below it:
def build_models():
return {
"Mean baseline": DummyRegressor(strategy="mean"),
"Linear regression": Pipeline(
steps=[
("preprocessor", build_preprocessor()),
("model", LinearRegression()),
]
),
}
How does the pipeline prevent leakage?
- The Mean baseline entry preserves the constant predictor from the previous step.
- The Linear regression entry places preprocessing before the model.
- Calling fit on this pipeline learns the numeric scales from X_train.
- The same training call learns the categorical vocabulary from X_train.
- Save regression_project.py.
- Confirm that the model dictionary parses by running:
python regression_project.py
What does this run prove?
You should see the Mean baseline row complete again. This confirms that Python can construct the new function definition without a syntax error.
Does the model dictionary fail to load?
Check that the dictionary closes after the Pipeline definition. Confirm that each pipeline step contains a name followed by its transformer or model.
help me fix my model dictionary.
Compare both models on the same rows
A fair comparison holds the data split constant. Both models must receive the same X_train rows and face the same X_test rows.
The main loop now evaluates every model returned by build_models. One rounded table keeps their errors directly comparable.
- Select the existing main function in regression_project.py.
- Replace the selected function with this comparison loop:
def main():
data = load_and_check_data(DATA_PATH)
X_train, X_test, y_train, y_test = split_data(data)
metric_rows = []
for name, model in build_models().items():
metrics, _ = evaluate_model(
name,
model,
X_train,
X_test,
y_train,
y_test,
)
metric_rows.append(metrics)
metrics_table = pd.DataFrame(metric_rows)
print("\nModel comparison:")
print(metrics_table.round(2).to_string(index=False))
How does the comparison work?
- The loop sends each model through evaluate_model with the same training split.
- The loop also tests each model with the same held-out split.
- The metric_rows list collects one result dictionary per model.
- The final DataFrame prints MAE in charge units. RMSE gives larger misses more influence. R2 measures improvement over constant mean predictions.
- Save regression_project.py.
Before you run the comparison, which model do you expect to have the lower test error?
- Run the completed comparison with this command:
python regression_project.py
What should the comparison show?
You should see one table under Model comparison: with rows for Mean baseline and Linear regression.
The Linear regression row should improve on the baseline's test error. Its metrics come from predictions that use all six input features.
Missing a model row?
Confirm that the loop calls build_models().items(). Check that metric_rows.append(metrics) sits inside the loop.
If categorical conversion fails, confirm that LinearRegression sits after build_preprocessor() in the same pipeline.
help me debug my model comparison table.
That is your first feature-aware model working. Its preprocessing learns only from the training rows while the test set remains a fair final check.
✔️ Awesome, I've got everything!
Great. Make sure regression_project.py is saved before moving on.
ⓧ I'd like to double check the full code
Here is the complete regression_project.py file for this step. Compare it with your saved file:
from pathlib import Path
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.dummy import DummyRegressor
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error, r2_score, root_mean_squared_error
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
DATA_PATH = Path("data") / "medical-charges.csv"
OUTPUT_PATH = Path("outputs") / "model_diagnostics.png"
TARGET = "charges"
RANDOM_STATE = 42
EXPECTED_COLUMNS = [
"age",
"sex",
"bmi",
"children",
"smoker",
"region",
"charges",
]
def load_and_check_data(path):
if not path.exists():
raise FileNotFoundError(
f"Could not find {path}. Place medical-charges.csv inside the data folder."
)
data = pd.read_csv(path)
missing_columns = set(EXPECTED_COLUMNS) - set(data.columns)
unexpected_columns = set(data.columns) - set(EXPECTED_COLUMNS)
if missing_columns or unexpected_columns:
raise ValueError(
"Dataset columns do not match the expected schema. "
f"Missing: {sorted(missing_columns)}. "
f"Unexpected: {sorted(unexpected_columns)}."
)
print(f"Dataset shape: {data.shape}")
print("\nFirst five rows:")
print(data.head())
print("\nData types:")
print(data.dtypes)
print("\nMissing values by column:")
print(data.isna().sum())
print(f"\nDuplicate rows: {data.duplicated().sum()}")
return data
def split_data(data):
X = data.drop(columns=[TARGET])
y = data[TARGET]
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=RANDOM_STATE,
)
print(f"\nTraining rows: {len(X_train)}")
print(f"Test rows: {len(X_test)}")
return X_train, X_test, y_train, y_test
def build_preprocessor():
numeric_features = ["age", "bmi", "children"]
categorical_features = ["sex", "smoker", "region"]
return ColumnTransformer(
transformers=[
("numeric", StandardScaler(), numeric_features),
(
"categorical",
OneHotEncoder(handle_unknown="ignore"),
categorical_features,
),
],
sparse_threshold=0,
)
def build_models():
return {
"Mean baseline": DummyRegressor(strategy="mean"),
"Linear regression": Pipeline(
steps=[
("preprocessor", build_preprocessor()),
("model", LinearRegression()),
]
),
}
def evaluate_model(name, model, X_train, X_test, y_train, y_test):
model.fit(X_train, y_train)
train_predictions = model.predict(X_train)
test_predictions = model.predict(X_test)
metrics = {
"Model": name,
"Train MAE": mean_absolute_error(y_train, train_predictions),
"Test MAE": mean_absolute_error(y_test, test_predictions),
"Test RMSE": root_mean_squared_error(y_test, test_predictions),
"Train R2": r2_score(y_train, train_predictions),
"Test R2": r2_score(y_test, test_predictions),
}
return metrics, test_predictions
def main():
data = load_and_check_data(DATA_PATH)
X_train, X_test, y_train, y_test = split_data(data)
metric_rows = []
for name, model in build_models().items():
metrics, _ = evaluate_model(
name,
model,
X_train,
X_test,
y_train,
y_test,
)
metric_rows.append(metrics)
metrics_table = pd.DataFrame(metric_rows)
print("\nModel comparison:")
print(metrics_table.round(2).to_string(index=False))
if __name__ == "__main__":
main()
Your baseline and linear pipeline now compete on identical unseen rows. Next, you will turn the linear model's errors into plots that reveal where its predictions fall short.
Reveal the Linear Model's Blind Spots
Your leakage-safe Linear regression now beats the Mean baseline on unseen rows. That comparison gives you a credible benchmark.
Aggregate scores compress every prediction into a few numbers. Held-out residual plots expose systematic errors that those numbers can hide.
In this step, get ready to:
- Import the tools required for held-out prediction diagnostics.
- Create actual-versus-predicted and residual panels for Linear regression.
- Inspect the saved figure for systematic errors.
Add the diagnostic plotting tools
A metric summarizes the size of your model's errors. A diagnostic figure keeps every held-out prediction visible.
Matplotlib renders the figure. scikit-learn's PredictionErrorDisplay converts the test predictions into two complementary views.
- At the top of regression_project.py, replace the entire import section with the following code:
from pathlib import Path
import matplotlib.pyplot as plt
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.dummy import DummyRegressor
from sklearn.linear_model import LinearRegression
from sklearn.metrics import (
PredictionErrorDisplay,
mean_absolute_error,
r2_score,
root_mean_squared_error,
)
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
What do these imports add?
- The matplotlib.pyplot import provides the figure layout and display tools.
- The PredictionErrorDisplay import creates regression diagnostics from actual values and predictions.
- The grouped metric imports keep the longer import list readable.
- Save regression_project.py.
- Confirm the plotting imports load in your activated terminal by running this command:
python regression_project.py
What should you see?
The script should print the existing Mean baseline and Linear regression comparison table. A clean run proves the new plotting imports load inside sklearn-env.
Plotting import failed?
Check that the terminal prompt still shows the activated sklearn-env environment. The pinned Matplotlib package was installed inside that environment.
Confirm that your import block matches the code above. Pay close attention to the parentheses around the metric imports.
help me fix the plotting import in regression_project.py.
Good, your plotting stack now loads inside the same environment as the model. You can turn the held-out predictions into visible evidence.
- In regression_project.py, add the following function after evaluate_model and before main:
def plot_diagnostics(y_test, predictions):
fig, axes = plt.subplots(1, 2, figsize=(12, 5))
PredictionErrorDisplay.from_predictions(
y_true=y_test,
y_pred=predictions["Linear regression"],
kind="actual_vs_predicted",
ax=axes[0],
scatter_kwargs={"alpha": 0.6},
)
axes[0].set_title("Linear regression: actual vs predicted")
PredictionErrorDisplay.from_predictions(
y_true=y_test,
y_pred=predictions["Linear regression"],
kind="residual_vs_predicted",
ax=axes[1],
scatter_kwargs={"alpha": 0.6},
)
axes[1].set_title("Linear regression: residuals")
fig.tight_layout()
OUTPUT_PATH.parent.mkdir(parents=True, exist_ok=True)
fig.savefig(OUTPUT_PATH, dpi=150, bbox_inches="tight")
print(f"\nSaved diagnostic figure to {OUTPUT_PATH}")
plt.show()
What does this function do?
- The fig and axes objects create two panels inside one figure.
- The first display compares each actual test charge with its Linear regression prediction.
- The second display plots each prediction against its residual error.
- The figure is saved to OUTPUT_PATH before plt.show() displays it.
- Save regression_project.py.
- Confirm Python can parse the new function by running this command:
python regression_project.py
What should you see?
The existing comparison table should print cleanly again. This confirms Python parsed the new diagnostic function without a syntax problem.
Script stopped after adding the function?
Compare the indentation inside plot_diagnostics with the reference above. Every statement in the function must remain indented by four spaces.
Check that both PredictionErrorDisplay.from_predictions calls have matching parentheses.
help me fix the plot_diagnostics function.
Save and inspect the evidence
The diagnostic function needs the Linear regression test predictions produced inside main. A predictions dictionary keeps those held-out values available until the model table has printed.
- In regression_project.py, replace the existing main function with the following version:
def main():
data = load_and_check_data(DATA_PATH)
X_train, X_test, y_train, y_test = split_data(data)
metric_rows = []
predictions = {}
for name, model in build_models().items():
metrics, test_predictions = evaluate_model(
name,
model,
X_train,
X_test,
y_train,
y_test,
)
metric_rows.append(metrics)
predictions[name] = test_predictions
metrics_table = pd.DataFrame(metric_rows)
print("\nModel comparison:")
print(metrics_table.round(2).to_string(index=False))
plot_diagnostics(y_test, predictions)
How does the data reach the plots?
- The predictions dictionary stores each model's predictions for the untouched test rows.
- The loop uses the model name as the key for each prediction array.
- The plot_diagnostics call passes the actual test charges and the stored predictions into the figure.
- Save regression_project.py.
✔️ Awesome, I've got everything!
Your updated regression_project.py now retains held-out predictions and sends them to the diagnostic function.
ⓧ I'd like to double check the full code
Compare your complete regression_project.py with this cumulative version:
from pathlib import Path
import matplotlib.pyplot as plt
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.dummy import DummyRegressor
from sklearn.linear_model import LinearRegression
from sklearn.metrics import (
PredictionErrorDisplay,
mean_absolute_error,
r2_score,
root_mean_squared_error,
)
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
DATA_PATH = Path("data") / "medical-charges.csv"
OUTPUT_PATH = Path("outputs") / "model_diagnostics.png"
TARGET = "charges"
RANDOM_STATE = 42
EXPECTED_COLUMNS = [
"age",
"sex",
"bmi",
"children",
"smoker",
"region",
"charges",
]
def load_and_check_data(path):
if not path.exists():
raise FileNotFoundError(
f"Could not find {path}. Place medical-charges.csv inside the data folder."
)
data = pd.read_csv(path)
missing_columns = set(EXPECTED_COLUMNS) - set(data.columns)
unexpected_columns = set(data.columns) - set(EXPECTED_COLUMNS)
if missing_columns or unexpected_columns:
raise ValueError(
"Dataset columns do not match the expected schema. "
f"Missing: {sorted(missing_columns)}. "
f"Unexpected: {sorted(unexpected_columns)}."
)
print(f"Dataset shape: {data.shape}")
print("\nFirst five rows:")
print(data.head())
print("\nData types:")
print(data.dtypes)
print("\nMissing values by column:")
print(data.isna().sum())
print(f"\nDuplicate rows: {data.duplicated().sum()}")
return data
def split_data(data):
X = data.drop(columns=[TARGET])
y = data[TARGET]
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=RANDOM_STATE,
)
print(f"\nTraining rows: {len(X_train)}")
print(f"Test rows: {len(X_test)}")
return X_train, X_test, y_train, y_test
def build_preprocessor():
numeric_features = ["age", "bmi", "children"]
categorical_features = ["sex", "smoker", "region"]
return ColumnTransformer(
transformers=[
("numeric", StandardScaler(), numeric_features),
(
"categorical",
OneHotEncoder(handle_unknown="ignore"),
categorical_features,
),
],
sparse_threshold=0,
)
def build_models():
return {
"Mean baseline": DummyRegressor(strategy="mean"),
"Linear regression": Pipeline(
steps=[
("preprocessor", build_preprocessor()),
("model", LinearRegression()),
]
),
}
def evaluate_model(name, model, X_train, X_test, y_train, y_test):
model.fit(X_train, y_train)
train_predictions = model.predict(X_train)
test_predictions = model.predict(X_test)
metrics = {
"Model": name,
"Train MAE": mean_absolute_error(y_train, train_predictions),
"Test MAE": mean_absolute_error(y_test, test_predictions),
"Test RMSE": root_mean_squared_error(y_test, test_predictions),
"Train R2": r2_score(y_train, train_predictions),
"Test R2": r2_score(y_test, test_predictions),
}
return metrics, test_predictions
def plot_diagnostics(y_test, predictions):
fig, axes = plt.subplots(1, 2, figsize=(12, 5))
PredictionErrorDisplay.from_predictions(
y_true=y_test,
y_pred=predictions["Linear regression"],
kind="actual_vs_predicted",
ax=axes[0],
scatter_kwargs={"alpha": 0.6},
)
axes[0].set_title("Linear regression: actual vs predicted")
PredictionErrorDisplay.from_predictions(
y_true=y_test,
y_pred=predictions["Linear regression"],
kind="residual_vs_predicted",
ax=axes[1],
scatter_kwargs={"alpha": 0.6},
)
axes[1].set_title("Linear regression: residuals")
fig.tight_layout()
OUTPUT_PATH.parent.mkdir(parents=True, exist_ok=True)
fig.savefig(OUTPUT_PATH, dpi=150, bbox_inches="tight")
print(f"\nSaved diagnostic figure to {OUTPUT_PATH}")
plt.show()
def main():
data = load_and_check_data(DATA_PATH)
X_train, X_test, y_train, y_test = split_data(data)
metric_rows = []
predictions = {}
for name, model in build_models().items():
metrics, test_predictions = evaluate_model(
name,
model,
X_train,
X_test,
y_train,
y_test,
)
metric_rows.append(metrics)
predictions[name] = test_predictions
metrics_table = pd.DataFrame(metric_rows)
print("\nModel comparison:")
print(metrics_table.round(2).to_string(index=False))
plot_diagnostics(y_test, predictions)
if __name__ == "__main__":
main()
How to use this reference
This is the complete cumulative file for this step. It preserves the two existing models while adding held-out diagnostic plots.
Before you run this check, consider whether the residuals will form a random cloud around zero.
- Run the completed diagnostic workflow in your activated terminal with this command:
python regression_project.py
What should you see?
- The terminal prints the Mean baseline and Linear regression comparison table.
- The terminal confirms that the diagnostic figure was saved to the configured output path.
- A Matplotlib window displays one actual-versus-predicted panel and one residual panel.
- Inspect the Linear regression: actual vs predicted panel for points that move away from the diagonal.
- Inspect the Linear regression: residuals panel for changing spread around the zero-residual line.
- Look for curved structure in the residual points.
- Check the higher-charge predictions for larger misses.
You should see visible departures from the diagonal at higher charges. You should also see structure or changing spread around the zero-residual line.
That structure is the intended finding. The linear model's straight-line assumptions still leave systematic error behind.
- In the VS Code Explorer sidebar, expand the outputs folder to confirm the saved artifact.
You should see model_diagnostics.png inside outputs.
Plots or image missing?
If the figure window does not open, check that plot_diagnostics(y_test, predictions) remains inside main.
If the image is missing, confirm that the script reaches fig.savefig before the figure window opens.
help me debug the missing diagnostic plots.
You now have a saved picture of where Linear regression falls short. Next, you'll test a Random forest on the same held-out rows.
Capture Nonlinear Patterns with a Random Forest
Your held-out plots exposed curved residual structure. They also revealed large misses at higher charges.
Additive straight-line effects cannot capture every pattern in this data. A Random forest averages decision trees to learn nonlinear thresholds plus feature interactions.
In this step, get ready to:
- Train a Random forest on the same training rows as the existing models.
- Select the model with the lowest held-out test MAE.
- Expand the saved diagnostics into a four-panel model comparison.
Add the nonlinear challenger
A random forest fits many decision trees. It averages their predictions to capture patterns that a single straight line misses.
The existing Pipeline keeps preprocessing attached to the model. This gives every challenger the same leakage-safe training process.
- In regression_project.py, find the existing DummyRegressor import.
- Add the Random forest import directly below it by pasting this line:
from sklearn.ensemble import RandomForestRegressor
What does this import provide?
The imported estimator trains a collection of decision trees. Its averaged prediction can represent nonlinear changes in medical charges.
- Find the existing build_models() function in regression_project.py.
- Replace the entire function with this expanded version:
def build_models():
return {
"Mean baseline": DummyRegressor(strategy="mean"),
"Linear regression": Pipeline(
steps=[
("preprocessor", build_preprocessor()),
("model", LinearRegression()),
]
),
"Random forest": Pipeline(
steps=[
("preprocessor", build_preprocessor()),
(
"model",
RandomForestRegressor(
n_estimators=200,
min_samples_leaf=3,
random_state=RANDOM_STATE,
n_jobs=-1,
),
),
]
),
}
What does this model configuration do?
- The new Pipeline applies the same column-specific preprocessor used by Linear regression.
- The forest averages 200 trees to produce each prediction.
- The min_samples_leaf=3 setting requires each leaf to represent at least three training rows.
- The shared RANDOM_STATE keeps the experiment reproducible.
- The n_jobs=-1 setting lets the estimator use the available processors.
- Save regression_project.py.
Before you run the script, how many model rows do you expect in the comparison table?
- Test the expanded model dictionary in the activated terminal by running this command:
python regression_project.py
What should I see?
The comparison table now contains rows for Mean baseline, Linear regression, plus Random forest.
The existing two-panel Linear regression figure still opens. The new table proves that the Random forest has trained on the shared split.
Missing the Random forest row?
Check that RandomForestRegressor is imported above LinearRegression.
Confirm the Random forest entry sits inside the dictionary returned by build_models().
help me fix the missing Random forest model
Compare numeric and visual evidence
The shared table makes test MAE comparable across all three models. Selecting its smallest value keeps the decision focused on unseen rows.
- In main(), find the line print(metrics_table.round(2).to_string(index=False)).
- Paste this selection block directly below that line:
best_model = min(metric_rows, key=lambda row: row["Test MAE"])
print(
f"\nLowest test MAE: {best_model['Model']} "
f"({best_model['Test MAE']:.2f})"
)
How does the selection work?
The min() call searches the collected metric rows using Test MAE as its key. The next line prints the winning model plus its held-out error.
The current diagnostic function only draws Linear regression. A two-column layout lets you compare both fitted models against identical held-out targets.
- Select the entire plot_diagnostics() function from its definition through plt.show().
- Replace the selected function with this four-panel version:
def plot_diagnostics(y_test, predictions):
model_names = ["Linear regression", "Random forest"]
fig, axes = plt.subplots(2, 2, figsize=(12, 9))
for column, model_name in enumerate(model_names):
y_pred = predictions[model_name]
PredictionErrorDisplay.from_predictions(
y_true=y_test,
y_pred=y_pred,
kind="actual_vs_predicted",
ax=axes[0, column],
scatter_kwargs={"alpha": 0.6},
)
axes[0, column].set_title(f"{model_name}: actual vs predicted")
PredictionErrorDisplay.from_predictions(
y_true=y_test,
y_pred=y_pred,
kind="residual_vs_predicted",
ax=axes[1, column],
scatter_kwargs={"alpha": 0.6},
)
axes[1, column].set_title(f"{model_name}: residuals")
fig.tight_layout()
OUTPUT_PATH.parent.mkdir(parents=True, exist_ok=True)
fig.savefig(OUTPUT_PATH, dpi=150, bbox_inches="tight")
print(f"\nSaved diagnostic figure to {OUTPUT_PATH}")
plt.show()
How does the comparison figure work?
- The model_names list assigns one figure column to each fitted model.
- The top row compares actual charges with predicted charges.
- The bottom row compares each prediction with its residual.
- Every panel uses y_test plus the corresponding held-out predictions.
- The figure overwrites outputs/model_diagnostics.png with the expanded layout.
- Save regression_project.py.
- Compare your file with the cumulative reference below.
✔️ Awesome, I've got everything!
Your Random forest pipeline plus four-panel diagnostic function are ready. Keep regression_project.py saved for the final run.
ⓧ I'd like to double check the full code
Compare your saved regression_project.py with this complete version:
from pathlib import Path
import matplotlib.pyplot as plt
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.dummy import DummyRegressor
from sklearn.ensemble import RandomForestRegressor
from sklearn.linear_model import LinearRegression
from sklearn.metrics import (
PredictionErrorDisplay,
mean_absolute_error,
r2_score,
root_mean_squared_error,
)
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
DATA_PATH = Path("data") / "medical-charges.csv"
OUTPUT_PATH = Path("outputs") / "model_diagnostics.png"
TARGET = "charges"
RANDOM_STATE = 42
EXPECTED_COLUMNS = [
"age",
"sex",
"bmi",
"children",
"smoker",
"region",
"charges",
]
def load_and_check_data(path):
if not path.exists():
raise FileNotFoundError(
f"Could not find {path}. Place medical-charges.csv inside the data folder."
)
data = pd.read_csv(path)
missing_columns = set(EXPECTED_COLUMNS) - set(data.columns)
unexpected_columns = set(data.columns) - set(EXPECTED_COLUMNS)
if missing_columns or unexpected_columns:
raise ValueError(
"Dataset columns do not match the expected schema. "
f"Missing: {sorted(missing_columns)}. "
f"Unexpected: {sorted(unexpected_columns)}."
)
print(f"Dataset shape: {data.shape}")
print("\nFirst five rows:")
print(data.head())
print("\nData types:")
print(data.dtypes)
print("\nMissing values by column:")
print(data.isna().sum())
print(f"\nDuplicate rows: {data.duplicated().sum()}")
return data
def split_data(data):
X = data.drop(columns=[TARGET])
y = data[TARGET]
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=RANDOM_STATE,
)
print(f"\nTraining rows: {len(X_train)}")
print(f"Test rows: {len(X_test)}")
return X_train, X_test, y_train, y_test
def build_preprocessor():
numeric_features = ["age", "bmi", "children"]
categorical_features = ["sex", "smoker", "region"]
return ColumnTransformer(
transformers=[
("numeric", StandardScaler(), numeric_features),
(
"categorical",
OneHotEncoder(handle_unknown="ignore"),
categorical_features,
),
],
sparse_threshold=0,
)
def build_models():
return {
"Mean baseline": DummyRegressor(strategy="mean"),
"Linear regression": Pipeline(
steps=[
("preprocessor", build_preprocessor()),
("model", LinearRegression()),
]
),
"Random forest": Pipeline(
steps=[
("preprocessor", build_preprocessor()),
(
"model",
RandomForestRegressor(
n_estimators=200,
min_samples_leaf=3,
random_state=RANDOM_STATE,
n_jobs=-1,
),
),
]
),
}
def evaluate_model(name, model, X_train, X_test, y_train, y_test):
model.fit(X_train, y_train)
train_predictions = model.predict(X_train)
test_predictions = model.predict(X_test)
metrics = {
"Model": name,
"Train MAE": mean_absolute_error(y_train, train_predictions),
"Test MAE": mean_absolute_error(y_test, test_predictions),
"Test RMSE": root_mean_squared_error(y_test, test_predictions),
"Train R2": r2_score(y_train, train_predictions),
"Test R2": r2_score(y_test, test_predictions),
}
return metrics, test_predictions
def plot_diagnostics(y_test, predictions):
model_names = ["Linear regression", "Random forest"]
fig, axes = plt.subplots(2, 2, figsize=(12, 9))
for column, model_name in enumerate(model_names):
y_pred = predictions[model_name]
PredictionErrorDisplay.from_predictions(
y_true=y_test,
y_pred=y_pred,
kind="actual_vs_predicted",
ax=axes[0, column],
scatter_kwargs={"alpha": 0.6},
)
axes[0, column].set_title(f"{model_name}: actual vs predicted")
PredictionErrorDisplay.from_predictions(
y_true=y_test,
y_pred=y_pred,
kind="residual_vs_predicted",
ax=axes[1, column],
scatter_kwargs={"alpha": 0.6},
)
axes[1, column].set_title(f"{model_name}: residuals")
fig.tight_layout()
OUTPUT_PATH.parent.mkdir(parents=True, exist_ok=True)
fig.savefig(OUTPUT_PATH, dpi=150, bbox_inches="tight")
print(f"\nSaved diagnostic figure to {OUTPUT_PATH}")
plt.show()
def main():
data = load_and_check_data(DATA_PATH)
X_train, X_test, y_train, y_test = split_data(data)
metric_rows = []
predictions = {}
for name, model in build_models().items():
metrics, test_predictions = evaluate_model(
name,
model,
X_train,
X_test,
y_train,
y_test,
)
metric_rows.append(metrics)
predictions[name] = test_predictions
metrics_table = pd.DataFrame(metric_rows)
print("\nModel comparison:")
print(metrics_table.round(2).to_string(index=False))
best_model = min(metric_rows, key=lambda row: row["Test MAE"])
print(
f"\nLowest test MAE: {best_model['Model']} "
f"({best_model['Test MAE']:.2f})"
)
plot_diagnostics(y_test, predictions)
if __name__ == "__main__":
main()
How to use this reference
This is the cumulative script for the completed model comparison. Check the import order plus the three model definitions.
Also check the winner selection plus the two-by-two diagnostic layout. Matching these sections gives you the intended final artifact.
Select the model with held-out evidence
Training performance shows how closely a model fits rows it has already seen. Model selection should use held-out performance.
Use MAE as the primary criterion. Use RMSE, R2, plus residual structure as supporting evidence.
Before you run the final comparison, which model do you expect to have the lowest test MAE?
- Run the completed experiment in the activated terminal with this command:
python regression_project.py
What should I see?
The terminal prints one comparison table containing Mean baseline, Linear regression, plus Random forest.
Below the table, the Lowest test MAE line identifies the winner. Its value is shown in charge units.
The displayed figure contains four labeled panels. The saved outputs/model_diagnostics.png file contains the same two-by-two comparison.
- Compare the Test MAE values across all three model rows.
- Check the winning model's Test RMSE value for evidence about larger misses.
- Check the winning model's Test R2 value for improvement over the mean predictor.
- Inspect the Random forest panels for tighter diagonal alignment plus less structured residuals.
- Confirm outputs/model_diagnostics.png now shows two rows with two model columns.
Missing the four-panel figure?
Check that model_names contains both model labels exactly as they appear in build_models().
Confirm the subplot call uses plt.subplots(2, 2, figsize=(12, 9)).
help me debug the four-panel diagnostics
You have turned the Linear regression shortfall into a defensible model comparison. Your final script now backs its winning model with held-out metrics plus visible residual evidence.
Secret mission
Cross-Validate the Winning Model
One held-out split can make a model look unusually strong or weak. Test the Random forest across five folds to measure how stable its MAE, RMSE, and R2 scores remain.
Clean Up Your Resources
Clean Up Your Resources
Choose whether to keep, pause, or delete your local project. Keeping it has no ongoing cost.
Resources you used:
- Local insurance-regression folder containing every project artifact.
Keep everything running
No action needed. Choose this if you plan to revisit the model or its diagnostics.
- Keep the insurance-regression folder unchanged.
- Keep data/medical-charges.csv in place so the modeling scripts can load it.
- Keep outputs/model_diagnostics.png as the saved visual record of your model comparison.
Pause - I'll come back to this later
End the current local session while preserving every project file.
- Close the diagnostic figure window if it remains open.
- Close the activated terminal to end the current sklearn-env session.
- Close VS Code after your files are saved.
- Leave the insurance-regression folder in place.
Delete - I don't want to use this again
Deleting the folder is permanent. Only the local insurance-regression folder is removed.
- Close the diagnostic figure window if it remains open.
- Remove the entire project folder from its parent directory by running these commands in the activated terminal from earlier:
cd ..
rmdir /s /q insurance-regression
What Do These Commands Remove?
The first line moves the terminal to the folder that contains insurance-regression. The second line permanently removes insurance-regression with everything inside it.
- Return to the file list in VS Code to confirm the insurance-regression project files have disappeared.
Folder Still Present?
- Close any program that is still using a file inside insurance-regression.
- Run the deletion commands above again in the terminal from earlier.
Help me diagnose why the folder will not delete.
- Close VS Code.
Nice Work!
Nice Work!
You did it! Your scikit-learn project now predicts medical insurance charges through a leakage-safe supervised regression workflow.
You've learned how to:
- Validate a real Kaggle dataset before modeling. You can protect a held-out test set. You can establish a Mean baseline for every trained model to beat.
- Train mixed numeric and categorical features inside a leakage-safe Pipeline. You can compare Linear regression with a Random forest using MAE, RMSE, and R2.
- Save prediction and residual diagnostics to outputs/model_diagnostics.png. You can select the winning model using test MAE as the primary evidence.
- Secret Mission: Use five-fold cross-validation to measure the Random forest model's performance stability across different held-out folds.
Ready to quiz yourself?