Build a QLoRA Training Lab

Train a 4-bit QLoRA adapter on Kaggle and publish it to Hugging Face Hub.

Introduction

30 Second Summary

An experiment can look successful while its starting point and proof remain scattered across a temporary session. Once that session closes, it becomes hard to explain what changed or repeat the result.

In this project, you will build a reusable QLoRA training lab in Kaggle Notebooks. The lab takes an untuned Qwen model to a published Hugging Face Hub adapter with saved evidence of what changed.

What You'll Build

Your finished lab will tell a complete tuning story, from a hardware report through repeatable model responses to a compact adapter you can inspect online.

By the end of this project, you'll have:

  • A hardware report showing every visible GPU. It records VRAM, CUDA capability, and the selected compute dtype.
  • A reusable schema normalizer that accepts three common dataset layouts. It trains a compact LoRA adapter with 4-bit NF4 quantization.
  • A published adapter repository containing the adapter files, tokenizer, configuration, metrics, and model card. You can show its evaluation loss, peak memory use, and before-and-after responses.
  • Secret Mission: Reuse the same functions with HuggingFaceTB/SmolLM3-3B without rewriting the training pipeline.

Are there any prerequisites?

You need existing Kaggle and Hugging Face accounts. Basic Python skills plus familiarity with notebook cells and LoRA concepts will help you follow the training workflow.

Before We Start

A reusable QLoRA lab starts with a clear model behavior to inspect and repeatable baseline prompts. This commitment gives you a fair way to compare untuned and fine-tuned responses later.

Set Up Kaggle and Hugging Face

Your reusable QLoRA lab needs Kaggle Notebooks to provide a GPU runtime. The notebook also needs Internet access to download models and datasets.

Hugging Face Hub supplies the model files that the lab uses. Write authentication lets the finished adapter return to the Hub as a reusable repository.

In this step, get ready to:
  • Create a Kaggle notebook with Internet access and a GPU accelerator.
  • Install the pinned training packages from requirements.txt.
  • Authenticate the notebook with a Hugging Face write token.
Configure the Kaggle notebook

The notebook runtime controls whether your lab can reach external model files. Its accelerator setting determines whether later training code can use CUDA.

Protect Your GPU Quota

Kaggle GPU use is free within its allocated quota. An idle active session still consumes that limited allocation.

Changing the accelerator restarts the notebook session. Make this change before creating project files.

  • Sign in to Kaggle with the account you already have.
  • Create a new Kaggle notebook from your account.
  • Find the Settings pane in the notebook editor.
  • Enable Internet.
  • Select GPU under Accelerator.
  • Choose GPU T4 x2 when that option is available.

You should see Internet enabled. You should also see a GPU selected under Accelerator.

Install the pinned training stack

Pinned dependencies give every notebook session the same package versions. This keeps changes in package APIs from silently changing your training pipeline.

  • Create requirements.txt in /kaggle/working using the notebook file area.
  • Paste the pinned dependency list into requirements.txt:
transformers==5.19.0
peft==0.21.2
accelerate==1.15.0
bitsandbytes==0.50.2
datasets==5.1.0
huggingface-hub==2.2.0

What Does This File Install?

  • transformers loads the tokenizer and causal language model.
  • peft attaches the trainable LoRA adapter.
  • accelerate supports the training runtime.
  • bitsandbytes provides 4-bit quantization and the 8-bit optimizer.
  • datasets loads and transforms the training records.
  • huggingface-hub handles authentication and publishing.
  • Save requirements.txt.

✔️ Awesome, I've Got Everything

Your pinned dependency file is ready. Double-check that requirements.txt is saved before installing it.

ⓧ I'd Like to Double Check the Full Code

Compare your complete requirements.txt file with this reference:

transformers==5.19.0
peft==0.21.2
accelerate==1.15.0
bitsandbytes==0.50.2
datasets==5.1.0
huggingface-hub==2.2.0
  • Add a new code cell to your Kaggle notebook.
  • Install the packages from requirements.txt by running this cell:
!pip install -r requirements.txt

What Does This Command Do?

pip reads each package pin from requirements.txt. It installs those exact versions into the active notebook environment.

The installation can take a few minutes while Kaggle downloads and replaces packages. Continuous installation output means the process is still working.

The installation output should finish without reporting that a package failed to install.

Package Installation Failed?

Confirm that requirements.txt is saved in /kaggle/working. Check that every package line matches the double-check tab.

Confirm that Internet remains enabled in the notebook settings.

Help me diagnose my Kaggle package installation.

  • Restart the notebook session using Kaggle's session controls.

Before you check the environment, which six package versions do you expect the restarted session to report?

  • Verify the installed package versions by running this cell:
!pip show transformers peft bitsandbytes datasets accelerate huggingface-hub

What Should I See?

  • transformers should show version 5.19.0.
  • peft should show version 0.21.2.
  • bitsandbytes should show version 0.50.2.
  • datasets should show version 5.1.0.
  • accelerate should show version 1.15.0.
  • huggingface-hub should show version 2.2.0.

✔️ I See All Six Pinned Versions

That is the dependency stack locked in. Your restarted session now uses the versions this lab expects.

ⓧ I See Different Versions

  • Return to the package installation cell above.
  • Run the installation cell again.
  • Restart the notebook session again.
  • Run the package verification cell again.

ⓧ One or More Packages Are Missing

  • Confirm that requirements.txt contains all six package lines.
  • Return to the package installation cell above.
  • Run the installation cell again.
  • Restart the notebook session after the installation finishes.
  • Run the package verification cell again.
Authenticate with Hugging Face

The finished adapter needs write access before it can be published. A Hugging Face write token grants that permission without placing the credential inside your project files.

Keep the Token Private

This token is a live credential with permission to create or update repository content. Anyone who obtains it can use that access from another device.

The checkpoint below asks how you stored the token. It never asks you to reveal or upload the token itself.

  • Sign in to Hugging Face with the account you already have.
  • Open your Hugging Face settings.
  • Select Access Tokens.
  • Select New token.
  • Enter kaggle-qlora-lab as the token name.
  • Select the write role.
  • Finish creating the token.
  • Copy the token value while it is available.
  • Store the token in your password manager.
  • Return to the Kaggle notebook from earlier.
  • Add a new code cell to the notebook.

Before you run the cell, what do you expect the notebook to request before it confirms authentication?

  • Start the notebook login flow by running this cell:
from huggingface_hub import notebook_login

notebook_login()

What Does This Code Do?

  • notebook_login is imported from the installed Hub package.
  • notebook_login() displays an authentication prompt inside the notebook.
  • The completed login is cached for later publishing commands.
  • Paste the write token into the notebook authentication prompt.
  • Submit the authentication prompt.

You should see confirmation that the notebook is authenticated. The token value should no longer be visible.

Authentication Not Completing?

Confirm that the token was created with the write role. A token with another role cannot provide the publishing access this project needs.

Create a replacement token if the original value was copied incorrectly. Store the replacement before closing its creation screen.

Help me troubleshoot Hugging Face notebook authentication.

You have cleared the setup work: the GPU notebook can reach the Internet and use the pinned training stack. Next, you will profile the available hardware and capture the untuned baseline.

Profile the GPU and Capture a Baseline

Your pinned training stack is ready inside the GPU-enabled Kaggle notebook. You can now check whether the runtime supports the training choices used by your reusable lab.

Before optimizing QLoRA training, you need evidence of the available hardware. You also need an untuned result that gives the later comparison a repeatable starting point.

In this step, get ready to:
  • Build a script that reports the available GPU hardware.
  • Load Qwen/Qwen3-1.7B with 4-bit NF4 quantization on GPU 0.
  • Save two untuned responses to /kaggle/working/baseline.json.
Build the hardware profiler

The available GPU determines which mixed-precision dtype your lab can use. PyTorch can inspect the runtime directly before the model consumes GPU memory.

  • Switch back to the GPU-enabled Kaggle notebook from the previous step.
  • Create baseline.py inside /kaggle/working by using the new-file control in the notebook's file browser.
  • Paste this initial hardware profiler into baseline.py:
import json
from pathlib import Path

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig

def main():
    print(f"PyTorch: {torch.__version__}")
    if not torch.cuda.is_available():
        raise RuntimeError("No CUDA GPU is available. Enable a Kaggle GPU accelerator.")

    gpu_count = torch.cuda.device_count()
    print(f"Visible GPUs: {gpu_count}")
    for index in range(gpu_count):
        properties = torch.cuda.get_device_properties(index)
        total_gib = properties.total_memory / (1024**3)
        print(
            f"GPU {index}: {torch.cuda.get_device_name(index)} | "
            f"VRAM {total_gib:.2f} GiB | "
            f"CUDA capability {properties.major}.{properties.minor}"
        )

    bf16_supported = torch.cuda.is_bf16_supported()
    compute_dtype = torch.bfloat16 if bf16_supported else torch.float16
    print(f"BF16 supported on GPU 0: {bf16_supported}")
    print(f"Selected compute dtype: {compute_dtype}")


if __name__ == "__main__":
    main()

What does this code do?

  • The CUDA availability check stops the script when the notebook has no usable GPU.
  • The device loop reports every visible GPU. Each report includes its name, VRAM, and CUDA capability.
  • The BF16 support check selects torch.bfloat16 when the runtime supports it.
  • The fallback selects torch.float16 for hardware without BF16 support.
  • Save baseline.py.

Before you run the profiler, how many GPUs do you expect the notebook to report?

  • Run the hardware profiler from a notebook cell with this command:
!python baseline.py

What should I see?

You should see the PyTorch version followed by the number of visible GPUs. Each GPU has its own hardware line.

The final lines show BF16 support and the selected compute dtype. That report gives your lab a hardware-aware precision choice.

No GPU report?

If the script says that no CUDA GPU is available, confirm that this is the GPU-enabled notebook session from the previous step. Restarting a session can change the active accelerator state.

If Python cannot find baseline.py, confirm that the file is stored directly under /kaggle/working.

Help me diagnose the missing Kaggle GPU report.

Add the quantized baseline

A useful baseline uses fixed prompts and one generation function. Repeating that contract after training makes changes in the adapter's responses easier to inspect.

  • Place the cursor on the blank line immediately above def main(): in baseline.py.
  • Add the model settings and generation function at that location:
MODEL_ID = "Qwen/Qwen3-1.7B"
OUTPUT_PATH = Path("/kaggle/working/baseline.json")
PROMPTS = [
    "Write a three-bullet checklist for reviewing a pull request.",
    "Explain gradient accumulation in one short paragraph.",
]


def generate_response(model, tokenizer, prompt):
    messages = [{"role": "user", "content": prompt}]
    text = tokenizer.apply_chat_template(
        messages,
        tokenize=False,
        add_generation_prompt=True,
        enable_thinking=False,
    )
    inputs = tokenizer(text, return_tensors="pt").to(model.device)
    generated = model.generate(
        **inputs,
        max_new_tokens=96,
        do_sample=True,
        temperature=0.7,
        top_p=0.8,
        top_k=20,
    )
    new_tokens = generated[0][inputs["input_ids"].shape[1] :]
    return tokenizer.decode(new_tokens, skip_special_tokens=True).strip()

How does generation stay repeatable?

  • PROMPTS stores the two questions used before training and after training.
  • apply_chat_template() turns each role-content message into the format expected by the model's chat template.
  • enable_thinking=False requests the model's non-thinking response format.
  • The sampling settings match the model's suggested non-thinking configuration.
  • Save baseline.py.
  • Confirm the new function parses by running the script again:
!python baseline.py

What does this check prove?

A complete hardware report confirms that Python parsed the new constants and generation function successfully. It also confirms that the GPU remains available before model loading begins.

Seeing a Python traceback?

Compare the new function against the code block above. Pay close attention to indentation inside apply_chat_template() and model.generate().

Help me find the syntax problem in my generation function.

The baseline model now needs to fit beside its generation inputs on GPU 0. Quantization reduces the base model's memory footprint before any responses are generated.

  • Find the line that prints Selected compute dtype inside main().
  • Add the quantization and model-loading code directly below that line:
    quantization_config = BitsAndBytesConfig(
        load_in_4bit=True,
        bnb_4bit_quant_type="nf4",
        bnb_4bit_use_double_quant=True,
        bnb_4bit_compute_dtype=compute_dtype,
    )
    tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
    if tokenizer.pad_token is None:
        tokenizer.pad_token = tokenizer.eos_token

    model = AutoModelForCausalLM.from_pretrained(
        MODEL_ID,
        quantization_config=quantization_config,
        dtype=compute_dtype,
        device_map=0,
    )
    footprint_gib = model.get_memory_footprint() / (1024**3)
    print(f"Quantized model footprint: {footprint_gib:.2f} GiB")

How does the model fit?

  • BitsAndBytesConfig configures the Transformers model loader to use 4-bit NF4 weights.
  • bnb_4bit_use_double_quant=True enables nested quantization.
  • bnb_4bit_compute_dtype uses the dtype selected by the hardware profiler.
  • device_map=0 places the complete model on GPU 0.
  • get_memory_footprint() turns the loaded model into a measurable memory result.

The first model download can take several minutes. The notebook may print download progress while it caches the model files.

  • Save baseline.py.
  • Load the quantized model by running the script again:
!python baseline.py

What should model loading show?

You should see the hardware report followed by a line beginning with Quantized model footprint:.

That footprint confirms that Qwen/Qwen3-1.7B loaded with the selected dtype and 4-bit configuration on GPU 0.

Model not loading?

Confirm that Internet remains enabled in the current Kaggle session. The model loader needs network access on its first run.

If the runtime reports insufficient GPU memory, stop other active work in the notebook before retrying the script.

Help me debug the quantized model load.

Run and inspect the baseline

The model is now small enough to run on the selected GPU. The final section generates both responses and stores the evidence as JSON.

  • Find the line that prints Quantized model footprint: inside main().
  • Add the response loop and file-writing code directly below that line:
    responses = []
    for prompt in PROMPTS:
        response = generate_response(model, tokenizer, prompt)
        responses.append({"prompt": prompt, "response": response})
        print(f"\nPrompt: {prompt}\nBaseline: {response}\n")

    OUTPUT_PATH.parent.mkdir(parents=True, exist_ok=True)
    OUTPUT_PATH.write_text(
        json.dumps(
            {
                "model_id": MODEL_ID,
                "compute_dtype": str(compute_dtype),
                "responses": responses,
            },
            indent=2,
        ),
        encoding="utf-8",
    )
    print(f"Saved baseline to {OUTPUT_PATH}")

What does this code preserve?

  • The loop sends both fixed prompts through generate_response().
  • Each prompt is stored beside its untuned response.
  • The output records the model ID and selected compute dtype.
  • write_text() saves the complete baseline to /kaggle/working/baseline.json.
  • Save baseline.py.

✔️ Awesome, I've got everything!

Your complete hardware profiler and baseline generator are saved in baseline.py.

ⓧ I'd like to double check the full code

Compare your baseline.py file with this complete version:

import json
from pathlib import Path

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig

MODEL_ID = "Qwen/Qwen3-1.7B"
OUTPUT_PATH = Path("/kaggle/working/baseline.json")
PROMPTS = [
    "Write a three-bullet checklist for reviewing a pull request.",
    "Explain gradient accumulation in one short paragraph.",
]


def generate_response(model, tokenizer, prompt):
    messages = [{"role": "user", "content": prompt}]
    text = tokenizer.apply_chat_template(
        messages,
        tokenize=False,
        add_generation_prompt=True,
        enable_thinking=False,
    )
    inputs = tokenizer(text, return_tensors="pt").to(model.device)
    generated = model.generate(
        **inputs,
        max_new_tokens=96,
        do_sample=True,
        temperature=0.7,
        top_p=0.8,
        top_k=20,
    )
    new_tokens = generated[0][inputs["input_ids"].shape[1] :]
    return tokenizer.decode(new_tokens, skip_special_tokens=True).strip()


def main():
    print(f"PyTorch: {torch.__version__}")
    if not torch.cuda.is_available():
        raise RuntimeError("No CUDA GPU is available. Enable a Kaggle GPU accelerator.")

    gpu_count = torch.cuda.device_count()
    print(f"Visible GPUs: {gpu_count}")
    for index in range(gpu_count):
        properties = torch.cuda.get_device_properties(index)
        total_gib = properties.total_memory / (1024**3)
        print(
            f"GPU {index}: {torch.cuda.get_device_name(index)} | "
            f"VRAM {total_gib:.2f} GiB | "
            f"CUDA capability {properties.major}.{properties.minor}"
        )

    bf16_supported = torch.cuda.is_bf16_supported()
    compute_dtype = torch.bfloat16 if bf16_supported else torch.float16
    print(f"BF16 supported on GPU 0: {bf16_supported}")
    print(f"Selected compute dtype: {compute_dtype}")

    quantization_config = BitsAndBytesConfig(
        load_in_4bit=True,
        bnb_4bit_quant_type="nf4",
        bnb_4bit_use_double_quant=True,
        bnb_4bit_compute_dtype=compute_dtype,
    )
    tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
    if tokenizer.pad_token is None:
        tokenizer.pad_token = tokenizer.eos_token

    model = AutoModelForCausalLM.from_pretrained(
        MODEL_ID,
        quantization_config=quantization_config,
        dtype=compute_dtype,
        device_map=0,
    )
    footprint_gib = model.get_memory_footprint() / (1024**3)
    print(f"Quantized model footprint: {footprint_gib:.2f} GiB")

    responses = []
    for prompt in PROMPTS:
        response = generate_response(model, tokenizer, prompt)
        responses.append({"prompt": prompt, "response": response})
        print(f"\nPrompt: {prompt}\nBaseline: {response}\n")

    OUTPUT_PATH.parent.mkdir(parents=True, exist_ok=True)
    OUTPUT_PATH.write_text(
        json.dumps(
            {
                "model_id": MODEL_ID,
                "compute_dtype": str(compute_dtype),
                "responses": responses,
            },
            indent=2,
        ),
        encoding="utf-8",
    )
    print(f"Saved baseline to {OUTPUT_PATH}")


if __name__ == "__main__":
    main()

Before you run the completed script, what hardware details and baseline evidence do you expect it to produce?

  • Run the completed baseline script from a notebook cell with this command:
!python baseline.py

What should the completed run show?

You should see every visible GPU with its name, VRAM, and CUDA capability. You should also see BF16 support and the selected compute dtype.

The script then prints the quantized model footprint and two baseline responses. Its final line confirms that the evidence was saved to /kaggle/working/baseline.json.

  • Select baseline.json under /kaggle/working in the notebook's file browser.
  • Confirm that the file contains model_id, compute_dtype, and two records under responses.

You now have a hardware report and two saved untuned responses. This baseline gives the later adapter comparison evidence it can reuse.

Missing a response or JSON file?

Confirm that the response loop remains indented inside main(). The file-writing section must use the same indentation.

If generation stops during model loading, rerun the completed script after confirming that the GPU session is still active.

Help me debug my missing baseline output.

Your reusable lab now knows its hardware and has a baseline to beat. Next, you'll expose the dataset-schema assumption that makes a rigid training pipeline fail.

Expose the Dataset Schema Problem

You have a solid checkpoint now. Your hardware profile and untuned baseline are saved in /kaggle/working/baseline.json for every later comparison.

Training data can describe the same conversation with different key layouts. In this step, you will test a rigid messages validator against three fixtures to expose the reuse problem.

In this step, get ready to:
  • Create representative records for three dataset schemas.
  • Pass each fixture through a rigid messages-only validator.
  • Define the canonical message contract for the reusable pipeline.
Create three schema fixtures

A dataset schema defines the fields used to store each record. These fixtures express similar training content through three common field layouts.

  • In your Kaggle notebook from earlier, create /kaggle/working/schema_probe.py with the notebook file editor.
  • Paste the three fixtures at the top of schema_probe.py using this code:
FIXTURES = {
    "conversational": {
        "messages": [
            {"role": "user", "content": "Name one benefit of unit tests."},
            {"role": "assistant", "content": "They catch regressions early."},
        ]
    },
    "prompt_completion": {
        "prompt": "Name one benefit of unit tests.",
        "completion": "They catch regressions early.",
    },
    "instruction_input_output": {
        "instruction": "Summarize the input.",
        "input": "Unit tests help catch regressions before release.",
        "output": "Unit tests catch pre-release regressions.",
    },
}

What do the fixtures represent?

  • The conversational fixture stores a list of messages with role fields plus content fields.
  • The prompt_completion fixture separates the request from its answer.
  • The instruction_input_output fixture separates the task from its source text plus its expected output.
  • Save schema_probe.py.
  • Confirm that schema_probe.py appears under /kaggle/working in the notebook file list.

File missing from the notebook?

  • Check that the file path is /kaggle/working/schema_probe.py.
  • Save the file again from the notebook editor.

Help me find or save schema_probe.py in my Kaggle notebook.

Add the rigid validator

The validator models a tempting shortcut. It treats the presence of one exact key as the complete dataset contract.

  • Add require_messages() below FIXTURES by pasting this code:
def require_messages(example):
    if "messages" not in example:
        raise ValueError(
            "Expected a messages column; received keys "
            f"{sorted(example.keys())}"
        )
    return example["messages"]

What does this validator do?

  • The if condition checks for an exact messages key.
  • The ValueError reports the keys that arrived when the required key is missing.
  • The final line returns the message list when the check succeeds.
  • Save schema_probe.py.
  • Confirm that require_messages() appears below the closing brace of FIXTURES.

Validator code not lining up?

  • Check that def require_messages(example): starts at the left edge of the file.
  • Keep the indented lines inside the function aligned with four spaces.

Help me check the indentation in require_messages() without changing its behavior.

The validator needs a repeatable way to inspect every fixture. A loop gives each record the same test while keeping the accepted path separate from the rejected path.

  • Add the fixture loop below require_messages() using this code:
for name, fixture in FIXTURES.items():
    try:
        messages = require_messages(fixture)
        print(f"{name}: accepted with {len(messages)} messages")
    except ValueError as error:
        print(f"{name}: rejected: {error}")

How does the probe work?

  • The loop sends each fixture through require_messages() once.
  • The try path prints the accepted fixture name plus its message count.
  • The except path catches the validation failure plus prints its explanation.
  • Save schema_probe.py.

✔️ Awesome, I've got everything!

Great. Double-check that you saved schema_probe.py before running the probe.

ⓧ I'd like to double check the full code

Compare your complete schema_probe.py file with this version:

FIXTURES = {
    "conversational": {
        "messages": [
            {"role": "user", "content": "Name one benefit of unit tests."},
            {"role": "assistant", "content": "They catch regressions early."},
        ]
    },
    "prompt_completion": {
        "prompt": "Name one benefit of unit tests.",
        "completion": "They catch regressions early.",
    },
    "instruction_input_output": {
        "instruction": "Summarize the input.",
        "input": "Unit tests help catch regressions before release.",
        "output": "Unit tests catch pre-release regressions.",
    },
}


def require_messages(example):
    if "messages" not in example:
        raise ValueError(
            "Expected a messages column; received keys "
            f"{sorted(example.keys())}"
        )
    return example["messages"]


for name, fixture in FIXTURES.items():
    try:
        messages = require_messages(fixture)
        print(f"{name}: accepted with {len(messages)} messages")
    except ValueError as error:
        print(f"{name}: rejected: {error}")
Run the probe and define the contract

Before you run the probe, predict how many fixtures the rigid validator will accept.

  • Run the completed probe in a new notebook cell using this command:
!python schema_probe.py

What should I see?

  • The conversational fixture prints accepted with two messages.
  • The prompt_completion fixture prints rejected. Its incompatible keys are completion plus prompt.
  • The instruction_input_output fixture prints rejected. Its incompatible keys are input, instruction, and output.

You have caught the rigid assumption in action. Two usable records are discarded because their keys differ.

Probe output looks different?

  • Confirm that the command runs from the Kaggle notebook containing /kaggle/working/schema_probe.py.
  • Compare each fixture name against the full-code tab above.
  • Check that require_messages() tests for the singular messages key.

Help me debug the output from schema_probe.py without changing the intended rigid behavior.

The failure shows where normalization belongs in a reusable pipeline. Every source record needs one canonical shape before a chat template formats it for the model.

  • Record that every accepted source record must become a non-empty list.
  • Record that every item in the list must contain a role field.
  • Record that every item in the list must contain a content field.
  • Record that normalization must finish before chat templating.

The schema failure is now repeatable and visible. Next, you will replace the rigid assumption with a normalizer that prepares all three record styles for QLoRA training.

Normalize Data and Train the Adapter

The schema probe exposed the weak point in a rigid training pipeline. Your conversational fixture worked while two common record shapes were rejected.

You will now add schema normalization before chat templating. You will use QLoRA to train a compact adapter while the quantized base model stays frozen.

In this step, get ready to:
  • Create a reusable module that normalizes three dataset schemas.
  • Build a hardware-aware training lab with 4-bit NF4 quantization.
  • Train the LoRA adapter and preserve its metrics.
Create the reusable training module

A normalizer gives every source record the same message contract. The rest of the pipeline can then apply one chat template without knowing the original dataset schema.

  • Select the new-file control in your Kaggle notebook's file panel.
  • Create qlora_lab.py inside /kaggle/working.
  • Copy the complete module from the second tab below into qlora_lab.py.

✔️ I created the training module

Great. Keep qlora_lab.py open while you confirm that it is saved under /kaggle/working.

ⓧ I'd like to double check the full code

Compare your qlora_lab.py file with this complete module:

import os

os.environ["CUDA_VISIBLE_DEVICES"] = "0"

import json
import random
from pathlib import Path

import torch
from datasets import Dataset, load_dataset
from packaging.version import Version
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training
from transformers import (
    AutoModelForCausalLM,
    AutoTokenizer,
    BitsAndBytesConfig,
    DataCollatorForLanguageModeling,
    Trainer,
    TrainingArguments,
)

DEFAULT_MODEL_ID = "Qwen/Qwen3-1.7B"
DEFAULT_DATASET_ID = "trl-lib/Capybara"
DEFAULT_OUTPUT_DIR = "/kaggle/working/qwen3-1.7b-capybara-qlora"
BASELINE_PATH = Path("/kaggle/working/baseline.json")
PROMPTS = [
    "Write a three-bullet checklist for reviewing a pull request.",
    "Explain gradient accumulation in one short paragraph.",
]


def validate_messages(messages):
    if not isinstance(messages, list) or not messages:
        raise ValueError("messages must be a non-empty list")
    normalized = []
    for message in messages:
        if not isinstance(message, dict):
            raise ValueError("each message must be a dictionary")
        role = message.get("role")
        content = message.get("content")
        if role not in {"system", "user", "assistant"}:
            raise ValueError(f"unsupported role: {role}")
        if not isinstance(content, str) or not content.strip():
            raise ValueError("message content must be a non-empty string")
        normalized.append({"role": role, "content": content.strip()})
    return normalized


def normalize_example(example):
    if "messages" in example:
        return {"messages": validate_messages(example["messages"])}

    if "prompt" in example and "completion" in example:
        prompt = example["prompt"]
        completion = example["completion"]
        if isinstance(prompt, list) and isinstance(completion, list):
            return {"messages": validate_messages(prompt + completion)}
        return {
            "messages": validate_messages(
                [
                    {"role": "user", "content": str(prompt)},
                    {"role": "assistant", "content": str(completion)},
                ]
            )
        }

    if "instruction" in example and "output" in example:
        instruction = str(example["instruction"]).strip()
        extra_input = str(example.get("input", "")).strip()
        user_content = instruction
        if extra_input:
            user_content = f"{instruction}\n\nInput:\n{extra_input}"
        return {
            "messages": validate_messages(
                [
                    {"role": "user", "content": user_content},
                    {"role": "assistant", "content": str(example["output"])},
                ]
            )
        }

    raise ValueError(f"Unsupported dataset keys: {sorted(example.keys())}")


def generate_responses(model, tokenizer):
    model.eval()
    responses = []
    for prompt in PROMPTS:
        messages = [{"role": "user", "content": prompt}]
        text = tokenizer.apply_chat_template(
            messages,
            tokenize=False,
            add_generation_prompt=True,
            enable_thinking=False,
        )
        inputs = tokenizer(text, return_tensors="pt").to(model.device)
        generated = model.generate(
            **inputs,
            max_new_tokens=96,
            do_sample=True,
            temperature=0.7,
            top_p=0.8,
            top_k=20,
        )
        new_tokens = generated[0][inputs["input_ids"].shape[1] :]
        response = tokenizer.decode(new_tokens, skip_special_tokens=True).strip()
        responses.append({"prompt": prompt, "response": response})
    return responses


def build_lab(
    model_id=DEFAULT_MODEL_ID,
    dataset_id=DEFAULT_DATASET_ID,
    output_dir=DEFAULT_OUTPUT_DIR,
    repo_id="your-hugging-face-username/qwen3-1.7b-capybara-qlora",
    train_samples=256,
    eval_samples=32,
    max_length=512,
    max_steps=30,
):
    if Version(torch.__version__.split("+")[0]) <= Version("2.5"):
        raise RuntimeError("Transformers 5.19.0 requires PyTorch newer than 2.5.")
    if not torch.cuda.is_available():
        raise RuntimeError("No CUDA GPU is available. Enable a Kaggle GPU accelerator.")

    output_path = Path(output_dir)
    output_path.mkdir(parents=True, exist_ok=True)

    bf16_supported = torch.cuda.is_bf16_supported()
    compute_dtype = torch.bfloat16 if bf16_supported else torch.float16
    print(f"Training GPU: {torch.cuda.get_device_name(0)}")
    print(f"Selected compute dtype: {compute_dtype}")

    tokenizer = AutoTokenizer.from_pretrained(model_id)
    if tokenizer.pad_token is None:
        tokenizer.pad_token = tokenizer.eos_token
    tokenizer.padding_side = "right"

    raw_dataset = load_dataset(dataset_id, split="train")
    required = train_samples + eval_samples
    if len(raw_dataset) < required:
        raise ValueError(
            f"Dataset has {len(raw_dataset)} rows, but {required} are required."
        )

    rows = [normalize_example(raw_dataset[index]) for index in range(required)]
    random.Random(42).shuffle(rows)
    train_dataset = Dataset.from_list(rows[:train_samples])
    eval_dataset = Dataset.from_list(rows[train_samples:required])

    def tokenize_example(example):
        text = tokenizer.apply_chat_template(
            example["messages"],
            tokenize=False,
            add_generation_prompt=False,
        )
        return tokenizer(
            text,
            truncation=True,
            max_length=max_length,
            return_special_tokens_mask=True,
        )

    train_dataset = train_dataset.map(
        tokenize_example,
        remove_columns=["messages"],
    )
    eval_dataset = eval_dataset.map(
        tokenize_example,
        remove_columns=["messages"],
    )

    quantization_config = BitsAndBytesConfig(
        load_in_4bit=True,
        bnb_4bit_quant_type="nf4",
        bnb_4bit_use_double_quant=True,
        bnb_4bit_compute_dtype=compute_dtype,
    )
    model = AutoModelForCausalLM.from_pretrained(
        model_id,
        quantization_config=quantization_config,
        dtype=compute_dtype,
        device_map=0,
    )
    model.config.use_cache = False
    model = prepare_model_for_kbit_training(
        model,
        use_gradient_checkpointing=True,
        gradient_checkpointing_kwargs={"use_reentrant": False},
    )

    lora_config = LoraConfig(
        r=16,
        lora_alpha=32,
        lora_dropout=0.05,
        bias="none",
        task_type="CAUSAL_LM",
        target_modules="all-linear",
    )
    model = get_peft_model(model, lora_config)
    model.print_trainable_parameters()
    footprint_gib = model.get_memory_footprint() / (1024**3)
    print(f"Model footprint before training: {footprint_gib:.2f} GiB")

    training_args = TrainingArguments(
        output_dir=output_dir,
        max_steps=max_steps,
        per_device_train_batch_size=1,
        per_device_eval_batch_size=1,
        gradient_accumulation_steps=8,
        learning_rate=1e-4,
        warmup_steps=max(1, max_steps // 10),
        optim="adamw_8bit",
        logging_strategy="steps",
        logging_steps=5,
        logging_first_step=True,
        eval_strategy="steps",
        eval_steps=max(1, max_steps // 2),
        save_strategy="no",
        report_to="none",
        prediction_loss_only=True,
        gradient_checkpointing=True,
        gradient_checkpointing_kwargs={"use_reentrant": False},
        use_cache=False,
        fp16=not bf16_supported,
        bf16=bf16_supported,
        seed=42,
        data_seed=42,
        skip_memory_metrics=False,
        hub_model_id=repo_id,
    )
    collator = DataCollatorForLanguageModeling(
        tokenizer=tokenizer,
        mlm=False,
        pad_to_multiple_of=8,
    )
    trainer = Trainer(
        model=model,
        args=training_args,
        train_dataset=train_dataset,
        eval_dataset=eval_dataset,
        data_collator=collator,
        processing_class=tokenizer,
    )

    return {
        "model_id": model_id,
        "dataset_id": dataset_id,
        "output_dir": output_path,
        "repo_id": repo_id,
        "compute_dtype": str(compute_dtype),
        "tokenizer": tokenizer,
        "trainer": trainer,
    }


def train_lab(lab):
    trainer = lab["trainer"]
    torch.cuda.reset_peak_memory_stats()
    train_result = trainer.train()
    trainer.save_metrics("train", train_result.metrics)

    eval_metrics = trainer.evaluate()
    trainer.save_metrics("eval", eval_metrics)
    trainer.save_model(str(lab["output_dir"]))

    peak_gib = torch.cuda.max_memory_allocated() / (1024**3)
    lab["eval_metrics"] = eval_metrics
    lab["peak_memory_gib"] = peak_gib
    print(f"Peak allocated GPU memory: {peak_gib:.2f} GiB")
    print(f"Saved adapter to {lab['output_dir']}")
    return train_result.metrics


def evaluate_and_publish(lab):
    trainer = lab["trainer"]
    model = trainer.model
    tokenizer = lab["tokenizer"]

    model.gradient_checkpointing_disable()
    model.config.use_cache = True
    fine_tuned_responses = generate_responses(model, tokenizer)

    baseline = None
    if BASELINE_PATH.exists():
        baseline = json.loads(BASELINE_PATH.read_text(encoding="utf-8"))

    comparison = {
        "model_id": lab["model_id"],
        "dataset_id": lab["dataset_id"],
        "compute_dtype": lab["compute_dtype"],
        "peak_memory_gib": lab.get("peak_memory_gib"),
        "eval_metrics": lab.get("eval_metrics", {}),
        "baseline": baseline,
        "fine_tuned": fine_tuned_responses,
    }
    comparison_path = lab["output_dir"] / "comparison.json"
    comparison_path.write_text(
        json.dumps(comparison, indent=2),
        encoding="utf-8",
    )
    print(json.dumps(comparison, indent=2))
    print(f"Saved comparison to {comparison_path}")

    if lab["repo_id"].startswith("your-hugging-face-username/"):
        raise ValueError("Replace the repository placeholder before publishing.")

    trainer.push_to_hub(
        commit_message="Upload QLoRA adapter, tokenizer, metrics, and model card",
        dataset_tags=lab["dataset_id"],
        finetuned_from=lab["model_id"],
        tasks="text-generation",
    )
    print(f"Published to https://huggingface.co/{lab['repo_id']}")
    return comparison

How is the module organized?

  • The validate_messages() function enforces non-empty role-content messages.
  • The normalize_example() function converts each supported source schema into that contract.
  • The build_lab() function samples 288 records before shuffling them with seed 42.
  • The same function assigns 256 records to training. It assigns 32 records to evaluation.
  • The train_lab() function trains the adapter. It also preserves metrics and peak GPU memory.
  • Save qlora_lab.py.
  • Confirm the notebook's file panel lists qlora_lab.py under /kaggle/working.

Cannot find the module file?

Check that the filename is exactly qlora_lab.py. Confirm that it sits directly inside /kaggle/working.

Compare your file with the complete version above if the editor reports a syntax problem.

Help me check my training module.

Build the hardware-aware lab

The build_lab() function turns the reusable module into a configured training run. It uses PyTorch to choose BF16 only when the GPU supports it.

The function loads Qwen/Qwen3-1.7B with 4-bit NF4 nested quantization. PEFT keeps only the all-linear LoRA adapter trainable.

  • Record your Hugging Face username here: your Hugging Face username.
  • Add a code cell below your earlier notebook cells.
  • Paste the lab setup into the new cell using this code:
from qlora_lab import build_lab, train_lab

lab = build_lab(
    repo_id="your-hugging-face-username/qwen3-1.7b-capybara-qlora",
)

What does this code prepare?

  • The import loads the two functions used in this training step.
  • The repo_id value records the Hugging Face Hub destination for the adapter.
  • The defaults select trl-lib/Capybara as the dataset.
  • The defaults cap the run at max_steps=30.
  • Replace your-hugging-face-username with your Hugging Face username.

This first build can take several minutes while the dataset and model are downloaded. The cell is still working while Kaggle shows active output.

Before you run the cell, which compute dtype do you expect after seeing your earlier hardware report?

  • Run the completed lab setup cell.

You will see the training GPU and selected compute dtype. You will also see the trainable-parameter count followed by the model footprint.

Why does this fit on one GPU?

NF4 stores the frozen base model in 4-bit form. Nested quantization reduces that storage footprint further.

The LoRA adapter remains trainable in full precision. Gradient checkpointing trades extra computation for lower activation memory.

Lab setup stopped early?

  • Confirm your Kaggle session still has a GPU accelerator if the message mentions unavailable CUDA hardware.
  • Restart the notebook session if an import still points to an older installed package.
  • Compare qlora_lab.py with the full-code tab if Python reports a name or indentation problem.

Help me debug the lab setup.

Train and verify the adapter

The Trainer now has normalized training rows and evaluation rows. Gradient accumulation combines eight one-record batches before each optimizer update.

Expect the training cell to run for several minutes. Kaggle will print progress logs while the run advances toward its configured step limit.

Before you run this cell, do you expect the frozen base model or the adapter parameters to receive training updates?

  • Train the configured adapter by running this cell:
train_lab(lab)

What happens during training?

  • The trainer updates only the adapter parameters exposed by PEFT.
  • Evaluation runs during the short training schedule.
  • Training metrics are saved under /kaggle/working/qwen3-1.7b-capybara-qlora.
  • The adapter and tokenizer are saved to the same directory.
  • Peak allocated GPU memory is stored in lab["peak_memory_gib"].

You will see the progress reach step 30 before evaluation metrics are printed. The final lines report Peak allocated GPU memory and the saved adapter path.

That is the main training milestone complete. Your quantized base model has produced a reusable adapter with recorded evaluation evidence.

  • Confirm the notebook's file panel contains /kaggle/working/qwen3-1.7b-capybara-qlora.
  • Confirm the final output reports evaluation metrics.
  • Confirm the final output reports peak allocated GPU memory.

Training did not reach step 30?

  • Confirm the GPU session is still active if training stops during model execution.
  • Confirm that the earlier setup cell finished creating lab before rerunning the training cell.
  • Restart the session only if the GPU runtime becomes unavailable. Rerun the authentication and setup cells before training again.

Help me diagnose the interrupted training run.

Your adapter is trained and stored with its metrics. Next, you will compare its responses with the saved baseline before publishing the evidence.

Compare, Save, and Publish the Adapter

Your trained QLoRA adapter is saved in Kaggle. The checkpoint proves that the 30-step training run finished.

A finished run becomes portfolio evidence when its metrics sit beside repeatable outputs. evaluate_and_publish() creates that comparison from the same prompts used for the baseline.

The function also publishes the adapter to Hugging Face Hub. The repository gives other people one place to inspect the adapter plus its model card.

In this step, get ready to:
  • Generate fine-tuned responses for the two baseline prompts.
  • Inspect the saved comparison evidence.
  • Publish the trained adapter to Hugging Face Hub.
Package the training evidence

The trained model still carries the memory-saving behavior used during optimization. The finishing function switches the model back to generation mode before collecting evidence.

  • Add a new code cell below the training output in your open Kaggle notebook.
  • Paste the finishing call into the new cell:
from qlora_lab import evaluate_and_publish

evaluate_and_publish(lab)

What does the finishing function do?

  • The function disables gradient checkpointing so the trained model is ready for response generation.
  • The model.config.use_cache = True line restores the model cache used during generation.
  • The generate_responses() function runs the same two prompts used in the untuned baseline.
  • The function stores the baseline plus the fine-tuned responses in a JSON comparison file.
  • The Transformers Trainer publishes the saved adapter to the repository stored in lab["repo_id"].

Before you run the cell, think about which saved evidence you expect to appear in the printed comparison.

  • Run the finishing cell.
  • Keep the cell running until it returns the published repository URL.

You should see the complete comparison object printed beneath the cell. It includes the saved baseline plus two fine-tuned responses.

  • Confirm the printed comparison contains baseline, fine_tuned, and eval_metrics.
  • Confirm the final output names /kaggle/working/qwen3-1.7b-capybara-qlora/comparison.json as the saved comparison path.
  • Confirm the final output prints your Hugging Face Hub repository URL.

That is the training run packaged as evidence. The comparison is saved locally while the adapter is available from its Hub repository.

Adapter not publishing?

  • If you see Replace the repository placeholder before publishing., return to the cell where you created lab.
  • Rebuild lab with a repository ID containing your actual Hugging Face username.
  • If the upload reports an authentication problem, return to the notebook login cell from Step 1.
  • Authenticate again with your write token through notebook_login().

Your local comparison.json remains available when an upload fails because the function writes it before publishing.

Help me fix the adapter publication from my Kaggle notebook.

Compare the saved responses

The baseline and fine_tuned entries use the same two prompts. Their shared prompt set makes changes in structure, relevance, or concision visible.

A 30-step run may improve one response while weakening another. Judge the observable differences without assuming that every fine-tuned response must be better.

  • Open /kaggle/working/qwen3-1.7b-capybara-qlora/comparison.json from the file browser in your Kaggle notebook.
  • Match the two objects under baseline.responses with the two objects under fine_tuned using their prompt values.
  • Compare the response structure for each prompt.
  • Compare the relevance of each response to its prompt.
  • Compare the concision of each response.
  • Review eval_metrics for the saved evaluation results.
  • Confirm comparison.json sits alongside the existing adapter, tokenizer, and metric files.

You should see model_id, dataset_id, compute_dtype, and peak_memory_gib at the top level. The same object should contain eval_metrics, baseline, and fine_tuned.

Both response collections should cover the pull-request checklist prompt plus the gradient-accumulation prompt.

Missing the saved baseline?

A null value under baseline means the finishing function could not read /kaggle/working/baseline.json.

  • Confirm /kaggle/working/baseline.json still exists in the notebook file browser.
  • Rerun the finishing cell after restoring the baseline file.

Help me restore the baseline comparison.

Verify the published repository

The Hub repository is the inspectable version of your training result. Its metadata connects the adapter to its source model, training dataset, and text-generation task.

Before you open the repository, think about which local artifacts you expect the upload to preserve.

  • Copy https://huggingface.co/your-hugging-face-username/qwen3-1.7b-capybara-qlora into your browser address bar.
  • Replace your-hugging-face-username with your Hugging Face username.
  • Press Enter to load the repository.
  • Confirm the repository file list contains the adapter, tokenizer, configuration, metrics, and model card.
  • Confirm the model card metadata identifies trl-lib/Capybara as the dataset.
  • Confirm the model card metadata identifies Qwen/Qwen3-1.7B as the source model.
  • Confirm the repository metadata identifies text-generation as the task.

You should see the published adapter plus the files needed to load it. The repository should also display the generated model card.

That closes the loop: your QLoRA run now has local comparison evidence plus a published adapter that other people can inspect.

Secret mission

Swap in SmolLM3 Without Rewriting the Pipeline

Your Qwen adapter proved that the training lab works for one model. Reuse the same data contract and QLoRA functions with HuggingFaceTB/SmolLM3-3B to prove that the lab transfers to a larger model.

Clean Up Your Resources

Clean Up Your Resources

Choose whether to keep your work available, pause the GPU session, or delete the artifacts entirely. This project stays at no ongoing cost through the quota-limited GPU allocation in Kaggle Notebooks plus hosting on Hugging Face Hub.

Resources you used:

  • A Kaggle notebook session with generated files under /kaggle/working.
  • Hugging Face Hub repository your-hugging-face-username/qwen3-1.7b-capybara-qlora containing the published Qwen QLoRA adapter, tokenizer, configuration, metrics, model card, and comparison evidence.
  • Hugging Face Hub repository your-hugging-face-username/smollm3-3b-capybara-qlora containing the published SmolLM3 adapter, tokenizer, configuration, metrics, model card, and comparison evidence.

Keep everything running

No action needed. Choose this if you are still comparing the two adapters or continuing your experiments.

  • Save a Kaggle notebook version to preserve the generated files plus cell outputs beyond the current session.
  • Keep both Hugging Face Hub repositories available for future inspection or reuse.
  • Use the active Kaggle GPU only while you need the training runtime.

Pause - I'll come back to this later

Shut down the active GPU session to protect your Kaggle allocation. Your published adapters stay available for later.

  • Save a Kaggle notebook version to preserve the outputs you want to revisit.
  • Stop the active Kaggle session with the notebook session controls.
  • Leave both Hugging Face Hub repositories in place for your next session.

Delete - I don't want to use this again

The delete path gives you a clean slate. Anything you download first stays safe after the published adapters and Kaggle files are removed.

  • Download any adapter files or comparison evidence you want to retain.
  • Delete your-hugging-face-username/qwen3-1.7b-capybara-qlora from its repository settings.
  • Delete your-hugging-face-username/smollm3-3b-capybara-qlora from its repository settings.
  • Delete any saved Kaggle notebook versions you no longer need.
  • Delete any saved Kaggle notebook outputs you no longer need.
  • Remove all generated files from /kaggle/working by running this command in a notebook cell:
!rm -rf /kaggle/working/*

What Does This Command Remove?

The command permanently clears the generated contents under /kaggle/working. This includes both adapter directories plus the baseline and comparison artifacts stored there.

  • Stop the active Kaggle session with the notebook session controls.

Nice Work!

Nice Work!

You pulled it off! Your reusable QLoRA training lab now produces evidence-backed adapters on Kaggle for publication to Hugging Face Hub.

What you learned:

  • Built a hardware profiler that inspects the available GPU. Selected the supported mixed-precision dtype before capturing repeatable untuned responses.
  • Exposed a rigid dataset schema through a visible failure. Added a reusable message normalizer that accepts three record formats.
  • Trained a 4-bit NF4 QLoRA adapter for Qwen/Qwen3-1.7B. Recorded evaluation metrics plus repeatable response comparisons. Published the adapter repository to Hugging Face Hub.
  • Secret Mission: Reused the unchanged pipeline with HuggingFaceTB/SmolLM3-3B. Published a second adapter. Compared it with Qwen under the same experiment contract.

Ready to quiz yourself?