Build a CPU Transformer Inference Engine

Build a local CPU text generator with manual decoding and KV-cache benchmarks.

Introduction

30 Second Summary

You type a prompt. A continuation appears without showing how each new piece was chosen.

In this project, you will build a local CPU inference engine with Python plus PyTorch. You will compare full-context recomputation with KV caching through a repeatable benchmark.

What You'll Build

From PowerShell, you will enter a prompt to produce matching DistilGPT2 continuations from two decode paths before comparing their measured speeds.

By the end of this project, you'll have:

  • A manual autoregressive decode loop that turns a terminal prompt into a DistilGPT2 continuation one token at a time.
  • A custom sampler you can toggle from deterministic greedy decoding to temperature-controlled top-k sampling.
  • A benchmark report that shows Token match: True when both paths agree. It also reports elapsed time plus tokens per second for each path.
  • Secret Mission: Stream each decoded token from the cached path while generation is still running.

Are there any prerequisites?

You need Visual Studio Code plus Python 3.10 or newer on Windows. Internet access is required for the first DistilGPT2 download.

Before We Start

Before the hands-on work begins, lock in what you are building and why comparing its two decoding paths matters.

Set Up the Windows Inference Lab

Your manual Transformer decode loop depends on library APIs that stay consistent from one run to the next. Unpinned packages can change those APIs before you inspect the logits or cache state.

An isolated virtual environment inside Visual Studio Code gives this native-Windows lab its own packages. PowerShell then runs the same Python installation throughout the project.

In this step, get ready to:
  • Prepare a native-Windows workspace that runs Python 3.10 or newer.
  • Isolate the project's dependencies inside .venv.
  • Install pinned PyTorch and Hugging Face Transformers packages from requirements.txt.
Prepare the Windows workspace

An opened folder gives the editor and terminal one shared project location. You will keep the entire lab on your Desktop so its files remain easy to find.

Why native Windows?

Python and Visual Studio Code are already available on your Windows computer. Staying in that environment keeps your attention on the inference loop.

WSL2 adds a separate shell boundary. Docker Desktop adds container setup that this local CPU project does not need.

  • Press the Windows key to open Windows search.
  • Type Visual Studio Code into the search field.
  • Press Enter to open Visual Studio Code.

Visual Studio Code opens with an empty editor window. You are ready to give the inference engine its own folder.

  • Select File from the top menu.
  • Select Open Folder....
  • Choose your Desktop as the folder location.
  • Select New Folder.
  • Enter cpu-inference-engine as the folder name.
  • Select Select Folder.

You should see cpu-inference-engine in the Explorer sidebar. Your lab now has a visible home on the Desktop.

  • Select Manage in the Restricted Mode banner if it appears.
  • Review the displayed folder path to confirm it points to cpu-inference-engine.
  • Select Trust to enable the editor features for your own new folder.
  • Select View from the top menu.
  • Select Terminal.
  • Open the terminal dropdown if the terminal tab does not show PowerShell.
  • Select PowerShell from the available shells.

The PowerShell prompt should include cpu-inference-engine in its path. This confirms that commands run inside the project folder.

  • Check the Python version available to PowerShell by running this command:
python --version

What does this command do?

The --version flag prints the installed Python version. The pinned libraries require Python 3.10 or newer.

✔️ I see version 3.10 or higher

That check is complete. Your Python installation supports the pinned packages used in this lab.

ⓧ I see an older version

The installed Python version is below the minimum required by the pinned packages. Update Python before creating the virtual environment.

  • Visit the official Python downloads page for Windows.
  • Install a stable Windows release of Python 3.10 or newer.
  • Restart Visual Studio Code from Windows search after the installation finishes.
  • Return to the cpu-inference-engine folder through File and Open Folder....
  • Open a PowerShell terminal through View and Terminal.
  • Recheck the installed version by running this command:
python --version

What should this confirm?

PowerShell should now report Python 3.10 or newer. That result confirms the updated runtime is available inside Visual Studio Code.

ⓧ Command not found

PowerShell cannot currently access a Python installation. Install a supported Windows release before continuing.

  • Visit the official Python downloads page for Windows.
  • Install a stable Windows release of Python 3.10 or newer.
  • Restart Visual Studio Code from Windows search after the installation finishes.
  • Return to the cpu-inference-engine folder through File and Open Folder....
  • Open a PowerShell terminal through View and Terminal.
  • Confirm that PowerShell can access Python by running this command:
python --version

What should this confirm?

PowerShell should print a Python version of 3.10 or newer. Your Windows runtime is now available to the project.

Isolate the project environment

A project-local environment keeps this lab's packages separate from every other Python project on your computer. The .venv folder stores that isolated installation inside cpu-inference-engine.

  • Create the .venv environment by running this command:
py -m venv .venv

What does this command do?

The Windows Python launcher runs the built-in virtual environment module. It creates a project-specific Python installation inside .venv.

  • Confirm that .venv appears beneath cpu-inference-engine in the Explorer sidebar.

The new .venv folder proves that the isolated environment exists.

Virtual environment missing?

  • Confirm that the PowerShell path includes cpu-inference-engine.
  • Confirm that the earlier version check reported Python 3.10 or newer.
  • Retry the environment command after correcting the terminal location.

Need another pair of eyes? Help me diagnose why the virtual environment was not created.

  • Activate the new environment by running this command:
.venv\Scripts\activate

What does activation change?

Activation points package commands at the project-local environment. The PowerShell prompt identifies .venv as the active environment.

  • Check the start of the PowerShell prompt for the .venv environment name.

Your prompt now shows that .venv is active. Future package commands target the isolated lab environment.

Activation not showing?

  • Confirm that the terminal remains inside cpu-inference-engine.
  • Confirm that .venv exists in the Explorer sidebar.
  • Follow the PowerShell script policy approved for your Windows device if activation is blocked.

Still stuck? Help me diagnose the virtual environment activation problem.

Pin and install the dependencies

A requirements file records the exact package versions used by the lab. Reinstalling from that file reproduces the same dependency choices later.

  • Select the new-file control beside cpu-inference-engine in the Explorer sidebar.
  • Enter requirements.txt as the file name.
  • Press Enter to create the file.
  • Populate requirements.txt with the pinned packages by pasting this content:
torch==2.14.1
transformers==5.19.0

What do these pins do?

  • The torch==2.14.1 entry fixes the PyTorch package used for tensor operations and model execution.
  • The transformers==5.19.0 entry fixes the Transformers package used to load the tokenizer and model later.
  • Save requirements.txt.
  • Confirm that the saved file appears beneath cpu-inference-engine in the Explorer sidebar.
  • Confirm that the editor shows exactly two package lines.

That dependency file is locked in. Every install now starts from the same two versions.

Package lines look different?

  • Confirm that the file name is exactly requirements.txt.
  • Remove any extra package lines from the file.
  • Compare both version pins with the full-file reference below.

Need help checking the file? Help me compare my requirements file with the expected package pins.

✔️ Awesome, I've got everything!

Your saved requirements.txt now pins both libraries for repeatable installs.

ⓧ I'd like to double check the full code

torch==2.14.1
transformers==5.19.0

The first install may take several minutes. A busy terminal means PowerShell is still downloading or installing packages.

Before you run the installer, what result would prove that the pinned environment is ready?

  • Install the pinned dependencies into the active environment by running this command:
py -m pip install -r requirements.txt

What does this command do?

The Windows Python launcher runs the package installer against requirements.txt. The active virtual environment directs the installed packages into .venv.

The terminal should complete the installation without an error. Its log should report the pinned PyTorch and Transformers packages.

That is the lab locked down. Your PowerShell session now has the dependencies required for the manual decode loop.

Dependency installation failed?

  • Confirm that the PowerShell prompt identifies .venv as active.
  • Confirm that requirements.txt matches the full-file reference.
  • Confirm that your internet connection is available for the package downloads.

Need help reading the terminal output? Help me troubleshoot the pinned dependency installation.

Your native-Windows inference lab is ready. Next, you will turn model logits into a visible text continuation with your own decode loop.

Build the Naive Decode Loop

Your isolated Windows lab now has PyTorch and Hugging Face Transformers installed inside .venv.

A manual autoregressive decoding loop exposes the logits behind every generated token. This first version keeps the model inputs visible on every pass.

In this step, get ready to:
  • Configure DistilGPT2 for local CPU inference.
  • Build greedy decoding plus temperature and top-k sampling.
  • Run a manual decode loop that generates a visible text continuation.
Configure model settings and token selection

DistilGPT2 produces vocabulary scores for the next token. Your sampler converts those scores into one token ID using greedy decoding or temperature plus top-k sampling.

  • In the Visual Studio Code file sidebar, select the cpu-inference-engine folder.
  • Select the new-file icon at the top of the file sidebar.
  • Type engine.py as the file name.
  • Press Enter to create the file.
  • Add the shared engine configuration by pasting this code into engine.py:
import time

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL_ID = "distilbert/distilgpt2"
CACHE_DIR = ".model-cache"
MAX_NEW_TOKENS = 30
DO_SAMPLE = False
TEMPERATURE = 0.8
TOP_K = 40
SEED = 42

What does this configuration control?

  • The imports provide timing tools plus the model and tokenizer APIs.
  • The MODEL_ID value identifies the pretrained model.
  • The CACHE_DIR value keeps downloaded model files inside .model-cache.
  • The generation settings control the token limit and sampling behavior.
  • The SEED value supports repeatable sampling later.
  • Save engine.py.
  • Confirm that engine.py is listed beside requirements.txt in the file sidebar.

Cannot find engine.py?

  • Confirm that the file name ends with .py.
  • Confirm that the file sits directly inside cpu-inference-engine.

Help me locate or create engine.py.

Greedy decoding always chooses the highest-scoring token. Sampling scales the scores with TEMPERATURE before choosing from the strongest TOP_K candidates.

  • Add token selection below SEED = 42 by pasting this function:
def select_next_token(logits):
    if not DO_SAMPLE:
        return torch.argmax(logits, dim=-1, keepdim=True)

    top_values, top_indices = torch.topk(
        logits / TEMPERATURE,
        k=TOP_K,
        dim=-1,
    )
    probabilities = torch.softmax(top_values, dim=-1)
    sampled_position = torch.multinomial(probabilities, num_samples=1)
    return top_indices.gather(dim=-1, index=sampled_position)

How does token selection work?

  • The DO_SAMPLE check selects deterministic greedy decoding by default.
  • The sampling path scales the logits with TEMPERATURE.
  • The top-k operation keeps the strongest candidate scores.
  • Softmax converts those scores into probabilities.
  • Multinomial sampling selects one candidate token.
  • Save engine.py.
  • Check the file for syntax problems by running:
py engine.py

What should you see?

PowerShell returns to its prompt without a traceback. That clean exit confirms Python can parse the configuration and sampler.

Seeing a Python error?

  • Confirm that each indented line in select_next_token uses four spaces.
  • Compare the parentheses in your function with the code block above.

Help me debug the select_next_token function.

Generate tokens with the full sequence

Each decode pass needs the prompt IDs plus the tokens selected so far. The loop also grows an attention mask so every position remains available to the model.

  • Add the naive generation function below select_next_token by pasting this code:
@torch.inference_mode()
def generate_naive(model, prompt_ids, prompt_attention_mask, eos_token_id):
    generated_ids = prompt_ids
    attention_mask = prompt_attention_mask
    started_at = time.perf_counter()

    for _ in range(MAX_NEW_TOKENS):
        outputs = model(
            input_ids=generated_ids,
            attention_mask=attention_mask,
            use_cache=False,
        )
        next_token = select_next_token(outputs.logits[:, -1, :])
        generated_ids = torch.cat([generated_ids, next_token], dim=-1)
        attention_mask = torch.cat(
            [attention_mask, torch.ones_like(next_token)],
            dim=-1,
        )

        if next_token.item() == eos_token_id:
            break

    elapsed = time.perf_counter() - started_at
    return generated_ids, elapsed

What does the naive loop do?

  • Inference mode removes gradient tracking from the generation loop.
  • The model receives the current generated_ids tensor on every pass.
  • The slice outputs.logits[:, -1, :] isolates the scores at the latest sequence position.
  • The selected token extends both the generated IDs and the attention mask.
  • The end-of-sequence check can stop generation before the token limit.
  • The elapsed value measures time spent inside the decode loop.
  • Save engine.py.
  • Check the expanded file for syntax problems by running:
py engine.py

What did this check prove?

PowerShell returns to its prompt without a traceback. The complete generate_naive function now parses successfully.

Does the function fail to parse?

  • Confirm that @torch.inference_mode() sits directly above generate_naive.
  • Check the indentation inside the token loop.
  • Check the closing parentheses around the attention-mask update.

Help me debug the generate_naive function.

Load the model and run the generator

The main path connects user input to tokenization and generation. It also decodes the resulting token IDs into readable text.

The first run downloads the model and tokenizer into .model-cache. Give the download a few minutes if your connection is busy.

  • Add the main path below generate_naive by pasting this code:
def main():
    prompt = input("Prompt: ").strip() or "The future of local AI is"

    print("Loading DistilGPT2 on CPU...")
    tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, cache_dir=CACHE_DIR)
    model = AutoModelForCausalLM.from_pretrained(MODEL_ID, cache_dir=CACHE_DIR)

    encoded = tokenizer(prompt, return_tensors="pt")
    prompt_ids = encoded["input_ids"]
    prompt_attention_mask = torch.ones_like(prompt_ids)

    torch.manual_seed(SEED)
    naive_ids, _ = generate_naive(
        model,
        prompt_ids,
        prompt_attention_mask,
        model.config.eos_token_id,
    )

    naive_tokens = naive_ids.shape[1] - prompt_ids.shape[1]
    print("\nNaive output:")
    print(tokenizer.decode(naive_ids[0], skip_special_tokens=True))
    print(f"Generated tokens: {naive_tokens}")


if __name__ == "__main__":
    main()

How does the main path connect everything?

  • The prompt falls back to a default sentence when the submitted input is empty.
  • The tokenizer and model use the same MODEL_ID.
  • The cache_dir arguments keep downloaded files inside the project.
  • The tokenizer returns the prompt IDs consumed by the model.
  • The seed fixes the random state before generation.
  • The final calculation counts only the newly generated tokens.
  • Save engine.py.

✔️ Awesome, I've got everything!

Your engine.py file now contains the model settings, sampler, naive decode loop, and command-line entry point.

ⓧ I'd like to double check the full code

import time

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL_ID = "distilbert/distilgpt2"
CACHE_DIR = ".model-cache"
MAX_NEW_TOKENS = 30
DO_SAMPLE = False
TEMPERATURE = 0.8
TOP_K = 40
SEED = 42


def select_next_token(logits):
    if not DO_SAMPLE:
        return torch.argmax(logits, dim=-1, keepdim=True)

    top_values, top_indices = torch.topk(
        logits / TEMPERATURE,
        k=TOP_K,
        dim=-1,
    )
    probabilities = torch.softmax(top_values, dim=-1)
    sampled_position = torch.multinomial(probabilities, num_samples=1)
    return top_indices.gather(dim=-1, index=sampled_position)


@torch.inference_mode()
def generate_naive(model, prompt_ids, prompt_attention_mask, eos_token_id):
    generated_ids = prompt_ids
    attention_mask = prompt_attention_mask
    started_at = time.perf_counter()

    for _ in range(MAX_NEW_TOKENS):
        outputs = model(
            input_ids=generated_ids,
            attention_mask=attention_mask,
            use_cache=False,
        )
        next_token = select_next_token(outputs.logits[:, -1, :])
        generated_ids = torch.cat([generated_ids, next_token], dim=-1)
        attention_mask = torch.cat(
            [attention_mask, torch.ones_like(next_token)],
            dim=-1,
        )

        if next_token.item() == eos_token_id:
            break

    elapsed = time.perf_counter() - started_at
    return generated_ids, elapsed


def main():
    prompt = input("Prompt: ").strip() or "The future of local AI is"

    print("Loading DistilGPT2 on CPU...")
    tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, cache_dir=CACHE_DIR)
    model = AutoModelForCausalLM.from_pretrained(MODEL_ID, cache_dir=CACHE_DIR)

    encoded = tokenizer(prompt, return_tensors="pt")
    prompt_ids = encoded["input_ids"]
    prompt_attention_mask = torch.ones_like(prompt_ids)

    torch.manual_seed(SEED)
    naive_ids, _ = generate_naive(
        model,
        prompt_ids,
        prompt_attention_mask,
        model.config.eos_token_id,
    )

    naive_tokens = naive_ids.shape[1] - prompt_ids.shape[1]
    print("\nNaive output:")
    print(tokenizer.decode(naive_ids[0], skip_special_tokens=True))
    print(f"Generated tokens: {naive_tokens}")


if __name__ == "__main__":
    main()

How to read this reference

This is the complete engine.py file for this step. Its order matches the configuration, sampler, decode loop, and main path you added.

The script pauses at Prompt: before loading the model. The submitted prompt becomes the starting sequence for generation.

Before you run this, do you expect the loop to produce one token or a full continuation?

  • Start the naive generator by running:
py engine.py

What does this command start?

The command starts engine.py inside your activated environment. The script waits at its prompt before model loading begins.

  • Type The future of local AI is at the prompt.
  • Press Enter to start generation.

You'll first see Loading DistilGPT2 on CPU... while the model becomes ready. You'll then see Naive output followed by a continuation and a Generated tokens: count.

That is your first local continuation running through a decode loop you wrote yourself.

The naive loop repeats work

The generator works, but every pass supplies the entire growing generated_ids tensor to the model. The prompt and every earlier generated token go through the model again.

The line use_cache=False makes this full-context recomputation explicit. This is the intended limitation of the first working version.

Generator not reaching Naive output?

  • Confirm that your internet connection is available for the first model download.
  • Confirm that the PowerShell prompt still shows the activated .venv environment.
  • Compare your MODEL_ID and CACHE_DIR values with the full-file reference.

Help me diagnose why py engine.py is not completing.

Your naive generator now traces a prompt from token IDs to a decoded continuation. Next, you'll change what the model receives after its first pass.

Add a KV-Cached Decode Path

Your naive DistilGPT2 generator now turns a prompt into a continuation. Each new token still sends the full prompt plus every earlier token through the model.

A KV cache preserves the earlier attention keys and values. After the prefill, the model receives only the unprocessed token.

In this step, get ready to:
  • Add a cached generation path that begins with the full prompt.
  • Carry the cache plus the full attention mask through every decode iteration.
  • Run the naive path plus the cached path from the same prompt.
Add the cached generation loop

The first model call processes the complete prompt. Every later call reuses the saved cache with one new token.

  • In engine.py, find the final return generated_ids, elapsed line inside generate_naive().
  • Add the cached function below that line by copying this code:
@torch.inference_mode()
def generate_cached(model, prompt_ids, prompt_attention_mask, eos_token_id):
    generated_ids = prompt_ids
    attention_mask = prompt_attention_mask
    model_inputs = prompt_ids
    past_key_values = None
    started_at = time.perf_counter()

    for _ in range(MAX_NEW_TOKENS):
        outputs = model(
            input_ids=model_inputs,
            attention_mask=attention_mask,
            past_key_values=past_key_values,
            use_cache=True,
        )
        next_token = select_next_token(outputs.logits[:, -1, :])
        past_key_values = outputs.past_key_values
        generated_ids = torch.cat([generated_ids, next_token], dim=-1)

        if next_token.item() == eos_token_id:
            break

        attention_mask = torch.cat(
            [attention_mask, torch.ones_like(next_token)],
            dim=-1,
        )
        model_inputs = next_token

    elapsed = time.perf_counter() - started_at
    return generated_ids, elapsed

What does this code do?

  • The initial model_inputs value contains the full prompt for the prefill pass.
  • Setting use_cache=True asks the model to return reusable attention state.
  • The past_key_values variable carries that state into the next model call.
  • The attention_mask continues to represent the full sequence length.
  • The next model_inputs value contains only the token that the cache has not processed.
  • Save engine.py.
  • Check that the new function preserves the naive generator by running this command:
py engine.py

What does this run check?

The command starts the script from your activated environment. The current main path still calls generate_naive(), so its output confirms the new function parses without breaking the baseline.

  • Enter The future of local AI is when Prompt: appears.

The terminal can sit quietly while your CPU produces the continuation. You should then see Naive output plus a generated-token count.

Script stopping after the edit?

Compare the indentation of the loop with the function above it. Make sure generate_cached() begins after generate_naive() has completely ended.

If the naive result disappears, confirm that you added the new function without replacing generate_naive().

Help me debug the cached generation function.

Call both generation paths

A function only runs when main() calls it. The main path now needs to invoke both generators before decoding their results.

  • In engine.py, find the block that begins with naive_tokens = naive_ids.shape[1] - prompt_ids.shape[1].
  • Delete that block through the print(f"Generated tokens: {naive_tokens}") line.
  • Replace the deleted block by copying this code:
    torch.manual_seed(SEED)
    cached_ids, _ = generate_cached(
        model,
        prompt_ids,
        prompt_attention_mask,
        model.config.eos_token_id,
    )

    print("\nNaive output:")
    print(tokenizer.decode(naive_ids[0], skip_special_tokens=True))
    print("\nCached output:")
    print(tokenizer.decode(cached_ids[0], skip_special_tokens=True))

How does the main path change?

  • The second torch.manual_seed(SEED) call resets token selection before cached decoding.
  • The generate_cached() call receives the same prompt tensors plus the same end-of-sequence token ID.
  • The two decode calls print separate continuations under Naive output plus Cached output.
  • Save engine.py.

✔️ Awesome, I've got everything!

Your saved file now contains both generation paths. The cached path is ready for a complete run.

ⓧ I'd like to double check the full code

import time

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL_ID = "distilbert/distilgpt2"
CACHE_DIR = ".model-cache"
MAX_NEW_TOKENS = 30
DO_SAMPLE = False
TEMPERATURE = 0.8
TOP_K = 40
SEED = 42


def select_next_token(logits):
    if not DO_SAMPLE:
        return torch.argmax(logits, dim=-1, keepdim=True)

    top_values, top_indices = torch.topk(
        logits / TEMPERATURE,
        k=TOP_K,
        dim=-1,
    )
    probabilities = torch.softmax(top_values, dim=-1)
    sampled_position = torch.multinomial(probabilities, num_samples=1)
    return top_indices.gather(dim=-1, index=sampled_position)


@torch.inference_mode()
def generate_naive(model, prompt_ids, prompt_attention_mask, eos_token_id):
    generated_ids = prompt_ids
    attention_mask = prompt_attention_mask
    started_at = time.perf_counter()

    for _ in range(MAX_NEW_TOKENS):
        outputs = model(
            input_ids=generated_ids,
            attention_mask=attention_mask,
            use_cache=False,
        )
        next_token = select_next_token(outputs.logits[:, -1, :])
        generated_ids = torch.cat([generated_ids, next_token], dim=-1)
        attention_mask = torch.cat(
            [attention_mask, torch.ones_like(next_token)],
            dim=-1,
        )

        if next_token.item() == eos_token_id:
            break

    elapsed = time.perf_counter() - started_at
    return generated_ids, elapsed


@torch.inference_mode()
def generate_cached(model, prompt_ids, prompt_attention_mask, eos_token_id):
    generated_ids = prompt_ids
    attention_mask = prompt_attention_mask
    model_inputs = prompt_ids
    past_key_values = None
    started_at = time.perf_counter()

    for _ in range(MAX_NEW_TOKENS):
        outputs = model(
            input_ids=model_inputs,
            attention_mask=attention_mask,
            past_key_values=past_key_values,
            use_cache=True,
        )
        next_token = select_next_token(outputs.logits[:, -1, :])
        past_key_values = outputs.past_key_values
        generated_ids = torch.cat([generated_ids, next_token], dim=-1)

        if next_token.item() == eos_token_id:
            break

        attention_mask = torch.cat(
            [attention_mask, torch.ones_like(next_token)],
            dim=-1,
        )
        model_inputs = next_token

    elapsed = time.perf_counter() - started_at
    return generated_ids, elapsed


def main():
    prompt = input("Prompt: ").strip() or "The future of local AI is"

    print("Loading DistilGPT2 on CPU...")
    tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, cache_dir=CACHE_DIR)
    model = AutoModelForCausalLM.from_pretrained(MODEL_ID, cache_dir=CACHE_DIR)

    encoded = tokenizer(prompt, return_tensors="pt")
    prompt_ids = encoded["input_ids"]
    prompt_attention_mask = torch.ones_like(prompt_ids)

    torch.manual_seed(SEED)
    naive_ids, _ = generate_naive(
        model,
        prompt_ids,
        prompt_attention_mask,
        model.config.eos_token_id,
    )

    torch.manual_seed(SEED)
    cached_ids, _ = generate_cached(
        model,
        prompt_ids,
        prompt_attention_mask,
        model.config.eos_token_id,
    )

    print("\nNaive output:")
    print(tokenizer.decode(naive_ids[0], skip_special_tokens=True))
    print("\nCached output:")
    print(tokenizer.decode(cached_ids[0], skip_special_tokens=True))


if __name__ == "__main__":
    main()

How to use this reference

Compare the placement of generate_cached() with your file. Check that main() calls both generators before printing their decoded outputs.

Run the cached path

This run checks the complete cache lifecycle. The first cached iteration receives the prompt before later iterations receive one new token each.

  • Predict whether the naive path plus the cached path produce the same continuation with greedy selection.
  • Start both generation paths by running this command:
py engine.py

What does this run prove?

The script invokes the full-context baseline before invoking the cached path. The cached path retains outputs.past_key_values before changing model_inputs to the newly selected token.

  • Enter The future of local AI is when Prompt: appears.

You should see Naive output followed by a generated continuation. You should then see Cached output followed by the cached continuation.

Cached output missing?

Confirm that main() calls generate_cached() before the two output headings.

If the continuations differ, check that DO_SAMPLE = False remains unchanged. Confirm that both paths receive the same prompt tensors.

Help me debug the cached output.

You now have a working cache lifecycle with one full-prompt prefill followed by single-token decode inputs. That is the key optimization behind your second inference path.

Next, you'll time both paths and prove that caching preserves the generated token sequence.

Prove the Cache with a Benchmark

Both decode paths now generate text from DistilGPT2. The cached path reuses earlier attention state through a KV cache.

A cache is useful only when it preserves the generated tokens while changing performance. A reproducible benchmark gives each path the same settings plus seed.

The final report measures observed performance on your CPU. It also checks whether both paths return identical token IDs.

In this step, get ready to:
  • Exclude startup work from the measured decode loops with a warm-up pass.
  • Measure elapsed time plus tokens per second for each path.
  • Compare naive output with cached output across greedy decoding plus seeded sampling.
Warm up the model before timing

The first model call can include one-time startup work. A warm-up performs that work before either decode timer starts.

  • In engine.py, locate generate_naive() below select_next_token().
  • Add this function between them:
@torch.inference_mode()
def warm_up(model, prompt_ids, prompt_attention_mask):
    model(
        input_ids=prompt_ids,
        attention_mask=prompt_attention_mask,
        use_cache=False,
    )

What does this code do?

  • The warm_up() function performs one forward pass with the tokenized prompt.
  • The use_cache=False setting keeps this pass independent from the cached benchmark.
  • The @torch.inference_mode() decorator removes gradient-tracking overhead during inference.
  • In main(), find the line that creates prompt_attention_mask.
  • Add this call directly below it:
    warm_up(model, prompt_ids, prompt_attention_mask)

Why place the call here?

The model plus prompt are ready at this point. The warm-up finishes before either generation function starts its internal timer.

  • Save engine.py.

Prediction time: will both existing outputs still appear after the warm-up.

The script pauses at Prompt: so you can supply test text.

  • Return to the activated PowerShell terminal.
  • Rerun the engine by running this command:
py engine.py

What does this command do?

The py engine.py command starts the inference engine with the active Windows environment.

  • Enter The future of local AI is at the prompt.

You'll see Naive output followed by Cached output. Both generation paths still complete after the new warm-up pass.

That is the timing foundation in place: both decode paths still run after an untimed warm-up.

Warm-up call failing?

Check that warm_up() sits above generate_naive(). Confirm that its call sits below the line that creates prompt_attention_mask.

Help me debug the warm-up function in my inference engine.

Build the benchmark report

Each generation function already returns its elapsed time. The current calls discard those values with underscores.

The report needs those durations to calculate throughput. It also needs the number of tokens added beyond the original prompt.

  • In engine.py, locate the two generation calls inside main().
  • Find this current block:
    torch.manual_seed(SEED)
    naive_ids, _ = generate_naive(
        model,
        prompt_ids,
        prompt_attention_mask,
        model.config.eos_token_id,
    )

    torch.manual_seed(SEED)
    cached_ids, _ = generate_cached(
        model,
        prompt_ids,
        prompt_attention_mask,
        model.config.eos_token_id,
    )

What is missing here?

Each underscore discards the elapsed time returned by its generation function. The repeated torch.manual_seed(SEED) calls already give both paths the same random starting state.

  • Replace that block with this timed version:
    torch.manual_seed(SEED)
    naive_ids, naive_elapsed = generate_naive(
        model,
        prompt_ids,
        prompt_attention_mask,
        model.config.eos_token_id,
    )

    torch.manual_seed(SEED)
    cached_ids, cached_elapsed = generate_cached(
        model,
        prompt_ids,
        prompt_attention_mask,
        model.config.eos_token_id,
    )

What changes in this version?

The variables naive_elapsed plus cached_elapsed retain each measured duration. Resetting SEED before each path keeps sampled comparisons reproducible.

  • Find the two current output sections below the cached generation call.
  • Replace those output sections with this benchmark report:
    naive_tokens = naive_ids.shape[1] - prompt_ids.shape[1]
    cached_tokens = cached_ids.shape[1] - prompt_ids.shape[1]
    naive_rate = naive_tokens / naive_elapsed
    cached_rate = cached_tokens / cached_elapsed

    print("\nNaive output:")
    print(tokenizer.decode(naive_ids[0], skip_special_tokens=True))
    print("\nCached output:")
    print(tokenizer.decode(cached_ids[0], skip_special_tokens=True))

    print("\nBenchmark:")
    print(f"Mode: {'top-k sampling' if DO_SAMPLE else 'greedy'}")
    print(f"Token match: {torch.equal(naive_ids, cached_ids)}")
    print(
        f"Naive:  {naive_tokens} tokens in {naive_elapsed:.2f}s "
        f"({naive_rate:.2f} tokens/s)"
    )
    print(
        f"Cached: {cached_tokens} tokens in {cached_elapsed:.2f}s "
        f"({cached_rate:.2f} tokens/s)"
    )

How does the benchmark work?

  • The token counts subtract the prompt length from each completed sequence length.
  • The rate calculations divide newly generated tokens by the measured decode duration.
  • The torch.equal() check confirms that both tensors have identical shapes plus values.
  • The report prints each continuation before showing its mode plus measured throughput.
  • Save engine.py.

Prediction time: will greedy decoding produce a token match before you inspect the timing numbers.

  • Return to the activated PowerShell terminal.
  • Run the completed benchmark with this command:
py engine.py

What does this check prove?

This run exercises the untimed warm-up before measuring each decode path. It then compares the complete token tensors produced from the same prompt.

  • Enter The future of local AI is at the prompt.

You'll see Naive output plus Cached output. The Benchmark: section shows Token match: True.

Each path also reports elapsed seconds plus tokens/s based on your CPU's observed performance.

Strong result. Your engine now proves that the cache preserves deterministic decoding while changing the work performed for each new token.

Token match showing False?

Confirm that DO_SAMPLE = False remains in the settings. Check that torch.manual_seed(SEED) appears immediately before each generation call.

Help me debug mismatched naive and cached token IDs.

Exercise top-k sampling

Greedy decoding always chooses the highest-scoring token. Top-k sampling draws from the strongest candidates after applying temperature.

Your sampler uses a temperature of 0.8 plus a top-k value of 40. Resetting the same seed before each path keeps this sampled comparison reproducible.

  • Switch back to engine.py in VS Code.
  • Find DO_SAMPLE = False near the generation settings.
  • Replace False with True.
  • Save engine.py.

Prediction time: will resetting the seed let both sampled paths choose the same token sequence.

  • Return to the activated PowerShell terminal.
  • Run the sampling benchmark with this command:
py engine.py

What changes in this run?

The same manual sampler now applies temperature before selecting from the top candidates. Both paths start from the same seed before making their token choices.

  • Enter The future of local AI is at the prompt.

You'll see Mode: top-k sampling in the report. The report should still show Token match: True because each path begins with the same seed.

Interpret your timing

CPU timing varies with prompt length plus background load. Use token equality as the correctness check while treating the measured rates as observations from this run.

  • Switch back to engine.py in VS Code.
  • Replace True with False on the DO_SAMPLE line.
  • Save engine.py.

✔️ Awesome, I've got everything!

Your file now contains the warm-up pass plus both timed decode paths. It finishes with the token comparison plus throughput report.

ⓧ I'd like to double check the full code

import time

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL_ID = "distilbert/distilgpt2"
CACHE_DIR = ".model-cache"
MAX_NEW_TOKENS = 30
DO_SAMPLE = False
TEMPERATURE = 0.8
TOP_K = 40
SEED = 42


def select_next_token(logits):
    if not DO_SAMPLE:
        return torch.argmax(logits, dim=-1, keepdim=True)

    top_values, top_indices = torch.topk(
        logits / TEMPERATURE,
        k=TOP_K,
        dim=-1,
    )
    probabilities = torch.softmax(top_values, dim=-1)
    sampled_position = torch.multinomial(probabilities, num_samples=1)
    return top_indices.gather(dim=-1, index=sampled_position)


@torch.inference_mode()
def warm_up(model, prompt_ids, prompt_attention_mask):
    model(
        input_ids=prompt_ids,
        attention_mask=prompt_attention_mask,
        use_cache=False,
    )


@torch.inference_mode()
def generate_naive(model, prompt_ids, prompt_attention_mask, eos_token_id):
    generated_ids = prompt_ids
    attention_mask = prompt_attention_mask
    started_at = time.perf_counter()

    for _ in range(MAX_NEW_TOKENS):
        outputs = model(
            input_ids=generated_ids,
            attention_mask=attention_mask,
            use_cache=False,
        )
        next_token = select_next_token(outputs.logits[:, -1, :])
        generated_ids = torch.cat([generated_ids, next_token], dim=-1)
        attention_mask = torch.cat(
            [attention_mask, torch.ones_like(next_token)],
            dim=-1,
        )

        if next_token.item() == eos_token_id:
            break

    elapsed = time.perf_counter() - started_at
    return generated_ids, elapsed


@torch.inference_mode()
def generate_cached(model, prompt_ids, prompt_attention_mask, eos_token_id):
    generated_ids = prompt_ids
    attention_mask = prompt_attention_mask
    model_inputs = prompt_ids
    past_key_values = None
    started_at = time.perf_counter()

    for _ in range(MAX_NEW_TOKENS):
        outputs = model(
            input_ids=model_inputs,
            attention_mask=attention_mask,
            past_key_values=past_key_values,
            use_cache=True,
        )
        next_token = select_next_token(outputs.logits[:, -1, :])
        past_key_values = outputs.past_key_values
        generated_ids = torch.cat([generated_ids, next_token], dim=-1)

        if next_token.item() == eos_token_id:
            break

        attention_mask = torch.cat(
            [attention_mask, torch.ones_like(next_token)],
            dim=-1,
        )
        model_inputs = next_token

    elapsed = time.perf_counter() - started_at
    return generated_ids, elapsed


def main():
    prompt = input("Prompt: ").strip() or "The future of local AI is"

    print("Loading DistilGPT2 on CPU...")
    tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, cache_dir=CACHE_DIR)
    model = AutoModelForCausalLM.from_pretrained(MODEL_ID, cache_dir=CACHE_DIR)

    encoded = tokenizer(prompt, return_tensors="pt")
    prompt_ids = encoded["input_ids"]
    prompt_attention_mask = torch.ones_like(prompt_ids)

    warm_up(model, prompt_ids, prompt_attention_mask)

    torch.manual_seed(SEED)
    naive_ids, naive_elapsed = generate_naive(
        model,
        prompt_ids,
        prompt_attention_mask,
        model.config.eos_token_id,
    )

    torch.manual_seed(SEED)
    cached_ids, cached_elapsed = generate_cached(
        model,
        prompt_ids,
        prompt_attention_mask,
        model.config.eos_token_id,
    )

    naive_tokens = naive_ids.shape[1] - prompt_ids.shape[1]
    cached_tokens = cached_ids.shape[1] - prompt_ids.shape[1]
    naive_rate = naive_tokens / naive_elapsed
    cached_rate = cached_tokens / cached_elapsed

    print("\nNaive output:")
    print(tokenizer.decode(naive_ids[0], skip_special_tokens=True))
    print("\nCached output:")
    print(tokenizer.decode(cached_ids[0], skip_special_tokens=True))

    print("\nBenchmark:")
    print(f"Mode: {'top-k sampling' if DO_SAMPLE else 'greedy'}")
    print(f"Token match: {torch.equal(naive_ids, cached_ids)}")
    print(
        f"Naive:  {naive_tokens} tokens in {naive_elapsed:.2f}s "
        f"({naive_rate:.2f} tokens/s)"
    )
    print(
        f"Cached: {cached_tokens} tokens in {cached_elapsed:.2f}s "
        f"({cached_rate:.2f} tokens/s)"
    )


if __name__ == "__main__":
    main()

How to use this reference

This reference shows the exact final state of engine.py for this step.

Prediction time: will the restored greedy benchmark report identical token IDs plus two measured rates.

  • Return to the activated PowerShell terminal.
  • Run the final verification with this command:
py engine.py

What does the final check cover?

This run verifies the saved greedy configuration. It covers the warm-up pass plus both timed generation paths.

  • Enter The future of local AI is at the prompt.

You'll see Naive output followed by Cached output. The report shows Mode: greedy plus Token match: True.

The naive line plus cached line each show a generated-token count. They also show elapsed seconds plus tokens per second.

You have a reproducible CPU benchmark that checks correctness before comparing performance.

Secret mission

Stream Tokens as They Are Generated

Turn the cached decode path into a streaming experience that prints each token as soon as it is selected. Keep the naive baseline intact so the final benchmark can still prove that both paths generate matching token IDs.

Clean Up Your Resources

Clean Up Your Resources

Your inference engine runs entirely on your Windows computer. Keeping the project folder creates no ongoing cost.

Resources you used:

  • The requirements.txt file with pinned PyTorch and Transformers dependencies.
  • The engine.py file with your manual decode loops and benchmark.
  • The .venv folder containing your local Python environment and installed packages.
  • The .model-cache folder containing the downloaded tokenizer and DistilGPT2 files.

Keep everything running

No action needed. Choose this if you want to rerun the benchmark or keep experimenting with token streaming.

  • Leave cpu-inference-engine intact so later runs can reuse .venv and .model-cache.

Pause - I'll come back to this later

Shut down the open tools to pause your work. No server or background process remains running after the terminal closes.

  • Close the activated Windows PowerShell terminal in VS Code.
  • Close VS Code.
  • Leave the cpu-inference-engine folder intact for your next session.

Delete - I don't want to use this again

Deleting the folder removes every local artifact from this project. The cleanup stays contained because the project has no remote resources.

  • Close the activated Windows PowerShell terminal in VS Code.
  • Close VS Code.
  • Press the Windows key to open Windows search.
  • Type File Explorer into the search field.
  • Press Enter to open File Explorer.
  • Navigate to the location containing cpu-inference-engine.
  • Right-click the cpu-inference-engine folder.
  • Select Delete.

This removes requirements.txt, engine.py, .venv, and .model-cache with the project folder.

Nice Work!

Nice Work!

Outstanding work! Your local CPU inference engine now turns prompts into DistilGPT2 continuations through a decode loop you wrote yourself.

What you learned:

  • Built a manual autoregressive decoding loop that selects each new token from DistilGPT2 logits. Added greedy decoding plus temperature and top-k sampling inside select_next_token().
  • Exposed full-context recomputation in the naive path. Added KV caching through past_key_values to process only unprocessed token IDs after prefill.
  • Verified cache correctness with Token match: True under a shared seed. Reported elapsed seconds plus tokens per second for both inference paths.
  • Secret Mission: Extended generate_cached() with a token_callback so the cached continuation streams token by token before the benchmark report.

Ready to quiz yourself?