Run Ollama On Your Own Machine

Install Ollama and run your first local AI model on your own machine

Introduction

⚡️ 30 Second Summary

What if you could have your own AI assistant that runs entirely on your computer, with no internet required, no API costs, and your data never leaving your machine?

In this project, you'll install Ollama and run your first local LLM, chatting with an AI that runs entirely on your computer.

What You'll Build

You'll set up a complete local AI environment using Ollama, a tool that makes running language models on your own machine as simple as running a single command.

By the end of this project, you'll have:

  • 🦙 Ollama installed and running as a local service on your machine.
  • 🤖 The qwen2.5:0.5b model downloaded and ready to use.
  • 💬 An interactive chat experience with a local AI through your terminal.
  • 💎 Secret Mission: Create a custom AI persona with its own personality using a Modelfile.

Want a complete demo of how to do this project, from start to finish? Check out our 🎬 walkthrough with Pano

Are there any prerequisites?

This is Part 1 of the RAG API series. No prerequisites required, this is where you start!

Not sure if this project is right for you? Check if it matches your goals

If you're up for a bit of a challenge, quiz yourself on the key concepts up ahead in this project.

  1. Part 1: You are here!
  2. Part 2: Build a RAG API with FastAPI

Install Ollama

To set up our local AI environment, we first need to install Ollama, a tool that lets you run large language models (LLMs) directly on your own machine.

Running AI locally means no API costs, complete privacy, and the ability to work offline. In this step, you'll get Ollama installed and running.

In this step, get ready to:

  • Install Ollama on your machine.
  • Verify the Ollama service is running.

Install Ollama

🆕 I need to install Ollama

🍎 macOS

  • Click Download.
  • Click Download for macOS
  • Open the downloaded .dmg file.
  • Drag Ollama to your Applications folder.
  • Open Ollama from your Applications folder.

Ollama will start running as a service in your menu bar. Close the chat window if it opens, we'll use the terminal directly!

🖼️ Windows

  • Click Download for Windows
  • Open the downloaded .exe file from your Downloads folder.
  • Follow the installation prompts.

Ollama will start running as a service in your system tray. Close the chat window if it opens, we'll use the terminal directly!

🐧 Linux

  • Run the following command in your terminal:
curl -fsSL https://ollama.com/install.sh | sh
What does this command do?

This command downloads and runs the Ollama installer:

  • curl -fsSL downloads the install script silently
  • | pipes the script to the next command
  • sh executes the script

How is Ollama different to OpenAI or Anthropic?

Ollama is an open-source tool that makes it easy to run large language models locally. Instead of sending your prompts to a cloud API like OpenAI or Anthropic, Ollama runs the AI model directly on your computer. This gives you complete privacy, zero API costs, and the ability to work offline.

✔️ I already have Ollama

Great, you're ahead of the game.

  • Make sure Ollama is open and running on your machine.
  • Continue to Verify the Ollama Service below to make sure it's running.

Verify the Ollama Service

Let's make sure Ollama is running and ready to accept requests.

  • Open your terminal and run the following command:

🍎 macOS

  • Press Cmd + Space and type Terminal.
  • Run the following command:
curl http://localhost:11434

🖼️ Windows

  • Press the Win key and type PowerShell.
  • Run the following command:
curl.exe http://localhost:11434

🐧 Linux

  • Open your terminal app.
  • Run the following command:
curl http://localhost:11434

What does this command do?

curl sends a request to a URL. Here, you're sending a request to localhost:11434, which is where Ollama's local server listens. If the server is running, it'll respond with a confirmation message.

✔️ Ollama is running

You should see the response:

Ollama is running

What's happening here?

Ollama runs as a local server on port 11434. When you send a request to http://localhost:11434, you're checking if the server is up and ready to process AI requests. This is the same endpoint that applications use to interact with your local AI models.

You're all set! Next up, you'll download your first AI model and start chatting with it.

ⓧ Connection refused

If you see an error like Connection refused or Failed to connect, Ollama isn't running yet.

Try these fixes:

🍎 macOS

  • Open Ollama from your Applications folder.
  • Look for the llama icon in your menu bar.

🖼️ Windows

  • Check your system tray for the Ollama icon.
  • If it's not there, search for "Ollama" in the Start menu and launch it.

🐧 Linux

  • Start the Ollama service with:
ollama serve

After starting Ollama, run the curl command again to verify it's working.

Still stuck?

Get help with your error or share your error with the NextWork community!

With Ollama running, you're ready to download your first AI model and start chatting with it.

Pull and Test Your AI Model

With Ollama installed and running, it's time to download an AI model and start chatting with it. Unlike cloud-based AI services like ChatGPT or Claude, you'll be running this model entirely on your own machine. It will be your own personal AI!

In this step, get ready to:

  • Check your current models.
  • Pull the qwen2.5:0.5b model.
  • Chat with your local AI.

Check Your Current Models

Before downloading a new model, let's see what models (if any) are already installed on your machine.

  • Run the following command in your terminal:
ollama list

✔️ Empty list

If this is a fresh Ollama installation, you'll see an empty list or a header with no models listed. That's expected! Continue to the next section to pull your first model.

✔️ I see models listed

If you already have models installed from previous use, that's fine! You can still follow along with this project. You'll add the qwen2.5:0.5b model to your collection.

Pull Your First AI Model

Now let's download your first AI model. In AI, a "model" isn't a version number. It's the actual program that has been trained to understand and generate text. Models are what power tools like ChatGPT and Claude behind the scenes.

There are many models to choose from, and you can browse them all at Ollama's model library. We'll use one called qwen2.5:0.5b.

What does qwen2.5:0.5b mean?

qwen2.5 is a model developed by Alibaba. We chose it because it's free, lightweight, and great for learning. The 0.5b is the size: 0.5 billion parameters.

Parameters are the patterns the model picked up during training. Think of them as the model's "brain cells." More parameters means smarter, but also bigger and slower. For context, 0.5b is tiny (~400MB), 7b is medium (~4GB), and 70b is large (40GB+).

We're starting small so it runs easily on any machine. What do these model sizes mean?

  • Run the following command:
ollama pull qwen2.5:0.5b

This downloads the qwen2.5:0.5b model to your machine. The download is approximately 400MB, so it may take a few minutes depending on your internet connection.

✔️ Download completed

  • Verify the model is installed:
ollama list

You should now see qwen2.5:0.5b in your list of available models.

ⓧ I see an error

If the download fails, here are common issues and solutions:

Network timeout or connection error:

  • Check your internet connection.
  • Try running the command again. Downloads can resume where they left off.

Disk space error:

  • The model needs approximately 400MB of free space.
  • Free up disk space and try again.

Permission denied:

  • On Linux, you may need to run with sudo or fix permissions.
  • Run the pull command again:
ollama pull qwen2.5:0.5b

Still stuck?

Get help with your error or share your error with the NextWork community!

Chat with Your Local AI

Now for the fun part. Let's have a conversation with your local AI model! You can start an interactive chat session right from the terminal.

  • Run the following command:
ollama run qwen2.5:0.5b

What does ollama run do?

While ollama pull only downloads a model, ollama run loads the model into memory and starts an interactive chat session. Think of pull as installing an app, and run as opening it.

You'll see a prompt where you can type messages to the AI.

  • Enter the following prompt:
What is the capital of France?

The AI should respond with information about Paris. You're now chatting with an AI running entirely on your own machine!

Notice anything different from ChatGPT or Claude?

This AI is running entirely on your machine. No internet connection needed and no data sent to any company. The response might be simpler than what you'd get from ChatGPT, because we're using a tiny model (0.5b parameters). Cloud AI services use much larger models (hundreds of billions of parameters) running on powerful servers, which is why they cost money to use.

  • Enter the following to exit the chat:
/bye

You've successfully downloaded a model and chatted with your local AI. Next, let's explore what local AI models can and can't do.

Understand Local AI Limitations

You've seen how your local AI can answer general knowledge questions, but how much does it actually know? In this step, you'll discover an important limitation of local AI models and learn about a technique called RAG (Retrieval Augmented Generation) that addresses it.

In this step, get ready to:

  • Discover what local AI models can and can't do.
  • Learn about RAG.

Discover What Local AI Can't Do

Let's demonstrate an important limitation of local AI models.

  • Run the following command to start a new chat session:
ollama run qwen2.5:0.5b

Why are we starting a new chat?

Each ollama run session starts fresh with no memory of previous conversations. This is important to know, local AI models don't retain context between sessions unless you build that capability yourself.

  • Enter the following prompt:
What's my name?

Notice how the AI didn't actually answer your question. It doesn't know your name, but instead of saying "I don't know," it tried to be helpful anyway. This is common with smaller AI models.

Why doesn't the AI know about me?

Large language models like qwen2.5 are trained on general knowledge from the internet, but they don't have any information about you personally. They can't access your files, your emails, or any private data.

This limitation is exactly why techniques like RAG (Retrieval-Augmented Generation) exist. It's a way to give AI access to your own data. What is RAG?

  • Enter the following to exit the chat:
/bye

Your local AI environment is now fully set up. You've installed Ollama, downloaded a model, and experienced both its capabilities and limitations. Ready to take it further? The Secret Mission awaits!

Secret mission

You've seen how qwen2.5 responds to questions, but what if you could give it a specific personality? With Ollama, you can create custom AI personas using Modelfiles and system prompts, no retraining required.

In this secret mission, get ready to:

  • Pull a second AI model (gemma3:1b).
  • Create a Modelfile with a custom system prompt.
  • Build your custom AI persona.
  • Compare responses between base and custom models.

Create a Custom AI Persona

Clean Up Your Resources

Clean Up Your Resources

Now that you've set up your local AI environment and created a custom persona, let's review the resources on your machine.

Heads up!

Your Ollama setup and custom AI persona are great additions to your portfolio. Consider keeping them to showcase your local AI skills to employers.

Resources to manage:

  • Ollama installation
  • Downloaded AI models (qwen2.5:0.5b, gemma3:1b)
  • Custom model (coding-tutor)
  • modelfiles directory

🟢 Keep everything running

Your models and Ollama are stored locally on your machine. There are no ongoing costs, just disk space (~1.2GB total for both models).

No action needed, keep experimenting!

🟡 Remove models but keep Ollama

If you want to free up disk space but keep Ollama installed for future projects:

  • Remove the downloaded models:
ollama rm qwen2.5:0.5b
ollama rm gemma3:1b
ollama rm coding-tutor
  • Delete the modelfiles directory:

🍎 macOS / Linux

rm -rf ~/modelfiles

🖼️ Windows

Remove-Item -Recurse -Force "$env:USERPROFILE\modelfiles"

You can always pull new models later with ollama pull.

🔴 Delete everything

If you're completely done with local AI:

  • Remove all models:
ollama rm qwen2.5:0.5b
ollama rm gemma3:1b
ollama rm coding-tutor
  • Delete the modelfiles directory:

🍎 macOS / Linux

rm -rf ~/modelfiles

🖼️ Windows

Remove-Item -Recurse -Force "$env:USERPROFILE\modelfiles"
  • Uninstall Ollama:

🍎 macOS

Move Ollama from your Applications folder to the Trash, then empty the Trash.

🖼️ Windows

Open Settings > Apps > Installed apps, find Ollama, and click Uninstall.

🐧 Linux

sudo rm $(which ollama)

Should I uninstall Ollama?

If you're continuing with the RAG API series, you'll need Ollama for upcoming projects. We recommend keeping it installed.

That's a wrap!

That's a wrap!

Nice work! 🚀 You've just set up your own local AI environment with Ollama and chatted with an AI running entirely on your machine.

You've learned how to:

  • 🦙 Install Ollama and run it as a local service on your machine.
  • 🤖 Pull AI models like qwen2.5:0.5b from Ollama's model library.
  • 💬 Start interactive chat sessions with your local AI through the terminal.
  • 🔒 Understand why local AI models can't access personal data (and why that's a feature!).
  • 💎 Create a custom AI persona using Modelfiles and system prompts.

Congratulations! 🎉

You now have the foundation to run AI locally with no API costs, complete privacy, and the ability to work offline.

Ready to quiz yourself? 💪

What's Next in the DevOps × AI Series?

You've completed Part 1 of the series! Continue building your local AI skills:

  • 📚 Project 2: Build a RAG Chatbot - Give your local AI access to your own documents using Retrieval-Augmented Generation.
  • 🔧 Project 3: Create an AI-Powered CLI Tool - Build a command-line assistant that combines your custom persona with RAG.

p.s. Does it say "Still tasks to complete!" at the bottom of the screen?

This means you still have screenshots left to upload, or questions left to answer!

  1. Press Ctrl+F (Windows) or Command+F (Mac) on your keyboard.
  2. Search for the text Return to later.
  3. Jump straight to your incomplete tasks!
  4. 🙋‍♀️ Still stuck? Ask the community!