Skip to content

Ollama — Local LLM ​

Ollama lets you run open-source large language models directly on your Mac. There are no API fees, no account required, and your data never leaves your machine.

NOTE

Ollama requires an Apple Silicon Mac (M1/M2/M3/M4) or an Intel Mac for best performance. Models run entirely offline.


What Is Ollama? ​

Ollama is a free, open-source tool that manages and serves LLMs locally. It exposes a simple HTTP API (compatible with the OpenAI API format) that DuRT can connect to directly. Think of it as a private, local version of the OpenAI API.


Step 1 — Install Ollama ​

The easiest way is via Homebrew:

bash
brew install ollama

Alternatively, download the macOS app from ollama.com/download.

After installation, start the Ollama service:

bash
ollama serve

TIP

If you installed the macOS app, Ollama starts automatically in the menu bar. You don't need to run ollama serve manually.


Step 2 — Pull a Model ​

Download a model with the ollama pull command. For example:

bash
# A well-rounded general-purpose model (~2 GB)
ollama pull llama3.2

# A compact multilingual model — great for Chinese/Japanese/Korean (~5 GB)
ollama pull qwen2.5

# A sharp, fast European-language model (~7 GB)
ollama pull mistral-nemo

To see all models you've downloaded:

bash
ollama list

NOTE

Model files are stored in ~/.ollama/models. Make sure you have enough free disk space before pulling large models.


Step 3 — Add Ollama in DuRT ​

  1. Open DuRT and go to Settings → Services.
  2. Click + Add Service and choose Ollama.
  3. Set the Base URL to:
    http://localhost:11434
  4. Enter the Model Name exactly as it appears in ollama list (e.g. llama3.2, qwen2.5).
  5. Leave the API Key field blank — Ollama does not require authentication by default.
  6. Click Save.

TIP

You can add multiple Ollama services in DuRT — one per model — and switch between them easily.


ModelSizeBest For
llama3.2~2 GBGeneral use, English tasks, quick responses
qwen2.5~5 GBMultilingual (especially Chinese), reasoning
mistral-nemo~7 GBEuropean languages, instruction following
llama3.2:1b~1 GBUltra-fast, low RAM usage, simple tasks

TIP

Start with llama3.2 if you're unsure. It's small, fast, and capable for most everyday tasks.


Privacy & Local Execution ​

  • Zero data upload — all inference runs locally on your CPU/GPU.
  • No account needed — no sign-up, no API keys, no terms of service to worry about.
  • Works offline — once a model is downloaded, no internet connection is required.
  • Free forever — no usage limits, no billing.

IMPORTANT

Performance depends on your Mac's RAM and chip. Models larger than your available RAM will be slower due to memory swapping. 8 GB RAM supports models up to ~4 GB; 16 GB+ is recommended for 7B+ parameter models.

Released under the MIT License.