Ollama — Local LLM
Ollama lets you run open-source large language models directly on your Mac. There are no API fees, no account required, and your data never leaves your machine.
NOTE
Ollama requires an Apple Silicon Mac (M1/M2/M3/M4) or an Intel Mac for best performance. Models run entirely offline.
What Is Ollama?
Ollama is a free, open-source tool that manages and serves LLMs locally. It exposes a simple HTTP API (compatible with the OpenAI API format) that DuRT can connect to directly. Think of it as a private, local version of the OpenAI API.
Step 1 — Install Ollama
The easiest way is via Homebrew:
brew install ollamaAlternatively, download the macOS app from ollama.com/download.
After installation, start the Ollama service:
ollama serveTIP
If you installed the macOS app, Ollama starts automatically in the menu bar. You don't need to run ollama serve manually.
Step 2 — Pull a Model
Download a model with the ollama pull command. For example:
# A well-rounded general-purpose model (~2 GB)
ollama pull llama3.2
# A compact multilingual model — great for Chinese/Japanese/Korean (~5 GB)
ollama pull qwen2.5
# A sharp, fast European-language model (~7 GB)
ollama pull mistral-nemoTo see all models you've downloaded:
ollama listNOTE
Model files are stored in ~/.ollama/models. Make sure you have enough free disk space before pulling large models.
Step 3 — Add Ollama in DuRT
- Open DuRT and go to Settings → Services.
- Click + Add Service and choose Ollama.
- Set the Base URL to:
http://localhost:11434 - Enter the Model Name exactly as it appears in
ollama list(e.g.llama3.2,qwen2.5). - Leave the API Key field blank — Ollama does not require authentication by default.
- Click Save.
TIP
You can add multiple Ollama services in DuRT — one per model — and switch between them easily.
Recommended Models
| Model | Size | Best For |
|---|---|---|
llama3.2 | ~2 GB | General use, English tasks, quick responses |
qwen2.5 | ~5 GB | Multilingual (especially Chinese), reasoning |
mistral-nemo | ~7 GB | European languages, instruction following |
llama3.2:1b | ~1 GB | Ultra-fast, low RAM usage, simple tasks |
TIP
Start with llama3.2 if you're unsure. It's small, fast, and capable for most everyday tasks.
Privacy & Local Execution
- Zero data upload — all inference runs locally on your CPU/GPU.
- No account needed — no sign-up, no API keys, no terms of service to worry about.
- Works offline — once a model is downloaded, no internet connection is required.
- Free forever — no usage limits, no billing.
IMPORTANT
Performance depends on your Mac's RAM and chip. Models larger than your available RAM will be slower due to memory swapping. 8 GB RAM supports models up to ~4 GB; 16 GB+ is recommended for 7B+ parameter models.
