# Ollama

Published January 1, 2026 · Updated August 10, 2026

Run large language models locally on your Mac

Category: Developer Tools · Free · Replaces ChatGPT Plus ($20/month) · [Official website](https://ollama.com)

## Install with Homebrew

`brew install --cask ollama-app`

## Quick take

Ollama is still the default local LLM runtime on Mac in August 2026—now with a credible cloud tier and MLX-era Apple Silicon speed. Stay on free local for private coding assistants; buy Pro when you need hosted open models and concurrency. LM Studio is the friendly GUI alternative; raw MLX is the enthusiast path. For most developers, install Ollama first.

Best for: Mac developers running private local assistants, Agent builders who need a localhost OpenAI-compatible API, Apple Silicon users chasing MLX performance, Teams evaluating open-model cloud with seat billing

## What is Ollama?

Ollama is the simplest way to run open large language models on a Mac. Install the app, run `ollama pull llama3.1` (or whatever model you need), then `ollama run`—or use the desktop UI and local OpenAI-compatible API on localhost. Weights live on your machine; prompts do not need to leave your network for local models.

August 2026 reality: Ollama is very much alive, well funded, and dual-mode. Local inference remains free and unlimited on your hardware. Cloud models and higher concurrency are metered through Free / Pro ($20/mo or $200/yr) / Max ($100/mo; new Max sign-ups were paused while capacity expands) plans, plus Team ($25/seat/mo, 5-seat minimum) and Enterprise. Official pricing emphasizes that running models on your own hardware is always unlimited; cloud usage has session and weekly limits that scale by plan.

The Mac story of 2026 is MLX. Starting with the March 2026 preview (Ollama 0.19 era) and continuing through June performance drops (GGUF+MLX work, highest MLX performance updates, faster Gemma 4 MTP paths), Ollama on Apple Silicon leans on Apple's MLX framework and unified memory—especially strong on M5-class GPUs with Neural Accelerators. Community and official posts report large prefill/decode speedups versus older llama.cpp Metal paths. Keep models in modern formats the engine expects; dusty quantizations may not inherit every speedup.

Funding note: Ollama's July 2026 financing coverage (including the $88M round referenced in industry notes) signals cloud capacity build-out, not abandonment of local. Competitors include LM Studio (GUI-first local), MLX-native tools and llama.cpp front-ends, and cloud gateways. Ollama wins on CLI ergonomics, library UX (`ollama.com/library`), and the local+cloud combo from one account.

For Mac developers building agents, RAG demos, or private coding assistants on Tahoe 26.x Apple Silicon, Ollama is still the default on-ramp. Start local, add Pro cloud when you need larger hosted open models without buying a 128 GB Mac.

## Features

- **One-Command Local Models.** `ollama run` pulls and serves models with a chat REPL. Ideal for trying Llama, Gemma, Qwen, Mistral, DeepSeek, and other open weights without Python env pain.
- **Apple Silicon MLX Engine.** Mac builds use MLX for high-throughput inference on unified memory, with major 2026 speed work (including M5 Neural Accelerator paths). Local runs stay private on-device.
- **OpenAI-Compatible Local API.** Point apps at localhost endpoints to swap cloud providers for local models in coding agents, note tools, and scripts.
- **Model Library and Modelfiles.** Browse official library tags, pin versions, and customize system prompts/parameters with Modelfiles for repeatable assistant setups.
- **Cloud Models (Free/Pro/Max).** Sign in to run hosted open models when local VRAM/unified memory is not enough. Free is light usage (1 concurrent); Pro $20/mo adds 50x usage and 3 concurrent; Max $100/mo targets heavy agent workloads (10 concurrent; new sign-ups may be paused).
- **Desktop Apps, CLI, and Integrations.** Native clients plus 40,000+ community integrations. Works beside Open WebUI, coding agents, and automation tools.
- **Team and Enterprise Controls.** Team seats ($25/seat/mo, 5 minimum) add shared billing and admin; Enterprise adds custom terms, security reviews, and deployment planning.
- **Private by Default Locally.** Local weights and prompts stay on disk/RAM you control. Cloud plans document no training on prompts/responses and zero-retention partner requirements.

## How to Install Ollama on Mac

Ollama installs cleanly via Homebrew and runs as a background service. The entire setup takes under two minutes on a typical broadband connection.

1. **Install via Homebrew.** Run `brew install --cask ollama-app` in your terminal. This installs the Ollama app, CLI, and sets it up as a launchd service that starts automatically.
2. **Start the Server.** Run `ollama serve` to start the API server, or if installed via Homebrew, it may already be running as a background service. Check with `ollama list` to verify it responds.
3. **Pull Your First Model.** Run `ollama pull llama3.3:8b` to download the 8B parameter Llama 3.3 model (about 4.9GB). For coding tasks, try `ollama pull deepseek-coder-v2:16b`.
4. **Start Chatting.** Run `ollama run llama3.3:8b` to open an interactive chat session right in your terminal. Type a question, hit enter, and see the response stream in real-time.

## Pros

- Fastest path from zero to a local LLM on a Mac
- MLX-backed Apple Silicon performance gains through 2026
- Free unlimited local runs; cloud is optional
- OpenAI-compatible API plugs into existing agent tools
- Huge model library and community integrations
- Serious company funding/cloud capacity without killing local-first

## Cons

- Large models still need lots of unified memory—8 GB Macs are hobby-tier
- Cloud limits and Max pause create confusion for heavy remote users
- GUI is simpler than LM Studio for some browse-and-test workflows
- Quality varies wildly by model/quant—defaults are not magic
- Team features (SSO/MDM) still maturing relative to big AI SaaS admins

## Deep Dive: How Ollama Turned Local LLMs Into a One-Liner

A look at Ollama's architecture, its role in the local AI ecosystem, and why it became the default tool for running open-source models on Apple Silicon.

## FAQ

### How much RAM do I need to run Ollama?

It depends on the model size. As a rough guide: 7B-8B models need about 8GB of available RAM, 13B models need about 16GB, 30B-34B models need about 32GB, and 70B models need 48-64GB. On Apple Silicon, this is unified memory shared with the GPU. A baseline M1 MacBook Air with 16GB can run 8B models comfortably. For the best experience with larger models, get a Mac with 36GB or more.

### Can I use Ollama with Cursor or VS Code?

Yes. Cursor supports custom OpenAI-compatible endpoints—point it at http://localhost:11434/v1 and select your model. For VS Code, use the Continue.dev extension which has native Ollama support. Both give you AI code completions and chat powered by your local models, with zero data leaving your machine.

### How does Ollama compare to running models with llama.cpp directly?

Ollama is built on top of llama.cpp internally but wraps it with model management, an API server, and automatic hardware optimization. Using llama.cpp directly gives you more control over quantization parameters and sampling settings, but you have to manage everything yourself. Ollama is llama.cpp made practical for daily use.

### Is Ollama just for chat, or can I use it for embeddings?

Both. Ollama supports embedding models like nomic-embed-text and mxbai-embed-large. Use the /api/embeddings endpoint to generate vector embeddings for RAG pipelines, semantic search, or document similarity. It's a complete local AI toolkit, not just a chatbot.

### Does Ollama support vision models?

Yes. Models like LLaVA and Llama 3.2 Vision can process images alongside text. Pass an image path in the API request or use the CLI with `ollama run llava` and paste an image path. The model describes what it sees. Useful for automated image tagging, screenshot analysis, and accessibility tools.

### Can I run Ollama on an Intel Mac?

Yes, but performance will be significantly worse since Intel Macs lack the Metal GPU acceleration that Apple Silicon provides. Models run on CPU only, which means a 7B model might generate around 2-5 tokens per second instead of 30-60 on an M-series chip. It works, but it's not a great experience for anything beyond small models.

### How do I update models when new versions come out?

Run `ollama pull <model>` again and it downloads only the changed layers, similar to Docker image updates. Ollama tracks model manifests and uses delta downloads to minimize bandwidth. You can also set up a cron job to periodically pull updates for your most-used models.

## Sources

- [Ollama Pricing](https://ollama.com/pricing)
- [Ollama powered by MLX on Apple Silicon](https://ollama.com/blog/mlx)
- [Ollama highest MLX performance on Apple Silicon](https://ollama.com/blog/mlx-performance)
- [Faster Gemma 4 MLX MTP](https://ollama.com/blog/faster-gemma-4-mlx-mtp)
- [Ollama Cloud docs](https://docs.ollama.com/cloud)
- [Ollama Cloud Pricing analysis 2026](https://checkthat.ai/brands/ollama/pricing)
- [All aboard open models - Ollama blog](https://ollama.com/blog/all-aboard-open-models)

## Related

- [Cursor](https://bundl.run/apps/cursor)
- [Claude Code](https://bundl.run/apps/claude-code)
- [ChatGPT](https://bundl.run/apps/chatgpt)
- [Claude](https://bundl.run/apps/claude)
- [Codex](https://bundl.run/apps/codex)
- [Windsurf](https://bundl.run/apps/windsurf)
- [Ollama vs LM Studio](https://bundl.run/compare/ollama-vs-lm-studio)
- [Free alternative to ChatGPT Plus](https://claude.com)

```json
{
  "@context": "https://schema.org",
  "@graph": [
    {
      "@type": "Organization",
      "@id": "https://bundl.run/#organization",
      "name": "Bundl.run",
      "url": "https://bundl.run",
      "logo": {
        "@type": "ImageObject",
        "url": "https://bundl.run/og-image.png",
        "width": 1200,
        "height": 630
      },
      "description": "The Ninite for Mac. Install all your essential Mac apps with one terminal command.",
      "sameAs": [
        "https://github.com/abhiofficial/bundl-mac-setup",
        "https://x.com/bundlrun",
        "https://www.producthunt.com/products/bundl-run"
      ],
      "foundingDate": "2024",
      "contactPoint": {
        "@type": "ContactPoint",
        "contactType": "customer support",
        "url": "https://bundl.run/faq"
      }
    },
    {
      "@type": "WebSite",
      "@id": "https://bundl.run/#website",
      "name": "Bundl.run",
      "url": "https://bundl.run",
      "description": "The Ninite for Mac. Install all your essential Mac apps with one terminal command.",
      "publisher": {
        "@id": "https://bundl.run/#organization"
      },
      "inLanguage": "en-US"
    },
    {
      "@type": "Person",
      "@id": "https://bundl.run/authors/alex-chen#person",
      "name": "Alex Chen",
      "jobTitle": "Senior Developer Tools Specialist",
      "url": "https://bundl.run/authors/alex-chen",
      "worksFor": {
        "@id": "https://bundl.run/#organization"
      },
      "description": "Alex Chen has been evaluating developer tools and productivity software for over 12 years, with deep expertise in code editors, terminal emulators, and development environments. As a former software engineer at several Bay Area startups, Alex brings hands-on experience with the real-world workflows these tools are meant to enhance. Alex tests each application extensively on both Intel and Apple Silicon Macs, documenting performance metrics, integration capabilities, and workflow efficiency. When not reviewing software, Alex contributes to open-source projects and writes technical tutorials for the developer community.",
      "knowsAbout": [
        "Code Editors & IDEs",
        "Terminal Emulators",
        "Version Control Tools",
        "DevOps & CI/CD",
        "API Development",
        "Performance Benchmarking"
      ],
      "image": "https://bundl.run/authors/alex-chen.svg"
    },
    {
      "@type": "BreadcrumbList",
      "@id": "https://bundl.run/apps/ollama#breadcrumb",
      "itemListElement": [
        {
          "@type": "ListItem",
          "position": 1,
          "name": "Home",
          "item": "https://bundl.run"
        },
        {
          "@type": "ListItem",
          "position": 2,
          "name": "Apps",
          "item": "https://bundl.run/apps"
        },
        {
          "@type": "ListItem",
          "position": 3,
          "name": "Ollama",
          "item": "https://bundl.run/apps/ollama"
        }
      ]
    },
    {
      "@type": "WebPage",
      "@id": "https://bundl.run/apps/ollama",
      "url": "https://bundl.run/apps/ollama",
      "name": "Ollama for Mac",
      "description": "Run large language models locally on your Mac",
      "isPartOf": {
        "@id": "https://bundl.run/#website"
      },
      "publisher": {
        "@id": "https://bundl.run/#organization"
      },
      "inLanguage": "en-US",
      "datePublished": "2026-01-01T00:00:00Z",
      "dateModified": "2026-08-10T13:34:04.000Z",
      "author": {
        "@id": "https://bundl.run/authors/alex-chen#person"
      },
      "mainEntity": {
        "@id": "https://bundl.run/apps/ollama#software"
      },
      "breadcrumb": {
        "@id": "https://bundl.run/apps/ollama#breadcrumb"
      }
    },
    {
      "@type": "SoftwareApplication",
      "@id": "https://bundl.run/apps/ollama#software",
      "name": "Ollama",
      "description": "Run large language models locally on your Mac",
      "applicationCategory": "DeveloperApplication",
      "operatingSystem": "macOS",
      "url": "https://bundl.run/apps/ollama",
      "image": {
        "@type": "ImageObject",
        "url": "https://img.logo.dev/ollama.com",
        "caption": "Ollama app icon for Mac"
      },
      "sameAs": [
        "https://ollama.com/",
        "https://formulae.brew.sh/cask/ollama-app"
      ],
      "offers": {
        "@type": "Offer",
        "price": "0",
        "priceCurrency": "USD",
        "availability": "https://schema.org/InStock"
      },
      "review": {
        "@type": "Review",
        "author": {
          "@id": "https://bundl.run/authors/alex-chen#person"
        },
        "reviewBody": "Ollama is still the default local LLM runtime on Mac in August 2026—now with a credible cloud tier and MLX-era Apple Silicon speed. Stay on free local for private coding assistants; buy Pro when you need hosted open models and concurrency. LM Studio is the friendly GUI alternative; raw MLX is the enthusiast path. For most developers, install Ollama first."
      },
      "about": {
        "@type": "Thing",
        "name": "Ollama",
        "description": "Run large language models locally on your Mac"
      },
      "isPartOf": {
        "@id": "https://bundl.run/#website"
      }
    },
    {
      "@type": "Article",
      "@id": "https://bundl.run/apps/ollama#article",
      "headline": "Ollama for Mac — Full Review & Installation Guide 2026",
      "description": "Ollama is still the default local LLM runtime on Mac in August 2026—now with a credible cloud tier and MLX-era Apple Silicon speed. Stay on free local for private coding assistants; buy Pro when you need hosted open models and concurrency. LM Studio is the friendly GUI alternative; raw MLX is the enthusiast path. For most developers, install Ollama first.",
      "image": "https://img.logo.dev/ollama.com",
      "author": {
        "@id": "https://bundl.run/authors/alex-chen#person"
      },
      "publisher": {
        "@id": "https://bundl.run/#organization"
      },
      "datePublished": "2026-01-01T00:00:00Z",
      "dateModified": "2026-08-10T13:34:04.000Z",
      "mainEntityOfPage": {
        "@type": "WebPage",
        "@id": "https://bundl.run/apps/ollama"
      },
      "articleSection": "DeveloperApplication",
      "speakable": {
        "@type": "SpeakableSpecification",
        "cssSelector": [
          "h1",
          ".key-facts"
        ]
      },
      "about": {
        "@type": "SoftwareApplication",
        "name": "Ollama",
        "url": "https://bundl.run/apps/ollama"
      },
      "mentions": [
        {
          "@type": "SoftwareApplication",
          "name": "ChatGPT Plus",
          "url": "https://claude.com/"
        }
      ]
    },
    {
      "@type": "FAQPage",
      "@id": "https://bundl.run/apps/ollama#faq",
      "mainEntity": [
        {
          "@type": "Question",
          "name": "How much RAM do I need to run Ollama?",
          "acceptedAnswer": {
            "@type": "Answer",
            "text": "It depends on the model size. As a rough guide: 7B-8B models need about 8GB of available RAM, 13B models need about 16GB, 30B-34B models need about 32GB, and 70B models need 48-64GB. On Apple Silicon, this is unified memory shared with the GPU. A baseline M1 MacBook Air with 16GB can run 8B models comfortably. For the best experience with larger models, get a Mac with 36GB or more."
          }
        },
        {
          "@type": "Question",
          "name": "Can I use Ollama with Cursor or VS Code?",
          "acceptedAnswer": {
            "@type": "Answer",
            "text": "Yes. Cursor supports custom OpenAI-compatible endpoints—point it at http://localhost:11434/v1 and select your model. For VS Code, use the Continue.dev extension which has native Ollama support. Both give you AI code completions and chat powered by your local models, with zero data leaving your machine."
          }
        },
        {
          "@type": "Question",
          "name": "How does Ollama compare to running models with llama.cpp directly?",
          "acceptedAnswer": {
            "@type": "Answer",
            "text": "Ollama is built on top of llama.cpp internally but wraps it with model management, an API server, and automatic hardware optimization. Using llama.cpp directly gives you more control over quantization parameters and sampling settings, but you have to manage everything yourself. Ollama is llama.cpp made practical for daily use."
          }
        },
        {
          "@type": "Question",
          "name": "Is Ollama just for chat, or can I use it for embeddings?",
          "acceptedAnswer": {
            "@type": "Answer",
            "text": "Both. Ollama supports embedding models like nomic-embed-text and mxbai-embed-large. Use the /api/embeddings endpoint to generate vector embeddings for RAG pipelines, semantic search, or document similarity. It's a complete local AI toolkit, not just a chatbot."
          }
        },
        {
          "@type": "Question",
          "name": "Does Ollama support vision models?",
          "acceptedAnswer": {
            "@type": "Answer",
            "text": "Yes. Models like LLaVA and Llama 3.2 Vision can process images alongside text. Pass an image path in the API request or use the CLI with `ollama run llava` and paste an image path. The model describes what it sees. Useful for automated image tagging, screenshot analysis, and accessibility tools."
          }
        },
        {
          "@type": "Question",
          "name": "Can I run Ollama on an Intel Mac?",
          "acceptedAnswer": {
            "@type": "Answer",
            "text": "Yes, but performance will be significantly worse since Intel Macs lack the Metal GPU acceleration that Apple Silicon provides. Models run on CPU only, which means a 7B model might generate around 2-5 tokens per second instead of 30-60 on an M-series chip. It works, but it's not a great experience for anything beyond small models."
          }
        },
        {
          "@type": "Question",
          "name": "How do I update models when new versions come out?",
          "acceptedAnswer": {
            "@type": "Answer",
            "text": "Run `ollama pull <model>` again and it downloads only the changed layers, similar to Docker image updates. Ollama tracks model manifests and uses delta downloads to minimize bandwidth. You can also set up a cron job to periodically pull updates for your most-used models."
          }
        }
      ],
      "isPartOf": {
        "@id": "https://bundl.run/apps/ollama"
      }
    }
  ]
}
```