Run Microsoft Phi-4 Locally With Ollama
Phi-4 14B runs on a 12 GB GPU via Ollama at 8.2 GB (MIT, Q4_K_M). Pull phi4 or the 2.5 GB phi4-mini, connect to Typilot, and keep every Polish command and Notes summary fully private.
Microsoft Phi-4 is a 14-billion-parameter open-weight model that runs on a 12 GB GPU via Ollama at 8.2 GB (Q4_K_M, MIT licence). Pull it with ollama pull phi4, connect it to Typilot, and every Polish command and Notes summary runs entirely on your hardware - audio and text never leave the device.
Here is the hardware breakdown, how to run it, and where Phi-4 sits for writing tasks.
What Phi-4 is
Phi-4 is Microsoft's highest-quality open model for the 14B parameter tier, released in December 2024 under the MIT licence. It was designed for reasoning, analysis, and structured text tasks: MMLU benchmark score is 84.8%, which beats models five times its size on several evaluations (localaimaster.com, 2026). Ollama ships it as the phi4 tag; phi4-mini is the lighter 3.8B companion model at 2.5 GB.
The core tradeoff is context. Phi-4 14B has a 16K token context window - roughly 12,000 words - which is sufficient for email drafts, meeting notes, and document summaries. Phi-4-mini is the right pick if you need CPU-only inference or have hardware with less than 12 GB of memory.
Hardware requirements
| Model | Size (Q4_K_M) | GPU / RAM needed | Context |
|---|---|---|---|
| Phi-4-mini (3.8B) | 2.5 GB | Any 8 GB+ system RAM or GPU | 128K |
| Phi-4 14B | 8.2 GB | RTX 3060 12 GB or better | 16K |
| Qwen3 32B (comparison) | ~19 GB | RTX 4090 24 GB or M-series Mac | 128K |
Phi-4 14B at 8.2 GB fits on an RTX 3060 12 GB with memory to spare. The same applies to any 12 GB+ GPU: RTX 3080, RTX 4060 Ti 16 GB, and the full Apple Silicon Mac line (8 GB through 192 GB unified memory). On Apple Silicon, Metal acceleration applies automatically - no separate build is needed.
CPU-only inference on Phi-4 14B produces around 5-8 tokens per second on a modern laptop, which is usable but slow. For CPU-only use, Phi-4-mini at 2.5 GB delivers 4-12 tok/s and handles most writing tasks without a GPU (source: promptquorum.com, 2026).
Running it in Ollama
Current Ollama stable is v0.35.0 (September 28, 2026). Both phi4 and phi4-mini have been in the Ollama library since April 2026.
# Pull Phi-4 14B (~8.2 GB download)
ollama pull phi4
# Pull the lighter 3.8B model (~2.5 GB)
ollama pull phi4-mini
# Run interactively
ollama run phi4
# Quick structured summary test
ollama run phi4 "Summarise this in three bullet points: [paste text]"
After pulling, connect to Typilot: open Preferences, go to Models, select Phi-4 from the list of downloaded Ollama models, and set it as the default for Polish or Notes commands.
Writing quality: where Phi-4 fits
Phi-4's architecture is optimised for structured reasoning rather than open-ended prose generation. It handles email drafts, meeting summaries, formatted notes, and code comment dictation reliably. Long-form creative prose is its weaker point - for documents longer than a few thousand words, Qwen3 32B or Muse Glimmer 30B produce more coherent output, though both require substantially more memory.
| Task | Phi-4 14B | Qwen3 32B local | Claude Sonnet / Opus (cloud) |
|---|---|---|---|
| Short email draft | Very good | Very good | Excellent |
| Meeting note cleanup | Very good | Very good | Excellent |
| Structured document summary | Very good | Very good | Excellent |
| Code comment dictation | Very good | Good | Excellent |
| Long creative prose (2,000+ words) | Good | Very good | Excellent |
| Audio and text leave device | Never | Never | Always |
| Works offline | Yes | Yes | No |
| Per-request cost | $0 | $0 | Pay-per-token |
Phi-4 14B scores 84.8% MMLU at the 14B tier - a result that beats much larger models (localaimaster.com, 2026). For Typilot Polish and Notes commands on typical text - email, meeting notes, code comments - Phi-4 handles the task reliably. If you need flowing long-form prose and have a 24 GB GPU or M4 Pro Mac, Qwen3.8-27B or Nemotron 3.5 Lightning raise the quality ceiling. Phi-4 is the practical choice for a 12 GB card.
Why local inference matters here
A cloud model sends every prompt to a vendor server on every request - that is true regardless of the vendor's privacy policy. Phi-4 running via Ollama processes everything at localhost:11434. The weights, your prompts, and every Polish command result stay on your machine.
Verify it: pull the model, disconnect your internet, and run a Polish command in Typilot. The model keeps working.
Using it with Typilot
Typilot routes all AI commands through Ollama at localhost:11434. Once you have pulled the model, it appears in Preferences > Models alongside any other downloaded Ollama models. You can:
- Set it as the default model for Polish and Notes commands in Preferences > General > Default model.
- Use Phi-4-mini as the fallback when working on battery or on hardware without a dedicated GPU.
- Switch between models per command type - for example, Phi-4 for meeting notes, Phi-4-mini for quick email fixes.
For a step-by-step Ollama setup, see how to run a local AI assistant with Ollama and Typilot and the Ollama setup docs.
For the full RAM and VRAM breakdown across model sizes, see how much RAM you need to run a local LLM.
If you want a stronger local prose writer with a 24 GB GPU, Qwen3.8-27B at ~16.8 GB fits M4 Pro and RTX 4090. For the most capable open agent model on similar hardware, Nemotron 3.5 Lightning at 17.8 GB on Apple Silicon brings 30B-class capability to the same tier of Mac hardware.
The short version
Phi-4 14B is Microsoft's best open model for a 12 GB GPU: MIT licence, 8.2 GB Q4_K_M on Ollama, ollama pull phi4. It scores 84.8% MMLU and handles structured writing tasks - email, notes, document summaries - reliably at 25-32 tok/s on an RTX 3060. The 16K context window covers most day-to-day writing. Phi-4-mini (2.5 GB) covers CPU-only and low-RAM setups.
Both models run entirely on your device at localhost:11434 with nothing sent to a server. That is what private AI assistance looks like. Try Typilot with Phi-4 free for three days at /download.
Common questions.
Can I run Phi-4 14B on an RTX 3060 12 GB?+
Yes. Phi-4 14B at Q4_K_M quantisation weighs 8.2 GB on Ollama, well within the RTX 3060's 12 GB of VRAM. You can expect around 25-32 tokens per second on that GPU. The model also runs on any 12 GB+ card, all Apple Silicon Macs, and on CPU alone (5-8 tok/s) - though Phi-4-mini at 2.5 GB is a better choice for CPU-only inference.
What is the context window of Phi-4?+
Phi-4 14B has a 16,000-token context window, which covers roughly 12,000 words - enough for email drafts, meeting notes, and document summaries. Phi-4-mini has a larger context window. If you need very long documents in context, Qwen3 8B or Qwen3 32B (both with 128K context) are better options.
Is Phi-4 good for writing assistance and note-taking?+
Yes, with an honest caveat. Phi-4 14B scores 84.8% MMLU and handles structured writing tasks - email drafts, meeting summaries, code comments, formatted notes - reliably. It is not optimised for long-form creative prose; for documents over a few thousand words, Qwen3 32B or Muse Glimmer 30B produce more coherent output. For everyday Typilot Polish and Notes commands, Phi-4 is a strong choice.
Is Phi-4 private to run locally?+
Yes, structurally. When you run Phi-4 via Ollama, inference happens entirely at localhost:11434. No prompt, response, or metadata is sent to Microsoft or any server after the initial model download. The model weights are MIT-licensed with no telemetry - privacy is architectural, not policy-dependent.