Local models · PC and Mac本機模型 · PC 與 Mac
The best local model is the largest one that fits and still answers.最好的本機模型,是放得下、而且還答得出來的最大那個。
Coding, office agents, research, and creative work do not want the same weights. A browser cannot see your GPU, so the presets below cover the machines people actually own. The form is for the machine you have.程式、辦公室智能體、研究與創作不想要同一套權重。瀏覽器看不到你的顯示卡,所以下面的預設覆蓋人們真正擁有的機器。表格則是給你手上的那一台。
If you do not want to type specs如果不想填規格
01
Each card is a full recommendation from the same rules as the form. "Use this config" copies it into the advisor so you can change one number.每張卡片用與表格相同的規則給出完整建議。「用這組設定」會把它抄進顧問,方便你改一個數字。
Deployment methods部署方式
03
| Method方式 |
PC | Mac |
Use it for適合 |
Skip it when略過,當 |
| Ollama | Yes是 | Yes是 |
The fastest start. One local API on port 11434 that agents already understand.最快的起點。連接埠 11434 上的一個本機 API,智能體已經看得懂。 |
You need many simultaneous users. Use vLLM on a Linux GPU server instead.你需要很多同時使用者。改在 Linux GPU 伺服器上用 vLLM。 |
| LM Studio | Yes是 | Yes是 |
A window, a model browser, and a local server, with no terminal required.一個視窗、模型瀏覽器與本機伺服器,不必開終端機。 |
The machine is a headless server.機器是沒有畫面的伺服器。 |
| llama.cpp | Yes是 | Yes是 |
Maximum control of quants, GPU layers, and CPU offload.對量化、GPU 層與 CPU 卸載有最大控制。 |
You wanted a button. Start in LM Studio, which uses the same family of code.你想要一個按鈕。從 LM Studio 開始,它用的是同一族程式。 |
| MLX | No否 | Yes是 |
The highest tokens per second on Apple Silicon, once you are past the first afternoon.過了第一個下午之後,Apple Silicon 上最高的每秒 token。 |
You are on Windows or NVIDIA. MLX will not help.你在 Windows 或 NVIDIA 上。MLX 幫不上。 |
| vLLM or SGLang | Linux + NVIDIA | No否 |
Serving an office or a lab from one GPU box.用一台 GPU 主機服務一間辦公室或實驗室。 |
You have a laptop. The overhead is not worth it.你只有筆電。額外開銷不值得。 |
| Open WebUI | Yes是 | Yes是 |
A shared chat window and a document folder on top of Ollama.疊在 Ollama 上的共用聊天視窗與文件資料夾。 |
You think it is the model. It is the interface.你以為它是模型。它是介面。 |
| ComfyUI | Yes是 | Yes是 |
Local images and draft video. FLUX for stills, LTX or Wan for short clips.本機圖像與影片草稿。靜態用 FLUX,短片用 LTX 或 Wan。 |
You only need chat. It is a graph editor, not a writing app.你只需要聊天。它是節點編輯器,不是寫作應用。 |
| Draw Things | No否 | Yes是 |
Local images on a Mac without building a node graph.在 Mac 上做本機圖像,不必建節點圖。 |
You are on Windows. Use ComfyUI there.你在 Windows 上。那裡用 ComfyUI。 |
Patterns that hold on both platforms兩個平台都成立的做法
04
Coding程式
Q4 is the default quant. Leave a gigabyte or two free so the context cache has a home. A 24 GB NVIDIA card or a 48 GB Mac is where a local coding agent starts to feel like a tool rather than a demo. Point Continue, Cline, or Aider at http://localhost:11434/v1.預設量化是 Q4。留出一兩個 GB,讓上下文快取有地方住。24 GB 的 NVIDIA 卡或 48 GB 的 Mac,本機程式智能體才開始像工具而不是展示。把 Continue、Cline 或 Aider 指向 http://localhost:11434/v1。
Office agents辦公室智能體
Prefer a Qwen-family model when you need tool calls. Run Hermes if the procedure should become a skill, OpenClaw if the agent must sit in chat apps. Give it one folder. Bind it to localhost. The model being local does not make an unrestricted shell safe.需要工具呼叫時,優先 Qwen 家族。流程應該變成技能就跑 Hermes;智能體必須坐在聊天應用裡就用 OpenClaw。只給一個資料夾。綁在 localhost。模型在本機,並不會讓不受限的殼層變安全。
Research研究
Do not paste a book into the prompt. Run a small embedding model (bge or nomic class, a couple of gigabytes) and retrieve passages into Open WebUI or AnythingLLM. Spend the big model on the answer. Cap local context around 8K–32K even if the card says 1M. The cache is what runs you out of memory.不要把一本書貼進提示。跑一個小的嵌入模型(bge 或 nomic 等級,幾個 GB),把段落取進 Open WebUI 或 AnythingLLM。大模型花在回答上。就算卡片寫 1M,本機上下文也限制在大約 8K–32K。把你記憶體用完的是快取。
Images and video圖像與影片
Stills become reasonable at 16 GB of usable accelerator memory and comfortable at 24 GB, with FLUX-class open weights in ComfyUI. Video stays a cloud job (Kling, Seedance, Veo, Hailuo) unless you are making drafts on a 48 GB-class machine. The media page names the cloud models.可用加速器記憶體 16 GB 時靜態圖開始合理,24 GB 才舒服,在 ComfyUI 用 FLUX 等級的開放權重。影片仍是雲端工作(Kling、Seedance、Veo、Hailuo),除非你在 48 GB 等級的機器上做草稿。雲端模型的名字在影像頁。
Names in the advisor are families, not eternal Ollama tags. Search the library for the current tag before you script a deploy. A point release in a family you already trust beats a new brand you saw in a chart.顧問裡的名字是家族,不是永恆的 Ollama 標籤。寫部署腳本之前,先在資料庫搜目前的標籤。你已經信任的家族出一個小版本,勝過圖表上看到的新品牌。