Sources來源
Where a number came from.這些數字從哪裡來。
This edition was closed on 1 October 2026. Index scores in the model table were copied only from Artificial Analysis articles read for this desk. The weekly ranks are a separate copy of the public boards named here. News headlines that sit behind a paywall are linked as a news search, not as a fabricated URL. If a figure is not on this page, do not treat it as checked.這一版在 2026 年 10 月 1 日定稿。模型表的指數只從本台讀過的 Artificial Analysis 文章抄來。每週名次是另一份,抄自這裡列出的公開榜。付費牆後面的新聞標題連到新聞搜尋,不是編造的網址。數字若不在這一頁,就不要當成已核對。
Scores and prices分數與價格
01- LMArena — blind preference votes. The models page copies the latest leaderboard dataset once a week: Text overall, with style control and without it. The three names on the home page are the style-control top three. The score shown is that board's rating, rounded. Dataset note: Arena Leaderboard Dataset.盲測偏好。模型頁每週從最新的 排行榜資料集抄一次:文字總榜,有 style control,也有沒有的。首頁三個名字是風格控制榜的前三名。顯示的分數是該榜評分,四捨五入。資料說明:Arena Leaderboard Dataset。
- LiveBench — checked tasks, contamination-resistant by design. The models page copies the latest release that site publishes and averages categories the same way its leaderboard does. The release date is shown on the row. No second score is invented.有標準答案的題組,設計上避免題目外洩。模型頁抄該站公布的最新一次,並用它自己的方式做分類平均。列上會寫題組日期。不另造第二個分數。
- Artificial Analysis — the intelligence index this desk quotes by hand. The publisher's data feed needs an API key, so the weekly job does not copy it. Open the article and read the effort level before comparing a new index with the table.本站手動引用的智力指數。發布者的資料介面需要 API 金鑰,所以每週作業不抄它。拿新的指數和表上的數字比之前,先打開文章,讀努力等級。
- Claude Opus 5.5 — index 58 at max, $4 / $20 per million, cache reads $0.20, 1M context. 22 September 2026.max 指數 58,每百萬 $4 / $20,快取讀取 $0.20,1M 上下文。2026 年 9 月 22 日。
- Claude Sonnet 5.5 — index 56 at max, two points behind Opus, $2 / $10 sticker, high output volume. 28 September 2026.max 指數 56,落後 Opus 兩分,$2 / $10 牌價,輸出量大。2026 年 9 月 28 日。
- Gemini 4 Argon — index 53 at high, tied with GPT-6 Astra (max, 53), one point ahead of GPT-6.1 Sol (max, 52). Discounted $2 / $10 on a $4 / $20 list. Limited rollout. Gemini 3.8 Flash (high) is stated as 12 points lower, which is the 41 in the table. 30 September 2026.high 指數 53,與 GPT-6 Astra(max,53)持平,領先 GPT-6.1 Sol(max,52)一分。$4 / $20 牌價上的折扣價 $2 / $10。開放有限。Gemini 3.8 Flash(high)被寫成低 12 分,也就是表上的 41。2026 年 9 月 30 日。
- GPT-6.1 Sol — near-Astra index at under a quarter of Astra's cost per task. 29 September 2026.指數接近 Astra,單次任務成本不到 Astra 的四分之一。2026 年 9 月 29 日。
- AA-AgentPerf-Local — laptop and workstation agent speed, 29 September 2026. Active parameter count dominated. RTX 5090 led when the model fit in 32 GB.筆電與工作站的智能體速度,2026 年 9 月 29 日。活躍參數量主導。模型放進 32 GB 時,RTX 5090 領先。
- Artificial Analysis leaderboard — use it the day you buy. It is not one of the weekly copied tables.要花錢的那天,先看這裡。它不在每週抄寫的表裡。
- OpenAI pricing and與 Anthropic pricing — confirm stickers. Luna's $0.10 / $0.50 is September reporting, not a fresh scrape of the price page.確認牌價。Luna 的 $0.10 / $0.50 是九月報導,不是價格頁的新抓取。
Grok 4.7, Muse Spark, Qwen3.8 Max, Kimi K3, GLM-5.3, DeepSeek V4.1 Flash, MiMo, and Mistral Medium 3.5 are in the table because they are the models a buyer will be offered. Their index cells are blank on purpose. Secondary leaderboards disagreed with each other in the same week, so those ranks were not copied.Grok 4.7、Muse Spark、Qwen3.8 Max、Kimi K3、GLM-5.3、DeepSeek V4.1 Flash、MiMo 與 Mistral Medium 3.5 在表上,是因為買方會被推這些模型。它們的指數格故意空白。同一週的次級排行榜彼此不合,所以那些名次沒有抄。
Labs實驗室
02- OpenAI — GPT-6 family, Codex, Dots, GPT Image.GPT-6 家族、Codex、Dots、GPT Image。
- Anthropic — Claude 5 family, Claude Code.Claude 5 家族、Claude Code。
- Google DeepMind — Gemini, Veo, Nano Banana.Gemini、Veo、Nano Banana。
- xAI — Grok, Grok Imagine.Grok、Grok Imagine。
- Meta — Muse and the Llama open-weight line. They are different products.Muse 與 Llama 開放權重線。它們是不同產品。
- Alibaba Qwen, DeepSeek, Kimi, Z.ai — Chinese frontier and open weights.中國前沿與開放權重。
- Mistral — European open weights.歐洲開放權重。
- Black Forest Labs — FLUX.
Agents智能體
03- NousResearch/hermes-agent
- openclaw/openclaw
- Aider and與 Cline
- Architecture comparisons read while writing the agents page included OpenClaw's docs, Hermes' docs, and neutral writeups from the last week of September 2026. Characterizations on this site follow the public design each project claims: a learning loop versus a gateway. Confirm details in the repo before you install. Both projects' security advisories are on their GitHub security tabs.寫智能體頁時讀過的架構比較,包括 OpenClaw 的文件、Hermes 的文件,以及 2026 年 9 月最後一週的中立文章。本站的描述跟隨各專案公開主張的設計:學習迴圈對上閘道。安裝前到儲存庫確認細節。兩個專案的安全通報都在它們的 GitHub 安全分頁。
Local deployment本機部署
04- Ollama, LM Studio, llama.cpp, MLX.
- vLLM, SGLang, Open WebUI.
- ComfyUI for local image and video graphs.用來跑本機圖像與影片。 Draw Things on Mac.在 Mac 上。
- Concrete files linked from the local file table and the media download list: Qwen2.5 and Qwen2.5-Coder on Ollama and Hugging Face, Llama 3.3 on Ollama, SDXL from Stability AI, FLUX.1 from Black Forest Labs, Wan from Wan-AI, LTX-Video from Lightricks. Advisor names such as Qwen3.6 are families. If the publisher has a newer tag in that size, use the newer tag. Do not treat a forum filename as a source.本機檔案表與影像下載清單連出去的具體檔案:Qwen2.5 與 Qwen2.5-Coder 在 Ollama 和 Hugging Face,Llama 3.3 在 Ollama,SDXL 來自 Stability AI,FLUX.1 來自 Black Forest Labs,Wan 來自 Wan-AI,LTX-Video 來自 Lightricks。顧問裡的 Qwen3.6 這類名字是家族。發布者若有同一大小的較新標籤,用較新的。論壇上的檔名不要當成來源。
- Fit estimates in the advisor are Q4-class weight sizes plus a small margin, rounded to families. They are planning numbers, not a benchmark certificate. Context length eats the margin.顧問裡的容量估計,是 Q4 等級的權重大小再加上一點餘裕,然後歸到某個家族。這些是規劃用的數字,不是跑分證書。上下文一長,餘裕就沒了。
- The prompt page shows the shape of a useful answer. Text results were written for this desk. Image results are example stills. Video results describe the clip. None of them is a transcript from a named model, and none of them is a benchmark.提示頁示範的是一份有用的回答長什麼樣子。文字結果是為本站寫的。圖像是示例靜態圖。影片寫的是片段裡該出現什麼。它們都不是某個模型的逐字稿,也不是跑分。
News and policy, this week本週新聞與政策
05Six linked stories from the past seven days of Google News, preferring a major desk when one is in the feed. The server refreshes this list weekly. Headlines stay in the language of the article.過去七天的 Google 新聞,列出六則有連結的報導。供稿裡若有主要媒體,優先採用。伺服器每週更新這份清單。標題保留文章原文。
- Loading this week's headlines…正在讀取本週標題…
- NIST post-quantum cryptography — ML-KEM, ML-DSA, SLH-DSA. Standing source for the quantum section. This one does not rotate.ML-KEM、ML-DSA、SLH-DSA。量子一節的固定來源。這一則不輪替。
The home page headline strip is the ten newest Google News stories from the past day. The time beside each one uses the clock on your device. A saved copy can lag by a few minutes. Under that strip, ten stories from the past seven days refresh weekly, major desks first. If that weekly list does not load, the 1 October briefing is shown instead. The three cards at the top refresh on their own clocks: the model names weekly, the US story daily, the China story daily. A China headline is the original wording, including simplified characters when the article uses them.首頁標題條是過去一天裡最新的十則 Google 新聞。旁邊的時間用你這台裝置的時鐘。存下的副本可能慢幾分鐘。那條底下的十則,是過去七天的報導,每週更新,主要媒體優先。每週清單沒有載入時,改顯示 10 月 1 日的簡報。最上方三張卡各自計時:模型名稱每週、美國報導每天、中國報導每天。中國標題用文章原文,文章若是簡體,就保留簡體。
How to read “this model is the best”怎麼讀「這個模型最強」
06- Name the board. A blind vote, a set of checked tasks, and a lab demo are three different questions. “Best” without the board is an advertisement.先問是哪一張榜。盲測投票、有標準答案的題組、實驗室自己的展示,問的是三件不同的事。不講榜單的「最強」,就是廣告。
- Read the effort. Max and high are different products and different bills. The score in a headline belongs to that setting only.看努力等級。Max 和 high 是不同的產品,也是不同的帳單。標題裡的分數只屬於那個設定。
- Price the finished job. A cheaper sticker can cost more if the model writes more tokens, or thinks longer.算完成一件工作的價錢。牌價較低的模型,如果寫出更多 token,或想得更久,單次工作可能更貴。
- A rank does not know your files, your language, or your computer. Open weights on a chart still have to fit in memory. That check is the local page.名次不知道你的檔案、你的語言、你的電腦。榜上的開放權重,還是得放得進記憶體。這件事看本機頁。
- An Artificial Analysis index on this desk was copied only after the article was opened and the effort level was read. A new index is not comparable until you have done the same.本站表上的 Artificial Analysis 指數,是打開文章、讀過努力等級之後才抄的。新的指數,在你做過同樣的事之前,不能直接比。