History沿革

How the field got to this desk.這個領域怎麼走到這張參考台。

A working timeline, not a museum. The point is to see which ideas are old, which products are new, and which 2026 claims are a price change wearing a new name.這是工作用的時間線,不是博物館。重點是看出哪些想法很舊、哪些產品很新,以及哪些 2026 年的說法只是換了名字的價格變動。

Timeline時間線

01
  1. 1950 · Foundations1950 · 奠基

    Turing's question圖靈的問題

    Alan Turing's "Computing Machinery and Intelligence" asks how you would tell a machine's answers from a person's. The imitation game is a test of behavior, not a definition of understanding. That gap is still the argument.艾倫·圖靈的〈計算機器與智能〉問:你要怎麼分辨機器的回答與人的回答。模仿遊戲測的是行為,不是理解的定義。這個缺口仍是爭論本身。

  2. 1956 · Foundations1956 · 奠基

    The field gets a name領域得到名字

    The Dartmouth workshop proposes "artificial intelligence" as a research program: machines that use language, form abstractions, and solve problems reserved for humans. The ambition arrives before the methods that could carry it.達特茅斯研討會把「人工智慧」提出為研究計畫:會使用語言、形成抽象、解決原本留給人的問題的機器。抱負先到,扛得起它的方法還沒到。

  3. 1958–1986 · Learning1958–1986 · 學習

    Perceptrons, a winter, then backpropagation感知器、一次寒冬,然後是反向傳播

    Rosenblatt's perceptron shows a trainable network. Limits of simple perceptrons, and thin results, feed the first funding winter. In 1986 Rumelhart, Hinton, and Williams make backpropagation the practical training story. Expert systems boom and then disappoint. A second winter follows.羅森布拉特的感知器展示可訓練的網路。簡單感知器的限制與單薄的結果,餵出第一次經費寒冬。1986 年 Rumelhart、Hinton 與 Williams 讓反向傳播成為實際的訓練故事。專家系統先熱後冷。第二次寒冬跟著來。

  4. 1997 · Foundations1997 · 奠基

    Deep Blue beats Kasparov深藍擊敗卡斯帕羅夫

    A narrow system, search plus chess knowledge, beats the world champion. It is a real result and a bad metaphor. Winning a closed game did not produce a general assistant.一個窄系統,搜尋加上西洋棋知識,擊敗世界冠軍。這是真結果,也是壞比喻。贏下一盤封閉的棋,並沒有產生通用助理。

  5. 2012 · Learning2012 · 學習

    AlexNet

    A deep convolutional net trained on GPUs wins ImageNet by a margin that ends the argument about deep learning for vision. Hardware, data, and depth become the recipe.在 GPU 上訓練的深度卷積網路,以結束爭論的差距贏下 ImageNet 的視覺深度學習之辯。硬體、資料與深度成為配方。

  6. 2014 · Generative2014 · 生成

    GANs

    Goodfellow and co-authors describe generative adversarial networks. Machines start making images people will look at, not only labels.Goodfellow 與共同作者描述生成對抗網路。機器開始做出人願意看的圖像,不只是標籤。

  7. 2016 · Learning2016 · 學習

    AlphaGo

    DeepMind's system beats Lee Sedol at Go. Reinforcement learning plus search, on a game long treated as too wide for brute force. The public story shifts from "software" to "a lab".DeepMind 的系統在圍棋擊敗李世乭。強化學習加上搜尋,用在長期被認為寬到不能硬算的遊戲。公眾故事從「軟體」轉成「一間實驗室」。

  8. 2017 · Generative2017 · 生成

    The TransformerTransformer

    "Attention Is All You Need" replaces recurrence with attention. Almost every model on the comparison page is a descendant of this paper. The architecture is old enough to vote. The products are not.〈Attention Is All You Need〉用注意力取代遞迴。比較頁上幾乎每個模型都是這篇論文的後代。架構老到可以投票。產品還不老。

  9. 2018–2020 · Generative2018–2020 · 生成

    BERT, GPT, then GPT-3

    Pretraining on a huge text pile, then a small amount of instruction, becomes the pattern. GPT-3 shows that scale changes what a text model can do in one prompt. The API exists before the mass-market product.先在巨大文本上預訓練,再加少量指令,成為模式。GPT-3 顯示規模改變文字模型在一次提示裡能做的事。API 先於大眾產品存在。

  10. 2022 · Generative2022 · 生成

    ChatGPT and public image modelsChatGPT 與公開的圖像模型

    ChatGPT (November 2022) is the moment non-specialists change their work. Stable Diffusion and its cousins make image generation something you can run, not only rent. The industry's unit of progress becomes the demo people forward.ChatGPT(2022 年 11 月)是非專業者改變工作的時刻。Stable Diffusion 與其同類讓圖像生成變成你可以跑的東西,不只是租。產業的進度單位變成人們會轉傳的展示。

  11. 2023 · Generative2023 · 生成

    GPT-4, Claude, Llama

    Closed frontier models get multimodal and much more reliable. Meta's Llama line makes serious weights downloadable. "Which API?" and "can I run it here?" become separate questions. They still are.封閉前沿模型變成多模態,也可靠得多。Meta 的 Llama 線讓認真的權重可以下載。「用哪個 API?」與「我能在這裡跑嗎?」變成兩個問題。現在仍是。

  12. 2024 · Agents2024 · 智能體

    Long context, tools, and the first coding agents people keep長上下文、工具,以及人們會留著的第一批程式智能體

    Context windows stretch. Models call tools. IDE agents stop being party tricks for a subset of programmers. The failure mode shifts from bad prose to a bad action taken with confidence.上下文視窗拉長。模型呼叫工具。IDE 智能體不再只是一部分程式人的派對把戲。失敗模式從壞文章,轉成自信地做了一個壞動作。

  13. 2025 · Agents2025 · 智能體

    Reasoning models and open weights that close the gap推理模型,以及縮小差距的開放權重

    Labs train models to spend tokens thinking. Open-weight labs, DeepSeek among them, force the closed labs to compete on price as well as quality. Local models become good enough for real drafts on a single GPU.實驗室訓練模型把 token 花在思考上。開放權重實驗室,包括 DeepSeek,迫使封閉實驗室在價格與品質上同時競爭。本機模型在單張 GPU 上已足以寫真正的草稿。

  14. 2026 · Agents2026 · 智能體

    The agent is the product智能體就是產品

    By October the frontier has names like Claude Opus 5.5, GPT-6 Astra and Sol, and Gemini 4 Argon. The thing users install is often not the chat window. It is Claude Code, Codex, Cursor, Hermes, OpenClaw, or OpenAI's new Dots. Quality still differs. The invoice differs more.到了十月,前沿的名字是 Claude Opus 5.5、GPT-6 Astra 與 Sol、Gemini 4 Argon。使用者安裝的常常不是聊天視窗,而是 Claude Code、Codex、Cursor、Hermes、OpenClaw,或 OpenAI 新的 Dots。品質仍有差。帳單差更多。

What is actually new in October 20262026 年 10 月真正新的是什麼

02

The index leader is not the value leader指數領先不是價效領先

Opus 5.5 at 58 leads the checked Artificial Analysis index. GPT-6.1 Sol at 52 costs about a fifth of Astra and does most office and coding work. Pick with a task, not a podium.Opus 5.5 的 58 分領先已核對的 Artificial Analysis 指數。GPT-6.1 Sol 的 52 分大約是 Astra 價格的五分之一,並做掉多數辦公室與程式工作。用任務選,不要用頒獎台選。

China is in the same conversation中國在同一場對話裡

Qwen, DeepSeek, Kimi, GLM, and Xiaomi's MiMo are not a side chart. They lead on price, on open weights, or on both. Compare them on your prompts and your data rules.Qwen、DeepSeek、Kimi、GLM 與小米的 MiMo 不是附圖。它們在價格、開放權重,或兩邊領先。用你的提示與資料規則比較。

Local is a hardware question本機是硬體問題

A 16 GB laptop and a 128 GB Mac Studio are different products. The advisor exists so you stop reading a benchmark written for an 8-GPU server.16 GB 筆電與 128 GB Mac Studio 是不同產品。顧問存在,是為了讓你不要再讀為 8 張 GPU 伺服器寫的基準。

Control is the scarce feature稀缺的功能是控制

Models will browse, click, and send. The week's policy news — an FTC probe, a voluntary White House accord, California's worker rules — is about that, not about spelling.模型會瀏覽、點擊、送出。這一週的政策新聞——FTC 調查、白宮自願協議、加州勞工規則——談的是這個,不是拼字。

Next: the model table, or what AGI does and does not mean here.下一頁:模型表,或AGI 在這裡是什麼、不是什麼。