🧠
第二大腦
🌏國際視野2026-06-08

Paving the way for agents in biology(為 AI Agent 鋪平生物學之路)

#部落格#英文#Anthropic#AI Agent#生物資訊#基礎建設

Paving the way for agents in biology(為 AI Agent 鋪平生物學之路)

來源: Anthropic 原文日期: 2026-06-08

中文摘要

Anthropic 研究員 Laura Luebbert 指出,生物學 AI Agent 發展的最大瓶頸不是模型推理能力,而是生物數據基礎設施對 agent 極不友好。研究團隊讓 Claude、Biomni OSS、Edison Analysis 和 GPT 等 agent 從 NCBI Virus 病毒数据库中擷取序列數據,結果最強模型的準確率也不一致(16.9%–91.3%),且同樣問題三次跑出的結果差異巨大。團隊開發了「gget virus」確定性檢索層後,準確率一舉提升至近 100%。

關鍵洞察

  • 生物學不等於軟體:軟體有結構化 API、版本控制、package manager 等「高速公路」基础设施;生物數據則像義大利小鎮的窄路,充滿異質格式、分散資料庫和一次性腳本
  • Karpathy 比喻的延伸:Andrej Karpathy 抱怨「程式碼最簡單,都在瀏覽器裡點擊」——生物學家用手動\>20 年的痛點,agent 也無法倖免
  • Ebola 疫情案例:2026 年 5 月剛果爆發 Bundibugyo 病毒 Ebola 疫情,研究人員需手動點擊 NCBI Virus 介面才能取得歷史基因組數據進行比對,嚴重阻礙防疫反應
  • VirBench 基準測試:120 個真實病毒序列查詢、40 種病原體,Claude Sonnet 4 在相同 prompt 三次跑出 106、15、5 條序列(期望值 266 條)
  • 確定性檢索層是解方:gget virus 協調 REST、Datasets、E-utilities 等多個 API,能正確重現網頁介面的篩選邏輯
  • 原文關鍵句(英文保留)

    > "Even the strongest models did not consistently achieve the level of accuracy required for reliable dataset construction. But accuracy rose to nearly 100% once she and her team added gget virus, a deterministic retrieval layer."

    > "Software infrastructure was basically made for the needs of cars (agents): paved roads, clear lanes, standardized signals. Using AI agents to navigate biological data infrastructure is like driving through an old city that was designed before cars."

    > "The bottleneck for biological agents is not only reasoning but the absence of widespread deterministic execution layers for querying biological data."

    對 CJ 哥的價值

  • ESG/公共健康:AI agent 在傳染病监测中的應用案例,說明為何「數據基礎設施」是 AI 落地於永續目標的關鍵前提
  • 企業 AI 轉型啟示:企業要讓 AI agent 在必領域(如研發、法务、合規)可靠運作,必須先建立「確定性執行層」——不能只靠 LLM 的推理能力
  • 內容生產:以「程式 agent vs. 生物 agent」對比切入各行業 AI 轉型的數據基礎需求,是非常有傳播力的角度
  • 相關文章

    當管理層的大腦還留在十九世紀,企業憑什麼駕馭二十一世紀的AI?

    # 當管理層的大腦還留在十九世紀,企業憑什麼駕馭二十一世紀的AI? > 來源:https://home.gamer.com.tw/artwork.php?sn=6393385|作者:劍心san(sanboy289)|2026-09-05|巴哈姆特創作 > 存入:2026-09-05|分類:管理心理學 ## 一句話 管理思維還停留在工業革命早期的企業——用羞辱究責取代系統分析——不可能駕馭 A...

    Stratechery:Anthropic Fable 5.1 與企業 AI 資料政策轉向

    # Stratechery:Anthropic Fable 5.1 與企業 AI 資料政策轉向 **來源:** Stratechery(Ben Thompson) **日期:** 2026-09-02 ## 摘要 Ben Thompson 分析 Anthropic 發布 Fable 5.1 的產業意義:最值得關注的不是模型能力,而是 Anthropic 大幅修改了引發巨大爭議的資料保留政策(F...

    AI 代理與客戶資料之爭:資料基礎建設的未來(A16Z)

    # AI 代理與客戶資料之爭:資料基礎建設的未來 **來源:** A16Z Podcast(Martin Casado × Fivetran 共同創辦人 George Fraser) **日期:** 2026-06-05 ## 摘要 Martin Casado 與 Fivetran CEO George Fraser 對談 AI 時代資料基礎建設的未來。涵蓋 Fivetran 與 dbt 的合...

    全球首例雙盲 AI 評測:用密碼學環境建立基準信任(Google DeepMind)

    # 全球首例雙盲 AI 評測:用密碼學環境建立基準信任 **來源:** Google DeepMind(William Isaac、Sol Messing、Kristian Lum) **日期:** 2026-08-27 ## 摘要 DeepMind 推出全球首例「專有前沿模型的雙盲評測(double-blind evaluation)」,把外部評測侷限在密碼學「盒子」內,防止模型事先看到題目...