← 知识图谱 ← 返回学习站

计算机使用

Introducing Computer Use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
Anthropic · 2024-10-22

Today we're announcing an upgraded Claude 3.5 Sonnet and a new model, Claude 3.5 Haiku. The upgraded Sonnet delivers across-the-board improvements, with particularly significant gains in coding. We're also introducing a groundbreaking new capability in public beta: computer use. Developers can direct Claude to use computers the way people do — by looking at a screen, moving a cursor, clicking buttons, and typing text. Claude 3.5 Sonnet is the first frontier AI model to offer computer use in public beta. At this stage it is still experimental — at times cumbersome and error-prone.

今天 Anthropic 发布了升级版 Claude 3.5 Sonnet 和新模型 Claude 3.5 Haiku。升级后的 Sonnet 在各方面都有提升,尤其在编码上进展显著。同时推出了一项进入公开测试的突破性新能力:计算机使用(computer use)。开发者可以让 Claude 像人类一样使用电脑——看屏幕、移动光标、点击按钮、输入文字。Claude 3.5 Sonnet 是第一个提供"计算机使用"公开测试的前沿 AI 模型。需要强调的是,这个阶段它仍是实验性的——有时笨拙、容易出错。

💡 AI 解读

这是 AI Agent 领域的一个里程碑事件。此前 Agent 只能调用预先定义好的 API/函数;而"计算机使用"让模型直接操作图形界面——这意味着任何"人能用软件完成的任务"理论上都能交给 Claude 自动化。它把 Agent 的可操作空间从"有限的工具集"扩展到了"整个操作系统"。Anthropic 坦诚它"笨拙、易错",这种诚实的预期管理本身就是工程成熟度的体现。

We're releasing computer use early for feedback from developers, and expect the capability to improve rapidly. Asana, Canva, Cognition, DoorDash, Replit, and The Browser Company have begun exploring these possibilities, carrying out tasks requiring dozens or even hundreds of steps. Replit uses Claude 3.5 Sonnet's computer use and UI navigation to evaluate apps as they're being built. Developers can build with the computer use beta on the Anthropic API, Amazon Bedrock, and Google Cloud's Vertex AI.

Anthropic 提前发布"计算机使用"以收集开发者反馈,并预期该能力会快速改进。Asana、Canva、Cognition、DoorDash、Replit、The Browser Company 等公司已经开始探索这些可能性,完成需要几十甚至上百个步骤的任务。例如 Replit 利用 Claude 3.5 Sonnet 的计算机使用和 UI 导航能力,开发出在应用构建过程中实时评估应用的关键功能。开发者现在可以在 Anthropic API、Amazon Bedrock 和 Google Cloud Vertex AI 上用这个公开测试版构建应用。

Claude 3.5 Sonnet:业界领先的软件工程能力

The updated Claude 3.5 Sonnet shows wide-ranging improvements on industry benchmarks, with particularly strong gains in agentic coding and tool use. On coding, it improves SWE-bench Verified from 33.4% to 49.0% — higher than all publicly available models, including reasoning models like OpenAI o1-preview. It also improves on TAU-bench (agentic tool use): retail 62.6%→69.2%, airline 36.0%→46.0%. These advancements come at the same price and speed as its predecessor.

升级版 Claude 3.5 Sonnet 在行业基准上展现了全面提升,在 Agentic 编码和工具使用任务上进展尤为突出。编码方面,它在 SWE-bench Verified 上从 33.4% 提升到 49.0%——高于当时所有公开可用的模型,包括 OpenAI o1-preview 这类推理模型。它还在 TAU-bench(Agentic 工具使用任务)上有所提升:零售领域从 62.6% 到 69.2%,更具挑战性的航空领域从 36.0% 到 46.0%。而且这些进步在与上一代相同的价格和速度下实现。

💡 AI 解读

SWE-bench Verified 衡量模型解决真实 GitHub issue 的能力,49% 是当时公开模型中的 SOTA。但比绝对数字更重要的是趋势:评估正在从"模型会不会答题"转向"模型会不会调用工具完成多步任务"(TAU-bench)。这正是 Agent 时代的评估范式——能干、会用工具、能在长链条里保持正确,比单次推理的聪明更重要。

Early customer feedback suggests a significant leap for AI-powered coding. GitLab found stronger reasoning (up to 10% across use cases) with no added latency. Cognition uses it for autonomous AI evaluations and saw substantial improvements in coding, planning, and problem-solving. The Browser Company noted it outperformed every model they'd tested for automating web-based workflows. Joint pre-deployment testing was conducted by the US AI Safety Institute (US AISI) and UK AISI; we evaluated catastrophic risks and found the ASL-2 Standard in our Responsible Scaling Policy remains appropriate.

早期客户反馈表明,升级版 Claude 3.5 Sonnet 是 AI 编码能力的一次重大飞跃。GitLab 在 DevSecOps 任务上测试发现它推理更强(各用例提升最高 10%)且无额外延迟,非常适合驱动多步软件开发流程。Cognition 用它做自主 AI 评估,在编码、规划、问题解决上比上一版有大幅提升。The Browser Company 用它自动化网页工作流,称其超越了他们此前测试过的所有模型。作为与外部专家合作的一部分,美国 AI 安全研究所(US AISI)和英国安全研究所(UK AISI)对新模型进行了联合部署前测试。Anthropic 还评估了灾难性风险,确认《负责任扩展政策》中的 ASL-2 标准对该模型依然适用。

Claude 3.5 Haiku:SOTA 与经济性和速度的结合

Claude 3.5 Haiku is the next generation of our fastest model. At a speed similar to Claude 3 Haiku, it improves across every skill set and surpasses even Claude 3 Opus on many intelligence benchmarks. It's particularly strong on coding — 40.6% on SWE-bench Verified, outperforming many agents using publicly available SOTA models, including the original Claude 3.5 Sonnet and GPT-4o. With low latency, improved instruction following, and more accurate tool use, it's well suited for user-facing products, specialized sub-agent tasks, and generating personalized experiences from huge data volumes.

Claude 3.5 Haiku 是 Anthropic 最快模型的下一代。在与 Claude 3 Haiku 相似的速度下,它在每项技能上都取得提升,并在很多智能基准上甚至超越了上一代最大的 Claude 3 Opus。它在编码上尤为出色——SWE-bench Verified 得分 40.6%,超过了许多使用公开 SOTA 模型的 Agent,包括原版 Claude 3.5 Sonnet 和 GPT-4o。凭借低延迟、更强的指令遵循和更准确的工具使用,它非常适合面向用户的产品、专门的子智能体任务,以及从海量数据(购买历史、定价、库存记录等)中生成个性化体验。

💡 AI 解读

Haiku 的定位很清晰:当"主力"太贵时用它。一个 40.6% SWE-bench 的小模型,速度极快、价格极低——这让它成为多智能体系统里"子智能体"的理想选择(呼应 Anthropic 多智能体文章中"主控用 Opus、子智能体用 Sonnet"的思路)。模型矩阵化(大/中/小各司其职)是 Agent 工程控制成本的关键手段。

教 Claude 负责任地导航计算机

With computer use, we're trying something fundamentally new. Instead of making specific tools for individual tasks, we're teaching Claude general computer skills — letting it use standard tools and software designed for people. We've built an API that lets Claude perceive and interact with computer interfaces, translating instructions (e.g. "use data from my computer and online to fill out this form") into computer commands (check a spreadsheet; move the cursor to open a browser; navigate to pages; fill the form). On OSWorld, which evaluates models' ability to use computers like people, Claude 3.5 Sonnet scored 14.9% in the screenshot-only category — notably better than the next-best AI system's 7.8%. With more steps, Claude scored 22.0%.

通过"计算机使用",Anthropic 在尝试一些根本性的新东西。不是为每个任务制作专用工具,而是教 Claude 通用的计算机技能——让它能使用为人类设计的各种标准工具和软件程序。开发者可以用这个能力自动化重复流程、构建和测试软件、执行研究等开放式任务。为此 Anthropic 构建了一个 API,让 Claude 能感知并与计算机界面交互,把指令(如"用我电脑和网上的数据填好这张表")翻译成一连串计算机命令(检查电子表格 → 移动光标打开浏览器 → 导航到相关网页 → 用这些页面的数据填表)。在评估模型像人类一样使用电脑能力的 OSWorld 基准上,Claude 3.5 Sonnet 在"仅截图"类别中得分 14.9%——明显优于次优 AI 系统的 7.8%。当被允许用更多步骤完成任务时,Claude 得分 22.0%。

💡 AI 解读

OSWorld 是理解"计算机使用"含金量的关键——它让模型操作真实的操作系统完成实际任务。14.9% vs 次优 7.8%,意味着 Claude 当时几乎领先一倍。但绝对值仍低(22%),说明这还远不是成熟能力。注意设计思路的转变:从专用 API(为每件事做一个工具)到通用 GUI 操作(像人一样看屏幕点击)。后者覆盖面无限大,但精度和可靠性是巨大挑战。这也是为什么 Anthropic 反复强调"从低风险任务开始"。

While we expect this to improve rapidly, Claude's current ability to use computers is imperfect. Some actions people perform effortlessly — scrolling, dragging, zooming — currently present challenges; we encourage developers to begin with low-risk tasks. Because computer use may provide a new vector for familiar threats (spam, misinformation, fraud), we're taking a proactive approach to safe deployment. We've developed classifiers that identify when computer use is being used and whether harm is occurring.

虽然 Anthropic 预期这项能力会在未来几个月快速提升,但Claude 当前使用计算机的能力并不完美。一些人类轻而易举的动作——滚动、拖拽、缩放——目前对 Claude 仍是挑战,因此建议开发者从低风险任务开始探索。同时,"计算机使用"可能为垃圾信息、虚假信息、欺诈等熟悉的威胁提供新的载体,所以 Anthropic 对安全部署采取了主动策略,开发了新的分类器来识别"计算机使用"是否被调用、是否正在造成危害。

💡 AI 解读

安全这段值得深思。让 AI 直接控制电脑,安全风险是数量级的提升——它能发邮件、转账、发帖、提交表单,任何人类能在电脑上做的坏事它都能加速做。Anthropic 的做法:① 实时分类器监测滥用 ② 劝开发者从低风险起步 ③ 与政府安全机构联合测试。对使用者而言,核心原则是最小权限 + 人在回路:永远不要一上来就给它联网的、能花钱的、能发消息的全权限。

展望

Learning from the initial deployments of this technology — still in its earliest stages — will help us understand both the potential and implications of increasingly capable AI systems. We're excited for you to explore our new models and the public beta of computer use, and welcome your feedback. We believe these developments will open up new possibilities for how you work with Claude, and look forward to seeing what you'll create.

从这项仍处于最早阶段的技术初步部署中学习,将帮助 Anthropic 更好地理解日益强大的 AI 系统的潜力与影响。Anthropic 期待开发者探索新模型和"计算机使用"公开测试版,并欢迎反馈。我们相信这些进展会为"你如何与 Claude 协作"打开新的可能,并期待看到大家创造出什么。

💡 AI 解读

回顾性来看(此文发表于 2024 年 10 月),"计算机使用"是 Agent 从"调 API"走向"操作真实环境"的标志性一步。它的意义在于把 Agent 的能力上限从"开发者能提供的工具数"解放到"软件本身能做的事"。后续的 Claude Computer Use、OpenAI Operator、各类浏览器 Agent 都沿这条路线演进。但对开发者而言,最重要的教训是:通用 GUI 操作的可靠性远低于专用 API——能用 API 就别用点击。计算机使用是"最后一公里的万能钥匙",但只在 API 不存在时才该用。