# Anthropic Claude Messages API > 来源: > - https://platform.claude.com/docs/en/api/messages > - https://platform.claude.com/docs/en/build-with-claude/streaming > - https://platform.claude.com/docs/en/api/errors > - https://platform.claude.com/docs/en/build-with-claude/handling-stop-reasons > - https://platform.claude.com/docs/en/build-with-claude/prompt-caching > - https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview > > 采集于 2026-08-12。注意 `docs.anthropic.com` 已 301 到 `platform.claude.com`。 端点:`POST https://api.anthropic.com/v1/messages` 必需 header:`x-api-key`、`anthropic-version: 2023-06-01`、`content-type: application/json`。 --- ## 1. 请求参数 必填:`model`、`messages`、**`max_tokens`**(与 OpenAI 不同,这里是必填)。 | 参数 | 类型 | 说明 | | --- | --- | --- | | `model` | string | 如 `claude-opus-5` | | `messages` | array\ | | | `max_tokens` | number | **必填** | | `system` | string \| array\ | **顶层独立字段**,不是 messages 里的一条 | | `temperature` | number | **0.0 ~ 1.0**,默认 1.0(不是 OpenAI 的 0~2) | | `top_p` | number | | | `top_k` | number | OpenAI 无此参数 | | `stop_sequences` | array\ | | | `stream` | boolean | | | `tools` | array\ | | | `tool_choice` | ToolChoice | | | `thinking` | ThinkingConfigParam | 见 §5 | | `metadata` | Metadata | | | `cache_control` | CacheControlEphemeral | 顶层写法 = 自动缓存 | | `container` | string | 代码执行容器 | | `inference_geo` | string | 推理地理区域 | | `output_config` | OutputConfig | 含 `effort`、`format`(结构化输出) | | `service_tier` | `"auto"` \| `"standard_only"` | | Header 参数:`anthropic-user-profile-id`(可选)。 ### MessageParam ```json {"role": "user" | "assistant" | "system", "content": "string" | [ContentBlockParam]} ``` --- ## 2. Content Block 类型 ### 输入侧(ContentBlockParam) | type | 关键字段 | | --- | --- | | `text` | `text`, `cache_control`, `citations` | | `image` | `source`: `{type:"base64", media_type: image/jpeg\|png\|gif\|webp, data}` 或 `{type:"url", url}` | | `document` | `source`: base64(PDF) / text / url / content;`title`, `context`, `citations:{enabled}` | | `tool_use` | `id`, `name`, `input`, `caller` | | `tool_result` | `tool_use_id`, `content`, `is_error` | | `thinking` | `thinking`, `signature` | | `search_result` | `title`, `source`, `content[]`, `citations` | | `server_tool_use` | `id`, `name`(web_search/web_fetch/code_execution/bash_code_execution/text_editor_code_execution/tool_search_tool_regex/tool_search_tool_bm25), `input` | | `mid_conv_system` | `content: [TextBlockParam]` — 会话中途插入 system 指令 | | `container_upload` | `file_id` | ### 输出侧(ContentBlock) `text`、`thinking`、`redacted_thinking`(`{type, data}`)、`tool_use`、`server_tool_use`、`web_search_tool_result`、`web_fetch_tool_result`、`code_execution_tool_result`、`bash_code_execution_tool_result`、`text_editor_code_execution_tool_result`、`tool_search_tool_result`、`container_upload`。 --- ## 3. 响应结构 ```json { "id": "msg_...", "type": "message", "role": "assistant", "model": "claude-opus-5", "content": [ContentBlock], "stop_reason": "end_turn", "stop_sequence": null, "stop_details": {"type": "refusal", "category": "...", "explanation": "..."}, "usage": {...}, "container": {"id": "...", "expires_at": "..."} } ``` ### stop_reason | 值 | 含义 | 处理 | | --- | --- | --- | | `end_turn` | 自然结束 | 直接用 | | `max_tokens` | 达到 `max_tokens` | 截断,可续写 | | `stop_sequence` | 命中 `stop_sequences` | 见 `stop_sequence` 字段 | | `tool_use` | 调用了**客户端**工具 | 执行后回 `tool_result` | | `pause_turn` | **服务端**工具采样循环达到迭代上限(默认 10 次) | 把整个 response.content 作为 assistant 消息发回去继续,**工具定义必须保持一致** | | `refusal` | 安全拒绝 | 读 `stop_details.category` | | `model_context_window_exceeded` | 响应填满上下文窗口(先于 max_tokens 触发) | 视作截断 | `stop_details.category` 取值:`cyber` | `bio` | `frontier_llm` | `reasoning_extraction` | `general_harms`。 ### usage ```json { "input_tokens": 50, "output_tokens": 503, "cache_creation_input_tokens": 0, "cache_read_input_tokens": 100000, "cache_creation": {"ephemeral_5m_input_tokens": 0, "ephemeral_1h_input_tokens": 0}, "output_tokens_details": {"thinking_tokens": 0}, "server_tool_use": {"web_search_requests": 0, "web_fetch_requests": 0}, "inference_geo": "...", "service_tier": "standard" | "priority" | "batch" } ``` **`input_tokens` 只统计最后一个缓存断点之后的未缓存 token**: `total = input_tokens + cache_read_input_tokens + cache_creation_input_tokens`。 直接把 `input_tokens` 映射到 OpenAI 的 `prompt_tokens` 会严重低报。 ### Token 计数端点 `POST /v1/messages/count_tokens`,参数同 create(不需要 `max_tokens`),返回 `input_tokens`、`cache_creation_input_tokens`、`cache_read_input_tokens`。 --- ## 4. 流式(SSE) 事件流固定顺序: 1. `message_start` — 含一个 `content` 为空的 Message 对象 2. 若干组 content block:`content_block_start` → N × `content_block_delta` → `content_block_stop`,每组带 `index` 3. 一个或多个 `message_delta` — 顶层变更(`stop_reason`、`usage`) 4. `message_stop` 其间可穿插任意数量的 `ping` 事件,以及 `error` 事件。 ### delta 类型 | delta.type | 字段 | 用于 | | --- | --- | --- | | `text_delta` | `text` | text block | | `input_json_delta` | `partial_json` | tool_use 的 `input`,**是部分 JSON 字符串**,累加后在 `content_block_stop` 时解析 | | `thinking_delta` | `thinking` | thinking block | | `signature_delta` | `signature` | thinking block,**在 `content_block_stop` 之前**发一次 | ### 流式怪癖 - **`message_delta.usage` 的 token 计数是累计值**,不是增量。做统计时不要相加。 - 流式返回 200 之后仍可能中途发 `error` 事件(例如 `overloaded_error`,对应非流式的 HTTP 529)。错误处理不能只看 HTTP 状态码。 - `content_block_delta` 的 `index` 对应最终 `content` 数组下标,**不保证连续到达顺序之外的假设**,但同一 block 的 delta 是有序的。 - **server-side fallback** 时会在模型边界插入一个 `fallback` content block:只有 `content_block_start` + `content_block_stop`,中间**没有任何 delta**。解析器不能假设每个 block 至少有一个 delta。 - 工具调用时,模型一次只能吐出一个完整的 key-value,因此**事件之间可能有明显停顿**。 - `thinking` 配置为 `display: "omitted"` 时,thinking block 会开启、只收到一个 `signature_delta`、然后关闭,**没有 `thinking_delta`**。 - 新事件类型会随时新增,解析器必须优雅忽略未知事件(官方 versioning policy 明确要求)。 ### 流式中断恢复 - Claude 4.5 及更早:把已收到的部分内容作为 assistant 消息前缀(prefill)续写。 - **Claude 4.6 及更晚:不支持 prefill**,改为加一条 user 消息说明"上次回复被中断于 X,请继续"。 - `tool_use` 和 thinking block **无法部分恢复**,只能从最后一个完整 text block 续。 --- ## 5. Thinking 的怪癖(最大坑区) Anthropic 在不同模型代际上换过三种 thinking 配置,且**互相排斥**: | 模型 | 可用配置 | 报错情况 | | --- | --- | --- | | Claude 4.5 及更早 | `thinking: {"type":"enabled", "budget_tokens":N}` | 传 `type:"adaptive"` → 400 `adaptive thinking is not supported on this model` | | Claude 4.7 及以后 | `thinking: {"type":"adaptive", "display":...}` + `output_config.effort` | 传 `type:"enabled"` → 400 `"thinking.type.enabled" is not supported for this model` | | Claude Fable 5 / Mythos 5 / Mythos Preview | thinking **恒开**,无法关闭 | 传 `type:"disabled"` → 400。且该错误信息建议的 `"thinking.type.enabled"` 在 Fable 5 / Mythos 5 上**同样被拒**——正确做法是**完全省略 `thinking` 参数** | `display` 取值:`summarized` | `omitted`(`omitted` 用于"开启思考但不返回思考内容")。 ### signature 的真实机制(关键) > 来源:https://platform.claude.com/docs/en/build-with-claude/thinking 的 "Thinking encryption" 章节 **`signature` 名为签名,实为加密载荷。** 官方原文: > "Full thinking content is encrypted and returned in the `signature` field on each thinking block. The API uses the signature to verify that thinking blocks were generated by Claude when you pass them back." 也就是说它身兼两职:**承载完整原始思考内容的密文** + **来源真实性校验**。 由此推出几个反直觉的事实: 1. **`thinking` 字段里的文本永远不是原始思考链**,而是摘要(summarized)。官方明确:"what you see is never the raw chain of thought"。**没有任何 `display` 设置能返回原始 CoT。** 2. **真正的推理内容在 `signature` 里。** 回传时服务端**解密 `signature` 来重建原始 thinking** 用于构造 prompt——它不读你传的 `thinking` 文本。 3. **`display: "omitted"` 时 `thinking` 为空串,但 `signature` 照常携带完整加密思考。** 且 `signature` 的值在 `summarized` 和 `omitted` 两种模式下**完全相同**,中途切换 `display` 是被支持的。 4. 因此 `display: "omitted"` 省的是**延迟不是钱**——完整思考 token 照常计费,只是服务端跳过流式下发思考文本,首个 text token 更快到达。 其他约束: - **`signature` 是 opaque 的,官方明令不要解析或解释它。** - Claude 4 及以后模型的 `signature` **显著更长**。若中转站对单字段长度有限制(数据库列宽、日志截断),这是个真实的踩坑点。 - **`signature` 跨平台通用**:Claude API、Amazon Bedrock、Google Vertex AI 三方生成的值可以互换使用。对多云路由的中转站是好消息。 - `redacted_thinking` 是**另一种 block type**(安全红线内容被编辑时返回),加密内容在 `data` 字段,同样 opaque、同样必须原样回传。它与 `display: "omitted"` 是两回事。 - Claude Fable 5 / Mythos 5 **从不返回原始思考链**,返回的是常规 `thinking` block 而非 `redacted_thinking`。 ### 回传规则(分级,不是一刀切) 我最初的理解过严。官方的实际分级是: | 场景 | 要求 | | --- | --- | | **工具使用轮次内** | **必须**回传 thinking block | | 跨轮次(非工具) | **建议**全部回传 | | 非工具场景 | **允许**省略历史轮次的 thinking | 但只要你选择回传,就**不能改**:最近一条 assistant 消息里连续的 `thinking` block 序列必须与模型原始生成的完全一致——不能重排、编辑、部分丢弃,`redacted_thinking` 也算在内。否则 400: ``` `thinking` or `redacted_thinking` blocks in the latest assistant message cannot be modified ``` **唯一的例外**:`display: "omitted"` 的块里,你写进那个空 `thinking` 字段的任何文本会被**忽略**而不是报错(因为服务端根本不读它,只解密 signature)。 **不需要自己修剪历史 thinking。** 全部传回去,API 会自动过滤,只保留维持推理连续性所需的块,并且**只对实际展示给 Claude 的块计费**。 ### 保留策略按模型分两档 | 档位 | 模型 | | --- | --- | | **保留所有历史轮次** | Opus 4.5 及以后的 Opus、Sonnet 4.6 及以后的 Sonnet、Fable 5、Mythos 5、Mythos Preview | | **只保留最后一轮** | 更早的 Opus/Sonnet,以及**所有 Haiku**(含 Haiku 4.5)。传回更早的块 API 会自动剥离 | 在 keep-all 模型上,长对话的历史 thinking block 会**持续占用输入 token**(从缓存读取时也计入 usage)。可用 `clear_thinking_20251015` 这个 context-editing 策略覆盖默认行为。 ### 换模型必须剥离 thinking **thinking block 与生成它的模型绑定。** 会话中途切换模型(例如 refusal fallback 之后)时,必须从历史 assistant 轮次中剥离 `thinking` 和 `redacted_thinking`。 危险之处:**其他模型不会拒绝,而是静默忽略**——但被忽略的块**照样计入输入 token**。中转站做模型路由/故障转移时如果不清理,用户会为完全无用的内容付费,且不会收到任何错误提示。 ### 对中转站的含义 好消息比我最初判断的多:signature 是 opaque 且跨平台通用的 blob,中转站**只需原样搬运,不需要理解它**。 可行的做法: - 网关侧按会话存储上游返回的原始 content block 数组(thinking 的 `signature` 和 redacted 的 `data` 作为 opaque blob),下游只暴露脱敏后的摘要文本;下一轮用存储的原始块重建上游请求。 - 非工具场景可以直接丢弃 thinking 块,合法且省上下文。 - **绝不能**用"把 thinking 文本转成 `reasoning_content` 字符串再转回来"的方式——转回来的是摘要文本、且没有 signature,Anthropic 侧必然 400 或丢失推理。 --- ## 6. Prompt Caching 的怪癖 缓存前缀层级顺序固定:**`tools` → `system` → `messages`**。改动某一层会使该层**及之后所有层**失效。 | 改动 | tools 失效 | system 失效 | messages 失效 | | --- | --- | --- | --- | | 工具定义 | ✘ | ✘ | ✘ | | web search / citations 开关 | ✓保留 | ✘ | ✘ | | speed 设置 | ✓保留 | ✘ | ✘ | | `tool_choice` | ✓保留 | ✓保留 | ✘ | | 增删图片 | ✓保留 | ✓保留 | ✘ | | thinking 配置 | 模型相关 | 模型相关 | ✘ | - **最多 4 个显式 cache breakpoint**。顶层 `cache_control` 的自动缓存**占用 1 个槽位**,所以自动+显式并用时显式最多 3 个。 - **最小可缓存 token 数按模型不同**,低于阈值时静默不缓存、**不报错**: | 模型 | 最小 token | | --- | --- | | Opus 5 / Fable 5 / Mythos 5 | 512 | | Opus 4.8 / Sonnet 5 / Sonnet 4.6 / Sonnet 4.5 / Opus 4.1 | 1,024 | | Opus 4.7 / Mythos Preview / Haiku 3.5 | 2,048 | | Opus 4.6 / Opus 4.5 / Haiku 4.5 | 4,096 | - TTL:默认 `5m`,可设 `{"type":"ephemeral","ttl":"1h"}`。**1h 的断点必须出现在 5m 断点之前**。 - 定价倍率(相对基础输入价):5m 写入 **1.25x**,1h 写入 **2.0x**,缓存读取 **0.1x**。 - 缓存查找**最多回溯 20 个 block**,超出则必须重建缓存。 - 并发请求时,缓存条目要等第一个响应**开始**后才可用——同时打多个相同前缀的请求不会共享缓存。 - `thinking` block 不能显式打 `cache_control`,但会随其他内容一起被缓存。 --- ## 7. Tool use 的怪癖 - 工具定义用 **`input_schema`**(不是 OpenAI 的 `parameters`),且没有 `function` 外层包装:`{name, description, input_schema}`。 - `tool_result` 必须放在 **user 消息**的 content 数组里,通过 `tool_use_id` 关联。OpenAI 用的是独立的 `role: "tool"` 消息。 - `tool_choice` 取值:`{"type":"auto"}` | `{"type":"any"}` | `{"type":"tool","name":"x"}` | `{"type":"none"}`,另有 `disable_parallel_tool_use: true`(对应 OpenAI 的 `parallel_tool_calls: false`)。注意是 **`any`** 不是 OpenAI 的 `required`。 - 错误用 `tool_result` 的 `is_error: true` 表达,不是抛异常。 - **不要在 `tool_result` 之后同一条 user 消息里追加 text block**——官方明确指出这会导致 Claude 返回空响应(`stop_reason: end_turn` 但 content 为空)。 - `strict: true` 可加在自定义工具定义上以保证 schema 一致(与 OpenAI 的 strict 语义相近,但配置位置不同)。 - **tools 会消耗额外的系统提示 token**,随模型和 `tool_choice` 而变。例如 Opus 5:`auto`/`none` 为 286 tokens,`any`/`tool` 为 406 tokens。中转站做 token 预估时必须计入。 - 服务端工具(`web_search` 等)会在**同一次响应内**返回结果 block,无需客户端执行;但若它和客户端工具在同一组并行调用里,行为会变(见 `pause_turn`)。 --- ## 8. 其他常见 400 - **Prefill 不支持**:Claude 4.6 及以后、Mythos Preview 不允许最后一条是 assistant 消息作为前缀。 ``` This model does not support assistant message prefill. The conversation must end with a user message. ``` 替代方案:结构化输出 / system 指令 / `output_config.format`。 这条对中转站影响很大——**DeepSeek 的 prefix completion、OpenAI 的部分模式无法映射到新 Claude 模型**。 - `temperature` 上限是 **1.0**。OpenAI 侧 0~2 的值需要 clamp(Anthropic 官方兼容层的做法是 >1 一律截断为 1)。