This commit is contained in:
2026-08-23 02:12:08 +08:00
commit 5987b2a1f2
19 changed files with 4174 additions and 0 deletions
+306
View File
@@ -0,0 +1,306 @@
# Anthropic Claude Messages API
> 来源:
> - https://platform.claude.com/docs/en/api/messages
> - https://platform.claude.com/docs/en/build-with-claude/streaming
> - https://platform.claude.com/docs/en/api/errors
> - https://platform.claude.com/docs/en/build-with-claude/handling-stop-reasons
> - https://platform.claude.com/docs/en/build-with-claude/prompt-caching
> - https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview
>
> 采集于 2026-08-12。注意 `docs.anthropic.com` 已 301 到 `platform.claude.com`。
端点:`POST https://api.anthropic.com/v1/messages`
必需 header`x-api-key``anthropic-version: 2023-06-01``content-type: application/json`
---
## 1. 请求参数
必填:`model``messages`、**`max_tokens`**(与 OpenAI 不同,这里是必填)。
| 参数 | 类型 | 说明 |
| --- | --- | --- |
| `model` | string | 如 `claude-opus-5` |
| `messages` | array\<MessageParam\> | |
| `max_tokens` | number | **必填** |
| `system` | string \| array\<TextBlockParam\> | **顶层独立字段**,不是 messages 里的一条 |
| `temperature` | number | **0.0 ~ 1.0**,默认 1.0(不是 OpenAI 的 0~2 |
| `top_p` | number | |
| `top_k` | number | OpenAI 无此参数 |
| `stop_sequences` | array\<string\> | |
| `stream` | boolean | |
| `tools` | array\<ToolUnion\> | |
| `tool_choice` | ToolChoice | |
| `thinking` | ThinkingConfigParam | 见 §5 |
| `metadata` | Metadata | |
| `cache_control` | CacheControlEphemeral | 顶层写法 = 自动缓存 |
| `container` | string | 代码执行容器 |
| `inference_geo` | string | 推理地理区域 |
| `output_config` | OutputConfig | 含 `effort``format`(结构化输出) |
| `service_tier` | `"auto"` \| `"standard_only"` | |
Header 参数:`anthropic-user-profile-id`(可选)。
### MessageParam
```json
{"role": "user" | "assistant" | "system", "content": "string" | [ContentBlockParam]}
```
---
## 2. Content Block 类型
### 输入侧(ContentBlockParam
| type | 关键字段 |
| --- | --- |
| `text` | `text`, `cache_control`, `citations` |
| `image` | `source`: `{type:"base64", media_type: image/jpeg\|png\|gif\|webp, data}``{type:"url", url}` |
| `document` | `source`: base64(PDF) / text / url / content`title`, `context`, `citations:{enabled}` |
| `tool_use` | `id`, `name`, `input`, `caller` |
| `tool_result` | `tool_use_id`, `content`, `is_error` |
| `thinking` | `thinking`, `signature` |
| `search_result` | `title`, `source`, `content[]`, `citations` |
| `server_tool_use` | `id`, `name`(web_search/web_fetch/code_execution/bash_code_execution/text_editor_code_execution/tool_search_tool_regex/tool_search_tool_bm25), `input` |
| `mid_conv_system` | `content: [TextBlockParam]` — 会话中途插入 system 指令 |
| `container_upload` | `file_id` |
### 输出侧(ContentBlock
`text``thinking``redacted_thinking``{type, data}`)、`tool_use``server_tool_use``web_search_tool_result``web_fetch_tool_result``code_execution_tool_result``bash_code_execution_tool_result``text_editor_code_execution_tool_result``tool_search_tool_result``container_upload`
---
## 3. 响应结构
```json
{
"id": "msg_...",
"type": "message",
"role": "assistant",
"model": "claude-opus-5",
"content": [ContentBlock],
"stop_reason": "end_turn",
"stop_sequence": null,
"stop_details": {"type": "refusal", "category": "...", "explanation": "..."},
"usage": {...},
"container": {"id": "...", "expires_at": "..."}
}
```
### stop_reason
| 值 | 含义 | 处理 |
| --- | --- | --- |
| `end_turn` | 自然结束 | 直接用 |
| `max_tokens` | 达到 `max_tokens` | 截断,可续写 |
| `stop_sequence` | 命中 `stop_sequences` | 见 `stop_sequence` 字段 |
| `tool_use` | 调用了**客户端**工具 | 执行后回 `tool_result` |
| `pause_turn` | **服务端**工具采样循环达到迭代上限(默认 10 次) | 把整个 response.content 作为 assistant 消息发回去继续,**工具定义必须保持一致** |
| `refusal` | 安全拒绝 | 读 `stop_details.category` |
| `model_context_window_exceeded` | 响应填满上下文窗口(先于 max_tokens 触发) | 视作截断 |
`stop_details.category` 取值:`cyber` | `bio` | `frontier_llm` | `reasoning_extraction` | `general_harms`
### usage
```json
{
"input_tokens": 50,
"output_tokens": 503,
"cache_creation_input_tokens": 0,
"cache_read_input_tokens": 100000,
"cache_creation": {"ephemeral_5m_input_tokens": 0, "ephemeral_1h_input_tokens": 0},
"output_tokens_details": {"thinking_tokens": 0},
"server_tool_use": {"web_search_requests": 0, "web_fetch_requests": 0},
"inference_geo": "...",
"service_tier": "standard" | "priority" | "batch"
}
```
**`input_tokens` 只统计最后一个缓存断点之后的未缓存 token**:
`total = input_tokens + cache_read_input_tokens + cache_creation_input_tokens`
直接把 `input_tokens` 映射到 OpenAI 的 `prompt_tokens` 会严重低报。
### Token 计数端点
`POST /v1/messages/count_tokens`,参数同 create(不需要 `max_tokens`),返回 `input_tokens``cache_creation_input_tokens``cache_read_input_tokens`
---
## 4. 流式(SSE
事件流固定顺序:
1. `message_start` — 含一个 `content` 为空的 Message 对象
2. 若干组 content block`content_block_start` → N × `content_block_delta``content_block_stop`,每组带 `index`
3. 一个或多个 `message_delta` — 顶层变更(`stop_reason``usage`
4. `message_stop`
其间可穿插任意数量的 `ping` 事件,以及 `error` 事件。
### delta 类型
| delta.type | 字段 | 用于 |
| --- | --- | --- |
| `text_delta` | `text` | text block |
| `input_json_delta` | `partial_json` | tool_use 的 `input`,**是部分 JSON 字符串**,累加后在 `content_block_stop` 时解析 |
| `thinking_delta` | `thinking` | thinking block |
| `signature_delta` | `signature` | thinking block**在 `content_block_stop` 之前**发一次 |
### 流式怪癖
- **`message_delta.usage` 的 token 计数是累计值**,不是增量。做统计时不要相加。
- 流式返回 200 之后仍可能中途发 `error` 事件(例如 `overloaded_error`,对应非流式的 HTTP 529)。错误处理不能只看 HTTP 状态码。
- `content_block_delta``index` 对应最终 `content` 数组下标,**不保证连续到达顺序之外的假设**,但同一 block 的 delta 是有序的。
- **server-side fallback** 时会在模型边界插入一个 `fallback` content block:只有 `content_block_start` + `content_block_stop`,中间**没有任何 delta**。解析器不能假设每个 block 至少有一个 delta。
- 工具调用时,模型一次只能吐出一个完整的 key-value,因此**事件之间可能有明显停顿**。
- `thinking` 配置为 `display: "omitted"` 时,thinking block 会开启、只收到一个 `signature_delta`、然后关闭,**没有 `thinking_delta`**。
- 新事件类型会随时新增,解析器必须优雅忽略未知事件(官方 versioning policy 明确要求)。
### 流式中断恢复
- Claude 4.5 及更早:把已收到的部分内容作为 assistant 消息前缀(prefill)续写。
- **Claude 4.6 及更晚:不支持 prefill**,改为加一条 user 消息说明"上次回复被中断于 X,请继续"。
- `tool_use` 和 thinking block **无法部分恢复**,只能从最后一个完整 text block 续。
---
## 5. Thinking 的怪癖(最大坑区)
Anthropic 在不同模型代际上换过三种 thinking 配置,且**互相排斥**
| 模型 | 可用配置 | 报错情况 |
| --- | --- | --- |
| Claude 4.5 及更早 | `thinking: {"type":"enabled", "budget_tokens":N}` | 传 `type:"adaptive"` → 400 `adaptive thinking is not supported on this model` |
| Claude 4.7 及以后 | `thinking: {"type":"adaptive", "display":...}` + `output_config.effort` | 传 `type:"enabled"` → 400 `"thinking.type.enabled" is not supported for this model` |
| Claude Fable 5 / Mythos 5 / Mythos Preview | thinking **恒开**,无法关闭 | 传 `type:"disabled"` → 400。且该错误信息建议的 `"thinking.type.enabled"` 在 Fable 5 / Mythos 5 上**同样被拒**——正确做法是**完全省略 `thinking` 参数** |
`display` 取值:`summarized` | `omitted``omitted` 用于"开启思考但不返回思考内容")。
### signature 的真实机制(关键)
> 来源:https://platform.claude.com/docs/en/build-with-claude/thinking 的 "Thinking encryption" 章节
**`signature` 名为签名,实为加密载荷。** 官方原文:
> "Full thinking content is encrypted and returned in the `signature` field on each thinking block. The API uses the signature to verify that thinking blocks were generated by Claude when you pass them back."
也就是说它身兼两职:**承载完整原始思考内容的密文** + **来源真实性校验**
由此推出几个反直觉的事实:
1. **`thinking` 字段里的文本永远不是原始思考链**,而是摘要(summarized)。官方明确:"what you see is never the raw chain of thought"。**没有任何 `display` 设置能返回原始 CoT。**
2. **真正的推理内容在 `signature` 里。** 回传时服务端**解密 `signature` 来重建原始 thinking** 用于构造 prompt——它不读你传的 `thinking` 文本。
3. **`display: "omitted"``thinking` 为空串,但 `signature` 照常携带完整加密思考。** 且 `signature` 的值在 `summarized``omitted` 两种模式下**完全相同**,中途切换 `display` 是被支持的。
4. 因此 `display: "omitted"` 省的是**延迟不是钱**——完整思考 token 照常计费,只是服务端跳过流式下发思考文本,首个 text token 更快到达。
其他约束:
- **`signature` 是 opaque 的,官方明令不要解析或解释它。**
- Claude 4 及以后模型的 `signature` **显著更长**。若中转站对单字段长度有限制(数据库列宽、日志截断),这是个真实的踩坑点。
- **`signature` 跨平台通用**Claude API、Amazon Bedrock、Google Vertex AI 三方生成的值可以互换使用。对多云路由的中转站是好消息。
- `redacted_thinking` 是**另一种 block type**(安全红线内容被编辑时返回),加密内容在 `data` 字段,同样 opaque、同样必须原样回传。它与 `display: "omitted"` 是两回事。
- Claude Fable 5 / Mythos 5 **从不返回原始思考链**,返回的是常规 `thinking` block 而非 `redacted_thinking`
### 回传规则(分级,不是一刀切)
我最初的理解过严。官方的实际分级是:
| 场景 | 要求 |
| --- | --- |
| **工具使用轮次内** | **必须**回传 thinking block |
| 跨轮次(非工具) | **建议**全部回传 |
| 非工具场景 | **允许**省略历史轮次的 thinking |
但只要你选择回传,就**不能改**:最近一条 assistant 消息里连续的 `thinking` block 序列必须与模型原始生成的完全一致——不能重排、编辑、部分丢弃,`redacted_thinking` 也算在内。否则 400
```
`thinking` or `redacted_thinking` blocks in the latest assistant message cannot be modified
```
**唯一的例外**`display: "omitted"` 的块里,你写进那个空 `thinking` 字段的任何文本会被**忽略**而不是报错(因为服务端根本不读它,只解密 signature)。
**不需要自己修剪历史 thinking。** 全部传回去,API 会自动过滤,只保留维持推理连续性所需的块,并且**只对实际展示给 Claude 的块计费**。
### 保留策略按模型分两档
| 档位 | 模型 |
| --- | --- |
| **保留所有历史轮次** | Opus 4.5 及以后的 Opus、Sonnet 4.6 及以后的 Sonnet、Fable 5、Mythos 5、Mythos Preview |
| **只保留最后一轮** | 更早的 Opus/Sonnet,以及**所有 Haiku**(含 Haiku 4.5)。传回更早的块 API 会自动剥离 |
在 keep-all 模型上,长对话的历史 thinking block 会**持续占用输入 token**(从缓存读取时也计入 usage)。可用 `clear_thinking_20251015` 这个 context-editing 策略覆盖默认行为。
### 换模型必须剥离 thinking
**thinking block 与生成它的模型绑定。** 会话中途切换模型(例如 refusal fallback 之后)时,必须从历史 assistant 轮次中剥离 `thinking``redacted_thinking`
危险之处:**其他模型不会拒绝,而是静默忽略**——但被忽略的块**照样计入输入 token**。中转站做模型路由/故障转移时如果不清理,用户会为完全无用的内容付费,且不会收到任何错误提示。
### 对中转站的含义
好消息比我最初判断的多:signature 是 opaque 且跨平台通用的 blob,中转站**只需原样搬运,不需要理解它**。
可行的做法:
- 网关侧按会话存储上游返回的原始 content block 数组(thinking 的 `signature` 和 redacted 的 `data` 作为 opaque blob),下游只暴露脱敏后的摘要文本;下一轮用存储的原始块重建上游请求。
- 非工具场景可以直接丢弃 thinking 块,合法且省上下文。
- **绝不能**用"把 thinking 文本转成 `reasoning_content` 字符串再转回来"的方式——转回来的是摘要文本、且没有 signatureAnthropic 侧必然 400 或丢失推理。
---
## 6. Prompt Caching 的怪癖
缓存前缀层级顺序固定:**`tools``system``messages`**。改动某一层会使该层**及之后所有层**失效。
| 改动 | tools 失效 | system 失效 | messages 失效 |
| --- | --- | --- | --- |
| 工具定义 | ✘ | ✘ | ✘ |
| web search / citations 开关 | ✓保留 | ✘ | ✘ |
| speed 设置 | ✓保留 | ✘ | ✘ |
| `tool_choice` | ✓保留 | ✓保留 | ✘ |
| 增删图片 | ✓保留 | ✓保留 | ✘ |
| thinking 配置 | 模型相关 | 模型相关 | ✘ |
- **最多 4 个显式 cache breakpoint**。顶层 `cache_control` 的自动缓存**占用 1 个槽位**,所以自动+显式并用时显式最多 3 个。
- **最小可缓存 token 数按模型不同**,低于阈值时静默不缓存、**不报错**:
| 模型 | 最小 token |
| --- | --- |
| Opus 5 / Fable 5 / Mythos 5 | 512 |
| Opus 4.8 / Sonnet 5 / Sonnet 4.6 / Sonnet 4.5 / Opus 4.1 | 1,024 |
| Opus 4.7 / Mythos Preview / Haiku 3.5 | 2,048 |
| Opus 4.6 / Opus 4.5 / Haiku 4.5 | 4,096 |
- TTL:默认 `5m`,可设 `{"type":"ephemeral","ttl":"1h"}`。**1h 的断点必须出现在 5m 断点之前**。
- 定价倍率(相对基础输入价):5m 写入 **1.25x**1h 写入 **2.0x**,缓存读取 **0.1x**
- 缓存查找**最多回溯 20 个 block**,超出则必须重建缓存。
- 并发请求时,缓存条目要等第一个响应**开始**后才可用——同时打多个相同前缀的请求不会共享缓存。
- `thinking` block 不能显式打 `cache_control`,但会随其他内容一起被缓存。
---
## 7. Tool use 的怪癖
- 工具定义用 **`input_schema`**(不是 OpenAI 的 `parameters`),且没有 `function` 外层包装:`{name, description, input_schema}`
- `tool_result` 必须放在 **user 消息**的 content 数组里,通过 `tool_use_id` 关联。OpenAI 用的是独立的 `role: "tool"` 消息。
- `tool_choice` 取值:`{"type":"auto"}` | `{"type":"any"}` | `{"type":"tool","name":"x"}` | `{"type":"none"}`,另有 `disable_parallel_tool_use: true`(对应 OpenAI 的 `parallel_tool_calls: false`)。注意是 **`any`** 不是 OpenAI 的 `required`
- 错误用 `tool_result``is_error: true` 表达,不是抛异常。
- **不要在 `tool_result` 之后同一条 user 消息里追加 text block**——官方明确指出这会导致 Claude 返回空响应(`stop_reason: end_turn` 但 content 为空)。
- `strict: true` 可加在自定义工具定义上以保证 schema 一致(与 OpenAI 的 strict 语义相近,但配置位置不同)。
- **tools 会消耗额外的系统提示 token**,随模型和 `tool_choice` 而变。例如 Opus 5`auto`/`none` 为 286 tokens`any`/`tool` 为 406 tokens。中转站做 token 预估时必须计入。
- 服务端工具(`web_search` 等)会在**同一次响应内**返回结果 block,无需客户端执行;但若它和客户端工具在同一组并行调用里,行为会变(见 `pause_turn`)。
---
## 8. 其他常见 400
- **Prefill 不支持**Claude 4.6 及以后、Mythos Preview 不允许最后一条是 assistant 消息作为前缀。
```
This model does not support assistant message prefill. The conversation must end with a user message.
```
替代方案:结构化输出 / system 指令 / `output_config.format`。
这条对中转站影响很大——**DeepSeek 的 prefix completion、OpenAI 的部分模式无法映射到新 Claude 模型**。
- `temperature` 上限是 **1.0**。OpenAI 侧 0~2 的值需要 clampAnthropic 官方兼容层的做法是 >1 一律截断为 1)。