feat: revamp analytics dashboards and harden database migrations

Add dashboard and overview analytics, health monitoring, provider expense tracking, and announcement updates across the gateway and frontend.

Keep schema migrations free of historical backfills while preserving automatic backfill execution. Bound migration deadlines, run schema preparation before Compose replacement, and anonymize deleted dashboard users.

Include the current documentation cleanup and regression coverage.
This commit is contained in:
elky
2026-10-01 11:48:17 +08:00
parent 60b89cc840
commit 066ea87d72
327 changed files with 31728 additions and 20645 deletions
-222
View File
@@ -1,222 +0,0 @@
# Embeddings API
Aether supports OpenAI compatible embedding requests through `POST /v1/embeddings`. Embedding requests are separate from chat and responses requests. They use `input`, never `messages`, and they are always non streaming.
## Quick Start
Run this against your Aether gateway URL with a user API key that can access the model and the `openai:embedding` API format.
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: application/json" \
-d '{
"model": "text-embedding-3-small",
"input": ["hello", "world"],
"encoding_format": "float"
}'
```
## Public Endpoint
| Method | Path | Client API format | Route kind |
| --- | --- | --- | --- |
| `POST` | `/v1/embeddings` | `openai:embedding` | `embedding` |
The gateway classifies this endpoint as an OpenAI family embedding route with endpoint signature `openai:embedding`. It is not handled as chat or responses.
## Request Body
Required fields:
| Field | Type | Notes |
| --- | --- | --- |
| `model` | string | Must name a model allowed for the API key and user. Blank strings are rejected. |
| `input` | string, string array, integer token array, nested integer token arrays, or multimodal object array | Must be non empty. Empty strings, empty arrays, empty token arrays, and empty multimodal objects are rejected. |
Optional fields that pass through the embedding conversion path when supported by the provider:
| Field | Notes |
| --- | --- |
| `encoding_format` | Passed to OpenAI compatible providers. |
| `dimensions` | Passed to providers whose embedding request shape supports it. |
| `parameters` | Provider-specific embedding parameters. For Aliyun DashScope this maps to DashScope `parameters`; `dimensions` is emitted as `parameters.dimension` unless `parameters.dimension` is already set. |
| `user` | Passed to OpenAI compatible providers. |
| `task` | Passed to Jina and OpenAI compatible embedding requests. Jina defaults to `text-matching` when no task is supplied. |
Accepted `input` shapes:
```json
{ "model": "text-embedding-3-small", "input": "hello" }
```
```json
{ "model": "text-embedding-3-small", "input": ["hello", "world"] }
```
```json
{ "model": "text-embedding-3-small", "input": [1, 2, 3] }
```
```json
{ "model": "text-embedding-3-small", "input": [[1, 2], [3, 4]] }
```
```json
{
"model": "qwen3-vl-embedding",
"input": [
{ "text": "white running shoes" },
{ "image": "https://dashscope.oss-cn-beijing.aliyuncs.com/images/256_1.png" }
],
"parameters": { "enable_fusion": true }
}
```
Use string or string array input when routing to Gemini or Doubao embedding providers. Token arrays are accepted by the OpenAI compatible public endpoint, but Gemini, Doubao, and Aliyun provider request emitters require text or multimodal content input.
## Provider Format Mapping
Embedding routes can select only embedding provider API formats. Chat, responses, image, and generation formats are not valid provider targets for this request type.
| Provider API format | Upstream path shape | Provider request shape |
| --- | --- | --- |
| `openai:embedding` | `/v1/embeddings` | OpenAI compatible `{ "model", "input" }` payload. |
| `jina:embedding` | `/v1/embeddings` | OpenAI compatible payload with a Jina `task`. Defaults to `text-matching` if omitted. |
| `gemini:embedding` | `models/{model}:embedContent` | Single text input uses `content.parts[].text`. Multiple text inputs use `requests[].content.parts[].text`. |
| `doubao:embedding` | `/embeddings/multimodal` | Text input is emitted as `input` items like `{ "type": "text", "text": "..." }`. |
| `aliyun:multimodal_embedding` | `/api/v1/services/embeddings/multimodal-embedding/multimodal-embedding` | Text and multimodal inputs are emitted as DashScope `input.contents`. Supports `text`, `image`, `video`, `multi_images`, `parameters.enable_fusion`, `parameters.res_level`, and `parameters.max_video_frames`. Alias: `dashscope:multimodal_embedding`. |
Custom provider endpoint paths are available when the endpoint is configured for an embedding API format. Gemini custom paths can use `{model}` and `{action}`. For `gemini:embedding`, `{action}` expands to `embedContent`.
## Model And Catalog Requirements
To use embeddings through the gateway:
1. The global model should include embedding metadata, for example `supported_capabilities: ["embedding"]`, `config.model_type: "embedding"`, or `config.api_formats` with one of the embedding formats.
2. The provider model or mapping must expose an embedding API format, one of `openai:embedding`, `gemini:embedding`, `jina:embedding`, `doubao:embedding`, or `aliyun:multimodal_embedding`.
3. The user and API key must be allowed to access the model and the `openai:embedding` client API format.
4. Public and admin catalog responses expose `supports_embedding` so clients can display embedding capability separately from chat.
Billing fails closed for embedding global models. A model marked as embedding capable must define either `default_price_per_request` or `default_tiered_pricing.tiers[].input_price_per_1m`. Missing request pricing and missing input token pricing cause the model record to be rejected instead of treated as free.
No schema migration is needed for embedding metadata. Existing model capability, config, provider mapping, API format, and pricing fields carry the data.
## Aliyun Qwen3-VL Examples
Text request through Aether:
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-vl-embedding",
"input": "white running shoes",
"dimensions": 1024
}'
```
Image and text fusion request:
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-vl-embedding",
"input": [
{ "text": "white running shoes, lightweight and breathable" },
{ "image": "https://dashscope.oss-cn-beijing.aliyuncs.com/images/256_1.png" }
],
"parameters": { "enable_fusion": true }
}'
```
Video request:
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-vl-embedding",
"input": [
{ "video": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20250107/lbcemt/new+video.mp4" }
],
"parameters": { "max_video_frames": 64 }
}'
```
Multi-image fusion request:
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-vl-embedding",
"input": [
{ "text": "product photos from multiple angles" },
{ "multi_images": [
"https://example.com/front.png",
"https://example.com/side.png"
] }
],
"parameters": { "enable_fusion": true }
}'
```
## Failure Behavior
The gateway validates deterministic request errors before local execution or provider transport.
| Case | Example request body or setup | Status | Error detail |
| --- | --- | --- | --- |
| Invalid JSON | `{` | `400` | `Embedding request JSON body is invalid` |
| Missing model | `{ "input": "hello" }` | `400` | `Embedding request model is required` |
| Empty input | `{ "model": "text-embedding-3-small", "input": [] }` | `400` | `Embedding request input is required` |
| Chat `messages` payload | `{ "model": "text-embedding-3-small", "messages": [] }` | `400` | `Embedding request must use input, not chat messages` |
| Streaming requested | `{ "model": "text-embedding-3-small", "input": "hello", "stream": true }` | `400` | `Embedding requests do not support streaming` |
| Non JSON content type | `Content-Type: text/plain` with an embedding JSON body | `400` | `Embedding request content-type must be application/json` |
| Chat only model | API key allows `text-embedding-3-small`, request uses `gpt-5` | `403` | The key is not allowed to access that model. |
| Chat only API format | API key allows `openai:chat` but not `openai:embedding` | `403` | The key is not allowed to access `openai:embedding`. |
Failure examples:
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: application/json" \
-d '{"model":"text-embedding-3-small","messages":[]}'
```
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: application/json" \
-d '{"input":"hello"}'
```
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: application/json" \
-d '{"model":"text-embedding-3-small","input":[]}'
```
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: application/json" \
-d '{"model":"text-embedding-3-small","input":"hello","stream":true}'
```
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: text/plain" \
-d '{"model":"text-embedding-3-small","input":"hello"}'
```
If a valid embedding request passes local validation but no usable provider transport is available, the gateway can return a provider or service availability error. That is different from the deterministic request validation errors above.
-238
View File
@@ -1,238 +0,0 @@
# Format Conversion Audit
Last audited: 2026-06-03
This audit tracks `source format -> Canonical -> target format` behavior. It is intentionally stricter than historical best-effort conversion.
Statuses:
- `native`: emitted as a target-native field without semantic change.
- `mapped`: converted through canonical/provider-specific mapping.
- `extension-preserved`: preserved in same-format canonical roundtrip or target-approved extension namespace.
- `transport-only`: audited client transport metadata intentionally omitted when the target has no compatible transport channel.
- `unaudited`: rejected because the source field is not in the audited provider schema inventory for cross-format conversion.
- `unsupported`: rejected because the request/field shape is outside the supported conversion surface, independent of schema drift.
- `lossy-blocked`: conversion fails closed.
- `invalid-enum`: conversion fails closed because a provider enum value is not valid for the target mapping.
Full schema field coverage is tracked in `docs/api/format-field-coverage-matrix.md`. That matrix is generated from the schema inventory in `docs/api/provider-interface-definitions.md` by `python3 docs/api/generate_format_field_coverage.py` and gives every documented OpenAI, Claude, and Gemini schema field a handling status. “Handled” means mapped, same-format/native preserved, extension-preserved, blocked with a structured error, or explicitly marked outside the canonical conversion surface.
Provider schema refresh is not a runtime dependency. Same-format runtime paths do not use this matrix; they bypass canonical conversion. Same-format canonical roundtrip must preserve unrecognized provider fields through provider extension namespaces. Cross-format conversion is capability-based: only explicitly mapped fields are emitted, and newly discovered or unknown provider fields fail closed with `UnauditedField` until a lossless mapping is audited.
## Implemented Boundary Changes
| Area | Current behavior |
| --- | --- |
| Pure conversion API | `convert_request_pure` and response equivalents do not apply model override or stream policy. |
| Legacy conversion API | `convert_request` / `convert_response` are retained for migration and may still use legacy context behavior. |
| Same-format provider path | Bypasses canonical conversion and copies the parsed JSON object before transport edits. |
| Cross-format same-format-provider path | Uses `convert_request_pure`, then applies model/body/stream edits in transport. |
| Conversion errors | Added `UnauditedField`, `UnsupportedField`, `InvalidEnumValue`, `LossyConversionBlocked`, and `InvalidTargetField`. |
| Reporting | Added `ConversionReport` with field statuses. Runtime reports remain conversion-operation oriented; exhaustive nested schema coverage is enforced by `format-field-coverage-matrix.md`. |
| Source schema coverage | Cross-format request conversion rejects unknown source root fields before emit. Every documented schema field is covered by the field coverage matrix. |
| Schema drift handling | Official schema changes are detected by regenerating the inventory/matrix. Runtime same-format remains passthrough; cross-format unknowns return `UnauditedField` until deliberately mapped. |
| Tool schema roundtrip | Claude `input_schema` and Gemini `functionDeclarations.parameters` preserve raw same-format schema through provider-specific extensions. |
| Tool result ids | Chat `tool_call_id`, Responses `call_id`, Claude `tool_use_id`, and Gemini `functionResponse.id` are mapped through canonical tool IDs. |
## OpenAI Chat -> OpenAI Responses
| Chat field | Canonical handling | Responses output | Status |
| --- | --- | --- | --- |
| `model` | request identity | `model` | native |
| `messages` | canonical messages/instructions | `input`, `instructions` | mapped |
| `max_tokens` | generation max tokens | `max_output_tokens` | mapped |
| `max_completion_tokens` | generation max tokens | `max_output_tokens` | mapped |
| `temperature` | generation | `temperature` | native |
| `top_p` | generation | `top_p` | native |
| `top_logprobs` | generation | `top_logprobs` | native |
| `n` | generation but no Responses equivalent | none | lossy-blocked |
| `stop` | generation but no Responses equivalent | none | lossy-blocked |
| `presence_penalty` | generation but no Responses equivalent | none | lossy-blocked |
| `frequency_penalty` | generation but no Responses equivalent | none | lossy-blocked |
| `seed` | generation but no Responses equivalent | none | lossy-blocked |
| `logprobs` | generation but no Responses equivalent | none | lossy-blocked |
| `stream` | OpenAI extension | `stream` | mapped if explicit |
| `stream_options` | Chat-specific extension | none | lossy-blocked |
| `tools[].function.name` | canonical tool | `tools[].name` | mapped |
| `tools[].function.description` | canonical tool | `tools[].description` | mapped |
| `tools[].function.parameters` | canonical tool | `tools[].parameters` | mapped |
| `tools[].function.strict` | canonical tool strict | `tools[].strict` | mapped, implemented |
| assistant `tool_calls[].id` | canonical tool use id | `function_call.call_id` | mapped, implemented |
| tool message `tool_call_id` | canonical tool result id | `function_call_output.call_id` | mapped, implemented |
| `tool_choice` | canonical tool choice | `tool_choice` | mapped |
| `parallel_tool_calls` | canonical bool | `parallel_tool_calls` | native |
| `metadata` | canonical metadata | `metadata` | native |
| `response_format` | canonical response format | `text.format` | mapped |
| `reasoning_effort` | OpenAI enum | `reasoning.effort` | mapped; invalid enum blocked |
| `verbosity` | OpenAI extension | `text.verbosity` | mapped |
| `store` | OpenAI extension | `store` | extension-preserved |
| `service_tier` | OpenAI extension | `service_tier` | extension-preserved |
| `safety_identifier` | OpenAI extension | `safety_identifier` | extension-preserved |
| `prompt_cache_key` | OpenAI extension | `prompt_cache_key` | extension-preserved |
| `user` | legacy Chat user field | none | lossy-blocked |
| unknown top-level fields | source schema guard | none | unaudited |
## OpenAI Responses -> OpenAI Chat
| Responses field | Canonical handling | Chat output | Status |
| --- | --- | --- | --- |
| `model` | request identity | `model` | native |
| `input` | canonical messages/content/tool I/O | `messages` | mapped |
| `instructions` | canonical instruction/system | `messages` system/developer | mapped |
| `max_output_tokens` | generation max tokens | `max_completion_tokens` | mapped |
| `temperature` | generation | `temperature` | native |
| `top_p` | generation | `top_p` | native |
| `top_logprobs` | generation | `top_logprobs` | native |
| `metadata` | canonical metadata | `metadata` | native |
| `client_metadata` | Responses client transport metadata | none | transport-only; omitted |
| `parallel_tool_calls` | canonical bool | `parallel_tool_calls` | native |
| `text.format` | canonical response format | `response_format` | mapped |
| `text.verbosity` | Responses extension | `verbosity` | mapped |
| `tools[].type=function` | canonical tool | `tools[].type=function` | mapped |
| `tools[].name` | canonical tool | `tools[].function.name` | mapped |
| `tools[].parameters` | canonical tool | `tools[].function.parameters` | mapped |
| `tools[].strict` | canonical tool strict | `tools[].function.strict` | mapped, implemented |
| `function_call.call_id` | canonical tool use id | `tool_calls[].id` | mapped, implemented |
| `function_call_output.call_id` | canonical tool result id | tool message `tool_call_id` | mapped, implemented |
| `tools[].type=custom` | raw Responses tool | none | lossy-blocked to Chat |
| `tools[].type=web_search*` | raw Responses tool | none | lossy-blocked to Chat |
| `tool_choice` | canonical tool choice | `tool_choice` | mapped |
| `reasoning.effort` | OpenAI enum | `reasoning_effort` | mapped; invalid enum blocked |
| `reasoning.summary` | Responses-only | none | lossy-blocked |
| `reasoning.budget_tokens` | Responses-only | none | lossy-blocked |
| `stream` | Responses request transport policy | none | lossy-blocked; target stream policy is transport-owned |
| `include` | Responses-only | none | lossy-blocked; legacy emitter no longer leaks |
| `previous_response_id` | Responses-only | none | lossy-blocked; legacy emitter no longer leaks |
| `truncation` | Responses-only | none | lossy-blocked |
| `prompt` | Responses-only | none | lossy-blocked |
| `conversation` | Responses-only | none | lossy-blocked |
| `background` | Responses-only | none | lossy-blocked |
| `max_tool_calls` | Responses-only | none | lossy-blocked |
| unknown top-level fields | source schema guard | none | unaudited |
## Claude Messages <-> OpenAI Chat / Responses
Claude to OpenAI Chat, Claude to OpenAI Responses, and the reverse directions are included in the field coverage matrix. Runtime strict guards cover request root fields, provider extension namespaces, thinking/cache/tool-result hazards, and target generation-field gaps. Fields without a lossless target equivalent fail closed instead of being dropped.
High-risk fields:
| Claude field | OpenAI target risk | Required status |
| --- | --- | --- |
| `system` with cache blocks | Chat/Responses system instructions | same-format preserved; cross-format `cache_control` loss is blocked |
| `thinking` | OpenAI reasoning | Claude request-level thinking config maps to OpenAI reasoning; message-level thinking blocks are blocked for Responses |
| `cache_control` | OpenAI content/tool extensions | same-format preserved; cross-format blocked when no target equivalent exists |
| `tools[].input_schema` | OpenAI tool parameters | mapped; raw same-format schema preservation implemented |
| `tool_choice.disable_parallel_tool_use` | OpenAI `parallel_tool_calls` | mapped, implemented |
| `tool_result` multi-block content | OpenAI tool output/content | same-format preserved; cross-format to Chat/Responses is lossy-blocked |
| `metadata` | OpenAI metadata | mapped when the target has metadata |
| `container`, `inference_geo`, `service_tier` | OpenAI target has no audited equivalent | lossy-blocked unless a target-approved mapping is added |
## Gemini GenerateContent <-> OpenAI Chat / Responses / Claude
Gemini to OpenAI Chat, Gemini to OpenAI Responses, Gemini to Claude, and reverse generation paths are included in the field coverage matrix. Gemini-only request fields are preserved same-format and blocked cross-format unless the target mapping is explicitly audited.
High-risk fields:
| Gemini field | Target risk | Required status |
| --- | --- | --- |
| `contents[].parts[].thoughtSignature` | OpenAI/Claude thinking | Chat/Claude preserve; Responses cross-format is lossy-blocked |
| `tools[].functionDeclarations` | OpenAI/Claude tool schema | mapped; raw same-format `parameters` preservation implemented |
| `toolConfig.functionCallingConfig.allowedFunctionNames` | OpenAI/Claude tool choice | single-name mapping implemented; multi-name input is lossy-blocked |
| `toolConfig.functionCallingConfig.mode` | OpenAI/Claude tool choice enum | valid enum required; invalid values fail with `InvalidEnumValue` |
| `generationConfig.thinkingConfig.thinkingLevel` | OpenAI/Claude reasoning effort | low/medium/high mapping implemented; invalid values fail closed |
| `safetySettings` | OpenAI/Claude no direct equivalent | lossy-blocked |
| `cachedContent` | OpenAI/Claude no direct equivalent | lossy-blocked |
| `codeExecution` | OpenAI/Claude tool/builtin mismatch | lossy-blocked |
| `generationConfig.responseModalities` | OpenAI/Claude modality mismatch | lossy-blocked |
| `functionResponse.id` | tool result id | conversion preserves id; Gemini upstream cleanup is transport-layer edit only |
## Embedding And Rerank
Embedding and rerank request parse/emit capability and strict target guards are implemented. Provider schema fields outside these canonical conversion surfaces are marked `not-in-conversion-surface` in the field coverage matrix instead of being left implicit.
Embedding source capability:
| Source format | Parsed request shape | Canonical fields | Status |
| --- | --- | --- | --- |
| OpenAI Embedding | `model`, `input`, `encoding_format`, `dimensions`, `user`, `parameters`, `task` | OpenAI-like embedding | mapped |
| Jina Embedding | OpenAI-like plus provider extension namespace | OpenAI-like embedding | mapped |
| Doubao Embedding | OpenAI-like `model` + text `input` | OpenAI-like embedding | mapped |
| Gemini Embedding | single `content.parts[].text` or batch `requests[]` | text input, `dimensions`, `task` | mapped |
| Aliyun Multimodal Embedding | `input.contents[]`, `parameters.dimension` | text/multimodal input, `dimensions`, `parameters` | mapped |
Embedding target guards:
| Target format | Accepted canonical fields | Blocked fields/cases | Status |
| --- | --- | --- | --- |
| OpenAI Embedding | text or token input, `encoding_format`, `dimensions`, `user` | multimodal input, `task`, generic `parameters` | lossy-blocked |
| Jina Embedding | text input, `dimensions`, `task`, `parameters` | token/multimodal input, `encoding_format`, `user` | lossy-blocked |
| Gemini Embedding | text input, `dimensions`, valid `taskType` | token/multimodal input, `encoding_format`, `user`, generic `parameters`, invalid `taskType` | lossy-blocked / invalid-enum |
| Doubao Embedding | text input, `dimensions` | token/multimodal input, `encoding_format`, `user`, `task`, generic `parameters` | lossy-blocked |
| Aliyun Multimodal Embedding | text or multimodal input, `dimensions`, `parameters` | token input, `encoding_format`, `user`, `task` | lossy-blocked |
Cross-format embedding invariants:
- Embedding formats can only convert to embedding formats.
- Unknown provider-specific embedding extension namespaces are blocked cross-format unless the namespace matches the target.
- Aliyun `parameters.dimension` maps to canonical `dimensions` and is not treated as generic `parameters`.
- Gemini batch embedding parse requires every batch item to share the same model, dimensions, and task.
Rerank first pass:
| Area | Current behavior | Status |
| --- | --- | --- |
| Source formats | OpenAI Rerank and Jina Rerank parse OpenAI-like `model`, `query`, `documents`, `top_n`, `return_documents` | mapped |
| Target formats | OpenAI Rerank and Jina Rerank emit OpenAI-like rerank bodies | mapped |
| Boundary | Rerank formats can only convert to rerank formats | lossy-blocked |
| Validation | Empty query/documents and `top_n=0` fail closed | invalid-target-field |
| Extensions | Unknown provider-specific rerank extension namespaces are blocked cross-format | unsupported |
## Sync Response Conversion
Cross-format sync response conversion now validates source stop/finish/status
enums before emitting a target body. Same-format runtime response passthrough is
still outside canonical conversion.
| Source field | Target risk | Current behavior | Status |
| --- | --- | --- | --- |
| Same-format response raw stop/status fields | canonical emitters would otherwise normalize unknown enum/status to default target stop values | raw OpenAI Chat `finish_reason`, OpenAI Responses `status`, Claude `stop_reason`/`stop_sequence`, and Gemini `finishReason` are preserved through provider extension metadata | extension-preserved |
| OpenAI Chat `choices[].finish_reason` | unknown value would otherwise emit as target normal stop | valid Chat enum required; unknown values fail with `InvalidEnumValue` | invalid-enum |
| OpenAI Responses `status` | `queued`, `in_progress`, and `cancelled` have no sync target equivalent | non-terminal valid states fail with `LossyConversionBlocked`; invalid states fail with `InvalidEnumValue` | lossy-blocked / invalid-enum |
| OpenAI Responses `incomplete_details.reason=content_filter` | previously mapped to max tokens/`length` | maps to canonical content filter and emits Chat `content_filter` / Claude `content_filtered` / Gemini `SAFETY` | mapped |
| Claude `stop_reason` | unknown value would otherwise emit as target normal stop | valid known stop enum required for cross-format conversion | invalid-enum |
| Gemini `candidates[].finishReason` | known-but-unmappable reasons would otherwise emit as target normal stop | mappable safety/max/stop reasons convert; known unmappable values such as `OTHER`, `MALFORMED_FUNCTION_CALL`, `UNEXPECTED_TOOL_CALL`, `MISSING_THOUGHT_SIGNATURE`, and `MALFORMED_RESPONSE` fail with `LossyConversionBlocked`; future unknown values fail with `InvalidEnumValue` | lossy-blocked / invalid-enum |
| Canonical `Unknown` stop reason | target emitters default to normal stop values | cross-format response conversion blocks canonical unknown stop reasons | lossy-blocked |
## Stream Conversion
Sixth batch first pass is implemented for unknown event handling and runtime
same-format boundaries. Sync response finish/status parity has a first strict
pass; stream finish-reason guardrails are implemented for unknown/unmappable
terminal reasons. Stream event schema fields are covered in the field coverage
matrix; provider-by-provider fixtures cover the runtime event behavior.
Current stream behavior:
| Area | Current behavior | Status |
| --- | --- | --- |
| Provider parsers | OpenAI Chat, OpenAI Responses, Claude, and Gemini unknown stream payloads become `CanonicalStreamEvent::UnknownEvent` | mapped |
| Cross-format stream matrix | Unknown canonical stream events emit a target-format error SSE with `unsupported_stream_event` and terminate conversion | lossy-blocked |
| Stream finish reason guard | Unknown OpenAI finish reasons, unknown Claude `stop_reason`, and Gemini known-but-unmappable `finishReason` values such as `OTHER` are preserved as raw canonical finish strings, then blocked by the matrix with `unsupported_finish_reason` | lossy-blocked |
| OpenAI Responses stream target | Canonical `length` and `content_filter` terminal reasons emit `response.incomplete` with `incomplete_details.reason=max_output_tokens` or `content_filter` instead of `response.completed` | mapped |
| Terminal observer | Unknown provider stream events increment `unknown_event_count`; OpenAI Responses failed events mark terminal error state | mapped |
| Stream -> sync aggregate | Unknown OpenAI Chat, OpenAI Responses, Claude, and Gemini stream events make the runtime finalize checked path return an error and block `body_json` fallback; legacy public aggregate helpers keep `Option` compatibility | lossy-blocked |
| Runtime strict fallback guard | `UnauditedField`, `InvalidEnumValue`, `UnsupportedField`, `LossyConversionBlocked`, and `InvalidTargetField` from registry response conversion are not allowed to fall through legacy conversion helpers | lossy-blocked |
| Runtime same-format stream | Same-format stream passthrough remains outside canonical conversion; stream policy edits are transport-layer only | native |
Stream fixture coverage:
| Provider stream | Covered fixture areas |
| --- | --- |
| OpenAI Chat | sync aggregation for text, tool call IDs/names/argument deltas, finish reason, and usage; cross-format unknown finish/event blocking |
| OpenAI Responses | text snapshot de-duplication, multi-part messages, reasoning/items, function calls, image generation calls, same-family stream sync, unknown event blocking |
| Claude Messages | thinking signatures, tool input deltas, cache/usage aggregation, media emission, unknown stop/event blocking |
| Gemini GenerateContent | text/media/signature aggregation, function calls/results, safety finish mapping, unknown parts/events, and unmappable finish reason blocking |
Matrix-level interception remains the authoritative runtime path for cross-format
unknown events; direct client emitters are covered only as provider/client
building blocks.
-145
View File
@@ -1,145 +0,0 @@
# Format Enum Mapping
Status values used below:
- `native`: same semantic value exists in the target format.
- `mapped`: explicit provider-specific mapping is required.
- `blocked`: no lossless target value; conversion must fail closed.
- `preserve-same-format`: unknown/raw values are preserved only when source and target format are the same.
## OpenAI Reasoning Effort
Provider-specific types: `OpenAiChatReasoningEffort` for Chat `reasoning_effort`, and `OpenAiResponsesReasoningEffort` for Responses `reasoning.effort`. They are intentionally separate even when their current value sets overlap; a value accepted by one field is not treated as valid for the other unless that field's own enum accepts it.
| Source field | Source value | Target field | Target value | Status |
| --- | --- | --- | --- | --- |
| Chat `reasoning_effort` | `none` | Responses `reasoning.effort` | `none` | native |
| Chat `reasoning_effort` | `minimal` | Responses `reasoning.effort` | `minimal` | native when the target model supports it; blocked for GPT-5.6 |
| Chat `reasoning_effort` | `low` | Responses `reasoning.effort` | `low` | native |
| Chat `reasoning_effort` | `medium` | Responses `reasoning.effort` | `medium` | native |
| Chat `reasoning_effort` | `high` | Responses `reasoning.effort` | `high` | native |
| Chat `reasoning_effort` | `xhigh` | Responses `reasoning.effort` | `xhigh` | native |
| Chat `reasoning_effort` | `max` | Responses `reasoning.effort` | `max` | native for GPT-5.6; blocked for models that do not publish `max` |
| Responses `reasoning.effort` | `none` | Chat `reasoning_effort` | `none` | native |
| Responses `reasoning.effort` | `minimal` | Chat `reasoning_effort` | `minimal` | native when the target model supports it; blocked for GPT-5.6 |
| Responses `reasoning.effort` | `low` | Chat `reasoning_effort` | `low` | native |
| Responses `reasoning.effort` | `medium` | Chat `reasoning_effort` | `medium` | native |
| Responses `reasoning.effort` | `high` | Chat `reasoning_effort` | `high` | native |
| Responses `reasoning.effort` | `xhigh` | Chat `reasoning_effort` | `xhigh` | native |
| Responses `reasoning.effort` | `max` | Chat `reasoning_effort` | `max` | native for GPT-5.6; blocked for models that do not publish `max` |
| Responses `reasoning.summary` | any | Chat | none | blocked |
| Responses `reasoning.budget_tokens` | any | Chat | none | blocked |
Internal model directive values:
| Internal value | OpenAI Chat | OpenAI Responses | Claude output effort | Gemini thinking level | Notes |
| --- | --- | --- | --- | --- | --- |
| `none` | `none` | `none` | `low` | `low` | Budget maps to `0`. |
| `minimal` | `minimal` | `minimal` | `low` | `low` | Budget maps to `512`. |
| `low` | `low` | `low` | `low` | `low` | Budget maps to `1280`. |
| `medium` | `medium` | `medium` | `medium` | `medium` | Budget maps to `2048`. |
| `high` | `high` | `high` | `high` | `high` | Budget maps to `4096`. |
| `xhigh` | `xhigh` | `xhigh` | `xhigh` | `high` | Budget maps to `8192`. |
| `max` | `max` for GPT-5.6 | `max` for GPT-5.6 | `max` | `high` | OpenAI emission is capability-gated by the resolved model. |
GPT-5.6 (`gpt-5.6`, `gpt-5.6-sol`, `gpt-5.6-terra`, and `gpt-5.6-luna`) publishes `none`, `low`, `medium`, `high`, `xhigh`, and `max`, and does not support `minimal`. Additional non-empty effort values advertised by a model are preserved verbatim across OpenAI Chat and Responses conversion. `ultra` is a Codex client preset that resolves to `max` before transmission and is not an OpenAI wire effort. Known effort capabilities are validated against the resolved provider model, so aliases mapped to GPT-5.6 receive the GPT-5.6 contract while concrete model families keep their published constraints.
## Tool Choice
| Canonical | OpenAI Chat | OpenAI Responses | Claude Messages | Gemini GenerateContent |
| --- | --- | --- | --- | --- |
| auto | `"auto"` | `"auto"` | `{"type":"auto"}` | unset / function calling config auto |
| none | `"none"` | `"none"` | `{"type":"none"}` | mode none |
| required | `"required"` | `"required"` | `{"type":"any"}` | mode any |
| named function | `{"type":"function","function":{"name":...}}` | `{"type":"function","name":...}` | `{"type":"tool","name":...}` | allowed function name |
Implemented guardrails:
- Claude `output_config.effort=max` maps to OpenAI `xhigh`; unknown Claude effort enums are blocked cross-format.
- Claude `tool_choice.disable_parallel_tool_use` maps inversely to OpenAI `parallel_tool_calls`.
- Gemini `allowedFunctionNames` maps to canonical named tool choice and emits back to `allowedFunctionNames`.
- Gemini `thinkingLevel` maps `low|medium|high` to OpenAI reasoning effort `low|medium|high`; unknown values are blocked cross-format.
- Responses `custom`, `web_search*`, and other built-in tools are blocked when converting to Chat unless a target raw passthrough is explicitly added.
## Tool Definition Kind
| Source | Target | Mapping |
| --- | --- | --- |
| OpenAI Chat `tools[].function.name` | Responses `tools[].name` | mapped |
| OpenAI Chat `tools[].function.parameters` | Responses `tools[].parameters` | mapped |
| OpenAI Chat `tools[].function.strict` | Responses `tools[].strict` | mapped, implemented |
| Responses `tools[].strict` | Chat `tools[].function.strict` | mapped, implemented |
| OpenAI Chat assistant `tool_calls[].id` | Responses `function_call.call_id` | mapped, implemented |
| OpenAI Chat tool `tool_call_id` | Responses `function_call_output.call_id` | mapped, implemented |
| Responses `function_call.call_id` | Chat `tool_calls[].id` | mapped, implemented |
| Responses `function_call_output.call_id` | Chat tool `tool_call_id` | mapped, implemented |
| Gemini `functionCall.id` | OpenAI/Claude canonical tool use id | mapped, implemented |
| Gemini `functionResponse.id` | OpenAI `tool_call_id` / Responses `call_id` / Claude `tool_use_id` | mapped, implemented |
| Claude `tools[].input_schema` | OpenAI `parameters` | mapped; raw same-format schema preserved |
| Gemini `functionDeclarations[].parameters` | OpenAI `parameters` | mapped; raw same-format schema preserved |
## Roles
| Canonical role | OpenAI Chat | OpenAI Responses | Claude Messages | Gemini |
| --- | --- | --- | --- | --- |
| system | `system` | `instructions` or system input item | top-level `system` | `systemInstruction` |
| developer | `developer` | `instructions` or developer input item | extension-preserved | systemInstruction extension |
| user | `user` | `message.role=user` | `user` | `user` |
| assistant | `assistant` | `message.role=assistant` / output item | `assistant` | `model` |
| tool | `tool` | `function_call_output` | `tool_result` inside user message | `functionResponse` |
Known lossy risks:
- Multiple system/developer instruction ordering needs full golden fixtures.
- Provider-specific role extensions must be preserved in same-format roundtrip and blocked cross-format if no target equivalent exists.
## Finish Reasons
| Canonical | OpenAI Chat | OpenAI Responses | Claude | Gemini |
| --- | --- | --- | --- | --- |
| stop | `stop` | completed output | `end_turn` | `STOP` |
| length | `length` | `status=incomplete`, `incomplete_details.reason=max_output_tokens` | `max_tokens` | `MAX_TOKENS` |
| tool calls | `tool_calls` | output contains `function_call` | `tool_use` | `functionCall` part, usually with `STOP` |
| content filter/safety | `content_filter` | `status=incomplete`, `incomplete_details.reason=content_filter` | `content_filtered` or refusal-compatible stops | `SAFETY`, `RECITATION`, `LANGUAGE`, `BLOCKLIST`, `PROHIBITED_CONTENT`, `SPII`, image safety/recitation stops |
| unknown | preserve-same-format | preserve-same-format | preserve-same-format | preserve-same-format |
Implemented response guardrails:
- Cross-format sync response conversion validates source finish/status enums before emitting the target response.
- Same-format canonical response roundtrip preserves raw OpenAI Chat `finish_reason`, OpenAI Responses `status`, Claude `stop_reason`, and Gemini `finishReason` values through provider extension metadata.
- OpenAI Chat unknown `choices[].finish_reason` fails with `InvalidEnumValue`.
- OpenAI Responses non-terminal `status` values (`queued`, `in_progress`, `cancelled`) are valid provider states but are blocked for sync response conversion because target sync formats cannot represent them losslessly.
- Runtime sync finalize does not fall back to legacy conversion when registry response conversion reports strict errors such as invalid enums, unsupported fields, lossy blocks, or invalid target fields.
- Stream terminal reasons now follow the same strict policy: unknown OpenAI / Claude raw finish reasons and Gemini known-but-unmappable values such as `OTHER` surface as `unsupported_finish_reason`, while OpenAI Responses `length` and `content_filter` stream finals emit `response.incomplete`.
- Gemini known but unmappable finish reasons such as `OTHER`, `MALFORMED_FUNCTION_CALL`, `UNEXPECTED_TOOL_CALL`, `MISSING_THOUGHT_SIGNATURE`, and `MALFORMED_RESPONSE` are blocked with `LossyConversionBlocked`; unknown future Gemini values fail with `InvalidEnumValue`.
- Stream finish reason guardrails are covered with provider-specific fixtures for usage, tool calls, reasoning signatures, media, and unknown payloads. Unknown provider stream events are not mapped as finish reasons; cross-format runtime conversion emits a target-format `unsupported_stream_event` error and terminates.
## Embedding Task Types
Provider-specific type: `GeminiEmbeddingTaskType`.
Gemini embedding task values are stored in canonical `embedding.task` only after
source parsing. They are emitted to Gemini as `taskType` and validated before
cross-format conversion to a Gemini target.
| Canonical task input | Gemini `taskType` output | Status |
| --- | --- | --- |
| `QUERY` | `RETRIEVAL_QUERY` | mapped alias |
| `RETRIEVAL_QUERY` | `RETRIEVAL_QUERY` | native |
| `DOCUMENT` | `RETRIEVAL_DOCUMENT` | mapped alias |
| `RETRIEVAL_DOCUMENT` | `RETRIEVAL_DOCUMENT` | native |
| `TEXT_MATCHING` | `SEMANTIC_SIMILARITY` | mapped alias |
| `SEMANTIC_SIMILARITY` | `SEMANTIC_SIMILARITY` | native |
| `CLASSIFICATION` | `CLASSIFICATION` | native |
| `CLUSTERING` | `CLUSTERING` | native |
| `QUESTION_ANSWERING` | `QUESTION_ANSWERING` | native |
| `FACT_VERIFICATION` | `FACT_VERIFICATION` | native |
| `CODE_RETRIEVAL_QUERY` | `CODE_RETRIEVAL_QUERY` | native |
| `TASK_TYPE_UNSPECIFIED` | `TASK_TYPE_UNSPECIFIED` | native |
| unknown value | none | blocked with `InvalidEnumValue` when targeting Gemini |
Cross-provider rules:
- Gemini `taskType` to OpenAI Embedding is blocked because OpenAI has no equivalent task field.
- Gemini `taskType` to Doubao/Aliyun is blocked for the same reason.
- Jina `task` may carry through canonical and emit as Jina `task`; when targeting Gemini it must match the valid Gemini task set above.
File diff suppressed because it is too large Load Diff
-85
View File
@@ -1,85 +0,0 @@
# Format Passthrough Contract
Last audited: 2026-06-03
This document defines the boundary between runtime passthrough, canonical roundtrip tests, and cross-format conversion.
## Runtime Same-Format Path
Runtime same-format provider paths must not call canonical conversion.
Current implementation:
- `crates/aether-provider/transport/src/same_format_provider/mod.rs` checks `api_format_alias_matches(client_api_format, provider_api_format)`.
- When formats match, the provider body is built by copying the parsed JSON object field-for-field.
- When formats differ, the provider body is built through `aether_ai_formats::convert_request_pure`.
- Model override, body rules, model directives, Claude Code sanitization, Gemini function-response id stripping, and stream policy are applied only after the passthrough/conversion branch in provider transport.
Important limitation:
- The current transport helper receives `body_json: &serde_json::Value`, not raw request bytes. It therefore guarantees no canonical conversion and JSON value preservation at this layer, but it cannot preserve original whitespace or object key order by itself.
- True byte-level passthrough for requests with no transport edits requires a higher-level raw-body path that can forward the original bytes directly. Until that raw-body plumbing exists, tests should assert "conversion module not called" and JSON value equivalence for this helper, not byte-for-byte serialization equivalence.
Provider schema drift does not change this rule. If OpenAI, Claude, or Gemini add a new field, same-format runtime routing must still forward it as part of the original provider body. The schema inventory and field coverage matrix are audit aids, not the runtime allowlist for same-format traffic.
## Canonical Same-Format Roundtrip
Canonical same-format roundtrip is only a test/audit mode:
```text
source format -> Canonical -> same source format
```
Required behavior:
- JSON-normalized equality, ignoring object field order and whitespace.
- Field values, array order, unknown fields, extension namespaces, and unknown enum strings must be preserved.
- This path may parse and emit; it is not the runtime path.
- Unknown provider fields are carried in provider extension namespaces and replayed when emitting the same provider format.
## Cross-Format Conversion
Cross-format conversion is strict:
```text
source format -> Canonical -> target format
```
Required behavior:
- Emit only fields valid for the target provider format.
- Map provider-specific enum values through explicit provider enum types.
- Preserve source fields only when the target has an equivalent field or documented extension passthrough.
- Fail closed with `FormatError::UnauditedField`, `FormatError::LossyConversionBlocked`, `FormatError::UnsupportedField`, `FormatError::InvalidEnumValue`, or `FormatError::InvalidTargetField` when no lossless mapping exists.
- Do not use `None` or silent omission to represent conversion failure.
- Newly added provider fields follow the same rule as other unknown fields: preserve same-format, fail closed cross-format with `UnauditedField`. A code change is required only when Aether intentionally supports a new cross-format semantic mapping.
## Pure Conversion Interface
Pure conversion lives in `crates/aether-ai/formats` and is limited to:
- parse
- emit
- provider-specific field/enum mapping
- `ConversionReport`
Pure conversion must not:
- override `model`
- add, remove, or force `stream`
- apply body rules
- apply model directives
- read the original request body to patch missing target fields
- perform provider transport policy edits
Current pure entrypoints:
- `parse_request_pure`
- `emit_request_pure`
- `convert_request_pure`
- `convert_request_pure_with_context`
- `parse_response_pure`
- `emit_response_pure`
- `convert_response_pure`
`convert_request` and `convert_response` remain legacy wrappers for existing callers that still need mapped model/report-context behavior during migration.
-548
View File
@@ -1,548 +0,0 @@
#!/usr/bin/env python3
"""Generate the provider schema field coverage matrix.
The input inventory is docs/api/provider-interface-definitions.md. Existing
coverage rows are reused so audited status/notes survive regeneration. Newly
introduced provider fields get conservative same-format/native and cross-format
fail-closed defaults until a human audits whether they deserve an explicit
mapping.
"""
from __future__ import annotations
import argparse
import dataclasses
from collections import Counter, defaultdict
from pathlib import Path
from typing import Iterable
ROOT = Path(__file__).resolve().parents[2]
DEFAULT_DEFINITIONS = ROOT / "docs/api/provider-interface-definitions.md"
DEFAULT_MATRIX = ROOT / "docs/api/format-field-coverage-matrix.md"
@dataclasses.dataclass(frozen=True)
class SourceField:
provider: str
schema: str
field: str
required: str
field_type: str
@dataclasses.dataclass(frozen=True)
class CoverageStatus:
surface: str
same_format_runtime: str
canonical_roundtrip: str
cross_format: str
notes: str
OPENAI_CHAT_MAPPED = {
"model",
"messages",
"max_tokens",
"max_completion_tokens",
"temperature",
"top_p",
"top_logprobs",
"tools",
"tool_choice",
"parallel_tool_calls",
"metadata",
"response_format",
"reasoning_effort",
"verbosity",
"store",
"service_tier",
"safety_identifier",
"prompt_cache_key",
"prompt_cache_retention",
"stream",
}
OPENAI_CHAT_BLOCKED = {
"n",
"stop",
"presence_penalty",
"frequency_penalty",
"seed",
"logprobs",
"stream_options",
"user",
"function_call",
"functions",
"logit_bias",
"modalities",
"prediction",
"audio",
"web_search_options",
}
OPENAI_RESPONSES_MAPPED = {
"model",
"input",
"instructions",
"max_output_tokens",
"temperature",
"top_p",
"top_logprobs",
"metadata",
"parallel_tool_calls",
"text",
"tools",
"tool_choice",
"reasoning",
"store",
"service_tier",
"safety_identifier",
"prompt_cache_key",
"prompt_cache_retention",
}
OPENAI_RESPONSES_BLOCKED = {
"include",
"previous_response_id",
"truncation",
"prompt",
"conversation",
"background",
"max_tool_calls",
"user",
"context_management",
"stream",
"stream_options",
}
CLAUDE_MAPPED_FIELDS = {
"id",
"type",
"role",
"text",
"content",
"source",
"name",
"description",
"input",
"input_schema",
"messages",
"model",
"max_tokens",
"system",
"temperature",
"top_p",
"top_k",
"stop_sequences",
"tool_choice",
"tools",
"metadata",
"thinking",
"output_config",
"usage",
"stop_reason",
"stop_sequence",
}
CLAUDE_PROVIDER_ONLY_FIELDS = {
"cache_control",
"container",
"inference_geo",
"service_tier",
"allowed_callers",
"allowed_domains",
"blocked_domains",
"defer_loading",
"max_uses",
"strict",
"user_location",
"citations",
"context",
"title",
"file_id",
"document_index",
"document_title",
"cited_text",
"caller",
}
def split_markdown_row(line: str) -> list[str]:
cells: list[str] = []
current: list[str] = []
escaped = False
for char in line:
if char == "|" and not escaped:
cells.append("".join(current).strip())
current.clear()
else:
current.append(char)
escaped = char == "\\" and not escaped
if escaped and char != "\\":
escaped = False
cells.append("".join(current).strip())
return cells
def strip_markdown_code(value: str) -> str:
value = value.strip()
if value.startswith("`") and value.endswith("`"):
value = value[1:-1]
return value.replace("\\|", "|")
def escape_markdown_cell(value: str) -> str:
return value.replace("|", "\\|")
def parse_schema_heading(line: str) -> str | None:
if not line.startswith("### `"):
return None
rest = line[len("### `") :]
schema, _, _ = rest.partition("`")
return schema or None
def parse_provider_definition_fields(definitions: str) -> list[SourceField]:
provider: str | None = None
schema: str | None = None
fields: list[SourceField] = []
for line in definitions.splitlines():
if line.startswith("## "):
if "OpenAI Schema" in line:
provider = "OpenAI"
elif "Claude / Anthropic TypeScript" in line:
provider = "Claude"
elif "Gemini Schema" in line:
provider = "Gemini"
else:
provider = None
schema = None
continue
if provider is None:
continue
if heading := parse_schema_heading(line):
schema = heading
continue
if schema is None or not line.startswith("| `"):
continue
cells = split_markdown_row(line)
if len(cells) < 4 or cells[2] not in {"是", "否"}:
continue
fields.append(
SourceField(
provider=provider,
schema=schema,
field=strip_markdown_code(cells[1]),
required=cells[2],
field_type=strip_markdown_code(cells[3]),
)
)
return fields
def parse_existing_coverage(
matrix: str,
) -> tuple[dict[tuple[str, str, str], CoverageStatus], dict[tuple[str, str], list[CoverageStatus]]]:
existing: dict[tuple[str, str, str], CoverageStatus] = {}
profiles: dict[tuple[str, str], list[CoverageStatus]] = defaultdict(list)
for line in matrix.splitlines():
if not line.startswith("| "):
continue
cells = split_markdown_row(line)
if len(cells) < 11 or cells[1] not in {"OpenAI", "Claude", "Gemini"}:
continue
status = CoverageStatus(
surface=cells[6],
same_format_runtime=cells[7],
canonical_roundtrip=cells[8],
cross_format=cells[9],
notes=cells[10],
)
provider = cells[1]
schema = strip_markdown_code(cells[2])
field = strip_markdown_code(cells[3])
existing[(provider, schema, field)] = status
profiles[(provider, schema)].append(status)
return existing, profiles
def most_common(values: Iterable[str]) -> str | None:
values = list(values)
if not values:
return None
return Counter(values).most_common(1)[0][0]
def openai_surface(schema: str) -> str:
if "CreateChatCompletion" in schema or "ChatCompletion" in schema:
return "openai:chat standard"
if "CreateEmbedding" in schema or "Embedding" in schema:
return "openai:embedding"
if any(token in schema for token in ("CreateImage", "EditImage", "Image", "Images")):
return "openai:image native-only"
if any(token in schema for token in ("Compact", "Compaction")):
return "openai:responses:compact native-only"
if any(
token in schema
for token in (
"Response",
"Input",
"Output",
"Tool",
"Reasoning",
"WebSearch",
"FileSearch",
"Computer",
"MCP",
"CodeInterpreter",
"Function",
"Custom",
"EasyInput",
"Prompt",
"Conversation",
"Annotation",
"Citation",
"LogProb",
"TopLogProb",
"Metadata",
"ServiceTier",
"Verbosity",
"TextResponse",
"ResponseFormat",
"Include",
"Modalities",
"ParallelToolCalls",
"StopConfiguration",
)
):
return "openai:responses standard"
return "openai auxiliary / not-in-conversion-surface"
def openai_default_status(field: SourceField, profile: list[CoverageStatus]) -> CoverageStatus:
surface = most_common(status.surface for status in profile) or openai_surface(field.schema)
if "not-in-conversion-surface" in surface or "native-only" in surface:
return CoverageStatus(
surface=surface,
same_format_runtime="native",
canonical_roundtrip="not-in-conversion-surface",
cross_format="not-in-conversion-surface",
notes="not part of current canonical cross-format conversion; same-format runtime path remains provider-native when routed directly",
)
if field.schema == "CreateChatCompletionRequest":
if field.field in OPENAI_CHAT_MAPPED:
return CoverageStatus(
surface=surface,
same_format_runtime="native",
canonical_roundtrip="mapped",
cross_format="mapped",
notes="Chat request field maps provider-specifically; target-incompatible cases fail closed",
)
if field.field in OPENAI_CHAT_BLOCKED:
return CoverageStatus(
surface=surface,
same_format_runtime="native",
canonical_roundtrip="extension-preserved",
cross_format="lossy-blocked",
notes="Chat-only or provider-specific field has no audited lossless target equivalent",
)
if field.schema == "CreateResponse":
if field.field in OPENAI_RESPONSES_MAPPED:
return CoverageStatus(
surface=surface,
same_format_runtime="native",
canonical_roundtrip="mapped",
cross_format="mapped",
notes="Responses request field maps provider-specifically; target-incompatible cases fail closed",
)
if field.field in OPENAI_RESPONSES_BLOCKED:
return CoverageStatus(
surface=surface,
same_format_runtime="native",
canonical_roundtrip="extension-preserved",
cross_format="lossy-blocked",
notes="Responses-only field has no audited lossless Chat/Claude/Gemini target equivalent",
)
if profile:
cross_format = most_common(status.cross_format for status in profile) or "lossy-blocked"
return CoverageStatus(
surface=surface,
same_format_runtime="native",
canonical_roundtrip=most_common(status.canonical_roundtrip for status in profile)
or "extension-preserved",
cross_format=cross_format,
notes=next(
(status.notes for status in profile if status.cross_format == cross_format),
"schema-level handling inherited from audited sibling fields",
),
)
return CoverageStatus(
surface=surface,
same_format_runtime="native",
canonical_roundtrip="extension-preserved",
cross_format="lossy-blocked",
notes="OpenAI documented field is preserved same-format; cross-format requires explicit target mapping or fails closed",
)
def claude_default_status(field: SourceField) -> CoverageStatus:
if "CountTokens" in field.schema:
return CoverageStatus(
surface="claude:messages/count_tokens native-only",
same_format_runtime="native",
canonical_roundtrip="not-in-conversion-surface",
cross_format="not-in-conversion-surface",
notes="count_tokens schemas are provider-native and outside canonical generation conversion",
)
if field.field in CLAUDE_MAPPED_FIELDS:
return CoverageStatus(
surface="claude:messages standard",
same_format_runtime="native",
canonical_roundtrip="mapped",
cross_format="mapped/lossy-blocked",
notes="Claude field maps where canonical and target support an equivalent; otherwise conversion fails closed",
)
if (
field.field in CLAUDE_PROVIDER_ONLY_FIELDS
or field.field.endswith("_tokens_details")
or "cache" in field.field
):
return CoverageStatus(
surface="claude:messages standard",
same_format_runtime="native",
canonical_roundtrip="extension-preserved",
cross_format="lossy-blocked",
notes="Claude provider-specific field is preserved same-format and blocked cross-format without an audited target equivalent",
)
return CoverageStatus(
surface="claude:messages standard",
same_format_runtime="native",
canonical_roundtrip="extension-preserved",
cross_format="lossy-blocked",
notes="Claude nested/provider-specific field is same-format preserved; cross-format requires explicit mapping or fails closed",
)
def gemini_default_status(field: SourceField, profile: list[CoverageStatus]) -> CoverageStatus:
if profile:
cross_format = most_common(status.cross_format for status in profile) or "lossy-blocked"
return CoverageStatus(
surface=most_common(status.surface for status in profile)
or "gemini:generate_content standard",
same_format_runtime="native",
canonical_roundtrip=most_common(status.canonical_roundtrip for status in profile)
or "extension-preserved",
cross_format=cross_format,
notes=next(
(status.notes for status in profile if status.cross_format == cross_format),
"Gemini field follows schema-level handling",
),
)
return CoverageStatus(
surface="gemini:generate_content standard",
same_format_runtime="native",
canonical_roundtrip="extension-preserved",
cross_format="lossy-blocked",
notes="Gemini documented field is preserved same-format; cross-format requires explicit mapping or fails closed",
)
def default_status(field: SourceField, profile: list[CoverageStatus]) -> CoverageStatus:
if field.provider == "OpenAI":
return openai_default_status(field, profile)
if field.provider == "Claude":
return claude_default_status(field)
if field.provider == "Gemini":
return gemini_default_status(field, profile)
raise ValueError(f"unsupported provider: {field.provider}")
def render_matrix(
fields: list[SourceField],
existing: dict[tuple[str, str, str], CoverageStatus],
profiles: dict[tuple[str, str], list[CoverageStatus]],
) -> str:
rows: list[str] = [
"# Format Field Coverage Matrix",
"",
"Last generated: 2026-06-03",
"",
"This file is generated from the schema inventory in `docs/api/provider-interface-definitions.md` and gives every documented schema field an explicit handling status. “处理到” here means the field is either mapped, preserved in same-format paths, rejected with a structured fail-closed error, or explicitly outside the current conversion surface. It does not mean every field can be cross-format converted.",
"",
"Provider schema updates do not require immediate conversion-code changes for runtime safety. Same-format runtime paths bypass canonical conversion, and same-format canonical roundtrip preserves provider extension fields. Cross-format conversion only enables fields with an audited semantic mapping; newly discovered or unknown provider fields default to structured fail-closed behavior until mapped.",
"",
"Regenerate with: `python3 docs/api/generate_format_field_coverage.py`.",
"",
"Statuses used in this matrix: `native`, `mapped`, `mapped/lossy-blocked`, `extension-preserved`, `unaudited`, `unsupported`, `invalid-enum`, `lossy-blocked`, `not-in-conversion-surface`.",
"",
"| Provider | Schema | Field | Required | Type | Surface | Same-Format Runtime | Canonical Roundtrip | Cross-Format | Notes |",
"| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |",
]
for field in fields:
status = existing.get(
(field.provider, field.schema, field.field),
default_status(field, profiles[(field.provider, field.schema)]),
)
rows.append(
"| "
+ " | ".join(
[
field.provider,
f"`{escape_markdown_cell(field.schema)}`",
f"`{escape_markdown_cell(field.field)}`",
field.required,
f"`{escape_markdown_cell(field.field_type)}`",
escape_markdown_cell(status.surface),
escape_markdown_cell(status.same_format_runtime),
escape_markdown_cell(status.canonical_roundtrip),
escape_markdown_cell(status.cross_format),
escape_markdown_cell(status.notes),
]
)
+ " |"
)
rows.extend(["", f"Total covered schema fields: {len(fields)}."])
return "\n".join(rows) + "\n"
def main() -> int:
parser = argparse.ArgumentParser()
parser.add_argument("--definitions", type=Path, default=DEFAULT_DEFINITIONS)
parser.add_argument("--matrix", type=Path, default=DEFAULT_MATRIX)
parser.add_argument("--check", action="store_true")
args = parser.parse_args()
definitions = args.definitions.read_text()
current_matrix = args.matrix.read_text() if args.matrix.exists() else ""
fields = parse_provider_definition_fields(definitions)
existing, profiles = parse_existing_coverage(current_matrix)
next_matrix = render_matrix(fields, existing, profiles)
if args.check:
if current_matrix != next_matrix:
print(
f"{args.matrix} is not up to date; run "
"`python3 docs/api/generate_format_field_coverage.py`",
)
return 1
return 0
args.matrix.write_text(next_matrix)
print(f"wrote {len(fields)} field coverage rows to {args.matrix}")
return 0
if __name__ == "__main__":
raise SystemExit(main())
-53
View File
@@ -1,53 +0,0 @@
# 提供商管理的端点健康度
以下管理接口使用相同的健康度汇总规则:
- `GET /api/admin/providers/summary`
- `GET /api/admin/providers/{provider_id}/summary`
## 统计规则
沿用 `v0.7.13`(`535ee098c`)的默认健康规则:已启用密钥缺少该格式的健康记录时,
按 `1.0` 参与统计,而不是要求先有一次请求或探测才能显示健康。
- 仅统计启用端点下、支持该端点 API 格式的启用密钥。
- 每个密钥读取其 `health_by_format[api_format].health_score`,不跨格式借用分数。
- 对上述启用密钥求算术平均;缺少有效分数的密钥沿用旧版默认值 `1.0`。
- 分数范围为 `0` 到 `1`,沿用调度器的分数读取与范围约束。
- 停用端点或没有启用密钥时,端点的 `health_score` 返回 `null`。
- `avg_health_score` 仅平均有启用密钥的启用端点,包含按默认值计算的端点;没有此类端点时返回 `null`。
- `unhealthy_endpoints` 仅统计上述端点中健康度低于 `0.5` 的数量,不把未知状态算作故障。
`total_keys` 和 `active_keys` 仍反映密钥配置数量,不因缺少健康数据而减少。
Codex、Kiro、Gemini CLI、Antigravity 等固定提供商的账号按照各提供商的认证规则,
继承其启用端点的 API 格式;账号的 `api_formats` 为 `null` 或空数组,不代表没有配置账号。
继承格式决定账号归属;缺少对应格式的分数时同样使用默认值 `1.0`,不借用其他格式的异常分数。
## 数据读取
摘要从密钥的轻量投影读取 API 格式、启用状态和 `health_by_format`。
PostgreSQL 投影中的凭据字段使用 `summary` / `{}` 等脱敏占位值,并非真实密文,
因此摘要读取不执行凭据解密、认证或迁移。完整密钥读取仍保留原有的凭据安全校验。
若将这些占位值送入凭据校验,读取会失败,旧的摘要聚合还会将其当作空密钥列表,
导致已配置密钥的端点也被错误显示为灰色。默认健康分数只能用于成功读取的启用密钥,
不能用于掩盖查询或凭据投影错误。
提供商、端点或密钥摘要查询失败时,接口返回 `503`,不能将失败当作空列表并返回零账号。
单个提供商确实不存在时仍返回 `404`。页面刷新失败保留已有列表,并显示加载错误。
## 页面展示
桌面表格、网格卡片和手机卡片使用相同规则:
- 有健康数据时显示百分比;有效的零分显示 `0%`。
- 有启用密钥、但这些密钥尚无该格式的健康记录时,显示绿色 `100%`,与 `v0.7.13` 一致。
- 端点停用、未配置密钥或没有启用密钥时显示灰色占位条和对应状态提示。
账号详情保留原有默认 `100%` 的规则。已有观测仍取各格式最低分;端点分数则只聚合对应格式,
两者统计范围不同,不要求百分比完全相等。调度器原有的缺省健康策略不变。
此分数是密钥当前健康状态的聚合,包含默认健康值,不是某个时间窗口内的请求成功率,
也不表示已经执行过主动探测。相比 `v0.7.13`,仍保留停用账号/端点不参与健康聚合、
凭据脱敏与查询失败显式报错等修复,不整体回退旧版代码。
File diff suppressed because it is too large Load Diff
-58
View File
@@ -1,58 +0,0 @@
# Rerank API
Aether exposes an OpenAI-compatible rerank surface at `POST /v1/rerank` and can route it to providers configured as `openai:rerank` or `jina:rerank`.
## Request
```http
POST /v1/rerank
Authorization: Bearer <aether-api-key>
Content-Type: application/json
```
```json
{
"model": "bge-reranker-base",
"query": "What document discusses gateway routing?",
"documents": [
"Aether routes public AI requests through the Rust gateway.",
"This document discusses unrelated content."
],
"top_n": 1,
"return_documents": true
}
```
Fields:
| Field | Required | Notes |
| --- | --- | --- |
| `model` | Yes | Aether global model name. |
| `query` | Yes | Non-empty query string. |
| `documents` | Yes | Non-empty array of strings or provider-native document objects. |
| `top_n` | No | Positive integer. |
| `return_documents` | No | Provider-compatible flag for including matched documents. |
## Response
Aether forwards the provider JSON response. OpenAI-compatible and Jina-compatible rerank providers commonly return `results[]`:
```json
{
"model": "bge-reranker-base",
"results": [
{
"index": 0,
"relevance_score": 0.98,
"document": {
"text": "Aether routes public AI requests through the Rust gateway."
}
}
],
"usage": {
"total_tokens": 32
}
}
```
Rerank requests must be JSON and do not support `stream` or chat `messages` payloads.