feat: revamp analytics dashboards and harden database migrations

Add dashboard and overview analytics, health monitoring, provider expense tracking, and announcement updates across the gateway and frontend.

Keep schema migrations free of historical backfills while preserving automatic backfill execution. Bound migration deadlines, run schema preparation before Compose replacement, and anonymize deleted dashboard users.

Include the current documentation cleanup and regression coverage.
This commit is contained in:
elky
2026-10-01 11:48:17 +08:00
parent 60b89cc840
commit 066ea87d72
327 changed files with 31728 additions and 20645 deletions
-428
View File
@@ -1,428 +0,0 @@
# WebSocket transports
Aether exposes several independent WebSocket surfaces. They share transport
machinery, but not request schemas or continuation state:
| Public route | API format | Protocol |
| --- | --- | --- |
| `GET /v1/responses` | `openai:responses` | Responses WebSocket mode; every turn starts with `response.create`. |
| `GET /v1/realtime?model=...` | `openai:realtime` | OpenAI Realtime JSON events, including Base64 audio events. |
| `GET /v1/realtime?model=...` with a first-party Codex `originator` | `codex:live` | Current Codex Realtime v2 direct WebSocket transport. |
| `POST /v1/realtime/calls`, `GET /v1/realtime?intent=quicksilver&call_id=...` | `codex:live` | Codex Realtime v1 AVAS WebRTC call creation and sideband transport. |
| `GET/POST /v1/live[/{call_id}]` | `codex:live` | Legacy Codex Frameless direct and WebRTC compatibility transport. |
Do not point one surface at an endpoint configured for another. In particular,
a Realtime or Live event is not passed through the Responses
`response.create` state machine.
## Responses WebSocket mode
The Responses API supports a WebSocket mode for long-running, tool-call-heavy workflows. In this mode, you keep a persistent connection to `/v1/responses` and continue each turn by sending only new input items plus `previous_response_id`.
WebSocket mode is compatible with both Zero Data Retention (ZDR) and `store=false`.
## OpenAI Realtime WebSocket bridge
Configure an active `openai:realtime` provider endpoint, then connect to:
```text
wss://<aether-host>/v1/realtime?model=<authorized-global-model>
```
Aether authenticates and plans the request before returning the downstream
WebSocket upgrade. The global model alias is replaced in the upstream query,
while safe non-credential query parameters, provider authentication,
`header_rules`, and proxy settings continue to apply. Client credentials in
the query string are rejected or removed rather than forwarded upstream.
After the handshake, Aether relays text, binary, ping, pong, and close frames
one at a time. JSON events, Base64 audio payloads, and unknown future fields are
not rebuilt or coalesced. When an upstream `response.done` contains an
authoritative `response.usage`, its text/audio token counters are accumulated
for the connection's usage record. A session that closes without authoritative
usage is recorded as `usage_available=false`; Aether does not estimate token
counts, audio duration, or cost from frame sizes.
`response.done` covers Realtime Response usage. Optional input transcription
is reported by a different event and can use a different transcription model;
it is not folded into the Response model's session row or priced as if it used
that model. Finite-balance Realtime access therefore remains fail-closed until
multi-event, multi-model settlement is implemented.
The upstream handshake is completed before Aether sends HTTP 101 to the
client. A provider authentication, TLS, proxy, or upgrade failure therefore
returns an ordinary bounded HTTP error instead of opening a socket that fails
immediately.
See the official [OpenAI Realtime WebSocket guide](https://developers.openai.com/api/docs/guides/realtime-websocket)
for the current event contract.
## Experimental Codex Live bridge
Aether also exposes the Codex Frameless Bidi V3 transport used by current
Codex clients. It is related to the OpenAI Realtime API, but it is not the
Responses WebSocket protocol and never enters Aether's `response.create`
state machine:
- Current WebRTC call creation: `POST /v1/realtime/calls?intent=quicksilver&architecture=avas`
with bounded `sdp` and `session` multipart parts. Aether normalizes those
selectors, applies the global-to-provider mapping, and rewrites the upstream
`Location` to `/v1/realtime/calls/<call-id>`.
- Current WebRTC sideband: `GET /v1/realtime?intent=quicksilver&call_id=<call-id>`.
The unique `intent=quicksilver` selector classifies it as Codex Live;
`call_id` alone remains on the separate OpenAI Realtime surface.
- Current direct WebSocket (Realtime v2):
`GET /v1/realtime?model=<authorized-global-model>`. V2 intentionally omits
`intent=quicksilver`; Aether uses the first-party Codex `originator` header
to distinguish it from ordinary OpenAI Realtime. The first client event is
still the opaque `session.update` frame generated by Codex; Aether forwards
it without converting it into a Responses `response.create` event.
- Realtime v1 direct WebSocket uses
`GET /v1/realtime?intent=quicksilver&model=<authorized-global-model>` and
sends `openai-alpha: quicksilver=v1` upstream. V2 does not send that header.
- Legacy direct WebSocket: `GET /v1/live?model=<global-model>`. The first client text
frame must be `session.update`; later text, binary, ping, pong, and close
frames are relayed opaquely.
- Legacy WebRTC call creation: `POST /v1/live` with bounded `sdp` and `session`
multipart parts. Aether applies the existing global-to-provider model
mapping and rewrites the upstream `Location` to `/v1/live/<call-id>`.
- Legacy WebRTC sideband: `GET /v1/live/<call-id>`. Frameless sideband attaches to an
already initialized call, so Aether neither waits for nor sends a second
`session.update` frame.
The provider must expose an active, dedicated `codex:live` endpoint. Fixed
Codex providers receive this endpoint from the managed provider template;
custom providers can add it in the endpoint editor. The
`responses_websocket.enabled` provider option belongs only to
`openai:responses` WebSocket mode and is not reused as the Live permission.
Codex Live also requires an authorized model mapping; adding the endpoint alone
is not enough. For WebRTC, Codex sends the selected model in the multipart
`session.model`. Aether treats that value as the downstream global model,
selects an existing mapping whose provider endpoint is `codex:live`, and
rewrites only `session.model` to the mapped upstream model. Configure Codex's
realtime model selection to an authorized global alias (the current Codex
default may be `gpt-realtime-1.5`, but that value is client-version dependent),
or create an authorized mapping for the alias the client already sends. Aether
does not invent a Live model name or add a bundled hard-coded model merely
because the endpoint is enabled. For Codex providers only, an existing model mapping scoped to
`openai:responses` (including the historical `/v1/responses` alias) can be
reused for Live. The provider endpoint and key must still explicitly allow
`codex:live`; OpenAI and custom providers do not receive this compatibility
rule.
For Codex Desktop/app-server, point both the call-creation and sideband
overrides at the same Aether origin when using a custom provider. Explicitly
setting both avoids a client-version-dependent fallback to the OpenAI origin:
```toml
model_provider = "aether"
experimental_realtime_webrtc_call_base_url = "https://<aether-host>/v1"
experimental_realtime_ws_base_url = "https://<aether-host>/v1"
experimental_realtime_ws_model = "<authorized-global-live-alias>"
[model_providers.aether]
name = "Aether"
base_url = "https://<aether-host>/v1"
wire_api = "responses"
```
The two experimental overrides are optional when the selected provider's
`base_url` already points at Aether, but if either is set they must resolve to
the same deployment so the authenticated call binding can be found. A missing
call-create request in Aether means the client did not select this provider or
failed before gateway routing; a call-create request without the matching
sideband usually means the sideband origin or credential differs. These
settings may change with newer Codex releases.
API-key and bearer providers can use direct WebSocket or WebRTC. ChatGPT OAuth
uses the official Codex backend for WebRTC call creation and the OpenAI Live
origin for its sideband; direct OAuth WebSocket and custom OAuth backend
origins fail closed. The call binding fixes the authenticated downstream
principal, provider/endpoint/key, mapped model, auth mode, account/FedRAMP
identity, session identity, and upstream origin. Raw call IDs are hashed in
RuntimeState keys, records expire after two hours, each principal retains at
most 64 call bindings, and one call permits only one renewable sideband
attachment at a time. The memory RuntimeState backend loses these bindings on
restart. The two-hour binding TTL and 64-record cap bound routing state and
abuse; they are not provider-concurrency reservations.
Frameless V3 currently has no stable usage object that Aether can settle into
its wallet pipeline. Aether therefore enables Live only for principals without
a finite `balance_remaining`; finite-balance keys receive an explicit local
error instead of unmetered service. Aether writes one lifecycle record for each
relayed direct or sideband WebSocket connection, with frame/byte counts and
`usage_available=false`; it does not create one database row per audio frame.
The synchronous WebRTC call-creation exchange keeps its ordinary HTTP record
and is also marked usage-unavailable. The WebRTC media leg itself does not
traverse Aether after call creation, so Aether cannot observe or invent a
separate audio-session usage record, token count, duration, or cost for it.
Aether-relayed direct and sideband WebSocket connections are limited to 60
minutes. The provider-pool and admission leases cover only the synchronous
HTTP call-creation exchange and are released after its SDP response. Aether
cannot infer media lifetime from the binding TTL or sideband lifetime, so a
created call that never attaches a sideband is not held against provider
concurrency after call creation.
For the public GA Realtime API's connection and session concepts, see the
[OpenAI Realtime guide](https://developers.openai.com/api/docs/guides/realtime),
[WebRTC connection guide](https://developers.openai.com/api/docs/guides/realtime-webrtc),
and [server-side controls guide](https://developers.openai.com/api/docs/guides/realtime-server-controls).
OpenAI's current WebSocket service supports named `stream_id` lanes: requests on
the same lane are FIFO, while different lanes may run concurrently. Aether's
bridge currently exposes only the implicit default lane and deliberately
rejects `response.create.stream_id` until per-lane binding, ordering, timeout,
usage, and error routing are implemented end to end. Use separate WebSocket
connections for parallel runs through Aether. A syntactically valid named
`stream_id` is rejected with `responses_websocket_named_stream_unsupported`;
the error event echoes the validated ID so the client can associate the error
with its attempted lane. Invalid or untrusted IDs are not echoed.
## Why use WebSocket mode
WebSocket mode is most useful when a workflow involves many model-tool round trips (for example, agentic coding or orchestration loops with repeated tool calls).
Because the connection stays open and each turn sends only incremental input, WebSocket mode reduces per-turn continuation overhead and improves end-to-end latency across long chains. The [OpenAI WebSocket-mode guide](https://developers.openai.com/api/docs/guides/websocket-mode) reports up to roughly 40% faster end-to-end execution for workloads with 20 or more tool calls; this is an upstream product claim, not an Aether benchmark.
## Connect and create responses
In WebSocket mode, start each turn by sending a `response.create` event from the client. The payload mirrors the normal [Responses create body](https://developers.openai.com/api/reference/resources/responses/methods/create), except that transport-specific fields like `stream` and `background` are not used.
```python
from websocket import create_connection
import json
import os
ws = create_connection(
"wss://api.openai.com/v1/responses",
header=[
f"Authorization: Bearer {os.environ['OPENAI_API_KEY']}",
],
)
ws.send(
json.dumps(
{
"type": "response.create",
"model": "gpt-5.6",
"store": False,
"input": [
{
"type": "message",
"role": "user",
"content": [{"type": "input_text", "text": "Find fizz_buzz()"}],
}
],
"tools": [],
}
)
)
```
Clients can optionally warm up request state by sending `response.create` with `generate: false`. This is useful when you already know the tools, instructions, and/or custom messages you plan to send with an upcoming turn. `generate: false` does not return a model output, but prepares request state so the next generated turn can start faster. The warmup request returns a response ID that you can chain from with `previous_response_id`, including on later turns in a response chain. The next section explains how to continue a session using `previous_response_id` and incremental inputs.
## Continue with incremental inputs
To continue a run, send another `response.create` with:
- `previous_response_id` set to the prior response ID.
- `input` containing only new items (for example, tool outputs and the next user message).
```python
ws.send(
json.dumps(
{
"type": "response.create",
"model": "gpt-5.6",
"store": False,
"previous_response_id": "resp_123",
"input": [
{
"type": "function_call_output",
"call_id": "call_123",
"output": "tool result",
},
{
"type": "message",
"role": "user",
"content": [{"type": "input_text", "text": "Now optimize it."}],
},
],
"tools": [],
}
)
)
```
## How continuation works
WebSocket mode uses the same `previous_response_id` chaining semantics as HTTP mode, but it adds a lower-latency continuation path on the active socket.
On an active Aether WebSocket connection, the selected upstream keeps the
previous-response state for the single default lane in its connection-local
cache. Continuing from that most recent response is fast because the service
can reuse connection-local state. Because the previous-response state is
retained only in memory and is not written to disk, you can use WebSocket mode
in a way that is compatible with `store=false` and Zero Data Retention (ZDR).
If a `previous_response_id` is not in the upstream connection's in-memory
cache, behavior depends on whether the upstream stored the response:
- With `store=true`, the upstream service may hydrate older response IDs from its persisted state when available. Continuation can still work, but it usually loses the in-memory latency benefit.
- With `store=false` (including ZDR), there is no persisted fallback. If the ID is uncached, the request returns `previous_response_not_found`.
For a new downstream WebSocket connection, Aether also has to prove that the
response belongs to the currently authenticated user/API key and to the exact
provider endpoint, key, credential generation, transport, adapter, model, and
normalization contract. Aether records this ownership only when the effective
provider `response.create`, after Aether's body rules and framing, explicitly
has `store=true` and a successful
`response.completed`, `response.done`, or non-error `response.incomplete`
terminal supplies a valid response ID. `store=false`, an omitted/overridden
`store`, failures, cancellations, malformed IDs, and ZDR turns never create
this registry state. This explicit-true rule is intentionally conservative:
Aether does not infer a provider default for an omitted `store` field.
The RuntimeState key contains a SHA-256 digest over length-delimited live
`user_id`, `api_key_id`, and the opaque response ID; raw response IDs and
credentials are not stored in keys or values. Records expire after 24 hours,
are capped at the 1,024 most recently registered IDs per user/API-key pair,
and are bounded in serialized size. The registry stores ownership/routing
proof and contract digests only; it does not store response contents and is
not a replacement for the upstream's `store=true` persistence. A registry
write failure does not turn a successful provider response into a failure, so
that terminal can reach the client but cannot later resume on a new socket.
With the Redis RuntimeState backend, ownership is shared across gateway
instances for the record TTL, subject to that Redis deployment's own
availability and persistence configuration. The memory backend is
process-local, is not shared between instances, and loses the registry on
restart. Expiry, per-principal eviction, a RuntimeState outage, or a memory
backend restart causes the first continuation on a new connection to fail
closed with `previous_response_not_found`, even if the upstream might still
retain the response. Aether never falls back to the ordinary scheduler for
such a miss and never sends the opaque response ID to a different provider or
key.
PII-redaction restore mappings intentionally remain connection-local and are
not persisted. If a stored response chain contains Aether PII sentinels, Aether
rejects cross-connection continuation rather than risk exposing those
sentinels without the original restore mapping. Start a new response with the
complete required context in that case.
If a continuation on the same lane fails (`4xx` or `5xx`), the service evicts
the referenced `previous_response_id` from the connection-local cache. Aether
only supports the implicit default lane, so this same-lane rule applies to all
continuations it currently accepts. The upstream service preserves a shared
parent when a cross-lane fork fails, but Aether does not yet expose that named
lane behavior.
The continuation must keep the model selected for the response chain. Aether
rejects a model change with status `409` and code
`responses_continuation_model_change_unsupported`.
## Compaction and creating new responses
If you are using compaction, there are two different continuation patterns:
### Server-side compaction (`context_management`)
When you enable server-side compaction (`context_management` with `compact_threshold`), compaction happens during normal `/responses` generation. In WebSocket mode, you continue the same way you normally do: send the next `response.create` with the latest `previous_response_id` and only new input items.
### Standalone `/responses/compact`
The standalone [`/responses/compact` endpoint](https://developers.openai.com/api/reference/resources/responses/methods/compact) returns a new compacted input window, not a response ID. After compaction, create a new response on your WebSocket connection using the compacted window as `input` (plus the next user/tool items).
Start a new chain by omitting `previous_response_id` or setting it to `null`. Pass the compacted output as-is; do not prune the returned window.
```python
# Compact your current window (HTTP call)
compacted = client.responses.compact(
model="gpt-5.6",
input=long_input_items_array,
)
# Start a new response on the WebSocket using the compacted window
ws.send(
json.dumps(
{
"type": "response.create",
"model": "gpt-5.6",
"store": False,
"input": [
*compacted.output,
{
"type": "message",
"role": "user",
"content": [{"type": "input_text", "text": "Continue from here."}],
},
],
"tools": [],
}
)
)
```
## Connection behavior and limits
- Server events and ordering match the existing Responses streaming event model.
- A single Aether WebSocket connection can receive multiple `response.create` messages over its lifetime, but the client must wait for a terminal event before sending the next one. Aether does not queue overlapping creates and returns `response_already_in_progress` while a turn is active.
- Named `stream_id` multiplexing is not exposed by Aether yet. Use multiple connections if you need parallel runs.
- The upstream OpenAI service allows at most 16 active/in-flight responses on one connection; additional `response.create` events are queued. It also allows at most 32 distinct named `stream_id` values per connection, and the implicit default lane does not count toward that 32-lane limit. These describe upstream multiplexing limits, not capabilities exposed by Aether's current single-lane bridge.
- Connection duration is limited to 60 minutes. Reconnect when the limit is reached.
- Aether binds each upstream WebSocket to one selected provider key. A provider must explicitly enable the standard Responses WebSocket capability and expose an `openai:responses` endpoint before it is eligible for this bridge.
- The Codex adapter additionally watches Codex quota events. A `usage_limit_reached` terminal error immediately marks the bound account unavailable. If the client has not received a standard `response.*` event and the request has no `previous_response_id`, Aether retries that one turn once on another eligible key without closing the public socket.
- After a standard response event has reached the client, after a retry has already been attempted, or for a request using `previous_response_id`, Aether forwards the provider terminal error and detaches only the exhausted upstream. If the upstream closes immediately after the quota signal, Aether emits a recoverable gateway error instead. The public WebSocket stays open so a later independent `response.create` can select another key.
- Aether does not transparently move an existing response chain to another provider key. Connection-local `previous_response_id` state cannot be transferred safely, especially with `store=false`/ZDR; send a new request with complete input after an exhausted continuation.
## Reconnect and recover
When a connection closes (or hits the 60-minute limit), open a new WebSocket
connection and continue with one of these patterns:
1. If the prior response is persisted (`store=true`) and its response ID remains valid, continue with `previous_response_id` and only the new input items.
2. If the chain cannot be hydrated (for example, `store=false`/ZDR or `previous_response_not_found`), start a new response by setting `previous_response_id` to `null` (or omitting it) and send the complete input context needed for the next turn.
3. If you compacted context with `/responses/compact`, use the returned compacted window as the base `input` for that new response, then append the latest user/tool items.
## Errors to handle
`previous_response_not_found`
```json
{
"type": "error",
"status": 400,
"error": {
"type": "invalid_request_error",
"code": "previous_response_not_found",
"message": "Previous response with id 'resp_abc' not found.",
"param": "previous_response_id"
}
}
```
`websocket_connection_limit_reached`
```json
{
"type": "error",
"error": {
"type": "invalid_request_error",
"code": "websocket_connection_limit_reached",
"message": "Responses websocket connection limit reached (60 minutes). Create a new websocket connection to continue."
},
"status": 400
}
```
## Related guides
- [Conversation state](https://developers.openai.com/api/docs/guides/conversation-state)
- [Streaming API responses](https://developers.openai.com/api/docs/guides/streaming-responses)
- [Responses streaming events reference](https://developers.openai.com/api/reference/resources/responses)
- [Responses WebSocket events reference](https://developers.openai.com/api/reference/resources/responses/websocket-events)
-222
View File
@@ -1,222 +0,0 @@
# Embeddings API
Aether supports OpenAI compatible embedding requests through `POST /v1/embeddings`. Embedding requests are separate from chat and responses requests. They use `input`, never `messages`, and they are always non streaming.
## Quick Start
Run this against your Aether gateway URL with a user API key that can access the model and the `openai:embedding` API format.
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: application/json" \
-d '{
"model": "text-embedding-3-small",
"input": ["hello", "world"],
"encoding_format": "float"
}'
```
## Public Endpoint
| Method | Path | Client API format | Route kind |
| --- | --- | --- | --- |
| `POST` | `/v1/embeddings` | `openai:embedding` | `embedding` |
The gateway classifies this endpoint as an OpenAI family embedding route with endpoint signature `openai:embedding`. It is not handled as chat or responses.
## Request Body
Required fields:
| Field | Type | Notes |
| --- | --- | --- |
| `model` | string | Must name a model allowed for the API key and user. Blank strings are rejected. |
| `input` | string, string array, integer token array, nested integer token arrays, or multimodal object array | Must be non empty. Empty strings, empty arrays, empty token arrays, and empty multimodal objects are rejected. |
Optional fields that pass through the embedding conversion path when supported by the provider:
| Field | Notes |
| --- | --- |
| `encoding_format` | Passed to OpenAI compatible providers. |
| `dimensions` | Passed to providers whose embedding request shape supports it. |
| `parameters` | Provider-specific embedding parameters. For Aliyun DashScope this maps to DashScope `parameters`; `dimensions` is emitted as `parameters.dimension` unless `parameters.dimension` is already set. |
| `user` | Passed to OpenAI compatible providers. |
| `task` | Passed to Jina and OpenAI compatible embedding requests. Jina defaults to `text-matching` when no task is supplied. |
Accepted `input` shapes:
```json
{ "model": "text-embedding-3-small", "input": "hello" }
```
```json
{ "model": "text-embedding-3-small", "input": ["hello", "world"] }
```
```json
{ "model": "text-embedding-3-small", "input": [1, 2, 3] }
```
```json
{ "model": "text-embedding-3-small", "input": [[1, 2], [3, 4]] }
```
```json
{
"model": "qwen3-vl-embedding",
"input": [
{ "text": "white running shoes" },
{ "image": "https://dashscope.oss-cn-beijing.aliyuncs.com/images/256_1.png" }
],
"parameters": { "enable_fusion": true }
}
```
Use string or string array input when routing to Gemini or Doubao embedding providers. Token arrays are accepted by the OpenAI compatible public endpoint, but Gemini, Doubao, and Aliyun provider request emitters require text or multimodal content input.
## Provider Format Mapping
Embedding routes can select only embedding provider API formats. Chat, responses, image, and generation formats are not valid provider targets for this request type.
| Provider API format | Upstream path shape | Provider request shape |
| --- | --- | --- |
| `openai:embedding` | `/v1/embeddings` | OpenAI compatible `{ "model", "input" }` payload. |
| `jina:embedding` | `/v1/embeddings` | OpenAI compatible payload with a Jina `task`. Defaults to `text-matching` if omitted. |
| `gemini:embedding` | `models/{model}:embedContent` | Single text input uses `content.parts[].text`. Multiple text inputs use `requests[].content.parts[].text`. |
| `doubao:embedding` | `/embeddings/multimodal` | Text input is emitted as `input` items like `{ "type": "text", "text": "..." }`. |
| `aliyun:multimodal_embedding` | `/api/v1/services/embeddings/multimodal-embedding/multimodal-embedding` | Text and multimodal inputs are emitted as DashScope `input.contents`. Supports `text`, `image`, `video`, `multi_images`, `parameters.enable_fusion`, `parameters.res_level`, and `parameters.max_video_frames`. Alias: `dashscope:multimodal_embedding`. |
Custom provider endpoint paths are available when the endpoint is configured for an embedding API format. Gemini custom paths can use `{model}` and `{action}`. For `gemini:embedding`, `{action}` expands to `embedContent`.
## Model And Catalog Requirements
To use embeddings through the gateway:
1. The global model should include embedding metadata, for example `supported_capabilities: ["embedding"]`, `config.model_type: "embedding"`, or `config.api_formats` with one of the embedding formats.
2. The provider model or mapping must expose an embedding API format, one of `openai:embedding`, `gemini:embedding`, `jina:embedding`, `doubao:embedding`, or `aliyun:multimodal_embedding`.
3. The user and API key must be allowed to access the model and the `openai:embedding` client API format.
4. Public and admin catalog responses expose `supports_embedding` so clients can display embedding capability separately from chat.
Billing fails closed for embedding global models. A model marked as embedding capable must define either `default_price_per_request` or `default_tiered_pricing.tiers[].input_price_per_1m`. Missing request pricing and missing input token pricing cause the model record to be rejected instead of treated as free.
No schema migration is needed for embedding metadata. Existing model capability, config, provider mapping, API format, and pricing fields carry the data.
## Aliyun Qwen3-VL Examples
Text request through Aether:
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-vl-embedding",
"input": "white running shoes",
"dimensions": 1024
}'
```
Image and text fusion request:
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-vl-embedding",
"input": [
{ "text": "white running shoes, lightweight and breathable" },
{ "image": "https://dashscope.oss-cn-beijing.aliyuncs.com/images/256_1.png" }
],
"parameters": { "enable_fusion": true }
}'
```
Video request:
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-vl-embedding",
"input": [
{ "video": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20250107/lbcemt/new+video.mp4" }
],
"parameters": { "max_video_frames": 64 }
}'
```
Multi-image fusion request:
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-vl-embedding",
"input": [
{ "text": "product photos from multiple angles" },
{ "multi_images": [
"https://example.com/front.png",
"https://example.com/side.png"
] }
],
"parameters": { "enable_fusion": true }
}'
```
## Failure Behavior
The gateway validates deterministic request errors before local execution or provider transport.
| Case | Example request body or setup | Status | Error detail |
| --- | --- | --- | --- |
| Invalid JSON | `{` | `400` | `Embedding request JSON body is invalid` |
| Missing model | `{ "input": "hello" }` | `400` | `Embedding request model is required` |
| Empty input | `{ "model": "text-embedding-3-small", "input": [] }` | `400` | `Embedding request input is required` |
| Chat `messages` payload | `{ "model": "text-embedding-3-small", "messages": [] }` | `400` | `Embedding request must use input, not chat messages` |
| Streaming requested | `{ "model": "text-embedding-3-small", "input": "hello", "stream": true }` | `400` | `Embedding requests do not support streaming` |
| Non JSON content type | `Content-Type: text/plain` with an embedding JSON body | `400` | `Embedding request content-type must be application/json` |
| Chat only model | API key allows `text-embedding-3-small`, request uses `gpt-5` | `403` | The key is not allowed to access that model. |
| Chat only API format | API key allows `openai:chat` but not `openai:embedding` | `403` | The key is not allowed to access `openai:embedding`. |
Failure examples:
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: application/json" \
-d '{"model":"text-embedding-3-small","messages":[]}'
```
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: application/json" \
-d '{"input":"hello"}'
```
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: application/json" \
-d '{"model":"text-embedding-3-small","input":[]}'
```
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: application/json" \
-d '{"model":"text-embedding-3-small","input":"hello","stream":true}'
```
```bash
curl -sS "http://localhost:8084/v1/embeddings" \
-H "Authorization: Bearer sk-your-aether-key" \
-H "Content-Type: text/plain" \
-d '{"model":"text-embedding-3-small","input":"hello"}'
```
If a valid embedding request passes local validation but no usable provider transport is available, the gateway can return a provider or service availability error. That is different from the deterministic request validation errors above.
-238
View File
@@ -1,238 +0,0 @@
# Format Conversion Audit
Last audited: 2026-06-03
This audit tracks `source format -> Canonical -> target format` behavior. It is intentionally stricter than historical best-effort conversion.
Statuses:
- `native`: emitted as a target-native field without semantic change.
- `mapped`: converted through canonical/provider-specific mapping.
- `extension-preserved`: preserved in same-format canonical roundtrip or target-approved extension namespace.
- `transport-only`: audited client transport metadata intentionally omitted when the target has no compatible transport channel.
- `unaudited`: rejected because the source field is not in the audited provider schema inventory for cross-format conversion.
- `unsupported`: rejected because the request/field shape is outside the supported conversion surface, independent of schema drift.
- `lossy-blocked`: conversion fails closed.
- `invalid-enum`: conversion fails closed because a provider enum value is not valid for the target mapping.
Full schema field coverage is tracked in `docs/api/format-field-coverage-matrix.md`. That matrix is generated from the schema inventory in `docs/api/provider-interface-definitions.md` by `python3 docs/api/generate_format_field_coverage.py` and gives every documented OpenAI, Claude, and Gemini schema field a handling status. “Handled” means mapped, same-format/native preserved, extension-preserved, blocked with a structured error, or explicitly marked outside the canonical conversion surface.
Provider schema refresh is not a runtime dependency. Same-format runtime paths do not use this matrix; they bypass canonical conversion. Same-format canonical roundtrip must preserve unrecognized provider fields through provider extension namespaces. Cross-format conversion is capability-based: only explicitly mapped fields are emitted, and newly discovered or unknown provider fields fail closed with `UnauditedField` until a lossless mapping is audited.
## Implemented Boundary Changes
| Area | Current behavior |
| --- | --- |
| Pure conversion API | `convert_request_pure` and response equivalents do not apply model override or stream policy. |
| Legacy conversion API | `convert_request` / `convert_response` are retained for migration and may still use legacy context behavior. |
| Same-format provider path | Bypasses canonical conversion and copies the parsed JSON object before transport edits. |
| Cross-format same-format-provider path | Uses `convert_request_pure`, then applies model/body/stream edits in transport. |
| Conversion errors | Added `UnauditedField`, `UnsupportedField`, `InvalidEnumValue`, `LossyConversionBlocked`, and `InvalidTargetField`. |
| Reporting | Added `ConversionReport` with field statuses. Runtime reports remain conversion-operation oriented; exhaustive nested schema coverage is enforced by `format-field-coverage-matrix.md`. |
| Source schema coverage | Cross-format request conversion rejects unknown source root fields before emit. Every documented schema field is covered by the field coverage matrix. |
| Schema drift handling | Official schema changes are detected by regenerating the inventory/matrix. Runtime same-format remains passthrough; cross-format unknowns return `UnauditedField` until deliberately mapped. |
| Tool schema roundtrip | Claude `input_schema` and Gemini `functionDeclarations.parameters` preserve raw same-format schema through provider-specific extensions. |
| Tool result ids | Chat `tool_call_id`, Responses `call_id`, Claude `tool_use_id`, and Gemini `functionResponse.id` are mapped through canonical tool IDs. |
## OpenAI Chat -> OpenAI Responses
| Chat field | Canonical handling | Responses output | Status |
| --- | --- | --- | --- |
| `model` | request identity | `model` | native |
| `messages` | canonical messages/instructions | `input`, `instructions` | mapped |
| `max_tokens` | generation max tokens | `max_output_tokens` | mapped |
| `max_completion_tokens` | generation max tokens | `max_output_tokens` | mapped |
| `temperature` | generation | `temperature` | native |
| `top_p` | generation | `top_p` | native |
| `top_logprobs` | generation | `top_logprobs` | native |
| `n` | generation but no Responses equivalent | none | lossy-blocked |
| `stop` | generation but no Responses equivalent | none | lossy-blocked |
| `presence_penalty` | generation but no Responses equivalent | none | lossy-blocked |
| `frequency_penalty` | generation but no Responses equivalent | none | lossy-blocked |
| `seed` | generation but no Responses equivalent | none | lossy-blocked |
| `logprobs` | generation but no Responses equivalent | none | lossy-blocked |
| `stream` | OpenAI extension | `stream` | mapped if explicit |
| `stream_options` | Chat-specific extension | none | lossy-blocked |
| `tools[].function.name` | canonical tool | `tools[].name` | mapped |
| `tools[].function.description` | canonical tool | `tools[].description` | mapped |
| `tools[].function.parameters` | canonical tool | `tools[].parameters` | mapped |
| `tools[].function.strict` | canonical tool strict | `tools[].strict` | mapped, implemented |
| assistant `tool_calls[].id` | canonical tool use id | `function_call.call_id` | mapped, implemented |
| tool message `tool_call_id` | canonical tool result id | `function_call_output.call_id` | mapped, implemented |
| `tool_choice` | canonical tool choice | `tool_choice` | mapped |
| `parallel_tool_calls` | canonical bool | `parallel_tool_calls` | native |
| `metadata` | canonical metadata | `metadata` | native |
| `response_format` | canonical response format | `text.format` | mapped |
| `reasoning_effort` | OpenAI enum | `reasoning.effort` | mapped; invalid enum blocked |
| `verbosity` | OpenAI extension | `text.verbosity` | mapped |
| `store` | OpenAI extension | `store` | extension-preserved |
| `service_tier` | OpenAI extension | `service_tier` | extension-preserved |
| `safety_identifier` | OpenAI extension | `safety_identifier` | extension-preserved |
| `prompt_cache_key` | OpenAI extension | `prompt_cache_key` | extension-preserved |
| `user` | legacy Chat user field | none | lossy-blocked |
| unknown top-level fields | source schema guard | none | unaudited |
## OpenAI Responses -> OpenAI Chat
| Responses field | Canonical handling | Chat output | Status |
| --- | --- | --- | --- |
| `model` | request identity | `model` | native |
| `input` | canonical messages/content/tool I/O | `messages` | mapped |
| `instructions` | canonical instruction/system | `messages` system/developer | mapped |
| `max_output_tokens` | generation max tokens | `max_completion_tokens` | mapped |
| `temperature` | generation | `temperature` | native |
| `top_p` | generation | `top_p` | native |
| `top_logprobs` | generation | `top_logprobs` | native |
| `metadata` | canonical metadata | `metadata` | native |
| `client_metadata` | Responses client transport metadata | none | transport-only; omitted |
| `parallel_tool_calls` | canonical bool | `parallel_tool_calls` | native |
| `text.format` | canonical response format | `response_format` | mapped |
| `text.verbosity` | Responses extension | `verbosity` | mapped |
| `tools[].type=function` | canonical tool | `tools[].type=function` | mapped |
| `tools[].name` | canonical tool | `tools[].function.name` | mapped |
| `tools[].parameters` | canonical tool | `tools[].function.parameters` | mapped |
| `tools[].strict` | canonical tool strict | `tools[].function.strict` | mapped, implemented |
| `function_call.call_id` | canonical tool use id | `tool_calls[].id` | mapped, implemented |
| `function_call_output.call_id` | canonical tool result id | tool message `tool_call_id` | mapped, implemented |
| `tools[].type=custom` | raw Responses tool | none | lossy-blocked to Chat |
| `tools[].type=web_search*` | raw Responses tool | none | lossy-blocked to Chat |
| `tool_choice` | canonical tool choice | `tool_choice` | mapped |
| `reasoning.effort` | OpenAI enum | `reasoning_effort` | mapped; invalid enum blocked |
| `reasoning.summary` | Responses-only | none | lossy-blocked |
| `reasoning.budget_tokens` | Responses-only | none | lossy-blocked |
| `stream` | Responses request transport policy | none | lossy-blocked; target stream policy is transport-owned |
| `include` | Responses-only | none | lossy-blocked; legacy emitter no longer leaks |
| `previous_response_id` | Responses-only | none | lossy-blocked; legacy emitter no longer leaks |
| `truncation` | Responses-only | none | lossy-blocked |
| `prompt` | Responses-only | none | lossy-blocked |
| `conversation` | Responses-only | none | lossy-blocked |
| `background` | Responses-only | none | lossy-blocked |
| `max_tool_calls` | Responses-only | none | lossy-blocked |
| unknown top-level fields | source schema guard | none | unaudited |
## Claude Messages <-> OpenAI Chat / Responses
Claude to OpenAI Chat, Claude to OpenAI Responses, and the reverse directions are included in the field coverage matrix. Runtime strict guards cover request root fields, provider extension namespaces, thinking/cache/tool-result hazards, and target generation-field gaps. Fields without a lossless target equivalent fail closed instead of being dropped.
High-risk fields:
| Claude field | OpenAI target risk | Required status |
| --- | --- | --- |
| `system` with cache blocks | Chat/Responses system instructions | same-format preserved; cross-format `cache_control` loss is blocked |
| `thinking` | OpenAI reasoning | Claude request-level thinking config maps to OpenAI reasoning; message-level thinking blocks are blocked for Responses |
| `cache_control` | OpenAI content/tool extensions | same-format preserved; cross-format blocked when no target equivalent exists |
| `tools[].input_schema` | OpenAI tool parameters | mapped; raw same-format schema preservation implemented |
| `tool_choice.disable_parallel_tool_use` | OpenAI `parallel_tool_calls` | mapped, implemented |
| `tool_result` multi-block content | OpenAI tool output/content | same-format preserved; cross-format to Chat/Responses is lossy-blocked |
| `metadata` | OpenAI metadata | mapped when the target has metadata |
| `container`, `inference_geo`, `service_tier` | OpenAI target has no audited equivalent | lossy-blocked unless a target-approved mapping is added |
## Gemini GenerateContent <-> OpenAI Chat / Responses / Claude
Gemini to OpenAI Chat, Gemini to OpenAI Responses, Gemini to Claude, and reverse generation paths are included in the field coverage matrix. Gemini-only request fields are preserved same-format and blocked cross-format unless the target mapping is explicitly audited.
High-risk fields:
| Gemini field | Target risk | Required status |
| --- | --- | --- |
| `contents[].parts[].thoughtSignature` | OpenAI/Claude thinking | Chat/Claude preserve; Responses cross-format is lossy-blocked |
| `tools[].functionDeclarations` | OpenAI/Claude tool schema | mapped; raw same-format `parameters` preservation implemented |
| `toolConfig.functionCallingConfig.allowedFunctionNames` | OpenAI/Claude tool choice | single-name mapping implemented; multi-name input is lossy-blocked |
| `toolConfig.functionCallingConfig.mode` | OpenAI/Claude tool choice enum | valid enum required; invalid values fail with `InvalidEnumValue` |
| `generationConfig.thinkingConfig.thinkingLevel` | OpenAI/Claude reasoning effort | low/medium/high mapping implemented; invalid values fail closed |
| `safetySettings` | OpenAI/Claude no direct equivalent | lossy-blocked |
| `cachedContent` | OpenAI/Claude no direct equivalent | lossy-blocked |
| `codeExecution` | OpenAI/Claude tool/builtin mismatch | lossy-blocked |
| `generationConfig.responseModalities` | OpenAI/Claude modality mismatch | lossy-blocked |
| `functionResponse.id` | tool result id | conversion preserves id; Gemini upstream cleanup is transport-layer edit only |
## Embedding And Rerank
Embedding and rerank request parse/emit capability and strict target guards are implemented. Provider schema fields outside these canonical conversion surfaces are marked `not-in-conversion-surface` in the field coverage matrix instead of being left implicit.
Embedding source capability:
| Source format | Parsed request shape | Canonical fields | Status |
| --- | --- | --- | --- |
| OpenAI Embedding | `model`, `input`, `encoding_format`, `dimensions`, `user`, `parameters`, `task` | OpenAI-like embedding | mapped |
| Jina Embedding | OpenAI-like plus provider extension namespace | OpenAI-like embedding | mapped |
| Doubao Embedding | OpenAI-like `model` + text `input` | OpenAI-like embedding | mapped |
| Gemini Embedding | single `content.parts[].text` or batch `requests[]` | text input, `dimensions`, `task` | mapped |
| Aliyun Multimodal Embedding | `input.contents[]`, `parameters.dimension` | text/multimodal input, `dimensions`, `parameters` | mapped |
Embedding target guards:
| Target format | Accepted canonical fields | Blocked fields/cases | Status |
| --- | --- | --- | --- |
| OpenAI Embedding | text or token input, `encoding_format`, `dimensions`, `user` | multimodal input, `task`, generic `parameters` | lossy-blocked |
| Jina Embedding | text input, `dimensions`, `task`, `parameters` | token/multimodal input, `encoding_format`, `user` | lossy-blocked |
| Gemini Embedding | text input, `dimensions`, valid `taskType` | token/multimodal input, `encoding_format`, `user`, generic `parameters`, invalid `taskType` | lossy-blocked / invalid-enum |
| Doubao Embedding | text input, `dimensions` | token/multimodal input, `encoding_format`, `user`, `task`, generic `parameters` | lossy-blocked |
| Aliyun Multimodal Embedding | text or multimodal input, `dimensions`, `parameters` | token input, `encoding_format`, `user`, `task` | lossy-blocked |
Cross-format embedding invariants:
- Embedding formats can only convert to embedding formats.
- Unknown provider-specific embedding extension namespaces are blocked cross-format unless the namespace matches the target.
- Aliyun `parameters.dimension` maps to canonical `dimensions` and is not treated as generic `parameters`.
- Gemini batch embedding parse requires every batch item to share the same model, dimensions, and task.
Rerank first pass:
| Area | Current behavior | Status |
| --- | --- | --- |
| Source formats | OpenAI Rerank and Jina Rerank parse OpenAI-like `model`, `query`, `documents`, `top_n`, `return_documents` | mapped |
| Target formats | OpenAI Rerank and Jina Rerank emit OpenAI-like rerank bodies | mapped |
| Boundary | Rerank formats can only convert to rerank formats | lossy-blocked |
| Validation | Empty query/documents and `top_n=0` fail closed | invalid-target-field |
| Extensions | Unknown provider-specific rerank extension namespaces are blocked cross-format | unsupported |
## Sync Response Conversion
Cross-format sync response conversion now validates source stop/finish/status
enums before emitting a target body. Same-format runtime response passthrough is
still outside canonical conversion.
| Source field | Target risk | Current behavior | Status |
| --- | --- | --- | --- |
| Same-format response raw stop/status fields | canonical emitters would otherwise normalize unknown enum/status to default target stop values | raw OpenAI Chat `finish_reason`, OpenAI Responses `status`, Claude `stop_reason`/`stop_sequence`, and Gemini `finishReason` are preserved through provider extension metadata | extension-preserved |
| OpenAI Chat `choices[].finish_reason` | unknown value would otherwise emit as target normal stop | valid Chat enum required; unknown values fail with `InvalidEnumValue` | invalid-enum |
| OpenAI Responses `status` | `queued`, `in_progress`, and `cancelled` have no sync target equivalent | non-terminal valid states fail with `LossyConversionBlocked`; invalid states fail with `InvalidEnumValue` | lossy-blocked / invalid-enum |
| OpenAI Responses `incomplete_details.reason=content_filter` | previously mapped to max tokens/`length` | maps to canonical content filter and emits Chat `content_filter` / Claude `content_filtered` / Gemini `SAFETY` | mapped |
| Claude `stop_reason` | unknown value would otherwise emit as target normal stop | valid known stop enum required for cross-format conversion | invalid-enum |
| Gemini `candidates[].finishReason` | known-but-unmappable reasons would otherwise emit as target normal stop | mappable safety/max/stop reasons convert; known unmappable values such as `OTHER`, `MALFORMED_FUNCTION_CALL`, `UNEXPECTED_TOOL_CALL`, `MISSING_THOUGHT_SIGNATURE`, and `MALFORMED_RESPONSE` fail with `LossyConversionBlocked`; future unknown values fail with `InvalidEnumValue` | lossy-blocked / invalid-enum |
| Canonical `Unknown` stop reason | target emitters default to normal stop values | cross-format response conversion blocks canonical unknown stop reasons | lossy-blocked |
## Stream Conversion
Sixth batch first pass is implemented for unknown event handling and runtime
same-format boundaries. Sync response finish/status parity has a first strict
pass; stream finish-reason guardrails are implemented for unknown/unmappable
terminal reasons. Stream event schema fields are covered in the field coverage
matrix; provider-by-provider fixtures cover the runtime event behavior.
Current stream behavior:
| Area | Current behavior | Status |
| --- | --- | --- |
| Provider parsers | OpenAI Chat, OpenAI Responses, Claude, and Gemini unknown stream payloads become `CanonicalStreamEvent::UnknownEvent` | mapped |
| Cross-format stream matrix | Unknown canonical stream events emit a target-format error SSE with `unsupported_stream_event` and terminate conversion | lossy-blocked |
| Stream finish reason guard | Unknown OpenAI finish reasons, unknown Claude `stop_reason`, and Gemini known-but-unmappable `finishReason` values such as `OTHER` are preserved as raw canonical finish strings, then blocked by the matrix with `unsupported_finish_reason` | lossy-blocked |
| OpenAI Responses stream target | Canonical `length` and `content_filter` terminal reasons emit `response.incomplete` with `incomplete_details.reason=max_output_tokens` or `content_filter` instead of `response.completed` | mapped |
| Terminal observer | Unknown provider stream events increment `unknown_event_count`; OpenAI Responses failed events mark terminal error state | mapped |
| Stream -> sync aggregate | Unknown OpenAI Chat, OpenAI Responses, Claude, and Gemini stream events make the runtime finalize checked path return an error and block `body_json` fallback; legacy public aggregate helpers keep `Option` compatibility | lossy-blocked |
| Runtime strict fallback guard | `UnauditedField`, `InvalidEnumValue`, `UnsupportedField`, `LossyConversionBlocked`, and `InvalidTargetField` from registry response conversion are not allowed to fall through legacy conversion helpers | lossy-blocked |
| Runtime same-format stream | Same-format stream passthrough remains outside canonical conversion; stream policy edits are transport-layer only | native |
Stream fixture coverage:
| Provider stream | Covered fixture areas |
| --- | --- |
| OpenAI Chat | sync aggregation for text, tool call IDs/names/argument deltas, finish reason, and usage; cross-format unknown finish/event blocking |
| OpenAI Responses | text snapshot de-duplication, multi-part messages, reasoning/items, function calls, image generation calls, same-family stream sync, unknown event blocking |
| Claude Messages | thinking signatures, tool input deltas, cache/usage aggregation, media emission, unknown stop/event blocking |
| Gemini GenerateContent | text/media/signature aggregation, function calls/results, safety finish mapping, unknown parts/events, and unmappable finish reason blocking |
Matrix-level interception remains the authoritative runtime path for cross-format
unknown events; direct client emitters are covered only as provider/client
building blocks.
-145
View File
@@ -1,145 +0,0 @@
# Format Enum Mapping
Status values used below:
- `native`: same semantic value exists in the target format.
- `mapped`: explicit provider-specific mapping is required.
- `blocked`: no lossless target value; conversion must fail closed.
- `preserve-same-format`: unknown/raw values are preserved only when source and target format are the same.
## OpenAI Reasoning Effort
Provider-specific types: `OpenAiChatReasoningEffort` for Chat `reasoning_effort`, and `OpenAiResponsesReasoningEffort` for Responses `reasoning.effort`. They are intentionally separate even when their current value sets overlap; a value accepted by one field is not treated as valid for the other unless that field's own enum accepts it.
| Source field | Source value | Target field | Target value | Status |
| --- | --- | --- | --- | --- |
| Chat `reasoning_effort` | `none` | Responses `reasoning.effort` | `none` | native |
| Chat `reasoning_effort` | `minimal` | Responses `reasoning.effort` | `minimal` | native when the target model supports it; blocked for GPT-5.6 |
| Chat `reasoning_effort` | `low` | Responses `reasoning.effort` | `low` | native |
| Chat `reasoning_effort` | `medium` | Responses `reasoning.effort` | `medium` | native |
| Chat `reasoning_effort` | `high` | Responses `reasoning.effort` | `high` | native |
| Chat `reasoning_effort` | `xhigh` | Responses `reasoning.effort` | `xhigh` | native |
| Chat `reasoning_effort` | `max` | Responses `reasoning.effort` | `max` | native for GPT-5.6; blocked for models that do not publish `max` |
| Responses `reasoning.effort` | `none` | Chat `reasoning_effort` | `none` | native |
| Responses `reasoning.effort` | `minimal` | Chat `reasoning_effort` | `minimal` | native when the target model supports it; blocked for GPT-5.6 |
| Responses `reasoning.effort` | `low` | Chat `reasoning_effort` | `low` | native |
| Responses `reasoning.effort` | `medium` | Chat `reasoning_effort` | `medium` | native |
| Responses `reasoning.effort` | `high` | Chat `reasoning_effort` | `high` | native |
| Responses `reasoning.effort` | `xhigh` | Chat `reasoning_effort` | `xhigh` | native |
| Responses `reasoning.effort` | `max` | Chat `reasoning_effort` | `max` | native for GPT-5.6; blocked for models that do not publish `max` |
| Responses `reasoning.summary` | any | Chat | none | blocked |
| Responses `reasoning.budget_tokens` | any | Chat | none | blocked |
Internal model directive values:
| Internal value | OpenAI Chat | OpenAI Responses | Claude output effort | Gemini thinking level | Notes |
| --- | --- | --- | --- | --- | --- |
| `none` | `none` | `none` | `low` | `low` | Budget maps to `0`. |
| `minimal` | `minimal` | `minimal` | `low` | `low` | Budget maps to `512`. |
| `low` | `low` | `low` | `low` | `low` | Budget maps to `1280`. |
| `medium` | `medium` | `medium` | `medium` | `medium` | Budget maps to `2048`. |
| `high` | `high` | `high` | `high` | `high` | Budget maps to `4096`. |
| `xhigh` | `xhigh` | `xhigh` | `xhigh` | `high` | Budget maps to `8192`. |
| `max` | `max` for GPT-5.6 | `max` for GPT-5.6 | `max` | `high` | OpenAI emission is capability-gated by the resolved model. |
GPT-5.6 (`gpt-5.6`, `gpt-5.6-sol`, `gpt-5.6-terra`, and `gpt-5.6-luna`) publishes `none`, `low`, `medium`, `high`, `xhigh`, and `max`, and does not support `minimal`. Additional non-empty effort values advertised by a model are preserved verbatim across OpenAI Chat and Responses conversion. `ultra` is a Codex client preset that resolves to `max` before transmission and is not an OpenAI wire effort. Known effort capabilities are validated against the resolved provider model, so aliases mapped to GPT-5.6 receive the GPT-5.6 contract while concrete model families keep their published constraints.
## Tool Choice
| Canonical | OpenAI Chat | OpenAI Responses | Claude Messages | Gemini GenerateContent |
| --- | --- | --- | --- | --- |
| auto | `"auto"` | `"auto"` | `{"type":"auto"}` | unset / function calling config auto |
| none | `"none"` | `"none"` | `{"type":"none"}` | mode none |
| required | `"required"` | `"required"` | `{"type":"any"}` | mode any |
| named function | `{"type":"function","function":{"name":...}}` | `{"type":"function","name":...}` | `{"type":"tool","name":...}` | allowed function name |
Implemented guardrails:
- Claude `output_config.effort=max` maps to OpenAI `xhigh`; unknown Claude effort enums are blocked cross-format.
- Claude `tool_choice.disable_parallel_tool_use` maps inversely to OpenAI `parallel_tool_calls`.
- Gemini `allowedFunctionNames` maps to canonical named tool choice and emits back to `allowedFunctionNames`.
- Gemini `thinkingLevel` maps `low|medium|high` to OpenAI reasoning effort `low|medium|high`; unknown values are blocked cross-format.
- Responses `custom`, `web_search*`, and other built-in tools are blocked when converting to Chat unless a target raw passthrough is explicitly added.
## Tool Definition Kind
| Source | Target | Mapping |
| --- | --- | --- |
| OpenAI Chat `tools[].function.name` | Responses `tools[].name` | mapped |
| OpenAI Chat `tools[].function.parameters` | Responses `tools[].parameters` | mapped |
| OpenAI Chat `tools[].function.strict` | Responses `tools[].strict` | mapped, implemented |
| Responses `tools[].strict` | Chat `tools[].function.strict` | mapped, implemented |
| OpenAI Chat assistant `tool_calls[].id` | Responses `function_call.call_id` | mapped, implemented |
| OpenAI Chat tool `tool_call_id` | Responses `function_call_output.call_id` | mapped, implemented |
| Responses `function_call.call_id` | Chat `tool_calls[].id` | mapped, implemented |
| Responses `function_call_output.call_id` | Chat tool `tool_call_id` | mapped, implemented |
| Gemini `functionCall.id` | OpenAI/Claude canonical tool use id | mapped, implemented |
| Gemini `functionResponse.id` | OpenAI `tool_call_id` / Responses `call_id` / Claude `tool_use_id` | mapped, implemented |
| Claude `tools[].input_schema` | OpenAI `parameters` | mapped; raw same-format schema preserved |
| Gemini `functionDeclarations[].parameters` | OpenAI `parameters` | mapped; raw same-format schema preserved |
## Roles
| Canonical role | OpenAI Chat | OpenAI Responses | Claude Messages | Gemini |
| --- | --- | --- | --- | --- |
| system | `system` | `instructions` or system input item | top-level `system` | `systemInstruction` |
| developer | `developer` | `instructions` or developer input item | extension-preserved | systemInstruction extension |
| user | `user` | `message.role=user` | `user` | `user` |
| assistant | `assistant` | `message.role=assistant` / output item | `assistant` | `model` |
| tool | `tool` | `function_call_output` | `tool_result` inside user message | `functionResponse` |
Known lossy risks:
- Multiple system/developer instruction ordering needs full golden fixtures.
- Provider-specific role extensions must be preserved in same-format roundtrip and blocked cross-format if no target equivalent exists.
## Finish Reasons
| Canonical | OpenAI Chat | OpenAI Responses | Claude | Gemini |
| --- | --- | --- | --- | --- |
| stop | `stop` | completed output | `end_turn` | `STOP` |
| length | `length` | `status=incomplete`, `incomplete_details.reason=max_output_tokens` | `max_tokens` | `MAX_TOKENS` |
| tool calls | `tool_calls` | output contains `function_call` | `tool_use` | `functionCall` part, usually with `STOP` |
| content filter/safety | `content_filter` | `status=incomplete`, `incomplete_details.reason=content_filter` | `content_filtered` or refusal-compatible stops | `SAFETY`, `RECITATION`, `LANGUAGE`, `BLOCKLIST`, `PROHIBITED_CONTENT`, `SPII`, image safety/recitation stops |
| unknown | preserve-same-format | preserve-same-format | preserve-same-format | preserve-same-format |
Implemented response guardrails:
- Cross-format sync response conversion validates source finish/status enums before emitting the target response.
- Same-format canonical response roundtrip preserves raw OpenAI Chat `finish_reason`, OpenAI Responses `status`, Claude `stop_reason`, and Gemini `finishReason` values through provider extension metadata.
- OpenAI Chat unknown `choices[].finish_reason` fails with `InvalidEnumValue`.
- OpenAI Responses non-terminal `status` values (`queued`, `in_progress`, `cancelled`) are valid provider states but are blocked for sync response conversion because target sync formats cannot represent them losslessly.
- Runtime sync finalize does not fall back to legacy conversion when registry response conversion reports strict errors such as invalid enums, unsupported fields, lossy blocks, or invalid target fields.
- Stream terminal reasons now follow the same strict policy: unknown OpenAI / Claude raw finish reasons and Gemini known-but-unmappable values such as `OTHER` surface as `unsupported_finish_reason`, while OpenAI Responses `length` and `content_filter` stream finals emit `response.incomplete`.
- Gemini known but unmappable finish reasons such as `OTHER`, `MALFORMED_FUNCTION_CALL`, `UNEXPECTED_TOOL_CALL`, `MISSING_THOUGHT_SIGNATURE`, and `MALFORMED_RESPONSE` are blocked with `LossyConversionBlocked`; unknown future Gemini values fail with `InvalidEnumValue`.
- Stream finish reason guardrails are covered with provider-specific fixtures for usage, tool calls, reasoning signatures, media, and unknown payloads. Unknown provider stream events are not mapped as finish reasons; cross-format runtime conversion emits a target-format `unsupported_stream_event` error and terminates.
## Embedding Task Types
Provider-specific type: `GeminiEmbeddingTaskType`.
Gemini embedding task values are stored in canonical `embedding.task` only after
source parsing. They are emitted to Gemini as `taskType` and validated before
cross-format conversion to a Gemini target.
| Canonical task input | Gemini `taskType` output | Status |
| --- | --- | --- |
| `QUERY` | `RETRIEVAL_QUERY` | mapped alias |
| `RETRIEVAL_QUERY` | `RETRIEVAL_QUERY` | native |
| `DOCUMENT` | `RETRIEVAL_DOCUMENT` | mapped alias |
| `RETRIEVAL_DOCUMENT` | `RETRIEVAL_DOCUMENT` | native |
| `TEXT_MATCHING` | `SEMANTIC_SIMILARITY` | mapped alias |
| `SEMANTIC_SIMILARITY` | `SEMANTIC_SIMILARITY` | native |
| `CLASSIFICATION` | `CLASSIFICATION` | native |
| `CLUSTERING` | `CLUSTERING` | native |
| `QUESTION_ANSWERING` | `QUESTION_ANSWERING` | native |
| `FACT_VERIFICATION` | `FACT_VERIFICATION` | native |
| `CODE_RETRIEVAL_QUERY` | `CODE_RETRIEVAL_QUERY` | native |
| `TASK_TYPE_UNSPECIFIED` | `TASK_TYPE_UNSPECIFIED` | native |
| unknown value | none | blocked with `InvalidEnumValue` when targeting Gemini |
Cross-provider rules:
- Gemini `taskType` to OpenAI Embedding is blocked because OpenAI has no equivalent task field.
- Gemini `taskType` to Doubao/Aliyun is blocked for the same reason.
- Jina `task` may carry through canonical and emit as Jina `task`; when targeting Gemini it must match the valid Gemini task set above.
File diff suppressed because it is too large Load Diff
-85
View File
@@ -1,85 +0,0 @@
# Format Passthrough Contract
Last audited: 2026-06-03
This document defines the boundary between runtime passthrough, canonical roundtrip tests, and cross-format conversion.
## Runtime Same-Format Path
Runtime same-format provider paths must not call canonical conversion.
Current implementation:
- `crates/aether-provider/transport/src/same_format_provider/mod.rs` checks `api_format_alias_matches(client_api_format, provider_api_format)`.
- When formats match, the provider body is built by copying the parsed JSON object field-for-field.
- When formats differ, the provider body is built through `aether_ai_formats::convert_request_pure`.
- Model override, body rules, model directives, Claude Code sanitization, Gemini function-response id stripping, and stream policy are applied only after the passthrough/conversion branch in provider transport.
Important limitation:
- The current transport helper receives `body_json: &serde_json::Value`, not raw request bytes. It therefore guarantees no canonical conversion and JSON value preservation at this layer, but it cannot preserve original whitespace or object key order by itself.
- True byte-level passthrough for requests with no transport edits requires a higher-level raw-body path that can forward the original bytes directly. Until that raw-body plumbing exists, tests should assert "conversion module not called" and JSON value equivalence for this helper, not byte-for-byte serialization equivalence.
Provider schema drift does not change this rule. If OpenAI, Claude, or Gemini add a new field, same-format runtime routing must still forward it as part of the original provider body. The schema inventory and field coverage matrix are audit aids, not the runtime allowlist for same-format traffic.
## Canonical Same-Format Roundtrip
Canonical same-format roundtrip is only a test/audit mode:
```text
source format -> Canonical -> same source format
```
Required behavior:
- JSON-normalized equality, ignoring object field order and whitespace.
- Field values, array order, unknown fields, extension namespaces, and unknown enum strings must be preserved.
- This path may parse and emit; it is not the runtime path.
- Unknown provider fields are carried in provider extension namespaces and replayed when emitting the same provider format.
## Cross-Format Conversion
Cross-format conversion is strict:
```text
source format -> Canonical -> target format
```
Required behavior:
- Emit only fields valid for the target provider format.
- Map provider-specific enum values through explicit provider enum types.
- Preserve source fields only when the target has an equivalent field or documented extension passthrough.
- Fail closed with `FormatError::UnauditedField`, `FormatError::LossyConversionBlocked`, `FormatError::UnsupportedField`, `FormatError::InvalidEnumValue`, or `FormatError::InvalidTargetField` when no lossless mapping exists.
- Do not use `None` or silent omission to represent conversion failure.
- Newly added provider fields follow the same rule as other unknown fields: preserve same-format, fail closed cross-format with `UnauditedField`. A code change is required only when Aether intentionally supports a new cross-format semantic mapping.
## Pure Conversion Interface
Pure conversion lives in `crates/aether-ai/formats` and is limited to:
- parse
- emit
- provider-specific field/enum mapping
- `ConversionReport`
Pure conversion must not:
- override `model`
- add, remove, or force `stream`
- apply body rules
- apply model directives
- read the original request body to patch missing target fields
- perform provider transport policy edits
Current pure entrypoints:
- `parse_request_pure`
- `emit_request_pure`
- `convert_request_pure`
- `convert_request_pure_with_context`
- `parse_response_pure`
- `emit_response_pure`
- `convert_response_pure`
`convert_request` and `convert_response` remain legacy wrappers for existing callers that still need mapped model/report-context behavior during migration.
-548
View File
@@ -1,548 +0,0 @@
#!/usr/bin/env python3
"""Generate the provider schema field coverage matrix.
The input inventory is docs/api/provider-interface-definitions.md. Existing
coverage rows are reused so audited status/notes survive regeneration. Newly
introduced provider fields get conservative same-format/native and cross-format
fail-closed defaults until a human audits whether they deserve an explicit
mapping.
"""
from __future__ import annotations
import argparse
import dataclasses
from collections import Counter, defaultdict
from pathlib import Path
from typing import Iterable
ROOT = Path(__file__).resolve().parents[2]
DEFAULT_DEFINITIONS = ROOT / "docs/api/provider-interface-definitions.md"
DEFAULT_MATRIX = ROOT / "docs/api/format-field-coverage-matrix.md"
@dataclasses.dataclass(frozen=True)
class SourceField:
provider: str
schema: str
field: str
required: str
field_type: str
@dataclasses.dataclass(frozen=True)
class CoverageStatus:
surface: str
same_format_runtime: str
canonical_roundtrip: str
cross_format: str
notes: str
OPENAI_CHAT_MAPPED = {
"model",
"messages",
"max_tokens",
"max_completion_tokens",
"temperature",
"top_p",
"top_logprobs",
"tools",
"tool_choice",
"parallel_tool_calls",
"metadata",
"response_format",
"reasoning_effort",
"verbosity",
"store",
"service_tier",
"safety_identifier",
"prompt_cache_key",
"prompt_cache_retention",
"stream",
}
OPENAI_CHAT_BLOCKED = {
"n",
"stop",
"presence_penalty",
"frequency_penalty",
"seed",
"logprobs",
"stream_options",
"user",
"function_call",
"functions",
"logit_bias",
"modalities",
"prediction",
"audio",
"web_search_options",
}
OPENAI_RESPONSES_MAPPED = {
"model",
"input",
"instructions",
"max_output_tokens",
"temperature",
"top_p",
"top_logprobs",
"metadata",
"parallel_tool_calls",
"text",
"tools",
"tool_choice",
"reasoning",
"store",
"service_tier",
"safety_identifier",
"prompt_cache_key",
"prompt_cache_retention",
}
OPENAI_RESPONSES_BLOCKED = {
"include",
"previous_response_id",
"truncation",
"prompt",
"conversation",
"background",
"max_tool_calls",
"user",
"context_management",
"stream",
"stream_options",
}
CLAUDE_MAPPED_FIELDS = {
"id",
"type",
"role",
"text",
"content",
"source",
"name",
"description",
"input",
"input_schema",
"messages",
"model",
"max_tokens",
"system",
"temperature",
"top_p",
"top_k",
"stop_sequences",
"tool_choice",
"tools",
"metadata",
"thinking",
"output_config",
"usage",
"stop_reason",
"stop_sequence",
}
CLAUDE_PROVIDER_ONLY_FIELDS = {
"cache_control",
"container",
"inference_geo",
"service_tier",
"allowed_callers",
"allowed_domains",
"blocked_domains",
"defer_loading",
"max_uses",
"strict",
"user_location",
"citations",
"context",
"title",
"file_id",
"document_index",
"document_title",
"cited_text",
"caller",
}
def split_markdown_row(line: str) -> list[str]:
cells: list[str] = []
current: list[str] = []
escaped = False
for char in line:
if char == "|" and not escaped:
cells.append("".join(current).strip())
current.clear()
else:
current.append(char)
escaped = char == "\\" and not escaped
if escaped and char != "\\":
escaped = False
cells.append("".join(current).strip())
return cells
def strip_markdown_code(value: str) -> str:
value = value.strip()
if value.startswith("`") and value.endswith("`"):
value = value[1:-1]
return value.replace("\\|", "|")
def escape_markdown_cell(value: str) -> str:
return value.replace("|", "\\|")
def parse_schema_heading(line: str) -> str | None:
if not line.startswith("### `"):
return None
rest = line[len("### `") :]
schema, _, _ = rest.partition("`")
return schema or None
def parse_provider_definition_fields(definitions: str) -> list[SourceField]:
provider: str | None = None
schema: str | None = None
fields: list[SourceField] = []
for line in definitions.splitlines():
if line.startswith("## "):
if "OpenAI Schema" in line:
provider = "OpenAI"
elif "Claude / Anthropic TypeScript" in line:
provider = "Claude"
elif "Gemini Schema" in line:
provider = "Gemini"
else:
provider = None
schema = None
continue
if provider is None:
continue
if heading := parse_schema_heading(line):
schema = heading
continue
if schema is None or not line.startswith("| `"):
continue
cells = split_markdown_row(line)
if len(cells) < 4 or cells[2] not in {"是", "否"}:
continue
fields.append(
SourceField(
provider=provider,
schema=schema,
field=strip_markdown_code(cells[1]),
required=cells[2],
field_type=strip_markdown_code(cells[3]),
)
)
return fields
def parse_existing_coverage(
matrix: str,
) -> tuple[dict[tuple[str, str, str], CoverageStatus], dict[tuple[str, str], list[CoverageStatus]]]:
existing: dict[tuple[str, str, str], CoverageStatus] = {}
profiles: dict[tuple[str, str], list[CoverageStatus]] = defaultdict(list)
for line in matrix.splitlines():
if not line.startswith("| "):
continue
cells = split_markdown_row(line)
if len(cells) < 11 or cells[1] not in {"OpenAI", "Claude", "Gemini"}:
continue
status = CoverageStatus(
surface=cells[6],
same_format_runtime=cells[7],
canonical_roundtrip=cells[8],
cross_format=cells[9],
notes=cells[10],
)
provider = cells[1]
schema = strip_markdown_code(cells[2])
field = strip_markdown_code(cells[3])
existing[(provider, schema, field)] = status
profiles[(provider, schema)].append(status)
return existing, profiles
def most_common(values: Iterable[str]) -> str | None:
values = list(values)
if not values:
return None
return Counter(values).most_common(1)[0][0]
def openai_surface(schema: str) -> str:
if "CreateChatCompletion" in schema or "ChatCompletion" in schema:
return "openai:chat standard"
if "CreateEmbedding" in schema or "Embedding" in schema:
return "openai:embedding"
if any(token in schema for token in ("CreateImage", "EditImage", "Image", "Images")):
return "openai:image native-only"
if any(token in schema for token in ("Compact", "Compaction")):
return "openai:responses:compact native-only"
if any(
token in schema
for token in (
"Response",
"Input",
"Output",
"Tool",
"Reasoning",
"WebSearch",
"FileSearch",
"Computer",
"MCP",
"CodeInterpreter",
"Function",
"Custom",
"EasyInput",
"Prompt",
"Conversation",
"Annotation",
"Citation",
"LogProb",
"TopLogProb",
"Metadata",
"ServiceTier",
"Verbosity",
"TextResponse",
"ResponseFormat",
"Include",
"Modalities",
"ParallelToolCalls",
"StopConfiguration",
)
):
return "openai:responses standard"
return "openai auxiliary / not-in-conversion-surface"
def openai_default_status(field: SourceField, profile: list[CoverageStatus]) -> CoverageStatus:
surface = most_common(status.surface for status in profile) or openai_surface(field.schema)
if "not-in-conversion-surface" in surface or "native-only" in surface:
return CoverageStatus(
surface=surface,
same_format_runtime="native",
canonical_roundtrip="not-in-conversion-surface",
cross_format="not-in-conversion-surface",
notes="not part of current canonical cross-format conversion; same-format runtime path remains provider-native when routed directly",
)
if field.schema == "CreateChatCompletionRequest":
if field.field in OPENAI_CHAT_MAPPED:
return CoverageStatus(
surface=surface,
same_format_runtime="native",
canonical_roundtrip="mapped",
cross_format="mapped",
notes="Chat request field maps provider-specifically; target-incompatible cases fail closed",
)
if field.field in OPENAI_CHAT_BLOCKED:
return CoverageStatus(
surface=surface,
same_format_runtime="native",
canonical_roundtrip="extension-preserved",
cross_format="lossy-blocked",
notes="Chat-only or provider-specific field has no audited lossless target equivalent",
)
if field.schema == "CreateResponse":
if field.field in OPENAI_RESPONSES_MAPPED:
return CoverageStatus(
surface=surface,
same_format_runtime="native",
canonical_roundtrip="mapped",
cross_format="mapped",
notes="Responses request field maps provider-specifically; target-incompatible cases fail closed",
)
if field.field in OPENAI_RESPONSES_BLOCKED:
return CoverageStatus(
surface=surface,
same_format_runtime="native",
canonical_roundtrip="extension-preserved",
cross_format="lossy-blocked",
notes="Responses-only field has no audited lossless Chat/Claude/Gemini target equivalent",
)
if profile:
cross_format = most_common(status.cross_format for status in profile) or "lossy-blocked"
return CoverageStatus(
surface=surface,
same_format_runtime="native",
canonical_roundtrip=most_common(status.canonical_roundtrip for status in profile)
or "extension-preserved",
cross_format=cross_format,
notes=next(
(status.notes for status in profile if status.cross_format == cross_format),
"schema-level handling inherited from audited sibling fields",
),
)
return CoverageStatus(
surface=surface,
same_format_runtime="native",
canonical_roundtrip="extension-preserved",
cross_format="lossy-blocked",
notes="OpenAI documented field is preserved same-format; cross-format requires explicit target mapping or fails closed",
)
def claude_default_status(field: SourceField) -> CoverageStatus:
if "CountTokens" in field.schema:
return CoverageStatus(
surface="claude:messages/count_tokens native-only",
same_format_runtime="native",
canonical_roundtrip="not-in-conversion-surface",
cross_format="not-in-conversion-surface",
notes="count_tokens schemas are provider-native and outside canonical generation conversion",
)
if field.field in CLAUDE_MAPPED_FIELDS:
return CoverageStatus(
surface="claude:messages standard",
same_format_runtime="native",
canonical_roundtrip="mapped",
cross_format="mapped/lossy-blocked",
notes="Claude field maps where canonical and target support an equivalent; otherwise conversion fails closed",
)
if (
field.field in CLAUDE_PROVIDER_ONLY_FIELDS
or field.field.endswith("_tokens_details")
or "cache" in field.field
):
return CoverageStatus(
surface="claude:messages standard",
same_format_runtime="native",
canonical_roundtrip="extension-preserved",
cross_format="lossy-blocked",
notes="Claude provider-specific field is preserved same-format and blocked cross-format without an audited target equivalent",
)
return CoverageStatus(
surface="claude:messages standard",
same_format_runtime="native",
canonical_roundtrip="extension-preserved",
cross_format="lossy-blocked",
notes="Claude nested/provider-specific field is same-format preserved; cross-format requires explicit mapping or fails closed",
)
def gemini_default_status(field: SourceField, profile: list[CoverageStatus]) -> CoverageStatus:
if profile:
cross_format = most_common(status.cross_format for status in profile) or "lossy-blocked"
return CoverageStatus(
surface=most_common(status.surface for status in profile)
or "gemini:generate_content standard",
same_format_runtime="native",
canonical_roundtrip=most_common(status.canonical_roundtrip for status in profile)
or "extension-preserved",
cross_format=cross_format,
notes=next(
(status.notes for status in profile if status.cross_format == cross_format),
"Gemini field follows schema-level handling",
),
)
return CoverageStatus(
surface="gemini:generate_content standard",
same_format_runtime="native",
canonical_roundtrip="extension-preserved",
cross_format="lossy-blocked",
notes="Gemini documented field is preserved same-format; cross-format requires explicit mapping or fails closed",
)
def default_status(field: SourceField, profile: list[CoverageStatus]) -> CoverageStatus:
if field.provider == "OpenAI":
return openai_default_status(field, profile)
if field.provider == "Claude":
return claude_default_status(field)
if field.provider == "Gemini":
return gemini_default_status(field, profile)
raise ValueError(f"unsupported provider: {field.provider}")
def render_matrix(
fields: list[SourceField],
existing: dict[tuple[str, str, str], CoverageStatus],
profiles: dict[tuple[str, str], list[CoverageStatus]],
) -> str:
rows: list[str] = [
"# Format Field Coverage Matrix",
"",
"Last generated: 2026-06-03",
"",
"This file is generated from the schema inventory in `docs/api/provider-interface-definitions.md` and gives every documented schema field an explicit handling status. “处理到” here means the field is either mapped, preserved in same-format paths, rejected with a structured fail-closed error, or explicitly outside the current conversion surface. It does not mean every field can be cross-format converted.",
"",
"Provider schema updates do not require immediate conversion-code changes for runtime safety. Same-format runtime paths bypass canonical conversion, and same-format canonical roundtrip preserves provider extension fields. Cross-format conversion only enables fields with an audited semantic mapping; newly discovered or unknown provider fields default to structured fail-closed behavior until mapped.",
"",
"Regenerate with: `python3 docs/api/generate_format_field_coverage.py`.",
"",
"Statuses used in this matrix: `native`, `mapped`, `mapped/lossy-blocked`, `extension-preserved`, `unaudited`, `unsupported`, `invalid-enum`, `lossy-blocked`, `not-in-conversion-surface`.",
"",
"| Provider | Schema | Field | Required | Type | Surface | Same-Format Runtime | Canonical Roundtrip | Cross-Format | Notes |",
"| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |",
]
for field in fields:
status = existing.get(
(field.provider, field.schema, field.field),
default_status(field, profiles[(field.provider, field.schema)]),
)
rows.append(
"| "
+ " | ".join(
[
field.provider,
f"`{escape_markdown_cell(field.schema)}`",
f"`{escape_markdown_cell(field.field)}`",
field.required,
f"`{escape_markdown_cell(field.field_type)}`",
escape_markdown_cell(status.surface),
escape_markdown_cell(status.same_format_runtime),
escape_markdown_cell(status.canonical_roundtrip),
escape_markdown_cell(status.cross_format),
escape_markdown_cell(status.notes),
]
)
+ " |"
)
rows.extend(["", f"Total covered schema fields: {len(fields)}."])
return "\n".join(rows) + "\n"
def main() -> int:
parser = argparse.ArgumentParser()
parser.add_argument("--definitions", type=Path, default=DEFAULT_DEFINITIONS)
parser.add_argument("--matrix", type=Path, default=DEFAULT_MATRIX)
parser.add_argument("--check", action="store_true")
args = parser.parse_args()
definitions = args.definitions.read_text()
current_matrix = args.matrix.read_text() if args.matrix.exists() else ""
fields = parse_provider_definition_fields(definitions)
existing, profiles = parse_existing_coverage(current_matrix)
next_matrix = render_matrix(fields, existing, profiles)
if args.check:
if current_matrix != next_matrix:
print(
f"{args.matrix} is not up to date; run "
"`python3 docs/api/generate_format_field_coverage.py`",
)
return 1
return 0
args.matrix.write_text(next_matrix)
print(f"wrote {len(fields)} field coverage rows to {args.matrix}")
return 0
if __name__ == "__main__":
raise SystemExit(main())
-53
View File
@@ -1,53 +0,0 @@
# 提供商管理的端点健康度
以下管理接口使用相同的健康度汇总规则:
- `GET /api/admin/providers/summary`
- `GET /api/admin/providers/{provider_id}/summary`
## 统计规则
沿用 `v0.7.13`(`535ee098c`)的默认健康规则:已启用密钥缺少该格式的健康记录时,
按 `1.0` 参与统计,而不是要求先有一次请求或探测才能显示健康。
- 仅统计启用端点下、支持该端点 API 格式的启用密钥。
- 每个密钥读取其 `health_by_format[api_format].health_score`,不跨格式借用分数。
- 对上述启用密钥求算术平均;缺少有效分数的密钥沿用旧版默认值 `1.0`。
- 分数范围为 `0` 到 `1`,沿用调度器的分数读取与范围约束。
- 停用端点或没有启用密钥时,端点的 `health_score` 返回 `null`。
- `avg_health_score` 仅平均有启用密钥的启用端点,包含按默认值计算的端点;没有此类端点时返回 `null`。
- `unhealthy_endpoints` 仅统计上述端点中健康度低于 `0.5` 的数量,不把未知状态算作故障。
`total_keys` 和 `active_keys` 仍反映密钥配置数量,不因缺少健康数据而减少。
Codex、Kiro、Gemini CLI、Antigravity 等固定提供商的账号按照各提供商的认证规则,
继承其启用端点的 API 格式;账号的 `api_formats` 为 `null` 或空数组,不代表没有配置账号。
继承格式决定账号归属;缺少对应格式的分数时同样使用默认值 `1.0`,不借用其他格式的异常分数。
## 数据读取
摘要从密钥的轻量投影读取 API 格式、启用状态和 `health_by_format`。
PostgreSQL 投影中的凭据字段使用 `summary` / `{}` 等脱敏占位值,并非真实密文,
因此摘要读取不执行凭据解密、认证或迁移。完整密钥读取仍保留原有的凭据安全校验。
若将这些占位值送入凭据校验,读取会失败,旧的摘要聚合还会将其当作空密钥列表,
导致已配置密钥的端点也被错误显示为灰色。默认健康分数只能用于成功读取的启用密钥,
不能用于掩盖查询或凭据投影错误。
提供商、端点或密钥摘要查询失败时,接口返回 `503`,不能将失败当作空列表并返回零账号。
单个提供商确实不存在时仍返回 `404`。页面刷新失败保留已有列表,并显示加载错误。
## 页面展示
桌面表格、网格卡片和手机卡片使用相同规则:
- 有健康数据时显示百分比;有效的零分显示 `0%`。
- 有启用密钥、但这些密钥尚无该格式的健康记录时,显示绿色 `100%`,与 `v0.7.13` 一致。
- 端点停用、未配置密钥或没有启用密钥时显示灰色占位条和对应状态提示。
账号详情保留原有默认 `100%` 的规则。已有观测仍取各格式最低分;端点分数则只聚合对应格式,
两者统计范围不同,不要求百分比完全相等。调度器原有的缺省健康策略不变。
此分数是密钥当前健康状态的聚合,包含默认健康值,不是某个时间窗口内的请求成功率,
也不表示已经执行过主动探测。相比 `v0.7.13`,仍保留停用账号/端点不参与健康聚合、
凭据脱敏与查询失败显式报错等修复,不整体回退旧版代码。
File diff suppressed because it is too large Load Diff
-58
View File
@@ -1,58 +0,0 @@
# Rerank API
Aether exposes an OpenAI-compatible rerank surface at `POST /v1/rerank` and can route it to providers configured as `openai:rerank` or `jina:rerank`.
## Request
```http
POST /v1/rerank
Authorization: Bearer <aether-api-key>
Content-Type: application/json
```
```json
{
"model": "bge-reranker-base",
"query": "What document discusses gateway routing?",
"documents": [
"Aether routes public AI requests through the Rust gateway.",
"This document discusses unrelated content."
],
"top_n": 1,
"return_documents": true
}
```
Fields:
| Field | Required | Notes |
| --- | --- | --- |
| `model` | Yes | Aether global model name. |
| `query` | Yes | Non-empty query string. |
| `documents` | Yes | Non-empty array of strings or provider-native document objects. |
| `top_n` | No | Positive integer. |
| `return_documents` | No | Provider-compatible flag for including matched documents. |
## Response
Aether forwards the provider JSON response. OpenAI-compatible and Jina-compatible rerank providers commonly return `results[]`:
```json
{
"model": "bge-reranker-base",
"results": [
{
"index": 0,
"relevance_score": 0.98,
"document": {
"text": "Aether routes public AI requests through the Rust gateway."
}
}
],
"usage": {
"total_tokens": 32
}
}
```
Rerank requests must be JSON and do not support `stream` or chat `messages` payloads.
@@ -1,39 +0,0 @@
# 调度策略:取消请求立即打断
管理入口:**调度策略配置 → 系统配置 → 取消请求立即打断**。
配置保存在当前策略的 `config_json.default_policy.cancel_on_client_disconnect`,默认 `false`。
旧策略缺失此字段也按关闭处理;不需要数据库结构迁移,不读取同名全局系统设置。
策略选定后,请求沿用该次解析的配置,不因管理员随后修改策略而改变断连处理。
| 配置 | 客户端取消/断连时 | 计费 |
| --- | --- | --- |
| 关闭(默认) | 已选定策略的请求继续执行,后台读取响应直到正常完成或原有超时/上游错误 | 按实际完成结果正常结算 |
| 开启 | 中止仍在执行的请求、停止读取上游响应 | Token、缓存 Token 和图片产出费用不收取;配置了 `price_per_request` 的请求保留一次请求费用及对应倍率 |
已经取得上游终态的请求不会因最后一跳投递失败而撤销已完成的结算。
取消按次收费的记录仍显示 `cancelled`/499,但计费状态可为 `settled`;不要仅凭请求状态判断是否收费。
## 执行边界
- HTTP 同步、SSE、同格式直通和 CF 心跳响应共用请求生命周期保护。
- Responses WebSocket 断连后只完成当前进行中的 turn,不无限维持空闲连接;开启开关则直接终止当前 turn。
- 鉴权、请求体接收等尚未选定策略的阶段仍可直接取消。
- Live/Realtime 长连接会话关闭、异步任务显式取消不是有限 HTTP/Responses 请求的断连续跑,不转成后台常驻会话。
- 原有上游总超时、首字节超时及故障处理仍有效;这里不创建持久化后台作业,进程退出不能保证继续执行。
## 代码调整
- `request_lifecycle.rs` 将 HTTP 请求 Future 和响应 Body 的所有权与客户端连接解耦。正常连接保留原来的逐帧路径、响应头、长度提示和 trailers;仅断连时启动后台接管,逐帧丢弃待发送内容,不聚合完整响应。
- 请求准入凭证随原 Future/Body 保留至完成,防止断连提前释放并发额度;诊断上下文同时保留。
- CF 心跳后台执行显式监听响应接收端关闭,避免打开开关后仍继续执行。
- Responses WebSocket 将“客户端已断开”和“上游已完成”分别处理,复用既有终态观察、超时和结算逻辑。
- 计费统一在 billing enrichment 中计算取消请求的单次费用。`cancelled_request_fee` 由服务端计算后标记,贯穿审计持久化、钱包结算及套餐成本预留结算,避免有费用却仍写为 `void` 或释放成本预留。
## 回归覆盖
- 新旧配置默认关闭、布尔校验、策略保存及模型级调度编辑保留开关。
- 响应头前断连、首帧前/首帧后断连、立即取消、并发额度、诊断信息、响应头/trailers 透传。
- 真实流式执行链路的完成/取消 usage 与候选状态,以及心跳取消。
- 取消时按次/按 Token/混合定价、图片请求单次费、倍率、审计状态、钱包及成本预留结算。
- Responses WebSocket 使用真实网关和临时 PostgreSQL 验证断连继续完成、开启后不收费,以及按次费用实际结算。
@@ -1,190 +0,0 @@
# Codex Responses WebSocket probe
`aether-codex-ws-probe` is a P0 compatibility probe for a Codex-compatible
Responses WebSocket upstream. It verifies two sequential `response.create`
warmups on one socket, with the second request continuing from the first
response ID.
The command, environment variables, JSON report shape, and Codex-specific
handshake headers remain stable. It now shares only the protocol-driving core
with the separate [OpenAI Responses WebSocket probe](openai-responses-websocket-probe.md);
the two probes intentionally retain independent authentication profiles and
provider-specific assertions.
The probe is intentionally not a production proxy. It does not persist,
refresh, log, or print credentials, account IDs, response IDs, request bodies,
or response bodies.
## Prerequisites
Use a dedicated, non-production Codex test account. Rotate any credential that
has been pasted into a chat, terminal history, issue, or source file before
using this probe.
Set these values only in the process environment or your secret manager:
```bash
export AETHER_CODEX_WS_PROBE_URL='wss://your-codex-upstream.example/backend-api/codex/responses'
export AETHER_CODEX_WS_PROBE_ACCESS_TOKEN='your-short-lived-access-token'
export AETHER_CODEX_WS_PROBE_ACCOUNT_ID='your-account-id'
export AETHER_CODEX_WS_PROBE_MODEL='your-codex-model'
```
The endpoint must use `ws://` or `wss://`, with no credentials, query string,
or fragment. The access token is accepted only through
`AETHER_CODEX_WS_PROBE_ACCESS_TOKEN`; there is deliberately no command-line
flag for it.
## Run
```bash
cargo run -p aether-gateway --bin aether-codex-ws-probe
```
Use `--url` to override only the endpoint and `--timeout-secs` to set a
per-turn receive timeout (1–120 seconds):
```bash
cargo run -p aether-gateway --bin aether-codex-ws-probe -- \
--url 'wss://your-codex-upstream.example/backend-api/codex/responses' \
--timeout-secs 30
```
The probe emits one JSON line. A successful run has
`"continuation_confirmed":true`; its event and header fields contain names
only, never values. A failure emits a stable error code such as
`"handshake_failed"`, `"upstream_error_event"`, or
`"response_id_not_observed"`.
## Interpretation
A successful probe establishes that the selected upstream accepts the
Responses WebSocket handshake and retains continuation state on one socket.
It does not establish that all Codex models, account plans, or tunnel egress
paths are supported. In particular, the current `aether-tunnel` HTTP relay
does not forward WebSocket upgrades, so a successful direct probe is a
prerequisite rather than tunnel support.
## Gateway bridge
The gateway exposes WebSocket mode at the same public Responses path:
```text
wss://<aether-gateway>/v1/responses
```
It is disabled by default per provider. In **添加提供商** or **编辑提供商**,
enable **Responses WebSocket 模式** under **功能开关** only after the selected
upstream has passed a compatible WebSocket probe. The setting takes effect for
new WebSocket connections without a gateway restart. It is available to every
provider type; candidate planning still requires a selected
`openai:responses` endpoint.
Authenticate the upgrade request with the normal Aether API key. The first
client frame must be a text JSON `response.create` containing a non-empty
`model`. Aether then applies its regular Responses candidate selection, but
accepts only an eligible, WebSocket-enabled endpoint using `openai:responses`.
It opens an upstream WebSocket using the selected provider key.
The selected provider's model mapping and request headers are applied to every
turn, along with the rest of that candidate's provider-body normalization:
model-directive patches, endpoint body rules, and the Codex body contract
(unsupported-field stripping, its HTTP `store: false` default, and
`tool_choice` defaulting). An explicitly supplied WebSocket `store` value is
restored unchanged after that HTTP-oriented normalization. A
continuation turn with a non-null `previous_response_id` is revalidated through
the current scheduler and normalized against its pinned binding. It can never
move to another provider key, and it is rejected if that exact candidate is no
longer eligible or its physical binding changed.
`store`, `previous_response_id`, and `generate` are re-applied after
normalization because they are WebSocket protocol state that the provider body
contract may otherwise rewrite or strip. `stream` and `background` are removed
because they are HTTP transport fields, not WebSocket-mode fields. Every
independent `response.create` (one without a non-null `previous_response_id`)
runs access checks and candidate planning again, even when the public model is unchanged.
It keeps the existing upstream when the same target remains eligible, or
transparently replaces the upstream between responses when the selected target
changes. Overlapping responses on one client socket remain rejected.
Each `response.create` is tracked as an independent Aether logical request:
it receives its own request/candidate identity, usage lifecycle, and terminal
audit record. `response.completed`, `response.failed`,
`response.incomplete`, `response.cancelled`, client disconnects, and upstream
transport failures all settle that turn through the existing stream reporting
path.
Example client setup:
```python
from websocket import create_connection
import json
import os
ws = create_connection(
"wss://gateway.example/v1/responses",
header=[f"Authorization: Bearer {os.environ['AETHER_API_KEY']}"],
)
ws.send(json.dumps({
"type": "response.create",
"model": "your-public-model",
"store": False,
"input": "Explain this repository.",
}))
```
### Operating limits
- Maximum frame and message size: 16 MiB.
- An idle connection must send its first `response.create` within 60 seconds.
- A connection is closed after 60 minutes; reconnect before then for long runs.
- Each `response.create` must receive its first upstream event within the
selected provider's `stream_first_byte_timeout` (30 seconds by default),
and finish within its `request_timeout` (20 minutes by default). Aether
sends `responses_websocket_first_event_timeout` or
`responses_websocket_turn_timeout` and closes the bound socket when either
deadline expires.
- Responses are sequential; no multiplexing is supported on one socket.
- Each `response.create` consumes the normal Aether user/API-key RPM budget.
- Continuations with a non-null `previous_response_id` stay on the bound
provider key. Independent turns are re-authorized and re-planned each time;
they reuse the socket only when planning selects the same physical target.
- Direct provider proxy settings are honored through the selected transport
profile. Tunnel-mode proxy nodes are not supported for this bridge yet.
### DNS and proxy behavior
The gateway's plain and browser-profile WebSocket clients share the HTTP/SSE
provider DNS resolver. Provider hostname answers are not filtered by address
range, including Fake-IP answers such as `198.18.0.0/15`. DNS lookup timeout,
answer-count limits, and rejection of empty answers still apply. Resolution
happens during connection establishment, not while building the client; DNS
failures therefore surface as upstream handshake failures rather than invalid
upstream URLs.
This policy applies only to configured provider hostnames. URL validation still
rejects credentials, fragments, and literal private/reserved IP targets (except
loopback `ws://`), and the gateway frontdoor self-loop guard remains active.
Tunnel owner-relay DNS address filtering is unchanged.
Configure an explicit HTTP(S) or SOCKS proxy on the provider when needed; these
clients do not automatically use system proxy environment variables. A proxy
connection does not trigger a separate gateway-side lookup of the provider
hostname. Existing SOCKS remote-DNS normalization remains in effect.
The [DNS egress audit](dns-egress-audit-2026-09-08.md) documents the separate
tunnel and untrusted-download policies, regression coverage, and remaining
deployment limitations; not every outbound path uses the provider DNS policy.
### Usage and logging
Usage and audit finalization now runs for every accepted `response.create`.
Existing usage body-capture and header-redaction policies apply to the resulting
records. Newly created WebSocket usage records expose `is_websocket=true`, and
the usage-record type column renders them as `WS`. For diagnosis, enable debug
logging for `aether_gateway::handlers::proxy::responses_ws`; event logs contain
only the event type and frame size, never request or response contents. Codex
quota-extension logs remain under `aether_gateway::handlers::proxy::codex_ws`.
Every WebSocket-specific log carries `transport="websocket"` and
`websocket=true`; keep `log_type` for its existing access/event/ops
classification, and render the transport flag as a `WS` label in a log viewer
if desired.
@@ -1,526 +0,0 @@
# Aether 并发设计审查
- 日期:2026-09-09
- 代码基线:`361952ada`
- 背景:RPM 增加后服务出现卡死现象。
- 范围:当前代码、默认配置、隔离本地复现。尚未取得故障实例、实际 RPM、线上配置或卡死时指标;以下是确认的代码问题及条件性风险,不代表已经确认本次生产故障根因。
- 初次审查仅新增报告;后续本地修复状态见下文。未部署、未修改生产配置、未向线上发起压测。
## 第一轮修复状态
以下问题描述及行号对应审查基线 `361952ada`,不是修复后的代码位置。
- **P1-1 已修复:** 压缩上传按声明的压缩大小预留,未知长度上传从一个额度单元开始,随缓冲容量增长申请预算;解压计入同时存活的输入、中间输出和最终输出。扩容额度不足立即返回 503,避免多个请求各持部分额度互相等待;取消和失败释放额度。请求体完整读取默认超时改为 120 秒,显式配置 0 仍可关闭。该预算覆盖读取和解压的显式缓冲,不等于整个请求生命周期或进程 RSS 上限。
- **P1-3 已修复:** SQL 改为 `FOR UPDATE OF user_plan_entitlements`,不同用户不再争抢共享套餐行,同一 entitlement 的扣费仍串行。套餐 overage 配置按当前语句快照读取,后续语句可读取已提交的配置变更。
- **P1-4 已修复:** reqwest、h2c、wreq 和 tunnel 的流统一接入空闲读取期限;执行配置 `read_ms` 优先,否则使用 `AETHER_GATEWAY_UPSTREAM_STREAM_IDLE_TIMEOUT_MS`,默认 300 秒,显式 0 可关闭。首包期限独立保留,网关 keepalive 和空数据帧不重置空闲计时;已收到成功终止事件的请求不会因随后空闲被改记为失败。
- **后续待处理:** P1-2 Redis 调度扫描、P1-5 响应捕获全局预算、P1-6 审计压缩及 P1-7 长期额度聚合和 SQL 期限,本轮未修改。
本地验证:
- `cargo test -p aether-gateway-frontdoor`:34 项通过。
- 网关请求体、解压、超时、流结束及模块边界的针对性回归:73 项通过;另跑 Anthropic 原生流及直通兼容性回归:24 项通过。
- `cargo test -p aether-data-postgres --lib settlement::tests`:6 项通过,1 项需要数据库的测试默认忽略;该测试另在隔离 PostgreSQL 14.17 中显式执行通过,覆盖不同用户并行、同用户阻塞、扣费余额及配置并发更新,临时实例已停止。
- 修改文件格式检查和 `git diff --check` 通过。未执行全工作区测试或阶梯吞吐压测,尚不能据此给出修复后的 RPM 容量。
## 第二轮修复状态
- **P1-2 已处理调度中的管理扫描和无关指标读取:** 增加调度专用运行态入口,跳过全池 sticky 会话扫描、会话计数及 cooldown TTL;sticky 直达只查询当前绑定。成本窗口仅在成本限额、`cost_first` 或 `quota_balanced` 启用时读取,延迟窗口仅在 `latency_first` 启用时读取。默认 64 个候选且未启用这些策略时,消除原先 128 次历史窗口查询;这是代码路径比较,不是压测吞吐结论。管理查询仍保留原统计,调度成本检查不再套用管理显示的 key 数量截断。
- **P1-5 已增加流式诊断捕获共享预算:** `AETHER_GATEWAY_STREAM_CAPTURE_MEMORY_BUDGET_BYTES` 默认 128 MiB,覆盖 provider/client 捕获的实际容量及扩容时的新旧分配,显式 0 关闭此类捕获。预算不足时保留连续前缀并标为截断,不阻塞客户端传输;取消和终态释放额度。计费观察独立消费完整数据,补充主解析器停用后的用量恢复,并保留同步 JSON 转流的终态摘要。完成语义恢复后及时释放预读副本。
- **P1-6 已处理单条及 pending 批量写入的审计压缩:** 准备工作移到事务前的 blocking 任务,进程内最多 4 个工作任务、32 个含等待的准入任务,等待工作槽最多 1 秒、等待执行结果最多 30 秒。超时和容量不足返回可重试 `TimedOut`;取消后的运行任务继续持有许可,且只准备数据,不会自行写数据库。序列化直接写入带 8 KiB 缓冲的 gzip,避免完整 JSON 中间副本及普通路径一次额外 body 克隆。事务内仍保留依赖旧记录的生命周期、清空、恢复和幂等判断。
本地验证:
- PostgreSQL usage 模块:124 项通过,13 项需要外部条件的测试默认忽略。
- 隔离 PostgreSQL 14.17 的 5 项真实回归通过,覆盖单条/批量完整审计读写、重复事件计数、过期终态 no-op 和捕获清空;临时数据库已停止。日志:`/tmp/aether-usage-concurrency.68ZESf/test.log`。
- 第一批网关回归 68 项通过,包含真实 Redis 命令计数、管理统计保留、策略读取矩阵、超过管理显示上限的成本检查及调度结果一致性。
- 最终流式模块及相关超时回归 187 项全部通过,覆盖捕获预算耗尽、并发预算释放、用量回退、累计更新及显式归零、协议转换、同步 JSON 桥接和非对象字段兼容。日志:`/tmp/aether-concurrency-round2-stream-final.log`。
- 修改过的 Rust 文件格式检查和 `git diff --check` 通过。
仍需后续处理的边界:
- 启用成本/延迟策略时仍读取原始窗口,尚未改成增量聚合或合并刷新。
- 128 MiB 不包含独立协议/计费解析缓冲、终态 base64 编码或 usage 队列副本。计费用量回退保留的单条协议记录仍有 Basic 5 MiB / Full 64 MiB 上限,记录完成即释放;大量并发超长单记录仍可能放大内存。审计准备许可约束任务数,也不是任意大 usage 记录的字节预算。
- P1-7 长期额度聚合和 SQL/锁等待期限,以及独立实例阶梯压测,本轮未处理。
## 第三轮修复状态
- **P1-7 普通 SQL 等待期限:** PostgreSQL 连接默认设置 `statement_timeout=30000ms`、`lock_timeout=3000ms`,由 `AETHER_GATEWAY_DATA_POSTGRES_STATEMENT_TIMEOUT_MS` 和 `AETHER_GATEWAY_DATA_POSTGRES_LOCK_TIMEOUT_MS` 覆盖,显式 0 关闭,非法值在创建连接池时拒绝。覆盖直接连接池查询和事务;这是单条语句期限,不是整笔事务总期限。超时 SQLSTATE 仍按原可重试错误处理。
- **P1-7 多窗口精确查询:** 长期请求额度和成本额度将多个窗口的独立扫描合并为一条带 `FILTER` 的聚合查询,保留用户锁、时间边界、有效预留过滤、当前事件排除、拒绝顺序和幂等语义。未引入近似额度或缓存;仍需扫描最大覆盖范围内的历史记录。
- **维护任务隔离:** schema migration 和历史 backfill 使用关闭普通期限的专用连接,所有退出路径均关闭连接,避免配置泄漏回请求池。日/小时统计、钱包每日聚合、用量统计重建与 VACUUM 使用 5 分钟语句期限和 30 秒锁等待期限。
- **P1-2 有界 Redis 聚合:** 单个窗口最多 512 条记录时在 Redis 内精确聚合,只返回总量和正值样本数;每批最多 16 个独立脚本。更大的窗口或聚合失败回到原完整查询;没有缓存金额,也没有无界 Lua 聚合。超过 512 条的成本窗口仍有原查询开销,并增加一次计数探测。
- **P1-5/P1-6 事件正文保留:** 已构造的 usage 事件在同步保存计费相关字段后、进入终态队列或等待正文策略前,按四份诊断 JSON 正文的堆内存估算申请共享预算。`AETHER_USAGE_EVENT_CAPTURE_MEMORY_BUDGET_BYTES` 默认 128 MiB,显式 0 关闭正文保留。预算申请不等待,额度不足舍弃正文并标记 `Truncated`,保留费用、用量、终态和引用;事件副本单独申请额度,重试随事件保留额度,取消及释放事件时归还。异步策略读取仍在原来的终态执行和准入位置,维持提交、排序与异常隔离语义。
- **减少中间副本:** usage envelope 和死信队列序列化改为借用正文及字段,避免序列化前的完整深拷贝。新增 `usage_runtime_event_capture_memory_budget_bytes`、`usage_runtime_event_capture_memory_retained_bytes` 和 `usage_runtime_event_capture_memory_downgraded_total` 指标;降级次数也可能包含随后按 Basic 策略清除的正文。
本地验证:
- Redis 运行态 3 项真实实例回归通过,覆盖 512/513 条边界、多批 Key、`u64` 精度和饱和、脚本重载、窗口变化及已完成的并发写入。日志:`/tmp/aether-round3-redis-tests.log`。
- PostgreSQL 全量普通测试 229 项通过、24 项默认忽略;跨模块 `cargo check -p aether-data` 通过,维护 future 的 `Send` 回归通过。隔离 PostgreSQL 14.17 显式验证空库迁移和 usage 读写、超时回滚与期限隔离、精确多窗口准入和幂等、共享套餐锁,共 4 项通过。临时库已停止;日志:`/tmp/aether-postgres-deadlines.7sZ8w1/`。
- 用量模块完整回归 264 项全部通过,含 13 项正文预算测试及终态队列容量、排序、异常隔离回归。日志:`/tmp/aether-round3-usage-agent-tests.log`。
- 最终网关调度、真实 Redis 窗口回退及指标回归 61 项全部通过;最终网关、mock upstream 和 seed 工具构建通过。日志:`/tmp/aether-round3-gateway-tests.log`、`/tmp/aether-round3-verified-build.log`、`/tmp/aether-round3-seed-final-build.log`。
- 修复压力 seed 工具按网关密钥加密 provider/client 凭据,并用已有凭据比较更新接口处理随机密文的重复初始化;隔离空库连续初始化两次已通过。mock chat 流按 `stream_options.include_usage=true` 输出终态用量,保留故障截断和未请求用量时的原行为,9 项测试全部通过。日志:`/tmp/aether-round3-mock-tests.log`;最终 mock 构建日志:`/tmp/aether-round3-mock-final-build.log`。
### 隔离端到端验收
使用本机 debug 构建、独立空 PostgreSQL 14.17 和 Redis、4 个 Tokio worker、入口并发上限 128、8 个客户端 API key。上游为单个零价 mock provider,每请求 20 个 256 字节内容块,首字节延迟 30 ms、块间隔 100 ms,完整流约 2 秒。请求启用流式 usage,客户端必须收到完整响应及 `[DONE]`。各档采样间隔 500 ms,停止流量后最多等待 45 秒排空;指标通过管理员会话采集,避免管理 Token 使用计数干扰排空判断。
| 并发 | 请求数 | 成功数 | 实测 RPS | 首字节 P95 (ms) | 总耗时 P95 (ms) | 网关峰值 RSS (MiB) | 排空 (s) |
| --- | --- | --- | --- | --- | --- | --- | --- |
| 8 | 64 | 64 | 4.01 | 117 | 2069 | 107.0 | 8.7 |
| 24 | 192 | 192 | 12.04 | 67 | 2015 | 122.3 | 9.9 |
| 48 | 384 | 384 | 23.91 | 87 | 2036 | 143.3 | 9.8 |
- 三档共 640 个请求,全部 HTTP 200 且完整 SSE 结束,无请求错误。数据库恰有 640 条 `completed`,每条 input/output/total tokens 均为 `1/20/21`,合计 `640/12800/13440`,所有首字节时间有值。
- 三档排空所需指标齐全,最终 usage 队列 lag、pending、计数 outbox、终态提交和有序生命周期待处理数均为 0;采样未观察到 PostgreSQL 锁等待,未出现 usage worker 处理失败。最终事件正文预算保留量为 0,未触发正文预算降级。
- RSS 随并发档位上升,排空时未立即回落;不能据此证明长期无内存增长。最高档约 1435 RPM 是此短时、约 2 秒 mock 流的实测值,不代表生产容量,也没有修复前的同条件吞吐基线。真实长流、大请求、多候选账号池、长期额度和现金扣费负载未在本次阶梯压测覆盖;额度和结算语义由前述独立数据库回归验证。
- 最终产物:`/tmp/aether-round3-pressure.T3sq7J/`,包含每档 JSON、原始日志、最终指标及 `usage-integrity.txt`;脚本:`/tmp/aether-concurrency-round3-pressure.sh`,总日志:`/tmp/aether-round3-pressure-final.log`。临时网关、mock、采样代理、Redis 和 PostgreSQL 均已停止。修改文件格式检查及 `git diff --check` 通过,未执行全工作区测试,未部署线上。
本轮仍未覆盖的内存与计算边界:
- 事件正文预算是 JSON 堆内存估算,不是进程 RSS 上限。仍不包含计费提取前的 seed、Redis 消费批次、数据库写入 DTO、压缩和序列化字符串,以及协议观察缓冲。
- Redis 超大成本窗口仍走原始历史查询;长期 SQL 额度仍需扫描历史,尚未改为带迁移和对账的增量账本聚合。
## 第四轮修复状态
- **受限等待期间减少重复候选准备:** 普通指定模型和无指定模型能力查询保留 150 ms 等待窗口,分页候选路径保留首次受限后的 100 ms 窗口。首次完整准备确认仅被客户端 API Key 并发限制阻塞后,等待期间只读取并发判断所需的近期候选;恢复或到达期限后重新执行完整动态校验,不直接复用旧候选放行。持续阻塞且首次准备未耗尽期限时,完整准备收敛为首次和最终两次;若短暂恢复后重新受限,仍可在原期限内重新准备。等待窗口不重置,也不是包含数据库 I/O 的端到端硬超时。
- **探测任务创建前合并:** 同一 AppState 的请求补探测按 provider 在 `spawn` 前合并;运行期间的重复触发只保留一次后续执行信号,完成和取消时清理本地条目。AppState clone 共享协调器,切换到不同 RuntimeState 时隔离。最多同时保留 1024 个活跃 provider,超过容量跳过此次 best-effort 补探测,不创建等待任务;周期基础探测继续运行,但不等价于请求触发的 Burst 补探测。
- **跨实例退出交接:** 保留 Redis pending 和带 token 的锁协议,释放锁后复查新 pending,避免另一实例在退出窗口提交的补探测无人接手。读取 pending 失败时释放锁并停止本轮处理,避免残留 pending 导致重新争锁的忙循环。
- **可观测性:** 增加 `pool_quota_probe_replenish_provider_capacity`、`pool_quota_probe_replenish_active_providers`、`pool_quota_probe_replenish_started_total`、`pool_quota_probe_replenish_coalesced_total` 和 `pool_quota_probe_replenish_capacity_rejected_total`。
本地验证:
- `cargo check --locked -p aether-gateway --lib` 通过。日志:`/tmp/aether-round4-check.log`。
- 网关调度、分页候选、探测、模块边界、指标及真实 HTTP 并发等待回归共 376 项全部通过,包含本轮新增的 11 项调度等待和 8 项探测协调测试。日志:`/tmp/aether-round4-gateway-tests.log`。
- 8 个同时受限的请求,通过真实候选筛选函数的依赖调用计数,确认完整候选读取共 16 次;恢复后重新加载被替换或移除的候选,轻量读取及完整重验的错误均正常传播。分页持续受限仅重启一次扫描,首次查询耗时计入原重试预算,取消后停止查询。
- 单 provider 的探测运行中并发触发 64 次,仅创建一个任务并额外执行一次补探测。测试覆盖不同 provider 并行、100 次本地退出交接竞争、未首次执行即取消、运行中取消、panic、容量释放和 runtime 绑定隔离。两个本地协调器共享 Memory Runtime 模拟跨实例 pending/锁交接,精确覆盖旧任务最终检查后、解锁前的新触发;没有声称本轮运行了多进程 Redis 压测。
- `cargo build --locked -p aether-gateway --bin aether-gateway` 通过,最终可执行文件已更新;日志:`/tmp/aether-round4-build.log`。修改文件格式和 `git diff --check` 通过,测试进程已退出。本轮未重新执行阶梯吞吐压测,第三轮的 640 请求结果仅代表其当时构建;未执行全工作区测试,未部署线上。
边界:
- API Key 并发检查仍使用全局最近 128 条候选中的同 Key 活跃候选行,依赖状态持久化及原 300 秒活跃窗口;未引入原子分布式许可,也未改变候选行计数为请求去重计数。
- 探测协调为 best-effort;取消后的 Redis 锁仍按原 30 秒 TTL 释放,未增加续期或 exactly-once 保证。同步日志、超大历史窗口增量聚合、未覆盖的内存副本以及目标环境长期压测仍待处理。
## 第五轮修复状态
- **运行日志 I/O 脱离请求线程:** stdout 和滚动文件各使用独立专用写线程,保留 Pretty/JSON、Stdout/File/Both、动态日志过滤及原有文件权限检查。文件写入、flush 和轮转重开均在写线程执行;定期文件清理移到 blocking 任务,单次完成后才安排下一次清理。启动时仍同步校验日志目录和目标文件,配置错误正常拒绝启动。
- **日志过载保护:** 每个目标最多排队 4096 条事件,复制后的日志正文最多保留 8 MiB,单条事件上限 256 KiB。字节预算包含生产者已预留的复制、排队和正在写入的正文;队列满、预算不足或单条过大时整条舍弃,不等待设备、不在请求线程同步回退 stderr。Both 两个目标独立接收和降级,可能保留不同的事件集合。
- **退出排空:** 网关、隧道、两个运行示例及 13 个基准工具的最外层入口持有日志 guard,先结束 Tokio runtime 再关闭队列。升级/回滚的显式进程退出也调用关闭接口。关闭先停止全部目标接收,再在共用的 2 秒预算内等待已接收事件及最终 flush;阻塞中的系统 I/O 无法强制取消,超时后不无限 join 写线程。
- **信号处理:** 网关 ready 后收到 SIGTERM/SIGINT 会结束运行函数并经过日志 guard;现有连接未被统一追踪,此路径仍按进程终止处理业务,不声称请求或用量队列优雅排空。启动过程中尚未进入信号等待时,以及 SIGKILL/abort,不保证排空。
- **日志健康指标:** 网关与隧道指标端点增加 `logging_stdout_*` / `logging_file_*`,记录队列和字节上限、当前正文保留量、接收数、按原因区分的丢弃数、写入错误、线程 panic、关闭超时及线程状态。沿用服务指标命名空间前缀。关闭 API 返回成功仅说明线程处理完队列且最终 flush 成功,之前的写入错误仍需查看指标,不代表 fsync 持久化成功。
本地验证:
- 网关、隧道、loadtools 和 integration 的所有 binary/example 入口通过 `cargo check --locked`,未增加新的第三方依赖;integration 的依赖清单补充已有共享 runtime。日志:`/tmp/aether-round5-entrypoints-check.log`。
- 共享 runtime 44 项单元测试及 2 项进程集成测试全部通过,包含 11 项新 writer 测试、2 项轮转回归、12 种初始化/输出格式/动态过滤场景,以及真实 stdout 堵塞测试。8 个生产者同时写入时,慢设备不阻塞生产者;4096 条和字节预算、整条拒绝、错误/panic 回收、退出竞争、stdout/file 双向隔离及阻塞 flush 均已覆盖。每种实际格式并发写入 512 条事件,记录完整且无重复。日志:`/tmp/aether-round5-runtime-final-tests.log`。
- 真实 stdout 测试保持子进程管道完全不读,确认 stdout 队列触发丢弃后,文件仍收到完整 JSON 尾记录;guard 超时返回后,标准进程退出也成功完成。该测试耗时约 2.36 秒,包含日志的 2 秒退出等待;没有关闭管道来人为解除阻塞。
- 网关入口 61 项、隧道 197 项回归全部通过,日志:`/tmp/aether-round5-service-tests.log`。以上合计 304 项测试通过,没有计入重复运行的首批测试。
- Linux root/capabilities 专用日志 fixture 已适配异步排空,当前 macOS 环境未执行;Unix 文件权限、符号链接、多硬链接及安全轮转的普通测试已通过。
- 网关和隧道最终二进制构建通过,日志:`/tmp/aether-round5-service-build.log`。隔离 PostgreSQL 和新网关实例的 Both/JSON 验收通过,stdout 和文件各保留 15 条完整日志,均包含唯一的 starting、ready 和 shutdown 事件;实际指标端点两路队列容量、接收及运行状态正常,丢弃、写入错误、panic 和关闭超时计数为 0。SIGTERM 后约 16 ms 正常退出,该耗时仅代表健康设备和无在途代理请求的此次烟测。产物:`/tmp/aether-logging-smoke-qvzN9U/result.json`,脚本:`/tmp/aether-nonblocking-logging-smoke.mjs`。
- 临时网关和 PostgreSQL 已全部停止,测试与构建进程均已退出。修改文件格式检查及 `git diff --check` 通过;本轮未重跑业务阶梯吞吐压测,未执行全工作区测试,未部署线上。
本轮边界:
- 日志格式化、字段 Debug 展开和 JSON 序列化仍发生在调用线程,订阅器线程局部字符串也可能保留历史容量;8 MiB 预算只约束交给后台写入的正文,不包含这些临时对象、队列元数据或进程 RSS。
- 运行日志在过载、I/O 错误或关闭超时下可能丢失,不能作为可靠计费账本;账务持久化流程不使用这条日志队列。2 秒仅限制日志关闭等待,不是整个服务退出期限,既有业务或 blocking 任务清理仍可能更久。
- 超大历史窗口增量聚合、尚未覆盖的内存副本、完整请求优雅排空及目标环境长期压测仍需后续处理。本轮不调整线上配置,也不据此给出生产 RPM 容量。
## 第六轮修复状态
- **补齐诊断正文预算的所有权传递:** 将纯预算令牌放到 data contracts,运行时仍使用原环境变量、128 MiB 默认值和指标。同步及流式终态 seed 在等待提交前申请预算;Redis 解码后的事件重新纳入本进程预算;事件生成的数据库写入 DTO 继续持有对应额度。事件和受管理 DTO 的正文副本分别申请额度,释放正文后才归还;序列化跳过令牌,不改变队列协议或数据库字段。
- **保留完整计费与终态语义:** 正常终态构建仍在 blocking 任务执行;预算不足时同步执行既有纯构建逻辑,先解析 token、显式 cache=0、图像估算、错误及终态,再舍弃诊断正文,随后进入原有有序提交和准入路径。保留原始终态观测时间,构建异常仍按原终态失败路径隔离。已有 `None`、`Disabled`、`Unavailable` 状态不被预算降级改写为 `Truncated`,避免破坏清空指令。
- **旧队列消息兼容:** 解码前的原始字段继续保留到记录或死信处理完成,重试不提前 ACK,DLQ 保存原始字段。旧消息缺正文且缺 typed state 时保留元数据内已有缓存 TTL、tier 和请求事实,显式 `None` 仍清空,避免预算接线改变后续计费输入。
- **数据库准备减少正文副本:** 单条与 pending 批量准备先移走四份正文和 headers,再复制两个存储投影需要的少量元数据,避免原先为清洗而深拷贝整份 DTO。预算随输入进入 blocking 压缩闭包,调用方取消不会提前释放仍存活的正文额度;压缩结束后原始 JSON 释放,压缩结果继续走原事务、审计和正文存储流程。
本地验证:
- data contracts 226 项、PostgreSQL 231 项、data runtime 356 项、usage runtime 282 项普通测试通过,合计 1095 项;其中本轮新增 27 项,覆盖队列积压、并发 seed、预算拒绝与复制、token/图像计费事实、typed clear、legacy metadata、重试/DLQ,以及取消和 panic 后额度释放。另有 25 项需要外部条件的测试默认忽略。日志:`/tmp/aether-round6-data-tests.log`、`/tmp/aether-round6-runtime-tests.log`。最终时间采样和反向 DTO 转事件的所有权修正后,usage runtime 282 项再次全部通过:`/tmp/aether-round6-usage-final-tests.log`,不重复计入总数。
- 隔离 PostgreSQL 14.17 的 5 项真实回归通过。完整审计测试现覆盖受预算管理的单条、普通 pending 批量及同 request 重复批量写入,四份正文读取一致,准备结束后只保留调用方原对象的额度,最终释放为 0;同时验证过期终态 no-op、辅助计数幂等和 typed `None` 清空。日志:`/tmp/aether-round6-postgres.K8Aq5W/`;临时数据库已停止。
- 网关、隧道、loadtools 和 integration 的 binary/example 入口通过 `cargo check --locked`,日志:`/tmp/aether-round6-entrypoints-check.log`;修改文件格式及 diff 检查通过。最终 `cargo build --locked -p aether-gateway --bin aether-gateway` 通过,网关可执行文件已更新,日志:`/tmp/aether-round6-gateway-build.log`。全部测试与构建进程已退出;尚未部署,未重跑业务阶梯吞吐压测,未执行全工作区测试。
本轮边界:
- 预算覆盖上述运行时链路持有的四份诊断 JSON 正文及副本,不是进程 RSS 上限。原始 Redis RESP、批次字段字符串及反序列化临时分配仍不受此额度约束;直接通过契约自行构造或反序列化的 DTO 默认不启用运行时预算。
- seed 进入本链路前的 JSON/base64 解析、构建过程的临时正文复制、序列化/压缩结果和 SQL bind 缓冲未纳入。预算不足时的纯构建会使用调用线程 CPU;它保留计费兼容性,没有消除解析和估算开销。
- 协议观察器及用量恢复缓冲、超大历史窗口增量聚合、完整请求优雅排空和目标环境长期压测仍待后续处理。第三轮吞吐数据不能作为本轮构建或生产容量结论。
## 第七轮修复状态
- **Redis 回复转移正文缓冲:** `XREADGROUP` 采用当前 redis crate 支持 owned conversion 的底层容器结构,避开 `StreamReadReply` 的借用转换;字段正文直接由 RESP `BulkString` 转成 `String`。`XAUTOCLAIM` 的消息正文、游标和删除 ID 同样使用所有权转移。保留 RESP2/RESP3、nil、重复字段覆盖、无效 UTF-8 和错误分类等既有解析语义,没有改变队列协议。
- **减少队列处理副本:** Memory 队列读取直接遍历待交付条目,去掉全局队列锁内的整批临时正文克隆;队列、PEL 和调用方继续独立持有自己的数据。usage worker 在处理完一条消息后先释放原始字段,再等待 ACK/DELETE,避免已处理的大正文跨确认 I/O 继续存活。失败和死信路径仍保留所需原文到处理结束。
- **流式用量恢复收敛保留字段:** Claude 的跨记录状态只累计 mapper 使用的 token 和缓存字段,以及非空用量出现标记,未知大字段不再随流长度累积,也不再每次复制完整累计对象。候选和 chunks 数组省略没有用量或 tier 信息的空元素,保留 Gemini 的首候选位置和倒序查找规则。完整且未超限的 SSE 记录直接借用当前输入切片,跨分片与超限记录继续沿用原 carry 和恢复规则。
- **投影兼容修复:** 显式 `usage: null` / `usageMetadata: null` 保留字段存在性,避免错误回退到旧快照;独立 null 不新增清零信号,带其他非空用量的嵌套结构遵循原 mapper 的优先级。图片回复仅提取 `data/result` 数量以恢复 `request_count/image_count`,不复制图片正文。
本地验证:
- runtime state 76 项、usage runtime 282 项回归通过,合计 358 项。Redis parser 的 7 项新增测试用正文原指针和容量断言验证缓冲转移,另有 3 项 Memory 所有权和队列生命周期回归。日志:`/tmp/aether-round7-queue-tests.log`。
- 新增真实 Redis 大消息回归,两种协议各 24 条消息、每条约 512 KiB 正文,3 个消费者同时分批读取,再以最多 5 条重领和 ACK/DELETE;48 条消息原文逐字节一致,无重复交付到不同读取结果,最终 pending、lag 和 stream length 都为 0。日志:`/tmp/aether-round7-large-redis-test.log`。测试自建 Redis 已关闭;此项已包含在前述 76 项中,不重复计数。
- 网关流式、stream pump 和空闲读取期限 186 项回归全部通过,日志:`/tmp/aether-round7-stream-final-tests.log`。包含本轮新增的 9 项投影与借用测试,覆盖原始 Claude 累计对象对照、256 条带不同未知大字段的长流、10000 个空数组元素、正文复制计数、CR/LF/CRLF 分片及上限边界、显式 null 的实际 mapper 对照,以及仅有图片数量的回复。连同队列侧共 544 项通过,本轮新增 20 项;首批 parser、真实 Redis 和流式聚焦测试不重复计数。
- 修改文件格式及 diff 检查通过,最终 `cargo build --locked -p aether-gateway --bin aether-gateway` 通过,网关可执行文件已更新,日志:`/tmp/aether-round7-gateway-build.log`。所有测试、构建进程和临时 Redis 均已退出;尚未部署,未执行新的业务吞吐压测或全工作区测试。
本轮边界:
- 原始 Redis 消费批次仍按配置的条数读取,不是字节预算。默认每批最多 128 条、最多 32 个 worker;本轮减少重复分配,没有限制任意大消息或整体批次的最大驻留字节,也没有通过丢弃账务消息缩小批次。当前同 worker 的 read/reclaim 已串行,已处理条目原本就逐条释放。
- 流式恢复仍保留必要的跨分片单记录,Basic 5 MiB / Full 64 MiB 的原限制未改变;多行 `data:` 拼接、JSON 解码临时值、非空用量数组和主协议观察器不在全局正文预算内。不能因减少诊断副本而直接停用计费恢复。
- 主协议解析器的累计文本、原始队列批次字节控制、超大窗口增量聚合、完整请求优雅排空及目标环境长期压测仍待后续处理。
## 第八轮修复状态
- **主协议用量观察器取消正文累计:** `StreamingStandardTerminalObserver` 为 OpenAI Chat、Responses(包括 compact)及 Gemini 选用现有 provider parser 的终态观察模式。协议转换仍使用完整模式;观察器沿用同一事件分类和用量解析,只保留摘要需要的状态。Claude 当前没有正文累计,独立图片观察器也继续使用原实现。
- **长流正文和工具参数:** Responses 不再保存文本、双份推理文本、工具参数、工具结果和图片项正文;保留工具索引、名称及 namespace 校验,它们影响未知事件计数和工具调用结束原因。Chat 不再等待迟到的工具名称或 ID 而持续缓存参数。Gemini 不再保存累计文本、推理、签名、媒体、工具参数和结果,只记录是否见过工具调用以保留结束原因。
- **opaque 项去重:** Responses 观察模式将未知扩展项的去重键改为固定 32 字节 SHA-256 摘要,按原始键的相同字节增量计算,避免加密内容或整个序列化项成为常驻键。默认转换和客户端 emitter 的原键及完整输出保持不变。
本地验证:
- `aether-ai-formats` 全部 922 项测试通过,包含本轮新增 16 项。逐个输入前缀及提前 EOF 对比完整解析器与观察模式的摘要,覆盖身份时点、用量、显式零、tier、错误、未知项去重、namespace、工具调用和 SSE / WebSocket 结构化入口。原有格式转换、会话历史、图片及同步转流回归均通过。日志:`/tmp/aether-round8-formats-final-tests.log`。
- 长流回归在 2048 轮 1 KiB 文本、推理及工具参数输入后,直接断言 Responses 的正文缓冲容量仍为 0、Chat 工具缓冲为空;Gemini 在 32 轮多种 16 KiB 字段输入后所有正文状态 map 仍为空,并与完整模式核对终态帧。大 completed 项与 opaque 键的摘要字节、去重语义另有测试。此处验证内容保留行为,没有测量生产 RSS 或吞吐上限。
- 网关流式、stream pump 和读取期限 186 项通过;Responses WebSocket 会话、上游和观察器 209 项通过。连同格式 crate 共 1317 项通过。日志:`/tmp/aether-round8-gateway-stream-tests.log`、`/tmp/aether-round8-gateway-websocket-final-tests.log`。WebSocket 回归直接运行同一份已编译测试程序;其间因现有 build script 监听 worktree 中不存在的 `.git/HEAD` 而触发的一次重复 Cargo 编译已主动停止,没有把中止当成测试通过。
- 修改文件格式和 diff 检查通过,最终 `cargo build --locked -p aether-gateway --bin aether-gateway` 通过,日志:`/tmp/aether-round8-gateway-build.log`。本机内存压力较高,网关测试目标编译耗时 11 分 21 秒、最终构建 8 分 54 秒;这些是本地编译耗时,不是业务延迟。全部测试、构建进程已退出,未遗留本轮临时服务。尚未部署,未执行新的业务吞吐压测或全工作区测试。
本轮边界与后续:
- Responses 的工具身份和 opaque 摘要集合仍按不同逻辑项数增长;单条 JSON 解码、未知事件错误载荷及部分 Gemini 分类 helper 仍有临时分配。完整格式转换和会话历史依赖的正文仍保留,本轮不宣称整个解析器或进程具有硬性 RSS 上限。
- Redis 原始批次仍没有硬性字节上限。现有 redis 连接的超时或 future 取消不会保证后台立即停止接收已发命令的回复,仅读取后套预算或减少 COUNT 不能解决任意大消息。建议下一步先为新生产的完整队列 envelope 实施精确序列化字节上限,超限诊断降级须保留计费事实并沿现有失败路径重试;旧消息和 PEL 仍需兼容排空。真正的接收预算还需要读取前预留及受控连接/解码器,不能直接依赖当前 usage body blob 表作为入队旁路,该表依赖已存在的 usage 父记录。
- 超大窗口增量聚合、完整请求优雅排空及目标环境长期压测继续待处理。尚未部署。
## 第九轮修复状态
- **新增队列消息的完整字节上限:** `AETHER_GATEWAY_USAGE_QUEUE_PAYLOAD_MAX_BYTES` 默认 1 MiB,启用 usage runtime 时显式 `0` 非法。`UsageQueue::enqueue` 在发送 Redis 命令前,以有界 writer 编码完整 v1 JSON envelope,包含 UTF-8、转义、metadata、正文、headers 及其他字段;不会先生成无限制的完整 JSON 字符串再检查长度。普通消息保持原格式,原公开编码接口及历史消息解码继续兼容。
- **诊断降级保留计费语义:** 完整消息超限后,先借用检查去掉四份正文和四份 headers 的核心字段大小,核心可容纳才克隆 metadata 并生成诊断投影;复用同一字节缓冲。保留 token、费用、显式零、错误存在性、身份、时间、终态、正文引用、预留 token 和计费维度。按完整 v1 消费者规则保留请求档位、推理参数、响应实际档位及缓存 TTL,并标记被移除的正文为 `Truncated`。显式 `None`、`Disabled`、`Unavailable` 不改为可回退的状态;JSON null 与非对象正文分别按旧解码和权威规则处理。
- **无法安全编码时的失败路径:** 核心仍超限,或去掉正文无法保留原缓存 TTL 计费语义时,返回 `InvalidInput`。终态沿现有有并发限制的数据库路径使用原事件回退;数据库不可用、受压或写入失败时明确返回 `Failed`,保留 first-byte 状态,不误报已入队或已缓冲。这类输入错误不打开 Redis 熔断,也不会无限重试。重试接收前再次校验,覆盖主路径已熔断或入队槽耗尽的旁路;重试 worker 也会终止单条永久失败并继续处理后续条目。
- **指标与临时分配:** 导出队列 payload 上限、诊断降级、编码拒绝及永久重试失败计数。payload 计数是进程级编码尝试,包含入队与重试预校验,不是唯一事件数。超长 tier/reasoning 字符串先检查已有的 64 字节限制,再执行大小写规范化,避免明知非法仍复制整个字符串。
本地验证:
- 用量模块全部 298 项通过,包含本轮新增的 9 项编码、6 项失败路径及 1 项配置回归。精确覆盖 UTF-8/转义字节边界、编码缓冲复用、字段透传、源事件不变、完整 v1 消费者对照,以及超限后的数据库回退、first-byte 保留、熔断/准入旁路拒绝和同重试分片继续排空。日志:`/tmp/aether-round9-usage-final-tests.log`。
- 计费模块全部 87 项、数据契约全部 229 项通过。本轮新增 7 项真实队列读回计费对照与 3 项超长字段规范化回归;对照覆盖请求档位、缓存 TTL、显式零、未知价格、错误/取消、图片矩阵维度,以及 15 组非对象/null 正文组合。计费结果以原完整 v1 消息经旧解码路径后的行为为基线。连同用量模块共 614 项通过,不重复计入首批验证。日志:`/tmp/aether-round9-billing-contracts-final-tests.log`。
- 网关指标、工作区模块边界、流式链路及 Responses WebSocket 共 438 项通过;网关入口配置全部 62 项通过,包括本轮新增的 payload 配置与指标回归。连同核心模块共 1114 项通过,本轮新增 28 项测试。直接复用本次构建生成的测试程序执行,日志:`/tmp/aether-round9-gateway-lib-tests.log`、`/tmp/aether-round9-gateway-main-tests.log`。
- `cargo build --locked -j 1 -p aether-gateway --all-targets` 通过,最终网关可执行文件已更新;一次构建同时生成普通程序与测试目标,本地耗时 25 分 20 秒。日志:`/tmp/aether-round9-gateway-build.log`,产物清单:`/tmp/aether-round9-gateway-artifacts.jsonl`。修改文件格式及 diff 检查通过,所有本轮测试和构建进程均已退出;未创建临时服务。尚未部署,未重新执行业务吞吐压测或全工作区测试。
本轮边界:
- 本次限制针对新生产的单条 JSON payload,不包含 RESP 外壳、完整读取批次、历史消息、PEL、DLQ 或进程总 RSS。原始 JSON 树、headers、metadata 的存活内存及解析临时值不因此获得统一字节预算;默认 128 条批次、32 个 worker 仍可能同时接收很多消息。真正的接收预算仍需要读取前预留和受控连接/解码器。
- 对过大的原 metadata 采用保守拒绝,避免为判断最终能否缩小而先深拷贝整个对象;即使后续规范化理论上可使其变小,也使用原事件回退。模型分类 helper 的超长输入临时分配仍需后续审查,本轮不宣称所有计费提取都有硬性内存上限。
- 没有新增持久化超限旁路。若消息无法编码且受限数据库回退也失败,用量落库会失败,并通过日志和指标暴露;不能把有限本地缓冲称为可靠落盘。超大窗口增量聚合、完整请求优雅排空及目标环境长期压测仍待处理。尚未部署。
## 第十轮修复状态
- **阻塞读取改为独占连接:** 原 blocking stream lane 通过轮询复用 `ConnectionManager`,快 worker 再次读取时可能命中其他 worker 正在执行 `BLOCK` 的连接,造成队头阻塞;取消调用 future 后,后台 driver 仍可能继续执行旧命令。现改为有固定容量和信号量的独占池,先取得空闲槽位再发命令,等待者取消不会影响现有读取。连接和未 spawn 的 driver 由同一查询持有,取消、超时及解析失败一起丢弃,下一次使用该槽位时重建;完整解析成功才回池复用。保留原 lane 数量、认证、数据库、RESP 配置及超时/延迟统计。
- **积压重领继续扫描游标:** 原运行时丢弃 `XAUTOCLAIM` 返回的下一扫描位置,worker 每次从 `0-0` 开始。当大量近期活动的待确认消息占据前段,Redis 单次扫描可能返回空页,后段过期消息长期得不到重领。新增兼容的分页接口,保留下一位置和已删除消息 ID;worker 在成功响应后推进游标,空页和删除页也推进,读取错误保持原位置,扫描结束再从头开始。写入失败的消息不提前确认,仍留在 PEL,回绕或 worker 重启后可再次重领。原返回消息列表的公开接口和旧队列实现继续兼容。
- **内存队列并发入队顺序:** 将序号分配移入队列插入的同一个锁范围。此前等待插入的生产者可能先取得较小 ID,其他生产者先插入并被消费后,较小 ID 会被读取游标永久跳过;现在 ID 顺序与实际插入顺序一致。内存后端同时实现重领分页,按数值序号推进。
本地验证:
- `cargo check --locked -j 1 -p aether-runtime-state -p aether-usage-runtime` 通过。日志:`/tmp/aether-round10-core-check.log`。
- 运行时状态全部 84 项、用量模块全部 300 项测试通过,共 384 项,包含本轮新增 10 项:连接池 2 项、内存队列 2 项、worker 游标 2 项、真实 Redis 4 项。覆盖连接容量、等待取消、多任务竞争、锁等待下的 ID 顺序、空页/删除页、读取失败重试、写入失败后回绕及原有记账流程。日志:`/tmp/aether-round10-queue-tests.log`。
- 真实 Redis 测试另以 `--nocapture` 运行同一份测试程序,确认 6 次隔离实例就绪、没有跳过,4 项全部通过,不重复计入上述 384 项。使用本机 Redis、ACL 认证及数据库 7,覆盖 RESP2/RESP3、满池等待不发新命令、其他 lane 继续服务、快消费者复用空闲连接、取消/超时后服务端旧连接消失、重连保留认证/数据库,以及空扫描页后的 PEL 尾部和删除项恢复。日志:`/tmp/aether-round10-redis-receive-tests.log`。
- 最终 `cargo build --locked -j 1 -p aether-gateway --bin aether-gateway` 在运行约 18 分 36 秒后因本机资源压力主动停止,退出码 143,不计为构建通过;停止前没有编译错误,日志:`/tmp/aether-round10-gateway-build.log`。期间清理了三份当前构建未使用的旧增量缓存,将可用磁盘从约 3 GiB 恢复至 7 GiB,但内存压力仍使编译持续缓慢。本轮没有修改网关入口、配置或指标,新增逻辑已在上述核心模块中验证;未重复编译网关测试目标,现有网关程序仍是上一轮产物。
- 修改文件格式和 diff 检查通过。本轮测试、编译和临时 Redis 进程均已退出。尚未部署,未重新执行业务吞吐压测或全工作区测试;资源允许时仍需补跑最终网关构建。
本轮边界:
- 本轮修复连接占用和积压恢复,不是完整读取批次的字节预算。历史超大消息、RESP 解码、默认 128 条批次和多个 worker 的总驻留内存仍需后续控制;上一轮新消息 1 MiB 上限保持有效。
- 连接生命周期的取消控制目前仅用于阻塞 `XREADGROUP`。非阻塞 stream 命令及 `XAUTOCLAIM` 仍沿用原共享连接;关闭连接不能撤销 Redis 已执行的读取,已进入 PEL 的消息仍依赖重领。内存后端分页仍扫描并排序符合条件的待确认条目,本轮未建立该扫描的硬性内存上限。
- 超大窗口增量聚合、完整请求优雅排空及目标环境长期压测继续待处理。尚未部署,测试结果不能用来推断生产 RPM 上限。
## 第十一轮修复状态
- **读取批次的共享预留:** 新增进程级 usage worker payload 预留,默认总量 128 MiB、单批目标 8 MiB。根据当前 `queue_payload_max_bytes` 推导实际 `COUNT`,默认由 128 条降到最多 8 条;读取与重领、所有 worker 和自动扩容后的新 worker 共用同一份额度。许可在发命令前取得,并保留到原始字段处理、记账及 ACK 完成;取消、失败和空响应释放,预算不足时在原任务内等待,不新增缓存任务。收到消息后按字段值实际长度缩减多余预留。
- **兼容历史消息与配置:** `AETHER_USAGE_QUEUE_READ_PAYLOAD_BUDGET_BYTES`、`AETHER_USAGE_QUEUE_READ_BATCH_PAYLOAD_BYTES` 分别控制总额与单批目标;0/非法值回退默认,极大值收敛到约 4 GiB 的有效额度,单批目标不超过总额。单条配置上限大于总额时明确返回配置错误,避免等待永远拿不到的许可。历史消息或其他生产者的大消息继续原记账流程,不在持有部分额度时等待追加,也不因超估算而删除或死信。
- **重领连接与停止:** `XAUTOCLAIM` 改用上一轮的独占连接池,完整解析成功才回收连接。取消/超时会同时释放查询和 driver,避免归还读取预留后旧命令仍在后台接收。等待池容量也纳入原命令期限,指标归入实际使用的 `blocking_stream` lane;worker 取得重领响应前可被 shutdown 取消,取得响应后完成原处理与确认。读取与重领之间不嵌套持有池连接。
- **扩缩容与观测:** worker 上报实际请求的 `COUNT`,防止读满 8 条却按 128 条误判为未满、压制扩容。新增 8 个 `usage_runtime_queue_read_*` 指标,涵盖总额、单批目标、当前预留、等待者、等待次数、累计字段字节及超估算条目/批次。累计字段字节包含字段名和值;预留及超估算判定使用值长度,合法 payload 恰好达到上限不会因字段名 `payload` 多出 7 字节而被误报。次数包含再次重领,不是唯一事件数。
- **计费查询失败的恢复语义:** 原 worker 吞掉计费补全错误后继续写入/结算;暂时的数据库超时可能被后续“缺少实际费用”的永久错误覆盖,导致错误死信和 ACK。现在先传播原始错误,失败条目留在 PEL,后续重领再尝试;真正返回成功的无价格事件沿用原行为。三个直接落库入口也统一在补全失败时停止写入,不能误报已持久化或完成有序终态;已有受限数据库回退本来就正确停止,继续沿用。永久错误归档前先释放已解码事件,减少与原字段、死信编码的同时持有。
本地验证:
- 运行时状态全部 87 项测试通过,包含本轮新增 3 项真实 Redis 重领回归。日志:`/tmp/aether-round11-state-tests.log`。
- 用量模块最终全部 321 项通过,包含本轮新增 9 项预留、8 项 worker、4 项直接落库回归;连同运行时状态共 408 项,本轮新增 24 项。覆盖跨 Queue/worker 共享额度、16 个并发任务竞争、极大配置不溢出或永久等待、读/重领实际 COUNT、慢写/ACK 持有、错误和停止释放、历史超估算消息继续处理,以及价格查询超时不提前写库或确认、后续成功尝试使用准确费用。日志:`/tmp/aether-round11-usage-final-tests.log`。直接落库回归中的恢复是后续显式调用,不是新增自动重试。
- 直接运行本轮已编译的真实 Redis 接收测试程序,7 项全部通过,确认 9 次隔离实例就绪、没有跳过;不重复计入上述测试数。RESP2/RESP3、ACL、数据库 7 下验证读/重领取消和超时、服务端旧连接释放、PEL/游标/删除项恢复,以及共池时等待期限、等待取消、释放单个槽位后继续读取和重领。日志:`/tmp/aether-round11-redis-receive-tests.log`。
- 修改文件格式与 diff 检查通过;网关 `cargo check --locked -j 1 -p aether-gateway --bin aether-gateway` 通过,耗时 3 分 14 秒,日志:`/tmp/aether-round11-gateway-check.log`。本轮测试、检查和临时 Redis 进程均已退出。本轮没有重复执行上一轮因本机资源压力中止的完整代码生成与链接,最终网关程序仍需在资源允许时构建;未部署,未执行业务 RPM 压测或全工作区测试。
本轮边界:
- 这是按当前生产配置估算的逻辑 payload 预留,不是网络接收字节或进程 RSS 的硬上限。滚动发布、其他实例使用更高上限、历史消息和直接写入额外字段都可超过估算。字段结构、字符串容量、Redis RESP 解码、连接缓冲高水位、解码 JSON 及死信序列化不由此获得硬上限;旧公开 Vec 读取接口保持兼容,不携带处理阶段许可。
- 死信仍先追加再确认源消息,异常重试可能重复归档;历史大消息的死信 JSON 编码也没有独立硬字节预算。后续需要按源消息身份幂等的转移及受控超大消息恢复,不能以截掉原始账务字段或提前 ACK 代替。
- 直接落库失败不会因此新增持久化重试渠道。超大窗口增量聚合、完整请求优雅排空及目标环境 RPM 压测仍待处理;尚未部署。
## 第十二轮修复状态
- **死信原子转移:** 内置 Redis 后端以一次 Lua 调用检查指定消费组的精确 PEL 身份,再追加完整死信、ACK 并删除源 ID;并发重领或提交成功但响应丢失后的重复调用不再次追加。Memory 后端在同一队列锁内完成转移,序号也在锁内分配。新增可选 trait 接口,默认返回 `None` 且无副作用;旧外部实现仍可使用原来的非原子追加后确认,内置转移报错不会降级为非原子写入。公开 `push_dead_letter` 保留原追加行为及 `{entry_id, fields, error}` JSON 格式。
- **先检查,再写入:** 校验完整 canonical `u64-u64` ID、不同的源与目标键及非空字段;Redis 在首写前检查 PEL、目标类型和 XADD/XACK/XDEL 权限,避免可预见的脚本错误导致部分归档。Lua 不将 ID 转为浮点数,保留大整数精度。正文已被 trim 但 PEL 仍存在时,可使用调用者持有的完整字段归档。操作使用既有独占连接和 owned driver,等待池容量也计入超时,完整解析成功才回池;取消不能撤销已经完成的 Redis 命令,重试依靠 PEL 状态避免重复。
- **独立的编码预留:** 新增 `AETHER_USAGE_DLQ_ENCODING_BUDGET_BYTES`,默认 64 MiB,及 `AETHER_USAGE_DLQ_ENCODING_MAX_JOBS`,默认 4。按原文字符串长度、JSON 最坏 6 倍转义及字段结构分隔符一次预留逻辑字节,使用 checked 运算和有界 writer;预留成功后才复制兼容接口输入或启动后台编码。worker 直接移动现有原始字段,永久记录错误时先释放解码事件。额度不足或单条过大立即失败,原消息保留在 PEL,不截断账务字段。后台任务取消等待后仍持有自己的许可,许可随编码结果保留到队列写入完成;与读取预留独立,避免持有部分读取额度再等待追加。
- **worker 状态与指标:** 成功原子转移直接使用实际 ACK 数并报告一次死信,不再加入批次 ACK;返回 `NotPending` 不报告已归档,也不额外删除源消息。普通记录与旧后端追加仍批量确认,改为报告实际返回的 ACK 数。某条消息的存储转移失败时,只确认此前已成功处理的前缀,失败条目及尚未处理的后缀等待重领。编码预算拒绝及编码失败单独返回延期状态,保留坏消息并继续同批其他条目,只确认成功项,批次末尾仍报告失败;避免永远超过总预算的坏消息在每次重领时持续挡住正常账务。归档成功日志移到写入返回之后。新增 7 个 `usage_runtime_dlq_encoding_*` 指标:总额、最大/当前任务、预留字节、容量拒绝、超限拒绝及编码次数;次数包括重复尝试,编码成功不等于归档成功。
本地验证:
- `cargo test --locked -j 1 -p aether-runtime-state` 全部 102 项通过;本轮新增 15 项,包括 8 项 Memory 转移、6 项真实 Redis 和 1 项返回解析验证。测试编译耗时 13.62 秒,执行 2.45 秒;日志:`/tmp/aether-round12-state-tests.log`。
- `cargo test --locked -j 1 -p aether-usage-runtime` 全部 341 项通过;本轮新增 20 项,包括 8 项编码预算、3 项队列接线、8 项 worker、1 项配置验证。测试编译耗时 30.79 秒,执行 1.16 秒;日志:`/tmp/aether-round12-usage-tests.log`。连同状态模块共 443 项通过,本轮新增 35 项,无编译警告。覆盖完整原文及所有转义、精确上界和溢出、并发饱和、未 poll/后台运行/存储等待取消、panic 释放、原子失败不回退、ACK 实际计数、成功前缀确认、超大坏消息延期后正常后缀继续、后续调高预算恢复,以及旧后端和公开追加兼容。
- 直接运行已编译的 `redis_dead_letter_transfer_ --nocapture`,6 项全部通过,确认 RESP2/RESP3、ACL 认证及数据库 5 下共 8 次隔离实例就绪,没有跳过;不重复计入上述 443 项。覆盖 16 个并发调用只追加一份、丢弃成功结果后重试、分别拒绝 XADD/XACK/XDEL 时无部分写且修复后可恢复、错误目标类型/消费组/ID/同键/空字段、正文 trim 后仍保留 PEL、超过 Lua 整数精度的 ID。日志:`/tmp/aether-round12-redis-transfer-tests.log`。取消/超时连接释放同时由状态模块中的前两轮 owned-driver 回归覆盖。
- `cargo check --locked -j 1 -p aether-gateway --bin aether-gateway` 通过,耗时 2 分 23 秒,无警告;日志:`/tmp/aether-round12-gateway-check.log`。修改文件格式及 diff 检查通过;测试、检查和本轮临时 Redis 进程均已退出,现有服务未变更。本轮未重复执行前轮因资源压力中止的完整代码生成和链接,网关可执行产物仍需后续完整构建;未部署,未进行目标环境业务 RPM 压测或全工作区测试。
本轮边界:
- 原子性和重复抑制按同一源 stream、消费组及源 ID 生效,不是永久去重索引。`NotPending` 只说明 PEL 不存在,外部 ACK、trim/delete、消费组销毁重建及其他消费者都可能改变这个状态,不能据此推断曾经归档。Redis XDEL 后其他消费组仍可能保留 PEL,Memory 沿用删除全部组 PEL 的既有语义;本轮不提供跨消费组的全局归档去重。
- Redis 转移要求 7+ 的 `redis.acl_check_cmd` 和 `EVAL/TYPE/XPENDING/XADD/XACK/XDEL` 权限。Lua 错误本身没有回滚能力;脚本预检当前已知可失败的前置条件,不替代 Redis 持久化及故障恢复保证。Cluster 两键必须同 slot,当前默认 stream 键没有自动改名或迁移;缺能力、权限或跨 slot 失败会保留源消息,不能自动降级。
- 编码预算覆盖任务所持原文长度与 JSON 上界,包含存储等待阶段,仍不是进程 RSS 或网络缓冲硬上限。字段容器、字符串多余容量、内存后端持久副本、Redis Cmd/packed-command/连接缓冲副本不计入;公开追加及旧后端可能沿用共享 driver。0/非法环境值回退默认,bytes 最大约 4 GiB,jobs 最大 128。保守 6 倍估算会拒绝实际编码较小的超大原文;这类存量需调高额度后重试或受控恢复,本轮未新增自动大消息通道。
- 已提交但响应丢失时可以避免重复归档,当前进程的成功计数可能少计;普通 ACK 成功后 XDEL 失败也可能少计 ACK,指标不是持久化账本。死信总量及保留期限仍需运营管理,本轮没有静默修剪死信。完整请求优雅排空、超大窗口增量聚合与目标环境业务 RPM 压测仍待处理;尚未部署。
## 第十三轮修复状态
- **TCP 层先限量:** 原入口每次 accept 后都创建独立连接任务,HTTP 请求限流尚未生效的握手及空闲 keep-alive 不受请求 gate 保护。新增 frontdoor `HttpConnectionBudget`,二进制入口的全部监听分片和指标使用同一个 Arc;accept 后立即尝试取得许可,满额直接关闭刚接入的 socket 并让出执行,不创建 HTTP 任务或额度等待队列。先 accept 再准入,避免无流量的 reuseport 分片预占许可、饿住有流量的分片。
- **许可跟随 socket:** 将许可放入底层 AsyncRead/AsyncWrite 包装,完整透传读、写、flush、shutdown 及 vectored write。socket 先释放、许可随后归还;HTTP/1 upgrade 会携带整个 IO 包装,所以连接 future 返回后 WebSocket 仍占一个许可,直到升级连接释放。HTTP/2 同一 TCP 内多路流共用一个许可,原 HTTP 请求和 WebSocket 会话 gate 保持独立;取消、解析失败及首个请求头超时也会释放底层 IO。
- **accept 错误恢复:** 原 `listener.accept().await?` 会退出任意出错的监听任务,主入口随后停止其他监听分片。改用共同的 accept helper:ConnectionRefused/Aborted/Reset 重试,其他错误记录后退避一秒再试,资源耗尽不会触发无等待的重复 accept。兼容 `serve_tcp` 入口也使用相同预算和恢复逻辑。
- **容量与观测:** 新增 `AETHER_GATEWAY_MAX_HTTP_CONNECTIONS` / `--max-http-connections`。未配置或 0 时按请求上限与 WebSocket 上限之和自动推导;自动和显式配置都限制在 1 至 65536,已知 FD soft limit 时再限制为 `max(1, (FD - 256) / 2)`。启动日志记录实际值,新增 `gateway_http_connections_limit/in_flight/high_watermark/rejected_total/accept_errors_total` 五项指标。
本地验证:
- `cargo test --locked -j 1 -p aether-gateway-frontdoor` 全部 43 项通过,无警告,包含本轮新增 9 项连接回归。覆盖默认/显式/0/FD 极值、不溢出的容量推导、IO 先于许可释放、拒绝立即消费 socket、读写及 vectored write/半关闭透传、两个真实 listener 共享 limit=1 后恢复、首次 poll 前取消、真实 HTTP/1 首头超时和解析失败、101 upgrade 后 HTTP future 已结束但升级 echo 仍持许可、真实 HTTP/2 两条并行请求共用一个连接许可,以及注入 accept 错误后继续服务、资源错误每次精确退避一秒。日志:`/tmp/aether-round13-frontdoor-tests.log`。暂停时钟的退避测试未真实耗尽进程 FD。
- `cargo check --locked -j 1 -p aether-gateway --all-targets` 最终通过,耗时 3 分 42 秒,无警告,覆盖主程序、库及测试/示例等目标;日志:`/tmp/aether-round13-gateway-final-check.log`。本轮另新增 2 项 AppState 共享预算和指标接线测试,已随全部目标完成编译检查,但没有重新生成及执行大型网关测试程序,不计入上述 43 项执行通过数。初次检查发现新增测试误用了不存在的 `for_tests` 方法,已按现有测试模式修正为 `AppState::new()`;初次退出码 101,不计为通过。
- 修改文件格式及 diff 检查通过。frontdoor 测试新增三个已锁定版本的开发依赖引用,Cargo.lock 仅增加对应依赖名,没有升级依赖版本。本轮测试、检查及临时 TCP 连接/任务均已退出,现有服务未变更;没有重复进行网关完整代码生成和链接,也没有部署或执行业务 RPM 压测。
本轮边界与后续发现:
- 这是已准入的入站 socket 数量上限,不是 HTTP/2 请求任务、全部 FD 或进程 RSS 上限。每个监听分片在 accept 与同步准入之间可短暂持有一个待判定 socket,kernel backlog、上游、Redis、数据库及其他入口不计入。容量规划仍需实测;健康检查使用同一入口,也可能在连接满额时被拒绝。
- 满额发生在 HTTP 解析前,客户端看到连接关闭而不是 429/503;客户端和反向代理的重试应有退避。成功请求结束后 keep-alive 连接继续占额度,直到连接释放;本轮没有新增空闲连接回收、强制流期限或完整请求优雅排空。现有超时及流式传输语义保持原样。
- 公共兼容 `serve_tcp` 每次调用使用独立预算,默认 4096,可用同名环境变量覆盖,最大 65536;它没有二进制入口的请求/WS 容量配置和 FD 探测,也未将预算绑定到其默认 AppState 的五项指标。二进制入口拥有 FD 约束、跨分片共享及指标接线。
- 另确认 usage 窗口 Lua 每次准入全量清理过期成员,集中到期可能阻塞 Redis。不能直接换成分批删除加当前窗口 ZCOUNT:合法并发请求的时间参数可能乱序到达,后来的较早 cutoff 会重新计入此前逻辑上已过期但未物理删除的成员。保持精确额度需配套单调清理水位等状态设计;本轮没有修改这段逻辑。相同终态的事件级及 stored-record 级成本预留协调可能重复取得 PostgreSQL 用户行锁,后续可研究携带匹配身份的首次结果,不能无条件删除任一调用。
- usage 独立 runtime 的本地重试渠道仍缺少完整停止与排空生命周期;当前数量有界,但进程退出前未入 Redis 的缓冲终态仍需统一处理。本轮没有修改停机流程、迁移架构或部署,业务 RPM 与长期排空压测仍待执行。
## 第十四轮修复状态
- **消除同次写入的重复成本协调:** worker 和直接落库原来先按事件协调成本预留,upsert 返回后又按存储记录协调一次,正常终态重复获取 PostgreSQL 用户行及预留行锁。现在首次调用返回的持久化结果同时匹配 request ID、用户、服务端预留 token、实际费用单位、终态和完成秒数时,携带一个仅限本次写入的内部结果;存储记录再次匹配全部字段才省去第二次协调。钱包结算、原用户级锁、费用检查及幂等性继续执行。
- **保留冲突和重试语义:** 首次返回 `None`、返回其他终态或身份、upsert 后记录变化,都继续按存储记录重新协调。首次协调出错仍在 upsert 前停止;upsert 失败后的新尝试重新协调,不跨请求或重试缓存结果。公开协调与结算 API 签名不变;没有更改数据库表或放宽账务条件。
- **修正验证工具:** PostgreSQL 结算基线补齐有余额的钱包及所属用户,并核对结算记录、快照、outbox、钱包消费和供应商累计值,检查失败返回非零退出码。此前无钱包的全量拒绝不能算吞吐成功。网关夹具补齐启动流程会创建的默认路由,通过 testkit 对象读取受保护的指标,缺失必要指标直接报错;容量报告保留 HTTP 状态和错误样本。探针新增排空基线与最终必要指标字段,判定阈值不变。网关阶梯夹具使用与主程序相同的 8 MiB Tokio worker 栈,避免 debug 路由调用链在默认小栈上溢出。
- **隧道夹具按当前入口运行:** 仅在 `testkit` 功能下增加适配入口,将测试请求交给现有网关签名验证、正文完整性校验和临时 spool 流程,再进入内部 relay;没有开放跳过鉴权的 HTTP 路由。每条压力请求使用独立签名和 nonce。每个档位先检查合法签名成功并返回完整正文,无签名、篡改正文及重放均被拒绝;模拟节点通过并发计时等待响应,避免旧夹具在单个 WebSocket 读循环逐条 sleep 造成串行瓶颈。每个档位结束后显式取消并等待模拟节点任务。
本地验证:
- `cargo test --locked -j1 -p aether-usage-runtime`:349 项全部通过,本轮新增 8 项,包含多字段不匹配矩阵、失败/取消/零费用、失败重试、256 个同用户并发写入和 32 次重复投递。重复投递使用真实 Memory 结算仓库,余额只扣一次。日志:`/tmp/aether-round14-usage-tests.log`。`gateway_pressure_probe` 既有 21 项测试全部通过,日志:`/tmp/aether-round14-probe-tests.log`。
- 本轮已成功完整构建并链接网关、容量基线、结算基线和压力探针,补上第十至十三轮只做编译检查而未更新可执行文件的验证。主要日志:`/tmp/aether-round14-full-build.log`、`/tmp/aether-round14-final-harness-build.log`;无编译警告。
- 主网关真实 TCP 冒烟:2 个监听分片共享连接上限 4,保持 4 条空闲连接后,另外 32 条连接在共 7 ms 内被关闭;释放后健康检查恢复 200,认证指标确认 high-watermark=4、rejected=32。网关正常退出,日志:`/tmp/aether-round14-connection-smoke.log`。
- 隔离 PostgreSQL 14 结算热点:同一用户、同一钱包,100 并发完成 2000 笔真实结算,0 失败;2000 条 usage、结算快照及已处理 outbox 完整,待处理为 0,钱包消费与供应商费用均为 2.00 USD。耗时 2332 ms,857 次/秒,P95=251 ms;采样锁等待最高 62 个、最长 549 ms,同一钱包仍需串行扣费。报告:`/tmp/aether-round14-postgres-settlement-final.json`。
- 网关同步/流式、执行运行时同步/流式和隧道共 5 个阶梯场景,在 8/32/128/256 的 gate 上限各执行 8 倍请求,共 16928 条全部 HTTP 200,无拒绝或读取失败。前四场景各 3392 条,隧道共 3360 条,隧道并发为上限减 1,因为 WebSocket 会话占一个许可。最终前四场景在途归零,隧道采样时只剩该会话,high-watermark 均达到对应上限。每个档位的签名、正文和重放预检通过。报告:`/tmp/aether-round14-capacity-final.json`;最终构建日志:`/tmp/aether-round14-capacity-auth-final-build.log`。这些场景是 Memory 仓库/模拟节点功能和容量验证,与下面真实数据库业务压测分开解读。
- 主网关 + 隔离 PostgreSQL/Redis + 本地模拟上游,8 个 API key、4 个网关 Tokio worker、2 个监听分片、HTTP 请求上限 256、TCP 上限 512、数据库池上限 24,执行以下约 2 秒的流式请求,要求完整读取且出现 SSE `[DONE]`:
| 并发 | 请求数 | 成功 / 失败 | 实测 RPS | 首字节 P95 | 完整响应 P95 | 排空判定耗时 |
| --- | --- | --- | --- | --- | --- | --- |
| 16 | 128 | 128 / 0 | 8.01 | 123 ms | 2062 ms | 9.93 s |
| 64 | 512 | 512 / 0 | 32.09 | 69 ms | 2046 ms | 9.86 s |
| 128 | 1024 | 1024 / 0 | 57.34 | 106 ms | 2074 ms | 119.21 s |
业务压测报告目录:`/tmp/aether-round14-pressure.UYCeSq`。全部 1664 个请求为 HTTP 200;SQL 验证全部 completed,input/output/total tokens 分别为 1664/33280/34944,没有缺失首字节时间或单条 token 不匹配。所有档位必要排空指标齐全并通过连续安静期检查,最终用量队列、PEL、DLQ、outbox、请求在途和数据库锁等待均为 0。128 并发档网关 RSS 峰值约 190 MiB、末值约 147 MiB,FD 从峰值 338 降至 62,Tokio 活跃任务降至 85;没有要求 RSS 或常驻后台任务归零。
本轮边界:
- 本地 debug 构建、模拟上游和短时压测不能推导生产 RPM 上限,也没有同环境修复前后的性能对照。模拟上游费用为 0,非零钱包金额另由上述 PostgreSQL 热点及 Memory 幂等测试验证。临时 PostgreSQL 夹具关闭 fsync、同步提交及 full-page writes,热点吞吐数不代表开启生产持久化后的性能。
- 阶梯夹具使用 75 ms 模拟同步/隧道响应或 3 次 25 ms 流式间隔。256 并发时网关同步/流式 P95 分别为 770/808 ms,超过该夹具的 300 ms 延迟预算,吞吐从 128 并发的 596/754 RPS 降至 537/524 RPS;全部请求成功不等于此档位容量合格。执行运行时 P95 为 104/96 ms,隧道为 115 ms。需在目标环境定位网关层 CPU、调度及后台写入开销后决定并发上限,不应直接按本地峰值放大配置。
- 128 并发后的完整排空约需两分钟;此前 45 秒检查及一次 120 秒检查未通过,不能作为快速回收的证据。当前复测确认最终回收,但仍需目标环境下长期压力、连接空闲回收和停机排空验证,不能归因为已确认的泄漏或宣称完全消除积压风险。
- 第十三轮发现的 Redis 过期窗口集中清理和 usage 本地缓冲优雅停止仍待后续处理。尚未部署、未调整生产配置,未执行全工作区测试。
本轮构建、测试、压测及临时服务均已退出,原有开发服务未改动;格式及 diff 检查通过。验证过程中出现的缺失默认路由、旧指标入口、未签名 relay、夹具小栈溢出和流 ID 类型编译错误均已修正;此前失败或全量拒绝的报告不计入上述通过结果。
## 第十五轮修复状态
- **整窗过期时异步释放:** Redis 用量准入脚本原来逐条同步删除整个过期 sorted set,集中到期的大集合会长时间占用命令线程。现在对超过 256 条的集合检查首尾时间:首条未过期时跳过清理,末条也已过期时在同一次 Lua 调用内 `UNLINK` 整个键,让 Redis 在后台释放旧对象,随后仍按原规则准入并写入新事件。小集合及混合新旧记录的窗口继续精确删除,空集合直接使用计数 0。
- **保留时间和事务语义:** 所有规则先清理,再检查,最后全部消费;任一规则拒绝时仍不写入新事件。后续规则的过期清理不会被前面规则的拒绝跳过。没有新增清理水位、辅助键或改变事件时间戳;乱序时间、重放、释放补偿、`Retry-After` 和键 TTL 沿用旧语义,无需迁移 Redis 数据。`UNLINK` 不可用或被 ACL 拒绝时退回原同步删除,不把未清理的旧记录当作已清理。
- **减去请求内重复工作:** 清理后计数只在当前 Lua 调用内复用,不再重新 `ZCARD`;Rust 端用 `OnceLock` 复用不可变脚本对象,避免每个请求复制脚本文本和计算 SHA1。每次调用仍单独构造键和参数;Redis 脚本缓存丢失时仍由现有客户端处理 `NOSCRIPT` 并重新加载。
本地验证:
- `cargo test --locked -j1 -p aether-runtime-state -p aether-usage-runtime`:状态模块 110 项、用量模块 349 项全部通过,共 459 项,0 失败;1 项较大数据计时基线默认忽略,已另行显式执行通过。日志:`/tmp/aether-round15-regression-tests.log`。
- 新增 8 项真实 Redis 回归及 1 项计时基线。单独执行新增回归时确认 11 次隔离实例就绪,无 fixture 跳过,覆盖 Redis 8.0.3、RESP2/RESP3、ACL 认证及数据库 6。验证 10 万条记录一次 detach 后键复用、截止边界及 Lua 精确整数上界、混合窗口保留有效项、首规则拒绝时后续窗口仍清理、拒绝 UNLINK 后兼容、64 个并发争抢只放行 8 个、`SCRIPT FLUSH` 后幂等重载,以及 2 个协议各 512 步新旧脚本差分对照。对照逐步核对判定及完整剩余记录,包含时间回退、重复事件、额度变化和释放补偿。日志:`/tmp/aether-round15-redis-verified.log`。
- 隔离 Redis 的 3 次对照各装入 30 万条同时过期记录。旧准入调用耗时 51.968 / 49.395 / 51.431 ms,新路径为 0.576 / 0.222 / 0.320 ms;另外装入 4096 条有效记录,各执行 1000 次持续准入,旧路径 P50/P95=86/123 us,新路径为 74/111 us,最终记录一致。日志:`/tmp/aether-round15-cleanup-timing-final.log`。测试验证逻辑和数据一致性,不以容易受本机负载影响的耗时阈值作为断言。
- `cargo build --locked -j1 -p aether-gateway --bin aether-gateway` 完整构建与链接通过,耗时 4 分 15 秒,无警告;日志:`/tmp/aether-round15-gateway-build.log`。全工作区格式检查、新增 include 测试文件的独立格式检查及 diff 检查通过;本轮构建、测试和临时 Redis 实例均已退出,原有开发服务未改动。
本轮边界:
- 这是整窗过期的优化,不是所有清理操作的硬耗时上限。混合窗口仍会同步删除其全部过期项;若大量旧记录与有效项同时存在,阻塞风险仍在。分批逻辑清理需要额外状态与完整的乱序、重试和滚动升级设计,不能直接以当前 cutoff 计数代替物理删除。
- Redis 自动 TTL 过期的释放由服务端过期策略决定,可能先于本脚本发生,不由此改为异步。UNLINK 加速也依赖服务端支持及 ACL 许可;兼容回退仍有原同步释放成本。异步释放期间旧对象仍可能占内存,本轮没有提供 Redis 内存或 RSS 硬上限。
- 数字来自本机 debug 构建、隔离 Redis 和单窗口局部对照,不代表网关端到端吞吐或生产 RPM 容量;未修改生产 Redis 配置、未部署,未重复上一轮业务压测。usage 停机排空及前轮观察到的高并发延迟仍待处理。
## 第十六轮修复状态
- **混合窗口的稀疏存活项重建:** 窗口超过 1024 条、有效项为 1–256 条且过期项超过有效项四倍时,先复制有效成员及原始 score 到临时键,再在同一次 Lua 内 `UNLINK` 旧集合并替换。先验证权限并完整构建临时集合,保留原 TTL;临时键已存在、Redis 缺少 ACL 检查接口或可选命令受限时继续精确同步清理。无清理水位或持久辅助数据,乱序、重放、多规则拒绝和释放语义不变。有效项超过 256 条的混合窗口仍走同步路径,不能据此宣称所有窗口清理已有硬耗时上限。
- **请求与用量统一停机:** 收到退出信号后停止所有监听分片接收请求,HTTP/1 和 HTTP/2 等待已接收响应;默认 30 秒到期后取消剩余连接任务,并使升级后的连接读写返回关闭错误。对可能脱离 HTTP 生命周期的流式终态、取消回调及 WebSocket 审计,在创建后台任务前登记 producer。另将整条代理请求及响应体纳入登记,覆盖响应头前断开、内联响应体及断开后的后台消费,避免先关闭用量入口再提交终态的竞态。路由的 `cancel_on_client_disconnect=false` 策略保持继续完成上游请求,后台处理受后续用量排空期限约束;配置为 true 时按原规则取消并持久化取消终态。
- **排空本地用量缓冲:** 等待 producer 结束后关闭生命周期入口,立即冲刷延迟事件并唤醒重试退避,等待各阶段与候选记录写入完成。已确认写入 Redis 的事件可留待下次消费;Memory 队列必须实际消费,存在未恢复的本地 DLQ 不判定为成功。worker 完成当前写入及 ACK 后停止,随后结束空闲分发任务和专用 Tokio runtime。超时明确报错;运行库层面保留未完成工作供再次调用停机,主进程不会把超时打印成排空成功。
- **缩短上游空闲连接滞留:** reqwest 和浏览器传输的每客户端、每 origin 空闲连接默认上限由 1024 降至 32,默认空闲期限为 15 秒;H2C 池默认上限由 512 降至 32,并补上主动清理所需 timer。活动请求及响应流不受空闲期限限制。配置项为 `AETHER_GATEWAY_UPSTREAM_POOL_MAX_IDLE_PER_HOST` 和 `AETHER_GATEWAY_UPSTREAM_POOL_IDLE_TIMEOUT_MS`;HTTP、用量停止期限分别用 `AETHER_GATEWAY_HTTP_SHUTDOWN_TIMEOUT_MS`、`AETHER_GATEWAY_USAGE_SHUTDOWN_TIMEOUT_MS` 配置,均默认 30000 ms。
- **保留恢复过程证据:** 压测探针记录有界的逐次排空观测,包含用量 producer、延迟事件、业务队列、Tokio 任务、连接和 FD。必要指标、任务基线容差以及连续 7 秒安静期保持原判定,新增本地待处理指标有值时也必须归零。
已完成的本地验证:
- frontdoor 45 项、用量运行库 360 项、负载工具库 17 项、压力探针 21 项、状态模块 113 项、网关主程序 66 项、请求生命周期 8 项、传输模块 115 项及流式断开策略 1 项,共 746 项通过。状态模块另有 1 项默认忽略的计时基线,已显式执行通过。新增停机验证覆盖 producer 竞态、并发终态、重试失败后再次停止、延迟事件、多个 worker、写入途中停止、Memory 队列、响应头前断开、后台响应体、HTTP/1 和 HTTP/2 挂起处理器的实际释放,以及升级连接。日志:`/tmp/aether-round16-final-regressions.log`、`/tmp/aether-round16-state-regression.log`、`/tmp/aether-round16-main-final-tests.log`、`/tmp/aether-round16-request-lifecycle-tests.log`、`/tmp/aether-round16-transport-tests.log`。流式策略测试首次使用默认小栈时发生栈溢出;按网关入口相同的 8 MiB 栈设置 `RUST_MIN_STACK=8388608` 后通过,日志:`/tmp/aether-round16-disconnect-policy-test.log`。
- Redis 新增混合重建测试覆盖 RESP2/RESP3、1/16/256/257 个有效项、TTL、非过期键、重复提交、已有临时键,以及 ZCOUNT/PTTL/EXISTS/UNLINK/RENAME/PEXPIRE 权限受限后的精确回退。定向 11 项测试确认 15 次隔离 Redis 实例就绪,无 fixture 跳过;原有 1024 步新旧脚本差分测试继续通过。日志:`/tmp/aether-round16-redis-tests.log`。
- 同机 Redis 8.0.3 对照,30 万条过期记录分别混合 1/64/256 个有效项,旧脚本为 48.009/54.660/52.924 ms,新脚本为 0.320/0.377/0.496 ms,最终成员完全一致。整窗过期 3 次为旧脚本 56.802/51.804/50.886 ms、新脚本 0.744/0.258/0.318 ms。4096 条有效记录的持续调用 P50/P95 为旧脚本 87/131 us、新脚本 76/123 us;计时不作为断言阈值。日志:`/tmp/aether-round16-redis-timing.log`。
- 网关主程序和压力探针完整构建、链接通过,无警告,日志:`/tmp/aether-round16-final-build.log`。补齐请求生命周期登记后再次完整构建主程序通过,日志:`/tmp/aether-round16-final-gateway-build.log`。
真实入口压力和停机验证:
- 使用隔离 PostgreSQL 14、Redis 8.0.3 和约 2 秒的流式模拟上游,8 个 API key、4 个网关 worker、2 个监听分片、请求上限 384、TCP 上限 768、数据库池上限 24。要求完整读取且出现 SSE `[DONE]`。本轮 PostgreSQL 未关闭持久化,启动对照实查 `fsync`、`synchronous_commit`、`full_page_writes` 均为 on;Redis 夹具关闭 AOF/RDB,因此只验证保留 Redis 进程时的网关重启恢复。
- 连接池单变量对照使用同一构建,仅设置旧空闲上限/期限 1024/90000 ms 或新默认 32/15000 ms。旧配置的 128、256 并发均全部 HTTP 200,业务待处理指标分别约 2.8/5.1 秒归零,但两档都未通过 150 秒恢复检查,最后 Tokio 活跃任务为 229/350、FD 为 206/327;脚本返回失败,不记为恢复通过。新配置分别在 22.217/26.542 秒通过原有恢复判定,FD 降至 65/63。这确认本地长时间恢复的主要滞留来自空闲连接及相关任务,而不是业务队列持续积压。对照目录分别为 `/tmp/aether-round16-pressure.3xsT7P`、`/tmp/aether-round16-pressure.iysPcj`;不把一次对照的吞吐或首字节波动视为可靠容量提升。
- 最后补齐整个代理请求的 producer 登记后,使用最终构建再次执行新配置,结果如下。必要排空指标齐全,最终队列 lag、PEL、DLQ、outbox、本地用量待处理及数据库锁等待均归零,无 worker 处理失败;恢复判定包含连续 7 秒安静期。
| 并发 | 请求数 | 成功 / 失败 | RPS | 首字节 P95 | 完整响应 P95 | 恢复判定耗时 | 最终 FD / Tokio 任务 |
| --- | --- | --- | --- | --- | --- | --- | --- |
| 128 | 1024 | 1024 / 0 | 56 | 150 ms | 2141 ms | 22.325 s | 65 / 87 |
| 256 | 2048 | 2048 / 0 | 115 | 155 ms | 2192 ms | 26.243 s | 63 / 85 |
最终报告目录:`/tmp/aether-round16-pressure.JI8vPF`,执行日志:`/tmp/aether-round16-pressure-final-new.log`,验证脚本:`/tmp/aether-round16-pressure.sh`。该轮 256 并发 FD 峰值 595、最终 63;RSS 峰值约 283 MiB,检查结束仍约 282 MiB,没有把短时间 RSS 不下降判为泄漏或声称内存已回到启动值。
- 正常 SIGTERM:32 条已开始输出的流全部完整结束,网关成功退出;重启同一数据库与 Redis 后,核对 3104 条 completed,逐条 input/output/total tokens 为 1/20/21,首字节时间无缺失。
- 将 HTTP 停止期限缩短为 50 ms:默认继续完成策略下,32 个客户端都观察到流中断,网关继续完成其后台请求后成功退出;恢复后 3136 条全部 completed,token 合计 3136/62720/65856,逐条仍为 1/20/21。随后仅在测试数据库启用断开取消策略,另外 32 条流中断后全部落为 cancelled、HTTP 499、billing_status=void、token 为 0,首字节时间保留。最终数据库为 3136 条 completed 加 32 条 cancelled,未把默认继续完成误判为取消。
- 最终脚本对 SQL 完整性失败显式退出。早期验证脚本误将默认断开策略预期为取消,并遇到旧 Bash 的条件失败未中止问题;已修正断言和预期,上述停机终态与 token 结论只采用最终目录的复测结果。
全工作区格式检查、新增 include 测试文件的独立格式检查及 diff 检查通过。本轮构建、测试、压测、临时 PostgreSQL/Redis 和模拟服务均已退出,原有开发服务未改动。
边界:
- 新连接池设置只限制空闲缓存,不能推导整个进程的 FD 或 RSS 上限,也不能替代活动请求并发预算。减少空闲连接会增加突发流量之间重新建立连接的次数;需要在目标环境观察握手成本及连接复用。
- 有界停机无法保证 SIGKILL、依赖持续故障或超过停止期限时的本地未持久数据不丢失;本轮未引入磁盘日志或更换消息架构。进程管理器的强杀期限应覆盖两阶段停止和额外收尾时间。
- 第十四轮短响应夹具在 256 并发的延迟超预算尚不能判定已解决;本地 debug、模拟上游和短时压力结果不代表生产容量。未部署或修改生产配置,未执行全工作区测试。
## 第十七轮修复状态
本轮继续处理混合过期 Redis 大窗口和短响应网关的历史记录读取开销,未部署生产。
- **Redis 混合大窗口:** 当过期项超过 4096、有效项超过 256 时,先执行只读规划,随后在独占连接上 WATCH 全部规则,每条命令最多复制 512 个有效成员;通过 MULTI/EXEC 校验期间没有源数据变化,最后在一个 Lua 调用中替换窗口并完成原有多规则检查与消费。旧进程的 ZADD/ZREM、TTL 变化也会使 WATCH 失效。复制期间不提前清理其他规则,避免乱序请求、重复事件和跨规则拒绝改变原有额度语义。
- **维护资源限制:** 每个 Redis runtime router 最多两条按需维护连接,最多尝试八次,包含连接池等待的总耗时受原有命令超时约束;未配置超时时使用 30 秒上限。取消或失败会丢弃仍有 WATCH/MULTI 状态的独占连接,未提交副本有 60 秒 TTL。成功提交后没有额外的可失败网络收尾操作。缺少所需 ACL 能力或 Redis 不支持 ACL 预检时保留原有精确清理路径。
- **候选内存仓库索引:** 请求更新不再全表查找逻辑主键,按 request_id 查候选;最近记录查询改用创建时间索引,在复制之前取 limit。三个索引在同一写锁中维护,覆盖 ID 替换、相同时间排序和删除空索引。所有写入入口保留清洗,读取不再重复重建已经清洗的 JSON。
- **公共候选诊断清洗:** 提取借用 JSON 对象的内部函数,移除持久化清洗前的整份输入克隆、清洗结果克隆及诊断对象的再次克隆;字段白名单、诊断大小限制、管理员与公开投影语义不变。PostgreSQL 与 Memory 都调用这一公共函数。
- **调度读取精简投影:** 新增 `list_recent_runtime` 读取身份、状态、计数及时间,Memory 从时间索引直接构建不带诊断的记录,PostgreSQL 查询不读取诊断和能力 JSON。调度与自适应 RPM 观察使用此路径,管理员和请求可观测性继续读取完整记录。精简前后并发计数、RPM 和失败冷却判断保持一致。
已完成 241 项检查:状态模块 119 项、候选仓库与数据契约 28 项、调度核心 92 项、隔离 PostgreSQL 14 精简投影检查 1 项,以及单独执行的 Redis 大窗口计时检查 1 项。PostgreSQL 检查使用独立临时实例和连接级临时表,包含 32 KiB 诊断字段,验证精简投影与完整记录的运行时字段一致,且完整读取仍保留诊断。最终网关压测程序编译通过;临时数据库均已停止。日志:`/tmp/aether-round17-state-complete-tests.log`、`/tmp/aether-round17-projection-tests.log`、`/tmp/aether-round17-scheduler-tests.log`、`/tmp/aether-round17-postgres-live.log`。
首次短响应采样确认调度等待 `list_recent` 的全局读锁,读路径原本复制全部历史记录再逐条重做 JSON 清洗;仅修索引后最近 128 条完整 JSON 的复制与清洗仍是热点,因此继续将调度读取改为精简投影。最终采样仍有候选仓库读写锁等待,但没有原先全量历史复制的放大行为,其他开销分布在请求处理、凭据解密和 JSON 构建。本机其他进程负载较高,不能直接与第十四轮数字比较。
短响应对照使用本轮修改前保留的二进制与最终二进制,同为本机 debug 构建、Memory 网关夹具、模拟上游约 75 ms、每点请求量为并发数的八倍。最终版本及随后复跑的对照版各执行 15344 个请求,全部成功,无失败或拒绝。网关 P95 单位 ms:
| 场景 | 并发 | 修改前复测 P95 | 修改后 P95 | 修改后 RPS |
| --- | --- | --- | --- | --- |
| 同步 | 128 | 1025 | 222 | 857 |
| 同步 | 256 | 2049 | 408 | 800 |
| 流式 | 128 | 569 | 263 | 678 |
| 流式 | 256 | 1564 | 451 | 794 |
128 并发达到该夹具的 300 ms 预算;256 并发两种网关路径仍超过预算。最终版本中独立执行器和隧道曲线的 P95 均低于预算。结果位于 `/tmp/aether-round17-capacity-final.json`、`/tmp/aether-round17-capacity-control-recheck.json`。本机对照时系统 CPU 约 92%–99%,同一进程包含压测客户端、网关与模拟上游,不能将这里的 RPS 视为生产容量。用于 CPU 采样的额外长测有采样器干扰,未用于上述性能对照。
Redis 8.0.3 同机前后对照,均含 300000 条过期记录,单位 ms:
| 有效项 | 旧整次调用 | 新整次调用 | 旧最长命令 | 新最长命令 |
| --- | --- | --- | --- | --- |
| 257 | 64.750 | 2.199 | 64.351 | 0.218 |
| 4096 | 77.099 | 7.331 | 76.301 | 0.681 |
| 32768 | 62.999 | 45.464 | 62.583 | 0.914 |
| 150000 | 77.200 | 242.083 | 76.840 | 1.585 |
最终成员与旧脚本一致;最长命令来自隔离 Redis 的 SLOWLOG,包含 EVAL/EXEC。有效项多时整次清理更慢,但复制批次之间允许其他命令执行,降低对同一 Redis 其他请求的连续阻塞。原有整窗过期和不超过 256 个有效项的快速路径继续通过。4096 条有效项持续调用 1000 次,旧 P50/P95 为 155/286 us,新为 152/310 us;未将计时作为测试断言。记录:`/tmp/aether-round17-redis-timing.log`。
边界:复制需要临时保存有效成员,额外内存与有效项数量相关;频繁并发修改可能触发重试或达到命令期限,未提交的复制不会替换原始窗口。提交成功但回复丢失仍具有现有 Redis 命令的结果不确定性,应使用同一事件 ID 重试。ACL 不足或不支持预检的 Redis 仍走同步回退,不应把本优化描述为所有配置下的 Redis 总阻塞硬上限。代码层面的已复现放大路径已修复,但不能据此宣称所有容量问题均已解决;256 并发的最终验收仍需在代表性环境运行 release 构建并结合持续压测确认。未部署生产。
## 主要判断
系统已经有入口并发门、请求体内存预算、认证缓存合并、前后台数据库池隔离及多种有界队列。问题主要在于部分请求路径仍执行全量统计、锁范围过大,以及资源预算和超时没有覆盖完整生命周期。
这些问题会相互放大:每请求成本增加,使请求停留更久;在途请求增加后继续竞争 Redis、数据库和内存,客户端重试又增加负载。只提高入口并发数或数据库连接数不能消除这些瓶颈。
平均在途请求约为 `RPM / 60 × 平均请求持续秒数`。例如 600 RPM、平均持续 60 秒,约有 600 个在途请求;该例是容量计算,不是当前实例测量值。
## P1-1:压缩或未知长度请求一次占满全局请求体预算
**代码:**
- `crates/aether-gateway/frontdoor/src/body.rs:392`:带非 identity Content-Encoding,或缺少 Content-Length 时,预留 `min(单请求上限, 全局预算)`。
- `apps/aether-gateway/src/state/app.rs:40`:默认全局预算 256 MiB;`apps/aether-gateway/src/headers.rs:20` 的单请求默认上限同为 256 MiB。
- `apps/aether-gateway/src/state/app.rs:195`、`crates/aether-gateway/frontdoor/src/body.rs:201`:默认没有请求体完整读取超时。
- `crates/aether-gateway/frontdoor/src/body.rs:154`:等不到预算则返回过载;默认等待预算 250 ms。
**结果:** 即使压缩后的请求只有 1 KiB,也会预留全部 256 MiB。压缩上传或未知长度上传的读取阶段因此被串行化;一条一直未传完的请求可以长期占住预算,其他需要读取请求体的请求陆续收到 503。这里的 256 MiB 是预留额度,并非立即分配的实际内存。
**验证:** 使用实际 `BodyBufferPolicy` 的独立程序复现:1 KiB gzip 请求预留 268435456 字节,剩余许可为 0;第二个普通 1 KiB 请求等待约 252 ms 后被拒绝,释放第一个预留后立即恢复。程序显式采用默认参数,验证的是预算准入,不是完整 HTTP 慢上传压测。源码位于 `/tmp/aether-frontdoor-budget-repro.rs`,可执行程序位于 `/tmp/aether-frontdoor-budget-repro`。`cargo test -p aether-gateway-frontdoor body::tests -- --nocapture` 的 13 项既有请求体测试通过。
**建议:** 设置适合实际上传大小和速率的读取期限;根据业务明确单请求大小上限,使一个普通上传不能耗尽全局预算。进一步改为分段、有界读取及解压预算,或为大型上传分配独立额度。增量预留必须避免多个请求各持部分预算、同时等待扩容导致死锁,不能简单取消当前保护。
## P1-2:调度请求执行 Redis 全库扫描和历史样本聚合
**代码:**
- `apps/aether-gateway/src/dispatch/pool_scheduler.rs:142`、`:985`:候选分页和 sticky 路径调用管理侧 runtime 统计函数。
- `apps/aether-gateway/src/handlers/admin/provider/pool/runtime/reads.rs:126`:cache affinity 且 sticky TTL 非零时,扫描所有匹配会话,再 MGET 全部结果。
- `crates/aether-runtime/state/src/redis/runtime.rs:289`:SCAN 循环直到游标归零,COUNT 200 是每轮提示,不是总量上限。
- `apps/aether-gateway/src/dispatch/pool_scheduler.rs:1984`:扫描生成的会话总数和按 key 统计并不进入实际调度状态。
- `apps/aether-gateway/src/handlers/admin/provider/pool/runtime/reads.rs:208`、`:225`:无条件对每个候选 key 拉取成本和延迟窗口的原始成员,即使相应排序或限额未启用。
**结果:** 一页 64 个不同候选 key 就产生 128 次窗口查询,另加会话扫描等操作。成本窗口默认 5 小时,历史样本按时间修剪;随着 RPM 增加,每个请求需要读取和聚合的历史数据也增加。`join_all` 并发等待不会消除 Redis 执行量及返回数据量。SCAN 和窗口读取都使用 Admin lane,管理请求也可能受到影响。
**建议:** 为调度建立独立读取接口,只读取当前 sticky 绑定及调度需要的状态;会话总数留给管理统计。按启用策略读取成本或延迟,改用增量计数、时间桶或后台维护的短时快照;保留严格额度检查的原子性。对相同池的刷新合并,避免每个请求重复拉取历史。
## P1-3:结算错误地锁住公共套餐行,跨用户串行
**代码:** `crates/aether-data/adapters/postgres/src/settlement.rs:482` 的 `user_plan_entitlements JOIN billing_plans ... FOR UPDATE` 没有限定锁的表。
**触发:** 普通用户的非零费用结算,且用户有有效套餐。即使套餐没有 daily_quota,查询也先锁行,之后才判断 `grants.is_empty`。
**结果:** 不同用户只要使用同一个套餐,就会争抢同一条 `billing_plans` 行。锁持续到事务结束,其间还可能汇总每日账本、写额度流水、更新钱包和结算快照。首先影响后台结算和队列消化速度,不能据此断言前台数据库池必然同时耗尽。
**验证:** 在隔离 PostgreSQL 14.17 中,两用户拥有不同 entitlement、共享一个 plan。事务 A 持原 SQL 锁时,事务 B 因 500 ms lock_timeout 失败,错误明确指向 `relation "billing_plans"`;对照使用 `FOR UPDATE OF user_plan_entitlements` 后事务 B 成功。临时 PostgreSQL 已停止。该验证证明锁冲突,不是完整结算吞吐压测;部署 Compose 使用 PostgreSQL 15。
复现脚本:`/tmp/aether-plan-lock-repro-20260909.sh`;本次日志:`/tmp/aether-plan-lock-repro.XoYxfO/`。
**建议:** 将锁范围限定为需要更新的用户 entitlement;明确套餐配置并发变更的一致性规则。添加真实双连接回归:同用户额度不能重复扣,不同用户同套餐不能互相阻塞。
## P1-4:首包之后的流缺少空闲期限
**代码:**
- `apps/aether-gateway/src/execution_runtime/stream/execution.rs:2699`:默认 inline 直通路径仅在首包前使用 timeout,首包后直接 `upstream.next().await`。
- `apps/aether-gateway/src/execution_runtime/transport.rs:3375`:stream 不使用总请求 timeout。
- `apps/aether-gateway/src/execution_runtime/transport.rs:4212`:该 reqwest 客户端配置连接超时,没有设置读取超时。
**结果:** 上游已经发送首包、随后不再发送数据也不关闭连接时,只要下游继续保持连接,该流就可能长期占据请求名额、provider 并发守卫及缓冲。坏流逐渐积累会压缩可用容量。target permit 在首次向客户端 yield 时释放,不能把它算作整条流一直占用的资源。
**建议:** 增加可按 provider 配置的上游空闲期限;计时以真实上游活动为依据,网关自己的 keepalive 不应重置它。超时后关闭上游并走一次终态结算,确保断开、取消和超时都释放请求及 provider 守卫。对合法长思考模型采用匹配其行为的阈值。
## P1-5:并发上限与实际常驻内存不匹配
**代码:**
- `apps/aether-gateway/src/main.rs:433`、`:467`:自动入口上限为每 CPU 1024,结合 FD 下调,但没有内存预算。
- `apps/aether-gateway/src/execution_runtime/stream/execution.rs:167`、`:400`:Basic 模式每份流分析缓冲上限 5 MiB;Full 为 64 MiB。
- 同文件 `:1993`、`:2031`:分别累积 provider 和 client 两份 body,包括直通路径。
**结果:** 两份捕获缓冲达到上限时,Basic 单流约 10 MiB,Full 约 128 MiB,尚未计入请求 JSON、转换状态、队列及其他内存。1000 条都达到 Basic 捕获上限的流,仅这两份内容就约 9.77 GiB。这是达到上限时的预算估算,不是普通小回复的固定内存或已测 RSS。
**建议:** 用实测每请求内存和 cgroup/物理内存确定入口容量;为响应捕获建立全局字节预算和截断策略。Basic 优先使用增量 usage/error 解析,直通时避免重复捕获同一内容。请求体读取预算不能充当整个流生命周期的内存保护。
## P1-6:审计压缩在 Tokio 线程和数据库事务内同步执行
**代码:** `crates/aether-data/adapters/postgres/src/usage/mod.rs:8457` 开始事务并锁 request;`:8522` 同步准备审计内容;`:12362` 序列化 JSON,`:12376` 执行 gzip level 6。`:12478` 附近最多处理四份请求/响应 body;inline 阈值为 0。
**结果:** 有 body 捕获,尤其大上下文、高并发时,CPU 压缩同时占据 Tokio 工作线程、数据库连接和事务锁。前后台连接池隔离不能隔离同一进程内的 CPU 和内存争用。
**建议:** 在开启事务前完成可独立准备的序列化和压缩,放到有并发和字节预算的 blocking worker;事务内只保留必须原子执行的读写。检查相同 body 的去重,避免单纯增加 worker 数量导致 CPU 和内存进一步饱和。
## P1-7:长期额度准入按用户串行扫描,且默认无锁等待期限
**代码:** `crates/aether-data/adapters/postgres/src/settlement.rs:276` 锁 `users` 行;`:611`、`:850` 在准入中取得该锁后,分别逐窗口 COUNT 请求预留或 SUM 成本预留。释放及对账还会争同一用户锁。`crates/aether-data/adapters/postgres/src/tx.rs:15`、`:116` 的默认读写事务不设置 lock_timeout 或 statement_timeout。
**触发:** 配置长期请求额度或成本上限;不能把普通短期 RPM 规则一概算入这条路径。
**结果:** 同一用户多个 key 的准入串行;窗口内历史越多,锁内工作越多。等待锁的事务还占用连接。池 acquire_timeout 只管取得连接之前的等待;Compose 的 idle_in_transaction_session_timeout 也不能终止正在执行的锁等待 SQL。
**建议:** 按用户和额度窗口维护原子聚合及预留,避免准入反复扫描明细。为前台和后台事务分别设定锁等待和 SQL 期限,失败时回滚并限制重试;精确额度和幂等结算约束必须保留。
## 次要放大器
- **P2,过载后重复计算:** `apps/aether-gateway/src/ai_serving/planner/state/scheduler.rs:100` 在 API key 并发受限时短间隔重做候选读取和排序。建议在昂贵规划前做准入,使用有界等待或通知。
- **P2,探测前置任务未合并:** `apps/aether-gateway/src/maintenance/runtime/pool_quota_probe.rs:1675` 先 spawn,再去 Redis 去重及争锁。池恶化时仍随请求量创建任务。建议在 spawn 前按 provider 合并触发信号。
- **P2,同步日志输出:** `crates/aether-runtime/base/src/tracing.rs:403` 直接写 stdout,`:695` 的文件写持全局 Mutex。磁盘或日志收集变慢时可能阻塞 Tokio worker。改为有界日志队列,并明确队列满时策略;当前没有证据证明这是本次故障主因。
## 修改和验收顺序
1. 先修公共套餐锁范围、读取期限及压缩上传预算问题;这些都有明确且局部的触发条件。
2. 从请求调度移走管理扫描,按需读取运行态,避免每请求重算历史统计。
3. 修流空闲期限,建立响应捕获全局预算,将压缩移出事务及异步工作线程。
4. 优化长期额度聚合,并为数据库锁等待、SQL 执行和过载重试设置预算。
5. 用独立测试实例及可控制延迟的 mock upstream 做阶梯压测;每档记录吞吐、P95/P99 首包及总耗时、RSS、CPU、队列积压,停止流量后检查资源能否回落。
生产定位至少需要:故障实例和版本、CPU/内存/FD 配额、实际 RPM 和平均流时长、是否使用压缩上传/账号池/套餐/完整 body 捕获。采集 `pool_runtime_state` 阶段延迟、Redis Admin lane 延迟、数据库 checked-out/lock waiting、usage 队列 lag 和 Tokio 任务数。CPU 低且锁等待高、Redis 延迟高、RSS 持续涨、请求体大量 503 分别对应不同路径,不能仅凭“卡死”选择一个原因。
现有 `crates/aether-testing/loadtools/src/bin/gateway_pressure_probe.rs` 可用于受控环境的压力和排空观测;不要用健康检查接口的吞吐代替真实 AI 调度及结算链路的容量。
@@ -1,42 +0,0 @@
# 格式转换失败诊断导出
## 使用方法
在请求详情的失败或跳过节点中,点击「失败诊断」面板的复制按钮。
复制时才会读取已采集的正文;页面预览本身不会批量加载正文。
请分享整个 JSON,而不是只分享 `summary`。复制成功标志仅在剪贴板写入成功后出现。
## Schema v2
- `diagnostic`:错误码、请求/响应/流式阶段、源/目标格式、转换器标识、完整 JSON 路径及期望约束/实际值。
- `path_source`:`structured` 是后端结构化路径,`message_inference` 是历史文案推断,`protocol_inference` 是根据上游协议推断的原始字段路径,`unavailable` 表示没有可靠字段路径。通用 `$.finish_reason` 会按协议定位到具体原始字段,同时保留 `reported_path`。
- `stage_source`:区分后端阶段信息与历史记录推断。请求转换方向为客户端到上游;响应/流式转换方向相反。
- `versions`:前端版本、导出时网关版本、失败时运行版本。历史记录没有运行版本时保留 `null`,不能把导出版本当作失败版本。
- `request` / `node`:请求、候选、重试、模型及时间等定位信息。
- `reproduction.sources`:脱敏正文片段、字段样本和流式失败事件窗口。数组通配路径的样本带具体下标。
- `reproduction.missing_context`:未采集、无权限、正文过大、读取失败、缺少失败事件或路径等缺口。
后端只为明确匹配 `candidate_id` 的候选提供上游正文记录。历史记录仅有候选索引时,不把最后一次重试的正文猜成当前失败的正文。
原始客户端请求可在同一请求内共享,但不会把其他候选的上游请求/响应当作失败现场。
`body_ref` 不是下载地址;前端不访问其中的 URL,而是使用现有、受权限保护的正文接口。
## 完整性与安全边界
- `not_loaded`:尚未补取上下文。
- `sanitized_context`:必要来源已取得,但仍然经过脱敏、大小限制或事件窗口裁剪。
- `insufficient_context`:还缺少明确列出的信息,不能据此假设能够完整复现。
- `replay_ready: false`:导出的是供排查的证据包,不是可以无条件自动执行的请求。修改前应根据样本建立最小回归测试。
每份正文的下载/解码处理上限为 1 MiB,读取超时为 5 秒;导出 JSON 上限为 64 Ki 字符。
字符串、数组、对象深度与节点数量也有限制。流式窗口保留匹配失败的事件、帧序号及邻近事件;匹配不到时明确标记,而不是认定流尾就是故障点。
默认移除常见认证头、密钥、令牌、Cookie、密码、签名 URL 参数、正文文本和二进制数据。
脱敏是规则化处理,分享前仍需检查自定义字段和错误消息是否含业务敏感信息。
不会为了诊断绕过正文采集策略、授权或存储限制;也不会把 `error` 或未知结束原因映射成正常成功。
## 建议处理流程
1. 检查 `diagnostic.stage`、`path_source` 和 `missing_context`,区分转换器缺陷、合法的无损转换拒绝和上游失败。
2. 对照源/目标格式以及字段样本,建立最小失败输入;流式问题同时保留必要的前序事件。
3. 先补失败回归测试,再修复转换规则。
4. 验证原有正常映射、失败闭合以及凭据脱敏没有回退。
@@ -1,171 +0,0 @@
# DNS 与出站连接审计(2026-09-08)
## 范围与结论边界
本轮基于当前工作区检查 `apps/`、`crates/` 中的 DNS 查询、IP 地址校验、
客户端构建、代理配置以及 TCP/WebSocket 连接入口,沿调用关系区分供应商请求、
身份认证、任意 URL 下载和隧道转发。不是仅搜索 `WebSocket` 或 `chatgpt.com`。
第一轮 WS 修复不足以说明所有路径已经一致。本轮又发现连接测试的旧地址过滤、
URL 形式 IPv6 误入 DNS、两处 DNS 答案静默截断,以及附件下载只保留首个地址。
这些问题,以及后续复核发现的 SMTP 无界解析和隧道缺少显式远程 DNS 模式,
已在工作区修复。本文不表示生产服务器已经部署,也不保证真实上游的
DNS、TLS、TUN 路由或出口代理一定可用。
## 已修复问题
### 1. 普通供应商出口的策略重复
- 普通 HTTP/SSE、浏览器指纹 HTTP、H2C 已使用供应商解析策略;普通 WS 仍有独立过滤,
已在第一轮改为使用 `ExecutionSafeDnsResolver`。
- `/v1/test-connection` 的本地快捷路径仍逐项拒绝私网/保留 DNS 答案,导致同一个
配置好的供应商正式请求能成功、连接测试却失败。本轮删除该重复策略,复用正式
HTTP 客户端的解析器。
- HTTP、WS 和连接测试统一使用 `validate_execution_upstream_url` 校验供应商 URL。
域名的 DNS 答案不按地址段过滤,不等于允许在 URL 中直接填写任意私网 IP。
- URL 中的凭据、fragment、私网/保留 IP 字面地址仍被拒绝;HTTP/WS 字面 loopback
保持正式执行路径已有的兼容策略。禁用重定向、供应商显式代理和 WS 自循环检查保留。
相关文件:
- `apps/aether-gateway/src/execution_runtime/transport.rs`
- `apps/aether-gateway/src/handlers/proxy/websocket/transport.rs`
- `apps/aether-gateway/src/handlers/public/support/test_connection/route.rs`
### 2. IPv6 字面地址被当作域名
`Url::host_str()` 可提供 `[::1]` 形式的主机名,而 `IpAddr::from_str` 和
`lookup_host((host, port))` 的原有调用没有正确消化这个形式。修复前新增测试实际失败,
错误为 `failed to lookup address information: nodename nor servname provided, or not known`。
- 公共解析器现在先识别 IPv4、裸 IPv6 和合法的方括号 IPv6,直接生成 socket 地址。
- 不接受 `[localhost]`、`[127.0.0.1]` 等伪造的方括号主机名。
- relay 的 loopback 判断也使用同一解析函数,防止解析成功后又误判 `[::1]`。
- 隧道 SOCKS 地址编码复用该函数;IPv6 字面地址不再在远程 DNS 模式下被当作域名发送。
- 私网过滤仍由每个调用方的安全策略决定,公共解析器本身不扩大地址权限。
相关文件:`crates/aether-http/src/dns.rs`、
`apps/aether-tunnel/src/egress_proxy.rs`、网关执行传输模块。
### 3. DNS 答案静默截断
Bark 推送和 ChatGPT-Web 图片解析仍直接调用系统 DNS,然后仅取前 32 个答案。
本轮改用共享的 `lookup_host_with_limits`:保留原超时预算,超过 32 个答案直接报错,
不再静默忽略剩余答案。两条路径的私网校验、地址固定及官方来源 Fake-IP 例外不变。
owner gateway 转发的独立实现已取第 33 个答案并拒绝超限,因此不是同类遗漏。
### 4. Grok 附件下载缺少多地址回退
原实现校验全部 DNS 答案后只固定第一个公网地址;首个地址不可连接时,客户端无法
尝试 DNS 返回的其它公网地址。本轮改为把全部经过校验的地址交给客户端,保留双栈和
多地址回退能力。空答案、Fake-IP、私网及公网/私网混合答案仍整体拒绝。
### 5. SMTP DNS 不受连接超时控制
SMTP 发送和连接探测现在都先用共享异步解析器建立 TCP 连接,再把已连接的 blocking
socket 交给原来的 SMTP/TLS 协议实现,不在阻塞任务中重新解析或连接。
- DNS 最长 10 秒;DNS 与 TCP 尝试共用 30 秒总预算。
- 保留全部不超过 32 个答案,不再静默截断为前 16 个;超限直接报错。
- TCP 地址竞争取首个成功连接并释放其它尝试,避免首个黑洞地址吃完整个预算,
使后续可达地址根本没有机会建连。SMTP/TLS 仅在最终选中的连接上运行。
- 解析失败、空答案、解析超时及总连接超时有明确错误,不把底层 DNS 细节暴露给调用方。
- TLS 继续验证原始 SMTP 主机名;读写超时仍为 30 秒,内网邮件服务器策略不变。
相关文件:`apps/aether-gateway/src/email_delivery.rs`。
### 6. 隧道供应商出口可显式委托代理解析
新增默认关闭的 `upstream_proxy_remote_dns`(CLI `--upstream-proxy-remote-dns`,环境变量
`AETHER_TUNNEL_UPSTREAM_PROXY_REMOTE_DNS`,setup 的 `Proxy Remote DNS` 开关)。
- 默认路径仍本地解析、执行 ACL、固定 IP,`socks5h://` 本身不改变既有安全策略。
- 显式启用后,HTTP CONNECT 或 SOCKS5h 接收原始域名,不再预先查询隧道本机 DNS;
Host、SNI 和证书校验仍保留原域名。
- 必须配置 HTTP/SOCKS5h 代理;无代理或 `socks5://` 会在启动和客户端构建时拒绝。
- 仍执行端口白名单、URL 校验以及 IP 字面地址/`localhost` 限制;字面 IP 保持固定。
- 远程 DNS 和固定 IP 使用不同连接池键;远程模式解析器明确拒绝本地 DNS 回退。
- 代理 DNS、TCP、CONNECT/SOCKS 和 TLS 握手共同受上游连接超时限制。
- 启用时打印安全提示:**域名目标的最终 IP ACL 由受信任代理负责**。普通 CONNECT/
SOCKS5 协议无法让隧道校验代理最终连接的 IP,不能宣称远程解析仍保留本地逐 IP 检查。
配置示例及部署边界见 `apps/aether-tunnel/README.md` 的“上游 HTTP 请求”章节。
该文档中 DNS 缓存和连接超时等环境变量误写的 `_SECS` 后缀也已纠正,避免按文档
设置后实际未被程序读取;CLI/TOML 参数名称不变。
## 必须保留的策略差异
| 路径 | DNS / 代理策略 | 本轮处理 |
| --- | --- | --- |
| 普通供应商 HTTP/SSE、浏览器指纹、H2C、WS、连接测试 | 域名答案不按地址段过滤;显式代理优先;供应商客户端不自动使用系统代理环境变量 | 统一遗留分支 |
| 供应商操作类 OAuth、模型获取 | 经执行计划进入供应商运行时;不能与用户登录的身份 OAuth 混为一谈 | 核对调用关系 |
| 身份 OAuth / 管理端 OAuth 探测 | 独立敏感出口;校验目标、固定地址;部分内置官方来源允许窄范围 Fake-IP | 保留,不全局放开 |
| Grok 用户附件、公共视频 URL | 不可信 URL;保留公网限制和固定地址,不能套用供应商域名策略 | Grok 保留全部安全地址 |
| ChatGPT-Web 图片下载与上传 | 普通 URL 严格过滤;可信存储来源有专门 Fake-IP 例外 | 公共有界解析器 |
| 支付出口 | 独立公网校验;固定 Stripe 来源有专门 Fake-IP 例外 | 保留 |
| 系统更新、外部模型目录、Server Chan、Bark | 各自的可信来源例外;自定义目的地不能自动获得同样权限 | Bark 公共有界解析器 |
| gateway owner / internal relay | 独立私网策略、可信 relay 配置和地址固定 | 保留;修复 IPv6 判断 |
| 隧道承载的供应商 HTTP 流量 | 默认本地端口/IP ACL 和固定 IP;显式远程模式委托受信任代理解析及执行域名 IP ACL | 新增默认关闭的远程 DNS 模式 |
| 隧道到 gateway 的控制连接 | 与隧道供应商出口分开;可配置专门出口代理和 IP family | IPv6 / SOCKS 编码复用公共函数 |
| 独立 Responses WS probe | 独立直连诊断程序,不使用供应商代理配置 | 不应当作生产代理路径的等价验证 |
## 仍需注意的实际限制
1. **Fake-IP 只是地址,不提供路由。** 取消供应商 DNS 地址过滤后,进程所在网络仍必须
能通过对应的 TUN/透明代理处理 Fake-IP;否则会变成 TCP 超时,而不是过滤报错。
2. **隧道代理不等于网关直连代理。** 默认仍先本地解析并通过 ACL;只有显式启用
`upstream_proxy_remote_dns` 才委托代理解析供应商域名。代理端点自身的域名仍需
本地 DNS;本地解析完全不可用时,代理 URL 应使用可达 IP。
3. 隧道默认 `AETHER_TUNNEL_ALLOW_PRIVATE_TARGETS=false` 仍会拒绝 Fake-IP。
该开关是扩大内网访问权限,不是建议普遍启用的 DNS 修复;应优先让隧道主机得到
可路由的真实 DNS 答案,或配置受信任代理的显式远程 DNS 模式。
4. 更新客户端支持自己的代理环境变量和 `NO_PROXY`;不能把这一点推广到供应商请求。
网关供应商 SOCKS 配置会归一化为远程 DNS 语义,但其它明确区分 SOCKS5/SOCKS5h
的独立工具仍遵循各自配置。
5. SMTP 等辅助服务不是供应商解析器的调用方。SMTP 已修复 DNS 超时和答案截断,
但不自动继承供应商出口代理。系统解析器由 Tokio 阻塞池承载,异步超时会停止等待,
不等于操作系统正在执行的 DNS 调用能够被强制终止。
6. 本轮没有使用生产凭据、发送真实模型请求、修改系统 DNS、关闭 TUN 或重启服务器。
生产验证仍需在实际容器/进程的网络命名空间内进行。
## 回归验证
- 公共 DNS:合法/非法 IPv6 主机形式、端口、零超时、答案上限、超限拒绝。
- 供应商 DNS:Fake-IP 保留,HTTP 与 wreq 解析结果一致;relay 私网过滤不变。
- WS:普通与浏览器指纹客户端,经本机 HTTP、SOCKS5、SOCKS5h mock 代理,使用
`provider-dns.invalid` 完成真实 WS upgrade;SOCKS mock 断言收到域名而非本地解析 IP。
这是本地明文 WS 的代理路径测试,不代替生产 WSS 的 TLS/SNI 检查。
- 连接测试:供应商 URL 校验一致,无预解析构建请求,重定向不转发凭据。
- 附件与辅助出口:多地址保留、私网/混合答案拒绝及官方 Fake-IP 例外回归。
- 隧道:IPv6 SOCKS 编码、地址 ACL、缓存策略隔离、默认固定 IP 代理连接;新增远程
DNS 模式的 HTTP CONNECT/SOCKS5h 域名握手、Host/SNI 主机名、禁止本地解析回退、
连接池隔离、代理/TLS 握手超时,以及 CLI/TOML/TUI 配置验证和持久化。
- SMTP:DNS 超时、空答案/错误信息、前 16 个地址不可用时使用第 17 个地址,以及
本机 mock SMTP 探测和邮件投递;不发送真实邮件。多地址测试先复现串行建连超时,
改成有界地址竞争后通过。
最终重新执行结果:**809 项测试通过,0 失败**。
- 网关:594 项,覆盖完整 WS 模块、执行传输、Grok、ChatGPT-Web 图片、连接测试,
DNS/Fake-IP/解析地址校验,以及 SMTP 发送、探测和相关配置回归。
- 隧道:197 项,全量单元/本机集成测试。
- `aether-http`:18 项,包含先失败、后修复通过的方括号 IPv6 回归。
- 单独执行的 13 项 SMTP 回归及前轮测试均为上述集合的子集,不重复计入总数。
- Rust 格式检查与 `git diff --check` 通过。
涉及监听器的测试在获准的沙箱外绑定本机回环端口;最初沙箱内的端口权限失败不作为
功能失败,也未通过跳过测试来规避。所有 Cargo 测试使用 `--offline`,代理上游是
本机 mock,不使用生产凭据。
复现命令:
```bash
cargo test -p aether-gateway --lib --offline -- --quiet \
handlers::proxy::websocket:: execution_runtime::transport::tests \
execution_runtime::grok::tests execution_runtime::chatgpt_web_image::tests \
test_connection bark_push::tests server_chan_push::tests \
dns fake_ip benchmarking resolved_addrs email_delivery smtp
cargo test -p aether-tunnel --offline -- --quiet
cargo test -p aether-http --offline
```
@@ -1,47 +0,0 @@
# 历史权限空值升级修复
## 原因
旧版 API Key、用户、用户组及上游 Key 允许把 JSON 字面量 `null`、字符串 `"null"`(忽略大小写和首尾空白)以及空字符串作为未设置的权限。严格权限读取启用后,这些值会触发 `contains JSON null; use SQL NULL for an unset policy` 等错误,影响 API Key 鉴权、列表和用户、用户组读取。管理令牌的 IP 限制和权限也曾接受 JSON 字面量 `null`,但不接受字符串空值;严格校验同样会使这些旧令牌失效。
增量迁移 `20260908000000_normalize_legacy_policy_nulls.sql` 将这些已知的旧版空值转换成 SQL `NULL`,不修改既有迁移及其校验和。新安装和已有数据库升级均通过同一迁移机制执行。
## 覆盖范围
| 表 | 字段 |
| --- | --- |
| `api_keys` | `allowed_providers`、`allowed_api_formats`、`allowed_models`、`ip_rules` |
| `users` | `allowed_providers`、`allowed_api_formats`、`allowed_models` |
| `user_groups` | `allowed_providers`、`allowed_api_formats`、`allowed_models` |
| `provider_api_keys` | `api_formats`、`allowed_models` |
| `management_tokens` | `allowed_ips`、`permissions`(仅 JSON 字面量 `null`) |
共 5 张表、14 个字段,同时兼容 `json` 和 `jsonb` 列。迁移可重复执行,且只更新命中旧版空值的字段:
- 保留 SQL `NULL`、空数组 `[]`、正常名单、字符串化的名单和单字符串策略。
- 保留 `specific`、`deny_all`、`inherit` 等权限模式,不修改用户组成员关系。
- 管理令牌只转换 JSON 字面量 `null`,恢复旧版既有的未设置 IP 限制或 `legacy_full` 权限语义;字符串 `"null"`、空字符串、空数组和其他非法权限不转换,避免把原本无效的令牌权限变成旧版全权限。
- 不清理数组内部的 `null`、空白元素、数字、对象或其他异常权限;它们仍由严格读取逻辑拒绝,不能借迁移变成无限制访问。
- 不修改其他 JSON 字段,例如 `metadata` 内的 JSON `null`。
## 升级方式
先备份数据库,部署包含此迁移的新网关二进制或镜像,并保留原来的数据库连接配置与密钥。仅重启不包含此迁移的旧版本不会修复数据。
默认 `AETHER_GATEWAY_DATABASE_MODE=auto` 会在启动时执行挂起迁移,再进入正常服务。使用 `verify-only` 的部署需在相同数据库连接配置下先执行新版本的准备命令,再重启服务:
```sh
aether-gateway db prepare
```
迁移只更新上述权限列,不删除业务记录、不替换用户或 Key、不重建数据库。不要把空数组或任意异常 JSON 统一改成 SQL `NULL`,也不要通过关闭严格权限校验绕过问题。
升级后可以只读检查迁移记录:
```sql
SELECT version, description, success
FROM public._sqlx_migrations
WHERE version = 20260908000000;
```
该记录应存在且 `success = true`。若仍有权限解码错误,核对实际报错实例连接的数据库,以及是否有旧进程或外部工具继续写入旧格式;不要将其他类型的权限错误直接作为空值清除。
@@ -1,61 +0,0 @@
# OpenAI Responses WebSocket probe
`aether-openai-responses-ws-probe` verifies the official OpenAI Responses
WebSocket protocol using standard API-key Bearer authentication. It sends two
sequential `response.create` warmups on one socket, chaining the second from
the first response ID with `previous_response_id`.
It shares its protocol-driving core with the Codex probe, but it does **not**
send Codex account headers or require Codex quota events. This makes it the
compatibility gate for Aether's standard Responses WebSocket adapter, rather
than a replacement for the Codex probe.
## Prerequisites
Use a dedicated API project and a model that your key can access. Keep values
only in your process environment or secret manager:
```bash
export AETHER_OPENAI_WS_PROBE_API_KEY='your-api-key'
export AETHER_OPENAI_WS_PROBE_MODEL='your-openai-model'
```
The default endpoint is the official Responses WebSocket endpoint:
```text
wss://api.openai.com/v1/responses
```
To test a compatible endpoint explicitly, set
`AETHER_OPENAI_WS_PROBE_URL` or pass `--url`. The endpoint must use `ws://` or
`wss://` and may not contain credentials, a query string, or a fragment. The
API key has no command-line flag and is never printed.
## Run
```bash
cargo run -p aether-gateway --bin aether-openai-responses-ws-probe
```
For an explicit endpoint and timeout:
```bash
cargo run -p aether-gateway --bin aether-openai-responses-ws-probe -- \
--url 'wss://api.openai.com/v1/responses' \
--timeout-secs 30
```
The probe uses `generate:false`, so the warmups prepare continuation state but
do not request model output. A successful JSON report contains
`"continuation_confirmed":true`; header and event arrays contain names only,
never credentials, response IDs, request bodies, or response bodies.
## Interpretation
Success establishes that this key, model, and endpoint support the Responses
WebSocket handshake plus an in-socket continuation. It does not establish
support for every model, tool, service tier, proxy path, or Aether provider
configuration. Treat a successful direct probe as a prerequisite before
enabling **Responses WebSocket mode** for the matching Aether provider.
For protocol details, see the official [WebSocket Mode guide](https://developers.openai.com/api/docs/guides/websocket-mode).
-150
View File
@@ -1,150 +0,0 @@
# Runtime Redis Operations Runbook
This runbook covers Aether runtime Redis connection pressure incidents. It is
not a substitute for fixing application-level connection churn.
## Persistence Policy
The bundled `docker-compose.yml` treats Redis as a low-latency runtime
coordination layer by default: locks, cache affinity, semaphores, and runtime
streams. Postgres remains the source of truth. The default Redis persistence
policy is passed directly to `redis-server` in `docker-compose.yml`:
```sh
--dir /tmp --appendonly no --save ""
```
This avoids request-path latency spikes from AOF fsync and background snapshot
forks. The trade-off is that Redis runtime state can be lost if the Redis
container or host crashes before workers have flushed queued records to the
database. The default `dir /tmp` also prevents old files in the mounted
data directory from being loaded as stale runtime state; the persistence
disable itself is `--appendonly no` and `--save ""`.
Only deployments that intentionally want Redis runtime streams to survive a
crash should restore persistence in the Redis command:
```sh
--dir /data --appendonly yes --appendfsync everysec --save 60 1000
```
Expect higher tail latency when Redis persistence shares disks with Postgres or
application logs.
### OpenAI Responses continuation history
When an OpenAI Responses request is converted to an OpenAI Chat provider,
Aether stores the completed continuation transcript in `RuntimeState` under the
`ai:responses:history:v1` namespace. Records are immutable, scoped by a hashed
API key identity, limited to 8 MiB, and expire after six hours. Redis `SET` with
TTL makes completion writes atomic and idempotent.
All gateway instances must use the same Redis URL and key prefix. This allows a
continuation request to land on another instance and allows gateway processes
to restart without losing history. `AETHER_RUNTIME_BACKEND=memory` remains a
single-process development mode and cannot provide either guarantee; multi-node
startup rejects it.
The bundled non-persistent Redis policy survives gateway restarts but not a
Redis container or host restart. Deployments that require continuation history
to survive Redis restarts must enable the AOF/RDB policy above and mount `/data`,
or use an externally managed persistent Redis service. Monitor
`openai_response_history_read_failed`, `openai_response_history_write_failed`,
and `openai_response_history_invalid` events for backend or payload failures.
## Latency Triage
Redis `INFO commandstats` reports `latency_percentiles_usec_*` values in
microseconds. For example `p99=2007` means about 2 ms, not 2 seconds.
Use these checks before attributing app stalls to Redis:
```sh
redis-cli -p 6379 -a "$REDIS_PASSWORD" LATENCY DOCTOR
redis-cli -p 6379 -a "$REDIS_PASSWORD" LATENCY LATEST
redis-cli -p 6379 -a "$REDIS_PASSWORD" SLOWLOG GET 20
redis-cli -p 6379 -a "$REDIS_PASSWORD" INFO persistence
redis-cli -p 6379 -a "$REDIS_PASSWORD" INFO commandstats
redis-cli -p 6379 -a "$REDIS_PASSWORD" INFO clients
```
For immediate mitigation on an existing container that is running with AOF
enabled:
```sh
redis-cli -p 6379 -a "$REDIS_PASSWORD" CONFIG SET appendfsync no
redis-cli -p 6379 -a "$REDIS_PASSWORD" CONFIG SET appendonly no
redis-cli -p 6379 -a "$REDIS_PASSWORD" CONFIG SET save ""
```
Active defrag can help when `mem_fragmentation_ratio` is high, but it is not an
AOF fsync fix. Enable it only after confirming the Redis build supports it:
```sh
redis-cli -p 6379 -a "$REDIS_PASSWORD" CONFIG SET activedefrag yes
redis-cli -p 6379 -a "$REDIS_PASSWORD" CONFIG SET active-defrag-ignore-bytes 50mb
redis-cli -p 6379 -a "$REDIS_PASSWORD" CONFIG SET active-defrag-threshold-lower 10
```
## Normal Expectations
- Each `RuntimeState` Redis backend initializes a fixed set of long-lived
connection lanes: fast, stream, blocking stream, and admin.
- `connected_clients` should stay near a small fixed number per app instance,
plus health checks and ad hoc admin clients.
- `total_connections_received` should not grow linearly with request volume.
- Large TIME_WAIT spikes between app and Redis indicate a regression or a
separate process repeatedly opening Redis connections.
## Emergency Mitigation
1. Disable the retry source first, such as expired Codex/OAuth keys causing a
retry storm.
2. Restart the app to stop continued connection creation:
```sh
docker compose restart app
```
3. On a Linux host, temporarily widen the ephemeral port range and enable safe
TIME_WAIT reuse:
```sh
sudo sysctl -w net.ipv4.ip_local_port_range="10000 65535"
sudo sysctl -w net.ipv4.tcp_tw_reuse=1
```
4. Do not enable `tcp_tw_recycle`; it is obsolete and unsafe with NAT.
Docker Desktop on macOS runs containers inside a Linux VM. Host-level macOS
`sysctl` changes do not necessarily affect the VM network namespace.
## Checks
Use Redis `INFO clients` and `INFO stats` to inspect:
- `connected_clients`
- `total_connections_received`
Use OS socket tooling on the Redis host or container namespace to inspect
TIME_WAIT counts. Persistent growth after the runtime Redis refactor means a
different code path or process is still opening short-lived Redis connections.
## File Descriptor Limits
Aether's compose files intentionally do not set container `ulimits.nofile`.
Redis connection churn must be fixed in application code, not hidden by larger
file descriptor limits.
For high-concurrency production hosts, set file descriptor policy at the
runtime or service-manager layer instead:
- Docker daemon default ulimit, for example `default-ulimits` in
`/etc/docker/daemon.json`.
- systemd service limits such as `LimitNOFILE=` for Docker or the process
supervisor.
- Managed container platform resource settings, when Docker daemon settings are
not available.
Keep Redis `maxclients` below the effective Redis process `nofile` limit with
room for persistence files, replicas, and admin connections.
-57
View File
@@ -1,57 +0,0 @@
# 调度策略级故障转移
调度策略的 `default_policy` 支持跨提供商的转移预算与错误规则。它们跟随当前请求的 `routing_execution_policy` 快照进入执行器,不依赖运行中修改全局系统设置。已有策略缺省为不限次数、不限累计时间、无全局错误规则。
```json
{
"default_policy": {
"sticky_key_attempts": 2,
"max_transfer_count": 3,
"max_transfer_timeout_seconds": 90,
"failover_rules": {
"success_failover_patterns": [
{ "pattern": "(?i)capacity.*exhausted" }
],
"error_stop_patterns": [
{ "status_codes": [400, 413], "pattern": "invalid.*parameter" },
{ "status_codes": [422] }
]
}
}
}
```
## 预算语义
- `sticky_key_attempts` 是首个粘性候选上的总尝试次数,`2` 表示首次请求加一次同 Key 重试。该行为保持不变。
- 全局 `max_transfer_count` 统计切换候选的次数,首次尝试不计数。同一提供商、端点、Key 上的重试不计数;改变该组合计一次。`3` 最多允许首次候选之后再切换三次。
- 全局 `max_transfer_timeout_seconds` 从首次候选开始执行时计时,覆盖后续重试与切换间的累计耗时。它在准备下一次尝试时检查,不会强制打断已经执行中的调用或已提交给客户端的流;单次连接、首字节、读取和非流式完整调用超时仍独立生效。
- 两个全局预算的 `0` 都表示不限制。被筛除、禁用或被提供商级预算跳过而未执行的候选不计数。
- 提供商自身的转移次数和时间预算继续生效。提供商预算耗尽只跳过该提供商,仍可尝试其他提供商;全局预算耗尽则不再执行任何提供商的新尝试。
- 预算仅约束当前请求,不能重置或替代客户端取消策略、权限校验及本地执行异常的终止行为。
## 错误规则
先匹配调度策略的全局显式规则;未匹配时继续使用提供商的规则及协议默认行为。本地执行函数真正返回 `Err` 时仍然终止,不做兜底重放。
- **成功转移规则**:仅当 HTTP 200 响应匹配配置的正则时继续转移,不是对所有 200 进行重试。非流式请求匹配响应体;流式请求只匹配尚未交付业务输出的有界预读取内容。
- **结构化错误优先**:标准流式请求统一在预读取阶段先解析完整 SSE 事件或 JSON 中的错误,再判断成功正则;不在半截错误载荷上提前触发成功转移,避免旧的 JSON 提前探测绕过错误终止规则。普通文本响应仍支持跨分片匹配。
- **图片成功保护**:`openai:image` 的成功响应保留不重放行为,不因全局或提供商的成功正则再次生成图片;正常错误响应仍按错误规则处理。
- **错误终止规则**:适用于 400–599 错误。状态码与正则都填写时要求同时满足;只填状态码表示该状态一律终止;只填正则表示在所有错误状态上匹配。流内错误使用解析后的错误状态,而不是外层 200。
- **网络错误**:没有上游 HTTP 状态的连接、TLS、DNS、提交前超时等错误统一继续转移;提供商级别若单独配置了停止规则,仍按提供商规则处理。
- 正则使用 Rust `regex` 语法,支持 `(?i)` 等内联标志。服务端拒绝无效正则、无意义的空规则以及错误状态范围。每组最多 64 条,每条表达式最多 4096 字节。
## 流式 200 的恢复窗口
上游 HTTP 200 响应头不再默认关闭标准文本 SSE 的恢复窗口。执行器先缓冲协议开场事件,例如 Responses 的 `response.created`、Chat 的 role-only 增量、Anthropic 的空 `message_start` / 文本块起始事件。`openai:image` 图片专用流保留原来的响应头提交行为,成功正则不打开重放窗口。
首个业务内容之前的结构化错误、过早 EOF、首字节超时以及 200 正则命中会进入统一故障转移判断。真实文本、思考、工具调用或正常结束事件确定后,缓冲内容按原顺序交付;之后发生的错误保持终止,不重新执行原请求。
预读取受单次首字节时间和既有字节上限约束。达到字节上限时保守提交,避免无限缓冲;这意味着不能承诺识别响应任意位置的错误或正则。原始 HTTP 状态与最终执行结果是不同观测值,不能因流内失败而伪改已经发送的 HTTP 状态码。
## 排查
- `routing_transfer_limit_reached`:当前调度策略的累计次数或时间预算耗尽。
- `provider_transfer_limit_reached`:提供商自身的预算耗尽。
- `local_stream_candidate_retry_scheduled`:输出前的流内错误或 200 规则触发了继续调度。
- `local_stream_transport_retry_scheduled`:输出前传输错误触发了继续调度。
-28
View File
@@ -1,28 +0,0 @@
# 按策略配置模型调度
管理端「调度策略 → 调度配置」按配置组织模型,不再逐个模型编辑和保存。
1. 先选择调度范围「全部模型」或「区分模型」。「全部模型」只显示一套调度设置,不显示模型选择和「添加配置」,保存后自动包含以后新增的模型。
2. 「区分模型」下,在表单内展开适用模型列表并勾选模型,再设置调度优先级(Provider / Key)和调度策略(缓存亲和、负载均衡、固定顺序)。搜索只过滤模型列表,其他配置已占用的模型默认隐藏;通过取消勾选移除模型。选择框显示已选名称或数量,不重复显示已选标签。
3. 设置此配置共用的提供商 / Key 排序。实际请求仍只使用该模型可用的候选,不会因为共用排序而启用不支持该模型的提供商。
4. 「区分模型」下,点击调度配置标题旁的「添加配置」,为剩余模型选择不同策略。收起的配置直接显示适用模型名称,已分配给其他指定配置的模型不能重复选择。
5. 点击页面顶部「保存」统一生效,不需要逐个模型另存草稿。指定范围为空时不能保存。
## 范围语义
- 「全部模型」是独立的动态范围,不是勾选当前列表的快捷操作,保存后也自动适用于之后新增的模型。「区分模型」的列表不提供「全部模型」选项。
- 「全选当前」与「全选结果」只批量勾选当前可选模型,仍属于指定模型范围,不自动包含以后新增的模型。
- 编辑期间切换范围会分别保留两种模式的草稿,切回后恢复原有选择和设置。首次切换时继承当前配置的调度和排序;从多配置首次切到全部模型时,优先继承默认配置,否则使用第一条配置。保存只写入当前模式,全部模型模式会清除模型级的自动生成规则和独立排序。
- 旧数据中的指定模型与全部模型混合配置按「区分模型」加载,原全部模型条目显示为「默认配置」。默认排序仍可被各模型继承并覆盖,且不妨碍为剩余模型添加配置。
- 没有全部模型配置时,未指定的模型继续使用已保存的默认调度设置,不会被禁用。先修改配置再改为指定范围,不会把修改后的调度模式应用到未选择的模型。
- 移除模型或删除配置会同步移除对应的排序覆盖和自动生成的调度规则。
- 故障转移、首个候选重试次数和客户端断开处理等仍作用于整个调度策略,不随模型范围拆分。
## 存储兼容
无需数据库迁移,继续使用 `default_policy`、`model_policies` 和 `rules`:
- 全部模型的调度模式写入 `default_policy`,排序写入 `model: "*"` 的模型策略。
- 指定范围的共用排序展开为各模型的 `model_policies`;同一配置使用一条 `ui_scheduling_policy:` 前缀的调度规则,通过 `conditions.any` 匹配适用模型,并使用 `set_scheduling` 设置调度模式。
- 旧的逐模型排序和 `ui_model_scheduling:` 规则仍可读取;编辑调度配置时转换为新结构。等价的旧模型配置可合并显示,已有自定义规则和全局故障转移设置保留。
- 新建的不同配置即使调度设置相同,也保留独立的配置范围,重新打开页面后仍可分别编辑。
File diff suppressed because it is too large Load Diff
@@ -1,98 +0,0 @@
# 安全加固兼容性复核(2026-09-07)
后续扩展到全仓的检查范围、新发现、逐文件清单及限制见 `security-hardening-full-audit.md` 和 `security-hardening-audit-manifest.tsv`。
## 范围
本轮针对 `579f2c7cc`(2026-09-04)及其后续修复进行复核。该提交涉及 1019 个文件,混合了安全边界、订阅计费、数据持久化和运行时变更,不适合整体 revert。
本轮重点检查管理端脱敏与编辑往返、Provider OAuth 默认配置、已有 DNS/代理兼容性修复、管理员错误诊断、SMTP/LDAP/OAuth 配置、导入导出及请求大小限制。以下是已确认的问题及处理,不代表对全部文件逐行审计或对生产环境的完整验证。
## 新增:管理员按需查看规则原值
- 端点管理的请求/响应规则工具栏增加“查看原值”。只读弹窗展示服务端已保存的请求头、请求体和响应头规则,不覆盖未保存的草稿。
- 新接口:`GET /api/admin/endpoints/{endpoint_id}/rules/reveal`。
- 仅管理员可访问;管理令牌需要 `admin:endpoints_manage:admin`,只读和普通写权限不足。
- 返回 `Cache-Control: no-store`、`Pragma: no-cache`,记录 `admin_endpoint_rules_revealed` 审计事件。响应仅包含三组规则,不附带端点的其他配置或凭据。
- 关闭弹窗、切换端点或卸载组件时取消请求并清除明文;旧请求的迟到结果不会重新显示。
## 本轮修复的六类误伤
| 问题 | 影响 | 处理 |
| --- | --- | --- |
| 普通协议头也全部脱敏 | `Content-Type`、`User-Agent`、版本/客户端信息及相应条件无法正常查看;配置导出也受影响 | 恢复明确的常用非凭据头的显示;认证头、Cookie 和未知自定义头继续默认隐藏。旧版本返回的保留标记仍可正常保存 |
| 重复规则无法恢复原值 | 同一头或路径配置多条条件规则后,保存或重排可能将原值变成 `***` | 在相同操作和目标范围内按可区分的规则结构匹配;未修改的重复规则保留原顺序与原值。无法唯一定位的脱敏编辑明确报错,禁止静默写入占位符 |
| 嵌套条件组丢失原值 | 多层 `all`/`any` 在修改其他条件或重排后,内部脱敏值无法恢复 | 增加递归条件组匹配、未修改条件组往返保留和保留值校验 |
| URL 类型判断过宽 | `image_url` 等对象/数组被变成 `null`,内嵌 `data:` 图片 URL 也受影响,继续编辑可能破坏请求规则 | 保留结构化 URL 和内嵌数据;递归处理网络 URL,继续隐藏并在保存时恢复网络凭据 |
| 保存后的编辑状态没有同步 | 后台返回脱敏数据后,界面仍保留提交前明文,持续显示“未保存” | 使用保存响应重新建立编辑基准;保存期间的新编辑不被覆盖,关闭后清除草稿,忽略跨弹窗的旧响应 |
| Gemini CLI 默认 OAuth 被误移除 | 未额外配置环境变量时,原有默认客户端不能授权或刷新 | 恢复加固前内置 native-app 客户端的配套默认值;显式配置优先,自定义客户端 ID 仍必须提供自己的凭据,调试输出继续脱敏 |
明确可直接显示的头包括 `Accept`/编码/语言、`Content-Type`/编码、`Cache-Control`、`User-Agent`、Anthropic/OpenAI 协议版本与 beta 标记,以及明确列出的 `x-stainless-*` 客户端运行时字段;不是按前缀放行任意自定义头。
## 补充修复:显式 `full` 请求记录被禁用
原则:安全处理应约束未授权访问、未配置时的默认行为和真实凭据泄露,不应静默覆盖管理员明确启用的功能。
本次确认不是单纯的前端隐藏,而是同一链路上的多重清空:
- `request_record_level=full` 在运行时被强制解析为 `basic`。
- 独立 HTTP 审计存储的输入投影清除了所有请求头、正文、正文引用和采集状态。
- 内存仓库每次更新都丢弃正文与引用;PostgreSQL 的单条写入把正文 blob 同步改成删除,并拒绝 HTTP 审计内容。
处理:恢复显式 `full`(含旧配置名 `request_log_level`)的采集与独立存储,保留当前配置名优先;恢复请求、上游请求、上游响应和客户端响应各方向已有的采集内容,流转状态更新不再无条件删除它们。PostgreSQL 正文仍写入原有压缩 blob 表,HTTP 头与引用仍写入独立审计表,不重新塞回计费主表。
保留:缺失、无效配置和读取配置失败时默认 `basic`;`basic` 不记录正文;OAuth 令牌交换等内部凭据请求不采集正文;认证头及未知自定义头继续脱敏;正文引用仍校验所属请求和字段;管理端权限、审计及无缓存策略不变。普通同格式非流式响应仍只保存一份原有响应体,前端沿用回退显示,不制造重复的客户端副本。
新增回归覆盖配置解析、内存生命周期、流式/非流式管理端正文读取、旧正文大小配置不覆盖 `full`、PostgreSQL 单条/批量写入与正文读回、前端按需加载与无正文时的展示。真实数据库测试使用本机临时隔离 PostgreSQL,不访问现有业务数据库。
此修复只能恢复之后新采集的记录;此前未采集或已删除的正文无法由代码补回。运行时沿用原有的 30 秒采集策略缓存,刚切换记录级别时需等待缓存刷新。
扩大执行原有、默认忽略的 PostgreSQL 测试时,另发现 5 个旧用例的测试数据/类型与当前 schema 不兼容:4 个使用超过 `varchar(36)` 限制的 Provider Key ID,1 个直接按 `f64` 解码 `NUMERIC` 列。它们并非本次正文修复引入,也不作为本次已通过项;未修改相关业务逻辑或测试夹具。
## 继续复核:视频任务链路的四类回归
继续沿“业务字段被当作诊断数据清空”“读取投影与更新条件不一致”检查,另确认以下四类问题,均可定位到本次加固新增的清空或校验条件,不是用户配置错误:
| 问题 | 实际影响 | 本轮修复 |
| --- | --- | --- |
| 任务业务信息被清空 | 提示词、用户名和客户端 Key 名称丢失,管理页显示空白或 `Unknown`;提示词被写成 `NULL` 还与 PostgreSQL 的非空约束冲突,可能直接阻止任务入库 | 保留这些明确用于任务展示的业务字段,不再作为敏感诊断一律删除 |
| 下载地址被清空或破坏签名 | OpenAI 任务的 `video_url` 被无条件丢弃;Gemini 地址仅保留 `alt=media`,使下载签名、有效期等参数失效 | 保留 HTTP(S) 产物地址及完整查询串,不重排重复参数、不重编码签名;继续拒绝非法协议和 URL userinfo,移除 fragment |
| Gemini 完成结果和重载丢失 | 完成轮询直接清空元数据;数据库重建与加密文件加载再次清空结果 URI,出现任务已完成但无法取得视频 | 仅保留下载所需的最小结果 URI,不保留上游完整响应;从数据库的业务字段重建必要的提示词、尺寸信息和结果 URI,保证再次保存不会丢失 |
| PostgreSQL 轮询更新条件自相矛盾 | 领取任务时不读出时长、分辨率、宽高比和尺寸,而加固后的更新要求这些不可变字段逐项相等,导致正常轮询更新被拒绝、任务停留在处理中 | 领取时读回必要业务字段,保留原有归属/身份校验和并发领取锁,不靠放宽更新条件绕过问题 |
保留原有播放/下载接口与鉴权语义:签名产物链接是有权限用户访问任务结果所需的业务数据,不等同于可随意删除的调试凭据。OpenAI 直接下载不附带 Provider 认证头;Gemini 仍通过网关文件接口访问,并校验上游地址同源后才附加 Provider Key。私网/保留地址拦截、跨域凭据隔离、任务归属校验及管理操作审计不撤销。
任务数据库仍不保存完整请求体、原始错误消息或含认证头的传输快照;本地文件仍使用原有认证加密格式,调试输出继续脱敏。回归覆盖任务保存后重读、签名参数顺序与编码、加密文件重载、管理列表/详情、带审计的下载,以及轮询前后提示词和尺寸不丢失。
真实 PostgreSQL 回归使用临时 Unix socket 实例,并从正式迁移后的 schema 复制会话隔离的任务表。OpenAI/Gemini 两种任务均验证了入库、领取、正常完成更新、拒绝不匹配的尺寸更新和重新查询签名地址;不连接现有业务数据库。
这四类是本轮已经确认的新增回归,不代表对混合提交全部 1019 个文件给出“零遗漏”保证。原有用户策略字段停用早于本次加固,不据此回退;后台任务和缓存中的裁剪未发现足以确认本次功能回归的调用链,不做猜测性恢复。数据库仍保存的历史值可以重新读取;已经写空、未入库或已过期的资源不能凭本次源码修复重建。
## 已存在的后续修复:保留,不重复撤销
- `522b97905`:已移除普通 Provider 代理的 DNS 地址过滤和相关白名单设置。
- `696273122`:已恢复 Antigravity 默认 native-app OAuth 配套凭据;本轮将同类遗漏补到 Gemini CLI。
- `a26680f46`:已恢复旧 SMTP 密码迁移。
- `062e111c0`:已恢复管理员查看上游错误诊断的能力,同时保留普通用户侧脱敏。
- `14f96c9fa` 等:已修复脱敏后的密钥健康状态摘要,不将其退回旧实现。
## 保留的安全边界
- 管理员/普通用户与管理令牌的权限隔离;凭据查看接口的审计和无缓存要求。
- 凭据加密存储及凭据与目标/身份绑定;未知头和真实敏感字段的默认脱敏。
- 登录 OAuth、支付、隧道中继等独立网络边界,TLS 校验和隧道防重放。
- 有界请求/响应缓冲、备份恢复模式、导入校验、计费与配额一致性;本轮没有因“安全加固”标签撤回这些功能。
## 回归覆盖与使用限制
回归用例覆盖规则投影/恢复、重复与嵌套规则、占位符拒绝、结构化 URL、保存期间的并发编辑、弹窗关闭/切换时的迟到响应、接口权限/审计/无缓存,以及 Gemini CLI 和 Antigravity 的默认配置与显式覆盖。
早期兼容性修复的阶段性验证记录(最终全量结果见 `security-hardening-full-audit.md`):
- 视频数据契约 12 项、视频核心 33 项(含加密文件重载)、内存视频仓库 11 项、PostgreSQL 视频仓库 9 项、Provider 视频传输 9 项均通过。
- 新增的真实 PostgreSQL 回归单独执行通过,同时覆盖 OpenAI 和 Gemini;临时数据库实例已停止并清理。
- 网关 `video` 101 项、`async_task::` 20 项、`usage` 302 项、端点规则原值查看 3 项、管理令牌权限 28 项均通过;筛选结果可能重叠,不合计为独立用例总数。
- 前端 199 个测试文件、1411 项测试通过;补齐保存测试的类型夹具后,该文件 4 项测试及 ESLint 再次通过;Rust 格式检查和 `git diff --check` 通过。
- 阶段性检查曾发现 461 条既有类型诊断,与未修改 HEAD 的同依赖基线一致。2026-09-07 的“全部修复”续轮已修正这些诊断,并将 `npm run type-check` 接入 `vue-tsc -b --force --pretty false`;当前真实全量类型检查为 0 错误。最新全量测试、构建及剩余验证边界以 `security-hardening-full-audit.md` 为准,未关闭严格模式或排除测试。
上述修复与回归不涉及生产服务器部署或现有业务数据库更改;源码提交、推送不代表生产环境已经部署或完成验证。若历史版本已经把真实值覆盖为字面量 `***`,查看接口不能重建丢失的原值,需重新填写或从可靠备份恢复。
@@ -1,134 +0,0 @@
# 安全加固全仓兼容性复核(2026-09-07)
## 范围与方法
- 加固基线:`579f2c7cc`(2026-09-04);复核时 HEAD:`a5c3699ae`。保留已有未提交的兼容性修复,不整体回退该混合提交,也不撤销后续修复。
- 历史提交涉及 1019 个文件,其中 907 个路径仍存在,112 个已在后续重构中移除。逐路径清单见 `security-hardening-audit-manifest.tsv`;已移除的旧数据库适配器、未发布迁移等不凭历史清单重新引入。
- 对复核时 Git 跟踪的 2843 个 Rust、TypeScript、Vue、JavaScript、Python、SQL 和 Shell 源文件进行清单、读取与模式扫描,并检查加固差异、后续修订及关键生产调用链。正文清空、无条件禁用、脱敏、URL 拒绝、权限与数据归属等是重点搜索模式。
- 清单中的 `heuristic_added_risk_lines` 是启发式候选行数,含初始化代码和内联测试,**不是 Bug 数、漏洞数或逐行人工审计完成标记**。全仓扫描、模块复核、回归测试是不同的覆盖层次,不等于人工精读所有源文件,也不承诺零遗漏。
- 判断标准:恢复管理员明确启用的功能和必要业务数据;保留未授权访问阻断、默认保护、凭据隔离和有界资源使用。不因一条旧测试要求“所有内容都为空”,就把正常功能重新禁用。
之前的规则、OAuth、`full` 记录及视频四类修复详见 `security-hardening-compatibility-review.md`。本报告只把本次扩展检查的新发现单独计数。
## 本次新增确认并修复的两类加固回归
### 1. 正文的“压缩保留”被替换成删除
涉及 `crates/aether-data/adapters/postgres/src/usage/cleanup.rs`。
原本分为详细正文、压缩正文、请求头、计费记录等独立保留周期。加固后,详细正文清理不再压缩迁移,而是删除正文、独立 blob 与审计引用;选择条件还包含已经独立存储的正文。因此采用默认 7/30 天设置时,明确启用 `full` 后保存的正文也会在约 7 天后提前失去,而不是按压缩正文的 30 天保留。
修复内容:
- 详细正文到期只处理主表中的旧内联/压缩数据,将四个方向的正文迁入独立 gzip blob,并保留 HTTP 审计引用;不选择已经迁出的正文反复处理。
- 旧 metadata 中的正文引用按当前请求和字段校验后迁入审计表,删除旧 metadata 键,但不删除实际正文。无效、跨请求及跨字段引用不恢复。
- 每条迁移在事务中锁定并重读来源行。正文写入或引用更新失败时回滚,不先清空原值;其他 metadata 内容保留。
- 预览统计与实际清理选择条件一致。压缩正文到期仍删除;明确选择“立即清理正文”的操作仍执行删除,未把隐私清理功能改成永远保留。
真实 PostgreSQL 回归使用正式迁移后的 schema 与会话隔离的临时表,覆盖四个方向的内联 JSON、旧 gzip、独立 blob、旧引用迁移、外部请求引用拒绝、7/30 天区间、过期删除、重复清理幂等、显式即时删除和中途写入失败后的事务回滚。不会读取或修改业务数据库。
同时修复了该迁移路径原有的 SQL 错误:`usage` 表没有 `updated_at` 列,旧更新语句却写入它,真实 schema 下会导致迁移失败。这个错误在加固前已经存在,**不混算成第三类加固新增回归**。
### 2. 合法支付会话链接的 fragment 被误拒绝
涉及 `frontend/src/utils/paymentUrl.ts`,调用方包括钱包充值和订阅购买。
加固把包含 `#fragment` 的 HTTPS 支付链接一律拒绝。支付会话链接可能依赖该片段,不能将其等同于脚本协议或 URL 内嵌凭据。现在保留原有片段、查询参数及编码,同时继续拒绝非 HTTPS、相对地址、反斜杠和 URL 用户名/密码。
对应工具函数新增会话片段和编码保留测试,危险协议及凭据注入用例继续执行。服务器端支付 API 基址的同源、TLS 和无片段限制不变;没有放宽 webhook 验签、支付金额、订单归属或实际支付状态校验。此次不进行真实扣款测试。
## 按模块的复核边界
| 模块 | 复核内容与处理 |
| --- | --- |
| 公共入口、认证、会话、管理令牌 | 核对路由分类、认证缓存刷新、跨节点撤权和令牌权限目录;保留身份头防伪、用户/管理员隔离及敏感动作的权限要求。 |
| 端点配置及编辑往返 | 保留前序规则原值查看、重复/嵌套规则恢复、占位符拒绝和并发编辑保护,避免只恢复显示而破坏再次保存。 |
| 普通 Provider 连接与代理 | 对照后续 DNS/FakeIP/SOCKS 修复,不重新引入已移除的普通 Provider 地址过滤;凭据专用连接与普通业务代理分开判断。 |
| OAuth、SMTP、LDAP 与密钥 | 检查默认客户端、显式覆盖、凭据迁移与目标绑定;保留 Gemini CLI/Antigravity 兼容性恢复及既有 SMTP 修复,不撤销传输凭据隔离。 |
| 请求记录、候选诊断及保留策略 | 核对配置读取、采集、投影、单条/批量写入、读取和清理链路;修复本报告的提前删除,保留显式 `full` 及必要诊断,不把原始正文重新塞回计费主表。 |
| 视频与异步任务 | 保留前序提示词/展示名、签名产物 URL、Gemini 结果 URI、数据库领取字段修复;身份不可变条件、任务归属、加密文件和下载凭据隔离不撤销。 |
| 模型、调度、配额和池状态 | 对照模型获取、手动模型、健康摘要、候选选择、额度预留及状态流转;不回退混入该提交的订阅/计费功能。 |
| 钱包、订阅、兑换与支付 | 检查业务 URL 到前端跳转链路并修复 fragment 误拦;交易归属、金额、并发更新、幂等及回调校验维持原边界。 |
| 流式、WebSocket、格式转换 | 检查请求限制、认证头转发、取消/重定向、会话延续和正文捕获之间的关系;修正两处与恢复后的日志语义冲突的旧断言,不退回禁用功能的实现。 |
| 系统导入、导出、备份和恢复 | 区分交互式导入、恢复备份、回滚模式;保留加密备份、用途绑定、导入锁和凭据保留规则,不把安全导出误当作完整灾备。 |
| 数据契约、PostgreSQL 和迁移 | 检查脱敏投影与业务字段消费者是否矛盾;旧驱动不复活,使用真实迁移 schema 验证正文及视频的关键写入链路。 |
| 管理前端与外链 | 复核脱敏状态、权限显示、URL/导航校验、支付调用方,并执行全部前端测试与真实项目类型检查。 |
| 安装、更新、容器和发布流程 | 执行归档、链接、目录、写入目标、来源信任、compose 参数及发布供应链的隔离 Shell 测试;不运行实际部署。 |
## “全部修复”续轮完成项
以下问题已按根因处理,不再作为上一轮的未解决清单。历史类型问题、测试夹具与辅助器缺陷不统一归因为本次加固,功能恢复也不以撤销所有保护为代价。
### 1. 真实前端类型检查及相关运行时缺陷
- 修复原先 116 个文件中的 461 条类型诊断,将 `npm run type-check` 改为 `vue-tsc -b --force --pretty false`,确实检查引用的应用和工具项目,不再依赖顶层空项目的成功退出。
- 按现有接口契约补齐响应泛型、可空字段、配置 schema、图表时间轴、用户角色、事件参数及测试夹具。保留严格检查、ES2021、未知数据边界与动态配置扩展字段,未通过 `any`、忽略诊断、排除测试或降低配置绕过错误。
- 修复请求缓存和登录校验在 fetcher 同步抛错后不能正确清理 in-flight 状态的问题;失败后可按原有退避策略重试,仍隔离不同身份和旧请求。
- 修复 Provider 余额重试的迟到响应覆盖新加载结果,以及卸载后更新状态的问题;新增回归同时覆盖合法的零余额与签到失败值,不因真假值判断隐藏正常数据。
- 修复钱包模板中的刷新调用,并在用户密钥删除、路由配置保存等异步流程中固定操作目标,避免确认期间切换页面后作用到另一个对象。
### 2. 手动保留天数取了更激进的截止点
`usage_cleanup_window_with_override` 原来使用 `max` 合并截止时间,与界面“在策略内取更保守时间点”的设计相反:删除条件是记录时间早于截止点,选更晚的时间会扩大删除范围。
现改为每一保留层级分别取 `min`。详细正文、压缩正文、请求头、完整日志均不短于既有策略,也不短于本次手动指定天数。测试覆盖 0、5、30、180、400 天及“选中记录是原策略子集”的关系。显式立即清理模式不受此改动影响。
### 3. 可信原始数据库恢复可显式保留凭据
原始 JSONL 导入及数据库复制曾无条件替换密码哈希、API Key 和管理令牌,导致可信恢复/迁移也不能继续使用原凭据。新增 CLI `--preserve-credentials`,仅由操作者显式选择,不能由导入文件中的字段自行启用:
```sh
aether-gateway --database-driver postgres --postgres-url "$TARGET_DATABASE_URL" import --input /path/to/trusted.jsonl --preserve-credentials
aether-gateway copy --source-driver postgres --source-url "$SOURCE_DATABASE_URL" --target-driver postgres --target-url "$TARGET_DATABASE_URL" --preserve-credentials
```
- 不加该参数时仍默认撤销导入的身份凭据,并在操作前给出警告;原有库函数入口也保持该默认值。
- 显式保留时,只保留导入文件中用户密码、API Key 和管理令牌原有字段/状态;不会把原本停用的凭据重新启用。
- 导入的登录会话仍撤销,代理隧道仍重置在线状态和代际。外键、身份归属、OAuth 绑定一致性、输入大小及事务校验不变。
- 仅适用于可信且由操作者控制的备份/来源。目标实例须使用与来源兼容的加密配置;此开关不会解密、重新加密或重建已经丢失的值。复制链接的 TLS 默认保护不变。
- 这是原始数据库工具的选项,不改变 HTTP 配置导入权限,也不替代已有的加密备份恢复模式。操作示例会写入指定目标,执行前须自行核对目标并备份;本轮仅在临时隔离库中验证。
真实 PostgreSQL 回归分别验证默认撤销与显式保留后的密码哈希、密钥哈希、密文和启用状态;单元回归另覆盖管理令牌、会话及隧道边界。
### 4. 数据库回归夹具与测试资源回收
- 修正历史测试的超长 ID、NUMERIC/f64 解码、依赖预填充数据库的统计夹具,以及使用不在持久化契约内的 metadata 字段。使用真实 UUID、显式 SQL 类型转换、自建自清理数据和合法 trace 字段,未放宽生产 schema 或脱敏规则来迁就测试。
- 保留现有 usage 导出时间戳以秒计的兼容契约,未因旧字段名包含 `_ms` 就改变导入导出单位。
- 临时 PostgreSQL 辅助器改为 `pg_ctl stop -m fast` 并等待子进程退出,再清理自有目录;关闭失败时保留进程与目录的所有权,允许重试,不强杀后直接删数据。
- 新增两个真实 PostgreSQL 回归,覆盖打开连接下四次停止/重启、数据保留、重复停止、关闭失败重试及析构清理。专项运行前后共享内存段清单未新增残留;未清除其他业务或历史进程的 IPC 资源。
- 续轮全量回归进一步定位到迁移、回填测试各自复制的旧辅助器,同样在强杀后泄漏资源,导致后续网关测试无法初始化数据库。两份实现已合并到仅测试使用的 `postgres_test_support.rs`,采用相同的正常关闭与失败重试规则,另补两个回归;没有只清理环境而保留泄漏代码。
- 用 `AETHER_REQUIRE_LOCAL_POSTGRES_TESTS=1` 强制迁移、回填和共用辅助器的真实数据库用例执行,生命周期筛选结果为 63 通过、0 失败、1 忽略;该忽略项为已在隔离数据库中单独通过的导入测试。此次完整生命周期运行前后 IPC 清单一致。
- 辅助器支持 `AETHER_PG_CTL_BIN`,默认查找 `AETHER_POSTGRES_BIN` 同目录下的 `pg_ctl`;初始化失败也回收本次拥有的工作目录。
## 验证结果
| 验证层 | 结果 |
| --- | --- |
| Rust 工作区所有 library 与 binary 测试目标 | 最终 `cargo test --workspace --lib --bins --locked -- --test-threads=4`:65 个目标,8944 通过、0 失败、19 默认忽略;其中 41 个 library 目标 8645 通过,24 个 binary 目标 299 通过。设置 `AETHER_REQUIRE_LOCAL_POSTGRES_TESTS=1`,迁移/回填测试不能因环境问题静默跳过。 |
| 最终网关全量 | 随工作区 feature 合并执行:5124 通过,0 失败,含此前因数据库初始化失败的两项跨节点认证回归。使用仓库既有的 `RUST_MIN_STACK=16777216` 测试配置;不直接运行遗漏该配置的产物,也不改生产配置规避问题。 |
| PostgreSQL 适配器全量含真实数据库回归 | `AETHER_TEST_DATABASE_URL=<隔离库> cargo test -p aether-data-postgres --lib --locked -- --include-ignored --test-threads=1`:232 通过,0 失败,0 忽略,包含原先默认忽略的 16 项。 |
| 原始数据库导入/导出 | 设置独立的 `AETHER_TEST_POSTGRES_URL` 并执行 `lifecycle::export::tests --include-ignored`:18 通过,0 失败,0 忽略,包含迁移后数据库读取及两种凭据策略的真实往返。 |
| 临时 PostgreSQL 关闭/重启 | `aether-testkit --features postgres` 的两个真实数据库专项回归均通过;四次打开连接下重启、关闭失败重试、目录与共享内存回收均已验证。 |
| WebSocket 真实网关集成 | 最终重新执行 11 通过、0 失败,含跨连接继续对话、连续计费、长连接撤权、认证头隔离、未知字段透传、断开结算、额度重试和 PII 恢复。使用临时 PostgreSQL 与模拟上游。 |
| 可执行入口 | 299 项通过中包括网关主入口 61、隧道 185、备份恢复 CLI 10、两种 WebSocket 探针 3/4,以及压力测试种子/探针、模拟上游等入口测试。没有启动这些工具的实际生产操作。 |
| 额外入口与身份隔离集成 | `aether-data --test public_entrypoints` 3 项、`aether-gateway --test admin_unsigned_identity_headers` 1 项均重新执行通过。 |
| 前端全量与构建 | 最终 `npm run test:run`:200 个测试文件、1419 项测试通过;`npm run build:with-typecheck` 构建通过。 |
| 类型与格式 | `npm run type-check` 实际运行 `vue-tsc -b --force --pretty false`:0 错误,原先 461 条诊断均已修正;142 个改动/新增前端源码文件 ESLint 为 0 错误、0 警告。`cargo fmt --all -- --check` 和 `git diff --check` 通过。 |
| 提交前 Clippy | 本地按 `.github/workflows/rust-ci.yml` 的 Gateway、Data、其余工作区三组范围执行,均使用 `--locked` 和 `-D warnings`,全部通过;未关闭或放宽 lint 规则。 |
| 安装/发布脚本 | 10 份隔离 Shell 测试全部通过;`python3 tests/compose_database_config_test.py` 通过,只渲染配置,不启动容器。未实际安装、升级或部署。 |
工作区默认忽略的 19 项全部另行显式执行通过:PostgreSQL 适配器 16 项、真实导入 1 项、测试辅助器 2 项。默认命令中的忽略不计作通过;工作区 feature、单包测试及过滤测试的范围不同且存在重叠,上表不累加成独立用例总数。此前的规则、完整正文单条/批量持久化和视频真实数据库回归结果保留在兼容性复核报告中。
中间轮次曾因旧测试辅助器留下的 IPC 残留导致 2 项网关测试和随后 10 项 WebSocket 测试初始化失败;没有把这些命令记为成功。现已修正三处辅助器的关闭路径,并只清理本轮确认创建、无连接且创建进程已退出的 11 个段,未清理此前已有的 21 个段,也未修改内核共享内存限制。
最终工作区、WebSocket、公开入口和身份隔离测试顺序运行前后的 IPC 清单一致,均为此前已有的 21 个段,无新增残留。单独用于数据库适配器及导入往返的临时 PostgreSQL 已正常停止并删除本次自有目录。
脚本范围:`tests/deploy_state_safety_test.sh`、`tests/install_archive_safety_test.sh`、`tests/install_container_runtime_security_test.sh`、`tests/install_current_release_link_test.sh`、`tests/install_local_bundle_safety_test.sh`、`tests/install_privileged_write_safety_test.sh`、`tests/install_source_trust_test.sh`、`tests/release_supply_chain_test.sh`、`tests/tunnel_installer_config_security_test.sh`、`tests/update_compose_safety_test.sh`。
## 交付限制
本轮已确认、可复现且可在本仓库修复的剩余问题均已处理,未保留已确认却未修复的本轮源码问题;这一结论不等于“所有代码及所有部署环境绝无未知缺陷”。
这是本地源码、静态检查与隔离回归,不是对生产配置、第三方账户或所有部署组合的认证。上述验证未执行服务器部署或改动现有业务数据库;后续源码提交、推送不代表已经部署或完成生产环境验证。
已经被旧代码写空、删除、未采集的正文或视频字段不能由修复自动重建;仍在数据库中的旧内联/压缩正文可在修正后的保留链路中迁移。真实第三方 OAuth、支付、对象存储和远程隧道服务未使用生产凭据进行端到端验证。
@@ -1,259 +0,0 @@
# Aether 系统、测试与 CI 瘦身审计
- 审计日期:2026-09-08(Asia/Shanghai)
- 源码基线:`8b766930b`
范围:仓库结构、依赖图、本地构建产物、GitHub Actions 实际日志、测试与发布边界。
本次只新增审计报告,没有修改业务代码、测试或工作流,没有清理文件,也没有访问生产服务器。没有采集生产 RSS、CPU、数据库体积或请求延迟,因此下文不能作为生产内存泄漏或吞吐退化的结论。
## 一、结论
优先减掉的是**重复构建、过重的测试夹具、没有实际用例的构建任务和失效的模块边界**,不是先删功能或减少回归断言。
- 普通 Rust CI 最近 30 条记录中,20 次成功运行的总耗时中位数约 **16 分 33 秒**,范围 **11 分 42 秒~19 分 03 秒**。样本覆盖 2026-09-05~2026-09-08,包含 push 和 PR;没有将失败、取消或未完成运行计入。
- 成功样本中 Gateway 是关键路径:一次耗时 **11 分 39 秒**,另一次 **16 分 46 秒**;既有巨型测试目标编译,也有数分钟测试执行,不能只归咎于缓存或测试数量。
- “Workspace Rest” 虽然排除了 Gateway 测试目标,仍经由 Tunnel 的开发依赖编译完整 Gateway。分 job 没有实现真正的依赖隔离。
- Nightly 文档测试花了 **252 秒**,41 个库实际执行 **0 个 doctest**;发布默认编译 4 个二进制,只上传其中 1 个。
- 本地 `target/` 约 **168GB**,其中 `debug/incremental/` 约 **116GB**。这是构建缓存,不是生产镜像或业务源码体积。
所有改进收益都需要前后对照验证。本文不会把并行 job 的节省简单相加为流水线总时长,也不会承诺尚未测量的提速比例。
## 二、实测基线
### 2.1 CI 运行记录
数据来自 `fawney19/Aether` 的 GitHub Actions API、各 job 的步骤时间戳及原始日志。
| 运行 ID | 类型、源码 | 观察结果 |
| --- | --- | --- |
| `34174603131` | Rust CI,`7113d04f`,成功 | 总耗时 11:55;17 个 job 累计执行 39:19 |
| `34153166516` | Rust CI,`099b810a`,成功 | 总耗时 17:02;17 个 job 累计执行 46:10 |
| `34132961583` | 正式发布,`v0.7.17`,成功 | 总耗时 33:25;最慢为 macOS Intel 构建 |
| `34163371099` | Nightly,`099b810a`,失败 | 总耗时 47:42;检查和编译成功,GHCR 发布 job 在启动阶段失败 |
普通 CI 的总耗时按 `updated_at - run_started_at` 统计,包含调度和收尾;各 job 耗时按自身开始、完成时间统计,累计值不是计费分钟。详细成功样本与本地基线的 Rust CI、Nightly、Release、Cargo profile 和 Vitest 配置没有差异,业务改动会影响测试数量与耗时。
`34174603131` 中各主要测试步骤:
| 任务 | 编译/链接 | 测试执行 | 步骤或 job 耗时 |
| --- | --- | --- | --- |
| Gateway lib | 4:57 | 5,139 个测试,213.807 秒 | Test lib 步骤 514 秒 |
| Gateway bins | 2:27 | 78 个测试,0.270 秒 | Test bins 步骤 151 秒 |
| Workspace Rest | 5:36 | 3,377 个测试,26.357 秒,另有 16 个跳过 | Test 步骤 366 秒 |
| Integration Scenarios | 5:16 | 15 个 bin 内单测及 11 个 E2E 用例,约 8 秒 | Test 步骤 325 秒 |
| Data | 未进一步拆分 | 保留数据库相关保障 | Test 步骤 55 秒,job 88 秒 |
另一轮 `34153166516` 的 Gateway lib 编译 6:22、执行 390.039 秒,bins 编译 3:05、执行 0.329 秒。运行环境和缓存差异明显,不能只用最快一轮估算收益。
### 2.2 源码与磁盘
源码只统计 Git 跟踪文件,行数为物理行,包含注释和测试,不等于生产代码行数。
| 项目 | 规模 |
| --- | --- |
| Cargo workspace | 42 个 package |
| Rust 源码 | 1,874 个文件,1,114,527 行 |
| Gateway package 的 Rust 文件 | 1,176 个文件,679,538 行,约占全部 Rust 行数 61% |
| 以 tests/test/testkit 等命名识别的 Rust 测试及支持内容 | 214,086 行;未加上大量内联 `#[cfg(test)]` 模块 |
| Gateway 架构守卫测试 | 13 个文件,14,918 行,208 个 `#[test]` |
| 前端测试 | 208 个文件,31,945 行;Nightly 实跑 1,486 个用例 |
| 本地 `target/debug/incremental/` | 约 116GB |
| 本地 `target/debug/deps/` | 约 49GB |
| 本地 `frontend/node_modules/` / `frontend/dist/` | 约 311MB / 8MB |
| 本地历史 `htmlcov/` / `logs/` | 约 75MB / 121MB,均被 Git 忽略 |
## 三、优先处理:不降低保障的浪费
### P1-1:解除 Tunnel 测试对完整 Gateway 的反向依赖
**证据:** `apps/aether-tunnel/Cargo.toml:48` 将带 `testkit` 的 Gateway 列为 dev-dependency。`.github/workflows/rust-ci.yml:208` 和 `.github/workflows/rust-ci.yml:394` 虽然排除 Gateway 目标,但 Rest 的真实日志仍出现编译 `aether-gateway`。Gateway、Rest、Integration 因此在独立 runner 中重复付出重型构建成本。
**建议:** 将 Tunnel 中需要完整 Gateway 的端到端场景迁到独立集成测试目标;Tunnel 的协议、状态机、配置等单测只依赖轻量契约和测试支持。迁移之后比较测试清单,确保没有丢失端到端场景。
**验收:** 普通 Rest 单测及 Clippy 的依赖闭包不再包含 Gateway;Tunnel 跨端集成场景仍在专门任务执行。此项主要降低累计 runner 工作量,是否缩短总时长取决于 Gateway 关键路径是否也得到优化。
### P1-2:发布只编译真正发布的二进制
**证据:** Cargo metadata 显示 Gateway 有 4 个 bin:服务主程序、backup-restore、两个 WebSocket probe。`.github/workflows/release.yml:227`、`.github/workflows/release.yml:229` 和 `.github/workflows/nightly.yml:296` 未选择具体 bin,但 `.github/workflows/release.yml:236` 只上传主程序。
2026-09-07 的正式发布中,macOS Intel job 耗时 31:24,编译/链接日志耗时 28:52;Nightly 对应 job 耗时 33:45。发布慢不能全算在测试头上。
**建议:** 正式发布与 Nightly 的 `cargo build` / `cross build` 添加 `--bin aether-gateway`;probe 和其他运维程序保留独立检查、测试或按需打包入口。保留当前 release 优化策略,先测减少目标的收益,再讨论 LTO 调整。
**边界:** 没有逐 bin 链接计时,不能声称这会让整个发布缩短四分之三。
### P1-3:取消空 doctest 构建,合并重复的驱动检查
- `.github/workflows/nightly.yml:100` 的 workspace doctest 步骤耗时 252 秒,41 个库合计 0 个用例。建议对确实无 doctest、且不计划承载文档示例的库显式管理 `doctest`,或通过独立清单检查是否存在可执行文档示例后再调度。未来新增示例必须能重新纳入检查,不能永久盲目跳过全部文档测试。
- `crates/aether-data/runtime/Cargo.toml:10` 中 `default = ["postgres"]`,`all-drivers = ["postgres"]`,源码没有单独依赖 `all-drivers` 的条件分支。`.github/workflows/rust-ci.yml:334` 的两个 feature job 实际没有覆盖两个不同数据库驱动。
- `.github/workflows/rust-ci.yml:394` 的 Rest 已执行 Postgres adapter 的测试,`.github/workflows/rust-ci.yml:403` 又运行一次。若保留独立 adapter job,应从 Rest 排除它;否则直接以 Rest 承担这份覆盖。
**收益口径:** 在样本中,去掉空 doctest 可省约 4:12 的该 job 时间;feature 与 adapter 重复工作为几十秒量级。它们多数不在普通 CI 关键路径上,不应当作主 CI 总时长的等额收益。
### P1-4:区分集成测试与压测工具的构建
**证据:** `crates/aether-testing/integration/Cargo.toml:1` 所属 package 自动发现 14 个场景 bin;其中 11 个没有测试函数。`.github/workflows/rust-ci.yml:467` 对整个 package 执行 `--bins --tests`,耗时 5:25,实际测试约 8 秒。
这里并没有执行那些压测程序的 `main()` 来验证容量、恢复时间或性能;空测试 harness 的成功不能视为压测成功。
**建议:** 将 E2E 测试与 benchmark/probe 工具分开;把三个工具内的 15 个单测迁入适合的库或独立目标。工具源码继续接受 Clippy/check,真实压测由手动或计划任务运行。必要时使用 `required-features` 门控工具 bin。
**注意:** 仅给 bin 添加 `test = false` 不足以阻止所有额外构建;Cargo 构建 integration test 时还可能自动构建同 package 的普通 bin。应从目标和 package 边界解决,而不是仅换命令拼写。
### P1-5:重做按变更类型路由,同时补覆盖缺口
**证据:** `.github/workflows/rust-ci.yml:9` 将 README、安装脚本、Compose 与 Rust 源码共同触发整条 Rust 流水线;job 内没有进一步区分。当前五份工作流没有 frontend PR 工作流;前端完整检查在 Nightly 执行。
另一个现存缺口:`apps/aether-gateway/tests/admin_unsigned_identity_headers.rs:16` 的普通 integration test 不在 Gateway 的 `--lib`、`--bins` 两条 nextest 命令内;独立 Integration Scenarios job 选择的是另一个 package。Nightly 的 `check --all-targets` 只检查编译,不会替代执行此安全用例。
**建议:**
1. 安装脚本、Compose、发布工作流改动优先运行对应安全 fixtures;纯 README 文档修改不必编译完整 Gateway。
2. Rust 改动执行相关测试;Cargo、工具链、公共契约和测试基础设施变更应保守扩大到完整检查。
3. 前端改动运行自己的类型检查和测试,不必等待 Nightly。
4. 显式加入 Gateway integration test,包括上述身份头安全用例。
5. 保留稳定的最终 `check` 门禁,并验证预期跳过的 job。不要让路径过滤造成 required check 永久 pending,或把失败当成允许跳过。
工具链触发项还应补查 `rust-toolchain.toml`、`.cargo/**`;这些文件目前不在该工作流的路径列表中。
## 四、关键路径:Gateway 测试要减“重量”而非减断言
### P1-6:改造昂贵夹具与不必要的全应用初始化
成功样本中最慢的用例包括:
| 用例 | 时间 | 代码位置 |
| --- | --- | --- |
| 跳过大量 blocked account 的扫描预算 | 22.668 秒 | `apps/aether-gateway/src/dispatch/pool_scheduler.rs:3940` |
| v1 备份兼容及历史密钥尝试 | 14.324 秒 | `apps/aether-gateway/src/backup/executor.rs:1223` |
| 跳过大量 exhausted account | 8.481 秒 | `apps/aether-gateway/src/dispatch/pool_scheduler.rs:3859` |
| 大池 LRU 和动态跳过 | 8.147 秒 | `apps/aether-gateway/src/dispatch/pool_scheduler.rs:4424` |
池调度测试会构造 1,700 个账号,并在 `apps/aether-gateway/src/dispatch/pool_scheduler.rs:4868` 为每个账号调用真实凭据封装。可以将扫描预算、分页、跳过逻辑用轻量 repository/credential fixture 验证,另保留少量真实加解密联调用例和大规模边界场景;不要简单把大池规模缩小到失去原来的回归条件。
`crates/aether-crypto/src/python_fernet.rs:302` 的历史密钥派生包含进程内缓存及 100,000 次 PBKDF2。真实 nextest 用例是分进程执行的,跨用例不能指望共享这份缓存。是否构成主要开销仍需针对性计时;可为不测试历史派生逻辑的夹具选择合法的固定测试密钥或预制密文,历史兼容和真实密码学用例必须保留生产强度。
**不要做:** 为通过 CI 下调生产加密迭代数、删掉备份兼容测试、跨测试共享可变 AppState、将关键安全测试统一 ignore。
### P1-7:把巨大单一测试目标拆成真正独立的边界
`apps/aether-gateway/src/lib.rs:234` 将广泛的内部测试树纳入同一个 lib test binary。单纯把大文件拆成几个 `mod` 文件,不会使它们成为独立 Cargo 编译单元。
建议优先迁出无需访问私有业务状态的架构守卫测试,再逐步将调度策略、协议转换、计费纯函数测试迁到所属 crate;HTTP 行为与跨模块场景放到有明确支持 API 的 integration target。避免为了迁移测试而把所有内部类型公开。
架构守卫现有 208 个测试、约 1.49 万行,很多是源码字符串和依赖规则检查。例如 `apps/aether-gateway/src/tests/architecture/workspace_tiers.rs:3` 按 manifest 字符串判断依赖边界。它们应继续存在,但可进入不依赖 Gateway 的小工具/测试目标;依赖规则优先检查 Cargo metadata,行为正确性继续由行为测试负责。
若采用 nextest 分片,应先复用一次构建产物,再分发执行;直接给 N 个 runner 各自重新编译巨型 Gateway 会放大成本。先衡量编译与执行占比,再确定是否分片。
**特别注意:** 仅将 `--lib` 与 `--bins` 合为一条命令,并不消除普通库与 `cfg(test)` 测试库的两种构建。不能把样本中 2:27 的 bins 编译时间直接记作可全部省掉。
### P1-8:缓存按实际构建方式组织
所有 Rust job 使用同一个 `shared-key`。实测 Gateway、Rest、Integration 恢复了同一份约 380MB 的缓存;日志显示 `cache-workspace-crates: false`。但 Gateway 的 mold `RUSTFLAGS` 只存在于测试 step,见 `.github/workflows/rust-ci.yml:265`,其余任务又使用不同的 Clippy/check/test 和 feature 组合。
Gateway 样本的 Rust sccache 命中率达到 95.48%,仍花了数分钟构建,并存在 76 次标记为 `crate-type` 的不可缓存调用。这说明“再装一个缓存”不是根治;同时也不能据此断言缓存无效。
建议把影响缓存选择的环境提前到 job 级,按 lint/check 与 test、target、toolchain、必要 feature 区分缓存用途;相同构建尽量统一。先测恢复、保存、不可缓存调用与编译时间,不要给每个细碎目标无限新建 cache key,也不要直接缓存完整的巨型 `target/`。
现有 nextest、sccache、Gateway mold 和 CI 的关闭 debug 信息配置已经到位,不列为“尚未实施”的建议。
## 五、源码和依赖的长期瘦身
### P2-1:继续完成已有模块边界,而不是继续增加空壳 crate
Gateway 仍承载约 67.95 万行 Rust。部分现有边界很薄:Gateway execution crate 82 行、control crate 100 行、provider core 91 行、usage core 110 行,而核心实现仍留在应用中。
值得分批治理的集中点:
| 文件 | 物理行数 | 建议拆分依据 |
| --- | --- | --- |
| `apps/aether-gateway/src/execution_runtime/stream/execution.rs:1` | 15,088 | 流状态机、传输适配、计费收尾、对应测试 |
| `crates/aether-data/adapters/postgres/src/usage/mod.rs:1` | 13,833 | 写入、查询、统计聚合、审计存储 |
| `crates/aether-usage/runtime/src/runtime.rs:1` | 12,809 | 状态推进、结算策略、持久化适配、测试 |
| `apps/aether-gateway/src/handlers/admin/request/system/import.rs:1` | 9,490 | 导入校验、版本兼容、执行与回滚 |
这些数字包含内联测试,不是生产实现行数。目标应是缩小依赖闭包、变更影响面和测试目标,而非追求拆出更多文件。
`apps/aether-gateway/src/state/app.rs:376` 的 AppState 有 91 个字段,其中 28 个字段名包含 cache;`crates/aether-data/contracts/src/repository/usage/types.rs:1984` 的 UpsertUsageRecord 有 67 个字段。这反映了测试构造和模块依赖面较宽,但不能据此直接判定运行时占用过大。可以引入按职责的窄上下文和统一 fixture builder,避免每个测试复制完整对象。
### P2-2:从基础契约中剥离重型格式实现
`crates/aether-data/contracts/Cargo.toml:10` 依赖整个 `aether-ai-formats`;后者约 8.23 万行 Rust。contracts 中实际使用包括格式权限、别名与少量 usage 元数据策略,见 `crates/aether-data/contracts/src/repository/auth.rs:295`。
建议把稳定的格式标识、权限和小型元数据契约下沉到现有基础契约层,完整 request/response/stream 转换留在 formats。避免为了一个格式权限判断,让数据库契约持续依赖整个转换实现。这比任意合并 crate 更有价值。
### P2-3:依赖体积优化必须保留兼容能力
`cargo tree --offline --locked` 确认 Gateway 同时包含:
- 主请求链路的 reqwest 0.12 与 `object_store` 引入的 reqwest 0.13。
- rustls 的 ring 和 aws-lc 路径,以及 wreq 的 boring2 路径。
这些是进一步分析构建和二进制体积的候选,不是已证实可直接删除的依赖。`aws-lc-rs` 还有直接密码学用途,wreq 承担专门传输能力;移除前必须核查调用和握手、指纹、代理兼容测试。
先获取 release 的 Cargo timings 和二进制符号/section 体积,再评估版本统一或可选能力 profile;不要仅根据依赖名字或锁文件重复条目盲目替换。
## 六、前端测试与构建
### P1/P2:测试环境按需要加载
Nightly 前端测试实测 106.01 秒,208 个文件、1,486 个用例全部通过。Vitest 报告 environment 146.40 秒、import 51.29 秒、tests 72.78 秒;这些分项含并行累计时间,不能相加当作墙钟时间。
`frontend/vitest.config.ts:10` 为全部测试使用 jsdom;静态扫描有 126/208 个测试文件未直接出现常见 DOM 操作标记,但这并不证明它们的传递依赖不需要 DOM。`frontend/src/tests/vitest.setup.ts:61` 还在每个用例前加载 i18n 并重设语言。
建议通过独立 Vitest project 或显式环境标记,将已确认的纯函数/解析器/数据转换测试放到 node 环境,DOM 组件保留 jsdom;按测试类别拆分 setup,不要全局取消隔离。先迁一组、验证测试数与结果,再扩展。
### P1:同一 web 项目重复构建
`.github/workflows/release.yml:84` 和 `.github/workflows/deploy-pages.yml:58` 已单独构建 VSCodex web;随后 frontend 的 `prebuild` 又调用 `frontend/scripts/sync-vscodex.mjs:64` 无条件重建一次。
建议分开“安装/构建嵌入 web”与“复制已构建产物”,在同一个任务内只构建一次,消费明确来源的 artifact。不要仅通过 `dist` 存在就认定源码和产物一致。Nightly 前端这里没有相同的预先重复 build,不应错误地宣称所有流水线都重复。
### P2:低风险依赖清理候选
当前前端源码未发现 `three` 的模块导入,但 package 声明了 `three` 和 `@types/three`;本地二者合计约 36MB。确认无动态/外部消费者后可移除并更新锁文件,验证 type-check、测试和构建。
这主要减少安装和维护成本;不能保证生产 bundle 同样减少 36MB。现有 chart、pinyin、Stripe 均发现使用,不能一并判作无用依赖。
MarkdownViewer 使用完整 `highlight.js`,而 CodeHighlight 已按语言导入,可统一按需高亮策略;现有前端产物约 8MB,优先级低于 Rust 重复构建和测试初始化。
## 七、本地构建产物治理
当前最值得清理的是 116GB 的 `target/debug/incremental/`,其次是 49GB 的 `target/debug/deps/`。全量 `cargo clean` 虽能回收空间,也会迫使下次重建全部依赖。
建议先确认没有进行中的 Cargo/rustc 任务,对长期未使用的增量会话、历史目标产物做定期清理;稳定本地 profile、工具链和 `RUSTFLAGS`,避免频繁产生不同构建组合。保留仍在使用的依赖缓存,不要每次构建前清空 target。
`htmlcov/` 属于历史 Python 覆盖率产物,可在确认不再需要后清理,但仅几十 MB,不是主要收益。不要误删仍有效的 Python 安装/Compose 安全测试,也不要把日志、备份或数据目录当构建缓存删除。
生产 Dockerfile 已使用预构建二进制、前端产物和 distroless runtime,并通过 `.dockerignore` 排除 target、node_modules 等开发内容;不建议把“换更小基础镜像”列为当前第一优先级。生产镜像实际体积还需单独测量。
## 八、建议实施顺序与验收
| 批次 | 改造 | 验收标准 |
| --- | --- | --- |
| A:小改动去空转 | 发布指定主 bin;管理空 doctest;消除重复 adapter 与等价 feature 任务;嵌入 web 只构建一次 | 保留现有有效用例与工件;对照任务耗时和测试清单 |
| B:加快反馈并补漏 | 变更路由、稳定 gate、前端 PR 检查、Gateway integration 安全用例 | 纯文档/脚本不编译全栈;公共变更仍完整检查;安全用例实际执行 |
| C:解除编译耦合 | Tunnel 跨端测试迁移;工具与 E2E 分离;架构守卫轻量化 | Rest 依赖闭包无 Gateway;无用 bin 不参与 E2E 构建;守卫持续有效 |
| D:减少测试初始化 | 重型夹具拆层;纯前端测试切 node;窄上下文和统一 builder | 慢用例时间降低,测试数与断言目的不减少,无新增污染或 flaky |
| E:结构与依赖治理 | 真实职责迁入现有 crate;基础格式契约下沉;审慎统一依赖 | 普通修改触发的重编译面缩小,性能和兼容性基线不回退 |
每批先使用相同 SHA、runner 类型、toolchain 和 feature 集合做对照,至少区分冷缓存/热缓存与 job 执行/流水线总耗时。持续保存编译 timings、nextest 测试清单和结果、慢测试列表、缓存统计、frontend environment 时间、发布工件体积。
首要结果指标:普通 PR 更快得到正确反馈、累计重复构建下降、测试保障不退化。不要把删除测试数量、增大并发数或缩短单个非关键任务当作最终目标。
## 九、复查入口
本次没有重新运行全量构建或全量测试,使用了真实 CI 日志、离线 Cargo metadata/tree 和只读源码统计。临时 API JSON、job 日志与依赖树保存在本机 `/tmp/aether-slim-audit/`,未纳入 Git;临时目录可能被系统清理,运行 ID 可用于再次取证。
可重复的只读命令:
```sh
cargo metadata --offline --locked --no-deps --format-version 1
cargo tree --offline --locked -p aether-gateway -e normal -d
gh api 'repos/fawney19/Aether/actions/runs/34174603131/jobs?per_page=100'
gh api 'repos/fawney19/Aether/actions/jobs/101901417322/logs'
gh api 'repos/fawney19/Aether/actions/jobs/101869558414/logs'
du -h -d 1 target/debug
```
Cargo 关于默认构建目标及 integration test 自动构建 bin 的语义,另对照本机 Rust 1.95.0 随附的 Cargo `cargo-build`、`cargo-test`、`cargo-targets` 官方文档。并行测试时间、缓存命中率和源码体积均按各自定义解释,未将其混用为生产性能结论。
@@ -1,94 +0,0 @@
# TLS Fingerprint Capture
Aether stores per-request TLS capture under `usage.request_metadata.tls_fingerprint`.
```json
{
"tls_fingerprint": {
"incoming": {
"source": "forwarded_header",
"ja3": "...",
"ja3_hash": "...",
"ja4": "...",
"protocol": "TLSv1.3",
"cipher": "TLS_AES_128_GCM_SHA256",
"sni": "api.example.com",
"alpn": "h2"
},
"outgoing": {
"source": "aether_transport_config",
"observed": false,
"transport_path": "direct",
"backend": "reqwest_rustls",
"http_mode": "auto",
"tls_stack": "rustls",
"tls_versions_offered": ["TLS1.3", "TLS1.2"],
"alpn_offered": ["h2", "http/1.1"]
}
}
}
```
`incoming` is the client-to-Aether TLS fingerprint. It can be populated by Aether native TLS capture in direct deployments or by trusted reverse-proxy headers when TLS terminates before Aether.
`outgoing` is the Aether-to-provider TLS transport record. The current gateway records the exact transport configuration it controls. It sets `observed: false` because reqwest/rustls does not expose the emitted ClientHello bytes on the direct path. A future connector-level ClientHello capture or probe result can reuse the same object with `observed: true` plus `ja3`, `ja3_hash`, and `ja4`.
## Nginx TLS Termination
When nginx terminates HTTPS and proxies HTTP to Aether, Aether cannot see the original ClientHello. Configure nginx to forward the TLS fields it can observe:
```nginx
server {
listen 443 ssl http2;
server_name api.example.com;
ssl_certificate /etc/letsencrypt/live/api.example.com/fullchain.pem;
ssl_certificate_key /etc/letsencrypt/live/api.example.com/privkey.pem;
location / {
proxy_pass http://127.0.0.1:3000;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_set_header X-Aether-TLS-Source nginx;
proxy_set_header X-Aether-TLS-Protocol $ssl_protocol;
proxy_set_header X-Aether-TLS-Cipher $ssl_cipher;
proxy_set_header X-Aether-TLS-SNI $ssl_server_name;
}
}
```
Stock nginx does not provide JA3/JA4 variables. The forwarded record is still useful, but it is not a complete TLS fingerprint. To forward JA3/JA4 through nginx, use an nginx build/module or edge layer that computes them and set:
```nginx
proxy_set_header X-Aether-TLS-JA3 $ja3;
proxy_set_header X-Aether-TLS-JA3-Hash $ja3_hash;
proxy_set_header X-Aether-TLS-JA4 $ja4;
```
Only accept these headers from trusted infrastructure. Do not expose Aether directly to public clients while also trusting client-supplied `X-Aether-TLS-*` headers.
## Nginx TCP Passthrough
If Aether terminates TLS itself, nginx can pass TCP through without decrypting:
```nginx
stream {
map $ssl_preread_server_name $aether_backend {
api.example.com 127.0.0.1:3443;
default 127.0.0.1:3443;
}
server {
listen 443;
proxy_pass $aether_backend;
ssl_preread on;
}
}
```
In this mode nginx cannot inject HTTP headers because it never sees HTTP. Aether native TLS capture is responsible for populating `tls_fingerprint.incoming`.
-65
View File
@@ -1,65 +0,0 @@
# 请求正文查看与性能边界
## 按需读取
请求详情的轻量读取仍使用 `GET /api/admin/usage/{id}?include_bodies=false`,返回正文可用性与记录概要,不解析正文。
查看正文时,管理界面使用 `GET /api/admin/usage/{id}?include_bodies=true&body_field={field}&body_format=raw`。`body_field` 仅允许以下值:
- `request_body`:客户端请求体。
- `provider_request_body`:提供商请求体。
- `response_body`:提供商响应体。
- `client_response_body`:客户端响应体。
网页正文读取不再经过服务器 JSON 解码链路。数据库中的 `payload_gzip` 原样返回,历史压缩列同样直传;历史内联 JSON 返回 JSON 字节。接口仍经过管理权限校验、正文捕获状态检查、引用归属校验及审计,不加载用户名称和提供商名称等无关详情。
二进制响应的 `X-Aether-Body-Encoding` 为 `gzip` 或 `json`,另含 `X-Aether-Usage-Id`、`X-Aether-Body-Field`,前端验证记录与字段匹配。使用 `application/octet-stream`、`Content-Encoding: identity`,避免 HTTP 中间件重复压缩或浏览器提前解压;`Cache-Control: no-store, no-transform` 避免缓存敏感正文及代理改写。跨域允许凭证时显式暴露上述协议头。错误通过 HTTP 状态及 `X-Aether-Body-Error` 返回,不要求主线程解析二进制错误响应。
未指定 `body_format=raw` 的服务端 JSON 接口保持原有行为,但网页不再调用它加载正文;前端详情 API 默认也只取概要。非法字段、空字段、raw 模式缺少字段或同时设置 `include_bodies=false` 返回 HTTP 400。
前端按请求和正文字段保留 Worker 句柄,最多缓存两份已解析正文,并按解压字节总量 64 MiB 预算淘汰最久未访问的 Worker。该预算不是浏览器进程内存上限:解析对象、字符串及加载中的数据仍有额外开销。切换正文来源或标签会取消未完成的下载与解压;关闭抽屉、切换记录或组件卸载会终止全部 Worker。迟到结果不显示、不入缓存。请求从进行中变为完成或失败时,正文缓存失效,并重新读取当前选中的正文。
## 渲染与资源控制
- 仅挂载当前标签的内容,不再后台渲染隐藏标签。
- JSON 与对话均连续虚拟滚动,不需要点击上一页或下一页。接近底部时自动读取后续内容,向上滚动可重新查看先前内容。
- 折叠节点不遍历其子节点;点击括号展开节点时完整展开该子树,无需逐层点击。JSON 内部按 50 个显示片段批量读取,视口与预读区域最多挂载 4 批(200 个片段)。离屏内容用高度占位,缓存仅保留视口附近最多 6 批,不随滚动积累 DOM 或正文副本。
- JSON 不再有“显示更多”或“继续显示”按钮:长字符串和长键名在 Worker 中分段转义,随滚动自动显示全部字符。纯文本及非 JSON 响应同样自动衔接完整内容。分段保持 Unicode 字符完整,续段不重复显示 JSON 行号。
- 虚拟显示范围不改变正文长度;复制按钮仍复制完整内容。JSON 顺序读取复用遍历游标,避免每次滚动都从正文起点重新遍历。
- 压缩数据以 transferable ArrayBuffer 交给 Worker;解压、UTF-8 解码、JSON 解析、JSON 遍历及对话解析均在 Worker 中进行,完整对象不返回页面主线程。
- 对话内部按每批最多 10 个顶层块预览并自动衔接,限制嵌套块和文本传输量,长内容明确提示并支持增加预览。完整复制在用户点击后由 Worker 生成,不受虚拟显示范围影响。
- Worker 解压时逐块检查 64 MiB 上限,损坏 gzip、无效 UTF-8 和无效 JSON 给出明确错误。后台任务 30 秒超时会终止 Worker;不支持 Worker 或原生 gzip 解压的浏览器提示升级,不回退到主线程解析。
- 服务端最多 4 个后台解码任务的保护仍用于复制 cURL、请求重放等实际内部调用,不是网页解压的兼容回退。存储读取实现复用,不维护两套数据库读取逻辑。
- 数据库读取先按存储字节过滤:压缩正文最大 65 MiB,未压缩正文最大 64 MiB,超限不将完整载荷读入网关;浏览器还会独立校验解压后的大小。
## 错误定位
二进制接口的 `X-Aether-Body-Error` 和 JSON 接口的 `body_load_error_codes` 使用以下存储错误码;浏览器解压失败也映射为相同的明确提示:
| 错误码 | 含义 |
| --- | --- |
| `too_large` | 解压后正文超过 64 MiB 的安全读取上限。 |
| `decode_failed` | 正文解压或 JSON 解析失败。 |
| `missing` | 记录标记正文可用,但未能解析到实际存储内容。 |
| `storage_unavailable` | 其他存储读取错误;内部连接信息不会返回给页面。 |
前端分别显示请求超时、网络失败、HTTP 错误及上述存储错误。`too_large` 和 `decode_failed` 不提供无意义的重复重试;其他错误可手动重试。正文请求仍沿用现有 API 超时配置。
完整记录模式与已有记录不做静默截断;64 MiB 解压安全上限没有放宽。因此,超出上限的历史正文仍不能在线预览,但页面会明确说明限制,不再仅显示笼统的加载失败。
前后端需要一同更新。前端收到缺少二进制协议头、记录不匹配或字段不匹配的响应时会拒绝显示,并提示检查前后端版本;不会自动退回耗资源的完整 JSON 正文接口。构建前端时必须同时发布生成的 Worker JavaScript 资源。
## 针对性验证
```sh
cd frontend
npm run test:run -- src/features/usage src/api/__tests__/dashboard-body-loading.spec.ts
npm run type-check
```
```sh
cargo test -p aether-data-postgres usage_body_decode --lib --offline
cargo test -p aether-gateway admin_usage_detail --lib --offline
cargo test -p aether-gateway raw_body --lib --offline
cargo test -p aether-gateway body_load_errors_expose_safe_codes --lib --offline
```
-17
View File
@@ -1,17 +0,0 @@
# Usage header capture
Usage HTTP captures preserve original header values for client requests,
provider requests, provider responses, and client responses. Header maps are
validated as JSON objects but their values are not redacted. Capture settings
and existing access controls still apply.
These records can contain credentials such as Authorization, API keys, and
cookies, as well as session identifiers and client metadata. Restrict access
to usage details, database records, exports, and backups accordingly.
Previously stored `[redacted]` values cannot be recovered. Original values
are available only for new captures after deploying this change.
The request detail drawer reads these values directly from the administrator
usage-detail API, including when bodies are not loaded. No frontend setting
can recover header values that were already replaced during capture.