feat(gateway): Responses WebSocket 连通性探针

新增 aether-codex-ws-probe 与 aether-openai-responses-ws-probe 两个
二进制,用于在不暴露凭据的前提下验证上游 WebSocket 端点可用性:凭据
只从环境变量读取,不写入日志。公共流程放在
bin/support/responses_ws_probe.rs,各 profile 只负责自己的鉴权与
请求头要求。
This commit is contained in:
AAEE86
2026-08-17 14:51:18 +08:00
committed by ZheFox
parent 71b54070e8
commit a498875591
5 changed files with 1014 additions and 0 deletions
@@ -0,0 +1,158 @@
# Codex Responses WebSocket probe
`aether-codex-ws-probe` is a P0 compatibility probe for a Codex-compatible
Responses WebSocket upstream. It verifies two sequential `response.create`
warmups on one socket, with the second request continuing from the first
response ID.
The command, environment variables, JSON report shape, and Codex-specific
handshake headers remain stable. It now shares only the protocol-driving core
with the separate [OpenAI Responses WebSocket probe](openai-responses-websocket-probe.md);
the two probes intentionally retain independent authentication profiles and
provider-specific assertions.
The probe is intentionally not a production proxy. It does not persist,
refresh, log, or print credentials, account IDs, response IDs, request bodies,
or response bodies.
## Prerequisites
Use a dedicated, non-production Codex test account. Rotate any credential that
has been pasted into a chat, terminal history, issue, or source file before
using this probe.
Set these values only in the process environment or your secret manager:
```bash
export AETHER_CODEX_WS_PROBE_URL='wss://your-codex-upstream.example/backend-api/codex/responses'
export AETHER_CODEX_WS_PROBE_ACCESS_TOKEN='your-short-lived-access-token'
export AETHER_CODEX_WS_PROBE_ACCOUNT_ID='your-account-id'
export AETHER_CODEX_WS_PROBE_MODEL='your-codex-model'
```
The endpoint must use `ws://` or `wss://`, with no credentials, query string,
or fragment. The access token is accepted only through
`AETHER_CODEX_WS_PROBE_ACCESS_TOKEN`; there is deliberately no command-line
flag for it.
## Run
```bash
cargo run -p aether-gateway --bin aether-codex-ws-probe
```
Use `--url` to override only the endpoint and `--timeout-secs` to set a
per-turn receive timeout (1120 seconds):
```bash
cargo run -p aether-gateway --bin aether-codex-ws-probe -- \
--url 'wss://your-codex-upstream.example/backend-api/codex/responses' \
--timeout-secs 30
```
The probe emits one JSON line. A successful run has
`"continuation_confirmed":true`; its event and header fields contain names
only, never values. A failure emits a stable error code such as
`"handshake_failed"`, `"upstream_error_event"`, or
`"response_id_not_observed"`.
## Interpretation
A successful probe establishes that the selected upstream accepts the
Responses WebSocket handshake and retains continuation state on one socket.
It does not establish that all Codex models, account plans, or tunnel egress
paths are supported. In particular, the current `aether-tunnel` HTTP relay
does not forward WebSocket upgrades, so a successful direct probe is a
prerequisite rather than tunnel support.
## Gateway bridge
The gateway exposes WebSocket mode at the same public Responses path:
```text
wss://<aether-gateway>/v1/responses
```
It is disabled by default per provider. In **添加提供商** or **编辑提供商**,
enable **Responses WebSocket 模式** under **功能开关** only after the selected
upstream has passed a compatible WebSocket probe. The setting takes effect for
new WebSocket connections without a gateway restart. It is available to every
provider type; candidate planning still requires a selected
`openai:responses` endpoint.
Authenticate the upgrade request with the normal Aether API key. The first
client frame must be a text JSON `response.create` containing a non-empty
`model`. Aether then applies its regular Responses candidate selection, but
accepts only an eligible, WebSocket-enabled endpoint using `openai:responses`.
It opens an upstream WebSocket using the selected provider key.
The selected provider's model mapping and request headers are applied to every
turn, along with the rest of that candidate's provider-body normalization:
model-directive patches, endpoint body rules, and the Codex body contract
(unsupported-field stripping, `store: false`, `tool_choice` defaulting). A
continuation turn is normalized against the binding it is pinned to rather than
being re-planned, so it can never move to another provider key.
`previous_response_id` and `generate` are re-applied after normalization
because they are WebSocket protocol state that the provider body contract
otherwise strips. `stream` and `background` are removed because they are HTTP
transport fields, not WebSocket-mode fields. If a later `response.create` changes the
public model, Aether runs access checks and candidate planning again. It keeps
the existing upstream when the same target remains eligible, or transparently
replaces the upstream between responses when the selected target changes.
Overlapping responses on one client socket remain rejected.
Each `response.create` is tracked as an independent Aether logical request:
it receives its own request/candidate identity, usage lifecycle, and terminal
audit record. `response.completed`, `response.failed`,
`response.incomplete`, `response.cancelled`, client disconnects, and upstream
transport failures all settle that turn through the existing stream reporting
path.
Example client setup:
```python
from websocket import create_connection
import json
import os
ws = create_connection(
"wss://gateway.example/v1/responses",
header=[f"Authorization: Bearer {os.environ['AETHER_API_KEY']}"],
)
ws.send(json.dumps({
"type": "response.create",
"model": "your-public-model",
"store": False,
"input": "Explain this repository.",
}))
```
### Operating limits
- Maximum frame and message size: 16 MiB.
- An idle connection must send its first `response.create` within 60 seconds.
- A connection is closed after 60 minutes; reconnect before then for long runs.
- Each `response.create` must receive its first upstream event within the
selected provider's `stream_first_byte_timeout` (30 seconds by default),
and finish within its `request_timeout` (20 minutes by default). Aether
sends `responses_websocket_first_event_timeout` or
`responses_websocket_turn_timeout` and closes the bound socket when either
deadline expires.
- Responses are sequential; no multiplexing is supported on one socket.
- Each `response.create` consumes the normal Aether user/API-key RPM budget.
- Same-model turns stay on the bound provider key. A model change is planned
again and can rebind the upstream between completed turns when necessary.
- Direct provider proxy settings are honored through the selected transport
profile. Tunnel-mode proxy nodes are not supported for this bridge yet.
Usage and audit finalization now runs for every accepted `response.create`.
Existing usage body-capture and header-redaction policies apply to the resulting
records. Newly created WebSocket usage records expose `is_websocket=true`, and
the usage-record type column renders them as `WS`. For diagnosis, enable debug
logging for `aether_gateway::handlers::proxy::responses_ws`; event logs contain
only the event type and frame size, never request or response contents. Codex
quota-extension logs remain under `aether_gateway::handlers::proxy::codex_ws`.
Every WebSocket-specific log carries `transport="websocket"` and
`websocket=true`; keep `log_type` for its existing access/event/ops
classification, and render the transport flag as a `WS` label in a log viewer
if desired.
@@ -0,0 +1,61 @@
# OpenAI Responses WebSocket probe
`aether-openai-responses-ws-probe` verifies the official OpenAI Responses
WebSocket protocol using standard API-key Bearer authentication. It sends two
sequential `response.create` warmups on one socket, chaining the second from
the first response ID with `previous_response_id`.
It shares its protocol-driving core with the Codex probe, but it does **not**
send Codex account headers or require Codex quota events. This makes it the
compatibility gate for Aether's standard Responses WebSocket adapter, rather
than a replacement for the Codex probe.
## Prerequisites
Use a dedicated API project and a model that your key can access. Keep values
only in your process environment or secret manager:
```bash
export AETHER_OPENAI_WS_PROBE_API_KEY='your-api-key'
export AETHER_OPENAI_WS_PROBE_MODEL='your-openai-model'
```
The default endpoint is the official Responses WebSocket endpoint:
```text
wss://api.openai.com/v1/responses
```
To test a compatible endpoint explicitly, set
`AETHER_OPENAI_WS_PROBE_URL` or pass `--url`. The endpoint must use `ws://` or
`wss://` and may not contain credentials, a query string, or a fragment. The
API key has no command-line flag and is never printed.
## Run
```bash
cargo run -p aether-gateway --bin aether-openai-responses-ws-probe
```
For an explicit endpoint and timeout:
```bash
cargo run -p aether-gateway --bin aether-openai-responses-ws-probe -- \
--url 'wss://api.openai.com/v1/responses' \
--timeout-secs 30
```
The probe uses `generate:false`, so the warmups prepare continuation state but
do not request model output. A successful JSON report contains
`"continuation_confirmed":true`; header and event arrays contain names only,
never credentials, response IDs, request bodies, or response bodies.
## Interpretation
Success establishes that this key, model, and endpoint support the Responses
WebSocket handshake plus an in-socket continuation. It does not establish
support for every model, tool, service tier, proxy path, or Aether provider
configuration. Treat a successful direct probe as a prerequisite before
enabling **Responses WebSocket mode** for the matching Aether provider.
For protocol details, see the official [WebSocket Mode guide](https://developers.openai.com/api/docs/guides/websocket-mode).