The partial-preface test raced a 5ms first-request deadline against a
10ms hyper header_read_timeout. On a slow runner both timers expire
before the next poll and tokio::select! may pick the connection branch,
surfacing hyper's header-timeout error instead of the clean deadline
close. Push hyper's timeout out to 30s so only the deadline can fire.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Verified against a live Antigravity + xAI deployment:
- Gemini and Claude clients calling a forced-stream Responses upstream
(xAI, Codex) without streaming always failed: the aggregated body echoes
request metadata (parallel_tool_calls, tools, encrypted reasoning) that
the strict cross-format check refuses, and the gateway then wrapped the
raw SSE capture in a client error body sent with HTTP 200. Project the
validated aggregate to every client format, as the Chat path already
does, and return 502 instead of raw provider bytes when a successful
cross-format response still cannot be converted.
- Gemini stream decoding keyed tool calls by part position, so parallel
calls arriving in separate chunks (all at parts[0]) merged into one call
with concatenated arguments. Key them by arrival order; ids cannot be
used because they are optional and the Antigravity envelope synthesizes
per-chunk ids that repeat across chunks. Generated call_auto_N ids now
follow arrival order.
- Non-stream Responses output reported truncated or filtered cross-format
answers as completed; derive incomplete + incomplete_details from the
canonical stop reason.
- Gemini request parsing ignored parametersJsonSchema and
responseJsonSchema and passed OpenAPI upper-case type names (OBJECT,
STRING) through to JSON Schema targets, which xAI rejects.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Gemini runs `googleSearch` inside Google. The search leaves no
client-visible tool call, and the evidence arrives only as
`candidates[].groundingMetadata`. Every cross-format target dropped it
wholesale, so a grounded answer reached OpenAI- and Claude-shaped clients
as prose that names its sources with nothing structured behind it: no
`annotations`, no `citations`, no `url_citation`. Callers that verify
grounding — the common "did this model actually search?" check — saw a
200 with no evidence and had to treat the answer as ungrounded.
Adapters now normalise `groundingMetadata` into neutral citations and
each target renders its own family's standard shape: `url_citation`
annotations for `openai:chat` and `openai:responses`, and
`web_search_result_location` citations on the text block for
`claude:messages`. Gemini reports segment bounds as UTF-8 byte offsets
while both targets count characters, so the bounds are converted rather
than copied.
Streaming is covered too, since that is what grounded traffic actually
uses. A new `CanonicalStreamEvent::Citations` carries the neutral list
once the answer text is whole — the offsets index into the finished
answer, so it rides just ahead of `Finish` rather than as a delta per
chunk — and each client emitter renders it: `delta.annotations` chunks,
`response.output_text.annotation.added` events (also kept on the finished
message item so clients that only read `response.completed` see them),
and `citations_delta` content block deltas.
For reference, CLIProxyAPI projects grounding only in its
antigravity→Claude translator, and only when the client declared a typed
`web_search_*` tool; its OpenAI and plain Gemini translators have no
grounding handling at all. The citation shape here matches theirs, but
the coverage is deliberately wider: all three targets, streaming and
non-streaming, with no dependency on a declared tool.
Co-Authored-By: Claude Opus 5 <[email protected]>
The transport boundary rewrote `googleSearch` into the Gemini 1.5-era
`googleSearchRetrieval` spelling before every v1internal call, on the stated
grounds that the private backend rejects `googleSearch` when it is combined
with function declarations. That rewrite breaks grounding on Gemini 3.
Observed on stabey-124 against daily-cloudcode-pa.googleapis.com. A controlled
pair, same model and keys, 5 seconds apart:
- no `web_search_options` -> 200
- with `web_search_options` -> 502 on all three candidates
The outgoing body carried `tools: [{"googleSearchRetrieval": {}}]` and no
function declarations at all, so the documented mixed-tool rationale did not
apply. `request_candidates.error_message` holds what the backend actually
said:
Malformed function call: call:google_search{query:current UTC date time}
Malformed function call: call:google:search{query:current UTC date}
Malformed function call: call:google_search{queries:[current UTC date]}
The model reaches for `google_search`, the legacy declaration binds nothing,
and the turn dies unparsed. CLIProxyAPI sends `googleSearch` to this same
v1internal surface, including alongside function declarations.
Keep folding the snake_case `google_search` alias into the canonical
`googleSearch` key, and leave a request that already spells the tool
`googleSearchRetrieval` untouched.
Co-Authored-By: Claude Opus 5 <[email protected]>
canonical_tools_to_gemini promoted any tool whose name normalized to
"websearch" / "googlesearch" / "websearchpreview" into Gemini's server-side
builtin, dropping it from functionDeclarations. Claude Code declares an
ordinary client-side `WebSearch` tool with a full input_schema, so every
/v1/messages request routed to a Gemini model lost that declaration and gained
`googleSearch` (rewritten to `googleSearchRetrieval` at the Anti Gravity
transport boundary) instead.
Two consequences, both observed on stabey-124 against gemini-3.8-flash:
- the model can never emit a `WebSearch` tool_use, so the client's own web
search is dead on that route;
- when the model does reach for the injected server-side search, the v1internal
backend answers `finishReason: MALFORMED_FUNCTION_CALL` /
"Function call is empty - no input to parse." and the turn fails.
Promote a tool to a builtin only when it is a bare marker carrying no schema.
A declared schema means the caller intends to execute the call itself, which
matches CLIProxyAPI: it keys builtins off Claude's `type: web_search_*` or an
explicit `google_search` tool key and never off a function name.
Co-Authored-By: Claude Opus 5 <[email protected]>
OpenRouter reports reasoning under `reasoning` and `reasoning_details`
rather than the DeepSeek-style `reasoning_content` this crate recognized.
Its streaming reasoning phase sends chunks whose `delta.content` is an
empty string, so those chunks were dropped and Responses clients saw
nothing after `response.in_progress` until they timed the stream out.
The sync aggregator kept only content and tool calls, so a stream
downgraded to a sync response lost the reasoning entirely.
Read all three spellings through one helper. `reasoning_details` wins
because only it carries the block index, and OpenRouter repeats the same
text in both fields, so exactly one source is read per object. Entries
typed `reasoning.encrypted` carry opaque provider state rather than
readable text and are skipped. A change of block index closes the open
part so downstream summaries keep the provider's segmentation.
Co-Authored-By: Claude Opus 5 <[email protected]>
Expose the xAI Imagine image and video surfaces on top of the `xai`
provider, and make the shared OpenAI video-task layer survive the
production configuration they need.
Native video requests live under /v1 (generations, edits, extensions,
with /v1/videos as a creation alias that only selects xAI candidates);
the OpenAI-compatible adapter stays under /openai/v1/videos and maps
`seconds` / `size` onto numeric duration, aspect ratio and resolution.
Clients receive an opaque Aether task ID scoped to the owning user;
polling uses the upstream task ID and the original credential, and
completed downloads fetch the returned media URL without forwarding
provider authorization to the media host.
Three fixes to the shared video layer are required for this to work
outside tests:
- OpenAI/xAI task persistence now supplies a stable 16-character
short_id, which the PostgreSQL schema requires. Existing rows keep
their original value across reconstruction, so no schema change or
historical rewrite is needed.
- Task retrieval and content downloads are admitted by the production
GET execution gate, and reconstructed tasks resolve proxy nodes,
system proxy defaults, tunnel affinity and transport profiles through
the same deployment resolver used for creation. A configured proxy
route no longer silently becomes a direct request after restart.
- When the gateway also serves the frontend, /openai/v1/videos and its
subpaths bypass the static SPA handler. Otherwise a video query
returns HTTP 200 with text/html instead of the task JSON.
Co-Authored-By: Claude Opus 5 <[email protected]>
Add a separate `xai` provider type for xAI Grok CLI subscription accounts.
It is independent of the existing `grok` provider, which reverse-proxies
grok.com with browser cookies; behavior of `grok` is unchanged.
Account binding uses the xAI device code flow, so no local callback
listener is needed and headless deployments can bind accounts. Refresh
tokens can also be imported individually or in batches, and are rotated
on refresh.
OAuth requests default to the cli-chat-proxy Responses API; API keys and
compact stay on api.x.ai. Explicit custom gateways are preserved. Only
`openai:responses` and `openai:responses:compact` are exposed; Chat,
Claude and Gemini clients reach the provider through Aether's existing
cross-format conversion rather than new native endpoints.
Upstream Responses payloads are sanitized for what xAI actually rejects:
`previous_response_id` and `metadata.user_id` are dropped, hosted
`tool_choice` is rewritten, `web_search` is restored for converted
clients, `image_generation` is stripped on older Grok conversation
models, unsupported reasoning effort is removed, and requested
`reasoning.encrypted_content` is preserved with a replay policy keyed on
the configured provider type rather than the model name.
Quota refresh reads /user and /billing?format=credits and stores a
structured usage snapshot; a prepaid balance keeps an account selectable
after the weekly allowance is exhausted. API-key accounts skip the
subscription billing surface. The admin UI shows remaining weekly quota
as a labeled bar in the provider drawer and the pool list.
Co-Authored-By: Claude Opus 5 <[email protected]>
The gemini:generate_content URL hook rewrites any Antigravity endpoint to
/v1internal:generateContent, but only the same-format passthrough and the two
OpenAI decision paths ever built the matching envelope. A Claude Messages or
Gemini client therefore reached the standard family planner, picked up the
rewritten URL, and posted a bare Gemini body that upstream rejects with
"Invalid JSON payload received. Unknown name \"contents\"" -- four retries
across every account, then a 503 that names none of this.
Route the standard family through the shared v1internal builder the same way
gemini_cli already is, so the URL and the body come from one decision. The
OpenAI-image-to-Gemini path cannot carry an envelope at all, so it now skips
Antigravity candidates instead of sending a request upstream can only reject.
Also stop treating a configured proxy as locally unsupported. The execution
plan carries the proxy itself, and the generic and Vertex gates moved to
transport_proxy_is_locally_supported long ago; Antigravity kept rejecting on
proxy.is_some(), which no longer matches how the local runtime executes. A
proxy that resolves to no route still disqualifies the request, and transport
profiles stay unsupported because the v1internal payload never carries one.
Co-Authored-By: Claude Opus 5 <[email protected]>
A local stream attempt writes its `usage` row and its `request_candidates`
slot as `pending` in `execute_execution_runtime_stream_inner`, then awaits
the provider's response headers. Everything after that point runs inside
the downstream request future, so a client disconnect drops it: the
dispatch `.await` never resumes and nothing settles either row. The stream
finalizer that already covers this only exists once upstream headers have
arrived, so the pre-first-byte window has no owner at all. Both rows stay
`pending` until the maintenance sweeper rewrites them as a 504 timeout ten
minutes later, losing the real outcome, the real latency, and the 499.
`AttemptCancellationGuard` takes that window. It is created disarmed, so
an attempt dropped before it owns any row does not grow a settlement row
it never had; it is armed as soon as the attempt owns its non-terminal
rows, and the stream wrappers disarm it the moment the attempt returns,
from where settlement belongs to the transport. On a cancelling drop it
settles the candidate slot through the same snapshot writer the `pending`
write above it uses, and the usage row through a terminal `Cancelled`
event.
The guard outlives the request future, so what it captures is retained for
the whole attempt. It therefore holds no request body: the plan carries the
provider request body and the report context carries the client request
body, and keeping both would double the request-body residency of every
in-flight stream attempt to serve a path that almost never runs. Simply
omitting them is not safe either, because a terminal write is
body-capture-authoritative: with both absent the seed carries the typed
`none` marker, which clears the stored capture rather than leaving it
alone. `build_usage_event_data_seed_describing_request_bodies` is the third
option -- it derives every capture state, body reference, request type and
derived request fact from the real plan and report context, and leaves out
only the two body values -- so the guard's snapshot is small and its
terminal write preserves the capture the `pending` write recorded.
The stream candidate first-byte watchdog also drops the attempt future, but
it settles the attempt itself through `build_transport_error_stop_response`.
It now marks the attempt abandoned before returning so the guard stands down
instead of racing a 499 against the watchdog's 504.
Co-Authored-By: Claude Opus 5 <[email protected]>
`parse_codex_websocket_usage_limit_error` backfills `primary_reset_at` and
`primary_reset_after_seconds` from the error body whenever the embedded headers
did not supply them. Now that a named per-model limit no longer claims the
unprefixed window headers, that absence is exactly what a Spark 429 produces —
and `resets_at` on such an error is the Spark window's reset, so the backfill
put model-scoped timing back on the account's own quota.
Gate the backfill on the same ownership rule the window parsing uses.
Reported by Cursor Bugbot on the fork PR.
Co-Authored-By: Claude Opus 5 <[email protected]>
`parse_codex_usage_headers` reads the unprefixed `x-codex-primary/secondary-*`
headers as the account's own quota. They are not: they carry whichever limit
governed the request, and `x-codex-active-limit` names it — `premium` for the
plan's own limit, or a metered feature such as `codex_bengalfox` for a named
per-model limit. On a request billed against a named limit the unprefixed
headers repeat that limit's windows verbatim.
So a single request to a model with its own limit writes that model's windows
into the account slots. The paid-window swap then makes it worse: a named
limit's secondary window is active, unlike the plan's disabled one, so the swap
promotes the model's weekly window into the account's weekly slot — the slot
the UI labels and the scheduler reads through `quota_usage_ratio`.
It also sticks. Both weekly windows share `window_minutes`, so
`codex_quota_same_window_identity` treats them as one window, and
`codex_quota_merge_same_window` drops an observation whose deadline is earlier
than the stored one. The two weeks start at different instants, so every later
account observation looks like a stale sample of a window that already rolled
over and is discarded until the model window's own deadline passes.
Observed on a `pro` key running both model families: one `gpt-5.3-codex-spark`
request replaced the account weekly window with the Spark weekly one, and the
~3000 plan-limit responses over the next 100 minutes were all discarded. The
account's real weekly usage never landed, and its reset time was reported nine
hours late.
The header set describes itself — every named limit announces
`x-codex-<feature>-limit-name` and carries its windows under the same prefix —
so parse the named families directly and only claim the unprefixed windows for
the account when no announced limit owns them. Responses without
`x-codex-active-limit` keep the previous behaviour.
This also stops the Spark windows from going stale: they were only ever written
by the `wham/usage` admin probe even though every response carries them.
Refs #746
Co-Authored-By: Claude Opus 5 <[email protected]>
The response.completed fallback rebuilt every call item with responsesCallInput(), which returns '{}' for a function_call lacking arguments. Since '{}' is truthy, ensureToolCall overwrote arguments already collected from streamed delta events. Guard the completed branch with responsesCallHasInput (matching the output_item.done branch) so empty/default inputs no longer clobber streamed args, and align its dedupe key with the streaming phase to avoid duplicate tool-call rendering when an item has no id. Drop the now-dead '工具调用' fallbacks since responsesCallName never returns empty.
The stale-pending cleanup task previously hardcoded status_code=504 and a
generic timeout message for every usage row it finalized. When a request
had already been observed as failing — e.g. upstream Connection reset by
peer, watchdog 504, or an authenticated 4xx — the cleanup overwrote that
context with a misleading "服务器超时" outcome and 504 status, hiding the
real cause from the dashboards and customer.
Pull the most recent failed/cancelled candidate per stale request_id and,
if present, finalize the usage row with the candidate's status_code
(defaulting to 502 when none was recorded) and error_message. Requests
that have no terminal candidate (truly stuck pending/streaming) keep the
existing 504 + timeout-message behavior, since they really are timeouts
from the cleanup's perspective. Applied to all three SQL backends with
parameterized UPDATE statements.
The Postgres failed-candidate lookup orders by
COALESCE(finished_at, started_at, created_at) DESC, matching the MySQL
and SQLite ORDER BY clauses so the three backends pick the same
"most recent terminal candidate" under every NULL combination of timing
columns.
Co-Authored-By: Claude Opus 4.7 <[email protected]>
When the endpoint forces upstream_stream_policy=force_non_stream while
the client streams, the local stream candidate watchdog still preferred
timeouts.first_byte_ms — a non-stream upstream produces no early first
byte, so the watchdog fired before the HTTP request_timeout and aborted
otherwise-healthy attempts at ~300s.
Read upstream_is_stream from report_context and invert the priority:
non-stream upstreams use total_ms first, falling back to first_byte_ms
and then the default; streaming upstreams keep the previous order.
Co-Authored-By: Claude Opus 4.7 <[email protected]>
The "upstream_is_stream" JSON key flows from the AI execution report
context producer (aether-ai-serving::report_context) through several
consumers — usage runtime metadata copy/move, gateway watchdog, sync
execution decision, observability handlers, and the per-driver usage
repositories. Each site spelled the key as a bare string literal, so a
producer-side rename would silently degrade every consumer to its
fallback (typically assuming streaming) with no compile-time signal.
Introduce a single pub const UPSTREAM_IS_STREAM_KEY in
aether-ai-formats (the lowest crate every consumer already depends on),
re-export from the crate root, and route producer + all map-style
consumers through it. The change is purely a string-literal → constant
swap; behaviour is identical.
Sites left as literals (intentional):
- `json!({"upstream_is_stream": ...})` macro keys, which must be string
literals at the macro layer; these are also API-response payload
field names (an external contract that should not silently track
internal report-context renames).
- SQL column accessors (`try_get::<...>("upstream_is_stream")`), which
refer to the database schema column, not the JSON key.
- Test fixtures and assertions, which validate the on-the-wire contract
and should keep verifying the actual string.
Co-Authored-By: Claude Opus 4.7 <[email protected]>