The partial-preface test raced a 5ms first-request deadline against a
10ms hyper header_read_timeout. On a slow runner both timers expire
before the next poll and tokio::select! may pick the connection branch,
surfacing hyper's header-timeout error instead of the clean deadline
close. Push hyper's timeout out to 30s so only the deadline can fire.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Verified against a live Antigravity + xAI deployment:
- Gemini and Claude clients calling a forced-stream Responses upstream
(xAI, Codex) without streaming always failed: the aggregated body echoes
request metadata (parallel_tool_calls, tools, encrypted reasoning) that
the strict cross-format check refuses, and the gateway then wrapped the
raw SSE capture in a client error body sent with HTTP 200. Project the
validated aggregate to every client format, as the Chat path already
does, and return 502 instead of raw provider bytes when a successful
cross-format response still cannot be converted.
- Gemini stream decoding keyed tool calls by part position, so parallel
calls arriving in separate chunks (all at parts[0]) merged into one call
with concatenated arguments. Key them by arrival order; ids cannot be
used because they are optional and the Antigravity envelope synthesizes
per-chunk ids that repeat across chunks. Generated call_auto_N ids now
follow arrival order.
- Non-stream Responses output reported truncated or filtered cross-format
answers as completed; derive incomplete + incomplete_details from the
canonical stop reason.
- Gemini request parsing ignored parametersJsonSchema and
responseJsonSchema and passed OpenAPI upper-case type names (OBJECT,
STRING) through to JSON Schema targets, which xAI rejects.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A provider model can be addressed by its upstream name or by any of its
`provider_model_mappings` entries, and neither name is published in the
model catalog, which lists global model names only. Resolution ran per API
format, so an alias could win a format the real global model had no
provider in: a `claude:messages` client asking for `gemini-3.8-flash`
landed on the provider that merely renames its own `gemini-3.8-flash-cursor`
model to `gemini-3.8-flash` on the way upstream, and the separate global
model stopped distinguishing the two routes.
Treat global model names as a reserved namespace instead: when the request
names an active global model, only rows bound to it may serve it, whatever
API format they sit in. Rows in hand answer that question for free whenever
one of them is bound to a global model of that name, so the lookup stays off
the path ordinary requests take. Authorization resolves the same way, so an
API key's allowed models cannot be satisfied through a resolution candidate
planning will no longer make.
A request naming a model that is not a global model keeps every matching
rule, so addressing a provider variant by its upstream name still works.
Co-Authored-By: Claude Opus 5 <[email protected]>
Merge upstream PR #837, preserving its original commits and resolving the UsageModelDisplay layout conflict after #841. Cover response models alongside dynamic service-tier badges in detail tests.