The partial-preface test raced a 5ms first-request deadline against a
10ms hyper header_read_timeout. On a slow runner both timers expire
before the next poll and tokio::select! may pick the connection branch,
surfacing hyper's header-timeout error instead of the clean deadline
close. Push hyper's timeout out to 30s so only the deadline can fire.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Verified against a live Antigravity + xAI deployment:
- Gemini and Claude clients calling a forced-stream Responses upstream
(xAI, Codex) without streaming always failed: the aggregated body echoes
request metadata (parallel_tool_calls, tools, encrypted reasoning) that
the strict cross-format check refuses, and the gateway then wrapped the
raw SSE capture in a client error body sent with HTTP 200. Project the
validated aggregate to every client format, as the Chat path already
does, and return 502 instead of raw provider bytes when a successful
cross-format response still cannot be converted.
- Gemini stream decoding keyed tool calls by part position, so parallel
calls arriving in separate chunks (all at parts[0]) merged into one call
with concatenated arguments. Key them by arrival order; ids cannot be
used because they are optional and the Antigravity envelope synthesizes
per-chunk ids that repeat across chunks. Generated call_auto_N ids now
follow arrival order.
- Non-stream Responses output reported truncated or filtered cross-format
answers as completed; derive incomplete + incomplete_details from the
canonical stop reason.
- Gemini request parsing ignored parametersJsonSchema and
responseJsonSchema and passed OpenAPI upper-case type names (OBJECT,
STRING) through to JSON Schema targets, which xAI rejects.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Merge upstream PR #837, preserving its original commits and resolving the UsageModelDisplay layout conflict after #841. Cover response models alongside dynamic service-tier badges in detail tests.
Raw chain-of-thought was written to both `content` (`reasoning_text`) and
`summary` (`summary_text`), and the stream emitter sent the same delta on
`response.reasoning_text.delta` *and* `response.reasoning_summary_text.delta`.
Clients that render both channels therefore printed every thinking chunk
twice — most visibly the Codex CLI, whose thinking panel repeated itself.
OpenAI keeps the two channels distinct: `content` carries the raw CoT while
`summary` is the summarised view. Emit the thinking on `content` only:
- `openai_responses_reasoning_text_fields` becomes
`openai_responses_reasoning_text_parts`, returning just the `content` array;
reasoning items keep `summary: []` (or a provider-supplied summary).
- The Responses stream emitter emits `response.reasoning_text.delta` /
`.done` and no longer mirrors them onto the summary events. The reasoning
`output_item.added` no longer announces a `reasoning_summary_part`.
- The provider-state reasoning reader accepts `content` (`reasoning_text`)
first and falls back to `summary`, so it also understands items produced by
older Aether versions; its state field is renamed accordingly.
- Non-streaming builders (Chat -> Responses, manual Responses response, Grok
gateway) place the thinking on `content` and leave `summary` empty.
Tests cover the raw thinking appearing exactly once in the emitted stream.
Move Antigravity schema-preservation policy into the format crate and expose provider-aware Chat and Responses builders through the existing gateway root seam. Preserve legacy conversion and scoped Responses history behavior without weakening architecture tests.
Validated: 61 standalone architecture tests, 947 format tests, 527 transport tests, 2 actual-source planner tests, and gateway cargo check.
Preserve tool schemas across provider-scoped format conversion and OpenAI gateway planners until the Antigravity boundary. Bound reference expansion, safely merge schema constraints, and lower unsupported Claude unions.
Drop unsigned historical Claude thinking without changing Gemini behavior. Add cross-format regression fixtures and retain the project's original error output policy.
Validation: 945 format tests and 527 transport tests passed; gateway cargo check passed.
OpenRouter reports reasoning under `reasoning` and `reasoning_details`
rather than the DeepSeek-style `reasoning_content` this crate recognized.
Its streaming reasoning phase sends chunks whose `delta.content` is an
empty string, so those chunks were dropped and Responses clients saw
nothing after `response.in_progress` until they timed the stream out.
The sync aggregator kept only content and tool calls, so a stream
downgraded to a sync response lost the reasoning entirely.
Read all three spellings through one helper. `reasoning_details` wins
because only it carries the block index, and OpenRouter repeats the same
text in both fields, so exactly one source is read per object. Entries
typed `reasoning.encrypted` carry opaque provider state rather than
readable text and are skipped. A change of block index closes the open
part so downstream summaries keep the provider's segmentation.
Co-Authored-By: Claude Opus 5 <[email protected]>
Expose the xAI Imagine image and video surfaces on top of the `xai`
provider, and make the shared OpenAI video-task layer survive the
production configuration they need.
Native video requests live under /v1 (generations, edits, extensions,
with /v1/videos as a creation alias that only selects xAI candidates);
the OpenAI-compatible adapter stays under /openai/v1/videos and maps
`seconds` / `size` onto numeric duration, aspect ratio and resolution.
Clients receive an opaque Aether task ID scoped to the owning user;
polling uses the upstream task ID and the original credential, and
completed downloads fetch the returned media URL without forwarding
provider authorization to the media host.
Three fixes to the shared video layer are required for this to work
outside tests:
- OpenAI/xAI task persistence now supplies a stable 16-character
short_id, which the PostgreSQL schema requires. Existing rows keep
their original value across reconstruction, so no schema change or
historical rewrite is needed.
- Task retrieval and content downloads are admitted by the production
GET execution gate, and reconstructed tasks resolve proxy nodes,
system proxy defaults, tunnel affinity and transport profiles through
the same deployment resolver used for creation. A configured proxy
route no longer silently becomes a direct request after restart.
- When the gateway also serves the frontend, /openai/v1/videos and its
subpaths bypass the static SPA handler. Otherwise a video query
returns HTTP 200 with text/html instead of the task JSON.
Co-Authored-By: Claude Opus 5 <[email protected]>