Raw chain-of-thought was written to both `content` (`reasoning_text`) and
`summary` (`summary_text`), and the stream emitter sent the same delta on
`response.reasoning_text.delta` *and* `response.reasoning_summary_text.delta`.
Clients that render both channels therefore printed every thinking chunk
twice — most visibly the Codex CLI, whose thinking panel repeated itself.
OpenAI keeps the two channels distinct: `content` carries the raw CoT while
`summary` is the summarised view. Emit the thinking on `content` only:
- `openai_responses_reasoning_text_fields` becomes
`openai_responses_reasoning_text_parts`, returning just the `content` array;
reasoning items keep `summary: []` (or a provider-supplied summary).
- The Responses stream emitter emits `response.reasoning_text.delta` /
`.done` and no longer mirrors them onto the summary events. The reasoning
`output_item.added` no longer announces a `reasoning_summary_part`.
- The provider-state reasoning reader accepts `content` (`reasoning_text`)
first and falls back to `summary`, so it also understands items produced by
older Aether versions; its state field is renamed accordingly.
- Non-streaming builders (Chat -> Responses, manual Responses response, Grok
gateway) place the thinking on `content` and leave `summary` empty.
Tests cover the raw thinking appearing exactly once in the emitted stream.
Move Antigravity schema-preservation policy into the format crate and expose provider-aware Chat and Responses builders through the existing gateway root seam. Preserve legacy conversion and scoped Responses history behavior without weakening architecture tests.
Validated: 61 standalone architecture tests, 947 format tests, 527 transport tests, 2 actual-source planner tests, and gateway cargo check.
Preserve tool schemas across provider-scoped format conversion and OpenAI gateway planners until the Antigravity boundary. Bound reference expansion, safely merge schema constraints, and lower unsupported Claude unions.
Drop unsigned historical Claude thinking without changing Gemini behavior. Add cross-format regression fixtures and retain the project's original error output policy.
Validation: 945 format tests and 527 transport tests passed; gateway cargo check passed.
Gemini runs `googleSearch` inside Google. The search leaves no
client-visible tool call, and the evidence arrives only as
`candidates[].groundingMetadata`. Every cross-format target dropped it
wholesale, so a grounded answer reached OpenAI- and Claude-shaped clients
as prose that names its sources with nothing structured behind it: no
`annotations`, no `citations`, no `url_citation`. Callers that verify
grounding — the common "did this model actually search?" check — saw a
200 with no evidence and had to treat the answer as ungrounded.
Adapters now normalise `groundingMetadata` into neutral citations and
each target renders its own family's standard shape: `url_citation`
annotations for `openai:chat` and `openai:responses`, and
`web_search_result_location` citations on the text block for
`claude:messages`. Gemini reports segment bounds as UTF-8 byte offsets
while both targets count characters, so the bounds are converted rather
than copied.
Streaming is covered too, since that is what grounded traffic actually
uses. A new `CanonicalStreamEvent::Citations` carries the neutral list
once the answer text is whole — the offsets index into the finished
answer, so it rides just ahead of `Finish` rather than as a delta per
chunk — and each client emitter renders it: `delta.annotations` chunks,
`response.output_text.annotation.added` events (also kept on the finished
message item so clients that only read `response.completed` see them),
and `citations_delta` content block deltas.
For reference, CLIProxyAPI projects grounding only in its
antigravity→Claude translator, and only when the client declared a typed
`web_search_*` tool; its OpenAI and plain Gemini translators have no
grounding handling at all. The citation shape here matches theirs, but
the coverage is deliberately wider: all three targets, streaming and
non-streaming, with no dependency on a declared tool.
Co-Authored-By: Claude Opus 5 <[email protected]>
The transport boundary rewrote `googleSearch` into the Gemini 1.5-era
`googleSearchRetrieval` spelling before every v1internal call, on the stated
grounds that the private backend rejects `googleSearch` when it is combined
with function declarations. That rewrite breaks grounding on Gemini 3.
Observed on stabey-124 against daily-cloudcode-pa.googleapis.com. A controlled
pair, same model and keys, 5 seconds apart:
- no `web_search_options` -> 200
- with `web_search_options` -> 502 on all three candidates
The outgoing body carried `tools: [{"googleSearchRetrieval": {}}]` and no
function declarations at all, so the documented mixed-tool rationale did not
apply. `request_candidates.error_message` holds what the backend actually
said:
Malformed function call: call:google_search{query:current UTC date time}
Malformed function call: call:google:search{query:current UTC date}
Malformed function call: call:google_search{queries:[current UTC date]}
The model reaches for `google_search`, the legacy declaration binds nothing,
and the turn dies unparsed. CLIProxyAPI sends `googleSearch` to this same
v1internal surface, including alongside function declarations.
Keep folding the snake_case `google_search` alias into the canonical
`googleSearch` key, and leave a request that already spells the tool
`googleSearchRetrieval` untouched.
Co-Authored-By: Claude Opus 5 <[email protected]>
canonical_tools_to_gemini promoted any tool whose name normalized to
"websearch" / "googlesearch" / "websearchpreview" into Gemini's server-side
builtin, dropping it from functionDeclarations. Claude Code declares an
ordinary client-side `WebSearch` tool with a full input_schema, so every
/v1/messages request routed to a Gemini model lost that declaration and gained
`googleSearch` (rewritten to `googleSearchRetrieval` at the Anti Gravity
transport boundary) instead.
Two consequences, both observed on stabey-124 against gemini-3.8-flash:
- the model can never emit a `WebSearch` tool_use, so the client's own web
search is dead on that route;
- when the model does reach for the injected server-side search, the v1internal
backend answers `finishReason: MALFORMED_FUNCTION_CALL` /
"Function call is empty - no input to parse." and the turn fails.
Promote a tool to a builtin only when it is a bare marker carrying no schema.
A declared schema means the caller intends to execute the call itself, which
matches CLIProxyAPI: it keys builtins off Claude's `type: web_search_*` or an
explicit `google_search` tool key and never off a function name.
Co-Authored-By: Claude Opus 5 <[email protected]>
OpenRouter reports reasoning under `reasoning` and `reasoning_details`
rather than the DeepSeek-style `reasoning_content` this crate recognized.
Its streaming reasoning phase sends chunks whose `delta.content` is an
empty string, so those chunks were dropped and Responses clients saw
nothing after `response.in_progress` until they timed the stream out.
The sync aggregator kept only content and tool calls, so a stream
downgraded to a sync response lost the reasoning entirely.
Read all three spellings through one helper. `reasoning_details` wins
because only it carries the block index, and OpenRouter repeats the same
text in both fields, so exactly one source is read per object. Entries
typed `reasoning.encrypted` carry opaque provider state rather than
readable text and are skipped. A change of block index closes the open
part so downstream summaries keep the provider's segmentation.
Co-Authored-By: Claude Opus 5 <[email protected]>
Expose the xAI Imagine image and video surfaces on top of the `xai`
provider, and make the shared OpenAI video-task layer survive the
production configuration they need.
Native video requests live under /v1 (generations, edits, extensions,
with /v1/videos as a creation alias that only selects xAI candidates);
the OpenAI-compatible adapter stays under /openai/v1/videos and maps
`seconds` / `size` onto numeric duration, aspect ratio and resolution.
Clients receive an opaque Aether task ID scoped to the owning user;
polling uses the upstream task ID and the original credential, and
completed downloads fetch the returned media URL without forwarding
provider authorization to the media host.
Three fixes to the shared video layer are required for this to work
outside tests:
- OpenAI/xAI task persistence now supplies a stable 16-character
short_id, which the PostgreSQL schema requires. Existing rows keep
their original value across reconstruction, so no schema change or
historical rewrite is needed.
- Task retrieval and content downloads are admitted by the production
GET execution gate, and reconstructed tasks resolve proxy nodes,
system proxy defaults, tunnel affinity and transport profiles through
the same deployment resolver used for creation. A configured proxy
route no longer silently becomes a direct request after restart.
- When the gateway also serves the frontend, /openai/v1/videos and its
subpaths bypass the static SPA handler. Otherwise a video query
returns HTTP 200 with text/html instead of the task JSON.
Co-Authored-By: Claude Opus 5 <[email protected]>
Add a separate `xai` provider type for xAI Grok CLI subscription accounts.
It is independent of the existing `grok` provider, which reverse-proxies
grok.com with browser cookies; behavior of `grok` is unchanged.
Account binding uses the xAI device code flow, so no local callback
listener is needed and headless deployments can bind accounts. Refresh
tokens can also be imported individually or in batches, and are rotated
on refresh.
OAuth requests default to the cli-chat-proxy Responses API; API keys and
compact stay on api.x.ai. Explicit custom gateways are preserved. Only
`openai:responses` and `openai:responses:compact` are exposed; Chat,
Claude and Gemini clients reach the provider through Aether's existing
cross-format conversion rather than new native endpoints.
Upstream Responses payloads are sanitized for what xAI actually rejects:
`previous_response_id` and `metadata.user_id` are dropped, hosted
`tool_choice` is rewritten, `web_search` is restored for converted
clients, `image_generation` is stripped on older Grok conversation
models, unsupported reasoning effort is removed, and requested
`reasoning.encrypted_content` is preserved with a replay policy keyed on
the configured provider type rather than the model name.
Quota refresh reads /user and /billing?format=credits and stores a
structured usage snapshot; a prepaid balance keeps an account selectable
after the weekly allowance is exhausted. API-key accounts skip the
subscription billing surface. The admin UI shows remaining weekly quota
as a labeled bar in the provider drawer and the pool list.
Co-Authored-By: Claude Opus 5 <[email protected]>
- delete pool member scores when a provider key is deactivated
- fall through to the catalog key scan when the pool score page is empty
- log skipped local candidates with a skip_reason for diagnosis
- add regressions for inactive keys with stale available scores
Show already checked / already associated models first in the 关联模型
list so they are easy to find and uncheck. Keep that order inside the
current search results.
Let admins select one or more models in a provider's model list and
delete them together, using the existing single-model delete API and
the same confirm-danger pattern as global model batch delete.
OpenAI Responses treats reasoning.content as the raw chain-of-thought
and summary as a skim view. Aether was dumping thinking into summary
and leaving content null, which hid the thinking panel in desktop UIs.
Put reasoning_content / equivalent text into reasoning_text content
parts, and copy the same text into summary_text so CLI clients still
work. Stream emitters now send both reasoning_text and summary events.