cli-chat-proxy.grok.com started rejecting every request on 2026-10-01
with HTTP 426 "Your Grok CLI version (0.2.120) is outdated. Please
update to version 1.0.13 or later", because x-grok-client-version and
the xai-grok-workspace user agent were pinned to 0.2.120.
Replace the pin with a runtime-published version (built-in fallback
1.0.46) and add a gateway worker, modelled on the Codex profile worker,
that prewarms at startup and refreshes every 3h:
- read the official stable channel https://x.ai/cli/stable, falling
back to npm @xai-official/grok/latest (deployments that cannot reach
x.ai directly), requiring all six platform binaries at one version;
- never roll back, persist the verified version in runtime KV and
restore it on restart;
- AETHER_XAI_CLIENT_VERSION pins a version, and
AETHER_XAI_CLIENT_PROFILE_REFRESH=off disables the network check.
Endpoint header rules still win over the injected identity headers.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
Verified against a live Antigravity + xAI deployment:
- Gemini and Claude clients calling a forced-stream Responses upstream
(xAI, Codex) without streaming always failed: the aggregated body echoes
request metadata (parallel_tool_calls, tools, encrypted reasoning) that
the strict cross-format check refuses, and the gateway then wrapped the
raw SSE capture in a client error body sent with HTTP 200. Project the
validated aggregate to every client format, as the Chat path already
does, and return 502 instead of raw provider bytes when a successful
cross-format response still cannot be converted.
- Gemini stream decoding keyed tool calls by part position, so parallel
calls arriving in separate chunks (all at parts[0]) merged into one call
with concatenated arguments. Key them by arrival order; ids cannot be
used because they are optional and the Antigravity envelope synthesizes
per-chunk ids that repeat across chunks. Generated call_auto_N ids now
follow arrival order.
- Non-stream Responses output reported truncated or filtered cross-format
answers as completed; derive incomplete + incomplete_details from the
canonical stop reason.
- Gemini request parsing ignored parametersJsonSchema and
responseJsonSchema and passed OpenAPI upper-case type names (OBJECT,
STRING) through to JSON Schema targets, which xAI rejects.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Merge upstream PR #837, preserving its original commits and resolving the UsageModelDisplay layout conflict after #841. Cover response models alongside dynamic service-tier badges in detail tests.
Raw chain-of-thought was written to both `content` (`reasoning_text`) and
`summary` (`summary_text`), and the stream emitter sent the same delta on
`response.reasoning_text.delta` *and* `response.reasoning_summary_text.delta`.
Clients that render both channels therefore printed every thinking chunk
twice — most visibly the Codex CLI, whose thinking panel repeated itself.
OpenAI keeps the two channels distinct: `content` carries the raw CoT while
`summary` is the summarised view. Emit the thinking on `content` only:
- `openai_responses_reasoning_text_fields` becomes
`openai_responses_reasoning_text_parts`, returning just the `content` array;
reasoning items keep `summary: []` (or a provider-supplied summary).
- The Responses stream emitter emits `response.reasoning_text.delta` /
`.done` and no longer mirrors them onto the summary events. The reasoning
`output_item.added` no longer announces a `reasoning_summary_part`.
- The provider-state reasoning reader accepts `content` (`reasoning_text`)
first and falls back to `summary`, so it also understands items produced by
older Aether versions; its state field is renamed accordingly.
- Non-streaming builders (Chat -> Responses, manual Responses response, Grok
gateway) place the thinking on `content` and leave `summary` empty.
Tests cover the raw thinking appearing exactly once in the emitted stream.
Usage records show a reasoning badge next to the model name for OpenAI and
Claude requests, but Gemini requests never got one. The extraction only read
the OpenAI/Claude shapes (`reasoning_effort`, `reasoning.effort`,
`output_config.effort`), while Gemini states its reasoning depth inside
`generationConfig.thinkingConfig` — so nothing was written to the usage
metadata and the list and detail views had no badge to render.
Read the Gemini shape too, as a fallback after the existing three so the
OpenAI and Claude paths are untouched:
- `thinkingLevel` / `thinking_level` wins when present, trimmed and
lowercased, with the protobuf enum prefix stripped so `THINKING_LEVEL_HIGH`
resolves like `high`.
- Otherwise `thinkingBudget` / `thinking_budget` goes through the existing
shared budget ladder, yielding the same `low|medium|high|xhigh` vocabulary
the badge already understands.
- Both camelCase and snake_case spellings are read, so a captured client body
and a converted provider body resolve to the same label.
- `includeThoughts` alone is a visibility flag, not a depth, and produces no
badge.
Two cases are handled explicitly rather than through the shared ladder:
- `thinkingBudget: 0` disables reasoning outright. The shared ladder maps
`0..=1664` to `low`, which would report an explicitly disabled request as a
shallow one, so it reports `none` instead.
- `THINKING_LEVEL_UNSPECIFIED` is the enum's "no explicit level" member, not a
depth; it is rejected rather than surfaced as an `unspecified` badge.
The frontend needs no change: `UsageModelDisplay` already renders the badge
whenever the fields are present, and keeps the `high -> xhigh` mapping format
when the requested and upstream efforts differ.
Move Antigravity schema-preservation policy into the format crate and expose provider-aware Chat and Responses builders through the existing gateway root seam. Preserve legacy conversion and scoped Responses history behavior without weakening architecture tests.
Validated: 61 standalone architecture tests, 947 format tests, 527 transport tests, 2 actual-source planner tests, and gateway cargo check.
Preserve tool schemas across provider-scoped format conversion and OpenAI gateway planners until the Antigravity boundary. Bound reference expansion, safely merge schema constraints, and lower unsupported Claude unions.
Drop unsigned historical Claude thinking without changing Gemini behavior. Add cross-format regression fixtures and retain the project's original error output policy.
Validation: 945 format tests and 527 transport tests passed; gateway cargo check passed.
Gemini runs `googleSearch` inside Google. The search leaves no
client-visible tool call, and the evidence arrives only as
`candidates[].groundingMetadata`. Every cross-format target dropped it
wholesale, so a grounded answer reached OpenAI- and Claude-shaped clients
as prose that names its sources with nothing structured behind it: no
`annotations`, no `citations`, no `url_citation`. Callers that verify
grounding — the common "did this model actually search?" check — saw a
200 with no evidence and had to treat the answer as ungrounded.
Adapters now normalise `groundingMetadata` into neutral citations and
each target renders its own family's standard shape: `url_citation`
annotations for `openai:chat` and `openai:responses`, and
`web_search_result_location` citations on the text block for
`claude:messages`. Gemini reports segment bounds as UTF-8 byte offsets
while both targets count characters, so the bounds are converted rather
than copied.
Streaming is covered too, since that is what grounded traffic actually
uses. A new `CanonicalStreamEvent::Citations` carries the neutral list
once the answer text is whole — the offsets index into the finished
answer, so it rides just ahead of `Finish` rather than as a delta per
chunk — and each client emitter renders it: `delta.annotations` chunks,
`response.output_text.annotation.added` events (also kept on the finished
message item so clients that only read `response.completed` see them),
and `citations_delta` content block deltas.
For reference, CLIProxyAPI projects grounding only in its
antigravity→Claude translator, and only when the client declared a typed
`web_search_*` tool; its OpenAI and plain Gemini translators have no
grounding handling at all. The citation shape here matches theirs, but
the coverage is deliberately wider: all three targets, streaming and
non-streaming, with no dependency on a declared tool.
Co-Authored-By: Claude Opus 5 <[email protected]>
The transport boundary rewrote `googleSearch` into the Gemini 1.5-era
`googleSearchRetrieval` spelling before every v1internal call, on the stated
grounds that the private backend rejects `googleSearch` when it is combined
with function declarations. That rewrite breaks grounding on Gemini 3.
Observed on stabey-124 against daily-cloudcode-pa.googleapis.com. A controlled
pair, same model and keys, 5 seconds apart:
- no `web_search_options` -> 200
- with `web_search_options` -> 502 on all three candidates
The outgoing body carried `tools: [{"googleSearchRetrieval": {}}]` and no
function declarations at all, so the documented mixed-tool rationale did not
apply. `request_candidates.error_message` holds what the backend actually
said:
Malformed function call: call:google_search{query:current UTC date time}
Malformed function call: call:google:search{query:current UTC date}
Malformed function call: call:google_search{queries:[current UTC date]}
The model reaches for `google_search`, the legacy declaration binds nothing,
and the turn dies unparsed. CLIProxyAPI sends `googleSearch` to this same
v1internal surface, including alongside function declarations.
Keep folding the snake_case `google_search` alias into the canonical
`googleSearch` key, and leave a request that already spells the tool
`googleSearchRetrieval` untouched.
Co-Authored-By: Claude Opus 5 <[email protected]>