cli-chat-proxy.grok.com started rejecting every request on 2026-10-01
with HTTP 426 "Your Grok CLI version (0.2.120) is outdated. Please
update to version 1.0.13 or later", because x-grok-client-version and
the xai-grok-workspace user agent were pinned to 0.2.120.
Replace the pin with a runtime-published version (built-in fallback
1.0.46) and add a gateway worker, modelled on the Codex profile worker,
that prewarms at startup and refreshes every 3h:
- read the official stable channel https://x.ai/cli/stable, falling
back to npm @xai-official/grok/latest (deployments that cannot reach
x.ai directly), requiring all six platform binaries at one version;
- never roll back, persist the verified version in runtime KV and
restore it on restart;
- AETHER_XAI_CLIENT_VERSION pins a version, and
AETHER_XAI_CLIENT_PROFILE_REFRESH=off disables the network check.
Endpoint header rules still win over the injected identity headers.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
Responses output turned every tool call named `web_search` or
`web_search_preview` into a hosted `web_search_call`, regardless of what the
client declared. OMP declares its own `{"type":"function","name":"web_search"}`
tool, so when gemini-3.8-flash called it the client got a hosted item it
cannot execute. OMP then echoed that `web_search_call` back as the last input
item with no output, the Gemini request body could not be built, and every
retry failed with 503 "上游请求体转换失败" (provider_request_body_build_failed).
Observed on stabey-124 on 2026-09-29 (request c6270ef6 and five retries after
b794e7ca returned functionCall web_search / call_109312).
Decide the hosted mapping in one place, NamespaceToolAliases::
emits_hosted_web_search_call, used by both the sync builder and the stream
emitter: emit `web_search_call` only when the name is not a namespaced child
and the client did not declare a function or custom tool of that name. When
both a hosted tool and a function share the name, the function wins.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
Verified against a live Antigravity + xAI deployment:
- Gemini and Claude clients calling a forced-stream Responses upstream
(xAI, Codex) without streaming always failed: the aggregated body echoes
request metadata (parallel_tool_calls, tools, encrypted reasoning) that
the strict cross-format check refuses, and the gateway then wrapped the
raw SSE capture in a client error body sent with HTTP 200. Project the
validated aggregate to every client format, as the Chat path already
does, and return 502 instead of raw provider bytes when a successful
cross-format response still cannot be converted.
- Gemini stream decoding keyed tool calls by part position, so parallel
calls arriving in separate chunks (all at parts[0]) merged into one call
with concatenated arguments. Key them by arrival order; ids cannot be
used because they are optional and the Antigravity envelope synthesizes
per-chunk ids that repeat across chunks. Generated call_auto_N ids now
follow arrival order.
- Non-stream Responses output reported truncated or filtered cross-format
answers as completed; derive incomplete + incomplete_details from the
canonical stop reason.
- Gemini request parsing ignored parametersJsonSchema and
responseJsonSchema and passed OpenAPI upper-case type names (OBJECT,
STRING) through to JSON Schema targets, which xAI rejects.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Merge upstream PR #837, preserving its original commits and resolving the UsageModelDisplay layout conflict after #841. Cover response models alongside dynamic service-tier badges in detail tests.
Raw chain-of-thought was written to both `content` (`reasoning_text`) and
`summary` (`summary_text`), and the stream emitter sent the same delta on
`response.reasoning_text.delta` *and* `response.reasoning_summary_text.delta`.
Clients that render both channels therefore printed every thinking chunk
twice — most visibly the Codex CLI, whose thinking panel repeated itself.
OpenAI keeps the two channels distinct: `content` carries the raw CoT while
`summary` is the summarised view. Emit the thinking on `content` only:
- `openai_responses_reasoning_text_fields` becomes
`openai_responses_reasoning_text_parts`, returning just the `content` array;
reasoning items keep `summary: []` (or a provider-supplied summary).
- The Responses stream emitter emits `response.reasoning_text.delta` /
`.done` and no longer mirrors them onto the summary events. The reasoning
`output_item.added` no longer announces a `reasoning_summary_part`.
- The provider-state reasoning reader accepts `content` (`reasoning_text`)
first and falls back to `summary`, so it also understands items produced by
older Aether versions; its state field is renamed accordingly.
- Non-streaming builders (Chat -> Responses, manual Responses response, Grok
gateway) place the thinking on `content` and leave `summary` empty.
Tests cover the raw thinking appearing exactly once in the emitted stream.
Usage records show a reasoning badge next to the model name for OpenAI and
Claude requests, but Gemini requests never got one. The extraction only read
the OpenAI/Claude shapes (`reasoning_effort`, `reasoning.effort`,
`output_config.effort`), while Gemini states its reasoning depth inside
`generationConfig.thinkingConfig` — so nothing was written to the usage
metadata and the list and detail views had no badge to render.
Read the Gemini shape too, as a fallback after the existing three so the
OpenAI and Claude paths are untouched:
- `thinkingLevel` / `thinking_level` wins when present, trimmed and
lowercased, with the protobuf enum prefix stripped so `THINKING_LEVEL_HIGH`
resolves like `high`.
- Otherwise `thinkingBudget` / `thinking_budget` goes through the existing
shared budget ladder, yielding the same `low|medium|high|xhigh` vocabulary
the badge already understands.
- Both camelCase and snake_case spellings are read, so a captured client body
and a converted provider body resolve to the same label.
- `includeThoughts` alone is a visibility flag, not a depth, and produces no
badge.
Two cases are handled explicitly rather than through the shared ladder:
- `thinkingBudget: 0` disables reasoning outright. The shared ladder maps
`0..=1664` to `low`, which would report an explicitly disabled request as a
shallow one, so it reports `none` instead.
- `THINKING_LEVEL_UNSPECIFIED` is the enum's "no explicit level" member, not a
depth; it is rejected rather than surfaced as an `unspecified` badge.
The frontend needs no change: `UsageModelDisplay` already renders the badge
whenever the fields are present, and keeps the `high -> xhigh` mapping format
when the requested and upstream efforts differ.