cli-chat-proxy.grok.com started rejecting every request on 2026-10-01
with HTTP 426 "Your Grok CLI version (0.2.120) is outdated. Please
update to version 1.0.13 or later", because x-grok-client-version and
the xai-grok-workspace user agent were pinned to 0.2.120.
Replace the pin with a runtime-published version (built-in fallback
1.0.46) and add a gateway worker, modelled on the Codex profile worker,
that prewarms at startup and refreshes every 3h:
- read the official stable channel https://x.ai/cli/stable, falling
back to npm @xai-official/grok/latest (deployments that cannot reach
x.ai directly), requiring all six platform binaries at one version;
- never roll back, persist the verified version in runtime KV and
restore it on restart;
- AETHER_XAI_CLIENT_VERSION pins a version, and
AETHER_XAI_CLIENT_PROFILE_REFRESH=off disables the network check.
Endpoint header rules still win over the injected identity headers.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
Add a separate `xai` provider type for xAI Grok CLI subscription accounts.
It is independent of the existing `grok` provider, which reverse-proxies
grok.com with browser cookies; behavior of `grok` is unchanged.
Account binding uses the xAI device code flow, so no local callback
listener is needed and headless deployments can bind accounts. Refresh
tokens can also be imported individually or in batches, and are rotated
on refresh.
OAuth requests default to the cli-chat-proxy Responses API; API keys and
compact stay on api.x.ai. Explicit custom gateways are preserved. Only
`openai:responses` and `openai:responses:compact` are exposed; Chat,
Claude and Gemini clients reach the provider through Aether's existing
cross-format conversion rather than new native endpoints.
Upstream Responses payloads are sanitized for what xAI actually rejects:
`previous_response_id` and `metadata.user_id` are dropped, hosted
`tool_choice` is rewritten, `web_search` is restored for converted
clients, `image_generation` is stripped on older Grok conversation
models, unsupported reasoning effort is removed, and requested
`reasoning.encrypted_content` is preserved with a replay policy keyed on
the configured provider type rather than the model name.
Quota refresh reads /user and /billing?format=credits and stores a
structured usage snapshot; a prepaid balance keeps an account selectable
after the weekly allowance is exhausted. API-key accounts skip the
subscription billing surface. The admin UI shows remaining weekly quota
as a labeled bar in the provider drawer and the pool list.
Co-Authored-By: Claude Opus 5 <[email protected]>
Preserve exact request payloads and model client surface and API operation explicitly.
Add Anthropic compatibility profiles, bounded stream commitment, and scoped OAuth retry behavior across provider transports.