Commit Graph
1029 Commits
Author SHA1 Message Date
ZheFox 0948f29da5 Merge pull request #869 from AAEE86/fix-pool-quota-reset-display
fix(pool): 配额倒计时归零后进度条与列表恢复显示 100%
2026-10-05 12:05:03 +08:00
ZheFox 8a767c4309 Merge pull request #872 from stabey/fix/responses-web-search-function-call
fix(ai-formats): keep client-declared web_search as a Responses function_call
2026-10-05 12:04:07 +08:00
ZheFox bd35e885a7 Merge pull request #877 from stabey/fix/xai-grok-cli-version
fix(xai): track the official Grok CLI version for cli-chat-proxy
2026-10-05 12:03:19 +08:00
ZheFox 8ffd8188e4 Merge pull request #879 from Kayphoon/feat/gemini-usage-reasoning-badges
feat(usage): surface Gemini thinkingConfig as reasoning effort
2026-10-05 12:03:02 +08:00
MMEXA e1dadf5b06 修复 Codex 记忆协议的 formats 入口边界与 CI 架构检查 2026-10-02 00:25:07 +08:00
MMEXA 2a63bafd20 按端点根地址契约校正原生操作回归用例 2026-10-01 18:22:43 +08:00
MMEXA ea6b739fd6 保留官方公开压缩兼容标记并明确原生错误测试策略 2026-10-01 18:16:13 +08:00
MMEXA 14befeda2c 对齐 Codex CLI 0.159.3 的通用画像、模型能力与原生协议 2026-10-01 17:54:04 +08:00
stabeyandClaude Opus 5.5 2257c3959f fix(xai): track the official Grok CLI version for cli-chat-proxy
cli-chat-proxy.grok.com started rejecting every request on 2026-10-01
with HTTP 426 "Your Grok CLI version (0.2.120) is outdated. Please
update to version 1.0.13 or later", because x-grok-client-version and
the xai-grok-workspace user agent were pinned to 0.2.120.

Replace the pin with a runtime-published version (built-in fallback
1.0.46) and add a gateway worker, modelled on the Codex profile worker,
that prewarms at startup and refreshes every 3h:

- read the official stable channel https://x.ai/cli/stable, falling
  back to npm @xai-official/grok/latest (deployments that cannot reach
  x.ai directly), requiring all six platform binaries at one version;
- never roll back, persist the verified version in runtime KV and
  restore it on restart;
- AETHER_XAI_CLIENT_VERSION pins a version, and
  AETHER_XAI_CLIENT_PROFILE_REFRESH=off disables the network check.

Endpoint header rules still win over the injected identity headers.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-01 14:53:26 +08:00
stabeyandClaude Opus 5.5 db5d2dfbb5 fix(ai-formats): keep client-declared web_search as a Responses function_call
Responses output turned every tool call named `web_search` or
`web_search_preview` into a hosted `web_search_call`, regardless of what the
client declared. OMP declares its own `{"type":"function","name":"web_search"}`
tool, so when gemini-3.8-flash called it the client got a hosted item it
cannot execute. OMP then echoed that `web_search_call` back as the last input
item with no output, the Gemini request body could not be built, and every
retry failed with 503 "上游请求体转换失败" (provider_request_body_build_failed).
Observed on stabey-124 on 2026-09-29 (request c6270ef6 and five retries after
b794e7ca returned functionCall web_search / call_109312).

Decide the hosted mapping in one place, NamespaceToolAliases::
emits_hosted_web_search_call, used by both the sync builder and the stream
emitter: emit `web_search_call` only when the name is not a namespaced child
and the client did not declare a function or custom tool of that name. When
both a hosted tool and a function share the name, the function wins.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-30 10:25:44 +08:00
ZheFox 54fbcc25a1 Merge pull request #871 from dalamudx/feat/claude-code-usage-quota
feat(claude-code): 支持查询 Claude Code 账号 5H/周额度并在号池和提供商详情展示
2026-09-30 09:00:53 +08:00
dalamudx 9cc4018a37 fix(claude-code): 额度查询携带 x-app/claude-cli UA 以获取 cedar_ember,并在无重置机会时显式清空 2026-09-30 00:29:18 +08:00
dalamudx b49f5c0fd7 feat(claude-code): 只读展示重置机会(cedar_ember)并补齐额度文案国际化 2026-09-29 23:55:27 +08:00
dalamudx 125cd40aa5 feat(claude-code): 被动采样 anthropic-ratelimit-unified 响应头更新额度 2026-09-29 22:57:39 +08:00
dalamudx 8093899b5a feat(claude-code): 支持通过 /api/oauth/usage 查询账号 5H/周额度并在号池展示 2026-09-29 22:40:44 +08:00
dalamudx e8ee7b4ecf fix(claude-code): 升级伪装的 Claude Code 版本到 2.1.284
上游按模型校验 Claude Code 最低版本,claude-opus-5-5 要求 >= 2.1.280,
而 profile 里写死的 2.1.161 会被拒绝(claude_code_version_too_old, 400),
且 Aether 会统一改写 UA,真实 Claude Code 客户端也会受影响。

将 cli_version 升级为 2.1.284,并同步 stainless 包版本 (0.112.1) 与
node 运行时版本 (v26.3.0),保持指纹一致;相关测试断言同步更新。
2026-09-29 18:47:59 +08:00
dalamudx c1aa5d618d feat(claude-code): 为 claude_code provider 补全 Claude Code 请求体特征
非 Claude Code 客户端(如 pi)经 OAuth 的 claude_code provider 转发时,
只有请求头被伪装成 Claude Code,请求体仍是客户端原样,被上游以
429 rate_limit_error 拒绝(响应无 ratelimit 额度头,并非真实限流)。

参考 sub2api 的 OAuth 请求体伪装,在传输层补齐请求体:
- system 重写为计费头 + 身份句 + 通用提示词三块(Fable 仅保留前两块)
- 原 system 迁入 messages 开头,避免丢失客户端指令
- 缺失时补 metadata.user_id,device/session id 基于 key 稳定派生
- 缺失时补 tools/temperature/max_tokens,并限制 cache_control 不超过 4 个
- 已带计费块且有 metadata.user_id 的真实 Claude Code 请求原样放行,
  重复应用幂等

接入点覆盖跨格式路径(apply_transport_request_body_semantics)和
原生 claude:messages 同格式路径,仅作用于 provider_type=claude_code。
2026-09-29 18:35:23 +08:00
AAEE86 309f507ef4 fix(pool): 配额倒计时归零后进度条与列表恢复显示 100%
- 后端读取层 provider_key_status_snapshot_payload 复用调度侧同一判定
  provider_pool_reset_deadline_elapsed,对已到期的 Codex 配额窗口归一化为
  used_ratio=0 / remaining_ratio=1,并清除窗口级耗尽标记;跳过
  window_minutes=0 与无用量观测的窗口
- 前端展示层在倒计时归零(isExpired)时兜底显示 100%,并隐藏重置前的旧用量
  文本,保证归零瞬间即时恢复,与后端读取口径、调度口径一致
- 补充后端 2 个单测(到期归一化、遗留耗尽状态清理)与前端 1 个回归用例
  (负向验证可复现旧行为)

验证:aether-provider-pool 71/71;aether-gateway lib 回归通过(仅 3 个依赖
PostgreSQL 服务端的用例因本机环境缺失失败,与本改动无关);前端 222 文件 /
1718 用例全部通过;clippy -D warnings、cargo fmt、vue-tsc、eslint 均通过。
2026-09-29 16:52:26 +08:00
ZheFox fb7e3fc224 Merge pull request #838 from dalamudx/feat/user-group-provider-stats
fix(stats): scope group usage by providers and add ungrouped view
2026-09-29 12:23:27 +08:00
ZheFox cafa05c4cb Merge pull request #859 from stabey/fix/protocol-conversion-live-fixes
fix(formats): repair live cross-format conversion gaps
2026-09-29 10:10:42 +08:00
ZheFox 49ec53cbd9 Merge pull request #858 from stabey/codex/fix-sse-prefetch-handoff
fix(stream): preserve parser state across SSE prefetch handoff
2026-09-29 10:06:04 +08:00
ZheFox 73f1d79637 Merge pull request #856 from Kayphoon/fix/manual-cleanup-buffered-body
fix(admin): buffer request body for manual cleanup, smtp test, and system update routes
2026-09-29 10:05:33 +08:00
ZheFox 1b8f78a992 Merge pull request #846 from RWDai/review/pr-02-user-bulk-balance
feat(admin): batch adjust user wallet balances
2026-09-29 10:05:01 +08:00
RWDai 74072e5007 test(postgres): decode ledger amount as float8 2026-09-28 18:23:20 +08:00
RWDai 4c07d9fcfb fix(admin): floor batch deductions at zero 2026-09-28 17:56:27 +08:00
stabeyandClaude Opus 5.5 75bc32cfe9 fix(formats): repair live cross-format conversion gaps
Verified against a live Antigravity + xAI deployment:

- Gemini and Claude clients calling a forced-stream Responses upstream
  (xAI, Codex) without streaming always failed: the aggregated body echoes
  request metadata (parallel_tool_calls, tools, encrypted reasoning) that
  the strict cross-format check refuses, and the gateway then wrapped the
  raw SSE capture in a client error body sent with HTTP 200. Project the
  validated aggregate to every client format, as the Chat path already
  does, and return 502 instead of raw provider bytes when a successful
  cross-format response still cannot be converted.
- Gemini stream decoding keyed tool calls by part position, so parallel
  calls arriving in separate chunks (all at parts[0]) merged into one call
  with concatenated arguments. Key them by arrival order; ids cannot be
  used because they are optional and the Antigravity envelope synthesizes
  per-chunk ids that repeat across chunks. Generated call_auto_N ids now
  follow arrival order.
- Non-stream Responses output reported truncated or filtered cross-format
  answers as completed; derive incomplete + incomplete_details from the
  canonical stop reason.
- Gemini request parsing ignored parametersJsonSchema and
  responseJsonSchema and passed OpenAPI upper-case type names (OBJECT,
  STRING) through to JSON Schema targets, which xAI rejects.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 01:49:11 +08:00
stabey d4bc058c2f fix(stream): preserve parser state across SSE prefetch handoff 2026-09-27 04:02:12 +08:00
Kayphoon 3541ccfe29 perf(usage): truncate usage_body_blobs in before-now cleanup to reclaim physical disk space 2026-09-25 06:02:14 +00:00
ZheFox d30268f80f Merge pull request #850 from AAEE86/ci/gateway-test-slim-batch1
ci(gateway): reduce Test (Gateway) runtime without duplicate execution
2026-09-24 12:22:27 +08:00
ZheFox 31d5e2d172 fix(codex): restore recovered quota status and stabilize lifecycle tests 2026-09-24 10:33:45 +08:00
AAEE86 27ae884759 ci: split gateway cache and tunnel integration scope 2026-09-24 09:35:44 +08:00
RWDai 390b73d4b6 test(data): include new migration in version expectation 2026-09-23 20:34:25 +08:00
RWDai bcd121d447 fix(admin): make bulk wallet batches idempotent 2026-09-23 19:43:18 +08:00
ZheFox 595b8e4e05 Merge pull request #848 from AAEE86/codex-dynamic-client-profile
feat(codex): add dynamic CLI client profile
2026-09-23 16:09:01 +08:00
AAEE86 7e033d0571 feat(codex): add dynamic CLI client profile 2026-09-23 15:57:37 +08:00
ZheFox 1a4eba1005 feat(codex): add optional minimum quota reserve for pool scheduling 2026-09-23 14:44:29 +08:00
RWDai e25e240d16 feat(admin): batch adjust user wallet balances 2026-09-23 11:30:33 +08:00
ZheFox 7f5e1a64fe merge: integrate usage response models with service tier badges
Merge upstream PR #837, preserving its original commits and resolving the UsageModelDisplay layout conflict after #841. Cover response models alongside dynamic service-tier badges in detail tests.
2026-09-22 23:31:22 +08:00
ZheFox 69930a6059 fix(codex): preserve explicit service tiers and adapt usage badges 2026-09-22 22:19:52 +08:00
AAEE86 07cb401fd4 fix: extract nested provider response models 2026-09-22 11:54:58 +08:00
wangpengxiang 593327c803 refactor(stats): pass query object to raw usage summary 2026-09-21 13:07:25 +08:00
wangpengxiang a9a7c64e5d fix(stats): scope group usage by providers and add ungrouped view 2026-09-21 12:49:48 +08:00
AAEE86 70d1a4ab74 fix: import usage body capture state in tests 2026-09-21 10:40:38 +08:00
AAEE86 e3c01fb554 fix: avoid usage payload json recursion overflow 2026-09-21 09:58:46 +08:00
AAEE86 f960bbd2c8 feat: expose upstream response model in usage records 2026-09-21 09:51:33 +08:00
dalamudx 906baae88e fix(gemini): preserve tool thought signatures 2026-09-20 01:01:26 +08:00
ZheFox ba7c9f8b27 Merge pull request #834 from dalamudx/feat/user-group-stats
feat(stats): add user group usage views
2026-09-19 20:43:45 +08:00
Kayphoon 166de33355 fix(responses): keep raw reasoning on content only
Raw chain-of-thought was written to both `content` (`reasoning_text`) and
`summary` (`summary_text`), and the stream emitter sent the same delta on
`response.reasoning_text.delta` *and* `response.reasoning_summary_text.delta`.

Clients that render both channels therefore printed every thinking chunk
twice — most visibly the Codex CLI, whose thinking panel repeated itself.

OpenAI keeps the two channels distinct: `content` carries the raw CoT while
`summary` is the summarised view. Emit the thinking on `content` only:

- `openai_responses_reasoning_text_fields` becomes
  `openai_responses_reasoning_text_parts`, returning just the `content` array;
  reasoning items keep `summary: []` (or a provider-supplied summary).
- The Responses stream emitter emits `response.reasoning_text.delta` /
  `.done` and no longer mirrors them onto the summary events. The reasoning
  `output_item.added` no longer announces a `reasoning_summary_part`.
- The provider-state reasoning reader accepts `content` (`reasoning_text`)
  first and falls back to `summary`, so it also understands items produced by
  older Aether versions; its state field is renamed accordingly.
- Non-streaming builders (Chat -> Responses, manual Responses response, Grok
  gateway) place the thinking on `content` and leave `summary` empty.

Tests cover the raw thinking appearing exactly once in the emitted stream.
2026-09-18 18:44:28 +00:00
wangpengxiang a95f0d2488 feat(stats): add user group usage views 2026-09-18 16:24:03 +08:00
Kayphoon b296d46e97 feat(usage): surface Gemini thinkingConfig as reasoning effort
Usage records show a reasoning badge next to the model name for OpenAI and
Claude requests, but Gemini requests never got one. The extraction only read
the OpenAI/Claude shapes (`reasoning_effort`, `reasoning.effort`,
`output_config.effort`), while Gemini states its reasoning depth inside
`generationConfig.thinkingConfig` — so nothing was written to the usage
metadata and the list and detail views had no badge to render.

Read the Gemini shape too, as a fallback after the existing three so the
OpenAI and Claude paths are untouched:

- `thinkingLevel` / `thinking_level` wins when present, trimmed and
  lowercased, with the protobuf enum prefix stripped so `THINKING_LEVEL_HIGH`
  resolves like `high`.
- Otherwise `thinkingBudget` / `thinking_budget` goes through the existing
  shared budget ladder, yielding the same `low|medium|high|xhigh` vocabulary
  the badge already understands.
- Both camelCase and snake_case spellings are read, so a captured client body
  and a converted provider body resolve to the same label.
- `includeThoughts` alone is a visibility flag, not a depth, and produces no
  badge.

Two cases are handled explicitly rather than through the shared ladder:

- `thinkingBudget: 0` disables reasoning outright. The shared ladder maps
  `0..=1664` to `low`, which would report an explicitly disabled request as a
  shallow one, so it reports `none` instead.
- `THINKING_LEVEL_UNSPECIFIED` is the enum's "no explicit level" member, not a
  depth; it is rejected rather than surfaced as an `unspecified` badge.

The frontend needs no change: `UsageModelDisplay` already renders the badge
whenever the fields are present, and keeps the `high -> xhigh` mapping format
when the requested and upstream efforts differ.
2026-09-18 07:15:59 +00:00