Commit Graph
1017 Commits
Author SHA1 Message Date
ZheFox 54fbcc25a1 Merge pull request #871 from dalamudx/feat/claude-code-usage-quota
feat(claude-code): 支持查询 Claude Code 账号 5H/周额度并在号池和提供商详情展示
2026-09-30 09:00:53 +08:00
dalamudx 9cc4018a37 fix(claude-code): 额度查询携带 x-app/claude-cli UA 以获取 cedar_ember,并在无重置机会时显式清空 2026-09-30 00:29:18 +08:00
dalamudx b49f5c0fd7 feat(claude-code): 只读展示重置机会(cedar_ember)并补齐额度文案国际化 2026-09-29 23:55:27 +08:00
dalamudx 125cd40aa5 feat(claude-code): 被动采样 anthropic-ratelimit-unified 响应头更新额度 2026-09-29 22:57:39 +08:00
dalamudx 8093899b5a feat(claude-code): 支持通过 /api/oauth/usage 查询账号 5H/周额度并在号池展示 2026-09-29 22:40:44 +08:00
dalamudx e8ee7b4ecf fix(claude-code): 升级伪装的 Claude Code 版本到 2.1.284
上游按模型校验 Claude Code 最低版本,claude-opus-5-5 要求 >= 2.1.280,
而 profile 里写死的 2.1.161 会被拒绝(claude_code_version_too_old, 400),
且 Aether 会统一改写 UA,真实 Claude Code 客户端也会受影响。

将 cli_version 升级为 2.1.284,并同步 stainless 包版本 (0.112.1) 与
node 运行时版本 (v26.3.0),保持指纹一致;相关测试断言同步更新。
2026-09-29 18:47:59 +08:00
dalamudx c1aa5d618d feat(claude-code): 为 claude_code provider 补全 Claude Code 请求体特征
非 Claude Code 客户端(如 pi)经 OAuth 的 claude_code provider 转发时,
只有请求头被伪装成 Claude Code,请求体仍是客户端原样,被上游以
429 rate_limit_error 拒绝(响应无 ratelimit 额度头,并非真实限流)。

参考 sub2api 的 OAuth 请求体伪装,在传输层补齐请求体:
- system 重写为计费头 + 身份句 + 通用提示词三块(Fable 仅保留前两块)
- 原 system 迁入 messages 开头,避免丢失客户端指令
- 缺失时补 metadata.user_id,device/session id 基于 key 稳定派生
- 缺失时补 tools/temperature/max_tokens,并限制 cache_control 不超过 4 个
- 已带计费块且有 metadata.user_id 的真实 Claude Code 请求原样放行,
  重复应用幂等

接入点覆盖跨格式路径(apply_transport_request_body_semantics)和
原生 claude:messages 同格式路径,仅作用于 provider_type=claude_code。
2026-09-29 18:35:23 +08:00
ZheFox fb7e3fc224 Merge pull request #838 from dalamudx/feat/user-group-provider-stats
fix(stats): scope group usage by providers and add ungrouped view
2026-09-29 12:23:27 +08:00
ZheFox cafa05c4cb Merge pull request #859 from stabey/fix/protocol-conversion-live-fixes
fix(formats): repair live cross-format conversion gaps
2026-09-29 10:10:42 +08:00
ZheFox 49ec53cbd9 Merge pull request #858 from stabey/codex/fix-sse-prefetch-handoff
fix(stream): preserve parser state across SSE prefetch handoff
2026-09-29 10:06:04 +08:00
ZheFox 73f1d79637 Merge pull request #856 from Kayphoon/fix/manual-cleanup-buffered-body
fix(admin): buffer request body for manual cleanup, smtp test, and system update routes
2026-09-29 10:05:33 +08:00
ZheFox 1b8f78a992 Merge pull request #846 from RWDai/review/pr-02-user-bulk-balance
feat(admin): batch adjust user wallet balances
2026-09-29 10:05:01 +08:00
RWDai 74072e5007 test(postgres): decode ledger amount as float8 2026-09-28 18:23:20 +08:00
RWDai 4c07d9fcfb fix(admin): floor batch deductions at zero 2026-09-28 17:56:27 +08:00
stabeyandClaude Opus 5.5 75bc32cfe9 fix(formats): repair live cross-format conversion gaps
Verified against a live Antigravity + xAI deployment:

- Gemini and Claude clients calling a forced-stream Responses upstream
  (xAI, Codex) without streaming always failed: the aggregated body echoes
  request metadata (parallel_tool_calls, tools, encrypted reasoning) that
  the strict cross-format check refuses, and the gateway then wrapped the
  raw SSE capture in a client error body sent with HTTP 200. Project the
  validated aggregate to every client format, as the Chat path already
  does, and return 502 instead of raw provider bytes when a successful
  cross-format response still cannot be converted.
- Gemini stream decoding keyed tool calls by part position, so parallel
  calls arriving in separate chunks (all at parts[0]) merged into one call
  with concatenated arguments. Key them by arrival order; ids cannot be
  used because they are optional and the Antigravity envelope synthesizes
  per-chunk ids that repeat across chunks. Generated call_auto_N ids now
  follow arrival order.
- Non-stream Responses output reported truncated or filtered cross-format
  answers as completed; derive incomplete + incomplete_details from the
  canonical stop reason.
- Gemini request parsing ignored parametersJsonSchema and
  responseJsonSchema and passed OpenAPI upper-case type names (OBJECT,
  STRING) through to JSON Schema targets, which xAI rejects.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 01:49:11 +08:00
stabey d4bc058c2f fix(stream): preserve parser state across SSE prefetch handoff 2026-09-27 04:02:12 +08:00
Kayphoon 3541ccfe29 perf(usage): truncate usage_body_blobs in before-now cleanup to reclaim physical disk space 2026-09-25 06:02:14 +00:00
ZheFox d30268f80f Merge pull request #850 from AAEE86/ci/gateway-test-slim-batch1
ci(gateway): reduce Test (Gateway) runtime without duplicate execution
2026-09-24 12:22:27 +08:00
ZheFox 31d5e2d172 fix(codex): restore recovered quota status and stabilize lifecycle tests 2026-09-24 10:33:45 +08:00
AAEE86 27ae884759 ci: split gateway cache and tunnel integration scope 2026-09-24 09:35:44 +08:00
RWDai 390b73d4b6 test(data): include new migration in version expectation 2026-09-23 20:34:25 +08:00
RWDai bcd121d447 fix(admin): make bulk wallet batches idempotent 2026-09-23 19:43:18 +08:00
ZheFox 595b8e4e05 Merge pull request #848 from AAEE86/codex-dynamic-client-profile
feat(codex): add dynamic CLI client profile
2026-09-23 16:09:01 +08:00
AAEE86 7e033d0571 feat(codex): add dynamic CLI client profile 2026-09-23 15:57:37 +08:00
ZheFox 1a4eba1005 feat(codex): add optional minimum quota reserve for pool scheduling 2026-09-23 14:44:29 +08:00
RWDai e25e240d16 feat(admin): batch adjust user wallet balances 2026-09-23 11:30:33 +08:00
ZheFox 7f5e1a64fe merge: integrate usage response models with service tier badges
Merge upstream PR #837, preserving its original commits and resolving the UsageModelDisplay layout conflict after #841. Cover response models alongside dynamic service-tier badges in detail tests.
2026-09-22 23:31:22 +08:00
ZheFox 69930a6059 fix(codex): preserve explicit service tiers and adapt usage badges 2026-09-22 22:19:52 +08:00
AAEE86 07cb401fd4 fix: extract nested provider response models 2026-09-22 11:54:58 +08:00
wangpengxiang 593327c803 refactor(stats): pass query object to raw usage summary 2026-09-21 13:07:25 +08:00
wangpengxiang a9a7c64e5d fix(stats): scope group usage by providers and add ungrouped view 2026-09-21 12:49:48 +08:00
AAEE86 70d1a4ab74 fix: import usage body capture state in tests 2026-09-21 10:40:38 +08:00
AAEE86 e3c01fb554 fix: avoid usage payload json recursion overflow 2026-09-21 09:58:46 +08:00
AAEE86 f960bbd2c8 feat: expose upstream response model in usage records 2026-09-21 09:51:33 +08:00
dalamudx 906baae88e fix(gemini): preserve tool thought signatures 2026-09-20 01:01:26 +08:00
ZheFox ba7c9f8b27 Merge pull request #834 from dalamudx/feat/user-group-stats
feat(stats): add user group usage views
2026-09-19 20:43:45 +08:00
Kayphoon 166de33355 fix(responses): keep raw reasoning on content only
Raw chain-of-thought was written to both `content` (`reasoning_text`) and
`summary` (`summary_text`), and the stream emitter sent the same delta on
`response.reasoning_text.delta` *and* `response.reasoning_summary_text.delta`.

Clients that render both channels therefore printed every thinking chunk
twice — most visibly the Codex CLI, whose thinking panel repeated itself.

OpenAI keeps the two channels distinct: `content` carries the raw CoT while
`summary` is the summarised view. Emit the thinking on `content` only:

- `openai_responses_reasoning_text_fields` becomes
  `openai_responses_reasoning_text_parts`, returning just the `content` array;
  reasoning items keep `summary: []` (or a provider-supplied summary).
- The Responses stream emitter emits `response.reasoning_text.delta` /
  `.done` and no longer mirrors them onto the summary events. The reasoning
  `output_item.added` no longer announces a `reasoning_summary_part`.
- The provider-state reasoning reader accepts `content` (`reasoning_text`)
  first and falls back to `summary`, so it also understands items produced by
  older Aether versions; its state field is renamed accordingly.
- Non-streaming builders (Chat -> Responses, manual Responses response, Grok
  gateway) place the thinking on `content` and leave `summary` empty.

Tests cover the raw thinking appearing exactly once in the emitted stream.
2026-09-18 18:44:28 +00:00
wangpengxiang a95f0d2488 feat(stats): add user group usage views 2026-09-18 16:24:03 +08:00
wangpengxiang 4124749a7d fix(ai-serving): route provider-aware normalization through root seams
Move Antigravity schema-preservation policy into the format crate and expose provider-aware Chat and Responses builders through the existing gateway root seam. Preserve legacy conversion and scoped Responses history behavior without weakening architecture tests.

Validated: 61 standalone architecture tests, 947 format tests, 527 transport tests, 2 actual-source planner tests, and gateway cargo check.
2026-09-17 15:38:27 +08:00
wangpengxiang 5a55116b62 fix(antigravity): harden tool schemas and Claude thought replay
Preserve tool schemas across provider-scoped format conversion and OpenAI gateway planners until the Antigravity boundary. Bound reference expansion, safely merge schema constraints, and lower unsupported Claude unions.

Drop unsigned historical Claude thinking without changing Gemini behavior. Add cross-format regression fixtures and retain the project's original error output policy.

Validation: 945 format tests and 527 transport tests passed; gateway cargo check passed.
2026-09-17 15:38:27 +08:00
stabeyandClaude Opus 5 bcb2308000 feat(ai-formats): deliver Gemini grounding to every client as native citations
Gemini runs `googleSearch` inside Google. The search leaves no
client-visible tool call, and the evidence arrives only as
`candidates[].groundingMetadata`. Every cross-format target dropped it
wholesale, so a grounded answer reached OpenAI- and Claude-shaped clients
as prose that names its sources with nothing structured behind it: no
`annotations`, no `citations`, no `url_citation`. Callers that verify
grounding — the common "did this model actually search?" check — saw a
200 with no evidence and had to treat the answer as ungrounded.

Adapters now normalise `groundingMetadata` into neutral citations and
each target renders its own family's standard shape: `url_citation`
annotations for `openai:chat` and `openai:responses`, and
`web_search_result_location` citations on the text block for
`claude:messages`. Gemini reports segment bounds as UTF-8 byte offsets
while both targets count characters, so the bounds are converted rather
than copied.

Streaming is covered too, since that is what grounded traffic actually
uses. A new `CanonicalStreamEvent::Citations` carries the neutral list
once the answer text is whole — the offsets index into the finished
answer, so it rides just ahead of `Finish` rather than as a delta per
chunk — and each client emitter renders it: `delta.annotations` chunks,
`response.output_text.annotation.added` events (also kept on the finished
message item so clients that only read `response.completed` see them),
and `citations_delta` content block deltas.

For reference, CLIProxyAPI projects grounding only in its
antigravity→Claude translator, and only when the client declared a typed
`web_search_*` tool; its OpenAI and plain Gemini translators have no
grounding handling at all. The citation shape here matches theirs, but
the coverage is deliberately wider: all three targets, streaming and
non-streaming, with no dependency on a declared tool.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-17 09:38:04 +08:00
stabeyandClaude Opus 5 6c92db2ba5 fix(antigravity): send googleSearch instead of the Gemini 1.5 retrieval tool
The transport boundary rewrote `googleSearch` into the Gemini 1.5-era
`googleSearchRetrieval` spelling before every v1internal call, on the stated
grounds that the private backend rejects `googleSearch` when it is combined
with function declarations. That rewrite breaks grounding on Gemini 3.

Observed on stabey-124 against daily-cloudcode-pa.googleapis.com. A controlled
pair, same model and keys, 5 seconds apart:

- no `web_search_options` -> 200
- with `web_search_options` -> 502 on all three candidates

The outgoing body carried `tools: [{"googleSearchRetrieval": {}}]` and no
function declarations at all, so the documented mixed-tool rationale did not
apply. `request_candidates.error_message` holds what the backend actually
said:

    Malformed function call: call:google_search{query:current UTC date time}
    Malformed function call: call:google:search{query:current UTC date}
    Malformed function call: call:google_search{queries:[current UTC date]}

The model reaches for `google_search`, the legacy declaration binds nothing,
and the turn dies unparsed. CLIProxyAPI sends `googleSearch` to this same
v1internal surface, including alongside function declarations.

Keep folding the snake_case `google_search` alias into the canonical
`googleSearch` key, and leave a request that already spells the tool
`googleSearchRetrieval` untouched.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-17 09:23:01 +08:00
stabeyandClaude Opus 5 fdf55525f5 fix(ai-formats): keep client-declared search tools as Gemini function declarations
canonical_tools_to_gemini promoted any tool whose name normalized to
"websearch" / "googlesearch" / "websearchpreview" into Gemini's server-side
builtin, dropping it from functionDeclarations. Claude Code declares an
ordinary client-side `WebSearch` tool with a full input_schema, so every
/v1/messages request routed to a Gemini model lost that declaration and gained
`googleSearch` (rewritten to `googleSearchRetrieval` at the Anti Gravity
transport boundary) instead.

Two consequences, both observed on stabey-124 against gemini-3.8-flash:

- the model can never emit a `WebSearch` tool_use, so the client's own web
  search is dead on that route;
- when the model does reach for the injected server-side search, the v1internal
  backend answers `finishReason: MALFORMED_FUNCTION_CALL` /
  "Function call is empty - no input to parse." and the turn fails.

Promote a tool to a builtin only when it is a bare marker carrying no schema.
A declared schema means the caller intends to execute the call itself, which
matches CLIProxyAPI: it keys builtins off Claude's `type: web_search_*` or an
explicit `google_search` tool key and never off a function name.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-17 09:23:01 +08:00
ZheFox 4ff4129034 Merge pull request #823 from stabey/fix/openrouter-reasoning-fields
fix(ai-formats): OpenRouter 推理字段在 Chat 转换链路中丢失
2026-09-16 23:57:54 +08:00
ZheFox 5842c7232e fix(usage): preserve full bodies before queue truncation 2026-09-16 13:10:40 +08:00
ZheFox fe1723d87c fix(responses): bound upstream tool call IDs 2026-09-16 11:51:56 +08:00
stabeyandClaude Opus 5 03496c46c5 fix(ai-formats): carry OpenRouter reasoning fields through chat conversion
OpenRouter reports reasoning under `reasoning` and `reasoning_details`
rather than the DeepSeek-style `reasoning_content` this crate recognized.
Its streaming reasoning phase sends chunks whose `delta.content` is an
empty string, so those chunks were dropped and Responses clients saw
nothing after `response.in_progress` until they timed the stream out.
The sync aggregator kept only content and tool calls, so a stream
downgraded to a sync response lost the reasoning entirely.

Read all three spellings through one helper. `reasoning_details` wins
because only it carries the block index, and OpenRouter repeats the same
text in both fields, so exactly one source is read per object. Entries
typed `reasoning.encrypted` carry opaque provider state rather than
readable text and are skipped. A change of block index closes the open
part so downstream summaries keep the provider's segmentation.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-16 00:39:20 +08:00
ZheFox e5ab73bf35 test(usage): preserve terminal build release notification 2026-09-15 17:29:28 +08:00
ZheFox 5a6692ade0 Merge pull request #818 from wanzhao-ysy/fix/antigravity-omit-agent-request-type
fix(antigravity): omit agent requestType from v1internal envelope
2026-09-15 10:02:55 +08:00
ZheFox 6e431e2ff6 Merge pull request #820 from Kayphoon/cursor/responses-reasoning-content-68bd
fix(responses): put raw reasoning in content, keep summary for CLI
2026-09-15 10:02:29 +08:00