Commit Graph
3015 Commits
Author SHA1 Message Date
stabeyandClaude Opus 5.5 45678d9419 test(gateway): decouple first-request deadline test from hyper header timeout
The partial-preface test raced a 5ms first-request deadline against a
10ms hyper header_read_timeout. On a slow runner both timers expire
before the next poll and tokio::select! may pick the connection branch,
surfacing hyper's header-timeout error instead of the clean deadline
close. Push hyper's timeout out to 30s so only the deadline can fire.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:46:43 +08:00
ZheFox d30268f80f Merge pull request #850 from AAEE86/ci/gateway-test-slim-batch1
ci(gateway): reduce Test (Gateway) runtime without duplicate execution
2026-09-24 12:22:27 +08:00
ZheFox 5679375f71 Merge pull request #853 from AAEE86/feat/codex-fingerprint-help
feat(frontend): improve Codex fingerprint setting guidance
2026-09-24 12:22:12 +08:00
AAEE86 3465d23db3 feat(frontend): improve Codex fingerprint setting guidance 2026-09-24 11:33:41 +08:00
ZheFox e3333f1aef Merge pull request #852 from zhefox/main
fix(codex): restore recovered quota status and stabilize lifecycle tests
2026-09-24 11:16:05 +08:00
ZheFox 31d5e2d172 fix(codex): restore recovered quota status and stabilize lifecycle tests 2026-09-24 10:33:45 +08:00
AAEE86 bb9f2eed8e ci: avoid rerunning gateway lib tests 2026-09-24 10:32:59 +08:00
AAEE86 926f5cc928 test: slim gateway startup helpers and configure nextest 2026-09-24 10:13:01 +08:00
AAEE86 27ae884759 ci: split gateway cache and tunnel integration scope 2026-09-24 09:35:44 +08:00
AAEE86 834eb9c308 ci(gateway): slim pool scheduler fixtures and backup candidates 2026-09-24 08:58:55 +08:00
AAEE86 3394a51278 ci(gateway): slim Test (Gateway) batch 1 — split architecture guards, pin build fingerprint
第一批减负(对应 docs/operations/gateway-ci-timeout-reduction-plan.md):

A1 架构守卫迁出
- 将 src/tests/architecture/**(208 项 / 约 1.5 万行字符串断言)迁至
  tests/architecture/,入口 tests/architecture_guard.rs
- 从 lib cfg(test) 巨型编译单元移除,压低 rustc 峰值与 OOM 风险
- 断言逻辑不变;helper 可见性改为 pub(crate)

A2 构建指纹统一
- test_gateway 的 mold RUSTFLAGS / RUST_MIN_STACK / sccache 上移至 job 级 env
- rust-ci.yml 全部 toolchain 钉住 1.95.0(与 rust-toolchain.toml、fmt/clippy 一致)
- Gateway 独立 cache key,避免 mold 指纹与无 mold job 互相污染

A4 补安全集成测试
- 新增 Test integration targets:cargo nextest run -p aether-gateway --tests
- 覆盖 architecture_guard + 此前未执行的 admin_unsigned_identity_headers

验证:architecture_guard + admin_unsigned 209 passed;
cargo check -p aether-gateway --lib --tests 通过;cargo fmt --check 通过。

不创建 PR,仅本地分支提交。
2026-09-23 17:51:45 +08:00
ZheFox 57f53903f5 Merge pull request #849 from AAEE86/codex-dynamic-client-profile
fix(codex): confine codex profile api to ai_serving root seams
2026-09-23 17:25:10 +08:00
AAEE86 81788d3a64 fix(codex): confine codex profile api to ai_serving root seams 2026-09-23 16:47:45 +08:00
ZheFox 595b8e4e05 Merge pull request #848 from AAEE86/codex-dynamic-client-profile
feat(codex): add dynamic CLI client profile
2026-09-23 16:09:01 +08:00
AAEE86 7e033d0571 feat(codex): add dynamic CLI client profile 2026-09-23 15:57:37 +08:00
ZheFox 5745442ed7 Merge pull request #847 from zhefox/main
feat(codex): add an optional 1% quota reserve for pool scheduling
2026-09-23 15:17:31 +08:00
ZheFox 1a4eba1005 feat(codex): add optional minimum quota reserve for pool scheduling 2026-09-23 14:44:29 +08:00
ZheFox 2a9d8d3b25 Merge pull request #845 from RWDai/review/pr-01-user-api-key-ip-rules
fix(admin): return user API key IP rules
2026-09-23 13:31:23 +08:00
RWDai cd765f2c2f fix(admin): return user API key IP rules 2026-09-23 11:03:45 +08:00
ZheFox ec95989e02 Merge pull request #842 from zhefox/fix/usage-response-model-conflicts
feat(usage): integrate response models and resolve conflicts after #841
2026-09-22 23:36:58 +08:00
ZheFox 7f5e1a64fe merge: integrate usage response models with service tier badges
Merge upstream PR #837, preserving its original commits and resolving the UsageModelDisplay layout conflict after #841. Cover response models alongside dynamic service-tier badges in detail tests.
2026-09-22 23:31:22 +08:00
ZheFox f86dd10467 Merge pull request #841 from zhefox/fix/codex-service-tier-passthrough
fix(codex): preserve explicit service tiers and adapt usage badges
2026-09-22 22:28:20 +08:00
ZheFox 69930a6059 fix(codex): preserve explicit service tiers and adapt usage badges 2026-09-22 22:19:52 +08:00
AAEE86 07cb401fd4 fix: extract nested provider response models 2026-09-22 11:54:58 +08:00
AAEE86 70d1a4ab74 fix: import usage body capture state in tests 2026-09-21 10:40:38 +08:00
ZheFox 0b7c7f94ac Merge pull request #833 from AAEE86/feat/usage-skipped-candidates
feat(usage): 展示调度跳过候选及原因并补齐手机端提示
2026-09-21 10:36:48 +08:00
ZheFox 67d0414483 Merge pull request #836 from dalamudx/fix/gemini-thought-signature-replay
fix(gemini): preserve tool thought signatures
2026-09-21 10:36:09 +08:00
AAEE86 e3c01fb554 fix: avoid usage payload json recursion overflow 2026-09-21 09:58:46 +08:00
AAEE86 f960bbd2c8 feat: expose upstream response model in usage records 2026-09-21 09:51:33 +08:00
dalamudx 906baae88e fix(gemini): preserve tool thought signatures 2026-09-20 01:01:26 +08:00
ZheFox ba7c9f8b27 Merge pull request #834 from dalamudx/feat/user-group-stats
feat(stats): add user group usage views
2026-09-19 20:43:45 +08:00
ZheFox 37e3a36680 Merge pull request #835 from Kayphoon/fix/responses-reasoning-content-only
fix(responses): keep raw reasoning on content only
2026-09-19 20:43:09 +08:00
Kayphoon 166de33355 fix(responses): keep raw reasoning on content only
Raw chain-of-thought was written to both `content` (`reasoning_text`) and
`summary` (`summary_text`), and the stream emitter sent the same delta on
`response.reasoning_text.delta` *and* `response.reasoning_summary_text.delta`.

Clients that render both channels therefore printed every thinking chunk
twice — most visibly the Codex CLI, whose thinking panel repeated itself.

OpenAI keeps the two channels distinct: `content` carries the raw CoT while
`summary` is the summarised view. Emit the thinking on `content` only:

- `openai_responses_reasoning_text_fields` becomes
  `openai_responses_reasoning_text_parts`, returning just the `content` array;
  reasoning items keep `summary: []` (or a provider-supplied summary).
- The Responses stream emitter emits `response.reasoning_text.delta` /
  `.done` and no longer mirrors them onto the summary events. The reasoning
  `output_item.added` no longer announces a `reasoning_summary_part`.
- The provider-state reasoning reader accepts `content` (`reasoning_text`)
  first and falls back to `summary`, so it also understands items produced by
  older Aether versions; its state field is renamed accordingly.
- Non-streaming builders (Chat -> Responses, manual Responses response, Grok
  gateway) place the thinking on `content` and leave `summary` empty.

Tests cover the raw thinking appearing exactly once in the emitted stream.
2026-09-18 18:44:28 +00:00
wangpengxiang a95f0d2488 feat(stats): add user group usage views 2026-09-18 16:24:03 +08:00
AAEE86 0486435f16 feat(usage): 展示调度跳过候选及原因并补齐手机端提示 2026-09-18 14:24:43 +08:00
ZheFox fb25dde4c9 Merge pull request #832 from dalamudx/fix/antigravity-schema-thought-replay
fix(antigravity): harden tool schemas and Claude thought replay
2026-09-17 18:39:04 +08:00
wangpengxiang 4124749a7d fix(ai-serving): route provider-aware normalization through root seams
Move Antigravity schema-preservation policy into the format crate and expose provider-aware Chat and Responses builders through the existing gateway root seam. Preserve legacy conversion and scoped Responses history behavior without weakening architecture tests.

Validated: 61 standalone architecture tests, 947 format tests, 527 transport tests, 2 actual-source planner tests, and gateway cargo check.
2026-09-17 15:38:27 +08:00
wangpengxiang 5a55116b62 fix(antigravity): harden tool schemas and Claude thought replay
Preserve tool schemas across provider-scoped format conversion and OpenAI gateway planners until the Antigravity boundary. Bound reference expansion, safely merge schema constraints, and lower unsupported Claude unions.

Drop unsigned historical Claude thinking without changing Gemini behavior. Add cross-format regression fixtures and retain the project's original error output policy.

Validation: 945 format tests and 527 transport tests passed; gateway cargo check passed.
2026-09-17 15:38:27 +08:00
ZheFox 681ce56c4f Merge pull request #831 from stabey/fix/gemini-native-search
fix(gemini): 原生 googleSearch 搜索从请求到响应全链路不可用
2026-09-17 11:34:12 +08:00
stabeyandClaude Opus 5 bcb2308000 feat(ai-formats): deliver Gemini grounding to every client as native citations
Gemini runs `googleSearch` inside Google. The search leaves no
client-visible tool call, and the evidence arrives only as
`candidates[].groundingMetadata`. Every cross-format target dropped it
wholesale, so a grounded answer reached OpenAI- and Claude-shaped clients
as prose that names its sources with nothing structured behind it: no
`annotations`, no `citations`, no `url_citation`. Callers that verify
grounding — the common "did this model actually search?" check — saw a
200 with no evidence and had to treat the answer as ungrounded.

Adapters now normalise `groundingMetadata` into neutral citations and
each target renders its own family's standard shape: `url_citation`
annotations for `openai:chat` and `openai:responses`, and
`web_search_result_location` citations on the text block for
`claude:messages`. Gemini reports segment bounds as UTF-8 byte offsets
while both targets count characters, so the bounds are converted rather
than copied.

Streaming is covered too, since that is what grounded traffic actually
uses. A new `CanonicalStreamEvent::Citations` carries the neutral list
once the answer text is whole — the offsets index into the finished
answer, so it rides just ahead of `Finish` rather than as a delta per
chunk — and each client emitter renders it: `delta.annotations` chunks,
`response.output_text.annotation.added` events (also kept on the finished
message item so clients that only read `response.completed` see them),
and `citations_delta` content block deltas.

For reference, CLIProxyAPI projects grounding only in its
antigravity→Claude translator, and only when the client declared a typed
`web_search_*` tool; its OpenAI and plain Gemini translators have no
grounding handling at all. The citation shape here matches theirs, but
the coverage is deliberately wider: all three targets, streaming and
non-streaming, with no dependency on a declared tool.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-17 09:38:04 +08:00
stabeyandClaude Opus 5 6c92db2ba5 fix(antigravity): send googleSearch instead of the Gemini 1.5 retrieval tool
The transport boundary rewrote `googleSearch` into the Gemini 1.5-era
`googleSearchRetrieval` spelling before every v1internal call, on the stated
grounds that the private backend rejects `googleSearch` when it is combined
with function declarations. That rewrite breaks grounding on Gemini 3.

Observed on stabey-124 against daily-cloudcode-pa.googleapis.com. A controlled
pair, same model and keys, 5 seconds apart:

- no `web_search_options` -> 200
- with `web_search_options` -> 502 on all three candidates

The outgoing body carried `tools: [{"googleSearchRetrieval": {}}]` and no
function declarations at all, so the documented mixed-tool rationale did not
apply. `request_candidates.error_message` holds what the backend actually
said:

    Malformed function call: call:google_search{query:current UTC date time}
    Malformed function call: call:google:search{query:current UTC date}
    Malformed function call: call:google_search{queries:[current UTC date]}

The model reaches for `google_search`, the legacy declaration binds nothing,
and the turn dies unparsed. CLIProxyAPI sends `googleSearch` to this same
v1internal surface, including alongside function declarations.

Keep folding the snake_case `google_search` alias into the canonical
`googleSearch` key, and leave a request that already spells the tool
`googleSearchRetrieval` untouched.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-17 09:23:01 +08:00
stabeyandClaude Opus 5 fdf55525f5 fix(ai-formats): keep client-declared search tools as Gemini function declarations
canonical_tools_to_gemini promoted any tool whose name normalized to
"websearch" / "googlesearch" / "websearchpreview" into Gemini's server-side
builtin, dropping it from functionDeclarations. Claude Code declares an
ordinary client-side `WebSearch` tool with a full input_schema, so every
/v1/messages request routed to a Gemini model lost that declaration and gained
`googleSearch` (rewritten to `googleSearchRetrieval` at the Anti Gravity
transport boundary) instead.

Two consequences, both observed on stabey-124 against gemini-3.8-flash:

- the model can never emit a `WebSearch` tool_use, so the client's own web
  search is dead on that route;
- when the model does reach for the injected server-side search, the v1internal
  backend answers `finishReason: MALFORMED_FUNCTION_CALL` /
  "Function call is empty - no input to parse." and the turn fails.

Promote a tool to a builtin only when it is a bare marker carrying no schema.
A declared schema means the caller intends to execute the call itself, which
matches CLIProxyAPI: it keys builtins off Claude's `type: web_search_*` or an
explicit `google_search` tool key and never off a function name.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-17 09:23:01 +08:00
ZheFox 4ff4129034 Merge pull request #823 from stabey/fix/openrouter-reasoning-fields
fix(ai-formats): OpenRouter 推理字段在 Chat 转换链路中丢失
2026-09-16 23:57:54 +08:00
ZheFox 364692da55 Merge pull request #830 from zhefox/fix/usage-full-body-retention
fix(usage): preserve full bodies before queue truncation
2026-09-16 13:22:30 +08:00
ZheFox 5842c7232e fix(usage): preserve full bodies before queue truncation 2026-09-16 13:10:40 +08:00
ZheFox 72a4bf3408 Merge pull request #829 from zhefox/fix/responses-call-id-length
fix(responses): bound upstream tool call IDs
2026-09-16 11:54:29 +08:00
ZheFox fe1723d87c fix(responses): bound upstream tool call IDs 2026-09-16 11:51:56 +08:00
ZheFox 03b198d5ab Merge pull request #825 from AAEE86/fix-scheduling-model-providers
fix(routing): 按调度配置所选模型筛选提供商
2026-09-16 10:25:33 +08:00
ZheFox 53562fd9de Merge pull request #828 from zhefox/main
test(frontend): make cross-tab refresh retry timing deterministic
2026-09-16 09:43:57 +08:00
ZheFox e66dd00b84 test(frontend): make cross-tab refresh retry timing deterministic 2026-09-16 09:39:00 +08:00