Commit Graph
3077 Commits
Author SHA1 Message Date
stabey d4bc058c2f fix(stream): preserve parser state across SSE prefetch handoff 2026-09-27 04:02:12 +08:00
Kayphoon 3541ccfe29 perf(usage): truncate usage_body_blobs in before-now cleanup to reclaim physical disk space 2026-09-25 06:02:14 +00:00
Kayphoon bd83cff58f fix(admin): buffer request body for manual cleanup, smtp test, and system update routes 2026-09-24 19:19:18 +00:00
AAEE86 85c04335c6 ci: guard rust scope detection against false green
- changes 脚本加 set -euo pipefail,git fetch/diff 失败即中止,避免写出
  rust=false/shell=false 让下游误判为“无需测试”。
- changed_paths 为空(异常事件)时保守置 rust=true/shell=true,宁可多跑不漏测。
- check 与 data_db_smoke 增加 needs.changes.result 兜底。
- 恢复 push/pull_request 的 paths 白名单,并补齐 rust-toolchain.toml、
  .cargo/**、*.sql,workflow 触发规则与分类脚本对齐,避免“分类正确但
  workflow 未启动”的漏测。
- nextest 删除硬编码 test-threads,改用默认 num-cpus,避免在大规格 runner
  上主动压低并发;保留 slow-timeout 卡死保护。
- 补真实 TCP smoke test,覆盖管理员安全接口的监听端口与 HTTP/JSON 链路。
2026-09-24 16:50:33 +08:00
AAEE86 c08497c963 ci: make nightly rust scope explicit 2026-09-24 13:37:25 +08:00
ZheFox d30268f80f Merge pull request #850 from AAEE86/ci/gateway-test-slim-batch1
ci(gateway): reduce Test (Gateway) runtime without duplicate execution
2026-09-24 12:22:27 +08:00
ZheFox 5679375f71 Merge pull request #853 from AAEE86/feat/codex-fingerprint-help
feat(frontend): improve Codex fingerprint setting guidance
2026-09-24 12:22:12 +08:00
AAEE86 3465d23db3 feat(frontend): improve Codex fingerprint setting guidance 2026-09-24 11:33:41 +08:00
ZheFox e3333f1aef Merge pull request #852 from zhefox/main
fix(codex): restore recovered quota status and stabilize lifecycle tests
2026-09-24 11:16:05 +08:00
ZheFox 31d5e2d172 fix(codex): restore recovered quota status and stabilize lifecycle tests 2026-09-24 10:33:45 +08:00
AAEE86 bb9f2eed8e ci: avoid rerunning gateway lib tests 2026-09-24 10:32:59 +08:00
AAEE86 926f5cc928 test: slim gateway startup helpers and configure nextest 2026-09-24 10:13:01 +08:00
AAEE86 27ae884759 ci: split gateway cache and tunnel integration scope 2026-09-24 09:35:44 +08:00
AAEE86 834eb9c308 ci(gateway): slim pool scheduler fixtures and backup candidates 2026-09-24 08:58:55 +08:00
RWDai 390b73d4b6 test(data): include new migration in version expectation 2026-09-23 20:34:25 +08:00
RWDai bcd121d447 fix(admin): make bulk wallet batches idempotent 2026-09-23 19:43:18 +08:00
AAEE86 3394a51278 ci(gateway): slim Test (Gateway) batch 1 — split architecture guards, pin build fingerprint
第一批减负(对应 docs/operations/gateway-ci-timeout-reduction-plan.md):

A1 架构守卫迁出
- 将 src/tests/architecture/**(208 项 / 约 1.5 万行字符串断言)迁至
  tests/architecture/,入口 tests/architecture_guard.rs
- 从 lib cfg(test) 巨型编译单元移除,压低 rustc 峰值与 OOM 风险
- 断言逻辑不变;helper 可见性改为 pub(crate)

A2 构建指纹统一
- test_gateway 的 mold RUSTFLAGS / RUST_MIN_STACK / sccache 上移至 job 级 env
- rust-ci.yml 全部 toolchain 钉住 1.95.0(与 rust-toolchain.toml、fmt/clippy 一致)
- Gateway 独立 cache key,避免 mold 指纹与无 mold job 互相污染

A4 补安全集成测试
- 新增 Test integration targets:cargo nextest run -p aether-gateway --tests
- 覆盖 architecture_guard + 此前未执行的 admin_unsigned_identity_headers

验证:architecture_guard + admin_unsigned 209 passed;
cargo check -p aether-gateway --lib --tests 通过;cargo fmt --check 通过。

不创建 PR,仅本地分支提交。
2026-09-23 17:51:45 +08:00
RWDai 75471ae4a4 fix(admin): classify bulk wallet adjustment failures 2026-09-23 17:27:34 +08:00
ZheFox 57f53903f5 Merge pull request #849 from AAEE86/codex-dynamic-client-profile
fix(codex): confine codex profile api to ai_serving root seams
2026-09-23 17:25:10 +08:00
AAEE86 81788d3a64 fix(codex): confine codex profile api to ai_serving root seams 2026-09-23 16:47:45 +08:00
RWDai 30bb0c3130 fix: harden bulk wallet balance adjustment 2026-09-23 16:29:47 +08:00
ZheFox 595b8e4e05 Merge pull request #848 from AAEE86/codex-dynamic-client-profile
feat(codex): add dynamic CLI client profile
2026-09-23 16:09:01 +08:00
AAEE86 7e033d0571 feat(codex): add dynamic CLI client profile 2026-09-23 15:57:37 +08:00
ZheFox 5745442ed7 Merge pull request #847 from zhefox/main
feat(codex): add an optional 1% quota reserve for pool scheduling
2026-09-23 15:17:31 +08:00
ZheFox 1a4eba1005 feat(codex): add optional minimum quota reserve for pool scheduling 2026-09-23 14:44:29 +08:00
ZheFox 2a9d8d3b25 Merge pull request #845 from RWDai/review/pr-01-user-api-key-ip-rules
fix(admin): return user API key IP rules
2026-09-23 13:31:23 +08:00
RWDai e25e240d16 feat(admin): batch adjust user wallet balances 2026-09-23 11:30:33 +08:00
RWDai cd765f2c2f fix(admin): return user API key IP rules 2026-09-23 11:03:45 +08:00
ZheFox ec95989e02 Merge pull request #842 from zhefox/fix/usage-response-model-conflicts
feat(usage): integrate response models and resolve conflicts after #841
2026-09-22 23:36:58 +08:00
ZheFox 7f5e1a64fe merge: integrate usage response models with service tier badges
Merge upstream PR #837, preserving its original commits and resolving the UsageModelDisplay layout conflict after #841. Cover response models alongside dynamic service-tier badges in detail tests.
2026-09-22 23:31:22 +08:00
ZheFox f86dd10467 Merge pull request #841 from zhefox/fix/codex-service-tier-passthrough
fix(codex): preserve explicit service tiers and adapt usage badges
2026-09-22 22:28:20 +08:00
ZheFox 69930a6059 fix(codex): preserve explicit service tiers and adapt usage badges 2026-09-22 22:19:52 +08:00
AAEE86 07cb401fd4 fix: extract nested provider response models 2026-09-22 11:54:58 +08:00
wangpengxiang 593327c803 refactor(stats): pass query object to raw usage summary 2026-09-21 13:07:25 +08:00
wangpengxiang a9a7c64e5d fix(stats): scope group usage by providers and add ungrouped view 2026-09-21 12:49:48 +08:00
AAEE86 70d1a4ab74 fix: import usage body capture state in tests 2026-09-21 10:40:38 +08:00
ZheFox 0b7c7f94ac Merge pull request #833 from AAEE86/feat/usage-skipped-candidates
feat(usage): 展示调度跳过候选及原因并补齐手机端提示
2026-09-21 10:36:48 +08:00
ZheFox 67d0414483 Merge pull request #836 from dalamudx/fix/gemini-thought-signature-replay
fix(gemini): preserve tool thought signatures
2026-09-21 10:36:09 +08:00
AAEE86 e3c01fb554 fix: avoid usage payload json recursion overflow 2026-09-21 09:58:46 +08:00
AAEE86 f960bbd2c8 feat: expose upstream response model in usage records 2026-09-21 09:51:33 +08:00
dalamudx 906baae88e fix(gemini): preserve tool thought signatures 2026-09-20 01:01:26 +08:00
ZheFox ba7c9f8b27 Merge pull request #834 from dalamudx/feat/user-group-stats
feat(stats): add user group usage views
2026-09-19 20:43:45 +08:00
ZheFox 37e3a36680 Merge pull request #835 from Kayphoon/fix/responses-reasoning-content-only
fix(responses): keep raw reasoning on content only
2026-09-19 20:43:09 +08:00
Kayphoon 166de33355 fix(responses): keep raw reasoning on content only
Raw chain-of-thought was written to both `content` (`reasoning_text`) and
`summary` (`summary_text`), and the stream emitter sent the same delta on
`response.reasoning_text.delta` *and* `response.reasoning_summary_text.delta`.

Clients that render both channels therefore printed every thinking chunk
twice — most visibly the Codex CLI, whose thinking panel repeated itself.

OpenAI keeps the two channels distinct: `content` carries the raw CoT while
`summary` is the summarised view. Emit the thinking on `content` only:

- `openai_responses_reasoning_text_fields` becomes
  `openai_responses_reasoning_text_parts`, returning just the `content` array;
  reasoning items keep `summary: []` (or a provider-supplied summary).
- The Responses stream emitter emits `response.reasoning_text.delta` /
  `.done` and no longer mirrors them onto the summary events. The reasoning
  `output_item.added` no longer announces a `reasoning_summary_part`.
- The provider-state reasoning reader accepts `content` (`reasoning_text`)
  first and falls back to `summary`, so it also understands items produced by
  older Aether versions; its state field is renamed accordingly.
- Non-streaming builders (Chat -> Responses, manual Responses response, Grok
  gateway) place the thinking on `content` and leave `summary` empty.

Tests cover the raw thinking appearing exactly once in the emitted stream.
2026-09-18 18:44:28 +00:00
wangpengxiang a95f0d2488 feat(stats): add user group usage views 2026-09-18 16:24:03 +08:00
Kayphoon b296d46e97 feat(usage): surface Gemini thinkingConfig as reasoning effort
Usage records show a reasoning badge next to the model name for OpenAI and
Claude requests, but Gemini requests never got one. The extraction only read
the OpenAI/Claude shapes (`reasoning_effort`, `reasoning.effort`,
`output_config.effort`), while Gemini states its reasoning depth inside
`generationConfig.thinkingConfig` — so nothing was written to the usage
metadata and the list and detail views had no badge to render.

Read the Gemini shape too, as a fallback after the existing three so the
OpenAI and Claude paths are untouched:

- `thinkingLevel` / `thinking_level` wins when present, trimmed and
  lowercased, with the protobuf enum prefix stripped so `THINKING_LEVEL_HIGH`
  resolves like `high`.
- Otherwise `thinkingBudget` / `thinking_budget` goes through the existing
  shared budget ladder, yielding the same `low|medium|high|xhigh` vocabulary
  the badge already understands.
- Both camelCase and snake_case spellings are read, so a captured client body
  and a converted provider body resolve to the same label.
- `includeThoughts` alone is a visibility flag, not a depth, and produces no
  badge.

Two cases are handled explicitly rather than through the shared ladder:

- `thinkingBudget: 0` disables reasoning outright. The shared ladder maps
  `0..=1664` to `low`, which would report an explicitly disabled request as a
  shallow one, so it reports `none` instead.
- `THINKING_LEVEL_UNSPECIFIED` is the enum's "no explicit level" member, not a
  depth; it is rejected rather than surfaced as an `unspecified` badge.

The frontend needs no change: `UsageModelDisplay` already renders the badge
whenever the fields are present, and keeps the `high -> xhigh` mapping format
when the requested and upstream efforts differ.
2026-09-18 07:15:59 +00:00
AAEE86 0486435f16 feat(usage): 展示调度跳过候选及原因并补齐手机端提示 2026-09-18 14:24:43 +08:00
ZheFox fb25dde4c9 Merge pull request #832 from dalamudx/fix/antigravity-schema-thought-replay
fix(antigravity): harden tool schemas and Claude thought replay
2026-09-17 18:39:04 +08:00
wangpengxiang 4124749a7d fix(ai-serving): route provider-aware normalization through root seams
Move Antigravity schema-preservation policy into the format crate and expose provider-aware Chat and Responses builders through the existing gateway root seam. Preserve legacy conversion and scoped Responses history behavior without weakening architecture tests.

Validated: 61 standalone architecture tests, 947 format tests, 527 transport tests, 2 actual-source planner tests, and gateway cargo check.
2026-09-17 15:38:27 +08:00
wangpengxiang 5a55116b62 fix(antigravity): harden tool schemas and Claude thought replay
Preserve tool schemas across provider-scoped format conversion and OpenAI gateway planners until the Antigravity boundary. Bound reference expansion, safely merge schema constraints, and lower unsupported Claude unions.

Drop unsigned historical Claude thinking without changing Gemini behavior. Add cross-format regression fixtures and retain the project's original error output policy.

Validation: 945 format tests and 527 transport tests passed; gateway cargo check passed.
2026-09-17 15:38:27 +08:00