Commit Graph
227 Commits
Author SHA1 Message Date
elky 911c7f8875 feat: add selectable routing groups and composite billing
Support per-model provider enablement and compact model editing. Capture request-time billing factors, charge customer costs separately, and preserve historical statistics without backfills.
2026-10-07 14:49:57 +08:00
elky 466c7918a1 feat: unify provider scheduling workspace 2026-10-07 00:34:18 +08:00
ZheFox e7de935e61 Merge pull request #857 from stabey/codex/fix-global-model-name-reservation
fix(routing): reserve global model names across provider aliases
2026-10-06 22:15:36 +08:00
dalamudx ac7e3abab3 fix: route Gemini client identity through transport facade 2026-10-05 16:01:41 +08:00
dalamudx 7acbaab82b refactor: unify provider client identity profiles and node synchronization 2026-10-05 15:18:36 +08:00
elky d4ed774423 Merge origin/main into main 2026-10-05 12:07:04 +08:00
elky cb7b9c9ecd feat: unify user analytics and optimize overview aggregation
Merge user accounts and usage reporting into one page with a combined ranking and account table, shared precise time ranges, and simpler range labels.

Parse overview metadata once through a schema-only view migration and disable JIT locally for bucket rebuilds. Preserve automatic backfills.

Add redacted OAuth refresh diagnostics, bucket failure context, and regression coverage. Resolve strict Clippy warnings.
2026-10-05 00:28:31 +08:00
MMEXA e1dadf5b06 修复 Codex 记忆协议的 formats 入口边界与 CI 架构检查 2026-10-02 00:25:07 +08:00
MMEXA 14befeda2c 对齐 Codex CLI 0.159.3 的通用画像、模型能力与原生协议 2026-10-01 17:54:04 +08:00
dalamudx 5d5880d75e fix(claude-code): 让 openai:responses/chat 转 claude 的路径也应用请求体伪装
openai:responses 与 openai:chat 的决策构建器各自构建上游 body,不经过
apply_transport_request_body_semantics,导致跨格式请求转发到 claude_code
provider 时请求体伪装没有生效,上游仍返回 429。

- responses/decision/request.rs:finalize 之后应用伪装
- chat/decision/request.rs:finalize 成功后应用伪装
- 同格式路径改为经 ai_serving::transport 门面引用,满足架构守卫
  (ai_serving 不得直接依赖 crate::provider_transport)
- 更新 chat -> claude_code 用例断言为新的请求体形态
2026-09-29 20:18:03 +08:00
dalamudx c1aa5d618d feat(claude-code): 为 claude_code provider 补全 Claude Code 请求体特征
非 Claude Code 客户端(如 pi)经 OAuth 的 claude_code provider 转发时,
只有请求头被伪装成 Claude Code,请求体仍是客户端原样,被上游以
429 rate_limit_error 拒绝(响应无 ratelimit 额度头,并非真实限流)。

参考 sub2api 的 OAuth 请求体伪装,在传输层补齐请求体:
- system 重写为计费头 + 身份句 + 通用提示词三块(Fable 仅保留前两块)
- 原 system 迁入 messages 开头,避免丢失客户端指令
- 缺失时补 metadata.user_id,device/session id 基于 key 稳定派生
- 缺失时补 tools/temperature/max_tokens,并限制 cache_control 不超过 4 个
- 已带计费块且有 metadata.user_id 的真实 Claude Code 请求原样放行,
  重复应用幂等

接入点覆盖跨格式路径(apply_transport_request_body_semantics)和
原生 claude:messages 同格式路径,仅作用于 provider_type=claude_code。
2026-09-29 18:35:23 +08:00
stabeyandClaude Opus 5 2516e51b4e fix(routing): keep global model names out of provider alias reach
A provider model can be addressed by its upstream name or by any of its
`provider_model_mappings` entries, and neither name is published in the
model catalog, which lists global model names only. Resolution ran per API
format, so an alias could win a format the real global model had no
provider in: a `claude:messages` client asking for `gemini-3.8-flash`
landed on the provider that merely renames its own `gemini-3.8-flash-cursor`
model to `gemini-3.8-flash` on the way upstream, and the separate global
model stopped distinguishing the two routes.

Treat global model names as a reserved namespace instead: when the request
names an active global model, only rows bound to it may serve it, whatever
API format they sit in. Rows in hand answer that question for free whenever
one of them is bound to a global model of that name, so the lookup stays off
the path ordinary requests take. Authorization resolves the same way, so an
API key's allowed models cannot be satisfied through a resolution candidate
planning will no longer make.

A request naming a model that is not a global model keeps every matching
rule, so addressing a provider variant by its upstream name still works.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-27 04:01:44 +08:00
AAEE86 7e033d0571 feat(codex): add dynamic CLI client profile 2026-09-23 15:57:37 +08:00
wangpengxiang 4124749a7d fix(ai-serving): route provider-aware normalization through root seams
Move Antigravity schema-preservation policy into the format crate and expose provider-aware Chat and Responses builders through the existing gateway root seam. Preserve legacy conversion and scoped Responses history behavior without weakening architecture tests.

Validated: 61 standalone architecture tests, 947 format tests, 527 transport tests, 2 actual-source planner tests, and gateway cargo check.
2026-09-17 15:38:27 +08:00
wangpengxiang 5a55116b62 fix(antigravity): harden tool schemas and Claude thought replay
Preserve tool schemas across provider-scoped format conversion and OpenAI gateway planners until the Antigravity boundary. Bound reference expansion, safely merge schema constraints, and lower unsupported Claude unions.

Drop unsigned historical Claude thinking without changing Gemini behavior. Add cross-format regression fixtures and retain the project's original error output policy.

Validation: 945 format tests and 527 transport tests passed; gateway cargo check passed.
2026-09-17 15:38:27 +08:00
ZheFox 5a6692ade0 Merge pull request #818 from wanzhao-ysy/fix/antigravity-omit-agent-request-type
fix(antigravity): omit agent requestType from v1internal envelope
2026-09-15 10:02:55 +08:00
ZheFox e7864e5611 Merge pull request #822 from stabey/upstream-pr/xai-media
feat(providers): 新增 xAI Provider(设备码 OAuth + 原生图像/视频)
2026-09-15 09:58:55 +08:00
hkxiaoyao 01acff0774 fix(routing): preserve fixed order for streaming chat 2026-09-15 08:13:45 +08:00
stabeyandClaude Opus 5 04c4a97766 feat(xai): add native image and video endpoints
Expose the xAI Imagine image and video surfaces on top of the `xai`
provider, and make the shared OpenAI video-task layer survive the
production configuration they need.

Native video requests live under /v1 (generations, edits, extensions,
with /v1/videos as a creation alias that only selects xAI candidates);
the OpenAI-compatible adapter stays under /openai/v1/videos and maps
`seconds` / `size` onto numeric duration, aspect ratio and resolution.
Clients receive an opaque Aether task ID scoped to the owning user;
polling uses the upstream task ID and the original credential, and
completed downloads fetch the returned media URL without forwarding
provider authorization to the media host.

Three fixes to the shared video layer are required for this to work
outside tests:

- OpenAI/xAI task persistence now supplies a stable 16-character
  short_id, which the PostgreSQL schema requires. Existing rows keep
  their original value across reconstruction, so no schema change or
  historical rewrite is needed.
- Task retrieval and content downloads are admitted by the production
  GET execution gate, and reconstructed tasks resolve proxy nodes,
  system proxy defaults, tunnel affinity and transport profiles through
  the same deployment resolver used for creation. A configured proxy
  route no longer silently becomes a direct request after restart.
- When the gateway also serves the frontend, /openai/v1/videos and its
  subpaths bypass the static SPA handler. Otherwise a video query
  returns HTTP 200 with text/html instead of the task JSON.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-14 21:17:21 +08:00
stabeyandClaude Opus 5 e83399db2f feat(providers): add xAI provider with device code OAuth
Add a separate `xai` provider type for xAI Grok CLI subscription accounts.
It is independent of the existing `grok` provider, which reverse-proxies
grok.com with browser cookies; behavior of `grok` is unchanged.

Account binding uses the xAI device code flow, so no local callback
listener is needed and headless deployments can bind accounts. Refresh
tokens can also be imported individually or in batches, and are rotated
on refresh.

OAuth requests default to the cli-chat-proxy Responses API; API keys and
compact stay on api.x.ai. Explicit custom gateways are preserved. Only
`openai:responses` and `openai:responses:compact` are exposed; Chat,
Claude and Gemini clients reach the provider through Aether's existing
cross-format conversion rather than new native endpoints.

Upstream Responses payloads are sanitized for what xAI actually rejects:
`previous_response_id` and `metadata.user_id` are dropped, hosted
`tool_choice` is rewritten, `web_search` is restored for converted
clients, `image_generation` is stripped on older Grok conversation
models, unsupported reasoning effort is removed, and requested
`reasoning.encrypted_content` is preserved with a replay policy keyed on
the configured provider type rather than the model name.

Quota refresh reads /user and /billing?format=credits and stores a
structured usage snapshot; a prepaid balance keeps an account selectable
after the weekly allowance is exhausted. API-key accounts skip the
subscription billing surface. The admin UI shows remaining weekly quota
as a labeled bar in the provider drawer and the pool list.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-14 21:09:03 +08:00
hkxiaoyao 7daf355e65 fix(pool): keep schedulable keys when stale inactive scores exist
- delete pool member scores when a provider key is deactivated
- fall through to the catalog key scan when the pool score page is empty
- log skipped local candidates with a skip_reason for diagnosis
- add regressions for inactive keys with stale available scores
2026-09-14 20:42:47 +08:00
wanzhao-ysy ea24d61910 fix(antigravity): omit agent requestType from v1internal envelope 2026-09-13 18:35:57 +08:00
AAEE86 30e36cd09a fix(deepseek): 仅按官方地址识别思考兼容 2026-09-10 18:42:28 +08:00
AAEE86 28cd77eb5e fix(deepseek): 完整保留思考内容并移除空值补齐 2026-09-10 18:19:48 +08:00
elky ecc16673eb fix: harden concurrency limits and high-RPM runtime paths
Bound request, stream, queue, and shutdown resource lifetimes. Reduce scheduler and Redis hot-path work and isolate database maintenance. Include regression coverage, load probes, and concurrency audit results.
2026-09-10 08:14:58 +08:00
elky 361952ada9 fix: resolve workspace lint and regression test failures 2026-09-09 11:34:45 +08:00
elky f2839ae6a7 feat(routing): add strategy failover controls 2026-09-09 09:12:09 +08:00
elky b5ed802277 chore(codex): bump default client version to 0.153.4 2026-09-06 19:58:35 +08:00
elky 9d7a0665c0 Merge remote-tracking branch 'origin/main' into codex/provider-policy-hardening 2026-09-05 16:21:10 +08:00
elky e29442a06a merge(main): resolve subscription usage policy conflicts 2026-09-05 14:21:16 +08:00
ZheFox c7676d567d Merge pull request #801 from zhefox/fix/codex-admin-model-catalog
fix(codex): refresh fingerprints and unify management model catalogs
2026-09-05 13:18:47 +08:00
ZheFox 5ca4f87951 fix(codex): unify management model catalogs and refresh fingerprints 2026-09-05 13:14:24 +08:00
stabeyandClaude Opus 5 a6dc43d5f6 fix(antigravity): wrap cross-format requests in the v1internal envelope
The gemini:generate_content URL hook rewrites any Antigravity endpoint to
/v1internal:generateContent, but only the same-format passthrough and the two
OpenAI decision paths ever built the matching envelope. A Claude Messages or
Gemini client therefore reached the standard family planner, picked up the
rewritten URL, and posted a bare Gemini body that upstream rejects with
"Invalid JSON payload received. Unknown name \"contents\"" -- four retries
across every account, then a 503 that names none of this.

Route the standard family through the shared v1internal builder the same way
gemini_cli already is, so the URL and the body come from one decision. The
OpenAI-image-to-Gemini path cannot carry an envelope at all, so it now skips
Antigravity candidates instead of sending a request upstream can only reject.

Also stop treating a configured proxy as locally unsupported. The execution
plan carries the proxy itself, and the generic and Vertex gates moved to
transport_proxy_is_locally_supported long ago; Antigravity kept rejecting on
proxy.is_some(), which no longer matches how the local runtime executes. A
proxy that resolves to no route still disqualifies the request, and transport
profiles stay unsupported because the v1internal payload never carries one.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-04 22:53:09 +08:00
elky 579f2c7cc1 feat(security): harden gateway boundaries and usage policies
Consolidate subscription usage policy enforcement, privacy-safe persistence, and gateway security hardening into one reviewable change.

Includes bounded HTTP and execution envelopes, header and protocol guards, DNS and relay validation, authentication and secret projection hardening, secure backup/install paths, and regression coverage.
2026-09-04 03:45:52 +08:00
ZheFox 4c6bafe255 Merge remote-tracking branch 'upstream/main' into codex/fix-antigravity-quota 2026-09-03 15:50:22 +08:00
elky 9309ad844f test(gateway): align routing fixtures with strategy policies 2026-09-03 14:11:15 +08:00
ZheFox 587486ab0c Merge remote-tracking branch 'upstream/main' into codex/fix-antigravity-quota 2026-09-03 13:24:39 +08:00
fawney 2cb4d554aa feat(routing): consolidate scheduling strategy configuration 2026-09-03 11:05:59 +08:00
ZheFox f822df6cce Merge remote-tracking branch 'origin/main' into codex/fix-antigravity-quota 2026-09-03 10:46:49 +08:00
ZheFox 45c840b8d3 fix(providers): refresh Antigravity grouped quotas 2026-09-03 10:46:44 +08:00
elky 7323d41fbe feat(routing): move sticky-key retries into routing policy with lazy attempts
Replace the provider/endpoint max_retries fields as the source of same-key
retries with a routing policy setting, sticky_key_attempts (default 2). Only
the first-ranked candidate is retried on the same key; every failover
candidate gets a single attempt so failover keeps advancing instead of
retrying each fallback key.

Materialize exactly one attempt per candidate and derive same-key retries in
the attempt loop after a candidate-scoped failure, so the retry budget no
longer inflates up-front materialization and needs no upper bound. The budget
travels in the report context; retries reuse the plan with a fresh candidate
id and incremented retry index. Pool groups only retry their first key within
the retry-index stride.

Expose the setting in the routing profile editor and the set_scheduling rule
action, and drop the max_retries input from the provider form.
2026-09-02 20:48:40 +08:00
elky 415b2da81b feat(routing): make routing profiles the sole scheduler policy source
Bootstrap an enabled system-default routing group from the legacy
scheduler config keys on startup, resolve the default ordering config
from that group before falling back to the legacy keys, and stop merging
keep_priority_on_conversion with the legacy flag when a policy is
resolved. Thread the policy-derived ordering config into candidate
preselection so it no longer reads system config independently.

Add per-API-format key priority overrides so a key serving several
formats keeps independent ordering, matching the legacy
global_priority_by_format semantics. Expose keep_priority_on_conversion
in the routing profile editor and read the effective policy in the
model routing preview, monitoring metrics and provider page badge.
2026-09-02 17:04:04 +08:00
zhefox 7ae984df4b fix(gateway): detect DeepSeek custom relay models 2026-09-01 22:49:35 +08:00
elky d5f34b2ee2 feat(codex): add provider outbound policy boundary 2026-09-01 21:21:42 +08:00
elky d07dc86376 refactor(codex): generalize fingerprint convergence 2026-09-01 17:05:54 +08:00
elky 3e540ce589 fix(gateway): route Codex context through transport facade 2026-09-01 15:53:47 +08:00
elky a39048ecce feat(codex): stabilize identity across retries 2026-09-01 15:33:40 +08:00
zhefox 56395945c0 fix(gateway): route Responses compaction only to Responses providers 2026-08-29 12:28:55 +08:00
ZheFox 5bcdcca784 fix(formats): normalize Responses additional tools for Chat 2026-08-28 23:03:56 +08:00
ZheFox 2c89202001 feat(gateway): add Codex Live and OpenAI Realtime
Implement preflighted Live/Realtime WebSocket transports, protocol-aware authentication, usage auditing, UI filtering, and legacy Codex permission migration.
2026-08-21 04:27:34 +08:00