ZheFox
f822df6cce
Merge remote-tracking branch 'origin/main' into codex/fix-antigravity-quota
2026-09-03 10:46:49 +08:00
elky
7323d41fbe
feat(routing): move sticky-key retries into routing policy with lazy attempts
...
Replace the provider/endpoint max_retries fields as the source of same-key
retries with a routing policy setting, sticky_key_attempts (default 2). Only
the first-ranked candidate is retried on the same key; every failover
candidate gets a single attempt so failover keeps advancing instead of
retrying each fallback key.
Materialize exactly one attempt per candidate and derive same-key retries in
the attempt loop after a candidate-scoped failure, so the retry budget no
longer inflates up-front materialization and needs no upper bound. The budget
travels in the report context; retries reuse the plan with a fresh candidate
id and incremented retry index. Pool groups only retry their first key within
the retry-index stride.
Expose the setting in the routing profile editor and the set_scheduling rule
action, and drop the max_retries input from the provider form.
2026-09-02 20:48:40 +08:00
elky
415b2da81b
feat(routing): make routing profiles the sole scheduler policy source
...
Bootstrap an enabled system-default routing group from the legacy
scheduler config keys on startup, resolve the default ordering config
from that group before falling back to the legacy keys, and stop merging
keep_priority_on_conversion with the legacy flag when a policy is
resolved. Thread the policy-derived ordering config into candidate
preselection so it no longer reads system config independently.
Add per-API-format key priority overrides so a key serving several
formats keeps independent ordering, matching the legacy
global_priority_by_format semantics. Expose keep_priority_on_conversion
in the routing profile editor and read the effective policy in the
model routing preview, monitoring metrics and provider page badge.
2026-09-02 17:04:04 +08:00
zhefox
dbbe7b22ab
fix(pool): isolate dynamic model quota buckets and 429 scheduling
2026-09-02 15:23:23 +08:00
elky
4148ab1931
fix(routing): harden routed pool scheduling
2026-07-27 22:06:28 +08:00
MMEXA
20b27a13b2
feat(codex): 按操作语义路由 Responses V2 压缩
...
(cherry picked from commit 2fc604e047 )
2026-07-16 00:34:42 +08:00
MMEXA
dfa121dd5b
feat(openai): align GPT-5.6 and Codex request contracts
2026-07-11 07:40:12 +08:00
elky
6f00e9fc67
Improve gateway transport and usage runtime
2026-06-25 22:36:27 +08:00
elky
d336d1a7fa
Improve gateway scheduling and runtime admission
2026-06-24 01:53:45 +08:00
elky
f75894acbb
perf: reduce gateway db pressure under load
2026-06-22 00:08:48 +08:00
elky
68038c182b
Distinguish expired OAuth token status
2026-06-12 19:43:59 +08:00
fawney19
1658925f52
Disable key circuit breaker for pool providers
2026-05-29 15:16:06 +08:00
fawney19
42e723ff7c
Clarify API key concurrency skip reasons
2026-05-27 17:00:09 +08:00
fawney19
f76bbaab52
feat: support api key ip restriction rules
2026-05-20 16:11:49 +08:00
RWDai
2a3593d9c5
Update planner auth snapshot fixtures
2026-05-18 20:48:14 +08:00
fawney19
fa73655134
refactor: isolate dispatch scheduling core
2026-05-12 13:15:41 +08:00
AAEE86
67ca47afd4
fix(gateway): 修复正则模型映射未生效到出站请求
...
修复 key allowed_models 通过 global_model_mappings 正则命中时,
候选选择仍使用 provider model mapping 作为出站模型的问题。
正则命中后将 allowed model 写入 selected_provider_model_name,
确保最终 execution runtime 请求体中的 model 使用映射后的模型。
补充三层回归测试:
- scheduler-core 正则映射解析
- gateway candidate/data 候选选择输出
- gateway ai_execute 出站请求 model 捕获
2026-05-11 15:58:13 +08:00
fawney19
247ea9d1bd
Restrict scheduler affinity to cache affinity mode
2026-05-11 14:06:49 +08:00
fawney19
cc4512fbbb
Refine load balance candidate ranking
2026-05-11 02:32:11 +08:00
fawney19
e3574e1918
Track scheduler affinity epochs and key sorting
2026-05-11 01:45:49 +08:00
fawney19
7b81c77424
Fix regex model mapping direction
2026-05-11 01:43:14 +08:00
fawney19
fd44906bb7
fix provider pool quota status handling
2026-05-07 11:28:44 +08:00
fawney19
10e679bb2f
Add provider key cycle stats reset handling
2026-05-07 03:23:20 +08:00
fawney19
012f8bcdf7
feat: scope model mappings by endpoint
2026-05-07 00:48:15 +08:00
RWDai
657f9f0be1
fix: carry session affinity through Gemini files
2026-05-05 12:25:07 +08:00
RWDai
9ef7952da5
fix: keep scheduler affinity out of runtime checks
2026-05-05 11:23:54 +08:00
RWDai
eb8878695a
feat: thread session affinity through scheduler selection
2026-05-05 11:23:54 +08:00
RWDai
c596d82dec
feat: pass session scope into scheduler cache
2026-05-05 11:23:54 +08:00
fawney19
6f00cabe96
Lazy load requested model candidates
2026-05-03 20:56:31 +08:00
fawney19
a24e4a793d
Refactor pool candidate scheduling
2026-05-03 20:14:29 +08:00
fawney19
c4ea042eb4
feat: add model directive management
2026-05-03 14:48:25 +08:00
Kayphoon and fawney19
fe27fb17fb
feat: add provider-key concurrent limit ( #352 )
...
* feat: add provider-key concurrent limit
* Fix provider key concurrent limit checks
---------
Co-authored-by: fawney19 <[email protected] >
2026-05-03 01:55:23 +08:00
fawney19
c130d0e2c9
refactor ai serving modules and crates
2026-05-02 13:23:54 +08:00
fawney19
07a319259b
Normalize canonical API formats
2026-04-29 10:20:41 +08:00
AAEE86
712b484bc8
fix(scheduler): 恢复正则模型映射作为上游模型 ( #355 )
...
保留 provider_model_mappings 的优先级选择逻辑,但恢复
GlobalModel model_mappings 正则命中后的行为:使用命中的
allowed_model 作为实际上游 mapped_model。
修复了 #354 请求绕过基于正则表达式的模型映射的问题
2026-04-28 10:36:50 +08:00
Kayphoon
5311eb0da1
Fix provider model mapping selection ( #354 )
2026-04-28 09:23:39 +08:00
fawney19
032dcf40e7
fix: filter candidate rows to resolved global model name only
2026-04-28 01:12:27 +08:00
fawney19
218e73e324
Use core affinity helpers directly
2026-04-27 17:24:32 +08:00
fawney19
9da47ceb22
Remove duplicate capability candidate sorting
2026-04-27 17:24:32 +08:00
fawney19
e9f03d8d29
Remove ranked minimal selection compatibility helper
2026-04-27 17:24:32 +08:00
fawney19
e5e09b49f4
Remove ranked candidate data read wrappers
2026-04-27 17:24:32 +08:00
fawney19
dda6f34a07
Rename ranked minimal candidate reads
2026-04-27 17:24:32 +08:00
fawney19
2101a4ecc7
Drop legacy candidate ordering helpers
2026-04-27 17:24:31 +08:00
fawney19
3b542434a2
Unify candidate ranking pipeline
2026-04-27 17:24:31 +08:00
fawney19
ea3dc3257e
clean up legacy openai cli adapter names
2026-04-26 23:59:53 +08:00
fawney19
5b914aa78c
migrate ai format conversion to responses adapters
2026-04-26 23:59:53 +08:00
Entropy.Xu
e8763b3bbd
fix(kiro): 放行可刷新 OAuth 调度候选 ( #340 )
2026-04-26 23:49:55 +08:00
fawney19
77aac74590
feat(oauth): 完善账号异常识别并在调度/展示层拦截失效 OAuth 密钥
...
- 新增 aether-admin provider status 模块,统一解析账号状态(禁用/工作区停用等)
- 调度器 runtime 增加 oauth_invalid 判定,跳过刷新失败或已撤销的 OAuth 密钥(REQUEST_FAILED 保留可选)
- gateway state 在 local oauth 刷新返回 4xx 时持久化失败原因并同步状态快照
- admin pool 列表/详情回填 account 状态与 scheduling 阻塞原因(account_blocked)
- 共享 catalog 的 status_snapshot payload 附加 account 字段
2026-04-19 20:50:31 +08:00
RWDai and fawney19
b8702ae124
Fix/api key concurrency runtime miss ( #309 )
...
* test(cli): 覆盖 API key 并发等待与超时路径
* feat(scheduler): API key 并发饱和时等待可用槽位
* fix(proxy): 区分 API key 并发受限与真正的 runtime miss
* fix(outcome): runtime miss 仅归因真实执行候选
* feat(api-keys): 统一 concurrent_limit 默认值与校验辅助
* feat(admin): 独立 Key 接口支持 concurrent_limit
* feat(admin): 用户 API Key 路由支持 concurrent_limit
* feat(public): 自助 API Key 路由支持 concurrent_limit
* feat(import): 导入与存储层持久化 concurrent_limit
* feat(frontend): 同步 API Key concurrent_limit 类型定义
* feat(frontend): 独立 Key 表单支持 concurrent_limit
* feat(frontend): 管理员用户 API Key 表单支持 concurrent_limit
* feat(frontend): 自助 API Key 页面支持 concurrent_limit
* chore(fmt): 统一 runtime 归因相关 Rust 格式
* chore(fmt): 统一 admin API key 路由 Rust 格式
* chore(fmt): 统一 public 路由与相关测试 Rust 格式
* fix(test): 对齐 no-execution usage 归因断言
* test(middleware): 固定 access log tracing 用例线程模型
* fix(frontend): 提取用户 API Key payload 默认并发辅助
* fix(frontend): 保留用户 Key 的 concurrent_limit 默认值
* fix(api-keys): remove hardcoded concurrent limit default
---------
Co-authored-by: fawney19 <[email protected] >
2026-04-17 14:21:43 +08:00
Entropy.Xu and fawney19
ac1a126756
fix(kiro,pool,model): 对齐 Kiro 管理链路并修复全局模型删除行为 ( #305 )
...
* feat(pool): 号池支持跳过额度耗尽账号
- 新增 pool_advanced.skip_exhausted_accounts 开关及高级设置 UI, 默认关闭并兼容旧配置
- 为 Codex/Kiro 增加额度耗尽判定, 接入请求侧候选跳过并新增 account_quota_exhausted skip reason
- 号池列表将额度耗尽账号标记为 blocked/额度耗尽, 并补充前后端相关测试
* fix(kiro): 对齐账号管理与 provider-query 的 Rust 行为
- 修复 Kiro 单条导入误走 import-refresh-token 的前端分流, 并为误用路径返回明确错误提示
- 为 Kiro 导入与本地请求链补齐 bearer 兼容, 同步放开账号启停等 Key 更新操作的 auth_type 校验
- 实现 Kiro provider-query 本地模型测试与 failover 执行链, 并修复结果弹窗在无 trace 时无法展示 attempts/响应体的问题
* fix(model): 删除全局模型时级联清理关联提供商模型
- 对齐 Python 版本删除逻辑, GlobalModel 删除前先在事务内清理关联的 Provider Model 记录
- 修复已绑定 Provider 的模型在 Rust SQL 仓库下会被外键约束拦住、无法正常删除的问题
- 增加管理端回归测试, 覆盖绑定 Provider Model 的 GlobalModel 删除场景
* fix(kiro,ci): 恢复 Kiro OAuth 持久化并修复 Rust CI
* Fix oauth-managed provider key semantics
---------
Co-authored-by: fawney19 <[email protected] >
2026-04-17 12:57:06 +08:00