Commit Graph
25 Commits
Author SHA1 Message Date
stabeyandClaude Opus 5 04c4a97766 feat(xai): add native image and video endpoints
Expose the xAI Imagine image and video surfaces on top of the `xai`
provider, and make the shared OpenAI video-task layer survive the
production configuration they need.

Native video requests live under /v1 (generations, edits, extensions,
with /v1/videos as a creation alias that only selects xAI candidates);
the OpenAI-compatible adapter stays under /openai/v1/videos and maps
`seconds` / `size` onto numeric duration, aspect ratio and resolution.
Clients receive an opaque Aether task ID scoped to the owning user;
polling uses the upstream task ID and the original credential, and
completed downloads fetch the returned media URL without forwarding
provider authorization to the media host.

Three fixes to the shared video layer are required for this to work
outside tests:

- OpenAI/xAI task persistence now supplies a stable 16-character
  short_id, which the PostgreSQL schema requires. Existing rows keep
  their original value across reconstruction, so no schema change or
  historical rewrite is needed.
- Task retrieval and content downloads are admitted by the production
  GET execution gate, and reconstructed tasks resolve proxy nodes,
  system proxy defaults, tunnel affinity and transport profiles through
  the same deployment resolver used for creation. A configured proxy
  route no longer silently becomes a direct request after restart.
- When the gateway also serves the frontend, /openai/v1/videos and its
  subpaths bypass the static SPA handler. Otherwise a video query
  returns HTTP 200 with text/html instead of the task JSON.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-14 21:17:21 +08:00
elky ecc16673eb fix: harden concurrency limits and high-RPM runtime paths
Bound request, stream, queue, and shutdown resource lifetimes. Reduce scheduler and Redis hot-path work and isolate database maintenance. Include regression coverage, load probes, and concurrency audit results.
2026-09-10 08:14:58 +08:00
elky e29442a06a merge(main): resolve subscription usage policy conflicts 2026-09-05 14:21:16 +08:00
ZheFox 5ca4f87951 fix(codex): unify management model catalogs and refresh fingerprints 2026-09-05 13:14:24 +08:00
elky f5ec76c5c8 fix(models): isolate legacy catalog rows during fetch 2026-09-04 21:31:57 +08:00
elky b72b6ab137 fix(workers): isolate legacy catalog credentials 2026-09-04 19:42:10 +08:00
elky 579f2c7cc1 feat(security): harden gateway boundaries and usage policies
Consolidate subscription usage policy enforcement, privacy-safe persistence, and gateway security hardening into one reviewable change.

Includes bounded HTTP and execution envelopes, header and protocol guards, DNS and relay validation, authentication and secret projection hardening, secure backup/install paths, and regression coverage.
2026-09-04 03:45:52 +08:00
fawney 2cb4d554aa feat(routing): consolidate scheduling strategy configuration 2026-09-03 11:05:59 +08:00
ZheFox b13d9b9b40 fix(codex): restore upstream model discovery 2026-08-15 19:36:28 +08:00
zhefox 810c3dfe2b fix(codex): serve versioned dynamic model catalogs 2026-08-14 18:41:44 +08:00
elky fc92c4f431 perf(gateway): scale request hot paths for 20k streams
Shard and singleflight hot-path caches, batch and prioritize candidate and usage lifecycle persistence, and extend database and pressure-test instrumentation for 20k concurrent streams.
2026-07-22 02:11:08 +08:00
MMEXA dfa121dd5b feat(openai): align GPT-5.6 and Codex request contracts 2026-07-11 07:40:12 +08:00
MMEXA 6c4e730e60 修复 Antigravity OAuth 配额复检缺 project 2026-06-19 22:21:22 +08:00
Mas0nShi 9ad9858ac2 Merge origin/main into fix/gemini-cli-v1internal 2026-05-28 11:58:00 +08:00
fawney19 d18b13a91a fix: harden frontdoor and usage ingestion 2026-05-22 23:57:38 +08:00
Mas0nShi 3e6ce6cf4a fix: hydrate Gemini CLI project metadata 2026-05-21 17:10:17 +08:00
Kayphoon 2958041dc7 feat(gateway): add reversible chat pii redaction 2026-05-13 18:25:13 +08:00
fawney19 9057537ab8 fix: preserve model associations on refresh 2026-05-11 14:52:50 +08:00
fawney19 e3574e1918 Track scheduler affinity epochs and key sorting 2026-05-11 01:45:49 +08:00
fawney19 6247ac3edc refactor: extract runtime state backends 2026-05-08 00:18:12 +08:00
fawney19 861ae81ff0 feat(admin): 完善代理节点与 OAuth 授权管理 2026-04-14 14:09:24 +08:00
AAEE86 ab82841426 feat(model-fetch): 对齐 Rust 上游模型抓取行为到 Python 语义
将 Rust 版上游可用模型抓取逻辑收敛到 Python 版行为,统一后台自动抓模
与管理员 provider-query 的模型发现路径,消除标准 /models、固定模型目录、
Antigravity、Vertex AI 等 provider 在两端实现上的分叉。

核心变更:
- 在 aether-model-fetch 中引入统一抓模策略层
- 覆盖标准 /models、Vertex API Key、Vertex Service Account、
  Antigravity fetchAvailableModels、固定模型目录五类抓模路径
- 将 provider-query 与后台自动抓模都切换到共享抓模入口,避免重复拼接
  URL、headers 和 provider 特判逻辑

标准模型抓取对齐:
- 按 Python 语义调整抓模优先级:
  openai:chat > openai:cli > openai:compact
  claude:chat > claude:cli
  gemini:chat > gemini:cli
- 从抓模候选中移除 openai:responses
- 为 openai:cli/openai:compact、claude:cli、gemini:* 补齐 Python 同款
  User-Agent / 浏览器指纹请求头
- Claude 抓模保留 after_id 分页语义
- Gemini 抓模统一为 v1beta/models?key=... 语义

provider-query 对齐:
- 返回结果改为按 model id 聚合,并合并/排序 api_formats
- 最终模型列表按 model id 排序,行为与 Python 保持一致
- 固定目录 provider(codex/kiro/claude_code/gemini_cli)不再依赖活跃
  endpoint,即使无 endpoint 也能返回预设模型目录
- Antigravity 多 key 查询改为按账户可用性 + tier 排序,首个成功结果即
  停止,并接入 provider 级缓存
- 仅配置 openai:responses 的 provider 不再被视为抓模成功路径

自动抓模对齐:
- 自动抓模成功时写入 allowed_models、upstream_models cache,并同步
  upstream_metadata
- upstream_metadata 合并逻辑对齐 Python,对 quota_by_model 做模型级合并,
  并保留已有 reset_time
- 自动抓模失败时不覆盖已有 allowed_models
- 固定目录 provider 在无 endpoint 场景下也可成功更新 allowed_models

Antigravity 对齐:
- 使用 POST /v1internal:fetchAvailableModels 抓取可用模型
- 按 Python 规则处理 URL fallback 和 429/404/408/5xx fallback 状态
- 强制要求 auth_config.project_id
- 过滤 Python 黑名单模型
- 解析并持久化 upstream_metadata.antigravity.quota_by_model

Vertex AI 对齐:
- API Key 模式仅抓取 publishers/google/models
- Service Account 模式新增 JWT token exchange,并按 Python region 顺序
  抓取 google + anthropic publishers
- 模型 owned_by / display_name / api_format 推断与 Python 对齐
- 软 404 处理行为与 Python 收敛

Gemini CLI / 固定目录对齐:
- Gemini CLI 改为返回 Python 预设模型目录
- 在可用时通过 loadCodeAssist 补充 plan_type/project_id 元数据
- Codex/Kiro/Claude Code 改为共享固定模型目录实现

测试:
- 扩展 aether-model-fetch 单元测试,覆盖格式优先级、openai:responses 排除、
  请求头、Claude 分页、Gemini query auth、固定目录与 metadata 合并
- 调整 provider-query 控制面测试到 Python 语义
- 新增自动抓模运行时测试,覆盖固定目录成功、Antigravity metadata 合并、
  失败保留旧 allowed_models

验证:
- cargo nextest run -p aether-model-fetch --lib
- cargo nextest run -p aether-gateway control::admin::provider_query model_fetch::runtime::tests
2026-04-12 10:14:29 +08:00
fawney19 5014e2f5fd refactor: 抽离 AI pipeline 与调度共享能力逻辑 2026-04-10 01:46:14 +08:00
fawney19 5d96d6673b refactor: 大规模模块拆分与代码精简,新增 ai-pipeline/data-contracts 独立 crate
- 新增 aether-ai-pipeline 和 aether-data-contracts crate,将 pipeline 逻辑与数据契约从 gateway 中解耦
- 重构 admin handlers:拆分单体模块为 auth/billing/endpoint/features/model/observability/provider/system 等独立子模块
- 合并 chat/cli 重复代码路径:精简 conversion、finalize、planner 中的 sync/chat/cli 分支
- 重构 scheduler/executor/data 层,引入 facade 模式降低模块间耦合
- 移除冗余的 intent 模块,将 plan_fallback/policy/stream_path/sync_path 迁移至 executor
- 前端适配:调整 admin API 调用和 provider 模型测试对话框
2026-04-07 02:50:19 +08:00
fawney19 763ff03a7b refactor: 拆分 gateway 单体为独立 crate,新增 systemd 部署方案
将 gateway 内部的 model-fetch、provider-transport、scheduler-core、
usage-runtime、video-tasks-core 模块提取为独立 crate;重构 gateway
内部模块结构(state/router/cache/data/query 等);移除大量遗留模块
文件;新增 systemd 二进制部署骨架及相关文档;更新前端 usage 相关
API 和组件。
2026-04-05 20:23:16 +08:00