Commit Graph
28 Commits
Author SHA1 Message Date
stabeyandClaude Opus 5 04c4a97766 feat(xai): add native image and video endpoints
Expose the xAI Imagine image and video surfaces on top of the `xai`
provider, and make the shared OpenAI video-task layer survive the
production configuration they need.

Native video requests live under /v1 (generations, edits, extensions,
with /v1/videos as a creation alias that only selects xAI candidates);
the OpenAI-compatible adapter stays under /openai/v1/videos and maps
`seconds` / `size` onto numeric duration, aspect ratio and resolution.
Clients receive an opaque Aether task ID scoped to the owning user;
polling uses the upstream task ID and the original credential, and
completed downloads fetch the returned media URL without forwarding
provider authorization to the media host.

Three fixes to the shared video layer are required for this to work
outside tests:

- OpenAI/xAI task persistence now supplies a stable 16-character
  short_id, which the PostgreSQL schema requires. Existing rows keep
  their original value across reconstruction, so no schema change or
  historical rewrite is needed.
- Task retrieval and content downloads are admitted by the production
  GET execution gate, and reconstructed tasks resolve proxy nodes,
  system proxy defaults, tunnel affinity and transport profiles through
  the same deployment resolver used for creation. A configured proxy
  route no longer silently becomes a direct request after restart.
- When the gateway also serves the frontend, /openai/v1/videos and its
  subpaths bypass the static SPA handler. Otherwise a video query
  returns HTTP 200 with text/html instead of the task JSON.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-14 21:17:21 +08:00
stabeyandClaude Opus 5 e83399db2f feat(providers): add xAI provider with device code OAuth
Add a separate `xai` provider type for xAI Grok CLI subscription accounts.
It is independent of the existing `grok` provider, which reverse-proxies
grok.com with browser cookies; behavior of `grok` is unchanged.

Account binding uses the xAI device code flow, so no local callback
listener is needed and headless deployments can bind accounts. Refresh
tokens can also be imported individually or in batches, and are rotated
on refresh.

OAuth requests default to the cli-chat-proxy Responses API; API keys and
compact stay on api.x.ai. Explicit custom gateways are preserved. Only
`openai:responses` and `openai:responses:compact` are exposed; Chat,
Claude and Gemini clients reach the provider through Aether's existing
cross-format conversion rather than new native endpoints.

Upstream Responses payloads are sanitized for what xAI actually rejects:
`previous_response_id` and `metadata.user_id` are dropped, hosted
`tool_choice` is rewritten, `web_search` is restored for converted
clients, `image_generation` is stripped on older Grok conversation
models, unsupported reasoning effort is removed, and requested
`reasoning.encrypted_content` is preserved with a replay policy keyed on
the configured provider type rather than the model name.

Quota refresh reads /user and /billing?format=credits and stores a
structured usage snapshot; a prepaid balance keeps an account selectable
after the weekly allowance is exhausted. API-key accounts skip the
subscription billing surface. The admin UI shows remaining weekly quota
as a labeled bar in the provider drawer and the pool list.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-14 21:09:03 +08:00
elky b5ed802277 chore(codex): bump default client version to 0.153.4 2026-09-06 19:58:35 +08:00
elky e29442a06a merge(main): resolve subscription usage policy conflicts 2026-09-05 14:21:16 +08:00
ZheFox 5ca4f87951 fix(codex): unify management model catalogs and refresh fingerprints 2026-09-05 13:14:24 +08:00
elky 579f2c7cc1 feat(security): harden gateway boundaries and usage policies
Consolidate subscription usage policy enforcement, privacy-safe persistence, and gateway security hardening into one reviewable change.

Includes bounded HTTP and execution envelopes, header and protocol guards, DNS and relay validation, authentication and secret projection hardening, secure backup/install paths, and regression coverage.
2026-09-04 03:45:52 +08:00
ZheFox 45c840b8d3 fix(providers): refresh Antigravity grouped quotas 2026-09-03 10:46:44 +08:00
ZheFox b13d9b9b40 fix(codex): restore upstream model discovery 2026-08-15 19:36:28 +08:00
zhefox 810c3dfe2b fix(codex): serve versioned dynamic model catalogs 2026-08-14 18:41:44 +08:00
MMEXA dfa121dd5b feat(openai): align GPT-5.6 and Codex request contracts 2026-07-11 07:40:12 +08:00
fawney19 b59c724455 Normalize endpoint API root handling 2026-05-29 02:29:33 +08:00
zhefox 1173a4d9d5 fix(provider): split partial model fetch warnings from errors 2026-05-26 17:45:06 +08:00
zhefox aa409a8a9c fix(provider): refine endpoint default paths for openai and claude roots 2026-05-26 16:50:36 +08:00
zhefox d28a389a93 fix(provider): support unversioned API roots in model fetch 2026-05-26 15:39:56 +08:00
Entropy.Xu 129c7c90c0 fix(provider): 修复 Windsurf 原生工具桥接 2026-05-21 01:02:02 +08:00
Entropy.Xu 9466d92a7a fix(provider): 接入 Windsurf 模型拉取和格式转换 2026-05-21 01:02:02 +08:00
mayrain cbfe1d378f feat(grok): add provider pool and transport support 2026-05-16 21:15:38 +08:00
AAEE86 2543676437 feat(model-fetch): fetch Kiro models from upstream
- add Kiro ListAvailableModels request planning and headers
- route Kiro model refresh through upstream fetch
- normalize Kiro model payloads and default model metadata
- remove profileArn requirement from ListAvailableModels
2026-05-08 13:24:27 +08:00
fawney19 1d62722d47 Improve Codex model fetching 2026-05-04 01:16:56 +08:00
fawney19 07a319259b Normalize canonical API formats 2026-04-29 10:20:41 +08:00
fawney19 4ec591fbf2 centralize openai responses alias handling 2026-04-26 23:59:53 +08:00
fawney19 5b914aa78c migrate ai format conversion to responses adapters 2026-04-26 23:59:53 +08:00
fawney19andEntropy.Xu fa328e18a1 feat: provider api_formats 可空继承、OpenAI 图片 edit/variation 与用量配额多项补强
- 鉴权: provider_api_keys.api_formats 改为可空,OAuth 托管 key 自动继承 provider endpoints 激活格式,相关 handler/测试同步更新
- 图片 planner: OpenAI 图片路由新增 edit/variation 操作并完善参数校验、响应合并与流式处理
- 用量: user me usage 返回区分 client_requested_stream/upstream_is_stream,前端 usage 列表筛选与展示增强
- 统计: stats_daily_model 新增 cache_creation_ephemeral_5m/1h tokens 字段与回填链路
- 配额/observability: quota repository 新增内存与 SQL 扩展,admin observability usage 字段扩充
- 其它: OAuth 导入/轮询收敛、provider 汇总与 pool admin 读写链路小修、新增 system_config 缓存与 provider template handler

Closes #318

Co-authored-by: Entropy.Xu <[email protected]>
2026-04-23 14:42:51 +08:00
Entropy.Xu f55f22d2e8 feat(codex-image): 封装 GPT Image 2 图片接口并收紧错误处理
- 新增 openai:image 路由、planner 与 finalize,内部通过 Codex responses image_generation tool 执行生图

- 补充 Codex OAuth/header 兼容、图片 success report 本地处理与相关前后端/集成测试

- 禁止 chat/completions 使用 gpt-image-2,图片接口限制 n=1,并移除 Provider 模型页的图片能力开关
2026-04-22 22:46:28 +08:00
AAEE86 ab82841426 feat(model-fetch): 对齐 Rust 上游模型抓取行为到 Python 语义
将 Rust 版上游可用模型抓取逻辑收敛到 Python 版行为,统一后台自动抓模
与管理员 provider-query 的模型发现路径,消除标准 /models、固定模型目录、
Antigravity、Vertex AI 等 provider 在两端实现上的分叉。

核心变更:
- 在 aether-model-fetch 中引入统一抓模策略层
- 覆盖标准 /models、Vertex API Key、Vertex Service Account、
  Antigravity fetchAvailableModels、固定模型目录五类抓模路径
- 将 provider-query 与后台自动抓模都切换到共享抓模入口,避免重复拼接
  URL、headers 和 provider 特判逻辑

标准模型抓取对齐:
- 按 Python 语义调整抓模优先级:
  openai:chat > openai:cli > openai:compact
  claude:chat > claude:cli
  gemini:chat > gemini:cli
- 从抓模候选中移除 openai:responses
- 为 openai:cli/openai:compact、claude:cli、gemini:* 补齐 Python 同款
  User-Agent / 浏览器指纹请求头
- Claude 抓模保留 after_id 分页语义
- Gemini 抓模统一为 v1beta/models?key=... 语义

provider-query 对齐:
- 返回结果改为按 model id 聚合,并合并/排序 api_formats
- 最终模型列表按 model id 排序,行为与 Python 保持一致
- 固定目录 provider(codex/kiro/claude_code/gemini_cli)不再依赖活跃
  endpoint,即使无 endpoint 也能返回预设模型目录
- Antigravity 多 key 查询改为按账户可用性 + tier 排序,首个成功结果即
  停止,并接入 provider 级缓存
- 仅配置 openai:responses 的 provider 不再被视为抓模成功路径

自动抓模对齐:
- 自动抓模成功时写入 allowed_models、upstream_models cache,并同步
  upstream_metadata
- upstream_metadata 合并逻辑对齐 Python,对 quota_by_model 做模型级合并,
  并保留已有 reset_time
- 自动抓模失败时不覆盖已有 allowed_models
- 固定目录 provider 在无 endpoint 场景下也可成功更新 allowed_models

Antigravity 对齐:
- 使用 POST /v1internal:fetchAvailableModels 抓取可用模型
- 按 Python 规则处理 URL fallback 和 429/404/408/5xx fallback 状态
- 强制要求 auth_config.project_id
- 过滤 Python 黑名单模型
- 解析并持久化 upstream_metadata.antigravity.quota_by_model

Vertex AI 对齐:
- API Key 模式仅抓取 publishers/google/models
- Service Account 模式新增 JWT token exchange,并按 Python region 顺序
  抓取 google + anthropic publishers
- 模型 owned_by / display_name / api_format 推断与 Python 对齐
- 软 404 处理行为与 Python 收敛

Gemini CLI / 固定目录对齐:
- Gemini CLI 改为返回 Python 预设模型目录
- 在可用时通过 loadCodeAssist 补充 plan_type/project_id 元数据
- Codex/Kiro/Claude Code 改为共享固定模型目录实现

测试:
- 扩展 aether-model-fetch 单元测试,覆盖格式优先级、openai:responses 排除、
  请求头、Claude 分页、Gemini query auth、固定目录与 metadata 合并
- 调整 provider-query 控制面测试到 Python 语义
- 新增自动抓模运行时测试,覆盖固定目录成功、Antigravity metadata 合并、
  失败保留旧 allowed_models

验证:
- cargo nextest run -p aether-model-fetch --lib
- cargo nextest run -p aether-gateway control::admin::provider_query model_fetch::runtime::tests
2026-04-12 10:14:29 +08:00
fawney19 b0b40c16ff feat: 全栈功能增强 - 扩展 provider/pool 管理、完善调度与数据层、重构前端 Pool 页面
后端:
- 扩展 pool_admin payloads 和 provider query models,增强 endpoint key 管理
- 完善 scheduler-core 候选排序与请求候选逻辑
- 增强 usage-runtime 写入、provider-transport 网络层与 OAuth 刷新
- 改进 AI pipeline 响应转换与流式处理
- 扩展 global_models/provider_catalog 数据层查询能力
- 增强 video-tasks-core 多 provider 支持
- 新增大量集成测试覆盖 pool/keys/provider_query/frontdoor

前端:
- 重构 PoolManagement 页面,拆分状态管理/对话框逻辑到独立模块
- 新增 poolAdvancedDialog/poolSchedulingDialog/poolManagementState/poolMobilePresentation 工具函数及测试
- 改进 Dialog 组件与 provider tabs 显示

部署:
- 更新 Rust CI workflow 和 Dockerfile 构建配置

Closes #275
Co-authored-by: AAEE86 <[email protected]>
2026-04-09 13:51:50 +08:00
fawney19 5d96d6673b refactor: 大规模模块拆分与代码精简,新增 ai-pipeline/data-contracts 独立 crate
- 新增 aether-ai-pipeline 和 aether-data-contracts crate,将 pipeline 逻辑与数据契约从 gateway 中解耦
- 重构 admin handlers:拆分单体模块为 auth/billing/endpoint/features/model/observability/provider/system 等独立子模块
- 合并 chat/cli 重复代码路径:精简 conversion、finalize、planner 中的 sync/chat/cli 分支
- 重构 scheduler/executor/data 层,引入 facade 模式降低模块间耦合
- 移除冗余的 intent 模块,将 plan_fallback/policy/stream_path/sync_path 迁移至 executor
- 前端适配:调整 admin API 调用和 provider 模型测试对话框
2026-04-07 02:50:19 +08:00
fawney19 763ff03a7b refactor: 拆分 gateway 单体为独立 crate,新增 systemd 部署方案
将 gateway 内部的 model-fetch、provider-transport、scheduler-core、
usage-runtime、video-tasks-core 模块提取为独立 crate;重构 gateway
内部模块结构(state/router/cache/data/query 等);移除大量遗留模块
文件;新增 systemd 二进制部署骨架及相关文档;更新前端 usage 相关
API 和组件。
2026-04-05 20:23:16 +08:00