fix: 治理 Prometheus 指标基数爆炸和内存缓存无界增长

- 移除 token/latency Prometheus 指标的 model 标签,避免 provider x model 笛卡尔积
- HealthMonitor 滑动窗口从 DB JSON 迁移至进程内存,减少写放大
- ModelCostService 三层缓存增加 500 条上限,超限时清空
- StickyPriority 粘性缓存和健康状态字典增加容量淘汰
- AffinityManager 请求锁字典增加 500 条上限,淘汰空闲锁
- 配额刷新/探测查询使用 defer/load_only 避免加载大 JSON 列
- Alembic 迁移清理 DB 中遗留的 request_results_window 数据
- 同步更新测试适配 batch_get_cooldowns 返回值和批量删除异步化
This commit is contained in:
fawney19
2026-03-10 10:40:50 +08:00
parent c9f0685b40
commit 7c580e843f
20 changed files with 241 additions and 85 deletions

View File

@@ -53,7 +53,7 @@ async def test_create_endpoint_injects_default_body_rules_when_missing(
monkeypatch.setattr(
routes,
"get_default_body_rules_for_endpoint",
lambda _fmt: [{"action": "drop", "path": "max_output_tokens"}],
lambda _fmt, **_kw: [{"action": "drop", "path": "max_output_tokens"}],
)
db = _FakeDB(
@@ -81,7 +81,7 @@ async def test_create_endpoint_keeps_user_body_rules_when_provided(
monkeypatch.setattr(
routes,
"get_default_body_rules_for_endpoint",
lambda _fmt: [{"action": "drop", "path": "max_output_tokens"}],
lambda _fmt, **_kw: [{"action": "drop", "path": "max_output_tokens"}],
)
user_rules = [{"action": "set", "path": "metadata.source", "value": "user"}]