fix(server): count quota by the key actually used, and more review fixes

Found by the PR review, each with a test that failed first:
- Quota: any key header skipped it, even one the provider never reads
  (x-aws-access-key-id with OpenAI), so a request ran on the server's
  key without being counted. The check now runs after the model is
  resolved and uses usesServerCredentials. On main already.
- usesServerCredentials read the raw base URL; "/" cleans up to none, so
  an Ollama request ran on the server's key past the server-model check.
- SGLang's default 127.0.0.1:8000 only fills the settings form. Chat and
  the model list used it as a real address, so the server called its own
  machine even with private URLs blocked. Now a base URL is required.
- With a user's OpenAI key and no base URL, the SDK read the server's
  OPENAI_BASE_URL. The official endpoint is now passed. On main already.
- The Test button refused nothing on the server's keys (Ollama Cloud),
  and a 15 s timeout reported "connected, no tool call".
- The model list for Ollama without a base URL came from ollama.com while
  chat went to the server's Ollama.
- Bedrock's "Too many tokens, please wait" counted as context too long.
- On the server's keys the provider's error text stays in the server log;
  it can name the server's AWS account, role or internal hosts.
- Desktop app: the preset keys are the user's own (NEXT_AI_DRAWIO_DESKTOP),
  so Max Output Tokens can be raised and keyless models in settings work
  again. A launch that found the remembered port taken no longer replaces
  it, which hid the user's chats and settings for good.
This commit is contained in:
dayuan.jiang
2026-10-04 23:04:21 +09:00
parent 504d2fa812
commit 0855b35ff2
14 changed files with 382 additions and 54 deletions
+44
View File
@@ -74,6 +74,50 @@ describe("POST /api/validate-model", () => {
expect(typeof data.responseTime).toBe("number")
})
it("reports a model that did not answer in time", async () => {
// The 15 s timeout has fired: the SDK ends the stream with an
// abort part instead of throwing
const timedOut = AbortSignal.abort(
new DOMException("The operation timed out.", "TimeoutError"),
)
const timeout = vi
.spyOn(AbortSignal, "timeout")
.mockReturnValue(timedOut)
vi.stubGlobal(
"fetch",
vi.fn(async () => {
throw timedOut.reason
}),
)
try {
const data = await testGlm()
expect(data.valid).toBe(false)
expect(data.code).toBe("timeout")
} finally {
timeout.mockRestore()
}
})
it("does not run on the server's keys", async () => {
process.env.OLLAMA_API_KEY = "server-ollama-key"
try {
const res = await validateModel(
new Request("http://localhost/api/validate-model", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
provider: "ollama",
modelId: "any-cloud-model",
}),
}),
)
expect(res.status).toBe(400)
expect((await res.json()).error).toMatch(/API key/)
} finally {
delete process.env.OLLAMA_API_KEY
}
})
it("warns when the model answers without a tool call", async () => {
streamReply({ role: "assistant", content: "OK" })
const data = await testGlm()