Found by the second PR review:
- With AWS_BEARER_TOKEN_BEDROCK set on the server, a request with the
user's AWS keys ran on the server's token: the Bedrock SDK prefers it.
Checked with Bedrock: invalid user keys used to get an answer.
- An OpenAI key with the official URL filled in (the settings form does
that) went to the Responses API. Back to main's rule: a configured base
URL uses Chat Completions.
- A user's Ollama key went to the server's OLLAMA_BASE_URL, for chat and
for the model list. Like every other provider, it goes to the user's
base URL or Ollama Cloud.
- The server's keyless Ollama and EdgeOne were not counted in the quota.
- AI_MODEL models on the server's keys ran on any provider with a server
key, not only on AI_PROVIDER.
- A user's Azure key without a base URL used the server's resource name.
- The admin panel's Test button failed whenever access codes were set.
- DeepSeek's errors in the stream (plain text) were shown as they were,
without a hint and also on the server's keys. Bedrock's throttling in
the stream was not recognised as a rate limit.
- The EdgeOne function accepted text/plain; x=application/json, which
other sites can send without a CORS preflight.
- Desktop app: a launch that found the old port taken for a moment (the
previous version still quitting after an update) remembered the new
port for good. The new port is kept only when Windows reserves the old
one. A failed read of the presets file moved it aside as corrupt, and a
save could then replace the presets. Switching presets on the same port
now reloads the page. The dev launcher no longer misses a preset change
made before or during a restart.
Found by the PR review, each with a test that failed first:
- Quota: any key header skipped it, even one the provider never reads
(x-aws-access-key-id with OpenAI), so a request ran on the server's
key without being counted. The check now runs after the model is
resolved and uses usesServerCredentials. On main already.
- usesServerCredentials read the raw base URL; "/" cleans up to none, so
an Ollama request ran on the server's key past the server-model check.
- SGLang's default 127.0.0.1:8000 only fills the settings form. Chat and
the model list used it as a real address, so the server called its own
machine even with private URLs blocked. Now a base URL is required.
- With a user's OpenAI key and no base URL, the SDK read the server's
OPENAI_BASE_URL. The official endpoint is now passed. On main already.
- The Test button refused nothing on the server's keys (Ollama Cloud),
and a 15 s timeout reported "connected, no tool call".
- The model list for Ollama without a base URL came from ollama.com while
chat went to the server's Ollama.
- Bedrock's "Too many tokens, please wait" counted as context too long.
- On the server's keys the provider's error text stays in the server log;
it can name the server's AWS account, role or internal hosts.
- Desktop app: the preset keys are the user's own (NEXT_AI_DRAWIO_DESKTOP),
so Max Output Tokens can be raised and keyless models in settings work
again. A launch that found the remembered port taken no longer replaces
it, which hid the user's chats and settings for good.
- getAIModel resolves credentials (client key, server env vars, the
existing SSRF rules) and createModel builds the model by SDK. The
provider-by-provider switch shrinks from 24 cases to the few that
differ (lib/ai-providers.ts 1531 -> 1106 lines)
- /api/validate-model calls getAIModel instead of its own 24-case switch
(503 -> 175 lines), which had drifted from the chat: it built Azure
with createOpenAI, Kimi and MiMo with createOpenAI instead of
createDeepSeek, and the official OpenAI endpoint with Chat Completions.
A passing test now means the chat works
- Plain OpenAI-compatible providers (SiliconFlow, SGLang, ModelScope,
GLM, Qwen, Qiniu, Novita, Atlas Cloud, EdgeOne, Doubao, MiniMax in
OpenAI mode, AIHubMix on a custom URL) use @ai-sdk/openai-compatible,
which reads reasoning_content, so their reasoning shows, and accepts
SGLang's stream as is (its 95-line stream rewrite is gone).
includeUsage keeps token usage for quotas. <think> tags in their text
become reasoning (extractReasoningMiddleware)
- SGLang without a base URL used OpenAI's endpoint; it now defaults to
http://127.0.0.1:8000/v1 like the Test button did
- Chat requests to a client base URL refuse redirects, as the Test
button already did (redirectGuardedFetch moves to lib/ssrf-protection)
- The Test button streams like the chat (the ModelScope special case is
gone), times out after 15 s, does not retry, asks the model to call a
ping tool and warns when it answers without one, and tests all models
at once. The time each test took shows on its check mark
- Unknown provider names are rejected with Object.hasOwn, and the error
texts list providers from PROVIDER_INFO instead of hand-kept lists