mirror of
https://github.com/DayuanJiang/next-ai-draw-io.git
synced 2026-10-11 12:09:53 +08:00
fix(server): count quota by the key actually used, and more review fixes
Found by the PR review, each with a test that failed first: - Quota: any key header skipped it, even one the provider never reads (x-aws-access-key-id with OpenAI), so a request ran on the server's key without being counted. The check now runs after the model is resolved and uses usesServerCredentials. On main already. - usesServerCredentials read the raw base URL; "/" cleans up to none, so an Ollama request ran on the server's key past the server-model check. - SGLang's default 127.0.0.1:8000 only fills the settings form. Chat and the model list used it as a real address, so the server called its own machine even with private URLs blocked. Now a base URL is required. - With a user's OpenAI key and no base URL, the SDK read the server's OPENAI_BASE_URL. The official endpoint is now passed. On main already. - The Test button refused nothing on the server's keys (Ollama Cloud), and a 15 s timeout reported "connected, no tool call". - The model list for Ollama without a base URL came from ollama.com while chat went to the server's Ollama. - Bedrock's "Too many tokens, please wait" counted as context too long. - On the server's keys the provider's error text stays in the server log; it can name the server's AWS account, role or internal hosts. - Desktop app: the preset keys are the user's own (NEXT_AI_DRAWIO_DESKTOP), so Max Output Tokens can be raised and keyless models in settings work again. A launch that found the remembered port taken no longer replaces it, which hid the user's chats and settings for good.
This commit is contained in:
+19
-3
@@ -75,6 +75,19 @@ async function getJson(
|
||||
return response.json()
|
||||
}
|
||||
|
||||
/**
|
||||
* Where to list from without the user's base URL: where chat goes then. For
|
||||
* Ollama that is the server's Ollama, else the SDK's local default; a local
|
||||
* default in PROVIDER_INFO (SGLang's) only fills the settings form.
|
||||
*/
|
||||
function listFallbackUrl(provider: ProviderName): string {
|
||||
if (provider === "ollama") {
|
||||
return process.env.OLLAMA_BASE_URL || "http://127.0.0.1:11434/api"
|
||||
}
|
||||
const url = PROVIDER_INFO[provider].defaultBaseUrl
|
||||
return url?.startsWith("https://") ? url : ""
|
||||
}
|
||||
|
||||
/**
|
||||
* The provider's chat models, with tool support from the provider's own
|
||||
* data or else models.dev. Only the client's key is used, so the server's
|
||||
@@ -85,9 +98,7 @@ export async function listProviderModels(
|
||||
{ apiKey, baseUrl }: { apiKey?: string; baseUrl?: string },
|
||||
fetchFn: typeof fetch = fetch,
|
||||
): Promise<ListedModel[]> {
|
||||
const base = normalizeBaseUrl(
|
||||
baseUrl || PROVIDER_INFO[provider].defaultBaseUrl || "",
|
||||
)
|
||||
const base = normalizeBaseUrl(baseUrl || listFallbackUrl(provider))
|
||||
const bearer: Record<string, string> = apiKey
|
||||
? { Authorization: `Bearer ${apiKey}` }
|
||||
: {}
|
||||
@@ -170,6 +181,11 @@ export async function listProviderModels(
|
||||
break
|
||||
}
|
||||
default: {
|
||||
if (!base) {
|
||||
throw new Error(
|
||||
`${PROVIDER_INFO[provider].label} needs a base URL to list its models.`,
|
||||
)
|
||||
}
|
||||
const data = await getJson(`${base}/models`, bearer, fetchFn)
|
||||
models = (data.data ?? [])
|
||||
.map((m: { id: string }) => ({ id: m.id }))
|
||||
|
||||
Reference in New Issue
Block a user