mirror of
https://github.com/DayuanJiang/next-ai-draw-io.git
synced 2026-10-04 08:47:45 +08:00
fix(chat): close credential leaks and harden the chat route
- Vertex: a client-supplied base URL only works with the client's own Vertex key - Accept only data: URLs for file parts in every message, so the server never downloads them - Output budget retry accounts for the thinking budget Bedrock/Anthropic add, and reads Volcengine, DashScope, SGLang and vLLM rejections; falls back to 16000 once - x-max-output-tokens can only lower the budget on server credentials - On server credentials only server models or AI_MODEL entries can be used - Drop tool results together with the invalid tool calls they belong to - Count quota tokens as input + output (cached tokens were counted twice) - Private-URL check for custom base URLs, end Langfuse traces on error/abort/early return - Fix repairToolCall ordering and placeholder, align edit_diagram prompt with operations - Panel Bedrock keys are read from ADMIN_AWS_*; forward the access code to EdgeOne - isMinimalDiagram only treats root cells as an empty canvas
This commit is contained in:
+2
-1
@@ -12,7 +12,8 @@ AI_PROVIDER=bedrock
|
||||
AI_MODEL=global.anthropic.claude-sonnet-4-5-20250929-v1:0
|
||||
|
||||
# Output limit, all providers (default: 64000). Shared by reasoning and the diagram XML,
|
||||
# so a thinking model can spend it all before the tool call. Users can override it in Settings.
|
||||
# so a thinking model can spend it all before the tool call. Users can lower it in Settings,
|
||||
# and raise it only when they use their own API key, so this also caps cost on server keys.
|
||||
# If a model's own ceiling is lower, the request is retried with that ceiling automatically.
|
||||
# MAX_OUTPUT_TOKENS=64000
|
||||
|
||||
|
||||
Reference in New Issue
Block a user