MCP preview after the server lost a session (it expired, or the MCP
process restarted):
- Every server state has an id, made when the state is created. The tab
notices a new id even when the version numbers happen to match, and
every push names the state it was based on, so one based on a lost state
is refused, also when it comes before the tab's first poll (the server
recovers the saved file first).
- The tab keeps the newest canvas XML, saved or not. When the server knows
nothing (no file) or exactly what the tab last saved, the canvas wins and
is saved, so edits made while the server was down are kept. Otherwise
the server's diagram (an AI write the tab missed, a cleared document
that was saved) is shown and the tab's copy goes to History.
- Late answers to an old state's push or poll are dropped; a failed push
says the server is unreachable; Download as .drawio saves the canvas.
Settings and server:
- Saved providers this version does not know stay in storage with their
keys, and sending no longer trips over them.
- The desktop "Ollama (Local)" preset with a key goes to local Ollama
again; a server model's Ollama URL variable is read; the admin panel
writes Ollama Cloud's URL for a key without one.
- Provider error texts show again in the desktop app and for EdgeOne.
- .env: a quoted value followed by a comment ending in a quote is read as
dotenv reads it; unquoted values are unchanged.
- Desktop app: the next launch opens the port where a chat was last
saved; a launch elsewhere that saves nothing does not move it, and a
page with no chats lets the next launch try the other port once.
- The Test button no longer stays busy after another tab changed the key.
- A completed append_diagram is no longer undone by an earlier failed
edit's preview; a file read once in vain is saved again once it is read
or gone.
From the first batch's review:
- The admin panel's Test of an entry without a URL now tests the server's
<P>_BASE_URL, where chat sends the entry's key; chat is unchanged (the
first fix rerouted working setups).
- The model list ends downloads that are too large, accepts answers
without a body, and keeps the "redirects are not allowed" explanation.
- A test covers the preview's History rendering.
- A call the server runs (get_shape_library) still reaches the browser's
tool handler, and it dropped the stored diagram of an earlier broken
edit, whose preview then stayed. Only the tools that draw take it now.
- A display_diagram whose final XML fails the checks loads the diagram
from before its preview again, as a failed edit does.
- After Stop, a tool result that arrives later (a screenshot check still
running) no longer sends a new request; Stop also skips calls the tool
handler already took.
- New chat and opening another chat kept nothing of a diagram drawn
without messages when it could not be saved; now they stay on it.
- The settings dialog drops a model list or test result whose provider
credentials changed meanwhile, also in another tab.
- A saved provider this version does not know crashed the whole page on
load; it is skipped.
- The input emptied a moment after the message showed in the chat, so it
briefly appeared twice (seen as a flaky e2e test).
- The desktop app's preset switch on the same port refetches the server
models instead of reloading the page, which lost unsent attachments.
Found by the second PR review:
- After a streamed edit, an older render of the stream stored the edit's
original diagram again, and the next failed request (no quota, a lost
connection) put that old diagram back on the canvas. The tool handler
now marks its call as handled, so the preview code leaves it alone.
- An edit applied before the UI showed an earlier broken edit's error was
erased when that error undid its preview, or was built on that preview.
The handler now starts from the diagram before all unhandled previews,
and reads the diagram state that updates at once.
- A failed or stopped display_diagram left its half drawn diagram on the
canvas. Its preview is undone now, like an edit's.
- "New chat" cleared a chat that could not be saved (storage full).
- The settings dialog showed a model list, a fetch error or a test result
on the provider that was opened after the request started, and marked a
model id changed during the test as tested.
- A tool call with broken JSON was shown as cut off by the output limit.
- History entries and session thumbnails could pair with a later diagram
when draw.io answered an export late.
- A server model saved before non-ASCII provider names got into the id
was reset to the default model.
Found by the PR review, each with a test that failed first:
- get_diagram during a page export returned the one-page projection on
screen as the whole document (6 of 6 times when timed so). The preview
page no longer answers a sync while a projection shows, and syncs after
reloading, so the poll that restores the real document exports it.
- Exports are numbered on the server too: a late result of an export that
timed out was saved as the next export's file.
- In Chrome, a new_xml with a syntax error counted the <parsererror>
element as a second cell, so the web app rejected edits that auto-fix
repairs ("must contain exactly one cell").
- hasCells missed single-quoted ids, so screenshot_diagram called such a
diagram empty and auto-save never created its file.
- A literal \n directly under a <diagram> that has a model passed
validation; only text-only pages are compressed data.
- A wrapped mxCell repeating its UserObject's id took the wrapper's place
in edits, so delete and update left an empty or nested wrapper.
- Bare cells with a shape or edge id of "0" or "1" are rejected with a
clear message instead of being renamed, which broke their edges.
- DRAWIO_DATA_DIR expands ~, which JSON configs pass on as it is.
Found by the PR review:
- The built-in examples showed a finished card and an empty canvas. They
are answered in the browser, never reach the tool handler, and relied
on the final redraw that an earlier commit removed. The example branch
now loads its diagram itself.
- When the request failed while an edit was streaming (a provider error,
a lost connection), its preview stayed on the canvas. The error handler
now restores the diagram from before the preview.
- The model picker could not scroll with the wheel or touch: the settings
dialog blocks those events outside itself, and the picker is rendered
outside it. The popover is modal now.
- A fetch error and the open picker stayed when switching providers.
- Editing a model id kept the old test warning and response time, which
also hid the "may not be able to draw" hint for the new id.
- Claude Opus 5.5 sent an edit with invalid JSON, then the same edit
again. The first call's streamed preview was never undone: its input
has no operations, and the undo sat behind that check. The second edit
then started from the preview, failed on a duplicate id, and the model
had to try a third time. The undo now runs first, and an edit that
starts in the same render uses the undone diagram.
- The SDK passes an invalid tool call's error as a string, which was
wrapped as a provider error. streamErrorText keeps it as the text the
model reads.
- Bedrock's "on-demand throughput isn't supported" gets the model id hint.
- The thinking header uses the page language ("Thought for 1 second" in
English), from the dictionary entries that were already there.
- An error object sent inside the stream (OpenRouter's { code, message })
showed as "[object Object]"; its message and status code are read now.
- A problem+json "detail" is added to the message: NVIDIA only said
"Gone" for a retired model. 410 counts as model not found.
- "Cannot connect to API" from the SDK gets the connection hint.
- Text that is only whitespace (Kimi K2.6 sends a space before a tool
call) no longer shows an empty bubble.
- allowSystemInMessages stops the warning on every request. Our system
messages carry cache points; a client's own system messages are already
dropped by the empty-content filter.
The "Fetch models" button asks the provider for its models (OpenAI-style
/models, Anthropic, Google, Ollama, OpenRouter, Vercel Gateway, AIHubMix)
and shows them in a searchable picker. This replaces the route that only
worked for AIHubMix.
A snapshot of models.dev (MIT) says which models support tool calls.
Models without them get a "no tool calls" badge in the picker and a hint
in the model list, since drawing needs tool calls. Refresh the snapshot
with scripts/update-model-catalog.mjs.
- lib/llm-errors.ts sorts an error into about a dozen kinds (key
rejected, no access, unknown model, no credit, rate limited, context
too long, no image input, no tool calls, output cut off, provider down,
cannot connect, timeout): first texts that name the cause precisely,
then the HTTP status code, then general texts. It unwraps RetryError and
hides keys and Bearer tokens in the provider's message
- The chat route uses it for errors before the stream and, through
toUIMessageStreamResponse's onError, for errors in the stream. Errors
of the model's own tool call stay as they are: the same text goes back
to the model so it can fix the call
- The chat shows the hint in the user's language, then the provider's
message; a rejected key, missing access or unknown model adds an "Open
model settings" button. The Test button shows the same hints
- Fixes: our message "API key is required when using a custom base URL"
was replaced by "Authentication failed" because it contains "key"; a
provider's "Rate limit exceeded" opened this site's quota toast; an
error body like {"error": ...} was shown as raw JSON; the Test button
matched "401" in the message, where providers rarely put it
- Remove the string matching fallbacks in the chat panel
- A "Get API key" link next to the API Key field for the 19 providers
that have a key page (from env.example and the providers' docs). 17
answered 200 to curl; OpenAI's is behind a Cloudflare challenge and
DeepSeek's behind a regional block, both checked in Chrome
- Base URLs drop spaces, trailing slashes and a pasted endpoint path
(/chat/completions, /completions, /messages, /responses), which the
SDK would otherwise append a second time and get a 404. getAIModel does
this for the chat and the Test button; the field does it on blur and
shows the URL requests go to
- getAIModel resolves credentials (client key, server env vars, the
existing SSRF rules) and createModel builds the model by SDK. The
provider-by-provider switch shrinks from 24 cases to the few that
differ (lib/ai-providers.ts 1531 -> 1106 lines)
- /api/validate-model calls getAIModel instead of its own 24-case switch
(503 -> 175 lines), which had drifted from the chat: it built Azure
with createOpenAI, Kimi and MiMo with createOpenAI instead of
createDeepSeek, and the official OpenAI endpoint with Chat Completions.
A passing test now means the chat works
- Plain OpenAI-compatible providers (SiliconFlow, SGLang, ModelScope,
GLM, Qwen, Qiniu, Novita, Atlas Cloud, EdgeOne, Doubao, MiniMax in
OpenAI mode, AIHubMix on a custom URL) use @ai-sdk/openai-compatible,
which reads reasoning_content, so their reasoning shows, and accepts
SGLang's stream as is (its 95-line stream rewrite is gone).
includeUsage keeps token usage for quotas. <think> tags in their text
become reasoning (extractReasoningMiddleware)
- SGLang without a base URL used OpenAI's endpoint; it now defaults to
http://127.0.0.1:8000/v1 like the Test button did
- Chat requests to a client base URL refuse redirects, as the Test
button already did (redirectGuardedFetch moves to lib/ssrf-protection)
- The Test button streams like the chat (the ModelScope special case is
gone), times out after 15 s, does not retry, asks the model to call a
ping tool and warns when it answers without one, and tests all models
at once. The time each test took shows on its check mark
- Unknown provider names are rejected with Object.hasOwn, and the error
texts list providers from PROVIDER_INFO instead of hand-kept lists
- edit_diagram runs the MCP server's editDiagram: every new_xml is checked
first, one cell per operation, and after the edit only the target page
is checked, rejecting only errors this edit introduced. An unrelated
problem elsewhere in the document no longer blocks every edit. The
error lists each failed operation
- The streaming edit preview uses the MCP applyDiagramOperations
- Delete applyDiagramOperations (292 lines) and wrapWithMxFile from
lib/utils.ts, and the unused hand-copied scripts/test-diagram-operations.mjs
- One blank document (BLANK_MXFILE) for the web app and the MCP preview,
replacing four copies
- Saving a .drawio wraps a bare model with normalizeToMxfile
- The empty-diagram check uses hasCells, which also counts cells wrapped
in a UserObject/object
- DiagramOperation is the MCP type
- The wrapped-cell and empty-diagram tests now run against the MCP code
- New e2e test: edit_diagram changes the canvas, and a failing edit
leaves it as it was
- Delete the web app's own copy of the XML checks and repairs from
lib/utils.ts (1,074 lines). loadDiagram now uses the MCP server's
validateAndFixXml without the strict checks, because the XML may hold
the user's own diagram
- display_diagram and append_diagram prepare the model's XML with the new
shared prepareNewDiagram, also used by the MCP create_new_diagram: wrap,
validate strictly and auto-fix while it is still a bare model (where
duplicate ids are renamed), then turn it into an mxfile
- The streaming preview of display_diagram no longer redraws the model's
raw cells after the tool handler loaded the checked diagram, and drops
a queued preview once the input is complete. That redraw lost
auto-fixes and UserObject/object wrappers, so a linked cell lost its
label; it also showed a second error toast
- The web repair regression tests now run against the MCP functions
- New e2e test checks the canvas content after display_diagram
- Fix the e2e upload tests, whose file input locator also matched the
template import input
* fix: raise the output budget so reasoning models reach the tool call
A reasoning model spends the output budget in order: thinking first, then prose,
then the tool call. With 16000 the thinking alone can consume all of it, so the
turn ends with finishReason "length" before display_diagram is ever called. The
canvas stays empty and nothing surfaces in the UI, because no tool call means no
tool error, and the client never reads finishReason.
Measured on openrouter deepseek/deepseek-v4-flash, the model from the report:
- max_tokens=800 with reasoning on returns reasoning_tokens=800, empty content,
finish_reason length. So reasoning is billed against this budget, not exempt.
- refining an existing diagram (19k chars of XML in the input) produced 49142
chars of reasoning, zero tool calls, finishReason "length" at 16000
- the same request at 40000 finished and called edit_diagram with 12 operations
64000 cannot just be sent to every model: bedrock claude-3-haiku caps at 4096,
nova-lite at 10000, and the openrouter deepseek-r1 endpoint counts input and
output against one 64000 ceiling. All three name the real limit in the 400, so
parse it and retry once. Verified: nova-lite logs "64000 rejected, retrying with
10000" and then completes its tool call.
Also expose the budget in Settings. It is sent as a header rather than read from
env only, so desktop users can raise it themselves without an env file.
vercel.json goes back to the 300s it had before #238 traded it for $2-4/month.
That is now Vercel's own default, and billing pauses while the function waits on
the model, so the saving that motivated 120s no longer applies. edgeone.json is
left alone: its 120 may be that platform's actual ceiling.
* fix: only reinterpret an error as a budget rejection when it says so
Review of the first commit found the retry could fire on errors that have
nothing to do with the budget, which would replace a readable provider error
with a truncated response: exactly the symptom this PR exists to remove.
- Drop the generic "lower than N" pattern. For the Bedrock message it was dead
code, since "model limit of N" matches first with the same number. Left live,
it would read a number out of any message shaped like "must be lower than 2".
- Skip errors whose status is not 400 or 422, so auth and rate-limit failures
are never reinterpreted.
- Require the parsed ceiling to be at least 1024. Below that a diagram cannot
come out whole, so retrying would hide the error behind broken XML.
- Validate MAX_OUTPUT_TOKENS from env the same way as the header, so a stray
"-1" falls back instead of reaching the provider.
Adds tests for the retry wrapper itself, which had none: it retries once with
the named ceiling, leaves a 401 alone, does not retry when the ceiling is not
smaller, propagates a second rejection, and preserves the other call options.
Re-verified against the live APIs: bedrock nova-lite still logs "64000 rejected,
retrying with 10000" and completes its tool call, and deepseek-v4-flash still
finishes normally at 64000.
* fix(e2e): resolve strict mode violation in iframe test
Use .first() with [title*="Diagram"] selector to avoid matching multiple elements.
Fixes CI failure in E2E Tests job.
* fix(e2e): use .first() to resolve strict mode violation
* style: fix biome formatting in iframe test
* style: fix biome formatting in iframe test
* fix(e2e): use .or().first() to handle both text and title selectors
* fix(e2e): increase timeout for draw.io toolbar visibility check
* fix(e2e): filter visible elements to avoid selecting hidden toolbar
The Lint & Unit Tests PR status check is failing
at the "Run lint" step over a half-dozen issues.
This is causing all PR's to fail the Lint & Unit tets check.
This fix resolves those issues.
Signed-off-by: Bryon Nevis <[email protected]>
* test: add Vitest and Playwright testing infrastructure
- Add Vitest for unit tests (39 tests)
- cached-responses.test.ts
- ai-providers.test.ts
- chat-helpers.test.ts
- utils.test.ts
- Add Playwright for E2E tests (3 smoke tests)
- Homepage load
- Japanese locale
- Settings dialog
- Add CI workflow (.github/workflows/test.yml)
- Add vitest.config.mts and playwright.config.ts
- Update .gitignore for test artifacts
* test: add more E2E tests for UI components
- Chat panel tests (interactive elements, iframe)
- Settings tests (dark mode, language, draw.io theme)
- Save dialog tests (buttons exist)
- History dialog tests
- Model config tests
- Keyboard interaction tests
- Upload area tests
Total: 15 E2E tests, all passing
* test: fix E2E test issues from review
Fixes based on Gemini and Codex review:
- Remove brittle nth(1) selector in keyboard tests
- Remove waitForTimeout(500) race condition
- Remove if(isVisible) silent skip patterns
- Add proper assertions instead of no-op checks
- Remove expect(count >= 0) that always passes
- Remove unused hasProviderUI variable
All 14 E2E tests and 39 unit tests pass.
* style: auto-format with Biome
* fix: resolve lint errors for CI
* test(e2e): add diagram generation tests with mocked AI responses
- Add tests for generate, edit, and append diagram operations
- Use SSE mocked responses matching AI SDK UI message stream format
- Generate mxCell XML directly in tests for deterministic assertions
- Tests verify tool card rendering and 'Complete' badge state
* test: add comprehensive E2E tests for all major features
- Error handling tests (API errors, rate limits, network timeout, truncated XML)
- Multi-turn conversation tests (sequential requests, history preservation)
- File upload tests (upload button, file preview, sending with message)
- Theme switching tests (dark mode toggle, persistence, system preference)
- Language switching tests (EN/JA/ZH, persistence, locale URLs)
- Iframe interaction tests (draw.io loading, toolbar, diagram rendering)
- Copy/paste tests (chat input, XML input, special characters)
- History restore tests (new chat, persistence, browser navigation)
* refactor: extract shared test helpers and improve error assertions
- Create tests/e2e/lib/helpers.ts with shared SSE mock functions
- Add proper error UI assertions to error-handling.spec.ts
- Remove waitForTimeout calls in favor of real assertions
- Update 6 test files to use shared helpers
* docs: add testing section to CONTRIBUTING.md
* fix: improve test infrastructure based on PR review
- Fix double build in CI: remove redundant build from playwright webServer
- Export chat helpers from shared module for proper unit testing
- Replace waitForTimeout with explicit waits in E2E tests
- Add data-testid attributes to settings and new chat buttons
- Add list reporter for CI to show failures in logs
- Add Playwright browser caching to speed up CI
- Add vitest coverage configuration
- Fix conditional test assertions to use test.skip() instead of silent pass
- Remove unused variables flagged by linter
* fix: improve E2E test assertions and remove silent skips
- Replace silent test.skip() with explicit conditional skips
- Add actual persistence assertion after page reload
- Use data-testid selector for new chat button test
* refactor: add shared fixtures and test.step() patterns
- Add tests/e2e/lib/fixtures.ts with shared test helpers
- Add tests/e2e/fixtures/diagrams.ts with XML test data
- Add expectBeforeAndAfterReload() helper for persistence tests
- Add test.step() for better test reporting in complex tests
- Consolidate mock helpers into fixtures module
- Reduce code duplication across 17 test files
* fix: make persistence tests more reliable
- Remove expectBeforeAndAfterReload from mocked API tests
- Add explicit test.step() for before/after reload checks
- Add retry config for flaky clipboard tests
- Add sleep after reload for language persistence test
* test: remove flaky XML paste test
* docs: run both unit and e2e tests before PR
* chore: add type check and unit test git hooks
---------
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>