Commit Graph
17 Commits
Author SHA1 Message Date
dayuan.jiang c75f74a5a0 fix: what the third round broke, and the first batch's review
MCP preview after the server lost a session (it expired, or the MCP
process restarted):
- Every server state has an id, made when the state is created. The tab
  notices a new id even when the version numbers happen to match, and
  every push names the state it was based on, so one based on a lost state
  is refused, also when it comes before the tab's first poll (the server
  recovers the saved file first).
- The tab keeps the newest canvas XML, saved or not. When the server knows
  nothing (no file) or exactly what the tab last saved, the canvas wins and
  is saved, so edits made while the server was down are kept. Otherwise
  the server's diagram (an AI write the tab missed, a cleared document
  that was saved) is shown and the tab's copy goes to History.
- Late answers to an old state's push or poll are dropped; a failed push
  says the server is unreachable; Download as .drawio saves the canvas.

Settings and server:
- Saved providers this version does not know stay in storage with their
  keys, and sending no longer trips over them.
- The desktop "Ollama (Local)" preset with a key goes to local Ollama
  again; a server model's Ollama URL variable is read; the admin panel
  writes Ollama Cloud's URL for a key without one.
- Provider error texts show again in the desktop app and for EdgeOne.
- .env: a quoted value followed by a comment ending in a quote is read as
  dotenv reads it; unquoted values are unchanged.
- Desktop app: the next launch opens the port where a chat was last
  saved; a launch elsewhere that saves nothing does not move it, and a
  page with no chats lets the next launch try the other port once.
- The Test button no longer stays busy after another tab changed the key.
- A completed append_diagram is no longer undone by an earlier failed
  edit's preview; a file read once in vain is saved again once it is read
  or gone.

From the first batch's review:
- The admin panel's Test of an entry without a URL now tests the server's
  <P>_BASE_URL, where chat sends the entry's key; chat is unchanged (the
  first fix rerouted working setups).
- The model list ends downloads that are too large, accepts answers
  without a body, and keeps the "redirects are not allowed" explanation.
- A test covers the preview's History rendering.
2026-10-05 17:33:36 +09:00
dayuan.jiang 4731394f32 fix(security): check request sources, regions and endpoints
- Bedrock: a request's AWS region must be a region name. It becomes part
  of the endpoint's host name, so a value such as
  "us-east-1.attacker.example/" sent the server's bearer token or signed
  request to another host.
- MCP preview server: only the preview page itself (Origin equal to the
  Host) or a non-browser client may call it; a page on another localhost
  port could replace the diagram with a plain text POST. History builds
  its thumbnails element by element and shows only SVG data images, so a
  stored value can no longer run script in the preview.
- chat, validate-model, validate-diagram, provider-models and parse-url
  take JSON bodies only, so another website cannot make the user's own
  server (the desktop app, a local install) run models with their keys;
  the desktop app also refuses a foreign Host (DNS rebinding).
- The model list reads at most 2 MB, also through the Gateway SDK, and
  answers only with its own error texts: the URL is the caller's and may
  be an internal address.
- An admin panel provider with its own key and no URL no longer inherits
  the global <P>_BASE_URL, which may be a proxy for another key; OpenAI
  then gets the official endpoint, as its Test. Azure keeps the server's
  resource.
2026-10-05 17:02:06 +09:00
dayuan.jiang cfbf17f97e fix(mcp-server): recover sessions in one place, and more fixes from the third review
- A session whose state expired was recovered from its auto-save file only
  for the preview page; the tools built on their older copy and then
  overwrote the file. They now recover it first (restoreSavedSession).
- A preview tab that missed the last AI write pushed its older diagram
  over the recovered one after a restart. It now shows the recovered
  diagram and keeps its own copy in History.
- An empty record of what the model has seen (after load_diagram or a
  page tool on unseen changes) no longer lets one page of a multi-page
  document count for all, and get_diagram counts a page only once found.
- History: a thumbnail goes only to the entry it shows, the cached image
  never belongs to an older diagram, a re-serialized copy adds no entry,
  and a cleared document with its own pages is kept before a restore.
- The root cell id check reads attributes one by one: rack-id="1" or id
  text inside a label no longer counts.
- A compressed page counts as having cells; a saved file that could not
  be read is never written over.
2026-10-05 13:04:32 +09:00
dayuan.jiang 080f44716f fix(mcp-server): count a one-page view only for that page, and more review fixes
Found by the second PR review:
- get_diagram with a page selector, or a rejected edit's error, counted
  the whole document as seen, so an edit on another page could overwrite
  the user's change there. A one-page view now counts for all pages only
  if the others are unchanged; otherwise the reply says to get them.
- add_page accepted shapes with the root cell ids "0" and "1" and renamed
  them, breaking their edges. The check also missed ids on UserObject
  wrappers and ids written with spaces around the "=".
- Root cells written over two lines were kept as an extra layer, cells
  with id = "a" did not count as cells, and CDATA text before a page's
  model passed the check although draw.io cannot open the page.
- Auto-save cleanup deleted the user's own files that start with mcp-.
  Only names in the session id format are removed now.
- Restoring a history entry dropped edits made in the browser since the
  last entry. They are added to history first.
- A session whose state expired showed a blank page, and the next change
  overwrote its auto-save file. The saved file is loaded instead.
- An edit on a page export's one-page projection, made before the real
  document was back, replaced the whole document.
- A late sync reply could overwrite a newer edit: each sync export is
  numbered, and the server ignores replies older than the current state.
- screenshot_diagram could return another session's image after
  start_session ran during its retries.
2026-10-05 10:52:20 +09:00
dayuan.jiang 504d2fa812 fix(mcp-server): keep both pages when get_diagram meets a page export, and more review fixes
Found by the PR review, each with a test that failed first:
- get_diagram during a page export returned the one-page projection on
  screen as the whole document (6 of 6 times when timed so). The preview
  page no longer answers a sync while a projection shows, and syncs after
  reloading, so the poll that restores the real document exports it.
- Exports are numbered on the server too: a late result of an export that
  timed out was saved as the next export's file.
- In Chrome, a new_xml with a syntax error counted the <parsererror>
  element as a second cell, so the web app rejected edits that auto-fix
  repairs ("must contain exactly one cell").
- hasCells missed single-quoted ids, so screenshot_diagram called such a
  diagram empty and auto-save never created its file.
- A literal \n directly under a <diagram> that has a model passed
  validation; only text-only pages are compressed data.
- A wrapped mxCell repeating its UserObject's id took the wrapper's place
  in edits, so delete and update left an empty or nested wrapper.
- Bare cells with a shape or edge id of "0" or "1" are rejected with a
  clear message instead of being renamed, which broke their edges.
- DRAWIO_DATA_DIR expands ~, which JSON configs pass on as it is.
2026-10-04 23:04:21 +09:00
dayuan.jiang 6f5f7b668b fix(mcp-server): reject text between tags, which draw.io cannot open
draw.io reads any text inside a page as compressed page data, so a stray
text node makes the whole page fail with an atob error. gpt-5-mini sends
new cells with a literal "\n" between the tags; the edit card said
Complete while draw.io showed the error and kept the old diagram.

Validation now reports text between tags, and auto-fix turns a literal
\n, \t or \r between tags into whitespace. Other text goes back to the
model as an error. The compressed data directly under <diagram> is fine.
2026-10-04 20:15:44 +09:00
dayuan.jiang b7c543ca70 refactor(mcp-server): make the diagram modules usable from the web app
The web app will reuse the MCP server's XML engine instead of its own
copy in lib/utils.ts, so these modules now run in the browser too.

- Relative imports end in .ts, rewritten to .js by tsc
  (rewriteRelativeImportExtensions); Next.js resolves them directly
- Every module uses the global DOMParser/XMLSerializer: native in the
  browser, linkedom in Node via installDomPolyfill. pages.ts parsed with
  linkedom but serialized with the global serializer, which throws in
  the browser
- The saxes syntax check moves to xml-syntax.ts, so the browser does not
  pull in linkedom; it now also rejects undeclared prefixes such as
  xlink:, as the browser does
- Page decompression uses pako and atob instead of node:zlib and Buffer
- hasCells moves to pages.ts, away from the file system code
- The duplicate cell id check counts UserObject/object ids
- wrapCellsInModel drops comments and text before the first cell, which
  the web app accepts today
- validateAndFixXml takes { strict: false } for diagrams with user content
- Web tests run these modules with a browser DOM (jsdom)
- saxes becomes a direct dependency of the web app
2026-10-04 12:25:53 +09:00
dayuan.jiang 899924ba98 fix(mcp-server): fix duplicate page exports and auto-save deleting user files
- Preview page: keep an MCP export open until the server has its result.
  A poll answered before that still saw the request and started the same
  export again, so a parallel page export could write the previous
  page's image into its file
- Auto-save only removes its own mcp-*.drawio files, so a DRAWIO_DATA_DIR
  that also holds the user's diagrams keeps them
- screenshot_diagram captures a page that has no id attribute by loading
  just that page, like export_diagram
- An empty <Array as="points"/> no longer hides orphan mxPoints that
  come after it
- POST /api/state refuses a push without xml, which used to wipe the
  stored diagram
- Clear exportOptions when an export ends, reuse hasCells for the empty
  diagram check, and reword two log lines
2026-10-04 07:38:17 +09:00
dayuan.jiang f52f95025f refactor(mcp-server): move the preview page into src/preview
The 580-line page template in http-server.ts becomes index.html,
preview.css and preview.js, copied to dist/preview by the build and
filled at request time. The rendered page is unchanged apart from the
session id and draw.io origin now coming from a small config script.
Biome skips the folder because of the {{placeholders}}, as it never
linted the old template string either.
2026-10-04 06:46:09 +09:00
dayuan.jiang 667f678eb3 feat(mcp-server): auto-save each session's diagram to a .drawio file
- Save the latest diagram of every session 1 second after each change
  (AI write, browser edit, history restore) to ~/.next-ai-drawio/<id>.drawio,
  keep the newest 50, flush on shutdown; DRAWIO_DATA_DIR changes the folder
  and "off" disables it, like the web app's IndexedDB sessions
- start_session names the file, so a resumed conversation can reopen the
  diagram with load_diagram after the MCP process restarted
- Fix PNG/SVG exports randomly timing out: a previous export's 10 second
  timer cleared the export in progress, and a late reply could be taken
  for the current one; exports are now numbered
2026-10-04 06:42:44 +09:00
dayuan.jiang 81ad317375 feat(mcp-server): add screenshot_diagram so the model can check its render
- New read-only screenshot_diagram tool returns the rendered page as a PNG
  plus the web app's visual checklist (overlaps, edges crossing shapes,
  readability, layout, rendering errors), replacing the web app's
  separate vision model with the host model's own vision
- PNG exports use draw.io's width and pageId options: screenshots stay
  under ~140,000 base64 characters and page exports no longer swap the
  page on screen
- Fail fast with a clear message when the preview tab stopped polling
  (browsers throttle background tabs)
- Mention the screenshot step in the drawing guide and instructions
2026-10-04 06:29:33 +09:00
dayuan.jiang 9de281627e feat(mcp-server): bring the web app's drawing knowledge to MCP
- Add a drawing guide adapted from the web system prompt (layout, edge
  routing, styles, minimal style, editing rules), returned by
  start_session, a new get_drawing_guide tool and the diagram-workflow prompt
- Add get_shape_library with the 30 icon libraries; the build copies
  docs/shape-libraries into dist and CI checks the packed files
- Accept bare mxCell lists in create_new_diagram and add_page; the server
  adds the wrapper and root cells
- Send server instructions, shorten create_new_diagram's description to
  fit Claude Code's 2,048 character limit, and annotate every tool
- Fix dead links and the totals in docs/shape-libraries/README.md
2026-10-04 06:15:46 +09:00
dayuan.jiang 6d67a0ec69 fix(mcp-server): make edit_diagram all-or-nothing and fix preview sync races
- edit_diagram applies nothing when any operation fails, rejects invalid or
  multi-cell new_xml, validates only the target page, and returns the
  current page XML on every rejection (including stale edits)
- Fix get_diagram reading the old diagram right after an AI write: the
  preview pushed its sync reply with a newer version than it was taken at
- Keep a user edit that loses the race with an AI write in history and
  tell the user in the preview
- Autofix removes only exact foreign tags (a stray <mxGraph/> deleted
  <mxGraphModel>), fixes tag case, drops orphan <mxPoint>s, and rejects
  unknown element names in model XML
- Edit empty and compressed pages; PNG exports use the page on screen;
  tag download exports; reload from the server after a page export
- Expand ~ in paths, tell the model when the browser sync timed out,
  use registerPrompt, require SDK ^1.31.0
2026-10-03 22:03:44 +09:00
dayuan.jiang a46787c1b8 fix(mcp-server): fix XSS and crashes, make XML validation strict
- Validate and escape the mcp session id; only serve localhost Host/Origin
- Malformed URLs and session ids return errors instead of crashing the process
- Strict XML syntax check with saxes (linkedom never reports parse errors)
- autoFixXml no longer corrupts valid XML; attribute newlines serialized as entities
- Sessions stay alive while polled; browser pushes carry a base version (409 on conflict)
- Page tools respect the edit gate; UTF-8 bodies decoded correctly
- Export replies matched to requests and serialized; xml sync export handled
- UserObject/object cells addressable by id; history restored by stable id; logs off stdout
2026-10-03 17:45:41 +09:00
Dayuan Jiang 4b07228320 feat(mcp): add load_diagram tool to load .drawio files into the session (#893)
* feat(mcp): add load_diagram tool to load .drawio files into the session

Loading a file previously required the agent to read the file itself and
pass the entire XML through create_new_diagram - wasteful for large
diagrams and impossible for draw.io's compressed save format.

load_diagram takes a file path; the server reads it, decompresses any
compressed pages (base64 -> raw deflate -> URI-decode, per page), and
replaces the session document. The loaded XML is deliberately NOT marked
as seen by the edit gate: the model only supplied a path, so it must
call get_diagram once before editing.

* chore(mcp): version 0.2.3

* fix(mcp): report package.json version in the MCP handshake

The McpServer metadata version was a separate hardcoded string that
never matched the published version (stuck at 0.1.2, then 0.3.0 while
npm shipped 0.2.x). Read it from package.json at startup instead —
works from both src/ (tsx) and dist/ (published build).
2026-07-12 19:54:42 +09:00
NgoQuocViet2001anddayuan.jiang f3a85558d8 fix(mcp): replace edit_diagram 30s time gate with content comparison (#890)
* fix(mcp): keep diagram context valid during edits

Closes #885

* fix(mcp): replace edit_diagram time gate with content comparison

The 30s wall-clock gate rejected slow-but-correct clients (#885).
Instead of a timeout, remember the exact state-store XML the model
last saw (get_diagram / create_new_diagram / edit_diagram / page CRUD)
and reject edit_diagram only when the live browser state differs -
i.e. the user made edits the model hasn't seen yet. Slow reasoning
no longer trips the gate, while unseen manual edits still do.

* docs(mcp): align edit_diagram/get_diagram descriptions with content-based gate

The 'You MUST call get_diagram BEFORE this tool' requirement and the
'Skipping get_diagram WILL cause user's changes to be LOST' warning no
longer match server behavior: a stale edit is rejected with no side
effects, never silently applied. Describe the freshness check instead,
and direct get_diagram usage at its real purpose - learning the current
diagram content when the model doesn't already know it.

* fix(mcp): compare diagram content structurally in the edit gate

draw.io re-serialises the document when pushing state back (attribute
order, pretty-printing, regenerated diagram ids, viewport attributes,
mxfile host), so byte comparison could flag an unchanged diagram as
stale. Fingerprint what a user can actually change instead - page set,
page names, and each page's root cell tree with sorted attributes -
keeping byte equality as the fast path. A bare mxGraphModel now also
fingerprints identically to its single-page mxfile wrapping.

* fix(mcp): don't compare page names against bare mxGraphModel pushes

A bare <mxGraphModel> pushed by the embed/sync path carries no page name,
so normalizeToMxfile invents "Page-1" — falsely reading any custom page
name as a content change and re-triggering the stale rejection on every
edit. When either side of the gate comparison is a bare mxGraphModel,
fingerprint cell trees only; full-mxfile comparisons still detect renames.

* chore(mcp): bump version to 0.2.2

* chore(mcp): sync package-lock.json version to 0.2.2

---------

Co-authored-by: dayuan.jiang <[email protected]>
2026-07-12 15:33:20 +09:00
Siddhant Shekharanddayuan.jiang 5c884766a8 feat(mcp): add multi-page (mxfile) support to MCP server (#862)
* feat(mcp): add multi-page (mxfile) support

The MCP server's write path could only address a single drawio page even
though the underlying .drawio file format and the embedded editor both
natively support multi-page documents. A user asking for "a second page
with a CNN diagram" would hit the validator with the error
"Expected closing tag </root> but found </mxCell>" because the validator
assumed input was a bare <mxGraphModel> and could not walk past the
<mxfile><diagram>...</diagram></mxfile> wrapper.

This patch closes the gap end to end:

* New helper module `pages.ts` centralises page CRUD (normalize, parse,
  list, find, add, rename, delete) so every layer agrees that the
  canonical in-memory shape is always <mxfile>. normalizeToMxfile and
  addPageToDoc both strip any leading <?xml ?> declaration before
  embedding a fragment inside <diagram> (the declaration is only valid
  at document start). addPageToDoc explicitly rejects full <mxfile>
  inputs so a caller cannot accidentally nest a document inside a page.
* `xml-validation.ts` now detects an <mxfile> root and scopes the
  duplicate-id check per <diagram>. The legacy regex check would
  otherwise reject every multi-page doc, because cells "0" and "1"
  repeat in each page's <root> by design. The DOM-parse path is gated
  by a cheap regex pre-check so legacy bare <mxGraphModel> callers
  don't pay any extra cost. The autoFix duplicate-id rename step is
  also guarded against mxfile inputs — renaming those sentinel cells
  would silently break drawio's parent references.
* `diagram-operations.ts` accepts an optional PageSelector. For
  <mxfile> input it resolves the page first and scopes all
  querySelectorAll calls to that page's <root>, so a delete on page 2's
  cell "2" no longer touches page 1's cell "2".
* `create_new_diagram` accepts either a bare <mxGraphModel> (legacy,
  auto-wrapped into a single-page mxfile) or a full <mxfile> with N
  diagrams. All existing single-page callers keep working unchanged.
* `edit_diagram`, `get_diagram`, and `export_diagram` gain optional
  `page_id` / `page_name` / `page_index` parameters. When omitted they
  target the first page — the "active by convention" default. Tool
  handlers with all-optional input schemas coalesce missing arguments
  via `input ?? {}` so a no-args MCP invocation can't crash on
  destructure before reaching the session-existence check.
* New tools: `list_pages`, `add_page`, `rename_page`, `delete_page`.
* Page-targeted PNG/SVG export uses a "load + export + restore" dance:
  the server projects the target page into a single-page <mxfile>,
  pushes it into the transient state so the browser reloads the iframe
  with just that page, waits for drawio to render (~3s), triggers the
  export, captures the data, and then restores the original multi-page
  document. The dance is wrapped in `try/finally` so the restore runs
  unconditionally — even if an exception is thrown mid-dance, the
  user's multi-tab view is recovered before the function returns.
  The earlier attempt to use drawio's `selectPage` postMessage was a
  no-op because drawio's JSON embed protocol does not expose that
  action — silently exporting whatever tab happened to be active. The
  load-export-restore approach trades a brief visible tab-flicker for
  correctness: the exported image is guaranteed to match the requested
  page.
* Tool description strings reflect the multi-page semantics so the LLM
  client learns the new contract.
* Package version bumped 0.2.0 → 0.3.0 (additive surface — four new
  tools, three extended input schemas, canonical XML shape change).
* CI: `.github/workflows/test.yml` gains an explicit install + vitest
  run for the mcp-server package so the new multi-page invariants are
  covered by automation, not just local runs.

Backward compatibility: every existing single-page caller continues to
work without modification. The session.xml shape is normalised on every
write, removing the wrapper-injection hack from the .drawio download
path.

Tests: 43 unit tests under `packages/mcp-server/tests/multi-page.test.ts`
pin the validator's mxfile path, the page-scoped operations, the XML
declaration-prefix handling for both normalizeToMxfile and addPageToDoc,
addPageToDoc's rejection of full <mxfile> inputs, the single-page
projection used by export_diagram (a direct regression test for the
selectPage bug — two distinct page selectors must produce visually
different projections), and the Transformer + CNN motivating scenario.
A `tests/smoke.mjs` smoke test drives the built `dist/index.js` over
JSON-RPC and asserts all 9 tools register with the right input schemas.
Root vitest suite (107 tests) still green.

* fix(mcp): rewrite page-targeted export browser-side; harden edit/get

The page-targeted PNG/SVG export never worked: export_diagram swapped the
live session to a single-page projection, slept 3s, then wrote the export
flag onto a state object that setState() had already replaced in the store
Map — so the browser never saw the request and every such export timed out.
The swap+restore also clobbered concurrent edits.

Move the projection entirely browser-side: requestExport() hands a single
-page <mxfile> to the bridge via state.exportXml; the bridge loads it,
lets draw.io render, exports, then reloads the user's real document. The
canonical session state is never mutated, so there is no restore race and
no fixed-delay guessing. The export poll now re-reads the live store entry
each tick instead of a captured reference. autosave is suppressed and the
version-bump reload is skipped while a projection is on screen; if no real
document was captured, restore forces a server reload rather than leaving
the iframe stuck on the projection.

Also:
- edit_diagram now returns isError on a page-level failure (selector matched
  no page / page has no <root>) instead of reporting success-with-warnings
  and persisting a no-op; the pre-edit history snapshot is taken only after
  that gate so a failed edit leaves no phantom undo entry.
- edit_diagram/get_diagram re-normalise browser-pushed xml to mxfile so a
  bare <mxGraphModel> can't silently strip a multi-page document.
- get_diagram now errors (instead of silently returning the full doc) when a
  selector is given but the session isn't a parseable mxfile.
- page_id / page_name / add_page.id get .min(1) so empty strings can't
  silently target the first page.
- Extract pages.ts:projectPage(), collapsing three copies of the
  parse→find→serialise projection logic in index.ts.
- Replace the never-in-CI tests/smoke.mjs with tests/server-wiring.test.ts,
  which boots the server from source via tsx and runs under the existing
  vitest CI step.

* chore(mcp): set version to 0.2.1 for release

---------

Co-authored-by: dayuan.jiang <[email protected]>
2026-06-16 09:15:50 +09:00