* feat(mcp): add multi-page (mxfile) support
The MCP server's write path could only address a single drawio page even
though the underlying .drawio file format and the embedded editor both
natively support multi-page documents. A user asking for "a second page
with a CNN diagram" would hit the validator with the error
"Expected closing tag </root> but found </mxCell>" because the validator
assumed input was a bare <mxGraphModel> and could not walk past the
<mxfile><diagram>...</diagram></mxfile> wrapper.
This patch closes the gap end to end:
* New helper module `pages.ts` centralises page CRUD (normalize, parse,
list, find, add, rename, delete) so every layer agrees that the
canonical in-memory shape is always <mxfile>. normalizeToMxfile and
addPageToDoc both strip any leading <?xml ?> declaration before
embedding a fragment inside <diagram> (the declaration is only valid
at document start). addPageToDoc explicitly rejects full <mxfile>
inputs so a caller cannot accidentally nest a document inside a page.
* `xml-validation.ts` now detects an <mxfile> root and scopes the
duplicate-id check per <diagram>. The legacy regex check would
otherwise reject every multi-page doc, because cells "0" and "1"
repeat in each page's <root> by design. The DOM-parse path is gated
by a cheap regex pre-check so legacy bare <mxGraphModel> callers
don't pay any extra cost. The autoFix duplicate-id rename step is
also guarded against mxfile inputs — renaming those sentinel cells
would silently break drawio's parent references.
* `diagram-operations.ts` accepts an optional PageSelector. For
<mxfile> input it resolves the page first and scopes all
querySelectorAll calls to that page's <root>, so a delete on page 2's
cell "2" no longer touches page 1's cell "2".
* `create_new_diagram` accepts either a bare <mxGraphModel> (legacy,
auto-wrapped into a single-page mxfile) or a full <mxfile> with N
diagrams. All existing single-page callers keep working unchanged.
* `edit_diagram`, `get_diagram`, and `export_diagram` gain optional
`page_id` / `page_name` / `page_index` parameters. When omitted they
target the first page — the "active by convention" default. Tool
handlers with all-optional input schemas coalesce missing arguments
via `input ?? {}` so a no-args MCP invocation can't crash on
destructure before reaching the session-existence check.
* New tools: `list_pages`, `add_page`, `rename_page`, `delete_page`.
* Page-targeted PNG/SVG export uses a "load + export + restore" dance:
the server projects the target page into a single-page <mxfile>,
pushes it into the transient state so the browser reloads the iframe
with just that page, waits for drawio to render (~3s), triggers the
export, captures the data, and then restores the original multi-page
document. The dance is wrapped in `try/finally` so the restore runs
unconditionally — even if an exception is thrown mid-dance, the
user's multi-tab view is recovered before the function returns.
The earlier attempt to use drawio's `selectPage` postMessage was a
no-op because drawio's JSON embed protocol does not expose that
action — silently exporting whatever tab happened to be active. The
load-export-restore approach trades a brief visible tab-flicker for
correctness: the exported image is guaranteed to match the requested
page.
* Tool description strings reflect the multi-page semantics so the LLM
client learns the new contract.
* Package version bumped 0.2.0 → 0.3.0 (additive surface — four new
tools, three extended input schemas, canonical XML shape change).
* CI: `.github/workflows/test.yml` gains an explicit install + vitest
run for the mcp-server package so the new multi-page invariants are
covered by automation, not just local runs.
Backward compatibility: every existing single-page caller continues to
work without modification. The session.xml shape is normalised on every
write, removing the wrapper-injection hack from the .drawio download
path.
Tests: 43 unit tests under `packages/mcp-server/tests/multi-page.test.ts`
pin the validator's mxfile path, the page-scoped operations, the XML
declaration-prefix handling for both normalizeToMxfile and addPageToDoc,
addPageToDoc's rejection of full <mxfile> inputs, the single-page
projection used by export_diagram (a direct regression test for the
selectPage bug — two distinct page selectors must produce visually
different projections), and the Transformer + CNN motivating scenario.
A `tests/smoke.mjs` smoke test drives the built `dist/index.js` over
JSON-RPC and asserts all 9 tools register with the right input schemas.
Root vitest suite (107 tests) still green.
* fix(mcp): rewrite page-targeted export browser-side; harden edit/get
The page-targeted PNG/SVG export never worked: export_diagram swapped the
live session to a single-page projection, slept 3s, then wrote the export
flag onto a state object that setState() had already replaced in the store
Map — so the browser never saw the request and every such export timed out.
The swap+restore also clobbered concurrent edits.
Move the projection entirely browser-side: requestExport() hands a single
-page <mxfile> to the bridge via state.exportXml; the bridge loads it,
lets draw.io render, exports, then reloads the user's real document. The
canonical session state is never mutated, so there is no restore race and
no fixed-delay guessing. The export poll now re-reads the live store entry
each tick instead of a captured reference. autosave is suppressed and the
version-bump reload is skipped while a projection is on screen; if no real
document was captured, restore forces a server reload rather than leaving
the iframe stuck on the projection.
Also:
- edit_diagram now returns isError on a page-level failure (selector matched
no page / page has no <root>) instead of reporting success-with-warnings
and persisting a no-op; the pre-edit history snapshot is taken only after
that gate so a failed edit leaves no phantom undo entry.
- edit_diagram/get_diagram re-normalise browser-pushed xml to mxfile so a
bare <mxGraphModel> can't silently strip a multi-page document.
- get_diagram now errors (instead of silently returning the full doc) when a
selector is given but the session isn't a parseable mxfile.
- page_id / page_name / add_page.id get .min(1) so empty strings can't
silently target the first page.
- Extract pages.ts:projectPage(), collapsing three copies of the
parse→find→serialise projection logic in index.ts.
- Replace the never-in-CI tests/smoke.mjs with tests/server-wiring.test.ts,
which boots the server from source via tsx and runs under the existing
vitest CI step.
* chore(mcp): set version to 0.2.1 for release
---------
Co-authored-by: dayuan.jiang <[email protected]>
Next AI Draw.io
A Next.js web application that integrates AI capabilities with draw.io diagrams. Create, modify, and enhance diagrams through natural language commands and AI-assisted visualization.
Note: Thanks to
ByteDance Doubao sponsorship, the demo site now uses the powerful glm-4.7 model!
https://github.com/user-attachments/assets/9d60a3e8-4a1c-4b5e-acbb-26af2d3eabd1
Table of Contents
- Next AI Draw.io
Examples
Here are some example prompts and their generated diagrams:
Features
- LLM-Powered Diagram Creation: Leverage Large Language Models to create and manipulate draw.io diagrams directly through natural language commands
- Image-Based Diagram Replication: Upload existing diagrams or images and have the AI replicate and enhance them automatically
- PDF & Text File Upload: Upload PDF documents and text files to extract content and generate diagrams from existing documents
- AI Reasoning Display: View the AI's thinking process for supported models (OpenAI o1/o3, Gemini, Claude, etc.)
- Diagram History: Comprehensive version control that tracks all changes, allowing you to view and restore previous versions of your diagrams before the AI editing.
- Interactive Chat Interface: Communicate with AI to refine your diagrams in real-time
- Cloud Architecture Diagram Support: Specialized support for generating cloud architecture diagrams (AWS, GCP, Azure)
- Animated Connectors: Create dynamic and animated connectors between diagram elements for better visualization
MCP Server
Use Next AI Draw.io with AI agents like Claude Desktop, Cursor, and VS Code via MCP (Model Context Protocol).
{
"mcpServers": {
"drawio": {
"command": "npx",
"args": ["@next-ai-drawio/mcp-server@latest"]
}
}
}
Claude Code CLI
claude mcp add drawio -- npx @next-ai-drawio/mcp-server@latest
Then ask Claude to create diagrams:
"Create a flowchart showing user authentication with login, MFA, and session management"
The diagram appears in your browser in real-time!
See the MCP Server README for VS Code, Cursor, and other client configurations.
Getting Started
Try it Online
No installation needed! Try the app directly on our demo site:
Bring Your Own API Key: You can use your own API key to bypass usage limits on the demo site. Click the Settings icon in the chat panel to configure your provider and API key. Your key is stored locally in your browser and is never stored on the server.
Desktop Application
Download the native desktop app for your platform from the Releases page:
Supported platforms: Windows, macOS, Linux.
Run with Docker
Installation
- Clone the repository:
git clone https://github.com/DayuanJiang/next-ai-draw-io
cd next-ai-draw-io
npm install
cp env.example .env.local
See the Provider Configuration Guide for detailed setup instructions for each provider.
- Run the development server:
npm run dev
- Open http://localhost:6002 in your browser to see the application.
Deployment
Deploy to EdgeOne Pages
You can deploy with one click using Tencent EdgeOne Pages.
Deploy by this button:
Check out the Tencent EdgeOne Pages documentation for more details.
Additionally, deploying through Tencent EdgeOne Pages will also grant you a daily free quota for DeepSeek models.
Deploy on Vercel
The easiest way to deploy is using Vercel, the creators of Next.js. Be sure to set the environment variables in the Vercel dashboard as you did in your local .env.local file.
See the Next.js deployment documentation for more details.
Deploy on Cloudflare Workers
Multi-Provider Support
- ByteDance Doubao
- AWS Bedrock (default)
- OpenAI
- Anthropic
- Google AI
- Google Vertex AI
- Azure OpenAI
- Ollama
- OpenRouter
- AIHubMix
- DeepSeek
- SiliconFlow
- ModelScope
- SGLang
- Vercel AI Gateway
All providers except AWS Bedrock and OpenRouter support custom endpoints.
📖 Detailed Provider Configuration Guide - See setup instructions for each provider.
Server-Side Multi-Model Configuration
Administrators can configure multiple server-side models that are available to all users without requiring personal API keys. Configure via AI_MODELS_CONFIG environment variable (JSON string) or ai-models.json file. For a single-provider quick setup, list comma-separated model IDs in AI_MODEL.
Admin Panel
Set the ADMIN_PASSWORD environment variable and visit /admin to manage server settings (models, access codes, features, observability, quota) from a web panel instead of hand-editing .env.
📖 Admin Panel Guide — setup, precedence rules, and notes.
Model Requirements: This task requires strong model capabilities for generating long-form text with strict formatting constraints (draw.io XML). Recommended models include Claude Sonnet 4.5, GPT-5.1, Gemini 3 Pro, and DeepSeek V3.2/R1.
Note that the claude series has been trained on draw.io diagrams with cloud architecture logos like AWS, Azure, GCP. So if you want to create cloud architecture diagrams, this is the best choice.
How It Works
The application uses the following technologies:
- Next.js: For the frontend framework and routing
- Vercel AI SDK (
ai+@ai-sdk/*): For streaming AI responses and multi-provider support - react-drawio: For diagram representation and manipulation
Diagrams are represented as XML that can be rendered in draw.io. The AI processes your commands and generates or modifies this XML accordingly.
Support & Contact
Special thanks to ByteDance Doubao for sponsoring the API token usage of the demo site! Register on the ARK platform to get 500K free tokens for all models!
If you find this project useful, please consider sponsoring to help me host the live demo site!
For support or inquiries, please open an issue on the GitHub repository or contact the maintainer at:
- Email: me[at]jiang.jp
FAQ
See FAQ for common issues and solutions.