Skip to content

Latest commit

 

History

History
835 lines (672 loc) · 45.6 KB

File metadata and controls

835 lines (672 loc) · 45.6 KB

Architecture

This document describes the architecture of NodeTool — a visual AI workflow platform that runs across desktop, web, and mobile. It covers both the TypeScript backend (the packages/ monorepo) and the React frontend (the web/ application), along with the Electron desktop shell and React Native mobile app.


Table of Contents


High-Level Overview

┌──────────────────────────────────────────────────────────────────┐
│                          Clients                                 │
│  ┌──────────┐   ┌──────────┐   ┌──────────┐   ┌──────────┐     │
│  │  Web UI  │   │ Electron │   │  Mobile  │   │   CLI    │     │
│  │ (React)  │   │ (Desktop)│   │ (Expo)   │   │  (Ink)   │     │
│  └────┬─────┘   └────┬─────┘   └────┬─────┘   └────┬─────┘     │
│       │              │              │              │             │
│       └──────────────┴──────┬───────┴──────────────┘             │
│                             │                                    │
│                     HTTP + WebSocket                             │
│                             │                                    │
├─────────────────────────────┼────────────────────────────────────┤
│                       Backend Server                             │
│                             │                                    │
│  ┌──────────────────────────┴──────────────────────────────┐     │
│  │              WebSocket + HTTP Server                     │     │
│  │           (packages/websocket)                          │     │
│  └──────┬────────────┬────────────┬────────────┬───────────┘     │
│         │            │            │            │                 │
│  ┌──────┴──┐  ┌──────┴──┐  ┌─────┴────┐ ┌────┴─────┐           │
│  │ Kernel  │  │ Agents  │  │  Models  │ │  Auth    │           │
│  │ (DAG    │  │ (LLM    │  │ (SQLite  │ │ (JWT     │           │
│  │ Runner) │  │ Tasks)  │  │ + ORM)   │ │ Tokens)  │           │
│  └────┬────┘  └────┬────┘  └──────────┘ └──────────┘           │
│       │            │                                            │
│  ┌────┴────┐  ┌────┴────────────────────────────────────┐       │
│  │Node SDK │  │  Tools (100+)                           │       │
│  │+ Nodes  │  │  Search, Code, File, Browser, PDF, ...  │       │
│  └────┬────┘  └─────────────────────────────────────────┘       │
│       │                                                         │
│  ┌────┴────────────────────────────────────────────────┐        │
│  │  Runtime  │  Protocol  │  Config  │  Security       │        │
│  │  Storage  │  VectorDB  │  Code Runners              │        │
│  └─────────────────────────────────────────────────────┘        │
└─────────────────────────────────────────────────────────────────┘

Repository Layout

nodetool/
├── packages/               # TypeScript backend monorepo (58 packages)
│   ├── protocol/           #   Shared types & message definitions (Zod)
│   ├── config/             #   Environment & settings management
│   ├── security/           #   Encryption, secrets, master key
│   ├── auth/               #   JWT authentication & user management
│   ├── storage/            #   Pluggable asset storage (file, S3, Supabase)
│   ├── models/             #   Database models (SQLite + Drizzle ORM)
│   ├── runtime/            #   Processing context & LLM providers
│   ├── kernel/             #   DAG orchestration & workflow runner
│   ├── node-sdk/           #   BaseNode class & node registry
│   ├── base-nodes/         #   100+ built-in node types
│   ├── agents/             #   Agent system with task planning & tools
│   ├── chat/               #   Chat protocol & runtime integration
│   ├── vectorstore/        #   SQLite-vec vector database
│   ├── dsl/                #   Type-safe workflow DSL
│   ├── websocket/          #   HTTP + WebSocket server (entry point)
│   ├── cli/                #   Terminal UI (React + Ink)
│   ├── deploy/             #   Deployment utilities (Docker, SSH, RunPod, GCP)
│   ├── huggingface/        #   HuggingFace model discovery & cache
│   ├── replicate-nodes/    #   Replicate AI integration nodes
│   ├── replicate-codegen/  #   Code generator for Replicate nodes
│   ├── fal-nodes/          #   FAL AI integration nodes
│   ├── fal-codegen/        #   Code generator for FAL nodes
│   ├── elevenlabs-nodes/   #   ElevenLabs voice synthesis nodes
│   ├── minimax-nodes/      #   MiniMax TTS, music, image, and video nodes
│   └── ...                 #   Additional provider packs (kie, topaz, reve,
│                           #   atlascloud, together), sandbox, compute, gpu,
│                           #   timeline, sdk, workflow-runner, and node packs
├── web/                    # React web application (Vite + MUI)
├── electron/               # Electron desktop shell
├── mobile/                 # React Native mobile app (Expo)
├── docs/                   # Jekyll documentation site
├── scripts/                # Build, release, and dev helper scripts
├── examples/               # Example workflows
├── package.json            # npm workspace configuration (build/test/lint scripts)
└── tsconfig.base.json      # Shared TypeScript configuration

Design Principles

  1. Streaming-first execution — Results are streamed to clients as they are produced. Workflows, chat, and agent tasks all emit incremental updates over WebSocket so users see progress in real time and can cancel at any point.

  2. Unified runtime — The same workflow JSON graph runs identically on desktop (Electron), headless server, RunPod GPU cloud, or Google Cloud Run. No platform-specific code is required.

  3. Pluggable execution strategies — Nodes can run in-process (fast iteration), in subprocesses (isolation), or in Docker containers (deployment). The strategy is chosen at runtime without changing the workflow definition.

  4. Layered package architecture — Packages are organized into layers (foundational → infrastructure → domain → feature → application) with strict dependency direction. Lower layers never import from higher layers.

  5. Type safety end-to-end — The protocol package defines Zod schemas shared between backend and frontend. The API client is generated from OpenAPI specs. TypeScript strict mode is enforced across all packages.


Backend Architecture

The backend is a TypeScript monorepo of 58 npm workspace packages. Each package has a focused responsibility and explicit dependencies.

Package Dependency Graph

Application Layer
  └── websocket ─────────────────────────── Server entry point
  └── cli ───────────────────────────────── Terminal UI
  └── deploy ────────────────────────────── Deployment tooling

Feature Layer
  ├── base-nodes ────────────────────────── Built-in node types
  ├── dsl ───────────────────────────────── Workflow DSL
  ├── replicate-nodes, fal-nodes ────────── Provider integrations
  ├── elevenlabs-nodes ──────────────────── Voice synthesis
  ├── minimax-nodes ─────────────────────── Speech, music, image, video
  ├── huggingface ───────────────────────── Model discovery
  └── chat ──────────────────────────────── Chat protocol

Domain Layer
  ├── kernel ────────────────────────────── DAG runner & actors
  ├── node-sdk ──────────────────────────── Node framework
  ├── agents ────────────────────────────── Agent task planning
  └── models ────────────────────────────── Database persistence

Infrastructure Layer
  ├── runtime ───────────────────────────── Processing context
  ├── vectorstore ───────────────────────── Vector search
  └── security ──────────────────────────── Encryption & secrets

Foundational Layer
  ├── protocol ──────────────────────────── Types & messages (Zod)
  ├── config ────────────────────────────── Environment management
  ├── storage ───────────────────────────── Asset storage adapters
  └── auth ──────────────────────────────── Authentication providers

Foundational Layer

These packages have no internal dependencies and form the base of the stack.

protocol — Defines all shared types using Zod schemas: ProcessingMessage, NodeDescriptor, Edge, JobUpdate, NodeUpdate, OutputUpdate, and asset reference types (ImageRef, AudioRef, VideoRef). Provides TypeMetadata for complex type parsing (e.g., list[dict[str, int]]), JSON schema generation for tool definitions, and type validation/coercion utilities.

config — Environment and settings management built on dotenv. Exports loadEnvironment(), getEnv(), requireEnv() for environment variables, registerSetting()/getSettings() for a typed settings registry, createLogger() using pino for structured logging, and diagnoseEnvironment() for startup validation.

storage — Pluggable asset storage with a shared AbstractStorage interface (store, retrieve, exists). Four backend implementations: FileStorage (local filesystem), MemoryStorage (tests), S3Storage (AWS S3 via dynamic import), and SupabaseStorage (Supabase cloud). SDKs are loaded lazily to keep them as optional runtime-only dependencies.

auth — JWT-based authentication with multiple provider implementations: LocalAuthProvider (file-based users), StaticTokenProvider (fixed tokens), MultiUserAuthProvider, and SupabaseAuthProvider. Provides middleware (createAuthMiddleware), token extraction, and a UserManager abstraction with role-based access control.

Infrastructure Layer

Built on top of the foundational packages.

runtime — The central ProcessingContext class that every node receives during execution. It provides a message queue for emitting ProcessingMessage events, cache adapter interface (get/set with TTL), storage adapter interface for assets, and LLM provider abstractions (BaseProvider) for OpenAI, Anthropic, Gemini, and others. Also includes StreamingInputs/StreamingOutputs types, a Python bridge for interop with Python nodes, and OpenTelemetry tracing.

security — Cryptography and secret management. Uses AES-256-GCM encryption with a master key derived from OS keychain (keytar), AWS KMS, or environment variables. Provides getSecret()/hasSecret() for database-backed encrypted credential access with caching. Includes runStartupChecks() to validate encryption and database connectivity at boot.

vectorstore — SQLite-vec backed vector database for semantic search. Exports SqliteVecStore with collection-based organization, multiple embedding providers (OpenAI, Gemini, Ollama, Mistral), and document splitting with configurable chunk size and overlap.

Domain Layer

Core business logic for workflows, nodes, agents, and persistence.

kernel — The workflow orchestration engine. Key components:

  • Graph — DAG data structure with O(1) node/edge lookup via Map-based indexing, edge type validation, cycle detection, and topological sort (Kahn's algorithm).
  • WorkflowRunner — Coordinates execution: graph initialization, NodeInbox creation for message buffering, actor spawning for concurrent node execution, edge counter tracking for end-of-stream propagation, and output collection.
  • NodeActor — Executes individual nodes by resolving a NodeExecutor implementation, managing output routing, and handling streaming/batched inputs.
  • NodeInbox — Per-node message buffer that tracks upstream completion and supports different sync modes (zip_all waits for all inputs, on_any fires immediately).

node-sdk — The framework for defining custom nodes. Exports BaseNode (abstract class with static metadata, property/output declarations, and serialization), NodeRegistry (central type registry), and TypeScript decorators (@node(), @output(), @property()) for declarative node definition. Nodes implement process(context, values) returning an output record.

agents — Multi-step LLM agent system with layered execution:

  • Agent — Entry point. Takes an objective (plus skills, tools, and optional outputSchema); orchestrates planning, execution, and final synthesis. With a pre-built task it skips planning and runs a single task. With useGraphPlanner it builds and runs a workflow DAG via authorGraph + executeAgentGraph.
  • TaskPlanner — LLM-driven decomposition of an objective into a TaskPlan DAG of tasks and steps.
  • authorGraph — a CodeAct sub-agent writes the workflow with the generated, typed @nodetool-ai/sandbox-dsl pack, discovering nodes with search_nodes / get_node_info / list_nodes / find_model and checking its work with validate_workflow; it hands the graph back through finish(). Produces GraphData ready for the kernel.
  • ParallelTaskExecutorTaskExecutorStepExecutor — Runs TaskPlan tasks concurrently, each task's steps sequentially (or in parallel when independent).
  • CompilerAgent — Final synthesis pass after ParallelTaskExecutor finishes; reads accumulated context.memory and produces the deliverable (schema-conformant JSON when outputSchema is set, otherwise prose).
  • executeAgentGraph (src/execute-agent-graph.ts) — Hands a GraphData graph to WorkflowRunner (kernel) and yields the run's ProcessingMessages. Planner-authored nodetool.agents.Agent nodes carry no model; applyRunPolicy stamps the run's configured provider+model onto them, and the live tool set is injected into a child ProcessingContext.
  • Tool system with 100+ tools across categories: search (Google, DataForSEO), code execution, file I/O, browser automation (Playwright), email, image generation, PDF processing, vector search, workflow management, and MCP (Model Context Protocol) integration.

models — Database persistence layer using Drizzle ORM over SQLite. Defines tables for: workflows (DAG definitions), jobs (execution records), messages / threads (chat history), assets (file metadata), secrets (encrypted credentials), workspaces, workflowVersions, oauthCredentials, predictions (usage/cost tracking), and runEvents.

Feature Layer

Node implementations and provider integrations.

base-nodes — 100+ built-in node types organized by category: input/output, data processing (lists, dicts, transforms), text manipulation, code execution (Python/JS/TS via isolated-vm), document processing (PDF extraction, DOCX/Excel/Markdown conversion), image/audio/video processing, web scraping (Playwright), email, search, agent execution, vector store operations, and LLM model integration (Gemini, Anthropic, OpenAI).

dsl — Type-safe TypeScript DSL for programmatically defining workflows. Converts between DSL representation and kernel graph format.

chat — Chat protocol and runtime integration for conversational AI interactions.

replicate-nodes, fal-nodes, elevenlabs-nodes, minimax-nodes — Provider-specific node packs for Replicate AI, FAL AI, ElevenLabs voice synthesis, and MiniMax (speech, music, image, video). Each extends BaseNode from node-sdk.

huggingface — HuggingFace model discovery, cache scanning, and artifact inspection for local model management.

Application Layer

Entry points that wire everything together.

websocket — The main server package. Runs an HTTP + WebSocket server on port 7777 (default).

HTTP API routes (40+ endpoints):

Route prefix Purpose
/api/workflows/* Workflow CRUD, tools, examples
/api/jobs/* Job status, triggers, cancellation
/api/messages/*, /api/threads/* Chat messages and threads
/api/assets/* Asset storage, search, thumbnails
/api/nodes/* Node metadata and validation
/api/settings/* Configuration and secrets
/api/collections/* Vector store collections
/api/models/* LLM provider model listings
/api/users/*, /api/workspaces/* User and workspace management
/v1/* OpenAI-compatible API endpoints
/api/oauth/* OAuth flows

WebSocket commands handled by WebSocketClientSession:

Command Purpose
run_job Execute a workflow DAG
chat_message Chat with agent mode
inference Direct LLM inference
cancel_job Stop execution
reconnect_job / resume_job Resume existing or suspended jobs
stream_input / end_input_stream Push streaming data to input nodes
get_status Query job status

The server also registers all node types (base + provider nodes), initializes the Python bridge for HuggingFace/MLX node execution, and sets up tool registries.

cli — Interactive terminal UI built with React + Ink. Provides nodetool-chat for conversational AI and nodetool for workflow management, connecting to the WebSocket server.

deploy — Deployment utilities supporting Docker image build/push, SSH remote deployment, RunPod GPU cloud, and Google Cloud Run/Compute Engine. Uses YAML configuration files.

Workflow Execution Pipeline

Client sends "run_job" via WebSocket
        │
        ▼
WebSocketClientSession (packages/websocket)
  ├── Creates ProcessingContext (storage, cache, secrets, LLM providers)
  ├── Provides resolveExecutor() for kernel
  │
  ▼
WorkflowRunner (packages/kernel)
  ├── Graph.fromDict() — validates, indexes, detects cycles, sorts topologically
  ├── Analyzes "streaming paths" — which edges carry streaming data
  ├── Creates NodeInbox per node (tracks upstream source counts)
  ├── Calls NodeActor.initialize() for all nodes
  ├── Dispatches input values to InputNodes
  ├── Spawns NodeActors concurrently (one per node)
  │
  ▼
NodeActor (packages/kernel) — runs per node
  ├── Pulls messages from NodeInbox (per-handle FIFO queue)
  ├── Resolves NodeExecutor via resolveExecutor():
  │     TypeScript NodeRegistry → PythonNodeExecutor (fallback)
  ├── Execution mode (chosen by node metadata):
  │     • Buffered:          process(inputs) → outputs (one shot)
  │     • Streaming output:  genProcess(inputs) → yields partial outputs
  │     • Streaming I/O:     run(StreamingInputs, StreamingOutputs)
  │     • Controlled:        caches inputs, awaits control event before firing
  │
  ▼
NodeExecutor.process() / .genProcess() / .run()
  ├── TypeScript nodes: BaseNode.process(context, values)
  ├── Python nodes:     PythonNodeExecutor → Python bridge (msgpack over stdio)
  ├── Agent nodes:      AgentNode → provider loop (LLM + tools)
  ├── Emits ProcessingMessages via context (node_update, chunk, etc.)
  │
  ▼
Output Routing
  ├── Results sent along edges to downstream NodeInboxes
  ├── Edge counters track EOS propagation per edge
  ├── Streaming paths propagate EOS immediately; buffered paths wait
  ├── OutputUpdate messages streamed to client via WebSocket
  │
  ▼
Client receives streaming updates (JobUpdate, NodeUpdate, OutputUpdate)

NodeExecutor interface — all executor types implement one or more of:

interface NodeExecutor {
  // Buffered: gather all inputs, call once
  process(inputs: Record<string, unknown>, ctx?: ProcessingContext): Promise<Record<string, unknown>>;

  // Streaming output: yields partial results progressively
  genProcess?(inputs: Record<string, unknown>, ctx?: ProcessingContext): AsyncGenerator<Record<string, unknown>>;

  // Streaming I/O: node controls its own input/output flow
  run?(inputs: StreamingInputs, outputs: StreamingOutputs, ctx?: ProcessingContext): Promise<void>;

  // Lifecycle hooks
  initialize?(): Promise<void>;
  preProcess?(): Promise<void>;
  finalize?(): Promise<void>;
}

Agent System

The agent system has a single entry point — Agent — which dispatches to one of three execution paths based on its options:

Path 1: Pre-built task — Caller supplies a Task directly; Agent skips planning and runs it through TaskExecutorStepExecutor.

Path 2: useGraphPlanner: trueauthorGraph builds a workflow graph (DAG of typed nodes) which executeAgentGraph hands to the kernel's WorkflowRunner. Every node — including the nodetool.agents.Agent reasoning steps — resolves through the normal NodeRegistry; the runner supplies those steps' model and tools. This is the hybrid path: LLM-driven reasoning nodes run alongside deterministic nodes in the same kernel.

Path 3: Default (plan)TaskPlanner decomposes the objective into a TaskPlan DAG; ParallelTaskExecutor runs independent tasks concurrently via TaskExecutorStepExecutor; CompilerAgent synthesizes the final deliverable from accumulated context.memory.

Agent.execute(context)
  ├── Loads skills from filesystem (auto-matched to objective)
  ├── Recalls long-term memory (if configured) and folds into system prompt
  │
  ├─── pre-built task ────────────────────────────────────┐
  │    TaskExecutor → StepExecutor                         │
  │      ├── LLM call with tool schemas                    │
  │      ├── tool_call → execute → result → LLM            │
  │      └── Repeats until finish_step tool called         │
  │                                                        │
  ├─── useGraphPlanner ───────────────────────────────────┤
  │    authorGraph                                         │
  │      ├── LLM discovers nodes (search_nodes,           │
  │      │   get_node_info, find_model), then writes the  │
  │      │   graph with the typed sandbox-dsl pack        │
  │      └── validate_workflow checks it; finish(graph)   │
  │          emits GraphData                              │
  │    executeAgentGraph                                   │
  │      └── WorkflowRunner (kernel) with custom resolver:│
  │            Agent nodes: model+tools from runner        │
  │                               └── StepExecutor (LLM)  │
  │            Other nodes → NodeRegistry (deterministic)  │
  │                                                        │
  └─── default plan ──────────────────────────────────────┘
       TaskPlanner
         └── LLM → TaskPlan (tasks[] with depends_on DAG)
       ParallelTaskExecutor
         └── Independent tasks run concurrently
             TaskExecutor (per task)
               └── StepExecutor (per step)
                    └── results stored in context.memory
       CompilerAgent
         └── Reads memory → produces final deliverable

Results streamed to client via WebSocket as AsyncGenerator<ProcessingMessage>

Three levels of concurrency:

  1. Task-levelParallelTaskExecutor runs independent tasks from the TaskPlan DAG concurrently
  2. Step-levelTaskExecutor can run independent steps within a task in parallel
  3. Node-levelWorkflowRunner runs all nodes with satisfied dependencies concurrently via NodeActor

Node System

Nodes are the fundamental units of computation. Each node:

  • Extends BaseNode from node-sdk
  • Declares input properties with types and defaults
  • Declares output slots with type information
  • Implements process(context, values) returning output values
  • Can be synchronous, streaming-output (genProcess), or full streaming I/O (run)
@node({ namespace: "math", title: "Add" })
class AddNode extends BaseNode {
  @property({ type: "float" }) a: number = 0;
  @property({ type: "float" }) b: number = 0;
  @output({ type: "float" })  result: number = 0;

  async process(context: ProcessingContext, values: { a: number; b: number }) {
    return { result: values.a + values.b };
  }
}

Node types are registered in NodeRegistry. At runtime the kernel resolves each node through an executor chain:

  1. TypeScript NodeRegistry — registered BaseNode subclasses run in-process
  2. PythonNodeExecutor — fallback for nodes declared in Python; communicates via the PythonStdioBridge (msgpack over stdin/stdout)
  3. AgentNode — the nodetool.agents.Agent node; runs a provider tool-calling loop and feeds its result back into the kernel graph

The executor type used also controls which NodeActor execution mode fires (buffered / streaming output / streaming I/O / controlled).


Frontend Architecture

Technology Stack

Layer Technology Version
UI Framework React 19.2
Language TypeScript 5.9
Bundler Vite 8.0
Component Library Material-UI (MUI) v7
State Management Zustand 5.0
Server State TanStack Query (React Query) v5
Graph Editor @xyflow/react (React Flow) v12
Routing React Router v7
Styling Emotion (CSS-in-JS) v11
Code Editor Monaco Editor
Rich Text Lexical
3D Visualization React Three Fiber + Three.js
Panel Layout Dockview
Testing Jest + React Testing Library 29.7
E2E Testing Playwright 1.60

Application Shell & Routing

The app entry point (web/src/index.tsx) wraps the application in a provider stack:

<ErrorBoundary>
  <ThemeProvider theme={ThemeNodetool}>
    <CssBaseline />
    <QueryClientProvider>
      <MobileClassProvider>
        <MenuProvider>
          <WorkflowManagerProvider>
            <KeyboardProvider>
              <RouterProvider routes={...} />
            </KeyboardProvider>
          </WorkflowManagerProvider>
        </MenuProvider>
      </MobileClassProvider>
    </QueryClientProvider>
  </ThemeProvider>
</ErrorBoundary>

Routes:

Path Component Description
/ NavigateToStart Redirects to /workspace, or /login when auth is required
/workspace WorkspaceShell The tabbed workspace; its empty state is the new-project surface
/dashboard, /welcome, /chat/:thread_id? Legacy routes, all redirect to /workspace
/settings SettingsRedirect Opens Settings as a workspace tab
/editor/:workflow WorkflowEditorRedirect Opens the workflow as a workspace tab
/miniapp/:workflowId LegacyAppRedirect Standalone app runner for one workflow
/a/:token PublicAppPage A mini app published to a hidden URL
/assets AssetExplorer Asset browser
/collections CollectionsExplorer Vector collection browser
/examples ExamplesPage Template workflow gallery
/models ModelsPage Model browser
/timeline/:sequenceId TimelineEditor Multi-track video editor
/sketch/:documentId SketchEditorPage Image document editor
/login Login Authentication page

All routes except /login, /a/:token, and dev routes are wrapped in ProtectedRoute.

State Management

The frontend uses a multi-layer state management approach:

Zustand stores (90+) manage client-side state. Major stores:

Store Responsibility
NodeStore Per-tab graph state (nodes, edges) with @xyflow/react integration and temporal undo/redo via zundo
WorkflowManagerStore Open workflows, create/load/copy/delete via API, localStorage persistence
WorkflowRunner Per-workflow execution state machine (idle → connecting → running → completed/error)
GlobalChatStore Chat threads, messages, WebSocket streaming, tool calls
ResultsStore Workflow execution results and previews
AssetStore Asset management with TanStack Query integration
MetadataStore Node type metadata registry
SettingsStore User settings with localStorage persistence
ModelDownloadStore HuggingFace model downloads with progress tracking
NotificationStore / ErrorStore User-facing notifications and errors
PanelStore / BottomPanelStore / RightPanelStore UI layout and panel visibility
KeyPressedStore Keyboard event tracking

Store patterns:

  • Selector-based subscriptions to prevent unnecessary re-renders: useStore(state => state.value)
  • Middleware: persist() for localStorage, temporal() for undo/redo
  • Per-workflow stores created and managed by WorkflowManagerContext

TanStack Query manages server state (data fetching, caching, mutations):

  • useWorkflow(), useAssets(), useWorkflowVersions() — query hooks in serverState/
  • Hierarchical query keys: ['workflows', workflowId]
  • Automatic cache invalidation after mutations

React Contexts coordinate cross-cutting concerns:

  • WorkflowManagerContext — manages all per-workflow NodeStore instances for the tabbed editor, persists open workflows to localStorage, provides validation helpers
  • NodeContext — provides node-specific data within node component trees
  • EditorInsertionContext — manages editor insertion/creation points

Component Architecture

web/src/components/
├── node_editor/       # Core graph editor (canvas, drag-drop, selection)
├── node/              # Individual node rendering and property editing
├── node_menu/         # Node search and quick-add menu
├── panels/            # App layout (Header, PanelLeft, PanelRight, PanelBottom)
├── workspace/         # The tabbed workspace shell and its tab content
├── projects/          # Project list, overview, and the new-project surface
├── chat/              # Chat UI (messages, streaming)
├── assets/            # Asset explorer and editor
├── workflows/         # Workflow templates and grid views
├── collections/       # Vector collection management
├── hugging_face/      # Model browser and download manager
├── textEditor/        # Rich text editor (Lexical)
├── audio/             # Audio player and waveform components
├── video/             # Video player components
├── asset_viewer/      # Asset preview (PDF, images, 3D models)
├── inputs/            # Form inputs and input widgets
├── widgets/           # Node input widgets (color pickers, sliders, etc.)
├── buttons/           # Reusable button components
├── dialogs/           # Modal dialogs
├── menus/             # Context menus and dropdowns
├── themes/            # MUI theme configuration
├── ui_primitives/     # Low-level UI primitives
├── ui/                # Generic shared UI components
└── common/            # Shared components (loaders, errors)

WebSocket Communication

The frontend communicates with the backend via a single WebSocket connection managed by WebSocketManager (in web/src/lib/websocket/):

  • Binary protocol — Messages are encoded with msgpack (not JSON) for performance
  • State machine — Connection states: disconnectedconnectingconnectedreconnectingfailed
  • Auto-reconnect — Exponential backoff with configurable decay (1.5×) and max attempts (10). Auth/policy errors (codes 1008–1011, 4000–4003) are not retried.
  • Message queuing — Outbound messages are queued during connection and flushed when connected
  • Global singletonGlobalWebSocketManager provides a shared connection for all stores

Messages flow:

WorkflowRunner store ──┐
GlobalChatStore ───────┴──▶ GlobalWebSocketManager ──▶ WebSocket ──▶ Server
                                      ▲
                                      │
                              message events
                                      │
                              ◀───────┘

API Client

REST calls go through restFetch (web/src/lib/rest-fetch.ts), a thin fetch wrapper:

  • Base URLBASE_URL (web/src/stores/BASE_URL.ts) from VITE_API_URL, empty in local dev so requests stay relative
  • Auth header — a Supabase Bearer token is attached only when the server reports authMode: "supabase" (isAuthRequired() in web/src/lib/runtimeConfig.ts); in local auth mode no header is sent
  • Environment detectionisLocalhost / isElectron / isProduction flags in web/src/lib/env.ts, derived from the hostname (localhost, 127.0.0.1, or a dev. prefix) and the Electron preload bridge
  • Types — request/response types come from @nodetool-ai/protocol, re-exported through web/src/stores/ApiTypes.ts
  • Dev proxy — Vite proxies /api, /ws, /trpc, and /storage (rewritten to /api/storage) to localhost:7777 during development

Code Splitting & Performance

Vite production builds use manual chunk splitting:

Chunk Contents
vendor-buffer buffer, base64-js, ieee754
vendor-react React, ReactDOM, React Router
vendor-mui Material-UI, Emotion
vendor-flow @xyflow/react
vendor-query TanStack Query, tRPC, msgpack
vendor-supabase Supabase client
vendor-utils cmdk, chroma-js, uuid, zod

Everything else — Monaco, Lexical, Plotly, Three.js, pdf.js, WaveSurfer — is feature-only and stays unnamed on purpose, so the boot path does not pull it in.

Heavy components (PanelLeft, PanelRight, PanelBottom, WorkspaceShell, Chat, Editor) are lazy-loaded with React.lazy() and <Suspense> boundaries.


Electron Desktop App

The Electron shell (electron/) wraps the web UI for desktop use:

  • Electron 39 with Vite for the renderer process
  • Uses contextBridge for secure IPC — nodeIntegration is disabled
  • electron-builder for packaging (Windows, macOS, Linux including Flatpak)
  • electron-updater for auto-updates
  • electron-log for structured logging
  • Hash-based routing (#/path) for compatibility with file:// protocol
  • Preload scripts bridge frontend tool calls to main process capabilities

Mobile App

The React Native mobile app (mobile/) enables users to browse and run mini-apps:

  • Expo SDK 56 with React Native 0.85
  • React Navigation (native-stack) for screen transitions
  • Zustand for state management (WorkflowRunner, ChatStore)
  • Axios for HTTP, custom WebSocketManager with msgpack for real-time chat
  • AsyncStorage for persistent settings (server URL configuration)
  • Screens: MiniAppsListScreen, MiniAppScreen, SettingsScreen, ChatScreen

See mobile/ARCHITECTURE.md for detailed mobile architecture documentation.


Communication Protocols

WebSocket Protocol

All WebSocket messages use msgpack binary encoding.

Client → Server commands:

run_job          { job_id, workflow_id, graph, params }
chat_message     { thread_id, content, model, tools }
inference        { model, messages, tools }
cancel_job       { job_id }
reconnect_job    { job_id }
resume_job       { job_id }
stream_input     { job_id, node_id, data }
end_input_stream { job_id, node_id }
get_status       { job_id }

Server → Client events:

job_update       { job_id, status }           # Job lifecycle (started, completed, failed, cancelled)
node_update      { job_id, node_id, status }  # Node execution state changes
output_update    { job_id, node_id, value }   # Streaming output values
node_progress    { job_id, node_id, progress } # Progress percentage
task_update      { task_id, status, result }  # Agent task progress
planning_update  { plan }                     # Agent planning status
prediction       { model, tokens, cost }      # LLM usage tracking
error            { message, details }         # Error notifications

HTTP API

REST endpoints follow standard conventions:

  • JSON request/response bodies
  • Bearer token authentication
  • OpenAPI spec for type generation
  • Vite dev proxy for local development

Data Flow Examples

Workflow Execution

1. User clicks "Run" in the editor
2. WorkflowRunner store serializes the graph from NodeStore
3. Sends "run_job" command via GlobalWebSocketManager (msgpack)
4. Server creates ProcessingContext with storage and secrets
5. WorkflowRunner (kernel) validates and indexes the graph
6. NodeActors process nodes concurrently following the DAG topology
7. Each node's BaseNode.process() emits ProcessingMessages
8. Outputs route along edges to downstream NodeInboxes
9. Server streams job_update, node_update, output_update to client
10. WorkflowRunner store updates execution state
11. ResultsStore updates output previews
12. NodeStore reflects node status (running, completed, error)

Chat / Agent Interaction

1. User sends a message in the chat UI
2. GlobalChatStore sends "chat_message" via WebSocket
3. Server creates or retrieves the thread
4. Agent loads skills from filesystem (auto-matched to objective) and recalls long-term memory
5. Path selection:
   pre-built task → TaskExecutor → StepExecutor: LLM ↔ tools until finish_step
   useGraphPlanner → authorGraph: LLM writes the node graph with the typed DSL pack
                    executeAgentGraph → WorkflowRunner (kernel)
                      Agent nodes → AgentNode → provider loop
                      Other nodes    → NodeRegistry (deterministic)
   default (plan) → TaskPlanner → TaskPlan DAG
                    ParallelTaskExecutor: concurrent tasks, each via StepExecutor
                    CompilerAgent: synthesize final result from context.memory
6. All paths yield AsyncGenerator<ProcessingMessage>
7. Streaming chunks (chunk, task_update, node_update) sent as produced
8. GlobalChatStore appends chunks to the current message
9. Final message stored in database (models package)

Storage & Persistence

Data Storage Location
Workflows, jobs, threads, messages SQLite (Drizzle ORM) packages/models
Assets (images, audio, video, files) Pluggable: filesystem, S3, or Supabase packages/storage
Encrypted secrets SQLite + AES-256-GCM packages/security + packages/models
Vector embeddings SQLite-vec packages/vectorstore
User settings (frontend) localStorage web/src/stores/SettingsStore
Open workflows (frontend) localStorage web/src/stores/WorkflowManagerStore
Server URL (mobile) AsyncStorage mobile/src/services/api.ts

Authentication

Authentication is handled by packages/auth with pluggable providers:

Provider Use Case
LocalAuthProvider Single-user local development (file-based)
StaticTokenProvider Fixed API token for headless/CI usage
MultiUserAuthProvider Multi-user with role-based access (admin, user)
SupabaseAuthProvider Production cloud authentication

The frontend uses Supabase JS client for session management, token refresh, and attaches Bearer tokens via API client middleware.


Deployment

Supported deployment targets (via packages/deploy):

Target Description
Local Docker container on local machine
Docker Container build and push to registry
SSH Remote deployment to self-hosted servers
RunPod GPU cloud platform for ML workloads
GCP Google Cloud Run / Compute Engine

Deployments are configured via YAML files and managed through the CLI. The same server image runs in all environments — only the storage adapter and auth provider change.


Build System

Top-Level Make Targets

npm install          # Install all dependencies (web, electron, mobile)
npm run build            # Build all packages + web
npm run typecheck        # TypeScript type-check all packages
npm run lint             # Lint all packages
npm run lint:fix         # Auto-fix linting issues
npm run test             # Run all tests
npm run check            # typecheck + lint + test (quality gate)
npm run clean            # Remove dependencies and build artifacts

Package Build Order

Packages must be built in dependency order. The root build:packages script handles this:

protocol → config → security → storage → auth
    → runtime → vectorstore → models → kernel → node-sdk
    → agents → base-nodes → dsl → chat → huggingface
    → replicate-nodes → fal-nodes → elevenlabs-nodes → minimax-nodes
    → websocket → cli → deploy

Development Mode

npm run dev              # Starts websocket server + web dev server (concurrent)
npm run electron         # Builds web, starts Electron desktop app

Vite dev server runs on port 3000 with hot module replacement. API calls are proxied to the backend on port 7777.


Testing Strategy

Layer Framework Location
Backend packages Vitest packages/*/tests/
Web unit tests Jest + React Testing Library web/src/__tests__/, web/src/components/__tests__/
Web E2E tests Playwright web/tests/ (e2e-runner/, journeys/, smoke/, debug-harness/, benchmarks/)
Electron unit tests Jest electron/src/__tests__/
Mobile tests Jest mobile/src/**/__tests__/

Backend packages use Vitest with per-package vitest.config.ts files. The web app uses Jest with jsdom environment and extensive module mocks. E2E tests use Playwright with automatic server startup.

# Run all tests
npm run test

# Package-specific
cd packages/kernel && npx vitest run
cd web && npm test
cd web && npm run test:e2e

CI/CD is managed through 41 GitHub Actions workflows covering unit tests, E2E tests, code quality, security auditing, accessibility, performance monitoring, and release automation.