Warning
Active development: FlintCode is under active development. APIs, configuration, workflows, and the TUI may change without notice. It is not yet recommended for production use.
Strike code without the cloud. — 燧石取火,离线生码。
FlintCode is a secure harness for AI coding agents in air-gapped and compliance-restricted environments. It keeps Agent choice separate from the trusted TeaQL validation and remote execution boundary: Agent-specific extensions are thin adapters, while all project operations pass through the policy-enforcing Runner.
Built with Rust. Powered by TeaQL.
See the component layout for the Runner, generic Agent Bridge, thin Agent integrations, pipeline, and legacy TUI boundaries.
Every major AI coding tool today — Cursor, Copilot, Devin, Windsurf — requires cloud connectivity. For organizations bound by regulatory, security, or data sovereignty constraints, none of them are usable.
FlintCode fills this gap:
| Cloud Coding Agents | FlintCode | |
|---|---|---|
| Network | Requires internet | Air-gapped / offline |
| Data privacy | Code sent to cloud | Data never leaves premises |
| Model size | 100B+ parameters | Works with small local models |
| Correctness | Hope it compiles | Compiler-verified |
| Compliance | ❌ GDPR / HIPAA / 等保 | ✅ Fully compliant |
FlintCode doesn't try to replace every general-purpose coding agent. Instead, it wraps compatible Agents with a deterministic validation pipeline and a policy-enforcing remote Runner:
Business Requirements
│
▼
┌──────────────────┐
│ LLM: Generate │ ← Small model (Nemotron-3-Super, Llama, etc.)
│ KSML Domain Model│
└────────┬─────────┘
│
┌────▼────┐
│ L1-L2 │ Local XML validation
│ L3 │ TeaQL domain validation (real semantics check)
│ L4 │ cargo teaql rust-lib-core (generate runtime code)
│ L5 │ cargo check (compiler verification ✓)
│ L6 │ cargo teaql assist (extract real API signatures)
│ L7 │ LLM: Generate business logic (guided by real APIs)
│ L8 │ cargo check (final compilation ✓)
└────┬────┘
│
▼
Verified, Compilable
Production Rust Code
The key insight: We don't let the LLM guess APIs. TeaQL's assist system feeds the model real, compiler-generated API signatures. Combined with multi-level validation, even a small 8B model can produce code that compiles on the first try.
This section is the source of truth for the isolation architecture. It records both what the current prototype already supports and what is still planned.
Status legend:
- Implemented in the current codebase
- Planned or not yet wired into the production execution path
The intended trust boundary is strict: the local FlintCode process is a control plane, not a code execution environment. It may persist Agent memory, task state, Skills, RAG results, and audit metadata, but it must not read or write a project workspace or execute project commands. A disposable remote machine is the only execution plane.
- Local Agent state machine, TUI, model client, context budgeting, and Skill orchestration foundations exist.
- Enterprise RAG client and retrieval status are represented in the current application.
- Continuous follow-up tasks retain bounded context, validation summaries, and a session ledger.
- A versioned SSH runner MVP now provides content-addressed bootstrap upload, persistent remote sessions, structured file/process operations, monotonic policy narrowing, cancellation, and reconnect by operation ID.
- Deterministic Cargo/Maven verification and real command exit status run through the durable SSH execution session.
- CLI and TUI project work require a strict SSH execution profile and stop on infrastructure failure; there is no automatic local project fallback.
- The production Pipeline uses one durable SSH runner workspace for initial generation, repair, validation, and continuous follow-ups.
- Structured remote operations cover stat/hash, ranged reads, directory list/walk, literal search, snapshots, atomic/CAS writes, artifact chunks, structured execution, replay, and cancellation.
- Provision every task into a fresh VM or microVM with a fixed TTL. Containers may be supported where the customer accepts the weaker isolation boundary.
- Run the remote workspace as an unprivileged user under a fixed root such as
/workspace/project, without host mounts or host filesystem visibility. - Canonicalize every runner-side path and reject traversal, device files, sockets, and symlink escapes outside the workspace.
- Disable public outbound networking and cloud metadata access. Allow only explicitly configured internal Git, package mirror, and artifact endpoints.
- Use pinned SSH host identities and short-lived, task-scoped credentials; never bake long-lived credentials into an image.
- Install all required toolchains in the remote image, including exactly
cargo-teaql2.0.11, and validate image capabilities before cloning code. - Confirm that the resulting commit exists in the enterprise Git service before the remote environment is destroyed.
- Destroy the environment on success, failure, cancellation, or TTL expiry, and retain only approved audit metadata and result identifiers locally.
CLI run/evaluate and the TUI require --execution-config; select a named
target with --execution-target and reattach a durable workspace with
--resume-session. See SSH Runner MVP for the full
configuration schema and the ca-mini deploy/probe/run procedure.
customer-controlled network
Local control plane Enterprise knowledge plane
+----------------------+ +--------------------------+
| model orchestration | <-----> | read-only RAG knowledge |
| memory and context | | APIs, rules, examples |
| Skills and task FSM | +--------------------------+
| audit metadata |
+----------+-----------+
|
| authenticated SSH / structured runner protocol
v
+----------------------+ +--------------------------+
| disposable execution | ------> | internal Git / artifacts |
| project files | commit | source of durable results|
| build and test tools | + push +--------------------------+
+----------+-----------+
|
v
destroyed
- Introduce one production
RemoteExecutionBackend; keep any local backend available only to isolated tests and explicit developer fixtures. - Give the model logical remote paths only. Local filesystem paths must never appear in model-visible tool results.
- Keep RAG access on the control plane and send only bounded, task-relevant knowledge excerpts to the model. The disposable machine does not receive unrestricted access to the enterprise knowledge base.
- Define a versioned runner capability handshake covering protocol version, workspace root, OS, architecture, installed toolchains, resource limits, and network policy.
- Implement an explicit lifecycle: provision, connect, preflight, clone, code, validate, commit, push, confirm, destroy.
- Make every operation idempotent where practical and attach task, environment, image digest, and operation identifiers to every audit event.
- Support reconnect and cancellation without losing the durable Agent memory or accidentally starting the task on a different workspace.
- Treat the enterprise Git commit SHA as the durable result. stdout, patches, and local workspace snapshots are not the primary handoff mechanism.
- Models cannot be relied on to obey prompt-only filesystem boundaries. An Agent
with a local shell can inspect
/,$HOME, parent directories, or symlinked paths even when instructed not to do so. - A disposable VM makes the machine itself the security boundary. If the Agent explores its root filesystem, it sees only a short-lived environment prepared for that task.
- Keeping source code and build tools off the control machine reduces the blast radius of model-generated commands and makes cleanup deterministic.
- Keeping RAG on the control plane prevents a disposable worker from obtaining broad access to enterprise knowledge and allows context to be filtered before it reaches the model.
- Prebuilt images avoid shipping multi-gigabyte Rust, Java, Maven, TeaQL, and dependency caches with the FlintCode application.
- Commit-and-push before destruction produces an auditable, reproducible result that survives the temporary machine.
- A versioned environment contract separates FlintCode from any specific cloud, hypervisor, container runtime, or physical distribution medium.
FlintCode should consume a capability contract rather than depend on one image format. The same standard environment may be delivered through several channels:
- OCI image in an internal registry for Kubernetes or approved container platforms.
- OVA/OVF image for VMware-based enterprise environments.
- QCOW2 image for KVM and OpenStack, including microVM-based task workers.
- Cloud marketplace image for customer-owned VPC/VNet deployments.
- Encrypted ISO, appliance disk, or other physical media for fully isolated sites without a connected image registry.
- Preinstalled physical-machine pool for sites that cannot create VMs on demand, with secure wiping and re-imaging after every task.
- Signed image manifest containing image digest, SBOM, protocol version, toolchain versions, supported capabilities, and compatibility range.
- Offline update bundles containing only changed image layers or packages, with signature and checksum verification.
- Customer-operated provisioner adapters so the same lifecycle can target VMware, OpenStack, Kubernetes, a cloud marketplace image, or a physical pool.
The application package should remain small: it contains the control plane, configuration, Skills, and environment contracts. Compilers, test toolchains, TeaQL tooling, and dependency caches belong in the separately distributed and replaceable execution image.
30 business objects, full pipeline — from natural language to compilable Rust:
| Step | Duration | Description |
|---|---|---|
| KSML Generation | ~315s | 10 LLM calls, ~5000 tokens |
| Domain Validation | ~1s | TeaQL semantic check |
| Code Generation | ~6s | cargo teaql for lib + app |
| Compilation | ~22s | Two cargo check passes |
| Assist Discovery | ~18s | 8 entity API signatures |
| Business Logic | ~30s | LLM writes query functions |
| Total | ~7 min | vs. 1-2 weeks manual |
The commands below use the current SSH-only project execution path. The local process remains the model/control plane; project files, generation, build, and validation stay in the selected durable runner session.
- Local LLM inference endpoint (NIM, vLLM, Ollama, etc.)
- Rust 1.75+ with Cargo for the local control-plane build
- A configured SSH runner target; see SSH Runner MVP
cargo-teaqlexactly 2.0.11 on the remote execution host:cargo install cargo-teaql --version 2.0.11 cargo-teaql install-links
git clone https://github.com/teaql/klint-code.git
cd klint-code
# Build
cargo build --release
# Run one backend code-generation task
LOCAL_API_KEY="<key-if-required>" \
./target/release/klintcode run \
--task benchmarks/tasks/school-service-rust \
--profile profiles/local-qwen.toml \
--build-target rust-lib-core \
--output runs/ \
--execution-config /absolute/path/to/remote-execution.toml \
--execution-target ca-mini
# Queue tasks and repeat the queue; each primary queue entry gets its own session
MIMO_API_KEY="<key>" \
./target/release/klintcode run \
--task benchmarks/tasks/school-service-rust \
--task benchmarks/tasks/simple-greeting \
--profile profiles/mimo-v2.5-pro.toml \
--build-target rust-lib-core \
--repeat 2 \
--output runs/ \
--execution-config /absolute/path/to/remote-execution.toml \
--execution-target ca-mini
# Evaluation-grade continuation with a typed, machine-verifiable contract
MIMO_API_KEY="<key>" \
./target/release/klintcode run \
--task benchmarks/tasks/moving-company-platform \
--profile profiles/mimo-v2.5-pro.toml \
--build-target rust-lib-core \
--follow-up benchmarks/tasks/moving-company-platform/followup.md \
--follow-up-acceptance benchmarks/tasks/moving-company-platform/followup-acceptance.json \
--output runs/ \
--execution-config /absolute/path/to/remote-execution.toml \
--execution-target ca-mini
# Legacy interactive TUI
./target/release/flintcode-tui-legacy \
--profile profiles/mimo-v2.5-pro.toml \
--execution-config /absolute/path/to/remote-execution.toml \
--execution-target ca-mini
# Ordinary composer input starts a task by default; later input continues it:
# Create a small school model
# Add school registration
# Add school information updates
# Use /ask <question> for one lightweight answer without leaving the task.
# Use /done to leave task mode, or /new to detach and start a fresh session.
# /task benchmarks/tasks/moving-company-platform remains available for packages.
# The moving-company suite runs its declared follow-up and strict contract
MIMO_API_KEY="<key>" \
./target/release/klintcode evaluate \
--plan benchmarks/moving-company-suite.toml \
--profile profiles/mimo-v2.5-pro.toml \
--output runs/evaluations \
--execution-config /absolute/path/to/remote-execution.toml \
--execution-target ca-miniEach repeated --task adds a fresh primary run to the queue. --repeat N
replays the complete queue, while repeated --follow-up values reuse the
workspace and bounded session history created by each primary run. A follow-up
may be literal instruction text or a path to a text file. Queued primary runs
continue after failures by default; add --fail-fast to stop at the first
failure. --skill applies an explicit modeling skill to every primary task.
The CLI requires one --follow-up-acceptance JSON file per --follow-up by
default. For a task package with a standard followup-acceptance.json sidecar,
the first continuation loads that contract automatically in both CLI and TUI.
The contract checks exact changed files, Rust Q/E expression
chains through the parsed AST, test counts, runtime exit status, and required or
forbidden output markers. Environment values are referenced by name and are
redacted from acceptance evidence; unrelated parent variables such as
MIMO_API_KEY are not inherited by generated build/test/runtime processes.
--allow-unverified-follow-up is an explicit
escape hatch for ad-hoc work and should never be used for evaluations. Without
an explicit contract, a follow-up that reaches its iteration limit is always a
failure.
health performs an authenticated /models request, parses the response, and
requires the configured model ID to be present. It exits non-zero on transport,
authentication, protocol, or model-selection failures:
MIMO_API_KEY="<key>" \
./target/release/klintcode health \
--profile profiles/mimo-v2.5-pro.tomlprobe runs explicit live conformance checks. Chat, stream, and tool checks
send small generation requests and may consume provider quota. Use --json
for machine-readable output, or select checks with repeated/comma-separated
--check values:
MIMO_API_KEY="<key>" \
./target/release/klintcode probe \
--profile profiles/mimo-v2.5-pro.toml \
--check models,chat,stream,tools \
--jsonRun a benchmark plan with per-case timeouts, token accounting, JSON output, Markdown output, and a non-zero exit status when any case fails:
MIMO_API_KEY="<key>" \
./target/release/klintcode evaluate \
--plan benchmarks/rust-build-suite.toml \
--profile profiles/mimo-v2.5-pro.toml \
--output runs/evaluation/The TUI opens with its prompt composer focused. Type a task and press Enter
to submit it through the normal Pipeline. Use Shift+Enter or Alt+Enter for
a new line. The composer grows from one to eight visible lines; additional
lines remain editable while the viewport follows the latest eight. Use Esc
for dashboard shortcuts and i to return to the composer.
The default surface stays minimal; enter /stats for Plan, Validation,
Context, and Token details, then /main to return.
Use the persistent protocol probe to test whether an OpenAI-compatible model can generate a task-specific plan, publish explicit plan status transitions, call constrained tools, edit an isolated Rust fixture, and pass its tests:
MIMO_API_KEY="<key>" \
FLINTCODE_ENDPOINT="https://example.com/v1" \
FLINTCODE_MODEL="model-id" \
node scripts/model-agent-probe.mjsThe probe never edits the repository. It writes a detailed report.json to
its temporary workspace.
Test whether MiMo can turn one immutable task goal and a bounded candidate skill set into a dependency-safe execution plan:
MIMO_API_KEY="<key>" \
FLINTCODE_MODEL="mimo-v2.5-pro" \
node scripts/mimo-skill-planner-probe.mjs \
--output /tmp/mimo-skill-plan-report.jsonThe probe includes phase-incompatible and unauthorized distractor skills. It
checks goal preservation, skill identity, phase and permission constraints,
input/output bindings, DAG validity, and acceptance-criterion coverage. Use
--dry-run to inspect the static test payload without contacting a model.
Add --with-tools to run the hierarchical DAG test. After the model creates
the Skill TaskGraph, a second bounded request expands every selected Skill into
a Tool ExecutionGraph. The local validator checks Skill-to-Tool allowlists,
permissions, bindings, exports, traceability, cycles, and the merged topological
execution waves:
MIMO_API_KEY="<key>" \
node scripts/mimo-skill-planner-probe.mjs \
--with-tools \
--output /tmp/mimo-hierarchical-dag-report.jsonProfiles are stored in profiles/. Example:
[model]
name = "nemotron-3-super"
endpoint = "http://localhost:8000/v1"
max_context = 65536
temperature = 0.1flint-code/
├── apps/
│ ├── flintcode-agent-bridge/ # Generic Agent-to-Runner bridge
│ ├── klintcode-cli/ # Headless CLI (compatibility name)
│ └── flintcode-tui-legacy/ # Maintenance-only Ratatui TUI
├── runner/ # Protocol, SSH client, bootstrap, runner server
├── integrations/
│ └── pi/extension/ # Thin Pi adapter for the generic bridge
├── crates/
│ ├── agent-core/ # State machine, reducer, event loop
│ ├── agent-bridge-core/ # Agent-neutral bridge protocol and dispatch
│ ├── execution-policy/ # Shared command and workspace policy
│ ├── pipeline/ # Evaluation suite runner, build validation
│ ├── model-vllm/ # LLM client (OpenAI-compatible)
│ ├── validation/ # Multi-level validation engine
│ ├── context-builder/ # Prompt construction, token budgeting
│ ├── artifact-store/ # Run output and artifact management
│ └── workspace-guard/ # Workspace isolation
├── benchmarks/
│ ├── tasks/ # Test cases (school-service, moving-company, etc.)
│ ├── rust-build-suite.toml # Quick validation suite
│ └── rust-full-30obj-bench.toml # Full 30-object benchmark
└── profiles/ # Model/hardware configuration
FlintCode's multi-level validation is what makes small models reliable:
| Level | What | How |
|---|---|---|
| L1 | XML well-formedness | Local parser |
| L2 | Schema conformance | Structure check |
| L3 | Domain semantics | cargo teaql evaluate (real TeaQL rules) |
| L4 | Code generation | cargo teaql rust-lib-core |
| L5 | Compilation | cargo check (Rust compiler) |
| L6 | API discovery | cargo teaql assist (real API signatures) |
| L7 | Business logic | LLM + assist output → query functions |
| L8 | Full compilation | cargo check on complete workspace |
If any level fails, the pipeline triggers an automatic repair loop — the LLM receives the specific error and regenerates.
FlintCode is hardware-agnostic but optimized for local inference:
| Device | Model Size | Speed | Context |
|---|---|---|---|
| DGX Spark | ≤ 70B (quantized) | ~16 tok/s | 64K |
| DGX Station | ≤ 1T | ~150+ tok/s | 128K+ |
| Any GPU server | Varies | Varies | Varies |
FlintCode follows the TeaQL Agent Kit rules:
- Never guess method names — use
assistoutput - Never edit generated files in
rust-lib-core/ - Every query:
.purpose("why")and.comment("what") - Every save:
.audit_as("description") - Required:
cargo-teaql2.0.11
Licensed under either of
at your option.
- TeaQL — Deterministic execution for non-deterministic AI
- teaql-agent-kit — Evaluation framework for AI coding agents
- ratatui — Rust terminal UI framework
