| title | NVIDIA NeMo Fabric Documentation |
|---|---|
| slug | /about-nemo-fabric/overview |
| description | Configure, plan, run, and observe agent harnesses and custom agents through one typed execution contract. |
| template-library-version | 1.0.0 |
{/* SPDX-FileCopyrightText: Copyright (c) 2026, NVIDIA CORPORATION & AFFILIATES. All rights reserved. SPDX-License-Identifier: Apache-2.0 */}
NVIDIA NeMo Fabric is the integration layer that turns agent harnesses and custom agents into one configurable, observable execution surface. Applications use the same versioned config, lifecycle, result, artifact, and telemetry contracts whether the selected Adapter Target is Hermes Agent, Codex SDK, a registered workflow behind a shared framework adapter, or a dedicated custom agent.
NeMo Fabric owns the seam between an application and its Adapter Target. It resolves configuration, selects an adapter, drives the runtime lifecycle, and returns normalized evidence without leaking target-specific control code into the caller.
Construct a complete, versioned `FabricConfig` in Python. Applications create variants with ordinary functions and typed copies. Plan and invoke harnesses and custom agents through one Rust core, CLI, and Python SDK instead of embedding target launch logic in every consumer. Resolve configs, inspect capabilities, run single-invocation jobs, and hold multi-turn runtimes with typed requests, plans, handles, and results. Collect output, errors, lifecycle events, artifact manifests, and telemetry references in stable contracts suitable for platforms and evaluations.Use the following table to choose the NeMo Fabric interface that best fits how your application works with Adapter Targets:
| Interface | Use it when | Start with |
|---|---|---|
| Python SDK | Your application owns job config, runtime lifecycle, or multi-turn state | Client API |
| Runtime API | You need multiple ordered turns over one live harness runtime | Runtime |
| Streaming API | You need live ATOF records generated by NeMo Relay during a runtime turn | Streaming |
nemo-fabric CLI |
You are experimenting with harnesses, running maintained examples, or troubleshooting configs | Experimentation CLI |
| JSON Schema | You are building editors, validation, code generation, or another language binding | Committed schemas in the repository |
Use FabricConfig as the canonical configuration contract. CLI selectors
obtain complete typed configs from built-in presets or maintained examples.
Application or evaluation harness
|
| Python SDK or nemo-fabric experimentation CLI
v
NeMo Fabric Rust core
config -> plan -> lifecycle
|
| resolved adapter contract
v
Agent harness | shared framework target | dedicated custom agent
|
v
RunResult + artifacts + events + telemetry references
- Configure a typed
FabricConfigwith a harness adapter, environment, models, tools, skills, MCP, and telemetry. - Create variants from deep copies to vary harness, model, environment, or observability settings without mutating the base config.
- Plan and diagnose to resolve the adapter and check capabilities and requirements before spending work on a runtime.
- Run or start a runtime through the shared lifecycle contract. To consume
live ATOF records, enable NeMo Relay, start the runtime with
streaming=True, and callRuntime.invoke_stream(). - Consume evidence from
RunResult: output, structured failure details, artifacts, events, and telemetry references.
Adapter authors can follow the adapter contract overview to choose a harness adapter, shared framework adapter, or dedicated custom-agent adapter and implement the minimum lifecycle.
Continue exploring NeMo Fabric through these resources.
- Installation — Installation to set up the runtime and adapters.
- Quickstart — Quickstart to build from source and run the maintained SDK example.
- Python SDK — Python SDK for planning, diagnostics, typed requests, and multi-turn runtimes.
- API Reference — Client API to resolve, plan, diagnose, run, and start stateful runtimes.