Skip to content

Latest commit

 

History

History
106 lines (90 loc) · 5 KB

File metadata and controls

106 lines (90 loc) · 5 KB
title NVIDIA NeMo Fabric Documentation
slug /about-nemo-fabric/overview
description Configure, plan, run, and observe agent harnesses and custom agents through one typed execution contract.
template-library-version 1.0.0

{/* SPDX-FileCopyrightText: Copyright (c) 2026, NVIDIA CORPORATION & AFFILIATES. All rights reserved. SPDX-License-Identifier: Apache-2.0 */}

NVIDIA NeMo Fabric is the integration layer that turns agent harnesses and custom agents into one configurable, observable execution surface. Applications use the same versioned config, lifecycle, result, artifact, and telemetry contracts whether the selected Adapter Target is Hermes Agent, Codex SDK, a registered workflow behind a shared framework adapter, or a dedicated custom agent.

NeMo Fabric owns the seam between an application and its Adapter Target. It resolves configuration, selects an adapter, drives the runtime lifecycle, and returns normalized evidence without leaking target-specific control code into the caller.

What NeMo Fabric Gives You

Construct a complete, versioned `FabricConfig` in Python. Applications create variants with ordinary functions and typed copies. Plan and invoke harnesses and custom agents through one Rust core, CLI, and Python SDK instead of embedding target launch logic in every consumer. Resolve configs, inspect capabilities, run single-invocation jobs, and hold multi-turn runtimes with typed requests, plans, handles, and results. Collect output, errors, lifecycle events, artifact manifests, and telemetry references in stable contracts suitable for platforms and evaluations.

Choose Your Interface

Use the following table to choose the NeMo Fabric interface that best fits how your application works with Adapter Targets:

Interface Use it when Start with
Python SDK Your application owns job config, runtime lifecycle, or multi-turn state Client API
Runtime API You need multiple ordered turns over one live harness runtime Runtime
Streaming API You need live ATOF records generated by NeMo Relay during a runtime turn Streaming
nemo-fabric CLI You are experimenting with harnesses, running maintained examples, or troubleshooting configs Experimentation CLI
JSON Schema You are building editors, validation, code generation, or another language binding Committed schemas in the repository

Use FabricConfig as the canonical configuration contract. CLI selectors obtain complete typed configs from built-in presets or maintained examples.

Core Workflow

Application or evaluation harness
        |
        |  Python SDK or nemo-fabric experimentation CLI
        v
NeMo Fabric Rust core
  config -> plan -> lifecycle
        |
        |  resolved adapter contract
        v
Agent harness | shared framework target | dedicated custom agent
        |
        v
RunResult + artifacts + events + telemetry references
  1. Configure a typed FabricConfig with a harness adapter, environment, models, tools, skills, MCP, and telemetry.
  2. Create variants from deep copies to vary harness, model, environment, or observability settings without mutating the base config.
  3. Plan and diagnose to resolve the adapter and check capabilities and requirements before spending work on a runtime.
  4. Run or start a runtime through the shared lifecycle contract. To consume live ATOF records, enable NeMo Relay, start the runtime with streaming=True, and call Runtime.invoke_stream().
  5. Consume evidence from RunResult: output, structured failure details, artifacts, events, and telemetry references.

Adapter authors can follow the adapter contract overview to choose a harness adapter, shared framework adapter, or dedicated custom-agent adapter and implement the minimum lifecycle.

Learn More

Continue exploring NeMo Fabric through these resources.

  • InstallationInstallation to set up the runtime and adapters.
  • QuickstartQuickstart to build from source and run the maintained SDK example.
  • Python SDKPython SDK for planning, diagnostics, typed requests, and multi-turn runtimes.
  • API ReferenceClient API to resolve, plan, diagnose, run, and start stateful runtimes.