Skip to content

Commit c5b54ae

Browse files
committed
feat(recipes): OKE RDMA fabric wiring (L40S RoCE + GB200 IB)
Upstream the OKE network fabric, closing gb200-oke-training's 'NET/RDMA intentionally left out until OCI-specific pod RDMA exposure is verified on the testbed' carve-out — the exposure below is validated on a production BM.GPU.GB200.4 NVL72 rack and a BM.GPU.L40S.4 RoCE cluster. - network-operator on both OKE training chains, NicClusterPolicy supplied by manifest (chart deployCR off). L40S (RoCE): SR-IOV VF device plugin advertising nvidia.com/mlnxnics (ConnectX VF device IDs 101a/101e) plus nv-ipam and multus. GB200 (IB): rdmaSharedDevicePlugin over the NVL72 east-west rdma0-3 netdevs, same nvidia.com/mlnxnics resource name; no SR-IOV/nv-ipam. Neither deploys ofedDriver: OCI nodes carry host MOFED in every image. Present in every gpuStack value (fabric is orthogonal to driver/plugin ownership); incompatible with Oracle's opt-in NvidiaNetworkOperator add-on. - GB200 kernel-module-params wiring (NVreg_GrdmaPciTopoCheckOverride=1): dma-buf attach over the IB fabric — GPUDirect RDMA without nvidia-peermem, whose chroot modprobe fails against the -64k Grace kernel. - nccl-all-reduce-bw-net (>= 40, matching gb200-eks-training) added to the gb200-oke training chain; supportedNCCLCombinations[variantNET] gains oke/gb200 with the ported testdata/gb200/oke/runtime-net.yaml TrainingRuntime (IB via the shared HCAs; NVLS/MNNVL forced off). - NicClusterPolicy image digest exemptions (repository/image/version triplet CRD schema, same as the AKS entries). Stock-render golden and BOM regenerated. Signed-off-by: Atif Mahmood <atif1996@users.noreply.github.com>
1 parent fe764f3 commit c5b54ae

14 files changed

Lines changed: 468 additions & 13 deletions

docs/user/container-images.md

Lines changed: 6 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -20,7 +20,7 @@ A machine-readable **CycloneDX 1.6 JSON** companion to this page is produced by
2020
## Summary
2121

2222
- Components: **43**
23-
- Unique images: **97**
23+
- Unique images: **101**
2424
- Distinct registries: **11**
2525

2626
Registries: `602401143452.dkr.ecr.us-west-2.amazonaws.com`, `cr.agentgateway.dev`, `docker.io`, `gcr.io`, `ghcr.io`, `gke.gcr.io`, `nvcr.io`, `public.ecr.aws`, `quay.io`, `registry.k8s.io`, `us-docker.pkg.dev`
@@ -55,7 +55,7 @@ _Rendering fidelity:_ `catalog-parity: charts are rendered with the shared recip
5555
| kueue | helm | kueue | 0.18.2 | 1 |
5656
| mariadb-operator | helm | mariadb-operator | 26.6.0 | 1 |
5757
| mariadb-operator-crds | helm | mariadb-operator-crds | 26.6.0 | 0 |
58-
| network-operator | helm | nvidia/network-operator | 26.4.1 | 5 |
58+
| network-operator | helm | nvidia/network-operator | 26.4.1 | 9 |
5959
| network-operator-ocp | manifest ||| 0 |
6060
| network-operator-ocp-olm | manifest ||| 0 |
6161
| nfd | helm | node-feature-discovery | 0.19.0 | 1 |
@@ -233,6 +233,10 @@ _No images extracted._
233233
### network-operator
234234

235235
- `docker.io/library/busybox:1.38.0@sha256:dc2d74b28e4cf8984fa52af1f39bc7c3d9c73760b41a74d629f5d11b1ab28616`
236+
- `ghcr.io/k8snetworkplumbingwg/multus-cni:v4.2.1`
237+
- `ghcr.io/k8snetworkplumbingwg/plugins:v1.6.2-update.1`
238+
- `ghcr.io/k8snetworkplumbingwg/sriov-network-device-plugin:v3.9.0`
239+
- `ghcr.io/mellanox/nvidia-k8s-ipam:v0.2.0`
236240
- `nvcr.io/nvidia/cloud-native/network-operator:v26.4.1`
237241
- `nvcr.io/nvidia/doca/doca_telemetry:1.22.5-doca3.1.0-host`
238242
- `nvcr.io/nvidia/mellanox/doca-driver:doca3.2.0-25.10-1.2.8.0-2`

pkg/bundler/testdata/stock_render_golden.yaml

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -17,7 +17,7 @@ gb200-eks-ubuntu-inference-dynamo: faed7a104d1bc1ee252d77b971df7b8afc71495eb2292
1717
gb200-eks-ubuntu-training-kubeflow: 5b393b53870521ec1ad82012080bdda259276b1fafafde18cc89f6d3422b612b
1818
gb200-eks-ubuntu-training-slurm: df81536ab91019c2cacb69ad747a3103a7829bb5b90c4f34395ec78e4c36b858
1919
gb200-oke-ubuntu-inference-dynamo: 89524fb57cca12f10bab5747bd7bf584c5a93e35fa29df6b2ff169303f345216
20-
gb200-oke-ubuntu-training-kubeflow: 606e95bf8e55120e2df56e7f6715e8f52d717198bb014dfccacd7a809a055dc1
20+
gb200-oke-ubuntu-training-kubeflow: 86c522cd4280fb25fef783d6a03c2ca05ce8ebcaf3472db4fdfd1ee950044340
2121
h100-aks-ubuntu-inference-dynamo: 966031acc3abedf7a0289530b9a37af29596d5896831030ff372f6bf1bdc4f6f
2222
h100-aks-ubuntu-training-kubeflow: e2ae0d0021d961ebef8e8d857aab32b9fe2015f5c42172b8ccfad1fee0fb9539
2323
h100-aks-ubuntu-training-slurm: 01507beb65f458fcb1bbf1eee688d92b199912fa7af108ea1e4d0727f6ce1a32
@@ -38,7 +38,7 @@ h200-eks-inference: e9c38a77ae6067ce8a8ffcffde56f87a0bfe3c311a73197d95766aa73c9e
3838
h200-eks-training: a7951808f09ed1aa30300b0f62954657df7063996fff4455d559d7e5b4f5ae0a
3939
l40s-any: 068ed1d8149883225e51e2554a2feb5632992aac90c07cc0dc458e7d7983d13f
4040
l40s-oke-inference: d396a6d8b01065a3f64b12ee4a2801b71d4a0031c6e5831000993da721342331
41-
l40s-oke-training: 2da8c3f72fe690d8ee6a3b01ae089b0627bb8d6505ec000333f7b35a5063f275
41+
l40s-oke-training: ee6caf6926b6ac1b1d83368d36e17bb4a8be7f1a3980cceaafe55979c8d6f6d6
4242
monitoring-hpa: 832b485a6dbd53b9cbb275305415ea77e9f2bef5ec5e70382330b9904571bc72
4343
ocp-inference-nim: bb00cdb191823b32da334bea70826c8a92c11b62d0e096918d43a8e1b043c361
4444
ocp-training: d7a213263630f2c25982d6f4a144df7d9d5784428d296ccda1b4dab5a42b98bb

pkg/recipe/performance_goals_oke_test.go

Lines changed: 6 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -34,22 +34,25 @@ func TestOKEPerformanceGoalsFollowTrainingInferencePattern(t *testing.T) {
3434
}{
3535
{
3636
name: "gb200-oke-training",
37-
wantChecks: []string{"nccl-all-reduce-bw-nvls"},
37+
wantChecks: []string{"nccl-all-reduce-bw-net", "nccl-all-reduce-bw-nvls"},
3838
wantConstraints: map[string]string{
39+
"nccl-all-reduce-bw-net": ">= 40",
3940
"nccl-all-reduce-bw-nvls": ">= 500",
4041
},
4142
},
4243
{
4344
name: "gb200-oke-ubuntu-training",
44-
wantChecks: []string{"nccl-all-reduce-bw-nvls"},
45+
wantChecks: []string{"nccl-all-reduce-bw-net", "nccl-all-reduce-bw-nvls"},
4546
wantConstraints: map[string]string{
47+
"nccl-all-reduce-bw-net": ">= 40",
4648
"nccl-all-reduce-bw-nvls": ">= 500",
4749
},
4850
},
4951
{
5052
name: "gb200-oke-ubuntu-training-kubeflow",
51-
wantChecks: []string{"nccl-all-reduce-bw-nvls"},
53+
wantChecks: []string{"nccl-all-reduce-bw-net", "nccl-all-reduce-bw-nvls"},
5254
wantConstraints: map[string]string{
55+
"nccl-all-reduce-bw-net": ">= 40",
5356
"nccl-all-reduce-bw-nvls": ">= 500",
5457
},
5558
},

pkg/recipe/testdata/catalog_parity_golden.yaml

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -17,7 +17,7 @@ gb200-eks-ubuntu-inference-dynamo: 9e6bf776ce05caaf8072cedeaa4cdb46110c91409b365
1717
gb200-eks-ubuntu-training-kubeflow: 8f680335b3480755ba9386922372422a74ed6b563eba868b952aed603056a6ed
1818
gb200-eks-ubuntu-training-slurm: 2f575bbbc126ae761d0333c133dae2dc068deba7289637b59d811c5ebeffcd45
1919
gb200-oke-ubuntu-inference-dynamo: abc880c4e0fcefcb5902ec34f94c5d3b421bc4eda9a5b33a4b5a29e3dc2e048d
20-
gb200-oke-ubuntu-training-kubeflow: 75ad2dd381a86a636525de7dadc61923e74ef90715dd417c191f1e8bf979b0cb
20+
gb200-oke-ubuntu-training-kubeflow: ed509ad8ccd4951fe0198b11033a8f312b41cd12fa1e312cfad42badcc50225f
2121
h100-aks-ubuntu-inference-dynamo: 623d87a7206ef064ec4bb088d082febeedbfbbfd3ba1e7a0edc30cd75e9f796d
2222
h100-aks-ubuntu-training-kubeflow: 4f557906f610b6c71eea7c48ab356a8b60b3e7da7a7a3537093f0633418c58ca
2323
h100-aks-ubuntu-training-slurm: 3918af534fd888957654dea0e89f08aea2fc31e57316773a37eff24b23c231cc
@@ -38,7 +38,7 @@ h200-eks-inference: cb67a2c82c7e5c4766ad74e2184d4714c226853d84272196369c007835d4
3838
h200-eks-training: f622509c221285e5fff866b6311b0f10c0be827d3de014b57b8e78e843b7dfca
3939
l40s-any: 594d6a6ad6b7a943e4c400fdd52f0f7fb4cddb59214812f23a3b158e49afc6a5
4040
l40s-oke-inference: 17053c54993f726338082dc541ae34a3308893450e582e2b479d79337cc4aa9f
41-
l40s-oke-training: 8db430710ae40b810c5ed356f8acc038af855616a0cd2402bfce2e0fdc07ea77
41+
l40s-oke-training: c2ad43da9182132a5a449d9ebae91f587b019f39a5aaf6296f79b26c3c759bcf
4242
monitoring-hpa: f281c5b4c34b0aa0a5505baf24378fdad7afb06269e939e9004793c240035cf3
4343
ocp-inference-nim: f57147ace807d49644443fdbed18bd6e6c20fa028f2144eb2d79989eb4271d5b
4444
ocp-training: d48c15a49e3c8b4f59aff6936c054bd714f54812141be7679961b7b669e3351d
Lines changed: 48 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,48 @@
1+
# NicClusterPolicy for GB200 OKE (OCI) — rdmaSharedDevicePlugin over InfiniBand.
2+
#
3+
# Mirrors the AOR OCI GB200 config validated on gb200-ew. No ofedDriver (host
4+
# MOFED), no SR-IOV: the NVL72 east-west fabric is IB on the rdma0-3 netdevs
5+
# (oci_hpc.rdma_device_names_mode=2 kernel cmdline names them deterministically).
6+
#
7+
# The IB devices are advertised as nvidia.com/mlnxnics — the same resource
8+
# name the L40S SR-IOV path uses, so workloads request RDMA uniformly
9+
# across OKE fabrics.
10+
apiVersion: mellanox.com/v1alpha1
11+
kind: NicClusterPolicy
12+
metadata:
13+
name: nic-cluster-policy
14+
annotations:
15+
helm.sh/hook: post-install,post-upgrade
16+
helm.sh/hook-weight: "5"
17+
helm.sh/hook-delete-policy: before-hook-creation
18+
labels:
19+
app.kubernetes.io/managed-by: {{ .Release.Service }}
20+
helm.sh/chart: {{ printf "%s-%s" .Chart.Name .Chart.Version | replace "+" "_" | trunc 63 | trimSuffix "-" }}
21+
spec:
22+
rdmaSharedDevicePlugin:
23+
image: k8s-rdma-shared-dev-plugin
24+
repository: nvcr.io/nvidia/mellanox
25+
version: network-operator-v26.4.1
26+
config: |
27+
{
28+
"configList": [
29+
{
30+
"resourcePrefix": "nvidia.com",
31+
"resourceName": "mlnxnics",
32+
"rdmaHcaMax": 63,
33+
"selectors": {
34+
"linkTypes": ["infiniband"],
35+
"ifNames": ["rdma0", "rdma1", "rdma2", "rdma3"]
36+
}
37+
}
38+
]
39+
}
40+
deploymentTolerations:
41+
- key: CriticalAddonsOnly
42+
operator: Exists
43+
tolerations:
44+
# RDMA DaemonSets must land on tainted GPU nodes.
45+
- key: nvidia.com/gpu
46+
operator: Exists
47+
- key: CriticalAddonsOnly
48+
operator: Exists
Lines changed: 73 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,73 @@
1+
# NicClusterPolicy for L40S OKE (OCI) SR-IOV RoCE.
2+
#
3+
# The network-operator Helm chart installs the operator + CRD but does not template
4+
# a NicClusterPolicy CR (values-oke-l40s.yaml sets deployCR: false). This manifest
5+
# creates it so the operator reconciles the RoCE fabric stack. Hand-rendered from
6+
# AOR's network-operator/nicclusterpolicy.yaml.tmpl (provider: oci branch, with
7+
# network.type == roce → nvIpam + secondaryNetwork included).
8+
#
9+
# OCI specifics (vs Forge IB): NO ofedDriver — OCI nodes carry host MOFED, consumed
10+
# by the GPU Operator driver via driver.rdma.useHostMofed (l40s-oke-ubuntu leaf). One
11+
# sriovDevicePlugin resource, nvidia.com/mlnxnics, selecting the OCI ConnectX VF
12+
# device IDs (101a = ConnectX-5 Ex VF, 101e = mlx5Gen VF). RoCE also needs nv-ipam
13+
# (VF IP allocation) + secondaryNetwork/multus (attach the VF into workload pods).
14+
# vendor 15b3 = Mellanox.
15+
apiVersion: mellanox.com/v1alpha1
16+
kind: NicClusterPolicy
17+
metadata:
18+
name: nic-cluster-policy
19+
annotations:
20+
helm.sh/hook: post-install,post-upgrade
21+
helm.sh/hook-weight: "5"
22+
helm.sh/hook-delete-policy: before-hook-creation
23+
labels:
24+
app.kubernetes.io/managed-by: {{ .Release.Service }}
25+
helm.sh/chart: {{ printf "%s-%s" .Chart.Name .Chart.Version | replace "+" "_" | trunc 63 | trimSuffix "-" }}
26+
spec:
27+
# RoCE: allocate IPs for the RDMA VFs and wire them into pods via multus.
28+
nvIpam:
29+
image: nvidia-k8s-ipam
30+
repository: ghcr.io/mellanox
31+
version: v0.2.0
32+
enableWebhook: false
33+
containerResources:
34+
- name: nv-ipam-node
35+
requests:
36+
cpu: 500m
37+
memory: 1Gi
38+
limits:
39+
cpu: "1"
40+
memory: 2Gi
41+
secondaryNetwork:
42+
cniPlugins:
43+
image: plugins
44+
repository: ghcr.io/k8snetworkplumbingwg
45+
version: v1.6.2-update.1
46+
multus:
47+
image: multus-cni
48+
repository: ghcr.io/k8snetworkplumbingwg
49+
version: v4.2.1
50+
sriovDevicePlugin:
51+
image: sriov-network-device-plugin
52+
repository: ghcr.io/k8snetworkplumbingwg
53+
version: v3.9.0
54+
config: |
55+
{
56+
"resourceList": [
57+
{
58+
"resourcePrefix": "nvidia.com",
59+
"resourceName": "mlnxnics",
60+
"selectors": {"isRdma":true,"vendors":["15b3"],"devices":["101a","101e"]}
61+
}
62+
]
63+
}
64+
# Operator DaemonSet placement: system/monitoring nodes only (matches AOR).
65+
deploymentTolerations:
66+
- key: CriticalAddonsOnly
67+
operator: Exists
68+
tolerations:
69+
# RDMA DaemonSets must land on tainted GPU nodes.
70+
- key: nvidia.com/gpu
71+
operator: Exists
72+
- key: CriticalAddonsOnly
73+
operator: Exists
Lines changed: 25 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,25 @@
1+
# network-operator Helm values for GB200 OKE (OCI) InfiniBand.
2+
#
3+
# OCI GB200 NVL72 model (vs L40S RoCE / Forge IB): NO ofedDriver — nodes carry host
4+
# MOFED — and no SR-IOV/nv-ipam/multus either. East-west is InfiniBand (rdma0-3),
5+
# served by rdmaSharedDevicePlugin from the post-install NicClusterPolicy manifest,
6+
# NOT the chart. deployCR off so that manifest CR is authoritative.
7+
# nfd.enabled: false — GPU Operator's NFD is used; no second NFD.
8+
deployCR: false
9+
nvIpam:
10+
enabled: false
11+
secondaryNetwork:
12+
deploy: false
13+
nfd:
14+
enabled: false
15+
operator:
16+
resources:
17+
limits:
18+
cpu: "1"
19+
memory: 2Gi
20+
requests:
21+
cpu: 500m
22+
memory: 2Gi
23+
# Operator placement comes from the bundler's system-node scheduling
24+
# injection (registry nodeScheduling: operator.nodeSelector /
25+
# operator.tolerations) — no hardcoded affinity here.
Lines changed: 34 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,34 @@
1+
# network-operator Helm values for L40S OKE (OCI) SR-IOV RoCE.
2+
# Hand-rendered from AOR's network-operator/values.yaml.tmpl (provider: oci) +
3+
# nicclusterpolicy.yaml.tmpl (oci branch, network.type == roce).
4+
#
5+
# OCI model (vs Forge IB / Mistral DOCA): NO ofedDriver — OCI bare-metal nodes carry
6+
# host MOFED, so the GPU Operator uses it via driver.rdma.useHostMofed (set on the
7+
# l40s-oke-ubuntu leaf). network-operator's job here is the SR-IOV VF device plugin
8+
# (advertises nvidia.com/mlnxnics RDMA VFs) plus nv-ipam + secondaryNetwork (multus)
9+
# for RoCE — all supplied by the post-install NicClusterPolicy manifest, NOT the chart.
10+
#
11+
# deployCR/nvIpam/secondaryNetwork: AICR's wrapper defaults are on (deployCR: true,
12+
# nvIpam.enabled: true, secondaryNetwork.deploy: true) — they template the wrapper's
13+
# own NicClusterPolicy. Turn deployCR off so our manifest CR is authoritative (it is
14+
# the only place the OCI VF selectors 101a/101e can be expressed); the operator
15+
# reconciles nv-ipam + secondaryNetwork + sriovDevicePlugin from that CR regardless.
16+
# nfd.enabled: false — GPU Operator's NFD is used; no second NFD.
17+
deployCR: false
18+
nvIpam:
19+
enabled: false
20+
secondaryNetwork:
21+
deploy: false
22+
nfd:
23+
enabled: false
24+
operator:
25+
resources:
26+
limits:
27+
cpu: "1"
28+
memory: 2Gi
29+
requests:
30+
cpu: 500m
31+
memory: 2Gi
32+
# Operator placement comes from the bundler's system-node scheduling
33+
# injection (registry nodeScheduling: operator.nodeSelector /
34+
# operator.tolerations) — no hardcoded affinity here.

recipes/manifest_images_test.go

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -100,6 +100,13 @@ var imageDigestExemptions = map[string]string{
100100
// Skyhook Package `containerSHA` field (issue #1031), folded into the
101101
// extracted image ref as `@sha256:...` by pkg/bom.ExtractImagesFromYAML.
102102
"ghcr.io/nvidia/skyhook-packages/shellscript:1.1.1": "Skyhook Package CRD does not accept image digests; tracked via #745 and NVIDIA/nodewright#224",
103+
104+
// NicClusterPolicy (network-operator OKE): same repository/image/version
105+
// triplet schema as the AKS entries above — no digest field in the CRD.
106+
"ghcr.io/mellanox/nvidia-k8s-ipam:v0.2.0": "NicClusterPolicy CRD does not accept image digests; tracked via #745 and Mellanox/network-operator#2555",
107+
"ghcr.io/k8snetworkplumbingwg/multus-cni:v4.2.1": "NicClusterPolicy CRD does not accept image digests; tracked via #745 and Mellanox/network-operator#2555",
108+
"ghcr.io/k8snetworkplumbingwg/plugins:v1.6.2-update.1": "NicClusterPolicy CRD does not accept image digests; tracked via #745 and Mellanox/network-operator#2555",
109+
"ghcr.io/k8snetworkplumbingwg/sriov-network-device-plugin:v3.9.0": "NicClusterPolicy CRD does not accept image digests; tracked via #745 and Mellanox/network-operator#2555",
103110
}
104111

105112
// TestComponentManifestImagesAreDigestPinned asserts that every image

recipes/overlays/gb200-oke-training.yaml

Lines changed: 32 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -39,30 +39,59 @@ spec:
3939
value: ">= 1.34"
4040

4141
componentRefs:
42-
# GB200-specific GPU Operator overrides (inherits valuesFile from oke-training)
42+
# GB200-specific GPU Operator overrides (inherits valuesFile from oke-training).
43+
# kernel-module-params sets NVreg_GrdmaPciTopoCheckOverride=1, required
44+
# for dma-buf attach over the IB fabric (GPUDirect RDMA without
45+
# nvidia-peermem, whose chroot modprobe fails to build against the -64k
46+
# Grace kernel).
4347
- name: gpu-operator
4448
type: Helm
49+
preManifestFiles:
50+
- components/gpu-operator/manifests/kernel-module-params.yaml
4551
dependencyRefs:
4652
- nfd
4753
- cert-manager
4854
- kube-prometheus-stack
4955
overrides:
5056
gdrcopy:
5157
enabled: true
58+
driver:
59+
kernelModuleConfig:
60+
name: nvidia-kernel-module-params
5261

5362
- name: nfd
5463
type: Helm
5564
overrides:
5665
topologyUpdater:
5766
enable: true
5867

68+
# InfiniBand east-west fabric (NVL72 rdma0-3). rdmaSharedDevicePlugin
69+
# advertises the shared HCAs as nvidia.com/mlnxnics; no SR-IOV/nv-ipam
70+
# (that is the L40S RoCE path) and no ofedDriver (OCI nodes carry host
71+
# MOFED). NicClusterPolicy is manifest-supplied (chart deployCR off).
72+
# Present in every gpuStack value; incompatible with Oracle's opt-in
73+
# NvidiaNetworkOperator add-on.
74+
- name: network-operator
75+
type: Helm
76+
valuesFile: components/network-operator/values-oke-gb200.yaml
77+
manifestFiles:
78+
- components/network-operator/manifests/nic-cluster-policy-oke-gb200.yaml
79+
dependencyRefs:
80+
- nfd
81+
- cert-manager
82+
5983
validation:
6084
performance:
61-
# NVLS runtime support is OKE-specific. NET/RDMA is intentionally left
62-
# out until OCI-specific pod RDMA exposure is verified on the testbed.
85+
# Both transport variants: NVLS (MNNVL across the NVL72 IMEX domain)
86+
# and NET (the IB east-west fabric this leaf's NicClusterPolicy
87+
# exposes — validated on a BM.GPU.GB200.4 NVL72 rack). Constraints
88+
# match gb200-eks-training.
6389
checks:
90+
- nccl-all-reduce-bw-net
6491
- nccl-all-reduce-bw-nvls
6592
constraints:
93+
- name: nccl-all-reduce-bw-net
94+
value: ">= 40"
6695
- name: nccl-all-reduce-bw-nvls
6796
value: ">= 500"
6897
conformance:

0 commit comments

Comments
 (0)