Skip to content

Commit 90dd642

Browse files
committed
Merge PR #8: CPU bucket cache + selective weight sync (zhenyulincs)
Conflict resolution: - coordinator.py: kept _nemo_rl_model_update_service (NeMo RL service, correct interface) and tgt_dp_ranks call signature; used PR #8 more descriptive error message - test_nemo_rl_pipeline.py: kept HEAD F5/F6 pipeline tests intact - test_bucket_cache_lifecycle.py: extracted PR #8 BucketCacheLifecycle tests to new file - test_gap_ratio.py: minor variable rename (g -> gpu for clarity)
2 parents 62c9426 + fe59782 commit 90dd642

43 files changed

Lines changed: 9608 additions & 110 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.
Lines changed: 44 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,44 @@
1+
name: Claude Code Review
2+
3+
on:
4+
pull_request:
5+
types: [opened, synchronize, ready_for_review, reopened]
6+
# Optional: Only run on specific file changes
7+
# paths:
8+
# - "src/**/*.ts"
9+
# - "src/**/*.tsx"
10+
# - "src/**/*.js"
11+
# - "src/**/*.jsx"
12+
13+
jobs:
14+
claude-review:
15+
# Optional: Filter by PR author
16+
# if: |
17+
# github.event.pull_request.user.login == 'external-contributor' ||
18+
# github.event.pull_request.user.login == 'new-developer' ||
19+
# github.event.pull_request.author_association == 'FIRST_TIME_CONTRIBUTOR'
20+
21+
runs-on: ubuntu-latest
22+
permissions:
23+
contents: read
24+
pull-requests: read
25+
issues: read
26+
id-token: write
27+
28+
steps:
29+
- name: Checkout repository
30+
uses: actions/checkout@v4
31+
with:
32+
fetch-depth: 1
33+
34+
- name: Run Claude Code Review
35+
id: claude-review
36+
uses: anthropics/claude-code-action@v1
37+
with:
38+
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
39+
plugin_marketplaces: 'https://github.com/anthropics/claude-code.git'
40+
plugins: 'code-review@claude-code-plugins'
41+
prompt: '/code-review:code-review ${{ github.repository }}/pull/${{ github.event.pull_request.number }}'
42+
# See https://github.com/anthropics/claude-code-action/blob/main/docs/usage.md
43+
# or https://code.claude.com/docs/en/cli-reference for available options
44+

.github/workflows/claude.yml

Lines changed: 50 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,50 @@
1+
name: Claude Code
2+
3+
on:
4+
issue_comment:
5+
types: [created]
6+
pull_request_review_comment:
7+
types: [created]
8+
issues:
9+
types: [opened, assigned]
10+
pull_request_review:
11+
types: [submitted]
12+
13+
jobs:
14+
claude:
15+
if: |
16+
(github.event_name == 'issue_comment' && contains(github.event.comment.body, '@claude')) ||
17+
(github.event_name == 'pull_request_review_comment' && contains(github.event.comment.body, '@claude')) ||
18+
(github.event_name == 'pull_request_review' && contains(github.event.review.body, '@claude')) ||
19+
(github.event_name == 'issues' && (contains(github.event.issue.body, '@claude') || contains(github.event.issue.title, '@claude')))
20+
runs-on: ubuntu-latest
21+
permissions:
22+
contents: read
23+
pull-requests: read
24+
issues: read
25+
id-token: write
26+
actions: read # Required for Claude to read CI results on PRs
27+
steps:
28+
- name: Checkout repository
29+
uses: actions/checkout@v4
30+
with:
31+
fetch-depth: 1
32+
33+
- name: Run Claude Code
34+
id: claude
35+
uses: anthropics/claude-code-action@v1
36+
with:
37+
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
38+
39+
# This is an optional setting that allows Claude to read CI results on PRs
40+
additional_permissions: |
41+
actions: read
42+
43+
# Optional: Give a custom prompt to Claude. If this is not specified, Claude will perform the instructions specified in the comment that tagged it.
44+
# prompt: 'Update the pull request description to include a summary of changes.'
45+
46+
# Optional: Add claude_args to customize behavior and configuration
47+
# See https://github.com/anthropics/claude-code-action/blob/main/docs/usage.md
48+
# or https://code.claude.com/docs/en/cli-reference for available options
49+
# claude_args: '--allowed-tools Bash(gh pr *)'
50+

.gitignore

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -67,6 +67,7 @@ output/
6767
# External dependencies (allow git submodules)
6868
external/*
6969
!external/ROLL
70+
!external/NeMo
7071

7172
## Internal / personal files (not for release)
7273
#CLAUDE.md

.gitmodules

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -2,3 +2,6 @@
22
path = external/ROLL
33
url = https://github.com/rlops/ROLL.git
44
branch = rlix
5+
[submodule "external/NeMo"]
6+
path = external/NeMo
7+
url = https://github.com/zhenyulincs/RL.git

TASK2.md

Lines changed: 254 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,254 @@
1+
# Task 2 — CPU Bucket Cache + 选择性权重同步 (F4, F6-transport)
2+
3+
**规格文档**: [nemorl-port-plan.md](https://github.com/rlops/rlix/blob/nemo/plans/nemorl-port-plan.md) — Feature 4 + Feature 6
4+
**Gate**: 2.5 — 全部 6 个 GPU 集成测试通过(4× RTX A5000)
5+
**代码分支**: `task2-bucket-cache` (rlix) · `rlix-task2` / `main` (NeMo 子模块)
6+
7+
---
8+
9+
## Feature 4 — 训练侧 CPU Bucket Cache
10+
11+
### 规格要求 → 实现位置
12+
13+
| 规格要求 | 实现文件 | 说明 |
14+
|---------|---------|------|
15+
| 所有 TP/PP/CP/EP rank 参与 gather,只有 cache owner 存储 | `external/NeMo/nemo_rl/models/policy/workers/megatron_policy_worker.py``build_latest_bucket_cache()` | owner = pp0/dp0/tp0/cp0,非 owner drain iterator 但不存储 |
16+
| 打包为 canonical `List[BucketRecord]`(512字节对齐 uint8) | `rlix/pipeline/bucket_cache.py``BucketRecord`, `_bucket_named_tensors()` | 包含 `param_names`, `shapes`, `dtypes`, `offsets`, `used_bytes`, `cpu_uint8_bucket` |
17+
| 接收侧 unpack 还原各 tensor | `rlix/pipeline/bucket_cache.py``unpack_bucket_record()` |`torch.empty(0, dtype=dtype).element_size()` 计算字节宽度,避免 uint8 slice 非法 view |
18+
| `_cache_ready_step` 原子更新(版本指针) | `rlix/pipeline/bucket_cache.py``VersionedBucketCache.promote()` | 两指针设计:`_latest_cached` / `_active_cached`,promote 后 GC 旧版本 |
19+
| 生命周期追踪 | `rlix/pipeline/bucket_cache_lifecycle.py``BucketCacheLifecycle` | `build_latest_bucket_cache.remote()``promote_active_checkpoint.remote()``mark_promoted()` |
20+
| `bucket_size_bytes` 必须显式配置,禁止隐式默认 | `megatron_policy_worker.py``_rlix_get_bucket_size_bytes()` | 未配置则 `raise RuntimeError`,读取 `RLIX_BUCKET_SIZE_BYTES``worker.cfg['rlix']['bucket_size_bytes']` |
21+
| 单个 param > bucket_size_bytes → fail fast | `megatron_policy_worker.py``build_latest_bucket_cache()` | append 前检查,匹配 ROLL `send_recv_utils.py` 的 assert 模式 |
22+
| host RAM 检查:2 × model_bytes < 80% available | `megatron_policy_worker.py``build_latest_bucket_cache()` | 用实际打包后的 `total_bytes`,而非 per-bucket 大小 |
23+
| `_cache_lock` 贯穿 cache lookup → transport → NCCL teardown | `megatron_policy_worker.py``selective_sync_active_cache()` | `with cache._cache_lock:` 覆盖整个 bucket 循环 + sender 侧 NCCL destroy |
24+
| Pipeline 层 init / post-train 调用序列 | `rlix/pipeline/full_finetune_pipeline.py` | init: `build_latest_bucket_cache(-1)``promote_active_checkpoint(version=-1)``mark_promoted(-1)` |
25+
26+
### 关键设计决策
27+
28+
- **两指针缓存**`_latest_cached` / `_active_cached`):比规格要求的单槽 `_cache_ready_step` 更安全,防止并发 build/promote 竞争
29+
- **receiver 侧 IPC 路径不走 CPU 中转**`cuda_ipc` 模式直接 `rebuild_cuda_tensor()` 得到 GPU tensor,无 GPU→CPU→GPU roundtrip
30+
- **receiver rank mask 用 `self.rank`**:不用 `dist.get_rank()`,因为 ipc_local_ranks 是 vLLM worker 本地 rank,非分布式 rank
31+
32+
---
33+
34+
## Feature 6 — 选择性权重同步(两条刷新路径)
35+
36+
### 规格要求 → 实现位置
37+
38+
| 规格要求 | 实现文件 | 说明 |
39+
|---------|---------|------|
40+
| `coordinator.sync_base_weights_to_active()` — training loop 刷新 active ranks | `rlix/pipeline/coordinator.py` + `rlix/protocol/coordinator.py` |`_resize_sync_lock`,snapshot `_active_infer_dp_ranks`,直接调 `ModelUpdateService.sync_selected_workers()` |
41+
| `_expand_workers()` — expand 时刷新 woken ranks | `rlix/pipeline/full_finetune_pipeline.py``_expand_workers()` | 顺序:sync → finalize → **version publish(先于 routing 激活)** → expand_sampler |
42+
| ModelUpdateService 6-phase 同步流程 | `rlix/pipeline/model_update_service.py``sync_selected_workers()` | Phase 1: NCCL setup / Phase 2: sender dispatch / Phase 3: receiver teardown / Phase 4: verify |
43+
| IPC vs NCCL broadcast 路由分类 | `model_update_service.py``_build_comm_plan_for_sender()` | 按 (node_rank, gpu_rank) 判断是否同一物理 GPU,同 GPU → IPC,跨 GPU → NCCL |
44+
| **CUDA IPC**(同一物理 GPU,不能建 NCCL group) | `megatron_policy_worker.py``selective_sync_active_cache()` | `get_handle_from_tensor(staging_buf)` 产生 IPC handle,随 payload 发给 receiver |
45+
| **CUDA IPC receiver**(零拷贝) | `external/NeMo/nemo_rl/models/generation/vllm/vllm_backend.py``update_parameter_in_bucket()` | `rebuild_cuda_tensor(*ipc_args)` 直接拿到 GPU tensor,无 CPU 中转 |
46+
| **NCCL broadcast**(跨 GPU,tp > 1) | `megatron_policy_worker.py``selective_sync_active_cache()` | stage CPU→GPU → `dist.broadcast(staging_buf, group=nccl_group)` |
47+
| 动态 NCCL group 创建/销毁 | `megatron_policy_worker.py``setup_collective_group()` / `destroy_collective_group()` | sender 在 `_cache_lock` 内 destroy;receiver 侧由 ModelUpdateService Phase 3 触发 |
48+
| 全部 6 个 receiver API | `vllm_backend.py` + `vllm_generation.py` | `setup_collective_group`, `update_parameter_in_bucket`, `broadcast_parameter`, `destroy_collective_group`, `verify_model`, `finalize_weight_update` |
49+
| vllm_generation pass-through 必须 await sub-worker | `vllm_generation.py` 全部 6 个方法 | 每个方法内 `ray.get(futures)` 确保 outer barrier 语义正确 |
50+
| **finalize_weight_update** — pipeline 所有,worker 执行 | `full_finetune_pipeline.py` | sync 返回后,pipeline 对每个 synced rank 调 `finalize_weight_update.remote()`;ModelUpdateService 不调 |
51+
| version publish 必须在 routing 激活**之前** | `full_finetune_pipeline.py``_expand_workers()` | `set_weight_version.remote(v)``expand_sampler(skip_load=True)` 顺序固定 |
52+
| trajectory collector 版本通知 | `vllm_backend.py` / `grpo.py` / `full_finetune_pipeline.py` | grpo.py 将 collector 注册为命名 Ray actor `rlix:trajectory_collector:{id}`;pipeline 通过 `_get_trajectory_collector()` 懒加载后调 `set_weight_version` |
53+
| port claim 在 teardown 完成后释放,失败时故意泄漏 | `model_update_service.py` | receiver teardown(Phase 3)完成后才 `_release_master_port_claim()`,异常时 finally 不 release |
54+
55+
### 版本号语义
56+
57+
```
58+
train step 3 完成: _cache_ready_step = 3
59+
active refresh: _current_weight_version = 3 (无 bump)
60+
collector.set_weight_version(3)
61+
later expand: collector.set_weight_version(3) (同一版本,无 bump)
62+
```
63+
64+
两条路径刷新的权重相同,版本号相同,避免双重递增。
65+
66+
### transport 模式选择
67+
68+
| 模式 | 场景 | 机制 |
69+
|------|------|------|
70+
| `cuda_ipc` | 同物理 GPU(colocated) | `get_handle_from_tensor()` → IPC handle → `rebuild_cuda_tensor()` |
71+
| `cpu_serialize` | 跨 GPU(默认) | CPU uint8 bucket dict → Ray RPC → `pin_memory().to(device)` |
72+
| NCCL broadcast | 跨 GPU,tp > 1 | `dist.broadcast()` on dynamic group `[sender] + [infer_ranks]` |
73+
74+
> **规格约束**(line 316):NCCL 无法在同一物理 GPU 的两个进程之间建组。同 GPU 的 colocated worker **必须** 走 CUDA IPC,这是正确性要求,不是性能优化。
75+
76+
---
77+
78+
## 文件索引
79+
80+
### rlix 主仓库(`zhenyulincs/rlix`
81+
82+
```
83+
rlix/pipeline/bucket_cache.py BucketRecord, VersionedBucketCache, pack/unpack
84+
rlix/pipeline/bucket_cache_lifecycle.py BucketCacheLifecycle(版本追踪)
85+
rlix/pipeline/model_update_service.py ModelUpdateService(Ray actor,6-phase 同步)
86+
rlix/pipeline/coordinator.py sync_base_weights_to_active()(具体实现)
87+
rlix/pipeline/full_finetune_pipeline.py _expand_workers, finalize, version publish
88+
rlix/protocol/coordinator.py 抽象协议接口
89+
```
90+
91+
### NeMo 子模块(`zhenyulincs/RL`,分支 `rlix-task2` / `main`
92+
93+
```
94+
nemo_rl/models/policy/workers/megatron_policy_worker.py
95+
build_latest_bucket_cache() — 所有 rank gather,owner 打包存储
96+
promote_active_checkpoint() — 切换 active 指针
97+
selective_sync_active_cache() — sender 主逻辑(IPC + NCCL)
98+
setup_collective_group() — 加入动态 NCCL group
99+
destroy_collective_group() — 销毁动态 NCCL group
100+
101+
nemo_rl/models/generation/vllm/vllm_backend.py
102+
update_parameter_in_bucket() — receiver IPC 路径(CUDA IPC / cpu_serialize)
103+
broadcast_parameter() — receiver NCCL broadcast 路径
104+
finalize_weight_update() — post-bucket hook(FP8 等)
105+
verify_model() — 可选验证
106+
setup_collective_group() — receiver 侧加入 NCCL group
107+
destroy_collective_group() — receiver 侧销毁 NCCL group
108+
109+
nemo_rl/models/generation/vllm/vllm_generation.py
110+
(以上 6 个方法的 Ray actor pass-through,每个内部 ray.get(futures) 确保 barrier)
111+
112+
nemo_rl/algorithms/grpo.py
113+
trajectory_collector 注册为命名 Ray actor: rlix:trajectory_collector:{pipeline_id}
114+
```
115+
116+
---
117+
118+
## 测试文件说明
119+
120+
### 第 0 步:环境检查(每次新机器必跑,其他测试之前)
121+
122+
```bash
123+
# 检查 setup.py / pyproject.toml VCS hash 一致性、子模块初始化、核心模块可导入
124+
python tests/test_env_install.py
125+
```
126+
127+
### 单元测试(无 GPU / Ray)
128+
129+
```bash
130+
python -m pytest tests/test_bucket_cache.py # BucketRecord pack/unpack
131+
python -m pytest tests/test_bucket_cache_lifecycle.py # 版本追踪、promote、GC
132+
python -m pytest tests/test_model_update_service.py # comm plan、finalize 归属
133+
python -m pytest tests/test_nemo_rl_pipeline.py # _expand_workers 顺序
134+
# 期望:53 passed
135+
```
136+
137+
### Gate 2.5 集成测试(需要 4× GPU,torchrun)
138+
139+
```bash
140+
export NCCL_P2P_DISABLE=1 NCCL_SHM_DISABLE=1 # PCIe 硬件(无 NVLink)
141+
142+
# 1. NCCL destroy/re-init 稳定性(2 GPU)
143+
torchrun --nproc-per-node=2 tests/integration/test_gate2_5_nccl_destroy.py
144+
145+
# 2. NCCL proper-subset group broadcast(4 GPU)
146+
# 验证: group=[0,2,3] 是 world=[0,1,2,3] 的真子集,不会 hang
147+
torchrun --nproc-per-node=4 tests/integration/test_gate2_5_selective_sync.py
148+
149+
# 3. Megatron TP=2 训练 + per-shard NCCL 同步(4 GPU)
150+
# group[0,2] 同步 shard0,group[1,3] 同步 shard1
151+
torchrun --nproc-per-node=4 tests/integration/test_gate2_5_megatron_tp.py
152+
153+
# 4. Qwen2.5-0.5B 真实模型训练 + 同步(4 GPU,需 HF 缓存)
154+
HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 \
155+
torchrun --nproc-per-node=4 tests/integration/test_gate2_5_qwen_train_sync.py
156+
157+
# 5. 双 pipeline 交替同步,A≠B 权重隔离(4 GPU)
158+
HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 \
159+
torchrun --nproc-per-node=4 tests/integration/test_gate2_5_full.py
160+
161+
# 6. F6 顺序验证:sync→finalize→version_publish→activate(4 GPU)
162+
torchrun --nproc-per-node=4 tests/integration/test_gate2_5_feature6.py
163+
```
164+
165+
全部 6 个应输出 `ALL GATE 2.5 * CHECKS PASSED`,exit 0。
166+
167+
### F6.3 / F4.4 / F6.6 专项测试(单 GPU)
168+
169+
```bash
170+
# CUDA IPC 跨进程零拷贝传输
171+
python tests/integration/test_gate2_5_cuda_ipc.py
172+
173+
# bucket_size_bytes 配置检查(未配置 → RuntimeError;过大 → RAM fail-fast)
174+
python tests/integration/test_gate2_5_bucket_size_guard.py
175+
176+
# version publish 顺序验证(set_weight_version 在 expand_sampler 之前)
177+
python tests/integration/test_gate2_5_trajectory_collector.py
178+
```
179+
180+
### 快速使用示例
181+
182+
```python
183+
# 在测试或调试时手动构造 bucket cache 并验证 pack/unpack
184+
import torch, importlib.util, sys
185+
from pathlib import Path
186+
187+
# 直接加载文件(避免 rlix package __init__ 的重依赖)
188+
def _load(name, path):
189+
spec = importlib.util.spec_from_file_location(name, path)
190+
mod = importlib.util.module_from_spec(spec)
191+
sys.modules[name] = mod; spec.loader.exec_module(mod); return mod
192+
193+
repo = Path(__file__).parent # rlix repo root
194+
bc = _load("rlix.pipeline.bucket_cache", repo / "rlix/pipeline/bucket_cache.py")
195+
_bucket_named_tensors = bc._bucket_named_tensors
196+
unpack_bucket_record = bc.unpack_bucket_record
197+
VersionedBucketCache = bc.VersionedBucketCache
198+
199+
# 1. 打包
200+
named_tensors = [("fc1.weight", torch.randn(256, 256)),
201+
("fc2.weight", torch.randn(256, 256))]
202+
record = _bucket_named_tensors(named_tensors)
203+
print(f"packed: {record.cpu_uint8_bucket.numel()} bytes")
204+
205+
# 2. 缓存
206+
cache = VersionedBucketCache()
207+
cache.build_latest(step=1, buckets=[record])
208+
cache.promote(version=1)
209+
210+
# 3. 读取(持锁)
211+
with cache._cache_lock:
212+
buckets = cache.get_active_buckets()
213+
214+
# 4. 解包还原
215+
for bucket in buckets:
216+
for name, tensor in unpack_bucket_record(bucket):
217+
print(f" {name}: {tensor.shape}, {tensor.dtype}")
218+
219+
# 5. 验证 bit-exact
220+
import hashlib
221+
def h(t): return hashlib.sha256(t.cpu().contiguous().view(torch.uint8).numpy().tobytes()).hexdigest()[:8]
222+
223+
orig = {name: h(t) for name, t in named_tensors}
224+
recv = {name: h(t) for name, t in unpack_bucket_record(buckets[0])}
225+
assert orig == recv, f"mismatch: {orig} vs {recv}"
226+
print("bit-exact ✓")
227+
```
228+
229+
---
230+
231+
## 已知待实现项
232+
233+
| 项目 | 原因 |
234+
|------|------|
235+
| `wake_up_partial()` / `activate_dp_ranks()` | Feature 2(VllmGeneration sleep/wake API)尚未实现,当前用 ROLL 的 `expand_sampler(skip_load=True)` 等效替代 |
236+
| ZMQ ping-pong 双缓冲 IPC | NeMo RL 环境未安装 `zmq`;Ray RPC 实现等效功能 |
237+
| `_cache_ready_step` 在 sender `_cache_lock` 下发布 | 跨 Ray actor 架构约束:training worker 锁 ≠ pipeline 的 lifecycle 锁,不可共享 |
238+
239+
---
240+
241+
## 环境配置
242+
243+
```bash
244+
# 克隆(含子模块)
245+
git clone https://github.com/zhenyulincs/rlix.git --recurse-submodules
246+
cd rlix
247+
248+
# 安装依赖
249+
pip install uv && uv sync
250+
251+
# 必须显式配置(无隐式默认值)
252+
export RLIX_BUCKET_SIZE_BYTES=$((256 * 1024 * 1024)) # 256 MB per bucket
253+
export RLIX_MODEL_UPDATE_TRANSPORT=cpu_serialize # 或 cuda_ipc(同 GPU colocated)
254+
```

0 commit comments

Comments
 (0)