Skip to content

Commit c698bdd

Browse files
Merge pull request #156 from flagos-ai/update/modelscope-docs-20260308-125442
ModelScope Documentation Update - 2026-03-08 12:54
2 parents 31fa9ce + 4da212c commit c698bdd

16 files changed

Lines changed: 94 additions & 133 deletions

docs/flagrelease_en/model_readmes/FlagRelease_Emu3.5-FlagOS.md

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -73,14 +73,14 @@ modelscope download --model FlagRelease/Emu3.5-FlagOS --local_dir /data/Emu3.5-F
7373
### Download FlagOS Image
7474

7575
```bash
76-
docker pull harbor.baai.ac.cn/flagrelease-public/flagrelease_nvidia_emu3p5
76+
docker pull harbor.baai.ac.cn/flagrelease-public/flagrelease-nvidia-release-model_emu3.5-tree_none-gems_4.1-scale_none-cx_none-python_3.12.3-torch_2.8.0-pcp_cuda12.9-gpu_nvidia003-arc_amd64-driver_570.158.01:2511191157
7777
```
7878

7979
### Start the inference service
8080

8181
```bash
8282
#Container Startup
83-
docker run --rm --init --detach --net=host --uts=host --ipc=host --security-opt=seccomp=unconfined --privileged=true --ulimit stack=67108864 --ulimit memlock=-1 --ulimit nofile=1048576:1048576 --shm-size=32G -v /data/Emu3.5-FlagOS:/share --gpus all --name flagos harbor.baai.ac.cn/flagrelease-public/flagrelease_nvidia_emu3p5 sleep infinity
83+
docker run --rm --init --detach --net=host --uts=host --ipc=host --security-opt=seccomp=unconfined --privileged=true --ulimit stack=67108864 --ulimit memlock=-1 --ulimit nofile=1048576:1048576 --shm-size=32G -v /data/Emu3.5-FlagOS:/share --gpus all --name flagos harbor.baai.ac.cn/flagrelease-public/flagrelease-nvidia-release-model_emu3.5-tree_none-gems_4.1-scale_none-cx_none-python_3.12.3-torch_2.8.0-pcp_cuda12.9-gpu_nvidia003-arc_amd64-driver_570.158.01:2511191157 sleep infinity
8484
```
8585

8686
### Use Emu3.5

docs/flagrelease_en/model_readmes/FlagRelease_Hunyuan-A13B-Instruct-FlagOS.md

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -59,7 +59,7 @@ FlagEval (Libra)** is a comprehensive evaluation system and open platform for la
5959
| Docker Version | Docker version 27.5.1, build 27.5.1-0ubuntu3~22.04.2|
6060
| Operating System | Description: Ubuntu 22.04.4 LTS |
6161
| FlagScale | Version: 0.8.0 |
62-
| FlagGems | Version: 4.1 |
62+
| FlagGems | Version: 4.2.1rc0 |
6363

6464
## Operation Steps
6565

@@ -74,7 +74,7 @@ modelscope download --model FlagRelease/Hunyuan-A13B-Instruct-FlagOS --local_dir
7474
### Download FlagOS Image
7575

7676
```bash
77-
docker pull harbor.baai.ac.cn/flagrelease-public/flagrelease-nvidia-release-model_hunyuan-a13b-instruct-tree_none-gems_4.1-scale_0.8.0-cx_none-python_3.12.3-torch_2.9.0-pcp_cuda13.0-gpu_nvidia004-arc_amd64-driver_535.183.06:2512041016
77+
docker pull harbor.baai.ac.cn/flagrelease-public/flagrelease-nvidia-release-model_hunyuan-a13b-instruct-tree_none-gems_4.1-scale_0.8.0-cx_none-python_3.12.3-torch_2.9.0-pcp_cuda13.0-gpu_nvidia004-arc_amd64-driver_535.183.06:260226
7878
```
7979

8080
### Start the inference service
@@ -85,7 +85,7 @@ docker run --init --detach --net=host --user 0 --ipc=host \
8585
-v /data:/data --security-opt=seccomp=unconfined \
8686
--privileged --ulimit=stack=67108864 --ulimit=memlock=-1 \
8787
--shm-size=512G --gpus all -e USE_FLAGGEMS=1 \
88-
--name flagos harbor.baai.ac.cn/flagrelease-public/flagrelease-nvidia-release-model_hunyuan-a13b-instruct-tree_none-gems_4.1-scale_0.8.0-cx_none-python_3.12.3-torch_2.9.0-pcp_cuda13.0-gpu_nvidia004-arc_amd64-driver_535.183.06:2512041016 sleep infinity
88+
--name flagos harbor.baai.ac.cn/flagrelease-public/flagrelease-nvidia-release-model_hunyuan-a13b-instruct-tree_none-gems_4.1-scale_0.8.0-cx_none-python_3.12.3-torch_2.9.0-pcp_cuda13.0-gpu_nvidia004-arc_amd64-driver_535.183.06:260226 sleep infinity
8989
```
9090

9191
### Serve

docs/flagrelease_en/model_readmes/FlagRelease_Kimi-K2-Thinking-FlagOS.md

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -75,7 +75,7 @@ modelscope download --model FlagRelease/Kimi-K2-Thinking-FlagOS --local_dir /sha
7575
### Download FlagOS Image
7676

7777
```bash
78-
docker pull harbor.baai.ac.cn/flagrelease-public/flagrelease_nvidia_kimi_k2_thinking
78+
docker pull harbor.baai.ac.cn/flagrelease-public/flagrelease-nvidia-release-model_kimi-k2-thinking-tree_none-gems_4.1-scale_0.8.0-cx_none-python_3.12.3-torch_2.9.0-pcp_cuda12.9-gpu_nvidia003-arc_amd64-driver_570.158.01:2512151813
7979
```
8080

8181
### Start the inference service
@@ -86,7 +86,7 @@ docker run --init --detach --net=host --user 0 --ipc=host \
8686
-v /share:/share --security-opt=seccomp=unconfined \
8787
--privileged --ulimit=stack=67108864 --ulimit=memlock=-1 \
8888
--shm-size=512G --gpus all -e USE_FLAGGEMS=1 \
89-
--name flagos harbor.baai.ac.cn/flagrelease-public/flagrelease_nvidia_kimi_k2_thinking sleep infinity
89+
--name flagos harbor.baai.ac.cn/flagrelease-public/flagrelease-nvidia-release-model_kimi-k2-thinking-tree_none-gems_4.1-scale_0.8.0-cx_none-python_3.12.3-torch_2.9.0-pcp_cuda12.9-gpu_nvidia003-arc_amd64-driver_570.158.01:2512151813 sleep infinity
9090
```
9191

9292
### Serve

docs/flagrelease_en/model_readmes/FlagRelease_MiniCPM-V-4-metax-FlagOS.md

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -63,7 +63,7 @@ FlagEval (Libra)** is a comprehensive evaluation system and open platform for la
6363
| Docker Version | Docker version 28.0.4, build b8034c0|
6464
| Operating System | Ubuntu 22.04.4 LTS |
6565
| FlagScale | Version: 0.8.0 |
66-
| FlagGems | Version: 3.0 |
66+
| FlagGems | Version: 4.2.1rc0 |
6767

6868
## Operation Steps
6969

@@ -78,14 +78,14 @@ modelscope download --model OpenBMB/MiniCPM-V-4 --local_dir /nfs/MiniCPM-V-4
7878
### Download FlagOS Image
7979

8080
```bash
81-
docker pull harbor.baai.ac.cn/flagrelease-public/flagrelease-metax-release-model_minicpm-v-4-tree_none-gems_3.0-scale_0.8.0-cx_none-python_3.10.10-torch_2.6.0_metax3.0.0.3-pcp_maca3.0.0.8-gpu_metax001-arc_amd64-driver_3.3.12:2510141031
81+
docker pull harbor.baai.ac.cn/flagrelease-public/flagrelease-metax-release-model_minicpm-v-4-tree_none-gems_3.0-scale_0.8.0-cx_none-python_3.10.10-torch_2.6.0_metax3.0.0.3-pcp_maca3.0.0.8-gpu_metax001-arc_amd64-driver_3.3.12:2602271358
8282
```
8383

8484
### Start the inference service
8585

8686
```bash
8787
#Container Startup
88-
docker run -it --device=/dev/dri --device=/dev/mxcd --group-add video --name flagos --device=/dev/mem --network=host --security-opt seccomp=unconfined --security-opt apparmor=unconfined --shm-size '100gb' --ulimit memlock=-1 -v /usr/local/:/usr/local/ -v /nfs:/nfs harbor.baai.ac.cn/flagrelease-public/flagrelease-metax-release-model_minicpm-v-4-tree_none-gems_3.0-scale_0.8.0-cx_none-python_3.10.10-torch_2.6.0_metax3.0.0.3-pcp_maca3.0.0.8-gpu_metax001-arc_amd64-driver_3.3.12:2510141031 /bin/bash
88+
docker run -it --device=/dev/dri --device=/dev/mxcd --group-add video --name flagos --device=/dev/mem --network=host --security-opt seccomp=unconfined --security-opt apparmor=unconfined --shm-size '100gb' --ulimit memlock=-1 -v /usr/local/:/usr/local/ -v /nfs:/nfs docker pull harbor.baai.ac.cn/flagrelease-public/flagrelease-metax-release-model_minicpm-v-4-tree_none-gems_3.0-scale_0.8.0-cx_none-python_3.10.10-torch_2.6.0_metax3.0.0.3-pcp_maca3.0.0.8-gpu_metax001-arc_amd64-driver_3.3.12:2602271358 /bin/bash
8989
```
9090

9191
### Serve

docs/flagrelease_en/model_readmes/FlagRelease_MiniMax-M2-FlagOS.md

Lines changed: 3 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -74,7 +74,7 @@ modelscope download --model FlagRelease/MiniMax-M2-FlagOS --local_dir /share/Min
7474
### Download FlagOS Image
7575

7676
```bash
77-
docker pull harbor.baai.ac.cn/flagrelease-public/flagrelease_nvidia_minimaxm2
77+
docker pull harbor.baai.ac.cn/flagrelease-public/flagrelease-nvidia-release-model_minimax-m2-tree_none-gems_3.0-scale_0.8.0-cx_none-python_3.12.3-torch_2.8.0a0_5228986c39.nv25.6-pcp_cuda12.9-gpu_nvidia003-arc_amd64-driver_570.158.01:2511041437
7878
```
7979

8080
### Start the inference service
@@ -85,7 +85,8 @@ docker run --init --detach --net=host --user 0 --ipc=host \
8585
-v /share:/share --security-opt=seccomp=unconfined \
8686
--privileged --ulimit=stack=67108864 --ulimit=memlock=-1 \
8787
--shm-size=512G --gpus all -e USE_FLAGGEMS=1 \
88-
--name flagos harbor.baai.ac.cn/flagrelease-public/flagrelease_nvidia_minimaxm2 sleep infinity
88+
--name flagos harbor.baai.ac.cn/flagrelease-public/flagrelease-nvidia-release-model_minimax-m2-tree_none-gems_3.0-scale_0.8.0-cx_none-python_3.12.3-torch_2.8.0a0_5228986c39.nv25.6-pcp_cuda12.9-gpu_nvidia003-arc_amd64-driver_570.158.01:2511041437 sleep infinity
89+
docker exec -it flagos /bin/bash
8990
```
9091

9192
### Serve

docs/flagrelease_en/model_readmes/FlagRelease_Qwen2-7B-Instruct-FlagOS.md

Lines changed: 7 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -67,25 +67,26 @@ FlagEval (Libra)** is a comprehensive evaluation system and open platform for la
6767

6868
```bash
6969
pip install modelscope
70-
modelscope download --model Qwen/Qwen2-7B-Instruct --local_dir /nfs/models/Qwen2-7B-Instruct
70+
modelscope download --model Qwen/Qwen2-7B-Instruct --local_dir /data/Qwen2-7B-Instruct
7171

7272
```
7373

7474
### Download FlagOS Image
7575

7676
```bash
77-
docker pull harbor.baai.ac.cn/flagrelease-public/flagrelease_nvidia_qwen2_7b
77+
docker pull harbor.baai.ac.cn/flagrelease-public/flagrelease-nvidia-release-model_qwen2-7b-instruct-tree_none-gems_3.0-scale_0.8.0-cx_none-python_3.12.3-torch_2.8.0-pcp_cuda12.9-gpu_nvidia004-arc_amd64-driver_535.183.06:2512151909
7878
```
7979

8080
### Start the inference service
8181

8282
```bash
8383
#Container Startup
84-
docker run -d -it --net=host --uts=host --ipc=host -e USE_FLAGGEMS=1 \
84+
docker run -d -it --net=host --uts=host --ipc=host -e USE_FLAGGEMS=1 --gpus all \
8585
--privileged=true --group-add video --shm-size 100gb --ulimit memlock=-1 \
8686
--security-opt seccomp=unconfined --security-opt apparmor=unconfined --device=/dev/dri \
87-
--device=/dev/mxcd -v /usr/share/zoneinfo/Asia/Shanghai:/etc/localtime:ro -v /nfs/:/nfs/\
88-
--name qwen2_7b_release harbor.baai.ac.cn/flagrelease-public/flagrelease_nvidia_qwen2_7b bash
87+
--device=/dev/mxcd -v /usr/share/zoneinfo/Asia/Shanghai:/etc/localtime:ro -v /data:/data/\
88+
--name flagos harbor.baai.ac.cn/flagrelease-public/flagrelease-nvidia-release-model_qwen2-7b-instruct-tree_none-gems_3.0-scale_0.8.0-cx_none-python_3.12.3-torch_2.8.0-pcp_cuda12.9-gpu_nvidia004-arc_amd64-driver_535.183.06:2512151909 bash
89+
docker exec -it flagos /bin/bash
8990
```
9091

9192
### Serve
@@ -103,7 +104,7 @@ flagscale serve qwen2
103104
import openai
104105
openai.api_key = "EMPTY"
105106
openai.base_url = "http://<server_ip>:8000/v1/"
106-
model = "/nfs/models/Qwen2-7B-Instruct/"
107+
model = "/data/Qwen2-7B-Instruct"
107108
messages = [
108109
{"role": "system", "content": "You are a helpful assistant."},
109110
{"role": "user", "content": "What's the weather like today?"}

docs/flagrelease_en/model_readmes/FlagRelease_Qwen3-235B-A22B-FlagOS-nvidia.md

Lines changed: 24 additions & 45 deletions
Original file line numberDiff line numberDiff line change
@@ -52,14 +52,14 @@ We use a variety of Triton-implemented operation kernels to run the Qwen3-235B-
5252
```bash
5353

5454
pip install modelscope
55-
modelscope download --model Qwen/Qwen3-235B-A22B --local_dir /nfs/Qwen3-235B-A22B
55+
modelscope download --model Qwen/Qwen3-235B-A22B --local_dir /share/Qwen3-235B-A22B
5656

5757
```
5858

5959
### Download the FlagOS image
6060

6161
```bash
62-
docker pull flagrelease-registry.cn-beijing.cr.aliyuncs.com/flagrelease/flagrelease:flagrelease_nv_robobrain2_32b
62+
docker pull harbor.baai.ac.cn/flagrelease-public/flagrelease-nvidia-release-model_qwen3-235b-a22b-tree_none-gems_2.2-scale_0.8.0-cx_none-python_3.12.10-torch_2.7.0-pcp_cuda12.8-gpu_nvidia003-arc_amd64-driver_570.158.01:2508011525
6363
```
6464

6565
### Start the inference service
@@ -73,61 +73,40 @@ docker run --rm --init --detach \
7373
--ulimit memlock=-1 \
7474
--ulimit nofile=1048576:1048576 \
7575
--shm-size=32G \
76-
-v /nfs:/nfs \
76+
-v /share:/share \
7777
--gpus all \
7878
--name flagos \
79-
flagrelease-registry.cn-beijing.cr.aliyuncs.com/flagrelease/flagrelease:flagrelease_nv_robobrain2_32b \
79+
harbor.baai.ac.cn/flagrelease-public/flagrelease-nvidia-release-model_qwen3-235b-a22b-tree_none-gems_2.2-scale_0.8.0-cx_none-python_3.12.10-torch_2.7.0-pcp_cuda12.8-gpu_nvidia003-arc_amd64-driver_570.158.01:2508011525 \
8080
sleep infinity
8181

8282
docker exec -it flagos bash
8383
```
8484

85-
### Modify serve configuration
86-
87-
1. Use `pip show flag_scale` to find the location of installed flag_scale framework like `/usr/local/lib/python3.12/dist-packages`
88-
2. `cd /usr/local/lib/python3.12/dist-packages/flag_scale/examples/`
89-
3. Use `mv robobrain2 qwen3` to modify the config name
90-
4. `vim qwen3/serve/32b.yaml` to replace the content of RoboBrain2-32B to this model(Qwen3):
91-
92-
```
93-
- serve_id: vllm_model
94-
engine: vllm
95-
engine_args:
96-
model: /share/project/jiyuheng/ckpt/32b_stage2_K1
97-
served_model_name: RoboBrain2.0-32B-nvidia-flagos
98-
tensor_parallel_size: 8
99-
pipeline_parallel_size: 1
100-
gpu_memory_utilization: 0.9
101-
limit_mm_per_prompt: image=18 # should be customized, 18 images/request is enough for most scenarios
102-
port: 9010
103-
trust_remote_code: true
104-
enforce_eager: false # set true if use FlagGems
105-
enable_chunked_prefill: true
85+
### Serve
10686

87+
```bash
88+
flagscale serve qwen3
10789
```
90+
# Service Invocation
10891

109-
you should modify the 32b.yaml to:
92+
## API-based Invocation Script
11093

11194
```
112-
- serve_id: vllm_model
113-
engine: vllm
114-
engine_args:
115-
model: /nfs/Qwen3-235B-A22B
116-
served_model_name: Qwen3-235B-A22B-nvidia-flagos
117-
tensor_parallel_size: 8
118-
pipeline_parallel_size: 1
119-
gpu_memory_utilization: 0.9
120-
port: 9010
121-
trust_remote_code: true
122-
enforce_eager: false # set true if use FlagGems
123-
enable_chunked_prefill: true
124-
125-
```
126-
127-
### Serve
128-
129-
```bash
130-
flagscale serve qwen3
95+
import openai
96+
openai.api_key = "EMPTY"
97+
openai.base_url = "http://<server_ip>:9010/v1/"
98+
model = "Qwen3-235B-A22B-FlagOS-nvidia"
99+
messages = [
100+
{"role": "system", "content": "You are a helpful assistant."},
101+
{"role": "user", "content": "What's the weather like today?"}
102+
]
103+
response = openai.chat.completions.create(
104+
model=model,
105+
messages=messages,
106+
stream=False,
107+
)
108+
for item in response:
109+
print(item)
131110
```
132111

133112
# Contributing

docs/flagrelease_en/model_readmes/FlagRelease_Qwen3-30B-A3B-FlagOS-nvidia.md

Lines changed: 24 additions & 45 deletions
Original file line numberDiff line numberDiff line change
@@ -52,14 +52,14 @@ We use a variety of Triton-implemented operation kernels to run the Qwen3-30B-A
5252
```bash
5353

5454
pip install modelscope
55-
modelscope download --model Qwen/Qwen3-30B-A3B --local_dir /nfs/Qwen3-30B-A3B
55+
modelscope download --model Qwen/Qwen3-30B-A3B --local_dir /share/Qwen3-30B-A3B
5656

5757
```
5858

5959
### Download the FlagOS image
6060

6161
```bash
62-
docker pull flagrelease-registry.cn-beijing.cr.aliyuncs.com/flagrelease/flagrelease:flagrelease_nv_robobrain2_32b
62+
docker pull harbor.baai.ac.cn/flagrelease-public/flagrelease-nvidia-release-model_qwen3-30b-a3b-tree_none-gems_2.2-scale_0.8.0-cx_none-python_3.12.10-torch_2.7.0-pcp_cuda12.2-gpu_nvidia004-arc_amd64-driver_535.183.06:2508011525
6363
```
6464

6565
### Start the inference service
@@ -73,61 +73,40 @@ docker run --rm --init --detach \
7373
--ulimit memlock=-1 \
7474
--ulimit nofile=1048576:1048576 \
7575
--shm-size=32G \
76-
-v /nfs:/nfs \
76+
-v /share:/share \
7777
--gpus all \
7878
--name flagos \
79-
flagrelease-registry.cn-beijing.cr.aliyuncs.com/flagrelease/flagrelease:flagrelease_nv_robobrain2_32b \
79+
harbor.baai.ac.cn/flagrelease-public/flagrelease-nvidia-release-model_qwen3-30b-a3b-tree_none-gems_2.2-scale_0.8.0-cx_none-python_3.12.10-torch_2.7.0-pcp_cuda12.2-gpu_nvidia004-arc_amd64-driver_535.183.06:2508011525 \
8080
sleep infinity
8181

8282
docker exec -it flagos bash
8383
```
8484

85-
### Modify serve configuration
86-
87-
1. Use `pip show flag_scale` to find the location of installed flag_scale framework like `/usr/local/lib/python3.12/dist-packages`
88-
2. `cd /usr/local/lib/python3.12/dist-packages/flag_scale/examples/`
89-
3. Use `mv robobrain2 qwen3` to modify the config name
90-
4. `vim qwen3/serve/32b.yaml` to replace the content of RoboBrain2-32B to this model(Qwen3):
91-
92-
```
93-
- serve_id: vllm_model
94-
engine: vllm
95-
engine_args:
96-
model: /share/project/jiyuheng/ckpt/32b_stage2_K1
97-
served_model_name: RoboBrain2.0-32B-nvidia-flagos
98-
tensor_parallel_size: 8
99-
pipeline_parallel_size: 1
100-
gpu_memory_utilization: 0.9
101-
limit_mm_per_prompt: image=18 # should be customized, 18 images/request is enough for most scenarios
102-
port: 9010
103-
trust_remote_code: true
104-
enforce_eager: false # set true if use FlagGems
105-
enable_chunked_prefill: true
85+
### Serve
10686

87+
```bash
88+
flagscale serve qwen3
10789
```
90+
# Service Invocation
10891

109-
you should modify the 32b.yaml to:
92+
## API-based Invocation Script
11093

11194
```
112-
- serve_id: vllm_model
113-
engine: vllm
114-
engine_args:
115-
model: /nfs/Qwen3-30B-A3B
116-
served_model_name: Qwen3-30B-A3B-nvidia-flagos
117-
tensor_parallel_size: 8
118-
pipeline_parallel_size: 1
119-
gpu_memory_utilization: 0.9
120-
port: 9010
121-
trust_remote_code: true
122-
enforce_eager: false # set true if use FlagGems
123-
enable_chunked_prefill: true
124-
125-
```
126-
127-
### Serve
128-
129-
```bash
130-
flagscale serve qwen3
95+
import openai
96+
openai.api_key = "EMPTY"
97+
openai.base_url = "http://<server_ip>:9010/v1/"
98+
model = "Qwen3-30B-A3B-FlagOS-nvidia"
99+
messages = [
100+
{"role": "system", "content": "You are a helpful assistant."},
101+
{"role": "user", "content": "What's the weather like today?"}
102+
]
103+
response = openai.chat.completions.create(
104+
model=model,
105+
messages=messages,
106+
stream=False,
107+
)
108+
for item in response:
109+
print(item)
131110
```
132111

133112
# Contributing

docs/flagrelease_en/model_readmes/FlagRelease_Qwen3-4B-FlagOS-Ascend.md

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -80,7 +80,7 @@ modelscope download --model Qwen/Qwen3-4B --local_dir /data/weights/Qwen3-4B/
8080
#Download the image for the A3 chip
8181
docker pull harbor.baai.ac.cn/flagrelease-public/flagrelease-ascend-release-model_qwen3-4b-tree_none-gems_2.2-scale_0.8.0-cx_none-python_3.11.11-torch_npu2.6.0rc1-pcp_cann8.2rc1.alpha002-gpu_ascend001-arc_arm64-driver_25.2.0:2512101714
8282
#Download the image for the A2 chip
83-
docker pull flagrelease-registry.cn-beijing.cr.aliyuncs.com/flagrelease/flagrelease:flagopen-910b-ubuntu24.04.2-py311
83+
docker pull harbor.baai.ac.cn/flagrelease-public/flagopen-910b-ubuntu24.04.2-py311_ascend_a2:latest
8484
```
8585

8686
### Start the inference service (A3 chip)
@@ -169,7 +169,7 @@ docker run -itd --name gems_test \
169169
-e PYTORCH_NPU_ALLOC_CONF=max_split_size_mb:256 \
170170
-e USE_FLAGGEMS=true \
171171
-e ASCEND_RT_VISIBLE_DEVICES=7 \
172-
flagrelease-registry.cn-beijing.cr.aliyuncs.com/flagrelease/flagrelease:flagopen-910b-ubuntu24.04.2-py311 bash
172+
harbor.baai.ac.cn/flagrelease-public/flagopen-910b-ubuntu24.04.2-py311_ascend_a2:latest bash
173173

174174
#Enter the container
175175
docker exec -it flagos bash

docs/flagrelease_en/model_readmes/FlagRelease_Qwen3-4B-FlagOS-Metax.md

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -59,7 +59,7 @@ FlagEval (Libra)** is a comprehensive evaluation system and open platform for la
5959
| Docker Version | Docker version 28.0.4, build b8034c0|
6060
| Operating System | Description: Ubuntu 22.04.4 LTS |
6161
| FlagScale | Version: 0.6.0 |
62-
| FlagGems | Version: 2.2 |
62+
| FlagGems | Version: 4.2.1rc0 |
6363

6464
## Operation Steps
6565

@@ -74,19 +74,19 @@ modelscope download --model Qwen/Qwen3-4B --local_dir /nfs/Qwen3-4B
7474
### Download FlagOS Image
7575

7676
```bash
77-
docker pull harbor.baai.ac.cn/flagrelease-public/flagrelease-metax-release-model_qwen3-4b-tree_none-gems_2.2-scale_0.6.0-cx_none-python_3.10.10-torch_2.4.0_metax2.29.2.6-pcp_maca2.29.2.7-gpu_metax001-arc_amd64-driver_3.3.12:2508011527
77+
docker pull harbor.baai.ac.cn/flagrelease-public/flagrelease-metax-release-model_qwen3-4b-tree_none-gems_2.2-scale_0.6.0-cx_none-python_3.10.10-torch_2.4.0_metax2.29.2.6-pcp_maca2.29.2.7-gpu_metax001-arc_amd64-driver_3.3.12:260225
7878
```
7979

8080
### Start the inference service
8181

8282
```bash
8383
#Container Startup
84-
docker run -it --device=/dev/dri --device=/dev/mxcd --group-add video \
84+
docker run -it --device=/dev/dri --device=/dev/mxcd -e USE_FLAGGEMS=1 --group-add video \
8585
--name flagos --device=/dev/mem --network=host \
8686
--security-opt seccomp=unconfined --security-opt apparmor=unconfined \
8787
--shm-size '100gb' --ulimit memlock=-1 \
8888
-v /usr/local/:/usr/local/ -v /nfs:/nfs \
89-
harbor.baai.ac.cn/flagrelease-public/flagrelease-metax-release-model_qwen3-4b-tree_none-gems_2.2-scale_0.6.0-cx_none-python_3.10.10-torch_2.4.0_metax2.29.2.6-pcp_maca2.29.2.7-gpu_metax001-arc_amd64-driver_3.3.12:2508011527 /bin/bash
89+
harbor.baai.ac.cn/flagrelease-public/flagrelease-metax-release-model_qwen3-4b-tree_none-gems_2.2-scale_0.6.0-cx_none-python_3.10.10-torch_2.4.0_metax2.29.2.6-pcp_maca2.29.2.7-gpu_metax001-arc_amd64-driver_3.3.12:260225 /bin/bash
9090
```
9191

9292
### Serve

0 commit comments

Comments
 (0)