Skip to content

Latest commit

 

History

History
67 lines (51 loc) · 4.96 KB

File metadata and controls

67 lines (51 loc) · 4.96 KB

Training Results

This document summarizes the training results of our models enhanced with Verl-Tool across various agentic RL tasks.

Model Checkpoints

All these models are available in our Huggingface Collection.

Math TIR Models

Model Link Wandb
Qwen-2.5-Math-1.5B-Verl-tool 🤗 📈
Qwen-2.5-Math-7B-Verl-tool 🤗 📈

Search-R1 Models

Model Link Wandb
VT-Search-zero-3B (GRPO) 🤗 📈
VT-Search-zero-3B (DAPO) 🤗 📈
VT-Search-zero-7B (GRPO) 🤗 📈
VT-Search-zero-7B (DAPO) 🤗 📈

NL2SQL Models

Model Link Wandb
VT-SQL-7B 🤗 📈

SWE Models

Model Link Wandb
VT-SWE-8B 🤗 📈

Visual Reasoner Models

Model Link Wandb
VT-VisualReasoner-7B 🤗 📈

DeepSearch Models

Model Link Wandb
VT-DeepSearch-8B 🤗 📈

Math Benchmark Results

1.5B Model Performance across challenging mathematical benchmarks:

Model Name Tool GSM8K MATH 500 Minerva Math Olympiad Bench AIME24 AMC23 Avg
Qwen2.5-Math-1.5B 39.50 34.80 8.10 23.00 13.30 35.00 25.62
Qwen2.5-Math-1.5B-Instruct 84.90 74.20 26.80 39.00 10.00 57.50 48.70
Qwen2.5-Math-1.5B-Instruct + SimpleRL-Zoo 81.90 70.20 20.60 33.90 20.00 55.00 46.90
Qwen-2.5-Math-1.5B-Instruct-TIR 83.70 76.20 24.30 41.30 26.70 55.00 51.20
ToRL-1.5B 85.60 77.80 29.80 44.00 26.70 67.50 55.23
Qwen-2.5-Math-1.5B + Verl-Tool 85.10 77.40 28.30 44.00 33.30 65.00 55.52

7B Model Performance across challenging mathematical benchmarks:

Model Name Tool GSM8K MATH 500 Minerva Math Olympiad Bench AIME24 AMC23 Avg
Qwen-2.5-Math-7B 65.50 63.60 12.50 25.80 13.30 42.50 37.20
Qwen2.5-Math-7B-Instruct 95.20 83.00 37.10 41.60 16.70 70.00 57.27
Qwen-2.5-Math-7B + SimpleRL-Zoo 88.80 80.20 26.80 41.60 30.00 52.50 53.30
Qwen-2.5-Math-7B-Instruct-TIR 94.60 82.40 29.00 50.50 30.00 62.50 58.17
TORL-7B 92.70 82.20 33.50 49.90 43.30 65.00 61.10
Qwen-2.5-Math-7B + Verl-Tool 91.40 83.40 29.80 50.20 40.00 72.50 61.22