-
Notifications
You must be signed in to change notification settings - Fork 15
v2.0.0 모델별 benchmark 결과
gannim edited this page Jul 30, 2025
·
1 revision
| model | judge type | singlecall | dialog | calldecision | overall score |
|---|---|---|---|---|---|
| claude-opus-4 | LLM judge | 0.9320 | 0.9800 | 0.9950 | 0.9690 |
| Human judge | 0.9160 | 0.9800 | 0.9950 | 0.9637 | |
| gemini-2.5-pro-preview-05-06 | LLM judge | 0.9300 | 0.9550 | 0.9769 | 0.9540 |
| Human judge | 0.9300 | 0.9450 | 0.9769 | 0.9506 | |
| Qwen3-32B(think) | LLM judge | 0.9280 | 0.9550 | 0.9670 | 0.9500 |
| Human judge | 0.9220 | 0.9550 | 0.9653 | 0.9474 | |
| gpt-4.1-2025-04-14 | LLM judge | 0.8980 | 0.9300 | 0.9092 | 0.9124 |
| Human judge | 0.9040 | 0.9200 | 0.8927 | 0.9056 | |
| gpt-4o-mini-2024-07-18 | LLM judge | 0.7820 | 0.8950 | 0.8564 | 0.8445 |
| Human judge | 0.7840 | 0.8950 | 0.8465 | 0.8418 |
| LLM judge rank | model | judge type | exact (100) | 4_random (100) | 4_close (100) | 8_random (100) | 8_close (100) | SUM (500) | AVG (micro) |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Qwen3-32B(think) | LLM judge | 0.9300 | 0.9300 | 0.9400 | 0.9400 | 0.9200 | 466 | 0.9320 |
| Human judge | 0.9100 | 0.9000 | 0.9300 | 0.9200 | 0.9200 | 458 | 0.9160 | ||
| 2 | claude-opus-4 | LLM judge | 0.9300 | 0.9500 | 0.9500 | 0.9300 | 0.8900 | 465 | 0.9300 |
| Human judge | 0.9400 | 0.9400 | 0.9400 | 0.9300 | 0.9000 | 465 | 0.9300 | ||
| 3 | gemini-2.5-pro-preview-05-06 | LLM judge | 0.9300 | 0.9200 | 0.9200 | 0.9400 | 0.9300 | 464 | 0.9280 |
| Human judge | 0.9400 | 0.9000 | 0.9100 | 0.9300 | 0.9300 | 461 | 0.9220 | ||
| 4 | gpt-4.1-2025-04-14 | LLM judge | 0.9000 | 0.9100 | 0.9000 | 0.8900 | 0.8900 | 449 | 0.8980 |
| Human judge | 0.9200 | 0.9000 | 0.9100 | 0.8900 | 0.9000 | 452 | 0.9040 | ||
| 5 | gpt-4o-mini-2024-07-18 | LLM judge | 0.8500 | 0.7700 | 0.7700 | 0.7700 | 0.7500 | 391 | 0.7820 |
| Human judge | 0.8500 | 0.7700 | 0.7700 | 0.7700 | 0.7600 | 392 | 0.7840 |
| LLM judge rank | model | judge type | CALL (70) | COMPLETION (71) | SLOT (36) | RELEVANCE (23) | SUM (200) | AVG (micro) |
|---|---|---|---|---|---|---|---|---|
| 1 | claude-opus-4 | LLM judge | 67 (0.9571) | 71 (1.0000) | 35 (0.9722) | 23 (1.0000) | 196 | 0.9800 |
| Human judge | 67 (0.9571) | 71 (1.0000) | 35 (0.9722) | 23 (1.0000) | 196 | 0.9800 | ||
| 2 | gpt-4.1-2025-04-14 | LLM judge | 66 (0.9429) | 68 (0.9577) | 36 (1.0000) | 21 (0.9130) | 191 | 0.9550 |
| Human judge | 66 (0.9429) | 67 (0.9437) | 36 (1.0000) | 20 (0.8696) | 189 | 0.9450 | ||
| 3 | gemini-2.5-pro-preview-05-06 | LLM judge | 65 (0.9286) | 67 (0.9437) | 36 (1.0000) | 23 (1.0000) | 191 | 0.9550 |
| Human judge | 65 (0.9286) | 68 (0.9577) | 35 (0.9722) | 23 (1.0000) | 191 | 0.9550 | ||
| 4 | Qwen3-32B(think) | LLM judge | 68 (0.9714) | 68 (0.9577) | 32 (0.8889) | 18 (0.7826) | 186 | 0.9300 |
| Human judge | 68 (0.9714) | 66 (0.9296) | 33 (0.9167) | 17 (0.7391) | 184 | 0.9200 | ||
| 5 | gpt-4o-mini-2024-07-18 | LLM judge | 63 (0.9000) | 70 (0.9859) | 34 (0.9444) | 12 (0.5217) | 179 | 0.8950 |
| Human judge | 63 (0.9000) | 70 (0.9859) | 33 (0.9167) | 13 (0.5652) | 179 | 0.8950 |
| LLM judge rank | model | judge type | CALL (100) | REJECT (100) | SLOT-all (100) | SLOT-some (306) | SUM (606) | AVG (micro) |
|---|---|---|---|---|---|---|---|---|
| 1 | claude-opus-4 | LLM judge | 100 (1.0000) | 100 (1.0000) | 100 (1.0000) | 303 (0.9902) | 603 | 0.9950 |
| Human judge | 100 (1.0000) | 100 (1.0000) | 100 (1.0000) | 303 (0.9902) | 603 | 0.9950 | ||
| 2 | gemini-2.5-pro-preview-05-06 | LLM judge | 96 (0.9600) | 98 (0.9800) | 96 (0.9600) | 302 (0.9876) | 592 | 0.9769 |
| Human judge | 97 (0.9700) | 95 (0.9500) | 98 (0.9800) | 302 (0.9876) | 592 | 0.9769 | ||
| 3 | Qwen3-32B(think) | LLM judge | 94 (0.9400) | 96 (0.9600) | 98 (0.9800) | 298 (0.9745) | 586 | 0.9670 |
| Human judge | 99 (0.9900) | 93 (0.9300) | 97 (0.9700) | 296 (0.9673) | 585 | 0.9653 | ||
| 4 | gpt-4.1-2025-04-14 | LLM judge | 93 (0.9300) | 65 (0.6500) | 100 (1.0000) | 293 (0.9575) | 551 | 0.9092 |
| Human judge | 94 (0.9400) | 58 (0.5800) | 100 (1.0000) | 289 (0.9444) | 541 | 0.8927 | ||
| 5 | gpt-4o-mini-2024-07-18 | LLM judge | 87 (0.8700) | 46 (0.4600) | 100 (1.0000) | 286 (0.9346) | 519 | 0.8564 |
| Human judge | 87 (0.8700) | 41 (0.4100) | 100 (1.0000) | 285 (0.9314) | 513 | 0.8465 |