| Benchmark Category | Benchmark | Temperature | Recommended max tokens | Recommended runs | Top-p | Others (e.g. test log) |
|---|---|---|---|---|---|---|
| Multi-modal | MMMU-Pro | 1.0 | max tokens = 96k | 3 | top\_p=0.95 | thinking= |
| MMMU-Pro w/ python | 1.0 | per step tokens = 64k; total max tokens = 256k |
3 | top\_p=0.95 | Recommended max steps = 50 thinking= |
|
| CharXiv (RQ) | 1.0 | max tokens = 96k | 3 | top\_p=0.95 | thinking= | |
| CharXiv (RQ) w/ python | 1.0 | per step tokens = 64k; total max tokens = 256k |
3 | top\_p=0.95 | Recommended max steps = 50 thinking= |
|
| MathVision | 1.0 | max tokens = 96k | 3 | top\_p=0.95 | thinking= | |
| MathVision w/ python | 1.0 | per step tokens = 64k; total max tokens = 256k |
3 | top\_p=0.95 | Recommended max steps = 50 thinking= |
|
| V\* w/ python | 1.0 | per step tokens = 64k; total max tokens = 256k |
3 | top\_p=0.95 | Recommended max steps = 50 thinking= |
|
| Agent | HLE-Full w/ tools | 1.0 | per step tokens = 48k; total max tokens = 256k |
1 | top\_p=0.95 | Recommended max steps = 300 thinking= |
| BrowseComp | 1.0 | per step tokens = 48k; total max tokens = 256k |
1 | top\_p=0.95 | Recommended max steps = 300 thinking= |
|
| DeepSearchQA | 1.0 | per step tokens = 48k; total max tokens = 256k |
1 | top\_p=0.95 | Recommended max steps = 300 thinking= |
|
| WideSearch | 1.0 | per step tokens = 48k; total max tokens = 256k |
4 | top\_p=0.95 | Recommended max steps = 300 thinking= |
|
| Toolathlon | 1.0 | per step tokens = 48k; total max tokens = 256k |
4 | top\_p=0.95 | Recommended max steps = 300 thinking= |
|
| MCPMark | 1.0 | per step tokens = 48k; total max tokens = 256k |
4 | top\_p=0.95 | Recommended max steps = 300 thinking= |
|
| Claw Eval | 1.0 | per step tokens = 48k; total max tokens = 256k |
4 | top\_p=0.95 | Recommended max steps = 300 thinking= |
|
| APEX-Agents | 1.0 | per step tokens = 48k; total max tokens = 256k |
4 | top\_p=0.95 | Recommended max steps = 300 thinking= |
|
| Coding | Terminal-Bench 2.0 (Terminus-2) | 1.0 | max tokens = 256k | 3 | top\_p=0.95 | thinking= |
| SWE-Bench Pro | 1.0 | per step tokens = 32k; total max tokens = 256k |
5 | top\_p=0.95 | Recommended max steps = 300 thinking= |
|
| SWE-Bench Multilingual | 1.0 | per step tokens = 32k; total max tokens = 256k |
5 | top\_p=0.95 | Recommended max steps = 300 thinking= |
|
| SWE-Bench Verified | 1.0 | per step tokens = 32k; total max tokens = 256k |
5 | top\_p=0.95 | Recommended max steps = 300 thinking= |
|
| SciCode | 1.0 | max tokens = 96k | 4 | top\_p=0.95 | thinking= | |
| OJBench (python) | 1.0 | max tokens = 96k | 8 | top\_p=0.95 | thinking= | |
| LiveCodeBench (v6) | 1.0 | max tokens = 96k | 1 | top\_p=0.95 | thinking= | |
| Math | AIME 2026 | 1.0 | max tokens = 96k | 32 | top\_p=0.95 | thinking= |
| HMMT 2026 (Feb) | 1.0 | max tokens = 96k | 32 | top\_p=0.95 | thinking= | |
| IMO-AnswerBench | 1.0 | max tokens = 96k | 4 | top\_p=0.95 | thinking= | |
| Knowledge | HLE-Full | 1.0 | max tokens = 96k | 1 | top\_p=0.95 | thinking= |
| GPQA-Diamond | 1.0 | max tokens = 96k | 8 | top\_p=0.95 | thinking= |
| Benchmark Category | Benchmark | Temperature | Recommended max tokens | Recommended runs | Top-p | Others (e.g. test log) |
|---|---|---|---|---|---|---|
| Multi-modal | MMMU-Pro | 1.0 | max tokens = 64k | 3 | top\_p=0.95 | thinking= |
| CharXiv (RQ) | 1.0 | max tokens = 64k | 3 | top\_p=0.95 | thinking= | |
| MathVision | 1.0 | max tokens = 64k | 3 | top\_p=0.95 | thinking= | |
| MathVista | 1.0 | max tokens = 64k | 3 | top\_p=0.95 | thinking= | |
| OCRBench | 1.0 | max tokens = 64k | 3 | top\_p=0.95 | thinking= | |
| ZeroBench | 1.0 | max tokens = 64k | 3 | top\_p=0.95 | thinking= | |
| WorldVQA | 1.0 | max tokens = 64k | 3 | top\_p=0.95 | thinking= | |
| InfoVQA (val) | 1.0 | max tokens = 64k | 3 | top\_p=0.95 | thinking= | |
| SimpleVQA | 1.0 | max tokens = 64k | 3 | top\_p=0.95 | thinking= | |
| ZeroBench w/ tools | 1.0 | max tokens = 64k | 3 | top\_p=0.95 | Recommended max steps = 30 thinking= |
|
| Code | SWE Series | 1.0 | per step tokens = 16k; total max tokens = 256k |
5 | top\_p=0.95 | thinking= |
| Lcb + OJBench | 1.0 | max tokens = 128k | 1 | top\_p=0.95 | thinking= | |
| TerminalBench | 1.0 | max tokens = 128k | 3 | top\_p=0.95 | thinking= | |
| Reasoning | AIME2025 no tools | 1.0 | total max tokens = 96k | 32 | top\_p=0.95 | thinking= |
| AIME2025 w/ tools | 1.0 | per turn tokens = 96k; total max tokens = 96k |
32 | top\_p=0.95 |
thinking=
Recommended max steps = 120 |
|
| HLE no tools | 1.0 | max tokens = 96k | 1 | top\_p=0.95 | thinking= | |
| HLE w/ tools | 1.0 | total max tokens = 128k; per step tokens = 48k |
1 | top\_p=0.95 |
thinking=
Recommended max steps = 120 |
|
| HLE heavy | 1.0 | total max tokens = 128k; per step tokens = 48k |
1 | top\_p=0.95 |
thinking=
Recommended max steps = 200 parallel n=8 |
|
| HMMT2025 no tools | 1.0 | max tokens = 96k | 32 | top\_p=0.95 | thinking= | |
| HMMT2025 w/tools | 1.0 | per step tokens = 96k; total tokens = 96k |
32 | top\_p=0.95 |
thinking=
Recommended max steps = 120 |
|
| IMO-AnswerBench | 1.0 | max tokens = 96k | 3 | top\_p=0.95 | thinking= | |
| GPQA-Diamond | 1.0 | max tokens = 96k | 8 | top\_p=0.95 | thinking= | |
| Agentic Search Task | BrowseComp / BrowseComp-ZH / Seal-0 / Frames | 1.0 | per step tokens = 24k; total max tokens = 256k |
4 | top\_p=0.95 |
thinking=
Recommended max steps = 250 Recommend using a context management mechanism to prevent overly long context and ensure enough tool calls Include today's date in the system prompt and let the model search when it is uncertain |
| Agentic Task | Tau | 1.0 | >=16k | 4 | top\_p=0.95 |
thinking=
Recommended max steps = 100 |
| Category | Benchmark | Temperature | Max token | Suggested runs | Notes |
|---|---|---|---|---|---|
| Code | SWE | 0.7(recommended) 1.0 (ok) |
per step tokens = 16k; total max token = 256k |
5 | |
| Lcb + OJBench | 1.0 | max tokens = 128k | 1 | ||
| TerminalBench | 1.0 | max tokens = 128k | 3 | ||
| Reasoning | AIME2025 no tools | 1.0 | total max tokens = 96k | 32 | |
| AIME2025 w/ tools | 1.0 | per step tokens = 48k; total max tokens = 128k |
16 | max steps = 120 | |
| HLE no tools | 1.0 | max tokens = 96k | 1 | ||
| HLE w/ tools | 1.0 | total max tokens = 128k; per step tokens = 48k |
1 | max steps = 120 | |
| HLE heavy | 1.0 | total max tokens = 128k; per step tokens = 48k |
1 | max steps = 200 parallel n=8 |
|
| HMMT2025 no tools | 1.0 | max tokens = 96k | 32 | ||
| HMMT2025 w/tools | 1.0 | per step tokens = 96k; total tokens = 96k |
32 | max steps = 120 | |
| IMO-AnswerBench | 1.0 | max tokens = 96k | 3 | ||
| GPQA-Diamond | 1.0 | max tokens = 96k | 8 | ||
| Agentic Search Task | BrowseComp/ BrowseComp-ZH/Seal-0/ Frames | 1.0 | per step tokens = 24k; total max tokens = 256k |
4 | max steps = 250 Enable context management to prevent context overflow and ensure enough tool calls. Include today's date in the system prompt, and tell the model to search when unsure. |
| Agentic Task | Tau | 0.0 | >=16k | 4 | max steps = 100 |