Playable
explicit result recorded

Measured decode throughput for every serve behind the arcade — aggregate tok/s, per-agent tok/s, and how both scale with concurrent agents on the NVIDIA GB10 and the Intel Arc Pro B70. Every number links back to its source CSV or run report.
Only an explicit playable result earns a green sticker. Static checks alone do not count.
explicit result recorded
recorded problem
no playability result
| Model | Road Hopper | Robot-Filled Maze Shooter | Ribbit Rush |
|---|---|---|---|
| AEON QWEN | 1 not checked 1 recorded artifact | 1 not checked 1 recorded artifact | — |
| Agents-A1 | 1 playable 1 recorded artifact | 1 playable 1 recorded artifact | — |
| AgentWorld-35B | 6 playable1 needs fixing2 not checked 9 recorded artifacts | 6 playable2 not checked 8 recorded artifacts | — |
| Claude Opus 4.8 | 2 not checked 2 recorded artifacts | 2 not checked 2 recorded artifacts | — |
| Claude Sonnet 5 | — | — | 1 not checked 1 recorded artifact |
| DeepSeek V4 Pro | 1 not checked 1 recorded artifact | 1 not checked 1 recorded artifact | — |
| Fable | 1 not checked 1 recorded artifact | 1 not checked 1 recorded artifact | — |
| Frosty Ornith-35B | 1 needs fixing2 not checked 3 recorded artifacts | 3 needs fixing1 not checked 4 recorded artifacts | — |
| Gemma-4 current | 1 not checked 1 recorded artifact | 1 not checked 1 recorded artifact | — |
| Gemma-4-26B | 1 not checked 1 recorded artifact | 1 not checked 1 recorded artifact | — |
| GLM-5.2 | 1 not checked 1 recorded artifact | 1 not checked 1 recorded artifact | — |
| GPT-5.5 xhigh | 1 not checked 1 recorded artifact | 1 not checked 1 recorded artifact | — |
| GPT-5.6 Sol | 1 not checked 1 recorded artifact | 1 not checked 1 recorded artifact | — |
| GPT-5.6 Terra | 1 not checked 1 recorded artifact | 1 not checked 1 recorded artifact | — |
| Laguna XS 2.1 | 1 not checked 1 recorded artifact | 2 not checked 2 recorded artifacts | — |
| Leanstral NVFP4 | 1 needs fixing 1 recorded artifact | 1 needs fixing 1 recorded artifact | — |
| LongCat-2.0 | 6 not checked 6 recorded artifacts | 2 not checked 2 recorded artifacts | — |
| Nemotron 3 Super | 1 playable 1 recorded artifact | 1 playable 1 recorded artifact | — |
| Nemotron Puzzle 75B | 1 playable 1 recorded artifact | 1 playable 1 recorded artifact | — |
| NEX2 | 1 playable1 not checked 2 recorded artifacts | 1 needs fixing2 not checked 3 recorded artifacts | — |
| Ornith-AEON | 1 not checked 1 recorded artifact | 1 not checked 1 recorded artifact | — |
| Quest Turbo | — | 1 needs fixing 1 recorded artifact | — |
| Qwen3 Coder 30B | 1 playable 1 recorded artifact | 1 playable 1 recorded artifact | — |
| Qwen3.6-35B Base | 1 not checked 1 recorded artifact | — | — |
| Qwen3.6-35B NVFP4 | 1 playable 1 recorded artifact | 1 playable 1 recorded artifact | — |
| SIQ-1-35B | 2 playable 2 recorded artifacts | 1 playable1 needs fixing 2 recorded artifacts | — |
| StepFun Step-3.7 | 1 playable 1 recorded artifact | 1 playable 1 recorded artifact | — |
| Status | Model / condition | Game | Variant | Static checks | Review score | Build wall | Evidence |
|---|---|---|---|---|---|---|---|
| Needs fixing | AgentWorld-35BIntel Arc Pro B70 | Road Hopper | B70 live teamwork eval | 6/6 | 31 | 1,262.7 s | ArtifactRun notes |
| Needs fixing | Frosty Ornith-35BIntel Arc Pro B70 | Road Hopper | B70 live teamwork eval | 6/6 | 33 | 1,694.6 s | Artifact |
| Needs fixing | Frosty Ornith-35BIntel Arc Pro B70 | Robot-Filled Maze Shooter | B70 failed attempt | — | — | — | Artifact |
| Needs fixing | Frosty Ornith-35BIntel Arc Pro B70 | Robot-Filled Maze Shooter | B70 live teamwork eval | 7/7 | 34 | 1,078.3 s | Artifact |
| Needs fixing | Frosty Ornith-35BIntel Arc Pro B70 | Robot-Filled Maze Shooter | B70 stock baseline | 6/6 | — | 989.8 s | ArtifactRun notes |
| Needs fixing | Leanstral NVFP4DGX Spark / NVIDIA GB10 | Road Hopper | Leanstral NVFP4 Road Hopper eval | 8/11 | — | 228.5 s | ArtifactRun notes |
| Needs fixing | Leanstral NVFP4DGX Spark / NVIDIA GB10 | Robot-Filled Maze Shooter | Leanstral NVFP4 Robot Maze eval | 8/11 | — | 75.8 s | ArtifactRun notes |
| Needs fixing | NEX2Intel Arc Pro B70 | Robot-Filled Maze Shooter | B70 live teamwork eval | 7/7 | 31 | 2,670.2 s | Artifact |
| Needs fixing | Quest TurboIntel Arc Pro B70 | Robot-Filled Maze Shooter | B70 live teamwork eval | 7/7 | 29 | 3,115.7 s | Artifact |
| Needs fixing | SIQ-1-35BIntel Arc Pro B70 | Robot-Filled Maze Shooter | B70 live teamwork eval | 7/7 | 30 | 1,064.7 s | Artifact |
| Playable | Agents-A1DGX Spark / NVIDIA GB10 | Road Hopper | Agents-A1 BF16 Road Hopper eval | 6/6 | — | 1,436.5 s | Artifact |
| Playable | Agents-A1DGX Spark / NVIDIA GB10 | Robot-Filled Maze Shooter | Agents-A1 BF16 Maze eval | 7/7 | — | 1,678.5 s | Artifact |
| Playable | AgentWorld-35BIntel Arc Pro B70 | Road Hopper | B70 Q5_K_M quant eval | 6/6 | 34 | 908.4 s | ArtifactRun notes |
| Playable | AgentWorld-35BIntel Arc Pro B70 | Road Hopper | B70 Q6_K quant eval run 1 | 6/6 | 34 | 971.3 s | ArtifactRun notes |
| Playable | AgentWorld-35BIntel Arc Pro B70 | Road Hopper | B70 Q6_K quant eval run 2 | 6/6 | 35 | 1,598.6 s | ArtifactRun notes |
| Playable | AgentWorld-35BIntel Arc Pro B70 | Road Hopper | CrewAI N=12 | 7/7 | — | 396 s | Artifact |
| Playable | AgentWorld-35BIntel Arc Pro B70 | Road Hopper | CrewAI N=6 | 7/7 | — | 232 s | Artifact |
| Playable | AgentWorld-35BIntel Arc Pro B70 | Road Hopper | Solo | 7/7 | — | 87 s | Artifact |
| Playable | AgentWorld-35BIntel Arc Pro B70 | Robot-Filled Maze Shooter | B70 live teamwork eval | 7/7 | 40 | 1,038.1 s | ArtifactRun notes |
| Playable | AgentWorld-35BIntel Arc Pro B70 | Robot-Filled Maze Shooter | B70 Q5_K_M quant eval | 7/7 | 40 | 1,105.0 s | ArtifactRun notes |
| Playable | AgentWorld-35BIntel Arc Pro B70 | Robot-Filled Maze Shooter | B70 Q6_K quant eval run 1 | 7/7 | 40 | 1,816.9 s | ArtifactRun notes |
| Playable | AgentWorld-35BIntel Arc Pro B70 | Robot-Filled Maze Shooter | B70 Q6_K quant eval run 2 | 7/7 | 40 | 1,292.4 s | ArtifactRun notes |
| Playable | AgentWorld-35BIntel Arc Pro B70 | Robot-Filled Maze Shooter | CrewAI N=12 | 7/7 | — | 416 s | Artifact |
| Playable | AgentWorld-35BIntel Arc Pro B70 | Robot-Filled Maze Shooter | Solo | 6/7 | — | 69 s | Artifact |
| Playable | Nemotron 3 SuperDGX Spark / NVIDIA GB10 | Road Hopper | Nemotron 3 Super Spark eval | 6/6 | — | 1,200.7 s | Artifact |
| Playable | Nemotron 3 SuperDGX Spark / NVIDIA GB10 | Robot-Filled Maze Shooter | Nemotron 3 Super Spark eval | 7/7 | — | 1,202.8 s | Artifact |
| Playable | Nemotron Puzzle 75BDGX Spark / NVIDIA GB10 | Road Hopper | Nemotron Puzzle 75B Spark eval | 6/6 | — | 1,200.9 s | Artifact |
| Playable | Nemotron Puzzle 75BDGX Spark / NVIDIA GB10 | Robot-Filled Maze Shooter | Nemotron Puzzle 75B Spark eval | 7/7 | — | 1,201.4 s | Artifact |
| Playable | NEX2Intel Arc Pro B70 | Road Hopper | B70 live teamwork eval | 6/6 | 37 | 1,316.8 s | Artifact |
| Playable | Qwen3 Coder 30BDGX Spark / NVIDIA GB10 | Road Hopper | Qwen3 Coder BF16 long-context teamwork eval | 6/6 | — | 1,326.8 s | Artifact |
| Playable | Qwen3 Coder 30BDGX Spark / NVIDIA GB10 | Robot-Filled Maze Shooter | Qwen3 Coder BF16 long-context teamwork eval | 7/7 | — | 2,235.6 s | Artifact |
| Playable | Qwen3.6-35B NVFP4DGX Spark / NVIDIA GB10 | Road Hopper | Qwen3.6-35B NVFP4 Road Hopper eval | 6/6 | — | 1,265 s | Artifact |
| Playable | Qwen3.6-35B NVFP4DGX Spark / NVIDIA GB10 | Robot-Filled Maze Shooter | Qwen3.6-35B NVFP4 Maze eval | 7/7 | — | 1,511 s | Artifact |
| Playable | SIQ-1-35BIntel Arc Pro B70 | Road Hopper | B70 live teamwork eval | 6/6 | 42 | 2,480.9 s | Artifact |
| Playable | SIQ-1-35BDGX Spark / NVIDIA GB10 | Road Hopper | GB10 teamwork eval | 6/6 | — | 1,367.1 s | Artifact |
| Playable | SIQ-1-35BDGX Spark / NVIDIA GB10 | Robot-Filled Maze Shooter | GB10 teamwork eval | 7/7 | — | 1,355.3 s | Artifact |
| Playable | StepFun Step-3.7DGX Spark / NVIDIA GB10 | Road Hopper | Step37 teamwork eval | 6/6 | — | 3,284.3 s | Artifact |
| Playable | StepFun Step-3.7DGX Spark / NVIDIA GB10 | Robot-Filled Maze Shooter | Step37 teamwork eval | 7/7 | — | 3,380.8 s | Artifact |
| Not checked | AEON QWENDGX Spark / NVIDIA GB10 | Road Hopper | teamwork eval | 6/6 | — | 1,753.4 s | Artifact |
| Not checked | AEON QWENDGX Spark / NVIDIA GB10 | Robot-Filled Maze Shooter | teamwork eval | 7/7 | — | 1,256.1 s | Artifact |
| Not checked | AgentWorld-35BIntel Arc Pro B70 | Road Hopper | B70 stock baseline | 6/6 | — | 1,834.2 s | ArtifactRun notes |
| Not checked | AgentWorld-35BDGX Spark / NVIDIA GB10 | Road Hopper | teamwork eval | 6/6 | — | 1,607.5 s | Artifact |
| Not checked | AgentWorld-35BIntel Arc Pro B70 | Robot-Filled Maze Shooter | B70 stock baseline | 5/6 | — | 2,044.4 s | ArtifactRun notes |
| Not checked | AgentWorld-35BDGX Spark / NVIDIA GB10 | Robot-Filled Maze Shooter | teamwork eval | 7/7 | — | 1,502.7 s | Artifact |
| Not checked | Claude Opus 4.8Anthropic API | Road Hopper | FROGGAI B70 control | teamwork/validate_game.js ok | 47/50 | — | ArtifactRun notes |
| Not checked | Claude Opus 4.8Anthropic API | Road Hopper | FroggAI control (gold-standard baseline) | 7/7 | — | 300.0 s | Artifact |
| Not checked | Claude Opus 4.8Anthropic API | Robot-Filled Maze Shooter | MAZERUNNER B70 control | teamwork/validate_game.js ok | 47/50 | — | ArtifactRun notes |
| Not checked | Claude Opus 4.8Anthropic API | Robot-Filled Maze Shooter | MAZERUNNER control (gold-standard baseline) | 7/7 | — | 300.0 s | Artifact |
| Not checked | Claude Sonnet 5Anthropic API | Ribbit Rush | Ribbit Rush standalone game | — | — | — | Artifact |
| Not checked | DeepSeek V4 ProDeepSeek API | Road Hopper | FroggAI deepseek eval | 7/7 | — | 120.0 s | Artifact |
| Not checked | DeepSeek V4 ProDeepSeek API | Robot-Filled Maze Shooter | Maze deepseek eval | 7/7 | — | 150.0 s | Artifact |
| Not checked | Fableexternal submission | Road Hopper | Fable Road Hopper | — | — | — | Artifact |
| Not checked | Fableexternal submission | Robot-Filled Maze Shooter | Fable Maze Shooter | — | — | — | Artifact |
| Not checked | Frosty Ornith-35BIntel Arc Pro B70 | Road Hopper | B70 stock baseline | 6/6 | — | 1,081.7 s | ArtifactRun notes |
| Not checked | Frosty Ornith-35BDGX Spark / NVIDIA GB10 | Road Hopper | teamwork eval | 6/6 | — | 1,784.4 s | Artifact |
| Not checked | Frosty Ornith-35BDGX Spark / NVIDIA GB10 | Robot-Filled Maze Shooter | teamwork eval | 7/7 | — | 1,486.2 s | Artifact |
| Not checked | Gemma-4 currentDGX Spark / NVIDIA GB10 | Road Hopper | current plain vLLM eval | 6/6 | — | 1,604.1 s | Artifact |
| Not checked | Gemma-4 currentDGX Spark / NVIDIA GB10 | Robot-Filled Maze Shooter | current plain vLLM eval | 6/7 | — | 1,663.0 s | Artifact |
| Not checked | Gemma-4-26BDGX Spark / NVIDIA GB10 | Road Hopper | teamwork eval | 6/6 | — | 1,506.2 s | Artifact |
| Not checked | Gemma-4-26BDGX Spark / NVIDIA GB10 | Robot-Filled Maze Shooter | teamwork eval | 7/7 | — | 1,316.6 s | Artifact |
| Not checked | GLM-5.2API | Road Hopper | GLM-5.2 OpenCode eval | — | — | — | Artifact |
| Not checked | GLM-5.2API | Robot-Filled Maze Shooter | GLM-5.2 OpenCode eval | — | — | — | Artifact |
| Not checked | GPT-5.5 xhighOpenAI API | Road Hopper | FroggAI xhigh eval | 7/7 | — | 210.0 s | Artifact |
| Not checked | GPT-5.5 xhighOpenAI API | Robot-Filled Maze Shooter | Maze xhigh eval | 7/7 | — | 246.0 s | Artifact |
| Not checked | GPT-5.6 SolOpenAI API via Codex Desktop | Road Hopper | GPT-5.6 Sol road-and-river one-shot | — | — | — | Artifact |
| Not checked | GPT-5.6 SolOpenAI API via Codex Desktop | Robot-Filled Maze Shooter | GPT-5.6 Sol maze-shooter one-shot | — | — | — | Artifact |
| Not checked | GPT-5.6 TerraOpenAI API via Codex Desktop | Road Hopper | GPT-5.6 Terra road-and-river one-shot | — | — | — | Artifact |
| Not checked | GPT-5.6 TerraOpenAI API via Codex Desktop | Robot-Filled Maze Shooter | GPT-5.6 Terra maze-shooter one-shot | — | — | — | Artifact |
| Not checked | Laguna XS 2.1Intel Arc Pro B70 | Road Hopper | Laguna XS 2.1 Q4_K_M B70 live teamwork eval | 6/6 | — | 1,362.1 s | ArtifactRun notes |
| Not checked | Laguna XS 2.1Intel Arc Pro B70 | Robot-Filled Maze Shooter | Laguna XS 2.1 Q4_K_M B70 live teamwork eval | 6/7 | — | 1,805.5 s | ArtifactRun notes |
| Not checked | Laguna XS 2.1DGX Spark / NVIDIA GB10 | Robot-Filled Maze Shooter | v2 | — | — | — | Artifact |
| Not checked | LongCat-2.0Meituan hosted API | Road Hopper | Classic pixel-art Frogger style | teamwork/validate_game.js ok | — | — | Artifact |
| Not checked | LongCat-2.0Meituan hosted API | Road Hopper | LongCat chain3 Road Hopper | not run: jsdom missing | — | — | Artifact |
| Not checked | LongCat-2.0Meituan hosted API | Road Hopper | LongCat one-shot Road Hopper | not run: jsdom missing | — | — | Artifact |
| Not checked | LongCat-2.0Meituan hosted API | Road Hopper | Noir rain skyline lantern-glow | teamwork/validate_game.js ok | — | — | Artifact |
| Not checked | LongCat-2.0Meituan hosted API | Road Hopper | Synthwave neon | teamwork/validate_game.js ok | — | — | Artifact |
| Not checked | LongCat-2.0Meituan hosted API | Road Hopper | Terminal ASCII green-on-black | teamwork/validate_game.js ok | — | — | Artifact |
| Not checked | LongCat-2.0Meituan hosted API | Robot-Filled Maze Shooter | LongCat chain3 Maze | teamwork/validate_game.js ok | — | — | Artifact |
| Not checked | LongCat-2.0Meituan hosted API | Robot-Filled Maze Shooter | LongCat one-shot Maze | teamwork/validate_game.js ok | — | — | Artifact |
| Not checked | NEX2DGX Spark / NVIDIA GB10 | Road Hopper | teamwork eval | 6/6 | — | 1,863.3 s | Artifact |
| Not checked | NEX2Turbo via Hydra | Robot-Filled Maze Shooter | Hydra Salvage Run | — | — | 3,600.0 s | Artifact |
| Not checked | NEX2DGX Spark / NVIDIA GB10 | Robot-Filled Maze Shooter | teamwork eval | 7/7 | — | 1,209.5 s | Artifact |
| Not checked | Ornith-AEONDGX Spark / NVIDIA GB10 | Road Hopper | AEON Ornith/Qwen eval | 6/6 | — | 1,594.3 s | Artifact |
| Not checked | Ornith-AEONDGX Spark / NVIDIA GB10 | Robot-Filled Maze Shooter | AEON Ornith/Qwen eval | 7/7 | — | 1,449.0 s | Artifact |
| Not checked | Qwen3.6-35B BaseDGX Spark / NVIDIA GB10 | Road Hopper | Qwen3.6 base GGUF Road Hopper scout | — | — | — | Artifact |
Batched generation throughput as the number of concurrent agents grows (llama.cpp, GGUF-NVFP4). Aggregate generation reaches 372.28 t/s at 256 agents; the separate total-throughput column reaches 848.67 t/s. Per-agent speed and single-reply latency degrade as concurrency rises. The highlighted 32-agent row is the recorded fleet selection.
| Agents | Prompt t/s | Agg gen t/s | Per-agent t/s | Total t/s | 256-tok reply (s) |
|---|---|---|---|---|---|
| 1 | 2286.40 | 68.98 | 68.98 | 195.15 | 3.711 |
| 8 | 2372.40 | 206.48 | 25.81 | 527.60 | 9.919 |
| 16 | 2355.98 | 225.74 | 14.11 | 568.31 | 18.145 |
| 32 | 2352.62 | 277.15 | 8.66 | 672.91 | 29.558 |
| 64 | 2351.41 | 326.31 | 5.10 | 766.25 | 50.210 |
| 96 | 2355.60 | 342.59 | 3.57 | 796.18 | 71.736 |
| 128 | 2353.23 | 350.62 | 2.74 | 810.37 | 93.458 |
| 192 | 2357.97 | 361.53 | 1.88 | 830.05 | 135.957 |
| 256 | 2356.25 | 372.28 | 1.45 | 848.67 | 176.038 |


NVFP4 MTP-XS on vLLM with z-lab DFlash speculative decode. DFlash makes single streams fast (~46–48 t/s), so aggregate plateaus early — the knee is 8 (91% of peak at 2× the per-agent speed of 16). The highlighted 8-agent row is its build-challenge fleet.
| agents | prompt_t/s | agg_gen_t/s | per_agent_gen_t/s | ttft_p50_s | ttft_p99_s | wall_s |
|---|---|---|---|---|---|---|
| 1 | 12.5 (cold-start) | 4.4 | 4.4 | 55.41 | 55.41 | 58.4 |
| 2 | 915.5 | 100.7 | 52.0 | 1.53 | 1.53 | 5.1 |
| 4 | 2417.4 | 165.6 | 47.5 | 1.16 | 1.16 | 6.2 |
| 8 | 2057.5 | 218.3 | 29.9 | 2.74 | 2.74 | 9.4 |
| 16 | 1676.6 | 232.6 | 16.4 | 6.78 | 6.78 | 17.6 |
| 24 | 1581.6 | 251.1 | 11.4 | 10.77 | 10.81 | 24.5 |
| 32 | 1628.1 | 252.1 | 8.6 | 12.52 | 14.02 | 32.5 |
conc=1 is cold-start (first request compiles CUDA graphs); true single-stream is the rate at conc 2–4.
Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4 on the local KV-fix vLLM image with z-lab DFlash drafter n=6. This is measured through the OpenAI-compatible API using real usage token counts and batch wall-clock time; the selected build fleet is the capped 16-agent serve. The highlighted 16-agent row is its build-challenge fleet.
| agents | prompt_t/s | agg_gen_t/s | per_agent_gen_t/s | wall_s |
|---|---|---|---|---|
| 1 | 19.9 | 76.2 | 76.2 | 3.36 |
| 2 | 36.4 | 139.2 | 69.6 | 3.68 |
| 4 | 54.2 | 206.9 | 53.0 | 4.95 |
| 8 | 89.0 | 340.0 | 43.9 | 6.02 |
| 16 | 122.7 | 466.1 | 30.5 | 8.79 |
Generic Qwen3.6-35B-A3B UD-Q6_K GGUF on Spark/GB10 via llama.cpp-qwen36, -c 32768 -np 32, q8 KV, FlashAttention, and fixed 256 generated tokens per agent using min_tokens=256 + ignore_eos. This is a GGUF llama.cpp condition, not the vLLM/NVFP4 stack. The highlighted 32-agent row is its build-challenge fleet.
| agents | prompt_t/s | agg_gen_t/s | per_agent_gen_t/s | wall_s |
|---|---|---|---|---|
| 1 | 14.7 | 56.1 | 56.2 | 4.56 |
| 2 | 13.8 | 52.9 | 26.4 | 9.68 |
| 4 | 14.0 | 53.5 | 13.4 | 19.16 |
| 8 | 19.4 | 74.0 | 9.3 | 27.67 |
| 16 | 32.5 | 123.4 | 7.7 | 33.19 |
| 32 | 64.7 | 244.6 | 7.6 | 33.49 |
Qwen3.6-35B-A3B-UD-Q5_K_XL GGUF served by local Turbo/llama.cpp SYCL on Intel Arc Pro B70 with a 12-slot unified KV configuration and 262144 context. This live OpenAI-compatible run peaks at 158.5 decode tok/s at the 12-agent serve ceiling, with 16.3 tok/s per agent. The highlighted 12-agent row is the 12-slot B70 live serve ceiling.
| agents | prompt_t/s | agg_gen_t/s | per_agent_gen_t/s | wall_s |
|---|---|---|---|---|
| 1 | 307.93 | 68.15 | 79.12 | 3.76 |
| 4 | 116.59 | 127.97 | 37.70 | 8.00 |
| 8 | 78.41 | 156.72 | 24.10 | 13.07 |
| 12 | 61.75 | 158.48 | 16.31 | 19.38 |
Qwen/Qwen-AgentWorld-35B-A3B local GGUF_NVFP4.official artifact on Spark/GB10 via llama.cpp-qwen36, -c 32768 -np 32, q8 KV, FlashAttention, and fixed 256 generated tokens per agent using min_tokens=256 + ignore_eos. This is the controlled Spark rerun of AgentWorld, kept separate from the older site AgentWorld CSV. The highlighted 32-agent row is its build-challenge fleet.
| agents | prompt_t/s | agg_gen_t/s | per_agent_gen_t/s | wall_s |
|---|---|---|---|---|
| 1 | 13.4 | 51.3 | 51.3 | 4.99 |
| 2 | 22.1 | 84.3 | 42.2 | 6.07 |
| 4 | 32.5 | 124.2 | 31.1 | 8.24 |
| 8 | 44.8 | 171.1 | 21.4 | 11.97 |
| 16 | 53.1 | 201.7 | 12.6 | 20.31 |
| 32 | 64.9 | 245.4 | 7.7 | 33.38 |
Official Qwen/Qwen3-Coder-30B-A3B-Instruct BF16 safetensors on Spark/GB10 via vLLM 0.23.0, max_model_len=32768, max_num_seqs=32, FlashAttention, and unquantized FlashInfer CUTLASS MoE. No REAP, no DFlash, no drafter, no quantized derivative. Fixed 256 generated tokens per agent using min_tokens=256 + ignore_eos. The highlighted 32-agent row is its build-challenge fleet.
| agents | prompt_t/s | agg_gen_t/s | per_agent_gen_t/s | wall_s |
|---|---|---|---|---|
| 1 | 7.6 | 30.3 | 30.3 | 8.46 |
| 2 | 13.2 | 52.6 | 26.3 | 9.73 |
| 4 | 20.1 | 80.3 | 20.1 | 12.75 |
| 8 | 30.7 | 122.8 | 15.4 | 16.67 |
| 16 | 53.1 | 211.2 | 13.2 | 19.39 |
| 32 | 87.0 | 344.2 | 10.8 | 23.80 |
Official NVIDIA NVFP4 safetensors on Spark/GB10 via vLLM 0.23.0, max_model_len=8192, max_num_seqs=1, kv_cache_memory_bytes=4G, and eager execution. This is a single-agent scout condition, not a wide concurrency Pareto, because this Spark serve was constrained to one active sequence. The highlighted 1-agent row is its build-challenge fleet.
| agents | prompt_t/s | agg_gen_t/s | per_agent_gen_t/s | wall_s |
|---|---|---|---|---|
| 1 | 4.9 | 13.9 | 13.9 | 13.80 |
Official NVIDIA Puzzle-75B-A9B NVFP4 checkpoint on Spark/GB10 via vLLM 0.23.0, max_model_len=8192, MTP speculative decode num_speculative_tokens=1, max_num_seqs=32, kv_cache_memory_bytes=16G, FlashInfer attention/Mamba, ModelOpt mixed NVFP4/FP8, and eager execution. Fixed-token sweep used min_tokens=192 + ignore_eos; the selected build fleet is the measured 32-agent ceiling, while game builds used the separate context-calibrated 1-agent team. The highlighted 32-agent row is its build-challenge fleet.
| agents | prompt_t/s | agg_gen_t/s | per_agent_gen_t/s | wall_s |
|---|---|---|---|---|
| 1 | 8.8 | 24.9 | 24.9 | 7.72 |
| 2 | 14.5 | 41.0 | 20.8 | 9.37 |
| 4 | 25.5 | 72.1 | 18.3 | 10.65 |
| 8 | 45.6 | 128.7 | 16.3 | 11.94 |
| 16 | 72.3 | 202.9 | 13.0 | 15.14 |
| 24 | 93.1 | 260.6 | 11.1 | 17.68 |
| 32 | 114.4 | 319.8 | 10.3 | 19.21 |
GGUF-NVFP4 (qwen35moe) on llama.cpp at -np 32 — same architecture and size as Ornith, and its curve tracks Ornith's, so the knee sits at 32. The highlighted 32-agent row is its build-challenge fleet.
| agents | prompt_t/s | agg_gen_t/s | per_agent_gen_t/s | ttft_p50_s | ttft_p99_s | wall_s |
|---|---|---|---|---|---|---|
| 1 | 1015.6 | 46.8 | 46.8 | 0.68 | 0.68 | 5.5 |
| 2 | 1452.6 | 71.9 | 36.1 | 0.73 | 0.96 | 7.1 |
| 4 | 1398.8 | 99.4 | 25.0 | 1.95 | 2.01 | 10.3 |
| 8 | 1299.2 | 121.5 | 15.3 | 3.40 | 4.34 | 16.9 |
| 16 | 1230.8 | 132.2 | 8.4 | 5.98 | 9.24 | 31.0 |
| 24 | 1085.2 | 120.0 | 5.1 | 8.76 | 15.76 | 51.2 |
| 32 | 1055.7 | 145.4 | 4.7 | 11.87 | 21.62 | 56.3 |
NVFP4 weights with z-lab DFlash speculative decode on vLLM — single-stream 227 t/s (DFlash), aggregate peaks 807 t/s at 64 agents. The knee sits at 32 (96% of peak at 2× the per-agent speed of 64). This is the model used for the 20-min teamwork builds. The highlighted 32-agent row is its build-challenge fleet.
| agents | prompt_t/s | agg_gen_t/s | per_agent_gen_t/s | ttft_p50_s | ttft_p99_s | wall_s |
|---|---|---|---|---|---|---|
| 1 | 3548.3 | 227.5 | 227.6 | 0.20 | 0.20 | 1.1 |
| 2 | 1462.9 | 131.6 | 96.9 | 0.96 | 0.96 | 3.9 |
| 4 | 6262.5 | 401.6 | 128.0 | 0.45 | 0.45 | 2.5 |
| 8 | 6432.1 | 482.3 | 83.8 | 0.88 | 0.88 | 4.2 |
| 16 | 6383.4 | 495.8 | 48.4 | 1.78 | 1.78 | 8.3 |
| 24 | 6507.1 | 743.1 | 40.6 | 2.63 | 2.63 | 8.3 |
| 32 | 6803.0 | 778.2 | 33.9 | 3.35 | 3.36 | 10.5 |
| 48 | 5903.3 | 741.2 | 20.9 | 5.48 | 5.81 | 16.6 |
| 64 | 6449.4 | 807.3 | 16.3 | 5.52 | 7.09 | 20.3 |
Current live gemma4 plain vLLM serve on Spark/GB10, max_model_len=32768, max_num_seqs=48, no DFlash/speculative decode. The 48-agent row is the measured server ceiling and the aggregate decode peak in this sweep; game builds use the separate 5-agent context-calibrated team size. The highlighted 48-agent row is its build-challenge fleet.
| agents | prompt_t/s | agg_gen_t/s | per_agent_gen_t/s | wall_s |
|---|---|---|---|---|
| 1 | 10.7 | 28.6 | 28.6 | 6.71 |
| 2 | 21.9 | 58.4 | 29.3 | 6.58 |
| 4 | 40.3 | 107.5 | 27.0 | 7.14 |
| 8 | 69.5 | 185.4 | 23.2 | 8.28 |
| 16 | 120.1 | 318.6 | 20.0 | 9.64 |
| 24 | 172.8 | 457.1 | 19.1 | 10.08 |
| 32 | 213.1 | 562.9 | 17.6 | 10.91 |
| 48 | 298.0 | 785.9 | 16.4 | 11.73 |
NVIDIA nvidia/Qwen3.6-35B-A3B-NVFP4 on Spark/GB10 via vLLM 0.23.0, modelopt NVFP4 (W4A16) + FP8 attention, MARLIN MoE, max_model_len=32768. This is a 4-bit quantized model — roughly 3× smaller in VRAM than the BF16 base, so more KV-cache budget remains and aggregate decode is 660 t/s at only 32 agents. The sweep stops at 32 because the Pareto knee is there; higher concurrency only degrades per-agent speed (20.6 t/s → lower), it does not raise aggregate further on a single GB10. The highlighted 32-agent row is its build-challenge fleet.
| agents | prompt_t/s | agg_gen_t/s | per_agent_gen_t/s | wall_s |
|---|---|---|---|---|
| 1 | 20.2 | 77.3 | 77.3 | 3.31 |
| 2 | 32.3 | 123.4 | 61.7 | 4.15 |
| 4 | 50.9 | 194.4 | 48.7 | 5.27 |
| 8 | 88.3 | 337.5 | 42.2 | 6.07 |
| 16 | 128.8 | 489.3 | 30.6 | 8.37 |
| 32 | 174.6 | 660.3 | 20.6 | 12.41 |
InternScience/Agents-A1 — 35B MoE agentic model, BF16 safetensors on Spark/GB10 via vLLM 0.23.0, FlashInfer CUTLASS MoE, max_model_len=131072. Full-precision (BF16) means a 65.5 GiB model footprint, leaving a smaller KV-cache than the NVFP4 condition, so aggregate is 349 t/s at 32 agents — within 1.4% of the Qwen3-Coder-30B-A3B BF16 baseline (344 t/s), as expected for the same architecture class. The sweep stops at 32 (Pareto knee); the GPU is saturated beyond that. The highlighted 32-agent row is its build-challenge fleet.
| agents | prompt_t/s | agg_gen_t/s | per_agent_gen_t/s | wall_s |
|---|---|---|---|---|
| 1 | 7.9 | 30.4 | 30.4 | 8.43 |
| 2 | 14.6 | 55.7 | 27.8 | 9.20 |
| 4 | 23.6 | 90.3 | 22.6 | 11.34 |
| 8 | 36.1 | 137.9 | 17.2 | 14.86 |
| 16 | 60.0 | 228.0 | 14.3 | 17.97 |
| 32 | 92.2 | 348.9 | 10.9 | 23.48 |
Single B70, llama.cpp SYCL, same-architecture Qwen3.5-35B-A3B fine-tunes plus DREAM-v0-7B. Date: 2026-06-27.
| Model | Arch | Params | Quant | Size | Decode |
|---|---|---|---|---|---|
| AgentWorld-35B-A3B | qwen35moe | 34.7B / ~3B active | Q5_K_M | 24.7 GB | autoregressive |
| NEX2 | qwen35moe | 34.7B / ~3B active | Q5_K_M | 24.7 GB | autoregressive |
| SIQ-1-35B | qwen35moe | 34.7B / ~3B active | Q5_K_M | 24.7 GB | autoregressive |
| Ornith-1.0-35B | qwen35moe | 34.7B / ~3B active | Q5_K_M | 24.7 GB | autoregressive |
| QUEST-30B-RL | qwen3moe | 30B / ~3B active | Q5_K_M | 21.7 GB | autoregressive |
| Laguna XS 2.1 | laguna moe | 33B / ~3B active | Q4_K_M | official GGUF | autoregressive |
| DREAM-v0-Instruct-7B | dream | 7B dense | Q4_K_M | 4.7 GB | masked diffusion |
These bars are B70-only. The 35B models share the same qwen35moe topology and converge on the same speed curve: ~75 tok/s single stream, ~151-160 tok/s at the 32-agent production knee, and ~169-178 tok/s at the 48-agent peak. The 56-agent row is the memory cliff.






| Agents | AgentWorld | NEX2 | SIQ | Ornith |
|---|---|---|---|---|
| 1 | 76.4 | 75.8 | 75.5 | 75.9 |
| 8 | 107.8 | 103.3 | 106.5 | 106.8 |
| 16 | 127.8 | 121.7 | 126.0 | 126.3 |
| 32 | 159.9 | 151.5 | 157.6 | 157.8 |
| 48 | 178.2 | 168.9 | 175.5 | 175.5 |
| 56 | 10.1 | cliff | cliff | cliff |
Peak is 48 agents; 56 agents is the 30 GiB memory cliff. Production knee is 32 agents.
| Agents | AgentWorld | Ornith | NEX2 | SIQ | Quest Turbo | Laguna XS 2.1 | Qwen3.6 Q5 Turbo |
|---|---|---|---|---|---|---|---|
| 1 | 62.6 | 63.1 | 62.7 | 63.1 | 66.5 | 63.7 | 68.2 |
| 4 | — | — | — | — | — | — | 128.0 |
| 8 | 141.5 | 142.2 | 142.3 | 141.9 | 164.7 | 136.0 | 156.7 |
| 12 | — | — | — | — | — | — | 158.5 |
| 16 | 163.3 | 163.4 | 163.0 | 162.8 | 223.0 | 183.1 | — |
| 32 | 223.9 | 224.9 | 224.3 | 223.7 | 335.5 | 283.1 | — |
Live values are simultaneous OpenAI-style requests through llama-server; Quest is qwen3moe and uses a 131k safe context ceiling.
| Per-agent depth | Max agents | AgentWorld | NEX2 | SIQ | Ornith |
|---|---|---|---|---|---|
| 512 | 48 | 209.0 | 201.2 | 206.4 | 206.7 |
| 2,048 | 48 | 177.7 | 169.4 | 175.6 | 175.9 |
| 8,192 | 28 | 105.2 | 101.1 | 104.8 | 104.8 |
| 16,384 | 15 | 65.2 | 63.4 | 65.2 | 65.2 |
| 32,768 | 6 | 34.4 | 33.9 | 34.4 | 34.4 |
| 65,536 | 3 | 22.4 | 22.2 | 22.4 | 22.3 |
Concurrency and per-agent context trade against the same B70 memory budget.
| Serve | Prompt tokens | Prefill tok/s | Decode tok/s | Status |
|---|---|---|---|---|
| np=1 | 805 | 655.8 | 53.8 | OK |
| np=1 | 14563 | 917.0 | 8.8 | OK |
| np=1 | 61341 | 283.0 | 1.0 | OK |
| np=1 | 129325 | 139.0 | 0.2 | OK |
| np=1 | 0 | 0.0 | 0.0 | ERR(1/1) |
| np=8 | 805 | 279.8 | 100.2 | OK |
| np=8 | 6963 | 406.0 | 22.4 | OK |
| np=8 | 14563 | 291.9 | 9.4 | OK |
| np=8 | 0 | 0.0 | 0.0 | ERR(8/8) |
| Steps | Tok/s | ms/step | Spot-check quality |
|---|---|---|---|
| 8 | 45.9 | 348.6 | incoherent |
| 16 | 26.2 | 304.8 | broken |
| 32 | 13.9 | 287.3 | partial |
| 64 | 7.3 | 275.4 | partial |
| 128 | 3.7 | 273.2 | terse |
| 256 | 1.8 | 271.2 | coherent |
| Game | Model | Variant | Chars | Quality /50 | Asserts | Playable | Wall | Tokens |
|---|---|---|---|---|---|---|---|---|
| Road Hopper | Laguna XS 2.1 | Laguna XS 2.1 Q4_K_M B70 live teamwork eval | 14,196 | 6/6 | 1,362.1 s | |||
| Robot-Filled Maze Shooter | Laguna XS 2.1 | Laguna XS 2.1 Q4_K_M B70 live teamwork eval | 18,440 | 6/7 | 1,805.5 s | |||
| Road Hopper | AgentWorld-35B | Solo | 17,970 | 7/7 | yes | 87 s | 6,470 | |
| Road Hopper | AgentWorld-35B | CrewAI N=6 | 13,684 | 7/7 | yes | 232 s | 28,668 | |
| Road Hopper | AgentWorld-35B | CrewAI N=12 | 23,809 | 7/7 | yes | 396 s | 47,065 | |
| Robot-Filled Maze Shooter | AgentWorld-35B | Solo | 11,894 | 6/7 | yes | 69 s | 5,668 | |
| Robot-Filled Maze Shooter | AgentWorld-35B | CrewAI N=12 | 24,923 | 7/7 | yes | 416 s | 48,036 | |
| Road Hopper | Frosty Ornith-35B | B70 live teamwork eval | 29,268 | 33 | 6/6 | source-review fail: unwinnable | 1,694.6 s | |
| Robot-Filled Maze Shooter | Frosty Ornith-35B | B70 live teamwork eval | 22,724 | 34 | 7/7 | source-review fail: cannot shoot | 1,078.3 s | |
| Robot-Filled Maze Shooter | Frosty Ornith-35B | B70 failed attempt | no | |||||
| Road Hopper | NEX2 | B70 live teamwork eval | 28,146 | 37 | 6/6 | yes | 1,316.8 s | |
| Robot-Filled Maze Shooter | NEX2 | B70 live teamwork eval | 30,749 | 31 | 7/7 | source-review fail: invisible/unshootable robots | 2,670.2 s | |
| Road Hopper | SIQ-1-35B | B70 live teamwork eval | 40,828 | 42 | 6/6 | yes | 2,480.9 s | |
| Robot-Filled Maze Shooter | SIQ-1-35B | B70 live teamwork eval | 19,185 | 30 | 7/7 | source-review fail: frozen player | 1,064.7 s | |
| Robot-Filled Maze Shooter | Quest Turbo | B70 live teamwork eval | 20,597 | 29 | 7/7 | source-review fail: overlay never hides | 3,115.7 s | |
| Road Hopper | AgentWorld-35B | B70 stock baseline | 27,520 | 6/6 | 1,834.2 s | |||
| Robot-Filled Maze Shooter | AgentWorld-35B | B70 stock baseline | 29,933 | 5/6 | 2,044.4 s | |||
| Road Hopper | Frosty Ornith-35B | B70 stock baseline | 17,891 | 6/6 | 1,081.7 s | |||
| Robot-Filled Maze Shooter | Frosty Ornith-35B | B70 stock baseline | 19,692 | 6/6 | no; runtime validator failure | 989.8 s | ||
| Road Hopper | AgentWorld-35B | B70 live teamwork eval | 40,439 | 31 | 6/6 | source-review fail: unwinnable | 1,262.7 s | |
| Robot-Filled Maze Shooter | AgentWorld-35B | B70 live teamwork eval | 27,483 | 40 | 7/7 | yes | 1,038.1 s | |
| Road Hopper | AgentWorld-35B | B70 Q5_K_M quant eval | 19,345 | 34 | 6/6 | yes | 908.4 s | |
| Robot-Filled Maze Shooter | AgentWorld-35B | B70 Q5_K_M quant eval | 24,135 | 40 | 7/7 | yes | 1,105.0 s | |
| Road Hopper | AgentWorld-35B | B70 Q6_K quant eval run 1 | 15,227 | 34 | 6/6 | yes | 971.3 s | |
| Robot-Filled Maze Shooter | AgentWorld-35B | B70 Q6_K quant eval run 1 | 23,030 | 40 | 7/7 | yes | 1,816.9 s | |
| Road Hopper | AgentWorld-35B | B70 Q6_K quant eval run 2 | 27,993 | 35 | 6/6 | yes | 1,598.6 s | |
| Robot-Filled Maze Shooter | AgentWorld-35B | B70 Q6_K quant eval run 2 | 21,541 | 40 | 7/7 | yes | 1,292.4 s |
Raw B70 report: CROSS_MODEL_B70.md · CROSS_MODEL_LIVE_B70.md · ORNITH_LIVE_EVALS.md · NEX2_LIVE_EVALS.md · SIQ_LIVE_EVALS.md · QUEST_LIVE_EVALS.md · LAGUNA_XS21_LIVE_EVALS.md.
results_throughput.csv · results_throughput_aeon.csv · results_throughput_aeon_ornith.csv · results_throughput_qwen36_35b.csv · results_throughput_qwen36_35b_q5_turbo.csv · results_throughput_agentworld_spark.csv · results_throughput_qwen3_coder_30b_a3b_instruct_bf16.csv · results_throughput_nemotron3_super.csv · results_throughput_nemotron_labs_3_puzzle_75b_a9b_nvfp4.csv · results_throughput_agentworld.csv · results_throughput_gemma.csv · results_throughput_gemma4_nvfp4_current.csv · results_throughput_qwen36_35b_nvfp4.csv · results_throughput_agents_a1.csv