Aircade retro glove controller header

Benchmarks & token stats

Measured decode throughput for every serve behind the arcade — aggregate tok/s, per-agent tok/s, and how both scale with concurrent agents on the NVIDIA GB10 and the Intel Arc Pro B70. Every number links back to its source CSV or run report.

Benchmark Binder

21 throughput sweeps · NVIDIA GB10 + Intel Arc Pro B70 · 85 game artifacts
All modelsClick a bar, a line, or a table row to isolate one serve across every page.
Page 1 · game results

How did the games do?

Only an explicit playable result earns a green sticker. Static checks alone do not count.

85 recorded artifacts
28

Playable

explicit result recorded

10

Needs fixing

recorded problem

47

Not checked yet

no playability result

Open model × game grid · 27 models
ModelRoad HopperRobot-Filled Maze ShooterRibbit Rush
AEON QWEN
1 not checked
1 recorded artifact
1 not checked
1 recorded artifact
Agents-A1
1 playable
1 recorded artifact
1 playable
1 recorded artifact
AgentWorld-35B
6 playable1 needs fixing2 not checked
9 recorded artifacts
6 playable2 not checked
8 recorded artifacts
Claude Opus 4.8
2 not checked
2 recorded artifacts
2 not checked
2 recorded artifacts
Claude Sonnet 5
1 not checked
1 recorded artifact
DeepSeek V4 Pro
1 not checked
1 recorded artifact
1 not checked
1 recorded artifact
Fable
1 not checked
1 recorded artifact
1 not checked
1 recorded artifact
Frosty Ornith-35B
1 needs fixing2 not checked
3 recorded artifacts
3 needs fixing1 not checked
4 recorded artifacts
Gemma-4 current
1 not checked
1 recorded artifact
1 not checked
1 recorded artifact
Gemma-4-26B
1 not checked
1 recorded artifact
1 not checked
1 recorded artifact
GLM-5.2
1 not checked
1 recorded artifact
1 not checked
1 recorded artifact
GPT-5.5 xhigh
1 not checked
1 recorded artifact
1 not checked
1 recorded artifact
GPT-5.6 Sol
1 not checked
1 recorded artifact
1 not checked
1 recorded artifact
GPT-5.6 Terra
1 not checked
1 recorded artifact
1 not checked
1 recorded artifact
Laguna XS 2.1
1 not checked
1 recorded artifact
2 not checked
2 recorded artifacts
Leanstral NVFP4
1 needs fixing
1 recorded artifact
1 needs fixing
1 recorded artifact
LongCat-2.0
6 not checked
6 recorded artifacts
2 not checked
2 recorded artifacts
Nemotron 3 Super
1 playable
1 recorded artifact
1 playable
1 recorded artifact
Nemotron Puzzle 75B
1 playable
1 recorded artifact
1 playable
1 recorded artifact
NEX2
1 playable1 not checked
2 recorded artifacts
1 needs fixing2 not checked
3 recorded artifacts
Ornith-AEON
1 not checked
1 recorded artifact
1 not checked
1 recorded artifact
Quest Turbo
1 needs fixing
1 recorded artifact
Qwen3 Coder 30B
1 playable
1 recorded artifact
1 playable
1 recorded artifact
Qwen3.6-35B Base
1 not checked
1 recorded artifact
Qwen3.6-35B NVFP4
1 playable
1 recorded artifact
1 playable
1 recorded artifact
SIQ-1-35B
2 playable
2 recorded artifacts
1 playable1 needs fixing
2 recorded artifacts
StepFun Step-3.7
1 playable
1 recorded artifact
1 playable
1 recorded artifact
Open all 85 artifact records
StatusModel / conditionGameVariantStatic checksReview scoreBuild wallEvidence
Needs fixingAgentWorld-35BIntel Arc Pro B70Road HopperB70 live teamwork eval6/6311,262.7 s
Needs fixingFrosty Ornith-35BIntel Arc Pro B70Road HopperB70 live teamwork eval6/6331,694.6 s
Needs fixingFrosty Ornith-35BIntel Arc Pro B70Robot-Filled Maze ShooterB70 failed attempt
Needs fixingFrosty Ornith-35BIntel Arc Pro B70Robot-Filled Maze ShooterB70 live teamwork eval7/7341,078.3 s
Needs fixingFrosty Ornith-35BIntel Arc Pro B70Robot-Filled Maze ShooterB70 stock baseline6/6989.8 s
Needs fixingLeanstral NVFP4DGX Spark / NVIDIA GB10Road HopperLeanstral NVFP4 Road Hopper eval8/11228.5 s
Needs fixingLeanstral NVFP4DGX Spark / NVIDIA GB10Robot-Filled Maze ShooterLeanstral NVFP4 Robot Maze eval8/1175.8 s
Needs fixingNEX2Intel Arc Pro B70Robot-Filled Maze ShooterB70 live teamwork eval7/7312,670.2 s
Needs fixingQuest TurboIntel Arc Pro B70Robot-Filled Maze ShooterB70 live teamwork eval7/7293,115.7 s
Needs fixingSIQ-1-35BIntel Arc Pro B70Robot-Filled Maze ShooterB70 live teamwork eval7/7301,064.7 s
PlayableAgents-A1DGX Spark / NVIDIA GB10Road HopperAgents-A1 BF16 Road Hopper eval6/61,436.5 s
PlayableAgents-A1DGX Spark / NVIDIA GB10Robot-Filled Maze ShooterAgents-A1 BF16 Maze eval7/71,678.5 s
PlayableAgentWorld-35BIntel Arc Pro B70Road HopperB70 Q5_K_M quant eval6/634908.4 s
PlayableAgentWorld-35BIntel Arc Pro B70Road HopperB70 Q6_K quant eval run 16/634971.3 s
PlayableAgentWorld-35BIntel Arc Pro B70Road HopperB70 Q6_K quant eval run 26/6351,598.6 s
PlayableAgentWorld-35BIntel Arc Pro B70Road HopperCrewAI N=127/7396 s
PlayableAgentWorld-35BIntel Arc Pro B70Road HopperCrewAI N=67/7232 s
PlayableAgentWorld-35BIntel Arc Pro B70Road HopperSolo7/787 s
PlayableAgentWorld-35BIntel Arc Pro B70Robot-Filled Maze ShooterB70 live teamwork eval7/7401,038.1 s
PlayableAgentWorld-35BIntel Arc Pro B70Robot-Filled Maze ShooterB70 Q5_K_M quant eval7/7401,105.0 s
PlayableAgentWorld-35BIntel Arc Pro B70Robot-Filled Maze ShooterB70 Q6_K quant eval run 17/7401,816.9 s
PlayableAgentWorld-35BIntel Arc Pro B70Robot-Filled Maze ShooterB70 Q6_K quant eval run 27/7401,292.4 s
PlayableAgentWorld-35BIntel Arc Pro B70Robot-Filled Maze ShooterCrewAI N=127/7416 s
PlayableAgentWorld-35BIntel Arc Pro B70Robot-Filled Maze ShooterSolo6/769 s
PlayableNemotron 3 SuperDGX Spark / NVIDIA GB10Road HopperNemotron 3 Super Spark eval6/61,200.7 s
PlayableNemotron 3 SuperDGX Spark / NVIDIA GB10Robot-Filled Maze ShooterNemotron 3 Super Spark eval7/71,202.8 s
PlayableNemotron Puzzle 75BDGX Spark / NVIDIA GB10Road HopperNemotron Puzzle 75B Spark eval6/61,200.9 s
PlayableNemotron Puzzle 75BDGX Spark / NVIDIA GB10Robot-Filled Maze ShooterNemotron Puzzle 75B Spark eval7/71,201.4 s
PlayableNEX2Intel Arc Pro B70Road HopperB70 live teamwork eval6/6371,316.8 s
PlayableQwen3 Coder 30BDGX Spark / NVIDIA GB10Road HopperQwen3 Coder BF16 long-context teamwork eval6/61,326.8 s
PlayableQwen3 Coder 30BDGX Spark / NVIDIA GB10Robot-Filled Maze ShooterQwen3 Coder BF16 long-context teamwork eval7/72,235.6 s
PlayableQwen3.6-35B NVFP4DGX Spark / NVIDIA GB10Road HopperQwen3.6-35B NVFP4 Road Hopper eval6/61,265 s
PlayableQwen3.6-35B NVFP4DGX Spark / NVIDIA GB10Robot-Filled Maze ShooterQwen3.6-35B NVFP4 Maze eval7/71,511 s
PlayableSIQ-1-35BIntel Arc Pro B70Road HopperB70 live teamwork eval6/6422,480.9 s
PlayableSIQ-1-35BDGX Spark / NVIDIA GB10Road HopperGB10 teamwork eval6/61,367.1 s
PlayableSIQ-1-35BDGX Spark / NVIDIA GB10Robot-Filled Maze ShooterGB10 teamwork eval7/71,355.3 s
PlayableStepFun Step-3.7DGX Spark / NVIDIA GB10Road HopperStep37 teamwork eval6/63,284.3 s
PlayableStepFun Step-3.7DGX Spark / NVIDIA GB10Robot-Filled Maze ShooterStep37 teamwork eval7/73,380.8 s
Not checkedAEON QWENDGX Spark / NVIDIA GB10Road Hopperteamwork eval6/61,753.4 s
Not checkedAEON QWENDGX Spark / NVIDIA GB10Robot-Filled Maze Shooterteamwork eval7/71,256.1 s
Not checkedAgentWorld-35BIntel Arc Pro B70Road HopperB70 stock baseline6/61,834.2 s
Not checkedAgentWorld-35BDGX Spark / NVIDIA GB10Road Hopperteamwork eval6/61,607.5 s
Not checkedAgentWorld-35BIntel Arc Pro B70Robot-Filled Maze ShooterB70 stock baseline5/62,044.4 s
Not checkedAgentWorld-35BDGX Spark / NVIDIA GB10Robot-Filled Maze Shooterteamwork eval7/71,502.7 s
Not checkedClaude Opus 4.8Anthropic APIRoad HopperFROGGAI B70 controlteamwork/validate_game.js ok47/50
Not checkedClaude Opus 4.8Anthropic APIRoad HopperFroggAI control (gold-standard baseline)7/7300.0 s
Not checkedClaude Opus 4.8Anthropic APIRobot-Filled Maze ShooterMAZERUNNER B70 controlteamwork/validate_game.js ok47/50
Not checkedClaude Opus 4.8Anthropic APIRobot-Filled Maze ShooterMAZERUNNER control (gold-standard baseline)7/7300.0 s
Not checkedClaude Sonnet 5Anthropic APIRibbit RushRibbit Rush standalone game
Not checkedDeepSeek V4 ProDeepSeek APIRoad HopperFroggAI deepseek eval7/7120.0 s
Not checkedDeepSeek V4 ProDeepSeek APIRobot-Filled Maze ShooterMaze deepseek eval7/7150.0 s
Not checkedFableexternal submissionRoad HopperFable Road Hopper
Not checkedFableexternal submissionRobot-Filled Maze ShooterFable Maze Shooter
Not checkedFrosty Ornith-35BIntel Arc Pro B70Road HopperB70 stock baseline6/61,081.7 s
Not checkedFrosty Ornith-35BDGX Spark / NVIDIA GB10Road Hopperteamwork eval6/61,784.4 s
Not checkedFrosty Ornith-35BDGX Spark / NVIDIA GB10Robot-Filled Maze Shooterteamwork eval7/71,486.2 s
Not checkedGemma-4 currentDGX Spark / NVIDIA GB10Road Hoppercurrent plain vLLM eval6/61,604.1 s
Not checkedGemma-4 currentDGX Spark / NVIDIA GB10Robot-Filled Maze Shootercurrent plain vLLM eval6/71,663.0 s
Not checkedGemma-4-26BDGX Spark / NVIDIA GB10Road Hopperteamwork eval6/61,506.2 s
Not checkedGemma-4-26BDGX Spark / NVIDIA GB10Robot-Filled Maze Shooterteamwork eval7/71,316.6 s
Not checkedGLM-5.2APIRoad HopperGLM-5.2 OpenCode eval
Not checkedGLM-5.2APIRobot-Filled Maze ShooterGLM-5.2 OpenCode eval
Not checkedGPT-5.5 xhighOpenAI APIRoad HopperFroggAI xhigh eval7/7210.0 s
Not checkedGPT-5.5 xhighOpenAI APIRobot-Filled Maze ShooterMaze xhigh eval7/7246.0 s
Not checkedGPT-5.6 SolOpenAI API via Codex DesktopRoad HopperGPT-5.6 Sol road-and-river one-shot
Not checkedGPT-5.6 SolOpenAI API via Codex DesktopRobot-Filled Maze ShooterGPT-5.6 Sol maze-shooter one-shot
Not checkedGPT-5.6 TerraOpenAI API via Codex DesktopRoad HopperGPT-5.6 Terra road-and-river one-shot
Not checkedGPT-5.6 TerraOpenAI API via Codex DesktopRobot-Filled Maze ShooterGPT-5.6 Terra maze-shooter one-shot
Not checkedLaguna XS 2.1Intel Arc Pro B70Road HopperLaguna XS 2.1 Q4_K_M B70 live teamwork eval6/61,362.1 s
Not checkedLaguna XS 2.1Intel Arc Pro B70Robot-Filled Maze ShooterLaguna XS 2.1 Q4_K_M B70 live teamwork eval6/71,805.5 s
Not checkedLaguna XS 2.1DGX Spark / NVIDIA GB10Robot-Filled Maze Shooterv2
Not checkedLongCat-2.0Meituan hosted APIRoad HopperClassic pixel-art Frogger styleteamwork/validate_game.js ok
Not checkedLongCat-2.0Meituan hosted APIRoad HopperLongCat chain3 Road Hoppernot run: jsdom missing
Not checkedLongCat-2.0Meituan hosted APIRoad HopperLongCat one-shot Road Hoppernot run: jsdom missing
Not checkedLongCat-2.0Meituan hosted APIRoad HopperNoir rain skyline lantern-glowteamwork/validate_game.js ok
Not checkedLongCat-2.0Meituan hosted APIRoad HopperSynthwave neonteamwork/validate_game.js ok
Not checkedLongCat-2.0Meituan hosted APIRoad HopperTerminal ASCII green-on-blackteamwork/validate_game.js ok
Not checkedLongCat-2.0Meituan hosted APIRobot-Filled Maze ShooterLongCat chain3 Mazeteamwork/validate_game.js ok
Not checkedLongCat-2.0Meituan hosted APIRobot-Filled Maze ShooterLongCat one-shot Mazeteamwork/validate_game.js ok
Not checkedNEX2DGX Spark / NVIDIA GB10Road Hopperteamwork eval6/61,863.3 s
Not checkedNEX2Turbo via HydraRobot-Filled Maze ShooterHydra Salvage Run3,600.0 s
Not checkedNEX2DGX Spark / NVIDIA GB10Robot-Filled Maze Shooterteamwork eval7/71,209.5 s
Not checkedOrnith-AEONDGX Spark / NVIDIA GB10Road HopperAEON Ornith/Qwen eval6/61,594.3 s
Not checkedOrnith-AEONDGX Spark / NVIDIA GB10Robot-Filled Maze ShooterAEON Ornith/Qwen eval7/71,449.0 s
Not checkedQwen3.6-35B BaseDGX Spark / NVIDIA GB10Road HopperQwen3.6 base GGUF Road Hopper scout
21 sweeps

Bars and table rows use each serve's operating point; curves show the full recorded sweep. Blank cells mean the source did not record it.

Loose-leaf appendix · full tables, old charts, methodology, and downloads

Ornith-1.0-35B — agent-concurrency throughput Pareto

Batched generation throughput as the number of concurrent agents grows (llama.cpp, GGUF-NVFP4). Aggregate generation reaches 372.28 t/s at 256 agents; the separate total-throughput column reaches 848.67 t/s. Per-agent speed and single-reply latency degrade as concurrency rises. The highlighted 32-agent row is the recorded fleet selection.

AgentsPrompt t/sAgg gen t/sPer-agent t/sTotal t/s256-tok reply (s)
12286.4068.9868.98195.153.711
82372.40206.4825.81527.609.919
162355.98225.7414.11568.3118.145
322352.62277.158.66672.9129.558
642351.41326.315.10766.2550.210
962355.60342.593.57796.1871.736
1282353.23350.622.74810.3793.458
1922357.97361.531.88830.05135.957
2562356.25372.281.45848.67176.038
Throughput scaling vs concurrent agents — aggregate climbs, per-agent decays.
Throughput scaling vs concurrent agents — aggregate climbs, per-agent decays.
Agent Pareto frontier — aggregate throughput vs single-reply latency (labels = #agents).
Agent Pareto frontier — aggregate throughput vs single-reply latency (labels = #agents).

AEON QWEN — concurrency Pareto

NVFP4 MTP-XS on vLLM with z-lab DFlash speculative decode. DFlash makes single streams fast (~46–48 t/s), so aggregate plateaus early — the knee is 8 (91% of peak at 2× the per-agent speed of 16). The highlighted 8-agent row is its build-challenge fleet.

agentsprompt_t/sagg_gen_t/sper_agent_gen_t/sttft_p50_sttft_p99_swall_s
112.5 (cold-start)4.44.455.4155.4158.4
2915.5100.752.01.531.535.1
42417.4165.647.51.161.166.2
82057.5218.329.92.742.749.4
161676.6232.616.46.786.7817.6
241581.6251.111.410.7710.8124.5
321628.1252.18.612.5214.0232.5

conc=1 is cold-start (first request compiles CUDA graphs); true single-stream is the rate at conc 2–4.

Ornith-AEON — concurrency Pareto

Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4 on the local KV-fix vLLM image with z-lab DFlash drafter n=6. This is measured through the OpenAI-compatible API using real usage token counts and batch wall-clock time; the selected build fleet is the capped 16-agent serve. The highlighted 16-agent row is its build-challenge fleet.

agentsprompt_t/sagg_gen_t/sper_agent_gen_t/swall_s
119.976.276.23.36
236.4139.269.63.68
454.2206.953.04.95
889.0340.043.96.02
16122.7466.130.58.79

Qwen3.6-35B-A3B base — concurrency Pareto

Generic Qwen3.6-35B-A3B UD-Q6_K GGUF on Spark/GB10 via llama.cpp-qwen36, -c 32768 -np 32, q8 KV, FlashAttention, and fixed 256 generated tokens per agent using min_tokens=256 + ignore_eos. This is a GGUF llama.cpp condition, not the vLLM/NVFP4 stack. The highlighted 32-agent row is its build-challenge fleet.

agentsprompt_t/sagg_gen_t/sper_agent_gen_t/swall_s
114.756.156.24.56
213.852.926.49.68
414.053.513.419.16
819.474.09.327.67
1632.5123.47.733.19
3264.7244.67.633.49

Qwen3.6-35B-A3B Q5 Turbo on Intel Arc Pro B70 — concurrency Pareto

Qwen3.6-35B-A3B-UD-Q5_K_XL GGUF served by local Turbo/llama.cpp SYCL on Intel Arc Pro B70 with a 12-slot unified KV configuration and 262144 context. This live OpenAI-compatible run peaks at 158.5 decode tok/s at the 12-agent serve ceiling, with 16.3 tok/s per agent. The highlighted 12-agent row is the 12-slot B70 live serve ceiling.

agentsprompt_t/sagg_gen_t/sper_agent_gen_t/swall_s
1307.9368.1579.123.76
4116.59127.9737.708.00
878.41156.7224.1013.07
1261.75158.4816.3119.38

Qwen AgentWorld-35B-A3B Spark — concurrency Pareto

Qwen/Qwen-AgentWorld-35B-A3B local GGUF_NVFP4.official artifact on Spark/GB10 via llama.cpp-qwen36, -c 32768 -np 32, q8 KV, FlashAttention, and fixed 256 generated tokens per agent using min_tokens=256 + ignore_eos. This is the controlled Spark rerun of AgentWorld, kept separate from the older site AgentWorld CSV. The highlighted 32-agent row is its build-challenge fleet.

agentsprompt_t/sagg_gen_t/sper_agent_gen_t/swall_s
113.451.351.34.99
222.184.342.26.07
432.5124.231.18.24
844.8171.121.411.97
1653.1201.712.620.31
3264.9245.47.733.38

Qwen3-Coder-30B-A3B-Instruct BF16 — concurrency Pareto

Official Qwen/Qwen3-Coder-30B-A3B-Instruct BF16 safetensors on Spark/GB10 via vLLM 0.23.0, max_model_len=32768, max_num_seqs=32, FlashAttention, and unquantized FlashInfer CUTLASS MoE. No REAP, no DFlash, no drafter, no quantized derivative. Fixed 256 generated tokens per agent using min_tokens=256 + ignore_eos. The highlighted 32-agent row is its build-challenge fleet.

agentsprompt_t/sagg_gen_t/sper_agent_gen_t/swall_s
17.630.330.38.46
213.252.626.39.73
420.180.320.112.75
830.7122.815.416.67
1653.1211.213.219.39
3287.0344.210.823.80

NVIDIA Nemotron-3-Super-120B-A12B-NVFP4 — concurrency Pareto

Official NVIDIA NVFP4 safetensors on Spark/GB10 via vLLM 0.23.0, max_model_len=8192, max_num_seqs=1, kv_cache_memory_bytes=4G, and eager execution. This is a single-agent scout condition, not a wide concurrency Pareto, because this Spark serve was constrained to one active sequence. The highlighted 1-agent row is its build-challenge fleet.

agentsprompt_t/sagg_gen_t/sper_agent_gen_t/swall_s
14.913.913.913.80

NVIDIA Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4 — concurrency Pareto

Official NVIDIA Puzzle-75B-A9B NVFP4 checkpoint on Spark/GB10 via vLLM 0.23.0, max_model_len=8192, MTP speculative decode num_speculative_tokens=1, max_num_seqs=32, kv_cache_memory_bytes=16G, FlashInfer attention/Mamba, ModelOpt mixed NVFP4/FP8, and eager execution. Fixed-token sweep used min_tokens=192 + ignore_eos; the selected build fleet is the measured 32-agent ceiling, while game builds used the separate context-calibrated 1-agent team. The highlighted 32-agent row is its build-challenge fleet.

agentsprompt_t/sagg_gen_t/sper_agent_gen_t/swall_s
18.824.924.97.72
214.541.020.89.37
425.572.118.310.65
845.6128.716.311.94
1672.3202.913.015.14
2493.1260.611.117.68
32114.4319.810.319.21

AgentWorld-35B-A3B — concurrency Pareto

GGUF-NVFP4 (qwen35moe) on llama.cpp at -np 32 — same architecture and size as Ornith, and its curve tracks Ornith's, so the knee sits at 32. The highlighted 32-agent row is its build-challenge fleet.

agentsprompt_t/sagg_gen_t/sper_agent_gen_t/sttft_p50_sttft_p99_swall_s
11015.646.846.80.680.685.5
21452.671.936.10.730.967.1
41398.899.425.01.952.0110.3
81299.2121.515.33.404.3416.9
161230.8132.28.45.989.2431.0
241085.2120.05.18.7615.7651.2
321055.7145.44.711.8721.6256.3

Gemma-4-26B-A4B — concurrency Pareto

NVFP4 weights with z-lab DFlash speculative decode on vLLM — single-stream 227 t/s (DFlash), aggregate peaks 807 t/s at 64 agents. The knee sits at 32 (96% of peak at 2× the per-agent speed of 64). This is the model used for the 20-min teamwork builds. The highlighted 32-agent row is its build-challenge fleet.

agentsprompt_t/sagg_gen_t/sper_agent_gen_t/sttft_p50_sttft_p99_swall_s
13548.3227.5227.60.200.201.1
21462.9131.696.90.960.963.9
46262.5401.6128.00.450.452.5
86432.1482.383.80.880.884.2
166383.4495.848.41.781.788.3
246507.1743.140.62.632.638.3
326803.0778.233.93.353.3610.5
485903.3741.220.95.485.8116.6
646449.4807.316.35.527.0920.3

Gemma-4-26B-A4B NVFP4 current plain vLLM — concurrency Pareto

Current live gemma4 plain vLLM serve on Spark/GB10, max_model_len=32768, max_num_seqs=48, no DFlash/speculative decode. The 48-agent row is the measured server ceiling and the aggregate decode peak in this sweep; game builds use the separate 5-agent context-calibrated team size. The highlighted 48-agent row is its build-challenge fleet.

agentsprompt_t/sagg_gen_t/sper_agent_gen_t/swall_s
110.728.628.66.71
221.958.429.36.58
440.3107.527.07.14
869.5185.423.28.28
16120.1318.620.09.64
24172.8457.119.110.08
32213.1562.917.610.91
48298.0785.916.411.73

Qwen3.6-35B-A3B NVFP4 — concurrency Pareto

NVIDIA nvidia/Qwen3.6-35B-A3B-NVFP4 on Spark/GB10 via vLLM 0.23.0, modelopt NVFP4 (W4A16) + FP8 attention, MARLIN MoE, max_model_len=32768. This is a 4-bit quantized model — roughly 3× smaller in VRAM than the BF16 base, so more KV-cache budget remains and aggregate decode is 660 t/s at only 32 agents. The sweep stops at 32 because the Pareto knee is there; higher concurrency only degrades per-agent speed (20.6 t/s → lower), it does not raise aggregate further on a single GB10. The highlighted 32-agent row is its build-challenge fleet.

agentsprompt_t/sagg_gen_t/sper_agent_gen_t/swall_s
120.277.377.33.31
232.3123.461.74.15
450.9194.448.75.27
888.3337.542.26.07
16128.8489.330.68.37
32174.6660.320.612.41

InternScience Agents-A1 BF16 — concurrency Pareto

InternScience/Agents-A1 — 35B MoE agentic model, BF16 safetensors on Spark/GB10 via vLLM 0.23.0, FlashInfer CUTLASS MoE, max_model_len=131072. Full-precision (BF16) means a 65.5 GiB model footprint, leaving a smaller KV-cache than the NVFP4 condition, so aggregate is 349 t/s at 32 agents — within 1.4% of the Qwen3-Coder-30B-A3B BF16 baseline (344 t/s), as expected for the same architecture class. The sweep stops at 32 (Pareto knee); the GPU is saturated beyond that. The highlighted 32-agent row is its build-challenge fleet.

agentsprompt_t/sagg_gen_t/sper_agent_gen_t/swall_s
17.930.430.48.43
214.655.727.89.20
423.690.322.611.34
836.1137.917.214.86
1660.0228.014.317.97
3292.2348.910.923.48

Intel Arc Pro B70 — cross-model tables

Single B70, llama.cpp SYCL, same-architecture Qwen3.5-35B-A3B fine-tunes plus DREAM-v0-7B. Date: 2026-06-27.

B70 model configs

ModelArchParamsQuantSizeDecode
AgentWorld-35B-A3Bqwen35moe34.7B / ~3B activeQ5_K_M24.7 GBautoregressive
NEX2qwen35moe34.7B / ~3B activeQ5_K_M24.7 GBautoregressive
SIQ-1-35Bqwen35moe34.7B / ~3B activeQ5_K_M24.7 GBautoregressive
Ornith-1.0-35Bqwen35moe34.7B / ~3B activeQ5_K_M24.7 GBautoregressive
QUEST-30B-RLqwen3moe30B / ~3B activeQ5_K_M21.7 GBautoregressive
Laguna XS 2.1laguna moe33B / ~3B activeQ4_K_Mofficial GGUFautoregressive
DREAM-v0-Instruct-7Bdream7B denseQ4_K_M4.7 GBmasked diffusion

B70 concurrency graph — solo / production knee / peak

Y: modelX: relative aggregate decode tok/sValue: solo / 32-agent / 48-agent tok/s
AgentWorld
178.2 peak
159.9 knee
76.4 solo
NEX2
168.9 peak
151.5 knee
75.8 solo
SIQ
175.5 peak
157.6 knee
75.5 solo
Ornith
175.5 peak
157.8 knee
75.9 solo

These bars are B70-only. The 35B models share the same qwen35moe topology and converge on the same speed curve: ~75 tok/s single stream, ~151-160 tok/s at the 32-agent production knee, and ~169-178 tok/s at the 48-agent peak. The 56-agent row is the memory cliff.

B70 figures

B70 cross-model concurrency overlay.
B70 cross-model concurrency overlay.
B70 cross-model context overlay.
B70 cross-model context overlay.
B70 live server concurrency overlay.
B70 live server concurrency overlay.
B70 live teamwork code-quality comparison.
B70 live teamwork code-quality comparison.
B70 collaborative game-build comparison.
B70 collaborative game-build comparison.
B70 DREAM diffusion step Pareto.
B70 DREAM diffusion step Pareto.

Concurrency Pareto — aggregate decode tok/s

AgentsAgentWorldNEX2SIQOrnith
176.475.875.575.9
8107.8103.3106.5106.8
16127.8121.7126.0126.3
32159.9151.5157.6157.8
48178.2168.9175.5175.5
5610.1cliffcliffcliff

Peak is 48 agents; 56 agents is the 30 GiB memory cliff. Production knee is 32 agents.

B70 live server fleet — tool prompt aggregate decode tok/s

AgentsAgentWorldOrnithNEX2SIQQuest TurboLaguna XS 2.1Qwen3.6 Q5 Turbo
162.663.162.763.166.563.768.2
4128.0
8141.5142.2142.3141.9164.7136.0156.7
12158.5
16163.3163.4163.0162.8223.0183.1
32223.9224.9224.3223.7335.5283.1

Live values are simultaneous OpenAI-style requests through llama-server; Quest is qwen3moe and uses a 131k safe context ceiling.

Context Pareto — aggregate decode tok/s

Per-agent depthMax agentsAgentWorldNEX2SIQOrnith
51248209.0201.2206.4206.7
2,04848177.7169.4175.6175.9
8,19228105.2101.1104.8104.8
16,3841565.263.465.265.2
32,768634.433.934.434.4
65,536322.422.222.422.3

Concurrency and per-agent context trade against the same B70 memory budget.

Quest Turbo live context ceiling

ServePrompt tokensPrefill tok/sDecode tok/sStatus
np=1805655.853.8OK
np=114563917.08.8OK
np=161341283.01.0OK
np=1129325139.00.2OK
np=100.00.0ERR(1/1)
np=8805279.8100.2OK
np=86963406.022.4OK
np=814563291.99.4OK
np=800.00.0ERR(8/8)

DREAM diffusion Pareto

Y: diffusion stepsX: relative tokens per secondValue: tok/s
8 steps
45.9 tok/s
16 steps
26.2 tok/s
32 steps
13.9 tok/s
64 steps
7.3 tok/s
128 steps
3.7 tok/s
256 steps
1.8 tok/s
StepsTok/sms/stepSpot-check quality
845.9348.6incoherent
1626.2304.8broken
3213.9287.3partial
647.3275.4partial
1283.7273.2terse
2561.8271.2coherent

B70 game-build comparison

GameModelVariantCharsQuality /50AssertsPlayableWallTokens
Road HopperLaguna XS 2.1Laguna XS 2.1 Q4_K_M B70 live teamwork eval14,1966/61,362.1 s
Robot-Filled Maze ShooterLaguna XS 2.1Laguna XS 2.1 Q4_K_M B70 live teamwork eval18,4406/71,805.5 s
Road HopperAgentWorld-35BSolo17,9707/7yes87 s6,470
Road HopperAgentWorld-35BCrewAI N=613,6847/7yes232 s28,668
Road HopperAgentWorld-35BCrewAI N=1223,8097/7yes396 s47,065
Robot-Filled Maze ShooterAgentWorld-35BSolo11,8946/7yes69 s5,668
Robot-Filled Maze ShooterAgentWorld-35BCrewAI N=1224,9237/7yes416 s48,036
Road HopperFrosty Ornith-35BB70 live teamwork eval29,268336/6source-review fail: unwinnable1,694.6 s
Robot-Filled Maze ShooterFrosty Ornith-35BB70 live teamwork eval22,724347/7source-review fail: cannot shoot1,078.3 s
Robot-Filled Maze ShooterFrosty Ornith-35BB70 failed attemptno
Road HopperNEX2B70 live teamwork eval28,146376/6yes1,316.8 s
Robot-Filled Maze ShooterNEX2B70 live teamwork eval30,749317/7source-review fail: invisible/unshootable robots2,670.2 s
Road HopperSIQ-1-35BB70 live teamwork eval40,828426/6yes2,480.9 s
Robot-Filled Maze ShooterSIQ-1-35BB70 live teamwork eval19,185307/7source-review fail: frozen player1,064.7 s
Robot-Filled Maze ShooterQuest TurboB70 live teamwork eval20,597297/7source-review fail: overlay never hides3,115.7 s
Road HopperAgentWorld-35BB70 stock baseline27,5206/61,834.2 s
Robot-Filled Maze ShooterAgentWorld-35BB70 stock baseline29,9335/62,044.4 s
Road HopperFrosty Ornith-35BB70 stock baseline17,8916/61,081.7 s
Robot-Filled Maze ShooterFrosty Ornith-35BB70 stock baseline19,6926/6no; runtime validator failure989.8 s
Road HopperAgentWorld-35BB70 live teamwork eval40,439316/6source-review fail: unwinnable1,262.7 s
Robot-Filled Maze ShooterAgentWorld-35BB70 live teamwork eval27,483407/7yes1,038.1 s
Road HopperAgentWorld-35BB70 Q5_K_M quant eval19,345346/6yes908.4 s
Robot-Filled Maze ShooterAgentWorld-35BB70 Q5_K_M quant eval24,135407/7yes1,105.0 s
Road HopperAgentWorld-35BB70 Q6_K quant eval run 115,227346/6yes971.3 s
Robot-Filled Maze ShooterAgentWorld-35BB70 Q6_K quant eval run 123,030407/7yes1,816.9 s
Road HopperAgentWorld-35BB70 Q6_K quant eval run 227,993356/6yes1,598.6 s
Robot-Filled Maze ShooterAgentWorld-35BB70 Q6_K quant eval run 221,541407/7yes1,292.4 s

Raw B70 report: CROSS_MODEL_B70.md · CROSS_MODEL_LIVE_B70.md · ORNITH_LIVE_EVALS.md · NEX2_LIVE_EVALS.md · SIQ_LIVE_EVALS.md · QUEST_LIVE_EVALS.md · LAGUNA_XS21_LIVE_EVALS.md.

Downloads

results_throughput.csv · results_throughput_aeon.csv · results_throughput_aeon_ornith.csv · results_throughput_qwen36_35b.csv · results_throughput_qwen36_35b_q5_turbo.csv · results_throughput_agentworld_spark.csv · results_throughput_qwen3_coder_30b_a3b_instruct_bf16.csv · results_throughput_nemotron3_super.csv · results_throughput_nemotron_labs_3_puzzle_75b_a9b_nvfp4.csv · results_throughput_agentworld.csv · results_throughput_gemma.csv · results_throughput_gemma4_nvfp4_current.csv · results_throughput_qwen36_35b_nvfp4.csv · results_throughput_agents_a1.csv

UNDER RAPID CONSTRUCTION - WILL CHANGE
why

Local agents need public test loops

Working on local agents matters because these tests are the first step toward an R&D base-model trust score: a practical signal for choosing which architectures and models deserve deeper research and development.

Agentic Arcade is a playable hello world for that work. Anyone can open the games, review the artifacts, and compare model behavior. The mission is maximal distribution of intelligence to all people.