Aircade retro glove controller header
← All models

Qwen3-Coder-30B-A3B-Instruct BF16

Qwen3 Coder 30B model image

Qwen3-Coder-30B-A3B-Instruct BF16

credit: Qwen
input tok/s87.0
decode tok/s344.2
per-agent tok/s10.8
agent Pareto32 · NVIDIA GB10

BF16 safetensors · vLLM BF16, 131k context for game builds · context-calibrated team of 6 · 20-min timed production loop  ·  Download on Hugging Face ↗.

model dossier

Model sheet

Playable artifacts and benchmark context for this model page.

Road Hopper
1 builds
Robot-Filled Maze Shooter
1 builds
Ribbit Rush
0 builds
Builds
2 recorded outputs
Method
team 6 / 20-min build
Runtime
vLLM BF16, 131k context for game builds
Quant
BF16 safetensors
Bench
32 agents / 344.2 decode tok/sec
benchmark sheet

Throughput profile

input tok/s87.0
decode tok/s344.2
per-agent tok/s10.8
agent Pareto32 · NVIDIA GB10
aggregate decode tok/secper-agent tok/sechighlight = selected fleet
1 agents
30.3 dec30.3/agent · 7.6 in
2 agents
52.6 dec26.3/agent · 13.2 in
4 agents
80.3 dec20.1/agent · 20.1 in
8 agents
122.8 dec15.4/agent · 30.7 in
16 agents
211.2 dec13.2/agent · 53.1 in
32 agents
344.2 dec10.8/agent · 87.0 in

Highlighted row is the selected build fleet / Pareto knee.

Raw sweep table
agentsprompt_t/sagg_gen_t/sper_agent_gen_t/swall_s
17.630.330.38.46
213.252.626.39.73
420.180.320.112.75
830.7122.815.416.67
1653.1211.213.219.39
3287.0344.210.823.80

Generated game outputs

Versions are listed first for selection. Embedded outputs remain below for direct review.

Road Hopper 1 versions

road_hopper_qwen3_coder_30b_bf16_teamwork · Road Hopper · team buildfullscreen ↗ · compare
Road Hopper team build Qwen3 Coder 30B road_hopper_qwen3_coder_30b_bf16_teamworkfullscreen ↗

Robot-Filled Maze Shooter 1 versions

maze_qwen3_coder_30b_bf16_teamwork · Robot-Filled Maze Shooter · team buildfullscreen ↗ · compare
Robot-Filled Maze Shooter team build Qwen3 Coder 30B maze_qwen3_coder_30b_bf16_teamworkfullscreen ↗
UNDER RAPID CONSTRUCTION - WILL CHANGE
why

Local agents need public test loops

Working on local agents matters because these tests are the first step toward an R&D base-model trust score: a practical signal for choosing which architectures and models deserve deeper research and development.

Agentic Arcade is a playable hello world for that work. Anyone can open the games, review the artifacts, and compare model behavior. The mission is maximal distribution of intelligence to all people.