Aircade retro glove controller header
← All models

NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4

Nemotron Puzzle 75B model image

NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4

credit: NVIDIA
input tok/s114.4
decode tok/s319.8
per-agent tok/s10.3
agent Pareto32 · NVIDIA GB10

NVFP4 + FP8 ModelOpt safetensors · vLLM 0.23.0, 8k context, MTP n=1, max_num_seqs=32 · context-calibrated team of 1 · 20-min timed production loop  ·  Download on Hugging Face ↗.

model dossier

Model sheet

Playable artifacts and benchmark context for this model page.

Road Hopper
1 builds
Robot-Filled Maze Shooter
1 builds
Ribbit Rush
0 builds
Builds
2 recorded outputs
Method
team 1 / 20-min build
Runtime
vLLM 0.23.0, 8k context, MTP n=1, max_num_seqs=32
Quant
NVFP4 + FP8 ModelOpt safetensors
Bench
32 agents / 319.8 decode tok/sec
benchmark sheet

Throughput profile

input tok/s114.4
decode tok/s319.8
per-agent tok/s10.3
agent Pareto32 · NVIDIA GB10
aggregate decode tok/secper-agent tok/sechighlight = selected fleet
1 agents
24.9 dec24.9/agent · 8.8 in
2 agents
41.0 dec20.8/agent · 14.5 in
4 agents
72.1 dec18.3/agent · 25.5 in
8 agents
128.7 dec16.3/agent · 45.6 in
16 agents
202.9 dec13.0/agent · 72.3 in
24 agents
260.6 dec11.1/agent · 93.1 in
32 agents
319.8 dec10.3/agent · 114.4 in

Highlighted row is the selected build fleet / Pareto knee.

Raw sweep table
agentsprompt_t/sagg_gen_t/sper_agent_gen_t/swall_s
18.824.924.97.72
214.541.020.89.37
425.572.118.310.65
845.6128.716.311.94
1672.3202.913.015.14
2493.1260.611.117.68
32114.4319.810.319.21
production sheet

Teamwork build

No fan-out: a 1-agent team collaborates on ONE game for a 20-minute timed loop — a lead assigns focus areas, workers improve in parallel, a merger integrates the best, validated each round so it never regresses. Team size comes from the context/teamwork Pareto (single-agent footprint × 1.10 → agents that fit the window). Methodology ↗.

GameScoreSeed → FinalAssertionsRoundsSeed → Final sizeBuild time
Road Hopper61.061.0 → 61.06/637913 → 13 KB20 min
Robot-Filled Maze Shooter71.071.0 → 71.07/727114 → 14 KB20 min

Generated game outputs

Versions are listed first for selection. Embedded outputs remain below for direct review.

Road Hopper 1 versions

road_hopper_nemotron_labs_3_puzzle_75b_a9b_nvfp4_teamwork · Road Hopper · team buildfullscreen ↗ · compare
Road Hopper team build Nemotron Puzzle 75B road_hopper_nemotron_labs_3_puzzle_75b_a9b_nvfp4_teamworkfullscreen ↗

Robot-Filled Maze Shooter 1 versions

maze_nemotron_labs_3_puzzle_75b_a9b_nvfp4_teamwork · Robot-Filled Maze Shooter · team buildfullscreen ↗ · compare
Robot-Filled Maze Shooter team build Nemotron Puzzle 75B maze_nemotron_labs_3_puzzle_75b_a9b_nvfp4_teamworkfullscreen ↗
UNDER RAPID CONSTRUCTION - WILL CHANGE
why

Local agents need public test loops

Working on local agents matters because these tests are the first step toward an R&D base-model trust score: a practical signal for choosing which architectures and models deserve deeper research and development.

Agentic Arcade is a playable hello world for that work. Anyone can open the games, review the artifacts, and compare model behavior. The mission is maximal distribution of intelligence to all people.