Aircade retro glove controller header
← All models

NVIDIA Nemotron-3-Super-120B-A12B-NVFP4

Nemotron 3 Super model image

NVIDIA Nemotron-3-Super-120B-A12B-NVFP4

credit: NVIDIA
input tok/s4.9
decode tok/s13.9
per-agent tok/s13.9
agent Pareto1 · NVIDIA GB10

NVFP4 safetensors · vLLM 0.23.0, 8k context, max_num_seqs=1 · context-calibrated team of 1 · 20-min timed production loop  ·  Download on Hugging Face ↗.

model dossier

Model sheet

Playable artifacts and benchmark context for this model page.

Road Hopper
1 builds
Robot-Filled Maze Shooter
1 builds
Ribbit Rush
0 builds
Builds
2 recorded outputs
Method
team 1 / 20-min build
Runtime
vLLM 0.23.0, 8k context, max_num_seqs=1
Quant
NVFP4 safetensors
Bench
1 agents / 13.9 decode tok/sec
benchmark sheet

Throughput profile

input tok/s4.9
decode tok/s13.9
per-agent tok/s13.9
agent Pareto1 · NVIDIA GB10
aggregate decode tok/secper-agent tok/sechighlight = selected fleet
1 agents
13.9 dec13.9/agent · 4.9 in

Highlighted row is the selected build fleet / Pareto knee.

Raw sweep table
agentsprompt_t/sagg_gen_t/sper_agent_gen_t/swall_s
14.913.913.913.80
production sheet

Teamwork build

No fan-out: a 1-agent team collaborates on ONE game for a 20-minute timed loop — a lead assigns focus areas, workers improve in parallel, a merger integrates the best, validated each round so it never regresses. Team size comes from the context/teamwork Pareto (single-agent footprint × 1.10 → agents that fit the window). Methodology ↗.

GameScoreSeed → FinalAssertionsRoundsSeed → Final sizeBuild time
Road Hopper61.061.0 → 61.06/620014 → 14 KB20 min
Robot-Filled Maze Shooter71.071.0 → 71.07/723911 → 11 KB20 min

Generated game outputs

Versions are listed first for selection. Embedded outputs remain below for direct review.

Road Hopper 1 versions

road_hopper_nemotron3_super_teamwork · Road Hopper · team buildfullscreen ↗ · compare
Road Hopper team build Nemotron 3 Super road_hopper_nemotron3_super_teamworkfullscreen ↗

Robot-Filled Maze Shooter 1 versions

maze_nemotron3_super_teamwork · Robot-Filled Maze Shooter · team buildfullscreen ↗ · compare
Robot-Filled Maze Shooter team build Nemotron 3 Super maze_nemotron3_super_teamworkfullscreen ↗
UNDER RAPID CONSTRUCTION - WILL CHANGE
why

Local agents need public test loops

Working on local agents matters because these tests are the first step toward an R&D base-model trust score: a practical signal for choosing which architectures and models deserve deeper research and development.

Agentic Arcade is a playable hello world for that work. Anyone can open the games, review the artifacts, and compare model behavior. The mission is maximal distribution of intelligence to all people.