Aircade retro glove controller header
← All models

InternScience/Agents-A1 (35B MoE, agentic)

Agents-A1 model image

InternScience/Agents-A1 (35B MoE, agentic)

credit: InternScience
input tok/s92.2
decode tok/s348.9
per-agent tok/s10.9
agent Pareto32 · NVIDIA GB10

BF16 (unquantized) · vLLM 0.23.0, BF16, FlashInfer CUTLASS MoE · context-calibrated team of 32 · 20-min timed production loop  ·  Download on Hugging Face ↗.

model dossier

Model sheet

Playable artifacts and benchmark context for this model page.

Road Hopper
1 builds
Robot-Filled Maze Shooter
1 builds
Ribbit Rush
0 builds
Builds
2 recorded outputs
Method
team 32 / 20-min build
Runtime
vLLM 0.23.0, BF16, FlashInfer CUTLASS MoE
Quant
BF16 (unquantized)
Bench
32 agents / 348.9 decode tok/sec
benchmark sheet

Throughput profile

input tok/s92.2
decode tok/s348.9
per-agent tok/s10.9
agent Pareto32 · NVIDIA GB10
aggregate decode tok/secper-agent tok/sechighlight = selected fleet
1 agents
30.4 dec30.4/agent · 7.9 in
2 agents
55.7 dec27.8/agent · 14.6 in
4 agents
90.3 dec22.6/agent · 23.6 in
8 agents
137.9 dec17.2/agent · 36.1 in
16 agents
228.0 dec14.3/agent · 60.0 in
32 agents
348.9 dec10.9/agent · 92.2 in

Highlighted row is the selected build fleet / Pareto knee.

Raw sweep table
agentsprompt_t/sagg_gen_t/sper_agent_gen_t/swall_s
17.930.430.48.43
214.655.727.89.20
423.690.322.611.34
836.1137.917.214.86
1660.0228.014.317.97
3292.2348.910.923.48
production sheet

Teamwork build

No fan-out: a 32-agent team collaborates on ONE game for a 20-minute timed loop — a lead assigns focus areas, workers improve in parallel, a merger integrates the best, validated each round so it never regresses. Team size comes from the context/teamwork Pareto (single-agent footprint × 1.10 → agents that fit the window). Methodology ↗.

GameScoreSeed → FinalAssertionsRoundsSeed → Final sizeBuild time
Road Hopper61.061.0 → 61.06/6213 → 17 KB23 min
Robot-Filled Maze Shooter71.071.0 → 71.07/7214 → 14 KB27 min

Generated game outputs

Versions are listed first for selection. Embedded outputs remain below for direct review.

Road Hopper 1 versions

road_hopper_agents_a1_teamwork · Road Hopper · team buildfullscreen ↗ · compare
Road Hopper team build Agents-A1 road_hopper_agents_a1_teamworkfullscreen ↗

Robot-Filled Maze Shooter 1 versions

maze_agents_a1_teamwork · Robot-Filled Maze Shooter · team buildfullscreen ↗ · compare
Robot-Filled Maze Shooter team build Agents-A1 maze_agents_a1_teamworkfullscreen ↗
UNDER RAPID CONSTRUCTION - WILL CHANGE
why

Local agents need public test loops

Working on local agents matters because these tests are the first step toward an R&D base-model trust score: a practical signal for choosing which architectures and models deserve deeper research and development.

Agentic Arcade is a playable hello world for that work. Anyone can open the games, review the artifacts, and compare model behavior. The mission is maximal distribution of intelligence to all people.