Aircade retro glove controller header
← All models

StepFun Step-3.7-Flash

StepFun Step-3.7 model image

StepFun Step-3.7-Flash

credit: StepFun
input tok/s94.2
decode tok/s33.6
per-agent tok/s11.3
agent Pareto3 · NVIDIA GB10

UD-IQ4_NL GGUF · llama.cpp + Q8 MTP draft · context-calibrated team of 3 · 20-min timed production loop.

model dossier

Model sheet

Playable artifacts and benchmark context for this model page.

Road Hopper
1 builds
Robot-Filled Maze Shooter
1 builds
Ribbit Rush
0 builds
Builds
2 recorded outputs
Method
team 3 / 20-min build
Runtime
llama.cpp + Q8 MTP draft
Quant
UD-IQ4_NL GGUF
Bench
3 agents / 33.6 decode tok/sec
benchmark sheet

Throughput profile

input tok/s94.2
decode tok/s33.6
per-agent tok/s11.3
agent Pareto3 · NVIDIA GB10
aggregate decode tok/secper-agent tok/sechighlight = selected fleet
1 agents
23.0 dec23.0/agent · 81.9 in
2 agents
29.3 dec14.8/agent · 94.5 in
3 agents
33.6 dec11.3/agent · 94.2 in

Highlighted row is the selected build fleet / Pareto knee.

Raw sweep table
agentsprompt_t/sagg_gen_t/sper_agent_gen_t/sttft_p50_sttft_p99_swall_s
181.923.023.08.478.4711.1
294.529.314.89.5614.7117.5
394.233.611.320.1422.1722.8
production sheet

Teamwork build

No fan-out: a 3-agent team collaborates on ONE game for a 20-minute timed loop — a lead assigns focus areas, workers improve in parallel, a merger integrates the best, validated each round so it never regresses. Team size comes from the context/teamwork Pareto (single-agent footprint × 1.10 → agents that fit the window). Methodology ↗.

GameScoreSeed → FinalAssertionsRoundsSeed → Final sizeBuild time
Road Hopper61.061.0 → 61.06/6117 → 21 KB54 min
Robot-Filled Maze Shooter71.071.0 → 71.07/7129 → 35 KB56 min

Generated game outputs

Versions are listed first for selection. Embedded outputs remain below for direct review.

Road Hopper 1 versions

road_hopper_step37_teamwork · Road Hopper · team buildfullscreen ↗ · compare
Road Hopper team build StepFun Step-3.7 road_hopper_step37_teamworkfullscreen ↗

Robot-Filled Maze Shooter 1 versions

maze_step37_teamwork · Robot-Filled Maze Shooter · team buildfullscreen ↗ · compare
Robot-Filled Maze Shooter team build StepFun Step-3.7 maze_step37_teamworkfullscreen ↗
UNDER RAPID CONSTRUCTION - WILL CHANGE
why

Local agents need public test loops

Working on local agents matters because these tests are the first step toward an R&D base-model trust score: a practical signal for choosing which architectures and models deserve deeper research and development.

Agentic Arcade is a playable hello world for that work. Anyone can open the games, review the artifacts, and compare model behavior. The mission is maximal distribution of intelligence to all people.