Every thUMBox runs on hardware you own — fixed cost, your data on your box. AgentBOX runs deterministic n8n workflows driving small local models for the everyday work, and reaches the cloud only when a task needs higher-level reasoning. Pick the hardware that fits the job; the same membership works across all of it.
You own the economics and the data.
AgentBOX is the flagship appliance — a local agent, its memory, a desktop control surface, and the MailBOX email pipeline, co-resident on one box. Power it on and it works. It runs on the Jetson Orin Nano Super, 67 TOPS in a 15-watt envelope.
On the 8GB Jetson, AgentBOX hosts one pipeline at a time. Need two or more pipelines running together? That's multi-pipeline hosting — and it lives on qualified hardware (T3+), below.
Your membership says how many pipelines you're entitled to. Your hardware says how many host at once. AgentBOX runs deterministic n8n paths on small local models and reaches the cloud only for higher-level reasoning, so it rides light on the box. Tier tells you capability; hardware tells you concurrency.
Launch is AgentBOX on the T2 Jetson. Below it, run the same platform on hardware you already own or a low-cost Pi; above it, the ladder is NVIDIA-native the whole way up — same software, more headroom, more pipelines at once, and more reasoning running on-box. Tap any tier for the detail.
The sovereignty path — and the cheapest. Install the full thUMBox stack — local model server, vector store, n8n workflow engine, and dashboard — on a machine you already own: a desktop, mini PC, laptop, or Apple-silicon / NVIDIA box. One command brings the containers up; the model downloads once and stays local. It runs on the free Community plan, no card required. Capability scales with the machine — 8 GB runs a single pipeline comfortably, and 16 GB and up hosts more than one at once and brings more of the reasoning on-box. Your metal, your models, your tokens — nothing phones home.
The cheapest path to a dedicated, always-on node. A Raspberry Pi 5 — or any equivalent single-board computer — runs thUMBox on ARM64 at a few watts, around the clock, for the price of a nice dinner. Inference is CPU-paced and sized to small models, so this is the floor of the ladder: think a private on-device agent and light email triage rather than heavy drafting. What you trade in speed you get back in cost and footprint — a tiny box on a shelf that's wholly yours, with no per-token meter and nothing leaving your network.
Available now. The Jetson Orin Nano Super fits 67 TOPS into a 15-watt envelope — purpose-built silicon that runs a 4B-class model on-device at ~16 tokens/sec. It ships as AgentBOX: the agent loop, gBrain memory, a desktop control surface, and the MailBOX email pipeline, co-resident on one box. Deterministic n8n workflows drive small local models for the everyday work, and the box reaches the cloud only when a task needs higher-level reasoning — so the 8 GB envelope goes to the runtime and your local data; pipelines run one at a time, with a Plus membership hot-swapping between your entitled pipelines. Every AgentBOX includes a year of Base membership.
The natural step up, on the same NVIDIA stack as your launch box. The Jetson AGX Orin brings 64 GB and up to 275 TOPS — the headroom to host several pipelines concurrently, run larger models, and keep more of the reasoning on-box instead of reaching the cloud, with no change to the software you already know. AgentBOX runs here exactly as it does on the launch box, with far more room to grow. This is the primary T3, and the box receptionBOX re-platforms onto.
The same T3-class headroom on a non-Jetson Linux box — chosen when you want a standard x86-64 or ARM64 server rather than NVIDIA's Jetson form factor. It delivers the same multi-pipeline capability and larger models, runs AgentBOX unchanged, and fits the racks and machines a team already standardizes on. The tiers differ by form factor and headroom, not capability; the exact SKU is finalizing.
On the roadmap. NVIDIA's DGX Spark puts a Grace Blackwell GB10 superchip and 128 GB of unified memory in a desktop box — roughly a petaFLOP of FP4 compute, enough to run models up to 200B parameters entirely on-device (or larger with two linked over ConnectX). With this much silicon the cloud reach-out all but disappears — the reasoning runs on-box first. The most capable single desktop in the ladder, running the same AgentBOX software you already know — for a power user or small team that wants data-center-class reasoning on hardware they own.
On the roadmap. A multi-GPU workstation with 64 GB and up runs 14–30B models and orchestrates several agents at once — enough for a small team to share a single box. With this much silicon, most of the reasoning runs on-box and the cloud reach-out becomes the exception. Built for the moment one operator's box becomes a team's shared brain.
On the roadmap. Rack- or fleet-scale hardware with 128 GB and up runs 30–70B models under fleet-coordinated policy — intelligence that spans an organization while still living on metal you own. Access control is enforced fleet-wide. The top of the ladder: data-center-class capability with zero data-center dependency.