Thor’s
Built on the NVIDIA Jetson AGX Thor module
A Jetson AGX Thor module on a PCIe card that fits one slot. 128 GB resident on the accelerator, at 130 W.
128GB
130W
1slot
$6,000
Against the alternatives
| RTX PRO 4000 | DGX Spark | Thor’s Hammer | RTX PRO 5000 | RTX PRO 6000 | |
|---|---|---|---|---|---|
| Memory | 24 GB | 128 GB | 128 GB | 48 GB | 96 GB |
| Bandwidth | 432 GB/s | 273 GB/s | 273 GB/s | 1.34 TB/s | 1.8 TB/s |
| AI compute, FP4 | 770 TFLOPS | 1,000 TFLOPS | 2,070 TFLOPS | 2,300 TFLOPS | 4,000 TFLOPS |
| Power | 140 W | 240 W | 130 W | 300 W | 600 W |
| Footprint | 1 slot | a desk | 1 slot | 2 slots | 2 slots |
| Price | $10,2003 × $3,400 | $4,699 | $6,000 | $13,2002 × $6,600 | $16,000 |
| Tokens/sLaguna S 2.1, NVFP4 | 305 | 64 | 64 | 632 | 424 |
| Context availableFP16 cache | 264K | 1Mfull window | 1Mfull window | 753K | 753K |
| Cost per GB | $142 | $37 | $47 | $138 | $167 |
Poolside’s Laguna S 2.1 at NVFP4 is 59 GB of weights and 4.2 GB read per token. Where that does not fit on one board the model is sharded, and the price is what those boards cost together. Tokens/s is single-stream decode — bandwidth ÷ bytes per token. Context is what the memory left over holds, capped at the model’s million-token window.
What it is for
Image & video generation
Diffusion is compute-bound, and 128 GB holds a whole ComfyUI graph at once — transformer, encoders, VAE, upscalers, nothing swapping.
Prefill & long context
Prompt processing is compute-bound too: RAG ingestion, code understanding, embeddings, reranking.
Vision & multimodal
Few weights per frame, tensor cores saturated, camera interfaces already on the module.
Async batch serving
At batch, throughput tracks compute rather than bandwidth. Queued work, not chat.
Many models resident
A dozen 7B models side by side, or Laguna at FP8 where a $16,000 RTX PRO 6000 needs two boards.
Rugged, remote & defense
130 W in one slot: sealed, on a vehicle, or air-gapped.
Not for one person waiting on one dense 70B. That is a latency problem — buy the RTX PRO 6000.
Reading the table
- Which model
- Decode speed is bandwidth over bytes read. A dense 70B reads all 35 GB of itself per token. Laguna S 2.1 reads 4.2 GB, and its entire million-token context fits beside it on one card with 43 GB spare — a slow worker that never forgets.
- Against a Spark
- Same memory, same bandwidth, $1,300 less. Buy one if it can sit on a desk. Ours is twice the FP4 compute at half the power, in a slot.
- Why the gap holds
- GDDR7 took the RTX PRO 6000 from $8,565 to $16,000 in seventeen months. LPDDR5X comes out of the mobile supply chain, not the graphics one. That is $47 a gigabyte against $167.
- The limit
- PCIe Gen5 x8 and no NVLink, so splitting a model across cards would be slow. We sized one card not to have to.
Compute
- Module
- NVIDIA Jetson AGX Thor T5000
- GPU
- NVIDIA Blackwell — 2,560 CUDA cores, 96 fifth-generation Tensor Cores
- AI performance
- 2,070 TFLOPS FP4 (sparse)
- CPU
- 14-core Arm Neoverse-V3AE, 1 MB L2 per core, 16 MB L3
Memory
- Capacity
- 128 GB LPDDR5X, unified
- Bandwidth
- 273 GB/s across a 256-bit bus
- Storage
- On-module NVMe, plus an M.2 site on the carrier
Power & thermal
- Envelope
- 130 W board power
- Supply
- 75 W from the slot, remainder over a single 8-pin aux
- Cooling
- Single-width fin stack, blower-fed front to back
Mechanical & host
- Form factor
- Single slot, full height, 20.32 mm slot pitch
- Length
- Half-length card — clears the drive cage in a standard tower
- Host interface
- PCIe Gen5 x8
- Bracket I/O
- Display out and a network port; the rest of the module's I/O breaks out on-board
- Software
- NVIDIA JetPack — CUDA, TensorRT, the usual container stack
In development. Module figures are NVIDIA’s published Jetson AGX Thor T5000 numbers; card-level power, cooling and I/O are ours and move as the board does, as does the price.