Northern Tech Circuit Start a project

Single-slot PCIe accelerator · in development

Thor’s

Built on the NVIDIA Jetson AGX Thor module

A Jetson AGX Thor module on a PCIe card that fits one slot. 128 GB resident on the accelerator, at 130 W.

PCIe Gen5 x8 · 1 SLOT
Unified memory

128GB

Board power

130W

Form factor

1slot

MSRP

$6,000

Against the alternatives

RTX PRO 4000 DGX Spark Thor’s Hammer RTX PRO 5000 RTX PRO 6000
Memory 24 GB 128 GB 128 GB 48 GB 96 GB
Bandwidth 432 GB/s 273 GB/s 273 GB/s 1.34 TB/s 1.8 TB/s
AI compute, FP4 770 TFLOPS 1,000 TFLOPS 2,070 TFLOPS 2,300 TFLOPS 4,000 TFLOPS
Power 140 W 240 W 130 W 300 W 600 W
Footprint 1 slot a desk 1 slot 2 slots 2 slots
Price $10,2003 × $3,400 $4,699 $6,000 $13,2002 × $6,600 $16,000
Tokens/sLaguna S 2.1, NVFP4 305 64 64 632 424
Context availableFP16 cache 264K 1Mfull window 1Mfull window 753K 753K
Cost per GB $142 $37 $47 $138 $167

Poolside’s Laguna S 2.1 at NVFP4 is 59 GB of weights and 4.2 GB read per token. Where that does not fit on one board the model is sharded, and the price is what those boards cost together. Tokens/s is single-stream decode — bandwidth ÷ bytes per token. Context is what the memory left over holds, capped at the model’s million-token window.

What it is for

Image & video generation

Diffusion is compute-bound, and 128 GB holds a whole ComfyUI graph at once — transformer, encoders, VAE, upscalers, nothing swapping.

Prefill & long context

Prompt processing is compute-bound too: RAG ingestion, code understanding, embeddings, reranking.

Vision & multimodal

Few weights per frame, tensor cores saturated, camera interfaces already on the module.

Async batch serving

At batch, throughput tracks compute rather than bandwidth. Queued work, not chat.

Many models resident

A dozen 7B models side by side, or Laguna at FP8 where a $16,000 RTX PRO 6000 needs two boards.

Rugged, remote & defense

130 W in one slot: sealed, on a vehicle, or air-gapped.

Not for one person waiting on one dense 70B. That is a latency problem — buy the RTX PRO 6000.

Reading the table

Which model
Decode speed is bandwidth over bytes read. A dense 70B reads all 35 GB of itself per token. Laguna S 2.1 reads 4.2 GB, and its entire million-token context fits beside it on one card with 43 GB spare — a slow worker that never forgets.
Against a Spark
Same memory, same bandwidth, $1,300 less. Buy one if it can sit on a desk. Ours is twice the FP4 compute at half the power, in a slot.
Why the gap holds
GDDR7 took the RTX PRO 6000 from $8,565 to $16,000 in seventeen months. LPDDR5X comes out of the mobile supply chain, not the graphics one. That is $47 a gigabyte against $167.
The limit
PCIe Gen5 x8 and no NVLink, so splitting a model across cards would be slow. We sized one card not to have to.

Compute

Module
NVIDIA Jetson AGX Thor T5000
GPU
NVIDIA Blackwell — 2,560 CUDA cores, 96 fifth-generation Tensor Cores
AI performance
2,070 TFLOPS FP4 (sparse)
CPU
14-core Arm Neoverse-V3AE, 1 MB L2 per core, 16 MB L3

Memory

Capacity
128 GB LPDDR5X, unified
Bandwidth
273 GB/s across a 256-bit bus
Storage
On-module NVMe, plus an M.2 site on the carrier

Power & thermal

Envelope
130 W board power
Supply
75 W from the slot, remainder over a single 8-pin aux
Cooling
Single-width fin stack, blower-fed front to back

Mechanical & host

Form factor
Single slot, full height, 20.32 mm slot pitch
Length
Half-length card — clears the drive cage in a standard tower
Host interface
PCIe Gen5 x8
Bracket I/O
Display out and a network port; the rest of the module's I/O breaks out on-board
Software
NVIDIA JetPack — CUDA, TensorRT, the usual container stack

In development. Module figures are NVIDIA’s published Jetson AGX Thor T5000 numbers; card-level power, cooling and I/O are ours and move as the board does, as does the price.

Reply within one business day · NDA on request

Need a Hammer?