Skip to content

Hardware Acceleration

An instance’s media container runs a software pipeline unless a launch asks for an accelerator. norsk-ctl knows two: an NVIDIA GPU (the nvidia profile) and a NETINT Quadra transcoder card (the quadra profile). The profile is a per-launch choice — --hardware on the CLI, the accelerator control in the launch form — layered onto the product template’s compose by the runner, so the same template runs on a GPU box, a Quadra box, and a CPU-only VM.

Three things follow from a profile, and this page covers each: how the host is detected, what a launch reserves, and how a product mandates (or forbids) an accelerator rather than leaving the operator to guess.

The daemon probes for accelerators on Linux only; macOS always reports none. norsk-ctl get hardware (and the status view) shows two fields:

  • profile — a single winner by precedence: nvidia beats quadra, and a host with neither is none. This is what the runner assumes when nothing else says otherwise.
  • accelerators — every accelerator the host actually has. A box with both an NVIDIA card and a Quadra reports nvidia as its profile but lists both here, which is what the launch form needs to offer a choice.

Each profile has one probe, and the same probe answers “is it here?” for a mandate and “what do I map?” at launch:

ProfileProbeWhy that probe
nvidiaThe device node /dev/nvidia0 existsCreated by the NVIDIA kernel driver; present-or-not
quadraAn NVMe controller under /sys/class/nvme whose PCI vendor is 0x1d82 (NETINT), with its n1 nodeThe Quadra is driven as an NVMe block device, /dev/nvmeXn1. The kernel assigns X afresh each boot, so the card cannot be found by a fixed path — only by asking sysfs who made it

The Quadra probe counts cards; every one it finds is mapped into the instance.

The runner adds a GPU reservation to the media service. Two wirings exist, because Docker has two ways of exposing an NVIDIA GPU, and the runner picks from docker info at launch:

  • Legacy — the host has the NVIDIA container runtime installed. The media service gets the classic deploy.resources.reservations request for all GPUs with the nvidia driver. This is what most existing hosts have, and it is the fallback whenever the probe cannot classify the host.
  • CDI — no NVIDIA runtime, but the engine advertises a nvidia.com/gpu device through the Container Device Interface. The legacy form fails there with could not select device driver "nvidia", so the runner injects nvidia.com/gpu=all as a device instead.

Detection is automatic. If a host is misclassified, the daemon’s NORSK_CTL_GPU_MODE environment variable forces legacy or cdi.

GPU sideload. Norsk media images are published slim: the CUDA-dependent libraries for NVENC/NVDEC encoding and GPU compose/resize ship as a separate gpu-video bundle rather than bloating every image on every CPU-only host. When a launch reserves the GPU, the runner reads the bundle manifest baked into the media image, fetches the matching bundle from Norsk’s published index (or a mirror — see NORSK_CTL_SIDELOAD_INDEX_URL in the environment reference), caches it under the daemon’s state directory, and mounts it read-only at /opt/norsk in the media container. The bundle is matched by content hash and a mismatch is refused rather than mounted: the CUDA libraries are ABI-coupled to the engine build, and a wrong bundle would crash the engine instead of failing cleanly. Products can declare further bundles (GPU inference, model weights) in their manifest; those follow the same path.

The Quadra is a PCIe transcoder that NETINT’s driver (libxcoder) reaches as an NVMe block device. Three things happen at launch, all on the media service:

  • Every NETINT NVMe node the probe found is mapped into the container at the same path (/dev/nvme1n1:/dev/nvme1n1, and so on for each card). The launch is refused with HARDWARE_NOT_FOUND if the probe finds none — mapping a guessed path that Docker cannot open would fail later and less clearly.
  • The host’s /dev/shm is mounted over media’s private one. libxcoder keeps the card’s resource table and lock files at a hard-coded /dev/shm and expects the host and every container using a card to share the one copy; that table is how sessions across processes coordinate load. This is the single exception to the private-shm rule, and --shm-size does not govern on this profile. Shared Memory states the trade in full.
  • The container user’s group becomes the host disk group. The /dev/nvme*n1 nodes are owned by root:disk, so the media process must carry that gid to open them. The suggested container user the daemon reports (suggestedContainerUser) is the daemon’s own uid with the disk gid whenever the resolved profile is quadra, falling back to the daemon’s gid if the host has no disk group. Studio shares the gid harmlessly.

A launch that states --hardware quadra decides the gid ahead of host detection: a Quadra mandate on a host whose profile winner is nvidia still gets the disk group.

Left to the operator, an accelerator is a guess: a product built around NVENC launches happily on a CPU-only host and fails at runtime as a media error nobody traces back to the missing reservation. A product template can instead declare what it needs in its manifest, under requirements.accelerator, in one of three modes:

DeclarationMeaning
{ mode: fixed, value: nvidia }A mandate. The operator cannot override it (asking for anything else is refused as HARDWARE_PINNED), and a host without the accelerator is refused as HARDWARE_UNAVAILABLE rather than launched unreserved. fixed with value none is how a CPU-only product forbids a reservation
{ mode: default, value: nvidia }A preference the product can live without. The operator may change it; a host that lacks it falls back to no reservation instead of refusing. An operator who explicitly asks for an accelerator the host lacks is still refused
{ mode: required }The product cannot know — the operator must choose at launch, and a launch with no choice is refused as HARDWARE_REQUIRED. Whatever is chosen must be present on the host
absentNothing is reserved unless the operator asks, and the host is not probed. The pre-mandate contract

The launch form reads the declaration against what the host carries. A fixed mandate is shown but not offered — including when this host cannot meet it, because a hidden mandate would turn into a refused launch with no visible cause. A default or required declaration becomes a choice that lists only none plus the accelerators actually present, so the form never offers an entry the daemon would refuse. A default on a host with no accelerator at all shows nothing: there is nothing to choose between.

Each refusal names its cause, so the fix follows from the message:

CodeCauseFix
HARDWARE_NOT_FOUNDquadra was selected but the probe found no NETINT NVMe controllerCheck ls /sys/class/nvme and each controller’s device/vendor; the card should show 0x1d82. A missing driver leaves the card unseen
HARDWARE_UNAVAILABLEThe product requires (or the operator asked for) an accelerator the host lacksLaunch on a host that has it, or, for a default declaration, pick none
HARDWARE_PINNEDThe operator asked for a different accelerator than the product’s fixed oneDrop the --hardware flag; the product decides
HARDWARE_REQUIREDThe product leaves the choice to the operator and none was madePass --hardware <profile> (or none)

A launch that passes these checks can still fail inside Docker. could not select device driver "nvidia" at compose time means the host is CDI-only and was classified as legacy; set NORSK_CTL_GPU_MODE=cdi on the daemon and relaunch.