Skip to content
HN On Hacker News ↗

GitHub - nestrilabs/virtio-nvgpu: [Experimental] A virtio device for near-native NVIDIA GPU access in KVM virtual machines.

▲ 155 points • 65 comments • by WanjohiRyan • 2w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is AI.

97 %

AI likelihood · overall

AI
0% human-written 100% AI-generated
SEGMENTS · HUMAN 0 of 1
SEGMENTS · AI 1 of 1
WORD COUNT 1,593
PEAK AI % 97% · §1
Analyzed
Sep 24
backend: pangram/v3.3
Segments scanned
1 windows
avg 1593 words each
Distribution
0 / 100%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 1,593 words · 1 segments analyzed

Human AI-generated
§1 AI · 97%

Near-native NVIDIA GPU access inside a KVM guest. A guest renders within 2% of the machine it is running on, and costs the same CPU. virtio-nvgpu forwards NVIDIA kernel driver ioctls between a Linux guest and the host at the driver ABI level, bypassing API-level translation entirely. The guest runs NVIDIA's own user-mode drivers, unmodified — the same libraries, the same Vulkan and NVENC, talking to the same card. The target is headless streaming: a compositor inside the VM renders, composites and encodes frames on the GPU, then sends compressed video out. The VM has no monitor, and the host keeps the card. Where it stands It works, and it has been measured. A Wayland client presents inside a guest, the capture layer encodes on the game's own device, and the H.264 comes out the other side — 618 frames that ffmpeg decodes without an error. Measured on an RTX 3060 (driver 595.99.02), guest against the same host, bare metal, with an identical headless Vulkan load: what the host takes for one frame guest frame time 39 ms −0.4% faster than bare metal, within noise 9.9 ms −0.7% 2.0 ms +1.7% 0.5 ms +7.1% a wake costs ~0.02 ms, and the frame is half of one 0.05 ms +40.8% Above about 2 ms a frame — which is every frame a game draws — a guest is within 2% of bare metal. Below that, the cost of waiting for the GPU starts to show against a frame that barely exists. CPU is the other half of it, because a shared GPU is only worth sharing if the guests are cheap. Unpaced at ~100 fps for 12 s, one guest: CPU used host, bare metal 0.40 s guest 0.37 s A guest costs what the host costs. Nothing is spent on forwarding in a render loop, because nothing is forwarded: NVIDIA's user-mode driver submits through memory it has mapped, and that memory is the host's. Over 813,691 frames the backend served 13,792 messages — one crossing per 59 frames, nearly all of it device setup. Full method, raw runs and the things these numbers do not support: BENCHMARKS.md. Several guests on one card Four guests on one RTX 3060, the same load in each: 25.84, 26.49, 25.57, 25.79 fps — 103.7 together, against 102.9 for a single guest — with p50 frame times of 39.165, 39.164, 39.168 and 39.165 ms. The total does not move as guests are added, and the split is even to four decimal places. All four render correctly at the same time, and four of them encode H.264 at once, each paced at exactly 60 Hz, with no NVENC session limit reached. Four is what was run, not a limit found. Driver versions Measured on 595.99.02; an A2000 on 615.71.09 renders but is not benchmarked. ABI profiles shipped: 535.129.03, 580.178.04, 595.71.05, matched by range, with anything older than the first refused. Details below. What is known to work a guest enumerates the card — nvidia-smi reports real power and memory, and the deviceUUID is the host's Vulkan renders: vulkaninfo exits 0, offscreen draws are pixel-correct a Wayland client presents through a compositor in the guest NVENC through Vulkan Video, encoding on the client's own device imported buffers are the host's memory, mapped through a shared window What is not done more than four guests, or guests doing anything heavier than vkcube at 720p. Four share the card evenly; eight has not been tried. two cards, two driver versions. RTX 3060 / 595.99.02 is where the numbers come from; an RTX A2000 / 615.71.09 has rendered but is not benchmarked. CUDA is forwarded but untested beyond enumeration; the jailer, per-version driver shares and the multi-tenant envelope are unbuilt. Repository layout Four components, three license zones. The split is deliberate: the guest half must be GPL to touch kernel symbols, the host half should be permissive so that other people can build on it, and the definitions both halves share must be includable from both. directory license what it is driver/ GPL-2.0 Guest kernel module. Registers /dev/nvidia*, forwards ioctl and mmap over the virtqueue. Deliberately not ABI-aware. device/ Apache-2.0 The virtio device, as a Rust crate with no VMM in its dependency list. Every VMM concern is a trait. isolate/ Apache-2.0 A design note, not code yet. The sandboxed per-guest helper that will hold the real device FDs. Today the backend holds them itself, in the VMM's own process. gen/ — Generated ABI tables. Checked in and reproducible. protocol/ BSD-3-Clause OR GPL-2.0+ Wire format and ABI definitions shared by both halves. Dual licensed so the GPL driver and the Apache crate can include the same headers. The layout follows chromeos/virtio-media, which solves the same problem — one repository holding a GPL guest driver beside a permissively licensed, VMM-agnostic device crate. Using it from a VMM device/ depends on no virtual machine monitor. A VMM adopts the device by implementing a small set of traits — descriptor chains as Read/Write, an event queue, guest memory mapping, host memory mapping — and gets the whole device without patching the crate. Optional capabilities degrade rather than fail to build, so a VMM can adopt it before supporting every feature. Buffer and window bookkeeping lives in device/. The VMM supplies raw map and unmap and nothing more. One thing that will not be a trait: the isolate. The intended design runs one sandboxed helper process per guest process, so adopting it eventually means inheriting a process model, not just a library dependency. That helper is not written — the backend holds the device descriptors itself today — and isolate/ is where the design lives until it is. Why The streaming pipeline we want Guest VM (headless, no physical display) ────────────────────────────────────────── Game / application │ Vulkan or OpenGL ▼ Wayland compositor (guest-side) │ composites all windows │ CUDA zero-copy import of composed frame ▼ NVENC hardware encoder (guest-side) │ H.264 / H.265 bitstream (~100 KB per frame) ▼ Stream to remote client The entire render → composite → encode pipeline runs on the GPU, inside the guest. Only the compressed bitstream leaves. This requires the guest to have real, driver-level access to GPU resources: buffer handles, fences, CUDA device pointers, NVENC sessions. Why existing approaches fall short virtio-gpu + Venus (API-level translation). Venus serializes every Vulkan or OpenGL call in the guest, transports it over virtio, and replays it host-side. Three problems for this use case: Latency compounds on draw-call-heavy workloads. Games issue 1,000–5,000 draw calls per frame plus binds, descriptor updates and render pass transitions, each serialized and replayed individually. At 60 fps the frame budget is 16.6 ms; 1–3 ms of serialization is 6–18% gone before any GPU work. CPU overhead is significant. Serialization, transport and replay burn host CPU the application needs. Where compute is billed and finite, that waste is the product. Guest-side encoding is not viable. GPU buffers are owned by the host. The guest compositor cannot see or import them, so there is no practical path to a CUdeviceptr in the guest pointing at a Venus-managed buffer — which means no NVENC without a full CPU readback and copy. DRM native context (Intel / AMD). The guest runs the real Mesa driver, builds command buffers locally, and only submissions cross the boundary. Guest-side buffer ownership and encoding work correctly. This does not exist for NVIDIA. VFIO passthrough. Native performance and a complete driver stack in the guest, but it dedicates the whole GPU to one VM. In multi-tenant environments that is often not an option. What virtio-nvgpu does differently Translation happens at the kernel driver level (ioctls to /dev/nvidia*), not the graphics API level. The guest runs NVIDIA's real user-mode libraries, which build GPU command buffers locally in the guest — individual draw calls are never serialized: Venus virtio-nvgpu ────────────── ────────────────────── Per draw call: serialize + local function call transport + (no VM exit) deserialize + replay Per frame ~2,000 messages ~5–20 messages boundary (one per API call) (queue submits + allocs) crossings GPU command generated on HOST generated in GUEST buffers after replay by NVIDIA's own compiler CPU overhead serialization + near zero for rendering deserialization (only ioctl forwarding) Guest buffer HOST owns buffers GUEST owns buffers ownership compositor can't compositor has full track them visibility and control Guest NVENC not viable works (real CUDA interop) How it works Guest kernel driver. Registers /dev/nvidiactl, /dev/nvidia0…N and /dev/nvidia-uvm. On ioctl() it serializes the request onto a control virtqueue. On mmap() it maps the appropriate shared-memory region into the calling process with the correct caching attributes. It copies raw bytes and makes no ABI decisions. Device crate. Receives requests, maps guest handles to host device file descriptors, performs ABI-aware translation of ioctl parameters — rewriting embedded pointers and file descriptors — and issues them against the host's devices. Buffer and window bookkeeping lives here. Events. A second virtqueue runs the other way. The host watches each descriptor it has opened and says when one becomes readable, which is how a guest waiting for the GPU is woken. Without it the guest cannot wait at all — it polls a descriptor the kernel reports as permanently ready, and spins. Isolate — not built yet. The plan is a sandboxed helper per guest process, holding the real device FDs and issuing the ioctl(2) calls unprivileged. Today the backend does that itself, inside the VMM's process. isolate/ holds the design and no code. ┌─ Guest ─────────────────────────────────────────────────┐ │ Application → NVIDIA Vulkan / GL / CUDA │ │ │ ioctl(/dev/nvidia*) │ │ driver/ (GPL) ▼ │