CH1What the FPGA presents
The DHO's link runs at 5 GT/s with two lanes (the hardware page has why the device tree's four lanes do not apply), the MHO's at 5 GT/s with four. The FPGA enumerates as Xilinx's XDMA endpoint, 10ee:7022 on the DHO and 10ee:7024 on the MHO, with BAR0 holding the XDMA engine's own 64 KiB of registers and BAR1 the acquisition logic's registers (32 KiB on the DHO; the app maps 16 MiB on the MHO). Behind it are one host-to-card and one card-to-host DMA channel, both AXI-Stream.
Rigol's scope app uses exactly two of the driver's device nodes. It maps the register window once and then does every register access as a plain 32-bit load or store. It reads samples from the card-to-host stream with ordinary read() calls. It issues no ioctl, uses no interrupt and never uses the driver's ring mode. That small surface is what let us change everything underneath it.
CH2The driver
Rigol ships Xilinx's 2019.2 reference driver built for its 4.4 kernel, with two changes in the binary: polling is the default, and the poll loop spins with preemption off, without ever yielding, for up to 300 jiffies. Upstream's 2019.2 loop called schedule() while holding a spinlock, a sleep-in-atomic bug that Xilinx fixed for arm64 in February 2020 by making the lock a mutex. Porting the 2019 source to Linux 6.12 would mean rewriting its DMA mapping calls (the pci_map_sg API is gone since 5.18), access_ok, class_create and the VMA flag handling, all for code the app never calls.
NovaOS instead builds a 2025 release of Xilinx's driver against its 6.12 kernel, with one small patch that gives it the stock module's defaults and parameter names: polling on, card-to-host credit mode on, 10 s timeouts. A test reads the parameter defaults straight out of both module binaries and compares them. Credit mode is one place the newer code is stricter: the engine's credit count is a 10-bit field, the 2019 code could post up to 2048 descriptors for one read (about 8 MiB in 4 KiB pages), and the current driver splits a transfer at 1023.
Polling has a cost. At 2 µs/div we measured the acquisition engine at 8 % of a CPU and the driver's completion thread at 6 %. The endpoint does have a legacy interrupt line, but interrupt mode has never been run with this FPGA design, so NovaOS keeps polling.
CH3Asking for waveforms
The FPGA records into its own DDR3 and sends nothing until it is told to play stored frames. Getting a batch takes three steps over the register window: arm a record sequence (which frames to record into, then run), wait until the frame-ready flag (bit 30 of register 0x4080) is set, then send the play command. Only after that does the engine call read(), once, for every frame recorded. On our first live NovaOS boot the engine skipped the play step, and every read sat with zero descriptors completed until the 10 s timeout; the same three steps done by hand over the register window streamed a frame at once.
Each frame is a 16-byte header, whose last six bytes hold the trigger time as a 48-bit count of 800 ps ticks, followed by little-endian 16-bit sample codes with the channels interleaved sample by sample. The read has to ask for exactly the bytes the FPGA will send: at the 2 µs/div screen length of 7,000 points a frame is 14,016 bytes with one channel, 28,016 with two and 56,016 with four. A larger request times out. The driver ends a stream read early at the stream's end marker only when the device was opened with O_TRUNC, and neither Rigol's app nor ours opens it that way.
Deep memory comes out the same way, in windows. Each read returns at most 1,000,000 interleaved points (2 MB) with no frame header, so a 125 Mpts channel is 125 reads. Saves and SCPI downloads go window by window, and a deep record is never held in RAM whole.
CH4Where the bytes land, and what it costs
Rigol's app first looks for a DMA-buffer device that its device tree names but no driver provides, logs "USE CPU Memory copy", and reads into a 100 MiB page-aligned buffer. NovaOS's acquisition engine reads each batch straight into one slot of a four-slot ring in a shared-memory file. The driver pins those pages and the DMA writes into them, so the CPU never copies a sample; the UI maps the same file, as described in the stack note. A four-channel batch at 2 µs/div is about 2 MB, about seventy times a second.
The link is not the bottleneck. Play plus read takes 59 µs for a 12.5 KB frame at 2 µs/div, about 210 MB/s, where two Gen2 lanes carry about 1 GB/s after 8b/10b encoding. The limit is the FPGA's play path, which reads the whole window out of its memory at about 130 MSa/s whatever it sends: a frame compressed to 97 points still costs 48 µs, against 54 µs for all 6,250. The per-frame figures are on the FPGA page.
Two ways to hang an RK3399
A read across a PCIe link that is no longer up is never completed, and on this SoC the CPU waits for it forever: no oops, no message. We met that twice.
Lanes switched off after training. On the first NovaOS boot with a running FPGA, the link trained, the root port enumerated, and the SoC stopped dead at the first configuration read across the link, the vendor ID of the endpoint. The hardware watchdog reset it about 30 s later. Between training and that read, the 6.12 BSP's host driver powers off every PHY lane its lane-map register does not name, a step Rigol's 4.4 driver never takes, and it does not check the link again before scanning. We moved the device tree to the stock driver's single-PHY model, which leaves all four lanes on, and added a guard to the host driver: after training the link must reach L0 within 500 ms and stay there for 100 ms, or the probe fails with -ENOLINK; configuration accesses are refused while the link reports down. The next image trained and enumerated. We changed both at once, so which one mattered is not proven.
An endpoint that resets under a bound driver. A boot command reboots the whole Zynq, so its endpoint vanishes, and any register access by the XDMA driver or the app at that moment hangs the same way. NovaOS unloads the XDMA driver and the PCIe host before anything resets the Zynq. When we tried probing 1 s after the boot command, the retry path unloaded the driver while the Zynq was resetting, and the driver's close took a synchronous external abort on 3 of about 12 boots. The details of that probe timing are in the FPGA bring-up note, and the PCIe PHY fix that made repeated probes possible at all is in the boot note.
Measured on our own DHO924S and MHO934 in October 2026. The stock driver's behaviour is read from its module binary and from the stock unit's own logs and parameters.