28-30 September 2026
America/Toronto timezone

LLM inference on the Raspberry Pi 5 via Vulkan

28 Sep 2026, 11:20
45m
Talk (full slot) Talk (full slot) Main Track

Speaker

José María Casanova Crespo (Igalia)

Description

The Raspberry Pi 5's Broadcom V3D 7.1 GPU was designed for graphics
workloads, but the wave of on-device LLM frameworks now directs pure
compute workloads at the V3DV Vulkan driver. This talk reports what
we learned extending V3DV for these workloads since early 2026.

ML inference stresses the driver in ways graphics never does: thousands
of tiny buffer fills and copies per token, heavy buffer-object allocation
against V3D's 4 GB memory ceiling, shaders with high register pressure,
and a dependence on fast fp16 support.

We walk through the driver changes this new workload required — enough that
it is now possible to run open-weight models with inference frameworks
such as Ollama/llama.cpp and LiteRT-LM.

Finally, the talk discusses future work and the challenges that remain
to run inference on V3DV.

In-person or virtual presentation In-person
GSoC, EVoC or Outreachy No
Code of Conduct Yes

Primary author

Presentation Materials

There are no materials yet.
2026 Platinum Sponsor
Arm
2026 Gold Sponsors
AMD
Collabora
Microsoft
NVIDIA
Raspberry Pi
2026 Silver Sponsors
CodeWeavers
Igalia
Netflix
Qualcomm
Specs
The Linux Foundation
2026 Bronze Sponsors
Imagination Technologies
Khronos Group
Libre Computer
Linaro
LunarG