BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//CERN//INDICO//EN
BEGIN:VEVENT
SUMMARY:LLM inference on the Raspberry Pi 5 via Vulkan
DTSTART;VALUE=DATE-TIME:20260928T152000Z
DTEND;VALUE=DATE-TIME:20260928T160500Z
DTSTAMP;VALUE=DATE-TIME:20260904T190940Z
UID:indico-contribution-529@indico.freedesktop.org
DESCRIPTION:Speakers: José María Casanova Crespo (Igalia)\nThe Raspberry
  Pi 5's Broadcom V3D 7.1 GPU was designed for graphics\nworkloads\, but th
 e wave of on-device LLM frameworks now directs pure\ncompute workloads at 
 the V3DV Vulkan driver. This talk reports what\nwe learned extending V3DV 
 for these workloads since early 2026.\n\nML inference stresses the driver 
 in ways graphics never does: thousands\nof tiny buffer fills and copies pe
 r token\, heavy buffer-object allocation\nagainst V3D's 4 GB memory ceilin
 g\, shaders with high register pressure\,\nand a dependence on fast fp16 s
 upport.\n\nWe walk through the driver changes this new workload required 
 — enough that\nit is now possible to run open-weight models with inferen
 ce frameworks\nsuch as Ollama/llama.cpp and LiteRT-LM.\n\nFinally\, the ta
 lk discusses future work and the challenges that remain\nto run inference 
 on V3DV.\n\nhttps://indico.freedesktop.org/event/12/contributions/529/
LOCATION:
URL:https://indico.freedesktop.org/event/12/contributions/529/
END:VEVENT
END:VCALENDAR
