Speaker
Description
Scale-up accelerator fabrics are already here: amdgpu has exposed xGMI topology through its own sysfs for years, AMD recently posted patches of UALink pod infrastructure, and Intel's XeLink/IAF support was posted and never became shared infrastructure. Every one of these describes the same thing: which accelerators are directly connected, through which ports, in what state. And everyone describes it differently.
The drm/fabric RFC proposes one vendor-neutral model for that: fabric → endpoint → port → peer, over Generic Netlink, with the core recording direct adjacency only and leaving routing, transport and hardware programming to the drivers. It ships with a synthetic provider and no real hardware backend on purpose, so the model can be argued about before it's tied to anyone's silicon.
I'll spend five minutes on the model and then ask the room the question the RFC actually needs answered: is this the right minimum common representation, and will vendors implement it, or do we accept per-driver topology forever?
| GSoC, EVoC or Outreachy | No |
|---|---|
| In-person or virtual presentation | In-person |
| Code of Conduct | Yes |
















