Local transcription and translation with GStreamer

whisper.cpp + llama.cpp

GStreamer Conference 2026

10 October 2026

Sebastian Dröge <sebastian@centricular.com>

Why local?

  • Privacy
  • Offline and edge processing
  • Cost
  • Fun and owning your computing experience

We have all the pieces in place!

But LLMs require huge data centres?

  • Following examples only need a GPU with 4GB of free VRAM

  • And these are not even the smallest models:

    • Qwen3.5 4B (4 bit)
    • Whisper large v3 turbo (8 bit)
  • Mostly idle during real-time playback

The partial pipeline graph

center

Translation example

center

Top text: transcription | Bottom text: German translation

Not perfect, and not as good as human transcription and translation
but a lot better than not having any at all!

Vision input for llama.cpp

  • Scene descriptions, OCR, scene change descriptions, ...

  • Your imagination is the limit

  • In a local branch for now (→ hackfest!)

Vision input example

The first frame shows a duck with a black head,
orange bill, white body, and brown chest
standing on grass, with a white bowl visible in
the upper right corner.

The second frame is a close-up shot that crops out
that duck's body and the background bowl,
focusing instead on its head and upper neck.

Future Work

  • Merge vision branch, explore audio/multi-modal inputs

  • Integration with general inference / gstanalytics design

    • How? Needs some thought!
  • audio.cpp: https://github.com/0xShug0/audio.cpp

    • 110+ audio model families
    • TTS, voice cloning, ASR, speech-to-speech, music generation, ...
  • Integration into a video player?

Questions? Comments?

Talk to me later!

Code available in the main branch at