Infrastructure from existing cloud-based elements
textaccumulate
textwrap
New whisper.cpp transcription element (gst-plugins-rs)
New llama.cpp text transformation element (gst-plugins-rs)
Following examples only need a GPU with 4GB of free VRAM
And these are not even the smallest models:
Mostly idle during real-time playback
Top text: transcription | Bottom text: German translation
Not perfect, and not as good as human transcription and translationbut a lot better than not having any at all!
Scene descriptions, OCR, scene change descriptions, ...
Your imagination is the limit
In a local branch for now (→ hackfest!)
The first frame shows a duck with a black head, orange bill, white body, and brown chest standing on grass, with a white bowl visible in the upper right corner.
The second frame is a close-up shot that crops out that duck's body and the background bowl, focusing instead on its head and upper neck.
Merge vision branch, explore audio/multi-modal inputs
Integration with general inference / gstanalytics design
audio.cpp: https://github.com/0xShug0/audio.cpp
Integration into a video player?
Talk to me later!
Code available in the main branch at