Transform your media pipelines with AI-powered insights! π₯π€
Report Bug
β’
Request Feature
Ever wondered what a GStreamer pipeline would say if it could talk? Now it can (almost)! With GstVlmVision (formerly GstGeminiVision), you can inject the power of any OpenAI-compatible vision-language model directly into your GStreamer media pipelines. Turn your video streams into insightful descriptions, automate content analysis, or just have some fun making your videos self-aware! π€π¬
vlm_vision_demo.mp4
GstVlmVision is a GStreamer plugin that acts as a bridge between your live video or image streams and any OpenAI-compatible vision-language model API. It periodically captures frames, sends them to a VLM endpoint via /v1/chat/completions, and then makes the generated description available through GstMeta, GObject signals, or GstBus messages.
Supported providers include OpenAI, Google Gemini (via OpenAI compatibility), vLLM, Ollama, and any /v1/chat/completions-compatible server.
Imagine:
- Generating visual descriptions for accessibility purposes.
- Creating a security camera that describes what it sees.
- Building interactive art installations that react to visual input.
- Automating video content moderation and tagging.
- ...the possibilities are as vast as your imagination (and the VLM's capabilities)!
- OpenAI Chat Completions Multimodal -- standard
/v1/chat/completionswith text + image_url content blocks - Multiple Output Modes -- bitmask-selectable: GstMeta, GObject signal, GstBus message (details)
- Template Override Mode --
{{PLACEHOLDER}}templates for nonstandard API providers (details) - Async HTTP with Concurrency Control -- threaded dispatch with configurable
max-inflightandtimeout - JPEG Frame Encoding -- raw video frames encoded to JPEG before sending
- Docker Support -- pre-configured build environment (details)
- Python Bindings -- GObject Introspection via
gi.repository.GstVlmVision
- Show a
videotestsrcpipeline running with the plugin. - Display the console output from
vlm_vision_example.pyorvlm_vision_example.cshowing the descriptions. - Bonus: If you have a more complex demo (e.g., overlaying text on video), showcase that!
gst-launch-1.0 -m videotestsrc ! videoconvert ! \
vlmvision base-url="https://api.openai.com" api-key="sk-..." \
model="gpt-4o" profile="openai" \
user-prompt="What do you see?" output-mode=4 \
! videoconvert ! autovideosinkSetting pipeline to PLAYING state...
Pipeline running...
Press Ctrl+C to quit
Pipeline state changed from NULL to READY
Pipeline state changed from READY to PAUSED
Pipeline state changed from PAUSED to PLAYING
=================================
Frame time: 0.000000000 (PTS: 0)
Description: That's a color bars test pattern, used to adjust color settings on television screens.
=================================
^CInterrupt received, stopping...
Cleaning up...
Pipeline stopped.gst-launch-1.0 -m videotestsrc ! videoconvert ! \
vlmvision base-url="http://localhost:8765" api-key="dummy" \
model="gpt-4o" profile="openai" \
user-prompt="What do you see?" output-mode=4 \
! videoconvert ! autovideosinkReady to give your GStreamer pipelines a voice? Let's go!
- GStreamer: Core GStreamer libraries and development files (version 1.16+ recommended).
- Build Tools:
meson,ninja,gcc(or your C compiler),pkg-config. - Dependencies for the Plugin:
libglib2.0-devlibgstreamer-plugins-base1.0-devlibcurl4-openssl-dev(or your system's cURL dev package)libjson-c-devlibjpeg-devlibgirepository1.0-dev&gobject-introspection(for GObject Introspection, used by Python bindings)
- Python 3 (for the Python example):
python3-gipython3-gst-1.0
-
Clone the repository (if you haven't already):
git clone https://github.com/Armaggheddon/GstVlmVision.git cd GstVlmVision -
Navigate to the plugin directory:
cd gst-vlm-plugin -
Configure and build with Meson & Ninja: First, set up the build directory using Meson. The
--prefix=/usrensures that a subsequent install places files in standard system locations.meson setup build --prefix=/usr --buildtype=release
Then, compile the plugin:
ninja -C build
Your compiled plugin shared object (e.g.,
libgstvlmvision.so) will be located in thegst-vlm-plugin/build/src/directory. -
Install the Plugin (Optional, but Recommended for System-Wide Access): To make the plugin and its development files available system-wide, run the install command (this usually requires root privileges):
sudo ninja -C build install
This command will copy the necessary files to standard system locations. Based on a typical installation with
--prefix=/usr, the files will be placed as follows:- The plugin library:
libgstvlmvision.soto/usr/lib/x86_64-linux-gnu/gstreamer-1.0/ - GObject Introspection data:
GstVlmVision-1.0.girto/usr/share/gir-1.0/GstVlmVision-1.0.typelibto/usr/lib/x86_64-linux-gnu/girepository-1.0/
- Pkg-config file:
gstvlmvision.pcto/usr/lib/x86_64-linux-gnu/pkgconfig/
(Note: The exact paths like
x86_64-linux-gnumight vary slightly based on your Linux distribution's multiarch setup.)After installation, GStreamer should be able to automatically discover the plugin. Verify with:
gst-inspect-1.0 vlmvision
- The plugin library:
For development without system install, you can set environment variables instead:
export GST_PLUGIN_PATH=$(pwd)/build:$GST_PLUGIN_PATH
export GI_TYPELIB_PATH=$(pwd)/build:$GI_TYPELIB_PATHSet at minimum an API key:
# OpenAI
export VLM_API_KEY=sk-...
# or Gemini via OpenAI endpoint
export VLM_API_KEY=<google-api-key>
export VLM_BASE_URL=https://generativelanguage.googleapis.com/v1beta/openai
export VLM_MODEL=gemini-2.0-flash
# or local vLLM (no key needed)
export VLM_BASE_URL=http://localhost:8000
export VLM_MODEL=llava-v1.6-34bNavigate to the examples directory and compile/run:
cd examples && make all
./bin/vlm_vision_example
./bin/vlm_vision_example2 [video-uri]cd examples
python3 vlm_vision_example.pyYou should see descriptions from the VLM printed to the console! π
For a full, detailed list of all properties, their types, default values, ranges, and descriptions, please refer to the output of gst-inspect-1.0:
β‘οΈ View Full Plugin Details (gst-inspect-1.0 output) β¬ οΈ
You can also generate this information yourself by running:
gst-inspect-1.0 vlmvisionHere's a summary of the key properties:
| Property | Type | Default | Description |
|---|---|---|---|
base-url |
string | https://api.openai.com |
Base URL for the API endpoint |
api-key |
string | NULL | API key; NULL = no auth header sent |
model |
string | "" (must set) |
Model name (e.g. gpt-4o, gemini-2.0-flash) |
system-prompt |
string | NULL | System prompt |
user-prompt |
string | "Describe what you see..." |
User prompt about the image(s) |
profile |
string | openai |
API compatibility profile |
analysis-interval |
double | 5.0 | Seconds between analyses (0.1--3600) |
frames-per-request |
int | 1 | Images per request (1--10) |
request-mode |
enum | chat-completions |
API request type |
stop-sequences |
GStrv | NULL | Stop generation strings |
temperature |
double | 1.0 | Sampling temperature (0.0--2.0) |
max-output-tokens |
int | 800 | Max tokens to generate |
top-p |
double | 0.8 | Nucleus sampling (0.0--1.0) |
output-mode |
flags | 3 | Bitmask: 1=meta, 2=signal, 4=bus, 8=json |
timeout |
int | 30 | HTTP timeout in seconds (1--600) |
max-inflight |
int | 1 | Max concurrent requests (1--16) |
queue-policy |
enum | drop |
Backpressure: drop or block |
error-policy |
enum | skip |
Error handling: skip, bus, or signal |
template-body |
string | NULL | File path for body template override |
template-headers |
string | NULL | File path for headers template override |
template-response |
string | NULL | JSON path for response extraction |
Full details with usage examples: docs/properties.md
output-mode is a bitmask combining:
| Value | Flag | Effect |
|---|---|---|
| 1 | META | Attach GstVlmVisionMeta to downstream buffers |
| 2 | SIGNAL | Emit description-received GObject signal |
| 4 | BUS | Post vlmvision-result GstBus message |
| 8 | JSON | Include raw JSON in bus message |
Default: 3 (meta + signal). See docs/output-modes.md for signal signatures, bus message structures, and code examples.
| Policy | Value | Behaviour |
|---|---|---|
skip |
default | Silently skip API errors |
bus |
-- | Post vlmvision-error GstBus message with { message, http-status } |
signal |
-- | Emit analysis-error(error_message, http_status) GObject signal |
For APIs that don't follow the standard OpenAI format, use template files to override request serialization and response parsing.
| Property | Purpose |
|---|---|
template-body |
File path -- JSON body with {{PLACEHOLDER}} substitutions |
template-headers |
File path -- per-line header templates |
template-response |
JSON path (e.g. $.choices[0].message.content) |
Supported placeholders: {{BASE_URL}}, {{MODEL}}, {{SYSTEM_PROMPT}}, {{USER_PROMPT}}, {{IMAGE_0_BASE64}}, {{IMAGE_0_DATA_URL}}, {{IMAGE_ARRAY_JSON}}, {{TEMPERATURE}}, {{MAX_TOKENS}}, {{TOP_P}}, {{SCHEMA_JSON}}, {{REQUEST_ID}}, {{STOP_SEQUENCES_JSON}}.
Full guide with examples: docs/template-mode.md
Want to dive straight into the action without wrestling with dependencies? Our Docker setup is your golden ticket! ποΈ It's like having a pre-configured media lab, ready to build, test, and run GstVlmVision with just a few commands. No more "it works on my machine" β it'll work in this machine!
Step 1: Build the All-Powerful Docker Image
First, conjure up your Docker image. This image contains all the tools and magic needed. From your GstVlmVision project root:
docker build -t gst-vlm-vision .(Psst! If you've already built it, you can skip this step unless you've made changes to the Dockerfile or the plugin build process itself.)
Step 2: Unleash the Entrypoint Script!
The Docker image comes with a super-handy entrypoint.sh script that acts as your mission control. You tell it what to do, and it handles the nitty-gritty. Here are your commands, Captain:
-
build(Default Action): Compile the Mighty Plugin! Just want to build the mainvlmvisionplugin? This is your command. It compiles the plugin but doesn't install it system-wide in the container. Perfect for a quick compilation check.# Run from your GstVlmVision project root docker run --rm \ --volume $(pwd)/gst-vlm-plugin:/builder \ gst-vlm-vision build
-
build-examples: Build the Plugin & The Examples! This action first ensures the mainvlmvisionplugin is built and installed inside the container. Then, it gallops over to your examples directory (/examplesin the container) and builds them (either using the Makefile or compiling C files directly).# Run from your GstVlmVision project root docker run --rm \ --volume $(pwd)/gst-vlm-plugin:/builder \ --volume $(pwd)/examples:/examples \ gst-vlm-vision build-examples
-
test-examples <example_script_name>: The Grand Showcase! This is where the real fun begins! This action:- If you specify a C example (e.g.,
vlm_vision_example.c), it compiles it on the fly if not already built. - Runs your chosen example script (C or Python)!
β¨ Requires
VLM_API_KEY! β¨
# Example for the C script (vlm_vision_example.c): # Run from your GstVlmVision project root docker run --rm \ -e VLM_API_KEY="YOUR_ACTUAL_API_KEY" \ -e DISPLAY=$DISPLAY -v /tmp/.X11-unix:/tmp/.X11-unix \ --volume $(pwd)/gst-vlm-plugin:/builder \ --volume $(pwd)/examples:/examples \ gst-vlm-vision test-examples vlm_vision_example.c # Example for the Python script (vlm_vision_example.py): # Run from your GstVlmVision project root docker run --rm \ -e VLM_API_KEY="YOUR_ACTUAL_API_KEY" \ -e DISPLAY=$DISPLAY -v /tmp/.X11-unix:/tmp/.X11-unix \ --volume $(pwd)/gst-vlm-plugin:/builder \ --volume $(pwd)/examples:/examples \ gst-vlm-vision test-examples vlm_vision_example.py
Don't forget to replace
"YOUR_ACTUAL_API_KEY"! The X11 forwarding lines are for examples that pop up a video window. - If you specify a C example (e.g.,
-
shell: Your Personal Command Deck! Want to poke around inside the container? Need to run some custom commands or debug something? Theshellaction drops you right into an interactive command line.# Run from your GstVlmVision project root docker run -it --rm \ -e VLM_API_KEY="YOUR_ACTUAL_API_KEY" `# Optional, but good to have if you plan to test` \ -e DISPLAY=$DISPLAY -v /tmp/.X11-unix:/tmp/.X11-unix \ --volume $(pwd)/gst-vlm-plugin:/builder \ --volume $(pwd)/examples:/examples \ gst-vlm-vision shell
Inside the shell, your plugin source will be at
/builderand examples at/examples. The main plugin won't be installed by default with this action alone, but theentrypoint.shscript itself is available at/entrypoint.shif you want to manually trigger parts of its logic, or use this shell after runningtest-examplesto inspect a fully set-up environment.
Note
Using -e DISPLAY=$DISPLAY -v /tmp/.X11-unix:/tmp/.X11-unix allows GUI elements to display on your host machine. This requires running the command xhost + on your host Linux machine to allow the Docker container to access your display. If you're using a different display server or setup, you might need to adjust these flags accordingly.
Important Notes for Docker Adventures:
- Volume Mounts are Key: The
--volume $(pwd)/...:/...flags map directories from your computer into the Docker container./builder: Points to yourgst-vlm-plugindirectory. This is where the main plugin source code lives./examples: Points to yourexamplesdirectory.
- API Key: For
test-examples, theVLM_API_KEYenvironment variable (-e) is crucial. The plugin won't talk to the VLM without it! - GUI Display: If your examples use
autovideosinkor any other element that creates a window, you'll need the-e DISPLAY=$DISPLAY -v /tmp/.X11-unix:/tmp/.X11-unixlines (and sometimesxhost +local:dockeron your host Linux machine) to see the output.
With these commands, you're all set to explore the wonders of GstVlmVision without breaking a sweat over setup!
| Old | New |
|---|---|
Element: geminivision |
Element: vlmvision |
model-name |
model |
prompt |
user-prompt |
output-metadata (bool) |
output-mode (bitmask) |
top-k |
Removed (not in OpenAI API) |
| URL hardcoded to Google | Configurable base-url |
Before:
geminivision api-key="YOUR_KEY" prompt="Describe this" model-name="gemini-2.0-flash"After:
vlmvision base-url=https://generativelanguage.googleapis.com/v1beta/openai \
api-key="YOUR_KEY" user-prompt="Describe this" model="gemini-2.0-flash" \
profile=gemini-openaivideo frame -> jpeg-encoder -> VlmRequest -> serializer -> http-client (libcurl POST)
VlmResult <- response-parser <- HTTP response
|
+-> GstVlmVisionMeta (buffer metadata)
+-> description-received (GObject signal)
+-> vlmvision-result (GstBus message)
Key source modules in gst-vlm-plugin/src/: gstvlmvision.c (element), gstvlmvisionmeta.c (metadata), serializer-openai-chat.c, serializer-template.c, response-parser-openai-chat.c, response-parser-template.c, template-engine.c, http-client.c, jpeg-encoder.c, queue-policy.c.
Contributions welcome -- bug fixes, features, documentation. Open an issue or submit a PR.
MIT License -- see the LICENSE file.
Happy Hacking and may your pipelines be ever insightful! π‘


