Skip to content

Add support for vLLM, Gemma4, and local models to GeminiImageCaptioning - #437

Open
adsarver wants to merge 88 commits into
developfrom
gemma-image-captioning
Open

Add support for vLLM, Gemma4, and local models to GeminiImageCaptioning#437
adsarver wants to merge 88 commits into
developfrom
gemma-image-captioning

Conversation

@adsarver

@adsarver adsarver commented Aug 10, 2026

Copy link
Copy Markdown

ZachCafego and others added 30 commits March 13, 2025 11:51
Improve subprocess error handling.
Add data NOTICE.
Adjust test.
adsarver added 28 commits May 20, 2026 14:01
…to motion emphasizer, general linting of spaces
- Now (by default) samples from the video at 1 FPS (hard limit for Gemma4)
- Guarantees frames with moving objects are included and viewable by the model
- Delivers similar performance to previous methods (high FPS, many duplicate frames) while reducing compute
- Maps sampled frames to real frames for timestamp tracking
- Added new motion threshold param (functional equivalent to motion percentile param)
…mpeg to the config, changed how motion is profiled, added support for audio tracks
…d pipeline default model, changed pipeline names
…iner, fixed typo in vllm dockerfile, added optional testing at image build without GPU
… old folder for easy comparisons, will remove in next commit
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants