Video Open-Source Projects
Browse selected Video open-source AI projects from the current public snapshot, with verified use cases, deployment notes, limitations, and sources.
Showing 27 indexable projects from the current snapshot.
Robbyant/lingbot-map
LingBot-Map is a feed-forward 3D foundation model for reconstructing scenes from streaming image or video data. It unifies coordinate grounding, dense geometric cues, and long-range drift correction to support efficient streaming 3D recons…
bradautomates/claude-video
This project adds video comprehension to AI coding assistants and chat agents by downloading content, extracting frames, and transcribing audio. It operates as a portable skill unit across over 50 agent hosts without modification, giving m…
Comfy-Org/ComfyUI
A modular node graph interface and backend for designing generative AI workflows across images, videos, audio, and 3D models without writing code. It operates fully offline with smart memory management that supports GPUs with as low as 1GB…
harry0703/MoneyPrinterTurbo
MoneyPrinterTurbo automates short video creation from a single topic or keyword, handling scripting, material sourcing, voiceover synthesis, subtitle generation, and HD video composition. It is a self-hostable application for creators, wit…
huggingface/diffusers
Diffusers provides modular inference pipelines, interchangeable schedulers, and pretrained building-block models for diffusion-based image, audio, and video generation in PyTorch. It is intended for developers and researchers who need to r…
opencv/opencv
OpenCV is an open-source library that provides algorithms and functions for computer vision, deep learning, and image processing. It operates locally as a dependency, handling image and video inputs to produce processed media and feature d…
ultralytics/ultralytics
Ultralytics provides SOTA YOLO models for real-time computer vision across object detection, segmentation, classification, pose estimation, and tracking. It is accessible via a unified CLI and Python API for model training, evaluation, and…
roboflow/supervision
Supervision provides reusable Python building blocks for computer vision workflows, including model-agnostic connectors for detection and segmentation results and customizable visualization annotators. It is intended for developers who nee…
calesthio/OpenMontage
OpenMontage provides an end-to-end video production pipeline that automates research, scripting, asset generation, and editing via AI coding assistants. It orchestrates structured workflows from natural language prompts or reference videos…
ATH-MaaS/Pixelle-Video
The project automates end-to-end short video creation from a single text prompt by combining scriptwriting, AI media generation, voiceover synthesis, and video composition. It is self-hosted via a web GUI and targets creators and general u…
Anil-matcha/Open-Generative-AI
A free, open-source, self-hostable generative AI studio for creating images, video, and audio using 200+ models without built-in content filters or subscription fees. Local inference is limited to the desktop application and requires suffi…
jianchang512/pyvideotrans
A one-click video translation and audio transcription tool that handles speech recognition, subtitle translation, multi-role AI dubbing, and video synthesis. It supports both local offline deployment and various online APIs.
lllyasviel/FramePack
framepack is an application that makes video diffusion practical by compressing frame contexts so generation workload stays invariant to video length. It can process many frames with 13B models on laptop GPUs, with a documented minimum of…
camenduru/stable-diffusion-webui-colab
A collection of pre-configured Google Colab notebooks that deploy AUTOMATIC1111's Stable Diffusion WebUI with models, extensions, and training templates already installed. It targets creators and general users who want to generate images a…
WEIFENG2333/VideoCaptioner
A desktop GUI and CLI application that consolidates video subtitling tasks—transcription, optimization, translation, synthesis, and dubbing—into a single pipeline. It includes no-configuration free options alongside LLM-based processing fo…
AliaksandrSiarohin/first-order-model
first-order-model animates a still source image by transferring motion from a driving video, providing source code and pre-trained checkpoints for image animation and motion retargeting across various datasets.
waooAI/waoowaoo
An AI video-production application that automatically converts novel text into storyboards, character and scene images, voices, and complete videos. It is self-hosted via Docker Compose and requires an external AI service API key.
HBAI-Ltd/Toonflow-app
An open-source AI tool that provides a text-to-video pipeline for turning novels and scripts into animated short dramas, organizing the workflow on an infinite canvas. It requires external AI service interfaces for large language models, v…
zai-org/CogVideo
Open-source diffusion models for text-to-video and image-to-video generation. Requires GPU hardware for inference, with support for running on lower-memory devices via quantized methods.
palmier-io/palmier-pro
PalmierPro is a free, open-source Swift-native macOS video editor that enables collaborative timeline editing between users and AI coding agents. It runs locally on Apple Silicon, exposes an MCP server so external agents can drive video ed…
hua1995116/awesome-ai-painting
This project is a read-only curated list of AI painting resources, including platform directories, tutorials, prompt databases, and self-deployment guides. It is intended for general users and creators seeking organized references for tool…
HKUDS/ViMax
hkuds/vimax is an agent-based framework that converts concepts or screenplays into multi-scene videos through scriptwriting, storyboarding, character creation, and video generation. It requires command-line usage, Python configuration, and…
ostris/ai-toolkit
ai-toolkit is a free, open-source training suite for fine-tuning diffusion models on image, video, and audio generation tasks. It is designed to run on consumer-grade hardware via a web interface or command line.
TEN-framework/ten-framework
ten-framework is an open-source framework for developing and deploying real-time multimodal conversational AI agents. It provides extensive examples and deployment options while addressing the need to integrate various models and services…
voxel51/fiftyone
fiftyone is an open-source tool for building high-quality datasets and visual AI models. It enables users to visualize and label data, evaluate models, and improve data and model quality.
magic-research/magic-animate
magic-animate generates temporally consistent human image animations from a reference image and a motion sequence using a diffusion model. It supports inference on both single and multiple GPUs and offers a local Gradio web interface for i…
Baiyuetribe/paper2gui
Paper2GUI packages more than 50 AI models into install-free desktop GUI tools for image, video, and audio tasks. It targets non-technical users, but operating-system support varies by tool and some TTS functions may rely on third-party ser…