Verified project record
modelscope/FunClip
FunClip automates video transcription, subtitle generation, and clipping locally using speech recognition and optional LLM-based content selection. It supports multi-segment clipping, speaker diarization, and hotword customization through a Gradio web UI or command line interface.
Project overview
FunClip automates video transcription, subtitle generation, and clipping locally using speech recognition and optional LLM-based content selection. It supports multi-segment clipping, speaker diarization, and hotword customization through a Gradio web UI or command line interface.
- Project type
- Video · Audio & Speech
- Use cases
- Translation & Subtitles · Video Creation
- Deployment
- Refer to project documentation
- License
- MIT
Best for
- Creators who need to process video files locally for transcription and multi-segment clipping.
- Developers or general users requiring automated text and LLM-based video highlight extraction.
Key capabilities
- Performs speech recognition on uploaded videos using Paraformer, Fun-ASR-Nano, or SenseVoice models to transcribe audio into text with timestamps.
- Extracts video highlights by combining ASR transcripts with LLMs like GPT or Qwen to automatically find and clip relevant segments.
- Integrates the CAM++ speaker recognition model to identify speakers, allowing users to clip segments from a specific speaker ID.
- Allows users to specify entity words or names as hotwords during the speech recognition process using SeACo-Paraformer to enhance recognition accuracy.
- Supports clipping multiple selected text segments freely and automatically returning the full video SRT subtitles and target segment SRT subtitles.
- Provides a Gradio-based web user interface that can be run locally or deployed on a server, accessed via a browser.
- Supports recognizing and clipping English audio files using the Paraformer English model or Whisper (planned).
- Provides a CLI tool to recognize and clip videos via commands.
Limitations and risks
- The Fun-ASR-Nano checkpoint does not provide reliable character-level timestamps; Paraformer is recommended for precise text-based clipping.
- Supporting the Whisper model for English requires massive GPU memory (planned).
- Clipping video files with embedded subtitles requires optional manual installation of ffmpeg and imagemagick.
- Using LLM-based clipping or the TwelveLabs API may incur costs or require paid API keys.
Getting started
- Clone the repository, install requirements via pip, and run the Gradio UI using python funclip/launch.py. The tool provides a simple pip installation and an automatic Gradio UI.
Evidence and sources
- GitHub project description: FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
- README: **FunClip** is a fully open-source, locally deployed automated video clipping tool.
- README: The functionalities are realized through Gradio interaction, offering simple installation and ease of use. It can also be deployed on a server and accessed via a browser.
- README: 🔥2024/05/13 FunClip v2.0.0 now supports smart clipping with large language models, integrating models from the qwen series, GPT series, etc., providing default prompts. You can al…
- README: FunClip now supports Fun-ASR-Nano and SenseVoice models. The `fun-asr-nano` option loads the flagship Fun-ASR-Nano-2512 checkpoint for Mandarin, English, Japanese, 7 Chinese diale…
AI 搜索
把需求说清楚,让项目选择更有依据
告诉我们你要解决什么、运行在哪里、哪些条件不能妥协。雷达会从已核验项目中给出主推荐、备选和采用前检查。
目标你最终想完成什么
环境本地、云端或现有技术栈
硬条件部署、界面、语言与 License
从一个真实需求开始点击只会填入搜索框,你可以继续修改
今日榜单
0