video-use
Edit videos with coding agents
video-use is an agent skill that enables AI coding agents to edit videos through natural language commands. Designed for Claude Code, Codex, Hermes, and similar shell-enabled agents, it processes raw footage by cutting filler words and silence, applying color grading, burning subtitles, and generating animation overlays—all through conversation. The skill uses text-based transcripts with word-level timestamps instead of frame analysis, allowing agents to make precise cuts without processing millions of image tokens. Outputs are self-evaluated before delivery, and session memory persists across editing sessions.
Key Features
Use Cases
- 01Editing YouTube talking-head videos by removing umms, uhs, and long pauses automatically
- 02Creating polished tutorial videos from raw screen recordings with auto-generated subtitles
- 03Producing marketing launch videos from multiple raw takes with color grading and overlays
- 04Editing interview footage with speaker diarization and seamless audio transitions
- 05Assembling travel montages with automated color correction and custom animations
- 06Running always-on video editing workflows from a VPS or Telegram via Browser Use Box
Related Skills
View moresuperpowers
An agentic skills framework & software development methodology that works.
skills
Skills for Real Engineers. Straight from my .agents directory.
skills
Public repository for Agent Skills
ponytail
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
video-use — FAQ
What is video-use?+
video-use is an agent skill that allows AI coding agents like Claude Code, Codex, or Hermes to edit videos through conversational commands. It transcribes footage into word-level text, then lets the agent cut, grade, subtitle, and composite without manually processing frames.
How do I install video-use?+
Paste the setup prompt from the README into your AI coding agent, which will clone the repository, install dependencies (ffmpeg and Python packages), and register the skill. Alternatively, manually clone the repo, symlink it into your agent's skills directory (e.g., ~/.claude/skills/video-use), run 'uv sync' or 'pip install -e .', install ffmpeg, and configure your ElevenLabs API key in .env.
Which AI clients work with video-use?+
video-use works with Claude Code, Codex, Hermes, Openclaw, and any AI coding agent that supports shell access and skill registration. For always-on workflows, it can run through Browser Use Box on a VPS or Telegram.
Do I need an API key to use video-use?+
Yes, you need an ElevenLabs API key for audio transcription with word-level timestamps and speaker diarization. You can obtain one at elevenlabs.io/app/settings/api-keys. The skill prompts you to add it during setup.
Is video-use free and open source?+
Yes, video-use is 100% open source and available on GitHub. However, you will need an ElevenLabs API subscription for transcription services, which has its own pricing.
How does video-use make editing decisions without watching the video?+
It uses a text-first approach: ElevenLabs Scribe provides word-level transcripts with timestamps, speaker labels, and audio events in a ~12KB markdown file. The agent reads this and only generates visual composites (filmstrip + waveform + labels) at specific decision points, avoiding the token overhead of processing every frame.
How do I install video-use?+
Open the source repository on GitHub and follow its README. video-use is a skill — MCP Agents Market links you directly to the official repo.
Is video-use free?+
video-use is an open-source project hosted on GitHub. Check the repository for its license and any usage requirements.