Clip Crafter AI
AI platform converting text into fully rendered short videos
Overview
What this is
Clip Crafter AI turns a text prompt into a fully rendered short video — script, voiceover, captions, and visuals generated and assembled automatically.
Problem
Why it needed building
Producing a short video from a text idea normally means stitching together several unrelated tools by hand — a script writer, a text-to-speech engine, a captioning tool, and a video renderer — with nothing connecting them into one pipeline.
Solution
How it works
Built a scalable system integrating Gemini for script generation, Google Text-to-Speech for voiceover, AssemblyAI for captions, and ClipDrop for visuals, assembled into a final video with Remotion, all behind a single async processing pipeline.
Architecture
How it's put together
- →Async processing pipeline chaining script generation (Gemini), voiceover (Google TTS), captioning (AssemblyAI), and visuals (ClipDrop)
- →Video assembly and rendering handled by Remotion
- →Neon PostgreSQL with Drizzle ORM for structured data
- →Firebase Storage for generated media assets, Clerk for authentication
Challenges
The hard part
Keeping the pipeline resilient when any one of four external AI services is the one that's slow or fails — a single stuck step shouldn't fail the entire video generation job.
Trade-offs
What I gave up on purpose
Chaining four separate external services gives more creative control over each step than a single all-in-one video API would, at the cost of more integration surface area to maintain.
Performance
What changed, measurably
Produces a fully rendered short video — script through final render — from a single text prompt via one async pipeline.
Next
What I'd build next
Add a queue-based retry mechanism per pipeline step so a single failed API call doesn't require regenerating the entire video from scratch.