← Guides

Guide

Make Videos with Your AI Agent

Send your hosted agent a message like “make me a 30-second video about our launch” and a few minutes later a finished MP4 arrives in the same chat — an animated explainer with title cards, stat reveals, and charts, rendered at 1080p. This guide covers how the video pipeline works, how to enable it, what the videos look like, and how to add AI-generated footage with your own key.

How It Works: Your Agent Is the Director

The pipeline splits the work in two. Your agent does the creative part: it reads your brief, decides the story arc, and writes a scene plan — which cards appear, in what order, with what text, numbers, and timing. That scene plan then goes to a dedicated render service, which draws every frame and encodes the MP4 with a real video toolchain (built on the open-source OpenMontage pipeline, with Remotion and FFmpeg under the hood). The heavy rendering runs on separate infrastructure, so your agent stays responsive while the video is produced.

Because the agent writes the scene plan itself, you can iterate conversationally: “make the stats bigger”, “add a comparison of the two plans”, “shorten it to 20 seconds” — each revision is just a new plan and a new render.

Enable It: Install the Make Video Skill

Open your dashboard Skills page, find Make Video in the Creative category, and install it on your instance. The install drops a render credential scoped to that one instance — no keys to paste, nothing to configure. From then on, just ask for a video in any channel your bot is connected to (web chat, Telegram, Discord, WhatsApp, and so on).

A good brief names the topic, the length, and anything concrete you want on screen: “a 45-second explainer about our Q3 results — revenue grew 23%, headcount doubled, close with our logo colors.” Numbers are gold: the renderer has dedicated scene types for stat reveals, KPI grids, and bar, line, and pie charts.

What the Videos Look Like

The default composition is a clean, motion-graphics explainer style: an opening title card, a sequence of content scenes, and a closing card. Scene types include text cards and callouts, quotes with attribution, big animated numbers, KPI grids, charts, side-by-side comparisons, and even terminal-style demo scenes for technical content. Videos are best kept between 20 and 90 seconds — long enough to tell one story, short enough to hold attention.

Rendering takes roughly ten seconds of processing per second of finished video, so a 30-second video is typically ready in about five minutes. Your agent tells you when the render starts and delivers the MP4 file into the chat when it completes.

Optional: AI-Generated Footage with Your Own Key

Out of the box, videos are text- and data-driven — no external services involved. If you want real imagery, add a fal API key on your API Keys page (fal is an aggregator that fronts many image and video models with one key). With a key configured, your agent can generate still images and short video clips and weave them into the scenes as full-screen shots. The agent defaults to inexpensive models — fractions of a cent per image, a few cents per clip — and only reaches for premium video models when you explicitly ask for them.

Plan Limits

Video rendering is included with paid plans: Lite includes 2 videos per day and Pro includes 10 per day, enforced per instance. Free trials don't include rendering — see pricing if you want to try it. AI-generated footage is billed by your own fal account, not by us; the built-in text and chart scenes cost nothing beyond your plan.

Tips for Better Videos

Give the agent real content, not just a topic — paste the stats, the quotes, the feature list you want on screen. Ask for one clear message per video rather than five. If a render fails, the agent reads the error and fixes its own scene plan; if it fails twice it will show you what went wrong. And if you want a specific look, say so: background colors, pacing, aspect ratio for vertical platforms.

FAQ

Do I need any API keys to make videos?

No. The default explainer videos — titles, text, stats, charts — render entirely on managed infrastructure with no external services. A fal key is only needed if you want AI-generated images or clips inside your videos.

How long does a video take?

About ten seconds of rendering per second of finished video, plus queue time. A 30-second video is usually ready in around five minutes; your agent delivers it automatically when done.

Is video rendering available on the free trial?

No — rendering is part of paid plans (Lite: 2 per day, Pro: 10 per day). Everything else about the skill can be explored on a trial, and the agent will explain the limit if you hit it.

Can the videos include voiceover or music?

Not yet — current output is silent, designed for captions-on viewing and social feeds. Voiceover and music tracks are on the roadmap.

Ask Your Agent for a Video

Install the Make Video skill and turn briefs into finished MP4s — titles, stats, charts, and optional AI footage.

Get Started