Back to glossary

Text-to-Video

Definition

You write a description. A model reads it and produces a video with pictures, narration, music, and captions already in place. No footage to shoot and no timeline to cut. The description can be a single sentence or a full article, and what comes back is a finished file rather than a project to edit.

The phrase covers two different products, and mixing them up is how people end up disappointed.

One kind is the raw clip model. Sora, Kling, Veo and their peers take a prompt and hand back a few seconds of footage, usually capped somewhere between five and ten seconds per generation, with no sound and no structure. They are remarkable at the shot level and useless on their own, because a shot is not a video.

The other kind is a pipeline. It writes the script first, decides the scene breakdown, calls a clip model or an image model per scene, records the narration, scores it, times the captions to the voice, and hands back a file you can upload. Orange is the second kind and it calls the first kind underneath, routing each scene to a frontier clip or image model and billing by target runtime rather than per scene.

Input length is more flexible than the name suggests. A sentence works. So does a 2,000-word article: paste a URL and the scraper pulls the body text, strips the navigation, and hands the writer a distilled brief instead of raw HTML. Long input usually produces a better script than short input, because the writer has real specifics to quote rather than gaps to invent around.

What you should not expect is an edit that needs no judgment. The first pass gets the pacing roughly right and the emphasis roughly wrong. Read the script before you spend credits on the render, swap the two scenes that are in the wrong order, and cut the line that says nothing. That review costs no credits at all, and it is the difference between a video you post and one you quietly delete.

Try Orange

Create AI videos from text descriptions. Free to start.

Get started