Skip to content

AI Video for Contractors: What It Can and Can't Do Yet

RenderSplash Team7 min read

A labeled diagram of one shot in the pipeline: a rendered still pool image on the left labeled 'What you upload', an arrow to a middle frame labeled 'Photoreal frame we draw', and a short film strip on the right labeled 'What the video model sees', showing that the original 3D render never reaches the video model directly.

What "AI video" means here

Every software vendor at a trade show this year has an AI pitch. Before getting into what a tool like ours can and can't do, it's worth being specific about which "AI video" we mean, because the term covers two very different things.

One kind starts from a text prompt: type "backyard pool with a stone patio" and the model invents a yard, a house, a pool shape, all of it, from nothing you drew. The other kind, which is what RenderSplash builds and what this article is mostly about, starts from a photograph or a rendered still you already have and tries to make that specific scene look real and move the way a camera would move through it. It's a narrower job. It also has a narrower and more honest set of limits, and this article is an attempt to state them plainly rather than let a contractor find out the hard way in front of a client.

What it's genuinely good at right now

Given a clean rendered still from your design software, this category of tool is good at one specific trick: taking a scene that already has the right geometry, the right camera angle, and the right layout, and making the surfaces in it (water, plaster, stone, planting, light) look closer to a photograph than a computer graphic. It is not drawing a new pool. It is repainting the one you already designed and priced.

That distinction matters for everything below, because almost every limit this technology has follows from one fact: it is generating pixels that look plausible, not simulating a physical scene.

What no AI video tool does well yet

This part isn't specific to us. It's true of every serious video model on the market right now, the well-funded ones and the cheap ones alike.

Clips stay short for a reason. Every model in this category shows real quality loss once a shot runs past roughly twenty to twenty-five seconds of continuous generation, and most tools fall apart well before a minute of usable footage (is4.ai, "The State of AI Video Generation in 2026"). That's not a business decision. It's a ceiling on how long these systems can hold a scene together before it drifts.

The drift has a specific cause. Most video models only remember a fixed window of recent frames while generating the next one. Once something falls outside that window, the model is working from a compressed memory of it instead of the real thing, which is why a color, a shape, or a light source can quietly shift partway through a clip (iMerit, "Solving Temporal Drift in AI-Generated Video").

None of it runs on an actual physics engine. These models learned what motion, light, and reflection usually look like by watching enormous amounts of video, not by simulating gravity, water, or optics. One detailed writeup of a leading model's public demo reels catalogued exactly this kind of failure: weapons that "flickered, bent, and morphed into odd shapes," a lion at a watering hole that ended up appearing to "drink directly from the dirt rather than the water," animals that came out "half teddy bear, half inflatable balloon" (Notebookcheck, "Sora 2 is OpenAI's consistently inconsistent AI video creator"). A separate review of AI tools in architectural visualization put the underlying problem in one line: these models are "optimized for aesthetic reward, not physical accuracy," trained to produce whatever statistically resembles a good rendering rather than a physically correct one (Dan Cumberland Labs, "AI Rendering in Architecture: The 'Too Perfect' Problem"). That is precisely why directing these systems in words instead of a real 3D scene is fragile: on our own pipeline, a shot whose instructions described the wrong end of a camera move once caused the model to fill in ground and structure that were never in the original design, because nothing told it not to.

It usually takes more than one try. The advertised price per second of AI video is not what a usable clip actually costs. One cost breakdown written for production teams puts it plainly: every generated attempt is billed whether it passes review or not, so the honest way to price a shot is total spend divided by the clips you actually keep, not the sticker rate (Seeles.ai, "AI Video Generation Cost in 2026"). Anyone telling you a single click reliably produces a finished, client-ready shot isn't being straight with you.

Where our own output has real limits

Our own pipeline inherits every limit above, because it's built on the same kind of models everyone else is. Here's where it specifically runs into a ceiling.

It's short. A finished video today is a matter of seconds, built from a handful of short shots stitched together, not a multi-minute cinematic walkthrough of the whole property. If you're picturing a two-minute film of the entire yard from every angle, that isn't what ships. What ships is a tight clip built for the moment you hit play at the kitchen table, not a feature film.

It only works with what you give it. The video is drawn from your own rendered stills, not invented from a description. If a shot's still doesn't convert cleanly into a photoreal frame, that shot gets dropped from the final cut rather than faked; if none of the shots convert, the video isn't made at all, and you aren't charged for it. That's a deliberate trade against the alternative, which is shipping something that looks finished but shows a pool that isn't the one you actually designed.

Camera movement is simple on purpose. Each shot gets one move: a push in, a pull back, a tilt, described in words alongside the picture. This is also where the general problem above shows up most directly for us. If the words describing the move and the picture itself disagree about which end of the move the still represents, the system fills the gap rather than leaving it blank, the same way every model in this category does. Clean, matching stills and a simple move description produce a far more reliable result than an ambitious one.

Shot-to-shot consistency isn't pixel-perfect. Each shot's photoreal frame is generated on its own before the clip is filmed, so the exact shade of the water or the precise grain of a travertine deck can shift slightly between shots in the same video. It won't turn your pool a different color, but it isn't frame-matched the way a single continuous camera pass would be.

We don't train on your uploads, and that has a cost too. Your renders and finished videos are never used to train our models or anyone else's. The trade-off is that the system doesn't get more tailored to your particular drawing style or your usual angles the more you use it. Every job starts from the same baseline.

Higher resolutions take longer and are newer. Draft and HD generation are the well-worn path. Pushing to 1080p or 4K, or sharpening a finished clip up the resolution ladder, is newer, slower, and has had less real-world mileage behind it than the standard path.

How to work with the limits instead of around them

None of this means the output isn't worth showing a client. It means treating it as what it is: a video that shows design intent, not a construction document, and not a guarantee of the finished build's materials or dimensions. A few practical habits follow directly from the limits above.

Pick your cleanest, most complete renders as the source, because the video can only be as accurate as what you feed it. Keep the camera moves you request simple and matched to what the still actually shows, rather than describing a move the picture doesn't support. And check a shot as a low-cost draft before you spend the credits on the finished version. That's exactly why a draft costs half of what a finished HD video does: it's built for catching a bad shot before it counts against you, not after.

The bottom line

A photoreal video made from your own design is a genuinely different thing from a rendered still, and it isn't magic. It can't invent a design element you never drew, it can't hold a scene together for minutes at a time, and it can't guarantee the exact shade of blue in shot three matches shot one down to the pixel. What it can do is take the render you already paid your design software to produce and make it feel like footage instead of a screenshot, in about the time it takes to walk a homeowner to the kitchen table. If you want to see what that looks like on real output rather than a claim, the gallery has full case pages, source renders included.

Get updates as we ship