AI Video Creation Still Begins With A Simple Text Prompt
By 813 Staff
The demo reel was already looping on screens behind the stage when the first eyebrows went up. At a packed launch event in San Francisco on September 24, a company building generative video tools showed off a workflow that skips the familiar text box entirely. Instead of typing a prompt and waiting for a model to interpret it, users pointed a camera, sketched a rough storyboard, or dropped in a reference clip, and the system filled in the rest. Engineers close to the project say the shift is deliberate: the team concluded that prompting, for all its flexibility, remains a bottleneck for anyone who does not think in paragraphs.
The timing was not accidental. Internal documents show the company has been quietly testing the approach with a few hundred creators since late summer, and the feedback was blunt. Testers liked the output but hated the ritual of describing what they wanted in words. Elias (@iam_elias1) captured the frustration succinctly in a post the same day, noting that so much AI video creation still begins with a prompt and questioning whether that starting point makes sense anymore. His tweet, brief and open-ended, landed as the industry is already circling the same question.
The rollout has been anything but smooth. Two people familiar with the engineering say latency spikes during the live demo forced a staffer to restart a render mid-presentation, and the reference-clip feature was quietly disabled for some accounts afterward. The company has not confirmed those details publicly, and it declined to say when the capability would reach general availability. What is clear is that the underlying pitch — that the next wave of AI video tools will be shaped by gestures, sketches, and existing footage rather than typed instructions — is gaining traction among investors who have grown weary of prompt-engineering tutorials as a product category.
For everyday users, the stakes are practical. If the approach works, the barrier to making a short film, an ad, or a social clip drops again, and the skill of writing a good prompt becomes less valuable than the ability to frame a shot. Rivals are watching closely; at least two competing teams are reportedly prototyping camera-first interfaces of their own.
What happens next depends on whether the company can stabilize the system and ship it broadly. A wider beta is expected in the coming weeks, though no firm date has been set. Until then, the question Elias raised lingers: if the prompt disappears, what replaces it — and who gets to decide?
