Most enterprise AI advice starts in the same sensible place: pick a narrow task, give the model reliable context and a limited set of tools, then measure the result. I agree with that advice.
Which makes my choice for Showspring look slightly irrational.
Generative video is expensive and unpredictable. One episode crosses writing, performance, visual continuity, editing, sound, publishing, and audience response. Each stage can produce something technically impressive that still feels lifeless. Viewers may not name every defect, but they recognize the category immediately: AI slop.
That was the attraction. If the system fails, I cannot hide the failure in a dashboard. The audience sees it.
If an agentic system can help create a recurring show that people actually want to watch, it has solved much more than video generation.
I wanted to see how far an agentic system could go when the assignment was not narrow. To make a recurring show, it has to coordinate specialized work, preserve memory across a long process, ask for approval at the right moments, retry intelligently, and evaluate its own output without mistaking automation for taste. Video makes all of that visible. It is a stress test for applied AI that happens to end in a story.
A story has nowhere to hide
Many valuable AI systems succeed quietly: a classification is correct, a document reaches the right queue, a forecast improves, or a support answer resolves the question. These are good places to create near-term value because the task can be isolated and the error can usually be measured.
The character has to feel like the same character from one shot to the next. A line has to sound like something that character would say. The image, motion, voice, timing, music, and edit all have to support the same intent. One wrong prop, one uncanny pause, one generic line, or one voice assigned to the wrong speaker can puncture the illusion.
Building Showspring, I kept seeing the same pattern: models do not fail in isolation. Their errors compound. A weak idea creates a weak script. A vague script creates a generic image. That image gives the video model nothing specific to preserve. By the time the problem reaches the edit, I can make the result shorter, but I cannot make it meaningful.
How do you make AI-generated content feel naturally assisted? The viewer should be glad AI made the work possible, not simply notice that it became cheap enough to publish.
What “human” means in a production system
Calling a prompt “warmer,” “funnier,” or “more authentic” does not make the result human. Those instructions rarely survive an entire production pipeline.
I learned quickly that humanization is the accumulation of small decisions. The show bible remembers how a character behaves. A readout catches a bad line before I turn it into 40 shots. Sometimes the cleanest image is the wrong image. A pause can matter more than another line, and an imperfection can be the thing that makes a performance believable.
I do not see the choice as human or AI. The AI can propose, generate, compare, and revise. The person supplies the intent, taste, and final judgment. The machine can absorb more of the production burden without becoming the author.
Some tasks clearly do get automated. Showspring can synthesize a voice, build a start frame, generate motion, find a cut, draft metadata, and prepare a publish package. I am comfortable with that. The part I care about is lowering the cost floor while keeping the creator in charge of what gets made.
A creator no longer needs to master every tool, buy every piece of equipment, hire every specialist, or turn a hobby into a full-time production company before telling an ambitious story. AI does not eliminate the cost of creativity. It changes which costs are fixed, which are variable, and where human attention creates the most value.
The DNA that prevents every show from becoming the same show
One recent Showspring failure taught me more than a clean demo would have: the shows were starting to converge.
The system began as the production engine for The Doodle Cast. When I expanded it to other shows, changing the cast and premise was not enough. Two-host banter, familiar pacing, the same escalation pattern, and eventually similar visual preferences kept resurfacing. The names changed, but the machinery pulled each new show toward the first format that had worked.
That is a recognizable enterprise AI problem. A shared platform creates leverage, but its defaults quietly become policy. If the system has one strong exemplar, every new use case starts to resemble it.
So I built a per-show Style DNA: a versioned identity contract derived from the creator brief, show bible, format, and cast. Calling a show “warm,” “funny,” or “cinematic” was not enough. A model can flatten those adjectives into the same generic output. The DNA has to demonstrate the show through signature lines, sample beats, generic-to-in-voice rewrites, cadence, dialogue density, favored language, forbidden patterns, visual grammar, and rules for when silence should carry the scene.
I did not arrive at that design in one pass. I first replaced the inherited Doodle Cast formula with neutral production rules for other shows. Then Style DNA became a binding contract that outranks generic writing guidance. Detailed creator briefs became authoritative instead of being softened into suggestions. Visual-first shows gained actual visual and action-only clips instead of being forced to fill every shot with dialogue. The same idea eventually moved into the image system, where each show can specify its rendered medium, color grade, lighting, camera language, texture, and the looks it must avoid.
That contract now reaches much further than the script. It shapes the pitch grammar used during idea generation; character voice and cadence; dialogue, narration, hybrid, and silent clip structures; and the visual direction applied to scenes, character portraits, locations, props, guest assets, podcast artwork, and consistency boards. Evaluation checks whether the output is both compliant and distinct, rather than merely polished.
The DNA is derived once during show setup and persisted, not regenerated in the middle of every request. When an episode enters production, Showspring records the DNA version and a stable content hash. Text DNA and visual DNA are frozen separately, so a later visual adjustment does not make an older script irreproducible, and a voice change does not silently alter images already in production.
That separation lets identity and optimization do different jobs. Audience analytics can influence which premise to explore next, where a format lost attention, or which program deserves another episode. It should not automatically mutate the voice, values, or visual identity of the show. Data informs the next decision. A person approves any change to what the show is.
A learning system should become more responsive without becoming less itself.
From a pipeline to a living loop
I have had a circle in my head for years. Most media diagrams were pipelines: content enters on the left and distribution happens on the right. The old Adobe Primetime lifecycle used a wheel, and that image stayed with me. A media business is not finished when a file is delivered. Viewing creates signals, those signals shape programming, and programming starts the next cycle.
Showspring applies that idea to agentic production. It spans audience and program analysis, idea generation, scripting, readout, images, video, editing, audio, rendering, publishing, and the watch experience. The finished episode is not the end state. It is the next piece of evidence.
Two feedback systems run inside that circle.
- The production loop evaluates the current artifact. Does the pitch fit the show? Does the script preserve canon? Does the character look right? Does the motion match the shot? Does the mix serve the scene?
- The audience loop evaluates what happened after release. Where did viewers stay? What did they replay? Which premise attracted attention? Which format earned another episode?
Neither loop should become an autopilot. A retention curve can tell you where attention changed; it cannot tell you what your show should mean. Self-evaluation can catch a continuity error; it cannot own the consequence of publishing. Data informs direction. It does not replace it.
Agents building agents that create video
One slightly strange part of this project is that agents help me build and harden Showspring, while another set of agents inside Showspring runs the creative workflow: idea agents, a script agent, canon and location resolvers, image and video generation, a multimodal audio planner, a shorts planner, and publishing preparation. Put simply, I am using agents to help build agents that create video.
But recursion is not the same as unrestricted autonomy. The build agents work inside repository rules, tests, release gates, and human-controlled deployment. The production agents work inside project scope, model policies, cost controls, persisted state, review points, and human-controlled publishing. I wrote about that operating model in AI Velocity Is a Control Problem.
The lesson carries across both layers: the more capable the agent, the more explicit the boundary has to become. Broad ambition works when the individual actions are narrow, observable, and reversible.
The output needs a real place to live
A production tool can easily become fascinated with its own machinery. Viewers do not care how many agents, models, queues, or fallbacks were involved. They care whether the result deserves the next minute of their attention.
That is why I built the watch experience into the system. The catalog makes the output concrete: recurring shows, releases, episodes, shorts, and a visible record of whether an idea became something worth watching. I am not pretending the quality is uniformly where I want it yet. Putting the work in front of viewers keeps the experiment honest in a way an internal demo never could.
Where I still want a person making the call
“Human in the loop” has become an easy phrase to say and a weak control to implement. I do not want a person clicking Approve at the end of a process they did not shape. I want the person at the points where judgment actually changes the outcome.
- Premise and point of view. What is worth making, and why should this show exist?
- Taste. Which of several plausible outputs is actually right?
- Truth and consequence. What claims can be made, what material can be used, and what should never ship?
- Exception handling. When does the process need to stop, change direction, or deliberately break its own pattern?
- Final authority. Publishing is a decision, not a queue state.
The machine can become better at presenting options and surfacing risk. It can even critique its own work before a person sees it. But self-critique is most useful as a filter for human attention, not as proof that the system has acquired taste.
The models will change. The system has to survive them.
Generative systems are still raw, and the models are moving fast. I do not know which provider will lead each stage a year from now. I do know that some of today’s workarounds will disappear and new failure modes will replace them. I have also learned not to treat an update as an automatic improvement: a more capable model can still be less faithful to a particular character or prop.
The durable advantage is the learning system around the models: interchangeable engines, retained canon, saved decisions, versioned prompts, evaluation, cost attribution, audience signals, and human review. It can absorb a model improvement, route around a regression, and use viewer response to inform the next program decision without letting that response dictate it.
The goal is not to control every generated pixel. It is to build enough control around uncertainty that creative intent survives the journey.
That is the practical reason I chose generative AI, and video in particular, as a service area. It forces orchestration, governance, evaluation, cost, memory, user feedback, change management, and human authority into the same system. The output happens to be a show. The operating model is much bigger than media.
And when it works, the reward is obvious: someone can attempt a story that previously required too much money, time, equipment, or specialist knowledge. The machine did not become the author. The author simply gained more reach.
The machine can take on more of the production crew. The human still owns the director’s chair and the result.
See the current output in the Showspring watch catalog, or explore how the production system works.
Discussion
Be the first to comment