IQBAL.EXE started with a simple, slightly ridiculous question: what if I didn't generate the images at all, and an AI wrote every frame as code instead?
No video model. No image model. Just a description of the scene, and an AI writing the code that draws it, frame by frame.
Why do it this way?
Because code gives you something generated video doesn't: total control and total consistency. BOT looks exactly the same in every shot, because BOT is the same code in every shot. Nothing drifts and nothing melts.
The workflow, step by step
- Write the episode as a plain script: scenes, lines, and what each moment needs to feel like.
- Describe one scene at a time to the AI: characters, positions, movement, timing.
- Let it write the code that draws and animates that scene.
- Render, watch, and correct. "Move BOT left." "Slower blink." "The text is clipped." Each note is a small change, not a re-roll.
- Voice and sound last, then the edit.
The blank page is only scary when you try to fill all of it at once. Fill one frame.
What I'd tell anyone trying this
- Start with one character and one gag. Not a world.
- Keep a style sheet: colours, sizes, fonts. Paste it into every request.
- Check every frame at phone size. Clipped text is the most common bug.
Episode 1 started from a single sentence:
Every episode since has been built the same way: one scene, one frame, one fix at a time.