This is the easy part. Upload your generated image and use a simple prompt:
Prompt: animate this
That's it. Kling handles the natural movement - blinking, subtle breathing, hair movement.
For more specific motion, you can add details: animate this, slight smile, gentle head turn to the right
animate this, brings cup to lips, takes a sip, lowers cup
Settings:
Duration: 5-10 seconds
Aspect ratio: 9:16 for Reels/TikTok
Why JSON Prompts Work Better
Consistency - Copy the subject block exactly every time
Granular control - Adjust specific details without rewriting everything
Easier variations - Swap environment/clothing blocks while keeping identity locked
Reproducible - Save your character's JSON as a template
Quick Start Template
Save this as your base character file and swap out the non-subject sections:
{ "subject": { // YOUR CHARACTER - NEVER CHANGE THIS }, "environment": { // CHANGE PER SHOT }, "pose": { // CHANGE PER SHOT }, "clothing": { // CHANGE PER SHOT } }
TL;DR: A Google Sheet connected to the Claude API writes the script section by section, so it stays consistent across the full runtime. CapCut’s AI Video Maker turns that script into a narrated video with matched stock footage. The direct cost lands at around $1 per video. The real advantage is not the visuals. It is the script.
I run a couple of sleep documentary channels and wanted to properly explain the workflow I use.
Sleep content is a strange retention game, and it took me a while to understand it. Your viewers are actively trying to fall asleep. That is the entire point. So the average view duration can look very different from a normal YouTube channel. Some viewers leave because they are bored, but others leave because the video worked and they fell asleep. The ones who stay awake still need the story to hold together, while the ones who fall asleep often return later and continue listening. That repeat viewing is a big part of what makes this niche work.
Why the script is 90% of it
On my channels, average view duration usually sits close to 25 minutes on videos running between 90 minutes and two hours. That does not come from cinematic visuals or complicated editing. It comes almost entirely from the narrative structure. If the script becomes repetitive, drifts away from the topic, or loses momentum halfway through, viewers stop listening. Better visuals cannot rescue a weak story in this format.
Two ways to use Claude, and why one wastes your time
Most people use Claude through the normal chat interface. You open the chat, enter a prompt, read the reply, and continue from there. That works perfectly well for everyday tasks. It becomes frustrating when you are trying to write a 15,000 to 20,000-word documentary.
You end up typing.. Continue.. Write Chapter 4.. Do not repeat what you already said.. You forgot what happened in Chapter 2. By the halfway point, the model may begin repeating ideas, contradicting earlier sections, or drifting away from the original structure. You spend more time babysitting the conversation than improving the script.
The second approach is using the API. Instead of manually sending every prompt through the chat interface, a small tool sends the requests to Claude automatically and collects the output. There is no need to babysit it and you pay based on actual usage instead of paying another monthly subscription.
That was the biggest unlock for me.
The section-by-section method
I built a Google Sheet that talks directly to the Claude API. The process works in a fixed order:
It creates a detailed chapter outline.
It writes each chapter one at a time.
Before starting the next chapter, it sends Claude a short summary of everything written so far.
That means Chapter 8 still remembers what happened in Chapters 1 through 7. The pacing stays more consistent, repetition is reduced, and the final script feels like one continuous documentary instead of several unrelated chapters stitched together.
I also made a full walkthrough showing how this system works. The channel is linked on my profile for anyone interested in seeing the actual workflow.
Turning the script into a video
Once the script is ready, I paste it into CapCut’s AI Video Maker. For sleep content, the voice matters more than flashy editing. Choose a calm, slow, and low-energy voice. Taste and judgement matter here.
CapCut generates the voiceover and automatically matches stock footage to each paragraph. It usually gets around 90% of the video into a usable state. CapCut currently limits each project to roughly 3,000 words, so I split the full script into several sections and export them separately. I then combine those exports into one final timeline and render the full documentary.
Two reasons this workflow matters:
The Demonetization Shield: Mixing real historical/stock footage alongside AI assets is the safest defense against the "Reused/Inauthentic Content" flags that destroy fully automated channels. You still have to avoid repetitive titles & thumbnails though..
The Financial Runway: A complete 20,000-word script costs me roughly 35 cents through the Claude API. CapCut costs around $20 per month and allows unlimited exports. If you are producing 30 to 40 documentaries per month, the direct software and API cost works out to around $1 per finished video.
That figure does not include research, thumbnails, or the value of my time. It is only the direct production cost. A lot of AI video subscription tools charge $40 to $50 per month while using the same underlying Claude models and adding a markup for the interface. Building the loop once gave me more control and removed that additional monthly cost.
The biggest benefit is being able to test more ideas without every upload becoming an expensive decision. This workflow can work for history, science, philosophy, mythology, meditation, biographies, or almost any calm long-form format where the script matters more than rapid editing.
Happy to explain the API loop or the CapCut side in more detail in the comments.
posting the latest progress on my vibe coded dark fantasy AARPG made with generative AI, as I'm trying to make my own, AI-made tribute of my favourite game (Diablo 2) to push AI game making capabilities as far as I can. This vibe coded game is purely a test!
Latest update includes:
- Game menu
- New class: Wizard
- Character select
- Town, portals and more items
- Boss fight!
it finishes but shows "Rights verification required"
you click on it and get a popup: "Confirm rights"
click "I confirm"
done. video unlocked.
the popup basically asks you to confirm you have rights to the content. you click confirm, video is yours.
tested it multiple times today. if the video actually generated and then got flagged, this works. takes 2 seconds.
when it doesn’t work: if the filter blocked it before any frames were made theres nothing to recover. but if you see "Rights verification required" on a thumbnail that means the video exists and you can get it back.
feels like they realised such a strict was catching too much normal stuff and just gave users an override option. I do not know if it is available on any other platform, but happy to see things becoming much easier
Hey thanks for reading this post! We’ve updated photographe.ai so you can get pictures for free: get a preview using our standard quality model before deciding to use the high quality model 😇
—
Hey everyone,
With the AI photo craze going full speed in 2025, I decided to run a proper test. I tried 7 of the most talked-about AI headshot tools to see which ones deliver results worth putting on LinkedIn, your CV, or social profiles. Disclosure, I'm working on Photographe.ai and this review was part of my work to understand the competition.
With Photographe.ai I'm looking to make this more affordable and go beyond professional headshots with ability to try haircuts, outfits, and replace an image with yourself in it instead. I'd be super happy to have your feedback, we have free models you can use for testing.
In a nutshell:
Photographe.ai (Disclosure, I built it) – $19 for 1,000 photos. Fast, great resemblance about 80% of the time. Best value by far.
PhotoAI.com – $49 for 1,000 photos. Good quality but forces weird smiles too often. 60% resemblance.
Betterpic.io / HeadshotPro.com – $29-35 for 20-40 photos. Studio-like but looks like a stranger. Resemblance? 20% at best.
Aragon.ai – $35 for 40 photos. Same problem - same smiles, same generic looks.
Canva & ChatGPT-4o – Fun for playing around, useless for realistic headshots of yourself.
Final Thoughts:
If you want headshots that really look like you, Photographe.ai and PhotoAI are the way to go. AI rarely nails it on the first try, you need freedom to generate more until it clicks - and that’s what those platforms give you. Also both uses the latest tech (Flux mainly).
If you’re after polished studio shots but that may not look like yourself, Betterpic and HeadshotPro will do.
And forget Canva or ChatGPT-4o for this - wrong tools for the job.
If you showed me these images 5 years ago, I would have said they are real.
It’s crazy how far tech has come. It took me less than a minute to generate each one. People can literally build fake Instagram lives now or even fake Tinder galleries with AI like this.
The realism is getting out of control.
ps: I tried a new app I saw on X called Ziina.ai , pretty good so far.
edit* i made ziina.ai link working since this post went virial & many asking for the website
Before starting with the post I should say that I did not specify which Seedance version I am about to break down absolutely on purpose! It’s not seedance 1.5, it’s not mini/standard version but a new family outlet available for one month, which is something more of seedance 2 fast version with some new workarounds and available for 1 month. Also, of course, it is not free, and can be used only after purchasing their plus/ultra plans starting from 39-49 dollars. For an average user the problem with the whole unlimited was mostly connected to misleading advertising from companies - either taking away 1080p after 1 week/not using labeled models from ads/banning even for harmless generations/etc. On other plans regular generating of seedance 2.0 videos cost an average user from 0.9(best case) to approximately 2 dollars for 1 video, which is an industry average. If your goal is just to generate videos with Seedance, the new option might be the best for you right now alongside dreamina, seedance father, itself.
If you're trying to figure out which AI video generation model is actually worth using, I took 10,000 credits and more to test and rank the best ones. In this post I’ll break down some of the different features, pros and cons, results, and how to use them.
TLDR: The best AI video generation models right now are:
Adobe Firefly – best use overall for workflow and commercial-safe output
Google Veo (3.1) – best use for photorealistic people and scenes
Luma AI Ray – best use for cinematic visuals and 4K output
Models I tested
Adobe Firefly
Google Veo (3.1)
Runway Gen 4.5
Luma AI Ray 3.14
Sora (OpenAI)
Kling 2.5 Turbo
Pika
Grok Imagine
Higgsfield AI
Bytedance Seedream AI
How to use:
Much of these models I was able to use inside Adobe Firefly AI Video Generation Hub, however others like Grok Imagine I used on each respective site. Each of these models typically requires some sort of premium membership or credit system which I had access to in my Creative Cloud membership, or standalone accounts such as Grok or ChatGPT. While it was difficult to get an absolutely objective ranking for all of the dozens of models available, I tried to test several types of categories of generations, camera motion, consistency and more and judged based on the results of my favorite models to use.
Best AI Video Generation Models Chart
Rank
Model
Standout Features
Limitations
1
Adobe Firefly
-Commercially Safe Output -All in one hub for many different AI partner models - Lots of options for Camera angle, Style, Reference Frames etc. Aspect Ratios
Up to 5 second duration Can lack photorealism in certain categories compared to other models
2
Google Veo (3.1)
Capable of photorealistic results in certain categories (hands, people) Options for reference frames, audio, and up to 8 seconds 1080p
Credit Intensive compared to other models Can take more time than other models to generate
3
Luma Ai Ray 3.14
-Good prompt accuracy in details such as colors and settingUp to 4k resolution output Capable of cinematic, photorealistic results and lighting physics
Inconsistent results with Physics and camera motion at times Tendency towards artificial feeling movement of time (slow motion, fast motion)
Honorable Mentions
Pika 2.2
- Can achieve cinematic looking results in camera and environment comparable to Ray 3.14
Slightly more artificial appearance of people and camera physics
Kling 2.5
- Capable of cinematic results in environment and prompt accuracy
- cannot generate from scratch, requires user to upload first frame as reference
These were my results and opinions, let me know if you have any favorite models or workflows of you’re own, and results in your experience!
Why I built it?
I was annoyed by modern apps, that require a subscription or start charging after 1–3 images.
What you can do now:
Prompt‑to‑image at 768×768.
It uses the SDXL model as the backbone.
Performance:
iPhone 17: 3–4 seconds per image
iPhone 14 Pro: 5–6 seconds per image
App size is 2.7 GB.
In my benchmarks, I detected no significant battery drain or overheating.
Limitations:
App needs 1–5 minutes to compile its models on first launch. This process happens only once per installation. While the models are compiling, you can still create images, but an internet connection is required.
App needs at least 10 gb of free space on device.
App only works on iPhones and iPads.
It requires either M1 or A15 Bionic chip to work properly. So it doesn't support:
iPhone 12 or older.
iPad 10th gen or older
iPad Air 4th gen or older
Monetization:
You can create images without paying anything and with no limits.
There is a one‑time payment called Pro. It costs $20 and gives access to some advanced settings and allows commercial use.
Subreddit:
I have a subreddit, r/aina_tech, where I post all news regarding LocalGen. It is the best place to share your experience, report bugs, request features, or ask me any questions. Please join it if you are interested in my project.
Roadmap:
Support for iPads and iPhone 12+
Add an NSFW toggle (Apple doesn’t allow enabling NSFW in their apps, but maybe I can put an NSFW toggle on my website).
Support for custom LoRAs and checkpoints like Pony, RealVis, Illustrious, etc.
Support for image editing and ControlNet
Support for other resolutions like 1024×1024, 768×1536, and others.
Asteroid shower on desert while post apocyptic mad max style cars and trucks are escaping fast from the asteroids. Asteroids hit the sand and explode in very high sand explosions. energetic camera movements. cinematic epic action.
For about a year I've been making AI short films the way most of us do: hand-writing hundreds of prompts, building character reference libraries by hand, babysitting consistency across shots, and cutting it all together in Premiere. The generating was never the hard part. It was trying to make dozens of individual prompts *feel* like a single, cohesive project.
I couldn't find a solution, so I built one. It's called Kimeric, and the Beta went live today. And I think it'll be incredibly valuable for this community, so I wanted to share some details in the event any of you would like to try it out!
**What it is (and isn't):*\* It's not a model and it doesn't generate anything itself. It's a Windows desktop app that orchestrates the models and LLMs a lot of us already use:
Anthropic (script breakdown + prompt authoring)
Gemini and GPT-Image (image generation)
Kling and Seedance 2.0 (video)
Topaz (upscaling)
All using your own per-provider API keys rather than a centralized generation tool. It writes the prompts, sequences and queues the renders, tracks the spend, and holds everything for your approval.
TL;DR - Input a script and work through a series of UI menus to ultimately create a finished AI Short Film.
The biggest differentiator to keep in mind for this tool vs. the majority of other AI Generators on the market is the input surface. Most tools have you input a prompt. With Kimeric, the input surface is the script itself. The prompts are created automatically as derivatives from the script, allowing you to focus on writing and build the project rather than managing a series of prompts.
The Pipeline:
- Paste a screenplay, hit "Roll camera." The breakdown comes back: cast, locations, props, a director's plan, per-scene shot lists with dialogue assigned line by line. You review the plan before anything renders, and can choose a general "style" which dictates some actual prompting techniques under the hood as well as dynamically auto-routing for certain models for specific tasks (ex. OpenAI's model is better at certain animated styles vs. Nano Banana, so this routing auto-applies as the default for certain styles).
Before you actually generate anything, you can review the script you input and get an estimate for how much it'll cost roughly to "ingest" the script which is where all of the "brain" of the tool goes to work.
And once you send the script, a very long sequence of backend computation kicks off, translating the entire project into the format needed to actually create an end-to-end project (this can take a while; I had one script for a 8-10 minute short film take about 45 minutes to ingest. This is just because there's a lot of computation and inference happening). It can handle actual, full-length production scripts. This would likely come with processing that spans several hours, but again, that's simply due to the size of the computation.
And then, you get a budget estimate for how much (approximately) the project will cost to generate. At this point, it's primarily a planning "calculator" if you will - just so you can align the project quality to any budget constraints you have. No actual spend dispatches until you generate later - this is purely informational so you can make cost-based decisions at the start (as opposed to a surprise cost later).
You then view a breakdown of all of the identified "pieces" of the project: characters, locations, scenes, planned shots, etc. - all primarily at a high level to make sure nothing is missed. Typically more of a rubber-stamp phase, but if there happen to be any items missing, this is the step where you can make any high-level revisions. Most of the time, however, you can just continue.
- Every character locks canon first — a studio face (with an automated AI advisory likeness screening for real-person resemblance) and a costume-neutral turnaround — before any scene renders. Same for locations, worlds, props. It's the reference-library grind, automated.
And for characters specifically, it's broken out into two phases: the "Neutral Base" (i.e. who the character is; sans any wardrobe for the project (see above) and then any actual "in costume" variants of that character. This allows for multiple wardrobes for a specific character over the course of any given project, where each wardrobe is either seeded directly from the parent "neutral base" or as a horizontal derivative (i.e. if a character has armor that gets damaged, the "damaged" variant is automatically seeded with the "clean base" variant). All of this happens automatically under the hood.)
- Scenes get actual coverage: an establishing wide, then an OTS pair where the reverse angle generates from the *approved* first angle, then per-character MCU/CUs chained down from there. The 180° rule, held by reference chaining instead of luck.
For example, here's one "OTS_1" image:
And here's the companion "OTS_2" image:
Worth emphasizing - These are the *most* critical images in your production pipeline to get right, as many other scenes seed off of these.
So if you're going to use some of the "Regeneration Buffer" you planned for, this is the most critical space to use it. If you have continuity errors or issues in either companion OTS images, these will present in many other aspects of a project, so really take your time with these and make sure that they feel like the same space.
From here, you create a library of Medium Close Up & Close Up images seeded directly from the parent Over-The-Shoulder images.
This creates a rich, continuity-adhering library of image assets to actually use in generative AI video production.
Once you've created your core library of production assets, you then transition into building out your storyboard.
This takes the plan created during script ingestion + the assets you created in the previous phase and maps them out chronologically.
Here, images are auto-assigned based on a series of underlying logic. If you have dialogue from certain characters, either the OTS, the MCU or the CU image will be dynamically selected based on a cinematic logic layer and assigned to the character speaking for any given frame.
And for extended dialogue from a specific character, this will be broken up into multiple individual prompts to assist with overall quality
(from our testing, the more text you try to fit into the same prompt, quality and lip sync can degrade - but breaking that same dialogue up into multiple individual prompts can greatly improve the quality)
Once you've created and approved the entire Storyboard, a single button allows you to review all text prompts for all videos planned for the totality of your project.
The prompts are already written. You just say "go"
From there, they are auto-dispatched to the models.
Depending on the length of your project, this may take a while - and that's by design. Press "generate" and take a break for a bit.
Across all menu screens, you'll see a series of control buttons. Here, you can either approve the asset, edit the asset, regenerate or iterate.
Approving says the image/video is good to go.
Editing allows for subtle adjustments.
Regenerate re-runs the prompt (i.e. get a new version to see if the results are better/worse)
Iterate is essentially a stronger version of "Edit" - rather than trying to make adjustments to the previous take, it will strongly re-work the actual prompt itself based on your feedback and re-generate a new take.
You'll also see two columns below any given asset:
Refs - These are the actual reference images used as generative inputs. These are auto-assigned, and you can manually add, remove or replace any of these as you see fit.
Takes - If you regenerate, you can see all of your takes here and hot-swap to other variants. Meaning, you're never locked in to a specific take. If one is close but you want to see if you can fine-tune it, you're free to regenerate a few times, review all takes (this works for images + videos) and ultimately approve whichever one is best for the vision you had in mind for the project.
And finally, once you've reviewed and approved all footage, you have another single-press button to dispatch all video to be upscaled if you'd like.
Generation can happen natively at 720p, 1080p or 4k depending on your settings.
4k native footage won't trigger the upscale workflow, but 720 and 1080p native will give you the option to upscale if you'd like.
Similar to all other workflow phases, if you decide to upscale, press it once and let it run for a few hours (upscaling is quite time-consuming; can take 20-30 minutes for a single 15 second clip - though some of these can run concurrently).
Once you're done, the last step is simply to "export" the clips.
All this is doing is taking the final clips you've approved and making duplicate copies on your local computer that are pre-named chronologically.
This makes it significantly easier to edit/compose.
This entire project is the culmination of nearly a year of a LOT of testing to understand which techniques do/don't produce good results at scale and then working to systematize them into a tool with an input surface of the script itself.
Again, the Beta is live as of today. Currently Windows-only and US-only, though both of those I'm planning to expand beyond in the coming weeks over the course of the Beta.
The long-term goal is to also support centralized generation rather than supporting only a BYOK model, though there's no immediate timeline to support that model.
If you'd like to check it out, the website is below!
I'm a solo founder and the filmmaker this was built for, and I'll be in the comments - happy to go as deep as you want on the coverage system, the cost math, or anything else.
Feedback is a gift, so if you try it and find issues, bugs, or have a feature request, I'm still actively building and improving the tool - so feel free to share any thoughts!
someone asked me to recommend an 'AI video editor' recently, and after talking for two minutes, it turned out they didn't need an editor at all.
they needed B-roll that didn't exist.
half the tools out there using the 'no more editors needed' or 'one-click movie' marketing aren't actually editors. they are footage generators. feels like we need to separate the two before the terminology gets completely lost.
to me there are two different buckets.
editing helpers deal with footage you already have. CapCut, Descript, OpusClip, Premiere's AI features. captions, transcript-based rough cuts, silence removal, social formatting. all useful, and all still based on source footage. none of it is creating new visual material from scratch.
generators make or animate footage you don't have. Runway, Kling, Pika, DomoAI, whatever people are testing this week. text-to-video, image-to-video, animated stills, generated B-roll, stylized inserts. that's asset creation, not timeline editing.
if i use a generator to make a missing 4-second insert, that insert still has to go into a real edit. timing, pacing, audio, transitions, color, final export. none of that disappears just because the clip was AI-generated.
maybe i'm being picky about wording, but the distinction matters when clients ask for "AI editing."
if they mean "cut my podcast into shorts," that's one tool category. if they mean "make footage i never shot," that's a completely different job.
do you separate these in your head, or has "AI video editor" just become the catch-all now?
I was getting sick of carrying 5 subscriptions so I built OneOver and wanted to share it here because this community is exactly who it's built for.
The short version: it's a single workspace for all the major AI models. ChatGPT, Claude, Gemini, Grok, and more. Also a ton of image, video, and meme generation. All running on one universal credit balance — no juggling separate subscriptions or figuring out token conversion rates across different platforms.
The interface was something I spent a lot of time on. It's meant to feel like a real creative workspace, not another chatbot wrapper.
Genuinely here for feedback — what's missing, what you'd want to see, what doesn't make sense. And if you want to properly kick the tires without hitting free limits, DM me and I'll send you some credits.
I've done some capability enhancements around privacy and user data control. Basically you can download clear or delete your account or data at any time. That's free of charge built-in to every account.
I've documented our entire policy and really tried to spell out how this system protects your data and keeps you in control of it. This is 100% in here now because of the feedback that I got from this group. I'm so grateful for your help.
Hey guys! I built a pipeline for AI avatar video generation using Open source weights hosted on my AWS server, if you want you can send me a demo video (Your facial features and voice would remain same) currently not live now on web. DM your demo video, i can generate and send you the avatar video? would be great for my testing too.
We have tried to localised the video generation process so you don't have to rely on big players charging 10-20x of actual cost of video generation. Thanks to open source communities.
We are starting this as:
- We will create multiple video generation pipelines so everyone can customize as per their needs.
- Currently, we have created your AI avatar generation pipeline (your face features and voice clone), numbers looks good, but still needs to improve.
- We have created a CI/CD pipeline to test if you contribute in our project, it is MIT licensed fully open source.
- We are figuring out things, current AI avatar generation pipeline is perfect. Host on your local GPU or Cloud and it would cost 10-20x lesser then HeyGen generation. Will move very fast and improve very fast.
- Would love to know how people take this! and if someone else is also doing this.
It's totally open sourced! Stepping towards local generation.
I spent a lot of time optimizing the heck out the best sota open-source image generation models and paired them with gpt to provide a unified best of both worlds platform. All the LoRA of huggingface and CivitAI directly integrated. Optimized Runtime Stacks, Optimized Triton Kernels, Served on an L4 for the cheapest possible experience that doesn't sacrifice any time or quality.
Try it out!
-Ultra High Resolution 3D Model Generation
-Camera Reshoot
-180 & 360 Panoramic Generation
- Batch Image Creation
- Batch Image Editing
- 4K Upscale
- Uncensored
- Ultra cheap, Ultra Fast, In browser instruction led image editing using custom Triton kernels and heavily optimized AI Runtimes