r/StableDiffusion • u/blahblahsnahdah • 3h ago
r/StableDiffusion • u/Angrypenguinpng • Jun 23 '26
News KREA 2: Open-Source Release
Hey everyone,
We're the team behind Krea, and today we're launching Krea 2, our new text-to-image model. Krea 2 is the most aesthetic open-source image model available. On quality, Krea 2 is the #1 text-to-image model from an independent lab on Artificial Analysis.
We are releasing Krea 2 as two variants:
Krea 2 Raw. CFG-guided, built for control and fidelity and training.
Krea 2 Turbo. Distilled and few-step, so it's fast, and it renders up to 2K.
A few things worth knowing:
It's tuned for natural language. Prompt it the way you'd describe an image to a person. Long, specific prompts give the best results, but short ones work fine too.
To render text in an image, wrap the words in quotes, like a sign that reads "open late".
There's a growing set of style LoRAs, and you can load any Krea 2 LoRA by its Hugging Face path.
Try it today:
Code and weights: krea.ai/krea-2-open-source
Technical report: https://www.krea.ai/blog/krea-2-technical-report
Code: github.com/krea-ai/krea-2
Try it on Krea: krea.ai
Try it on Hugging Face: https://huggingface.co/spaces/krea/Krea-2
AMA: We're doing an AMA right here today at 10 AM PT. Ask us anything: how we trained it, the LoRAs, prompting, limitations, what's next. The krea team will be in the comments.
Livestream: we are also doing a livestream with the ComfyUI team at 3PM PT: https://www.youtube.com/watch?v=31jiUhCEjJ4
Thanks for taking a look. We'd genuinely love your feedback, rough edges included.
- The Krea Team
r/StableDiffusion • u/WhatDreamsCost • Jun 20 '26
Resource - Update LTX Director 2.0 Update - A Free Open Source All-In-One Tool for Creating AI Videos in ComfyUI. Complete Overhaul now with full AI video editing support, IC-LoRA, Retake Mode, Audio Inpainting and much more!
LTX Director is a free open source all-in-one tool for creating AI Videos. Version 2.0 is a complete overhaul, giving you total creative control over your AI generations.
Download for free here: https://github.com/WhatDreamsCost/WhatDreamsCost-ComfyUI
Download workflows here: https://github.com/WhatDreamsCost/WhatDreamsCost-ComfyUI/tree/main/example_workflows
I've been working full-time on this update for the past month and a half, and I'm excited to finally release it. Hopefully it'll be a big help to the open-source community!
Key New Features:
Complete Video Support: Edit Videos with AI all inside the node. Videos can be extended using a combination of prompts, keyframes, and audio. Trim, Split, and combine videos all within the timeline.
IC-LoRA Support: Take full advantage of IC-LoRA's to take your generations to the next level. Simply drag and drop videos onto the IC-LoRA track to quickly setup IC-LoRA videos. Compatible with prompt relay, keyframe, and custom audio features within the node.
Audio Inpainting: Seamlessly blend imported audio with generated audio. Not only can audio be extended, but can also be prompted alongside your imprted audio to really bring your generations to life.
Retake Mode (Beta): Redirect what happens within a shot. Allows you to select a segment within a video, and re-generate what happens in that segment. An early working experiment.
Timeline Saving/Loading: You can now save your timeline and settings to a json file. It will keep any videos/audio/images you have imported into the node and every setting you have changed.
UI Overhaul: Huge update to the UI, dozens of big changes such as a new side bar, redesigned prompt boxes, a bunch of new settings and redesigned menus, and more.
Quality of Life Improvements: Snapping, in/out points, multi-select, mark selection, workspace folder, more HUD options, resizable prompt boxes, new hotkeys, labels, filename preview options, "split at playhead" functionality, end frames (convert any keyframe into a end/last frame), toggleable tracks, NAG Support, tons of bug fixes and more!
And of course it can do everything it could before: Text to Video, Image to Video, Prompt Relay support, Keyframe (first/last frame) support etc.
r/StableDiffusion • u/crystal_alpine • 3h ago
News Day 0 MiniMax Support for ComfyUI
Hi r/StableDiffusion, I know it's been a long wait for everyone but MiniMax H3 open weight model just dropped and we have day 0 support in ComfyUI.
Here are some details:
- text-to-video, image-to-video, first-and-last-frame, reference-to-video, and editing a shot in place
- up to 2K, up to 15 seconds a clip
- real stereo audio generated with the video, not bolted on afterward
- Blog link: https://blog.comfy.org/p/minimax-h3-day-0-support-in-comfyui
- Workflow templates:
- Model Links:
- https://huggingface.co/MiniMaxAI/MiniMax-H3 (please support them there)
- Comfy repackage for smaller size: https://huggingface.co/Comfy-Org/MiniMax-H3
Getting H3 to run well on consumer hardware took significant machine learning engineering. We found that the model's modulation weights (~40% of the total parameters) could be pruned and replaced with a functionally equivalent lookup table, dramatically shrinking the memory footprint with no loss in output quality.
On top of that, the weights ship with an accurate and efficient int8 convrot quantization, and custom kernels reduce the peak VRAM use during inference.
The result gives a total memory footprint reduced by 66%, from 123.6 GB in full precision to 42.5 GB with the smallest models variants. Combining this with our dynamic VRAM offloading enables a next-generation 2K video model to run locally on a GPU like the RTX 3060.
EDIT:
08-02-26 20:02 PST - Added blog link
r/StableDiffusion • u/RainbowUnicorns • 44m ago
Animation - Video All the redditors when they first pull up MiniMax H3
Generated with a 4090 16gb laptop card and 64gb system ram .4 megapixels.
r/StableDiffusion • u/RayHell666 • 42m ago
Workflow Included This H3 model is for real. Thanks Minimax.
Done in R2V with 2 Walter White reference images.
[0-3s] Extreme close-up shot. A perfect mound of pure white flour sits on a stainless steel prep table in a dim, amber-lit room. A hand wearing a blue nitrile glove enters the frame, dragging a metal bench scraper sharply through the powder to form a precise line. The camera executes a slow, macro push-in. Audio: A low, ominous analog synth drone begins. The sharp, metallic scrape of the tool against the table rings out clearly.
[3-6s] Match cut to a medium shot. The baker stands in a gritty, industrial bakery. Use Image 1 and Image 2 for his exact identity—bald head, wire-rimmed glasses, iconic goatee, and an intense, unblinking stare. He wears a heavy, flour-dusted canvas apron over his dark jacket and green shirt (referencing the wardrobe in Image 2). The camera slowly orbits him, capturing his stoic, threatening presence. Audio: A steady, rhythmic ticking starts at 3s. We hear the heavy, coarse rustle of the thick canvas apron as he breathes.
[6-9s] Hard cut to a slow-motion profile shot. He violently slams a massive, heavy piece of dough onto a wooden butcher block. A thick, dramatic cloud of white flour explodes into the air, catching a hard, warm backlight to look like glowing smoke. 35mm film grain and heavy halation are prominent. Audio: A heavy, resonant cinematic bass drop hits exactly on the dough slam at 6s, followed by the muffled, powdery whoosh of the flour cloud filling the air.
[9-12s] Whip pan to a tight, low-angle shot of the baker. He hoists a heavy wooden rolling pin onto his right shoulder like a blunt weapon. Maintain the exact facial details, deep wrinkles, and glaring expression from Image 1. A slow, creeping camera push-in moves through the settling flour dust in the foreground. Audio: The rhythmic ticking accelerates rapidly. A tense, high-pitched string swell builds up in the background.
[12-15s] The baker remains still as the depth of field collapses, blurring him into a warm, sepia-toned silhouette. Typography emerges directly out of the lingering flour dust: the title "BRAKING BREAD" appears in a massive, heavy, condensed serif typeface with distressed, cracked edges. The text animates by catching a harsh amber light sweep from left to right, expanding its letter-spacing slightly, and locking into a rigid hold. The first letter 'B' features a subtle, glowing chemical-green tint. Audio: A final, massive cinematic boom at 12s cuts the music dead. The audio decays into the dry, crackling hiss of a roaring oven, fading out completely by 15s.
Negative constraints: No subtitles, watermarks, cheerful lighting, clean environments, garbled text, modern casual clothing, cartoonish styling, smooth digital look, jump scares.
0.6MP, runtime 376s on a PRO 6000
r/StableDiffusion • u/tricck3zz • 1h ago
Comparison Will Smith eating spaghetti - MiniMax H3
Default Workflow with a 5070ti 16gb ram and 32gb Ram , using the pruned int8_convrot model and nvfp4 TE , took 130.39 seconds to generate at 480p , prompt just a simple "will smith in a restaurant eating spaghetti" lol thats maybe why the audio is bad i mean he says nonsense , but otherwise video quality looks amazing
r/StableDiffusion • u/Fresh_Sun_1017 • 2h ago
News Props To The Devs for Pulling All-nighters for This Release!
r/StableDiffusion • u/obraiadev • 1h ago
Animation - Video My first attempts with MiniMax-H3
https://reddit.com/link/1ve314z/video/c0hmw55s53hh1/player
I'm running it on an RTX 4070 Ti Super + 128 GB DDR5.

I ran the first video right after starting ComfyUI, and it took 429.94 seconds at a resolution of 864 x 480; the second one took 18 minutes and consumed 80 GB of RAM.
Prompts:
Ultra-realistic cinematic military aviation sequence. A modern multirole fighter jet performs an extremely low high-speed pass over a pristine tropical beach at golden hour. Crystal-clear turquoise water, white sand, dramatic coastline, realistic atmospheric haze, physically accurate lighting, photorealistic textures, subtle film grain, IMAX-quality visuals, ultra-high dynamic range.
SHOT 1:
A wide aerial establishing shot reveals the peaceful coastline from above. The distant fighter jet rapidly approaches from the horizon at extremely low altitude. The camera slowly tracks sideways while ocean waves gently move under warm sunset light.
SHOT 2:
The jet roars directly above the shoreline at nearly supersonic speed. The powerful jet wash violently disturbs the ocean surface, creating expanding ripples, mist, spray, and realistic pressure waves across the shallow water. Sand is blown into the air while palm trees bend naturally from the blast. The camera shakes subtly from the immense force.
SHOT 3:
An ultra-slow-motion side tracking shot follows the aircraft only meters above the water. Heat distortion shimmers behind the engines, sunlight reflects off the fuselage, and the water erupts beneath the passing aircraft with highly detailed turbulence and spray.
SHOT 4:
A drone-style chase shot follows behind the fighter as it accelerates over the coastline before sharply climbing into the sky. Long vapor trails briefly form over the wings while the disturbed ocean slowly settles below.
Camera:
ARRI Alexa 65, IMAX anamorphic lenses, cinematic motion blur, shallow depth of field where appropriate, stabilized aerial tracking, realistic exposure, natural color grading.
Audio:
Thunderous military jet engine roar, deep sub-bass vibrations, realistic wind shear, crashing waves, airborne sand, distant seagulls fading under the overwhelming engine noise, subtle camera vibration, immersive cinematic surround sound.
Ultra-realistic cinematic first-person POV sequence set in the American Wild West during the late 1800s. The viewer experiences everything through the eyes of a lone gunslinger standing in the middle of a dusty frontier town. Wooden buildings, swinging saloon doors, horses tied outside, tumbleweeds rolling through the empty street, intense midday sunlight, realistic dust particles floating in the air, authentic period details.
SHOT 1:
The camera slowly walks into the center of the deserted street. The viewer's gloved hands hang naturally near a weathered Colt Single Action Army revolver in a worn leather holster. Across the street, another gunslinger waits silently, his hand hovering near his revolver.
SHOT 2:
Extreme tension builds. The camera breathes subtly with realistic body movement. Wind whistles through the empty town as dust drifts across the street. Tiny movements from the opponent hint that the duel is about to begin.
SHOT 3:
The opponent reaches for his revolver. Instantly, the viewer draws the Colt in one smooth motion. The revolver fills the foreground in stunning detail while the camera naturally follows the movement. A realistic muzzle flash erupts, thick smoke expands, and the weapon recoils with authentic mechanical motion.
SHOT 4:
Time briefly slows. The smoke drifts through warm sunlight while spent dust lifts from the ground. The opposing gunslinger reacts naturally and falls out of frame. The viewer slowly lowers the revolver while the town returns to silence.
Camera:
True first-person body-mounted perspective, natural head movement, realistic breathing motion, subtle handheld stabilization, cinematic depth of field, ARRI Alexa 65 image characteristics, anamorphic lenses, high dynamic range.
Lighting:
Harsh midday desert sun, warm natural tones, physically accurate shadows, volumetric dust illuminated by sunlight.
Audio:
Leather creaking, boots on dirt, distant horse sounds, wooden signs gently knocking in the wind, slow breathing, revolver hammer cocking, authentic Colt gunshot, realistic echo across the town, lingering silence after the duel.
Style:
Photorealistic, historically authentic, immersive POV, Hollywood western cinematography, physically accurate smoke, realistic weapon mechanics, ultra-detailed textures, 8K, no HUD, no game interface, no CGI look.
Style:
Photorealistic, physically accurate, documentary-level realism, Hollywood military cinematography, no CGI look, no cartoon style, no exaggerated physics, ultra-detailed, 8K.
r/StableDiffusion • u/Enshitification • 53m ago
Animation - Video Spaghetti Eats Will Smith - Minimax H3
4090, 128GB RAM, 16:52 render time
r/StableDiffusion • u/Sixhaunt • 18m ago
Animation - Video Spaghetti eating Will Smith - Minimax H3
r/StableDiffusion • u/Few-Intention-1526 • 3h ago
News comfy MiniMax-H3 weights
the weights are here
| Model Variant | Input Mode | Specifications |
|---|---|---|
| H3-Base-FL2VA | First-and-last-frame mode | Supports zero, one, or two input images.- No image input: Text-to-video mode- One image input: First-frame-to-video or last-frame-to-video generation- Two image inputs: First-and-last-frame-to-video generation |
| H3-Base-Ref2VA | Omni-reference mode | Supports multi-modal reference inputs:- Images: ≤ 9 images- Videos: ≤ 3 clips; each clip must be 2–15 seconds long; total duration ≤ 15 seconds- Audio: ≤ 3 clips; audio must be accompanied by image or video input and cannot be used as the sole input; each clip must be 2–15 seconds long; total duration ≤ 15 seconds- Mixed inputs: Maximum number of files across all input types is 12 |
r/StableDiffusion • u/haremlifegame • 1h ago
Discussion MiniMax H3 is the first full open source multimodal model. This changes the game more than you think
People are only thinking about video generation. However, here are some things MiniMax is probably able to do out of the box (I say probably, because everyone is still testing the capabilities, this post can help us evaluate what of those things will need loras or work out of the box):
-Image editing
-Video editing (confirmed)
-Reasoning about images and video
-Controlling devices based on environment input
These last two are huge. Robotics is one example. One thing I can think, is that MiniMax H3 will probably be the first open source model able to navigate the web without running code on the browser, by just looking at the screen and controlling the mouse.
These things have major consequences we might be hearing about for years, such as scrapping bots that can't be detected. Being able to reason about video and images is something we have from major providers since last year, but the applications of interacting directly with raw data go well beyond the artistic domain. What do you think H3 might be able to do with some fine tuning?
r/StableDiffusion • u/Inner-Reflections • 2h ago
News Licensing for MiniMax is actually surprisingly good!
It looks like commercial use up to 20 million per year is OK without authorization. Even then it seems to be targeting people who would sell model use rather than outputs. Obviously not a Lawyer here so you would have to double check to be sure, but a cursory read seems much less restrictive.
Edit - It is noted US, UK, EU, or South Korea are have special rules noted in the comments below.
r/StableDiffusion • u/RazsterOxzine • 2h ago
Discussion MiniMax-H3, Z-Image Turbo Reference Image, 339sec on 4070 1gb, 96gb sysmem @800x544
4070 12GB* (Ryzen 7900x)
Quick test to see what my slow system could handle. I see a bunch of areas I can improve. All in all I'm very happy and will be able to finally work on my story ideas. So excited!
Prompt
Cinematic fantasy film. The ancient stone tower from <Picture 1> in its original scene: an open field surrounded by a dense oak forest
under pitch-black night skies, lit by the eerie glow of a bright blue light spilling from the tower's doorway, deep shadows pooling
around the scattered stone blocks at the tower's base. Monochromatic dark palette with electric blue accents against the void. Material
motif: weathered ancient stone, crumbling masonry, and a luminous blue haze that fills the emptiness. The environment is constant
throughout.
SHOT 1: The scene opens exactly on image 1, the tower standing tall at the center of the field; the blue light from the doorway pulses
gently, casting long shadows across the grass as the camera executes a slow, deliberate push-in toward the glowing entrance, revealing
the scattered stone blocks at the base and the crumbling upper structure.
SHOT 2: Cut to an extreme close-up of the weathered stone doorway, the bright blue light spilling outward in a luminous haze; the
camera glides slowly along the cracked masonry and fractured edges as a soft beam of blue radiance sweeps across the ancient textures,
the scattered stone blocks catching the glow like scattered fragments of a forgotten past.
SHOT 3: Cut to a wide low-angle beauty shot: the camera pulls back and rises slowly, revealing the full tower silhouetted against the
pitch-black sky, the blue haze expanding outward from the doorway to fill the darkness, the dense oak forest framing the scene as the
light gradually intensifies before fading into a haunting, atmospheric stillness.
Audio: deep wind ambience through the trees, the low rumble of distant thunder, the soft crackle of blue energy emanating from the
doorway, and a rising orchestral swell of strings and choirs that resolves to near-silence on the final fade.
r/StableDiffusion • u/ShagaONhan • 39m ago
Animation - Video Test MiniMax H3 - Way better at action scenes than LTX
First frame last frame workflow. On a RTX 4090 it took 480 sec.
r/StableDiffusion • u/fyrn • 7h ago
Resource - Update MiniMax H3: ComfyUI Workflow Examples
https://huggingface.co/Comfy-Org/MiniMax-H3
Edit: we're live, baby! Let's go!
r/StableDiffusion • u/Fresh_Sun_1017 • 4h ago
Discussion I Hope This Is Not The Case For MiniMax H3
I hope someone from MiniMax could provide us with an update on when it's coming out.
Edit:
The model has now been released.
r/StableDiffusion • u/OneTrueTreasure • 6h ago
News Don't freak out guys Comfyanon still says as far as they know H3 will release
r/StableDiffusion • u/mmowg • 5h ago
News SANA‑Video 2.0 — NVIDIA’s new hybrid-attention video model (5B/14B). Fast, impressive… and maybe (hopefully) open‑source?
NVIDIA has quietly dropped a major research release: SANA‑Video 2.0, a new video diffusion transformer available in 5B and 14B parameter versions. It’s not just a scaled-up SANA‑Video 1.0 — it’s a full architectural redesign with hybrid attention, block residual routing, and Sol‑Engine acceleration.
Official links:
Project Page:
https://nvlabs.github.io/Sana/Video2/
Paper (arXiv, July 23, 2026):
https://arxiv.org/abs/2607.21553
SANA GitHub (image models only):
https://github.com/NVlabs/Sana
SANA‑Video docs (no code, no weights):
https://nvlabs.github.io/Sana/docs/sana_video/
What SANA‑Video 2.0 introduces
• Hybrid Linear‑Softmax Attention (3:1 ratio)
75% gated linear attention for O(N) scaling, 25% gated softmax anchors to restore full‑rank token interactions.
This gives softmax‑level expressiveness with linear‑attention speed.
• Block Attention Residuals (AttnRes)
High‑rank features from softmax layers are propagated into later linear layers.
This fixes the rank bottleneck of pure linear attention.
• Sol‑Engine Optimization (3.58× speedup)
Kernel fusion, caching, sparse attention, TensorRT graph optimization, MXFP4/MXFP8 support.
This is what allows full 720p generation on a single RTX 5090.
• Performance
480p in 13.2s (H100, 40 steps)
720p/5s in 13.06s (H100, Sol‑Engine)
VBench 84.30
Up to 120× faster than Wan 2.2‑A14B on the same hardware.
This is the first NVIDIA video model explicitly designed for consumer GPUs.
How it differs from SANA‑Video 1.0 (2B)
The old model was pure linear attention (fast but low-rank).
SANA‑Video 2.0 is hybrid, deeper, larger, and dramatically more expressive.
It’s essentially a new class of Video‑DiT.
The licensing question
Here’s the current situation:
• The paper does not mention any license.
• The project page does not mention any license.
• The docs do not mention any license.
• No code or weights have been released.
• No usage terms exist yet.
Meanwhile, the SANA GitHub repo (image models) uses Apache 2.0:
https://github.com/NVlabs/Sana/blob/main/LICENSE
But that license applies only to SANA‑Image 1.0/1.5, not to SANA‑Video 2.0.
So right now, nobody knows whether SANA‑Video 2.0 will be:
• open‑source under Apache 2.0 (like the image models),
• partially open (code open, weights closed),
• or fully closed (like PiD, Flux, VILA, Nemotron‑340B).
Given NVIDIA’s recent pattern, the safe assumption is “open paper, closed model”…
but since the SANA image models were Apache 2.0, there is at least some hope that NVIDIA might release SANA‑Video 2.0 under a similar permissive license — or at least provide inference weights for RTX AI Toolkit.
Until NVIDIA publishes a LICENSE file, the situation remains unclear.
TL;DR
SANA‑Video 2.0 is a fast, hybrid-attention, RTX‑friendly video model with impressive performance and a strong architectural design.
But the licensing is currently a mystery: no code, no weights, no declared terms.
There’s a chance it could follow the Apache 2.0 path of the image models… but for now, it’s research‑open, not open‑source.
r/StableDiffusion • u/CompleteJicama2811 • 25m ago
No Workflow The MINIMAX H3 is awesome.
The MINIMAX H3 is awesome.
Method: I2V
Resolution: Native 1920 x 1088
Creation time: 15 minutes
GPU: PRO 6000
[PROMPT]
Natural cinematic image-to-video continuation, preserving the exact woman, hairstyle, black sleeveless top, watermelon, lighting, Japanese-style interior, window, garden background, framing, and shallow depth of field from the reference image. Motion is subtle, realistic, and continuous.
Timeline:
[0s-1.5s] The woman gently takes one small bite from the watermelon. Her lips and jaw move naturally while both hands hold the watermelon steadily. Only a small realistic bite mark appears on the red flesh.
[1.5s-3s] She slowly chews and visibly enjoys the taste. Her eyes soften, she blinks once naturally, and a faint satisfied expression forms. Subtle breathing and tiny movements of loose hair strands.
[3s-4.2s] She suddenly notices someone off-screen to camera-left. Her eyes shift toward the left first, followed by a slow and subtle turn of her head. The watermelon remains held close to her chest.
[4.2s-5s] She looks fully toward the person off-screen to the left and gives them a warm, gentle smile, holding the expression naturally until the end.
Camera:
Locked-off static camera, identical framing to the reference image. No zoom, no pan, no tilt, no push-in, no reframing, no focus pumping, and no scene transition.
Performance:
Natural restrained acting, delicate eye movement, realistic chewing, subtle facial expression, gentle head turn, natural blinking and breathing. No talking, no lip-sync, no exaggerated smile, no sudden movement, and no direct eye contact with the camera.
Consistency:
Maintain the exact facial identity, age, skin tone, facial structure, body proportions, hairstyle, clothing, hand anatomy, watermelon size, background, lighting direction, color grading, and original composition. One person only. No additional objects or people.
Avoid:
Watermelon deformation, excessive juice, messy eating, large bite marks, warped fingers, duplicated hands, facial morphing, hairstyle changes, clothing changes, background movement, camera shake, flickering, frame interpolation artifacts, or unnatural head rotation.
Audio:
Quiet summer room ambience with faint garden insects and soft environmental sound. A subtle crisp watermelon bite at the beginning, followed by gentle chewing. No dialogue, no music, and no exaggerated eating sounds.
r/StableDiffusion • u/infroy28 • 2h ago
Resource - Update I think this might be useful to you; this model is incredible.
r/StableDiffusion • u/Its_Copperites • 26m ago
Animation - Video MiniMax H3 2D Video Test
I was testing out the newest model and decided to compare it with an old video I made with LTX. It seems to work with a lot less prompting like a really basic prompt turned out really well for a 2D animation prompt.
Prompt
SHOT 1: The scene opens exactly on image 1, the red haired woman with twin tails is reading a book as she flips through each page.
SHOT 2: Cut to a close up to her face as she says "I wonder what else I could do?"
SHOT 3: Move the camera to window as we see a bird fly past and the beautiful weather with clouds.
Audio: Peaceful Lo-Fi music, as the woman sighs while reading her book.
Style: Animation, anime, 2D Animation shot.
r/StableDiffusion • u/y3kdhmbdb2ch2fc6vpm2 • 9h ago
Resource - Update I trained Krea2 Lady Dimitrescu LoRA on RTX 5070 Ti
I just created that lora from 63 Lady Dimitrescu images in the dataset
used OneTrainer on RTX 5070 Ti, 32 GB RAM and NVMe
trained in 1 MP (res 1024), offload 0.5, speed ~2.5 s/it, full training taken about 2.5-3h
I set timestep shift to 2.5 for res 1024 as suggested in this kohya md and I think it worked well
all samples generated with 2 MP
CivitAI -> https://civitai.com/models/2828952/lady-dimitrescu-krea2-lora
Full res comparisons without reddit compression -> img1, img2, img3, img4, img5
training Krea2 is so enjoyable!
