Images, video and music, all made with models on home hardware. A few posts like that have been scrolling past on X lately. Here they are next to the tools' official documentation: a video whose images, footage and music each came from a locally run model, a comic whose camera angle you can drag around, and an anime where the camera flies around inside frozen time.
The numbers and impressions in these posts are the creators' own reports, so they are attributed to them rather than stated as fact.
Images, video and music on a home GPU
First, a video from superalesha.
This video was created entirely at home on 1× RTX 3090.
— Alexey Fateev (@superalesha) October 8, 2026
Images: Qwen Image 2.1, FLUX dev
Video: MiniMax H3
Music: YuE2 pic.twitter.com/SPsXt2OGHt
The creator writes that the video was made entirely at home on 1 RTX 3090 1. The credits are spelled out: images from Qwen Image 2.1 and FLUX dev, video from MiniMax H3, and music from YuE2 2. Each stage gets its own model. To us, that combination is what stands out most.
Here is what the official documentation says about each tool.
- MiniMax H3 has open weights, and according to ComfyUI's documentation they let you run the model locally [[F2]]. ComfyUI's page also says it runs locally in ComfyUI [[F4]].
- The Qwen Image 2.1 README says its developers have open-sourced the model [[F32]].
- FLUX.1 [dev] has open weights [[F62]]. Its license is the FLUX.1 [dev] Non-Commercial License, which allows use for non-commercial purposes only [[F61]] [[F63]].
- YuE2 has open weights [[F69]], plus ComfyUI nodes and an official workflow [[F73]]. The model weights are under CC BY-NC 4.0 with additional creator permission [[F68]]. Individuals, creators and musicians can monetize outputs for free, while companies are asked to contact the authors about a commercial license for the weights [[F74]].
If you plan to publish or monetize what you make, the license differs from model to model, so read each one's official page first. For example, commercial use of locally generated MiniMax H3 outputs requires a MiniMax commercial license 3.
A comic you can drag around
Next is a comic from monoradio. In a reply, the creator writes (in Japanese, our translation) "you can actually try it here" 4, linking a page on claude.ai whose title translates as "Lighthouse in the Desert." The page is captioned, in Japanese, with words that translate as "drag a panel to spin it" 5.
MiniMax-H3 + ORB360 LoRAで画角を自由に動かせる漫画を作ってみた。
— monoradio (@monoradio102) October 11, 2026
キャラを参照画像で渡せば、顔も固定できる。 pic.twitter.com/FHfYf8TrZA
On how the comic was made, the creator writes (in Japanese) that it was made with MiniMax-H3 and the ORB360 LoRA, and that the camera angle can be moved freely 6. Passing the character in as a reference image can also keep the face fixed, according to the post 7. In a reply, the creator describes the workflow (in Japanese) as storyboard sketches, then image generation, then animation 8.
Now the LoRA's own documentation. According to its README, you load it with the standard Load LoRA node on a MiniMax-H3 Ref2VA 9. It is released under the MiniMax H3 Community License Agreement 10. The idea of moving the camera angle "freely" (our translation) is the creator's wording, not a specification written on the LoRA's page.
What makes this example fun is how close it sits to an illustrator's normal work: make images from the storyboard, then set them in motion. If your starting point is a hand-drawn character sheet or storyboard, a pen tablet makes that flow easier.
The camera flies through frozen time
Last is an anime from Sawa (AI_swwww). The creator writes (in Japanese) that a prompt can produce a scene where the camera flies around inside frozen time, in a clip of 15 seconds 11.
AIアニメ、まだ静止画をちょっと動かすだけで終わってませんか。
— さわ|AI画像生成で月3桁継続中 (@AI_swwww) October 6, 2026
いまは時間停止の中をカメラが飛び回る演出まで、プロンプトで作れます。
しかも1本15秒。
やり方とプロンプトはリプ欄から👇 pic.twitter.com/3nGiJ7gEPm
The method is in the creator's reply.
【やり方】
— さわ|AI画像生成で月3桁継続中 (@AI_swwww) October 6, 2026
やることはこの4つです。
① 設定画:GPT Image 2.5でオリキャラの設定画を1枚(正面・斜め・横・後ろ、表情3つ、小物)
② 動画:Seedance 2.5(Higgsfield)の参照モードに設定画を渡す。16:9・15秒
③ プロンプト:「走る→転ぶ→時間停止→カメラが飛ぶ→時間が戻る」を秒単位で書く…
The steps are as follows (translated from Japanese).
- Character sheet: first make a character sheet for an original character with GPT Image 2.5 (front, three-quarter, side and back views, 3 expressions, and props) [[C66]]
- Video: give the sheet to the reference mode of Seedance 2.5 (via Higgsfield), at 16:9 and 15 seconds [[C67]]
- Prompt: write the sequence "run, fall, time stops, camera flies, time resumes" second by second [[C68]]
- Tip: while time is stopped, state that nothing except the camera moves, and freeze blinking, hair and clothing sway as well [[C69]]
The creator names services such as Seedance 2.5 via Higgsfield; the post does not describe a locally run setup 12.
What follows is the writer's opinion. Pinning down the look with a character sheet and writing out the flow of time second by second are ideas that could carry over to locally run video generation too. They also seem to pair well with passing a reference image to keep a character's face consistent, as in the MiniMax H3 example above.
If you want to do it at your own desk
The examples above mix a setup where the creator names a home GPU, one made through online services, and one whose hardware is not described. If you are thinking about doing it all at your own desk, here is what NVIDIA says about RTX Spark.
- According to NVIDIA's product page, you can generate images and video right on your device, as often as you like, with no subscription and no per-prompt fees [[F101]].
- On the same page NVIDIA says running locally gives you instant iteration, no API bills, and data that stays on your machine [[F102]].
- According to NVIDIA's blog, RTX Spark runs the full NVIDIA CUDA platform [[F104]]. For creators, it lists 5th-generation Tensor Cores with NVFP4 support, hardware-accelerated AV1, and 4:2:2 video encode and decode [[F105]].
Some of this is still planned. NVIDIA says RTX Video Frame Generation will be coming to ComfyUI, that Adobe is rearchitecting Photoshop and Premiere, and that these updates arrive this fall with RTX Spark 13. For Adobe Premiere, NVIDIA says it is working with Adobe to redesign it for RTX Spark 14. A comment from ComfyUI co-founder Yannik Marek, published on NVIDIA's blog, is also written in the future tense: ComfyUI users will be able to run very complex multimodal workflows and generate ultra-high-resolution images and video on portable devices at unprecedented speed 15.
None of the examples above is described as made on an RTX Spark machine. If you want to use MiniMax H3 or YuE2 on one, check each tool's official documentation for supported environments first.
Here are some RTX Spark laptops.
Once you start making video, you quickly need somewhere to keep the clips. If you move between images, video and music as these creators do, keeping a separate drive for exports helps.
Read next
To go deeper on the music side, see the YuE2 local music examples.
For a different direction on the same make-it-yourself theme, see the article on letting an AI agent handle Blender.
