Some people run large AI models on their own hardware instead of through a cloud service, and share what they do with them on X. This article looks at several of those posts, mostly from creators using NVIDIA DGX Spark, DGX Spark-compatible machines and RTX 5090 GPUs, plus one post that summarizes a Microsoft demonstration. For each, we look at what was combined and what came out of it.
A note on scope. DGX Spark and the Windows PCs built on RTX Spark are different products. Most of the examples below come from DGX Spark, compatible machines and an RTX 5090; the last one is a summary of a demonstration on a Windows PC. Whether the same things can be done on RTX Spark machines depends on what each tool supports, and is not covered here. Most numbers and impressions are also the creators' own reports, so they are attributed to them rather than stated as established fact.
Continuous talking video, generated locally
The first example pairs the MiniMax H3 video generation model with Irodori, a text-to-speech tool. The result is a demo in the style of an AI streamer: according to the creator, a character that speaks with a voice is generated continuously and played back without gaps.
ローカルMiniMax H3で、声つきの約7秒動画をリアルタイムに連続生成してみました。RTX 5090+DGX Spark互換機3台を使いました。
— 鈴木憂一 | Highdrama (@yu_ichi_suzuki) October 5, 2026
動画の完成間隔は平均約6.9秒。RTX 5090が動画を生成している間に事前・事後の処理をDGX Sparkが行い、途切れずに再生します。
RTX… pic.twitter.com/ogjmQUKHJn
The creator writes (in Japanese) that the setup used an RTX 5090 plus DGX Spark-compatible machines, three of them 1. Note that these are compatible machines, not DGX Spark itself. The creator describes a split of work: while the RTX 5090 generates video, the DGX Spark side (the compatible machines) handles the processing before and after, so playback does not stop 2. According to the post, a new video segment is finished about every 6.9 seconds on average 3.
The creator says the segments are now about 7 seconds long, down from 15 seconds when the RTX 5090 worked alone 4. Because the PDMD 2step LoRA had problems with audio, the creator makes the voice with Irodori first 5. For lip sync, the H3 audio VAE converts that voice into an "audio latent" (an internal representation of the sound), and that is passed in as guidance during generation 6. The creator also reports raising the resolution from 320×576 to 384×704 7.
On the tools themselves: according to ComfyUI's documentation, MiniMax H3 has open weights that allow it to be run locally 8. ("Open weights" means the trained model files are published for download.) It is released under the MiniMax H3 Community License Agreement 9, and commercial use of locally generated outputs requires a MiniMax commercial license 10. The code of Irodori-TTS is under the MIT license 11.
The creator is candid about the limits. The comments and replies in the demo were prepared in advance 12. Producing a single reply takes about 14.4 seconds, and in this demo about 16 seconds including the wait for playback, according to the post 13. The creator calls the response delay a remaining issue. The creator also writes that running an actual AI streamer would probably need another PC to run one more model 14, and that the first two greetings were generated in advance and the pixel-art look is added in post-processing 15.
Making video generation lighter and faster
The next post is an announcement from the developers themselves: FastH3, a faster and smaller variant of MiniMax H3, from the FastVideo project at Hao AI Lab.
1/8) FastVideo FastH3 now runs on a single consumer machine: @NVIDIA RTX GPUs 💚, DGX Spark and @Apple Silicon 🍎
— Hao AI Lab (@haoailab) October 6, 2026
FastH3 V2 beats base @MiniMax_AI H3 in 8 steps. A 5 s 480p clip with audio takes 15 s on one RTX 5090.
We also release FastH3 Trim: 4.2× smaller than base H3, and… pic.twitter.com/Ma9NZsfr3t
According to the post, a clip of 5 seconds at 480p with audio takes 15 seconds on a single RTX 5090 with FastH3 V2 16. FastH3 Trim, a smaller version of H3, is said to run in as little as 8 GB of GPU memory 17. The project's README describes Trim as an experimental pruned model that runs in as little as 8 GB of GPU memory 18. The post says every clip shown was generated locally 19.
The README says FastH3 V2 runs on RTX 5090 and other GPUs, DGX Spark and Apple Silicon 20, and says FastH3 runs on DGX Spark through CUDA 13 21. FastVideo itself is under the Apache-2.0 license 22, and it has installation instructions for DGX Spark 23. DGX Spark is an ARM64 machine, and according to those instructions, some components for that platform have to be built from source 24.
Letting one machine think for hours
Next is a single DGX Spark running Qwen3.8-27B, in a demo where it was asked to design a pagoda garden and keep implementing new ideas.
Running Qwen3.8-27B (the dense model, not Qwen Flash) on a single DGX Spark.
— 0xBakeer (@0xBakeer) September 30, 2026
Thanks to optimized caching algorithms and the StairCut method, the ~150k context prefill takes just milliseconds, and decoding speeds hit around 50–90 tok/s depending on the task.
In this clip, I… pic.twitter.com/4Mhn9578CH
The creator says that thanks to optimized caching and a method called StairCut, processing (prefilling) a context of roughly 150k tokens takes just milliseconds 25. Output speed is given as around 50 to 90 tok/s, depending on the task 26. Tok/s means tokens per second, a measure of how quickly the model writes text. The clip was recorded after 5 hours of continuous running, according to the post 27.
Qwen3.8-27B is an open-weights model under the Apache-2.0 license 28 29. The clip is one example of a local machine left working on a long task.
Turning several machines into a server for every device
The next post is less about a model and more about how to use it.
single most useful thing for my local ai setup that made life easy is my 2x DGX Spark becoming one serving box for every device i own
— Sudo su (@sudoingX) September 30, 2026
here is how i do it:
> 1. the two sparks run one vLLM server together with one endpoint, i use dgx sparks, you can use any nodes that can run an… https://t.co/bqNaCMQ2Ps pic.twitter.com/lC4m9AE0Wv
The creator runs a vLLM server across two DGX Spark units, sharing a single endpoint 30. Tailscale puts all of their machines and devices on the same private network 31. The creator writes that the endpoint speaks the OpenAI chat, OpenAI Responses and Anthropic Messages formats 32, so a chat app, a coding agent or a phone bot pointed at it all talk to the same model 33.
The server also answers to the name "local". Clients can ask for "local" instead of a specific model name, so when the model is swapped on the server, nothing needs to change on the devices 34. At the time of the post it was serving GLM 5.3-Flash 35. The creator also notes that their setup is currently set to handle one request at a time, so when every device sends a request at once, the requests queue up 36.
vLLM is under the Apache-2.0 license 37, and NVIDIA publishes instructions for installing it on DGX Spark 38.
A large model on a Windows PC
The last example is different. It is not a first-hand report from someone who ran the model. It is a third party, Rohan Paul, summarizing a Microsoft demonstration.
Microsoft today showcased DeepSeek V4 Flash, a 284B-parameter open-weight model, runs locally on a Windows PC once quantized to 1.6 bits, which brings it down to about 60GB of memory.
— Rohan Paul (@rohanpaul_ai) October 7, 2026
Their new Surface Laptop Ultra, built on Nvidia's RTX Spark chip with up to 128GB of unified… pic.twitter.com/SGhAz9ehaT
According to that summary, Microsoft showed DeepSeek V4 Flash, an open-weights model with 284B parameters, running locally on a Windows PC once quantized to 1.6 bits, which brings it down to about 60 GB of memory 39. Quantizing means storing the model's numbers at lower precision so that it takes less memory. The post adds that the Surface Laptop Ultra, built on the RTX Spark chip, has room for it in its higher-memory configurations 40.
DeepSeek V4 Flash is itself released under the MIT license 41. But the post only says there is room for the model, so it tells us little about what it is like to use.
If you are thinking of trying this yourself
Taken together, the examples point to a few different needs:
- Splitting a pipeline, such as video and voice generation, across several machines
- Handing a long task to a machine and leaving it running
- Running a large model and sharing it with the devices you already own
Windows PCs with RTX Spark include laptops and small desktops. Whether the examples above run on them as they are depends on the support of each tool, so it is worth checking each tool's official documentation before you choose.
Several of the examples above were run on DGX Spark and compatible machines.
Whichever machine you pick, the deciding questions are whether it has enough memory for the model you want to run, and whether the tools you use support it. The speeds and waiting times the creators report differ from setup to setup.
