日本語

Showcase · Oct 10, 2026

What People Run on Their Own Machines: Local AI Demos on DGX Spark, Compatible Machines and a Windows PC

Illustration of a small unbranded desktop computer on a desk, with floating panels suggesting video frames, sound and conversation

Some people run large AI models on their own hardware instead of through a cloud service, and share what they do with them on X. This article looks at several of those posts, mostly from creators using NVIDIA DGX Spark, DGX Spark-compatible machines and RTX 5090 GPUs, plus one post that summarizes a Microsoft demonstration. For each, we look at what was combined and what came out of it.

A note on scope. DGX Spark and the Windows PCs built on RTX Spark are different products. Most of the examples below come from DGX Spark, compatible machines and an RTX 5090; the last one is a summary of a demonstration on a Windows PC. Whether the same things can be done on RTX Spark machines depends on what each tool supports, and is not covered here. Most numbers and impressions are also the creators' own reports, so they are attributed to them rather than stated as established fact.

Continuous talking video, generated locally

The first example pairs the MiniMax H3 video generation model with Irodori, a text-to-speech tool. The result is a demo in the style of an AI streamer: according to the creator, a character that speaks with a voice is generated continuously and played back without gaps.

The creator writes (in Japanese) that the setup used an RTX 5090 plus DGX Spark-compatible machines, three of them 1. Note that these are compatible machines, not DGX Spark itself. The creator describes a split of work: while the RTX 5090 generates video, the DGX Spark side (the compatible machines) handles the processing before and after, so playback does not stop 2. According to the post, a new video segment is finished about every 6.9 seconds on average 3.

The creator says the segments are now about 7 seconds long, down from 15 seconds when the RTX 5090 worked alone 4. Because the PDMD 2step LoRA had problems with audio, the creator makes the voice with Irodori first 5. For lip sync, the H3 audio VAE converts that voice into an "audio latent" (an internal representation of the sound), and that is passed in as guidance during generation 6. The creator also reports raising the resolution from 320×576 to 384×704 7.

On the tools themselves: according to ComfyUI's documentation, MiniMax H3 has open weights that allow it to be run locally 8. ("Open weights" means the trained model files are published for download.) It is released under the MiniMax H3 Community License Agreement 9, and commercial use of locally generated outputs requires a MiniMax commercial license 10. The code of Irodori-TTS is under the MIT license 11.

The creator is candid about the limits. The comments and replies in the demo were prepared in advance 12. Producing a single reply takes about 14.4 seconds, and in this demo about 16 seconds including the wait for playback, according to the post 13. The creator calls the response delay a remaining issue. The creator also writes that running an actual AI streamer would probably need another PC to run one more model 14, and that the first two greetings were generated in advance and the pixel-art look is added in post-processing 15.

Making video generation lighter and faster

The next post is an announcement from the developers themselves: FastH3, a faster and smaller variant of MiniMax H3, from the FastVideo project at Hao AI Lab.

According to the post, a clip of 5 seconds at 480p with audio takes 15 seconds on a single RTX 5090 with FastH3 V2 16. FastH3 Trim, a smaller version of H3, is said to run in as little as 8 GB of GPU memory 17. The project's README describes Trim as an experimental pruned model that runs in as little as 8 GB of GPU memory 18. The post says every clip shown was generated locally 19.

The README says FastH3 V2 runs on RTX 5090 and other GPUs, DGX Spark and Apple Silicon 20, and says FastH3 runs on DGX Spark through CUDA 13 21. FastVideo itself is under the Apache-2.0 license 22, and it has installation instructions for DGX Spark 23. DGX Spark is an ARM64 machine, and according to those instructions, some components for that platform have to be built from source 24.

Letting one machine think for hours

Next is a single DGX Spark running Qwen3.8-27B, in a demo where it was asked to design a pagoda garden and keep implementing new ideas.

The creator says that thanks to optimized caching and a method called StairCut, processing (prefilling) a context of roughly 150k tokens takes just milliseconds 25. Output speed is given as around 50 to 90 tok/s, depending on the task 26. Tok/s means tokens per second, a measure of how quickly the model writes text. The clip was recorded after 5 hours of continuous running, according to the post 27.

Qwen3.8-27B is an open-weights model under the Apache-2.0 license 28 29. The clip is one example of a local machine left working on a long task.

Turning several machines into a server for every device

The next post is less about a model and more about how to use it.

The creator runs a vLLM server across two DGX Spark units, sharing a single endpoint 30. Tailscale puts all of their machines and devices on the same private network 31. The creator writes that the endpoint speaks the OpenAI chat, OpenAI Responses and Anthropic Messages formats 32, so a chat app, a coding agent or a phone bot pointed at it all talk to the same model 33.

The server also answers to the name "local". Clients can ask for "local" instead of a specific model name, so when the model is swapped on the server, nothing needs to change on the devices 34. At the time of the post it was serving GLM 5.3-Flash 35. The creator also notes that their setup is currently set to handle one request at a time, so when every device sends a request at once, the requests queue up 36.

vLLM is under the Apache-2.0 license 37, and NVIDIA publishes instructions for installing it on DGX Spark 38.

A large model on a Windows PC

The last example is different. It is not a first-hand report from someone who ran the model. It is a third party, Rohan Paul, summarizing a Microsoft demonstration.

According to that summary, Microsoft showed DeepSeek V4 Flash, an open-weights model with 284B parameters, running locally on a Windows PC once quantized to 1.6 bits, which brings it down to about 60 GB of memory 39. Quantizing means storing the model's numbers at lower precision so that it takes less memory. The post adds that the Surface Laptop Ultra, built on the RTX Spark chip, has room for it in its higher-memory configurations 40.

DeepSeek V4 Flash is itself released under the MIT license 41. But the post only says there is room for the model, so it tells us little about what it is like to use.

If you are thinking of trying this yourself

Taken together, the examples point to a few different needs:

  • Splitting a pipeline, such as video and voice generation, across several machines
  • Handing a long task to a machine and leaving it running
  • Running a large model and sharing it with the devices you already own

Windows PCs with RTX Spark include laptops and small desktops. Whether the examples above run on them as they are depends on the support of each tool, so it is worth checking each tool's official documentation before you choose.

Several of the examples above were run on DGX Spark and compatible machines.

Whichever machine you pick, the deciding questions are whether it has enough memory for the model you want to run, and whether the tools you use support it. The speeds and waiting times the creators report differ from setup to setup.

Sources

  1. Post by @yu_ichi_suzukiRTX 5090+DGX Spark互換機3台を使いました。
  2. Post by @yu_ichi_suzukiRTX 5090が動画を生成している間に事前・事後の処理をDGX Sparkが行い、途切れずに再生します。
  3. Post by @yu_ichi_suzuki動画の完成間隔は平均約6.9秒。
  4. Post by @yu_ichi_suzukiRTX 5090単体でやったときは15秒単位でしたが、7秒単位になったことで
  5. Post by @yu_ichi_suzukiPDMD 2step LoRAは音声に問題があるため、先にIrodoriで音声を作り
  6. Post by @yu_ichi_suzuki先にIrodoriで音声を作り、H3の音声VAEで「音声latent」に変換して生成時のガイドとして渡し、口の動きを合わせています。
  7. Post by @yu_ichi_suzuki解像度も320×576から384×704(画素数約1.5倍)に上げられたので
  8. MiniMax H3: ComfyUI MiniMax H3 Video Generation Guide - ComfyUIH3’s open weights let you run the model locally.
  9. MiniMax H3: https://huggingface.co/MiniMaxAI/MiniMax-H3/raw/main/README.mdMiniMax H3 is released under the [MiniMax H3 Community License Agreement](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE).
  10. MiniMax H3: ComfyUI MiniMax H3 Video Generation Guide - ComfyUICommercial use of locally generated outputs requires a MiniMax commercial license
  11. Irodori-TTS: https://raw.githubusercontent.com/Aratako/Irodori-TTS/main/README.md**Code**: [MIT License](LICENSE)
  12. Post by @yu_ichi_suzuki※コメントと返答文は事前に用意したデモです。
  13. Post by @yu_ichi_suzuki1つの返事を作る処理には約14.4秒、今回のデモでは再生待ちを含めて約16秒かかります。
  14. Post by @yu_ichi_suzuki実際にAITuberをやるならもう一つモデルを動かすPCが必要になりそうです。
  15. Post by @yu_ichi_suzuki冒頭2本の挨拶は事前生成、ドット絵化は後処理です。
  16. Post by @haoailabA 5 s 480p clip with audio takes 15 s on one RTX 5090.
  17. Post by @haoailabit runs in as little as 8 GB of GPU memory.
  18. FastH3: https://raw.githubusercontent.com/hao-ai-lab/FastVideo/main/README.mdan experimental pruned model that is 4.2× smaller than base H3 and runs in as little as 8 GB of GPU memory
  19. Post by @haoailabEvery clip here was generated locally.
  20. FastH3: https://raw.githubusercontent.com/hao-ai-lab/FastVideo/main/README.mdFastH3 V2 now runs on a single consumer machine: NVIDIA RTX 5090, RTX 4090 and RTX PRO 6000 GPUs, DGX Spark and Apple Silicon.
  21. FastH3: https://raw.githubusercontent.com/hao-ai-lab/FastVideo/main/README.mdFastH3 now runs locally on Apple Silicon through MLX and on NVIDIA DGX Spark through CUDA 13, including two-Spark inference.
  22. FastVideo: https://raw.githubusercontent.com/hao-ai-lab/FastVideo/main/LICENSEVersion 2.0, January 2004
  23. FastVideo: https://raw.githubusercontent.com/hao-ai-lab/FastVideo/main/docs/getting_started/installation/spark.mdInstructions to install FastVideo on an **NVIDIA DGX Spark**
  24. FastVideo: https://raw.githubusercontent.com/hao-ai-lab/FastVideo/main/docs/getting_started/installation/spark.mdThe Spark is **ARM64 (`aarch64`) with CUDA 13**
  25. Post by @0xBakeerThanks to optimized caching algorithms and the StairCut method, the ~150k context prefill takes just milliseconds
  26. Post by @0xBakeerdecoding speeds hit around 50–90 tok/s depending on the task.
  27. Post by @0xBakeerI recorded after 5 hours of continuous running
  28. Qwen3.8-27B: https://huggingface.co/Qwen/Qwen3.8-27B/raw/main/README.mdlicense: apache-2.0
  29. Qwen3.8-27B: https://huggingface.co/Qwen/Qwen3.8-27B/raw/main/README.mdThis repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format.
  30. Post by @sudoingXthe two sparks run one vLLM server together with one endpoint
  31. Post by @sudoingXinstall tailscale to have all your nodes and machines on the same tailnet
  32. Post by @sudoingXit speaks OpenAI chat, OpenAI Responses and Anthropic Messages
  33. Post by @sudoingXso any chat app, coding agent or phone bot you point at it talks to the same model
  34. Post by @sudoingXthe server also answers to the name "local", so clients can ask for "local" instead of a model name and when i swap the model on the servers nothing on the devices needs a config change
  35. Post by @sudoingXright now it's serving GLM 5.3-Flash
  36. Post by @sudoingXmine is set to one at a time right now, so when every device fires together they queue up
  37. vLLM: https://raw.githubusercontent.com/vllm-project/vllm/main/LICENSEVersion 2.0, January 2004
  38. vLLM: https://raw.githubusercontent.com/NVIDIA/dgx-spark-playbooks/main/nvidia/vllm/README.mdInstall and use vLLM on DGX Spark
  39. Post by @rohanpaul_aia 284B-parameter open-weight model, runs locally on a Windows PC once quantized to 1.6 bits, which brings it down to about 60GB of memory.
  40. Post by @rohanpaul_aihas room for it in its higher-memory configurations.
  41. DeepSeek-V4-Flash: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/raw/main/README.mdThis repository and the model weights are licensed under the [MIT License](LICENSE).

Embedded posts belong to their authors.