Hybrid Automatic Video Colorizer (HAVC) server that exposes a GPU-accelerated colorization pipeline for black-and-white images and video frames based on Diffusion Transformer (DiT) models. 4 backends, one API : pick the one that fits your hardware:
- nunchaku-qwen: SVDQuant FP4/INT4 transformer via Nunchaku : 4 sec/frameΒΉ, requires RTX 30/40/50 (16GB+ VRAM , 64GB RAM) & CUDA 13.0
- gguf-qwen: ComfyUI-native GGUF pipeline (Q3_K_S, Q4_K_S, Q5_K_M, Q6_K, Q8_0) : 12 sec/frameΒ², runs on RTX 30/40/50 (12GB+ VRAM, 32GB+ RAM), zero ComfyUI GUI dependency
- longcat-gguf: LongCat-Image-Edit-Turbo GGUF pipeline (Q3_K_MβQ8_0) : ~12 sec/frameΒ², runs on RTX 30/40/50 (12GB+ VRAM, 32GB+ RAM), better image quality than gguf-qwen, zero ComfyUI GUI dependency
- qwen21-viggle: Qwen-Image-2.1 (native ComfyUI int8 ConvRot UNet) + Viggle-Turbo LoRA : ~6 sec/frameΒΉ (
steps=2) or ~8-11 sec/frame (steps=6, single-image, see What's New), runs on RTX 30/40/50 (14GB+ VRAM, 32GB+ RAM), optionalenhance_prompt(Qwen3-VL image-aware prompt rewriting), zero ComfyUI GUI dependency
ΒΉ Measured with Fast Pipeline (paired inference, two frames per forward pass) at the backend's fastest recommended step count. Β²
gguf-qwen/longcat-ggufdon't support paired inference (fall back to per-image processing, see What's New) β their figure is a genuine single-image time, not directly comparable to the Fast Pipeline figures above.
Recommended: nunchaku-qwen and qwen21-viggle are both recommended for production use, nunchaku-qwen at the fastest usable step count (
steps=2) has an inference speed of about 4 sec/frame using Fast Pipeline, qwen21-viggle at fastest usable step count (steps=2) has an inference speed of about 8 sec/frame; using Fast Pipeline the speed improves to about 6 sec/frame (not 4 β the pair-mode working resolution was deliberately raised for this backend to avoid a color artifact, see theβ οΈ Fast Pipelinenote below).qwen21-viggleneeds meaningfully less hardware (14GB+ VRAM / 32GB+ RAM vs. 16GB+ VRAM / 64GB+ RAM).longcat-ggufremain the choice for VRAM-constrained setups where neither of the above fits, at a real speed cost (see theβ οΈ Experimentalnote under GGUF below).
β οΈ Fast Pipeline +qwen21-viggle: paired inference uses a higher working resolution for this backend specifically (1280vs.1024for single images and for the other backends, since What's New 2026-09-27) β this fixes a color artifact previously seen on fine detail near the merge boundary (e.g. a hand rendered in tones close to the surrounding foliage), at the cost of some speed (~6 sec/frame instead of 4). A milder residual effect can still appear on secondary, color-ambiguous details (e.g. a flower's petals taking a noticeably different but still plausible hue between runs) β not a defect on the same order as the original artifact, more of the same color-hedging behavior described elsewhere in this README. The Viggle-Turbo LoRA is still an experimental release, and this residual effect may be a limitation of the LoRA itself. If maximum consistency matters more than speed, disable Fast Pipeline forqwen21-viggle(~8 sec/frame, no longer speed-competitive withnunchaku-qwen) or spot-check the output before a long batch run.Color stability vs. variety: based on real-world use across thousands of frames,
nunchaku-qwentends to show more color variability between similar frames β can look more vivid, but with weaker frame-to-frame consistency β whileqwen21-viggleis more conservative in its color choices and more stable, likely a consequence of the Viggle-Turbo LoRA's aggressive step-distillation, which tends to narrow the range of plausible outputs toward "safe" choices. For video work, where flickering color between consecutive frames is a visible defect, this makesqwen21-viggle's conservatism a practical advantage rather than just a stylistic difference β worth factoring in alongside the speed/hardware trade-offs above. The same pattern shows up specifically in Fast Pipeline (paired inference): when both frames share an object,qwen21-viggleconsistently colors it the same way in both halves, whilenunchaku-qwenis less reliable at this β the exact cause (the LoRA itself vs. something more general about the two pipelines) is not established.
If you already have the
.venvwith CUDA 13.0 and just need to update the project to the latest version, follow these steps:
Shortcut:
quick_update.cmdautomates all of the steps below (including thecomfy-kitchen/comfy-aimdopin) β activate the.venvfirst, then double-click it or run it from a terminal. The manual steps are documented here for reference and for non-Windows setups.
# 1) Pull the latest code
git pull
# 2) Activate the virtual environment
.venv\Scripts\activate
# 3) Install / update the GUI dependencies (if new packages were added)
pip install -r GUI\requirements.txt
# 4) Update vscmnet2 (if a newer wheel is available in packages/)
pip install packages\vscmnet2-1.1.0-py3-none-any.whl
# 5) Re-apply the Nunchaku patch
python patch_nunchaku.py
# 6) Required for qwen21-viggle (see What's New, 2026-09-26): pin
# comfy-kitchen and comfy-aimdo to the tested versions β install.cmd
# only sets these for a fresh install, an existing .venv needs this
# explicitly
pip install comfy-kitchen==0.2.35
pip install comfy-aimdo==0.5.5
# 7) Verify everything is up-to-date
pip show torch # Expected: 2.10.0+cu130
pip show nunchaku # Expected: 1.2.1+cu13.0torch2.10
pip show comfy-kitchen # Expected: 0.2.35
pip show comfy-aimdo # Expected: 0.5.5Note: steps 4β5 are only needed if
packages/orpatch_nunchaku.pyhave changed. Step 6 is only needed to useqwen21-viggleβ the other three backends work with the oldercomfy-kitchen/comfy-aimdoversions. Checkgit log --oneline -5to see what was updated.
If you already created the
.venvwith a previous version (CUDA 12.8, PyTorch 2.9.1, Nunchaku cu12.8torch2.9), upgrade to get these benefits:
| Improvement | Before (12.8) | After (13.0) |
|---|---|---|
| CUDA allocator | native (slower reallocation) |
cudaMallocAsync (async, ~10 % faster memory ops) |
| comfy-kitchen CUDA | disabled: True (fallback to eager) |
disabled: False (native dequantization kernels) |
| Warning | You need pytorch with cu130 or higher |
gone (build matches Nunchaku) |
Upgrade steps:
# 1) Deactivate and reactivate the venv to ensure a clean shell
deactivate
.venv\Scripts\activate
# 2) Upgrade PyTorch to 2.10 + CUDA 13.0
pip install torch==2.10.0+cu130 torchvision==0.25.0+cu130 torchaudio==2.10.0+cu130 --index-url https://download.pytorch.org/whl/cu130 --force-reinstall
# 3) Upgrade Nunchaku (CUDA 13.0 + PyTorch 2.10 build)
pip install https://github.com/nunchaku-ai/nunchaku/releases/download/v1.2.1/nunchaku-1.2.1+cu13.0torch2.10-cp312-cp312-win_amd64.whl --force-reinstall
# 4) Re-pin PyTorch (Nunchaku may have upgraded it to 2.12)
pip install torch==2.10.0+cu130 torchvision==0.25.0+cu130 torchaudio==2.10.0+cu130 --index-url https://download.pytorch.org/whl/cu130 --force-reinstall
# 5) Re-apply the Nunchaku patch
python patch_nunchaku.py
# 6) Verify
pip show torch # Expected: 2.10.0+cu130
pip show nunchaku # Expected: 1.2.1+cu13.0torch2.10Run Server (Tab 2, Colorization) now starts and stops the RPC server
itself instead of just launching a .cmd file in its own terminal window:
the server runs as a hidden child process controlled by the GUI, and its
full output streams live into a new Server Log tab β right next to the
existing log, now labeled App Log β in the Dashboard. The button
becomes Stop Server while it's running, and a status line next to it
tracks the sequence: Starting server on ... β running on ... (once the
server actually reports it's listening) β stopped.
Once the server reports it's ready, the GUI connects automatically β no need to also click Connect on Tab 2. Closing the GUI (or clicking Stop Server) always shuts the process down cleanly; the RPC connection indicator resets to Disconnected at the same time, since the server it was talking to is gone.
A new External console checkbox next to the button restores the previous behavior exactly (a separate visible console window, started and left running independently of the GUI) for anyone who prefers it or needs to keep an eye on the raw console.
A Local DiT Server frame on the Dashboard mirrors the Run Server button and its status text, so the server can be started/stopped without switching to Tab 2.
START PIPELINE also uses this: if the 3. Colorize Frames (AI) task is enabled and the client isn't connected yet, the GUI starts the DiT server automatically (same as clicking Run Server) and holds the pipeline start until it reports it's actually online, instead of just failing with "not connected". This only applies when Run Server is GUI-managed (External console unchecked) β with an external console there's no way for the GUI to know when that separate process is ready, so the previous behavior (an error asking to connect manually) still applies there.
A new optional Dashboard task, 2. Select Reference Frames, has been added
between Extract and Colorize β every task after it, is renumbered
(Colorize/Encode/Merge become tasks 3/4/5). It deduplicates the reference
frames extracted in Step 1 by semantic similarity (DINOv3-based, via
vscmnet2.vs_select_reference_frames()), reducing redundant near-identical
candidates before they reach colorization β useful for long or slow-changing
scenes where scene-change detection alone still produces many visually
similar frames.
The task renames ref_tht10/ (produced by Extraction) to ref_tht10_temp/,
then writes the deduplicated representative frames back to a freshly created
ref_tht10/ β the same folder Colorize already reads from, so no other step
changes. ref_tht10_temp/ is kept as a full backup of every extracted
candidate unless Move Files is checked. If ref_tht10/ is missing/empty,
or ref_tht10_temp/ already exists from an interrupted previous run, the
task stops the entire pipeline with an error rather than guessing or
overwriting anything.
New Selection Settings frame in the GUI's Extraction tab exposes
similarity_threshold, select_window, and the Dry Run/Debug HTML/
Move Files options. If Debug HTML is checked, in the output folder is written the file cluster_debug.html. This files contains all the reference clusters as shown in the image below
for example in the Cluster 2, the reference frame #000145 was selected to represent all the references included in the Cluster 2. If the parameter similarity threshold is set above 0.95 will be selected smaller clusters, vice-versa if the threshold is set below 0.95 the similarity clusters will be bigger (will be available less reference frame to colorize).
This deduplication of keyframes will improve color consistency and accelerate the coloring process, as fewer images will need to be colored.
See GUI README: Tab 1 for the full workflow and recovery steps if a run is interrupted.
Existing
gui_cmnet2_settings.jsonfiles are migrated automatically on next load β no manual action needed.
The Fix Colors tab (Tab 5) now exposes a Backbone combo (dinov3 /
dinov2), passed as the backbone parameter of vscmnet2.pil_cmnet2_colorize()
β the same choice already available in Encode/Merge (Tab 3) and Fix
Video (Tab 6), now consistent across all three tabs that drive CMNET2.
Previously the tab always used the vscmnet2 default (dinov3) with no way
to select the legacy DINOv2 backbone. Applies in both single-image and batch
mode. Persisted in gui_cmnet2_settings.json as fixc_backbone.
2026-09-29 β qwen21-viggle: GGUF+mmproj CLIP as new default, clip_mmproj generalized, known limitation documented
New default CLIP for qwen21-viggle: Qwen3-VL-8B-Instruct-UD
(GGUF+mmproj, unsloth/Qwen3-VL-8B-Instruct-GGUF), loaded through a
vendored ComfyUI-GGUF-Reboot custom node (the standard ComfyUI-GGUF
does not support merging a separate mmproj file for the qwen3vl
architecture β only qwen2vl, used by gguf-qwen/longcat-gguf).
Replaces the .safetensors CLIP options evaluated (int8_convrot,
fp8_scaled, w4a8) as the default: same disk footprint as the lightest
of those (w4a8, ~5.9GB) but without a chromatic-drift issue found on
subjects with a strong color convention (w4a8 occasionally converged on
the wrong hue where the other options didn't). The int8_convrot/w4a8
files remain valid alternatives β see Pipeline Configuration.
clip_mmproj/clip_mmproj_hf_name (config fields) generalized from
qwen21-viggle to gguf-qwen/longcat-gguf too, replacing the old
mmproj_gguf key (which was never actually read by any code β dead
documentation only). This also closed a real gap: longcat-gguf never had
any mechanism to auto-download its own mmproj file β it only worked
because the file was already present from gguf-qwen sharing the same
folder. A from-scratch longcat-gguf-only installation would have loaded
its CLIP without a working vision tower.
Known limitation, extensively investigated: on some frames, a human
body part near a visually similar background (e.g. a hand close to
foliage) can be rendered with the wrong color (blended into the
background) instead of a natural skin tone β confirmed across the entire
Qwen-Image-2.1/Viggle-Turbo family, including Viggle's own official demo
app and longcat-gguf, and independent of which CLIP quantization/variant
is used (int8, fp8, w4a8, and several GGUF text-encoder builds were
tested). This appears to be a genuine limitation of the underlying models
for this kind of ambiguous content, not a bug in this integration. If a
frame is affected, nunchaku-qwen/gguf-qwen are unaffected by the same
issue and can be used as a fallback.
Paired inference (Fast Pipeline) for qwen21-viggle now uses a working
resolution of 1280 instead of the usual 1024 (single-image and every
other backend are unaffected). This fixes a color artifact found on frames
with fine detail near the merge boundary β a hand, held up close to the
camera, could be rendered in tones nearly indistinguishable from the
background foliage instead of a natural skin tone. The cause was the
reduced working resolution from paired inference combined with
Viggle-Turbo's own limits on fine detail; raising it to 1280 resolves the
artifact in every case tested, at a real but modest speed cost (~6
sec/frame instead of 4 β see the recommendation note near the top of this
README). 1536 was
tested too and fixes the same artifact slightly more completely, at
roughly double the extra cost; 1280 was chosen as the better trade-off
after validation on thousands of real frames.
Updated to vscmnet2 1.1.0. The proximity-weighted memory matching feature added in 1.0.9 (see below) is no longer installation-wide only: vs_cmnet2 now accepts enable_proximity_bias/proximity_bias_alpha directly, taking precedence over vsslib/models.json when passed explicitly for a single call.
clip = vs_cmnet2(
clip,
clip_ref=ref_clip,
method=0,
enable_proximity_bias=True,
proximity_bias_alpha=0.5,
)vs_cmnet2_recolor/vs_cmnet2dit are unchanged β they still only pick up the installation-wide default from vsslib/models.json. Exposed in the GUI's Encode/Merge tab only (the tab backed by vs_cmnet2), in a new CMNET2 Backbone frame grouping Backbone, Proximity Bias and Alpha together: the latter two are automatically disabled when Backbone = dinov2 (DINOv3-only feature) and re-enabled on switching back to dinov3. Unchecked always forces enable_proximity_bias=False for that run (an explicit override, not "leave it to models.json"). Not added to Fix Video (backed by vs_cmnet2_recolor, which doesn't accept these parameters).
Existing installations, action needed: the shipped DINOv3 checkpoint was renamed
DINOv3FeatureV6_LocalAtten_p372402.pthβDINOv3FeatureV6_LocalAtten_p374099.pth(defaultproximity_bias_alphaalso changed 0.7 β 0.5). Re-download the checkpoint under the new name β see DINOv3 backbone weights. If the old file is left in place,vscmnet2fails fast at init with a clear error listing the files actually present in the weights directory.
A fourth model backend has been added: qwen21-viggle (Qwen-Image-2.1 native ComfyUI weights + Viggle-Turbo LoRA). Uses native int8 ConvRot quantized weights (not GGUF) β GGUF quantization was evaluated and works, but is ~2Γ slower for this model, so it was not adopted. Runs on 14GB+ VRAM GPUs, ~8-11 sec/frame (the time scales little with the number of steps β a fixed text-encoding/VAE-decode cost dominates over sampling).
To use this backend on an existing installation:
git pullthen runquick_update.cmdβ this pulls in the requiredcomfy-kitchen==0.2.35/comfy-aimdo==0.5.5versions (see Quick Update) alongside the rest of the qwen21-viggle code.
Four model files are required (auto-downloaded on first run):
| File | Size | Source |
|---|---|---|
unet/qwen_image_2.1_int8_convrot.safetensors |
~6.8 GB | Comfy-Org/Qwen-Image-2.1 |
clip/qwen3vl_8b_int8_convrot.safetensors |
~8.7 GB | Comfy-Org/Qwen-Image-2.1 |
vae/qwen_image_2.1_vae_bf16.safetensors |
~0.7 GB | Comfy-Org/Qwen-Image-2.1 |
loras/Qwen-Image-2.1-viggle-turbo-v0.2.1-6step-lora-r128.safetensors |
~0.7 GB | Viggle/Qwen-Image-2.1-viggle-turbo |
Launch via run_server_qwen21.cmd (no arguments needed β a single config
is available). Config: config/qwen21_viggle.json.
Steps: the Viggle-Turbo LoRA is distilled for 6-step inference (its
native step count). Four precomputed sigma schedules are available β
2, 4, 6 (native), 8 β selected via the usual steps parameter; any
other value falls back to the 6-step schedule with a warning in the logs.
2/4/8 are experimental (the LoRA author only documents 5-7 step schedules).
enhance_prompt (optional, default False, all colorization RPC
methods): rewrites the prompt using Qwen3-VL as an image-aware "observer"
before colorizing β useful when the plain prompt leaves color ambiguous on
recognizable subjects (e.g. a costume with a well-known color) and the
model resolves the ambiguity inconsistently. Adds ~15-20s per frame (a
second Qwen3-VL generation pass). A direct, explicitly anti-hedging prompt
often achieves the same result without the extra cost β see the
suggested prompt below before reaching for
enhance_prompt by default.
Prerequisite: 12 GB+ VRAM and 32GB+ RAM. Uses
comfy_bridge's native ComfyUI runtime (no external ComfyUI checkout needed) β the same asgguf-qwen/longcat-gguf, extended with native Qwen-Image-2.1/Qwen3-VL support.
The GUI Tab 2 (Colorization) supports this backend: selecting
qwen21-viggle from Model Name auto-disables the (unused) Precision
combo and reads model paths from config/qwen21_viggle.json. Run
Server manages the equivalent of run_server_qwen21.cmd directly (see
What's New, 2026-09-30) β or launches that same .cmd file
in its own console when the External console checkbox is ticked.
An Enhance Prompt checkbox is available in Tab 2 and Tab 4 (Fix Image).
Added x264 as a third software encoder choice in the GUI's Encode/Merge tab, alongside the existing x265 (software, 10-bit) and Nvenc (GPU hardware, H.265). x264 is an 8-bit H.264 CPU encoder β useful when H.265 decoding/compatibility is a constraint. The x264.exe binary is located next to x265.exe, same convention already used for NVEncC64.exe β no extra path field needed in the GUI. Now included in the Release 1.0.0 tools.zip alongside x265.exe/mkvmerge.exe. See GUI README: install external tools.
Updated to vscmnet2 1.0.9, which add proximity-weighted memory matching. By default, permanent-memory candidates are ranked purely by content similarity, with no notion of when in the video a reference frame was captured relative to the frame being colorized β with a wide max_memory_frames window holding several visually similar but differently-colored references, this can wash the result toward gray. Unlike backbone, this is not exposed as a filter parameter on vs_cmnet2: it is configured once for the whole installation via the enable_proximity_bias/proximity_bias_alpha keys in vsslib/models.json (see above). Off by default. To permanently enable it (useful for permanent memory window size > 50) it is necessary to set enable_proximity_bias=true in the configuration file stored in: vsslib/models.json as shown in the example below:
{
"cmnet2": {
"dinov3": {
"checkpoint": "DINOv3FeatureV6_LocalAtten_p374099.pth",
"weights_dir": "dinov3-vitb16",
"enable_proximity_bias": true,
"proximity_bias_alpha": 0.5
},
"dinov2": {
"checkpoint": "DINOv2FeatureV6_LocalAtten_s2_154000.pth"
}
}
}Updated to vscmnet2 1.0.8, which switches CMNET2 to a DINOv3 ViT-B/16 key-encoder backbone by default (previously DINOv2 ViT-S/14), improving colorization quality. The legacy DINOv2 backbone remains available via a backbone parameter.
A new Backbone combo (dinov3 / dinov2) has been added to the GUI in both tabs that drive CMNET2 through VapourSynth:
- Encode/Merge (Tab 3) β
GUI/scripts/encode_cmnet2.vpy - Fix Video (Tab 6) β
GUI/scripts/encode_cmnet2_recolor.vpy
Both scripts now pass the selected backbone to vs_cmnet2() / vs_cmnet2_recolor() via a Backbone VapourSynth argument, alongside the existing RenderSpeed and MemoryFrames parameters.
Prerequisite: the DINOv3 backbone requires new weight files β see GUI README: DINOv3 backbone weights for download links and install steps. The
dinov2option remains available for installations that only have the legacy DINOv2 weights.
A third model backend has been added: longcat-gguf (LongCat-Image-Edit-Turbo GGUF). Uses quantized UNet (Q4_K_M) and CLIP GGUF files β runs on 12 GB VRAM GPUs. It achieves excellent colorization quality (~12 s/frame via the RPC server) β richer colors, more natural skin tones, and better detail preservation than the gguf-qwen model β and is accessible from all existing tabs and RPC endpoints.
Three new model files are required (auto-downloaded on first run):
| File | Size | Source |
|---|---|---|
unet/LongCat-Image-Edit-Turbo-Q4_K_M.gguf |
~5.4 GB | vantagewithai/LongCat-Image-Edit-Turbo-GGUF |
clip/Qwen2.5-VL-7B-Instruct-Q4_K_M.gguf |
~4.6 GB | unsloth/Qwen2.5-VL-7B-Instruct-GGUF |
vae/lct_vae.safetensors |
~160 MB | meituan-longcat/LongCat-Image-Edit-Turbo |
Launch via run_server_longcat.cmd (Q4_K_M) or start_server.cmd longcat (Q4), longcat-q3, longcat-q5, longcat-q6, longcat-q8. Pre-made configs for all quantizations are in the config/ folder.
Prerequisite: 12 GB+ VRAM and 32GB+ RAM. Uses GGUF quantized models. The pipeline uses the comfy_bridge runtime (no external ComfyUI checkout needed). Custom nodes
CFGNorm,FluxKontextMultiReferenceLatentMethod, andTextEncodeQwenImageEditPlusare included incomfy_bridge/comfy_extras/.
The GUI Tab 2 (Colorization) now includes a Run Server button that manages
the server for the selected Model Name + Precision directly β see
What's New, 2026-09-30 for how it's started/stopped/logged, and
the External console checkbox for opening a plain terminal window instead.
This replaces the need to manually find and run the right .cmd file.
The Fix Image (Tab 4) and Fix Colors (Tab 5) tabs now support batch processing of multiple images:
Fix Image (Tab 4):
- New Enable batch processing checkbox β toggles between single-image and batch mode
- The image field is now a ComboBox showing all loaded images (drag & drop / Browse appends)
- Colorize processes all images sequentially against the DiT RPC server
- Overwrite overwrites all originals; Save As proposes a
*_colorizedwildcard mask - Outputs are kept in memory until explicitly saved; errors on single images are skipped
- Swap Output is automatically disabled in batch mode
Fix Colors (Tab 5):
- Same batch logic on the Target Image field β multiple targets against one color reference
- ComboBox, Clear, sequential colorization via local CMNET2, wildcard save
- Copy β Fix Image disabled in batch mode
Both tabs share the same interaction pattern for batch mode β ComboBox list,
sequential processing with progress counter, wildcard * save mask, memory-only
outputs until explicit save β while each tab uses its own backend (DiT RPC for
Fix Image, local CMNET2 for Fix Colors).
A standalone Fix Colors tab (Tab 5) has been added to the desktop GUI (GUI/CMNET2_colorize_client_GUI.py).
It colorizes a B&W or colorized target image using a color reference image via the local CMNET2 model (exemplar-based color propagation) β no RPC server required:
- Load a color reference image (drag & drop or Browse)
- Load a B&W target image (drag & drop or Browse)
- Colorize β runs
vscmnet2.pil_cmnet2_colorize()in a background thread. It allows to propagate the reference colors to target image.
Key features:
- Three preview panels: reference, target, and output sideβbyβside
- Copy β Fix Image: sends the output directly to Tab 4 (Fix Image) for a twoβstage pipeline (CMNET2 β DiT RPC)
- Save / Overwrite: save the colorized result as PNG/JPG or overwrite the original target file
- Full-resolution preservation: images are always kept at original resolution in memory; resizing only applies to previews
- Delayed import:
vscmnet2is imported only when Colorize is clicked (does not block GUI startup) - Backbone selection (
dinov3/dinov2, since 2026-09-28): passed tovscmnet2.pil_cmnet2_colorize()β same combo already available in Encode/Merge (Tab 3) and Fix Video (Tab 6)
Prerequisite:
vscmnet2must be installed with model weights and checkpoints present (see GUI README). No RPC connection needed.
The tab order has been updated: 1. Extraction β 2. Colorization β 3. Encode/Merge β 4. Fix Image β 5. Fix Colors β 6. Fix Video.
A standalone Fix Video tab (Tab 6) has been added to the desktop GUI (GUI/CMNET2_colorize_client_GUI.py).
It runs a VapourSynth + NVEnc pipeline to recolor a video using two reference images:
- Select a video and an encode VPY script
- Load two reference images (First / Last) via drag-and-drop or Browse
- Recolor β runs the VapourSynth β NVEnc pipeline and produces
_dt-recolor.mkv
Key features:
- NVEnc-only: uses GPU hardware encoding (NVEncC64.exe required)
- RefStart / RefEnd: reference images passed to the VapourSynth script as parameters
- RefDir auto-detection: set to the folder of the first reference image
- Configurable: FPS, VBR Quality, Memory Frames, Render Speed, Backbone
- MKV output:
.h265intermediate automatically muxed to.mkvand deleted - Pre-flight check: verifies NVEncC64.exe exists before starting
The Fix Video tab is independent of the batch pipeline and does not require the RPC server. Only the frames between RefStart / RefEnd will be recolored.
Prerequisite: NVEncC must be installed in
tools\NVEncC\(see GUI README).
A standalone Fix Image tab (Tab 4) has been added to the desktop GUI (GUI/CMNET2_colorize_client_GUI.py).
It allows single-image colorization with seed control, drag-and-drop file loading, and preview:
- Load a B&W image via drag-and-drop (
GUI/load_image_DtD_GUI.py) or the Browse button - Colorize with fixed seed (42) or random seed for variation
- Save the colorized result as PNG / JPG
The Fix Image tab is independent of the batch video pipeline and does not require VapourSynth.
If you already created the
.venvwith a previous version install the package tkinterDnD2 to add drag-and-drop support to tkinter
# Windows
.venv\Scripts\activate
(.venv) pip install tkinterDnD2Changed the GGUF configuration files. The pipeline Qwen-Image-Edit-2511 + Qwen-Image-Edit-2511-Lightning-4steps has substituted by the pipeline with Qwen-Image-Edit-2509 + Qwen-Image-Edit-2511-Lightning-4steps. This change has removed the artifacts problem which affected the colored images with the GGUF models and improved the overall quality of the colored images. It should be noted that, despite these improvements, the Nunchaku model remains the best and is the one recommended for production use (for systems with limited hardware resources, it is recommended to use the GGUFs of LongCat-Image-Edit-Turbo added on July 10th, 2027).
A FreeSimpleGUI desktop client (GUI/CMNET2_colorize_client_GUI.py) has been added to the project.
It orchestrates the full video colorization pipeline from a single graphical interface:
- Extract reference frames via VapourSynth + scene-change detection
- Colorize frames via the HAVC DiT Server (standard or paired inference)
- Encode the result as H.265 (x265 or NVEnc) or H.264 (x264)
- Merge the AI output with an existing color clip (optional, luminance-guided chroma blend)
See GUI/README_GUI.md for installation, setup, and usage instructions.
Prerequisite: the HAVC DiT Server must be running before the GUI can colorize frames.
- π¦ 4 backends, one API : nunchaku-qwen (FP4/INT4, 4 sec/frame) for speed, gguf-qwen and longcat-gguf (Q3, β¦, Q8, 12 sec/frame) for lower VRAM, qwen21-viggle (int8 ConvRot UNet, ~8-11 sec/frame) with optional Qwen3-VL prompt rewriting
- π¨ Batch colorization : process entire directories of B&W images via filesystem paths
- πΌοΈ Paired inference : colorize two images in a single forward pass (faster, temporally consistent)
- π‘ In-memory RPC : pass raw PNG frames over XML-RPC without touching the filesystem (ideal for video pipelines)
- β‘ 4-step lightning model : SVDQuant FP4 quantized transformer for maximum throughput
- π Thread-safe : pipeline loading and stop control are protected by locks; every RPC call runs in its own thread
- βοΈ Startup preload : optional
--load-pipelineflag loads the model at boot from a JSON config file - π Shared memory transport : zero-copy image transfer for same-host deployments (~23% faster than standard RPC)
Choose the backend that matches your hardware:
| Requirement | Details |
|---|---|
| GPU | NVIDIA RTX 30/40/50 (16 GB+ VRAM) |
| RAM | 64 GB+ |
| CUDA | 13.0 or newer |
| CUDA Toolkit | Must match the PyTorch build |
RTX 30/40-Series (Ampere / Ada): use
"model_precision": "int4". FP4 requires Blackwell (RTX 50). Requires Nunchaku 1.2.1 anddiffusers==0.37.0.dev0(wheel included inpackages/).
| Requirement | Details |
|---|---|
| GPU | NVIDIA RTX 30/40/50 (12 GB+ VRAM) |
| RAM | 32 GB+ |
| CUDA | 13.0+ (or CPU-only: slower, zero VRAM) |
Q3_K_S fits in 12 GB VRAM. Q4_K_S (default) balances quality and VRAM. Q5_K_M / Q6_K improve fidelity at higher VRAM cost. Q8_0 is near-lossless. Uses ComfyUI-native code : no ComfyUI GUI installation needed. Pre-made configs for all quantizations are in the
config/folder.
| Requirement | Details |
|---|---|
| GPU | NVIDIA RTX 30/40/50 (12 GB+ VRAM) |
| RAM | 32 GB+ |
| CUDA | 13.0+ |
LongCat-Image-Edit-Turbo delivers noticeably better colorization than the gguf-qwen model β richer colors, more natural skin tones, and better detail preservation β at the same ~12 s/frame speed.
Uses GGUF quantized UNet (Q3_K_M to Q8_0) + CLIP Q4_K_M. The UNet is distributed in five quantization levels to fit different VRAM budgets. All files are auto-downloaded on first run.
See
config/longcat_gguf_q*.jsonβ the general rule: lower quant = less VRAM. Launch withrun_server_longcat.cmd(Q4_K_M) orstart_server.cmd longcat|longcat-q3|....
| Requirement | Details |
|---|---|
| GPU | NVIDIA RTX 30/40/50 (14 GB+ VRAM) |
| RAM | 32 GB+ |
| CUDA | 13.0+ |
Native ComfyUI int8 ConvRot weights for the UNet (not GGUF β a GGUF UNet was evaluated but is ~2Γ slower for this model). The CLIP/text encoder, unlike the UNet, uses GGUF+mmproj by default (
Qwen3-VL-8B- Instruct-UD, see What's New) β a.safetensorsCLIP (int8_convrot/w4a8) remains a valid, simpler alternative, see Pipeline Configuration. Requirescomfy-kitchen==0.2.35andcomfy-aimdo==0.5.5exactly (pinned, not a minimum β both are compiled packages and an untested newer build is not assumed safe). A freshinstall.cmdrun sets these; an existing.venvneeds an explicit upgrade, see Quick Update. All files are auto-downloaded on first run β see What's New. Launch withrun_server_qwen21.cmd.
| Requirement | Details |
|---|---|
| OS | Windows 10/11 or Linux |
| Python | 3.12 |
Before setting up the project environment, make sure both Git and Python 3.12 are installed on your system.
Windows: download and install Git for Windows.
Accept the default options : in particular keep core.autocrlf=true (the default),
which ensures correct line endings for .cmd files.
Linux:
sudo apt install git # Debian / Ubuntu
sudo dnf install git # Fedora / RHELVerify: git --version
Windows: download the installer from python.org/downloads.
During installation, check "Add Python to PATH" : without this, python will not be
recognized in the terminal.
Linux:
sudo apt install python3.12 python3.12-venv # Debian / Ubuntu
sudo dnf install python3.12 # Fedora / RHELVerify: python --version (Windows) or python3.12 --version (Linux)
Clone the repository with git : this ensures correct line endings for all files
(.gitattributes is applied automatically at checkout):
git clone https://github.com/dan64/HAVCServerDiT.git
cd HAVCServerDiTThen create and activate the virtual environment inside the project directory:
python -m venv .venv
# Windows
.venv\Scripts\activate
# Linux / macOS
source .venv/bin/activateWindows quick-start: once the venv is active you can run
install.cmdto execute steps 2β6 automatically instead of running them one by one.
Use the stable build for all GPU generations (RTX 30 / 40 / 50):
pip install torch==2.10.0+cu130 torchvision==0.25.0+cu130 torchaudio==2.10.0+cu130 \
--index-url https://download.pytorch.org/whl/cu130Verify the installation:
python -c "import torch; print(torch.__version__); print(torch.cuda.is_available())"
# Expected: 2.10.0+cu130, True
β οΈ Do NOT usepip install nunchaku: that installs an unrelated package from PyPI with the same name that will fail withModuleNotFoundError: No module named 'nunchaku.models'.
Install the correct MIT Han Lab build directly from the GitHub release:
# Windows / Python 3.12 / CUDA 13.0 / PyTorch 2.10
pip install https://github.com/nunchaku-ai/nunchaku/releases/download/v1.2.1/nunchaku-1.2.1+cu13.0torch2.10-cp312-cp312-win_amd64.whlFor other platforms or Python versions, browse the full list of available wheels on the Nunchaku releases page and replace the filename accordingly.
Nunchaku pulls
torch>=2.0as a dependency (viaaccelerate) and may upgrade PyTorch to a newer version. After installing Nunchaku, re-pin PyTorch:
pip install torch==2.10.0+cu130 torchvision==0.25.0+cu130 torchaudio==2.10.0+cu130 --index-url https://download.pytorch.org/whl/cu130 --force-reinstallVerify the correct package is installed :
pip show nunchaku
# Version: 1.2.1+cu13.0torch2.10
pip show torch
# Version: 2.10.0+cu130Nunchaku 1.2.1 contains a bug in its transformer forward pass: txt_seq_lens is always
None at the point where it is passed to pos_embed, causing a ValueError with
diffusers >= 0.37.0.dev0. The included patch_nunchaku.py fixes this by deriving
max_txt_seq_len directly from encoder_hidden_states:
python patch_nunchaku.pyOn Windows you can also double-click patch_nunchaku.cmd or run it from a terminal:
patch_nunchaku.cmd # apply the patch
patch_nunchaku.cmd --check # check status without modifying files
patch_nunchaku.cmd --revert # revert to original (.bak backup)
You can verify the patch status at any time:
python patch_nunchaku.py --checkAnd revert to the original if needed (a .bak backup is created automatically):
python patch_nunchaku.py --revert
β οΈ Do NOT install diffusers from GitHub (pip install git+https://...). Nunchaku 1.2.1 requires exactly0.37.0.dev0. Later dev builds (β₯ 0.39.0) changed theQwenEmbedRopeAPI in a way that is incompatible even after the nunchaku patch.
A tested compatible wheel is included in the packages/ folder.
Install it directly:
pip install packages\diffusers-0.37.0.dev0-py3-none-any.whlVerify:
python -c "import diffusers; print(diffusers.__version__)"
# Expected: 0.37.0.dev0Pin the versions to match the tested working environment:
pip install \
transformers==4.57.6 \
accelerate==1.12.0 \
"huggingface_hub>=0.26.0" \
"Pillow>=10.0.0" \
scipy \
av \
torchsde \
gguf \
comfy-aimdo==0.5.5 \
comfy-kitchen==0.2.35Nunchaku users:
diffuserswas already installed in step 5 as the compatible0.37.0.dev0wheel. Do NOT upgrade it : nunchaku 1.2.1 requires exactly that version.
safetensorsis pulled automatically by diffusers.
scipy,av, andtorchsdeare required by the diffusers pipeline.gguf,comfy-aimdo, andcomfy-kitchenare required by the GGUF backends (gguf-qwen/longcat-gguf) and byqwen21-viggle. The pinned versions here (0.5.5/0.2.35) are required specifically forqwen21-viggleβ exact pins, not just a minimum, to avoid drifting to an untested newer build of these compiled packages.
dit-colorize-rpc/
βββ dit_rpc_server.py # XML-RPC server (entry point)
βββ dit_colorize_main.py # Colorization pipeline and image utilities
βββ dit_client_example.py # Example RPC client : single frame
βββ dit_client_pair_example.py # Example RPC client : paired inference
βββ patch_nunchaku.py # Compatibility patch for nunchaku 1.2.1
βββ config/ # Pipeline configs (nunchaku FP4/INT4, gguf/longcat Q3βQ8, qwen21-viggle)
βββ comfy_bridge/ # Self-contained ComfyUI runtime (gguf-qwen/longcat-gguf/qwen21-viggle)
βββ install.cmd # Windows automated installer
βββ start_server.cmd # Windows launcher : server (nunchaku/gguf/longcat)
βββ run_server_longcat.cmd # Windows launcher : LongCat server
βββ run_server_qwen21.cmd # Windows launcher : qwen21-viggle server
βββ run_client_example.cmd # Windows launcher : single frame example
βββ run_client_pair_example.cmd # Windows launcher : paired inference example
βββ patch_nunchaku.cmd # Windows launcher : nunchaku patch
βββ assets/
β βββ santa_bw.png # Sample B&W image (single frame test)
β βββ sample1_bw.jpg # Sample B&W image 1 (paired inference test)
β βββ sample2_bw.jpg # Sample B&W image 2 (paired inference test)
βββ packages/
β βββ diffusers-0.37.0.dev0-py3-none-any.whl # Tested compatible diffusers build
βββ README.md
Ready-to-use config files for both backends are in the config/ folder.
Pick the one that matches your hardware and pass it to --pipeline-config.
{
"model_name": "nunchaku-qwen",
"model_precision": "fp4",
"model_rank": "32",
"model_inference_steps": "4",
"cache_dir": "",
"full_model_path": ""
}{
"model_name": "nunchaku-qwen",
"model_precision": "int4",
"model_rank": "32",
"model_inference_steps": "4",
"cache_dir": "",
"full_model_path": ""
}
β οΈ model_precision: use"fp4"only on RTX 50-Series (Blackwell). On RTX 30 / 40-Series use"int4": FP4 kernels require sm_120 and will fail on older architectures.
GGUF Backend : config/qwen_gguf_q3.json β¦ qwen_gguf_q8.json / config/longcat_gguf_q3.json β¦ longcat_gguf_q8.json
Five quantization levels are available. All share the same structure with
model_name: "gguf-qwen" and a quant field that selects the quantization:
| Config file | quant |
UNet | CLIP |
|---|---|---|---|
qwen_gguf_q3.json |
"q3" |
β¦Q3_K_S.gguf |
β¦Q3_K_S.gguf |
qwen_gguf_q4.json |
"q4" |
β¦Q4_K_S.gguf |
β¦Q4_K_S.gguf |
qwen_gguf_q5.json |
"q5" |
β¦Q5_K_M.gguf |
β¦Q5_K_M.gguf |
qwen_gguf_q6.json |
"q6" |
β¦Q6_K.gguf |
β¦Q6_K.gguf |
qwen_gguf_q8.json |
"q8" |
β¦Q8_0.gguf |
β¦Q8_0.gguf |
Q4 is the recommended default : good quality/VRAM balance, but even Q3 is capable of delivering frames with acceptable colors. All quants share the same VAE, mmproj, and LoRA files (auto-downloaded from HuggingFace).
β οΈ The GGUF backend is experimental. In some cases the frames colors may be faded or little colored. For production use, prefernunchaku-qwen(FP4/INT4) orqwen21-viggle, which are not affected by such problems β see the recommendation note at the top of this README.
Config example (config/qwen_gguf_q4.json):
{
"model_name": "gguf-qwen",
"quant": "q4",
"unet_gguf": "models/unet/qwen-image-edit-2511-Q4_K_S.gguf",
"clip_gguf": "models/clip/Qwen2.5-VL-7B-Instruct-Q4_K_S.gguf",
"clip_mmproj": "models/clip/Qwen2.5-VL-7B-Instruct-mmproj-BF16.gguf",
"clip_mmproj_hf_name": "mmproj-BF16.gguf",
"vae_name": "qwen_image_vae.safetensors",
"lora_path": "models/loras/Qwen-Image-Edit-2511-Lightning-4steps-V1.0-bf16.safetensors",
"steps": 4,
"hf_unet": "unsloth/Qwen-Image-Edit-2511-GGUF",
"hf_clip": "unsloth/Qwen2.5-VL-7B-Instruct-GGUF",
"hf_vae": "Comfy-Org/Qwen-Image_ComfyUI",
"hf_lora": "lightx2v/Qwen-Image-Edit-2511-Lightning"
}
clip_mmproj_hf_nameexists because the mmproj file's name on HuggingFace (a genericmmproj-BF16.gguf, shared across many unrelated repos) rarely matches the locally-prefixed name you actually want on disk β it tells the downloader what to fetch,clip_mmprojis where it ends up and what the loader looks for locally. Omit it and the downloader falls back to usingclip_mmproj's own filename as the remote name too, which only works if they happen to match.
The LoRA file Qwen-Image-Edit-2511-Lightning-4steps-V1.0-bf16.safetensors enables 4-step inference (down from 20-50 steps without LoRA). It is a ComfyUI-format LoRA that gets merged directly into the transformer at load time.
- With LoRA: call
colorize_image(..., steps=4): fast, same quality - Without LoRA: set
full_model_pathto""and usesteps=20or higher
The LoRA is merged statically (not applied as an adapter), so there is no runtime overhead.
A single config file β this backend has no quantization variants for the
UNet (int8 ConvRot only). The CLIP, unlike the UNet, can be either a
.safetensors file or a GGUF+mmproj pair β the default uses GGUF+mmproj
(see What's New):
{
"model_name": "qwen21-viggle",
"unet_name": "models/unet/qwen_image_2.1_int8_convrot.safetensors",
"clip_name": "models/clip/Qwen3-VL-8B-Instruct-UD-Q4_K_XL.gguf",
"clip_mmproj": "models/clip/Qwen3-VL-8B-Instruct-mmproj-BF16.gguf",
"clip_mmproj_hf_name": "mmproj-BF16.gguf",
"vae_name": "qwen_image_2.1_vae_bf16.safetensors",
"lora_path": "models/loras/Qwen-Image-2.1-viggle-turbo-v0.2.1-6step-lora-r128.safetensors",
"steps": 6,
"hf_unet": "Comfy-Org/Qwen-Image-2.1",
"hf_clip": "unsloth/Qwen3-VL-8B-Instruct-GGUF",
"hf_vae": "Comfy-Org/Qwen-Image-2.1",
"hf_lora": "Viggle/Qwen-Image-2.1-viggle-turbo"
}Note the different key names from the GGUF-backend format above:
unet_name/clip_name(notunet_gguf/clip_gguf) β but unlike the GGUF backend,clip_namehere can point to either a.safetensorsfile or a.gguffile (the loader picks the right code path from the extension).clip_mmproj/clip_mmproj_hf_nameonly apply whenclip_nameis a.gguffile β omit both to use a.safetensorsCLIP instead:"clip_name": "models/text_encoders/qwen3vl_8b_int8_convrot.safetensors",
steps: 6here only documents the LoRA's native step count for anyone reading the file; the actual number of steps used at inference time is thestepsargument passed per-call to the colorization RPC methods (see Suggested Inference Steps and RPC API Reference), same as every other backend.
| Key | Required | Description |
|---|---|---|
model_name |
β | "nunchaku-qwen", "gguf-qwen", "longcat-gguf", or "qwen21-viggle" |
quant |
GGUF only: quantization level ("q3", "q4", "q5", "q6", "q8"). Default: "q4" |
|
model_precision |
β | Nunchaku: "fp4" (RTX 50) or "int4" (RTX 30/40). GGUF/qwen21-viggle: not used |
unet_gguf / clip_gguf |
β | GGUF only: local paths to the GGUF model files |
unet_name / clip_name |
β | qwen21-viggle only: local paths to the model files β unet_name is always .safetensors, clip_name can be .safetensors or .gguf |
clip_mmproj / clip_mmproj_hf_name |
GGUF/qwen21-viggle-with-GGUF-CLIP: local path to the mmproj (vision tower) file / its filename on HuggingFace if different from the local one. Required for a GGUF CLIP to see images at all β without it the vision tower silently isn't loaded | |
model_rank |
Nunchaku: SVD rank ("32"). GGUF/qwen21-viggle: not used |
|
model_inference_steps |
Nunchaku: diffusion steps ("4"). GGUF/qwen21-viggle: not used at load time |
|
cache_dir |
HuggingFace cache directory. Leave empty to use the default ~/.cache/huggingface |
|
full_model_path |
Nunchaku: local path to the transformer checkpoint. GGUF: not used | |
lora_path |
GGUF/qwen21-viggle: path to the LoRA (.safetensors). Omit for GGUF to skip LoRA merging |
|
steps |
GGUF: inference steps (4 with LoRA, 20 without). qwen21-viggle: documents the native step count only, not consumed at load time |
|
vae_name |
GGUF/qwen21-viggle only: VAE filename | |
hf_* |
GGUF/qwen21-viggle only: HuggingFace repo names for auto-download |
python dit_rpc_server.py# RTX 50-Series
python dit_rpc_server.py --load-pipeline --pipeline-config config/qwen_nunchaku_fp4.json
# RTX 30 / 40-Series
python dit_rpc_server.py --load-pipeline --pipeline-config config/qwen_nunchaku_int4.json
# GGUF (any quantization)
python dit_rpc_server.py --load-pipeline --pipeline-config config/qwen_gguf_q3.jsonOn Windows you can also use the provided start_server.cmd (see Windows launch script).
usage: dit_rpc_server.py [-h] [--host HOST] [--port PORT]
[--logfile LOGFILE] [--module-dir MODULE_DIR]
[--load-pipeline] [--pipeline-config CONFIG.json]
options:
--host HOST Address to listen on (default: 127.0.0.1)
--port PORT TCP port (default: 8765)
--logfile LOGFILE Optional path for a log file
--module-dir MODULE_DIR Directory containing dit_colorize_main.py
(default: same directory as this script)
--load-pipeline Load the colorization pipeline at startup
--pipeline-config CONFIG.json
Path to the JSON pipeline config file
(required when --load-pipeline is set)
Connect from any Python client using xmlrpc.client:
import xmlrpc.client
proxy = xmlrpc.client.ServerProxy("http://127.0.0.1:8765/", use_builtin_types=True)All methods return a dict with at least {"ok": bool, "msg": str}.
| Method | Returns | Description |
|---|---|---|
ping() |
"pong" |
Connectivity check |
| Method | Returns | Description |
|---|---|---|
load_pipeline(model_name, model_precision, model_rank, model_inference_steps, cache_dir="", full_model_path="", vae_name="", hf_unet="", hf_clip="", hf_vae="", hf_lora="") |
{"ok", "msg"} |
Load the model into VRAM. The vae_name/hf_* arguments are only meaningful for gguf-qwen/longcat-gguf/qwen21-viggle β omit for nunchaku-qwen |
is_pipeline_loaded() |
bool |
True if the pipeline is ready |
unload_pipeline() |
{"ok", "msg"} |
Release VRAM |
| Method | Returns | Description |
|---|---|---|
request_stop() |
bool |
Ask the server to refuse new colorization calls |
clear_stop() |
bool |
Reset the stop flag before a new batch |
is_stop_requested() |
bool |
Check the current stop flag |
| Method | Returns | Description |
|---|---|---|
colorize_image(in_path, out_path, prompt, img_size=0, steps=2, enhance_prompt=False) |
{"ok", "elapsed", "skipped", "msg"} |
Single image, paths on the server filesystem |
colorize_image_pair(img1_path, img2_path, out_dir, prompt, gap_px=8, steps=2, enhance_prompt=False) |
{"ok", "elapsed", "msg"} |
Two images, single inference pass |
colorize_single_image(img_path, out_dir, prompt, steps=2, enhance_prompt=False) |
{"ok", "elapsed", "msg"} |
Single image fallback (odd batch end) |
| Method | Returns | Description |
|---|---|---|
colorize_frame(img_data, prompt, img_size=0, steps=2, seed=42, skip_bw=False, enhance_prompt=False) |
{"ok", "data", "elapsed", "skipped", "msg"} |
Single frame as raw PNG bytes |
colorize_frame_pair(img1_data, img2_data, prompt, gap_px=8, steps=2, enhance_prompt=False) |
{"ok", "data1", "data2", "elapsed", "skipped1", "skipped2", "msg"} |
Two frames, single inference pass |
skipped=Truemeans the frame was too dark to colorize (average brightness < 9/255). The returneddatafield contains the unchanged input in that case.
| Method | Returns | Description |
|---|---|---|
colorize_frame_shm(shm_in, shm_out, h, w, prompt, img_size=0, steps=2, seed=42, skip_bw=False, enhance_prompt=False) |
{"ok", "elapsed", "skipped", "msg"} |
Single frame via shared memory |
colorize_frame_pair_shm(shm_in1, shm_out1, h1, w1, shm_in2, shm_out2, h2, w2, prompt, gap_px=8, steps=4, enhance_prompt=False) |
{"ok", "elapsed", "skipped1", "skipped2", "msg"} |
Two frames via shared memory, single inference pass |
enhance_prompt(all methods above, defaultFalse) rewrites the prompt via Qwen3-VL before colorizing β only meaningful forqwen21-viggle; silently has no effect on the other backends. See What's New. See Shared Memory Transport for usage details.
Both clients support two transport modes selectable via --use-shm:
| Mode | Flag | When to use | Measured speed (1480Γ1080 px pair) |
|---|---|---|---|
| Standard RPC | (default) | Any deployment, including remote server | ~5.25s/image |
| Shared memory | --use-shm |
Server and client on the same host only | ~4.06s/image (~23% faster) |
The pipeline must be loaded on the server before running the clients. Start the server with
--load-pipeline --pipeline-config CONFIG.json.
Colorizes assets/santa_bw.png and saves the result as assets/santa_colorized.png.
# standard RPC : works with local and remote server
python dit_client_example.py
# shared memory : same-host only, lower latency
python dit_client_example.py --use-shmWindows: run_client_example.cmd
To enable shared memory edit run_client_example.cmd and set USE_SHM=1.
Colorizes assets/sample1_bw.jpg and assets/sample2_bw.jpg in a single forward
pass, saving assets/sample1_colorized.jpg and assets/sample2_colorized.jpg.
Paired inference places the two images side-by-side and runs one inference instead of two, roughly halving the per-image cost (~5.25s/image vs ~11s standalone). Combined with shared memory transport this reaches ~4.06s/image.
# standard RPC
python dit_client_pair_example.py
# shared memory : same-host only
python dit_client_pair_example.py --use-shmWindows: run_client_pair_example.cmd
To enable shared memory edit run_client_pair_example.cmd and set USE_SHM=1.
--host HOST Server host (default: 127.0.0.1)
--port PORT Server port (default: 8765)
--prompt PROMPT Text prompt for the model
--steps N Number of steps for inference (default:4)
--use-shm Use shared memory transport (same-host only)
Additional argument for the paired client:
--gap-px N Separator width in pixels between the two
images in the merged input (default: 8)
The standard RPC transport serializes each image as a PNG byte stream, encodes it in Base64, sends it over a TCP socket, and decodes it on the other side. For a 1480Γ1080 frame this is roughly 4β5 MB per round trip.
The shared memory transport bypasses the network entirely. The client writes the raw pixel array directly into a shared memory segment; the server attaches to the same segment and reads the pixels without any copy. Only the metadata (segment name, dimensions, prompt) travels over the XML-RPC socket.
Requirement: server and client must run on the same machine.
If the server is on a dedicated GPU machine and the client is on a separate workstation,
shared memory is not available : use the standard RPC transport instead (default).
The clients detect this automatically: passing --use-shm when the host is not
127.0.0.1 / localhost prints a warning and falls back to standard RPC.
Measured on a 1480Γ1080 pixel pair (RTX 5070 Ti, FP4, paired inference):
| Transport | Per-image time | Round-trip overhead |
|---|---|---|
| Standard RPC (PNG) | ~5.25s | ~1.1s |
| Shared memory | ~4.06s | ~0.16s |
| Gain | ~23% faster | ~7Γ less overhead |
The round-trip overhead with shared memory is essentially zero : the 0.16s gap between inference time and wall-clock time is just Python function call and numpy overhead.
On a 100k-frame video processed as pairs (50k inference calls) the cumulative saving is:
(5.25 - 4.06) Γ 50,000 β 16.5 hours
The client owns and manages all shared memory segments. The server is fully stateless with respect to shared memory : it only attaches, reads/writes, and detaches.
Client Server
β β
β create shm_in (h Γ w Γ 3 bytes) β
β create shm_out (h Γ w Γ 3 bytes) β
β write raw RGB pixels β shm_in β
β β
β RPC(shm_in_name, shm_out_name, h, w, β¦) ββΊβ
β β attach shm_in β PIL Image
β β inference
β β result β shm_out
βββ return {elapsed, skipped, β¦} ββββββββββββ
β β detach both segments
β read shm_out β PIL Image β
β unlink shm_in + shm_out β
From the command line:
python dit_client_pair_example.py --use-shm
python dit_client_example.py --use-shmFrom the Windows .cmd launchers, edit the user configuration block and set:
set USE_SHM=1The banner will confirm the active transport:
Transport : 1 (0=RPC 1=shared memory)
And the Python client will print:
[INFO] Transport: shared memory
import uuid
import numpy as np
from multiprocessing.shared_memory import SharedMemory
from PIL import Image
def colorize_pair_shm(proxy, img1: Image.Image, img2: Image.Image, prompt: str):
arr1, arr2 = np.array(img1), np.array(img2)
h1, w1 = arr1.shape[:2]
h2, w2 = arr2.shape[:2]
uid = uuid.uuid4().hex[:12]
# Create all four segments (client owns them)
segs = {
tag: SharedMemory(name=f"dit_{tag}_{uid}", create=True, size=h*w*3)
for tag, h, w in [("in1",h1,w1),("out1",h1,w1),("in2",h2,w2),("out2",h2,w2)]
}
try:
np.ndarray((h1,w1,3), dtype=np.uint8, buffer=segs["in1"].buf)[:] = arr1
np.ndarray((h2,w2,3), dtype=np.uint8, buffer=segs["in2"].buf)[:] = arr2
result = proxy.colorize_frame_pair_shm(
segs["in1"].name, segs["out1"].name, h1, w1,
segs["in2"].name, segs["out2"].name, h2, w2,
prompt, 8, # gap_px
)
out1 = Image.fromarray(
np.ndarray((h1,w1,3), dtype=np.uint8, buffer=segs["out1"].buf).copy())
out2 = Image.fromarray(
np.ndarray((h2,w2,3), dtype=np.uint8, buffer=segs["out2"].buf).copy())
return result, out1, out2
finally:
for shm in segs.values():
shm.close(); shm.unlink()start_server.cmd is a ready-to-use launcher for Windows.
Edit the variables at the top of the file to match your setup, then double-click it or run it from a terminal.
start_server.cmd [q3|q4|q5|q6|q8|fp4|int4|longcat]
| Argument | Backend | Quantization | VRAM |
|---|---|---|---|
| (none) | GGUF | Q4_K_S | 12 GB |
q3 |
GGUF | Q3_K_S | 12 GB |
q4 |
GGUF | Q4_K_S | 12 GB |
q5 |
GGUF | Q5_K_M | 16 GB |
q6 |
GGUF | Q6_K | 18 GB |
q8 |
GGUF | Q8_0 | 22 GB |
fp4 |
Nunchaku | FP4 | 16 GB |
int4 |
Nunchaku | INT4 | 16 GB |
longcat |
LongCat | Q4_K_M | 12 GB |
If no argument is passed it defaults to q4 (Q4_K_S). Use int4 for RTX 30 / 40-Series Nunchaku:
start_server.cmd int4
Convenience wrappers β double-click or run from terminal without arguments:
| File | Equivalent command | Backend |
|---|---|---|
run_server_q3.cmd |
start_server.cmd q3 |
GGUF Q3_K_S |
run_server_fp4.cmd |
start_server.cmd fp4 |
Nunchaku FP4 |
run_server_int4.cmd |
start_server.cmd int4 |
Nunchaku INT4 |
run_server_longcat.cmd |
start_server.cmd longcat |
LongCat Q4_K_M |
run_server_qwen21.cmd is a separate, standalone launcher for
qwen21-viggle β it does not take an argument (start_server.cmd qwen21-viggle is not a thing), it always launches with
config/qwen21_viggle.json (the only config available for this backend).
GUI shortcut: From the desktop GUI, go to Tab 2 (Colorization), pick a Model + Precision, and click Run Server β the GUI starts the server itself, with live output in the Server Log tab and auto-connect once it's ready (see What's New, 2026-09-30). Tick External console first to instead open a plain terminal window with the correct
.cmdfile/arguments for the selected Model Name (run_server_qwen21.cmdwhenqwen21-viggleis selected,start_server.cmdwith the right arguments otherwise).
| Model Family | Recommended Steps | Notes |
|---|---|---|
| Qwen (nunchaku fp4/int4) | 2 | Good results with 2 steps when using lightning LoRA |
| Qwen (gguf q3βq8) | 2 | Default in config files; 4 steps possible but slower |
| LongCat (longcat-gguf) | 8 | Calibrated for 8 steps; best quality at 8 steps; 4 steps possible but colors are faded |
| Qwen-Image-2.1 (qwen21-viggle) | 6 | LoRA's native step count. 2/4/8 are experimental alternate schedules β 2 in particular trades a little brightness accuracy for ~25% less time, worth trying |
Prompt tip (qwen21-viggle): on subjects with a strong color convention (e.g. a well-known costume), the model can leave the color ambiguous and resolve it inconsistently between runs. Before reaching for
enhance_prompt(which adds ~15-20s/frame), try a direct, explicitly anti-hedging prompt β it solves the same problem for free in most cases:"Add color to this black-and-white image without hesitation regarding the appropriate colors. For any subject, garment, object, or setting where the color is common knowledge or established by convention, confidently apply the expected color; otherwise use natural colors. Color the image by strictly preserving all shapes, outlines, and background details."
Avoid naming specific example subjects in this prompt (e.g. "like a stop sign") β the model may render that literal example into the scene instead of just using it as a color reference.
CUDA out of memory
Close other GPU applications. On 16 GB cards the server automatically enables sequential CPU offload for layers that do not fit in VRAM.
dit_colorize_main.py NOT FOUND
Use --module-dir to point the server to the directory that contains dit_colorize_main.py:
python dit_rpc_server.py --module-dir /path/to/dit_colorize_mainModel 'xxx' is not supported
Supported values for model_name are "nunchaku-qwen" (FP4/INT4), "gguf-qwen" (Q3_K_S, Q4_K_S, Q5_K_M, Q6_K, Q8_0), "longcat-gguf", and "qwen21-viggle". For "gguf-qwen", the quantization is selected via the quant field in the config (e.g. "q4").
Pipeline takes a long time to load
Nunchaku: on the first run the model weights (~15β30 GB) are downloaded from HuggingFace.
Subsequent runs load from the local cache.
GGUF: only the VAE and tokenizer (~320 MB) are downloaded from HuggingFace; the UNet and CLIP are loaded directly from the local .gguf files. Set cache_dir in the config to control where the cache is stored.
- Model: Qwen/Qwen-Image-Edit-2511, Qwen/Qwen-Image-2.1, LongCat-Image-Edit-Turbo
- VapourSynth filter for video colorization with CMNET2: vs-cmnet2
- Viggle-Turbo LoRA: Viggle/Qwen-Image-2.1-viggle-turbo
- Nunchaku quantization: Nunchaku / SVDQuant
- GGUF dequantization kernels: adapted from ComfyUI-GGUF (Apache 2.0), Qwen3-VL mmproj support from the ComfyUI-GGUF-Reboot fork (molbal)
- Pipeline: Hugging Face Diffusers




