Skip to content

Latest commit

Β 

History

95 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

HAVC Server DiT

Hybrid Automatic Video Colorizer (HAVC) server that exposes a GPU-accelerated colorization pipeline for black-and-white images and video frames based on Diffusion Transformer (DiT) models. 4 backends, one API : pick the one that fits your hardware:

  • nunchaku-qwen: SVDQuant FP4/INT4 transformer via Nunchaku : 4 sec/frameΒΉ, requires RTX 30/40/50 (16GB+ VRAM , 64GB RAM) & CUDA 13.0
  • gguf-qwen: ComfyUI-native GGUF pipeline (Q3_K_S, Q4_K_S, Q5_K_M, Q6_K, Q8_0) : 12 sec/frameΒ², runs on RTX 30/40/50 (12GB+ VRAM, 32GB+ RAM), zero ComfyUI GUI dependency
  • longcat-gguf: LongCat-Image-Edit-Turbo GGUF pipeline (Q3_K_M–Q8_0) : ~12 sec/frameΒ², runs on RTX 30/40/50 (12GB+ VRAM, 32GB+ RAM), better image quality than gguf-qwen, zero ComfyUI GUI dependency
  • qwen21-viggle: Qwen-Image-2.1 (native ComfyUI int8 ConvRot UNet) + Viggle-Turbo LoRA : ~6 sec/frameΒΉ (steps=2) or ~8-11 sec/frame (steps=6, single-image, see What's New), runs on RTX 30/40/50 (14GB+ VRAM, 32GB+ RAM), optional enhance_prompt (Qwen3-VL image-aware prompt rewriting), zero ComfyUI GUI dependency

ΒΉ Measured with Fast Pipeline (paired inference, two frames per forward pass) at the backend's fastest recommended step count. Β² gguf-qwen/longcat-gguf don't support paired inference (fall back to per-image processing, see What's New) β€” their figure is a genuine single-image time, not directly comparable to the Fast Pipeline figures above.

Recommended: nunchaku-qwen and qwen21-viggle are both recommended for production use, nunchaku-qwen at the fastest usable step count (steps=2) has an inference speed of about 4 sec/frame using Fast Pipeline, qwen21-viggle at fastest usable step count (steps=2) has an inference speed of about 8 sec/frame; using Fast Pipeline the speed improves to about 6 sec/frame (not 4 β€” the pair-mode working resolution was deliberately raised for this backend to avoid a color artifact, see the ⚠️ Fast Pipeline note below). qwen21-viggle needs meaningfully less hardware (14GB+ VRAM / 32GB+ RAM vs. 16GB+ VRAM / 64GB+ RAM). longcat-gguf remain the choice for VRAM-constrained setups where neither of the above fits, at a real speed cost (see the ⚠️ Experimental note under GGUF below).

⚠️ Fast Pipeline + qwen21-viggle: paired inference uses a higher working resolution for this backend specifically (1280 vs. 1024 for single images and for the other backends, since What's New 2026-09-27) β€” this fixes a color artifact previously seen on fine detail near the merge boundary (e.g. a hand rendered in tones close to the surrounding foliage), at the cost of some speed (~6 sec/frame instead of 4). A milder residual effect can still appear on secondary, color-ambiguous details (e.g. a flower's petals taking a noticeably different but still plausible hue between runs) β€” not a defect on the same order as the original artifact, more of the same color-hedging behavior described elsewhere in this README. The Viggle-Turbo LoRA is still an experimental release, and this residual effect may be a limitation of the LoRA itself. If maximum consistency matters more than speed, disable Fast Pipeline for qwen21-viggle (~8 sec/frame, no longer speed-competitive with nunchaku-qwen) or spot-check the output before a long batch run.

Color stability vs. variety: based on real-world use across thousands of frames, nunchaku-qwen tends to show more color variability between similar frames β€” can look more vivid, but with weaker frame-to-frame consistency β€” while qwen21-viggle is more conservative in its color choices and more stable, likely a consequence of the Viggle-Turbo LoRA's aggressive step-distillation, which tends to narrow the range of plausible outputs toward "safe" choices. For video work, where flickering color between consecutive frames is a visible defect, this makes qwen21-viggle's conservatism a practical advantage rather than just a stylistic difference β€” worth factoring in alongside the speed/hardware trade-offs above. The same pattern shows up specifically in Fast Pipeline (paired inference): when both frames share an object, qwen21-viggle consistently colors it the same way in both halves, while nunchaku-qwen is less reliable at this β€” the exact cause (the LoRA itself vs. something more general about the two pipelines) is not established.


πŸ“¦ Quick Update (existing installation)

If you already have the .venv with CUDA 13.0 and just need to update the project to the latest version, follow these steps:

Shortcut: quick_update.cmd automates all of the steps below (including the comfy-kitchen/comfy-aimdo pin) β€” activate the .venv first, then double-click it or run it from a terminal. The manual steps are documented here for reference and for non-Windows setups.

# 1) Pull the latest code
git pull

# 2) Activate the virtual environment
.venv\Scripts\activate

# 3) Install / update the GUI dependencies (if new packages were added)
pip install -r GUI\requirements.txt

# 4) Update vscmnet2 (if a newer wheel is available in packages/)
pip install packages\vscmnet2-1.1.0-py3-none-any.whl

# 5) Re-apply the Nunchaku patch
python patch_nunchaku.py

# 6) Required for qwen21-viggle (see What's New, 2026-09-26): pin
#    comfy-kitchen and comfy-aimdo to the tested versions β€” install.cmd
#    only sets these for a fresh install, an existing .venv needs this
#    explicitly
pip install comfy-kitchen==0.2.35
pip install comfy-aimdo==0.5.5

# 7) Verify everything is up-to-date
pip show torch          # Expected: 2.10.0+cu130
pip show nunchaku       # Expected: 1.2.1+cu13.0torch2.10
pip show comfy-kitchen  # Expected: 0.2.35
pip show comfy-aimdo    # Expected: 0.5.5

Note: steps 4–5 are only needed if packages/ or patch_nunchaku.py have changed. Step 6 is only needed to use qwen21-viggle β€” the other three backends work with the older comfy-kitchen/comfy-aimdo versions. Check git log --oneline -5 to see what was updated.


πŸ”„ Upgrading from CUDA 12.8 to 13.0

If you already created the .venv with a previous version (CUDA 12.8, PyTorch 2.9.1, Nunchaku cu12.8torch2.9), upgrade to get these benefits:

Improvement Before (12.8) After (13.0)
CUDA allocator native (slower reallocation) cudaMallocAsync (async, ~10 % faster memory ops)
comfy-kitchen CUDA disabled: True (fallback to eager) disabled: False (native dequantization kernels)
Warning You need pytorch with cu130 or higher gone (build matches Nunchaku)

Upgrade steps:

# 1) Deactivate and reactivate the venv to ensure a clean shell
deactivate
.venv\Scripts\activate

# 2) Upgrade PyTorch to 2.10 + CUDA 13.0
pip install torch==2.10.0+cu130 torchvision==0.25.0+cu130 torchaudio==2.10.0+cu130 --index-url https://download.pytorch.org/whl/cu130 --force-reinstall

# 3) Upgrade Nunchaku (CUDA 13.0 + PyTorch 2.10 build)
pip install https://github.com/nunchaku-ai/nunchaku/releases/download/v1.2.1/nunchaku-1.2.1+cu13.0torch2.10-cp312-cp312-win_amd64.whl --force-reinstall

# 4) Re-pin PyTorch (Nunchaku may have upgraded it to 2.12)
pip install torch==2.10.0+cu130 torchvision==0.25.0+cu130 torchaudio==2.10.0+cu130 --index-url https://download.pytorch.org/whl/cu130 --force-reinstall

# 5) Re-apply the Nunchaku patch
python patch_nunchaku.py

# 6) Verify
pip show torch       # Expected: 2.10.0+cu130
pip show nunchaku    # Expected: 1.2.1+cu13.0torch2.10

πŸ“’ What's New

2026-10-02 β€” Run Server managed by the GUI, with a live Server Log (GUI)

Run Server (Tab 2, Colorization) now starts and stops the RPC server itself instead of just launching a .cmd file in its own terminal window: the server runs as a hidden child process controlled by the GUI, and its full output streams live into a new Server Log tab β€” right next to the existing log, now labeled App Log β€” in the Dashboard. The button becomes Stop Server while it's running, and a status line next to it tracks the sequence: Starting server on ... β†’ running on ... (once the server actually reports it's listening) β†’ stopped.

Once the server reports it's ready, the GUI connects automatically β€” no need to also click Connect on Tab 2. Closing the GUI (or clicking Stop Server) always shuts the process down cleanly; the RPC connection indicator resets to Disconnected at the same time, since the server it was talking to is gone.

A new External console checkbox next to the button restores the previous behavior exactly (a separate visible console window, started and left running independently of the GUI) for anyone who prefers it or needs to keep an eye on the raw console.

A Local DiT Server frame on the Dashboard mirrors the Run Server button and its status text, so the server can be started/stopped without switching to Tab 2.

START PIPELINE also uses this: if the 3. Colorize Frames (AI) task is enabled and the client isn't connected yet, the GUI starts the DiT server automatically (same as clicking Run Server) and holds the pipeline start until it reports it's actually online, instead of just failing with "not connected". This only applies when Run Server is GUI-managed (External console unchecked) β€” with an external console there's no way for the GUI to know when that separate process is ready, so the previous behavior (an error asking to connect manually) still applies there.

2026-10-01 β€” Select Reference Frames task (GUI)

A new optional Dashboard task, 2. Select Reference Frames, has been added between Extract and Colorize β€” every task after it, is renumbered (Colorize/Encode/Merge become tasks 3/4/5). It deduplicates the reference frames extracted in Step 1 by semantic similarity (DINOv3-based, via vscmnet2.vs_select_reference_frames()), reducing redundant near-identical candidates before they reach colorization β€” useful for long or slow-changing scenes where scene-change detection alone still produces many visually similar frames.

The task renames ref_tht10/ (produced by Extraction) to ref_tht10_temp/, then writes the deduplicated representative frames back to a freshly created ref_tht10/ β€” the same folder Colorize already reads from, so no other step changes. ref_tht10_temp/ is kept as a full backup of every extracted candidate unless Move Files is checked. If ref_tht10/ is missing/empty, or ref_tht10_temp/ already exists from an interrupted previous run, the task stops the entire pipeline with an error rather than guessing or overwriting anything.

New Selection Settings frame in the GUI's Extraction tab exposes similarity_threshold, select_window, and the Dry Run/Debug HTML/ Move Files options. If Debug HTML is checked, in the output folder is written the file cluster_debug.html. This files contains all the reference clusters as shown in the image below

Reference Selection

for example in the Cluster 2, the reference frame #000145 was selected to represent all the references included in the Cluster 2. If the parameter similarity threshold is set above 0.95 will be selected smaller clusters, vice-versa if the threshold is set below 0.95 the similarity clusters will be bigger (will be available less reference frame to colorize).

This deduplication of keyframes will improve color consistency and accelerate the coloring process, as fewer images will need to be colored.

See GUI README: Tab 1 for the full workflow and recovery steps if a run is interrupted.

Existing gui_cmnet2_settings.json files are migrated automatically on next load β€” no manual action needed.

2026-09-30 β€” Fix Colors: Backbone selection (GUI)

The Fix Colors tab (Tab 5) now exposes a Backbone combo (dinov3 / dinov2), passed as the backbone parameter of vscmnet2.pil_cmnet2_colorize() β€” the same choice already available in Encode/Merge (Tab 3) and Fix Video (Tab 6), now consistent across all three tabs that drive CMNET2. Previously the tab always used the vscmnet2 default (dinov3) with no way to select the legacy DINOv2 backbone. Applies in both single-image and batch mode. Persisted in gui_cmnet2_settings.json as fixc_backbone.

2026-09-29 β€” qwen21-viggle: GGUF+mmproj CLIP as new default, clip_mmproj generalized, known limitation documented

New default CLIP for qwen21-viggle: Qwen3-VL-8B-Instruct-UD (GGUF+mmproj, unsloth/Qwen3-VL-8B-Instruct-GGUF), loaded through a vendored ComfyUI-GGUF-Reboot custom node (the standard ComfyUI-GGUF does not support merging a separate mmproj file for the qwen3vl architecture β€” only qwen2vl, used by gguf-qwen/longcat-gguf). Replaces the .safetensors CLIP options evaluated (int8_convrot, fp8_scaled, w4a8) as the default: same disk footprint as the lightest of those (w4a8, ~5.9GB) but without a chromatic-drift issue found on subjects with a strong color convention (w4a8 occasionally converged on the wrong hue where the other options didn't). The int8_convrot/w4a8 files remain valid alternatives β€” see Pipeline Configuration.

clip_mmproj/clip_mmproj_hf_name (config fields) generalized from qwen21-viggle to gguf-qwen/longcat-gguf too, replacing the old mmproj_gguf key (which was never actually read by any code β€” dead documentation only). This also closed a real gap: longcat-gguf never had any mechanism to auto-download its own mmproj file β€” it only worked because the file was already present from gguf-qwen sharing the same folder. A from-scratch longcat-gguf-only installation would have loaded its CLIP without a working vision tower.

Known limitation, extensively investigated: on some frames, a human body part near a visually similar background (e.g. a hand close to foliage) can be rendered with the wrong color (blended into the background) instead of a natural skin tone β€” confirmed across the entire Qwen-Image-2.1/Viggle-Turbo family, including Viggle's own official demo app and longcat-gguf, and independent of which CLIP quantization/variant is used (int8, fp8, w4a8, and several GGUF text-encoder builds were tested). This appears to be a genuine limitation of the underlying models for this kind of ambiguous content, not a bug in this integration. If a frame is affected, nunchaku-qwen/gguf-qwen are unaffected by the same issue and can be used as a fallback.

2026-09-28 β€” qwen21-viggle: higher working resolution for Fast Pipeline

Paired inference (Fast Pipeline) for qwen21-viggle now uses a working resolution of 1280 instead of the usual 1024 (single-image and every other backend are unaffected). This fixes a color artifact found on frames with fine detail near the merge boundary β€” a hand, held up close to the camera, could be rendered in tones nearly indistinguishable from the background foliage instead of a natural skin tone. The cause was the reduced working resolution from paired inference combined with Viggle-Turbo's own limits on fine detail; raising it to 1280 resolves the artifact in every case tested, at a real but modest speed cost (~6 sec/frame instead of 4 β€” see the recommendation note near the top of this README). 1536 was tested too and fixes the same artifact slightly more completely, at roughly double the extra cost; 1280 was chosen as the better trade-off after validation on thousands of real frames.

2026-09-27 β€” vscmnet2 1.1.0 (proximity bias now a per-call parameter)

Updated to vscmnet2 1.1.0. The proximity-weighted memory matching feature added in 1.0.9 (see below) is no longer installation-wide only: vs_cmnet2 now accepts enable_proximity_bias/proximity_bias_alpha directly, taking precedence over vsslib/models.json when passed explicitly for a single call.

clip = vs_cmnet2(
    clip,
    clip_ref=ref_clip,
    method=0,
    enable_proximity_bias=True,
    proximity_bias_alpha=0.5,
)

vs_cmnet2_recolor/vs_cmnet2dit are unchanged β€” they still only pick up the installation-wide default from vsslib/models.json. Exposed in the GUI's Encode/Merge tab only (the tab backed by vs_cmnet2), in a new CMNET2 Backbone frame grouping Backbone, Proximity Bias and Alpha together: the latter two are automatically disabled when Backbone = dinov2 (DINOv3-only feature) and re-enabled on switching back to dinov3. Unchecked always forces enable_proximity_bias=False for that run (an explicit override, not "leave it to models.json"). Not added to Fix Video (backed by vs_cmnet2_recolor, which doesn't accept these parameters).

Existing installations, action needed: the shipped DINOv3 checkpoint was renamed DINOv3FeatureV6_LocalAtten_p372402.pth β†’ DINOv3FeatureV6_LocalAtten_p374099.pth (default proximity_bias_alpha also changed 0.7 β†’ 0.5). Re-download the checkpoint under the new name β€” see DINOv3 backbone weights. If the old file is left in place, vscmnet2 fails fast at init with a clear error listing the files actually present in the weights directory.

2026-09-26 β€” qwen21-viggle Backend

A fourth model backend has been added: qwen21-viggle (Qwen-Image-2.1 native ComfyUI weights + Viggle-Turbo LoRA). Uses native int8 ConvRot quantized weights (not GGUF) β€” GGUF quantization was evaluated and works, but is ~2Γ— slower for this model, so it was not adopted. Runs on 14GB+ VRAM GPUs, ~8-11 sec/frame (the time scales little with the number of steps β€” a fixed text-encoding/VAE-decode cost dominates over sampling).

To use this backend on an existing installation: git pull then run quick_update.cmd β€” this pulls in the required comfy-kitchen==0.2.35/ comfy-aimdo==0.5.5 versions (see Quick Update) alongside the rest of the qwen21-viggle code.

Four model files are required (auto-downloaded on first run):

File Size Source
unet/qwen_image_2.1_int8_convrot.safetensors ~6.8 GB Comfy-Org/Qwen-Image-2.1
clip/qwen3vl_8b_int8_convrot.safetensors ~8.7 GB Comfy-Org/Qwen-Image-2.1
vae/qwen_image_2.1_vae_bf16.safetensors ~0.7 GB Comfy-Org/Qwen-Image-2.1
loras/Qwen-Image-2.1-viggle-turbo-v0.2.1-6step-lora-r128.safetensors ~0.7 GB Viggle/Qwen-Image-2.1-viggle-turbo

Launch via run_server_qwen21.cmd (no arguments needed β€” a single config is available). Config: config/qwen21_viggle.json.

Steps: the Viggle-Turbo LoRA is distilled for 6-step inference (its native step count). Four precomputed sigma schedules are available β€” 2, 4, 6 (native), 8 β€” selected via the usual steps parameter; any other value falls back to the 6-step schedule with a warning in the logs. 2/4/8 are experimental (the LoRA author only documents 5-7 step schedules).

enhance_prompt (optional, default False, all colorization RPC methods): rewrites the prompt using Qwen3-VL as an image-aware "observer" before colorizing β€” useful when the plain prompt leaves color ambiguous on recognizable subjects (e.g. a costume with a well-known color) and the model resolves the ambiguity inconsistently. Adds ~15-20s per frame (a second Qwen3-VL generation pass). A direct, explicitly anti-hedging prompt often achieves the same result without the extra cost β€” see the suggested prompt below before reaching for enhance_prompt by default.

Prerequisite: 12 GB+ VRAM and 32GB+ RAM. Uses comfy_bridge's native ComfyUI runtime (no external ComfyUI checkout needed) β€” the same as gguf-qwen/longcat-gguf, extended with native Qwen-Image-2.1/Qwen3-VL support.

The GUI Tab 2 (Colorization) supports this backend: selecting qwen21-viggle from Model Name auto-disables the (unused) Precision combo and reads model paths from config/qwen21_viggle.json. Run Server manages the equivalent of run_server_qwen21.cmd directly (see What's New, 2026-09-30) β€” or launches that same .cmd file in its own console when the External console checkbox is ticked. An Enhance Prompt checkbox is available in Tab 2 and Tab 4 (Fix Image).

2026-09-25 β€” x264 encoder option (GUI)

Added x264 as a third software encoder choice in the GUI's Encode/Merge tab, alongside the existing x265 (software, 10-bit) and Nvenc (GPU hardware, H.265). x264 is an 8-bit H.264 CPU encoder β€” useful when H.265 decoding/compatibility is a constraint. The x264.exe binary is located next to x265.exe, same convention already used for NVEncC64.exe β€” no extra path field needed in the GUI. Now included in the Release 1.0.0 tools.zip alongside x265.exe/mkvmerge.exe. See GUI README: install external tools.

2026-09-24 β€” vscmnet2 1.0.9 (proximity-weighted memory matching)

Updated to vscmnet2 1.0.9, which add proximity-weighted memory matching. By default, permanent-memory candidates are ranked purely by content similarity, with no notion of when in the video a reference frame was captured relative to the frame being colorized β€” with a wide max_memory_frames window holding several visually similar but differently-colored references, this can wash the result toward gray. Unlike backbone, this is not exposed as a filter parameter on vs_cmnet2: it is configured once for the whole installation via the enable_proximity_bias/proximity_bias_alpha keys in vsslib/models.json (see above). Off by default. To permanently enable it (useful for permanent memory window size > 50) it is necessary to set enable_proximity_bias=true in the configuration file stored in: vsslib/models.json as shown in the example below:

{
  "cmnet2": {
    "dinov3": {
      "checkpoint": "DINOv3FeatureV6_LocalAtten_p374099.pth",
      "weights_dir": "dinov3-vitb16",
      "enable_proximity_bias": true,
      "proximity_bias_alpha": 0.5
    },
    "dinov2": {
      "checkpoint": "DINOv2FeatureV6_LocalAtten_s2_154000.pth"
    }
  }
}

2026-09-17 β€” vscmnet2 1.0.8 (DINOv3 Backbone)

Updated to vscmnet2 1.0.8, which switches CMNET2 to a DINOv3 ViT-B/16 key-encoder backbone by default (previously DINOv2 ViT-S/14), improving colorization quality. The legacy DINOv2 backbone remains available via a backbone parameter.

A new Backbone combo (dinov3 / dinov2) has been added to the GUI in both tabs that drive CMNET2 through VapourSynth:

  • Encode/Merge (Tab 3) β€” GUI/scripts/encode_cmnet2.vpy
  • Fix Video (Tab 6) β€” GUI/scripts/encode_cmnet2_recolor.vpy

Both scripts now pass the selected backbone to vs_cmnet2() / vs_cmnet2_recolor() via a Backbone VapourSynth argument, alongside the existing RenderSpeed and MemoryFrames parameters.

Prerequisite: the DINOv3 backbone requires new weight files β€” see GUI README: DINOv3 backbone weights for download links and install steps. The dinov2 option remains available for installations that only have the legacy DINOv2 weights.

2026-07-10 β€” LongCat GGUF Backend

A third model backend has been added: longcat-gguf (LongCat-Image-Edit-Turbo GGUF). Uses quantized UNet (Q4_K_M) and CLIP GGUF files β€” runs on 12 GB VRAM GPUs. It achieves excellent colorization quality (~12 s/frame via the RPC server) β€” richer colors, more natural skin tones, and better detail preservation than the gguf-qwen model β€” and is accessible from all existing tabs and RPC endpoints.

Three new model files are required (auto-downloaded on first run):

File Size Source
unet/LongCat-Image-Edit-Turbo-Q4_K_M.gguf ~5.4 GB vantagewithai/LongCat-Image-Edit-Turbo-GGUF
clip/Qwen2.5-VL-7B-Instruct-Q4_K_M.gguf ~4.6 GB unsloth/Qwen2.5-VL-7B-Instruct-GGUF
vae/lct_vae.safetensors ~160 MB meituan-longcat/LongCat-Image-Edit-Turbo

Launch via run_server_longcat.cmd (Q4_K_M) or start_server.cmd longcat (Q4), longcat-q3, longcat-q5, longcat-q6, longcat-q8. Pre-made configs for all quantizations are in the config/ folder.

Prerequisite: 12 GB+ VRAM and 32GB+ RAM. Uses GGUF quantized models. The pipeline uses the comfy_bridge runtime (no external ComfyUI checkout needed). Custom nodes CFGNorm, FluxKontextMultiReferenceLatentMethod, and TextEncodeQwenImageEditPlus are included in comfy_bridge/comfy_extras/.

The GUI Tab 2 (Colorization) now includes a Run Server button that manages the server for the selected Model Name + Precision directly β€” see What's New, 2026-09-30 for how it's started/stopped/logged, and the External console checkbox for opening a plain terminal window instead. This replaces the need to manually find and run the right .cmd file.

2026-07-01 β€” Batch Processing for Fix Image & Fix Colors (GUI)

The Fix Image (Tab 4) and Fix Colors (Tab 5) tabs now support batch processing of multiple images:

Fix Image (Tab 4):

  • New Enable batch processing checkbox β€” toggles between single-image and batch mode
  • The image field is now a ComboBox showing all loaded images (drag & drop / Browse appends)
  • Colorize processes all images sequentially against the DiT RPC server
  • Overwrite overwrites all originals; Save As proposes a *_colorized wildcard mask
  • Outputs are kept in memory until explicitly saved; errors on single images are skipped
  • Swap Output is automatically disabled in batch mode

Fix Colors (Tab 5):

  • Same batch logic on the Target Image field β€” multiple targets against one color reference
  • ComboBox, Clear, sequential colorization via local CMNET2, wildcard save
  • Copy β†’ Fix Image disabled in batch mode

Both tabs share the same interaction pattern for batch mode β€” ComboBox list, sequential processing with progress counter, wildcard * save mask, memory-only outputs until explicit save β€” while each tab uses its own backend (DiT RPC for Fix Image, local CMNET2 for Fix Colors).

2026-06-20 β€” Fix Colors Tab (GUI)

A standalone Fix Colors tab (Tab 5) has been added to the desktop GUI (GUI/CMNET2_colorize_client_GUI.py).

GUI Tab #5

It colorizes a B&W or colorized target image using a color reference image via the local CMNET2 model (exemplar-based color propagation) β€” no RPC server required:

  1. Load a color reference image (drag & drop or Browse)
  2. Load a B&W target image (drag & drop or Browse)
  3. Colorize β€” runs vscmnet2.pil_cmnet2_colorize() in a background thread. It allows to propagate the reference colors to target image.

Key features:

  • Three preview panels: reference, target, and output side‑by‑side
  • Copy β†’ Fix Image: sends the output directly to Tab 4 (Fix Image) for a two‑stage pipeline (CMNET2 β†’ DiT RPC)
  • Save / Overwrite: save the colorized result as PNG/JPG or overwrite the original target file
  • Full-resolution preservation: images are always kept at original resolution in memory; resizing only applies to previews
  • Delayed import: vscmnet2 is imported only when Colorize is clicked (does not block GUI startup)
  • Backbone selection (dinov3 / dinov2, since 2026-09-28): passed to vscmnet2.pil_cmnet2_colorize() β€” same combo already available in Encode/Merge (Tab 3) and Fix Video (Tab 6)

Prerequisite: vscmnet2 must be installed with model weights and checkpoints present (see GUI README). No RPC connection needed.

The tab order has been updated: 1. Extraction β†’ 2. Colorization β†’ 3. Encode/Merge β†’ 4. Fix Image β†’ 5. Fix Colors β†’ 6. Fix Video.

2026-06-17 β€” Fix Video Tab (GUI)

A standalone Fix Video tab (Tab 6) has been added to the desktop GUI (GUI/CMNET2_colorize_client_GUI.py). It runs a VapourSynth + NVEnc pipeline to recolor a video using two reference images:

GUI Tab #6

  1. Select a video and an encode VPY script
  2. Load two reference images (First / Last) via drag-and-drop or Browse
  3. Recolor β€” runs the VapourSynth β†’ NVEnc pipeline and produces _dt-recolor.mkv

Key features:

  • NVEnc-only: uses GPU hardware encoding (NVEncC64.exe required)
  • RefStart / RefEnd: reference images passed to the VapourSynth script as parameters
  • RefDir auto-detection: set to the folder of the first reference image
  • Configurable: FPS, VBR Quality, Memory Frames, Render Speed, Backbone
  • MKV output: .h265 intermediate automatically muxed to .mkv and deleted
  • Pre-flight check: verifies NVEncC64.exe exists before starting

The Fix Video tab is independent of the batch pipeline and does not require the RPC server. Only the frames between RefStart / RefEnd will be recolored.

Prerequisite: NVEncC must be installed in tools\NVEncC\ (see GUI README).

2026-06-12 β€” Fix Image Tab (GUI)

A standalone Fix Image tab (Tab 4) has been added to the desktop GUI (GUI/CMNET2_colorize_client_GUI.py). It allows single-image colorization with seed control, drag-and-drop file loading, and preview:

GUI Tab #4

  1. Load a B&W image via drag-and-drop (GUI/load_image_DtD_GUI.py) or the Browse button
  2. Colorize with fixed seed (42) or random seed for variation
  3. Save the colorized result as PNG / JPG

The Fix Image tab is independent of the batch video pipeline and does not require VapourSynth.

If you already created the .venv with a previous version install the package tkinterDnD2 to add drag-and-drop support to tkinter

# Windows
.venv\Scripts\activate
(.venv) pip install tkinterDnD2

2026-06-09 β€” Improved GGUF

Changed the GGUF configuration files. The pipeline Qwen-Image-Edit-2511 + Qwen-Image-Edit-2511-Lightning-4steps has substituted by the pipeline with Qwen-Image-Edit-2509 + Qwen-Image-Edit-2511-Lightning-4steps. This change has removed the artifacts problem which affected the colored images with the GGUF models and improved the overall quality of the colored images. It should be noted that, despite these improvements, the Nunchaku model remains the best and is the one recommended for production use (for systems with limited hardware resources, it is recommended to use the GGUFs of LongCat-Image-Edit-Turbo added on July 10th, 2027).

2026-06-07 β€” Desktop GUI for Batch Video Processing

A FreeSimpleGUI desktop client (GUI/CMNET2_colorize_client_GUI.py) has been added to the project. It orchestrates the full video colorization pipeline from a single graphical interface:

  1. Extract reference frames via VapourSynth + scene-change detection
  2. Colorize frames via the HAVC DiT Server (standard or paired inference)
  3. Encode the result as H.265 (x265 or NVEnc) or H.264 (x264)
  4. Merge the AI output with an existing color clip (optional, luminance-guided chroma blend)

GUI Tab #2

See GUI/README_GUI.md for installation, setup, and usage instructions.

Prerequisite: the HAVC DiT Server must be running before the GUI can colorize frames.


✨ Features

  • πŸ“¦ 4 backends, one API : nunchaku-qwen (FP4/INT4, 4 sec/frame) for speed, gguf-qwen and longcat-gguf (Q3, …, Q8, 12 sec/frame) for lower VRAM, qwen21-viggle (int8 ConvRot UNet, ~8-11 sec/frame) with optional Qwen3-VL prompt rewriting
  • 🎨 Batch colorization : process entire directories of B&W images via filesystem paths
  • πŸ–ΌοΈ Paired inference : colorize two images in a single forward pass (faster, temporally consistent)
  • πŸ“‘ In-memory RPC : pass raw PNG frames over XML-RPC without touching the filesystem (ideal for video pipelines)
  • ⚑ 4-step lightning model : SVDQuant FP4 quantized transformer for maximum throughput
  • πŸ”’ Thread-safe : pipeline loading and stop control are protected by locks; every RPC call runs in its own thread
  • βš™οΈ Startup preload : optional --load-pipeline flag loads the model at boot from a JSON config file
  • πŸš€ Shared memory transport : zero-copy image transfer for same-host deployments (~23% faster than standard RPC)

πŸ“‹ Prerequisites

Choose the backend that matches your hardware:

nunchaku-qwen : 4 sec/frame (FP4/INT4)

Requirement Details
GPU NVIDIA RTX 30/40/50 (16 GB+ VRAM)
RAM 64 GB+
CUDA 13.0 or newer
CUDA Toolkit Must match the PyTorch build

RTX 30/40-Series (Ampere / Ada): use "model_precision": "int4". FP4 requires Blackwell (RTX 50). Requires Nunchaku 1.2.1 and diffusers==0.37.0.dev0 (wheel included in packages/).

gguf-qwen : 14 sec/frame (Q3, Q4, Q5, Q6, Q8)

Requirement Details
GPU NVIDIA RTX 30/40/50 (12 GB+ VRAM)
RAM 32 GB+
CUDA 13.0+ (or CPU-only: slower, zero VRAM)

Q3_K_S fits in 12 GB VRAM. Q4_K_S (default) balances quality and VRAM. Q5_K_M / Q6_K improve fidelity at higher VRAM cost. Q8_0 is near-lossless. Uses ComfyUI-native code : no ComfyUI GUI installation needed. Pre-made configs for all quantizations are in the config/ folder.

longcat-gguf : 12 sec/frame (Q3–Q8) β€” Best Quality

Requirement Details
GPU NVIDIA RTX 30/40/50 (12 GB+ VRAM)
RAM 32 GB+
CUDA 13.0+

LongCat-Image-Edit-Turbo delivers noticeably better colorization than the gguf-qwen model β€” richer colors, more natural skin tones, and better detail preservation β€” at the same ~12 s/frame speed.

Uses GGUF quantized UNet (Q3_K_M to Q8_0) + CLIP Q4_K_M. The UNet is distributed in five quantization levels to fit different VRAM budgets. All files are auto-downloaded on first run.

See config/longcat_gguf_q*.json β€” the general rule: lower quant = less VRAM. Launch with run_server_longcat.cmd (Q4_K_M) or start_server.cmd longcat|longcat-q3|....

qwen21-viggle : ~8-11 sec/frame (Qwen-Image-2.1 int8 ConvRot)

Requirement Details
GPU NVIDIA RTX 30/40/50 (14 GB+ VRAM)
RAM 32 GB+
CUDA 13.0+

Native ComfyUI int8 ConvRot weights for the UNet (not GGUF β€” a GGUF UNet was evaluated but is ~2Γ— slower for this model). The CLIP/text encoder, unlike the UNet, uses GGUF+mmproj by default (Qwen3-VL-8B- Instruct-UD, see What's New) β€” a .safetensors CLIP (int8_convrot/w4a8) remains a valid, simpler alternative, see Pipeline Configuration. Requires comfy-kitchen==0.2.35 and comfy-aimdo==0.5.5 exactly (pinned, not a minimum β€” both are compiled packages and an untested newer build is not assumed safe). A fresh install.cmd run sets these; an existing .venv needs an explicit upgrade, see Quick Update. All files are auto-downloaded on first run β€” see What's New. Launch with run_server_qwen21.cmd.

All backends

Requirement Details
OS Windows 10/11 or Linux
Python 3.12

πŸ› οΈ Installing Git and Python

Before setting up the project environment, make sure both Git and Python 3.12 are installed on your system.

Git

Windows: download and install Git for Windows. Accept the default options : in particular keep core.autocrlf=true (the default), which ensures correct line endings for .cmd files.

Linux:

sudo apt install git        # Debian / Ubuntu
sudo dnf install git        # Fedora / RHEL

Verify: git --version


Python 3.12

Windows: download the installer from python.org/downloads. During installation, check "Add Python to PATH" : without this, python will not be recognized in the terminal.

Linux:

sudo apt install python3.12 python3.12-venv   # Debian / Ubuntu
sudo dnf install python3.12                   # Fedora / RHEL

Verify: python --version (Windows) or python3.12 --version (Linux)


βš™οΈ Environment Setup

1 : Clone the repository and create a virtual environment

Clone the repository with git : this ensures correct line endings for all files (.gitattributes is applied automatically at checkout):

git clone https://github.com/dan64/HAVCServerDiT.git
cd HAVCServerDiT

Then create and activate the virtual environment inside the project directory:

python -m venv .venv
# Windows
.venv\Scripts\activate
# Linux / macOS
source .venv/bin/activate

Windows quick-start: once the venv is active you can run install.cmd to execute steps 2–6 automatically instead of running them one by one.


2 : Install PyTorch 2.10.0 + CUDA 13.0

Use the stable build for all GPU generations (RTX 30 / 40 / 50):

pip install torch==2.10.0+cu130 torchvision==0.25.0+cu130 torchaudio==2.10.0+cu130 \
    --index-url https://download.pytorch.org/whl/cu130

Verify the installation:

python -c "import torch; print(torch.__version__); print(torch.cuda.is_available())"
# Expected: 2.10.0+cu130, True

3 : Install Nunchaku

⚠️ Do NOT use pip install nunchaku : that installs an unrelated package from PyPI with the same name that will fail with ModuleNotFoundError: No module named 'nunchaku.models'.

Install the correct MIT Han Lab build directly from the GitHub release:

# Windows / Python 3.12 / CUDA 13.0 / PyTorch 2.10
pip install https://github.com/nunchaku-ai/nunchaku/releases/download/v1.2.1/nunchaku-1.2.1+cu13.0torch2.10-cp312-cp312-win_amd64.whl

For other platforms or Python versions, browse the full list of available wheels on the Nunchaku releases page and replace the filename accordingly.

Nunchaku pulls torch>=2.0 as a dependency (via accelerate) and may upgrade PyTorch to a newer version. After installing Nunchaku, re-pin PyTorch:

pip install torch==2.10.0+cu130 torchvision==0.25.0+cu130 torchaudio==2.10.0+cu130 --index-url https://download.pytorch.org/whl/cu130 --force-reinstall

Verify the correct package is installed :

pip show nunchaku
# Version: 1.2.1+cu13.0torch2.10
pip show torch
# Version: 2.10.0+cu130

4 : Patch Nunchaku

Nunchaku 1.2.1 contains a bug in its transformer forward pass: txt_seq_lens is always None at the point where it is passed to pos_embed, causing a ValueError with diffusers >= 0.37.0.dev0. The included patch_nunchaku.py fixes this by deriving max_txt_seq_len directly from encoder_hidden_states:

python patch_nunchaku.py

On Windows you can also double-click patch_nunchaku.cmd or run it from a terminal:

patch_nunchaku.cmd            # apply the patch
patch_nunchaku.cmd --check    # check status without modifying files
patch_nunchaku.cmd --revert   # revert to original (.bak backup)

You can verify the patch status at any time:

python patch_nunchaku.py --check

And revert to the original if needed (a .bak backup is created automatically):

python patch_nunchaku.py --revert

5 : Install Diffusers

⚠️ Do NOT install diffusers from GitHub (pip install git+https://...). Nunchaku 1.2.1 requires exactly 0.37.0.dev0. Later dev builds (β‰₯ 0.39.0) changed the QwenEmbedRope API in a way that is incompatible even after the nunchaku patch.

A tested compatible wheel is included in the packages/ folder. Install it directly:

pip install packages\diffusers-0.37.0.dev0-py3-none-any.whl

Verify:

python -c "import diffusers; print(diffusers.__version__)"
# Expected: 0.37.0.dev0

6 : Install remaining dependencies

Pin the versions to match the tested working environment:

pip install \
    transformers==4.57.6 \
    accelerate==1.12.0 \
    "huggingface_hub>=0.26.0" \
    "Pillow>=10.0.0" \
    scipy \
    av \
    torchsde \
    gguf \
    comfy-aimdo==0.5.5 \
    comfy-kitchen==0.2.35

Nunchaku users: diffusers was already installed in step 5 as the compatible 0.37.0.dev0 wheel. Do NOT upgrade it : nunchaku 1.2.1 requires exactly that version.

safetensors is pulled automatically by diffusers.

scipy, av, and torchsde are required by the diffusers pipeline. gguf, comfy-aimdo, and comfy-kitchen are required by the GGUF backends (gguf-qwen/longcat-gguf) and by qwen21-viggle. The pinned versions here (0.5.5/0.2.35) are required specifically for qwen21-viggle β€” exact pins, not just a minimum, to avoid drifting to an untested newer build of these compiled packages.

πŸ“‚ Project Structure

dit-colorize-rpc/
β”œβ”€β”€ dit_rpc_server.py            # XML-RPC server (entry point)
β”œβ”€β”€ dit_colorize_main.py         # Colorization pipeline and image utilities
β”œβ”€β”€ dit_client_example.py        # Example RPC client : single frame
β”œβ”€β”€ dit_client_pair_example.py   # Example RPC client : paired inference
β”œβ”€β”€ patch_nunchaku.py            # Compatibility patch for nunchaku 1.2.1
β”œβ”€β”€ config/                      # Pipeline configs (nunchaku FP4/INT4, gguf/longcat Q3–Q8, qwen21-viggle)
β”œβ”€β”€ comfy_bridge/                # Self-contained ComfyUI runtime (gguf-qwen/longcat-gguf/qwen21-viggle)
β”œβ”€β”€ install.cmd                  # Windows automated installer
β”œβ”€β”€ start_server.cmd             # Windows launcher : server (nunchaku/gguf/longcat)
β”œβ”€β”€ run_server_longcat.cmd       # Windows launcher : LongCat server
β”œβ”€β”€ run_server_qwen21.cmd        # Windows launcher : qwen21-viggle server
β”œβ”€β”€ run_client_example.cmd       # Windows launcher : single frame example
β”œβ”€β”€ run_client_pair_example.cmd  # Windows launcher : paired inference example
β”œβ”€β”€ patch_nunchaku.cmd           # Windows launcher : nunchaku patch
β”œβ”€β”€ assets/
β”‚   β”œβ”€β”€ santa_bw.png             # Sample B&W image (single frame test)
β”‚   β”œβ”€β”€ sample1_bw.jpg           # Sample B&W image 1 (paired inference test)
β”‚   └── sample2_bw.jpg           # Sample B&W image 2 (paired inference test)
β”œβ”€β”€ packages/
β”‚   └── diffusers-0.37.0.dev0-py3-none-any.whl  # Tested compatible diffusers build
└── README.md

πŸ”§ Pipeline Configuration

Ready-to-use config files for both backends are in the config/ folder. Pick the one that matches your hardware and pass it to --pipeline-config.

Nunchaku Backend : config/qwen_nunchaku_fp4.json & qwen_nunchaku_int4.json

config/qwen_nunchaku_fp4.json : RTX 50-Series (Blackwell)

{
    "model_name":            "nunchaku-qwen",
    "model_precision":       "fp4",
    "model_rank":            "32",
    "model_inference_steps": "4",
    "cache_dir":             "",
    "full_model_path":       ""
}

config/qwen_nunchaku_int4.json : RTX 30 / 40-Series (Ampere / Ada Lovelace)

{
    "model_name":            "nunchaku-qwen",
    "model_precision":       "int4",
    "model_rank":            "32",
    "model_inference_steps": "4",
    "cache_dir":             "",
    "full_model_path":       ""
}

⚠️ model_precision: use "fp4" only on RTX 50-Series (Blackwell). On RTX 30 / 40-Series use "int4" : FP4 kernels require sm_120 and will fail on older architectures.

GGUF Backend : config/qwen_gguf_q3.json … qwen_gguf_q8.json / config/longcat_gguf_q3.json … longcat_gguf_q8.json

Five quantization levels are available. All share the same structure with model_name: "gguf-qwen" and a quant field that selects the quantization:

Config file quant UNet CLIP
qwen_gguf_q3.json "q3" …Q3_K_S.gguf …Q3_K_S.gguf
qwen_gguf_q4.json "q4" …Q4_K_S.gguf …Q4_K_S.gguf
qwen_gguf_q5.json "q5" …Q5_K_M.gguf …Q5_K_M.gguf
qwen_gguf_q6.json "q6" …Q6_K.gguf …Q6_K.gguf
qwen_gguf_q8.json "q8" …Q8_0.gguf …Q8_0.gguf

Q4 is the recommended default : good quality/VRAM balance, but even Q3 is capable of delivering frames with acceptable colors. All quants share the same VAE, mmproj, and LoRA files (auto-downloaded from HuggingFace).

⚠️ The GGUF backend is experimental. In some cases the frames colors may be faded or little colored. For production use, prefer nunchaku-qwen (FP4/INT4) or qwen21-viggle, which are not affected by such problems β€” see the recommendation note at the top of this README.

Config example (config/qwen_gguf_q4.json):

{
    "model_name":       "gguf-qwen",
    "quant":            "q4",
    "unet_gguf":        "models/unet/qwen-image-edit-2511-Q4_K_S.gguf",
    "clip_gguf":        "models/clip/Qwen2.5-VL-7B-Instruct-Q4_K_S.gguf",
    "clip_mmproj":         "models/clip/Qwen2.5-VL-7B-Instruct-mmproj-BF16.gguf",
    "clip_mmproj_hf_name": "mmproj-BF16.gguf",
    "vae_name":         "qwen_image_vae.safetensors",
    "lora_path":        "models/loras/Qwen-Image-Edit-2511-Lightning-4steps-V1.0-bf16.safetensors",
    "steps":            4,
    "hf_unet":          "unsloth/Qwen-Image-Edit-2511-GGUF",
    "hf_clip":          "unsloth/Qwen2.5-VL-7B-Instruct-GGUF",
    "hf_vae":           "Comfy-Org/Qwen-Image_ComfyUI",
    "hf_lora":          "lightx2v/Qwen-Image-Edit-2511-Lightning"
}

clip_mmproj_hf_name exists because the mmproj file's name on HuggingFace (a generic mmproj-BF16.gguf, shared across many unrelated repos) rarely matches the locally-prefixed name you actually want on disk β€” it tells the downloader what to fetch, clip_mmproj is where it ends up and what the loader looks for locally. Omit it and the downloader falls back to using clip_mmproj's own filename as the remote name too, which only works if they happen to match.

LoRA (Lightning 4-step)

The LoRA file Qwen-Image-Edit-2511-Lightning-4steps-V1.0-bf16.safetensors enables 4-step inference (down from 20-50 steps without LoRA). It is a ComfyUI-format LoRA that gets merged directly into the transformer at load time.

  • With LoRA: call colorize_image(..., steps=4) : fast, same quality
  • Without LoRA: set full_model_path to "" and use steps=20 or higher

The LoRA is merged statically (not applied as an adapter), so there is no runtime overhead.

qwen21-viggle Backend : config/qwen21_viggle.json

A single config file β€” this backend has no quantization variants for the UNet (int8 ConvRot only). The CLIP, unlike the UNet, can be either a .safetensors file or a GGUF+mmproj pair β€” the default uses GGUF+mmproj (see What's New):

{
    "model_name":          "qwen21-viggle",
    "unet_name":           "models/unet/qwen_image_2.1_int8_convrot.safetensors",
    "clip_name":           "models/clip/Qwen3-VL-8B-Instruct-UD-Q4_K_XL.gguf",
    "clip_mmproj":         "models/clip/Qwen3-VL-8B-Instruct-mmproj-BF16.gguf",
    "clip_mmproj_hf_name": "mmproj-BF16.gguf",
    "vae_name":            "qwen_image_2.1_vae_bf16.safetensors",
    "lora_path":           "models/loras/Qwen-Image-2.1-viggle-turbo-v0.2.1-6step-lora-r128.safetensors",
    "steps":               6,
    "hf_unet":             "Comfy-Org/Qwen-Image-2.1",
    "hf_clip":             "unsloth/Qwen3-VL-8B-Instruct-GGUF",
    "hf_vae":              "Comfy-Org/Qwen-Image-2.1",
    "hf_lora":             "Viggle/Qwen-Image-2.1-viggle-turbo"
}

Note the different key names from the GGUF-backend format above: unet_name/ clip_name (not unet_gguf/clip_gguf) β€” but unlike the GGUF backend, clip_name here can point to either a .safetensors file or a .gguf file (the loader picks the right code path from the extension). clip_mmproj/clip_mmproj_hf_name only apply when clip_name is a .gguf file β€” omit both to use a .safetensors CLIP instead:

    "clip_name":   "models/text_encoders/qwen3vl_8b_int8_convrot.safetensors",

steps: 6 here only documents the LoRA's native step count for anyone reading the file; the actual number of steps used at inference time is the steps argument passed per-call to the colorization RPC methods (see Suggested Inference Steps and RPC API Reference), same as every other backend.

Key reference

Key Required Description
model_name βœ… "nunchaku-qwen", "gguf-qwen", "longcat-gguf", or "qwen21-viggle"
quant GGUF only: quantization level ("q3", "q4", "q5", "q6", "q8"). Default: "q4"
model_precision βœ… Nunchaku: "fp4" (RTX 50) or "int4" (RTX 30/40). GGUF/qwen21-viggle: not used
unet_gguf / clip_gguf βœ… GGUF only: local paths to the GGUF model files
unet_name / clip_name βœ… qwen21-viggle only: local paths to the model files β€” unet_name is always .safetensors, clip_name can be .safetensors or .gguf
clip_mmproj / clip_mmproj_hf_name GGUF/qwen21-viggle-with-GGUF-CLIP: local path to the mmproj (vision tower) file / its filename on HuggingFace if different from the local one. Required for a GGUF CLIP to see images at all β€” without it the vision tower silently isn't loaded
model_rank Nunchaku: SVD rank ("32"). GGUF/qwen21-viggle: not used
model_inference_steps Nunchaku: diffusion steps ("4"). GGUF/qwen21-viggle: not used at load time
cache_dir HuggingFace cache directory. Leave empty to use the default ~/.cache/huggingface
full_model_path Nunchaku: local path to the transformer checkpoint. GGUF: not used
lora_path GGUF/qwen21-viggle: path to the LoRA (.safetensors). Omit for GGUF to skip LoRA merging
steps GGUF: inference steps (4 with LoRA, 20 without). qwen21-viggle: documents the native step count only, not consumed at load time
vae_name GGUF/qwen21-viggle only: VAE filename
hf_* GGUF/qwen21-viggle only: HuggingFace repo names for auto-download

πŸš€ Usage

Start the server (no preload : pipeline loaded later via RPC)

python dit_rpc_server.py

Start the server with pipeline preloaded at boot

# RTX 50-Series
python dit_rpc_server.py --load-pipeline --pipeline-config config/qwen_nunchaku_fp4.json

# RTX 30 / 40-Series
python dit_rpc_server.py --load-pipeline --pipeline-config config/qwen_nunchaku_int4.json

# GGUF (any quantization)
python dit_rpc_server.py --load-pipeline --pipeline-config config/qwen_gguf_q3.json

On Windows you can also use the provided start_server.cmd (see Windows launch script).

Full list of CLI arguments

usage: dit_rpc_server.py [-h] [--host HOST] [--port PORT]
                         [--logfile LOGFILE] [--module-dir MODULE_DIR]
                         [--load-pipeline] [--pipeline-config CONFIG.json]

options:
  --host HOST                  Address to listen on (default: 127.0.0.1)
  --port PORT                  TCP port (default: 8765)
  --logfile LOGFILE            Optional path for a log file
  --module-dir MODULE_DIR      Directory containing dit_colorize_main.py
                               (default: same directory as this script)
  --load-pipeline              Load the colorization pipeline at startup
  --pipeline-config CONFIG.json
                               Path to the JSON pipeline config file
                               (required when --load-pipeline is set)

πŸ“‘ RPC API Reference

Connect from any Python client using xmlrpc.client:

import xmlrpc.client
proxy = xmlrpc.client.ServerProxy("http://127.0.0.1:8765/", use_builtin_types=True)

All methods return a dict with at least {"ok": bool, "msg": str}.

Health

Method Returns Description
ping() "pong" Connectivity check

Pipeline management

Method Returns Description
load_pipeline(model_name, model_precision, model_rank, model_inference_steps, cache_dir="", full_model_path="", vae_name="", hf_unet="", hf_clip="", hf_vae="", hf_lora="") {"ok", "msg"} Load the model into VRAM. The vae_name/hf_* arguments are only meaningful for gguf-qwen/longcat-gguf/qwen21-viggle β€” omit for nunchaku-qwen
is_pipeline_loaded() bool True if the pipeline is ready
unload_pipeline() {"ok", "msg"} Release VRAM

Stop control

Method Returns Description
request_stop() bool Ask the server to refuse new colorization calls
clear_stop() bool Reset the stop flag before a new batch
is_stop_requested() bool Check the current stop flag

Colorization : filesystem-based

Method Returns Description
colorize_image(in_path, out_path, prompt, img_size=0, steps=2, enhance_prompt=False) {"ok", "elapsed", "skipped", "msg"} Single image, paths on the server filesystem
colorize_image_pair(img1_path, img2_path, out_dir, prompt, gap_px=8, steps=2, enhance_prompt=False) {"ok", "elapsed", "msg"} Two images, single inference pass
colorize_single_image(img_path, out_dir, prompt, steps=2, enhance_prompt=False) {"ok", "elapsed", "msg"} Single image fallback (odd batch end)

Colorization : in-memory (PNG bytes over RPC)

Method Returns Description
colorize_frame(img_data, prompt, img_size=0, steps=2, seed=42, skip_bw=False, enhance_prompt=False) {"ok", "data", "elapsed", "skipped", "msg"} Single frame as raw PNG bytes
colorize_frame_pair(img1_data, img2_data, prompt, gap_px=8, steps=2, enhance_prompt=False) {"ok", "data1", "data2", "elapsed", "skipped1", "skipped2", "msg"} Two frames, single inference pass

skipped=True means the frame was too dark to colorize (average brightness < 9/255). The returned data field contains the unchanged input in that case.

Colorization : shared memory (same-host only, zero-copy)

Method Returns Description
colorize_frame_shm(shm_in, shm_out, h, w, prompt, img_size=0, steps=2, seed=42, skip_bw=False, enhance_prompt=False) {"ok", "elapsed", "skipped", "msg"} Single frame via shared memory
colorize_frame_pair_shm(shm_in1, shm_out1, h1, w1, shm_in2, shm_out2, h2, w2, prompt, gap_px=8, steps=4, enhance_prompt=False) {"ok", "elapsed", "skipped1", "skipped2", "msg"} Two frames via shared memory, single inference pass

enhance_prompt (all methods above, default False) rewrites the prompt via Qwen3-VL before colorizing β€” only meaningful for qwen21-viggle; silently has no effect on the other backends. See What's New. See Shared Memory Transport for usage details.


πŸ§ͺ Example Clients

Both clients support two transport modes selectable via --use-shm:

Mode Flag When to use Measured speed (1480Γ—1080 px pair)
Standard RPC (default) Any deployment, including remote server ~5.25s/image
Shared memory --use-shm Server and client on the same host only ~4.06s/image (~23% faster)

The pipeline must be loaded on the server before running the clients. Start the server with --load-pipeline --pipeline-config CONFIG.json.

Single frame : dit_client_example.py

Colorizes assets/santa_bw.png and saves the result as assets/santa_colorized.png.

# standard RPC : works with local and remote server
python dit_client_example.py

# shared memory : same-host only, lower latency
python dit_client_example.py --use-shm

Windows: run_client_example.cmd To enable shared memory edit run_client_example.cmd and set USE_SHM=1.


Paired inference : dit_client_pair_example.py

Colorizes assets/sample1_bw.jpg and assets/sample2_bw.jpg in a single forward pass, saving assets/sample1_colorized.jpg and assets/sample2_colorized.jpg.

Paired inference places the two images side-by-side and runs one inference instead of two, roughly halving the per-image cost (~5.25s/image vs ~11s standalone). Combined with shared memory transport this reaches ~4.06s/image.

# standard RPC
python dit_client_pair_example.py

# shared memory : same-host only
python dit_client_pair_example.py --use-shm

Windows: run_client_pair_example.cmd To enable shared memory edit run_client_pair_example.cmd and set USE_SHM=1.

Full list of arguments (both clients)

  --host HOST                  Server host (default: 127.0.0.1)
  --port PORT                  Server port (default: 8765)
  --prompt PROMPT              Text prompt for the model
  --steps N                    Number of steps for inference (default:4)
  --use-shm                    Use shared memory transport (same-host only)

Additional argument for the paired client:

  --gap-px N                   Separator width in pixels between the two
                               images in the merged input (default: 8)

πŸš€ Shared Memory Transport (same-host only)

What it is

The standard RPC transport serializes each image as a PNG byte stream, encodes it in Base64, sends it over a TCP socket, and decodes it on the other side. For a 1480Γ—1080 frame this is roughly 4–5 MB per round trip.

The shared memory transport bypasses the network entirely. The client writes the raw pixel array directly into a shared memory segment; the server attaches to the same segment and reads the pixels without any copy. Only the metadata (segment name, dimensions, prompt) travels over the XML-RPC socket.

When you can use it

Requirement: server and client must run on the same machine.

If the server is on a dedicated GPU machine and the client is on a separate workstation, shared memory is not available : use the standard RPC transport instead (default). The clients detect this automatically: passing --use-shm when the host is not 127.0.0.1 / localhost prints a warning and falls back to standard RPC.

Performance

Measured on a 1480Γ—1080 pixel pair (RTX 5070 Ti, FP4, paired inference):

Transport Per-image time Round-trip overhead
Standard RPC (PNG) ~5.25s ~1.1s
Shared memory ~4.06s ~0.16s
Gain ~23% faster ~7Γ— less overhead

The round-trip overhead with shared memory is essentially zero : the 0.16s gap between inference time and wall-clock time is just Python function call and numpy overhead.

On a 100k-frame video processed as pairs (50k inference calls) the cumulative saving is:

(5.25 - 4.06) Γ— 50,000 β‰ˆ 16.5 hours

How the protocol works

The client owns and manages all shared memory segments. The server is fully stateless with respect to shared memory : it only attaches, reads/writes, and detaches.

Client                                     Server
  β”‚                                           β”‚
  β”‚  create shm_in  (h Γ— w Γ— 3 bytes)         β”‚
  β”‚  create shm_out (h Γ— w Γ— 3 bytes)         β”‚
  β”‚  write raw RGB pixels β†’ shm_in            β”‚
  β”‚                                           β”‚
  β”‚  RPC(shm_in_name, shm_out_name, h, w, …) ─►│
  β”‚                                           β”‚  attach shm_in  β†’ PIL Image
  β”‚                                           β”‚  inference
  β”‚                                           β”‚  result β†’ shm_out
  │◄─ return {elapsed, skipped, …} ───────────│
  β”‚                                           β”‚  detach both segments
  β”‚  read shm_out β†’ PIL Image                 β”‚
  β”‚  unlink shm_in + shm_out                  β”‚

Enabling shared memory

From the command line:

python dit_client_pair_example.py --use-shm
python dit_client_example.py      --use-shm

From the Windows .cmd launchers, edit the user configuration block and set:

set USE_SHM=1

The banner will confirm the active transport:

Transport   : 1 (0=RPC 1=shared memory)

And the Python client will print:

[INFO] Transport: shared memory

Implementing shared memory in your own client

import uuid
import numpy as np
from multiprocessing.shared_memory import SharedMemory
from PIL import Image

def colorize_pair_shm(proxy, img1: Image.Image, img2: Image.Image, prompt: str):
    arr1, arr2 = np.array(img1), np.array(img2)
    h1, w1 = arr1.shape[:2]
    h2, w2 = arr2.shape[:2]
    uid = uuid.uuid4().hex[:12]

    # Create all four segments (client owns them)
    segs = {
        tag: SharedMemory(name=f"dit_{tag}_{uid}", create=True, size=h*w*3)
        for tag, h, w in [("in1",h1,w1),("out1",h1,w1),("in2",h2,w2),("out2",h2,w2)]
    }
    try:
        np.ndarray((h1,w1,3), dtype=np.uint8, buffer=segs["in1"].buf)[:] = arr1
        np.ndarray((h2,w2,3), dtype=np.uint8, buffer=segs["in2"].buf)[:] = arr2

        result = proxy.colorize_frame_pair_shm(
            segs["in1"].name, segs["out1"].name, h1, w1,
            segs["in2"].name, segs["out2"].name, h2, w2,
            prompt, 8,  # gap_px
        )

        out1 = Image.fromarray(
            np.ndarray((h1,w1,3), dtype=np.uint8, buffer=segs["out1"].buf).copy())
        out2 = Image.fromarray(
            np.ndarray((h2,w2,3), dtype=np.uint8, buffer=segs["out2"].buf).copy())
        return result, out1, out2
    finally:
        for shm in segs.values():
            shm.close(); shm.unlink()

πŸͺŸ Windows Launch Script

start_server.cmd is a ready-to-use launcher for Windows. Edit the variables at the top of the file to match your setup, then double-click it or run it from a terminal.

start_server.cmd [q3|q4|q5|q6|q8|fp4|int4|longcat]
Argument Backend Quantization VRAM
(none) GGUF Q4_K_S 12 GB
q3 GGUF Q3_K_S 12 GB
q4 GGUF Q4_K_S 12 GB
q5 GGUF Q5_K_M 16 GB
q6 GGUF Q6_K 18 GB
q8 GGUF Q8_0 22 GB
fp4 Nunchaku FP4 16 GB
int4 Nunchaku INT4 16 GB
longcat LongCat Q4_K_M 12 GB

If no argument is passed it defaults to q4 (Q4_K_S). Use int4 for RTX 30 / 40-Series Nunchaku:

start_server.cmd int4

Convenience wrappers β€” double-click or run from terminal without arguments:

File Equivalent command Backend
run_server_q3.cmd start_server.cmd q3 GGUF Q3_K_S
run_server_fp4.cmd start_server.cmd fp4 Nunchaku FP4
run_server_int4.cmd start_server.cmd int4 Nunchaku INT4
run_server_longcat.cmd start_server.cmd longcat LongCat Q4_K_M

run_server_qwen21.cmd is a separate, standalone launcher for qwen21-viggle β€” it does not take an argument (start_server.cmd qwen21-viggle is not a thing), it always launches with config/qwen21_viggle.json (the only config available for this backend).

GUI shortcut: From the desktop GUI, go to Tab 2 (Colorization), pick a Model + Precision, and click Run Server β€” the GUI starts the server itself, with live output in the Server Log tab and auto-connect once it's ready (see What's New, 2026-09-30). Tick External console first to instead open a plain terminal window with the correct .cmd file/arguments for the selected Model Name (run_server_qwen21.cmd when qwen21-viggle is selected, start_server.cmd with the right arguments otherwise).


🎯 Suggested Inference Steps

Model Family Recommended Steps Notes
Qwen (nunchaku fp4/int4) 2 Good results with 2 steps when using lightning LoRA
Qwen (gguf q3–q8) 2 Default in config files; 4 steps possible but slower
LongCat (longcat-gguf) 8 Calibrated for 8 steps; best quality at 8 steps; 4 steps possible but colors are faded
Qwen-Image-2.1 (qwen21-viggle) 6 LoRA's native step count. 2/4/8 are experimental alternate schedules β€” 2 in particular trades a little brightness accuracy for ~25% less time, worth trying

Prompt tip (qwen21-viggle): on subjects with a strong color convention (e.g. a well-known costume), the model can leave the color ambiguous and resolve it inconsistently between runs. Before reaching for enhance_prompt (which adds ~15-20s/frame), try a direct, explicitly anti-hedging prompt β€” it solves the same problem for free in most cases:

"Add color to this black-and-white image without hesitation regarding the appropriate colors. For any subject, garment, object, or setting where the color is common knowledge or established by convention, confidently apply the expected color; otherwise use natural colors. Color the image by strictly preserving all shapes, outlines, and background details."

Avoid naming specific example subjects in this prompt (e.g. "like a stop sign") β€” the model may render that literal example into the scene instead of just using it as a color reference.


πŸ”§ Troubleshooting

CUDA out of memory Close other GPU applications. On 16 GB cards the server automatically enables sequential CPU offload for layers that do not fit in VRAM.

dit_colorize_main.py NOT FOUND Use --module-dir to point the server to the directory that contains dit_colorize_main.py:

python dit_rpc_server.py --module-dir /path/to/dit_colorize_main

Model 'xxx' is not supported Supported values for model_name are "nunchaku-qwen" (FP4/INT4), "gguf-qwen" (Q3_K_S, Q4_K_S, Q5_K_M, Q6_K, Q8_0), "longcat-gguf", and "qwen21-viggle". For "gguf-qwen", the quantization is selected via the quant field in the config (e.g. "q4").

Pipeline takes a long time to load Nunchaku: on the first run the model weights (~15–30 GB) are downloaded from HuggingFace. Subsequent runs load from the local cache. GGUF: only the VAE and tokenizer (~320 MB) are downloaded from HuggingFace; the UNet and CLIP are loaded directly from the local .gguf files. Set cache_dir in the config to control where the cache is stored.


πŸ”— Credits

About

Hybrid Automatic Video Colorizer (HAVC) server that exposes a GPU-accelerated colorization pipeline for B&W images and video frames based on Diffusion Transformer (DiT) models: Qwen-Image-Edit-2511 (with Nunchaku transformer) and LongCat ImageEdit Turbo. Colorized reference frames are propagated using CMNET2.

Topics

Resources

Stars

14 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages