viralvfx GitHub avatar

ComfyUI-INT4-Fast

viralvfx

ComfyUI-INT4-Fast is a custom node package that enables loading, running, and saving diffusion models in INT4 format with Tensor Core acceleration.

Stars

33

7-day growth

No data

Forks

3

Open issues

3

License

AGPL-3.0

Last updated

2026-07-10

AI repository intelligence
FR-AI / ANALYSIS

Why it is worth attention

It brings native INT4 inference to ComfyUI, leveraging Tensor Cores for speed and memory efficiency, with built-in mixed-precision handling and dynamic LoRA patching.

Who it is for

  • ComfyUI users wanting to reduce model memory footprint
  • AI artists and developers working with diffusion models
  • Researchers interested in quantization techniques
  • Users with GPU Tensor Cores seeking faster inference

Use cases

  • Run large diffusion models with reduced VRAM usage
  • Quantize float checkpoints to INT4 on the fly
  • Save quantized models for reuse in ComfyUI
  • Integrate LoRA adapters with quantized models

Strengths

  • Ultra-fast INT4 inference via Tensor Cores
  • Mixed-precision support for sensitive layers (first/last patches)
  • On-the-fly quantization from BF16/FP16/FP32
  • Dynamic LoRA weight rotation for compatibility

Considerations

  • First model generation run requires extra compilation time
  • Depends on comfy-kitchen package for execution layouts
  • Only one model (Krea2 Turbo INT4) has been verified
  • Requires a GPU with Tensor Cores (e.g., NVIDIA RTX series)

README quick start

Installation & Setup

Description

Fast INT4 model inference custom node for ComfyUI leveraging Tensor Cores.

Related repositories

Similar projects matched by category, topics, and programming language.

lopopolo
Featured
lopopolo GitHub avatar

harness-engineering

Harness Engineering is a methodology for improving coding agent outputs by carefully crafting the environment around them—providing curated context, tools, and executable constraints that encode an organization’s nonfunctional requirements and cumulative lessons.

AI & Machine LearningAI Agents
2,390
slvDev
Featured
slvDev GitHub avatar

esp32-ai

A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.

AI & Machine LearningLarge Language Models
1,960
littledivy
Featured
littledivy GitHub avatar

mimic

mimic captures traffic from any iOS or web app and automatically generates a Python client library that lets you call the app's API like a regular library.

AI & Machine Learning
1,482