drowzeys GitHub avatar

keys-GLM5.2-Quantrio-INT4-INT8-Mixed-Abliterated-C1-30toks-4x-DGX-Sparks

drowzeys

This is an archived repository that redirects users to the canonical recipe for the GLM-5.2 Quantrio INT4/INT8 mixed abliterated model with DFlash speculative decoding.

Stars

12

7-day growth

No data

Forks

1

Open issues

1

License

NOASSERTION

Last updated

2026-07-18

AI repository intelligence
FR-AI / ANALYSIS

Why it is worth attention

It provides a clear, authoritative pointer to the only supported stack for a specific high-performance LLM configuration, preventing installation of an outdated version.

Who it is for

  • ML engineers deploying GLM-5.2 with quantization
  • Researchers working with abliterated models
  • Users needing DFlash speculative decoding specs

Use cases

  • Locating the correct model weights and draft model
  • Understanding the DFlash profile (k=12, 128k context, etc.)
  • Avoiding installation of deprecated software versions

Strengths

  • Explicit superseding notice and link to canonical recipe
  • Detailed DFlash profile specifications (k=12, 128k, UTIL 0.85, KV pin 10 GiB)
  • Direct URLs to Hugging Face weights and draft model

Considerations

  • The repository itself contains no code or model files
  • Intended only as a redirect; users must follow the external link
  • Limited to one specific model variant (GLM-5.2 Quantrio)

README quick start

⚠️ SUPERSEDED — do not install from this repository

Canonical (only) recipe:

https://github.com/drowzeys/keys-latest-GLM-5.2-Quantrio-INT4-INT8mixed-Abliterated-DFlash-4x-DGX-Sparks

This older C1-30toks title is archived to prevent agents/humans from installing the wrong stack.

Description

GLM-5.2 Quantrio INT4/INT8 Mixed Abliterated — SPEED=1 C1≈30 tok/s @ 128k on 4x DGX Spark. Step-by-step recipe, image bake, results.

Related repositories

Similar projects matched by category, topics, and programming language.

lopopolo
Featured
lopopolo GitHub avatar

harness-engineering

Harness Engineering is a methodology for improving coding agent outputs by carefully crafting the environment around them—providing curated context, tools, and executable constraints that encode an organization’s nonfunctional requirements and cumulative lessons.

AI & Machine LearningAI Agents
2,390
slvDev
Featured
slvDev GitHub avatar

esp32-ai

A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.

AI & Machine LearningLarge Language Models
1,960
littledivy
Featured
littledivy GitHub avatar

mimic

mimic captures traffic from any iOS or web app and automatically generates a Python client library that lets you call the app's API like a regular library.

AI & Machine Learning
1,482