drowzeys GitHub avatar

keys-GLM5.2-Quantrio-INT4-INT8-Mixed-Abliterated-C1-30toks-4x-DGX-Sparks

drowzeys

这是一个已归档的仓库,用于将用户重定向至 GLM-5.2 Quantrio INT4/INT8 混合消融模型及 DFlash 推测解码的规范资源。

Stars

12

7 天增长

暂无数据

Fork 数

1

开放 Issue

1

开源协议

NOASSERTION

最近更新

2026-07-18

AI 仓库情报摘要
FR-AI / ANALYSIS

为什么值得关注

它明确指出了唯一支持的模型栈,并提供了完整的 DFlash 配置参数,避免用户安装过时版本。

适合谁使用

  • 部署 GLM-5.2 量化模型的机器学习工程师
  • 研究消融模型的研究人员
  • 需要 DFlash 推测解码规格的用户

典型使用场景

  • 查找正确的模型权重和草稿模型
  • 了解 DFlash 配置(k=12, 128k 上下文, UTIL 0.85, KV pin 10 GiB)
  • 避免安装已弃用的软件版本

项目优势

  • 明确的弃用声明及指向规范仓库的链接
  • 详尽的 DFlash 参数表
  • 直接提供 Hugging Face 上的权重和草稿模型地址

使用前须知

  • 本仓库不含任何代码或模型文件
  • 仅作为重定向用途,用户需访问外部链接
  • 仅针对 GLM-5.2 Quantrio 这一特定模型变体

README 快速开始

⚠️ SUPERSEDED — do not install from this repository

Canonical (only) recipe:

https://github.com/drowzeys/keys-latest-GLM-5.2-Quantrio-INT4-INT8mixed-Abliterated-DFlash-4x-DGX-Sparks

This older C1-30toks title is archived to prevent agents/humans from installing the wrong stack.

项目描述

GLM-5.2 Quantrio INT4/INT8 Mixed Abliterated — SPEED=1 C1≈30 tok/s @ 128k on 4x DGX Spark. Step-by-step recipe, image bake, results.

相关仓库与替代方案

根据分类、Topic 和编程语言匹配的相似项目。

lopopolo
精选
lopopolo GitHub avatar

harness-engineering

Harness Engineering is a methodology for improving coding agent outputs by carefully crafting the environment around them—providing curated context, tools, and executable constraints that encode an organization’s nonfunctional requirements and cumulative lessons.

AI 与机器学习AI 智能体
2,390
slvDev
精选
slvDev GitHub avatar

esp32-ai

A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.

AI 与机器学习大语言模型
1,960
littledivy
精选
littledivy GitHub avatar

mimic

mimic captures traffic from any iOS or web app and automatically generates a Python client library that lets you call the app's API like a regular library.

AI 与机器学习
1,482