open model lab
Big models,
small machines.
Sapid Labs quantizes and prunes open LLMs so they run on the hardware in front of you. Laptops, single GPUs, home rigs. We ship the weights on Hugging Face, free to use.
what we do
quantize
4-bit, 8-bit, GGUF, AWQ. We squeeze weights down so big models fit in the memory you actually have.
prune
Expert pruning and REAP on MoE models. Fewer experts, same feel, a fraction of the footprint.
benchmark
Every release ships with real tok/s and quality numbers on real local hardware. No hand waving.
latest models
all on hugging faceGLM-5.2-NVFP4-attn-experimental
quantized from zai-org/GLM-5.2-FP8
330updated 6d ago
GLM-5.2-2bit-MoE-planes-pruned208-tp2
quantized from zai-org/GLM-5.2-FP8
expert pruningMoEtext-generation
01updated 7d ago
GLM-5.2-moe-w2-planes
quantized from zai-org/GLM-5.2-FP8
01updated 9d ago
DeepSeek-V4-Flash-moe-w2-planes
quantized from deepseek-ai/DeepSeek-V4-Flash
00updated 10d ago
Hy3-REAP-48e
quantized from tencent/Hy3
REAP pruningexpert pruningpruningMoEtext-generation
2744updated 12d ago
Nemotron-3-Nano-30B-A3B-REAP-64e
quantized from nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
REAP pruningexpert pruningpruningMoEtext-generation
4001updated 12d ago