open model lab
Big models,
small machines.
Sapid Labs quantizes and prunes open LLMs so they run on the hardware in front of you. Laptops, single GPUs, home rigs. We ship the weights on Hugging Face, free to use.
what we do
quantize
4-bit, 8-bit, GGUF, AWQ. We squeeze weights down so big models fit in the memory you actually have.
prune
Expert pruning and REAP on MoE models. Fewer experts, same feel, a fraction of the footprint.
benchmark
Every release ships with real tok/s and quality numbers on real local hardware. No hand waving.
latest models
all on hugging faceSparkulator-Qwen3.6-27B
quantized from nvidia/Qwen3.6-27B-NVFP4
Sparkulator-Gemma-4-26B-A4B
quantized from google/gemma-4-26B-A4B-it-assistant
Sparkulator-Laguna-S-2.1
quantized from poolside/Laguna-S-2.1-DFlash-FP8
Sparkulator-Laguna-S-2.1-NVFP4
quantized from poolside/Laguna-S-2.1-DFlash-NVFP4
Sparkulator-Laguna-XS-2.1
quantized from poolside/Laguna-XS-2.1-DFlash
Sparkulator-GLM-5.2
quantized from RedHatAI/GLM-5.2-speculator.dspark
GLM-5.2-NVFP4-attn-experimental
quantized from zai-org/GLM-5.2-FP8
GLM-5.2-2bit-MoE-planes-pruned208-tp2
quantized from zai-org/GLM-5.2-FP8
GLM-5.2-moe-w2-planes
quantized from zai-org/GLM-5.2-FP8
DeepSeek-V4-Flash-moe-w2-planes
quantized from deepseek-ai/DeepSeek-V4-Flash
Hy3-REAP-48e
quantized from tencent/Hy3
Nemotron-3-Nano-30B-A3B-REAP-64e
quantized from nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16