[{"data":1,"prerenderedAt":34},["ShallowReactive",2],{"chapter:teenygrad\u002Fdeployment\u002Fteeny-quant.json":3},{"project":4,"route":5,"title":6,"titleHtml":7,"navTitle":6,"part":8,"sourcePath":9,"editUrl":10,"html":11,"toc":12,"hasMermaid":26,"prev":27,"next":31},"teenygrad","\u002Fteenygrad\u002Fdeployment\u002Fteeny-quant","teeny-quant and Model Quantization","\u003Ccode>teeny-quant\u003C\u002Fcode> and Model Quantization","Deployment","deployment\u002Fteeny-quant.md","https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fteenygrad\u002Fedit\u002Fmain\u002Fbooks\u002Fteenygrad\u002Fsrc\u002Fdeployment\u002Fteeny-quant.md","\u003Cp>\u003Ca href=\"\u002Fapi\u002Fteenygrad\u002Fteenygrad\u002Fteeny_quant\u002F\">\u003Ccode>teeny-quant\u003C\u002Fcode>\u003C\u002Fa> quantizes \u003Ccode>.safetensors\u003C\u002Fcode> model\ncheckpoints for deployment — smaller weights, faster inference — as both a reusable library and a\n\u003Ccode>teeny-quant\u003C\u002Fcode> binary. It’s initially being validated against Ultralytics YOLO models.\u003C\u002Fp>\n\u003Cblockquote>\n\u003Cp>⚠️ \u003Cstrong>Weight-only today.\u003C\u002Fstrong> \u003Ccode>teeny-quant\u003C\u002Fcode> currently does post-training quantization from the\ncheckpoint’s weights alone. Static \u003Cem>activation\u003C\u002Fem> quantization (calibrating scales from a forward\npass over sample inputs, TensorRT-INT8-calibration style) is planned but not implemented yet —\nit needs a model forward pass, which means running an ONNX export through\n\u003Ca href=\"\u002Fapi\u002Fteenygrad\u002Fteenygrad\u002Fteeny_onnx\u002F\">\u003Ccode>teeny-onnx\u003C\u002Fcode>\u003C\u002Fa> and \u003Ccode>teeny-compiler\u003C\u002Fcode>’s \u003Ccode>ndarray\u003C\u002Fcode>\nbackend, and op coverage there hasn’t been verified for a CNN detection model’s op set (\u003Ccode>Conv\u003C\u002Fcode>,\n\u003Ccode>BatchNormalization\u003C\u002Fcode>, \u003Ccode>Concat\u003C\u002Fcode>, …).\u003C\u002Fp>\n\u003C\u002Fblockquote>\n\u003Ch2 id=\"schemes-and-granularity\">Schemes and granularity\u003C\u002Fh2>\n\u003Cp>Three quantization schemes:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>INT8\u003C\u002Fstrong> — symmetric or asymmetric affine quantization.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>INT4\u003C\u002Fstrong> — same affine math as INT8, nibble-packed two values per \u003Ccode>U8\u003C\u002Fcode> byte (\u003Ccode>safetensors\u003C\u002Fcode> has\nno native 4-bit dtype).\u003C\u002Fli>\n\u003Cli>\u003Cstrong>FP8\u003C\u002Fstrong> — \u003Ccode>F8_E4M3\u003C\u002Fcode> or \u003Ccode>F8_E5M2\u003C\u002Fcode>, both natively supported \u003Ccode>safetensors\u003C\u002Fcode> dtypes.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Independently of scheme, pick a \u003Cstrong>granularity\u003C\u002Fstrong>: \u003Ccode>tensor\u003C\u002Fcode> (one scale for the whole tensor),\n\u003Ccode>channel\u003C\u002Fcode> (one scale per index along an axis, reduced over every other axis — the usual choice\nfor per-output-channel weight quantization), or \u003Ccode>group\u003C\u002Fcode> (GPTQ\u002FAWQ-style: subdivide one axis into\nfixed-size chunks, with every other axis getting its own independent set of groups). \u003Ccode>channel\u003C\u002Fcode> and\n\u003Ccode>group\u003C\u002Fcode> are \u003Cem>not\u003C\u002Fem> the same iteration pattern — see the crate’s \u003Ccode>quant::groups\u003C\u002Fcode> module docs if\nyou’re calling the library directly rather than the CLI.\u003C\u002Fp>\n\u003Ch2 id=\"cli\">CLI\u003C\u002Fh2>\n\u003Cpre data-lang=\"bash\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#7F877D;font-style:italic\"># INT8, per-channel (the default granularity).\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#7FB6D9\">cargo\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> run\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> -p\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny-quant\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --bin\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny-quant\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> quantize\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> \\\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#B79AD4\">  --input\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> model.safetensors\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --output\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> model-int8.safetensors\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --scheme\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> int8\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#7F877D;font-style:italic\"># INT4, group-wise along the reduction axis (axis 1 for a [out, in] weight matrix).\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#7FB6D9\">cargo\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> run\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> -p\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny-quant\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --bin\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny-quant\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> quantize\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> \\\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#B79AD4\">  --input\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> model.safetensors\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --output\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> model-int4.safetensors\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> \\\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#B79AD4\">  --scheme\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> int4\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --granularity\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> group\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --axis\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 1\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --group-size\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 128\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#7F877D;font-style:italic\"># What's in a checkpoint (plain or quantized) -- tensor names\u002Fdtypes\u002Fshapes, plus the\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#7F877D;font-style:italic\"># quantization_config if present.\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#7FB6D9\">cargo\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> run\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> -p\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny-quant\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --bin\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny-quant\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> inspect\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> model-int8.safetensors\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#7F877D;font-style:italic\"># Per-tensor quantization error: max abs error, mean abs error, SQNR (dB).\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#7FB6D9\">cargo\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> run\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> -p\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny-quant\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --bin\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny-quant\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> validate\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> \\\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#B79AD4\">  --original\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> model.safetensors\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --quantized\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> model-int8.safetensors\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Tensors with rank &lt; 2 (biases, norm weights) are left unquantized and listed in the output’s\n\u003Ccode>ignore\u003C\u002Fcode> metadata, rather than quantized — the usual default for PTQ tooling.\u003C\u002Fp>\n\u003Ch2 id=\"output-format\">Output format\u003C\u002Fh2>\n\u003Cp>Output follows the \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fcompressed-tensors\" target=\"_blank\" rel=\"noopener noreferrer\">compressed-tensors\u003C\u002Fa>\nconvention layered on plain \u003Ccode>.safetensors\u003C\u002Fcode>, so quantized checkpoints stay loadable by existing\nHF\u002FvLLM-side tooling for INT8\u002FFP8: a quantized \u003Ccode>foo.weight\u003C\u002Fcode> keeps its name, gains a\n\u003Ccode>foo.weight_scale\u003C\u002Fcode> (and \u003Ccode>foo.weight_zero_point\u003C\u002Fcode> for asymmetric schemes) sibling tensor, and the\nfile’s \u003Ccode>__metadata__\u003C\u002Fcode> header carries a \u003Ccode>quantization_config\u003C\u002Fcode> JSON blob describing the scheme,\ngranularity, and which tensors were left unquantized. INT4’s nibble-packing isn’t bit-compatible\nwith compressed-tensors’ own int32-based \u003Ccode>pack-quantized\u003C\u002Fcode> layout — that’s \u003Ccode>teeny-quant\u003C\u002Fcode>’s own,\ndocumented convention (see \u003Ccode>quant::pack4\u003C\u002Fcode>).\u003C\u002Fp>\n\u003Ch2 id=\"relationship-to-teeny-cores-dtype-system\">Relationship to \u003Ccode>teeny-core\u003C\u002Fcode>’s dtype system\u003C\u002Fh2>\n\u003Cp>\u003Ca href=\"\u002Fteenygrad\u002Fcore-concepts\u002Fdtype-system\">\u003Ccode>teeny-core::dtype\u003C\u002Fcode>\u003C\u002Fa> defines \u003Ccode>F8E4M3FN\u003C\u002Fcode>\u002F\u003Ccode>BF16\u003C\u002Fcode>\u002F\u003Ccode>I4\u003C\u002Fcode> as marker\ntraits for a future typed kernel dtype system, but has no concrete implementations yet.\n\u003Ccode>teeny-quant\u003C\u002Fcode> doesn’t depend on or wait for that — it works directly on raw \u003Ccode>safetensors\u003C\u002Fcode> bytes\n(\u003Ccode>safetensors::Dtype\u003C\u002Fcode>, not \u003Ccode>teeny_core::dtype::Dtype\u003C\u002Fcode>), including its own from-scratch FP8\nbit-conversion codec, since neither \u003Ccode>teeny-core\u003C\u002Fcode> nor the \u003Ccode>half\u003C\u002Fcode> crate (f16\u002Fbf16 only) has one.\u003C\u002Fp>\n",[13,17,20,23],{"id":14,"text":15,"level":16},"schemes-and-granularity","Schemes and granularity",2,{"id":18,"text":19,"level":16},"cli","CLI",{"id":21,"text":22,"level":16},"output-format","Output format",{"id":24,"text":25,"level":16},"relationship-to-teeny-cores-dtype-system","Relationship to teeny-core’s dtype system",false,{"title":28,"titleHtml":29,"route":30},"teeny-llm","\u003Ccode>teeny-llm\u003C\u002Fcode>","\u002Fteenygrad\u002Fcli-and-aot\u002Fteeny-llm",{"title":32,"titleHtml":32,"route":33},"Contributing to Teenygrad","\u002Fteenygrad\u002Fcontributing\u002Fcontributing",1786271829542]