[{"data":1,"prerenderedAt":31},["ShallowReactive",2],{"chapter:vision-rs\u002Fkernels-and-performance\u002Fteenyc-toolchain.json":3},{"project":4,"route":5,"title":6,"titleHtml":7,"navTitle":6,"part":8,"sourcePath":9,"editUrl":10,"html":11,"toc":12,"hasMermaid":23,"prev":24,"next":27},"vision-rs","\u002Fvision-rs\u002Fkernels-and-performance\u002Fteenyc-toolchain","The teenyc Toolchain","The \u003Ccode>teenyc\u003C\u002Fcode> Toolchain","Kernels & Performance","kernels-and-performance\u002Fteenyc-toolchain.md","https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fvision-rs\u002Fedit\u002Fmain\u002Fbook\u002Fsrc\u002Fkernels-and-performance\u002Fteenyc-toolchain.md","\u003Cp>vision-rs’s kernels (see \u003Ca href=\"\u002Fvision-rs\u002Fkernels-and-performance\u002Fcustom-kernels\">Custom Kernels\u003C\u002Fa>) are\nordinary-looking \u003Ccode>#[kernel]\u003C\u002Fcode>-annotated Rust functions, but they’re compiled\nto PTX\u002FMLIR by a separate, custom compiler fork — \u003Ccode>teenyc\u003C\u002Fcode> — not by the\nstable \u003Ccode>rustc\u003C\u002Fcode> you build the rest of vision-rs with. See\n\u003Ca href=\"\u002Fvision-rs\u002Fgetting-started\u002Finstallation\">Installation &amp; Toolchain\u003C\u002Fa> for how to\ninstall it.\u003C\u002Fp>\n\u003Ch2 id=\"two-compilation-modes\">Two compilation modes\u003C\u002Fh2>\n\u003Cp>\u003Cstrong>JIT (just-in-time)\u003C\u002Fstrong>: during normal development — running \u003Ccode>cargo test\u003C\u002Fcode>,\nthe \u003Ccode>yolo26\u003C\u002Fcode> example, or anything else that touches the \u003Ccode>cuda\u003C\u002Fcode> feature —\nkernels are compiled by \u003Ccode>teenyc\u003C\u002Fcode> at runtime, the first time each kernel is\ninvoked, then cached. This is what \u003Ccode>TEENYC_PATH\u003C\u002Fcode> is for: the process needs\nto be able to shell out to the compiler on demand.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>AOT (ahead-of-time)\u003C\u002Fstrong>: \u003Ccode>cargo teeny package\u003C\u002Fcode>\u002F\u003Ccode>aot\u003C\u002Fcode> cross-compiles kernels\n\u003Cem>before\u003C\u002Fem> deployment, producing a \u003Ccode>cache\u002F\u003C\u002Fcode> directory of precompiled PTX that\nships alongside the binary. This is how vision-rs runs on a Jetson Orin\nNano with no \u003Ccode>teenyc\u003C\u002Fcode> (or even Rust toolchain) installed on the device — see\n\u003Ca href=\"\u002Fvision-rs\u002Fdeployment\u002Fpackaging\">Packaging a Deployable Bundle\u003C\u002Fa>. The binary\nauto-detects a sibling \u003Ccode>cache\u002F\u003C\u002Fcode> directory at runtime and uses it instead of\ntrying to JIT-compile, which would simply fail on a device without\n\u003Ccode>teenyc\u003C\u002Fcode>.\u003C\u002Fp>\n\u003Ch2 id=\"why-a-separate-compiler\">Why a separate compiler\u003C\u002Fh2>\n\u003Cp>\u003Ccode>teenyc\u003C\u002Fcode> exists because the kernel DSL compiles through an MLIR backend\nthat isn’t part of upstream \u003Ccode>rustc\u003C\u002Fcode>. If a kernel fails to compile with an\nerror that doesn’t obviously trace back to anything in vision-rs or\nteenygrad’s Rust-level code, the root cause may be in this compiler fork\nrather than in either of those — see\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fteeny\" target=\"_blank\" rel=\"noopener noreferrer\">teenygrad\u002Fteeny\u003C\u002Fa>.\u003C\u002Fp>\n\u003Ch2 id=\"compute-capability\">Compute capability\u003C\u002Fh2>\n\u003Cp>AOT-compiled kernels are compiled for a specific GPU compute capability\n(e.g. \u003Ccode>sm_87\u003C\u002Fcode> for the Jetson Orin Nano’s Ampere GPU), passed via\n\u003Ccode>--options capability=sm_87\u003C\u002Fcode> to \u003Ccode>cargo teeny package\u003C\u002Fcode>. JIT compilation\ninstead targets whatever capability the local device reports at runtime.\nSee \u003Ca href=\"\u002Fvision-rs\u002Fdeployment\u002Fpackaging\">Packaging a Deployable Bundle\u003C\u002Fa> for the\nfull \u003Ccode>package\u003C\u002Fcode> invocation and what \u003Ccode>ptx-version\u003C\u002Fcode> overrides are for.\u003C\u002Fp>\n",[13,17,20],{"id":14,"text":15,"level":16},"two-compilation-modes","Two compilation modes",2,{"id":18,"text":19,"level":16},"why-a-separate-compiler","Why a separate compiler",{"id":21,"text":22,"level":16},"compute-capability","Compute capability",false,{"title":25,"titleHtml":25,"route":26},"Custom Kernels","\u002Fvision-rs\u002Fkernels-and-performance\u002Fcustom-kernels",{"title":28,"titleHtml":29,"route":30},"Benchmarking & Profiling","Benchmarking &amp; Profiling","\u002Fvision-rs\u002Fkernels-and-performance\u002Fbenchmarking",1786271830534]