[{"data":1,"prerenderedAt":35},["ShallowReactive",2],{"chapter:kernels\u002Fin-a-model\u002Fcross-building.json":3},{"project":4,"route":5,"title":6,"titleHtml":6,"navTitle":6,"part":7,"sourcePath":8,"editUrl":9,"html":10,"toc":11,"hasMermaid":28,"prev":29,"next":32},"kernels","\u002Fkernels\u002Fin-a-model\u002Fcross-building","Building for Another Target","Kernels in a Real Model","in-a-model\u002Fcross-building.md","https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fteenygrad\u002Fedit\u002Fmain\u002Fbooks\u002Fkernels\u002Fsrc\u002Fin-a-model\u002Fcross-building.md","\u003Cp>The machine you develop on and the machine that runs your model are often not\nthe same. A Jetson on a robot has an Arm CPU, a different GPU generation, and no\nappetite for compiling anything.\u003C\u002Fp>\n\u003Cp>Two separate problems, and it helps to keep them apart:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>The CPU code\u003C\u002Fstrong> — your program — must be compiled for the board’s\narchitecture.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>The GPU code\u003C\u002Fstrong> — your kernels — must be compiled for the board’s compute\ncapability.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>\u003Ccode>cargo-teeny\u003C\u002Fcode> handles both, with a different subcommand each.\u003C\u002Fp>\n\u003Ch2 id=\"capability-not-architecture\">Capability, not architecture\u003C\u002Fh2>\n\u003Cp>Chapter 4 listed the compute capabilities. The one that matters here is \u003Ccode>sm_87\u003C\u002Fcode>,\nJetson Orin — because it is the case where a developer’s desktop and the target\ndiffer in a way that silently produces a binary that will not run.\u003C\u002Fp>\n\u003Cp>PTX gives you some slack. It is an intermediate form, and the driver compiles it\nto machine code at load time, so PTX built for an older capability generally\nruns on a newer card. What it will not do is use instructions the older\ncapability did not have. Build for \u003Ccode>sm_75\u003C\u002Fcode> and run on an \u003Ccode>sm_90\u003C\u002Fcode> and you get a\nworking kernel that leaves the newer Tensor Cores idle.\u003C\u002Fp>\n\u003Cp>So: build for the capability you will run on.\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> target \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> Target\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">new\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">Capability\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">Sm87\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> ptx_path \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\"> compile_kernel\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(&#x26;\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">kernel\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> &#x26;\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">target\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> false\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">)?;\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Nothing about this needs an \u003Ccode>sm_87\u003C\u002Fcode> device present. Compiling for a capability\nand having one are unrelated — which is what makes cross-building possible at\nall.\u003C\u002Fp>\n\u003Ch2 id=\"cross-compiling-the-program\">Cross-compiling the program\u003C\u002Fh2>\n\u003Cpre data-lang=\"bash\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#7FB6D9\">cargo\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> build\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --target\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> jetson-orin-nano\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#7FB6D9\">cargo\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> check\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --target\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> jetson-orin-nano\u003C\u002Fspan>\u003Cspan style=\"color:#7F877D;font-style:italic\">          # faster feedback\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#7FB6D9\">cargo\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> clippy\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --target\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> jetson-orin-nano\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#7FB6D9\">cargo\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> build\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --target\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> jetson-orin-nano\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --example\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> yolo26\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>This delegates to \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fcross-rs\u002Fcross\" target=\"_blank\" rel=\"noopener noreferrer\">\u003Ccode>cross\u003C\u002Fcode>\u003C\u002Fa>, which builds\ninside a container holding the target’s toolchain, and handles two things\n\u003Ccode>cross\u003C\u002Fcode> does not do on its own:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>It resolves the teenygrad workspace root from your \u003Ccode>[patch.crates-io]\u003C\u002Fcode> entries\nand mounts it, because \u003Ccode>cross\u003C\u002Fcode> mounts individual crate directories and that is\nnot enough for workspace inheritance.\u003C\u002Fli>\n\u003Cli>It mounts the host’s CUDA aarch64 target directory where the board’s image\nexpects it.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>\u003Ccode>cargo teeny check --target jetson-orin-nano\u003C\u002Fcode> is a genuinely useful thing to run\non a laptop. It type-checks everything, including your kernels, for a board you\ndo not own.\u003C\u002Fp>\n\u003Cp>If the board needs libraries you do not have locally, \u003Ccode>cargo teeny sysroot\u003C\u002Fcode> lays\nout an FHS-style tree and can \u003Ccode>rsync\u003C\u002Fcode> the real thing off the device:\u003C\u002Fp>\n\u003Cpre data-lang=\"bash\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#7FB6D9\">cargo\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> sysroot\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --host\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> aarch64-unknown-linux-gnu\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --path\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> .\u002Fsysroot\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> \\\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#B79AD4\">  --type\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> jetson-orin-nano\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --rsync-from\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> ubuntu@jetson\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Ch2 id=\"compiling-kernels-ahead-of-time\">Compiling kernels ahead of time\u003C\u002Fh2>\n\u003Cp>Chapter 3 said kernel compilation happens at run time, on the first call, cached\non disk. On a development machine that is a one-off cost you never notice.\u003C\u002Fp>\n\u003Cp>On a deployed board it is a problem. The first inference pays for compiling every\nkernel in the model, \u003Ccode>teenyc\u003C\u002Fcode> has to be installed on the board, and a read-only\nor space-constrained filesystem may not have anywhere to put a cache.\u003C\u002Fp>\n\u003Cp>Ahead-of-time compilation moves that to build time:\u003C\u002Fp>\n\u003Cpre data-lang=\"bash\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#7FB6D9\">cargo\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> aot\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --example\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> yolo26\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --device\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> cuda\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --options\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> \"\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\">capability=sm_87,ptx-version=82\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">\"\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>The mechanism is worth understanding, because it is unusual. \u003Ccode>aot\u003C\u002Fcode> builds your\nbinary \u003Cstrong>for the host\u003C\u002Fstrong> — not cross-compiled — and \u003Cem>runs\u003C\u002Fem> it, with flags telling\nit to compile its kernels for the named capability and stop. So the program\nitself does the compiling; \u003Ccode>cargo-teeny\u003C\u002Fcode> forwards \u003Ccode>--device\u003C\u002Fcode>, \u003Ccode>--options\u003C\u002Fcode>,\n\u003Ccode>--cache-dir\u003C\u002Fcode> and \u003Ccode>--force\u003C\u002Fcode> verbatim without parsing them.\u003C\u002Fp>\n\u003Cp>What lands in the cache directory is PTX for \u003Ccode>sm_87\u003C\u002Fcode>, produced on your desktop.\u003C\u002Fp>\n\u003Ch2 id=\"packaging-and-deploying\">Packaging and deploying\u003C\u002Fh2>\n\u003Cpre data-lang=\"bash\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#7FB6D9\">cargo\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> package\u003C\u002Fspan>\u003Cspan style=\"color:#7F877D;font-style:italic\">    # cross-compiles the binary + AOT-compiles its kernels\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#7FB6D9\">cargo\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> deploy\u003C\u002Fspan>\u003Cspan style=\"color:#7F877D;font-style:italic\">     # ships the result over SSH\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>\u003Ccode>package\u003C\u002Fcode> runs both of the previous steps in one go, forcing the AOT cache\ndirectory to \u003Ccode>&lt;dest&gt;\u002Fcache\u003C\u002Fcode> so the layout is right — \u003Ccode>cache\u002F\u003C\u002Fcode> beside \u003Ccode>bin\u002F\u003C\u002Fcode>.\u003C\u002Fp>\n\u003Cp>That layout is not arbitrary. \u003Ccode>default_cache_dir\u003C\u002Fcode> looks for a \u003Ccode>cache\u002F\u003C\u002Fcode> directory\nnext to the executable and uses it if present, falling back to\n\u003Ccode>TEENYC_CACHE_DIR\u003C\u002Fcode> or \u003Ccode>\u002Ftmp\u002Fteenyc_cache\u003C\u002Fcode>. So a packaged binary finds its\nprecompiled kernels with no environment variables set — and a normal \u003Ccode>cargo run\u003C\u002Fcode>\nduring development, whose executable lives under \u003Ccode>target\u002Fdebug\u002F\u003C\u002Fcode> with no \u003Ccode>cache\u002F\u003C\u002Fcode>\nsibling, is unaffected.\u003C\u002Fp>\n\u003Ch2 id=\"what-to-check-before-you-ship\">What to check before you ship\u003C\u002Fh2>\n\u003Cp>\u003Cstrong>The capability matches.\u003C\u002Fstrong> Building \u003Ccode>sm_89\u003C\u002Fcode> PTX for an Orin gets you a runtime\nfailure, not a compile error.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>The PTX version is one the board’s driver accepts.\u003C\u002Fstrong> This is the failure that\ncatches people out. On some Blackwell cards \u003Ccode>teenyc\u003C\u002Fcode>’s default PTX version is\nrejected outright:\u003C\u002Fp>\n\u003Cpre class=\"code-panel\" data-lang=\"text\">\u003Ccode>PTX .version 8.6 does not support .target sm_120a\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>The fix is \u003Ccode>TEENYC_PTX_VERSION=87\u003C\u002Fcode>, or \u003Ccode>ptx-version=\u003C\u002Fcode> in the AOT options. It is\na \u003Ccode>teenyc\u003C\u002Fcode>-side default, not something the SDK can work around, and this tree’s\nbenches carry a note about it.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>The cache shipped.\u003C\u002Fstrong> A packaged binary with an empty \u003Ccode>cache\u002F\u003C\u002Fcode> will try to\ncompile at run time and fail, because \u003Ccode>teenyc\u003C\u002Fcode> is not installed on the board.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>The dtypes match.\u003C\u002Fstrong> A model quantised on the host must be loaded as the dtype\nit was written as. Chapter 19 covers what silently goes wrong when it is not.\u003C\u002Fp>\n\u003Cp>Next: how much of any of this transfers.\u003C\u002Fp>\n",[12,16,19,22,25],{"id":13,"text":14,"level":15},"capability-not-architecture","Capability, not architecture",2,{"id":17,"text":18,"level":15},"cross-compiling-the-program","Cross-compiling the program",{"id":20,"text":21,"level":15},"compiling-kernels-ahead-of-time","Compiling kernels ahead of time",{"id":23,"text":24,"level":15},"packaging-and-deploying","Packaging and deploying",{"id":26,"text":27,"level":15},"what-to-check-before-you-ship","What to check before you ship",false,{"title":30,"titleHtml":30,"route":31},"Training: The Backward Kernel","\u002Fkernels\u002Fin-a-model\u002Fbackward",{"title":33,"titleHtml":33,"route":34},"What Is Portable","\u002Fkernels\u002Fin-a-model\u002Fportability",1786271830114]