[{"data":1,"prerenderedAt":19},["ShallowReactive",2],{"chapter:vision-rs\u002Fdeployment\u002Fpackaging.json":3},{"project":4,"route":5,"title":6,"titleHtml":6,"navTitle":6,"part":7,"sourcePath":8,"editUrl":9,"html":10,"toc":11,"hasMermaid":12,"prev":13,"next":16},"vision-rs","\u002Fvision-rs\u002Fdeployment\u002Fpackaging","Packaging a Deployable Bundle","Cross-Compilation & Deployment","deployment\u002Fpackaging.md","https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fvision-rs\u002Fedit\u002Fmain\u002Fbook\u002Fsrc\u002Fdeployment\u002Fpackaging.md","\u003Cp>\u003Ccode>cargo teeny package\u003C\u002Fcode> combines cross-compiling the binary\u002Fexample for the\ntarget board with ahead-of-time-compiling its GPU kernels on the host, into\none self-contained directory you can copy straight to the device — no\n\u003Ccode>teenyc\u003C\u002Fcode>, CUDA toolkit, or Rust install needed on the Jetson itself (see\n\u003Ca href=\"\u002Fvision-rs\u002Fkernels-and-performance\u002Fteenyc-toolchain\">The \u003Ccode>teenyc\u003C\u002Fcode> Toolchain\u003C\u002Fa> for\nwhy AOT compilation is what makes this possible).\u003C\u002Fp>\n\u003Cpre data-lang=\"bash\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#7FB6D9\">cargo\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> package\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> \\\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#B79AD4\">  --target\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> jetson-orin-nano\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> \\\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#B79AD4\">  --example\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> yolo26\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> \\\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#B79AD4\">  --dest\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> .\u002Fdist\u002Fyolo26-orin\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> \\\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#B79AD4\">  --device\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> cuda\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> \\\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#B79AD4\">  --options\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> \"\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\">capability=sm_87,ptx-version=82,sm-count=8\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">\"\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cul>\n\u003Cli>\u003Ccode>--options capability=sm_87\u003C\u002Fcode> is the Jetson Orin Nano’s GPU compute\ncapability (Ampere). \u003Ccode>ptx-version=82\u003C\u002Fcode> overrides \u003Ccode>teenyc\u003C\u002Fcode>’s otherwise\nconservative default PTX ISA floor for \u003Ccode>sm_87\u003C\u002Fcode> — keep this pinned unless\nyour target device is on a materially different CUDA version.\u003C\u002Fli>\n\u003Cli>\u003Ccode>sm-count=8\u003C\u002Fcode> is the Jetson Orin Nano’s actual SM count (1024 CUDA cores \u002F\n128 per SM) — enables shape-adaptive conv kernel tile-size selection, so\ndeep\u002Fsmall-spatial layers pick a smaller tile size (more thread blocks)\ninstead of under-occupying this GPU’s 8 SMs at the default tile size.\nOmit it to keep the previous fixed-tile-size behavior.\u003C\u002Fli>\n\u003Cli>Use \u003Ccode>--bin &lt;name&gt;\u003C\u002Fcode> instead of \u003Ccode>--example &lt;name&gt;\u003C\u002Fcode> when packaging a binary\ncrate rather than an example.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>This produces:\u003C\u002Fp>\n\u003Cpre class=\"code-panel\" data-lang=\"text\">\u003Ccode>dist\u002Fyolo26-orin\u002F\n  bin\u002Fyolo26        # cross-compiled binary\n  cache\u002F            # AOT-compiled GPU kernels\n  conf\u002F             # provenance marker (target\u002Fdevice\u002Foptions\u002Fcommit\u002Fbuild time)\n  data\u002F             # empty — populate with models\u002Fdatasets separately\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>The binary auto-detects \u003Ccode>cache\u002F\u003C\u002Fcode> as its sibling directory at runtime (no\nextra env var needed on the device — see\n\u003Ccode>teeny_compiler::compiler::default_cache_dir()\u003C\u002Fcode>), so it uses the\nprecompiled kernels instead of trying to JIT-compile, which would fail:\nthere’s no \u003Ccode>teenyc\u003C\u002Fcode> on the Jetson.\u003C\u002Fp>\n\u003Cp>Next: \u003Ca href=\"\u002Fvision-rs\u002Fdeployment\u002Fdeploying\">Deploying to the Device\u003C\u002Fa>.\u003C\u002Fp>\n",[],false,{"title":14,"titleHtml":14,"route":15},"Building for Jetson Orin Nano","\u002Fvision-rs\u002Fdeployment\u002Fcross-compilation",{"title":17,"titleHtml":17,"route":18},"Deploying to the Device","\u002Fvision-rs\u002Fdeployment\u002Fdeploying",1786271830570]