[{"data":1,"prerenderedAt":46},["ShallowReactive",2],{"chapter:kernels\u002Ffast\u002Fmeasuring.json":3},{"project":4,"route":5,"title":6,"titleHtml":6,"navTitle":6,"part":7,"sourcePath":8,"editUrl":9,"html":10,"toc":11,"hasMermaid":39,"prev":40,"next":43},"kernels","\u002Fkernels\u002Ffast\u002Fmeasuring","Measuring","Making It Fast","fast\u002Fmeasuring.md","https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fteenygrad\u002Fedit\u002Fmain\u002Fbooks\u002Fkernels\u002Fsrc\u002Ffast\u002Fmeasuring.md","\u003Cp>Every chapter in this part has ended by telling you to measure. This is how, and\nmore importantly, how to get a number you can trust.\u003C\u002Fp>\n\u003Ch2 id=\"why-this-is-harder-than-it-looks\">Why this is harder than it looks\u003C\u002Fh2>\n\u003Cp>Four things will give you a wrong number, and the first two will give you one\nthat is wrong by an order of magnitude.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>The compile happens on the first call.\u003C\u002Fstrong> \u003Ccode>compile_kernel\u003C\u002Fcode> shells out to\n\u003Ccode>teenyc\u003C\u002Fcode> and caches the result. Time that call and you have measured the\ncompiler. Compile outside the timing loop, always.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Launches are asynchronous — but this API already handles it.\u003C\u002Fstrong> \u003Ccode>cuLaunchKernel\u003C\u002Fcode>\nreturns before the kernel has run, so in CUDA generally a timing loop that\nlaunches and stops the clock measures the cost of asking, not of doing.\u003C\u002Fp>\n\u003Cp>\u003Ccode>Device::launch\u003C\u002Fcode> calls \u003Ccode>cuCtxSynchronize\u003C\u002Fcode> immediately after launching, in both\nits typed and arg-packed paths. Its comment says why, and it is not about\ntiming: synchronising there makes a GPU-side fault surface as a CUDA error code\nrather than as a SIGSEGV somewhere later.\u003C\u002Fp>\n\u003Cp>So a loop around \u003Ccode>device.launch\u003C\u002Fcode> measures real kernel time, and you do not need\nyour own barrier. Worth knowing in both directions — it also means you cannot\noverlap launches or hide latency behind the host, because every launch waits.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>The first run is never representative.\u003C\u002Fstrong> Caches are cold, clocks have not\nboosted, memory is not resident. Warm up.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Clocks move.\u003C\u002Fstrong> GPUs throttle when hot and boost when idle. A benchmark run\nimmediately after another one is not measuring the same machine.\u003C\u002Fp>\n\u003Cp>\u003Ccode>criterion\u003C\u002Fcode>, which this tree uses, handles warm-up and repetition and gives you a\ndistribution rather than a single number. It does not handle the first two — the\ncompile and the synchronisation are yours.\u003C\u002Fp>\n\u003Ch2 id=\"the-harness-in-this-tree\">The harness in this tree\u003C\u002Fh2>\n\u003Cp>\u003Ccode>kernels\u002Fteeny-kernels\u002Fbenches\u002Fconv2d_bn_silu.rs\u003C\u002Fcode> is the pattern to copy. Its\nstructure:\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Compile once, outside the loop.\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> kernel \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> Conv2dBnSiluForward\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">new\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">kh\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> kw\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> stride_h\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> stride_w\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> pad_h\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> pad_w\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> BLOCK_OW_SCALAR\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> ptx \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> std\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">fs\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">read\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">compile_kernel\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(&#x26;\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">kernel\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> target\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> false\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">)?)?;\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> program \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> testing\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">load_program_from_ptx\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::&#x3C;\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">Conv2dBnSiluForward\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">>(&#x26;\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">ptx\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">)?;\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Note \u003Ccode>force: false\u003C\u002Fcode>. The cache is wanted here — you are not benchmarking\ncompilation.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Allocate and fill once, outside the loop.\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#FF5F9E\"> mut\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> x_buf \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> device\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">buffer\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::&#x3C;\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">f32\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">>(\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">shape\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">nb \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">*\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> shape\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">c_in \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">*\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> shape\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">hh \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">*\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> shape\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">ww\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">)?;\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#E6E8E3\">x_buf\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">to_device\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(&#x26;\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">shape\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">x_host\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">())?;\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Host-to-device copies are slow and are not what you are measuring.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Only the launch is inside.\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#E6E8E3\">group\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">bench_function\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">format!\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">\"\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\">scalar\u002F\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">{}\"\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> shape\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">label\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">),\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> |\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">b\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">|\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> {\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#E6E8E3\">    b\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">iter\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(||\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> {\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> device\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">launch\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(&#x26;\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">program\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> &#x26;\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">cfg\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> (...))\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> })\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#8A9088\">});\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>\u003Cstrong>Use deterministic inputs.\u003C\u002Fstrong> The bench generates them arithmetically:\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">fn\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\"> x_host\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(&#x26;\u003C\u002Fspan>\u003Cspan style=\"color:#FF5F9E\">self\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">)\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> ->\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> Vec\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">&#x3C;\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">f32\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">>\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> {\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#8A9088\">    (\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">0\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">..\u003C\u002Fspan>\u003Cspan style=\"color:#FF5F9E\">self\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">nb \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">*\u003C\u002Fspan>\u003Cspan style=\"color:#FF5F9E\"> self\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">c_in \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">*\u003C\u002Fspan>\u003Cspan style=\"color:#FF5F9E\"> self\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">hh \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">*\u003C\u002Fspan>\u003Cspan style=\"color:#FF5F9E\"> self\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">ww\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">)\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#8A9088\">        .\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">map\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(|\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">i\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">|\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> (\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">i \u003C\u002Fspan>\u003Cspan style=\"color:#FF5F9E\">as\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> f32\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> %\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 17\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">0\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> -\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 8\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">0\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">)\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> *\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 0\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">)\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#8A9088\">        .\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">collect\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">()\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#8A9088\">}\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Not random. Two runs get identical data, and any difference between them is the\nkernel.\u003C\u002Fp>\n\u003Ch2 id=\"what-to-compare-against\">What to compare against\u003C\u002Fh2>\n\u003Cp>A number alone means nothing. \u003Ccode>142 µs\u003C\u002Fcode> is neither good nor bad.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Against the alternative you would otherwise ship.\u003C\u002Fstrong> The fused kernel against\nthe three unfused ones. This is the comparison that decides whether the work was\nworth it.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Against the other implementations of the same thing.\u003C\u002Fstrong> The conv bench times\nthree kernels across shapes chosen to straddle the dispatch thresholds — so the\nmeasurement answers “does the lowering still pick the right one?”, not just “how\nfast is this?”.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Against the hardware’s limit.\u003C\u002Fstrong> For a memory-bound kernel, divide the bytes\nmoved by the elapsed time and compare with the card’s peak bandwidth. At 80% you\nare close to done. At 15% something is wrong, and Chapter 17 is where to look.\nThis is the most useful single check available, and it needs no reference\nimplementation.\u003C\u002Fp>\n\u003Ch2 id=\"choosing-shapes\">Choosing shapes\u003C\u002Fh2>\n\u003Cp>Benchmark the shapes you run, not round numbers.\u003C\u002Fp>\n\u003Cp>Powers of two are the friendliest case: no ragged tail, no masked lanes, tiles\nthat divide evenly. A kernel benchmarked only at 1024×1024 can be much worse at\n1000×1000, and 1000 is the realistic one.\u003C\u002Fp>\n\u003Cp>The conv bench picks shapes deliberately either side of the dispatch thresholds,\nwhich is the right instinct: benchmark where behaviour \u003Cem>changes\u003C\u002Fem>, not where it\nis comfortable.\u003C\u002Fp>\n\u003Ch2 id=\"running-it\">Running it\u003C\u002Fh2>\n\u003Cpre data-lang=\"bash\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#7FB6D9\">cargo\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> bench\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> -p\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny-kernels\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --features\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> cuda,training\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --bench\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> conv2d_bn_silu\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>On Blackwell you may need the PTX-version workaround from Chapter 4:\u003C\u002Fp>\n\u003Cpre data-lang=\"bash\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#E6E8E3\">TEENYC_PTX_VERSION\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\">87\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\"> cargo\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> bench\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> -p\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny-kernels\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --features\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> cuda,training\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --bench\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> conv2d_bn_silu\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>\u003Ccode>criterion\u003C\u002Fcode> writes an HTML report and, on a second run, compares against the\nprevious one — which makes “did my change help?” a question it answers directly.\u003C\u002Fp>\n\u003Ch2 id=\"recording-a-result\">Recording a result\u003C\u002Fh2>\n\u003Cp>A measurement without its context is not reproducible. Record:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>The card\u003C\u002Fstrong>, by name and compute capability.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>The shapes.\u003C\u002Fstrong>\u003C\u002Fli>\n\u003Cli>\u003Cstrong>The block sizes and tile shapes.\u003C\u002Fstrong>\u003C\u002Fli>\n\u003Cli>\u003Cstrong>The date.\u003C\u002Fstrong> Driver and toolchain versions move.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2 id=\"a-worked-result\">A worked result\u003C\u002Fh2>\n\u003Cp>The conv bench, run on an \u003Cstrong>RTX 5070 (sm_120), CUDA 13.3, driver 610.43.02\u003C\u002Fstrong>.\nCriterion means; the kernel the lowering actually picks for each shape is\nmarked \u003Cstrong>✓\u003C\u002Fstrong>.\u003C\u002Fp>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Shape\u003C\u002Fth>\n\u003Cth>scalar\u003C\u002Fth>\n\u003Cth>tiled\u003C\u002Fth>\n\u003Cth>gemm\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\n\u003Ctr>\n\u003Ctd>1×1, \u003Ccode>c_out\u003C\u002Fcode>=8, 32×32\u003C\u002Ftd>\n\u003Ctd>\u003Cstrong>11.8 µs ✓\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>13.3 µs\u003C\u002Ftd>\n\u003Ctd>12.7 µs\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>1×1, \u003Ccode>c_out\u003C\u002Fcode>=16, 32×32\u003C\u002Ftd>\n\u003Ctd>16.7 µs\u003C\u002Ftd>\n\u003Ctd>\u003Cstrong>13.2 µs ✓\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>12.4 µs\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>1×1, \u003Ccode>c_out\u003C\u002Fcode>=32, 40×40\u003C\u002Ftd>\n\u003Ctd>42.4 µs\u003C\u002Ftd>\n\u003Ctd>14.5 µs\u003C\u002Ftd>\n\u003Ctd>\u003Cstrong>12.8 µs ✓\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>3×3, \u003Ccode>c_out\u003C\u002Fcode>=32, 40×40\u003C\u002Ftd>\n\u003Ctd>871.4 µs\u003C\u002Ftd>\n\u003Ctd>\u003Cstrong>92.0 µs ✓\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>n\u002Fa\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\n\u003C\u002Ftable>\n\u003Cp>Read it as the bench’s author intended — as a check on whether the dispatch\nthresholds from Chapter 12 still hold.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Three of the four are right.\u003C\u002Fstrong> At \u003Ccode>c_out\u003C\u002Fcode>=8 the scalar kernel wins and is\nchosen. At \u003Ccode>c_out\u003C\u002Fcode>=32 the GEMM kernel wins and is chosen, 3.3× faster than\nscalar. On the 3×3 convolution, where GEMM does not apply, tiled beats scalar by\n\u003Cstrong>9.5×\u003C\u002Fstrong> — the single biggest number here, and a good illustration of why the\nnaive kernel from Chapter 11 is not the one you ship.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>One is not.\u003C\u002Fstrong> At \u003Ccode>c_out\u003C\u002Fcode>=16 the lowering picks the tiled kernel at 13.2 µs,\nbut the GEMM kernel does it in 12.4 µs — about 6% faster, with non-overlapping\nconfidence intervals. The GEMM threshold is 32; on this card it could come down\nto 16.\u003C\u002Fp>\n\u003Cp>That is a real finding from one bench run, and it is worth being careful with\nit: 6% on one card at one shape is not enough to move a threshold that has to\nserve every card. What it justifies is measuring the same point on the other\ntargets, which is exactly the argument a hand-picked constant needs to survive.\u003C\u002Fp>\n\u003Ch2 id=\"recording-a-result-1\">Recording a result\u003C\u002Fh2>\n\u003Cp>The table above has everything a measurement needs: the card, the shapes, what\nwas compared, and what was chosen. Add the date when you record one — driver\nand toolchain versions move, and these numbers came from a specific pairing.\u003C\u002Fp>\n\u003Cp>An invented number is worse than a blank one. A blank prompts someone to\nmeasure; an invented number gets quoted.\u003C\u002Fp>\n\u003Ch2 id=\"profiling\">Profiling\u003C\u002Fh2>\n\u003Cp>\u003Ccode>criterion\u003C\u002Fcode> tells you a kernel is slow. It does not tell you why.\u003C\u002Fp>\n\u003Cp>For that you need a profiler — NVIDIA’s Nsight Compute reports achieved\nbandwidth, occupancy, warp stall reasons and instruction mix per kernel, which\nis the level at which “why” gets answered.\u003C\u002Fp>\n\u003Cp>Nothing in this SDK integrates with it, so this book cannot teach reading one.\nIt has been used on these kernels, though, and the result is a good example of\nwhat a profiler tells you that a stopwatch does not:\u003C\u002Fp>\n\u003Cblockquote>\n\u003Cp>Profiling YOLO26n’s conv layers showed occupancy as low as \u003Cstrong>8%\u003C\u002Fstrong> on deep\nlayers with small spatial extent and many output channels. Theoretical\noccupancy was 100% in every case — so it was not register or shared-memory\npressure. It was purely a grid-size problem: the fixed tile size produced as\nfew as 40 thread blocks, which cannot fill the machine no matter how good the\nkernel is.\u003C\u002Fp>\n\u003C\u002Fblockquote>\n\u003Cp>That distinction — achieved 8%, theoretical 100% — is the whole value of the\ntool. A benchmark would have told you the layer was slow. Only the profiler\ntold you it was slow because there was not enough work to go around, which\npoints at the tile size rather than at the kernel body.\u003C\u002Fp>\n\u003Cp>The fix was \u003Ccode>sm_count\u003C\u002Fcode>, described in Chapter 16.\u003C\u002Fp>\n\u003Cp>What you \u003Cem>can\u003C\u002Fem> do without a profiler, and should do first:\u003C\u002Fp>\n\u003Col>\n\u003Cli>\u003Cstrong>Compute achieved bandwidth by hand.\u003C\u002Fstrong> Bytes moved ÷ time. Compare with the\ncard’s specification.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Read the MLIR.\u003C\u002Fstrong> Chapter 9. Count the loads and stores; check nothing is\nloaded twice.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Vary one thing at a time.\u003C\u002Fstrong> Block size, tile shape, layout. The measurement\ntells you which mattered.\u003C\u002Fli>\n\u003C\u002Fol>\n\u003Cp>That covers most of what a first profiler session would have told you.\u003C\u002Fp>\n\u003Cp>Next: the numbers themselves.\u003C\u002Fp>\n",[12,16,19,22,25,28,31,34,36],{"id":13,"text":14,"level":15},"why-this-is-harder-than-it-looks","Why this is harder than it looks",2,{"id":17,"text":18,"level":15},"the-harness-in-this-tree","The harness in this tree",{"id":20,"text":21,"level":15},"what-to-compare-against","What to compare against",{"id":23,"text":24,"level":15},"choosing-shapes","Choosing shapes",{"id":26,"text":27,"level":15},"running-it","Running it",{"id":29,"text":30,"level":15},"recording-a-result","Recording a result",{"id":32,"text":33,"level":15},"a-worked-result","A worked result",{"id":35,"text":30,"level":15},"recording-a-result-1",{"id":37,"text":38,"level":15},"profiling","Profiling",false,{"title":41,"titleHtml":41,"route":42},"Memory Coalescing and Tensor Layout","\u002Fkernels\u002Ffast\u002Flayout",{"title":44,"titleHtml":44,"route":45},"Numerics","\u002Fkernels\u002Ffast\u002Fnumerics",1786271830059]