<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://page-fault.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://page-fault.io/" rel="alternate" type="text/html" /><updated>2026-07-29T11:03:51+00:00</updated><id>https://page-fault.io/feed.xml</id><title type="html">VMExit [Leaving Sandbox Zone…]</title><subtitle>Just another Blog about Low Level Programming: Operating Systems, Security Platform Firmwares, Binary Analysis, Algorithms and other boring stuff...</subtitle><author><name>Maciej Grochowski</name></author><entry><title type="html">tlp-tool and the Next Phase of PCIe Debugging</title><link href="https://page-fault.io/hardware/2026/07/04/tlp-tool-and-the-next-phase-of-pcie-debugging.html" rel="alternate" type="text/html" title="tlp-tool and the Next Phase of PCIe Debugging" /><published>2026-07-04T12:00:00+00:00</published><updated>2026-07-04T12:00:00+00:00</updated><id>https://page-fault.io/hardware/2026/07/04/tlp-tool-and-the-next-phase-of-pcie-debugging</id><content type="html" xml:base="https://page-fault.io/hardware/2026/07/04/tlp-tool-and-the-next-phase-of-pcie-debugging.html"><![CDATA[<blockquote>
  <p><strong>Note:</strong> This piece started as a pitch to LWN. The editors passed on it, so it’s
going up here instead.</p>
</blockquote>

<p>Peripheral Component Interconnect Express (PCIe) failures are often reported in a form that is precise, information-dense, and nearly unreadable: a Transaction Layer Packet (TLP) header dumped as four raw 32-bit words in an Advanced Error Reporting (AER) log. Understanding what those bytes mean requires knowing the TLP format — and the format just changed.</p>

<h2 id="what-a-tlp-header-looks-like">What a TLP header looks like</h2>

<p>When a PCIe device causes an error, the hardware records the offending packet’s header in the AER Header Log register. A typical kernel log or <code class="language-plaintext highlighter-rouge">lspci -vv</code> output shows it as four DWORDs:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>HeaderLog: 00000001 0000220f 01070000 9eece789
</code></pre></div></div>

<p>For PCIe 1.0 through 5.0, the layout of those bytes has been stable. DW0 — the first four bytes — encodes the packet type, traffic class, and payload length. Everything else follows from what DW0 says:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>DW0:  0x00000001
      Byte 0: 0x00 = 0b00000000
        [7:5] FMT   = 000   → 3DW header, no data
        [4:0] TYPE  = 00000 → Memory Read Request
      Byte 1: 0x00  → TC=0, Attr[2]=0
      Byte 2: 0x00  → TD=0, EP=0 (not poisoned), Attr[1:0]=00
      Byte 3: 0x01  → Length = 1 DW (4 bytes)
</code></pre></div></div>

<p>The FMT and TYPE fields together form a matrix that determines both the packet kind and the header length (3DW or 4DW). Once you know the format, you know where every subsequent field lives:</p>

<table>
  <thead>
    <tr>
      <th>DW</th>
      <th>Memory Request</th>
      <th>Config Request</th>
      <th>Completion</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>0</td>
      <td>FMT+TYPE, TC, Length</td>
      <td>FMT+TYPE, TC, Length</td>
      <td>FMT+TYPE, TC, Length</td>
    </tr>
    <tr>
      <td>1</td>
      <td>Requester ID, Tag, Byte Enables</td>
      <td>Requester ID, Tag, Byte Enables</td>
      <td>Completer ID, Status, Byte Count</td>
    </tr>
    <tr>
      <td>2</td>
      <td>Address[31:2] (3DW) or Address[63:32] (4DW)</td>
      <td>Bus/Dev/Func, Ext Reg, Register</td>
      <td>Requester ID, Tag, Lower Address</td>
    </tr>
  </tbody>
</table>

<p>That table is the core of traditional TLP parsing. The same FMT+TYPE matrix, the same field positions, the same bit extraction — unchanged from 2003 to 2019. Sixteen years of stability.</p>

<h2 id="then-flit-mode-happened">Then Flit Mode happened</h2>

<p>PCIe 6.0 introduced Flit Mode to support PAM4 signaling at 64.0 GT/s. Instead of variable-length packets delimited by STP/END framing tokens, TLPs are packed into fixed 256-byte containers called flits, each with its own forward error correction. That is a data-link layer change, but it reached up into the TLP format itself.</p>

<p>DW0 in Flit Mode carries a different set of fields. The FMT+TYPE matrix — the thing that has been the entry point for TLP parsing since PCIe 1.0 — is gone. In its place, there are thirteen type codes in a flat enumeration and a new set of header fields:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Classic DW0:  FMT[7:5] | TYPE[4:0] | TC | Attr | Length
Flit DW0:     Type[4:0] | T9 | TC | OHC-A | TH | Attr | AT | Length
</code></pre></div></div>

<p>TLP Prefixes — the extension mechanism in non-Flit framing — are replaced by Optional Header Components (OHC). The type encoding is not a subset or superset of the old one; it is a different numbering. The same raw DW0 bytes decode to a different packet type depending on whether the link was using Flit framing or not.</p>

<p>A handful of representative DW0 byte 0 values, side by side, makes the divergence concrete:</p>

<table>
  <thead>
    <tr>
      <th>DW0 byte 0</th>
      <th>Non-Flit framing</th>
      <th>Flit Mode framing</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">0x00</code></td>
      <td>Memory Read (3DW)</td>
      <td><strong>NOP</strong></td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">0x03</code></td>
      <td>undefined</td>
      <td>Memory Read (32-bit)</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">0x04</code></td>
      <td>Type 0 Configuration Read</td>
      <td>undefined</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">0x22</code></td>
      <td>undefined</td>
      <td>UIO Memory Read (64-bit)</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">0x40</code></td>
      <td>Memory Write (3DW)</td>
      <td>Memory Write (32-bit)</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">0x4C</code></td>
      <td>FetchAdd AtomicOp</td>
      <td>FetchAdd AtomicOp (32-bit)</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">0x4E</code></td>
      <td>CompareSwap AtomicOp</td>
      <td>CompareSwap AtomicOp (32-bit)</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">0x70</code></td>
      <td>Message with Data → RC</td>
      <td>Message with Data → RC</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">0x8D</code></td>
      <td>Local TLP Prefix</td>
      <td>Local TLP Prefix</td>
    </tr>
  </tbody>
</table>

<p>The first row is the pathological case: a byte that has meant “Memory Read, 3DW header” since 2003 now means “no operation” on a Flit Mode link. The <code class="language-plaintext highlighter-rouge">0x04</code> row is almost as severe — a Type 0 Configuration Read in classic PCIe is not a defined type at all in Flit Mode. Other rows align by design — Memory Write at <code class="language-plaintext highlighter-rouge">0x40</code>, the AtomicOps at <code class="language-plaintext highlighter-rouge">0x4C</code>/<code class="language-plaintext highlighter-rouge">0x4E</code>, the routed Message-with-Data at <code class="language-plaintext highlighter-rouge">0x70</code>, and the TLP Prefix at <code class="language-plaintext highlighter-rouge">0x8D</code> — because the spec writers preserved semantic continuity where they could. But neither the alignments nor the divergences are something the bytes themselves announce; the framing context is what selects the table.</p>

<p>The header does not tell you which interpretation is correct — the framing context does. The worked example in the next section will run <code class="language-plaintext highlighter-rouge">04000001</code> through the tool as a Type 0 Configuration Read in classic mode; on a Flit Mode link, those same bytes would not decode as a defined type at all.</p>

<h2 id="no-one-expects-you-to-memorize-this">No one expects you to memorize this</h2>

<p>The field-level detail in those tables is useful for understanding the problem, but nobody should be decoding TLP headers by hand for every AER event. The PCIe maintainers agree — which is why the kernel’s PCI documentation now points users at <strong><a href="https://github.com/mmpg-x86/tlp-tool">tlp-tool</a></strong>, whose binary is named <code class="language-plaintext highlighter-rouge">rtlp-tool</code>.</p>

<p>On March 23, Lukas Wunner <a href="https://lore.kernel.org/linux-pci/20260323165038.GA830530@bhelgaas/T/#t">posted a patch</a> titled <strong>“Documentation: PCI: Document decoding of TLP Header in AER messages”</strong>, arguing that the TLP Header hints at the root cause of an error but is often ignored because of its seeming opaqueness. Wunner noted that the obvious alternative — wireshark TLP dissectors — expects complete TLPs rather than just headers, and cannot consume the hex format the kernel emits directly; tlp-tool was, in his words, “the most cut and dried solution out there.” Bjorn Helgaas <a href="https://lore.kernel.org/linux-pci/20260323165038.GA830530@bhelgaas/T/#t">applied the patch</a> to <code class="language-plaintext highlighter-rouge">pci/for-linus</code> for the 7.0 merge window the same day, tweaking the commit log on the way in to note that the Header Log lives in the AER Capability — which may be present on any PCIe function, not just root ports. The result is modest but significant: the kernel’s PCIe AER documentation now recommends tlp-tool as a way to decode the TLP Header into human-readable form, and <a href="https://git.kernel.org/torvalds/c/8af4fad545fa4df358c8e4d12f269e460717e514">ships a worked example</a> — piping a kernel commit URL through <code class="language-plaintext highlighter-rouge">rtlp-tool --aer</code> — to make the recommendation directly executable. Wunner’s cover letter also floated going one step further — pointing users at the tool from a <code class="language-plaintext highlighter-rouge">printk_once()</code> message when the first AER error fires — but kept the patch itself to documentation only, calling that “probably sufficient” for now.</p>

<p>The lowest-level entry point is to hand the tool the four header words directly:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>rtlp-tool <span class="nt">-i</span> <span class="s2">"04000001 00200a03 05010000 00050100"</span>
</code></pre></div></div>

<p>That is useful when a header is already in hand. The two modes that matter for real debugging map to where Header Logs actually surface in the kernel: <code class="language-plaintext highlighter-rouge">dmesg</code> AER messages, and the <code class="language-plaintext highlighter-rouge">HeaderLog:</code> line inside <code class="language-plaintext highlighter-rouge">lspci -vv</code> output.</p>

<p><strong><code class="language-plaintext highlighter-rouge">--aer</code>: kernel log input.</strong> A TLP Header copied straight from <code class="language-plaintext highlighter-rouge">dmesg</code> goes through <code class="language-plaintext highlighter-rouge">--aer</code>:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nb">echo</span> <span class="s1">'TLP Header: 04000001 00200a03 05010000 00050100'</span> | rtlp-tool <span class="nt">--aer</span>
</code></pre></div></div>

<p>The output makes the bit extraction explicit:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>+----------+------------------+--------------------+
| TLP Type | ConfType0ReadReq | 3DW no Data Header |
+----------+------------------+--------------------+
+------------+--------+--------+-------+
| Field Name | Offset | Length | Value |
|            | (bits) | (bits) |       |
+------------+--------+--------+-------+
| Fmt        | 0      | 3      | 0     |
| Type       | 3      | 5      | 4     |
| T9         | 8      | 1      | 0     |
| TC         | 9      | 3      | 0     |
| T8         | 12     | 1      | 0     |
| Attr_b2    | 13     | 1      | 0     |
| LN         | 14     | 1      | 0     |
| TH         | 15     | 1      | 0     |
| TD         | 16     | 1      | 0     |
| Ep         | 17     | 1      | 0     |
| Attr       | 18     | 2      | 0     |
| AT         | 20     | 2      | 0     |
| Length     | 22     | 10     | 1     |
+------------+--------+--------+-------+
+-----------------+------------------------------------------------+
| TLP:            | 3DW no Data Header                             |
+-----------------+------------------------------------------------+
| Req ID          | 0x20                                           |
| Tag             | 0xA                                            |
| First DW BE     | 0x3 (bytes 0-1)                                |
| Last DW BE      | 0x0 (none)                                     |
| Target BDF      | 05:00.1                                        |
| Bus             | 0x5                                            |
| Device          | 0x0                                            |
| Function        | 0x1                                            |
| Register Offset | 0x000                                          |
| Register Name   | Vendor ID / Device ID                          |
| Ext Reg Nr      | 0x0                                            |
| Reg Nr          | 0x0                                            |
| Operation       | Read Vendor ID / Device ID register at 05:00.1 |
+-----------------+------------------------------------------------+
</code></pre></div></div>

<p>Three things to read out of that. First, the type: <code class="language-plaintext highlighter-rouge">ConfType0ReadReq</code> — a Configuration Read Request, Type 0 (terminating at the addressed device, not forwarded), 3DW header with no data payload. Second, the requester: ID <code class="language-plaintext highlighter-rouge">0x0020</code>, which decodes as Bus 0, Device 4, Function 0 — a Root Complex device, almost certainly a Root Port issuing the read. Third, the <code class="language-plaintext highlighter-rouge">Operation</code> line spells out the rest directly: a read of the Vendor ID / Device ID register on Function 1 of the target at 05:00.1. tlp-tool did not always name the register; earlier releases stopped at the raw register number and left the lookup to the reader. A register-name table for standard PCI config space now resolves the common cases — Vendor/Device ID, Command, Status, Class Code, and so on — to a name instead of a hex offset.</p>

<p>That target is the interesting part. Probing Function 1 with a config read of register 0 is exactly what the kernel does during PCI enumeration to discover whether a multi-function device exposes a second function. If Function 1 does not exist, the device responds with an Unsupported Request, and the kernel logs an AER event for it. What looks like a scary uncorrectable PCIe error in <code class="language-plaintext highlighter-rouge">dmesg</code> turns out to be a benign enumeration miss. Without the decode, all you would see is the AER summary; with it, you can tell this case apart from a real failure in seconds.</p>

<p><strong><code class="language-plaintext highlighter-rouge">--lspci</code>: device-attributed input.</strong> When the AER Header Log is read directly out of device capabilities, <code class="language-plaintext highlighter-rouge">--lspci</code> parses the entire <code class="language-plaintext highlighter-rouge">lspci -vv</code> block, picks up every non-zero <code class="language-plaintext highlighter-rouge">HeaderLog</code>, and annotates each TLP with the BDF of the device it came from:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>lspci <span class="nt">-vv</span> <span class="nt">-s</span> 05:00.1 | rtlp-tool <span class="nt">--lspci</span>
</code></pre></div></div>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>+----------+------------------------------------------------------------+--------------------+
| TLP Type | ConfType0ReadReq                                           | 3DW no Data Header |
+----------+------------------------------------------------------------+--------------------+
| Source   | 05:00.1 Non-Volatile memory controller: Phison Electronics |                    |
+----------+------------------------------------------------------------+--------------------+
+------------+--------+--------+-------+
| Field Name | Offset | Length | Value |
|            | (bits) | (bits) |       |
+------------+--------+--------+-------+
| Fmt        | 0      | 3      | 0     |
| Type       | 3      | 5      | 4     |
| T9         | 8      | 1      | 0     |
| TC         | 9      | 3      | 0     |
| T8         | 12     | 1      | 0     |
| Attr_b2    | 13     | 1      | 0     |
| LN         | 14     | 1      | 0     |
| TH         | 15     | 1      | 0     |
| TD         | 16     | 1      | 0     |
| Ep         | 17     | 1      | 0     |
| Attr       | 18     | 2      | 0     |
| AT         | 20     | 2      | 0     |
| Length     | 22     | 10     | 1     |
+------------+--------+--------+-------+
+-----------------+------------------------------------------------+
| TLP:            | 3DW no Data Header                             |
+-----------------+------------------------------------------------+
| Req ID          | 0x20                                           |
| Tag             | 0xA                                            |
| First DW BE     | 0x3 (bytes 0-1)                                |
| Last DW BE      | 0x0 (none)                                     |
| Target BDF      | 05:00.1                                        |
| Bus             | 0x5                                            |
| Device          | 0x0                                            |
| Function        | 0x1                                            |
| Register Offset | 0x000                                          |
| Register Name   | Vendor ID / Device ID                          |
| Ext Reg Nr      | 0x0                                            |
| Reg Nr          | 0x0                                            |
| Operation       | Read Vendor ID / Device ID register at 05:00.1 |
+-----------------+------------------------------------------------+
</code></pre></div></div>

<p>The leading <code class="language-plaintext highlighter-rouge">Source</code> row is the practical advantage over <code class="language-plaintext highlighter-rouge">--aer</code>. When several PCIe devices report errors at once — a noisy enclosure, a flaky switch, an enumeration error storm after hotplug — a single <code class="language-plaintext highlighter-rouge">lspci -vv</code> capture aggregates everything, and <code class="language-plaintext highlighter-rouge">--lspci</code> decodes every Header Log and tells you which BDF produced each one. That distinction matters when you are trying to figure out whether one device is producing all the noise or several devices are independently misbehaving.</p>

<p>The two modes also differ in how they pick up Flit framing — and that is what the next section is about.</p>

<h2 id="the-flit-detection-problem">The Flit detection problem</h2>

<p>On March 24, the day after the patch was applied, the linux-pci discussion surfaced a harder question. Maciej Grochowski — the author of this article and the maintainer of tlp-tool — <a href="https://lore.kernel.org/linux-pci/20260323165038.GA830530@bhelgaas/T/#t">pointed out</a> an awkward PCIe 6.0 reality: Flit Mode is mandatory at 64.0 GT/s but supported at all link speeds, so a PCIe 6.x link operating below 64.0 GT/s may still be using Flit framing. You cannot simply check the link speed and assume the framing model — and this is not a hypothetical worry. The same message noted that the framing ambiguity is “already a concern among switch and device vendors working through the transition” to PCIe 6.x.</p>

<p>The raw TLP header bytes do not encode which framing produced them — the same bytes decode differently depending on whether they came from classic PCIe or from Flit Mode. The same-bytes-different-meaning ambiguity from the FMT+TYPE comparison above is no longer hypothetical; it is now a real condition in production logs.</p>

<p>Ilpo Järvinen then <a href="https://lore.kernel.org/linux-pci/20260323165038.GA830530@bhelgaas/T/#t">added the caveat</a> that makes this a real debugging hazard. The Flit Mode Status bit in Link Status 2 is useful only while the link is up — which is not guaranteed in failure scenarios. The kernel tries to compensate by explicitly indicating Flit Mode in its log messages, but the spec creates an asymmetry the kernel cannot fully paper over: AER carries an explicit flag indicating which framing each TLP Log was captured under, while Downstream Port Containment’s (DPC) TLP logging path does not. Järvinen called the DPC omission “botched in the PCIe spec.” To work around it, the kernel saves off Link Status 2 contents and hopes the cached value is still valid when DPC brings the link down — “relatively likely to remain valid,” as he put it, “but fundamentally racy.”</p>

<p>That discussion led directly to tool changes. Grochowski <a href="https://lore.kernel.org/linux-pci/20260323165038.GA830530@bhelgaas/T/#t">replied</a> that tlp-tool would add auto-detection of the kernel’s <code class="language-plaintext highlighter-rouge">(Flit)</code> suffix (introduced in kernel commit <code class="language-plaintext highlighter-rouge">7e077e6707b3</code>, v6.15+) in <code class="language-plaintext highlighter-rouge">--aer</code> mode so that mixed Flit and non-Flit TLPs in the same log could be decoded correctly without forcing a global <code class="language-plaintext highlighter-rouge">--flit</code> switch. The <code class="language-plaintext highlighter-rouge">--lspci</code> mode would read the <code class="language-plaintext highlighter-rouge">Flit+</code> indicator from <code class="language-plaintext highlighter-rouge">LnkSta2</code> per device, while keeping <code class="language-plaintext highlighter-rouge">--flit</code> as a manual override for logs that lack markers. The <a href="https://github.com/mmpg-x86/tlp-tool/releases/tag/v0.5.1">v0.5.1 release notes</a> reflect those changes.</p>

<h2 id="why-this-matters-now">Why this matters now</h2>

<p>The good news first: PCIe 1.0 through 5.0 is stable. The header layout that was true in 2003 is still true on hardware shipping today, and tools that decode it work the same way they always have. If your debugging stays within that envelope, nothing has changed and nothing needs to change.</p>

<p>The harder news is what is now sharing the bus next to it. PCIe 6.0’s Flit Mode is the protocol-level shift this article has been about, but the broader pressure on PCIe is the proliferation of accelerators — GPUs, custom AI silicon, FPGAs, smart NICs — putting more devices, more lanes, and more error volume on every system. Users at the edge of that wave will spend more time staring at AER messages, not less; and the maintainers who triage incoming bug reports will receive more raw <code class="language-plaintext highlighter-rouge">HeaderLog</code> lines from people who do not yet read TLPs by eye.</p>

<p>tlp-tool is useful on both sides of that exchange. For a user filing a bug report, it turns an opaque four-DWORD line in <code class="language-plaintext highlighter-rouge">dmesg</code> into a structured report that names the requester, the target, and the nature of the access — the kind of detail that makes a bug actionable instead of a guess. For a maintainer reading that report, it removes a step: the bit extraction is already done, and the conversation can start at “this is a Type 0 Configuration Read of register 0 on Function 1” instead of “let me decode that for you.” Small per-report, large in aggregate.</p>

<p>The kernel maintainers pointing users at the tool is a quiet acknowledgment that the era of casually reading TLP headers by eye is ending — partly because the protocol made it harder, and partly because the population of people who now need to read them is much larger than it used to be.</p>]]></content><author><name>Maciej Grochowski</name></author><category term="hardware" /><summary type="html"><![CDATA[A kernel 7.0 documentation patch pointed PCIe AER debugging at tlp-tool. The linux-pci discussion that followed surfaced a real Flit Mode debugging hazard -- and the tool changed in response.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://page-fault.io/assets/social-preview.png" /><media:content medium="image" url="https://page-fault.io/assets/social-preview.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Versions, Callgraphs, and other ways to compare what changed</title><link href="https://page-fault.io/reverse-engineering/firmware/mips/2026/07/04/versions-callgraphs-and-other-ways-to-compare-what-changed.html" rel="alternate" type="text/html" title="Versions, Callgraphs, and other ways to compare what changed" /><published>2026-07-04T12:00:00+00:00</published><updated>2026-07-04T12:00:00+00:00</updated><id>https://page-fault.io/reverse-engineering/firmware/mips/2026/07/04/versions-callgraphs-and-other-ways-to-compare-what-changed</id><content type="html" xml:base="https://page-fault.io/reverse-engineering/firmware/mips/2026/07/04/versions-callgraphs-and-other-ways-to-compare-what-changed.html"><![CDATA[<p>In <a href="/reverse-engineering/firmware/mips/2026/05/31/symbols-scripts-and-other-linking-nightmares.html">Part 4</a>, we tried to recompile 191 decompiled C files and hit 2,896 unresolved symbols. We learned why firmware linking is fundamentally different from userspace linking, and concluded that binary patching – the four-byte NOP from Part 3 – was the pragmatic fix.</p>

<p>But there’s a question we haven’t asked: <strong>what did the vendor do?</strong></p>

<p>If there’s a newer firmware version, comparing old and new tells you exactly what changed – and whether the vendor agreed with your diagnosis. In this post, we’ll put two firmware versions side by side, decompile the same function from each, and see how a real fix differs from a four-byte patch. Then we’ll turn to a second technique: using call frequency statistics to automatically identify core utility functions in firmware where every name starts with <code class="language-plaintext highlighter-rouge">FUN_</code>.</p>

<hr />

<h2 id="two-binaries-one-function">Two Binaries, One Function</h2>

<p>Suppose the vendor ships a firmware update. We now have two flat binaries:</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">$ ls -la firmware*.bin
-rw-r--r-- 1 user user  8496 Apr 25 firmware.bin       # v2.4.1-rc3 (buggy)
-rw-r--r-- 1 user user  8712 Apr 25 firmware_v2.bin    # v3.0.0 (fixed)</code></pre></figure>

<p>The new version is 216 bytes larger. That is suspicious – a NOP patch does not add bytes. Whatever changed, it was not a one-instruction fix.</p>

<p>The first thing we do with any new firmware version is the same thing we did in Part 2: <code class="language-plaintext highlighter-rouge">strings</code>. But this time, we diff the output:</p>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash"><span class="nv">$ </span>diff &lt;<span class="o">(</span>strings firmware.bin | <span class="nb">sort</span><span class="o">)</span> &lt;<span class="o">(</span>strings firmware_v2.bin | <span class="nb">sort</span><span class="o">)</span></code></pre></figure>

<figure class="highlight"><pre><code class="language-text" data-lang="text">&lt; bridge routes deferred to bridge handler
---
&gt; bridge
&gt; bridge entries registered
&gt; forwarding
&gt; forwarding table updated
&gt; registering route modules
&lt; route: bridge flag set, skipping
&gt; route: registering bridge entry
&lt; status
&lt; uptime
&lt; v2.4.1-rc3
---
&gt; v3.0.0</code></pre></figure>

<p>Several things jump out:</p>

<ol>
  <li><strong>Version changed</strong>: <code class="language-plaintext highlighter-rouge">v2.4.1-rc3</code> to <code class="language-plaintext highlighter-rouge">v3.0.0</code> – a major version bump, not a point release</li>
  <li><strong>New strings</strong>: “forwarding”, “bridge”, “registering route modules” – these look like module names</li>
  <li><strong>Changed messages</strong>: “route: bridge flag set, skipping” became “route: registering bridge entry” – the bridge path is not skipping routes anymore</li>
  <li><strong>Removed string</strong>: “bridge routes deferred to bridge handler” is gone – the old deferral logic was removed entirely</li>
</ol>

<p>And two lines that look like removed commands but are not: <code class="language-plaintext highlighter-rouge">status</code> and <code class="language-plaintext highlighter-rouge">uptime</code> disappear from the v3.0.0 output entirely. Did the vendor remove the <code class="language-plaintext highlighter-rouge">status</code> and <code class="language-plaintext highlighter-rouge">uptime</code> CLI commands? No – <code class="language-plaintext highlighter-rouge">nm firmware_v2.elf</code> still lists both handlers. What disappeared is the standalone string literal. In v2.4.1-rc3, the command name <code class="language-plaintext highlighter-rouge">"status"</code> happens to be a suffix of the help text <code class="language-plaintext highlighter-rouge">"Show system status"</code>, and <code class="language-plaintext highlighter-rouge">"uptime"</code> a suffix of <code class="language-plaintext highlighter-rouge">"Show system uptime"</code>. The linker’s mergeable string sections do suffix merging: if one string is a tail of another, the shorter one gets folded into the longer one’s bytes instead of getting its own copy. In v3.0.0 that merge just happened to trigger for these two commands, so <code class="language-plaintext highlighter-rouge">strings</code> – which only reports null-terminated runs – no longer sees <code class="language-plaintext highlighter-rouge">"status"</code> or <code class="language-plaintext highlighter-rouge">"uptime"</code> as independent strings. It is a linker artifact of that specific build, not a vendor change. Cross-version string diffs will occasionally show noise like this; the fix is to check <code class="language-plaintext highlighter-rouge">nm</code> before concluding a command was removed.</p>

<p>The string diff alone tells a story: the route engine was restructured from “check flags, skip bridge” to “register modules, dispatch independently.” Let’s confirm this by decompiling both versions.</p>

<hr />

<h2 id="side-by-side-decompilation">Side-by-Side Decompilation</h2>

<p>Here’s <code class="language-plaintext highlighter-rouge">iterate_active_routes</code> from the old firmware (v2.4.1-rc3). We saw this in Part 3 – the function that contains our priority inversion bug:</p>

<figure class="highlight"><pre><code class="language-c" data-lang="c"><span class="cm">/* v2.4.1-rc3 — iterate_active_routes() — flag-checking loop */</span>
<span class="kt">void</span> <span class="nf">iterate_active_routes</span><span class="p">(</span><span class="kt">int</span> <span class="n">stack_idx</span><span class="p">)</span>
<span class="p">{</span>
    <span class="kt">int</span> <span class="n">i</span><span class="p">;</span>
    <span class="kt">int</span> <span class="n">programmed</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span>
    <span class="kt">int</span> <span class="n">skipped</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span>

    <span class="k">for</span> <span class="p">(</span><span class="n">i</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="n">i</span> <span class="o">&lt;</span> <span class="n">MAX_ROUTES</span><span class="p">;</span> <span class="n">i</span><span class="o">++</span><span class="p">)</span> <span class="p">{</span>
        <span class="n">route_entry_t</span> <span class="o">*</span><span class="n">entry</span> <span class="o">=</span> <span class="o">&amp;</span><span class="n">g_route_table_entries</span><span class="p">[</span><span class="n">i</span><span class="p">];</span>

        <span class="k">if</span> <span class="p">(</span><span class="n">entry</span><span class="o">-&gt;</span><span class="n">flags</span> <span class="o">==</span> <span class="mi">0</span><span class="p">)</span>
            <span class="k">continue</span><span class="p">;</span>

        <span class="k">if</span> <span class="p">(</span><span class="n">entry</span><span class="o">-&gt;</span><span class="n">flags</span> <span class="o">&amp;</span> <span class="n">ROUTE_FLAG_BRIDGE</span><span class="p">)</span> <span class="p">{</span>        <span class="cm">/* ← checks bridge FIRST */</span>
            <span class="n">log_msg</span><span class="p">(</span><span class="n">MOD_ROUTE</span><span class="p">,</span> <span class="s">"route: bridge flag set, skipping"</span><span class="p">);</span>
            <span class="n">skipped</span><span class="o">++</span><span class="p">;</span>
            <span class="k">continue</span><span class="p">;</span>                                    <span class="cm">/* ← never reaches active check! */</span>
        <span class="p">}</span>

        <span class="k">if</span> <span class="p">(</span><span class="n">entry</span><span class="o">-&gt;</span><span class="n">flags</span> <span class="o">&amp;</span> <span class="n">ROUTE_FLAG_ACTIVE</span><span class="p">)</span> <span class="p">{</span>
            <span class="n">log_msg</span><span class="p">(</span><span class="n">MOD_ROUTE</span><span class="p">,</span> <span class="s">"route: programming forwarding entry"</span><span class="p">);</span>
            <span class="n">programmed</span><span class="o">++</span><span class="p">;</span>
        <span class="p">}</span>
    <span class="p">}</span>

    <span class="k">if</span> <span class="p">(</span><span class="n">programmed</span> <span class="o">==</span> <span class="mi">0</span><span class="p">)</span>
        <span class="n">log_msg</span><span class="p">(</span><span class="n">MOD_ROUTE</span><span class="p">,</span> <span class="s">"WARNING: no routes programmed for stack"</span><span class="p">);</span>
    <span class="k">if</span> <span class="p">(</span><span class="n">skipped</span> <span class="o">&gt;</span> <span class="mi">0</span><span class="p">)</span>
        <span class="n">log_msg</span><span class="p">(</span><span class="n">MOD_ROUTE</span><span class="p">,</span> <span class="s">"bridge routes deferred to bridge handler"</span><span class="p">);</span>
<span class="p">}</span></code></pre></figure>

<p>The structure is a single loop with inline flag checks. The <code class="language-plaintext highlighter-rouge">if/else</code> ordering means BRIDGE (0x04) is checked before ACTIVE (0x01), and the <code class="language-plaintext highlighter-rouge">continue</code> means routes with both flags never reach the active check.</p>

<p>Now here’s the same function from v3.0.0:</p>

<figure class="highlight"><pre><code class="language-c" data-lang="c"><span class="cm">/* v3.0.0 — iterate_active_routes() — table-driven dispatch */</span>

<span class="k">typedef</span> <span class="nf">void</span> <span class="p">(</span><span class="o">*</span><span class="n">route_handler_t</span><span class="p">)(</span><span class="n">route_entry_t</span> <span class="o">*</span><span class="n">entry</span><span class="p">,</span> <span class="kt">int</span> <span class="o">*</span><span class="n">count</span><span class="p">);</span>

<span class="k">typedef</span> <span class="k">struct</span> <span class="p">{</span>
    <span class="k">const</span> <span class="kt">char</span> <span class="o">*</span><span class="n">name</span><span class="p">;</span>
    <span class="n">route_handler_t</span> <span class="n">handler</span><span class="p">;</span>
    <span class="kt">uint32_t</span> <span class="n">route_mask</span><span class="p">;</span>
<span class="p">}</span> <span class="n">route_module_t</span><span class="p">;</span>

<span class="k">static</span> <span class="kt">void</span> <span class="nf">handle_forwarding_routes</span><span class="p">(</span><span class="n">route_entry_t</span> <span class="o">*</span><span class="n">entry</span><span class="p">,</span> <span class="kt">int</span> <span class="o">*</span><span class="n">count</span><span class="p">)</span> <span class="p">{</span>
    <span class="k">if</span> <span class="p">(</span><span class="n">entry</span><span class="o">-&gt;</span><span class="n">flags</span> <span class="o">&amp;</span> <span class="n">ROUTE_FLAG_ACTIVE</span><span class="p">)</span> <span class="p">{</span>
        <span class="n">log_msg</span><span class="p">(</span><span class="n">MOD_ROUTE</span><span class="p">,</span> <span class="s">"route: programming forwarding entry"</span><span class="p">);</span>
        <span class="p">(</span><span class="o">*</span><span class="n">count</span><span class="p">)</span><span class="o">++</span><span class="p">;</span>
    <span class="p">}</span>
<span class="p">}</span>

<span class="k">static</span> <span class="kt">void</span> <span class="nf">handle_bridge_routes</span><span class="p">(</span><span class="n">route_entry_t</span> <span class="o">*</span><span class="n">entry</span><span class="p">,</span> <span class="kt">int</span> <span class="o">*</span><span class="n">count</span><span class="p">)</span> <span class="p">{</span>
    <span class="k">if</span> <span class="p">(</span><span class="n">entry</span><span class="o">-&gt;</span><span class="n">flags</span> <span class="o">&amp;</span> <span class="n">ROUTE_FLAG_BRIDGE</span><span class="p">)</span> <span class="p">{</span>
        <span class="n">log_msg</span><span class="p">(</span><span class="n">MOD_ROUTE</span><span class="p">,</span> <span class="s">"route: registering bridge entry"</span><span class="p">);</span>
        <span class="p">(</span><span class="o">*</span><span class="n">count</span><span class="p">)</span><span class="o">++</span><span class="p">;</span>
    <span class="p">}</span>
<span class="p">}</span>

<span class="k">const</span> <span class="n">route_module_t</span> <span class="n">g_route_modules</span><span class="p">[]</span> <span class="o">=</span> <span class="p">{</span>
    <span class="p">{</span> <span class="s">"forwarding"</span><span class="p">,</span> <span class="n">handle_forwarding_routes</span><span class="p">,</span> <span class="n">ROUTE_FLAG_ACTIVE</span> <span class="p">},</span>
    <span class="p">{</span> <span class="s">"bridge"</span><span class="p">,</span>     <span class="n">handle_bridge_routes</span><span class="p">,</span>     <span class="n">ROUTE_FLAG_BRIDGE</span> <span class="p">},</span>
<span class="p">};</span>

<span class="kt">void</span> <span class="nf">iterate_active_routes</span><span class="p">(</span><span class="kt">int</span> <span class="n">stack_idx</span><span class="p">)</span>
<span class="p">{</span>
    <span class="kt">int</span> <span class="n">i</span><span class="p">,</span> <span class="n">m</span><span class="p">;</span>
    <span class="kt">int</span> <span class="n">counts</span><span class="p">[</span><span class="n">NUM_ROUTE_MODULES</span><span class="p">];</span>

    <span class="k">for</span> <span class="p">(</span><span class="n">m</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="n">m</span> <span class="o">&lt;</span> <span class="n">NUM_ROUTE_MODULES</span><span class="p">;</span> <span class="n">m</span><span class="o">++</span><span class="p">)</span>
        <span class="n">counts</span><span class="p">[</span><span class="n">m</span><span class="p">]</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span>

    <span class="k">for</span> <span class="p">(</span><span class="n">i</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="n">i</span> <span class="o">&lt;</span> <span class="n">MAX_ROUTES</span><span class="p">;</span> <span class="n">i</span><span class="o">++</span><span class="p">)</span> <span class="p">{</span>
        <span class="n">route_entry_t</span> <span class="o">*</span><span class="n">entry</span> <span class="o">=</span> <span class="o">&amp;</span><span class="n">g_route_table_entries</span><span class="p">[</span><span class="n">i</span><span class="p">];</span>

        <span class="k">if</span> <span class="p">(</span><span class="n">entry</span><span class="o">-&gt;</span><span class="n">flags</span> <span class="o">==</span> <span class="mi">0</span><span class="p">)</span>
            <span class="k">continue</span><span class="p">;</span>

        <span class="cm">/* Dispatch to EVERY registered module — no flag priority */</span>
        <span class="k">for</span> <span class="p">(</span><span class="n">m</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="n">m</span> <span class="o">&lt;</span> <span class="n">NUM_ROUTE_MODULES</span><span class="p">;</span> <span class="n">m</span><span class="o">++</span><span class="p">)</span> <span class="p">{</span>
            <span class="n">g_route_modules</span><span class="p">[</span><span class="n">m</span><span class="p">].</span><span class="n">handler</span><span class="p">(</span><span class="n">entry</span><span class="p">,</span> <span class="o">&amp;</span><span class="n">counts</span><span class="p">[</span><span class="n">m</span><span class="p">]);</span>
        <span class="p">}</span>
    <span class="p">}</span>

    <span class="k">if</span> <span class="p">(</span><span class="n">counts</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span> <span class="o">==</span> <span class="mi">0</span><span class="p">)</span>
        <span class="n">log_msg</span><span class="p">(</span><span class="n">MOD_ROUTE</span><span class="p">,</span> <span class="s">"WARNING: no routes programmed for stack"</span><span class="p">);</span>
    <span class="k">if</span> <span class="p">(</span><span class="n">counts</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span> <span class="o">&gt;</span> <span class="mi">0</span><span class="p">)</span>
        <span class="n">log_msg</span><span class="p">(</span><span class="n">MOD_ROUTE</span><span class="p">,</span> <span class="s">"forwarding table updated"</span><span class="p">);</span>
    <span class="k">if</span> <span class="p">(</span><span class="n">counts</span><span class="p">[</span><span class="mi">1</span><span class="p">]</span> <span class="o">&gt;</span> <span class="mi">0</span><span class="p">)</span>
        <span class="n">log_msg</span><span class="p">(</span><span class="n">MOD_ROUTE</span><span class="p">,</span> <span class="s">"bridge entries registered"</span><span class="p">);</span>
<span class="p">}</span></code></pre></figure>

<p>This is a <strong>complete architectural rewrite</strong>. The function no longer checks flags inline. Instead:</p>

<ol>
  <li>Route handlers are registered in a <strong>module table</strong> (array of <code class="language-plaintext highlighter-rouge">route_module_t</code>)</li>
  <li>Each handler is a <strong>separate function</strong> that decides independently whether to process an entry</li>
  <li>The dispatch loop calls <strong>every</strong> handler for <strong>every</strong> non-empty entry</li>
  <li>A route with both ACTIVE and BRIDGE flags gets processed by <strong>both</strong> handlers</li>
</ol>

<p>There’s no <code class="language-plaintext highlighter-rouge">if/else</code> between the bridge check and the active check. There’s no <code class="language-plaintext highlighter-rouge">continue</code> that skips one path. The two handlers run independently, in separate functions, dispatched from a data table. The priority inversion that existed in v2.4.1-rc3 is <strong>structurally impossible</strong> in v3.0.0.</p>

<hr />

<h2 id="what-changed-and-why-it-matters">What Changed and Why It Matters</h2>

<p>Let’s make the comparison precise:</p>

<table>
  <thead>
    <tr>
      <th>Aspect</th>
      <th>v2.4.1-rc3 (buggy)</th>
      <th>v3.0.0 (fixed)</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><strong>Algorithm</strong></td>
      <td>Single loop with inline flag checks</td>
      <td>Table-driven module dispatch</td>
    </tr>
    <tr>
      <td><strong>Flag handling</strong></td>
      <td><code class="language-plaintext highlighter-rouge">if/else</code> chain: BRIDGE before ACTIVE</td>
      <td>Each handler checks its own flag independently</td>
    </tr>
    <tr>
      <td><strong>Bridge + Active route</strong></td>
      <td>Bridge check fires, <code class="language-plaintext highlighter-rouge">continue</code> skips active</td>
      <td>Both handlers run, both succeed</td>
    </tr>
    <tr>
      <td><strong>Data structures</strong></td>
      <td>None (flags checked inline)</td>
      <td><code class="language-plaintext highlighter-rouge">route_module_t</code> table with function pointers</td>
    </tr>
    <tr>
      <td><strong>Handler isolation</strong></td>
      <td>None (all logic in one function)</td>
      <td>Each handler is a separate function</td>
    </tr>
    <tr>
      <td><strong>New route types</strong></td>
      <td>Requires adding <code class="language-plaintext highlighter-rouge">else if</code> branch</td>
      <td>Requires adding table entry</td>
    </tr>
  </tbody>
</table>

<p>The v2.4.1-rc3 bug was in the <em>ordering</em> of two <code class="language-plaintext highlighter-rouge">if</code> statements inside one function. Our NOP patch in Part 3 fixed the <em>symptom</em> by disabling the bridge branch. The vendor fixed the <em>class of bug</em> by making the ordering irrelevant.</p>

<p>This is a common pattern in firmware evolution. The first version implements something quickly with inline checks. A bug is found. The fix isn’t “swap two lines” – it’s “restructure the function so that the bug category can’t exist.” If you see a major version bump and a function that went from 20 lines to 60 lines, this is probably what happened.</p>

<p>For a reverse engineer, cross-version comparison gives you three things:</p>

<ol>
  <li><strong>Confirmation</strong>: The vendor changed <code class="language-plaintext highlighter-rouge">iterate_active_routes</code>, confirming our bug diagnosis from Part 3</li>
  <li><strong>Understanding</strong>: The vendor didn’t just fix the symptom – they rewrote the dispatch architecture</li>
  <li><strong>Future-proofing</strong>: When analyzing the new version, we know to look for the module table instead of inline flag checks</li>
</ol>

<p>The general technique: when you have two firmware versions and a known bug location, decompile the same function from both, and diff. The changes tell you exactly what the vendor considered broken and how they chose to fix it.</p>

<hr />

<h2 id="call-graph-analysis-naming-the-unnamed">Call Graph Analysis: Naming the Unnamed</h2>

<p>Let’s shift to a different problem. In Part 2, we found 35 functions by scanning for MIPS prologues. In Part 4, we saw what Ghidra produces when it decompiles them: <code class="language-plaintext highlighter-rouge">port_validate</code> renamed to <code class="language-plaintext highlighter-rouge">FUN_80000c84</code>, calling a still-unidentified helper, <code class="language-plaintext highlighter-rouge">FUN_800002e4</code>, twice along the way. Meaningful names replaced by addresses.</p>

<p>If you have debug symbols, <code class="language-plaintext highlighter-rouge">nm</code> gives you every name instantly. But vendor firmware ships stripped. You’re staring at 35 functions (or 4,426 in the real project) and need to figure out which is which.</p>

<p>Here’s the insight: <strong>not all functions are equal</strong>. Some functions are called once. Others are called from everywhere. The ones called from everywhere are utilities – and their call frequency alone tells you what they are.</p>

<p>Let’s count how many times each function is called as a JAL target in our firmware:</p>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash"><span class="nv">$ </span>mipsel-linux-gnu-objdump <span class="nt">-d</span> firmware.elf <span class="se">\</span>
    | <span class="nb">grep</span> <span class="nt">-P</span> <span class="s1">'\tjal\t'</span> <span class="se">\</span>
    | <span class="nb">sed</span> <span class="s1">'s/.*jal\s*[0-9a-f]*//'</span> <span class="se">\</span>
    | <span class="nb">sort</span> | <span class="nb">uniq</span> <span class="nt">-c</span> | <span class="nb">sort</span> <span class="nt">-rn</span> | <span class="nb">head</span> <span class="nt">-10</span></code></pre></figure>

<figure class="highlight"><pre><code class="language-text" data-lang="text">    111 FUN_800002e4
     15 FUN_80000790
     15 FUN_80000530
     12 FUN_80000b74
      6 FUN_800002c4
      4 FUN_80000fec
      4 FUN_80000aa8
      2 FUN_80000d00
      2 FUN_80000c84
      2 FUN_80000280</code></pre></figure>

<p>(I’ve replaced the actual symbol names with Ghidra-style <code class="language-plaintext highlighter-rouge">FUN_</code> addresses – in a real stripped binary, this is what you’d see. Note the <code class="language-plaintext highlighter-rouge">\tjal\t</code> filter: a plain <code class="language-plaintext highlighter-rouge">grep 'jal'</code> also matches <code class="language-plaintext highlighter-rouge">jalr</code>, the indirect-call-through-register instruction, and pollutes the count with a bogus entry. That distinction matters beyond cleanliness – v3.0.0’s table-driven dispatch calls each route handler through a function pointer, which compiles to <code class="language-plaintext highlighter-rouge">jalr</code>, not <code class="language-plaintext highlighter-rouge">jal</code>. A JAL-only call count is structurally blind to every call the new dispatch table makes; that dispatch traffic simply will not show up in this kind of analysis, no matter how the grep is tuned.)</p>

<p>The distribution is wildly uneven. <code class="language-plaintext highlighter-rouge">FUN_800002e4</code> is called <strong>111 times</strong> – more than all other functions combined. That’s not business logic. That’s infrastructure. Let’s figure out what it is.</p>

<hr />

<h2 id="from-counts-to-names">From Counts to Names</h2>

<p><strong>FUN_800002e4 (111 calls)</strong>: Called from nearly every function in the binary. It takes a single pointer argument. In the string cross-reference analysis from Part 2, every call to this function passes a string address. A function called 111 times that takes a string? That’s <code class="language-plaintext highlighter-rouge">uart_puts</code> – the serial output function.</p>

<p><strong>FUN_80000790 (15 calls)</strong> and <strong>FUN_80000530 (15 calls)</strong>: Both called frequently, both appear in the same functions as <code class="language-plaintext highlighter-rouge">uart_puts</code>. One takes a 32-bit integer and outputs hex characters (look for the <code class="language-plaintext highlighter-rouge">0123456789ABCDEF</code> string reference). The other outputs decimal digits. These are <code class="language-plaintext highlighter-rouge">uart_puthex</code> and <code class="language-plaintext highlighter-rouge">uart_putdec</code>.</p>

<p><strong>FUN_80000b74 (12 calls)</strong>: Called by every subsystem initialization function. Takes two arguments: an integer (always a small constant like 1, 2, 3, 4) and a string pointer. The strings are log messages (“system starting”, “initializing ports”). This is <code class="language-plaintext highlighter-rouge">log_msg</code> – the logging function with a module ID and message string.</p>

<p><strong>FUN_80000aa8 (4 calls)</strong>: Takes three arguments: an integer, a string that ends in <code class="language-plaintext highlighter-rouge">.c</code>, and another integer. The strings are source filenames: <code class="language-plaintext highlighter-rouge">route.c</code>, <code class="language-plaintext highlighter-rouge">flash.c</code>, <code class="language-plaintext highlighter-rouge">diag.c</code>. A function that takes <code class="language-plaintext highlighter-rouge">(code, "file.c", line_number)</code> is an assert handler. This is <code class="language-plaintext highlighter-rouge">fw_assert</code>.</p>

<p>Five functions identified from call counts alone:</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">Call Frequency Analysis — Top Functions
========================================

Address       Calls  Identification      Evidence
──────────────────────────────────────────────────────────
FUN_800002e4    111  uart_puts           Every function, takes string arg
FUN_80000790     15  uart_putdec         Takes int, outputs decimal digits
FUN_80000530     15  uart_puthex         Takes int, refs "0123456789ABCDEF"
FUN_80000b74     12  log_msg             Takes (module_id, string), called by init fns
FUN_80000aa8      4  fw_assert           Takes (code, "file.c", line), halts on fatal
FUN_800002c4      6  uart_putchar        Takes single char, called by uart_puts</code></pre></figure>

<p>In a 35-function firmware, we just named 6 functions (17%) from call frequency and string cross-references. Those 6 functions account for <strong>163 of 184 total calls</strong> (89%). Name the infrastructure and you understand most of the binary’s call graph.</p>

<hr />

<h2 id="the-heuristics-at-scale">The Heuristics at Scale</h2>

<p>In our sample firmware with 35 functions, this is a nice exercise. In real vendor firmware with thousands of functions, it’s essential. The real project had 4,426 functions. Call graph analysis identified:</p>

<ul>
  <li><strong>fw_assert</strong> at the top with <strong>497 inbound calls</strong> – unmistakable as the assert handler</li>
  <li><strong>log_msg</strong> with <strong>470 callers</strong> – the logging function with severity and module ID</li>
  <li><strong>irq_disable/irq_enable</strong> with <strong>345/277 callers</strong> – interrupt lock pairs (always appear together)</li>
  <li><strong>mem_alloc</strong> with <strong>330 callers</strong> – the firmware’s memory allocator</li>
  <li><strong>spinlock_acquire/release</strong> with <strong>240/204 callers</strong> – synchronization primitives</li>
</ul>

<p>The heuristics that work at scale:</p>

<ol>
  <li><strong>Highest call count + string argument</strong> → output function (print, log, puts)</li>
  <li><strong>Called with <code class="language-plaintext highlighter-rouge">"file.c"</code> + integer</strong> → assert or error handler</li>
  <li><strong>Always called in pairs</strong> (enable/disable, acquire/release) → lock/unlock primitives</li>
  <li><strong>Called at the start of many functions</strong> → initialization or lock acquisition</li>
  <li><strong>Called at the end of many functions</strong> → cleanup or lock release</li>
</ol>

<p>Once you’ve named the top 10-15 utility functions, Ghidra’s decompiled output transforms. Instead of:</p>

<figure class="highlight"><pre><code class="language-c" data-lang="c"><span class="n">FUN_c0060934</span><span class="p">(</span><span class="mh">0x8000</span><span class="p">,</span> <span class="s">"psc_port.c"</span><span class="p">,</span> <span class="mi">342</span><span class="p">);</span>
<span class="n">FUN_c0098e38</span><span class="p">(</span><span class="mi">1</span><span class="p">,</span> <span class="mh">0x53</span><span class="p">,</span> <span class="mi">3</span><span class="p">,</span> <span class="mi">0</span><span class="p">,</span> <span class="mi">0</span><span class="p">,</span> <span class="mi">0</span><span class="p">,</span> <span class="mi">0</span><span class="p">);</span>
<span class="n">FUN_c01714e0</span><span class="p">();</span></code></pre></figure>

<p>You read:</p>

<figure class="highlight"><pre><code class="language-c" data-lang="c"><span class="n">fw_assert</span><span class="p">(</span><span class="n">FATAL</span><span class="p">,</span> <span class="s">"route.c"</span><span class="p">,</span> <span class="mi">342</span><span class="p">);</span>
<span class="n">log_msg</span><span class="p">(</span><span class="n">ERROR</span><span class="p">,</span> <span class="n">MOD_ROUTE</span><span class="p">,</span> <span class="n">MSG_BRIDGE_FLAG</span><span class="p">);</span>
<span class="n">irq_disable</span><span class="p">();</span></code></pre></figure>

<p>The code hasn’t changed. Your understanding of it has.</p>

<hr />

<h2 id="the-full-picture">The Full Picture</h2>

<p>Over five posts, we’ve built a complete firmware reverse engineering toolkit:</p>

<ol>
  <li><strong>Part 1</strong>: Understood the flat binary format – no headers, no symbols, raw instructions at byte zero</li>
  <li><strong>Part 2</strong>: Found strings, traced cross-references, discovered hidden CLI commands</li>
  <li><strong>Part 3</strong>: Decompiled the route engine, found the priority inversion bug, patched four bytes</li>
  <li><strong>Part 4</strong>: Tried to recompile, hit 2,896 unresolved symbols, learned why firmware linking is hard</li>
  <li><strong>Part 5</strong>: Compared two firmware versions to understand architectural fixes, used call graphs to name functions automatically</li>
</ol>

<p>Every technique came from a real project. Every pattern is something we found in production firmware. The domain was changed (network gateway instead of the actual device), but the methods are universal. Cross-version comparison works on any firmware with multiple releases. Call graph analysis works on any binary with function calls.</p>

<p>The firmware still doesn’t want you to read it. But now you have five ways to read it anyway.</p>

<hr />

<h2 id="whats-next-a-new-series">What’s Next: A New Series</h2>

<p>This series used an 8 KB sample firmware to keep things approachable. But real firmware is bigger, more complex, and full of patterns we haven’t touched yet – cryptographic algorithm identification from lookup tables, struct recovery from raw memory offsets, hardware peripheral mapping from MMIO access patterns, and the kind of runtime architecture (RTOS schedulers, pool allocators, rule engines) that makes embedded devices work.</p>

<p>For the next series – <em>“Crypto Tables, Struct Ghosts, and Other Secrets Hidden in Firmware”</em> – we’ve built a <strong>38 KB firmware</strong> with AES-128 encryption, SHA-256 hashing, a cooperative RTOS, a three-mode pool allocator, and seven MMIO hardware blocks. We’ll also finally introduce <strong>radare2</strong> and compare it head-to-head against Ghidra and Capstone.</p>

<p>If you’ve followed along this far, you have the foundation. The next series goes deeper.</p>

<hr />

<h2 id="references">References</h2>

<ul>
  <li><a href="https://ghidra-sre.org/">Ghidra</a> – Decompiler used for cross-version comparison</li>
  <li><a href="https://www.zynamics.com/bindiff.html">BinDiff</a> – Google/Zynamics tool for binary diffing (automates what we did manually)</li>
  <li><a href="https://github.com/joxeankoret/diaphora">Diaphora</a> – Open-source Ghidra/IDA binary diffing plugin</li>
  <li><a href="https://www.capstone-engine.org/">Capstone Engine</a> – Disassembly framework used for call graph extraction</li>
  <li><a href="https://s3-eu-west-1.amazonaws.com/downloads-mips/documents/MD00086-2B-MIPS32BIS-AFP-6.06.pdf">MIPS Instruction Reference</a> – JAL encoding for call target extraction</li>
  <li><a href="/reverse-engineering/firmware/mips/2026/04/22/headers-symbols-and-other-things-you-wont-find.html">Part 1: Headers, Symbols, and other things you won’t find</a></li>
  <li><a href="/reverse-engineering/firmware/mips/2026/04/26/opcodes-prologues-and-other-hidden-patterns.html">Part 2: Opcodes, Prologues, and other hidden patterns</a></li>
  <li><a href="/reverse-engineering/firmware/mips/2026/05/03/decompilers-annotations-and-other-ways-to-read-the-unreadable.html">Part 3: Decompilers, Annotations, and other ways to read the unreadable</a></li>
  <li><a href="/reverse-engineering/firmware/mips/2026/05/31/symbols-scripts-and-other-linking-nightmares.html">Part 4: Symbols, Scripts, and other linking nightmares</a></li>
</ul>]]></content><author><name>Maciej Grochowski</name></author><category term="reverse-engineering" /><category term="firmware" /><category term="mips" /><summary type="html"><![CDATA[In Part 4, we tried to recompile 191 decompiled C files and hit 2,896 unresolved symbols. We learned why firmware linking is fundamentally different from userspace linking, and concluded that binary patching – the four-byte NOP from Part 3 – was the pragmatic fix.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://page-fault.io/assets/social-preview.png" /><media:content medium="image" url="https://page-fault.io/assets/social-preview.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Symbols, Scripts, and other linking nightmares</title><link href="https://page-fault.io/reverse-engineering/firmware/mips/2026/05/31/symbols-scripts-and-other-linking-nightmares.html" rel="alternate" type="text/html" title="Symbols, Scripts, and other linking nightmares" /><published>2026-05-31T12:00:00+00:00</published><updated>2026-05-31T12:00:00+00:00</updated><id>https://page-fault.io/reverse-engineering/firmware/mips/2026/05/31/symbols-scripts-and-other-linking-nightmares</id><content type="html" xml:base="https://page-fault.io/reverse-engineering/firmware/mips/2026/05/31/symbols-scripts-and-other-linking-nightmares.html"><![CDATA[<p>In <a href="/reverse-engineering/firmware/mips/2026/05/03/decompilers-annotations-and-other-ways-to-read-the-unreadable.html">Part 3</a>, we patched a four-byte bug in the firmware’s route engine and recalculated the CRC. Four bytes – problem solved. But as we noted at the end: what if the fix was not that simple? What if we needed to restructure a function, add error handling, or change a data structure?</p>

<p>You cannot do that with a hex editor. You would need to recompile from source.</p>

<p>We have Ghidra’s decompiled C. We have the function names, the string references, the module structure. We have <code class="language-plaintext highlighter-rouge">gcc</code> and a linker script. How hard can it be?</p>

<p>Very.</p>

<hr />

<h2 id="the-illusion-of-having-source-code">The Illusion of Having Source Code</h2>

<p>Ghidra produces C code that <em>reads</em> like source but is not source. Here is what Ghidra generates for <code class="language-plaintext highlighter-rouge">port_validate</code>:</p>

<figure class="highlight"><pre><code class="language-c" data-lang="c"><span class="n">undefined4</span> <span class="nf">FUN_80000c84</span><span class="p">(</span><span class="kt">int</span> <span class="n">param_1</span><span class="p">)</span>
<span class="p">{</span>
    <span class="n">undefined4</span> <span class="n">uVar1</span><span class="p">;</span>

    <span class="k">if</span> <span class="p">((</span><span class="n">param_1</span> <span class="o">&lt;</span> <span class="mi">0</span><span class="p">)</span> <span class="o">||</span> <span class="p">(</span><span class="mi">7</span> <span class="o">&lt;</span> <span class="n">param_1</span><span class="p">))</span> <span class="p">{</span>
        <span class="n">FUN_800002e4</span><span class="p">(</span><span class="n">s_ERROR_invalid_port_number_80001bb0</span><span class="p">);</span>
        <span class="n">FUN_800002e4</span><span class="p">(</span><span class="o">&amp;</span><span class="n">DAT_80001a88</span><span class="p">);</span>
        <span class="n">uVar1</span> <span class="o">=</span> <span class="mh">0xffffffff</span><span class="p">;</span>
    <span class="p">}</span>
    <span class="k">else</span> <span class="k">if</span> <span class="p">((</span><span class="o">*</span><span class="p">(</span><span class="n">uint</span> <span class="o">*</span><span class="p">)(</span><span class="n">param_1</span> <span class="o">*</span> <span class="mh">0x20</span> <span class="o">+</span> <span class="mh">0x80040200</span><span class="p">)</span> <span class="o">&amp;</span> <span class="mi">1</span><span class="p">)</span> <span class="o">==</span> <span class="mi">0</span><span class="p">)</span> <span class="p">{</span>
        <span class="n">FUN_800002e4</span><span class="p">(</span><span class="n">s_ERROR_port_is_not_enabled_80001b94</span><span class="p">);</span>
        <span class="n">FUN_800002e4</span><span class="p">(</span><span class="o">&amp;</span><span class="n">DAT_80001a88</span><span class="p">);</span>
        <span class="n">uVar1</span> <span class="o">=</span> <span class="mh">0xffffffff</span><span class="p">;</span>
    <span class="p">}</span>
    <span class="k">else</span> <span class="p">{</span>
        <span class="n">uVar1</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span>
    <span class="p">}</span>
    <span class="k">return</span> <span class="n">uVar1</span><span class="p">;</span>
<span class="p">}</span></code></pre></figure>

<p>The structure is correct. The logic is correct. But look at what is missing:</p>

<ul>
  <li><strong>Types are wrong</strong>: <code class="language-plaintext highlighter-rouge">undefined4</code> instead of <code class="language-plaintext highlighter-rouge">int</code>. No <code class="language-plaintext highlighter-rouge">port_config_t</code> struct.</li>
  <li><strong>Names are gone</strong>: <code class="language-plaintext highlighter-rouge">FUN_80000c84</code> instead of <code class="language-plaintext highlighter-rouge">port_validate</code>. <code class="language-plaintext highlighter-rouge">param_1</code> instead of <code class="language-plaintext highlighter-rouge">port_num</code>.</li>
  <li><strong>Data access is raw</strong>: <code class="language-plaintext highlighter-rouge">*(uint *)(param_1 * 0x20 + 0x80040200)</code> instead of <code class="language-plaintext highlighter-rouge">g_ports[port_num].flags</code>.</li>
  <li><strong>Symbols are addresses</strong>: <code class="language-plaintext highlighter-rouge">s_ERROR_invalid_port_number_80001bb0</code> and <code class="language-plaintext highlighter-rouge">DAT_80001a88</code> are Ghidra-generated names that encode the binary address.</li>
  <li><strong>No headers</strong>: No <code class="language-plaintext highlighter-rouge">#include</code>, no type definitions, no struct layouts.</li>
</ul>

<p>You can fix all of this manually. Rename <code class="language-plaintext highlighter-rouge">FUN_80000c84</code> to <code class="language-plaintext highlighter-rouge">port_validate</code>, change <code class="language-plaintext highlighter-rouge">param_1</code> to <code class="language-plaintext highlighter-rouge">port_num</code>, define the <code class="language-plaintext highlighter-rouge">port_config_t</code> struct. But now multiply that by every function in the firmware.</p>

<p>In the real project, Ghidra decompiled 4,426 functions across 1.5 MB of code. We organized them into 191 C source files matching the original module structure. Every single file compiled individually with <code class="language-plaintext highlighter-rouge">gcc -c</code>. It looked like we were close to a complete build.</p>

<p>Then we tried to link.</p>

<hr />

<h2 id="the-linking-wall">The Linking Wall</h2>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash"><span class="nv">$ </span>mipsel-linux-gnu-ld <span class="nt">-T</span> linker.ld <span class="nt">-o</span> firmware.elf <span class="k">*</span>.o
startup.o: undefined reference to <span class="sb">`</span>firmware_main<span class="s1">'
main.o: undefined reference to `s_system_starting_80001ee0'</span>
main.o: undefined reference to <span class="sb">`</span>s_initializing_ports_80001cf0<span class="s1">'
route.o: undefined reference to `DAT_80040008'</span>
cli.o: undefined reference to <span class="sb">`</span>s_Available_commands_800018b4<span class="s1">'
port.o: undefined reference to `_DAT_B8000008'</span>
... 2,896 more ...</code></pre></figure>

<p>2,896 unresolved symbols. Every single one needs to be resolved before the linker will produce a binary.</p>

<p>Why? When you compile with <code class="language-plaintext highlighter-rouge">gcc -c</code>, each source file compiles in isolation. The compiler trusts your <code class="language-plaintext highlighter-rouge">extern</code> declarations: “this function exists, this variable exists, the linker will find them.” Compilation succeeds because the compiler does not check. Linking fails because the linker <em>does</em>.</p>

<p>In userspace, the linker resolves symbols against libraries (<code class="language-plaintext highlighter-rouge">libc.so</code>, <code class="language-plaintext highlighter-rouge">libm.so</code>, etc.) and other object files. In firmware, there are no libraries. Every symbol must come from your own object files, your linker script, or it does not exist.</p>

<p>If you are following along with the sample firmware, you will not hit this wall yourself – it is a single translation unit that links cleanly from its own source. This wall belongs to the real project: 191 separately compiled object files, each trusting the others to supply the symbols it referenced. The sample reproduces the shape of the problem, not its scale.</p>

<hr />

<h2 id="the-symbol-taxonomy">The Symbol Taxonomy</h2>

<p>In the real project, the 2,896 unresolved symbols broke into six categories, each needing a different resolution strategy:</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">Unresolved Symbol Breakdown
============================

Category           Count   Example                          Resolution
─────────────────────────────────────────────────────────────────────────
String literals    1,172   s_system_starting_80001ee0       Extract from binary
Const data           993   DAT_C0012345                     Extract from binary
MMIO registers       444   _DAT_B8000008                    PROVIDE() in linker script
RAM globals          107   DAT_80240100                     BSS/data section placement
Code labels           91   LAB_C0012345                     Switch table extraction
Compiler runtime      15   func_0xC0170000                  Extract from binary
Other                 74   builtin_strncpy, _gp_disp        Stubs + linker magic
─────────────────────────────────────────────────────────────────────────
Total              2,896</code></pre></figure>

<hr />

<h2 id="string-literals-the-biggest-category">String Literals: The Biggest Category</h2>

<p>1,172 symbols are strings. Ghidra names them after their address: <code class="language-plaintext highlighter-rouge">s_system_starting_80001ee0</code> means “the string starting at address <code class="language-plaintext highlighter-rouge">0x80001EE0</code>.” The string data lives in the original binary’s <code class="language-plaintext highlighter-rouge">.rodata</code> section, which <code class="language-plaintext highlighter-rouge">objcopy</code> happily included in the flat binary.</p>

<p>The fix: extract the strings from the original binary and assemble them into a linkable object file:</p>

<figure class="highlight"><pre><code class="language-python" data-lang="python"><span class="c1"># Extract strings from firmware.bin at known addresses
# Generate assembly that the linker can resolve
</span>
<span class="k">with</span> <span class="nf">open</span><span class="p">(</span><span class="sh">'</span><span class="s">firmware.bin</span><span class="sh">'</span><span class="p">,</span> <span class="sh">'</span><span class="s">rb</span><span class="sh">'</span><span class="p">)</span> <span class="k">as</span> <span class="n">f</span><span class="p">:</span>
    <span class="n">binary</span> <span class="o">=</span> <span class="n">f</span><span class="p">.</span><span class="nf">read</span><span class="p">()</span>

<span class="nf">print</span><span class="p">(</span><span class="sh">'</span><span class="s">.section .rodata</span><span class="sh">'</span><span class="p">)</span>
<span class="k">for</span> <span class="n">name</span><span class="p">,</span> <span class="n">addr</span> <span class="ow">in</span> <span class="n">string_symbols</span><span class="p">:</span>
    <span class="n">offset</span> <span class="o">=</span> <span class="n">addr</span> <span class="o">-</span> <span class="n">BASE_ADDR</span>
    <span class="n">string</span> <span class="o">=</span> <span class="nf">read_cstring</span><span class="p">(</span><span class="n">binary</span><span class="p">,</span> <span class="n">offset</span><span class="p">)</span>
    <span class="nf">print</span><span class="p">(</span><span class="sa">f</span><span class="sh">'</span><span class="s">.global </span><span class="si">{</span><span class="n">name</span><span class="si">}</span><span class="sh">'</span><span class="p">)</span>
    <span class="nf">print</span><span class="p">(</span><span class="sa">f</span><span class="sh">'</span><span class="si">{</span><span class="n">name</span><span class="si">}</span><span class="s">: .asciz </span><span class="sh">"</span><span class="si">{</span><span class="nf">escape</span><span class="p">(</span><span class="n">string</span><span class="p">)</span><span class="si">}</span><span class="sh">"'</span><span class="p">)</span></code></pre></figure>

<p>Assemble this into <code class="language-plaintext highlighter-rouge">strings.o</code>, add it to the link – 1,172 symbols resolved. This works because the strings are <em>data</em>, not code. We are literally pulling bytes out of the original binary and giving them labels.</p>

<hr />

<h2 id="mmio-registers-addresses-without-content">MMIO Registers: Addresses Without Content</h2>

<p>444 symbols are hardware register addresses – UART, GPIO, timer, system control. They look like <code class="language-plaintext highlighter-rouge">_DAT_B8000008</code> (UART status register at <code class="language-plaintext highlighter-rouge">0xB8000008</code>). These are not in the binary at all – they are memory-mapped I/O addresses that the CPU accesses at runtime.</p>

<p>The linker does not need <em>content</em> for these. It just needs to know the address. That is what <code class="language-plaintext highlighter-rouge">PROVIDE()</code> does:</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">/* linker.ld -- MMIO register definitions */
SECTIONS
{
    /* ... code and data sections ... */
}

/* Hardware register addresses -- no content, just locations */
PROVIDE(_DAT_B8000000 = 0xB8000000);   /* UART TX */
PROVIDE(_DAT_B8000004 = 0xB8000004);   /* UART RX */
PROVIDE(_DAT_B8000008 = 0xB8000008);   /* UART Status */
PROVIDE(_DAT_B800000C = 0xB800000C);   /* UART Control */
PROVIDE(_DAT_B8010000 = 0xB8010000);   /* GPIO Data */
/* ... 439 more ... */</code></pre></figure>

<p><code class="language-plaintext highlighter-rouge">PROVIDE()</code> tells the linker: “if nobody else defines this symbol, it lives at this address.” The generated code will use <code class="language-plaintext highlighter-rouge">LUI</code>/<code class="language-plaintext highlighter-rouge">ADDIU</code> to load the address and then read/write the hardware register. No content needed in the binary – MMIO registers exist in the hardware, not in flash.</p>

<p>In the real project, 444 MMIO registers spanned five address blocks (<code class="language-plaintext highlighter-rouge">0xB4</code>, <code class="language-plaintext highlighter-rouge">0xB6</code>, <code class="language-plaintext highlighter-rouge">0xB9</code>, <code class="language-plaintext highlighter-rouge">0xBD</code>, <code class="language-plaintext highlighter-rouge">0xBF</code>), covering the switch core, packet processor, PHY interface, DMA engine, and management controller. Each one got a <code class="language-plaintext highlighter-rouge">PROVIDE()</code> line.</p>

<hr />

<h2 id="the-global-pointer-problem">The Global Pointer Problem</h2>

<p>On MIPS, frequently accessed global variables use <strong>GP-relative addressing</strong>. Instead of a two-instruction <code class="language-plaintext highlighter-rouge">LUI</code>/<code class="language-plaintext highlighter-rouge">ADDIU</code> pair, the compiler emits a single instruction:</p>

<figure class="highlight"><pre><code class="language-asm" data-lang="asm">lw  v0, -32000(gp)    # Load global at GP - 32000</code></pre></figure>

<p>This saves one instruction per access – significant in a firmware with thousands of global accesses. But it only works if <code class="language-plaintext highlighter-rouge">$gp</code> (the global pointer register) is set to the right value, and all code agrees on what that value is.</p>

<p>The problem: the decompiled code contains hardcoded GP offsets baked in by the original compiler. If the recompiled binary places globals at different addresses – even slightly – every GP-relative access breaks.</p>

<figure class="highlight"><pre><code class="language-c" data-lang="c"><span class="cm">/* Ghidra decompiled code contains things like: */</span>
<span class="o">*</span><span class="p">(</span><span class="kt">int</span> <span class="o">*</span><span class="p">)(</span><span class="n">gp</span> <span class="o">+</span> <span class="o">-</span><span class="mh">0x7c50</span><span class="p">)</span> <span class="o">=</span> <span class="mi">1</span><span class="p">;</span>    <span class="cm">/* original GP offset, hardcoded */</span></code></pre></figure>

<p>Three ways to fix it. Match GP exactly: set it to the original compiler’s value and reproduce the full original global layout, which requires perfect section placement. Disable GP-relative addressing entirely with <code class="language-plaintext highlighter-rouge">-G0 -mno-gpopt</code> and rewrite all accesses as full <code class="language-plaintext highlighter-rouge">LUI</code>/<code class="language-plaintext highlighter-rouge">ADDIU</code> pairs – the optimization is gone, but at least it links. Or declare all globals properly and let a fresh compiler recalculate offsets from scratch, which requires getting every global’s type and size right first.</p>

<p>Our sample firmware uses option two (<code class="language-plaintext highlighter-rouge">-G0 -mno-gpopt</code> in the Makefile) – every global uses the full <code class="language-plaintext highlighter-rouge">LUI</code>/<code class="language-plaintext highlighter-rouge">ADDIU</code> pair, so there is no GP dependency. Real vendor firmware is rarely that considerate.</p>

<hr />

<h2 id="the-custom-linker-script">The Custom Linker Script</h2>

<p>In userspace, the linker uses a default script that puts <code class="language-plaintext highlighter-rouge">.text</code> at a dynamic address, <code class="language-plaintext highlighter-rouge">.data</code> after it, and <code class="language-plaintext highlighter-rouge">.bss</code> at the end. The kernel’s ELF loader handles the rest. In firmware, you need total control:</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">/* linker.ld -- Firmware memory layout */
MEMORY
{
    ROM (rx)  : ORIGIN = 0x80000000, LENGTH = 256K
    RAM (rw)  : ORIGIN = 0x80040000, LENGTH = 64K
}

SECTIONS
{
    /* Exception vectors MUST be at 0x80000000 */
    .vectors 0x80000000 : {
        *(.vectors)
    } &gt; ROM

    /* Code follows immediately */
    .text : ALIGN(4) {
        *(.text)
        *(.text.*)
    } &gt; ROM

    /* Read-only data (strings, tables) */
    .rodata : ALIGN(16) {
        *(.rodata)
        *(.rodata.*)
    } &gt; ROM

    _etext = .;

    /* Globals in RAM (zeroed at boot by startup code) */
    .bss (NOLOAD) : ALIGN(4) {
        __bss_start = .;
        *(.bss)
        *(COMMON)
        __bss_end = .;
    } &gt; RAM

    /* Stack at top of RAM */
    _stack_top = ORIGIN(RAM) + LENGTH(RAM);
}</code></pre></figure>

<p>It is clean because we built it from scratch. The <code class="language-plaintext highlighter-rouge">linker.ld</code> bundled with the sample firmware is a fuller version of this – it adds explicit <code class="language-plaintext highlighter-rouge">.rodata.cli_*</code> ordering (so the hidden CLI commands land after the public ones), a real <code class="language-plaintext highlighter-rouge">.data</code> section loaded from ROM into RAM, and a <code class="language-plaintext highlighter-rouge">_gp</code> assignment – all elided here for clarity. In a real reverse engineering project, you would need to derive the whole thing from the binary’s memory layout – figuring out where each section starts and ends by analyzing address patterns in the code.</p>

<p>The <code class="language-plaintext highlighter-rouge">MEMORY</code> block defines two regions: ROM (where the flat binary lives in flash) and RAM (where globals and stack live at runtime). The exception vectors <em>must</em> be at <code class="language-plaintext highlighter-rouge">0x80000000</code> because that is where the MIPS CPU looks for them at reset. There is no flexibility here – get it wrong by one byte and the device crashes on boot.</p>

<p>Compare this with <a href="/kernel/2018/07/28/Elfs_Linkers_Other.html">the 2018 ELF post</a>, where we relied on the kernel’s ELF loader to parse program headers and map segments. In firmware, the linker script <em>is</em> the loader. Everything the kernel did for ELF – segment mapping, BSS zeroing, stack setup – you do yourself in a 20-line startup assembly file and a linker script.</p>

<hr />

<h2 id="the-compilation-gap">The Compilation Gap</h2>

<p>Here is what full recompilation actually requires:</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">Full Firmware Recompilation Checklist
=====================================

1. Decompile all functions              ✓  (Ghidra automated this)
2. Fix Ghidra output artifacts          ✗  (undefined4, unaff_gp, raw casts)
3. Reconstruct struct definitions       ✗  (Ghidra doesn't recover struct layouts)
4. Reconstruct enum/constant values     ✗  (flag values inferred from usage)
5. Extract strings from binary          ~  (automated but needs validation)
6. Extract const data from binary       ~  (automated, 993 symbols)
7. Map MMIO registers to PROVIDE()      ~  (automated, 444 addresses)
8. Resolve GP-relative globals          ✗  (requires matching GP layout)
9. Stub/extract compiler runtime        ~  (15 functions from vendor compiler)
10. Handle switch tables (LAB_*)        ✗  (91 computed gotos need manual work)
11. Write startup code                  ✗  (exception vectors, BSS clear, GP setup)
12. Write linker script                 ~  (derive from binary memory map)
13. Integration test                    ✗  (need hardware or accurate emulator)

Estimated effort: 100-170 hours for 1.5 MB of firmware</code></pre></figure>

<p>Step 1 is essentially free – Ghidra decompiles thousands of functions in minutes. Steps 5-7 can be automated with scripts. Steps 2-4, 8, and 10-13 require manual analysis and domain knowledge – which is where the time actually goes.</p>

<p>A note on using an LLM for this: tools like Claude Opus 4.8 can compress the annotation work substantially – renaming functions from Ghidra output, reconstructing struct definitions, generating <code class="language-plaintext highlighter-rouge">PROVIDE()</code> lists from address patterns. In practice, the 100-170 hour estimate drops by roughly a factor of five, to something in the 20-30 hour range for comparable firmware. But it does not eliminate the need for supervision. The model will confidently reconstruct a struct layout that is subtly wrong, or generate a section offset that is plausible but off by four bytes. Every output still needs a human to verify against the binary. The hours go down; the attention required does not.</p>

<p>The decompilation-to-compilation gap: decompiled C is a <em>description</em> of the binary’s behavior, not a <em>specification</em> of how to build it. That is like having a photograph of a building – you can see every room, but you do not have the blueprints, the materials list, or the building code compliance.</p>

<hr />

<h2 id="structural-verification-compiles-doesnt-link">Structural Verification: Compiles, Doesn’t Link</h2>

<p>In the real project, we stopped before linking. The goal was <strong>structural verification</strong>: every source file compiles, the function signatures are correct, the call graph is preserved – but no linked binary.</p>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash"><span class="c"># Compile every file independently</span>
<span class="k">for </span>f <span class="k">in </span>src/<span class="k">*</span>.c<span class="p">;</span> <span class="k">do
    </span>mipsel-linux-gnu-gcc-14 <span class="nt">-mips32r2</span> <span class="nt">-EL</span> <span class="nt">-O1</span> <span class="nt">-ffreestanding</span> <span class="se">\</span>
        <span class="nt">-nostdlib</span> <span class="nt">-mno-abicalls</span> <span class="nt">-fno-pic</span> <span class="nt">-G0</span> <span class="nt">-mno-gpopt</span> <span class="se">\</span>
        <span class="nt">-c</span> <span class="nt">-o</span> <span class="s2">"</span><span class="k">${</span><span class="nv">f</span><span class="p">%.c</span><span class="k">}</span><span class="s2">.o"</span> <span class="s2">"</span><span class="nv">$f</span><span class="s2">"</span>
<span class="k">done</span>

<span class="c"># Result: 191/191 compile successfully</span>
<span class="c"># But: 2,896 symbols remain unresolved at link time</span></code></pre></figure>

<p>This is not failure – it is a deliberate stopping point. Every file compiling confirms the decompilation is syntactically correct, function signatures match across the 191-file tree, and no conflicting type definitions exist. It does not prove the binary will boot, but for reading the code and finding bugs, it is more than enough.</p>

<p>For the bug fix we did in Part 3, this was exactly that. We understood the code, found the bug, and patched it in the binary. Full recompilation would have been the “right” way to fix it, but binary patching achieved the same result in an afternoon instead of several months.</p>

<hr />

<h2 id="when-full-recompilation-makes-sense">When Full Recompilation Makes Sense</h2>

<p>Binary patching is a scalpel. Recompilation is a complete rebuild. You choose based on the scope of changes:</p>

<table>
  <thead>
    <tr>
      <th>Change scope</th>
      <th>Approach</th>
      <th>Time</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Fix one branch instruction</td>
      <td>Binary patch (hex editor)</td>
      <td>Hours</td>
    </tr>
    <tr>
      <td>Change a constant or string</td>
      <td>Binary patch + CRC update</td>
      <td>Hours</td>
    </tr>
    <tr>
      <td>Add a new function</td>
      <td>Recompilation (probably)</td>
      <td>Weeks</td>
    </tr>
    <tr>
      <td>Restructure a module</td>
      <td>Recompilation (definitely)</td>
      <td>Months</td>
    </tr>
    <tr>
      <td>Port to new hardware</td>
      <td>Full rewrite from decompiled source</td>
      <td>Months+</td>
    </tr>
  </tbody>
</table>

<p>The real project chose binary patching for the immediate fix and structural verification as a long-term investment. If the vendor never provides an official fix, having 191 compilable C files is a foundation for future work – even if linking them into a single binary is still a 100+ hour project.</p>

<hr />

<h2 id="it-linked-it-still-didnt-boot">It Linked. It Still Didn’t Boot.</h2>

<p>Producing a correctly linked binary is not the finish line.</p>

<p>In a separate project – a 4+ MB firmware on a different embedded device – we reached the point this sample firmware’s analysis stops short of: every symbol resolved, the linker produced a binary, the image packed cleanly with valid CRC checksums. We flashed it. The device entered the primary bootloader, the secondary bootloader ran – and then silence. No UART output. No register activity. The firmware died without a word.</p>

<p>The first problem is the bootloader. As we covered in <a href="/reverse-engineering/firmware/mips/2026/04/22/headers-symbols-and-other-things-you-wont-find.html">Part 1</a>, the firmware image has separate partitions. The primary bootloader lives at a fixed flash address, is often write-protected, and cannot be replaced – it is the first code that runs at reset. That bootloader was compiled by the vendor with their exact toolchain, and it has specific expectations about the handoff to main firmware: entry point address, initial stack state, BSS layout, exception vector locations. A recompiled binary that differs even slightly from those expectations will fail at the handoff, silently.</p>

<p>The second problem is the compiler. Vendor firmware is frequently built with a proprietary SDK – a vendor-patched GCC, IAR, Green Hills MULTI, or the MIPS Technologies compiler itself. These make choices that <code class="language-plaintext highlighter-rouge">mipsel-linux-gnu-gcc-14</code> does not: interrupt calling conventions, stack alignment under exception handlers, runtime initialization order, startup register state. None of these surface as linker errors. They show up as a device that boots partway and then goes quiet.</p>

<p>Weeks went into educated guesses – adjusting the entry point, rewriting the startup assembly, matching section alignment, trying to reconstruct what the original runtime initialization looked like. The firmware died after the secondary bootloader every single time. At some point the cost of continued guessing exceeded the value of the exercise.</p>

<p>Binary patching was the right answer all along. It was just less satisfying to admit.</p>

<hr />

<h2 id="the-full-circle">The Full Circle</h2>

<p>We started this series with an ELF binary on an x86-64 Linux system – sections, symbols, relocations, a well-documented format, a kernel loader that handles everything. Four posts later, we have been to the other end of the spectrum: a flat binary on a MIPS32 embedded device – no headers, no symbols, no loader, a custom container format with CRC checksums.</p>

<p>The journey so far:</p>

<ol>
  <li><strong>Part 1</strong>: Understood what is missing in flat firmware vs ELF</li>
  <li><strong>Part 2</strong>: Found strings, traced functions, discovered hidden CLI commands</li>
  <li><strong>Part 3</strong>: Decompiled the route engine, found a priority inversion bug, patched it</li>
  <li><strong>Part 4</strong>: Tried to recompile, hit 2,896 unresolved symbols, got 191 files to compile – then learned that a linked binary still might not boot</li>
</ol>

<p>Every technique came from a real project. Every pattern is something we found in production firmware. The domain was changed (network gateway instead of the actual device), but the methods are universal. If you can find strings in a binary, trace function calls, decompile with Ghidra, and write a linker script, you can reverse engineer any firmware on any architecture.</p>

<p>The firmware doesn’t want you to read it. Read it anyway.</p>

<p>But there is one more question we have not asked: what happens when the vendor ships an update? In Part 5: “Versions, Callgraphs, and other ways to compare what changed”, we will put two firmware versions side by side, compare how the vendor chose to fix the same bug we patched, and use call graph statistics to automatically name functions in stripped binaries.</p>

<hr />

<h2 id="references">References</h2>

<ul>
  <li><a href="https://sourceware.org/binutils/docs/ld/Scripts.html">GNU ld Linker Scripts</a> – <code class="language-plaintext highlighter-rouge">MEMORY</code>, <code class="language-plaintext highlighter-rouge">SECTIONS</code>, <code class="language-plaintext highlighter-rouge">PROVIDE()</code>, and everything else in linker script language</li>
  <li><a href="https://ghidra-sre.org/">Ghidra</a> – Decompiler that produces the “almost-source” we tried to compile</li>
  <li><a href="https://training.mips.com/basic_mips/PDF/Caches_GP_Rel.pdf">MIPS GP-Relative Addressing</a> – Why <code class="language-plaintext highlighter-rouge">$gp</code> makes firmware linking painful</li>
  <li><a href="/kernel/2018/07/28/Elfs_Linkers_Other.html">ELF’s Linker’s and other magical creatures</a> – Where this all started, eight years ago</li>
  <li><a href="/reverse-engineering/firmware/mips/2026/04/22/headers-symbols-and-other-things-you-wont-find.html">Part 1: Headers, Symbols, and other things you won’t find</a> – Building the firmware and the image packager</li>
  <li><a href="/reverse-engineering/firmware/mips/2026/04/26/opcodes-prologues-and-other-hidden-patterns.html">Part 2: Opcodes, Prologues, and other hidden patterns</a> – Finding strings, functions, and the CLI table</li>
  <li><a href="/reverse-engineering/firmware/mips/2026/05/03/decompilers-annotations-and-other-ways-to-read-the-unreadable.html">Part 3: Decompilers, Annotations, and other ways to read the unreadable</a> – Decompiling the route engine and patching the binary</li>
</ul>]]></content><author><name>Maciej Grochowski</name></author><category term="reverse-engineering" /><category term="firmware" /><category term="mips" /><summary type="html"><![CDATA[In Part 3, we patched a four-byte bug in the firmware’s route engine and recalculated the CRC. Four bytes – problem solved. But as we noted at the end: what if the fix was not that simple? What if we needed to restructure a function, add error handling, or change a data structure?]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://page-fault.io/assets/social-preview.png" /><media:content medium="image" url="https://page-fault.io/assets/social-preview.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Decompilers, Annotations, and other ways to read the unreadable</title><link href="https://page-fault.io/reverse-engineering/firmware/mips/2026/05/03/decompilers-annotations-and-other-ways-to-read-the-unreadable.html" rel="alternate" type="text/html" title="Decompilers, Annotations, and other ways to read the unreadable" /><published>2026-05-03T12:00:00+00:00</published><updated>2026-05-03T12:00:00+00:00</updated><id>https://page-fault.io/reverse-engineering/firmware/mips/2026/05/03/decompilers-annotations-and-other-ways-to-read-the-unreadable</id><content type="html" xml:base="https://page-fault.io/reverse-engineering/firmware/mips/2026/05/03/decompilers-annotations-and-other-ways-to-read-the-unreadable.html"><![CDATA[<p>In <a href="/reverse-engineering/firmware/mips/2026/04/26/opcodes-prologues-and-other-hidden-patterns.html">Part 2</a>, we mapped the major functions, string references, and 16 CLI commands using nothing but <code class="language-plaintext highlighter-rouge">strings</code>, pattern matching, and a Python script. That left us with a suspicious route-engine string:</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">WARNING: no routes programmed for stack</code></pre></figure>

<p>Part 2 could show that the string exists and that it belongs to the route subsystem. It could not prove when that path runs, or whether it is the exact runtime symptom. For that we need structure. In this post, we will decompile the firmware, trace the execution path through the route engine, find why two boot routes silently miss the forwarding table, patch the binary, and recalculate the firmware image’s CRC. Four bytes is enough to restore forwarding in this sample.</p>

<hr />

<h2 id="from-disassembly-to-c-what-decompilers-give-you">From Disassembly to C: What Decompilers Give You</h2>

<p>Reading MIPS assembly works for short functions, but once you are staring at a 60-instruction function with nested loops and multiple branches, you want C. Decompilers like <a href="https://ghidra-sre.org/">Ghidra</a> take machine code and produce readable (if imperfect) C.</p>

<p>Here is what Ghidra produces for our <code class="language-plaintext highlighter-rouge">iterate_active_routes</code> function. This is the raw, unannotated output – no cleanup, exactly what the tool generates:</p>

<figure class="highlight"><pre><code class="language-c" data-lang="c"><span class="kt">void</span> <span class="nf">FUN_80000efc</span><span class="p">(</span><span class="kt">int</span> <span class="n">param_1</span><span class="p">)</span>
<span class="p">{</span>
    <span class="kt">int</span> <span class="o">*</span><span class="n">piVar1</span><span class="p">;</span>
    <span class="kt">int</span> <span class="n">iVar2</span><span class="p">;</span>
    <span class="kt">int</span> <span class="n">iVar3</span><span class="p">;</span>
    <span class="kt">int</span> <span class="n">iVar4</span><span class="p">;</span>

    <span class="n">iVar3</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span>
    <span class="n">iVar4</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span>
    <span class="n">piVar1</span> <span class="o">=</span> <span class="p">(</span><span class="kt">int</span> <span class="o">*</span><span class="p">)</span><span class="mh">0x80040008</span><span class="p">;</span>
    <span class="k">do</span> <span class="p">{</span>
        <span class="n">iVar2</span> <span class="o">=</span> <span class="o">*</span><span class="n">piVar1</span><span class="p">;</span>
        <span class="k">if</span> <span class="p">(</span><span class="n">iVar2</span> <span class="o">!=</span> <span class="mi">0</span><span class="p">)</span> <span class="p">{</span>
            <span class="k">if</span> <span class="p">((</span><span class="n">iVar2</span> <span class="o">&amp;</span> <span class="mi">4</span><span class="p">)</span> <span class="o">!=</span> <span class="mi">0</span><span class="p">)</span> <span class="p">{</span>
                <span class="n">FUN_80000b74</span><span class="p">(</span><span class="mi">4</span><span class="p">,</span> <span class="n">s_route_bridge_flag_set_skipping_80001f70</span><span class="p">);</span>
                <span class="n">iVar3</span> <span class="o">=</span> <span class="n">iVar3</span> <span class="o">+</span> <span class="mi">1</span><span class="p">;</span>
            <span class="p">}</span>
            <span class="k">else</span> <span class="k">if</span> <span class="p">((</span><span class="n">iVar2</span> <span class="o">&amp;</span> <span class="mi">1</span><span class="p">)</span> <span class="o">!=</span> <span class="mi">0</span><span class="p">)</span> <span class="p">{</span>
                <span class="n">FUN_80000b74</span><span class="p">(</span><span class="mi">4</span><span class="p">,</span> <span class="n">s_route_programming_forwarding_entry_80001f94</span><span class="p">);</span>
                <span class="n">iVar4</span> <span class="o">=</span> <span class="n">iVar4</span> <span class="o">+</span> <span class="mi">1</span><span class="p">;</span>
            <span class="p">}</span>
        <span class="p">}</span>
        <span class="n">piVar1</span> <span class="o">=</span> <span class="n">piVar1</span> <span class="o">+</span> <span class="mi">4</span><span class="p">;</span>
    <span class="p">}</span> <span class="k">while</span> <span class="p">(</span><span class="n">piVar1</span> <span class="o">!=</span> <span class="p">(</span><span class="kt">int</span> <span class="o">*</span><span class="p">)</span><span class="mh">0x80040208</span><span class="p">);</span>
    <span class="k">if</span> <span class="p">(</span><span class="n">iVar4</span> <span class="o">==</span> <span class="mi">0</span><span class="p">)</span> <span class="p">{</span>
        <span class="n">FUN_80000b74</span><span class="p">(</span><span class="mi">4</span><span class="p">,</span> <span class="n">s_WARNING_no_routes_programmed_80001fb8</span><span class="p">);</span>
    <span class="p">}</span>
    <span class="k">if</span> <span class="p">(</span><span class="mi">0</span> <span class="o">&lt;</span> <span class="n">iVar3</span><span class="p">)</span> <span class="p">{</span>
        <span class="n">FUN_80000b74</span><span class="p">(</span><span class="mi">4</span><span class="p">,</span> <span class="n">s_bridge_routes_deferred_80001fe0</span><span class="p">);</span>
    <span class="p">}</span>
    <span class="k">return</span><span class="p">;</span>
<span class="p">}</span></code></pre></figure>

<p>Ghidra got the structure right. You can see the loop, the flag checks, the log messages. But everything has auto-generated names: <code class="language-plaintext highlighter-rouge">FUN_80000efc</code>, <code class="language-plaintext highlighter-rouge">param_1</code>, <code class="language-plaintext highlighter-rouge">piVar1</code>, <code class="language-plaintext highlighter-rouge">iVar2</code>. The string references are the one lifeline – they tell you what each branch <em>does</em>, even when the variable names are meaningless.</p>

<p>This is what decompilation gives you: <strong>structure without semantics</strong>. You get the control flow and data flow right, but understanding what it <em>means</em> requires annotation.</p>

<hr />

<h2 id="annotation-from-pivar1-to-entry-flags">Annotation: From piVar1 to entry-&gt;flags</h2>

<p>Let us clean up Ghidra’s output. Using the strings, the function name we discovered in <a href="/reverse-engineering/firmware/mips/2026/04/26/opcodes-prologues-and-other-hidden-patterns.html">Part 2</a>, and our knowledge of the route table structure, we can annotate everything:</p>

<figure class="highlight"><pre><code class="language-c" data-lang="c"><span class="cm">/* iterate_active_routes() at 0x80000EFC
 * Walks the route table and programs active routes into the
 * hardware forwarding table. Bridge-flagged routes are deferred.
 */</span>
<span class="kt">void</span> <span class="nf">iterate_active_routes</span><span class="p">(</span><span class="kt">int</span> <span class="n">stack_idx</span><span class="p">)</span>
<span class="p">{</span>
    <span class="n">route_entry_t</span> <span class="o">*</span><span class="n">entry</span><span class="p">;</span>
    <span class="kt">int</span> <span class="n">flags</span><span class="p">;</span>
    <span class="kt">int</span> <span class="n">skipped</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span>       <span class="cm">/* bridge-deferred routes */</span>
    <span class="kt">int</span> <span class="n">programmed</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span>    <span class="cm">/* routes sent to forwarding table */</span>

    <span class="n">entry</span> <span class="o">=</span> <span class="o">&amp;</span><span class="n">g_route_table_entries</span><span class="p">[</span><span class="mi">0</span><span class="p">];</span>   <span class="cm">/* Ghidra's raw piVar1 points at entry[0].flags */</span>
    <span class="k">do</span> <span class="p">{</span>
        <span class="n">flags</span> <span class="o">=</span> <span class="n">entry</span><span class="o">-&gt;</span><span class="n">flags</span><span class="p">;</span>
        <span class="k">if</span> <span class="p">(</span><span class="n">flags</span> <span class="o">!=</span> <span class="mi">0</span><span class="p">)</span> <span class="p">{</span>                    <span class="cm">/* skip empty slots */</span>
            <span class="k">if</span> <span class="p">((</span><span class="n">flags</span> <span class="o">&amp;</span> <span class="mh">0x04</span><span class="p">)</span> <span class="o">!=</span> <span class="mi">0</span><span class="p">)</span> <span class="p">{</span>       <span class="cm">/* ROUTE_FLAG_BRIDGE = 0x04 */</span>
                <span class="n">log_msg</span><span class="p">(</span><span class="n">MOD_ROUTE</span><span class="p">,</span> <span class="s">"route: bridge flag set, skipping"</span><span class="p">);</span>
                <span class="n">skipped</span><span class="o">++</span><span class="p">;</span>
            <span class="p">}</span>
            <span class="k">else</span> <span class="k">if</span> <span class="p">((</span><span class="n">flags</span> <span class="o">&amp;</span> <span class="mh">0x01</span><span class="p">)</span> <span class="o">!=</span> <span class="mi">0</span><span class="p">)</span> <span class="p">{</span>  <span class="cm">/* ROUTE_FLAG_ACTIVE = 0x01 */</span>
                <span class="n">log_msg</span><span class="p">(</span><span class="n">MOD_ROUTE</span><span class="p">,</span> <span class="s">"route: programming forwarding entry"</span><span class="p">);</span>
                <span class="n">programmed</span><span class="o">++</span><span class="p">;</span>
            <span class="p">}</span>
        <span class="p">}</span>
        <span class="n">entry</span><span class="o">++</span><span class="p">;</span>
    <span class="p">}</span> <span class="k">while</span> <span class="p">(</span><span class="n">entry</span> <span class="o">!=</span> <span class="o">&amp;</span><span class="n">g_route_table_entries</span><span class="p">[</span><span class="mi">32</span><span class="p">]);</span>

    <span class="k">if</span> <span class="p">(</span><span class="n">programmed</span> <span class="o">==</span> <span class="mi">0</span><span class="p">)</span> <span class="p">{</span>
        <span class="n">log_msg</span><span class="p">(</span><span class="n">MOD_ROUTE</span><span class="p">,</span> <span class="s">"WARNING: no routes programmed for stack"</span><span class="p">);</span>
    <span class="p">}</span>
    <span class="k">if</span> <span class="p">(</span><span class="n">skipped</span> <span class="o">&gt;</span> <span class="mi">0</span><span class="p">)</span> <span class="p">{</span>
        <span class="n">log_msg</span><span class="p">(</span><span class="n">MOD_ROUTE</span><span class="p">,</span> <span class="s">"bridge routes deferred to bridge handler"</span><span class="p">);</span>
    <span class="p">}</span>
<span class="p">}</span></code></pre></figure>

<p>Now it reads like source code. And now the bug jumps out.</p>

<hr />

<h2 id="the-bug-priority-inversion">The Bug: Priority Inversion</h2>

<p>Look at the two if-statements in the loop:</p>

<figure class="highlight"><pre><code class="language-c" data-lang="c"><span class="k">if</span> <span class="p">((</span><span class="n">flags</span> <span class="o">&amp;</span> <span class="mh">0x04</span><span class="p">)</span> <span class="o">!=</span> <span class="mi">0</span><span class="p">)</span> <span class="p">{</span>       <span class="cm">/* check BRIDGE flag first */</span>
    <span class="n">skipped</span><span class="o">++</span><span class="p">;</span>
    <span class="k">continue</span><span class="p">;</span>
<span class="p">}</span>
<span class="k">else</span> <span class="nf">if</span> <span class="p">((</span><span class="n">flags</span> <span class="o">&amp;</span> <span class="mh">0x01</span><span class="p">)</span> <span class="o">!=</span> <span class="mi">0</span><span class="p">)</span> <span class="p">{</span>  <span class="cm">/* check ACTIVE flag second */</span>
    <span class="n">programmed</span><span class="o">++</span><span class="p">;</span>
<span class="p">}</span></code></pre></figure>

<p>The <strong>bridge flag</strong> (0x04) is checked <strong>before</strong> the <strong>active flag</strong> (0x01). This creates an <code class="language-plaintext highlighter-rouge">if/else</code> where a route that has <em>both</em> flags set hits the bridge check first and gets diverted – the active check never runs.</p>

<p>Let us trace what happens during boot. The firmware adds four routes:</p>

<figure class="highlight"><pre><code class="language-c" data-lang="c"><span class="n">route_add</span><span class="p">(</span><span class="mh">0x0A01</span><span class="p">,</span> <span class="mh">0x01</span><span class="p">,</span> <span class="n">ACTIVE</span> <span class="o">|</span> <span class="n">STATIC</span><span class="p">);</span>           <span class="cm">/* 10.1.x.x → port 0 */</span>
<span class="n">route_add</span><span class="p">(</span><span class="mh">0x0A02</span><span class="p">,</span> <span class="mh">0x02</span><span class="p">,</span> <span class="n">ACTIVE</span> <span class="o">|</span> <span class="n">STATIC</span><span class="p">);</span>           <span class="cm">/* 10.2.x.x → port 1 */</span>
<span class="n">route_add</span><span class="p">(</span><span class="mh">0xC0A8</span><span class="p">,</span> <span class="mh">0x04</span><span class="p">,</span> <span class="n">ACTIVE</span> <span class="o">|</span> <span class="n">BRIDGE</span><span class="p">);</span>           <span class="cm">/* 192.168.x → port 2 */</span>
<span class="n">route_add</span><span class="p">(</span><span class="mh">0xAC10</span><span class="p">,</span> <span class="mh">0x08</span><span class="p">,</span> <span class="n">ACTIVE</span> <span class="o">|</span> <span class="n">STATIC</span> <span class="o">|</span> <span class="n">BRIDGE</span><span class="p">);</span>  <span class="cm">/* 172.16.x → port 3 */</span></code></pre></figure>

<p>Routes 1 and 2 have flags <code class="language-plaintext highlighter-rouge">0x03</code> (ACTIVE + STATIC). The BRIDGE check fails (<code class="language-plaintext highlighter-rouge">0x03 &amp; 0x04 = 0</code>), so they fall through to the ACTIVE check (<code class="language-plaintext highlighter-rouge">0x03 &amp; 0x01 = 1</code>) and get programmed. Good.</p>

<p>Routes 3 and 4 have flags <code class="language-plaintext highlighter-rouge">0x05</code> and <code class="language-plaintext highlighter-rouge">0x07</code> respectively – both include BRIDGE (0x04). The bridge check succeeds first: <code class="language-plaintext highlighter-rouge">0x05 &amp; 0x04 = 4</code>, which is non-zero, so the bridge path wins. The route’s ACTIVE bit exists, but the programming path never gets to use it.</p>

<p>The result? Only 2 of 4 routes get programmed into the forwarding table. Traffic to 192.168.x.x and 172.16.x.x silently blackholes. What the final boot scan actually produces:</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">[MOD:4] route: programming forwarding entry    ← route 1 (10.1.x)
[MOD:4] route: programming forwarding entry    ← route 2 (10.2.x)
[MOD:4] route: bridge flag set, skipping       ← route 3 (192.168.x) SKIPPED!
[MOD:4] route: bridge flag set, skipping       ← route 4 (172.16.x) SKIPPED!
[MOD:4] bridge routes deferred to bridge handler</code></pre></figure>

<p><code class="language-plaintext highlighter-rouge">iterate_active_routes</code> is called by <code class="language-plaintext highlighter-rouge">route_add</code> after every insertion, so the scan runs four times during boot – not once at the end. Calls 1 and 2 only see ACTIVE+STATIC entries and program them cleanly. Call 3 hits route 3’s BRIDGE flag for the first time. Call 4 does the same for route 4.</p>

<p>Routes 1 and 2 get re-programmed on every rescan (idempotent), so <code class="language-plaintext highlighter-rouge">programmed</code> is never zero – the <code class="language-plaintext highlighter-rouge">WARNING: no routes programmed</code> guard never fires for this route set. It is a safety net for a completely empty forwarding table, not for a partially-broken one. That is the correction decompilation gives us over the string-only view from Part 2: the warning string was a clue to inspect this function, while the actual bug is the quieter “bridge flag set, skipping” path.</p>

<p>This is a <strong>priority inversion</strong> – the bridge flag check has higher priority than the active flag check, but the developer intended them to be independent. A route should be programmed if it is ACTIVE, regardless of whether it is also BRIDGE. The bridge handler should get a <em>copy</em> of bridge-flagged routes, not steal them from the forwarding table entirely.</p>

<p>In the real vendor firmware project, we found exactly this pattern: a flag check at instruction <code class="language-plaintext highlighter-rouge">0xC012936C</code> tested a P2P flag before the active-capability flag, causing an entire subsystem to silently fail whenever a specific hardware configuration was present. The symptom was identical – “no rules programmed” – and it took weeks of trace analysis to find those four bytes.</p>

<hr />

<h2 id="the-disassembly-four-instructions">The Disassembly: Four Instructions</h2>

<p>Let us look at the exact machine code. The bug lives in four instructions at offsets <code class="language-plaintext highlighter-rouge">0x0F6C</code>-<code class="language-plaintext highlighter-rouge">0x0F78</code>:</p>

<figure class="highlight"><pre><code class="language-asm" data-lang="asm">80000f64:   8e020000    lw      v0, 0(s0)       # v0 = entry-&gt;flags
80000f68:   1040fffb    beqz    v0, next_entry  # if flags == 0, skip

80000f6c:   30430004    andi    v1, v0, 0x4     # v1 = flags &amp; BRIDGE  ← THE BUG
80000f70:   1460fff5    bnez    v1, bridge_skip # if BRIDGE set → branch to bridge_skip
80000f74:   30420001    andi    v0, v0, 0x1     # delay slot: executes even when branch is taken
80000f78:   1040fff7    beqz    v0, next_entry  # ← never reached when branch taken</code></pre></figure>

<p>The instruction at <code class="language-plaintext highlighter-rouge">0x80000F70</code> is the problem: <code class="language-plaintext highlighter-rouge">bnez v1, bridge_skip</code>. When the BRIDGE bit is set this branch fires – and because <code class="language-plaintext highlighter-rouge">0x80000F74</code> is the branch delay slot, the <code class="language-plaintext highlighter-rouge">andi</code> there executes either way. What never executes is <code class="language-plaintext highlighter-rouge">0x80000F78</code>: the branch that tests the ACTIVE result before the programming path. Routes with both BRIDGE and ACTIVE flags are treated as bridge-only, not as active routes that also participate in forwarding.</p>

<hr />

<h2 id="binary-patching-the-four-byte-fix">Binary Patching: The Four-Byte Fix</h2>

<p>The simplest forwarding fix: NOP out the branch at <code class="language-plaintext highlighter-rouge">0x0F70</code>. Replace <code class="language-plaintext highlighter-rouge">bnez v1, bridge_skip</code> with <code class="language-plaintext highlighter-rouge">nop</code>, so the bridge flag check still runs but never diverts execution. All non-empty routes reach the ACTIVE check.</p>

<figure class="highlight"><pre><code class="language-python" data-lang="python"><span class="kn">import</span> <span class="n">struct</span>

<span class="k">with</span> <span class="nf">open</span><span class="p">(</span><span class="sh">'</span><span class="s">firmware.bin</span><span class="sh">'</span><span class="p">,</span> <span class="sh">'</span><span class="s">rb</span><span class="sh">'</span><span class="p">)</span> <span class="k">as</span> <span class="n">f</span><span class="p">:</span>
    <span class="n">data</span> <span class="o">=</span> <span class="nf">bytearray</span><span class="p">(</span><span class="n">f</span><span class="p">.</span><span class="nf">read</span><span class="p">())</span>

<span class="c1"># The bug: bnez instruction at offset 0x0F70
</span><span class="n">bug_offset</span> <span class="o">=</span> <span class="mh">0x0F70</span>
<span class="n">original</span> <span class="o">=</span> <span class="n">struct</span><span class="p">.</span><span class="nf">unpack_from</span><span class="p">(</span><span class="sh">'</span><span class="s">&lt;I</span><span class="sh">'</span><span class="p">,</span> <span class="n">data</span><span class="p">,</span> <span class="n">bug_offset</span><span class="p">)[</span><span class="mi">0</span><span class="p">]</span>
<span class="nf">print</span><span class="p">(</span><span class="sa">f</span><span class="sh">"</span><span class="s">Original: 0x</span><span class="si">{</span><span class="n">original</span><span class="si">:</span><span class="mi">08</span><span class="n">X</span><span class="si">}</span><span class="s">  (bnez v1, bridge_skip)</span><span class="sh">"</span><span class="p">)</span>

<span class="c1"># The fix: NOP (0x00000000)
</span><span class="n">struct</span><span class="p">.</span><span class="nf">pack_into</span><span class="p">(</span><span class="sh">'</span><span class="s">&lt;I</span><span class="sh">'</span><span class="p">,</span> <span class="n">data</span><span class="p">,</span> <span class="n">bug_offset</span><span class="p">,</span> <span class="mh">0x00000000</span><span class="p">)</span>
<span class="nf">print</span><span class="p">(</span><span class="sa">f</span><span class="sh">"</span><span class="s">Patched:  0x00000000  (nop)</span><span class="sh">"</span><span class="p">)</span>

<span class="k">with</span> <span class="nf">open</span><span class="p">(</span><span class="sh">'</span><span class="s">firmware_patched.bin</span><span class="sh">'</span><span class="p">,</span> <span class="sh">'</span><span class="s">wb</span><span class="sh">'</span><span class="p">)</span> <span class="k">as</span> <span class="n">f</span><span class="p">:</span>
    <span class="n">f</span><span class="p">.</span><span class="nf">write</span><span class="p">(</span><span class="n">data</span><span class="p">)</span></code></pre></figure>

<p>Four bytes. That is the entire forwarding patch. The bridge flag still gets tested by the <code class="language-plaintext highlighter-rouge">andi</code> instruction, but the result is never acted on – all non-empty routes proceed to the ACTIVE check.</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">After patching:
80000f6c:   30430004    andi    v1, v0, 0x4     # v1 = flags &amp; BRIDGE (still runs)
80000f70:   00000000    nop                      # ← PATCHED (was: bnez v1, bridge_skip)
80000f74:   30420001    andi    v0, v0, 0x1     # former delay slot; already executed before patch
80000f78:   1040fff7    beqz    v0, next_entry  # ← now always reached</code></pre></figure>

<p>Now routes with <code class="language-plaintext highlighter-rouge">ACTIVE | BRIDGE</code> (flags 0x05 or 0x07) hit the NOP, the former delay-slot <code class="language-plaintext highlighter-rouge">andi</code> runs as before, but execution falls through to <code class="language-plaintext highlighter-rouge">0x80000F78</code> – the <code class="language-plaintext highlighter-rouge">beqz</code> that was previously unreachable. The ACTIVE flag is set, the branch is not taken, and the route gets programmed. The forwarding table gets all four boot routes instead of two.</p>

<hr />

<h2 id="crc-recalculation-you-break-it-you-fix-it">CRC Recalculation: You Break It, You Fix It</h2>

<p>Remember the firmware image from <a href="/reverse-engineering/firmware/mips/2026/04/22/headers-symbols-and-other-things-you-wont-find.html">Part 1</a>? Each partition has a CRC-32 checksum, and the image header has a global checksum. Patching even a single byte invalidates both:</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">Original firmware.bin CRC-32:  0x9F25E0DC
Patched firmware.bin CRC-32:   0x156FE139   ← completely different!</code></pre></figure>

<p>If we packed the patched binary into a firmware image without updating the CRC, the device’s bootloader would reject it. We need to repack:</p>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash"><span class="nv">$ </span>python3 fwpack.py firmware_patched.bin firmware_patched.img

Packed firmware image: firmware_patched.img
  Total size:    9152 bytes
  Image CRC-32:  0xC5E88C64
  Partitions:    3

  <span class="o">[</span>bootloader]
    Offset:  0x0120
    Size:    32 bytes
    CRC-32:  0x4F7A6ACA
  <span class="o">[</span>main_fw]
    Offset:  0x0160
    Size:    8496 bytes
    CRC-32:  0x156FE139    ← updated automatically
  <span class="o">[</span>config]
    Offset:  0x22B0
    Size:    256 bytes
    CRC-32:  0x2683AC5B</code></pre></figure>

<p>The packager recalculates all checksums from the payload data. In practice, reverse engineering the vendor’s CRC algorithm (or identifying it as standard CRC-32) is a necessary step before you can deploy any binary patch. Some vendors use non-standard polynomials, or CRC the payload with a salt, or layer RSA signatures on top. Each additional layer makes patching harder – but the fundamental workflow is the same: patch the code, recalculate the integrity checks, repack.</p>

<hr />

<h2 id="what-we-changed-and-what-we-didnt">What We Changed (And What We Didn’t)</h2>

<p>Let us be precise about what this patch does:</p>

<table>
  <thead>
    <tr>
      <th> </th>
      <th>Before patch</th>
      <th>After patch</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Routes with ACTIVE only</td>
      <td>Programmed</td>
      <td>Programmed</td>
    </tr>
    <tr>
      <td>Routes with BRIDGE only</td>
      <td>Skipped (deferred log)</td>
      <td>Skipped (silent)</td>
    </tr>
    <tr>
      <td>Routes with ACTIVE + BRIDGE</td>
      <td><strong>Skipped</strong> (bug!)</td>
      <td><strong>Programmed</strong> (fixed!)</td>
    </tr>
    <tr>
      <td>Routes with no flags</td>
      <td>Skipped (empty)</td>
      <td>Skipped (empty)</td>
    </tr>
  </tbody>
</table>

<p>The ACTIVE + BRIDGE path is fixed. The ACTIVE-only path is unchanged. But the BRIDGE-only path changes: before the patch, a BRIDGE-only route incremented <code class="language-plaintext highlighter-rouge">skipped</code> and triggered the “bridge routes deferred” log. After the NOP, the branch that drove that path is gone – the route falls through to the ACTIVE check, which fails (<code class="language-plaintext highlighter-rouge">flags &amp; 0x01 == 0</code>), and exits via <code class="language-plaintext highlighter-rouge">next_entry</code> with no counter and no log.</p>

<p>In this firmware, that is fine. The “bridge handler” is a log message; there is no real deferred queue behind it. If there were, NOPping the branch would trade one bug for another. In a real project you would replace the branch with proper independent checks – test BRIDGE and ACTIVE separately rather than as a chain. The NOP is the minimal patch that fixes the symptom we can measure; the right fix is a code restructure.</p>

<hr />

<h2 id="the-bigger-picture">The Bigger Picture</h2>

<p>We went from suspicious route strings to a root cause and a binary patch in one post. Let us trace back the full path:</p>

<ol>
  <li><strong>Part 1</strong>: Built the flat binary – 8,496 bytes, no symbols, no sections</li>
  <li><strong>Part 2</strong>: <code class="language-plaintext highlighter-rouge">strings</code> found suspicious route messages; xref scanning pointed us at the route subsystem</li>
  <li><strong>Part 3</strong>: Decompilation revealed the flag priority inversion in <code class="language-plaintext highlighter-rouge">iterate_active_routes</code> at <code class="language-plaintext highlighter-rouge">0x80000EFC</code></li>
  <li><strong>Binary patch</strong>: NOP at offset <code class="language-plaintext highlighter-rouge">0x0F70</code> (4 bytes)</li>
  <li><strong>CRC update</strong>: Repack firmware image with new checksums</li>
</ol>

<p>That is the real-world firmware RE workflow. Find a clue in the strings. Trace it to a function. Decompile the function. Understand the logic. Patch the binary. Update the checksums.</p>

<p>In the real project, this same workflow – from a warning message in a UART log to a binary patch at a specific address – took weeks. Most of that time was spent building the tools and understanding the firmware’s architecture. The actual bug was four bytes, just like this one.</p>

<p>But here is the thing: we patched four bytes. What if the fix was more complex – restructuring a function, adding new code, changing data structures? You cannot do that with a hex editor. You would need to recompile from source.</p>

<p>And that raises a question: we have Ghidra’s decompiled C for every function. We have the source filenames from the assert strings. Can’t we just… compile it?</p>

<p>In Part 4: “Symbols, Scripts, and other linking nightmares”, we will try. And we will discover why that simple question has a very complicated answer.</p>

<hr />

<h2 id="references">References</h2>

<ul>
  <li><a href="https://ghidra-sre.org/">Ghidra</a> – NSA’s Software Reverse Engineering Framework (decompiler used in this post)</li>
  <li><a href="https://s3-eu-west-1.amazonaws.com/downloads-mips/documents/MD00086-2B-MIPS32BIS-AFP-6.06.pdf">MIPS Instruction Reference</a> – <code class="language-plaintext highlighter-rouge">ANDI</code>, <code class="language-plaintext highlighter-rouge">BNEZ</code>, <code class="language-plaintext highlighter-rouge">NOP</code> instruction encoding</li>
  <li><a href="https://en.wikipedia.org/wiki/Cyclic_redundancy_check">CRC-32</a> – The checksum algorithm used in firmware image integrity verification</li>
  <li><a href="/reverse-engineering/firmware/mips/2026/04/22/headers-symbols-and-other-things-you-wont-find.html">Part 1: Headers, Symbols, and other things you won’t find</a> – Building the firmware and the image packager</li>
  <li><a href="/reverse-engineering/firmware/mips/2026/04/26/opcodes-prologues-and-other-hidden-patterns.html">Part 2: Opcodes, Prologues, and other hidden patterns</a> – Finding strings, functions, and the CLI table</li>
</ul>]]></content><author><name>Maciej Grochowski</name></author><category term="reverse-engineering" /><category term="firmware" /><category term="mips" /><summary type="html"><![CDATA[In Part 2, we mapped the major functions, string references, and 16 CLI commands using nothing but strings, pattern matching, and a Python script. That left us with a suspicious route-engine string:]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://page-fault.io/assets/social-preview.png" /><media:content medium="image" url="https://page-fault.io/assets/social-preview.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Opcodes, Prologues, and other hidden patterns</title><link href="https://page-fault.io/reverse-engineering/firmware/mips/2026/04/26/opcodes-prologues-and-other-hidden-patterns.html" rel="alternate" type="text/html" title="Opcodes, Prologues, and other hidden patterns" /><published>2026-04-26T12:00:00+00:00</published><updated>2026-04-26T12:00:00+00:00</updated><id>https://page-fault.io/reverse-engineering/firmware/mips/2026/04/26/opcodes-prologues-and-other-hidden-patterns</id><content type="html" xml:base="https://page-fault.io/reverse-engineering/firmware/mips/2026/04/26/opcodes-prologues-and-other-hidden-patterns.html"><![CDATA[<p>In <a href="/reverse-engineering/firmware/mips/2026/04/22/headers-symbols-and-other-things-you-wont-find.html">Part 1</a>, we built a flat MIPS32 firmware binary and stared at 8,496 bytes with no headers, no symbols, and no sections. <code class="language-plaintext highlighter-rouge">file</code> said “data.” <code class="language-plaintext highlighter-rouge">readelf</code> said nothing. We left off with a question: now what?</p>

<p>The temptation is to fire up a disassembler and start reading opcodes. Resist it. The first useful thing you can do with an unknown firmware binary isn’t disassembly – it’s reading the text that the developers left behind.</p>

<blockquote>
  <p><strong>Code and binary:</strong> the <code class="language-plaintext highlighter-rouge">firmware.bin</code> analyzed below, the build sources, and every script in this post are bundled in <a href="https://res.cloudinary.com/gotocco/raw/upload/v1777182332/sample_firmware.tar.gz">sample_firmware.tar.gz</a>. Grab it and follow along on your own copy:</p>
</blockquote>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash">curl <span class="nt">-L</span> <span class="nt">-O</span> https://res.cloudinary.com/gotocco/raw/upload/v1777182332/sample_firmware.tar.gz
<span class="nb">tar </span>xf sample_firmware.tar.gz
<span class="nb">cd </span>sample_firmware</code></pre></figure>

<hr />

<h2 id="strings-the-first-foothold">Strings: The First Foothold</h2>

<p>Every firmware binary contains embedded strings – log messages, error texts, command prompts, version information. The developers who wrote the firmware needed to print things to the serial console, log diagnostic messages, and display help text. All of that text lives in the <code class="language-plaintext highlighter-rouge">.rodata</code> section, which survives the <code class="language-plaintext highlighter-rouge">objcopy</code> to flat binary untouched.</p>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash"><span class="nv">$ </span>strings firmware.bin</code></pre></figure>

<p>111 strings come pouring out. Let’s categorize what we find:</p>

<p><strong>Boot sequence messages:</strong></p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">system starting
initializing ports
initializing route table
WARNING: no routes programmed for stack
bridge routes deferred to bridge handler
system ready</code></pre></figure>

<p>These tell us the firmware has a boot sequence: it starts up, initializes a port subsystem, sets up a routing table, and then something goes wrong with route programming before declaring “system ready.” That WARNING is interesting – we’ll come back to it.</p>

<p><strong>CLI interface:</strong></p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">gw&gt;
Available commands:
Unknown command. Type 'help' for list.</code></pre></figure>

<p>There’s a command-line interface with a <code class="language-plaintext highlighter-rouge">gw&gt;</code> prompt. This is a gateway device.</p>

<p><strong>Command names:</strong></p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">help                    version                 uptime
status                  portdump                routedump
reset                   loopback                memtest
log</code></pre></figure>

<p>And then, further down in the binary, a second cluster:</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">dbglvl                  peek                    poke
flashid                 crashme                 regdump</code></pre></figure>

<p>Re-running <code class="language-plaintext highlighter-rouge">strings</code> with <code class="language-plaintext highlighter-rouge">-t x</code> to see file offsets makes the layout visible:</p>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash"><span class="nv">$ </span>strings <span class="nt">-t</span> x firmware.bin | <span class="nb">grep</span> <span class="nt">-E</span> <span class="s2">"^</span><span class="se">\s</span><span class="s2">+[0-9a-f]+ (help|version|uptime|status|portdump|routedump|reset|loopback|memtest|log|dbglvl|peek|poke|flashid|crashme|regdump)$"</span>
   17bc memtest
   17e0 loopback
   1800 reset
   181c routedump
   1844 portdump
   1864 status
   1880 <span class="nb">uptime
   </span>18a0 version
   18c0 <span class="nb">help
   </span>1958 regdump
   197c crashme
   199c flashid
   19bc poke
   19d8 peek
   19fc dbglvl</code></pre></figure>

<p>The two clusters sit in different regions of <code class="language-plaintext highlighter-rouge">.rodata</code>: the familiar-looking names are packed contiguously between <code class="language-plaintext highlighter-rouge">0x17bc</code> and <code class="language-plaintext highlighter-rouge">0x18c0</code>, the rest live ~0x80 bytes further on between <code class="language-plaintext highlighter-rouge">0x1958</code> and <code class="language-plaintext highlighter-rouge">0x19fc</code>, and the two groups never interleave. That’s not how a compiler lays out a single array. (<code class="language-plaintext highlighter-rouge">log</code> is missing from this grep because <code class="language-plaintext highlighter-rouge">strings</code> defaults to a minimum length of 4 characters and <code class="language-plaintext highlighter-rouge">log</code> is three; pass <code class="language-plaintext highlighter-rouge">-n 3</code> to see it – it sits at <code class="language-plaintext highlighter-rouge">0x17a8</code>, inside the first cluster.)</p>

<p>Wait. <code class="language-plaintext highlighter-rouge">peek</code>? <code class="language-plaintext highlighter-rouge">poke</code>? <code class="language-plaintext highlighter-rouge">crashme</code>? Those don’t sound like commands you’d put in a user manual.</p>

<p>We don’t yet know which of these are reachable from the user-facing CLI and which aren’t – the layout is suggestive, but suggestion isn’t proof. For now we have 16 command-shaped strings split across two clusters. Keep that in mind.</p>

<p><strong>Error messages and source filenames:</strong></p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">ERROR: invalid port number (0-7)
ERROR: routing table full
ASSERT FAILED
main.c    cli.c    port.c    route.c    diag.c    flash.c</code></pre></figure>

<p>The firmware was built from six source files. The assert macro includes the filename – a gift for reverse engineers, because now we know the firmware’s module structure without reading a single instruction.</p>

<p><strong>Firmware identity:</strong></p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">NetGW-MIPS32
v2.4.1-rc3
Apr 2026
(c) Embedded Network Systems</code></pre></figure>

<p>In under a minute, using nothing but <code class="language-plaintext highlighter-rouge">strings</code>, we now know:</p>
<ul>
  <li>This is a <strong>network gateway</strong> device (<code class="language-plaintext highlighter-rouge">NetGW-MIPS32</code>, prompt <code class="language-plaintext highlighter-rouge">gw&gt;</code>)</li>
  <li>It manages <strong>ports</strong> and <strong>routes</strong> with a forwarding table</li>
  <li>It has a <strong>CLI</strong> with at least 16 command-shaped strings, in two physically separated clusters (we don’t yet know which are user-visible)</li>
  <li>It was built from <strong>6 source modules</strong> (main, cli, port, route, diag, flash)</li>
  <li>Something might be <strong>wrong with route programming</strong> during boot</li>
  <li>The second cluster has names like <code class="language-plaintext highlighter-rouge">peek</code>, <code class="language-plaintext highlighter-rouge">poke</code>, <code class="language-plaintext highlighter-rouge">crashme</code> – shapes of developer commands, not user ones</li>
</ul>

<p>All from one command. No disassembly required. In the real vendor firmware project this series is based on, <code class="language-plaintext highlighter-rouge">strings</code> was our very first tool – and it revealed over 4,400 embedded strings that mapped directly to every subsystem in a 1.5 MB binary.</p>

<hr />

<h2 id="mips32-survival-guide">MIPS32 Survival Guide</h2>

<p>You don’t need to learn MIPS assembly to follow this series. You need exactly five patterns. If you can spot these five things in a disassembly listing, you can trace function calls, find function boundaries, and follow data references across the entire binary.</p>

<p><strong>1. Function prologue</strong> – “I’m starting a function”</p>

<figure class="highlight"><pre><code class="language-asm" data-lang="asm">addiu   sp, sp, -24       # Allocate 24 bytes of stack space
sw      ra, 20(sp)        # Save return address
sw      s0, 16(sp)        # Save callee-saved register</code></pre></figure>

<p>Every function begins by growing the stack downward. The size tells you how many local variables and saved registers it needs. <code class="language-plaintext highlighter-rouge">-24</code> is a small function. <code class="language-plaintext highlighter-rouge">-184</code> is a big one with lots of locals.</p>

<p><strong>2. Function epilogue</strong> – “I’m returning to my caller”</p>

<figure class="highlight"><pre><code class="language-asm" data-lang="asm">lw      ra, 20(sp)        # Restore return address
lw      s0, 16(sp)        # Restore saved register
jr      ra                # Jump to return address
addiu   sp, sp, 24        # Deallocate stack (delay slot!)</code></pre></figure>

<p>Note the <code class="language-plaintext highlighter-rouge">addiu sp, sp, 24</code> after <code class="language-plaintext highlighter-rouge">jr ra</code>. On MIPS, the instruction after a jump always executes (the <strong>branch delay slot</strong>). The stack cleanup happens “while” the CPU is jumping back.</p>

<p><strong>3. Function call</strong> – “Call another function”</p>

<figure class="highlight"><pre><code class="language-asm" data-lang="asm">jal     80000b74          # Jump And Link — call function at 0x80000b74
                          # Sets ra = address of instruction after delay slot</code></pre></figure>

<p><code class="language-plaintext highlighter-rouge">JAL</code> is the MIPS “call” instruction. It saves the return address in register <code class="language-plaintext highlighter-rouge">$ra</code> and jumps to the target. Every <code class="language-plaintext highlighter-rouge">jal</code> in the binary is a function call.</p>

<p><strong>4. Load a 32-bit address</strong> – “Point at something”</p>

<figure class="highlight"><pre><code class="language-asm" data-lang="asm">lui     a1, 0x8000        # Load Upper Immediate: a1 = 0x80000000
addiu   a1, a1, 0x20e0    # Add lower 16 bits:    a1 = 0x800020e0</code></pre></figure>

<p>MIPS instructions are 32 bits wide, so you can’t load a 32-bit address in one instruction. The compiler uses a <code class="language-plaintext highlighter-rouge">LUI</code>/<code class="language-plaintext highlighter-rouge">ADDIU</code> pair: <code class="language-plaintext highlighter-rouge">LUI</code> loads the upper 16 bits, <code class="language-plaintext highlighter-rouge">ADDIU</code> adds the lower 16. This two-instruction pattern is how every string reference, every global variable access, and every table pointer gets loaded.</p>

<p><strong>5. Register convention</strong> – “Who’s who”</p>

<table>
  <thead>
    <tr>
      <th>Register</th>
      <th>Name</th>
      <th>Purpose</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">$a0-$a3</code></td>
      <td>Arguments</td>
      <td>First 4 function arguments</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">$v0-$v1</code></td>
      <td>Values</td>
      <td>Return values</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">$ra</code></td>
      <td>Return Address</td>
      <td>Where to return after <code class="language-plaintext highlighter-rouge">jr ra</code></td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">$sp</code></td>
      <td>Stack Pointer</td>
      <td>Current stack top</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">$s0-$s7</code></td>
      <td>Saved</td>
      <td>Callee-saved (preserved across calls)</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">$t0-$t9</code></td>
      <td>Temporary</td>
      <td>Caller-saved (may be clobbered by calls)</td>
    </tr>
  </tbody>
</table>

<p>That’s it. With these five patterns you can read any MIPS disassembly listing well enough to trace program flow.</p>

<hr />

<h2 id="from-strings-to-code-finding-where-strings-live">From Strings to Code: Finding Where Strings Live</h2>

<p>We found 111 strings. Now we need to connect them to the code that uses them. The question: which function prints “system starting”?</p>

<p>One convention before we go further. The flat binary starts at file offset <code class="language-plaintext highlighter-rouge">0</code> but the firmware was linked to run at <code class="language-plaintext highlighter-rouge">0x80000000</code> (we set that base in the Part 1 linker script – it’s KSEG0, the canonical bootable region on MIPS32). So <strong>runtime address = file offset + 0x80000000</strong>, and that’s the only translation we’ll need for the rest of the post. When we say “the string at <code class="language-plaintext highlighter-rouge">0x800020E0</code>,” that’s the same byte as “file offset <code class="language-plaintext highlighter-rouge">0x20E0</code>.”</p>

<p>Here’s the key insight. When the compiler generates code that references a string, it emits a <code class="language-plaintext highlighter-rouge">LUI</code>/<code class="language-plaintext highlighter-rouge">ADDIU</code> pair to load the string’s address into a register. If we know the string’s address in the binary, we can search for the <code class="language-plaintext highlighter-rouge">LUI</code>/<code class="language-plaintext highlighter-rouge">ADDIU</code> pair that loads it.</p>

<p>Let’s trace it manually. The string “system starting” is at file offset <code class="language-plaintext highlighter-rouge">0x20E0</code> in our binary:</p>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash"><span class="nv">$ </span>strings <span class="nt">-t</span> x firmware.bin | <span class="nb">grep</span> <span class="s2">"system starting"</span>
   20e0 system starting</code></pre></figure>

<p>Since our binary is loaded at <code class="language-plaintext highlighter-rouge">0x80000000</code>, the runtime address of this string is <code class="language-plaintext highlighter-rouge">0x800020E0</code>. To load this address, the compiler emits:</p>

<figure class="highlight"><pre><code class="language-asm" data-lang="asm">lui     a1, 0x8000        # a1 = 0x80000000  (upper 16 bits)
addiu   a1, a1, 0x20e0    # a1 = 0x800020E0  (add lower 16 bits)</code></pre></figure>

<p>We can automate this directly against the instruction encoding – <code class="language-plaintext highlighter-rouge">LUI</code> is opcode <code class="language-plaintext highlighter-rouge">0x0F</code> (top six bits) and <code class="language-plaintext highlighter-rouge">ADDIU</code> is opcode <code class="language-plaintext highlighter-rouge">0x09</code>, so each instruction is a 32-bit word with a known shape, and pattern-matching the bytes is enough to find every pair:</p>

<figure class="highlight"><pre><code class="language-python" data-lang="python"><span class="kn">import</span> <span class="n">struct</span>

<span class="n">BASE</span> <span class="o">=</span> <span class="mh">0x80000000</span>

<span class="k">def</span> <span class="nf">find_string_xrefs</span><span class="p">(</span><span class="n">data</span><span class="p">,</span> <span class="n">string_addr</span><span class="p">):</span>
    <span class="sh">"""</span><span class="s">Find all LUI/ADDIU pairs that load a given address.</span><span class="sh">"""</span>
    <span class="n">hi16</span> <span class="o">=</span> <span class="p">(</span><span class="n">string_addr</span> <span class="o">&gt;&gt;</span> <span class="mi">16</span><span class="p">)</span> <span class="o">&amp;</span> <span class="mh">0xFFFF</span>
    <span class="n">lo16</span> <span class="o">=</span> <span class="n">string_addr</span> <span class="o">&amp;</span> <span class="mh">0xFFFF</span>

    <span class="c1"># Sign extension: if lo16 &gt;= 0x8000, LUI loads hi16+1
</span>    <span class="k">if</span> <span class="n">lo16</span> <span class="o">&gt;=</span> <span class="mh">0x8000</span><span class="p">:</span>
        <span class="n">lui_imm</span> <span class="o">=</span> <span class="p">(</span><span class="n">hi16</span> <span class="o">+</span> <span class="mi">1</span><span class="p">)</span> <span class="o">&amp;</span> <span class="mh">0xFFFF</span>
    <span class="k">else</span><span class="p">:</span>
        <span class="n">lui_imm</span> <span class="o">=</span> <span class="n">hi16</span>

    <span class="n">xrefs</span> <span class="o">=</span> <span class="p">[]</span>
    <span class="k">for</span> <span class="n">reg</span> <span class="ow">in</span> <span class="nf">range</span><span class="p">(</span><span class="mi">1</span><span class="p">,</span> <span class="mi">32</span><span class="p">):</span>  <span class="c1"># search all registers
</span>        <span class="n">lui_word</span> <span class="o">=</span> <span class="mh">0x3C000000</span> <span class="o">|</span> <span class="p">(</span><span class="n">reg</span> <span class="o">&lt;&lt;</span> <span class="mi">16</span><span class="p">)</span> <span class="o">|</span> <span class="n">lui_imm</span>
        <span class="n">lui_bytes</span> <span class="o">=</span> <span class="n">struct</span><span class="p">.</span><span class="nf">pack</span><span class="p">(</span><span class="sh">'</span><span class="s">&lt;I</span><span class="sh">'</span><span class="p">,</span> <span class="n">lui_word</span><span class="p">)</span>

        <span class="n">pos</span> <span class="o">=</span> <span class="mi">0</span>
        <span class="k">while</span> <span class="bp">True</span><span class="p">:</span>
            <span class="n">pos</span> <span class="o">=</span> <span class="n">data</span><span class="p">.</span><span class="nf">find</span><span class="p">(</span><span class="n">lui_bytes</span><span class="p">,</span> <span class="n">pos</span><span class="p">)</span>
            <span class="k">if</span> <span class="n">pos</span> <span class="o">==</span> <span class="o">-</span><span class="mi">1</span><span class="p">:</span>
                <span class="k">break</span>
            <span class="c1"># Search forward for matching ADDIU within 64 instructions
</span>            <span class="n">addiu_word</span> <span class="o">=</span> <span class="mh">0x24000000</span> <span class="o">|</span> <span class="p">(</span><span class="n">reg</span> <span class="o">&lt;&lt;</span> <span class="mi">21</span><span class="p">)</span> <span class="o">|</span> <span class="p">(</span><span class="n">reg</span> <span class="o">&lt;&lt;</span> <span class="mi">16</span><span class="p">)</span> <span class="o">|</span> <span class="p">(</span><span class="n">lo16</span> <span class="o">&amp;</span> <span class="mh">0xFFFF</span><span class="p">)</span>
            <span class="n">addiu_bytes</span> <span class="o">=</span> <span class="n">struct</span><span class="p">.</span><span class="nf">pack</span><span class="p">(</span><span class="sh">'</span><span class="s">&lt;I</span><span class="sh">'</span><span class="p">,</span> <span class="n">addiu_word</span><span class="p">)</span>
            <span class="n">addiu_pos</span> <span class="o">=</span> <span class="n">data</span><span class="p">.</span><span class="nf">find</span><span class="p">(</span><span class="n">addiu_bytes</span><span class="p">,</span> <span class="n">pos</span><span class="p">,</span> <span class="n">pos</span> <span class="o">+</span> <span class="mi">256</span><span class="p">)</span>
            <span class="k">if</span> <span class="n">addiu_pos</span> <span class="o">!=</span> <span class="o">-</span><span class="mi">1</span><span class="p">:</span>
                <span class="n">xrefs</span><span class="p">.</span><span class="nf">append</span><span class="p">((</span><span class="n">BASE</span> <span class="o">+</span> <span class="n">pos</span><span class="p">,</span> <span class="n">BASE</span> <span class="o">+</span> <span class="n">addiu_pos</span><span class="p">))</span>
            <span class="n">pos</span> <span class="o">+=</span> <span class="mi">4</span>

    <span class="k">return</span> <span class="n">xrefs</span></code></pre></figure>

<p>The sign-extension detail matters more than it looks: MIPS <code class="language-plaintext highlighter-rouge">ADDIU</code> sign-extends its 16-bit immediate. If the lower half of the address is &gt;= <code class="language-plaintext highlighter-rouge">0x8000</code>, the <code class="language-plaintext highlighter-rouge">LUI</code> has to load one more than the upper half to compensate. Get this wrong and you’ll miss half your cross-references. (In the real project, getting this right was the difference between finding 2,000 strings and finding all 4,471.)</p>

<p>The runnable script in the bundle (<code class="language-plaintext highlighter-rouge">tools/find_string_xrefs.py</code>) takes the same approach but uses <a href="https://www.capstone-engine.org/">Capstone</a> to decode instructions instead of matching byte patterns – mostly so it can track register state across the LUI/ADDIU window and reject pairs where an intervening write has invalidated the upper half. The companion scripts <code class="language-plaintext highlighter-rouge">tools/find_word.py</code>, <code class="language-plaintext highlighter-rouge">tools/find_function_prologues.py</code>, and <code class="language-plaintext highlighter-rouge">tools/find_cli_tables.py</code> cover the data-pointer scan and the prologue/table walks we’ll use in §5 and §6. (<code class="language-plaintext highlighter-rouge">apt install python3-capstone</code> on Debian-family systems, or <code class="language-plaintext highlighter-rouge">pip install capstone</code> inside a venv elsewhere.)</p>

<hr />

<h2 id="finding-function-boundaries">Finding Function Boundaries</h2>

<p>Now that we can locate where a string is referenced, we need to find which <strong>function</strong> contains that reference. The answer: scan backward for a function prologue.</p>

<figure class="highlight"><pre><code class="language-python" data-lang="python"><span class="k">def</span> <span class="nf">find_function_prologues</span><span class="p">(</span><span class="n">data</span><span class="p">):</span>
    <span class="sh">"""</span><span class="s">Find all ADDIU SP, SP, -N instructions (function starts).</span><span class="sh">"""</span>
    <span class="n">prologues</span> <span class="o">=</span> <span class="p">[]</span>
    <span class="k">for</span> <span class="n">off</span> <span class="ow">in</span> <span class="nf">range</span><span class="p">(</span><span class="mi">0</span><span class="p">,</span> <span class="nf">len</span><span class="p">(</span><span class="n">data</span><span class="p">)</span> <span class="o">-</span> <span class="mi">3</span><span class="p">,</span> <span class="mi">4</span><span class="p">):</span>
        <span class="n">word</span> <span class="o">=</span> <span class="n">struct</span><span class="p">.</span><span class="nf">unpack_from</span><span class="p">(</span><span class="sh">'</span><span class="s">&lt;I</span><span class="sh">'</span><span class="p">,</span> <span class="n">data</span><span class="p">,</span> <span class="n">off</span><span class="p">)[</span><span class="mi">0</span><span class="p">]</span>
        <span class="nf">if </span><span class="p">(</span><span class="n">word</span> <span class="o">&gt;&gt;</span> <span class="mi">16</span><span class="p">)</span> <span class="o">==</span> <span class="mh">0x27BD</span><span class="p">:</span>           <span class="c1"># ADDIU $sp, $sp, imm
</span>            <span class="n">imm</span> <span class="o">=</span> <span class="n">word</span> <span class="o">&amp;</span> <span class="mh">0xFFFF</span>
            <span class="k">if</span> <span class="n">imm</span> <span class="o">&gt;=</span> <span class="mh">0x8000</span><span class="p">:</span>                <span class="c1"># negative = stack allocation
</span>                <span class="n">frame_size</span> <span class="o">=</span> <span class="n">imm</span> <span class="o">-</span> <span class="mh">0x10000</span>
                <span class="n">prologues</span><span class="p">.</span><span class="nf">append</span><span class="p">((</span><span class="n">BASE</span> <span class="o">+</span> <span class="n">off</span><span class="p">,</span> <span class="n">frame_size</span><span class="p">))</span>
    <span class="k">return</span> <span class="n">prologues</span></code></pre></figure>

<p>Running this on our firmware.bin:</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">Found 35 function prologues:
  0x80000248: addiu sp, sp, -32
  0x800002F0: addiu sp, sp, -32
  0x80000354: addiu sp, sp, -40
  ...
  0x80000EFC: addiu sp, sp, -48    ← iterate_active_routes?
  ...
  0x800015C8: addiu sp, sp, -184   ← big function (CLI main loop?)
  0x800016B0: addiu sp, sp, -24    ← last function (firmware_main?)</code></pre></figure>

<p>35 function prologues. The real count is higher – leaf functions that never call anything else don’t need to save <code class="language-plaintext highlighter-rouge">$ra</code> and often skip the standard prologue entirely, so they slip past this scan. But 35 is enough to map the major functions, and that’s what we need to follow the strings.</p>

<p>We can also count <code class="language-plaintext highlighter-rouge">JAL</code> calls to build a call graph:</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">JAL call frequency:
  0x800002E4: called 111 times   ← most-called function (uart_puts?)
  0x80000530: called  15 times   ← probably uart_puthex
  0x80000790: called  15 times   ← probably uart_putdec
  0x80000B74: called  12 times   ← log_msg (used in every subsystem)
  0x800002C4: called   6 times   ← uart_putchar
  0x80000AA8: called   4 times   ← fw_assert
  0x80000FEC: called   4 times   ← route_add (called 4 times in main!)</code></pre></figure>

<p>The function at <code class="language-plaintext highlighter-rouge">0x800002E4</code>, called 111 times, is almost certainly the UART print function – every command handler, every log message, every error path calls it. That matches what we’d expect from looking at the source filenames: a CLI-heavy firmware spends most of its time printing strings.</p>

<hr />

<h2 id="the-cli-table-following-pointer-chains">The CLI Table: Following Pointer Chains</h2>

<p>Now for the detective work. We know the firmware has CLI commands because <code class="language-plaintext highlighter-rouge">strings</code> found “help”, “version”, “routedump”, and friends. But strings are just bytes – they don’t dispatch themselves. When the user types <code class="language-plaintext highlighter-rouge">help</code>, some function has to compare the input against every known command name and call the matching handler.</p>

<p>Here’s where the xref scanner from §4 actually pays off. Pick a command-name string and ask: which <code class="language-plaintext highlighter-rouge">LUI</code>/<code class="language-plaintext highlighter-rouge">ADDIU</code> pair in the binary loads its address? For <code class="language-plaintext highlighter-rouge">"system starting"</code> (<code class="language-plaintext highlighter-rouge">0x800020e0</code>) the answer is one xref at <code class="language-plaintext highlighter-rouge">0x800016bc</code> – some function loads it and prints it. But for <code class="language-plaintext highlighter-rouge">"help"</code> (<code class="language-plaintext highlighter-rouge">0x800018c0</code>)?</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">$ python3 tools/find_string_xrefs.py firmware.bin 0x800018c0
LUI/ADDIU pairs loading 0x800018c0: 0</code></pre></figure>

<p>Zero. Same for every other command name in both clusters. Yet the strings clearly exist – <code class="language-plaintext highlighter-rouge">strings</code> found them. So how is the firmware getting at them?</p>

<p>A second scanner answers the complementary question: scan the binary for any 32-bit word that <em>equals</em> the address we’re looking for, regardless of how it got there.</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">$ python3 tools/find_word.py firmware.bin 0x800018c0
0x800018c8: 32-bit word == 0x800018c0</code></pre></figure>

<p>The address of <code class="language-plaintext highlighter-rouge">"help"</code> shows up as a <em>pointer</em>, sitting in <code class="language-plaintext highlighter-rouge">.rodata</code> itself, eight bytes after the string. No code loads it; data points at it. Combine the two scanners and you get a sharp diagnostic for any string in the binary:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: right">code xrefs</th>
      <th style="text-align: right">data hits</th>
      <th>what the string is</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">&gt; 0</code></td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">0</code></td>
      <td>passed to a function (e.g. <code class="language-plaintext highlighter-rouge">"system starting"</code>)</td>
    </tr>
    <tr>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">0</code></td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">&gt;= 1</code></td>
      <td>referenced through a table (e.g. <code class="language-plaintext highlighter-rouge">"help"</code> in a CLI table)</td>
    </tr>
    <tr>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">0</code></td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">0</code></td>
      <td>dead string, never used at runtime</td>
    </tr>
  </tbody>
</table>

<p><code class="language-plaintext highlighter-rouge">"help"</code> lands in the second row. So do all 16 command names. That’s the signal: when a string has no code xref but its address appears as a 32-bit word inside <code class="language-plaintext highlighter-rouge">.rodata</code>, it’s referenced through a <strong>table</strong>, not printed inline.</p>

<p>The natural shape of such a table is an array of structs, one per command:</p>

<figure class="highlight"><pre><code class="language-c" data-lang="c"><span class="k">struct</span> <span class="n">cli_entry</span> <span class="p">{</span>
    <span class="k">const</span> <span class="kt">char</span> <span class="o">*</span><span class="n">name</span><span class="p">;</span>       <span class="cm">/* "help", "version", etc. */</span>
    <span class="kt">void</span> <span class="p">(</span><span class="o">*</span><span class="n">handler</span><span class="p">)(</span><span class="kt">int</span><span class="p">,</span> <span class="k">const</span> <span class="kt">char</span> <span class="o">**</span><span class="p">);</span>   <span class="cm">/* function pointer */</span>
    <span class="k">const</span> <span class="kt">char</span> <span class="o">*</span><span class="n">help</span><span class="p">;</span>       <span class="cm">/* "Show this help message" */</span>
<span class="p">};</span></code></pre></figure>

<p>That’s 12 bytes per entry (three 32-bit pointers). We can scan <code class="language-plaintext highlighter-rouge">.rodata</code> for sequences of 12-byte entries where the first and third words are pointers to printable strings, and the second word is a pointer into the <code class="language-plaintext highlighter-rouge">.text</code> section.</p>

<figure class="highlight"><pre><code class="language-python" data-lang="python"><span class="c1"># Scan .rodata for CLI table pattern: [string_ptr, code_ptr, string_ptr] × N
</span><span class="k">for</span> <span class="n">off</span> <span class="ow">in</span> <span class="nf">range</span><span class="p">(</span><span class="n">rodata_start</span><span class="p">,</span> <span class="n">rodata_end</span><span class="p">,</span> <span class="mi">12</span><span class="p">):</span>
    <span class="n">name_ptr</span>  <span class="o">=</span> <span class="nf">read_u32</span><span class="p">(</span><span class="n">off</span><span class="p">)</span>
    <span class="n">handler</span>   <span class="o">=</span> <span class="nf">read_u32</span><span class="p">(</span><span class="n">off</span> <span class="o">+</span> <span class="mi">4</span><span class="p">)</span>
    <span class="n">help_ptr</span>  <span class="o">=</span> <span class="nf">read_u32</span><span class="p">(</span><span class="n">off</span> <span class="o">+</span> <span class="mi">8</span><span class="p">)</span>

    <span class="nf">if </span><span class="p">(</span><span class="nf">is_string_pointer</span><span class="p">(</span><span class="n">name_ptr</span><span class="p">)</span> <span class="ow">and</span>
        <span class="nf">is_code_pointer</span><span class="p">(</span><span class="n">handler</span><span class="p">)</span> <span class="ow">and</span>
        <span class="nf">is_string_pointer</span><span class="p">(</span><span class="n">help_ptr</span><span class="p">)):</span>
        <span class="c1"># Found a table entry!</span></code></pre></figure>

<p>Running this on our firmware binary, we find <strong>two</strong> tables:</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">Table at 0x800018C8 (10 entries):
  "help"         handler=0x80000354  "Show this help message"
  "version"      handler=0x800003FC  "Show firmware version"
  "uptime"       handler=0x80000838  "Show system uptime"
  "status"       handler=0x80000874  "Show system status"
  "portdump"     handler=0x80000E18  "Dump port status [port_num]"
  "routedump"    handler=0x800008FC  "Dump routing table"
  "reset"        handler=0x8000047C  "Reset the system"
  "loopback"     handler=0x8000135C  "Run loopback test &lt;port&gt;"
  "memtest"      handler=0x800012A4  "Run memory test"
  "log"          handler=0x800004A4  "Show recent log entries"

Table at 0x80001A04 (6 entries):
  "dbglvl"       handler=0x800009FC  "Set debug verbosity level"
  "peek"         handler=0x800005E4  "Read memory address"
  "poke"         handler=0x800004F4  "Write memory address"
  "flashid"      handler=0x800010C4  "Read SPI flash JEDEC ID"
  "crashme"      handler=0x80000B40  "Trigger deliberate crash"
  "regdump"      handler=0x800006CC  "Dump hardware registers"</code></pre></figure>

<p>Two tables – one with 10 entries, one with 6. Now the obvious question: if all 16 commands are in the binary, why does <code class="language-plaintext highlighter-rouge">help</code> only show 10 of them?</p>

<p>Run the xref scanner against each table’s base address:</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">$ python3 tools/find_string_xrefs.py firmware.bin 0x800018c8   # public table
LUI/ADDIU pairs loading 0x800018c8: 3
LUI/ADDIU @ 0x80000388 (inside function 0x80000354)
LUI/ADDIU @ 0x80001468 (inside function 0x80001424)
LUI/ADDIU @ 0x800014f4 (inside function 0x80001424)

$ python3 tools/find_string_xrefs.py firmware.bin 0x80001a04   # hidden table
LUI/ADDIU pairs loading 0x80001a04: 1
LUI/ADDIU @ 0x80001498 (inside function 0x80001424)</code></pre></figure>

<p>Two functions reference the public table; only one references the hidden table – and the public-only function is the asymmetry the help command rests on.</p>

<p>The function at <code class="language-plaintext highlighter-rouge">0x80000354</code> is the public-only one. From the table itself we know it’s <code class="language-plaintext highlighter-rouge">cmd_help</code>, the handler for the <code class="language-plaintext highlighter-rouge">help</code> command: a small function that walks one array and prints its entries. The function at <code class="language-plaintext highlighter-rouge">0x80001424</code> references <strong>both</strong> tables – that’s the dispatcher: read a line of input, walk the public table looking for a match, then walk the hidden table, and call the handler if either hit. (The dispatcher loads the public table address twice – once when the search loop sets up its iterator, then again after the matched handler returns and the table base needs to be re-fetched to read the function pointer. The compiler didn’t bother keeping the address live across the call. Routine scheduling, not a second logical reference.)</p>

<p>Nothing was “removed.” The 6 commands in the second table aren’t dead code, aren’t disabled, aren’t gated behind a flag. They just don’t get listed by <code class="language-plaintext highlighter-rouge">cmd_help</code>, because <code class="language-plaintext highlighter-rouge">cmd_help</code> only iterates one of the two arrays the dispatcher accepts. Whoever wrote the firmware wanted <code class="language-plaintext highlighter-rouge">peek</code>, <code class="language-plaintext highlighter-rouge">poke</code>, <code class="language-plaintext highlighter-rouge">crashme</code>, and friends reachable from the prompt without advertising them in the help output – and they implemented that with the simplest possible mechanism: two arrays, one printed, both dispatched.</p>

<p>In the real vendor firmware, the same xref-the-table-base trick discovered <strong>148 CLI commands</strong> across multiple dispatch tables – including handlers that could read and write arbitrary hardware registers, dump internal routing tables, and trigger diagnostic modes not documented in any user manual. None of them were hidden. They were just iterated by a different function than the one printing the help text. The strings gave every single one away the moment we asked which functions referenced each table.</p>

<hr />

<h2 id="putting-it-together-the-subsystem-map">Putting It Together: The Subsystem Map</h2>

<p>It’s worth naming what we’ve actually been doing for the last few sections. A “symbol,” in the linker’s sense, is a triple: a <strong>name</strong>, an <strong>address</strong>, and a <strong>kind</strong> (function, data, etc.). The flat binary on disk has none of them – they were dropped on the way from ELF to <code class="language-plaintext highlighter-rouge">objcopy -O binary</code>. What we’ve been doing is rebuilding that triple ourselves: prologue scanning gave us addresses, JAL-frequency and string xrefs let us guess the kind (UART printer, log function, command handler, leaf), and the strings themselves – <code class="language-plaintext highlighter-rouge">"help"</code>, <code class="language-plaintext highlighter-rouge">"system starting"</code>, the source filenames – gave us names. None of them are real symbols. All of them are good enough to navigate by. That’s the move that makes static RE work on stripped firmware: when the symbol table isn’t there, you reconstruct it.</p>

<p>Starting from nothing but a flat binary and the <code class="language-plaintext highlighter-rouge">strings</code> command, we’ve built a surprisingly complete picture of this firmware:</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">Firmware Subsystem Map (built from static analysis only)
========================================================

Identity:   NetGW-MIPS32 v2.4.1-rc3, network gateway device
Source:     6 modules (main.c, cli.c, port.c, route.c, diag.c, flash.c)
Interface:  UART serial console, prompt "gw&gt;"
Functions:  35 prologue-detected (plus an unknown number of leaf functions
            that don't save $ra and slip past the standard prologue scan)
Strings:    111 embedded
Commands:   16 total — all 16 dispatched, only 10 listed by cmd_help

Module Map:
  MOD_CORE  (1) — System init, boot sequence
  MOD_CLI   (2) — Command dispatch, input processing
  MOD_PORT  (3) — 8-port forwarding engine
  MOD_ROUTE (4) — Route table management
  MOD_DMA   (5) — DMA engine (referenced but minimal)
  MOD_DIAG  (6) — Diagnostic tests (loopback, memtest)
  MOD_FLASH (7) — SPI flash operations

Suspicious:
  - "WARNING: no routes programmed for stack" during boot
  - "bridge flag set, skipping" — routes being skipped?
  - "bridge routes deferred to bridge handler" — are they ever handled?</code></pre></figure>

<p>All of this without running the firmware, without a disassembler GUI, and without symbols. Just <code class="language-plaintext highlighter-rouge">strings</code>, some Python scripting, and pattern recognition.</p>

<p>But those suspicious messages are nagging. The boot sequence says routes are being skipped because of a “bridge flag,” and then warns that no routes were programmed. That sounds like a bug – routes that should be programmed into the forwarding table are getting silently dropped.</p>

<p>In Part 3, we’ll throw this binary at Ghidra to decompile it into readable C, trace the execution path through the route engine function by function, and find exactly where – and why – those routes are disappearing. Then we’ll fix it. With a hex editor.</p>

<hr />

<h2 id="references">References</h2>

<ul>
  <li><a href="https://res.cloudinary.com/gotocco/raw/upload/v1777182332/sample_firmware.tar.gz">sample_firmware.tar.gz</a> – <code class="language-plaintext highlighter-rouge">firmware.bin</code>, build sources (<code class="language-plaintext highlighter-rouge">main.c</code>, <code class="language-plaintext highlighter-rouge">startup.S</code>, <code class="language-plaintext highlighter-rouge">linker.ld</code>, <code class="language-plaintext highlighter-rouge">Makefile</code>), and the four Python scripts used in this post under <code class="language-plaintext highlighter-rouge">tools/</code>:
    <ul>
      <li><code class="language-plaintext highlighter-rouge">find_string_xrefs.py</code> – find LUI/ADDIU pairs that load a given address (uses Capstone)</li>
      <li><code class="language-plaintext highlighter-rouge">find_word.py</code> – find a 32-bit word stored as data anywhere in the binary</li>
      <li><code class="language-plaintext highlighter-rouge">find_function_prologues.py</code> – find function starts via the <code class="language-plaintext highlighter-rouge">addiu sp, sp, -N</code> pattern (uses Capstone)</li>
      <li><code class="language-plaintext highlighter-rouge">find_cli_tables.py</code> – scan <code class="language-plaintext highlighter-rouge">.rodata</code> for <code class="language-plaintext highlighter-rouge">{name_ptr, code_ptr, help_ptr}</code> runs</li>
    </ul>
  </li>
  <li><a href="https://www.capstone-engine.org/">Capstone Engine</a> – multi-architecture disassembly framework. <code class="language-plaintext highlighter-rouge">apt install python3-capstone</code> (Debian/Ubuntu) or <code class="language-plaintext highlighter-rouge">pip install capstone</code> (in a venv elsewhere)</li>
  <li><a href="https://s3-eu-west-1.amazonaws.com/downloads-mips/documents/MD00086-2B-MIPS32BIS-AFP-6.06.pdf">MIPS32 Architecture For Programmers Vol. II</a> – Instruction set reference (LUI, ADDIU, JAL encoding)</li>
  <li><a href="https://en.wikipedia.org/wiki/Calling_convention#MIPS">MIPS Calling Convention</a> – Register usage: <code class="language-plaintext highlighter-rouge">$a0-$a3</code>, <code class="language-plaintext highlighter-rouge">$v0</code>, <code class="language-plaintext highlighter-rouge">$ra</code>, <code class="language-plaintext highlighter-rouge">$sp</code></li>
  <li><code class="language-plaintext highlighter-rouge">strings(1)</code> – GNU binutils string finder; use <code class="language-plaintext highlighter-rouge">-t x</code> for hex offsets</li>
  <li><a href="/reverse-engineering/firmware/mips/2026/04/22/headers-symbols-and-other-things-you-wont-find.html">Part 1: Headers, Symbols, and other things you won’t find</a> – Building the sample firmware</li>
</ul>]]></content><author><name>Maciej Grochowski</name></author><category term="reverse-engineering" /><category term="firmware" /><category term="mips" /><summary type="html"><![CDATA[In Part 1, we built a flat MIPS32 firmware binary and stared at 8,496 bytes with no headers, no symbols, and no sections. file said “data.” readelf said nothing. We left off with a question: now what?]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://page-fault.io/assets/social-preview.png" /><media:content medium="image" url="https://page-fault.io/assets/social-preview.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Headers, Symbols, and other things you won’t find</title><link href="https://page-fault.io/reverse-engineering/firmware/mips/2026/04/22/headers-symbols-and-other-things-you-wont-find.html" rel="alternate" type="text/html" title="Headers, Symbols, and other things you won’t find" /><published>2026-04-22T12:00:00+00:00</published><updated>2026-04-22T12:00:00+00:00</updated><id>https://page-fault.io/reverse-engineering/firmware/mips/2026/04/22/headers-symbols-and-other-things-you-wont-find</id><content type="html" xml:base="https://page-fault.io/reverse-engineering/firmware/mips/2026/04/22/headers-symbols-and-other-things-you-wont-find.html"><![CDATA[<p>Eight years ago I wrote <a href="/kernel/2018/07/28/Elfs_Linkers_Other.html">ELF’s Linker’s and other magical creatures</a> – a walkthrough of the ELF binary format, relocations, segments, and even live code injection through <code class="language-plaintext highlighter-rouge">/proc/pid/mem</code>. That post ended in a comfortable place: <code class="language-plaintext highlighter-rouge">.text</code>, <code class="language-plaintext highlighter-rouge">.data</code>, <code class="language-plaintext highlighter-rouge">.bss</code>; <code class="language-plaintext highlighter-rouge">ld</code> resolving symbols; the kernel’s ELF loader mapping segments into memory; <code class="language-plaintext highlighter-rouge">gdb</code> poking at a running process. The civilized world of userspace binaries.</p>

<p>Then someone hands you a 1.5 MB file. No extension, no headers, no documentation. <code class="language-plaintext highlighter-rouge">file</code> says “data.” <code class="language-plaintext highlighter-rouge">readelf</code> says “not an ELF file.” <code class="language-plaintext highlighter-rouge">objdump</code> refuses to open it. Welcome to firmware reverse engineering.</p>

<p>This is Part 1 of a five-part series that picks up where the ELF post left off, but for bare-metal firmware on MIPS32 embedded devices. Everything you learned about ELF structure, symbol tables, and linking is still useful – mostly as a reference for what <em>isn’t</em> in front of you anymore.</p>

<p>The techniques come from a real project analyzing production firmware on an embedded network device. I can’t name the vendor (lawyers), but every pattern, every tool, every trick here is something we actually used in anger. To keep it reproducible, we’ll build a sample MIPS32 firmware – a small network gateway with a UART serial console – that recreates the same structures.</p>

<blockquote>
  <p><strong>Bonus – the sample firmware bundled:</strong> if you’d rather skip the build and start poking at bytes, the same sources, the resulting <code class="language-plaintext highlighter-rouge">firmware.bin</code> and <code class="language-plaintext highlighter-rouge">firmware.elf</code>, the multi-partition image, and the small <code class="language-plaintext highlighter-rouge">tools/</code> directory used in Part 2 are all bundled in <a href="https://res.cloudinary.com/gotocco/raw/upload/v1777182332/sample_firmware.tar.gz">sample_firmware.tar.gz</a>. Grab it and follow along on your own copy:</p>

  <div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>curl <span class="nt">-L</span> <span class="nt">-O</span> https://res.cloudinary.com/gotocco/raw/upload/v1777182332/sample_firmware.tar.gz
<span class="nb">tar </span>xf sample_firmware.tar.gz
<span class="nb">cd </span>sample_firmware
</code></pre></div>  </div>
</blockquote>

<hr />

<h2 id="whats-missing-elf-vs-flat-binary">What’s Missing: ELF vs Flat Binary</h2>

<p>The quickest way to understand flat firmware is to compare it with what you already know. Here’s what a standard ELF binary looks like at byte zero:</p>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash"><span class="nv">$ </span>xxd <span class="nt">-l</span> 64 firmware.elf
00000000: 7f45 4c46 0101 0100 0500 0000 0000 0000  .ELF............
00000010: 0200 0800 0100 0000 0000 0080 3400 0000  ............4...
00000020: a431 0100 0110 0070 3400 2000 0500 2800  .1.....p4. ...<span class="o">(</span><span class="nb">.</span>
00000030: 0d00 0c00 0300 0070 0821 0100 0821 0080  .......p.!...!..</code></pre></figure>

<p>You can spot the <code class="language-plaintext highlighter-rouge">7F 45 4C 46</code> magic immediately – <code class="language-plaintext highlighter-rouge">.ELF</code>. After that comes the class (32-bit), endianness (little), the machine type (MIPS), the entry point, program header offset, section header offset. Everything the kernel’s loader needs to map this binary into memory.</p>

<p>Now here’s what a firmware binary looks like at byte zero:</p>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash"><span class="nv">$ </span>xxd <span class="nt">-l</span> 64 firmware.bin
00000000: 0580 1d3c 0000 bd37 0480 1c3c 0000 9c37  ...&lt;...7...&lt;...7
00000010: 8000 0008 0000 0000 0000 0000 0000 0000  ................
00000020: 0000 0000 0000 0000 0000 0000 0000 0000  ................
00000030: 0000 0000 0000 0000 0000 0000 0000 0000  ................</code></pre></figure>

<p>No magic bytes. No headers. The very first byte of this file <em>is</em> a machine instruction. If you happen to read MIPS, you can decode it on sight: <code class="language-plaintext highlighter-rouge">lui sp, 0x8005</code> – loading the upper half of the stack pointer, the canonical first move of a freshly-booted CPU. Once the boot ROM (or first-stage bootloader) copies the payload to RAM and jumps to its entry address, this is where execution begins. (Register conventions come in Part 2. For now, <code class="language-plaintext highlighter-rouge">sp</code>, <code class="language-plaintext highlighter-rouge">ra</code>, <code class="language-plaintext highlighter-rouge">a0..a3</code>, and <code class="language-plaintext highlighter-rouge">s0..s7</code> are the names you’ll keep seeing.)</p>

<p>Here’s what each format gives you:</p>

<table>
  <thead>
    <tr>
      <th>Feature</th>
      <th>ELF</th>
      <th>Flat firmware</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Magic bytes</td>
      <td><code class="language-plaintext highlighter-rouge">7F 45 4C 46</code></td>
      <td>None (first instruction)</td>
    </tr>
    <tr>
      <td>Section headers</td>
      <td><code class="language-plaintext highlighter-rouge">.text</code>, <code class="language-plaintext highlighter-rouge">.data</code>, <code class="language-plaintext highlighter-rouge">.bss</code>, …</td>
      <td>No</td>
    </tr>
    <tr>
      <td>Symbol table</td>
      <td>Function names, variables</td>
      <td>No</td>
    </tr>
    <tr>
      <td>Relocation entries</td>
      <td>Linker fixups</td>
      <td>No</td>
    </tr>
    <tr>
      <td>Entry point</td>
      <td><code class="language-plaintext highlighter-rouge">e_entry</code> field</td>
      <td>Byte 0 (or known fixed address)</td>
    </tr>
    <tr>
      <td>Load method</td>
      <td>Kernel parses headers, maps segments</td>
      <td>Bootloader copies to RAM, jumps</td>
    </tr>
  </tbody>
</table>

<p>Remember <code class="language-plaintext highlighter-rouge">readelf -S</code> showing section headers? <code class="language-plaintext highlighter-rouge">nm</code> listing every function by name? <code class="language-plaintext highlighter-rouge">objdump -tT</code> dumping the symbol table? None of that works here. The flat binary has no metadata. The hex dump is the documentation.</p>

<hr />

<h2 id="why-firmware-is-flat">Why Firmware Is Flat</h2>

<p>This isn’t an accident or a limitation – it’s a deliberate engineering choice. Here’s why embedded firmware ships as raw binary blobs:</p>

<p><strong>Bootloader simplicity.</strong> The first code that runs when a MIPS chip powers on is a tiny bootrom, often a few hundred bytes burned into mask ROM. It does not have an ELF parser. It loads a fixed payload from flash into RAM and jumps to a known entry address – in our sample, <code class="language-plaintext highlighter-rouge">0x80000000</code>. Any complexity in the binary format becomes complexity in silicon, and silicon has to be right on the first wafer.</p>

<p><strong>Speed.</strong> A router that takes 30 seconds to come up after a power cycle is a router nobody deploys. Flat binaries boot in one <code class="language-plaintext highlighter-rouge">memcpy</code> – no parsing, no relocations, no dynamic linking. Copy the bytes, jump to the entry point, you’re running.</p>

<p><strong>Deterministic layout.</strong> The developer controls exactly where every byte lands in memory. The linker script says “put the exception vectors at <code class="language-plaintext highlighter-rouge">0x80000000</code>, put <code class="language-plaintext highlighter-rouge">.text</code> right after, put <code class="language-plaintext highlighter-rouge">.bss</code> at <code class="language-plaintext highlighter-rouge">0x80040000</code>,” and that is exactly what happens. No ASLR, no PIE, no loader surprises. (You miss this guarantee the first time you stop missing it.)</p>

<p><strong>The objcopy pipeline.</strong> Here’s how a flat binary is produced from source:</p>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash"><span class="c"># Step 1: Compile to object files</span>
mipsel-linux-gnu-gcc-14 <span class="nt">-mips32r2</span> <span class="nt">-EL</span> <span class="nt">-c</span> <span class="nt">-o</span> startup.o startup.S
mipsel-linux-gnu-gcc-14 <span class="nt">-mips32r2</span> <span class="nt">-EL</span> <span class="nt">-O1</span> <span class="nt">-ffreestanding</span> <span class="nt">-nostdlib</span> <span class="nt">-c</span> <span class="nt">-o</span> main.o main.c

<span class="c"># Step 2: Link into ELF (intermediate)</span>
mipsel-linux-gnu-ld <span class="nt">-T</span> linker.ld <span class="nt">-o</span> firmware.elf startup.o main.o

<span class="c"># Step 3: Strip to flat binary -- THIS is the key step</span>
mipsel-linux-gnu-objcopy <span class="nt">-O</span> binary firmware.elf firmware.bin</code></pre></figure>

<p>Step 3 is where everything disappears. <code class="language-plaintext highlighter-rouge">objcopy -O binary</code> takes the ELF, extracts only the loadable segments (the raw bytes that need to be in memory at runtime), and writes them sequentially to a file. Section headers? Gone. Symbol table? Gone. String table? Gone. Every piece of metadata that made the ELF navigable goes in the bin.</p>

<p>The developer still has the ELF – they need it to debug – but what ships to the customer is the flat binary. And when you are the one doing the reverse engineering, the flat binary is usually all you get.</p>

<hr />

<h2 id="memory-layout-where-everything-lives">Memory Layout: Where Everything Lives</h2>

<p>Every flat firmware binary has a base address – the memory address where the bootloader copies it. For MIPS32 devices, this is typically in KSEG0 (<code class="language-plaintext highlighter-rouge">0x80000000</code> - <code class="language-plaintext highlighter-rouge">0x9FFFFFFF</code>), which is cached, unmapped kernel space. Our sample firmware uses <code class="language-plaintext highlighter-rouge">0x80000000</code>:</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">Virtual Memory Map
==================

0x80000000 +---------------------------+
           | Exception vectors (640 B) |  Startup + exception handlers
0x80000280 +---------------------------+
           | .text (5,392 bytes)       |  All firmware functions
0x80001790 +---------------------------+
           | .rodata (2,396 bytes)     |  Strings, CLI tables, constants
0x800020EC +---------------------------+
           | (end of flash content)    |
           ...
0x80040000 +---------------------------+
           | .bss (1,024 bytes)        |  Globals (zeroed at boot)
0x80040400 +---------------------------+
           | (free RAM)               |
           ...
0x80050000 +---------------------------+
           | Stack top                 |  Stack grows downward
           +---------------------------+

MMIO Registers (not in binary):
  0xB8000000  UART (TX, RX, status, control)
  0xB8010000  GPIO (data, direction)
  0xB8020000  Timer (count, compare, control)
  0xB8030000  System (reset, clock, version)</code></pre></figure>

<p>For this sample, the address translation is trivial: once you’ve extracted the raw payload, any byte at file offset <code class="language-plaintext highlighter-rouge">N</code> lives at virtual address <code class="language-plaintext highlighter-rouge">0x80000000 + N</code>. No page tables, no ASLR, no segment mapping – one addition. If you find an interesting string at file offset <code class="language-plaintext highlighter-rouge">0x18CC</code>, you know it sits at <code class="language-plaintext highlighter-rouge">0x800018CC</code> in the running firmware.</p>

<p>Compare with ELF, where the loader reads program headers to figure out which file offset maps to which virtual address, with each segment at its own alignment. Here, you can do the math in your head.</p>

<hr />

<h2 id="build-your-own-the-sample-firmware">Build Your Own: The Sample Firmware</h2>

<p>Before reverse engineering bytes, it helps to watch where the bytes come from. The build is short, and once you’ve seen a flat binary fall out the end of <code class="language-plaintext highlighter-rouge">objcopy</code>, the rest of the series stops feeling like archaeology and starts feeling like accounting.</p>

<p>Our sample firmware simulates a small <strong>MIPS32 network gateway</strong> – think of an embedded router or bridge device managed over a UART serial console. This is a common pattern: millions of MIPS-based routers and switches run firmware exactly like this (Broadcom, MediaTek, and Qualcomm Atheros SoCs all use MIPS32 cores).</p>

<p>The firmware has:</p>

<ul>
  <li><strong>UART serial console</strong> with a command-line interface</li>
  <li><strong>8-port forwarding engine</strong> with route table management</li>
  <li><strong>Public commands</strong>: <code class="language-plaintext highlighter-rouge">help</code>, <code class="language-plaintext highlighter-rouge">version</code>, <code class="language-plaintext highlighter-rouge">uptime</code>, <code class="language-plaintext highlighter-rouge">status</code>, <code class="language-plaintext highlighter-rouge">portdump</code>, <code class="language-plaintext highlighter-rouge">routedump</code>, <code class="language-plaintext highlighter-rouge">reset</code>, <code class="language-plaintext highlighter-rouge">loopback</code>, <code class="language-plaintext highlighter-rouge">memtest</code>, <code class="language-plaintext highlighter-rouge">log</code></li>
  <li><strong>Hidden commands</strong>: <code class="language-plaintext highlighter-rouge">peek</code>, <code class="language-plaintext highlighter-rouge">poke</code>, <code class="language-plaintext highlighter-rouge">flashid</code>, <code class="language-plaintext highlighter-rouge">crashme</code>, <code class="language-plaintext highlighter-rouge">regdump</code>, <code class="language-plaintext highlighter-rouge">dbglvl</code> (more on these later)</li>
  <li><strong>MMIO register access</strong> for UART, GPIO, Timer, and System registers</li>
  <li><strong>Assert/logging framework</strong> with source filename references</li>
</ul>

<p>You’ll need the MIPS cross-compilation toolchain:</p>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash"><span class="c"># Debian/Ubuntu</span>
<span class="nb">sudo </span>apt <span class="nb">install </span>gcc-mipsel-linux-gnu binutils-mipsel-linux-gnu</code></pre></figure>

<p>Build and inspect:</p>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash"><span class="nv">$ </span>make clean <span class="o">&amp;&amp;</span> make
<span class="nb">rm</span> <span class="nt">-f</span> <span class="k">*</span>.o <span class="k">*</span>.elf firmware.bin
mipsel-linux-gnu-gcc-14 <span class="nt">-mips32r2</span> <span class="nt">-EL</span> <span class="nt">-fno-pic</span> <span class="nt">-mno-abicalls</span> <span class="nt">-c</span> <span class="nt">-o</span> startup.o startup.S
mipsel-linux-gnu-gcc-14 <span class="nt">-mips32r2</span> <span class="nt">-EL</span> <span class="nt">-O1</span> <span class="nt">-std</span><span class="o">=</span>gnu89 <span class="nt">-fno-builtin</span> <span class="nt">-ffreestanding</span> <span class="se">\</span>
    <span class="nt">-nostdlib</span> <span class="nt">-mno-abicalls</span> <span class="nt">-fno-pic</span> <span class="nt">-G0</span> <span class="nt">-mno-gpopt</span> <span class="nt">-Wall</span> <span class="nt">-c</span> <span class="nt">-o</span> main.o main.c
mipsel-linux-gnu-ld <span class="nt">-T</span> linker.ld <span class="nt">--no-warn-rwx-segments</span> <span class="nt">-o</span> firmware.elf startup.o main.o
mipsel-linux-gnu-objcopy <span class="nt">-O</span> binary firmware.elf firmware.bin

<span class="o">========================================</span>
  firmware.bin: 8496 bytes
  Format: Raw MIPS32 LE flat binary
  Base:   0x80000000
<span class="o">========================================</span></code></pre></figure>

<p>8,496 bytes of raw MIPS32 instructions and data. No headers, no symbols, no sections. Just bytes.</p>

<p>Try running <code class="language-plaintext highlighter-rouge">file</code> on it:</p>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash"><span class="nv">$ </span>file firmware.bin
firmware.bin: data

<span class="nv">$ </span>file firmware.elf
firmware.elf: ELF 32-bit LSB executable, MIPS, MIPS32 rel2 version 1 <span class="o">(</span>SYSV<span class="o">)</span>,
              statically linked, not stripped</code></pre></figure>

<p><code class="language-plaintext highlighter-rouge">file</code> does not even know what <code class="language-plaintext highlighter-rouge">firmware.bin</code> is. It just says “data.” That is what you are up against in firmware RE. Meanwhile, the intermediate ELF is recognized on sight – and that ELF is the one thing that never ships to customers.</p>

<p>The firmware has a UART interface that prints boot messages. With a serial cable, you would watch it initialize ports, set up routing tables, and present a <code class="language-plaintext highlighter-rouge">gw&gt;</code> prompt. For this series, we don’t need the hardware – the binary is enough.</p>

<p>If you <em>do</em> want to see it run, there are two options. You can emulate it under <strong>QEMU</strong> with the Malta MIPS board (<code class="language-plaintext highlighter-rouge">qemu-system-mipsel -M malta -kernel firmware.elf -serial stdio -nographic</code>) – the build includes a QEMU-compatible UART variant, and using the ELF here is just a convenience for emulation and debugging. The shipped firmware artifact is still the flat <code class="language-plaintext highlighter-rouge">firmware.bin</code> payload inside the image. Or if you have real hardware, a <strong>Microchip PIC32MX470 Curiosity</strong> board (DM320103) has a MIPS32 M4K core with UART-over-USB and can run similar bare-metal firmware. Neither is required to follow along – everything in this series works purely through static analysis of the binary.</p>

<p>In Part 2, we’ll learn how to extract information from those bytes without ever running the firmware.</p>

<hr />

<h2 id="firmware-is-more-than-one-file">Firmware Is More Than One File</h2>

<p>When you download a firmware update from a vendor, you don’t get a bare <code class="language-plaintext highlighter-rouge">.bin</code>. You get a <strong>firmware image</strong> – a container wrapping one or more payloads with headers, checksums, and a partition table. The flat binary is the payload we care about, but vendors rarely distribute it naked.</p>

<p>Our sample includes a packager that produces exactly this kind of image:</p>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash"><span class="nv">$ </span>python3 fwpack.py firmware.bin firmware.img

Packed firmware image: firmware.img
  Total size:    9152 bytes
  Image CRC-32:  0x0FF01071
  Partitions:    3

  <span class="o">[</span>bootloader]
    Offset:  0x0120
    Size:    32 bytes
    CRC-32:  0x4F7A6ACA
  <span class="o">[</span>main_fw]
    Offset:  0x0160
    Size:    8496 bytes
    CRC-32:  0x9F25E0DC
  <span class="o">[</span>config]
    Offset:  0x22B0
    Size:    256 bytes
    CRC-32:  0x2683AC5B</code></pre></figure>

<p>Notice the two magic strings (<code class="language-plaintext highlighter-rouge">FWPK</code> for the image header, <code class="language-plaintext highlighter-rouge">FWND</code> at the trailer), the per-partition CRCs, and the global image CRC – those are the three things every custom container format reinvents in some form.</p>

<p>The structure looks like this:</p>

<figure class="highlight"><pre><code class="language-text" data-lang="text">Firmware Image Layout
=====================

0x0000 +---------------------------+
       | Image Header (256 bytes)  |  Magic "FWPK", version, CRC-32,
       |                           |  firmware name, build date, board ID,
       |                           |  partition count, entry point
0x0100 +---------------------------+
       | Partition 0: Bootloader   |  32B header + 32B payload
       |  (tiny MIPS stub)         |  Sets SP, jumps to main FW
0x0140 +---------------------------+
       | Partition 1: Main FW      |  32B header + 8,496B payload
       |  (our firmware.bin)       |  The actual firmware code
0x2290 +---------------------------+
       | Partition 2: Config       |  32B header + 256B payload
       |  (default settings)       |  Hostname, IP, routes, serial config
0x23B0 +---------------------------+
       | Trailer (16 bytes)        |  End magic "FWND", size, CRC-32
0x23C0 +---------------------------+</code></pre></figure>

<p>Each partition has its own CRC-32, and the image has a global CRC in the header. Two reasons this matters: the device’s bootloader verifies them before flashing (a corrupted update bricks the device), and if you ever binary-patch the firmware, you have to recompute every CRC the bootloader checks – or the device rejects your modified image and refuses to boot.</p>

<p>Every vendor invents their own container format. Some build on well-known structures (U-Boot’s <code class="language-plaintext highlighter-rouge">uImage</code>). Others ship custom headers with proprietary magic numbers, encryption layers, or RSA signature verification. Your first RE task is always defeating the container before you can read the code inside it.</p>

<p>We can verify our image and extract the partitions back out:</p>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash"><span class="nv">$ </span>python3 fwpack.py <span class="nt">--info</span> firmware.img

Firmware Image: firmware.img
  Magic:       FWPK
  Version:     1.0.0.0
  Total size:  9152 bytes <span class="o">(</span>0x23c0<span class="o">)</span>
  Partitions:  3
  Entry point: 0x80000000
  Image CRC:   0x0FF01071
  FW name:     NetGW-MIPS32 v2.4.1-rc3
  Build <span class="nb">date</span>:  Apr 2026
  Board ID:    MIPS32-GW-DEV

  CRC verified OK

Partition Table:
  <span class="c">#  Type          Offset     Size       CRC-32      Load Addr</span>
  <span class="nt">--------------------------------------------------------------------</span>
  0  bootloader    0x0120         32     0x4F7A6ACA  0x80000000  <span class="o">[</span>OK]
  1  main_fw       0x0160       8496     0x9F25E0DC  0x80000000  <span class="o">[</span>OK]
  2  config        0x22B0        256     0x2683AC5B  0x80000000  <span class="o">[</span>OK]</code></pre></figure>

<p>Remember: ELF is a standard. Documented, parsed by every tool you own. Here? Custom format, custom magic, custom checksums. <code class="language-plaintext highlighter-rouge">file</code> won’t help. <code class="language-plaintext highlighter-rouge">readelf</code> won’t help. You and a hex editor, on a Friday night.</p>

<hr />

<h2 id="debug-symbols-the-cheat-code">Debug Symbols: The Cheat Code</h2>

<p>Before we dive deeper into the raw binary (that’s Part 2), one more comparison to make the stakes concrete. Same function – <code class="language-plaintext highlighter-rouge">firmware_main</code>, the entry point of our gateway – once without symbols and once with.</p>

<p><strong>Without symbols</strong> (the flat binary, as shipped to customers):</p>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash"><span class="nv">$ </span>mipsel-linux-gnu-objdump <span class="nt">-D</span> <span class="nt">-b</span> binary <span class="nt">-m</span> mips firmware.bin
...
    16b0:   e8ffbd27    addiu   sp,sp,-24
    16b4:   1400bfaf    sw      ra,20<span class="o">(</span>sp<span class="o">)</span>
    16b8:   1000b0af    sw      s0,16<span class="o">(</span>sp<span class="o">)</span>
    16bc:   0080053c    lui     a1,0x8000
    16c0:   e01ea524    addiu   a1,a1,7904
    16c4:   dd02000c    jal     0xb74
    16c8:   01000424    li      a0,1
    ...
    16f0:   f502000c    jal     0xbd4
    16f4:   00000000    nop
    16f8:   ad03000c    jal     0xeb4
    16fc:   00000000    nop
    1700:   03000624    li      a2,3
    1704:   01000524    li      a1,1
    1708:   fb03000c    jal     0xfec
    170c:   010a0424    li      a0,2561
    ...
    1728:   fb03000c    jal     0xfec
    172c:   a8c00434    li      a0,0xc0a8
    ...
    1784:   7205000c    jal     0x15c8</code></pre></figure>

<p>What can you read here? “Call <code class="language-plaintext highlighter-rouge">0xb74</code>. Call <code class="language-plaintext highlighter-rouge">0xbd4</code>. Call <code class="language-plaintext highlighter-rouge">0xeb4</code>. Call <code class="language-plaintext highlighter-rouge">0xfec</code> four times with different arguments. Call <code class="language-plaintext highlighter-rouge">0x15c8</code>.” That is the entirety of the information. Just addresses. You have no idea what any of those functions do, or whether two of them do the same thing.</p>

<p><strong>With symbols</strong> (the ELF, kept for debugging):</p>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash"><span class="nv">$ </span>mipsel-linux-gnu-objdump <span class="nt">-d</span> firmware.elf
...
800016b0 &lt;firmware_main&gt;:
800016b0:   27bdffe8    addiu   sp,sp,-24
800016b4:   afbf0014    sw      ra,20<span class="o">(</span>sp<span class="o">)</span>
800016b8:   afb00010    sw      s0,16<span class="o">(</span>sp<span class="o">)</span>
800016bc:   3c058000    lui     a1,0x8000
800016c0:   24a51ee0    addiu   a1,a1,7904
800016c4:   0c0002dd    jal     80000b74 &lt;log_msg&gt;
800016c8:   24040001    li      a0,1
...
800016f0:   0c0002f5    jal     80000bd4 &lt;port_init&gt;
800016f4:   00000000    nop
800016f8:   0c0003ad    jal     80000eb4 &lt;route_init&gt;
800016fc:   00000000    nop
80001700:   24060003    li      a2,3
80001704:   24050001    li      a1,1
80001708:   0c0003fb    jal     80000fec &lt;route_add&gt;
8000170c:   24040a01    li      a0,2561
...
80001728:   0c0003fb    jal     80000fec &lt;route_add&gt;
8000172c:   3404c0a8    li      a0,0xc0a8
...
80001784:   0c000572    jal     800015c8 &lt;cli_main_loop&gt;</code></pre></figure>

<p>Now you can read the whole boot sequence: log “system starting”, initialize ports, initialize the route table, add four routes (10.1.x.x, 10.2.x.x, 192.168.x.x, 172.16.x.x), log “system ready”, print banner, enter the CLI main loop.</p>

<p>Same machine code, same addresses, same behavior. The only difference is that the ELF carries a symbol table mapping addresses to names:</p>

<figure class="highlight"><pre><code class="language-bash" data-lang="bash"><span class="nv">$ </span>mipsel-linux-gnu-nm firmware.elf | <span class="nb">head</span> <span class="nt">-20</span>
80000000 T _start
80000248 T _exception_handler
800002e4 T uart_puts
80000b74 T log_msg
80000bd4 T port_init
80000eb4 T route_init
80000efc T iterate_active_routes
80000fec T route_add
800013d0 T cli_process_command
800015c8 T cli_main_loop
800016b0 T firmware_main
...</code></pre></figure>

<p>93 symbols in our toy firmware. The vendor firmware we worked on had over 4,400 functions across 1.5 MB of code. Picture 4,400 nameless functions – just addresses and raw MIPS instructions, with no obvious starting point and no obvious end. That is why recovering function names and code structure takes weeks of manual work.</p>

<p>Symbols turn weeks into hours. Vendor firmware almost never ships with them. The rest of this series is about getting them back.</p>

<hr />

<h2 id="what-were-up-against">What We’re Up Against</h2>

<p>Let me summarize where we stand. We have:</p>

<ul>
  <li>A <strong>flat binary</strong> – raw MIPS32 instructions with no headers, no sections, no symbols</li>
  <li>A <strong>firmware image container</strong> – custom format with magic numbers, CRC checksums, and multiple partitions</li>
  <li>A <strong>trivial address mapping</strong> – file offset + base address = virtual address</li>
  <li><strong>109 embedded strings</strong> – the main human-readable foothold in this sample binary</li>
</ul>

<p>And we know what we’re missing:</p>

<ul>
  <li>No function boundaries (where does one function end and another begin?)</li>
  <li>No function names (what does the code at <code class="language-plaintext highlighter-rouge">0xbd4</code> do?)</li>
  <li>No variable names (what’s stored at <code class="language-plaintext highlighter-rouge">0x80040340</code>?)</li>
  <li>No type information (is that 32-bit value an integer, a pointer, or a bitfield?)</li>
  <li>No cross-references (who calls this function? Where is this string used?)</li>
</ul>

<p>The tools from our ELF toolkit – <code class="language-plaintext highlighter-rouge">readelf</code>, <code class="language-plaintext highlighter-rouge">nm</code>, <code class="language-plaintext highlighter-rouge">objdump -tT</code> – are useless against a flat binary. We need a different approach.</p>

<p>Here’s what we’re going to do across the rest of the series:</p>

<p><strong>Part 2: “Opcodes, Prologues, and other hidden patterns”</strong> – We’ll learn to read MIPS32 assembly just enough to find function boundaries, trace calls, and discover a hidden CLI with commands that don’t appear in any help output. Our entry point? Those 109 strings.</p>

<p><strong>Part 3: “Decompilers, Annotations, and other ways to read the unreadable”</strong> – We’ll throw the binary at Ghidra to decompile all 8,496 bytes into readable C, then follow the execution path through the route engine to discover a priority inversion bug that silently drops traffic. Then we’ll binary-patch the fix – and recalculate the CRC.</p>

<p><strong>Part 4: “Symbols, Scripts, and other linking nightmares”</strong> – We’ll try to go from decompiled C back to a working binary. We’ll discover why that’s far harder than it sounds, build custom linker scripts, and learn why firmware linking is fundamentally different from anything in userspace.</p>

<p><strong>Part 5: “Versions, Callgraphs, and other ways to compare what changed”</strong> – We’ll compare two firmware versions side by side, decompile the same functions from each, and use diffs and callgraph patterns to see how a vendor fix evolves between releases.</p>

<p>Every technique here comes from a real project. Every pattern is something we hit in production firmware. The sample recreates those patterns so you can follow along hands-on, without anyone’s lawyer getting involved.</p>

<p>So far we’ve established what isn’t in the binary. In Part 2, we start with the only thing that <em>is</em> readable – the strings – and from there pull function boundaries, a call graph, and a hidden CLI out of raw bytes.</p>

<hr />

<h2 id="references">References</h2>

<ul>
  <li><a href="/kernel/2018/07/28/Elfs_Linkers_Other.html">ELF’s Linker’s and other magical creatures</a> – The 2018 predecessor to this series</li>
  <li><a href="https://s3-eu-west-1.amazonaws.com/downloads-mips/documents/MD00086-2B-MIPS32BIS-AFP-6.06.pdf">MIPS32 Architecture For Programmers</a> – MIPS instruction set reference</li>
  <li><a href="https://sourceware.org/binutils/docs/">GNU binutils documentation</a> – <code class="language-plaintext highlighter-rouge">objcopy</code>, <code class="language-plaintext highlighter-rouge">objdump</code>, <code class="language-plaintext highlighter-rouge">readelf</code>, <code class="language-plaintext highlighter-rouge">nm</code></li>
  <li><a href="https://ghidra-sre.org/">Ghidra</a> – NSA’s reverse engineering framework</li>
  <li><a href="https://www.capstone-engine.org/">Capstone Engine</a> – Lightweight disassembly framework</li>
  <li><a href="https://en.wikipedia.org/wiki/Calling_convention#MIPS">MIPS Calling Conventions</a> – Register usage and stack frame layout</li>
</ul>]]></content><author><name>Maciej Grochowski</name></author><category term="reverse-engineering" /><category term="firmware" /><category term="mips" /><summary type="html"><![CDATA[Eight years ago I wrote ELF’s Linker’s and other magical creatures – a walkthrough of the ELF binary format, relocations, segments, and even live code injection through /proc/pid/mem. That post ended in a comfortable place: .text, .data, .bss; ld resolving symbols; the kernel’s ELF loader mapping segments into memory; gdb poking at a running process. The civilized world of userspace binaries.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://page-fault.io/assets/social-preview.png" /><media:content medium="image" url="https://page-fault.io/assets/social-preview.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Before the BSD Kernel Starts: Part Two ARMv8</title><link href="https://page-fault.io/kernel/2026/03/21/Before-BSD-Kernel-Part2.html" rel="alternate" type="text/html" title="Before the BSD Kernel Starts: Part Two ARMv8" /><published>2026-03-21T10:00:00+00:00</published><updated>2026-03-21T10:00:00+00:00</updated><id>https://page-fault.io/kernel/2026/03/21/Before-BSD-Kernel-Part2</id><content type="html" xml:base="https://page-fault.io/kernel/2026/03/21/Before-BSD-Kernel-Part2.html"><![CDATA[<blockquote>
  <p><strong>Note:</strong> This article was written in 2020 as a companion to
<a href="/kernel/2020/11/21/Before-BSD-Kernel-Part1.html">Part One (AMD64)</a>
but was never published. The technical content reflects the state of NetBSD and the
RK3399 platform at that time. Published here now without major changes.</p>
</blockquote>

<hr />

<h2 id="introduction">Introduction</h2>

<p>In Part One we covered legacy initialization of an AMD64-based system. This time we look
at ARM processors.</p>

<p>ARM architecture is well known from smaller devices — embedded systems, IoT, smartphones —
but ARMv8 already powers modern servers and laptops. The architecture has many revisions
with backward compatibility considerations. To keep things focused, we will concentrate
on ARMv8 (64-bit), which can also execute the 32-bit instruction set in compatibility mode.
We will run 64-bit NetBSD throughout.</p>

<p>There are two main ways manufacturers obtain ARM processors: via RTL design (ARM sells the
core design directly) or as an IP license (the company buys the specification). This business
model makes the architecture highly flexible and implementation-specific, which also makes
learning it harder — there are many chip-specific details that the architecture itself leaves
undefined.</p>

<p>For this article we look at generic, high-level ARMv8 details first, then go deeper using a
specific hardware platform. I originally considered Raspberry Pi 3 or 4, but the Broadcom GPU
plays an essential and poorly-documented role in its boot process — after reset it is the GPU
that starts first and later releases the ARM cores. The lack of public documentation makes it
a poor teaching example.</p>

<p>Instead I use the RK3399Pro from Rockchip, the chip behind PINE64 devices such as the
ROCKPro64 single-board computer and the Pinebook Pro laptop. It has fully open-sourced
documentation and real-world relevance.</p>

<p>This article intentionally avoids the Secure Boot process and UEFI internals — those will
get their own dedicated part.</p>

<hr />

<h2 id="the-bigger-picture">The Bigger Picture</h2>

<h3 id="exception-levels">Exception Levels</h3>

<p>ARMv8 introduces four privilege levels called Exception Levels (EL), numbered 0 to 3.
Every implementation must support EL0 and EL1; EL2 and EL3 are optional.</p>

<p>The name “exception level” can be confusing — “privilege level” is the clearer term, and
was used in earlier ARM specifications. The best way to understand them is by what runs at
each level:</p>

<ul>
  <li><strong>EL0</strong> — unprivileged user-space applications. No direct access to system registers,
page tables, or hardware devices.</li>
  <li><strong>EL1</strong> — the OS kernel. Responsible for hardware configuration, page table setup,
and managing EL0 processes.</li>
  <li><strong>EL2</strong> — hypervisor. UEFI also runs here during boot — not because UEFI itself needs
hypervisor privileges, but because running at EL2 <em>preserves</em> EL2 availability for the
OS. If UEFI ran at EL1, the OS would have nowhere to put a hypervisor later.</li>
  <li><strong>EL3</strong> — Secure Monitor. The most privileged level, used by ARM Trusted Firmware.</li>
</ul>

<p>EL0 and EL1 are the minimum required to run a modern OS such as NetBSD or Linux.</p>

<p>ARMv8 only allows a <em>decrease</em> in execution state going toward less privileged levels: a
64-bit EL2 can host a 64-bit or 32-bit EL1, a 64-bit EL1 can run 64-bit or 32-bit EL0
applications, but a 32-bit OS cannot run 64-bit applications. The execution state of each
level is controlled by bits in the higher level’s configuration registers: <code class="language-plaintext highlighter-rouge">SCR_EL3.RW</code>
determines whether EL2 runs AArch64 or AArch32, and <code class="language-plaintext highlighter-rouge">HCR_EL2.RW</code> determines the same for
EL1. The reset execution state of EL3 itself is controlled by <code class="language-plaintext highlighter-rouge">RMR_EL3.AA64</code> — but that
register affects the state after a <em>warm reset</em>, not the running state of lower ELs.</p>

<p><img src="https://res.cloudinary.com/gotocco/image/upload/v1610250036/Articles/System_boot_part2_ARMv8/ARMv8_ExceptionLevels.svg" alt="Exception Levels" /></p>

<h3 id="out-of-reset">Out of Reset</h3>

<p>Unlike AMD64 where the reset vector is a fixed physical address (<code class="language-plaintext highlighter-rouge">0xFFFFFFF0</code>), ARMv8
processors do not have a fixed reset vector. After reset, the processor reads <code class="language-plaintext highlighter-rouge">RVBAR_ELx</code>
(where <code class="language-plaintext highlighter-rouge">x</code> is the highest implemented exception level, typically EL3) to obtain the
implementation-defined reset vector address, then fetches instructions from that address.</p>

<p>For the rest of this article, “EL3” refers to EL3 if implemented, or the highest available
exception level otherwise. Most ARMv8 chips do implement EL3, but it is not architecturally
mandatory.</p>

<p>After reset, the processor state is architecturally defined as follows:</p>
<ul>
  <li>All interrupt masks are set (DAIF = 0b1111 — all asynchronous exceptions masked)</li>
  <li>MMU is off; instruction and data caches are disabled (<code class="language-plaintext highlighter-rouge">SCTLR_ELx.M</code>, <code class="language-plaintext highlighter-rouge">.I</code>, <code class="language-plaintext highlighter-rouge">.C</code> = 0)</li>
  <li>The CPU executes physical addresses directly</li>
  <li>TLB and cache contents are implementation-defined (boot ROM typically invalidates both)</li>
  <li>Memory is unconfigured; the interrupt controller state is unknown</li>
</ul>

<hr />

<h2 id="general-boot-sequence">General Boot Sequence</h2>

<p>Many architectures — including ARMv7 — start code execution from the first entry of an
exception table (the reset vector). In ARMv8 the reset vector is no longer part of the
exception table. Instead, after reset the processor fetches from the implementation-defined
address in <code class="language-plaintext highlighter-rouge">RVBAR_ELx</code> (where <code class="language-plaintext highlighter-rouge">x</code> is the highest implemented level: 3–1).</p>

<p>The first instructions executed are typically initialization code in a small on-chip ROM
(or OTP — One-Time Programmable memory). Most peripherals — memory controllers, bus
controllers — are disabled right after reset, so the processor needs some immediately
accessible memory. That memory must be inside the chip itself, which makes it expensive
and therefore small (order of kilobytes).</p>

<p>Boot ROM prepares the processor and handles chip-specific details: memory and cache
initialization, branch predictor configuration. Since TLB and cache state after reset is
implementation-defined, boot ROM typically performs explicit invalidations before enabling
them.</p>

<h3 id="boot-stages">Boot Stages</h3>

<p>The ARMv8 boot process typically has three stages:</p>

<ol>
  <li><strong>First stage — Boot ROM:</strong> read-only code on-chip, executes immediately after reset</li>
  <li><strong>Second stage — Boot Loader:</strong> loaded by Boot ROM from external storage</li>
  <li><strong>Third stage — Advanced Boot Loader:</strong> loads the operating system (where UEFI lives)</li>
</ol>

<p><img src="https://res.cloudinary.com/gotocco/image/upload/v1610250036/Articles/System_boot_part2_ARMv8/ARM_General_Boot.svg" alt="Boot Stages" /></p>

<h3 id="bootrom-vs-bootloader">BootROM vs. Bootloader</h3>

<p><strong>BootROM</strong> is the small program in read-only memory (ROM or OTP), typically provided by
the chip manufacturer since it handles low-level, platform-specific initialization. Its
main job is to prepare the CPU to execute external code from storage (flash, SD card, NVMe).
It may also provide cryptographic features for a chain of trust and handle partition layout
parsing. ARM Trusted Firmware defines requirements for this stage, but for non-secure boot
there is no mandated BootROM interface standard.</p>

<p><strong>Bootloader</strong> is a program (or chain of programs) whose primary task is to load the OS.
It typically lives on external storage, is modifiable, and can be upgraded. Some systems
have a BootROM advanced enough to load a kernel or UEFI image directly; others have a
minimal BootROM that can only fetch and execute from a fixed memory location, requiring the
bootloader to be split into multiple stages due to size constraints.</p>

<p>The second-stage bootloader, loaded by BootROM from external storage, is platform-specific
and sometimes split further into sub-programs due to hardware constraints. Its goal is to
initialize system RAM so the third-stage bootloader has full platform access.</p>

<p>The third stage prepares the entire platform to boot the OS. UEFI operates here at EL2.
DRAM is operational at this point. The bootloader reads the OS from storage, verifies the
image if secure boot is active, passes boot parameters, and jumps to the kernel entry point.</p>

<h3 id="before-the-os-device-tree-and-platform-differences">Before the OS: Device Tree and Platform Differences</h3>

<p>ARM hardware covers a wide range — embedded devices through enterprise servers — and the
boot environment differs accordingly.</p>

<p>For embedded devices there is usually no universal hardware description standard. What is
common is a <strong>device tree</strong>: a structured description of the peripherals connected to the
chip (I2C, SPI, UART, CAN buses and their devices) that the OS cannot discover at runtime.
The kernel reads this binary at boot and configures peripherals accordingly. The
<a href="https://github.com/ARM-software/ebbr">Embedded Base Boot Requirements (EBBR)</a> specification
defines the format.</p>

<p>For server-class systems, the <a href="https://developer.arm.com/documentation/den0044/e/">Server Base Boot Requirements (SBBR)</a>
apply instead. SBBR-compliant systems must not expose a device tree to the OS — hardware
discovery follows ACPI and UEFI conventions.</p>

<hr />

<h2 id="rockpro64-and-rk3399">ROCKPro64 and RK3399</h2>

<p><img src="https://res.cloudinary.com/gotocco/image/upload/v1610250036/Articles/System_boot_part2_ARMv8/RK3399_diagram.svg" alt="RK3399 Diagram" /></p>

<p>The Rockchip RK3399 is a System-on-Chip with two ARM processor clusters:</p>

<ul>
  <li>Dual-core Cortex-A72 (big cores)</li>
  <li>Quad-core Cortex-A53 (little cores)</li>
</ul>

<p>The clusters are connected via ARM Coherent Interconnect (CCI). Both have L1 and L2 caches
and support all exception levels (EL0–EL3). The SoC contains a main interconnect and a
peripheral interconnect (with low- and high-performance domains), connected through a
128-bit/64-bit/32-bit multi-layer AXI/AHB/APB bus architecture.</p>

<p>Embedded SRAM comes in two units: 8 KB (accessible via the PMU for power management) and
192 KB (accessible via the peripheral bus, protected by TZMA for security). The larger
192 KB block is remapped and accessible at <code class="language-plaintext highlighter-rouge">0xFFFF_0000–0xFFFF_FFFF</code> and is where the
bootloader stages execute.</p>

<h3 id="rockpro64-boot-sequence">RockPro64 Boot Sequence</h3>

<p>RK3399 supports normal and secure boot. This article covers normal boot only.</p>

<p>After power-on reset, the RK3399 Boot ROM starts at <code class="language-plaintext highlighter-rouge">0xFFFF0000</code>. It searches for a valid
ID block header at the following locations, in order:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Offset 0          in SPI flash
Offset 0x8000     on eMMC   (sector 64, where a sector is 512 bytes)
Offset 0x8000     on SD card (sector 64)
</code></pre></div></div>

<p>These are fixed offsets — not partition-table entries. The Boot ROM looks for an ID header
at each location and executes from the first one that matches.</p>

<p>Full boot chain using U-Boot:</p>

<ol>
  <li>Boot ROM loads <strong>U-Boot TPL</strong> into SRAM. TPL initializes main system RAM.</li>
  <li>Control returns from TPL to Boot ROM (<code class="language-plaintext highlighter-rouge">return-to-bootrom</code>).</li>
  <li>Boot ROM loads <strong>U-Boot SPL</strong>.</li>
  <li>SPL loads <strong>ARM Trusted Firmware (ATF)</strong> and U-Boot into main memory.</li>
  <li>ATF runs U-Boot.</li>
  <li>U-Boot loads <code class="language-plaintext highlighter-rouge">bootaa64.efi</code> from the FAT boot partition (<code class="language-plaintext highlighter-rouge">\EFI\BOOT\</code>).</li>
  <li>The UEFI loader finds and executes the NetBSD kernel.</li>
</ol>

<p><img src="https://res.cloudinary.com/gotocco/image/upload/v1615077833/RK3399_BootSequence_NetBSD_zbviux.svg" alt="Boot Sequence" /></p>

<p><strong>Glossary for this chain:</strong></p>

<ul>
  <li><strong>TPL (Tertiary Program Loader):</strong> A minimal subset of SPL. Some boards have tight size
constraints on early-stage code; TPL handles DDR initialization only, then hands off to SPL.</li>
  <li><strong>SPL (Secondary Program Loader):</strong> A small binary generated from U-Boot source that fits
in SRAM and loads the main U-Boot into system RAM.</li>
  <li><strong>MLO (Memory Loader):</strong> A broader term for any second-stage program that loads the next
bootloader into memory. SPL is the U-Boot-specific variant; some boards use their own MLO.</li>
  <li><strong>ATF (ARM Trusted Firmware):</strong> The reference implementation for secure boot on ARM.
Divides the boot process into stages BL1 (Boot ROM), BL2 (verified by BL1), and BL3
(loads U-Boot/GRUB or similar). Establishes the chain of trust.</li>
</ul>

<p><strong>A note on <code class="language-plaintext highlighter-rouge">return-to-bootrom</code>:</strong> this is a Rockchip-specific mechanism. The Boot ROM
leaves a function pointer in memory before transferring control to TPL. TPL can call this
pointer — roughly equivalent to a <code class="language-plaintext highlighter-rouge">longjmp</code> — to return control to the Boot ROM, which then
continues to the next stage. If the second-stage loader fails and returns to Boot ROM with
the appropriate error code, Boot ROM enters an upgrade mode accessible over USB. In the
normal boot path, TPL uses this mechanism to hand off to SPL cleanly.</p>

<h3 id="disk-layout">Disk Layout</h3>

<p>U-Boot is flexible and can boot from many sources. For NetBSD on RockPro64 we format the
boot medium with an MBR partition table containing two partitions: a FAT32 system partition
holding the EFI loader, and a NetBSD partition (ID <code class="language-plaintext highlighter-rouge">0xa9</code>) holding the kernel and root filesystem.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>file ./NetBSD-9-aarch64-202101091800Z-pinebook-pro.img:
DOS/MBR boot sector;
partition 1 : ID=0xc, active, start-CHS (0x2,10,9),  end-CHS (0xc,60,48),   startsector 32768, 163840 sectors;
partition 2 : ID=0xa9,        start-CHS (0xc,60,49), end-CHS (0x92,123,41), startsector 196608, 2156672 sectors
</code></pre></div></div>

<p>The layout of sectors on the medium before the partitions:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Start (sectors) | Size (sectors) | Name                  | Description
----------------|----------------|-----------------------|---------------------------
64              | 16320          | IDBLoader             | SoC init code (TPL + SPL)
16384           | 8192           | OS loader             | Main U-Boot
24576           | 8192           | TrustedFirmware-A     | ARM Trusted Firmware
32768           | 163840         | FAT32 system partition | EFI loader
196608          | 2156672        | NetBSD partition       | FFS filesystem with kernel
</code></pre></div></div>

<p>The EFI loader (provided in NetBSD sources) finds the kernel file, loads it into RAM, and
calls <code class="language-plaintext highlighter-rouge">aarch64_exec_kernel</code>. Before transferring control, it invalidates instruction and data
caches to ensure the kernel image is coherently placed in RAM, and disables the MMU — the
kernel starts with physical addresses and sets up its own address space from scratch.</p>

<hr />

<h2 id="starts-locores-and-machine-dependent-code">start.S, locore.S and Machine-Dependent Code</h2>

<h3 id="from-uefi-to-the-kernel">From UEFI to the Kernel</h3>

<p>When a multiprocessor SoC powers on, all cores become active after the power-on reset.
Early-stage firmware designates one as the boot core and holds the others with a
<code class="language-plaintext highlighter-rouge">wfi</code> (wait for interrupt) instruction. NetBSD’s early initialization assumes exactly this:
one active core, the rest idle.</p>

<p>Execution begins in <code class="language-plaintext highlighter-rouge">start.S</code> inside <code class="language-plaintext highlighter-rouge">/sys/arch/aarch64</code>. This file is a thin proxy to
<code class="language-plaintext highlighter-rouge">aarch64_start</code> in <code class="language-plaintext highlighter-rouge">locore.S</code>; its main job is extracting boot arguments.</p>

<p>The processor arrives here at EL1 — the EFI loader has already dropped from EL2 to EL1
before jumping to the kernel. This drop is done via the <code class="language-plaintext highlighter-rouge">ERET</code> instruction: the EFI loader
sets <code class="language-plaintext highlighter-rouge">SPSR_EL2</code> to describe the target PSTATE (EL1h, all interrupts masked), sets <code class="language-plaintext highlighter-rouge">ELR_EL2</code>
to the kernel entry address, and executes <code class="language-plaintext highlighter-rouge">ERET</code>. The processor atomically switches to EL1
and begins fetching from the kernel entry point. The MMU is off at this point — the EFI
loader disabled it before the jump, so the kernel executes physical addresses until it
builds its own page tables.</p>

<p>The steps ahead:</p>

<ol>
  <li>Initialize system registers</li>
  <li>Build MMU translation tables</li>
  <li>Enable the MMU</li>
  <li>Jump into virtual address space</li>
</ol>

<h3 id="the-pstate-register">The PState Register</h3>

<p>PState (Processing Element State) holds the processor’s current execution state, including
the active exception level. In AArch64, PState is not a single accessible register — it is
a collection of fields readable through dedicated system registers (<code class="language-plaintext highlighter-rouge">CurrentEL</code>, <code class="language-plaintext highlighter-rouge">SPSR_ELx</code>,
<code class="language-plaintext highlighter-rouge">DAIF</code>, etc.). Each exception level has its own copies of the relevant registers.</p>

<p><img src="https://res.cloudinary.com/gotocco/image/upload/v1616909952/PState_Explained_zlhqmd.svg" alt="PState Register" /></p>

<p>The overall sequence in <code class="language-plaintext highlighter-rouge">locore.S</code> before <code class="language-plaintext highlighter-rouge">main</code> is called:</p>

<pre><code class="language-asm">        /* Disable MMU (ensure clean state) */
        bl      mmu_disable

        /* Initialize system registers and build MMU tables */
        bl      init_sysregs
        bl      init_mmutable
        bl      save_ttbrs

        /* Enable MMU */
        bl      mmu_enable

        /* Load virtual address of vstart into x20 */
        ldr     x20, =vstart

        /* Jump to virtual address space */
        br      x20

/*
 * vstart executes in kernel virtual address space
 */
vstart:
</code></pre>

<p><code class="language-plaintext highlighter-rouge">mmu_disable</code> is called first even though the EFI loader already disabled the MMU — it is
a defensive measure in case the kernel is entered through a different path (e.g. a kexec-like
mechanism or a bootloader that does not follow the EFI handoff convention).</p>

<p><code class="language-plaintext highlighter-rouge">save_ttbrs</code> stores the translation table base registers <code class="language-plaintext highlighter-rouge">TTBR0_EL1</code> and <code class="language-plaintext highlighter-rouge">TTBR1_EL1</code> after
they are configured. ARMv8 uses two TTBRs: TTBR0 covers the lower virtual address range
(user space, configured per-process), and TTBR1 covers the upper range (kernel space, shared).
Saving them at this point lets secondary CPUs reuse the same page tables when they are
brought online later.</p>

<p><strong>ARMv8 branch instructions quick reference:</strong></p>
<ul>
  <li><code class="language-plaintext highlighter-rouge">B label</code> — unconditional branch to a PC-relative 26-bit signed offset</li>
  <li><code class="language-plaintext highlighter-rouge">BL label</code> — branch and link; saves return address in <code class="language-plaintext highlighter-rouge">x30</code></li>
  <li><code class="language-plaintext highlighter-rouge">BR xN</code> — branch to address in register <code class="language-plaintext highlighter-rouge">xN</code> (not a subroutine return; use <code class="language-plaintext highlighter-rouge">BLR</code> for that)</li>
</ul>

<h3 id="initializing-system-registers">Initializing System Registers</h3>

<p><code class="language-plaintext highlighter-rouge">init_sysregs</code> configures several EL1 system registers before the MMU is enabled:</p>

<pre><code class="language-asm">init_sysregs:
        stp     x0, lr, [sp, #-16]!

        /* Configure debug event register */
        ldr     x0, mdscr_setting
        msr     mdscr_el1, x0

        /* Unlock OS lock (allows external debugger access) */
        msr     oslar_el1, xzr

        /* Clear context ID register */
        msr     contextidr_el1, xzr

        /* Trap FP/SIMD at both EL0 and EL1 until FP context is ready */
        msr     cpacr_el1, xzr

        /* Allow EL0 to read the virtual counter and frequency */
        mrs     x0, cntkctl_el1
        orr     x0, x0, #CNTKCTL_EL0VCTEN
        msr     cntkctl_el1, x0

        /* Unmask all exceptions */
        msr     daif, xzr

        ldp     x0, lr, [sp], #16
        ret
</code></pre>

<p>Key registers touched here:</p>

<table>
  <thead>
    <tr>
      <th>Register</th>
      <th>Purpose</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">MDSCR_EL1</code></td>
      <td>Monitor Debug System Control — configures debug event behavior</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">OSLAR_EL1</code></td>
      <td>OS Lock Access Register — writing zero clears the OS Lock, enabling external debugger access to OS save/restore registers</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">CONTEXTIDR_EL1</code></td>
      <td>Process context identifier used by debug and trace infrastructure; zeroed here as no process context exists yet</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">CPACR_EL1</code></td>
      <td>Architectural Feature Access Control — <code class="language-plaintext highlighter-rouge">FPEN=0b00</code> traps all FP/SIMD instructions at both EL0 and EL1 until the kernel sets up FP context management</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">CNTKCTL_EL1</code></td>
      <td>Counter-timer Kernel Control — setting <code class="language-plaintext highlighter-rouge">EL0VCTEN</code> allows user-space to read the virtual counter directly without trapping to EL1</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">DAIF</code></td>
      <td>Interrupt mask bits (Debug, SError, IRQ, FIQ) — zeroing unmasks all</td>
    </tr>
  </tbody>
</table>

<p>A note on <code class="language-plaintext highlighter-rouge">CPACR_EL1</code>: setting it to zero means <code class="language-plaintext highlighter-rouge">FPEN = 0b00</code>, which causes any FP or SIMD
instruction at either EL0 or EL1 to trap. This is intentional — until the kernel has set
up FP register save/restore in its exception handlers, allowing FP instructions would
silently corrupt the FP register file across context switches. The kernel re-enables FP
access later once the infrastructure is ready.</p>

<p>A note on <code class="language-plaintext highlighter-rouge">CONTEXTIDR_EL1</code>: this register holds a process identifier for debug and trace
hardware (e.g. ETM). It does not control ASID-based TLB tagging — ASIDs live in bits
[63:48] of <code class="language-plaintext highlighter-rouge">TTBR0_EL1</code> and <code class="language-plaintext highlighter-rouge">TTBR1_EL1</code>.</p>

<h3 id="enabling-the-mmu">Enabling the MMU</h3>

<p>Out of reset (and after the EFI loader’s handoff), the MMU is off. The CPU executes physical
addresses directly with no access permission enforcement and no memory attribute
differentiation. Both instruction and data caches are architecturally disabled after reset
(<code class="language-plaintext highlighter-rouge">SCTLR_EL1.I = 0</code>, <code class="language-plaintext highlighter-rouge">SCTLR_EL1.C = 0</code>). The instruction cache can be enabled independently
of the MMU, but enabling the data cache without the MMU requires careful setup — without address
translation, the cache cannot enforce memory ordering or access permissions correctly,
which makes it unsafe in any general-purpose OS context.</p>

<p><code class="language-plaintext highlighter-rouge">mmu_enable</code> performs the following steps:</p>

<pre><code class="language-asm">mmu_enable:
        dsb     sy

        /* Invalidate all EL1 TLB entries */
        dsb     ishst
        tlbi    vmalle1
        dsb     ish
        isb

        /* Set memory attributes */
        ldr     x0, mair_setting
        msr     mair_el1, x0

        /* Configure translation control; set IPS from hardware capability */
        ldr     x0, tcr_setting
        mrs     x1, id_aa64mmfr0_el1
        bfi     x0, x1, #32, #3
        msr     tcr_el1, x0

        /* Configure and enable via SCTLR_EL1 */
        mrs     x0, sctlr_el1
        ldr     x1, sctlr_clear
        bic     x0, x0, x1
        ldr     x1, sctlr_pac       /* disable PAC */
        bic     x0, x0, x1
        ldr     x1, sctlr_set
        orr     x0, x0, x1

        ldr     x1, sctlr_ee
#ifdef __AARCH64EB__
        orr     x0, x0, x1          /* big-endian */
#else
        bic     x0, x0, x1          /* little-endian */
#endif
        msr     sctlr_el1, x0       /* M bit set — MMU now enabled */
        isb

        ret
</code></pre>

<p>Breaking down the key parts:</p>

<p><strong>MAIR_EL1 (Memory Attribute Indirection Register):</strong> defines up to 8 memory attribute
slots that page table entries reference by index (<code class="language-plaintext highlighter-rouge">AttrIndx</code> field). NetBSD sets up five:</p>

<pre><code class="language-asm">mair_setting:
        .quad (                                          \
            __SHIFTIN(MAIR_NORMAL_WB,     MAIR_ATTR0) | \
            __SHIFTIN(MAIR_NORMAL_NC,     MAIR_ATTR1) | \
            __SHIFTIN(MAIR_NORMAL_WT,     MAIR_ATTR2) | \
            __SHIFTIN(MAIR_DEVICE_MEM,    MAIR_ATTR3) | \
            __SHIFTIN(MAIR_DEVICE_MEM_SO, MAIR_ATTR4))
</code></pre>

<p>Normal Write-Back, Normal Non-Cacheable, Normal Write-Through, Device memory
(<code class="language-plaintext highlighter-rouge">Device-nGRE</code>), and <code class="language-plaintext highlighter-rouge">MAIR_DEVICE_MEM_SO</code> — the most restrictive device type,
corresponding to AArch64 <code class="language-plaintext highlighter-rouge">Device-nGnRnE</code> (non-Gathering, non-Reordering,
non-Early-write-acknowledgement). This is the AArch64 equivalent of what ARMv7 called
“Strongly Ordered.” Page table descriptors point to these slots via their <code class="language-plaintext highlighter-rouge">AttrIndx[2:0]</code>
field.</p>

<p><strong>TCR_EL1 (Translation Control Register):</strong> controls granule size, address space size, and
the intermediate physical address size (IPS) at bits [34:32]. The instruction
<code class="language-plaintext highlighter-rouge">bfi x0, x1, #32, #3</code> inserts bits [2:0] of <code class="language-plaintext highlighter-rouge">ID_AA64MMFR0_EL1.PARange</code> into bits [34:32]
of <code class="language-plaintext highlighter-rouge">TCR_EL1</code> — matching the IPS configuration to the physical address range the hardware
actually supports. Getting this wrong would cause translation faults on systems with more
than 32-bit physical address space.</p>

<p><strong>SCTLR_EL1 (System Control Register):</strong> the <code class="language-plaintext highlighter-rouge">M</code> bit (bit 0) enables the MMU. The code
also clears PAC (Pointer Authentication Code) bits and configures endianness before writing
the final value. The <code class="language-plaintext highlighter-rouge">isb</code> instruction barrier after the write ensures the pipeline is
flushed and the MMU is fully active before any subsequent instruction fetch proceeds.</p>

<p>After <code class="language-plaintext highlighter-rouge">mmu_enable</code> returns, the CPU is translating virtual addresses. The <code class="language-plaintext highlighter-rouge">br x20</code>
instruction then jumps to <code class="language-plaintext highlighter-rouge">vstart</code>, which is a virtual address — this is the moment the
kernel fully enters its own virtual address space.</p>

<hr />

<h2 id="inside-virtual-memory">Inside Virtual Memory</h2>

<p>Once the MMU is enabled and execution is in virtual address space, <code class="language-plaintext highlighter-rouge">vstart</code> completes
the remaining initialization before calling <code class="language-plaintext highlighter-rouge">main</code>:</p>

<ul>
  <li><strong>Exception Vector</strong> — installs the kernel exception vector table by writing to <code class="language-plaintext highlighter-rouge">VBAR_EL1</code></li>
  <li><strong>Process 0 stack</strong> — sets up the initial kernel stack for the idle/init process</li>
  <li><strong>PAC setup</strong> — enables Pointer Authentication if <code class="language-plaintext highlighter-rouge">ID_AA64ISAR1_EL1</code> indicates support</li>
  <li><strong>CPU topology</strong> — <code class="language-plaintext highlighter-rouge">arm_cpu_topology_set</code> reads MPIDR_EL1 to determine cluster and core IDs</li>
  <li><strong>Cache info</strong> — <code class="language-plaintext highlighter-rouge">aarch64_getcacheinfo</code> reads <code class="language-plaintext highlighter-rouge">CLIDR_EL1</code> and <code class="language-plaintext highlighter-rouge">CCSIDR_EL1</code> to discover
cache geometry (levels, associativity, line size)</li>
  <li><strong><code class="language-plaintext highlighter-rouge">initarm</code></strong> — machine-dependent initialization: GIC interrupt controller, clocks,
device tree parsing, early console</li>
  <li><strong><code class="language-plaintext highlighter-rouge">main</code></strong> — the C entry point for the rest of kernel initialization</li>
</ul>

<hr />

<h2 id="summary">Summary</h2>

<p>This part covered the ARMv8 boot process from reset to the kernel’s C entry point, using
NetBSD on the RK3399-based ROCKPro64 as a concrete example.</p>

<p>The key structural difference from AMD64 is that ARMv8 provides a formal privilege hierarchy
(exception levels) from the start, and the boot chain reflects this: BootROM runs at EL3,
ATF establishes secure world services, UEFI runs at EL2, and the OS kernel settles into EL1.
The transition between each level is explicit — done via <code class="language-plaintext highlighter-rouge">ERET</code> with carefully prepared
SPSR and ELR registers. On AMD64, privilege rings exist but the boot chain does not engage
them until the OS sets them up.</p>

<p>The other major difference is the flexibility — and the resulting complexity. Rockchip’s
<code class="language-plaintext highlighter-rouge">return-to-bootrom</code> mechanism, the TPL/SPL split, the fixed-offset sector layout: none of
this is mandated by the ARM architecture. Each vendor solves these problems differently,
which is both ARM’s strength and the main reason ARM boot sequences are notoriously hard
to document generically.</p>

<p>In Part Three we will look at ARM Secure Boot and ATF in more detail, and after that
compare both architectures in the context of UEFI.</p>

<hr />

<h2 id="resources">Resources</h2>

<ol>
  <li><a href="https://github.com/ARM-software/ebbr">Embedded Base Boot Requirements (EBBR) Specification</a></li>
  <li><a href="https://developer.arm.com/documentation/den0044/e/">Server Base Boot Requirements (SBBR) v1.2</a></li>
  <li><a href="https://man7.org/linux/man-pages/man3/longjmp.3p.html">longjmp man page</a></li>
  <li><a href="https://www.rock-chips.com/a/en/index.html">Rockchip RK3399 Technical Reference Manual</a></li>
  <li><a href="https://developer.arm.com/documentation/ddi0487/latest">ARM Architecture Reference Manual — ARMv8</a></li>
  <li><a href="https://trustedfirmware-a.readthedocs.io/">ARM Trusted Firmware documentation</a></li>
  <li><a href="https://github.com/NetBSD/src">NetBSD source code</a></li>
</ol>]]></content><author><name>Maciej Grochowski</name></author><category term="kernel" /><summary type="html"><![CDATA[Note: This article was written in 2020 as a companion to Part One (AMD64) but was never published. The technical content reflects the state of NetBSD and the RK3399 platform at that time. Published here now without major changes.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://page-fault.io/assets/social-preview.png" /><media:content medium="image" url="https://page-fault.io/assets/social-preview.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">rust_tlplib and tlp-tool 0.5.0: PCIe 6.0 Flit Mode</title><link href="https://page-fault.io/hardware/2026/03/17/rust-tlplib-tlp-tool-0.5.0.html" rel="alternate" type="text/html" title="rust_tlplib and tlp-tool 0.5.0: PCIe 6.0 Flit Mode" /><published>2026-03-17T12:00:00+00:00</published><updated>2026-03-17T12:00:00+00:00</updated><id>https://page-fault.io/hardware/2026/03/17/rust-tlplib-tlp-tool-0.5.0</id><content type="html" xml:base="https://page-fault.io/hardware/2026/03/17/rust-tlplib-tlp-tool-0.5.0.html"><![CDATA[<p>The <a href="https://github.com/mmpg-x86/rust_tlplib/releases/tag/v0.5.0">rust_tlplib</a> and
<a href="https://github.com/mmpg-x86/tlp-tool/releases/tag/v0.5.0">rtlp-tool</a> libraries had not seen
a release in a long time. After getting Gen 4 TLP parsing to a point where it covered
everything I needed day-to-day, I left it there. Good enough was good enough.</p>

<p>Then Gen 6 started showing up everywhere — conferences, hardware announcements, specs dropping
into the mailbox. Flit mode is not a minor revision; it is a different framing model entirely,
and the TLP format that goes inside it was reworked to match. I had to spend a real amount of
time working through the spec before I understood what had actually changed and what had
stayed the same. The PCIe 6.0 spec calls the relevant section “TLP Format Revisited.” That
name carries more weight than it looks.</p>

<p>The other thing that slows you down when you return to old code is the old code itself. I
remembered writing some of it under time pressure and taking shortcuts that made sense at the
time. Seeing them again — as someone who had to then build on top of them — was educational
in the way that finding your own bugs always is. So the API rename pass in this release is
partly Gen 6 groundwork and partly old debts getting paid.</p>

<p>Both tracks took longer than expected. The release is the result.</p>

<p>If you have been using either library for Gen 1–5 parsing, nothing changes — the existing
paths are untouched. What follows is what is new.</p>

<h2 id="flit-mode-what-actually-changed">Flit Mode: What Actually Changed</h2>

<p>The <a href="/hardware/2022/08/07/how-to-parse-pcie-tlps.html">previous post on parsing TLPs</a> covers the
traditional format: DW0 encodes FMT[2:0] + TYPE[4:0], the rest of the header follows a
layout determined by that combination, and parsing is a matter of reading the right bits
from the right dwords.</p>

<p>PCIe 6.0 changes the framing entirely. To support PAM4 signaling at 64 GT/s, the spec
introduced flits — fixed 256-byte containers with their own forward error correction and CRC.
TLPs are no longer delimited by STP/END sequences at the data link layer; they are packed into
flits. The TLP format itself was revised to match: DW0 in flit mode carries a different set
of fields, there are thirteen type codes rather than the FMT+TYPE matrix of earlier
generations, and Optional Header Components (OHC) replace TLP prefixes as the extension
mechanism.</p>

<p>None of this is obvious from a quick read. The spec’s framing of it as “revisited” undersells
how much the mental model needs to shift.</p>

<h2 id="library-changes">Library Changes</h2>

<p>The new types in <code class="language-plaintext highlighter-rouge">rust_tlplib</code>:</p>

<ul>
  <li><code class="language-plaintext highlighter-rouge">FlitDW0</code> — flit-mode DW0 with the Gen 6 field layout</li>
  <li><code class="language-plaintext highlighter-rouge">FlitTlpType</code> — the thirteen flit type codes (NOPs, memory, atomics, messages)</li>
  <li><code class="language-plaintext highlighter-rouge">FlitOhcA</code> — Optional Header Component type A, the primary OHC variant</li>
  <li><code class="language-plaintext highlighter-rouge">FlitStreamWalker</code> — walks a byte stream extracting TLPs packed inside a flit</li>
</ul>

<p>Mode dispatch is now on <code class="language-plaintext highlighter-rouge">TlpPacket</code> directly:</p>

<div class="language-rust highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">match</span> <span class="n">packet</span><span class="nf">.mode</span><span class="p">()</span> <span class="p">{</span>
    <span class="nn">TlpMode</span><span class="p">::</span><span class="nf">Standard</span><span class="p">(</span><span class="n">tlp</span><span class="p">)</span> <span class="k">=&gt;</span> <span class="p">{</span> <span class="cm">/* Gen 1–5 handling */</span> <span class="p">}</span>
    <span class="nn">TlpMode</span><span class="p">::</span><span class="nf">Flit</span><span class="p">(</span><span class="n">flit</span><span class="p">)</span>    <span class="k">=&gt;</span> <span class="p">{</span> <span class="cm">/* Gen 6 handling */</span>   <span class="p">}</span>
<span class="p">}</span>
</code></pre></div></div>

<p>The release also includes the API rename pass mentioned above. Existing <code class="language-plaintext highlighter-rouge">get_*</code> accessors are
deprecated in favor of canonical names — the library ships a migration table in the release
notes, and the old names still compile with deprecation warnings for now. 212 tests pass on
Rust 1.85.</p>

<h2 id="rtlp-tool">rtlp-tool</h2>

<p>The command-line tool gains a <code class="language-plaintext highlighter-rouge">--flit</code> flag:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>$ rtlp-tool --flit -i "03 00 00 01 01 00 0A FF AB CD 12 34"
</code></pre></div></div>

<p>All output modes — table, JSON, CSV — support the flag. JSON output now includes
<code class="language-plaintext highlighter-rouge">"flit_mode": true</code> to make it unambiguous which parser was used.</p>

<p>This release also cleans up several bugs that had been sitting in the non-flit path: a DW0
byte-slice extraction reading from the wrong offset, a crash on odd-nibble hex input, missing
mutual exclusion between <code class="language-plaintext highlighter-rouge">--aer</code> and <code class="language-plaintext highlighter-rouge">--lspci</code>, and JSON output that was generating malformed
escapes for certain field values.</p>

<h2 id="installing">Installing</h2>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>cargo install rtlp_tool
</code></pre></div></div>

<p>Binaries for Linux (x86_64, aarch64), macOS (Intel and Apple Silicon), Windows, and FreeBSD
are on the <a href="https://github.com/mmpg-x86/tlp-tool/releases/tag/v0.5.0">release page</a>.
Debian/Ubuntu <code class="language-plaintext highlighter-rouge">.deb</code> and RPM packages are also available if you would rather not go through
Cargo.</p>]]></content><author><name>Maciej Grochowski</name></author><category term="hardware" /><summary type="html"><![CDATA[The rust_tlplib and rtlp-tool libraries had not seen a release in a long time. After getting Gen 4 TLP parsing to a point where it covered everything I needed day-to-day, I left it there. Good enough was good enough.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://page-fault.io/assets/social-preview.png" /><media:content medium="image" url="https://page-fault.io/assets/social-preview.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">A Transformer Block in CUDA</title><link href="https://page-fault.io/ml/2025/12/28/transformer-block-cuda.html" rel="alternate" type="text/html" title="A Transformer Block in CUDA" /><published>2025-12-28T21:00:00+00:00</published><updated>2025-12-28T21:00:00+00:00</updated><id>https://page-fault.io/ml/2025/12/28/transformer-block-cuda</id><content type="html" xml:base="https://page-fault.io/ml/2025/12/28/transformer-block-cuda.html"><![CDATA[<p>December, somewhere over the Pacific on the way to Tokyo. The <em>letGPU</em> challenge
open in one browser tab, Ro Salaverry’s <em>The Scaling Era: An Oral History of AI,
2019–2025</em> on my Kindle. No plan beyond getting a feel for the challenge — I
poked at it for a while, read a few chapters, and fell asleep somewhere east of
the dateline. Woke up on approach to Narita.</p>

<p>The book stuck with me, though. It is a collection of firsthand accounts from
the people who built the systems that now run most of the software industry,
and the recurring theme is the same across a dozen voices: the transformer
architecture arrived, and then everything else followed from scale.</p>

<p>That is what pulled me back to the challenge once I had a hotel and a power
outlet. Not because I do this for a living — machine learning is not my area —
but because the underlying question is squarely in territory I care about:
what does this look like in hardware, and why does it work the way it does?</p>

<p>A <a href="/ml/2025/11/15/letgpu-gpt2-transformer-block.html">previous post</a>
covered the GPT-2 decoder block at a conceptual level: what the LetGPU challenge
is asking for, how data flows through the block, what each parameter group does,
and why the structure is what it is. This post assumes that context — either you
have read it, or you already know GPT-2 well enough that you do not need it. Here
we go deeper: the actual CUDA implementation, kernel by kernel, with the concrete
hardware trade-offs that make each design decision interesting. The code implements
a single transformer block at GPT-2 small scale. No PyTorch, no cuBLAS — just
kernels.</p>

<h2 id="the-block-concretely">The Block, Concretely</h2>

<p>The implementation is GPT-2 small: hidden dimension D=768, twelve attention
heads (H=12), head dimension Dh=64, feed-forward dimension FF=3072. One
forward pass through a block takes a token sequence of shape <code class="language-plaintext highlighter-rouge">(seq_len, 768)</code>
and produces the same shape.</p>

<p>This is a Pre-LN architecture — layer norm applied <em>before</em> each sublayer, not
after. The original paper used Post-LN (norm after the residual add). The
difference matters: Pre-LN feeds normalized inputs into the attention and FFN,
which keeps gradient flow cleaner in deeper models. GPT-2 and most subsequent
work moved to Pre-LN; the original formulation is now mostly a historical
artifact.</p>

<p>The ten steps of the forward pass:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>x  ──►  LN1  ──►  QKV proj  ──►  MHA  ──►  O proj  ──►  add(x, ·)  ──►  x1
x1 ──►  LN2  ──►  FC proj   ──►  GELU  ──►  out proj  ──►  add(x1, ·)  ──►  output
</code></pre></div></div>

<p>In kernel terms:</p>

<ol>
  <li><code class="language-plaintext highlighter-rouge">layernorm768_kernel(x)</code> → <code class="language-plaintext highlighter-rouge">ln1</code></li>
  <li><code class="language-plaintext highlighter-rouge">matmul_bias_tiled(ln1, W_qkv)</code> → <code class="language-plaintext highlighter-rouge">qkv</code> — shape <code class="language-plaintext highlighter-rouge">(seq_len, 2304)</code>, packing Q, K, V</li>
  <li><code class="language-plaintext highlighter-rouge">mha_nocausal_kernel(qkv)</code> → <code class="language-plaintext highlighter-rouge">attn</code> — shape <code class="language-plaintext highlighter-rouge">(seq_len, 768)</code></li>
  <li><code class="language-plaintext highlighter-rouge">matmul_bias_tiled(attn, W_attn)</code> → <code class="language-plaintext highlighter-rouge">proj</code></li>
  <li><code class="language-plaintext highlighter-rouge">add_kernel(x, proj)</code> → <code class="language-plaintext highlighter-rouge">x1</code> — first residual</li>
  <li><code class="language-plaintext highlighter-rouge">layernorm768_kernel(x1)</code> → <code class="language-plaintext highlighter-rouge">ln2</code></li>
  <li><code class="language-plaintext highlighter-rouge">matmul_bias_tiled(ln2, W_fc)</code> → <code class="language-plaintext highlighter-rouge">ff1</code> — shape <code class="language-plaintext highlighter-rouge">(seq_len, 3072)</code></li>
  <li><code class="language-plaintext highlighter-rouge">gelu_kernel(ff1)</code> — in-place</li>
  <li><code class="language-plaintext highlighter-rouge">matmul_bias_tiled(ff1, W_proj)</code> → <code class="language-plaintext highlighter-rouge">ff2</code> — shape <code class="language-plaintext highlighter-rouge">(seq_len, 768)</code></li>
  <li><code class="language-plaintext highlighter-rouge">add_kernel(x1, ff2)</code> → <code class="language-plaintext highlighter-rouge">output</code> — second residual</li>
</ol>

<p>The rest of this post walks each step in turn.</p>

<h2 id="block-reduction">Block Reduction</h2>

<p>Almost every kernel needs a single value derived from all threads in a block —
a sum for layer norm, a maximum for softmax numerical stability. The standard
tool is a tree reduction over shared memory:</p>

<div class="language-cuda highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">__device__</span> <span class="k">__forceinline__</span> <span class="kt">float</span> <span class="nf">block_reduce_sum</span><span class="p">(</span><span class="kt">float</span> <span class="n">v</span><span class="p">)</span> <span class="p">{</span>
    <span class="k">__shared__</span> <span class="kt">float</span> <span class="n">smem</span><span class="p">[</span><span class="mi">1024</span><span class="p">];</span>
    <span class="kt">int</span> <span class="n">tid</span> <span class="o">=</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">;</span>
    <span class="n">smem</span><span class="p">[</span><span class="n">tid</span><span class="p">]</span> <span class="o">=</span> <span class="n">v</span><span class="p">;</span>
    <span class="n">__syncthreads</span><span class="p">();</span>
    <span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">stride</span> <span class="o">=</span> <span class="n">blockDim</span><span class="p">.</span><span class="n">x</span> <span class="o">/</span> <span class="mi">2</span><span class="p">;</span> <span class="n">stride</span> <span class="o">&gt;</span> <span class="mi">0</span><span class="p">;</span> <span class="n">stride</span> <span class="o">&gt;&gt;=</span> <span class="mi">1</span><span class="p">)</span> <span class="p">{</span>
        <span class="k">if</span> <span class="p">(</span><span class="n">tid</span> <span class="o">&lt;</span> <span class="n">stride</span><span class="p">)</span> <span class="n">smem</span><span class="p">[</span><span class="n">tid</span><span class="p">]</span> <span class="o">+=</span> <span class="n">smem</span><span class="p">[</span><span class="n">tid</span> <span class="o">+</span> <span class="n">stride</span><span class="p">];</span>
        <span class="n">__syncthreads</span><span class="p">();</span>
    <span class="p">}</span>
    <span class="k">return</span> <span class="n">smem</span><span class="p">[</span><span class="mi">0</span><span class="p">];</span>
<span class="p">}</span>
</code></pre></div></div>

<p>Each thread writes its local value to shared memory, then the loop halves the
active thread count on each iteration: threads 0–511 add from threads 512–1023,
then 0–255 add from 256–511, and so on. After log₂(blockDim.x) iterations,
<code class="language-plaintext highlighter-rouge">smem[0]</code> holds the total. Every thread in the block reads that result.</p>

<p>The cost is the <code class="language-plaintext highlighter-rouge">__syncthreads()</code> calls — a block-wide barrier that stalls
until every thread in the block reaches it. Here there are log₂(256) = eight
barriers per reduction, and layer norm alone runs two reductions. Not free, but
unavoidable: the alternative is letting threads read values their neighbors
have not written yet.</p>

<p><code class="language-plaintext highlighter-rouge">block_reduce_max</code> is identical with <code class="language-plaintext highlighter-rouge">fmaxf</code> in place of addition.</p>

<h2 id="layer-normalization">Layer Normalization</h2>

<p>One block per token. All threads in the block cooperate on a single
768-element normalization:</p>

<div class="language-cuda highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1">// Pass 1: mean</span>
<span class="kt">float</span> <span class="n">sum</span> <span class="o">=</span> <span class="mf">0.0f</span><span class="p">;</span>
<span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">i</span> <span class="o">=</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">;</span> <span class="n">i</span> <span class="o">&lt;</span> <span class="mi">768</span><span class="p">;</span> <span class="n">i</span> <span class="o">+=</span> <span class="n">blockDim</span><span class="p">.</span><span class="n">x</span><span class="p">)</span> <span class="n">sum</span> <span class="o">+=</span> <span class="n">x</span><span class="p">[</span><span class="n">i</span><span class="p">];</span>
<span class="n">sum</span> <span class="o">=</span> <span class="n">block_reduce_sum</span><span class="p">(</span><span class="n">sum</span><span class="p">);</span>
<span class="kt">float</span> <span class="n">mean</span> <span class="o">=</span> <span class="n">sum</span> <span class="o">/</span> <span class="mf">768.0f</span><span class="p">;</span>

<span class="c1">// Pass 2: variance, then normalize</span>
<span class="kt">float</span> <span class="n">vsum</span> <span class="o">=</span> <span class="mf">0.0f</span><span class="p">;</span>
<span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">i</span> <span class="o">=</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">;</span> <span class="n">i</span> <span class="o">&lt;</span> <span class="mi">768</span><span class="p">;</span> <span class="n">i</span> <span class="o">+=</span> <span class="n">blockDim</span><span class="p">.</span><span class="n">x</span><span class="p">)</span> <span class="p">{</span>
    <span class="kt">float</span> <span class="n">d</span> <span class="o">=</span> <span class="n">x</span><span class="p">[</span><span class="n">i</span><span class="p">]</span> <span class="o">-</span> <span class="n">mean</span><span class="p">;</span> <span class="n">vsum</span> <span class="o">+=</span> <span class="n">d</span> <span class="o">*</span> <span class="n">d</span><span class="p">;</span>
<span class="p">}</span>
<span class="n">vsum</span> <span class="o">=</span> <span class="n">block_reduce_sum</span><span class="p">(</span><span class="n">vsum</span><span class="p">);</span>
<span class="kt">float</span> <span class="n">inv_std</span> <span class="o">=</span> <span class="n">rsqrtf</span><span class="p">(</span><span class="n">vsum</span> <span class="o">/</span> <span class="mf">768.0f</span> <span class="o">+</span> <span class="mf">1e-5f</span><span class="p">);</span>

<span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">i</span> <span class="o">=</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">;</span> <span class="n">i</span> <span class="o">&lt;</span> <span class="mi">768</span><span class="p">;</span> <span class="n">i</span> <span class="o">+=</span> <span class="n">blockDim</span><span class="p">.</span><span class="n">x</span><span class="p">)</span>
    <span class="n">y</span><span class="p">[</span><span class="n">i</span><span class="p">]</span> <span class="o">=</span> <span class="p">(</span><span class="n">x</span><span class="p">[</span><span class="n">i</span><span class="p">]</span> <span class="o">-</span> <span class="n">mean</span><span class="p">)</span> <span class="o">*</span> <span class="n">inv_std</span> <span class="o">*</span> <span class="n">gamma</span><span class="p">[</span><span class="n">i</span><span class="p">]</span> <span class="o">+</span> <span class="n">beta</span><span class="p">[</span><span class="n">i</span><span class="p">];</span>
</code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">rsqrtf</code> is the reciprocal square root — a single hardware instruction on any
modern GPU, cheaper than computing <code class="language-plaintext highlighter-rouge">sqrtf</code> and then dividing. The epsilon
<code class="language-plaintext highlighter-rouge">1e-5f</code> keeps the denominator away from zero. <code class="language-plaintext highlighter-rouge">gamma</code> and <code class="language-plaintext highlighter-rouge">beta</code> are the
learned affine parameters applied after normalization.</p>

<h2 id="tiled-matrix-multiply">Tiled Matrix Multiply</h2>

<p>The projection steps — QKV, output, FFN up and down — all go through one
templated kernel. The core idea is tiling: loading a 16×16 submatrix of A and
a 16×16 submatrix of B into shared memory before computing their product, so
the 256 threads in the block see each value once from slow global memory and
reuse it sixteen times from fast shared memory.</p>

<div class="language-cuda highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">template</span><span class="o">&lt;</span><span class="kt">int</span> <span class="n">TILE</span><span class="p">&gt;</span>
<span class="k">__global__</span> <span class="kt">void</span> <span class="nf">matmul_bias_tiled</span><span class="p">(</span><span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="k">__restrict__</span> <span class="n">A</span><span class="p">,</span>
                                 <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="k">__restrict__</span> <span class="n">B</span><span class="p">,</span>
                                 <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="k">__restrict__</span> <span class="n">bias</span><span class="p">,</span>
                                 <span class="kt">float</span><span class="o">*</span> <span class="k">__restrict__</span> <span class="n">C</span><span class="p">,</span>
                                 <span class="kt">int</span> <span class="n">M</span><span class="p">,</span> <span class="kt">int</span> <span class="n">K</span><span class="p">,</span> <span class="kt">int</span> <span class="n">N</span><span class="p">)</span> <span class="p">{</span>
    <span class="k">__shared__</span> <span class="kt">float</span> <span class="n">As</span><span class="p">[</span><span class="n">TILE</span><span class="p">][</span><span class="n">TILE</span><span class="p">];</span>
    <span class="k">__shared__</span> <span class="kt">float</span> <span class="n">Bs</span><span class="p">[</span><span class="n">TILE</span><span class="p">][</span><span class="n">TILE</span><span class="p">];</span>
    <span class="kt">float</span> <span class="n">acc</span> <span class="o">=</span> <span class="mf">0.0f</span><span class="p">;</span>

    <span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">k0</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="n">k0</span> <span class="o">&lt;</span> <span class="n">K</span><span class="p">;</span> <span class="n">k0</span> <span class="o">+=</span> <span class="n">TILE</span><span class="p">)</span> <span class="p">{</span>
        <span class="n">As</span><span class="p">[</span><span class="n">threadIdx</span><span class="p">.</span><span class="n">y</span><span class="p">][</span><span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">]</span> <span class="o">=</span> <span class="cm">/* A[row, k0+threadIdx.x] */</span><span class="p">;</span>
        <span class="n">Bs</span><span class="p">[</span><span class="n">threadIdx</span><span class="p">.</span><span class="n">y</span><span class="p">][</span><span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">]</span> <span class="o">=</span> <span class="cm">/* B[k0+threadIdx.y, col] */</span><span class="p">;</span>
        <span class="n">__syncthreads</span><span class="p">();</span>

        <span class="cp">#pragma unroll
</span>        <span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">k</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="n">k</span> <span class="o">&lt;</span> <span class="n">TILE</span><span class="p">;</span> <span class="o">++</span><span class="n">k</span><span class="p">)</span>
            <span class="n">acc</span> <span class="o">+=</span> <span class="n">As</span><span class="p">[</span><span class="n">threadIdx</span><span class="p">.</span><span class="n">y</span><span class="p">][</span><span class="n">k</span><span class="p">]</span> <span class="o">*</span> <span class="n">Bs</span><span class="p">[</span><span class="n">k</span><span class="p">][</span><span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">];</span>
        <span class="n">__syncthreads</span><span class="p">();</span>
    <span class="p">}</span>
    <span class="n">C</span><span class="p">[</span><span class="n">row</span> <span class="o">*</span> <span class="n">N</span> <span class="o">+</span> <span class="n">col</span><span class="p">]</span> <span class="o">=</span> <span class="n">acc</span> <span class="o">+</span> <span class="n">bias</span><span class="p">[</span><span class="n">col</span><span class="p">];</span>
<span class="p">}</span>
</code></pre></div></div>

<p>The thread at <code class="language-plaintext highlighter-rouge">(threadIdx.y, threadIdx.x)</code> is responsible for one output
element at <code class="language-plaintext highlighter-rouge">(row, col)</code>. It accumulates the dot product of a row of A with a
column of B, advancing through K in strides of TILE=16. The two
<code class="language-plaintext highlighter-rouge">__syncthreads()</code> per iteration are load-fence and store-fence: the first
ensures all threads have finished loading into shared memory before any thread
starts reading from it; the second ensures the compute is done before the next
iteration overwrites the tiles.</p>

<p><code class="language-plaintext highlighter-rouge">#pragma unroll</code> on the inner loop tells the compiler to unroll the 16
iterations at compile time, generating 16 independent FMAs rather than a loop
with a branch and counter update on each pass.</p>

<p><code class="language-plaintext highlighter-rouge">__restrict__</code> on the pointer arguments is a promise to the compiler that A, B,
and C do not alias — the compiler can assume no write to C affects what is read
from A or B, enabling more aggressive load scheduling.</p>

<h2 id="multi-head-attention">Multi-Head Attention</h2>

<p>This is the interesting one. One CUDA block per <code class="language-plaintext highlighter-rouge">(token, head)</code> pair — the
launch grid is <code class="language-plaintext highlighter-rouge">(seq_len, 12)</code>. Each block computes one row of the attention
output for one head.</p>

<p>The attention score for query token <code class="language-plaintext highlighter-rouge">t</code> against key token <code class="language-plaintext highlighter-rouge">s</code> in head <code class="language-plaintext highlighter-rouge">h</code> is:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>score(t, s) = dot(Q_t_h, K_s_h) / sqrt(64)
</code></pre></div></div>

<p>Then softmax across all <code class="language-plaintext highlighter-rouge">s</code>, then weighted sum of the value vectors <code class="language-plaintext highlighter-rouge">V_s_h</code>.</p>

<p>Softmax has a numerical problem: if scores are large, <code class="language-plaintext highlighter-rouge">exp(score)</code> overflows.
The standard fix is to subtract the maximum before exponentiating. That
requires knowing the maximum, which requires a pass over all scores. So the
kernel runs three passes:</p>

<div class="language-cuda highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1">// Pass 1: compute raw scores, find max</span>
<span class="kt">float</span> <span class="n">local_max</span> <span class="o">=</span> <span class="o">-</span><span class="mf">1e30f</span><span class="p">;</span>
<span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">s</span> <span class="o">=</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">;</span> <span class="n">s</span> <span class="o">&lt;</span> <span class="n">seq_len</span><span class="p">;</span> <span class="n">s</span> <span class="o">+=</span> <span class="n">blockDim</span><span class="p">.</span><span class="n">x</span><span class="p">)</span> <span class="p">{</span>
    <span class="c1">// compute dot(Q[t,h], K[s,h]) and scale</span>
    <span class="n">scores</span><span class="p">[</span><span class="n">s</span><span class="p">]</span> <span class="o">=</span> <span class="n">dot</span> <span class="o">*</span> <span class="n">inv_sqrt_dh</span><span class="p">;</span>
    <span class="n">local_max</span> <span class="o">=</span> <span class="n">fmaxf</span><span class="p">(</span><span class="n">local_max</span><span class="p">,</span> <span class="n">scores</span><span class="p">[</span><span class="n">s</span><span class="p">]);</span>
<span class="p">}</span>
<span class="kt">float</span> <span class="n">max_sc</span> <span class="o">=</span> <span class="n">block_reduce_max</span><span class="p">(</span><span class="n">local_max</span><span class="p">);</span>

<span class="c1">// Pass 2: exp(score - max), accumulate sum</span>
<span class="kt">float</span> <span class="n">local_sum</span> <span class="o">=</span> <span class="mf">0.0f</span><span class="p">;</span>
<span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">s</span> <span class="o">=</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">;</span> <span class="n">s</span> <span class="o">&lt;</span> <span class="n">seq_len</span><span class="p">;</span> <span class="n">s</span> <span class="o">+=</span> <span class="n">blockDim</span><span class="p">.</span><span class="n">x</span><span class="p">)</span> <span class="p">{</span>
    <span class="kt">float</span> <span class="n">e</span> <span class="o">=</span> <span class="n">expf</span><span class="p">(</span><span class="n">scores</span><span class="p">[</span><span class="n">s</span><span class="p">]</span> <span class="o">-</span> <span class="n">max_sc</span><span class="p">);</span>
    <span class="n">scores</span><span class="p">[</span><span class="n">s</span><span class="p">]</span> <span class="o">=</span> <span class="n">e</span><span class="p">;</span>         <span class="c1">// reuse the shared buffer</span>
    <span class="n">local_sum</span> <span class="o">+=</span> <span class="n">e</span><span class="p">;</span>
<span class="p">}</span>
<span class="kt">float</span> <span class="n">inv_sum</span> <span class="o">=</span> <span class="mf">1.0f</span> <span class="o">/</span> <span class="n">block_reduce_sum</span><span class="p">(</span><span class="n">local_sum</span><span class="p">);</span>

<span class="c1">// Pass 3: weighted sum of V (first 64 threads only)</span>
<span class="k">if</span> <span class="p">(</span><span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span> <span class="o">&lt;</span> <span class="n">Dh</span><span class="p">)</span> <span class="p">{</span>
    <span class="kt">float</span> <span class="n">acc</span> <span class="o">=</span> <span class="mf">0.0f</span><span class="p">;</span>
    <span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">s</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="n">s</span> <span class="o">&lt;</span> <span class="n">seq_len</span><span class="p">;</span> <span class="o">++</span><span class="n">s</span><span class="p">)</span>
        <span class="n">acc</span> <span class="o">+=</span> <span class="n">scores</span><span class="p">[</span><span class="n">s</span><span class="p">]</span> <span class="o">*</span> <span class="n">inv_sum</span> <span class="o">*</span> <span class="n">V</span><span class="p">[</span><span class="n">s</span><span class="p">][</span><span class="n">h</span><span class="p">][</span><span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">];</span>
    <span class="n">attn_out</span><span class="p">[</span><span class="n">t</span> <span class="o">*</span> <span class="n">D</span> <span class="o">+</span> <span class="n">h</span> <span class="o">*</span> <span class="n">Dh</span> <span class="o">+</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">]</span> <span class="o">=</span> <span class="n">acc</span><span class="p">;</span>
<span class="p">}</span>
</code></pre></div></div>

<p>The <code class="language-plaintext highlighter-rouge">scores</code> buffer lives in dynamic shared memory (<code class="language-plaintext highlighter-rouge">extern __shared__</code>),
allocated at launch time as <code class="language-plaintext highlighter-rouge">seq_len * sizeof(float)</code>. The size is not known
at compile time; the kernel receives it as the third argument to the <code class="language-plaintext highlighter-rouge">&lt;&lt;&lt;&gt;&gt;&gt;</code>
launch syntax: <code class="language-plaintext highlighter-rouge">mha_nocausal_kernel&lt;&lt;&lt;grid, block, seq_len * sizeof(float)&gt;&gt;&gt;</code>.
Note that shared memory is limited per block (48 KB on most devices, up to 96 KB
with explicit opt-in). For large <code class="language-plaintext highlighter-rouge">seq_len</code> this allocation will exceed the limit —
the implementation as written is suitable for the GPT-2 Small context lengths
(up to 1024 tokens → 4 KB), not for arbitrarily long sequences.</p>

<p>The QKV layout is interleaved by position: each token’s Q, K, and V are
contiguous. The head dimension is packed within each section. Pointer
arithmetic for Q at token <code class="language-plaintext highlighter-rouge">t</code>, head <code class="language-plaintext highlighter-rouge">h</code> is:
<code class="language-plaintext highlighter-rouge">qkv + t * 2304 + 0 * 768 + h * 64</code>, and similarly for K (offset <code class="language-plaintext highlighter-rouge">1*768</code>)
and V (offset <code class="language-plaintext highlighter-rouge">2*768</code>).</p>

<p>Pass 3 only uses the first 64 threads (one per head dimension). The remaining
threads in the block are idle for that phase. This is a deliberate simplicity
trade-off — for a challenge implementation it is fine; production kernels would
use the idle threads to pipeline across heads or fuse operations.</p>

<h2 id="gelu">GELU</h2>

<p>The feed-forward sublayer uses GELU as its activation function. The true
definition involves the error function <code class="language-plaintext highlighter-rouge">erf</code>, which is expensive. The standard
approximation:</p>

<div class="language-cuda highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">static</span> <span class="k">__device__</span> <span class="k">__forceinline__</span> <span class="kt">float</span> <span class="nf">gelu_tanh</span><span class="p">(</span><span class="kt">float</span> <span class="n">x</span><span class="p">)</span> <span class="p">{</span>
    <span class="k">const</span> <span class="kt">float</span> <span class="n">k</span> <span class="o">=</span> <span class="mf">0.7978845608028654f</span><span class="p">;</span> <span class="c1">// sqrt(2/pi)</span>
    <span class="kt">float</span> <span class="n">x3</span> <span class="o">=</span> <span class="n">x</span> <span class="o">*</span> <span class="n">x</span> <span class="o">*</span> <span class="n">x</span><span class="p">;</span>
    <span class="k">return</span> <span class="mf">0.5f</span> <span class="o">*</span> <span class="n">x</span> <span class="o">*</span> <span class="p">(</span><span class="mf">1.0f</span> <span class="o">+</span> <span class="n">tanhf</span><span class="p">(</span><span class="n">k</span> <span class="o">*</span> <span class="p">(</span><span class="n">x</span> <span class="o">+</span> <span class="mf">0.044715f</span> <span class="o">*</span> <span class="n">x3</span><span class="p">)));</span>
<span class="p">}</span>
</code></pre></div></div>

<p>This is a Padé approximant to <code class="language-plaintext highlighter-rouge">erf(x/sqrt(2))</code>. The maximum absolute error
relative to the true GELU is small enough to be irrelevant for inference
(on the order of 10⁻⁴ in the active range — the exact bound depends on the
measurement interval, but PyTorch uses this same approximation in production).
<code class="language-plaintext highlighter-rouge">tanhf</code> is hardware-accelerated on every CUDA-capable device; the full
expression is five floating-point operations. The kernel applies it elementwise
across the <code class="language-plaintext highlighter-rouge">(seq_len, 3072)</code> intermediate tensor.</p>

<h2 id="putting-it-together">Putting It Together</h2>

<p>The <code class="language-plaintext highlighter-rouge">solve</code> function allocates eight intermediate tensors on the device,
launches the kernels in sequence, and frees them. The weight layout is a flat
device buffer. Offsets are computed from the architecture constants:</p>

<table>
  <thead>
    <tr>
      <th>Weights</th>
      <th>Offset</th>
      <th>Size</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>γ₁, β₁</td>
      <td>0</td>
      <td>768 + 768</td>
    </tr>
    <tr>
      <td>W_qkv</td>
      <td>1,536</td>
      <td>768 × 2,304</td>
    </tr>
    <tr>
      <td>b_qkv</td>
      <td>1,771,008</td>
      <td>2,304</td>
    </tr>
    <tr>
      <td>W_attn</td>
      <td>1,773,312</td>
      <td>768 × 768</td>
    </tr>
    <tr>
      <td>b_attn</td>
      <td>2,363,136</td>
      <td>768</td>
    </tr>
    <tr>
      <td>γ₂, β₂</td>
      <td>2,363,904</td>
      <td>768 + 768</td>
    </tr>
    <tr>
      <td>W_fc</td>
      <td>2,365,440</td>
      <td>768 × 3,072</td>
    </tr>
    <tr>
      <td>b_fc</td>
      <td>4,724,736</td>
      <td>3,072</td>
    </tr>
    <tr>
      <td>W_proj</td>
      <td>4,727,808</td>
      <td>3,072 × 768</td>
    </tr>
    <tr>
      <td>b_proj</td>
      <td>7,087,104</td>
      <td>768</td>
    </tr>
  </tbody>
</table>

<p>About 7.1 million parameters per block. GPT-2 small stacks twelve of them on
top of embeddings — 85 million parameters total. The numbers in <em>The Scaling
Era</em> that stuck with me were the orders of magnitude larger: GPT-3 at 175
billion, models that followed at trillions. The architecture is the same. The
arithmetic just runs longer.</p>

<p>The implementation here is not production code. It is not fused, not tuned for
tensor cores, and does three passes through memory where a flash attention
implementation would do one. But the patterns — shared memory tiling, tree
reductions, online softmax — are the same patterns production kernels use,
without the extra complexity that comes from optimizing for throughput at scale.</p>

<p>Sometimes it helps to see the skeleton before studying the muscle.</p>

<hr />

<h2 id="full-code">Full Code</h2>

<div class="language-cuda highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="cp">#include</span> <span class="cpf">&lt;cuda_runtime.h&gt;</span><span class="cp">
#include</span> <span class="cpf">&lt;math_constants.h&gt;</span><span class="cp">
#include</span> <span class="cpf">&lt;cmath&gt;</span><span class="cp">
</span>
<span class="cp">#define CHECK_CUDA(x) do { cudaError_t err = (x); if (err != cudaSuccess) return; } while(0)
</span>
<span class="k">static</span> <span class="k">__device__</span> <span class="k">__forceinline__</span> <span class="kt">float</span> <span class="nf">gelu_tanh</span><span class="p">(</span><span class="kt">float</span> <span class="n">x</span><span class="p">)</span> <span class="p">{</span>
    <span class="c1">// GELU(x) = 0.5*x*(1 + tanh(sqrt(2/pi)*(x + 0.044715*x^3)))</span>
    <span class="k">const</span> <span class="kt">float</span> <span class="n">k</span> <span class="o">=</span> <span class="mf">0.7978845608028654f</span><span class="p">;</span> <span class="c1">// sqrt(2/pi)</span>
    <span class="kt">float</span> <span class="n">x3</span> <span class="o">=</span> <span class="n">x</span> <span class="o">*</span> <span class="n">x</span> <span class="o">*</span> <span class="n">x</span><span class="p">;</span>
    <span class="k">return</span> <span class="mf">0.5f</span> <span class="o">*</span> <span class="n">x</span> <span class="o">*</span> <span class="p">(</span><span class="mf">1.0f</span> <span class="o">+</span> <span class="n">tanhf</span><span class="p">(</span><span class="n">k</span> <span class="o">*</span> <span class="p">(</span><span class="n">x</span> <span class="o">+</span> <span class="mf">0.044715f</span> <span class="o">*</span> <span class="n">x3</span><span class="p">)));</span>
<span class="p">}</span>

<span class="c1">// ---------------------------</span>
<span class="c1">// Reductions (block-wide)</span>
<span class="c1">// ---------------------------</span>
<span class="k">__device__</span> <span class="k">__forceinline__</span> <span class="kt">float</span> <span class="nf">block_reduce_sum</span><span class="p">(</span><span class="kt">float</span> <span class="n">v</span><span class="p">)</span> <span class="p">{</span>
    <span class="c1">// assumes blockDim.x &lt;= 1024</span>
    <span class="k">__shared__</span> <span class="kt">float</span> <span class="n">smem</span><span class="p">[</span><span class="mi">1024</span><span class="p">];</span>
    <span class="kt">int</span> <span class="n">tid</span> <span class="o">=</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">;</span>
    <span class="n">smem</span><span class="p">[</span><span class="n">tid</span><span class="p">]</span> <span class="o">=</span> <span class="n">v</span><span class="p">;</span>
    <span class="n">__syncthreads</span><span class="p">();</span>
    <span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">stride</span> <span class="o">=</span> <span class="n">blockDim</span><span class="p">.</span><span class="n">x</span> <span class="o">/</span> <span class="mi">2</span><span class="p">;</span> <span class="n">stride</span> <span class="o">&gt;</span> <span class="mi">0</span><span class="p">;</span> <span class="n">stride</span> <span class="o">&gt;&gt;=</span> <span class="mi">1</span><span class="p">)</span> <span class="p">{</span>
        <span class="k">if</span> <span class="p">(</span><span class="n">tid</span> <span class="o">&lt;</span> <span class="n">stride</span><span class="p">)</span> <span class="n">smem</span><span class="p">[</span><span class="n">tid</span><span class="p">]</span> <span class="o">+=</span> <span class="n">smem</span><span class="p">[</span><span class="n">tid</span> <span class="o">+</span> <span class="n">stride</span><span class="p">];</span>
        <span class="n">__syncthreads</span><span class="p">();</span>
    <span class="p">}</span>
    <span class="k">return</span> <span class="n">smem</span><span class="p">[</span><span class="mi">0</span><span class="p">];</span>
<span class="p">}</span>

<span class="k">__device__</span> <span class="k">__forceinline__</span> <span class="kt">float</span> <span class="nf">block_reduce_max</span><span class="p">(</span><span class="kt">float</span> <span class="n">v</span><span class="p">)</span> <span class="p">{</span>
    <span class="k">__shared__</span> <span class="kt">float</span> <span class="n">smem</span><span class="p">[</span><span class="mi">1024</span><span class="p">];</span>
    <span class="kt">int</span> <span class="n">tid</span> <span class="o">=</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">;</span>
    <span class="n">smem</span><span class="p">[</span><span class="n">tid</span><span class="p">]</span> <span class="o">=</span> <span class="n">v</span><span class="p">;</span>
    <span class="n">__syncthreads</span><span class="p">();</span>
    <span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">stride</span> <span class="o">=</span> <span class="n">blockDim</span><span class="p">.</span><span class="n">x</span> <span class="o">/</span> <span class="mi">2</span><span class="p">;</span> <span class="n">stride</span> <span class="o">&gt;</span> <span class="mi">0</span><span class="p">;</span> <span class="n">stride</span> <span class="o">&gt;&gt;=</span> <span class="mi">1</span><span class="p">)</span> <span class="p">{</span>
        <span class="k">if</span> <span class="p">(</span><span class="n">tid</span> <span class="o">&lt;</span> <span class="n">stride</span><span class="p">)</span> <span class="n">smem</span><span class="p">[</span><span class="n">tid</span><span class="p">]</span> <span class="o">=</span> <span class="n">fmaxf</span><span class="p">(</span><span class="n">smem</span><span class="p">[</span><span class="n">tid</span><span class="p">],</span> <span class="n">smem</span><span class="p">[</span><span class="n">tid</span> <span class="o">+</span> <span class="n">stride</span><span class="p">]);</span>
        <span class="n">__syncthreads</span><span class="p">();</span>
    <span class="p">}</span>
    <span class="k">return</span> <span class="n">smem</span><span class="p">[</span><span class="mi">0</span><span class="p">];</span>
<span class="p">}</span>

<span class="c1">// ---------------------------</span>
<span class="c1">// LayerNorm over 768 features</span>
<span class="c1">// One block per token</span>
<span class="c1">// ---------------------------</span>
<span class="k">__global__</span> <span class="kt">void</span> <span class="nf">layernorm768_kernel</span><span class="p">(</span><span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="k">__restrict__</span> <span class="n">in</span><span class="p">,</span>
                                   <span class="kt">float</span><span class="o">*</span> <span class="k">__restrict__</span> <span class="n">out</span><span class="p">,</span>
                                   <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="k">__restrict__</span> <span class="n">gamma</span><span class="p">,</span>
                                   <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="k">__restrict__</span> <span class="n">beta</span><span class="p">,</span>
                                   <span class="kt">int</span> <span class="n">seq_len</span><span class="p">)</span> <span class="p">{</span>
    <span class="kt">int</span> <span class="n">t</span> <span class="o">=</span> <span class="n">blockIdx</span><span class="p">.</span><span class="n">x</span><span class="p">;</span>
    <span class="k">if</span> <span class="p">(</span><span class="n">t</span> <span class="o">&gt;=</span> <span class="n">seq_len</span><span class="p">)</span> <span class="k">return</span><span class="p">;</span>

    <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="n">x</span> <span class="o">=</span> <span class="n">in</span> <span class="o">+</span> <span class="n">t</span> <span class="o">*</span> <span class="mi">768</span><span class="p">;</span>
    <span class="kt">float</span><span class="o">*</span> <span class="n">y</span> <span class="o">=</span> <span class="n">out</span> <span class="o">+</span> <span class="n">t</span> <span class="o">*</span> <span class="mi">768</span><span class="p">;</span>

    <span class="c1">// sum</span>
    <span class="kt">float</span> <span class="n">sum</span> <span class="o">=</span> <span class="mf">0.0f</span><span class="p">;</span>
    <span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">i</span> <span class="o">=</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">;</span> <span class="n">i</span> <span class="o">&lt;</span> <span class="mi">768</span><span class="p">;</span> <span class="n">i</span> <span class="o">+=</span> <span class="n">blockDim</span><span class="p">.</span><span class="n">x</span><span class="p">)</span> <span class="n">sum</span> <span class="o">+=</span> <span class="n">x</span><span class="p">[</span><span class="n">i</span><span class="p">];</span>
    <span class="n">sum</span> <span class="o">=</span> <span class="n">block_reduce_sum</span><span class="p">(</span><span class="n">sum</span><span class="p">);</span>
    <span class="kt">float</span> <span class="n">mean</span> <span class="o">=</span> <span class="n">sum</span> <span class="o">/</span> <span class="mf">768.0f</span><span class="p">;</span>

    <span class="c1">// var</span>
    <span class="kt">float</span> <span class="n">vsum</span> <span class="o">=</span> <span class="mf">0.0f</span><span class="p">;</span>
    <span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">i</span> <span class="o">=</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">;</span> <span class="n">i</span> <span class="o">&lt;</span> <span class="mi">768</span><span class="p">;</span> <span class="n">i</span> <span class="o">+=</span> <span class="n">blockDim</span><span class="p">.</span><span class="n">x</span><span class="p">)</span> <span class="p">{</span>
        <span class="kt">float</span> <span class="n">d</span> <span class="o">=</span> <span class="n">x</span><span class="p">[</span><span class="n">i</span><span class="p">]</span> <span class="o">-</span> <span class="n">mean</span><span class="p">;</span>
        <span class="n">vsum</span> <span class="o">+=</span> <span class="n">d</span> <span class="o">*</span> <span class="n">d</span><span class="p">;</span>
    <span class="p">}</span>
    <span class="n">vsum</span> <span class="o">=</span> <span class="n">block_reduce_sum</span><span class="p">(</span><span class="n">vsum</span><span class="p">);</span>
    <span class="kt">float</span> <span class="n">var</span> <span class="o">=</span> <span class="n">vsum</span> <span class="o">/</span> <span class="mf">768.0f</span><span class="p">;</span>
    <span class="kt">float</span> <span class="n">inv_std</span> <span class="o">=</span> <span class="n">rsqrtf</span><span class="p">(</span><span class="n">var</span> <span class="o">+</span> <span class="mf">1e-5f</span><span class="p">);</span>

    <span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">i</span> <span class="o">=</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">;</span> <span class="n">i</span> <span class="o">&lt;</span> <span class="mi">768</span><span class="p">;</span> <span class="n">i</span> <span class="o">+=</span> <span class="n">blockDim</span><span class="p">.</span><span class="n">x</span><span class="p">)</span> <span class="p">{</span>
        <span class="kt">float</span> <span class="n">n</span> <span class="o">=</span> <span class="p">(</span><span class="n">x</span><span class="p">[</span><span class="n">i</span><span class="p">]</span> <span class="o">-</span> <span class="n">mean</span><span class="p">)</span> <span class="o">*</span> <span class="n">inv_std</span><span class="p">;</span>
        <span class="n">y</span><span class="p">[</span><span class="n">i</span><span class="p">]</span> <span class="o">=</span> <span class="n">n</span> <span class="o">*</span> <span class="n">gamma</span><span class="p">[</span><span class="n">i</span><span class="p">]</span> <span class="o">+</span> <span class="n">beta</span><span class="p">[</span><span class="n">i</span><span class="p">];</span>
    <span class="p">}</span>
<span class="p">}</span>

<span class="c1">// ---------------------------</span>
<span class="c1">// Simple tiled GEMM (row-major)</span>
<span class="c1">// C = A(M,K) * B(K,N) + bias(N)</span>
<span class="c1">// TILE=16; assumes N,K multiples of 16 for best perf (true for your shapes)</span>
<span class="c1">// ---------------------------</span>
<span class="k">template</span><span class="o">&lt;</span><span class="kt">int</span> <span class="n">TILE</span><span class="p">&gt;</span>
<span class="k">__global__</span> <span class="kt">void</span> <span class="nf">matmul_bias_tiled</span><span class="p">(</span><span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="k">__restrict__</span> <span class="n">A</span><span class="p">,</span>
                                 <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="k">__restrict__</span> <span class="n">B</span><span class="p">,</span>
                                 <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="k">__restrict__</span> <span class="n">bias</span><span class="p">,</span>
                                 <span class="kt">float</span><span class="o">*</span> <span class="k">__restrict__</span> <span class="n">C</span><span class="p">,</span>
                                 <span class="kt">int</span> <span class="n">M</span><span class="p">,</span> <span class="kt">int</span> <span class="n">K</span><span class="p">,</span> <span class="kt">int</span> <span class="n">N</span><span class="p">)</span> <span class="p">{</span>
    <span class="kt">int</span> <span class="n">row</span> <span class="o">=</span> <span class="n">blockIdx</span><span class="p">.</span><span class="n">y</span> <span class="o">*</span> <span class="n">TILE</span> <span class="o">+</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">y</span><span class="p">;</span>
    <span class="kt">int</span> <span class="n">col</span> <span class="o">=</span> <span class="n">blockIdx</span><span class="p">.</span><span class="n">x</span> <span class="o">*</span> <span class="n">TILE</span> <span class="o">+</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">;</span>

    <span class="k">__shared__</span> <span class="kt">float</span> <span class="n">As</span><span class="p">[</span><span class="n">TILE</span><span class="p">][</span><span class="n">TILE</span><span class="p">];</span>
    <span class="k">__shared__</span> <span class="kt">float</span> <span class="n">Bs</span><span class="p">[</span><span class="n">TILE</span><span class="p">][</span><span class="n">TILE</span><span class="p">];</span>

    <span class="kt">float</span> <span class="n">acc</span> <span class="o">=</span> <span class="mf">0.0f</span><span class="p">;</span>

    <span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">k0</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="n">k0</span> <span class="o">&lt;</span> <span class="n">K</span><span class="p">;</span> <span class="n">k0</span> <span class="o">+=</span> <span class="n">TILE</span><span class="p">)</span> <span class="p">{</span>
        <span class="kt">float</span> <span class="n">a</span> <span class="o">=</span> <span class="mf">0.0f</span><span class="p">;</span>
        <span class="kt">float</span> <span class="n">b</span> <span class="o">=</span> <span class="mf">0.0f</span><span class="p">;</span>

        <span class="k">if</span> <span class="p">(</span><span class="n">row</span> <span class="o">&lt;</span> <span class="n">M</span> <span class="o">&amp;&amp;</span> <span class="p">(</span><span class="n">k0</span> <span class="o">+</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">)</span> <span class="o">&lt;</span> <span class="n">K</span><span class="p">)</span>
            <span class="n">a</span> <span class="o">=</span> <span class="n">A</span><span class="p">[</span><span class="n">row</span> <span class="o">*</span> <span class="n">K</span> <span class="o">+</span> <span class="p">(</span><span class="n">k0</span> <span class="o">+</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">)];</span>
        <span class="k">if</span> <span class="p">(</span><span class="n">col</span> <span class="o">&lt;</span> <span class="n">N</span> <span class="o">&amp;&amp;</span> <span class="p">(</span><span class="n">k0</span> <span class="o">+</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">y</span><span class="p">)</span> <span class="o">&lt;</span> <span class="n">K</span><span class="p">)</span>
            <span class="n">b</span> <span class="o">=</span> <span class="n">B</span><span class="p">[(</span><span class="n">k0</span> <span class="o">+</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">y</span><span class="p">)</span> <span class="o">*</span> <span class="n">N</span> <span class="o">+</span> <span class="n">col</span><span class="p">];</span>

        <span class="n">As</span><span class="p">[</span><span class="n">threadIdx</span><span class="p">.</span><span class="n">y</span><span class="p">][</span><span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">]</span> <span class="o">=</span> <span class="n">a</span><span class="p">;</span>
        <span class="n">Bs</span><span class="p">[</span><span class="n">threadIdx</span><span class="p">.</span><span class="n">y</span><span class="p">][</span><span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">]</span> <span class="o">=</span> <span class="n">b</span><span class="p">;</span>
        <span class="n">__syncthreads</span><span class="p">();</span>

        <span class="cp">#pragma unroll
</span>        <span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">k</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="n">k</span> <span class="o">&lt;</span> <span class="n">TILE</span><span class="p">;</span> <span class="o">++</span><span class="n">k</span><span class="p">)</span>
            <span class="n">acc</span> <span class="o">+=</span> <span class="n">As</span><span class="p">[</span><span class="n">threadIdx</span><span class="p">.</span><span class="n">y</span><span class="p">][</span><span class="n">k</span><span class="p">]</span> <span class="o">*</span> <span class="n">Bs</span><span class="p">[</span><span class="n">k</span><span class="p">][</span><span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">];</span>

        <span class="n">__syncthreads</span><span class="p">();</span>
    <span class="p">}</span>

    <span class="k">if</span> <span class="p">(</span><span class="n">row</span> <span class="o">&lt;</span> <span class="n">M</span> <span class="o">&amp;&amp;</span> <span class="n">col</span> <span class="o">&lt;</span> <span class="n">N</span><span class="p">)</span> <span class="p">{</span>
        <span class="n">acc</span> <span class="o">+=</span> <span class="p">(</span><span class="n">bias</span> <span class="o">?</span> <span class="n">bias</span><span class="p">[</span><span class="n">col</span><span class="p">]</span> <span class="o">:</span> <span class="mf">0.0f</span><span class="p">);</span>
        <span class="n">C</span><span class="p">[</span><span class="n">row</span> <span class="o">*</span> <span class="n">N</span> <span class="o">+</span> <span class="n">col</span><span class="p">]</span> <span class="o">=</span> <span class="n">acc</span><span class="p">;</span>
    <span class="p">}</span>
<span class="p">}</span>

<span class="c1">// ---------------------------</span>
<span class="c1">// Add residual: y = a + b</span>
<span class="c1">// ---------------------------</span>
<span class="k">__global__</span> <span class="kt">void</span> <span class="nf">add_kernel</span><span class="p">(</span><span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="k">__restrict__</span> <span class="n">a</span><span class="p">,</span>
                           <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="k">__restrict__</span> <span class="n">b</span><span class="p">,</span>
                           <span class="kt">float</span><span class="o">*</span> <span class="k">__restrict__</span> <span class="n">y</span><span class="p">,</span>
                           <span class="kt">int</span> <span class="n">n</span><span class="p">)</span> <span class="p">{</span>
    <span class="kt">int</span> <span class="n">i</span> <span class="o">=</span> <span class="n">blockIdx</span><span class="p">.</span><span class="n">x</span> <span class="o">*</span> <span class="n">blockDim</span><span class="p">.</span><span class="n">x</span> <span class="o">+</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">;</span>
    <span class="k">if</span> <span class="p">(</span><span class="n">i</span> <span class="o">&lt;</span> <span class="n">n</span><span class="p">)</span> <span class="n">y</span><span class="p">[</span><span class="n">i</span><span class="p">]</span> <span class="o">=</span> <span class="n">a</span><span class="p">[</span><span class="n">i</span><span class="p">]</span> <span class="o">+</span> <span class="n">b</span><span class="p">[</span><span class="n">i</span><span class="p">];</span>
<span class="p">}</span>

<span class="c1">// ---------------------------</span>
<span class="c1">// GELU elementwise</span>
<span class="c1">// ---------------------------</span>
<span class="k">__global__</span> <span class="kt">void</span> <span class="nf">gelu_kernel</span><span class="p">(</span><span class="kt">float</span><span class="o">*</span> <span class="k">__restrict__</span> <span class="n">x</span><span class="p">,</span> <span class="kt">int</span> <span class="n">n</span><span class="p">)</span> <span class="p">{</span>
    <span class="kt">int</span> <span class="n">i</span> <span class="o">=</span> <span class="n">blockIdx</span><span class="p">.</span><span class="n">x</span> <span class="o">*</span> <span class="n">blockDim</span><span class="p">.</span><span class="n">x</span> <span class="o">+</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">;</span>
    <span class="k">if</span> <span class="p">(</span><span class="n">i</span> <span class="o">&lt;</span> <span class="n">n</span><span class="p">)</span> <span class="n">x</span><span class="p">[</span><span class="n">i</span><span class="p">]</span> <span class="o">=</span> <span class="n">gelu_tanh</span><span class="p">(</span><span class="n">x</span><span class="p">[</span><span class="n">i</span><span class="p">]);</span>
<span class="p">}</span>

<span class="c1">// ---------------------------</span>
<span class="c1">// Attention kernel:</span>
<span class="c1">// One block per (token t, head h)</span>
<span class="c1">// Input: qkv [seq, 2304] = [Q|K|V] each 768</span>
<span class="c1">// Output: attn_out [seq, 768] (concatenated heads)</span>
<span class="c1">// No causal mask</span>
<span class="c1">//</span>
<span class="c1">// Shared memory: scores[seq_len] floats</span>
<span class="c1">// blockDim.x recommended 256</span>
<span class="c1">// ---------------------------</span>
<span class="k">__global__</span> <span class="kt">void</span> <span class="nf">mha_nocausal_kernel</span><span class="p">(</span><span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="k">__restrict__</span> <span class="n">qkv</span><span class="p">,</span>
                                   <span class="kt">float</span><span class="o">*</span> <span class="k">__restrict__</span> <span class="n">attn_out</span><span class="p">,</span>
                                   <span class="kt">int</span> <span class="n">seq_len</span><span class="p">)</span> <span class="p">{</span>
    <span class="kt">int</span> <span class="n">t</span> <span class="o">=</span> <span class="n">blockIdx</span><span class="p">.</span><span class="n">x</span><span class="p">;</span>      <span class="c1">// query token</span>
    <span class="kt">int</span> <span class="n">h</span> <span class="o">=</span> <span class="n">blockIdx</span><span class="p">.</span><span class="n">y</span><span class="p">;</span>      <span class="c1">// head 0..11</span>
    <span class="k">if</span> <span class="p">(</span><span class="n">t</span> <span class="o">&gt;=</span> <span class="n">seq_len</span> <span class="o">||</span> <span class="n">h</span> <span class="o">&gt;=</span> <span class="mi">12</span><span class="p">)</span> <span class="k">return</span><span class="p">;</span>

    <span class="k">extern</span> <span class="k">__shared__</span> <span class="kt">float</span> <span class="n">shmem</span><span class="p">[];</span> <span class="c1">// size &gt;= seq_len floats</span>
    <span class="kt">float</span><span class="o">*</span> <span class="n">scores</span> <span class="o">=</span> <span class="n">shmem</span><span class="p">;</span>

    <span class="k">const</span> <span class="kt">int</span> <span class="n">D</span> <span class="o">=</span> <span class="mi">768</span><span class="p">;</span>
    <span class="k">const</span> <span class="kt">int</span> <span class="n">Dh</span> <span class="o">=</span> <span class="mi">64</span><span class="p">;</span>
    <span class="k">const</span> <span class="kt">float</span> <span class="n">inv_sqrt_dh</span> <span class="o">=</span> <span class="mf">1.0f</span> <span class="o">/</span> <span class="n">sqrtf</span><span class="p">((</span><span class="kt">float</span><span class="p">)</span><span class="n">Dh</span><span class="p">);</span>

    <span class="c1">// Q pointer for this (t,h)</span>
    <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="n">Q</span> <span class="o">=</span> <span class="n">qkv</span> <span class="o">+</span> <span class="n">t</span> <span class="o">*</span> <span class="p">(</span><span class="mi">3</span> <span class="o">*</span> <span class="n">D</span><span class="p">)</span> <span class="o">+</span> <span class="mi">0</span> <span class="o">*</span> <span class="n">D</span> <span class="o">+</span> <span class="n">h</span> <span class="o">*</span> <span class="n">Dh</span><span class="p">;</span>

    <span class="c1">// Pass 1: compute scores[s] = dot(Q, K_s) / sqrt(Dh)</span>
    <span class="kt">float</span> <span class="n">local_max</span> <span class="o">=</span> <span class="o">-</span><span class="mf">1e30f</span><span class="p">;</span>
    <span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">s</span> <span class="o">=</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">;</span> <span class="n">s</span> <span class="o">&lt;</span> <span class="n">seq_len</span><span class="p">;</span> <span class="n">s</span> <span class="o">+=</span> <span class="n">blockDim</span><span class="p">.</span><span class="n">x</span><span class="p">)</span> <span class="p">{</span>
        <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="n">K</span> <span class="o">=</span> <span class="n">qkv</span> <span class="o">+</span> <span class="n">s</span> <span class="o">*</span> <span class="p">(</span><span class="mi">3</span> <span class="o">*</span> <span class="n">D</span><span class="p">)</span> <span class="o">+</span> <span class="mi">1</span> <span class="o">*</span> <span class="n">D</span> <span class="o">+</span> <span class="n">h</span> <span class="o">*</span> <span class="n">Dh</span><span class="p">;</span>
        <span class="kt">float</span> <span class="n">dot</span> <span class="o">=</span> <span class="mf">0.0f</span><span class="p">;</span>
        <span class="cp">#pragma unroll
</span>        <span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">i</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="n">i</span> <span class="o">&lt;</span> <span class="n">Dh</span><span class="p">;</span> <span class="o">++</span><span class="n">i</span><span class="p">)</span> <span class="n">dot</span> <span class="o">+=</span> <span class="n">Q</span><span class="p">[</span><span class="n">i</span><span class="p">]</span> <span class="o">*</span> <span class="n">K</span><span class="p">[</span><span class="n">i</span><span class="p">];</span>
        <span class="kt">float</span> <span class="n">sc</span> <span class="o">=</span> <span class="n">dot</span> <span class="o">*</span> <span class="n">inv_sqrt_dh</span><span class="p">;</span>
        <span class="n">scores</span><span class="p">[</span><span class="n">s</span><span class="p">]</span> <span class="o">=</span> <span class="n">sc</span><span class="p">;</span>
        <span class="n">local_max</span> <span class="o">=</span> <span class="n">fmaxf</span><span class="p">(</span><span class="n">local_max</span><span class="p">,</span> <span class="n">sc</span><span class="p">);</span>
    <span class="p">}</span>
    <span class="kt">float</span> <span class="n">max_sc</span> <span class="o">=</span> <span class="n">block_reduce_max</span><span class="p">(</span><span class="n">local_max</span><span class="p">);</span>
    <span class="n">__syncthreads</span><span class="p">();</span>

    <span class="c1">// Pass 2: exp(scores - max), sum</span>
    <span class="kt">float</span> <span class="n">local_sum</span> <span class="o">=</span> <span class="mf">0.0f</span><span class="p">;</span>
    <span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">s</span> <span class="o">=</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">;</span> <span class="n">s</span> <span class="o">&lt;</span> <span class="n">seq_len</span><span class="p">;</span> <span class="n">s</span> <span class="o">+=</span> <span class="n">blockDim</span><span class="p">.</span><span class="n">x</span><span class="p">)</span> <span class="p">{</span>
        <span class="kt">float</span> <span class="n">e</span> <span class="o">=</span> <span class="n">expf</span><span class="p">(</span><span class="n">scores</span><span class="p">[</span><span class="n">s</span><span class="p">]</span> <span class="o">-</span> <span class="n">max_sc</span><span class="p">);</span>
        <span class="n">scores</span><span class="p">[</span><span class="n">s</span><span class="p">]</span> <span class="o">=</span> <span class="n">e</span><span class="p">;</span>
        <span class="n">local_sum</span> <span class="o">+=</span> <span class="n">e</span><span class="p">;</span>
    <span class="p">}</span>
    <span class="kt">float</span> <span class="n">sum_sc</span> <span class="o">=</span> <span class="n">block_reduce_sum</span><span class="p">(</span><span class="n">local_sum</span><span class="p">);</span>
    <span class="kt">float</span> <span class="n">inv_sum</span> <span class="o">=</span> <span class="mf">1.0f</span> <span class="o">/</span> <span class="n">sum_sc</span><span class="p">;</span>
    <span class="n">__syncthreads</span><span class="p">();</span>

    <span class="c1">// Pass 3: weighted sum of V</span>
    <span class="k">if</span> <span class="p">(</span><span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span> <span class="o">&lt;</span> <span class="n">Dh</span><span class="p">)</span> <span class="p">{</span>
        <span class="kt">int</span> <span class="n">i</span> <span class="o">=</span> <span class="n">threadIdx</span><span class="p">.</span><span class="n">x</span><span class="p">;</span>
        <span class="kt">float</span> <span class="n">acc</span> <span class="o">=</span> <span class="mf">0.0f</span><span class="p">;</span>
        <span class="k">for</span> <span class="p">(</span><span class="kt">int</span> <span class="n">s</span> <span class="o">=</span> <span class="mi">0</span><span class="p">;</span> <span class="n">s</span> <span class="o">&lt;</span> <span class="n">seq_len</span><span class="p">;</span> <span class="o">++</span><span class="n">s</span><span class="p">)</span> <span class="p">{</span>
            <span class="kt">float</span> <span class="n">p</span> <span class="o">=</span> <span class="n">scores</span><span class="p">[</span><span class="n">s</span><span class="p">]</span> <span class="o">*</span> <span class="n">inv_sum</span><span class="p">;</span>
            <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="n">V</span> <span class="o">=</span> <span class="n">qkv</span> <span class="o">+</span> <span class="n">s</span> <span class="o">*</span> <span class="p">(</span><span class="mi">3</span> <span class="o">*</span> <span class="n">D</span><span class="p">)</span> <span class="o">+</span> <span class="mi">2</span> <span class="o">*</span> <span class="n">D</span> <span class="o">+</span> <span class="n">h</span> <span class="o">*</span> <span class="n">Dh</span><span class="p">;</span>
            <span class="n">acc</span> <span class="o">+=</span> <span class="n">p</span> <span class="o">*</span> <span class="n">V</span><span class="p">[</span><span class="n">i</span><span class="p">];</span>
        <span class="p">}</span>
        <span class="n">attn_out</span><span class="p">[</span><span class="n">t</span> <span class="o">*</span> <span class="n">D</span> <span class="o">+</span> <span class="n">h</span> <span class="o">*</span> <span class="n">Dh</span> <span class="o">+</span> <span class="n">i</span><span class="p">]</span> <span class="o">=</span> <span class="n">acc</span><span class="p">;</span>
    <span class="p">}</span>
<span class="p">}</span>

<span class="c1">// ---------------------------</span>
<span class="c1">// Entry point (device pointers)</span>
<span class="c1">// x:      (seq_len, 768)</span>
<span class="c1">// output: (seq_len, 768)</span>
<span class="c1">// weights packed per your offsets</span>
<span class="c1">// ---------------------------</span>
<span class="k">extern</span> <span class="s">"C"</span> <span class="kt">void</span> <span class="nf">solve</span><span class="p">(</span><span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="n">x</span><span class="p">,</span> <span class="kt">float</span><span class="o">*</span> <span class="n">output</span><span class="p">,</span> <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="n">weights</span><span class="p">,</span> <span class="kt">int</span> <span class="n">seq_len</span><span class="p">)</span> <span class="p">{</span>
    <span class="k">const</span> <span class="kt">int</span> <span class="n">D</span> <span class="o">=</span> <span class="mi">768</span><span class="p">;</span>
    <span class="k">const</span> <span class="kt">int</span> <span class="n">FF</span> <span class="o">=</span> <span class="mi">3072</span><span class="p">;</span>
    <span class="k">const</span> <span class="kt">int</span> <span class="n">H</span> <span class="o">=</span> <span class="mi">12</span><span class="p">;</span>

    <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="n">gamma1</span> <span class="o">=</span> <span class="n">weights</span> <span class="o">+</span> <span class="mi">0</span><span class="p">;</span>
    <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="n">beta1</span>  <span class="o">=</span> <span class="n">weights</span> <span class="o">+</span> <span class="mi">768</span><span class="p">;</span>

    <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="n">W_qkv</span>  <span class="o">=</span> <span class="n">weights</span> <span class="o">+</span> <span class="mi">1536</span><span class="p">;</span>
    <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="n">b_qkv</span>  <span class="o">=</span> <span class="n">weights</span> <span class="o">+</span> <span class="mi">1771008</span><span class="p">;</span>

    <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="n">W_attn</span> <span class="o">=</span> <span class="n">weights</span> <span class="o">+</span> <span class="mi">1773312</span><span class="p">;</span>
    <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="n">b_attn</span> <span class="o">=</span> <span class="n">weights</span> <span class="o">+</span> <span class="mi">2363136</span><span class="p">;</span>

    <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="n">gamma2</span> <span class="o">=</span> <span class="n">weights</span> <span class="o">+</span> <span class="mi">2363904</span><span class="p">;</span>
    <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="n">beta2</span>  <span class="o">=</span> <span class="n">weights</span> <span class="o">+</span> <span class="mi">2364672</span><span class="p">;</span>

    <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="n">W_fc</span>   <span class="o">=</span> <span class="n">weights</span> <span class="o">+</span> <span class="mi">2365440</span><span class="p">;</span>
    <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="n">b_fc</span>   <span class="o">=</span> <span class="n">weights</span> <span class="o">+</span> <span class="mi">4724736</span><span class="p">;</span>

    <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="n">W_proj</span> <span class="o">=</span> <span class="n">weights</span> <span class="o">+</span> <span class="mi">4727808</span><span class="p">;</span>
    <span class="k">const</span> <span class="kt">float</span><span class="o">*</span> <span class="n">b_proj</span> <span class="o">=</span> <span class="n">weights</span> <span class="o">+</span> <span class="mi">7087104</span><span class="p">;</span>

    <span class="kt">float</span> <span class="o">*</span><span class="n">ln1</span> <span class="o">=</span> <span class="nb">nullptr</span><span class="p">,</span> <span class="o">*</span><span class="n">qkv</span> <span class="o">=</span> <span class="nb">nullptr</span><span class="p">,</span> <span class="o">*</span><span class="n">attn</span> <span class="o">=</span> <span class="nb">nullptr</span><span class="p">,</span> <span class="o">*</span><span class="n">proj</span> <span class="o">=</span> <span class="nb">nullptr</span><span class="p">,</span> <span class="o">*</span><span class="n">x1</span> <span class="o">=</span> <span class="nb">nullptr</span><span class="p">;</span>
    <span class="kt">float</span> <span class="o">*</span><span class="n">ln2</span> <span class="o">=</span> <span class="nb">nullptr</span><span class="p">,</span> <span class="o">*</span><span class="n">ff1</span> <span class="o">=</span> <span class="nb">nullptr</span><span class="p">,</span> <span class="o">*</span><span class="n">ff2</span> <span class="o">=</span> <span class="nb">nullptr</span><span class="p">;</span>

    <span class="kt">size_t</span> <span class="n">bytes_ln</span>  <span class="o">=</span> <span class="p">(</span><span class="kt">size_t</span><span class="p">)</span><span class="n">seq_len</span> <span class="o">*</span> <span class="n">D</span>  <span class="o">*</span> <span class="k">sizeof</span><span class="p">(</span><span class="kt">float</span><span class="p">);</span>
    <span class="kt">size_t</span> <span class="n">bytes_qkv</span> <span class="o">=</span> <span class="p">(</span><span class="kt">size_t</span><span class="p">)</span><span class="n">seq_len</span> <span class="o">*</span> <span class="p">(</span><span class="mi">3</span> <span class="o">*</span> <span class="n">D</span><span class="p">)</span> <span class="o">*</span> <span class="k">sizeof</span><span class="p">(</span><span class="kt">float</span><span class="p">);</span>
    <span class="kt">size_t</span> <span class="n">bytes_ff1</span> <span class="o">=</span> <span class="p">(</span><span class="kt">size_t</span><span class="p">)</span><span class="n">seq_len</span> <span class="o">*</span> <span class="n">FF</span> <span class="o">*</span> <span class="k">sizeof</span><span class="p">(</span><span class="kt">float</span><span class="p">);</span>

    <span class="n">CHECK_CUDA</span><span class="p">(</span><span class="n">cudaMalloc</span><span class="p">(</span><span class="o">&amp;</span><span class="n">ln1</span><span class="p">,</span>  <span class="n">bytes_ln</span><span class="p">));</span>
    <span class="n">CHECK_CUDA</span><span class="p">(</span><span class="n">cudaMalloc</span><span class="p">(</span><span class="o">&amp;</span><span class="n">qkv</span><span class="p">,</span>  <span class="n">bytes_qkv</span><span class="p">));</span>
    <span class="n">CHECK_CUDA</span><span class="p">(</span><span class="n">cudaMalloc</span><span class="p">(</span><span class="o">&amp;</span><span class="n">attn</span><span class="p">,</span> <span class="n">bytes_ln</span><span class="p">));</span>
    <span class="n">CHECK_CUDA</span><span class="p">(</span><span class="n">cudaMalloc</span><span class="p">(</span><span class="o">&amp;</span><span class="n">proj</span><span class="p">,</span> <span class="n">bytes_ln</span><span class="p">));</span>
    <span class="n">CHECK_CUDA</span><span class="p">(</span><span class="n">cudaMalloc</span><span class="p">(</span><span class="o">&amp;</span><span class="n">x1</span><span class="p">,</span>   <span class="n">bytes_ln</span><span class="p">));</span>
    <span class="n">CHECK_CUDA</span><span class="p">(</span><span class="n">cudaMalloc</span><span class="p">(</span><span class="o">&amp;</span><span class="n">ln2</span><span class="p">,</span>  <span class="n">bytes_ln</span><span class="p">));</span>
    <span class="n">CHECK_CUDA</span><span class="p">(</span><span class="n">cudaMalloc</span><span class="p">(</span><span class="o">&amp;</span><span class="n">ff1</span><span class="p">,</span>  <span class="n">bytes_ff1</span><span class="p">));</span>
    <span class="n">CHECK_CUDA</span><span class="p">(</span><span class="n">cudaMalloc</span><span class="p">(</span><span class="o">&amp;</span><span class="n">ff2</span><span class="p">,</span>  <span class="n">bytes_ln</span><span class="p">));</span>

    <span class="c1">// 1) LN1</span>
    <span class="n">layernorm768_kernel</span><span class="o">&lt;&lt;&lt;</span><span class="n">seq_len</span><span class="p">,</span> <span class="mi">256</span><span class="o">&gt;&gt;&gt;</span><span class="p">(</span><span class="n">x</span><span class="p">,</span> <span class="n">ln1</span><span class="p">,</span> <span class="n">gamma1</span><span class="p">,</span> <span class="n">beta1</span><span class="p">,</span> <span class="n">seq_len</span><span class="p">);</span>

    <span class="c1">// 2) QKV = ln1 * W_qkv + b_qkv   (M=seq_len, K=768, N=2304)</span>
    <span class="p">{</span>
        <span class="k">const</span> <span class="kt">int</span> <span class="n">TILE</span> <span class="o">=</span> <span class="mi">16</span><span class="p">;</span>
        <span class="kt">dim3</span> <span class="n">block</span><span class="p">(</span><span class="n">TILE</span><span class="p">,</span> <span class="n">TILE</span><span class="p">);</span>
        <span class="kt">dim3</span> <span class="n">grid</span><span class="p">((</span><span class="mi">2304</span> <span class="o">+</span> <span class="n">TILE</span><span class="o">-</span><span class="mi">1</span><span class="p">)</span><span class="o">/</span><span class="n">TILE</span><span class="p">,</span> <span class="p">(</span><span class="n">seq_len</span> <span class="o">+</span> <span class="n">TILE</span><span class="o">-</span><span class="mi">1</span><span class="p">)</span><span class="o">/</span><span class="n">TILE</span><span class="p">);</span>
        <span class="n">matmul_bias_tiled</span><span class="o">&lt;</span><span class="n">TILE</span><span class="o">&gt;&lt;&lt;&lt;</span><span class="n">grid</span><span class="p">,</span> <span class="n">block</span><span class="o">&gt;&gt;&gt;</span><span class="p">(</span><span class="n">ln1</span><span class="p">,</span> <span class="n">W_qkv</span><span class="p">,</span> <span class="n">b_qkv</span><span class="p">,</span> <span class="n">qkv</span><span class="p">,</span> <span class="n">seq_len</span><span class="p">,</span> <span class="mi">768</span><span class="p">,</span> <span class="mi">2304</span><span class="p">);</span>
    <span class="p">}</span>

    <span class="c1">// 3) MHA: attn = (seq_len, 768)</span>
    <span class="p">{</span>
        <span class="kt">dim3</span> <span class="n">grid</span><span class="p">(</span><span class="n">seq_len</span><span class="p">,</span> <span class="n">H</span><span class="p">);</span>
        <span class="kt">size_t</span> <span class="n">shmem</span> <span class="o">=</span> <span class="p">(</span><span class="kt">size_t</span><span class="p">)</span><span class="n">seq_len</span> <span class="o">*</span> <span class="k">sizeof</span><span class="p">(</span><span class="kt">float</span><span class="p">);</span>
        <span class="n">mha_nocausal_kernel</span><span class="o">&lt;&lt;&lt;</span><span class="n">grid</span><span class="p">,</span> <span class="mi">256</span><span class="p">,</span> <span class="n">shmem</span><span class="o">&gt;&gt;&gt;</span><span class="p">(</span><span class="n">qkv</span><span class="p">,</span> <span class="n">attn</span><span class="p">,</span> <span class="n">seq_len</span><span class="p">);</span>
    <span class="p">}</span>

    <span class="c1">// 4) proj = attn * W_attn + b_attn</span>
    <span class="p">{</span>
        <span class="k">const</span> <span class="kt">int</span> <span class="n">TILE</span> <span class="o">=</span> <span class="mi">16</span><span class="p">;</span>
        <span class="kt">dim3</span> <span class="n">block</span><span class="p">(</span><span class="n">TILE</span><span class="p">,</span> <span class="n">TILE</span><span class="p">);</span>
        <span class="kt">dim3</span> <span class="n">grid</span><span class="p">((</span><span class="mi">768</span> <span class="o">+</span> <span class="n">TILE</span><span class="o">-</span><span class="mi">1</span><span class="p">)</span><span class="o">/</span><span class="n">TILE</span><span class="p">,</span> <span class="p">(</span><span class="n">seq_len</span> <span class="o">+</span> <span class="n">TILE</span><span class="o">-</span><span class="mi">1</span><span class="p">)</span><span class="o">/</span><span class="n">TILE</span><span class="p">);</span>
        <span class="n">matmul_bias_tiled</span><span class="o">&lt;</span><span class="n">TILE</span><span class="o">&gt;&lt;&lt;&lt;</span><span class="n">grid</span><span class="p">,</span> <span class="n">block</span><span class="o">&gt;&gt;&gt;</span><span class="p">(</span><span class="n">attn</span><span class="p">,</span> <span class="n">W_attn</span><span class="p">,</span> <span class="n">b_attn</span><span class="p">,</span> <span class="n">proj</span><span class="p">,</span> <span class="n">seq_len</span><span class="p">,</span> <span class="mi">768</span><span class="p">,</span> <span class="mi">768</span><span class="p">);</span>
    <span class="p">}</span>

    <span class="c1">// 5) x1 = x + proj</span>
    <span class="p">{</span>
        <span class="kt">int</span> <span class="n">n</span> <span class="o">=</span> <span class="n">seq_len</span> <span class="o">*</span> <span class="n">D</span><span class="p">;</span>
        <span class="n">add_kernel</span><span class="o">&lt;&lt;&lt;</span><span class="p">(</span><span class="n">n</span><span class="o">+</span><span class="mi">255</span><span class="p">)</span><span class="o">/</span><span class="mi">256</span><span class="p">,</span> <span class="mi">256</span><span class="o">&gt;&gt;&gt;</span><span class="p">(</span><span class="n">x</span><span class="p">,</span> <span class="n">proj</span><span class="p">,</span> <span class="n">x1</span><span class="p">,</span> <span class="n">n</span><span class="p">);</span>
    <span class="p">}</span>

    <span class="c1">// 6) LN2</span>
    <span class="n">layernorm768_kernel</span><span class="o">&lt;&lt;&lt;</span><span class="n">seq_len</span><span class="p">,</span> <span class="mi">256</span><span class="o">&gt;&gt;&gt;</span><span class="p">(</span><span class="n">x1</span><span class="p">,</span> <span class="n">ln2</span><span class="p">,</span> <span class="n">gamma2</span><span class="p">,</span> <span class="n">beta2</span><span class="p">,</span> <span class="n">seq_len</span><span class="p">);</span>

    <span class="c1">// 7) ff1 = ln2 * W_fc + b_fc   (M=seq_len, K=768, N=3072)</span>
    <span class="p">{</span>
        <span class="k">const</span> <span class="kt">int</span> <span class="n">TILE</span> <span class="o">=</span> <span class="mi">16</span><span class="p">;</span>
        <span class="kt">dim3</span> <span class="n">block</span><span class="p">(</span><span class="n">TILE</span><span class="p">,</span> <span class="n">TILE</span><span class="p">);</span>
        <span class="kt">dim3</span> <span class="n">grid</span><span class="p">((</span><span class="mi">3072</span> <span class="o">+</span> <span class="n">TILE</span><span class="o">-</span><span class="mi">1</span><span class="p">)</span><span class="o">/</span><span class="n">TILE</span><span class="p">,</span> <span class="p">(</span><span class="n">seq_len</span> <span class="o">+</span> <span class="n">TILE</span><span class="o">-</span><span class="mi">1</span><span class="p">)</span><span class="o">/</span><span class="n">TILE</span><span class="p">);</span>
        <span class="n">matmul_bias_tiled</span><span class="o">&lt;</span><span class="n">TILE</span><span class="o">&gt;&lt;&lt;&lt;</span><span class="n">grid</span><span class="p">,</span> <span class="n">block</span><span class="o">&gt;&gt;&gt;</span><span class="p">(</span><span class="n">ln2</span><span class="p">,</span> <span class="n">W_fc</span><span class="p">,</span> <span class="n">b_fc</span><span class="p">,</span> <span class="n">ff1</span><span class="p">,</span> <span class="n">seq_len</span><span class="p">,</span> <span class="mi">768</span><span class="p">,</span> <span class="mi">3072</span><span class="p">);</span>
    <span class="p">}</span>

    <span class="c1">// 8) GELU(ff1)</span>
    <span class="p">{</span>
        <span class="kt">int</span> <span class="n">n</span> <span class="o">=</span> <span class="n">seq_len</span> <span class="o">*</span> <span class="n">FF</span><span class="p">;</span>
        <span class="n">gelu_kernel</span><span class="o">&lt;&lt;&lt;</span><span class="p">(</span><span class="n">n</span><span class="o">+</span><span class="mi">255</span><span class="p">)</span><span class="o">/</span><span class="mi">256</span><span class="p">,</span> <span class="mi">256</span><span class="o">&gt;&gt;&gt;</span><span class="p">(</span><span class="n">ff1</span><span class="p">,</span> <span class="n">n</span><span class="p">);</span>
    <span class="p">}</span>

    <span class="c1">// 9) ff2 = ff1 * W_proj + b_proj  (M=seq_len, K=3072, N=768)</span>
    <span class="p">{</span>
        <span class="k">const</span> <span class="kt">int</span> <span class="n">TILE</span> <span class="o">=</span> <span class="mi">16</span><span class="p">;</span>
        <span class="kt">dim3</span> <span class="n">block</span><span class="p">(</span><span class="n">TILE</span><span class="p">,</span> <span class="n">TILE</span><span class="p">);</span>
        <span class="kt">dim3</span> <span class="n">grid</span><span class="p">((</span><span class="mi">768</span> <span class="o">+</span> <span class="n">TILE</span><span class="o">-</span><span class="mi">1</span><span class="p">)</span><span class="o">/</span><span class="n">TILE</span><span class="p">,</span> <span class="p">(</span><span class="n">seq_len</span> <span class="o">+</span> <span class="n">TILE</span><span class="o">-</span><span class="mi">1</span><span class="p">)</span><span class="o">/</span><span class="n">TILE</span><span class="p">);</span>
        <span class="n">matmul_bias_tiled</span><span class="o">&lt;</span><span class="n">TILE</span><span class="o">&gt;&lt;&lt;&lt;</span><span class="n">grid</span><span class="p">,</span> <span class="n">block</span><span class="o">&gt;&gt;&gt;</span><span class="p">(</span><span class="n">ff1</span><span class="p">,</span> <span class="n">W_proj</span><span class="p">,</span> <span class="n">b_proj</span><span class="p">,</span> <span class="n">ff2</span><span class="p">,</span> <span class="n">seq_len</span><span class="p">,</span> <span class="mi">3072</span><span class="p">,</span> <span class="mi">768</span><span class="p">);</span>
    <span class="p">}</span>

    <span class="c1">// 10) output = x1 + ff2</span>
    <span class="p">{</span>
        <span class="kt">int</span> <span class="n">n</span> <span class="o">=</span> <span class="n">seq_len</span> <span class="o">*</span> <span class="n">D</span><span class="p">;</span>
        <span class="n">add_kernel</span><span class="o">&lt;&lt;&lt;</span><span class="p">(</span><span class="n">n</span><span class="o">+</span><span class="mi">255</span><span class="p">)</span><span class="o">/</span><span class="mi">256</span><span class="p">,</span> <span class="mi">256</span><span class="o">&gt;&gt;&gt;</span><span class="p">(</span><span class="n">x1</span><span class="p">,</span> <span class="n">ff2</span><span class="p">,</span> <span class="n">output</span><span class="p">,</span> <span class="n">n</span><span class="p">);</span>
    <span class="p">}</span>

    <span class="n">cudaFree</span><span class="p">(</span><span class="n">ln1</span><span class="p">);</span> <span class="n">cudaFree</span><span class="p">(</span><span class="n">qkv</span><span class="p">);</span>  <span class="n">cudaFree</span><span class="p">(</span><span class="n">attn</span><span class="p">);</span> <span class="n">cudaFree</span><span class="p">(</span><span class="n">proj</span><span class="p">);</span>
    <span class="n">cudaFree</span><span class="p">(</span><span class="n">x1</span><span class="p">);</span>  <span class="n">cudaFree</span><span class="p">(</span><span class="n">ln2</span><span class="p">);</span>  <span class="n">cudaFree</span><span class="p">(</span><span class="n">ff1</span><span class="p">);</span>  <span class="n">cudaFree</span><span class="p">(</span><span class="n">ff2</span><span class="p">);</span>
<span class="p">}</span>
</code></pre></div></div>]]></content><author><name>Maciej Grochowski</name></author><category term="ml" /><summary type="html"><![CDATA[December, somewhere over the Pacific on the way to Tokyo. The letGPU challenge open in one browser tab, Ro Salaverry’s The Scaling Era: An Oral History of AI, 2019–2025 on my Kindle. No plan beyond getting a feel for the challenge — I poked at it for a while, read a few chapters, and fell asleep somewhere east of the dateline. Woke up on approach to Narita.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://page-fault.io/assets/social-preview.png" /><media:content medium="image" url="https://page-fault.io/assets/social-preview.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">The GPT-2 Decoder Block: Before Writing Any Code</title><link href="https://page-fault.io/ml/2025/11/15/letgpu-gpt2-transformer-block.html" rel="alternate" type="text/html" title="The GPT-2 Decoder Block: Before Writing Any Code" /><published>2025-11-15T12:00:00+00:00</published><updated>2025-11-15T12:00:00+00:00</updated><id>https://page-fault.io/ml/2025/11/15/letgpu-gpt2-transformer-block</id><content type="html" xml:base="https://page-fault.io/ml/2025/11/15/letgpu-gpt2-transformer-block.html"><![CDATA[<p>At university, long before the current AI period, neural networks were still niche. Academic
appeal, limited hardware, results that rarely justified the compute. We wrote simple recurrent
networks in C, used them in robotics and image-processing experiments, and spent most of the
time fighting the hardware constraints rather than thinking about the model. I liked that part
more than I probably should have admitted.</p>

<p>Then around 2017 and 2018 transformers arrived, and the field reorganized itself around them.
Not a replacement for neural networks — a different way to structure them, with a mechanism
for sequence modeling that turned out to scale in a way nothing before it had.</p>

<p>When LetGPU published a challenge to implement a GPT-2 Small decoder block, the old interest
surfaced immediately. Before writing a single line of code I wanted to understand what the
block actually is, how data flows through it, and what the challenge is really asking for.
This post is that understanding. The implementation — in CUDA, on the
<a href="/ml/2025/12/28/transformer-block-cuda.html">December trip to Tokyo</a> —
comes separately.</p>

<h2 id="what-a-transformer-block-does">What a Transformer Block Does</h2>

<p>At a high level, a neural network is a system that repeatedly transforms a representation into
a better one. In language models the representation is a vector per token. Each block takes
those vectors, processes them, and passes a refined version forward.</p>

<p>The key difference from older sequence models is attention. RNNs process tokens one at a time
and carry state forward; long-range dependencies get compressed into whatever the hidden state
can hold. Transformers let every token look at every other token directly and decide which ones
matter for its current representation. That is not a minor optimization — it changes what the
model can learn and how fast it can train.</p>

<p><img src="https://res.cloudinary.com/gotocco/image/upload/v1774206004/transformer_block_intro_kljls3.svg" alt="Generic Transformer Block" /></p>

<p><em>Figure 1. A simplified transformer block. Attention lets tokens exchange information; the
feed-forward network refines each representation independently. Residual connections add the
original signal back after each sublayer; LayerNorm keeps the values numerically stable.</em></p>

<p>The block does two jobs in sequence. First, attention: tokens share context with each other.
Second, a feed-forward network: each token’s representation is deepened independently. Residual
connections run around both, so the block learns updates rather than replacements. The output
shape matches the input — same number of tokens, same hidden width.</p>

<h2 id="the-letgpu-challenge">The LetGPU Challenge</h2>

<p>The task is to implement one GPT-2 Small decoder block. You are given an input tensor <code class="language-plaintext highlighter-rouge">x</code> of
shape <code class="language-plaintext highlighter-rouge">(seq_len, 768)</code> and a packed flat buffer containing all parameters for the block. Compute
the output — same shape as the input.</p>

<p>That sounds compact. It is not. One block contains:</p>

<ul>
  <li>two LayerNorms (each with learned scale and bias)</li>
  <li>a combined QKV projection (768 → 2304)</li>
  <li>12-head self-attention</li>
  <li>an output projection (768 → 768)</li>
  <li>two residual connections</li>
  <li>a feed-forward network: 768 → 3072 → GELU → 768</li>
</ul>

<p>About 7.1 million parameters in one block. GPT-2 Small stacks twelve of them.</p>

<p><img src="https://res.cloudinary.com/gotocco/image/upload/v1774206004/gpt2_decoder_block_challenge_ngh87q.svg" alt="GPT-2 Small Decoder Block Challenge Flow" /></p>

<p><em>Figure 2. The GPT-2 Small decoder block from the LetGPU challenge. Pre-norm: LayerNorm runs
before attention and before the feed-forward network. Hidden size 768, twelve heads of 64
dimensions each, feed-forward expansion to 3072.</em></p>

<h2 id="pre-norm-and-the-flow">Pre-Norm and the Flow</h2>

<p>GPT-2 uses <strong>pre-norm</strong>: LayerNorm is applied before each sublayer, not after. The original
<em>Attention Is All You Need</em> paper used post-norm (normalize after the residual add). The
difference matters for training stability in deep models and is worth knowing before reading
the weight layout.</p>

<p>The full sequence through the block:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>x → LN1 → QKV projection → multi-head attention → output projection → add(x, ·) → x1
x1 → LN2 → FC projection → GELU → output projection → add(x1, ·) → output
</code></pre></div></div>

<p>Ten operations, two residual paths, one block.</p>

<h2 id="attention-q-k-v">Attention: Q, K, V</h2>

<p>Each token is projected into three vectors — Query, Key, Value — via a single combined weight
matrix that is then split. The intuition that actually holds up: the Query describes what this
token is looking for, the Key describes what it offers, and the Value is what gets passed along
if the match is strong.</p>

<p>Attention weights come from comparing queries against keys (scaled by <code class="language-plaintext highlighter-rouge">1/√64</code> to keep
the dot-product variance at 1 regardless of dimension — without this, the softmax
inputs grow with <code class="language-plaintext highlighter-rouge">d_k</code> and push gradients into near-zero saturation regions), then
those weights determine how values are mixed. Each token ends
up with a weighted sum of values from the entire sequence — its representation updated by
whatever context the model has learned to attend to.</p>

<p>Multi-head attention runs this in parallel across twelve subspaces of 64 dimensions each.
<code class="language-plaintext highlighter-rouge">12 × 64 = 768</code>. Each head can specialize in a different kind of dependency; the results are
concatenated and projected back to 768.</p>

<h2 id="feed-forward-refine-independently">Feed-Forward: Refine Independently</h2>

<p>After attention, the feed-forward network processes each token on its own — no cross-token
interaction here. It expands from 768 to 3072, applies GELU, and projects back to 768. The
expansion gives the model room to build richer feature combinations before compressing back
down. GELU acts as a smooth gate: it suppresses near-zero activations continuously rather than
with a hard threshold, which is why it works better than ReLU for this use case.</p>

<h2 id="the-packed-weight-buffer">The Packed Weight Buffer</h2>

<p>The parameters are not handed over as named tensors. They arrive as one flat device buffer and
you are expected to know the layout:</p>

<table>
  <thead>
    <tr>
      <th>Parameters</th>
      <th>Size</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>γ₁, β₁ (LN1)</td>
      <td>768 + 768</td>
    </tr>
    <tr>
      <td>W_qkv, b_qkv</td>
      <td>768×2304 + 2304</td>
    </tr>
    <tr>
      <td>W_attn, b_attn</td>
      <td>768×768 + 768</td>
    </tr>
    <tr>
      <td>γ₂, β₂ (LN2)</td>
      <td>768 + 768</td>
    </tr>
    <tr>
      <td>W_fc, b_fc</td>
      <td>768×3072 + 3072</td>
    </tr>
    <tr>
      <td>W_proj, b_proj</td>
      <td>3072×768 + 768</td>
    </tr>
  </tbody>
</table>

<p>Working out the byte offsets is part of the exercise. It forces you to understand the model
structure at the level of actual memory, not just diagram boxes.</p>

<p>That is what I found interesting about the challenge. The transformer block is not complicated
once you stop treating it as a black box. Attention shares context; the feed-forward network
deepens each representation; residuals and LayerNorm keep the whole thing trainable. The packed
buffer just makes you prove you understood the structure before the hardware will let you run it.</p>

<p>The CUDA implementation is in the
<a href="/ml/2025/12/28/transformer-block-cuda.html">follow-up post</a>.</p>]]></content><author><name>Maciej Grochowski</name></author><category term="ml" /><summary type="html"><![CDATA[At university, long before the current AI period, neural networks were still niche. Academic appeal, limited hardware, results that rarely justified the compute. We wrote simple recurrent networks in C, used them in robotics and image-processing experiments, and spent most of the time fighting the hardware constraints rather than thinking about the model. I liked that part more than I probably should have admitted.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://page-fault.io/assets/social-preview.png" /><media:content medium="image" url="https://page-fault.io/assets/social-preview.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry></feed>