mirror of
https://github.com/odin-lang/Odin.git
synced 2026-09-02 10:13:35 +00:00
A ZA tile printed as the bare number it is encoded as -- `addha #0, p0/m, p0/m, z0.s` -- and a tile slice printed as one too, where the syntax is `za0h.b[w12, 0]`: a tile, taken along its rows or columns, addressed by one of W12..W15 plus an offset. Neither was anything an assembler would take. Tiles are a register class now (ZA0..ZA15, viewed at an element size), so they print through the same path as every other register. A slice is its own operand kind holding the four things it is made of, rather than one packed immediate that only the encoder understood. Getting that right needed the field layout, and the layout is not what the encoding table implied: the tile number and the offset share the low nibble, and how it splits follows the element size -- a byte tile has no tile bits at all and four of offset, while a quadword tile is all tile and none. Reading a fixed four bits as the tile made every byte slice come back as tile 4. LD1Q/ST1Q scale their index by 16, which the mnemonic's last letter does not spell the way B/H/W/D do; they were left unscaled when the other 38 forms were fixed. That also wanted a .q element shape, which nothing had needed before. SVE/SME2 decode entries against llvm-mc: 559 byte-exact and 0 mismatched, from 211 and 78 at the start of the session. What is left is mostly one structural limit: SME's outer products take five operands (`umopa za0.s, p0/m, p1/m, z0.b, z1.b`) and Instruction holds four. Widening it would fit -- five Operands is 55 bytes of the 64 -- but it reaches through Encoding, the table, both codecs and the builders, so it wants its own change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018UmHLRF11EoWwNWCJ7JGaA