Commit Graph

11 Commits

Author SHA1 Message Date
Brendan Punsky
c67282dfbd rexcode/arm64: an instruction can hold five operands
SME's outer products take five -- `smopa za0.s, p0/m, p1/m, z0.b, z1.b`
-- and Instruction held four, so the 14 forms in that family could not
be represented at all, let alone printed. They named one vector where
the instruction multiplies two.

Five Operands is 55 bytes, which is odd, so Mnemonic's alignment costs
one more; three bytes of padding still land Instruction on exactly one
64-byte cache line, as before. Encoding and Decode_Entry each grow by
two.

The tile number was wrong as well: it sits in the low bits, as wide as
the element size leaves room for -- three for .d down to none for .b --
not at bits 23:22 where the encoding read it. Every MOP form named a
tile it was not writing.

Note for anyone regenerating: tablegen re-emits tables.odin from a
template inside gen.odin, so a struct change there has to go in the
template, and the package has to compile before tablegen can run at all.
Same for the builders.

SVE/SME2 against llvm-mc: 697 byte-exact and 0 mismatched, of 704.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018UmHLRF11EoWwNWCJ7JGaA
2026-08-28 19:29:20 -04:00
Brendan Punsky
d40d9e687a rexcode/arm64: model SME's ZA tiles and tile slices
A ZA tile printed as the bare number it is encoded as -- `addha #0,
p0/m, p0/m, z0.s` -- and a tile slice printed as one too, where the
syntax is `za0h.b[w12, 0]`: a tile, taken along its rows or columns,
addressed by one of W12..W15 plus an offset. Neither was anything an
assembler would take.

Tiles are a register class now (ZA0..ZA15, viewed at an element size),
so they print through the same path as every other register. A slice is
its own operand kind holding the four things it is made of, rather than
one packed immediate that only the encoder understood.

Getting that right needed the field layout, and the layout is not what
the encoding table implied: the tile number and the offset share the low
nibble, and how it splits follows the element size -- a byte tile has no
tile bits at all and four of offset, while a quadword tile is all tile
and none. Reading a fixed four bits as the tile made every byte slice
come back as tile 4.

LD1Q/ST1Q scale their index by 16, which the mnemonic's last letter
does not spell the way B/H/W/D do; they were left unscaled when the
other 38 forms were fixed. That also wanted a .q element shape, which
nothing had needed before.

SVE/SME2 decode entries against llvm-mc: 559 byte-exact and 0
mismatched, from 211 and 78 at the start of the session.

What is left is mostly one structural limit: SME's outer products take
five operands (`umopa za0.s, p0/m, p1/m, z0.b, z1.b`) and Instruction
holds four. Widening it would fit -- five Operands is 55 bytes of the 64
-- but it reaches through Encoding, the table, both codecs and the
builders, so it wants its own change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018UmHLRF11EoWwNWCJ7JGaA
2026-08-28 16:58:33 -04:00
Brendan Punsky
7664ab61cd rexcode/arm64: SVE predicate sizes, gather/scatter addressing, INDEX
Predicates are written with an element size wherever they are an operand
rather than the governing mask -- `zip1 p0.b, p1.b, p2.b`, and `and
p0.b, p1/z, p2.b, p3.b`, where the same instruction has both. 59 source
operands printed bare. Two paths were missing it outright: .PN never
picked up the size at all, and the WHILE family had one form standing in
for all four element sizes, so `whilelt` could only ever print `p0`.

SVE gather and scatter name the element size of their vector index and
how the base extends it (`[x0, z0.s, uxtw]`); none of that was printed.
The two lay their fields out differently and the difference is not
cosmetic: a gather has the extend at bit 22 and takes the index width
from its opcode, while a scatter has the width at 22 and the extend at
14. Reading them the same way meant the .s scatters carried the .d
encoding -- ST1B, ST1H and ST1W all had two forms with identical bits,
so one of each pair was dead.

Memory had no room left for the index's element size (16 + 16 + 23 + 3 +
3 + 3 is exactly 64; the "1 bit spare" comment was stale), so it travels
in the operand's own size field, which memory operands do not otherwise
use.

INDEX had its two operands in each other's slots: SVE_IMM5 is bits
20:16, which is the *second* operand, so the first was written to the
wrong field and the second was not written at all -- `index z0.b, #0,`
with a trailing empty operand. It was also .b-only, and its register
operands are W below .d rather than X.

SVE LDR/STR move a whole register and take no element size; an earlier
pass had given their operands one.

SVE/SME2 decode entries against llvm-mc: 525 byte-exact and 0
mismatched, from 211 and 78 at the start of the session.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018UmHLRF11EoWwNWCJ7JGaA
2026-08-28 16:48:22 -04:00
Brendan Punsky
4bd417b63f rexcode/arm64: 48 instructions could not be decoded at all
An entry whose `bits` set a bit its `mask` does not cover can never
match anything -- `word & mask == bits` is unsatisfiable -- so those
instructions were absent from the decoder entirely. There were 48, and
LDXR/LDAXR were among them: a plain `ldxr w0, [x1]` disassembled to
nothing.

They divide cleanly. Most are the SVE predicated *unary* ops
(ABS/CLS/CLZ/CNT/FABS/FNEG/FSQRT/NEG), which keep their opcode at bits
20:16 -- the same field that was wrong for the binary ops, except here
the stray bits sat in `bits` rather than being left free, so the entry
was dead instead of over-matching. The LDXR family has Rs = 11111 in the
same place. FCADD, SETE, SETM and UZP each had one fixed bit outside
their mask. Every one of them was verified against llvm-mc before
widening: the patterns were right, only the masks were too narrow.

Four needed more than a wider mask:

  - BTI was modelled as taking a hint immediate, but its variants are
    already their own mnemonics (BTI_C/BTI_J/BTI_JC) with exact
    patterns. The bare row is plain `bti` and never took an operand.

  - LUTI2/LUTI4 named a Z pair where the architecture has ZT0, SME2's
    lookup table, and put a register at bits 20:16 where the table index
    lives. ZT0 is modelled now -- one register, no bits -- and the index
    is a real lane index at bits 16:15. Both now cover all three element
    sizes rather than one.

  - MOVA's mask missed bit 17.

Two more the newly-assemblable output exposed: ST1D's scalar+scalar form
carried the .q encoding, and LD1SH's was labelled .s while holding the
.d pattern, with the .s form missing outright.

SVE scalar+scalar addressing scales its index by the access size and an
assembler wants that spelled (`[x0, x0, lsl #2]`), which 38 forms did
not print. The amount is fixed by the form -- and by the *access* size,
not the destination, so LD1SB's is 0 even though it writes halfwords.

SVE/SME2 decode entries against llvm-mc: 438 byte-exact and 0
mismatched, from 211 and 78 at the start of the session. Nothing decodes
to `invalid` any more.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018UmHLRF11EoWwNWCJ7JGaA
2026-08-28 16:14:56 -04:00
Brendan Punsky
049f439c7a rexcode/arm64: SVE predicates, SME2 pairs and quads
Chasing the SME2 gap turned up that the thing blocking it was much
larger than SME2. Every predicated SVE instruction printed its predicate
bare -- `p0` where the syntax needs `p0/z` or `p0/m` -- and an assembler
rejects that outright. 357 forms carried one.

A predicate's governing qualifier is fixed by the form, so it comes from
the operand type and rides in the operand as a marker the printer reads.
Predicated SVE loads and stores also write their vector as a list, so
those 42 forms go through the same one-register list path the NEON work
added: `ld1b { z0.b }, p0/z, [x0, x0]`.

SME2 then needed three things it did not have. A predicate-as-counter
register class -- SME2 governs with pn8..pn15, numbered from 8, sharing
the field but not the register bank. An element size on the pair and
quad operands, which cannot ride on the encoding the way the list length
does, because it is what separates LD1B from LD1H. And the list length
itself, which the pair/quad encodings now carry.

The sweep that verified this found two real encoding bugs behind it:

  - 97 SVE predicated binary ops read Zm from bits 20:16, where the
    architecture has the opcode. `add z0.b, p0/m, z0.b, z1.b` encoded
    0x04010000, which is SUB. Their masks left that opcode field free
    too, so each mnemonic's pattern also matched its siblings'.

  - AND/ORR/EOR/BIC predicated had one form apiece, labelled .d but
    encoding .b, since the size field at bits 23:22 was never in the
    pattern. Split into the four sizes.

SVE/SME2 decode entries against llvm-mc: 312 byte-exact and 1
mismatched, from 211 and 78. The one left is XAR, whose four forms are
legitimately bit-identical -- the element size shares a field with the
shift.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018UmHLRF11EoWwNWCJ7JGaA
2026-08-28 00:37:56 -04:00
Brendan Punsky
1d4887ccb6 rexcode/arm64: vector builders relabelled the caller's register
inst_add(X0, X1, X2) encoded `add v0.16b, v1.16b, v2.16b`. It built the
operand as op_v_16b(u8(reg_hw(dst))), which throws away the register's
class and rebuilds a V register from the bare number -- so an X register
became a V register before the matcher, whose whole job is to reject
that, ever saw it. It encoded, round-tripped, and printed cleanly. Seven
mnemonics with both a scalar and a vector three-register form were
affected: add, and, bic, eor, orn, orr, sub.

The vector constructors now take the register the caller actually has
and put the arrangement in op.size, so the class survives. A wrong class
matches no form and encode reports it; the right class picks the right
form, which is what the matcher was always supposed to do.

The same laundering hid a second bug. Because every arrangement built
the same Odin signature, all of a mnemonic's arrangements collapsed onto
one builder name and only the first survived -- ADD has seven NEON forms
and six were unreachable. 229 mnemonics were in that state. The
arrangement is now part of the builder name (inst_add_v8b_v8b_v8b), so
they are all reachable: 992 builders becomes 1847. The overload group is
unchanged, since it still dedups by Odin signature, so inst_add(V0, V1,
V2) still means .16b as before.

Making them reachable exposed two pre-existing bugs, both fixed here:
CMLE/CMLT/FCMLE/FCMLT compare against zero and the zero is part of the
syntax rather than an encoded operand, so their disassembly was missing
the trailing `#0`/`#0.0` and no assembler would take it; and BFCVTN was
typed .8h at the destination where the architecture says .4h (BFCVTN2 is
the .8h one, and was already right).

Verified against llvm-mc: of the 874 all-register vector builders, 803
are byte-exact and none disagree. The 10 that llvm cannot assemble are
the already-known modelling gaps -- SM3TT lane indices, TBL/TBX register
lists, and PMULL's .1q destination, which the arrangement encoding
cannot represent since 1 lane * 16 bytes collides with 16B. The other 61
are the harness passing V registers where a scalar B/H/S/D/Q view is
required, which is the class check doing its job.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018UmHLRF11EoWwNWCJ7JGaA
2026-08-27 21:26:11 -04:00
Brendan Punsky
dcaab1aa85 rexcode/arm64: system registers get a type instead of being bare i64
They were plain i64 constants handed to op_imm, so any integer typed as
one and `inst_mrs(X0, 999999)` compiled fine. Worse, the printer could
not tell a system register from an immediate and had to recover the
distinction by mnemonic and slot -- MSR's other form holds a PSTATE
field selector in the same position, so it keyed off whether operand 1
was a register.

System_Register is now its own type with its own Operand_Kind, union
member and op_sysreg constructor, exactly as Cond is. The printer's slot
logic is gone: the operand knows what it is, so naming it is a case in
the same switch that prints every other operand kind. MSR's PSTATE
selector is typed PSTATE_FIELD, which is what it always was.

It cannot join `Register` itself: that is a u16 with the class in its
high byte, leaving 8 bits for the number, and a system register needs
15. Widening it would break `Memory`, which packs two registers plus a
displacement and a mode into exactly 64 bits.

The constants are also reorganised. They had accreted into overlapping
sections -- two "ID registers" groups, three cache groups, a "Batch 5:
comprehensive sysreg sweep" banner, and a "hmm let me recompute" note
left in a comment. All 231 are now grouped by architectural function
(18 groups, alphabetical within each) with their five fields aligned.

Verified unchanged against llvm-mc: 222 registers byte-exact through
MRS, 8 write-only through MSR, and the PSTATE form still decodes as an
immediate rather than a register.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018UmHLRF11EoWwNWCJ7JGaA
2026-08-27 21:13:12 -04:00
Brendan Punsky
4406ce69db rexcode/arm64: B.cond is sixteen mnemonics, not one with a condition operand
B_COND was a straight contradiction of the rule the rest of the enum follows.
No assembler has a mnemonic called `b.cond`; it has `b.eq`, `b.le`, `b.ne` and
thirteen more, and x86 in this same library already models exactly this shape
the right way -- JE, JNE, JG, JLE are sixteen separate mnemonics with the
condition in the opcode.

Two things were wrong with folding it into an operand.

The condition is not an operand. It is four bits of the opcode, no different
from x86 putting Jcc's condition in the low nibble of 0x7_. Modelling it as
one has to be papered over everywhere: the printer special-cased B_COND to
rebuild `b.eq` out of operand 0, sbprint had to skip that operand so it did
not print twice, and the public mnemonic_to_string -- which has no instruction
to read the operand from -- returned "b_cond", a string no assembler takes.
All three of those are now gone; the name prints itself.

More importantly it destroyed the flag data. x86 records per condition which
status bits it consults: JE reads {ZF}, JLE reads {ZF, SF, OF}. One B_COND
entry could not say that, so it claimed nzcv_rd = {N, Z, C, V} -- all four
flags, for every condition. Every entry was wrong. Split apart they carry what
they actually read:

  b.eq/b.ne          Z          b.hi/b.ls          Z, C
  b.cs/b.cc          C          b.ge/b.lt          N, V
  b.mi/b.pl          N          b.gt/b.le          N, Z, V
  b.vs/b.vc          V          b.al/b.nv          none

which is the data the compiler's asm checker reads for flag liveness, now that
arm64 feeds it.

The mask also covers the condition field for the first time (0xFF000010 ->
0xFF00001F): with the condition in an operand, four opcode bits sat outside
the mask.

BC.cond gets the same treatment. Builders come out per condition, so `b.le` is
`inst_b_le(label)` rather than `inst_b_cond(.LE, label)`. Cond stays exactly as
it is -- CSEL, CSINC, CSINV, CSNEG, CCMP, CCMN and FCSEL take a real condition
operand, and the compiler's OP_COND handling is untouched.

All sixteen verified by disassembling our own output with llvm-mc: b.eq, b.ne,
b.hs, b.lo, b.mi, b.pl, b.vs, b.vc, b.hi, b.ls, b.ge, b.lt, b.gt, b.le, b.al,
b.nv. Rows follow the table's current column formatting. All rexcode suites
pass and the generators stay idempotent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 08:28:39 -04:00
Brendan Punsky
dd925287ad rexcode/arm64: fix the encode, decode and print bugs the mnemonic pass exposed
Encoder
  * LSR/ASR by immediate used the generic IMM12 encoding, which writes bits
    10-21 -- straight into the imms field the UBFM/SBFM base pattern already
    fills, so immr stayed 0 and every shift encoded as #0 (`asr x0,x1,#7`
    gave 9340fc20, not 9347fc20). They need immr alone, since imms is the
    constant 31/63 fixed in the form: new ENC_SHIFT_IMMR.
  * LDP/STP and friends borrowed the single-register addressing encodings,
    which put an UNSCALED 9-bit displacement at bits 20:12 and OR a pre/post
    marker into bits 11:10. The pair forms want a SCALED 7-bit value at
    21:15, and bits 11:10 are part of Rt2 -- so `ldp x0,x1,[x2,#16]!` came
    back with Rt2=3. New OFFSET_PAIR_4/8/16 (the scale does not follow from
    the register type: LDPSW pairs X registers but loads words, STGP scales
    by 16), with the addressing mode read from bits[24:23] where the
    architecture keeps it. 26 forms retargeted.

Decoder
  * Vector operands came back with size=4 always, so a decoded V register
    lost its arrangement and disassembly printed a bare `v0` that no
    assembler would take. Reconstruct it from the form's operand type.
  * Vd/Vn/Vm/Va hardcoded REG_V, but SVE forms use those same slots with
    Z_REG_* operands -- `add z0.d, z0.d, z0.d` decoded as a V register.
    Take the class from the operand type, as every other slot already does.

Printer
  * V/Z registers now print their arrangement (`add v0.4s, v1.4s, v2.4s`,
    `add z0.d, ...`). Element views (op_v_elem_*) moved from 1/2/4/8 to odd
    codes 1/3/5/7, because an element-D view and an 8B arrangement were both
    size 8 and could not be told apart.
  * MOVZ/MOVN/MOVK print the hw index as `lsl #16`, omitted when zero.
  * BC_COND folds its condition into the mnemonic like B_COND already did,
    instead of printing it twice.

Table (each bit pattern re-derived from llvm-mc)
  * BTI_J and BTI_C had each other's encodings.
  * FCMLA's mask left size bit 22 free, so .4s and .2d were indistinguishable
    and .2d decoded as .4s.
  * BFDOT carried the Q=0 pattern for its .4s/.8h form; PMULLB/PMULLT were
    missing the size field; TLBI PAALL/PAALLOS had the wrong CRm/op2.
  * RDSVL's imm6 sits at bits 10:5, not where IMM6 puts it: ENC_IMM6_LO.
  Nine test expectations that asserted the wrong values were corrected.

specgen.lua
  Was already dead before the mnemonic work -- it wrote to encoding_table.odin
  and spliced a SPECGEN region, neither of which survived the merge into
  instruction_table.odin. Retargeted, taught the canonical names, and made it
  emit Form literals (Encoding + Clobber). It can no longer own whole
  `.MNEM = { ... }` blocks either, since ADD now holds integer, NEON and SVE
  forms together, so it MERGES: a form is added only when no (bits, mask)
  match exists, and existing rows are never rewritten -- their hand-maintained
  Clobber data has to survive a regeneration.

Verified: all 11 rexcode suites match HEAD exactly (arm64 461/461); the three
generator stages stay idempotent; a 73-case differential against llvm-mc is
byte-exact for both encode and decode round-trip. Over the whole decode table,
canonical-form disassembly re-assembled by llvm-mc goes from 594 byte-exact /
1818 unassemblable to 1737 / 678. Re-running specgen re-derives 1130 forms
from llvm-mc and finds every one already present, which independently confirms
those bit patterns.

Still open: multi-vector register lists ({z0.b, z1.b}) and lane indices
(v0.s[2]) are not modelled, so those forms print without them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 18:24:26 -04:00
Brendan Punsky
f4bd6d74f4 rexcode/arm64: mnemonics are assembler mnemonics, not per-encoding names
The Mnemonic enum had one member per encoding form -- ADD_IMM, ADD_SR,
ADD_ER, ADD_V for what an assembler just calls ADD; LDR, LDR_LIT, LDR_PRE,
LDR_POST, LDR_REG, LDR_V for LDR; SVE_ADD_Z / SVE_ADD_PRED / SVE_AND_P for
names SVE spells ADD and AND. The encoder never needed that: like x86, it
already resolves a mnemonic by scanning its run of forms and matching
operand types, so the split bought nothing and cost a printer that had to
strip suffixes back off at runtime -- incompletely, so ADD_V printed
"add.v", LDR_PRE "ldr.pre" and FCVT_H_S "fcvt.h.s".

Collapse the enum to the names assemblers accept: 1104 -> 785 mnemonics,
with the variants becoming forms under one name (ADD now has 13, LDR 15).
LSLV/LSRV/ASRV/RORV fold into LSL/LSR/ASR/ROR. Form order within a run is
precedence, and the original declaration order is already the order an
assembler resolves: "add w0,w1,w2" takes the shifted-register form, and
only the extended form can encode SP.

This needed one structural change. The matcher was blind to addressing
mode -- `case .MEM: return op.kind == .MEMORY` -- which is precisely why
LDR/LDR_PRE/LDR_POST/LDR_REG had to be separate mnemonics; all 20 merge
collisions were this and nothing else. Split Operand_Type.MEM into
mode-specific types (MEM_OFFSET/PRE/POST/REG/EXT plus four SVE), matching
how W_REG/W_SHIFTED/W_EXTENDED are already distinct types over one
register class. The decoder derives Address_Mode from `enc`, so it is
unaffected.

Encodings are unchanged: the multiset of (ops, enc, bits, mask, feature,
flags) over all forms is identical before and after except for two entries
deliberately dropped. NOT_V_ALIAS duplicated NOT_V byte for byte, and
MOV_V_ALIAS was wrong -- it encoded VN where the ORR-based MOV alias needs
VN_VM_DUP, so "mov v1.8b, v2.8b" would have emitted "orr v1.8b, v2.8b,
v0.8b".

AMX_* keeps its prefix: Apple's coprocessor is undocumented with no
assembler spelling, so there is no canonical name to collapse to and bare
"set"/"clr"/"ldx" would mislead. The two-token system instructions keep
theirs too and print with a space (dc zva, tlbi vae1, bti j).

Verified: all 11 rexcode suites match HEAD exactly (arm64 461/461); the
three generator stages round-trip idempotently; 754 of 785 mnemonics are
accepted by llvm-mc, the rest being AMX (24), TME (4, no +tme in this LLVM
build), B_COND/BC_COND and TBL2; and a 39-case encode/print differential
against llvm-mc matches 34, with the 5 others confirmed byte-identical at
HEAD (pre-existing LSR/ASR immediate and LDP pre-index packing bugs, and
printer gaps for vector arrangements and MOVZ/MOVK shifts).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 17:51:36 -04:00
Brendan Punsky
95df04fbe1 rexcode: re-house ISA packages under core:rexcode/isa/<arch>
Move all ten ISA packages (x86, arm32, arm64, mips, riscv, ppc, ppc_vle,
rsp, mos6502, mos65816) from core/rexcode/<arch> to core/rexcode/isa/<arch>,
so the import pattern is now `import "core:rexcode/isa/x86"`. The shared
core stays at core:rexcode/isa.

Mechanical: relative `import "../isa"` / "../../isa" -> absolute
"core:rexcode/isa" (the only path that survives the move; the "../" and
"../.." self/generated imports move with their packages). build.lua now
builds paths as <root>/isa/<name>; stale `cd <arch>` hints in the verify
tools and the doc.odin paths updated.

WASM stays at core/rexcode/wasm for now -- it is an IR, not an ISA, and
will move under the forthcoming core:rexcode/ir once that layer lands.

All 10 arches gen/builders/check/test green; import core:rexcode/isa/x86
verified working; wasm still compiles.
2026-06-18 19:03:27 -04:00