Commit Graph

7 Commits

Author SHA1 Message Date
Brendan Punsky
a99f21922c rexcode/arm64: close the last ten disassembly gaps
The three remaining shapes llvm-mc would not accept back, now that
making every NEON arrangement reachable had exposed them.

PMULL/PMULL2 wrote their destination as .2d where the architecture says
.1q. That arrangement had no operand type because the size marker is
lanes*elem-bytes and 1*16 collides with 16B, so V_1Q takes the next free
multiple of 8 instead. The encodings were already right; only the label
was wrong.

TBL/TBX write their table register as a list, `{v1.16b}`. The braces
belong to the operand rather than the mnemonic, so V_LIST_16B carries
them and the printer stays generic. Only one-register lists are modelled
-- LD1-4/ST1-4 need a count, which is still open.

SM3TT1A/1B/2A/2B were missing their lane index entirely. The mask
already left imm2 free at bits 13:12; the operand simply was not in the
table, so every one of them decoded as index 0 and printed `v2.s` with
no index at all.

Printing that index needed the piece that was never there: a lane index
is its own immediate operand, so it printed as a separate `#2` rather
than glued to the register it indexes. It now carries a marker and the
printer writes `v2.s[3]`. The marker is set from the ENCODING, not the
operand type -- EXT shares .VEC_INDEX for a byte index that really is
written `#3`.

That last part reaches further than these ten: of the 51 lane-indexed
forms, all 51 used to print an index that no assembler would take. 15
are now byte-exact against llvm-mc (SM3TT, DUP, INS) and the other 36
are LD1-4/ST1-4, which additionally need the register-list braces.

The vector sweep is now 809 byte-exact with nothing mismatched and
nothing llvm cannot assemble, from 803/10 before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018UmHLRF11EoWwNWCJ7JGaA
2026-08-27 23:44:13 -04:00
Brendan Punsky
1d4887ccb6 rexcode/arm64: vector builders relabelled the caller's register
inst_add(X0, X1, X2) encoded `add v0.16b, v1.16b, v2.16b`. It built the
operand as op_v_16b(u8(reg_hw(dst))), which throws away the register's
class and rebuilds a V register from the bare number -- so an X register
became a V register before the matcher, whose whole job is to reject
that, ever saw it. It encoded, round-tripped, and printed cleanly. Seven
mnemonics with both a scalar and a vector three-register form were
affected: add, and, bic, eor, orn, orr, sub.

The vector constructors now take the register the caller actually has
and put the arrangement in op.size, so the class survives. A wrong class
matches no form and encode reports it; the right class picks the right
form, which is what the matcher was always supposed to do.

The same laundering hid a second bug. Because every arrangement built
the same Odin signature, all of a mnemonic's arrangements collapsed onto
one builder name and only the first survived -- ADD has seven NEON forms
and six were unreachable. 229 mnemonics were in that state. The
arrangement is now part of the builder name (inst_add_v8b_v8b_v8b), so
they are all reachable: 992 builders becomes 1847. The overload group is
unchanged, since it still dedups by Odin signature, so inst_add(V0, V1,
V2) still means .16b as before.

Making them reachable exposed two pre-existing bugs, both fixed here:
CMLE/CMLT/FCMLE/FCMLT compare against zero and the zero is part of the
syntax rather than an encoded operand, so their disassembly was missing
the trailing `#0`/`#0.0` and no assembler would take it; and BFCVTN was
typed .8h at the destination where the architecture says .4h (BFCVTN2 is
the .8h one, and was already right).

Verified against llvm-mc: of the 874 all-register vector builders, 803
are byte-exact and none disagree. The 10 that llvm cannot assemble are
the already-known modelling gaps -- SM3TT lane indices, TBL/TBX register
lists, and PMULL's .1q destination, which the arrangement encoding
cannot represent since 1 lane * 16 bytes collides with 16B. The other 61
are the harness passing V registers where a scalar B/H/S/D/Q view is
required, which is the class check doing its job.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018UmHLRF11EoWwNWCJ7JGaA
2026-08-27 21:26:11 -04:00
Brendan Punsky
dcaab1aa85 rexcode/arm64: system registers get a type instead of being bare i64
They were plain i64 constants handed to op_imm, so any integer typed as
one and `inst_mrs(X0, 999999)` compiled fine. Worse, the printer could
not tell a system register from an immediate and had to recover the
distinction by mnemonic and slot -- MSR's other form holds a PSTATE
field selector in the same position, so it keyed off whether operand 1
was a register.

System_Register is now its own type with its own Operand_Kind, union
member and op_sysreg constructor, exactly as Cond is. The printer's slot
logic is gone: the operand knows what it is, so naming it is a case in
the same switch that prints every other operand kind. MSR's PSTATE
selector is typed PSTATE_FIELD, which is what it always was.

It cannot join `Register` itself: that is a u16 with the class in its
high byte, leaving 8 bits for the number, and a system register needs
15. Widening it would break `Memory`, which packs two registers plus a
displacement and a mode into exactly 64 bits.

The constants are also reorganised. They had accreted into overlapping
sections -- two "ID registers" groups, three cache groups, a "Batch 5:
comprehensive sysreg sweep" banner, and a "hmm let me recompute" note
left in a comment. All 231 are now grouped by architectural function
(18 groups, alphabetical within each) with their five fields aligned.

Verified unchanged against llvm-mc: 222 registers byte-exact through
MRS, 8 write-only through MSR, and the PSTATE form still decodes as an
immediate rather than a register.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018UmHLRF11EoWwNWCJ7JGaA
2026-08-27 21:13:12 -04:00
Brendan Punsky
5627cc94d6 rexcode/arm64: add the cset/csetm/cinc/cinv/cneg aliases
These are what an assembler writes -- and what one prints back -- for
CSINC/CSINV/CSNEG with the condition inverted, and none of the five
existed. A disassembly of `cset w0, eq` came out as
`csinc w0, wzr, wzr, ne`.

Two constraints the table could not state before:

  - The condition is stored inverted, so COND_HI_INV packs `cond ~ 1`
    and reads it back the same way. The printer needs nothing; the
    decoder hands it a plain condition operand.

  - cinc/cinv/cneg are only the alias when Rn == Rm, which is a
    cross-field equality no mask expresses. One operand fills both
    slots on the way in (RN_RM), and decode checks the two fields agree
    before accepting the entry -- reached only on a mask match, so it
    costs nothing in the scan.

The aliases also require cond != 111x. That one *is* expressible: the
14 legal values are covered exactly by three masked patterns (0xxx,
10xx, 110x), so AL and NV fall through to the underlying instruction
the way llvm-mc does. COND_NOT_AL rejects them on the encode side.

Verified against llvm-mc across all five mnemonics, both widths and all
14 conditions: 140/140 of our printed strings assemble to exactly our
bytes. Disassembly agrees except for cs/hs and cc/lo, which is the
package's existing spelling of those two conditions and shows up on
CSEL and B.cond alike. AL/NV and Rn != Rm both fall through correctly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018UmHLRF11EoWwNWCJ7JGaA
2026-08-27 08:59:59 -04:00
Brendan Punsky
a2cd94f406 rexcode/arm64: catch the tooling up to per-condition branch mnemonics
Splitting B_COND left three places still describing the old model. The
generated builders were already right -- inst_b_le(label), inst_bc_ne(label),
one per condition, all 32 verified to encode and round-trip -- but the
hand-written scaffolding around them was not.

verify_against_llvm normalised our mnemonic by truncating at the first
underscore, which turned B_COND into "b". LLVM prints b.eq/b.ne/..., so the
tool carried 32 alias rows pairing "b" with each of them to stop the mismatch
being reported. That truncation now collapses all sixteen B_* onto "b" and
makes every condition compare equal to every other -- the check would pass
whatever the table said. Keep the condition instead (B_LE -> "b.le") and the
32 alias rows are unnecessary; they are gone.

specgen's canonicalizer kept B_COND and BC_COND off its rename path by name.
Those names no longer exist, so replace the entry with a rule that matches the
shape (BC?_%u%u), which is what the intent was.

And the note in instructions.odin still pointed at inst_b_cond. It now says
what is actually true: a conditional branch is one builder per condition
because the condition is part of the mnemonic, while the select/compare family
-- CSEL, CSINC, CSINV, CSNEG, CCMP, CCMN, FCSEL -- really does take a
condition operand and keeps one.

specgen still re-derives its 1130 forms from llvm-mc and finds every one
already present, all suites pass, and all 13 packages build.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 08:33:23 -04:00
Brendan Punsky
f4bd6d74f4 rexcode/arm64: mnemonics are assembler mnemonics, not per-encoding names
The Mnemonic enum had one member per encoding form -- ADD_IMM, ADD_SR,
ADD_ER, ADD_V for what an assembler just calls ADD; LDR, LDR_LIT, LDR_PRE,
LDR_POST, LDR_REG, LDR_V for LDR; SVE_ADD_Z / SVE_ADD_PRED / SVE_AND_P for
names SVE spells ADD and AND. The encoder never needed that: like x86, it
already resolves a mnemonic by scanning its run of forms and matching
operand types, so the split bought nothing and cost a printer that had to
strip suffixes back off at runtime -- incompletely, so ADD_V printed
"add.v", LDR_PRE "ldr.pre" and FCVT_H_S "fcvt.h.s".

Collapse the enum to the names assemblers accept: 1104 -> 785 mnemonics,
with the variants becoming forms under one name (ADD now has 13, LDR 15).
LSLV/LSRV/ASRV/RORV fold into LSL/LSR/ASR/ROR. Form order within a run is
precedence, and the original declaration order is already the order an
assembler resolves: "add w0,w1,w2" takes the shifted-register form, and
only the extended form can encode SP.

This needed one structural change. The matcher was blind to addressing
mode -- `case .MEM: return op.kind == .MEMORY` -- which is precisely why
LDR/LDR_PRE/LDR_POST/LDR_REG had to be separate mnemonics; all 20 merge
collisions were this and nothing else. Split Operand_Type.MEM into
mode-specific types (MEM_OFFSET/PRE/POST/REG/EXT plus four SVE), matching
how W_REG/W_SHIFTED/W_EXTENDED are already distinct types over one
register class. The decoder derives Address_Mode from `enc`, so it is
unaffected.

Encodings are unchanged: the multiset of (ops, enc, bits, mask, feature,
flags) over all forms is identical before and after except for two entries
deliberately dropped. NOT_V_ALIAS duplicated NOT_V byte for byte, and
MOV_V_ALIAS was wrong -- it encoded VN where the ORR-based MOV alias needs
VN_VM_DUP, so "mov v1.8b, v2.8b" would have emitted "orr v1.8b, v2.8b,
v0.8b".

AMX_* keeps its prefix: Apple's coprocessor is undocumented with no
assembler spelling, so there is no canonical name to collapse to and bare
"set"/"clr"/"ldx" would mislead. The two-token system instructions keep
theirs too and print with a space (dc zva, tlbi vae1, bti j).

Verified: all 11 rexcode suites match HEAD exactly (arm64 461/461); the
three generator stages round-trip idempotently; 754 of 785 mnemonics are
accepted by llvm-mc, the rest being AMX (24), TME (4, no +tme in this LLVM
build), B_COND/BC_COND and TBL2; and a 39-case encode/print differential
against llvm-mc matches 34, with the 5 others confirmed byte-identical at
HEAD (pre-existing LSR/ASR immediate and LDP pre-index packing bugs, and
printer gaps for vector arrangements and MOVZ/MOVK shifts).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 17:51:36 -04:00
Brendan Punsky
95df04fbe1 rexcode: re-house ISA packages under core:rexcode/isa/<arch>
Move all ten ISA packages (x86, arm32, arm64, mips, riscv, ppc, ppc_vle,
rsp, mos6502, mos65816) from core/rexcode/<arch> to core/rexcode/isa/<arch>,
so the import pattern is now `import "core:rexcode/isa/x86"`. The shared
core stays at core:rexcode/isa.

Mechanical: relative `import "../isa"` / "../../isa" -> absolute
"core:rexcode/isa" (the only path that survives the move; the "../" and
"../.." self/generated imports move with their packages). build.lua now
builds paths as <root>/isa/<name>; stale `cd <arch>` hints in the verify
tools and the doc.odin paths updated.

WASM stays at core/rexcode/wasm for now -- it is an IR, not an ISA, and
will move under the forthcoming core:rexcode/ir once that layer lands.

All 10 arches gen/builders/check/test green; import core:rexcode/isa/x86
verified working; wasm still compiles.
2026-06-18 19:03:27 -04:00