Go to file
Brendan Punsky dd925287ad rexcode/arm64: fix the encode, decode and print bugs the mnemonic pass exposed
Encoder
  * LSR/ASR by immediate used the generic IMM12 encoding, which writes bits
    10-21 -- straight into the imms field the UBFM/SBFM base pattern already
    fills, so immr stayed 0 and every shift encoded as #0 (`asr x0,x1,#7`
    gave 9340fc20, not 9347fc20). They need immr alone, since imms is the
    constant 31/63 fixed in the form: new ENC_SHIFT_IMMR.
  * LDP/STP and friends borrowed the single-register addressing encodings,
    which put an UNSCALED 9-bit displacement at bits 20:12 and OR a pre/post
    marker into bits 11:10. The pair forms want a SCALED 7-bit value at
    21:15, and bits 11:10 are part of Rt2 -- so `ldp x0,x1,[x2,#16]!` came
    back with Rt2=3. New OFFSET_PAIR_4/8/16 (the scale does not follow from
    the register type: LDPSW pairs X registers but loads words, STGP scales
    by 16), with the addressing mode read from bits[24:23] where the
    architecture keeps it. 26 forms retargeted.

Decoder
  * Vector operands came back with size=4 always, so a decoded V register
    lost its arrangement and disassembly printed a bare `v0` that no
    assembler would take. Reconstruct it from the form's operand type.
  * Vd/Vn/Vm/Va hardcoded REG_V, but SVE forms use those same slots with
    Z_REG_* operands -- `add z0.d, z0.d, z0.d` decoded as a V register.
    Take the class from the operand type, as every other slot already does.

Printer
  * V/Z registers now print their arrangement (`add v0.4s, v1.4s, v2.4s`,
    `add z0.d, ...`). Element views (op_v_elem_*) moved from 1/2/4/8 to odd
    codes 1/3/5/7, because an element-D view and an 8B arrangement were both
    size 8 and could not be told apart.
  * MOVZ/MOVN/MOVK print the hw index as `lsl #16`, omitted when zero.
  * BC_COND folds its condition into the mnemonic like B_COND already did,
    instead of printing it twice.

Table (each bit pattern re-derived from llvm-mc)
  * BTI_J and BTI_C had each other's encodings.
  * FCMLA's mask left size bit 22 free, so .4s and .2d were indistinguishable
    and .2d decoded as .4s.
  * BFDOT carried the Q=0 pattern for its .4s/.8h form; PMULLB/PMULLT were
    missing the size field; TLBI PAALL/PAALLOS had the wrong CRm/op2.
  * RDSVL's imm6 sits at bits 10:5, not where IMM6 puts it: ENC_IMM6_LO.
  Nine test expectations that asserted the wrong values were corrected.

specgen.lua
  Was already dead before the mnemonic work -- it wrote to encoding_table.odin
  and spliced a SPECGEN region, neither of which survived the merge into
  instruction_table.odin. Retargeted, taught the canonical names, and made it
  emit Form literals (Encoding + Clobber). It can no longer own whole
  `.MNEM = { ... }` blocks either, since ADD now holds integer, NEON and SVE
  forms together, so it MERGES: a form is added only when no (bits, mask)
  match exists, and existing rows are never rewritten -- their hand-maintained
  Clobber data has to survive a regeneration.

Verified: all 11 rexcode suites match HEAD exactly (arm64 461/461); the three
generator stages stay idempotent; a 73-case differential against llvm-mc is
byte-exact for both encode and decode round-trip. Over the whole decode table,
canonical-form disassembly re-assembled by llvm-mc goes from 594 byte-exact /
1818 unassemblable to 1737 / 678. Re-running specgen re-derives 1130 forms
from llvm-mc and finds every one already present, which independently confirms
those bit patterns.

Still open: multi-vector register lists ({z0.b, z1.b}) and lane indices
(v0.s[2]) are not modelled, so those forms print without them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 18:24:26 -04:00
2025-12-11 11:11:53 +00:00
2025-10-28 13:26:56 +00:00
2018-12-27 10:51:15 +00:00
2026-08-09 20:23:34 +02:00
2022-11-04 11:40:07 +00:00
2026-08-12 16:17:49 +01:00
2025-09-26 12:05:16 +02:00
2025-03-18 15:39:18 +00:00
2026-05-17 13:18:48 +01:00
2020-11-03 07:40:17 -03:00
2025-04-09 07:46:44 +11:00
2025-03-28 18:38:08 +01:00

Odin logo
The Data-Oriented Language for Sane Software Development.


The Odin Programming Language

Odin is a general-purpose programming language with distinct typing, built for high performance, modern systems, and built-in data-oriented data types. The Odin Programming Language, the C alternative for the joy of programming.

Website: https://odin-lang.org/

package main

import "core:fmt"

main :: proc() {
	program := "+ + * 😃 - /"
	accumulator := 0

	for token in program {
		switch token {
		case '+': accumulator += 1
		case '-': accumulator -= 1
		case '*': accumulator *= 2
		case '/': accumulator /= 2
		case '😃': accumulator *= accumulator
		case: // Ignore everything else
		}
	}

	fmt.printf("The program \"%s\" calculates the value %d\n",
	           program, accumulator)
}

Documentation

Getting Started

Instructions for downloading and installing the Odin compiler and libraries.

Nightly Builds

Get the latest nightly builds of Odin.

Learning Odin

Overview of Odin

An overview of the Odin programming language.

Frequently Asked Questions (FAQ)

Answers to common questions about Odin.

Packages

Documentation for all the official packages part of the core and vendor library collections.

Examples

Examples on how to write idiomatic Odin code. Shows how to accomplish specific tasks in Odin, as well as how to use packages from core and vendor.

Odin Documentation

Documentation for the Odin language itself.

Odin Discord

Get live support and talk with other Odin programmers on the Odin Discord.

Articles

The Odin Blog

The official blog of the Odin programming language, featuring announcements, news, and in-depth articles by the Odin team and guests.

Warnings

  • The Odin compiler is still in development.
Languages
Odin 82.3%
C++ 12.3%
C 4.6%
Python 0.3%
Lua 0.2%
Other 0.2%