- benchmark: avoid buffers to avoid a memcpy
- build: keep frame pointers on macOS. There was some debug changes from
Zig 0.15 and this helps. Also, Apple actually requires/expects x29 to
always be a frame pointer.
- build/macos: force libSystem symbols instead of compiler-rt
- global: add InitOpts.tool so that ghostty-gen/bench can parse their
own actions in `+action`
- quirks: provide our own vectorized memset. see the comment for more
details why.
- synthetic: fix UB by accessing global.io before it was initialized
- terminal/hash_map: force inline for unique repr types. Zig 0.15
inlined and 0.16 doesn't, measured a huge slowdown in hyperlink
benchmarks.
- terminal: add explicit `@Vector` usage for storing a run of identical cells
as well as for scanning printable cells. This auto-vectorized in Zig
0.15 but not in Zig 0.16. This produces the same assembly.
- unicode: properties and LUT need power-of-two backing integer to avoid
bad LLVM codegen
Profiling terminal-stream on a 2.6 GB recording of real terminal
sessions showed ~9% of total time inside the UTF-8 decode stage,
and most of it was not the decode itself: real streams contain an
escape sequence every ~18 bytes, so utf8DecodeUntilControlSeq is
called on short printable runs, and each call paid simdutf setup
plus its scalar rewind_and_convert_with_errors tail (which handles
the last partial SIMD block of every conversion) for only a
handful of bytes. The scalar tail alone accounted for ~3.4% of
total time.
Terminal input is also overwhelmingly ASCII, for which UTF-8 to
UTF-32 "decoding" is just widening each byte to 32 bits. This
fuses the two passes: while scanning each chunk for ESC we also
check for bytes >= 0x80 and widen pure-ASCII chunks straight into
the output vector via PromoteTo, never touching simdutf. The first
non-ASCII byte hands the remainder of the run (up to the next ESC)
to the existing simdutf-based path, so non-ASCII text takes
exactly the same code as before. Inputs shorter than one vector
are handled by a scalar byte loop that likewise skips simdutf for
ASCII.
The widening store needs a dedicated path for the HWY_SCALAR
fallback target (compiled on targets without guaranteed SIMD, e.g.
arm-linux-androideabi): its single-lane vectors cannot be halved
so the one lane is widened directly.
The new differential fuzz test verifies the SIMD implementation
still matches the scalar reference exactly. Measured with
ghostty-bench terminal-stream (2.6 GB real-session corpus, 87%
printable ASCII / 5.5% ESC / 5.6% UTF-8, 120x80, M4 Max,
ReleaseFast, hyperfine means):
| stream | before | after | change |
|-------------------|-----------------|-----------------|--------|
| real 2.6 GB corpus | 9.582 s (272 MB/s) | 9.090 s (287 MB/s) | +5.4% |
The scalar fallback of utf8DecodeUntilControlSeq (used when SIMD is
disabled, e.g. wasm builds) treated a valid-so-far but incomplete
UTF-8 sequence at the end of its decode region as pending input in
all cases: it stopped without consuming the bytes so a future chunk
could complete the sequence. That is correct when the region ends
at the end of the input, but the region can also be bounded by an
ESC byte. In that case the sequence can never be completed (the
next byte is already known to be ESC), and the SIMD implementation,
via simdutf, replaces the ill-formed prefix with U+FFFD and
consumes up to the ESC. The two implementations disagreed on both
the consumed count and the decoded output for inputs like
"\xC2\x1B[0m".
The divergence is invisible at the stream level (the pending bytes
take the scalar nextUtf8 path which also emits a replacement
character once it sees the ESC) but it means the scalar decoder is
not a faithful reference for the SIMD one.
This makes the scalar decoder treat a partial sequence bounded by
an ESC as a maximal subpart per Unicode 3-7: one U+FFFD, consumed
through the end of the region. Truncation at the true end of input
still leaves the bytes pending. Also adds a differential fuzz test
that runs 10k random mixtures of ASCII, escapes, controls, and
valid/invalid UTF-8 through both implementations and requires
identical results, which is what caught this.
This updates simdutf to my fork which has a SIMDUTF_NO_LIBCXX option
that removes all libc++ and libc++ ABI dependencies.
From there, the hand-written simd code we have has been updated to also
no longer use any libc++ features. Part of this required removing utfcpp
since it depended on libc++ (`<iterator>`).
libghostty-vt now only depends on libc.
The vendored Highway package was being built with libc++ even though
Ghostty only uses its runtime target selection and dispatch support.
That pulled in extra C++ runtime baggage from upstream support files
such as abort, timer, print, and benchmark helpers.
Build Highway in HWY_NO_LIBCXX mode, only compile the target dispatch
sources we actually need, and compile Ghostty's SIMD translation units
with the same define so the header ABI stays consistent. Replace the
upstream abort implementation with a small local bridge that provides
Highway's Warn/Abort hooks and the target-query shim without depending
on libc++.
This keeps the Highway archive down to the dispatch pieces Ghostty
uses while preserving the existing dynamic dispatch behavior. The
bridge is documented so it is clear why Ghostty carries this small
local replacement.
Zig's bundled libc++/libc++abi conflicts with the MSVC C++ runtime
headers (vcruntime_typeinfo.h, vcruntime_exception.h, etc.) when
targeting native-native-msvc. This caused compilation failures in
the SIMD C++ code due to -nostdinc++ suppressing MSVC headers and
libc++ types clashing with MSVC runtime types.
Skip linkLibCpp() for MSVC targets across all packages (highway,
simdutf, utfcpp) and the main build (SharedDeps, GhosttyZig) since
MSVC provides its own C++ standard library natively. Also add
missing <iterator> and <cstddef> includes that were previously
pulled in transitively through libc++ headers but are not
guaranteed by MSVC's headers.