This should release our ReleaseFast bundle from 1.1MB to ~800KB.
The major win is disabling logging in ReleaseFast wasm builds (~200KB).
The next is removing aggressive inlining in paths that don't make sense
for performance. Verified with benchmarks on native to not affect
anything really.
The third was really dumb: `var buf: [4096]u32 = @splat(c)` in LLVM
releasefast for wasm was lowering to 4096 separate `i32.store`...
like... 30KB of code. Replacing this with a for loop reduced by 30KB and
made REP (a rare sequence) 11x faster lol.
Resizing a terminal whose buffer is heavy with wide characters (CJK,
emoji) is now 3-5x faster on the reflow path.
This was relatively simple work. We already have a bulk fast path for
same-style cells. We previously omitted ANY wide characters from this.
We relaxed this by making it work with complete wide pairs (wide
followed by spacer tail).
### Benchmarks
Resize dance (13 column resizes, 80 -> 40 -> 132 and back and forth
again) over a ~2,000-row scrollback buffer.
| Workload | Before | After | Speedup |
|---|---|---|---|
| cjk | 0.938 | 0.296 | 3.2x |
| emoji | 0.875 | 0.191 | 4.6x |
| mixed build-log | 0.306 | 0.155 | 2.0x |
| grapheme | 2.449 | 1.831 | 1.3x |
| ascii-short | 0.115 | 0.117 | 1.0x |
| ascii-long | 0.109 | 0.112 | 1.0x |
| latin | 0.093 | 0.095 | 1.0x |
| sgr-truecolor | 0.939 | 0.965 | 1.0x |
**AI usage:** Profiled, implemented, benchmarked, and written by Fable.
Plan validated by me before doing it, I wrote all the
comments/commits/blah.
Resizing a terminal whose buffer is heavy with wide characters (CJK,
emoji) is now 3-5x faster on the reflow path.
This was relatively simple work. We already have a bulk fast path for
same-style cells. We previously omitted ANY wide characters from this.
We relaxed this by making it work with complete wide pairs (wide followed
by spacer tail).
### Benchmarks
Resize dance (13 column resizes, 80 -> 40 -> 132 and back and forth again)
over a ~2,000-row scrollback buffer.
| Workload | Before | After | Speedup |
|---|---|---|---|
| cjk | 0.938 | 0.296 | 3.2x |
| emoji | 0.875 | 0.191 | 4.6x |
| mixed build-log | 0.306 | 0.155 | 2.0x |
| grapheme | 2.449 | 1.831 | 1.3x |
| ascii-short | 0.115 | 0.117 | 1.0x |
| ascii-long | 0.109 | 0.112 | 1.0x |
| latin | 0.093 | 0.095 | 1.0x |
| sgr-truecolor | 0.939 | 0.965 | 1.0x |
**AI usage:** Profiled, implemented, benchmarked, and written by Fable.
Plan validated by me before doing it, I wrote all the
comments/commits/blah.
Processing grapheme-heavy input (ZWJ sequences, emoji modifiers, flags,
combining marks) through is now almost 3x faster.
### Primary Change: PageList Capacity Projection
This workload was heavily bound by `PageList.increaseCapacity` because
pathological cases of single-dimensional growth cause repeated page
capacity doublings which get increasingly expensive because each time we
do a full allocation + clone.
So the major change is that for grapheme bytes in particular, when we
reach a capacity limit, we take the current usage for the current set of
rows and project it out to the remaining capacity of rows. Basically, we
assume that a similar workload will continue. So rather than doubling,
we're _guessing_ how much you're going to need.
In the real world, I'm not really sure if this matters at all. There are
no regressions on any regular corpus streams (asciinema, wikipedia
dumps, etc.).
### Other Changes
There are some other changes here, all found on the path to improving
grapheme IO throughput:
* The bitmap allocator now maintains `search_start` hint we update on
every allocation so that future free-scans are much faster. This is the
lowest possible place we don't have a full bitmap.
* For wasm32, we use an alternate hashing structure for small keys since
Wyhash's 64bit * 64bit multiplication is very very slow because wasm has
no widening instruction.
* Terminal `printSlice` now checks the fast path compatibility once up
front rather than on every fast-path attempt.
### Benchmarks
Data: ZWJ family/profession sequences, skin-tone modifiers, flags, and
combining marks streamed in 64 KiB chunks into an 80x24 terminal,
default modes.
| Benchmark | Before | After | Speedup |
|---|---|---|---|
| wasm, V8, 16 MiB stream | 52 MB/s | 151 MB/s | 2.9x |
| native, terminal-stream, 64 MiB | 894 ms | 305 ms | 2.9x |
Sorry the native stuff is in ms, that's how our native `ghostty-bench`
does things versus the custom little V8 harness.
**AI usage:** Developed alongside Fable: profiling, implementation, and
benchmarks. All human language messages written myself. Validated
myself.
Processing grapheme-heavy input (ZWJ sequences, emoji modifiers, flags,
combining marks) through is now almost 3x faster.
### Primary Change: PageList Capacity Projection
This workload was heavily bound by `PageList.increaseCapacity` because
pathological cases of single-dimensional growth cause repeated page
capacity doublings which get increasingly expensive because each time we
do a full allocation + clone.
So the major change is that for grapheme bytes in particular, when we
reach a capacity limit, we take the current usage for the current set of
rows and project it out to the remaining capacity of rows. Basically, we
assume that a similar workload will continue. So rather than doubling,
we're _guessing_ how much you're going to need.
In the real world, I'm not really sure if this matters at all. There are
no regressions on any regular corpus streams (asciinema, wikipedia dumps, etc.).
### Other Changes
There are some other changes here, all found on the path to improving
grapheme IO throughput:
* The bitmap allocator now maintains `search_start` hint we update on
every allocation so that future free-scans are much faster. This is
the lowest possible place we don't have a full bitmap.
* For wasm32, we use an alternate hashing structure for small keys
since Wyhash's 64bit * 64bit multiplication is very very slow because
wasm has no widening instruction.
* Terminal `printSlice` now checks the fast path compatibility once up front
rather than on every fast-path attempt.
### Benchmarks
Data: ZWJ family/profession sequences, skin-tone modifiers,
flags, and combining marks streamed in 64 KiB chunks into an 80x24
terminal, default modes.
| Benchmark | Before | After | Speedup |
|---|---|---|---|
| wasm, V8, 16 MiB stream | 52 MB/s | 151 MB/s | 2.9x |
| native, terminal-stream, 64 MiB | 894 ms | 305 ms | 2.9x |
Sorry the native stuff is in ms, that's how our native `ghostty-bench`
does things versus the custom little V8 harness.
**AI usage:** Developed alongside Fable: profiling, implementation, and
benchmarks. All human language messages written myself. Validated myself.
This makes the `ghostty_render_state_*` C API significantly faster on
wasm32-freestanding, measured in V8 via Node for Chrome. Also verified
in `jsc` for Safari.
The major change is a new bulk row read API that makes full-screen cell
reads roughly 10x faster for wasm embedders. This should help any
embedder with high FFI overhead, such as Go, Python, etc. too.
Non-wasm performance is not impacted, all benchmarks were run on my mac
too w/ no regressions (two of the changes are native wins as well).
## Changes
* color: the "vectorized" palette conversion loop was silently
scalarized by LLVM into per-byte ops because it loaded/stored through
array-typed pointers. Zig 0.16 disables the LLVM loop vectorizer, so
manually vectorized loops must go through vector-typed pointers.
* C styles: major optimizations to converting Zig styles to C styles.
This is a heavy operation for render state.
* render: `endUpdate`'s style-run fill (`@memset` with a struct value)
re-loaded its source every iteration and stored field by field. Now
manually vectorized.
* render: new `GHOSTTY_RENDER_STATE_ROW_DATA_CELLS_RAW` returns a
borrowed `GhosttyCellsView` of the current row's raw cell values, valid
until the next update. One call per row instead of 3-6 calls per cell.
## Benchmarks
| Benchmark | Before | After | Speedup |
|---|---|---|---|
| colors_get | 114 ns | 35 ns | 3.3x |
| style get, per styled cell | 7.8 ns | 6.7 ns | 1.2x |
| raw+style read, per cell | 8.6 ns | 7.7 ns | 1.1x |
| full-screen text read, per cell | 7.5 ns | 0.7 ns | 10.7x |
| full-screen text+style read, per cell | 8.6 ns | 1.7 ns | 5.1x |
| render state update, styled full frame | 3.4 us | 2.6 us | 1.3x |
**AI usage:** Fable did the implementation and benchmarking and drafted
this message. Comments were partially rewritten by me.
This makes the `ghostty_render_state_*` C API significantly faster on
wasm32-freestanding, measured in V8 via Node for Chrome. Also verified
in `jsc` for Safari.
The major change is a new bulk row read API that makes full-screen cell reads
roughly 10x faster for wasm embedders. This should help any embedder with
high FFI overhead, such as Go, Python, etc. too.
Non-wasm performance is not impacted, all benchmarks were run on my mac
too w/ no regressions (two of the changes are native wins as well).
## Changes
* color: the "vectorized" palette conversion loop was silently
scalarized by LLVM into per-byte ops because it loaded/stored through
array-typed pointers. Zig 0.16 disables the LLVM loop vectorizer, so
manually vectorized loops must go through vector-typed pointers.
* C styles: major optimizations to converting Zig styles to C styles.
This is a heavy operation for render state.
* render: `endUpdate`'s style-run fill (`@memset` with a struct value)
re-loaded its source every iteration and stored field by field. Now
manually vectorized.
* render: new `GHOSTTY_RENDER_STATE_ROW_DATA_CELLS_RAW` returns a
borrowed `GhosttyCellsView` of the current row's raw cell values, valid
until the next update. One call per row instead of 3-6 calls per cell.
## Benchmarks
| Benchmark | Before | After | Speedup |
|---|---|---|---|
| colors_get | 114 ns | 35 ns | 3.3x |
| style get, per styled cell | 7.8 ns | 6.7 ns | 1.2x |
| raw+style read, per cell | 8.6 ns | 7.7 ns | 1.1x |
| full-screen text read, per cell | 7.5 ns | 0.7 ns | 10.7x |
| full-screen text+style read, per cell | 8.6 ns | 1.7 ns | 5.1x |
| render state update, styled full frame | 3.4 us | 2.6 us | 1.3x |
**AI usage:** Fable did the implementation and benchmarking and drafted
this message. Comments were partially rewritten by me.
This builds and publishes `ghostty-vt.wasm` binaries into our tip GitHub
releases. These are built with the proper optimization, `simd128` CPU
feature set, and run through `wasm-opt`.
This allows wasm consumers to use libghostty without a Zig toolkit.
Published two: `ghostty-vt.wasm` and `ghostty-vt-small.wasm`. The latter
is ReleaseSmall, but is 10 to 20% slower. Users choice.
This builds and publishes `ghostty-vt.wasm` binaries into our tip
GitHub releases. These are built with the proper optimization, `simd128`
CPU feature set, and run through `wasm-opt`.
This allows wasm consumers to use libghostty without a Zig toolkit.
Published two: `ghostty-vt.wasm` and `ghostty-vt-small.wasm`. The latter
is ReleaseSmall, but is 10 to 20% slower. Users choice.
This makes `ghostty_terminal_vt_write` on wasm32-freestanding anywhere
from 1.4x to 13x faster depending on the input, measured in V8 via Node
for Chrome as well as `jsc` for Safari.
This changes the default Wasm build to default to enabling the `simd128`
CPU feature because baseline doesn't have that and every major browser
has supported it for years. This results in massive performance
improvements (like, 50%+ on all streams).
Non-wasm performance is not impacted, all benchmarks were run on my mac
too w/ no regressions.
## Changes
* stream: the batched parse path (bulk UTF-8 decode, print_slice runs)
is used even when `build_options.simd` is false. The per-byte loop is
now debug-only.
* simd/vt: the scalar `utf8DecodeUntilControlSeq` gets a vectorized
ASCII bulk path that is compatible with wasm simd128.
* style: on wasm, `Style.eql` compares canonical `PackedStyle` forms
which is faster by like 11%. On native its slower so we only do this for
wasm.
* build: wasm targets now default to the `simd128` CPU feature since
every browser engine has supported it for years. Opt out with
`-Dcpu=generic`.
* PACKAGING.md documents the wasm build, including `wasm-opt` notes.
## Benchmarks
| Workload | Before | After | Speedup |
|---|---|---|---|
| ascii | 85 MB/s | 1070 MB/s | 12.5x |
| ascii-wrap | 84 MB/s | 1103 MB/s | 13.1x |
| clear-redraw | 85 MB/s | 913 MB/s | 10.7x |
| scroll | 79 MB/s | 304 MB/s | 3.8x |
| cursor | 120 MB/s | 255 MB/s | 2.1x |
| utf8 | 99 MB/s | 169 MB/s | 1.7x |
| sgr16 | 81 MB/s | 133 MB/s | 1.6x |
| sgr-truecolor | 62 MB/s | 88 MB/s | 1.4x |
End result: wasm at roughly 50-85% of the native ReleaseFast+SIMD build
on the same workloads. Plain ASCII was at 6% of native before.
**AI usage:** Lots of Fable help. As always, the human language stuff
like this commit and comments were rewritten by me.
This makes `ghostty_terminal_vt_write` on wasm32-freestanding anywhere
from 1.4x to 13x faster depending on the input, measured in V8 via Node
for Chrome as well as `jsc` for Safari.
## Changes
* stream: the batched parse path (bulk UTF-8 decode, print_slice runs)
is used even when `build_options.simd` is false. The per-byte loop
is now debug-only.
* simd/vt: the scalar `utf8DecodeUntilControlSeq` gets a vectorized
ASCII bulk path that is compatible with wasm simd128.
* style: on wasm, `Style.eql` compares canonical `PackedStyle` forms
which is faster by like 11%. On native its slower so we only do this
for wasm.
* build: wasm targets now default to the `simd128` CPU feature since
every browser engine has supported it for years. Opt out with
`-Dcpu=generic`.
* PACKAGING.md documents the wasm build, including `wasm-opt` notes.
## Benchmarks
| Workload | Before | After | Speedup |
|---|---|---|---|
| ascii | 85 MB/s | 1070 MB/s | 12.5x |
| ascii-wrap | 84 MB/s | 1103 MB/s | 13.1x |
| clear-redraw | 85 MB/s | 913 MB/s | 10.7x |
| scroll | 79 MB/s | 304 MB/s | 3.8x |
| cursor | 120 MB/s | 255 MB/s | 2.1x |
| utf8 | 99 MB/s | 169 MB/s | 1.7x |
| sgr16 | 81 MB/s | 133 MB/s | 1.6x |
| sgr-truecolor | 62 MB/s | 88 MB/s | 1.4x |
End result: wasm at roughly 50-85% of the native ReleaseFast+SIMD build
on the same workloads. Plain ASCII was at 6% of native before.
**AI usage:** Lots of Fable help. As always, the human language stuff
like this commit and comments were rewritten by me.
This improves the performance of render state plus C API reads. I
specifically benchmarked the C API call and found a lot of overhead in
the C API layer which this cleans up. The impact of these changes will
be less visible to Zig consumers but moderately improve there.
All benchmark numbers below are via the C API.
Highlights:
- full rebuilds are **1.71x faster (11.4µs to 6.6µs per 120x80 frame)**
- single-dirty-row updates (e.g. the TUI/prompt steady state) are
**1.44x faster**
- full-frame reads through the C API are **1.2x to 1.8x faster**
## Changes
* endUpdate skips unchanged style runs.
* `GRAPHEMES_UTF8` getter gets a fast path for single ASCII codepoints
(the overwhelming majority of cells).
* The bg/fg color getters no longer copy the full 28-byte style.
Instead, they switch directly on the one color field they need.
* The `get_multi` variants validate the handle and position once per
batch instead of per key.
* Iterator positions are sentinel values instead of Zig optionals. The
optional tagging overhead was showing up in benchmarks.
* `colors_get` reads through a pointer instead of copying the ~1KB
colors struct to the stack per call.
* The palette conversion is vectorized. The 4-byte padded RGB to 3-byte
was not being auto-vectorized. Explicitly vectorize it. Something like a
4x speedup on NEON.
## Benchmarks
| Benchmark | Before | After | Speedup |
|---|---|---|---|
| update (forced full rebuild) | 11.4 µs/frame | 6.6 µs/frame | 1.71x |
| update (single dirty row) | 143 ns | 99 ns | 1.44x |
| read cell style/bg/fg/selected | 10.3 ns/cell | 8.8 ns/cell | 1.17x |
| read cell via get_multi | 9.6 ns/cell | 6.9 ns/cell | 1.40x |
|read cell UTF-8 text | 4.9 ns/cell | 2.7 ns/cell | 1.78x |
| colors_get + palette | 213 ns/call | 45 ns/call | 4.58x |
Clean updates (no terminal changes) and the raw cell read paths are
unchanged.
**AI usage:** Driven by Fable primarily, reviewed everything and rewrote
all human-language (comments) since Fable in particular does really bad
at that. This commit message too.
This improves the performance of render state plus C API reads. I
specifically benchmarked the C API call and found a lot of overhead in
the C API layer which this cleans up. The impact of these changes will be
less visible to Zig consumers but moderately improve there.
All benchmark numbers below are via the C API.
Highlights:
- full rebuilds are **1.71x faster (11.4µs to 6.6µs per 120x80 frame)**
- single-dirty-row updates (e.g. the TUI/prompt steady state) are *1.44x
faster**
- full-frame reads through the C API are **1.2x to 1.8x faster**
## Changes
* endUpdate skips unchanged style runs.
* `GRAPHEMES_UTF8` getter gets a fast path for single ASCII codepoints (the
overwhelming majority of cells).
* The bg/fg color getters no longer copy the full 28-byte style. Instead,
they switch directly on the one color field they need.
* The `get_multi` variants validate the handle and position once per batch
instead of per key.
* Iterator positions are sentinel values instead of Zig optionals. The
optional tagging overhead was showing up in benchmarks.
* `colors_get` reads through a pointer instead of copying the ~1KB colors
struct to the stack per call.
* The palette conversion is vectorized. The 4-byte padded RGB to 3-byte
was not being auto-vectorized. Explicitly vectorize it. Something like
a 4x speedup on NEON.
## Benchmarks
| Benchmark | Before | After | Speedup |
|---|---|---|---|
| update (forced full rebuild) | 11.4 µs/frame | 6.6 µs/frame | 1.71x |
| update (single dirty row) | 143 ns | 99 ns | 1.44x |
| read cell style/bg/fg/selected | 10.3 ns/cell | 8.8 ns/cell | 1.17x |
| read cell via get_multi | 9.6 ns/cell | 6.9 ns/cell | 1.40x |
| read cell UTF-8 text | 4.9 ns/cell | 2.7 ns/cell | 1.78x |
| colors_get + palette | 213 ns/call | 45 ns/call | 4.58x |
Clean updates (no terminal changes) and the raw cell read paths
are unchanged.
**AI usage:** Driven by Fable primarily, reviewed everything and rewrote
all human-language (comments) since Fable in particular does really bad
at that. This commit message too.
Cell contents used our ArrayListCollection container to manage per-row
foreground lists. This was the only place ArrayListCollection was used.
We now own the row list slice directly, initialize cursor capacity to
exactly one cell, and reallocate the contiguous background buffer in
place when possible. Foreground rows still use exact sizes so resizes
(which are infrequent) do not retain the high-water mark.
Cell contents used our ArrayListCollection container to manage per-row
foreground lists. This was the only place ArrayListCollection was used.
We now own the row list slice directly, initialize cursor capacity to
exactly one cell, and reallocate the contiguous background buffer in
place when possible. Foreground rows still use exact sizes so resizes
(which are infrequent) do not retain the high-water mark.
We shouldn't hold a closing surface view when sending notifications and
waiting to dismiss that notification. This happens rarely, but it's the
right thing to do.
### AI Disclosure
Found by Claude when judging other branches, I applied the changes
myself.
> Forgot that after force pushing, you can't reopen#13787 , linking it
here for the review history.
This PR also updates old translations to keep better consistency.
AI Disclaimer: I translated manually all strings and then used an agent
to review consistency and legibility and applied many suggestions.
Minor things as I was just auditing the state of our headers:
* Make sure all subheaders like `point.h` can be included standalone
* Support Clang/GCC extensions for typed enums if we can detect it
* Add missing structs to the `ghostty_type_json` function
* Fix `GHOSTTY_INIT_SIZED` for C++ mode
Use an immediately invoked lambda for GHOSTTY_INIT_SIZED in C++ so the
macro value-initializes every field before setting the ABI size. The
previous C compound literal and designated initializer required compiler
extensions in C++17 and C++20.
Keep the existing standard compound literal for C callers.
Use fixed int enum types for C++11, C23, Clang's fixed-enum extension,
and GCC 13 or newer. Previously only finalized C23 mode selected an explicit
underlying type, leaving C++ and common older C modes with
implementation-defined enum types.
Dynamic tabstop storage treated the number of columns above the inline
capacity as a byte count even though each byte stores eight stops. This
was wasteful, although in practice this is in the order of just bytes.
Round the extra column count up to whole storage units and grow the
existing slice with realloc, preserve existing stops.
Fixes#13774
Simple one-line fix: added the missing flush to the function.
- `zig build` passed
- `zig build test` passed
- Confirmed template appears after fix in `~/Library/Application
Support/com.mitchellh.ghostty/config.ghostty`
---
**AI Disclosure**
As mentioned before - I used OpenAI Codex (Sol 5.6 high):
a) to see if the template still exists & a function exists that tries to
use the template
b) find the moment in history this behavior changed.
Additionally:
c) reviewing my work and draft a commit message matching your preferred
style
I implemented the change, ran the build/test, manually verified the
behavior, and understand the fix.