Reworks the terminal PAGE grid wire format and optimize both
encode/decode. Example improvements for 1MB of VT input w/ full
scrollback: ~30x smaller wire size, ~45x faster encoding and decoding.
> [!IMPORTANT]
>
> **Snapshot version 1 is still explicitly a work-in-progress format, so
this breaks wire compatibility**.
The original snapshot version I merged favored simplicity over
optimization. This was the format used a proof-of-concept in my own
projects, but I knew it wasn't what I wanted to ship. This PR looks at
the record formats and trades simplicity for performance, a fair trade
for a performance-sensitive binary format.
Overview of changes:
- **8-byte grid cells.** Cells are now one 64-bit word whose layout
deliberately coincides with the native cell. Previously, cells were 16
bytes each and in our 1MB corpus 97% of the data was `0`. Lol.
- **Blank trailing cells are not written.** Rows declare an encoded cell
count so trailing blank cells cost nothing.
- **Hardware CRC32C.** Added `src/crc32c.zig` that uses inline-asm on
aarch64/x86_64 to get hardware speeds for CRC32. Zig's stdlib is 0.56
GB/s, aarch64 hardware is 10 GB/s on my computer.
- **Variable-width cells.** Each row declares how many bytes transport
each cell word: 1, 2, 4, or 8 depending on the widest row cell.
## Format
Grid layout, per PAGE record:
```
old new
+--------------------------+ +--------------------------+
| row 0 | | row 0 |
| flags (1) | | flags + width (1) |
| cols * 16B cells with | | encoded cell count (2) |
| inline suffixes | | count * width cells |
+--------------------------+ +--------------------------+
| ... | | ... |
+--------------------------+ +--------------------------+
| row (rows - 1) | | row (rows - 1) |
+--------------------------+ +--------------------------+
| grapheme suffix section |
+--------------------------+
```
Every row previously carried exactly `cols` cells; now it carries cells
only through its last non-default cell, and the cells past the count are
implicitly zero. The row flag byte gains the encoded cell width in its
previously reserved bits:
```
bit 0 wrap bit 2-3 semantic prompt
bit 1 wrap continuation bit 4-5 encoded cell width (log2 bytes)
```
The cell itself, old fixed 16-byte header versus the new single word:
```
old (16 bytes + inline suffixes) new (one u64 word)
+--------+---------+--------+ bit 0 +------------------+
| kind 1 | width 1 | flags 1| | content kind 2b |
+--------+---------+--------+ bit 2 +------------------+
| zero 1 | style id 2 | | content 24b |
+--------+------------------+ bit 26 +------------------+
| hyperlink id 2 | | style ID 16b |
+---------------------------+ bit 42 +------------------+
| value 4 | | width kind 2b |
+---------------------------+ bit 44 +------------------+
| grapheme count 4 | | protected 1b |
+---------------------------+ bit 45 +------------------+
| grapheme cps 4 * count | | hyperlink 1b |
+---------------------------+ bit 46 +------------------+
| semantic 2b |
bit 48 +------------------+
| hyperlink ID 16b |
bit 64 +------------------+
```
The word's bit layout intentionally matches the native cell (with the
wire hyperlink ID in the native padding), so full-width rows are a
straight copy of page memory. The row's encoded width then transports
each word truncated, and decode is the matching zero-extension:
```
width | bytes | admitted cells
------+-------+------------------------------------------------
0 | 1 | codepoint <= U+00FF, nothing else set
1 | 2 | codepoint <= U+FFFF, nothing else set
2 | 4 | any content kind/codepoint, style IDs 1-63,
| | narrow, no flags, no hyperlink
3 | 8 | everything
```
Grapheme suffixes were inline after each cell, which forced per-cell
framing decisions; they are now one section after the rows, so a
grapheme-free page (the overwhelming case) pays 4 bytes total:
```
old: ... | cell | cp cp | cell | ... (inline, per cell)
new: +----------------+----------------------------------+
| entry count 4 | entries: row 2, col 2, count 2, |
| | count * codepoint 4 |
+----------------+----------------------------------+
```
## Performance
Setup: `ghostty-bench +terminal-snapshot`, 80x24 terminal with unlimited
scrollback fed 1 MB of VT input.
Per-commit improvements:
| change | wire size | encode | decode |
|-------------------------------|-----------|---------|----------|
| baseline (v1 before this PR) | 34.16 MB | 92.8 ms | 119.8 ms |
| 8-byte cells + blank elision | 7.66 MB | 18.2 ms | 28.0 ms |
| hardware CRC32C | 7.66 MB | 5.8 ms | 15.5 ms |
| gate page verification | 7.66 MB | 5.8 ms | 12.2 ms |
| staged PAGE payload decoding | 7.66 MB | 5.8 ms | 8.1 ms |
| variable-width cells | 1.03 MB | 2.0 ms | 2.7 ms |
Final result across various inputs:
| corpus | wire size | encode | decode |
|----------------------------|------------------------|----------------------|-----------------------|
| ascii lines 1-70 | 34.16 -> 1.03 MB (33x) | 92.8 -> 2.0 ms (46x) |
119.8 -> 2.7 ms (44x) |
| ascii full-width wrap | 16.01 -> 1.04 MB (15x) | 43.6 -> 1.3 ms (34x)
| 56.2 -> 1.8 ms (31x) |
| utf8 (wide/grapheme heavy) | 4.33 -> 1.89 MB (2.3x) | 12.4 -> 1.7 ms
(7x) | 19.0 -> 2.2 ms (9x) |
### Relationship with Compression
I expect that users of this will wrap everything in compression, so I
also benchmarked all my changes against a caller-owned zstd compressor
to ensure we're making the write tradeoffs. Less bytes means less time
in a compressor, even if a ton of 0s compresses really well.
My results: `zstd -1` over the `lines` snapshot drops from 12.7 ms to
0.8 ms, and the compressed artifact shrinks from 1.35 MB to 0.86 MB. So
the end state is a win-win.
Add a per-row encoded cell width to the PAGE grid format. Rows
previously always spent eight bytes per cell, but a plain text cell
carries only a codepoint: on line-shaped scrollback most encoded
bytes were predictable zeros that still had to pass through CRC32C,
BLAKE3, both codecs, and any transport compression the caller
applies.
Each row now declares one of four widths in previously reserved row
flag bits, chosen canonically as the smallest width admitted by the
bitwise OR of the row cell words: one or two bytes transport a bare
codepoint, four bytes transport the low word half (any content kind,
style IDs up to sixty-three, no wide or flag or hyperlink bits), and
eight bytes remain the full word. Every width is a truncation on
encode and a zero-extension on decode, so narrow rows encode and
decode as vectorizable integer loops, one and two byte rows need at
most surrogate replacement and skip cell normalization entirely, and
full-width rows keep the existing bulk copy. Decoders use the
declared width for framing and accept rows encoded wider than
necessary. Rows containing wide characters, hyperlinks, semantic
content, or large style IDs still use the full width, which leaves
CJK-heavy content unchanged.
Benchmark deltas at this commit (terminal-snapshot, M-series,
ReleaseFast, 1 MB corpora):
ascii lines 1-70: 7.66 MB -> 1.03 MB (7.4x)
encode 5.8 -> 2.0 ms, decode 8.1 -> 2.7 ms
ascii full-wrap: 8.04 MB -> 1.04 MB (7.7x)
encode 5.4 -> 1.3 ms, decode 7.2 -> 1.8 ms
utf8: unchanged (wide cells keep rows at full width)
For a caller compressing the stream, the lines snapshot end to end
with zstd -1: encode plus compress 18.5 -> 2.8 ms, decompress plus
decode 15.8 -> 3.6 ms, and the compressed size itself drops from
1.35 MB to 0.86 MB because the packed stream is denser for the
entropy coder.
PAGE payloads were decoded through a stack of stream adapters:
a CRC32C-hashing reader over a length-limited reader over the
BLAKE3-hashing snapshot reader. Every row paid several adapter
crossings and both hashes were fed row-sized chunks, which kept
BLAKE3 out of its efficient many-block path and made adapter
overhead about a quarter of decode time.
Decode now reads the remaining payload into a scratch buffer with
one bulk read, so each hash sees the payload as a single update, and
then parses the tables and grid from a flat in-memory reader. Row
headers are also read as one three-byte read instead of two calls.
Staging is capped at 8 MiB, far above any standard-capacity page
payload, so a hostile declared length cannot force a large
allocation; larger payloads fall back to the streaming path. CRC
validation and exact-exhaustion checks are unchanged, with the
staged reader checked for leftover bytes to preserve
PayloadNotExhausted semantics.
Benchmark deltas at this commit (terminal-snapshot, 1 MB corpora):
ascii lines 1-70: decode 12.2 -> 8.1 ms (encode unchanged)
ascii full-wrap: decode 11.1 -> 7.2 ms
utf8: decode 3.1 -> 2.1 ms
Relative to the previous wire format and codecs, the series is a
16.0x encode and 14.8x decode improvement on line-shaped scrollback
at 4.5x smaller wire size.
PAGE decoding verified the complete native integrity of every decoded
page unconditionally, building per-cell reference maps that accounted
for roughly a fifth of decode time. The decoder normalizes every
semantic value while decoding, so a completed decode upholds page
invariants by construction and the verification only defends against
decoder bugs. Follow the native page policy instead: assertIntegrity
and friends run full verification only when slow runtime safety is
enabled, which keeps the check in debug and test builds where those
bugs are caught.
Benchmark deltas at this commit (terminal-snapshot, 1 MB corpora):
ascii lines 1-70: decode 15.5 -> 12.2 ms (encode unchanged)
Rework the PAGE grid encoding for codec speed and size. This is a
breaking change to the work-in-progress version 1 wire format.
Cells were previously a fixed 16-byte header plus inline grapheme
suffixes: one byte each for content kind, width, and flags, a
reserved byte, 16-bit style and hyperlink IDs, a 32-bit value, and an
always-present 32-bit suffix count that was almost always zero. Cells
are now one 64-bit little-endian word with a documented bit registry
that carries the hyperlink ID in its high bits. The layout
deliberately coincides with the native cell so clean rows encode as a
straight copy of page memory and decode as one bulk read plus an
in-place normalization pass; a comptime check falls back to a
portable field-by-field codec if the native layout ever diverges.
Each row header also gains an encoded cell count so trailing default
cells are elided instead of spending 16 bytes apiece encoding
nothing: on typical shell output most of every row is blank, and
measurement showed 97% of encoded snapshot bytes were zero. Grapheme
suffixes move out of the cell stream into a per-grid section of
(row, column, codepoints) entries, which keeps row decoding
fixed-stride and bulk-copyable.
Decode ID remapping switches from hash maps to direct-indexed tables
sized by the 16-bit encoded ID space, removing per-styled-cell hash
lookups.
Benchmark deltas at this commit (terminal-snapshot, M-series,
ReleaseFast, 1 MB corpora):
ascii lines 1-70: 34.16 MB -> 7.66 MB (4.5x)
encode 92.8 -> 18.2 ms, decode 119.8 -> 28.0 ms
ascii full-wrap: 16.01 MB -> 8.04 MB (2.0x)
encode 43.6 -> 18.4 ms, decode 56.2 -> 25.6 ms
utf8: 4.33 MB -> 1.90 MB (2.3x)
encode 12.4 -> 4.9 ms, decode 19.0 -> 9.6 ms
Previously if you select "**Install and Relaunch**" in the update pill,
there's still a confirmation alert about killing active process, but it
will just relaunch regardlessly(which is the intended behaviour) when
you do it in the command palette.
Using UpdateState to check so Ghostty will relaunch immediately when
user chooses "Install and Relaunch".
> For the auto update case, this will be handled a bit differently in
the future. If an update is already installed and waiting for relaunch
(that the user is not aware of), quitting Ghostty will still remind them
if there's active processes.
OSC 52 clipboard writes decoded their base64 payload with the scalar
std implementation while Kitty graphics payloads already used the SIMD
decoder in src/simd.
Move clipboard path to use the same SIMD decoder. The encode side
of the read reply stays scalar since the codebase has no SIMD
encoder. 3.3x faster on a 4KB-payload decode micro-benchmark.
Measures the terminal binary snapshot codecs in both directions
against the same terminal state. Setup feeds a pre-generated VT
stream (for example from ghostty-gen ascii) to a terminal outside
the timed region.
Baseline measurements at this commit (M-series, ReleaseFast, 80x24,
unlimited scrollback, 1 MB corpora, per-loop time with setup
subtracted):
ascii lines 1-70: 34.16 MB encode 92.8 ms decode 119.8 ms
ascii full-wrap: 16.01 MB encode 43.6 ms decode 56.2 ms
utf8: 4.33 MB encode 12.4 ms decode 19.0 ms
The ascii generator emits an unbroken stream of printable bytes, which
exercises terminal wrapping but produces only full-width rows. Add
line-min and line-max options that emit CR LF-terminated lines with a
uniformly distributed printable length so generated corpora can also
model shell-like output where most rows end well before the last
column. The default behavior is unchanged.
In the html formatter every page is formatted in a div. When the div
closes it causes a newline in the html rendering. In order to fix this a
newline is now removed whenever the div is closed (if there are any
newlines waiting to be rendered - as far as I could see in my testing
there was always one).
I thought of trying to add a test but could not think of a way to do so
without adding a massive blob of html into the file.
Before (orange was added by me to show where the div ends):
<img width="1278" height="586" alt="image"
src="https://github.com/user-attachments/assets/473ba28d-bec0-481f-9a89-a6a72c9a3657"
/>
After:
<img width="959" height="439" alt="image"
src="https://github.com/user-attachments/assets/27bae0d3-cc82-4dc8-aa72-c4a8b0f7d424"
/>
No AI was used in this pr.
### What
Comment-only cleanup in `src/termio/message.zig`:
- Fixes a subject-verb agreement error in the `Message` union's doc
comment: "the number of messages ... are also very few" -> "is also
very small"
- Capitalizes "pty" to "PTY" in several doc comments
### Why
Caught while reading through `message.zig`. No functional changes,
doc comments only.
Builds on #13544
This adds a new CONTINUATION record type that is sent before READY.
CONTINUATION contains the bytes (if any) that will bring a ground-state
VT state machine up to the same state.
This allows snapshotting a terminal instance that is, for example,
blocked waiting for a caller to complet an in-flight Kitty graphics
protocol send. In practice, I think this will be rare. But in theory, it
avoids a DoS-type attack.
The continuation state must be the MINIMAL set of bytes that will move
the virtual terminal state from a ground to non-ground state. The reason
it must be minimal is because any extra bytes can duplicate work into
the terminal that might already exist.
Builds on #13544
This adds a new CONTINUATION record type that is sent before READY.
CONTINUATION contains the bytes (if any) that will bring a ground-state
VT state machine up to the same state.
This allows snapshotting a terminal instance that is, for example,
blocked waiting for a caller to complet an in-flight Kitty graphics
protocol send. In practice, I think this will be rare. But in theory, it
avoids a DoS-type attack.
The continuation state must be the MINIMAL set of bytes that will move
the virtual terminal state from a ground to non-ground state. The reason
it must be minimal is because any extra bytes can duplicate work into the
terminal that might already exist.
- "the number of messages we send to the IO thread are also very few"
had a subject-verb agreement issue; reworded to "is also very small"
- Capitalized "pty" -> "PTY"
This adds opt-in continuation tracking to `terminal.Stream` that allows
any caller to call `writeContinuation` in order to get the minimum bytes
necessary from a grounded parser state to the identical state.
This enables reliable stream restart across serialization states, which
could be used for local restart, networked terminals, etc. For me, this
is used for multiplexers. :)
**LLM usage:** I wrote the continuation tracker myself, used a mix of
5.6+Fable to review it for me, applied their feedback directly. Only
place with predominantly AI code are tests, which I reviewed. Commit and
PR message written myself.
## Implementation
The implementation of this was really carefully done to avoid any
negative performance impact particularly when continuation tracking is
_off_.
The way this work is simple:
1. ESC is the only char that leaves the ground state and most ESC
sequences are short. So if we're in a non-ground state, we do a
backwards vectorized search to find the last `ESC` in the input slice.
If one doesn't exist, we assume we found it previously and store the
whole slice (rare, since ESC sequences are usually short like I said).
2. If we're in the ground state that means we only have a potential
incomplete UTF-8 codepoint, so we find the lead UTF-8 byte.
3. When writing, we normalize the suffix to drop things like BEL
commands that would've already been handled to avoid double-calling.
## Performance
No real impact.
Via `ghostty-bench +terminal-stream`.
Corpus | Main | PR w/ Tracking Off | PR w/ Tracking On
-- | -- | -- | --
Plain ASCII (256 MiB) | 175.4 ms | 175.8 ms | 175.5 ms
UTF-8 (32 MiB) | 268.4 ms | 270.0 ms | 269.9 ms
5% invalid UTF-8 (32 MiB) | 316.4 ms | 317.8 ms | 320.2 ms
CSI-heavy (32 MiB) | 145.0 ms | 146.3 ms | 145.9 ms
OSC (32 MiB) | 1621.0 ms | 1627.5 ms | 1638.3 ms
Kitty APC (128 MiB) | 95.5 ms | 96.9 ms | 96.3 ms
Mixed traffic (32 MiB) | 172.4 ms | 172.0 ms | 172.3 ms
Giant APC (128 MiB) | 38.3 ms | 38.3 ms | 40.8 ms
This adds opt-in continuation tracking to `terminal.Stream` that allows
any caller to call `writeContinuation` in order to get the minimum bytes
necessary from a grounded parser state to the identical state.
This enables reliable stream restart across serialization states, which
could be used for local restart, networked terminals, etc. For me, this
is used for multiplexers. :)
## Implementation
The implementation of this was really carefully done to avoid any
negative performance impact particularly when continuation tracking is
_off_.
The way this work is simple:
1. ESC is the only char that leaves the ground state and most
ESC sequences are short. So if we're in a non-ground state, we
do a backwards vectorized search to find the last `ESC` in the
input slice. If one doesn't exist, we assume we found it previously
and store the whole slice (rare, since ESC sequences are usually
short like I said).
2. If we're in the ground state that means we only have a potential
incomplete UTF-8 codepoint, so we find the lead UTF-8 byte.
3. When writing, we normalize the suffix to drop things like BEL
commands that would've already been handled to avoid
double-calling.
## Performance
Via `ghostty-bench +terminal-stream`
Corpus main tracking off tracking on
plain ASCII (256 MiB) 175.4ms 175.8ms 175.5ms
UTF-8 (32 MiB) 268.4ms 270.0ms 269.9ms
5% invalid UTF-8 (32 MiB) 316.4ms 317.8ms 320.2ms
CSI-heavy (32 MiB) 145.0ms 146.3ms 145.9ms
OSC (32 MiB) 1621.0ms 1627.5ms 1638.3ms
Kitty APC (128 MiB) 95.5ms 96.9ms 96.3ms
mixed traffic (32 MiB) 172.4ms 172.0ms 172.3ms
giant APC (128 MiB) 38.3ms 38.3ms 40.8ms
Every page is formatted in a div, and when the div closes it creates a newline in the html rendering.
In order to fix this a newline is now removed whenever the div is closed (if there are any newlines waiting to be rendered).
## Summary
- normalize circular-buffer metadata after every resize
- retain the oldest values when shrinking below the current length
- cover partial shrink, exact-length shrink, and empty-to-zero
boundaries
## Root cause
`resize` rotated live values to index zero before reallocating, but only
repaired `head` and `full` when capacity grew. Shrinking a partially
filled buffer could therefore leave `head` beyond the new allocation and
report a length greater than capacity. A later append could index
outside the resized storage.
## Validation
- `zig fmt --check src/datastruct/circ_buf.zig`
- `zig test src/circ_buf_test.zig --test-filter 'CircBuf resize'` using
a temporary import harness: 8 tests passed
Adds "Copy" and "Export to file" buttons to the Terminal IO inspector so
recorded VT events can be saved outside the app for sharing or analysis.
Found myself needing/wishing for this while I was debugging my tmux fork
with libghostty-vt.
Disclaimer: I haven't considered performance at all, so please lmk if
there are anything here you would like me to optimize.
~~Maybe I have too many exclamation marks, let me know if I should
metaphorically calm down.~~ Fixed now.
My wording is intentionally biased toward languages spoken less
(**edit**: not nearly as much anymore), but I specifically do not
disallow members who know each other regardless of language popularity.
Quoting myself from Discord[^convo]:
> people who know each other are more likely to have more similar tastes
or quirks in their language use by virtue of (perhaps subconsciously)
stealing off each other, and there's also the whole “eh it's good enough
i trust that you thought it through” thing that's more likely if you
know the other translators already
I don't believe it's *necessarily worse*, and if you have more than two
members then the issue greatly diminishes too, but I don't want people
to see this and go “oh no I need to get my translations in before
Ghostty 1.4 that releases while I'm sleeping tomorrow so I should ask my
bestie to help”, and to instead be willing to be more patient, at least
for a reasonable amount of time (which I consider to be ≤ 2 months).
[^convo]: @trag1c and I chatted about this prior to this PR in
`#maintainers` on the Ghostty Discord server. If you have access to that
channel, check out these links:
[1](https://discord.com/channels/1005603569187160125/1337443701403815999/1504241367357063188),
[2](https://discord.com/channels/1005603569187160125/1337443701403815999/1504465642361979021),
[3](https://discord.com/channels/1005603569187160125/1337443701403815999/1505255908676993135).
Improves the time to resize mixed content w/ wrapping 120x80, 10k lines
of scrollback by about ~6x.
I'm working on deferred resize in another branch, but it still follows
roughly the same logic so I instead decided to shift course and look at
the existing full-pagelist resize+reflow and found many places to
improve while keeping understanding.
The optimizations here match the general patterns of other recent
optimizations: cache some stuff, reuse some pages, bring in our
`page.Mask` helper and add vectorized ops. Nothing exotic we haven't
been doing recently.
Also note the Neovim project brought this up as a noticeable issue and I
believe this will help mitigate their issues until we get proper
deferred reflow in.
**LLM notes:** The optimizations were produced with Fable 5 using a
profile-driven approach (macOS `sample` plus disassembly-level
attribution of the hot loops at each step). I then requested each be
split into its own measurable commit, reviewed each in isolation, and
modified most of the commit messages. This PR message is hand-written.
The masked-compare scan that finds bulk-copyable cell runs still
processed one cell per iteration and remained the largest single
cost in a column reflow.
Scan whole groups of cells at a time using the group variants of the
Mask helper: a group that fully matches the run pattern (and, for
text runs, contains no Kitty virtual placeholder, via eqlAny)
extends the run by the whole group, and any mismatch falls through
to the scalar loop which finds the exact end of the run within it.
The group length comes from the shared simd.lanes helper where the
target has SIMD support and falls back to a plain unrolled group
elsewhere.
1.19x faster on ghostty-bench +terminal-resize --mode=cols (120x80
terminal, 10k-line scrollback, shrink/grow column reflow cycles).
Combined with the preceding reflow optimizations, resize with reflow
is 5.8x faster than before the series.
Finding the length of a bulk-copyable cell run evaluated the
field-wise bulkCopyable predicate plus a style compare per cell,
which compiles to a chain of extracts and branches and had become
the hottest loop in a column reflow.
Once the first cell passes the full predicate, a cell continues the
run iff it matches the first cell in content tag, style id, wide
property, and hyperlink flag, so the continuation test is now a
masked compare of the raw cell bits via the Mask helper, plus a
masked equality test against the Kitty virtual placeholder codepoint
for text runs (placeholders must set a row flag so they take the
slow path). This is slightly stricter than the predicate (a bg-color
cell no longer extends an unstyled text run), which only splits a
copy into multiple runs and remains correct.
1.21x faster on ghostty-bench +terminal-resize --mode=cols (120x80
terminal, 10k-line scrollback, shrink/grow column reflow cycles).
Two small extensions to the Mask helper, both motivated by the
reflow bulk run scan in the next commits.
fieldMask now accepts dot-separated field paths so a mask can cover
a nested field of a packed struct or packed union member, e.g.
"content.codepoint.data" covers exactly the codepoint bits of a cell
without its padding. Packed union members all share bit offset zero.
Mask gains eqlAny, the "any" counterpart to eql: it returns whether
any value in a group has masked fields equal to the expected
pattern. This supports run scans that must stop when a sentinel
value appears anywhere in a group, such as the Kitty virtual
placeholder codepoint which requires slow-path handling.
In `resizeCols`, stash the most recently finished source node
instead of destroying it, so we can recycle it without a bunch of
syscalls.
1.30x faster on ghostty-bench +terminal-resize --mode=cols (120x80
terminal, 10k-line scrollback, shrink/grow column reflow cycles),
with system time dropping from 34ms to 8ms per run.
reflowRow computed the capacity for prospective destination pages on
every source row via Capacity.adjust, which performs a full page
layout calculation to find the available grid space.
The result only depends on the source page, and reflow visits source pages
sequentially and never revisits one, so memoize the adjustment per
source page so we only do this once.
1.06x faster on ghostty-bench +terminal-resize --mode=cols (120x80
terminal, 10k-line scrollback, shrink/grow column reflow cycles).
Reflow copied every cell through a per-cell state machine
(writeCell) that dispatches on content tag, wide property, grapheme,
hyperlink, and style handling, and advances the destination cursor
one cell at a time. The vast majority of cells in practice are
narrow text or bg-color cells with no managed memory that share a
single style across long runs.
reflowRow now scans ahead for the run of such cells bounded by the
remaining space in the destination row, copies the run with a single
memcpy, and adjusts the style ref count once for the whole run via
useMultiple. Wide characters, spacers, graphemes, hyperlinks, Kitty
placeholders, and rows containing tracked pins all take the original
per-cell path, and a style set failure falls back to writeCell which
handles growing page capacity.
2.19x faster on ghostty-bench +terminal-resize --mode=cols (120x80
terminal, 10k-line scrollback, shrink/grow column reflow cycles).
Memoize the most recent style mapping and when there is a reuse
bump the ref with `use()`. This avoids a lookup (`addWithId`) on
every single styled cell.
1.25x faster on ghostty-bench +terminal-resize --mode=cols (120x80
terminal, 10k-line scrollback, shrink/grow column reflow cycles).
Reflow scanned the full tracked pin list for every source cell it
copied, twice per cell in the wide-character case, even though pins
are rare and at most a handful exist. Each check also went through
node.page(), which can restore a compressed page just to compare
pointers.
reflowRow now determines once per row whether any tracked pin is on
the source row and skips the per-cell pin scans entirely when there
is none, which is the overwhelmingly common case. The comparisons
use node identity instead of pages: a node owns exactly one page, so
they are equivalent, and this avoids the restore hazard.
1.09x faster on ghostty-bench +terminal-resize --mode=cols (120x80
terminal, 10k-line scrollback, shrink/grow column reflow cycles).
A state directory with the wrong permissions left the terminfo cache
failing with errors that named no path, so there was nothing to act on:
$ ghostty +ssh-cache --add=user@host
Error: Unable to add 'user@host' to cache. Error: error.AccessDenied
Every +ssh-cache failure now names its cache file, and +ssh no longer
swallows cache-related errors.
Error messages in these actions are also lowercased after the "Error: "
prefix and append the error with ": {t}" rather than a second "Error: ".
Ref:
https://github.com/ghostty-org/ghostty/issues/9393#issuecomment-5145799368
This adds the first version of a binary snapshot format for terminal
state.
Use cases: replay software (like asciinema), multiplexers (like zmx),
scrollback-saving on disk, etc.
The intention of the binary snapshot format is to be able to fully
encode and decode terminal state across mediums such as network and
disk. You can also encode partial terminal state (e.g. only one screen
or even one page of contents). Long term, the intention is to also
support streaming state while a live terminal is running, but this
initial PR focuses on the full snapshot first (with some design choices
to get to the streaming state in the future).
The format is documented in the Zig code, but I also did a
[Kaitai](https://kaitai.io/) descriptor and both the Zig and Kaitai spec
verify they can parse committed fixtures. This helps identify drift in
the format or encoder/decoders in any way since this must ultimately be
a fixed format.
> [!NOTE]
>
> **On reviewability:** this is a massive PR that I don't expect anyone
to reasonably review. I'm going through it line-by-line (again) but I
purposely extracted any changes that affect other parts of Ghostty out
to other already-merged PRs. This one is isolated purely to a package
that isn't called by any client software. **So the plan is if this rough
shape looks good I'll merge it and we'll iterate from there.**
> [!WARNING]
>
> **Experimental.** The format can and will change. And we may also
decide that binary snapshotting in this way isn't the right direction
altogether (although, I'm pretty confident it is). It'd be impossible to
get a single large perfect PR because it'd be even larger than this by
multiples. So instead, we'll iterate on main so long as this work is not
touching any production code, which it isn't!
## Example
Encode:
```zig
const terminal = @import("terminal/main.zig");
var file_buffer: [16 * 1024]u8 = undefined;
var file_writer = file.writer(io, &file_buffer);
try terminal.snapshot.encode(alloc, &file_writer.interface, &t);
try file_writer.interface.flush();
```
Decode a full terminal:
```zig
var t = try terminal.snapshot.decode(&reader, io, alloc);
defer t.deinit(alloc);
```
## Future
This PR purposely only supports a synchronous encode/decode. I wanted to
get the large groundwork in before iterating further. Some iterations in
the future:
* Kitty graphics
* Live terminal snapshotting
* PTY stream continuation records (so VT state machines can stay in
sync)
* Performance work (encoding and decoding, maybe size)
* Configurable limits to prevent DoS
* C API
* etc...
## Wire format
The "robustness principle" is a guiding principle: "be conservative in
what you do, be liberal in what you accept from others." Our encoders
have a lot of extra validation, our decoders massage invalid data into
reasonable defaults (e.g. invalid styles become unstyled text).
> [!IMPORTANT]
>
> **Version 1 has no compatibility promise.** We use version 1 in the
envelope header. We will absolutely break this format as needed as we
iterate and improve on it...
### Envelope
Every snapshot starts with a fixed ten-byte envelope:
| Offset | Size | Field |
| ---: | ---: | :--- |
| 0 | 8 | Magic: `GHOSTSNP` |
| 8 | 2 | Snapshot version: `1` |
### Record framing
After the envelope, records are concatenated back-to-back. Every record
has a fixed header:
| Offset | Size | Field |
| ---: | ---: | :--- |
| 0 | 2 | Record tag |
| 2 | 4 | Payload length |
| 6 | 4 | CRC32C |
| 10 | variable | Payload |
CRC32C covers the encoded tag, payload length, and payload. The
payload-length boundary prevents a malformed record decoder from
consuming bytes belonging to the next record.
The registered record tags are:
| Value | Tag | Purpose |
| ---: | :--- | :--- |
| 1 | `TERMINAL` | Terminal-wide state and declared screens |
| 2 | `SCREEN` | One screen's live state and active page manifest |
| 3 | `PAGE` | One self-contained set of rows and cells |
| 4 | `HISTORY` | One screen's historical page manifest |
| 5 | `READY` | Digest of the renderable prefix |
| 6 | `FINISH` | Digest of the complete snapshot |
To view the format of each record, read its corresponding
`terminal/snapshot/<type>.zig` file.
### Complete record sequence
```text
+----------------------------------------+
| Envelope |
+----------------------------------------+
| TERMINAL |
+----------------------------------------+
| SCREEN * terminal.screen_count |
| PAGE * each screen.page_count |
+----------------------------------------+
| READY |
+----------------------------------------+
| HISTORY * terminal.screen_count |
| PAGE * each history.page_count |
+----------------------------------------+
| FINISH |
+----------------------------------------+
| Optional containing-transport bytes |
+----------------------------------------+
```
`SCREEN` and `HISTORY` groups are routed by their encoded screen key and
may arrive in either key order.
## Checkpoints and validation
Each record has an independent CRC32C, but per-record checksums cannot
detect a valid record being reordered, omitted, or duplicated. `READY`
and `FINISH` therefore contain BLAKE3-256 digests over exact snapshot
prefixes:
- `READY` covers the envelope, `TERMINAL`, and all live `SCREEN`/`PAGE`
sequences. It does not include itself.
- `FINISH` covers that same prefix, the complete `READY` record, and all
`HISTORY`/`PAGE` sequences. It does not include itself.
This gives the format two useful integrity boundaries:
```text
envelope ... active pages | READY | history pages | FINISH
<------ renderable ------->
<------------- complete snapshot --------------->
```
## Performance
### Size, Compression Recommended
We intentionally use a simple grid over something like RLE (run-length
encoding). So every row contains exactly `columns` cells and each is
16-bytes! This is large! A 80x24, 10,000 line scrollback terminal
uncompressed would be ~13MB. However, with zstd level 1 compression that
goes down to 260K.
### Speed
We haven't benchmarked encoding or decoding speed yet. This PR focused
on getting a format in place. This will be heavily optimized later. I
suspect its probably pretty darn slow, actually.
## Kaitai Struct
I added a `snapshot.ksy` Kaita Struct spec that independently describes
the complete format. This is used by us for format validation but it can
also be used to programmatically generate parsers. For example, our test
fixture in the Kaita Struct web IDE decodes to:
<img width="523" height="779" alt="image"
src="https://github.com/user-attachments/assets/cd199c73-b6d6-4b35-8957-0cfe3d1a18f2"
/>
**AI Usage:** This work was done in concert with various models and
agents. Writing full encoders/decoders is tedious so it took a lot of
that way. A lot of review was done by AI (trying to find holes, issues,
inconsistencies). The actual binary protocol design and iteration was
done by me. This PR message was written by me.
Exclude annotated snapshot fixture hex files from typo checking. Their arbitrary binary byte sequences can otherwise be misidentified as misspelled words.