This adds opt-in continuation tracking to `terminal.Stream` that allows
any caller to call `writeContinuation` in order to get the minimum bytes
necessary from a grounded parser state to the identical state.
This enables reliable stream restart across serialization states, which
could be used for local restart, networked terminals, etc. For me, this
is used for multiplexers. :)
## Implementation
The implementation of this was really carefully done to avoid any
negative performance impact particularly when continuation tracking is
_off_.
The way this work is simple:
1. ESC is the only char that leaves the ground state and most
ESC sequences are short. So if we're in a non-ground state, we
do a backwards vectorized search to find the last `ESC` in the
input slice. If one doesn't exist, we assume we found it previously
and store the whole slice (rare, since ESC sequences are usually
short like I said).
2. If we're in the ground state that means we only have a potential
incomplete UTF-8 codepoint, so we find the lead UTF-8 byte.
3. When writing, we normalize the suffix to drop things like BEL
commands that would've already been handled to avoid
double-calling.
## Performance
Via `ghostty-bench +terminal-stream`
Corpus main tracking off tracking on
plain ASCII (256 MiB) 175.4ms 175.8ms 175.5ms
UTF-8 (32 MiB) 268.4ms 270.0ms 269.9ms
5% invalid UTF-8 (32 MiB) 316.4ms 317.8ms 320.2ms
CSI-heavy (32 MiB) 145.0ms 146.3ms 145.9ms
OSC (32 MiB) 1621.0ms 1627.5ms 1638.3ms
Kitty APC (128 MiB) 95.5ms 96.9ms 96.3ms
mixed traffic (32 MiB) 172.4ms 172.0ms 172.3ms
giant APC (128 MiB) 38.3ms 38.3ms 40.8ms
- benchmark: avoid buffers to avoid a memcpy
- build: keep frame pointers on macOS. There was some debug changes from
Zig 0.15 and this helps. Also, Apple actually requires/expects x29 to
always be a frame pointer.
- build/macos: force libSystem symbols instead of compiler-rt
- global: add InitOpts.tool so that ghostty-gen/bench can parse their
own actions in `+action`
- quirks: provide our own vectorized memset. see the comment for more
details why.
- synthetic: fix UB by accessing global.io before it was initialized
- terminal/hash_map: force inline for unique repr types. Zig 0.15
inlined and 0.16 doesn't, measured a huge slowdown in hyperlink
benchmarks.
- terminal: add explicit `@Vector` usage for storing a run of identical cells
as well as for scanning printable cells. This auto-vectorized in Zig
0.15 but not in Zig 0.16. This produces the same assembly.
- unicode: properties and LUT need power-of-two backing integer to avoid
bad LLVM codegen
This commit represents the majority of the work necessary to upgrade
Ghostty to use Zig 0.16.0.
Key parts:
* In addition to its previous responsibilities, the global state now
houses state for global I/O implementations and the process
environment. It is now also utilized in the main application along
with the C library. Where necessary, global state is isolated from key
parts of the implementation (e.g., in libghostty subsystems), and it's
expected that this list will grow.
* We currently manage our own C translation layer where necessary. In
these cases, cImport has been removed in favor of the new external
translate-c package. Due to fixes that have needed be made to properly
translate the dependencies that were swapped out, as mentioned, we
have had to backport fixes from the current translate-c package (and
the upstream Arocc dependency). We will host this ourselves until Zig
0.17.0 is released with these fixes.
* Where necessary (only a small number of cases), some stdlib code from
0.15.2 (and even from 0.17.0) has been taken, adopted, and vendored in
lib/compat.
Co-authored-by: Leah Amelia Chen <hi@pluie.me>
Kitty graphics payloads are dispatched in bulk, but finding each slice
boundary still examines every byte with a scalar loop. This leaves large
direct base64 image transmissions parser-bound.
Scan ordinary APC bytes using the vector width recommended for the compile
target. Keep the scalar scan both as the tail and as the full fallback when
the target has no recommended vector width. Test state-machine boundaries
against byte-at-a-time parsing.
A ReleaseFast APC parser benchmark over the same 64 MiB Kitty graphics
corpus, with 10 warmups and 30 measured runs, produced:
mean median
scalar 37.6 ms 32.6 ms
vectorized 22.3 ms 19.0 ms
Hyperfine reports the vectorized version as 1.69 times faster overall, with
the median runtime improving by approximately 42 percent.
APC payloads such as Kitty graphics images can be megabytes of base64
data, but every byte was dispatched individually: through the VT state
machine table, an apc_put action, the stream handler, the APC protocol
handler, and finally a per-byte ArrayList append in the Kitty command
parser. Five layers of dispatch per byte made large image transfers
far slower than they needed to be.
Add a bulk fast path alongside the existing CSI fast paths in
consumeUntilGround: scan the longest run of apc_put bytes (stopping
at any byte the parse table doesn't treat as APC payload: CAN, SUB,
ESC, and most C1 bytes exit or abort the string state, and 0xA0-0xFF
are ignored by it) and dispatch the run as a single new apc_put_slice
action. The APC handler identifies the protocol from the first few
bytes as before, then passes the remainder of each slice to the
protocol parser in bulk; the Kitty parser appends payload data with a
single appendSlice. Ignored/unknown APC sequences now drop each slice
in O(1) instead of per-byte dispatch.
The fast path is guarded the same way as the CSI fast paths: handlers
with a vtRaw hook (the inspector) keep receiving per-byte apc_put
actions, and the scalar next() path is unchanged.
Also add benchmark support: a `ghostty-gen +kitty` synthetic generator
emitting well-formed Kitty graphics transmit commands with 4 KiB
random base64 payloads (not valid image data; the corpus exercises
the parsing paths, not image decoding), and a `ghostty-bench
+apc-parser` benchmark that measures the stream -> APC -> Kitty parse
path without image decode/storage.
Benchmarks on a 64 MiB corpus (hyperfine, ReleaseFast, x86_64 Linux,
baseline is identical source with only the fast path disabled):
apc-parser: 1.061 s -> 43 ms (~25x)
terminal-stream (kitty): 1.163 s -> 72 ms (~16x)
terminal-stream (ascii): no change
The ascii case was verified with retired instruction counts (perf
stat, pinned to one core) since wall time on the test machine has
4-7 ms of noise: 988,030,458 vs 988,045,833 instructions (+0.0016%),
a fixed startup-size delta; the ground-state hot loop never reaches
the new branch.
Profiling terminal-stream on a 2.6 GB recording of real terminal
sessions showed ~5% of total time under writev, all of it log
output: the recording triggers ~120k warnings, dominated by a few
repeated messages ("unimplemented mode: 34", "invalid device
attributes command", "invalid C0 character") that some program in
the recorded session re-emitted on every frame or every prompt.
Each occurrence pays formatting plus a blocking write syscall,
and repeats add no diagnostic value beyond the first: the message
already includes the offending value.
These messages are emitted in response to input that the terminal
application controls, so a misbehaving or merely chatty program
can flood the log indefinitely. This adds a logUnsupportedOnce
helper that suppresses repeats per (call site, value): each site
tracks the distinct keys it has logged (the mode number, final
byte, or first parameter, depending on the site) in a small fixed
table of 16 u32 slots, 64 bytes per site. Real streams only ever
produce a handful of distinct unsupported values per site, so if a
table fills, new values are suppressed too; by then the log
already shows the problem class and unbounded distinct values
would flood it anyway. Slots are claimed with 32-bit atomics
(native on wasm32) and never change afterwards, so lookups are a
lock-free scan and the worst case race is a duplicate message.
The OSC 1 change-icon message moves from info to warn to match the
other unsupported-input messages the helper covers.
Measured with ghostty-bench terminal-stream (2.6 GB real-session
corpus, 120x80, M4 Max, ReleaseFast, hyperfine means of 5 runs,
stderr to /dev/null which undersells the cost of a real log sink):
| stream | before | after | change |
|----------------------------|---------|---------|--------|
| real 2.6 GB session corpus | 7.916 s | 7.674 s | +3.2% |
System time drops from 0.49 s to 0.22 s from the eliminated
writev calls.
Profiling terminal-stream on a 2.6 GB recording of real terminal
sessions showed ~7% of time in nextNonUtf8 self, and most calls
were for the structural bytes of CSI sequences: the '[' after ESC
and the single byte spent in the csi_entry state (a digit, private
marker, or final byte). Real streams contain tens of millions of
CSI sequences, and each paid two to three function calls just to
advance the parser through those states before the bulk parameter
loop could take over.
This lifts both transitions into the consumeUntilGround loop: the
"ESC [" prefix is matched directly, and the csi_entry byte is
handled by a shared csiEntryByte helper that both the loop and
nextNonUtf8 use (the logic previously lived only in nextNonUtf8).
A typical CSI sequence now parses entirely within
consumeUntilGround/consumeCsiParams without any per-byte calls.
Handlers with a vtRaw hook keep the general path since csiEntryByte
dispatches finals directly.
Measured with ghostty-bench terminal-stream (120x80, M4 Max,
ReleaseFast, hyperfine means of 5 runs). nextNonUtf8 self time
drops from ~7% to ~3% of the profile:
| stream | before | after | change |
|----------------------------|---------|---------|--------|
| real 2.6 GB session corpus | 9.097 s | 8.854 s | +2.7% |
| csi mix (SGR/CUP, 100 MB) | 695 ms | 674 ms | +3.1% |
After the CSI dispatch fast paths, profiling showed the remaining
escape-sequence cost was the per-byte plumbing itself: for every
parameter byte of a sequence like "ESC [ 38;2;10;20;30 m" the
stream re-entered nextNonUtf8, re-checked the parser state, and
re-dispatched through the fast-path switch, paying call and state
check overhead per digit.
consumeUntilGround now hands whole input slices to a new
consumeCsiParams loop whenever the parser is in the csi_param
state. It consumes runs of digits and separators with the parser
accumulator state held in locals, dispatches directly when it
reaches the final byte, and returns to the general path on the
first byte it doesn't understand (C0 controls, intermediates,
etc.), guaranteeing byte-for-byte identical semantics with the
per-byte fast path it hoists. Like the dispatch fast paths, this is
disabled at comptime for handlers that declare vtRaw so the
inspector continues to observe every action.
Throughput measured with ghostty-bench terminal-stream (full
terminal handler, 100 MB deterministic corpora, 120x80, M4 Max,
ReleaseFast, hyperfine means of 10 runs):
| stream | before | after | change |
|--------|--------|--------|--------|
| csi | 525 ms | 407 ms | +29% |
| sgr | 414 ms | 294 ms | +41% |
Combined with the previous commit, CSI-heavy streams are 1.5-1.7x
faster end to end than before this series.
Profiling escape-heavy streams showed the dominant remaining cost
was Parser.next: every byte routed through it copies a [3]?Action
return value that is ~240 bytes (the action union is sized by
osc.Command). A typical CSI sequence paid this twice: once for the
first byte after "ESC [" (csi_entry has no fast path, so even the
first parameter digit went through the table machine) and once for
the final byte that dispatches the sequence.
This extends the existing stream fast paths to cover both. The
csi_param fast path now handles final bytes (0x40-0x7E) by
finalizing parameters and dispatching the CSI directly via a new
csiDispatchFinal, which replicates the parser's csi_dispatch action
(MAX_PARAMS overflow drop, trailing parameter finalization, and the
colon-separator validation for non-'m' finals) without constructing
the action array. A new csi_entry fast path handles the byte right
after "ESC [": first parameter digit, empty first parameter,
private markers (0x3C-0x3F), and parameterless finals. Everything
else (C0 controls, intermediates, the csi_entry colon edge case)
still defers to the state machine.
Because these paths dispatch without going through Parser.next,
they would bypass a handler's vtRaw hook, so they are disabled at
comptime for handlers that declare one (the inspector). Those
handlers keep the exact previous behavior.
Throughput measured with ghostty-bench terminal-stream (full
terminal handler, 100 MB deterministic corpora, 120x80, M4 Max,
ReleaseFast, hyperfine means of 10 runs). The csi corpus is a
realistic mix of SGR, cursor movement, erases, and mode changes
with short text runs; sgr is a doom-fire-like stream of truecolor
SGRs and cell pairs:
| stream | before | after | change |
|--------|--------|--------|--------|
| csi | 618 ms | 525 ms | +18% |
| sgr | 486 ms | 414 ms | +17% |
#13209
After #13209 the IO pipeline delivers the parse thread's full
measured capacity, so IO throughput is now bound by VT processing.
Profiling `terminal-stream` on plain text showed ~85% of wall time
inside Terminal.print: every printable codepoint paid the full
per-character cost (right margin computation, grapheme clustering
checks, width lookup, wrap/insert mode checks, charset mapping,
per-cell style bookkeeping, dirty marking, cursor advance) even
though for typical bulk output every one of those answers is the
same for thousands of consecutive characters.
This adds a new print_slice stream action carrying a run of
printable codepoints, emitted whenever the SIMD ground-state path
decodes multiple codepoints at once, plus Terminal.printSlice which
processes such runs in batch. Since action dispatch is comptime,
delivering a slice through the existing vt handler interface has
the same codegen as a dedicated entry point; handlers that don't
care about batching can simply loop and treat each codepoint as a
print action.
printSlice hoists all run-invariant checks (status display, insert
and wraparound modes, charset state, hyperlink state) out of the
loop and then fills cells row by row. A single masked u64 compare
classifies each destination cell as "simple" (plain codepoint cell,
narrow, no hyperlink, style already matching the cursor); runs of
simple cells are written with a branch-free store loop, style-only
mismatches are handled inline with the same ref-counting printCell
does, and anything needing real cleanup (wide spacers, grapheme
data, hyperlinks) exits the fast path with the cursor positioned on
the offending cell so print() handles that one codepoint with full
generality. Dirty marking, previous_char, and cursor advancement
happen once per row instead of once per character.
The fast path handles both narrow and wide codepoints (CJK/emoji are
written as wide+spacer_tail pair fills, including spacer-head
handling at the right edge) and stays exact under grapheme
clustering (mode 2027): a codepoint only joins a run if it is width
1 or 2 and is a grapheme break from the previously written
codepoint, so print() would never have attached it to the previous
cell. The first codepoint of a batch defers to print() whenever the
previous cell could carry cluster state we can't cheaply reason
about (including a pending wrap, where print attaches to the
pending cell instead of wrapping).
Correctness is verified by a new differential fuzz test that runs
the same operations through per-codepoint print and randomly
chunked printSlice, comparing full screen dumps, cursor state, and
page integrity (style refcounts, grapheme maps) after every
operation, across wraps, margins, mode toggles, hyperlinks,
charsets, and wide/combining/ZWJ/RI/jamo codepoints.
Throughput measured with ghostty-bench terminal-stream (full
terminal handler, 100 MB deterministic corpora, 120x80, M4 Max,
ReleaseFast, hyperfine means of 10 runs; ~15ms process startup
included in all numbers):
| stream | before | after | change |
|---------------------------|--------|--------|--------|
| ascii (no newlines) | 784 ms | 138 ms | 5.7x |
| ascii lines | 833 ms | 198 ms | 4.2x |
| unicode mixed-script | 779 ms | 320 ms | 2.4x |
| CJK (all wide) | 424 ms | 126 ms | 3.4x |
| unicode, mode 2027 on | 807 ms | 367 ms | 2.2x |
| CJK, mode 2027 on | 495 ms | 198 ms | 2.5x |
Adds an OSC 72 parser following the kitty drag and drop protocol
specification. Parses metadata and payload into a Command.kitty_dnd_protocol
variant. Reassembly of chunked transfers and any action handling are
intentionally out of scope here; stream.zig logs the command as
unimplemented for now.
Includes a walkthrough document covering the design and each touched file.
Previously every file in the terminal package independently imported
build_options and ../lib/main.zig, then computed the same
lib_target constant. This was repetitive and meant each file needed
both imports just to get the target.
Introduce src/terminal/lib.zig which computes the target once and
re-exports the commonly used lib types (Enum, TaggedUnion, Struct,
String, checkGhosttyHEnum, structSizedFieldFits). All terminal
package files now import lib.zig and use lib.target instead of the
local lib_target constant, removing the per-file boilerplate.
Introduce a dedicated device_attributes.zig module that consolidates
all device attribute types and encoding logic. This moves
DeviceAttributeReq out of ansi.zig and adds structured response
types for DA1 (primary), DA2 (secondary), and DA3 (tertiary) with
self-encoding methods.
Primary DA uses a ConformanceLevel enum covering VT100-series
per-model values and VT200+ conformance levels, plus a Feature
enum with all known xterm DA1 attribute codes (132-col, printer,
sixel, color, clipboard, etc.) as a simple slice. Secondary DA
uses a DeviceType enum matching the xterm decTerminalID values.
Tertiary DA encodes the DECRPTUI unit ID as a u32 formatted to
8 hex digits.
This is preparatory work for exposing device attributes through
the stream_terminal Effects callback system.
Rename stream_readonly.zig to stream_terminal.zig and its exported
types from ReadonlyStream/ReadonlyHandler to TerminalStream. The
"readonly" name is now wrong since the handler now supports
settable effects callbacks. The new name better reflects that this
is a stream handler for updating terminal state.
Move MouseEvent and MouseFormat out of Terminal.zig and MouseShape out
of mouse_shape.zig into a new mouse.zig file. The types are named
without the Mouse prefix inside the module (Event, Format, Shape) and
re-exported with the prefix from terminal/main.zig for external use.
Update all call sites (mouse_encode.zig, surface_mouse.zig, stream.zig)
to import through terminal/main.zig or directly from mouse.zig. Remove
the now-unused mouse_shape.zig.
The terminal.Stream next/nextSlice functions can now no longer fail.
All prior failure modes were fully isolated in the handler `vt`
callbacks. As such, vt callbacks are now required to not return an error
and handle their own errors somehow.
Allowing streams to be fallible before was an incorrect design. It
caused problematic scenarios like in `nextSlice` early terminating
processing due to handler errors. This should not be possible.
There is no safe way to bubble up vt errors through the stream because
if nextSlice is called and multiple errors are returned, we can't
coalesce them. We could modify that to return a partial result but its
just more work for stream that is unnecessary. The handler can do all of
this.
This work was discovered due to cleanups to prepare for more C APIs.
Less errors make C APIs easier to implement! And, it helps clean up our
Zig, too.
A fuzz crash found that CSI g with a parameter that saturates to
u16 max (65535) causes @enumFromInt to panic when narrowing to
TabClear (enum(u8)). Use std.meta.intToEnum instead, which safely
returns an error for out-of-range values.
CSI @ (ICH) with an explicit parameter of 0 should be clamped to 1,
matching xterm behavior. Previously, a zero count reached
Terminal.insertBlanks which called clearCells with an empty slice,
triggering an out-of-bounds panic.
Fix the stream dispatch to clamp 0 to 1 via @max, and add a defensive
guard in insertBlanks for count == 0. Found by AFL++ stream fuzzer.
CSI ? W (cursor tabulation control) accessed input.params[0] without
first checking that params.len > 0, causing an index out-of-bounds
panic when the sequence had an intermediate but no parameters.
Add a params.len == 1 guard before accessing params[0].
Found by AFL++ fuzzing.
Implements parsing for OSC 3008, which allows terminal emulators to keep track of the stack of processes that have current control over the tty. The implementation mirrors existing `semantic_prompt.zig` architecture and natively maps UAPI definitions to Zig structures with lazy evaluation for optional metadata.
Fixes#10900
These can be unambiguously invoked in certain parser states, and as such
we need to handle them. In real world use they are extremely rare, hence
the branch hint. Without this, we get illegal behavior by trying to cast
the value to the 7-bit C0 enum.
This adds a new formatter that can be used with standard Zig `{f}`
formatting that emits any portion of the terminal screen as VT
sequences. In addition to simply styling, this can emit the entire
terminal/screen state such as cursor positions, active style, terminal
modes, etc.
To do this, I've extracted all formatting to a dedicated `formatter`
package within `terminal`. This handles all formatting types (currently
plaintext and VT formatting, but can imagine things like HTML in the
future). Presently, we have "formatting" split out across a variety of
places in Terminal, Screen, PageList, and Page. I didn't remove this
code yet but I intend to unify it all on formatter in the future.
This also doesn't expose this functionality in any user-facing way yet.
This PR just adds it to the ghostty-vt Zig module and unit tests it.
Ghostty app changes will come later.
**This also improves the readonly stream** to handle OSC color
operations for _setting_ but it doesn't emit any responses of course,
since its readonly.
This adds a new stream handler implementation that updates terminal
state in reaction to VT sequences, but doesn't perform any of the
actions that would require responses (e.g. queries).
This is exposed in two ways: first, as a standalone `ReadonlyStream` and
`ReadonlyHandler` type that contains all the implementation. Second, as
a convenience func on `Terminal` as `vtStream` and `vtHandler` which
return their respective types preconfigured to update the calling
terminal state.
This dramatically simplifies libghostty-vt usage from Zig (and will
eventually be exposed to C, too) since a Terminal on its own is ready to
go as a full VT parser and state machine without needing to build any
custom types!
There's a second big bonus here which is that our `stream_readonly.zig`
tests are true end-to-end tests for raw bytes to terminal state. This
will let us test a wider variety of situations more broadly. To start,
there are only a handful of tests implemented here.
**AI disclosure:** Amp wrote basically this whole thing, but I reviewed
it. https://ampcode.com/threads/T-3490efd2-1137-4112-96f6-4bf8a0141ff5