Commit Graph

9495 Commits

Author SHA1 Message Date
Mitchell Hashimoto
461562ca4f terminal: enable idle scrollback compression
Scrollback compression was available only through explicit PageList
calls, so normal renderer-backed surfaces never reclaimed cold history.

Debounce renderer wakeups with a one-shot timer and run bounded
compression steps only when the terminal lock is immediately available.
Timer state is isolated in the renderer and becomes idle once PageList
reports that the pass is complete.

Keep inspector reads representation-preserving by decoding compressed
pages into temporary copies. Update the scrollback configuration and
benchmark documentation to describe the production behavior and its
memory accounting.
2026-07-09 09:47:13 -07:00
Mitchell Hashimoto
f8c217e557 terminal: track pending scrollback compression
PageList previously kept only traversal position, so callers had no
central signal for deciding when an incremental compression pass should
run. Scheduling policy had to infer work from output and UI activity.

Track compression dirtiness alongside the PageList continuation state.
Growth preserves valid progress while marking work pending, and resize
or viewport transitions restart from the oldest page. A no-work
verification pass clears the state.

Expose Terminal helpers which report whether compression is required and
run compression against primary scrollback even while the alternate
screen is active. Unsupported retained-memory targets report no work.
2026-07-09 09:47:12 -07:00
Mitchell Hashimoto
0d015b2fce terminal: preserve compressed pages during search
Search previously used the normal page access boundary while formatting
history and checking soft-wrap boundaries. Inspecting compressed history
therefore restored its retained mapping and undid the memory
reclamation.

Format through preserved page snapshots and copy row counts and wrap
state into the sliding window's owned metadata. Overlap decisions reuse
the same snapshot, so compressed pages are decoded at most once per
append and remain compressed after matching.

Add a cross-page regression which searches compressed history and
verifies both source nodes retain their compressed representation.
2026-07-09 09:47:12 -07:00
Mitchell Hashimoto
e7969ed436 terminal: make PageList compression self-contained
Incremental compression previously exposed its traversal state to
callers, requiring them to coordinate the cursor with PageList lifetime
and topology changes. Read-only consumers also had to restore compressed
nodes to inspect their contents.

Move the continuation state into PageList and expose a single mode-based
compression entry point. Incremental passes restart safely across
mutations and verify a no-work pass before becoming idle, while full
passes leave incremental state fresh.

Add preserved page reads which decode compressed nodes into caller-owned
storage without changing their representation, and migrate the
scrollback compression benchmark to the new API.
2026-07-09 09:47:12 -07:00
Mitchell Hashimoto
70e788f066 terminal: add incremental scrollback compression
Cold-history compression previously required scanning every eligible
page in one call, which makes it unsuitable for an idle-time scheduler.
The inspector also restored compressed pages while traversing collapsed
entries, hiding the representation and undoing reclamation.

Add caller-owned serial state that resumes compression without retaining
node pointers. Each invocation inspects at most eight candidates and
attempts one resident page. List mutations restart safely, while
unsupported or historical viewports stop work. Keep a stateless
whole-history operation for measurement.

Expose metadata-only storage and memory accounting for diagnostics,
update the inspector to restore only expanded pages, and add an
incremental live scrollback benchmark. This remains disconnected from
production scheduling.
2026-07-09 09:47:12 -07:00
Mitchell Hashimoto
9156ada169 terminal: add cold-history compression pass
PageList could compress individual nodes but had no policy-level
operation for selecting pages that are safe to reclaim. The compressed
state therefore remained reachable only from tests.

Add a stateless pass that considers only complete history pages while
the viewport follows the active area. It gates work on supported
retained-mapping reclamation, reports attempts and retained bytes, and
leaves restoration lazy when a resize pulls compressed history back into
the active area.

Add a live scrollback-compression benchmark for measuring complete
PageList compression and restoration against saved VT corpora. The pass
still has no production callers, and ReleaseFast terminal-stream
comparisons remain within the existing throughput guardrail.
2026-07-09 09:47:12 -07:00
Mitchell Hashimoto
421fe8dabe terminal: integrate compressed pages into PageList
PageList nodes previously exposed Page directly, so introducing a
compressed representation would require every consumer to understand its
state and ownership.

Add resident and compressed node states behind a page access boundary.
Content access transparently recommits and restores retained mappings,
while metadata traversal stays compressed and lifecycle paths can
discard encoded contents without decoding. Compression borrows pool
memory for standard scratch and uses temporary aligned storage for
oversized pages.

Migrate terminal, rendering, search, formatting, and C API consumers to
the new boundary. Hot grow, scroll, and print paths reuse resolved
pages, with an explicit resident-only accessor where live cursor
pointers prove that the page cannot be compressed.

The compression entry point remains private with no production callers,
so normal terminal behavior and scrollback accounting are unchanged.
ReleaseFast terminal-stream comparisons across bulk output, scrolling,
redraw, and erase workloads remain within 2% of the parent revision.
2026-07-09 09:47:11 -07:00
Mitchell Hashimoto
f6fd4cb087 terminal: generalize page memory reclamation
PageList's virtual-memory helpers were tied to pool items even though
the underlying decommit and recommit operations also apply to retained
page mappings.

Move the helpers into terminal/mem.zig and express their different
failure contracts as generic modes. Zero mode preserves the existing
pool invariant and fallbacks, while strict mode only succeeds when the
operating system accepts reclamation and avoids touching memory that
will be restored.

PageList now uses the shared zero mode for its page pool. The strict
path is tested in isolation and remains unused, so this does not enable
page compression yet.
2026-07-09 09:47:11 -07:00
Mitchell Hashimoto
ebc3ffd222 terminal: add compressed page representation
The standalone LZ4 codec had no representation for terminal page
ownership or metadata, so PageList integration would otherwise need to
reconstruct every Page field independently.

Add compress.Page, which embeds the complete terminal Page while
retaining its original virtual mapping and owns only an exact-sized
encoded block. Compression is kept only when the encoded state is
strictly smaller, and scratch output is capped at that profitability
boundary so a future PageList caller can borrow a standard pool item.

Extend the page-compression benchmark with a store mode that measures
the encoded copy, allocation, bounded retention, and eviction path.
Nothing uses compression from PageList yet; this remains isolated
groundwork.
2026-07-09 09:47:11 -07:00
Mitchell Hashimoto
c62c159841 terminal: add standalone LZ4 block codec
Scrollback compression needs a codec that can be used from libghostty-vt
without pulling in libc, and we need to measure it before integrating it
with terminal page ownership.

This adds an allocation-free raw LZ4 block codec in scalar Zig.  Callers
provide the input, output, and fixed-size scratch table. The decoder
uses an exact-size output contract so page metadata mismatches fail
cleanly. Compatibility vectors, boundary cases, random round trips, and
fuzz coverage exercise the block format.

Also adds a page-compression benchmark that operates on reusable raw
page corpora. Compression and decompression have separate modes with
setup outside the timed region, plus a ratio report and no-op baseline.
Nothing uses compression in the terminal yet; this is the isolated codec
and measurement groundwork.
2026-07-09 09:47:11 -07:00
Mitchell Hashimoto
896aca4990 terminal: return free-listed page memory to the OS
The PageList page pool never returns memory to the OS: destroyed pages
are zeroed and free-listed until the surface exits. Any operation that
shrinks the page count (clearing scrollback, pruning churn, resets,
reflow) therefore retains its high-water RSS forever. 

Clearing a full scrollback keeps all of it resident, which at the default 10MB
scrollback-limit is 10MB per terminal of memory that can never be
used again for anything else.

Lots of memes on the internet about this, and it turns out operating
systems give us an answer for this (both Linux and macOS at least),
so let's do it kids.

Pool items can't be individually freed since they live inside arena
chunks, but they are page-aligned and page-multiple sized, so we can
decommit them while they sit in the free list. Our page-aligned
allocation pays off, again!

On Linux, madvise(MADV_DONTNEED) reclaims the pages immediately and
guarantees zero-fill on the next touch, which also lets us skip the
zeroing memset entirely, making destroy cheaper.

On macOS, we zero in place and mark the item MADV_FREE_REUSABLE, which
removes it from the process footprint immediately. Reuse is paired
with MADV_FREE_REUSE when the pool hands a buffer back out so that
footprint accounting stays correct. The zero invariant required by
page reuse holds either way: reusable page contents are either
preserved (our zeroes) or reclaimed and zero-filled by the kernel.

Other platforms and test builds keep the existing memset behavior.

## LLM Notes

Fable 5 found the retention behavior while re-analyzing scrollback
memory, wrote the change and tests, and verified the madvise semantics
empirically with memory probes on macOS and a real Linux kernel, plus
before/after throughput benchmarks on both. I reviewed the analysis,
the diff, rewrote the code to be more idiomatic Zig, and wrote this
commit message you're reading.
2026-07-07 20:52:44 -07:00
Mitchell Hashimoto
c442eb4b30 terminal: document the pool reuse zeroing invariant
PageList skips zeroing pooled page buffers in release builds, relying
on the OS page allocator handing out zeroed pages and destroyNodeExt
zeroing buffers before returning them to the pool. There is a hidden
exception: std.heap.MemoryPool writes its free list node into the
first pointer-size bytes of a free-listed buffer, so a reused buffer
is not fully zero. This is only safe because the page rows array is
laid out at offset 0, a page always has at least one row, and initBuf
fully rewrites every row, overwriting the stale free list pointer.

None of that was written down or checked anywhere, so a future layout
reorder (or a zero-row page) would corrupt pages in release builds
only, in a way that depends on pool reuse patterns. This adds a
comptime assert that a Row covers at least a pointer, a runtime assert
that pages always have at least one row, and comments tying the
invariant together at the layout, initBuf, and pool reuse sites.

Also fixes stale doc comments: deinit referenced a clonePool function
that no longer exists, and Screen tests referenced increaseCapacity by
its old adjustCapacity name.
2026-07-07 20:05:09 -07:00
Mitchell Hashimoto
e44f5cb0fa terminal: guard RefCountedSet lookups against zero-capacity sets
Layout.init(0) is an explicitly supported special case that produces a
valid zero-capacity set with a zero-size table. But lookupContext had
no guard for it: probing computes `table[hash & 0]` and reads whatever
memory follows the set in its backing buffer, treating those bytes as
an item ID which is then used to index the (also zero-size) items
array, an out-of-bounds read reaching arbitrarily far past the set.

This has never fired in practice because no production page carries a
zero-capacity set today, and where one could occur the adjacent bytes
happen to be zero (which reads as an empty bucket and returns null).
Page.exactRowCapacity legitimately produces zero capacities for pages
without styled or hyperlinked cells though, so any page compaction
work makes this reachable with nonzero adjacent memory: in a Page
layout the styles set can be followed by the grapheme bitmap, which
is initialized to all ones.

Lookups on a zero-capacity set now return null without touching the
table. This also covers add, which looks up before inserting and
already handles the zero capacity correctly after that point by
returning OutOfMemory. All other entry points assert on valid IDs,
which a zero-capacity set cannot have.
2026-07-07 20:05:09 -07:00
Mitchell Hashimoto
8307349ec5 terminal: fix increaseCapacity growth from zero-capacity dimensions
PageList.increaseCapacity grows a capacity dimension by doubling it.
If the dimension is zero, doubling "succeeds" without growing: the
page is reallocated and recloned with an identical capacity, violating
the documented guarantee that we always increase by at least one unit.
Every unbounded retry site (startHyperlink, cursorSetHyperlink, the
reflow probes, insertLines/deleteLines) then loops forever reallocating
a page per iteration, and the single-retry sites (styles, graphemes)
fail their retry and silently drop data.

No production page has a zero dimension today, which is why this has
never fired: standard capacities are nonzero and doubling keeps them
nonzero. But exactRowCapacity legitimately returns zero for dimensions
with no content (a compacted plain text page has zero styles, grapheme,
string, and hyperlink capacity), so any compaction work makes this
reachable.

Growth from zero now jumps straight to the standard default for the
dimension rather than doubling. The default is what every standard
page starts with, so single-retry callers are guaranteed enough room
for their pending allocation, whereas doubling from a minimum unit
could still come up short (a single grapheme can need multiple chunks,
and a style set below capacity 3 cannot store anything).
2026-07-07 20:05:09 -07:00
Mitchell Hashimoto
16e4b5e98f terminal: track page ownership explicitly on pagelist nodes
PageList decided whether a page's backing memory belongs to the memory
pool or the heap by comparing its length against std_size. This was super
error-prone and the source of many bugs historically. We locked it down
but its bothered me and has gotten in the way of another feature I've
wanted to do: memory compaction.

First, this commit records the ownership explicitly on each node and uses it
everywhere ownership was previously inferred from size. 

Second, createPage gains an exact_size option that forces an exact-size 
heap allocation even when the layout would fit a pool item.

Third, compact() now uses it to shrink any page to its minimum size,
including standard pages. Nothing calls compact [YET!] but this is going
to be the key to compressing scrollback history.

This is groundwork for a ton of memory savings. Coming soon.
2026-07-07 15:33:47 -07:00
Mitchell Hashimoto
b953bb3463 terminal: fix bitmap allocator chunk region sizing
BitmapAllocator.layout takes a capacity in bytes, but sized its chunk
region as `aligned_cap * chunk_size`, reserving chunk_size times more
memory than the bitmaps can ever address. As a result, the grapheme
region of every standard page reserved 128 KiB with only 8 KiB
reachable, and the string region 64 KiB with only 2 KiB reachable.
About ~180 KiB of waste in every 576 KiB page.

Results for a standard page:

| region | before | after |
|---|---|---|
| grapheme allocator | 131,136 B | 8,256 B |
| string allocator | 65,544 B | 2,056 B |
| page total | 589,824 B (576 KiB) | 409,600 B (400 KiB) |

30% less memory for every standard page in every terminal,
including the preheated pages in the PageList pool.
2026-07-07 13:42:48 -07:00
Mitchell Hashimoto
3ff6d08fad lib-vt: report OSC and Kitty color queries (#13239)
Hooks up responses to OSC and Kitty color queries if `write_pty` is set
for libghostty.

Also found a memory leak: the OSC parser now releases color operation
request lists during reset.
2026-07-07 13:10:16 -07:00
Mitchell Hashimoto
14c8298830 terminal: report OSC color queries in lib-vt
libghostty-vt already tracked OSC color state but ignored color queries in the standalone stream handler. This meant embedders that installed write_pty still received no response for OSC 4/10/11/12 or Kitty OSC 21 queries.

Resolve the current terminal colors through shared Terminal helpers and encode replies through the write_pty effect. Xterm queries use the fixed 16-bit rgb form, preserve the request terminator, and fall back from cursor to foreground when no cursor color is set. Kitty color queries now report supported terminal-backed keys and return empty values for unset dynamic colors.

Add RGB wire encoders and tests covering the stream handler and C API. The OSC parser now releases color operation request lists during reset, fixing an allocation leak exposed by multi-query OSC color tests.
2026-07-07 13:03:19 -07:00
Elias Andualem
20a1bfa5fd fix: pass RGB color inputs by pointer 2026-07-08 02:57:17 +08:00
Mitchell Hashimoto
bb0ac4c723 termio: don't bridge pty reads while the parser is idle
Fixes a frame time regression reported with fortio's `fps -fire`
benchmark (fortio.org/terminal/fps): frame times nearly doubled at
typical grid sizes after #13209 (the pipelined pty reader), and were
more than 5x worse at small grids.

## The problem

The `fps -fire` program follows a request/response pattern: write a
frame, end it with a cursor position query (CSI 6n), and block until
the reply arrives before starting the next frame. 

The gather stage treats any burst of 1 KiB or more as a saturated
stream. When its spin retries come up dry, it sleeps in a 1ms poll
expecting the writer's next refill so it can publish fewer, larger
batches. But a frame-synced writer will never refill here: it is
blocked waiting for a reply to a query that is sitting inside the
very batch the gather stage is holding back. The poll always sleeps
its full timeout, adding ~1.2ms to every round trip. Ouch!

## The fix

Sleeping for a refill gap is only free while the parse stage is busy,
since the wait hides behind parse time. Once the parser is idle,
every additional microsecond spent bridging is added straight to
output latency. So:

1. When the spin retries exhaust and the parse stage is idle,
   deliver the batch immediately instead of polling.

2. When the parse stage is busy, use a `pipe2` pipe to allow the parser
   to notify the gather thread it is idle. In the middle of the `poll`
   loop and sleep, we can get interrupted immediately and deliver
   the batch.

The pipe is only written while the gather stage is actively polling, so an
interactive terminal never pays the syscalls, and a saturated stream
never idles the parser, preserving full batching and throughput for
bulk output. Win, win, win!

## Benchmarks

| workload | pre-#13209 | before | after |
|---|---|---|---|
| fps -fire 80x24 | 0.262 ms / 3674 fps | 1.435 ms / 694 fps | 0.234 ms / 4106 fps |
| fps -fire 160x45 | 1.012 ms / 975 fps | 1.823 ms / 545 fps | 0.701 ms / 1405 fps |
| fps -fire 284x68 | 2.443 ms / 407 fps | 2.391 ms / 414 fps | 1.351 ms / 732 fps |
| cat 19.3 MB | 0.204 s (95 MB/s) | 0.086 s (224 MB/s) | 0.082 s (236 MB/s) |

## LLM Notes

Bisect script written by hand and found the offending commit. Fable 5
found the likely cause. I manually came up with the proposed solution and
wrote it out. Fable helped benchmark for me to verify my assumptions 
conceptually and in the real world.
2026-07-07 11:35:50 -07:00
Ēriks Remess
bed47168ca termio: bound POSIX read-ahead on non-Darwin
The pipelined POSIX pty reader uses multiple large gather buffers to avoid
stalling on Darwin, where pty reads are capped around 1KiB. On Linux this can
let bulk terminal producers run too far ahead of terminal parsing/rendering.

Frame-style terminal apps such as DOOM-fire can then report very high producer
FPS while Ghostty displays stale frames from the buffered stream.

Keep the Darwin-tuned 4 x 64KiB pipeline, but reduce non-Darwin read-ahead to
2 x 8KiB so the pty can apply backpressure before multiple frames are queued.
2026-07-07 20:24:06 +03:00
Mitchell Hashimoto
77190bd023 terminal: handful of scroll region optimizations
This optimizes scrolling inside a scroll region (DECSTBM). 

## The changes

1. **Stop creating scrollback for top-anchored regions on screens that don't 
  retain scrollback.** `index()` routed any full-width region with `top == 0` 
  through `cursorScrollAbove()`, which pushes the scrolled-out row into 
  scrollback. Every scroll paid `PageList.grow()` plus amortized page pruning, 
  which includes a 512 KB `memset` each time a page is recycled. These now 
  use the in-place region scroll instead. CSI S gets the same routing fix. 
  **Result: 1.05x-1.49x on the bottom-anchored region workloads, 1.25x on 
  alt-screen full-screen scrolling.**

2. **Add a specialized `Screen.cursorScrollRegionUp()` for the region scroll 
   hot path.** The previous fast path (`PageList.eraseRowBounded`) paid 
   per-scroll bookkeeping that exceeded the actual row work.
   The new function is built around the invariant that the cursor sits on the 
   bottom row of a full-width region.
   **Result: 1.23x-1.24x on the top-anchored region workloads.**

## Benchmarks

| workload | region (80 rows) | before | after | change |
|---|---|---|---|---|
| scrolling (control) | primary screen, no region | 237 ms | 235 ms | 1.0x |
| scrolling_bottom_region | alt, rows 1-79 | 243 ms | 231 ms | 1.05x |
| scrolling_bottom_small_region | alt, rows 1-40 | 311 ms | 208 ms | 1.49x |
| scrolling_top_region | alt, rows 2-80 | 283 ms | 229 ms | 1.23x |
| scrolling_top_small_region | alt, rows 40-80 | 258 ms | 208 ms | 1.24x |
| alt screen full-screen scrolling | alt, no region | 288 ms | 230 ms | 1.25x |

## LLM Notes

Assisted by Fable 5: it diagnosed the vtebench gap, wrote the benchmark 
harness payloads, profiled, and proposed hot paths. I manually wrote the
hot path replacements and had it judge my work.
2026-07-06 21:13:42 -07:00
Mitchell Hashimoto
cabbdee32b Fix adjust-cursor-height regression (#13225)
The `cursor-height` metric (and corresponding `adjust-`) was introduced
by #3062 but the sprite face rework in #7732 accidentally removed the
logic that it relied on. I've moved the logic to live inside of the
sprite face itself (which, I think, was my plan while writing the
rework, I just forgot to actually do it lol), and added a test for the
height metric being respected and the re-centering being performed
correctly.

This problem came to my attention because of #13221, which didn't go
about doing the fix the right way, but did make me realize that it was a
problem in the first place (since I had thought that I had already
implemented this logic when doing the rework!)

### Verification

https://github.com/user-attachments/assets/4074690b-846e-442d-8ec0-91a34042f6eb
2026-07-06 20:18:11 -07:00
Mitchell Hashimoto
446f80f4ed terminal: render state update optimizations (~2.7x to ~11x less lock hold)
This optimizes `RenderState.update`, the per-frame call that snapshots
terminal state for the renderer and is the main reason the renderer
thread holds the terminal lock. 

Lock hold time is reduced ~2.7x to ~11x depending on the frame.

## The changes

1. iterate page chunks instead of rows in `update`
2. classify cells with masked vector compares. 
3. split the update into `beginUpdate`/`endUpdate` phases. There's a 
   lot to be gained by accumulating data with the lock held and then
   processing it out of the lock.
4. generalize the masked-compare scans into `page.Mask`. This is just
   a really common pattern we're doing now and it yields a ton of great
   value. Its error prone so lets make it a tested helper.

## Benchmarks

Measured with the new `ghostty-bench +screen-clone` modes (`render`,
`render-locked`, `render-clean`, `render-partial`), 120x80 terminal, M4
Max, macOS 26, ReleaseFast, hyperfine means of 10+ runs, per-update
times derived from fixed-count update loops with process startup
subtracted. "Lock held" is the time the terminal lock must be held per
update; "before" held the lock for the entire update.

| scenario | before (lock held) | after (lock held) | after (total) | lock change |
|----------|--------------------|-------------------|---------------|-------------|
| clean frame (nothing dirty) | 202 ns | 19 ns | 19 ns | 10.9x |
| partial frame (1 dirty row) | 290 ns | 54 ns | 54 ns | 5.4x |
| full rebuild, lightly styled | 6.9 µs | 2.5 µs | 3.0 µs | 2.7x |
| full rebuild, fully styled | 9.3 µs | 2.4 µs | 8.0 µs | 3.8x |
| full rebuild, fully styled, 250x150 | 49.9 µs | 9.4 µs | 31.6 µs | 5.3x |
| full rebuild, plain text | 1.9 µs | 1.9 µs | 1.9 µs | 1.0x (memcpy floor) |

The clean and partial cases are the steady-state frame costs (cursor
blink, mouse movement, typing). The full-rebuild cases are the contended
ones: colored scrolling output (build logs, htop, vim) moves the
viewport pin every frame, forcing a full rebuild exactly when the IO
thread is busiest, so that row of the table is where lock contention
actually hurts. Plain text was already at the memcpy floor and is
unchanged.

## LLM Notes

This work was driven by Fable 5: benchmarks, optimizations, the property
test, and the measurements above. I reviewed every line, simplified the
design in a few places (API naming, the Mask helper shape), and re-ran
the verifications myself.
2026-07-06 19:57:04 -07:00
Mitchell Hashimoto
b5053153f4 terminal: log unsupported-input messages once per distinct value
Profiling terminal-stream on a 2.6 GB recording of real terminal
sessions showed ~5% of total time under writev, all of it log
output: the recording triggers ~120k warnings, dominated by a few
repeated messages ("unimplemented mode: 34", "invalid device
attributes command", "invalid C0 character") that some program in
the recorded session re-emitted on every frame or every prompt.
Each occurrence pays formatting plus a blocking write syscall,
and repeats add no diagnostic value beyond the first: the message
already includes the offending value.

These messages are emitted in response to input that the terminal
application controls, so a misbehaving or merely chatty program
can flood the log indefinitely. This adds a logUnsupportedOnce
helper that suppresses repeats per (call site, value): each site
tracks the distinct keys it has logged (the mode number, final
byte, or first parameter, depending on the site) in a small fixed
table of 16 u32 slots, 64 bytes per site. Real streams only ever
produce a handful of distinct unsupported values per site, so if a
table fills, new values are suppressed too; by then the log
already shows the problem class and unbounded distinct values
would flood it anyway. Slots are claimed with 32-bit atomics
(native on wasm32) and never change afterwards, so lookups are a
lock-free scan and the worst case race is a duplicate message.

The OSC 1 change-icon message moves from info to warn to match the
other unsupported-input messages the helper covers.

Measured with ghostty-bench terminal-stream (2.6 GB real-session
corpus, 120x80, M4 Max, ReleaseFast, hyperfine means of 5 runs,
stderr to /dev/null which undersells the cost of a real log sink):

| stream                     | before  | after   | change |
|----------------------------|---------|---------|--------|
| real 2.6 GB session corpus | 7.916 s | 7.674 s | +3.2%  |

System time drops from 0.49 s to 0.22 s from the eliminated
writev calls.
2026-07-06 13:51:35 -07:00
Mitchell Hashimoto
8d663a76e9 terminal: release style refs per run instead of per cell in clearCells
clearCells released the style reference of every styled cell
individually: an array index, a ref decrement, and a liveness
check per cell. Styled cells overwhelmingly come in runs sharing
the same style id (a colored status bar, a highlighted region, a
full row painted in one color), so most of that work is repeated
bookkeeping on the same style entry.

This groups consecutive cells with the same style id and releases
each run with a single releaseMultiple call. Rows with alternating
styles degrade to the same per-cell cost as before; uniform rows,
the common case, do one ref-count update per run. The
releaseMultiple assertion that the ref count is at least the run
length holds by construction since every cell in the run held a
reference.

Measured with ghostty-bench terminal-stream (120x80, M4 Max,
ReleaseFast, hyperfine means of 5 runs). The erase corpus paints a
full screen of styled rows and erases it with ED 2 in a loop,
which is the pattern full-screen TUIs produce on clear/redraw:

| stream                     | before  | after   | change |
|----------------------------|---------|---------|--------|
| real 2.6 GB session corpus | 8.055 s | 7.965 s | +1.1%  |
| styled paint + ED 2 (100 MB) | 260 ms | 123 ms | 2.1x   |
2026-07-06 13:51:35 -07:00
Mitchell Hashimoto
cb2d785871 terminal: fill style-only cell runs in bulk in printSliceFill
Profiling terminal-stream on a 2.6 GB recording of real terminal
sessions showed printSliceFill as the single largest item (~25% of
total time), and disassembly showed the time split across three
scalar loops: the run-eligibility scan over codepoints, the
simple-cell check that guards the branch-free fill, and the general
path that fixes up style ref counts one cell at a time. The store
loop itself was already auto-vectorized by LLVM, but the two scans
are early-exit search loops that LLVM does not vectorize, and the
general path turns out to be the common case in real traffic:
styled text constantly overwrites cells holding a different style
(TUI redraws, scrolling colored output), so every such cell failed
the simple check and paid a release/use pair.

Three changes, which only pay off together (vectorizing the scans
without the bulk path makes mismatch-heavy rows slower because the
wider check re-runs for every cell the general path consumes):

The run-eligibility scan handles the narrow class, codepoints in
[0x10, 0xFF], eight lanes at a time. The simple-cell check compares
four masked cells per iteration. And a new bulk path handles runs
of cells that differ from the expected simple cell only by style
id: one vector scan finds the extent of the uniformly-styled run,
the ref counts are fixed with a single releaseMultiple/useMultiple
pair, and the cells are filled with the same branch-free store
loop as the simple case. Cells with graphemes, hyperlinks, or wide
content still fall back to print().

Measured with ghostty-bench terminal-stream (120x80, M4 Max,
ReleaseFast, hyperfine means of 5 runs). The redraw corpus is a
full-screen 80-row styled repaint whose span color rotates every
frame, so every cell is overwritten with a different style:

| stream                     | before  | after   | change |
|----------------------------|---------|---------|--------|
| real 2.6 GB session corpus | 8.826 s | 7.955 s | +11%   |
| TUI redraw (100 MB)        | 348 ms  | 287 ms  | +21%   |
2026-07-06 13:51:35 -07:00
Mitchell Hashimoto
300f42c7a9 terminal: handle CSI entry bytes inline in consumeUntilGround
Profiling terminal-stream on a 2.6 GB recording of real terminal
sessions showed ~7% of time in nextNonUtf8 self, and most calls
were for the structural bytes of CSI sequences: the '[' after ESC
and the single byte spent in the csi_entry state (a digit, private
marker, or final byte). Real streams contain tens of millions of
CSI sequences, and each paid two to three function calls just to
advance the parser through those states before the bulk parameter
loop could take over.

This lifts both transitions into the consumeUntilGround loop: the
"ESC [" prefix is matched directly, and the csi_entry byte is
handled by a shared csiEntryByte helper that both the loop and
nextNonUtf8 use (the logic previously lived only in nextNonUtf8).
A typical CSI sequence now parses entirely within
consumeUntilGround/consumeCsiParams without any per-byte calls.
Handlers with a vtRaw hook keep the general path since csiEntryByte
dispatches finals directly.

Measured with ghostty-bench terminal-stream (120x80, M4 Max,
ReleaseFast, hyperfine means of 5 runs). nextNonUtf8 self time
drops from ~7% to ~3% of the profile:

| stream                     | before  | after   | change |
|----------------------------|---------|---------|--------|
| real 2.6 GB session corpus | 9.097 s | 8.854 s | +2.7%  |
| csi mix (SGR/CUP, 100 MB)  | 695 ms  | 674 ms  | +3.1%  |
2026-07-06 13:51:35 -07:00
Mitchell Hashimoto
083d9709be terminal: decode ASCII inline in the SIMD scan for ESC
Profiling terminal-stream on a 2.6 GB recording of real terminal
sessions showed ~9% of total time inside the UTF-8 decode stage,
and most of it was not the decode itself: real streams contain an
escape sequence every ~18 bytes, so utf8DecodeUntilControlSeq is
called on short printable runs, and each call paid simdutf setup
plus its scalar rewind_and_convert_with_errors tail (which handles
the last partial SIMD block of every conversion) for only a
handful of bytes. The scalar tail alone accounted for ~3.4% of
total time.

Terminal input is also overwhelmingly ASCII, for which UTF-8 to
UTF-32 "decoding" is just widening each byte to 32 bits. This
fuses the two passes: while scanning each chunk for ESC we also
check for bytes >= 0x80 and widen pure-ASCII chunks straight into
the output vector via PromoteTo, never touching simdutf. The first
non-ASCII byte hands the remainder of the run (up to the next ESC)
to the existing simdutf-based path, so non-ASCII text takes
exactly the same code as before. Inputs shorter than one vector
are handled by a scalar byte loop that likewise skips simdutf for
ASCII.

The widening store needs a dedicated path for the HWY_SCALAR
fallback target (compiled on targets without guaranteed SIMD, e.g.
arm-linux-androideabi): its single-lane vectors cannot be halved
so the one lane is widened directly.

The new differential fuzz test verifies the SIMD implementation
still matches the scalar reference exactly. Measured with
ghostty-bench terminal-stream (2.6 GB real-session corpus, 87%
printable ASCII / 5.5% ESC / 5.6% UTF-8, 120x80, M4 Max,
ReleaseFast, hyperfine means):

| stream            | before          | after           | change |
|-------------------|-----------------|-----------------|--------|
| real 2.6 GB corpus | 9.582 s (272 MB/s) | 9.090 s (287 MB/s) | +5.4% |
2026-07-06 13:51:35 -07:00
Qwerasd
e8f3f6c438 font/sprite: add regression test for cursor-height metric 2026-07-06 15:48:23 -04:00
Qwerasd
dac341cad5 font/sprite: make cursor height respect adjust-cursor-height
This was a regression caused by the sprite face rework (#7732), I'm
surprised it went unnoticed for so long.
2026-07-06 15:48:23 -04:00
Mitchell Hashimoto
cb4c49fbf2 terminal: scalar UTF-8 decode consumes partial sequences cut off by ESC
The scalar fallback of utf8DecodeUntilControlSeq (used when SIMD is
disabled, e.g. wasm builds) treated a valid-so-far but incomplete
UTF-8 sequence at the end of its decode region as pending input in
all cases: it stopped without consuming the bytes so a future chunk
could complete the sequence. That is correct when the region ends
at the end of the input, but the region can also be bounded by an
ESC byte. In that case the sequence can never be completed (the
next byte is already known to be ESC), and the SIMD implementation,
via simdutf, replaces the ill-formed prefix with U+FFFD and
consumes up to the ESC. The two implementations disagreed on both
the consumed count and the decoded output for inputs like
"\xC2\x1B[0m".

The divergence is invisible at the stream level (the pending bytes
take the scalar nextUtf8 path which also emits a replacement
character once it sees the ESC) but it means the scalar decoder is
not a faithful reference for the SIMD one.

This makes the scalar decoder treat a partial sequence bounded by
an ESC as a maximal subpart per Unicode 3-7: one U+FFFD, consumed
through the end of the region. Truncation at the true end of input
still leaves the bytes pending. Also adds a differential fuzz test
that runs 10k random mixtures of ASCII, escapes, controls, and
valid/invalid UTF-8 through both implementations and requires
identical results, which is what caught this.
2026-07-06 12:25:58 -07:00
Mitchell Hashimoto
4c5b1d5d52 bench: terminal-stream reads 64KiB chunks to match the IO thread
The terminal-stream benchmark fed the stream in 4KiB chunks while
the real IO thread reads from the pty into 64KiB buffers (see
buffer_capacity in termio/Exec.zig) and hands those to the stream
whole. Chunk size affects measurement in two ways: it determines
how often the stream crosses a chunk boundary (partial UTF-8
sequences, escape sequences split mid-parse) and how many read
syscalls the harness itself performs (a 2.6 GB corpus is ~636k
pread calls at 4KiB versus ~40k at 64KiB).

This bumps the benchmark read and dispatch buffers to 64KiB so the
stream is exercised with realistic chunk sizes. Measured with
ghostty-bench terminal-stream on a 2.6 GB recording of real
terminal sessions (120x80, M4 Max, ReleaseFast, hyperfine means):

| harness      | time            | throughput |
|--------------|-----------------|------------|
| 4KiB chunks  | 9.651 s ± 0.013 | 270 MB/s   |
| 64KiB chunks | 9.582 s ± 0.101 | 272 MB/s   |

The stream itself is barely chunk-size sensitive (most time is in
parsing and terminal state updates), but the harness now matches
what the IO thread actually does, and later commits are measured
against this configuration.
2026-07-06 12:25:20 -07:00
Mitchell Hashimoto
cee35cabf6 terminal: skip style map update when SGR leaves style unchanged
Profiling the csi benchmark showed ~20% of time in the style
ref-counted set (hash, probe, release/use churn) driven by
manualStyleUpdate, which runs after every SGR attribute even when
the attribute didn't actually change the cursor style. Real
programs re-assert the same style constantly (per span, per line,
or on every refresh of a mostly static screen), so a large share of
these updates are no-ops.

Screen.setAttribute already snapshots the old style to restore it
on failure, so this compares the style after applying the attribute
and returns early when it's unchanged: the current style ID is
already correct and no release/lookup/use is needed.

The tradeoff is one extra Style.eql on every style-changing
attribute. Measured with ghostty-bench terminal-stream (full
terminal handler, 100 MB deterministic corpora, 120x80, M4 Max,
ReleaseFast, hyperfine means of 10 runs) across corpora with
different repeated style rates (the csi/sgr corpora draw random
colors from a palette so nearly every SGR changes the style, which
is the worst case for this change; the redraw corpora model TUI
refreshes that re-assert the current style for 70% / 95% of SGRs):

| stream              | before | after  | change |
|---------------------|--------|--------|--------|
| redraw (95% same)   | 277 ms | 260 ms | +7%    |
| redraw (70% same)   | 302 ms | 291 ms | +4%    |
| csi (~0% same)      | 407 ms | 414 ms | -2%    |
| sgr (~0% same)      | 295 ms | 303 ms | -3%    |

Real-world SGR traffic is far closer to the redraw corpora than to
the adversarial random-color ones, so this trades a small worst
case regression for a solid win on the common pattern.
2026-07-06 08:51:52 -07:00
Mitchell Hashimoto
253e4f9c3c terminal: bulk-parse CSI parameter bytes at the slice level
After the CSI dispatch fast paths, profiling showed the remaining
escape-sequence cost was the per-byte plumbing itself: for every
parameter byte of a sequence like "ESC [ 38;2;10;20;30 m" the
stream re-entered nextNonUtf8, re-checked the parser state, and
re-dispatched through the fast-path switch, paying call and state
check overhead per digit.

consumeUntilGround now hands whole input slices to a new
consumeCsiParams loop whenever the parser is in the csi_param
state. It consumes runs of digits and separators with the parser
accumulator state held in locals, dispatches directly when it
reaches the final byte, and returns to the general path on the
first byte it doesn't understand (C0 controls, intermediates,
etc.), guaranteeing byte-for-byte identical semantics with the
per-byte fast path it hoists. Like the dispatch fast paths, this is
disabled at comptime for handlers that declare vtRaw so the
inspector continues to observe every action.

Throughput measured with ghostty-bench terminal-stream (full
terminal handler, 100 MB deterministic corpora, 120x80, M4 Max,
ReleaseFast, hyperfine means of 10 runs):

| stream | before | after  | change |
|--------|--------|--------|--------|
| csi    | 525 ms | 407 ms | +29%   |
| sgr    | 414 ms | 294 ms | +41%   |

Combined with the previous commit, CSI-heavy streams are 1.5-1.7x
faster end to end than before this series.
2026-07-06 08:51:51 -07:00
Mitchell Hashimoto
1a88f3622b terminal: dispatch CSI finals directly from stream fast paths
Profiling escape-heavy streams showed the dominant remaining cost
was Parser.next: every byte routed through it copies a [3]?Action
return value that is ~240 bytes (the action union is sized by
osc.Command). A typical CSI sequence paid this twice: once for the
first byte after "ESC [" (csi_entry has no fast path, so even the
first parameter digit went through the table machine) and once for
the final byte that dispatches the sequence.

This extends the existing stream fast paths to cover both. The
csi_param fast path now handles final bytes (0x40-0x7E) by
finalizing parameters and dispatching the CSI directly via a new
csiDispatchFinal, which replicates the parser's csi_dispatch action
(MAX_PARAMS overflow drop, trailing parameter finalization, and the
colon-separator validation for non-'m' finals) without constructing
the action array. A new csi_entry fast path handles the byte right
after "ESC [": first parameter digit, empty first parameter,
private markers (0x3C-0x3F), and parameterless finals. Everything
else (C0 controls, intermediates, the csi_entry colon edge case)
still defers to the state machine.

Because these paths dispatch without going through Parser.next,
they would bypass a handler's vtRaw hook, so they are disabled at
comptime for handlers that declare one (the inspector). Those
handlers keep the exact previous behavior.

Throughput measured with ghostty-bench terminal-stream (full
terminal handler, 100 MB deterministic corpora, 120x80, M4 Max,
ReleaseFast, hyperfine means of 10 runs). The csi corpus is a
realistic mix of SGR, cursor movement, erases, and mode changes
with short text runs; sgr is a doom-fire-like stream of truecolor
SGRs and cell pairs:

| stream | before | after  | change |
|--------|--------|--------|--------|
| csi    | 618 ms | 525 ms | +18%   |
| sgr    | 486 ms | 414 ms | +17%   |
2026-07-06 08:51:51 -07:00
Mitchell Hashimoto
47e26df60f terminal: batch printed codepoint runs into direct row fills
#13209

After #13209 the IO pipeline delivers the parse thread's full
measured capacity, so IO throughput is now bound by VT processing.
Profiling `terminal-stream` on plain text showed ~85% of wall time
inside Terminal.print: every printable codepoint paid the full
per-character cost (right margin computation, grapheme clustering
checks, width lookup, wrap/insert mode checks, charset mapping,
per-cell style bookkeeping, dirty marking, cursor advance) even
though for typical bulk output every one of those answers is the
same for thousands of consecutive characters.

This adds a new print_slice stream action carrying a run of
printable codepoints, emitted whenever the SIMD ground-state path
decodes multiple codepoints at once, plus Terminal.printSlice which
processes such runs in batch. Since action dispatch is comptime,
delivering a slice through the existing vt handler interface has
the same codegen as a dedicated entry point; handlers that don't
care about batching can simply loop and treat each codepoint as a
print action.

printSlice hoists all run-invariant checks (status display, insert
and wraparound modes, charset state, hyperlink state) out of the
loop and then fills cells row by row. A single masked u64 compare
classifies each destination cell as "simple" (plain codepoint cell,
narrow, no hyperlink, style already matching the cursor); runs of
simple cells are written with a branch-free store loop, style-only
mismatches are handled inline with the same ref-counting printCell
does, and anything needing real cleanup (wide spacers, grapheme
data, hyperlinks) exits the fast path with the cursor positioned on
the offending cell so print() handles that one codepoint with full
generality. Dirty marking, previous_char, and cursor advancement
happen once per row instead of once per character.

The fast path handles both narrow and wide codepoints (CJK/emoji are
written as wide+spacer_tail pair fills, including spacer-head
handling at the right edge) and stays exact under grapheme
clustering (mode 2027): a codepoint only joins a run if it is width
1 or 2 and is a grapheme break from the previously written
codepoint, so print() would never have attached it to the previous
cell. The first codepoint of a batch defers to print() whenever the
previous cell could carry cluster state we can't cheaply reason
about (including a pending wrap, where print attaches to the
pending cell instead of wrapping).

Correctness is verified by a new differential fuzz test that runs
the same operations through per-codepoint print and randomly
chunked printSlice, comparing full screen dumps, cursor state, and
page integrity (style refcounts, grapheme maps) after every
operation, across wraps, margins, mode toggles, hyperlinks,
charsets, and wide/combining/ZWJ/RI/jamo codepoints.

Throughput measured with ghostty-bench terminal-stream (full
terminal handler, 100 MB deterministic corpora, 120x80, M4 Max,
ReleaseFast, hyperfine means of 10 runs; ~15ms process startup
included in all numbers):

| stream                    | before | after  | change |
|---------------------------|--------|--------|--------|
| ascii (no newlines)       | 784 ms | 138 ms | 5.7x   |
| ascii lines               | 833 ms | 198 ms | 4.2x   |
| unicode mixed-script      | 779 ms | 320 ms | 2.4x   |
| CJK (all wide)            | 424 ms | 126 ms | 3.4x   |
| unicode, mode 2027 on     | 807 ms | 367 ms | 2.2x   |
| CJK, mode 2027 on         | 495 ms | 198 ms | 2.5x   |
2026-07-06 08:51:51 -07:00
Mitchell Hashimoto
258de36d15 benchmark: terminal-stream uses the full terminal handler
The terminal-stream benchmark previously used a simplified handler
that handles print actions and drops everything else. That was
originally intended to isolate parse and print throughput, but it
understates the cost of escape-heavy streams: no terminal state is
updated for CSI/OSC/ESC sequences, and because actions are
dispatched at comptime, the unhandled action arms are eliminated
entirely, so the benchmark measures dispatch code that doesn't
exist in the real app.

This switches the benchmark to the full readonly terminal stream
handler (terminal.TerminalStream). Every escape sequence now
updates real terminal state (styles, cursor movement, erases,
modes, etc.), closely mirroring the work the real IO thread does
per byte. This is the handler used to measure the VT throughput
changes in the following commits.

Parser-in-isolation measurement remains covered by the separate
terminal-parser and osc-parser benchmarks, and print throughput is
identical under both handlers since printing flows into the same
Terminal call either way.
2026-07-06 06:44:04 -07:00
Mitchell Hashimoto
2f0e6659dd termio: pipeline pty reads to overlap parsing with draining
This replaces the single-threaded pty read loop on posix systems with
a two-stage pipeline: a new `io-gather` thread drains the pty into a
small ring of large buffers while the existing `io-reader` thread
parses the previous batch concurrently.

The motivating discovery was actually found by Fable, but the
resulting code was hand-written and hand-verified (in addition to
model-verified as an extra check): on macOS the kernel tty output queue
caps every read on the pty master at about 1 KiB regardless of the
read buffer size. Instrumenting a pty with a 64 KB buffer while
streaming a 6.49 MB file produced 6,337 reads where every read was exactly 
1024 bytes. 

This immediately made me realize two things about the old loop that we've 
had since like 2023 which called processOutput once per read: all per-call 
overhead (terminal lock, render wakeup, cursor timestamp) was paid per
kilobyte of bulk output, and the child could never run more than 1 KiB
ahead of us, so while we parsed the child (e.g. `cat`) sat blocked on a full
kernel queue instead of producing. 

In 2023, I justified this architecture by saying "reads are generally small"
but I didn't understand then that reads are generally small because the
kernel makes them small even if there is a lot of data.

To preserve latency for the more typical
small-reads-that-are-actually-small, sub-1 KB payloads deliver on the first
EAGAIN with no added latency.

End-to-end throughput was measured by timing `cat file > /dev/ttysN`
against a fresh app instance (M4 Max, macOS 26, ReleaseFast, medians
of repeated interleaved A/B runs):

| stream                  | before       | after         | change  |
|-------------------------|--------------|---------------|---------|
| ascii.txt (6.5 MB)      | 91-92 MB/s   | 114-123 MB/s  | +25-30% |
| unicode.txt (8 MB)      | 116-117 MB/s | 180-183 MB/s  | +55%    |
| DOOM-Fire-Zig           | 530 fps      | 770 fps       | +45%    |

The pipeline now delivers the parse stage's full measured capacity
(the parse thread is pegged while gather spends ~33% of a core, so
any IO throughput improvements are now fully parser-bound).

**Linux note:** This needs to be verified on Linux. I think broadly
architecture is better and should never be worse. But its possible
some of the magic constants need to be tuned differently. Would love
more testing there.
2026-07-05 21:49:35 -07:00
Mitchell Hashimoto
63e75e86c2 lib-vt: many more color utility APIs (#13206)
Embedders that render theme editors, palette pickers, or custom settings
UI need to use the same color semantics as Ghostty.

This moves the shared parsing paths into terminal/color and exposes them
through libghostty-vt. Config color and palette parsing now delegate to
the same helpers, so CLI/config behavior and the C ABI stay in lockstep.

From C:

    GhosttyColorRgb rgb;
    ghostty_color_parse("ForestGreen", 11, &rgb);

    uint8_t index;
    ghostty_color_parse_palette_entry(
        "0x10=#282c34", 12, &index, &rgb);

    const GhosttyColorX11Entry* names =
        ghostty_color_x11_names();

The exported color API is:

    ghostty_color_parse
    ghostty_color_parse_x11
    ghostty_color_parse_palette_entry
    ghostty_color_palette_default
    ghostty_color_palette_generate
    ghostty_color_luminance
    ghostty_color_perceived_luminance
    ghostty_color_contrast
    ghostty_color_x11_names
    ghostty_color_x11_name_count

The X11 name table is parsed once at comptime into null-terminated
entries in rgb.txt order. The existing case-insensitive map keeps the
same behavior for RGB.parse and +list-colors, while bindings can walk a
static table without allocations.

This doesn't add any more binary size since all of this was already used
by terminal internals.
2026-07-05 13:15:17 -07:00
Mitchell Hashimoto
2970e9a2a8 lib-vt: many more color utility APIs
Embedders that render theme editors, palette pickers, or custom
settings UI need to use the same color semantics as Ghostty.

This moves the shared parsing paths into terminal/color and exposes them
through libghostty-vt. Config color and palette parsing now delegate to
the same helpers, so CLI/config behavior and the C ABI stay in lockstep.

From C:

    GhosttyColorRgb rgb;
    ghostty_color_parse("ForestGreen", 11, &rgb);

    uint8_t index;
    ghostty_color_parse_palette_entry(
        "0x10=#282c34", 12, &index, &rgb);

    const GhosttyColorX11Entry* names =
        ghostty_color_x11_names();

The exported color API is:

    ghostty_color_parse
    ghostty_color_parse_x11
    ghostty_color_parse_palette_entry
    ghostty_color_palette_default
    ghostty_color_palette_generate
    ghostty_color_luminance
    ghostty_color_perceived_luminance
    ghostty_color_contrast
    ghostty_color_x11_names
    ghostty_color_x11_name_count

The X11 name table is parsed once at comptime into null-terminated
entries in rgb.txt order. The existing case-insensitive map keeps the
same behavior for RGB.parse and +list-colors, while bindings can walk a
static table without allocations.
2026-07-05 13:11:11 -07:00
Mitchell Hashimoto
4a7cabc4fe lib-vt: add color scheme report encoder (#13192)
Add a shared encoder for CSI ? 997 ; Ps n color scheme reports and use
it for both CSI ? 996 n replies and unsolicited Termio reports. Export
the same encoder through the libghostty-vt C API with docs and an
example.

This is a really light API, arguably easy for consumers to hardcode, but
it didn't match the rest of our style in the libghostty API so we should
expose it.
2026-07-05 12:53:13 -07:00
Jeffrey C. Ollie
715ef6c154 fix: set max window clamp to current monitor size (#13171)
This PR fixes #7984. The issue was that GTK would clamp the window
itself based on the display it was opened on. We fix this by computing
the size based on the current display and then implicitly setting the
window size instead of relying on GTK to do it.

Claude Code w/ Opus 4.7 was used to investigate, fix and explain some of
the Ghostty architecture to me.
2026-07-05 14:12:49 -05:00
Mitchell Hashimoto
f00e906949 lib-vt: add color scheme report encoder
Add a shared encoder for CSI ? 997 ; Ps n color scheme reports and use 
it for both CSI ? 996 n replies and unsolicited Termio reports. Export the 
same encoder through the libghostty-vt C API with docs and an example.

This is a really light API, arguably easy for consumers to hardcode,
but it didn't match the rest of our style in the libghostty API so we 
should expose it.

Example: GHOSTTY_COLOR_SCHEME_DARK encodes to ESC [ ? 997 ; 1 n,
while GHOSTTY_COLOR_SCHEME_LIGHT encodes to ESC [ ? 997 ; 2 n.
2026-07-05 10:51:39 -07:00
yak
004c88e41e fix: set max window clamp to current monitor size 2026-07-05 10:41:06 -04:00
Mitchell Hashimoto
65e61282a6 lib-vt: add unicode grapheme width API
Embedders that render text outside the terminal grid need to predict
how many cells text will occupy once it is written to the terminal.
The existing codepoint width API exposes the table used by print, but
that is not enough for mode 2027 grapheme clustering: VS15/VS16, ZWJ
sequences, skin tone modifiers, and other continuation codepoints can
change the width of the whole cluster.

This exposes a single segment-and-measure API so callers use Ghostty
segmentation and width folding together:

    uint8_t width;
    size_t n = ghostty_unicode_grapheme_width(cps, len, &width);

From the Zig module:

    const vt = @import("ghostty-vt");
    const result = vt.unicode.graphemeWidth(u21, cps);

Callers loop until their string is consumed. The API is intentionally
not streaming: input must contain a complete first cluster or the
logical string end, so chunked readers should keep buffering when the
function consumes all available codepoints and more may arrive.

The terminal hot path now shares the width-decision func with the
API, the helper is inline and preserves the old branch structure. So
this doesn't change codegen at all.
2026-07-04 14:03:42 -07:00
Jeffrey C. Ollie
6887509035 Kitty dnd parser (#13029)
(#12852)
I opened a discussion to work on the new kitty dnd protocol and
implementing it for Ghostty. I was told to work on the parser but not to
hook up any actions to it yet. So, that's what I did! Largely based the
format on kitty_clipboard_protocol.zig, and used Claude Opus 4.8 (Claude
Code) for writing tests and some structural guidance early on. Would
love to get started on adding actions as well!
2026-07-04 02:21:56 -05:00
Mitchell Hashimoto
cca51729a1 lib-vt: add scroll-to-row viewport scrolling (#13179)
This adds a GHOSTTY_SCROLL_VIEWPORT_ROW tag with a `size_t row` member
in the value union. The row is an absolute offset from the top of the
scrollable area, clamped to the active area, in the same row space as
the scrollbar offset so thumb positions round-trip cleanly:

    ghostty_terminal_scroll_viewport(term,
        (GhosttyTerminalScrollViewport){
            .tag = GHOSTTY_SCROLL_VIEWPORT_ROW,
            .value = {.row = 42},
        });

The tag is appended to the existing enum and the union fits within the
reserved padding, so this is ABI compatible.

This also corrects the docs on GHOSTTY_TERMINAL_DATA_SCROLLBAR: the
getter is amortized O(1) (total is maintained incrementally, the offset
is cached), not "expensive". Since there is intentionally no change
callback, the docs now bless polling per frame or per write batch and
diffing, which is what Ghostty's own renderer does.

Motivation: Embedders building native scrollbars can already read scroll
state via GHOSTTY_TERMINAL_DATA_SCROLLBAR, but the write side only
exposed top/bottom/delta scrolling. Mapping a scrollbar thumb drag to an
absolute position required reading the current offset and computing a
delta, which is two calls that must be sequenced atomically by the
caller.

The core already supports absolute positioning and the macOS app uses it
for scroller drags via the scroll_to_row keybinding; this exposes the
same operation through the libghostty C API.
2026-07-03 21:29:38 -07:00
Mitchell Hashimoto
3a2e28329c lib-vt: add scroll-to-row viewport scrolling
This adds a GHOSTTY_SCROLL_VIEWPORT_ROW tag with a `size_t row` member
in the value union. The row is an absolute offset from the top of the
scrollable area, clamped to the active area, in the same row space as
the scrollbar offset so thumb positions round-trip cleanly:

    ghostty_terminal_scroll_viewport(term,
        (GhosttyTerminalScrollViewport){
            .tag = GHOSTTY_SCROLL_VIEWPORT_ROW,
            .value = {.row = 42},
        });

The tag is appended to the existing enum and the union fits within the
reserved padding, so this is ABI compatible.

This also corrects the docs on GHOSTTY_TERMINAL_DATA_SCROLLBAR: the
getter is amortized O(1) (total is maintained incrementally, the offset
is cached), not "expensive". Since there is intentionally no change
callback, the docs now bless polling per frame or per write batch and
diffing, which is what Ghostty's own renderer does.

Motivation: Embedders building native scrollbars can already read scroll state via
GHOSTTY_TERMINAL_DATA_SCROLLBAR, but the write side only exposed
top/bottom/delta scrolling. Mapping a scrollbar thumb drag to an
absolute position required reading the current offset and computing a
delta, which is two calls that must be sequenced atomically by the
caller. 

The core already supports absolute positioning and the macOS
app uses it for scroller drags via the scroll_to_row keybinding; this
exposes the same operation through the libghostty C API.
2026-07-03 21:26:59 -07:00
Mitchell Hashimoto
fc5a727729 lib-vt: add unicode codepoint width API
Embedders that render text outside the terminal grid need to predict
how many cells a codepoint will occupy once it is written to the
terminal. The immediate motivation is IME preedit overlay rendering:
measuring preedit text with font APIs (e.g. CoreText advances) can
disagree with the terminal's unicode table on ambiguous-width CJK and
emoji, causing the overlay to visibly jump when the composed text
commits and reflows through the real grid layout.

This exposes the exact width table the terminal print path already
uses, so overlays are column-accurate by construction. From C:

    uint8_t w = ghostty_unicode_codepoint_width(0x4E00); // 2

And from the Zig module:

    const vt = @import("ghostty-vt");
    const w = vt.unicode.codepointWidth(0x4E00); // 2

The function is total over its input: 0 for zero-width codepoints
(controls, combining marks, default-ignorables, surrogates), 2 for
wide codepoints (East Asian Wide/Fullwidth, regional indicators,
clamped at 2), and 1 for everything else, including invalid values
beyond U+10FFFF.

Perf: uses the LUT lookup we use for the main core terminal

Binary size: the width table was already linked into libghostty-vt
via the print path, so this adds only the exported wrapper.
2026-07-03 21:20:54 -07:00