mirror of
https://github.com/ghostty-org/ghostty.git
synced 2026-07-31 04:39:01 +00:00
This optimizes `RenderState.update`, the per-frame call that snapshots terminal state for the renderer and is the main reason the renderer thread holds the terminal lock. Lock hold time is reduced ~2.7x to ~11x depending on the frame. ## The changes 1. iterate page chunks instead of rows in `update` 2. classify cells with masked vector compares. 3. split the update into `beginUpdate`/`endUpdate` phases. There's a lot to be gained by accumulating data with the lock held and then processing it out of the lock. 4. generalize the masked-compare scans into `page.Mask`. This is just a really common pattern we're doing now and it yields a ton of great value. Its error prone so lets make it a tested helper. ## Benchmarks Measured with the new `ghostty-bench +screen-clone` modes (`render`, `render-locked`, `render-clean`, `render-partial`), 120x80 terminal, M4 Max, macOS 26, ReleaseFast, hyperfine means of 10+ runs, per-update times derived from fixed-count update loops with process startup subtracted. "Lock held" is the time the terminal lock must be held per update; "before" held the lock for the entire update. | scenario | before (lock held) | after (lock held) | after (total) | lock change | |----------|--------------------|-------------------|---------------|-------------| | clean frame (nothing dirty) | 202 ns | 19 ns | 19 ns | 10.9x | | partial frame (1 dirty row) | 290 ns | 54 ns | 54 ns | 5.4x | | full rebuild, lightly styled | 6.9 µs | 2.5 µs | 3.0 µs | 2.7x | | full rebuild, fully styled | 9.3 µs | 2.4 µs | 8.0 µs | 3.8x | | full rebuild, fully styled, 250x150 | 49.9 µs | 9.4 µs | 31.6 µs | 5.3x | | full rebuild, plain text | 1.9 µs | 1.9 µs | 1.9 µs | 1.0x (memcpy floor) | The clean and partial cases are the steady-state frame costs (cursor blink, mouse movement, typing). The full-rebuild cases are the contended ones: colored scrolling output (build logs, htop, vim) moves the viewport pin every frame, forcing a full rebuild exactly when the IO thread is busiest, so that row of the table is where lock contention actually hurts. Plain text was already at the memcpy floor and is unchanged. ## LLM Notes This work was driven by Fable 5: benchmarks, optimizations, the property test, and the measurements above. I reviewed every line, simplified the design in a few places (API naming, the Mask helper shape), and re-ran the verifications myself.