Commit Graph

2 Commits

Author SHA1 Message Date
Mitchell Hashimoto
e01e75bbb2 terminal: don't touch me! keep the page pool free list unobtrusive
This replaces the `std.heap.MemoryPool` used for page buffers with
a custom pool called `UntouchedPool`. This keeps its free list in a side
array and never reads/writes items until `create()`. This means that
demand-driven allocations (like mmaped pages) don't incur physical costs
until they're actually used.

The standard `std.heap.MemoryPool` uses an intrusive linked list for
its items which causes every item to be touched, which forces a full
page-in of memory.

It turns out we also had a lot of assertions and logic to work around
this in various ways (size of rows, asserting we overwrite the free
list entry, etc.) that we can now remove because of this.

For an 80x24 terminal on macOS (16 KB pages):

| Per terminal                 | Before   | After    |
|------------------------------|----------|----------|
| Page-list memory dirty       | 128 KiB  | 48 KiB   |
| Process phys_footprint delta | 143 KiB  | 62 KiB   |
| Page-list virtual size       | 2208 KiB | 1600 KiB |

The remaining 48 KB is the active page, because we sprinkle metadata
around the page which forces every page to be paged in. I'm going to
follow this up with some work trying to move all our metadata to the
front of the page so we only page one in until the rest is needed,
but not sure if its achievable.

Micro-benchmarks on the pool show that its twice the speed (slower) to
create/free due to the side list, but in an actual `+terminal-stream`
benchmark churning through pages, there is no measurable difference. I
think its a good trade.
2026-09-02 20:52:40 -07:00
Mitchell Hashimoto
492c26067c libghostty: use custom memory pool for Wasm
A custom memory pool for Wasm that grows by exactly one item size per
growth and shares the pool across the entire Wasm-module instead of 
per-terminal.

Some background on why `std.heap.MemoryPool` is considered harmful for 
WebAssembly:

First, the std.heap.MemoryPool grows 1.5x at each growth point. The backing 
allocator for that is usually a GPA which is the BrkAllocator for wasm. 
This grows by power-of-two big-allocation slots. If you pair these together 
you get a massive permanent linear memory growth. On non-wasm targets,
this doesn't matter because these are virtual memory mappings that don't
cost physical memory, but wasm doesn't work that way.

Second, we were using one pool per terminal. On wasm, this meant that 
we paid for the free list N times. On non-wasm, this makes sense because
the synchronization overhead has so far been measurable enough under
load to be prohibitive (although, I'm still skeptical about this and want
to look into it). On wasm, we build single-threaded modules, so we can use
a global free list without any extra overhead.

## Benchmarks

80x24 terminal with 1000-line scrollack processing 16MB of plain ASCII.

| Scenario                        |    Before |     After |
| ------------------------------- | --------: | --------: |
| Fresh instance                  |  0.56 MiB |  0.56 MiB |
| First `terminal_new` (delta)    | +3.44 MiB | +0.88 MiB |
| One filled terminal (total)     |  4.00 MiB |  1.88 MiB |
| Each additional filled terminal | +3.00 MiB | +0.44 MiB |
| 5 filled terminals (total)      | 16.00 MiB |  4.06 MiB |

Throughput numbers are unchanged on wasm and native (to be expected in
the latter because this is all gated on wasm).
2026-08-16 20:55:09 -07:00