parse_object_body allocates an object key, then may fail in parse_colon or
parse_value before that key is ever inserted into the object. Its cleanup defer
only walks `obj`, so a key that never got there is unreachable to it. The caller
cannot free it either -- a failed parse returns a nil Value -- so it leaks.
The same applies to the parsed element on the duplicate-key path, and to both on
the out-of-memory path.
JSON5 makes this reachable from ordinary malformed input, because an unquoted
ident is a legal key and anything other than a colon after it fails. Plain JSON
leaks it too, via a quoted key.
before, measured with a tracking allocator over 8 inputs x 2 specs:
LEAK JSON5 colon fails after unquoted key 1 alloc / 7 bytes
LEAK JSON colon fails after quoted key 1 alloc / 2 bytes
LEAK JSON5 colon fails after quoted key 1 alloc / 2 bytes
LEAK JSON value fails after key 1 alloc / 2 bytes
LEAK JSON5 value fails after key 1 alloc / 2 bytes
LEAK JSON nested value fails 2 alloc / 4 bytes
LEAK JSON5 nested value fails 2 alloc / 4 bytes
LEAK JSON deep nesting fails 3 alloc / 6 bytes
LEAK JSON5 deep nesting fails 3 alloc / 6 bytes
LEAK JSON array element fails 1 alloc / 2 bytes
LEAK JSON5 array element fails 1 alloc / 2 bytes
total leaked allocations: 17
after, same probe:
total leaked allocations: 0
The leak scales with nesting depth -- one orphaned key per enclosing object -- so
a service parsing untrusted JSON leaks a little on every malformed request.
The fix marks the key and the element as owned by the loop iteration until they
are stored, and frees them otherwise. The duplicate-key path loses its explicit
delete, which the same mechanism now covers.
Found via odinfmt, which reported a 7-byte leak in a downstream test that parses
`{ broken not json` to check that invalid input is rejected.
Regression test added to tests/core/encoding/json: it reports
`17 leaks and 0 bad frees` without this change and passes with it. The existing
11 tests pass unchanged under -define:ODIN_TEST_FAIL_ON_BAD_MEMORY=true.
i was working on this when i stumbled on this bug, so i did bunch of different characters to see which ones cause issue, hence test has many different options. I could make it either simpler or split it, as needed.
`decode_grapheme_iterate` already computed the text of each grapheme
cluster, but `decode_grapheme_clusters` discarded it, forcing callers to
reconstruct each cluster's range by looking ahead to the next grapheme's
`byte_index`. Store it on `Grapheme` instead; it is a slice of the input
string, so no allocation is involved.
`core:text/edit` was using `Grapheme.width` -- the number of monospace
cells -- as a byte count when moving the caret by grapheme, so it landed
in the middle of a cluster for anything wider than one byte per cell. Use
`len(Grapheme.text)`, and add a test package for the translations.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>