combine utf-16 surrogate pairs in cupsJSONImportString - #169
tanjiroK-coder wants to merge 1 commit into
Conversation
michaelrsweet
left a comment
There was a problem hiding this comment.
I don't really think a specific unit test for this is necessary.
Will look at the rest but I'm inclined to refactor the code a bit first.
|
Also, I don't even think that surrogates are technically valid here - they are a UTF-16 encoding side-effect using a block of reserved Unicode code points while |
8c82f81 to
e2d60b9
Compare
|
Dropped the On the surrogates: RFC 8259 §7 defines |
cupsJSONImportStringdecodes each\uXXXXescape on its own and never combines a UTF-16 surrogate pair, so"\uD834\uDD1E"(U+1D11E) decodes to two 3-byte CESU-8 sequences (ed a0 b4 ed b4 9e) instead of the 4-byte UTF-8f0 9d 84 9e, and a lone surrogate leaves a rawed a0 b4in the value. The JSON here is untrusted:cupsJSONImportURL/cupsOAuthGetTokensfeed it from an OAuth/OIDC endpoint andcupsJWTImportStringruns it over token contents, so a hostile server can drop invalid UTF-8 into strings that later get compared, re-encoded or logged. I spotted it reading the\ubranch after noticingdnssd.calready folds surrogate pairs correctly. The decoder now pairs a high surrogate (D800-DBFF) with the following low surrogate (DC00-DFFF) into one code point emitted as 4-byte UTF-8, and substitutes U+FFFD for an unpaired half rather than emitting a bare surrogate. The pre-scan already reserves five bytes per\uXXXX, so a combined pair uses four of the ten reserved bytes and the allocation is unchanged.testjsongets two cases that fail on the current code.Assisted-by: Claude Code:claude-opus-4-8