Json.parse decodes \uXXXX to UTF-8 - a surrogate pair joined, a lone surrogate or bad hex as U+FFFD - and \t \r \b \f, where it dropped the backslash and kept the hex as text; examples/lang/json_unicode checks it byte by byte

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Orkun ÇAKILKAYA 2026-09-29 21:07:10 +03:00
parent 313f95cd1f
commit 4376ddf0ec
4 changed files with 97 additions and 0 deletions

5
changes/json-unicode.md Normal file
View file

@ -0,0 +1,5 @@
bump: patch
type: fix
**`Json.parse` decodes `\uXXXX`.** An escaped code point is UTF-8 now - a surrogate pair joined into one, a
lone surrogate or bad hex as U+FFFD - where the backslash was dropped and the hex kept as text ("iu015f"),
and `\t`, `\r`, `\b` and `\f` are the characters they name. Python's `json.dump` writes non-ASCII this way.