2 months ago
9a5229c* fix(utils/paths): truncate text by code point to avoid splitting surrogate pairs truncateText sliced at a UTF-16 code-unit index, so when an astral character (emoji, flags, CJK-ext) straddled the limit it emitted a lone surrogate — invalid UTF-16 that renders as U+FFFD mojibake. The model also saw no truncation marker and could treat the corrupted text as complete. Counting by code point keeps astral symbols whole and leaves BMP behavior identical. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(utils/paths): truncate by code point without full-string allocation Walk code points to find the UTF-16 end index and slice once, keeping the old O(1) BMP fast path allocation-free on hot call sites (tool output truncation) while still avoiding mid-surrogate-pair splits. Also rename a regression test whose "BMP emoji" label was wrong — 😀 is astral. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(utils/paths): validate low surrogate before truncating pair Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: mweinbach <maxweinbach5@gmail.com> Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Parent75d8ed7