ludic/ns-unicode.html
2026-09-17 22:10:03 +00:00

28 lines
No EOL
4.4 KiB
HTML
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Unicode — Ludic</title>
<meta name="description" content="Correct-by-default text over UTF-8.">
<link rel="stylesheet" href="base.css">
<link rel="stylesheet" href="docs.css">
</head>
<body>
<header class="nav"><div class="wrap nav-in"><a class="brand" href="index.html"><span class="logo">L</span> Ludic</a><button class="nav-toggle" aria-label="Toggle menu" aria-expanded="false">☰</button><nav class="nav-links"><a href="index.html">Home</a><a href="api.html">API Reference</a><a class="nav-cta" href="https://git.workshopsoft.io/workshopsoft/ludic">Source ↗</a></nav></div></header>
<main class="wrap item">
<div class="crumbs"><a href="api.html">API Reference</a> <span>›</span> <span class="here">Unicode</span></div>
<div class="item-head"><span class="kind-badge kind-namespace">namespace</span><h1 id="top">Unicode</h1></div>
<p class="ns-blurb">Correct-by-default text over UTF-8. Player names, translated menus, and chat arrive as UTF-8 bytes, and counting *bytes* gets non-ASCII text wrong — the wrong length, and truncation that slices a character in half into mojibake. The <a href="ns-Unicode"><code>Unicode</code></a> namespace works in <strong>code points</strong> and (approximately) <strong>grapheme clusters</strong> instead, so measuring, indexing, truncating, and case-mapping behave for every language.
This is the correctness layer, not a replacement: the byte-oriented <a href="ns-Text"><code>Text</code></a> operations stay for speed on ASCII and for raw byte work. Reach for <code>Unicode.*</code> whenever the text came from a human — a name, a message, a localized string.
Three notions of "length" matter, and the API keeps them distinct: <a href="unicode-byte_len"><code>byte_len</code></a> (storage), <a href="unicode-len"><code>len</code></a> (code points — Unicode scalar values), and <a href="unicode-grapheme_len"><code>grapheme_len</code></a> (user-perceived characters, where a base letter plus its combining accent, or a ZWJ emoji sequence, count as one).
**Coverage (v1):** decoding and <a href="unicode-is_valid_utf8"><code>validation</code></a> cover the full UTF-8 range. Case mapping covers ASCII and the Latin-1 letters — correct for Western-European text; wider scripts (Latin-Extended, Greek, Cyrillic), locale rules (Turkish <code>i</code>, German <code>ß</code>), and NFC normalization are follow-ups. <a href="unicode-grapheme_len"><code>grapheme_len</code></a> approximates UAX#29 for the cases real player text hits — combining marks, variation selectors, ZWJ sequences, and flag pairs.</p>
<div class="ns-methods"><a class="ns-method" href="unicode-len.html"><code class="nm-sig">Unicode.len(s) -&gt; int</code><span class="nm-tip">Number of code points in a string (not bytes).</span></a><a class="ns-method" href="unicode-byte_len.html"><code class="nm-sig">Unicode.byte_len(s) -&gt; int</code><span class="nm-tip">Number of bytes in a string.</span></a><a class="ns-method" href="unicode-is_valid_utf8.html"><code class="nm-sig">Unicode.is_valid_utf8(s) -&gt; bool</code><span class="nm-tip">Is a string well-formed UTF-8?</span></a><a class="ns-method" href="unicode-char_at.html"><code class="nm-sig">Unicode.char_at(s, i) -&gt; int</code><span class="nm-tip">The i-th code point of a string.</span></a><a class="ns-method" href="unicode-chars.html"><code class="nm-sig">Unicode.chars(s) -&gt; []int</code><span class="nm-tip">Every code point of a string, in order.</span></a><a class="ns-method" href="unicode-upper.html"><code class="nm-sig">Unicode.upper(s) -&gt; string</code><span class="nm-tip">Uppercase a string.</span></a><a class="ns-method" href="unicode-lower.html"><code class="nm-sig">Unicode.lower(s) -&gt; string</code><span class="nm-tip">Lowercase a string.</span></a><a class="ns-method" href="unicode-truncate.html"><code class="nm-sig">Unicode.truncate(s, n) -&gt; string</code><span class="nm-tip">First n code points of a string.</span></a><a class="ns-method" href="unicode-grapheme_len.html"><code class="nm-sig">Unicode.grapheme_len(s) -&gt; int</code><span class="nm-tip">Number of user-perceived characters.</span></a></div>
<a class="back" href="api.html">← All symbols</a>
</main>
<script src="ludic-highlight.js"></script>
<script>Ludic.installCards(); Ludic.flashTarget();</script>
</body></html>