28 lines
No EOL
4.4 KiB
HTML
28 lines
No EOL
4.4 KiB
HTML
<!doctype html>
|
||
<html lang="en">
|
||
<head>
|
||
<meta charset="utf-8">
|
||
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||
<title>Unicode — Ludic</title>
|
||
<meta name="description" content="Correct-by-default text over UTF-8.">
|
||
<link rel="stylesheet" href="base.css">
|
||
<link rel="stylesheet" href="docs.css">
|
||
</head>
|
||
<body>
|
||
<header class="nav"><div class="wrap nav-in"><a class="brand" href="index.html"><span class="logo">L</span> Ludic</a><button class="nav-toggle" aria-label="Toggle menu" aria-expanded="false">☰</button><nav class="nav-links"><a href="index.html">Home</a><a href="api.html">API Reference</a><a class="nav-cta" href="https://git.workshopsoft.io/workshopsoft/ludic">Source ↗</a></nav></div></header>
|
||
<main class="wrap item">
|
||
<div class="crumbs"><a href="api.html">API Reference</a> <span>›</span> <span class="here">Unicode</span></div>
|
||
<div class="item-head"><span class="kind-badge kind-namespace">namespace</span><h1 id="top">Unicode</h1></div>
|
||
<p class="ns-blurb">Correct-by-default text over UTF-8. Player names, translated menus, and chat arrive as UTF-8 bytes, and counting *bytes* gets non-ASCII text wrong — the wrong length, and truncation that slices a character in half into mojibake. The <a href="ns-Unicode"><code>Unicode</code></a> namespace works in <strong>code points</strong> and (approximately) <strong>grapheme clusters</strong> instead, so measuring, indexing, truncating, and case-mapping behave for every language.
|
||
|
||
This is the correctness layer, not a replacement: the byte-oriented <a href="ns-Text"><code>Text</code></a> operations stay for speed on ASCII and for raw byte work. Reach for <code>Unicode.*</code> whenever the text came from a human — a name, a message, a localized string.
|
||
|
||
Three notions of "length" matter, and the API keeps them distinct: <a href="unicode-byte_len"><code>byte_len</code></a> (storage), <a href="unicode-len"><code>len</code></a> (code points — Unicode scalar values), and <a href="unicode-grapheme_len"><code>grapheme_len</code></a> (user-perceived characters, where a base letter plus its combining accent, or a ZWJ emoji sequence, count as one).
|
||
|
||
**Coverage (v1):** decoding and <a href="unicode-is_valid_utf8"><code>validation</code></a> cover the full UTF-8 range. Case mapping covers ASCII and the Latin-1 letters — correct for Western-European text; wider scripts (Latin-Extended, Greek, Cyrillic), locale rules (Turkish <code>i</code>, German <code>ß</code>), and NFC normalization are follow-ups. <a href="unicode-grapheme_len"><code>grapheme_len</code></a> approximates UAX#29 for the cases real player text hits — combining marks, variation selectors, ZWJ sequences, and flag pairs.</p>
|
||
<div class="ns-methods"><a class="ns-method" href="unicode-len.html"><code class="nm-sig">Unicode.len(s) -> int</code><span class="nm-tip">Number of code points in a string (not bytes).</span></a><a class="ns-method" href="unicode-byte_len.html"><code class="nm-sig">Unicode.byte_len(s) -> int</code><span class="nm-tip">Number of bytes in a string.</span></a><a class="ns-method" href="unicode-is_valid_utf8.html"><code class="nm-sig">Unicode.is_valid_utf8(s) -> bool</code><span class="nm-tip">Is a string well-formed UTF-8?</span></a><a class="ns-method" href="unicode-char_at.html"><code class="nm-sig">Unicode.char_at(s, i) -> int</code><span class="nm-tip">The i-th code point of a string.</span></a><a class="ns-method" href="unicode-chars.html"><code class="nm-sig">Unicode.chars(s) -> []int</code><span class="nm-tip">Every code point of a string, in order.</span></a><a class="ns-method" href="unicode-upper.html"><code class="nm-sig">Unicode.upper(s) -> string</code><span class="nm-tip">Uppercase a string.</span></a><a class="ns-method" href="unicode-lower.html"><code class="nm-sig">Unicode.lower(s) -> string</code><span class="nm-tip">Lowercase a string.</span></a><a class="ns-method" href="unicode-truncate.html"><code class="nm-sig">Unicode.truncate(s, n) -> string</code><span class="nm-tip">First n code points of a string.</span></a><a class="ns-method" href="unicode-grapheme_len.html"><code class="nm-sig">Unicode.grapheme_len(s) -> int</code><span class="nm-tip">Number of user-perceived characters.</span></a></div>
|
||
<a class="back" href="api.html">← All symbols</a>
|
||
</main>
|
||
<script src="ludic-highlight.js"></script>
|
||
<script>Ludic.installCards(); Ludic.flashTarget();</script>
|
||
</body></html> |