Stage 4a: the rules of paths and texts on the Unicode tables

lib/src/pathrule.dart ports internal/pathrule of datekeys-go at c531e93
(spec §29.5, §29.5.1, §29.6): NFD with the canonical ordering and Hangul,
the case folding, the simple lowercase, Default_Ignorable with the
whitelist of R4, the best-fit projections of R6c, every rule of a path
(R2 to R6c and R10), the tree (R7 with its key and the two paths it names,
and R9), and the texts of the comment and the declared author, with the
texts of Go. canonicalTables is pathrule.Canonical, and a test recomputes
tablesDigest from the lists.

Go reads a string as bytes, and so does this port: the functions whose
name ends in Utf8 take the bytes of a Go string, where a byte that is not
valid UTF-8 is the rune U+FFFD, and the limits count bytes; the others
take a String as utf8Bytes writes it. Each rule returns its violation, as
in Go, and only the public functions throw.

tool/pathrule_go_vectors.go runs in an export of datekeys-go, since
internal/pathrule cannot be imported from outside its tree, and writes
test/vectors/pathrule_vectors.json and its Dart copy: the cases of the
tests of Go and of datekeys-ts, 1300 strings and 350 trees drawn from a
fixed seed (marks, Hangul, ignorables, emoji, best-fit look-alikes,
device names, 8.3 aliases, limits, texts and invalid UTF-8), the cases of
R9, code points, and for each plane the SHA-256 of one line per code
point of each function. Every code point of every plane gives the results
of Go: planes 0, 1 and 14 also on Node.js.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
v0.11
dev 4 days ago
parent b4927829eb
commit 03c4333ca8

@ -0,0 +1,930 @@
/// The rules of the paths and of the texts of the head of a format 3 capsule
/// (spec §29.5, §29.6), as internal/pathrule of datekeys-go: the same checks
/// in the same order, with the same texts. They use the fixed Unicode 18.0.0
/// and WindowsBestFit tables of pathrule_tables.dart (spec §29.5.1), which
/// the reference generates from the same pinned files, and never the
/// Unicode of the platform: no toLowerCase, no RegExp, nothing of dart:core
/// that knows Unicode, whose version changes with each runtime.
///
/// Go reads a string as bytes, and so do these rules:
/// - the functions whose name ends in `Utf8` take the bytes of a Go string,
/// each 0 to 255. Its runes are those of Go's `for range`: every byte that
/// is not part of valid UTF-8 is the rune U+FFFD, and the limits and the
/// comparisons are of bytes, never of UTF-16 code units, save the limit of
/// R3 on the NFD of a segment, which the spec gives in UTF-16;
/// - the other functions take a Dart String as its UTF-8 bytes, [utf8Bytes]
/// of bytes.dart, which writes a lone surrogate as the three bytes of its
/// code point, not valid UTF-8: each gives what the Go function gives for
/// those bytes.
///
/// A reader meets valid UTF-8 only, which the CBOR profile requires (spec
/// §58), and a writer rejects a malformed String before these rules, as the
/// writer of Go rejects a string that is not valid UTF-8. Bytes that are not
/// valid UTF-8 are read as Go reads them only so that every input gives the
/// result of the reference.
///
/// Internal: lib/datekeys.dart does not export it.
library;
import 'dart:typed_data';
import 'bytes.dart';
import 'pathrule_tables.dart';
/// The longest path, in bytes (R1), fixed with format 3 (spec §29.5). The
/// head checks R1 and R8 in its third layer, before these rules.
const maxPathLen = 1024;
/// The most segments of a path (R2).
const maxSegments = 32;
/// The longest segment, in bytes (R3).
const maxSegmentLen = 255;
/// The longest NFD of a segment, in UTF-16 code units, the limit of HFS+
/// (R3).
const maxSegmentUtf16 = 255;
/// The most implicit folders of a head (R9).
const maxImplicitDirs = 65535;
/// The longest comment of a head, in bytes (spec §29.4). The head checks it
/// in its third layer.
const maxCommentLen = 16384;
/// The longest declared author of a head, in bytes (spec §29.4). The head
/// checks it in its third layer.
const maxAuthorLen = 256;
// The first segment of a path must not have a key that starts with it
// (R10).
const _reservedDirPrefix = '.datekeys-';
/// The violation of one rule of spec §29.5 or §29.6: Go's pathrule.Error,
/// whose text is the same in every implementation, so that the message of a
/// rejected head does not depend on the reader. It has no normative code:
/// the head reports it as ERR_HEAD_INVALID, the public note as
/// ERR_EXTENSION_DATA_INVALID.
final class PathRuleException implements Exception {
/// A violation of [rule], described by [detail].
const PathRuleException(this.rule, this.detail, [this.paths = (0, 0)]);
/// "R2" to "R10", or "text".
final String rule;
/// What breaks the rule, such as `segment 1: control U+0009`.
final String detail;
/// The positions, from 1, of the two paths of an R7 violation, the later
/// one first, so that a writer can name them; (0, 0) for the other rules.
final (int, int) paths;
/// The text of Go's Error(): `rule: detail`.
String get message => '$rule: $detail';
@override
String toString() => message;
}
// The violation [e] with `prefix: ` before its detail, as the Go reference
// rewrites e.Detail.
PathRuleException _within(PathRuleException e, String prefix) =>
PathRuleException(e.rule, '$prefix: ${e.detail}', e.paths);
/// [r] as the spec writes a code point: U+ and at least four upper-case
/// hexadecimal digits, Go's `U+%04X`.
String codePointName(int r) =>
'U+${r.toRadixString(16).toUpperCase().padLeft(4, '0')}';
// ---------------------------------------------------------------------------
// Bytes and runes
// [s] as a Uint8List, without a copy when it is one.
Uint8List _bytes(List<int> s) => s is Uint8List ? s : Uint8List.fromList(s);
// The rune at [offset] of [s] and its size in bytes, as Go's `for range`
// reads it: a byte that is not part of valid UTF-8 is U+FFFD of one byte.
(int, int) _runeAt(List<int> s, int offset) {
final b = s[offset];
return b < 0x80 ? (b, 1) : decodeRune(s, offset);
}
// The runes of [s], as Go's `for range` reads a string.
List<int> _runes(List<int> s) {
final out = <int>[];
for (var i = 0; i < s.length;) {
final (r, size) = _runeAt(s, i);
out.add(r);
i += size;
}
return out;
}
// The number of runes of [s], Go's utf8.RuneCount: each byte that is not
// part of valid UTF-8 counts as one.
int _runeCount(List<int> s) {
var n = 0;
for (var i = 0; i < s.length; n++) {
i += _runeAt(s, i).$2;
}
return n;
}
// The size of the valid rune [r] in UTF-8, Go's utf8.RuneLen: three for
// U+FFFD, whatever the byte that it stands for.
int _runeLen(int r) => r < 0x80
? 1
: r < 0x800
? 2
: r < 0x10000
? 3
: 4;
// The UTF-8 bytes of [runes], which are Unicode scalar values.
Uint8List _encode(List<int> runes) {
var n = 0;
for (final r in runes) {
n += _runeLen(r);
}
final out = Uint8List(n);
var o = 0;
for (final r in runes) {
if (r < 0x80) {
out[o++] = r;
} else if (r < 0x800) {
out[o++] = 0xc0 | r >> 6;
out[o++] = 0x80 | r & 0x3f;
} else if (r < 0x10000) {
out[o++] = 0xe0 | r >> 12;
out[o++] = 0x80 | r >> 6 & 0x3f;
out[o++] = 0x80 | r & 0x3f;
} else {
out[o++] = 0xf0 | r >> 18;
out[o++] = 0x80 | r >> 12 & 0x3f;
out[o++] = 0x80 | r >> 6 & 0x3f;
out[o++] = 0x80 | r & 0x3f;
}
}
return out;
}
// The pieces of [s] between the bytes [sep], Go's strings.Split: one more
// than there are separators, so that "" is one empty piece.
List<Uint8List> _split(Uint8List s, int sep) {
final out = <Uint8List>[];
var start = 0;
for (var i = 0; i < s.length; i++) {
if (s[i] == sep) {
out.add(Uint8List.sublistView(s, start, i));
start = i + 1;
}
}
out.add(Uint8List.sublistView(s, start));
return out;
}
// Whether [s] is [n] dots, as bytes or as runes.
bool _isDots(List<int> s, int n) {
if (s.length != n) return false;
for (final c in s) {
if (c != 0x2e) return false;
}
return true;
}
// ---------------------------------------------------------------------------
// The tables
// The index of [v] in the sorted [list], or -1.
int _indexOf(List<int> list, int v) {
var lo = 0;
var hi = list.length;
while (lo < hi) {
final m = (lo + hi) >> 1;
if (list[m] < v) {
lo = m + 1;
} else {
hi = m;
}
}
return lo < list.length && list[lo] == v ? lo : -1;
}
// Whether [r] lies in one of the inclusive ranges lo, hi, lo, hi, ... of
// [ranges]: the first range whose hi is at least r starts at most at r.
bool _inRanges(List<int> ranges, int r) {
final n = ranges.length ~/ 2;
var lo = 0;
var hi = n;
while (lo < hi) {
final m = (lo + hi) >> 1;
if (ranges[2 * m + 1] < r) {
lo = m + 1;
} else {
hi = m;
}
}
return lo < n && ranges[2 * lo] <= r;
}
/// Whether [r] is an assigned code point of Unicode 18.0.0: one that is not
/// of general category Cn. Go's pathrule.Assigned.
bool isAssigned(int r) => _inRanges(assignedRanges, r);
/// Whether [r] has the property Default_Ignorable_Code_Point in Unicode
/// 18.0.0. Go's pathrule.DefaultIgnorable.
bool isDefaultIgnorable(int r) => _inRanges(ignorableRanges, r);
// The first code point with a non-zero combining class: every one before it
// has class 0.
final int _firstMark = cccTable.first >> 8;
// The canonical combining class of [r]. Each entry of cccTable is a code
// point times 256 plus its class, below 2^29: the shifts stay in 32 bits.
int _ccc(int r) {
if (r < _firstMark) return 0;
var lo = 0;
var hi = cccTable.length;
while (lo < hi) {
final m = (lo + hi) >> 1;
if ((cccTable[m] >> 8) < r) {
lo = m + 1;
} else {
hi = m;
}
}
return lo < cccTable.length && (cccTable[lo] >> 8) == r
? cccTable[lo] & 0xff
: 0;
}
// Hangul syllable decomposition (Unicode §3.12).
const _hangulS = 0xac00;
const _hangulL = 0x1100;
const _hangulV = 0x1161;
const _hangulT = 0x11a7;
const _hangulN = 21 * 28;
const _hangulC = 19 * _hangulN;
// NFD of [runes], as Go's pathrule.NFD: every code point replaced by its
// full canonical decomposition, Hangul syllables decomposed
// algorithmically, and each run of combining marks put in canonical order,
// a stable sort by combining class.
List<int> _nfd(List<int> runes) {
final out = <int>[];
for (final r in runes) {
if (r >= _hangulS && r < _hangulS + _hangulC) {
final i = r - _hangulS;
out
..add(_hangulL + i ~/ _hangulN)
..add(_hangulV + (i % _hangulN) ~/ 28);
final t = i % 28;
if (t != 0) out.add(_hangulT + t);
continue;
}
final k = _indexOf(decompositionKeys, r);
if (k < 0) {
out.add(r);
} else {
out.addAll(
decompositionData.getRange(
decompositionStart[k],
decompositionStart[k + 1],
),
);
}
}
// Canonical ordering: an insertion sort within each run of non-starters,
// stable, by combining class. A starter, class 0, stops every mark.
for (var i = 1; i < out.length; i++) {
final c = _ccc(out[i]);
if (c == 0) continue;
for (var j = i; j > 0; j--) {
if (_ccc(out[j - 1]) <= c) break;
final x = out[j - 1];
out[j - 1] = out[j];
out[j] = x;
}
}
return out;
}
// The case folding of [runes], as Go's pathrule.Fold: the entries C and F
// of CaseFolding.txt, and U+0131 to U+0069, as NTFS does.
List<int> _fold(List<int> runes) {
final out = <int>[];
for (final r in runes) {
if (r == 0x0131) {
out.add(0x0069);
continue;
}
final k = _indexOf(foldingKeys, r);
if (k < 0) {
out.add(r);
} else {
out.addAll(foldingData.getRange(foldingStart[k], foldingStart[k + 1]));
}
}
return out;
}
/// The simple lowercase mapping of [r] in Unicode 18.0.0, field 13 of
/// UnicodeData.txt, or [r] itself: Go's pathrule.Lower. The key of words
/// uses it (spec §38.1).
int lower(int r) {
final k = _indexOf(lowercaseKeys, r);
return k < 0 ? r : lowercaseData[lowercaseStart[k]];
}
// The whitelist of R4: the Default_Ignorable_Code_Point that a path may
// hold (spec §29.5).
const _zwnj = 0x200c;
const _zwj = 0x200d;
const _vs15 = 0xfe0e;
const _vs16 = 0xfe0f;
bool _whitelisted(int r) => r == _zwnj || r == _zwj || r == _vs15 || r == _vs16;
// [runes] without the whitelist of R4.
List<int> _strip(List<int> runes) => [
for (final r in runes)
if (!_whitelisted(r)) r,
];
// The key of R7 of the segment [s]: NFD(fold(NFD(s′))), s′ the segment
// without the whitelist of R4. It is removed before normalizing because its
// code points have combining class 0: between two combining marks, they
// would change their canonical order.
List<int> _key(List<int> s) => _nfd(_fold(_nfd(_strip(_runes(s)))));
// Whether [base] forms an emoji variation sequence with the selector [sel],
// U+FE0E or U+FE0F (R4b).
bool _allowsSelector(int base, int sel) =>
_indexOf(sel == _vs15 ? vs15Bases : vs16Bases, base) >= 0;
// ---------------------------------------------------------------------------
// Normalization
/// The canonical decomposition of [s], normalization form D of UAX #15
/// with the data of Unicode 18.0.0, Go's pathrule.NFD. It is not the
/// normalization of the platform.
String nfd(String s) => String.fromCharCodes(_nfd(_runes(utf8Bytes(s))));
/// [nfd] of the bytes [s], in UTF-8.
Uint8List nfdUtf8(List<int> s) => _encode(_nfd(_runes(s)));
/// The case folding of [s], Go's pathrule.Fold: the entries C and F of
/// CaseFolding.txt of Unicode 18.0.0, and U+0131 to U+0069, as NTFS does
/// (R7). It is not toLowerCase.
String fold(String s) => String.fromCharCodes(_fold(_runes(utf8Bytes(s))));
/// [fold] of the bytes [s], in UTF-8.
Uint8List foldUtf8(List<int> s) => _encode(_fold(_runes(s)));
/// The key of a segment in R7, NFD(fold(NFD(s′))) with s′ the segment
/// without ZWNJ, ZWJ, VS15 and VS16: Go's pathrule.Key. Two siblings of a
/// tree must not share one.
String pathKey(String segment) =>
String.fromCharCodes(_key(utf8Bytes(segment)));
/// [pathKey] of the bytes [segment], in UTF-8.
Uint8List pathKeyUtf8(List<int> segment) => _encode(_key(segment));
// ---------------------------------------------------------------------------
// Paths
//
// As in the Go reference, each rule returns its violation, or null when the
// input passes it, and only the public functions throw: a violation found
// costs no exception until it is the answer.
/// Checks one path of the head with the rules of the fourth layer that
/// concern it alone, in this order: R2, and then for each segment R3, R4,
/// R4b, R5, R6, R6b and R6c; then R10. R1 and R8 belong to the third layer,
/// and R7 and R9 to the whole tree ([checkTree]). Throws a
/// [PathRuleException] for the first violation. Go's pathrule.CheckPath.
void checkPath(String path) => checkPathUtf8(utf8Bytes(path));
/// [checkPath] of the bytes [path].
void checkPathUtf8(List<int> path) {
final e = _checkPath(_bytes(path));
if (e != null) throw e;
}
PathRuleException? _checkPath(Uint8List path) {
final segs = _split(path, 0x2f);
if (segs.length > maxSegments) {
return PathRuleException(
'R2',
'${segs.length} segments, more than $maxSegments',
);
}
for (var i = 0; i < segs.length; i++) {
if (segs[i].isEmpty) {
return PathRuleException('R2', 'segment ${i + 1} is empty');
}
}
for (var i = 0; i < segs.length; i++) {
final e = _checkSegment(segs[i]);
if (e != null) return _within(e, 'segment ${i + 1}');
}
if (_startsWithReserved(_key(segs[0]))) {
return PathRuleException(
'R10',
'the first segment starts with ${goQuote(utf8Bytes(_reservedDirPrefix))}',
);
}
return null;
}
// Whether the runes of a key start with those of the reserved prefix of
// R10, all ASCII: the byte prefix of Go's strings.HasPrefix.
bool _startsWithReserved(List<int> key) {
const p = _reservedDirPrefix;
if (key.length < p.length) return false;
for (var i = 0; i < p.length; i++) {
if (key[i] != p.codeUnitAt(i)) return false;
}
return true;
}
// R3, R4, R4b, R5, R6, R6b and R6c on a segment, in that order.
PathRuleException? _checkSegment(Uint8List s) =>
_checkR3(s) ??
_checkR4(s) ??
_checkPlacement('R4b', s) ??
_checkR5(s) ??
_checkR6(s) ??
_checkR6b(s) ??
_checkR6c(s);
PathRuleException? _checkR3(Uint8List s) {
if (s.length > maxSegmentLen) {
return PathRuleException(
'R3',
'${s.length} bytes, more than $maxSegmentLen',
);
}
if (_isDots(s, 1)) {
return const PathRuleException('R3', 'the segment is a dot');
}
if (_isDots(s, 2)) {
return const PathRuleException('R3', 'the segment is two dots');
}
final runes = _runes(s);
final stripped = _strip(runes);
if (stripped.isEmpty) {
return const PathRuleException(
'R3',
'the segment is empty without ZWNJ, ZWJ, VS15 and VS16',
);
}
if (_isDots(stripped, 1)) {
return const PathRuleException(
'R3',
'the segment is a dot without ZWNJ, ZWJ, VS15 and VS16',
);
}
if (_isDots(stripped, 2)) {
return const PathRuleException(
'R3',
'the segment is two dots without ZWNJ, ZWJ, VS15 and VS16',
);
}
final n = _utf16Length(_nfd(runes));
if (n > maxSegmentUtf16) {
return PathRuleException(
'R3',
'its NFD is $n UTF-16 code units, more than $maxSegmentUtf16',
);
}
return null;
}
// The number of UTF-16 code units of [runes].
int _utf16Length(List<int> runes) {
var n = 0;
for (final r in runes) {
n += r >= 0x10000 ? 2 : 1;
}
return n;
}
// The ASCII characters of R4 that are not controls: " * : < > ? \ |.
bool _forbiddenAscii(int r) => switch (r) {
0x22 || 0x2a || 0x3a || 0x3c || 0x3e || 0x3f || 0x5c || 0x7c => true,
_ => false,
};
// R4 on each code point of [s], in order.
PathRuleException? _checkR4(Uint8List s) {
for (var i = 0; i < s.length;) {
final (r, size) = _runeAt(s, i);
final e = _checkR4Rune(r);
if (e != null) return e;
i += size;
}
return null;
}
// R4: the code points that no segment holds.
PathRuleException? _checkR4Rune(int r) {
if (r <= 0x1f || (r >= 0x7f && r <= 0x9f)) {
return PathRuleException('R4', 'control ${codePointName(r)}');
}
if (_forbiddenAscii(r)) {
return PathRuleException('R4', 'character ${codePointName(r)}');
}
if (r == 0x2028 || r == 0x2029) {
return PathRuleException('R4', 'separator ${codePointName(r)}');
}
if (isDefaultIgnorable(r) && !_whitelisted(r)) {
return PathRuleException('R4', 'invisible ${codePointName(r)}');
}
if (r >= 0xf000 && r <= 0xf0ff) {
return PathRuleException('R4', 'private use ${codePointName(r)}');
}
if (!isAssigned(r)) {
return PathRuleException('R4', 'unassigned ${codePointName(r)}');
}
return null;
}
// R4b on [s], a segment or a line of a text, under the name [rule]: VS15
// and VS16 only right after a character that forms an emoji variation
// sequence with them, and ZWJ and ZWNJ never first, last or right after
// another ZWJ or ZWNJ.
//
// The end is found as the Go reference finds it: i adds utf8.RuneLen of each
// rune and is compared with the length of [s]. Over valid UTF-8, i is the
// offset of the rune; after a byte that is not, it runs ahead by two for
// each such byte, read as U+FFFD of three, and the result follows it.
PathRuleException? _checkPlacement(String rule, Uint8List s) {
var prev = -1;
var i = 0;
for (var o = 0; o < s.length;) {
final (r, size) = _runeAt(s, o);
o += size;
if (r == _vs15 || r == _vs16) {
if (prev < 0 || !_allowsSelector(prev, r)) {
return PathRuleException(
rule,
'${codePointName(r)} is not part of an emoji variation sequence',
);
}
} else if (r == _zwj || r == _zwnj) {
if (prev < 0) {
return PathRuleException(rule, '${codePointName(r)} at the start');
}
if (prev == _zwj || prev == _zwnj) {
return PathRuleException(
rule,
'${codePointName(r)} right after ${codePointName(prev)}',
);
}
if (i + _runeLen(r) == s.length) {
return PathRuleException(rule, '${codePointName(r)} at the end');
}
}
prev = r;
i += _runeLen(r);
}
return null;
}
PathRuleException? _checkR5(Uint8List s) {
if (s.isEmpty) return null;
if (s.first == 0x20) {
return const PathRuleException('R5', 'the segment starts with U+0020');
}
if (s.last == 0x20) {
return const PathRuleException('R5', 'the segment ends with U+0020');
}
if (s.last == 0x2e) {
return const PathRuleException('R5', "the segment ends with '.'");
}
return null;
}
// The device names of R6, upper case, by the Latin-1 String of their UTF-8
// bytes, the form in which a segment is compared with them. Six end in a
// superscript digit, U+00B9, U+00B2 or U+00B3.
final Map<String, String> _reserved = {
for (final name in [
'CON',
'PRN',
'AUX',
'NUL',
r'CONIN$',
r'CONOUT$',
'COM¹',
'COM²',
'COM³',
'LPT¹',
'LPT²',
'LPT³',
for (var d = 0; d <= 9; d++) ...['COM$d', 'LPT$d'],
])
String.fromCharCodes(utf8Bytes(name)): name,
};
// The longest device name, in bytes: a longer base is none of them.
final int _longestReserved = _reserved.keys
.map((k) => k.length)
.reduce((a, b) => a > b ? a : b);
// R6: the part before the first '.', without its final U+0020 and with the
// ASCII letters, and only those, in upper case, is not a device name.
PathRuleException? _checkR6(Uint8List s) {
var end = s.indexOf(0x2e);
if (end < 0) end = s.length;
while (end > 0 && s[end - 1] == 0x20) {
end--;
}
if (end > _longestReserved) return null;
final name =
_reserved[String.fromCharCodes([
for (var i = 0; i < end; i++)
s[i] >= 0x61 && s[i] <= 0x7a ? s[i] - 0x20 : s[i],
])];
if (name == null) return null;
return PathRuleException('R6', '$name is a reserved device name');
}
// R6b: not the three conditions of an 8.3 alias of NTFS together, counting
// runes: at most one '.', a base of 1 to 8 that matches ^[^.]*~[0-9]{1,6}$,
// and an extension of 0 to 3.
PathRuleException? _checkR6b(Uint8List s) {
var dots = 0;
for (final c in s) {
if (c == 0x2e) dots++;
}
if (dots > 1) return null;
final dot = s.indexOf(0x2e);
final base = dot < 0 ? s : Uint8List.sublistView(s, 0, dot);
final ext = dot < 0 ? Uint8List(0) : Uint8List.sublistView(s, dot + 1);
final n = _runeCount(base);
if (n < 1 || n > 8) return null;
if (_runeCount(ext) > 3 || !_isShortNameBase(base)) return null;
return const PathRuleException(
'R6b',
'the segment has the form of an 8.3 alias',
);
}
// Whether [base], which holds no '.', matches ^[^.]*~[0-9]{1,6}$: it ends in
// '~' and 1 to 6 ASCII digits. The digits are the whole run of digits at its
// end, since the byte before them must be '~'. An ASCII byte is always a rune
// of its own, also among bytes that are not valid UTF-8, as Go's regexp reads
// them.
bool _isShortNameBase(Uint8List base) {
var k = 0;
while (k < base.length &&
base[base.length - 1 - k] >= 0x30 &&
base[base.length - 1 - k] <= 0x39) {
k++;
}
return k >= 1 &&
k <= 6 &&
k < base.length &&
base[base.length - 1 - k] == 0x7e;
}
// R6c: for each best-fit table, the projection of the segment, where each
// code point of U+0080 and above that the table maps to an ASCII byte
// becomes that byte, does not hold '/', '\', ':' or U+0000, and passes R3,
// R5, R6 and R6b. A projection is in UTF-8: a byte that is not valid UTF-8
// becomes the three bytes of U+FFFD, as Go's strings.Builder writes it.
PathRuleException? _checkR6c(Uint8List s) {
final runes = _runes(s);
// No table maps an ASCII code point, so a segment of ASCII stays as it is.
if (!runes.any((r) => r >= 0x80)) return null;
for (final t in bestFitTables) {
List<int>? p;
for (var i = 0; i < runes.length; i++) {
final r = runes[i];
if (r < 0x80) continue;
final k = _indexOf(t.from, r);
if (k >= 0) (p ??= List.of(runes))[i] = t.to[k];
}
if (p == null) continue;
for (final r in p) {
if (r == 0x2f || r == 0x5c || r == 0x3a || r == 0) {
return PathRuleException(
'R6c',
'code page ${t.name} maps the segment to one with ${codePointName(r)}',
);
}
}
final projected = _encode(p);
for (final check in [_checkR3, _checkR5, _checkR6, _checkR6b]) {
final e = check(projected);
if (e != null) {
return PathRuleException(
'R6c',
'code page ${t.name} maps the segment to one that breaks '
'${e.rule}: ${e.detail}',
);
}
}
}
return null;
}
// ---------------------------------------------------------------------------
// Trees
/// Applies R7 and then R9 to the paths of a head, which have passed
/// [checkPath]: no two siblings with the same key ([pathKey]), no path that
/// is both a file and a folder, comparing keys, and at most
/// [maxImplicitDirs] folders. Paths are numbered from 1 in the messages, and
/// an R7 violation names the two in [PathRuleException.paths]. Go's
/// pathrule.CheckTree.
void checkTree(List<String> paths) =>
checkTreeUtf8([for (final p in paths) utf8Bytes(p)]);
/// [checkTree] of the bytes of [paths].
void checkTreeUtf8(List<List<int>> paths) {
// The children of each folder by their key, "" for the root; the key of a
// folder joins the keys of its segments with '/'. A key is the String of
// its runes: two are equal when their UTF-8 bytes are.
final children = <String, Map<String, _Node>>{};
for (var n = 0; n < paths.length; n++) {
final segs = _split(_bytes(paths[n]), 0x2f);
var parent = '';
for (var i = 0; i < segs.length; i++) {
final s = segs[i];
final k = String.fromCharCodes(_key(s));
final isDir = i < segs.length - 1;
final kids = children.putIfAbsent(parent, () => <String, _Node>{});
final old = kids[k];
if (old == null) {
kids[k] = (name: s, dir: isDir, path: n + 1);
} else if (!equalBytes(old.name, s)) {
throw PathRuleException(
'R7',
'path ${n + 1} collides with path ${old.path} in segment ${i + 1}',
(n + 1, old.path),
);
} else if (old.dir != isDir || !isDir) {
throw PathRuleException(
'R7',
'path ${n + 1} makes a file of path ${old.path} a folder, or the '
'reverse, in segment ${i + 1}',
(n + 1, old.path),
);
}
parent = parent.isEmpty ? k : '$parent/$k';
}
}
// The implicit folders: the distinct bytes before each '/' of each path,
// as Latin-1 Strings of those bytes.
final dirs = <String>{};
for (final path in paths) {
for (var i = 0; i < path.length; i++) {
if (path[i] == 0x2f) dirs.add(String.fromCharCodes(path, 0, i));
}
}
if (dirs.length > maxImplicitDirs) {
throw PathRuleException(
'R9',
'${dirs.length} folders, more than $maxImplicitDirs',
);
}
}
// A node of the tree of R7: the segment that first reached it, whether it
// is a folder, and the path, from 1, that reached it first.
typedef _Node = ({Uint8List name, bool dir, int path});
// ---------------------------------------------------------------------------
// Texts
/// Checks the comment of a head (spec §29.6): no control but TAB and LF, no
/// bidirectional control, U+2028, U+2029, U+FEFF or noncharacter, no
/// Default_Ignorable_Code_Point but the whitelist of R4, and R4b on each
/// line. Throws a [PathRuleException] of the rule "text". Go's
/// pathrule.CheckComment.
void checkComment(String text) => checkCommentUtf8(utf8Bytes(text));
/// [checkComment] of the bytes [text].
void checkCommentUtf8(List<int> text) {
final e = _checkText(_bytes(text), true);
if (e != null) throw e;
}
/// Checks the declared author of a head (spec §29.6), whose rules the
/// public note follows too (spec §24.1): those of the comment, without TAB
/// and LF, and with no U+0020 at either end. Throws a [PathRuleException] of
/// the rule "text". Go's pathrule.CheckAuthor.
void checkAuthor(String text) => checkAuthorUtf8(utf8Bytes(text));
/// [checkAuthor] of the bytes [text].
void checkAuthorUtf8(List<int> text) {
final s = _bytes(text);
final e = _checkText(s, false);
if (e != null) throw e;
if (s.isNotEmpty && (s.first == 0x20 || s.last == 0x20)) {
throw const PathRuleException(
'text',
'the declared author starts or ends with U+0020',
);
}
}
// The rules of spec §29.6, code point by code point, and then R4b on each
// line.
PathRuleException? _checkText(Uint8List s, bool comment) {
for (var i = 0; i < s.length;) {
final (r, size) = _runeAt(s, i);
i += size;
final String what;
if (r == 0x09 || r == 0x0a) {
if (comment) continue;
what = 'control ${codePointName(r)} in the declared author';
} else if (r <= 0x1f || (r >= 0x7f && r <= 0x9f)) {
what = 'control ${codePointName(r)}';
} else if ((r >= 0x202a && r <= 0x202e) ||
(r >= 0x2066 && r <= 0x2069) ||
r == 0x061c ||
r == 0x200e ||
r == 0x200f) {
what = 'bidirectional control ${codePointName(r)}';
} else if (r == 0x2028 || r == 0x2029) {
what = 'separator ${codePointName(r)}';
} else if (r == 0xfeff) {
what = 'byte order mark ${codePointName(r)}';
} else if (_isNoncharacter(r)) {
what = 'noncharacter ${codePointName(r)}';
} else if (isDefaultIgnorable(r) && !_whitelisted(r)) {
what = 'invisible ${codePointName(r)}';
} else {
continue;
}
return PathRuleException('text', what);
}
final lines = _split(s, 0x0a);
for (var i = 0; i < lines.length; i++) {
final e = _checkPlacement('text', lines[i]);
if (e != null) return _within(e, 'line ${i + 1}');
}
return null;
}
// The 66 noncharacters: U+FDD0 to U+FDEF, and the last two code points of
// each plane. A code point is below 2^21: the mask stays in 32 bits.
bool _isNoncharacter(int r) =>
(r >= 0xfdd0 && r <= 0xfdef) || r & 0xfffe == 0xfffe;
// ---------------------------------------------------------------------------
// The digest of the tables
String _hex(int v, int digits) =>
v.toRadixString(16).toUpperCase().padLeft(digits, '0');
/// The canonical text of the tables, whose SHA-256 is [tablesDigest]: Go's
/// pathrule.Canonical, line by line. Each value is in upper-case hexadecimal
/// of at least four digits, but the combining class and the code page, in
/// decimal, and the best-fit byte, in two hexadecimal digits.
String canonicalTables() {
final b = StringBuffer('unicode $unicodeVersion\n');
for (var i = 0; i < assignedRanges.length; i += 2) {
b.write(
'assigned ${_hex(assignedRanges[i], 4)} '
'${_hex(assignedRanges[i + 1], 4)}\n',
);
}
for (var i = 0; i < ignorableRanges.length; i += 2) {
b.write(
'ignorable ${_hex(ignorableRanges[i], 4)} '
'${_hex(ignorableRanges[i + 1], 4)}\n',
);
}
for (final e in cccTable) {
b.write('ccc ${_hex(e >> 8, 4)} ${e & 0xff}\n');
}
for (final (name, keys, start, data) in [
('decomp', decompositionKeys, decompositionStart, decompositionData),
('fold', foldingKeys, foldingStart, foldingData),
('lower', lowercaseKeys, lowercaseStart, lowercaseData),
]) {
for (var i = 0; i < keys.length; i++) {
b.write('$name ${_hex(keys[i], 4)}');
for (var j = start[i]; j < start[i + 1]; j++) {
b.write(' ${_hex(data[j], 4)}');
}
b.write('\n');
}
}
for (final r in vs15Bases) {
b.write('vs15 ${_hex(r, 4)}\n');
}
for (final r in vs16Bases) {
b.write('vs16 ${_hex(r, 4)}\n');
}
for (final t in bestFitTables) {
for (var i = 0; i < t.from.length; i++) {
b.write('bestfit ${t.name} ${_hex(t.from[i], 4)} ${_hex(t.to[i], 2)}\n');
}
}
return b.toString();
}

@ -0,0 +1,95 @@
// The exhaustive comparison of the rules of paths and texts with Go: for each
// code point of a plane, one line per function, hashed with SHA-256, against
// the digests of test/vectors/pathrule_vectors.json, which
// tool/pathrule_go_vectors.go computes with the same lines.
import 'dart:typed_data';
import 'package:datekeys/src/bytes.dart';
import 'package:datekeys/src/pathrule.dart';
import 'package:datekeys/src/sha256.dart';
import 'package:test/test.dart';
/// Lines of ASCII added to a SHA-256 through a buffer of 64 KiB.
final class LineHash {
final _hash = Sha256();
final _buffer = Uint8List(1 << 16);
var _length = 0;
void add(String line) {
if (_length + line.length > _buffer.length) _flush();
for (var i = 0; i < line.length; i++) {
final c = line.codeUnitAt(i);
if (c >= 0x80) throw ArgumentError.value(line, 'line', 'not ASCII');
_buffer[_length++] = c;
}
}
void _flush() {
_hash.add(_buffer, 0, _length);
_length = 0;
}
String finish() {
_flush();
return toHex(_hash.finish());
}
}
String _hexUpper(int r) => r.toRadixString(16).toUpperCase();
String _outcome(void Function() f) {
try {
f();
return 'ok';
} on PathRuleException catch (e) {
return e.message;
}
}
/// The planes whose code points pathrule_test.dart checks, also compiled to
/// JavaScript: the BMP, the SMP and plane 14, of the tags and of the
/// variation selectors. pathrule_vm_test.dart checks the others.
const sharedPlanes = [0, 1, 14];
/// The names of the digests of a plane, in the order of the lines.
const planeFunctions = [
'props',
'nfd',
'fold',
'key',
'path',
'comment',
'author',
];
/// Checks every code point of [plane] against [digests], the entry of the
/// plane in the vectors: for each code point r, its string is its UTF-8, or
/// U+FFFD for a surrogate, as Go's string(rune(r)) writes it.
void checkPlane(int plane, Map<String, Object?> digests) {
final replacement = Uint8List.fromList([0xef, 0xbf, 0xbd]);
final h = [for (final _ in planeFunctions) LineHash()];
for (var r = plane << 16; r <= (plane << 16 | 0xffff); r++) {
final s = r >= 0xd800 && r <= 0xdfff
? replacement
: utf8Bytes(String.fromCharCode(r));
final x = _hexUpper(r);
h[0].add(
'$x ${isAssigned(r)} ${isDefaultIgnorable(r)} ${_hexUpper(lower(r))}\n',
);
h[1].add('$x ${toHex(nfdUtf8(s))}\n');
h[2].add('$x ${toHex(foldUtf8(s))}\n');
h[3].add('$x ${toHex(pathKeyUtf8(s))}\n');
h[4].add('$x ${_outcome(() => checkPathUtf8(s))}\n');
h[5].add('$x ${_outcome(() => checkCommentUtf8(s))}\n');
h[6].add('$x ${_outcome(() => checkAuthorUtf8(s))}\n');
}
expect(digests['plane'], plane);
for (var i = 0; i < planeFunctions.length; i++) {
expect(
h[i].finish(),
digests[planeFunctions[i]],
reason: 'plane $plane, ${planeFunctions[i]}',
);
}
}

@ -0,0 +1,293 @@
// The rules of the paths and texts of a format 3 head (spec §29.5, §29.6)
// against test/vectors/pathrule_vectors.json, whose expected values the Go
// reference computed (tool/pathrule_go_vectors.go): the cases of the tests
// of internal/pathrule and of datekeys-ts, and strings and trees drawn from a
// fixed seed, bytes that are not valid UTF-8 among them. The vectors come
// from a Dart constant, so that these tests also run compiled to
// JavaScript, where a String counts UTF-16 code units and the integers are
// doubles. Every code point of three planes is checked here, those of the
// others and the shared vectors of testdata/ in pathrule_vm_test.dart.
import 'dart:convert';
import 'package:datekeys/src/bytes.dart';
import 'package:datekeys/src/pathrule.dart';
import 'package:datekeys/src/pathrule_tables.dart';
import 'package:datekeys/src/sha256.dart';
import 'package:test/test.dart';
import 'pathrule_planes.dart';
import 'vectors/pathrule_vectors.g.dart';
final Map<String, Object?> vectors =
jsonDecode(pathruleVectorsJson) as Map<String, Object?>;
List<Map<String, Object?>> cases(String name) =>
(vectors[name]! as List).cast<Map<String, Object?>>();
String cp(List<int> runes) => String.fromCharCodes(runes);
/// "ok", or the message of the [PathRuleException] that [f] throws, as the
/// vectors write the result of Go.
String outcome(void Function() f) {
try {
f();
return 'ok';
} on PathRuleException catch (e) {
return e.message;
}
}
/// The [PathRuleException] that [f] throws.
PathRuleException violation(void Function() f) {
try {
f();
} on PathRuleException catch (e) {
return e;
}
throw StateError('no violation');
}
/// A case of the vectors for a failure message: its name, or its bytes.
String label(Map<String, Object?> c) => c['name'] as String? ?? 'in ${c['in']}';
void main() {
group('the tables', () {
test('match their digest, recomputed from the lists as Go does', () {
final digest = toHex(sha256(utf8Bytes(canonicalTables())));
expect(digest, tablesDigest);
expect(vectors['tables_digest'], tablesDigest);
expect(unicodeVersion, vectors['unicode_version']);
expect(sources, hasLength(vectors['sources']));
expect(bestFitTables, hasLength(15));
});
test('the limits are those of Go', () {
expect(vectors['limits'], {
'max_author_len': maxAuthorLen,
'max_comment_len': maxCommentLen,
'max_implicit_dirs': maxImplicitDirs,
'max_path_len': maxPathLen,
'max_segment_len': maxSegmentLen,
'max_segment_utf16': maxSegmentUtf16,
'max_segments': maxSegments,
});
});
test('Assigned, DefaultIgnorable and Lower of code points', () {
final points = (vectors['code_points']! as List).cast<List<Object?>>();
expect(points, hasLength(greaterThan(1000)));
for (final p in points) {
final r = p[0]! as int;
final name = codePointName(r);
expect(isAssigned(r), p[1] == 1, reason: '$name assigned');
expect(isDefaultIgnorable(r), p[2] == 1, reason: '$name ignorable');
expect(lower(r), p[3], reason: '$name lower');
}
});
});
group('strings', () {
test('CheckPath, CheckComment, CheckAuthor, NFD and Key of Go', () {
final failures = <String>[];
void check(String what, Object? got, Object? want, String c) {
if (got != want) failures.add('$c: $what $got, want $want');
}
final all = cases('strings');
expect(all, hasLength(greaterThan(1400)));
for (final c in all) {
final input = fromHex(c['in']! as String);
final name = label(c);
final comment = c['comment'];
check('path', outcome(() => checkPathUtf8(input)), c['path'], name);
check('comment', outcome(() => checkCommentUtf8(input)), comment, name);
check(
'author',
outcome(() => checkAuthorUtf8(input)),
c['author'] ?? comment,
name,
);
check('nfd', toHex(nfdUtf8(input)), c['nfd'], name);
check('key', toHex(pathKeyUtf8(input)), c['key'], name);
}
expect(failures, isEmpty, reason: failures.take(20).join('\n'));
});
test('a String is its UTF-8, a lone surrogate the bytes of Go', () {
// Every valid UTF-8 input, as a String, gives the same results; and a
// lone surrogate is the three bytes of its code point, which Go reads
// as three U+FFFD.
var strings = 0;
for (final c in cases('strings')) {
final s = decodeUtf8(fromHex(c['in']! as String));
if (s == null) continue;
strings++;
final name = label(c);
expect(outcome(() => checkPath(s)), c['path'], reason: name);
expect(outcome(() => checkComment(s)), c['comment'], reason: name);
expect(
outcome(() => checkAuthor(s)),
c['author'] ?? c['comment'],
reason: name,
);
expect(utf8Bytes(nfd(s)), fromHex(c['nfd']! as String), reason: name);
expect(
utf8Bytes(pathKey(s)),
fromHex(c['key']! as String),
reason: name,
);
}
expect(strings, greaterThan(1000));
final surrogate = cases('strings')
.singleWhere((c) => c['name'] == 'a surrogate in UTF-8 bytes');
expect(surrogate['in'], 'eda080');
final lone = cp([0xd800]);
expect(outcome(() => checkPath(lone)), surrogate['path']);
expect(outcome(() => checkComment(lone)), surrogate['comment']);
expect(utf8Bytes(nfd(lone)), fromHex(surrogate['nfd']! as String));
expect(utf8Bytes(pathKey(lone)), fromHex(surrogate['key']! as String));
});
test('Fold maps each code point alone', () {
// The planes of pathrule_vm_test.dart check Fold code point by code
// point against Go; on a string it is their concatenation.
for (final c in cases('strings')) {
final input = fromHex(c['in']! as String);
final each = <int>[];
for (final r in decodeUtf8(nfdUtf8(input))!.runes) {
each.addAll(foldUtf8(utf8Bytes(cp([r]))));
}
expect(foldUtf8(nfdUtf8(input)), each, reason: label(c));
}
});
test('the named cases of Go, with their exact texts', () {
final named = {
for (final c in cases('strings'))
if (c['name'] != null) c['name']! as String: c,
};
expect(named, hasLength(greaterThan(140)));
String path(String name) => named[name]!['path']! as String;
// A few of them as the reader of a head sees them.
expect(outcome(() => checkPath('a${cp([9])}b')), path('TAB'));
expect(path('TAB'), 'R4: segment 1: control U+0009');
expect(
outcome(() => checkPath('CON.txt')),
'R6: segment 1: CON is a reserved device name',
);
expect(path('CON.txt'), 'R6: segment 1: CON is a reserved device name');
expect(
path('.datekeys-x'),
'R10: the first segment starts with ".datekeys-"',
);
expect(outcome(() => checkPath('.datekeys-x')), path('.datekeys-x'));
expect(
outcome(() => checkPath(cp([0x3000, 0x61]))),
path('U+3000 first'),
);
expect(
outcome(() => checkPath(cp(List.filled(127, 0x390)))),
path('127 times U+0390: 381 UTF-16 units after NFD'),
);
expect(
outcome(() => checkAuthor(' Ana')),
named['a leading space in an author']!['author'],
);
});
test('a violation carries its rule and detail apart', () {
final e = violation(() => checkPath('a/b${cp([0x202e])}'));
expect(e.rule, 'R4');
expect(e.detail, 'segment 2: invisible U+202E');
expect(e.message, 'R4: segment 2: invisible U+202E');
expect(e.toString(), e.message);
expect(e.paths, (0, 0));
final t = violation(() => checkComment('ok\na${cp([0xfe0f])}'));
expect(t.rule, 'text');
expect(
t.detail,
'line 2: U+FE0F is not part of an emoji variation sequence',
);
});
test('limits count UTF-8 bytes, not UTF-16 code units', () {
// 85 emoji are 170 UTF-16 code units and 340 bytes: R3 counts bytes.
final emoji = cp(List.filled(85, 0x1f600));
expect(emoji.length, 170);
expect(
outcome(() => checkPath(emoji)),
'R3: segment 1: 340 bytes, more than 255',
);
// R6b counts code points: an astral letter is one, not two.
final astral = '${cp([0x10400])}BCDEF~1';
expect(astral.length, 9);
expect(
outcome(() => checkPath(astral)),
'R6b: segment 1: the segment has the form of an 8.3 alias',
);
});
});
test('every code point of planes 0, 1 and 14 gives the results of Go', () {
final planes = cases('planes');
for (final plane in sharedPlanes) {
checkPlane(plane, planes[plane]);
}
});
group('trees', () {
test('CheckTree of Go, with the two paths of R7', () {
final all = cases('trees');
expect(all, hasLength(greaterThan(350)));
var r7 = 0;
for (final c in all) {
final paths = [
for (final p in (c['in'] as List?) ?? const <Object?>[])
fromHex(p! as String),
];
final name = label2(c);
final want = c['result']! as String;
expect(outcome(() => checkTreeUtf8(paths)), want, reason: name);
final pair = (c['paths']! as List).cast<int>();
if (want != 'ok') {
final e = violation(() => checkTreeUtf8(paths));
expect(e.paths, (pair[0], pair[1]), reason: name);
if (e.rule == 'R7') r7++;
} else {
expect(pair, [0, 0], reason: name);
}
final strings = [for (final p in paths) decodeUtf8(p)];
if (strings.every((s) => s != null)) {
expect(
outcome(() => checkTree([for (final s in strings) s!])),
want,
reason: name,
);
}
}
expect(r7, greaterThan(50));
});
test('R9 counts the distinct folders of every path', () {
for (final c in cases('folders')) {
final prefix = c['prefix']! as String;
final suffix = c['suffix']! as String;
final digits = c['digits']! as int;
final paths = [
for (var i = 0; i < (c['count']! as int); i++)
'$prefix${i.toString().padLeft(digits, '0')}$suffix',
];
expect(
outcome(() => checkTree(paths)),
c['result'],
reason: c['name']! as String,
);
}
});
});
}
/// A tree of the vectors for a failure message: its name, or its paths.
String label2(Map<String, Object?> c) =>
c['name'] as String? ?? 'paths ${c['in']}';

@ -0,0 +1,134 @@
// The rules of paths and texts against files: the copy of
// test/vectors/pathrule_vectors.json that pathrule_test.dart reads, a Dart
// constant for the tests compiled to JavaScript, is the JSON file byte for
// byte; the shared vectors of testdata/, paths.json and path_fold.json of
// datekeys-go, give their results; and every code point of every plane
// gives the results of Go, compared through the SHA-256 of one line per
// code point that tool/pathrule_go_vectors.go computes with the reference:
// here the planes that pathrule_test.dart leaves out.
@TestOn('vm')
library;
import 'dart:convert';
import 'dart:io';
import 'package:datekeys/src/bytes.dart';
import 'package:datekeys/src/pathrule.dart';
import 'package:datekeys/src/pathrule_tables.dart';
import 'package:test/test.dart';
import 'pathrule_planes.dart';
import 'vectors/pathrule_vectors.g.dart';
Map<String, Object?> readJson(String path) =>
jsonDecode(File(path).readAsStringSync()) as Map<String, Object?>;
String outcome(void Function() f) {
try {
f();
return 'ok';
} on PathRuleException catch (e) {
return e.message;
}
}
void main() {
test('pathrule_vectors.g.dart holds pathrule_vectors.json', () {
final file = File('test/vectors/pathrule_vectors.json').readAsStringSync();
expect(pathruleVectorsJson, file);
});
group('testdata', () {
final paths = readJson('testdata/vectors/paths.json');
final fold = readJson('testdata/vectors/path_fold.json');
test('name the tables of this library', () {
for (final f in [paths, fold]) {
expect(f['unicode_version'], unicodeVersion);
expect(f['tables_digest'], tablesDigest);
}
});
test('paths.json: the rules of one entry', () {
final cases = (paths['paths']! as List).cast<Map<String, Object?>>();
expect(cases, hasLength(greaterThan(80)));
for (final c in cases) {
expect(
outcome(() => checkPath(c['path']! as String)),
c['result'],
reason: c['name'] as String?,
);
}
});
test('paths.json: the paths of a head, from layer 3 to layer 4', () {
// A head checks R1 and R8 in its third layer (ERR_NON_CANONICAL_CBOR),
// comparing UTF-8 bytes, never UTF-16 code units; then each entry,
// with R2 to R10, and R7 and R9 over the tree (ERR_HEAD_INVALID).
// The decoder of the head, in stage 4c, does this with the CBOR; here
// it is done with the rules of pathrule.dart.
final trees = (paths['trees']! as List).cast<Map<String, Object?>>();
expect(trees, hasLength(greaterThan(10)));
var utf16Differs = false;
for (final c in trees) {
final name = c['name']! as String;
final list = (c['paths']! as List).cast<String>();
var result = 'ok';
var detail = '';
for (var i = 0; i < list.length && result == 'ok'; i++) {
final n = utf8Bytes(list[i]).length;
if (n < 1 ||
n > maxPathLen ||
(i > 0 &&
compareBytes(utf8Bytes(list[i - 1]), utf8Bytes(list[i])) >=
0)) {
result = 'ERR_NON_CANONICAL_CBOR';
}
if (i > 0 && list[i - 1].compareTo(list[i]) >= 0 && result == 'ok') {
utf16Differs = true;
}
}
for (var i = 0; i < list.length && result == 'ok'; i++) {
final o = outcome(() => checkPath(list[i]));
if (o != 'ok') {
result = 'ERR_HEAD_INVALID';
detail = 'file ${i + 1}: $o';
}
}
if (result == 'ok') {
final o = outcome(() => checkTree(list));
if (o != 'ok') {
result = 'ERR_HEAD_INVALID';
detail = o;
}
}
expect(result, c['result'], reason: name);
expect(detail, c['detail'] ?? '', reason: name);
}
// U+FF5E before U+1F600 is in UTF-8 order, not in UTF-16 order.
expect(utf16Differs, isTrue);
});
test('path_fold.json: the NFD and the key of R7', () {
final keys = (fold['keys']! as List).cast<Map<String, Object?>>();
expect(keys, hasLength(greaterThan(20)));
for (final c in keys) {
final s = c['segment']! as String;
expect(nfd(s), c['nfd'], reason: c['name'] as String?);
expect(pathKey(s), c['key'], reason: c['name'] as String?);
}
});
});
test('every code point of the other planes gives the results of Go', () {
// Planes 0, 1 and 14 run in pathrule_test.dart, also compiled to
// JavaScript.
final vectors = jsonDecode(pathruleVectorsJson) as Map<String, Object?>;
final planes = (vectors['planes']! as List).cast<Map<String, Object?>>();
expect(planes, hasLength(17));
for (final p in planes) {
final plane = p['plane']! as int;
if (!sharedPlanes.contains(plane)) checkPlane(plane, p);
}
});
}

File diff suppressed because it is too large Load Diff

File diff suppressed because it is too large Load Diff

File diff suppressed because it is too large Load Diff
Loading…
Cancel
Save

Powered by TurnKey Linux.