wordlist.dart: the random words of a key of words, as wordkey.Generate

A port of List, CheckList, Generate and Bits of package wordkey of
datekeys-go at 27a75ee (generate.go, spec 38.1), with the same checks in the
same order and the same texts, each error a WordKeyException:

- readWordList takes the file of a list that the app downloads or bundles
  only with the SHA-256 that wordListSha256 pins, "wordkey: the list "es"
  has the SHA-256 ..., not ...", and then reads it as List reads its text;
- checkWordList refuses fewer than 2048 words, a word that is not one word
  of 3 letters or more once normalized, a rune that checkWords refuses, a
  character outside the alphabet of the language, which only the code
  gives (es: a to z, the five vowels with an acute accent, u with diaeresis
  and n with tilde, lower case, NFC), and two words that are one once
  normalized;
- generateWords draws each index with randomIndex, as crypto/rand.Int, and
  draws again an index already drawn, so that the same bytes draw the same
  words as Go; what the RandomSource throws goes through;
- wordBits sums Go's math.Log2, ported with its Frexp and its Log, so that
  the double is Go's, on the VM and on the web.

wordkey.dart shares its check of each rune, checkWordRune, with the same
behaviour. lib/datekeys.dart does not export the new file yet.

test/vectors/wordlist_vectors.json, from tool/wordlist_go_vectors.go in an
export of datekeys-go at 27a75ee, holds what Go gives: List on 19 texts,
CheckList on 83 lists and on every code point of planes 0, 1 and 14,
Generate in 51 cases, on the Spanish list and on lists of 12 to 2^20 + 1
words, from seeded, counter and finite streams, and Bits bit for bit.
wordlist_test.dart runs them, also compiled to JavaScript, with the cases of
TestCheckList, TestGenerate and TestBits of Go; wordlist_vm_test.dart checks
the Dart copy of the vectors and every code point.

The seed of TestGenerate of Go draws two indices only, 1793 and 2081, so
that Generate runs out of bytes and the test compares two errors; the
vectors record it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
v0.15
dev 10 hours ago
parent faa2c4c891
commit a770d4d8b0

@ -116,6 +116,23 @@ void checkWordsUtf8(List<List<int>> words) {
final b = w[i];
final (r, size) = b < 0x80 ? (b, 1) : decodeRune(w, i);
i += size;
checkWordRune(r);
}
if (n >= minLetters) counted.add(String.fromCharCodes(w));
}
if (counted.length < minWords) {
throw WordKeyException(
'wordkey: a key of words needs at least $minWords different words of '
'$minLetters or more letters, not ${counted.length}',
);
}
}
/// Throws the [WordKeyException] of Go's wordkey.checkRunes when the rune
/// [r] of a word is a control, a Default_Ignorable_Code_Point or a code
/// point unassigned in Unicode 18.0.0: the check of each rune of
/// [checkWords], and of each word of a list in wordlist.dart.
void checkWordRune(int r) {
// Go's unicode.IsControl: C0 and C1, nothing above U+00FF.
if (r <= 0x1f || (r >= 0x7f && r <= 0x9f)) {
throw WordKeyException(
@ -124,23 +141,13 @@ void checkWordsUtf8(List<List<int>> words) {
}
if (isDefaultIgnorable(r)) {
throw WordKeyException(
'wordkey: the words hold the invisible character '
'${codePointName(r)}',
'wordkey: the words hold the invisible character ${codePointName(r)}',
);
}
if (!isAssigned(r)) {
throw WordKeyException(
'wordkey: the words hold ${codePointName(r)}, unassigned in '
'Unicode $unicodeVersion',
);
}
}
if (n >= minLetters) counted.add(String.fromCharCodes(w));
}
if (counted.length < minWords) {
throw WordKeyException(
'wordkey: a key of words needs at least $minWords different words of '
'$minLetters or more letters, not ${counted.length}',
'wordkey: the words hold ${codePointName(r)}, unassigned in Unicode '
'$unicodeVersion',
);
}
}

@ -0,0 +1,295 @@
/// The random words of a key of words (spec §38.1: at least 6 words of a
/// public list of 2048 or more, which a writer SHOULD offer), as List,
/// CheckList, Generate and Bits of package wordkey of datekeys-go
/// (generate.go): the word lists of datekeys-go, taken only with the
/// SHA-256 pinned here and checked against the alphabet of their language;
/// the words drawn from one, uniformly, with a [RandomSource]; and their
/// strength in bits.
///
/// Go builds its lists into the module; this library holds none. The app
/// downloads or bundles the file of a list, `wordkey/lists/<lang>.txt` of
/// datekeys-go (vendored in wordlists/ of this package), and gives its
/// bytes to [readWordList], which trusts none: their SHA-256 must be that of
/// [wordListSha256], and the words must pass [checkWordList]. The lists are
/// not normative: a reader does not need them, because the key is derived
/// from the normalized text of the words, whatever list they came from.
/// Each list has a license of its own, which its README records (es.txt is
/// CC BY-SA 4.0, an adaptation of FrequencyWords by Hermit Dave), unlike the
/// code.
///
/// As in wordkey.dart, a Dart String is taken as its UTF-8 bytes
/// ([utf8Bytes]), and the functions ending in `Utf8` take the bytes of a Go
/// string: every input gives the result and the text of Go.
library;
import 'dart:math' as math;
import 'dart:typed_data';
import 'bytes.dart';
import 'pathrule.dart' show codePointName;
import 'random.dart';
import 'sha256.dart';
import 'wordkey.dart';
/// The number of words that [generateWords] draws when the caller does not
/// ask for more: 7 words of a list of 7776 are about 90 bits ([wordBits]).
const defaultWordCount = 7;
/// The fewest words of a list that [checkWordList] accepts (spec §38.1: at
/// least 6 words of a list of 2048 or more).
const minListSize = 2048;
/// The SHA-256 of each word list of datekeys-go, by language, as its
/// wordkey/lists/README.md records it: [readWordList] takes a list only with
/// these bytes, so that a list changes only with its hash here.
const Map<String, String> wordListSha256 = {
'es': 'ff77b487765c000da97cca58fe94a2cdb947303e7a07460614d7d95d800034fe',
};
// The letters that a word of a list of each language may hold, as the list
// writes it: lower case, in NFC. A word with a letter of another script that
// looks like one of these, such as the Cyrillic U+0430 for the Latin a,
// would be written down and typed again with the letter of the keyboard, and
// the capsule would not open. The alphabet of a list comes from here, never
// from the list: Go's alphabets.
const _alphabets = {
'es': 'abcdefghijklmnopqrstuvwxyz\u00e1\u00e9\u00ed\u00f3\u00fa\u00fc\u00f1',
};
String _q(String s) => goQuote(utf8Bytes(s));
/// The languages of the lists of [wordListSha256], sorted: Go's Languages.
List<String> wordListLanguages() => wordListSha256.keys.toList()..sort();
/// The words of [file], the word list of the language [lang], one word per
/// line, as Go's List gives the list built into it: [file] must have the
/// SHA-256 of [wordListSha256], and its words must pass [checkWordList].
/// For a list that the app downloads or bundles, which it trusts no more
/// than any other download. Throws a [WordKeyException] with Go's text:
/// `wordkey: no word list for "xx"; the lists are es` for a language
/// without a list; `wordkey: the list "es" has the SHA-256 …, not …` for any
/// other bytes; and the error of [checkWordList] after `wordkey: the list
/// "es": `.
List<String> readWordList(String lang, List<int> file) {
final want = wordListSha256[lang];
if (want == null) {
throw WordKeyException(
'wordkey: no word list for ${_q(lang)}; the lists are '
'${wordListLanguages().join(', ')}',
);
}
final got = toHex(sha256(file));
if (got != want) {
throw WordKeyException(
'wordkey: the list ${_q(lang)} has the SHA-256 $got, not $want',
);
}
return parseWordList(lang, file);
}
/// [readWordList] without its two first checks, the language and the
/// SHA-256: the lines of [file], split as Go's List splits its text,
/// `strings.Split(strings.TrimSuffix(text, "\n"), "\n")`, and checked with
/// [checkWordListUtf8]. The bytes are taken as Go takes a string: a byte
/// that is not valid UTF-8 is U+FFFD, which no alphabet holds, so the check
/// refuses it with Go's text, and the words that pass are decoded strictly.
/// For the tests: the app reads a list with [readWordList].
List<String> parseWordList(String lang, List<int> file) {
final lines = _lines(file);
try {
checkWordListUtf8(lang, lines);
} on WordKeyException catch (e) {
throw WordKeyException('wordkey: the list ${_q(lang)}: ${e.message}');
}
return [for (final l in lines) decodeUtf8(l)!];
}
// The lines of [file], as Go's strings.Split(strings.TrimSuffix(text,
// "\n"), "\n"): one final LF goes, and the text is split at every LF, so
// that an empty text is one empty line. A byte 0x0A is never part of a
// longer UTF-8 sequence.
List<Uint8List> _lines(List<int> file) {
final b = file is Uint8List ? file : Uint8List.fromList(file);
final end = b.isNotEmpty && b.last == 0x0a ? b.length - 1 : b.length;
final lines = <Uint8List>[];
var start = 0;
for (var i = 0; i < end; i++) {
if (b[i] == 0x0a) {
lines.add(Uint8List.sublistView(b, start, i));
start = i + 1;
}
}
lines.add(Uint8List.sublistView(b, start, end));
return lines;
}
/// Checks that [words] can be a list of the language [lang] for
/// [generateWords], as Go's CheckList: the same checks, in the same order,
/// with Go's text in a [WordKeyException]. It refuses a language without an
/// alphabet here; fewer than [minListSize] words; a word that is not one
/// word of [minLetters] characters or more once normalized
/// ([normalizeWords]), that holds a character that [checkWords] refuses, or
/// one that is not a letter of the alphabet of [lang], as the list writes
/// it; and a word that is the same as an earlier one once normalized. Two
/// words such as «papa» and «papá» would be one word with less entropy than
/// the list promises.
void checkWordList(String lang, List<String> words) =>
checkWordListUtf8(lang, [for (final w in words) utf8Bytes(w)]);
/// [checkWordList] of the bytes of [words].
void checkWordListUtf8(String lang, List<List<int>> words) {
final alphabet = _alphabets[lang];
if (alphabet == null) {
throw WordKeyException('no alphabet for the language ${_q(lang)}');
}
if (words.length < minListSize) {
throw WordKeyException('${words.length} words, fewer than $minListSize');
}
final letters = alphabet.runes.toSet();
// The first word of the list that is each normalized word.
final seen = <String, List<int>>{};
for (var i = 0; i < words.length; i++) {
final w = words[i];
final line = i + 1;
final n = normalizeWordsUtf8(w);
// Go also refuses a word with white space at its ends, which a word of
// normalizeWords never has.
if (n.length != 1 || n[0].runes.length < minLetters) {
throw WordKeyException(
'line $line, ${goQuote(w)}, is not one word of $minLetters or more '
'letters',
);
}
try {
for (final r in n[0].runes) {
checkWordRune(r);
}
} on WordKeyException catch (e) {
throw WordKeyException('line $line: ${e.message}');
}
// The runes of the word as the list writes it, as Go's `for range`
// reads them: a byte that is not valid UTF-8 is U+FFFD.
for (var j = 0; j < w.length;) {
final (r, size) = w[j] < 0x80 ? (w[j], 1) : decodeRune(w, j);
if (!letters.contains(r)) {
throw WordKeyException(
'line $line, ${goQuote(w)}, holds ${codePointName(r)}, which is '
'not in the alphabet of ${_q(lang)}',
);
}
j += size;
}
final prev = seen[n[0]];
if (prev != null) {
throw WordKeyException(
'line $line, ${goQuote(w)}, is the same word as ${goQuote(prev)} '
'once normalized',
);
}
seen[n[0]] = w;
}
}
/// [n] different words of [list], drawn uniformly with [random], the CSPRNG
/// of the platform by default, as Go's Generate: each index as
/// crypto/rand.Int draws it ([randomIndex]), and an index already drawn is
/// drawn again, so that the same random bytes draw the same words as Go.
/// Each word adds log2 of the size of the list, a little less for each word
/// already drawn ([wordBits]). The list is taken as it is: it is one that
/// [readWordList] gave.
///
/// Throws a [WordKeyException] with Go's text when [n] is less than
/// [minWords] or more than half the list, before drawing anything. What
/// [random] throws goes through: a [RandomSource] does not run out, where
/// Go's io.Reader can, and Go's Generate then fails with `wordkey: EOF`.
List<String> generateWords(
List<String> list, [
int n = defaultWordCount,
RandomSource random = secureRandom,
]) {
if (n < minWords) {
throw WordKeyException(
'wordkey: a key of words needs at least $minWords words, not $n',
);
}
if (n > list.length ~/ 2) {
throw WordKeyException('wordkey: $n words of a list of ${list.length}');
}
final drawn = <int>{};
final words = <String>[];
while (words.length < n) {
final i = randomIndex(random, list.length);
if (drawn.add(i)) words.add(list[i]);
}
return words;
}
/// The strength of [count] words that [generateWords] draws from a list of
/// [size] words, as Go's Bits: log2 of the number of draws in order,
/// size·(size − 1)·…, which whoever knows the list must search, before the
/// rounds of PBKDF2. It is the sum of the log2 of size − i for each i below
/// [count], each by the math.Log2 of Go, so that the double is Go's, on the
/// VM and on the web: 7 words of 7776 are a little under 90.5 bits.
double wordBits(int size, int count) {
var b = 0.0;
for (var i = 0; i < count; i++) {
b += _log2((size - i).toDouble());
}
return b;
}
// Go's math.Log2 of [x], the one of every platform but s390x: x is
// frac·2^exp with frac in [0.5, 1), as math.Frexp splits it; a power of two
// gives exp − 1, exactly, and any other x Log(frac)·(1/Ln2) + exp. 0 gives
// −Infinity and a negative x NaN, as in Go.
double _log2(double x) {
if (x == 0) return double.negativeInfinity;
if (x < 0 || x.isNaN) return double.nan;
if (x.isInfinite) return x;
final (frac, exp) = _frexp(x);
if (frac == 0.5) return (exp - 1).toDouble();
// dart:math's log2e is Go's 1/Ln2 as a float64, 0x3ff71547652b82fe.
return _log(frac) * math.log2e + exp;
}
final _float = ByteData(8);
// Go's math.Frexp of [x], positive and normal: frac in [0.5, 1) and exp,
// with x = frac·2^exp, from the bits of x. Only the high 32 bits change, so
// that the bit operators are exact on the web too.
(double, int) _frexp(double x) {
_float.setFloat64(0, x);
final high = _float.getUint32(0);
_float.setUint32(0, (high & 0x000fffff) | 0x3fe00000);
return (_float.getFloat64(0), (high >> 20) - 1022);
}
// Go's math.Log of [x], positive and normal: its port of FreeBSD's
// e_log.c, with the operations of its code, and of its assembly for amd64,
// in their order, so that each double is Go's.
double _log(double x) {
const ln2Hi = 6.93147180369123816490e-01; // 3fe62e42 fee00000
const ln2Lo = 1.90821492927058770002e-10; // 3dea39ef 35793c76
const l1 = 6.666666666666735130e-01; // 3fe55555 55555593
const l2 = 3.999999999940941908e-01; // 3fd99999 9997fa04
const l3 = 2.857142874366239149e-01; // 3fd24924 94229359
const l4 = 2.222219843214978396e-01; // 3fcc71c5 1d8e78af
const l5 = 1.818357216161805012e-01; // 3fc74664 96cb03de
const l6 = 1.531383769920937332e-01; // 3fc39a09 d078c69f
const l7 = 1.479819860511658591e-01; // 3fc2f112 df3e5244
var (f1, ki) = _frexp(x);
if (f1 < math.sqrt2 / 2) {
f1 *= 2;
ki--;
}
final f = f1 - 1;
final k = ki.toDouble();
final s = f / (2 + f);
final s2 = s * s;
final s4 = s2 * s2;
final t1 = s2 * (l1 + s4 * (l3 + s4 * (l5 + s4 * l7)));
final t2 = s4 * (l2 + s4 * (l4 + s4 * l6));
final r = t1 + t2;
final hfsq = 0.5 * f * f;
return k * ln2Hi - ((hfsq - (s * (hfsq + r) + k * ln2Lo)) - f);
}

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

@ -0,0 +1,168 @@
// Helpers of the tests of the word lists: the base list of TestCheckList of
// Go and its edits, as test/vectors/wordlist_vectors.json writes them; a
// list of words that are their indices; and the random sources of the
// vectors of Generate. They read no file.
import 'dart:collection';
import 'dart:typed_data';
import 'package:datekeys/src/bytes.dart';
import 'package:datekeys/src/random.dart';
import 'package:datekeys/src/sha256.dart';
import 'package:datekeys/src/wordkey.dart' show WordKeyException;
typedef Json = Map<String, Object?>;
/// "ok", or the message of the [WordKeyException] that [f] throws.
String outcome(void Function() f) {
try {
f();
return 'ok';
} on WordKeyException catch (e) {
return e.message;
}
}
/// Word [i] of the base list of TestCheckList: "pal" and three letters,
/// palaaa, palaab, …, up to i = 17 575, palzzz.
String baseWord(int i) {
final letters = [i ~/ 676, i ~/ 26 % 26, i % 26];
return 'pal${String.fromCharCodes([for (final l in letters) 0x61 + l])}';
}
/// The words of a case of the vectors, as bytes: the first `size` words of
/// the base list, the word of each edit of `set` at its index, and the
/// words of `append` after them.
List<Uint8List> caseWords(Json c) {
final words = [
for (var i = 0; i < (c['size']! as int); i++) utf8Bytes(baseWord(i)),
];
for (final e in (c['set'] as List? ?? const []).cast<Json>()) {
words[e['at']! as int] = fromHex(e['word']! as String);
}
for (final w in (c['append'] as List? ?? const []).cast<String>()) {
words.add(fromHex(w));
}
return words;
}
/// [parts] joined by [separator].
Uint8List joinBytes(List<List<int>> parts, List<int> separator) => concatBytes([
for (var i = 0; i < parts.length; i++) ...[if (i > 0) separator, parts[i]],
]);
/// [bytes] repeated [count] times, as Go's bytes.Repeat.
Uint8List repeatBytes(List<int> bytes, int count) =>
concatBytes([for (var i = 0; i < count; i++) bytes]);
/// The 64 bits of [x], big-endian, in hexadecimal: Go's
/// `%016x` of math.Float64bits.
String doubleHex(double x) =>
toHex(Uint8List.sublistView(ByteData(8)..setFloat64(0, x)));
/// The double of the 64 bits [hex].
double hexDouble(String hex) =>
ByteData.sublistView(fromHex(hex)).getFloat64(0);
/// The words "0", "1", …, one less than [length]: a list whose words are
/// their indices, without storing them. generateWords reads only its length
/// and the words it draws.
final class IndexWords extends ListBase<String> {
IndexWords(this._length);
final int _length;
@override
int get length => _length;
@override
set length(int _) => throw UnsupportedError('a fixed length');
@override
String operator [](int i) {
RangeError.checkValidIndex(i, this);
return '$i';
}
@override
void operator []=(int i, String value) =>
throw UnsupportedError('an unmodifiable list');
}
/// A source that counts the bytes that another one gives.
final class CountingSource implements RandomSource {
CountingSource(this._inner);
final RandomSource _inner;
/// The bytes given so far.
int read = 0;
@override
void fill(Uint8List out) {
_inner.fill(out);
read += out.length;
}
}
/// The bytes of [bytes], then nothing: a fill past them gives nothing and
/// throws a [StateError], where Go's io.Reader gives io.EOF, or the bytes
/// left and io.ErrUnexpectedEOF.
final class ReplaySource implements RandomSource {
ReplaySource(List<int> bytes) : _bytes = Uint8List.fromList(bytes);
final Uint8List _bytes;
int _at = 0;
@override
void fill(Uint8List out) {
if (_at + out.length > _bytes.length) throw StateError('no more bytes');
out.setRange(0, out.length, _bytes, _at);
_at += out.length;
}
}
/// SHA-256 of [label] and a 32-bit big-endian counter from 0, block after
/// block: the counter stream of the vectors.
final class CounterSource implements RandomSource {
CounterSource(List<int> label) : _label = Uint8List.fromList(label);
final Uint8List _label;
int _counter = 0;
Uint8List _block = Uint8List(0);
int _at = 0;
@override
void fill(Uint8List out) {
for (var n = 0; n < out.length;) {
if (_at == _block.length) {
final input = Uint8List(_label.length + 4)..setAll(0, _label);
ByteData.sublistView(input).setUint32(_label.length, _counter++);
_block = sha256(input);
_at = 0;
}
final k = out.length - n < _block.length - _at
? out.length - n
: _block.length - _at;
out.setRange(n, n + k, _block, _at);
n += k;
_at += k;
}
}
}
/// The source of the stream of a case of generate: `seeded`, the keystream
/// of [SeededRandomSource] of the seed; `counter`, a [CounterSource] of the
/// label; `bytes`, hex repeated `repeat` times, a [ReplaySource].
RandomSource streamOf(Json s) => switch (s['kind']) {
'seeded' => SeededRandomSource(utf8Bytes(s['seed']! as String)),
'counter' => CounterSource(utf8Bytes(s['seed']! as String)),
'bytes' => ReplaySource(
repeatBytes(fromHex(s['hex'] as String? ?? ''), s['repeat'] as int? ?? 0),
),
final kind => throw ArgumentError.value(kind, 'kind'),
};
/// The bytes that crypto/rand.Int reads for each draw of an index below
/// [size]: those of size − 1.
int drawSize(int size) => ((size - 1).bitLength + 7) ~/ 8;

@ -0,0 +1,341 @@
// The word lists of the key of words (spec §38.1) against
// test/vectors/wordlist_vectors.json, whose expected values the Go reference
// computed (tool/wordlist_go_vectors.go): the constants and the pinned
// lists; the text of a list read as wordkey.List reads it; wordkey.CheckList
// with its texts, also on each code point up to U+017F; the words that
// wordkey.Generate draws from the same bytes; and wordkey.Bits, bit for bit.
// And the cases of TestCheckList, TestGenerate and TestBits of Go, as Go
// writes them. The vectors come from a Dart constant, so that these tests
// also run compiled to JavaScript; the Spanish list, the CSPRNG of the
// platform and the alphabet of every code point are in wordlist_vm_test.dart.
import 'dart:convert';
import 'dart:math' as math;
import 'dart:typed_data';
import 'package:datekeys/src/bytes.dart';
import 'package:datekeys/src/pathrule.dart' show codePointName;
import 'package:datekeys/src/random.dart';
import 'package:datekeys/src/sha256.dart';
import 'package:datekeys/src/wordkey.dart';
import 'package:datekeys/src/wordlist.dart';
import 'package:test/test.dart';
import 'random_support.dart' show seeded;
import 'vectors/wordlist_vectors.g.dart';
import 'wordkey_support.dart' show fromWtf8;
import 'wordlist_support.dart';
final Json vectors = jsonDecode(wordlistVectorsJson) as Json;
List<Json> cases(String name) => (vectors[name]! as List).cast<Json>();
void main() {
test('the constants and the pinned lists are those of Go', () {
expect(vectors['default_count'], defaultWordCount);
expect(vectors['min_list_size'], minListSize);
expect(vectors['min_words'], minWords);
expect(vectors['min_letters'], minLetters);
// Languages of Go, and the SHA-256 of the file of each of its lists.
final lists = cases('lists');
expect(wordListLanguages(), [for (final l in lists) l['lang']]);
expect(wordListSha256, {for (final l in lists) l['lang']: l['sha256']});
});
group('checkWordList', () {
test('the cases of TestCheckList of Go', () {
// pal followed by three letters: palaaa, palaab, …
final base = [for (var i = 0; i < minListSize; i++) baseWord(i)];
checkWordList('es', base);
expect(
outcome(() => checkWordList('xx', base)),
'no alphabet for the language "xx"',
);
List<String> withWord(int i, String w) => [...base]..[i] = w;
for (final (list, want) in [
(base.sublist(0, minListSize - 1), '2047 words, fewer than 2048'),
(
withWord(5, 'dos palabras'),
'line 6, "dos palabras", is not one word',
),
(withWord(5, ' '), 'is not one word'),
(withWord(5, 'mi'), '"mi", is not one word of 3 or more letters'),
(withWord(5, 'casa\u200b'), 'invisible character U+200B'),
// Only the letters of the alphabet of the language, as the list
// writes them: no capitals, no Cyrillic U+0441 that looks like a
// Latin c, no digits, no carriage return of a file with CRLF lines.
(
withWord(5, 'Palaaf'),
'line 6, "Palaaf", holds U+0050, which is not in the alphabet of '
'"es"',
),
(
withWord(5, '\u0441asa'),
'holds U+0441, which is not in the alphabet',
),
(withWord(5, 'pal1'), 'holds U+0031'),
(withWord(5, 'palaaf\r'), r'line 6, "palaaf\r", holds U+000D'),
(
withWord(5, base[4]),
'line 6, "palaae", is the same word as "palaae"',
),
(
withWord(5, 'pala\u00e1e'),
'line 6, "pala\u00e1e", is the same word as "palaae"',
),
(
[...withWord(0, 'pap\u00e1'), 'papa'],
'"papa", is the same word as "pap\u00e1"',
),
]) {
expect(outcome(() => checkWordList('es', list)), contains(want));
}
});
test('refuses lists as wordkey.CheckList, with its texts', () {
final all = cases('check_list');
var refused = 0;
var strings = 0;
for (final c in all) {
final name = c['name']! as String;
final lang = c['lang']! as String;
final want = c['result']! as String;
final words = caseWords(c);
expect(
toHex(sha256(joinBytes(words, const [0x0a]))),
c['sha256'],
reason: name,
);
expect(
outcome(() => checkWordListUtf8(lang, words)),
want,
reason: name,
);
if (want != 'ok') refused++;
// The same words as Strings, where they have one: the bytes of a
// surrogate are a lone surrogate, as utf8Bytes writes it.
final asStrings = [for (final w in words) fromWtf8(w)];
if (asStrings.every((s) => s != null)) {
strings++;
expect(
outcome(() => checkWordList(lang, [for (final s in asStrings) s!])),
want,
reason: name,
);
}
}
expect(all, hasLength(greaterThan(80)));
expect(refused, greaterThan(70));
// All but three, whose bytes no String has.
expect(strings, all.length - 3);
});
test('the alphabet of each code point up to U+017F', () {
final results = cases('alphabet_results');
expect(
[for (final c in results) c['rune']],
[for (var r = 0; r < 0x180; r++) r],
);
// The first word is pala, the code point and zz.
final words = caseWords(const {'size': minListSize});
final accepted = <int>[];
for (final c in results) {
final r = c['rune']! as int;
words[0] = utf8Bytes('pala${String.fromCharCode(r)}zz');
final got = outcome(() => checkWordListUtf8('es', words));
expect(got, c['result'], reason: codePointName(r));
if (got == 'ok') accepted.add(r);
}
// a to z, á, é, í, ñ, ó, ú and ü: all that Go accepts in planes 0, 1
// and 14.
expect(
String.fromCharCodes(accepted),
'abcdefghijklmnopqrstuvwxyz'
'\u00e1\u00e9\u00ed\u00f1\u00f3\u00fa\u00fc',
);
final ok = ((vectors['alphabet']! as Json)['ok']! as List).cast<String>();
expect([for (final s in ok) s.runes.single], accepted);
});
});
test('parseWordList reads the text of a list as wordkey.List', () {
for (final c in cases('parse')) {
final name = c['name']! as String;
final lang = c['lang']! as String;
final lines = caseWords(c);
final text = concatBytes([
fromHex(c['prefix']! as String),
joinBytes(lines, fromHex(c['sep']! as String)),
fromHex(c['suffix']! as String),
]);
expect(toHex(sha256(text)), c['sha256'], reason: name);
final want = c['result']! as String;
expect(outcome(() => parseWordList(lang, text)), want, reason: name);
if (want == 'ok') {
expect(lines, hasLength(c['words']), reason: name);
expect(parseWordList(lang, text), [
for (final l in lines) decodeUtf8(l),
], reason: name);
}
}
});
test('readWordList refuses a language without a list and other bytes', () {
expect(
outcome(() => readWordList('xx', const [])),
'wordkey: no word list for "xx"; the lists are es',
);
final text = utf8Bytes('palaaa\n');
expect(
outcome(() => readWordList('es', text)),
'wordkey: the list "es" has the SHA-256 ${toHex(sha256(text))}, not '
'${wordListSha256['es']}',
);
// The language first, as Go's List looks for its list first.
expect(
outcome(() => readWordList('ES', text)),
'wordkey: no word list for "ES"; the lists are es',
);
});
group('generateWords', () {
test('draws the words of wordkey.Generate from the same bytes', () {
final all = cases('generate');
var drawn = 0;
for (final c in all) {
final name = c['name']! as String;
final size = c['size']! as int;
final n = c['n']! as int;
final read = c['read']! as int;
final error = c['error'] as String?;
final source = CountingSource(streamOf(c['stream']! as Json));
final list = IndexWords(size);
if (error == null) {
final words = generateWords(list, n, source);
expect(
[for (final w in words) int.parse(w)],
c['indices'],
reason: name,
);
expect(source.read, read, reason: name);
final next = c['next'] as String?;
if (next != null) {
expect(toHex(randomBytes(source, 4)), next, reason: name);
}
drawn++;
} else if (error.endsWith('EOF')) {
// The bytes end: the source throws, and generateWords lets it
// through, after the draws that Go made in full.
expect(
() => generateWords(list, n, source),
throwsStateError,
reason: name,
);
expect(source.read, read - read % drawSize(size), reason: name);
} else {
// The arguments, refused before anything is drawn.
expect(
outcome(() => generateWords(list, n, source)),
error,
reason: name,
);
expect(source.read, 0, reason: name);
}
}
expect(all, hasLength(greaterThan(40)));
expect(drawn, greaterThan(30));
});
test('the cases of TestGenerate of Go', () {
final list = IndexWords(7776);
// The same random bytes draw the same words: generateWords reads
// nothing else.
final a = generateWords(list, 6, seeded('TestGenerate'));
expect(generateWords(list, 6, seeded('TestGenerate')), a);
expect(a.toSet(), hasLength(6));
// The seed of TestGenerate of Go draws two indices only, 1793 and
// 2081, again and again until it runs out: Go's Generate then fails
// with EOF, and its two results compare equal (see the vectors).
expect(
() => generateWords(
list,
6,
ReplaySource(repeatBytes(const [7, 1, 200, 33], 64)),
),
throwsStateError,
);
for (final (n, want) in [
(5, 'at least 6 words, not 5'),
(list.length, '7776 words of a list of 7776'),
]) {
expect(
outcome(() => generateWords(list, n, ReplaySource(const []))),
contains(want),
);
}
// Without random bytes: the source throws, where Go's Generate fails
// with EOF.
expect(
() => generateWords(list, 6, ReplaySource(const [])),
throwsStateError,
);
});
});
group('wordBits', () {
double goBits(int size, int count) => hexDouble(
cases(
'bits',
).singleWhere((c) => c['size'] == size && c['count'] == count)['hex']!
as String,
);
test('the cases of TestBits of Go', () {
// Go: Bits(7776, 1) is math.Log2(7776).
expect(wordBits(7776, 1), goBits(7776, 1));
expect(wordBits(7776, 1), closeTo(math.log(7776) / math.ln2, 1e-12));
// 7 words of 7776 are a little under 90.5 bits, and 6 of 2048, the
// fewest that generateWords draws, a little under 66.
for (final (size, count, low, high) in [
(7776, 0, 0.0, 0.0),
(7776, 7, 90.469, 90.470),
(7776, 8, 103.393, 103.394),
(2048, 6, 65.989, 65.990),
]) {
expect(
wordBits(size, count),
inInclusiveRange(low, high),
reason: '$size, $count',
);
}
});
test('is Bits of Go, bit for bit', () {
for (final c in cases('bits')) {
final size = c['size']! as int;
final count = c['count']! as int;
final want = c['hex']! as String;
final got = wordBits(size, count);
if (hexDouble(want).isNaN) {
// The NaN of Go, 7ff8000000000001, is not that of Dart.
expect(got.isNaN, isTrue, reason: '$size, $count');
} else {
expect(doubleHex(got), want, reason: '$size, $count: ${c['bits']}');
}
}
for (final d in cases('bits_digests')) {
final sizes = (d['sizes']! as List).cast<int>();
final counts = (d['counts']! as List).cast<int>();
final digest = Sha256();
final b = ByteData(8);
for (var size = sizes[0]; size <= sizes[1]; size++) {
for (var count = counts[0]; count <= counts[1]; count++) {
b.setFloat64(0, wordBits(size, count));
digest.add(Uint8List.sublistView(b));
}
}
expect(toHex(digest.finish()), d['sha256'], reason: '${d['name']}');
}
});
});
}

@ -0,0 +1,53 @@
// The word lists against files, and at length: the copy of
// test/vectors/wordlist_vectors.json that wordlist_test.dart reads, a Dart
// constant for the tests compiled to JavaScript, is the JSON file byte for
// byte; and checkWordList gives the result of Go's CheckList for every code
// point of planes 0, 1 and 14. On the VM only: they read a file or take a
// few seconds.
@TestOn('vm')
library;
import 'dart:convert';
import 'dart:io';
import 'package:datekeys/src/bytes.dart';
import 'package:datekeys/src/pathrule.dart' show codePointName;
import 'package:datekeys/src/sha256.dart';
import 'package:datekeys/src/wordlist.dart';
import 'package:test/test.dart';
import 'vectors/wordlist_vectors.g.dart';
import 'wordlist_support.dart';
void main() {
final vectors = jsonDecode(wordlistVectorsJson) as Json;
test('wordlist_vectors.g.dart holds wordlist_vectors.json', () {
final file = File('test/vectors/wordlist_vectors.json').readAsStringSync();
expect(wordlistVectorsJson, file);
});
test('the alphabet of every code point of planes 0, 1 and 14', () {
// The first word of the base list is pala, the code point and zz: the
// lines of every result, but for the surrogates, have Go's SHA-256.
final a = vectors['alphabet']! as Json;
final words = caseWords(const {'size': minListSize});
final digest = Sha256();
final accepted = <String>[];
var count = 0;
for (final p in (a['planes']! as List).cast<int>()) {
for (var r = p << 16; r <= (p << 16 | 0xffff); r++) {
if (r >= 0xd800 && r <= 0xdfff) continue;
final letter = String.fromCharCode(r);
words[0] = utf8Bytes('pala${letter}zz');
final result = outcome(() => checkWordListUtf8('es', words));
digest.add(utf8Bytes('${codePointName(r)} $result\n'));
count++;
if (result == 'ok') accepted.add(letter);
}
}
expect(count, a['count']);
expect(toHex(digest.finish()), a['sha256']);
expect(accepted, a['ok']);
});
}

@ -0,0 +1,727 @@
//go:build ignore
// Writes test/vectors/wordlist_vectors.json and the same JSON as a Dart
// constant, wordlist_vectors.g.dart: the results of the word lists of
// package wordkey of the Go reference (generate.go: List, CheckList,
// Generate and Bits, the random words of spec §38.1) for
// lib/src/wordlist.dart. Every expected value is computed here by the Go
// reference; none is written by hand.
//
// - lists: the built-in lists, with the SHA-256, the size and the number
// of words of each file of wordkey/lists, which List reads.
// - parse: the body of List, restated because List reads only its
// built-in lists, on texts of the base list of TestCheckList, "pal" and
// three letters, joined by a separator, with a prefix, a suffix and
// edits: with and without the final LF, CR LF lines, an empty last
// line, a byte order mark, an empty text, bytes that are not valid
// UTF-8. The restatement gives the words of List on every built-in list.
// - check_list: CheckList on the base list of 0 to 17 576 words, edited:
// the cases of TestCheckList; languages without an alphabet; white
// space, controls, invisible and unassigned code points; letters of
// other scripts that look like those of the alphabet, and marks; bytes
// that are not valid UTF-8; words that are one once normalized; and the
// order of the checks within a word and across words.
// - alphabet: CheckList of the base list of 2048 words whose first word
// is "pala", one code point and "zz": the result of every code point up
// to U+017F, and the SHA-256 of the lines "U+XXXX result\n" of every
// code point of planes 0, 1 and 14 but the surrogates.
// - generate: Generate on the Spanish list and on lists of 0 to
// 2^20 + 1 words, while it reads fixed streams instead of crypto/rand:
// the keystream of SeededRandomSource of lib/src/random.dart (ChaCha20
// under SHA-256(seed), zero nonce); SHA-256 counter blocks; and bytes
// that end, the seed of TestGenerate among them. Each case has the
// indices drawn (and the words, from the Spanish list), the bytes read
// and the next four of an endless stream; or the text of the error and
// the bytes read.
// - bits: Bits of sizes and counts, as the 64 bits of the float64 and its
// shortest decimal; and bits_digests, the SHA-256 of the 8 big-endian
// bytes of each Bits over ranges of sizes and counts, sizes outer.
//
// Binary values are lower-case hexadecimal, and the JSON is ASCII, so that
// the Dart constant is too. The output is the same on every run. It imports
// the package wordkey and reads wordkey/lists, so it runs in an export of
// datekeys-go made with git archive, without changing the repository, at
// 27a75ee, the branch v0.15 after the tag spec-v0.15:
//
// commit=$(git -C ../datekeys-go rev-parse 27a75ee)
// out=$PWD/test/vectors
// tmp=$(mktemp -d)
// git -C ../datekeys-go archive "$commit" | tar -x -C "$tmp"
// cp tool/wordlist_go_vectors.go "$tmp"
// (cd "$tmp" && go run ./wordlist_go_vectors.go -source "$commit" -out "$out")
// rm -rf "$tmp"
package main
import (
"bytes"
"crypto/sha256"
"encoding/binary"
"encoding/hex"
"encoding/json"
"flag"
"fmt"
"io"
"log"
"math"
"os"
"path/filepath"
"runtime"
"strconv"
"strings"
"golang.org/x/crypto/chacha20"
"g.activething.com/go/DateKeys/wordkey"
)
func check(err error) {
if err != nil {
log.Fatal(err)
}
}
func h(s string) string { return hex.EncodeToString([]byte(s)) }
func hexAll(ss []string) []string {
var out []string
for _, s := range ss {
out = append(out, h(s))
}
return out
}
func sum(s string) string {
b := sha256.Sum256([]byte(s))
return hex.EncodeToString(b[:])
}
// result is "ok", or the text of err.
func result(err error) string {
if err != nil {
return err.Error()
}
return "ok"
}
// ---------------------------------------------------------------------------
// Lists
// baseWord is word i of the base list of TestCheckList, "pal" and three
// letters: palaaa, palaab, …, up to i = 17 575, palzzz.
func baseWord(i int) string {
return "pal" + string([]rune{'a' + rune(i/676), 'a' + rune(i/26%26), 'a' + rune(i%26)})
}
// edit is the word at the index At of a base list replaced by Word, in hex.
type edit struct {
At int `json:"at"`
Word string `json:"word"`
}
func set(at int, word string) edit { return edit{at, h(word)} }
// edited is the base list of size words with the edits.
func edited(size int, sets []edit) []string {
l := make([]string, size)
for i := range l {
l[i] = baseWord(i)
}
for _, e := range sets {
b, err := hex.DecodeString(e.Word)
check(err)
l[e.At] = string(b)
}
return l
}
// listOf is the body of wordkey.List for the text of a list of the caller,
// restated because List reads only its built-in lists.
func listOf(lang, text string) ([]string, error) {
words := strings.Split(strings.TrimSuffix(text, "\n"), "\n")
if err := wordkey.CheckList(lang, words); err != nil {
return nil, fmt.Errorf("wordkey: the list %q: %w", lang, err)
}
return words, nil
}
type listInfo struct {
Lang string `json:"lang"`
Sha256 string `json:"sha256"`
Bytes int `json:"bytes"`
Words int `json:"words"`
}
func listInfos() []listInfo {
var out []listInfo
for _, lang := range wordkey.Languages() {
text, err := os.ReadFile(filepath.Join("wordkey", "lists", lang+".txt"))
check(err)
words, err := wordkey.List(lang)
check(err)
again, err := listOf(lang, string(text))
check(err)
if strings.Join(again, "\n") != strings.Join(words, "\n") {
log.Fatalf("wordkey/lists/%s.txt is not the list of List, or listOf is not its body", lang)
}
out = append(out, listInfo{lang, sum(string(text)), len(text), len(words)})
}
return out
}
type parseCase struct {
Name string `json:"name"`
Lang string `json:"lang"`
Size int `json:"size"`
Set []edit `json:"set,omitempty"`
Prefix string `json:"prefix"`
Sep string `json:"sep"`
Suffix string `json:"suffix"`
Sha256 string `json:"sha256"`
Words int `json:"words"`
Result string `json:"result"`
}
// parseOf reads with listOf the text prefix, the edited base list joined
// by sep, and suffix.
func parseOf(name, lang string, size int, sets []edit, prefix, sep, suffix string) parseCase {
text := prefix + strings.Join(edited(size, sets), sep) + suffix
words, err := listOf(lang, text)
return parseCase{name, lang, size, sets, h(prefix), h(sep), h(suffix), sum(text), len(words), result(err)}
}
func parseCases() []parseCase {
return []parseCase{
parseOf("2048 words, one per line, with a final LF", "es", 2048, nil, "", "\n", "\n"),
parseOf("without the final LF", "es", 2048, nil, "", "\n", ""),
parseOf("7776 words", "es", 7776, nil, "", "\n", "\n"),
parseOf("an empty last line: two final LF", "es", 2048, nil, "", "\n", "\n\n"),
parseOf("an empty first line", "es", 2048, nil, "\n", "\n", "\n"),
parseOf("CR LF lines", "es", 2048, nil, "", "\r\n", "\r\n"),
parseOf("CR LF lines, the last without", "es", 2048, nil, "", "\r\n", ""),
parseOf("CR lines", "es", 2048, nil, "", "\r", "\r"),
parseOf("a byte order mark", "es", 2048, nil, "\ufeff", "\n", "\n"),
parseOf("an empty text", "es", 0, nil, "", "\n", ""),
parseOf("an LF only", "es", 0, nil, "", "\n", "\n"),
parseOf("2047 words", "es", 2047, nil, "", "\n", "\n"),
parseOf("the words separated by spaces", "es", 2048, nil, "", " ", "\n"),
parseOf("a language without an alphabet", "xx", 2048, nil, "", "\n", "\n"),
parseOf("a byte that is not valid UTF-8", "es", 2048, []edit{set(6, "pal\xffaa")}, "", "\n", "\n"),
parseOf("a sequence cut at the end of the text", "es", 2048, []edit{set(2047, "palzz\xc3")}, "", "\n", ""),
parseOf("the same word twice", "es", 2048, []edit{set(2000, "palaaa")}, "", "\n", "\n"),
parseOf("a capital letter", "es", 2048, []edit{set(100, "Paldww")}, "", "\n", "\n"),
parseOf("letters of the alphabet", "es", 2048, []edit{set(0, "\u00e1\u00e9\u00ed\u00f3\u00fa\u00fc\u00f1")}, "", "\n", "\n"),
}
}
type checkCase struct {
Name string `json:"name"`
Lang string `json:"lang"`
Size int `json:"size"`
Set []edit `json:"set,omitempty"`
Append []string `json:"append,omitempty"`
Sha256 string `json:"sha256"`
Result string `json:"result"`
}
// checkOf runs CheckList on the edited base list of size words followed by
// appends. Sha256 is that of the words joined by LF.
func checkOf(name, lang string, size int, sets []edit, appends ...string) checkCase {
l := append(edited(size, sets), appends...)
return checkCase{name, lang, size, sets, hexAll(appends), sum(strings.Join(l, "\n")), result(wordkey.CheckList(lang, l))}
}
func checkCases() []checkCase {
at := func(i int, w string) []edit { return []edit{set(i, w)} }
out := []checkCase{
// TestCheckList, in its order.
checkOf("the base list", "es", 2048, nil),
checkOf("a language without an alphabet", "xx", 2048, nil),
checkOf("2047 words", "es", 2047, nil),
checkOf("two words in a line", "es", 2048, at(5, "dos palabras")),
checkOf("two spaces", "es", 2048, at(5, " ")),
checkOf("two letters", "es", 2048, at(5, "mi")),
checkOf("ZWSP", "es", 2048, at(5, "casa\u200b")),
checkOf("a capital", "es", 2048, at(5, "Palaaf")),
checkOf("a Cyrillic letter that looks like a Latin c", "es", 2048, at(5, "\u0441asa")),
checkOf("a digit", "es", 2048, at(5, "pal1")),
checkOf("the CR of a CR LF line", "es", 2048, at(5, "palaaf\r")),
checkOf("the same word", "es", 2048, at(5, baseWord(4))),
checkOf("the same word once normalized", "es", 2048, at(5, "pala\u00e1e")),
checkOf("papa after pap\u00e1", "es", 2048, at(0, "pap\u00e1"), "papa"),
// The language and the size.
checkOf("no words", "es", 0, nil),
checkOf("no words in a language without an alphabet", "xx", 0, nil),
checkOf("the empty language", "", 2048, nil),
checkOf("ES, in capitals", "ES", 2048, nil),
checkOf("es and a space", "es ", 2048, nil),
checkOf("espa\u00f1ol", "espa\u00f1ol", 2048, nil),
checkOf("2049 words", "es", 2049, nil),
checkOf("7776 words", "es", 7776, nil),
checkOf("17576 words", "es", 17576, nil),
// One word of three letters or more, once normalized.
checkOf("an empty word", "es", 2048, at(5, "")),
checkOf("a word of three letters", "es", 2048, at(5, "pal")),
checkOf("\u00f1u, two letters", "es", 2048, at(5, "\u00f1u")),
checkOf("\u00f1u\u00f1, three letters", "es", 2048, at(5, "\u00f1u\u00f1")),
checkOf("a space before", "es", 2048, at(5, " palaaf")),
checkOf("a space after", "es", 2048, at(5, "palaaf ")),
checkOf("a tab inside", "es", 2048, at(5, "pal\taf")),
checkOf("NBSP inside", "es", 2048, at(5, "pal\u00a0af")),
checkOf("NEL inside", "es", 2048, at(5, "pal\u0085af")),
checkOf("the ideographic space inside", "es", 2048, at(5, "pal\u3000af")),
checkOf("LS inside", "es", 2048, at(5, "pal\u2028af")),
checkOf("an LF inside", "es", 2048, at(5, "pal\naf")),
checkOf("two runes and ZWSP: the count before the runes", "es", 2048, at(5, "p\u200b")),
checkOf("three letters with a mark that goes: two", "es", 2048, at(5, "pa\u0301")),
checkOf("a\u0301b, two letters once normalized", "es", 2048, at(5, "a\u0301b")),
// The runes of the normalized word.
checkOf("a control", "es", 2048, at(5, "pal\x01af")),
checkOf("DEL", "es", 2048, at(5, "pal\x7faf")),
checkOf("a C1 control", "es", 2048, at(5, "pal\u0080af")),
checkOf("a control and a capital: the runes before the alphabet", "es", 2048, at(5, "Pal\x01af")),
checkOf("a soft hyphen", "es", 2048, at(5, "pa\u00ad")),
checkOf("ZWJ", "es", 2048, at(5, "pal\u200daf")),
checkOf("a byte order mark", "es", 2048, at(5, "\ufeffpalaaf")),
checkOf("an unassigned code point", "es", 2048, at(5, "pal\u0378")),
checkOf("a noncharacter", "es", 2048, at(5, "pal\ufffe")),
checkOf("a tag", "es", 2048, at(5, "pal\U000e0041af")),
// The alphabet, on the word as the list writes it.
checkOf("the letters of the alphabet", "es", 2048, at(5, "abcdefghijklmnopqrstuvwxyz\u00e1\u00e9\u00ed\u00f3\u00fa\u00fc\u00f1")),
checkOf("an acute accent in NFD", "es", 2048, at(5, "pala\u0301e")),
checkOf("a diaeresis in NFD", "es", 2048, at(5, "palaa\u0308")),
checkOf("a tilde in NFD", "es", 2048, at(5, "pan\u0303o")),
checkOf("\u00c1, a capital with an accent", "es", 2048, at(5, "\u00c1baco")),
checkOf("\u00d1, a capital", "es", 2048, at(5, "\u00d1and\u00fa")),
checkOf("a grave accent", "es", 2048, at(5, "pal\u00e0a")),
checkOf("a circumflex", "es", 2048, at(5, "pal\u00e2a")),
checkOf("\u00e4", "es", 2048, at(5, "pal\u00e4a")),
checkOf("\u00e7", "es", 2048, at(5, "pal\u00e7a")),
checkOf("\u00f6", "es", 2048, at(5, "pal\u00f6a")),
checkOf("\u00fd", "es", 2048, at(5, "pal\u00fda")),
checkOf("\u00df", "es", 2048, at(5, "pal\u00dfa")),
checkOf("a Cyrillic a", "es", 2048, at(5, "pal\u0430a")),
checkOf("a Greek omicron", "es", 2048, at(5, "pal\u03bfa")),
checkOf("a full-width a", "es", 2048, at(5, "pal\uff41a")),
checkOf("a mathematical bold a", "es", 2048, at(5, "pal\U0001d41aa")),
checkOf("the Kelvin sign", "es", 2048, at(5, "pal\u212aa")),
checkOf("a dotless i", "es", 2048, at(5, "pal\u0131a")),
checkOf("a hyphen", "es", 2048, at(5, "pal-af")),
checkOf("an apostrophe", "es", 2048, at(5, "pal'af")),
checkOf("a backtick and a brace, around a to z", "es", 2048, at(5, "pal`{")),
checkOf("an emoji", "es", 2048, at(5, "pal\U0001f600")),
// Bytes that are not valid UTF-8: U+FFFD, which no alphabet holds.
checkOf("a byte that is not valid UTF-8", "es", 2048, at(5, "pal\xffaa")),
checkOf("a surrogate in UTF-8 bytes", "es", 2048, at(5, "\xed\xa0\x80aaa")),
checkOf("a sequence cut short", "es", 2048, at(5, "pala\xc3")),
checkOf("an overlong NUL", "es", 2048, at(5, "\xc0\x80aaa")),
checkOf("U+FFFD itself", "es", 2048, at(5, "pal\ufffdaa")),
// The same word once normalized, and the order across words.
checkOf("a\u00f1o and ano", "es", 2048, []edit{set(5, "a\u00f1o"), set(6, "ano")}),
checkOf("ping\u00fcino and pinguino", "es", 2048, []edit{set(5, "pinguino"), set(9, "ping\u00fcino")}),
checkOf("the same word, the first of the list last", "es", 2048, nil, "palaaa"),
checkOf("a capital of a word of the list: the alphabet before the same word", "es", 2048, at(5, "PALAAE")),
checkOf("an error at line 4 and another at line 6", "es", 2048, []edit{set(3, "Mal"), set(5, "pa")}),
checkOf("an error at the last line", "es", 2048, at(2047, "x")),
checkOf("an empty last line", "es", 2048, nil, ""),
}
return out
}
// ---------------------------------------------------------------------------
// The alphabet of every code point
type runeResult struct {
Rune int `json:"rune"`
Result string `json:"result"`
}
type alphabetHead struct {
Word string `json:"word"`
Planes []int `json:"planes"`
Count int `json:"count"`
Sha256 string `json:"sha256"`
Ok []string `json:"ok"`
}
func alphabet() (alphabetHead, []runeResult) {
planes := []int{0, 1, 14}
l := edited(2048, nil)
digest := sha256.New()
var results []runeResult
var ok []string
count := 0
for _, p := range planes {
for r := rune(p << 16); r <= rune(p<<16|0xffff); r++ {
if r >= 0xd800 && r <= 0xdfff {
continue
}
l[0] = "pala" + string(r) + "zz"
res := result(wordkey.CheckList("es", l))
fmt.Fprintf(digest, "U+%04X %s\n", r, res)
count++
if r <= 0x17f {
results = append(results, runeResult{int(r), res})
}
if res == "ok" {
ok = append(ok, string(r))
}
}
}
head := alphabetHead{"pala, the code point and zz, the first word of the base list of 2048", planes, count, hex.EncodeToString(digest.Sum(nil)), ok}
return head, results
}
// ---------------------------------------------------------------------------
// Generate
// seeded is the keystream of SeededRandomSource of lib/src/random.dart:
// ChaCha20 under SHA-256(seed), with a zero nonce, from block 0.
type seeded struct{ c *chacha20.Cipher }
func newSeeded(seed string) *seeded {
key := sha256.Sum256([]byte(seed))
c, err := chacha20.NewUnauthenticatedCipher(key[:], make([]byte, chacha20.NonceSize))
check(err)
return &seeded{c}
}
func (s *seeded) Read(p []byte) (int, error) {
clear(p)
s.c.XORKeyStream(p, p)
return len(p), nil
}
// counter is SHA-256(label ‖ i) for i = 0, 1, …, a 32-bit big-endian
// counter, one block after the other.
type counter struct {
label []byte
i uint32
block []byte
}
func (c *counter) Read(p []byte) (int, error) {
for n := 0; n < len(p); {
if len(c.block) == 0 {
b := sha256.Sum256(binary.BigEndian.AppendUint32(bytes.Clone(c.label), c.i))
c.block = b[:]
c.i++
}
k := copy(p[n:], c.block)
c.block = c.block[k:]
n += k
}
return len(p), nil
}
// counting counts the bytes read from r.
type counting struct {
r io.Reader
n int
}
func (c *counting) Read(p []byte) (int, error) {
n, err := c.r.Read(p)
c.n += n
return n, err
}
// stream is what a case of generate reads instead of crypto/rand: kind
// seeded, the keystream of Seed; counter, the blocks of the label Seed; or
// bytes, Hex repeated Repeat times, which end.
type stream struct {
Kind string `json:"kind"`
Seed string `json:"seed,omitempty"`
Hex string `json:"hex,omitempty"`
Repeat int `json:"repeat,omitempty"`
}
func (s stream) reader() io.Reader {
switch s.Kind {
case "seeded":
return newSeeded(s.Seed)
case "counter":
return &counter{label: []byte(s.Seed)}
case "bytes":
b, err := hex.DecodeString(s.Hex)
check(err)
return bytes.NewReader(bytes.Repeat(b, s.Repeat))
}
log.Fatalf("stream kind %q", s.Kind)
return nil
}
func seed(s string) stream { return stream{Kind: "seeded", Seed: s} }
func given(hexBytes string, repeat int) stream {
return stream{Kind: "bytes", Hex: hexBytes, Repeat: repeat}
}
type generateCase struct {
Name string `json:"name"`
List string `json:"list,omitempty"`
Size int `json:"size"`
N int `json:"n"`
Stream stream `json:"stream"`
Indices []int `json:"indices,omitempty"`
Words []string `json:"words,omitempty"`
Read int `json:"read"`
Next string `json:"next,omitempty"`
Error string `json:"error,omitempty"`
}
// numbers is a list of size words, "0", "1", …: Generate reads only its
// length and the words it draws, whose indices they are.
func numbers(size int) []string {
l := make([]string, size)
for i := range l {
l[i] = strconv.Itoa(i)
}
return l
}
func generateOf(name, listName string, list []string, n int, s stream) generateCase {
r := &counting{r: s.reader()}
words, err := wordkey.Generate(list, n, r)
c := generateCase{Name: name, List: listName, Size: len(list), N: n, Stream: s, Read: r.n}
if err != nil {
c.Error = err.Error()
return c
}
index := make(map[string]int, len(list))
for i, w := range list {
index[w] = i
}
for _, w := range words {
c.Indices = append(c.Indices, index[w])
}
if listName != "" && n <= 24 {
c.Words = words
}
if s.Kind != "bytes" {
next := make([]byte, 4)
_, err := io.ReadFull(r.r, next)
check(err)
c.Next = hex.EncodeToString(next)
}
return c
}
func generateCases() []generateCase {
es, err := wordkey.List("es")
check(err)
var out []generateCase
add := func(c generateCase) { out = append(out, c) }
// The Spanish list, from the keystream of a seed and from SHA-256
// counter blocks.
for _, n := range []int{6, wordkey.DefaultCount, 8, 12, 24} {
add(generateOf(fmt.Sprintf("%d words of the Spanish list", n), "es", es, n, seed(fmt.Sprintf("wordkey.Generate es %d", n))))
}
for _, n := range []int{6, wordkey.DefaultCount, 12} {
add(generateOf(fmt.Sprintf("%d words of the Spanish list, SHA-256 counter blocks", n), "es", es, n, stream{Kind: "counter", Seed: fmt.Sprintf("wordkey.Generate counter %d", n)}))
}
add(generateOf("3888 words of the Spanish list, the most", "es", es, len(es)/2, seed("wordkey.Generate es 3888")))
// Lists of every number of bytes of a draw of crypto/rand.Int, with and
// without draws again.
for _, size := range []int{12, 13, 16, 17, 100, 255, 256, 257, 2047, 2048, 2049, 4096, 7775, 7776, 7777, 8192, 65535, 65536, 65537, 100000, 1<<20 + 1} {
add(generateOf(fmt.Sprintf("6 words of %d", size), "", numbers(size), 6, seed(fmt.Sprintf("wordkey.Generate %d", size))))
}
add(generateOf("12 words of 65537, SHA-256 counter blocks", "", numbers(65537), 12, stream{Kind: "counter", Seed: "wordkey.Generate counter 65537"}))
// The most words of small lists: indices drawn twice, again and again.
for _, size := range []int{12, 13, 17, 100} {
add(generateOf(fmt.Sprintf("%d words of %d, the most", size/2, size), "", numbers(size), size/2, seed(fmt.Sprintf("wordkey.Generate %d most", size))))
}
// Bytes that end.
add(generateOf("the seed of TestGenerate: two indices only, then EOF", "es", es, 6, given("0701c821", 64)))
add(generateOf("twelve bytes, six indices", "es", es, 6, given("000000010002000300040005", 1)))
add(generateOf("8191 and 8032 drawn again, 0 drawn twice, 7775 the last index", "es", es, 6, given("1fff00001f601e5f00000001000200030004", 1)))
add(generateOf("the three bits above the 13 of 7775 are cleared", "es", es, 6, given("e000ffff21002200230024002500", 1)))
add(generateOf("the bytes end inside a draw", "es", es, 6, given("0000000100", 1)))
add(generateOf("no bytes", "es", es, 6, given("", 0)))
// The arguments, refused before anything is read.
for _, n := range []int{5, 0, -1, len(es)/2 + 1, len(es)} {
add(generateOf(fmt.Sprintf("%d words of the Spanish list", n), "es", es, n, seed("never read")))
}
for _, c := range []struct{ size, n int }{{11, 6}, {0, 6}, {0, 5}, {12, 7}, {13, 7}} {
add(generateOf(fmt.Sprintf("%d words of %d", c.n, c.size), "", numbers(c.size), c.n, seed("never read")))
}
return out
}
// ---------------------------------------------------------------------------
// Bits
type bitsCase struct {
Size int `json:"size"`
Count int `json:"count"`
Bits string `json:"bits"`
Hex string `json:"hex"`
}
func bitsOf(size, count int) bitsCase {
b := wordkey.Bits(size, count)
return bitsCase{size, count, strconv.FormatFloat(b, 'g', -1, 64), fmt.Sprintf("%016x", math.Float64bits(b))}
}
func bitsCases() []bitsCase {
var out []bitsCase
for _, size := range []int{1, 2, 3, 5, 12, 13, 100, 255, 256, 257, 2047, 2048, 2049, 4096, 7775, 7776, 7777, 8192, 10000, 65535, 65536, 65537, 1 << 20, 1<<31 - 1, 1 << 31, 1<<32 - 1, 1 << 32, 1<<32 + 1, 1<<52 + 1, 1<<53 - 1, 1 << 53} {
for _, count := range []int{0, 1, 2, 6, 7, 8, 12} {
out = append(out, bitsOf(size, count))
}
}
// Go adds log2 of 0, −Inf, and of a negative number, NaN, and adds
// nothing for a negative count.
for _, c := range [][2]int{{0, 1}, {-1, 1}, {7776, -1}, {0, 0}} {
out = append(out, bitsOf(c[0], c[1]))
}
return out
}
type bitsDigest struct {
Name string `json:"name"`
Sizes [2]int `json:"sizes"`
Counts [2]int `json:"counts"`
Sha256 string `json:"sha256"`
}
func digestOf(name string, sizes, counts [2]int) bitsDigest {
d := sha256.New()
var b [8]byte
for size := sizes[0]; size <= sizes[1]; size++ {
for count := counts[0]; count <= counts[1]; count++ {
binary.BigEndian.PutUint64(b[:], math.Float64bits(wordkey.Bits(size, count)))
d.Write(b[:])
}
}
return bitsDigest{name, sizes, counts, hex.EncodeToString(d.Sum(nil))}
}
func bitsDigests() []bitsDigest {
return []bitsDigest{
digestOf("one word of 1 to 65536: the log2 of Go", [2]int{1, 65536}, [2]int{1, 1}),
digestOf("6 to 8 words of 2048 to 10000", [2]int{2048, 10000}, [2]int{6, 8}),
digestOf("0 to 600 words of 7776", [2]int{7776, 7776}, [2]int{0, 600}),
digestOf("1 word of 2^32 - 4096 to 2^32 + 4096", [2]int{1<<32 - 4096, 1<<32 + 4096}, [2]int{1, 1}),
}
}
// ---------------------------------------------------------------------------
// JSON
// ascii escapes every code point above U+007F of JSON as \uXXXX, in UTF-16
// for those above U+FFFF: the same JSON, in ASCII.
func ascii(b []byte) []byte {
var out bytes.Buffer
for _, r := range string(b) {
switch {
case r < 0x80:
out.WriteByte(byte(r))
case r < 0x10000:
fmt.Fprintf(&out, `\u%04x`, r)
default:
r -= 0x10000
fmt.Fprintf(&out, `\u%04x\u%04x`, 0xd800+(r>>10), 0xdc00+(r&0x3ff))
}
}
return out.Bytes()
}
func marshal(v any) []byte {
var b bytes.Buffer
enc := json.NewEncoder(&b)
enc.SetEscapeHTML(false)
check(enc.Encode(v))
return ascii(bytes.TrimSuffix(b.Bytes(), []byte("\n")))
}
// list writes the items of a JSON array one per line.
func list[T any](w *bytes.Buffer, key string, items []T, last bool) {
fmt.Fprintf(w, " %q: [\n", key)
for i, it := range items {
w.WriteString(" ")
w.Write(marshal(it))
if i < len(items)-1 {
w.WriteByte(',')
}
w.WriteByte('\n')
}
w.WriteString(" ]")
if !last {
w.WriteByte(',')
}
w.WriteByte('\n')
}
func main() {
outDir := flag.String("out", "", "the directory test/vectors of datekeys-dart")
source := flag.String("source", "", "the commit of datekeys-go")
flag.Parse()
if *outDir == "" || *source == "" {
log.Fatal("usage: go run wordlist_go_vectors.go -source COMMIT -out DIR")
}
lists := listInfos()
parse := parseCases()
checks := checkCases()
alpha, runes := alphabet()
gen := generateCases()
bits := bitsCases()
digests := bitsDigests()
var w bytes.Buffer
w.WriteString("{\n")
head := []struct {
key string
v any
}{
{"description", "The word lists of package wordkey of the Go reference (generate.go, spec \u00a738.1) for lib/src/wordlist.dart; see tool/wordlist_go_vectors.go. " +
"Binary values and words are lower-case hex. The base list is word i = pal and three letters, a + i/676, a + i/26%26 and a + i%26; a case of size n takes its first n words, replaces the word at each at of set, and adds those of append. " +
"lists: the SHA-256, bytes and words of each built-in list. parse: the body of List on the text prefix, the base list joined by sep, and suffix (sha256 is that of the text): the words read, and ok or the error. " +
"check_list: CheckList, ok or the error (sha256 is that of the words joined by LF). alphabet: CheckList of the base list of 2048 whose first word is pala, a code point and zz, for each code point of the planes but the surrogates: the SHA-256 of the lines U+XXXX result, the code points of the words it accepts, and alphabet_results, each result up to U+017F. " +
"generate: Generate on the list of the Spanish words (list es) or on a list of size words whose indices they are, reading a stream: seeded, the keystream of SeededRandomSource of the seed; counter, SHA-256 of the label and a 32-bit big-endian counter from 0, block after block; bytes, hex repeated repeat times, which end. The indices drawn, the bytes read and the next four of an endless stream, or the error. " +
"bits: Bits, its shortest decimal and the hex of its float64. bits_digests: the SHA-256 of the 8 big-endian bytes of each Bits of the sizes and counts, inclusive, sizes outer."},
{"generator", "tool/wordlist_go_vectors.go, " + runtime.Version()},
{"source", *source},
{"default_count", wordkey.DefaultCount},
{"min_list_size", wordkey.MinListSize},
{"min_words", wordkey.MinWords},
{"min_letters", wordkey.MinLetters},
{"alphabet", alpha},
}
for _, kv := range head {
fmt.Fprintf(&w, " %q: %s,\n", kv.key, marshal(kv.v))
}
list(&w, "lists", lists, false)
list(&w, "parse", parse, false)
list(&w, "check_list", checks, false)
list(&w, "alphabet_results", runes, false)
list(&w, "generate", gen, false)
list(&w, "bits", bits, false)
list(&w, "bits_digests", digests, true)
w.WriteString("}\n")
var parsed any
if err := json.Unmarshal(w.Bytes(), &parsed); err != nil {
log.Fatalf("the JSON does not parse: %v", err)
}
if bytes.Contains(w.Bytes(), []byte("'''")) {
log.Fatal("the JSON holds three quotes")
}
path := filepath.Join(*outDir, "wordlist_vectors.json")
check(os.WriteFile(path, w.Bytes(), 0o644))
fmt.Printf("wrote %s, %d bytes: %d lists, %d texts, %d lists checked, %d code points, %d draws, %d bits, %d digests\n",
path, w.Len(), len(lists), len(parse), len(checks), alpha.Count, len(gen), len(bits), len(digests))
dart := "// Generated by tool/wordlist_go_vectors.go from wordlist_vectors.json, for\n" +
"// the tests that also run compiled to JavaScript, where no file can be\n" +
"// read. Do not edit.\n\n" +
"/// The text of test/vectors/wordlist_vectors.json.\n" +
"const wordlistVectorsJson = r'''\n" + w.String() + "''';\n"
dpath := filepath.Join(*outDir, "wordlist_vectors.g.dart")
check(os.WriteFile(dpath, []byte(dart), 0o644))
fmt.Printf("wrote %s\n", dpath)
}
Loading…
Cancel
Save

Powered by TurnKey Linux.