Random words: readWordList, checkWordList, generateWords and wordBits

wordlist.ts does what wordkey.Generate, CheckList and Bits of Go do at
27a75ee, with their texts: generateWords draws different words, 7 by
default, with crypto.getRandomValues; wordBits is their strength;
checkWordList refuses a list of fewer than 2048 words, with two words that
are one once normalized or with a character outside the alphabet of its
language, which the code gives and not the list. readWordList takes a list
only with the SHA-256 pinned for its language, in UTF-8 and accepted by
checkWordList: no list is trusted, not even those of DateKeys.

wordkey.ts gains wordRules, which loads the Unicode tables once for a
caller that reads many words; normalizeWords and checkWords use it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
main
dev 14 hours ago
parent a077e85e47
commit 704838016b

@ -4,7 +4,12 @@ Cambios notables de la librería TypeScript y de la página. El proyecto usa ver
## 0.5.0 — sin publicar
Nada todavía.
### Las palabras al azar de la llave de palabras (07-10-2026)
El SHOULD de §38.1, ofrecer palabras al azar de una lista pública, como hace `datekeys encrypt -new-words` desde `c49c67c` de `datekeys-go`. No cambia ningún formato ni la derivación.
- **Las listas de `datekeys-go`.** `scripts/sync-testdata.mjs` copia del mismo commit `wordkey/lists` en `wordlists/`, con su `SOURCE.json`, y `check` comprueba las dos copias. `testdata` y `wordlists` están en `27a75ee` de `datekeys-go`, de la rama `v0.15` después del tag `spec-v0.15`: los ficheros de `testdata` son los del tag, y `wordlists` trae la lista española, de 7 776 palabras, un borrador aún sin revisar, con licencia CC BY-SA 4.0.
- **`wordlist.ts`**, como `wordkey.Generate`, `CheckList` y `Bits` de Go en `27a75ee`, con sus textos: `generateWords` sortea palabras distintas, 7 por defecto, con `crypto.getRandomValues`; `wordBits` da su fuerza, 90 bits para 7 de 7 776; `checkWordList` rechaza una lista de menos de 2 048 palabras, con dos que son una al normalizarlas o con un carácter que no es una letra del alfabeto de su idioma, que da el código y no la lista (para `es`, de la `a` a la `z`, `á`, `é`, `í`, `ó`, `ú`, `ü` y `ñ`); y `readWordList` solo acepta una lista con el SHA-256 fijado para su idioma, que sea UTF-8 y que `checkWordList` acepte. `wordkey.ts` gana `wordRules`, que carga una vez las tablas de Unicode para quien lee muchas palabras.
## 0.4.0 — 7 de octubre de 2026

@ -81,6 +81,8 @@ La inspección (pasos 1 a 8) no importa ninguna dependencia. Funciona en navegad
| `cms.ts` | La firma CMS (RFC 5652) de `alg` 2 y el token RFC 3161 del sello (§29.10, §29.11), con el orden de comprobaciones y los resultados de Go y la tabla cerrada de algoritmos: RSA PKCS #1 v1.5 y PSS con `BigInt`, de 2048 a 4096 bits y con un módulo impar, y ECDSA sobre P-256, P-384 y P-521 con la aritmética de `@noble/curves`, solo con el punto sin comprimir. Lee el certificado campo a campo con el perfil del §29.10 del borrador v0.12, como Go: a quién nombra (su `givenName` y su `surname` antes que su `commonName`), el emisor que dice (su `commonName` o su `organizationName`), su validez y su clave; nunca comprueba quién lo emitió ni si se revocó. Compara los OID por los bytes de su DER, acepta un SET OF que repite un elemento y cuenta como uno un certificado repetido, y un nombre conserva un U+FEFF inicial, como en Go | `internal/cms` |
| `securitycms.ts` | Los veredictos de una firma de `alg` 2, F1, F2, F5 y F6, con cada firmante requerido y ajeno nombrado como el §29.7 del borrador v0.12 (su nombre si cumple las reglas del autor declarado, tiene como mucho 64 puntos de código y no lleva dos espacios seguidos, y si no el SHA-256 del certificado; su emisor, o el SHA-256 de su `Name`; y la autoridad de su sello), y los de un sello, S1 a S5, con su autoridad y t. `encodeSigners` escribe `SIGNERS` para el escritor, ordenado y con los textos de Go | `capsule/signature2.go` |
| `note.ts` | La nota pública (§24.1): la extensión `datekeys.note` de la cabecera, de 1 a 1024 bytes de UTF-8 que cumplen las reglas del autor declarado, con los textos y el orden de `extension.CheckNote`. Una cadena con un sustituto suelto se rechaza: nunca se escribe con U+FFFD. `checkNoteData` comprueba los bytes de una nota en el orden de Go: la longitud, el UTF-8 y las reglas. `publicNote` la lee como Go, con un U+FEFF inicial que la deja inservible, y `unusableNote` distingue una nota inservible de ninguna | `extension` (`CheckNote`, `NewNote`, `Note`), `capsule` (`Header.UnusableNote`) |
| `wordkey.ts` | La llave de palabras (§38.1): `normalizeWords` (NFD con las tablas de Unicode 18.0.0, sin las marcas U+0300 a U+036F, la minúscula simple de cada punto de código, partido por los espacios de la lista), `checkWords` (al menos 6 palabras distintas de 3 letras o más, sin controles, invisibles ni puntos sin asignar), `wordRules` (las dos, con las tablas cargadas una vez) y `wordKey`, PBKDF2-SHA256 de 600 000 rondas de Web Crypto con la sal de la cadena, la ronda y `capsule_id`. `quickWords`, `hiddenCodePoint` y `countedWords` leen con las tablas de la plataforma, para un formulario que no puede esperar a las de Unicode 18.0.0 | `wordkey.Normalize`, `Check`, `Key` |
| `wordlist.ts` | Las palabras al azar de la llave de palabras (el SHOULD de §38.1), con los textos de Go: `generateWords` sortea `DEFAULT_WORD_COUNT` palabras distintas, 7, con `randomIndex` y `crypto.getRandomValues`; `wordBits` da su fuerza, 90 bits para 7 de 7 776; `checkWordList` rechaza una lista de menos de 2 048 palabras, con dos que son una al normalizarlas o con un carácter que no es una letra del alfabeto de su idioma, que da el código y no la lista (para `es`, de la `a` a la `z`, `á`, `é`, `í`, `ó`, `ú`, `ü` y `ñ`): una letra cirílica que parece latina se volvería a escribir con la latina, y la cápsula no se abriría; y `readWordList` solo acepta una lista con el SHA-256 fijado para su idioma en `WORD_LIST_SHA256`, que sea UTF-8 y que `checkWordList` acepte. Ninguna lista se da por buena, tampoco las de DateKeys | `wordkey.Generate`, `CheckList`, `Bits`, `List` |
| `pathrule.ts`, `pathrule-tables.ts` | Las reglas de las rutas y de los textos del head (§29.5, §29.6), de R1 a R10 con R4b, R6b, R6c y R9, NFD, el pliegue y la clave de R7, con los textos de error de Go. Nunca usa `normalize`, `toLowerCase`, `localeCompare`, `Intl` ni las clases `\p{…}`, cuya versión de Unicode cambia con el motor: las tablas de Unicode 18.0.0 y WindowsBestFit las genera `datekeys-go` (`pathrule/gen -ts`), y una prueba recalcula su digest | `internal/pathrule` |
| `open3.ts`, `sink.ts` | El paso 17 del formato 3 en sus subpasos 17.2 a 17.8, con la precedencia de Go: un fallo de `age` o un texto en claro que no mide P prevalecen, manda el primer subpaso que falla, y los códigos distintos de `ERR_INTEGRITY` solo se dan tras leer `PAYLOAD_AGE` hasta el final. `Sink` (`begin`, `create`, `commit`, `abort`) recibe los ficheros y solo los publica en el paso 18; `MemorySink` los guarda en memoria, que crece con los bytes recibidos y nunca con los tamaños que declara el head | `capsule/open3.go` |
| `framing.ts` | Prelude DKC1 (16 bytes), con el formato de la cápsula, de 1 a 3, y DKK1 (12 bytes) en el orden de §23 y §40, longitudes de 1 byte hasta los límites de §57, y troceo de secciones | `capsule/framing.go` |
@ -455,7 +457,7 @@ Todas las versiones se fijan exactas y `package-lock.json` se versiona. `.npmrc`
npm run testdata:sync
```
Copia `testdata/` de `../datekeys-go` en `HEAD`, leyendo los blobs con git para no arrastrar cambios sin commit, y escribe `testdata/SOURCE.json` con el commit completo y el SHA-256 de cada fichero. Para fijar otro commit: `node scripts/sync-testdata.mjs sync --commit <rev>`.
Copia `testdata/` de `../datekeys-go` en `HEAD`, leyendo los blobs con git para no arrastrar cambios sin commit, y escribe `testdata/SOURCE.json` con el commit completo y el SHA-256 de cada fichero. Del mismo commit copia `wordkey/lists` en `wordlists/`, con su `wordlists/SOURCE.json`: las listas de las palabras al azar, que publica la página. Para fijar otro commit: `node scripts/sync-testdata.mjs sync --commit <rev>`.
```bash
npm run testdata:check
@ -463,10 +465,12 @@ npm run testdata:check
Comprueba que los ficheros coinciden con `SOURCE.json`, sin faltantes ni sobrantes, y vuelve a leerlos del repositorio Go en el commit registrado. Sin la opción `--against`, `node scripts/sync-testdata.mjs check` verifica solo la copia local. Ambos comandos usan solo Node y git.
`.gitattributes` marca `testdata/**` como binario para que git no altere ningún byte.
`.gitattributes` marca `testdata/**` y `wordlists/**` como binarios para que git no altere ningún byte.
Copia actual: la de `testdata/SOURCE.json` (el tag `spec-v0.15` de `datekeys-go`, `fe50885`), que se sincroniza con `node scripts/sync-testdata.mjs sync --commit spec-v0.15`. Añade `vectors/release.json`, los ficheros de `releases/` y el campo `source` de `mutations.json`.
Copia actual: la de `testdata/SOURCE.json`, `datekeys-go` en `27a75ee`, de la rama `v0.15` después del tag `spec-v0.15`, que se sincroniza con `node scripts/sync-testdata.mjs sync --commit 27a75ee`. Su `testdata` es el del tag (`fe50885`), que añade `vectors/release.json`, los ficheros de `releases/` y el campo `source` de `mutations.json`; y trae `wordkey/lists`, con la lista española.
## Licencia
Apache-2.0 ([LICENSE](LICENSE)), como la librería Go de referencia. El código que se derive de terceros conserva su aviso de copyright y licencia en el propio fichero: el núcleo IBE (`ibe.ts`), derivado de `tlock-js` (Apache-2.0 OR MIT, usado bajo MIT), y `bech32.ts`, portado de `age` (MIT). El sitio publica esos avisos y los de sus paquetes npm en `licenses.txt`.
Las listas de palabras de `wordlists/` no son código de este proyecto ni tienen su licencia: la española es CC BY-SA 4.0, una adaptación de FrequencyWords de Hermit Dave. Su `README.md`, copiado de `datekeys-go`, dice de dónde sale, cómo se hizo y su licencia.

@ -42,6 +42,35 @@ function split(text: string, nfd: (s: string) => string, lower: (ch: string) =>
return words;
}
/** The reading and the checks of words that need the tables of Unicode 18.0.0, once they have loaded (wordRules). */
export interface WordRules {
/** The words of a text, as normalizeWords returns them. */
readonly normalize: (text: string) => string[];
/**
* The error of Go wordkey for the first code point of a word that a
* person cannot see: a control, a Default_Ignorable_Code_Point or a code
* point unassigned in Unicode 18.0.0; undefined when there is none.
*/
readonly hidden: (word: string) => string | undefined;
}
/** The rules of the words, with the tables of pathrule.ts loaded once, for a caller that reads many words. */
export async function wordRules(): Promise<WordRules> {
const [{ isAssigned, isDefaultIgnorable, lower, nfd }, { UNICODE_VERSION }] = await Promise.all([import('./pathrule.ts'), import('./pathrule-tables.ts')]);
return {
normalize: (text) => split(text, nfd, (ch) => String.fromCodePoint(lower(ch.codePointAt(0)!))),
hidden: (word) => {
for (const ch of word) {
const r = ch.codePointAt(0)!;
if (r <= 0x1f || (r >= 0x7f && r <= 0x9f)) return `wordkey: the words hold the control character U+${hex4(r)}`;
if (isDefaultIgnorable(r)) return `wordkey: the words hold the invisible character U+${hex4(r)}`;
if (!isAssigned(r)) return `wordkey: the words hold U+${hex4(r)}, unassigned in Unicode ${UNICODE_VERSION}`;
}
return undefined;
},
};
}
/**
* The words of a text, as Go wordkey.Normalize reads them: its NFD by the
* tables of pathrule.ts, without the combining marks U+0300 to U+036F, each
@ -49,8 +78,7 @@ function split(text: string, nfd: (s: string) => string, lower: (ch: string) =>
* white space.
*/
export async function normalizeWords(text: string): Promise<string[]> {
const { nfd, lower } = await import('./pathrule.ts');
return split(text, nfd, (ch) => String.fromCodePoint(lower(ch.codePointAt(0)!)));
return (await wordRules()).normalize(text);
}
/**
@ -94,14 +122,10 @@ export function countedWords(words: readonly string[]): number {
* wordkey.Check.
*/
export async function checkWords(words: readonly string[]): Promise<void> {
const [{ isAssigned, isDefaultIgnorable }, { UNICODE_VERSION }] = await Promise.all([import('./pathrule.ts'), import('./pathrule-tables.ts')]);
const { hidden } = await wordRules();
for (const w of words) {
for (const ch of w) {
const r = ch.codePointAt(0)!;
if (r <= 0x1f || (r >= 0x7f && r <= 0x9f)) throw new Error(`wordkey: the words hold the control character U+${hex4(r)}`);
if (isDefaultIgnorable(r)) throw new Error(`wordkey: the words hold the invisible character U+${hex4(r)}`);
if (!isAssigned(r)) throw new Error(`wordkey: the words hold U+${hex4(r)}, unassigned in Unicode ${UNICODE_VERSION}`);
}
const problem = hidden(w);
if (problem !== undefined) throw new Error(problem);
}
const n = countedWords(words);
if (n < MIN_WORDS) throw new Error(`wordkey: a key of words needs at least ${MIN_WORDS} different words of ${MIN_LETTERS} or more letters, not ${n}`);

@ -0,0 +1,153 @@
// Tests of wordlist.ts: the cases of generate_test.go of Go wordkey, with its
// texts; the list of datekeys-go, read with the SHA-256 that its README
// records; and draws that are uniform and never repeat a word.
import { readFileSync } from 'node:fs';
import { fileURLToPath } from 'node:url';
import { describe, expect, it } from 'vitest';
import { toHex, utf8Bytes } from './bytes.ts';
import type { RandomWords } from './random.ts';
import { checkWords, normalizeWords } from './wordkey.ts';
import { checkWordList, DEFAULT_WORD_COUNT, generateWords, MIN_LIST_SIZE, readWordList, WORD_LIST_SHA256, wordBits } from './wordlist.ts';
const listFile = (name: string): Uint8Array => new Uint8Array(readFileSync(fileURLToPath(new URL(`../../../wordlists/${name}`, import.meta.url))));
const cp = (r: number): string => String.fromCodePoint(r);
const hash = async (b: Uint8Array): Promise<string> => toHex(new Uint8Array(await crypto.subtle.digest('SHA-256', b as Uint8Array<ArrayBuffer>)));
// pal followed by three letters: palaaa, palaab, …, as in Go.
const base = Array.from({ length: MIN_LIST_SIZE }, (_, i) => `pal${String.fromCharCode(97 + Math.floor(i / 676), 97 + (Math.floor(i / 26) % 26), 97 + (i % 26))}`);
const withWord = (i: number, w: string): string[] => base.map((x, j) => (j === i ? w : x));
const fileOf = (words: readonly string[]): Uint8Array => utf8Bytes(`${words.join('\n')}\n`);
// A source of 32-bit words that repeats `seed`.
function repeating(seed: readonly number[]): RandomWords {
let i = 0;
return () => seed[i++ % seed.length]!;
}
describe('readWordList', () => {
it('reads the list es of datekeys-go, with the SHA-256 that its README records', async () => {
const readme = new TextDecoder().decode(listFile('README.md'));
expect(readme).toContain(`| \`es.txt\` | 7776 | \`${WORD_LIST_SHA256['es']}\` |`);
const words = await readWordList('es', listFile('es.txt'));
expect(words).toHaveLength(7776);
expect([words[0], words.at(-1)]).toEqual(['abad', 'útil']);
});
it('refuses a list of another SHA-256 and a language without a list, with the texts of Go', async () => {
const changed = listFile('es.txt');
changed[0]! ^= 1;
const refused = `wordkey: the list "es" has the SHA-256 ${await hash(changed)}, not ${WORD_LIST_SHA256['es']}`;
await expect(readWordList('es', changed)).rejects.toThrow(refused);
await expect(readWordList('xx', listFile('es.txt'))).rejects.toThrow(/^wordkey: no word list for "xx"; the lists are es$/);
await expect(readWordList('es', fileOf(base), { fr: '00', de: '11' })).rejects.toThrow(/^wordkey: no word list for "es"; the lists are de, fr$/);
});
it('refuses a list with its pin that is not UTF-8, or that checkWordList refuses, as Go wordkey.List', async () => {
const pin = async (lang: string, b: Uint8Array): Promise<Record<string, string>> => ({ [lang]: await hash(b) });
const ok = fileOf(base);
await expect(readWordList('es', ok, await pin('es', ok))).resolves.toEqual(base);
// Without the end of its last line, too.
const unended = ok.slice(0, -1);
await expect(readWordList('es', unended, await pin('es', unended))).resolves.toEqual(base);
const cyrillic = fileOf(withWord(5, `${cp(0x441)}asa`));
const err = (await readWordList('es', cyrillic, await pin('es', cyrillic)).catch((e: unknown) => e)) as Error;
expect(err.message).toBe(`wordkey: the list "es": line 6, "${cp(0x441)}asa", holds U+0441, which is not in the alphabet of "es"`);
expect((err.cause as Error).message).toBe(`line 6, "${cp(0x441)}asa", holds U+0441, which is not in the alphabet of "es"`);
// A BOM stays: it is an invisible character of the first word.
const bom = utf8Bytes(`${cp(0xfeff)}${base.join('\n')}\n`);
await expect(readWordList('es', bom, await pin('es', bom))).rejects.toThrow(/^wordkey: the list "es": line 1: wordkey: the words hold the invisible character U\+FEFF$/);
const latin1 = fileOf(base);
latin1[3] = 0xe1;
await expect(readWordList('es', latin1, await pin('es', latin1))).rejects.toThrow(/^wordkey: the list "es" is not UTF-8$/);
await expect(readWordList('xx', ok, await pin('xx', ok))).rejects.toThrow(/^wordkey: the list "xx": no alphabet for the language "xx"$/);
});
});
describe('checkWordList', () => {
it('accepts a list of 2048 different words of the alphabet, and refuses the cases of Go wordkey.TestCheckList', async () => {
await expect(checkWordList('es', base)).resolves.toBeUndefined();
await expect(checkWordList('xx', base)).rejects.toThrow(/^no alphabet for the language "xx"$/);
const cases: [readonly string[], string][] = [
[base.slice(0, MIN_LIST_SIZE - 1), '2047 words, fewer than 2048'],
[withWord(5, 'dos palabras'), 'line 6, "dos palabras", is not one word'],
[withWord(5, ' '), 'is not one word'],
[withWord(5, 'mi'), '"mi", is not one word of 3 or more letters'],
[withWord(5, `casa${cp(0x200b)}`), 'invisible character U+200B'],
// Only the letters of the alphabet of the language, as the list writes
// them: no capitals, no Cyrillic U+0441 that looks like a Latin c, no
// digits, no carriage return of a file with CRLF lines.
[withWord(5, 'Palaaf'), 'line 6, "Palaaf", holds U+0050, which is not in the alphabet of "es"'],
[withWord(5, `${cp(0x441)}asa`), 'holds U+0441, which is not in the alphabet'],
[withWord(5, 'pal1'), 'holds U+0031'],
[withWord(5, `palaaf${cp(0x0d)}`), `line 6, "palaaf${cp(0x5c)}r", holds U+000D`],
[withWord(5, base[4]!), 'line 6, "palaae", is the same word as "palaae"'],
[withWord(5, 'palaáe'), 'line 6, "palaáe", is the same word as "palaae"'],
[[...withWord(0, 'papá'), 'papa'], '"papa", is the same word as "papá"'],
];
for (const [list, want] of cases) await expect(checkWordList('es', list)).rejects.toThrow(want);
});
});
describe('generateWords', () => {
it('draws different words of the list that make a key, by default 7 with crypto.getRandomValues', async () => {
const list = await readWordList('es', listFile('es.txt'));
const words = generateWords(list);
expect(words).toHaveLength(DEFAULT_WORD_COUNT);
expect(new Set(words).size).toBe(DEFAULT_WORD_COUNT);
for (const w of words) expect(list).toContain(w);
await expect(checkWords(await normalizeWords(words.join(' ')))).resolves.toBeUndefined();
});
it('reads nothing but its random words, and draws a word only once', () => {
const seed = [7, 1, 200, 33, 0x9e3779b9, 12345];
expect(generateWords(base, 6, repeating(seed))).toEqual(generateWords(base, 6, repeating(seed)));
// Index 0 comes twice: the second is drawn again.
expect(generateWords(base, 6, repeating([0, 0, 1, 2, 3, 4, 5]))).toEqual(base.slice(0, 6));
expect(() =>
generateWords(base, 6, () => {
throw new Error('no more randomness');
}),
).toThrow('no more randomness');
});
it('refuses fewer than 6 words and more than half the list, with the texts of Go', () => {
expect(() => generateWords(base, 5)).toThrow(/^wordkey: a key of words needs at least 6 words, not 5$/);
expect(() => generateWords(base, 6.5)).toThrow(/^wordkey: a key of words needs at least 6 words, not 6\.5$/);
expect(() => generateWords(base, 1025)).toThrow(/^wordkey: 1025 words of a list of 2048$/);
expect(generateWords(base, 1024)).toHaveLength(1024);
});
// As TestGenerateUniform of Go: over 7776·40 draws of one word, each index
// falls in its bucket of 64 between 0.8 and 1.2 times the mean.
it('draws every word about as often', async () => {
const list = await readWordList('es', listFile('es.txt'));
const index = new Map(list.map((w, i) => [w, i]));
const buckets = 64;
const count = new Array<number>(buckets).fill(0);
let draws = 0;
while (draws < list.length * 40) {
for (const w of generateWords(list, 6)) {
count[Math.floor((index.get(w)! * buckets) / list.length)]!++;
draws++;
}
}
const mean = draws / buckets;
for (const c of count) {
expect(c).toBeGreaterThan(0.8 * mean);
expect(c).toBeLessThan(1.2 * mean);
}
});
});
describe('wordBits', () => {
it('is log2 of the draws in order, as Go wordkey.Bits', () => {
expect(wordBits(7776, 1)).toBe(Math.log2(7776));
expect(wordBits(7776, 0)).toBe(0);
// 7 words of 7776 are a little under 90.5 bits, and 6 of 2048, the
// fewest, a little under 66.
expect(wordBits(7776, 7)).toBeCloseTo(90.4698, 4);
expect(wordBits(7776, 8)).toBeCloseTo(103.3933, 4);
expect(wordBits(2048, 6)).toBeCloseTo(65.9894, 4);
});
});

@ -0,0 +1,126 @@
// The random words of a key of words (spec §38.1, the SHOULD to offer them by
// default): Go wordkey.Generate, CheckList and Bits, with their texts. A list
// is public, and its strength is its number of words, never its secrecy: a
// list whose words are one once normalized cuts that strength without a sign,
// and a word with a letter of another script that looks like one of the
// language, a Cyrillic U+0430 for a Latin a, is typed again with the letter
// of the keyboard and leaves the capsule shut. So no list is trusted, not
// even one of DateKeys: readWordList takes a list only with its pinned
// SHA-256, and checkWordList checks it against the alphabet of its language,
// which only this code gives. The lists are those of datekeys-go,
// wordkey/lists, which scripts/sync-testdata.mjs copies into wordlists/ with
// the test data. The words are drawn on the device, with
// crypto.getRandomValues by default, and never travel.
import { decodeUtf8, goQuote, sha256, toHex } from './bytes.ts';
import { cryptoWords, randomIndex, type RandomWords } from './random.ts';
import { MIN_LETTERS, MIN_WORDS, wordRules } from './wordkey.ts';
/** The number of words drawn when the caller does not ask for more: 7 of a list of 7776 are about 90 bits. */
export const DEFAULT_WORD_COUNT = 7;
/** The fewest words of a list (spec §38.1: at least 6 words of a list of 2048 or more). */
export const MIN_LIST_SIZE = 2048;
/**
* The SHA-256 of each list of datekeys-go, by language, as
* wordkey/lists/README.md records it: a list changes only with its hash
* here, so that the change is never silent.
*/
export const WORD_LIST_SHA256: Readonly<Record<string, string>> = {
es: 'ff77b487765c000da97cca58fe94a2cdb947303e7a07460614d7d95d800034fe',
};
// The letters that a word of a list of each language may hold, as the list
// writes it, lower case and in NFC: the alphabets of Go wordkey.
const ALPHABETS: Readonly<Record<string, string>> = {
es: 'abcdefghijklmnopqrstuvwxyzáéíóúüñ',
};
const hex4 = (r: number): string => r.toString(16).toUpperCase().padStart(4, '0');
/**
* The words of the list of the language `lang` in `file`, its bytes as
* datekeys-go has them: UTF-8, one word per line, the last one ended. The
* list must have the SHA-256 that `pins` gives for `lang`, WORD_LIST_SHA256
* by default, be UTF-8 and pass checkWordList: otherwise it throws, with the
* texts of Go wordkey.List for a language without a list and for a list
* refused.
*/
export async function readWordList(lang: string, file: Uint8Array, pins: Readonly<Record<string, string>> = WORD_LIST_SHA256): Promise<string[]> {
if (!Object.hasOwn(pins, lang)) throw new Error(`wordkey: no word list for ${goQuote(lang)}; the lists are ${Object.keys(pins).sort().join(', ')}`);
const got = toHex(await sha256(file));
if (got !== pins[lang]) throw new Error(`wordkey: the list ${goQuote(lang)} has the SHA-256 ${got}, not ${pins[lang]}`);
// A BOM stays, and checkWordList refuses it as an invisible character.
const text = decodeUtf8(file);
if (text === undefined) throw new Error(`wordkey: the list ${goQuote(lang)} is not UTF-8`);
const words = (text.endsWith('\n') ? text.slice(0, -1) : text).split('\n');
try {
await checkWordList(lang, words);
} catch (err) {
throw new Error(`wordkey: the list ${goQuote(lang)}: ${(err as Error).message}`, { cause: err });
}
return words;
}
/**
* Why `words` cannot be a list of the language `lang` for generateWords,
* with the errors of Go wordkey.CheckList and in its order: a language
* without an alphabet here; fewer than MIN_LIST_SIZE words; a word that is
* not one word of MIN_LETTERS characters or more once normalized, that holds
* a character that checkWords refuses or one that is not a letter of the
* alphabet of `lang`; or a word that is the same as another once normalized.
* Two words such as «papa» and «papá» would be one word with less entropy
* than the list promises.
*/
export async function checkWordList(lang: string, words: readonly string[]): Promise<void> {
if (!Object.hasOwn(ALPHABETS, lang)) throw new Error(`no alphabet for the language ${goQuote(lang)}`);
const alphabet = new Set(ALPHABETS[lang]);
if (words.length < MIN_LIST_SIZE) throw new Error(`${words.length} words, fewer than ${MIN_LIST_SIZE}`);
const { normalize, hidden } = await wordRules();
const seen = new Map<string, string>();
for (const [i, w] of words.entries()) {
const n = normalize(w);
if (n.length !== 1 || [...n[0]!].length < MIN_LETTERS) throw new Error(`line ${i + 1}, ${goQuote(w)}, is not one word of ${MIN_LETTERS} or more letters`);
const problem = hidden(n[0]!);
if (problem !== undefined) throw new Error(`line ${i + 1}: ${problem}`);
for (const ch of w) {
if (!alphabet.has(ch)) throw new Error(`line ${i + 1}, ${goQuote(w)}, holds U+${hex4(ch.codePointAt(0)!)}, which is not in the alphabet of ${goQuote(lang)}`);
}
const prev = seen.get(n[0]!);
if (prev !== undefined) throw new Error(`line ${i + 1}, ${goQuote(w)}, is the same word as ${goQuote(prev)} once normalized`);
seen.set(n[0]!, w);
}
}
/**
* `n` different words of `list`, drawn uniformly with `random`,
* crypto.getRandomValues by default, as Go wordkey.Generate: n must be
* MIN_WORDS or more, and at most half the list. Each word adds
* log2(list.length) bits, a little less for each word already drawn
* (wordBits).
*/
export function generateWords(list: readonly string[], n: number = DEFAULT_WORD_COUNT, random: RandomWords = cryptoWords()): string[] {
if (!Number.isSafeInteger(n) || n < MIN_WORDS) throw new Error(`wordkey: a key of words needs at least ${MIN_WORDS} words, not ${n}`);
if (n > Math.floor(list.length / 2)) throw new Error(`wordkey: ${n} words of a list of ${list.length}`);
const picked = new Set<number>();
const words: string[] = [];
while (words.length < n) {
const i = randomIndex(list.length, random);
if (picked.has(i)) continue;
picked.add(i);
words.push(list[i]!);
}
return words;
}
/**
* The strength of `count` words that generateWords draws from a list of
* `size` words, as Go wordkey.Bits: log2 of the number of draws in order,
* size·(size−1)·…, which whoever knows the list must search, before the
* rounds of PBKDF2.
*/
export function wordBits(size: number, count: number): number {
let bits = 0;
for (let i = 0; i < count; i++) bits += Math.log2(size - i);
return bits;
}

@ -68,6 +68,7 @@ export default defineConfig({
'src/lib/inspector/create-check.ts': { 100: true },
'src/lib/inspector/drand.ts': { 100: true },
'src/lib/dkc/wordkey.ts': { 100: true },
'src/lib/dkc/wordlist.ts': { 100: true },
// The locator of datekeys.capsule: its data, its addresses and their
// IP addresses, its plaintext, and the age files of its sealing and of
// the envelope.

Loading…
Cancel
Save

Powered by TurnKey Linux.