Custom Dialects & Libraries
You can teach rosetta-date a new formatting language with defineDialect, or model a new tool on
top of an existing dialect with defineLibrary.
The key idea: map to meaning
Map each token to a canonical symbol — what it means (year, 2-digit month, …), not another
dialect’s spelling. Canonical is the neutral vocabulary every built-in dialect routes through, so
a token you map to Canonical.YearNumeric converts to and from every other dialect for free.
defineDialect
import { Canonical, convert, defineDialect } from 'rosetta-date'
import { ldml } from 'rosetta-date/dialects'
const custom = defineDialect({
name: 'custom',
syntax: { kind: 'delimited', open: '{', close: '}' },
tokens: [
{ token: 'yr', canonical: Canonical.YearNumeric },
{ token: 'mo', canonical: Canonical.MonthTwoDigit },
{ token: 'dy', canonical: Canonical.DayOfMonthTwoDigit },
],
})
convert('yr-mo-dy', { from: custom, to: ldml }) // 'yyyy-MM-dd'The syntax chooses a tokenization family. The example above is delimited (letter tokens, with
{...} literals). For a %-directive grammar like strftime, use the directive family — token
spellings include the marker, and a literal marker is the marker doubled:
const percent = defineDialect({
name: 'percent',
syntax: { kind: 'directive', marker: '%' },
tokens: [
{ token: '%Y', canonical: Canonical.YearNumeric },
{ token: '%m', canonical: Canonical.MonthTwoDigit },
{ token: '%d', canonical: Canonical.DayOfMonthTwoDigit },
],
})
convert('%Y-%m-%d', { from: percent, to: ldml }) // 'yyyy-MM-dd'A dialect can also declare composites — parse-time macros where one spelling expands to a sub-pattern. They render as their expansion (never the composite), so a composite normalizes on a round trip:
const withComposites = defineDialect({
name: 'percent-composites',
syntax: { kind: 'directive', marker: '%' },
tokens: [/* … %H %M %S … */],
composites: [{ token: '%T', expandsTo: '%H:%M:%S' }],
})
convert('%T', { from: withComposites, to: ldml }) // 'HH:mm:ss'defineLibrary
A library extends a dialect — adding a tool’s own tokens via extends (and, in the built-ins,
declaring a supports set). Here we add an LDML-flavoured tool that understands the Unix epoch:
import { Canonical, convert, defineLibrary } from 'rosetta-date'
import { ldml } from 'rosetta-date/dialects'
import { dateFns } from 'rosetta-date/libraries'
const ldmlPlus = defineLibrary({
name: 'ldml-plus',
dialect: ldml,
extends: [{ token: 'u', canonical: Canonical.EpochSeconds }],
})
convert('u', { from: ldmlPlus, to: dateFns }) // 't' — bridged by the canonical symbolThe token u is not in the bare ldml dialect, but because it maps to Canonical.EpochSeconds, it
bridges straight to date-fns’s t.
Narrowing what a tool renders with supports
Where extends adds tokens, supports narrows a tool to a subset of the grammar. It is keyed
by canonical field, not spelling: you list the fields the tool renders, and each is emitted with
the dialect’s primary spelling. Omit it (as ldmlPlus does) to render the whole grammar.
import { Canonical, convert, defineLibrary } from 'rosetta-date'
import { moment } from 'rosetta-date/dialects'
const dayjsLike = defineLibrary({
name: 'dayjs-like',
dialect: moment,
supports: new Set([Canonical.YearNumeric, Canonical.MonthTwoDigit, Canonical.DayOfMonthTwoDigit]),
})
convert('MM', { from: moment, to: dayjsLike }) // 'MM' — a supported field
convert('Mo', { from: moment, to: dayjsLike }) // '[Mo]' — month ordinal isn't in `supports`, so it's flaggedValidation happens once
defineDialect and defineLibrary validate a definition once, at definition time — they
throw on an incoherent token table or syntax (including a directive token that omits its marker) — and return a stable object you reuse across
conversions. That object identity matters: per-dialect compilation is cached by object, so
constructing a fresh definition on every convert call silently misses the cache and recompiles
each time. Define once, reuse everywhere.
Canonical symbols
Canonical values are stable identifiers shaped as field/style, versioned by semver:
- Adding a symbol is a minor (additive) change.
- Renaming or removing one is breaking.
Prefer the named members (Canonical.YearNumeric) over the raw strings. The full vocabulary is
exported from the package root, and the TokenRule and TokenSyntax types are available for
annotating your token tables and syntax.
See the API Reference for the exact signatures and types.