A Rust implementation of the Markdoc language: parse, validate, transform, render, and format.
Markdoc is CommonMark plus a tag syntax that turns documents into structured, validatable content instead of pre-rendered HTML:
{% callout type="note" %}
Tags nest, take typed attributes, and are checked against a schema.
{% /callout %}Documentation and a live playground -- the engine compiled to WebAssembly, running in your browser.
cargo add accent-proustuse accent_proust::{builtins, parse, render, transform};
let document = parse::parse("# Title\n\nSome *text*.\n");
let tree = transform::transform(&document, &builtins::config());
assert_eq!(
render::render_all(&tree.into_vec()),
"<article><h1>Title</h1><p>Some <em>text</em>.</p></article>"
);A tag needs a schema before it validates or renders. render names the element
to emit; declared attributes reach the output, undeclared ones are an error.
use std::sync::Arc;
use accent_proust::validate::{self, MapSchemaSource, Schema, SchemaAttribute, ValidationType};
let mut schemas = MapSchemaSource::builtin();
schemas.insert_tag(
"callout",
Schema::new().render("div").attribute(
"type",
SchemaAttribute {
attribute_type: Some(ValidationType::String),
required: true,
..SchemaAttribute::default()
},
),
);
let config = builtins::config_with(Arc::new(schemas));
let document = parse::parse("{% callout type=\"note\" %}\nBody\n{% /callout %}\n");
assert!(validate::validate_tree(&document, &config).is_empty());
let tree = transform::transform(&document, &config);
assert_eq!(
render::render_all(&tree.into_vec()),
"<article><div type=\"note\"><p>Body</p></div></article>"
);Validation errors are data, not failures: you get a Vec, so an editor shows
every problem at once instead of the first one.
let document = parse::parse("{% callout %}\nBody\n{% /callout %}\n");
for error in validate::validate_tree(&document, &config) {
println!("{}: {}", error.error.id, error.error.message);
// attribute-missing-required: Missing required attribute: 'type'
}Error ids match upstream Markdoc exactly, so tooling written against its codes works unchanged.
format prints a tree as canonical Markdoc source. It normalises spacing inside
a tag and leaves your own spellings alone, so __bold__ stays __bold__.
use accent_proust::format;
let document = parse::parse("{% callout type=\"note\" %}\nBody\n{% /callout %}\n");
assert_eq!(
format::format(&document),
"{% callout type=\"note\" %}\nBody\n{% /callout %}\n"
);format(parse(s)) is idempotent, so a tool can rewrite a file in place, and
parse(format(ast)) gives back the same tree, so formatting loses nothing.
The same engine as a command, for a documentation repository that wants a CI gate and for anyone with a Markdoc file to tidy:
cargo install --path crates/accent-proust-cli # the binary is `accent-proust`
accent-proust fmt --check docs/*.md
accent-proust validate --config schema.yaml --partials docs/partials docs/*.md
accent-proust render --config schema.yaml --var channel=stable docs/page.mdfmt reprints canonical source and, with --check, prints a diff and exits 1
if anything would change. validate reports path:line:column: level[id]: message per error, or one JSON object per file with --format json, and
exits 1 on an error. render, transform and parse print HTML, the
renderable tree and the syntax tree. The configuration is a YAML or JSON file
in the same vocabulary the WebAssembly bindings read from an object, so a
schema declared for one host is accepted by the other; --partials is a
directory of files, which is the thing the browser cannot do. The crate's
README has the rest, exit codes
included.
The bundled tokenizer uses pulldown-cmark, behind the default
pulldown-cmark-tokenizer feature. Turn it off and implement Tokenizer if you
already parse CommonMark, or if you pin pulldown-cmark to a git revision --
Cargo treats that as a different package, so you would compile two CommonMark
parsers into one binary and render some documents through each.
accent-proust = { version = "*", default-features = false }A CI job builds and tests the crate in exactly that shape, so it is supported
rather than tolerated. Tokenizer is one of three seams; the crate does no I/O,
reads no configuration, and decides no HTML policy. SchemaSource answers where
a schema comes from, and TagRenderer owns escaping and HTML policy. All three
are yours.
Ported from upstream Markdoc v0.5.9 (revision afee1a4). The tag language and
the error ids are the contract. CommonMark edge behaviour is not: upstream builds
on markdown-it, this crate on pulldown-cmark. Every deliberate difference is
declared in DIVERGENCES.md, never emulated silently.
Upstream's 105-case corpus is vendored and run as the test suite. Nothing fails; "annotated" is a case exercising a declared divergence, counted apart so that giving something up stays visible.
cargo test --test conformance -- --nocapture
# conformance: 95 green, 10 annotated, 0 failing (of 105)The library's minimum supported Rust version is 1.96. Develop on stable, which the test suite needs. See AGENT.md for the gates and the workflow.
MIT. See LICENSE.
A compatible reimplementation derived from the MIT-licensed Markdoc source.
Marcel Proust composed A la recherche du temps perdu on strips of paper glued into the manuscript to extend it. Those paperoles are exactly what a formatter does: parse, mutate, and print canonical source.
accent-proust belongs to the family of Accent CMS
crates.