For each real or desired use case: is it enabled by the format as it stands (0.4), and if not, what is missing? The point is to separate three things that get conflated —
- work the format already enables (author it today),
- work that is a consumer concern the format deliberately does not own (the data-agnostic north
star — see
CLAUDE.md), so there is nothing to add, and - genuine gaps that need an additive format change.
Every gap below is tagged → RMn and collected in Roadmap items surfaced at the end; those are the
items to migrate into ROADMAP.md.
The design docs are stages of one loop; an idea moves left-to-right as it matures:
- Feedback — a consumer's field report. →
CONSUMER_SUGGESTIONS.md(open) /CONSUMER_SUGGESTIONS_HISTORY.md(answered), the runbook inCONSUMER_TRIAGE_LOOP.md. The pre-Snrounds 1 and 2 were their own threads and are both retired — their dispositions are recorded inhistory/CONSUMER_SUGGESTIONS_HISTORY_PRE_0_6.md - Usage → blockers → solvability — run each use case against the current bricks: is it enabled, consumer-side, or a gap; and is the gap closable additively? → this doc
- Means → draft schema → decision — for a gap worth closing, the proposed shape + a charter check
- the open questions to settle it. →
PROPOSAL_0_5.md(0.5),PROPOSAL_0_4_1.md(the 0.4.1 patch). (0.4's shipped decisions →CHANGELOG.md.)
- the open questions to settle it. →
- Conclusion — "how to do it now, with these bricks" — the distilled worked example once the shape
is settled. →
REFERENCE_EXAMPLES.md - Terminal, one of two:
- Fixed — schema + compiler shipped (the models;
COMPILER.mdmarks it validated/materialized), or - Deferred — a recognised gap parked as a roadmap item (
RMn→ROADMAP.md) when the means aren't worth building yet.
- Fixed — schema + compiler shipped (the models;
So this doc and REFERENCE_EXAMPLES.md are the same use cases at two points in the loop — here
they are questions (what blocks?), there they are answers (author it like this). An ENABLED /
schema-ready row here graduates to a REFERENCE_EXAMPLES entry; a GAP row graduates to the
current proposal (PROPOSAL_0_5 / PROPOSAL_0_4_1, if being closed now) or an RMn roadmap item
(if deferred). The loop is why a
"blocker" is never a dead end: it is either dissolved (it was consumer-side all along), closed
additively, or explicitly parked.
Verdict legend
- ENABLED — authorable/usable on 0.4 now, no format change.
- CONSUMER-SIDE — a consumer (runner/app) feature; the format correctly owns nothing here (per the north star). No format change — often no blocker at all.
- SCHEMA-READY / COMPILER-PENDING — the schema models express it; only the deferred compiler materialization (new parquet + round-trip) is missing. One known gap, not per-use-case.
- GAP — needs an additive format addition (the interesting rows).
A recurring result: most "can the format do X?" questions dissolve into CONSUMER-SIDE, because the
module is declarative annotation and the doing is a consumer's. The format earns its keep by the
properties it froze (declarative-not-code, integrity-as-identity, the unresolved/callability
contract), which are what make a consumer's X safe and reproducible — not by hosting X.
Verdict: CONSUMER-SIDE — no blocker, no format deliverable. This is the headline dogfooding use case and it is already enabled. Walk the requirements:
- A panel is a set of loci + expected genotypes/bins → already a module (a
variants.csvof genotype→conclusion rows, or a binning table of measure→phenotype). No new "panel type". - Deterministic extraction of the observed value from a VCF → the consumer's job, and
source_field(0.4) now names the exact VCFFORMAT/INFOfield to read, removing the last glue. - Trustworthy before/after (per-caller, ±liftover) diffs →
artifact.digestalready gives the module a content identity, so a report-card diff is a byte-level diff of two deterministic runs. - No-call ≠ mismatch → the mandatory
unresolvedoutcome + the callability contract already stop a "no-call under caller B" from masquerading as a "mismatch vs caller A".
The one thing that looked like a format artifact — a standardized evaluation-output / report-card
schema ({locus, observed, callability, bin_selected, verdict}) — is per-sample results, i.e. a
measurement, so the north star keeps it consumer-side. It belongs in just-dna-lite, not
just-dna-format. Nothing to schedule in the format. → RM7 records the consumer-side schema so
it is not mistaken for a format item.
Verdict: ENABLED at the format boundary; emission is CONSUMER-SIDE. A caller emits its niche
genotype (PER3 span, DAT1 motif-path, MAOA half-repeat) as a synthetic VCF record (<STR> with
INFO/RU, FORMAT/REPCN, custom evidence fields); a repeat_alleles.csv module consumes it via the
same source_field=REPCN path as ExpansionHunter. The format does not invent a representation — it
binds to the VCF one. Producing the augmented VCF is the caller's job (consumer-side). The only
format touch-point (source_field) shipped in 0.4. Consuming symbolic alleles (<STR n>,
<CNV>) at the count/dosage layer is enabled via the binning tables; representing a symbolic allele
inside a VariantRow.genotype is not (see §3b) → RM5.
- Callability three-state (covered-hom-ref vs no-call): CONSUMER-SIDE, derivable from VCF
DP/GQ/FT(or a gVCF ref-block). The format's part:requires_callable(reserved flag) marks rows where absence is informative; promoting it to a typed boolean column and reservingcallable_from(the DP,GQ,FT signal) are the format-side follow-ups→ RM6. - Phasing-aware panels: ENABLED — the
phasedflag + the phased genotype formA|G(0.3 item 5b) already let a runner do cis/trans for compound-het and star-allele phasing (*2x2/*4vs*2/*4x2). No gap. - Trio / multi-sample (Mendelian / de-novo assertions): CONSUMER-SIDE — VCF is natively
multi-sample; the assertion runner is a consumer. An optional declarative inheritance-expectation
field on a panel row would let the module carry the assertion as data rather than consumer lore —
small, additive, optional
→ RM10(only if a real module needs it).
- A canonical machine/LLM-facing authoring reference. Consumers (MCP servers, agents, docs)
hard-code prose summaries of the DSL that drift from the real schema. ADOPTED (RM8, shipped in
the 0.4 sample):
just_dna_format.reference.authoring_reference()returns a JSON-serialisable summary — every model's field list + all vocabularies + reserved names + the palette — generated from the live models, so it cannot drift;json_schemas()gives the full JSON Schema. A consumer'sget_spec_formatrenders this instead of a hand-maintained blob. - A recommended icon/color palette.
Displayvalidatesicon_set/colorbut shipped no recommended enumerated palette, so each authoring tool invented one. ADOPTED (RM9, shipped in the 0.4 sample):manifest.RECOMMENDED_COLORS/RECOMMENDED_ICONS(curatedsemantic-use → valuemaps, recommendation-only — not enforced), surfaced throughauthoring_reference().
Verdict: SCHEMA-READY (interface) / GAP (native materialization). GenePanelSpec (source,
reference, reference_sha256, genes, significance) already declares the panel and is recorded
verbatim in the manifest. Today an app-side adapter (just-dna-lite) enumerates the matching ClinVar
pathogenic/likely-pathogenic variants into variants.csv; the compiler does not resolve the panel
itself. Missing: native compile-time materialization (gene set + significance predicate →
weights.parquet) gated on a working, content-pinned ClinVar reference mixin → RM4. Until then the
use case is fully reachable through the app-side adapter — so it is enabled in practice, with the
native path a convenience/quality follow-up.
Verdict: ADOPTED (RM3, shipped in the 0.4 sample). A PharmGKB row maps a variant/diplotype → a
drug + a response/phenotype + a PharmGKB evidence level (1A…4, VALID_EVIDENCE_LEVELS)
— a different axis from a risk weight. Built as a dedicated PharmVariantRow (pharm_variants.csv)
for single-variant drug response (keeps the SNP core clean — one CSV = one concern), plus optional
drug/response/evidence_level columns on DiplotypeRow for the diplotype-keyed case. A PharmGKB
module has no empty variants.csv. evidence_level is a third significance-flavoured axis,
distinct from stat_significance/clin_sig (orthogonal-axes discipline, Principle 5). Materialization
deferred with the rest of 0.4.
Verdict: ENABLED (RM90 shipped), and the interesting part is what it refused. A curator with a
trait panel wants the published effect sizes beside their own annotations — partly as evidence, partly
because a consumer reported that hand-set weight values "construct nonsense" across a corpus and that
GWAS effects are often better grounded (S36).
The enabled shape: just-dna-enricher gwas fills gwas_effects.csv, one row per published
association, and the module ships it beside variants.csv. A consumer joins on variant_key and reads
effect_size with effect_unit and effect_allele.
The blocker that was not dissolved, stated so nobody re-proposes it. Filling weight from those
effects is barred (MODULE_LIFECYCLE § Stage 3), and a per-row precedence rule was refused as putting two
methodologies in one summable column. What closed the gap additively was two tables and a declaration,
not an exception to the rule.
And a caution the real data supplies better than the design note could. rs1800562's 186
associations span 62 EFO traits in 12 distinct effect units — three of them spellings of one unit, two
more differing only in case — with 138 rows in the Catalog's uninformative unit and 42 of 195 naming
no effect allele at all. A "GWAS effects are better than curator weights" pipeline that pools those is
worse than the weights it replaces. Read them per trait, and read manifest.gwas_effects.units before
pooling anything.
Verdict: ENABLED (RM1 + RM2 shipped). A module is a directory of CSVs carrying both
variants.csv (VariantRow) and pgs.csv (PgsRow), joined on the shared trait_efo_id (item 5) so
a variant panel and its PRS companion sit in one content-addressed unit. The compiler now materializes
every present table kind to parquet (round-trip lossless) and treats variants.csv as optional, so
composed and single-domain modules both compile. No blocker.
0.5 revisit: the 0.4 shape was modelled at the wrong grain, and real data proved it. RM3 was declared shipped against a hand-authored sample. Run against the actual ClinPGx corpus it does not hold, in two stages:
- A clinical annotation is published per genotype — the summary table names the variant and drug,
a child table gives one row per call, and 4,618 of 5,113 carry exactly three.
PharmVariantRowhad nogenotype, and the compiler deduped on(variant_key, drug), so authoring the real SLCO1B1/simvastatin annotation producedduplicate row for key ('rs4149056', 'simvastatin'). ~97% of the corpus was unauthorable. - Adding
genotypewas not enough. One variant and one drug carry several distinct annotations — rs4149056 + simvastatin is Metabolism/PK at 1A, Efficacy at 3 and Toxicity at 1A. 1,199 of 17,380 triples collide; 839 separate by phenotype category, 283 by neither category nor level.
Closed additively with genotype, phenotype_category (closed vocabulary) and annotation_id (a
source accession as identity, like PgsRow.pgs_id) → RM20. The lesson is the dogfood rule in
CLAUDE.md read from the other side: a shape validated against a sample rather than a corpus is not
validated. See reference_examples/pgx_slco1b1_simvastatin/.
Verdict: ENABLED, with the licensing made legible (RM21). The enricher can now cross-check a
module's allele_function.csv against PharmVar and CPIC and its pharm_variants.csv against a
ClinPGx snapshot, and resolution reaches pharm_variants.csv/haplotypes.csv so a PGx module gets
coordinates without carrying a variants.csv.
The blocker that turned up was not technical. Every pharmacogenomics upstream is copyleft and none
is sellable: ClinPGx, CPIC and PharmVar are each CC BY-SA 4.0 plus a separate contractual bar on
sale. api.pharmgkb.org was retired 2026-07-20, and CPIC sits inside the ClinPGx merger with its
licence page redirecting to the ClinPGx policy — so swapping sources does not escape the terms. That
is consumer-relevant rather than a format gap, but it is unrepresentable in a 0.4 module, so it was
closed additively as sources.csv (RM21) rather than left to a README nobody can query.
Composition principle (settled during the PharmGKB decision, now in CLAUDE.md): a module composes
from optional table kinds — one CSV = one concern — so the SNP core (variants.csv+studies.csv)
stays minimal and no module ever carries an empty variants.csv or a foreign domain's columns just to
host one table. This is the human-authorable half of the RM2 work.
Verdict: ENABLED for small ACGT indels; GAP for structural/symbolic. VariantRow alleles are
^[ACGT]+$ multi-base, so a small insertion/deletion is expressible today (ref=A, alts=AT,
genotype A/AT) on the same variants.csv as SNPs — a mixed SNP+indel panel is authorable now. What
is not expressible: symbolic/large structural alleles (<DEL>, <INS>, <DUP>, repeat
expansions as <STR>) — there is no symbolic-allele genotype. Those route through the copy-number /
repeat binning tables (dosage/count, not sequence) or await a symbolic-allele representation
→ RM5. So: everyday SNP+indel modules need nothing; SV-scale variation is the recognised gap.
Verdict: ENABLED (RM1 + RM2 shipped). The generalization of 3a: a personal/curated module mixing
variants.csv, pgs.csv, activity_phenotype.csv/diplotypes.csv, and copynumbers.csv, all joined
on trait_efo_id, compiles today — each present kind materializes to parquet with round-trip. This is
exactly the "personal module re-checked deterministically on every pipeline change" the verification
harness (§1a) wants — and RM1/RM2 are what unlocked it.
4a. just-module-validator — deterministic source-checks + provenance enrichment against public sources
Verdict: the validator is CONSUMER-SIDE (Principle 2 keeps it out of these libs); the two additive format anchors it needs shipped in 0.4 (RM11/RM12), and one requiredness fix waits for 1.0.
A proposed sibling library that is network-first: given a module it checks the authored claims against public sources and enriches them —
- validate every
pmidresolves in PubMed, and everyrsidresolves in dbSNP at the authoredchrom:start(flag coord/liftover drift); - cross-fill provenance ids — derive a
doifrom a PMID and vice-versa; - confirm a study's claim actually appears in the cited article's fulltext (imagine further source-checks in the same spirit).
By the data-agnostic north star and Principle 2 (no network; inject-only), the doing — every
fetch and lookup — is a consumer's, and can never live in just-dna-format/just-dna-compiler
(that would pull the network dependency the tiers forbid, Goal 2). So the validator is a new
consumer/enricher sibling to just-dna-lite, recorded as → RM13 so it is not mistaken for format
scope — exactly as the report-card harness is (§1a / RM7). Crucially, most of what it checks needs
nothing from the format: rsid, chrom, start already exist, so validating them against dbSNP
is pure consumer work — enabled today. Two things it wants to anchor are genuine additive format
gaps, and one is a 1.0 requiredness fix:
-
doias a provenance id — additive, shipped in 0.4 (RM11).StudyRowpreviously carried onlypmid(required, and it must contain ≥1 real PubMed id). DOI is wider: it covers preprints (bioRxiv/medRxiv), books, theses, and datasets that have no PMID. The optionaldoicolumn now lets the validator record and cross-fill it, and lets a module cite a DOI-bearing source. Purely additive → P3/P8 clean (new optional field; existing data still validates); validated against the DOI grammar and kept verbatim.→ RM11. -
A provenance locator — search-phrase/regex pointing at the passage in fulltext — additive, shipped in 0.4 (RM12). So the validator can answer "does the cited article's fulltext actually contain this claim?" in a yes/no manner, a study row now carries optional
provenance_quote(keyword phrase) andprovenance_regex. The regex sits squarely inside Principle 1's sanctioned escape hatch: a declarative pattern grammar is data, not code — the module ships the pattern, the consumer supplies the fulltext and runs the match, evaluated by a linear-time / ReDoS-safe engine (P1's explicit requirement; the compiler onlyre.compile-checks it at author time). It is the provenance analogue ofsource_field(0.4):source_fieldis a declarative pointer to where the measurement lives in a VCF; the locator is a declarative pointer to where the claim lives in the article. Neither holds the data it points at (north star ✓). Primarily an aid for LLM-authors (which can emit a precise pattern), yet a plain keyword phrase is legible enough to clear the human-authorability gate for a human author too.→ RM12. -
pmidis mandatory today — the DOI-only case cannot be closed additively (a 1.0 fix).pmid: stris required and must parse to a real PubMed id (extract_pmids), so a preprint/book/thesis with only a DOI is unauthorable right now — and demoting a required field to optional is precisely the move Principle 8 forbids within a major. Addingdoi(RM11) is necessary but not sufficient: whilepmidstays required-and-PMID-shaped, DOI-only provenance is still rejected. The full fix is doi-first at 1.0 — makepmidoptional/legacy and require at least one of{doi, pmid}("not every citation has a PMID, but every citation has a stable id" — the reverse of today's rule). That is a requiredness change → major-only, parked as a 1.0-cleanup candidate, not anRMn. Until 1.0, DOI-only provenance is an explicitly-parked gap.
5a. Structured per-version authorship — who created / edited / audited, and whether each is AI or a human expert
Verdict: SHIPPED in 0.4 (RM14) — an additive, digest-neutral authorship record; the old flat
fields stay for compat. This is the module-level companion to the
network-first validator (§4a): the validator — and a marketplace review queue, and a human auditor —
needs to route its scrutiny by who authored the version, because AI and human error-spectra
overlap but differ. An AI author fabricates plausible-but-wrong PMIDs / rsids / effect-sizes (exactly
the checks RM11–RM13 automate); a human expert makes transcription / off-by-one / stale-reference
slips. The format never performs the scrutiny (consumer-side, north star) — it must carry the
author-kind so the consumer can select the right profile, the same "annotate so the consumer's X is
safe and reproducible" contract as everywhere else.
What exists today, and why it does not cover it:
ModuleManifest.authors: list[str]— a flat list: no role (created/edited/audited), no kind (AI/human). The overloaded-axis anti-pattern (P5), at the list level.curator/method— single free-form strings;Defaults.curatoreven defaults to"ai-module-creator", smuggling author-kind into a string a consumer cannot reliably facet on. This is precisely the axis-overload Principle 5 exists to unwind.Provenance(generator/model/agent_version) + per-variantProvenanceItem.human_reviewed— captures AI-generation and per-variant human review, but not module-level role attribution (who edited vs. audited this version), and it names only the AI side.
So the axes were half-present and tangled. The shipped shape is a structured, per-version
authorship list (Contribution model) unbundling three orthogonal axes (P5): identity (who),
role (created | edited | audited | reviewed, a closed vocab), and kind — a multi-valued,
open tag set with a recommended seed: a human ladder of assurance human → human_expert →
human_certified (medically / board-certified, e.g. a clinical geneticist), or ai plus a scale tag
agent/team/swarm. There is deliberately no hybrid tag — it was rejected as non-explicit
(hybrid what?); a joint contribution is two entries (a human and an ai), each with its own
kind, so the mix is always spelled out. Each entry is optionally timestamped (at). "Per-version"
falls out of immutability (P4): a version's manifest records its own authorship, and cross-version
history is the union via aggregate_provenance.
Why it is cheap. artifact.digest is a Merkle root over the parquet files only — manifest
metadata (logs, provenance, logo, and this) is deliberately out of it. So two versions with
identical annotation content but different authorship keep the same content identity (correct: who
authored ≠ what the annotation is). Adding it is additive/optional (P3/P8), touches no parquet column,
and is digest-neutral even after 0.4 freezes. curator/authors/provenance stay working;
folding the flat authors into the structured record is a 1.0-cleanup candidate. authoring_reference()
picks up the new vocabularies automatically.
Charter check: data-agnostic ✓ (module metadata, not sample data); declarative ✓; P5 — this is
the axis-unbundling; P6 — role/kind are frozenset vocabularies; the human-authorability gate is
met by keeping the whole block optional and collapsing it to a single entry for the common
one-AI-author case, so a module never reads like an enterprise audit ledger. Like panel, it is
manifest metadata and is not reconstructed by the lossy parquet→spec reverse_module (which rebuilds
a content skeleton) — the durable per-version record is the manifest itself, which is correct and no P7
issue (P7 governs artifact columns). → RM14 (shipped).
Verdict: FIXED — was a silent round-trip GAP introduced by the variant_key column, closed by
keying annotations.parquet on the variant-effect pair (variant_key, conclusion, negatives). No
DSL change: the author still writes ordinary variants.csv rows. The fix is entirely in how the
compiler dedups and rejoins annotation.
The scenario. rs334 (HBB, GAG→GTG, β-globin Glu6Val) is the textbook antagonistic-pleiotropy
locus: the same variant produces categorically different phenotypes by genotype. The carrier is
malaria-resistant; the homozygote has sickle-cell disease. Authored, that is two informative genotype
rows at one locus:
rsid,genotype,state,conclusion,gene,phenotype,category
rs334,A/A,ref,No HbS allele — no sickle phenotype,HBB,Normal hemoglobin,hematologic
rs334,A/T,protective,Sickle-cell trait — resistance to severe P. falciparum malaria,HBB,Malaria resistance,infectious-disease
rs334,T/T,risk,Sickle-cell anemia (HbSS) — chronic hemolysis and vaso-occlusion,HBB,Sickle-cell disease,hematologicThe A/T and T/T rows share one variant_key (rs334) but carry different conclusion,
phenotype, and category — infectious-disease (a protective trait) versus hematologic (a
disease). The effects genuinely do not live in one category: category does not subsume them.
Why one-row-per-variant was wrong (the reasoning). weights.parquet is keyed on
(variant_key, genotype), so each genotype row is faithfully distinct there. But
annotations.parquet — which carries gene/phenotype/category and exists so a consumer can read a
variant's annotation without scanning every genotype row — was deduplicated on variant_key alone.
That silently asserts "a variant has one annotation." For a genuine poly-effect variant it is false:
the second row (T/T) collapsed onto the first met (A/T), and on reverse_module every rs334
row was rewritten with the surviving row's phenotype/category. The homozygote's Sickle-cell disease / hematologic became Malaria resistance / infectious-disease — a confident, silent
inversion of clinical meaning, and a Principle-7 (lossless round-trip) violation. This is not exotic:
the same shape recurs wherever developmental / neural loci are pleiotropic and a single category tag
cannot hold the effect. The bug was introduced with variant_key — before it, dedup keyed on rsid
and had the same latent flaw, just less visible.
The honest identity of an annotation-bearing row is therefore the variant-effect pair, not the
variant: variant + effect, where the effect is (conclusion, negatives). (It has to be conclusion,
not genotype: annotations.parquet is per-variant-effect, and two genotypes that share an effect
should still share one annotation row — dedup on the effect, not on the trigger.)
The mechanics (what actually changed).
- Dedup key.
_build_annotationsnow dedups on(variant_key, conclusion, negatives)— one row per genuine variant-effect pair (first occurrence wins). TheA/TandT/Teffects survive as two rows; a truly identical repeat still collapses. - Self-joinable table.
annotations.parquetnow carriesconclusionandnegatives(alongsidevariant_key), so the table can be rejoined toweights.parqueton the exact pair.weightsalready carriesvariant_key/conclusion/negativesper row, so no newweightscolumn is needed. - Reverse probes the same key.
reverse_modulerebuilds each variant row's(variant_key, conclusion, negatives)triple from itsweightsrow and looks up its own annotation — soT/TgetsSickle-cell disease/hematologicback, notA/T's. An older artifact whoseannotationslacks aconclusioncolumn falls back to the legacyvariant_key-only probe (backward-compatible read). - Digest.
artifact.digestmoved once becauseannotations.parquetgained two columns — free at the time, since 0.4 was still unpublished (Principle 4); determinism + round-trip are the held invariants. That window is gone: 0.5.0 published on 2026-08-07, so a further column would be 1.0.
Charter check: data-agnostic ✓ (still pure annotation — no measurement; the sample's genotype is
supplied by the consumer at query time); declarative ✓; P5 — this unbundles an overloaded identity
(variant ≠ variant-effect); P7 — the whole point is restoring lossless round-trip, proven by
test_poly_effect_annotation_survives_roundtrip (both effects survive and the digest is a fixed
point). Human-authorability gate ✓: the author writes plain genotype→conclusion rows and never sees the
key; the machinery is entirely compiler-side. See COMPILER.md §"Intentionally
unimplemented" item 5 (the reverse_module boundary) and the SNV example in
REFERENCE_EXAMPLES.md §1.
Four use cases that the SNP core, as of 0.4, could not serve at all. None is a gap in the annotation model: every one of them needs a reference number about a population, which a module had nowhere to put. They are the reason the frequency and gene-constraint sidecars exist.
6a. "Is this variant actually rare?" — offline carrier-frequency context. A carrier-screening module
lists pathogenic HBB alleles. A consumer showing a positive call wants to say how common that allele
is, and in which ancestry group — sickle-cell's HBB 11:5227002 T>A sits near 4.8% in African
ancestry and near zero in Finnish. Before 0.5 the module could carry the annotation but not the
frequency, so either the consumer fetched gnomAD at query time (network at read time, and a different
number than the curator saw) or it said nothing. Closed additively: frequencies.csv → one row per
(allele, ancestry group). Enabled.
6b. Reproducing an ACMG BA1/BS1 filter against the frequency the curator saw. BA1 ("allele frequency
too high for a Mendelian disorder") and BS1 are frequency thresholds, and applying them needs the
filtering allele frequency (faf95), not the point estimate. The blocker was never the threshold — a
consumer can apply that — it was that a re-run months later hit a different gnomAD release and silently
reclassified variants. Closed additively: faf95 is carried on its owning ancestry group's row, and
dataset names the release, so the filter is reproducible against the numbers the curator used rather
than against whatever the API serves today. This is why dataset is inside the fact set. Enabled.
6c. Out-of-ancestry caveats. A risk annotation derived from a European cohort applied to a South Asian sample may rest on an allele that is common in one group and absent in the other. The module cannot decide what to do about that — that is the consumer's disclosure policy, and the format never makes the call (the data-agnostic north star). What it can now do is carry the per-group numbers so the consumer has something to reason with. Enabled (format supplies the table; consumer supplies the policy).
6d. Gene-level triage on a cardio or cancer panel. A forty-variant panel spanning a dozen genes:
which genes are haploinsufficient, and which tolerate loss of function? LOEUF separates them (MYH7 at
0.64 is constrained; a tolerant gene sits above 1), and pLI and missense-Z refine it. Repeating a
gene-level fact on every variant row would be the wrong shape — same gene, forty copies, one axis
smeared across another. Closed additively: gene_metrics.csv, one row per gene, a separate table
(Principle 5). Enabled.
What stayed out. Sex-stratified counts (a second axis — folding nfe_XX into population would be
the state-overloading mistake again), and an offline frequency snapshot (58 GB exomes / 742 GB
genomes — not a thing that ships; parked in ROADMAP.md).
The gaps above, consolidated. Format-side items migrate into ROADMAP.md; the
consumer-side one is recorded so it is not mistaken for a format task.
| # | Item | Kind | Unblocks | Priority |
|---|---|---|---|---|
| RM1 | ✅ shipped — compiler materializes all 0.4 tables → parquet with lossless round-trip (generic _build_table/_write_table_csv over _TABLE_KINDS) |
format (compiler) | 3a, 3c, harness on binned loci | done |
| RM2 | ✅ shipped — composed modules: variants.csv optional, a module carries only the kinds it uses (no empty variants.csv); studies.csv required iff variants present |
format (compiler) | SNP+PRS, personal panels | done |
| RM3 | ✅ shipped in 0.4 sample — PharmVariantRow (pharm_variants.csv) + drug/response/evidence_level on DiplotypeRow |
format (schema) | 2b | done |
| RM20 | ✅ shipped in 0.5 — PharmGKB annotations are per-genotype and per-category: genotype, phenotype_category (closed vocab) and annotation_id on PharmVariantRow; duplicate key (variant_key, drug, genotype, phenotype_category, annotation_id). Corrects RM3, which was validated against a sample rather than the corpus. |
format (schema + compiler) | 2b, the real ClinPGx corpus | done |
| RM21 | ✅ shipped in 0.5 — Data-source licensing as data (sources.csv + manifest.sources): per (source, layer) licence, attribution, pinned license_sha256, tri-state share_alike/commercial_use, and the acquirer's declared_use. Compiler refuses annotation-layer content that forbids sale when no declaration is recorded; enricher refuses at acquisition. |
format (schema + compiler) + enricher | 2c, marketplace redistribution | done |
| RM22 | ✅ shipped in 0.5 — PGx tables join resolution: enrich() reads pharm_variants.csv and haplotypes.csv, so a module with no variants.csv gets coordinates (it previously enriched to an empty resolution.csv). |
enricher | 2c, 3c | done |
| RM4 | Native ClinVar gene-panel materialization + content-pinned reference mixin (item 7 follow-up) | format (compiler) + consumer ref | 2a (native path) | medium |
| RM5 | Symbolic/structural alleles (<S>/<L>/<DEL>/<INS>/<DUP>/<STR>; large indels) — a representation beyond ^[ACGT]+$. Motivating case: 5-HTTLPR (S/L not nucleotides → rejected today) |
format (schema) | 3b (SV), 1b (symbolic consume), 5-HTTLPR | medium |
| RM6 | Promote requires_callable to a typed boolean column; reserve/build callable_from (DP,GQ,FT three-state) |
format (schema) | 1c callability | low-medium |
| RM7 | Evaluation-output / report-card schema for the verification harness | consumer (just-dna-lite), NOT the format |
1a | — (not a format task) |
| RM8 | ✅ shipped in 0.4 sample — reference.authoring_reference() + json_schemas(), generated from the live models |
format (schema) | 1d drift | done |
| RM9 | ✅ shipped in 0.4 sample — manifest.RECOMMENDED_COLORS/RECOMMENDED_ICONS |
format (schema) | 1d palette | done |
| RM10 | Optional declarative inheritance-expectation field (trio/de-novo assertion as data) | format (schema) | 1c trio | low (only if needed) |
| RM11 | ✅ shipped in 0.4 — doi provenance column on StudyRow (optional; validated against the DOI grammar, kept verbatim) |
format (schema) | 4a | done |
| RM12 | ✅ shipped in 0.4 — Provenance locator: optional provenance_quote (keyword phrase) + provenance_regex (author-time-compiled, matched by a consumer-side linear-time engine — P1 pattern grammar) on StudyRow |
format (schema) | 4a | done |
| RM13 | just-module-validator — network-first source-check/enrichment library |
consumer (new sibling), NOT the format | 4a | — (not a format task) |
| RM18 | ✅ shipped in 0.5 — Population-frequency + gene-constraint sidecars (frequencies.csv, gene_metrics.csv), produced by the enricher's gnomAD v4.1 passes, compiled to their own optional parquets and fact-hashed. Retires the planned allele_frequency/af_population axes in favour of tables. |
format (schema + compiler) + enricher | 6a–6d | done |
| RM19 | ✅ shipped in 0.5 — GA4GH VRS allele identity: stdlib derive_vrs_allele_id, vrs_id/caid cross-reference columns, and variant_key deriving from the VA for a resolved substitution. Satisfies RM15's build-naming condition (GRCh38-only now; multi-build minting remains RM15). |
format (schema + compiler) + enricher | build-naming identity, cross-database joins | done |
| RM14 | ✅ shipped in 0.4 — Structured per-version authorship (authorship: [Contribution]): {who, role, kind, at}; role closed {created/edited/audited/reviewed}; kind open, seed = human ladder {human, human_expert, human_certified} / {ai}+scale {agent,team,swarm} (no hybrid — joint = two entries). Manifest metadata → digest-neutral. |
format (schema) | 4a validator, marketplace review | done |
Takeaway. The two load-bearing items — RM1 + RM2 (compiler materialization + composed
modules) — are now shipped: the frozen 0.4 shapes are runnable artifacts, and every
composite/personal module compiles with lossless round-trip. What remains open is small and clearly
scoped: RM3-adjacent extensions, RM5 (symbolic alleles / 5-HTTLPR), RM6/RM10 refinements, and the two
provenance anchors RM11/RM12 (doi + fulltext locator) that let a network-first validator scrutinise
a module without the format ever fetching. Notably,
the format's purpose expansion (the verification harness) still needs no format change — it rides
on the properties already frozen, now with the tables materialized under it.