docs(site): render the 4-system × Cat 1-9 cross-system comparison matrix - #231
docs(site): render the 4-system × Cat 1-9 cross-system comparison matrix#231jphein wants to merge 1 commit into
Conversation
JP asked 'where's the matrix comparing ALL memory systems using multipass?' — the cross-system data existed only in the markdown doc + prose, never as a rendered table. This adds it as a new article in the multipass region, complementary to the existing MemPalace-only single-substrate matrix. New table: rows = Cat 1/2c/3/4/5/6/7/7b/8/9a/9b; columns = MemPalace | OMEGA | Hindsight | Mem0-OSS. Cells carry the reading (final numbers) or N/A-with-reason. Matches the site's cat-tbl aesthetic + sortable Cat header (auto-wired by the existing table.cat-tbl sort init). Key framing made visible: - MemPalace + OMEGA take the full Cat 1-9 (both graph-native). - Hindsight + Mem0-OSS read N/A on structural cats 3-8 (no queryable graph) — distinct causes: Hindsight no graph endpoint; Mem0-OSS dropped its graph layer (hosted platform keeps it). 'N/A is a finding, not a blank.' - Capstone caveat: emergent contradiction/supersession is ~0 across BOTH graph-native systems (mp 0/0 edges; OMEGA contradiction recall 0.50, supersession 0.00) — the open frontier SME is alone in even asking. Uses the FINAL post-re-map mempalace numbers (Cat 4 26.83%/0.645/40 types, Cat 5 61.87%/22.2%, Cat 8 modularity 0.7961) — consistent with the single-substrate matrix. Headless-validated: html.parser clean, 50 articles / 20 tables balanced, new table has 11 data rows, per-cell scoreability faithful to baselines/cross_system_multipass_matrix_2026-05-30.json. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Code Review
This pull request adds a new "Cross-system multipass" comparison matrix to the benchmarks page, comparing four memory systems (mempalace, OMEGA, Hindsight, and Mem0-OSS) across categories 1–9. The reviewer feedback suggests improving visual consistency by replacing plain text labels (such as "QA deferred" and "emergent") in the table cells with their corresponding HTML badge elements defined in the scoreability legend.
| <td class="reading">verified · <strong>QA deferred</strong> <span class="src">extraction-cost-gated</span></td> | ||
| <td class="reading">verified · <strong>QA deferred</strong> <span class="src">extraction-cost-gated</span></td> |
There was a problem hiding this comment.
The scoreability legend defines the <span class="vbadge paper">QA-deferred</span> badge for verified + runnable, extraction-cost-gated runs, but the table cells use plain <strong>QA deferred</strong> text instead. Replacing the plain text with the defined badge improves visual consistency and matches the legend.
| <td class="reading">verified · <strong>QA deferred</strong> <span class="src">extraction-cost-gated</span></td> | |
| <td class="reading">verified · <strong>QA deferred</strong> <span class="src">extraction-cost-gated</span></td> | |
| <td class="reading">verified · <span class="vbadge paper">QA-deferred</span> <span class="src">extraction-cost-gated</span></td> | |
| <td class="reading">verified · <span class="vbadge paper">QA-deferred</span> <span class="src">extraction-cost-gated</span></td> |
| <td class="reading">verified · <strong>QA deferred</strong></td> | ||
| <td class="reading">verified · <strong>QA deferred</strong></td> |
There was a problem hiding this comment.
The scoreability legend defines the <span class="vbadge paper">QA-deferred</span> badge for verified + runnable, extraction-cost-gated runs, but the table cells use plain <strong>QA deferred</strong> text instead. Replacing the plain text with the defined badge improves visual consistency and matches the legend.
| <td class="reading">verified · <strong>QA deferred</strong></td> | |
| <td class="reading">verified · <strong>QA deferred</strong></td> | |
| <td class="reading">verified · <span class="vbadge paper">QA-deferred</span></td> | |
| <td class="reading">verified · <span class="vbadge paper">QA-deferred</span></td> |
| <td class="reading"><strong>0.00</strong> emergent <span class="src">0 <code>contradicts</code> edges in the live KG — no emergent contradiction structure (the prior +1.00 was a corpus-declared ceiling)</span></td> | ||
| <td class="reading"><strong>recall 0.50 / prec 0.25</strong> emergent <span class="src">auto-relate generated 4 contradicts edges; caught 1 of 2 ground-truth themes (good-dog)</span></td> |
There was a problem hiding this comment.
The scoreability legend defines the <span class="vbadge self">emergent</span> badge for system-generated edges, but the table cells use plain emergent text instead. Replacing the plain text with the defined badge improves visual consistency and matches the legend.
| <td class="reading"><strong>0.00</strong> emergent <span class="src">0 <code>contradicts</code> edges in the live KG — no emergent contradiction structure (the prior +1.00 was a corpus-declared ceiling)</span></td> | |
| <td class="reading"><strong>recall 0.50 / prec 0.25</strong> emergent <span class="src">auto-relate generated 4 contradicts edges; caught 1 of 2 ground-truth themes (good-dog)</span></td> | |
| <td class="reading"><strong>0.00</strong> <span class="vbadge self">emergent</span> <span class="src">0 <code>contradicts</code> edges in the live KG — no emergent contradiction structure (the prior +1.00 was a corpus-declared ceiling)</span></td> | |
| <td class="reading"><strong>recall 0.50 / prec 0.25</strong> <span class="vbadge self">emergent</span> <span class="src">auto-relate generated 4 contradicts edges; caught 1 of 2 ground-truth themes (good-dog)</span></td> |
| <td class="reading"><strong>0.00</strong> emergent <span class="src">0 <code>supersedes</code> edges in the live KG — completeness 0.00 (prior +1.00 was the declared ceiling)</span></td> | ||
| <td class="reading"><strong>0.00</strong> emergent <span class="src">0 supersedes; its temporal analogue <code>evolution</code> (17 edges) doesn’t normalize to supersedes</span></td> |
There was a problem hiding this comment.
The scoreability legend defines the <span class="vbadge self">emergent</span> badge for system-generated edges, but the table cells use plain emergent text instead. Replacing the plain text with the defined badge improves visual consistency and matches the legend.
| <td class="reading"><strong>0.00</strong> emergent <span class="src">0 <code>supersedes</code> edges in the live KG — completeness 0.00 (prior +1.00 was the declared ceiling)</span></td> | |
| <td class="reading"><strong>0.00</strong> emergent <span class="src">0 supersedes; its temporal analogue <code>evolution</code> (17 edges) doesn’t normalize to supersedes</span></td> | |
| <td class="reading"><strong>0.00</strong> <span class="vbadge self">emergent</span> <span class="src">0 <code>supersedes</code> edges in the live KG — completeness 0.00 (prior +1.00 was the declared ceiling)</span></td> | |
| <td class="reading"><strong>0.00</strong> <span class="vbadge self">emergent</span> <span class="src">0 supersedes; its temporal analogue <code>evolution</code> (17 edges) doesn’t normalize to supersedes</span></td> |
| <td class="reading">verified · <strong>QA deferred</strong> <span class="src">~150h strat150 (cost wall)</span></td> | ||
| <td class="reading">verified · <strong>QA deferred</strong> <span class="src">~18h strat150; extraction lossy by design</span></td> |
There was a problem hiding this comment.
The scoreability legend defines the <span class="vbadge paper">QA-deferred</span> badge for verified + runnable, extraction-cost-gated runs, but the table cells use plain <strong>QA deferred</strong> text instead. Replacing the plain text with the defined badge improves visual consistency and matches the legend.
| <td class="reading">verified · <strong>QA deferred</strong> <span class="src">~150h strat150 (cost wall)</span></td> | |
| <td class="reading">verified · <strong>QA deferred</strong> <span class="src">~18h strat150; extraction lossy by design</span></td> | |
| <td class="reading">verified · <span class="vbadge paper">QA-deferred</span> <span class="src">~150h strat150 (cost wall)</span></td> | |
| <td class="reading">verified · <span class="vbadge paper">QA-deferred</span> <span class="src">~18h strat150; extraction lossy by design</span></td> |
What this does
Renders the 4-system × Cat 1–9 cross-system comparison matrix on the site as a table — closing a real gap JP flagged: "where's the matrix comparing ALL memory systems using multipass?" The cross-system data existed only in
docs/benchmarks/2026-05-30-cross-system-multipass-matrix.md+ prose; the site's multipass matrix was MemPalace-only (single-substrate deep-dive). This adds the comparison table as a complementary article right after it.(Supersedes closed PR #230 — identical content, opened from a properly-named branch off current main to avoid the auto-generated-branch footgun.)
The table
cat-tblaesthetic; the sortable Cat header is auto-wired by the existingtable.cat-tblsort init (no JS change).The framing it makes visible
Numbers
Uses the FINAL post-re-map MemPalace numbers (Cat 4 26.83%/0.645/40 types, Cat 5 61.87%/22.2%, Cat 8 modularity 0.7961) — consistent with the single-substrate matrix. OMEGA/Hindsight/Mem0 cells from cassia-2's cross-system grid (#146).
Validation (headless — no browser launched)
python -m html.parserparses clean: 50<article>/ 20<table>balanced, new table has exactly 11 data rows, per-cell scoreability verified faithful tobaselines/cross_system_multipass_matrix_2026-05-30.json(Cat 1/2c/7 deferred, Cat 3/4/5/6/8 N/A-no-graph, Cat 9a/9b N/A-no-harness for the extraction systems). HTML-only diff.🤖 Generated with Claude Code