Skip to content

Commit 0f773ea

Browse files
bench: add extraction model A/B testing and results
Introduce a dedicated script to isolate and benchmark the extraction phase across different LLMs. To ensure a clean comparison, the bench and eval commands now support pinning the answer and judge models to fixed oracles, preventing variations in reasoning or scoring from masking the quality of the extracted facts.
1 parent a23756d commit 0f773ea

8 files changed

Lines changed: 347 additions & 2 deletions

bench/results/extraction_ab.log

Lines changed: 149 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,149 @@
1+
=== [22:13:40] EVAL phase (extraction recall/precision) ===
2+
--- [22:13:40] eval google/gemini-2.5-flash-lite ---
3+
running 18 fixtures, model=google/gemini-2.5-flash-lite, base=https://openrouter.ai/api/v1, k=5
4+
5+
# mneme eval results
6+
7+
model: `google/gemini-2.5-flash-lite` · embedder: `openai` · search k: 5
8+
9+
| version | recall | precision | specificity | search@k | dedup | aggregate |
10+
|---|---|---|---|---|---|---|
11+
| v1 | 0.94 | 1.00 | 0.93 | 0.94 | 0.81 | **0.93** |
12+
13+
## Per-fixture (version v1)
14+
15+
| fixture | recall | precision | spec | search | new-on-redup | aggregate |
16+
|---|---|---|---|---|---|---|
17+
| decision | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
18+
| health-fact | 0.00 | 1.00 | 0.00 | 0.00 | 2 | 0.20 |
19+
| meaning-preservation-direction | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
20+
| meaning-preservation-negation | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
21+
| multi-speaker | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
22+
| multi-topic-three | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
23+
| no-fabrication | 1.00 | 1.00 | — | 1.00 | 0 | 1.00 |
24+
| nothing-to-extract | 1.00 | 1.00 | — | — | 0 | 1.00 |
25+
| plans-and-dates | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
26+
| preferences | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
27+
| promotion-and-pet | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
28+
| pronoun-resolution | 1.00 | 1.00 | 1.00 | 1.00 | 1 | 0.80 |
29+
| question-not-fact | 1.00 | 1.00 | — | — | 0 | 1.00 |
30+
| relationships | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
31+
| relative-date-grounding | 1.00 | 1.00 | 1.00 | 1.00 | 1 | 0.80 |
32+
| single-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
33+
| specificity-numbers | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
34+
| specificity-product | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
35+
36+
wrote results to eval/results/extract_google_gemini-2.5-flash-lite.md
37+
eval google/gemini-2.5-flash-lite OK
38+
--- [22:16:40] eval google/gemini-2.5-flash ---
39+
running 18 fixtures, model=google/gemini-2.5-flash, base=https://openrouter.ai/api/v1, k=5
40+
41+
# mneme eval results
42+
43+
model: `google/gemini-2.5-flash` · embedder: `openai` · search k: 5
44+
45+
| version | recall | precision | specificity | search@k | dedup | aggregate |
46+
|---|---|---|---|---|---|---|
47+
| v1 | 1.00 | 0.94 | 1.00 | 1.00 | 1.00 | **0.99** |
48+
49+
## Per-fixture (version v1)
50+
51+
| fixture | recall | precision | spec | search | new-on-redup | aggregate |
52+
|---|---|---|---|---|---|---|
53+
| decision | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
54+
| health-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
55+
| meaning-preservation-direction | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
56+
| meaning-preservation-negation | 1.00 | 0.67 | 1.00 | 1.00 | 0 | 0.93 |
57+
| multi-speaker | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
58+
| multi-topic-three | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
59+
| no-fabrication | 1.00 | 1.00 | — | 1.00 | 0 | 1.00 |
60+
| nothing-to-extract | 1.00 | 1.00 | — | — | 0 | 1.00 |
61+
| plans-and-dates | 1.00 | 0.67 | 1.00 | 1.00 | 0 | 0.93 |
62+
| preferences | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
63+
| promotion-and-pet | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
64+
| pronoun-resolution | 1.00 | 0.67 | 1.00 | 1.00 | 0 | 0.93 |
65+
| question-not-fact | 1.00 | 1.00 | — | — | 0 | 1.00 |
66+
| relationships | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
67+
| relative-date-grounding | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
68+
| single-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
69+
| specificity-numbers | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
70+
| specificity-product | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
71+
72+
wrote results to eval/results/extract_google_gemini-2.5-flash.md
73+
eval google/gemini-2.5-flash OK
74+
--- [22:19:15] eval openai/gpt-5-mini ---
75+
running 18 fixtures, model=openai/gpt-5-mini, base=https://openrouter.ai/api/v1, k=5
76+
77+
# mneme eval results
78+
79+
model: `openai/gpt-5-mini` · embedder: `openai` · search k: 5
80+
81+
| version | recall | precision | specificity | search@k | dedup | aggregate |
82+
|---|---|---|---|---|---|---|
83+
| v1 | 0.94 | 0.83 | 1.00 | 1.00 | 1.00 | **0.94** |
84+
85+
## Per-fixture (version v1)
86+
87+
| fixture | recall | precision | spec | search | new-on-redup | aggregate |
88+
|---|---|---|---|---|---|---|
89+
| decision | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
90+
| health-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
91+
| meaning-preservation-direction | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
92+
| meaning-preservation-negation | 1.00 | 0.33 | 1.00 | 1.00 | 0 | 0.87 |
93+
| multi-speaker | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
94+
| multi-topic-three | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
95+
| no-fabrication | 1.00 | 1.00 | — | 1.00 | 0 | 1.00 |
96+
| nothing-to-extract | 1.00 | 1.00 | — | — | 0 | 1.00 |
97+
| plans-and-dates | 0.00 | 0.00 | 1.00 | 1.00 | 0 | 0.60 |
98+
| preferences | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
99+
| promotion-and-pet | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
100+
| pronoun-resolution | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
101+
| question-not-fact | 1.00 | 0.00 | — | — | 0 | 0.50 |
102+
| relationships | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
103+
| relative-date-grounding | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
104+
| single-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
105+
| specificity-numbers | 1.00 | 0.67 | 1.00 | 1.00 | 0 | 0.93 |
106+
| specificity-product | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
107+
108+
wrote results to eval/results/extract_openai_gpt-5-mini.md
109+
eval openai/gpt-5-mini OK
110+
--- [22:25:21] eval deepseek/deepseek-v3.2 ---
111+
running 18 fixtures, model=deepseek/deepseek-v3.2, base=https://openrouter.ai/api/v1, k=5
112+
113+
# mneme eval results
114+
115+
model: `deepseek/deepseek-v3.2` · embedder: `openai` · search k: 5
116+
117+
| version | recall | precision | specificity | search@k | dedup | aggregate |
118+
|---|---|---|---|---|---|---|
119+
| v1 | 1.00 | 0.89 | 1.00 | 1.00 | 0.88 | **0.92** |
120+
121+
## Per-fixture (version v1)
122+
123+
| fixture | recall | precision | spec | search | new-on-redup | aggregate |
124+
|---|---|---|---|---|---|---|
125+
| decision | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
126+
| health-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
127+
| meaning-preservation-direction | 1.00 | 1.00 | 1.00 | 1.00 | 1 | 0.80 |
128+
| meaning-preservation-negation | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
129+
| multi-speaker | 1.00 | 1.00 | 1.00 | 1.00 | 1 | 0.80 |
130+
| multi-topic-three | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
131+
| no-fabrication | 1.00 | 1.00 | — | 1.00 | 0 | 1.00 |
132+
| nothing-to-extract | 1.00 | 0.00 | — | — | 0 | 0.50 |
133+
| plans-and-dates | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
134+
| preferences | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
135+
| promotion-and-pet | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
136+
| pronoun-resolution | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
137+
| question-not-fact | 1.00 | 0.00 | — | — | 0 | 0.50 |
138+
| relationships | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
139+
| relative-date-grounding | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
140+
| single-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
141+
| specificity-numbers | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
142+
| specificity-product | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
143+
144+
wrote results to eval/results/extract_deepseek_deepseek-v3.2.md
145+
eval deepseek/deepseek-v3.2 OK
146+
=== [22:28:44] BENCH phase (end-to-end LoCoMo) ===
147+
--- [22:28:44] bench extract=google/gemini-2.5-flash (answer+judge=google/gemini-2.5-flash-lite) ---
148+
running locomo: 10 samples, 1986 questions, model=google/gemini-2.5-flash, embedder=openai, k=5, strategy=additive
149+
scored 1/10 samples scored 2/10 samples scored 3/10 samples scored 4/10 samples scored 5/10 samples

bench/run_extraction_ab.sh

Lines changed: 53 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,53 @@
1+
#!/usr/bin/env bash
2+
# run_extraction_ab.sh — isolate the EXTRACTION model. Everything else is held
3+
# constant: embedder=text-embedding-3-small, additive, k=5, and (for bench) the
4+
# answer + judge models are pinned to gemini-2.5-flash-lite so the only variable
5+
# is which model writes facts to the store. Two views:
6+
# eval/ — direct extraction recall/precision on 18 fixtures (judge pinned to
7+
# gemini-2.5-flash, a fixed oracle stronger than the models under test)
8+
# bench/ — end-to-end LoCoMo Judge (does better extraction => better QA?)
9+
# Sequential, resilient (one failure logs and continues), telegram on completion.
10+
set -uo pipefail
11+
cd "$(dirname "$0")/.."
12+
set -a; . ./.env; set +a
13+
14+
DATA="bench/data/locomo10.json"
15+
EMBED="text-embedding-3-small"
16+
ANSWER="google/gemini-2.5-flash-lite" # bench answer model, pinned
17+
EVAL_JUDGE="google/gemini-2.5-flash" # eval oracle, pinned (stronger, cheap on 18 fixtures)
18+
EO="eval/results"; BO="bench/results"; mkdir -p "$EO" "$BO"
19+
LOG="$BO/extraction_ab.log"; : > "$LOG"
20+
21+
# Extraction models under test.
22+
MODELS=( "google/gemini-2.5-flash-lite" "google/gemini-2.5-flash" "openai/gpt-5-mini" "deepseek/deepseek-v3.2" )
23+
24+
slug() { echo "$1" | tr '/' '_'; }
25+
26+
echo "=== [$(date +%H:%M:%S)] EVAL phase (extraction recall/precision) ===" | tee -a "$LOG"
27+
for m in "${MODELS[@]}"; do
28+
s=$(slug "$m")
29+
echo "--- [$(date +%H:%M:%S)] eval $m ---" | tee -a "$LOG"
30+
if MNEME_EMBED_MODEL="$EMBED" go run ./cmd/eval -embedder openai -k 5 \
31+
-model "$m" -judge-model "$EVAL_JUDGE" -out "$EO/extract_$s.md" >>"$LOG" 2>&1; then
32+
echo " eval $m OK" | tee -a "$LOG"
33+
else
34+
echo " eval $m FAILED (continuing)" | tee -a "$LOG"
35+
fi
36+
done
37+
38+
echo "=== [$(date +%H:%M:%S)] BENCH phase (end-to-end LoCoMo) ===" | tee -a "$LOG"
39+
# flash-lite extraction == existing bench/results/additive_3small.md, skip it.
40+
for m in "google/gemini-2.5-flash" "openai/gpt-5-mini" "deepseek/deepseek-v3.2"; do
41+
s=$(slug "$m")
42+
echo "--- [$(date +%H:%M:%S)] bench extract=$m (answer+judge=$ANSWER) ---" | tee -a "$LOG"
43+
if MNEME_EMBED_MODEL="$EMBED" go run ./cmd/bench -dataset locomo -path "$DATA" -k 5 -concurrency 8 \
44+
-strategy additive -model "$m" -answer-model "$ANSWER" -judge-model "$ANSWER" \
45+
-out "$BO/extract_$s.md" >>"$LOG" 2>&1; then
46+
echo " bench $m OK" | tee -a "$LOG"
47+
else
48+
echo " bench $m FAILED (continuing)" | tee -a "$LOG"
49+
fi
50+
done
51+
52+
echo "=== [$(date +%H:%M:%S)] EXTRACTION_AB COMPLETE ===" | tee -a "$LOG"
53+
python3 ~/telegram-bridge/tgctl.py notify "✅ mneme extraction-model A/B done! eval + bench results ready." --label mneme-bench 2>/dev/null || true

cmd/bench/main.go

Lines changed: 15 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -45,7 +45,9 @@ func run() error {
4545
dataset = flag.String("dataset", "locomo", "dataset: locomo | longmemeval")
4646
path = flag.String("path", "", "path to the dataset file (required; gitignored)")
4747
base = flag.String("base", envOr("MNEME_LLM_BASE_URL", "https://openrouter.ai/api/v1"), "OpenAI-compatible base URL")
48-
model = flag.String("model", envOr("MNEME_LLM_MODEL", "openai/gpt-4o-mini"), "extraction/answer/judge model")
48+
model = flag.String("model", envOr("MNEME_LLM_MODEL", "openai/gpt-4o-mini"), "extraction model under test (also answers/judges unless those are pinned below)")
49+
answerMdl = flag.String("answer-model", "", "model that answers questions from retrieved facts; defaults to -model")
50+
judgeMdl = flag.String("judge-model", "", "model for the semantic judge; defaults to -model. Pin both of these when A/B-ing extraction so answer+oracle stay constant")
4951
embedKind = flag.String("embedder", "openai", "embedder: openai | fake")
5052
k = flag.Int("k", 5, "search top-k facts fed to the answer model")
5153
strategy = flag.String("strategy", bench.StrategyAdditive, "write strategy: additive | consolidate")
@@ -86,6 +88,16 @@ func run() error {
8688
}
8789
llm := &openai.LLM{BaseURL: *base, APIKey: key, Model: *model}
8890

91+
// Pin answer + judge to fixed models when A/B-ing extraction, so the only
92+
// thing that varies between runs is the model that writes facts to the store.
93+
var answerLLM, judgeLLM provider.LLM
94+
if *answerMdl != "" && *answerMdl != *model {
95+
answerLLM = &openai.LLM{BaseURL: *base, APIKey: key, Model: *answerMdl}
96+
}
97+
if *judgeMdl != "" && *judgeMdl != *model {
98+
judgeLLM = &openai.LLM{BaseURL: *base, APIKey: key, Model: *judgeMdl}
99+
}
100+
89101
var embedder provider.Embedder
90102
embedLabel := *embedKind
91103
switch *embedKind {
@@ -127,6 +139,8 @@ func run() error {
127139

128140
report, err := bench.Run(ctx, samples, bench.Config{
129141
LLM: llm,
142+
AnswerLLM: answerLLM,
143+
Judge: judgeLLM,
130144
Embedder: embedder,
131145
Store: st,
132146
K: *k,

cmd/eval/main.go

Lines changed: 10 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -42,7 +42,8 @@ func run() error {
4242
var (
4343
fixturesDir = flag.String("fixtures", "eval/fixtures", "directory of *.json fixtures")
4444
base = flag.String("base", envOr("MNEME_LLM_BASE_URL", "https://openrouter.ai/api/v1"), "OpenAI-compatible base URL")
45-
model = flag.String("model", envOr("MNEME_LLM_MODEL", "openai/gpt-4o-mini"), "extraction/judge model")
45+
model = flag.String("model", envOr("MNEME_LLM_MODEL", "openai/gpt-4o-mini"), "extraction model under test")
46+
judgeModel = flag.String("judge-model", "", "model for the semantic judge; defaults to -model. Pin it (to a fixed model) when A/B-ing extraction models so the oracle is constant across runs")
4647
embedKind = flag.String("embedder", "fake", "embedder: fake | openai")
4748
k = flag.Int("k", 3, "search top-k for recall@k")
4849
out = flag.String("out", "", "optional results file to write (markdown)")
@@ -58,6 +59,13 @@ func run() error {
5859

5960
llm := &openai.LLM{BaseURL: *base, APIKey: key, Model: *model}
6061

62+
// Pin the judge to a fixed model when A/B-ing extraction models, so the
63+
// oracle scoring recall/precision is constant and the comparison is clean.
64+
var judgeLLM provider.LLM
65+
if *judgeModel != "" && *judgeModel != *model {
66+
judgeLLM = &openai.LLM{BaseURL: *base, APIKey: key, Model: *judgeModel}
67+
}
68+
6169
var embedder provider.Embedder
6270
switch *embedKind {
6371
case "fake":
@@ -94,6 +102,7 @@ func run() error {
94102

95103
reports, err := eval.Run(ctx, fixtures, eval.Config{
96104
LLM: llm,
105+
Judge: judgeLLM,
97106
Embedder: embedder,
98107
Store: st,
99108
K: *k,
Lines changed: 30 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,30 @@
1+
# mneme eval results
2+
3+
model: `deepseek/deepseek-v3.2` · embedder: `openai` · search k: 5
4+
5+
| version | recall | precision | specificity | search@k | dedup | aggregate |
6+
|---|---|---|---|---|---|---|
7+
| v1 | 1.00 | 0.89 | 1.00 | 1.00 | 0.88 | **0.92** |
8+
9+
## Per-fixture (version v1)
10+
11+
| fixture | recall | precision | spec | search | new-on-redup | aggregate |
12+
|---|---|---|---|---|---|---|
13+
| decision | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
14+
| health-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
15+
| meaning-preservation-direction | 1.00 | 1.00 | 1.00 | 1.00 | 1 | 0.80 |
16+
| meaning-preservation-negation | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
17+
| multi-speaker | 1.00 | 1.00 | 1.00 | 1.00 | 1 | 0.80 |
18+
| multi-topic-three | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
19+
| no-fabrication | 1.00 | 1.00 || 1.00 | 0 | 1.00 |
20+
| nothing-to-extract | 1.00 | 0.00 ||| 0 | 0.50 |
21+
| plans-and-dates | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
22+
| preferences | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
23+
| promotion-and-pet | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
24+
| pronoun-resolution | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
25+
| question-not-fact | 1.00 | 0.00 ||| 0 | 0.50 |
26+
| relationships | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
27+
| relative-date-grounding | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
28+
| single-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
29+
| specificity-numbers | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
30+
| specificity-product | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
Lines changed: 30 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,30 @@
1+
# mneme eval results
2+
3+
model: `google/gemini-2.5-flash-lite` · embedder: `openai` · search k: 5
4+
5+
| version | recall | precision | specificity | search@k | dedup | aggregate |
6+
|---|---|---|---|---|---|---|
7+
| v1 | 0.94 | 1.00 | 0.93 | 0.94 | 0.81 | **0.93** |
8+
9+
## Per-fixture (version v1)
10+
11+
| fixture | recall | precision | spec | search | new-on-redup | aggregate |
12+
|---|---|---|---|---|---|---|
13+
| decision | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
14+
| health-fact | 0.00 | 1.00 | 0.00 | 0.00 | 2 | 0.20 |
15+
| meaning-preservation-direction | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
16+
| meaning-preservation-negation | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
17+
| multi-speaker | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
18+
| multi-topic-three | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
19+
| no-fabrication | 1.00 | 1.00 || 1.00 | 0 | 1.00 |
20+
| nothing-to-extract | 1.00 | 1.00 ||| 0 | 1.00 |
21+
| plans-and-dates | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
22+
| preferences | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
23+
| promotion-and-pet | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
24+
| pronoun-resolution | 1.00 | 1.00 | 1.00 | 1.00 | 1 | 0.80 |
25+
| question-not-fact | 1.00 | 1.00 ||| 0 | 1.00 |
26+
| relationships | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
27+
| relative-date-grounding | 1.00 | 1.00 | 1.00 | 1.00 | 1 | 0.80 |
28+
| single-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
29+
| specificity-numbers | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
30+
| specificity-product | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
Lines changed: 30 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,30 @@
1+
# mneme eval results
2+
3+
model: `google/gemini-2.5-flash` · embedder: `openai` · search k: 5
4+
5+
| version | recall | precision | specificity | search@k | dedup | aggregate |
6+
|---|---|---|---|---|---|---|
7+
| v1 | 1.00 | 0.94 | 1.00 | 1.00 | 1.00 | **0.99** |
8+
9+
## Per-fixture (version v1)
10+
11+
| fixture | recall | precision | spec | search | new-on-redup | aggregate |
12+
|---|---|---|---|---|---|---|
13+
| decision | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
14+
| health-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
15+
| meaning-preservation-direction | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
16+
| meaning-preservation-negation | 1.00 | 0.67 | 1.00 | 1.00 | 0 | 0.93 |
17+
| multi-speaker | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
18+
| multi-topic-three | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
19+
| no-fabrication | 1.00 | 1.00 || 1.00 | 0 | 1.00 |
20+
| nothing-to-extract | 1.00 | 1.00 ||| 0 | 1.00 |
21+
| plans-and-dates | 1.00 | 0.67 | 1.00 | 1.00 | 0 | 0.93 |
22+
| preferences | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
23+
| promotion-and-pet | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
24+
| pronoun-resolution | 1.00 | 0.67 | 1.00 | 1.00 | 0 | 0.93 |
25+
| question-not-fact | 1.00 | 1.00 ||| 0 | 1.00 |
26+
| relationships | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
27+
| relative-date-grounding | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
28+
| single-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
29+
| specificity-numbers | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
30+
| specificity-product | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
Lines changed: 30 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,30 @@
1+
# mneme eval results
2+
3+
model: `openai/gpt-5-mini` · embedder: `openai` · search k: 5
4+
5+
| version | recall | precision | specificity | search@k | dedup | aggregate |
6+
|---|---|---|---|---|---|---|
7+
| v1 | 0.94 | 0.83 | 1.00 | 1.00 | 1.00 | **0.94** |
8+
9+
## Per-fixture (version v1)
10+
11+
| fixture | recall | precision | spec | search | new-on-redup | aggregate |
12+
|---|---|---|---|---|---|---|
13+
| decision | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
14+
| health-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
15+
| meaning-preservation-direction | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
16+
| meaning-preservation-negation | 1.00 | 0.33 | 1.00 | 1.00 | 0 | 0.87 |
17+
| multi-speaker | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
18+
| multi-topic-three | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
19+
| no-fabrication | 1.00 | 1.00 || 1.00 | 0 | 1.00 |
20+
| nothing-to-extract | 1.00 | 1.00 ||| 0 | 1.00 |
21+
| plans-and-dates | 0.00 | 0.00 | 1.00 | 1.00 | 0 | 0.60 |
22+
| preferences | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
23+
| promotion-and-pet | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
24+
| pronoun-resolution | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
25+
| question-not-fact | 1.00 | 0.00 ||| 0 | 0.50 |
26+
| relationships | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
27+
| relative-date-grounding | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
28+
| single-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
29+
| specificity-numbers | 1.00 | 0.67 | 1.00 | 1.00 | 0 | 0.93 |
30+
| specificity-product | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |

0 commit comments

Comments
 (0)