1+ === [22:13:40] EVAL phase (extraction recall/precision) ===
2+ --- [22:13:40] eval google/gemini-2.5-flash-lite ---
3+ running 18 fixtures, model=google/gemini-2.5-flash-lite, base=https://openrouter.ai/api/v1, k=5
4+
5+ # mneme eval results
6+
7+ model: `google/gemini-2.5-flash-lite` · embedder: `openai` · search k: 5
8+
9+ | version | recall | precision | specificity | search@k | dedup | aggregate |
10+ |---|---|---|---|---|---|---|
11+ | v1 | 0.94 | 1.00 | 0.93 | 0.94 | 0.81 | **0.93** |
12+
13+ ## Per-fixture (version v1)
14+
15+ | fixture | recall | precision | spec | search | new-on-redup | aggregate |
16+ |---|---|---|---|---|---|---|
17+ | decision | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
18+ | health-fact | 0.00 | 1.00 | 0.00 | 0.00 | 2 | 0.20 |
19+ | meaning-preservation-direction | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
20+ | meaning-preservation-negation | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
21+ | multi-speaker | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
22+ | multi-topic-three | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
23+ | no-fabrication | 1.00 | 1.00 | — | 1.00 | 0 | 1.00 |
24+ | nothing-to-extract | 1.00 | 1.00 | — | — | 0 | 1.00 |
25+ | plans-and-dates | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
26+ | preferences | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
27+ | promotion-and-pet | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
28+ | pronoun-resolution | 1.00 | 1.00 | 1.00 | 1.00 | 1 | 0.80 |
29+ | question-not-fact | 1.00 | 1.00 | — | — | 0 | 1.00 |
30+ | relationships | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
31+ | relative-date-grounding | 1.00 | 1.00 | 1.00 | 1.00 | 1 | 0.80 |
32+ | single-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
33+ | specificity-numbers | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
34+ | specificity-product | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
35+
36+ wrote results to eval/results/extract_google_gemini-2.5-flash-lite.md
37+ eval google/gemini-2.5-flash-lite OK
38+ --- [22:16:40] eval google/gemini-2.5-flash ---
39+ running 18 fixtures, model=google/gemini-2.5-flash, base=https://openrouter.ai/api/v1, k=5
40+
41+ # mneme eval results
42+
43+ model: `google/gemini-2.5-flash` · embedder: `openai` · search k: 5
44+
45+ | version | recall | precision | specificity | search@k | dedup | aggregate |
46+ |---|---|---|---|---|---|---|
47+ | v1 | 1.00 | 0.94 | 1.00 | 1.00 | 1.00 | **0.99** |
48+
49+ ## Per-fixture (version v1)
50+
51+ | fixture | recall | precision | spec | search | new-on-redup | aggregate |
52+ |---|---|---|---|---|---|---|
53+ | decision | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
54+ | health-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
55+ | meaning-preservation-direction | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
56+ | meaning-preservation-negation | 1.00 | 0.67 | 1.00 | 1.00 | 0 | 0.93 |
57+ | multi-speaker | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
58+ | multi-topic-three | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
59+ | no-fabrication | 1.00 | 1.00 | — | 1.00 | 0 | 1.00 |
60+ | nothing-to-extract | 1.00 | 1.00 | — | — | 0 | 1.00 |
61+ | plans-and-dates | 1.00 | 0.67 | 1.00 | 1.00 | 0 | 0.93 |
62+ | preferences | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
63+ | promotion-and-pet | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
64+ | pronoun-resolution | 1.00 | 0.67 | 1.00 | 1.00 | 0 | 0.93 |
65+ | question-not-fact | 1.00 | 1.00 | — | — | 0 | 1.00 |
66+ | relationships | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
67+ | relative-date-grounding | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
68+ | single-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
69+ | specificity-numbers | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
70+ | specificity-product | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
71+
72+ wrote results to eval/results/extract_google_gemini-2.5-flash.md
73+ eval google/gemini-2.5-flash OK
74+ --- [22:19:15] eval openai/gpt-5-mini ---
75+ running 18 fixtures, model=openai/gpt-5-mini, base=https://openrouter.ai/api/v1, k=5
76+
77+ # mneme eval results
78+
79+ model: `openai/gpt-5-mini` · embedder: `openai` · search k: 5
80+
81+ | version | recall | precision | specificity | search@k | dedup | aggregate |
82+ |---|---|---|---|---|---|---|
83+ | v1 | 0.94 | 0.83 | 1.00 | 1.00 | 1.00 | **0.94** |
84+
85+ ## Per-fixture (version v1)
86+
87+ | fixture | recall | precision | spec | search | new-on-redup | aggregate |
88+ |---|---|---|---|---|---|---|
89+ | decision | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
90+ | health-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
91+ | meaning-preservation-direction | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
92+ | meaning-preservation-negation | 1.00 | 0.33 | 1.00 | 1.00 | 0 | 0.87 |
93+ | multi-speaker | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
94+ | multi-topic-three | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
95+ | no-fabrication | 1.00 | 1.00 | — | 1.00 | 0 | 1.00 |
96+ | nothing-to-extract | 1.00 | 1.00 | — | — | 0 | 1.00 |
97+ | plans-and-dates | 0.00 | 0.00 | 1.00 | 1.00 | 0 | 0.60 |
98+ | preferences | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
99+ | promotion-and-pet | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
100+ | pronoun-resolution | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
101+ | question-not-fact | 1.00 | 0.00 | — | — | 0 | 0.50 |
102+ | relationships | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
103+ | relative-date-grounding | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
104+ | single-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
105+ | specificity-numbers | 1.00 | 0.67 | 1.00 | 1.00 | 0 | 0.93 |
106+ | specificity-product | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
107+
108+ wrote results to eval/results/extract_openai_gpt-5-mini.md
109+ eval openai/gpt-5-mini OK
110+ --- [22:25:21] eval deepseek/deepseek-v3.2 ---
111+ running 18 fixtures, model=deepseek/deepseek-v3.2, base=https://openrouter.ai/api/v1, k=5
112+
113+ # mneme eval results
114+
115+ model: `deepseek/deepseek-v3.2` · embedder: `openai` · search k: 5
116+
117+ | version | recall | precision | specificity | search@k | dedup | aggregate |
118+ |---|---|---|---|---|---|---|
119+ | v1 | 1.00 | 0.89 | 1.00 | 1.00 | 0.88 | **0.92** |
120+
121+ ## Per-fixture (version v1)
122+
123+ | fixture | recall | precision | spec | search | new-on-redup | aggregate |
124+ |---|---|---|---|---|---|---|
125+ | decision | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
126+ | health-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
127+ | meaning-preservation-direction | 1.00 | 1.00 | 1.00 | 1.00 | 1 | 0.80 |
128+ | meaning-preservation-negation | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
129+ | multi-speaker | 1.00 | 1.00 | 1.00 | 1.00 | 1 | 0.80 |
130+ | multi-topic-three | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
131+ | no-fabrication | 1.00 | 1.00 | — | 1.00 | 0 | 1.00 |
132+ | nothing-to-extract | 1.00 | 0.00 | — | — | 0 | 0.50 |
133+ | plans-and-dates | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
134+ | preferences | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
135+ | promotion-and-pet | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
136+ | pronoun-resolution | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
137+ | question-not-fact | 1.00 | 0.00 | — | — | 0 | 0.50 |
138+ | relationships | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
139+ | relative-date-grounding | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
140+ | single-fact | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
141+ | specificity-numbers | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
142+ | specificity-product | 1.00 | 1.00 | 1.00 | 1.00 | 0 | 1.00 |
143+
144+ wrote results to eval/results/extract_deepseek_deepseek-v3.2.md
145+ eval deepseek/deepseek-v3.2 OK
146+ === [22:28:44] BENCH phase (end-to-end LoCoMo) ===
147+ --- [22:28:44] bench extract=google/gemini-2.5-flash (answer+judge=google/gemini-2.5-flash-lite) ---
148+ running locomo: 10 samples, 1986 questions, model=google/gemini-2.5-flash, embedder=openai, k=5, strategy=additive
149+ scored 1/10 samples scored 2/10 samples scored 3/10 samples scored 4/10 samples scored 5/10 samples
0 commit comments