Skip to content

Latest commit

 

History

History
53 lines (42 loc) · 2.95 KB

File metadata and controls

53 lines (42 loc) · 2.95 KB

zep-ingest examples

Every example is self-contained and re-runnable: it creates a fresh example-*-<timestamp> graph (or user), sets the starter ontology from example_ontology.py, ingests the bundled sample data under data/, and finishes with a search (or user-context fetch) so you see the graph pay off immediately.

pip install zep-ingest
export ZEP_API_KEY=...   # Zep dashboard → Project → API Keys
cd examples
python slack_export_example.py

Each run creates a new graph; delete old example-* graphs and users from the dashboard when you're done.

The examples

Script Demonstrates Destination
email_example.py .eml files → text episodes dated by their Date: headers; alias canonicalization named graph
documents_example.py Markdown → ~500-char chunks; optional LLM contextualization (auto-detects ANTHROPIC_API_KEY/OPENAI_API_KEY) named graph
fact_triples_example.py molding a realistic directory export (no triple-shaped columns) into explicit fact triples; the manual create → set_ontology → seed lifecycle named graph
json_records_example.py structured records with identity-field mapping — Zep extracts the relationships named graph
slack_export_example.py free preview() first, then a Slack export (one episode per message, thread as document_id), skip_subtypes, and the opt-in risky-alias guard named graph
user_graph_example.py combined: profile fact triples → chat-thread backfill → a document, all on one user's graph user graph
thread_backfill_example.py historic chat history (JSONL) → threads that power thread.get_user_context() user graph

example_ontology.py is the starter ontology every example applies with client.graph.set_ontology(...) before ingesting — copy it and adapt the types to your domain. The ingestion package never sets an ontology itself; that is a one-time graph setup step you own.

The sample data

Everything under data/ follows one scenario centered on Alder Ridge Robotics. The same people, products, and projects recur across emails, a handbook, a directory export, a catalog, a Slack export, and chat histories so the resulting graph contains useful cross-source relationships.

Thread-message files (chat_history.jsonl, combined_threads.jsonl) are one JSON object per line with columns matching ThreadMessage; a JSON array with the same columns also works:

{"thread_id": "support-1", "role": "user", "name": "Morgan Lee",
 "content": "...", "created_at": "2025-04-10T15:02:00Z"}

Every row is validated client-side before the first API call — role, RFC3339 created_at, metadata limits — so a bad line 500 fails fast, not mid-run.