This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Muckrake is the reusable FollowTheMoney data-pipeline core, published to PyPI. Application-specific code — the OpenLobbying crawlers, FastAPI app, Svelte frontend, FtM schema extensions, and deployment assets — lives in the sibling ../openlobbying/ repo, which depends on this repo as an editable path dependency. Nothing OpenLobbying-specific belongs here.
Hard constraint 🔒: muckrake must stay standalone and project-agnostic (it is the "data plane" in the muckrake × UTI merge — see ../docs/projects/undertheinfluence/merge-plan.md). It must never import, reference, or know about Django, Popolo, Wagtail, or UTI; projections to consumer schemas live in consumer repos. Litmus test for every change: could a brand-new project use muckrake without pulling in any UTI/Django code?
Note: the org-id dependency needs a v0.1.1 PyPI release to pick up the Registry.default() cache fix (org-id#2, closes muckrake#21); until then rebuilt images keep the slow path. Work in this repo is tracked on GitHub project boards 2 ("To-do", this repo's issues) and 3 ("muckrake × UTI merge", issues in openlobbying/docs). Merge-driven workstreams that land here:
- Containerisation (docs#30, ✅ done — PR #20): multi-stage Dockerfile + a standalone compose (Postgres with app + published DBs, CLI runner), with the db host port offset to 5433 to avoid the UTI stack; no API service (the server moved to openlobbying). The standalone compose must have no knowledge of UTI (🔒).
- Tooling baseline + tests (docs#33, #34): adopt ruff/mypy/pre-commit (currently none configured) and build out test coverage across core — existing tests only cover CLI entities, entity writes, and SQLite storage; load, release, artifacts,
make_id, NER apply, and dedupe have essentially none. Hardening work should land with tests. - Data-plane hardening workstream (docs#23–#29), prerequisite for UTI-scale ingestion (155k actors); fixes belong in core so all crawlers inherit them:
- ✅ resilient fetch layer (PR #22): shared polite session in
src/muckrake/extract/fetch.py— retry + exponential backoff (Retry-After honoured, idempotent methods only) + opt-in per-host rate limit (MUCKRAKE_HTTP_MIN_INTERVAL) + timeout - crawl checkpointing + keep partial artifacts (
src/muckrake/crawl.pycurrently discards everything on failure) - stream spreadsheet/CSV rows to the statement writer instead of materialising whole sheets
- remove silent pagination caps; emit completeness warnings into the run summary
- batched/incremental xref (hard
--limit 50000today) + document complexity at 150k+ entities - atomic load in
src/muckrake/load.py(currently delete-all-then-insert; a midway failure leaves the dataset empty) - opt-in concurrency for I/O-bound fetches
- ✅ resilient fetch layer (PR #22): shared polite session in
- Data backfill ports (later): Companies House enrichment as a
nomenklatura.enrichenricher feeding the resolver (note: UTI holds no human CH match decisions to migrate — its 51k figure is orgs processed, mostly unreviewed; the enricher matches from scratch), and APPC-archive / ParlParse historical crawlers — generic GB data-plane capabilities, not UTI glue. - A generic, project-agnostic export interface (docs#6) is an open decision that gates the merge's projection work — don't preempt its shape in code.
Always run Python via uv (never python/python3/pip directly).
uv sync— install deps.uv run muckrake --help— full CLI reference.uv run pytest— run tests. Single test:uv run pytest tests/test_foo.py::test_bar.docker compose up -d db— optional local Postgres 18 on host port 5433 (offset so the undertheinfluence umbrella stack can hold 5432/8000); an init script creates themuckrake+muckrake_publisheddatabases. WithoutMUCKRAKE_DATABASE_URL, muckrake defaults to SQLite atdata/muckrake.db(no setup needed).docker compose run --rm muckrake <command>— the containerised CLI (image installs muckrake non-editable; the repo is bind-mounted at/work, the runner's working directory, so datasets/data/.env come from the host). The image is also consumed by the undertheinfluence umbrella compose, which builds it from this checkout and points the runner's working directory at the openlobbying checkout instead.
Dataset configs are discovered from ./datasets/ in the current working directory plus MUCKRAKE_DATASET_PATHS. This repo ships no datasets — to run pipeline commands (crawl, load, xref, dedupe, ner-extract, release-build/release-publish, …) against real data, run them from ../openlobbying/. The API server is also there (uv run openlobbying server); there is no muckrake server command.
Pipeline stages, in order, with on-disk/DB handoffs between them:
- Crawl (
src/muckrake/crawl.py,dataset.py) — runs a dataset'scrawl(dataset)function, records adataset_runsrow, writes an immutable artifact underMUCKRAKE_ARTIFACT_PATH(defaultdata/artifacts/), and mirrors the latest success todata/datasets/<name>/statements.pack.csv. - NER extract/review (
src/muckrake/extract/ner/, see its README) — splits composite text fields into FtM fragment candidates (delimiterorllmextractor) stored inner_candidates; onlyapprovedcandidates are applied at load time. - Load (
src/muckrake/load.py) — readsstatements.pack.csv(or a--run-idartifact), applies approved NER candidates, materialises entities/relationships into the working DB. - Dedupe (
src/muckrake/dedupe/, see its README) —xrefgenerates resolver suggestions vianomenklatura;dedupeis the review TUI;dedupe-edgescollapses duplicateRepresentationedges. Resolver state lives in the DB. - Release (
src/muckrake/release.py) —release-buildsnapshots dataset runs into an immutable release;release-publishwrites it intoMUCKRAKE_PUBLISHED_DATABASE_URL(the read-only serving DB consumers query). Neverloaddirectly into the published DB.
The CLI (src/muckrake/cli.py) also exposes manual entity CRUD — add, get, search, update — for the "documenting investigations" use case where users (or agents) build a graph one entity at a time.
src/muckrake/__init__.py runs _configure_ftm_model_path() at import time: if extension schema dirs exist (from MUCKRAKE_FTM_SCHEMA_PATHS or ./ftm_schema_ext/ in the cwd), it overlays their YAMLs onto the upstream followthemoney schema in a tempdir and sets FTM_MODEL_PATH so all subsequent FtM imports see the merged model. Anything that uses FtM must import muckrake first — this is why tests/conftest.py exists; don't skip it in scripts or notebooks. Schema extensions belong in consuming app repos, not here.
MUCKRAKE_DATABASE_URL (working DB; SQLite default when unset), MUCKRAKE_PUBLISHED_DATABASE_URL (defaults to working DB URL — use a separate DB when testing releases), MUCKRAKE_DATA_PATH / MUCKRAKE_ARTIFACT_PATH / MUCKRAKE_DATASET_PATHS overrides, OPENROUTER_API_KEY + LLM_MODEL for LLM-based NER. .env is loaded from the repo root of the current working directory (MUCKRAKE_ENV_FILE overrides); see .env.example.
- Code style: simple and tidy. No speculative abstractions, no backwards-compat shims, no defensive handling of conditions that don't occur.
- Typed Python. Be conservative about adding dependencies.
- Prefer existing
followthemoney/nomenklaturafunctions before writing new ones — chances are they exist. FtM schema/property reference: https://followthemoney.tech/explorer/ - Crawler code must crash loudly on ambiguous data — never emit a guessed value.
- Keep README.md / AGENTS.md accurate when changing behaviour.
- Don't commit, push, or open PRs unless asked.
ghCLI is available for read-only GitHub interactions.