This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
- Run all tests:
pytest - Run a single test:
pytest tests/path/to/test_file.py::test_function_name -v - Format code:
ruff format - Lint code:
ruff check --fix - Type check:
mypy --exclude tests/test_package src tests
-
Formatting: Follow Google style convention. Use ruff for formatting
-
Imports: Use isort order (enforced by ruff)
-
Types: Strict typing is required. All functions must have type annotations
-
Typed returns: When a function returns multiple values, prefer a
NamedTuple(or small dataclass) over a bare tuple. Adjacent same-typed slots (andbool/intadjacency) make positional mistakes invisible to the type checker; named fields keep construction sites keyword-checked and give call sites self-documenting attribute access. -
Naming: Use snake_case for variables, functions, methods; PascalCase for classes
-
Docstrings: Google-style docstrings required for public APIs
-
Comments at call sites: Don't describe what a function does at the call site — the function's name and docstring already document that, and the comment will drift if the function evolves. Document rationale in the function's docstring instead. A call-site comment is appropriate only when the reason this caller specifically invokes it isn't obvious from surrounding context (eg. an unusual ordering constraint, a workaround for a known bug in this code path). When in doubt, write the docstring and leave the call site uncommented.
-
Error Handling: Use appropriate exception types; include context in error messages
-
Testing: Write tests with pytest; maintain high coverage. See "Testing Async Code" below for async test conventions.
-
Async Concurrency: Use
inspect_ai._util._async.tg_collect()instead ofasyncio.gather()for running concurrent async tasks. Useinspect_ai.util.collect()only inside sample subtasks (it adds transcript span grouping). -
File Paths: All code that handles file paths must support
s3://URLs,file://URIs, and plain local paths. Usefilesystem()frominspect_ai._util.filefor filesystem operations andlocal_path()to resolvefile://URIs to local paths before passing to APIs that only accept local paths (e.g.ZipFile). -
**Respect existing code patterns when modifying files. Run linting before committing changes.
All async test functions automatically run under both asyncio and trio backends via anyio (applied by the pytest_pycollect_makeitem hook in tests/conftest.py). Trio variants are skipped by default; use --runtrio to enable them.
- Do NOT use
@pytest.mark.asyncio— it conflicts with anyio and is blocked by conftest. Just writeasync def test_...and the hook handles the rest. - Use
anyio.sleep()notasyncio.sleep()in tests;anyio.Event()notasyncio.Event();tg_collect()notasyncio.gather(). - Use
@skip_if_trio(fromtest_helpers.utils) for tests that cannot run under trio (e.g. they test asyncio-specific fallback paths). @pytest.mark.anyiois not required but harmless — use it to signal intentional dual-backend coverage.
Additional files provide context when working in specific areas:
design/ contains architecture notes, subsystem internals, and documentation of repo/CI/development processes and workflows. Browse it before diving into an unfamiliar area.
Write the PR description using the template at .github/pull_request_template.md (fill in its sections — the "This PR contains" checklist, current vs. new behavior, breaking changes, other info).
When asked to open a PR, don't stop at creation — monitor it afterward: watch its CI checks (e.g. gh pr checks <number> --repo <owner>/<repo> --watch) until they complete, report the outcome, and investigate/fix any failures. If the branch has fallen behind its base (out of date), update it — merge or rebase the base branch in and push — so CI runs against current code.
For changes to product functionality (not test-only or build-only changes), add a CHANGELOG entry: a single-line, single-sentence item in the ## Unreleased section at the top of CHANGELOG.md (create that section if it doesn't exist), grouped with similar existing items when there are any, otherwise appended to the list. After updating a branch against its base, re-check that the entry is still under ## Unreleased — a merge can relocate it under a released heading; move it back if so.
Opening a PR from an organization fork to its upstream via the GitHub API/CLI needs an explicit head_repo — without it GitHub can't resolve the org fork as the PR head and rejects it with {"field":"head","code":"invalid"} (an org fork requires head_repo; a personal fork resolves without it). gh pr create has no --head-repo flag (cli/cli#6462), so use gh api — e.g. from the meridianlabs-ai/inspect_ai fork to upstream UKGovernmentBEIS/inspect_ai:
gh api repos/UKGovernmentBEIS/inspect_ai/pulls -X POST \
-f title="<title>" -f base="main" \
-f head="<branch>" -f head_repo="meridianlabs-ai/inspect_ai" \
-F body=@<body.md>Once the upstream PR is open it's the system of record: close the corresponding org-fork PR, with a close comment linking to the upstream PR.