repowise indexes your codebase once, builds five foundational layers, then keeps them in sync on every commit. Four further layers are derived from those five. This document is the deep dive; the README gives the one-paragraph version of each and links here for the detail.
The layers compound. The graph locates what git flags, code health scores it, decisions explain why it is shaped that way, and the docs make all of it searchable in natural language.
Foundational
- Graph Intelligence
- Git Intelligence
- Documentation Intelligence
- Decision Intelligence
- Code Health Intelligence
Derived — built on the five, each with its own reference page
Cross-cutting
Everything below is computed without model calls. An LLM is an optional upgrade for prose quality in the docs layer, never a requirement for the index.
tree-sitter parses your source into a two-tier dependency graph: file nodes
and symbol nodes (functions, classes, methods). 18 languages parse to a full
AST; see LANGUAGE_SUPPORT.md for per-language tiers.
A 3-tier call resolver with confidence scoring handles import aliases,
barrel re-exports and namespace imports: same-file resolution at 0.95,
import-scoped at 0.90, globally-unique at 0.50, with language-specific pre-tiers
for Go packages, JVM same-package and C++ same-target. Heritage extraction
covers extends, implements and trait impls.
- Leiden community detection (Louvain fallback) finds logical modules even when your directory structure doesn't reflect them.
- PageRank, betweenness centrality, SCC analysis, and execution-flow tracing from entry points identify your most central, most coupled, and most traversed code.
- Framework-aware edges connect routes to handlers across 22 framework detectors, including Django, FastAPI, Flask, ASP.NET, Spring, Micronaut, Quarkus, Jakarta, Express/NestJS, Next.js App Router, Remix, tRPC, Hono, Gin/Echo/Chi, Axum/Actix/Rocket, Rails, Laravel, TYPO3, Flutter and the pytest/gtest test runners.
Every derived metric is defined in
COMPUTED_GLOSSARY.md.
repowise mines your git history to produce signals no static analysis can find.
Hotspots: files in the top quartile of decayed churn (exponential
half-life, so last month outweighs last year) that also clear minimum-activity
floors — at least 3 commits in 90 days and a meaningful temporal score. Churn
alone would flag every file a bulk refactor touched; the floors keep the list
about sustained activity. Surfaced by get_risk() before your agent edits.
Ownership: git blame aggregated into per-author percentages. Know who to
ping, and where knowledge silos are.
Co-change pairs: files that change together in the same commit without an
import link. Hidden coupling AST parsing cannot detect. get_context() surfaces
co-change partners alongside direct dependencies.
Bus factor: how many authors it takes to cover 80% of a file's commits. A
bus factor of 1 is the classic single-owner risk, surfaced in CLAUDE.md.
Significant commits: up to 50 meaningful commit messages per file — merges, dependency bumps, lint-only and bot commits filtered out — with the commit body retained when it carries decision intent. These feed generation prompts, so the wiki can explain why code is structured the way it is.
Contributor profiles: every author gets a page — modules they own, top files, co-authors, commit category mix, silo modules, bus-factor risk files, and dead-code burden.
Module health: a 0–100 composite per top-level module from silo penalty, hotspot density, dead-code percentage, average churn, doc coverage and median bus factor.
Reviewer suggestions: paste a PR file list into Blast Radius for a ranked reviewer list, scored by direct authorship (×1.0), co-change partners (×0.5) and recency (×0.4), capped at the 5 strongest co-change signals per file.
A wiki for every module and file, rebuilt incrementally on every commit. Deterministic templates render it with no model calls; supplying an LLM key upgrades the prose rather than enabling the layer.
- Coverage tracking: what's documented and what isn't.
- Freshness scoring per page, relative to the underlying code.
- Semantic search via RAG: hybrid retrieval merging full-text and vector results through Reciprocal Rank Fusion, with PageRank bias and a 1–2 hop graph expansion that walks the imports and projected-calls graph for flow-shaped questions.
A typical single-commit update regenerates only the handful of pages your change actually touched.
Architectural decisions mined at index time from five sources: ADR files
(Nygard/MADR), PR and squash-commit bodies, inline markers, git archaeology, and
centrality-bounded code comments. Two further capture paths exist for humans and
agents: repowise decision add, and mining your Claude Code or Codex session
transcripts.
# WHY: JWT chosen over sessions — API must be stateless for k8s horizontal scaling
# DECISION: All external API calls wrapped in CircuitBreaker after payment provider outages
# TRADEOFF: Accepted eventual consistency in preferences for write throughputEvery decision is evidence-backed: each rationale traces to a verbatim source span, and an anti-hallucination substring gate stamps it exact, fuzzy or unverified. Corroborating sources raise confidence rather than overwrite each other.
On lineage, precisely. Decisions have a typed-edge schema
(supersedes / refines / relates_to / conflicts_with), but no edges are
written today. The only detector that produced them scoped conflicts by
similarity, which does not scope, so it is disabled and the edges it wrote were
removed. Lineage stays empty until a structural detector replaces it. What does
work: the diff-driven pass on repowise update marks decisions a new commit
reversed, and repowise decision deprecate --superseded-by records the successor
on the record itself. See DECISIONS.md.
These records surface everywhere your agent already looks: get_why() for the
archaeology, governing decisions in get_context(), a governance_risk flag in
get_risk() PR review, a Key Decisions section in get_overview(), and the
ungoverned_hotspot / stale_governance / contradictory_decision findings in
the code-health layer.
repowise decision add # guided interactive capture
repowise decision confirm # review auto-proposed decisions
repowise decision health # stale, conflicting, ungoverned hotspotsThe "why" usually walks out the door: when a teammate leaves, or when you reopen your own repo six months later. This keeps it in the codebase.
repowise scores every file 1–10 on three co-equal signals — defect risk, maintainability, and performance risk — from a roster of 49 deterministic detectors, of which only 26 are permitted to move the defect number. Pure static analysis over tree-sitter and git data, budgeted (and CI-tested) to finish in under 30 seconds on a 3,000-file repo.
The defect weights are calibrated offline against a real bug corpus, not hand-tuned: every file scored at a commit preceding the bug window (no leakage), an L2-logistic fit with NLOC as an explicit control so a marker only earns weight for defect lift beyond file size. Only the learned constants ship; the runtime stays fully deterministic.
Validated leakage-free across 21 repositories, 9 languages, 2,826 files at mean ROC AUC 0.737, and 0.76–0.78 on the public PROMISE/jEdit dataset that played no part in calibration.
It does not stop at scoring — the layer closes the loop into concrete, graph-aware refactoring plans an agent can execute: Extract Class, Extract Method, Extract Helper, Move Method, Break Cycle and Split File, each with its plan, recovered impact and blast radius.
repowise health # KPIs + lowest-scoring files
repowise health --refactoring-targets # ranked by impact / effort
repowise health --trend # snapshots + declining-health alertsFull guide, the calibration story and the head-to-head against CodeScene:
CODE_HEALTH.md.
A calibrated 0–10 defect-risk score for a whole commit or base..head
range, computed from the diff's shape against the live checkout — no index
lookup, no model call. Distinct from get_risk(), which scores indexed files by
path. Lead with risk_percentile, which ranks the change against sampled recent
commits in the same repo.
repowise risk HEAD
repowise risk main..feature-branchReference: CHANGE_RISK.md.
Which tests actually exercise the code you changed, which changed files have no guarding test at all, and per-file coverage merged across every test that touches it. Coverage ingests from LCOV, Cobertura, Clover or normalized JSON and feeds the code-health coverage markers.
The practical payoff is tests_to_run in get_risk() PR mode: a
coverage-backed list rather than a guess, so an agent runs the tests that can
actually catch its change.
Reference: TEST_INTELLIGENCE.md.
Bug-fix commits attributed back to the files and functions they repaired, giving
each file a fix count, a last-fixed age, and a "bug magnet" flag. This is why the
generated CLAUDE.md orders its files that need care by bug-fix history
first, then churn — a file that keeps breaking is a better warning than a file
that merely changes often.
Reference: BUG_HISTORY.md.
A static security scan over the working tree and the full history, so a
finding carries the commit that introduced it and its author, not just a line
number. Findings are idempotent across re-scans and surface through
repowise security, the REST API and the dashboard.
Pure graph traversal and SQL. No model calls.
repowise dead-code
23 findings · 4 safe to delete
✓ utils/legacy_parser.ts file 1.00 safe to delete
✓ auth/session.ts file 0.92 safe to delete
✓ helpers/formatDate export 0.71 safe to delete
✗ analytics/v1/tracker.ts file 0.41 recent activity — review first
Conservative by design. safe_to_delete requires confidence ≥ 0.70 and excludes
17 dynamically-loaded naming patterns (*Plugin, *Handler, *Middleware,
register_*, on_*, *_route, *_callback and more). Dynamic-import
detection and a per-language framework-convention registry further cut false
positives. repowise surfaces candidates; engineers decide.
Reference: DEAD_CODE.md.
Most MCP tools are passive — the agent has to know to call them. repowise hooks
are active: they act on what the agent is already doing. No LLM calls, no
network, pure local SQLite queries. Installed automatically during
repowise init.
There is deliberately no unconditional pre-search injection. An earlier
version enriched every Grep/Glob before it ran and was removed: it added
noise on the majority of searches where the agent had already found what it
wanted. What ships now fires on evidence that the agent needs help:
- Zero-result rescue — a grep found nothing, so wiki full-text, fuzzy symbol and decision matches surface the closest real hit.
- Flood digest and triage — a search returning far too much is replaced with a compact per-file digest of match counts and anchor lines, plus the top files by PageRank.
- Wrong-path rescue — a failed
Read/Edit/Writepath is resolved against the index to the file you almost certainly meant. - Read skeleton — an unbounded read of a large indexed file is served as its skeleton with 1-indexed ranges, once per file per session.
- Stale-read notice — a file was edited after this session read it earlier.
- Decision injection — editing a file with a governing decision surfaces it in one line.
- Session start — index freshness, the trust protocol, and standing decisions.
- Post-commit staleness — after a successful commit, merge, rebase, cherry-pick or pull, a notice if the wiki has drifted from HEAD.
Claude Code and Codex are both supported.
Related capability: Distill reuses the index (symbol bounds, centrality, hotspots) to compress noisy command output and large reads before the agent sees them — built on the layers, not a layer.
| Method | Command | Best for |
|---|---|---|
| Post-commit hook | repowise hook install |
Set-and-forget local development |
| File watcher | repowise watch |
Active development without committing |
| GitHub webhook | Configure in repo settings | Teams, CI/CD |
| GitLab webhook | Configure in project settings | Teams, CI/CD |
| Polling fallback | Automatic with repowise serve |
Safety net for missed webhooks |
repowise hook install # post-commit hook (current repo)
repowise hook install --workspace # all workspace repos
repowise watch # or use the file watcherUpdates are incremental: only the pages your change actually touched are
regenerated. Full guide: AUTO_SYNC.md.
After every repowise init and repowise update, repowise regenerates your
CLAUDE.md from actual codebase intelligence, not a template. No LLM calls. An
AGENTS.md generator shares the same pipeline.
repowise generate-claude-mdThe generated section includes: index freshness and a _meta explainer, how to
work in this repo, a trust protocol stating when a served result may be used
without re-reading, the MCP tool table, an architecture summary, key modules,
entry points, files that need care (ordered by bug-fix history, then churn),
code health across all three signals, standing architectural decisions, and your
build/test commands. A user-owned section at the top is never touched.
<!-- REPOWISE:START — Do not edit below this line. Auto-generated by Repowise. -->
## Architecture
Monorepo with 4 packages. Entry points: api/server.ts, cli/index.ts.
## Files that need care (bug-fix history first, then churn)
- payments/processor.ts — 19 bug fixes, last fix 2 days ago (bug magnet); 23 commits/90d
## Standing decisions (ask get_why before diverging)
- JWT over sessions (auth/service.ts) — stateless required for k8s horizontal scaling
<!-- REPOWISE:END -->