Deep research you can inspect, resume, and trust.
Evident turns complex questions into conclusion-first reports backed by traceable web, academic, and private evidence.
Product · 8-Agent Workflow · Architecture · Run Locally · Quality
Evident is a general-purpose, multi-tenant research platform for technical, business, policy, market, and internal-knowledge questions. It combines an eight-agent LangGraph workflow with durable execution, source-level evidence inspection, private RAG, and explicit model selection.
The goal is not to produce the most confident-sounding answer. It is to make the path from question to conclusion visible and auditable.
The interface follows a deliberate information hierarchy:
- Read the conclusion first. The report remains the primary surface.
- Check research quality next. Citation coverage, source mix, conflicts, and review status are visible without reading internal traces.
- Open evidence when needed. The evidence panel stays out of the reading path until a citation or source deserves inspection.
| Product promise | What the user gets |
|---|---|
| Evidence over assertions | Inline citations resolve to the exact scored source pool used by the workflow. |
| Private knowledge, safely scoped | Uploaded TXT, Markdown, and PDF documents are isolated by tenant and can be selected per research run. |
| Provider control | Choose Claude or local Qwen per request; the choice is persisted with the run for auditability. |
| Durable work | Queued jobs, agent steps, checkpoints, leases, reports, sources, and audit events survive process restarts. |
| Honest uncertainty | Conflicts, unsupported claims, evidence gaps, high-risk domains, and citation-revision requirements are first-class states. |
| Inspectable cost and quality | The evaluation harness records quality, latency, token usage, and explicitly configured provider cost without inventing missing data. |
The primary demo was recorded against the AWS staging deployment, using Claude and the full research workflow for: “Compare HTTP/2 and HTTP/3 using current technical sources.”
Watch the full research recording.
The Private Knowledge demo exercises the path public-search demos usually skip: upload a document, embed and index it through the real Bedrock Titan V2 and Milvus pipeline, scope a run to that document, and verify that the final report cites the private source instead of confusing it with unrelated public results.
Watch the Private Knowledge recording.
The staging URL is d5llhn72jopyp.cloudfront.net when the stack is active. Staging is intentionally demo-on-demand, so the URL may be offline between review sessions to avoid idle cloud cost.
Simple questions take a fast direct-answer path. Questions requiring current, comparative, multi-source, or private evidence enter the canonical eight-agent workflow.
1. Intent Router
├── direct → Direct Answer
└── deep research
↓
2. Planner
↓
3. Web Scout ─────┐
4. Local Scout ────┴─ run in parallel
↓
5. Evidence Judge
↓
6. Analyst
↓
7. Reflect ── evidence gap + budget → targeted scout round
↓
8. Writer ─── citation validation and bounded repair → Report
| Agent | Responsibility |
|---|---|
| Intent Router | Chooses direct answer or deep research and flags high-risk domains. |
| Planner | Decomposes the request into research tasks and a domain-appropriate report outline. |
| Web Scout | Retrieves public web, academic, and optional MCP evidence. |
| Local Scout | Searches only the tenant's permitted private-document scope. |
| Evidence Judge | Normalizes, scores, deduplicates, and identifies gaps or conflicts. |
| Analyst | Converts evidence into structured findings tied to canonical source IDs. |
| Reflect | Applies the quality gate and can request one bounded supplementary round. |
| Writer | Produces a conclusion-first report and repairs invalid citation references within a fixed budget. |
All agents communicate through ResearchState; durable jobs additionally
persist node-level checkpoints and agent-step traces. See
docs/workflow.md for routing, state fields, and recovery
semantics.
The console is designed for failure and ambiguity, not only the happy path.
| State | Product behavior |
|---|---|
| Redis unavailable | Result-cache and progress publishing degrade without losing durable PostgreSQL work. Idempotency and rate limiting fail closed with 503 when their guarantees cannot be enforced. |
| SSE disconnected | The UI announces the lost live connection; the durable job may continue and its state can be rehydrated from the API. |
| Job failed | The run keeps an explicit terminal state, bounded error detail, checkpoints, and audit history. |
| Report unavailable | The user sees a deliberate unavailable state rather than an empty report surface. |
| Citation revision required | The report remains visible with a first-class quality warning instead of being presented as verified. |
| High-risk domain | Medical, legal, financial, and safety-critical research requires human review independently of citation coverage. |
| Cancellation requested | Only queued or running jobs transition to cancelled; late cancellation cannot rewrite a completed result. |
This separation is intentional: PostgreSQL is the source of truth, while Redis is used only where its speed or coordination semantics add value.
flowchart TB
researcher["Researcher"] --> console["Vue 3 + Vite console"]
console -->|"REST + SSE"| api["FastAPI application"]
api --> sessionBoundary["Session auth + tenant boundary"]
api --> jobManager["Durable job manager"]
jobManager --> researchWorkflow["LangGraph 8-agent workflow"]
researchWorkflow --> llmProviders["Claude / Qwen"]
researchWorkflow --> publicSources["Tavily + Semantic Scholar"]
researchWorkflow --> mcpTools["MCP tools"]
researchWorkflow --> privateRag["Private RAG"]
privateRag --> embeddingService["Ollama / Bedrock embeddings"]
embeddingService --> vectorDatabase["Milvus / Zilliz Cloud"]
api <--> postgresState["PostgreSQL<br/>runs, reports, sources, checkpoints, audit"]
api <--> redisState["Redis / Valkey<br/>cache, idempotency, locks, rate limits, progress"]
api --> objectStorage["Local storage / private S3<br/>documents and report exports"]
| Boundary | Implementations |
|---|---|
| LLM | Claude through Anthropic; Qwen through Ollama |
| Search | Tavily web search; Semantic Scholar academic search; optional MCP federation |
| Embeddings | Qwen embeddings through Ollama; Amazon Titan Text Embeddings V2 through Bedrock |
| Vector store | In-memory test store; Milvus / Zilliz Cloud |
| Object storage | Private local filesystem; private Amazon S3 |
| Durable state | PostgreSQL with SQLAlchemy, Alembic, and LangGraph's PostgreSQL checkpointer |
| Coordination | Redis locally; Amazon ElastiCache for Valkey in AWS staging |
For the detailed component and request-flow diagrams, see docs/architecture.md. For the schema and ERD, see docs/data-model.md.
- Authenticated tenant registration, login, profile, password change, and session logout.
- Synchronous research and durable asynchronous jobs with polling and SSE.
- Run history, progress, cancellation, report retrieval, scored sources, and Markdown/PDF export with numbered or footnote citations.
- Private Knowledge upload, list, detail, retry, selection, and deletion.
- Provider capability and readiness endpoints.
- MCP Streamable HTTP server for web search, private retrieval, source lookup, document ingestion, report persistence, history, and human-review requests.
Every research endpoint derives tenant and user identity from an authenticated session. Client-supplied identity headers are not trusted. See docs/security.md for the implemented boundary and the remaining account-management limitations.
The shortest path starts the complete local stack:
cp .env.example .env
docker compose up --build --detach --waitOpen http://localhost:3000. The Compose topology and smoke test are documented in docs/deployment.md.
python3.13 -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[dev]"
cp .env.example .env
uvicorn app.main:app --reloadIn a second terminal:
cd frontend
npm install
npm run devPDF rendering uses WeasyPrint and native Pango libraries. On macOS:
brew install weasyprint
export DYLD_FALLBACK_LIBRARY_PATH="$(brew --prefix)/lib${DYLD_FALLBACK_LIBRARY_PATH:+:$DYLD_FALLBACK_LIBRARY_PATH}"
python -m weasyprint --infocurl -c cookies.txt -X POST http://127.0.0.1:8000/auth/register \
-H 'Content-Type: application/json' \
-d '{"email":"researcher@example.com","password":"correct-horse-battery","tenant_name":"Example Research"}'
curl -b cookies.txt -X POST http://127.0.0.1:8000/research-runs \
-H 'Content-Type: application/json' \
-H 'Idempotency-Key: first-research-request' \
-d '{"query":"Compare PostgreSQL logical replication and change data capture.","llm_provider":"qwen"}'Use POST /research-runs/jobs for durable asynchronous execution. The response
returns the run, progress, events, and report URLs.
Default tests are deterministic and do not call paid or external services. Live integrations are explicit and opt-in.
ruff check .
mypy app tests scripts alembic/env.py alembic/versions
pytest -qThe Docker test target includes the native Linux PDF dependencies used by the application image and CI:
docker build --target test --tag evident-backend-test .
docker run --rm evident-backend-testcd frontend
npm run lint
npm run format:check
npm run typecheck
npm test
npm run build
npx playwright testThe verification strategy covers unit tests, API contracts, PostgreSQL and Redis integration, live provider smoke tests, reversible migrations, restart recovery, Playwright user journeys, Terraform validation, container packaging, and reproducible research evaluation. Current verified counts and known gaps live in docs/status.md; evaluation fixtures and published runs live in docs/evaluation.md.
Terraform defines the AWS staging stack: CloudFront, an Application Load Balancer, ECS Fargate, RDS PostgreSQL, ElastiCache for Valkey, private S3, Bedrock access, ECR, and GitHub OIDC deployment identity.
The stack has been applied and verified with a real public research run and a separate private-document Bedrock/Milvus round trip. It is operated on demand because the running resources incur cost.
# Deploy after configuring the Terraform backend and required secrets.
TF_VAR_budget_notification_email=you@example.com \
TF_VAR_anthropic_model=replace-with-supported-model-id \
TF_VAR_milvus_uri=https://replace-with-managed-milvus-endpoint \
ANTHROPIC_API_KEY=... TAVILY_API_KEY=... MILVUS_TOKEN=... \
scripts/aws-deploy.sh
# Destroy billable staging resources; the protected state bucket remains.
AWS_DESTROY_CONFIRM=destroy-staging scripts/aws-destroy.shReview the cost basis, networking, secret flow, and deployment lifecycle in docs/deployment.md before applying.
| Document | Purpose |
|---|---|
| Project charter | Product vision, scope, phases, and final acceptance criteria |
| Architecture | Components, provider boundaries, and request flows |
| Research workflow | Eight agents, routing, shared state, and durable execution |
| Data model | PostgreSQL schema and ERD |
| Security | Authentication, tenant isolation, prompt-injection handling, and secrets |
| Reliability | Retry, circuit breaking, checkpoint recovery, and failure isolation |
| Evaluation | Metrics, fixtures, published runs, limitations, and cost methodology |
| Trade-offs | Architectural decisions and rejected alternatives |
| Deployment | Local topology, CI/CD, AWS staging, and cost controls |
| Status | Evidence-backed implementation and verification log |
- Put the conclusion first, research quality second, and evidence on demand.
- Make provider choice, private scope, and human-review state explicit.
- Treat tenant isolation, idempotency, and durable ownership as correctness boundaries rather than UI details.
- Keep paid integrations opt-in and record only reproducible metrics.
- Keep provider SDKs behind application interfaces so the product is not tied to one model, search engine, vector store, or cloud runtime.
- Never present planned, mocked, local-only, or unavailable behavior as a verified production capability.






