Free, open-source autonomous research engine: auto research from a prompt to a source-grounded draft with verified citations.
19 specialized agents · CrossRef, OpenAlex, Semantic Scholar · PDF/DOCX/LaTeX export
This repository is the open-source engine, MIT-licensed and self-hostable.
| What it is | Open-source Python engine for AI-generated research drafts with verified citations |
| Best for | Literature reviews, research papers, thesis drafts, reproducible research workflows |
| Agents | 19 specialized AI agents (research, structure, writing, citation, polish, export) |
| Sources | CrossRef, OpenAlex, Semantic Scholar (200M+) |
| Languages | 57+ languages including English, Spanish, German, French, Chinese, Japanese |
| Export | PDF, Microsoft Word (.docx), LaTeX |
| Cost | Free and open source (MIT license); you bring your own model API keys. |
| Typical output | 5–80+ pages, 10k–20k+ words, 30–50+ citations (measured before multi-source confirmation was made the default) |
| Time to draft | 10–20 minutes |
| API cost per draft | ~$0.35 (Gemini Flash) to ~$3.00 (Claude Opus) |
- At a Glance
- What is OpenDraft?
- Why OpenDraft Exists
- OpenDraft for Open Source Maintainers
- What OpenDraft is NOT
- How It Works
- Citation verification
- Features
- Quick Start
- Which AI Model Should I Use?
- Example Output
- FAQ
- Tech Stack
- Contributing
- Links
OpenDraft is an open-source Python engine that generates source-grounded research drafts using 19 specialized AI agents. It is designed for academic researchers who need long-form documents (10,000–20,000+ words) with citations verified against real databases.
OpenDraft does not invent its citations. By default a source is only included once its DOI is held by at least two of CrossRef, OpenAlex and Semantic Scholar, and every citation records which databases confirmed it and which ones this engine re-queried itself. See Citation verification for exactly what that does and does not establish.
- OpenDraft is a command-line tool and Python library for drafting academic papers.
- Best for: Researchers drafting literature reviews, journal submissions, structured research papers, and thesis first drafts.
- License: 100% free and open source (MIT).
- Setup time: ~10 minutes for local installation.
OpenDraft is not just a drafting tool — it is a reproducible research-agent pipeline that open-source maintainers can extend, audit, and improve.
We use Codex and OpenAI models to maintain OpenDraft itself:
- Automated PR review — Codex reviews contributor changes for agent logic, prompt quality, and citation handling
- Regression test generation — AI-assisted tests for citation accuracy, source coverage, and draft coherence
- Issue triage — Codex suggests labels, duplicates, and fixes for bug reports
- Release workflow automation — Automated changelogs, version bumps, and eval runs before each release
- Contributor templates — Codex-assisted onboarding for adding new agents, validators, and export formats
See EVALUATION.md for the benchmark plan and CONTRIBUTING.md for maintainer guidelines.
We built OpenDraft after repeatedly encountering AI writing tools that produced confident-sounding research drafts with hallucinated or unverifiable citations.
Academic research requires trust, sources, and accountability.
OpenDraft explores a different approach: instead of a single general-purpose model, it uses multiple specialized agents, each responsible for a specific step in the research drafting process, grounded in real academic literature.
We open-sourced OpenDraft so researchers can inspect, critique, and improve how these systems actually work.
- Citation confidence — OpenDraft confirms each source's DOI against multiple scholarly databases and drops the ones it cannot confirm, rather than accepting model-asserted references.
- Long-form output — OpenDraft generates 20,000+ word research drafts.
- Academic structure — OpenDraft builds structured research outlines with proper chapter/section hierarchy.
- Export formats — OpenDraft exports to PDF, DOCX, and LaTeX with academic formatting.
- Open source — OpenDraft is fully open source under the MIT license and can be inspected and extended.
- Researchers preparing literature reviews, journal submissions, or structured first drafts.
- Open-source maintainers building tools on top of a reproducible research-drafting pipeline.
- Graduate students working on a master's thesis or PhD dissertation.
- Academics who want to verify that every citation in their AI-assisted draft links to a real paper.
- Developers extending the agent pipeline for custom research workflows, citation validators, and export formats.
OpenDraft is intentionally not designed for:
- One-click generation of final papers
- Cheating on assignments
- Inventing citations or bypassing peer review
- Replacing human researchers
It is a research assistance and drafting tool, not an autonomous author.
OpenDraft uses 19 specialized AI agents that work like a research team:
📚 RESEARCH PHASE → Finds candidate papers via CrossRef, OpenAlex, Semantic Scholar,
web search, then confirms each DOI in 2+ of
CrossRef/OpenAlex/Semantic Scholar and drops the rest
🏗️ STRUCTURE PHASE → Creates research outline with chapters
✍️ WRITING PHASE → Drafts each section with academic tone
🔍 CITATION PHASE → Dedupes, quality-filters, then checks each surviving
source is actually on-topic for the paper
✨ POLISH PHASE → Refines language and formatting
📄 EXPORT PHASE → Generates PDF, Word, or LaTeX
Result: A complete research draft in 10-20 minutes.
Citations are checked in two independent ways. They answer different questions and neither substitutes for the other.
Discovery may find a candidate through any source. Confirmation then looks the candidate's DOI up directly in each scholarly database and counts how many hold a record for it.
By default a citation is kept only if at least 2 of {CrossRef, OpenAlex,
Semantic Scholar} hold its DOI. A single-source result is dropped and the drop
is logged. Accepting single-source results is an explicit opt-out
(require_multi_source=False), not the default.
To be exact about what "2 databases hold it" means: one of the two may be the
database that returned the candidate in the first place, which is taken at its
word rather than re-queried. The others are looked up directly by DOI. The
engine tracks this distinction internally (confirming_sources versus
independently_confirmed_by) and verification_notes on each citation spells
out which database found it and which ones confirmed it.
| Setting | Default | Effect |
|---|---|---|
require_multi_source |
True |
Drop citations fewer than min_confirming_sources databases hold |
min_confirming_sources |
2 |
How many of the three must hold the DOI |
allow_unconfirmed_web_sources |
False |
Keep DOI-less web-search results (kept tagged if enabled) |
enable_llm_fallback |
False |
Let the LLM assert a citation when every lookup fails |
Every citation in bibliography.json carries its provenance:
verification_status |
Meaning |
|---|---|
multi_source_confirmed |
The DOI is held by min_confirming_sources or more databases, listed in verification_sources |
single_source |
Exactly one database holds the DOI. Dropped under the default settings |
unconfirmed |
The DOI carries no record in any scholarly database. Dropped under the default settings |
web_search_unconfirmed |
No DOI, so no scholarly database could be queried. Zero databases confirmed it |
llm_unverified |
Asserted by the LLM with no external lookup of any kind. Nothing checked that it exists |
not_checked |
Confirmation was disabled for this run |
verification_sources is written out even when it is empty, precisely so an
unconfirmed citation can never serialize to look like a confirmed one.
What this establishes, and what it does not. A confirmation means the DOI is registered and indexed in that many databases. It does not mean the work supports the sentence it is attached to, and it is not three separately sourced attestations of the same facts: OpenAlex and Semantic Scholar both ingest Crossref metadata, so the three are not fully independent of one another.
arXiv is not queried as a citation database. arxiv.org can appear as a
web-search result and Semantic Scholar exposes arXiv IDs, but there is no arXiv
API client in this engine.
A real, correctly cited, multi-source-confirmed paper can still be attached to a
claim it says nothing about. Existence checking cannot detect that, so
CitationClaimVerifier judges each source against the claim it is cited for and
returns RELEVANT, IRRELEVANT or UNCERTAIN.
In the citation phase this runs against the paper topic, because that phase
executes before any draft text exists and the topic is the only claim available
at that point. Sentence-level checking needs a draft and is available through
run_citation_claim_verification().
Reports are written to the research folder as
citation_claim_verification.md and .json. A citation judged IRRELEVANT is
removed only above a confidence floor (CLAIM_VERIFICATION_MIN_CONFIDENCE,
default 0.7), and the engine refuses to empty the bibliography outright.
| Env var | Default | Effect |
|---|---|---|
ENABLE_CLAIM_VERIFICATION |
true |
Run claim-level verification at all |
CLAIM_VERIFICATION_DROP_IRRELEVANT |
true |
Remove irrelevant citations rather than only reporting them |
CLAIM_VERIFICATION_MIN_CONFIDENCE |
0.7 |
Confidence needed before a removal happens |
These verdicts are language-model judgements, not proofs. The judge reads a
citation's title and abstract, not the paper's full text. UNCERTAIN means
unchecked, not passing. Treat the output as evidence for a human reviewer.
Requiring two independent confirmations necessarily lets fewer candidates through than accepting the first responder did. That is the intended trade: fewer citations, each one confirmed by more than one database.
A run can now fail where it previously produced a weak draft. Strict
confirmation, the strict quality filter and claim-level removal all shrink the
bibliography, and the pipeline raises PipelineValidationError if no citations
survive the citation phase. If you hit that, widen the search or relax the
settings deliberately rather than by accident.
Citation counts quoted elsewhere in this README and in EVALUATION.md were
measured before multi-source confirmation became the default and have not
been re-measured since. Treat them as historical. If you need the old
behaviour, set require_multi_source=False — and note that citations then
carry verification_status: not_checked rather than being labelled confirmed.
Neither check removes the need to read the draft. See What OpenDraft is NOT.
By default a citation is kept only if its DOI is held by at least two of CrossRef, OpenAlex and Semantic Scholar. A source only one database knows about is dropped, not quietly accepted. Every citation in bibliography.json carries the list of databases that confirmed it, so an unconfirmed source can never look like a confirmed one. See Citation verification.
- Research papers (5-15 pages)
- Literature reviews (20-40 pages)
- Thesis drafts (30-80 pages)
- Structured reports (10-100+ pages)
English, Spanish, German, French, Chinese, Japanese, Korean, Arabic, Portuguese, Italian, Dutch, Polish, Russian, and 40+ more.
- PDF - LaTeX-quality formatting
- Microsoft Word (.docx)
- LaTeX source - for journals
MIT license. Self-host with your own API keys.
OpenDraft includes two standalone tools for quickly understanding any research paper:
Generate a concise 5-bullet summary of any paper in seconds:
# As a subcommand
opendraft tldr paper.pdf
# Or standalone
opendraft-tldr paper.pdf
# Output to file
opendraft tldr paper.pdf -o summary.mdEach bullet follows academic structure: thesis, key finding, method, implication, limitation.
Generate a podcast-style audio summary you can listen to:
# Generate script + audio
opendraft digest paper.pdf
# Choose a different voice (rachel, adam, josh, elli, bella)
opendraft digest paper.pdf --voice adam
# Script only (no audio)
opendraft digest paper.pdf --no-audio
# Specify output directory
opendraft digest paper.pdf -o output/Requirements:
- Digest audio requires an ElevenLabs API key set as
ELEVENLABS_API_KEY - PDF reading requires the optional
pdfextra:pip install opendraft[pdf]
Both tools work with any academic paper (PDF, Markdown, or plain text), not just OpenDraft-generated documents.
Fetch research data from major statistical APIs directly into your workflow:
# Search for indicators
opendraft data search GDP
# Fetch World Bank data
opendraft data worldbank NY.GDP.MKTP.CD --countries USA;DEU --start 2020 --end 2023
# Fetch EU statistics (Eurostat)
opendraft data eurostat nama_10_gdp
# Fetch Our World in Data datasets
opendraft data owid covid-19Supported providers:
- World Bank - Development indicators (GDP, population, education, health)
- Eurostat - European Union statistics
- Our World in Data - Open research datasets
Data is saved as CSV files for use in your research.
Revise existing drafts with AI assistance:
# Revise a draft with natural language instructions
opendraft revise ./output "Make the introduction longer and add more context"
# The revised draft is saved as draft_v2.md (with PDF/DOCX exports)Features:
- Auto-detects draft files in output folders
- Preserves all citations during revision
- Automatic versioning (v2, v3, v4...)
- Quality scoring before/after
- PDF and DOCX export of revised version
Generate a quick research overview instead of a full draft:
opendraft "Neural Networks in Healthcare" --exposeThis produces a research expose with:
- Research Sources Overview - Number of sources, publication years, key journals
- Key Research Teams - Major authors and research groups in the field
- Structured Outline - Chapter/section structure for a full paper
- Complete Bibliography - All sources with DOIs and journal info
- Next Steps - Guidance for developing into a full draft
Use expose mode when you want to:
- Quickly scope a research topic
- Validate there's enough literature
- Get a structured starting point
- Review sources before committing to a full draft
Expose mode is ~3x faster than full draft generation.
Generate a 5-bullet summary of any academic paper in seconds:
# Summarize a PDF
opendraft tldr paper.pdf
# Summarize a markdown file
opendraft tldr draft.md
# Save to file
opendraft tldr paper.pdf --output summary.mdOutput:
📄 TL;DR: paper.pdf
• Main finding: Neural networks improve diagnostic accuracy by 23%
• Method: Retrospective analysis of 50,000 patient records
• Key limitation: Single-center study, needs external validation
• Implication: AI-assisted diagnosis could reduce misdiagnosis rates
• Future work: Multi-center trials planned for 2025
Works with any PDF, Markdown, or text file.
Generate a 60-second audio summary using ElevenLabs TTS:
# Generate audio digest (requires ElevenLabs API key)
opendraft digest paper.pdf
# Choose a voice
opendraft digest paper.pdf --voice adam
# Available voices: rachel (default), adam, josh, elli, bellaOutput: paper_digest.mp3 - a professional narration summarizing the key points.
Setup: Set ELEVENLABS_API_KEY in your environment or .env file.
- Python 3.10+
- A free Gemini API key
git clone https://github.com/federicodeponte/opendraft.git
cd opendraft
pip install -r requirements.txtCreate a .env file with your API key:
GOOGLE_API_KEY=your-gemini-api-keyfrom engine.draft_generator import DraftGenerator
generator = DraftGenerator()
draft = generator.generate(
topic="The Impact of AI on Academic Research",
paper_type="master", # research_paper, bachelor, master, phd
language="en"
)
# Export to different formats
draft.to_pdf("thesis.pdf")
draft.to_docx("thesis.docx")
draft.to_latex("thesis.tex")See engine/README.md for detailed API documentation.
| Model | Speed | Quality | Cost/Draft | Best For |
|---|---|---|---|---|
| Gemini 3 Flash | ⚡ Fast | Good | ~$0.35 | Most users |
| Gemini 3 Pro | Medium | Excellent | ~$1.40 | Important papers |
| GPT-5.2 | Medium | Excellent | ~$1.60 | OpenAI users |
| Claude Sonnet 4.5 | Medium | Excellent | ~$1.80 | Nuanced writing |
| Claude Opus 4.5 | Slow | Best | ~$3.00 | Maximum quality |
Recommendation: Start with Gemini 3 Flash for most use cases. Use Gemini 3 Pro or Claude Sonnet 4.5 for important papers.
See what OpenDraft produces:
Sample drafts with their verified bibliographies are in the examples/ directory.
Generated in ~15 minutes with verified citations from real academic papers.
opendraft/
├── engine/
│ ├── draft_generator.py # Main 19-agent pipeline
│ ├── config.py # Model & API settings
│ ├── prompts/ # Agent instruction templates
│ ├── utils/ # Citations, export, helpers
│ └── opendraft/ # Core agent modules
├── examples/ # Sample research outputs
├── requirements.txt # Python dependencies
└── README.md
Yes. OpenDraft is 100% free and open source under the MIT license. You self-host it with your own model API keys (a typical draft costs ~$0.35–$3 in API fees).
OpenDraft confirms each citation's DOI against at least two of CrossRef, OpenAlex and Semantic Scholar and drops the ones it cannot confirm. See Citation verification for exactly what that establishes.
OpenDraft can generate a complete first draft of a PhD dissertation (100+ pages) in 10–20 minutes. However, it is a drafting assistant, not an autonomous author. You must review, edit, and add your own analysis before submission.
Citations are not invented, and the engine records exactly how each one was established. By default a citation must have its DOI held by at least two of CrossRef, OpenAlex and Semantic Scholar; single-source results are dropped. One of the two may be the database that returned the candidate, which is taken at its word rather than re-queried; verification_independent_sources records the ones actually re-queried. The LLM-asserted fallback is off by default and, if you switch it on, everything it produces is permanently tagged llm_unverified. Note what this proves: that the cited work is registered and indexed, not that it supports the sentence it is attached to. A separate claim-level check covers that, and its verdicts are LLM judgements for a human reviewer. See Citation verification.
10–20 minutes for a full master's thesis (50–80 pages). A shorter research paper takes 5–10 minutes.
PDF, Microsoft Word (.docx), and LaTeX source.
Yes. The MIT license permits commercial use, modification, and distribution without restriction.
Yes. OpenDraft is 100% open source under the MIT license. Self-host with your own API keys. A typical research draft costs ~$0.35-$3 depending on the model.
OpenDraft confirms each citation's DOI in at least two of CrossRef, OpenAlex and Semantic Scholar before keeping it, and drops single-source or unconfirmed results.
OpenDraft generates research drafts—starting points you should review, edit, and build upon. Always:
- Verify all sources yourself
- Add your own analysis and insights
- Check your institution's AI policy
OpenDraft uses 19 specialized agents—one for research, one for citations, one for structure, one for export, etc.—instead of a single general-purpose prompt.
Yes. MIT license allows commercial use. Build products, offer services, modify the code—no restrictions.
- Engine: Python 3.10+, multi-agent orchestration
- Models: Google Gemini 3, Anthropic Claude Sonnet 4.5 / Opus 4.5, OpenAI GPT-5.5 / GPT-5
- Citations: CrossRef API, OpenAlex API, Semantic Scholar API
- Export: WeasyPrint (PDF), python-docx (Word)
Contributions welcome!
Ideas:
- Add new AI model support
- Improve citation accuracy
- Add export formats
- Translate prompts
Maintainer workflow docs:
- Push/auth runbook:
docs/MAINTAINER_PUSH_RUNBOOK.md - Automated push preflight:
scripts/push-preflight.sh
- 🌐 Hosted version: openpaper.dev
- 💬 Discussions: GitHub Discussions
- 🐛 Issues: Report Bug
- 🗒️ Changelog: CHANGELOG.md
- 📜 License: MIT
OpenDraft is a free, open-source Python engine for generating academic research drafts. It uses 19 specialized AI agents to create drafts whose citations are confirmed against real databases (CrossRef, OpenAlex, Semantic Scholar).
If OpenDraft helps your research, please star the repo!
⭐ Star on GitHub
