|
| 1 | +# Where the company goes next |
| 2 | + |
| 3 | +**Our ambition: you change one business decision, and the whole virtual company |
| 4 | +responds coherently.** |
| 5 | + |
| 6 | +This roadmap proposes the next capabilities. They are not shipped features or |
| 7 | +dated commitments. The founder owns the goals, mandate, and final decisions; |
| 8 | +CEO and CTO lead as peers. The founder should see what each department changed, |
| 9 | +why, and which decision needs attention next. |
| 10 | + |
| 11 | +**Future scenario:** a founder changes service pricing. Finance revisits the |
| 12 | +assumptions, Sales revises the offer, Marketing checks its claims, Legal reviews |
| 13 | +the terms, and developers and designers update the relevant implementation. |
| 14 | +CEO and CTO bring the unresolved business and technical tradeoffs to the founder |
| 15 | +in one decision brief. The proposed result is consistent work across departments; |
| 16 | +it does not predict how the market will respond. |
| 17 | + |
| 18 | +## The foundation today: v1.5.1 |
| 19 | + |
| 20 | +- Eight departments, 48 employee manuals and six staff manuals: 54 unique skill |
| 21 | + manuals, with nine registered agents. These are operating instructions, not |
| 22 | + 54 continuously running workers. |
| 23 | +- Native Claude Code execution with local projects, dependencies, artifacts, and |
| 24 | + recorded reviews. See the [project workspace guide](docs/project-workspace.md). |
| 25 | +- Optional business harnesses with task policies, lifetime submission budgets, |
| 26 | + revision checks, and explicit extensions. See [harness loops](docs/harness-loops.md). |
| 27 | +- A website directory, brief preparation, and read-only project snapshots, plus |
| 28 | + optional focused [Mission Studio recipes](docs/mission-studio.md). |
| 29 | + |
| 30 | +Local contract tests do not establish real model delivery quality or security. |
| 31 | +Reviewer labels are declarations, not authenticated identities. NVIDIA and |
| 32 | +scanner references in the manuals are neither installed checks nor certifications. |
| 33 | + |
| 34 | +## The proposed sequence |
| 35 | + |
| 36 | +Each milestone must earn the next through the evidence below. |
| 37 | + |
| 38 | +| Milestone | Status | Founder-visible benefit | Required proof | |
| 39 | +|---|---|---|---| |
| 40 | +| A. First delivery | NEXT | A useful result with visible handoffs | Repeated pilots against baselines | |
| 41 | +| B. Shared decisions | PLANNED AFTER A | Change direction without rebuilding everything | Traced changes across departments | |
| 42 | +| C. Earned staffing | PLANNED AFTER A/B | Teams suited to the actual work | Held-out comparisons under equal budgets | |
| 43 | +| D. Company forks | EXPLORATION, DEPENDS ON A–C | Compare alternatives before committing | Isolated, reproducible scenarios | |
| 44 | + |
| 45 | +## A. The first real company delivery |
| 46 | + |
| 47 | +**Next:** instrument one local Claude Code pilot for a real, consented project or |
| 48 | +an explicitly synthetic one. Start with one concrete project involving at least |
| 49 | +three departments. Bring CEO business scope and CTO technical and security |
| 50 | +review into the same decision brief; activating all 54 manuals is not the goal. |
| 51 | + |
| 52 | +The pilot must demonstrate interruption and resumption without repeating an |
| 53 | +already completed external action. Require explicit user authorization before |
| 54 | +external publishing, spending, or new data transmissions, and enforce declared |
| 55 | +limits for each run. |
| 56 | + |
| 57 | +Publish a sanitized brief, artifacts, chronology, failures, founder |
| 58 | +interventions, and actual usage where the host reports it. Missing usage stays |
| 59 | +unknown. Record model aliases and the resolved runtime identity observed for each |
| 60 | +run without pinning the configuration. |
| 61 | + |
| 62 | +**Gate:** three runs per configuration (company, single agent, and fixed team), |
| 63 | +using the same disclosed case, inputs, and budgets, with business acceptance |
| 64 | +criteria defined beforehand and human review. This is an initial pilot-sized |
| 65 | +comparison. Include at least one case outside a software business before |
| 66 | +generalizing beyond software. |
| 67 | + |
| 68 | +## B. One decision, every affected team |
| 69 | + |
| 70 | +**Planned after A:** explicitly link business assumptions and decisions to tasks, |
| 71 | +artifacts, and their consumers. A changed decision would compute impact from |
| 72 | +those links, flag superseded evidence, propose only affected rework, and route |
| 73 | +unresolved CEO/CTO tradeoffs to the founder. Accepted tasks remain immutable |
| 74 | +history; revisions create superseding work with traceable links. |
| 75 | + |
| 76 | +An optional local live company view would require an authenticated companion |
| 77 | +with scoped permissions. The static Pages website continues to prepare briefs |
| 78 | +and display snapshots; native execution remains in Claude Code. |
| 79 | + |
| 80 | +**Gate:** demonstrate a price change across departments and a requirement change. |
| 81 | +Identify all affected artifacts through declared links, preserve unaffected work, |
| 82 | +review incomplete links, and measure human arbitration. Missing links must remain |
| 83 | +visible uncertainty; the company cannot assume it sees every consequence. |
| 84 | + |
| 85 | +## C. Teams that earn their place |
| 86 | + |
| 87 | +**Planned after A/B:** propose staffing and business harness updates from the |
| 88 | +goal, project phase, risk, and previous failures. Stage changes separately from |
| 89 | +the current frozen policy. Keep the full company available while activating |
| 90 | +only useful skills. |
| 91 | + |
| 92 | +CTO skill admission would require pinned-source attribution and licensing, |
| 93 | +declared permissions, security scan coverage and unknowns, sandbox evaluation |
| 94 | +with and without the skill, a shadow trial, and explicit promotion and rollback. |
| 95 | +A candidate cannot approve itself or loosen a currently failed gate to pass. |
| 96 | +NVIDIA SkillSpector and SkillEvaluator are candidate integrations and references, |
| 97 | +not mandatory vendors or installed scanners. |
| 98 | + |
| 99 | +**Gate:** compare candidate and current teams or policies on versioned, held-out |
| 100 | +business cases under the same disclosed budgets. Require quality and safety |
| 101 | +thresholds declared in advance, report tradeoffs, and retain the current setup |
| 102 | +without a demonstrated gain. Lower usage is a possible finding, not a promise. |
| 103 | + |
| 104 | +## D. Fork the company. Compare the options. |
| 105 | + |
| 106 | +**Exploration, dependent on A–C:** branch a local business scenario from a |
| 107 | +checkpoint, such as subscription versus service fee. Compare assumptions, |
| 108 | +generated artifacts, work, risks, and observed usage; the founder selects a |
| 109 | +reconciled proposal. A scenario is not a forecast. |
| 110 | + |
| 111 | +Replay would read evidence or run an isolated simulation, with external writes |
| 112 | +off by default and no duplicated real actions. A portable company recipe and |
| 113 | +sanitized evidence/replay pack would let others reproduce the case. Import must |
| 114 | +allow permission and reference previews without implicit execution. |
| 115 | + |
| 116 | +Wider runtime support, bounded recurring operation, and portfolios of companies |
| 117 | +would follow only after equivalent capability, permission, budget, isolation, |
| 118 | +and resume conformance. |
| 119 | + |
| 120 | +## What earns a release |
| 121 | + |
| 122 | +The headline measure is the proportion of projects reaching a reviewed, useful |
| 123 | +result without the founder coordinating every handoff. Track interventions, |
| 124 | +reversions, stale work caught, elapsed time, actual host usage, and uncertainty. |
| 125 | +Run evidence should be opt-in, redacted, and reproducible. Stars come after people |
| 126 | +completing projects, returning, and sharing useful real cases. |
| 127 | + |
| 128 | +## Help prove the first milestone |
| 129 | + |
| 130 | +Three bounded starting contributions for gate A: |
| 131 | + |
| 132 | +- Define one baseline fixture and a human review rubric shared across all runs. |
| 133 | +- Build an evidence recorder with redaction and explicit unknown usage fields. |
| 134 | +- Add an interrupted-run recovery test that detects repeated external actions. |
| 135 | + |
| 136 | +Use the [contribution guide](CONTRIBUTING.md) for repository conventions. |
| 137 | + |
| 138 | +## Design references |
| 139 | + |
| 140 | +Primary sources checked 2026-09-08: |
| 141 | + |
| 142 | +- [Paperclip](https://github.com/paperclipai/paperclip) documents agent organizations, goals, and budgets. |
| 143 | +- [MetaGPT](https://github.com/FoundationAgents/MetaGPT) models software-company operating procedures. |
| 144 | +- [LangGraph persistence](https://docs.langchain.com/oss/python/langgraph/persistence) supports state checkpoints. |
| 145 | +- [Claude Code agent teams](https://code.claude.com/docs/en/agent-teams) document coordination and current limitations. |
| 146 | +- [NVIDIA SkillEvaluator](https://docs.nvidia.com/skills/skillevaluator) separates deterministic, semantic, and live evaluation. |
| 147 | + |
| 148 | +Our differentiating hypothesis is decision coherence across departments with |
| 149 | +observable outcomes. Teams and memory alone are established building blocks. |
0 commit comments