Guided builder — brief field reference
This is a builder-focused quick reference: fill order, section cheat sheet, and the interview sprint loop. For the full product guide — concepts, workflow, context lenses, reading results, and multi-day programs — see the user guide. For consultant-led interviews and evidence elicitation prompts, use the consultant interview guide.
The builder turns what you know about a company into a structured model. You do not need every field — 52 operator patterns fire as you add the facts that trigger them. On Review, green in operator coverage = fired; grey (dormant) ones tell you what to add next. On Evidence, the evidence preview gives a compact readout (active operators, top gaps, leading candidate) without the full board in your way. Dormant operators collapse by default — expand when you want the full unlock checklist.
For most companies, an honest model is a multi-person, multi-day discovery program, not a form you finish in one sitting. Work Source → Evidence → Review one step at a time: add facts → Analyze → close the top gaps → repeat. You get value on a partial model; dormant operators are a roadmap, not a failing grade.
Fast start
- Step 1 · Source — set Company context (SME lens, department lens,
data-maturity lens, peer lens) when they help; Load an example, Auto-map, or resume a saved model.
- Continue to Evidence — set Archetype and resource level in Company;
the builder reorders sections and adapts gap questions to your company stage.
- Use Build your evidence map (+ Offering, + Process, …) or load from auto-map
(`.txt`, `.md`, `.html`, `.pdf`, `.csv`, `.xlsx`, …), including annual and quarterly reports. New sources propose merge patches — review before apply. Auto-map is a starting draft only — do not stop there.
- Watch the evidence preview update as you edit (debounced Analyze, no LLM).
When ready, Analyze & Review board — see full operator coverage, ranked candidates, and Improve the model gaps. Add internal pains, data access, bottlenecks, and abstractions — sharper evidence produces sharper candidates. Use Map from evidence on high-priority gaps when facts live in CRM/ERP but shouldn't be uploaded raw. Use Generate full plan on Review when you want full use-case write-ups for the candidates you checked (after the highest-impact gaps are closed).
Auto-map: draft, not done
Auto-map (website crawl, pasted notes, or file upload) typically captures roughly half to two-thirds of a useful model from public-facing material. It rarely includes:
- Linked pains with severity and cost
- Accessible data flags on high-volume processes
- Bottleneck skills on real workflows
- Project abstractions — the reusable patterns behind defensible bets
Do not rely on auto-map alone. After import:
- Run Analyze immediately to see what fired and what is dormant.
- Work the top gaps in Improve the model — add internal facts, interview
owners, or use Map from evidence for privacy-preserving inputs.
- Re-Analyze after each enhancement session until high-severity gaps close.
Generic candidates on a thin auto-mapped graph are expected. Company-specific, defensible bets appear once the evidence map reflects how the business actually operates.
For `.csv` and `.xlsx` uploads, Auto-map can produce a first artifact before the full model is complete: deterministic file findings, a file-grounded AI use case, an Open live dashboard action, and downloadable dashboard/one-pager HTML. Use that when you want a week-one pilot that starts from the uploaded file alone.
Annual and quarterly reports can also produce a cited financial profile for review. Comparable asset-utilization, working-capital, and margin signals feed three dedicated operator patterns instead of being left as background reading.
Workstation capture: draft workflows, opt-in
After the first Analyze, Add more evidence → Workstation capture drafts `processes`, `assets`, and possible scarce `skills` from an employee's own app/window activity — opt-in, local-first, review-only. With ActivityWatch running locally, click Import from ActivityWatch for a one-click draft. For a hosted Foundry, run the tray agent (`python -m collector.capture_tray`) on each workstation; it pushes only a redacted summary, and your Workstation capture tab shows what has crossed the graduation gate (Ready to review) versus what is still accumulating evidence (Still watching). Same review-before-apply panel as Auto-map. Full setup, including denylists and the CLI path: `docs/WORKSTATION_CAPTURE.md` in the repository.
Multi-day discovery (what works today)
Use this loop across sessions and contributors:
| Step | What to do |
|---|---|
| Day 0 | Auto-map (website and/or file upload) + archetype + one department lens → first draft. Run Analyze — expect gaps, not a finished model. |
| Each session | Pick one department focus. Work only the top 3 gaps in Improve the model — add internal evidence auto-map missed. Use Map from evidence when you have metadata or redacted samples instead of full records. |
| Between sessions | Copy the YAML tab into a shared file (git, Drive, etc.). Reload it next time. |
| Before Generate | Close high-severity gaps: linked pains, accessible data on high-volume processes, project abstractions where they matter. |
Important: the model lives in the browser until you Save to server (Step 1 · Source → Saved models) or copy the YAML tab into a shared file. A refresh without saving loses in-progress work. When signed in, saved models are private to your account.
For the full multi-person workflow strategy and product roadmap, see the user guide and resources.
Recommended fill order
Work top to bottom, but prioritize facts that unlock the most operators. With an archetype set, the builder may reorder sections (e.g. segments first for greenfield startups).
| Order | Section | Why |
|---|---|---|
| 1 | Company | Name + risk appetite set the ranking tilt. |
| 2 | Offerings | Most external plays start here. Add components (`manual`, `rule_based`, `judgment`) and link segments. |
| 3 | Segments | Connect offerings via serves — unlocks Segment Expansion and Customer Success. |
| 4 | Assets | Accessible data is the main grounding signal. For tech assets, add a tech profile (EOL status, CVE backlog, test coverage, lock-in, incident frequency) to unlock software-stack operators. Mark `accessible` only if you truly have it today. |
| 5 | Processes | High volume + repetitive → Automate; scarce bottleneck skill → Augment; no data yet → Instrument First. Name workflows (finance close, supplier exceptions, hiring funnel, support escalations, …) and attach pains to fire domain orchestration operators. |
| 6 | Skills | Scarce skills power Augment, Training, and Knowledge Capture (with a talent constraint). |
| 7 | Projects | Write the abstraction (pattern stripped of industry words) to unlock Abstraction Transfer. |
| 8 | Pain & value interview → Pains | Attach pains to a process/offering — raises Impact so real problems rank above generic ideas. |
| 9 | Constraints | Not just penalties — they generate Compliance, Integration, Risk Guardrails, Knowledge Capture, and software-stack plays when paired with tech assets. |
Operator coverage (52 patterns)
After Analyze on Review, the operator coverage panel lists every pattern in the catalog:
| State | Meaning |
|---|---|
| Green / fired | Trigger matched — jump to candidates in the list |
| Grey / dormant | Not yet applicable — hover or expand for the unlock hint |
| Collapsed dormant | Dormant operators are grouped in a collapsible section so fired ones stay visible |
Categories include core internal plays (Automate, Augment, Forecasting, …), domain orchestration (financial close, supply chain, HR, healthcare, compliance, support, deal desk, marketing ops), physical & R&D (physical automation adoption, generative design exploration, simulation/digital twin, scientific discovery acceleration), analytics & revenue (self-serve analytics, dynamic pricing), financial-report signals (asset utilization, working capital, margin defense), software & stack (modernization, test harness, security, observability, autonomous remediation, vendor exit), constraint triggers, prerequisites, and external innovation operators. You do not need every box green — aim for the operators that match real leverage on this company.
Relationships (edges)
Links between sections are as important as the nodes themselves:
- Offering → serves → Segment — who buys what.
- Offering / Process / Project → uses → Asset — what data or tech grounds the work.
- Offering → reuses → Project — proven capability you productize.
- Process → bottlenecked_by → Skill — where expert judgment gates throughput.
- Process → uses → Data asset — required for Automate/Augment feasibility.
Tick the checkboxes in each section; the builder writes the graph edges for you.
Section cheat sheet
- Offerings — products/services you sell; delivery `bespoke` + data unlocks Sales Intelligence.
- Segments — customer groups; pair with offerings and data for retention and forecasting.
- Projects — past delivered work; the abstraction field is the reusable pattern.
- Processes — internal workflows; volume + repetitiveness drive automation; workflow names and pains unlock domain orchestration; data drives feasibility.
- Skills — people capabilities; `med`/`high` scarcity marks bottlenecks.
- Assets — data (records), tech (platforms), physical (equipment). For tech, fill tech profile fields when you know legacy, security, or reliability posture. Only tick accessible when usable now.
- Pains — where time/money leaks; always attach to a node when you can.
- Constraints — regulatory, budget, risk, talent, stack — each unlocks targeted operators (including software modernization and integration when stack is the bottleneck).
Hover any ⓘ icon in the form for field-level help.
Interview sprint (the loop that grounds the model)
Every Analyze returns an Interview Sprint panel — interviews are a first-class part of discovery, the mechanism that turns assumptions into grounded facts. It runs a repeatable loop:
`auto-map -> generate interview agenda -> conduct interviews -> extract/update evidence graph -> analyze -> fix gaps -> repeat`
- Agenda — ask first: the highest-leverage questions; **Copy agenda as
Markdown** to share before a call.
- Gap checklist: severity, why it matters, and the questions to ask. Click
the status pill to cycle Open → Captured → Done across a multi-day program.
- Open in builder: jump to the section that closes the gap.
- Conduct interview: open Map from evidence for that gap, capture
evidence, extract proposals, review, and apply.
Applying proposals re-Analyzes the model, marks the gap Captured, and shows progress. Repeat until high-severity gaps are gone, then Generate full plan. Analyze uses no LLM. Your model and sprint checklist autosave to this browser.
Use Open evidence map or Download evidence map at any point to create a self-contained HTML exhibit of reviewed graph facts, recorded source classes, cited financial metrics, open gaps, and the current ranked opportunities. It is viewable offline without the app; it is a proof-of-process export, not a complete review-history audit log.
High-severity gaps (no linked pain, a high-volume process with no accessible data, a project with no abstraction) hurt ranking the most — close those first. Candidate cards also show Evidence that would strengthen this plus an Add in builder shortcut, so you can improve a specific opportunity without hunting through the form.
Map from evidence (Private Evidence Discovery)
When facts live in internal systems, you may not want to upload raw CRM/ERP records. The Interview Sprint → Conduct interview → Map from evidence flow closes gaps from privacy-preserving inputs:
| Tier | What you share | Typical use |
|---|---|---|
| No-data guide | Local inspection + typed observations | First pass; nothing leaves the browser until Extract |
| Metadata only | Field names, objects, stages, report titles | Reveal processes/assets without record contents |
| Aggregate only | Counts, rates, medians, distributions | Surface pains and bottlenecks safely |
| Redacted sample | Anonymized snippets | Ground abstractions and edge cases |
Workflow: pick a gap → choose a tier → add observations → Extract → review proposals in Evidence review → apply only selected patches → re-Analyze. Raw evidence is not persisted; only human-approved facts merge into the graph.
For stricter boundaries, operators can run the LLM against a customer-controlled endpoint (`LLM_RUNTIME=vpc_local`) so extraction never crosses a public API.
Analyze vs Generate
| Action | LLM | Output |
|---|---|---|
| Analyze | Not used | Instant prescores (Impact · Feasibility · Moat · Quality), operator coverage, the Improve the model panel, and a grouped candidate list. |
| Generate full plan | Used | Above plus LLM-written use cases and business artifacts (`why_this_company`, `first_experiment`, …) for selected candidates, a narrative strategic bet with per-bet summaries, and a sequenced plan (quick wins + strategic bet). |
| Critique & rewrite | Used (with Generate) | Extra LLM pass that flags generic candidates before final ranking. |
| Opportunity dossiers | Used (with Generate) | Packages the selected quick wins + strategic bet into decision-ready one-pagers you can copy as Markdown. |
| Agent blueprints | Used (with Generate) | Turns the selected opportunities into parameterized agent specs + deployment-readiness kits (readiness gated by graph evidence), copyable as Markdown. |
| Agent scaffolds | Used (with Generate) | Turns the blueprints into pilot-ready scaffolds — runtime contract, eval suite, readiness gap workflow, and pilot package (auto-enables Agent blueprints), copyable as Markdown. |
| Pilot validations | Used (with Generate) | Validates each scaffold against mock connectors — validation manifest, mock connector pack, eval run plan, and a pilot readiness verdict (auto-enables Agent blueprints + scaffolds), copyable as Markdown. |
| Pilot executions | Used (with Generate) | Actually runs the mock eval suite for each ready validation against mock connector stubs — case results, unsafe-action check coverage, threshold result, and a mock pilot verdict (auto-enables Agent blueprints + scaffolds + Pilot validations), copyable as Markdown. |
| Human pilots | Used (with Generate) | Converts mock-eval-passed executions into a structured human pilot plan — protocol, evidence capture checklist, risk register, completion criteria, and a pilot readiness verdict (auto-enables all prior phases), copyable as Markdown. |
| Production deployments | Used (with Generate) | Converts ready human pilots into a production deployment readiness package — connector inventory, deployment package, operating model, production eval plan, and a handoff readiness verdict (auto-enables all prior phases), copyable as Markdown. |
Use the feasibility / moat slider (or Risk appetite) to tilt toward quick wins vs defensible bets.
After Analyze, check Include in Generate on the candidate cards you want to take forward. The default selection is the top five non-generic candidates; the selection bar gives you Top 5, All, and Clear shortcuts. Prerequisites for checked candidates are included automatically. Use the How far should Generate go? slider to decide whether you only need the roadmap or want to continue through blueprint, scaffold, validation, mock execution, human pilot, and production handoff.
What raises agent readiness
The Agent blueprints option attaches a deployment-readiness status to each selected opportunity, computed from your graph (not the LLM). The same facts that sharpen candidates also justify creating an agent — to move a blueprint from *needs discovery* toward *ready*, add:
- a linked pain on the workflow (grounds the impact case)
- an accessible data asset the agent would read at decision time
- specific, non-generic grounding (a named process/offering, not a category)
- a satisfied prerequisite (e.g. Instrument First done before Automate)
- and, in interviews, the process owner, approval boundary (what needs
human sign-off), and a few past examples to seed an eval set
Strategy/commercialization plays (e.g. Segment Expansion, Data Productization) are intentionally capped at *prototype only* — they are product ideas, not operational agent loops.
What improves scaffold quality
The Agent scaffolds option builds pilot-ready material (runtime contract, eval suite, gap workflow, pilot package) on top of the blueprints. Beyond what raises readiness above, the following evidence makes the scaffold concrete and safe to pilot:
- Sample records (a few redacted tickets, invoices, RFQs) — seed golden and
edge-case evals instead of purely synthetic ones.
- Named tools/systems the agent would touch (CRM, ticketing, ERP) — realistic
tool contracts, permission manifest, and mock connectors.
- Write vs draft-only permissions — the human-approval workflow and
least-privilege scopes hinge on this.
- Approval boundaries (what needs human sign-off) — shape approval steps and
unsafe-action checks.
- Past successful cases / resolutions — become golden cases and before/after
metrics for the pilot.
Scaffolds are pilot accelerators, not live deployments — scaffolding plus a validation loop to prove an agent before any production rollout.
What improves pilot-validation quality
The Pilot validations option checks whether a scaffold is actually complete and safe enough to run a mock pilot with — never against a real system. Beyond what improves scaffold quality above, the following evidence raises the readiness verdict from *needs evidence* toward *ready for mock eval*:
- A named mock connector per tool — closes the gap that otherwise leaves
a tool's stub "uncovered" in the mock connector pack.
- A pass/fail rubric alongside golden cases — without one, the golden
cases exist but can't be judged, so the eval run plan isn't runnable.
- A human-approval workflow for every write/action tool — this is the one
gap that alone forces the verdict to *unsafe to pilot*, regardless of everything else.
- Closed evidence tasks from the readiness gap workflow — open tasks keep
the verdict at *needs evidence* even when everything else checks out.
- Unsafe-action checks — expected wherever a write/action tool exists;
their absence is flagged as a gap in the eval suite.
Pilot validation never talks to a real customer system — it proves a scaffold is worth building against one.
What improves mock pilot execution quality
The Pilot executions option actually runs a `ready_for_mock_eval` validation's eval plan against mock connector stubs — offline only. Beyond what improves pilot-validation quality above, the following evidence raises the odds the mock run passes rather than fails or blocks:
- Fixture-shaped sample inputs and expected outputs on every case — a
case missing an expected output can't be graded and is scored `fail`, not skipped.
- A case that actually exercises every unsafe-action check — a check that
is declared but no case shares its keywords with shows up untested, and pairs with any write/action tool to force `unsafe to run`.
- A realistic, parseable `minimum_pilot_threshold` (e.g. "90% precision")
— without one, the run conservatively assumes a 100% pass requirement.
- A named mock connector for every tool — the same gap that leaves a
tool "uncovered" in the validation's connector pack blocks every case in the execution run, not just that one tool.
- Closed evidence tasks and an approved runtime contract — execution
refuses to run at all while the underlying validation is `needs_evidence` or `unsafe_to_pilot`; it reports why and stops there.
Mock pilot execution proves the scaffold's wiring — tool coverage, fixture completeness, safety-check coverage, threshold discipline — is sound. It does not grade real agent reasoning: a mock connector stub is defined to echo back the fixture-declared expected output, so this is a wiring check, not a substitute for testing against a live system.
What improves human pilot quality
The Human pilots option converts a `mock_eval_passed` execution into a structured pilot plan a technical owner can execute. Evidence that raises the odds the pilot plan clears all gates:
- A named pilot cohort (`users`) — the manifest blocks if no pilot
participants are identified. Add names or team identifiers to the pilot package in the scaffold.
- Sample data requirements — specify the scope and redaction requirements
for real records the pilot needs.
- Kill / rollback criteria — explicit stop conditions are required before
a human pilot can be scheduled.
- Security review checklist — the gate blocks without a completed or
scoped security checklist.
- Success metrics (`before_after_metrics`) — measurable before/after
outcomes let the pilot produce a defensible verdict.
- A named human approval owner for write/action tools — the manifest
requires `human_approval_workflow` when any write/action tool is declared; without it, no human pilot can start.
What improves production deployment quality
The Production deployments option converts a `ready_for_human_pilot` plan into a production readiness package. Evidence that clears the production gates:
- A completed human pilot — the gate requires `ready_for_human_pilot`;
anything else blocks.
- Rollback / kill switch criteria — required at the production stage, not
just the human pilot stage.
- Monitoring plan — define alerting and observability targets before the
package is assembled.
- Security review checklist — the same checklist required for the human
pilot must also be present at the production stage.
- Sign-off checklist — named owners who must sign off before production.
- Eval plan with golden cases and a rubric — the production eval plan
converts these directly into a regression suite. Without them, the production gate blocks.
- Approval owner for write/action tools — the same requirement from the
human pilot carries through; still required in the production package.
Downloading the deployable scaffold
Once a candidate reaches a production deployment, its card gets a Download deployable scaffold (code) button — a zip of real starter files (`agent/runtime.py`, tool adapters, `.env.example`, `Dockerfile`/`docker-compose.yml`, an eval harness, `OPERATING_MODEL.md`, and a handoff README) generated deterministically from that candidate's runtime contract and connector inventory. What's inside follows an honesty boundary per tool, not one mock skeleton:
| Tool shape | What ships |
|---|---|
| Spreadsheet-backed | The company's real bundled `.xlsx` data plus a real reader/writer — no mock. |
| Typed connector (`rest_api`, `http_webhook`, `sql`, `local_file`, `smtp`) | A real, config-only adapter — set the env vars in `.env.example`, no code changes. |
| `bespoke_api` | A real HTTP call, with one `build_request()` mapping function marked TODO — a proprietary API's contract can't be inferred from a company model alone. |
| Everything else | An offline fixture-echoing mock stub until you replace it. |
`agent/runtime.py` reasons for real by default — a genuine LLM tool-calling loop, or (for the ticket + FAQ spreadsheet shape) a real batch runtime that drafts replies from bundled data. Every write/action tool is blocked until a human sets `AGENT_APPROVE_WRITES=1`. This is a scaffold, not a live deployment: no credentials are created and nothing is provisioned — a technical owner still supplies real credentials, an LLM key, and verifies any remaining mock or best-effort adapters before running this unattended.
Candidate scorecard (Generate only)
Inside the recommended plan, Candidate scorecard rates each bet 1–5 on:
| Axis | What it means |
|---|---|
| Value | Business impact — real pain and plausible upside |
| Feasibility | Can you build it with today's data, tech, and skills? |
| Time to value | Speed to first results (5 = weeks, 1 = many quarters) |
| Risk | Execution / regulatory / reputational risk (5 = low risk) |
| Moat | Defensibility — uses something only this company has |
Color follows the score (green 4–5, amber 3, gray 1–2). These come from the plan selector LLM and may differ slightly from the I · F · M · Q badges on Analyze cards — read the Why these scores line when they disagree.
Tips for better output
- Be specific — "RFQ → quote turnaround" beats "operations."
- Ground with data — accessible datasets lift feasibility, moat, and quality.
- Record real pains — linked pains are the strongest impact signal.
- Watch operator coverage — expand dormant operators when planning interviews; aim for green on patterns that match real leverage, not every box ticked.
- YAML tab — save and reload the model between sessions; switching back from YAML does not re-parse unsaved edits.
- One department per session — use Company context so each contributor sees relevant prompts without drowning in the whole company.