This guide is for practitioners running AI opportunity discovery: innovation leads, transformation teams, consultants, product leaders, and domain experts who need a repeatable, evidence-backed process — not another brainstorm.
For field-by-field builder instructions, see Builder guide. For technical architecture, see methodology.
---
How to use this guide
- If you are new, read in order: What it does → Getting started → Discovery program.
- If you are actively running sessions, jump to Building the evidence map,
Analyze vs Plan, and Troubleshooting.
- Keep the Builder guide open beside this guide for field-level help.
What Use Case Foundry does
Use Case Foundry is an AI Opportunity Discovery Engine. You build a structured model of a company — offerings, processes, data, skills, pains, constraints — and the engine systematically generates and ranks AI use cases that *that company in particular* is positioned to win.
The output is not a long generic list. It is a short, ranked portfolio with:
- Transparent scoring on Impact · Feasibility · Moat · Quality
- Explicit prerequisites (e.g. instrument data before automating)
- Per-candidate evidence strengthening guidance and interview questions
- A sequenced plan: quick wins plus one strategic bet
- Optional decision-ready dossiers, agent blueprints, pilot packages, and deployable scaffold exports
- A downloadable evidence & provenance map that stakeholders can open offline
Most AI ideation fails because it is generic. Use Case Foundry shifts discovery from open-ended brainstorming to an evidence-based system where every recommendation traces back to facts you captured — and gaps you still need to close.
What good looks like
By the time you are ready to Plan, you should usually see:
- High-severity gaps mostly closed in Improve the model
- At least one high-impact pain linked to each priority workflow
- Accessible data marked for key high-volume processes
- Clear prerequisites and sequencing (not a flat idea list)
- Only the candidates you actually want to take forward checked for Include in Plan
Who this is for
| Role | What you get |
|---|---|
| Innovation / transformation lead | A defensible shortlist to fund, with prerequisites sequenced |
| Consultant or advisor | A structured discovery program you can run across client sessions |
| Functional leader (finance, ops, support, HR, …) | Department-scoped discovery without drowning in the whole company |
| Executive sponsor | Quick wins vs strategic bets, with business case framing from Develop use cases |
| Domain expert / SME | Contribute facts through interviews and evidence capture — no strategy doc required |
You do not need to understand the engine internals. You do need to treat discovery as a program over days, not a form you finish in one sitting.
Core concepts
Three layers
- Mapping — Turn what you know about a company into a structured evidence graph
(nodes: offerings, segments, accounts, processes, skills, assets, projects, pains, constraints, plus optional strategy/governance nodes; typed edges connect use, ownership, decisions, measurement, contracts, and locations).
- Generation — 101 evidence-backed plays are tested against reviewed graph
facts to produce company-specific candidate use cases across operations, commerce, product, strategy, data/AI governance, legal/contracts, cybersecurity/IT service management, finance planning, workforce lifecycle, facilities, software, physical work, and R&D.
- Prioritization — Candidates rank on Impact × Feasibility × Moat, with a
quality gate that demotes generic ideas. A Target Selector picks quick wins and one strategic bet, respecting prerequisites.
Hybrid engine
- Deterministic evidence checks decide *when* a use case should exist and compute base scores.
- LLM reasoning (only on Develop use cases) writes *how* it should be executed for this company.
- Analyze uses no LLM — instant feedback as you build the model.
This separation matters: you can iterate on the evidence map all day without API cost, and you always know which scores came from rules vs model output.
The evidence graph is the product
Facts live in nodes; value lives on edges:
- Offering → serves → Segment — who buys what; supports growth and retention analysis
- Process → uses → Data asset — what evidence can ground a workflow-level recommendation
- Process → bottlenecked_by → Skill — where expert judgment gates throughput
- Offering → reuses → Project — proven capability that may be packaged or reused
- Initiative → advances → Goal and Process → owned_by → Role — reviewed
strategic scope and accountability
- KPI → measures → Process and **Contract → creates_obligation →
Obligation** — explicit outcomes and post-signature duties
- Process → located_at → Location and **Offering → delivers_via →
Location/channel** — operating footprint and delivery route
- Pains attached to nodes — the strongest signal on the impact axis
An incomplete graph is expected at first. Missing evidence is a roadmap, not a failing grade.
The optional Strategy, governance & footprint builder area holds goals, initiatives, roles, decisions and approvals, KPIs, contracts, obligations, and locations. It stays collapsed until needed. When the Strategy & Capital pack is enabled, reviewed goals, initiatives, KPIs, owners, and decisions can unlock three strategy-specific operators. Reviewed Data & AI governance profiles on assets can unlock four data/model control loops; asset owners, approvals, and KPIs sharpen their prompts. Contracts, obligations, and locations still mainly enrich other packs. A named distributor or reseller is still an Account with channel kind, while `locations.kind: channel` describes a route-to-market type.
Cybersecurity & IT Service Management uses the same process, IT-role, approval, KPI, asset, account, contract, pain, and constraint evidence. Name the actual lifecycle—such as SOC alert investigation, identity entitlement review, phishing response, vendor security assessment, IT service request, CAB change, or problem/known-error management. Bare “monitoring,” “access review,” “incident,” and “change management” remain intentionally dormant.
Finance Planning, Treasury & Tax also uses the same graph plus the existing optional `financial_profile`. Name the governed workflow—such as driver-based financial planning, daily cash positioning, FX exposure and hedge review, a tax compliance calendar, or capital-project appraisal—and connect finance owners, approval decisions, KPIs, and source assets. Bare “budget,” “spreadsheet,” “variance,” “tax,” “cash,” “hedge,” “investment,” and “capital” remain dormant. Prediction stays in Forecasting; close/reconciliation stays in Financial Close; receivables/payables cycles stay in Working Capital; multi-initiative prioritization stays in Strategy & Capital.
Workforce Lifecycle uses aggregate or minimum-necessary reviewed facts in the same graph—roles, skills, initiatives, decisions, KPIs, processes, assets, obligations, locations, pains, and constraints. Name the governed workflow, such as strategic workforce planning, internal mobility, a performance or compensation control exception, aggregate retention action review, or a labor obligation control gap. Generic performance reviews, compensation words, skills searches, labor policy, fairness, bias, or support language stay dormant. Recruiting/onboarding remains People Ops; attrition/headcount prediction remains Forecasting; policy/audit evidence remains Compliance; bias monitoring reaches AI Eval & Drift only for an explicit model or AI system. No person-level sensitive workforce profile is collected.
Assets, Facilities, Energy & Sustainability also uses the same graph. Name the governed workflow—facility portfolio optimization, site energy performance, EHS control-gap remediation, emissions inventory/reporting, or capital- maintenance lifecycle planning—and connect named sites/physical assets to reviewed source evidence, owners, approvals, KPIs, contracts/obligations, and constraints. Optional facility and capital-maintenance profile values must be source-backed. Generic building, energy, safety, emissions, ESG, carbon, capital, or maintenance words stay dormant. The engine does not autonomously control assets/buildings, close safety cases, file or publish claims, alter leases, or commit capital.
Regulated-domain matches are also filtered by domain anti-triggers. Healthcare, legal, and similar operators stay silent on clinical, diagnostic, or counsel-practice language unless the named workflow is an explicit administrative or revenue-cycle surface.
Ranking axes
Every candidate carries four badges:
| Badge | Question it answers |
|---|---|
| I (Impact) | Does this matter? Driven by linked pains (severity × frequency, optional cost) |
| F (Feasibility) | Can we build and run it? Adjusted by constraints (budget, talent, risk, stack) |
| M (Moat) | Can competitors copy it? Raised by proprietary data, abstractions, scarce expertise |
| Q (Quality) | Is it specific and grounded? Generic candidates are demoted |
Candidates also show dependency lines: do first (prerequisites), unlocks (what they enable), overlaps with (related efforts on the same subgraph).
Each candidate now also includes an Economics section with deterministic annual value range status (`sourced`, `inferred`, or `needs baseline`), an effort-cost band, and (for quick-win lane candidates in a generated plan) payback-band months. Use Edit assumptions on the card to adjust addressable-share assumptions and linked pain baselines, then let Analyze recompute.
Tilt ranking with risk appetite (averse / balanced / aggressive) or the feasibility/moat slider — averse favors quick wins; aggressive favors defensible bets.
Getting started
Run the app
```bash pip install -r requirements.txt python -m engine.server --open # marketing site at /, app at /app ```
Hosted deployments serve the same UI. LLM keys are server-side environment variables only. Deployment setup is typically handled by your technical owner.
The three-step journey
The app organizes work into three focused steps. Only one step is visible at a time — use the top Source → Evidence → Review nav (or URLs `#source`, `#evidence`, `#review`, `#outside-in`) to move between them without losing context.
| Step | Name | What you do |
|---|---|---|
| 1 | Source | Load an example, auto-map from sources, generate a pre-assessment starter, or resume a saved model; optionally set Company context |
| 2 | Evidence | Add facts in the guided builder; any substantive model unlocks the Outside-in lane; after the first analysis, add capture or CX evidence only when useful |
| 3 | Review | Approve company-map proposals, run Analyze, read the board, close gaps, choose candidates, then Plan (Develop use cases and/or Plan adoption) |
You do not need every field before getting value. On Evidence, debounced Analyze (no LLM) keeps the preview fresh: supported patterns, high-priority gaps, and the leading candidate. The full coverage panel and grouped candidates live on Review.
Outside-in is an optional parallel lane (`#outside-in`) rather than a numbered step. It opens after the model has substance and keeps scanning, external-evidence review, demand-hypothesis derivation, accepted ledgers, and account plays on one page. The Review page shows a pending-item link back to the lane while retaining company-map proposals and the opportunity board.
The lane adapts to the business model: scan named customers, channels, competitors, and suppliers; discover association, regulator, and trade-news sources when named accounts are unavailable; or use aggregate consumer-voice and platform-policy sources for B2C companies. Competitor discovery and claims remain review-gated. Accepted evidence can sharpen ranking, open bounded demand-gap reviews, and support reviewable named-account plays, but it never silently becomes company truth or asserts private customer intent.
Step actions:
- Source → Evidence: Continue to Evidence (enabled once the model has a name or facts)
- Evidence → Review: Analyze & Review board re-runs analysis, then opens Review; Open Review board jumps to the last board without re-analyzing
- Evidence → Outside-in: Open Outside-in lane when the model has substance
- Pre-assessment starter: Generate it from the same source inputs, then stay in Source / Evidence until you either Save starter or Promote to validation workshop. The Review board stays locked while the model is still a starter.
- Back buttons return to the previous step; gap Open links on the preview jump back to the exact builder field
Your first session (30–60 minutes)
- Open `/app` and dismiss the welcome screen (or load an example from it).
- In Step 1 · Source, load Meridian Fab or Beacon Pay, then Continue to Evidence.
- In Step 2 · Evidence, use Build your evidence map (+ Offering, + Process, …) or change one field — watch the evidence preview react.
- Click Analyze & Review board (or Open Review board if you already analyzed) — read the coverage panel, top candidates, and Improve the model.
- Do not Plan yet. Note the top 3 high-severity gaps; use Open on a preview gap or Open in builder on Review to jump back to Evidence.
That session teaches the loop. Real discovery continues across days and contributors.
The fastest path: load a bundled demo
If you want to see the *end* of the pipeline before building your own model, load the Atlas Ledger Production Demo (finance back office) or Meera Home & Kitchen example, or choose the Beacon Pay Production Demo for a regulated chargeback workflow. All three ship with pre-built Develop use cases output that already cleared every deterministic gate — pilot validation, mock execution, human pilot, and production handoff — so Analyze & Review board → Develop use cases with Production deployments ticked returns a finished package in seconds, no LLM call required for the bundled output. Open the production deployment card and click Download deployable scaffold (code) to see the real starter files described below before you invest time modeling your own company. Viewing and downloading these demos needs no API key; running an exported agent against real systems still requires your own LLM endpoint, credentials, and technical review.
The discovery program
Honest company models are multi-person, multi-day work. The product is designed for that.
Recommended cadence
| Phase | Actions |
|---|---|
| Day 0 | Auto-map + set archetype + one department lens → first draft. Run Analyze — expect gaps. |
| Each session | Pick one department focus. Close only the top 3 gaps in Improve the model. |
| Between sessions | Save to server (when signed in) or copy the YAML tab to a shared file. |
| Before Plan | Close high-severity gaps: linked pains, accessible data on high-volume processes, project abstractions where they matter. |
The interview loop
Interviews are first-class — they turn assumptions into grounded facts:
``` auto-map → generate interview agenda → conduct interviews → extract/update evidence graph → analyze → fix gaps → repeat ```
Every Analyze returns an Interview Sprint panel:
- Agenda — ask first: highest-leverage questions from current gaps
- Ask in the interview: candidate-specific questions on the opportunities you are validating
- Gap checklist: severity, why it matters, questions to ask
- Status per gap: Open → Captured → Done (tracks multi-day progress)
- Open in builder: jumps to the section that closes the gap
- Conduct interview: opens Map from evidence for that gap
Use Copy agenda as Markdown to take questions into a call, Slack, or email. After applying evidence proposals, the model re-Analyzes automatically.
Questions created by Plan adoption appear in the same Interview Sprint panel, on the Adoption readiness filter. Discovery evidence keeps Show all gaps / Show priority gaps only. Adoption readiness uses question language instead. The plan card shows how many need attention and opens that filter with Resolve readiness questions. Adding a fact moves you to Evidence and shows Return to readiness questions; the question view also has Back to adoption plan. When all plan questions have been reviewed, choose Update adoption plan.
Pain Discovery Studio
After mapping at least one process, use the two paths in Evidence → Pain & value interview: Add known pain when the client can name the friction, or Discover hidden pain when a blank-sheet “What are your pain points?” question is unlikely to work. The latter saves the current evidence map, opens `/pain-discovery` with the selected company and workflow, and keeps the workspace-header Pain Discovery link as a secondary shortcut.
Selecting a company should show Loaded {name}. Continue to tools creates a session if you have not already. The toolkit pre-checks a recommended set from the map (always hypothesis recognition, consequence scan, last bad day, and evidence plan; plus workaround archaeology and decision-latency when processes exist; policy-to-practice when contracts or obligations exist; customer echo when accounts or segments exist). Uncheck anything you do not want, then Save tool selection.
- Context — company, function, optional workflow focus, and objective.
- Tools — choose complementary elicitation techniques from the 15-tool
catalog.
- Elicit — Generate up to six hypotheses from saved processes. Unmapped
workflows are asked first. Workflows that already have a recorded pain still get one adjacent-friction card that uses a *different* map fact (accessible data, dollar impact without an operational baseline, scarce skill, shadow workaround, or a handoff), so a fully mapped example such as Lumen Sales Ops or Cedar Facilities is not a dead end. Cards are not a shared ROI template. Run another selected technique picks the workflow from a dropdown of those saved processes so a later promotion can attach to the same node. Record recognition, consequences, severity, frequency, and the least-sensitive evidence rung.
- Validate — one target, with owner, evidence request, baseline, and stop
condition.
- Review — preview eligible pain proposals, then **Add reviewed pains to
model. That creates a new saved-model version. Choose Return to Evidence → Pains to inspect the reviewed findings in context, then Analyze again** so the ranking uses them. The return reloads the latest saved version rather than an older builder-tab draft. You can also export the session as Markdown.
The page labels hypotheses separately from facts. `Unsure` and `Prefer not to answer` do not count as refutation. Session findings do not affect ranking until explicitly promoted. UCF never infers annual cost or revenue from qualitative recognition.
Artifact-backed techniques can use the existing redacted evidence extractor, and Customer echo can review a CSV/XLSX customer-experience export. Other tools are facilitator-guided in this version. There are no participant links or live ERP/CRM trace connectors on this page.
Consultant interview guide
Consultants often need to extract critical evidence from partial memory and cross-functional context. Use the dedicated consultant interview guide for:
- Leading/searching prompts by graph section
- Memory-retrieval question funnels
- Edge-completion checklists
- Session templates and anti-patterns
Continuity and saving
| Method | Survives refresh? | Cross-device? |
|---|---|---|
| Browser autosave (in-progress draft) | Yes, same browser | No |
| Save to server (when auth enabled) | Yes | Yes, when signed in |
| YAML tab export | Yes, if you save the file | Yes, via git/Drive/etc. |
Important: if auth is not enabled and you have not exported YAML, treat the browser as your only copy. When signed in, saved models are private to your account.
Context lenses
The company graph is the truth. Context lenses change how it is elicited, ranked, and framed — they never overwrite facts you captured.
Set them in the four-question Company context card on Source. Detailed SME, department, data-readiness, and peer controls stay optional.
| Question | Controls |
|---|---|
| Who is the company? | Archetype + resources |
| What matters this cycle? | Near-term focus + longer-term focus + budget |
| Who reads the output? | Audience |
| Use market context? | Peer lens; detailed lenses under More context lenses |
Archetype
Set under Source → Company context. Shapes section order and empty-state guidance:
| Archetype | Typical starting point |
|---|---|
| Startup | Segments and pains first, then unfair advantage (data, skill, method) |
| Scaleup | Offerings, segments served, data assets behind them |
| Agency | Delivered projects and reusable abstractions |
| Enterprise / SMB | High-volume processes, data produced, where time or money leaks |
Also sets default resource level (capacity to build and run AI).
Department lens
Set under More context lenses. Scopes a session to one function: finance, operations, customer support, HR, sales, marketing, healthcare ops, supply chain, compliance, software/engineering, …
Use this when different contributors own different parts of the company. Each person sees relevant section hints and pain prompts without the full org at once.
Data-maturity lens
Set under More context lenses. Adjusts expectations for data readiness: spreadsheet/manual ops through to ML-native. Influences gap questions and feasibility framing — does not invent data you do not have.
SME lens (e.g. India SME)
Set under More context lenses. Market context for operating realities: spreadsheet-first data, owner-led approval, lean IT, short ROI horizons. Supplies defaults and framing.
Peer lens
Set under Company context → Use market context? Reference company archetypes (not real clients) that suggest which gaps to close first while your graph is thin. Modes: off, auto-match, or manual pick. Never copies another company's graph — only nudges gap ordering and ranking early on.
Strategic focus lens
Set under Company context → What matters this cycle? Near-term and longer-term focus areas re-rank candidates: cost, productivity, innovation, risk, customer experience, … Plus budget direction (flat / slashed / increased — fund quick wins vs bets).
Affects Analyze ranking and Develop use cases framing.
Audience lens
Set under Company context → Who will read the output? Shapes Develop use cases output only (not Analyze scoring): founder, transformation lead, functional VP, agency partner, … Same candidates, different dossier emphasis.
Building the evidence map
On Step 2 · Evidence, the guided builder elicits facts that support stronger recommendations. Use Build your evidence map quick-add buttons (+ Offering, + Process, …) to seed sections; each click opens the matching section below.
While you edit, the evidence preview (top of the step) summarizes the latest Analyze run: supported patterns, high-priority gaps, and the leading candidate. The full coverage panel (N/101 supported: green chips for grounded patterns; collapsed Not supported yet with section-level evidence guidance) appears on Review after you open the board.
Recommended fill order
Work top to bottom, prioritizing facts that improve the most downstream recommendations:
| Order | Section | Why |
|---|---|---|
| 1 | Company | Name + risk appetite set ranking tilt |
| 2 | Offerings | Most external plays start here; add components (`manual`, `rule_based`, `judgment`) |
| 3 | Segments | Link via serves — clarifies who each offering is for and where growth pressure sits |
| 4 | Assets | Accessible data is the main grounding signal; mark `accessible` only if usable today |
| 5 | Processes | Volume + repetitiveness often supports automation; scarce bottleneck skill often supports judgment tools; no data often points to instrumentation first |
| 6 | Skills | Scarce skills reveal where scaling judgment or preserving know-how matters most |
| 7 | Projects | Abstraction makes reusable capability visible beyond one delivery context |
| 8 | Pain & value interview → Pains | Attach pains to raise Impact |
| 9 | Constraints | Surface compliance, integration, guardrail, and knowledge-transfer work that may need to happen first |
Section order may reorder based on archetype. Hover any ⓘ icon for field help.
Detailed section cheat sheet: Builder guide
Pain & value interview
Structured prompts above the raw pains form — e.g. "Where is time wasted?", "Which workflows delay revenue?" — seed graded pains attached to nodes you select. Better pain inputs raise impact and quality measurability. Use this when the interviewee can already name the friction. Use Pain Discovery Studio when they cannot or should not be asked for a pain inventory.
YAML tab (power users)
The YAML tab remains for direct editing, git workflows, and round-tripping with the builder. Loading an example or auto-mapping populates the builder; edges the builder does not manage are preserved on round-trip.
Switching back from YAML without saving does not re-parse unsaved edits.
Auto-map: draft, not done
Auto-map drafts a starting model from:
- Website URLs (optional same-domain crawl)
- Pasted notes
- Uploaded files (`.txt`, `.md`, `.html`, `.pdf`, `.csv`, `.xlsx`, …), including annual and quarterly reports
It typically captures roughly half to two-thirds of a useful model from public material. It rarely includes linked pains, accessible data flags, bottleneck skills, or project abstractions.
For tabular uploads (`.csv`, `.xlsx`), Auto-map also computes uploaded-file insights before the LLM writes anything. When the file has useful operational signals, the app shows What your uploaded data shows with deterministic findings, a file-grounded AI use case, a week-one pilot shape, and buttons to Open live dashboard, Download dashboard (HTML), and Download one-pager. The first pilot is deliberately scoped to the uploaded file alone; every numeric claim is tied back to computed findings, and you can include that file-grounded use case in Develop use cases when you want a deployment package based only on the upload.
When an uploaded report contains comparable financial periods, Auto-map can also extract a cited financial profile for review. Accepted asset-utilization, working- capital, and margin signals can fire three dedicated opportunity patterns; they remain tied to their report period and source note.
For advisor-led discovery, the same source intake can instead generate a pre-assessment starter. That path stops before the ranked board and packages a first impression a skeptical advisor can take into the free, bounded Validation workshop:
- a company-specific first read (named offerings, workflows, and systems)
- 2-3 illustrative ways those facts could become AI use cases, showing where AI
would act, why it could matter, and what the workshop must prove
- 3-5 operational hypotheses worth falsifying on the first call — not generic
templates
- the evidence gaps most likely to change a funding decision, not schema leftovers
- the questions that decide whether an Opportunity sprint is warranted
The app saves the result automatically so it can be reopened in the same advisor workspace or client account. If the fast first read is usable but too thin, cited external pressure is checked in the background and added to that same pre-assessment. You can review the initial result immediately. Promote it only when you are ready to run the free workshop, validate plausibility, and unlock the Review board for a paid sprint decision.
Download working paper creates a standalone HTML artifact of the advisor card — useful for your own notes. Do not forward that file to the company. Prepare client send is the workflow for that: confirm your white-label brand, keep only the directions you can defend, copy a covering note, and Download client brief. That file talks to the company, calls the next step a working session rather than a scoped commercial proposal, and is safe to print to PDF.
Do not rely on auto-map alone. After a full auto-map import (not the starter path):
- Run Analyze
- Close top gaps in Improve the model
- Add internal evidence auto-map missed
- Re-Analyze until high-severity gaps close
On the starter path, stay with the first-read card until the advisor is ready to promote into a workshop. Analyze on unreviewed public pages is how the ranked board looks generic. The starter's AI-use-case directions are deliberately unranked: they make the path from evidence to AI work visible, while the paid workshop determines whether the problem, data, owner, and value mechanism are real.
Use Merge into current model when extending an existing draft with new sources. Review review notes after import — inferred fields are flagged for human edit.
Identifiers in pasted notes and uploads are redacted server-side before any LLM call. See Trust and evidence boundary.
Workstation capture (evidence from how work actually happens)
Auto-map reads what a company says about itself. Workstation capture drafts model facts from what employees actually *do* — opt-in, local-first, and review-only. After the first analysis, open Add more evidence → Workstation capture; it turns recurring app/window/SaaS activity into `processes` (with volume and repetitiveness), touched `assets`, `uses` edges between them, and possible scarce `skills` — plus constraint hints from domain words like audit, approval, or HIPAA that a human must confirm.
One-click, local (Foundry and ActivityWatch on the same machine):
- Keep ActivityWatch running in the background.
- Run Analyze once, then open Add more evidence → Workstation capture.
- Click Import from ActivityWatch — the app fetches, summarizes, and opens
the review panel. No export or paste required.
Hosted Foundry (capture agent): run a small tray app on each workstation (`python -m collector.capture_tray`). It keeps ActivityWatch and summarization local and pushes only a redacted summary to your Foundry URL. Signing in shows a Workstation capture tab with two zones — Ready to review (workflow candidates that crossed the graduation gate) and Still watching (candidates still accumulating sessions/days before they earn review attention, which you can dismiss if noisy).
Either path returns patch proposals reviewed in the same panel Auto-map uses — nothing merges into your model without a human accepting it. Denylists for apps, domains, and folders apply before any event is summarized, and raw activity never leaves the employee's machine — only a redacted aggregate summary (or, for the CLI/API path, a fully local one-shot summarize-then-push) does.
Capture micro-interviews include naming, pain confirmation, cost baseline, volume, and skill-coverage prompts. Cost answers flow into `pains.annual_cost` so economics ranges can move from `needs baseline` to `sourced`.
See Trust and evidence boundary for the full privacy design.
Map from evidence (private discovery)
When facts live in internal systems (CRM, ERP, ticketing, spreadsheets), users may not want to upload raw records. Map from evidence closes gaps without full data sharing.
Privacy tiers
| Tier | What you share | Typical use |
|---|---|---|
| No-data guide | Local inspection + typed observations | First pass; nothing leaves browser until Extract |
| Metadata only | Field names, objects, stages, report titles | Reveal processes/assets without record contents |
| Aggregate only | Counts, rates, medians, distributions | Surface pains and bottlenecks safely |
| Redacted sample | Anonymized snippets | Ground abstractions and edge cases |
Raw evidence is not persisted. Only human-approved graph patches merge into the model.
Workflow
- Run Analyze → open Improve the model
- Pick a high-severity gap → Map from evidence
- Choose privacy tier; enter observations
- Review graph patch proposals → apply selected changes
- Re-Analyze and repeat
For stricter boundaries (VPC-only LLM), ask your technical owner to run the private-runtime deployment profile.
Analyze vs Plan
| Action | LLM | Output |
|---|---|---|
| Analyze | Not used | Prescores (I·F·M·Q), deterministic economics ranges + effort bands, coverage panel, Improve the model panel, grouped candidates, Interview Sprint |
| Develop use cases | Used | Full use cases and business artifacts for the candidates you selected, plus sequenced quick wins and a strategic bet. Historical name: Generate full plan / Write-up |
| Plan adoption | GraphFoundry (optional LLM for swarm) | What must be established, in what order, and through which route: foundations, readiness, and review tickets. Continues the last saved adoption plan unless you start a new one. Historical name: Build adoption plan / Rollout |
| Critique & rewrite | Used (with Develop use cases) | Extra pass flagging generic candidates before final ranking |
| Opportunity dossiers | Used (with Develop use cases) | Decision-ready one-pagers: business case, deterministic economics summary, lane-aware payback framing, pilot scope, risks, metrics, kill criteria |
| Agent blueprints | Used (with Develop use cases) | Parameterized agent specs + deployment-readiness kits for the selected opportunities, gated by evidence confidence |
| Agent scaffolds | Used (with Develop use cases) | Pilot-ready scaffolds on top of the blueprints: runtime contract, eval suite, readiness gap workflow, and pilot package (auto-enables Agent blueprints) |
| Pilot validations | Used (with Develop use cases) | Mock-eval readiness report on top of the scaffolds: validation manifest, mock connector pack, eval run plan, and a pilot readiness verdict (auto-enables Agent blueprints + scaffolds) |
| Pilot executions | Used (with Develop use cases) | Actually runs the mock eval suite from a ready validation against mock connector stubs: case results, unsafe-action check coverage, a threshold result, and a mock pilot verdict (auto-enables Agent blueprints + scaffolds + Pilot validations) |
| Human pilots | Used (with Develop use cases) | Converts mock-eval-passed executions into a structured human pilot plan: protocol, evidence capture checklist, risk register, completion criteria, and a pilot readiness verdict (auto-enables all prior phases) |
| Production deployments | Used (with Develop use cases) | Converts ready human pilots into a production deployment readiness package: connector inventory, deployment package, operating model, production eval plan, and a handoff readiness verdict (auto-enables all prior phases) |
When to Analyze: constantly — as you edit, after every evidence apply, between interview sessions. It is free and instant.
When to Develop use cases: when high-severity gaps are closed and you need sponsor-ready use cases and packaging. It usually takes a few minutes and requires LLM configuration.
When to Plan adoption: when you have checked the use cases you want to sequence and need to know what must be established, in what order, and through which route. Plan adoption continues the last saved plan for the company; use New adoption plan to compose from the current evidence map. The completed result has a Download plan action that exports selected use cases, the recommended sequence, foundations, route comparison, ranked foundation sensitivity (rank, route, and damage score), review lenses, and findings as Markdown. If the plan raises missing foundations or diligence questions, follow the guided loop: Plan adoption → Resolve readiness questions → Add evidence → Update adoption plan.
After Analyze, every candidate card has an Include in Plan checkbox. The default is the top five non-generic candidates; use Top 5, All, or Clear in the selection bar to choose the set explicitly. Prerequisites for a selected opportunity are included automatically, and candidates you leave unchecked remain on the board as Analyze only context rather than consuming LLM time.
Use the How far should use-case development go? slider to choose the artifact depth: roadmap only, agent blueprint, scaffold, validation, mock execution, human pilot, or production handoff. The advanced checkboxes mirror that depth for teams that want exact stage control.
Next-step engagement paths
Use Review to understand and challenge the ranking. When you are ready to continue, choose the outcome rather than a feature preset:
- Validate the shortlist opens a dedicated Validation workshop page. It
explains the facilitated evidence review, who participates, what you receive, what is excluded, and how to scope it. The workshop stops at Analyze and the bounded shortlist validation record; selecting or booking it does not unlock paid product actions. The workshop does not rank opportunities, compare scorecards, or produce economics/adoption sequencing.
- Turn the shortlist into a decision plan opens the Opportunity sprint
page. Request Opportunity sprint records a request for the current workspace. It does not unlock anything until an operator activates it. An already-active Sprint or Custom entitlement cannot be replaced by requesting a different path; ask the operator to change it. Once active, Configure the sprint workflow returns to the same workspace and applies the standard sprint controls: system-data graph, critique, and dossiers on; agent stages off; depth at Roadmap. It does not run any action automatically.
Until the current workspace has an active Opportunity sprint or Custom entitlement, Develop use cases and Plan adoption remain disabled; Analyze remains available. Sprint access does not automatically include agent blueprint, scaffold, pilot, or production stages. After use cases are developed, choose Request pilot handoff on the specific candidate that is funded. Operator approval unlocks pilot stages for that saved-model candidate only; other candidates require separate requests and approvals. Advanced agent pipeline remains a bespoke workspace-wide exception. Browser settings, links, and presets cannot grant either entitlement.
The standard sprint includes Plan adoption, which compares rollout sequences. It does not include Pilot executions, which test agent behavior on mock cases. Agent-behavior simulation belongs in a later pilot handoff, not the sprint.
Without LLM configured, Auto-map and Develop use cases return ready-to-run prompts you can execute elsewhere.
Agent blueprints and deployment readiness
Because the engine already knows which patterns are already supported, it can take the selected opportunities one step further: tick Agent blueprints before Develop use cases to get a parameterized agent spec plus a deployment-readiness kit for each — mission, operating loop, inputs/outputs, tools and connectors, human approval points, then permissions, an eval plan, telemetry, rollout steps, and guardrails.
This is a specification, not an automated deployment. The hard parts of shipping an agent — connectors, permissions, data quality, approval boundaries, and evaluation — are surfaced honestly rather than hidden.
Crucially, graph evidence confidence drives whether an agent is justified, not just whether the opportunity is interesting. Each blueprint carries a readiness status, computed from your facts (not the LLM):
| Status | What it means |
|---|---|
| Ready | Strong evidence: linked pain, accessible data, specific grounding, no unmet prerequisite |
| Needs discovery | Worth building toward, but key evidence (data access, owners, scope) is still missing |
| Prototype only | A product/commercialization idea or thin evidence — explore as a prototype, not a deployed agent |
| Not recommended | The evidence does not justify creating an agent yet (no deployment kit is written) |
A use case can rank highly yet still be only "needs discovery" — high impact does not mean deployment-ready. Close the gaps the blockers list calls out (link a pain, mark data accessible, confirm an approval owner) and the status rises.
Agent scaffolds and the pilot loop
Tick Agent scaffolds to go one step further than the blueprint (this auto-enables Agent blueprints, since scaffolds build on them). For each agent you get four pilot-ready layers:
- Runtime contract — agent instructions, tool contracts, input/output schemas,
a least-privilege permission manifest, environment-variable names, mock connectors so the pilot runs offline, and a human-approval workflow.
- Eval suite — golden cases, synthetic edge cases, a pass/fail rubric,
unsafe-action checks, and the minimum threshold to clear before you trust it.
- Readiness gap workflow — each blocker turned into a concrete ask ("upload 10
redacted tickets", "confirm write vs draft-only access", "name the approval owner") so low readiness becomes an evidence loop, not a dead end.
- Pilot package — a scoped 2-week pilot: users, sample data, rollout steps,
monitoring, kill criteria, security review, sign-off checklist, and before/after metrics.
Each scaffold is classified by type — workflow agent, copilot agent, analytics agent, or prototype brief (for strategy/commercialization plays). These are pilot accelerators, not live deployments: scaffolding plus a validation loop to safely prove an agent before any production rollout. Testing expectation: run the eval suite against the mock connectors and clear the minimum threshold before the pilot, then use the gap workflow to close blockers.
Pilot validations and reading pass/fail readiness
Tick Pilot validations to go one step further than the scaffold (this auto-enables Agent blueprints and Agent scaffolds, since validation builds on both). Instead of just producing pilot material, this checks whether that material is actually complete and safe enough to run a mock pilot with — still never connecting to a real customer system:
- Validation manifest — what is testable right now vs. what is blocked by
missing evidence (e.g. no mock connector named for a tool, no pass/fail rubric for the golden cases).
- Mock connector pack — an offline stub for every declared tool: a
fixture-file name and the expected input/output shape, with write/action tools called out explicitly.
- Eval run plan — the golden cases, edge cases, unsafe-action checks, and
minimum threshold in one runnable plan, with placeholders and malformed entries already stripped out.
- Pilot readiness report — the verdict, in plain language:
| Readiness | What it means | |-----------|---------------| | Ready for mock eval | Tools, mock connectors, golden cases, a rubric, and (if needed) an approval workflow are all in place — run the mock eval suite | | Needs evidence | Something is missing or still open (evidence tasks, a rubric, connector coverage) — close it, then re-run | | Unsafe to pilot | A write/action tool has no approval boundary, or the underlying blueprint itself is not recommended — do not pilot yet |
A `prototype_brief` scaffold gets a prototype checklist instead of a mock-eval verdict — there is no runtime contract to validate, so pilot validation is skipped in favor of the smallest-prototype note.
This verdict is computed from the graph and the scaffold, not the LLM — the optional narrative (pilot summary, top risks, recommended next steps) can explain the verdict, but it cannot change it.
Pilot executions and reading mock-eval verdicts
Tick Pilot executions to go one step further than validation (this auto-enables Agent blueprints, Agent scaffolds, and Pilot validations, since execution builds on all three). Instead of just saying a pilot is ready to run, this actually runs the eval plan — offline, against mock connector stubs, never a real system:
- Execution manifest — whether this validation could actually be run, or
why it was skipped (not `ready_for_mock_eval`, a `prototype_brief`, no cases, or the mock connector pack isn't ready).
- Mock run inputs — the fixture-shaped input used for every case that ran.
- Case results — each golden/edge case scored pass (the mock stub
returned the fixture's declared expected output), fail (no expected output to grade against), or blocked (no connector coverage for that tool).
- Unsafe-action check results — whether every declared unsafe-action
check was actually exercised by at least one case. A check that is declared but never tested by any case shows untested — a real gap in the eval suite, not just a missing checklist item.
- Threshold result — cases passed / total against the declared minimum
pilot threshold (a missing or unreadable threshold conservatively assumes a 100% pass requirement rather than skipping the check).
- Mock pilot verdict, in plain language:
| Verdict | What it means | |---------|---------------| | Mock eval passed | Every executable case passed and the threshold was met | | Mock eval failed | The suite ran, but the pass rate fell short of the threshold | | Blocked | The validation wasn't ready to run, had no cases, or the mock connector pack isn't ready — see the execution manifest | | Unsafe to run | A write/action tool has an unsafe-action check that no case actually exercises, or the underlying validation itself was unsafe |
This phase is deliberately honest about what it can prove offline: a mock connector stub is defined to echo back the fixture-declared expected output, so a "pass" here confirms the scaffold's wiring — tool coverage, fixture completeness, safety-check coverage, threshold discipline — is sound. It does not, and cannot, grade genuine agent reasoning without a live LLM loop against real systems. The optional narrative (execution summary, failure analysis, recommended fixes) can explain the verdict, but it cannot change it.
Human pilots and reading pilot readiness
Tick Human pilots to go one step further than mock execution (auto-enables all prior phases). For each `mock_eval_passed` execution this generates a structured plan a customer technical owner can take into a real pilot:
- Pilot manifest — whether the human pilot is ready to run or what is
blocking it (missing pilot cohort, no sample data requirements, no kill criteria, missing security review checklist, no success metrics, or missing approval owner for write/action tools).
- Pilot protocol — duration, cohort, allowed workflows, human review
steps, escalation path, and success metrics.
- Evidence capture plan — what to collect during real pilot runs: agent
inputs/outputs, human decisions and overrides, unsafe-action attempts, latency, user feedback, and before/after metrics.
- Pilot risk register — residual risks, mitigations, named owners, and
stop conditions.
- Completion criteria — what "done" means for the pilot.
- Human pilot verdict, in plain language:
| Verdict | What it means | |---------|---------------| | Ready for human pilot | All gates cleared — the pilot can be scheduled | | Needs pilot setup | One or more required fields are missing — see the manifest | | Unsafe for human pilot | The underlying mock execution was `unsafe_to_run` — resolve safety issues first |
Production deployments and reading production readiness
Tick Production deployments to go one step further than human pilot planning (auto-enables all prior phases). For each `ready_for_human_pilot` plan this generates a production deployment readiness package:
- Readiness manifest — whether the production handoff is ready or what
is blocking it (incomplete human pilot, missing rollback plan, no monitoring, missing security review, no sign-off checklist, no eval plan, or missing approval owner for write/action tools).
- Deployment package — connector inventory with credential placeholders,
permission manifest, environment variables, rollout stages (shadow → supervised → production), rollback instructions, kill switch, and audit log requirements.
- Operating model — owner matrix for product, technical, security, support,
incident response, eval maintenance, and write-tool approval.
- Production eval plan — regression suite sourced from the mock eval cases,
shadow-mode checks, unsafe-action alert thresholds, and periodic review cadence.
- Production deployment verdict, in plain language:
| Verdict | What it means | |---------|---------------| | Ready for production handoff | All gates cleared — the deployment package can be handed to engineering | | Needs production setup | One or more required fields are missing — see the manifest | | Unsafe for production | The human pilot was `unsafe_for_human_pilot` — resolve safety issues first |
This phase produces a readiness package and gate, not an automated deployment. No credentials are created, no infrastructure is provisioned, and no external systems are called.
For concrete examples of this handoff path (ready, needs setup, unsafe), see [Evidence to deployment examples](EVIDENCE_TO_DEPLOYMENT_EXAMPLES.md) and `examples/evidence_to_deployment/`.
Each production deployment card also has a Download deployable scaffold (code) button — a zip of real starter files (`agent/runtime.py`, tool adapters, `.env.example`, `Dockerfile`/`docker-compose.yml`, an eval harness, `OPERATING_MODEL.md`, and a handoff README), generated deterministically from that candidate's runtime contract and connector inventory. This export requires an active Custom entitlement for the current workspace. What ends up inside depends on what each declared tool actually is — this is an honesty boundary, not a single mock skeleton:
| Export tier | What you get |
|---|---|
| Spreadsheet-native | A tool backed by a spreadsheet (e.g. a ticket queue or FAQ sheet) ships the company's real bundled `.xlsx` data plus a real reader/writer — no mock at all for that tool. |
| Typed connector | `rest_api`, `http_webhook`, `sql`, `local_file`, and `smtp` tools get a real, config-only adapter: set the environment variables in `.env.example` and it calls, queries, reads, or sends for real — no code changes. |
| Best-effort API | `bespoke_api` tools get a real HTTP call too, but the one `build_request()` mapping function is marked TODO, since a proprietary API's request/response contract can't be inferred from a company model alone. |
| Mock stub | Anything else falls back to an offline fixture-echoing stub — credentials are placeholders and nothing talks to a real system until you replace it. |
`agent/runtime.py` itself is real, not a placeholder — and its *structure* matches the agent's type, so packages differ in code, not just in the prompt:
- Workflow agents get a queue-processing runtime: batches are ranked and
deduped deterministically in code (bundled triage engine) before any LLM call, then processed per item in priority order.
- Copilot agents get a grounded, interactive runtime: every question is
grounded on the real files under `data/knowledge/` via bundled retrieval code, and citations come from retrieval, not the model.
- Analytics agents compute first and narrate second: bundled engines
produce the numbers in real code, and the LLM only explains them.
- The "AI drafting over spreadsheets" shape gets a real batch runtime that
reads the bundled `.xlsx` rows, calls an LLM to draft each reply, and writes drafts back for human review.
Each package also ships `agent/engines/` — real, stdlib-only implementations selected by the candidate shape (forecast baseline, schedule heuristic, significance testing, safe aggregation, record reconciliation, triage, retrieval) with selftests the bundled eval harness actually executes — and `agent/guards.py`, which enforces the operating model in code: a file-based kill switch, a per-run action budget, and an audit log of executed writes.
It's a scaffold, not a live deployment: a technical owner still supplies real credentials and an LLM key, and verifies any remaining mock or best-effort adapters before running this unattended in production.
Delivery support vs deployment stage. Each candidate also carries a delivery-support value (what UCF can deliver) and a delivery route (how to pursue it). That is separate from the evidence-backed deployment stage on the same card. A UCF native profile means a production profile exists — not that the package is production-ready. Download deployable scaffold (code) is offered only for the `build_with_ucf` route; otherwise the card shows the next action (collect prerequisites, commission an engine, use a partner, defer, or reject). High/Medium/Low tier tags are scaffold-confidence labels, not delivery support.
Reading the results
Pattern coverage
On Review, after Analyze. Shows N/101 supported. Green chips = patterns already grounded in reviewed graph facts. Expand Not supported yet for section-level evidence guidance on what to strengthen next (unsupported plays are listed as a distinct count; each evidence area then shows how many of those plays cite that area — the same play can appear in several areas, so those counts do not add up to the unsupported total).
Aim for grounded patterns that match the company's actual leverage — not every box ticked.
Candidate buckets
The review board groups candidates deterministically:
| Bucket | Meaning |
|---|---|
| Prerequisites | Do first — instrumentation, guardrails, test coverage, simulation, synthetic data, and other readiness work |
| Quick wins | High feasibility + meaningful impact |
| Strategic bets | High moat, plausibly feasible |
| Needs more evidence | Medium quality — gaps still hurt ranking |
| Dropped as generic | Failed quality gate |
Each card also shows Evidence that would strengthen this. Use Add in builder to jump to the section that would improve the candidate, or use the candidate-specific interview question when the missing evidence needs a human answer. After Develop use cases, cards are labeled Full plan generated or Analyze only so you can see which selected candidates received full artifacts.
Evidence & provenance map
From the builder or Review, choose Open evidence map or Download evidence map to create a self-contained HTML exhibit. It connects the reviewed company graph to source classes, cited financial metrics, open evidence gaps, and the ranked opportunities supported by those facts. The file works offline and does not require a Foundry login.
This is a proof-of-process export, not a full audit log: raw evidence, review history, outside-in signal ledgers, and complete candidate lineage are not embedded.
AI Opportunity Review (PDF)
On Review, Download PDF builds a customer-safe AI Opportunity Review. The server constructs a bounded report from the ranked board and a cited company digest, then renders PDF with Chromium. Sendability checks must pass. Export requires active Sprint or Custom access for the current workspace and fails closed rather than falling back to browser print.
The digest is the only set of company facts and figures the report may assert. It is a projection of the reviewed graph plus evidence, not a second crawl.
Developed use-case artifacts
Each generated use case may include:
- why_this_company — defensibility tied to graph facts
- first_experiment — smallest shippable pilot
- required_evidence — what to validate before scaling
- what_would_make_us_drop_this — kill criteria
With Opportunity dossiers enabled, selected quick wins and the strategic bet package into Markdown one-pagers you can copy to slides or memos.
Target Selector plan
The final plan surfaces:
- Quick wins — fund momentum (high impact + feasibility)
- One strategic bet — highest moat that is plausibly feasible
- Sequencing note — prerequisites honored; overlaps flagged
Quick wins and the strategic bet are never collapsed into one recommendation.
Roles in a multi-person program
| Contributor | Contribution | Tool surface |
|---|---|---|
| Program lead | Archetype, focus areas, when to Develop use cases | Source + Evidence + Review |
| Domain SME (finance, ops, …) | Processes, pains, data access | Evidence (department lens + builder) |
| Technical lead | Assets, constraints, stack | Evidence → Assets + Constraints |
| Executive sponsor | Risk appetite, budget direction | Source (Company context) + Review → dossiers |
| Consultant | YAML export, cross-session continuity | Evidence YAML tab + Source saved models |
Assign one department per session. Diagnostics = the sprint backlog — do not ask contributors to complete the whole form before they see value.
More workflow detail: the consultant interview guide
Authentication and saved models
When the server has `AUTH_SECRET` set:
- Sign in / register to use Analyze, Develop use cases, Auto-map, and saved models
- Saved models persist across devices, private to your account
- Pre-assessment starters are listed separately from normal company engagements until you promote them into the validation-workshop flow
- Marketing pages and the Builder guide remain public
When auth is disabled (typical local dev), all features work without login but models live in browser storage only.
Privacy and data handling
- Pasted notes, uploads, and evidence snippets are redacted before LLM calls
- Raw uploads are never persisted on the server
- Public website crawl content is treated as public (not redacted)
- Only human-approved graph facts enter the company model
- Workstation capture is opt-in; raw activity stays on the employee's
machine and only a redacted summary (or reviewable patch proposals) reaches Foundry
Full policy: Trust and evidence boundary
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| All candidates feel generic | Thin model — no linked pains, no accessible data | Close top gaps; add pains to high-volume processes |
| Large Not supported yet count | Missing edges or section facts | Expand the section for evidence guidance; see fill order |
| Auto-map returned sparse model | Public sources lack operational detail | Merge internal docs; use Map from evidence |
| Develop use cases fails or hangs | LLM not configured or timeout | Check `/api/status`; verify provider setup with your technical owner |
| Develop use cases says nothing is selected | All candidate checkboxes are cleared | Pick specific cards, or use Top 5 / All in the selection bar |
| Workstation capture shows no ActivityWatch events | Permissions not granted, or app not running | Confirm ActivityWatch is running and (macOS) Accessibility/Screen Recording permissions are granted, then retry |
| Rankings shift when I edit | Expected — Analyze re-runs on every edit | Use risk appetite / slider to stabilize tilt |
| Lost work after refresh | Not saved | Save to server or export YAML |
| YAML edits not reflected | Switched tabs without applying | Return to YAML tab and ensure content is saved |
| Pain Discovery stays on Working… after you pick a company | Stale studio page, or the load banner was not replaced | Hard-refresh `/pain-discovery`; a successful load should read Loaded {name} |
| Promoted pain missing in the builder | Older tab draft hid the new saved version | Open Builder from the studio (or `/app?model_id=…`), not the previous builder tab |
| Technique Workflow dropdown is empty | Saved model has no processes | Add named processes in the builder, save, then reload the studio |
Example companies
The app ships contrasting examples you can load to learn the model:
| Example | Profile |
|---|---|
| Meridian Fab | Custom fabrication; physical ops, scarce estimator judgment |
| Lumen Consulting | Data/analytics consultancy; services, regulated clients |
| Beacon Pay | Regulated fintech; data-rich, risk-averse |
| Atlas Ledger | Finance back office; month-end close, reconciliation |
| Harbor Link 3PL | Logistics; supplier delays, stockouts |
| Northstar People Ops | HR; recruiting funnel and onboarding bottlenecks |
| Summit Health System | Healthcare ops; prior-auth, denials, discharge |
| Cascade Support | B2B SaaS support; ticket triage, SLA recovery |
| Lumen Sales Ops | RevOps; pipeline hygiene, deal desk, RFPs |
| Nova Growth | Marketing ops; campaign lifecycle, attribution |
| Quay Commerce | Marketplace; matching through payout, fraud, liquidity |
| Northline Product | Product & portfolio; discovery through SKU rationalization |
| Helix Strategy | Strategy & capital; competing initiatives and KPI steering |
| Prism Data AI | Data & AI governance; quality, access, eval/drift |
| Pact Legal Ops | Legal; matter intake, obligations, discovery, IP |
| Aegis Security Ops | Cyber/ITSM; SOC, identity, change, problem |
| Keel Treasury | FP&A, cash, hedges, tax, capital appraisal |
| Ridge Workforce | Workforce planning, mobility, retention, labor controls |
| Cedar Facilities | Facilities; portfolio, energy, EHS, emissions, renewal |
| Atlas Ledger Production Demo | Bundled demo — AP invoice matching, pre-cleared through production handoff |
| Meera Home & Kitchen | Bundled demo — spreadsheet-native support triage with a real LLM-backed runtime |
| Beacon Pay Production Demo | Bundled demo — regulated chargeback evidence drafting with adjudicator approval gates |
Load one, change two fields, and watch the evidence preview react before modeling your own company. The three bundled deployment demos above are pre-graded: Develop use cases returns a finished, gate-cleared output instantly instead of calling an LLM.
Further reading
| Document | Audience | Contents |
|---|---|---|
| Builder guide | Anyone using the builder | Field reference, fill order, interview sprint, Analyze vs Plan |
| operator catalog | Anyone reading the board | Full 101-operator catalog |
| Plan adoption guide | Operators running Plan adoption | Foundations, routes, GraphFoundry contract |
| [SYSTEM_DATA_GRAPH.md](SYSTEM_DATA_GRAPH.md) | Operators adding execution topology | System-data sidecar model, API, and query workflow |
| AI Opportunity Review guide | Anyone sending a customer PDF | Cited digest + Opportunity Review rules |
| product glossary | All users | Canonical product vocabulary |
| methodology | Technical owners | Engine, API, scoring, module map |
| Pain Discovery Studio | Consultants | Hypothesis-led pain elicitation without a blank-sheet inventory |
| deployment resources | Operators | Hosting, env vars, auth |
Quick reference card
```
- Source → example | auto-map | optional Company context
- Evidence → add facts → stronger coverage → pains linked | Outside-in
- Discover → add known pain | run Pain Discovery on a mapped workflow
- Review → approve findings → Analyze → close top 3 gaps → repeat
- Plan → Develop use cases and/or Plan adoption
- Topology → Open system-data graph (optional) for execution/readiness analysis
- Export → Opportunity Review PDF | evidence map | YAML
- More input → capture | CX (after first Analyze)
- Save → server or YAML between sessions
```
Loop: `map workflows → add or discover pain → review evidence → analyze → close gaps → analyze → … → plan`
Remember: missing support = roadmap. Generic candidates = thin evidence. Defensible bets appear when the graph reflects how the business actually operates.