Playbook / FDE / Applied deploy / Enterprise discovery call — "VP wants a chatbot"

Enterprise discovery call — "VP wants a chatbot"

Expected question

"Role-play: I'm a VP who wants 'a chatbot' / Claude copilot for our company. Run the discovery call — do not open an architecture deck."

Variant forms

  • "I'm VP Engineering at a 5,000-person fintech evaluating your model for an internal copilot. Convince me."
  • "Honestly, I'm not sure your product does anything ChatGPT can't. Run discovery, don't pitch."
  • "Explain embeddings / RAG to our Chief Legal Officer in two minutes — brilliant, busy, non-technical."
  • "A customer demands a feature that won't solve their actual problem. What do you do?"
  • "Turn a skeptical stakeholder into a champion."
  • "Tell me about the most ambiguous project you've owned — what did you do in week one?"
  • "How do you learn an unfamiliar domain fast, and how do you validate you understood it?"
  • "Estimate LLM tokens/day for a Fortune-500 support org — then convert to dollars."

Where this actually gets asked

Anthropic Applied AI customer simulations often grade this harder than coding. Also OpenAI FDE and enterprise SE→FDE loops. Pitching in the first 90 seconds fails.

The question, as it might actually be asked

"Don't pitch. Diagnose. The word 'chatbot' is usually wrong."

The framework

30-second thesis

I'd replace “chatbot” with a job-to-be-done, a 90-day metric, constraints we can't change, and an irreversible-action boundary — then reflect that back as a wedge and non-goals before any architecture. If I open a deck in minute one, I've already lost.

2-minute method

JTBD discovery, not demo. How the first fifteen minutes sound:

Me: “Before we talk product — what workflow is broken today? Who's waiting on what decision?”

Me: “What have you already tried with AI, and why did it die?”

Me: “In ninety days, what number means this worked — and what's the baseline window?”

Me: “What can't we change — SSO, residency, change windows, legal, data that can't leave the VPC?”

Me: “What must never be autonomous on day one?”

Then reflect back before proposing anything:

So the real problem is X for role Y under constraint Z — not a generic chatbot. Closest wedge is A; non-goal is B. Did I get that wrong?

Only after they correct me do I propose verbally: read-only assist over an approved corpus with citations, SSO + ACL-aware retrieval, HITL on any write, weekly eval with a named owner. Architecture deep-dive later — not now.

Question sequence (first 15 minutes)

  1. Job: What workflow fails today — who waits, what decision?
  2. Failed AI history: What did you try? Why did it die?
  3. Success definition: What metric in 90 days means this worked?
  4. Constraints you can't change: SSO, residency, change windows, unions/legal, data that cannot leave the VPC.
  5. Irreversible actions: What must never be autonomous on day one?
  6. Systems of record: Where is truth written today?
  7. Users vs buyers: Who uses it daily vs who signs?
  8. Eval ownership: Who will judge quality weekly?

Two-minute RAG for Legal / Ops VP

It's an open-book exam over your approved documents, with citations — not a free-form genius. If the book doesn't contain the answer, a good system says it doesn't know.

Then stop. One risk (hallucination) and one control (decline + human review). Don't keep talking.

Fermi / tokens → dollars

Users × sessions × turns × tokens/turn × $/1M tokens; sanity-check against seat count and ticket volume; separate retrieval vs generation. Label estimates H. Invite their ticket volume — don't perform precision theater.

Requirements (call success criteria)

Functional

  • Restated problem in customer language.
  • Named wedge + non-goals.
  • Named sponsor metric and eval owner.
  • Explicit day-one autonomy boundary.

Non-functional

  • No architecture dump before reflect-back.
  • Honest competitive framing (ChatGPT vs grounded enterprise path).
  • Optional Fermi cost sanity check when asked.

Core entities / actors

  • Buyer VP: budget and politics.
  • Daily user: workflow pain.
  • Skeptic (Legal/Security): harm and residency.
  • Failed prior AI: the ghost in the room — ask early.
  • Wedge: first production path that is not “a chatbot.”

Process flow — discovery call

Rendering architecture diagram…

High-level design (only after reflect-back)

Propose verbally:

  1. Read-only assist over approved corpus with citations for role Y.
  2. SSO + ACL-aware retrieval.
  3. HITL for any write into system of record.
  4. Weekly eval with named owner.
  5. Architecture deep-dive scheduled after wedge agreement — not now.

Deep dive 1: “ChatGPT already does this”

VP: “Honestly, how is this different from ChatGPT?”

Me: “For general Q&A, maybe it isn't. What ChatGPT doesn't have is your identity, your ACLs, your systems of record, your audit trail, your change windows, and your irreversible-action policy. Proof isn't a bake-off slide — it's shadow on one workflow for two weeks with groundedness and wrong-action metrics.”

Same brand: governed agents, access-aware RAG, HITL, evals — not a new pitch deck per logo.

Deep dive 2: learning an unfamiliar domain fast

Week one: shadow three operators; write the workflow in their nouns; teach it back to validate; pick the metric they already watch. I don't invent industry jargon to sound smart.

Deep dive 3: Fermi cost without losing the room

Walk the formula aloud; show sensitivity (turns/session dominate); separate retrieval vs generation; tie to seat count. Offer to refine with their ticket volume. Collaboration, not precision theater.

Quantitative trade-offs

DecisionTrade-off and reversal evidenceEvidence class
Discovery-first vs early demoDemo can excite; reverse when buyer already saw demos and needs diagnosis of failed AIH/R
Read-only wedge vs “full copilot” promiseNarrower story; reverse only when irreversible writes are in-scope, HITL-staffed, and metric demands cycle-timeH/P
Schedule architecture follow-up vs design liveLive design risks pitch mode; reverse if VP explicitly asks for trust boundaries nowH

Migration / next-week plan

  1. Confirm metric + weekly eval owner on calendar.
  2. Access for one corpus slice + SSO staging.
  3. Shadow plan beside humans.
  4. Written non-goals shared to Legal/Security early.
  5. Architecture session only after wedge lock.

Org ownership

  • VP buyer owns funding and air cover.
  • Ops lead owns workflow truth and HITL reviewers.
  • Security/Legal own residency and high-harm veto.
  • FDE owns discovery notes, reflect-back doc, and next-week critical path.

Situation

Discovery calls match Lucid stakeholder rooms (P): Commerce / Supply Chain ask for outcomes (exceptions, cycle time), not “a chatbot.” Interview role-play puts you across from a busy VP who may be skeptical that your product beats ChatGPT.

Task

Run JTBD discovery: replace chatbot with workflow, metric, and non-goals; earn the right to propose a wedge — without opening an architecture deck.

Action

  1. Ask job, failed-AI history, and 90-day metric before product claims.
  2. Inventory constraints and irreversible actions.
  3. Reflect back problem + wedge + non-goals; invite correction.
  4. Explain RAG to non-technical skeptics in ≤2 minutes with one risk and one control.
  5. If challenged on ChatGPT, differentiate on identity/ACL/audit/SoR and propose a shadow proof.
  6. Leave a next-week plan with owners — architecture later.

Control story after discovery (O): access-aware RAG + HITL gateway — not before.

Result

Diagnosis over pitch; wedge clarity; trust with skeptics. You leave with a sponsor metric and non-goals — not a vague “AI chatbot initiative.”

The follow-up question you should expect

"So what do we build first?"
Read-only assist for the highest-hour workflow with citations under SSO/ACLs; HITL on any write; golden set seeded from last month’s hard cases; expand only after the weekly eval owner signs the gate.

What I'd ask them

  1. What AI attempt already failed here — and why?
  2. Who uses this daily vs who signs the PO?
  3. What must never be autonomous on day one?
  4. Who will sit in the weekly eval review — name and calendar?

Candidate-owned evidence prompts

  1. Which Lucid workflow nouns will you use if they ask for a real discovery story?
  2. Have you practiced the 2-minute Legal RAG explanation out loud?
  3. What Fermi assumptions will you state as H?
  4. What is your single reflect-back sentence template?

Author reference (do not memorize)

Role-play scripts are scaffolds. Validate domain understanding with teach-back; don’t fake industry expertise.

Staff+/Principal signal rubric

  • Mid-level: Asks a few questions then pitches RAG.
  • Senior: Maps users, data, and a v1 use case.
  • Staff+: Failed-AI history, irreversible boundary, 90-day metric, reflect-back, wedge + non-goals.
  • Principal: Turns discovery into account plan + product feedback signal.