Playbook / Principal search path / Signal artifact — one deep proof

Signal artifact — one deep proof

Rule: ship one lab-grade artifact while I’m on the bridge. Not a new product brand. Not five half-demos.

Interviewers at Anthropic / OpenAI / FAANG AI Platform have seen portfolio sprawl. “Seventeen live platforms” reads as breadth without judgment. What converts is depth, honesty about limitations, and production taste — especially around evals, isolation, and irreversible tools.

I can start drafting this in the war room. I ship it when I have a seat story or a clean docs-first proof. I don’t wait for chapter 4 to have a link.


Choose exactly one primary

OptionWhat it provesEffortPrefer when
A. Eval methodology write-upCI regression gates, false confidence, suite kindsDocs + screenshots of golden-eval CILab / reliability / Applied panels
B. NIST AI RMF one-pagerGovern/Map/Measure/Manage mapped to live controlsDocs + links to AegisAI/RAG/evalsEnterprise / risk / CoE panels
C. FDE wedge case studyDiscovery → scored wedge → HITL → handoffMarkdown + metrics I can defendApplied / FDE / forward-deployed seats
D. Thin multi-tenant Strict pathRunning negative tests / JWT Principal storySmall code + honesty labelsHM already grilled isolation

Default if unsure: A or B. Docs-first. Fastest. Lowest risk of looking like I rebuilt a product for the interview.

If a Rivian/VW-style HM already stressed tenancy, bias to D — still one artifact, not a SaaS rewrite.


Definition of done

  • One public URL (portfolio ADR/case study, Substack essay with proof links, or GitHub doc)
  • Links to ≤2 live systems (AegisAI and/or Enterprise RAG) — not the full catalog
  • Explicit limitation section (Demo vs Strict, free-tier, what I’d harden in week 1–4 on their stack)
  • Usable in a 90-second outreach sentence
  • Does not invent a new Forge product name
  • I can defend every metric and every diagram without notes

Tone

AvoidPrefer
“This comprehensive framework leverages…”“Here’s the control I actually run, and where it still lies.”
Perfect diagrams with no scarsOne incident or near-miss that changed the design
Claiming Lucid ran the public binaries“Same principles. Lucid is P. This is O.”
Vanity latency tables on free-tier demos“I won’t invent P95s — here’s what the eval gate catches instead.”

A limitations section is not humility theater. It’s how I show I won’t ship unsafe confidence into someone else’s production.


Outlines (steal structure, write in my voice)

A — Eval methodology

  1. What “good” means for one agent/RAG path (task, risk, user)
  2. Suite kinds I use and what each catches
  3. Where false confidence shows up (fluent wrong answers, retrieval miss, tool misuse)
  4. How CI fails the build — screenshot or log redaction
  5. What I still don’t measure
  6. What I’d change in their stack in week one

2-minute walk I’d actually say (once the URL exists):

Here’s one RAG or agent path, what “good” means, and the three ways it lies — fluent wrong, retrieval miss, unauthorized tool. The suite is wired so a regression fails CI, not a dashboard. Limitation I will name first: I won’t invent P95s off a free-tier demo, and the public spine is O, not Lucid’s binary. Week one on your stack I’d port the deny-path fixtures, not the UI.

B — NIST mapped to live controls

One page. Four boxes. Each box names a running control (gateway policy, authZ-before-retrieve, eval gate, audit). No posterware. End with gaps.

C — FDE wedge

Discovery notes (sanitized) → scored options → why this wedge → HITL boundary → handoff criteria → what I’d refuse to automate.

D — Multi-tenant Strict

Threat model → layering → negative test that must fail open access → Demo vs Strict labels → what a real enterprise SSO path still needs.


What I will not build

  • New VoiceForge / pattern repo / teaching stub
  • Always-on expensive cloud for vanity SLOs
  • Multi-tenant SaaS rewrite across all org repos
  • A third personal brand website
  • “Seventeen platforms, updated for north-star”

If bandwidth remains after the primary artifact: harden one Strict/JWT or negative-eval path only — still under option D, not a new brand.


Storytelling constraint

Deepen only two systems in verbal answers:

  1. AegisAI — governance / HITL / audit / fail-closed
  2. Enterprise RAG — access-before-ranking / decline / eval gates

VAP and Content Factory are supporting cast.

Reuse: isolation · deep path · STAR pack


Scripts

90-second outreach

I keep the public catalog narrow on purpose. The deep proof I’d point you at is [A/B/C/D title] — it shows [evals that fail builds / NIST mapped to live controls / wedge→HITL / tenant deny-path]. Happy to walk limitations first. That’s usually the useful part.

“Tell me about your portfolio”

I’ll skip the inventory. Two systems matter for this conversation: a governance control plane for agent side effects, and an access-aware RAG path with eval gates. Lucid is where I shipped production outcomes. The public spine is inspectable proof of the same principles — not a claim those exact binaries ran inside Lucid. Here’s the one artifact that goes deep: [link].

“Why not more platforms / why so few demos?”

Breadth without a deny path is a content farm. I’d rather be boring and deep on governance and RAG than impressive and shallow on seventeen URLs.

“Walk NIST / responsible AI.”

I don’t recite Govern-Map-Measure-Manage as a poster. Map is where retrieval sits relative to authZ. Measure is an eval that can fail CI. Manage is HITL on the scary writes. Govern is who can change policy. Here’s what I still don’t cover: [limitation].

“How do you know your agent is working?”

Fluency isn’t the metric. I want suites that catch false confidence — retrieval misses, unauthorized tool calls, regressions against golden tasks — and I want at least one gate that can fail CI. If we only have a dashboard nobody pages on, we don’t have quality. We have vibes.


Substack (optional, capped)

At most one essay / week. Each essay ends with a CTA to technical-review and maps to one capability-map row. If charter delivery slips, pause writing. Writing is not a substitute for the signal artifact.