Principal leadership — budget fights and FinOps
Expected question
"GPU and LLM spend is exploding. How do you fight for budget — and how do you cut without killing the product?"
Variant forms
- "Walk me through a FinOps conversation with a CFO who only sees the invoice."
- "How do you allocate shared inference cost across product lines?"
- "When do you say no to a model upgrade that marketing wants?"
- "Build vs buy a gateway / eval platform under a cost freeze."
Executive summary
30-second thesis
I'd meter real tokens and GPU-hours before I argue. Then I'd cut waste with quotas and routing — not with a blanket "use the cheap model" that nukes quality on the one path that makes money.
2-minute answer
Open with measured cost by tenant/product/model — guessed dashboards lose CFO trust forever. Propose three levers: (1) admission/quotas so runaways can't eat the fleet, (2) router to cheaper models where eval slices still pass, (3) kill shadow/canary spend that isn't gated.
What I'd fund anyway: eval gates and the side-effect gateway. Those look like cost centers until the first incident. What I'd defer: vanity multi-region training, unowned 10k eval suites, denser packing before isolation proofs.
Scar: optimizing on fake token costs. If metering isn't real, every savings claim is theater.
Spoken CFO frame (~60s)
Here's where the money goes by product and model. Here's the quality gate that blocks a cheaper route. I can save X by quotas and routing without touching the revenue path; I will not save Y by removing HITL on refunds. Here's the kill switch if we blow the monthly ceiling.
Skeptical follow-ups
- "Just use the smallest model." Show the slice that breaks; offer a routed compromise.
- "Why is platform spend so high?" Amortize incidents prevented and duplicated stacks not built — with numbers you can defend.
Staff+/Principal signal rubric
- Senior: knows cost knobs.
- Staff+: ties cuts to quality gates and quotas.
- Principal: runs the exec conversation with metering honesty and explicit refuse lines.