What these skills are
An AI feature, from the first idea to the weekly review after launch, as installable Claude skills.
This is AI product management that starts from the mistakes the feature can make, not the demo: Claude drafts the spec, the evals and the briefs, and a person who knows the domain makes the call.
The skills follow the job in the order it happens. You ask whether a model is needed at all and which mistakes would matter, write the spec as tests, brief the model and the agent, build the evals out of real outputs, ship it through gates that can stop it, keep it working once it is live, and end with what it costs, what it could charge and how to explain it to leaders.
The red line: Claude drafts the tests, never the verdict. It does not set the error your users can live with, sign off a launch, or report a score it did not see run on real outputs, and a person who knows the domain reads real transcripts every week.
Each skill is one named method or one artifact you walk away with: a use case canvas, a failure modes map, a behavior contract written as Given-When-Then cases, an agent spec with a spend budget, a golden dataset of real cases, a judge you can show agrees with your experts, launch gates with kill criteria, a cost per successful outcome. Each works from what you paste — the request for "an AI feature", the system prompt as it stands, last week's transcripts, the token counts from thirty real tasks — with no eval platform to buy, and each one ends with the decision a named person makes, by a date. Legal, privacy and regulatory points become questions for a qualified adviser rather than answers.
Mechanically, each skill is one folder with a SKILL.md file. The 36 are grouped into seven stages of the work, and each one gives full value on its own: run one, run a stage, or work through the lot in the order the job moves. The sources behind each method are in resources/evidence-and-sources.md.
Download all 36 skills. One zip: ready-to-install skill zips, readable SKILL.md files, and the sources behind every method. Free, no signup. The same 36 skills are open on GitHub: polar-bear-org/claude-skills.
The 36 skills you get
The set follows the work, in seven stages: deciding if AI fits, speccing it, briefing the model and the agent, building the evals, shipping it safely, running and improving it, then cost, price and explaining it.
1 · Decide if AI fits
1. AI Use Case Canvas
Use when: Someone wants "an AI feature" and nobody has said what decision it improves
Output: The canvas — prediction, judgment, action, outcome, input, feedback — a simpler-fix check, a data check, and a verdict
2. AI Failure Modes Map
Use when: Leaders expect it to be right every time and you need to show which mistakes matter
Output: A failure list per output, the cost of a false yes against a false no, severity and detection, a graceful failure, and the tolerable rate left blank for you
3. Workflow or Agent Decision
Use when: The team wants an agent and nobody asked whether a fixed workflow would do
Output: The candidate patterns, their cost, latency and failure trade-offs, the simplest pattern that passes, and the trigger for the next step up
4. AI Model Selection
Use when: You must justify the model tier and the bill before engineering commits
Output: A side-by-side test plan on real inputs, a quality, latency and cost table per tier, a pick per task, and a re-check date
5. AI Prototype Brief
Use when: Leadership saw a slick AI demo and thinks production is just more prompts
Output: What the demo tests, a pass line set in advance, a cherry-picked input check, and the gap-to-production list
2 · Spec it
6. AI PRD
Use when: The AI feature spec reads like a normal PRD and nobody can test it
Output: The problem, the inputs the model sees, the outputs, a behavior summary, fallbacks, data needs, the release bar, and the open questions
7. AI Success Criteria
Use when: The team says "make it good" and nobody can say what good means
Output: Measurable criteria per dimension, today's baseline, a target set by a named person, and how each one is measured
8. AI Behavior Contract
Use when: "It should be helpful and safe" is the whole spec
Output: Must, must-never and when-unsure lines written as Given-When-Then cases, a pass rate over repeated runs, and the guardrails
9. Human-in-the-Loop Design
Use when: A wrong move is costly and nobody has drawn where a person steps in
Output: The approve, edit and take-over points, their triggers, the handoff message, a response time a person can actually meet, and the log
10. AI Data Privacy Brief
Use when: Legal asks what user data reaches the model and nobody has drawn it
Output: The data flow from user to model to logs, the personal data in prompts, retention, region, and the questions for legal
3 · Brief the model and the agent
11. System Prompt Brief
Use when: The prompt grew by patches and nobody knows which line does what
Output: Role, audience, task, rules, examples, output format, refusals and handoffs, plus a critique of the prompt you have now
12. Context Engineering Brief
Use when: Answers are wrong because the model read stale or irrelevant documents
Output: The knowledge sources and their owners, freshness rules, what to leave out, retrieval spot checks, and a context budget
13. Agent Spec
Use when: The agent can act on its own and nobody has written where it must stop
Output: The goal, the allowed tools, permissions, stop conditions, a spend budget, the approval points, and the never-alone list
14. Tool Descriptions
Use when: The agent picks the wrong tool or misreads what a tool returns
Output: Per tool: name, purpose, inputs, outputs and error messages, plus namespacing, an overlap check, and three test tasks
4 · Build the evals
15. Error Analysis
Use when: Quality is judged by vibes and the dashboard tracks scores that match no real problem
Output: Notes on real outputs, open codes, failure types grouped, counts per type, and the three to fix first
16. Golden Dataset
Use when: Nobody can tell if a new prompt helped the real workflow or only the demo
Output: 20 to 50 real cases with their expected outcomes, coverage by failure type, the source and version, and add-and-retire rules
17. Synthetic Test Data
Use when: Real cases are too thin or too sensitive to cover the edge cases
Output: Generated cases for coverage gaps only, each marked synthetic, checked by a person, and kept out of the headline score
18. Eval Rubric
Use when: Scores on a 1 to 5 scale move and nobody knows what changed
Output: One pass-or-fail check per failure type, a grader per check, a pass and a fail example, and the blockers
19. LLM-as-a-Judge Prompt
Use when: You want automated grading but cannot prove the judge agrees with your experts
Output: A judge prompt for one failure type, a labelled set, agreement measured on a held-out split, and bias checks
20. AI Eval Plan
Use when: Everyone agrees evals matter and nobody owns them
Output: The capability and regression suites, when each one runs, what blocks a release, and the named owner of quality
5 · Ship it safely
21. AI Red Teaming Plan
Use when: An agent will read emails, tickets or web pages written by strangers
Output: Attack cases by risk class, who runs them, the pass line, and the fixes due before launch
22. AI UX Review
Use when: Users rephrase, give up or leave, and the design never said what happens when it is wrong
Output: 18 interaction guidelines checked by phase, an AI disclosure check, the correction and handoff paths, and the fixes ranked
23. AI Impact Assessment
Use when: A customer, a regulator or your legal team asks for an impact assessment
Output: Who is affected, intended use and misuse, the data, oversight, mitigations, the residual risk owner, and the questions for an adviser
24. AI Risk Register
Use when: Legal and security ask for the risks and you have a list of worries
Output: Cause, event and effect risks mapped to Map, Measure and Manage, each with an owner, a trigger and a response
25. Questions for Legal
Use when: You need legal sign-off and do not know what to ask or what to bring
Output: The questions for the legal and privacy team by topic, with the facts attached — it asks, never answers
26. AI Launch Checklist
Use when: Launch is close, the first gate has not opened, and nobody has checked the boring things
Output: A go-live list with an owner per line, from evals passed as run to the rollback rehearsed
27. AI Launch Gates
Use when: The feature is ready to go out and nobody has written what would stop it
Output: The stages, the eval and live numbers per gate, the kill criteria, the rollback path, and who signs each gate
6 · Run and improve
28. Weekly AI Quality Review
Use when: It worked for weeks, then broke, and nobody had been reading the transcripts
Output: A transcript sample plan, the review notes, the new failure types, the cases added to the golden dataset, and the decisions
29. AI Feedback Signals
Use when: Users rarely rate answers and you cannot see what they think
Output: A signal map — accept, edit, retry, rephrase, abandon, escalate, rating — what to log, a weekly readout, and the review triggers
30. Prompt Regression Test
Use when: A prompt edit fixed one case and quietly broke another
Output: The change note, the cases to re-run, before and after per failure type, and a ship-or-hold for a named person
31. Model Migration Plan
Use when: The model you depend on has a retirement date
Output: The deadline, the breaking changes, the eval rerun, the cost difference, the behavior diffs to read, and the rollout and fallback
32. AI Incident Response Plan
Use when: The AI told a customer something wrong and it is spreading
Output: Severity levels, containment, the customer correction, an evidence timeline, the fix carried into the golden dataset, and a blameless review
7 · Cost, price and explain
33. AI Unit Economics
Use when: The feature passed every eval and finance still wants to stop it
Output: Cost per call, per task and per successful outcome, the heavy-user case, the caching and batch levers, and what it costs today
34. AI Usage and Pricing Test
Use when: You need run cost and price options before you have real usage
Output: A small real-task test, a usage estimate at light, normal and heavy use, the price options to test, and what to re-measure
35. AI Feature Card
Use when: Legal, sales and support each ask what the feature does and where it fails
Output: Intended and out-of-scope use, the data, the eval results as run with dates, the known limits, the disclosures, and the owner
36. AI Exec Brief
Use when: Leaders ask "how accurate is it" and one number would mislead them
Output: One page, bottom line first: how often it is wrong by severity, what that costs, the run cost, and the decision being asked for
Where to start, and how the skills chain
There is no compulsory sequence. You bring one live piece of work — the request for an AI assistant that drafts replies, the system prompt that grew by patches, sixty transcripts from last week, the retirement notice for the model you depend on — and open the skill that matches where it is stuck. The order in the catalogue is the order the work moves in: whether a model is needed at all first, because a feature nobody could justify is what makes every later step wasted effort, then which mistakes would matter, which is what decides what you measure.
Most of them chain. aipm-ai-use-case-canvas says whether a model improves a decision, aipm-ai-failure-modes names the mistakes that would cost you, and aipm-workflow-or-agent and aipm-ai-model-selection settle the pattern and the tier. aipm-ai-prd, aipm-ai-success-criteria and aipm-behavior-contract turn that into a spec you can test, with aipm-human-in-the-loop drawing where a person steps in. aipm-system-prompt, aipm-context-engineering, aipm-agent-spec and aipm-tool-descriptions brief the model and the agent. The evals run from aipm-error-analysis, which produces the failure types that aipm-golden-dataset, aipm-eval-rubric and aipm-llm-judge are built around, and aipm-eval-plan gives them an owner. aipm-red-team-plan, aipm-ai-ux-review and aipm-launch-gates take it out, aipm-quality-review, aipm-feedback-signals and aipm-prompt-regression-test keep it honest, and aipm-unit-economics, aipm-feature-card and aipm-exec-brief answer finance, legal and leadership.
Setup guide
- Download the pack. One zip: an install folder with 36 ready-to-upload skill zips, a skills folder with the same 36 skills as readable SKILL.md files, and resources/evidence-and-sources.md, which names the source behind every method and where it stops.
- Install your skills. In Claude Code, add the marketplace and install ai-product-managers-pack, and all 36 load at once. In Claude, turn on code execution in Settings, Capabilities, then go to Customize, Skills and upload one zip per skill from the install folder. Prefer working from files? Add the SKILL.md files to your Project knowledge instead; it works, just less cleanly.
- Bring ten real outputs. No connector, no eval platform, no admin: every skill works from what you paste — the request for "an AI feature", the system prompt as it stands, sixty transcripts from last week, the token counts from thirty real tasks. Start with the feature on your desk and write "run aipm-ai-use-case-canvas". Where you have little, each skill starts from the minimum and marks the result as a first draft.
Where to start
| Your situation | Skill to run |
|---|---|
| Need to know if a model is needed at all? | AI Use Case Canvas |
| Need to show which mistakes matter? | AI Failure Modes Map |
| Need to choose between a workflow and an agent? | Workflow or Agent Decision |
| Need a demo not mistaken for the product? | AI Prototype Brief |
| Need a spec engineers can test? | AI PRD |
| Need "good" written as numbers? | AI Success Criteria |
| Need limits an agent cannot cross? | Agent Spec |
| Need failure types from real outputs? | Error Analysis |
| Need a test set built from real cases? | Golden Dataset |
| Need to attack it before strangers do? | AI Red Teaming Plan |
| Need to know a prompt edit broke nothing? | Prompt Regression Test |
| Need cost per successful outcome? | AI Unit Economics |
The quality bar
Every skill in the pack holds the same standard, the one we hold when we ship AI in our own product:
- One method, applied properly: each skill uses the real mechanics of its source — the canvas fields, the pattern list, the coding passes, the capability and regression split, the risk classes — not a generic "gather, analyse, recommend"
- Real outputs first: failure types, golden cases and judge checks come from outputs you paste, and synthetic cases fill gaps, are marked synthetic, and never carry the headline score
- No invented numbers: no made-up scores, prices, error rates or usage, and an eval result appears only when you pasted the run behind it
- Errors by severity and cost per successful outcome, never a single accuracy figure on its own or a cost per call without the retries and failures behind it
- Work is scored, people are not: skills score outputs, cases, risks and options, never reviewers, users or team members
- Legal, privacy and regulatory points go to a qualified adviser — the skills ask the questions and never answer them
- The red line: Claude drafts the tests, never the verdict — it does not set the error your users can live with, sign off a launch, or report a score it did not see run on real outputs, and a person who knows the domain reads real transcripts every week
Who made this
Polar Bear is a people ops consultancy for human-size teams (20 to 200 people). Built by ex-McKinsey founders with a dream to make AI work for People, not instead of them. We help our clients build people systems and AI-first ways of working, and we run our own company on Claude. This pack is the free, self-serve version of how we work.
The pack carries one AI product manager's work. When you want your whole team working this way, AI carrying the overhead so people do the part only people can do, across hiring, management and everyday operations, that's what we build with clients.
Meet Pauline. A demo that wowed and a feature nobody can grade? Bring ten real outputs and the launch you're weighing — we'll see what the evals say. Book a 30-minute call · Pauline on LinkedIn
Install in Claude Code
Two lines, and every skill in the pack loads at once.
/plugin marketplace add polar-bear-org/claude-skills /plugin install ai-product-managers-pack@polar-bear-skills
Using Claude on the web instead? Download the zip and upload each skill from its install folder under Customize → Skills.