Polar Bear / Claude Skills / Claude for AI Product Managers Pack

October 2026 · 9 min read

Claude for AI Product Managers: 36 Claude Skills to Test an AI Feature Before It Ships

By Pauline Bertry, ex-McKinsey Manager

It starts before the model: does this need AI at all, and which mistakes would matter. Then the spec written as tests, the prompt and the agent's limits, evals built from real outputs, red teaming and launch gates, and the weekly transcript read that keeps it honest.

What these skills are

An AI feature, from the first idea to the weekly review after launch, as installable Claude skills.

This is AI product management that starts from the mistakes the feature can make, not the demo: Claude drafts the spec, the evals and the briefs, and a person who knows the domain makes the call.

The skills follow the job in the order it happens. You ask whether a model is needed at all and which mistakes would matter, write the spec as tests, brief the model and the agent, build the evals out of real outputs, ship it through gates that can stop it, keep it working once it is live, and end with what it costs, what it could charge and how to explain it to leaders.

The red line: Claude drafts the tests, never the verdict. It does not set the error your users can live with, sign off a launch, or report a score it did not see run on real outputs, and a person who knows the domain reads real transcripts every week.

Each skill is one named method or one artifact you walk away with: a use case canvas, a failure modes map, a behavior contract written as Given-When-Then cases, an agent spec with a spend budget, a golden dataset of real cases, a judge you can show agrees with your experts, launch gates with kill criteria, a cost per successful outcome. Each works from what you paste — the request for "an AI feature", the system prompt as it stands, last week's transcripts, the token counts from thirty real tasks — with no eval platform to buy, and each one ends with the decision a named person makes, by a date. Legal, privacy and regulatory points become questions for a qualified adviser rather than answers.

Mechanically, each skill is one folder with a SKILL.md file. The 36 are grouped into seven stages of the work, and each one gives full value on its own: run one, run a stage, or work through the lot in the order the job moves. The sources behind each method are in resources/evidence-and-sources.md.

Download all 36 skills. One zip: ready-to-install skill zips, readable SKILL.md files, and the sources behind every method. Free, no signup. The same 36 skills are open on GitHub: polar-bear-org/claude-skills.

The 36 skills you get

The set follows the work, in seven stages: deciding if AI fits, speccing it, briefing the model and the agent, building the evals, shipping it safely, running and improving it, then cost, price and explaining it.

1 · Decide if AI fits

1. AI Use Case Canvas

Use when: Someone wants "an AI feature" and nobody has said what decision it improves

Output: The canvas — prediction, judgment, action, outcome, input, feedback — a simpler-fix check, a data check, and a verdict

2. AI Failure Modes Map

Use when: Leaders expect it to be right every time and you need to show which mistakes matter

Output: A failure list per output, the cost of a false yes against a false no, severity and detection, a graceful failure, and the tolerable rate left blank for you

3. Workflow or Agent Decision

Use when: The team wants an agent and nobody asked whether a fixed workflow would do

Output: The candidate patterns, their cost, latency and failure trade-offs, the simplest pattern that passes, and the trigger for the next step up

4. AI Model Selection

Use when: You must justify the model tier and the bill before engineering commits

Output: A side-by-side test plan on real inputs, a quality, latency and cost table per tier, a pick per task, and a re-check date

5. AI Prototype Brief

Use when: Leadership saw a slick AI demo and thinks production is just more prompts

Output: What the demo tests, a pass line set in advance, a cherry-picked input check, and the gap-to-production list

2 · Spec it

6. AI PRD

Use when: The AI feature spec reads like a normal PRD and nobody can test it

Output: The problem, the inputs the model sees, the outputs, a behavior summary, fallbacks, data needs, the release bar, and the open questions

7. AI Success Criteria

Use when: The team says "make it good" and nobody can say what good means

Output: Measurable criteria per dimension, today's baseline, a target set by a named person, and how each one is measured

8. AI Behavior Contract

Use when: "It should be helpful and safe" is the whole spec

Output: Must, must-never and when-unsure lines written as Given-When-Then cases, a pass rate over repeated runs, and the guardrails

9. Human-in-the-Loop Design

Use when: A wrong move is costly and nobody has drawn where a person steps in

Output: The approve, edit and take-over points, their triggers, the handoff message, a response time a person can actually meet, and the log

10. AI Data Privacy Brief

Use when: Legal asks what user data reaches the model and nobody has drawn it

Output: The data flow from user to model to logs, the personal data in prompts, retention, region, and the questions for legal

3 · Brief the model and the agent

11. System Prompt Brief

Use when: The prompt grew by patches and nobody knows which line does what

Output: Role, audience, task, rules, examples, output format, refusals and handoffs, plus a critique of the prompt you have now

12. Context Engineering Brief

Use when: Answers are wrong because the model read stale or irrelevant documents

Output: The knowledge sources and their owners, freshness rules, what to leave out, retrieval spot checks, and a context budget

13. Agent Spec

Use when: The agent can act on its own and nobody has written where it must stop

Output: The goal, the allowed tools, permissions, stop conditions, a spend budget, the approval points, and the never-alone list

14. Tool Descriptions

Use when: The agent picks the wrong tool or misreads what a tool returns

Output: Per tool: name, purpose, inputs, outputs and error messages, plus namespacing, an overlap check, and three test tasks

4 · Build the evals

15. Error Analysis

Use when: Quality is judged by vibes and the dashboard tracks scores that match no real problem

Output: Notes on real outputs, open codes, failure types grouped, counts per type, and the three to fix first

16. Golden Dataset

Use when: Nobody can tell if a new prompt helped the real workflow or only the demo

Output: 20 to 50 real cases with their expected outcomes, coverage by failure type, the source and version, and add-and-retire rules

17. Synthetic Test Data

Use when: Real cases are too thin or too sensitive to cover the edge cases

Output: Generated cases for coverage gaps only, each marked synthetic, checked by a person, and kept out of the headline score

18. Eval Rubric

Use when: Scores on a 1 to 5 scale move and nobody knows what changed

Output: One pass-or-fail check per failure type, a grader per check, a pass and a fail example, and the blockers

19. LLM-as-a-Judge Prompt

Use when: You want automated grading but cannot prove the judge agrees with your experts

Output: A judge prompt for one failure type, a labelled set, agreement measured on a held-out split, and bias checks

20. AI Eval Plan

Use when: Everyone agrees evals matter and nobody owns them

Output: The capability and regression suites, when each one runs, what blocks a release, and the named owner of quality

5 · Ship it safely

21. AI Red Teaming Plan

Use when: An agent will read emails, tickets or web pages written by strangers

Output: Attack cases by risk class, who runs them, the pass line, and the fixes due before launch

22. AI UX Review

Use when: Users rephrase, give up or leave, and the design never said what happens when it is wrong

Output: 18 interaction guidelines checked by phase, an AI disclosure check, the correction and handoff paths, and the fixes ranked

23. AI Impact Assessment

Use when: A customer, a regulator or your legal team asks for an impact assessment

Output: Who is affected, intended use and misuse, the data, oversight, mitigations, the residual risk owner, and the questions for an adviser

24. AI Risk Register

Use when: Legal and security ask for the risks and you have a list of worries

Output: Cause, event and effect risks mapped to Map, Measure and Manage, each with an owner, a trigger and a response

25. Questions for Legal

Use when: You need legal sign-off and do not know what to ask or what to bring

Output: The questions for the legal and privacy team by topic, with the facts attached — it asks, never answers

26. AI Launch Checklist

Use when: Launch is close, the first gate has not opened, and nobody has checked the boring things

Output: A go-live list with an owner per line, from evals passed as run to the rollback rehearsed

27. AI Launch Gates

Use when: The feature is ready to go out and nobody has written what would stop it

Output: The stages, the eval and live numbers per gate, the kill criteria, the rollback path, and who signs each gate

6 · Run and improve

28. Weekly AI Quality Review

Use when: It worked for weeks, then broke, and nobody had been reading the transcripts

Output: A transcript sample plan, the review notes, the new failure types, the cases added to the golden dataset, and the decisions

29. AI Feedback Signals

Use when: Users rarely rate answers and you cannot see what they think

Output: A signal map — accept, edit, retry, rephrase, abandon, escalate, rating — what to log, a weekly readout, and the review triggers

30. Prompt Regression Test

Use when: A prompt edit fixed one case and quietly broke another

Output: The change note, the cases to re-run, before and after per failure type, and a ship-or-hold for a named person

31. Model Migration Plan

Use when: The model you depend on has a retirement date

Output: The deadline, the breaking changes, the eval rerun, the cost difference, the behavior diffs to read, and the rollout and fallback

32. AI Incident Response Plan

Use when: The AI told a customer something wrong and it is spreading

Output: Severity levels, containment, the customer correction, an evidence timeline, the fix carried into the golden dataset, and a blameless review

7 · Cost, price and explain

33. AI Unit Economics

Use when: The feature passed every eval and finance still wants to stop it

Output: Cost per call, per task and per successful outcome, the heavy-user case, the caching and batch levers, and what it costs today

34. AI Usage and Pricing Test

Use when: You need run cost and price options before you have real usage

Output: A small real-task test, a usage estimate at light, normal and heavy use, the price options to test, and what to re-measure

35. AI Feature Card

Use when: Legal, sales and support each ask what the feature does and where it fails

Output: Intended and out-of-scope use, the data, the eval results as run with dates, the known limits, the disclosures, and the owner

36. AI Exec Brief

Use when: Leaders ask "how accurate is it" and one number would mislead them

Output: One page, bottom line first: how often it is wrong by severity, what that costs, the run cost, and the decision being asked for

Where to start, and how the skills chain

There is no compulsory sequence. You bring one live piece of work — the request for an AI assistant that drafts replies, the system prompt that grew by patches, sixty transcripts from last week, the retirement notice for the model you depend on — and open the skill that matches where it is stuck. The order in the catalogue is the order the work moves in: whether a model is needed at all first, because a feature nobody could justify is what makes every later step wasted effort, then which mistakes would matter, which is what decides what you measure.

Most of them chain. aipm-ai-use-case-canvas says whether a model improves a decision, aipm-ai-failure-modes names the mistakes that would cost you, and aipm-workflow-or-agent and aipm-ai-model-selection settle the pattern and the tier. aipm-ai-prd, aipm-ai-success-criteria and aipm-behavior-contract turn that into a spec you can test, with aipm-human-in-the-loop drawing where a person steps in. aipm-system-prompt, aipm-context-engineering, aipm-agent-spec and aipm-tool-descriptions brief the model and the agent. The evals run from aipm-error-analysis, which produces the failure types that aipm-golden-dataset, aipm-eval-rubric and aipm-llm-judge are built around, and aipm-eval-plan gives them an owner. aipm-red-team-plan, aipm-ai-ux-review and aipm-launch-gates take it out, aipm-quality-review, aipm-feedback-signals and aipm-prompt-regression-test keep it honest, and aipm-unit-economics, aipm-feature-card and aipm-exec-brief answer finance, legal and leadership.

Setup guide

  1. Download the pack. One zip: an install folder with 36 ready-to-upload skill zips, a skills folder with the same 36 skills as readable SKILL.md files, and resources/evidence-and-sources.md, which names the source behind every method and where it stops.
  2. Install your skills. In Claude Code, add the marketplace and install ai-product-managers-pack, and all 36 load at once. In Claude, turn on code execution in Settings, Capabilities, then go to Customize, Skills and upload one zip per skill from the install folder. Prefer working from files? Add the SKILL.md files to your Project knowledge instead; it works, just less cleanly.
  3. Bring ten real outputs. No connector, no eval platform, no admin: every skill works from what you paste — the request for "an AI feature", the system prompt as it stands, sixty transcripts from last week, the token counts from thirty real tasks. Start with the feature on your desk and write "run aipm-ai-use-case-canvas". Where you have little, each skill starts from the minimum and marks the result as a first draft.

Where to start

Your situationSkill to run
Need to know if a model is needed at all?AI Use Case Canvas
Need to show which mistakes matter?AI Failure Modes Map
Need to choose between a workflow and an agent?Workflow or Agent Decision
Need a demo not mistaken for the product?AI Prototype Brief
Need a spec engineers can test?AI PRD
Need "good" written as numbers?AI Success Criteria
Need limits an agent cannot cross?Agent Spec
Need failure types from real outputs?Error Analysis
Need a test set built from real cases?Golden Dataset
Need to attack it before strangers do?AI Red Teaming Plan
Need to know a prompt edit broke nothing?Prompt Regression Test
Need cost per successful outcome?AI Unit Economics

The quality bar

Every skill in the pack holds the same standard, the one we hold when we ship AI in our own product:

Who made this

Polar Bear is a people ops consultancy for human-size teams (20 to 200 people). Built by ex-McKinsey founders with a dream to make AI work for People, not instead of them. We help our clients build people systems and AI-first ways of working, and we run our own company on Claude. This pack is the free, self-serve version of how we work.

The pack carries one AI product manager's work. When you want your whole team working this way, AI carrying the overhead so people do the part only people can do, across hiring, management and everyday operations, that's what we build with clients.

Meet Pauline. A demo that wowed and a feature nobody can grade? Bring ten real outputs and the launch you're weighing — we'll see what the evals say. Book a 30-minute call · Pauline on LinkedIn

Install in Claude Code

Two lines, and every skill in the pack loads at once.

/plugin marketplace add polar-bear-org/claude-skills
/plugin install ai-product-managers-pack@polar-bear-skills

Using Claude on the web instead? Download the zip and upload each skill from its install folder under Customize → Skills.