# Claude for AI Product Managers: 36 Claude Skills to Test an AI Feature Before It Ships

> From "do we even need AI" to the weekly transcript read — 36 skills that spec an AI feature as tests, build its evals from real outputs and leave the verdict to a named person.

36 free Claude Skills (Agent Skills) by Pauline Bertry, Polar Bear. For Claude Code, Claude.ai and Claude Desktop.

- Page: https://meet-polar-bear.com/skills/ai-product-managers-pack
- Download: https://meet-polar-bear.com/skills/ai-product-managers-pack.zip
- Source: https://github.com/polar-bear-org/claude-skills/tree/main/packs/ai-product-managers-pack
- License: Free to use inside your company, not for resale.
- For: AI product managers, product managers adding AI features, founders shipping AI products

## Install

**Claude Code** — two lines, every skill in the pack at once:

```
/plugin marketplace add polar-bear-org/claude-skills
/plugin install ai-product-managers-pack@polar-bear-skills
```

**Claude.ai / Claude Desktop** — download https://meet-polar-bear.com/skills/ai-product-managers-pack.zip and upload each skill from its `install/` folder under Settings → Customize → Skills.

**Manually** — copy `packs/ai-product-managers-pack/skills/*` from https://github.com/polar-bear-org/claude-skills into `~/.claude/skills/` or a project's `.claude/skills/`.

## The 36 skills

### 1. Aipm AI use case canvas (`aipm-ai-use-case-canvas`)

Canvas (prediction, judgment, action, outcome, input, feedback), simpler-fix check, data check, verdict

Say: "run aipm-ai-use-case-canvas", "AI canvas", "do we even need AI for this", "leadership wants an AI feature", "is this a good AI use case"

### 2. Aipm AI failure modes (`aipm-ai-failure-modes`)

Failure list per output, false yes vs false no cost, severity and detection, graceful failure, tolerable rate left blank

Say: "run aipm-ai-failure-modes", "AI FMEA", "which AI mistakes matter", "leaders expect it to be right every time", "false positives vs false negatives"

### 3. Aipm workflow or agent (`aipm-workflow-or-agent`)

Candidate patterns, cost, latency and failure trade-offs, simplest pattern that passes, trigger for the next step up

Say: "run aipm-workflow-or-agent", "do we need an agent", "workflow vs agent", "agent or prompt chain", "agentic architecture choice"

### 4. Aipm AI model selection (`aipm-ai-model-selection`)

Side by side test plan on real inputs, quality, latency and cost table per tier, pick per task, re-check date

Say: "run aipm-ai-model-selection", "which model tier should we use", "justify the model choice", "compare models side by side", "cheaper model or stronger model"

### 5. Aipm AI prototype brief (`aipm-ai-prototype-brief`)

What the demo tests, pass line set in advance, cherry-picked input check, gap-to-production list

Say: "run aipm-ai-prototype-brief", "leadership thinks the demo is the product", "AI demo to production", "test card for an AI prototype", "what does this AI demo prove"

### 6. Aipm AI prd (`aipm-ai-prd`)

Problem, inputs the model sees, outputs, behavior summary, fallbacks, data needs, release bar, open questions

Say: "run aipm-ai-prd", "write a PRD for an AI feature", "AI feature spec", "spec for our copilot", "PRD template for an LLM feature"

### 7. Aipm AI success criteria (`aipm-ai-success-criteria`)

Measurable criteria per dimension, today's baseline, target set by a named person, how each is measured

Say: "run aipm-ai-success-criteria", "define success criteria for our AI feature", "what does good mean for this model", "make it good is not a spec", "acceptance criteria for an LLM"

### 8. Aipm behavior contract (`aipm-behavior-contract`)

Must, must never and when-unsure lines as Given-When-Then cases, pass rate over repeated runs, guardrails

Say: "run aipm-behavior-contract", "behavior spec for our AI", "must never rules for the assistant", "Given When Then for an LLM", "turn helpful and safe into tests"

### 9. Aipm human in the loop (`aipm-human-in-the-loop`)

Approve, edit or take-over points, triggers, handoff message, response time a person can meet, log

Say: "run aipm-human-in-the-loop", "where should a human approve", "human in the loop design", "AI handoff to a person", "approval step for the agent"

### 10. Aipm data privacy brief (`aipm-data-privacy-brief`)

Data flow from user to model to logs, personal data in prompts, retention, region, questions for legal

Say: "run aipm-data-privacy-brief", "what user data reaches the model", "data flow for our AI feature", "privacy review for an LLM feature", "legal asked about the model's data"

### 11. Aipm system prompt (`aipm-system-prompt`)

Role, audience, task, rules, examples, output format, refusals and handoffs, critique of the current prompt

Say: "run aipm-system-prompt", "write the system prompt", "clean up our prompt", "system prompt template", "which line of the prompt does what"

### 12. Aipm context engineering (`aipm-context-engineering`)

Knowledge sources and owners, freshness rules, what to leave out, retrieval spot checks, context budget

Say: "run aipm-context-engineering", "context engineering", "which documents should the model read", "the AI answers from stale docs", "knowledge base for our AI feature"

### 13. Aipm agent spec (`aipm-agent-spec`)

Goal, allowed tools, permissions, stop conditions, spend budget, approval points, never-alone list

Say: "run aipm-agent-spec", "agent spec", "write the limits for our agent", "agent permissions", "where should the agent stop"

### 14. Aipm tool descriptions (`aipm-tool-descriptions`)

Per tool name, purpose, inputs, outputs, error messages, namespacing, overlap check, three test tasks

Say: "run aipm-tool-descriptions", "write tool descriptions", "the agent picks the wrong tool", "tool definitions for our agent", "MCP tool descriptions"

### 15. Aipm error analysis (`aipm-error-analysis`)

Notes on real outputs, open codes, failure types grouped, counts per type, the three to fix first

Say: "run aipm-error-analysis", "error analysis", "what is actually going wrong with our AI", "read our transcripts", "build a failure taxonomy"

### 16. Aipm golden dataset (`aipm-golden-dataset`)

20 to 50 real cases with expected outcomes, coverage by failure type, source and version, add-and-retire rules

Say: "run aipm-golden-dataset", "golden dataset", "build an eval set", "test cases for our AI feature", "did the new prompt actually help"

### 17. Aipm synthetic test data (`aipm-synthetic-test-data`)

Generated cases for coverage gaps only, each marked synthetic, checked by a person, kept out of the headline score

Say: "run aipm-synthetic-test-data", "synthetic test data", "generate edge cases", "we do not have enough test cases", "test cases for rare inputs"

### 18. Aipm eval rubric (`aipm-eval-rubric`)

One pass or fail check per failure type, grader per check, pass and fail examples, blockers

Say: "run aipm-eval-rubric", "eval rubric", "pass fail criteria for our AI", "how should we grade outputs", "our 1 to 5 scores mean nothing"

### 19. Aipm llm judge (`aipm-llm-judge`)

Judge prompt for one failure type, labelled set, agreement on a held-out split, bias checks

Say: "run aipm-llm-judge", "LLM as a judge", "judge prompt", "automate our eval grading", "can we trust the LLM grader"

### 20. Aipm eval plan (`aipm-eval-plan`)

Capability and regression suites, when each runs, what blocks a release, named owner of quality

Say: "run aipm-eval-plan", "eval plan", "AI eval strategy", "who owns quality", "when should evals run"

### 21. Aipm red team plan (`aipm-red-team-plan`)

Attack cases by risk class, who runs them, pass line, fixes before launch

Say: "run aipm-red-team-plan", "red team our AI feature", "prompt injection test plan", "the agent reads emails from strangers", "OWASP LLM top 10 checklist"

### 22. Aipm AI ux review (`aipm-ai-ux-review`)

18 interaction guidelines checked by phase, AI disclosure check, correction and handoff paths, fixes ranked

Say: "run aipm-ai-ux-review", "AI UX review", "human-AI interaction guidelines", "users keep rephrasing", "what happens when the assistant is wrong"

### 23. Aipm AI impact assessment (`aipm-ai-impact-assessment`)

Who is affected, intended use and misuse, data, oversight, mitigations, residual risk owner, adviser questions

Say: "run aipm-ai-impact-assessment", "AI impact assessment", "algorithmic impact assessment", "the customer wants an impact assessment", "foreseeable misuse of our AI feature"

### 24. Aipm AI risk register (`aipm-ai-risk-register`)

Cause, event, effect risks mapped to Map, Measure, Manage, owner, trigger, response

Say: "run aipm-ai-risk-register", "AI risk register", "NIST AI RMF risks", "legal and security want our AI risks", "turn these worries into risks"

### 25. Aipm legal questions (`aipm-legal-questions`)

Questions for the legal and privacy team by topic, with the facts attached; asks, never answers

Say: "run aipm-legal-questions", "questions for legal", "prepare for the legal review", "what do I ask legal about our AI feature", "legal sign-off for AI"

### 26. Aipm AI launch checklist (`aipm-ai-launch-checklist`)

Go-live list with an owner per line, from evals passed as run to rollback rehearsed

Say: "run aipm-ai-launch-checklist", "AI launch checklist", "go-live checklist for an AI feature", "are we ready to launch", "launch readiness review"

### 27. Aipm launch gates (`aipm-launch-gates`)

Stages, eval and live numbers per gate, kill criteria, rollback path, who signs each gate

Say: "run aipm-launch-gates", "AI launch gates", "staged rollout for an AI feature", "canary release plan", "kill criteria for our AI launch"

### 28. Aipm quality review (`aipm-quality-review`)

Transcript sample plan, review notes, new failure types, cases added to the golden dataset, decisions

Say: "run aipm-quality-review", "weekly AI quality review", "read our AI transcripts", "it worked for weeks then broke", "review production conversations"

### 29. Aipm feedback signals (`aipm-feedback-signals`)

Signal map (accept, edit, retry, rephrase, abandon, escalate, rating), logging, weekly readout, review triggers

Say: "run aipm-feedback-signals", "users never rate our AI answers", "implicit feedback for an AI feature", "what should we log for our assistant", "thumbs up thumbs down is not enough"

### 30. Aipm prompt regression test (`aipm-prompt-regression-test`)

Change note, cases to re-run, before and after per failure type, ship or hold for a named person

Say: "run aipm-prompt-regression-test", "did this prompt change break anything", "regression test a prompt edit", "before and after eval for our prompt", "the fix broke another case"

### 31. Aipm model migration (`aipm-model-migration`)

Deadline, breaking changes, eval rerun, cost difference, behavior diffs to read, rollout and fallback

Say: "run aipm-model-migration", "our model is being retired", "deprecation notice for our model", "migrate to the replacement model", "model upgrade plan"

### 32. Aipm AI incident response (`aipm-ai-incident-response`)

Severity levels, containment, customer correction, evidence timeline, fix into the golden dataset, blameless review

Say: "run aipm-ai-incident-response", "our AI told a customer something wrong", "AI incident playbook", "the chatbot made a promise we cannot keep", "switch off the AI feature"

### 33. Aipm unit economics (`aipm-unit-economics`)

Cost per call, per task and per successful outcome, heavy-user case, caching and batch levers, today's cost

Say: "run aipm-unit-economics", "cost per task", "cost per successful outcome", "what does each AI answer cost us", "finance wants to stop the feature"

### 34. Aipm usage and pricing test (`aipm-usage-and-pricing-test`)

Small real-task test, usage estimate at light, normal and heavy use, price options to test, what to re-measure

Say: "run aipm-usage-and-pricing-test", "what will this feature cost to run", "estimate AI usage before launch", "seat or usage pricing for AI", "credits pricing test"

### 35. Aipm feature card (`aipm-feature-card`)

Intended and out-of-scope use, data, eval results as run with dates, known limits, disclosures, owner

Say: "run aipm-feature-card", "model card for our feature", "what does the AI feature do and where does it fail", "AI fact sheet for sales", "document the AI feature for legal"

### 36. Aipm exec brief (`aipm-exec-brief`)

One page, bottom line first: how often it is wrong by severity, what that costs, run cost, the decision asked

Say: "run aipm-exec-brief", "how accurate is it", "explain the AI to leadership", "one-pager for execs on the AI feature", "go or no-go memo for the AI"

