Skills tagged evaluation: 31 agent skills for Claude Code
ab-testing — When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.
advanced-evaluation — This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration.
agents-optimize — Use when measuring or improving agent quality and performance — set up evaluators, online monitoring, CI/CD quality gates, observability, or cost optimization.
attribution — When the user wants to figure out which marketing actually drives conversions and revenue, choose or interpret an attribution model.
book-sft-pipeline — This skill should be used for book-to-SFT pipelines: ePub extraction, literary segmentation, author-voice dataset construction, style-transfer training.
code-maturity-assessor — Systematic code maturity assessment using Trail of Bits' 9-category framework.
customer-research — When the user wants to conduct, analyze, or synthesize customer research.
design-style-picker — Batch-generate and compare visual design directions so a user can choose the style they actually want.
dfam-check — Measure mesh files against Design for Additive Manufacturing (DfAM) rules and report printability findings per process (FDM, SLS, SLA/DLP, metal PBF, MJF).
evaluation — This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates.
fine-tuning-expert — Use when fine-tuning LLMs, training custom models, or adapting foundation models for specific tasks.
gcode — gcode` from 3D mesh files by orchestrating real slicer CLIs.
interaction-design-board — Generate several genuinely different, runnable HTML interaction prototypes for one product surface, combine them in an interactive Design Board.
interpreting-culture-index — Interprets Culture Index (CI) surveys, behavioral profiles, and personality assessment data.
let-fate-decide — Draws the 12 Houses of the Zodiac Tarot spread to inject entropy into planning when prompts are vague, ambiguous, or casually delegated.
llm-eval-harness — Test/evaluate any LLM behind an OpenAI- or Anthropic-compatible endpoint: availability (max_tokens-aware).
marketing-council — When the user wants multiple expert perspectives on a marketing question — a simulated board of advisors staffed by legendary marketers (Seth Godin.
ml-pipeline — Designs and implements production-grade ML pipeline infrastructure: configures experiment tracking with MLflow or Weights & Biases.
prompt-engineer — Writes, refactors, and evaluates prompts for LLMs — generating optimized prompt templates, structured output schemas, evaluation rubrics, and test suites.
promptfoo-evaluation — Configures and runs LLM evaluation using Promptfoo framework.
rag-architect — Designs and implements production-grade RAG systems by chunking documents, generating embeddings, configuring vector stores, building hybrid search pipelines.
sdf — SDFormat/SDF model and world authoring, validation, and simulator handoff.
security-reviewer — Identifies security vulnerabilities, generates structured audit reports with severity ratings, and provides actionable remediation guidance.
sendcutsend — com orders using its ordering guide, catalog, and specs.
skill-creator — Create new skills, modify and improve existing skills, and measure skill performance.
skill-inspector — Review AI agent skills before installation using NVIDIA SkillSpector and source-aware semantic review.
srdf — MoveIt2 SRDF authoring, validation, and planning-semantics workflow.
test-master — Generates test files, creates mocking strategies, analyzes code coverage, designs test architectures.
the-fool — Use when challenging ideas, plans, decisions, or proposals using structured critical reasoning.
urdf — URDF robot description authoring and validation.