advanced-evaluation
An agent skill by muratcankoylan, from muratcankoylan/Agent-Skills-for-Context-Engineering. Tags: analytics, automation, evaluation, product-strategy, strategy.
What it does
This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated quality assessment.
Install
With the skills CLI, which installs into Claude Code, Codex, Cursor and other agents:
npx skills add muratcankoylan/Agent-Skills-for-Context-Engineering --skill advanced-evaluation
Or copy the skill folder into Claude Code's skills directory by hand (~/.claude/skills for every project, or .claude/skills inside one):
git clone --depth 1 https://github.com/muratcankoylan/Agent-Skills-for-Context-Engineering
cp -r Agent-Skills-for-Context-Engineering/skills/advanced-evaluation ~/.claude/skills/advanced-evaluation
Safety box score
Not rated yet. A safety box score grades what a skill and its scripts can reach on the machine of whoever installs it, across eight categories from shell execution to secrets access. Anyone can request one from this page; it is saved for everyone. How the score works.
Source
- Repository
- muratcankoylan/Agent-Skills-for-Context-Engineering (all skills from this repository)
- Path
- skills/advanced-evaluation/SKILL.md
- Branch
- main
- Updated
- 2026-09-19
Related skills
- project-development — This skill should be used for project-level decisions about LLM-powered systems: whether an LLM is the right primitive for the task at hand.
- evaluation — This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates.
- memory-systems — This skill should be used for persistent semantic memory in agent systems: cross-session knowledge retention, entity tracking, temporal validity.
- brainstorm-experiments-existing — Design experiments to test assumptions for an existing product — prototypes, A/B tests, spikes, and other low-effort validation methods.
- dfam-check — Measure mesh files against Design for Additive Manufacturing (DfAM) rules and report printability findings per process (FDM, SLS, SLA/DLP, metal PBF, MJF).
- metrics-dashboard — Define and design a product metrics dashboard with key metrics, data sources, visualization types, and alert thresholds.