evaluation
An agent skill by muratcankoylan, from muratcankoylan/Agent-Skills-for-Context-Engineering. Tags: analytics, automation, developer-tools, evaluation, performance.
What it does
This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and outcome measurement for agent pipelines.
Install
With the skills CLI, which installs into Claude Code, Codex, Cursor and other agents:
npx skills add muratcankoylan/Agent-Skills-for-Context-Engineering --skill evaluation
Or copy the skill folder into Claude Code's skills directory by hand (~/.claude/skills for every project, or .claude/skills inside one):
git clone --depth 1 https://github.com/muratcankoylan/Agent-Skills-for-Context-Engineering
cp -r Agent-Skills-for-Context-Engineering/skills/evaluation ~/.claude/skills/evaluation
Safety box score
Not rated yet. A safety box score grades what a skill and its scripts can reach on the machine of whoever installs it, across eight categories from shell execution to secrets access. Anyone can request one from this page; it is saved for everyone. How the score works.
Source
- Repository
- muratcankoylan/Agent-Skills-for-Context-Engineering (all skills from this repository)
- Path
- skills/evaluation/SKILL.md
- Branch
- main
- Updated
- 2026-09-19
Related skills
- agents-optimize — Use when measuring or improving agent quality and performance — set up evaluators, online monitoring, CI/CD quality gates, observability, or cost optimization.
- advanced-evaluation — This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration.
- context-compression — This skill should be used when long-running agent sessions need context compression, structured summarization, compaction, token-per-task optimization.
- context-degradation — This skill should be used for diagnosing and mitigating context degradation: lost-in-middle failures, context poisoning, context clash, context confusion.
- context-optimization — This skill should be used for improving context efficiency: context budgeting, observation masking, prefix or KV-cache strategy, partitioning.
- latent-briefing — This skill should be used when the user asks to "share memory between agents", "KV cache compaction for multi-agent", "orchestrator worker context".