Skills
Self-contained task procedures the Hermes agent can follow — each a single SKILL.md following the agentskills.io open standard.
2 results
Skill
evaluating-llms-harness
1.0.0Orchestra Research · mlops
lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.).
EvaluationLM Evaluation HarnessBenchmarking
0skill-creator
1.0.0anthropics · software-development
Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.
SkillsAuthoringMeta
0