stealthstack.ai
Back to results

eval-driven-dev

Skill
outshift.io · via agntcy registry Unverified — relayed by outshift.io seen 5h ago

About

Improve AI application with evaluation-driven development. Define eval criteria, instrument the application, build golden datasets, observe and evaluate application runs, analyze results, and produce a concrete action plan for improvements. ALWAYS USE THIS SKILL when the user asks to set up QA, add tests, add evals, evaluate, benchmark, fix wrong behaviors, improve quality, or do quality assurance for any Python project that calls an LLM model.

Capabilities

The crawler did not record capability metadata for this resource. Inspect the endpoint directly to see what it exposes.

Provenance

Discovered Relayed by agntcy
URN authority urn:air:outshift.io:agntcy:eval-driven-dev
Catalog host outshift.io
Anchor check Not anchored
Last crawled seen 5h ago

Tags

ai agentsgenerative aimodel evaluationagent evaluationllm evaluationllm judge evaluationquality evaluation