Authoring Guide

Build skills and agents that work across Claude Code and Copilot CLI

Platforms

Two Agent Systems

One plugin, two platforms -- Claude Code and VS Code Copilot consume the same files differently.

AspectClaude CodeCopilot CLI
File namingkebab-case.mdPascalCase.agent.md
InvocationTask(subagent_type: “name”)@AgentName in chat
Tools formatRead, Glob, Grep, Bashread, search, codebase, agent
MCP formatmcp__server__tool_name’server/tool_name’
Modelsonnet / opus / haikuFull model string or array
Sub-agentsFlat (Task tool)agents: field (nested, VS Code 1.113+)
HandoffsN/Ahandoffs: buttons
Model formatsonnet / opus / haikuString or array (fallback chain)
Memorymemory: projectNot supported

Key Insight

VS Code reads Claude’s formats (.claude/agents/, .claude/skills/). Claude does NOT read VS Code’s formats. Put shared agents in .claude/agents/ for both platforms.

Skills

Skill Authoring

YAML frontmatter + markdown body = a reusable command.

Complete skill frontmatter
---
name: my-skill                    # Required. Becomes /plugin:my-skill command
description: What it does...      # Required. Used for auto-invocation matching
argument-hint: "<expected args>"  # Optional. UI placeholder text
context: fork                     # Optional. "fork" = isolated subagent execution
agent: my-agent                   # Optional. Agent for the fork
model: sonnet                     # Optional. opus | sonnet | haiku -- never a pinned id
effort: high                      # Optional. low | medium | high | xhigh | max
paths: ["**/*.ts"]                # Optional. Limit auto-activation to file patterns
disable-model-invocation: false   # Optional. Prevent auto-invocation
---

Inline Skill

No context: or agent:. Runs directly in the user’s conversation. Good for most skills.

Forked Skill

Uses context: fork + agent:. Creates an isolated subagent. Main context stays clean from large MCP output.

Coordinator Skill

Uses Task tool in the body. Delegates to executor agents. Should ONLY delegate — never implement directly.

$ARGUMENTS

$ARGUMENTS is replaced with whatever the user typed after the command. Example: /dx-req 2416553 sets $ARGUMENTS = “2416553”.

Agents

Agent Authoring

Define who does the work -- model tier, tools, and permissions.

Claude Code agent frontmatter
---
name: my-agent
description: What it does
model: sonnet              # sonnet | opus | haiku | inherit
tools: Read, Write, Bash   # OMIT to inherit ALL tools (recommended for MCP)
disallowedTools: Write     # Block specific tools while inheriting rest
permissionMode: default    # default | acceptEdits | bypassPermissions
maxTurns: 30
memory: project            # user | project | local
isolation: worktree        # Isolated git worktree
---

The #1 Gotcha: Tool Inheritance

If you specify tools:, the agent can ONLY use those listed tools. MCP tools are BLOCKED unless explicitly listed. Either omit tools: entirely to inherit everything (recommended for MCP-dependent agents), or use disallowedTools: to block specific tools while keeping everything else.

Model Tier Strategy

Opus -- Deep Reasoning

Architecture, planning, code review, security analysis, root cause diagnosis. Cost: ≈15x haiku.

T1

Sonnet -- Standard Work

Implementation, MCP coordination, PR review, AEM inspection. Cost: ≈5x haiku.

T2

Haiku -- Fast Ops

File lookups, doc search, git commits, config reads. Cost: 1x (baseline).

T3

Write the alias, never a pinned model id

Use model: opus | sonnet | haiku. An alias follows the current default generation; a version-pinned id goes stale the moment a new model ships.

Effort, and when to escalate

Tiering applies at two levels: agents set model: in frontmatter, and skills can set model: and effort: to run directly without an agent. Opus already runs at high effort by default, so every model: opus skill gets deep reasoning without setting anything.

TierEffortUseWhere it is set today
OpusxhighEscalation above the high baseline — multi-file architectural review, complex verificationNothing today. Opt-in only
OpushighDeep reasoning — planning, review, verificationdx-plan, dx-pr-review, dx-security, dx-simplify, dx-step-verify
Sonnet(default)Execution — steps, inspectionsdx-step, dx-req, dx-step-fix
HaikulowSimple lookups — file and doc searchdx-help, dx-ticket-analyze, dx-pattern-extract

xhigh is not a free upgrade

Reach for effort: xhigh only when a step has demonstrably failed at high, or when reviewing more than five files of changes. It costs more and runs slower. No skill in the repo sets it today, so it has never been measured against high — re-check before assuming an escalation helped.

effort: was silently ignored before v2.1.267

That release (2026-09-09) fixed effort: on custom commands, skills and subagents being dropped on models whose default effort is pinned (Opus 4.7, Opus 4.8, Fable 5). Any effort: set before it was a no-op on those models, so this tiering is only now actually in force. The same release added a maxEffortLevel setting (top-level or per-model under modelSettings) that caps effort across every provider — useful as a cost ceiling for dx-automation pipeline runs.

Structure

File Structure and Naming

Where files live and how they are discovered.

Plugin directory layout
plugin/
├── .claude-plugin/plugin.json   # Plugin manifest (both platforms)
├── .mcp.json                    # MCP server config
├── agents/                      # Agent definitions (*.md with YAML frontmatter)
├── skills/                      # Skill directories (*/SKILL.md)
├── rules/                       # Default prompt templates
├── hooks/                       # Plugin hooks (hooks.json)
├── data/                        # Seed files copied to project by init
├── shared/                      # Reference files read by skills
└── templates/                   # Init-time file templates
  ├── rules/                   #   Convention rules -> .claude/rules/
  ├── instructions/            #   Detailed docs -> .github/instructions/
  └── agents/                  #   Copilot agents -> .github/agents/

Agent File Naming

  • Claude Code: agents/dx-code-reviewer.md
  • Copilot (.github): DxCodeReview.agent.md
  • Copilot CLI (plugin): Reads Claude format directly

Plugin Manifest

Never add agents or skills fields for default directories — this breaks Claude Code with a Zod validation error. Both platforms auto-discover agents/ and skills/ when these fields are omitted.

Config

Config-Driven Development

Skills never hardcode project-specific values.

Instead ofUse
https://myorg.visualstudio.com/scm.org from config
mvn clean installbuild.command from config
/apps/myproject/components/aem.component-path from config
http://localhost:4502aem.author-url from config
developscm.base-branch from config

Three-Layer Override System

.ai/rules/<topic>.md (full replacement) > config.yaml overrides: (quick tweaks) > plugin defaults (rules/*.md). Projects can also shadow entire skills by creating .claude/skills/<name>/SKILL.md.

Copilot

Copilot Agent Templates

Translating Claude agents for VS Code Chat.

Claude CodeCopilot Template
name: dx-code-reviewername: DxCodeReview (PascalCase)
model: opusOmit (VS Code selects model)
tools: Read, Glob, Grep, Bashtools: [read, search, search/codebase]
mcp__ado__search_code’ado/search_code’
Task tool delegationagents: [‘sub-agent-name’] + tools: [agent]
ToolSearch(“+AEM”)Remove — Copilot auto-loads MCP tools
model: opusmodel: [‘claude-opus-4-5’, ‘claude-sonnet-4-5’] (fallback array)

Handoffs

Copilot agents support handoffs: for workflow navigation buttons. Design them to follow natural workflow progression. Cross-plugin handoffs are supported — AEM agents can hand off to DX agents.

Manifest

Plugin Manifest -- Dual-Platform Format

One plugin.json, read by both Claude Code and Copilot CLI.

.claude-plugin/plugin.json
{
"name": "dx-core",
"version": "3.6.9",
"description": "ADO/Jira workflow for AI-assisted development",
"author": { "name": "..." }
}

Never list agents or skills for default directories

Adding agents or skills fields when using the default agents/ and skills/ directories breaks Claude Code with an “agents: Invalid input” validation error. Both platforms auto-discover those directories. Only add the fields for non-standard paths, e.g. {"agents": ["./specialized-agents"]}.

ManifestRead byNotes
.claude-plugin/plugin.jsonClaude Code, Copilot CLIAuto-discovers default directories
.cursor-plugin/plugin.jsonCursorNeeds explicit skill / agent / hook paths
gemini-extension.jsonGemini CLIRepo root, not per-plugin

Versioning is automated -- do not hand-edit

semantic-release bumps every version file on push to main from the conventional-commit prefix: all four plugin manifests (both .claude-plugin and .cursor-plugin), gemini-extension.json, marketplace.json, and the config template. The authoritative list is scripts/bump-versions.sh + .releaserc, and scripts/validate-structure.sh enforces they stay in sync. feat: is a minor bump, fix: a patch, BREAKING CHANGE: in the body a major; chore:, docs: and ci: trigger no release.

Verification

Verifying a Change

There is no build step -- so verification happens at four independent levels.

LevelWhat it checksHow
StructuralNaming, frontmatter, manifest/version consistency, collisionsscripts/validate-*.sh — CI, every PR
Helper scriptsDeterministic bash in skills/*/scripts/ and data/lib/Any run-tests.sh / *.test.sh — CI, every PR
Shared libsDeterministic JS in dx-core/data/lib/Any *.test.js via node —test — CI, every PR
Behavioral (evals)Whether the skill actually produces the right resultclaude plugin eval — not wired to CI (billed agent runs)

A structural check cannot tell you a script works

The validate-*.sh scripts read frontmatter, manifests and counts. They pass happily over a script that dies on its first call. Two habits follow: a change to anything under data/lib/ or skills/*/scripts/ needs its suite run, and node —check / bash -n are not verification — they parse syntax and cannot see a rejected argument, a schema change, or a wrong enum value. Run the thing.

CI discovers suites, it does not list them

validate.yml globs for run-tests.sh, *.test.sh and *.test.js under plugins/, so a new suite is picked up with no workflow change — a hardcoded list is how suites go unrun in the first place. The constraint this places on a new suite: it must be hermetic. No live ADO, AEM or network, and no dependence on ambient git config. Existing suites seed throwaway git repos and set user.email / user.name per repo; they pass with GIT_CONFIG_GLOBAL=/dev/null and an empty HOME. A check that needs a live service belongs in a skill, not a *.test.sh.

Behavioral evals

A skill is instructions for a model, so the same prompt can produce different output on different runs. An eval scores a skill over repeated runs against a fixture whose correct answer is known in advance — the opposite of an assertion, which expects one exact value. Suites live in plugins/<plugin>/evals/.

Run a suite
TMPDIR=/tmp claude plugin eval plugins/dx-core --case <case-name> \
--runs 3 --ablation none --scaffold --no-publish --allow-tools Write Bash

Fixtures carry a known defect

Plus correct distractors — so graders measure both recall (found the real problem) and precision (did not invent others).

Prefer deterministic graders

Use llm only when no exact oracle exists. A judge can be argued out of its rubric by a confident wrong answer; a regex cannot.

Sanity-check against known-bad input

A green suite means either the skill is good or the grader is blind, and you cannot tell which until something that should fail does. Each grader needs its own failing case.

Cost is real

Agent runs are billed; llm/baseline graders add 3 judge calls each, and —ablation with-without doubles the run count. Structural graders (regex, file_exists, tool_used, tool_order) are free. Pilot with —runs 1 and cap with —max-cost-usd.

Pin your models before comparing scores

Pin —model and —judge-model, or a model rollout reads as a regression. TMPDIR=/tmp keeps the sandbox from discovering ~/.claude/skills. Do not wire evals to every push — release tags or manual dispatch. Run artifacts land in plugins/*/evals/results/ and are gitignored.

Current coverage

One suite: dx-core/evals/plan-validate-finds-gap. Its README is a worked example — including a judge that passed a report its own rubric said to fail.

Checklist

PR Checklist

Verify before merging a new skill or agent.

Before You Merge

  • Skill has SKILL.md with proper frontmatter
  • No hardcoded project values (grep for org URLs, paths, names)
  • Config fields documented if new ones are introduced
  • Skill/agent catalog docs updated
  • Shell scripts are executable (chmod +x)
  • scripts/validate-*.sh pass
  • If data/lib/ or skills/*/scripts/ changed, its suite was run — not just bash -n
  • Any new test suite is hermetic (no live ADO/AEM/network, no ambient git config)
  • Tested with a real project
  • Copilot template added if user-facing
  • Version numbers not hand-edited — semantic-release owns them
KAI by Dragan Filipovic