> ## Documentation Index
> Fetch the complete documentation index at: https://docs.decimal.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Run the A/B benchmark suite for this skill's latest version

> Phase 3 — Benchmark runner trigger.

Synchronously runs every SkillTestCase for the skill's latest version
against the configured agent stub, with skill and without. Returns
the SkillTestRun including per-case results. Caller can poll the
results endpoint or just consume the returned payload.

``runs`` (default 1) re-runs the whole suite uniformly and averages the
per-case results by MEAN — the run-level successor to the retired
per-case pass^k ``trials`` field (ADR-0007).



## OpenAPI

````yaml /openapi.json post /api/v1/skills/{skill_name}/benchmark/run
openapi: 3.1.0
info:
  title: DecimalAI Platform
  description: Agent Dataset Lifecycle Platform — Backend API
  version: 0.1.0
servers: []
security:
  - BearerAuth: []
tags:
  - name: traces
    description: >-
      Ingest, search, and inspect agent execution traces. A trace is the atomic
      unit of the platform — every other feature (evaluation, compatibility
      scoring, dataset building) operates on traces. Each trace is auto-tagged
      with the manifest hash of the agent that produced it.
  - name: eval-scores
    description: >-
      Push, fetch, and aggregate per-trace evaluation scores. Scores carry
      source provenance (built-in, DeepEval, LangSmith, custom) and feed the
      unified decision engine, which produces a keep / repair / replay / drop
      verdict per trace.
  - name: skills
    description: >-
      Manage reusable instruction blocks (SKILL.md files) that modify agent
      behavior. Skills are first-class manifest components — changes to a skill
      appear in your regression-check impact reports. Use these endpoints to
      create, version, fork, and sync skills with your local SKILL.md files.
  - name: manifests
    description: >-
      Register and inspect manifest versions. A manifest is a snapshot of your
      agent's structural identity (tools, prompts, models, skills) at a point in
      time. The SDK registers manifests automatically; these endpoints expose
      the underlying records.
  - name: datasets
    description: >-
      Build versioned SFT/DPO training datasets from manifest-classified,
      eval-scored traces. Datasets are reproducible — each version locks the
      manifest and filter set used to build it. Export as JSONL, Parquet, or
      push to HuggingFace Hub.
  - name: replay
    description: >-
      Re-execute historical traces against a new manifest version. Replay is the
      deferred-future capability for behavioral verification of agent changes.
      The SDK creates a replay batch, your worker executes each task, then
      submits results back here for eval scoring.
  - name: registry
    description: >-
      Browse and install community skills from the public skills registry. Each
      registry skill carries a SkillScore — an evidence-tiered quality composite
      computed from real-world evals, benchmarks, and ratings across
      organizations. Installing creates a fork in your org — edits don't affect
      the public version.
  - name: agents
    description: >-
      List and inspect agents, plus their multi-agent topology (orchestrator →
      sub-agent edges discovered from traces). Agents are identified by a stable
      agent_name string set in your SDK init() call.
  - name: import
    description: >-
      Bulk-import historical traces from another observability platform or a
      JSONL backup. Useful for migrating from LangSmith / Braintrust / Langfuse.
      Duplicate trace IDs are silently skipped — re-running an import is safe.
paths:
  /api/v1/skills/{skill_name}/benchmark/run:
    post:
      tags:
        - skills
      summary: Run the A/B benchmark suite for this skill's latest version
      description: |-
        Phase 3 — Benchmark runner trigger.

        Synchronously runs every SkillTestCase for the skill's latest version
        against the configured agent stub, with skill and without. Returns
        the SkillTestRun including per-case results. Caller can poll the
        results endpoint or just consume the returned payload.

        ``runs`` (default 1) re-runs the whole suite uniformly and averages the
        per-case results by MEAN — the run-level successor to the retired
        per-case pass^k ``trials`` field (ADR-0007).
      operationId: run_skill_benchmark_api_v1_skills__skill_name__benchmark_run_post
      parameters:
        - name: skill_name
          in: path
          required: true
          schema:
            type: string
            title: Skill Name
        - name: triggered_by
          in: query
          required: false
          schema:
            type: string
            description: Where this run was kicked off from
            default: api
            title: Triggered By
          description: Where this run was kicked off from
        - name: model
          in: query
          required: false
          schema:
            anyOf:
              - type: string
              - type: 'null'
            description: >-
              Override the agent-arm model (e.g. claude-haiku-4-5 for the
              model-relative lane). Must be in GET /api/v1/auth/models. Default:
              the platform's configured model. The executed model is recorded as
              runner_model on the run.
            title: Model
          description: >-
            Override the agent-arm model (e.g. claude-haiku-4-5 for the
            model-relative lane). Must be in GET /api/v1/auth/models. Default:
            the platform's configured model. The executed model is recorded as
            runner_model on the run.
        - name: runs
          in: query
          required: false
          schema:
            type: integer
            maximum: 10
            minimum: 1
            description: >-
              Re-run the WHOLE suite this many times, uniformly (1-10; default
              1). Every (case, run) record enters the aggregate at equal weight,
              so the headline pass-rate is a MEAN over runs whose expected value
              does not depend on this count — more runs only narrow the error
              bars (ADR-0007). This is the run-level replacement for the retired
              per-case pass^k `trials`.
            default: 1
            title: Runs
          description: >-
            Re-run the WHOLE suite this many times, uniformly (1-10; default 1).
            Every (case, run) record enters the aggregate at equal weight, so
            the headline pass-rate is a MEAN over runs whose expected value does
            not depend on this count — more runs only narrow the error bars
            (ADR-0007). This is the run-level replacement for the retired
            per-case pass^k `trials`.
        - name: Authorization
          in: header
          required: false
          schema:
            type: string
            title: Authorization
        - name: decimal_session
          in: cookie
          required: false
          schema:
            anyOf:
              - type: string
              - type: 'null'
            title: Decimal Session
      responses:
        '200':
          description: Successful Response
          content:
            application/json:
              schema: {}
        '422':
          description: Validation Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/HTTPValidationError'
components:
  schemas:
    HTTPValidationError:
      properties:
        detail:
          items:
            $ref: '#/components/schemas/ValidationError'
          type: array
          title: Detail
      type: object
      title: HTTPValidationError
    ValidationError:
      properties:
        loc:
          items:
            anyOf:
              - type: string
              - type: integer
          type: array
          title: Location
        msg:
          type: string
          title: Message
        type:
          type: string
          title: Error Type
        input:
          title: Input
        ctx:
          type: object
          title: Context
      type: object
      required:
        - loc
        - msg
        - type
      title: ValidationError
  securitySchemes:
    BearerAuth:
      type: http
      scheme: bearer
      description: Enter your API key (e.g. dai_sk_test_key_001)

````