Framework instrumentation is
instrument() as of 0.10.2; on
0.10.0 and earlier it was install(), with identical arguments. The old
name keeps working and emits a DeprecationWarning. Unrelated to
SkillRouter.install() below, which is about skills and keeps its name — the
collision between those two is exactly why the framework one moved.How fast?
~250–350ms warm. Cache hit (multi-turn) is free. Adds ~10–25% to total turn latency.
What if it fails?
Every method degrades to an empty fragment + a warning log. Your agent never breaks on a Router hiccup.
What's the payoff?
Per-skill, per-model effectiveness scoring in the dashboard. Every routing decision is logged, so effectiveness is measured rather than guessed.
Before / after
Integrate in 3 steps
1
Initialize the SDK
2
Install the framework adapter
Pick your framework. Each adapter is a one-line install with the same flag:Skills are injected into every
- LangChain
- OpenAI Agents
- Pydantic AI
- Anthropic SDK
- No framework
BaseChatModel.invoke() / ainvoke() as a SystemMessage. Works with LangChain agents, LangGraph nodes, and bare LLM calls.3
Run your agent — telemetry flows automatically
Make a normal call. Open the DecimalAI dashboard. You’ll see:
- Which skills the Router offered for each user turn
- Which ones the LLM activated
- Per-skill pass rate, per-(skill, model) effectiveness, trend over time
routing_id on every trace so the join can close server-side.How it works (high level)
The single load-bearing detail: the samerouting_id propagates from the Router through every LLM call in a turn to the trace. That join is what produces per-skill effectiveness scoring.
Raw trace ingest: the manifest handshake
If you’re calling the hosted API directly (no SDK adapter), closing the loop takes three calls — the first once per agent version, the other two per turn:- Register the manifest (once per agent version, not per turn) —
POST /api/v1/manifestswith{"manifest_hash": "...", "agent_name": "...", "version_label": "...", "components": [...]}. Onlymanifest_hashis required. The response’smanifest_idis what you stamp on every trace from this agent version. - Route the turn —
POST /api/v1/skills/routewith the user query (the curl in Inspecting a routing decision directly below). Splice the returnedprompt_fragmentinto your system prompt and hold onto therouting_id. - Ingest the trace —
POST /api/v1/traceswith both ids on the payload:
manifest_id ties the trace to the agent version it ran as (compat reports, regression checks); routing_id closes the offered-vs-activated join. Neither is validated at ingest — a missing or unknown id doesn’t fail the call, the corresponding join just never finds a row, so effectiveness panels stay silently empty. To serve a skill’s full body mid-turn (the load_skill path), the raw endpoint is GET /api/v1/skills/{skill_name}/body — parameters in Progressive disclosure below. From Python without a framework, decimalai.ingest_raw(payload) posts the same dict with no SDK modeling.
Routing strategies
The framework adapters default to smart route when a query is detectable in context, full menu otherwise. Smart route has one more trick: when your whole eligible menu fits the description token budget (~1,500 estimated tokens, max 30 rows), it returns the full menu instead of picking top-K — showing everything beats selection when everything is cheap, and it skips the embedding round-trip. You’ll see
"strategy": "full_menu" in the response when this fires.
Response shape
Bothget_menu() and smart_route() return the same top-level envelope — strategy, skills, prompt_fragment, routing_id. A smart-routed response:
prompt_fragment is a ready-to-splice markdown block. routing_id is the join key you (or your adapter) attach to the resulting trace.
The rows differ by strategy: full-menu rows carry only name, description and version. Smart-routed rows add skill_id, relevance, performance, score, activation_count and trend. Neither returns a category field.
Full-menu responses additionally carry the description-tier budget accounting:
desc_tokens is the estimated token cost of the rows shown; truncated: true means rows were dropped to stay under the budget (rows_total is how many were eligible).
Policy controls
Plan limits
The Router itself is Free. What scales with your plan is skill count and a few publishing / analytics capabilities:smart_route is Free because its value scales with skill count — a user with 8 skills doesn’t benefit; a user with 80 does, and at 80 they’ve already left the Free tier.
Smart-route internals: how skills are picked
Smart-route internals: how skills are picked
The server runs a hybrid retrieval pipeline against your registry:
- Embed the query with the production embedding model (currently
gemini-embedding-2, 768-dim). - Hybrid retrieve — combine semantic similarity (vector cosine on
description_embedding) with lexical match (PostgreSQL full-text search onsearch_doc). Fuse with Reciprocal Rank Fusion (RRF) — parameter-free, no weight to tune. - Re-rank by effectiveness — blend retrieval score with each skill’s historical pass rate. Skills with proven track records bubble up; new skills aren’t penalised until they have enough activations to be judged.
- Persist a
routing_decisionrow so the offered-vs-activated join can close.
Progressive disclosure: how skill bodies reach the agent
Progressive disclosure: how skill bodies reach the agent
Descriptions are the cheap, always-on tier. Bodies (the actual instructions) load on demand:
- The menu surfaces name + description rows, budgeted at ~1,500 estimated tokens / 30 rows.
- On the
openai_agentsandpydantic_aiadapters, enabling the skill loader also auto-registers a nativeload_skill(name)tool on every agent — no wiring on your side. The model reads the menu, callsload_skill("refund-policy"), and the skill’s full body arrives as the tool result. Parallel calls for multiple skills work. - Body loads are budgeted per turn: at most 3 bodies, ~6,000 estimated tokens total, each body trimmed to 8 KB. Past the budget the tool returns a polite refusal the model can act on; re-loading an already-loaded skill is free.
anthropic and langchain adapters patch a single create() / invoke() call — there’s no tool loop to route a result through — so they stay on prompt injection. Their inject_skill_body=True path applies the same trim + budget to injected bodies.The raw endpoint behind the tool is GET /api/v1/skills/{skill_name}/body, which accepts:The response’s
version field is the concrete resolved version number. Skills you subscribe to from the public registry (Use) are fetchable through this endpoint like your own.Kill switch for the tool: decimalai.init(load_skill_tool=False) or DECIMALAI_LOAD_SKILL_TOOL=0.Caching: why multi-step agent loops still produce one routing decision
Caching: why multi-step agent loops still produce one routing decision
The Router caches
build_prompt_fragment() results in-process for 30 seconds by default. This matters because a single user turn often produces multiple LLM calls (tool-using agent loops, retries, sub-steps). Without caching, each call would re-route the same query, producing duplicate routing_decision rows and wasted embedding calls.With caching, every LLM call within a turn:- Hits the platform once (the first call)
- Reuses the cached fragment and the same
routing_idfor all subsequent calls - Produces exactly one
routing_decisionrow for the entire turn
Routing telemetry — what gets logged
Routing telemetry — what gets logged
Every routing call writes a record. Joined with traces, it produces:
Per-(skill, model) effectiveness is computed from your own routing telemetry, so it reflects your traffic rather than a generic benchmark.
Graceful degradation — what happens when something fails
Graceful degradation — what happens when something fails
Net effect: a Router failure never breaks your agent run. Worst case, the LLM falls back to whatever instructions it had pre-Router.
Inspecting a routing decision directly
Inspecting a routing decision directly
What the Router does NOT do
What the Router does NOT do
- Execute skills. Skills are markdown — they shape the LLM’s behavior. For code execution, see how skills differ from tools.
- Decide which model to use. Model routing is separate. The Router operates within your chosen model.
- Cache bodies across processes. The in-process cache is per-process. For shared caching, put a service in front.
Router vs disk auto-loading (pick one)
Some agent runtimes (Claude Code, Cursor) auto-discoverSKILL.md
files from .claude/skills/ and inject them into the system prompt
themselves. The Router also injects skills into the system prompt — from
the platform. Running both at once means the same skill ends up in
the prompt twice, which confuses the LLM and inflates token cost.
The fix is to pick one source of skill injection per agent process:
The SDK auto-detects known disk runtimes via environment variables
(
CLAUDECODE, CLAUDE_CODE_ENTRYPOINT, CURSOR_AGENT) and logs a
one-shot warning when enable_skill_loader=True fires inside one of
them. Silence the warning intentionally with
DECIMALAI_SUPPRESS_DISK_RUNTIME_WARNING=1 if you’ve consciously
chosen the setup (e.g., a benchmark).
disk_sync=False — Router is the only source
When you’re running a Python stack that should rely solely on the
platform for skill content, pass disk_sync=False to the framework
adapter:
- Local
SKILL.mdauto-discovery (no disk read) - Push of local skills to the platform (no
sync_skills) - Pull of platform-only skills to disk (no
pull_missing)
/api/v1/skills/route per turn;
that’s the runtime injector. Everything happens over the network —
nothing on disk.
Available on decimalai.openai_agents.instrument() and
decimalai.langchain.instrument(). The pydantic_ai and anthropic
adapters never touched disk, so they don’t need the flag.
Related
- Skills API endpoints — raw REST surface
SkillRouterPython class — full SDK reference with CRUD, sync, export-to-disk- Skills & Data Pipeline — conceptual model: skills vs tools vs prompts
- Registry API — browse, use, and fork public skills
- Skills Observability tutorial — measure skill impact with experiments