How Skill Synthesis Works
How Skill Synthesis Works
Section titled “How Skill Synthesis Works”A skill goes through three stages: capture, distillation, and a living library. The Skills tab surfaces each stage (Sessions → Recommended / auto-promoted → Library).
The stage that matters most for what you should expect is distillation, because it splits into two tracks with very different levels of human involvement. Read that section closely before the rest.
Stage 1 — Capture (Sessions)
Section titled “Stage 1 — Capture (Sessions)”Every session-end, idle gap, subagent stop, completed turn, and boot scan writes a cheap regex-only row into a durable queue — nothing is read by a model yet. A background drain later picks the row up and runs the real capture:
session ends / idle / subagent stops / turn completes / boot scan ↓Trajectory extractor → normalized turns + tool sequence + success signal ↓Prefilter (regex, free) → drop sessions with no signal at all ↓Trajectory-hash dedup → skip if this exact trajectory was already captured ↓Candidate row (Sessions, status: candidate) ↓Session archaeologist → reads the transcript, writes an intent/outcome/friction verdictThe prefilter is deliberately generous: a candidate passes if it edited code, used tools, ran a test command, or simply sustained a long back-and-forth — any one of those is “there’s something here.” That’s a deliberate widening: a debugging session that took three failed attempts and a correction is exactly the material worth capturing, and it’s often short on edits but dense with friction. See Background Learning for what the archaeologist does with that friction once it runs.
For each captured session, an LLM distills a first-pass { name, description, body } following skill-authoring best practices (a trigger-oriented description, an imperative body, no workspace-specific paths). On boot scans, or when no model is available, a template is used instead.
Stage 2 — Distillation: two tracks, one library
Section titled “Stage 2 — Distillation: two tracks, one library”Track 1 — direct promotion (fully automated)
Section titled “Track 1 — direct promotion (fully automated)”The “I did the exact same thing enough times” path. Once a candidate’s recorded success count reaches the threshold, promotion runs automatically:
successesToPromote reached ↓duplicate-of-active-skill check (cosine dedup) ↓judge scores the candidate — must be 'scored' and ≥ minJudgeScore ↓replay confidence check — blocks only if MEASURED below the floor ↓residency cap check (may demote the weakest active skill to dormant) ↓SKILL.md written, candidate → promotedThere is no “recommend it to the user” step anywhere in that chain. The Promote button in Sessions still exists, but it’s a manual override for jumping the threshold — not the only door in.
Track 2 — cluster → Recommended (human-in-the-loop)
Section titled “Track 2 — cluster → Recommended (human-in-the-loop)”The path for workflows that are similar but not identical — no single session repeats often enough to hit the threshold on its own, but several sessions cluster together. The Curator pass:
- Clusters candidates that look alike (at least
skillSynthesis.suggestionMinClusterSize, default 2) - Holds one cluster member back so nothing is judged against a session that trained it — see the replay note below
- Synthesizes one generalized, repo-agnostic skill from the rest
- Runs it past the same quality judge
- If it passes, proposes it in Recommended for you to review, edit, and Accept
Nothing here reaches your library without acceptSuggestion. This is the track the phrase “review your recommendations” has always correctly described — it just isn’t the only track.
The judge
Section titled “The judge”Before anything is promoted or recommended, the judge scores it 1–10 on five criteria and averages them:
| Criterion | Asks |
|---|---|
| novelty | Is this non-obvious versus what an agent already knows? |
| actionability | Are the steps concrete and ordered? |
| scope | Is it one well-defined workflow, not a trivial one-off? |
| generalization | Repo-agnostic and transferable, with no session-specific leftovers? |
| triggerClarity | Does the description clearly say when to use the skill? |
The average is compared against skillSynthesis.minJudgeScore (default 6.0).
The judge panel (weekly, deeper)
Section titled “The judge panel (weekly, deeper)”Once a week, candidates that reached the judge get a second, independent opinion — and the two panellists are asked different questions, which is the entire reason it’s worth a second call. Panellist A judges the artifact cold, same as the capture-time judge. Panellist B judges the same artifact but is also shown the candidate’s nearest description-neighbours among your active skills and whatever the empirical gates below have already measured for it. If the two disagree by more than skillSynthesis.judgePanel.disagreementThreshold on any single criterion, a third call reads both rationales and adjudicates.
This doesn’t gate the capture-time promotion decision — Track 1 can promote before its weekly panel ever runs. It’s what backs the scorecard you see for a candidate, and the foundation later gates build on.
The empirical gates — measuring instead of asking
Section titled “The empirical gates — measuring instead of asking”Two gates replace a model’s opinion with a measurement. Both run weekly.
Trigger eval asks: given prompts this skill should answer and prompts it shouldn’t, how often does its description actually come back from retrieval? One cheap LLM call generates the test prompts; everything after that — embedding, ranking, precision, recall — is local vector math, never a second model opinion. The result replaces the judge’s triggerClarity guess with a measured number wherever a candidate has one.
Replay validation asks something a rubric can’t: hold one session out of the cluster a skill was drafted from, hand the skill plus that session’s opening ask to a fresh model, and see whether the plan it produces resembles what that session actually did.
Both gates share the same rule for missing data: null means never measured, 0 means measured and got nothing. A skill nobody has replayed reads as “not measured,” never as a zero — the UI never turns an absent measurement into a number that looks like a bad score.
Before a candidate is created or counted, its embedding is compared against the active skill set. If cosine similarity to any active skill is ≥ skillSynthesis.dedupCosineThreshold (default 0.85), the trajectory is treated as already represented rather than creating a duplicate.
Stage 3 — A living library
Section titled “Stage 3 — A living library”Materialized skills, plus cloned agents and commands, live in the Library. Ptah records when each one is actually used — the Skill tool, slash-command/skill expansion, and subagent runs (by subagent_type). That usage signal drives auto-enhancement:
≥ 5 recorded runs and not in cooldown ↓Curator rewrites the skill against its recent usage (judge-gated) ↓previous version snapshotted to History → re-propagated ↓24h cooldownYou can also Enhance now to run it manually, or Revert to any History snapshot.
Residency
Section titled “Residency”Active skills are capped at skillSynthesis.maxActiveSkills (default 200). When the cap is exceeded, the weakest resident is demoted to dormant — it stays on disk and in the database but is skipped when skills are loaded into a session. Dormant skills are never deleted, and authored skills are exempt from demotion entirely.