Files
skillfactor-pipeline/pipeline/templates/QUALITY_BAR.md

2.3 KiB

QUALITY_BAR — acceptance checklist per occupation package

The bar every package must clear before its depth files count as done. Reference implementation: skills/recruitment-consultant/ (gold standard, reviewed 2026-07-07). The QA harness (pipeline/qa_sample.py) checks the mechanical items on a 2 % sample per batch; the judgment items are for Claude spot-checks and human review.

Mechanical (harness-checkable)

  • All four depth files exist: references/usecases.md, references/intake.md, references/quality.md, evals/eval-01..05.md
  • Section skeletons match the templates (headings greppable: # Use cases —, # Intake —, # Quality criteria —, # Eval 0)
  • usecases.md has 6-8 ## N. sections, each with an Output: line
  • intake.md has 10-14 numbered questions and a "Do not guess" marker
  • quality.md has 15-25 checkboxes (- [ ])
  • 5 eval files, each with ## Task and ## Pass criteria, <= 30 lines
  • One eval contains a marked trap (TRAP: in pass criteria)
  • Every market percentage cited in usecases.md exists verbatim in market.md (no numbers from memory)
  • No banned filler words (see templates/STYLE.md list), no TODO/TBD
  • No real PII / employer names from job ads in evals or examples
  • manifest.json updated (sources incl. evidence as-of date), PROVENANCE.md regenerated, sanitize gate passes

Judgment (Claude/human review)

  • Use cases are the LLM-doable slice of the occupation (paperwork around the craft for manual jobs, never the physical act)
  • Use-case ranking follows leverage (market share x LLM fit), not alphabetical or random order
  • Intake questions target this occupation's real failure modes, not generic project questions
  • quality.md encodes the occupation's non-negotiables as hard stops
  • Evals are fully specified (agent needs nothing beyond the file) and the trap actually tests refusal/flagging behaviour
  • Tone per STYLE.md: terse, evidence-based, no marketing filler

Extraction quality (per evidence batch, 2 % sample)

  • >= 90 % of sampled records faithful to the ad text (no invented items)
  • Seniority plausible vs. title/requirements
  • No cross-field duplicates (product in tools only)
  • Items normalized: english, singular, lowercase except products/acronyms