Qualify a capability for a specific job.
Certification Garage builds and qualifies job-specific agents, capability packs, validators, SDKs, and DSLs against approved requirements. A capability must produce real work and leave evidence that the approved checks can evaluate. The result applies to the approved job version and operating boundaries, not to every task the same model, agent, or package might attempt.
The capability cannot rewrite its own test.
The operator and client approve the job-specific contract before execution. That package defines what counts as evidence, which checks are authoritative, and when a person must decide. It also prevents the capability under evaluation from weakening the requirement, selecting only favorable checks, or broadening a narrow result into a general claim.
The job version, objective, and required artifact
The package names the exact work, expected deliverable, intended operating context, and version being qualified. Changing the objective or artifact creates a new qualification target rather than silently widening the result.
Deterministic checks and acceptance thresholds
Required tests, validators, schemas, source rules, tolerances, and pass thresholds are approved in advance. Their failures remain authoritative even when a model believes the output is acceptable.
Allowed tools, models, data, and execution boundaries
The package records which resources may be used, which actions require approval, what data may leave the environment, and which limits stop the run. Qualification applies only inside those boundaries.
Evidence, redaction, review, retry, and stop rules
The operator defines what the receipt must preserve, what must be removed before sharing, how many repair attempts are allowed, and which uncertainty requires an authorized human decision.
Prepare, run, evaluate, and record.
Preparation here means job-specific instructions, examples, tools, validation, and bounded repair. It does not automatically mean base-model fine-tuning. The goal is a repeatable capability under a known contract: an operator should be able to see what was configured, what ran, what evidence was collected, and why the recorded outcome followed.
Build the package
Turn the approved requirements into a versioned task and evaluation package. Include the input fixture, required artifact, checks, thresholds, evidence fields, action policy, and review rules so another authorized run can apply the same contract.
Prepare the capability
Configure the agent, capability pack, validator, SDK, or DSL for that exact job. Preparation may include instructions, examples, tools, schemas, and bounded repair behavior, but it does not change the approved acceptance threshold.
Run real work
Produce the required artifact under the approved tools, data boundaries, time limits, and evidence rules. The run should exercise the behavior being qualified rather than substituting a descriptive promise or a hand-authored sample.
Evaluate honestly
Run authoritative checks, preserve their raw result and relevant evidence, and route defined uncertainty to human review. A favorable narrative cannot reverse a required failure or erase a missing deliverable.
A certification result belongs to one approved job version.
These are job-specific certification outcomes, distinct from CR Lite delivery standing. Deterministic checks remain authoritative; a model review cannot reverse a required failure. The outcome names whether the capability met the approved qualification contract, while the receipt carries the artifact, test evidence, exceptions, reviewer decisions, and conditions that make that result interpretable.
Available as a controlled external engagement.
External-first. Not currently an Exchange routing or gating service.
Packages, attempts, evidence, and results are handled outside live Exchange routing while the lane is qualified. Public copy does not claim durable Exchange certification records or automated market gates. Engagements begin with a bounded job, approved evaluation package, available test fixture, and an agreed reviewer. Results can inform later product work, but they do not automatically change listings, provider access, order routing, or settlement behavior.
Start with one job and one approved result.
Define the work that matters, the operating boundaries, the evidence you need, and the decision the result must support. Begin narrowly enough that the package can be reviewed, executed, and repeated without turning one qualification result into a claim about unrelated jobs.