Specification
A frozen task, repository snapshot, acceptance criteria and environment identity.
Research program · Reliable and Human-Governed Agent Systems
I study the operational layer between a capable model and a trustworthy software change: task specification, skills, memory, tools, verification, release identity and human intervention.
Current material proves the existence of a longitudinal engineering corpus. It does not yet prove that skills improve success.
First paper
Which failures become reusable procedural skills, and which produce only one-off patches?
Do skills improve behavior-level acceptance on unseen tasks from the same domain?
When do version, environment, or abstraction mismatches create negative transfer?
Can scope, evidence, expiry, and independent verifiers outperform semantic retrieval alone?
Controlled comparison
The same model, tools, harness and budget, without task-specific skills.
Semantic retrieval without version, evidence, or expiry governance.
Domain-, version-, and environment-matched skills with provenance and validity.
Condition C plus an independent behavioral verifier that can reject completion.
Freeze 30 unseen tasks across at least five engineering domains. Use the pilot only to estimate baseline, variance, verifier reliability and execution cost. Power analysis and preregistration happen before the confirmatory study.
Auditable episode
A frozen task, repository snapshot, acceptance criteria and environment identity.
Tools, edits, failures, retries, intervention points and skills exposed to the agent.
Behavioral tests, public artifacts, release receipts, online smoke checks or rollback evidence.
Skill origin, reuse, conflict, expiry, negative transfer and later corrective incidents.
Publication boundary