TaGzxia
  • Home⌘H
  • Projects⌘P
  • Research⌘R
  • Contents⌘C
  • Milestone⌘M
  • Sponsor⌘S

Research program · Reliable and Human-Governed Agent Systems

When an agent says “done,” what evidence should survive?

Pilot-ready · not preregistered

I study the operational layer between a capable model and a trustworthy software change: task specification, skills, memory, tools, verification, release identity and human intervention.

Current material proves the existence of a longitudinal engineering corpus. It does not yet prove that skills improve success.

522
task sessions
2,568
indexed prompts
729
independent QA reports
200+
managed skills confirmed

First paper

From Incidents to Skills

A Longitudinal Field Study of Reusable Procedural Memory in Production Software Agents
  1. RQ1

    Which failures become reusable procedural skills, and which produce only one-off patches?

  2. RQ2

    Do skills improve behavior-level acceptance on unseen tasks from the same domain?

  3. RQ3

    When do version, environment, or abstraction mismatches create negative transfer?

  4. RQ4

    Can scope, evidence, expiry, and independent verifiers outperform semantic retrieval alone?

Controlled comparison

Same task. Same budget. Different memory discipline.

Condition order is randomized; graders do not see the assigned condition.
  1. ANoSkill

    The same model, tools, harness and budget, without task-specific skills.

  2. BSemanticSkill

    Semantic retrieval without version, evidence, or expiry governance.

  3. CGovernedSkill

    Domain-, version-, and environment-matched skills with provenance and validity.

  4. DGovernedSkill + Verifier

    Condition C plus an independent behavioral verifier that can reject completion.

Pilot gate

Freeze 30 unseen tasks across at least five engineering domains. Use the pilot only to estimate baseline, variance, verifier reliability and execution cost. Power analysis and preregistration happen before the confirmatory study.

Auditable episode

A patch is one artifact in a longer evidence chain.

01

Specification

A frozen task, repository snapshot, acceptance criteria and environment identity.

02

Trajectory

Tools, edits, failures, retries, intervention points and skills exposed to the agent.

03

Outcome

Behavioral tests, public artifacts, release receipts, online smoke checks or rollback evidence.

04

Lifecycle

Skill origin, reuse, conflict, expiry, negative transfer and later corrective incidents.

Publication boundary

Private traces stay private by default.

  • Raw trajectories remain in a local encrypted zone and never enter the public site or paper repository.
  • Credentials, cookies, hosts, paths, identities and free text are removed before research sampling.
  • Employer, customer, minor and consumer-media data require explicit written authorization or are excluded.
  • Only licensed tasks, structured trajectories, verifier outcomes and necessary minimal diffs may become public artifacts.
Current statusPilot-ready, not preregistered

Next artifact: a frozen data dictionary, task manifest, verifier specification and de-identification review.

Return to selected projects