Lifecycle Model

Proposed · draft 0.1Illustrative examples

The chain

A change usually travels from a problem to a corrective action, and the corrective action feeds the next change. The chain has two halves. Software engineering takes an idea from a problem to a release. Reliability engineering starts when a release breaks something and ends with a corrective action. Each stage is recorded as one or more events, and each event links back to the one before it. A stage can be skipped (a hotfix may have no proposal).

Software engineering

  1. Problemproblem.reportedA need, defect, or request. Usually where a story starts.
  2. Proposalproposal.recordedA proposed solution and its alternatives, including the rejected ones.
  3. Decisiondecision.madeWhat was decided, on what terms, and who owns the call.
  4. Implementationchange.implementedThe change itself, tied to the reasons it exists.
  5. Verificationverification.completedTests and checks, with the report as evidence.
  6. Releaserelease.deployedA version reaching production: the moment a change can affect users.

Reliability engineering

  1. Incidentincident.openedSomething broke. Linked to the release it followed, and to the change suspected of causing it.
  2. Mitigationmitigation.appliedA temporary measure that reduces impact, such as a feature flag.
  3. Rollbackrollback.executedA release reverted, then recovery checked.
  4. Postmortempostmortem.publishedThe review of an incident, so the team learns from it.
  5. Corrective actionaction.trackedWhat will be done so it does not recur. It feeds the whole of software engineering, and its outcome decides where it re-enters.

Stages

StageTypeWhoWhat
Problemproblem.reportedA person, or an issue trackerRecords a need, defect, or request. Usually the start of a story, so it has few links, or a relates_to link to a similar report. It gives every later event something to point back to.
Proposalinvestigation.completed, proposal.recordedAn AI agent on behalf of a personRecords a proposed solution and the alternatives. Links: investigates and motivated_by the problem. It keeps the rejected options, not only the one chosen.
Decisiondecision.madeA person, or a chat or tracker integration after a person confirmsRecords what was decided and on what terms. Link: decides a proposal. It shows who owns the call, so nobody digs through a thread.
Implementationchange.implementedA pull request or commit collector, with sources attachedRecords the change itself. Links: implements the proposal, relates_to the decision. It ties the code to the reasons it exists.
Verificationverification.completedCIRecords tests and checks, with the report as evidence. Link: verifies the change. A reader can trust the change without rerunning it.
Releaserelease.deployedA deployment systemRecords a version reaching an environment or a share of traffic. Link: releases a change. It marks the moment a change can affect users.
Incidentincident.openedAn alert or an incident toolRecords that something broke. Links: observed_after a release, caused_by a change (often a hypothesis at first). It points at the suspected cause while facts are still coming in.
Mitigationmitigation.appliedA personRecords a temporary measure that reduces impact, such as a feature flag. Link: mitigates an incident or a problem. It shows how impact was limited before a real fix.
Rollbackrollback.executed, recovery.verifiedA person, or a deployment systemRecords that a release was reverted, then that recovery was checked. Links: reverts the release, mitigates the incident, verifies the rollback. It shows the system recovered, not only that someone acted.
Postmortempostmortem.publishedA person, or a document collectorRecords the review of an incident. Link: follows_up the incident. It turns an incident into something the team learns from.
Corrective actionaction.trackedA person, tracked until a verified outcomeRecords what will be done so the problem does not recur, and whether it was done. Link: follows_up the incident or problem. It stops the next release from repeating the failure.

Confirm hypothesis

An agent can see a connection before anyone can prove it. Techlog records the guess honestly, and specific actions confirm it.

  1. GuessAn agent links an incident to a change as a hypothesis.
  2. ActReplay traffic, read a trace, inspect the code.
  3. ConfirmAn investigation links it again as confirmed, with evidence.
  4. FixThe corrective action follows up the incident.

Rules of thumb