Incident Management

This is the standard the Quality Gate's Reliability lens holds work to for the moment something goes wrong in production. It covers the whole arc of an incident: how anyone raises one, how we respond and meet our legal duties, and how AI helps us investigate without widening the blast radius. It builds on the Infrastructure Planning Policy (which gives us backups, RPO/RTO targets, and reversible deploys), the Security Policy (what we protect and the risks we assume), and the pull-request discipline of the Development Guide. Deviations are allowed, but โ€” as everywhere in the handbook โ€” they must be deliberate and justified in the project's design notes.

An incident is where OSBR's values are tested hardest. Be Nice: we keep stakeholders informed, on time, in plain language, and we never hand a tired on-call human a change they did not ask for. Be Kind: the person who reports an incident โ€” or caused one โ€” is doing us a favour by surfacing it, so the report is thanked and the post-mortem is blameless; we fix systems, not people. Be Strong: we face the incident directly, contain it, tell the client the truth even when it hurts, and know when to call for help. Humans and AI agents work an incident here as collaborators โ€” and that partnership carries a specific hazard, live production, that this policy names head-on (ยง3).

How to read this policy

1. Reporting an Incident

The goal of reporting is to make it fast, safe, and normal for any developer or collaborator to raise a suspected incident, so the Security Officer can act while the situation is still small. The behaviour we want is simple: report early, report often. A rumour of a problem, reported in two minutes, beats a confirmed breach found three days later. This section exists to remove every reason someone might hesitate โ€” uncertainty, embarrassment, fear of blame, or worry about "wasting" the officer's time.

1-1. The reporter does not decide whether it is an incident

This is the single most important rule of reporting:

1-2. Where to report โ€” speed beats formality

1-3. What to report โ€” the three points

Every report MUST include these three points. Keep it short; a few sentences each is enough. You are not expected to know the full answer to any of them โ€” "I don't know yet" is a valid and useful answer.

  1. Summary โ€” What did you see? A plain description of the symptom (e.g. "prod API is returning other users' data on /orders", "I think I pushed an AWS key to a public repo").
  2. Status & certainty โ€” Is it ongoing or over? How sure are you? Say so honestly โ€” "still happening", "not sure if it's real", "90% sure". Low certainty is not a reason to stay silent.
  3. Scope of impact โ€” Who or what looks affected, as far as you can tell? Which system, which data, which customers โ€” even a rough guess ("looks like staging only, but I can't confirm prod is safe").

1-4. Report regardless of scale or certainty

1-5. No blame for reporting in good faith

OSBR operates a blameless / just culture (Dekker), grounded in psychological safety (Edmondson): people only surface problems early when they are confident that doing so will not be held against them. Punishing reporters destroys exactly the early-warning signal we depend on. This is Be Kind made concrete.

2. Responding to an Incident

We do not invent our own incident model. We follow the practices the field has already settled on โ€” NIST SP 800-61, SANS PICERL (Prepare, Identify, Contain, Eradicate, Recover, Lessons-learned), and the incident-command patterns published by Google SRE and PagerDuty โ€” right-sized for a small team. The goal is to detect an incident quickly, contain the damage, meet every mandatory reporting obligation on time, restore verified service, and make sure the same thing cannot happen the same way twice.

2-1. One owner โ€” the Security Officer as Incident Commander

Every incident has exactly one owner: the on-duty Security Officer, who acts as Incident Commander (IC) in the sense used by Google SRE and PagerDuty. The IC is the single decision-maker; they need not do every task, but own that every task happens.

The Security Officer MUST:

The reporter SHOULD stay reachable to answer follow-up questions but is not responsible for running the response unless asked.

Roles scale down, they do not disappear. On a large incident the IC delegates two roles from the same playbooks: an Operations/Tech lead who actually changes the system, and a Communications lead who handles client and internal updates. On a small incident one person may wear all three hats โ€” but the roles still exist, so nothing is dropped.

2-2. Determine type and severity

Before acting, the Security Officer classifies the incident. Classification decides severity, which decides how fast we move and who we wake up.

Type โ€” classify against the risk categories in the Security Policy, Appendix 2 (accidents / human error, external attacks, insider threats), and note whether personal data is in scope, because that triggers the legal duties in ยง2-3. Typical types: personal-data breach; unauthorised access / account compromise; malware / ransomware; availability incident (ties into the SLO / RTO / RPO targets in the Infrastructure Planning Policy); accidental disclosure or data loss.

Severity โ€” the Security Officer MUST assign a severity at declaration and MUST revise it as facts change. Severity is about impact and reach, not blame, and is set by the officer, never by the reporter.

Sev Meaning Response
SEV1 Confirmed personal-data breach, active attacker, or major outage. Reporting clocks likely running. IC engages immediately; client notified; consider external assistance (ยง2-4).
SEV2 Contained or limited-blast-radius incident; potential (unconfirmed) data exposure. IC owns; assess reporting duties (ยง2-3) without delay.
SEV3 Minor / suspected incident, no evidence of data exposure. Handle in-hours; record and monitor.
SEV4 / near-miss Caught before impact. Record as a free lesson (ยง1-4); no urgent response.

When unsure between two severities, the officer MUST pick the higher one and de-escalate later. Under the PDPA, APPI, and GDPR the reporting clock can start on a suspected breach, not only a confirmed one (ยง2-3). Downgrading is cheap; a missed legal deadline is not. Do not under-classify to avoid paperwork.

2-3. Mandatory-reporting obligations

As part of analysis, the Security Officer MUST determine whether a legal notification duty applies and start the clock the moment a breach is reasonably suspected โ€” not when it is fully understood. If a duty applies, the Disclose flow (ยง2-4) becomes mandatory, not optional.

Malaysia โ€” PDPA (Personal Data Protection Act 2010, Act 709). OSBR's home law. Under the Personal Data Protection (Amendment) Act 2024 (Act A1727), a data controller must notify the Personal Data Protection Commissioner of a personal-data breach as soon as practicable, and notify affected data subjects where the breach is likely to cause significant harm. The regulator is the Personal Data Protection Commissioner / Personal Data Protection Department (JPDP).

Japan โ€” APPI (revised Act on the Protection of Personal Information). A business handling personal information must report a reportable breach to the Personal Information Protection Commission (PPC, ๅ€‹ไบบๆƒ…ๅ ฑไฟ่ญทๅง”ๅ“กไผš) and notify affected individuals. A breach is reportable when it involves any of: sensitive personal information (่ฆ้…ๆ…ฎๅ€‹ไบบๆƒ…ๅ ฑ); a risk of property damage (e.g. leaked payment data); a breach committed for an improper purpose (cyberattack / unauthorised access); or more than 1,000 data subjects. Reporting is two-stage โ€” a preliminary report (้€Ÿๅ ฑ) promptly, within about 3โ€“5 days of becoming aware, and a final report (็ขบๅ ฑ) within 30 days of awareness (extended to 60 days where the breach was for an improper purpose). If some required items are not yet known by the deadline, file with what is known and complete it as the facts are established.

EU / EEA โ€” GDPR (where applicable). GDPR applies where a project processes the personal data of individuals in the EU/EEA (confirm scope with the client; mind data residency per the Infrastructure Planning Policy). Article 33 โ€” notify the competent supervisory authority without undue delay and, where feasible, within 72 hours of becoming aware, unless the breach is unlikely to result in a risk to individuals; a missed 72-hour mark must be explained. Article 34 โ€” where the breach is likely to result in a high risk to individuals' rights and freedoms, communicate it to the affected data subjects without undue delay.

More than one may apply. A single incident can trigger the PDPA, APPI, and GDPR at once. Track each clock separately โ€” they have different recipients and different deadlines โ€” inside the incident record (ยง2-4).

2-4. The five response flows

Once type, severity, and reporting duties are set, the Security Officer executes among these five flows. They are not strictly sequential: Record runs throughout, Disclose runs on the legal clock, and Request external assistance can start at any point. The lifecycle maps OSBR's flows onto SANS PICERL and the NIST SP 800-61 phases; a real incident loops back as new facts arrive.

OSBR flow PICERL phase NIST SP 800-61 phase
โ€” (before the incident) Prepare Preparation
Record Identify Detection & Analysis
Prevent (contain) Contain Containment
Remediate Eradicate + Recover Eradication & Recovery
Disclose (across all phases) Post-Incident notification duties
Request external assistance (any phase, as needed) โ€”
Close-out (ยง2-5) Lessons-learned Post-Incident Activity

Record ยท Identify. Open the incident record the moment the incident is declared and keep it current โ€” it is the backbone of every later flow and of any regulator submission. The Security Officer MUST capture: a single, timestamped, append-only timeline of what was observed, decided, and done, by whom; type, severity, and scope (which systems, whose data, how many records); any reporting clocks started (ยง2-3), with their deadlines; and evidence preserved before it is destroyed โ€” logs, metrics, and communication history are protected assets (Security Policy, Appendix 1). Preserve first; do not tamper while containing. The record SHOULD live in the project's agreed incident location, access-limited to those who need it, with personal data masked per the Security Policy.

Prevent (Contain) ยท Contain. Stop the bleeding before cleaning up โ€” containment comes before eradication. The Security Officer MUST contain proportionally to severity: revoke or rotate compromised credentials and keys (immediately revoke any committed credential, per the Security Policy); isolate affected hosts, disable compromised accounts, or block malicious traffic at the WAF / edge; and where appropriate throttle or take a service offline rather than let a breach continue โ€” a short outage can beat an ongoing data leak. Containment actions MUST be written to the record as they happen.

Remediate ยท Eradicate + Recover. Once contained, remove the root cause and restore verified service. The Security Officer MUST: eradicate โ€” remove the attacker's foothold, malware, or the defect and close the vulnerability that allowed the incident, fixing the root cause, not the symptom; recover โ€” restore service from a known-good state and, where data was lost, restore from tested backups against the datastore's RPO / RTO (per the Infrastructure Planning Policy); and verify โ€” confirm the system is clean and healthy through monitoring before declaring recovery, watching for recurrence. Root-cause fixes that cannot ship during the incident become items in the recurrence-prevention plan (ยง2-5), each with an owner and a verification date.

Disclose ยท cross-phase, on the legal clock. Tell the people who need to know โ€” honestly, promptly, in plain language. This flow is mandatory whenever ยง2-3 applies, and good practice even when it does not. The Security Officer (or Communications lead) MUST: file the regulator notifications on their deadlines (PDPA notification to the Personal Data Protection Commissioner as soon as practicable; APPI ้€Ÿๅ ฑ / ็ขบๅ ฑ to the PPC; GDPR Art. 33 to the supervisory authority); notify affected individuals where required (PDPA where significant harm is likely; APPI; GDPR Art. 34 high-risk); keep the client informed from the start โ€” never let a client learn of their own incident from a regulator or the news; and give internal stakeholders honest, timely status updates (Be Nice). Disclosure SHOULD state what happened, what data was involved, what we have done, and what affected parties should do โ€” no spin, no minimising. Every external communication is logged in the record.

Request external assistance ยท any phase. Knowing when to call for help is Be Strong, not weakness. The Security Officer SHOULD bring in outside help when the incident exceeds the team's capacity or authority: legal counsel for reporting obligations and liability (the default for any SEV1/SEV2 personal-data breach); the cloud provider's security team, or a specialist DFIR (digital-forensics & incident-response) firm for serious intrusions; law enforcement, where the client and counsel agree it is warranted; and the relevant CSIRT / JPCERT-CC-style coordination body where appropriate. External parties are given least-privilege access and recorded in the incident record. Involving them never removes the Security Officer's ownership of the incident.

2-5. Close-out โ€” every incident ends the same way

An incident is not closed when service is restored โ€” it is closed when we have learned from it (the PICERL Lessons-learned phase; NIST Post-Incident Activity). The Security Officer MUST produce all three artifacts below.

3. AI-Assisted Production Investigation

The same structured logs, metrics, and traces that let an on-call engineer find a fault are what let an AI agent investigate one. This section is the access-control and data-handling contract that makes it safe to actually do that in production. The promise is asymmetric on purpose: we want AI to make investigation faster โ€” trace an error to its span, correlate a spike to a deploy, read the error budget, draft the root-cause narrative โ€” without making the blast radius any larger than a human investigator's already is. So AI gets exactly the observer's reach and no more.

Two boundaries are inviolable: every path AI uses to reach production is read-only, and PII is masked before it enters AI context. They are not tunable per project, per incident, or per urgency. A faster investigation is never a reason to widen either one; if a boundary is in the way, the answer is a better read-only view or a better mask, never an exception. Be Nice โ€” AI does the tedious correlation across a hundred thousand log lines so a tired on-call human does not. Be Kind โ€” the people in the telemetry are protected, because their personal data never reaches the model at all. Be Strong โ€” an investigator, human or AI, that can see the whole system reaches root cause faster.

We lean on published standards rather than inventing our own: the principle of least privilege and read-only / just-in-time access (AWS IAM best practices; NIST SP 800-207 Zero Trust), PII de-identification / masking (NIST SP 800-122), data minimisation (GDPR Art. 5(1)(c)), human oversight of AI (NIST AI RMF), and the LLM-specific failure mode of sensitive-information disclosure (OWASP Top 10 for LLM Applications).

3-1. Same outputs a human reaches โ€” no private backchannel

AI investigates through the same observation outputs a human on-call engineer uses โ€” the central log store, the metrics and trace backend โ€” defined by OSBR's observability discipline.

3-2. Read-only is the first inviolable boundary

Every path AI uses to reach production observation data is read-only, enforced by the permissions of the identity, so that even a confused, prompt-injected, or buggy agent cannot mutate production through its investigation path.

3-3. PII masked before context is the second inviolable boundary

Personal data is masked or removed before it enters AI context โ€” before the bytes reach the model, not after. This inherits from OSBR's rule that PII must be masked or omitted at the source before it is written to any log, span attribute, or metric label, and from the Security Policy's treatment of telemetry as a Protected Asset. AI adds a second reason the masking must already be done: data placed in a model's context is data minimisation you can no longer take back.

3-4. Mutating actions live on a separate, human-approved path

An investigation may conclude that production must change โ€” restart an instance, roll back a deploy, scale a pool, flip a flag, correct a record. That action does not happen on the investigation path; it happens on a separate path gated by explicit human approval, the way high-privilege and break-glass access is gated elsewhere in modern practice (Google Cloud Privileged Access Manager; AWS IAM).

3-5. Human-in-the-loop โ€” AI diagnoses, a human decides

The output of AI-assisted investigation is a proposal, not an executed decision. Keeping a human in the loop for the action is what lets us take the speed of AI investigation without inheriting the risk of an autonomous agent changing production (NIST AI RMF).

3-6. Attribution and audit

Both boundaries are only real if their use is visible.

Anti-patterns this section exists to prevent: handing the investigation agent a broad admin/deploy credential "so it can fix things too"; enforcing read-only only in the prompt while the underlying token can write; masking PII after it reaches the model, or "planning to add masking later"; pasting a raw, unmasked production log dump into an agent because it is faster; letting AI restart, roll back, or scale production on its own conclusion with no human approving; a break-glass change applied with no record of who, what, or why; and AI reaching production data through a backchannel (raw DB, host shell) that bypasses the observability pipeline's masking.

References

Incident-handling frameworks

Incident command & operations

Blameless culture & post-mortems

Legal / mandatory reporting

Least-privilege, read-only & human-approved access

PII masking & AI oversight

Related OSBR standards