Skip to content

docs: add incident resource support contract/RFC - #731

Open
michael-richey wants to merge 2 commits into
mainfrom
incident-rfc
Open

michael-richey wants to merge 2 commits into
mainfrom
incident-rfc

Conversation

@michael-richey

Copy link
Copy Markdown
Collaborator

Summary

This PR adds the Milestone 1 Contract/RFC for incident-resource synchronization support — a prerequisite gate for any incident-resource implementation work. No implementation code lands in this PR; it is a design/contract document that must be reviewed and approved before Wave 0 (generic durability prep) and Wave 1 (the vertical safety slice) can begin.

Why an RFC first

The Incidents API is complex and preview-gated, and the central safety constraint — a sync-cli-created incident must never page or notify anyone — cannot be satisfied by a single endpoint choice. Several architecture decisions (identity/recovery, barrier lifecycle, saga/durability backend, OBO owner strategy, incident-rule round-trip) must be settled and signed off before any model work, because they determine whether the design is viable and what most later PRs contain. This document captures the verified schema research and the open [SPIKE] questions that must be resolved in a controlled org before coding.

What is in the RFC

  • Endpoint inventory and disposition — every incident endpoint classified as syncable / singleton / action / write-only config / read-only / unsupported-fidelity. 17 syncable types; incident_configurations is destination-side safety state (no GET/LIST), not a synced resource.
  • Incident creation via POST /incidents/import — the import endpoint documented contract (Imported incidents do not execute integrations or notification rules) and why that guarantee covers only the import operation (later PATCHes, child creation, and rule installation can still notify/page).
  • Suppression barrier — the per-incident configuration with execute_integrations=false and execute_notification_rules=false established immediately after import, before any PATCH or child creation; the [SPIKE] configuration recovery contract; the barrier lifecycle policy; the tri-state side-effect matrix failing closed on unknown.
  • Ownership attribution (OBO) — verified creator paths per resource (incidents route by created_by_user, optional with fallback; incident_types.attributes.created_by is a UUID str, not a handle; postmortem templates have no created_by_user; attachments inherit the parent incident created-by owner); the OBO grouper redesign; API permissions vs ownership (fail/report, no silent principal switch).
  • Identity and unknown-POST recovery — creation_idempotency_key is response-only and does not close the crash-after-acceptance window; one concrete strategy required before coding (provenance marker / operator mappings / manual-reconciliation; no heuristic adoption).
  • Durability/saga backend — BaseStorage has no compare-and-swap/lease (Azure uses overwrite=True); the saga atomic intent + replay fencing cannot be built on the current abstraction; design choice required (conditional API / single-writer invariant / append-only winner).
  • Nested-resource discovery — get_resources_by_ids parent-ID fan-out (the team_memberships pattern); ImportState is write-only; command-scoped dual-meaning CLI allowlists.
  • Dependency graph and tier table — set-preserving integer tier insertion (no existing type dropped); ID-remap vs semantic-ordering vs safety edges.
  • Per-resource contracts — all 17 types with verified base paths, verbs, mapping keys, owner fields, readonly strips, and fidelity losses.
  • Incident-rule round-trip — IncidentRuleDataAttributesRequest accepts incident_type_uuid but the response exposes incident_settings_association_uuid with no relationships; conditional support pending a spike.
  • Rules — installed disabled on both create and update; canonicalize a copy before diff (no source mutation); activation is a distinct product workflow, not in this project; distinct activation policies for notification vs automation rules.
  • Complete field/relationship fidelity + environment matrix — UDF autocomplete metadata, integration metadata identifiers, incident-type config relationships, rule conditions/queries, template content/URLs.
  • Behavior-changing-config classification — inert / behavior-changing / action; behavior-changing behind separate opt-ins; preserve destination values by default.
  • Replication vs. sync — cleanup disabled initially means create/update-only; deletion qualification is a separate wave; docs/metrics say "replicated, deletion not converged."
  • Per-wave rollback — what is left in place; re-disable externally-enabled rules; barrier audit; global/type restore from snapshot; saga readable after downgrade; no automatic compensating delete.
  • Implementation sequence — Wave 0 (durability prep) then Wave 1 (vertical slice) then Wave 2 (inert config) then Wave 3 (children) then Wave 4 (environment-backed config) then Wave 5 (rules disabled) then Wave 6 (deletion qualification); a narrow DR-orchestration companion change lands with each OSS wave.
  • Open [SPIKE] questions — configuration recovery, identity/idempotency, saga backend, parent discovery, incident-rule type-scoping, global-handle update/delete, API permissions, side-effect experiments.

Verification

All schema claims were checked against datadog-api-client 2.61.0 (the version in the tox environment) and the live Incidents API documentation. The document contains no implementation code and no customer/internal data.

Test plan

  • RFC reviewed and approved by the incident API owners and the DR orchestration team.
  • The [SPIKE] questions (section 18) are resolved in a controlled org and signed off before Wave 1 coding begins.
  • The implementation sequence (section 17) is re-confirmed against the approved RFC (resource counts, tier indices, supported types) before each wave.

Add the Milestone 1 Contract/RFC for incident-resource synchronization
support. This document is a prerequisite gate for any incident-resource
implementation work; no implementation code may land until it is reviewed
and approved.

The RFC covers:
- Endpoint inventory and disposition (17 syncable types; incident_configurations
  is destination-side safety state, not syncable)
- Incident creation via POST /incidents/import (no-notification guarantee)
  and the defense-in-depth suppression barrier (per-incident configuration
  with execute_integrations/execute_notification_rules false)
- Ownership attribution (OBO) with verified creator paths and fallbacks
- Identity and unknown-POST recovery (creation_idempotency_key is
  response-only; concrete strategy required before coding)
- Durability/saga backend (BaseStorage has no CAS/lease; design choice
  required before model work)
- Nested-resource discovery (get_resources_by_ids parent-ID fan-out)
- Dependency graph and set-preserving integer tier table
- Per-resource contracts (types, fields, templates, handles, settings,
  incidents, impacts, integrations, todos, attachments, responders,
  timestamp overrides)
- Incident-rule round-trip (incident_type_uuid is request-only; conditional
  support pending spike)
- Rules installed disabled; activation as a distinct product workflow
- Complete field/relationship fidelity and environment matrix
- Behavior-changing-config classification
- Replication-vs-sync product mode (cleanup disabled initially)
- Per-wave rollback
- Implementation sequence conditioned on RFC approval
- Open [SPIKE] questions

All schema claims verified against datadog-api-client 2.61.0.
@michael-richey

Copy link
Copy Markdown
Collaborator Author

Review feedback:\n\nThanks for putting this together — the depth here is strong and I agree with gating implementation behind an RFC. I have a few contract-level concerns to address before using this as the implementation baseline:\n\n1. Public OSS scope/privacy needs tightening.\n The doc currently includes internal-operational details that don’t belong in the public contract (for example: verification against an internal orchestration branch in the header, references to a DR orchestration companion, real two-DC/OBO path requirements, managed-binary pinning). In this repo, please keep the RFC focused on sync-cli’s public behavior/contracts and either:\n - generalize those sections to neutral wording (e.g., “external orchestrator”), or\n - move internal rollout/topology/process details into a private companion doc.\n\n Relatedly, claims in a public RFC should be reproducible from public artifacts; “verified against internal branch” isn’t independently checkable by external reviewers.\n\n2. Postmortem-template owner policy is still ambiguous in a place that needs to be deterministic.\n In §§5/10.6 the policy is currently “service-account ownership OR last-modifier approximation.” This affects OBO grouping, replay behavior, and drift characteristics, so we should lock a single contract choice (or mark this explicitly as a [SPIKE] gate that must resolve before Wave 1) rather than carrying an OR into implementation.\n\n3. Minor wording nit.\n §8 says the command-scoped dual-meaning allowlists are “intentional and tested.” Since this PR is RFC-only, consider rephrasing to “intentional and must be tested in implementation” to avoid implying existing test coverage.\n\nHappy to re-review after updates.

@michael-richey
michael-richey marked this pull request as ready for review September 30, 2026 19:05
@michael-richey
michael-richey requested a review from a team as a code owner September 30, 2026 19:05
Tighten public OSS scope/privacy, lock postmortem-template owner
policy to a [SPIKE] gate, and fix a wording nit per review on PR #731.

1. Public OSS scope/privacy: generalize internal-operational references
   (internal DR orchestration service/branch, DR orchestration companion,
   two-DC/OBO path, managed-binary pinning) to neutral wording
   ("external orchestrator"). Scope the RFC to sync-cli's public
   behavior/contracts; move internal rollout topology/process details to a
   private companion doc. Make claims reproducible from public artifacts
   (published API docs + datadog-api-client package).

2. Postmortem-template owner policy: the OR (service-account OR
   last-modifier approximation) affects OBO grouping, replay, and drift, so
   lock it to a single deterministic [SPIKE] gate that must resolve before
   Wave 1 rather than carrying an OR into implementation.

3. Wording nit (section 8): "intentional and tested" -> "intentional
   and must be tested in implementation" to avoid implying existing test
   coverage in an RFC-only PR.
@michael-richey

Copy link
Copy Markdown
Collaborator Author

Thanks for the review — all three points addressed in commit 8065afa.

  1. Public OSS scope/privacy tightened. All internal-operational references (internal DR orchestration service/branch, DR orchestration companion, two-DC/OBO path, managed-binary pinning) are generalized to neutral wording ("external orchestrator"). The RFC is now scoped to sync-cli's public behavior/contracts; internal rollout topology/process details are noted as belonging in a private companion doc. The verification basis now states claims are reproducible from public artifacts (published API docs + the datadog-api-client package), with no reference to an internal branch.

  2. Postmortem-template owner policy locked to a [SPIKE] gate. In §§5 and 10.6 the OR is replaced with a single deterministic [SPIKE] gate: "must resolve to a single deterministic choice (service-account ownership OR last-modifier approximation) before Wave 1; an OR is not carried into implementation." This removes the ambiguity for OBO grouping/replay/drift.

  3. Wording nit fixed. §8 now reads "intentional and must be tested in implementation" to avoid implying existing test coverage in an RFC-only PR.

Happy to have another look.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant