docs: add incident resource support contract/RFC - #731
michael-richey wants to merge 2 commits into
Conversation
Add the Milestone 1 Contract/RFC for incident-resource synchronization support. This document is a prerequisite gate for any incident-resource implementation work; no implementation code may land until it is reviewed and approved. The RFC covers: - Endpoint inventory and disposition (17 syncable types; incident_configurations is destination-side safety state, not syncable) - Incident creation via POST /incidents/import (no-notification guarantee) and the defense-in-depth suppression barrier (per-incident configuration with execute_integrations/execute_notification_rules false) - Ownership attribution (OBO) with verified creator paths and fallbacks - Identity and unknown-POST recovery (creation_idempotency_key is response-only; concrete strategy required before coding) - Durability/saga backend (BaseStorage has no CAS/lease; design choice required before model work) - Nested-resource discovery (get_resources_by_ids parent-ID fan-out) - Dependency graph and set-preserving integer tier table - Per-resource contracts (types, fields, templates, handles, settings, incidents, impacts, integrations, todos, attachments, responders, timestamp overrides) - Incident-rule round-trip (incident_type_uuid is request-only; conditional support pending spike) - Rules installed disabled; activation as a distinct product workflow - Complete field/relationship fidelity and environment matrix - Behavior-changing-config classification - Replication-vs-sync product mode (cleanup disabled initially) - Per-wave rollback - Implementation sequence conditioned on RFC approval - Open [SPIKE] questions All schema claims verified against datadog-api-client 2.61.0.
|
Review feedback:\n\nThanks for putting this together — the depth here is strong and I agree with gating implementation behind an RFC. I have a few contract-level concerns to address before using this as the implementation baseline:\n\n1. Public OSS scope/privacy needs tightening.\n The doc currently includes internal-operational details that don’t belong in the public contract (for example: verification against an internal orchestration branch in the header, references to a DR orchestration companion, real two-DC/OBO path requirements, managed-binary pinning). In this repo, please keep the RFC focused on sync-cli’s public behavior/contracts and either:\n - generalize those sections to neutral wording (e.g., “external orchestrator”), or\n - move internal rollout/topology/process details into a private companion doc.\n\n Relatedly, claims in a public RFC should be reproducible from public artifacts; “verified against internal branch” isn’t independently checkable by external reviewers.\n\n2. Postmortem-template owner policy is still ambiguous in a place that needs to be deterministic.\n In §§5/10.6 the policy is currently “service-account ownership OR last-modifier approximation.” This affects OBO grouping, replay behavior, and drift characteristics, so we should lock a single contract choice (or mark this explicitly as a [SPIKE] gate that must resolve before Wave 1) rather than carrying an OR into implementation.\n\n3. Minor wording nit.\n §8 says the command-scoped dual-meaning allowlists are “intentional and tested.” Since this PR is RFC-only, consider rephrasing to “intentional and must be tested in implementation” to avoid implying existing test coverage.\n\nHappy to re-review after updates. |
Tighten public OSS scope/privacy, lock postmortem-template owner policy to a [SPIKE] gate, and fix a wording nit per review on PR #731. 1. Public OSS scope/privacy: generalize internal-operational references (internal DR orchestration service/branch, DR orchestration companion, two-DC/OBO path, managed-binary pinning) to neutral wording ("external orchestrator"). Scope the RFC to sync-cli's public behavior/contracts; move internal rollout topology/process details to a private companion doc. Make claims reproducible from public artifacts (published API docs + datadog-api-client package). 2. Postmortem-template owner policy: the OR (service-account OR last-modifier approximation) affects OBO grouping, replay, and drift, so lock it to a single deterministic [SPIKE] gate that must resolve before Wave 1 rather than carrying an OR into implementation. 3. Wording nit (section 8): "intentional and tested" -> "intentional and must be tested in implementation" to avoid implying existing test coverage in an RFC-only PR.
|
Thanks for the review — all three points addressed in commit 8065afa.
Happy to have another look. |
Summary
This PR adds the Milestone 1 Contract/RFC for incident-resource synchronization support — a prerequisite gate for any incident-resource implementation work. No implementation code lands in this PR; it is a design/contract document that must be reviewed and approved before Wave 0 (generic durability prep) and Wave 1 (the vertical safety slice) can begin.
Why an RFC first
The Incidents API is complex and preview-gated, and the central safety constraint — a sync-cli-created incident must never page or notify anyone — cannot be satisfied by a single endpoint choice. Several architecture decisions (identity/recovery, barrier lifecycle, saga/durability backend, OBO owner strategy, incident-rule round-trip) must be settled and signed off before any model work, because they determine whether the design is viable and what most later PRs contain. This document captures the verified schema research and the open
[SPIKE]questions that must be resolved in a controlled org before coding.What is in the RFC
incident_configurationsis destination-side safety state (no GET/LIST), not a synced resource.POST /incidents/import— the import endpoint documented contract (Imported incidents do not execute integrations or notification rules) and why that guarantee covers only the import operation (later PATCHes, child creation, and rule installation can still notify/page).execute_integrations=falseandexecute_notification_rules=falseestablished immediately after import, before any PATCH or child creation; the[SPIKE]configuration recovery contract; the barrier lifecycle policy; the tri-state side-effect matrix failing closed onunknown.created_by_user, optional with fallback;incident_types.attributes.created_byis a UUIDstr, not a handle; postmortem templates have nocreated_by_user; attachments inherit the parent incident created-by owner); the OBO grouper redesign; API permissions vs ownership (fail/report, no silent principal switch).creation_idempotency_keyis response-only and does not close the crash-after-acceptance window; one concrete strategy required before coding (provenance marker / operator mappings / manual-reconciliation; no heuristic adoption).BaseStoragehas no compare-and-swap/lease (Azure usesoverwrite=True); the saga atomic intent + replay fencing cannot be built on the current abstraction; design choice required (conditional API / single-writer invariant / append-only winner).get_resources_by_idsparent-ID fan-out (theteam_membershipspattern);ImportStateis write-only; command-scoped dual-meaning CLI allowlists.IncidentRuleDataAttributesRequestacceptsincident_type_uuidbut the response exposesincident_settings_association_uuidwith no relationships; conditional support pending a spike.[SPIKE]questions — configuration recovery, identity/idempotency, saga backend, parent discovery, incident-rule type-scoping, global-handle update/delete, API permissions, side-effect experiments.Verification
All schema claims were checked against
datadog-api-client2.61.0 (the version in the tox environment) and the live Incidents API documentation. The document contains no implementation code and no customer/internal data.Test plan
[SPIKE]questions (section 18) are resolved in a controlled org and signed off before Wave 1 coding begins.