Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co

**Current phase**: Implementation, under a locked pre-registered evaluation gate.

This repository is Populace dynamics: an open longitudinal microsimulation layer (paper at populace.dev/papers/dynamics) plus a working implementation in `src/populace_dynamics/`:
This repository is Microcosm Dynamics: an open longitudinal microsimulation layer (paper at microcosm.institute/dynamics/paper) plus a working implementation in `src/populace_dynamics/`:

- `harness/` — the population-view scoring harness (geometry blocks, PanelView trajectory windows, the moment battery in `moments.py`)
- `data/` — label-verified PSID readers (`family.py` builds the 1968-2022 head/spouse earnings panel with assignment flags; PSID files staged at `~/PolicyEngine/psid-data`)
Expand Down
2 changes: 1 addition & 1 deletion docs/_quarto.yml
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@ project:
output-dir: _book

book:
title: "Populace dynamics"
title: "Microcosm Dynamics"
subtitle: "An open, scored longitudinal layer for policy microsimulation — validated first on U.S. Social Security"
author:
- name: Max Ghenis
Expand Down
8 changes: 4 additions & 4 deletions docs/benchmark-model-component-matrix.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,7 +33,7 @@ That means:
This chapter covers five comparison objects:

1. **Our plan**
longitudinal `populace` plus PolicyEngine-US plus this Social
longitudinal Microcosm plus PolicyEngine-US plus this Social
Security application layer
2. **DYNASIM**
the main non-governmental dynamic benchmark
Expand Down Expand Up @@ -65,11 +65,11 @@ modern retirement-outcomes model being used for LTSS policy analysis

| Component | Our plan | DYNASIM | MINT | CBO / CBOLT | Morningstar |
|---|---|---|---|---|---|
| **Base population** | Public synthetic `populace`, PolicyEngine's ML-first microdata layer, extended longitudinally | SIPP-based starting sample with publicly documented 0.04% core and 0.4% expanded variants [@favreault2015; @urban2024dynasim4] | SIPP plus administrative earnings and program records for strong near-retirement credibility [@smith2010mint; @smith2021mint8; @ssa2024mint] | SSA Continuous Work History Sample foundation, with a 1-in-1,000 representative microsimulation sample and SIPP/CPS imputations for missing demographics and family structure [@cbo2018; @cbo2019replacementrates] | Household-oriented retirement simulation using current resources, projected longevity, healthcare, and retirement assets; public materials emphasize model outputs more than raw starting-file mechanics [@look2024retirementoutcomes; @morningstar2024modelpage] |
| **Historical earnings** | Synthetic reconstruction from panel data, calibrated against `populace`'s registry of administrative targets; benchmarked across candidate model families | Public record shows lifetime economic histories built from survey and linked administrative inputs, with annual updating and alignment [@favreault2015; @urban2024dynasim4] | Strongest public benchmark because administrative earnings are built in for many cohorts [@smith2010mint; @ssa2024mint] | Public record discusses lifetime earnings assumptions and fiscal outputs, but is much less explicit on the micro history-construction machinery [@cbo2004; @cbo2018; @cbo2024finances] | Public papers say the model estimates historical wages for each household member and then simulates accumulation and retirement adequacy; claim age is simplified in the inaugural analysis [@look2024retirementoutcomes] |
| **Base population** | Public synthetic Microcosm, PolicyEngine's ML-first microdata layer, extended longitudinally | SIPP-based starting sample with publicly documented 0.04% core and 0.4% expanded variants [@favreault2015; @urban2024dynasim4] | SIPP plus administrative earnings and program records for strong near-retirement credibility [@smith2010mint; @smith2021mint8; @ssa2024mint] | SSA Continuous Work History Sample foundation, with a 1-in-1,000 representative microsimulation sample and SIPP/CPS imputations for missing demographics and family structure [@cbo2018; @cbo2019replacementrates] | Household-oriented retirement simulation using current resources, projected longevity, healthcare, and retirement assets; public materials emphasize model outputs more than raw starting-file mechanics [@look2024retirementoutcomes; @morningstar2024modelpage] |
| **Historical earnings** | Synthetic reconstruction from panel data, calibrated against Microcosm's registry of administrative targets; benchmarked across candidate model families | Public record shows lifetime economic histories built from survey and linked administrative inputs, with annual updating and alignment [@favreault2015; @urban2024dynasim4] | Strongest public benchmark because administrative earnings are built in for many cohorts [@smith2010mint; @ssa2024mint] | Public record discusses lifetime earnings assumptions and fiscal outputs, but is much less explicit on the micro history-construction machinery [@cbo2004; @cbo2018; @cbo2024finances] | Public papers say the model estimates historical wages for each household member and then simulates accumulation and retirement adequacy; claim age is simplified in the inaugural analysis [@look2024retirementoutcomes] |
| **Family structure** | Explicit relationship-history layer with spouse links, widowhood, divorce duration, remarriage, and benefit-facing auxiliary states | Public documentation shows marriage, divorce, family structure, and spouse-related states are part of the annual simulation [@favreault2015; @urban2024dynasim4] | Public methodology supports spouse and survivor benefit analysis, but the public record is less explicit than DYNASIM on relationship-history mechanics [@smith2010mint; @ssa2024mint] | Public record is relatively thin on family-history construction at the record level [@cbo2004; @cbo2018] | Public outputs are household-based and broken out by family status, but the public record does not suggest a fully general spouse-former-spouse-child network like the one needed for detailed auxiliary-benefit analysis [@morningstar2024modelpage; @look2024retirementoutcomes] |
| **Disability and health** | Separate impairment, program-pathway, and claiming states, plus mortality and family interactions | Publicly documented health, disability, cognition, and work-limitation modules with yearly transitions [@favreault2015; @urban2024dynasim4] | Includes disability pathways but with publicly documented simplifications around adjudication and return-to-work rules [@ssa2024mint] | Public record is strong on aggregate Social Security finances and disability spending, weaker on record-level disability-state machinery [@cbo2024finances; @cbo2024longterm] | Public papers explicitly include healthcare costs, projected longevity, and LTSS states such as home healthcare and nursing home need, but not a public SSDI-style program pathway [@look2024retirementoutcomes; @look2025ltss] |
| **Wealth, assets, and LTSS** | Not phase-1 core, but preserved as a later extension track through longitudinal `populace` | Major documented strength: wealth, pensions, health spending, LTSS use, payer assignment, and Medicaid interaction [@favreault2015; @urban2024dynasim4; @favreault2020ltss] | Stronger than a Social Security-only model on pensions and SSI interactions, but not positioned publicly as a leading LTSS model [@smith2010mint; @ssa2024mint] | Public emphasis is fiscal outlook rather than household adequacy, wealth depletion, or LTSS risk pathways [@cbo2024finances; @cbo2024longterm] | Major strength: retirement assets, expenses, projected inadequacy, and recent LTSS and WISH analyses using the same model family [@look2024retirementoutcomes; @look2025ltss; @look2025wish] |
| **Wealth, assets, and LTSS** | Not phase-1 core, but preserved as a later extension track through longitudinal Microcosm | Major documented strength: wealth, pensions, health spending, LTSS use, payer assignment, and Medicaid interaction [@favreault2015; @urban2024dynasim4; @favreault2020ltss] | Stronger than a Social Security-only model on pensions and SSI interactions, but not positioned publicly as a leading LTSS model [@smith2010mint; @ssa2024mint] | Public emphasis is fiscal outlook rather than household adequacy, wealth depletion, or LTSS risk pathways [@cbo2024finances; @cbo2024longterm] | Major strength: retirement assets, expenses, projected inadequacy, and recent LTSS and WISH analyses using the same model family [@look2024retirementoutcomes; @look2025ltss; @look2025wish] |

## Matrix 2: benefit logic, behavior, and policy use

Expand Down
4 changes: 2 additions & 2 deletions docs/calibration-targets.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,11 +2,11 @@

## Overview

Calibration ensures that longitudinal `populace` matches known
Calibration ensures that longitudinal Microcosm matches known
population characteristics and that the Social Security application
layer matches system aggregates. This chapter specifies the targets
the project will use for calibration, their sources, and priority
weighting. `populace` maintains these administrative aggregates as a
weighting. Microcosm maintains these administrative aggregates as a
versioned target registry — signed facts with standard errors — so
that calibration runs against a single, consistent set of targets and
treats them as uncertainty-weighted evidence rather than exact hits.
Expand Down
12 changes: 6 additions & 6 deletions docs/data-sources.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,12 +3,12 @@
## Overview

Building a dynamic Social Security microsimulation model means
extending `populace` longitudinally and then using it for Social
extending Microcosm longitudinally and then using it for Social
Security analysis. That requires multiple data sources that capture
cross-sectional population characteristics, longitudinal earnings
dynamics, and demographic transitions.

PolicyEngine's `populace` stack assembles these primary sources,
PolicyEngine's Microcosm stack assembles these primary sources,
along with the administrative aggregates used as calibration targets,
and builds a calibrated synthetic population from them. This chapter
describes the primary sources that feed that pipeline.
Expand Down Expand Up @@ -42,7 +42,7 @@ describes the primary sources that feed that pipeline.
- Limited earning history (only current year)

**Our Use**:
- Core cross-sectional input to the current public `populace`
- Core cross-sectional input to the current public Microcosm
population layer
- Validation of age-earnings profiles
- Calibration targets for population characteristics
Expand Down Expand Up @@ -75,7 +75,7 @@ describes the primary sources that feed that pipeline.
- Public use files have restricted geographic detail

**Our Use**:
- **Primary source for longitudinal extension of `populace`**
- **Primary source for longitudinal extension of Microcosm**
- Training data for quantile regression forests
- Validation of lifetime earnings distributions
- Demographic transition modeling
Expand Down Expand Up @@ -336,7 +336,7 @@ state LTC pilot.

One reason LTC is hard to model is that no single public dataset adequately covers household populations, caregivers, and institutional residents at the same time. CPS and many other core household surveys exclude most institutional populations. An LTC-ready architecture therefore needs an explicit blended strategy:

1. Household base population from Populace's calibrated CPS-based core and allied surveys
1. Household base population from Microcosm's calibrated CPS-based core and allied surveys
2. Longitudinal aging and wealth dynamics from PSID and HRS
3. Care-need and caregiving detail from NHATS/NSOC and MCBS
4. Institutional population benchmarks from MDS and Medicaid administrative sources
Expand All @@ -355,7 +355,7 @@ Rather than treating each survey in isolation, we pursue a multi-survey fusion s

Our data integration follows a hierarchical structure:

1. **Base population**: Populace's CPS-based core providing a large sample with calibrated cross-sectional income
1. **Base population**: Microcosm's CPS-based core providing a large sample with calibrated cross-sectional income
2. **Longitudinal structure**: PSID for earnings trajectories and transition dynamics
3. **Income detail**: PUF for tax return variables and high-income tail corrections
4. **Validation**: SIPP for program participation; administrative aggregates
Expand Down
16 changes: 8 additions & 8 deletions docs/evaluation-and-model-selection.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ enough to justify the next stage of the build. This chapter therefore
defines the evaluation framework for deciding:

- which earnings architecture becomes the production path for
longitudinal `populace`
longitudinal Microcosm
- whether the resulting panel is good enough for benefit calculation
- whether the full project has earned the right to advance from stage 1
to stage 2
Expand Down Expand Up @@ -141,7 +141,7 @@ person-years.

### Cross-sectional anchor tests

Because the final use case starts from a cross-sectional `populace`
Because the final use case starts from a cross-sectional Microcosm
record, the project should also simulate that workflow directly:

1. collapse a held-out panel person to a pseudo-cross-section at a
Expand Down Expand Up @@ -274,7 +274,7 @@ These may include:
- zero-fraction error
- correlation preservation

They are useful as diagnostics, especially for comparing `populace`
They are useful as diagnostics, especially for comparing Microcosm
candidate families, but they are not the final decision rule.

### Operational metrics
Expand Down Expand Up @@ -317,7 +317,7 @@ For candidates that clear Gate 1, score them on:
- policy-output fit
- stability
- runtime and reproducibility
- architectural alignment with longitudinal `populace`
- architectural alignment with longitudinal Microcosm

The scorecard should be reported as a table, not just prose.

Expand All @@ -330,7 +330,7 @@ The winning architecture should be the one that:
3. is simple enough to explain and maintain publicly

That rule leaves open whether the winner is ZI-QDNN, ZI-MAF, a broader
`populace` sequence model, or a more transparent annual-state process.
Microcosm sequence model, or a more transparent annual-state process.

## Suggested numeric thresholds for stage 1

Expand Down Expand Up @@ -381,14 +381,14 @@ decision:
- which architecture deserves continued investment
- what the residual limitations are even if the answer is "yes"

## Relationship to the refreshed Populace evaluations
## Relationship to the refreshed Microcosm evaluations

The `populace` imputation evaluations should feed directly into this
The Microcosm imputation evaluations should feed directly into this
chapter, but they should not be the only evidence.

The right interpretation is:

- refreshed `populace` evals help narrow the candidate set
- refreshed Microcosm evals help narrow the candidate set
- Social-Security-specific benchmarks decide the production winner
- the proposal should remain architecture-agnostic until both pieces are
in hand
Expand Down
28 changes: 14 additions & 14 deletions docs/funder-summary.md
Original file line number Diff line number Diff line change
@@ -1,9 +1,9 @@
# Concept note: Populace dynamics
# Concept note: Microcosm Dynamics

## What this is

This concept note describes the design of an open, longitudinal
Dynamics layer for `populace`, PolicyEngine's country-agnostic
Dynamics layer for Microcosm, PolicyEngine's country-agnostic
microdata stack — validated first on the U.S. Social Security
system, and built so that every claim it makes can be scored against
reality. Social Security is the proving ground because it is the
Expand All @@ -15,7 +15,7 @@ benefit systems as PolicyEngine's country coverage grows.
The premise is George Box's, taken literally: all models are wrong,
and a model is useful only if it improves predictions. So this
project's product is not a brand-name simulator. The machinery lives
in `populace`, PolicyEngine's open microdata stack; the deliverable
in Microcosm, PolicyEngine's open microdata stack; the deliverable
is a versioned population artifact with a manifest and a public
scorecard; and this repository holds the Social Security application
and the validation program that grades it. Models made their names
Expand Down Expand Up @@ -156,7 +156,7 @@ combination:
backtests, and held-out moments — in place of fidelity-only
validation
- domains of validity as shipped metadata on every output
- a contribution rule inherited from `populace`: changes merge if
- a contribution rule inherited from Microcosm: changes merge if
and only if they improve the score on held-out facts, from any
contributor
- AI-callable interfaces from day one
Expand All @@ -167,8 +167,8 @@ No equivalent bundle exists for U.S. Social Security analysis.

The natural implementation is the PolicyEngine open-source stack.

**Populace** is PolicyEngine's rebuilt, open-source microdata stack
([github.com/PolicyEngine/populace](https://github.com/PolicyEngine/populace),
**Microcosm** is PolicyEngine's rebuilt, open-source microdata stack
([github.com/PolicyEngine/microcosm](https://github.com/PolicyEngine/microcosm),
MIT). It builds a calibrated synthetic population entirely from
primary-source government data (CPS/ASEC, IRS Public Use File,
Survey of Consumer Finances, SIPP, CPS outgoing-rotation groups,
Expand All @@ -179,7 +179,7 @@ PolicyEngine's enhanced CPS as the certified default U.S. microdata
in policyengine.py, after a matched, symmetric-refit comparison on
41,314 households with a 739-target holdout:

| Metric (lower is better) | Populace | enhanced CPS |
| Metric (lower is better) | Microcosm | enhanced CPS |
|---|---|---|
| Holdout loss (739 held-out targets) | 0.038 | 0.317 |
| Training loss | 0.190 | 1.089 |
Expand All @@ -188,13 +188,13 @@ in policyengine.py, after a matched, symmetric-refit comparison on

The asymmetry in the last row is published deliberately: the
enhanced CPS wins more individual targets narrowly, while its
largest misses are far larger — Populace's aggregate loss is an
largest misses are far larger — Microcosm's aggregate loss is an
order of magnitude lower on held-out targets. Publishing the number
that cuts against the headline is the discipline this whole project
runs on. (Source: the release manifest in the Populace repository.)
runs on. (Source: the release manifest in the Microcosm repository.)

**The longitudinal extension is designed, not improvised.**
Populace's charter names this project's direction explicitly and
Microcosm's charter names this project's direction explicitly and
specifies the kernel rules: one weight per trajectory, with
multi-period targets stacked as (target, period) constraint rows
over the same weight vector; entry and exit markers (birth, death,
Expand All @@ -212,12 +212,12 @@ forward are the same operator run in either direction.

**PolicyEngine-US** supplies the rules engine — OASDI benefit
calculation, benefit taxation, and means-tested interactions —
through Populace's rules-engine adapter, with Axiom's rules layer as
through Microcosm's rules-engine adapter, with Axiom's rules layer as
the next adapter: statute encoded declaratively and compiled to
Rust, a performance boundary that matters when benefit formulas run
over person-periods across hundreds of thousands of trajectories.
In that architecture PolicyEngine is a composition — Axiom rules,
Populace population, and a labeled behavioral scenario layer.
Microcosm population, and a labeled behavioral scenario layer.
**PolicyEngine-API** and the MCP server are the delivery surface.

The deliverable is a versioned artifact — `populace_us_panel_*` —
Expand Down Expand Up @@ -337,7 +337,7 @@ These are not phase-one commitments. They are reasons to design the
core architecture well.

The longitudinal machinery itself is generic and lives upstream in
`populace`, whose kernel is country-agnostic. The same extension can
Microcosm, whose kernel is country-agnostic. The same extension can
eventually serve other countries' pension and benefit systems;
Social Security is the first application, not the boundary.

Expand Down Expand Up @@ -377,7 +377,7 @@ makes the scorecard, and the case for trusting it, longer.
be co-owned through bilateral institutional agreements.
- Not a 75-year oracle — the long horizon ships as a sensitivity
surface, never a point forecast.
- Not a brand-name simulator — the machinery is Populace's, the
- Not a brand-name simulator — the machinery is Microcosm's, the
artifact is versioned, and the scorecard is the product.

## Open invitation
Expand Down
Loading
Loading