Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -19,3 +19,7 @@
**/_build
!policyengine_uk_data/storage/*.csv
**/version.json

# Build outputs with household-level weights (licensed FRS derivatives); never commit.
policyengine_uk_data/storage/local_geography_weights.csv.gz
.brma-cache/
1 change: 1 addition & 0 deletions changelog.d/brma-census-weights.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- Draw FRS households' Broad Rental Market Areas in proportion to census private-rented households by BRMA and bedrooms, instead of the row counts of the 2019-20 LHA list of rents, whose Scottish, Welsh and Northern Ireland lists were copies of English BRMAs' lists (uk-data#515). Remove `lha_list_of_rents.csv.gz`.
13 changes: 13 additions & 0 deletions docs/imputations.md
Original file line number Diff line number Diff line change
Expand Up @@ -211,6 +211,19 @@ Assigns student loan plan type based on age and reported repayments.

---

## Broad Rental Market Area Assignment

**Source:** Census private-rented households by BRMA and bedrooms (not QRF; applied when the base FRS dataset is built)

The FRS identifies only the region, but Local Housing Allowance rates vary by Broad Rental Market Area (BRMA). `datasets/brma.py` draws each benefit unit's BRMA within its region, in proportion to the private-rented households in each BRMA:
- shared-accommodation and one-bedroom LHA categories use one-bedroom homes;
- the two-, three- and four-or-more-bedroom categories use homes with that many bedrooms;
- Northern Ireland's census has no bedrooms, so its weights are the same for every category.

A household takes one of its benefit units' BRMAs, chosen at random. Sources, method and validation are in `storage/BRMA_DATA_SOURCES.md`.

---

## Calibration Targets

After imputation, household weights are calibrated to match aggregate statistics from:
Expand Down
106 changes: 106 additions & 0 deletions policyengine_uk_data/datasets/brma.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,106 @@
"""Broad Rental Market Area (BRMA) assignment for FRS benefit units.

The FRS identifies only a household's region, but Local Housing Allowance
rates vary by BRMA. Each benefit unit is given a BRMA drawn within its
region, in proportion to the number of private-rented households in each
BRMA with the number of bedrooms matching its LHA category (shared
accommodation and one-bedroom categories both use one-bedroom homes).

The counts come from the censuses (England and Wales 2021, Scotland 2022,
Northern Ireland 2021) mapped to BRMAs; ``storage/BRMA_DATA_SOURCES.md``
gives the sources, method and validation. A BRMA that crosses a region
boundary appears once per region, holding the households on that side.

These weights replace the row counts of the 2019-20 LHA list of rents. Its
Scottish, Welsh and Northern Ireland rows were copies of English BRMAs'
lists, so their counts said nothing about those nations' rental markets.
"""

from pathlib import Path

import numpy as np
import pandas as pd

from policyengine_uk_data.storage import STORAGE_FOLDER

BRMA_HOUSEHOLDS_PATH = STORAGE_FOLDER / "brma_private_rented_households.csv"

# Census bedroom band whose private-rented households weight each LHA category.
LHA_CATEGORY_BEDROOMS = {"A": "1", "B": "1", "C": "2", "D": "3", "E": "4+"}


def load_brma_weights(path: Path = BRMA_HOUSEHOLDS_PATH) -> pd.DataFrame:
"""Return BRMA sampling weights by region and LHA category.

Columns: ``region``, ``lha_category``, ``brma``, ``weight`` (private-rented
households). Every row has a positive weight. Northern Ireland's census
has no bedrooms question, so its rows (bedrooms ``all``) weight every
category.
"""
households = pd.read_csv(path, dtype={"bedrooms": str})
bands = pd.DataFrame(
LHA_CATEGORY_BEDROOMS.items(), columns=["lha_category", "bedrooms"]
)
by_band = bands.merge(households, on="bedrooms")
any_band = households[households.bedrooms == "all"].merge(
bands[["lha_category"]], how="cross"
)
weights = pd.concat([by_band, any_band], ignore_index=True)
weights = weights[weights.households > 0]
return weights.rename(columns={"households": "weight"})[
["region", "lha_category", "brma", "weight"]
].reset_index(drop=True)


def assign_brmas(
region: np.ndarray,
lha_category: np.ndarray,
rng: np.random.Generator,
weights: pd.DataFrame | None = None,
) -> np.ndarray:
"""Draw a BRMA for each benefit unit from its region × LHA category cell.

Args:
region: Region name of each benefit unit's household.
lha_category: LHA category (A-E) of each benefit unit.
rng: Generator for the draws.
weights: Output of ``load_brma_weights``; loaded from storage if omitted.

Raises:
ValueError: If a benefit unit's cell has no BRMA with a positive weight.
"""
if weights is None:
weights = load_brma_weights()
if hasattr(region, "decode_to_str"): # EnumArray of region codes
region = region.decode_to_str()
region = np.asarray(region).astype(str)
lha_category = np.asarray(lha_category).astype(str)
brma = np.empty(len(region), dtype=object)
for (cell_region, category), cell in weights.groupby(
["region", "lha_category"], sort=True
):
mask = (region == cell_region) & (lha_category == category)
if mask.any():
p = cell.weight.to_numpy(dtype=float)
brma[mask] = rng.choice(
cell.brma.to_numpy(), size=mask.sum(), p=p / p.sum()
)
missing = pd.isna(brma)
if missing.any():
cells = sorted(set(zip(region[missing], lha_category[missing])))
raise ValueError(f"No BRMA weights for region × LHA category cells {cells}.")
return brma


def pick_household_brmas(
brma: np.ndarray, household_id: np.ndarray, rng: np.random.Generator
) -> pd.Series:
"""Give each household the BRMA of one of its benefit units, chosen at random.

``brma`` and ``household_id`` are per benefit unit; returns a Series
indexed by household ID.
"""
units = pd.DataFrame({"brma": brma, "household_id": household_id})
return units.groupby("household_id").brma.aggregate(
lambda x: x.sample(n=1, random_state=rng).iloc[0]
)
46 changes: 12 additions & 34 deletions policyengine_uk_data/datasets/frs.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@
from policyengine_uk.variables.household.income.employment_status import (
EmploymentStatus,
)
from policyengine_uk_data.datasets.brma import assign_brmas, pick_household_brmas
from policyengine_uk_data.datasets.disability_benefits import (
add_disability_benefit_categories_from_reported_amounts,
add_disability_benefit_flags_from_reported_amounts,
Expand Down Expand Up @@ -1409,43 +1410,20 @@ def determine_education_level(fted_val, typeed2_val, age_val):

sim = Microsimulation(dataset=dataset)
region = sim.populations["benunit"].household("region", dataset.time_period)
lha_category = sim.calculate("LHA_category", year)
brma = np.empty(len(region), dtype=object)

# Sample from a random BRMA in the region, weighted by the number of observations in each BRMA.
# Use a seeded generator so the assignment is reproducible across builds;
# pandas .sample() otherwise draws from the unseeded global numpy RNG.
lha_list_of_rents = pd.read_csv(STORAGE_FOLDER / "lha_list_of_rents.csv.gz")
lha_list_of_rents = lha_list_of_rents.copy()
brma_rng = np.random.default_rng(0)

for possible_region in lha_list_of_rents.region.unique():
for possible_lha_category in lha_list_of_rents.lha_category.unique():
lor_mask = (lha_list_of_rents.region == possible_region) & (
lha_list_of_rents.lha_category == possible_lha_category
)
mask = (region == possible_region) & (lha_category == possible_lha_category)
brma[mask] = lha_list_of_rents[lor_mask].brma.sample(
n=len(region[mask]), replace=True, random_state=brma_rng
)

# Convert benunit-level BRMAs to household-level BRMAs (pick a random one)
lha_category = np.asarray(sim.calculate("LHA_category", year))

df = pd.DataFrame(
{
"brma": brma,
"household_id": sim.populations["benunit"].household(
"household_id", sim.dataset.time_period
),
}
)
# Draw each benefit unit's BRMA in proportion to the private-rented
# households in each of its region's BRMAs with the matching number of
# bedrooms. Use a seeded generator so the assignment is reproducible.
brma_rng = np.random.default_rng(0)
brma = assign_brmas(region, lha_category, brma_rng)

df = df.groupby("household_id").brma.aggregate(
lambda x: x.sample(n=1, random_state=brma_rng).iloc[0]
household_brma = pick_household_brmas(
brma,
sim.populations["benunit"].household("household_id", dataset.time_period),
brma_rng,
)
brmas = df[sim.calculate("household_id")].values

pe_household["brma"] = brmas
pe_household["brma"] = household_brma[sim.calculate("household_id")].values

pe_person = add_disability_benefit_flags_from_reported_amounts(
pe_person,
Expand Down
49 changes: 49 additions & 0 deletions policyengine_uk_data/storage/BRMA_DATA_SOURCES.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,49 @@
# BRMA private-rented households

File: `brma_private_rented_households.csv`, with columns `region, brma, bedrooms, households`. It holds private-rented households (private landlord or letting agency, plus other private rented) by region, Broad Rental Market Area (BRMA) and number of bedrooms (`1`, `2`, `3`, `4+`). Northern Ireland's 2021 census has no bedrooms question, so its rows have `bedrooms = all`. A BRMA that crosses a region boundary has one row per region, for the households on each side.

`datasets/brma.py` uses it to draw each FRS benefit unit's BRMA within its region. LHA categories A and B (shared and one-bedroom) use one-bedroom homes, C uses two-bedroom, D three-bedroom and E four-or-more.

## Sources

| Nation | Households | BRMA geography |
|---|---|---|
| England | Census 2021 (ONS). Private-rented households per LSOA from TS054, split by bedrooms using the LSOA's mix for "private rented or lives rent free" (custom table `hh_tenure_5a` × `number_bedrooms_5a`). This is the only tenure × bedrooms split ONS releases for every LSOA; rent-free households are 0.3-1.2% of the category. | VOA BRMA boundary layer, May 2020. Each LSOA is assigned by its population-weighted centroid. Five border LSOAs fall outside the English layer and take the Welsh BRMA that VOA's LHA lookup returns near their centroids. |
| Wales | As for England. | Rent Officers Wales BRMA layer (2012 boundaries, 2014 release). 11 LSOAs fall in West Cheshire. Cardiff Bay (W01002024) lies outside the layer and is assigned to Cardiff, as VOA's lookup returns for its postcodes. |
| Scotland | Scotland's Census 2022 (NRS), tenure by bedrooms by 2022 electoral ward. This is the finest geography at which the table is not suppressed. | Scottish Government BRMA polygons. Each ward is split across BRMAs by its output areas' household counts, placing each output area by its population-weighted centroid. Rent Service Scotland's postcode-to-BRMA lookup (FOI 202300368850) moves 61 Balloch output areas to West Dunbartonshire: 59 on unique postcode matches and 2 via the council's own lookup. |
| Northern Ireland | NISRA Census 2021 households by postcode district, times Northern Ireland's private-rented share (NISRA tenure by Data Zone, 17.2%). | NIHE's definition of each BRMA as a set of postcode districts. |

Northern Ireland's private-rented households are not split by area within the nation. Linking NISRA's tenure areas to postcode districts would need the ONS Postcode Directory's Northern Ireland records, whose licence (LPS end user licence) does not clearly allow publishing derived figures. Using every BRMA's all-tenure households instead moves BRMA shares by 1.5 percentage points on average; for example, Belfast gets 20.1% of NI's private renters where the private-rented split would give 23.5%.

Totals reconcile with the published national private-rented counts to within 0.03%:

| Nation | Private-rented households in this file | Published |
|---|---|---|
| England | 4,795,158 | 4,794,889 |
| Wales | 228,601 | 228,642 |
| Scotland | 323,001 | 323,042 |
| Northern Ireland | 132,449 | 132,436 |

Small-area census counts are perturbed for disclosure control, so their sums differ slightly from national tables.

## Validation

The check below correlates, within each region, the BRMA shares of each candidate weight with DWP's count of Universal Credit households whose housing costs are assessed under LHA ("LHA covers rent" plus "does not cover rent"). The DWP figures are the mean of April 2019 to November 2020 (UC statistics supplementary table 3.2, February 2021).

To make the comparison, each BRMA's households from this file are summed across bedroom bands and region parts, then placed in policyengine-uk's single region for that BRMA.

| Nation | This file | `lha_list_of_rents.csv.gz` row counts (previous weights) |
|---|---|---|
| England | 0.94 | 0.82 |
| Wales | 0.95 | 0.23 |
| Scotland | 0.82 | 0.21 |

The previous weights' Scottish, Welsh and Northern Ireland lists were copies of English BRMAs' lists (issue #515).

Against the genuine list-of-rents category counts, the census bedroom bands correlate as follows (VOA 2019-20 for England; Scottish Government FOI 202200303624 for Scotland):
- categories B-E: 0.82 to 0.94 in England, 0.95 to 0.98 in Scotland;
- category A, shared accommodation: 0.51 in England, 0.95 in Scotland. The census has no measure of room lets, and no bedroom band does better than 0.59.

## Rebuilding

`tools/brma_households/` (run from the repository root) downloads every source, checks each one against a pinned sha256 and rebuilds this file byte for byte; see its `README.md`.
Loading
Loading