This document explains the architecture and patterns used in the seed data generation system.
The prototype uses generators to create realistic test data for participants, episodes, appointments, and associated medical information. Data generation follows a hierarchical pattern where each generator can be configured, overridden, and tested independently.
The main orchestrator is generate-seed-data.js which calls generators in sequence:
generate-seed-data.js (main orchestrator)
│
├─> participant-generator.js
│ Creates test participants (with overrides from test scenarios)
│ Creates random participants
│ Can create additional participants on-demand during clinic generation
│
├─> clinic-generator.js
│ Creates clinics with slots for each BSU
│
├─> episode-generator.js + appointment-generator.js
│ Every appointment sits inside an episode (the screening round it
│ belongs to). For each filled slot: generateEpisode() creates the
│ episode first, then generateAppointment() creates the appointment,
│ and the two are linked both ways (episode.appointmentIds /
│ appointment.episodeId)
│ │
│ └─> medical-information/ generators (for completed and in-progress
│ appointments)
│ ├─> symptoms-generator.js
│ ├─> breast-density-factors-generator.js
│ ├─> other-medical-information-generator.js
│ ├─> breast-features-generator.js
│ └─> medical-history-generator.js
│
├─> reading-generator.js
│ Adds reading/outcome data to appointments
│ finaliseEpisodeStage() then sets each episode's stage/outcome from
│ its appointments and reads
│
├─> episode-generator.js (historic episodes)
│ generateHistoricEpisodes() adds summary-level past screening rounds
│ per participant (outcome-first: dates, outcome and a
│ mammogramSummary — no appointments or reads).
│ checkEpisodes() validates the combined result
│
└─> issue-generator.js
generateIssues() seeds open and resolved issues on current episodes
(see docs/issues.md). checkIssues() confirms every link resolves
Key points:
- Generators are called sequentially by the main orchestrator
- Every appointment belongs to an episode; episodes are created first
- Medical information is generated for completed and in-progress appointments
- Participants are mostly created upfront, but can be generated on-demand
- Reading data is added after all appointments are created, then episode stages are finalised
- Issues are generated last, from the finished episodes, and written to their own
issues.json - Historic episodes are summary-level only — see
docs/data-conventions.mdfor the episode data model
Participant level:
- Basic demographics
- NHS number, GP details
- Persistent information
Episode level:
- The screening round: stage (
scheduled|mammograms|reading|assessment|closed), stage history, final outcome - Links to its appointments (
appointmentIds) - Reading cases (
readingCases), one per set of images taken — the reads live here - Historic episodes: summary-level record of a past round (dates, outcome,
mammogramSummary) with no appointments
Appointment level:
- Medical information collected during appointments
- Session details (who, when, where)
- Mammogram images (
mammogramData) and prior mammograms - Appointment-specific data
Key principle: Medical information is stored at the appointment level, representing data collected during specific appointments. Round-level state (stage, outcome) and the reading of each image set live on the episode.
Because reads are written onto cases, generation runs in that order: episodes
open a case per image set (syncEpisodeReadingCases), then generateReadingData
writes reads into them, then finaliseEpisodeStage settles each episode's stage
from its latest case.
// app/lib/generators/[name]-generator.js
const { faker } = require('@faker-js/faker')
const weighted = require('weighted')
const generateId = require('../utils/id-generator')
/**
* Generate a single item
* @param {object} [options] - Generation options
* @returns {object} Generated item
*/
const generateItem = (options = {}) => {
return {
id: generateId(),
// ... fields here
dateAdded: new Date().toISOString(),
addedByUserId: options.addedByUserId || null
}
}
/**
* Generate multiple items with probability controls
* @param {object} [options] - Generation options
* @param {number} [options.probability] - Chance of having any items (0-1)
* @param {number} [options.maxItems] - Maximum number of items
* @returns {Array} Array of generated items
*/
const generateItems = (options = {}) => {
const { probability = 0.15, maxItems = 3 } = options
// Layer 1: Do they have any at all?
if (Math.random() > probability) {
return []
}
// Layer 2: How many?
const numberOfItems = weighted.select({
1: 0.7,
2: 0.2,
3: 0.1
})
// Generate items
return Array.from({ length: Math.min(numberOfItems, maxItems) }, () =>
generateItem(options)
)
}
module.exports = {
generateItem,
generateItems
}- Export both singular and plural functions - Flexibility for different use cases
- Accept options parameter - Enables configuration and overrides
- Use weighted randomization - Creates realistic distributions
- Use faker for realistic data - Convincing names, dates, text
- Generate IDs consistently - Use shared
generateId()utility
Use multiple probability layers to create realistic distributions:
// Layer 1: Overall probability - do they have ANY medical history?
if (Math.random() > 0.20) return {}
// Layer 2: Which types do they have?
const numberOfTypes = weighted.select({
1: 0.6,
2: 0.3,
3: 0.1
})
// Layer 3: How many of each type?
const numberOfItems = weighted.select({
1: 0.8,
2: 0.15,
3: 0.05
})const weighted = require('weighted')
// Select one item based on weights
const choice = weighted.select({
'option1': 0.7, // 70% chance
'option2': 0.2, // 20% chance
'option3': 0.1 // 10% chance
})
// Select multiple items
const items = weighted.select(
['item1', 'item2', 'item3'],
[0.5, 0.3, 0.2],
2 // select 2 items
)Set default probabilities in the generator, allow overrides via options:
const generateMedicalInformation = (options = {}) => {
const {
probabilityOfSymptoms = 0.85, // Default: 85%
probabilityOfHRT = 0.30, // Default: 30%
config // May contain overrides
} = options
// Use defaults unless overridden
const symptoms = generateSymptoms({
probability: probabilityOfSymptoms,
...options
})
}Test scenarios live in app/data/test-scenarios.js and allow creating specific test cases:
// app/data/test-scenarios.js
module.exports = [
{
participant: {
id: 'test-scenario-001',
demographicInformation: {
firstName: 'Sarah',
lastName: 'Test'
},
config: {
// Appointment configuration (ids are opaque strings, e.g. '5gpn41oi')
appointmentId: 'abc12345',
scheduling: {
whenRelativeToToday: 0,
status: 'complete'
},
// Generator overrides
forceMedicalHistoryTypes: ['breastCancer'],
probabilityOfSymptoms: 1.0 // Force symptoms
}
}
}
]Pattern 1: Complete Override - Replace all generated data
const generateItems = (options = {}) => {
// Allow complete override from test scenarios
if (options.items) {
return options.items
}
// Normal generation...
}Pattern 2: Forced Inclusion - Guarantee specific items appear
const generateItems = (options = {}) => {
const { forceTypes = [] } = options
const items = []
// First, generate forced types
forceTypes.forEach(type => {
items.push(generateItem({ type }))
})
// Then add random types if desired
if (Math.random() < someProbability) {
// Add additional items...
}
return items
}Pattern 3: Probability Override - Control likelihood
const generateItems = (options = {}) => {
const {
probability = 0.15, // Default
config
} = options
// Config can override probability
const actualProbability = config?.probabilityOverride ?? probability
if (Math.random() > actualProbability) {
return []
}
// Generate items...
}Configuration flows from test scenarios through the generator hierarchy:
// In appointment-generator.js
const appointment = generateAppointment({
participant, // Contains participant.config
// ...
})
// Inside generateAppointment, for completed appointments:
const medicalInfo = generateMedicalInformation({
addedByUserId: appointment.sessionDetails.startedBy,
config: participant.config // Pass config down
})
// In medical-information-generator.js
const symptoms = generateSymptoms({
...options,
forceSymptomTypes: config?.forceSymptomTypes // Use config
})For complex data with multiple sub-generators, use an umbrella generator to orchestrate:
// app/lib/generators/medical-information-generator.js
const { generateSymptoms } = require('./medical-information/symptoms-generator')
const {
generateBreastDensityFactors
} = require('./medical-information/breast-density-factors-generator')
// ... other generators
const generateMedicalInformation = (options = {}) => {
const {
addedByUserId,
probabilityOfSymptoms = 0.85,
probabilityOfHRT = 0.30,
config
} = options
const medicalInfo = {}
// Generate each component
const symptoms = generateSymptoms({
probability: probabilityOfSymptoms,
addedByUserId,
forceTypes: config?.forceSymptomTypes
})
if (symptoms.length > 0) {
medicalInfo.symptoms = symptoms
}
// Sub-generators can return several keys at once, merged into the parent object
const breastDensityFactors = generateBreastDensityFactors({
probabilityOfHrt: probabilityOfHRT
})
Object.assign(medicalInfo, breastDensityFactors)
// Only return if we generated anything
return medicalInfo
}Benefits:
- Single point of orchestration
- Default probabilities in one place
- Simplified integration with parent generators
- Easy to add new sub-generators
// Track used items to avoid duplicates
const usedTypes = new Set()
const items = []
while (items.length < numberOfItems) {
const availableTypes = allTypes.filter(type => !usedTypes.has(type))
if (availableTypes.length === 0) break
const type = weighted.select(weightsByType(availableTypes))
items.push(generateItem({ type }))
usedTypes.add(type)
}Medical information should be attributed to the user who collected it:
// In appointment-generator.js
if (isCompleted(appointmentStatus)) {
const medicalInfo = generateMedicalInformation({
addedByUserId: appointment.sessionDetails.startedBy // Who ran appointment
})
}
// In sub-generators
const generateItem = (options = {}) => {
return {
id: generateId(),
// ...
dateAdded: new Date().toISOString(),
addedByUserId: options.addedByUserId || null
}
}Only include fields when relevant:
const item = {
id: generateId(),
status: weighted.select({
'yes': 0.4,
'no-recently-stopped': 0.3,
'no': 0.3
})
}
// Conditional fields based on status
if (item.status === 'yes') {
item.duration = faker.helpers.arrayElement(['6 months', '2 years', '5 years'])
} else if (item.status === 'no-recently-stopped') {
item.durationSinceStopped = faker.helpers.arrayElement(['two weeks ago', 'one month ago'])
item.durationBeforeStopping = faker.helpers.arrayElement(['3 years', '5 years'])
}const dayjs = require('dayjs')
// Random past date
const date = faker.date.past({ years: 5 })
// Year only
const year = faker.number.int({ min: 2015, max: 2024 }).toString()
// Month and year object
const monthYear = {
month: faker.number.int({ min: 1, max: 12 }),
year: faker.number.int({ min: 2018, max: 2024 })
}
// Relative date
const relativeDate = faker.helpers.arrayElement([
'about 3 months ago',
'earlier this year',
'last summer'
])For user research testing:
- Higher probabilities to ensure features appear
- 30-40% chance of medical history (vs realistic 15-25%)
- More edge cases than would occur naturally
For stakeholder demos:
- Lower, more realistic probabilities
- Reflects actual NHS screening population
// Testing mode - high probability
const TESTING_MODE = {
probabilityOfSymptoms: 0.85, // 85%
probabilityOfMedicalHistory: 0.40 // 40%
}
// Realistic mode - lower probability
const REALISTIC_MODE = {
probabilityOfSymptoms: 0.15, // 15%
probabilityOfMedicalHistory: 0.20 // 20%
}Use test scenarios to create specific high-probability cases when needed.
app/lib/generators/
├── participant-generator.js # Generates participant records
├── appointment-generator.js # Generates appointment records
├── medical-information-generator.js # Umbrella generator
├── medical-information/ # Sub-generators
│ ├── symptoms-generator.js
│ ├── breast-density-factors-generator.js
│ └── medical-history-generator.js
└── [new]-generator.js # New generators here
Naming conventions:
generateItem()- Single itemgenerateItems()- Array of itemsgenerate[Type]()- Specific type/variant
const generateId = require('../utils/id-generator')
const item = {
id: generateId(), // Generates unique ID
// ...
}const { faker } = require('@faker-js/faker')
// Names
faker.person.firstName()
faker.person.lastName()
// Dates
faker.date.past({ years: 5 })
faker.date.future({ years: 2 })
// Selection
faker.helpers.arrayElement(['option1', 'option2', 'option3'])
faker.helpers.arrayElements(['a', 'b', 'c'], { min: 1, max: 2 })
// Text
faker.lorem.sentence()
faker.lorem.paragraph()# Generate seed data
node app/lib/generate-seed-data.js
# Generated data stored in
app/data/generated/Generated files are gitignored and created on-demand when the app starts.
// 1. Create test participants with overrides
const testParticipants = testScenarios.map((scenario) => {
return generateParticipant({
ethnicities,
breastScreeningUnits,
overrides: scenario.participant // Test scenario overrides
})
})
// 2. Create random participants
const randomParticipants = Array.from(
{ length: config.generation.numberOfParticipants },
() => generateParticipant({ ethnicities, breastScreeningUnits })
)
// 3. Generate clinics for each BSU and date
const clinics = generateClinicsForBSU({
date: date.toDate(),
breastScreeningUnit: unit
})
// 4. Generate appointments by allocating participants to slots
const appointment = generateAppointment({
slot,
participant,
clinic,
outcomeWeights: config.screening.outcomes[clinic.clinicType],
forceStatus: scenario?.participant?.config?.scheduling?.status
})
// 5. Generate reading data after all appointments created
const appointmentsWithReadingData = generateReadingData(
sortedAppointments,
users
)Medical information is generated for completed appointments (and again, similarly, for today's in-progress appointment):
// In appointment-generator.js
if (isCompleted(appointmentStatus)) {
// Generate medical information (symptoms, medical history, etc.)
const medicalInformation = generateMedicalInformation({
addedByUserId: appointment.sessionDetails.startedBy,
config: participant.config,
// Allow config to override probabilities
...(participant.config?.medicalInformation || {})
})
if (Object.keys(medicalInformation).length > 0) {
appointment.medicalInformation = medicalInformation
}
}If no participants are available for a slot, a new one is created:
// In generate-seed-data.js, inside generateClinicsForDay()
if (availableParticipants.length === 0) {
const newParticipant = generateParticipant({
ethnicities,
breastScreeningUnits: [unit],
riskLevel: selectedRiskLevel
})
participants.push(newParticipant)
availableParticipants.push(newParticipant)
}- Don't store appointment data on participants - Medical information goes on appointments
- Check
canHaveMultipleflags - Some types limited to single entry - Pass config through - Test scenario config must flow to sub-generators
- Attribute to correct user - Use
sessionDetails.startedByfor medical info - Layer probabilities - Avoid too many participants with rare conditions
- Match form structure - Generated data must match what forms expect
The generator system provides:
- Hierarchical structure - Generators call sub-generators
- Flexible configuration - Defaults, overrides, test scenarios
- Realistic distributions - Weighted probabilities and faker data
- Testability - Force specific cases when needed
- Maintainability - Consistent patterns and file organization
When creating new generators, follow these patterns for consistency and maintainability.