Skip to content

Latest commit

 

History

History
674 lines (535 loc) · 18 KB

File metadata and controls

674 lines (535 loc) · 18 KB

Data Generator Reference

This document explains the architecture and patterns used in the seed data generation system.

Overview

The prototype uses generators to create realistic test data for participants, episodes, appointments, and associated medical information. Data generation follows a hierarchical pattern where each generator can be configured, overridden, and tested independently.

Architecture

Data Flow

The main orchestrator is generate-seed-data.js which calls generators in sequence:

generate-seed-data.js (main orchestrator)
  │
  ├─> participant-generator.js
  │     Creates test participants (with overrides from test scenarios)
  │     Creates random participants
  │     Can create additional participants on-demand during clinic generation
  │
  ├─> clinic-generator.js
  │     Creates clinics with slots for each BSU
  │
  ├─> episode-generator.js + appointment-generator.js
  │     Every appointment sits inside an episode (the screening round it
  │     belongs to). For each filled slot: generateEpisode() creates the
  │     episode first, then generateAppointment() creates the appointment,
  │     and the two are linked both ways (episode.appointmentIds /
  │     appointment.episodeId)
  │     │
  │     └─> medical-information/ generators (for completed and in-progress
  │           appointments)
  │           ├─> symptoms-generator.js
  │           ├─> breast-density-factors-generator.js
  │           ├─> other-medical-information-generator.js
  │           ├─> breast-features-generator.js
  │           └─> medical-history-generator.js
  │
  ├─> reading-generator.js
  │     Adds reading/outcome data to appointments
  │     finaliseEpisodeStage() then sets each episode's stage/outcome from
  │     its appointments and reads
  │
  ├─> episode-generator.js (historic episodes)
  │     generateHistoricEpisodes() adds summary-level past screening rounds
  │     per participant (outcome-first: dates, outcome and a
  │     mammogramSummary — no appointments or reads).
  │     checkEpisodes() validates the combined result
  │
  └─> issue-generator.js
        generateIssues() seeds open and resolved issues on current episodes
        (see docs/issues.md). checkIssues() confirms every link resolves

Key points:

  • Generators are called sequentially by the main orchestrator
  • Every appointment belongs to an episode; episodes are created first
  • Medical information is generated for completed and in-progress appointments
  • Participants are mostly created upfront, but can be generated on-demand
  • Reading data is added after all appointments are created, then episode stages are finalised
  • Issues are generated last, from the finished episodes, and written to their own issues.json
  • Historic episodes are summary-level only — see docs/data-conventions.md for the episode data model

Storage Locations

Participant level:

  • Basic demographics
  • NHS number, GP details
  • Persistent information

Episode level:

  • The screening round: stage (scheduled | mammograms | reading | assessment | closed), stage history, final outcome
  • Links to its appointments (appointmentIds)
  • Reading cases (readingCases), one per set of images taken — the reads live here
  • Historic episodes: summary-level record of a past round (dates, outcome, mammogramSummary) with no appointments

Appointment level:

  • Medical information collected during appointments
  • Session details (who, when, where)
  • Mammogram images (mammogramData) and prior mammograms
  • Appointment-specific data

Key principle: Medical information is stored at the appointment level, representing data collected during specific appointments. Round-level state (stage, outcome) and the reading of each image set live on the episode.

Because reads are written onto cases, generation runs in that order: episodes open a case per image set (syncEpisodeReadingCases), then generateReadingData writes reads into them, then finaliseEpisodeStage settles each episode's stage from its latest case.

Generator Pattern

Standard Structure

// app/lib/generators/[name]-generator.js

const { faker } = require('@faker-js/faker')
const weighted = require('weighted')
const generateId = require('../utils/id-generator')

/**
 * Generate a single item
 * @param {object} [options] - Generation options
 * @returns {object} Generated item
 */
const generateItem = (options = {}) => {
  return {
    id: generateId(),
    // ... fields here
    dateAdded: new Date().toISOString(),
    addedByUserId: options.addedByUserId || null
  }
}

/**
 * Generate multiple items with probability controls
 * @param {object} [options] - Generation options
 * @param {number} [options.probability] - Chance of having any items (0-1)
 * @param {number} [options.maxItems] - Maximum number of items
 * @returns {Array} Array of generated items
 */
const generateItems = (options = {}) => {
  const { probability = 0.15, maxItems = 3 } = options

  // Layer 1: Do they have any at all?
  if (Math.random() > probability) {
    return []
  }

  // Layer 2: How many?
  const numberOfItems = weighted.select({
    1: 0.7,
    2: 0.2,
    3: 0.1
  })

  // Generate items
  return Array.from({ length: Math.min(numberOfItems, maxItems) }, () =>
    generateItem(options)
  )
}

module.exports = {
  generateItem,
  generateItems
}

Key Principles

  1. Export both singular and plural functions - Flexibility for different use cases
  2. Accept options parameter - Enables configuration and overrides
  3. Use weighted randomization - Creates realistic distributions
  4. Use faker for realistic data - Convincing names, dates, text
  5. Generate IDs consistently - Use shared generateId() utility

Probability Management

Layered Probabilities

Use multiple probability layers to create realistic distributions:

// Layer 1: Overall probability - do they have ANY medical history?
if (Math.random() > 0.20) return {}

// Layer 2: Which types do they have?
const numberOfTypes = weighted.select({
  1: 0.6,
  2: 0.3,
  3: 0.1
})

// Layer 3: How many of each type?
const numberOfItems = weighted.select({
  1: 0.8,
  2: 0.15,
  3: 0.05
})

Weighted Selection

const weighted = require('weighted')

// Select one item based on weights
const choice = weighted.select({
  'option1': 0.7,  // 70% chance
  'option2': 0.2,  // 20% chance
  'option3': 0.1   // 10% chance
})

// Select multiple items
const items = weighted.select(
  ['item1', 'item2', 'item3'],
  [0.5, 0.3, 0.2],
  2  // select 2 items
)

Default Probabilities

Set default probabilities in the generator, allow overrides via options:

const generateMedicalInformation = (options = {}) => {
  const {
    probabilityOfSymptoms = 0.85,  // Default: 85%
    probabilityOfHRT = 0.30,       // Default: 30%
    config  // May contain overrides
  } = options

  // Use defaults unless overridden
  const symptoms = generateSymptoms({
    probability: probabilityOfSymptoms,
    ...options
  })
}

Configuration & Overrides

Test Scenarios

Test scenarios live in app/data/test-scenarios.js and allow creating specific test cases:

// app/data/test-scenarios.js
module.exports = [
  {
    participant: {
      id: 'test-scenario-001',
      demographicInformation: {
        firstName: 'Sarah',
        lastName: 'Test'
      },
      config: {
        // Appointment configuration (ids are opaque strings, e.g. '5gpn41oi')
        appointmentId: 'abc12345',
        scheduling: {
          whenRelativeToToday: 0,
          status: 'complete'
        },

        // Generator overrides
        forceMedicalHistoryTypes: ['breastCancer'],
        probabilityOfSymptoms: 1.0  // Force symptoms
      }
    }
  }
]

Override Patterns

Pattern 1: Complete Override - Replace all generated data

const generateItems = (options = {}) => {
  // Allow complete override from test scenarios
  if (options.items) {
    return options.items
  }

  // Normal generation...
}

Pattern 2: Forced Inclusion - Guarantee specific items appear

const generateItems = (options = {}) => {
  const { forceTypes = [] } = options
  const items = []

  // First, generate forced types
  forceTypes.forEach(type => {
    items.push(generateItem({ type }))
  })

  // Then add random types if desired
  if (Math.random() < someProbability) {
    // Add additional items...
  }

  return items
}

Pattern 3: Probability Override - Control likelihood

const generateItems = (options = {}) => {
  const {
    probability = 0.15,  // Default
    config
  } = options

  // Config can override probability
  const actualProbability = config?.probabilityOverride ?? probability

  if (Math.random() > actualProbability) {
    return []
  }

  // Generate items...
}

Config Flow

Configuration flows from test scenarios through the generator hierarchy:

// In appointment-generator.js
const appointment = generateAppointment({
  participant,  // Contains participant.config
  // ...
})

// Inside generateAppointment, for completed appointments:
const medicalInfo = generateMedicalInformation({
  addedByUserId: appointment.sessionDetails.startedBy,
  config: participant.config  // Pass config down
})

// In medical-information-generator.js
const symptoms = generateSymptoms({
  ...options,
  forceSymptomTypes: config?.forceSymptomTypes  // Use config
})

Umbrella Generator Pattern

For complex data with multiple sub-generators, use an umbrella generator to orchestrate:

// app/lib/generators/medical-information-generator.js

const { generateSymptoms } = require('./medical-information/symptoms-generator')
const {
  generateBreastDensityFactors
} = require('./medical-information/breast-density-factors-generator')
// ... other generators

const generateMedicalInformation = (options = {}) => {
  const {
    addedByUserId,
    probabilityOfSymptoms = 0.85,
    probabilityOfHRT = 0.30,
    config
  } = options

  const medicalInfo = {}

  // Generate each component
  const symptoms = generateSymptoms({
    probability: probabilityOfSymptoms,
    addedByUserId,
    forceTypes: config?.forceSymptomTypes
  })

  if (symptoms.length > 0) {
    medicalInfo.symptoms = symptoms
  }

  // Sub-generators can return several keys at once, merged into the parent object
  const breastDensityFactors = generateBreastDensityFactors({
    probabilityOfHrt: probabilityOfHRT
  })

  Object.assign(medicalInfo, breastDensityFactors)

  // Only return if we generated anything
  return medicalInfo
}

Benefits:

  • Single point of orchestration
  • Default probabilities in one place
  • Simplified integration with parent generators
  • Easy to add new sub-generators

Common Patterns

Avoiding Duplicates

// Track used items to avoid duplicates
const usedTypes = new Set()
const items = []

while (items.length < numberOfItems) {
  const availableTypes = allTypes.filter(type => !usedTypes.has(type))
  if (availableTypes.length === 0) break

  const type = weighted.select(weightsByType(availableTypes))
  items.push(generateItem({ type }))
  usedTypes.add(type)
}

User Attribution

Medical information should be attributed to the user who collected it:

// In appointment-generator.js
if (isCompleted(appointmentStatus)) {
  const medicalInfo = generateMedicalInformation({
    addedByUserId: appointment.sessionDetails.startedBy  // Who ran appointment
  })
}

// In sub-generators
const generateItem = (options = {}) => {
  return {
    id: generateId(),
    // ...
    dateAdded: new Date().toISOString(),
    addedByUserId: options.addedByUserId || null
  }
}

Conditional Fields

Only include fields when relevant:

const item = {
  id: generateId(),
  status: weighted.select({
    'yes': 0.4,
    'no-recently-stopped': 0.3,
    'no': 0.3
  })
}

// Conditional fields based on status
if (item.status === 'yes') {
  item.duration = faker.helpers.arrayElement(['6 months', '2 years', '5 years'])
} else if (item.status === 'no-recently-stopped') {
  item.durationSinceStopped = faker.helpers.arrayElement(['two weeks ago', 'one month ago'])
  item.durationBeforeStopping = faker.helpers.arrayElement(['3 years', '5 years'])
}

Realistic Date Generation

const dayjs = require('dayjs')

// Random past date
const date = faker.date.past({ years: 5 })

// Year only
const year = faker.number.int({ min: 2015, max: 2024 }).toString()

// Month and year object
const monthYear = {
  month: faker.number.int({ min: 1, max: 12 }),
  year: faker.number.int({ min: 2018, max: 2024 })
}

// Relative date
const relativeDate = faker.helpers.arrayElement([
  'about 3 months ago',
  'earlier this year',
  'last summer'
])

Testing & Distributions

Testing-Friendly vs Realistic

For user research testing:

  • Higher probabilities to ensure features appear
  • 30-40% chance of medical history (vs realistic 15-25%)
  • More edge cases than would occur naturally

For stakeholder demos:

  • Lower, more realistic probabilities
  • Reflects actual NHS screening population

Adjusting Probabilities

// Testing mode - high probability
const TESTING_MODE = {
  probabilityOfSymptoms: 0.85,      // 85%
  probabilityOfMedicalHistory: 0.40  // 40%
}

// Realistic mode - lower probability
const REALISTIC_MODE = {
  probabilityOfSymptoms: 0.15,      // 15%
  probabilityOfMedicalHistory: 0.20  // 20%
}

Use test scenarios to create specific high-probability cases when needed.

File Organization

app/lib/generators/
├── participant-generator.js          # Generates participant records
├── appointment-generator.js                # Generates appointment records
├── medical-information-generator.js  # Umbrella generator
├── medical-information/              # Sub-generators
│   ├── symptoms-generator.js
│   ├── breast-density-factors-generator.js
│   └── medical-history-generator.js
└── [new]-generator.js                # New generators here

Naming conventions:

  • generateItem() - Single item
  • generateItems() - Array of items
  • generate[Type]() - Specific type/variant

Utilities

ID Generation

const generateId = require('../utils/id-generator')

const item = {
  id: generateId(),  // Generates unique ID
  // ...
}

Faker

const { faker } = require('@faker-js/faker')

// Names
faker.person.firstName()
faker.person.lastName()

// Dates
faker.date.past({ years: 5 })
faker.date.future({ years: 2 })

// Selection
faker.helpers.arrayElement(['option1', 'option2', 'option3'])
faker.helpers.arrayElements(['a', 'b', 'c'], { min: 1, max: 2 })

// Text
faker.lorem.sentence()
faker.lorem.paragraph()

Running Generators

# Generate seed data
node app/lib/generate-seed-data.js

# Generated data stored in
app/data/generated/

Generated files are gitignored and created on-demand when the app starts.

Integration Points

Main Orchestrator (generate-seed-data.js)

// 1. Create test participants with overrides
const testParticipants = testScenarios.map((scenario) => {
  return generateParticipant({
    ethnicities,
    breastScreeningUnits,
    overrides: scenario.participant  // Test scenario overrides
  })
})

// 2. Create random participants
const randomParticipants = Array.from(
  { length: config.generation.numberOfParticipants },
  () => generateParticipant({ ethnicities, breastScreeningUnits })
)

// 3. Generate clinics for each BSU and date
const clinics = generateClinicsForBSU({
  date: date.toDate(),
  breastScreeningUnit: unit
})

// 4. Generate appointments by allocating participants to slots
const appointment = generateAppointment({
  slot,
  participant,
  clinic,
  outcomeWeights: config.screening.outcomes[clinic.clinicType],
  forceStatus: scenario?.participant?.config?.scheduling?.status
})

// 5. Generate reading data after all appointments created
const appointmentsWithReadingData = generateReadingData(
  sortedAppointments,
  users
)

In appointment-generator.js

Medical information is generated for completed appointments (and again, similarly, for today's in-progress appointment):

// In appointment-generator.js
if (isCompleted(appointmentStatus)) {
  // Generate medical information (symptoms, medical history, etc.)
  const medicalInformation = generateMedicalInformation({
    addedByUserId: appointment.sessionDetails.startedBy,
    config: participant.config,
    // Allow config to override probabilities
    ...(participant.config?.medicalInformation || {})
  })

  if (Object.keys(medicalInformation).length > 0) {
    appointment.medicalInformation = medicalInformation
  }
}

On-Demand Participant Creation

If no participants are available for a slot, a new one is created:

// In generate-seed-data.js, inside generateClinicsForDay()
if (availableParticipants.length === 0) {
  const newParticipant = generateParticipant({
    ethnicities,
    breastScreeningUnits: [unit],
    riskLevel: selectedRiskLevel
  })
  participants.push(newParticipant)
  availableParticipants.push(newParticipant)
}

Common Pitfalls

  1. Don't store appointment data on participants - Medical information goes on appointments
  2. Check canHaveMultiple flags - Some types limited to single entry
  3. Pass config through - Test scenario config must flow to sub-generators
  4. Attribute to correct user - Use sessionDetails.startedBy for medical info
  5. Layer probabilities - Avoid too many participants with rare conditions
  6. Match form structure - Generated data must match what forms expect

Summary

The generator system provides:

  • Hierarchical structure - Generators call sub-generators
  • Flexible configuration - Defaults, overrides, test scenarios
  • Realistic distributions - Weighted probabilities and faker data
  • Testability - Force specific cases when needed
  • Maintainability - Consistent patterns and file organization

When creating new generators, follow these patterns for consistency and maintainability.