Skip to content

Multilayer-grating design workflow: survey + energy scan, with a web tab - #46

Merged
simonevadi merged 15 commits into
developfrom
feature/multilayer-grating-design-workflow
Sep 12, 2026
Merged

simonevadi merged 15 commits into
developfrom
feature/multilayer-grating-design-workflow

Conversation

@simonevadi

Copy link
Copy Markdown
Contributor

Replaces the unreleased three-stage multilayer-optimization workflow with a two-step design workflow, and gives it a full web UI.

The old workflow estimated the incident angle from a CFF value. That input is gone: the angle is now a result of a multilayer theta search run at every grid point, which is what the CFF was standing in for. gamma is held fixed and the anti-blaze angle defaults to 0; both are later optimizations.

Library — grax.MultilayerGratingDesigner

One frozen MultilayerDesignConfig, grouped into three banner-marked sections (shared / survey-only / energy-scan-only) so it is obvious which stage reads what. The rough/fine/final theta-search settings are a separate ThetaSearchScanSettings held on two independent fields, survey_scan_settings and energy_scan_settings — the survey runs many cheap single-energy searches, the scan runs few designs over many energies, and they want different budgets.

  • run_survey — a d-spacing × blaze-angle grid; every cell keeps the complete theta-search artifact bundle under survey/runs/d<d>nm/blaze<b>deg/ plus a search_parameters.json, so a suspicious cell can be inspected and the settings retuned. Headline outputs: optimal blaze vs d, max efficiency vs d, and a (d, blaze) efficiency heatmap with the optimal-blaze ridge.
  • run_energy_scan — sweeps chosen (d, blaze) designs over an energy range, one titled plot per design plus an overlay comparison when there are two or more.
  • evaluate_survey / evaluate_energy_scan — rebuild every aggregate from what is already on disk, no solver (--eval in the example scripts).

Runnable Ru/B4C second-order example at examples/simulation/multilayer_grating_design/.

Web app — the Multilayer design tab

Enter the parameters, run the survey, look at the plots, pick designs, scan them over energy. Specifically:

  • Interactive Plotly survey plots. Hover reads d, blaze and efficiency off any point. Clicking a heatmap cell selects that design for step 2 and marks it; clicking again removes it. The dropdown picker still works and stays in sync.
  • Live updates. Both stages draw themselves as they run — the survey rewrites its table after every cell (atomically, so a reader never catches a half-written file), and the energy scan reads the per-energy checkpoints it already writes. The page reloads once when a stage finishes.
  • Abort actually aborts. should_continue was only consulted between items, so a 200-energy scan was uninterruptible for its whole multi-hour run. run_multilayer_theta_search_sweep now takes a stop_event and terminates its worker pool. Because an in-process solve cannot be interrupted, passing a stop event forces worker-process execution even at max_workers=1; callers that pass none keep the cheaper in-process path. Measured on a live scan: 12 processes to 4 in under a second, with all 147 checkpointed energies intact.
  • Edit parameters after creation. A change a finished stage depended on marks that stage stale rather than deleting anything; cosmetic changes invalidate nothing.
  • Download script — the study as one standalone file: parameters as named constants, then both stages behind --survey / --energy-scan (plus --best, --pairs, --eval). Imports only grax and pandas, so it runs on a cluster with nothing else installed.

Verification

576 passed, 5 skipped. New coverage for the designer, the web routes, and the sweep's stop path (including that terminating the pool does not deadlock the executor's __exit__).

Beyond the suite, this was exercised against a real 99-cell Ru/C order-2 study throughout: survey, a 200-energy scan, a 4-design scan, abort/resume, and the generated script run end to end (--survey, --energy-scan --best, --pairs, both --eval paths) with its CONFIG compared field by field against the web app's.

pyproject.toml is deliberately untouched — no version bump; that is for the release.

🤖 Generated with Claude Code

simonevadi and others added 15 commits September 4, 2026 15:04
Exposes the three-stage d-spacing / gamma / blaze workflow in grax-web as its
own page. A study is a self-contained directory under
`<data_dir>/multilayer_studies/<id>/`; its stages write straight into it through
the library's own layout (0_d_spacing/, 1_gamma/, 2_blaze/, plot/,
optimization_state.json) and a study.json manifest on top tracks the shared
config plus each stage's status, inputs and suggestions.

Library:
- run_d_spacing_study / run_gamma_study / run_blaze_study gain optional
  progress_callback (called with a new grax.StageProgress per scanned item) and
  should_continue (checked per item; returning False stops the scan, computes
  the suggestion from the completed subset, and sets aborted=True on the result).
  Both default to None -- no change for existing callers/CLI/tests.
- The d-spacing scan now evaluates the geometry candidate first so an aborted
  run still yields its suggestion.

Web:
- New grax/web/multilayer_studies.py: MultilayerStudyStore (mirrors RunStore),
  the editable-field spec list, and config parse/build helpers.
- New routes in create_app: /multilayer (list + create + bulk delete),
  /multilayer/new, /multilayer/<id>, per-stage run / status / abort / reset, and
  /multilayer/<id>/delete. Each stage runs in a daemon thread reusing
  ActiveRunState + resource_manager; the status endpoint returns the shape
  initRunMonitor already consumes, and web.js now binds every
  [data-live-run-monitor] on a page.
- Re-running a stage flags later completed stages "stale" (results kept); a
  per-stage Reset deletes one stage's outputs and clears its state keys; Delete
  study removes the directory. Slug-guarded study ids reject path traversal.
- New templates (_multilayer_macros, multilayer_index / study_form /
  study_detail / stage_abort), .stage-card / .status-pill CSS, a nav entry, and
  a web-docs section.

Tests: tests/unit/test_web_multilayer.py (stage runners faked) covers create,
run, stale marking, reset, abort+discard, delete and the traversal guard;
tests/unit/test_multilayer_optimization.py gains progress_callback /
should_continue / partial-abort coverage. Full unit + smoke suites pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…nergy scan

Drops the unreleased three-stage run_d_spacing_study/run_gamma_study/
run_blaze_study workflow (and its web UI page) in favor of
grax.MultilayerGratingDesigner: a 2-D d-spacing x blaze-angle survey that
seeds every grating's incident angle from the multilayer theta search
itself (no CFF input), followed by per-design energy scans. Each solver
run keeps its full artifact bundle on disk, survey/energy-scan results
can be re-aggregated and re-plotted without re-solving (--eval), scan
settings are independent per step, and plot titles show the real
material compound.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
MultilayerDesignConfig.plot_dir is now "<output_dir>/plots" (was "plot"),
and EnergyScanResult.titled_plot_path is written there alongside the
survey's headline plots instead of inside each design's own
energy_scan/ folder. Since every design now shares one folder, the
filename itself carries the coating, diffraction order, d-spacing and
blaze angle: efficiency_vs_energy_<materials>_order<n>_d<d>nm_blaze<b>deg.png
(e.g. efficiency_vs_energy_Ru-B4C_order2_d3.102nm_blaze0.859deg.png).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Gives grax.MultilayerGratingDesigner a web UI, replacing the /multilayer
page removed with the old three-stage API. The form exposes every
MultilayerDesignConfig field in the dataclass's three sections, with the
nested scan settings behind an Advanced toggle and a live warning once the
survey grid gets large. After the survey its three headline plots render
and the energy scan is chosen with three buttons -- best only, the optimal
blaze at every d-spacing, or a manually built list of (d, blaze) cases; the
best-only choice can also be made up front so step 2 chains straight off
the survey.

Supporting library additions: plot_energy_scan_overlay for comparing
several scanned designs on one axis, and should_continue on
run_energy_scan so a long scan can be aborted between designs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The fifteen scan settings previously flowed as an undifferentiated block in
the fieldset grid. FieldSpec gains a `row`, and study_form_sections returns
(row_label, fields) pairs, so each theta-search pass claims its own labelled
line: rough, fine, final solve, then peak selection and roughness. With the
pass named in the row heading the field labels drop their prefix, leaving
"Half-width, deg" / "Points" / "Fourier orders" / "x resolution, nm".

Presentation only -- every input keeps its dotted config name.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
run_energy_scan reports progress once per design, so a single-design
scan sat at "running 0 / 1" for its entire multi-hour run -- visually
indistinguishable from a job that never started. The web monitor now
counts solved energies from each design's checkpoint file instead, and
names the design currently being scanned.

Because scans resume from checkpoints, the ETA divides the elapsed time
by the energies solved by *this* run only (ActiveRunState.resumed_points),
not by the resumed ones, which cost it no time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two changes to the Multilayer design tab.

The three step-1 plots become interactive Plotly charts, built client-side
from the survey grid the page already embeds for the design picker. The two
line plots now sit at equal width (.grid.halves) instead of the main-plus-
sidebar .grid.two, and clicking a heatmap cell selects that (d, blaze) for
step 2 -- opening the manual picker, adding the case and marking the cell,
with a second click removing it. The matplotlib PNGs are still written and
still serve as the fallback when plotly is not installed.

Aborting a stage now terminates the work in flight. should_continue was
only consulted between survey cells and between energy-scan designs, so a
200-energy scan could not be interrupted at all until the whole design
finished. run_multilayer_theta_search_sweep takes a stop_event and an
on_worker_pids_changed callback; setting the event stops submissions,
terminates the pool and returns stopped_early with the energies already
solved. Since an in-process solve cannot be interrupted, supplying a
stop_event forces worker-process execution even at max_workers=1 -- callers
that pass none keep the cheaper in-process path untouched.

A killed survey cell is discarded, because its header-only summary CSV
would make every later evaluate_survey() raise on .iloc[0]. A half-scanned
design is not returned, but its checkpoint keeps every solved energy, so
re-running resumes there. The web app records a deliberately killed stage
as aborted rather than failed, and reports the real pool size instead of a
hardcoded 1 worker.

Verified against a live 99-cell study: abort dropped 12 processes to 4 in
under a second, the stage read "aborted" with no error text, and all 147
checkpointed energies survived.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The stage monitor cleared its polling timers on completed/failed/aborted
but left the progress card on screen. Stage results are rendered
server-side, so a page opened while the stage was running could never show
them -- a finished energy scan looked like it had produced no plot until
the user reloaded by hand.

Monitors now reload once on a terminal state, opted into with
data-run-reload-on-finish so the shared run-detail monitor is unchanged.
The reloaded page renders the results in place of the monitor, so there is
no loop.

The energy-scan result plots also move to the survey's equal-width grid,
now auto-fit so a lone scanned design takes the full width rather than
half of it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Parameters could only be set when the study was created, so adjusting a
survey grid or an energy range meant starting over and abandoning the
results. "Edit parameters" on the study page now reopens the same form,
prefilled from the stored config, and posts to a new edit route.

Changing something a finished stage depended on marks that stage stale
instead of deleting anything, so the artifacts stay on disk until the user
re-runs or resets. stages_invalidated_by() decides which stages those are
from each field's section: shared and survey fields invalidate the survey
and everything downstream, energy-scan fields only the scan, and a small
cosmetic set (coating label, plot saving, checkpointing, worker count)
invalidates nothing. Editing is refused while a stage is running.

Checkboxes now render a hidden "0" companion, because an unchecked box
submits nothing at all -- which on an edit, where the parser overlays the
form onto the stored config, would have silently kept the old value.

Verified against a live study: submitting the form unchanged leaves the
config byte-identical, all four checkbox fields included.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The survey plots only existed once the stage finished, so a long survey
showed a progress bar and nothing else. run_survey now rewrites
survey/survey.csv after every cell and the design page polls a new
survey-options endpoint to redraw the three charts from the growing table;
the reload-on-finish already in place is what ends the polling.

That rewrite has to be atomic. A plain to_csv truncates before writing, and
the page reading the file in that window got a partial or empty table -- a
500 on the detail page, reproducible within seconds of starting a survey.
It now writes a temp file and os.replace()s it, and survey_design_options
answers an unreadable read with "nothing yet" rather than raising.

The new-study form also starts from the newest study's config instead of
the dataclass defaults, so tuned advanced scan settings carry over; the
form names the study it copied. With no studies it falls back to defaults.

Verified live: 20 polls of the detail page and the endpoint during a
running survey all returned 200 with the cell count climbing 0 to 9, and
the charts grew in the browser without a reload.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Each design's sweep already checkpoints one JSON record per solved energy,
so the live curve needs no new bookkeeping -- only a reader. A new
energy-scan-points endpoint turns those records into per-design
(energy, efficiency) series, sorted by energy because workers finish out
of order, skipping failed cases and the partial last line. The running
stage card polls it every few seconds and draws one curve per design;
the finished stage renders its saved plots as before.

The scan's subtitle names the coating and order but not the survey
energy -- that belongs to the survey, not to a scan across energies -- and
its legend sits below the axes, where a rising curve cannot hide it.

_energy_scan_checkpoint_progress now shares the path helper rather than
rebuilding the design directory name itself.

Verified against a live 4-design scan: the overlay drew the finished
200-point curve plus the design in progress, and the checkpoint grew from
3 to 11 points in 45 s across 8 workers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
"Download script" on the study page writes the study out as one
self-contained Python file: the parameters as named constants under the
same SHARED / SURVEY / ENERGY SCAN banners MultilayerDesignConfig groups
them by, the config built from them, then both stages behind --survey /
--energy-scan, with --best, --pairs and --eval matching the bundled
example. No stage flag runs both in order.

It imports only grax and pandas -- no sibling parameters module, nothing
from the web app -- so it can be copied to a cluster and run there. The
scan-settings blocks are emitted in dataclass order (rough, fine, final,
peak) rather than whatever order the stored JSON holds.

Verified by generating from a real study and running the result: --survey,
--energy-scan --best, --energy-scan --pairs and both --eval paths all
produced their artifacts, and the generated CONFIG compares equal to the
stored one field by field.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The Unreleased entries for this branch were newest-first, so a reader met
the offline script download before learning the design tab existed. They
now read foundation (library workflow) to UI (the tab, its plots) to
refinements (live updates, abort, editing, export).

Also drops a claim the abort work invalidated: the tab entry still said a
stage "can be aborted between items".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The tutorial covered the API but gave the bundled example four lines of
pointer. It now walks through it: what each of the four files is for, why
the parameters file is split into three banner sections, how to run it and
how to shrink the shipped 1400-cell grid for a first pass, how to read the
survey tree, how to pick designs, and how --eval and checkpointing let you
iterate without re-solving.

The part worth having written down is why the rough theta-search window has
to be wide. The search is seeded from the multilayer Bragg estimate, which
overshoots the true grating optimum for shallow inside orders: in the
example's best cell the seed is 1.356 deg and the search settles at
0.578 deg, so a 0.2 deg half-width would never reach the peak and every
cell would report an edge artefact. The numbers come from a real 11x9 run,
as does the efficiency-versus-d profile showing how sharp the resonance is.

The abort section now covers stop_event beside should_continue.

Also removes "There is no CFF input" from the design form and web docs. It
answered a question only someone who had used the removed three-stage
workflow would think to ask; the docs page now states the positive fact
instead.

Docs build clean (2 pre-existing unrelated toctree warnings); every grax
cross-reference on the page resolves.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… mentions

The tutorial explained the workflow before showing a single result. It now
opens with a Quick start: the two commands that run the bundled Ru/B4C
example, immediately followed by the real plots those commands produce --
the (d, blaze) efficiency heatmap for the survey, then the efficiency-
versus-energy curve for the best design. The four images are the actual
output of a real 50x28-cell survey and 1000-point energy scan, checked
into docs/tutorials/images/simulation/ (the project's existing convention
for checked-in result plots, matching e.g. fixed_angle_roughness's). The
two survey headline curves are placed lower, next to the Step 1 prose that
already describes them in detail, so nothing needed to be shown twice.

The in-depth walkthrough below it is otherwise the prior content, with its
example numbers corrected against the real full-grid run rather than a
smaller partial one: the Bragg seed sits 0.79 deg above the angle the
search settles on for the actual best cell (d = 3.102 nm, blaze = 0.859
deg), not the earlier partial-run estimate.

Also removes every "there is no CFF input" mention from the multilayer-
design workflow's documentation and source: the tutorial, the API
reference page, the library's own module docstring, and the bundled
example's parameter file. It was explaining the absence of an input a
now-removed workflow used to need, which reads as a non sequitur to
anyone who never saw that workflow -- the text now just states what the
search does.

Docs build clean (2 pre-existing unrelated toctree warnings only); all ten
grax cross-references and both {doc} links on the page resolve. Verified
by rendering the built HTML in a browser: both new plots and their
captions land exactly where the quick-start code block expects them, and
the two headline curves render inline where Step 1 discusses them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@simonevadi
simonevadi force-pushed the feature/multilayer-grating-design-workflow branch from 8d55e7f to b1dc562 Compare September 11, 2026 17:59
@simonevadi
simonevadi merged commit 6e6cc8f into develop Sep 12, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant