Multilayer-grating design workflow: survey + energy scan, with a web tab - #46
Merged
Merged
Conversation
Exposes the three-stage d-spacing / gamma / blaze workflow in grax-web as its own page. A study is a self-contained directory under `<data_dir>/multilayer_studies/<id>/`; its stages write straight into it through the library's own layout (0_d_spacing/, 1_gamma/, 2_blaze/, plot/, optimization_state.json) and a study.json manifest on top tracks the shared config plus each stage's status, inputs and suggestions. Library: - run_d_spacing_study / run_gamma_study / run_blaze_study gain optional progress_callback (called with a new grax.StageProgress per scanned item) and should_continue (checked per item; returning False stops the scan, computes the suggestion from the completed subset, and sets aborted=True on the result). Both default to None -- no change for existing callers/CLI/tests. - The d-spacing scan now evaluates the geometry candidate first so an aborted run still yields its suggestion. Web: - New grax/web/multilayer_studies.py: MultilayerStudyStore (mirrors RunStore), the editable-field spec list, and config parse/build helpers. - New routes in create_app: /multilayer (list + create + bulk delete), /multilayer/new, /multilayer/<id>, per-stage run / status / abort / reset, and /multilayer/<id>/delete. Each stage runs in a daemon thread reusing ActiveRunState + resource_manager; the status endpoint returns the shape initRunMonitor already consumes, and web.js now binds every [data-live-run-monitor] on a page. - Re-running a stage flags later completed stages "stale" (results kept); a per-stage Reset deletes one stage's outputs and clears its state keys; Delete study removes the directory. Slug-guarded study ids reject path traversal. - New templates (_multilayer_macros, multilayer_index / study_form / study_detail / stage_abort), .stage-card / .status-pill CSS, a nav entry, and a web-docs section. Tests: tests/unit/test_web_multilayer.py (stage runners faked) covers create, run, stale marking, reset, abort+discard, delete and the traversal guard; tests/unit/test_multilayer_optimization.py gains progress_callback / should_continue / partial-abort coverage. Full unit + smoke suites pass. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…nergy scan Drops the unreleased three-stage run_d_spacing_study/run_gamma_study/ run_blaze_study workflow (and its web UI page) in favor of grax.MultilayerGratingDesigner: a 2-D d-spacing x blaze-angle survey that seeds every grating's incident angle from the multilayer theta search itself (no CFF input), followed by per-design energy scans. Each solver run keeps its full artifact bundle on disk, survey/energy-scan results can be re-aggregated and re-plotted without re-solving (--eval), scan settings are independent per step, and plot titles show the real material compound. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
MultilayerDesignConfig.plot_dir is now "<output_dir>/plots" (was "plot"), and EnergyScanResult.titled_plot_path is written there alongside the survey's headline plots instead of inside each design's own energy_scan/ folder. Since every design now shares one folder, the filename itself carries the coating, diffraction order, d-spacing and blaze angle: efficiency_vs_energy_<materials>_order<n>_d<d>nm_blaze<b>deg.png (e.g. efficiency_vs_energy_Ru-B4C_order2_d3.102nm_blaze0.859deg.png). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Gives grax.MultilayerGratingDesigner a web UI, replacing the /multilayer page removed with the old three-stage API. The form exposes every MultilayerDesignConfig field in the dataclass's three sections, with the nested scan settings behind an Advanced toggle and a live warning once the survey grid gets large. After the survey its three headline plots render and the energy scan is chosen with three buttons -- best only, the optimal blaze at every d-spacing, or a manually built list of (d, blaze) cases; the best-only choice can also be made up front so step 2 chains straight off the survey. Supporting library additions: plot_energy_scan_overlay for comparing several scanned designs on one axis, and should_continue on run_energy_scan so a long scan can be aborted between designs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The fifteen scan settings previously flowed as an undifferentiated block in the fieldset grid. FieldSpec gains a `row`, and study_form_sections returns (row_label, fields) pairs, so each theta-search pass claims its own labelled line: rough, fine, final solve, then peak selection and roughness. With the pass named in the row heading the field labels drop their prefix, leaving "Half-width, deg" / "Points" / "Fourier orders" / "x resolution, nm". Presentation only -- every input keeps its dotted config name. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
run_energy_scan reports progress once per design, so a single-design scan sat at "running 0 / 1" for its entire multi-hour run -- visually indistinguishable from a job that never started. The web monitor now counts solved energies from each design's checkpoint file instead, and names the design currently being scanned. Because scans resume from checkpoints, the ETA divides the elapsed time by the energies solved by *this* run only (ActiveRunState.resumed_points), not by the resumed ones, which cost it no time. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two changes to the Multilayer design tab. The three step-1 plots become interactive Plotly charts, built client-side from the survey grid the page already embeds for the design picker. The two line plots now sit at equal width (.grid.halves) instead of the main-plus- sidebar .grid.two, and clicking a heatmap cell selects that (d, blaze) for step 2 -- opening the manual picker, adding the case and marking the cell, with a second click removing it. The matplotlib PNGs are still written and still serve as the fallback when plotly is not installed. Aborting a stage now terminates the work in flight. should_continue was only consulted between survey cells and between energy-scan designs, so a 200-energy scan could not be interrupted at all until the whole design finished. run_multilayer_theta_search_sweep takes a stop_event and an on_worker_pids_changed callback; setting the event stops submissions, terminates the pool and returns stopped_early with the energies already solved. Since an in-process solve cannot be interrupted, supplying a stop_event forces worker-process execution even at max_workers=1 -- callers that pass none keep the cheaper in-process path untouched. A killed survey cell is discarded, because its header-only summary CSV would make every later evaluate_survey() raise on .iloc[0]. A half-scanned design is not returned, but its checkpoint keeps every solved energy, so re-running resumes there. The web app records a deliberately killed stage as aborted rather than failed, and reports the real pool size instead of a hardcoded 1 worker. Verified against a live 99-cell study: abort dropped 12 processes to 4 in under a second, the stage read "aborted" with no error text, and all 147 checkpointed energies survived. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The stage monitor cleared its polling timers on completed/failed/aborted but left the progress card on screen. Stage results are rendered server-side, so a page opened while the stage was running could never show them -- a finished energy scan looked like it had produced no plot until the user reloaded by hand. Monitors now reload once on a terminal state, opted into with data-run-reload-on-finish so the shared run-detail monitor is unchanged. The reloaded page renders the results in place of the monitor, so there is no loop. The energy-scan result plots also move to the survey's equal-width grid, now auto-fit so a lone scanned design takes the full width rather than half of it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Parameters could only be set when the study was created, so adjusting a survey grid or an energy range meant starting over and abandoning the results. "Edit parameters" on the study page now reopens the same form, prefilled from the stored config, and posts to a new edit route. Changing something a finished stage depended on marks that stage stale instead of deleting anything, so the artifacts stay on disk until the user re-runs or resets. stages_invalidated_by() decides which stages those are from each field's section: shared and survey fields invalidate the survey and everything downstream, energy-scan fields only the scan, and a small cosmetic set (coating label, plot saving, checkpointing, worker count) invalidates nothing. Editing is refused while a stage is running. Checkboxes now render a hidden "0" companion, because an unchecked box submits nothing at all -- which on an edit, where the parser overlays the form onto the stored config, would have silently kept the old value. Verified against a live study: submitting the form unchanged leaves the config byte-identical, all four checkbox fields included. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The survey plots only existed once the stage finished, so a long survey showed a progress bar and nothing else. run_survey now rewrites survey/survey.csv after every cell and the design page polls a new survey-options endpoint to redraw the three charts from the growing table; the reload-on-finish already in place is what ends the polling. That rewrite has to be atomic. A plain to_csv truncates before writing, and the page reading the file in that window got a partial or empty table -- a 500 on the detail page, reproducible within seconds of starting a survey. It now writes a temp file and os.replace()s it, and survey_design_options answers an unreadable read with "nothing yet" rather than raising. The new-study form also starts from the newest study's config instead of the dataclass defaults, so tuned advanced scan settings carry over; the form names the study it copied. With no studies it falls back to defaults. Verified live: 20 polls of the detail page and the endpoint during a running survey all returned 200 with the cell count climbing 0 to 9, and the charts grew in the browser without a reload. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Each design's sweep already checkpoints one JSON record per solved energy, so the live curve needs no new bookkeeping -- only a reader. A new energy-scan-points endpoint turns those records into per-design (energy, efficiency) series, sorted by energy because workers finish out of order, skipping failed cases and the partial last line. The running stage card polls it every few seconds and draws one curve per design; the finished stage renders its saved plots as before. The scan's subtitle names the coating and order but not the survey energy -- that belongs to the survey, not to a scan across energies -- and its legend sits below the axes, where a rising curve cannot hide it. _energy_scan_checkpoint_progress now shares the path helper rather than rebuilding the design directory name itself. Verified against a live 4-design scan: the overlay drew the finished 200-point curve plus the design in progress, and the checkpoint grew from 3 to 11 points in 45 s across 8 workers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
"Download script" on the study page writes the study out as one self-contained Python file: the parameters as named constants under the same SHARED / SURVEY / ENERGY SCAN banners MultilayerDesignConfig groups them by, the config built from them, then both stages behind --survey / --energy-scan, with --best, --pairs and --eval matching the bundled example. No stage flag runs both in order. It imports only grax and pandas -- no sibling parameters module, nothing from the web app -- so it can be copied to a cluster and run there. The scan-settings blocks are emitted in dataclass order (rough, fine, final, peak) rather than whatever order the stored JSON holds. Verified by generating from a real study and running the result: --survey, --energy-scan --best, --energy-scan --pairs and both --eval paths all produced their artifacts, and the generated CONFIG compares equal to the stored one field by field. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The Unreleased entries for this branch were newest-first, so a reader met the offline script download before learning the design tab existed. They now read foundation (library workflow) to UI (the tab, its plots) to refinements (live updates, abort, editing, export). Also drops a claim the abort work invalidated: the tab entry still said a stage "can be aborted between items". Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The tutorial covered the API but gave the bundled example four lines of pointer. It now walks through it: what each of the four files is for, why the parameters file is split into three banner sections, how to run it and how to shrink the shipped 1400-cell grid for a first pass, how to read the survey tree, how to pick designs, and how --eval and checkpointing let you iterate without re-solving. The part worth having written down is why the rough theta-search window has to be wide. The search is seeded from the multilayer Bragg estimate, which overshoots the true grating optimum for shallow inside orders: in the example's best cell the seed is 1.356 deg and the search settles at 0.578 deg, so a 0.2 deg half-width would never reach the peak and every cell would report an edge artefact. The numbers come from a real 11x9 run, as does the efficiency-versus-d profile showing how sharp the resonance is. The abort section now covers stop_event beside should_continue. Also removes "There is no CFF input" from the design form and web docs. It answered a question only someone who had used the removed three-stage workflow would think to ask; the docs page now states the positive fact instead. Docs build clean (2 pre-existing unrelated toctree warnings); every grax cross-reference on the page resolves. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… mentions
The tutorial explained the workflow before showing a single result. It now
opens with a Quick start: the two commands that run the bundled Ru/B4C
example, immediately followed by the real plots those commands produce --
the (d, blaze) efficiency heatmap for the survey, then the efficiency-
versus-energy curve for the best design. The four images are the actual
output of a real 50x28-cell survey and 1000-point energy scan, checked
into docs/tutorials/images/simulation/ (the project's existing convention
for checked-in result plots, matching e.g. fixed_angle_roughness's). The
two survey headline curves are placed lower, next to the Step 1 prose that
already describes them in detail, so nothing needed to be shown twice.
The in-depth walkthrough below it is otherwise the prior content, with its
example numbers corrected against the real full-grid run rather than a
smaller partial one: the Bragg seed sits 0.79 deg above the angle the
search settles on for the actual best cell (d = 3.102 nm, blaze = 0.859
deg), not the earlier partial-run estimate.
Also removes every "there is no CFF input" mention from the multilayer-
design workflow's documentation and source: the tutorial, the API
reference page, the library's own module docstring, and the bundled
example's parameter file. It was explaining the absence of an input a
now-removed workflow used to need, which reads as a non sequitur to
anyone who never saw that workflow -- the text now just states what the
search does.
Docs build clean (2 pre-existing unrelated toctree warnings only); all ten
grax cross-references and both {doc} links on the page resolve. Verified
by rendering the built HTML in a browser: both new plots and their
captions land exactly where the quick-start code block expects them, and
the two headline curves render inline where Step 1 discusses them.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
simonevadi
force-pushed
the
feature/multilayer-grating-design-workflow
branch
from
September 11, 2026 17:59
8d55e7f to
b1dc562
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Replaces the unreleased three-stage multilayer-optimization workflow with a two-step design workflow, and gives it a full web UI.
The old workflow estimated the incident angle from a CFF value. That input is gone: the angle is now a result of a multilayer theta search run at every grid point, which is what the CFF was standing in for.
gammais held fixed and the anti-blaze angle defaults to 0; both are later optimizations.Library —
grax.MultilayerGratingDesignerOne frozen
MultilayerDesignConfig, grouped into three banner-marked sections (shared / survey-only / energy-scan-only) so it is obvious which stage reads what. The rough/fine/final theta-search settings are a separateThetaSearchScanSettingsheld on two independent fields,survey_scan_settingsandenergy_scan_settings— the survey runs many cheap single-energy searches, the scan runs few designs over many energies, and they want different budgets.run_survey— a d-spacing × blaze-angle grid; every cell keeps the complete theta-search artifact bundle undersurvey/runs/d<d>nm/blaze<b>deg/plus asearch_parameters.json, so a suspicious cell can be inspected and the settings retuned. Headline outputs: optimal blaze vs d, max efficiency vs d, and a(d, blaze)efficiency heatmap with the optimal-blaze ridge.run_energy_scan— sweeps chosen(d, blaze)designs over an energy range, one titled plot per design plus an overlay comparison when there are two or more.evaluate_survey/evaluate_energy_scan— rebuild every aggregate from what is already on disk, no solver (--evalin the example scripts).Runnable Ru/B4C second-order example at
examples/simulation/multilayer_grating_design/.Web app — the Multilayer design tab
Enter the parameters, run the survey, look at the plots, pick designs, scan them over energy. Specifically:
should_continuewas only consulted between items, so a 200-energy scan was uninterruptible for its whole multi-hour run.run_multilayer_theta_search_sweepnow takes astop_eventand terminates its worker pool. Because an in-process solve cannot be interrupted, passing a stop event forces worker-process execution even atmax_workers=1; callers that pass none keep the cheaper in-process path. Measured on a live scan: 12 processes to 4 in under a second, with all 147 checkpointed energies intact.--survey/--energy-scan(plus--best,--pairs,--eval). Imports onlygraxand pandas, so it runs on a cluster with nothing else installed.Verification
576 passed, 5 skipped. New coverage for the designer, the web routes, and the sweep's stop path (including that terminating the pool does not deadlock the executor's__exit__).Beyond the suite, this was exercised against a real 99-cell Ru/C order-2 study throughout: survey, a 200-energy scan, a 4-design scan, abort/resume, and the generated script run end to end (
--survey,--energy-scan --best,--pairs, both--evalpaths) with itsCONFIGcompared field by field against the web app's.pyproject.tomlis deliberately untouched — no version bump; that is for the release.🤖 Generated with Claude Code