# Reconstruction scan format Reconstructing a scan produces one HDF5 file per point and a shared catalog, `scan.h5`, in the output directory. The catalog lists the requested points, their acquisition metadata, and their outcomes. Under `points/`, each completed file contains the reconstructed depth stack and everything needed to inspect, index, or export it independently: physical depths, reference images, settings, geometry, and acquisition metadata. This page specifies the two file layouts, scan preparation, pixel values, frame selection, and square regions of interest (ROIs). For examples, see [Reconstruct a wire scan](../guides/reconstruction.md). {func}`~lauelab.reconstruct.prepare_scan` checks the inputs and publishes the first catalog. {func}`~lauelab.reconstruct.reconstruct_point` writes one point file, and {func}`~lauelab.reconstruct.reconstruct_scan` runs the full workflow using local worker processes. {class}`~lauelab.reconstruct.ScanReader` reads a catalog, {class}`~lauelab.reconstruct.PointReader` reads a point file, and {func}`~lauelab.reconstruct.validate_scan_file` checks a catalog. Writers, readers, and validators share the layout definitions in `lauelab/reconstruct/_scan_layout.py`. {class}`~lauelab.reconstruct.Reconstructor`, {func}`~lauelab.reconstruct.reconstruct_points`, {func}`~lauelab.reconstruct.reconstruct`, and their per-depth output do not change. The ROI rules on this page are implemented by `lauelab.reconstruct.inspection` and rendered by the depth-inspection builders in `lauelab.visualization`; the fixtures under `tests/data/reconstruction_contract/` pin them so that the Laue Portal is built against the same definition. Both files follow the [HDF5 file conventions](hdf5-conventions.md). The catalog's root `format` attribute is `lauelab-reconstruction-scan` and its `version` is 2. A point file's `format` is `lauelab-reconstruction-point` and its `version` is 2. Point version 1 is no longer supported. ## Scan directory ```text / scan.h5 points/ Twin2_wire_1.h5 Twin2_wire_2.h5 ``` For an input named `Twin2_wire_1.h5`, reconstruction writes `points/Twin2_wire_1.h5`. The input stem must be nonempty, must not start with `.`, and must not contain path separators or control characters; other names accepted by the file system are allowed. Inputs with the same stem, including names that differ only in letter case, would produce conflicting output paths. Preparation rejects these inputs before writing any files. This also applies to an input listed twice. Reconstruct inputs with conflicting names into separate scan directories. The scan directory must not exist or must be empty. A nonempty directory is rejected so that a new scan never mixes with old output; choose another directory or remove the old output deliberately. The catalog records each point file as a relative path with `/` separators, such as `points/Twin2_wire_1.h5`, resolved against the directory that contains `scan.h5`. Moving the whole directory keeps the scan readable. Individual point files can also be copied and used independently. Both layouts use ordinary datasets, without HDF5 external links or virtual datasets. ## NeXus structure Point files use the [NeXus base classes](https://manual.nexusformat.org/examples/python/simple_example_write1/index.html) to identify the reconstructed data and its axes. The catalog remains a lauelab HDF5 table. This format does not claim conformance to a specialized NeXus application definition, and copied acquisition metadata is not automatically reclassified. The root has `default="entry1"`. `/entry1` has `NX_class="NXentry"` and `default="data"`. `/entry1/data` has `NX_class="NXdata"`, `signal="data"`, and `axes=["depth", ".", "."]`. Its `depth` child is an internal hard link to `/entry1/depth`; the other two dimensions use implicit image indices. Physical depth remains in µm along the incident beam from the geometry's Si origin. Image indexing remains `data[depth_index, y, x]`. `/entry1/reconstruction` has `NX_class="NXprocess"`. Its `point`, `settings`, `execution`, `geometry`, `acquisition`, `detector`, and `normalization` groups have `NX_class="NXparameters"`. The same process group contains five `NXdata` groups, each with `signal="data"`: | Group under `/entry1/reconstruction` | Axes | Data shape | | --- | --- | --- | | `first_raw` | `[".", "."]` | `(rows, columns)` | | `sum_raw` | `[".", "."]` | `(rows, columns)` | | `sum_reconstructed` | `[".", "."]` | `(rows, columns)` | | `stored_depth_intensity` | `["depth"]` | `(n_depths,)` | | `computed_depth_intensity` | `["depth"]` | `(n_depths,)` | Each intensity group has a `depth` hard link to `/entry1/depth`. These links refer to the same dataset inside the file, add no copy of the depth values, and remain valid when the file is moved. Readers validate the NeXus attributes and depth links along with dataset shapes, dtypes, units, and value conventions. ## Identities Frames are selected by point and depth index. Catalog entries follow the input order, which is frozen when the scan is prepared. | Identity | Type | Meaning | | --- | --- | --- | | Point ID | string | Unique, non-empty identifier of a point within its scan. Defaults to the input stem. It is independent of the point filename. | | Manifest index | zero-based integer | Position of the point in the input order. Every catalog dataset uses this order. | | Point path | string | Location of the point file relative to the scan directory. It is assigned when the scan is prepared, whatever the point's outcome. | | Depth index | zero-based integer | Position of a frame in the point's `(n_depths, ny, nx)` stack. | | Physical depth | `numpy.float64`, µm | Depth along the incident beam from the Si origin of the geometry file. It is `depth_um[depth index]`. | The point ID and original manifest index are saved in the point file and remain available if you copy it elsewhere. Standalone points have no manifest index. Physical depths are stored in `/entry1/depth`. For example, depth index 0 corresponds to -25 µm when the reconstructed depth grid starts at -25 µm. A **frame reference**, {class}`~lauelab.indexing.ScanFrame`, selects one stored frame by the path of its point file and a depth index. The reference can be serialized and passed to indexing workers, which open the file locally. Indexing results preserve the point identity alongside the depth index. ## Preparation and tasks {func}`~lauelab.reconstruct.prepare_scan` takes the input files in manifest order, the scan directory, the geometry, and the reconstruction settings. Before writing, it validates the settings, detector slot, point IDs, output filenames, and destination. It then reads each input's metadata without loading frames and creates one {class}`~lauelab.reconstruct.PointTask` for each point whose input can be read. Input paths are made absolute against the current directory; symbolic links are not resolved. The first catalog is published before any point is computed. It lists every input. An input that cannot be read, or whose metadata is invalid, is recorded as a `failed` point with its error and has no task; it does not stop the other points. Each task includes the absolute input and output paths, point identity, detector slot, validated reconstruction settings, and complete geometry XML, together with the frame shape, depth count and bounds, pixel type, selected raw-slice range, scan number, sample position, and incident energy found during inspection. Missing acquisition values are `None` in the task and use the schema's missing values in HDF5. Tasks contain only strings, numbers, and plain containers, so they can be pickled or converted to JSON. The executing process opens its own input and builds its own native state. It checks the output description, raw-slice range, and acquisition values against the prepared task before creating an output file. A disagreement produces a failed point. These checks read metadata; they do not verify raw pixel content. Treat tasks as read-only. The coordinator keeps an independent copy of the prepared request and rejects changed tasks or completed files that disagree with it. Exactly one coordinator, the {class}`~lauelab.reconstruct.PreparedScan` returned by `prepare_scan`, writes `scan.h5`. Workers write only the point files of their tasks. A scheduler other than {func}`~lauelab.reconstruct.reconstruct_scan` uses the coordinator in this order: 1. Call {meth}`~lauelab.reconstruct.PreparedScan.record_dispatch` when handing a task to a worker. This sets its status to `writing`. 2. The worker calls {func}`~lauelab.reconstruct.reconstruct_point` with the task and returns the small {class}`~lauelab.reconstruct.PointOutcome`. Pixels stay in the worker. 3. Call {meth}`~lauelab.reconstruct.PreparedScan.record` as outcomes arrive, in any order. 4. Call {meth}`~lauelab.reconstruct.PreparedScan.finish` after every dispatched task has an outcome. Pass `cancelled=True` if some tasks were never dispatched; they become `unattempted`. The scheduler assigns each task once and records a failed outcome for a task whose worker was lost; the coordinator provides no retry. Leaving the coordinator's context without `finish` records the run as `failed`, dispatched points without an outcome as `interrupted`, and the other points as `unattempted`. A standalone point, reconstructed without a scan, is prepared with the same rules and written by the same writer; its manifest index is missing. ## Value semantics Pixel values are described at three stages of processing: | Kind | Dtype | Meaning | | --- | --- | --- | | Raw | dtype of the input frames | Detector counts as acquired, before the cosmic-ray filter and before any normalization. | | Computed | `numpy.float64` | Output of the reconstruction kernel. These are `ReconstructionResult.images`, and their per-depth totals are `ReconstructionResult.depth_intensity`. | | Stored | the point's output pixel type | The pixels in the file. Inspection, ROI traces, and downstream indexing use these values. | Before writing, the library multiplies computed intensities by `rescale` and converts them to the output pixel type. `rescale` is 1 unless exponent normalization writes an integer type; it is recorded as `/entry1/reconstruction/normalization/rescale` and is applied once. For a floating-point type, the conversion equals a NumPy cast. For an integer type, the conversion truncates toward zero and then saturates at the limits of the type. Positive and negative infinity saturate. NaN stores 0 in an integer type. The native library converts pixels and computes reductions in one pass. For finite values and infinities, conversion matches the HDF5 conversion used by per-depth output. Integer conversion of NaN depends on the HDF5 destination type and can differ from the scan format's defined value of zero. ```{warning} Reconstructed intensities can be negative with any `wire_edge` setting. An unsigned output type stores 0 for each negative pixel, so its stored totals are larger than the computed totals. For the recorded leading-edge reference, the `numpy.uint16` frames sum to 573801 and the computed frames sum to 512684.9. ``` ### Reductions Stored-pixel totals are calculated after conversion, using the same values you would read back from the file. - Integer stored pixels accumulate in `numpy.int64` and are exact. A sum of `n` values of at most 4 bytes cannot overflow while `n` is below 2³¹, which exceeds the pixel count of a 4096 by 4096 frame. - Floating-point stored pixels accumulate in `numpy.float64`. The summation order is not specified. Reconstructed pixels are signed, so a total can be much smaller than the pixels that produce it, and a tolerance relative to the total is not meaningful. Compare a reduction with a recomputed one at an absolute tolerance of 10⁻¹² times the sum of the absolute pixel values in the reduction. `/entry1/reconstruction/computed_depth_intensity/data` contains the totals before storage conversion, as reported in the LaueGo-compatible summary. `/entry1/reconstruction/stored_depth_intensity/data` contains the stored-pixel totals used for inspection. Conversion and `rescale` can therefore change the totals. ### Reference images Three reference images are saved with each point for inspection after the raw input becomes unavailable. Each has the same shape as a detector frame. | Image | Values | Definition | | --- | --- | --- | | `/entry1/reconstruction/first_raw/data` | raw | The first scan frame the reader selects. | | `/entry1/reconstruction/sum_raw/data` | raw | The sum of every selected scan frame. | | `/entry1/reconstruction/sum_reconstructed/data` | stored | The sum of the stored frames through depth. | `/entry1/reconstruction/acquisition/raw_slices` records the half-open range of stored input slices that was selected. For a 34-ID-E multi-image file with `n` stored slices the range is `[1, n - 1)`: slice 0 is bookkeeping, and the last slice is never differenced. See "How a point file is read" in [Reconstruct a wire scan](../guides/reconstruction.md). The raw stack itself is not copied. ## File layout In the tables below, `n_points` is the number of points in the scan. The depth count (`n_depths`), frame dimensions (`rows`, `columns`), and output pixel type can vary between points. The following dtype rules apply separately to each point: | Rule | Dtype | | --- | --- | | `stored` | The point's output pixel type. | | `input` | The dtype of the point's raw frames. | | `stored_accumulator` | `