Core concepts
Core data model
earthrs.scene.Scene, earthrs.dataset.Dataset, and earthrs.samples.Samples are immutable
(frozen, slotted) dataclasses:
Scene— a single immutable raster acquisition, with a normalised cloud mask/probability schema regardless of source product.Dataset— an immutable collection ofSceneobjects, withfilter_dates,filter_clouds, andsample_points.Samples— the canonical immutable feature-table type: an immutable tuple of row mappings plus metadata, with one row per observation.Dataset.sample_pointsand the functions inearthrs.samplingall returnSamples.
Methods that "change" one of these objects (Scene.add_history, Dataset.add_scene,
Dataset.filter_dates, Dataset.filter_clouds, Samples.with_metadata, and every processing
function) return a new instance and leave the original untouched:
from earthrs import Dataset, Scene
scene = Scene(data={"nir": [0.4], "red": [0.1]})
updated = scene.add_history("imported")
assert scene.history == () # original is unchanged
assert updated.history == ("imported",)
Cloud metadata normalisation
Cloud fields can arrive under different keys depending on the source product (for example
cloud, cloud_mask, or mask_cloud in a masks dict; cloud_mask or cloud_prob in a
metadata dict). Scene.__post_init__ runs earthrs.io.normalise_cloud_schema so that,
regardless of which field an importer populated, scene.cloud_mask, scene.cloud_probability,
scene.masks["cloud"], and scene.metadata["cloud_mask"] all agree. Downstream algorithms
(cloud masking, cloud filtering) read from this normalised schema rather than product-specific
field names.
Spectral indices
earthrs.indices provides ndvi, ndwi, and evi, each taking a Scene and returning
per-pixel index values computed from its named bands.
Sampling
earthrs.sampling provides sample_points, sample_polygons, and sample_transects (plus a
sample() dispatcher that infers the target from geometry type), all returning Samples — one
row per point, polygon, or transect vertex observation. Dataset.sample_points delegates to
earthrs.sampling.sample_points per scene and aggregates the results into a single Samples.