latent.datasets¶
Dataset loading and saving with filesystem persistence and MLflow lineage.
Inputs are loaded from filesystem (data/
Classes¶
CatalogAccessor¶
Dict subclass that adds a .stable property for accessing stable artifacts.
All existing dict operations (subscript, in, .items(), len(), etc.)
are preserved. The only addition is the stable property.
DatasetProxy¶
Lazy-loading proxy for a dataset.
StableAccessor¶
Attribute-style access to curated artifacts in data/stable/.
Supports loading files by stem name and returning paths: accessor.glossary -> loads data/stable/glossary.json via json.JSON accessor.glossary_path -> Path("/.../data/stable/glossary.json")
Functions¶
dataset_fingerprint¶
dataset_fingerprint(flow_name: str, dataset_name: str, catalog_config: dict, is_output: bool = False) -> str | None
A value that changes when this dataset's content changes, or None.
None means "cannot be fingerprinted", and every caller must treat that
as uncacheable rather than as "unchanged": a floating remote version, a
missing file, or an unreadable one could all be hiding a change, and
serving a cached result across one is the staleness this exists to prevent.
A file under _FINGERPRINT_READ_BUDGET is hashed, because size + mtime
alone collide on filesystems whose mtime resolution is a whole second.
Above the budget the read would cost more than the recompute it saves, so
size + mtime_ns stands in, and the two failure modes there point
opposite ways: a git checkout rewrites mtime without changing content,
which costs one recompute, while content changing with size and mtime
both preserved takes deliberate effort.
get_catalog_datasets¶
Get lazy-loading dataset proxies from catalog configuration.
This returns dataset proxy objects (not the data itself), similar to Kedro's DataCatalog behavior. Call .load() on a proxy to actually load the data.
Args: flow_name: Name of the flow. If None, auto-detects from current flow context.
Returns: Dict mapping dataset names to dataset proxy objects
Example: >>> # Inside a @flow decorated function >>> catalog = get_catalog_datasets() # Auto-detects flow name >>> # Or explicitly specify >>> catalog = get_catalog_datasets("my_flow") >>> df = catalog["my_dataset"].load() # Loads here >>> catalog["output"].save(df) # Save data
load_dataset¶
load_dataset(flow_name: str, dataset_name: str, catalog_config: dict, is_output: bool = False) -> pd.DataFrame | dict | list | str | Any
Load a dataset.
- Inputs (is_output=False): Load from filesystem at data/
/input/ - Outputs (is_output=True): Load from filesystem at data/
/output/, falling back to MLflow artifacts for backward compatibility.
Args: flow_name: Name of the flow dataset_name: Name of the dataset catalog_config: Catalog configuration for this dataset is_output: True for output datasets, False for input datasets
Returns: Loaded data (DataFrame, dict, list, or str)
Example: data = load_dataset("my_flow", "raw_data", config, is_output=False) results = load_dataset("upstream_flow", "results", config, is_output=True)
save_dataset¶
save_dataset(flow_name: str, dataset_name: str, data: Any, catalog_config: dict, is_output: bool = True) -> str
Save a dataset.
- Outputs (is_output=True): Save to filesystem at data/
/output/, with MLflow lineage metadata logged when a run is active. - Inputs (is_output=False): Save to filesystem at data/
/input/.
Args: flow_name: Name of the flow dataset_name: Name of the dataset data: Data to save (DataFrame, dict, list, or str) catalog_config: Catalog configuration for this dataset is_output: True for output datasets, False for input datasets
Returns: File path string where the dataset was saved
Example: save_dataset("my_flow", "results", df, config, is_output=True)
Methods¶
CatalogAccessor.stable¶
DatasetProxy.load¶
Load the dataset.
Args: is_output: True for output datasets, False for input datasets
Returns: Loaded data
DatasetProxy.save¶
Save data to this dataset.
Args: data: Data to save is_output: True for output datasets, False for input datasets