Skip to content

latent.datasets

Dataset loading and saving with filesystem persistence and MLflow lineage.

Inputs are loaded from filesystem (data//input/). Outputs are saved to filesystem (data//output/) and tracked via MLflow lineage metadata when an active run exists. Output loading reads from filesystem first, falling back to MLflow for backward compatibility.

Classes

CatalogAccessor

CatalogAccessor(args: Any = (), kwargs: Any = {})

Dict subclass that adds a .stable property for accessing stable artifacts.

All existing dict operations (subscript, in, .items(), len(), etc.) are preserved. The only addition is the stable property.

DatasetProxy

DatasetProxy(flow_name: str, dataset_name: str, config: dict)

Lazy-loading proxy for a dataset.

StableAccessor

StableAccessor(stable_dir: Path)

Attribute-style access to curated artifacts in data/stable/.

Supports loading files by stem name and returning paths: accessor.glossary -> loads data/stable/glossary.json via json.JSON accessor.glossary_path -> Path("/.../data/stable/glossary.json")

Functions

dataset_fingerprint

dataset_fingerprint(flow_name: str, dataset_name: str, catalog_config: dict, is_output: bool = False) -> str | None

A value that changes when this dataset's content changes, or None.

None means "cannot be fingerprinted", and every caller must treat that as uncacheable rather than as "unchanged": a floating remote version, a missing file, or an unreadable one could all be hiding a change, and serving a cached result across one is the staleness this exists to prevent.

A file under _FINGERPRINT_READ_BUDGET is hashed, because size + mtime alone collide on filesystems whose mtime resolution is a whole second. Above the budget the read would cost more than the recompute it saves, so size + mtime_ns stands in, and the two failure modes there point opposite ways: a git checkout rewrites mtime without changing content, which costs one recompute, while content changing with size and mtime both preserved takes deliberate effort.

get_catalog_datasets

get_catalog_datasets(flow_name: str | None = None) -> CatalogAccessor

Get lazy-loading dataset proxies from catalog configuration.

This returns dataset proxy objects (not the data itself), similar to Kedro's DataCatalog behavior. Call .load() on a proxy to actually load the data.

Args: flow_name: Name of the flow. If None, auto-detects from current flow context.

Returns: Dict mapping dataset names to dataset proxy objects

Example: >>> # Inside a @flow decorated function >>> catalog = get_catalog_datasets() # Auto-detects flow name >>> # Or explicitly specify >>> catalog = get_catalog_datasets("my_flow") >>> df = catalog["my_dataset"].load() # Loads here >>> catalog["output"].save(df) # Save data

load_dataset

load_dataset(flow_name: str, dataset_name: str, catalog_config: dict, is_output: bool = False) -> pd.DataFrame | dict | list | str | Any

Load a dataset.

  • Inputs (is_output=False): Load from filesystem at data//input/
  • Outputs (is_output=True): Load from filesystem at data//output/, falling back to MLflow artifacts for backward compatibility.

Args: flow_name: Name of the flow dataset_name: Name of the dataset catalog_config: Catalog configuration for this dataset is_output: True for output datasets, False for input datasets

Returns: Loaded data (DataFrame, dict, list, or str)

Example: data = load_dataset("my_flow", "raw_data", config, is_output=False) results = load_dataset("upstream_flow", "results", config, is_output=True)

save_dataset

save_dataset(flow_name: str, dataset_name: str, data: Any, catalog_config: dict, is_output: bool = True) -> str

Save a dataset.

  • Outputs (is_output=True): Save to filesystem at data//output/, with MLflow lineage metadata logged when a run is active.
  • Inputs (is_output=False): Save to filesystem at data//input/.

Args: flow_name: Name of the flow dataset_name: Name of the dataset data: Data to save (DataFrame, dict, list, or str) catalog_config: Catalog configuration for this dataset is_output: True for output datasets, False for input datasets

Returns: File path string where the dataset was saved

Example: save_dataset("my_flow", "results", df, config, is_output=True)

Methods

CatalogAccessor.stable

DatasetProxy.load

load(is_output: bool = False) -> Any

Load the dataset.

Args: is_output: True for output datasets, False for input datasets

Returns: Loaded data

DatasetProxy.save

save(data: Any, is_output: bool = True)

Save data to this dataset.

Args: data: Data to save is_output: True for output datasets, False for input datasets