Skip to content

latent.stats.sampling

Sampling utilities for building representative eval subsets.

Functions

difficulty_based_sample

difficulty_based_sample(df: pd.DataFrame, stratum_column: str, n_total: int, oversample_factor: float = 2.0, min_per_stratum: int = 10, seed: int | None = None) -> pd.DataFrame

Over-sample edge cases and under-represented categories.

Smaller strata receive a higher sampling rate, up to oversample_factor times the proportional rate.

Args: df: Source DataFrame to sample from. stratum_column: Column name defining strata. n_total: Total number of rows to sample. oversample_factor: Maximum multiplier applied to the proportional rate for the smallest strata. min_per_stratum: Minimum samples per stratum. seed: Random seed for reproducibility.

Returns: Sampled DataFrame preserving the index from the original.

stratified_sample

stratified_sample(df: pd.DataFrame, stratum_column: str, n_total: int, allocation: str = 'proportional', min_per_stratum: int = 0, seed: int | None = None) -> pd.DataFrame

Select a representative eval subset via stratified sampling.

Args: df: Source DataFrame to sample from. stratum_column: Column name defining strata. n_total: Total number of rows to sample. allocation: Allocation strategy - "proportional" or "equal". min_per_stratum: Minimum samples per stratum. seed: Random seed for reproducibility.

Returns: Sampled DataFrame preserving the index from the original.

Raises: ValueError: If n_total exceeds the DataFrame size or allocation is not recognized.