Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 10 additions & 1 deletion .github/workflows/pr.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -26,11 +26,17 @@ jobs:
run: make lint

test:
name: Test Python ${{ matrix.python-version }}
name: Test Python ${{ matrix.python-version }} / pandas ${{ matrix.pandas-version }}
runs-on: ubuntu-latest
strategy:
matrix:
python-version: ["3.9", "3.10", "3.11", "3.12", "3.13"]
pandas-version: ["latest"]
include:
- python-version: "3.13"
pandas-version: "2"
- python-version: "3.13"
pandas-version: "3"

steps:
- uses: actions/checkout@v6
Expand All @@ -46,6 +52,9 @@ jobs:
- name: Install dependencies
run: |
uv pip install -e ".[dev]" --system
- name: Select pandas major version
if: matrix.pandas-version != 'latest'
run: uv pip install --system "pandas==${{ matrix.pandas-version }}.*"
- name: Run tests with coverage
run: make test
- name: Upload coverage to Codecov
Expand Down
10 changes: 8 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,10 +31,16 @@ the two ways to get a believable-looking wrong answer.
## Key Features
- **MicroDataFrame**: A pandas DataFrame with an integrated weight column
- **MicroSeries**: A pandas Series with integrated weights
- **Weighted operations**: All aggregations (sum, mean, median, etc.) automatically use weights
- **Weighted operations**: Supported aggregations (sum, mean, median, etc.) use weights; unsupported estimators raise
- **Inequality metrics**: Built-in Gini coefficient calculation
- **Poverty analysis**: Integrated poverty rate and gap calculations

Weight preservation is limited to the tested operations in the
[support matrix](docs/support.md). Convert explicitly to `pd.Series(s)` or
`pd.DataFrame(df)` to request unweighted pandas behaviour. Use `microdf.concat`
to reject mixed weighted and plain inputs regardless of their order; pandas can
bypass microdf's checks when a plain input comes first in `pd.concat`.

## Installation
Install with:

Expand All @@ -57,7 +63,7 @@ df = pd.DataFrame(
# Create a MicroDataFrame
mdf_df = mdf.MicroDataFrame(df, weights="weights")

# All operations are weight-aware
# Supported estimators use the row weights
print(mdf_df.income.mean()) # Weighted mean
print(mdf_df.income.gini()) # Gini coefficient
```
Expand Down
1 change: 1 addition & 0 deletions changelog.d/333.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Preserve weights through aggregation, construction, pandas conversions and row operations; reject conflicting arithmetic weights and unsupported estimators. Add checked microdf.concat and a regression-backed support matrix.
1 change: 1 addition & 0 deletions changelog.d/334.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Clarify the units of poverty gap totals and test all poverty estimators with non-uniform and zero weights and threshold boundaries.
1 change: 1 addition & 0 deletions changelog.d/335.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Describe frequency-weighted covariance and Pearson correlation in the examples, with a checked NumPy example.
26 changes: 20 additions & 6 deletions docs/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,8 @@
vector; `MicroDataFrame` is a `pandas.DataFrame` carrying a weight column. Both
behave like their pandas counterparts, and the methods below either add a
weighted estimator or preserve weights through an operation that would otherwise
drop them.
drop them. The [support matrix](support.md) lists tested operations and explicit
rejections; arbitrary pandas operations do not necessarily preserve weights.

```python
import microdf as mdf
Expand All @@ -31,6 +32,8 @@ These have the same names as their pandas equivalents and return weighted result
| `cov` | `(other: Series, min_periods: Optional[int] = None, ddof: int = 1, *, skipna: bool = True) -> float` | Calculate frequency-weighted covariance with another Series. |
| `corr` | `(other: Series, method: str = 'pearson', min_periods: Optional[int] = None, *, ddof: int = 1, skipna: bool = True) -> float` | Calculate frequency-weighted Pearson correlation. |
| `rank` | `(pct: Optional[bool] = False) -> Series` | Weighted rank of each element. |
| `value_counts` | `(normalize=False, sort=True, ascending=False, bins=None, dropna=True) -> Series` | Sum weights by value; optionally divide by the included weight total. |
| `mode` | `(dropna: bool = True) -> Series` | Return values with the greatest positive total weight as a plain Series. |

### Weight-preserving operations

Expand All @@ -44,6 +47,7 @@ Operations that change the shape or type of the data, overridden so weights stay
| `clip` | `(lower: Optional[float] = None, upper: Optional[float] = None, axis: Optional[int] = None, inplace: Optional[bool] = False, *args, **kwargs) -> MicroSeries` | Trim values at the given thresholds, preserving weights. |
| `round` | `(decimals: Optional[int] = 0, *args, **kwargs) -> MicroSeries` | Round each value, preserving weights. |
| `repeat` | `(repeats, axis=None)` | Repeat elements, repeating their weights alongside. |
| `explode` | `(ignore_index: bool = False) -> MicroSeries` | Expand list entries, repeating each observation's weight. |
| `sqrt` | `() -> MicroSeries` | Element-wise square root, preserving weights. |
| `copy` | `(deep: Optional[bool] = True)` | Copy the series and its weights. |
| `equals` | `(other: MicroSeries) -> bool` | True when both the values and the weights are equal. |
Expand Down Expand Up @@ -101,6 +105,7 @@ the series can compute, including the Gini coefficient and quantiles.
| `sum` | `(axis: Union[int, str, NoneType] = 0, skipna: bool = True, numeric_only: bool = False, min_count: int = 0, **kwargs) -> Union[Series, MicroSeries, float]` | Sum numeric columns, weighting reductions across observations. |
| `cov` | `(min_periods: Optional[int] = None, ddof: int = 1, numeric_only: bool = False) -> DataFrame` | Pairwise frequency-weighted covariance of the columns. |
| `corr` | `(method: str = 'pearson', min_periods: int = 1, numeric_only: bool = False) -> DataFrame` | Pairwise frequency-weighted Pearson correlation of the columns. |
| `pivot_table` | `(values=None, index=None, columns=None, aggfunc='mean', fill_value=None, margins=False, dropna=True, margins_name='All', observed=True, sort=True, **kwargs)` | Build a pivot table by applying estimators to weighted column groups. |

### Weight-preserving operations

Expand All @@ -110,6 +115,8 @@ the series can compute, including the Gini coefficient and quantiles.
| `merge` | `(right, how='inner', on=None, left_on=None, right_on=None, left_index=False, right_index=False, sort=False, suffixes=('_x', '_y'), copy=True, indicator=False, validate=None)` | Database-style join that carries the weight column through. |
| `reset_index` | `(level: Optional[int] = None, drop: Optional[bool] = False, inplace: Optional[bool] = False, col_level: Optional[int] = 0, col_fill: Optional[str] = '', allow_duplicates: Optional[bool] = None, names: Optional[list[str]] = None) -> Optional[MicroDataFrame]` | Reset the index, keeping weights aligned to their rows. |
| `drop` | `(labels=None, axis=0, index=None, columns=None, level=None, inplace=False, errors='raise')` | Drop rows or columns, keeping weights aligned to the remaining rows. |
| `dropna` | `(*, axis=0, how=None, thresh=None, subset=None, inplace=False, ignore_index=False)` | Drop missing observations and their weights using row positions. |
| `apply` | `(func, axis=0, raw=False, result_type=None, args=(), **kwargs)` | Apply row functions while retaining row weights on the result. |
| `astype` | `(dtype, copy: Optional[bool] = True, errors: Optional[str] = 'raise') -> MicroDataFrame` | Convert MicroDataFrame to specified data type while preserving weights. |
| `copy` | `(deep: Optional[bool] = True) -> MicroDataFrame` | Copy the frame and its weights. |
| `equals` | `(other: MicroDataFrame) -> bool` | True when both the values and the weights are equal. |
Expand All @@ -118,12 +125,18 @@ the series can compute, including the Gini coefficient and quantiles.

| Method | Signature | Description |
|---|---|---|
| `poverty_rate` | `(income: str, threshold: str) -> float` | Calculate poverty rate, i.e., the population share with income below their poverty threshold. |
| `poverty_gap` | `(income: str, threshold: str) -> float` | Calculate poverty gap, i.e., the total gap between income and poverty thresholds for all people in poverty. |
| `poverty_rate` | `(income: str, threshold: str) -> float` | Return the weighted headcount share strictly below the poverty threshold. |
| `poverty_gap` | `(income: str, threshold: str) -> float` | Return the weighted aggregate poverty gap in income currency units. |
| `poverty_count` | `(income: Union[MicroSeries, str], threshold: Union[MicroSeries, str]) -> int` | Calculates the number of entities with income below a poverty threshold. |
| `deep_poverty_rate` | `(income: str, threshold: str) -> float` | Calculate deep poverty rate, i.e., the population share with income below half their poverty threshold. |
| `deep_poverty_gap` | `(income: str, threshold: str) -> float` | Calculate deep poverty gap, i.e., the total gap between income and half of poverty thresholds for all people in deep poverty. |
| `squared_poverty_gap` | `(income: str, threshold: str) -> float` | Calculate squared poverty gap, i.e., the total squared gap between income and poverty thresholds for all people in poverty. Also known as the poverty severity index. |
| `deep_poverty_rate` | `(income: str, threshold: str) -> float` | Return the weighted headcount share strictly below half the threshold. |
| `deep_poverty_gap` | `(income: str, threshold: str) -> float` | Return the weighted aggregate deep poverty gap in income currency units. |
| `squared_poverty_gap` | `(income: str, threshold: str) -> float` | Return the weighted aggregate squared gap in squared currency units. |

The gap methods return currency totals. Normalised FGT(1) (poverty gap index)
and FGT(2) (poverty severity index) are outside this API's scope. In those
indices, divide each positive gap by that row's threshold before raising to
the first or second power, then take the population-weighted mean. People at
the threshold contribute zero, and people with zero weight do not contribute.

### Weights

Expand All @@ -137,6 +150,7 @@ the series can compute, including the Gini coefficient and quantiles.

| Function | Description |
|---|---|
| `microdf.concat` | Concatenate Micro objects; reject plain pandas inputs in either order. |
| `microdf.replicate_variance` | Variance of a statistic from replicate weights. |
| `microdf.replicate_standard_error` | Square root of the above. |

Expand Down
27 changes: 24 additions & 3 deletions docs/build_api.py
Original file line number Diff line number Diff line change
Expand Up @@ -106,7 +106,8 @@ def build():
vector; `MicroDataFrame` is a `pandas.DataFrame` carrying a weight column. Both
behave like their pandas counterparts, and the methods below either add a
weighted estimator or preserve weights through an operation that would otherwise
drop them.
drop them. The [support matrix](support.md) lists tested operations and explicit
rejections; arbitrary pandas operations do not necessarily preserve weights.

```python
import microdf as mdf
Expand Down Expand Up @@ -135,6 +136,8 @@ def build():
"cov",
"corr",
"rank",
"value_counts",
"mode",
],
)
}
Expand All @@ -153,6 +156,7 @@ def build():
"clip",
"round",
"repeat",
"explode",
"sqrt",
"copy",
"equals",
Expand Down Expand Up @@ -202,14 +206,24 @@ def build():

### Weighted aggregation

{table(FRAME, ["sum", "cov", "corr"])}
{table(FRAME, ["sum", "cov", "corr", "pivot_table"])}

### Weight-preserving operations

{
table(
FRAME,
["groupby", "merge", "reset_index", "drop", "astype", "copy", "equals"],
[
"groupby",
"merge",
"reset_index",
"drop",
"dropna",
"apply",
"astype",
"copy",
"equals",
],
)
}

Expand All @@ -229,6 +243,12 @@ def build():
)
}

The gap methods return currency totals. Normalised FGT(1) (poverty gap index)
and FGT(2) (poverty severity index) are outside this API's scope. In those
indices, divide each positive gap by that row's threshold before raising to
the first or second power, then take the population-weighted mean. People at
the threshold contribute zero, and people with zero weight do not contribute.

### Weights

{table(FRAME, ["set_weights", "set_weight_col", "nullify_weights"])}
Expand All @@ -237,6 +257,7 @@ def build():

| Function | Description |
|---|---|
| `microdf.concat` | Concatenate Micro objects; reject plain pandas inputs in either order. |
| `microdf.replicate_variance` | Variance of a statistic from replicate weights. |
| `microdf.replicate_standard_error` | Square root of the above. |

Expand Down
7 changes: 7 additions & 0 deletions docs/build_support.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
"""Generate the support matrix from the executable regression cases."""

from pathlib import Path
from microdf.tests.test_fail_closed import support_matrix

if __name__ == "__main__":
Path(__file__).with_name("support.md").write_text(support_matrix())
26 changes: 18 additions & 8 deletions docs/examples.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,14 +41,24 @@ new Micro object with appropriate weights when that choice is intentional.
A single DataFrame row is a plain pandas `Series`, because its entries are
columns rather than weighted observations.

`MicroDataFrame.cov()` and `.corr()` retain pandas' unweighted calculations
and return plain pandas `DataFrame` matrices. Their rows describe columns,
so observation weights do not apply to the result or subsequent operations
such as `.sum()`. These methods accept the installed pandas version's
arguments and defaults, including missing-value handling and correlation
methods.

Use Micro objects for **every input** to `pd.concat`. A mixed concat raises
`MicroDataFrame.cov()` computes frequency-weighted covariance, and `.corr()`
computes frequency-weighted Pearson correlation. Each matrix cell applies the
corresponding `MicroSeries` estimator to the pair of columns, excluding missing
pairs and zero-weight rows. Integer weights agree with the replicated sample.
Both methods return plain pandas `DataFrame` matrices: their rows describe
columns, so subsequent matrix operations have no observation weights. Other
correlation methods are unsupported. The example below returns the covariance
matrix `[[8/3, 4/3], [4/3, 2]]`.

```python
import numpy as np

xy = MicroDataFrame({"x": [1, 3, 5], "y": [2, 5, 4]}, weights=[1, 2, 1])
assert np.allclose(xy.cov(), np.cov(xy.to_numpy().T, fweights=[1, 2, 1]))
```

Use `microdf.concat` to reject plain pandas inputs in either order, or use Micro
objects for **every input** to `pd.concat`. A mixed concat raises
`ValueError` when pandas calls the Micro object's hooks. If a plain pandas
object comes first, pandas can bypass those hooks and return an unweighted
object; microdf cannot intercept that dispatch. Convert each input to
Expand Down
1 change: 1 addition & 0 deletions docs/myst.yml
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@ project:
children:
- file: gini.ipynb
- file: api.md
- file: support.md
site:
options:
logo: microdf_logo.png
Expand Down
54 changes: 54 additions & 0 deletions docs/support.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,54 @@
# Supported operations

Generated from `microdf/tests/test_fail_closed.py` by `uv run python docs/build_support.py`.

The regression suite checks these contracts on pandas 2 and 3. Aggregated results carry no row weights; transformations retain independent copies of the input weights.

| Operation | Behaviour |
|---|---|
| groupby column mean | Weighted result / preserved weights |
| groupby dict agg | Weighted result / preserved weights |
| groupby named agg | Weighted result / preserved weights |
| groupby callable agg | Weighted result / preserved weights |
| SeriesGroupBy callable agg | Weighted result / preserved weights |
| callable pivot_table | Weighted result / preserved weights |
| row apply | Weighted result / preserved weights |
| frame numeric mean | Weighted result / preserved weights |
| frame numeric median | Weighted result / preserved weights |
| groupby numeric mean | Weighted result / preserved weights |
| groupby numeric median | Weighted result / preserved weights |
| frame constructor | Weighted result / preserved weights |
| series constructor | Weighted result / preserved weights |
| pd.cut preserves weights | Weighted result / preserved weights |
| pd.qcut preserves weights | Weighted result / preserved weights |
| pd.to_numeric preserves weights | Weighted result / preserved weights |
| series explode | Weighted result / preserved weights |
| np.average | Weighted result / preserved weights |
| np.mean | Weighted result / preserved weights |
| np.median | Weighted result / preserved weights |
| weighted value_counts | Weighted result / preserved weights |
| weighted mode | Weighted result / preserved weights |
| numeric frame dropna | Weighted result / preserved weights |
| `rolling` | Raises; use `pd.Series(s)` for unweighted pandas behaviour |
| `expanding` | Raises; use `pd.Series(s)` for unweighted pandas behaviour |
| `ewm` | Raises; use `pd.Series(s)` for unweighted pandas behaviour |
| `sem` | Raises; use `pd.Series(s)` for unweighted pandas behaviour |
| `skew` | Raises; use `pd.Series(s)` for unweighted pandas behaviour |
| `kurt` | Raises; use `pd.Series(s)` for unweighted pandas behaviour |
| `kurtosis` | Raises; use `pd.Series(s)` for unweighted pandas behaviour |
| `prod` | Raises; use `pd.Series(s)` for unweighted pandas behaviour |
| `product` | Raises; use `pd.Series(s)` for unweighted pandas behaviour |
| `idxmax` | Raises; use `pd.Series(s)` for unweighted pandas behaviour |
| `idxmin` | Raises; use `pd.Series(s)` for unweighted pandas behaviour |
| Arithmetic with conflicting weights | Raises `ValueError` in either operand order |

`pd.cut` and `pd.qcut` retain row weights, but choose bin edges using pandas' unweighted rules. Supply explicit bin edges for weighted quantile bins.

`pivot_table` grouping keys must name columns; external Series, callable
groupers and index-level groupers raise. Weighted `Series.value_counts` and
`Series.mode` return plain summary Series; frame and grouped variants raise.

`microdf.concat` rejects mixed weighted/plain inputs in either order. Direct
`pd.concat` still bypasses microdf when its first input is plain pandas; the
regression suite records this upstream dispatch limitation as an expected failure.
Use Micro objects for every input to `pd.concat`, or use `microdf.concat`.
2 changes: 2 additions & 0 deletions microdf/__init__.py
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
from importlib.metadata import PackageNotFoundError, version

from .concat import concat
from .microdataframe import MicroDataFrame, MicroDataFrameGroupBy
from .microseries import MicroSeries, MicroSeriesGroupBy
from .replication import replicate_standard_error, replicate_variance
Expand All @@ -14,6 +15,7 @@
__version__ = "unknown"

__all__ = [
"concat",
# microseries.py
"MicroSeries",
"MicroSeriesGroupBy",
Expand Down
Loading
Loading