Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions docs/docs/get_started.md
Original file line number Diff line number Diff line change
Expand Up @@ -75,6 +75,9 @@ plot(locations, dte, lower_bound, upper_bound, title="DTE of simple estimator")

![DTE of empirical estimator](assets/dte_empirical.png)

!!! note "Input requirements"
`treatment_arms`, `outcomes` and `strata` must not contain missing values (NaN); `fit` raises a `ValueError` otherwise. They may be 1-D arrays or single columns of shape `(n, 1)`. Missing values in `covariates` are not checked: simple estimators ignore covariates, while for adjusted estimators it depends on whether your base model accepts them (e.g. scikit-learn's `HistGradientBoostingClassifier` does, `LogisticRegression` does not).

To initialize the adjusted distribution function, the base model for conditional distribution function needs to be passed.
In the following example, Logistic Regression is used. Please make sure that your base model implements `fit` and `predict_proba` methods.

Expand Down Expand Up @@ -135,6 +138,8 @@ plot(quantiles, qte, lower_bound, upper_bound, title="QTE of adjusted estimator"
![QTE of adjusted estimator](assets/qte.png)

You can use any model with `predict_proba` or `predict` method to adjust the distribution function estimation.

The adjustment is meant for randomized experiments: it reduces the variance of the estimates (tighter confidence intervals) and does not correct for confounding. Cross-fitting assigns folds at random, so with very small samples a fold can end up with no training data for a treatment arm; in that case a `ValueError` is raised and you should reduce `folds` (e.g. `folds=2`).
For example, the following code use XGBoost classifier to estimate the conditional distribution.

```python
Expand Down
10 changes: 5 additions & 5 deletions docs/docs/tutorials/oregon.md
Original file line number Diff line number Diff line change
Expand Up @@ -531,7 +531,7 @@ plt.show()
- Simple: LDTE ≈ -0.55 at zero costs, converging to zero around $15,000-$20,000
- ML-Adjusted: LDTE ≈ -0.10 to -0.15 at zero costs, stable pattern with improved confidence intervals
- Much larger magnitude effects in the Simple estimator, indicating households with multiple members show substantially stronger treatment effects
- ML adjustment provides more conservative estimates, potentially controlling for confounding household characteristics
- ML adjustment provides more conservative estimates and tighter confidence intervals by exploiting household characteristics that predict the outcome (the lottery itself randomizes treatment, so this is variance reduction rather than confounding correction)

**2. Heterogeneity Across Strata**

Expand All @@ -540,7 +540,7 @@ The stratified analysis reveals substantial treatment effect heterogeneity:
- **"Signed self up" stratum**: Moderate effects (LDTE ≈ -0.18 to -0.20), suggesting single-person households have more modest increases in ED utilization
- **"Signed self up + others" stratum**: Large effects in Simple estimator (LDTE ≈ -0.55), suggesting multi-person households experience much greater increases in ED access when not adjusting for covariates
- The 3-4x larger effect in the "signed self up + others" group (Simple estimator) indicates that household composition is a critical moderator of insurance impact
- However, ML adjustment substantially reduces this estimate, suggesting that some of the observed effect may be attributable to observable household characteristics rather than pure treatment effects
- However, ML adjustment substantially reduces this estimate. Because treatment is randomized, the adjusted and unadjusted estimators target the same quantity, so a large gap is more likely sampling noise in the small stratum (reduced by the adjustment) than a difference in what is being estimated

**3. Comparison of Estimation Methods**

Expand All @@ -551,14 +551,14 @@ The stratified analysis reveals substantial treatment effect heterogeneity:
- The Simple estimator shows the largest treatment effects across all strata (LDTE ≈ -0.55)
- ML adjustment substantially reduces the estimated effect and stabilizes confidence intervals
- This divergence suggests that observable covariates (e.g., household size, age composition, baseline health status) explain a significant portion of the treatment effect heterogeneity
- The improved stability of ML-adjusted estimates indicates successful control for confounding factors that may have been correlated with both treatment assignment and outcomes
- The improved stability of ML-adjusted estimates reflects variance reduction from covariates that are predictive of the outcome

**4. Practical Implications**

- **Household structure matters**: Multi-person households show substantially larger treatment effects in unadjusted analyses, likely because insurance coverage enables care-seeking for multiple family members.
- **The role of covariates**: The difference between Simple and ML-adjusted estimates in the "signed self up + others" stratum highlights the importance of controlling for household characteristics. The unadjusted effect may overstate the pure treatment effect by conflating insurance provision with pre-existing household differences.
- **The role of covariates**: The difference between Simple and ML-adjusted estimates in the "signed self up + others" stratum highlights how much precision household characteristics can add. The unadjusted estimate is noisier in this small stratum and may overstate the effect by chance.
- **Stratification reveals hidden heterogeneity**: The overall population estimate masks substantial variation across household types, demonstrating the value of subgroup analysis.
- **Model specification considerations**: ML adjustment improves estimation stability in smaller strata and provides more defensible causal estimates by controlling for observable confounders. The convergence of all estimates to zero at higher cost levels confirms that the treatment primarily affects the lower tail of the cost distribution.
- **Model specification considerations**: ML adjustment improves estimation stability in smaller strata and provides more precise estimates by exploiting predictive covariates (it is not a correction for confounding, since treatment is randomized). The convergence of all estimates to zero at higher cost levels confirms that the treatment primarily affects the lower tail of the cost distribution.

### Conclusion

Expand Down
45 changes: 23 additions & 22 deletions dte_adj/base.py
Original file line number Diff line number Diff line change
Expand Up @@ -418,30 +418,31 @@ def _compute_qtes(
locations = np.sort(outcomes)

def find_quantile(quantile, arm):
# Smallest location y with F(y) >= quantile, i.e. the generalized inverse
# of the estimated CDF. Returns the largest location if no such y exists.
low, high = 0, locations.shape[0] - 1
result = -1
while low <= high:
mid = (low + high) // 2
# Temporarily store original strata and use the provided strata
original_strata = self.strata
self.strata = strata

val, _, _ = self._compute_cumulative_distribution(
arm,
np.full((1), locations[mid]),
covariates,
treatment_arms,
outcomes,
)

# Restore original strata
result = locations[-1]
# Temporarily use the provided strata when evaluating the CDF
original_strata = self.strata
self.strata = strata
try:
while low <= high:
mid = (low + high) // 2
val, _, _ = self._compute_cumulative_distribution(
arm,
np.full((1), locations[mid]),
covariates,
treatment_arms,
outcomes,
)

if val[0] >= quantile - 1e-12:
result = locations[mid]
high = mid - 1
else:
low = mid + 1
finally:
self.strata = original_strata

if val[0] <= quantile:
result = locations[mid]
low = mid + 1
else:
high = mid - 1
return result

result = np.zeros(quantiles.shape)
Expand Down
38 changes: 24 additions & 14 deletions dte_adj/local.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,8 @@
ArrayLike,
compute_ldte,
compute_lpte,
_convert_to_ndarray,
_to_1d,
_check_no_missing,
_infer_default_locations,
)

Expand Down Expand Up @@ -55,8 +56,13 @@ def fit(
Returns:
SimpleLocalDistributionEstimator: The fitted estimator.
"""
treatment_indicator = _convert_to_ndarray(treatment_indicator)
treatment_indicator = _to_1d("treatment_indicator", treatment_indicator)
super().fit(covariates, treatment_arms, outcomes, strata)
if treatment_indicator.shape[0] != self.covariates.shape[0]:
raise ValueError(
"The shape of covariates and treatment_indicator should be same"
)
_check_no_missing("treatment_indicator", treatment_indicator)
self.treatment_indicator = treatment_indicator

return self
Expand Down Expand Up @@ -219,9 +225,10 @@ class AdjustedLocalDistributionEstimator(AdjustedStratifiedDistributionEstimator
A class for computing Local Distribution Treatment Effects (LDTE) and Local Probability
Treatment Effects (LPTE) using machine learning adjustment.

This estimator combines the benefits of ML adjustment with local treatment effect estimation,
providing precise estimates of treatment effects that are weighted by treatment propensity
within each stratum. It uses cross-fitting to avoid overfitting issues.
This estimator combines ML adjustment with local treatment effect estimation, providing
variance-reduced estimates in randomized experiments with partial compliance. The ML
adjustment targets efficiency (tighter confidence intervals), not correction for
confounding. It uses cross-fitting to avoid overfitting issues.
"""

def fit(
Expand All @@ -245,8 +252,13 @@ def fit(
Returns:
AdjustedLocalDistributionEstimator: The fitted estimator.
"""
treatment_indicator = _convert_to_ndarray(treatment_indicator)
treatment_indicator = _to_1d("treatment_indicator", treatment_indicator)
super().fit(covariates, treatment_arms, outcomes, strata)
if treatment_indicator.shape[0] != self.covariates.shape[0]:
raise ValueError(
"The shape of covariates and treatment_indicator should be same"
)
_check_no_missing("treatment_indicator", treatment_indicator)
self.treatment_indicator = treatment_indicator

return self
Expand Down Expand Up @@ -288,13 +300,12 @@ def predict_ldte(
from sklearn.ensemble import RandomForestClassifier
from dte_adj import AdjustedLocalDistributionEstimator

# Generate confounded data with strata
# Generate data with strata from a randomized experiment
np.random.seed(42)
X = np.random.randn(1000, 5)
strata = np.random.choice([0, 1], size=1000)
# Treatment assignment depends on covariates
Z_prob = 1 / (1 + np.exp(-(X[:, 0] + X[:, 1] + strata)))
Z = np.random.binomial(1, Z_prob, 1000)
# Treatment assignment is random (independent of covariates)
Z = np.random.binomial(1, 0.5, 1000)
D = np.random.binomial(1, 0.3 + 0.4 * Z, 1000)
Y = X.sum(axis=1) + 2 * D + strata + np.random.randn(1000)

Expand Down Expand Up @@ -366,13 +377,12 @@ def predict_lpte(
from sklearn.linear_model import LogisticRegression
from dte_adj import AdjustedLocalDistributionEstimator

# Generate confounded data with strata
# Generate data with strata from a randomized experiment
np.random.seed(42)
X = np.random.randn(1000, 5)
strata = np.random.choice([0, 1], size=1000)
# Treatment assignment depends on covariates
Z_prob = 1 / (1 + np.exp(-(X[:, 0] + strata)))
Z = np.random.binomial(1, Z_prob, 1000)
# Treatment assignment is random (independent of covariates)
Z = np.random.binomial(1, 0.5, 1000)
D = np.random.binomial(1, 0.3 + 0.4 * Z, 1000)
Y = X.sum(axis=1) + 2 * D + strata + np.random.randn(1000)

Expand Down
56 changes: 26 additions & 30 deletions dte_adj/simple.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@
SimpleStratifiedDistributionEstimator,
AdjustedStratifiedDistributionEstimator,
)
from dte_adj.util import ArrayLike, _convert_to_ndarray
from dte_adj.util import ArrayLike, _prepare_fit_inputs


class SimpleDistributionEstimator(SimpleStratifiedDistributionEstimator):
Expand All @@ -15,8 +15,8 @@ class SimpleDistributionEstimator(SimpleStratifiedDistributionEstimator):

This estimator computes Distribution Treatment Effects (DTE), Probability Treatment Effects (PTE),
and Quantile Treatment Effects (QTE) without using machine learning models for adjustment.
It provides a baseline approach suitable when treatment assignment is random or when
covariate adjustment is not needed.
It provides a baseline approach for randomized experiments where covariate adjustment
is not needed.

Example:
```python
Expand Down Expand Up @@ -61,20 +61,14 @@ def fit(
Returns:
SimpleDistributionEstimator: The fitted estimator.
"""
covariates = _convert_to_ndarray(covariates)
treatment_arms = _convert_to_ndarray(treatment_arms)
outcomes = _convert_to_ndarray(outcomes)

if covariates.shape[0] != treatment_arms.shape[0]:
raise ValueError("The shape of covariates and treatment_arm should be same")

if covariates.shape[0] != outcomes.shape[0]:
raise ValueError("The shape of covariates and outcome should be same")
covariates, treatment_arms, outcomes, strata = _prepare_fit_inputs(
covariates, treatment_arms, outcomes
)

self.covariates = covariates
self.treatment_arms = treatment_arms
self.outcomes = outcomes
self.strata = np.zeros(len(self.covariates))
self.strata = strata

return self

Expand All @@ -83,21 +77,29 @@ class AdjustedDistributionEstimator(AdjustedStratifiedDistributionEstimator):
"""
A class for computing distribution treatment effects using machine learning adjustment.

This estimator uses cross-fitting with ML models to adjust for confounding when computing
Distribution Treatment Effects (DTE), Probability Treatment Effects (PTE), and
Quantile Treatment Effects (QTE). It provides more precise estimates when treatment
assignment depends on observed covariates.
This estimator uses cross-fitting with ML models of the conditional distribution given
covariates to compute Distribution Treatment Effects (DTE), Probability Treatment Effects
(PTE), and Quantile Treatment Effects (QTE) with reduced variance. It is designed for
randomized experiments, where treatment is assigned independently of the covariates: in
that setting the estimator stays consistent regardless of how well the ML model fits, and
a good fit yields tighter confidence intervals than ``SimpleDistributionEstimator``.

Note:
The method targets efficiency, not bias correction. It does not correct for
confounding in observational data; if treatment assignment depends on covariates, the
estimates are valid only under selection on observables, and ``fit`` assumes the
assignment probability is constant across observations (no propensity weighting).

Example:
```python
import numpy as np
from sklearn.ensemble import RandomForestClassifier
from dte_adj import AdjustedDistributionEstimator

# Generate confounded data
# Generate data from a randomized experiment: treatment is independent of X,
# while the outcome depends on X (this is what the ML adjustment exploits)
X = np.random.randn(1000, 5)
treatment_prob = 1 / (1 + np.exp(-(X[:, 0] + X[:, 1])))
D = np.random.binomial(1, treatment_prob, 1000)
D = np.random.binomial(1, 0.5, 1000)
Y = X.sum(axis=1) + 2 * D + np.random.randn(1000)

# Fit adjusted estimator
Expand Down Expand Up @@ -125,19 +127,13 @@ def fit(
Returns:
AdjustedDistributionEstimator: The fitted estimator.
"""
covariates = _convert_to_ndarray(covariates)
treatment_arms = _convert_to_ndarray(treatment_arms)
outcomes = _convert_to_ndarray(outcomes)

if covariates.shape[0] != treatment_arms.shape[0]:
raise ValueError("The shape of covariates and treatment_arm should be same")

if covariates.shape[0] != outcomes.shape[0]:
raise ValueError("The shape of covariates and outcome should be same")
covariates, treatment_arms, outcomes, strata = _prepare_fit_inputs(
covariates, treatment_arms, outcomes
)

self.covariates = covariates
self.treatment_arms = treatment_arms
self.outcomes = outcomes
self.strata = np.zeros(len(self.covariates))
self.strata = strata

return self
Loading
Loading