Skip to content
Back to skills

Alterlab Statsmodels

ASecurity

Statistical modeling in Python with statsmodels — OLS/WLS/GLS, GLM, discrete-choice and count models, mixed models, ARIMA/SARIMAX/VAR, with diagnostics, robust standard errors, and coefficient-level inference. Use when fitting specific model classes for econometrics, time series, or rigorous inference with coefficient tables and confidence intervals, or when updating code for statsmodels 0.15 (result_object named results, rng keyword). For guided statistical test selection with APA reporting ...

  • 68 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added May 27, 2026
data-aipythongobashtestinggitapi

Works with

  • api

Security analysis

A100/100

Pro scans all 6 files and shows the line behind each finding

Scanned September 23, 2026

npx -y skills add AlterLab-IEU/AlterLab-Academic-Skills --skill alterlab-statsmodels --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Alterlab Statsmodels?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Alterlab Statsmodels
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/alterlab-ieu-alterlab-statsmodels/badge)](https://www.skillsdirectory.com/skills/alterlab-ieu-alterlab-statsmodels)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: alterlab-statsmodels
description: Statistical modeling in Python with statsmodels — OLS/WLS/GLS, GLM, discrete-choice and count models, mixed models, ARIMA/SARIMAX/VAR, with diagnostics, robust standard errors, and coefficient-level inference. Use when fitting specific model classes for econometrics, time series, or rigorous inference with coefficient tables and confidence intervals, or when updating code for statsmodels 0.15 (result_object named results, rng keyword). For guided statistical test selection with APA reporting prefer alterlab-statistical-analysis. Part of the AlterLab Academic Skills suite.
license: MIT
allowed-tools: Read Write Edit Bash(python:*) Bash(uv:*)
compatibility: No API key required. Runs locally via `uv run python`; requires statsmodels >= 0.14 (current 0.15.0 as of 2026-09; Python >= 3.10). 0.15 adds formulaic as a required dependency and accepts Polars DataFrames.
metadata:
    skill-author: AlterLab
    version: "1.1.0"
    last_updated: "2026-09-23"
---

# Statsmodels: Statistical Modeling and Econometrics

## Overview

Statsmodels is Python's premier library for statistical modeling, providing tools for estimation, inference, and diagnostics across a wide range of statistical methods. Apply this skill for rigorous statistical analysis, from simple linear regression to complex time series models and econometric analyses.

## When to Use This Skill

This skill should be used when:
- Fitting regression models (OLS, WLS, GLS, quantile regression)
- Performing generalized linear modeling (logistic, Poisson, Gamma, etc.)
- Analyzing discrete outcomes (binary, multinomial, count, ordinal)
- Conducting time series analysis (ARIMA, SARIMAX, VAR, forecasting)
- Running statistical tests and diagnostics
- Testing model assumptions (heteroskedasticity, autocorrelation, normality)
- Detecting outliers and influential observations
- Comparing models (AIC/BIC, likelihood ratio tests)
- Estimating causal effects
- Producing publication-ready statistical tables and inference

### Does NOT Trigger

| Scenario | Use Instead |
|----------|-------------|
| Choosing which test fits the design, checking assumptions, and writing APA-style results | `alterlab-statistical-analysis` |
| Bayesian or hierarchical models with posterior distributions (PyMC, NUTS) | `alterlab-pymc` |
| Quasi-experimental causal designs (DiD, IV, RD, panel fixed effects, event studies) as a full workflow | `alterlab-causal-inference` |
| Predictive machine learning with cross-validated tuning rather than coefficient inference | `alterlab-scikit-learn` |
| Zero-shot forecasting with a pretrained foundation model instead of fitting ARIMA/ETS | `alterlab-timesfm` |

## statsmodels 0.15 Notes

statsmodels 0.15.0 (Aug 2026) is the first release since 0.14 (2023). Changes that affect everyday code:

- **Named results**: `adfuller`, `kpss`, `acf` (with `qstat`/`alpha`), `het_arch`, `acorr_lm`, `acorr_breusch_godfrey`, `het_goldfeldquandt` and a few others still return the legacy tuple but emit a `FutureWarning`; pass `result_object=True` to get the named result now (e.g. `adfuller(y, result_object=True).pvalue`, `het_arch(r, result_object=True).lmpval`). The default switches in 0.16.
- **Randomness**: `seed=` / `random_state=` arguments are deprecated in favor of `rng=` (SPEC 7).
- **Removed**: `grangercausalitytests(verbose=...)`, `AutoReg(old_names=...)`; `kpss(nlags=None)` now raises — pass `'auto'`, `'legacy'`, or an integer.
- **Formulas**: patsy remains the default engine; formulaic is now a required dependency and can be selected with `SM_FORMULA_ENGINE=formulaic`. Model and formula APIs also accept Polars DataFrames.
- **New**: Games-Howell post-hoc comparisons (`pairwise_tukeyhsd(..., use_var="unequal")`), the Leybourne-McCabe stationarity test (`statsmodels.tsa.stattools.leybourne`), `MultivariateLS`, and robust MM/S estimators.

## Quick Start

Copy-paste starting points for OLS, logistic regression, ARIMA, GLM/Poisson, and the R-style formula API are in `references/quickstart_examples.md`. Core rule: always `sm.add_constant()` for an intercept unless you deliberately want none.

## Core Statistical Modeling Capabilities

### 1. Linear Regression Models

Comprehensive suite of linear models for continuous outcomes with various error structures.

**Available models:**
- **OLS**: Standard linear regression with i.i.d. errors
- **WLS**: Weighted least squares for heteroskedastic errors
- **GLS**: Generalized least squares for arbitrary covariance structure
- **GLSAR**: GLS with autoregressive errors for time series
- **Quantile Regression**: Conditional quantiles (robust to outliers)
- **Mixed Effects**: Hierarchical/multilevel models with random effects
- **Recursive/Rolling**: Time-varying parameter estimation

**Key features:**
- Comprehensive diagnostic tests
- Robust standard errors (HC, HAC, cluster-robust)
- Influence statistics (Cook's distance, leverage, DFFITS)
- Hypothesis testing (F-tests, Wald tests)
- Model comparison (AIC, BIC, likelihood ratio tests)
- Prediction with confidence and prediction intervals

**When to use:** Continuous outcome variable, want inference on coefficients, need diagnostics

**Reference:** See `references/linear_models.md` for detailed guidance on model selection, diagnostics, and best practices.

### 2. Generalized Linear Models (GLM)

Flexible framework extending linear models to non-normal distributions.

**Distribution families:**
- **Binomial**: Binary outcomes or proportions (logistic regression)
- **Poisson**: Count data
- **Negative Binomial**: Overdispersed counts
- **Gamma**: Positive continuous, right-skewed data
- **Inverse Gaussian**: Positive continuous with specific variance structure
- **Gaussian**: Equivalent to OLS
- **Tweedie**: Flexible family for semi-continuous data

**Link functions:**
- Logit, Probit, Log, Identity, Inverse, Sqrt, CLogLog, Power
- Choose based on interpretation needs and model fit

**Key features:**
- Maximum likelihood estimation via IRLS
- Deviance and Pearson residuals
- Goodness-of-fit statistics
- Pseudo R-squared measures
- Robust standard errors

**When to use:** Non-normal outcomes, need flexible variance and link specifications

**Reference:** See `references/glm.md` for family selection, link functions, interpretation, and diagnostics.

### 3. Discrete Choice Models

Models for categorical and count outcomes.

**Binary models:**
- **Logit**: Logistic regression (odds ratios)
- **Probit**: Probit regression (normal distribution)

**Multinomial models:**
- **MNLogit**: Unordered categories (3+ levels)
- **Conditional Logit**: Choice models with alternative-specific variables
- **Ordered Model**: Ordinal outcomes (ordered categories)

**Count models:**
- **Poisson**: Standard count model
- **Negative Binomial**: Overdispersed counts
- **Zero-Inflated**: Excess zeros (ZIP, ZINB)
- **Hurdle Models**: Two-stage models for zero-heavy data

**Key features:**
- Maximum likelihood estimation
- Marginal effects at means or average marginal effects
- Model comparison via AIC/BIC
- Predicted probabilities and classification
- Goodness-of-fit tests

**When to use:** Binary, categorical, or count outcomes

**Reference:** See `references/discrete_choice.md` for model selection, interpretation, and evaluation.

### 4. Time Series Analysis

Comprehensive time series modeling and forecasting capabilities.

**Univariate models:**
- **AutoReg (AR)**: Autoregressive models
- **ARIMA**: Autoregressive integrated moving average
- **SARIMAX**: Seasonal ARIMA with exogenous variables
- **Exponential Smoothing**: Simple, Holt, Holt-Winters
- **ETS**: Innovations state space models

**Multivariate models:**
- **VAR**: Vector autoregression
- **VARMAX**: VAR with MA and exogenous variables
- **Dynamic Factor Models**: Extract common factors
- **VECM**: Vector error correction models (cointegration)

**Advanced models:**
- **State Space**: Kalman filtering, custom specifications
- **Regime Switching**: Markov switching models
- **ARDL**: Autoregressive distributed lag

**Key features:**
- ACF/PACF analysis for model identification
- Stationarity tests (ADF, KPSS)
- Forecasting with prediction intervals
- Residual diagnostics (Ljung-Box, heteroskedasticity)
- Granger causality testing
- Impulse response functions (IRF)
- Forecast error variance decomposition (FEVD)

**When to use:** Time-ordered data, forecasting, understanding temporal dynamics

**Reference:** See `references/time_series.md` for model selection, diagnostics, and forecasting methods.

### 5. Statistical Tests and Diagnostics

Extensive testing and diagnostic capabilities for model validation.

**Residual diagnostics:**
- Autocorrelation tests (Ljung-Box, Durbin-Watson, Breusch-Godfrey)
- Heteroskedasticity tests (Breusch-Pagan, White, ARCH)
- Normality tests (Jarque-Bera, Omnibus, Anderson-Darling, Lilliefors)
- Specification tests (RESET, Harvey-Collier)

**Influence and outliers:**
- Leverage (hat values)
- Cook's distance
- DFFITS and DFBETAs
- Studentized residuals
- Influence plots

**Hypothesis testing:**
- t-tests (one-sample, two-sample, paired)
- Proportion tests
- Chi-square tests
- Non-parametric tests (Mann-Whitney, Wilcoxon, Kruskal-Wallis)
- ANOVA (one-way, two-way, repeated measures)

**Multiple comparisons:**
- Tukey's HSD
- Bonferroni correction
- False Discovery Rate (FDR)

**Effect sizes and power:**
- Cohen's d, eta-squared
- Power analysis for t-tests, proportions
- Sample size calculations

**Robust inference:**
- Heteroskedasticity-consistent SEs (HC0-HC3)
- HAC standard errors (Newey-West)
- Cluster-robust standard errors

**When to use:** Validating assumptions, detecting problems, ensuring robust inference

**Reference:** See `references/stats_diagnostics.md` for comprehensive testing and diagnostic procedures.

## Formula API, Model Selection, Workflows

- **R-style formula API** (`smf.ols`, `smf.logit`, `smf.poisson`, interactions, `C()`, `I()`) → `references/quickstart_examples.md`.
- **Model selection and comparison** (AIC/BIC tables, likelihood ratio test for nested models, k-fold cross-validation) → `references/model_selection.md`.
- **Best practices, end-to-end workflows (OLS, logistic, count, time series), and common pitfalls** → `references/workflows_and_practices.md`.

## Routing Guidance

- Linear/continuous outcome, inference + diagnostics → Capability 1 + `references/linear_models.md`.
- Non-normal outcome, flexible link/variance → Capability 2 + `references/glm.md`.
- Binary, categorical, or count outcome → Capability 3 + `references/discrete_choice.md`.
- Time-ordered data, forecasting → Capability 4 + `references/time_series.md`.
- Validating assumptions, testing, robust inference → Capability 5 + `references/stats_diagnostics.md`.
- Choosing between candidate models → `references/model_selection.md`.

## References Index

- `references/quickstart_examples.md` — copy-paste OLS / Logit / ARIMA / GLM examples and the R-style formula API.
- `references/linear_models.md` — OLS, WLS, GLS, GLSAR, quantile, mixed effects, recursive/rolling; diagnostics, influence, robust SEs, hypothesis testing.
- `references/glm.md` — all distribution families, link functions, interpretation, pseudo R-squared, residual analysis.
- `references/discrete_choice.md` — binary (Logit/Probit), multinomial, count (Poisson/NB/ZIP/ZINB/hurdle), ordinal, marginal effects.
- `references/time_series.md` — AR/ARIMA/SARIMAX/ETS, VAR/VARMAX/dynamic factor, state space, stationarity, forecasting, Granger/IRF/FEVD.
- `references/stats_diagnostics.md` — residual diagnostics, influence/outliers, parametric and non-parametric tests, ANOVA, multiple comparisons, robust covariances, power/effect sizes.
- `references/model_selection.md` — AIC/BIC comparison, likelihood ratio test, cross-validation.
- `references/workflows_and_practices.md` — best practices, end-to-end workflows, common pitfalls, search patterns, official docs links.

Part of the AlterLab Academic Skills suite.

Files in this skill

  • SKILL.md19.3 KB
  • references/discrete_choice.md16.8 KB
  • references/glm.md16.2 KB
  • references/linear_models.md12.7 KB
  • references/stats_diagnostics.md20 KB
  • references/time_series.md18.5 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…