Skip to content
Back to skills

Regression To Cdf Smoothing

ASecurity

Convert a scalar regression prediction into a smoothed CDF over discrete bins using a linear ramp instead of a hard step

  • 61 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 12, 2026
testingpython

Works with

  • cli

Security analysis

A100/100

Scanned September 12, 2026

npx -y skills add wenmin-wu/ds-skills --skill regression-to-cdf-smoothing --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Regression To Cdf Smoothing?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Regression To Cdf Smoothing
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/wenmin-wu-regression-to-cdf-smoothing/badge)](https://www.skillsdirectory.com/skills/wenmin-wu-regression-to-cdf-smoothing)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: tabular-regression-to-cdf-smoothing
description: Convert a scalar regression prediction into a smoothed CDF over discrete bins using a linear ramp instead of a hard step
domain: tabular
---

# Regression to CDF Smoothing

## Overview

When the evaluation metric is CRPS on a discrete CDF but your model outputs a scalar prediction, convert it to a smooth CDF using a linear ramp (or sigmoid) centered on the predicted value. A hard step function is overconfident; a ramp with width W spreads probability over ±W bins, hedging against prediction error.

## Quick Start

```python
import numpy as np

def scalar_to_cdf(predictions, n_bins=199, offset=99, ramp_width=10):
    """Convert scalar predictions to smoothed CDFs.
    
    Args:
        predictions: array of scalar predictions
        n_bins: number of CDF bins
        offset: bin index for target=0
        ramp_width: half-width of linear ramp (bins)
    Returns:
        (N, n_bins) CDF array
    """
    cdf = np.zeros((len(predictions), n_bins))
    for i, pred in enumerate(predictions):
        center = int(round(pred)) + offset
        for j in range(n_bins):
            if j >= center + ramp_width:
                cdf[i, j] = 1.0
            elif j >= center - ramp_width:
                cdf[i, j] = (j - center + ramp_width) / (2 * ramp_width)
    return np.clip(cdf, 0, 1)

# Usage with LightGBM
y_pred_scalar = np.mean([m.predict(X_test) for m in models], axis=0)
y_pred_cdf = scalar_to_cdf(y_pred_scalar, ramp_width=10)
```

## Key Decisions

- **Ramp width**: 10 bins is conservative; optimize on validation CRPS — wider = safer, narrower = sharper
- **Linear vs sigmoid**: linear ramp is simple and effective; sigmoid is smoother but adds a hyperparameter
- **Physical clipping**: enforce domain bounds (e.g., can't gain more yards than field remaining)
- **Ensemble first**: average scalar predictions before converting to CDF for best results

## References

- Source: [nfl-simple-model-using-lightgbm](https://www.kaggle.com/code/hukuda222/nfl-simple-model-using-lightgbm)
- Competition: NFL Big Data Bowl

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…