Skip to content
Back to skills

Pandas

ASecurity

Analyze and transform tabular data with pandas: dataframes, cleaning, aggregation, joins, and time series. Use for any data analysis or ETL task.

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 29, 2026
ai-agentspythongobashapiperformance

Works with

  • api

Security analysis

A96/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 2 files and shows the line behind each finding

Scanned September 29, 2026

npx -y skills add ssrjkk/claude-skills --skill pandas --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Pandas?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Pandas
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/ssrjkk-pandas/badge)](https://www.skillsdirectory.com/skills/ssrjkk-pandas)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: pandas
description: "Analyze and transform tabular data with pandas: dataframes, cleaning, aggregation, joins, and time series. Use for any data analysis or ETL task."
category: data
tags: [pandas, python, data-analysis, dataframe, etl, csv, time-series]
models: [sonnet, opus, gpt-5, gemini-2.5, glm-4.6]
version: 1.0.0
created: 2026-09-20
updated: 2026-09-28
author: ssrjkk
---
# Pandas

> Analyzing and transforming tabular data with pandas.

## Quick Start
```bash
pip install pandas
python -c "import pandas as pd; print(pd.read_csv('data.csv').head())"
```

## When to Use
- Cleaning and reshaping tabular datasets
- Aggregations, joins, and pivots
- Time series analysis and resampling
- ETL steps before machine learning

## Best Practices

### Data Loading
- Specify `dtype`, `parse_dates`, and `usecols` to avoid surprises
- Use `read_parquet`/`read_csv` with explicit schema
- Chunk large files with `chunksize` or use Dask/Polars for scale
- Validate shape and nulls right after loading

### Cleaning
- Use vectorized ops; avoid row-wise loops
- Handle missing data deliberately (`fillna` vs `dropna`)
- Normalize types: datetime, categorical, numeric
- Keep an audit trail of transformations

### Performance
- Prefer vectorized operations over `.apply`/`.iterrows`
- Use `groupby` aggregations instead of loops
- Convert object columns to categorical for speed
- Set indexes for joins; avoid merge on object columns

## Dependencies
```bash
pip install pandas numpy
# scale: polars, dask, pyarrow
```

## Examples
```python
import pandas as pd

df = pd.read_csv("sales.csv", parse_dates=["date"], dtype={"qty": "int32"})
df.info()
print(df.isna().sum())
```
```python
# Cleaning
df["price"] = df["price"].str.replace("$", "").astype(float)
df = df.dropna(subset=["order_id"])
df["date"] = pd.to_datetime(df["date"], errors="coerce")
df = df[df["qty"] > 0]
```
```python
# Aggregation
by_region = (
    df.groupby("region", as_index=False)
      .agg(total=("amount", "sum"), orders=("order_id", "nunique"))
      .sort_values("total", ascending=False)
)
print(by_region.head())
```
```python
# Time series resampling
daily = df.set_index("date")["amount"].resample("D").sum()
monthly = daily.resample("ME").mean()
print(monthly.head())
```

## Step-by-Step
1. Load data with explicit dtypes and date parsing.
2. Profile: shape, columns, nulls, and unique counts.
3. Clean types and missing values deliberately.
4. Filter and transform with vectorized operations.
5. Aggregate with `groupby`; reshape with pivot/merge.
6. For time series, resample and window correctly.
7. Export with explicit schema (parquet preferred).
8. Verify outputs against the source invariants.

## Validation
1. No unexpected nulls or type errors after cleaning
2. Aggregations match manually computed totals
3. Joins don't multiply rows unexpectedly
4. Date ranges and resampling are correct
5. Results reproducible with a fixed seed (if sampling)

## Troubleshooting
- MemoryError: use dtypes, chunking, or Polars.
- Wrong joins: check key types and duplicates before merge.
- Slow loops: convert to vectorized or `groupby` operations.

Files in this skill

  • SKILL.md3.1 KB
  • SKILL.ru.md4.4 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…