Skip to content
Back to skills

Image Perf

ASecurity

Optimize image load/decode/processing throughput — find the bottleneck before changing code

  • 3 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 3, 2026
ai-agentspythongo

Security analysis

A100/100

Scanned September 3, 2026

npx -y skills add black141312/ada --skill image-perf --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Image Perf?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Image Perf
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/black141312-image-perf/badge)](https://www.skillsdirectory.com/skills/black141312-image-perf)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: image-perf
description: Optimize image load/decode/processing throughput — find the bottleneck before changing code
category: image
---

# Image Perf

Reach for this when image loading, decoding, or processing is too slow and you need to raise throughput without guessing.

1. Reproduce on a representative batch and measure end-to-end time plus a per-stage breakdown (download vs decode vs resize vs model/encode). Optimize the dominant stage — never the one you assume.
2. Determine if you're I/O-bound or CPU-bound: high wait/low CPU means I/O (network/disk), saturated cores means decode/compute. The diagnosis dictates the fix (concurrency vs faster codec/SIMD vs fewer pixels).
3. Decode at target size, not full res: use JPEG scaled decoding / draft mode / thumbnail-on-load so you never decode pixels you'll immediately throw away — often the single biggest win.
4. Parallelize the right way: overlap I/O with threads/async (GIL released during native decode), and use processes for CPU-heavy pure-Python work. Batch GPU/model calls instead of one-at-a-time.
5. Swap in faster building blocks where it matters: a SIMD/turbo JPEG decoder, a vectorized resize, or GPU decode/resize; pick the right interpolation (don't pay for bicubic when bilinear suffices).
6. Cache and avoid rework: memoize decoded/resized results, reuse buffers instead of reallocating per image, and skip redundant color conversions or copies between stages.
7. Re-measure after each change and keep only the ones that move the dominant stage; verify output is still correct (perf must not silently change pixels).

## Rules
- Profile first and per-stage — the bottleneck is rarely where intuition says.
- Decoding fewer pixels (scaled decode/downsample-on-load) usually beats any post-decode optimization.
- Classify I/O-bound vs CPU-bound before choosing threads, processes, or a faster codec.
- Threads overlap native decode (GIL released); use processes only for CPU-bound pure-Python.
- Verify pixels are unchanged after optimizing; faster but wrong is a regression.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…