Skip to content
Back to skills

Regex Pro

ASecurity

Write, test, and debug regular expressions with pattern recipes, engine quirks, and performance safety.

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 29, 2026
ai-agentsjavascriptpythonrustgojavaexpresstestingdebugginggitapi

Works with

  • api

Security analysis

A100/100

Scanned September 29, 2026

npx -y skills add aicodedecode/awesome-muse-skills --skill regex-pro --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Regex Pro?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Regex Pro
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-regex-pro/badge)](https://www.skillsdirectory.com/skills/aicodedecode-regex-pro)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: regex-pro
description: Write, test, and debug regular expressions with pattern recipes, engine quirks, and performance safety.
category: utilities
---

## Overview

Regular expressions are a precision instrument most people wield like a hammer. This skill covers
writing correct, readable, safe regex: the core syntax, pattern recipes for common tasks, testing
methodology, engine differences, and catastrophic backtracking — the performance trap that turns a
bad pattern into a denial of service.

## When to use

- Writing patterns for validation, extraction, or search-and-replace

- Debugging a regex that doesn't match (or matches too much)

- Choosing between regex and simpler string methods

- Reviewing regex for performance and safety

- Learning regex systematically rather than by Stack Overflow

## Core concepts

- - **The essential syntax.** Literals, `.` (any), `\d \w \s` (digit/word/space) and negations `\D
  \W \S`, quantifiers `* + ? {n,m}`, anchors `^ $ \b`, groups `()` with alternation `|`, character
  classes `[]` with ranges and negation `[^]`. This covers 90% of real patterns.
- - **Greedy vs lazy.** `.*` grabs as much as possible (greedy); `.*?` grabs as little as possible
  (lazy). `<.*>` on `<b>hi</b>` matches the whole string; `<.*?>` matches `<b>`. Default greed
  causes most "it matched too much" bugs.
- - **Groups and backreferences.** `(...)` captures; `(?:...)` groups without capturing (prefer for
  performance and clarity); `\1` refers back to group 1. Named groups `(?P<name>...)` (Python) make
  complex patterns readable.
- - **Anchors change everything.** `^\d+$` validates the entire string is digits; `\d+` without
  anchors finds digits anywhere. Validation without anchors isn't validation — it's searching.
- - **Engine differences.** PCRE, JavaScript, Python `re`, Go, Java — lookbehind support, named
  group syntax, and Unicode handling vary. Test in the target engine, not just a generic tester.
- - **Catastrophic backtracking.** Nested quantifiers like `(a+)+$` on non-matching input cause
  exponential blowup — a 30-character string can hang for hours. This is a security issue in
  servers. Recognize the shape: quantified group containing quantifiers, especially with overlapping
  character classes.

## Practical workflow

1. 1. **Define the language precisely.** What exactly should match? Write 5 positive and 5 negative
   examples before writing the pattern. Vague requirements produce vague patterns.
2. 2. **Build incrementally.** Start with the simplest core, test, then add one element at a time.
   `^\d{4}` → `^\d{4}-\d{2}` → `^\d{4}-\d{2}-\d{2}$`. Each step verified before the next.
3. 3. **Test both sides.** Every pattern gets positive tests (must match) and negative tests (must
   not match) — including adversarial inputs: empty strings, unicode, newlines, very long strings.
4. 4. **Prefer clarity.** Named groups, comments (`(?x)` verbose mode in Python), and breaking
   complex patterns into multiple simpler checks. A regex nobody can read is a bug waiting to
   happen.
5. 5. **Consider alternatives.** Email validation via regex is a famous tarpit — use a
   parser/library. HTML parsing with regex is a classic mistake — use a parser. Regex excels at:
   delimited formats, log lines, simple tokens, search/replace.
6. 6. **Performance-check risky patterns.** Test with pathological inputs (long non-matching
   strings). If it hangs, rewrite: possessive quantifiers where supported, atomic groups, or
   restructure to avoid nested quantifiers.

**Common recipes:**
- Email (pragmatic, not RFC-perfect): `^[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}$`

- URL-ish: `https?://[^\s/$.?#].[^\s]*`

- ISO date: `^\d{4}-(0[1-9]|1[0-2])-(0[1-9]|[12]\d|3[01])$`

- Whitespace trim: `^\s+|\s+$` (replace with empty)

- Quoted string (simple): `"([^"\\]|\\.)*"` (handles escapes, avoids catastrophic backtracking)

## Common pitfalls

- - **Unanchored validation.** `/\d+/` "validates" `abc123` as numeric. Anchor both ends for
  validation.
- - **Greedy dot.** `".*"` matching across multiple quoted strings. Use lazy `.*?` or negated
  classes `[^"]*`.
- - **Forgetting to escape.** `.` `(` `[` `+` `*` `?` in literals. In most languages, prefer raw
  strings (`r"\d+\.\d+"`) to avoid backslash-escaping hell.
- - **Catastrophic backtracking.** `(x+x+)+y` style patterns on user input. Audit any regex that
  processes untrusted input — it's a ReDoS vector.
- - **Over-engineering.** A 200-character email regex that's still wrong. Match the requirement's
  actual strictness; use libraries for genuinely complex grammars.
- - **Not testing the engine.** A pattern tested on regex101 with PCRE flavor failing in JavaScript
  (no lookbehind in older JS) or Go (no backreferences). Always verify in the deployment engine.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…