Skip to content
Back to skills

Translation Regex Postprocessing

ASecurity

Multi-rule regex pipeline to clean seq2seq translation outputs — deduplicate phrases, fix punctuation, remove artifacts

  • 61 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 12, 2026
developmentpython

Security analysis

A100/100

Scanned September 12, 2026

npx -y skills add wenmin-wu/ds-skills --skill translation-regex-postprocessing --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Translation Regex Postprocessing?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Translation Regex Postprocessing
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/wenmin-wu-translation-regex-postprocessing/badge)](https://www.skillsdirectory.com/skills/wenmin-wu-translation-regex-postprocessing)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: nlp-translation-regex-postprocessing
description: Multi-rule regex pipeline to clean seq2seq translation outputs — deduplicate phrases, fix punctuation, remove artifacts
domain: nlp
---

# Translation Regex Post-Processing

## Overview

Seq2seq models produce artifacts: repeated phrases, prompt leakage, trailing fragments, inconsistent punctuation. A cascaded regex pipeline fixes these systematically. Apply after decoding, before MBR or final submission.

## Quick Start

```python
import re

RULES = [
    # Remove leaked prompt prefix
    (re.compile(r'(?i)^translate \w+ to \w+:\s*'), ''),
    # Deduplicate repeated phrases (2-4 word spans)
    (re.compile(r'\b(\w+(?:\s+\w+){1,3})\s+\1\b'), r'\1'),
    # Collapse repeated single words
    (re.compile(r'\b(\w+)(\s+\1){2,}\b'), r'\1'),
    # Remove trailing short fragments
    (re.compile(r'\s+\w{1,3}$'), ''),
    # Normalize multiple spaces
    (re.compile(r'\s{2,}'), ' '),
    # Fix space before punctuation
    (re.compile(r'\s+([.,;:!?])'), r'\1'),
]

def postprocess_translation(text):
    text = text.strip()
    for pattern, repl in RULES:
        text = pattern.sub(repl, text)
    if text and text[-1] not in '.!?"':
        text += '.'
    return text.strip()
```

## Key Decisions

- **Order matters**: prompt removal first, dedup second, punctuation last
- **Sentence ending**: force period if missing — most metrics penalize incomplete sentences
- **Conservative dedup**: only exact phrase repeats, not paraphrases

## References

- Source: [hybrid-best-akkadian](https://www.kaggle.com/code/meenalsinha/hybrid-best-akkadian)
- Competition: Deep Past Challenge - Translate Akkadian to English

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…