Skip to content
Back to skills

Levenshtein Distance Metric

ASecurity

Evaluates image-to-sequence models using mean Levenshtein edit distance between predicted and ground-truth strings.

  • 61 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 12, 2026
datapython

Security analysis

A100/100

Scanned September 12, 2026

npx -y skills add wenmin-wu/ds-skills --skill levenshtein-distance-metric --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Levenshtein Distance Metric?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Levenshtein Distance Metric
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/wenmin-wu-levenshtein-distance-metric/badge)](https://www.skillsdirectory.com/skills/wenmin-wu-levenshtein-distance-metric)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: cv-levenshtein-distance-metric
description: >
  Evaluates image-to-sequence models using mean Levenshtein edit distance between predicted and ground-truth strings.
---
# Levenshtein Distance Metric

## Overview

For sequence generation tasks where output order matters (molecular formulas, OCR, LaTeX rendering), BLEU and accuracy are too coarse. Levenshtein (edit) distance counts the minimum insertions, deletions, and substitutions to transform the predicted string into the ground truth. Lower is better; 0 means exact match. Works at character or token level.

## Quick Start

```python
import Levenshtein
import numpy as np

def levenshtein_score(y_true, y_pred):
    scores = []
    for true, pred in zip(y_true, y_pred):
        scores.append(Levenshtein.distance(true, pred))
    return np.mean(scores)

# Usage:
preds = ["InChI=1S/C6H12O6", "InChI=1S/C2H6O"]
truth = ["InChI=1S/C6H12O6", "InChI=1S/C2H5OH"]
print(levenshtein_score(truth, preds))  # average edit distance
```

## Workflow

1. Generate predicted sequences via greedy/beam search
2. Decode token IDs back to strings (stop at `<eos>`)
3. Compute Levenshtein distance per sample
4. Report mean distance across the dataset

## Key Decisions

- **Normalized vs raw**: Divide by max(len(true), len(pred)) for 0-1 scale; raw is more interpretable
- **Character vs token level**: Character-level for formulas/OCR; token-level for word sequences
- **Library**: `python-Levenshtein` is C-optimized; `editdistance` is an alternative
- **Complementary metrics**: Report exact-match accuracy alongside mean edit distance

## References

- [InChI / Resnet + LSTM with attention / starter](https://www.kaggle.com/code/yasufuminakama/inchi-resnet-lstm-with-attention-starter)
- [InChI / Resnet + LSTM with attention / inference](https://www.kaggle.com/code/yasufuminakama/inchi-resnet-lstm-with-attention-inference)

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…