Skip to content
Back to skills

Image To Text

ASecurity

Extracts text and its structure from an image by the route that fits the picture: an OCR engine for clean printed text, the agent's own vision for photos, tables, forms and handwriting, or both, and then checks the result before handing it over. Use when someone asks to "read the text in this screenshot", "transcribe this photo", "turn this table image into CSV", "pull the fields from this receipt", "what does this label say", or needs amounts, dates and codes copied exactly. For engine optio...

  • 142 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added May 29, 2026
content-marketingpythongobashgitapi

Works with

  • terminal
  • api

Security analysis

A92/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 4 files and shows the line behind each finding

Scanned October 4, 2026

npx -y skills add TerminalSkills/skills --skill image-to-text --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Image To Text?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Image To Text
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/terminalskills-image-to-text/badge)](https://www.skillsdirectory.com/skills/terminalskills-image-to-text)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: image-to-text
description: >-
  Extracts text and its structure from an image by the route that fits the
  picture: an OCR engine for clean printed text, the agent's own vision for
  photos, tables, forms and handwriting, or both, and then checks the result
  before handing it over. Use when someone asks to "read the text in this
  screenshot", "transcribe this photo", "turn this table image into CSV", "pull
  the fields from this receipt", "what does this label say", or needs amounts,
  dates and codes copied exactly. For engine options see the tesseract skill;
  for scanned PDFs see pdf-ocr.
license: Apache-2.0
compatibility: "Python 3.8+ and the Tesseract 5 command line on PATH (run on 5.3.4); the crop helper needs Pillow. The vision route needs an agent whose model accepts images."
metadata:
  author: terminal-skills
  version: "2.0.0"
  category: data-ai
  tags: ["ocr", "text-extraction", "vision-model", "data-extraction", "verification"]
---

# Image to Text

## Overview

There are two ways to get text out of a picture, and they fail differently. An OCR engine recognizes characters: it is local, repeatable and reports a confidence per word, but it does not understand layout and it swaps look-alike characters. A vision model reads the way a person does: it follows tables, forms, angles and handwriting, but it gives no confidence and can quietly tidy, skip or invent text. Neither failure announces itself. In the runs behind this skill Tesseract returned a wrong date digit at confidence 83 and a wrong card digit at confidence 80. So the job has three parts: pick the route from what is in the image, read, then verify with a signal that does not depend on the route.

## Instructions

### 1. Look first, then choose the route

Open the image with the agent's file-reading tool so the model sees the picture. Note the kind of content, the layout, the quality (blur, angle, lighting, text height) and the language. Ask what the text is for, which values must be exact (amounts, dates, codes, names), and whether the image may be sent to a hosted model at all: identity documents, medical and financial papers often may not.

| What the image shows | Route |
|---|---|
| Screenshot or clean scan, text in lines or blocks | OCR, then compare with what is seen |
| Photo of a document: angle, shadow, soft focus | Vision for the whole text, OCR to cross-check every number |
| Table or form where the position of a value matters | Vision for the structure, OCR to cross-check the cell values |
| Handwriting, text over artwork, stylized type | Vision; mark unsure words; an OCR engine for print will not help |
| Chart or diagram | Vision for labels and printed values; values read off an axis are estimates and are labeled so |
| Codes, serials, reference numbers | Both routes, enlarged, character by character |
| Hundreds of similar images | OCR for all; send only low-confidence or failed-check images to vision |
| Image that must stay on the machine | OCR only, or a vision model that runs locally; a person reviews the doubtful words |

### 2. OCR route

Save as `ocr_words.py`:

```python
#!/usr/bin/env python3
"""OCR one image with Tesseract and show how sure the engine was.
ocr_words.py IMAGE [--psm 6] [--lang eng] [--below 80] [--text out.txt]"""
import argparse, csv, io, json, subprocess, sys

p = argparse.ArgumentParser()
p.add_argument('image')
p.add_argument('--psm', default='6', help='page segmentation mode: 6 one block, 4 one column, 11 scattered text, 7 one line')
p.add_argument('--lang', default='eng')
p.add_argument('--below', type=float, default=80, help='list words whose confidence is under this value')
p.add_argument('--text', help='also write the plain text to this file')
args = p.parse_args()

run = subprocess.run(['tesseract', args.image, 'stdout', '--psm', args.psm, '-l', args.lang, 'tsv'],
                     capture_output=True, text=True)
if run.returncode:
    sys.exit(run.stderr.strip())
lines, words = {}, []
for row in csv.DictReader(io.StringIO(run.stdout), delimiter='\t', quoting=csv.QUOTE_NONE):
    if row['level'] != '5' or not row['text'].strip():      # level 5 = word; the other rows are containers
        continue
    lines.setdefault((int(row['block_num']), int(row['par_num']), int(row['line_num'])), []).append(row['text'])
    words.append({'text': row['text'], 'conf': round(float(row['conf'])),
                  'box': [int(row[k]) for k in ('left', 'top', 'width', 'height')]})
confs = [w['conf'] for w in words]
text = '\n'.join(' '.join(v) for _, v in sorted(lines.items()))
if args.text:
    open(args.text, 'w', encoding='utf-8').write(text + '\n')
print(json.dumps({
    'text': text,
    'words': len(words),
    'mean_confidence': round(sum(confs) / len(confs), 1) if confs else None,
    'doubtful': [w for w in words if w['conf'] < args.below]}, indent=2, ensure_ascii=False))
```

Pick `--psm` from the layout seen in step 1: `3` (Tesseract's own default) for a page or screen made of separate blocks, `6` when each line must stay together across a gap (receipts, price lists, table rows), `11` for scattered words, `7` for a single cropped line. On the receipt in Example 1, mode 3 returned 57 words and lost the price column; mode 6 returned 67 and kept each price on its line. Installing the engine, language data and image clean-up are covered by the `tesseract` skill.

`doubtful` lists the words under the confidence limit with their boxes. Treat it as a list of places to look, not as the list of errors: a word above the limit can still be wrong.

### 3. Vision route

Transcribe from the image under these rules:

- Copy, do not improve. Keep spelling, capitals, punctuation, currency signs and line breaks as printed. Do not expand abbreviations or fix typos.
- Follow the visual order: top to bottom, and for a table row by row with an empty cell left empty.
- Write `[?]` for anything unreadable instead of the likely word. Never complete a number from context.
- Read values that must be exact a second time on their own, from an enlarged crop (step 4).
- State what was not transcribed: logos, stamps, signatures, a barcode.

### 4. Verify

Use at least one check that is independent of the route that produced the text.

**Compare the two routes.** Save each transcription to a file (`--text` writes the OCR one) and run `reconcile.py`:

```python
#!/usr/bin/env python3
"""Compare two transcriptions of the same image token by token: reconcile.py ocr.txt vision.txt"""
import difflib, re, sys

def tokens(path):   # words in reading order; Markdown table pipes and rule lines are dropped
    return [t for t in re.findall(r'[^\s|]+', open(path, encoding='utf-8').read()) if t.strip('-:')]

a, b = tokens(sys.argv[1]), tokens(sys.argv[2])
count = 0
for tag, i1, i2, j1, j2 in difflib.SequenceMatcher(None, a, b, autojunk=False).get_opcodes():
    if tag == 'equal':
        continue
    left, right = ' '.join(a[i1:i2]), ' '.join(b[j1:j2])
    kind = 'NUMBER/CODE' if re.search(r'\d', left + right) else 'word'
    print(f'{kind:12}{left!r:42}{right!r}')
    count += 1
print(f'{count} disagreement(s) across {max(len(a), len(b))} tokens')
sys.exit(1 if count else 0)
```

Every line it prints is a place where one reader is wrong. Lines tagged `NUMBER/CODE` are settled before anything else.

**Enlarge and re-read.** Cut the disputed word out with `zoom.py`, using the box from `doubtful` or coordinates estimated from the image, then view the crop and run `tesseract crop.png - --psm 7` on it. This script needs Pillow: `python3 -m venv .venv && .venv/bin/pip install pillow`, then run it as `.venv/bin/python zoom.py`.

```python
#!/usr/bin/env python3
"""Cut one word or field out of an image and enlarge it: zoom.py IMAGE X,Y,W,H OUT.png [SCALE]"""
import sys
from PIL import Image, ImageOps

image, box, out = sys.argv[1:4]
scale = int(sys.argv[4]) if len(sys.argv) > 4 else 4
x, y, w, h = (int(v) for v in box.split(','))
pad = max(8, h // 2)
im = ImageOps.exif_transpose(Image.open(image)).convert('L')
crop = im.crop((max(x - pad, 0), max(y - pad, 0), min(x + w + pad, im.width), min(y + h + pad, im.height)))
crop = ImageOps.autocontrast(crop.resize((crop.width * scale, crop.height * scale), Image.Resampling.LANCZOS))
ImageOps.expand(crop, border=20, fill=255).save(out)
```

**Check the content against itself.** Line items add up to the subtotal; quantity times unit price equals the line total; tax is the stated rate; dates exist; a table has the same number of cells in every row; identifiers with a check digit pass their checksum. A transcription that fails arithmetic is wrong somewhere even when both routes agree.

**Leave the rest open.** A disputed token is settled when the enlarged crop is plainly legible and one more signal agrees with it: the other route, the number of characters visible, or the arithmetic. If an amount, date or code is still disputed after enlarging, or a character is ambiguous in the typeface itself (O and 0, I and l), report it as unverified with both readings. Do not pick one.

### 5. Deliver

Return the text in the shape asked for (plain text, Markdown table, CSV, JSON fields), followed by:

```
Route: vision for layout, Tesseract 5.3.4 --psm 6 as cross-check
Checks: 7 disagreements between routes, 5 settled from enlarged crops; items sum to subtotal; subtotal + VAT = total
Unverified: auth code "08B15Z" (OCR "086152"), reference "RF7Q-08B1-Z5S2" (OCR "RF7Q-0881-25S2")
Not transcribed: none
```

## Examples

### Example 1: Fields from a photographed receipt

Priya Raman sends `receipt-photo.jpg`, a phone photo of a till receipt taken at an angle in poor light, and needs the lines and totals for an expense claim. The image shows one column with prices on the right, so OCR runs in mode 6 and the agent also transcribes it by eye.

```bash
python3 ocr_words.py receipt-photo.jpg --psm 6 --text receipt.ocr.txt
python3 reconcile.py receipt.ocr.txt receipt.vision.txt
```

OCR reports 67 words at a mean confidence of 77.1. `reconcile.py` prints:

```
NUMBER/CODE '75™'                                     '75mm'
NUMBER/CODE 'p1se'                                    'P180'
word        'Subtotel'                                'Subtotal'
NUMBER/CODE '4472'                                    '4471'
NUMBER/CODE '086152 Ihaturns wanin'                   '08B15Z Returns within'
word        'cays wien ts recelPs'                    'days with this receipt'
NUMBER/CODE 'RF7Q-0881-25S2'                          'RF7Q-08B1-Z5S2'
7 disagreement(s) across 67 tokens
```

All eight prices agree between the routes, and 7.90 + 6.45 + 4.35 + 5.80 + 3.15 = 27.65, 20% of that is 5.53, and the total is 33.18, so the amounts are confirmed twice. The card digits were read by OCR as `4472` with confidence 80, above the doubtful limit; `.venv/bin/python zoom.py receipt-photo.jpg 244,528,31,11 card.png` followed by `tesseract card.png - --psm 7` returns `4471`, matching the vision reading. The misread words (`75mm`, `P180`, `Subtotal`, the returns line) are plainly legible in their crops. The authorization code stays `086152` for OCR even when enlarged, and B and 8 look alike in the blurred crop, so both codes are delivered as unverified:

```json
{
  "merchant": "KESTREL HARDWARE",
  "date": "2026-09-18",
  "items": [
    {"description": "2 x Brass hinge 75mm", "amount": 7.90},
    {"description": "1 x Wood glue 500ml", "amount": 6.45},
    {"description": "3 x Sandpaper P180", "amount": 4.35},
    {"description": "1 x Masonry bit 8mm", "amount": 5.80},
    {"description": "1 x Cable ties (100)", "amount": 3.15}
  ],
  "subtotal": 27.65, "vat": 5.53, "total": 33.18, "currency": "GBP",
  "card_ending": "4471",
  "unverified": {"auth_code": ["08B15Z", "086152"], "reference": ["RF7Q-08B1-Z5S2", "RF7Q-0881-25S2"]}
}
```

### Example 2: A table screenshot to CSV

Marek Dvorak has `rates-table.png`, a screenshot of a shipping-rates table with five columns, and wants a CSV. OCR in mode 6 reads every value correctly, but an empty cell leaves no trace in its output:

```
Ireland 6.80 11.40 2-3 days
Rest of world 18.40 7-12 days
```

Nothing in those lines says which weight band is missing. The agent therefore takes the structure from the image and writes the table itself, then uses the OCR text only to confirm the values: `reconcile.py` reports one disagreement across 48 tokens, the header `Upto2kg` against `Up to 2 kg`, and no difference in any of the 22 filled cells.

```csv
Zone,Up to 2 kg,Up to 10 kg,Up to 30 kg,Transit
United Kingdom,4.20,7.95,14.60,1-2 days
Ireland,6.80,11.40,,2-3 days
EU zone 1,9.15,16.30,31.75,3-5 days
EU zone 2,11.60,19.85,38.20,4-6 days
Rest of world,18.40,,,7-12 days
```

The agent adds: "5 rows, 5 columns; Ireland has no 30 kg rate and Rest of world has rates up to 2 kg only, as in the image."

## Guidelines

- Confidence is a hint, not a verdict. On a soft 520 px photo of a shipping label, Tesseract read the date `29/09/2026` as `20/09/2026` with confidence 83; the enlarged crop read correctly.
- A vision reading that looks fluent proves nothing: a wrong reading is fluent too. Model vendors document mistakes on low-quality, rotated and very small images. That is why numbers are re-read from crops and checked by arithmetic.
- Some characters cannot be settled from the picture: O and 0, I, l and 1 in many sans-serif typefaces. Ask for the source record or a sharper image instead of choosing.
- Small text is where the errors were: most misread words in these runs had letters about 10 to 13 px tall. Enlarge the crop three or four times before reading; enlarging a whole page only spreads the blur.
- Models downscale large images, so a full-page photo can lose its small print before the model sees it. Crop sections and read them one at a time.
- Do not use OCR in mode 3 on a receipt or price list and then pair names with amounts by position: the columns come back as separate blocks.
- Handwriting was not part of the test runs for this skill. Expect a vision model to be the only workable route, and return doubtful words as `[?]`.
- For many files, engine tuning, other languages or searchable PDFs use the `tesseract` and `pdf-ocr` skills; for complex documents with handwriting and forms see `chandra-ocr`.

Files in this skill

  • SKILL.md3.1 KB
  • _scores.json1.5 KB
  • scripts/image-to-text.js1.2 KB
  • scripts/image-to-text.sh483 B

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…