Skip to content
Back to skills

Generate Image

ASecurity

Generate an original image from a text description, locally (Bonsai-Image 4B, ternary-quantized, on a long-lived studio server). The generated image is put on the user's canvas and attached to your own visual input, so the user sees it and you can SEE what came back, not only what you asked for. Use when the user wants a picture, illustration, avatar, or face created from a description that does not already exist on the web. For existing photos of real things, prefer image search instead; for...

  • 10 stars
  • 0 votes
  • 0 copies
  • 3 views
  • Added June 6, 2026
developmentnextjsfastapiapifrontendbackend

Works with

  • cli
  • api

Security analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned October 7, 2026

npx -y skills add bdambrosio/Cognitive_workbench --skill generate-image --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Generate Image?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Generate Image
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/bdambrosio-generate-image/badge)](https://www.skillsdirectory.com/skills/bdambrosio-generate-image)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: generate-image
description: Generate an original image from a text description, locally (Bonsai-Image 4B, ternary-quantized, on a long-lived studio server). The generated image is put on the user's canvas and attached to your own visual input, so the user sees it and you can SEE what came back, not only what you asked for. Use when the user wants a picture, illustration, avatar, or face created from a description that does not already exist on the web. For existing photos of real things, prefer image search instead; for simple diagrams or line drawings, prefer authoring inline SVG.
args:
  prompt: required string — what to depict, e.g. "a friendly cartoon robot face, soft smile, flat vector style". State qualities you want directly; the model follows prompts literally.
  steps: optional int (default 4) — denoising steps; Bonsai is distilled for 4, more buys little
  size: optional int (default 512) — square side length in pixels. 512 is the fast preset; 1024 is the quality preset (slower).
  seed: optional int — fix for reproducible output; omit for variety
  show: optional bool (default true) — put the image on the user's canvas. Pass false for an image the user should not see yet.
  attach: optional bool (default true) — attach the image to your visual input. Pass false when generating several images in one turn and you do not need to look at this one.
---

# generate-image

Produces an original image from a prompt using Bonsai-Image 4B (ternary
1.58-bit) running locally on the GPU via the prism-image-studio backend. The
backend is a long-lived server; the tool is a thin HTTP client that POSTs the
prompt and saves the returned PNG.

The tool talks to the FastAPI backend directly: `POST {base}/generate`, with a
`{base}/healthz` readiness probe. Base URL defaults to `http://localhost:4001`
(override with `GENERATE_IMAGE_URL`). If the backend isn't already running, the
tool launches it via the Bonsai-Image-Demo `serve.sh` and waits for `/healthz`
(a cold boot prewarms the model and can take a couple of minutes). To start it
by hand:

```
BACKEND_PORT=4001 FRONTEND_PORT=3100 ./scripts/serve.sh
```

`4001` is the **backend** port here. The tool bypasses the Next.js studio
frontend (`/api/generate`), so its port is irrelevant and its prompt-moderation
gate is skipped.

Best for invented/illustrative imagery — characters, faces, scenes, styled
graphics. It is a *generator*, not a search: it cannot reproduce a specific
real photograph or a named existing image. There is no separate negative
prompt; describe exactly what you want (including what to leave out) in the
`prompt`.

## Seeing the result

The image is attached to your visual input for the rest of this turn (no
separate "describe" step), labelled "generated image". Three limits:

- Only the most recent image from any tool is in view. A second generation, or
  a camera capture, replaces this one.
- It is not carried to your next turn; only what you wrote about it is.
- It adds its tokens to every remaining iteration of the turn, and a 1024
  image costs more of them than a 512 one.

## Showing the result

The image is put on the user's canvas by this tool; you do not call `display`
for it. The observation says "[shown on the user's canvas]" when that
happened. At the end of such a turn the canvas holds the picture, and your
reply is not mirrored over it.

Each image shown replaces the one before it on the canvas, so of several
generated in one turn the user sees the last. To show several together, or an
image with text, compose a page yourself: the observation gives an `<img>` tag
with a `/local` URL to pass to the `display` tool (`format=html`), and that
replaces what this tool showed.

## Examples

```json
{"thought": "the user wants a friendly assistant face drawn", "tool": "generate-image", "prompt": "a friendly cartoon robot assistant face, large round eyes, gentle smile, flat vector illustration, pastel background, centered"}
```

```json
{"thought": "generate a calm scene with no people or text", "tool": "generate-image", "prompt": "a quiet misty lake at dawn, soft pastel sky, minimalist, no people, no text", "seed": 7}
```

Files in this skill

  • Skill.md2.6 KB
  • tool.py9.2 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…