Skip to content
Back to skills

Use Omniparser For Vision Based Gui Parsing

ASecurity

Parse screenshots into structured UI elements so computer-use agents can reason about controls before acting.

  • 36 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 2, 2026
ai-agentspythongogitsecuritydocumentation

Security analysis

A92/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro shows the line behind each finding and how to fix it

Scanned September 2, 2026

npx -y skills add agentskillexchange/skills --skill use-omniparser-for-vision-based-gui-parsing --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Use Omniparser For Vision Based Gui Parsing?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Use Omniparser For Vision Based Gui Parsing
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/agentskillexchange-use-omniparser-for-vision-based-gui-parsing/badge)](https://www.skillsdirectory.com/skills/agentskillexchange-use-omniparser-for-vision-based-gui-parsing)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: "Use OmniParser for vision-based GUI parsing"
slug: "use-omniparser-for-vision-based-gui-parsing"
description: "Parse screenshots into structured UI elements so computer-use agents can reason about controls before acting."
github_stars: 24836
verification: "security_reviewed"
source: "https://github.com/microsoft/OmniParser"
author: "Microsoft"
publisher_type: "open_source"
category: "Browser Automation"
framework: "Multi-Framework"
tool_ecosystem:
  github_repo: "microsoft/OmniParser"
  github_stars: 24836
---

# Use OmniParser for vision-based GUI parsing

Parse screenshots into structured UI elements so computer-use agents can reason about controls before acting.

## Prerequisites

Python 3.12; conda; Hugging Face model weights; optional Gradio demo

## Installation

Use the upstream install or setup path that matches your environment:
- conda create -n "omni" python==3.12
- conda activate omni
- pip install -r requirements.txt

Requirements and caveats from upstream:
- python
- python weights/convert_safetensor_to_pt.py
- python gradio_demo.py

Basic usage or getting-started notes:
- First clone the repo, and then install environment:
- cd OmniParser
- Ensure you have the V2 weights downloaded in weights folder (ensure caption weights folder is called icon_caption_florence). If not download them with:

- Source: https://github.com/microsoft/OmniParser
- Extracted from upstream docs: https://raw.githubusercontent.com/microsoft/OmniParser/HEAD/README.md

## Documentation

- https://microsoft.github.io/OmniParser/

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/use-omniparser-for-vision-based-gui-parsing/)

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…