Skip to content
Back to skills

Mai Ui Agents

ASecurity

Scale GUI agents to real-world complexity via extended action space (user interaction, tool calls) and device-cloud collaboration. Online RL supports 500+ parallel environments with asynchronous handling; local agent monitors trajectory alignment and handoffs to cloud when drift detected—achieving 41.7% MobileWorld success with privacy-preserving delegation.

  • 6 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 9, 2026
ai-agentspythonperformance

Works with

  • cli
  • mcp

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add ADu2021/skillXiv --skill mai-ui-agents --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Mai Ui Agents?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Mai Ui Agents
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/adu2021-mai-ui-agents/badge)](https://www.skillsdirectory.com/skills/adu2021-mai-ui-agents)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: mai-ui-agents
title: "MAI-UI: Real-World Centric Foundation GUI Agents"
version: 0.0.2
engine: skillxiv-v0.0.2-claude-opus-4.6
license: MIT
url: https://arxiv.org/abs/2512.22047
keywords: [agents, gui-automation, reinforcement-learning, real-world, multi-modal]
description: "Scale GUI agents to real-world complexity via extended action space (user interaction, tool calls) and device-cloud collaboration. Online RL supports 500+ parallel environments with asynchronous handling; local agent monitors trajectory alignment and handoffs to cloud when drift detected—achieving 41.7% MobileWorld success with privacy-preserving delegation."
---

## Overview

MAI-UI addresses critical limitations in existing GUI agents through pragmatic design choices: extended actions enable richer interactions, device-cloud collaboration preserves privacy, and online RL at scale improves agentic reasoning.

## Core Technique

**Extended Action Space:**
Beyond pure UI operations, agents can request clarification and invoke tools.

```python
class ExtendedActionSpace:
    # Actions: click, type, scroll, + new ones
    user_ask = "ask_user"      # Request clarification
    mcp_call = "mcp_call"      # Use external tools
```

**Device-Cloud Collaboration:**
Local agent monitors alignment; cloud only handles complex cases.

```python
class HybridAgent:
    def should_handoff_to_cloud(self, trajectory, instruction):
        deviation = self.alignment_monitor.evaluate(trajectory)
        if deviation > threshold:
            return True  # Handoff to cloud
        return False  # Continue locally
```

## Key Performance

- 73.5% grounding (ScreenSpot-Pro)
- 76.7% mobile navigation (AndroidWorld)
- 41.7% real-world tasks (MobileWorld)
- 500+ parallel environments

## References

- Extended action space design
- Device-cloud collaboration architecture
- Online RL with asynchronous parallelism

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…