Skip to tool

On-Device Assistant Screen Context & Action Planner

Modern mobile AI assistants no longer just answer trivia—they perceive what's on your screen, cross-reference your personal contacts and calendar graph, and synthesize executable tool calls. Test screen grounding, OCR entity resolution, and structured action synthesis below.

Screen Perception & Grounding Sandbox On-Device Neural Engine

Active App: Mail

Hover/tap elements on screen to highlight resolved semantic tokens.

Grounded Screen Tokens

Multimodal Inference Pipeline

1
Visual Layout & OCR Pass
Extracted text tokens and spatial hierarchy in 12ms.
2
Cross-Reference Personal Graph
Matched user home location and Apple Maps traffic model.
3
Intent & Tool Dispatch
Synthesized Maps departure calculation & Calendar alert.
Synthesized Executable Tool Call (JSON) Maps.calculateTravelTime
Loading simulation...
Grounded 4 screen entities. Ready to execute actions.

1. On-Screen Visual Grounding

Instead of reading arbitrary text buffers blind, modern on-device vision models parse UI coordinate trees, iconography, and text elements directly from the GPU framebuffer, determining context without leaking data off-device.

2. Personal Knowledge Graph

A reference to "this flight" or "Maya" requires semantic linking. The on-device assistant resolves ambiguous pronouns by searching local SQLite indexes (CoreSpotlight) of calendar events, recent chats, and indexed contacts.

3. Direct Tool & App Intents

Rather than hallucinating answers in natural language, the assistant emits strict schema-validated App Intents—enabling background operations like setting flight reminders, dispatching rides, or populating grocery baskets.

Enjoy this tool? Build your own with Super