1. On-Screen Visual Grounding
Instead of reading arbitrary text buffers blind, modern on-device vision models parse UI coordinate trees, iconography, and text elements directly from the GPU framebuffer, determining context without leaking data off-device.