Engineering Applications of Artificial Intelligence journal

AI LLM

Jul 4 , 2024 read

Discover the first version of our scientific publication “Graphical user interface agents optimization for visual instruction grounding using multi-modal artificial intelligence systems” published in arxiv and submitted to the Engineering Applications of Artificial Intelligence journal. This article is already available to the public.

Thanks to the Novelis research team for their know-how and expertise.

Go to arXiv

Abstract

Most instance perception and image understanding solutions focus mainly on natural images. However, applications for synthetic images, and more specifically, images of Graphical User Interfaces (GUI) remain limited. This hinders the development of autonomous computer-vision-powered Artificial Intelligence (AI) agents. In this work, we present Search Instruction Coordinates or SIC, a multi-modal solution for object identification in a GUI. More precisely, given a natural language instruction and a screenshot of a GUI, SIC locates the coordinates of the component on the screen where the instruction would be executed. To this end, we develop two methods. The first method is a three-part architecture that relies on a combination of a Large Language Model (LLM) and an object detection model. The second approach uses a multi-modal foundation model.

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Download document

Recent blogs

All blogs

Graphical user interface agents optimization for visual instruction grounding using multi-modal Artificial Intelligence systems

Abstract

arXivLabs: experimental projects with community collaborators

Other topics that may interest you

Recent blogs

[White Paper] From RPA to Governed Agentic AI

PFE 2026 Tour: A Look Back at an Edition Full of Insights and Connections

Novelis takes part in DuoDay 2025

[White Paper] The Future of Automation: Strategies for Efficiency, Growth, and ROI

Recent blogs

[White Paper] From RPA to Governed Agentic AI

PFE 2026 Tour: A Look Back at an Edition Full of Insights and Connections

Novelis takes part in DuoDay 2025

[White Paper] The Future of Automation: Strategies for Efficiency, Growth, and ROI