Mobile-Agent: The Powerful GUI Agent Family
-
Updated
Jul 7, 2026 - Python
Mobile-Agent: The Powerful GUI Agent Family
[EMNLP-2024] Build multimodal language agents for fast prototype and production
AUITestAgent is the first automatic, natural language-driven GUI testing tool for mobile apps, capable of fully automating the entire process of GUI interaction and function verification.
MobileUse: an open-source mobile GUI agent for Android phone automation, AndroidWorld/AndroidLab evaluation, hierarchical reflection, and proactive exploration.
Official code for "MobileForge: Annotation-Free Adaptation for Mobile GUI Agents with Hierarchical Feedback-Guided Policy Optimization"
Knowledge-enhanced multimodal Fire Safety Agent with two-hop reasoning, evidence-tree RAG, open-set guardrails, and four-layer memory.
LightMem-Ego: Your AI Memory for Everyday Life
The Python Harness for Production AI Multi-Agent Systems
Vision-grounding plugin for browser-use agents with SoM, Florence-2, Vision–DOM alignment, adaptive visual context, and objective evaluation.
Claude Code in Docker. Drop-in OpenAI-compatible API, MCP server, Telegram bot, and CLI — five interfaces, one image. Persistent sessions, file ops, always-on skill injection, and a full dev toolchain (Go, Python, Node, K8s, Terraform, databases) or a minimal image with just the basics.
🖼️ Workshop: Build a multimodal AI agent with Haystack & GPT-4o — featuring image understanding, document retrieval, conversational memory, and human-in-the-loop safety controls
[COLM 2024] ExoViP: Step-by-step Verification and Exploration with Exoskeleton Modules for Compositional Visual Reasoning
A curated collection of multimodal agentic coding—vision as feedback for building, verifying, and repairing executable artifacts.
🧮 Multi-agent AI math tutor built with LangGraph — CRAG retrieval, episodic & semantic long-term memory, Tavily MCP web search, Google OAuth, and Neo4j-style memory graph. Powered by LLaMA 3.3 70B on Groq.
A persistent, emotionally reactive 3D Digital Persona powered by Gemini 2.5 Flash Native Audio. Features sub-100ms conversational latency and procedural ARKit emotive realism.
LangGraph-based multimodal paper reading agent for AI/ML papers, with PDF parsing, visual extraction, memory, and Markdown/PDF report generation.
Evidence-grounded video security agent powered by Qwen3-VL, FastAPI, adaptive sampling, policy guards, and auditable sessions
Cosight — A realtime multimodal AI agent runtime with composable roles and capabilities.
DeepSeek V4 Vision is HERE: Opus 4.8 Killer? (DeepSeek-V4-Flash-Vision-Exp Tested) - Automated multimodal AI agent workflow analyzing financial charts.
To associate your repository with the multimodal-agent topic, visit your repo's landing page and select "manage topics."