r/LocalLLaMA · · 1 min read

My Reading Library: Evaluating LLMs on Android Tasks

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

My Reading Library: Evaluating LLMs on Android Tasks

Can LLM agents actually get through a day in the life of a normal user?

That question got me reading papers on Android agents and mobile benchmarks over the past few months.

A few patterns kept showing up:

  • Most benchmarks run on emulators, making real-device metrics difficult to measure.
  • Important deployment metrics like battery, thermals, and temperature are often missing.
  • Everyday tasks are scattered across benchmarks, languages, and apps, rather than forming a consistent, globally relevant task set.
  • This makes it harder to evaluate whether an agent can actually work reliably on a real phone, for real users.

For now, I’ve put together a library of papers on benchmarking mobile/Android agents for you all to read!

Link: https://www.alphaxiv.org/shared/folder/01a070c6-29a0-77a9-a5b4-b670d5eee169

submitted by /u/East-Muffin-6472
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA