AI

On-device AI Agents With Google ADK for Kotlin

Build on-device AI agents with Google ADK for Kotlin: compare LiteRT-LM, Gemini Nano, and Firebase, run the latency math, and ship safely.
On-device AI Agents With Google ADK for Kotlin

Introduction

On-device AI agents with Google ADK for Kotlin let an Android app reason, call tools, and remember context without sending every request to a data center. Google made that practical in two steps, starting with the 0.1.0 release of ADK for Kotlin and ADK for Android in May 2026. That announcement noted that Gemini Nano is available on more than 140 million devices, which gives local agents an enormous potential audience. Version 1.0 followed in September 2026, and the general availability announcement for ADK for Kotlin 1.0 described full feature parity with the Python and Java editions. Kotlin developers now get hierarchical agents, sessions, memory, and compile-time tool calling in the language they already use for Android. This guide walks through architecture, model backends, working code, hardware limits, risks, ethics, and real deployments so you can decide where a local agent belongs in your product. By the end you will have a clear build path from an empty Gradle project to an agent that answers questions using a model stored on the phone.

Quick Answers on On-Device Agents With Google ADK for Kotlin

What are on-device AI agents with Google ADK for Kotlin?

They are tool-using agents built with Google’s open-source Agent Development Kit in Kotlin, where the language model runs locally on the phone or laptop through LiteRT-LM or Gemini Nano.

Which model backends can a Kotlin ADK agent use on Android?

ADK for Kotlin supports LiteRT-LM for local open models with tool calling, ML Kit for Gemini Nano in beta, and Firebase AI Logic for cloud Gemini models.

Does an on-device ADK agent work offline?

Yes. When the agent uses a local backend such as LiteRT-LM or Gemini Nano, inference needs no network, although cloud tools and remote agents still require connectivity.

Key Takeaways

  • ADK for Kotlin 1.0 brings agent orchestration, sessions, memory, and compile-time tool schemas to Android and the JVM with one Kotlin Multiplatform core.
  • Backend choice drives everything: LiteRT-LM supports tool calling today, while the ML Kit Gemini Nano backend is beta and lacks tool calling.
  • Phones can run small open models such as Gemma 4 E2B at roughly 52 decode tokens per second on a flagship GPU, so agent loops must stay short.
  • Confirmation flows, narrow tools, and honest evaluation matter more on a personal device than raw model size.

Table of contents

What Is an On-Device Agent in Google ADK for Kotlin?

On-device AI agents with Google ADK for Kotlin are tool-using agents whose language model runs locally on a phone or laptop, orchestrated by Google’s open-source Agent Development Kit written in Kotlin, so reasoning, tool calls, and memory stay on the user’s hardware.

An Interactive From AIplusInfo

How Long Will Your On-Device Agent Make Users Wait?

Choose a device, then set the steps, prompt size, and reply length of one agent task to estimate latency from published Gemma-4-E2B speeds.

Samsung S26 Ultra GPU

Flagship phoneSingle-board CPU

3

1 call8 calls

800

2004,000

120

20400

2.5 s

Time for one model call

7.6 s

Time for the whole task

52 tok/s

Decode speed on this device

Prefill 8 percent, decode 92 percent of each call.

Feels interactive for a foreground feature.

Speeds from the published LiteRT-LM benchmarks for the 2.58 GB Gemma-4-E2B model in the LiteRT-LM overview. Estimates exclude engine startup, tool execution, and thermal throttling.

Why Agents Are Moving From the Cloud to the Pocket

For most of the generative AI boom, an agent meant a loop running on a server that called a hosted model over the network. That design is simple to build, but it ties every interaction to connectivity, per-token billing, and a copy of the user's data leaving the device. Phones now ship with neural accelerators, and Google's own published numbers show small models reaching usable speeds on them. The practical question has shifted from whether a phone can run an agent to which parts of an agent should stay local. Teams that treat locality as a design variable, not a slogan, tend to end up with faster and cheaper products. This article uses that framing throughout, because it keeps every later decision about backends, tools, and memory grounded in a clear trade-off.

Three distinct forces explain the shift toward local execution in mobile software. Latency comes first, because a local model removes the network round trip and keeps the first token close to the touch that triggered it. Privacy comes second, since photos, messages, and locations never need to reach a server when a local model handles them. Cost comes third, because inference on hardware the user already owns costs the developer nothing per request. The Kakao Mobility case study makes that cost argument explicit later in this article. Together these forces explain why Google invested in an Android-first agent kit rather than leaving mobile to cloud-only frameworks.

Locality is not free, and an honest design weighs what it gives up. Small models reason less reliably than frontier cloud models, so complex planning still benefits from a larger brain. Devices also vary widely in memory, thermal limits, and accelerator support, which complicates testing and support. Download size and first-run setup add friction that a cloud call never has. The best architectures therefore split the work, keeping sensitive and latency-critical steps on the device while escalating hard reasoning to the cloud when the user allows it. That split is exactly what the hybrid patterns later in this guide implement.

How Google ADK for Kotlin Came Together

Looking back at the timeline helps explain why the kit is shaped the way it is. Google announced ADK for Kotlin and a separate Android artifact at version 0.1.0 in May 2026, positioning it as a way to build agents on Android and beyond. The 0.1.0 announcement described LlmAgent, workflow agents, custom agents, and multi-agent systems. It also listed session state, a memory service, MCP tools, agent-to-agent communication, plugins, and OpenTelemetry observability. Early guidance already warned developers to pick the Android core artifact on phones and never to ship API keys inside an app. Those two warnings still anchor the safest project setup today.

Version 1.0 arrived in September 2026 and tightened the story considerably. InfoQ's report on the ADK 1.0 release highlighted feature parity with Python and Java, first-class Java interoperability, context compaction, and session pause and restore. The 1.0 line also made on-device inference a first-class citizen through LiteRT-LM integration and an ML Kit adapter that was still marked beta. The same release added Room and AppSearch integrations so chat history and long-term memory can live in app-private storage. It also introduced skills defined in SKILL.md files with progressive disclosure, which keeps prompts small until a skill is actually needed. For a phone, where every token of context costs battery, that detail is more important than it first appears.

Inside the ADK Kotlin Architecture and Module Map

Building on that history, the repository layout shows how responsibilities are divided across the project. The google/adk-kotlin repository publishes every module under the com.google.adk group, with a core module that holds agents, models, tools, sessions, and memory. A processor module runs Kotlin Symbol Processing to turn annotated functions into tool definitions. A webserver module hosts a development UI and API for local testing. An a2a module handles remote agent communication, and an integrations module carries plugins for external services. Keeping these concerns in separate artifacts lets a mobile app depend on only what it needs.

Three modules matter most for on-device work, and each maps to a different inference path. The litertlm module runs local models through LiteRT-LM and, at the time of the README snapshot, required JDK 21 or newer for its JVM build. The mlkit-android module wraps Gemini Nano through ML Kit and is Android-only. The firebase-android module connects to Firebase AI Logic for cloud Gemini models, which is the recommended route when an app needs a cloud model without embedding a raw API key. You can include several of these modules at once and choose between them at runtime. That flexibility is what makes the hybrid designs described later possible.

Dependencies at version 1.0.0 are short and easy to audit. A typical Android project adds the core Android artifact and the KSP processor, then adds one or two backend modules. Use the Android core artifact on phones, and never add both the Android and JVM core artifacts to the same project. That rule comes straight from Google's Android documentation and prevents duplicate class errors that are tedious to diagnose. The Android guide also lists a minimum SDK of 24, a compile SDK of 34 or higher, and a Java 17 toolchain. Pinning those values early avoids a surprising number of Gradle sync failures on fresh projects.

The core is agnostic about backends, session stores, and memory systems, which is what lets the same agent definition run on a server and on a phone. InfoQ quotes an Android engineer who said compile-time tool schemas keep startup fast on mobile targets. That comment reflects a real architectural choice, because reflection-free tooling avoids both cold-start cost and shrinker surprises in release builds. Developers who have fought R8 rules for reflective libraries will appreciate the difference. The web server module also gives teams a browser-based dev interface, so prompts and tools can be tested on a laptop before any device is involved. Treat that interface as your fast feedback loop and reserve device testing for hardware-specific behavior.

Choosing a Model Backend for On-Device AI Agents With Google ADK for Kotlin

Stepping back from module names, the backend decision shapes capabilities more than any other choice. LiteRT-LM is the local engine that runs open models such as Gemma in a dedicated file format, and it supports tool calling with constrained decoding. The LiteRT-LM Android guide shows an Engine configured with a model path and a CPU, GPU, or NPU backend. It also warns that engine initialization can take about ten seconds. That delay means you should create the engine on a background coroutine and keep it alive across turns. Reloading the model for every request would erase the latency advantage of running locally.

The ML Kit path exposes Gemini Nano, the model that Android manages and updates for you. Because the system owns the model, your app ships no weights, and the Prompt API described in the ML Kit GenAI Prompt API announcement accepts both text and images. Today the ADK ML Kit adapter does not support tool calling, so it suits single-turn reasoning or sub-agents that only read and write text. Gemini Nano also performs best on newer devices such as the Pixel 10 series, so you should plan a fallback for older phones. The adapter carries a beta label in the 1.0 release, which is a signal to weigh before committing a production feature to it. Many teams will use it for summarization, classification, and extraction while another backend handles tool use.

Firebase AI Logic is the cloud leg of the triangle, and it offers full tool calling with Gemini models. Google's Android guidance recommends it, or your own backend, over embedding an API key in a client app. Many production designs combine paths, with a cloud model orchestrating and a local model handling privacy-sensitive steps. ADK's model abstraction makes that mixture a configuration decision instead of a rewrite. The same agent definition can point at a local model during development and a cloud model in a release build. That portability also protects you if a backend changes terms, pricing, or availability.

Tools, KSP, and Zero-Reflection Function Calling

Turning to the part that makes agents useful, tools are ordinary Kotlin functions with two annotations. You mark a function with the Tool annotation and describe each argument with the Param annotation. The KSP processor reads those annotations at build time and generates an extension function called generatedTools, which returns the tool list you pass to the agent. Because the schemas exist before the app runs, there is no reflection at runtime. Suspend functions work as tools too, so a tool can query a Room database or call a network API without blocking the main thread. The result feels like writing normal Kotlin, which lowers the learning curve for Android teams.

That design has several practical consequences when the agent runs on a phone. Startup stays fast because the app does not scan classes to discover tools, and release builds shrink cleanly because nothing depends on reflective lookups. Typed parameters mean the schema the model sees always matches the function signature you compile. Keep tool descriptions short and literal, because a small on-device model follows a crisp one-line description far better than a paragraph of nuance. Our explainer on function calling in LLMs covers how models choose and format calls. Treat each description as a tiny prompt that you can test and tune like any other.

Small models are also sensitive to tool count, and that sensitivity deserves a deliberate design response. Offering a local model twenty tools invites wrong picks and malformed arguments, while three or four focused tools raises reliability. LiteRT-LM adds constrained decoding, which forces output to match a JSON schema and removes a whole class of parser errors. Google also published a fine-tuning recipe in its FunctionGemma Mobile Actions guide that adapts a 270 million parameter model to phone actions such as contacts, email, and calendar events. That recipe shows a path for teams whose tool set is narrow and stable. Fine-tuning a tiny model on your own tools can beat prompting a larger one.

Sessions, Memory, and State on a Phone

Moving on from tools to context, an agent that forgets everything between turns is only a chatbot with extras. ADK separates short-term state, held in a session, from long-term recall, held in a memory service. On Android the 1.0 release added a Room-backed session service and an AppSearch-backed memory service, so history and facts persist in the app's private storage. Our own overview of agent memory architecture explains why that split matters. Keeping both stores local also means a user can wipe their agent's memory by clearing app data. That simple property is a meaningful privacy feature for ordinary users.

Context windows deserve special care when the agent runs on a battery-powered phone. Gemma 4 edge models advertise a 128K context window, but every extra token costs memory and prefill time on a battery-powered device. ADK's context compaction, which summarizes older turns, is therefore not an optimization but a requirement for long-running local agents. Sessions can also be paused, serialized, and restored, which lets an app survive process death without losing a half-finished task. Android can kill a background process at any moment, so this resilience is not optional. Design every long task so that it can resume from a saved session instead of restarting.

Orchestrating Multi-Agent Hierarchies Under Tight Resource Limits

Beyond single agents, ADK lets you compose a hierarchy in which a root agent delegates to specialists through a subAgents parameter. The 0.1.0 announcement also documented the disallowTransferToPeers and disallowTransferToParent flags, which stop a child from handing work sideways or back up the tree. On a server those flags are a convenience, but on a phone they are a safety valve against delegation loops that burn battery. Each hop in a hierarchy costs a model call, and a local model call costs seconds, not milliseconds. A three-level tree can therefore turn a simple request into a noticeably slow experience. Our guide to hierarchical coordination in multi-agent tasks describes the same trade-off in a broader setting.

Memory is the second major constraint on any agent hierarchy that runs locally. A 2.58 GB model file loaded twice would double the footprint, so design your agents to share one local model instance wherever the framework allows it. A flat design with one capable local agent and two or three narrow tools usually beats a deep tree of tiny specialists on a phone. Reserve extra agents for genuine boundaries, such as a privacy-sensitive sub-agent that must never see cloud-bound data. When in doubt, measure the end-to-end latency of the flat design first and add structure only when the numbers justify it.

The hybrid pattern in Google's Android documentation is the most practical multi-agent shape for mobile. A cloud Gemini model orchestrates the conversation, while an on-device sub-agent handles anything that touches personal data. That arrangement keeps the strong reasoning in the cloud and the sensitive content on the device, and it degrades gracefully offline. Agents can also reach shared services through the Model Context Protocol, which ADK supports through MCP tools. Treat every remote connection as optional so that the core experience still works without a network.

Hardware Reality: Memory, Tokens per Second, and Battery

Shifting from architecture to physics, the numbers published for LiteRT-LM show what a phone can really deliver. The LiteRT-LM overview lists the Gemma-4-E2B model at 2.58 GB, with 3,808 prefill tokens per second and 52 decode tokens per second on a Samsung S26 Ultra GPU. A MacBook with an M4 Max reaches 7,835 prefill and 160 decode tokens per second, and a Raspberry Pi 5 on CPU manages 133 and 7.6. A Qualcomm Dragonwing IQ8 NPU reaches 3,700 prefill and 31 decode tokens per second. Those figures give you a realistic range for planning, from comfortable on a flagship to painful on a single-board computer.

Turning those numbers into an agent budget takes only simple arithmetic. Imagine one agent step with an 800-token prompt and a 120-token reply. On the S26 Ultra GPU that step needs about 0.2 seconds of prefill and 2.3 seconds of decode, roughly 2.5 seconds in total. On the Raspberry Pi 5 the same step needs about 22 seconds, so a three-step task stretches past a minute. Agent loops multiply latency by the number of tool calls, which is why the interactive estimator earlier in this article lets you vary steps, prompt size, and reply length.

Memory deserves equal attention, because a model that does not fit in memory simply cannot run. The Gemma 4 edge announcement says the E2B model can run in under 1.5 GB of memory on some devices thanks to 2-bit and 4-bit weights. A separate LiteRT-LM post reports a 607 MB physical footprint on Apple mobile CPUs, because image and audio encoders load only when needed. Those savings depend on the device, the runtime, and the quantization you choose, so verify them on your own test phones. Budget for the model file itself as well, since a download of more than two gigabytes affects storage and data plans.

Setup details matter as much as raw speed when you ship to real users. The Android guide notes that GPU use requires native libraries such as libOpenCL declared in the app manifest. It also says the NPU path is still an early preview on Android and Windows. Engine initialization takes around ten seconds, so show a loading state and warm the engine before the user asks for anything. Multi-token prediction, enabled through an experimental flag before initialization, can speed GPU decoding by up to 2.2 times according to Google. Battery drain and heat follow the same curve as compute, so long agent loops on a warm phone will throttle, and your tests should include sustained runs.

Human Confirmation and Guardrails for Agents That Act

Given those constraints on speed, the cheapest safety feature is making the agent ask before it acts. ADK for Kotlin 1.0 supports human-in-the-loop flows through a requireConfirmation setting on tools, and the runtime pauses execution until the user answers. The announcement demonstrated this with a banking assistant whose transfer tool needs explicit approval, and it carried a disclaimer that the sample was not designed to meet compliance requirements. Mark every tool that spends money, sends messages, deletes data, or changes settings as requiring confirmation, and treat read-only tools as the default. That single habit prevents most of the dramatic failures people fear from autonomous mobile agents. Our overview of human-in-the-loop AI accuracy shows how review steps also improve results.

Confirmation is necessary but not sufficient, because a user can approve an action the agent has described misleadingly. Pair it with deterministic checks that run in plain Kotlin outside the model, such as amount limits, allow-lists of contacts, and rate limits per hour. A small local model cannot be trusted to enforce policy on itself, so put policy in code the model cannot rewrite. Our piece on deterministic guardrails for AI agents lays out that principle in detail. Log each confirmed action locally so the user can review what the agent did on their behalf.

Putting On-Device AI Agents With Google ADK for Kotlin to Work

Moving on from safeguards to delivery, the most reliable rollout starts in the cloud and ends on the device. Develop the agent on a laptop using the ADK dev server and a Gemini model, because iteration there is fast and the web interface exposes each tool call. Once prompts and tools behave, swap the model for a local backend and rerun the same scenarios to see what the smaller model gets wrong. This order separates logic bugs from model-capacity problems, which are very different to fix. Teams comparing languages for this work may find our note on Kotlin versus Python useful context.

Next comes routing, which decides which model sees which request. A simple policy sends anything containing personal data to the on-device model. It sends complex multi-step planning to a cloud model only when the user is online and has consented. When the device is offline, it falls back to a degraded local mode. Keep that policy in one Kotlin class with unit tests, because it is the most security-relevant code in the product. Run the agent from a ViewModel coroutine scope, collect the event flow, and update the interface incrementally so users see progress instead of a frozen screen. The engine should live in a long-lived holder that survives configuration changes, since recreating it costs several seconds.

Finally, plan carefully for the wide spread of device capabilities in the field. Query capabilities at startup and choose a backend per device tier. A GPU-capable flagship can run a larger local model, while a mid-range phone might use Gemini Nano for text-only tasks. An older device can fall back to a cloud model with the user's consent. Ship feature flags so you can disable a backend remotely if a firmware update breaks it. Our walkthrough on building custom AI agents for workflow automation offers general agent design advice that applies here as well. Start with one narrow, valuable workflow and expand only after real usage data arrives.

Testing and Evaluating Local Agents

Looking at quality next, an on-device agent needs an evaluation set before it needs any more features. The Kakao Mobility team evaluated Gemini Nano against roughly 200 labeled examples per feature and tracked accuracy, precision, recall, and F1 score while tuning prompts. That discipline is portable: collect a few hundred realistic requests, record the expected tool calls and answers, and rerun them after every prompt, model, or tool change. Evaluate tool selection and argument correctness separately from the quality of the final prose, because small models fail in different ways at each step. Store the results per device tier so regressions on mid-range phones do not hide behind flagship numbers.

Observability closes the loop once the agent is running in production. ADK supports OpenTelemetry, which lets you trace model calls, tool invocations, and latency without inventing a logging scheme. Collect only what privacy policy allows, and prefer aggregate metrics such as step counts and failure rates over raw user text. Our guide on how to measure AI agent performance covers metrics that translate well to local agents. Add a visible feedback control in the interface so real users can flag bad answers and feed your next evaluation round.

Privacy, Security, and the Risks of Local Agents

Turning to security, local execution improves privacy but does not make an agent safe, and the risks deserve plain language. The biggest threat is prompt injection through content the agent reads, such as a message, web page, or notification that contains hidden instructions. If that agent can send email or change settings, the injected text becomes an attack. Researchers have built benchmarks for this problem, including MobileSafetyBench, which evaluates the safety of autonomous agents that control mobile devices. Treat all external text as untrusted data, never as instructions, and give the agent only the minimum tools it needs.

Several other risks are specific to on-device deployments and deserve their own checklist. A downloaded model file is an executable asset, so verify its source and integrity and store it in app-private storage. Embedded API keys can be extracted from an APK, which is why Google's guidance points to Firebase AI Logic or a backend instead. The ML Kit backend is still beta at the 1.0 release. A commentary from TechJack Solutions noted that no public data describes real-world tool-calling reliability for open-weight models on phones. Plan for that uncertainty with conservative tool sets and staged rollouts.

Security frameworks for agents apply to phones just as they do to enterprises. Define what the agent may touch, log what it did, and make revocation easy. Our framework for securing agentic AI walks through identity, permissions, and monitoring in that spirit. Run a threat model before launch and repeat it whenever you add a tool. Keep the written results next to your evaluation set, so security and quality reviews use the same scenarios.

Turning to the human side, a phone is the most personal computer a person owns, so agent behavior carries ethical weight. Consent has to be specific, informed, and easy for the user to revisit later. Users should know which data the agent reads and which actions it can take. They should also know when a request uses a cloud model instead of the local one. A single permission dialog at install time is not enough for a system that may act on messages, photos, and calendars. Disclose clearly, in the interface itself, whenever a request leaves the device, because privacy by locality is only trustworthy if the boundary is honest. Offer a plain switch that forces local-only operation for users who want it.

Autonomy raises a second question about who is accountable when an agent errs. The user, the developer, and the model vendor each play a part, and the design should make responsibility traceable instead of diffuse. Confirmation prompts, action logs, and undo options are ethical features as much as engineering ones. They let people notice mistakes and correct them before harm spreads. Agents that act silently on a user's behalf erode the trust that makes local AI attractive in the first place. Real incidents show the stakes, as in our report on an agent flaw that opened an email attack vector.

Equity matters too, because capability differs sharply from one phone to the next. Gemini Nano performs best on recent flagship hardware, and a 2.58 GB local model is a real burden on a phone with limited storage or data plans. If the best experience reserves itself for expensive devices, the product widens an existing gap. Offer a lightweight mode, document hardware requirements openly, and avoid making essential services depend on premium silicon. Accessibility benefits deserve the same attention, since voice-driven local agents can help users with motor or vision impairments when designed with them.

Cost, Latency, and the Business Case Against Cloud-Only Agents

Looking at the economics, local inference changes who pays for each token. The Kakao Mobility results credit on-device processing with removing server costs for its address-entry feature, and its developer said the approach incurs no additional cost. Savings of that kind scale with usage, so the heaviest features gain the most from moving local. The honest comparison counts engineering effort, testing across devices, and model distribution alongside the saved cloud bill. Our primer on reducing LLM inference costs covers the cloud side of that ledger.

Latency tells a similar story, although it comes with more nuance than cost does. A local model removes the network round trip, but a 2.5-second step on a flagship phone is still slower than a fast cloud model on a good connection. Local wins when connectivity is poor, when responses are short, or when the interaction is a quick classification or extraction. Cloud wins for long reasoning chains that a small model cannot finish reliably. Telecom teams exploring the same balance have written about edge small language models for ingestion workloads, which shows the pattern is not limited to phones.

Common Mistakes and How to Avoid Them

In practice, most failures come from a short list of avoidable errors. Developers embed an API key in an APK or add both the Android and JVM core artifacts. Others initialize the LiteRT-LM engine on the main thread and freeze the interface for ten seconds. Others recreate the engine for every message, which discards the speed advantage and drains the battery. Create the engine once on a background coroutine, reuse it across turns, and release it deliberately when the feature is no longer needed. A review checklist built from these mistakes catches most problems before users do.

Design mistakes are just as common as setup mistakes and often cost more to fix. Teams give a small model too many tools, skip confirmation on sensitive actions, or trust the ML Kit backend for tool calling it does not yet support. Others test only on one flagship phone and discover performance cliffs after launch. Unbounded agent loops are another trap, so cap steps per task and fail gracefully. Experiments with a local AI coding stack show how quickly local setups surface configuration surprises that cloud services hide.

The Future of On-Device Agents on Android and Beyond

Looking ahead, three trends should make local agents steadily more capable. Accelerators are improving, and LiteRT-LM already lists NPU support as an early preview on Android and Windows, with a Qualcomm Dragonwing NPU reaching 31 decode tokens per second. Decoding techniques such as multi-token prediction promise up to 2.2 times faster generation on GPUs. Model families are also shrinking the gap, since Gemma 4 edge models pair a 128K context window with tool calling in a footprint that fits a phone. Expect the line between what must run in the cloud and what can run locally to keep moving toward the device.

The software layer is maturing in parallel with the hardware and the models. Google positions ADK for Kotlin as part of a broader Android effort. That effort includes an agent-to-UI protocol for rendering structured responses in Jetpack Compose, agent-to-agent communication, and skills defined as files. The ML Kit backend should graduate from beta and gain tool calling, which would let Gemini Nano power full agent loops with no model download. Those steps would make the hybrid pattern the default shape of mobile assistants. Our report on Google's offline AI for Android tracks the platform side of that story. Smaller weights matter here too, and our article on post-training quantization for edge AI explains the accuracy trade-offs behind them.

Open questions remain, and honest forecasting about on-device agents includes them. Nobody has published rigorous reliability data for open-weight tool calling on mid-range phones, and standards for agent consent and auditing are still forming. Regulators are likely to focus on actions taken without clear user approval, which favors the confirmation patterns described earlier. Developers who build evaluation sets, narrow tools, and transparent boundaries now will adapt fastest as rules and hardware evolve. The safest bet is a flexible architecture that can move work between device and cloud as the evidence changes.

Chart From AIplusInfo

How Fast Gemma-4-E2B Generates Text on Edge Hardware

Decode speed in tokens per second, higher is faster. Gemma-4-E2B, 2.58 GB model file.

Source: Google LiteRT-LM overview benchmarks for Gemma-4-E2B. Backends differ by device: GPU on the phone and Mac, NPU on the Dragonwing board, CPU on the Raspberry Pi.

How to Build Your First On-Device Agent With ADK for Kotlin

Step 1 - Prepare the Android project

Beyond the theory, start with an Android Studio project that targets a minimum SDK of 24 and a compile SDK of 34 or higher. Google's ADK for Android guide specifies a Java 17 toolchain and the KSP plugin at version 2.1.20-2.0.1 for the 0.1.0 release. Add the Android core artifact and the processor, and make sure you do not add the JVM core artifact alongside it. The 1.0 documentation uses version 1.0.0 for the same artifact names, so update the version strings to the current release. Sync the project and confirm the build succeeds before writing any agent code.

plugins {
    id("com.android.application")
    kotlin("android")
    id("com.google.devtools.ksp") version "2.1.20-2.0.1"
}

android {
    namespace = "com.example.agent"
    compileSdk = 34
    defaultConfig {
        applicationId = "com.example.agent"
        minSdk = 24
        targetSdk = 34
    }
}

dependencies {
    implementation("com.google.adk:google-adk-kotlin-core-android:1.0.0")
    ksp("com.google.adk:google-adk-kotlin-processor:1.0.0")
}

kotlin {
    jvmToolchain(17)
}

Step 2 - Define tools with annotations

Write each capability as a plain Kotlin function inside a service class. Mark the function with the Tool annotation and describe every argument with the Param annotation, keeping each description to one literal sentence of fewer than 15 words. The KSP processor will generate the schema at build time, so the model always sees exactly what the compiler sees. Start with a harmless read-only tool, because a time lookup proves the whole pipeline without any side effects. Aim for three or four tools in total, since small local models degrade quickly as the tool list grows. Small models call short, specific tools more reliably than broad ones, so resist the urge to bundle several actions into one function.

import com.google.adk.kt.annotations.Param
import com.google.adk.kt.annotations.Tool

class TimeService {
    @Tool
    fun getCurrentTime(
        @Param("Name of the city to get the time for") city: String
    ): Map<String, String> {
        return mapOf("city" to city, "time" to "The time is 10:30am.")
    }
}

Step 3 - Choose and configure the model backend

Pick the model backend that best matches the goal of your feature, and choose exactly 1 backend for the first build. For a fully local agent with tool calling, download a Gemma model in the LiteRT-LM format. Then point the model at its path with a CPU or GPU backend, as the ADK LiteRT-LM documentation shows. For text-only reasoning with a system-managed model, wrap an ML Kit GenerativeModel with the GenaiPrompt adapter instead. Create the model once and keep it in a long-lived holder, because initialization is slow. Load the model file from app-private storage, and download it with a progress indicator the first time the feature is used.

// LiteRtLmModel, EngineConfig, and Backend come from the litertlm modules.
val modelPath = File(context.filesDir, "gemma-4-E2B-it.litertlm").absolutePath
val localModel = LiteRtLmModel.create(
    EngineConfig(modelPath = modelPath, backend = Backend.CPU())
)

// Alternative: Gemini Nano through ML Kit (beta, no tool calling yet).
// val onDeviceModel = GenaiPrompt.create(
//     generativeModel = generativeModel,
//     name = "gemini-nano",
// )

Step 4 - Assemble the agent

Combine the model, a concise instruction, and the generated tools into an LlmAgent. The instruction should state the agent's role and tell it exactly when to use each tool, because small models benefit from explicit guidance. Limit the instruction to roughly 3 or 4 sentences, since every extra token adds prefill time on a phone. Call generatedTools on your service instance to obtain the tool list that the processor produced. Give the agent a clear name and description, which become important once you add sub-agents or remote callers. Keep the first version small and observable, then grow it as your evaluation set expands.

import com.google.adk.kt.agents.Instruction
import com.google.adk.kt.agents.LlmAgent

object HelloTimeAgent {
    val rootAgent = LlmAgent(
        name = "hello_time_agent",
        description = "Tells the current time in a specified city.",
        model = localModel,
        instruction = Instruction(
            "You are a helpful assistant that tells the current time in a city. " +
                "Use the 'getCurrentTime' tool for this purpose."
        ),
        tools = TimeService().generatedTools(),
    )
}

Step 5 - Run the agent with a runner

Wrap the agent in an InMemoryRunner with an InMemorySessionService, then call runAsync from a coroutine scope such as a ViewModel. The call returns a Flow of events, and you collect it to update the interface as the model produces text and tool results. Pass a stable user ID and session ID so conversation state carries across turns, and reuse the same 2 identifiers for every message in a conversation. Replace the in-memory session service with the Room-backed service when you want history to survive process death. Always collect the flow off the main thread, because local inference occupies the CPU or GPU heavily.

import com.google.adk.kt.runners.InMemoryRunner
import com.google.adk.kt.sessions.InMemorySessionService
import com.google.adk.kt.types.Content
import com.google.adk.kt.types.Part
import com.google.adk.kt.types.Role

val runner = InMemoryRunner(
    agent = HelloTimeAgent.rootAgent,
    sessionService = InMemorySessionService(),
)

scope.launch {
    runner.runAsync(
        userId = "user-123",
        sessionId = "session-123",
        newMessage = Content(
            role = Role.USER,
            parts = listOf(Part(text = "What time is it in New York?")),
        ),
    ).collect { event ->
        val text = event.content?.parts?.firstOrNull()?.text
        if (!text.isNullOrBlank()) {
            // Update the UI with the agent response.
        }
    }
}

Step 6 - Protect sensitive actions

Add a second tool only after the read-only path works, and mark it as requiring confirmation if it has side effects. The 1.0 release exposes a requireConfirmation setting on the Tool annotation, and the runtime pauses until the user approves. Build a small confirmation sheet in Compose that shows the exact action and arguments in plain language. Pro tip: validate amounts, recipients, and rate limits in ordinary Kotlin code before the tool body runs, so a persuasive prompt cannot talk the agent past a policy. Log each confirmed and declined action locally so the user can review the history.

class ReminderService {
    @Tool(requireConfirmation = true)
    fun createReminder(
        @Param("Short reminder text") text: String,
        @Param("ISO 8601 date and time") whenIso: String
    ): Map<String, String> {
        // Validate inputs here, then write the reminder.
        return mapOf("status" to "created", "text" to text)
    }
}

Step 7 - Test, measure, and harden

Rehearse the agent against a labeled set of realistic requests on at least 3 device tiers before shipping. Record tool-selection accuracy, argument correctness, step counts, and time per step, and compare them with the budget from the estimator in this article. Run a sustained session of ten minutes or more to expose thermal throttling that short tests hide. You can also run the ADK dev server on a laptop to inspect traces while you tune prompts. Finish with a staged rollout behind a remote feature flag so you can pull the backend if field data disagrees with your lab results.

source .env
gradle run -PmainClass=com.example.agent.WebMainKt

Recommended by AIplusInfo

Books to go deeper on Kotlin and local AI

Hand-picked titles that map to the Kotlin, evaluation, and model fundamentals behind the build steps above.

As an Amazon Associate, AIplusInfo earns from qualifying purchases.

Kotlin in Action, Second Edition

Book

Kotlin in Action, Second Edition

Written by JetBrains Kotlin leaders, it covers the coroutines and flows that every ADK for Kotlin agent loop depends on.

Buy on Amazon
AI Engineering: Building Applications with Foundation Models

Book

AI Engineering: Building Applications with Foundation Models

Covers evaluation, prompt engineering, fine-tuning, and agents, the same skills needed to test and harden a local agent.

Buy on Amazon
Build a Large Language Model (From Scratch)

Book

Build a Large Language Model (From Scratch)

Walks through building a GPT-style model on a laptop, which clarifies the token, context, and decoding limits behind on-device inference.

Buy on Amazon

Key Insights

Taken together, these figures describe an ecosystem that is fast enough for short, focused agent loops and still too slow for long chains of reasoning on modest hardware. Installed base and model efficiency keep improving, which widens the set of features that make sense to run locally. Latency scales with the number of model calls, so the number of steps matters more than the size of any single prompt. Measured business results from Kakao Mobility show that the payoff is real when a feature is narrow and well evaluated. The consistent lesson is to design small, confirmable tools, measure on real devices, and keep a cloud path ready for the hard cases.

DimensionLiteRT-LMML Kit (Gemini Nano)Firebase AI Logic
Where inference runsOn the device or JVM hostOn the Android deviceIn Google's cloud
Tool calling in ADKSupported, with constrained decodingNot supported yetFully supported
Model ownershipYou ship a Gemma model fileThe system manages and updates the modelGoogle hosts Gemini models
Offline useYes, after the model is downloadedYes, on supported devicesNo, a connection is required
Maturity at ADK 1.0General availabilityBetaGeneral availability
Platform scopeAndroid and JVMAndroid onlyAndroid
Hardware dependenceCPU, GPU, or NPU previewBest on recent devices such as Pixel 10Any device with a network
Privacy postureData stays on the deviceData stays on the deviceRequests leave the device
Best fitLocal agents that call toolsPrivate text and image reasoningHard reasoning and orchestration

On-Device Agent Examples in Practice

Gemma 4 E2B on Flagship Phones and Edge Boards

Beyond the abstract benchmarks, Google's published LiteRT-LM figures show what a real model does on real hardware. Engineers ran the 2.58 GB Gemma-4-E2B model on a Samsung S26 Ultra GPU and measured 3,808 prefill and 52 decode tokens per second, according to the LiteRT-LM overview. The same model ran on a Raspberry Pi 5 CPU at 133 prefill and 7.6 decode tokens per second, and on a Qualcomm Dragonwing IQ8 NPU at 3,700 and 31. That spread of roughly seven times between a flagship GPU and a single-board CPU defines the budget for any agent loop. A typical agent step with 120 output tokens takes about 2 seconds on the phone and more than 15 seconds on the board, a sevenfold increase in waiting time. The limitation is that these are raw model throughput figures, and they exclude tool execution, engine startup of about ten seconds, and thermal throttling during long sessions.

Looking at agentic behavior directly, Google's Gemma 4 announcement describes Agent Skills in the AI Edge Gallery app. Developers built skills that pull in Wikipedia knowledge, produce summaries and flashcards, or call companion models for speech and image generation. Google's Gemma 4 edge post reports that the runtime processed 4,000 input tokens across two distinct skills in under 3 seconds. The E2B model reportedly needs less than 1.5 GB of memory on some devices, well below its 2.58 GB file size. Because skills load on demand, the model reasoning ran on the device with no per-request inference bill, a reduction to zero in server cost for that step. The limitation is that the post publishes no accuracy or task-completion rates, so teams still have to measure reliability themselves.

LiteRT-LM Across Chrome, Pixel Watch, and Mobile Apps

Rounding out the examples, Google's LiteRT-LM team lists production surfaces that already rely on the engine. The LiteRT-LM engineering post says Google deployed it for on-device inference in Chrome and ChromeOS, on Pixel Watch, and in the AI Edge Gallery apps for Android and iOS. Engineers measured 52 decode tokens per second on Android with OpenCL and 56 tokens per second on iOS with Metal for the Gemma 4 E2B model. A web build on WebGPU reached 76 decode tokens per second on a MacBook Pro with an M4 Max. Enabling multi-token prediction raised GPU throughput by up to 2.2 times, an increase of 120 percent over baseline decoding. The limitation is that this speedup sits behind an experimental flag, and the post publishes no agent-level latency or reliability data.

Lessons From Case Studies in On-Device AI

Case Study: Kakao Mobility's Gemini Nano Address Extraction

Turning to documented deployments, Kakao Mobility offers the clearest on-device case with published numbers. The company faced a problem in its delivery service, where drivers had to type recipient names, addresses, and phone numbers from chat messages into order forms by hand. That manual step was slow and error-prone, which created friction for customers and drivers alike. The team built an address extraction feature on Gemini Nano through the ML Kit GenAI Prompt API, using few-shot prompting and about 200 evaluation examples to reduce hallucinations. A cloud fallback to Gemini Flash was added for robustness. According to the Android Developers case study, the feature runs on the device, so it needs no server inference.

The measured impact was substantial and covered several different business metrics. Order completion time fell by 24 percent, and conversion rose by 45 percent for new users and 6 percent for existing users. AI-powered orders increased by more than 200 percent during peak seasons. The team credits on-device processing with removing server costs and avoiding the need to upload sensitive photos or locations. The caveats are equally instructive, because the companion bike-parking feature had not launched yet and still needed image filtering work to handle difficult urban scenes. Both features required careful prompt tuning against roughly 200 labeled examples before reaching production quality. Teams planning similar projects should expect that evaluation effort and keep a cloud fallback ready.

Case Study: Google Meet's NPU-Accelerated Segmentation

Moving on to video, Google Meet faced a demanding challenge when it wanted better background segmentation without draining the battery. Running a sophisticated model on a phone for an extended call generates heat, and heat forces throttling. The team deployed NPU acceleration through LiteRT to run an enhanced segmentation model on the device. According to Google's LiteRT NPU case studies, the result was a model 25 times larger without sacrificing inference speed. The power footprint stayed consistent, which created thermal headroom for calls of 20 to 30 minutes and enabled higher-quality background replacement.

The lesson for agent builders is that accelerator choice can change what is feasible, not just how fast it runs. A larger model at the same speed means better quality within the same battery budget, which is the trade every local agent faces. The limitation of this evidence is that it comes from Google itself and describes a vision model, so it says nothing directly about tool-calling language agents. Independent measurements on language workloads are still scarce, and NPU support for LLMs remains an early preview. Treat the case as proof of the hardware direction and verify the language-model side on your own devices.

Case Study: Argmax Speech Recognition on NPUs

Looking at speech, Argmax needed to deploy frontier speech recognition models in apps for customers such as Heidi Health, where latency, app size, and battery life all mattered. The challenge was that long transcription sessions punish any model that is slow or power hungry. The company built its pipeline on LiteRT with ahead-of-time compilation and delivered models through AI Pack feature delivery on Google Play to keep downloads small. Google's NPU write-up reports an over 2x speedup, a gain of more than 100 percent, when moving from GPU to NPU acceleration. That change delivered industry-leading latency and mitigated battery impact for extended sessions. The limitation is that the best results required specific NPU-equipped devices, so teams must plan fallbacks for the rest of the install base.

Common Questions About On-Device AI Agents With Google ADK for Kotlin

What are on-device AI agents with Google ADK for Kotlin?

They are agents built with Google's open-source Agent Development Kit in Kotlin that run their language model locally on a phone or laptop. The agent can plan, call Kotlin tools, and keep session memory without a server round trip. Local backends include LiteRT-LM for open Gemma models and ML Kit for Gemini Nano. A cloud model can still join the design whenever the app needs stronger reasoning.

Is Google ADK for Kotlin ready for production use?

Google announced general availability of ADK for Kotlin 1.0 in September 2026, with feature parity against the Python and Java editions. The core, processor, LiteRT-LM, and Firebase modules reached 1.0.0, while the ML Kit adapter stayed in beta. That mix means most features are stable, but you should test the Gemini Nano path carefully. Staged rollouts and remote feature flags reduce the risk of any surprises.

Can an ADK for Kotlin agent run completely offline?

Yes, as long as every model call and tool uses local resources. A LiteRT-LM or Gemini Nano backend needs no network once the model is on the device. Tools that call web APIs or remote agents will fail offline, so design them to degrade gracefully. A local-only mode is a good default for privacy-sensitive features.

Which backend should I choose: LiteRT-LM, ML Kit, or Firebase AI Logic?

Choose LiteRT-LM when you need a local agent that calls tools, because it supports tool calling with constrained decoding. Choose ML Kit when a system-managed Gemini Nano model is enough for private text or image reasoning without tools. Choose Firebase AI Logic for demanding reasoning that justifies a network call. Many apps combine two or three of these backends behind one routing policy.

Does the ML Kit Gemini Nano backend support tool calling in ADK?

At the time of the 1.0 documentation, the ADK ML Kit adapter does not support tool calling. It is marked beta and is available only on Android. Use it for summarization, classification, extraction, and other tasks that need text in and text out. Pair it with a tool-capable backend when your agent must take actions.

How do tools work in ADK for Kotlin?

You write ordinary Kotlin functions and mark them with the Tool annotation, then describe each argument with the Param annotation. The KSP processor generates tool definitions at compile time and exposes them through a generatedTools function. Because nothing relies on reflection, startup stays fast and release builds shrink cleanly. Suspend functions are supported, so tools can perform asynchronous work.

What are the minimum requirements for ADK for Android?

Google's Android guide lists a minimum SDK of 24, a compile SDK of 34 or higher, and a Java 17 toolchain. You also need the KSP Gradle plugin and the Android core artifact rather than the JVM artifact. The LiteRT-LM JVM module has documented a requirement of JDK 21 or newer in the repository README. Check the current documentation before pinning versions in a production build.

How fast can a phone run a local model for agent loops?

Google's published figures show the 2.58 GB Gemma-4-E2B model reaching 52 decode tokens per second on a Samsung S26 Ultra GPU. A Raspberry Pi 5 CPU manages only 7.6 decode tokens per second for the same model. A step with 800 prompt tokens and 120 output tokens therefore takes roughly 2.5 seconds on the phone. Multiply that by the number of tool calls to estimate total task time.

How much memory and storage does a local Gemma model need?

The Gemma-4-E2B file is about 2.58 GB, so plan for a multi-gigabyte download and matching storage. Google reports that the model can run in under 1.5 GB of memory on some devices thanks to low-bit weights. A separate LiteRT-LM post cites a 607 MB footprint on Apple mobile CPUs. Always measure on your own target phones, because results vary by device and runtime.

How do I stop an on-device agent from taking harmful actions?

Mark every tool with side effects as requiring confirmation, and show the user exactly what will happen before it runs. Add deterministic checks in plain Kotlin, such as spending limits, contact allow-lists, and rate limits. Treat all text the agent reads as untrusted, because injected instructions can hide in messages and web pages. Keep a local log of confirmed and declined actions for review.

Can I combine on-device and cloud models in one app?

Yes, and Google's Android guidance describes this hybrid pattern directly. A cloud Gemini model can orchestrate the conversation while an on-device sub-agent handles privacy-sensitive steps. Route requests with a small, well-tested policy class that considers data sensitivity, connectivity, and user consent. Keep a clear indicator in the interface whenever a request leaves the device.

How should I store conversation history and memory on Android?

ADK for Kotlin 1.0 added a Room-backed session service for chat history and an AppSearch-backed memory service for indexed long-term recall. Both keep data in app-private storage, which supports privacy and simple deletion. Use context compaction to summarize older turns so prompts stay small. Sessions can also be paused and restored after the process is killed.

How long does LiteRT-LM engine initialization take, and how should I handle it?

Google's LiteRT-LM guide says engine initialization can take around ten seconds. Create the engine on a background coroutine, show a loading state, and keep it alive across turns. A writable cache directory can improve load times on later launches. Never recreate the engine for every message, because that wastes both time and battery.

Is ADK for Kotlin free and open source?

The google/adk-kotlin repository is published under the Apache 2.0 license, so you can use and modify it commercially. The framework itself costs nothing, but your backends may not. Cloud models bill per request, while local models cost engineering time, storage, and battery instead. Review each model's license as well, since open models carry their own terms.