hardMultiple Choice
Generative AI Leader Practice Question: A developer is building a mobile app that needs…
A developer is building a mobile app that needs to run an AI model on-device for low-latency inference even without internet. Which Gemini model variant is designed for on-device deployment?
⚠ Common exam trap
Watch out — candidates often confuse model capability tiers with deployment targets — candidates assume the 'best' model (Ultra or Pro) is always the answer, missing that Nano is the only variant purpose-built for on-device, offline use.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Gemini Nano
Gemini Nano is Google's smallest Gemini model variant, specifically designed to run locally on devices such as smartphones and laptops for on-device inference without network connectivity. It powers features like on-device summarization in Pixel devices and is optimized for low-latency, offline AI workloads.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Gemini Pro
Why it's wrong here
Gemini Pro targets data-centre and cloud API inference, requiring network connectivity and substantial compute, so it cannot run offline on a handset. It suits complex reasoning and multimodal tasks served through the Gemini API where latency and connectivity are acceptable.
- ✗
Gemini Ultra
Why it's wrong here
Gemini Ultra targets data-centre-scale deployment, requiring substantial compute and memory that mobile hardware cannot supply, so it cannot deliver offline on-device inference. It is tempting because Ultra is Google's highest-capability tier for complex reasoning, and it would be the right pick when maximum accuracy matters and workloads run in the cloud rather than on a handset.
- ✗
Gemini Flash
Why it's wrong here
Gemini Flash is optimised for low-latency, high-throughput cloud API serving, not embedded on-device execution; it still requires network access to Google's endpoints. It suits cost-efficient, fast responses for chat and summarisation workloads where connectivity exists.
- ✓
Gemini Nano
Why this is correct
Gemini Nano runs locally on-device, executing inference on the handset's own hardware rather than calling a remote endpoint. This satisfies the stem's offline, low-latency constraint, which cloud-hosted variants such as Pro and Ultra cannot meet.
Go deeper
Related to this question
About these practice questions
Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.