Courseiva
hardMultiple Choice

Generative AI Leader Practice Question: A developer is building a mobile app that needs…

A developer is building a mobile app that needs to run an AI model on-device for low-latency inference even without internet. Which Gemini model variant is designed for on-device deployment?

⚠ Common exam trap

Watch out — candidates often confuse model capability tiers with deployment targets — candidates assume the 'best' model (Ultra or Pro) is always the answer, missing that Nano is the only variant purpose-built for on-device, offline use.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Gemini Nano

Gemini Nano is Google's smallest Gemini model variant, specifically designed to run locally on devices such as smartphones and laptops for on-device inference without network connectivity. It powers features like on-device summarization in Pixel devices and is optimized for low-latency, offline AI workloads.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Gemini Pro

    Why it's wrong here

    Gemini Pro targets data-centre and cloud API inference, requiring network connectivity and substantial compute, so it cannot run offline on a handset. It suits complex reasoning and multimodal tasks served through the Gemini API where latency and connectivity are acceptable.

  • ✗

    Gemini Ultra

    Why it's wrong here

    Gemini Ultra targets data-centre-scale deployment, requiring substantial compute and memory that mobile hardware cannot supply, so it cannot deliver offline on-device inference. It is tempting because Ultra is Google's highest-capability tier for complex reasoning, and it would be the right pick when maximum accuracy matters and workloads run in the cloud rather than on a handset.

  • ✗

    Gemini Flash

    Why it's wrong here

    Gemini Flash is optimised for low-latency, high-throughput cloud API serving, not embedded on-device execution; it still requires network access to Google's endpoints. It suits cost-efficient, fast responses for chat and summarisation workloads where connectivity exists.

  • ✓

    Gemini Nano

    Why this is correct

    Gemini Nano runs locally on-device, executing inference on the handset's own hardware rather than calling a remote endpoint. This satisfies the stem's offline, low-latency constraint, which cloud-hosted variants such as Pro and Ultra cannot meet.

About these practice questions

Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.