Start with the feature, not the model
Most useful AI features in apps fall into a few patterns. Naming the pattern first makes the technical choice much easier.
- Summarise or rewrite text the user already has: notes, messages, documents.
- Extract structure from messy input: receipts, scanned pages, voice notes.
- Generate personalised content: workout or meal plans, study schedules, replies.
- Understand images: describe a photo, read text from a document, classify an item.
- Open-ended assistant or chat: the broadest and most expensive pattern to run and moderate.
On-device AI on iOS
Apple’s Foundation Models framework gives Swift apps direct access to the on-device model behind Apple Intelligence, on iOS 26, iPadOS 26 and macOS 26 and later. It supports text generation, guided generation, where the model fills in a Swift type you define, and tool calling, where the model can call functions in your app. Apple positions the on-device model for tasks such as summarisation, entity extraction, text and image understanding and refinement, and recommends Private Cloud Compute or a server model when you need more reasoning or a larger context.
The main constraint is availability: the framework requires a device that supports Apple Intelligence, with the feature enabled. Apple also states that apps in the App Store Small Business Program with fewer than 2 million total first-time downloads can use its next-generation models on Private Cloud Compute at no cloud API cost, which changes the economics for smaller apps.
- No per-request cost for on-device inference.
- User data stays on the device for on-device requests.
- Always check model availability at runtime and design a fallback.
On-device AI on Android
On Android, ML Kit’s GenAI APIs run Gemini Nano through AICore, a system service. Ready-made APIs cover summarisation, proofreading, rewriting and image description, and a Prompt API accepts custom text or multimodal prompts. Input, inference and output stay on the device, and there is no server cost per call.
The trade-offs are real. Support is limited to a list of recent devices, different Gemini Nano versions can produce different output for the same prompt, inference is only allowed while the app is in the foreground, and AICore applies quotas that can return a busy error. Treat on-device AI on Android as an enhancement for supported phones, with a cloud or non-AI fallback for the rest.
When the cloud is the better choice
Cloud models are larger, more capable and identical on every device, which matters when output quality has to be consistent. They suit long documents, multi-step reasoning, broad world knowledge, high-quality image understanding and anything that needs the same answer on a five-year-old phone and a new one.
Never ship a provider API key inside the app binary. Route requests through your own backend, or use a service designed for mobile clients. Firebase AI Logic, for example, offers Swift, Kotlin, Flutter and Unity SDKs for Gemini, keeps the key on the server side behind a proxy, supports App Check against unauthorised clients and applies per-user rate limits. Several of its SDKs can also try an on-device model first and fall back to the cloud.
- Use on-device for short, frequent, privacy-sensitive tasks.
- Use the cloud for quality-critical, long-context or cross-device features.
- A hybrid design, on-device first with cloud fallback, often gives the best balance.
Estimating cloud costs before you build
Cloud language models are priced per million input and output tokens, and the spread between models is large. On Google’s Gemini API pricing page at the time of writing, Gemini 2.5 Flash-Lite costs $0.10 per million input tokens and $0.40 per million output tokens on the paid tier, while Gemini 3.5 Flash costs $1.50 and $9.00. Prices change often, so recheck them before you commit.
A quick worked example: 10,000 daily active users making five requests a day, each with 1,500 input tokens and 300 output tokens, is 75 million input and 15 million output tokens per day. On the cheaper model that is about $13.50 per day, or roughly $400 per month. On the more capable model the same traffic is about $247 per day, or over $7,000 per month. Model choice is the biggest cost decision you will make.
- Keep prompts short and cache stable instructions where the provider supports it.
- Cap output length and the number of free requests per user.
- Use the smallest model that passes your quality tests, and route only hard requests to larger ones.
- Tie heavy AI usage to a paid plan so cost scales with revenue.
Privacy and store rules
In November 2025 Apple updated App Review Guideline 5.1.2(i). It now says you must clearly disclose where personal data will be shared with third parties, including with third-party AI, and obtain explicit permission before doing so. A line in a privacy policy is a weak way to meet that; a clear in-app explanation and consent step before the first request is safer.
Data protection law applies too: GDPR for users in the EU, and in Turkey the Personal Data Protection Law No. 6698, which regulates transfers of personal data abroad. Also read each provider’s terms. Google’s pricing page notes that content sent on the Gemini API free tier may be used to improve its products, while paid-tier content is not, which is a strong reason not to run production traffic on a free tier.
- Send the minimum data needed; strip names, emails and identifiers where possible.
- Explain in the interface what is sent, to whom and why, and ask before the first request.
- Update App Store privacy labels and the Google Play Data safety form to match.
- Log prompts and outputs only if you need them, and set a retention period.
Shipping AI features that hold up
Model output is probabilistic, so treat it like any other untrusted input. Validate structured output, show users that results are AI-generated where it matters, and give them a way to correct or regenerate. For health, finance or safety-related advice, add explicit guardrails and disclaimers. We have learned this first-hand building our own AI apps, including AI workout and diet plans, a Turkish-language chatbot and document scanning with text recognition.
- Build a small evaluation set of real inputs and expected outputs before launch.
- Measure latency on older devices and mobile networks, not just on a developer phone.
- Monitor cost per active user weekly in the first months after launch.