AI on your phone: how Ferrabelle's on-device assistant works
Ferrabelle has an optional AI assistant. You can ask it why a prediction looks the way it does, get a weekly summary, or log a day in plain words. I want to explain what “on-device” means in that sentence, because the phrase gets used loosely, and in a cycle app the difference is the whole point.
Off until you turn it on
First, the default. The assistant is off. With it off, and with pressure tracking off (also the default), the app makes no internet connections at all. I measured the release build this way: zero bytes left the phone. The assistant is a Pro feature, and you have to go into Settings and switch it on. Nothing I describe below happens unless you do.
Two ways to run it on the phone
When you turn it on, the app asks where the AI should run. Two of the options keep everything on the phone.
The first is your phone’s built-in AI. Some newer Android phones ship with Gemini Nano, running under an Android system service called AICore. Google’s own description: AICore is “an Android system service that enables on-device execution of GenAI foundation models”. Its documentation says “Input, inference, and output data is processed locally” and that “Functionality remains the same without reliable internet connection.” If your phone has it, Ferrabelle can use it, and your question never leaves the phone.
The second is a model you download once. Ferrabelle offers Gemma 4, Google’s family of open-weight models, whose smaller sizes Google describes as “built for ultra-mobile, edge, and browser deployment”. The recommended one, Gemma 4 E2B, is about a 2.6 GB download and needs about 8 GB of phone memory. A larger one, Gemma 4 E4B, is about 3.7 GB and is only offered on phones with 12 GB of memory or more. You can also bring your own model file. The app downloads the model from Hugging Face (no personal data is sent to get it), checks the file, and from then on runs it with LiteRT-LM, Google’s library for running language models on devices. Google’s own AI Edge Gallery app runs the same models “entirely offline using LiteRT-LM.” After the download, the model needs no internet.
In both cases, the question you type, the facts the app adds to it, and the answer the model writes all stay in the phone’s memory and storage. There is no server of mine, and no one else’s server either.
Cloud options exist, and they are your choice
I am not going to pretend the on-phone models are as capable as the big cloud ones. So the app also lets you use a server you run yourself, your own OpenRouter account, your own API key, or hand a question to the Claude, ChatGPT, Grok or Gemini app on your phone. If you pick any of those, your question goes to that provider under its terms. Before anything is sent, the app shows you exactly what will be sent, and it keeps a private log on the phone of when you sent something and to which provider.
That is the trade: better answers, in exchange for text leaving the phone. I would rather show the trade than hide it.
What the assistant can and cannot see
Whatever option you pick, the assistant works from a short set of facts the app assembles: your cycle status and the things you logged by hand that the question needs. It never sees raw Health Connect values. Your resting heart rate from last night, your sleep sessions, your steps: none of those go into a prompt. Your cycle predictions can, and Health Connect data can influence those, so I say that plainly rather than leave it out.
Predictions never come from AI
This gets its own section because it matters most.
Ferrabelle’s period and phase predictions do not come from a language model. They come from a statistical engine that runs on your phone, gives the same answer every time for the same data, and is measured against a public dataset on my accuracy page. The AI assistant only explains and answers questions. If you turn it off, every prediction is exactly the same. I did not want a product where the central number could change because a model felt differently that morning.
One thing that does send data while AI is on
While the assistant is on, whichever option you pick, Google’s on-phone AI library (ML Kit) can send Google usage and error statistics: device information, app version, an installation identifier and error codes. It never includes your question or your data. With the assistant off, nothing is sent.
You do not have to believe me on any of this. The app has a privacy meter, under Settings, Privacy, What left your phone, that shows the app’s actual network totals as Android counts them, today, over the last 7 days and since install. Android measures it, not me, and the app cannot change the numbers.
Why I built it this way
On-device AI on Android is not a privacy slogan for me. I built this app for one person first, and I did not want her questions about her own body sitting in a log on someone else’s computer. The on-phone path costs something. The models are smaller, and a 2.6 GB download is not nothing. But it means the hard part of the promise, that the app lives entirely on your phone, survives turning the assistant on.
If you want the cloud, you can have it, with your eyes open and the send button in your hand. If you do not, the phone in your pocket is enough.
Ferrabelle is in a closed test on Google Play right now. If you want to try it, you can join at ferrabelle.com/testers.
Sources
- Gemma 4 model overview, Google AI for Developers.
- LiteRT-LM Overview, Google AI Edge.
- Overview of the ML Kit GenAI APIs, ML Kit, Google.
- How We Measure Accuracy, Ferrabelle, 2026.
- Privacy Policy, Ferrabelle, 2026.