← All articles

Why I built Ferrabelle, and how it works under the hood

One day my wife Steffany and I got into an argument. You know the kind. Afterward I sat there knowing I was missing something, and not knowing what. Most husbands would apologize and move on. I told her I was going to build an app, just for her, so I could understand why she was angry with me and understand her better, full stop.

I am aware this is not the standard move. But I work with data for a living, and it was the only way I knew to listen harder.

Steffany lifts. Every cycle app she had tried treated her training like it did not exist, so the app learned both: her cycle and her barbell. It learned when her strong days tend to land and when a rough stretch is on its way. I did not plan what came next. One by one she dropped every other app she used to track her cycle. Today Ferrabelle is the only one she uses, and she says it is the most accurate she has ever had. Those are her words, not a measurement, and I will get to the measurements below.

Her favorite feature is not even for her. With her say-so, the app sends me short notes about how she is doing. She loves that I get them. Honestly, so do I.

That is the story, and it is on the creator page in more or less the same words. What I want to do here is the part I do not usually get to talk about: what is under the hood, and why I built it the way I did. If you are looking for a cycle tracking app for women who lift, you should know what you are actually installing.

It is a phone app in the plain sense

Ferrabelle is a React and TypeScript app running inside Capacitor, which wraps it as a native Android app. Your data lives in an SQLite database in the app’s private storage on your phone. There is no account and no server, so there is nothing for me to see and nothing for anyone to ask me for. The app is in English and Spanish, and the translations ship inside it, so switching language downloads nothing.

The prediction engine is its own package in the codebase, separate from the screens, with its own tests. That separation is not glamorous, but it is what lets me do the next thing.

An accuracy test, and a check you can run yourself

Most cycle apps describe their predictions with one unexplained number, or none. I wanted the same standard I would want at work: the metric defined, the dataset named, the comparison shown, the weaknesses admitted.

So I took the engine and replayed it against a public research dataset, the Fehring/Marquette menstrual cycle study, which Marquette University publishes openly. It holds 159 women and 1,665 cycles. For each woman, the engine predicted every period as of several days before it arrived, using only the data that existed at that moment. The method is called leave-future-out walk-forward validation, and the point is that there is no way for the engine to peek at the answer. In the end, 120 women and 3,282 predictions qualified for scoring.

The headline: the median error on the next period was 1 day. The mean error was 1.82 days. 75.6 percent of predictions landed within 2 days, and 56.8 percent within 1 day. On the same data, a plain calendar method, the kind that averages your last few cycles, had a mean error of 2.08 days.

Two things make that number honest. The engine was never trained or tuned on this dataset; it saw each cycle cold, the same way it sees yours. And the whole evaluation is deterministic and lives in the codebase, so it can be rerun from scratch any time the engine changes. The full write-up, including where it is weak, is on the accuracy page.

The weaknesses are real, and I would rather you hear them from me. Long cycles over 35 days predicted at about 2.8 days mean error against 1.79 for the regular range. The public dataset has no wearable temperature data, so these numbers show the engine without what should be its strongest signal. And hormonal birth control changes what a prediction even means, so the app has a separate mode for that rather than pretending otherwise.

Here is the part I am proudest of as a data person. You do not have to take my word for the engine. In the app, Settings, then “Check prediction accuracy,” replays 300 test cases on your own phone and confirms it gives the same answers as the engine I measured. I like being able to say that.

Predictions come from statistics, not AI

I want to be precise about this, because “AI” gets attached to everything now. Ferrabelle’s predictions come from a transparent statistical engine. Before you have logged anything, it starts from published population research covering over 600,000 cycles, conditioned on your age. From your first logged cycle, it personalizes to you, and your own data outweighs the population more and more as it accumulates. Once a week it recalibrates to your history, right on the phone. Nothing is pooled across users, because there is no server to pool it on.

There is an optional AI assistant, and it is a separate thing. It is off by default, it is part of Pro, and it only explains and answers questions. It never sees raw Health Connect values. You pick where it runs, and one of the options is a Gemma 4 model that downloads once to your phone and runs there through Google’s LiteRT-LM library, so your question never leaves the device. I wrote about that choice separately in the on-device AI article.

Partner notes that never share the details

Back to Steffany’s favorite feature, because the design of it is the design of the whole app in miniature.

By default, the notes a partner gets never mention fertile days, weight or symptoms, and they are off unless she turns them on. By default a note says only whether a period started, could start any day now, or is about a week or more away, plus a kind suggestion. She can choose to share more with a particular partner, such as her phase, her mood or symptoms she logged, and she sees a preview and confirms before anything new goes out. Fertile days, ovulation, temperature, weight and lab values are never sent, whatever she picks. The app also tells her plainly that over time, someone reading the countdown could guess roughly when her period comes. I would rather say that out loud than have her find it out.

Partner messages are a Pro feature and need the AI assistant on, since the assistant writes the note in her language from the few facts she allowed.

The idea, as the creator page puts it, is that the app never hands anyone the details. It just helps them show up.

Why it is built this way

I built it for one person first, and I still build every part of it the way I would want it built for her. I mean that practically. It is why there is no account, why the engine runs on the phone, why the accuracy test is something you can run yourself, and why the partner notes say less than they could.

Ferrabelle is in a closed test on Google Play right now. If you want to try it, you can join at ferrabelle.com/testers, and I would be glad to hear what your own logs end up saying.

Sources