Building a personal app that agents can actually use

Super App started as a health app with a place to write notes. About a week ago, I made it for my phone. It had Fitbit data and sections for food, movement, body measurements, and other health records. Pretty straightforward. Building it with AI agents, I added capabilities one at a time, and each addition changed what the underlying system needed to do.

A completely custom app, just for me, to track what I do and keep my data together. When I have something for my agent to record, I open it and leave a note. There are more capabilities now (Bluetooth scale readings, a personal wiki, agent updates, offline data and an export), but most of my use still starts there. A note from the app on my home screen.

Editing a note keeps the original, so I can clarify something later without losing what I first wrote. I want that freedom when I capture things. A meal, a workout, an observation, an idea for the app (all in the same note if that's how they come out); the agent can organize them afterwards.

Having to choose between 4 forms first would get in the way. As the app grew beyond health, I gave it its own project and added camera access, image attachments and editing. More ways to get something down while it's still in my head.

As more features arrived, I kept asking for compact controls and fewer labels. More room for the content. Some of the notes themselves became requests for improvements, because I was already using the app to capture what needed doing. I asked for pictures and reported that new notes weren't scrolling into view. When I asked an agent to implement those changes, it had the original requests.

The screens got a consistent look too: square edges, the dark Tokyo Night palette and shared imagery. Search and filters open when I need them. I wanted the app to do more while staying easy to open and use.

Body history and interactive charts, with measured weight and impedance kept separate from estimated body-composition values. The distinction matters when I look at a reading. After we connected the scale through Bluetooth, I stepped on it and the reading arrived in the app. We checked that it saved. No manufacturer's app, no extra account, no computer needed for a normal weigh-in.

Food records gained researched calorie and nutrient estimates (with their sources and assumptions attached), alongside Fitbit records. Notes give those numbers context; what I ate, how a workout felt and what was happening that day can help explain them. It already makes a day easier to review. I want to use that context across a longer stretch of time too.

When a sleep score was missing, the Fitbit code had been substituting deep-sleep minutes. Different measures, same field. We fixed the processing, and corrections to notes needed care too. A summary using an older version can show that it needs another look.

An agent could have explained that wrong number convincingly (that's what bothers me about the bug). I want to trace an answer back to its source, particularly once we're looking across months of records. An explanation can sound reasonable while starting from something that was never a sleep score.

Keep missing food logs missing, keep the assumptions attached to estimates, and avoid carrying mistakes forward. Collecting more data is only useful if we can trust what the records mean.

Saved progress gives another agent somewhere to resume when a session stops. We added records of completed changes too, so it can see what actually finished. Before an agent saves a review, it checks whether the notes changed while it was working. Then it reads the saved result back.

I have a phone interface, and the agents have direct tools for the records. Their work stays in the app after the conversation ends; procedures and results can be reused while the information behind them is still valid. When that information changes, the conclusions need checking.

I started treating agent use as part of the app's design. That's what I mean by agent-first. A notes review starts with finding new entries, reading the relevant material and preparing an update. The tools give the agent a way to carry that work through to a saved result.

To capture notes, inspect records and see more of the agents' work in one place, I brought their knowledge and messages into the app. I want to see what an agent says it finished; I also want to check the update is actually there.

The GBrain browser gives me access to the shared knowledge system my agents use, and the Grokbot inbox holds their messages. Connected, with different jobs. A message saying the work is done doesn't establish that the result was saved.

One note, several things it describes. One day, several notes. A finding with sources from different places. I wanted a wiki so I could follow those connections, and the agents could find the relevant pieces too.

It grew to more than 4,000 generated pages covering notes, days, datasets and findings. Views of records already there. I hadn't written thousands of articles.

For browsing, we separated opening saved content from preparing an export, which had been slowing the wiki down. Now the app opens the page first and checks for changes in the background. From a day to its notes and measurements. From a finding back to the information supporting it.

To find correlations I'd otherwise miss, I want a few months of connected history to work with. That's a big part of what I'm building this for.

An unusually difficult workout is one example to investigate. What did the sleep leading up to it look like? Does that combination keep appearing? We'd need wearable records, workout notes and a clear view of which days had both; a note entered late needs to link to the day it describes, while retaining when I actually wrote it. Otherwise we could be comparing the wrong days.

Bring the relevant evidence together through the wiki, then give me an answer I can inspect. I want to understand why the agent thinks a relationship deserves a closer look, with the records behind that judgment.

For any comparison, repeatable calculations and a clear account of the records included or missing. Those are things I want to see. A correlation gives us something to investigate. Working out what caused what needs more evidence.

Something really interesting, hopefully, over the next few months. Even a promising idea that doesn't hold up would be useful. The record needs time to grow before we can see what the comparisons are worth.

I want the exceptions too: the unusual week that explains most of a result, or another part of my routine changing at the same time. Show me those details. They can change what a pattern means.

Export readable pages, structured records and original evidence. That's available now, with GBrain and Grokbot kept outside the personal archive.

Without a connection, I still want access to my data. The phone keeps a copy of Fitbit data published by Halla, my server, and retains the last complete copy if Halla goes away.

To finish the archive, we still need to fill gaps in supporting evidence and do more work on larger archives and recovery. The app reports those gaps. I want to know what I actually have. Downloading a file shouldn't make it look as though every part of the record is there.

Capture in Home, review records in Health, find knowledge and agent updates in Brain. The wiki, offline copy and export in Data. Those are the 4 main areas now.

The original notes are still there, and everyday use still starts with opening the app, leaving a note and having the agent help organize it. Broader unattended analysis is still to build.

I made a custom app to track what I do. It now connects those records and gives agents a way to work with them. Over the next few months, I want to see which patterns are actually worth paying attention to. A thought left on my phone can stay useful long after I've forgotten writing it.