Skip to content
  • 6 min read
  • Building BrainDump

How I built AI into BrainDump, a notes app for messy thoughts

How I added AI to my iOS notes app BrainDump: asking questions about your own notes, turning them into summaries, to-do lists or emails, and the streaming, error-handling and usage limits behind it.

I built BrainDump because I needed a place to dump my thoughts quickly, before they disappear, without deciding first where they belong. The app opens straight to a blank page. Organizing can wait.

But "organizing can wait" has a cost: after a few weeks you have a pile of half-sentences, meeting scraps and ideas, and finding anything in it gets hard. That's the problem I wanted AI to solve. Not to write for you, but to help you get something back out of what you've already written.

This is how it works, and the decisions behind it.

Two jobs: ask and create

The AI in BrainDump does two things.

Ask. You pick a folder and ask a question about it: "What did I decide about the pricing page?", "Which ideas did I have for the trip?". The answer is built only from your notes, and it links back to the notes it used.

Create. You select one or more notes and turn them into something else: a summary, a to-do list, a meeting report, an email, a tweet, a blog post, a presentation outline, a report, study questions, flashcards, or a translation. There's also a rewrite mode with tones (professional, casual, inspirational, empathetic, technical).

A created result is saved as a new note, linked to the notes it came from, and an answer from Ask can be saved the same way. Your original thoughts are never overwritten. That was a non-negotiable for me: a brain dump is raw material, and AI output is a derivative of it, not a replacement.

Prompts live on the server, not in the app

BrainDump uses Gemini through Firebase AI Logic with server-side prompt templates. The app never contains a prompt. It only sends a template ID and the inputs:

let ai = FirebaseAI.firebaseAI(backend: .googleAI())
let model = ai.templateGenerativeModel()
let response = try await model.generateContent(templateID: template, inputs: inputs)

Each mode maps to its own template (summary-template, todo-template, ask-template, and so on), and the inputs are just the selected notes (id, title, date and text) plus mode parameters like the tone or the target language.

Why bother? Because prompts change much more often than app code. When a summary comes out too long or a to-do list misses items, I fix the template and every user gets the improvement immediately, without waiting for an App Store review or for people to update the app.

Answers that point back to your notes

For Ask, a plain text answer wasn't enough. If the AI says "you decided to drop the free trial", you want to see where you wrote that.

So the ask template returns a small JSON envelope, inside a fenced block, with the answer and the IDs of the notes it used:

{ "answer": "You decided to drop the free trial and…", "notes": ["7F3A…", "C21B…"] }

The app decodes it, looks up those notes and attaches them to the message. When decoding fails, which happens when a model drifts from the format, the user still gets the raw answer instead of an error, and the failure is reported with the template name and the decoding error so I can see when a template needs fixing.

Streaming text out of half-finished JSON

Answers stream in token by token, because waiting several seconds for a wall of text feels broken. But streaming and JSON don't mix well: halfway through, the model has produced something like {"answer": "You decided to dr, which isn't valid JSON yet.

I didn't want to show users curly braces, so a small progressive parser runs on every chunk. It looks for the "answer" field, reads its value up to the first unescaped quote and un-escapes \n, \t and \" as it goes. The UI shows that partial answer, and the full JSON is decoded only once the stream ends.

It's about 50 lines of code and it makes the difference between "this is a chat" and "this is a debug console".

Designing for failure

AI calls fail in more ways than normal network requests, so I spent as much time on the unhappy paths as on the happy one.

  • Timeouts. Every call races against a 45-second timer in a task group. A stalled request becomes a recoverable error instead of a spinner that never stops.
  • Cancel. Users can cancel a generation. Cancellation is treated as a normal outcome, not an error.
  • Honest error messages. Errors are split into two kinds: messages written by me (like "you've used your free requests") that are safe to show as-is, and raw errors from the network or the SDK, which are replaced with a generic message. Nobody should read a stack trace in a notes app.
  • Empty answers. When the model returns nothing, usually a safety filter or a broken template, it's reported separately from network errors, because the fix is completely different.
  • One stream at a time. The send button stays disabled for the whole stream, not just until the first chunk. Otherwise a second question sent mid-answer interleaves its text into the previous reply.
  • Main thread delivery. Every result, partial, final or error, is delivered on the main thread, because the SwiftUI views update their state directly from these callbacks.

Paying for it without ruining the free app

AI costs money per token, and BrainDump is free to download. The free tier has to stay genuinely useful, and AI has to stay sustainable.

The model I landed on:

  • Free users get 3 AI requests a month. It's a monthly allowance, not a lifetime one: someone who tries the AI in their first week should be able to try it again next month, instead of the feature silently disappearing forever.
  • Subscribers get a generous monthly token quota, which resets every month.
  • Top-ups ("ink drops") can be bought once and don't expire, so someone who only needs AI occasionally doesn't have to subscribe.
  • When you run out, the error says so and offers the upgrade path right there, instead of a dead end.

Usage is counted in tokens, using the count Gemini reports. One detail that bit me: when streaming, Gemini reports the total token count cumulatively on each chunk. My first version summed the chunks and counted the same tokens many times over. The fix was to keep the latest value instead.

Usage is also only counted when the answer actually decodes. If the model breaks the response format, the user doesn't pay for it.

What I'd tell myself before starting

  • Keep the AI in its own module. All of this lives in a small Swift package, AIEngine, that knows nothing about Core Data or SwiftUI. The app passes in notes through a tiny protocol and gets results back. It made the AI code easy to change without touching the rest of the app.
  • Put prompts where you can change them. Server-side templates were the single best decision.
  • Never overwrite the user's words. Generate new notes and link them to the source.
  • Spend real time on errors. With AI, the unhappy path is a normal path.

If you're curious, BrainDump is on the App Store, and you can read more about how the rest of the app is built in the BrainDump case study.

If you're planning an iOS app with AI features and want help building it, this is what I do.

  • #iOS
  • #SwiftUI
  • #AI
  • #Gemini
  • #Firebase AI Logic
  • #Product

Written by Francesco Leoni

Independent iOS & web developer in Bergamo, Italy. I design, build and publish my own apps, and build them for startups and companies.

Coda06 / 06

Let’s buildsomething..

Tell me about your idea and where you are today. A short email is all it takes to get started.

or write to leoni.francesco98@gmail.com