Skip to content
  • 6 min read
  • Building Don't Forget

How Don't Forget's voice input works entirely on device

How I built voice tasks for my iOS focus timer Don't Forget with Apple Speech and Apple's on-device Foundation Models: no audio or transcripts leave the phone, a rule-based parser as a safety net, and a strict 'never guess the time' policy.

Don't Forget is a task timer for people who switch between things all day. The moments when you need it most are also the moments when you have the least patience for typing: you're in the middle of something, a new task pops into your head, and you want it out of your head and into the app in a few seconds.

So I added voice. Tap the mic, say "Tomorrow morning at 10, plan the launch. First review the analytics, then check the screenshots", and the New Task form fills itself in: title, steps, reminder.

I had one hard requirement: nothing you say leaves your device. Not the audio, not the transcript. A to-do list says a lot about someone's life, and I didn't want it on anyone's server, including mine.

Here's how it works.

The pipeline

Voice input goes through four steps, each one small and testable on its own:

  1. Speech → text with Apple's Speech framework, on device.
  2. Text → task with two interpreters running side by side: Apple's on-device language model and a rule-based parser.
  3. Words → dates with a pure function that turns "tomorrow at nine" into an actual date, or asks when it can't.
  4. Review in the normal New Task form, where you check everything and save.

Step 1: on-device speech recognition

The transcriber uses SFSpeechRecognizer with one flag that matters more than any other:

request.requiresOnDeviceRecognition = true

With that set, the audio is never sent to Apple's servers. The trade-off is that not every language has an on-device model on every device. When it's missing, the app doesn't silently fall back to server recognition. It fails with a clear "on-device recognition isn't available" and offers typing instead. Privacy that disappears in edge cases isn't privacy.

Listening also has to know when you're done. Don't Forget stops after 1.8 seconds of silence, gives up after 8 seconds if nothing was said, and never listens for more than a minute. Too short and it cuts you off mid-sentence; too long and it feels like the app stopped listening.

Step 2: a model and a set of rules

Understanding "call John at nine about the new project" is a language problem, and on iOS 26 and macOS 26 Apple ships a small language model that runs entirely on device through the Foundation Models framework.

The important design choice is that the model never writes free text. It fills a typed Swift structure using @Generable:

@Generable
struct GeneratedTask {
  @Guide(description: "The main action: a short imperative task name, without any date or time. At most 10 words.")
  var title: String
 
  @Guide(description: "Ordered supporting steps, only if the user listed a sequence or checklist.", .maximumCount(12))
  var steps: [String]
 
  var notes: String
  var when: GeneratedWhen
  var folder: String
  var isSomeday: Bool
}

The session runs with greedy sampling, so the same sentence gives the same result every time, and the model is prewarmed while you're still speaking, so it's ready the moment you stop.

The instructions are strict about what the model must not do: never invent steps you didn't say, never repeat the title as a step, and only pick a folder that already exists. A voice feature that creates a folder because it misheard you is worse than no voice feature.

Why rules too?

The model isn't always there. It needs Apple Intelligence, a supported device and a supported language. It can also take a while to answer, or simply fail.

So a rule-based parser always runs too. It's a plain Swift parser with a word list per app language (English, Italian, German, Portuguese and Korean) that recognizes things like "tomorrow", "at nine", "first… then… finally…". It's less clever than the model, but it's instant and completely predictable.

The two race each other:

  • The rules finish immediately.
  • The model gets a 1.5-second grace period. If it answers in time, you see its result.
  • If it doesn't, you see the rules' result straight away, and the model keeps working in the background (up to 20 seconds). When it finishes, it quietly replaces the draft, but only if you haven't started editing it. Overwriting something you just typed would be the fastest way to make people stop trusting the feature.

There's one more split of responsibilities: time words always come from the rules, because their matches are literal. The model is great at "what is the task?"; for "when?", I wanted something that can't be creative.

Step 3: never guess the time

This is the rule I'm most strict about. If you say "call John at nine", do you mean 9 in the morning or 9 in the evening? The honest answer is: the app doesn't know.

So neither interpreter converts what you said into a date. They report time as spoken: the hour as a number, AM/PM only if you literally said it, and a part of the day ("morning", "evening") only if you said one. Then a pure function, SpokenDateResolver, does the date math with a few explicit rules:

  • "At nine" with no AM/PM → the form asks you to pick, and the task can't be scheduled until you choose (or skip the time).
  • A day without a time ("tomorrow") → the form asks for a time.
  • "This evening" → it proposes 19:00, and shows that it's a proposal.
  • A time that has already passed today ("at 8" said at 10) → it proposes tomorrow, with a note saying so.

The resolver has no dependencies, which made it easy to cover with unit tests. Date logic is exactly the kind of code where "it worked when I tried it" means nothing.

Step 4: the form is the review

Voice doesn't have its own editor. Once you stop talking, the result is merged into the same New Task form you'd use when typing, with a small banner showing what was heard, a "redo" button and, when something needs attention, "check the details".

Nothing is created until you tap save. This keeps the feature honest: a misheard word costs you one correction, not a wrong task in your list that you discover later.

Keeping it private, including analytics

Don't Forget has opt-out product analytics, and voice was the place where it would have been easiest to leak something. The rule is simple: the transcript is never reported. A dictation sends only which parts were recognized (a title, steps, a time, a folder), whether the model or the rules produced them, and whether something failed. Never what you said.

Testing something you have to say out loud

Testing voice by talking to a simulator gets old fast. The app accepts debug launch arguments that replace the microphone with a sentence, or force a specific failure:

-VoiceTranscript "Tomorrow at nine call John"
-VoiceFailure noSpeech
-VoiceModelFails

Together with unit tests for the rule parser and the date resolver, plus a flow test for the whole pipeline, that makes every path reproducible, including the ones that are hard to trigger by speaking, like the model failing halfway.

What I'd tell myself before starting

  • Make the model fill a structure, not write text. @Generable turns a language model into something you can reason about.
  • Always have a fallback that doesn't need the model. On-device AI isn't available everywhere yet, and it never will be on older devices.
  • Don't let AI make decisions it can't know. AM or PM is the user's call, not the model's.
  • Never overwrite what the user has touched.
  • Privacy has to hold in the edge cases too, including errors and analytics.

Don't Forget is on the App Store, and the case study covers the rest of the app.

If you're planning an iOS app and want it built with this kind of care, this is what I do.

  • #iOS
  • #SwiftUI
  • #Speech
  • #Foundation Models
  • #Apple Intelligence
  • #Privacy

Written by Francesco Leoni

Independent iOS & web developer in Bergamo, Italy. I design, build and publish my own apps, and build them for startups and companies.

Coda06 / 06

Let’s buildsomething..

Tell me about your idea and where you are today. A short email is all it takes to get started.

or write to leoni.francesco98@gmail.com