The Food Maestro, an iOS app that cooks around your allergies. Tell it what you can’t eat, and every recipe it writes has to get past that before you see it.
Consumer iOS · Food safety · Generative AI · Walkthrough on request (opens in a new tab)
- Role
- Lead engineer: the backend, the AI and the platform on my own, the iOS app with one other developer
- Stack
- .NET MAUI on iOS, FastAPI on async SQLAlchemy, Postgres, GPT-5 via OpenRouter, Stripe, ECS Fargate with GitHub OIDC, Sentry

Problem
Cooking around an allergy is a research job. Every recipe gets read twice, and a substitution you didn’t think about can put someone in hospital. Most recipe apps suggest anything; the ones that filter do it on tags, which is only as good as whoever typed the tags in. A language model makes it worse, because it will invent an ingredient that was never there.
Approach
You pick a chef, set what you can’t eat, and tell the app what’s in your kitchen. The model writes the recipe, and plain code decides whether it is safe. It checks every draft against your allergy list, including the wording of the steps, and sends it back if anything slipped through. Photos get the same treatment from a second model that only says what it can see.

By the numbers
- 2,363
- tests green at the release QA run, on a build I still recommended holding back
- 4 routes
- an allergic user could still have been served their allergen, each one reproduced against the API
- 30 to 190s
- real generation times, so generation became a job the app polls
- 1 of 113
- generated recipes that found a real photo, until I measured it and rebuilt the matching
- 10 users
- where the database pool ran dry, so the first pilot wave was capped there
- 200k
- tokens per person per day, enforced inside the generation loop
Architecture
- App
- .NET MAUI on iOS: XAML and MVVM, Shell routing, SQLite cache, Keychain-only tokens
- Networking
- Refit over Polly retries with backoff, an offline guard, idempotency keys on the four writes that could act twice
- API
- FastAPI on async SQLAlchemy, Postgres and Alembic, 157 endpoints
- Generation
- GPT-5 for recipes under a strict JSON schema, GPT-4o-mini for receipt scanning, Gemini 2.5 Flash Image to draw photos and 2.5 Flash to check them
- Safety
- Deterministic allergen gate, model-expanded aliases that can only add, allergens hidden from the model’s pantry
- Platform
- ECS Fargate behind an ALB, RDS Postgres, S3, Secrets Manager, GitHub OIDC with no static keys
- Billing and notifications
- Stripe subscriptions behind a signature-verified, idempotent webhook, JWT sessions over bcrypt hashes, Firebase push driven by a rule engine over user state
- Delivery
- pytest, mypy and ruff in CI, one image built per commit, ECS deploys serialised per branch so two merges cannot race, migrations run after the service update
- Release
- Signed store builds from a clean tree, symbols uploaded before the build is kept, consent-gated crash reporting
Decisions
- Rules decide whether you see a recipePlain code checks a draft recipe before anyone sees it. It looks for your allergens and their aliases in the wording of the steps as well as the ingredient list. A model expands “shellfish” into what counts as shellfish, but it can only add to that list.
- Three attempts, then it says soA failed draft comes back with the exact line that broke the rule, three attempts in total. The model never sees an allergen in your pantry either, so it can’t reach for one by accident.
- When the safety check started inventing problemsA second model reviewed recipes too, and it turned out to be the least trustworthy part. One tester’s run came back with all fifteen drafts vetoed across every chef, so there was nothing to show. Another’s drafts were refused for gluten, egg and dairy when the only things declared were shellfish, turkey, onions and bell peppers.
- Keeping the reviewer below the rulesThe deterministic check runs first and has the final say. The model has to name the terms it objected to, and a veto naming something the user never declared gets overruled. A veto that names nothing still wins, because I have no way to disprove it.
- Why generation failed in productionRecipe generation kept failing in production and nowhere else. Asking a model why gets you a plausible theory, so I measured the running system instead. The load balancer was still on its default sixty-second idle timeout while real generations took anywhere from thirty seconds to over three minutes, and the lock left behind meant every retry was refused too.
- Turning generation into a jobI raised the timeout to five minutes, but that only moved the limit. Generation starts a job now, the app polls it, and a dropped connection or a killed app picks the job back up.
- The photos nobody was gettingGenerated images take two calls. One draws the dish, and the other lists only what it can see. If anything on the plate is not in the recipe, the image is thrown away.
- Counting the photo matchesIn production I went further and served a small reviewed library instead of generating. When I counted, the library matched 1 recipe in 113. Its vocabulary came off a stock photo deck while the real recipes were one-pan Latin breakfasts, and skillet, which came up 27 times, was not a format it knew.
- Two deploys, four minutes apartTwo pull requests merged minutes apart, each baking its own task definition, and a feature flag ended up set by whichever Docker build finished last. Nothing logged it. I put a queue around the deploy job so a second one waits instead of cancelling, and made deploys read the flags already running instead of resetting them to defaults.
- Telling the client not to shipBefore the build went to outside testers I reviewed it end to end against the live API. I found four ways an allergic user could be shown, planned and cooked their own allergen without a warning, plus a production server still running a development configuration.
- Recommending a smaller first waveI wrote it up and recommended not shipping to allergy or medical testers, with a staged fix plan and a first wave capped at ten people.
- A build you can trace back to codeA build that exists only as edits on a laptop, or crash reports landing in an account the project does not own, means you can’t tell what is running or watch it fail. So every build is tied to a commit and every crash report lands in the project’s own account.
- Tying each release to a signed tagThe release script refuses to build from a dirty tree, commits the version bump itself, signs a tag on the exact commit it built, and throws away a build whose debug symbols failed to upload. It lands with the next build, so it has not stamped a shipped build yet.
- Working with agents as the only reviewerI ran the API, the AI and the AWS side alone and shared the iOS app with one other developer. That only worked because agents did the mechanical share, including scaffolding, migrations and the tests that carried the suite past two thousand.
- What the agents were not allowed to doThey get an approved plan before they edit anything, they run in parallel only where files don’t overlap, and the allergen logic stays plain code no model may rewrite. I don’t hand off review or the final diff, and that is how the four allergen routes turned up before testers did.


Outcome
The app is on TestFlight by invite: four chefs, 79 recipes, a pantry you fill by photographing a receipt, and meal plans built from what you already own. I owned the API, the AI and the AWS platform outright and the iOS app alongside one other developer. The suite stood at 2,363 at the release QA run, and I still recommended holding the build back.