Skip to content
Abhishek KolgeSenior AI Product Engineer
Index

The Food Maestro, an iOS app that cooks around your allergies. Tell it what you can’t eat, and every recipe it writes has to get past that before you see it.

Consumer iOS · Food safety · Generative AI · Walkthrough on request (opens in a new tab)

Role
Lead engineer: the backend, the AI and the platform on my own, the iOS app with one other developer
Stack
.NET MAUI on iOS, FastAPI on async SQLAlchemy, Postgres, GPT-5 via OpenRouter, Stripe, ECS Fargate with GitHub OIDC, Sentry
Home: today’s suggested lunch with its chef, cook time and calories, and shortcuts to Recipe Now and Meal Plan

Problem

Cooking around an allergy is a research job. Every recipe gets read twice, and a substitution you didn’t think about can put someone in hospital. Most recipe apps suggest anything; the ones that filter do it on tags, which is only as good as whoever typed the tags in. A language model makes it worse, because it will invent an ingredient that was never there.

Approach

You pick a chef, set what you can’t eat, and tell the app what’s in your kitchen. The model writes the recipe, and plain code decides whether it is safe. It checks every draft against your allergy list, including the wording of the steps, and sends it back if anything slipped through. Photos get the same treatment from a second model that only says what it can see.

Meet Our Chefs: four chefs, each with a cooking style, a speciality and its own recipe count

By the numbers

2,363
tests green at the release QA run, on a build I still recommended holding back
4 routes
an allergic user could still have been served their allergen, each one reproduced against the API
30 to 190s
real generation times, so generation became a job the app polls
1 of 113
generated recipes that found a real photo, until I measured it and rebuilt the matching
10 users
where the database pool ran dry, so the first pilot wave was capped there
200k
tokens per person per day, enforced inside the generation loop

Architecture

App
.NET MAUI on iOS: XAML and MVVM, Shell routing, SQLite cache, Keychain-only tokens
Networking
Refit over Polly retries with backoff, an offline guard, idempotency keys on the four writes that could act twice
API
FastAPI on async SQLAlchemy, Postgres and Alembic, 157 endpoints
Generation
GPT-5 for recipes under a strict JSON schema, GPT-4o-mini for receipt scanning, Gemini 2.5 Flash Image to draw photos and 2.5 Flash to check them
Safety
Deterministic allergen gate, model-expanded aliases that can only add, allergens hidden from the model’s pantry
Platform
ECS Fargate behind an ALB, RDS Postgres, S3, Secrets Manager, GitHub OIDC with no static keys
Billing and notifications
Stripe subscriptions behind a signature-verified, idempotent webhook, JWT sessions over bcrypt hashes, Firebase push driven by a rule engine over user state
Delivery
pytest, mypy and ruff in CI, one image built per commit, ECS deploys serialised per branch so two merges cannot race, migrations run after the service update
Release
Signed store builds from a clean tree, symbols uploaded before the build is kept, consent-gated crash reporting

Decisions

  • Rules decide whether you see a recipePlain code checks a draft recipe before anyone sees it. It looks for your allergens and their aliases in the wording of the steps as well as the ingredient list. A model expands “shellfish” into what counts as shellfish, but it can only add to that list.
  • Three attempts, then it says soA failed draft comes back with the exact line that broke the rule, three attempts in total. The model never sees an allergen in your pantry either, so it can’t reach for one by accident.
  • When the safety check started inventing problemsA second model reviewed recipes too, and it turned out to be the least trustworthy part. One tester’s run came back with all fifteen drafts vetoed across every chef, so there was nothing to show. Another’s drafts were refused for gluten, egg and dairy when the only things declared were shellfish, turkey, onions and bell peppers.
  • Keeping the reviewer below the rulesThe deterministic check runs first and has the final say. The model has to name the terms it objected to, and a veto naming something the user never declared gets overruled. A veto that names nothing still wins, because I have no way to disprove it.
  • Why generation failed in productionRecipe generation kept failing in production and nowhere else. Asking a model why gets you a plausible theory, so I measured the running system instead. The load balancer was still on its default sixty-second idle timeout while real generations took anywhere from thirty seconds to over three minutes, and the lock left behind meant every retry was refused too.
  • Turning generation into a jobI raised the timeout to five minutes, but that only moved the limit. Generation starts a job now, the app polls it, and a dropped connection or a killed app picks the job back up.
  • The photos nobody was gettingGenerated images take two calls. One draws the dish, and the other lists only what it can see. If anything on the plate is not in the recipe, the image is thrown away.
  • Counting the photo matchesIn production I went further and served a small reviewed library instead of generating. When I counted, the library matched 1 recipe in 113. Its vocabulary came off a stock photo deck while the real recipes were one-pan Latin breakfasts, and skillet, which came up 27 times, was not a format it knew.
  • Two deploys, four minutes apartTwo pull requests merged minutes apart, each baking its own task definition, and a feature flag ended up set by whichever Docker build finished last. Nothing logged it. I put a queue around the deploy job so a second one waits instead of cancelling, and made deploys read the flags already running instead of resetting them to defaults.
  • Telling the client not to shipBefore the build went to outside testers I reviewed it end to end against the live API. I found four ways an allergic user could be shown, planned and cooked their own allergen without a warning, plus a production server still running a development configuration.
  • Recommending a smaller first waveI wrote it up and recommended not shipping to allergy or medical testers, with a staged fix plan and a first wave capped at ten people.
  • A build you can trace back to codeA build that exists only as edits on a laptop, or crash reports landing in an account the project does not own, means you can’t tell what is running or watch it fail. So every build is tied to a commit and every crash report lands in the project’s own account.
  • Tying each release to a signed tagThe release script refuses to build from a dirty tree, commits the version bump itself, signs a tag on the exact commit it built, and throws away a build whose debug symbols failed to upload. It lands with the next build, so it has not stamped a shipped build yet.
  • Working with agents as the only reviewerI ran the API, the AI and the AWS side alone and shared the iOS app with one other developer. That only worked because agents did the mechanical share, including scaffolding, migrations and the tests that carried the suite past two thousand.
  • What the agents were not allowed to doThey get an approved plan before they edit anything, they run in parallel only where files don’t overlap, and the allergen logic stays plain code no model may rewrite. I don’t hand off review or the final diff, and that is how the four allergen routes turned up before testers did.
Recipe Library: 79 recipes, searchable and filterable by chef
Meal Plan: a week of slots you drag recipes onto, or generate in one go

Outcome

The app is on TestFlight by invite: four chefs, 79 recipes, a pantry you fill by photographing a receipt, and meal plans built from what you already own. I owned the API, the AI and the AWS platform outright and the iOS app alongside one other developer. The suite stood at 2,363 at the release QA run, and I still recommended holding the build back.

Résumé (opens in a new tab)