# The Food Maestro

> An iOS app that cooks around your allergies. Tell it what you can’t eat, and every recipe it writes has to get past that before you see it.

Consumer iOS · Food safety · Generative AI · [Walkthrough on request](https://cal.com/abhishek-kolge-yed7dg/intro)

Case study from https://abhishekkolge.dev. Canonical page: https://abhishekkolge.dev/work/food-maestro

- **Role:** Lead engineer: the backend, the AI and the platform on my own, the iOS app with one other developer
- **Stack:** .NET MAUI on iOS, FastAPI on async SQLAlchemy, Postgres, GPT-5 via OpenRouter, Stripe, ECS Fargate with GitHub OIDC, Sentry

## Problem

Cooking around an allergy is a research job. Every recipe gets read twice, and a substitution you didn’t think about can put someone in hospital. Most recipe apps suggest anything; the ones that filter do it on tags, which is only as good as whoever typed the tags in. A language model makes it worse, because it will invent an ingredient that was never there.

## Approach

You pick a chef, set what you can’t eat, and tell the app what’s in your kitchen. The model writes the recipe, and plain code decides whether it is safe. It checks every draft against your allergy list, including the wording of the steps, and sends it back if anything slipped through. Photos get the same treatment from a second model that only says what it can see.

## By the numbers

- **2,363**: tests green at the release QA run, on a build I still recommended holding back
- **4 routes**: an allergic user could still have been served their allergen, each one reproduced against the API
- **30 to 190s**: real generation times, so generation became a job the app polls
- **1 of 113**: generated recipes that found a real photo, until I measured it and rebuilt the matching
- **10 users**: where the database pool ran dry, so the first pilot wave was capped there
- **200k**: tokens per person per day, enforced inside the generation loop

## Architecture

- **App:** .NET MAUI on iOS: XAML and MVVM, Shell routing, SQLite cache, Keychain-only tokens
- **Networking:** Refit over Polly retries with backoff, an offline guard, idempotency keys on the four writes that could act twice
- **API:** FastAPI on async SQLAlchemy, Postgres and Alembic, 157 endpoints
- **Generation:** GPT-5 for recipes under a strict JSON schema, GPT-4o-mini for receipt scanning, Gemini 2.5 Flash Image to draw photos and 2.5 Flash to check them
- **Safety:** Deterministic allergen gate, model-expanded aliases that can only add, allergens hidden from the model’s pantry
- **Platform:** ECS Fargate behind an ALB, RDS Postgres, S3, Secrets Manager, GitHub OIDC with no static keys
- **Billing and notifications:** Stripe subscriptions behind a signature-verified, idempotent webhook, JWT sessions over bcrypt hashes, Firebase push driven by a rule engine over user state
- **Delivery:** pytest, mypy and ruff in CI, one image built per commit, ECS deploys serialised per branch so two merges cannot race, migrations run after the service update
- **Release:** Signed store builds from a clean tree, symbols uploaded before the build is kept, consent-gated crash reporting

## Decisions

### Rules decide whether you see a recipe

Plain code checks a draft recipe before anyone sees it. It looks for your allergens and their aliases in the wording of the steps as well as the ingredient list. A model expands “shellfish” into what counts as shellfish, but it can only add to that list.

### Three attempts, then it says so

A failed draft comes back with the exact line that broke the rule, three attempts in total. The model never sees an allergen in your pantry either, so it can’t reach for one by accident.

### When the safety check started inventing problems

A second model reviewed recipes too, and it turned out to be the least trustworthy part. One tester’s run came back with all fifteen drafts vetoed across every chef, so there was nothing to show. Another’s drafts were refused for gluten, egg and dairy when the only things declared were shellfish, turkey, onions and bell peppers.

### Keeping the reviewer below the rules

The deterministic check runs first and has the final say. The model has to name the terms it objected to, and a veto naming something the user never declared gets overruled. A veto that names nothing still wins, because I have no way to disprove it.

### Why generation failed in production

Recipe generation kept failing in production and nowhere else. Asking a model why gets you a plausible theory, so I measured the running system instead. The load balancer was still on its default sixty-second idle timeout while real generations took anywhere from thirty seconds to over three minutes, and the lock left behind meant every retry was refused too.

### Turning generation into a job

I raised the timeout to five minutes, but that only moved the limit. Generation starts a job now, the app polls it, and a dropped connection or a killed app picks the job back up.

### The photos nobody was getting

Generated images take two calls. One draws the dish, and the other lists only what it can see. If anything on the plate is not in the recipe, the image is thrown away.

### Counting the photo matches

In production I went further and served a small reviewed library instead of generating. When I counted, the library matched 1 recipe in 113. Its vocabulary came off a stock photo deck while the real recipes were one-pan Latin breakfasts, and skillet, which came up 27 times, was not a format it knew.

### Two deploys, four minutes apart

Two pull requests merged minutes apart, each baking its own task definition, and a feature flag ended up set by whichever Docker build finished last. Nothing logged it. I put a queue around the deploy job so a second one waits instead of cancelling, and made deploys read the flags already running instead of resetting them to defaults.

### Telling the client not to ship

Before the build went to outside testers I reviewed it end to end against the live API. I found four ways an allergic user could be shown, planned and cooked their own allergen without a warning, plus a production server still running a development configuration.

### Recommending a smaller first wave

I wrote it up and recommended not shipping to allergy or medical testers, with a staged fix plan and a first wave capped at ten people.

### A build you can trace back to code

A build that exists only as edits on a laptop, or crash reports landing in an account the project does not own, means you can’t tell what is running or watch it fail. So every build is tied to a commit and every crash report lands in the project’s own account.

### Tying each release to a signed tag

The release script refuses to build from a dirty tree, commits the version bump itself, signs a tag on the exact commit it built, and throws away a build whose debug symbols failed to upload. It lands with the next build, so it has not stamped a shipped build yet.

### Working with agents as the only reviewer

I ran the API, the AI and the AWS side alone and shared the iOS app with one other developer. That only worked because agents did the mechanical share, including scaffolding, migrations and the tests that carried the suite past two thousand.

### What the agents were not allowed to do

They get an approved plan before they edit anything, they run in parallel only where files don’t overlap, and the allergen logic stays plain code no model may rewrite. I don’t hand off review or the final diff, and that is how the four allergen routes turned up before testers did.

## Screens

- Home: today’s suggested lunch with its chef, cook time and calories, and shortcuts to Recipe Now and Meal Plan
- Meet Our Chefs: four chefs, each with a cooking style, a speciality and its own recipe count
- Recipe Library: 79 recipes, searchable and filterable by chef
- Meal Plan: a week of slots you drag recipes onto, or generate in one go

## Outcome

The app is on TestFlight by invite: four chefs, 79 recipes, a pantry you fill by photographing a receipt, and meal plans built from what you already own. I owned the API, the AI and the AWS platform outright and the iOS app alongside one other developer. The suite stood at 2,363 at the release QA run, and I still recommended holding the build back.

## Next

- [Hedged](https://abhishekkolge.dev/work/hedged.md)
