I start from the person and the problem, decide what the product must never do, and then build the real thing with AI coding agents until it can be argued with.
Press → or swipe to continue · Case study: Lull, 2026
Sleep is noisy. Any product that pretends otherwise gets caught the first day the baby wakes early. Precision is not the job; honesty about uncertainty is.
If the marketing page says one wake window and the app says another, the parent believes neither. In a health-adjacent product that is fatal.
Ratings, testimonials and "% of parents" convert once. When a tired parent notices the numbers were invented, everything else the product says becomes suspect.
So trust was the product. Everything else was scoped from that.
Sleep, feeds, nappies, solids and allergens, health and appointments, dressing for outside, growth, learning, and the parents themselves. The brief said "sleep tracker". The parent's day did not.
The tracker, every article and every age guide are free, forever. Premium is depth: the courses in full, the solids planner, long-range trends, export, more than one baby.
Predictions are windows with a stated confidence, never a time. Anything about medicine or safe sleep is said plainly. No invented social proof anywhere, enforced by a test.
The next nap and tonight's bedtime, as a window that narrows as the app learns this baby. On the dial, before the baby shows you.
Month-by-month age guides and a course staged by age, every claim cited, so "is this normal" has an answer that matches the engine.
First foods with preparation by age, the nine allergens one at a time on a ladder, a week of meals with swaps, and what to dress them in for today's weather.
A reaction logger that never triages, safe-sleep guidance that is always free, and a plain "when to call your doctor" that the wink never touches.
The tracker and the prediction engine. Two courses: sleep and starting solids. The solids planner and allergen ladder. Dressing by weather. Age-filtered guidance. A quiz-to-checkout funnel that runs end to end.
Prescribing a sleep-training method: the course explains them and the evidence, the app never asks you to do one. Assessing any symptom. A community feed. Anything that would need a clinician in the loop to be safe.
Every "out" is a place where being wrong would hurt a family. Those are not features to ship fast.
The dial shows an onset plus a ± that tightens with data, and a chip that names where the learning is: "Learning Mira's rhythm · day 2". Day zero falls back to age norms and still reads as progress. There is no empty state, because a tired parent on day one is exactly who needs the product to work.
Trust is kept by making uncertainty visible, not by hiding it behind a confident number.

The engine, the age guides on the public site and the course text all read the same per-month table. A page cannot quote a wake window the app disagrees with, because there is nowhere else for a number to come from.
The second "why" from slide three, turned into architecture.
When a food disagrees with a baby, the app offers the mild signs in the course's exact words, records the parent's choice, pauses the food, and stops. Severe signs go to the emergency number, never to a form. Contact irritation is carved out so families do not drop foods they never needed to.
No severity score, no triage question. A test asserts the wording still matches the course.
"Twelve months is a schedule problem. Lengthen the wake windows. Don't drop the nap."
The data marks which regression windows the evidence recognizes: four months is biological, eight to ten months developmental. A struggle at twelve months is classified as schedule, and the product says so. An opinion, stated and sourced, is more useful to a tired parent than a generic banner. It is also what makes the product worth paying for.




A working beta: public site, funnel, checkout, app and API, walkable end to end on a laptop in a minute.
What the product is, what it will never claim, how it speaks. Written before the first line of code, so the agents build against something.
Tone, no social proof, wording that must match the course: each is a test, not a memo. A build that breaks a rule fails.
Anything the product claims about itself is computed from the build. If a number cannot be derived, it does not go on the page. Same for this deck.
Every diff reviewed, every changed flow walked in a real browser. The agents are fast, confident and sometimes wrong; my job is making wrong visible early.
Quiz completion and email conversion. Share of babies whose prediction reaches "personalized" by day seven, which only happens if parents keep logging. Trial-to-paid. Course completion by module, to see where the audio earns its keep.
A clinician's review of the age table and the allergen method. A full visual redesign. The remaining audio lessons and final photography. Translations beyond the app strings. Payments, CRM, support and legal.
I keep this as a checklist with a status per line, and I would rather show it than hide it.
Lull is a pre-launch product. Screens are from the working beta; prices are omitted on purpose.