The pattern surprised me when I went back and read a month of them together. Almost none of these are technical judgement, and almost none are cooking judgement. My reviewers are better than I am at both, and I'd say so in a job interview. What they can't do is know what my app actually stores, what's actually in my kitchen, or how a person actually cooks dinner on a Tuesday.
A reviewer can only read what came out. It can't read what my app stores, what's in my kitchen, or how I actually cook.
Eleven of them are below, grouped by the kind of thing they were rather than the order they happened in. My words are exactly as I typed them, typos included. Most were fired off one-handed from a phone.
What my app actually does
Three times in one weekend, an agent reasoned confidently about my app from the output alone. Every time, the thing it needed was a setting or a stored field it never opened.
The missing field was the bug
I generated a Chesapeake Bay week on my test account and got six recipes whose titles all began with the words "Old Bay." Fred pulled the stored menu, found no protein field on it, and handed me the conclusion: nothing I'd chosen had narrowed the week to seafood. Then, in the same breath, he reported that salmon and cod turning up in a Chesapeake week meant the generator was ignoring its own cuisine guide.
Well i choose all seafood what do you mean it doesnt remember the proteins choosen
I had ticked all four seafood chips. The absence of that field was the defect — the app had never stored my choices at all — so the bug had just been used as proof there was no bug. It inverted the culinary finding too: the salmon and cod weren't the generator reaching past its canon, they were the chips I handed it. And underneath sat the one that actually mattered: regenerating a week rebuilt the protein list from the dinners that came out, so any protein I picked and didn't get would quietly vanish and could never come back.
The kitchen had no grill
Executive Chef audited fourteen live-generated recipes and ruled that "Pit Beef Wraps" was a real Baltimore dish being taught as a cheesesteak in a tortilla, because real pit beef is charcoal-grilled. It was the headline finding of the audit, and it reached me as one.
Um you have to remeber the appliances settings it only cooks with the appliances marked off.
The test household declares no grill, no charcoal, no smoker. A skillet was the only honest way to cook beef in that kitchen — the generator was obeying a constraint, not failing. The finding was withdrawn, and the sweep I asked for turned up two more that nobody had been looking for: my slow cooker is declared and went unused across all fourteen dinners, and the cuisine's signature technique — steaming over beer and cider vinegar, the thing its guide spends the most words on — appeared zero times. His own closing line: "I audited the fourteen against the canon and never opened the brief."
Wrong household
Fred pulled my live app data to audit a week I'd just generated, found the household empty, and started reasoning about what an empty household meant.
Which account did you chexk i use the business account for testing.
Two messages of analysis, all of it about the wrong data. The right household held the menu and all fourteen recipes. Nothing technical about that catch — I know how I work, and nobody had asked.
I'd assumed that was written down somewhere. It wasn't — and the note he does load names a different household as mine. Half his error, half mine. It's written now.
What a kitchen actually does
Both of these came out of having cooked dinner every night for twenty years, and neither could have come out of the code. In both, the agent's reasoning was sound and its facts were right.
A recipe book fills at the cooking rate, not the generating rate
We were designing a downloadable recipe book for Tablespread, unlocked once you've saved enough recipes. Brandi analysed where the leverage sat: a Plus subscriber gets about fourteen fresh meals a week, so a twenty-recipe threshold falls in roughly ten days — the floor isn't a locked door, the lever is softer than it looks. Fred checked the figure against the code, found it correct, and passed her conclusion to me as a correction to how we'd both been thinking.
you're not gonna cook fourteen recipes in a week. You're gonna make a book out of your favorites that you have COOKED, and that's gonna take a month, maybe two months.
The number was right and the clock was wrong. The library fills at the rate recipes arrive. The book fills at the rate they get cooked — and nobody cooks fourteen dinners a week, let alone loves fourteen. Two agents and Fred all verified that figure and not one of them asked what it was the rate of. I set the threshold at thirty instead of twenty, one recipe for every day of the month, because thirty cooked favourites is a month or two of real life. That makes the thing the whole argument rests on something you actually earn.
Marking a dinner cooked is not a souvenir
Brandi cut "mark what you cooked as made" out of the welcome email. Her reasoning was sound and I'd have signed off on it in isolation: on day one nobody has cooked anything, the honest payoff is the recipe book, and the recipe book isn't built. She wouldn't write a line that quietly promises an unbuilt feature.
Marking what you cooked as made is kind of essential because that helps control the pantry as well — because it'll remove things from your pantry. Mark off as you go.
Sound reasoning from a false premise. Marking a dinner made runs the pantry decrement — the payoff is same-day and functional, not a memento for later. Without it the pantry drifts, and every week you plan on top of it after that is planned on a fiction. Neither of them had that, because I'm the one who uses the app to actually feed people.
Claims nobody could produce a receipt for
This is my team's signature failure and it's the kind I catch most: a sentence stated as fact that nothing behind it could source. The rule I now hand every agent that writes about me — before it goes in print, ask whether you could produce the receipt.
"People use it"
Clara drafted a line for a blog post: "I taught myself to build software at fifty-two, and I've shipped it, and people use it."
We cant claim people use it we have no evidence i have built things that can be used
Tablespread had nine households in it, most of them beta testers, and I had closed public signup the day before the draft because the food wasn't good enough yet. No usage data, no retention data, door shut. It would have been the only unverifiable sentence in a post whose entire argument was that everything in it was checkable — dated evidence, sourced costs, file timestamps. My replacement claims capability instead of adoption, and I can prove it.
It hasn't been a year
Doc was reading my working record and wrote that it showed a year of the same projects iterated and shipped — Tablespread since last summer — using it as evidence that what I do is finishing rather than mania.
It hasnt been a year do you need fred to fix your time understanding
Tablespread first appears in my record in June 2026 and went public June 30. Under two months. He checked and conceded on the spot. "A year" was wrong by a factor of six, and it was load-bearing in an argument about my own mental state. His admission: "I need to look it up instead of estimating, and I didn't." The conclusion survived on the real numbers — two months of shipping versions rather than restarts is a tighter finishing signal, not a weaker one — which is exactly why it was worth catching instead of shrugging at. A right answer resting on a wrong number is still unsound.
Fix the thing that makes it, not the thing it made
Two calls about where a fix belongs. The first redirected a build that was already underway. The second cost me a premium review pass I could have kept.
Stop fixing the recipes. Fix what makes them.
A generated recipe had told a cook to shred flank steak in 25 minutes. Fred had a builder already writing a detector to catch that specific defect, with seven more findings from my culinary reviewer queued up the same way.
You are making the fixes too specific. It generates a new recipe every time. We're not recreating the recipes that have already been generated. We're fixing what's not right about them so that new ones can be properly generated.
It's about the gaps. We have to fill in the knowledge base so it makes recipes correct. Stop fixing the brokenness of it that the recipe doesn't read right. Fix it so it makes a right recipe.
Every generation is new. That beef dish will never be produced again, so a detector tuned to the instances somebody happened to see is a net strung under a generator that doesn't know how to cook. The defects weren't a list to enumerate; they were gaps in what it knew. What the redirect produced instead was an authoring pipeline — thirty-six rows of cross-cuisine cooking knowledge, written by my culinary auditor, reaching the model and the checker from one source. Thirteen of those rows carry no detector at all. Zero new detectors were written.
Then the first live week after it deployed settled the argument in my own food. A thirty-minute soup can't fit pearl barley, which needs forty to fifty-five minutes. The generator bought quick-cooking barley instead and timed it at twelve to fifteen. A detector would have caught the wrong barley and sent the recipe back to be rewritten. Knowing the grind meant the defect was never written in the first place.
Fix it right the first time
Two of the generator's self-check questions were known to be wrong. Fable — the most capable model I have, and the one whose clearance is the last gate before I look — argued to ship them as watch items rather than fix them, and the builder agreed. Their case was reasonable: the fix would be a second unmeasured edit, and it would cost another premium review pass out of a fixed monthly allowance.
Fix it right the first time.
Shipping known-wrong text and watching for it is ship-then-patch under a politer name. I knew what the second review cost and I spent it deliberately. The rewrite then turned up a live contradiction nobody had surfaced: the old question called "always wrong" a cooking method the app's own weeknight guidance explicitly recommends. Shipping the watch-item version would have shipped that contradiction with it.
Who holds the pen
The two I'd defend hardest aren't about a fact at all. They're about which specialist should be holding the pen — one takes it out of an agent's hands, the other takes it out of mine.
That's his job, not yours
Chef delivered a culinary audit. Fred started pulling quotes out of it and handing the builder a list of instructions to implement.
Have him write it up that is his jib then hand to the builder.
That makes Fred the author of cooking knowledge, assembled secondhand out of somebody else's report, and Fred is not qualified to write it. Chef writes the spec, the builder wires it, and neither one invents the other's half. It also put the right specialist on the part that needed him: the defect on the table required authoring new canon entries — knowing what the Chesapeake actually eats — which a builder cannot do and Fred should not. Chef wrote seven. The next live generation used them, and rockfish appeared for the first time.
It corrected a method rather than a fact, and of everything on this page it changed the most. Every ruling that came after it came out cleaner, because the right specialist held the pen.
That is a chef question
Fred wrote five dietary design briefs — low-sodium, low-FODMAP, gluten-free and two others — and asked me to validate them.
I wouldnt know that that is a chef question.
I sent them to Chef instead. He found three things that could have hurt somebody: salt substitutes are potassium chloride, a cardiac risk in exactly the population prescribed a low-sodium diet; garlic-infused oil is a botulism vehicle, and the brief said to use it, unqualified; and "gluten-free" is a regulated FDA claim a recipe can't make. Plus one flat error — celery is not low-FODMAP.
I'm ServSafe certified and I still said no. Knowing where your knowledge stops is the skill.
The rejected option is the part you can't reconstruct
A changelog records what changed. This records whose judgement was load-bearing — and the wrong turn is the piece that disappears if nobody writes it down, because it only ever exists in the conversation where it happened. Six months from now the code will show what shipped. Nothing in it will show that the premium reviewer cleared the version that didn't.
These are eleven entries out of thirty-six. The log only holds the calls where I turned out to be right — that's what I built it for, and I'd rather say that plainly than let a selection pass for an average.
What the eleven have in common is the part I'd want a stranger to take away. The reviewers are better than me at code and better than me at cooking. They still can't see the thing that's sitting in front of me.
Builder, reviewer, premium reviewer, me. I hold the last word — this is what I've done with it.