Streaks and push alerts can prompt another visit. They cannot make a static product useful twice.
AI can make a small team look larger than it is. The hard part is deciding which promises the product should own before a permanent product leader is worth hiring.
A working demo proves the model answered once. The spec has to explain what users should trust, review, or ignore when the next answer is partial, stale, or wrong.
Two months of evenings, four stages, three thousand lines of Python: what I learned building a multi-voice audiobook pipeline for my novel Cold Storage.
I built a harness to compare Claude and GPT on real PM scenarios, then used swap-testing to separate real quality differences from judge bias.
I set out to build a harness that would let an AI act as a neutral judge. The tricky part was realizing the judge wasn't actually neutral.