← Back to Blog

One Founder, One AI Stack: Three 2026 Cases, Fact-Checked

The Compass Team

March 14, 2026 · Updated September 8, 2026

One Founder, One AI Stack: Three 2026 Cases, Fact-Checked

The claim circulating all year is that one person with the right AI stack can now run a company. Three stories get cited most often. Here they are, checked against primary sources, including the parts the flattering profiles leave out.

Last verified: September 7, 2026. Every number, price and outcome below traces to a primary source or a major publication, linked inline.

In all three cases the machine produced the output and a human still had to decide whether the output was worth having. That gap is the whole story.

The Shopify CEO ran 120 experiments and the pull request is still open

On March 11, 2026, Tobi Lütke opened pull request #2056 against Shopify's Liquid template engine, from a branch named autoresearch/liquid-perf-2026-03-11.

The numbers in the PR: roughly 120 automated experiments, 93 commits, parse plus render time down 53% (7,469µs to 3,534µs), object allocations down 61% (62,620 to 24,530), 974 unit tests passing against a benchmark built from real Shopify templates.

He did not write those commits. He ran the Pi coding agent with a pi-autoresearch plugin, a variant of Andrej Karpathy's autoresearch pattern, where an agent proposes a change, runs the tests, benchmarks it, keeps what wins and discards the rest. Simon Willison's writeup names the load-bearing part: the 974-test suite is what made the loop safe to run unattended.

One correction to the version that spread in March. The PR was never merged. As of September 2026 it is still open, with 3 of 4,192 specs failing and readability regressions that at least one detailed critique called unmaintainable.

That is the real lesson. The loop generated weeks of candidate work in days. It could not judge whether the benchmark was the right target, and it could not agree to carry the maintenance cost.

The writer who ships his own software

Craig Mod walks across Japan and publishes books about it. Engineering has never been his job. In Software Bonkers, published March 2026, he calls himself "an OK-but-not-great coder" whose software opinions had stayed "locked in my noggin" for years.

Then he started building. A rebuilt Twitter that works the way he always thought it should: posts expire after 7 days, two posts a day, reverse-chronological, no algorithm. TaxBot2000, custom accounting software covering multiple currencies, foreign bank accounts and his real tax reconciliation, built in about five days. Auto-generated chapters for his members-only livestreams. A searchable video archive with indexed timestamps.

None of this is a company, which is why it matters. It shows the constraint moving. The bottleneck was never engineering capacity. It was having a specific opinion and no economical way to express it.

The two-person company doing $401 million, and its warning letter

Matthew Gallagher launched Medvi, a GLP-1 telehealth business, from his Los Angeles home in September 2024 with about $20,000 and no employees. Headcount today is two: him and his brother. The company reported $401 million in revenue in 2025 across 250,000 customers at a 16.2% net margin, with 2026 tracking toward $1.8 billion. He ran code, site copy, ad video and customer service through more than a dozen AI tools, treating each business function as a prompt.

One caveat belongs next to those numbers. The FDA sent Medvi a warning letter on February 20, 2026 over misbranding claims on its website, six weeks before the profile that made the company famous (Fortune, implicator.ai on the timeline).

You need both halves. Two people really did serve a quarter million customers, and the same automation shipped marketing claims that nobody in the loop checked against the rules. Nothing in the stack asked whether it should.

The system is the advantage, and so is its blast radius

None of these three worked harder than a team. They each built a loop.

Lütke's ran experiments while he did his actual job. Mod's turns an opinion into working software in an afternoon. Gallagher's acquires, prescribes and supports customers with two people on payroll.

Hustle scales linearly, because you only have so many hours. Loops compound, because each cycle feeds the next. That is why headcount stopped being a proxy for output.

It is also why judgment stopped being optional. Point a loop at the wrong benchmark and you get 93 commits nobody will merge. Point it at an ad channel with no compliance review and you get a warning letter.

The layer most solo founders skip

The 2026 conversation about solo founder AI tools is almost entirely about production: coding agents, design generators, deploy pipelines. Very little of it is about the thinking those tools multiply.

Where does the insight from the shower go? The pattern across three customer calls? The hypothesis you will forget by Monday?

Not a Slack channel that scrolls away. Not a Notion database you never reopen. Not a ChatGPT thread that vanishes into your history. Founders who journal with intention and reread last month's notes catch patterns earlier, and they pressure-test ideas in private first. That is what a founder's note system has to do.

The solo operator stack, with September 2026 prices

List prices in USD, checked September 7, 2026. Most rows have a usable free tier, so the realistic monthly floor is closer to $50 than $500.

Layer Options and list prices What to know
Coding agents Claude: Pro $20/mo ($17 annual), Max from $100/mo, Claude Code on every paid plan. Cursor: Hobby free, Pro $20/mo, Pro+ $60/mo, Ultra $200/mo. GitHub Copilot: Free, Pro $10/mo, Pro+ $39/mo, Max $100/mo. Paid tiers are credit pools, not seats. Heavy agent use hits the ceiling fast, which is why $100 to $200 tiers exist. A real test suite matters more than the model.
Design Figma: Starter free with 150 AI credits/day, Professional full seat $16/mo annual with 3,000 AI credits/mo, Dev seat $12/mo, Collab seat $3/mo. v0: free with $5 monthly credits, Plus $30/user/mo, Business $100/user/mo. v0 is faster from prompt to a working screen. Figma is where you fix what the generator got wrong.
Marketing and content Buffer: free for 3 channels, Essentials $5/channel/mo, Team $10/channel/mo. beehiiv: Launch free to 2,500 subscribers, Scale $43/mo, Max $96/mo. Cost tracks distribution, not headcount. Both free tiers carry you through a first launch.
Support Intercom Fin: from $0.99 per resolved outcome. Seats from $29/mo (Essential), $85/mo (Advanced), $132/mo (Expert). The one line that scales with customers rather than time. Budget it as variable cost.
Analytics PostHog: free every month for 1M events, 5,000 session replays and 1M feature flag requests, then usage-based. The free monthly allowance covers most products through early launch.
Notes and thinking Apple Notes: free. Obsidian: free for personal use, Sync $5/mo ($4/mo annual), commercial license $50/user/yr. Mem: free tier, Plus $9/mo, Pro $29/mo. Compass: early access (invite) on iPhone, 14-day trial then $9.99/mo or $99/yr for early-access members, $19.99/mo at public launch. Apple Notes wins on friction, Obsidian on ownership and plugins. Mem and Compass do the retrieval and pattern work for you. Compass is iPhone-only and founder-specific, so it is the narrowest option here.

On that last row, plainly: Compass captures by voice or text, sorts each note into founder domains like product, fundraising, hiring and go-to-market across 20-plus categories, remembers people and decisions across notes, surfaces patterns and weekly reflections, answers questions over your own notes, and turns private notes into X and LinkedIn drafts it can publish. It is also new, iOS 17+ only, and useless if you are not a founder. Apple Notes is free and already on your phone.

The playbook

Three moves, in this order.

  1. Capture the thinking. Every insight and hypothesis lands somewhere durable and rereadable. Not your head. Not a chat thread.
  2. Build loops, not habits. Habits need willpower. Loops run regardless. Design the cycle once: capture, review, decide, execute, measure.
  3. Automate execution, own the judgment. Agents will write the code, draft the ads and answer the tickets. They will not tell you the benchmark is wrong or the ad is a compliance problem.

Lütke got 93 commits and an open PR. Mod got the software he always wanted. Gallagher got $401 million and a warning letter. Same category of tool, three outcomes, and the difference each time was the human deciding what to point it at and what to accept back.

Everyone has the stack now. What differs is the thinking it multiplies.

Share this article