LeakyButton Sign in Start free

Stories

Six ways a button eats an AI budget.

Each one starts with code that works. That's the problem: a feature that works gets shipped, and nobody asks how many times it can run.

These are illustrative scenarios, not reports of real incidents. The companies are invented. Each story is a composite of patterns that are common in AI-built apps, and every number is worked out from the assumptions printed next to it, so you can check the arithmetic and swap in your own.
  1. 01 The weekend demo
  2. 02 The retry loop that ran all night
  3. 03 The key in the bundle
  4. 04 The suggestion that ran on every keystroke
  5. 05 The double tap
  6. 06 The free tier somebody found

The weekend demo

Pocket Portrait turns a selfie into a painted avatar. Its maker built it over two evenings with an AI app builder, put it behind a single "Generate again" button and posted it to a subreddit on Friday night.

The button called an image model on every press. It didn't lock while it worked and the server had no limit, because nothing in the prompt that built it had asked for one. People liked it. Some liked it a lot. On Saturday afternoon somebody noticed the endpoint took a plain JSON body and pointed a script at it, using it as a free image farm until Monday morning.

The bill arrived on Tuesday.

The arithmetic

Visitors over the weekend2,400
Average generations each6
Generations by people14,400
Script: one call every 2 s, for 10 h18,000
Total image calls32,400
Price per image (assumed)$0.04
Weekend total$1,296

How it gets caught

The first visitor past 15 presses in an hour raises "Generate again" has no limit, about forty minutes after the post went up. When the script starts, Automated sessions are spending on AI follows within a minute: no pointer, clicks exactly two seconds apart.

The fix

A guard switched on from the alert (10 generations per visitor per hour) stops the people at once, with no deploy. The prompt for the coding assistant adds the server limit and a free-account sign-in after three tries, which stops the script.

The retry loop that ran all night

Cartwheel sells a support assistant for small online shops. Its chat widget retried any failed request, forever, 600 milliseconds apart. That was fine until a database migration broke one column.

From then on the server asked the model for an answer, tried to save it, failed, and returned a 500. The widget asked again. Each retry paid for a complete answer that was thrown away. Twelve support agents had left the dashboard open in a tab when they went home.

The arithmetic

Tabs left open12
Retries per tab per hour (one per 0.6 s)6,000
Hours until someone looked7
Paid model calls504,000
Price per call (assumed)$0.002
One night$1,008

How it gets caught

A retry loop is hammering POST /api/assist appears in the first minute: a 500, then the same call within three seconds, again and again, with no click in between. Most calls to POST /api/assist fail follows on the hour.

The fix

Retry with exponential backoff, three attempts at most, never on a 4xx. On the server, save first and call the model second, or return the answer even when the save fails. A guard of 10 calls a minute per tab caps the damage while that ships.

The key in the bundle

Brightdesk summarizes meeting notes. To get the first version out, the model was called straight from the browser, with the API key in a public environment variable that the build tool happily put into the JavaScript bundle.

Nine days after launch, someone found it in the network tab and started using it from their own servers. Those calls never touched Brightdesk's pages, so nothing on the site looked unusual. The provider's monthly limit finally stopped it, two days later.

The arithmetic

Hours the key was used by someone else52
Spend per hour at their volume (assumed)$25
Before the cap kicked in$1,300
Plusa new key, a redeploy and a support week

How it gets caught

LeakyButton could not have seen the theft itself: those calls came from someone else's server. What it does see is the first page view on launch day, where the browser calls the AI provider directly. That raises Your page calls api.openai.com directly as critical, nine days before anyone found the key.

The fix

Rotate the key. Move the call behind a server route that reads the key from a server-side variable, and give that route a per-user limit. The fix prompt writes the route and finds every place the key appears in client code.

The suggestion that ran on every keystroke

Recipe Muse suggests ingredient swaps while you type a recipe. The suggestions came from a small model, fetched in an effect whose dependency was an object rebuilt on every render. Typing re-rendered the editor, so every keystroke asked the model again, mostly with text it had already seen.

Nobody noticed, because the suggestions were right and fast. The spend just grew with every new user.

The arithmetic

Sessions a day3,000
Keystrokes in a typical recipe400
Calls a day1,200,000
Price per call, small model (assumed)$0.0008
A day$960
After a 600 ms debounce (about 3 calls a session)$7.20 a day

How it gets caught

POST /api/suggest runs with nobody clicking: hundreds of calls per session with no press before them. Identical requests are paid for again lands next to it, because half the calls sent text that had already been answered.

The fix

Give the effect a stable dependency, debounce it by 600 ms, and cache answers by a hash of the text. The fix prompt names the endpoint and the page, so the assistant goes straight to the right effect.

The double tap

Lexi reviews contracts. Upload one, press "Analyze", get a summary of the risky clauses. Long documents make every analysis expensive.

The button didn't lock while the analysis ran, and on phones a lot of people tap twice when nothing visibly happens. Every second tap ran the whole analysis again, in parallel, and the second answer simply replaced the first.

The arithmetic

Analyses a day1,800
Share that ran twice22%
Duplicate analyses a day396
Price per analysis (assumed)$0.05
A month of duplicates$594

How it gets caught

"Analyze" sends POST /api/analyze twice: the same request body twice within half a second, from the same press. The Buttons page shows the button locks 0% of the time.

The fix

Lock the button while it runs and show progress. Send an idempotency key so the server answers a duplicate from the first call. Two small changes, and the fix prompt asks for both.

The free tier somebody found

Snapcaption writes captions for photos, free and without an account, which was the whole pitch. A scraper found it, opened it in a headless browser on three machines, and pressed "Caption" every 1.2 seconds for four days to caption its own image library.

Each call was cheap. There were simply a lot of them, and the endpoint had never once said no.

The arithmetic

Machines3
Calls per machine per day (one per 1.2 s)72,000
Days4
Calls864,000
Price per call (assumed)$0.001
Four days$864

How it gets caught

Automated sessions are spending on AI in the first hour: the browser reports it's automated, clicks come exactly 1.2 seconds apart, and the pointer never moves. POST /api/caption has never answered 429 explains why it kept working.

The fix

Keep the free tier, but limit guests per IP, add a challenge for automated browsers, and offer an account for heavier use. The spend alert would have flagged the first day's jump against the usual week.

Every one of these starts the same way.

A button that works, and nothing counting how often it runs. One script tag counts it, tells you when it looks wrong, and hands you the fix.

All six stories are illustrative composites. Names are invented; prices per call are stated assumptions in line with common model pricing at the time of writing, and will differ for your provider and prompts.