The App Started Guarding Before the Invoice Arrived - an AI Cost-Cap Spend Guard in AIO Helper
DEV Community

The App Started Guarding Before the Invoice Arrived - an AI Cost-Cap Spend Guard in AIO Helper

Conclusion

On April 16, 2026, I added a mechanism to the API server of AIO Helper, our SEO analysis SaaS, that stops processing before a site's monthly AI API spend exceeds its per-site cap (default $50). A request whose estimated cost exceeds the remaining budget is stopped before the AI is called. Usage is recorded after the AI call succeeds. I combined usage recording and the cap check into a single INSERT instead of splitting them into two steps. That said, the strict cap guarantee under concurrency has not been verified against real Postgres. Also, actual costs above the estimate, or overruns caused by concurrent requests, cannot be stopped by after-the-fact recording.
On April 23, I added a warning that fires when the remaining budget drops below 10% (240 tests passed after implementation - from the execution records at the time). The check covers the three features that call the text-generation AI (page-goal auto-generation, title and description suggestions, and batch suggestions). Batch suggestions re-check right before each page's call. AI calls that create search vectors were not covered at this point.

Main text (about 8 min read)

What I Was Building

The target is AIO Helper, our own SaaS for SEO operations. It ingests site data from sources such as Google Search Console and proposes improvements page by page. It has two parts: the admin UI that users operate, and the API server (seo-api) that does the processing. For an overview of the whole product, see the AIO Helper introduction.
The AI is called on the API server side. As of April 2026, three features called the text-generation AI API:

  • Page-goal auto-generation (POST /v1/page-goals/auto-generate). From a page's title and body, it drafts target keywords, the page's purpose, and the intended reader.
  • Title and description suggestions (POST /v1/suggest). Improvement proposals for one page.
  • Batch suggestions (POST /v1/suggest/batch). Proposals for many pages, starting from those with the most clicks.

There were also calls that create search vectors (embeddings, which turn text into a list of numbers) to find related documents: the ones used inside the two suggestion features, and the API that rebuilds the vectors (POST /v1/embeddings/rebuild). The April mechanism did not cover these (explained later under "What This Design Has Not Verified").

For users, the value is that analysis finishes fast and manual research time goes down. But an AI API costs money every time you call it. The bill grows not only when usage grows, but also when a bug repeats the same processing. Just showing the spend in an admin dashboard is not enough - by the time you notice, the cost has already been incurred. So I built two things:

  1. A mechanism that stops before spending. Each site has a monthly cap, and a request without enough remaining budget is never sent to the AI API.
  2. A ledger you can trace after spending. Site, model, endpoint, token counts, and cost are recorded for each text-generation call.

On April 16, a site's monthly cap was changed through the API (PUT /v1/sites/:siteId/budget); on April 23 I also added an "AI Budget" tab to the admin UI. The cap and the usage records live in the API server's database, and the three text-generation features read the same tables right before calling the AI.
The goal is not simply to stop at $50 a month. It is to be able to trace which feature spent how much, and to confirm whether costs actually went down after changing how we use it. I aimed for a state where "stopping" and "improving" both work from the same records.

Where to Stop a Single Request

The processing order is as follows:

  1. The API server receives a request for one of the three text-generation features.
  2. It calculates the cost from the model to be used and the expected input/output.
  3. It compares the month's spend and the estimated cost inside the database.
  4. If the remaining budget is insufficient, it never calls the external AI API.
    • Page-goal auto-generation and single-page suggestions return HTTP 402.
    • Batch suggestions skip the pages that don't fit the remaining budget and return HTTP 200 with the number of skipped pages.
  5. If the budget is sufficient, it calls the AI API, and after success it records the actuals in the usage ledger.
  6. At record time, the same SQL statement checks the cap again.

What matters in this order is not treating the stop decision as "cleanup after the AI API returns an error." Once the request has gone out, you cannot cancel the cost. Only by stopping at the app's entry point does the cap function as a control.
At the same time, this record is not a substitute for the invoice. The app estimates from the token counts it knows and its own price table. It will not necessarily match the provider's final bill exactly. The in-app ledger is the number for stopping early; the invoice is the number you finally pay - separate roles.

It Started with "There's a Display, but No Cap"

The trigger was a test audit. The AIO Helper admin panel had a widget showing "AI spend this month," but there were zero tests enforcing a cap. Even if you can calculate the spend, if you cannot stop before exceeding the cap, you keep calling pay-per-use generative AI models.
The first thing I did was separate the pricing calculation from the data model. I decided the price table would live statically in code, not be fetched from an external API. Three reasons:

  • The AI API response does not include cost.
  • AIO Helper calls the OpenAI API directly, and the response contains only token counts.
  • Prices don't change that often. A static price table is realistic to maintain.
    I didn't want billing decisions to depend on an external call. This avoids the loop where the API call made to learn the cost fails and the billing decision becomes impossible.
// ai-cost.ts - unknown models are deliberately estimated high
const FALLBACK_PRICE = {
  input : 0.02,
  output : 0.08
};
// calcCost() rounds to 6 digits, clamps negatives to 0, and never throws
// estimateCostFromBytes() uses 3 bytes/token (between English ~4 and Japanese ~1.5-2)
// and assumes output is 25% of input, for the preflight estimate

The point is that FALLBACK_PRICE is deliberately high ($0.02/$0.08 per 1K). If you add a new model and forget to register it in the price table, its cost is treated as 0 and it slips past the cap. So treating unknown models as "expensive" is the safe side.

Checking the Cap and Recording Usage in One Statement

The core of the spend stop lives in budget.ts. There are two tables:

-- seo_site_budgets: per-site monthly cap (default $50)
-- monthly_usd_cap NUMERIC(10,2)

-- seo_ai_usage: append-only ledger, one row per call
-- cost_usd NUMERIC(10,6), created_at timestamptz

I used NUMERIC instead of FLOAT for amounts to avoid rounding errors accumulating over millions of rows.
What I agonized over most was how to handle two requests arriving right at the cap at the same time. Instead of SELECT FOR UPDATE with a transaction, I chose a single INSERT statement that includes the cap check:

INSERT INTO seo_ai_usage (
  site_id,
  model,
  endpoint,
  input_tokens,
  output_tokens,
  cost_usd
)
SELECT
  $1 :: text,
  $2 :: text,
  $3 :: text,
  $4 :: int,
  $5 :: int,
  $6 :: numeric
WHERE (
  COALESCE (
    (SELECT monthly_usd_cap FROM seo_site_budgets WHERE site_id = $1),
    $7 :: numeric
  ) - COALESCE (
    (SELECT SUM(cost_usd) FROM seo_ai_usage WHERE site_id = $1 AND created_at >= date_trunc('month', NOW())),
    0
  )
) >= $6 :: numeric
RETURNING id;

$7 is the default cap ($50) used when the site has no budget row.
The upfront remaining-budget check is done by checkBudget(), and if the budget is insufficient, it stops without calling the upstream AI fetch (the two single-page features return HTTP 402 Payment Required).
recordUsage(), which runs after a successful AI call, returns { ok: false } if this INSERT returns 0 rows. At that point the cost has already been incurred, so the request is not failed; it writes a warning log. Because nothing goes into the ledger, the full cost of that call is missing from the ledger, and since the recorded total does not grow, later cap checks do not reflect it either.
cap = 0 works as a shutoff setting meaning "don't spend a cent."

What a Single INSERT Statement Can Verify

The cap check is not a separate SELECT statement; it is a correlated subquery inside the INSERT's WHERE clause. Within a single statement, no application-side work can slip in between the aggregation and the write. However, what I could verify independently was only the structure of this SQL. That two statements starting at the same time will always include each other's preceding rows in the aggregation has not been confirmed with contention tests against real Postgres. If you need a strict cap guarantee, you need to make per-site ordering explicit - locking the budget row, serializable isolation, or advisory locks.
This design also had a test environment problem. The tests ran on pg-mem (an in-memory Postgres), which did not implement date_trunc('month', NOW()), so according to the records at the time, 12 of the 18 tests failed with 500s.

  • 12 out of 18 tests in integration-ai-cost-cap.test.ts failed
  • Only the 3 pure calcCost unit tests pass

At the time of the record, the only tests passing were the calcCost unit tests that never touch the database. Some tests also failed for reasons other than the 500s; what the records show is a case where PUT budget returned 400 due to an input validation issue. According to the test records, registering date_trunc with db.public.registerFunction() cut the failures to 5. The rest were values inside the INSERT...SELECT whose types could not be resolved. node-postgres sends untyped values as strings, so I fixed it by adding explicit casts like $1::text.
The test records also include an experiment where I deliberately broke the date_trunc WHERE clause. 17 of the 18 tests still passed with the monthly extraction condition broken; only the "don't count last month" test failed. That test inserts a previous-month row dated 60 days back and confirms it is not aggregated. If someone accidentally deleted the monthly filter, this one test would be the only thing catching it. I learned that verification of a critical condition was concentrated in a single test.

April 23: Adding a Warning Below 10% Remaining

Stopping at the cap alone means users suddenly get a 402 without knowing why it stopped. So I added a soft warning that sets a flag once the remaining budget drops below 10%.

export const WARNING_REMAINING_PCT = 0.1;
function isNearLimit(cap: number, remaining: number): boolean {
  if (cap <= 0) return false; // a kill switch is not a "warning"
  return remaining / cap < WARNING_REMAINING_PCT;
}

Explicitly excluding cap <= 0 is so that "usage is shut off" and "the cap is near" are treated as separate notifications. According to the execution records, I first wrote 4 failing tests, and after implementing the warning condition, 240 tests passed. The change was committed as feat: AI budget soft-warning flag at <10% remaining.
Separating the warning from the shutoff also lets you separate the wording in the UI and notifications. If the remaining budget is low, the user can reduce their workload or ask an admin to raise the cap. If the cap is set to 0, an admin has intentionally stopped usage. Show both as the same yellow warning, and users will misread it as "wait and it comes back."

Don't Let the Ledger End as Just a Stop Mechanism

Notifying only that the cap was reached does not tell you what to fix next. For operations, you need at least these five things:

  1. How much was spent this month.
  2. The baseline number to compare against the monthly cap.
  3. How much it grew today. A sudden increase can mean not just more users but also retry loops or bugs.
  4. Which feature spent it. Without recording the endpoint, you cannot pick reduction candidates.
  5. Unit cost per call and per deliverable. More calls means something different if the number of suggestions is growing at the same rate.
  6. How things changed before and after a change. Record the day you changed the model or the input volume, and compare unit costs afterward.

The April seo_ai_usage table has the columns (site, model, endpoint, token counts, cost, timestamp) to answer the first three. There was no screen yet for daily growth or before/after comparisons. Still, as long as each call is recorded, you can aggregate later and compare reductions with numbers instead of gut feeling.

What This Design Has Not Verified

I cannot present this article as "a finished version that never exceeds the budget even under concurrency." Four reasons:

  1. Contention tests against real Postgres are not done.
  2. Collapsing everything into a single SQL statement is not the same as complete per-site ordering control.
  3. Some calls never reach the ledger. Usage is recorded only after a successful AI call, and failed calls are not recorded by design. However, in page-goal auto-generation, if reading the response fails after the AI has responded, it returns an error without recording. The cost is incurred, but it is not in the ledger.
  4. Creating search vectors was outside the cap. The embeddings created inside the two suggestion features and the API that rebuilds the vectors had neither a cap check nor ledger records. In September 2026, I added a cap check and recording to the rebuild API, and recording to the embeddings inside the suggestions.
    Behavior when the budget database is unavailable is also unverified. Keep processing and you cannot honor the cost cap; stop and you halt users' work. You have to decide which takes priority, and prepare failure-mode tests and notifications.
    So this implementation is not an endpoint. It is the first stage of treating cost as something to control. Follow-up articles move on to reconciling against billing data, identifying fixed costs, and re-measuring after changes. Skip this order, and you cannot explain reduction effects with numbers.

What Came After

(The article ends here; no further content provided.)

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.