maxOutputTokens Was Missing, So Usage Had No Ceiling
This article may contain affiliate links. Its content is not affected by advertising.
In short
A student chat reachable via anonymous magic link had its rate limit scoped to one Cloud Run instance, and the cookie could be regenerated, so output and calls had no cap.
Conclusion
The student chat callable from an anonymous magic link had no cap on either output tokens per response or total calls per day. A per-minute rate limit did exist, but it was scoped to a single Cloud Run instance’s in-memory state, so each running instance counted separately, and the device cookie — one of its identifying keys — could be regenerated client-side on every request. A 2026-10-01 security audit found this, and output-token and cross-instance daily caps went in the same day. The fix itself turned up two more review findings.
Not an incident so much as a documented limit
This wasn’t abuse discovered in production. It turned up inside a comment already written into the existing rate-limit implementation.
/**
* F06 (#42 slice 1): **dual-key rate limit** for student Q&A.
*
* Acceptance: "10 questions per minute per magic_link, 10 per minute per device cookie."
*
* **Scope and limits (stated honestly)**: state lives in a module-level Map, scoped
* per-instance within a single Cloud Run instance. A cross-instance total for
* multi-instance deployments needs #155 (distributed rate limiting / shared store).
* This limiter is a first line of defense against abuse/runaway calls within a single process.
*
* **Memory bound (#437 Low-4, public-endpoint DoS mitigation)**: magic_link / cookie keys
* are spoofable — an attacker could send a unique key on every request and grow the Map unbounded
*/
The implementer had already written down that this was a per-instance limit, that a system-wide cap needed a separate ticket, and that the cookie was spoofable. As a “first line of defense within one instance,” the limiter worked as designed. But it never answered the question of how far one leaked magic link could go once Cloud Run has scaled to multiple instances. At 10 calls per minute per instance, 5 running instances means 50 calls per minute combined, and regenerating the cookie on every request neutralizes the device-side cap entirely. On top of that, neither the output tokens nor the thinking tokens per response had any cap, so a single question like “repeat X ten thousand times” could drive one response all the way to the model’s own output ceiling. Put together, one leaked link meant effectively no ceiling on Gemini usage.
Cause
The 2026-10-01 security audit re-evaluated how much usage this “honestly documented limit” plus “unbounded output” could actually add up to, and found there was no structural cap. The fix split into two parts.
- Cap a single response’s output at the code level — add
CHAT_MAX_OUTPUT_TOKENSinchat-stream.tsand pass it intostreamText.
/**
* Max output tokens per response (2026-10-01 security audit — fix the usage
* ceiling structurally).
*
* Student chat is callable anonymously (magic link), so a question like
* "repeat X ten thousand times" could drive output to the model's own ceiling,
* making usage unbounded from a single leaked link. Short answers grounded in
* posted material (ADR-028) need 1024 (roughly 1,500 Japanese characters).
* Thinking is billed as output too, so thinking is pinned to 0.
*/
export const CHAT_MAX_OUTPUT_TOKENS = 1024;
const result = streamText({
model: vertex(modelId),
system: req.system,
prompt: req.user,
...toGenerationOptions({ maxOutputTokens: CHAT_MAX_OUTPUT_TOKENS, thinkingBudget: 0 }),
});
- Cap the total call count across all instances — the new
daily-quota.tscounts calls against a 300-per-school daily limit, shared across every instance, using the existing Cloud SQL table (ai_rate_limit_windows, ADR-027). No new table or migration was needed.
/**
* The per-minute limit (`rate-limit.ts`) is scoped to a single Cloud Run instance,
* and the device cookie can be regenerated by the client on every request. That
* means one leaked magic link could run (instance count × 10 calls/minute) to Gemini
* all day with no usage ceiling. This adds one **cross-instance shared cap**.
*/
export const CHAT_DAILY_LIMIT_PER_SCHOOL = 300;
The same comment leaves a worst-case estimate: at roughly 3,000 input tokens plus 1,024 output tokens per call (about ¥0.6), one school’s worst case comes to about ¥180/day, or about ¥5,500/month. Any failure to evaluate the quota (e.g. a DB connection failure) fails closed, returning a 503 without calling Vertex at all.
The fix (including two findings from review)
The fix itself was flagged twice in PR review. Both are cases where “the correct-looking fix” quietly created a new problem.
Finding 1: the daily-limit check was running inside the RLS transaction
The first implementation read the daily-limit table from inside the student chat’s open RLS transaction. That contradicted a general rule already written into packages/db/src/client.ts.
/**
* ## Usage discipline (pool exhaustion, consistency)
*
* - **Use fork only at "leaf query" granularity. Don't call fork from inside fork**:
* if an outer lane holds a connection while waiting for an inner lane's connection,
* once the pool (`createDbClient` max:10) is full, every request can end up waiting
* on every other request in a starvation deadlock.
*/
const sqlClient = postgres(url, { max: 10, onnotice: () => {} });
The pool is capped at 10 connections with no wait timeout. The student chat’s RLS transaction holds one connection until the response finishes, so going back to the pool for another connection to check the daily limit inside that same transaction meant that just 10 concurrent requests could leave every request waiting on every other one. Review flagged this, and the check was moved in sse-handler.ts to run before the transaction opens.
// 2.5) Per-school daily limit (daily-quota.ts, shared across instances). Count it
// **before opening the RLS tx** (the tx holds a connection until the response
// finishes, so grabbing another connection inside it can exhaust the pool — Reviewer #1401).
const quotaRejection = await checkChatDailyQuota(args.schoolId, question);
if (quotaRejection) {
return errorOnlySse(quotaRejection, args.setCookieHeader ?? null);
}
The rejection is returned the same way as other rejections — a single 200-status SSE error frame.
Finding 2: the out-of-scope check ran against different text than what actually reached Gemini
The daily limit is designed not to count questions that generate no usage: empty or overlong questions (rejected with a 400 by executeChat), and questions out of scope for study/career topics (flagged out_of_scope by classifyScope). The first implementation ran that out-of-scope check against the user’s raw, unprocessed text.
export function scopeTextAsExecuteChatSees(question: string): string {
return redactSuspectedNames(maskPII(question, []).masked);
}
export async function checkChatDailyQuota(
schoolId: string,
rawQuestion: string,
opts: { quota?: RateLimiter; nowMs?: number } = {},
): Promise<ChatQuotaRejection | null> {
const v = validateQuestion(rawQuestion);
if (!v.ok) return null;
if (classifyScope(scopeTextAsExecuteChatSees(v.question)).verdict === "out_of_scope") return null;
...
The text actually sent to Gemini has already been through PII masking and name redaction, on the executeChat side. Scoping the check against the raw text meant that a question made up only of terms masking strips out entirely — like homework@example.com — was misjudged as “out of scope, don’t count,” while the actual call to Gemini on the post-masking text still went through, letting it slip past the daily limit. The fix aligns the check to run scopeTextAsExecuteChatSees — the same masking executeChat applies (maskPII(q, []) then redactSuspectedNames) — before judging scope.
Why it went unnoticed
The per-minute and per-cookie rate limit’s comment already stated, at the time it was written, both “this is a per-instance limit” and “the cookie is spoofable.” Nothing was hidden. The fix was deferred because this limit was correctly doing its job as “a first line of defense against abuse within a single instance.” The conclusion that total usage was unbounded only emerges from combining three separately-known facts: the per-instance scope, the cookie’s spoofability, and the unbounded output tokens. Looking at any one of them alone, it reasonably looks like one layer of defense-in-depth. It took a security audit re-evaluating these three already-individually-known limits as a single path to put an actual number on what they add up to when combined.
Frequently asked questions
Q1Why wasn't the per-minute rate limit enough on its own?
It was scoped to a single Cloud Run instance's memory, so each instance counted separately. Since the device cookie, one of its identifying keys, could be regenerated client-side on every request, one leaked magic link was effectively enough to slip past it and keep calling indefinitely.
Q2Why was the daily-limit check moved outside the RLS transaction?
Checking it inside the transaction meant holding that connection while fetching another from the pool for the check. With the pool capped at 10 and no wait timeout, 10 concurrent requests alone could leave every request stuck waiting, so the check now runs before the transaction opens.
Q3Wasn't a cap on output tokens alone enough?
No. Even with a single response capped, total usage stays unbounded if nothing limits the call count itself. Alongside capping one response at 1024 output tokens, a separate cap on calls per day per school, shared across instances, was added.
Q4What exactly was the out-of-scope check bypass?
It ran against the raw, unprocessed question text. A question made up only of terms masking strips out entirely, like an email address, was misjudged as out-of-scope and not counted — yet the call to Gemini still went through, slipping past the limit.
Environment verified
- ai ^5.0.52 / @ai-sdk/google-vertex ^3.0.140 (Vertex Gemini, default model gemini-2.5-flash)
- Cloud SQL connections via postgres.js, pool capped at max:10 (packages/db/src/client.ts)
- Found in a production security audit on 2026-10-01 and fixed the same day (#1401)
What this article is based on
- TypeScript file lines 13-22commit 681ad2d
- TypeScript file lines 51-59commit 64feb38
- TypeScript file lines 75-80commit 64feb38
- TypeScript file lines 13-40commit 64feb38
- TypeScript file lines 70-94commit 64feb38
- TypeScript file lines 176-182commit 64feb38
- TypeScript file lines 133-137commit 4e80de2
Every claim in this article comes from the records above. The repositories we operate are private so we cannot link to them, but which file, which lines, and at which commit we read them is recorded for every article. Nothing here is written from guesswork.