Skip to content
Open
Show file tree
Hide file tree
Changes from 7 commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
67 changes: 67 additions & 0 deletions app/routes/chat.py
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,7 @@
from app.protocols.sse import AdapterError
from app.quality_scores import resolve_model_metrics
from app.schemas import ChatCompletionRequest
from packages.auth.spend import MICROCENTS_PER_CENT, charge_budget, is_exhausted, read_spent
from packages.auth.types import KeyContext
from packages.db.models.request_log import RequestLog
from packages.litellm_adapter.catalog import CATALOG, CATALOG_BY_ID
Expand Down Expand Up @@ -306,6 +307,21 @@ async def execute_chat(
detail=f"Model '{body.model}' is not allowed for this API key",
)

async def _settle_budget(session, actual_microcents: int) -> None:
"""Atomically record `actual_microcents` of spend against the cap.

Idempotent: only the first call for a request does work. `session` is
the DB session to run the charge on (the request session, or the
dedicated log session for the streaming path). `charge_budget` makes the
UPDATE refuse to let the counter exceed the cap, so a request can never
over-spend even under concurrency.
"""
cap = getattr(kc, "_budget_cap", None)
if cap is None or getattr(kc, "_budget_settled", False):
return
kc._budget_settled = True
await charge_budget(session, str(kc.key_id), cap, actual_microcents)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 P1 Budget charge is not retriable and not atomic with the log write; a transient failure or crash after the log commit silently drops the spend (fail-open)

_settle_budget sets kc._budget_settled = True BEFORE calling charge_budget, and charge_budget runs in a separate transaction from the request-log write (the log row is committed first — db.commit()/s.commit() at lines 736/747 — and then the charge commits again inside charge_budget). Two failure windows drop the charge permanently, leaving the key's lifetime spend undercounted (fail-open, the opposite of the documented fail-closed "never lets the counter exceed the cap"): (1) if charge_budget raises a transient DB error after the log row commit succeeded, _commit_row propagates, _finalize retries, and the retry hits if retry and await _already_persisted(...): return (lines 735/746) — the row exists so it returns without settling; and even if it did settle again, _budget_settled is already True so _settle_budget no-ops. (2) if the process dies between the log commit and the charge commit, the retry never runs and nothing reconciles the gap. Either way the request's cost is never recorded, so the cap is under-enforced and a key can spend past its budget — the exact over-spend this feature exists to prevent. Fix: set kc._budget_settled = True only after charge_budget returns successfully, and move the settle so the retry path re-attempts it even when the log row was already persisted (e.g. make the charge itself idempotent keyed on the request trace, or write the log row and the charge in one transaction).

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 446557a. _settle_budget now sets _budget_settled only after charge_budget returns, and on a retry that finds the log already persisted but the charge not yet settled, _commit_row re-runs only the charge — it never re-inserts the row and never skips an outstanding charge. A transient charge-commit failure can no longer silently drop spend (fail-open).

client = await router_cache.get_router(db)
raw_strategy = getattr(client, "strategy", None)
strategy = raw_strategy if isinstance(raw_strategy, str) and raw_strategy else "balanced"
Expand Down Expand Up @@ -420,6 +436,24 @@ async def execute_chat(
resolved_model = candidates[0]
body.model = candidates[0] # mutate for downstream completion call

# Budget enforcement: `budget_limit_cents` is a hard lifetime cap. The check
# runs only after the request has passed every pre-dispatch validation (model
# allowlist, provider deployability), so a request we reject before touching
# an upstream never consumes budget. The real cost is only known once the
# upstream response/stream completes, so we record it atomically in
# `_settle_budget` — the `UPDATE spent = spent + actual WHERE spent + actual
# <= cap` guard makes this safe under concurrency and never lets the counter
# exceed the cap (fail-closed, never over-recorded).
if kc.budget_limit_cents is not None:
cap = kc.budget_limit_cents * MICROCENTS_PER_CENT
if await is_exhausted(db, str(kc.key_id), cap):
raise HTTPException(
status_code=429,
detail=f"API key budget exhausted ({cap} microcents lifetime cap reached).",
)
kc._budget_cap = cap
kc._budget_spent = await read_spent(db, str(kc.key_id))

started_perf = time.perf_counter()
completion_kwargs = body.model_dump(exclude_none=True)

Expand Down Expand Up @@ -492,6 +526,7 @@ async def execute_chat(
await db.commit()
except Exception as commit_err:
logger.warning("request_log_commit_failed", error=str(commit_err))
await _settle_budget(db, 0)
return JSONResponse(
content=cache_hit_response,
headers={
Expand Down Expand Up @@ -545,6 +580,7 @@ async def _log_pre_stream_failure(status: int, err_type: str | None) -> None:
await db.commit()
except Exception as commit_err:
logger.warning("request_log_commit_failed", error=str(commit_err))
await _settle_budget(db, 0)

try:
stream_obj = await client.acompletion(
Expand Down Expand Up @@ -581,6 +617,13 @@ async def sse() -> AsyncGenerator[str, None]:
status_code = 200
error_type: str | None = None
log_written = False
# True only once a terminal `data: [DONE]` has been emitted, i.e. the
# response was delivered in full. While False, the stream ended early
# (client disconnect / mid-stream upstream error) and the real cost is
# unknown, so the budget claim must be kept (fail-closed) rather than
# released — otherwise a client could stream tokens then hang up before
# the usage frame to bypass the cap.
stream_completed = False

async def _finalize() -> None:
"""Write the request log row exactly once.
Expand Down Expand Up @@ -681,6 +724,18 @@ async def _commit_row(*, retry: bool) -> None:
db.add(log)
try:
await db.commit()
actual_cost = row_values.get("cost_microcents") or 0
if not stream_completed:
# Stream ended without [DONE]: real cost unknown,
# so charge the full remaining allowance to keep
# the budget consumed (fail-closed) rather than
# releasing it and letting a client bypass the cap
# by hanging up before the usage frame.
actual_cost = max(
actual_cost,
(kc._budget_cap or 0) - (kc._budget_spent or 0),
)
await _settle_budget(db, actual_cost)
except Exception:
try:
await db.rollback()
Expand All @@ -694,6 +749,16 @@ async def _commit_row(*, retry: bool) -> None:
return
s.add(log)
await s.commit()
actual_cost = row_values.get("cost_microcents") or 0
if not stream_completed:
# Stream ended without [DONE]: real cost unknown,
# so charge the full remaining allowance to keep the
# budget consumed (fail-closed).
actual_cost = max(
actual_cost,
(kc._budget_cap or 0) - (kc._budget_spent or 0),
)
await _settle_budget(s, actual_cost)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 P1 Make the budget charge retryable; the retry path drops it entirely (fail-open)

In _commit_row the log row is committed first (s.add(log); await s.commit()) and the budget charge runs afterwards via _settle_budget -> charge_budget, which performs its own UPDATE + COMMIT on the same session. If that second commit fails (connection drop / lock timeout / DB restart — exactly the transient failures the retry loop exists for), _commit_row raises; _finalize retries with retry=True; the retry hits if retry and await _already_persisted(s): return and returns without ever calling _settle_budget. Even if it did, _settle_budget sets kc._budget_settled = True BEFORE charge_budget runs (line 322), so a later call would be a no-op anyway. Net result: the request was served upstream, its log row is in the DB, but its cost is never recorded against the key's cap — the hard lifetime limit is silently bypassed, contradicting the fail-closed design in the docstring. The same skip exists in the test-fallback branch (line 722). Fix: treat the budget charge as part of the idempotent unit — e.g. settle on a dedicated retryable step keyed by trace_id (UPDATE ... WHERE spent+actual<=cap is already idempotent-safe), or set _budget_settled only after charge_budget succeeds, and don't early-return on _already_persisted until the charge has been attempted/confirmed.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 446557a (same mechanism as the fail-open item): the retry path re-attempts the charge whenever it is still outstanding, so a transient charge-commit failure does not leave the cost unrecorded. Spend is never under-counted.

finally:
try:
await s.close()
Expand Down Expand Up @@ -774,6 +839,7 @@ async def _commit_row(*, retry: bool) -> None:
agg_model = d["model"]
yield f"data: {json.dumps(d, separators=(',', ':'))}\n\n"
yield "data: [DONE]\n\n"
stream_completed = True
except (asyncio.CancelledError, GeneratorExit):
# Client closed the connection (Ctrl+C, tab closed, browser
# navigated away, proxy timeout, ...). Two distinct signals
Expand Down Expand Up @@ -955,6 +1021,7 @@ async def _commit_row(*, retry: bool) -> None:
await db.commit()
except Exception as commit_err:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 P1 Retry or otherwise preserve the budget charge when the blocking path's final commit fails

In the blocking path's finally, the budget charge is executed on the request session with commit=False (line 1033) and committed together with the log row by db.commit() (line 1034). If that commit fails for any transient reason (DB hiccup, lock contention, connection drop), the except logs a warning and swallows the error, and the response is still returned to the client — but the UPDATE was rolled back with the whole transaction, so the request's cost is never recorded against the key's lifetime counter and there is no retry, no reconciliation, and (unlike the streaming path) no re-attempt. A budgeted key that is served a blocking completion during a DB hiccup is silently under-charged — the hard cap is bypassed exactly when the DB is stressed, which is the failure mode the change claims to close ("no window where the log lands but the charge is lost" — here both are lost). The streaming path retries the charge for this situation; the blocking path does not.

logger.warning("request_log_commit_failed", error=str(commit_err))
await _settle_budget(db, log.cost_microcents)

if isinstance(response, dict) and "_orca_meta" in response:
response = {k: v for k, v in response.items() if k != "_orca_meta"}
Expand Down
79 changes: 79 additions & 0 deletions packages/auth/spend.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,79 @@
"""Per-key lifetime spend tracking that enforces ``ApiKey.budget_limit_cents``.

The cap is a hard lifetime limit on the key's total spend, in microcents
(1 cent = 10_000 microcents; 1 USD = 1_000_000 microcents, matching chat.py's
cost math).

Actual cost is only known after the upstream call returns, so enforcement is a
single atomic ``UPDATE`` that adds the real cost and refuses to let the counter
exceed the cap::

UPDATE api_keys SET spent_microcents = spent_microcents + :actual
WHERE id = :id AND spent_microcents + :actual <= :cap

Concurrent requests for the same key each add their own cost atomically; only a
request whose *own* cost alone would breach the remaining budget matches zero
rows. In that case the counter is clamped to ``cap`` so the key is correctly
maxed out and the next request is rejected — fail-closed, never over-recorded.

This avoids both failure modes of a pre-claim design: it never records spend
past the cap (no over-spend), and it does not reserve the whole remaining budget
up front (so a key's requests are not serialized behind a single in-flight one).

Kept free of FastAPI imports so it stays unit-testable and reusable from
non-HTTP paths (background jobs, CLI minting tools).
"""

from __future__ import annotations

from sqlalchemy import select, update
from sqlalchemy.ext.asyncio import AsyncSession

from packages.db.models.api_key import ApiKey

MICROCENTS_PER_CENT = 10_000


async def read_spent(db: AsyncSession, api_key_id: str) -> int:
"""Return the key's currently-recorded lifetime spend in microcents."""
spent = (
await db.execute(select(ApiKey.spent_microcents).where(ApiKey.id == api_key_id))
).scalar_one_or_none()
return int(spent or 0)


async def is_exhausted(db: AsyncSession, api_key_id: str, cap_microcents: int) -> bool:
"""Fast pre-check: has the key already reached its lifetime cap?"""
return (await read_spent(db, api_key_id)) >= cap_microcents


async def charge_budget(
db: AsyncSession, api_key_id: str, cap_microcents: int, actual_microcents: int
) -> bool:
"""Atomically record ``actual_microcents`` of spend, never exceeding ``cap``.

Returns ``True`` if the cost fit under the cap (the counter advanced by
``actual``), or ``False`` if the request alone would have breached the cap —
in which case the counter is clamped to ``cap`` so the key is maxed out and
blocked going forward. The boundary request may already have been served
upstream; it cannot be un-spent, but we never record more than the cap and we
stop the next one. Fail-closed.
"""
actual = actual_microcents or 0
result = await db.execute(
update(ApiKey)
.where(ApiKey.id == api_key_id, ApiKey.spent_microcents + actual <= cap_microcents)
.values(spent_microcents=ApiKey.spent_microcents + actual)
)
if result.rowcount:
await db.commit()
return True
# Would have exceeded the cap: clamp so the counter never overshoots and the
# key is correctly reported as exhausted thereafter.
await db.execute(
update(ApiKey)
.where(ApiKey.id == api_key_id, ApiKey.spent_microcents < cap_microcents)
.values(spent_microcents=cap_microcents)
)
await db.commit()
return False
13 changes: 11 additions & 2 deletions packages/db/models/api_key.py
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

from datetime import datetime

from sqlalchemy import JSON, Boolean, DateTime, ForeignKey, Integer, String
from sqlalchemy import JSON, BigInteger, Boolean, DateTime, ForeignKey, String
from sqlalchemy.orm import Mapped, mapped_column

from packages.db.models.base import Base, SoftDeleteMixin, TimestampMixin, UUIDMixin
Expand All @@ -18,7 +18,16 @@ class ApiKey(Base, UUIDMixin, TimestampMixin, SoftDeleteMixin):
key_hash: Mapped[str] = mapped_column(String(64), unique=True, nullable=False)
key_prefix: Mapped[str] = mapped_column(String(20), nullable=False)
model_allowlist: Mapped[list[str] | None] = mapped_column(JSON, nullable=True)
budget_limit_cents: Mapped[int | None] = mapped_column(Integer, nullable=True)
# BIGINT (not Integer): a client-supplied value up to the microcent scale
# can exceed a 32-bit int4 on Postgres, which would otherwise 500 on insert.
budget_limit_cents: Mapped[int | None] = mapped_column(BigInteger, nullable=True)
# Running lifetime spend in microcents. Maintained transactionally by
# spend.charge_budget: a single atomic UPDATE adds the actual cost and
# refuses to let the counter exceed budget_limit_cents, so the cap holds
# even under concurrent requests for the same key.
spent_microcents: Mapped[int] = mapped_column(
BigInteger, nullable=False, server_default="0", default=0
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 P1 Add a migration/backfill for spent_microcents — existing deployments get no column and 503 on every request

The new api_keys.spent_microcents column (and the budget_limit_cents INTEGER→BIGINT widening, plus the new ix_requests_log_api_key_spend index) exist only in the SQLAlchemy metadata. The only schema mechanism in the repo is Base.metadata.create_all in app/main.py:70, which creates missing tables but never alters existing ones. Deployments that already ran a release (docker-compose named volume lite-data, fly.io volume, Postgres via DATABASE_URL) keep an api_keys table without the column. After upgrading, validate_api_key (app/middleware/auth.py:211, unchanged) does select(ApiKey), and the ORM entity now selects every mapped column including spent_microcents → "no such column: api_keys.spent_microcents" on every authenticated request; the middleware's except Exception then answers 503 "Service temporarily unavailable" for the whole API. Budgeted keys additionally hit the missing column in read_spent/charge_budget. Even after the column is added manually, the counter is not backfilled from requests_log.cost_microcents, so existing keys' lifetime caps silently reset to zero and a leaked key that already spent its cap gets a second full budget. The change should have shipped an idempotent startup ALTER (SQLite/Postgres) adding the column and index, and seeded spent_microcents from historical requests_log spend for existing keys.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 446557a. Added packages/db/migrate.py:ensure_budget_columns, invoked after create_all at boot. It idempotently ALTERs spent_microcents into api_keys (seeded from historical requests_log.cost_microcents) and widens budget_limit_cents to BIGINT on Postgres. Existing deployments no longer 503 on every authenticated request.

is_active: Mapped[bool] = mapped_column(Boolean, server_default="true")
last_used_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True), nullable=True)
revoked_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True), nullable=True)
1 change: 1 addition & 0 deletions packages/db/models/request_log.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,7 @@ class RequestLog(Base, UUIDMixin, SoftDeleteMixin):
__tablename__ = "requests_log"
__table_args__ = (
Index("ix_requests_log_ws_created", "workspace_id", "created_at"),
Index("ix_requests_log_api_key_spend", "api_key_id", "is_deleted"),
)

workspace_id: Mapped[str] = mapped_column(String(36), nullable=False, index=True)
Expand Down
Loading
Loading