lundie.io Get In Touch

Phase 11: Metering & Cost Safety

February – March 2026

What Happens When Your Database Bills by the Read?

Firestore charges per document read and per write – a predictable cost model under normal use, but a punishing one when things go wrong.

A single graph traversal in the InPromptOut app could trigger several hundred reads, and a bug in an edge cascade could loop reads or writes indefinitely. Normal usage costs I had calculated early on – part of weighing Firestore against moving straight to a graph DB. A batching overhaul in late 2025 was the first signal that read volume mattered, but I hadn't modelled error-cost blowups – runaway loops, failure-path cascades – until this phase. An oversight, but a useful learning moment: the MVP had no read/write enforcement layer, which clearly needed to change.

This wasn't triggered by a bill shock or a production incident. It was defensive engineering – the healthy kind of paranoia. The safety net had to live in the infrastructure itself.

By the end of the phase, every Firestore read and write in the API flowed through a metered wrapper. Included were per-request budgets, per-user hourly limits and test-enforced guards that fail the suite if any code path bypasses them. The first meter caught a real (non-trivial) infinite-loop bug within days of going in – proof the work wasn't premature.

Why Firestore, Not a Graph DB?

A graph database was always the eventual target. As the data model is a graph, a native graph store would be a more natural fit long-term. Using Firestore for the MVP was a deliberate trade-off.

React was initially new to me on this project, stacked on top of a Python backend. Adding a new query language (Cypher, Gremlin, AQL) on top of that was a step too far for a solo build where time was limited. I had prior Firestore experience, the free tier suited solo-dev economics and Firebase Auth integrated cleanly.

The API was architected with strict separation of concerns from day one, so the data layer can be swapped to a graph store later without touching the service or domain layers. The accepted trade-off: Firestore bills per read with graph traversals getting expensive fast. This phase is where that trade-off got paid down.

The Read Meter (14 Feb)

The first circuit breaker. Every Firestore read operation – get(), stream(), aggregation queries – was routed through metered wrappers that incremented a per-request counter. If a request exceeded its budget, the middleware returned 429 Too Many Requests.

Default Budget
200 reads*
Graph Budget
1,000 reads*

* Budgets are tuned against observed ingress/egress patterns and shift with real traffic data.

It took three commits on the same day to get right. The first issue: broad except Exception handlers throughout the codebase were swallowing ReadMeterExceeded – the meter was firing but requests kept going. The second: Starlette's BaseHTTPMiddleware runs the handler in a child context, so plain ContextVar value mutations don't propagate back to the parent middleware. The counter would increment inside the handler but read as zero when the middleware checked it.

The Starlette ContextVar Gotcha

Meter Core – Mutable State in a ContextVar
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
class _MeterState:
    """Stored in a ContextVar so child contexts share the same object reference.
    Mutations to count/limit are visible across the context boundary."""
    __slots__ = ("count", "limit")

    def __init__(self, count: int = 0, limit: int = 0):
        self.count = count
        self.limit = limit

class Meter:
    def __init__(self, name: str, label: str, exceeded_exc: Type[MeterExceeded]):
        self._state: ContextVar[Optional[_MeterState]] = ContextVar(name, default=None)
        self._exceeded_exc = exceeded_exc

    def increment(self, count: int = 1) -> None:
        if count < 1:
            return
        state = self._state.get()
        if state is None:
            return  # No meter active (bypass path)
        state.count += count
        if state.limit > 0 and state.count > state.limit:
            raise self._exceeded_exc(state.count, state.limit)
    

The fix: store a mutable _MeterState object in a single ContextVar instead of two integer ContextVars. The child context gets the same object reference, so state.count += 1 is visible to the parent. Both the read and write meters inherit from this base class.

Metered Wrappers

firestore_reads.py – Every Read Counted
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
def metered_get(doc_ref, *, transaction=None, **kwargs):
    """Fetch a single document. Increments read meter by 1."""
    with allow_raw_reads():
        result = doc_ref.get(transaction=transaction, **kwargs)
    increment(1)
    return result

def metered_query(query):
    """Execute a query. Increments by 1 per document, raising
    ReadMeterExceeded mid-stream if the limit is hit."""
    docs = []
    with allow_raw_reads():
        for doc in query.stream():
            increment(1)
            docs.append(doc)
    return docs
    

Every raw .get() and list(q.stream()) call across the entire codebase was replaced with these wrappers – 10 call sites in the mutations service alone, plus every CRUD repo, the edge repo and the registry repo. By the end, only two raw reads remained: the metered wrapper itself and an admin-only aggregation query (annotated and accepted).

The First Catch: DeleteNodeOp Infinite Loop (16 Feb)

Within days of the read meter going live, it caught the bug that validated the entire phase. Deleting a node required cleaning up its edges, but the edge cleanup loop re-queried Firestore before the queued deletes were committed. If fewer than the batch threshold (450) were queued, the same edges were read indefinitely – spinning until the read meter tripped. Without the meter, the loop would have run until it exhausted the request budget on Firestore's side, not mine.

The fix was small: commit or apply deletes before the next query iteration. The lesson was larger. An egress audit followed – expected reads compared against actual reads in the Firestore console, with every mismatch traced back to code.

  • P1: DeleteNodeOp infinite loop fixed (commit before re-query)
  • P0: .limit() added to all unbounded edge queries – some had been running without limits since prototyping
  • P0: OpenAI API calls wrapped with asyncio.wait_for timeout – a hanging call could hold a request (and its read budget) open indefinitely
  • P0: Graph traversal capped at max depth and max node count – prototype code that should have been future-proofed earlier
  • P2: Graph traversal truncation flag surfaced to the frontend

AI was used to sweep the codebase for similar patterns after each finding – the same re-query-before-commit bug could have appeared anywhere a loop combined reads with batched writes.

The Write Meter (11 Mar)

Read protection alone was not enough. Writes cost more than reads and write failures are harder to clean up. The mutation queue also added significant complexity – writes are batched before a single batch.commit(). The write meter needed to count at queue time, not commit time, otherwise a runaway loop could queue hundreds of writes before a single commit and remain invisible to the meter.

Reserve-Before-Write with Refund-on-Failure

firestore_writes.py – Write Wrappers
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
def metered_set(doc_ref, data, **kwargs):
    """Write a single document. Reserves budget (+1), refunds on failure."""
    increment(1)
    try:
        with allow_raw_writes():
            return doc_ref.set(data, **kwargs)
    except Exception:
        decrement(1)
        raise

def metered_batch_commit(batch, count):
    """Commit a batch write. Reserves budget first, refunds on failure."""
    if count < 1:
        raise ValueError("metered_batch_commit count must be >= 1")
    increment(count)
    try:
        return batch.commit()
    except Exception:
        decrement(count)
        raise
    

Every standalone Firestore write was routed through metered wrappers. Batch writes via MutationContext were metered at commit time through metered_batch_commit():

MutationContext – Commit-Time Metering
1
2
3
4
5
6
7
8
9
10
11
12
13
def increment_writes(self, count: int = 1) -> None:
    """Increment queued write counter and auto-flush at threshold."""
    self.writes += count
    if self.writes >= AUTO_FLUSH_WRITE_THRESHOLD:
        self.commit_and_reset()

def commit_and_reset(self) -> None:
    """Commit current batch and open a fresh one."""
    if self.writes:
        metered_batch_commit(self.batch, self.writes)
    self.batch = self.db.batch()
    self.writes = 0
    

DRY Refactor: The Meter Base Class

The read and write meters started as separate implementations. When the write meter landed, the duplication was obvious – same ContextVar pattern, same increment/decrement logic, same exception structure – and they collapsed cleanly into a shared Meter base class at ~24 lines per concrete meter. Coding at pace with AI means duplication sometimes lands before the abstraction does; worth staying conscious of, but not a problem when the cleanup is part of the workflow.

Exception Propagation Fix (Phase 5B)

A subtle but critical cross-cutting fix. Every except ReadMeterExceeded: raise clause in the codebase only re-raised read meter exceptions. After introducing WriteMeterExceeded, any broad except Exception handler following those clauses would silently swallow write meter exceptions – defeating the circuit breaker. The fix: replace all specific catches with except MeterExceeded: raise using the shared base class, across every repo, service and router layer.

Test Guards: Enforcing the Contract

The metered wrappers were a convention – every Firestore call should go through them. The test guard system turned "should" into "must" by monkey-patching the Firestore SDK in the test suite. Any DocumentReference.set(), .delete(), .update(), or .get() call that bypassed the metered wrappers would throw an AssertionError.

This pattern – using tests to enforce architectural constraints, not just behaviour – evolved from a similar approach in MetaMaker, an earlier Python desktop project. There, an enforce_contract decorator validated signal/event payload types at runtime, catching integration violations that unit tests alone would miss. The Firestore guard was the same instinct applied to a different problem: if a future contributor adds a raw .set() call, tests fail automatically.

When the write guard was first switched on, it caught real violations – test fixtures and setup paths that weren't going through metered wrappers. Exactly what it was built to do.

Per-User Rate Limiting (1–14 Mar)

The meters capped individual requests. But a user sending thousands of small, perfectly-metered requests could still accumulate significant cost. The per-user rate limiter closed this gap with separate read and write buckets per authenticated user.

The suggestion endpoint came first – not because OpenAI costs were the primary concern (that billing is more controllable), but because rate limiting was needed and the limit parameter was designed to accept per-account values, preparing for potential account tiers.

Per-User Rate Limiter – In-Memory Sliding Window
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
_lock = threading.Lock()
_user_timestamps: dict[str, list[float]] = defaultdict(list)

class UserRequestRateLimitMiddleware(BaseHTTPMiddleware):
    async def dispatch(self, request, call_next):
        # ... auth + path checks ...

        if request.method == "GET" or is_exempt_post:
            bucket = "read"
            limit = USER_READ_REQUEST_RATE_LIMIT_PER_HOUR
        else:
            bucket = "write"
            limit = USER_REQUEST_RATE_LIMIT_PER_HOUR

        now = time.monotonic()
        bucket_key = f"{user_id}:{bucket}"

        with _lock:
            timestamps = _user_timestamps[bucket_key]
            _user_timestamps[bucket_key] = [
                ts for ts in timestamps if ts > window_start
            ]
            if len(_user_timestamps[bucket_key]) >= limit:
                return JSONResponse(status_code=429, content={...})
            _user_timestamps[bucket_key].append(now)
    

The in-memory approach – threading.Lock, time.monotonic, a dict of timestamps – was a deliberate KISS trade-off. It won't survive multi-instance deployment (each Cloud Run instance has its own memory), but it avoids adding more Firestore reads to a system built specifically to limit Firestore reads. A lightweight SQL sidecar or Redis instance would be the natural evolution if the project scales to multiple instances.

The Mutations God Function (27 Feb)

Buried in the metering work was a significant refactor: extracting the mutations service "God function" into a handler dispatch package. The original apply_mutations() was prototype code – an inline if/elif chain written to prove the batch-mutation flow end-to-end when only three op types existed. It held up while the surface stayed small, but once the DeleteNodeOp infinite-loop bug exposed how much shared mutable state the closure was carrying, the shape had outlived its purpose. The write meter needed clean entry points regardless – extracting each operation into its own handler module solved both problems in one pass.

Before · Inline Switch
flowchart TB
            A1[apply_mutations] --> B1["for each op:
if CreateNodeOp: …
elif CreateEdgeOp: …
elif DeleteNodeOp: …
(would have grown to 6 op types)"] B1 --> C1[(Firestore batch)] classDef entry fill:#2a2e3a,stroke:#88a,color:#eee,stroke-width:1px classDef monolith fill:#3a2a2a,stroke:#c77,color:#eee,stroke-width:1px classDef store fill:#1f1f1f,stroke:#888,color:#ccc,stroke-width:1px class A1 entry class B1 monolith class C1 store
After · Handler Dispatch
flowchart LR
            A2[apply_mutations] --> D{match op}
            D --> H1[handle_create_node]
            D --> H2[handle_create_edge]
            D --> H3[handle_delete_node]
            D --> H4[handle_delete_edge]
            D --> H5[handle_move_node]
            D --> H6[handle_update_node]
            H1 & H2 & H3 & H4 & H5 & H6 --> C2[(MutationContext
+ metered batch)] classDef entry fill:#2a2e3a,stroke:#88a,color:#eee,stroke-width:1px classDef dispatch fill:#3a3320,stroke:#cc8,color:#eee,stroke-width:1px classDef handler fill:#223a2d,stroke:#5a9,color:#eee,stroke-width:1px classDef store fill:#1f1f1f,stroke:#888,color:#ccc,stroke-width:1px class A2 entry class D dispatch class H1,H2,H3,H4,H5,H6 handler class C2 store

The refactor made the write meter wiring straightforward. Each handler became a clean entry point that could reserve write budget before queuing while the dispatch function in apply_mutations stayed minimal. It also made the handlers independently testable – each operation type could be exercised without standing up the full mutations pipeline.

The Middleware Stack

By mid-March, the API had a layered protection model. Each piece was always planned individually – kill switch, metering, rate limiting – but the stack assembled itself as the specifics of Firestore's billing model became clearer. The ordering was deliberate (and had to be fixed once after a likely merge regression):

main.py – Middleware Registration (Last-Added = Outermost)
1
2
3
4
5
6
7
8
app.add_middleware(WriteMeterMiddleware)          # 7. Per-request write caps
app.add_middleware(ReadMeterMiddleware)           # 6. Per-request read caps
app.add_middleware(UserRequestRateLimitMiddleware)# 5. Per-user hourly budgets
app.add_middleware(AuthMiddleware)                # 4. JWT verification
app.add_middleware(KillSwitchMiddleware)          # 3. Emergency stop (zero external calls)
app.add_middleware(RequestLoggingMiddleware)      # 2. Audit trail
app.add_middleware(CloudflareOnlyMiddleware)      # 1. Bot protection
    

Kill switch before auth means zero external calls when active – not even Firebase token verification. Maximum cost protection during an incident. The ordering matters operationally: a cost-control stack that runs after authentication is still exposing the auth surface under attack; running it earlier means a bad actor hits a 429 before costing anything.

Backend Commits
35+
Protection Layers
7
Metered Wrappers
9
Tests at Phase End
696

Key Commits from This Phase

API ab40033 2026-02-14
[Feat] Add per-request Firestore read meter (P0 circuit breaker)
API 7369597 2026-02-16
[Harden] Add .limit() to all unbounded edge queries and wire constants (P0)
API 864448d 2026-02-16
[Fix] Resolve DeleteNodeOp infinite loop caused by missing commit before re-query (P1)
API 8539927 2026-02-27
[Refactor] Extract mutations God function into handler dispatch package
API c8f6dda 2026-03-11
[Safety] Add per-request Firestore write meter (core, wrappers, middleware)
API f152d3b 2026-03-13
[Fix/Tests] Enforce write guards and fix read meter propagation
API 9d9165f 2026-03-14
[Feature][Agent] Add per-user request rate limiter middleware
API 6170a8d 2026-03-14
[Feature][Agent] Add per-user read burst protection to rate limiter

Get In Touch

Prefer using email? Say hi at hello@lundie.io