What Happens When Your Database Bills by the Read?
Firestore charges per document read and per write – a predictable cost model under normal use, but a punishing one when things go wrong.
A single graph traversal in the InPromptOut app could trigger several hundred reads, and a bug in an edge cascade could loop reads or writes indefinitely. Normal usage costs I had calculated early on – part of weighing Firestore against moving straight to a graph DB. A batching overhaul in late 2025 was the first signal that read volume mattered, but I hadn't modelled error-cost blowups – runaway loops, failure-path cascades – until this phase. An oversight, but a useful learning moment: the MVP had no read/write enforcement layer, which clearly needed to change.
This wasn't triggered by a bill shock or a production incident. It was defensive engineering – the healthy kind of paranoia. The safety net had to live in the infrastructure itself.
By the end of the phase, every Firestore read and write in the API flowed through a metered wrapper. Included were per-request budgets, per-user hourly limits and test-enforced guards that fail the suite if any code path bypasses them. The first meter caught a real (non-trivial) infinite-loop bug within days of going in – proof the work wasn't premature.
Why Firestore, Not a Graph DB?
A graph database was always the eventual target. As the data model is a graph, a native graph store would be a more natural fit long-term. Using Firestore for the MVP was a deliberate trade-off.
React was initially new to me on this project, stacked on top of a Python backend. Adding a new query language (Cypher, Gremlin, AQL) on top of that was a step too far for a solo build where time was limited. I had prior Firestore experience, the free tier suited solo-dev economics and Firebase Auth integrated cleanly.
The API was architected with strict separation of concerns from day one, so the data layer can be swapped to a graph store later without touching the service or domain layers. The accepted trade-off: Firestore bills per read with graph traversals getting expensive fast. This phase is where that trade-off got paid down.
The Read Meter (14 Feb)
The first circuit breaker. Every Firestore read operation – get(),
stream(), aggregation queries – was routed through metered wrappers
that incremented a per-request counter. If a request exceeded its budget, the
middleware returned 429 Too Many Requests.
* Budgets are tuned against observed ingress/egress patterns and shift with real traffic data.
It took three commits on the same day to get right. The first issue: broad
except Exception handlers throughout the codebase were swallowing
ReadMeterExceeded – the meter was firing but requests kept going.
The second: Starlette's BaseHTTPMiddleware runs the handler in a
child context, so plain ContextVar value mutations don't propagate
back to the parent middleware. The counter would increment inside the handler but
read as zero when the middleware checked it.
The Starlette ContextVar Gotcha
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
class _MeterState:
"""Stored in a ContextVar so child contexts share the same object reference.
Mutations to count/limit are visible across the context boundary."""
__slots__ = ("count", "limit")
def __init__(self, count: int = 0, limit: int = 0):
self.count = count
self.limit = limit
class Meter:
def __init__(self, name: str, label: str, exceeded_exc: Type[MeterExceeded]):
self._state: ContextVar[Optional[_MeterState]] = ContextVar(name, default=None)
self._exceeded_exc = exceeded_exc
def increment(self, count: int = 1) -> None:
if count < 1:
return
state = self._state.get()
if state is None:
return # No meter active (bypass path)
state.count += count
if state.limit > 0 and state.count > state.limit:
raise self._exceeded_exc(state.count, state.limit)
The fix: store a mutable _MeterState object in a single ContextVar
instead of two integer ContextVars. The child context gets the same object reference,
so state.count += 1 is visible to the parent. Both the read and write
meters inherit from this base class.
Metered Wrappers
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
def metered_get(doc_ref, *, transaction=None, **kwargs):
"""Fetch a single document. Increments read meter by 1."""
with allow_raw_reads():
result = doc_ref.get(transaction=transaction, **kwargs)
increment(1)
return result
def metered_query(query):
"""Execute a query. Increments by 1 per document, raising
ReadMeterExceeded mid-stream if the limit is hit."""
docs = []
with allow_raw_reads():
for doc in query.stream():
increment(1)
docs.append(doc)
return docs
Every raw .get() and list(q.stream()) call across the
entire codebase was replaced with these wrappers – 10 call sites in the mutations
service alone, plus every CRUD repo, the edge repo and the registry repo. By the
end, only two raw reads remained: the metered wrapper itself and an admin-only
aggregation query (annotated and accepted).
The First Catch: DeleteNodeOp Infinite Loop (16 Feb)
Within days of the read meter going live, it caught the bug that validated the entire phase. Deleting a node required cleaning up its edges, but the edge cleanup loop re-queried Firestore before the queued deletes were committed. If fewer than the batch threshold (450) were queued, the same edges were read indefinitely – spinning until the read meter tripped. Without the meter, the loop would have run until it exhausted the request budget on Firestore's side, not mine.
The fix was small: commit or apply deletes before the next query iteration. The lesson was larger. An egress audit followed – expected reads compared against actual reads in the Firestore console, with every mismatch traced back to code.
- P1: DeleteNodeOp infinite loop fixed (commit before re-query)
- P0:
.limit()added to all unbounded edge queries – some had been running without limits since prototyping - P0: OpenAI API calls wrapped with
asyncio.wait_fortimeout – a hanging call could hold a request (and its read budget) open indefinitely - P0: Graph traversal capped at max depth and max node count – prototype code that should have been future-proofed earlier
- P2: Graph traversal truncation flag surfaced to the frontend
AI was used to sweep the codebase for similar patterns after each finding – the same re-query-before-commit bug could have appeared anywhere a loop combined reads with batched writes.
The Write Meter (11 Mar)
Read protection alone was not enough. Writes cost more than reads and write failures
are harder to clean up. The mutation queue also added significant complexity – writes are batched
before a single batch.commit(). The write meter needed to count at
queue time, not commit time, otherwise a runaway loop could queue hundreds of writes
before a single commit and remain invisible to the meter.
Reserve-Before-Write with Refund-on-Failure
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
def metered_set(doc_ref, data, **kwargs):
"""Write a single document. Reserves budget (+1), refunds on failure."""
increment(1)
try:
with allow_raw_writes():
return doc_ref.set(data, **kwargs)
except Exception:
decrement(1)
raise
def metered_batch_commit(batch, count):
"""Commit a batch write. Reserves budget first, refunds on failure."""
if count < 1:
raise ValueError("metered_batch_commit count must be >= 1")
increment(count)
try:
return batch.commit()
except Exception:
decrement(count)
raise
Every standalone Firestore write was routed through metered wrappers. Batch writes
via MutationContext were metered at commit time through
metered_batch_commit():
1
2
3
4
5
6
7
8
9
10
11
12
13
def increment_writes(self, count: int = 1) -> None:
"""Increment queued write counter and auto-flush at threshold."""
self.writes += count
if self.writes >= AUTO_FLUSH_WRITE_THRESHOLD:
self.commit_and_reset()
def commit_and_reset(self) -> None:
"""Commit current batch and open a fresh one."""
if self.writes:
metered_batch_commit(self.batch, self.writes)
self.batch = self.db.batch()
self.writes = 0
DRY Refactor: The Meter Base Class
The read and write meters started as separate implementations. When the write meter
landed, the duplication was obvious – same ContextVar pattern, same increment/decrement
logic, same exception structure – and they collapsed cleanly into a shared
Meter base class at ~24 lines per concrete meter. Coding at pace with
AI means duplication sometimes lands before the abstraction does; worth staying
conscious of, but not a problem when the cleanup is part of the workflow.
Exception Propagation Fix (Phase 5B)
A subtle but critical cross-cutting fix. Every except ReadMeterExceeded: raise
clause in the codebase only re-raised read meter exceptions. After introducing
WriteMeterExceeded, any broad except Exception handler
following those clauses would silently swallow write meter exceptions – defeating
the circuit breaker. The fix: replace all specific catches with
except MeterExceeded: raise using the shared base class, across every
repo, service and router layer.
Test Guards: Enforcing the Contract
The metered wrappers were a convention – every Firestore call should go
through them. The test guard system turned "should" into "must" by monkey-patching
the Firestore SDK in the test suite. Any DocumentReference.set(),
.delete(), .update(), or .get() call that
bypassed the metered wrappers would throw an AssertionError.
This pattern – using tests to enforce architectural constraints, not just behaviour
– evolved from a similar approach in
MetaMaker, an earlier
Python desktop project. There, an enforce_contract decorator validated
signal/event payload types at runtime, catching integration violations that unit
tests alone would miss. The Firestore guard was the same instinct applied to a
different problem: if a future contributor adds a raw .set() call,
tests fail automatically.
When the write guard was first switched on, it caught real violations – test fixtures and setup paths that weren't going through metered wrappers. Exactly what it was built to do.
Per-User Rate Limiting (1–14 Mar)
The meters capped individual requests. But a user sending thousands of small, perfectly-metered requests could still accumulate significant cost. The per-user rate limiter closed this gap with separate read and write buckets per authenticated user.
The suggestion endpoint came first – not because OpenAI costs were the primary
concern (that billing is more controllable), but because rate limiting was needed
and the limit parameter was designed to accept per-account values,
preparing for potential account tiers.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
_lock = threading.Lock()
_user_timestamps: dict[str, list[float]] = defaultdict(list)
class UserRequestRateLimitMiddleware(BaseHTTPMiddleware):
async def dispatch(self, request, call_next):
# ... auth + path checks ...
if request.method == "GET" or is_exempt_post:
bucket = "read"
limit = USER_READ_REQUEST_RATE_LIMIT_PER_HOUR
else:
bucket = "write"
limit = USER_REQUEST_RATE_LIMIT_PER_HOUR
now = time.monotonic()
bucket_key = f"{user_id}:{bucket}"
with _lock:
timestamps = _user_timestamps[bucket_key]
_user_timestamps[bucket_key] = [
ts for ts in timestamps if ts > window_start
]
if len(_user_timestamps[bucket_key]) >= limit:
return JSONResponse(status_code=429, content={...})
_user_timestamps[bucket_key].append(now)
The in-memory approach – threading.Lock, time.monotonic,
a dict of timestamps – was a deliberate KISS trade-off. It won't survive
multi-instance deployment (each Cloud Run instance has its own memory), but it
avoids adding more Firestore reads to a system built specifically to limit Firestore
reads. A lightweight SQL sidecar or Redis instance would be the natural evolution
if the project scales to multiple instances.
The Mutations God Function (27 Feb)
Buried in the metering work was a significant refactor: extracting the mutations
service "God function" into a handler dispatch package. The original
apply_mutations() was prototype code – an inline if/elif
chain written to prove the batch-mutation flow end-to-end when only three op
types existed. It held up while the surface stayed small, but once the
DeleteNodeOp infinite-loop bug exposed how much shared mutable state
the closure was carrying, the shape had outlived its purpose. The write meter needed
clean entry points regardless – extracting each operation into its own handler
module solved both problems in one pass.
flowchart TB
A1[apply_mutations] --> B1["for each op:
if CreateNodeOp: …
elif CreateEdgeOp: …
elif DeleteNodeOp: …
(would have grown to 6 op types)"]
B1 --> C1[(Firestore batch)]
classDef entry fill:#2a2e3a,stroke:#88a,color:#eee,stroke-width:1px
classDef monolith fill:#3a2a2a,stroke:#c77,color:#eee,stroke-width:1px
classDef store fill:#1f1f1f,stroke:#888,color:#ccc,stroke-width:1px
class A1 entry
class B1 monolith
class C1 store
flowchart LR
A2[apply_mutations] --> D{match op}
D --> H1[handle_create_node]
D --> H2[handle_create_edge]
D --> H3[handle_delete_node]
D --> H4[handle_delete_edge]
D --> H5[handle_move_node]
D --> H6[handle_update_node]
H1 & H2 & H3 & H4 & H5 & H6 --> C2[(MutationContext
+ metered batch)]
classDef entry fill:#2a2e3a,stroke:#88a,color:#eee,stroke-width:1px
classDef dispatch fill:#3a3320,stroke:#cc8,color:#eee,stroke-width:1px
classDef handler fill:#223a2d,stroke:#5a9,color:#eee,stroke-width:1px
classDef store fill:#1f1f1f,stroke:#888,color:#ccc,stroke-width:1px
class A2 entry
class D dispatch
class H1,H2,H3,H4,H5,H6 handler
class C2 store
The refactor made the write meter wiring straightforward. Each handler became a
clean entry point that could reserve write budget before queuing while the dispatch
function in apply_mutations stayed minimal. It also made the handlers
independently testable – each operation type could be exercised without standing up
the full mutations pipeline.
The Middleware Stack
By mid-March, the API had a layered protection model. Each piece was always planned individually – kill switch, metering, rate limiting – but the stack assembled itself as the specifics of Firestore's billing model became clearer. The ordering was deliberate (and had to be fixed once after a likely merge regression):
1
2
3
4
5
6
7
8
app.add_middleware(WriteMeterMiddleware) # 7. Per-request write caps
app.add_middleware(ReadMeterMiddleware) # 6. Per-request read caps
app.add_middleware(UserRequestRateLimitMiddleware)# 5. Per-user hourly budgets
app.add_middleware(AuthMiddleware) # 4. JWT verification
app.add_middleware(KillSwitchMiddleware) # 3. Emergency stop (zero external calls)
app.add_middleware(RequestLoggingMiddleware) # 2. Audit trail
app.add_middleware(CloudflareOnlyMiddleware) # 1. Bot protection
Kill switch before auth means zero external calls when active – not even Firebase token verification. Maximum cost protection during an incident. The ordering matters operationally: a cost-control stack that runs after authentication is still exposing the auth surface under attack; running it earlier means a bad actor hits a 429 before costing anything.