From "It Works" to "It's Safe"
The MVP was feature-complete, but AI-assisted development had accelerated everything. Every line was reviewed manually, yet the velocity was higher than working without AI tools, which gave me pause. I was also still getting comfortable with React (new to me on this project), so a paranoid development style pushed me to audit the backend before shipping. The API is where the buck stops: if the frontend has a weakness, the backend needs to catch it.
Rather than fixing issues as they surfaced, the approach was structured: multiple AI models were run through the codebase, each given specific vulnerability checklists. Findings were triaged over about a week – not everything flagged was real – and organised into a priority list from P0 to P6.
The Defence-in-Depth Fix: Client-Supplied User IDs (8 Feb)
The most instructive finding was a defence-in-depth gap: mutation request bodies still
accepted a userId field from the client. The mutations service already
overwrote it with the authenticated identity from the JWT before any Firestore write,
so no live exploit existed – but the redundant field was an anti-pattern. If the
server-side injection were ever removed or bypassed, client-supplied values would
flow straight through to Firestore, enabling impersonation.
Before: Frontend Sent userId in Every Operation
1
2
3
4
5
6
// userId was included in every mutation operation
export type OpBase = {
opId: string;
userId: string; // ❌ Client-supplied – server should own this
ts: number;
};
After: Server Injects Identity
1
2
3
4
export type OpBase = {
opId: string; // unique idempotency key for this op
ts: number; // epoch ms – userId removed entirely
};
1
2
3
4
5
6
7
8
9
10
11
12
13
14
def apply_mutations(req: ApplyMutationsRequest, user_id: str) -> BatchApplyResponse:
"""Apply a batch of mutations atomically (injecting authenticated user_id)."""
db = get_firestore_client()
context = MutationContext(db=db, batch=db.batch(), writes=0, temp_id_map={})
for op in req.operations:
# SECURITY: Inject authenticated user_id, don't trust request body
op.user_id = user_id
match op:
case CreateNodeOp():
results.append(handle_create_node(op, context))
case CreateEdgeOp():
results.append(handle_create_edge(op, context))
# ... remaining handlers
Pydantic strict mode – extra="forbid" on the mutation schemas – was
already in place, which made the fix clean: once userId was stripped
from the request schemas, any client still sending it would be rejected with a 422
at the validation boundary. Both repos were updated in the same commit window.
Test Suite Realignment (10–11 Feb)
The schema changes broke a number of existing test fixtures. Tests had constructed
request objects directly with a user_id field, so once it was stripped
from the public request schemas, those fixtures had to switch to the new
EdgeCreateSchema shape. Separately, ownership checks added at the
repository update layer (finding #4 below) required a pass through repo and service
tests to supply the verified user_id explicitly.
At the same time, the test infrastructure moved from mocked Firestore clients to the Firestore emulator for real integration testing.
Security Test: Anti-Enumeration Pattern
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
TEST_USER = "user-owner"
OTHER_USER = "user-attacker"
def test_cross_user_404_prevents_enumeration(self, mock_get):
"""404 response for both 'not found' and 'wrong user' prevents enumeration."""
# Case 1: Node doesn't exist
mock_get.return_value = None
with pytest.raises(HTTPException) as not_found:
get_owned_registry_entry("nonexistent", TEST_USER)
# Case 2: Node exists but belongs to another user
entry = _make_registry_entry("node-1", OTHER_USER)
mock_get.return_value = entry
with pytest.raises(HTTPException) as wrong_user:
get_owned_registry_entry("node-1", TEST_USER)
# Both must return identical status and detail (prevents enumeration)
assert not_found.value.status_code == wrong_user.value.status_code == 404
assert not_found.value.detail == wrong_user.value.detail == "Node not found"
Test Priority Tiers
The security-first test suite was planned as a layered risk map: each tier rated how exposed the system would be if that layer failed, from the front door inward. This ordering drove which layers got coverage first.
1
2
3
4
5
6
7
8
9
"""P0 – Kill Switch Middleware Tests""" # Front-door emergency stop
"""P0 – Auth Middleware Security Tests""" # Token verification, 401/503 responses
"""P1 – Request Logging Middleware Tests""" # Audit trail
"""P2 – Public Endpoint Exposure Tests""" # Unauthenticated surface area
"""P3 – Authorization Service Tests""" # Ownership verification
"""P3.5 – Mutations Service Security Tests""" # Batch operation security
"""P4 – Repo Layer user_id Scoping Tests""" # Defence-in-depth at data layer
"""P5 – Auth Dependency Injection Tests""" # FastAPI DI chain
"""P6 – Cross-Layer Integration Tests""" # End-to-end auth flow
Documented Findings
The audit added 180+ security-focused tests across 9 priority tiers – a mix of new
test files for middleware, authorization and repo-layer scoping, plus substantial
extensions to existing service and repo test suites. Out of that sweep, 4 actual bugs
were caught – each tracked in a SECURITY_TEST_FINDINGS.md file with
severity, reproduction steps, affected code paths and test coverage references.
This sat alongside the project's broader documentation practices: architecture
decision records (ADRs), engineering journals and audit logs maintained in a
dedicated inpromptout-doc repo.
| # | Severity | Finding | Status |
|---|---|---|---|
| 1 | Critical | Auth middleware returning 500 instead of 401/503 (Starlette BaseHTTPMiddleware gotcha) |
Fixed |
| 2 | Medium | Broken test fixtures importing non-existent auth service module | Fixed |
| 3 | Medium | get_all_edges_from_db() had no user_id filter (potential data leak) |
Removed |
| 4 | Low | Repo update functions lacked ownership check (TOCTOU gap – low severity because the mutation pipeline, which enforced ownership, was the primary write path for all UI interactions) | Fixed |
Finding 1: The Starlette Middleware Gotcha
The most instructive finding was a framework-level issue. HTTPException
raised inside Starlette's BaseHTTPMiddleware.dispatch() is not caught by
FastAPI's exception handler – it propagates to ServerErrorMiddleware, which
converts it to a generic 500 Internal Server Error. Every auth failure was
returning 500 instead of 401.
1
2
3
4
5
class AuthMiddleware(BaseHTTPMiddleware):
async def dispatch(self, request, call_next):
# ❌ Starlette's BaseHTTPMiddleware swallows this
raise HTTPException(status_code=401, detail="Missing authentication token")
# Client receives: 500 Internal Server Error (not 401)
1
2
3
4
5
6
7
8
class AuthMiddleware(BaseHTTPMiddleware):
async def dispatch(self, request, call_next):
# ✅ Returns proper HTTP response, bypasses exception handling
return JSONResponse(
status_code=401,
content={"detail": "Missing authentication token"}
)
# Client receives: 401 with JSON error detail
The Invite System (1 Feb)
InPromptOut launched as invite-only – a controlled rollout while the metering system was still being built. Email validation with normalisation, wired up across both repos. An admin dashboard followed (7 Feb) for managing invites.
The Kill Switch (12 Feb)
A Firestore-backed configuration document, cached in-process with a 30-second TTL,
that can reject all traffic with a 503. Runs before authentication in
the middleware stack – when active, zero external calls are made.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
KILL_SWITCH_BYPASS_PATHS = {"/", "/health"}
class KillSwitchMiddleware(BaseHTTPMiddleware):
async def dispatch(self, request: Request, call_next):
path = request.url.path
normalized_path = path if path == "/" else path.rstrip("/")
if request.method == "OPTIONS" or normalized_path in KILL_SWITCH_BYPASS_PATHS:
return await call_next(request)
enabled, message = await get_kill_switch_state()
if enabled:
return JSONResponse(
status_code=503,
content={"detail": message},
)
return await call_next(request)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
CACHE_TTL_SECONDS = 30.0
_cached_enabled: bool | None = None
_cached_at: float = 0.0
async def get_kill_switch_state() -> Tuple[bool, str]:
global _cached_enabled, _cached_message, _cached_at
now = time.monotonic()
if _cached_enabled is not None and (now - _cached_at) < CACHE_TTL_SECONDS:
return _cached_enabled, _cached_message
try:
enabled, message = await _fetch_kill_switch_state()
_cached_enabled = enabled
_cached_at = time.monotonic()
return enabled, message
except Exception:
if _cached_enabled is not None:
return _cached_enabled, _cached_message # Use stale cache
# No cache available – fail closed
return True, DEFAULT_KILL_SWITCH_MESSAGE
The initial implementation was fail-open. That lasted about an hour before being changed to fail-closed. The reasoning: paranoia and a very low budget. If the system can't even read its own config, the safest default is to stop.