A control nobody has watched fail is not a control
Lessons from a signup-abuse incident on a platform we took over: the previous developer's abuse controls existed in code, passed review, and did nothing. What made them inert, how we fixed them, and the practices we kept.
A security control that has never been observed failing does not exist. It is a line of code that reads well. We relearned this recently on a platform we took over from its previous developer, when automated signups started creating accounts and sending emails to addresses nobody had verified. Every control meant to prevent that was present in the codebase we inherited. None of them worked, and the one that would have stopped most of it, the CAPTCHA, had been configured for a different subdomain from the one the signup form actually lived on.
This piece is about how the previous developer had built it, why each control was inert, how we fixed it, and the practices we kept afterwards. The details of the product are not the point; the pattern is, because it is the shape of most inherited systems we see.
What happened, in general terms
The platform has a public signup form. As built, submitting it created a durable account, a profile, and an email on the first request, before the address was verified. Bots found the form. Each hit left a real record in the database and sent real mail to an address the bot had chosen.
For the organization that owns the platform, this was not a data-hygiene problem. A trusted institution sending unsolicited email to harvested addresses is a reputation and data-protection problem. That framing decided the priority.
How the previous developer did it, and why each control was inert
Six separate things were true at once. Each looked fine on its own, and all of them had passed review before we arrived.
The CAPTCHA was pointed at the wrong subdomain. A bot challenge was integrated into the form and its keys were present in the configuration. They had been issued for a different hostname from the one the form was served on, so the challenge never verified anything on the live site. This was the single largest failure: one configuration value, never checked against production, removed the outermost defence entirely.
Account creation preceded verification. The form wrote permanent state first and asked for proof of ownership second. Anything that reaches the form can therefore create records.
The rate limits were real and did nothing. The application used a cache-backed rate limiter. In the environment where the limits were tested, the cache store was a null store, so every limit silently allowed everything. Code review saw four rate-limit calls and moved on. No test ever exercised them.
Behind a proxy, per-IP limits collapsed into one bucket. The application sat behind an edge network and a platform router with no trusted-proxy configuration. Every request resolved to the same upstream address, so a per-IP limit became a single global limit. IP is also the wrong identity for this product's users, many of whom sit behind one shared institutional address.
Validation lived in the browser. Profile completeness relied on HTML required attributes and a controller predicate. A direct request, or an administrator editing a record, could persist a "complete" profile with no name and no organization.
Lifecycle email had no eligibility gate. Welcome and follow-up emails keyed off the existence of an account rather than a verified, complete profile. Unverified accounts entered email journeys automatically.
None of these is exotic, and none of them was malicious. They are what happens when controls are designed, written, and reviewed, and never once watched doing their job in the environment where they matter.
The question that finds them
For every control, ask: what test would fail if this control were deleted? And for every control that depends on configuration: has anyone seen it reject something in production?
If the answer to either is "no", the control is unprotected. It may work today. It will not survive the next refactor, the next environment change, or the next person who does not know why it is there. Tests that assert wiring (the middleware is registered, the method is called) give false confidence. Tests that assert behaviour (the sixth request in a minute is rejected; a profile without a name cannot be saved through any path; the challenge fails on the real hostname without a valid token) are the only evidence a control exists.
We now treat "unauthenticated endpoint that can cause an email" as its own review category, with layered limits, a challenge verified server-side against the hostname it is served from, and a behavioural test for each layer.
How we fixed it
Every gate moved server-side, and every control became provable.
- The CAPTCHA keys were reissued for the hostname the form is actually served on, verification happens on the server, and a test asserts that a request without a valid token is rejected.
- Signup holds pending state only until the one-time code is verified. Nothing durable, and no email beyond the code itself, before that.
- Rate limits are tested with a real cache store, and are keyed on recipient and domain velocity and a global ceiling, not only on IP. The per-IP limit is documented as best-effort until the edge configuration is complete.
- Trusted proxies are configured explicitly, and the boot-time check that enforces them is achievable, not merely strict. A "fail closed" default that nobody can satisfy pushes people toward pasting broad ranges; safe defaults have to be reachable.
- Profile validation is enforced by the model, so the browser, the API, and the admin interface cannot disagree.
- Lifecycle email checks eligibility (verified, complete, correct role) at send time, not at signup.
Two things outside the code mattered as much.
Independent review found what a green build did not. Two read-only review passes over our own remediation surfaced a permanent lockout in the one-time-code path, an unbounded audit read, a document embed broken by a tightened content-security policy, and a cleanup selector that could have deactivated a legitimate user. The build was green throughout. Our fixes needed the same scrutiny as the code we inherited.
Real data disagreed with the fixtures. Running an import against the full third-party reference dataset, rather than the small fixture the previous developer had tested with, showed that public domain suffixes were being accepted as organizational evidence and that closed entities were being imported as active. Fixtures encode assumptions; the source encodes reality.
Cleaning up without making it worse
Removing the bot-created accounts was the most dangerous step, because the evidence for "this account was never real" initially lived in short-lived records that a retention job deletes daily. A legitimately verified user could have looked inactive.
The pattern we settled on is portable to any destructive maintenance task:
- Dry run by default, producing a reviewed digest.
- An exact list of record identifiers, approved by a person.
- Membership recomputed at execution time, so the approved list cannot drift.
- Soft state change only. No deletion, so every rollback is a state change rather than a restore.
- A durable, append-only audit entry with an allowlisted payload, so identifiers and codes cannot leak into the log.
Practices we kept
- Configuration is verified against production, not read. Keys, hostnames, and allowed origins are checked where they run. A value that is present is not a value that is correct.
- One seam per concern. A single session-creation path and a single audit interface, so a control cannot be half-applied across callers.
- Evidence is not entitlement. Reference data about organizations is labelled as evidence for a later review, never as a signup gate. Naming that boundary in code and documentation stops it from quietly becoming an access rule.
- Copy is part of correctness. A message that is right for sign-in ("if we have an account for this address…") is wrong for signup. Reused strings drift when flows change.
- Provenance from day one. Imported records carry where they came from and a status. Reconstructing that later is expensive, and guessing it lets imports overwrite manual corrections.
- Say plainly what remains external. Edge configuration, challenge keys, and monitoring credentials belong to the client's infrastructure. Naming them as prerequisites, in an operations runbook kept in the repository, turned the handover into a document rather than a conversation.
The position
Controls are claims until a test, a review, or production has watched them fail. When we take responsibility for a system someone else built, we look first for the controls that have never been exercised, and we assume they do not work until shown otherwise. It is why the codebase takeover checklist asks whether authorization is enforced on the server and whether anyone has seen each control reject a request, and why security is part of how we maintain software rather than a separate service.
Paul Bădărău