AW Dev Rethought

🕵️ Debugging is like being the detective in a crime movie where you are also the murderer - Filipe Fortes

Systems Realities: Systems Fail Where Assumptions Are Wrong


Introduction:

Every system is built on assumptions. Assumptions about how users will behave, how dependencies will perform, how data will be structured, and how infrastructure will fail. These assumptions are not weaknesses — they are necessary. Building a system without assumptions is impossible. Every architectural decision, every API contract, every data model encodes assumptions about the world the system will operate in.

The problem is not that systems have assumptions. The problem is that assumptions are rarely documented, rarely challenged, and rarely revisited as the system and its environment evolve. An assumption that was valid when a system was designed may become invalid as traffic grows, as upstream systems change, as user behaviour shifts, or as the system is deployed in contexts its designers never anticipated.

When an assumption that a system depends on turns out to be wrong, the system fails — often in ways that are difficult to diagnose precisely because the assumption was never made explicit. Debugging a failure caused by a violated assumption requires first recognising that an assumption exists before you can investigate whether it holds.


Assumptions About Scale Are the Most Commonly Violated:

Systems are typically designed for the scale that exists when they are built, with headroom for anticipated growth. The assumptions embedded in that design — about database query performance, about network bandwidth, about memory requirements, about the number of concurrent users — are validated against current conditions and may not hold at a scale the designers did not anticipate.

A database schema that performs acceptably with one million records may perform unacceptably with one hundred million. A synchronous API that handles one hundred concurrent requests may time out under one thousand. A caching strategy that reduces database load at current traffic levels may be insufficient when traffic doubles. Each of these is a scale assumption — valid when made, violated as the system grows.

Scale assumption violations are particularly insidious because they occur gradually. Performance degrades incrementally rather than failing abruptly, making the connection between growth and degradation difficult to see until the degradation is severe enough to cause visible failures.


Dependency Assumptions Create Hidden Fragility:

Every system that depends on external services makes assumptions about those services — assumptions about their availability, their latency, their data quality, and their API stability. These assumptions are often implicit, encoded in timeout values, retry policies, and data parsing logic rather than in explicit documentation.

A service that assumes its downstream dependency responds within 500 milliseconds will fail when that dependency starts responding in two seconds — not because the service has a bug but because its assumption about dependency latency is no longer valid. A service that assumes its upstream data source produces well-formed records will fail when that source starts producing records with missing fields — not because the service cannot handle missing fields but because it never assumed it would need to.

Dependency assumptions are violated by changes that happen outside the system — upstream schema changes, downstream performance degradation, third-party API modifications, and infrastructure changes that affect latency. The system has not changed but its assumptions have been violated by the environment it operates in.


Concurrency Assumptions Are Frequently Wrong:

Systems designed and tested in single-threaded or low-concurrency environments make assumptions about execution order that do not hold under concurrent load. A database update that seems atomic when tested sequentially may produce a race condition when two requests attempt it simultaneously. A cache invalidation strategy that works correctly under low traffic may produce stale reads under high concurrency when invalidation and repopulation overlap in unexpected ways.

Concurrency assumption violations are among the hardest failures to reproduce and diagnose. They occur non-deterministically, depend on timing that varies between environments, and often disappear when debugging tools are attached because the observation changes the timing. A system that fails intermittently under production load but passes all tests in a development environment is frequently exhibiting a concurrency assumption violation that testing did not surface.

The assumptions that cause concurrency failures are often invisible during design — the possibility of two requests interacting was simply not considered. Making concurrency assumptions explicit — documenting which operations are expected to be atomic, which state is expected to be consistent, and which operations can safely be interleaved — is the first step toward identifying which assumptions are at risk.


Infrastructure Assumptions Encode Environmental Expectations:

Systems make assumptions about the infrastructure they run on — assumptions about available memory, disk speed, network bandwidth, and the reliability of infrastructure components. These assumptions are often encoded implicitly in configuration values, timeout settings, and resource allocation decisions that reflect the infrastructure that existed when the system was designed.

A system designed for dedicated hardware makes different assumptions than a system designed for virtualised cloud infrastructure. Dedicated hardware provides predictable performance. Virtualised infrastructure introduces variability — CPU steal time from noisy neighbours, network latency that varies with cloud provider load, and storage performance that fluctuates in ways that dedicated hardware does not.

Systems migrated from dedicated hardware to cloud infrastructure without revisiting their infrastructure assumptions frequently encounter failures that appear random but are actually the result of cloud-specific variability violating assumptions that were valid on dedicated hardware. The system has not changed. The infrastructure assumptions it encodes are no longer valid.


Documenting Assumptions Makes Them Challengeable:

Assumptions that are never documented cannot be challenged. An engineer who inherits a system without documented assumptions must either reverse-engineer them from the code and configuration or discover them by violating them in production. Both paths are more expensive than maintaining explicit assumption documentation.

Documenting assumptions does not mean creating exhaustive specification documents. It means recording the key assumptions that the system depends on — the expected scale it was designed for, the performance characteristics of its dependencies, the concurrency model it assumes, and the infrastructure it was designed to run on — in a place where engineers who work on the system can find and challenge them.

A system whose key assumptions are documented is a system whose failure modes are partially predictable. When an assumption is challenged — by growth, by a dependency change, or by an infrastructure migration — the engineering team can identify which parts of the system depend on that assumption and assess the risk of violation before it causes a production failure.


Conclusion:

Systems fail where assumptions are wrong because assumptions are the foundation on which every design decision rests. When the foundation shifts — because scale increases, dependencies change, concurrency increases, or infrastructure evolves — the decisions built on top of it fail in ways that are difficult to diagnose without understanding the assumptions that were violated.

The engineering practice that reduces assumption-related failures is not eliminating assumptions — that is impossible. It is making assumptions explicit, documenting them where they can be found and challenged, revisiting them as the system and its environment evolve, and designing systems that degrade gracefully when the assumptions they depend on turn out to be wrong. Systems that do this fail less often, recover faster, and are understood more deeply by the engineers who operate them.


If this article helped you, you can support my work on AW Dev Rethought.


Rethought Relay:
Link copied!

Enjoyed this post?

Stay in the loop

New posts + weekly digest, straight to your inbox.

or

Create a free account

  • Save posts to your vault
  • Like posts & build history
  • New-post alerts

Comments

Add Your Comment

Comment Added!