AW Dev Rethought

✂️ Perfection is achieved not when there is nothing more to add, but when there is nothing left to take away - Antoine de Saint-Exupéry

AWS in Production: The Hidden Risk of Over-Reliance on Managed Services


Introduction:

Managed services are one of the most compelling value propositions in cloud computing. Hand off the operational burden of running databases, message queues, search clusters, and caching layers to a provider that specialises in operating them at scale. Eliminate the toil of patching, scaling, and recovering from failures. Focus engineering effort on building product rather than operating infrastructure.

The case for managed services is genuinely strong. For most organisations, most of the time, using a managed database rather than operating one is the correct decision. The operational expertise required to run a production database reliably — backup management, replication configuration, failover testing, performance tuning — is significant, specialised, and not a source of competitive advantage for organisations whose core business is not database operations.

But managed services introduce dependencies and constraints that are not always visible when the adoption decision is made. Understanding what those risks are — not to avoid managed services but to adopt them with appropriate awareness — is what separates organisations that use managed services effectively from organisations that discover their limitations during incidents.


Vendor Lock-In Is Real and Accumulates Gradually:

The most discussed risk of managed service adoption is vendor lock-in — the progressive accumulation of dependencies on provider-specific features, APIs, and behaviours that make migration increasingly expensive over time. This risk is real, and it accumulates in ways that are not always visible at the point where individual adoption decisions are made.

A team that adopts Amazon DynamoDB makes a reasonable decision based on its operational simplicity and scalability characteristics. A team that builds its data access patterns around DynamoDB's single-table design conventions, uses DynamoDB Streams for change data capture, and integrates with DynamoDB Accelerator for caching has accumulated a dependency that would require significant re-architecture to migrate away from — not because any individual decision was wrong but because each decision built on the previous one.

Lock-in risk is not a reason to avoid managed services. It is a reason to make adoption decisions with awareness of the migration cost that would be incurred if the service needed to be replaced — due to cost changes, capability gaps, or a provider relationship that no longer serves the organisation's needs.


Operational Knowledge Atrophies Without Exercise:

When organisations adopt managed services for infrastructure that their teams previously operated, the operational knowledge required to run that infrastructure does not transfer to the managed service — it atrophies. Engineers who would previously have understood how to tune a PostgreSQL instance, configure replication, and recover from a corrupted index learn instead how to configure RDS parameters and interpret CloudWatch metrics.

This is generally a good trade. But it creates a specific vulnerability when the managed service fails in ways that require understanding the underlying infrastructure to diagnose. An RDS instance that is experiencing performance degradation due to a storage throughput issue requires understanding how EBS volumes behave under load to diagnose correctly. An ElastiCache cluster that is evicting keys unexpectedly requires understanding Redis memory management to investigate effectively.

The engineers who are most effective at diagnosing managed service failures are the ones who understand what the managed service is doing beneath its management layer — not because they need to operate it themselves but because that understanding enables them to interpret the signals the managed service exposes when something goes wrong.


Service Limits Create Unexpected Ceilings:

Every managed service has limits — on throughput, on storage, on the number of resources that can be created, on the size of individual items, and on the rate at which configuration changes can be applied. These limits are documented, but they are frequently not evaluated against actual usage patterns until the system is operating at a scale where the limits become constraints.

DynamoDB table throughput limits that were adequate at launch become bottlenecks as traffic grows. Lambda concurrency limits that were never reached during development become constraints during traffic spikes. SQS message size limits that seemed generous for initial use cases become problems when payload sizes grow as requirements evolve.

Service limits are not failures of the managed service — they are design boundaries that reflect architectural trade-offs the provider made when building the service. But discovering them in production, under load, during an incident is significantly more disruptive than discovering them during capacity planning. Understanding the limits of every managed service in the critical path and monitoring usage against those limits is an operational discipline that managed service adoption requires but does not automatically provide.


Outages Are Outside Your Control and Inside Your SLA:

When a managed service experiences an outage, the engineering team responsible for the systems that depend on it has no ability to accelerate recovery. They cannot restart the service, roll back a problematic update, or apply a fix. They can only wait for the provider to restore service and communicate to stakeholders why their systems are unavailable.

This loss of control is the most fundamental risk of managed service adoption, and it is a risk that organisations frequently underweight when making adoption decisions. The provider's SLA commits to a specific availability level — typically expressed as a monthly uptime percentage — but SLA credits that compensate for downtime do not compensate for the business impact of the outage itself.

Designing for managed service outages means treating every managed service as a potential failure point and building systems that degrade gracefully when that service is unavailable — through caching layers that serve stale data during database outages, through circuit breakers that prevent cascading failures when downstream services are unavailable, and through architecture that identifies which managed services are on the critical path and ensures that their unavailability produces degraded functionality rather than complete unavailability.


Cost Models Change in Ways That Are Hard to Predict:

Managed service pricing is based on usage — requests, storage, data transfer, and compute time. This model aligns cost with value at low usage levels and can produce unexpected cost growth as usage scales in ways that were not anticipated when the service was adopted.

DynamoDB pricing that seemed negligible for a small application can become significant when read and write volumes scale by two orders of magnitude. API Gateway request pricing that was invisible at low traffic levels can become a meaningful cost centre as an API scales to millions of requests per day. Data transfer costs that were never considered during architecture design can become the largest line item in a cloud bill as data volumes grow.

Provider pricing changes introduce additional uncertainty. A service that is priced attractively when adopted may become significantly more expensive if the provider adjusts its pricing model — and the migration cost that lock-in has created may make switching to a cheaper alternative more expensive than absorbing the price increase.


Conclusion:

Managed services are the right choice for most infrastructure decisions in most organisations. The operational leverage they provide — eliminating the toil of running infrastructure in exchange for a service fee — is real and valuable. The risks they introduce are equally real and require explicit management rather than passive acceptance.

Organisations that use managed services most effectively are those that make adoption decisions with awareness of lock-in costs, maintain enough operational knowledge to diagnose failures in what the managed service abstracts, monitor usage against service limits before those limits become constraints, design for managed service outages rather than assuming the provider's SLA makes outages acceptable, and evaluate cost trajectories at scale rather than only at current usage. The goal is not to avoid managed services — it is to depend on them deliberately rather than by default.


If this article helped you, you can support my work on AW Dev Rethought.


Rethought Relay:
Link copied!

Enjoyed this post?

Stay in the loop

New posts + weekly digest, straight to your inbox.

or

Create a free account

  • Save posts to your vault
  • Like posts & build history
  • New-post alerts

Comments

Add Your Comment

Comment Added!