Www Independent coverage of news

API Design in Practice: Lessons From Real Deployments

By Michael Torres · · 1157 words
API Design in Practice: Lessons From Real Deployments

Consider observability specifically. The interesting number is not the average, it is the 99th percentile. Observability: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Every abstraction you add is a place where behaviour can differ from intent. That applies to observability as well.

Edge Caching: Periodic jobs should be safe to run twice, because they will be. Edge Caching: You rarely need a new component to fix a boundary problem. Edge Caching: The signal you want is often already logged, just not aggregated.

Cost Controls: If the rollback plan needs a meeting, it is not a rollback plan. Cost Controls: Small pages that stay small are easier to keep fast than large ones made fast. Cost Controls: Write the invariant down; otherwise it lives only in someone's memory.

Configurations should be reviewable in a diff, not only in a console. This is most visible in access control. Consider access control specifically. The best time to add an index is before the table gets large. Access Control: Failures are usually correlated, so plan for the shared dependency.

Teams working on rate limiting usually discover this the hard way. A design that cannot be rolled back is a design that cannot be changed safely. Latency budgets are easier to defend when every hop has a stated ceiling. This is most visible in rate limiting. Consider rate limiting specifically. Caching helps only until the invalidation rules become the bottleneck.

Queue Design: Periodic jobs should be safe to run twice, because they will be. You rarely need a new component to fix a boundary problem. That applies to queue design as well. In practice, queue design behaves differently: The signal you want is often already logged, just not aggregated.

Log Analysis: A queue smooths spikes but also hides how far behind you are. Log Analysis: Retries without jitter turn a small outage into a large one. Log Analysis: Separating the reads from the writes buys room to change either side.

The first thing to settle is the failure mode, not the happy path. This is most visible in storage tiers. Consider storage tiers specifically. Measurements taken once are anecdotes; you need a baseline that repeats. Storage Tiers: Costs usually concentrate in a small number of operations, so find those first.

Serving static bytes is the cheapest thing you can do at the edge. That applies to storage tiers as well. In practice, storage tiers behaves differently: A schema is an interface; changing it is a migration, not an edit. Track the denominator as carefully as the numerator. The same reasoning holds for storage tiers.

Monitoring Alerts: If the rollback plan needs a meeting, it is not a rollback plan. Monitoring Alerts: Small pages that stay small are easier to keep fast than large ones made fast. Monitoring Alerts: Write the invariant down; otherwise it lives only in someone's memory.

Crawl Budget: Configurations should be reviewable in a diff, not only in a console. Crawl Budget: The best time to add an index is before the table gets large. Crawl Budget: Failures are usually correlated, so plan for the shared dependency.

Observability: The interesting number is not the average, it is the 99th percentile. Observability: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Observability: Every abstraction you add is a place where behaviour can differ from intent.

You can often replace a coordination problem with an idempotency key. That applies to queue design as well. In practice, queue design behaves differently: Anything that grows without a bound will eventually hit one. Documentation that is not tested tends to describe the previous version. The same reasoning holds for queue design.

Queue Design: The interesting number is not the average, it is the 99th percentile. Queue Design: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Queue Design: Every abstraction you add is a place where behaviour can differ from intent.

Storage Tiers: Configurations should be reviewable in a diff, not only in a console. Storage Tiers: The best time to add an index is before the table gets large. Storage Tiers: Failures are usually correlated, so plan for the shared dependency.

Serving static bytes is the cheapest thing you can do at the edge. The same reasoning holds for content delivery. For content delivery, the constraint matters more than the feature list. A schema is an interface; changing it is a migration, not an edit. Teams working on content delivery usually discover this the hard way. Track the denominator as carefully as the numerator.

Queue Design: Periodic jobs should be safe to run twice, because they will be. Queue Design: You rarely need a new component to fix a boundary problem. Queue Design: The signal you want is often already logged, just not aggregated.

You can often replace a coordination problem with an idempotency key. That applies to content delivery as well. In practice, content delivery behaves differently: Anything that grows without a bound will eventually hit one. Documentation that is not tested tends to describe the previous version. The same reasoning holds for content delivery.

Periodic jobs should be safe to run twice, because they will be. This is most visible in storage tiers. Consider storage tiers specifically. You rarely need a new component to fix a boundary problem. Storage Tiers: The signal you want is often already logged, just not aggregated.

Cost Controls: The first thing to settle is the failure mode, not the happy path. Cost Controls: Measurements taken once are anecdotes; you need a baseline that repeats. Cost Controls: Costs usually concentrate in a small number of operations, so find those first.

Cost Controls: A queue smooths spikes but also hides how far behind you are. Cost Controls: Retries without jitter turn a small outage into a large one. Cost Controls: Separating the reads from the writes buys room to change either side.

In practice, load balancing behaves differently: If a metric has no owner, it will drift until it causes an incident. The cheapest optimisation is usually removing work nobody asked for. The same reasoning holds for load balancing. For load balancing, the constraint matters more than the feature list. Aggregating at write time trades flexibility for predictable read cost.

Schema Migration: The interesting number is not the average, it is the 99th percentile. Schema Migration: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Schema Migration: Every abstraction you add is a place where behaviour can differ from intent.

Access Control: A queue smooths spikes but also hides how far behind you are. Retries without jitter turn a small outage into a large one. That applies to access control as well. In practice, access control behaves differently: Separating the reads from the writes buys room to change either side.

Related reading