Www Independent coverage of news

Crawl Budget in Practice: Lessons From Real Deployments

By James Whitfield · · 1164 words
Crawl Budget in Practice: Lessons From Real Deployments

Cloud Infrastructure: A design that cannot be rolled back is a design that cannot be changed safely. Cloud Infrastructure: Latency budgets are easier to defend when every hop has a stated ceiling. Cloud Infrastructure: Caching helps only until the invalidation rules become the bottleneck.

Log Analysis: If the rollback plan needs a meeting, it is not a rollback plan. Log Analysis: Small pages that stay small are easier to keep fast than large ones made fast. Log Analysis: Write the invariant down; otherwise it lives only in someone's memory.

You can often replace a coordination problem with an idempotency key. That applies to content delivery as well. In practice, content delivery behaves differently: Anything that grows without a bound will eventually hit one. Documentation that is not tested tends to describe the previous version. The same reasoning holds for content delivery.

Schema Migration: A design that cannot be rolled back is a design that cannot be changed safely. Schema Migration: Latency budgets are easier to defend when every hop has a stated ceiling. Schema Migration: Caching helps only until the invalidation rules become the bottleneck.

For access control, the constraint matters more than the feature list. Periodic jobs should be safe to run twice, because they will be. Teams working on access control usually discover this the hard way. You rarely need a new component to fix a boundary problem. The signal you want is often already logged, just not aggregated. This is most visible in access control.

Schema Migration: Configurations should be reviewable in a diff, not only in a console. Schema Migration: The best time to add an index is before the table gets large. Schema Migration: Failures are usually correlated, so plan for the shared dependency.

Schema Migration: If the rollback plan needs a meeting, it is not a rollback plan. Schema Migration: Small pages that stay small are easier to keep fast than large ones made fast. Schema Migration: Write the invariant down; otherwise it lives only in someone's memory.

Data Pipelines: Configurations should be reviewable in a diff, not only in a console. Data Pipelines: The best time to add an index is before the table gets large. Data Pipelines: Failures are usually correlated, so plan for the shared dependency.

Edge Caching: The interesting number is not the average, it is the 99th percentile. Edge Caching: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Edge Caching: Every abstraction you add is a place where behaviour can differ from intent.

A design that cannot be rolled back is a design that cannot be changed safely. That applies to data pipelines as well. In practice, data pipelines behaves differently: Latency budgets are easier to defend when every hop has a stated ceiling. Caching helps only until the invalidation rules become the bottleneck. The same reasoning holds for data pipelines.

Serving static bytes is the cheapest thing you can do at the edge. That applies to data pipelines as well. In practice, data pipelines behaves differently: A schema is an interface; changing it is a migration, not an edit. Track the denominator as carefully as the numerator. The same reasoning holds for data pipelines.

Schema Markup: The first thing to settle is the failure mode, not the happy path. Schema Markup: Measurements taken once are anecdotes; you need a baseline that repeats. Schema Markup: Costs usually concentrate in a small number of operations, so find those first.

Cost Controls: The first thing to settle is the failure mode, not the happy path. Cost Controls: Measurements taken once are anecdotes; you need a baseline that repeats. Cost Controls: Costs usually concentrate in a small number of operations, so find those first.

Cost Controls: Configurations should be reviewable in a diff, not only in a console. Cost Controls: The best time to add an index is before the table gets large. Cost Controls: Failures are usually correlated, so plan for the shared dependency.

Content Delivery: Periodic jobs should be safe to run twice, because they will be. You rarely need a new component to fix a boundary problem. That applies to content delivery as well. In practice, content delivery behaves differently: The signal you want is often already logged, just not aggregated.

Teams working on search indexing usually discover this the hard way. A design that cannot be rolled back is a design that cannot be changed safely. Latency budgets are easier to defend when every hop has a stated ceiling. This is most visible in search indexing. Consider search indexing specifically. Caching helps only until the invalidation rules become the bottleneck.

In practice, rate limiting behaves differently: The first thing to settle is the failure mode, not the happy path. Measurements taken once are anecdotes; you need a baseline that repeats. The same reasoning holds for rate limiting. For rate limiting, the constraint matters more than the feature list. Costs usually concentrate in a small number of operations, so find those first.

Backup Strategy: You can often replace a coordination problem with an idempotency key. Backup Strategy: Anything that grows without a bound will eventually hit one. Backup Strategy: Documentation that is not tested tends to describe the previous version.

A queue smooths spikes but also hides how far behind you are. This is most visible in api design. Consider api design specifically. Retries without jitter turn a small outage into a large one. API Design: Separating the reads from the writes buys room to change either side.

Consider api design specifically. If the rollback plan needs a meeting, it is not a rollback plan. API Design: Small pages that stay small are easier to keep fast than large ones made fast. Write the invariant down; otherwise it lives only in someone's memory. That applies to api design as well.

Queue Design: If the rollback plan needs a meeting, it is not a rollback plan. Queue Design: Small pages that stay small are easier to keep fast than large ones made fast. Queue Design: Write the invariant down; otherwise it lives only in someone's memory.

In practice, cost controls behaves differently: A queue smooths spikes but also hides how far behind you are. Retries without jitter turn a small outage into a large one. The same reasoning holds for cost controls. For cost controls, the constraint matters more than the feature list. Separating the reads from the writes buys room to change either side.

Consider cost controls specifically. You can often replace a coordination problem with an idempotency key. Cost Controls: Anything that grows without a bound will eventually hit one. Documentation that is not tested tends to describe the previous version. That applies to cost controls as well.

Schema Markup: A design that cannot be rolled back is a design that cannot be changed safely. Schema Markup: Latency budgets are easier to defend when every hop has a stated ceiling. Schema Markup: Caching helps only until the invalidation rules become the bottleneck.

Related reading