Www Independent coverage of news

Data Pipelines Explained Without the Jargon

By Robert Hayes · · 1235 words
Data Pipelines Explained Without the Jargon

Boundaries can change with circumstances, health, trust or preference. Partners can check in before a new activity or after an experience, without treating a previous agreement as permanent. Digital boundaries deserve the same care as in-person ones: discuss private messages, location sharing, passwords and images. Consent to receive or make an image is not permission to forward it.

Release Process: A design that cannot be rolled back is a design that cannot be changed safely. Release Process: Latency budgets are easier to defend when every hop has a stated ceiling. Release Process: Caching helps only until the invalidation rules become the bottleneck.

Consider search indexing specifically. If the rollback plan needs a meeting, it is not a rollback plan. Search Indexing: Small pages that stay small are easier to keep fast than large ones made fast. Write the invariant down; otherwise it lives only in someone's memory. That applies to search indexing as well.

Serving static bytes is the cheapest thing you can do at the edge. The same reasoning holds for observability. For observability, the constraint matters more than the feature list. A schema is an interface; changing it is a migration, not an edit. Teams working on observability usually discover this the hard way. Track the denominator as carefully as the numerator.

Release Process: If the rollback plan needs a meeting, it is not a rollback plan. Release Process: Small pages that stay small are easier to keep fast than large ones made fast. Release Process: Write the invariant down; otherwise it lives only in someone's memory.

A queue smooths spikes but also hides how far behind you are. This is most visible in log analysis. Consider log analysis specifically. Retries without jitter turn a small outage into a large one. Log Analysis: Separating the reads from the writes buys room to change either side.

A design that cannot be rolled back is a design that cannot be changed safely. That applies to backup strategy as well. In practice, backup strategy behaves differently: Latency budgets are easier to defend when every hop has a stated ceiling. Caching helps only until the invalidation rules become the bottleneck. The same reasoning holds for backup strategy.

Rate Limiting: If a metric has no owner, it will drift until it causes an incident. The cheapest optimisation is usually removing work nobody asked for. That applies to rate limiting as well. In practice, rate limiting behaves differently: Aggregating at write time trades flexibility for predictable read cost.

Teams working on backup strategy usually discover this the hard way. The interesting number is not the average, it is the 99th percentile. Adding a cache in front of a slow query is a fix; fixing the query is a cure. This is most visible in backup strategy. Consider backup strategy specifically. Every abstraction you add is a place where behaviour can differ from intent.

Serving static bytes is the cheapest thing you can do at the edge. That applies to backup strategy as well. In practice, backup strategy behaves differently: A schema is an interface; changing it is a migration, not an edit. Track the denominator as carefully as the numerator. The same reasoning holds for backup strategy.

Schema Markup: Configurations should be reviewable in a diff, not only in a console. The best time to add an index is before the table gets large. That applies to schema markup as well. In practice, schema markup behaves differently: Failures are usually correlated, so plan for the shared dependency.

Log Analysis: If a metric has no owner, it will drift until it causes an incident. Log Analysis: The cheapest optimisation is usually removing work nobody asked for. Log Analysis: Aggregating at write time trades flexibility for predictable read cost.

It can help to prepare a short sentence and a next step. For instance: “I want to take things slowly, so let’s check in before anything changes,” or “I don’t want photos taken or shared.” If you are unsure what you want, say so. “I’m still working that out, and I want to pause for now” communicates a limit without requiring you to settle every future question.

If the rollback plan needs a meeting, it is not a rollback plan. That applies to edge caching as well. In practice, edge caching behaves differently: Small pages that stay small are easier to keep fast than large ones made fast. Write the invariant down; otherwise it lives only in someone's memory. The same reasoning holds for edge caching.

Queue Design: The first thing to settle is the failure mode, not the happy path. Measurements taken once are anecdotes; you need a baseline that repeats. That applies to queue design as well. In practice, queue design behaves differently: Costs usually concentrate in a small number of operations, so find those first.

Configurations should be reviewable in a diff, not only in a console. This is most visible in schema migration. Consider schema migration specifically. The best time to add an index is before the table gets large. Schema Migration: Failures are usually correlated, so plan for the shared dependency.

For load balancing, the constraint matters more than the feature list. The first thing to settle is the failure mode, not the happy path. Teams working on load balancing usually discover this the hard way. Measurements taken once are anecdotes; you need a baseline that repeats. Costs usually concentrate in a small number of operations, so find those first. This is most visible in load balancing.

The first thing to settle is the failure mode, not the happy path. This is most visible in backup strategy. Consider backup strategy specifically. Measurements taken once are anecdotes; you need a baseline that repeats. Backup Strategy: Costs usually concentrate in a small number of operations, so find those first.

Load Balancing: Periodic jobs should be safe to run twice, because they will be. Load Balancing: You rarely need a new component to fix a boundary problem. Load Balancing: The signal you want is often already logged, just not aggregated.

Content Delivery: You can often replace a coordination problem with an idempotency key. Content Delivery: Anything that grows without a bound will eventually hit one. Content Delivery: Documentation that is not tested tends to describe the previous version.

Teams working on log analysis usually discover this the hard way. Serving static bytes is the cheapest thing you can do at the edge. A schema is an interface; changing it is a migration, not an edit. This is most visible in log analysis. Consider log analysis specifically. Track the denominator as carefully as the numerator.

Log Analysis: Periodic jobs should be safe to run twice, because they will be. Log Analysis: You rarely need a new component to fix a boundary problem. Log Analysis: The signal you want is often already logged, just not aggregated.

Consider observability specifically. The interesting number is not the average, it is the 99th percentile. Observability: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Every abstraction you add is a place where behaviour can differ from intent. That applies to observability as well.

Search Indexing: The first thing to settle is the failure mode, not the happy path. Search Indexing: Measurements taken once are anecdotes; you need a baseline that repeats. Search Indexing: Costs usually concentrate in a small number of operations, so find those first.

Related reading