↓ Ir para o conteúdo principal

← todas as notas

📎 Webclip

Building robust distributed systems

The article says robust distributed systems come from limiting connections between components, assuming every component can fail, and leaving slack in the system. It groups the approach into three parts: minimize dependencies, isolate errors, and add buffers.

Reading notes
#

  • Reduce connections by moving data or functionality into the calling component when possible.
  • Duplicate frequently used data locally, cache data that changes over time, and store rarely changing data directly in the component.
  • Denormalize data inside a component to avoid looking across multiple entities.
  • Package remote functionality as a library when it is critical and heavily used, even if that brings upgrade tradeoffs.
  • Use SLAs so each component declares latency, error-rate, and concurrency limits.
  • Let callers time out, retry idempotent operations, and open circuit breakers when failures continue.
  • Add random backoff to retries so callers do not retry all at once and overload a recovering component.
  • Use backpressure to drop new requests when a component is close to breaching its SLA.
  • Use asynchronous communication such as message buses so callers do not depend on a tight SLA.
  • Add hardware capacity as a buffer when load grows and cost allows it.